VLDB 2026 Research / reviewers in the wild / expert
Zilong Wang 0002
dblp:42/898-2
· DBLP profile ↗
15ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-1614-0943ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's NestabstractMassive high-quality data, both pre-training raw texts and post-training annotations, have been carefully prepared to incubate advanced large language models (LLMs).In contrast, for information extraction (IE), pre-training data, such as BIO-tagged sequences, are hard to scale up.We show that IE models can act as free riders on LLM resources by reframing next-token prediction into extraction for tokens already present in the context.Specifically, our proposed next tokens extraction (NTE) paradigm learns a versatile IE model, Cuckoo 1 , with 102.6M extractive data converted from LLM's pre-training and post-training data.Under the few-shot setting, Cuckoo adapts effectively to traditional and complex instruction-following IE with better performance than existing pretrained IE models.As a free rider, Cuckoo can naturally evolve with the ongoing advancements in LLM data preparation, benefiting from improvements in LLM training pipelines without additional manual effort.2 Letian Peng, Zilong Wang 0002, Jingbo Shang |
ACL (1) | 2 |
| 2025 | Speculative RAG: Enhancing Retrieval Augmented Generation through DraftingabstractRetrieval augmented generation (RAG) combines the generative abilities of large language models (LLMs) with external knowledge sources to provide more accurate and up-to-date responses. Recent RAG advancements focus on improving retrieval outcomes through iterative LLM refinement or self-critique capabilities acquired through additional instruction tuning of LLMs. In this work, we introduce Speculative RAG - a framework that leverages a larger generalist LM to efficiently verify multiple RAG drafts produced in parallel by a smaller, distilled specialist LM. Each draft is generated from a distinct subset of retrieved documents, offering diverse perspectives on the evidence while reducing input token counts per draft. This approach enhances comprehension of each subset and mitigates potential position bias over long context. Our method accelerates RAG by delegating drafting to the smaller specialist LM, with the larger generalist LM performing a single verification pass over the drafts. Extensive experiments demonstrate that Speculative RAG achieves state-of-the-art performance with reduced latency on TriviaQA, MuSiQue, PopQA, PubHealth, and ARC-Challenge benchmarks. It notably enhances accuracy by up to 12.97% while reducing latency by 50.83% compared to conventional RAG systems on PubHealth. Zilong Wang 0002, Zifeng Wang 0002, Long T. Le, Huaixiu Steven Zheng, Swaroop Mishra, Vincent Perot, Yuwei Zhang 0001, Anush Mattapalli, Ankur Taly, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 1 |
| 2025 | Training Language Models to Generate Quality Code with Program Analysis FeedbackabstractCode generation with large language models (LLMs), often termed vibe coding, is increasingly adopted in production but fails to ensure code quality, particularly in security (e.g., SQL injection vulnerabilities) and maintainability (e.g., missing type annotations). Existing methods, such as supervised fine-tuning and rule-based post-processing, rely on labor-intensive annotations or brittle heuristics, limiting their scalability and effectiveness. We propose REAL (Reinforcement rEwards from Automated anaLysis), a reinforcement learning framework that trains LLMs to generate production-quality code using program analysis–guided feedback. Specifically, REAL integrates two automated signals: (1) static analyzers detecting security and maintainability defects and (2) unit tests ensuring functional correctness. Unlike prior work, our framework is prompt-agnostic and reference-free, enabling scalable supervision without manual intervention. Experiments across multiple datasets and model scales demonstrate that REAL outperforms state-of-the-art methods in simultaneous assessments of functionality and code quality. Our work bridges the gap between rapid prototyping and production-ready code, enabling LLMs to deliver both speed and quality. Zilong Wang 0002, Junxia Cui, Xiaohan Fu, Haohui Mai, Viswanathan Krishnan, Jianfeng Gao 0001, Jingbo Shang |
NeurIPS | 2 |
| 2024 | Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code GenerationabstractRecently, large language models (LLMs) have shown an extraordinary ability to understand natural language and generate programming code. It has been a common practice for software engineers to consult LLMs when encountering coding questions. Although efforts have been made to avoid syntax errors and align the code with the intended semantics, the reliability, and robustness of the code generation from LLMs have not yet been thoroughly studied. The executable code is not equivalent to reliable and robust code, especially in the context of real-world software development. For example, the misuse of APIs in the generated code could lead to severe problems, such as resource leaks, program crashes, etc. Existing code evaluation benchmarks and datasets focus on crafting small tasks such as programming questions in coding interviews, which, however, deviates from the problem that developers would ask LLM for real-world coding help. To fill the missing piece, in this work, we propose a dataset RobustAPI for evaluating the reliability and robustness of code generated by LLMs. We collect 1208 coding questions from Stack Overflow on 18 representative Java APIs. We summarize the common misuse patterns of these APIs and evaluate them on current popular LLMs. The evaluation results show that even for GPT-4, 62% of the generated code contains API misuses, which would cause unexpected consequences if the code is introduced into real-world software. Zilong Wang 0002 |
AAAI | 2 |
| 2024 | Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionabstractLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Letian Peng, Yuwei Zhang 0001, Zilong Wang 0002, Jayanth Srinivasa, Gaowen Liu, Zihan Wang 0001, Jingbo Shang |
ACL (1) | 3 |
| 2024 | Towards Few-shot Entity Recognition in Document Images: A Graph Neural Network Approach Robust to Image ManipulationabstractRecent advances of incorporating layout information, typically bounding box coordinates, into pre-trained language models have achieved significant performance in entity recognition from document images. Using coordinates can easily model the position of each token, but they are sensitive to manipulations in document images (e.g., shifting, rotation or scaling) which are common in real scenarios. Such limitation becomes even worse when the training data is limited in few-shot settings. In this paper, we propose a novel framework, LAGER, which leverages the topological adjacency relationship among the tokens through learning their relative layout information with graph neural networks. Specifically, we consider the tokens in the documents as nodes and formulate the edges based on the topological heuristics. Such adjacency graphs are invariant to affine transformations, making it robust to the common image manipulations. We incorporate these graphs into the pre-trained language model by adding graph neural network layers on top of the language model embeddings. Extensive experiments on two benchmark datasets show that LAGER significantly outperforms strong baselines under different few-shot settings and also demonstrate better robustness to manipulations. Prashant Krishnan, Zilong Wang 0002, Yangkun Wang, Jingbo Shang |
LREC/COLING | 2 |
| 2024 | Incubating Text Classifiers Following User Instruction with Nothing but LLMabstractIn this paper, we aim to generate text classification data given arbitrary class definitions (i.e., user instruction), so one can train a text classifier without any human annotation or raw corpus.Recent advances in large language models (LLMs) lead to pioneer attempts to individually generate texts for each class via prompting.In this paper, we propose Incubator, the first framework that can handle complicated and even mutually dependent classes (e.g., "TED Talk given by Educator" and "Other").Specifically, our Incubator is a fine-tuned LLM that takes the instruction of all class definitions as input, and in each inference, it can jointly generate one sample for every class.First, we tune Incubator on the instruction-to-data mappings that we obtained from classification datasets and descriptions on Hugging Face together with in-context augmentation by GPT-4.To emphasize the uniformity and diversity in generations, we refine Incubator by fine-tuning with the cluster centers of semantic textual embeddings of the generated samples.We compare Incubator on various classification tasks with strong baselines such as direct LLM-based inference and training data generation by prompt engineering.Experiments show Incubator is able to (1) outperform previous methods on traditional benchmarks, (2) take label interdependency and user preference into consideration, and (3) enable logical text mining by incubating multiple classifiers. Letian Peng, Zilong Wang 0002, Jingbo Shang |
EMNLP | 2 |
| 2024 | Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingabstractTable-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both free-form questions and semi-structured tabular data. Chain-of-Thought and its similar approaches incorporate the reasoning chain in the form of textual context, but it is still an open question how to effectively leverage tabular data in the reasoning chain. We propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts. Specifically, we guide LLMs using in-context learning to iteratively generate operations and update the table to represent a tabular reasoning chain. LLMs can therefore dynamically plan the next operation based on the results of the previous ones. This continuous evolution of the table forms a chain, showing the reasoning process for a given tabular problem. The chain carries structured information of the intermediate results, enabling more accurate and reliable predictions. Chain-of-Table achieves new state-of-the-art performance on WikiTQ, FeTaQA, and TabFact benchmarks across multiple LLM choices. Zilong Wang 0002, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang 0002, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 1 |
| 2024 | TableRAG: Million-Token Table Understanding with Language ModelsabstractRecent advancements in language models (LMs) have notably enhanced their ability to reason with tabular data, primarily through program-aided mechanisms that manipulate and analyze tables.
However, these methods often require the entire table as input, leading to scalability challenges due to the positional bias or context length constraints.
In response to these challenges, we introduce TableRAG, a Retrieval-Augmented Generation (RAG) framework specifically designed for LM-based table understanding.
TableRAG leverages query expansion combined with schema and cell retrieval to pinpoint crucial information before providing it to the LMs.
This enables more efficient data encoding and precise retrieval, significantly reducing prompt lengths and mitigating information loss.
We have developed two new million-token benchmarks from the Arcade and BIRD-SQL datasets to thoroughly evaluate TableRAG's effectiveness at scale.
Our results demonstrate that TableRAG's retrieval design achieves the highest retrieval quality, leading to the new state-of-the-art performance on large-scale table understanding. Si-An Chen, Lesly Miculicich, Julian Martin Eisenschlos, Zifeng Wang 0002, Zilong Wang 0002, Yanfei Chen, Yasuhisa Fujii, Hsuan-Tien Lin, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 5 |
| 2023 | VRDU: A Benchmark for Visually-rich Document UnderstandingabstractUnderstanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achieved impressive results, we find that existing benchmarks do not reflect the complexity of real documents seen in industry. In this work, we identify the desiderata for a more comprehensive benchmark and propose one we call Visually Rich Document Understanding (VRDU). VRDU contains two datasets that represent several challenges: rich schema including diverse data types as well as hierarchical entities, complex templates including tables and multi-column layouts, and diversity of different layouts (templates) within a single document type. We design few-shot and conventional experiment settings along with a carefully designed matching algorithm to evaluate extraction results. We report the performance of strong baselines and offer three observations: (1) generalizing to new document templates is still very challenging, (2) few-shot performance has a lot of headroom, and (3) models struggle with hierarchical fields such as line-items in an invoice. We plan to open source the benchmark and the evaluation toolkit. We hope this helps the community make progress on these challenging tasks in extracting structured data from visually rich documents. Zilong Wang 0002, Yichao Zhou 0001, Wei Wei 0019, Chen-Yu Lee, Sandeep Tata |
KDD | 1 |
| 2022 | MGDoc: Pre-training with Multi-granular Hierarchy for Document Image UnderstandingabstractZilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, Vlad Morariu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Zilong Wang 0002, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005, Jingbo Shang, Vlad I. Morariu |
EMNLP | 1 |
| 2021 | LayoutReader: Pre-training of Text and Layout for Reading Order DetectionabstractReading order detection is the cornerstone to understanding visually-rich documents (e.g., receipts and forms).Unfortunately, no existing work took advantage of advanced deep learning models because it is too laborious to annotate a large enough dataset.We observe that the reading order of WORD documents is embedded in their XML metadata; meanwhile, it is easy to convert WORD documents to PDFs or images.Therefore, in an automated manner, we construct ReadingBank, a benchmark dataset that contains reading order, text, and layout information for 500,000 document images covering a wide spectrum of document types.This first-ever large-scale dataset unleashes the power of deep neural networks for reading order detection.Specifically, our proposed LayoutReader captures the text and layout information for reading order prediction using the seq2seq model.It performs almost perfectly in reading order detection and significantly improves both open-source and commercial OCR engines in ordering text lines in their results in our experiments.The dataset and models are publicly available at https: //aka.ms/layoutreader. Zilong Wang 0002, Yiheng Xu, Lei Cui 0001, Jingbo Shang, Furu Wei |
EMNLP (1) | 1 |
| 2020 | Exploring Semantic Capacity of TermsabstractWe introduce and study semantic capacity of terms.For example, the semantic capacity of artificial intelligence is higher than that of linear regression since artificial intelligence possesses a broader meaning scope.Understanding semantic capacity of terms will help many downstream tasks in natural language processing.For this purpose, we propose a two-step model to investigate semantic capacity of terms, which takes a large text corpus as input and can evaluate semantic capacity of terms if the text corpus can provide enough cooccurrence information of terms.Extensive experiments in three fields demonstrate the effectiveness and rationality of our model compared with well-designed baselines and human-level evaluations. Jie Huang 0009, Zilong Wang 0002, Kevin Chen-Chuan Chang, Wen-Mei W. Hwu, Jinjun Xiong |
EMNLP (1) | 2 |
| 2020 | TransModality: An End2End Fusion Method with Transformer for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis is an important research area that predicts speaker’s sentiment tendency through features extracted from textual, visual and acoustic modalities. The central challenge is the fusion method of the multimodal information. A variety of fusion methods have been proposed, but few of them adopt end-to-end translation models to mine the subtle correlation between modalities. Enlightened by recent success of Transformer in the area of machine translation, we propose a new fusion method, TransModality, to address the task of multimodal sentiment analysis. We assume that translation between modalities contributes to a better joint representation of speaker’s utterance. With Transformer, the learned features embody the information both from the source modality and the target modality. We validate our model on multiple multimodal datasets: CMU-MOSI, MELD, IEMOCAP. The experiments show that our proposed method achieves the state-of-the-art performance. Zilong Wang 0002, Zhaohong Wan, Xiaojun Wan 0001 |
WWW | 1 |
| 2019 | BAB-QA: A New Neural Model for Emotion Detection in Multi-party Dialogue
Zilong Wang 0002, Zhaohong Wan, Xiaojun Wan 0001 |
PAKDD (1) | 1 |