VLDB 2026 Research / reviewers in the wild / expert
Zeqiu Wu
dblp:188/5861
· DBLP profile ↗
12ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Language models and text generation · 34% Question answering and dialogue systems · 28% Reinforcement learning · 25% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 42% Knowledge graphs · 35% Query processing and optimization · 22% |
Topics — the 29 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.4 | 2 | 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 Fine-Grained Human Feedback Gives Better Rewards for Language Model Training · NeurIPS 2023 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering › knowledge-grounded question answering
retrieval-augmented question answering |
1.0 | 2 | 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection · ICLR 2024 Training Language Models to Generate Text with Citations via Fine-grained Rewards · ACL (1) 2024 |
Natural language and speech › Language models and text generation › text generation › scientific text generation
citation generation |
0.8 | 1 | 2024 | Training Language Models to Generate Text with Citations via Fine-grained Rewards · ACL (1) 2024 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.8 | 1 | 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality |
0.8 | 1 | 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection · ICLR 2024 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.8 | 1 | 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection · ICLR 2024 |
Machine learning › Reinforcement learning
preference learning |
0.8 | 1 | 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.8 | 1 | 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection · ICLR 2024 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.8 | 1 | 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback · NeurIPS 2024 |
Natural language and speech › Language models and text generation
self-reflection |
0.8 | 1 | 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection · ICLR 2024 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.7 | 3 | 2018 | Indirect Supervision for Relation Extraction using Question-Answer Pairs · WSDM 2018 CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017 HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree Expansion · KDD 2018 |
Natural language and speech › Language models and text generation
alignment |
0.7 | 1 | 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model Training · NeurIPS 2023 |
Natural language and speech › Question answering and dialogue systems
long-form question answering |
0.7 | 1 | 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model Training · NeurIPS 2023 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback › learning from human feedback
RLHF |
0.7 | 1 | 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model Training · NeurIPS 2023 |
Natural language and speech › Question answering and dialogue systems › interactive question answering
conversational question answering |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Query processing and optimization
query rewriting |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Information retrieval › machine learning for information retrieval
reinforcement learning for retrieval |
0.6 | 1 | 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning · EMNLP 2022 |
Natural language and speech › Language models and text generation
controllable text generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems
knowledge-grounded dialogue |
0.5 | 1 | 2021 | DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document Contextualization · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
knowledge-grounded response generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Information retrieval › document retrieval
passage retrieval |
0.5 | 1 | 2021 | DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document Contextualization · EMNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.3 | 1 | 2018 | Indirect Supervision for Relation Extraction using Question-Answer Pairs · WSDM 2018 |
Machine learning › Learning paradigms › weakly supervised learning
indirect supervision |
0.3 | 1 | 2018 | Indirect Supervision for Relation Extraction using Question-Answer Pairs · WSDM 2018 |
Knowledge graphs
taxonomy construction |
0.3 | 1 | 2018 | HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree Expansion · KDD 2018 |
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction |
0.3 | 1 | 2017 | CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017 |
Knowledge graphs
distant supervision |
0.3 | 1 | 2017 | CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017 |
Knowledge graphs
knowledge graph construction |
0.3 | 1 | 2017 | CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › grounding
knowledge grounding |
0.1 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › relation extraction
weakly supervised relation extraction |
0.1 | 1 | 2018 | HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree Expansion · KDD 2018 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning from human feedback · 1.4reward function · 1.1reinforcement learning · 1.1joint embedding · 0.9self-reflection tokens · 0.8retrieval-augmented generation · 0.8fine-grained rewards · 0.8PPO · 0.8DPO · 0.8reward model · 0.7document structure encoding · 0.5auxiliary loss · 0.5weakly-supervised relation extraction · 0.3hierarchical tree expansion · 0.3partial-label loss · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Training Language Models to Generate Text with Citations via Fine-grained RewardsabstractWhile recent Large Language Models (LLMs) have proven useful in answering user queries, they are prone to hallucination, and their responses often lack credibility due to missing references to reliable sources.An intuitive solution to these issues would be to include in-text citations referring to external documents as evidence.While previous works have directly prompted LLMs to generate in-text citations, their performances are far from satisfactory, especially when it comes to smaller LLMs.In this work, we propose an effective training framework using fine-grained rewards to teach LLMs to generate highly supportive and relevant citations, while ensuring the correctness of their responses.We also conduct a systematic analysis of applying these fine-grained rewards to common LLM training strategies, demonstrating its advantage over conventional practices.We conduct extensive experiments on Question Answering (QA) datasets taken from the ALCE benchmark and validate the model's generalizability using EXPERTQA.On LLaMA-2-7B, the incorporation of fine-grained rewards achieves the best performance among the baselines, even surpassing that of GPT-3.5-turbo.1 Chengyu Huang 0003, Zeqiu Wu, Yushi Hu, Wenya Wang 0001 |
ACL (1) | 2 |
| 2024 | Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionabstractDespite their remarkable capabilities, large language models (LLMs) often produce responses containing factual inaccuracies due to their sole reliance on the parametric knowledge they encapsulate. Retrieval-Augmented Generation (RAG), an ad hoc approach that augments LMs with retrieval of relevant knowledge, decreases such issues. However, indiscriminately retrieving and incorporating a fixed number of retrieved passages, regardless of whether retrieval is necessary, or passages are relevant, diminishes LM versatility or can lead to unhelpful response generation. We introduce a new framework called **Self-Reflective Retrieval-Augmented Generation (Self-RAG)** that enhances an LM's quality and factuality through retrieval and self-reflection.
Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its generations using special tokens, called {\it reflection} tokens. Generating reflection tokens makes the LM controllable during the inference phase, enabling it to tailor its behavior to diverse task requirements.
Experiments show that Self-RAG (7B and 13B parameters) significantly outperforms state-of-the-art LLMs and retrieval-augmented models on a diverse set of tasks.
Specifically, Self-RAG outperforms ChatGPT and retrieval-augmented Llama2-chat on Open-domain QA, reasoning, and fact verification tasks, and it shows significant gains in improving factuality and citation accuracy for long-form generations relative to these models. Our code and trained models are available at https://selfrag.github.io/ Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, Hannaneh Hajishirzi |
ICLR | 2 |
| 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackabstractLearning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs). Despite its widespread use, the way preference-based learning is applied varies wildly, with differing data, learning algorithms, and evaluations used, making disentangling the impact of each aspect difficult. In this work, we identify four core aspects of preference-based learning: preference data, learning algorithm, reward model, and policy training prompts, systematically investigate the impact of these components on downstream model performance, and suggest a recipe for strong learning for preference feedback. Our findings indicate that all aspects are important for performance, with better preference data leading to the largest improvements, followed by the choice of learning algorithm, the use of improved reward models, and finally the use of additional unlabeled prompts for policy training. Notably, PPO outperforms DPO by up to 2.5% in math and 1.2% in general domains. High-quality preference data leads to improvements of up to 8% in instruction following and truthfulness. Despite significant gains of up to 5% in mathematical evaluation when scaling up reward models, we surprisingly observe marginal improvements in other categories. Hamish Ivison, Yizhong Wang, Jiacheng Liu 0010, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert 0001, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 4 |
| 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingabstractLanguage models (LMs) often exhibit undesirable text generation behaviors, including generating false, toxic, or irrelevant outputs.
Reinforcement learning from human feedback (RLHF)---where human preference judgments on LM outputs are transformed into a learning signal---has recently shown promise in addressing these issues. However, such holistic feedback conveys limited information on long text outputs; it does not indicate which aspects of the outputs influenced user preference; e.g., which parts contain what type(s) of errors. In this paper, we use fine-grained human feedback (e.g., which sentence is false, which sub-sentence is irrelevant) as an explicit training signal. We introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.g., a sentence) is generated; and (2) incorporating multiple reward models associated with different feedback types (e.g., factual incorrectness, irrelevance, and information incompleteness). We conduct experiments on detoxification and long-form question answering to illustrate how learning with this reward function leads to improved performance, supported by both automatic and human evaluation. Additionally, we show that LM behaviors can be customized using different combinations of fine-grained reward models. We release all data, collected human feedback, and codes at https://FineGrainedRLHF.github.io. Zeqiu Wu, Yushi Hu, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, Hannaneh Hajishirzi |
NeurIPS | 1 |
| 2023 | InSCIt: Information-Seeking Conversations with Mixed-Initiative InteractionsabstractAbstract In an information-seeking conversation, a user may ask questions that are under-specified or unanswerable. An ideal agent would interact by initiating different response types according to the available knowledge sources. However, most current studies either fail to or artificially incorporate such agent-side initiative. This work presents InSCIt, a dataset for Information-Seeking Conversations with mixed-initiative Interactions. It contains 4.7K user-agent turns from 805 human-human conversations where the agent searches over Wikipedia and either directly answers, asks for clarification, or provides relevant information to address user queries. The data supports two subtasks, evidence passage identification and response generation, as well as a human evaluation protocol to assess model performance. We report results of two systems based on state-of-the-art models of conversational knowledge identification and open-domain question answering. Both systems significantly underperform humans, suggesting ample room for improvement in future studies.1 Zeqiu Wu, Ryu Parish, Hao Cheng 0002, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, Hannaneh Hajishirzi |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningabstractCompared to standard retrieval tasks, passage retrieval for conversational question answering (CQA) poses new challenges in understanding the current user question, as each question needs to be interpreted within the dialogue context.Moreover, it can be expensive to retrain well-established retrievers such as search engines that are originally developed for nonconversational queries.To facilitate their use, we develop a query rewriting model CONQRR that rewrites a conversational question in the context into a standalone question.It is trained with a novel reward function to directly optimize towards retrieval using reinforcement learning and can be adapted to any off-theshelf retriever.CONQRR achieves state-ofthe-art results on a recent open-domain CQA dataset containing conversations from three different sources, and is effective for two different off-the-shelf retrievers.Our extensive analysis also shows the robustness of CON-QRR to out-of-domain dialogues as well as to zero query rewriting supervision. Zeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter, Hannaneh Hajishirzi, Mari Ostendorf, Gaurav Tomar |
EMNLP | 1 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 1 |
| 2021 | DIALKI: Knowledge Identification in Conversational Systems through Dialogue-Document ContextualizationabstractIdentifying relevant knowledge to be used in conversational systems that are grounded in long documents is critical to effective response generation.We introduce a knowledge identification model that leverages the document structure to provide dialogue-contextualized passage encodings and better locate knowledge relevant to the conversation.An auxiliary loss captures the history of dialogue-document connections.We demonstrate the effectiveness of our model on two document-grounded conversational datasets and provide analyses showing generalization to unseen documents and long dialogue contexts. Zeqiu Wu, Bo-Ru Lu, Hannaneh Hajishirzi, Mari Ostendorf |
EMNLP (1) | 1 |
| 2018 | HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree ExpansionabstractTaxonomies are of great value to many knowledge-rich applications. As the manual taxonomy curation costs enormous human effects, automatic taxonomy construction is in great demand. However, most existing automatic taxonomy construction methods can only build hypernymy taxonomies wherein each edge is limited to expressing the is-a relation. Such a restriction limits their applicability to more diverse real-world tasks where the parent-child may carry different relations. In this paper, we aim to construct a task-guided taxonomy from a domain-specific corpus, and allow users to input a seed taxonomy, serving as the task guidance. We propose an expansion-based taxonomy construction framework, namely HiExpan, which automatically generates key term list from the corpus and iteratively grows the seed taxonomy. Specifically, HiExpan views all children under each taxonomy node forming a coherent set and builds the taxonomy by recursively expanding all these sets. Furthermore, HiExpan incorporates a weakly-supervised relation extraction module to extract the initial children of a newly-expanded node and adjusts the taxonomy tree by optimizing its global structure. Our experiments on three real datasets from different domains demonstrate the effectiveness of HiExpan for building task-guided taxonomies. Zeqiu Wu, Dongming Lei, Chao Zhang 0014, Xiang Ren 0001, Michelle Vanni, Brian M. Sadler, Jiawei Han 0001 |
KDD | 2 |
| 2018 | Indirect Supervision for Relation Extraction using Question-Answer PairsabstractAutomatic relation extraction (E)for types of interest is of great importance for interpreting massive text corpora in an efficient manner. For example, we want to identify the relationship "president_of" between entities "Donald Trump" and "United States" in a sentence expressing such a relation. Traditional RE models have heavily relied on human-annotated corpus for training, which can be costly in generating labeled data and become obstacles when dealing with more relation types. Thus, more RE extraction systems have shifted to be built upon training data automatically acquired by linking to knowledge bases (distant supervision). However, due to the incompleteness of knowledge bases and the context-agnostic labeling, the training data collected via distant supervision (DS) can be very noisy. In recent years, as increasing attention has been brought to tackling question-answering (QA) tasks, user feedback or datasets of such tasks become more accessible. In this paper, we propose a novel framework, ReQuest, to leverage question-answer pairs as an indirect source of supervision for relation extraction, and study how to use such supervision to reduce noise induced from DS. Our model jointly embeds relation mentions, types, QA entity mention pairs and text features in two low-dimensional spaces (RE and QA), where objects with same relation types or semantically similar question-answer pairs have similar representations. Shared features connect these two spaces, carrying clearer semantic knowledge from both sources. ReQuest, then use these learned embeddings to estimate the types of test relation mentions. We formulate a global objective function and adopt a novel margin-based QA loss to reduce noise in DS by exploiting semantic evidence from the QA dataset. Our experimental results achieve an average of 11% improvement in F1 score on two public RE datasets combined with TREC QA dataset. Codes and datasets can be downloaded at https://github.com/ellenmellon/ReQuest. Zeqiu Wu, Xiang Ren 0001, Frank F. Xu, Jiawei Han 0001 |
WSDM | 1 |
| 2017 | SetExpan: Corpus-Based Set Expansion via Context Feature Selection and Rank Ensemble
Zeqiu Wu, Dongming Lei, Jingbo Shang, Xiang Ren 0001, Jiawei Han 0001 |
ECML/PKDD (1) | 2 |
| 2017 | CoType: Joint Extraction of Typed Entities and Relations with Knowledge BasesabstractExtracting entities and relations for types of interest from text is important for understanding massive text corpora. Traditionally, systems of entity relation extraction have relied on human-annotated corpora for training and adopted an incremental pipeline. Such systems require additional human expertise to be ported to a new domain, and are vulnerable to errors cascading down the pipeline. In this paper, we investigate joint extraction of typed entities and relations with labeled data heuristically obtained from knowledge bases (i.e., distant supervision). As our algorithm for type labeling via distant supervision is context-agnostic, noisy training data poses unique challenges for the task. We propose a novel domain-independent framework, called CoType, that runs a data-driven text segmentation algorithm to extract entity mentions, and jointly embeds entity mentions, relation mentions, text features and type labels into two low-dimensional spaces (for entity and relation mentions respectively), where, in each space, objects whose types are close will also have similar representations. CoType, then using these learned embeddings, estimates the types of test (unlinkable) mentions. We formulate a joint optimization problem to learn embeddings from text corpora and knowledge bases, adopting a novel partial-label loss function for noisy labeled data and introducing an object "translation" function to capture the cross-constraints of entities and relations on each other. Experiments on three public datasets demonstrate the effectiveness of CoType across different domains (e.g., news, biomedical), with an average of 25% improvement in F1 score compared to the next best method. Xiang Ren 0001, Zeqiu Wu, Wenqi He, Meng Qu, Clare R. Voss, Heng Ji 0001, Tarek F. Abdelzaher, Jiawei Han 0001 |
WWW | 2 |