VLDB 2026 Research / reviewers in the wild / expert
Xinya Du
dblp:200/8114
· DBLP profile ↗
27ranked-venue papers
10as first author
22since 2021 · last 2026
0000-0003-4255-8013ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from ExperienceabstractZhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001 |
ACL (1) | 6 |
| 2026 | Dr.V : A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-Grained Spatial-Temporal Grounding
Meng Luo 0010, Shengqiong Wu, Liqiang Jing, Tianjie Ju, Jinxiang Lai, Tianlong Wu, Xinya Du, Siyuan Yan, Jiebo Luo 0001, William Yang Wang, Hao Fei 0001, Mong-Li Lee, Wynne Hsu |
Int. J. Comput. Vis. | 8 |
| 2025 | SciEvent: Benchmarking Multi-domain Scientific Event ExtractionabstractScientific information extraction (SciIE) has primarily relied on entity-relation extraction in narrow domains, limiting its applicability to interdisciplinary research and struggling to capture the necessary context of scientific information, often resulting in fragmented or conflicting statements.In this paper, we introduce SciEvent 1 , a novel multi-domain benchmark of scientific abstracts annotated via a unified event extraction (EE) schema designed to enable structured and context-aware understanding of scientific content.It includes 500 abstracts across five research domains, with manual annotations of event segments, triggers, and fine-grained arguments.We define SciIE as a multi-stage EE pipeline: (1) segmenting abstracts into core scientific activities-Background, Method, Result, and Conclusion; and (2) extracting the corresponding triggers and arguments.Experiments with fine-tuned EE models, large language models (LLMs), and human annotators reveal a performance gap, with current models struggling in domains such as sociology and humanities.SciEvent serves as a challenging benchmark and a step toward generalizable, multi-domain SciIE. Dataset #Doc #Mentions Arg./Ent. Types Avg Bofu Dong, Pritesh Shah, Sumedh Sonawane, Tiyasha Banerjee, Erin Brady, Xinya Du |
EMNLP | 6 |
| 2025 | Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing UncertaintyabstractAgentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval.However, these systems often exhibit sub-optimal search behaviors like over-search (retrieving redundant information) and under-search (failing to initiate retrieval for necessary information), which hinder efficiency and reliability.This work formally defines and quantifies these behaviors, revealing their prevalence across multiple QA datasets and agentic RAG systems (e.g., one model could have avoided searching in 27.7% of its search steps).Furthermore, we demonstrate a crucial link between these inefficiencies and the models' uncertainty regarding their own knowledge boundaries, where response accuracy correlates with model's uncertainty or confidence in its search decisions.To address this, we propose β-GRPO, a reinforcement learning-based training method that incorporates confidence threshold to reward high-certainty search decisions.Experiments on seven QA benchmarks show that β-GRPO enable a 3B model with better agentic RAG ability, outperforming other strong baselines with a 4% higher average exact match score, with lower over-search and under-search rate 1 . Xinlu Zhang, Xinya Du, Zhiyu Chen 0002 |
EMNLP | 4 |
| 2025 | LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchabstractShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang, Daoyang Li, Liqiang Jing, Kaiyu He, Peilin Wu, Juntong Ni, George Michalopoulos, Yue Zhang, Ziyang Zhang, Mian Zhang, Zhiyu Chen, Xinya Du. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziming Luo, Daoyang Li, Liqiang Jing, Kaiyu He, Juntong Ni, George Michalopoulos, Zhiyu Chen 0002, Xinya Du |
EMNLP | 15 |
| 2025 | DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?abstractLarge Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI software engineers. Recently, many data science benchmarks have been proposed to investigate their performance in the data science domain. However, existing data science benchmarks still fall short when compared to real-world data science applications due to their simplified settings. To bridge this gap, we introduce DSBench, a comprehensive benchmark designed to evaluate data science agents with realistic tasks. This benchmark includes 466 data analysis tasks and 74 data modeling tasks, sourced from Eloquence and Kaggle competitions. DSBench offers a realistic setting by encompassing long contexts, multimodal task backgrounds, reasoning with large data files and multi-table structures, and performing end-to-end data modeling tasks. Our evaluation of state-of-the-art LLMs, LVLMs, and agents shows that they struggle with most tasks, with the best agent solving only 34.12% of data analysis tasks and achieving a 34.74% Relative Performance Gap (RPG). These findings underscore the need for further advancements in developing more practical, intelligent, and autonomous data science agents. Liqiang Jing, Zhehui Huang, Xiaoyang Wang 0001, Wenlin Yao, Wenhao Yu 0002, Kaixin Ma, Hongming Zhang 0009, Xinya Du, Dong Yu 0001 |
ICLR | 8 |
| 2025 | VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
Haojian Huang, Shengqiong Wu, Meng Luo 0010, Jinlan Fu, Xinya Du, Hanwang Zhang, Hao Fei 0001 |
ICML | 6 |
| 2025 | Tutorial Proposal: Hallucinations in Large Language Models and Large Vision-Language ModelsabstractMultimodal Large Language Models (MLLMs), Large Vision-Language Models (LVLMs), and Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, including text-based reasoning and multimodal content generation. However, these models frequently generate hallucinations-factually incorrect or misleading content-that pose significant challenges, particularly in high-stakes domains such as healthcare, law, and finance. This tutorial provides a comprehensive exploration of hallucinations in MLLMs, LVLMs, and LLMs, examining their causes, detection methods, and mitigation strategies. We discuss different types of hallucination evaluation and benchmarking and explore state-of-the-art techniques for hallucination mitigation. Liqiang Jing, Yue Zhang 0096, Xinya Du |
ICMR | 3 |
| 2024 | Making Natural Language Reasoning Explainable and FaithfulabstractNeural models, including large language models (LLMs), achieve superior performance on logical reasoning tasks such as question answering. To elicit reasoning capabilities from LLMs, recent works propose using the chain-of-thought (CoT) mechanism to generate both the reasoning chain and the answer, which enhances the model’s capabilities in conducting reasoning. However, due to LLM’s uninterpretable nature and the extreme flexibility of free-form explanations, several challenges remain: such as struggling with inaccurate reasoning, hallucinations, and not aligning with human preferences. In this talk, we will focus on (1) our design of leveraging structured information (that is grounded to the context), for the explainable complex question answering and reasoning; (2) our multi-module interpretable framework for inductive reasoning, which conducts step-wise faithful reasoning with iterative feedback. Xinya Du |
AAAI | 1 |
| 2024 | Language Models as Inductive ReasonersabstractZonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zonglin Yang 0001, Li Dong 0004, Xinya Du, Hao Cheng 0002, Erik Cambria, Xiaodong Liu 0003, Jianfeng Gao 0001, Furu Wei |
EACL (1) | 3 |
| 2024 | Proto-CLIP: Vision-Language Prototypical Network for Few-Shot LearningabstractWe propose a novel framework for few-shot learning by leveraging large-scale vision-language models such as CLIP [1]. Motivated by unimodal prototypical networks for few-shot learning, we introduce Proto-CLIP which utilizes image prototypes and text prototypes for few-shot learning. Specifically, Proto-CLIP adapts the image and text encoder embeddings from CLIP in a joint fashion using few-shot examples. The embeddings from the two encoders are used to compute the respective prototypes of image classes for classification. During adaptation, we propose aligning the image and text prototypes of the corresponding classes. Such alignment is beneficial for few-shot classification due to the reinforced contributions from both types of prototypes. Proto-CLIP has both training-free and fine-tuned variants. We demonstrate the effectiveness of our method by conducting experiments on benchmark datasets for few-shot learning, as well as in the real world for robot perception1. Jishnu Jaykumar, Kamalesh Palanisamy, Yu-Wei Chao, Xinya Du, Yu Xiang 0001 |
IROS | 4 |
| 2024 | IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringabstractTo evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on directly assessing the immediate responses generated by the models based on the given question and context. In the common use case of humans seeking AI assistant’s help in finding information, these non-interactive evaluations do not account for the dynamic nature of human-model conversations, and interaction-aware evaluations have shown that accurate models are not necessarily preferred by humans Lee et al. Recent works in human-computer interaction (HCI) have employed human evaluators to conduct interactions and evaluations, but they are often prohibitively expensive and time-consuming to scale. In this work, we introduce an automated evaluation framework IQA-EVAL to Interactive Question Answering Evaluations, more specifically, we introduce LLM-based Evaluation Agent (LEA) that can: (1) simulate human behaviors to generate interactions with IQA models; (2) automatically evaluate the generated interactions. Moreover, we propose assigning personas to LEAs to better simulate groups of real human evaluators. We show that: (1) our evaluation framework with GPT-4 (or Claude) as the backbone model achieves a high correlation with human evaluations on the IQA task; (2) assigning personas to LEA to better represent the crowd further significantly improves correlations. Finally, we use our automated metric to evaluate five recent LLMs with over 1000 questions from complex and ambiguous question answering tasks, which would cost $5k if evaluated by humans. Ruosen Li, Barry Wang, Xinya Du |
NeurIPS | 4 |
| 2024 | MEQA: A Benchmark for Multi-hop Event-centric Question Answering with ExplanationsabstractExisting benchmarks for multi-hop question answering (QA) primarily evaluate models based on their ability to reason about entities and the relationships between them. However, there's a lack of insight into how these models perform in terms of both events and entities. In this paper, we introduce a novel semi-automatic question generation strategy by composing event structures from information extraction (IE) datasets and present the first Multi-hop Event-centric Question Answering (MEQA) benchmark. It contains (1) 2,243 challenging questions that require a diverse range of complex reasoning over entity-entity, entity-event, and event-event relations; (2) corresponding multi-step QA-format event reasoning chain (explanation) which leads to the answer for each question. We also introduce two metrics for evaluating explanations: completeness and logical consistency. We conduct comprehensive benchmarking and analysis, which shows that MEQA is challenging for the latest state-of-the-art models encompassing large language models (LLMs); and how they fall short of providing faithful explanations of the event-centric reasoning process. Ruosen Li, Son Quoc Tran, Lei Xia 0004, Xinya Du |
NeurIPS | 5 |
| 2023 | End-to-end Case-Based Reasoning for Commonsense Knowledge Base CompletionabstractPretrained language models have been shown to store knowledge in their parameters and have achieved reasonable performance in commonsense knowledge base completion (CKBC) tasks.However, CKBC is knowledge-intensive and it is reported that pretrained language models' performance in knowledge-intensive tasks are limited because of their incapability of accessing and manipulating knowledge.As a result, we hypothesize that providing retrieved passages that contain relevant knowledge as additional input to the CKBC task will improve performance.In particular, we draw insights from Case-Based Reasoning (CBR) -which aims to solve a new problem by reasoning with retrieved relevant cases, and investigate the direct application of it to CKBC.On two benchmark datasets, we demonstrate through automatic and human evaluations that our End-to-end Case-Based Reasoning Framework (ECBRF) generates more valid knowledge than the state-of-the-art COMET model for CKBC in both the fully supervised and few-shot settings.From the perspective of CBR, our framework addresses a fundamental question on whether CBR methodology can be utilized to improve deep learning models. Zonglin Yang 0001, Xinya Du, Erik Cambria, Claire Cardie |
EACL | 2 |
| 2023 | POE: Process of Elimination for Multiple Choice ReasoningabstractLanguage models (LMs) are capable of conducting in-context learning for multiple choice reasoning tasks, but the options in these tasks are treated equally.As humans often first eliminate wrong options before picking the final correct answer, we argue a similar two-step strategy can make LMs better at these tasks.To this end, we present the Process of Elimination (POE), a two-step scoring method.In the first step, POE scores each option, and eliminates seemingly wrong options.In the second step, POE masks these wrong options, and makes the final prediction from the remaining options.Zero-shot experiments on 8 reasoning tasks illustrate the effectiveness of POE, and a following analysis finds our method to be especially performant on logical reasoning tasks.We further analyze the effect of masks, and show that POE applies to few-shot settings and large language models (LLMs) like ChatGPT. 1 Chenkai Ma, Xinya Du |
EMNLP | 2 |
| 2023 | Logical Entity Representation in Knowledge-Graphs for Differentiable Rule Learning
Chi Han, Qizheng He, Charles Yu, Xinya Du, Hanghang Tong, Heng Ji 0001 |
ICLR | 4 |
| 2022 | Automatic Error Analysis for Document-level Information ExtractionabstractAliva Das, Xinya Du, Barry Wang, Kejian Shi, Jiayuan Gu, Thomas Porter, Claire Cardie. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Aliva Das, Xinya Du, Barry Wang, Kejian Shi, Jiayuan Gu, Thomas Porter, Claire Cardie |
ACL (1) | 2 |
| 2022 | Dynamic Global Memory for Document-level Argument ExtractionabstractExtracting informative arguments of events from news articles is a challenging problem in information extraction, which requires a global contextual understanding of each document.While recent work on document-level extraction has gone beyond single-sentence and increased the cross-sentence inference capability of end-to-end models, they are still restricted by certain input sequence length constraints and usually ignore the global context between events.To tackle this issue, we introduce a new global neural generation-based framework for document-level event argument extraction by constructing a document memory store to record the contextual event information and leveraging it to implicitly and explicitly help with decoding of arguments for later events.Empirical results show that our framework outperforms prior methods substantially and it is more robust to adversarially annotated examples with our constrained decoding design. 1 Xinya Du, Heng Ji 0001 |
ACL (1) | 1 |
| 2022 | Retrieval-Augmented Generative Question Answering for Event Argument ExtractionabstractEvent argument extraction has long been studied as a sequential prediction problem with extractive-based methods, tackling each argument in isolation.Although recent work proposes generation-based methods to capture cross-argument dependency, they require generating and post-processing a complicated target sequence (template).Motivated by these observations and recent pretrained language models' capabilities of learning from demonstrations.We propose a retrieval-augmented generative QA model (R-GQA) for event argument extraction.It retrieves the most similar QA pair and augments it as prompt to the current example's context, then decodes the arguments as answers.Our approach outperforms substantially prior methods across various settings (i.e.fully supervised, domain transfer, and fewshot learning).Finally, we propose a clusteringbased sampling strategy (JointEnc) and conduct a thorough analysis of how different strategies influence the few-shot learning performance.1 Xinya Du, Heng Ji 0001 |
EMNLP | 1 |
| 2021 | GRIT: Generative Role-filler Transformers for Document-level Event Entity ExtractionabstractWe revisit the classic problem of documentlevel role-filler entity extraction (REE) for template filling.We argue that sentence-level approaches are ill-suited to the task and introduce a generative transformer-based encoderdecoder framework (GRIT) that is designed to model context at the document level: it can make extraction decisions across sentence boundaries; is implicitly aware of noun phrase coreference structure, and has the capacity to respect cross-role dependencies in the template structure.We evaluate our approach on the MUC-4 dataset, and show that our model performs substantially better than prior work.We also show that our modeling choices contribute to model performance, e.g., by implicitly capturing linguistic knowledge such as recognizing coreferent entity mentions. Xinya Du, Alexander M. Rush, Claire Cardie |
EACL | 1 |
| 2021 | Template Filling with Generative TransformersabstractTemplate filling is generally tackled by a pipeline of two separate supervised systemsone for role-filler extraction and another for template/event recognition.Since pipelines consider events in isolation, they can suffer from error propagation.We introduce a framework based on end-to-end generative transformers for this task (i.e., GTT).It naturally models the dependence between entities both within a single event and across the multiple events described in a document.Experiments demonstrate that this framework substantially outperforms pipeline-based approaches, and other neural end-to-end baselines that do not model between-event dependencies.We further show that our framework specifically improves performance on documents containing multiple events. Xinya Du, Alexander M. Rush, Claire Cardie |
NAACL-HLT | 1 |
| 2021 | Few-shot Intent Classification and Slot Filling with Retrieved ExamplesabstractDian Yu, Luheng He, Yuan Zhang, Xinya Du, Panupong Pasupat, Qi Li. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Luheng He, Yuan Zhang 0001, Xinya Du, Panupong Pasupat |
NAACL-HLT | 4 |
| 2020 | Document-Level Event Role Filler Extraction using Multi-Granularity Contextualized EncodingabstractFew works in the literature of event extraction have gone beyond individual sentences to make extraction decisions.This is problematic when the information needed to recognize an event argument is spread across multiple sentences.We argue that document-level event extraction is a difficult task since it requires a view of a larger context to determine which spans of text correspond to event role fillers.We first investigate how end-toend neural sequence models (with pre-trained language model representations) perform on document-level role filler extraction, as well as how the length of context captured affects the models' performance.To dynamically aggregate information captured by neural representations learned at different levels of granularity (e.g., the sentence-and paragraph-level), we propose a novel multi-granularity reader.We evaluate our models on the MUC-4 event extraction dataset, and show that our best system performs substantially better than prior work.We also report findings on the relationship between context length and neural model performance on the task. Xinya Du, Claire Cardie |
ACL | 1 |
| 2020 | Event Extraction by Answering (Almost) Natural QuestionsabstractThe problem of event extraction requires detecting the event trigger and extracting its corresponding arguments.Existing work in event argument extraction typically relies heavily on entity recognition as a preprocessing/concurrent step, causing the well-known problem of error propagation.To avoid this issue, we introduce a new paradigm for event extraction by formulating it as a question answering (QA) task that extracts the event arguments in an end-to-end manner.Empirical results demonstrate that our framework outperforms prior methods substantially; in addition, it is capable of extracting event arguments for roles not seen at training time (i.e., in a zeroshot learning setting).1 Xinya Du, Claire Cardie |
EMNLP (1) | 1 |
| 2018 | Harvesting Paragraph-level Question-Answer Pairs from WikipediaabstractWe study the task of generating from Wikipedia articles question-answer pairs that cover content beyond a single sentence.We propose a neural network approach that incorporates coreference knowledge via a novel gating mechanism.Compared to models that only take into account sentence-level information (Heilman and Smith, 2010; Du et al., 2017;Zhou et al., 2017), we find that the linguistic knowledge introduced by the coreference representation aids question generation significantly, producing models that outperform the current state-of-theart.We apply our system (composed of an answer span extraction system and the passage-level QG system) to the 10,000 top-ranking Wikipedia articles and create a corpus of over one million questionanswer pairs.We also provide a qualitative analysis for this large-scale generated corpus from Wikipedia. Xinya Du, Claire Cardie |
ACL (1) | 1 |
| 2017 | Learning to Ask: Neural Question Generation for Reading ComprehensionabstractWe study automatic question generation for sentences from text passages in reading comprehension.We introduce an attention-based sequence learning model for the task and investigate the effect of encoding sentence-vs.paragraph-level information.In contrast to all previous work, our model does not rely on hand-crafted rules or a sophisticated NLP pipeline; it is instead trainable end-to-end via sequenceto-sequence learning.Automatic evaluation results show that our system significantly outperforms the state-of-the-art rule-based system.In human evaluations, questions generated by our system are also rated as being more natural (i.e., grammaticality, fluency) and as more difficult to answer (in terms of syntactic and lexical divergence from the original text and reasoning needed to answer). Xinya Du, Junru Shao, Claire Cardie |
ACL (1) | 1 |
| 2017 | Identifying Where to Focus in Reading Comprehension for Neural Question GenerationabstractA first step in the task of automatically generating questions for testing reading comprehension is to identify questionworthy sentences, i.e. sentences in a text passage that humans find it worthwhile to ask questions about.We propose a hierarchical neural sentence-level sequence tagging model for this task, which existing approaches to question generation have ignored.The approach is fully data-driven -with no sophisticated NLP pipelines or any hand-crafted rules/features -and compares favorably to a number of baselines when evaluated on the SQuAD data set.When incorporated into an existing neural question generation system, the resulting end-to-end system achieves stateof-the-art performance for paragraph-level question generation for reading comprehension. Xinya Du, Claire Cardie |
EMNLP | 1 |