EDBT 2026 Demo / reviewers in the wild / expert
Xiangkun Hu
dblp:224/5990
· DBLP profile ↗
20ranked-venue papers
3as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and SummarizationabstractAs online platforms and recommendation algorithms evolve, people are increasingly trapped in echo chambers, leading to biased understandings of various issues. To combat this issue, we have introduced PerSphere, a benchmark designed to facilitate multi-faceted perspective retrieval and summarization, thus breaking free from these information silos. For each query within PerSphere, there are two opposing claims, each supported by distinct, non-overlapping perspectives drawn from one or more documents. Our goal is to accurately summarize these documents, aligning the summaries with the respective claims and their underlying perspectives. This task is structured as a two-step end-to-end pipeline that includes comprehensive document retrieval and multi-faceted summarization. Furthermore, we propose a set of metrics to evaluate the comprehensiveness of the retrieval and summarization content. Experimental results on various counterparts for the pipeline show that recent models struggle with such a complex task. Analysis shows that the main challenge lies in long context and perspective extraction, and we propose a simple but effective multi-agent summarization system, offering a promising solution to enhance performance on PerSphere. Yingjie Li 0008, Xiangkun Hu, Qinglin Qi, Qipeng Guo, Zheng Zhang 0001, Yue Zhang 0004 |
ACL (1) | 3 |
| 2025 | Common Learning Constraints Alter Interpretations of Direct Preference OptimizationabstractLarge language models in the past have typically relied on some form of reinforcement learning with human feedback (RLHF) to better align model responses with human preferences. However, because of oft-observed instabilities when implementing these RLHF pipelines, various reparameterization techniques have recently been introduced to sidestep the need for separately learning an RL reward model. Instead, directly fine-tuning for human preferences is achieved via the minimization of a single closed-form training objective, a process originally referred to as direct preference optimization (DPO). Although effective in certain real-world settings, we detail how the foundational role of DPO reparameterizations (and equivalency to applying RLHF with an optimal reward) may be obfuscated once inevitable optimization constraints are introduced during model training. This then motivates alternative derivations and analysis of DPO that remain intact even in the presence of such constraints. As initial steps in this direction, we re-derive DPO from a simple Gaussian estimation perspective, with strong ties to compressive sensing and classical constrained optimization problems involving noise-adaptive, concave regularization. Lemin Kong, Xiangkun Hu, Tong He 0002, David P. Wipf |
AISTATS | 2 |
| 2025 | Multi-Document Event Extraction Using Large and Small Language ModelsabstractMulti-document event extraction aims to aggregate event information from diverse sources for a comprehensive understanding of complex events.Despite its practical significance, this task has received limited attention in existing research.The inherent challenges include handling complex reasoning over long contexts and intricate event structures.In this paper, we propose a novel collaborative framework that integrates large language models for multi-step reasoning and fine-tuned small language models to handle key subtasks, guiding the overall reasoning process.We introduce a new benchmark for multi-document event extraction and propose an evaluation metric designed for comprehensive assessment of multiple aggregated events.Experimental results demonstrate that our approach significantly outperforms existing methods, providing new insights into collaborative reasoning to tackle the complexities of multi-document event extraction. Qingkai Min, Zitian Qu, Qipeng Guo, Xiangkun Hu, Yue Zhang 0004 |
EMNLP | 4 |
| 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsabstractLarge Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods-brittle prompt engineering or RAG-based reinforcement learning in controlled environments-fail to capture real-world complexities.In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions.Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web.We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges.Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents.Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers.Our results highlight that end-to-end training in realworld web environments is fundamental for developing robust research capabilities aligned with real-world applications.The source code for DeepResearcher is released at: https:// github.com/GAIR-NLP/DeepResearcher. Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu 0003 |
EMNLP | 3 |
| 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensabstractRecent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this gap, we introduce NovelQA, a benchmark tailored for evaluating LLMs with complex, extended narratives. NovelQA, constructed from English novels, offers a unique blend of complexity, length, and narrative coherence, making it an ideal tool for assessing deep textual understanding in LLMs. This paper details the design and construction of NovelQA, focusing on its comprehensive manual annotation process and the variety of question types aimed at evaluating nuanced comprehension. Our evaluation of long-context LLMs on NovelQA reveals significant insights into their strengths and weaknesses. Notably, the models struggle with multi-hop reasoning, detail-oriented questions, and handling extremely long inputs, averaging over 200,000 tokens. Results highlight the need for substantial advancements in LLMs to enhance their long-context comprehension and contribute effectively to computational literary analysis. Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng 0001, Guangsheng Bao, Xiangkun Hu, Zheng Zhang 0001, Yue Zhang 0004 |
ICLR | 8 |
| 2025 | Explicit Preference Optimization: No Need for an Implicit Reward ModelabstractThe generated responses of large language models (LLMs) are often fine-tuned to human preferences through a process called reinforcement learning from human feedback (RLHF). As RLHF relies on a challenging training sequence, whereby a separate reward model is independently learned and then later applied to LLM policy updates, ongoing research effort has targeted more straightforward alternatives. In this regard, direct preference optimization (DPO) and its many offshoots circumvent the need for a separate reward training step. Instead, through the judicious use of a reparameterization trick that induces an implicit reward, DPO and related methods consolidate learning to the minimization of a single loss function. And yet despite demonstrable success in some real-world settings, we prove that DPO-based objectives are nonetheless subject to sub-optimal regularization and counter-intuitive interpolation behaviors, underappreciated artifacts of the reparameterizations upon which they are based. To this end, we introduce an explicit preference optimization framework termed EXPO that requires no analogous reparameterization to achieve an implicit reward. Quite differently, we merely posit intuitively-appealing regularization factors from scratch that transparently avoid the potential pitfalls of key DPO variants, provably satisfying regularization desiderata that prior methods do not. Empirical results serve to corroborate our analyses and showcase the efficacy of EXPO. Xiangkun Hu, Lemin Kong, Tong He 0002, David P. Wipf |
ICML | 1 |
| 2025 | CCDFormer: A dual-backbone complex crack detection network with transformer
Xiangkun Hu, Hua Li 0019, Yixiong Feng, Songrong Qian, Shaobo Li 0001 |
Pattern Recognit. | 1 |
| 2024 | Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language ModelsabstractCross-document event coreference resolution (CDECR) involves clustering event mentions across multiple documents that refer to the same real-world events.Existing approaches utilize fine-tuning of small language models (SLMs) like BERT to address the compatibility among the contexts of event mentions.However, due to the complexity and diversity of contexts, these models are prone to learning simple co-occurrences.Recently, large language models (LLMs) like ChatGPT have demonstrated impressive contextual understanding, yet they encounter challenges in adapting to specific information extraction (IE) tasks.In this paper, we propose a collaborative approach for CDECR, leveraging the capabilities of both a universally capable LLM and a task-specific SLM.The collaborative strategy begins with the LLM accurately and comprehensively summarizing events through prompting.Then, the SLM refines its learning of event representations based on these insights during fine-tuning.Experimental results demonstrate that our approach surpasses the performance of both the large and small language models individually, forming a complementary advantage.Across various datasets, our approach achieves stateof-the-art performance, underscoring its effectiveness in diverse scenarios. Qingkai Min, Qipeng Guo, Xiangkun Hu, Songfang Huang |
ACL (1) | 3 |
| 2024 | Knowledge-Centric Hallucination DetectionabstractXiangkun Hu, Dongyu Ru, Lin Qiu, Qipeng Guo, Tianhang Zhang, Yang Xu, Yun Luo, Pengfei Liu, Yue Zhang, Zheng Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Xiangkun Hu, Dongyu Ru, Qipeng Guo, Tianhang Zhang |
EMNLP | 1 |
| 2024 | Can Language Models Learn to Skip Steps?abstractTrained on vast corpora of human language, language models demonstrate emergent human-like reasoning abilities. Yet they are still far from true intelligence, which opens up intriguing opportunities to explore the parallels of humans and model behaviors. In this work, we study the ability to skip steps in reasoning—a hallmark of human expertise developed through practice. Unlike humans, who may skip steps to enhance efficiency or to reduce cognitive load, models do not inherently possess such motivations to minimize reasoning steps. To address this, we introduce a controlled framework that stimulates step-skipping behavior by iteratively refining models to generate shorter and accurate reasoning paths. Empirical results indicate that models can develop the step skipping ability under our guidance. Moreover, after fine-tuning on expanded datasets that include both complete and skipped reasoning sequences, the models can not only resolve tasks with increased efficiency without sacrificing accuracy, but also exhibit comparable and even enhanced generalization capabilities in out-of-domain scenarios. Our work presents the first exploration into human-like step-skipping ability and provides fresh perspectives on how such cognitive abilities can benefit AI models. Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
NeurIPS | 3 |
| 2024 | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented GenerationabstractDespite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001 |
NeurIPS | 3 |
| 2023 | An AMR-based Link Prediction Approach for Document-level Event Argument ExtractionabstractRecent works have introduced Abstract Meaning Representation (AMR) for Document-level Event Argument Extraction (Doc-level EAE), since AMR provides a useful interpretation of complex semantic structures and helps to capture long-distance dependency.However, in these works AMR is used only implicitly, for instance, as additional features or training signals.Motivated by the fact that all event structures can be inferred from AMR, this work reformulates EAE as a link prediction problem on AMR graphs.Since AMR is a generic structure and does not perfectly suit EAE, we propose a novel graph structure, Tailored AMR Graph (TAG), which compresses less informative subgraphs and edge types, integrates span information, and highlights surrounding events in the same document.With TAG, we further propose a novel method using graph neural networks as a link prediction model to find event arguments.Our extensive experiments on WikiEvents and RAMS show that this simpler approach outperforms the state-of-the-art models by 3.63pt and 2.33pt F1, respectively, and do so with reduced 56% inference time.The code is available at https://github.com/ayyyq/TARA. Yuqing Yang 0004, Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
ACL (1) | 3 |
| 2023 | Dual Cache for Long Document Neural Coreference ResolutionabstractRecent works show the effectiveness of cachebased neural coreference resolution models on long documents.These models incrementally process a long document from left to right and extract relations between mentions and entities in a cache, resulting in much lower memory and computation cost compared to computing all mentions in parallel.However, they do not handle cache misses when high-quality entities are purged from the cache, which causes wrong assignments and leads to prediction errors.We propose a new hybrid cache that integrates two eviction policies to capture global and local entities separately, and effectively reduces the aggregated cache misses up to half as before, while improving F1 score of coreference by 0.7 ∼ 5.7pt.As such, the hybrid policy can accelerate existing cache-based models and offer a new long document coreference resolution solution.Results show that our method outperforms existing methods on four benchmarks while saving up to 83% of inference time against non-cache-based models.Further, we achieve a new state-of-the-art on a long document coreference benchmark, LitBank. Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
ACL (1) | 2 |
| 2023 | Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsabstractAs large language models (LLMs) have shown effectiveness with different prompting methods, such as Chain of Thought, Program of Thought, we find that these methods have formed a great complementarity to each other on math reasoning tasks.In this work, we propose XoT, an integrated problem solving framework by prompting LLMs with diverse reasoning thoughts.For each question, XoT always begins with selecting the most suitable method then executes each method iteratively.Within each iteration, XoT actively checks the validity of the generated answer and incorporates the feedback from external executors, allowing it to dynamically switch among different prompting methods.Through extensive experiments on 10 popular math reasoning datasets, we demonstrate the effectiveness of our proposed approach and thoroughly analyze the strengths of each module.Moreover, empirical results suggest that our framework is orthogonal to recent work that makes improvements on single reasoning methods and can further generalise to logical reasoning domain.By allowing method switching, XoT provides a fresh perspective on the collaborative integration of diverse reasoning thoughts in a unified framework. Tengxiao Liu, Qipeng Guo, Yuqing Yang 0004, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
EMNLP | 4 |
| 2023 | Evaluating Open-QA EvaluationabstractThis study focuses on the evaluation of the Open Question Answering (Open-QA) task, which can directly estimate the factuality of large language models (LLMs). Current automatic evaluation methods have shown limitations, indicating that human evaluation still remains the most reliable approach. We introduce a new task, QA Evaluation (QA-Eval) and the corresponding dataset EVOUNA, designed to assess the accuracy of AI-generated answers in relation to standard answers within Open-QA. Our evaluation of these methods utilizes human-annotated results to measure their performance. Specifically, the work investigates methods that show high correlation with human evaluations, deeming them more reliable. We also discuss the pitfalls of current methods and methods to improve LLM-based evaluators. We believe this new QA-Eval task and corresponding dataset EVOUNA will facilitate the development of more effective automatic evaluation tools and prove valuable for future research in this area. All resources are available at https://github.com/wangcunxiang/QA-Eval and it is under the Apache-2.0 License. Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Zhikun Xu, Yidong Wang 0003, Xiangkun Hu, Yue Zhang 0004 |
NeurIPS | 8 |
| 2022 | RLET: A Reinforcement Learning Based Approach for Explainable QA with Entailment TreesabstractInterpreting the reasoning process from questions to answers poses a challenge in approaching explainable QA.A recently proposed structured reasoning format, entailment tree, manages to offer explicit logical deductions with entailment steps in a tree structure.To generate entailment trees, prior single pass sequence-tosequence models lack visible internal decision probability, while stepwise approaches are supervised with extracted single step data and cannot model the tree as a whole.In this work, we propose RLET, a Reinforcement Learning based Entailment Tree generation framework, which is trained utilising the cumulative signals across the whole tree.RLET iteratively performs single step reasoning with sentence selection and deduction generation modules, from which the training signal is accumulated across the tree with elaborately designed aligned reward function that is consistent with the evaluation.To the best of our knowledge, we are the first to introduce RL into the entailment tree generation task.Experiments on three settings of the EntailmentBank dataset demonstrate the strength of using RL framework. Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Yue Zhang 0004, Xipeng Qiu, Zheng Zhang 0001 |
EMNLP | 3 |
| 2022 | BART-Reader: Predicting Relations Between Entities via Reading Their Document-Level Context Information
Hang Yan 0001, Yu Sun 0031, Junqi Dai, Xiangkun Hu, Qipeng Guo, Xipeng Qiu, Xuanjing Huang 0001 |
NLPCC (1) | 4 |
| 2021 | Co-Attention Memory Network for Multimodal Microblog's Hashtag RecommendationabstractHashtags are keywords describing a topic or a theme and are usually chosen by microblogging users. Hence, the hashtags can be used to categorize microblog posts. With the fast development of the social network, the task of recommending suitable hashtags has received considerable attention in recent years. Recently, most neural network methods have treated the task as a multi-class classification problem. In fact, users are constantly introducing new hashtags in a highly dynamic way. Treating the task as a multi-class classification problem with a fixed number of target categories does not allow the method to deal with the new hashtags. To address this problem, the task is reinterpreted as a matching problem and a novel co-attention memory network is proposed to represent the multimodal microblogs and hashtags. We utilize a co-attention mechanism to model the multimodal mircroblogs, and utilize the post history to represent the hashtags. Experimental results on a Twitter-based dataset demonstrated that the proposed method can achieve better performance than the current state-of-the-art methods that treat the task as a multi-class classification problem. Renfeng Ma, Xipeng Qiu, Qi Zhang 0001, Xiangkun Hu, Yu-Gang Jiang 0001, Xuanjing Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2019 | Hot Topic-Aware Retweet Prediction with Masked Self-attentive ModelabstractSocial media users create millions of microblog entries on various topics each day. Retweet behaviour play a crucial role in spreading topics on social media. Retweet prediction task has received considerable attention in recent years. The majority of existing retweet prediction methods are focus on modeling user preference by utilizing various information, such as user profiles, user post history, user following relationships, etc. Yet, the users exposures towards real-time posting from their followees contribute significantly to making retweet predictions, considering that the users may participate into the hot topics discussed by their followees rather than be limited to their previous interests. To make efficient use of hot topics, we propose a novel masked self-attentive model to perform the retweet prediction task by perceiving the hot topics discussed by the users' followees. We incorporate the posting histories of users with external memory and utilize a hierarchical attention mechanism to construct the users' interests. Hence, our model can be jointly hot-topic aware and user interests aware to make a final prediction. Experimental results on a dataset collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods. Renfeng Ma, Xiangkun Hu, Qi Zhang 0001, Xuanjing Huang 0001, Yu-Gang Jiang 0001 |
SIGIR | 2 |
| 2018 | Incorporating Argument-Level Interactions for Persuasion Comments Evaluation using Co-attention ModelabstractIn this paper, we investigate the issue of persuasiveness evaluation for argumentative comments. Most of the existing research explores different text features of reply comments on word level and ignores interactions between participants. In general, viewpoints are usually expressed by multiple arguments and exchanged on argument level. To better model the process of dialogical argumentation, we propose a novel co-attention mechanism based neural network to capture the interactions between participants on argument level. Experimental results on a publicly available dataset show that the proposed model significantly outperforms some state-of-the-art methods for persuasiveness evaluation. Further analysis reveals that attention weights computed in our model are able to extract interactive argument pairs from the original post and the reply. Lu Ji, Zhongyu Wei, Xiangkun Hu, Yang Liu 0004, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 3 |