Enrui Hu

dblp:297/0008 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0007-5427-6175ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Beyond the Surface: Measuring Self-Preference in LLM Judgments
abstract
Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by other models.Existing methods typically measure this bias by calculating the difference between the scores a judge model assigns to its own responses and those it assigns to responses from other models.However, this approach conflates self-preference bias with response quality, as higher-quality responses from the judge model may also lead to positive score differences, even in the absence of bias.To address this issue, we introduce gold judgments as proxies for the actual quality of responses and propose the DBG score, which measures self-preference bias as the difference between the scores assigned by the judge model to its own responses and the corresponding gold judgments.Since gold judgments reflect true response quality, the DBG score mitigates the confounding effect of response quality on bias measurement.Using the DBG score, we conduct comprehensive experiments to assess self-preference bias across LLMs of varying versions, sizes, and reasoning abilities.Additionally, we investigate two factors that influence and help alleviate self-preference bias: response text style and the post-training data of judge models.Finally, we explore potential underlying mechanisms of self-preference bias from an attention-based perspective.Our code and data are available at https://github.com/ zhiyuanc2001/self-preference.Model B Gold Model A Judge Model 31.8% 68.2% 33.9% 66.1% 44.4% 55.6% Model A: Llama-3.1-8BModel B: Qwen2.5
Zhi-Yuan Chen, Hao Wang 0139, Enrui Hu, Yankai Lin 0001
EMNLP4
2025 Generalizing Experience for Language Agents with Hierarchical MetaFlows
abstract
Recent efforts to employ large language models (LLMs) as agents have demonstrated promising results in a wide range of multi-step agent tasks. However, existing agents lack an effective experience reuse approach to leverage historical completed tasks. In this paper, we propose a novel experience reuse framework MetaFlowLLM, which constructs a hierarchical experience tree from historically completed tasks. Each node in this experience tree is presented as a MetaFlow which contains static execution workflow and subtask required by agents to complete dynamically. Then, we propose a Hierarchical MetaFlow Merging algorithm to construct the hierarchical experience tree. When accomplishing a new task, MetaFlowLLM can first retrieve the most relevant MetaFlow node from the experience tree and then execute it accordingly. To effectively generate valid MetaFlows from historical data, we further propose a reinforcement learning pipeline to train the MetaFlowGen. Extensive experimental results on AppWorld and WorkBench demonstrate that integrating with MetaFlowLLM, existing agents (e.g., ReAct, Reflexion) can gain substantial performance improvement with reducing execution costs. Notably, MetaFlowLLM achieves an average success rate improvement of 32.3% on AppWorld and 6.2% on WorkBench, respectively.
Shengda Fan, Xin Cong, Zhong Zhang 0004, Yuepeng Fu, Yesai Wu, Hao Wang 0139, Enrui Hu, Yankai Lin 0001
NeurIPS8
2024 Enhancing Code Representation Learning for Code Search with Abstract Code Semantics
abstract
Code representation learning is an important way to encode the semantics of source code through pre-training. The learned representation supports a variety of downstream tasks, such as natural language code search and code defect detection. Inspired by pre-trained models for natural language representation learning, existing approaches often treat the source code or its structural information (e.g., Abstract Syntax Tree or AST) as a plain token sequence. Unlike natural language, programming language has its unique code unit information (e.g., identifiers and expressions) and logic information (e.g., the functionality of a code snippet). To further explore those properties, we propose Abstract Code Embedding (AbCE), a self-supervised learning method that considers the abstract semantics of code logic. Instead of scattered tokens, AbCE treats an entire node or a subtree in an AST as a basic code unit during pre-training, which preserves the entirety of a coding unit. Moreover, AbCE learns the abstract semantics of AST nodes via a self-distillation way. Experimental results show that it achieves significant improvements over state-of-the-art baselines on code search tasks and comparable performance on code clone detection and defect detection tasks even without using contrastive learning or curriculum learning.
Yiwei Ding, Enrui Hu, Yue Yu 0001
IJCNN3
2022 Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering
abstract
Jiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Jiawei Zhou 0003, Lifeng Shang, Ke Zhan, Enrui Hu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002
ACL (1)6
2022 Triple-Fact Retriever: An explainable reasoning retrieval model for multi-hop QA problem
abstract
Nowadays, multi-hop question answer (QA) problem is challenging and not well solved in the QA community. The dominant bottleneck of the multi-hop QA problem is the need for a reasoning retriever to fetch a document path from an open-domain corpus (e.g., Wikipedia). A reasoning retriever aims to collect an evidence document from large corpora at one hop retrieval and aggregate the evidence for subsequent hop retrieval, which yields a document path after multi-hop retrieval. There exist two challenges, (1) to fetch the evidence document in an efficient and explainable way at one hop retrieval and (2) to update the question information by aggregating the evidence from the retrieved document after each hop retrieval. To address these two challenges, we propose a triple-fact-based retrieval model to effectively retrieve a related document path in an explainable way for each question. We extract a structured representation from the unstructured document and utilize the knowledge of pre-trained language model (PLM) to do the semantic-level matching between the question and document. We evaluate the proposed Triple-fact Retriever model on the recently proposed open-domain multi-hop QA dataset, HotpotQA, and a cross-document multi-step Reading Comprehension dataset, Wikihop. The results11The source code is available on our website: https://github.com/Rebaccamin/triple_retriever. demonstrate that the Triple-fact retriever outperforms the existing baseline retrieval works.
Chengmin Wu, Enrui Hu, Ke Zhan, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Lei Chen 0002
ICDE2
2022 Leveraging Multi-view Inter-passage Interactions for Neural Document Ranking
abstract
The configuration of 512 window size prevents transformers from being directly applicable to document ranking that requires larger context. Hence, recent works propose to estimate document relevance with fine-grained passage-level relevance signals. A limitation of such models, however, is that scoring each passage independently falls short in modeling inter-passage interactions and leads to unsatisfactory results. In this paper, we propose a Multiview inter-passage Interaction based Ranking model (MIR), to combine intra-passage interactions and inter-passage interactions in a complementary manner. The former captures local semantic relations inside each passage, whereas the latter draws global dependencies between different passages. Moreover, we represent inter-passage relationships via multi-view attention patterns, allowing information propagation at token, sentence, and passage-level. The representations at different levels of granularity, being aware of global context, are then aggregated into a document-level representation for ranking. Experimental results on two benchmarks show that modeling inter-passage interactions brings substantial improvements over existing passage-level methods.
Chengzhen Fu, Enrui Hu, Letian Feng, Zhicheng Dou, Yantao Jia, Lei Chen 0002, Fan Yu 0004, Zhao Cao
WSDM2
2021 Answer Complex Questions: Path Ranker Is All You Need
abstract
Currently, the most popular method for open-domain Question Answering (QA) adopts "Retriever and Reader" pipeline, where the retriever extracts a list of candidate documents from a large set of documents followed by a ranker to rank the most relevant documents and the reader extracts answer from the candidates. Existing studies take the greedy strategy in the sense that they only use samples for ranking at the current hop, and ignore the global information across the whole documents. In this paper, we propose a purely rank-based framework Thinking Path Re-Ranker (TPRR), which is comprised of Thinking Path Ranker (TPR) for generating document sequences called "a path" and External Path Reranker (EPR) for selecting the best path from candidate paths generated by TPR. Specifically, TPR leverages the scores of a dense model and conditional probabilities to score the full paths. Moreover, to further enhance the performance of the dense ranker in the iterative training, we propose a "thinking" negatives selection method that the top-K candidates treated as negatives in the current hop are adjusted dynamically through supervised signals. After achieving multiple supporting paths through TPR, the EPR component which integrates several fine-grained training tasks for QA is used to select the best path for answer extraction. We have tested our proposed solution on the multi-hop dataset "HotpotQA" with a full wiki set ting, and the results show that TPRR significantly outperforms the existing state-of-the-art models. Moreover, our method has won the first place in the HotpotQA official leaderboard since Feb 1, 2021 under the Fullwiki setting. Code is available at https://gitee.com/mindspore/mindspore/ tree/master/model_zoo/research/nlp/tprr.
Xinyu Zhang 0019, Ke Zhan, Enrui Hu, Chengzhen Fu, Hao Jiang 0022, Yantao Jia, Fan Yu 0004, Zhicheng Dou, Zhao Cao, Lei Chen 0002
SIGIR3