Sen Hu 0005

dblp:53/7997-5 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-4201-5919ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
abstract
Beyond scratch coding, exploiting large-scale code repositories (e.g., GitHub) for practical tasks is vital in real-world software development, yet current benchmarks rarely evaluate code agents in such authentic, workflow-driven scenarios. To bridge this gap, we introduce GitTaskBench, a benchmark designed to systematically assess this capability via 54 realistic tasks across 7 modalities and 7 domains. Each task pairs a relevant repository with an automated, human-curated evaluation harness specifying practical success criteria. Beyond measuring execution and task success, we also propose the alpha-value metric to quantify the economic benefit of agent performance, which integrates task success rates, token cost, and average developer salaries. Experiments across three state-of-the-art agent frameworks with multiple advanced LLMs show that leveraging code repositories for complex task solving remains challenging: even the best-performing system, OpenHands+Claude 3.7, solves only 48.15% of tasks. Error analysis attributes over half of failures to seemingly mundane yet critical steps like environment setup and dependency resolution, highlighting the need for more robust workflow management and increased timeout preparedness. By releasing GitTaskBench, we aim to drive progress and attention toward repository-aware code reasoning, execution, and deployment---moving agents closer to solving complex, end-to-end real-world tasks.
Ziyi Ni, Huacan Wang, Shuo Lu, Wang You, Zhenheng Tang, Sen Hu 0005, Bo Li 0117, Binxing Jiao, Daxin Jiang, Yuntao Du 0001
AAAI8
2026 Does Memory Need Graphs? A Unified Framework and Empirical Analysis for Long-Term Dialog Memory
abstract
Sen Hu, Yuxiang Wei, Jiaxin Ran, Xueran Han, Zhiyuan Yao, Huacan Wang, Ronghao Chen, Lei Zou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sen Hu 0005, Jiaxin Ran, Xueran Han, Huacan Wang, Ronghao Chen, Lei Zou 0001
ACL (1)1
2026 CloneMem: Benchmarking Long-Term Memory for AI Clones
abstract
Sen Hu, Zhiyu Zhang, Yuxiang Wei, Xueran Han, Zhenheng Tang, Ronghao Chen, Huacan Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sen Hu 0005, Xueran Han, Zhenheng Tang, Ronghao Chen, Huacan Wang
ACL (1)1
2026 KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
abstract
Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tingyu Wu, Zhisheng Chen 0003, Ziyan Weng, Shuhe Wang, Sen Hu 0005, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen
ACL (1)6
2025 SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
abstract
Large Language Model (LLM)-based agents have recently shown impressive capabilities in complex reasoning and tool use via multi-step interactions with their environments. While these agents have the potential to tackle complicated tasks, their problem-solving process—agents' interaction trajectory leading to task completion—remains underexploited. These trajectories contain rich feedback that can navigate agents toward the right directions for solving problems correctly. Although prevailing approaches, such as Monte Carlo Tree Search (MCTS), can effectively balance exploration and exploitation, they ignore the interdependence among various trajectories and lack the diversity of search spaces, which leads to redundant reasoning and suboptimal outcomes. To address these challenges, we propose SE-Agent, a Self-Evolution framework that enables Agents to optimize their reasoning processes iteratively. Our approach revisits and enhances former pilot trajectories through three key operations: revision, recombination, and refinement. This evolutionary mechanism enables two critical advantages: (1) it expands the search space beyond local optima by intelligently exploring diverse solution paths guided by previous trajectories, and (2) it leverages cross-trajectory inspiration to efficiently enhance performance while mitigating the impact of suboptimal reasoning paths. Through these mechanisms, SE-Agent achieves continuous self-evolution that incrementally improves reasoning quality. We evaluate SE-Agent on SWE-bench Verified to resolve real-world GitHub issues. Experimental results across five strong LLMs show that integrating SE-Agent delivers up to 55% relative improvement, achieving state-of-the-art performance among all open-source agents on SWE-bench Verified.
Yifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han, Sen Hu 0005, Ziyi Ni, Mingguang Chen
NeurIPS5
2025 RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
abstract
The ultimate goal of code agents is to solve complex tasks autonomously. Although large language models (LLMs) have made substantial progress in code generation, real-world tasks typically demand full-fledged code repositories rather than simple scripts. Building such repositories from scratch remains a major challenge. Fortunately, GitHub hosts a vast, evolving collection of open-source repositories, which developers frequently reuse as modular components for complex tasks. Yet, existing frameworks like OpenHands and SWE-Agent still struggle to effectively leverage these valuable resources. Relying solely on README files provides insufficient guidance, and deeper exploration reveals two core obstacles: overwhelming information and tangled dependencies of repositories, both constrained by the limited context windows of current LLMs. To tackle these issues, we propose RepoMaster, an autonomous agent framework designed to explore and reuse GitHub repositories for solving complex tasks. For efficient understanding, RepoMaster constructs function-call graphs, module-dependency graphs, and hierarchical code trees to identify essential components, providing only identified core elements to the LLMs rather than the entire repository. During autonomous execution, it progressively explores related components using our exploration tools and prunes information to optimize context usage. Evaluated on the adjusted MLE-bench, RepoMaster achieves a 110\% relative boost in valid submissions over the strongest baseline OpenHands. On our newly released GitTaskBench, RepoMaster lifts the task-pass rate from 40.7% to 62.9% while reducing token usage by 95%. Our code and demonstration materials are publicly available at https://github.com/QuantaAlpha/RepoMaster.
Huacan Wang, Ziyi Ni, Shuo Lu, Sen Hu 0005, Jiaye Lin, Yifu Guo, Yuntao Du 0001
NeurIPS5
2024 Are LLM-based Evaluators Confusing NLG Quality Criteria?
abstract
Some prior work has shown that LLMs perform well in NLG evaluation for different tasks.However, we discover that LLMs seem to confuse different evaluation criteria, which reduces their reliability.For further verification, we first consider avoiding issues of inconsistent conceptualization and vague expression in existing NLG quality criteria themselves.So we summarize a clear hierarchical classification system for 11 common aspects with corresponding different criteria from previous studies involved.Inspired by behavioral testing, we elaborately design 18 types of aspect-targeted perturbation attacks for fine-grained analysis of the evaluation behaviors of different LLMs.We also conduct human annotations beyond the guidance of the classification system to validate the impact of the perturbations.Our experimental results reveal confusion issues inherent in LLMs, as well as other noteworthy phenomena, and necessitate further research and improvements for LLM-based evaluation. * Equal contribution. Prompt:Your task is to evaluate the summary written for a dialogue on the given criterion.
Xinyu Hu 0001, Mingqi Gao 0002, Sen Hu 0005, Yicheng Chen 0001, Teng Xu 0007, Xiaojun Wan 0001
ACL (1)3
2023 CORD: A Three-Stage Coarse-to-Fine Framework for Relation Detection in Knowledge Base Question Answering
abstract
As a fundamental subtask of Knowledge Base Question Answering (KBQA), Relation Detection (KBQA-RD) plays a crucial role to detect the KB relations between entities or variables in natural language questions. It remains, however, a challenging task, particularly for significant large-scale relations and in the presence of easily confused relations. Recent state-of-the-art methods not only struggle with such scenarios, but often take into account only one facet and fail to incorporate the subtle discrepancy among the relations. In this paper, we propose a simple and efficient three-stage framework to exploit the coarse-to-fine paradigm. Specifically, we employ a natural clustering over all KB relations and perform a coarse-to-fine relation recognition process based on the relation clustering. In this way, our framework (i.e., CORD) refines the detection of relations, so as to scale well with large-scale relations. Experiments on both single-relation (i.e., SimpleQuestions (SQ)) and multi-relation (i.e., WebQSP (WQ)) benchmarks show that CORD not only achieves the outstanding relation detection performance in KBQA-RD subtask; but more importantly, further improves the accuracy of KBQA systems.
Yanzeng Li, Sen Hu 0005, Wenjuan Han, Lei Zou 0001
CIKM2
2023 S2M: Converting Single-Turn to Multi-Turn Datasets for Conversational Question Answering
abstract
Supplying data augmentation to conversational question answering (CQA) can effectively improve model performance. However, there is less improvement from single-turn datasets in CQA due to the distribution gap between single-turn and multi-turn datasets. On the other hand, while numerous single-turn datasets are available, we have not utilized them effectively. To solve this problem, we propose a novel method to convert single-turn datasets to multi-turn datasets. The proposed method consists of three parts, namely, a QA pair Generator, a QA pair Reassembler, and a question Rewriter. Given a sample consisting of context and single-turn QA pairs, the Generator obtains candidate QA pairs and a knowledge graph based on the context. The Reassembler utilizes the knowledge graph to get sequential QA pairs, and the Rewriter rewrites questions from a conversational perspective to obtain a multi-turn dataset S2M. Our experiments show that our method can synthesize effective training resources for CQA. Notably, S2M ranks 1st place on the QuAC leaderboard (https://quac.ai/) at the time of submission (Aug 24th, 2022).
Baokui Li, Wangshu Zhang, Yicheng Chen 0001, Changlin Yang, Sen Hu 0005, Teng Xu 0007, Siye Liu, Jiwei Li 0001
ECAI6
2023 Read Key Points: Dialogue-Grounded Knowledge Points Generation with Multi-Level Salience-Aware Mixture
abstract
Knowledge-grounded dialogue (KGD) has become increasingly essential for online services, enabling individuals to obtain desired information. While KGD contains knowledge information, most knowledge points are fragmented and repeated in dialogues, making it difficult for users to quickly grasp complete and key information from a collection of sessions. In this paper, we propose a novel task of dialogue-grounded knowledge points generation (DialKPG) to condense a collection of sessions on a topic into succinct and complete knowledge points. To enable empirical study, we create TopicDial and OpenDial corpus based on two existing knowledge-grounded dialogue corpus FaithDial and OpenDialKG by a Three-Stage Annotation Framework, and establish a novel approach for DialKPG task, namely MSAM (Multi-Level Salience-Aware Mixture). MSAM explicitly incorporates salient information at the token-level, utterance-level, and session-level to better guide knowledge points generation. Extensive experiments have verified the effectiveness of our method over competitive baselines. Furthermore, our analysis shows that the proposed model is particularly effective at handling long inputs and multiple sessions due to its strong capability of duplicated elimination and knowledge integration.
Baokui Li, Wangshu Zhang, Changlin Yang, Yicheng Chen 0001, Sen Hu 0005, Teng Xu 0007, Jiwei Li 0001
ECAI6
2023 A Data-centric Solution to Improve Online Performance of Customer Service Bots
abstract
The online performance of customer service bots is often less than satisfactory because of the gap between limited training data and real-world user questions. As a straightforward way to improve online performance, model iteration and re-deployment are time consuming and labor-intensive, and therefore difficult to sustain. To fix badcases and improve online performance of chatbots in a timely and continuous manner, we propose a data-centric solution consisting of three main modules: badcase detection, bad case correction, and answer extraction. By making full use of online model signals, implicit user feedback and artificial customer service log, the proposed solution can fix online badcases automatically. Our solution has been deployed and bringing consistently positive impacts for hundreds of customer service bots used by Alipay app.
Sen Hu 0005, Changlin Yang, Siye Liu, Teng Xu 0007, Wangshu Zhang
SIGIR1
2022 COSSUM: Towards Conversation-Oriented Structured Summarization for Automatic Medical Insurance Assessment
abstract
In medical insurance industry, a lot of human labor is required to collect information of claimants. Human assessors need to converse with claimants in order to record key information and organize it into a structured summary. With the purpose of helping save human labor, we propose the task of conversation-oriented structured summarization which aims to automatically produce the desired structured summary from a conversation automatically. One major challenge of the task is that the structured summary contains multiple fields of different types. To tackle this problem, we propose a unified approach COSSUM based on prompting to generate the values of all fields simultaneously. By learning all fields together, our approach can capture the inherent relationship between them. Moreover, we propose a specially designed curriculum learning strategy for model training. Both automatic and human evaluations are performed, and the results show the effectiveness of our proposed approach.
Xiaojun Wan 0001, Sen Hu 0005, Mengdi Zhou, Teng Xu 0007, Haitao Mi
KDD3
2022 Knowledge based natural answer generation via masked-graph transformer
Sen Hu 0005, Lei Zou 0001
World Wide Web2
2020 The Value of Paraphrase for Knowledge Base Predicates
abstract
Paraphrase, i.e., differing textual realizations of the same meaning, has proven useful for many natural language processing (NLP) applications. Collecting paraphrase for predicates in knowledge bases (KBs) is the key to comprehend the RDF triples in KBs. Existing works have published some paraphrase datasets automatically extracted from large corpora, but have too many redundant pairs or don't cover enough predicates, which cannot be improved by computer only and need the help of human beings. This paper shows a full process of collecting large-scale and high-quality paraphrase dictionaries for predicates in knowledge bases, which takes advantage of existing datasets and combines the technologies of machine mining and crowdsourcing. Our dataset comprises 2284 distinct predicates in DBpedia and 31130 paraphrase pairs in total, the quality of which is a great leap over previous works. Then it is demonstrated that such good paraphrase dictionaries can do great help to natural language processing tasks such as question answering and language generation. We also publish our own dictionary for further research.
Bingcong Xue, Sen Hu 0005, Lei Zou 0001, Jiashu Cheng
AAAI2
2019 An Interactive Mechanism to Improve Question Answering Systems via Feedback
abstract
Semantic parsing-based RDF question answering (QA) systems are to interpret users' natural language questions as query graphs and return answers over RDF repository. However, due to the complexity of linking natural phrases with specific RDF items (e.g., entities and predicates), it remains difficult to understand users' question sentences precisely, hence QA systems may not meet users' expectation, offering wrong answers and dismissing some correct answers. In this paper, we design an I nteractive M echanism aiming for PRO motion V ia users' fe edback to Q A systems (IMPROVE-QA), a whole framework to not only make existing QA systems return more precise answers based on a few feedbacks over the original answers given by RDF QA systems, but also enhance paraphrasing dictionaries to ensure a continuous-learning capability in improving RDF QA systems. To provide better interactivity and online performance, we design a holistic graph mining algorithm (HWspan) to automatically refine the query graph. Extensive experiments on both Freebase and DBpedia confirm the effectiveness and superiority of our approach.
Xinbo Zhang, Lei Zou 0001, Sen Hu 0005
CIKM3
2019 How Question Generation Can Help Question Answering over Knowledge Base
Sen Hu 0005, Lei Zou 0001, Zhanxing Zhu
NLPCC (1)1
2018 A State-transition Framework to Answer Complex Questions over Knowledge Base
abstract
Although natural language question answering over knowledge graphs have been studied in the literature, existing methods have some limitations in answering complex questions.To address that, in this paper, we propose a State Transition-based approach to translate a complex natural language question N to a semantic query graph (SQG) Q S , which is used to match the underlying knowledge graph to find the answers to question N .In order to generate Q S , we propose four primitive operations (expand, fold, connect and merge) and a learning-based state transition approach.Extensive experiments on several benchmarks (such as QALD, WebQuestions and ComplexQuestions) with two knowledge bases (DBpedia and Freebase) confirm the superiority of our approach compared with stateof-the-arts.
Sen Hu 0005, Lei Zou 0001, Xinbo Zhang
EMNLP1
2018 Answering Natural Language Questions by Subgraph Matching over Knowledge Graphs (Extended Abstract)
abstract
RDF question/answering (Q/A) allows users to ask questions in natural languages over a knowledge base represented by RDF. To answer a natural language question, the existing works focus on question understanding to deal with the disambiguation of phrases linking, which ignore the query composition and execution. In this paper, we propose a systematic framework to answer natural language questions over RDF repository (RDF Q/A) from a graph data-driven perspective. We propose the (super) semantic query graph to model the query intention in the natural language question in a structural way, based on which, RDF Q/A is reduced to subgraph matching problem. More importantly, we resolve the ambiguity both of phrases and structures at the time when matches of query are found. To build the super semantic query graph, we propose a node-first framework which has high robustness and can tackle with complex questions. Extensive experiments confirm that our method not only improves the precision but also speeds up query performance greatly.
Sen Hu 0005, Lei Zou 0001, Jeffrey Xu Yu, Haixun Wang, Dongyan Zhao 0001
ICDE1
2018 Answering Natural Language Questions by Subgraph Matching over Knowledge Graphs
abstract
RDF question/answering (Q/A) allows users to ask questions in natural languages over a knowledge base represented by RDF. To answer a natural language question, the existing work takes a two-stage approach: question understanding and query evaluation. Their focus is on question understanding to deal with the disambiguation of the natural language phrases. The most common technique is the joint disambiguation, which has the exponential search space. In this paper, we propose a systematic framework to answer natural language questions over RDF repository (RDF Q/A) from a graph data-driven perspective. We propose a semantic query graph to model the query intention in the natural language question in a structural way, based on which, RDF Q/A is reduced to subgraph matching problem. More importantly, we resolve the ambiguity of natural language questions at the time when matches of query are found. The cost of disambiguation is saved if there are no matching found. More specifically, we propose two different frameworks to build the semantic query graph, one is relation (edge)-first and the other one is node-first. We compare our method with some state-of-the-art RDF Q/A systems in the benchmark dataset. Extensive experiments confirm that our method not only improves the precision but also speeds up query performance greatly.
Sen Hu 0005, Lei Zou 0001, Jeffrey Xu Yu, Haixun Wang, Dongyan Zhao 0001
IEEE Trans. Knowl. Data Eng.1