EDBT 2026 Demo / reviewers in the wild / expert
Benfeng Xu
dblp:268/0859
· DBLP profile ↗
18ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-0976-1634ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsabstractThe Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP's growing adoption, existing benchmarks often fail to capture real-world agent performance within this new paradigm, leading to a distorted perception of their true operational value and an inability to reliably differentiate proficiencies. To bridge this critical evaluation gap, we introduce MCP-AgentBench—a comprehensive benchmark specifically engineered to rigorously assess language agent capabilities in MCP-mediated tool interactions. Core contributions of MCP-AgentBench include: the establishment of a robust MCP testbed comprising 33 operational servers with 188 distinct tools; the development of a benchmark featuring 600 systematically designed queries distributed across 6 distinct categories of varying interaction complexity; and the introduction of MCP-Eval, a novel outcome-oriented evaluation methodology prioritizing real-world task success. Through extensive empirical evaluation of leading language agents, we provide foundational insights. MCP-AgentBench aims to equip the research community with a standardized and reliable framework to build, validate, and advance agents capable of fully leveraging MCP's transformative benefits, thereby accelerating progress toward truly capable and interoperable AI systems. Zikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong, Zhendong Mao 0001 |
AAAI | 2 |
| 2026 | FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based AgentsabstractChiwei Zhu, Benfeng Xu, Mingxuan Du, Shaohan Wang, Xiaorui Wang, Zhendong Mao, Yongdong Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chiwei Zhu, Benfeng Xu, Mingxuan Du, Shaohan Wang, Zhendong Mao 0001, Yongdong Zhang 0001 |
ACL (1) | 2 |
| 2026 | GraphSynthQA: Knowledge-Graph-Guided Query Synthesis and Step-Level Preference Optimization for Web AgentsabstractWeb browsing—widely used for information retrieval and fact verification—has become a fundamental capability of recently emerged large language model (LLM) agents, which is often elicited by training on complex questions requiring web search. However, this task faces challenges with respect to data and training: existing QA datasets are mostly 1-3 hop over closed corpora (e.g., Wikipedia); meanwhile, outcome-based on-policy RL that used by recent works is inefficient and brittle in long-horizon, tool-heavy browsing environments. To address these challenges, we introduce GraphSynthQA, a knowledge-graph (KG)—guided synthesis framework in an open-web setting. Starting from Wikidata seed entities, GraphSynthQA iteratively retrieves and verifies evidence from the internet to expand a KG, then synthesizes complex, answer-verifiable queries grounded in multi-evidence dependencies. Building on the synthesized data, we train web-browsing agents with a compute-efficient two-stage recipe: (i) cold-start supervised fine-tuning on ReAct-style trajectories, and (ii) step-level Direct Preference Optimization (DPO), where preferences are constructed offline via single-step branched rollouts that contrast candidate actions by their downstream success rates, providing dense process supervision without expensive on-policy exploration. Experiments show that our approach consistently improves performance on challenging web-browsing benchmarks and remains competitive among models of similar size. Chiwei Zhu, Mingxuan Du, Benfeng Xu, Shengzhuo Zhang, Zhendong Mao 0001 |
SIGIR | 3 |
| 2025 | From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed GroundingabstractThe pursuit of diverse, complex, and largescale instruction data is crucial for automatically aligning large language models (LLMs).While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or rely on trivial extensions that fail to produce meaningful trajectories in terms of complexity.In contrast, instructions that benefit efficient alignment are typically crafted with cognitive insights and grounded in real-world use cases.In this paper, we synthesize such instructions using attributed grounding, which involves 1) a top-down attribution process that grounds a selective set of real instructions to situated users, and 2) a bottom-up synthesis process that leverages web documents to first generate a situation, then a meaningful instruction.This framework allows us to harvest diverse and complex instructions at scale, utilizing the vast range of web documents.Specifically, we construct a dataset of 1 million instructions, called SYNTHQUESTIONS, and demonstrate that models trained on it achieve leading performance on several common benchmarks, with improvements that continually scale with more web corpora.Data, models and codes will be available at https://github. com/Ignoramus0817/SynthQuestions. Chiwei Zhu, Benfeng Xu, Zhendong Mao 0001 |
ACL (1) | 2 |
| 2025 | MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool LearningabstractComplex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting erroneous trajectories in agentic workflows. However, existing approaches only exploit such capability in the post-action stage, where the agent observes the execution outcomes. We argue that, like humans, LLMs can also engage in reflection before action execution: the agent can anticipate undesirable outcomes from its own decisions, which not only provides a necessarily complementary perspective to evaluate the decision but also prevents the propagation of errors throughout the trajectory. In this paper, we propose MIRROR, a framework that consists of both intra-reflection, which critically assesses intended actions before execution, and inter-reflection, which further adjusts the trajectory based on observations. This design systematically leverages LLM reflection capabilities to eliminate and rectify erroneous actions on a more comprehensive scope. Evaluations on both the StableToolBench and TravelPlanner benchmarks demonstrate MIRROR's superior performance, achieving state-of-the-art results compared to existing approaches. Zikang Guo, Benfeng Xu, Zhendong Mao 0001 |
IJCAI | 2 |
| 2025 | PromptMetric: Prompt Recipe as an Automatic Metric for Evaluating Open-domain Question Answering SystemsabstractOpen-domain Question Answering (ODQA) has long been an NLP task receiving wide attention of researchers. Despite being utilized various domains and applications, the evaluation of ODQA systems remains a complicated problem, which is worsened by the wide usage of large language models(LLMs). As LLMs often generate free-form answers that do not follow certain format, traditional string-matching-driven evaluation metrics like Lexical Match are not feasible to accurately reflect the performance of LLM-based ODQA systems. In the meantime, LLM-as-a-Judge methods with simple prompts also display limited consistency with human annotators. To tackle above challenges, we propose a framework of developing effective evaluation prompts based on iterative test and optimization, which can be conducted either by human or LLMs. The resulting prompt, which we refer to as PromptMetric, shows considerable advantages over traditional evaluation methods and LLM evaluation methods with basic prompts. We also demonstrate the robustness of our methods on different models, and show that PromptMetric can be highly economical when applied in ODQA evaluation. Pengzhe Wang, Chiwei Zhu, Benfeng Xu, Zhendong Mao 0001, Yongdong Zhang 0001 |
IJCNN | 4 |
| 2024 | Benchmarking Large Language Models on Controllable Generation under Diversified InstructionsabstractWhile large language models (LLMs) have exhibited impressive instruction-following capabilities, it is still unclear whether and to what extent they can respond to explicit constraints that might be entailed in various instructions. As a significant aspect of LLM alignment, it is thus important to formulate such a specialized set of instructions as well as investigate the resulting behavior of LLMs. To address this vacancy, we propose a new benchmark CoDI-Eval to systematically and comprehensively evaluate LLMs' responses to instructions with various constraints. We construct a large collection of constraints-attributed instructions as a test suite focused on both generalization and coverage. Specifically, we advocate an instruction diversification process to synthesize diverse forms of constraint expression and also deliberate the candidate task taxonomy with even finer-grained sub-categories. Finally, we automate the entire evaluation process to facilitate further developments. Different from existing studies on controllable text generation, CoDI-Eval extends the scope to the prevalent instruction-following paradigm for the first time. We provide extensive evaluations of representative LLMs (e.g., ChatGPT, Vicuna) on CoDI-Eval, revealing their limitations in following instructions with specific constraints and there is still a significant gap between open-source and commercial closed-source LLMs. We believe this benchmark will facilitate research into improving the controllability of LLMs' responses to instructions. Our data and code are available at https://github.com/Xt-cyh/CoDI-Eval. Yihan Chen 0001, Benfeng Xu, Quan Wang 0002, Yi Liu 0148, Zhendong Mao 0001 |
AAAI | 2 |
| 2024 | Disentangled Learning with Synthetic Parallel Data for Text Style TransferabstractText style transfer (TST) is an important task in natural language generation, which aims to transfer the text style (e.g., sentiment) while keeping its semantic information.Due to the absence of parallel datasets for supervision, most existing studies have been conducted in an unsupervised manner, where the generated sentences often suffer from high semantic divergence and thus low semantic preservation.In this paper, we propose a novel disentanglementbased framework for TST named DisenTrans, where disentanglement means that we separate the attribute and content components in the natural language corpus and consider this task from these two perspectives.Concretely, we first create a disentangled Chain-of-Thought prompting procedure to synthesize parallel data and corresponding attribute components for supervision.Then we develop a disentanglement learning method with synthetic data, where two losses are designed to enhance the focus on attribute properties and constrain the semantic space, thereby benefiting style control and semantic preservation respectively.Instructed by the disentanglement concept, our framework creates valuable supervised information and utilizes it effectively in TST tasks.Extensive experiments on mainstream datasets present that our framework achieves significant performance with great sample efficiency. Jingxuan Han, Quan Wang 0002, Zikang Guo, Benfeng Xu, Licheng Zhang 0002, Zhendong Mao 0001 |
ACL (1) | 4 |
| 2024 | KNN-Instruct: Automatic Instruction Construction with K Nearest Neighbor DeductionabstractSupervised fine-tuning (SFT) is a critical procedure for aligning large language models.Despite its efficiency, the construction of SFT data often struggles with issues of quality, diversity, and scalability.Many existing methods, inspired by the SELF-INSTRUCT framework, typically generate synthetic instructions by prompting aligned proprietary models like ChatGPT.However, such process suffers from stale distribution, resulting in instructions that are merely trivial variations of existing ones.In this paper, we introduce a novel bootstrapping approach termed KNN-INSTRUCT, which incorporates KNN deduction to produce meaningful new instructions by effectively summarizing and learning from similar existing ones.We conduct an economical controlled experiment to preliminarily validate its effectiveness.In the further experiment, we construct a high-quality SFT dataset named KNN-INST-12K*.Applying the dataset to Qwen-2-7B, we get a MT-Bench score of 7.64, which outperforms all 7B models on the LMSYS leaderboard, including Starling-LM-7B (7.48), OpenChat-3.5 (7.06) and Zephyr-7B-beta (6.53).Our code and data are available at https://github.com/ CrossmodalGroup/KNN-Instruct/. Jianshang Kou, Benfeng Xu, Chiwei Zhu, Zhendong Mao 0001 |
EMNLP | 2 |
| 2024 | Curriculum Learning Driven Domain Adaptation for Low-Resource Machine Reading ComprehensionabstractAlthough the pre-trained language models have achieved great success on machine reading comprehension task, they often rely on large-scale annotated data, while only a little amount of data is available in the most real-world scenarios. To enhance the PTLMs' capabilities in low-resource scenario, we propose a curriculum learning driven domain adaptation method for low-resource machine reading comprehension, the basic paradigm of which is to train a source model with sufficient data and then adaptive it to our target domain. In the adapting procedure, we introduce the curriculum learning strategy, the core idea of which is arranging training examples from easy to difficult, to bridge the gap between source and target domains and enable the source model adapting to the target domain progressively. Specifically, before fine-tuning the well-trained source model using target data, we firstly calculate the loss of each target example using the source model to evaluating the example difficulty accurately. After that, we sample suitable batches based on an increasing sampling function at each fine-tuning step, allowing the source model to start learning from easy examples in the target domain and gradually transition to difficult ones. Experiments conducted on two public datasets have demonstrated the effectiveness of our method. Licheng Zhang 0002, Quan Wang 0002, Benfeng Xu, Yi Liu 0148, Zhendong Mao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation ExtractionabstractCurrent relation extraction methods suffer from the inadequacy of large-scale annotated data.While distant supervision alleviates the problem of data quantities, there still exists domain disparity in data qualities due to its reliance on domain-restrained knowledge bases. In this work, we propose S2ynRE, a framework of two-stage Self-training with Synthetic data for Relation Extraction.We first leverage the capability of large language models to adapt to the target domain and automatically synthesize large quantities of coherent, realistic training data.We then propose an accompanied two-stage self-training algorithm that iteratively and alternately learns from synthetic and golden data together.We conduct comprehensive experiments and detailed ablations on popular relation extraction datasets to demonstrate the effectiveness of the proposed framework. Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Dai Dai, Yongdong Zhang 0001, Zhendong Mao 0001 |
ACL (1) | 1 |
| 2023 | Modaldrop: Modality-Aware Regularization for Temporal-Spectral Fusion in Human Activity RecognitionabstractAlthough most of existing works for sensor-based Human Activity Recognition rely on the temporal view, we argue that the spectral view also provides complementary prior and accordingly benchmark a standard multi-view framework with extensive experiments to demonstrate its consistent superiority over single-view opponents. We then delve into the intrinsic mechanism of the multi-view representation fusion, and propose ModalDrop as a novel modality-aware regularization method to learn and exploit representations of both views effectively. We demonstrate its advantage over existing representation fusion alternatives with comprehensive experiments and ablations. The improvements are consistent for various settings and are orthogonal with different backbones. We also discuss its potential application for other related tasks regarding representation or modality fusion. The source code is available on https://github.com/studyzx/ModalDrop.git. Yiqiang Chen 0001, Benfeng Xu, Tengxiang Zhang |
ICASSP | 3 |
| 2023 | $k$NN Prompting: Beyond-Context Learning with Calibration-Free Nearest Neighbor Inference
Benfeng Xu, Quan Wang 0002, Zhendong Mao 0001, Yajuan Lyu, Qiaoqiao She, Yongdong Zhang 0001 |
ICLR | 1 |
| 2022 | UniRel: Unified Representation and Interaction for Joint Relational Triple ExtractionabstractRelational triple extraction is challenging for its difficulty in capturing rich correlations between entities and relations.Existing works suffer from 1) heterogeneous representations of entities and relations, and 2) heterogeneous modeling of entity-entity interactions and entity-relation interactions.Therefore, the rich correlations are not fully exploited by existing works.In this paper, we propose UniRel to address these challenges.Specifically, we unify the representations of entities and relations by jointly encoding them within a concatenated natural language sequence, and unify the modeling of interactions with a proposed Interaction Map, which is built upon the off-the-shelf self-attention mechanism within any Transformer block.With comprehensive experiments on two popular relational triple extraction datasets, we demonstrate that UniRel is more effective and computationally efficient.The source code is available at https://github.com/wtangdev/UniRel. Wei Tang 0015, Benfeng Xu, Yuyue Zhao, Zhendong Mao 0001, Yifeng Liu 0002, Yong Liao 0003, Haiyong Xie 0001 |
EMNLP | 2 |
| 2022 | EmRel: Joint Representation of Entities and Embedded Relations for Multi-triple ExtractionabstractBenfeng Xu, Quan Wang, Yajuan Lyu, Yabing Shi, Yong Zhu, Jie Gao, Zhendong Mao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Yabing Shi, Yong Zhu 0004, Zhendong Mao 0001 |
NAACL-HLT | 1 |
| 2021 | Entity Structure Within and Throughout: Modeling Mention Dependencies for Document-Level Relation ExtractionabstractEntities, as the essential elements in relation extraction tasks, exhibit certain structure. In this work, we formulate such entity structure as distinctive dependencies between mention pairs. We then propose SSAN, which incorporates these structural dependencies within the standard self-attention mechanism and throughout the overall encoding stage. Specifically, we design two alternative transformation modules inside each self-attention building block to produce attentive biases so as to adaptively regularize its attention flow. Our experiments demonstrate the usefulness of the proposed entity structure and the effectiveness of SSAN. It significantly outperforms competitive baselines, achieving new state-of-the-art results on three popular document-level relation extraction datasets. We further provide ablation and visualization to show how the entity structure guides the model for better relation extraction. Our code is publicly available. Benfeng Xu, Quan Wang 0002, Yajuan Lyu, Yong Zhu 0004, Zhendong Mao 0001 |
AAAI | 1 |
| 2021 | Review and Arrange: Curriculum Learning for Natural Language UnderstandingabstractWith the notable success of pretrained language models, the pretraining-fine-tuning paradigm has become a dominant solution for natural language understanding (NLU) tasks. Typically, the training instances of a target NLU task are introduced in a completely random order and treated equally at the fine-tuning stage. However, these instances can vary greatly in difficulty, and similar to human learning procedures, language models can benefit from an easy-to-difficult curriculum. Based on this concept, we propose a curriculum learning (CL) framework. Our framework consists of two stages, Review and Arrange, targeting the two main challenges in curriculum learning, i.e., how to define the difficulty of instances and how to arrange a curriculum based on the difficulty, respectively. In the first stage, we devise a cross-review (CR) method to train several teacher models first and then review the training set in a crossed manner to distinguish easy instances from difficult instances. In the second stage, two sampling algorithms, a coarse-grained arrangement (CGA) and a fine-grained arrangement (FGA), are proposed to arrange a curriculum for language models in which the learning materials start from the easiest instances, and more difficult instances are gradually added into the training procedure. Compared to previous heuristic CL methods, our framework can avoid the errors caused by a gap in difficulty between humans and machines and has strong generalization ability. We conduct comprehensive experiments, and the results show that our curriculum learning framework, without any manual model architecture design or use of external data, obtains significant and universal performance improvements on a wide range of NLU tasks in different languages. Licheng Zhang 0002, Zhendong Mao 0001, Benfeng Xu, Quan Wang 0002, Yongdong Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Curriculum Learning for Natural Language UnderstandingabstractWith the great success of pre-trained language models, the pretrain-finetune paradigm now becomes the undoubtedly dominant solution for natural language understanding (NLU) tasks.At the fine-tune stage, target task data is usually introduced in a completely random order and treated equally.However, examples in NLU tasks can vary greatly in difficulty, and similar to human learning procedure, language models can benefit from an easy-to-difficult curriculum.Based on this idea, we propose our Curriculum Learning approach.By reviewing the trainset in a crossed way, we are able to distinguish easy examples from difficult ones, and arrange a curriculum for language models.Without any manual model architecture design or use of external data, our Curriculum Learning approach obtains significant and universal performance improvements on a wide range of NLU tasks. Benfeng Xu, Licheng Zhang 0002, Zhendong Mao 0001, Quan Wang 0002, Hongtao Xie 0001, Yongdong Zhang 0001 |
ACL | 1 |