EDBT 2026 Demo / reviewers in the wild / expert
Dayuan Fu
dblp:331/3042
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0003-3614-6653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsabstractKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhao Shi, Mohan Jiang, Jie Sun 0030, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Weiye Si, Wenjie Li 0002, Dequan Wang, Pengfei Liu 0003 |
ACL (1) | 7 |
| 2026 | SelfCorrect-Agent: Toward robust and generalizable LLM-based agents
Keqing He 0001, Dayuan Fu, Lele Yang, Weiran Xu |
Neurocomputing | 2 |
| 2026 | Fine-grained inter-series dependency enhanced mining for multi-domain multivariate time series forecastsabstractMulti-domain multivariate time series (MTS) forecasting is increasingly important for large-scale and transferable time-series modeling. However, heterogeneous datasets usually contain different numbers of variables, making scalable inter-series dependency modeling challenging. Existing scalable forecasting paradigms often rely on channel-independent (CI) strategies to accommodate variable-dimensional datasets. Nevertheless, by sharing global parameters across independently processed channels, CI models may confuse heterogeneous dependency structures under multi-domain joint training, leading to the dependency confusion problem. In this paper, we argue that effective multi-domain MTS forecasting requires variable-number-agnostic channel-dependence modeling that can explicitly capture fine-grained inter-series dependencies (FID), including both local inter-series dependencies and cross-temporal inter-series dependencies. To this end, we propose the Fine-grained Inter-series Dependency Enhanced (FIDE) framework, a plug-and-play module for patch-based CI forecasters. FIDE introduces dependency prototypes to dynamically perceive input-specific inter-series dependency patterns and employs dependency transmission to propagate dependency information across temporal segments while remaining agnostic to the number of variables. Extensive experiments on eight real-world benchmarks with six representative CI forecasting models demonstrate that FIDE consistently improves forecasting accuracy, effectively mitigates dependency confusion, and achieves statistically significant gains over baseline methods. Qi Li 0053, Tianmu Sha, Zhenyu Zhang 0032, Xiaolei Hua, Dayuan Fu, Yinglei Teng, Yong Zhang 0025 |
Neurocomputing | 6 |
| 2025 | PreAct: Prediction Enhances Agent's Planning AbilityabstractAddressing the disparity between predictions and actual results can enable individuals to expand their thought processes and stimulate self-reflection, thus promoting accurate planning. In this research, we present PreAct, an agent framework that integrates prediction, reasoning, and action. By utilizing the information derived from predictions, the large language model (LLM) agent can provide a wider range and more strategically focused reasoning. This leads to more efficient actions that aid the agent in accomplishing intricate tasks. Our experimental results show that PreAct surpasses the ReAct method in completing complex tasks and that PreAct’s performance can be further improved when paired with other memory or selection strategy techniques. We presented the model with varying quantities of historical predictions and discovered that these predictions consistently enhance LLM planning. The variances in single-step reasoning between PreAct and ReAct indicate that PreAct indeed has benefits in terms of diversity and strategic orientation over ReAct. Dayuan Fu, Jianzhao Huang, Guanting Dong 0001, Yejie Wang, Keqing He 0001, Weiran Xu |
COLING | 1 |
| 2025 | DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsabstractLarge Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods-brittle prompt engineering or RAG-based reinforcement learning in controlled environments-fail to capture real-world complexities.In this paper, we introduce DeepResearcher, the first comprehensive framework for end-to-end training of LLM-based deep research agents through scaling reinforcement learning (RL) in real-world environments with authentic web search interactions.Unlike RAG approaches reliant on fixed corpora, DeepResearcher trains agents to navigate the noisy, dynamic open web.We implement a specialized multi-agent architecture where browsing agents extract relevant information from various webpage structures and overcoming significant technical challenges.Extensive experiments on open-domain research tasks demonstrate that DeepResearcher achieves substantial improvements of up to 28.9 points over prompt engineering-based baselines and up to 7.2 points over RAG-based RL agents.Our qualitative analysis reveals emergent cognitive behaviors from end-to-end RL training, such as planning, cross-validation, self-reflection for research redirection, and maintain honesty when unable to find definitive answers.Our results highlight that end-to-end training in realworld web environments is fundamental for developing robust research capabilities aligned with real-world applications.The source code for DeepResearcher is released at: https:// github.com/GAIR-NLP/DeepResearcher. Dayuan Fu, Xiangkun Hu, Xiaojie Cai, Lyumanshan Ye, Pengrui Lu, Pengfei Liu 0003 |
EMNLP | 2 |
| 2025 | AgentRefine: Enhancing Agent Generalization through Refinement TuningabstractLarge Language Model (LLM) based agents have proved their ability to perform complex tasks like humans. However, there is still a large gap between open-sourced LLMs and commercial models like the GPT series. In this paper, we focus on improving the agent generalization capabilities of LLMs via instruction tuning. We first observe that the existing agent training corpus exhibits satisfactory results on held-in evaluation sets but fails to generalize to held-out sets. These agent-tuning works face severe formatting errors and are frequently stuck in the same mistake for a long while. We analyze that the poor generalization ability comes from overfitting to several manual agent environments and a lack of adaptation to new situations. They struggle with the wrong action steps and can not learn from the experience but just memorize existing observation-action relations. Inspired by the insight, we propose a novel AgentRefine framework for agent-tuning. The core idea is to enable the model to learn to correct its mistakes via observation in the trajectory. Specifically, we propose an agent synthesis framework to encompass a diverse array of environments and tasks and prompt a strong LLM to refine its error action according to the environment feedback. AgentRefine significantly outperforms state-of-the-art agent-tuning work in terms of generalization ability on diverse agent tasks. It also has better robustness facing perturbation and can generate diversified thought in inference. Our findings establish the correlation between agent generalization and self-refinement and provide a new paradigm for future research. Dayuan Fu, Keqing He 0001, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng 0003, Jingang Wang, Weiran Xu |
ICLR | 1 |
| 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryabstractLarge language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round evaluation of the computer science field. To bridge this gap, we introduce CS-Bench, the first multilingual (English, Chinese, French, German) benchmark dedicated to evaluating the performance of LLMs in computer science. CS-Bench comprises approximately 10K meticulously curated test samples, covering 26 subfields across 4 key areas of computer science, encompassing various task forms and divisions of knowledge and reasoning. Utilizing CS-Bench, we conduct a comprehensive evaluation of over 30 mainstream LLMs, revealing the relationship between CS performance and model scales. We also quantitatively analyze the reasons for failures in existing LLMs and highlight directions for improvements, including knowledge supplementation and CS-specific reasoning. Further cross-capability experiments show a high correlation between LLMs' capabilities in computer science and their abilities in mathematics and coding. Moreover, expert LLMs specialized in mathematics and coding also demonstrate strong performances in several CS subfields. Looking ahead, we envision CS-Bench serving as a cornerstone for LLM applications in the CS field and paving new avenues in assessing LLMs' diverse reasoning capabilities. Our project homepage is available at https://csbench.github.io/. Xiaoshuai Song, Muxi Diao, Guanting Dong 0001, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Weihao Zeng 0003, Yejie Wang, Zhuoma Gongque, Jianing Yu 0001, Qiuna Tan, Weiran Xu |
ICLR | 8 |
| 2025 | FuseMind: Fusing reflection and prediction elevates agent's reasoning capabilities
Xiufa Ma, Xinxin Ge, Heyang Xu, Dayuan Fu, Zhexu Wang, Keqing He 0001, Weiran Xu |
Neurocomputing | 4 |
| 2024 | BootTOD: Bootstrap Task-oriented Dialogue Representations by Aligning Diverse ResponsesabstractPre-trained language models have been successful in many scenarios. However, their usefulness in task-oriented dialogues is limited due to the intrinsic linguistic differences between general text and task-oriented dialogues. Current task-oriented dialogue pre-training methods rely on a contrastive framework, which faces challenges such as selecting true positives and hard negatives, as well as lacking diversity. In this paper, we propose a novel dialogue pre-training model called BootTOD. It learns task-oriented dialogue representations via a self-bootstrapping framework. Unlike contrastive counterparts, BootTOD aligns context and context+response representations and dismisses the requirements of contrastive pairs. BootTOD also uses multiple appropriate response targets to model the intrinsic one-to-many diversity of human conversations. Experimental results show that BootTOD outperforms strong TOD baselines on diverse downstream dialogue tasks. Weihao Zeng 0003, Keqing He 0001, Yejie Wang, Dayuan Fu, Weiran Xu |
LREC/COLING | 4 |
| 2024 | MSI-Agent: Incorporating Multi-Scale Insight into Embodied Agents for Superior Planning and Decision-MakingabstractLong-term memory is significant for agents, in which insights play a crucial role.However, the emergence of irrelevant insight and the lack of general insight can greatly undermine the effectiveness of insight.To solve this problem, in this paper, we introduce Multi-Scale Insight Agent (MSI-Agent), an embodied agent designed to improve LLMs' planning and decision-making ability by summarizing and utilizing insight effectively across different scales.MSI achieves this through the experience selector, insight generator, and insight selector.Leveraging a three-part pipeline, MSI can generate task-specific and high-level insight, store it in a database, and then use relevant insight from it to aid in decisionmaking.Our experiments show that MSI outperforms another insight strategy when planning by GPT3.5.Moreover, We delve into the strategies for selecting seed experience and insight, aiming to provide LLM with more useful and relevant insight for better decision-making.Our observations also indicate that MSI exhibits better robustness when facing domainshifting scenarios. Dayuan Fu, Biqing Qi, Yihuai Gao, Che Jiang, Guanting Dong 0001, Bowen Zhou 0002 |
EMNLP | 1 |
| 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good DataabstractYejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu |
EMNLP | 3 |
| 2024 | On Large Language Models' Hallucination with Regard to Known FactsabstractChe Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, Jie Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Fandong Meng, Mo Yu, Bowen Zhou 0002, Jie Zhou 0016 |
NAACL-HLT | 4 |
| 2023 | A Multi-Task Semantic Decomposition Framework with Task-specific Pre-training for Few-Shot NERabstractThe objective of few-shot named entity recognition is to identify named entities with limited labeled instances. Previous works have primarily focused on optimizing the traditional token-wise classification framework, while neglecting the exploration of information based on NER data characteristics. To address this issue, we propose a Multi-Task Semantic Decomposition Framework via Joint Task-specific Pre-training (MSDP) for few-shot NER. Drawing inspiration from demonstration-based and contrastive learning, we introduce two novel pre-training tasks: Demonstration-based Masked Language Modeling (MLM) and Class Contrastive Discrimination. These tasks effectively incorporate entity boundary information and enhance entity representation in Pre-trained Language Models (PLMs). In the downstream main task, we introduce a multi-task joint optimization framework with the semantic decomposing method, which facilitates the model to integrate two different semantic information for entity classification. Experimental results of two few-shot NER benchmarks demonstrate that MSDP consistently outperforms strong baselines by a large margin. Extensive analyses validate the effectiveness and generalization of MSDP. Guanting Dong 0001, Zechen Wang, Jinxu Zhao, Daichi Guo, Dayuan Fu, Tingfeng Hui, Keqing He 0001, Xuefeng Li 0002, Liwen Wang 0007, Weiran Xu |
CIKM | 6 |
| 2023 | A Prototypical Semantic Decoupling Method via Joint Contrastive Learning for Few-Shot Named Entity RecognitionabstractFew-shot named entity recognition (NER) aims at identifying named entities based on only few labeled instances. Most existing prototype-based sequence labeling models tend to memorize entity mentions which would be easily confused by close prototypes. In this paper, we proposed a Prototypical Semantic Decoupling method via joint Contrastive learning (PSDC) for few-shot NER. Specifically, we decouple class-specific prototypes and contextual semantic prototypes by two masking strategies to lead the model to focus on two different semantic information for inference. Besides, we further introduce joint contrastive learning objectives to better integrate two kinds of decoupling information and prevent semantic collapse. Experimental results on two few-shot NER benchmarks demonstrate that PSDC consistently outperforms the previous SOTA methods in terms of overall performance. Extensive analysis further validates the effectiveness and generalization of PSDC. Guanting Dong 0001, Zechen Wang, Liwen Wang 0007, Daichi Guo, Dayuan Fu, Yuxiang Wu, Xuefeng Li 0002, Tingfeng Hui, Keqing He 0001, QiXiang Gao, Weiran Xu |
ICASSP | 5 |
| 2023 | Revisit Out-Of-Vocabulary Problem For Slot Filling: A Unified Contrastive Framework With Multi-Level Data AugmentationsabstractIn real dialogue scenarios, the existing slot filling model, which tends to memorize entity patterns, has a significantly reduced generalization facing Out-of-Vocabulary (OOV) problems. To address this issue, we propose an OOV robust slot filling model based on multi-level data augmentations to solve the OOV problem from both word and slot perspectives. We present a unified contrastive learning framework, which pull representations of the origin sample and augmentation samples together, to make the model resistant to OOV problems. We evaluate the performance of the model from some specific slots and carefully design test data with OOV word perturbation to further demonstrate the effectiveness of OOV words. Experiments on two datasets show that our approach outperforms the previous sota methods in terms of both OOV slots and words. Daichi Guo, Guanting Dong 0001, Dayuan Fu, Yuxiang Wu, Tingfeng Hui, Liwen Wang 0007, Xuefeng Li 0002, Zechen Wang, Keqing He 0001, Weiran Xu |
ICASSP | 3 |