EDBT 2026 Demo / reviewers in the wild / expert
Yifan Xu 0014
dblp:62/1662-14
· DBLP profile ↗
13ranked-venue papers
1as first author
13since 2021 · last 2026
0009-0009-0188-4075ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic TasksabstractXueqiao Sun, Xiao Liu, Bowen Lv, Hanchen Zhang, Bohao Jing, Zehan Qi, Yifan Xu, Yuxiao Dong, Jie Tang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueqiao Sun, Xiao Liu 0036, Bowen Lv, Hanchen Zhang, Bohao Jing, Zehan Qi, Yifan Xu 0014, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 7 |
| 2025 | AndroidGen: Building an Android Language Agent under Data ScarcityabstractLarge language models have opened up a world of possibilities for various NLP tasks, sparking optimism for the future.Despite their potential, LLMs have yet to be widely used as agents on real mobile devices.The main challenge is the need for high-quality data sources.Time constraints and labor intensity often hinder human annotation.On the other hand, existing LLMs exhibit inadequate completion rates and need a robust data filtration strategy.Given these challenges, we develop a framework called ANDROIDGEN to enhance the capabilities of LLM-based agents under data scarcity.In addition, we leverage AN-DROIDGEN to collect trajectories given human tasks and train open-source LLMs on these trajectories to develop an open-source mobile agent without manually labeled trajectories.We extensively evaluate ANDROIDGEN with AndroidWorld, AitW, and various popular applications, demonstrating its improvements and revealing potential areas for future improvement.Code, model, and data are available at https://github.com/THUDM/AndroidGen. Hanyu Lai, Xiao Liu 0036, Yifan Xu 0014, Shudan Zhang, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 4 |
| 2025 | A Survey of Post-Training Scaling in Large Language ModelsabstractHanyu Lai, Xiao Liu, Junjie Gao, Jiale Cheng, Zehan Qi, Yifan Xu, Shuntian Yao, Dan Zhang, Jinhua Du, Zhenyu Hou, Xin Lv, Minlie Huang, Yuxiao Dong, Jie Tang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hanyu Lai, Xiao Liu 0036, Zehan Qi, Yifan Xu 0014, Shuntian Yao, Jinhua Du, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 6 |
| 2025 | AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsabstractYifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, Yuxiao Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yifan Xu 0014, Xiao Liu 0036, Xueqiao Sun, Hao Yu 0030, Hanyu Lai, Shudan Zhang, Jie Tang 0001, Yuxiao Dong |
ACL (1) | 1 |
| 2025 | VisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsabstractLarge Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}. Xiao Liu 0036, Tianjie Zhang, Yu Gu 0016, Iat Long Iong, Xixuan Song, Yifan Xu 0014, Shudan Zhang, Hanyu Lai, Jiadai Sun, Zehan Qi, Shuntian Yao, Xueqiao Sun, Qinkai Zheng, Hao Yu 0030, Hanchen Zhang, Wenyi Hong, Ming Ding 0004, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su 0001, Yuxiao Dong, Jie Tang 0001 |
ICLR | 6 |
| 2025 | WebGLM: Towards an Efficient and Reliable Web-Enhanced Question-Answering SystemabstractWe present WebGLM, an enhanced Large Language Model (LLM)-based retrieval question-answering system based on the ChatGLM3-6B, offering significant improvements over previous systems. We aim to augment a pre-trained LLM with web search and reliable retrieval capabilities while being efficient for real-world deployments. Leveraging LLM’s in-context learning ability and a robust filter strategy, we create a high-quality training dataset and address the hallucination issue with a self-check mechanism. Our base model, ChatGLM3-6B, excels in extracting critical information and generating desired responses. We tackle the decline in retrieval effectiveness for complex queries with a keywording technique and incorporate more web content for references. We align with user preferences by training a human preference-aware scorer and employing DPO training for direct alignment. Extensive experiments, including human evaluations and the Turing test, demonstrate WebGLM’s superior performance against leading web-enhanced question-answering systems, significantly enhancing performance and efficiency. The code, demo, and data are at https://github.com/THUDM/WebGLM . Hanyu Lai, Xiao Liu 0036, Hao Yu 0030, Yifan Xu 0014, Iat Long Iong, Shuntian Yao, Aohan Zeng, Zhengxiao Du, Yuxiao Dong, Jie Tang 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | AlignBench: Benchmarking Chinese Alignment of Large Language ModelsabstractXiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Andrew Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, Xiaohan Zhang, Lichao Sun, Xiaotao Gu, Hongning Wang, Jing Zhang, Minlie Huang, Yuxiao Dong, Jie Tang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiao Liu 0036, Xuanyu Lei, Shengyuan Wang 0002, Yue Huang 0001, Andrew Feng, Bosi Wen, Pei Ke, Yifan Xu 0014, Weng Lam Tam, Lichao Sun 0001, Xiaotao Gu, Hongning Wang, Jing Zhang 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 9 |
| 2024 | A Cause-Effect Look at Alleviating Hallucination of Knowledge-grounded Dialogue GenerationabstractEmpowered by the large-scale pretrained language models, existing dialogue systems have demonstrated impressive performance conducting fluent and natural-sounding conversations. However, they are still plagued by the <b>hallucination</b> problem, causing unpredictable factual errors in the generated responses. Recently, knowledge-grounded dialogue generation models, that intentionally invoke external knowledge resources to more informative responses, are also proven to be effective in reducing hallucination. Following the idea of getting high-quality knowledge, a few efforts have achieved pretty good performance on this issue. As some inevitable knowledge noises may also lead to hallucinations, it is emergent to investigate the reason and future directions for building noise-tolerant methods in KGD tasks. In this paper, we analyze the causal story behind this problem with counterfactual reasoning methods. Based on the causal effect analysis, we propose a possible solution for alleviating the hallucination in KGD by exploiting the dialogue-knowledge interaction. Experimental results of our example implementation show that this method can reduce hallucination without disrupting other dialogue performance, while keeping adaptive to different generation models. We hope our efforts can support and call for more attention to developing lightweight techniques towards robust and trusty dialogue systems. Jifan Yu, Yifan Xu 0014, Xuanyu Lei, Zijun Yao 0002, Jing Zhang 0001, Lei Hou 0001, Juan-Zi Li |
LREC/COLING | 3 |
| 2024 | AgentBench: Evaluating LLMs as AgentsabstractThe potential of Large Language Model (LLM) as agents has been widely acknowledged recently.
Thus, there is an urgent need to quantitatively evaluate LLMs as agents on challenging tasks in interactive environments.
We present AgentBench, a multi-dimensional benchmark that consists of 8 distinct environments to assess LLM-as-Agent's reasoning and decision-making abilities.
Our extensive test over 29 API-based and open-sourced (OSS) LLMs shows that, while top commercial LLMs present a strong ability of acting as agents in complex environments, there is a significant disparity in performance between them and many OSS competitors that are no larger than 70B.
We identify the typical reasons of failures in environments and LLMs, showing that poor long-term reasoning, decision-making, and instruction following abilities are the main obstacles for developing usable LLM agents.
Improving instruction following and training on high quality multi-round alignment data could improve agent performance.
And different from existing assumptions, training on code present ambivalent impacts on different agent tasks.
Datasets, environments, and an integrated evaluation package for AgentBench are released at https://github.com/THUDM/AgentBench. Xiao Liu 0036, Hao Yu 0030, Hanchen Zhang, Yifan Xu 0014, Xuanyu Lei, Hanyu Lai, Yu Gu 0016, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng 0001, Aohan Zeng, Zhengxiao Du, Sheng Shen 0001, Tianjun Zhang, Yu Su 0001, Huan Sun 0001, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ICLR | 4 |
| 2023 | GOAL: A Challenging Knowledge-grounded Video Captioning Benchmark for Real-time Soccer Commentary GenerationabstractDespite the recent emergence of video captioning models, how to generate vivid, fine-grained video descriptions based on the background knowledge (i.e., long and informative commentary about the domain-specific scenes with appropriate reasoning) is still far from being solved, which however has great applications such as automatic sports narrative. Based on soccer game videos and synchronized commentary data, we present GOAL, a benchmark of over 8.9k soccer video clips, 22k sentences, and 42k knowledge triples for proposing a challenging new task setting as Knowledge-grounded Video Captioning (KGVC). We experimentally test existing state-of-the-art (SOTA) methods on this resource to demonstrate the future directions for improvement in this challenging task. We hope that our data resource (now available at https://github.com/THU-KEG/goal) can serve researchers and developers interested in knowledge-grounded cross-modal applications. Ji Qi 0003, Jifan Yu, Teng Tu 0002, Kunyu Gao, Yifan Xu 0014, Xiaozhi Wang, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li, Jie Tang 0001 |
CIKM | 5 |
| 2023 | GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu 0036, Zhengxiao Du, Hanyu Lai, Ming Ding 0004, Zhuoyi Yang, Yifan Xu 0014, Wendi Zheng, Weng Lam Tam, Zixuan Ma, Jidong Zhai, Zhiyuan Liu 0001, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001 |
ICLR | 8 |
| 2023 | WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human PreferencesabstractWe present WebGLM, a web-enhanced question-answering system based on the General Language Model (GLM). Its goal is to augment a pre-trained large language model (LLM) with web search and retrieval capabilities while being efficient for real-world deployments. To achieve this, we develop WebGLM with strategies for the LLM-augmented retriever, bootstrapped generator, and human preference-aware scorer. Specifically, we identify and address the limitations of WebGPT (OpenAI), through which WebGLM is enabled with accuracy, efficiency, and cost-effectiveness advantages. In addition, we propose systematic criteria for evaluating web-enhanced QA systems. We conduct multi-dimensional human evaluation and quantitative ablation studies, which suggest the outperformance of the proposed WebGLM designs over existing systems. WebGLM with the 10-billion-parameter GLM (10B) is shown to perform better than the similar-sized WebGPT (13B) and even comparably to WebGPT (175B) in human evaluation. The code, demo, and data are at https://github.com/THUDM/WebGLM. Xiao Liu 0036, Hanyu Lai, Hao Yu 0030, Yifan Xu 0014, Aohan Zeng, Zhengxiao Du, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001 |
KDD | 4 |
| 2022 | XDAI: A Tuning-free Framework for Exploiting Pre-trained Language Models in Knowledge Grounded Dialogue GenerationabstractLarge-scale pre-trained language models (PLMs) have shown promising advances on various downstream tasks, among which dialogue is one of the most concerned. However, there remain challenges for individual developers to create a knowledge-grounded dialogue system upon such big models because of the expensive cost of collecting the knowledge resources for supporting the system as well as tuning these large models for the task. To tackle these obstacles, we propose XDAI, a knowledge-grounded dialogue system that is equipped with the prompt-aware tuning-free PLM exploitation and supported by the ready-to-use open-domain external knowledge resources plus the easy-to-change domain-specific mechanism. With XDAI, the developers can leverage the PLMs without any fine-tuning cost to quickly create the open-domain dialogue systems as well as easily customize their own domain-specific systems. Extensive experiments including human evaluation, Turing test, and online evaluation have demonstrated the competitive performance of XDAI compared with the state-of-the-art general PLMs and specific PLMs for dialogue. XDAI pilots studies on the exploitation of PLMs and made intriguing findings which could be inspiring for the future research on other PLM-based applications. Jifan Yu, Yifan Xu 0014, Xuanyu Lei, Jing Zhang 0001, Lei Hou 0001, Juan-Zi Li, Jie Tang 0001 |
KDD | 3 |