EDBT 2026 Demo / reviewers in the wild / expert
Cheng Qian 0008
dblp:12/654-8
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0001-9913-820XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ShortageSim: Simulating Drug Shortages Under Information AsymmetryabstractDrug shortages pose critical risks to patient care and healthcare systems worldwide, yet the effectiveness of regulatory interventions remains poorly understood due to information asymmetries in pharmaceutical supply chains. We propose ShortageSim, which addresses this challenge by providing the first simulation framework that evaluates the impact of regulatory interventions on competition dynamics under information asymmetry. Using Large Language Model (LLM)-based agents, the framework models the strategic decisions of drug manufacturers and institutional buyers, in response to shortage alerts given by the regulatory agency. Unlike traditional game theory models that assume perfect rationality and complete information, ShortageSim simulates heterogeneous interpretations on regulatory announcements and the resulting decisions. Experiments on self-processed dataset of historical shortage events show that ShortageSim reduces the resolution lag for production disruption cases by up to 84%, achieving closer alignment to real-world trajectories than the zero-shot baseline. Our framework confirms the effect of regulatory alert in addressing shortages and introduces a new method for understanding competition in multi-stage environments under uncertainty. We open-source ShortageSim and a dataset of 2,925 FDA shortage events, providing a novel framework for future research on policy design and testing in supply chains under information asymmetry. Mingxuan Cui, Yilan Jiang, Duo Zhou, Cheng Qian 0008, Yuji Zhang 0002 |
AAAI | 4 |
| 2026 | PEARL: Self-Evolving Assistant for Time Management with Reinforcement LearningabstractBingxuan Li, Jeonghwan Kim, Cheng Qian, Xiusi Chen, Eitan Anzenberg, Niran Kundapur, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Xiusi Chen, Eitan Anzenberg, Niran Kundapur, Heng Ji 0001 |
ACL (1) | 3 |
| 2026 | From Word to World: Can Large Language Models be Implicit Text-based World Models?abstractYixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang, Cheng Qian, Zeping Li, Xiaoteng Ma, Guanhua Chen, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Hongru Wang 0003, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang 0001, Cheng Qian 0008, Zeping Li, Xiaoteng Ma, Guanhua Chen 0001, Heng Ji 0001 |
ACL (1) | 6 |
| 2026 | CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsabstractJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung 0001 |
ACL (1) | 2 |
| 2026 | Current Agents Fail to Leverage World Model as Tool for ForesightabstractCheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen, Yuji Zhang, Bingxiang He, Qinyu Luo, Gokhan Tur, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Emre Can Acikgoz, Xiusi Chen, Yuji Zhang 0002, Bingxiang He, Qinyu Luo, Gökhan Tür, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji 0001 |
ACL (1) | 1 |
| 2026 | WiNELL: Wikipedia Never-Ending Updating with LLM Agents
Revanth Gangi Reddy, Tanay Dixit, Jiaxin Qin, Cheng Qian 0008, Jiawei Han 0001, Kevin Small, Ruhi Sarikaya, Heng Ji 0001 |
WWW | 4 |
| 2025 | Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHubabstractBohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Cheng Qian, Zihe Wang, Yujia Qin, Yining Ye, Yaxi Lu, Chen Qian, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Bohan Lyu 0001, Xin Cong, Heyang Yu, Pan Yang 0022, Cheng Qian 0008, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang 0004, Yukun Yan, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 5 |
| 2025 | EscapeBench: Towards Advancing Creative Intelligence of Language Model AgentsabstractCheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He, Xiusi Chen, Yuji Zhang, Hongyi Du, Jiarui Yao, Xiaocheng Yang, Denghui Zhang, Yunzhu Li, Heng Ji. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Cheng Qian 0008, Peixuan Han, Qinyu Luo, Bingxiang He, Xiusi Chen, Yuji Zhang 0002, Hongyi Du, Jiarui Yao, Xiaocheng Yang, Yunzhu Li, Heng Ji 0001 |
ACL (1) | 1 |
| 2025 | MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsabstractLarge Language Models (LLMs) have shown remarkable capabilities as autonomous agents; yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario,and cognitive planning improves milestone achievement rates by 3%. Code and datasets are publicavailable at https://github.com/ulab-uiuc/MARBLE. Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhenhailong Wang, Cheng Qian 0008, Robert Tang, Heng Ji 0001, Jiaxuan You |
ACL (1) | 8 |
| 2025 | Aligning LLMs with Individual Preferences via InteractionabstractAs large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption. While previous research focuses on general alignment to principles such as helpfulness, harmlessness, and honesty, the need to account for individual and diverse preferences has been largely overlooked, potentially undermining customized human experiences. To address this gap, we train LLMs that can “interact to align”, essentially cultivating the meta-skill of LLMs to implicitly infer the unspoken personalized preferences of the current user through multi-turn conversations, and then dynamically align their following behaviors and responses to these inferred preferences. Our approach involves establishing a diverse pool of 3,310 distinct user personas by initially creating seed examples, which are then expanded through iterative self-generation and filtering. Guided by distinct user personas, we leverage multi-LLM collaboration to develop a multi-turn preference dataset containing 3K+ multi-turn conversations in tree structures. Finally, we apply supervised fine-tuning and reinforcement learning to enhance LLMs using this dataset. For evaluation, we establish the ALOE (ALign with custOmized prEferences) benchmark, consisting of 100 carefully selected examples and well-designed metrics to measure the customized alignment performance during conversations. Experimental results demonstrate the effectiveness of our method in enabling dynamic, personalized alignment via interaction. The code and dataset will be made public. Shujin Wu, Yi R. Fung 0001, Cheng Qian 0008, Dilek Hakkani-Tür, Heng Ji 0001 |
COLING | 3 |
| 2025 | Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate ContextabstractRushi Wang, Jiateng Liu, Cheng Qian, Yifan Shen, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji, Denghui Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Rushi Wang, Jiateng Liu, Cheng Qian 0008, Yanzhou Pan, Zhaozhuo Xu, Ahmed Abbasi, Heng Ji 0001 |
EMNLP | 3 |
| 2025 | Proactive Agent: Shifting LLM Agents from Reactive Responses to Active AssistanceabstractAgents powered by large language models have shown remarkable abilities in solving complex tasks. However, most agent systems remain reactive, limiting their effectiveness in scenarios requiring foresight and autonomous decision-making. In this paper, we tackle the challenge of developing proactive agents capable of anticipating and initiating tasks without explicit human instructions. We propose a novel data-driven approach for this problem. Firstly, we collect real-world human activities to generate proactive task predictions. These predictions are then labeled by human annotators as either accepted or rejected. The labeled data is used to train a reward model that simulates human judgment and serves as an automatic evaluator of the proactiveness of LLM agents. Building on this, we develop a comprehensive data generation pipeline to create a diverse dataset, ProactiveBench, containing 6,790 events. Finally, we demonstrate that fine-tuning models with the proposed ProactiveBench can significantly elicit the proactiveness of LLM agents. Experimental results show that our fine-tuned model achieves an F1-Score of 66.47% in proactively offering assistance, outperforming all open-source and close-source models. These results highlight the potential of our method in creating more proactive and effective agent systems, paving the way for future advancements in human-agent collaboration. Yaxi Lu, Shenzhi Yang, Cheng Qian 0008, Guirong Chen, Qinyu Luo, Yesai Wu, Xin Cong, Zhong Zhang 0004, Yankai Lin 0001, Weiwen Liu, Yasheng Wang, Zhiyuan Liu 0001, Fangming Liu, Maosong Sun 0001 |
ICLR | 3 |
| 2025 | EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsabstractLeveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluation frameworks. To bridge this gap, we introduce EmbodiedBench, an extensive benchmark designed to evaluate vision-driven embodied agents. EmbodiedBench features: (1) a diverse set of 1,128 testing tasks across four environments, ranging from high-level semantic tasks (e.g., household) to low-level tasks involving atomic actions (e.g., navigation and manipulation); and (2) six meticulously curated subsets evaluating essential agent capabilities like commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning. Through extensive experiments, we evaluated 24 leading proprietary and open-source MLLMs within EmbodiedBench. Our findings reveal that: MLLMs excel at high-level tasks but struggle with low-level manipulation, with the best model, GPT-4o, scoring only $28.9\%$ on average. EmbodiedBench provides a multifaceted standardized evaluation platform that not only highlights existing challenges but also offers valuable insights to advance MLLM-based embodied agents. Our code and dataset are available at [https://embodiedbench.github.io](https://embodiedbench.github.io). Rui Yang 0010, Hanyang Chen, Mark Zhao, Cheng Qian 0008, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, Marziyeh Movahedi, Manling Li, Heng Ji 0001, Huan Zhang 0001, Tong Zhang 0001 |
ICML | 5 |
| 2025 | ToolRL: Reward is All Tool Learning NeedsabstractCurrent Large Language Models (LLMs) often undergo supervised fine-tuning (SFT) to acquire tool use capabilities. However, SFT struggles to generalize to unfamiliar or complex tool use scenarios. Recent advancements in reinforcement learning (RL), particularly with R1-like models, have demonstrated promising reasoning and generalization abilities. Yet, reward design for tool use presents unique challenges: multiple tools may be invoked with diverse parameters, and coarse-grained reward signals, such as answer matching, fail to offer the finegrained feedback required for effective learning.
In this work, we present the first comprehensive study on reward design for tool selection and application tasks within the RL paradigm. We systematically explore a wide range of reward strategies, analyzing their types, scales, granularity, and temporal dynamics. Building on these insights, we propose a principled reward design tailored for tool use tasks and apply it to train LLMs using RL methods.
Empirical evaluations across diverse benchmarks demonstrate that our approach yields robust, scalable, and stable training, achieving a 17\% improvement over base models and a 15\% gain over SFT models. These results highlight the critical role of thoughtful reward design in enhancing the tool use capabilities and generalization performance of LLMs. All the codes are released to facilitate future research. Cheng Qian 0008, Emre Can Acikgoz, Hongru Wang 0011, Xiusi Chen, Dilek Hakkani-Tür, Gökhan Tür, Heng Ji 0001 |
NeurIPS | 1 |
| 2024 | Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven AgentsabstractCheng Qian, Bingxiang He, Zhong Zhuang, Jia Deng, Yujia Qin, Xin Cong, Zhong Zhang, Jie Zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Cheng Qian 0008, Bingxiang He, Zhong Zhuang, Yujia Qin, Xin Cong, Zhong Zhang 0004, Jie Zhou 0016, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 1 |
| 2024 | Toolink: Linking Toolkit Creation and Using through Chain-of-Solving on Open-Source ModelabstractCheng Qian, Chenyan Xiong, Zhenghao Liu, Zhiyuan Liu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Cheng Qian 0008, Chenyan Xiong, Zhenghao Liu 0001, Zhiyuan Liu 0001 |
NAACL-HLT | 1 |
| 2022 | Exploring Mode Connectivity for Pre-trained Language ModelsabstractRecent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP.From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found.Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations.In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path.We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity?(2) How does mode connectivity change during pretraining?(3) How does the PLM's task knowledge change along the path connecting two minima?In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation.The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM. Yujia Qin, Cheng Qian 0008, Jing Yi, Weize Chen, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
EMNLP | 2 |