VLDB 2026 Research / reviewers in the wild / expert
Zehan Qi
dblp:358/8851
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0007-5232-9130ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 46% Trustworthy machine learning · 30% Reinforcement learning · 10% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
1.6 | 2 | 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning · ICLR 2025 Knowledge Conflicts for LLMs: A Survey · EMNLP 2024 |
Machine learning › Trustworthy machine learning
fairness |
1.5 | 2 | 2024 | Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024 Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias · EMNLP 2024 |
Machine learning › Trustworthy machine learning
robustness |
1.5 | 2 | 2024 | Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias · EMNLP 2024 Knowledge Conflicts for LLMs: A Survey · EMNLP 2024 |
Machine learning › Reinforcement learning
LLM agent training |
1.0 | 1 | 2026 | KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic Tasks · ACL (1) 2026 |
Machine learning › Reinforcement learning
curriculum reinforcement learning |
0.9 | 1 | 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning · ICLR 2025 |
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent |
0.9 | 1 | 2025 | VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.9 | 1 | 2025 | A Survey of Post-Training Scaling in Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation › LLM agents
web agents |
0.9 | 1 | 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning · ICLR 2025 |
Natural language and speech › Language models and text generation › LLM agents › web agents
web agent training |
0.9 | 1 | 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
0.8 | 1 | 2024 | Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias · EMNLP 2024 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.8 | 1 | 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict |
0.8 | 1 | 2024 | Knowledge Conflicts for LLMs: A Survey · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
metareasoning |
0.8 | 1 | 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
parametric vs contextual knowledge |
0.8 | 1 | 2024 | Knowledge Conflicts for LLMs: A Survey · EMNLP 2024 |
Natural language and speech › Language models and text generation › large language model evaluation
reasoning benchmark |
0.8 | 1 | 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
stereotype bias evaluation |
0.8 | 1 | 2024 | Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
toxicity reduction |
0.8 | 1 | 2024 | Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias · EMNLP 2024 |
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
agentic reasoning |
0.3 | 1 | 2026 | KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic Tasks · ACL (1) 2026 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.3 | 1 | 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning · ICLR 2025 |
Human-AI interaction
GUI agent |
0.3 | 1 | 2025 | VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents · ICLR 2025 |
Natural language and speech › Language models and text generation
alignment |
0.2 | 1 | 2024 | Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency · NeurIPS 2024 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2024 | Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias · EMNLP 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.9program-based solver · 1.7human demonstration · 1.7agent bootstrapping · 1.7survey · 1.6self-evolving curriculum · 0.9statistical framework · 0.8self-correction · 0.8perspective-taking prompting · 0.8bias-variance decomposition · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic TasksabstractXueqiao Sun, Xiao Liu, Bowen Lv, Hanchen Zhang, Bohao Jing, Zehan Qi, Yifan Xu, Yuxiao Dong, Jie Tang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueqiao Sun, Xiao Liu 0036, Bowen Lv, Hanchen Zhang, Bohao Jing, Zehan Qi, Yifan Xu 0014, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 6 |
| 2025 | A Survey of Post-Training Scaling in Large Language ModelsabstractHanyu Lai, Xiao Liu, Junjie Gao, Jiale Cheng, Zehan Qi, Yifan Xu, Shuntian Yao, Dan Zhang, Jinhua Du, Zhenyu Hou, Xin Lv, Minlie Huang, Yuxiao Dong, Jie Tang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hanyu Lai, Xiao Liu 0036, Zehan Qi, Yifan Xu 0014, Shuntian Yao, Jinhua Du, Minlie Huang, Yuxiao Dong, Jie Tang 0001 |
ACL (1) | 5 |
| 2025 | VisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsabstractLarge Multimodal Models (LMMs) have ushered in a new era in artificial intelligence, merging capabilities in both language and vision to form highly capable \textbf{Visual Foundation Agents} that are postulated to excel across a myriad of tasks. However, existing benchmarks fail to sufficiently challenge or showcase the full potential of LMMs as visual foundation agents in complex, real-world environments. To address this gap, we introduce VisualAgentBench (VAB), a comprehensive and unified benchmark specifically designed to train and evaluate LMMs as visual foundation agents across diverse scenarios in one standard setting, including Embodied, Graphical User Interface, and Visual Design, with tasks formulated to probe the depth of LMMs' understanding and interaction capabilities. Through rigorous testing across 9 proprietary LMM APIs and 9 open models (18 in total), we demonstrate the considerable yet still developing visual agent capabilities of these models. Additionally, VAB explores the synthesizing of visual agent trajectory data through hybrid methods including Program-based Solvers, LMM Agent Bootstrapping, and Human Demonstrations, offering insights into obstacles, solutions, and trade-offs one may meet in developing open LMM agents. Our work not only aims to benchmark existing models but also provides an instrumental playground for future development into visual foundation agents. Code, train, and test data are available at \url{https://github.com/THUDM/VisualAgentBench}. Xiao Liu 0036, Tianjie Zhang, Yu Gu 0016, Iat Long Iong, Xixuan Song, Yifan Xu 0014, Shudan Zhang, Hanyu Lai, Jiadai Sun, Zehan Qi, Shuntian Yao, Xueqiao Sun, Qinkai Zheng, Hao Yu 0030, Hanchen Zhang, Wenyi Hong, Ming Ding 0004, Lihang Pan, Xiaotao Gu, Aohan Zeng, Zhengxiao Du, Chan Hee Song, Yu Su 0001, Yuxiao Dong, Jie Tang 0001 |
ICLR | 12 |
| 2025 | WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningabstractLarge language models (LLMs) have shown remarkable potential as autonomous agents, particularly in web-based tasks.
However, existing LLM web agents face significant limitations: high-performing agents rely on expensive proprietary LLM APIs, while open LLMs lack the necessary decision-making capabilities.
This paper introduces WebRL, a novel self-evolving online curriculum reinforcement learning framework designed to train high-performance web agents using open LLMs.
Our approach addresses key challenges in this domain, including the scarcity of training tasks, sparse feedback signals, and policy distribution drift in online learning.
WebRL incorporates a self-evolving curriculum that generates new tasks from unsuccessful attempts, a robust outcome-supervised reward model (ORM), and adaptive reinforcement learning strategies to ensure consistent improvement.
We apply WebRL to transform Llama-3.1 models into proficient web agents, achieving remarkable results on the WebArena-Lite benchmark.
Our Llama-3.1-8B agent improves from an initial 4.8\% success rate to 42.4\%, while the Llama-3.1-70B agent achieves a 47.3\% success rate across five diverse websites.
These results surpass the performance of GPT-4-Turbo (17.6\%) by over 160\% relatively and significantly outperform previous state-of-the-art web agents trained on open LLMs (AutoWebGLM, 18.2\%).
Our findings demonstrate WebRL's effectiveness in bridging the gap between open and proprietary LLM-based web agents, paving the way for more accessible and powerful autonomous web interaction systems. Zehan Qi, Xiao Liu 0036, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Jiadai Sun, Shuntian Yao, Wei Xu 0017, Jie Tang 0001, Yuxiao Dong |
ICLR | 1 |
| 2024 | Knowledge Conflicts for LLMs: A SurveyabstractThis survey provides an in-depth analysis of knowledge conflicts for large language models (LLMs), highlighting the complex challenges they encounter when blending contextual and parametric knowledge.Our focus is on three categories of knowledge conflicts: contextmemory, inter-context, and intra-memory conflict.These conflicts can significantly impact the trustworthiness and performance of LLMs, especially in real-world applications where noise and misinformation are common.By categorizing these conflicts, exploring the causes, examining the behaviors of LLMs under such conflicts, and reviewing available solutions, this survey aims to shed light on strategies for improving the robustness of LLMs, thereby serving as a valuable resource for advancing research in this evolving area. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang 0003, Yue Zhang 0004, Wei Xu 0039 |
EMNLP | 2 |
| 2024 | Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and BiasabstractThe common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm.Present solutions often demand whitebox access to the model or substantial training, which is impractical for cutting-edge commercial LLMs.Moreover, prevailing prompting methods depend on external tool feedback and fail to simultaneously lessen toxicity and bias.Motivated by social psychology principles, we propose a novel strategy named perspective-taking prompting (PET) that inspires LLMs to integrate diverse human perspectives and self-regulate their responses.This self-correction mechanism can significantly diminish toxicity (up to 89%) and bias (up to 73%) in LLMs' responses.Rigorous evaluations and ablation studies are conducted on two commercial LLMs (ChatGPT and GLM) and three open-source LLMs, revealing PET's superiority in producing less harmful responses, outperforming five strong baselines."Words kill, words give life; they're either poison or fruit-you choose."~Proverbs 18:21 (MSG) Rongwu Xu, Zi'an Zhou, Tianwei Zhang 0004, Zehan Qi, Su Yao, Ke Xu 0002, Wei Xu 0039, Han Qiu 0001 |
EMNLP | 4 |
| 2024 | Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation InconsistencyabstractWe present a novel statistical framework for analyzing stereotypes in large language models (LLMs) by systematically estimating the bias and variation in their generation. Current evaluation metrics in the alignment literature often overlook the randomness of stereotypes caused by the inconsistent generative behavior of LLMs. For example, this inconsistency can result in LLMs displaying contradictory stereotypes, including those related to gender or race, for identical professions across varied contexts. Neglecting such inconsistency could lead to misleading conclusions in alignment evaluations and hinder the accurate assessment of the risk of LLM applications perpetuating or amplifying social stereotypes and unfairness.This work proposes a Bias-Volatility Framework (BVF) that estimates the probability distribution function of LLM stereotypes. Specifically, since the stereotype distribution fully captures an LLM's generation variation, BVF enables the assessment of both the likelihood and extent to which its outputs are against vulnerable groups, thereby allowing for the quantification of the LLM's aggregated discrimination risk. Furthermore, we introduce a mathematical framework to decompose an LLM’s aggregated discrimination risk into two components: bias risk and volatility risk, originating from the mean and variation of LLM’s stereotype distribution, respectively. We apply BVF to assess 12 commonly adopted LLMs and compare their risk levels. Our findings reveal that: i) Bias risk is the primary cause of discrimination risk in LLMs; ii) Most LLMs exhibit significant pro-male stereotypes for nearly all careers; iii) Alignment with reinforcement learning from human feedback lowers discrimination by reducing bias, but increases volatility; iv) Discrimination risk in LLMs correlates with key sociol-economic factors like professional salaries. Finally, we emphasize that BVF can also be used to assess other dimensions of generation inconsistency's impact on LLM behavior beyond stereotypes, such as knowledge mastery. Ke Yang 0003, Zehan Qi, Yang Yu 0011, ChengXiang Zhai |
NeurIPS | 3 |
| 2024 | MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsabstractLarge language models (LLMs) have shown increasing capability in problem-solving and decision-making, largely based on the step-by-step chain-of-thought reasoning processes. However, evaluating these reasoning abilities has become increasingly challenging. Existing outcome-based benchmarks are beginning to saturate, becoming less effective in tracking meaningful progress. To address this, we present a process-based benchmark MR-Ben that demands a meta-reasoning skill, where LMs are asked to locate and analyse potential errors in automatically generated reasoning steps. Our meta-reasoning paradigm is especially suited for system-2 slow thinking, mirroring the human cognitive process of carefully examining assumptions, conditions, calculations, and logic to identify mistakes. MR-Ben comprises 5,975 questions curated by human experts across a wide range of subjects, including physics, chemistry, logic, coding, and more. Through our designed metrics for assessing meta-reasoning on this benchmark, we identify interesting limitations and weaknesses of current LLMs (open-source and closed-source models). For example, with models like the o1 series from OpenAI demonstrating strong performance by effectively scrutinizing the solution space, many other state-of-the-art models fall significantly behind on MR-Ben, exposing potential shortcomings in their training strategies and inference methodologies. Zhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li 0001, Pengguang Chen, Jianbo Dai, Rongwu Xu, Zehan Qi, Wanru Zhao, Linling Shen, Jianqiao Lu, Haochen Tan, Yukang Chen, Bailin Wang, Zhijiang Guo, Jiaya Jia |
NeurIPS | 9 |