VLDB 2026 Research / reviewers in the wild / expert
Yuanzhao Zhai
dblp:296/9009
· DBLP profile ↗
19ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0003-1385-0074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis GenerationabstractXiaoying Le, Pengfei Qian, Yuanzhao Zhai, Xu Zhang, Qian Liu, Feng Dawei, Bo Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoying Le, Pengfei Qian, Yuanzhao Zhai, Bo Ding 0001 |
ACL (1) | 3 |
| 2026 | DART: Empowering LLM-Based Agents on Domain-Specific Tasks through Dynamic Augmented Reasoning Trees
Genrui Zhang, Yuanzhao Zhai, Guangyuan Qian |
ICIC (23) | 2 |
| 2026 | Diagnosing LLM Benchmark: A Psychometric Analysis of Difficulty and DiscriminationabstractBenchmarks have established themselves as the standard for evaluating and tracking the progress of Large Language Models. However, the community increasingly faces inconsistent model rankings across nominally similar tasks, suggesting structural misalignments in evaluation instruments. Current paradigms, which rely heavily on aggregate scalar scores, often treat benchmarks as black boxes and fail to diagnose the root causes of these discrepancies. In this work, we propose a structural diagnostic framework grounded in Item Response Theory to evaluate the benchmarks themselves. By mapping items into a latent skill space, we assess benchmark quality along two critical dimensions: calibration, which measures the alignment of difficulty with model capabilities, and efficiency, which evaluates the discriminative power of items. Our analysis of two math benchmarks yields two findings. First, distinct difficulty topologies quantitatively explain why models exhibit conflicting rankings. Second, an observed negative exponential saturation pattern shows that larger item counts can yield diminishing reliability gains. These findings show that quantity does not equal quality, providing a principled roadmap for constructing leaner, more rigorous benchmarks by resolving difficulty mismatches and pruning redundant dimensions. Jiacheng Qin, Bo Ding 0001, Yuanzhao Zhai |
ICPR (6) | 5 |
| 2026 | Uncertainty-penalized reinforcement learning from human feedback with diversified reward LoRA ensembles
Yuanzhao Zhai, Han Zhang 0025, Yue Yu 0001, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
Inf. Process. Manag. | 1 |
| 2025 | Enhancing Decision-Making for LLM Agents via Step-Level Q-Value ModelsabstractAgents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific tasks is difficult when intermediate actions are neither appropriately rewarded nor penalized. In this paper, we propose leveraging a task-relevant Q-value model to guide action selection. Specifically, we first collect decision-making trajectories annotated with step-level Q values via Monte Carlo Tree Search (MCTS) and construct preference data. We then use another LLM to fit these preferences through step-level Direct Policy Optimization (DPO), which serves as the Q-value model. During inference, at each decision-making step, LLM agents select the action with the highest Q value before interacting with the environment. We apply our method to various open-source and API-based LLM agents, demonstrating that Q-value models significantly improve their performance. Notably, the performance of the agent built with Phi-3-mini-4k-instruct improved by 103% on WebShop and 75% on HotPotQA when enhanced with Q-value models, even surpassing GPT-4o-mini. Additionally, Q-value models offer several advantages, such as generalization to different LLM agents and seamless integration with existing prompting strategies. Yuanzhao Zhai, Tingkai Yang, Kele Xu, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
AAAI | 1 |
| 2025 | Correcting Large Language Model Behavior via Influence FunctionabstractRecent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate from contemporary human preferences and societal norms. Existing methodologies, either curation of new data for continual alignment or manual correction of outdated data for re-alignment, demand costly human resources. To address this, we propose a novel approach, LLM BehAvior Correction with INfluence FunCtion REcall and Post-Training (LANCET), which needs no human involvement. LANCET consists of two phases: (1) using a new method LinFAC to efficiently identify the training data that significantly impact undesirable model outputs, and (2) applying an novel Influence-driven Bregman Optimization (IBO) technique to adjust the model’s outputs based on these influence distributions. Our experiments show that LANCET effectively and efficiently corrects inappropriate behaviors of LLMs while preserving model utility. Further more, LANCET exhibits stronger generalization ability than all baselines under out-of-distribution harmful prompts, offering better interpretability and compatibility with real-world applications of LLMs. Han Zhang 0025, Zhuo Zhang 0007, Yi Zhang 0127, Yuanzhao Zhai, Hanyang Peng, Yue Yu 0001, Hui Wang 0030, Bin Liang 0004, Lin Gui 0003, Ruifeng Xu 0001 |
AAAI | 4 |
| 2025 | GRACE: Graph-Adapted Case-Augmented Execution for Tool Use in Large Language ModelsabstractLarge Language Models (LLMs) have shown remarkable capability. However, their static parametric memory inherently impedes alignment with rapidly evolving real-world tools and APIs. Existing tool-augmented LLMs, although effective, frequently fail to identify accurate and executable tool combinations within an ever-expanding and dynamically changing toolbox, leading to cascading failures and hallucinations. These limitations are further exacerbated by semantic misalignment between natural-language queries and symbolic tool specifications. In this paper, we present Graph-Adapted Case-Augmented Execution (GRACE), a lightweight yet effective framework that equips LLMs with precise, scalable, and continually adaptive tooluse capabilities. GRACE introduces a demand-driven retrieval mechanism in which the LLM first composes a structured demand vector that explicitly aligns natural-language intent with the symbolic vocabulary of tool APIs, thereby narrowing the semantic gap and enabling markedly more precise retrieval. The demand vector is scored against every node of the tool graph; the LLM then selects the highest-ranked tool together with its dependency-closed subgraph, ensuring that all prerequisite constraints are satisfied and every retrieved tool can be executed immediately. Additionally, GRACE maintains a case library that stores execution traces and enables long-term memory without model retraining. Whenever tools are added, deprecated, or modified, both the graph and case library are updated, preserving long-term memory alignment with the evolving external ecosystem. Extensive experiments on ToolSandbox and ToolQA-D demonstrate that GRACE consistently surpasses both proprietary prompting baselines and fine-tuned models, achieving 76.1 % exact-match accuracy on ToolSandbox and 70.7 % under the dynamic tool environment on ToolQA-D, thereby establishing new state-of-the-art results while maintaining parameter efficiency and plug-and-play deployment. Yuanzhao Zhai, Bo Ding 0001 |
ICPADS | 2 |
| 2025 | Extracting Reasoning Patterns from Knowledge Graph to Enhance LLMs' Reasoning CapabilityabstractLarge language models (LLMs) have demonstrated great potential across diverse fields. However, their reasoning capabilities face challenges, especially when dealing with complex tasks. Existing work on enhancing LLMs' reasoning abilities mostly focuses on fields like mathematics and coding. Due to the specificity of domain data, these methods have poor generalization in specific areas such as medicine and materials. In this work, we propose an approach to enhance LLMs' reasoning capabilities in these fields lacking reasoning datasets. We construct reasoning datasets based on the logical relationships contained in their knowledge graphs. First, we clean the metadata, focus on a specific sub-field, and optimize data quality to construct a specialized knowledge graph. Then, we design seven knowledge graph patterns to extract instances, transform them into natural-language questions, and sample using the Deepseek-R1 model to build a reasoning dataset. We fine-tune the Qwen2.5-32B model with this dataset. Experimental results show that in the Relevance dimension, the score of Qwen2.5-32B improves from 45 to 98.57, with a remarkable increase of about 119.04%. In the Logic dimension, the score rises from 76.87 to 87.2, representing an increase of approximately 13.44%. This validates the effectiveness of our method. In these two aspects, the fine-tuned model reaches or approaches the level of advanced LLMs such as Deepseek-R1 and OpenAI 01, indicating that our approach can effectively enhance LLMs' reasoning ability. Tingkai Yang, Yuanzhao Zhai, Huanxi Liu, Huaimin Wang 0001 |
JCC | 2 |
| 2025 | Preference-Strength-Aware Self-Improving Alignment with Generative Preference ModelsabstractSelf-improving alignment leveraging large language models (LLMs) to automatically generate synthetic preference data has garnered significant attention as a means of reducing reliance on human labelers. These methods typically employ the LLM-as-a-judge mechanism, where the LLM generates responses and then employs itself to judge which response best aligns with the given prompt for curating the binary self-preferred dataset. However, these methods encounter two major challenges: (1) LLM-as-a-judge often produces error-prone evaluations, resulting in low-quality preference annotation, and (2) their optimization strategies often overlook the strength of preferences within binary pairs, leading to overfitting. This paper proposes a novel method, Preference-Strength-aware Optimization (PSO), to address these issues. Specifically, PSO frames the preference annotation process as a judgment token prediction task using the generative preference model to produce reliable judgments. The predicted judgment token indicates the preferred response and its corresponding probability reflects the disparity between responses, referred to as preference strength. Based on this strength, we introduce a new preference-strength-aware loss to adaptively reweight the impact of different response pairs on optimization, concentrating the model's learning on high-quality response pairs. Our experiments demonstrate that PSO significantly improves performance in preference benchmarks, achieving stronger alignment with human preferences, reducing verbose responses, and mitigating overfitting. Furthermore, PSO exhibits robust generalization and sample efficiency, offering a scalable and promising solution for LLM alignment without relying on human-annotated preferences. Yuanzhao Zhai, Zhuo Zhang 0007, Cheng Yang 0004, Kele Xu, Yue Yu 0001, Wei Li 0022, Hui Wang 0030, Zenglin Xu, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 1 |
| 2025 | Empowering Large Language Model Agent through Step-Level Self-Critique and Self-TrainingabstractLarge Language Model (LLM) agents frequently produce sub-optimal actions when tackling complex, multi-step decision-making tasks. Employing self-critique to identify flaws and suggest enhancements is an effective strategy for refining actions. Although trajectory-level critique is commonly employed, it often fails to identify flawed steps accurately. In this paper, we introduce SLSC-MCTS, a method that integrates Monte Carlo Tree Search with Step-Level Self-Critique to enhance LLM agents during both testing and self-training phases. During decision tree expansion with SLSC-MCTS, the LLM agent initially generates an action, receives environmental feedback, and subsequently generates further actions via self-critique and refinement. Through multiple episodes of SLSC-MCTS, LLM agents can effectively utilize step-level critiques while disregarding ineffective ones based on node values, thereby incorporating the critiques more robustly. Additionally, our method further empowers LLM agents in a self-training manner, collecting training data from the constructed decision tree to iteratively fine-tune the LLM agents. The self-training data gathered via SLSC-MCTS is diverse and high-quality, which further enhances the reasoning, critiquing, and refining abilities of LLM agents. Experimental results demonstrate that SLSC-MCTS significantly improves LLM agents during testing, surpassing state-of-the-art baselines and achieving shorter task completion trajectories across information retrieval benchmarks such as WebShop and HotPotQA. After three iterations of self-training, LLM agents established by Llama-3.1-8B-Instruct show substantial improvement, even surpassing human experts in WebShop. Yuanzhao Zhai, Huanxi Liu, Zhuo Zhang 0007, Kele Xu, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 1 |
| 2024 | Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationabstractModel-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decision Process (P-MDP). However, the P-MDP discourages the policies from learning in out-of-distribution (OOD) regions beyond the support of offline datasets, which can under-utilize the generalization ability of dynamics models. In contrast, we propose constructing an Optimistic MDP (O-MDP). We initially observed the potential benefits of optimism brought by encouraging more OOD rollouts. Motivated by this observation, we present ORPO, a simple yet effective model-based offline RL framework. ORPO generates Optimistic model Rollouts for Pessimistic offline policy Optimization. Specifically, we train an optimistic rollout policy in the O-MDP to sample more OOD model rollouts. Then we relabel the sampled state-action pairs with penalized rewards, and optimize the output policy in the P-MDP. Theoretically, we demonstrate that the performance of policies trained with ORPO can be lower-bounded in linear MDPs. Experimental results show that our framework significantly outperforms P-MDP baselines by a margin of 30%, achieving state-of-the-art performance on the widely-used benchmark. Moreover, ORPO exhibits notable advantages in problems that require generalization. Yuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
AAAI | 1 |
| 2024 | Nuclear-Norm Maximization for Low-Rank UpdatesabstractPre-trained large language models exhibit significant potential in speech and language processing. Fine-tuning all parameters becomes impractical when confronted with numerous downstream tasks. To address this challenge, various low-rank adaptation techniques have been introduced for parameter-efficient fine-tuning, which freeze the over-parametrized models and learn incremental parameter updates within smaller subspaces. However, our observation reveals that most directions of the learned subspace play a minor role in the incremental updates. Consequently, fine-tuned models may not achieve optimal performance. To bridge this gap, we introduce NNM-LoRA, which strives to harness more meaningful singular directions. Through Nuclear Norm Maximization (NNM), we can better regulate the allocation of singular values. Accordingly, we propose a parameter-free plug-and-play regularizer for low-rank updates. This innovative approach allows us to utilize as many singular directions of the subspace as possible during the training of low-rank updates. To validate the effectiveness of NNM-LoRA, we conduct extensive experiments involving different pre-trained models on various natural language understanding tasks. Results demonstrate that NNM-LoRA exhibits significant improvements compared to baseline methods. Huanxi Liu, Yuanzhao Zhai, Kele Xu, Yiying Li |
ICASSP | 2 |
| 2024 | Iterative Regularized Policy Optimization with Imperfect DemonstrationsabstractImitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employed to stabilize the learning process. However, our investigation reveals that on the one hand, imperfect demonstrations can bias the online learning process, the KL regularization will further constrain the improvement of online policy exploration. To address the above issues, we propose Iterative Regularized Policy Optimization (IRPO), a framework that involves iterative offline imitation learning and online reinforcement exploration. Specifically, the policy learned online is used to serve as the demonstrator for successive learning iterations, with a demonstration boosting to consistently enhance the quality of demonstrations. Experimental validations conducted across widely used benchmarks and a novel fixed-wing UAV control task consistently demonstrate the effectiveness of IRPO in improving both the demonstration quality and the policy performance. Our code is available at https://github.com/GongXudong/IRPO. Xudong Gong, Kele Xu, Yuanzhao Zhai, Chengkang Yao, Bo Ding 0001, Huaimin Wang 0001 |
ICML | 4 |
| 2023 | Progressive Diversifying Policy for Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) has recently achieved promising performance in many collaborative decision making tasks. However, one of the main bottleneck challenges for MARL is the sparsity of the team reward, which can lead to the homogenization of agents’ behaviors. To address these issues, we propose a Progressive Diversifying Policy (PDP) algorithm in this paper. Specifically, we propose to actively amplify the diversity between agents’ policies during learning and exploit diversity as an additional intrinsic reward for MARL. Furthermore, we propose a progressive diversity boosting policy to find a better team policy. Leveraging the aforementioned improvements, our method can handle sparse team rewards and alleviate the homogeneous behaviors of agents. We conduct experiments on widely-used MARL environments and the results show that PDP can provide state-of-the-art performance while maintaining a competitive convergence speed. Shaoqi Sun, Yuanzhao Zhai, Kele Xu, Bo Ding 0001 |
ICASSP | 2 |
| 2023 | Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm RegularizationabstractThe use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challenge, we study the adjacency tensor of the communication graph. Then we define a new nuclear tensor rank and its convex surrogate, the normalized tensor nuclear norm to measure the homogeneity of message aggregation. Leveraging the norm, we further propose a plug-and-play regularizer on the adjacency tensor, named Normalized Tensor Nuclear Norm Regularization (NTNNR), to actively enrich the diversity of message aggregation during the training stage. NTNNR is agnostic to specific Comm-MARL algorithms and can be flexibly integrated with different graph-attention methods. Empirical results demonstrate that aggregating messages using NTNNR-enhanced GAT can improve the efficiency of the training and achieve higher asymptotic performance than existing message aggregation methods. Yuanzhao Zhai, Kele Xu, Bo Ding 0001, Zijian Gao, Huaimin Wang 0001 |
ICASSP | 1 |
| 2023 | Bi-level Multi-Agent Actor-Critic Methods with ransformersabstractRecently, deep multi-agent reinforcement learning methods have witnessed great progress, including multi-agent actor-critic methods. However, it’s worth noticing there is a performance gap between multi-agent actor-critic methods and state-of-the-art value-based methods. In this paper, we investigate the causes and attribute inferior performance to issues of contribution-mismatch and indiscriminate guidance. To overcome these problems, we introduce a novel bi-level multi-agent actorcritic reinforcement learning approach with transformers, called BMT. Specifically, we propose a simple but efficient bi-level optimization mechanism to learn both global critic and agentspecific critic, thus jointly guiding the policy update. In addition, we adopt the transformer-based model as the policy network to decouple complicated relationships and generate flexible policy. BMT is also general enough to be plugged into any actor-critic multi-agent reinforcement learning approach, such as MAPPO, and equips it with strong expression. On multiple benchmarks including multi-agent particle environments and a challenging set of StarCraft II micromanagement tasks, large-scale empirical experiments demonstrate that BMT-based multi-agent reinforcement learning methods achieve superior performance over both state-of-the-art actor-critic and value-based approaches. Tianjiao Wan, Haibo Mi, Zijian Gao, Yuanzhao Zhai, Bo Ding 0001 |
JCC | 4 |
| 2022 | Exploring Policy Diversity in Parallel Actor-Critic LearningabstractExploration is a critical challenge for deep reinforcement learning methods. Although existing works such as actor-critic algorithms have made much progress, most still suffer from the sample inefficiency problem in complex environments where rewards are sparse. Parallel sampling, which uses multiple actors with the same policy interacting with the environment, is an effective approach to improve sample efficiency. However, parallel parameter-sharing actors collect similar samples, which generally hinders the improvement of the overall exploration process. In this paper, we propose a Policy Diversity enhanced approach for parallel Actor-Critic (PDAC). Specifically, we extend the parallel actor-critic architecture to the PDAC framework composed of a shared critic and parallel distinct actors. Then we introduce the KL-divergence of the action probability distribution between parallel actors as the intrinsic reward to encourage actors to explore diverse strategies. We evaluate our approach in multiple challenging procedurally-generated tasks and compare it with state-of-the-art algorithms. Experiments show that PDAC makes significant progress in the comparison, in terms of cumulative rewards and sample efficiency. Yanqiang Zhang, Yuanzhao Zhai, Gongqian Zhou, Bo Ding 0001, Songwang Liu |
ICTAI | 2 |
| 2021 | Cloudroid Swarm: A QoS-Aware Framework for Multirobot Cooperation OffloadingabstractComputation offloading has been widely recognized as an effective way to promote the capabilities of resource‐constrained mobile devices. Recent years have seen a renewal of the importance of this technology in the emerging field of mobile robots, supporting resource‐intensive robot applications. However, cooperating to solve complex tasks in the physical world, which is a significant feature of a robot swarm compared to traditional mobile computing devices, has not received in‐depth attention in research concerned with traditional computation offloading. In this study, we propose an approach named cooperation offloading, which offloads the intensive communication among robots as well as the computation for compute‐intensive and data‐intensive tasks. We analyze the performance gain of cooperation offloading by formalizing multirobot cooperative models; in addition, we study offloading decisions. Based on this approach, we design a cloud robotic framework named Cloudroid Swarm and develop several QoS‐aware mechanisms to provide a general solution to cooperation offloading with QoS assurance in multirobot cooperative scenes. We implement Cloudroid Swarm to transparently migrate multirobot applications to cloud servers without any code modification. We evaluate our framework using three different multirobot cooperative applications. Our results show that Cloudroid Swarm can be applied to various robotic applications and real‐world environments and bring significant benefits in terms of both network optimization and task performance. Besides, our framework has good scalability and can do support as many as 256 robot entities simultaneously. Yuanzhao Zhai, Bo Ding 0001, Pengfei Zhang 0006 |
Wirel. Commun. Mob. Comput. | 1 |
| 2020 | Cooperative Offloading for Multiple Robot ApplicationsabstractComputation offloading has been widely recognized as an effective way to promote the capabilities of resource-constrained mobile devices. The past years have seen a renewed importance of this technology in the emerging field of mobile robots. However, a significant feature of robots compared to traditional mobile computing devices (e.g. smartphones) is that they must collaborate to solve complex tasks in the physical world in many cases, which implies intensive data exchange among the robot peers. This characteristic has not been dealt with in-depth in traditional computation offloading research. In this paper, we propose an approach called Cooperative Offloading, which takes into account the cooperation among mobile devices as well as the communication it brings in computation offloading. Firstly, we propose the offloading decision approach that deals with offloading from a set of robots that cooperate to finish a specific task. Secondly, we present a set of mechanisms to optimize the data transfer path on network topology in multi-robot applications, which can significantly reduce the bandwidth consumption of the robot wireless network. Based on the cooperative offloading approach, we realized a cloud robotics framework called Cloudroid Swarm. Evaluations based on real-life applications have shown that Cloudroid Swarm brings more than five times performance promotion compared to the setup without offloading or with individual computation offloading. Yuanzhao Zhai, Bo Ding 0001, Pengfei Zhang 0006, Qingtong Wu, Peichang Shi, Huaimin Wang 0001 |
JCC | 1 |