VLDB 2026 Research / reviewers in the wild / expert
Yaozhong Gan
dblp:234/8610
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARPO: A Reflective Policy Optimization for Multi-Agent Reinforcement LearningabstractWe propose Multi-Agent Reflective Policy Optimization MARPO to alleviate the issue of sample inefficiency in multi-agent reinforcement learning. MARPO consists of two key components: a reflection mechanism that leverages subsequent trajectories to enhance sample efficiency, and an asymmetric clipping mechanism that is derived from the KL divergence and dynamically adjusts the clipping range to improve training stability. We evaluate MARPO in classic multi-agent environments, where it consistently outperforms other methods. Cuiling Wu, Yaozhong Gan, Junliang Xing |
AAAI | 2 |
| 2026 | ORAL: Adaptive Gap Increasing for Advantage Learning via Occam's Razor PrincipleabstractBenefiting from the gap increasing between the optimal action and its competitors, the advantage learning (AL) operator is more robust to estimation errors in the approximated $Q$ -functions than the Bellman optimality operator in reinforcement learning (RL). However, our analysis reveals that its robustness and larger action gaps come at the cost of a worse performance loss bound, leading to slower convergence of value functions. To address this issue, we present a novel method, named Occam's Razor-based AL (ORAL), which follows Occam's Razor principle and takes the necessity into consideration when increasing the action gap. Specifically, our ORAL can adaptively increase the action gap for different state-action pairs, depending on the proximity of their $Q$ values to the optimal ones. We first propose a naive implementation of ORAL, employing a nonsmooth clipping function to realize the above idea, and then introduce a smooth version of ORAL aimed at achieving more stable learning. Furthermore, our methods can be easily plugged into other AL-based operators and extended to more complex continuous-control tasks. Theoretical analysis supports the feasibility of our approaches, demonstrating their ability to balance the gap increasing with fast convergence. Empirical results further validate its effectiveness, showing significant performance improvements across multiple benchmarks. Yongle Zhou, Yuyang Long, Jia Zhang 0019, Juanjuan Weng, Zhetao Li, Yaozhong Gan, Xiaoyang Tan |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001 |
ICCV | 3 |
| 2024 | PAE: Reinforcement Learning from External Knowledge for Efficient ExplorationabstractHuman intelligence is adept at absorbing valuable insights from external knowledge.
This capability is equally crucial for artificial intelligence.
In contrast, classical reinforcement learning agents lack such capabilities and often resort to extensive trial and error to explore the environment.
This paper introduces $\textbf{PAE}$: $\textbf{P}$lanner-$\textbf{A}$ctor-$\textbf{E}$valuator, a novel framework for teaching agents to $\textit{learn to absorb external knowledge}$.
PAE integrates the Planner's knowledge-state alignment mechanism, the Actor's mutual information skill control, and the Evaluator's adaptive intrinsic exploration reward to achieve 1) effective cross-modal information fusion, 2) enhanced linkage between knowledge and state, and 3) hierarchical mastery of complex tasks.
Comprehensive experiments across
11 challenging tasks from the BabyAI and MiniHack environment suites demonstrate PAE's superior exploration efficiency with good interpretability. Haofei Lu, Junliang Xing, Renye Yan, Yaozhong Gan, Yuanchun Shi |
ICLR | 6 |
| 2024 | Reflective Policy OptimizationabstractOn-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensive data per update, leading to sample inefficiency. This paper introduces Reflective Policy Optimization (RPO), a novel on-policy extension that amalgamates past and future state-action information for policy optimization. This approach empowers the agent for introspection, allowing modifications to its actions within the current state. Theoretical analysis confirms that policy performance is monotonically improved and contracts the solution space, consequently expediting the convergence procedure. Empirical results demonstrate RPO's feasibility and efficacy in two reinforcement learning benchmarks, culminating in superior sample efficiency. The source code of this work is available at https://github.com/Edgargan/RPO. Yaozhong Gan, Renye Yan, Junliang Xing |
ICML | 1 |
| 2024 | Autoencoder Reconstruction Model for Long-Horizon ExplorationabstractConventional reinforcement learning (RL) algorithms often necessitate millions of environment interactions to ascertain an efficacious policy. In stark contrast, humans, leveraging their curiosity mechanisms, can develop proficient policies with minimal effort. Drawing inspiration from this observation, we introduce the Autoencoder Reconstruction Model(ARM), a curiosity-driven RL model that significantly reduces interactions while enhancing policy effectiveness. ARM employs an autoencoder module, utilizing a deep neural network to learn feature representations from the environment. ARM utilizes its Curiosity Measurement Module to motivate RL agents for effective exploration, particularly in environments with sparse rewards. ARM also introduces an innovative mechanism to balance the exploration-exploitation dilemma. Theoretical analyses reveal that the reward shaping introduced by the ARM aligns with the potential-based reward shaping paradigm, thereby preserving the optimality of reinforcement learning. We will release the source code and trained models to facilitate further studies in this research direction. Renye Yan, Yaozhong Gan, Yunfan Yang, Zhaoke Yu, Zongxi Liu, Ling Liang 0003, Yimao Cai |
IJCNN | 3 |
| 2022 | Smoothing Advantage LearningabstractAdvantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the method tends to be unstable in the case of function approximation. In this paper, we propose a simple variant of AL, named smoothing advantage learning (SAL), to alleviate this problem. The key to our method is to replace the original Bellman Optimal operator in AL with a smooth one so as to obtain more reliable estimation of the temporal difference target. We give a detailed account of the resulting action gap and the performance bound for approximate SAL. Further theoretical analysis reveals that the proposed value smoothing technique not only helps to stabilize the training procedure of AL by controlling the trade-off between convergence rate and the upper bound of the approximation errors, but is beneficial to increase the action gap between the optimal and sub-optimal action value as well. Yaozhong Gan, Xiaoyang Tan |
AAAI | 1 |
| 2022 | Robust Action Gap Increasing with Clipped Advantage LearningabstractAdvantage Learning (AL) seeks to increase the action gap between the optimal action and its competitors, so as to improve the robustness to estimation errors. However, the method becomes problematic when the optimal action induced by the approximated value function does not agree with the true optimal action. In this paper, we present a novel method, named clipped Advantage Learning (clipped AL), to address this issue. The method is inspired by our observation that increasing the action gap blindly for all given samples while not taking their necessities into account could accumulate more errors in the performance loss bound, leading to a slow value convergence, and to avoid that, we should adjust the advantage value adaptively. We show that our simple clipped AL operator not only enjoys fast convergence guarantee but also retains proper action gaps, hence achieving a good balance between the large action gap and the fast convergence. The feasibility and effectiveness of the proposed method are verified empirically on several RL benchmarks with promising performance. Yaozhong Gan, Xiaoyang Tan |
AAAI | 2 |
| 2022 | Alleviating the estimation bias of deep deterministic policy gradient via co-regularization
Yuhui Wang 0004, Yaozhong Gan, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2021 | Stabilizing Q Learning Via Soft Mellowmax OperatorabstractLearning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal difference updates can theoretically cause instability for most linear or non-linear approximation schemes. Mellowmax is a recently proposed differentiable and non-expansion softmax operator that allows a convergent behavior in learning and planning. Unfortunately, the performance bound for the fixed point it converges to remains unclear, and in practice, its parameter is sensitive to various domains and has to be tuned case by case. Finally, the Mellowmax operator may suffer from oversmoothing as it ignores the probability being taken for each action when aggregating them. In this paper we address all the above issues with an enhanced Mellowmax operator, named SM2 (Soft Mellowmax). Particularly, the proposed operator is reliable, easy to implement, and has provable performance guarantee, while preserving all the advantages of Mellowmax. Furthermore, we show that our SM2 operator can be applied to the challenging multi-agent reinforcement learning scenarios, leading to stable value function approximation and state of the art performance. Yaozhong Gan, Xiaoyang Tan |
AAAI | 1 |
| 2019 | Trust Region-Guided Proximal Policy OptimizationabstractProximal policy optimization (PPO) is one of the most popular deep reinforcement learning (RL) methods, achieving state-of-the-art performance across a wide range of challenging tasks. However, as a model-free RL method, the success of PPO relies heavily on the effectiveness of its exploratory policy search. In this paper, we give an in-depth analysis on the exploration behavior of PPO, and show that PPO is prone to suffer from the risk of lack of exploration especially under the case of bad initialization, which may lead to the failure of training or being trapped in bad local optima. To address these issues, we proposed a novel policy optimization method, named Trust Region-Guided PPO (TRGPPO), which adaptively adjusts the clipping range within the trust region. We formally show that this method not only improves the exploration ability within the trust region but enjoys a better performance bound compared to the original PPO as well. Extensive experiments verify the advantage of the proposed method. Yuhui Wang 0004, Xiaoyang Tan, Yaozhong Gan |
NeurIPS | 4 |