Chenxiao Gao

dblp:356/7556 · also Chen-Xiao Gao · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
0009-0009-9169-8056ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 82% Robot manipulation · 6% Optimization for machine learning · 6%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
offline reinforcement learning
1.622025
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning · ICML 2025
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning · AAAI 2024
Machine learning › Reinforcement learning
policy optimization
1.232025
Policy Rehearsing: Training Generalizable Policies for Reinforcement Learning · ICLR 2024
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Efficient and Stable Offline-to-online Reinforcement Learning via Continual Policy Revitalization · IJCAI 2024
Machine learning › Optimization for machine learning
black-box optimization
0.912025
Reinforced In-Context Black-Box Optimization · IJCAI 2025
Robotics › Robot manipulation
diffusion policy
0.912025
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › meta-reinforcement learning
in-context reinforcement learning
0.912025
Reinforced In-Context Black-Box Optimization · IJCAI 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation
0.912025
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Reward Models in Deep Reinforcement Learning: A Survey · IJCAI 2025
Machine learning › Generative modeling
diffusion model
0.812024
Diffusion Spectral Representation for Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › generalization in reinforcement learning
generalizable policy
0.812024
Policy Rehearsing: Training Generalizable Policies for Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning
meta-reinforcement learning
0.812024
Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data Limitations · AAAI 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
0.812024
Policy Rehearsing: Training Generalizable Policies for Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning › meta-reinforcement learning
offline meta-reinforcement learning
0.812024
Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data Limitations · AAAI 2024
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.812024
Efficient and Stable Offline-to-online Reinforcement Learning via Continual Policy Revitalization · IJCAI 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
spectral representation
0.812024
Diffusion Spectral Representation for Reinforcement Learning · NeurIPS 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
task representation learning
0.812024
Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data Limitations · AAAI 2024
Machine learning › Reinforcement learning
value-based reinforcement learning
0.812024
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning · AAAI 2024

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.6sequence model · 0.9reward learning · 0.9regret-to-go tokens · 0.9kullback-leibler regularization · 0.9in-context learning · 0.9actor-critic · 0.9sequence modeling · 0.8in-sample value iteration · 0.8advantage estimation · 0.8
YearPublicationVenuePosition
2025 Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
abstract
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous exploitation of unseen actions. Nevertheless, existing literature on behavior-regularized RL primarily focuses on explicit policy parameterizations, such as Gaussian policies. Consequently, it remains unclear how to extend this framework to more advanced policy parameterizations, such as diffusion models. In this paper, we introduce BDPO, a principled behavior-regularized RL framework tailored for diffusion-based policies, thereby combining the expressive power of diffusion policies and the robustness provided by regularization. The key ingredient of our method is to calculate the Kullback-Leibler (KL) regularization analytically as the accumulated discrepancies in reverse-time transition kernels along the diffusion trajectory. By integrating the regularization, we develop an efficient two-time-scale actor-critic RL algorithm that produces the optimal policy while respecting the behavior constraint. Comprehensive evaluations conducted on synthetic 2D tasks and continuous control tasks from the D4RL benchmark validate its effectiveness and superior performance.
Chenxiao Gao, Chenyang Wu 0001, Mingjun Cao 0001, Chenjun Xiao, Yang Yu 0001, Zongzhang Zhang
ICML1
2025 Reinforced In-Context Black-Box Optimization
abstract
Black-Box Optimization (BBO) has found successful applications in many fields of science and engineering. Recently, there has been a growing interest in meta-learning particular components of BBO algorithms to speed up optimization and get rid of tedious hand-crafted heuristics. As an extension, learning the entire algorithm from data requires the least labor from experts and can provide the most flexibility. In this paper, we propose RIBBO, a method to reinforce-learn a BBO algorithm from offline data in an end-to-end fashion. RIBBO employs expressive sequence models to learn the optimization histories produced by multiple behavior algorithms and tasks, leveraging the in-context learning ability of large models to extract task information and make decisions accordingly. Central to our method is to augment the optimization histories with regret-to-go tokens, which are designed to represent the performance of an algorithm based on cumulative regret over the future part of the histories. The integration of regret-to-go tokens enables RIBBO to automatically generate sequences of query points that are positively correlated to the user-desired regret, verified by its universally good empirical performance on diverse problems, including BBO benchmark, hyper-parameter optimization, and robot control problems.
Chenxiao Gao, Ke Xue 0001, Chenyang Wu 0001, Dong Li 0016, Jianye Hao, Zongzhang Zhang, Chao Qian 0001
IJCAI2
2025 Reward Models in Deep Reinforcement Learning: A Survey
abstract
In reinforcement learning (RL), agents continually interact with the environment and use the feedback to refine their behavior. To guide policy optimization, reward models are introduced as proxies of the desired objectives, such that when the agent maximizes the accumulated reward, it also fulfills the task designer's intentions. Recently, significant attention from both academic and industrial researchers has focused on developing reward models that not only align closely with the true objectives but also facilitate policy optimization. In this survey, we provide a comprehensive review of reward modeling techniques within the RL literature. We begin by outlining the background and preliminaries in reward modeling. Next, we present an overview of recent reward modeling approaches, categorizing them based on the source, the mechanism, and the reward learning paradigm. Building on this understanding, we discuss various applications of these reward modeling techniques and review methods for evaluating reward models. Finally, we conclude by highlighting promising research directions in reward modeling. Altogether, this survey includes both established and emerging methods, filling the vacancy of a systematic review of reward models in current literature.
Shenghua Wan, Yucen Wang, Chenxiao Gao, Le Gan, Zongzhang Zhang, De-Chuan Zhan
IJCAI4
2024 ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning
abstract
Decision Transformer (DT), which employs expressive sequence modeling techniques to perform action generation, has emerged as a promising approach to offline policy optimization. However, DT generates actions conditioned on a desired future return, which is known to bear some weaknesses such as the susceptibility to environmental stochasticity. To overcome DT's weaknesses, we propose to empower DT with dynamic programming. Our method comprises three steps. First, we employ in-sample value iteration to obtain approximated value functions, which involves dynamic programming over the MDP structure. Second, we evaluate action quality in context with estimated advantages. We introduce two types of advantage estimators, IAE and GAE, which are suitable for different tasks. Third, we train an Advantage-Conditioned Transformer (ACT) to generate actions conditioned on the estimated advantages. Finally, during testing, ACT generates actions conditioned on a desired advantage. Our evaluation results validate that, by leveraging the power of dynamic programming, ACT demonstrates effective trajectory stitching and robust action generation in spite of the environmental stochasticity, outperforming baseline methods across various benchmarks. Additionally, we conduct an in-depth analysis of ACT's various design choices through ablation studies. Our code is available at https://github.com/LAMDA-RL/ACT.
Chenxiao Gao, Chenyang Wu 0001, Mingjun Cao 0001, Zongzhang Zhang, Yang Yu 0001
AAAI1
2024 Generalizable Task Representation Learning for Offline Meta-Reinforcement Learning with Data Limitations
abstract
Generalization and sample efficiency have been long-standing issues concerning reinforcement learning, and thus the field of Offline Meta-Reinforcement Learning (OMRL) has gained increasing attention due to its potential of solving a wide range of problems with static and limited offline data. Existing OMRL methods often assume sufficient training tasks and data coverage to apply contrastive learning to extract task representations. However, such assumptions are not applicable in several real-world applications and thus undermine the generalization ability of the representations. In this paper, we consider OMRL with two types of data limitations: limited training tasks and limited behavior diversity and propose a novel algorithm called GENTLE for learning generalizable task representations in the face of data limitations. GENTLE employs Task Auto-Encoder (TAE), which is an encoder-decoder architecture to extract the characteristics of the tasks. Unlike existing methods, TAE is optimized solely by reconstruction of the state transition and reward, which captures the generative structure of the task models and produces generalizable representations when training tasks are limited. To alleviate the effect of limited behavior diversity, we consistently construct pseudo-transitions to align the data distribution used to train TAE with the data distribution encountered during testing. Empirically, GENTLE significantly outperforms existing OMRL methods on both in-distribution tasks and out-of-distribution tasks across both the given-context protocol and the one-shot protocol.
Renzhe Zhou, Chenxiao Gao, Zongzhang Zhang, Yang Yu 0001
AAAI2
2024 Policy Rehearsing: Training Generalizable Policies for Reinforcement Learning
abstract
Human beings can make adaptive decisions in a preparatory manner, i.e., by making preparations in advance, which offers significant advantages in scenarios where both online and offline experiences are expensive and limited. Meanwhile, current reinforcement learning methods commonly rely on numerous environment interactions but hardly obtain generalizable policies. In this paper, we introduce the idea of \textit{rehearsal} into policy optimization, where the agent plans for all possible outcomes in mind and acts adaptively according to actual responses from the environment. To effectively rehearse, we propose ReDM, an algorithm that generates a diverse and eligible set of dynamics models and then rehearse the policy via adaptive training on the generated model set. Rehearsal enables the policy to make decision plans for various hypothetical dynamics and to naturally generalize to previously unseen environments. Our experimental results demonstrate that ReDM is capable of learning a valid policy solely through rehearsal, even with \emph{zero} interaction data. We further extend ReDM to scenarios where limited or mismatched interaction data is available, and our experimental results reveal that ReDM produces high-performing policies compared to other offline RL baselines.
Chengxing Jia, Chenxiao Gao, Fuxiang Zhang, Xiong-Hui Chen, Tian Xu 0003, Lei Yuan 0005, Zongzhang Zhang, Zhi-Hua Zhou, Yang Yu 0001
ICLR2
2024 Efficient and Stable Offline-to-online Reinforcement Learning via Continual Policy Revitalization
Chenyang Wu 0001, Chenxiao Gao, Zongzhang Zhang, Ming Li 0005
IJCAI3
2024 Diffusion Spectral Representation for Reinforcement Learning
abstract
Diffusion-based models have achieved notable empirical successes in reinforcement learning (RL) due to their expressiveness in modeling complex distributions. Despite existing methods being promising, the key challenge of extending existing methods for broader real-world applications lies in the computational cost at inference time, i.e., sampling from a diffusion model is considerably slow as it often requires tens to hundreds of iterations to generate even one sample. To circumvent this issue, we propose to leverage the flexibility of diffusion models for RL from a representation learning perspective. In particular, by exploiting the connection between diffusion models and energy-based models, we develop Diffusion Spectral Representation (Diff-SR), a coherent algorithm framework that enables extracting sufficient representations for value functions in Markov decision processes (MDP) and partially observable Markov decision processes (POMDP). We further demonstrate how Diff-SR facilitates efficient policy optimization and practical algorithms while explicitly bypassing the difficulty and inference cost of sampling from the diffusion model. Finally, we provide comprehensive empirical studies to verify the benefits of Diff-SR in delivering robust and advantageous performance across various benchmarks with both fully and partially observable settings.
Dmitry Shribak, Chenxiao Gao, Chenjun Xiao, Bo Dai 0001
NeurIPS2