Zhepeng Cen

dblp:254/6182 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 69% Trustworthy machine learning · 11% Generative modeling · 9%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
4.772024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Feasibility Consistent Representation Learning for Safe Reinforcement Learning · ICML 2024
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
offline reinforcement learning
2.232024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Learning from Sparse Offline Datasets via Conservative Density Estimation · ICLR 2024
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Trustworthy machine learning
robustness
1.322023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
On the Robustness of Safe Reinforcement Learning under Observational Perturbations · ICLR 2023
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
1.222023
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning · NeurIPS 2023
Constrained Variational Policy Optimization for Safe Reinforcement Learning · ICML 2022
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models
0.912025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Machine learning › Deep learning architectures and training
data augmentation
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Generative modeling
synthetic data generation
0.812024
OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy learning
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › offline reinforcement learning
decision transformer
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
0.712023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023
Machine learning › Reinforcement learning
policy optimization
0.612022
Constrained Variational Policy Optimization for Safe Reinforcement Learning · ICML 2022
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.312025
Behavior Injection: Preparing Language Models for Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.212023
Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data · ICML 2023
Mathematical optimization
multi-objective optimization
0.212023
Constrained Decision Transformer for Offline Safe Reinforcement Learning · ICML 2023

Methods — techniques the papers use, named apart from their topics

variational inference · 1.3convex optimization · 1.2supervised fine-tuning · 0.9reinforcement learning · 0.9data augmentation · 0.9self-supervised learning · 0.8representation learning · 0.8importance sampling · 0.8density estimation · 0.8conservative constraints · 0.8zero-shot adaptation · 0.7multi-objective optimization · 0.7decision transformer · 0.7
YearPublicationVenuePosition
2025 Behavior Injection: Preparing Language Models for Reinforcement Learning
abstract
Reinforcement learning (RL) has emerged as a powerful post-training technique to incentivize the reasoning ability of large language models (LLMs). However, LLMs can respond very inconsistently to RL finetuning: some show substantial performance gains, while others plateau or even degrade. To understand this divergence, we analyze the per-step influence of the RL objective and identify two key conditions for effective post-training: (1) RL-informative rollout accuracy, and (2) strong data co-influence, which quantifies how much the training data affects performance on other samples. Guided by these insights, we propose behavior injection, a task-agnostic data augmentation scheme applied prior to RL. Behavior injection enriches the supervised finetuning (SFT) data by seeding exploratory and exploitative behaviors, effectively making the model more RL-ready. We evaluate our method across two reasoning benchmarks with multiple base models. The results demonstrate that our theoretically motivated augmentation can significantly increase the performance gain from RL over the pre-RL model.
Zhepeng Cen, Yihang Yao, William Jongwon Han, Zuxin Liu, Ding Zhao
NeurIPS1
2024 Learning from Sparse Offline Datasets via Conservative Density Estimation
abstract
Offline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL.
Zhepeng Cen, Zuxin Liu, Zitong Wang 0005, Yihang Yao, Henry Lam, Ding Zhao
ICLR1
2024 Feasibility Consistent Representation Learning for Safe Reinforcement Learning
abstract
In the field of safe reinforcement learning (RL), finding a balance between satisfying safety constraints and optimizing reward performance presents a significant challenge. A key obstacle in this endeavor is the estimation of safety constraints, which is typically more difficult than estimating a reward metric due to the sparse nature of the constraint signals. To address this issue, we introduce a novel framework named Feasibility Consistent Safe Reinforcement Learning (FCSRL). This framework combines representation learning with feasibility-oriented objectives to identify and extract safety-related information from the raw state for safe RL. Leveraging self-supervised learning techniques and a more learnable safety metric, our approach enhances the policy learning and constraint estimation. Empirical evaluations across a range of vector-state and image-based tasks demonstrate that our method is capable of learning a better safety-aware embedding and achieving superior performance than previous representation learning baselines.
Zhepeng Cen, Yihang Yao, Zuxin Liu, Ding Zhao
ICML1
2024 OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning
abstract
Offline safe reinforcement learning (RL) aims to train a policy that satisfies con- straints using a pre-collected dataset. Most current methods struggle with the mismatch between imperfect demonstrations and the desired safe and rewarding performance. In this paper, we mitigate this issue from a data-centric perspective and introduce OASIS (cOnditionAl diStributIon Shaping), a new paradigm in offline safe RL designed to overcome these critical limitations. OASIS utilizes a conditional diffusion model to synthesize offline datasets, thus shaping the data dis- tribution toward a beneficial target domain. Our approach makes compliance with safety constraints through effective data utilization and regularization techniques to benefit offline safe RL training. Comprehensive evaluations on public benchmarks and varying datasets showcase OASIS’s superiority in benefiting offline safe RL agents to achieve high-reward behavior while satisfying the safety constraints, out- performing established baselines. Furthermore, OASIS exhibits high data efficiency and robustness, making it suitable for real-world applications, particularly in tasks where safety is imperative and high-quality demonstrations are scarce. More details are available at the website https://sites.google.com/view/saferl-oasis/home.
Yihang Yao, Zhepeng Cen, Wenhao Ding, Haohong Lin, Shiqi Liu 0005, Tingnan Zhang, Wenhao Yu 0003, Ding Zhao
NeurIPS2
2023 On the Robustness of Safe Reinforcement Learning under Observational Perturbations
Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Jie Tan 0001, Bo Li 0026, Ding Zhao
ICLR3
2023 Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data
abstract
Previous work demonstrates that the optimal safe reinforcement learning policy in a noise-free environment is vulnerable and could be unsafe under observational attacks. While adversarial training effectively improves robustness and safety, collecting samples by attacking the behavior agent online could be expensive or prohibitively dangerous in many applications. We propose the robuSt vAriational ofF-policy lEaRning (SAFER) approach, which only requires benign training data without attacking the agent. SAFER obtains an optimal non-parametric variational policy distribution via convex optimization and then uses it to improve the parameterized policy robustly via supervised learning. The two-stage policy optimization facilitates robust training, and extensive experiments on multiple robot platforms show the efficiency of SAFER in learning a robust and safe policy: achieving the same reward with much fewer constraint violations during training than on-policy baselines.
Zuxin Liu, Zijian Guo 0002, Zhepeng Cen, Huan Zhang 0001, Yihang Yao, Hanjiang Hu, Ding Zhao
ICML3
2023 Constrained Decision Transformer for Offline Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) trains a constraint satisfaction policy by interacting with the environment. We aim to tackle a more challenging problem: learning a safe policy from an offline dataset. We study the offline safe RL problem from a novel multi-objective optimization perspective and propose the $\epsilon$-reducible concept to characterize problem difficulties. The inherent trade-offs between safety and task performance inspire us to propose the constrained decision transformer (CDT) approach, which can dynamically adjust the trade-offs during deployment. Extensive experiments show the advantages of the proposed method in learning an adaptive, safe, robust, and high-reward policy. CDT outperforms its variants and strong offline safe RL baselines by a large margin with the same hyperparameters across all tasks, while keeping the zero-shot adaptation capability to different constraint thresholds, making our approach more suitable for real-world RL under constraints.
Zuxin Liu, Zijian Guo 0002, Yihang Yao, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao
ICML4
2023 Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications.
Yihang Yao, Zuxin Liu, Zhepeng Cen, Wenhao Yu 0003, Tingnan Zhang, Ding Zhao
NeurIPS3
2022 Constrained Variational Policy Optimization for Safe Reinforcement Learning
abstract
Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications. Previous primal-dual style approaches suffer from instability issues and lack optimality guarantees. This paper overcomes the issues from the perspective of probabilistic inference. We introduce a novel Expectation-Maximization approach to naturally incorporate constraints during the policy learning: 1) a provable optimal non-parametric variational distribution could be computed in closed form after a convex optimization (E-step); 2) the policy parameter is improved within the trust region based on the optimal variational distribution (M-step). The proposed algorithm decomposes the safe RL problem into a convex optimization phase and a supervised learning phase, which yields a more stable training performance. A wide range of experiments on continuous robotic tasks shows that the proposed method achieves significantly better constraint satisfaction performance and better sample efficiency than baselines. The code is available at https://github.com/liuzuxin/cvpo-safe-rl.
Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Steven Z. Wu, Bo Li 0026, Ding Zhao
ICML2