Shenao Zhang

dblp:253/4543 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 49% Language models and text generation · 31% Motion planning and robot control · 10%

Topics — the 24 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy optimization
policy gradient
2.132024
Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations · ICML 2024
Model-Based Reparameterization Policy Gradient Methods: Theory and Practical Algorithms · NeurIPS 2023
Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics · ICML 2023
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
1.622025
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs · ICML 2025
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model reasoning
1.622025
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning · ICML 2025
Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM Agents · ICML 2024
Machine learning › Reinforcement learning
policy optimization
1.322024
Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations · ICML 2024
Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement Learning · NeurIPS 2022
Machine learning › Reinforcement learning
model-based reinforcement learning
1.222023
Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration · NeurIPS 2023
Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement Learning · NeurIPS 2022
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning · ICML 2025
Natural language and speech › Language models and text generation › large language model training
post-training
0.912025
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning · ICML 2025
Natural language and speech › Language models and text generation › alignment
preference alignment
0.912025
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs · ICML 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs · ICML 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
agent planning
0.812024
Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM Agents · ICML 2024
Natural language and speech › Language models and text generation
alignment
0.812024
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer · NeurIPS 2024
Machine learning › Reinforcement learning
differentiable simulation
0.812024
Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations · ICML 2024
Robotics › Motion planning and robot control › dynamic modeling
non-smooth dynamics
0.812024
Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations · ICML 2024
Robotics › Motion planning and robot control › robot dynamics
contact dynamics
0.712023
Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics · ICML 2023
Robotics › Robot manipulation
contact-rich manipulation
0.712023
Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics · ICML 2023
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.712023
Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration · NeurIPS 2023
Robotics › Motion planning and robot control
robot control
0.712023
Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics · ICML 2023
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.712023
Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration · NeurIPS 2023
Machine learning › Optimization for machine learning
variance reduction
0.712023
Model-Based Reparameterization Policy Gradient Methods: Theory and Practical Algorithms · NeurIPS 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
zero-sum markov game
0.712023
Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration · NeurIPS 2023
Machine learning › Reinforcement learning › exploration
efficient exploration
0.612022
Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement Learning · NeurIPS 2022
Machine learning › Reinforcement learning
exploration
0.612022
Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement Learning · NeurIPS 2022
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.312025
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs · ICML 2025

Methods — techniques the papers use, named apart from their topics

reward shaping · 0.9reward conditioning · 0.9reinforcement learning · 0.9probabilistic graphical model · 0.9direct preference optimization · 0.9data relabeling · 0.9variance analysis · 0.8bellman equation · 0.8adaptive analytic gradient · 0.8actor-critic · 0.8
YearPublicationVenuePosition
2025 BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, yet generating reliable reasoning processes remains a significant challenge. We present a unified probabilistic framework that formalizes LLM reasoning through a novel graphical model incorporating latent thinking processes and evaluation signals. Our framework addresses two critical questions: (1) how to generate high-quality reasoning processes during inference automatically, and (2) how to integrate these processes into post-training. We propose the Bootstrapping Reinforced Thinking Process (BRiTE) algorithm and demonstrate its theoretical convergence at a rate of $1/T$, where $T$ is the number of iterations. The algorithm operates in two steps. First, it generates high-quality rationales by approximating the desired posterior distribution using a reinforcement learning approach with a novel reward shaping mechanism. Second, it fine-tunes the base LLM by maximizing the joint probability of rationale generation with respect to LLM parameters. Empirical evaluation on GSM8K and MATH benchmarks demonstrates that our approach consistently improves performance across different model sizes without requiring human-annotated thinking processes, outperforming standard chain-of-thought prompting while enhancing existing post-training methods.
Han Zhong 0001, Yutong Yin, Shenao Zhang, Yuanxin Liu, Yifei Zuo, Boyi Liu 0001, Sirui Zheng, Hongyi Guo, Liwei Wang 0001, Mingyi Hong 0001, Zhaoran Wang 0001
ICML3
2025 Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
abstract
Preference alignment in Large Language Models (LLMs) has significantly improved their ability to adhere to human instructions and intentions. However, existing direct alignment algorithms primarily focus on relative preferences and often overlook the qualitative aspects of responses, despite having access to preference data that includes reward scores from judge models during AI feedback. Striving to maximize the implicit reward gap between the chosen and the slightly inferior rejected responses can cause overfitting and unnecessary unlearning of the high-quality rejected responses. The unawareness of the reward scores also drives the LLM to indiscriminately favor the low-quality chosen responses and fail to generalize to optimal responses that are sparse in data. To overcome these shortcomings, our study introduces reward-conditioned LLM policies that discern and learn from the entire spectrum of response quality within the dataset, helping extrapolate to more optimal regions. We propose an effective yet simple data relabeling method that conditions the preference pairs on quality scores to construct a reward-augmented dataset. The experiments across various benchmarks and diverse models demonstrate that our approach consistently boosts DPO by a considerable margin. Through comprehensive ablation studies, we demonstrate that our method not only maximizes the utility of preference data but also mitigates the issue of unlearning, demonstrating its broad effectiveness beyond mere data expansion. Our code is available at https://github.com/shenao-zhang/reward-augmented-preference.
Shenao Zhang, Boyi Liu 0001, Yufeng Zhang 0007, Yingxiang Yang, Yongfei Liu, Liyu Chen, Zhaoran Wang 0001
ICML1
2024 Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable Simulations
abstract
Recent advancements in differentiable simulators highlight the potential of policy optimization using simulation gradients. Yet, these approaches are largely contingent on the continuity and smoothness of the simulation, which precludes the use of certain simulation engines, such as Mujoco. To tackle this challenge, we introduce the adaptive analytic gradient. This method views the Q function as a surrogate for future returns, consistent with the Bellman equation. By analyzing the variance of batched gradients, our method can autonomously opt for a more resilient Q function to compute the gradient when encountering rough simulation transitions. We also put forth the Adaptive-Gradient Policy Optimization (AGPO) algorithm, which leverages our proposed method for policy learning. On the theoretical side, we demonstrate AGPO’s convergence, emphasizing its stable performance under non-smooth dynamics due to low variance. On the empirical side, our results show that AGPO effectively mitigates the challenges posed by non-smoothness in policy learning through differentiable simulation.
Liangzhi Shi, Shenao Zhang, Zhaoran Wang 0001, Yi Wu 0013
ICML3
2024 Reason for Future, Act for Now: A Principled Architecture for Autonomous LLM Agents
abstract
Large language models (LLMs) demonstrate impressive reasoning abilities, but translating reasoning into actions in the real world remains challenging. In particular, it is unclear how to complete a given task provably within a minimum number of interactions with the external environment, e.g., through an internal mechanism of reasoning. To this end, we propose the first framework with provable regret guarantees to orchestrate reasoning and acting, which we call reason for future, act for now (RAFA). Specifically, we design a prompt template for reasoning that learns from the memory buffer and plans a future trajectory over a long horizon (reason for future). At each step, the LLM agent takes the initial action of the planned trajectory (act for now), stores the collected feedback in the memory buffer, and reinvokes the reasoning routine to replan the future trajectory from the new state. The key idea is to cast reasoning in LLMs as learning and planning in Bayesian adaptive Markov decision processes (MDPs). Correspondingly, we prompt LLMs with the memory buffer to estimate the unknown environment (learning) and generate an optimal trajectory for multiple future steps that maximize a value function (planning). The learning and planning subroutines are performed in an in-context manner to emulate the actor-critic update for MDPs. Our theoretical analysis establishes a $\sqrt{T}$ regret, while our experimental validation demonstrates superior empirical performance.
Shenao Zhang, Hongyi Guo, Shuqi Ke, Boyi Liu 0001, Zhaoran Wang 0001
ICML3
2024 Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
abstract
Aligning generative models with human preference via RLHF typically suffers from overoptimization, where an imperfectly learned reward model can misguide the generative model to output even undesired responses. We investigate this problem in a principled manner by identifying the source of the issue as the distributional shift and uncertainty of human preference in dataset. To mitigate overoptimization, we first propose a theoretical algorithm which optimizes the policy against an adversarially chosen reward model, one that simultaneously minimizes its MLE loss and a reward penalty term. The penalty pessimistically biases the uncertain rewards so as to prevent the policy from choosing actions with spursiouly high proxy rewards, resulting in provable sample efficiency of the algorithm under a partial coverage style condition. Moving from theory to practice, the proposed algorithm further enjoys an equivalent but surprisingly easy to implement form. With a clever usage of the equivalence between reward models and the corresponding optimal policy, the algorithm features a simple objective that combines (i) a preference optimization loss that directly aligns the policy with human preference, and (ii) a supervised learning loss which explicitly imitates the policy with a baseline distribution. In the context of aligning large language models (LLM), this objective fuses the direct preference optimization (DPO) loss with the supervised fune-tuning (SFT) loss to help mitigate the overoptimization towards undesired responses, for which we name the algorithm Regularized Preference Optimization (RPO). Experiments of aligning LLMs demonstrate the improved performance of our method when compared with DPO baselines. Our work sheds light on the interplay between preference optimization and SFT in tuning LLMs with both theoretical guarantees and empirical evidence.
Miao Lu, Shenao Zhang, Boyi Liu 0001, Hongyi Guo, Yingxiang Yang, Jose H. Blanchet, Zhaoran Wang 0001
NeurIPS3
2023 Adaptive Barrier Smoothing for First-Order Policy Gradient with Contact Dynamics
abstract
Differentiable physics-based simulators have witnessed remarkable success in robot learning involving contact dynamics, benefiting from their improved accuracy and efficiency in solving the underlying complementarity problem. However, when utilizing the First-Order Policy Gradient (FOPG) method, our theory indicates that the complementarity-based systems suffer from stiffness, leading to an explosion in the gradient variance of FOPG. As a result, optimization becomes challenging due to chaotic and non-smooth loss landscapes. To tackle this issue, we propose a novel approach called Adaptive Barrier Smoothing (ABS), which introduces a class of softened complementarity systems that correspond to barrier-smoothed objectives. With a contact-aware adaptive central-path parameter, ABS reduces the FOPG gradient variance while controlling the gradient bias. We justify the adaptive design by analyzing the roots of the system’s stiffness. Additionally, we establish the convergence of FOPG and show that ABS achieves a reasonable trade-off between the gradient variance and bias by providing their upper bounds. Moreover, we present a variant of FOPG based on complementarity modeling that efficiently fits the contact dynamics by learning the physical parameters. Experimental results on various robotic tasks are provided to support our theory and method.
Shenao Zhang, Wanxin Jin, Zhaoran Wang 0001
ICML1
2023 Maximize to Explore: One Objective Function Fusing Estimation, Planning, and Exploration
abstract
In reinforcement learning (RL), balancing exploration and exploitation is crucial for achieving an optimal policy in a sample-efficient way. To this end, existing sample- efficient algorithms typically consist of three components: estimation, planning, and exploration. However, to cope with general function approximators, most of them involve impractical algorithmic components to incentivize exploration, such as data-dependent level-set constraints or complicated sampling procedures. To address this challenge, we propose an easy-to-implement RL framework called Maximize to Explore (MEX), which only needs to optimize unconstrainedly a single objective that integrates the estimation and planning components while balancing exploration and exploitation automatically. Theoretically, we prove that the MEX achieves a sublinear regret with general function approximators and is extendable to the zero-sum Markov game setting. Meanwhile, we adapt deep RL baselines to design practical versions of MEX in both the model-based and model-free settings, which outperform baselines in various MuJoCo environments with sparse reward by a stable margin. Compared with existing sample-efficient algorithms with general function approximators, MEX achieves similar sample efficiency while also enjoying a lower computational cost and is more compatible with modern deep RL methods.
Miao Lu, Wei Xiong 0015, Han Zhong 0001, Shenao Zhang, Sirui Zheng, Zhuoran Yang, Zhaoran Wang 0001
NeurIPS6
2023 Model-Based Reparameterization Policy Gradient Methods: Theory and Practical Algorithms
abstract
ReParameterization (RP) Policy Gradient Methods (PGMs) have been widely adopted for continuous control tasks in robotics and computer graphics. However, recent studies have revealed that, when applied to long-term reinforcement learning problems, model-based RP PGMs may experience chaotic and non-smooth optimization landscapes with exploding gradient variance, which leads to slow convergence. This is in contrast to the conventional belief that reparameterization methods have low gradient estimation variance in problems such as training deep generative models. To comprehend this phenomenon, we conduct a theoretical examination of model-based RP PGMs and search for solutions to the optimization difficulties. Specifically, we analyze the convergence of the model-based RP PGMs and pinpoint the smoothness of function approximators as a major factor that affects the quality of gradient estimation. Based on our analysis, we propose a spectral normalization method to mitigate the exploding variance issue caused by long model unrolls. Our experimental results demonstrate that proper normalization significantly reduces the gradient variance of model-based RP PGMs. As a result, the performance of the proposed method is comparable or superior to other gradient estimators, such as the Likelihood Ratio (LR) gradient estimator. Our code is available at https://github.com/agentification/RP_PGM.
Shenao Zhang, Boyi Liu 0001, Zhaoran Wang 0001, Tuo Zhao
NeurIPS1
2022 Conservative Dual Policy Optimization for Efficient Model-Based Reinforcement Learning
abstract
Provably efficient Model-Based Reinforcement Learning (MBRL) based on optimism or posterior sampling (PSRL) is ensured to attain the global optimality asymptotically by introducing the complexity measure of the model. However, the complexity might grow exponentially for the simplest nonlinear models, where global convergence is impossible within finite iterations. When the model suffers a large generalization error, which is quantitatively measured by the model complexity, the uncertainty can be large. The sampled model that current policy is greedily optimized upon will thus be unsettled, resulting in aggressive policy updates and over-exploration. In this work, we propose Conservative Dual Policy Optimization (CDPO) that involves a Referential Update and a Conservative Update. The policy is first optimized under a reference model, which imitates the mechanism of PSRL while offering more stability. A conservative range of randomness is guaranteed by maximizing the expectation of model value. Without harmful sampling procedures, CDPO can still achieve the same regret as PSRL. More importantly, CDPO enjoys monotonic policy improvement and global optimality simultaneously. Empirical results also validate the exploration efficiency of CDPO.
Shenao Zhang
NeurIPS1