Shutong Ding

dblp:359/0842 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 43% Generative modeling · 25% Deep learning architectures and training · 12%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.432025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024
Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024
Machine learning › Reinforcement learning
policy optimization
2.332025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024
Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023
Robotics › Robot manipulation
diffusion policy
1.622025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024
Machine learning › Reinforcement learning › policy optimization
diffusion policy optimization
0.912025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
on-policy reinforcement learning
0.912025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Robotics › Motion planning and robot control
robot learning
0.912025
GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.812024
Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024
Machine learning › Reinforcement learning › online decision making
online reinforcement learning
0.812024
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance
0.812024
Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024
Machine learning › Reinforcement learning
constrained reinforcement learning
0.712023
Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023
Machine learning › Reinforcement learning
continuous control
0.712023
Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model
0.712023
Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023
Machine learning › Reinforcement learning › constrained reinforcement learning
hard constraints
0.712023
Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023
Machine learning › Optimization for machine learning
homotopy methods
0.712023
Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023
Machine learning › Deep learning architectures and training › deep generative model
implicit models
0.712023
Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.712023
Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023
Mathematical optimization
constrained optimization
0.422024
Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024
Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

spherical gaussian constraint · 1.5manifold deviation analysis · 1.5lagrangian relaxation · 1.3generalized reduced gradient · 1.3action projection · 1.3exact diffusion inversion · 0.9doubled dummy action mechanism · 0.9PPO · 0.9q-weighted variational loss · 0.8entropy regularization · 0.8behavior policy · 0.8newton's method · 0.7homotopy continuation · 0.7
YearPublicationVenuePosition
2025 GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning
abstract
Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks like PPO remains underexplored. This gap is particularly significant given the widespread use of large-scale parallel GPU-accelerated simulators, such as IsaacLab, which are optimized for on-policy RL algorithms and enable rapid training of complex robotic tasks. A key challenge lies in computing state-action log-likelihoods under diffusion policies, which is straightforward for Gaussian policies but intractable for flow-based models due to irreversible forward-reverse processes and discretization errors (e.g., Euler-Maruyama approximations). To bridge this gap, we propose GenPO, a generative policy optimization framework that leverages exact diffusion inversion to construct invertible action mappings. GenPO introduces a novel doubled dummy action mechanism that enables invertibility via alternating updates, resolving log-likelihood computation barriers. Furthermore, we also use the action log-likelihood for unbiased entropy and KL divergence estimation, enabling KL-adaptive learning rates and entropy regularization in on-policy updates. Extensive experiments on eight IsaacLab benchmarks, including legged locomotion (Ant, Humanoid, Anymal-D, Unitree H1, Go2), dexterous manipulation (Shadow Hand), aerial control (Quadcopter), and robotic arm tasks (Franka), demonstrate GenPO’s superiority over existing RL baselines. Notably, GenPO is the first method to successfully integrate diffusion policies into on-policy RL, unlocking their potential for large-scale parallelized training and real-world robotic deployment.
Shutong Ding, Haoyang Luo, Weinan Zhang 0001, Jingya Wang 0001, Ye Shi 0001
NeurIPS1
2024 Guidance with Spherical Gaussian Constraint for Conditional Diffusion
abstract
Recent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step sizes, leading to longer sampling processes. This paper reveals that the fundamental issue lies in the manifold deviation during the sampling process when loss guidance is employed. We theoretically show the existence of manifold deviation by establishing a certain lower bound for the estimation error of the loss guidance. To mitigate this problem, we propose Diffusion with Spherical Gaussian constraint (DSG), drawing inspiration from the concentration phenomenon in high-dimensional Gaussian distributions. DSG effectively constrains the guidance step within the intermediate data manifold through optimization and enables the use of larger guidance steps. Furthermore, we present a closed-form solution for DSG denoising with the Spherical Gaussian constraint. Notably, DSG can seamlessly integrate as a plugin module within existing training-free conditional diffusion methods. Implementing DSG merely involves a few lines of additional code with almost no extra computational overhead, yet it leads to significant performance improvements. Comprehensive experimental results in various conditional generation tasks validate the superiority and adaptability of DSG in terms of both sample quality and time efficiency.
Lingxiao Yang, Shutong Ding, Jingyi Yu 0001, Jingya Wang 0001, Ye Shi 0001
ICML2
2024 Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
abstract
Diffusion models have garnered widespread attention in Reinforcement Learning (RL) for their powerful expressiveness and multimodality. It has been verified that utilizing diffusion policies can significantly improve the performance of RL algorithms in continuous control tasks by overcoming the limitations of unimodal policies, such as Gaussian policies. Furthermore, the multimodality of diffusion policies also shows the potential of providing the agent with enhanced exploration capabilities. However, existing works mainly focus on applying diffusion policies in offline RL, while their incorporation into online RL has been less investigated. The diffusion model's training objective, known as the variational lower bound, cannot be applied directly in online RL due to the unavailability of 'good' samples (actions). To harmonize the diffusion model with online RL, we propose a novel model-free diffusion-based online RL algorithm named Q-weighted Variational Policy Optimization (QVPO). Specifically, we introduce the Q-weighted variational loss and its approximate implementation in practice. Notably, this loss is shown to be a tight lower bound of the policy objective. To further enhance the exploration capability of the diffusion policy, we design a special entropy regularization term. Unlike Gaussian policies, the log-likelihood in diffusion policies is inaccessible; thus this entropy term is nontrivial. Moreover, to reduce the large variance of diffusion policies, we also develop an efficient behavior policy through action selection. This can further improve its sample efficiency during online interaction. Consequently, the QVPO algorithm leverages the exploration capabilities and multimodality of diffusion policies, preventing the RL agent from converging to a sub-optimal policy. To verify the effectiveness of QVPO, we conduct comprehensive experiments on MuJoCo continuous control benchmarks. The final results demonstrate that QVPO achieves state-of-the-art performance in terms of both cumulative reward and sample efficiency.
Shutong Ding, Kan Ren, Weinan Zhang 0001, Jingyi Yu 0001, Jingya Wang 0001, Ye Shi 0001
NeurIPS1
2023 Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation
abstract
Deep Equilibrium Models (DEQs) and Neural Ordinary Differential Equations (Neural ODEs) are two branches of implicit models that have achieved remarkable success owing to their superior performance and low memory consumption. While both are implicit models, DEQs and Neural ODEs are derived from different mathematical formulations. Inspired by homotopy continuation, we establish a connection between these two models and illustrate that they are actually two sides of the same coin. Homotopy continuation is a classical method of solving nonlinear equations based on a corresponding ODE. Given this connection, we proposed a new implicit model called HomoODE that inherits the property of high accuracy from DEQs and the property of stability from Neural ODEs. Unlike DEQs, which explicitly solve an equilibrium-point-finding problem via Newton's methods in the forward pass, HomoODE solves the equilibrium-point-finding problem implicitly using a modified Neural ODE via homotopy continuation. Further, we developed an acceleration method for HomoODE with a shared learnable initial point. It is worth noting that our model also provides a better understanding of why Augmented Neural ODEs work as long as the augmented part is regarded as the equilibrium point to find. Comprehensive experiments with several image classification tasks demonstrate that HomoODE surpasses existing implicit models in terms of both accuracy and memory consumption.
Shutong Ding, Tianyu Cui, Jingya Wang 0001, Ye Shi 0001
NeurIPS1
2023 Reduced Policy Optimization for Continuous Control with Hard Constraints
abstract
Recent advances in constrained reinforcement learning (RL) have endowed reinforcement learning with certain safety guarantees. However, deploying existing constrained RL algorithms in continuous control tasks with general hard constraints remains challenging, particularly in those situations with non-convex hard constraints. Inspired by the generalized reduced gradient (GRG) algorithm, a classical constrained optimization technique, we propose a reduced policy optimization (RPO) algorithm that combines RL with GRG to address general hard constraints. RPO partitions actions into basic actions and nonbasic actions following the GRG method and outputs the basic actions via a policy network. Subsequently, RPO calculates the nonbasic actions by solving equations based on equality constraints using the obtained basic actions. The policy network is then updated by implicitly differentiating nonbasic actions with respect to basic actions. Additionally, we introduce an action projection procedure based on the reduced gradient and apply a modified Lagrangian relaxation technique to ensure inequality constraints are satisfied. To the best of our knowledge, RPO is the first attempt that introduces GRG to RL as a way of efficiently handling both equality and inequality hard constraints. It is worth noting that there is currently a lack of RL environments with complex hard constraints, which motivates us to develop three new benchmarks: two robotics manipulation tasks and a smart grid operation control task. With these benchmarks, RPO achieves better performance than previous constrained RL algorithms in terms of both cumulative reward and constraint violation. We believe RPO, along with the new benchmarks, will open up new opportunities for applying RL to real-world problems with complex constraints.
Shutong Ding, Jingya Wang 0001, Yali Du 0001, Ye Shi 0001
NeurIPS1