EDBT 2026 Demo / reviewers in the wild / expert
Shutong Ding
dblp:359/0842
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 43% Generative modeling · 25% Deep learning architectures and training · 12% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.4 | 3 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024 Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024 |
Machine learning › Reinforcement learning
policy optimization |
2.3 | 3 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024 Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023 |
Robotics › Robot manipulation
diffusion policy |
1.6 | 2 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024 |
Machine learning › Reinforcement learning › policy optimization
diffusion policy optimization |
0.9 | 1 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 |
Machine learning › Reinforcement learning
on-policy reinforcement learning |
0.9 | 1 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 |
Robotics › Motion planning and robot control
robot learning |
0.9 | 1 | 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement Learning · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.8 | 1 | 2024 | Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024 |
Machine learning › Reinforcement learning › online decision making
online reinforcement learning |
0.8 | 1 | 2024 | Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › guided diffusion
training-free guidance |
0.8 | 1 | 2024 | Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.7 | 1 | 2023 | Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023 |
Machine learning › Reinforcement learning
continuous control |
0.7 | 1 | 2023 | Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model |
0.7 | 1 | 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023 |
Machine learning › Reinforcement learning › constrained reinforcement learning
hard constraints |
0.7 | 1 | 2023 | Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023 |
Machine learning › Optimization for machine learning
homotopy methods |
0.7 | 1 | 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › deep generative model
implicit models |
0.7 | 1 | 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023 |
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
0.7 | 1 | 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy Continuation · NeurIPS 2023 |
Mathematical optimization
constrained optimization |
0.4 | 2 | 2024 | Guidance with Spherical Gaussian Constraint for Conditional Diffusion · ICML 2024 Reduced Policy Optimization for Continuous Control with Hard Constraints · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
spherical gaussian constraint · 1.5manifold deviation analysis · 1.5lagrangian relaxation · 1.3generalized reduced gradient · 1.3action projection · 1.3exact diffusion inversion · 0.9doubled dummy action mechanism · 0.9PPO · 0.9q-weighted variational loss · 0.8entropy regularization · 0.8behavior policy · 0.8newton's method · 0.7homotopy continuation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GenPO: Generative Diffusion Models Meet On-Policy Reinforcement LearningabstractRecent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks like PPO remains underexplored. This gap is particularly significant given the widespread use of large-scale parallel GPU-accelerated simulators, such as IsaacLab, which are optimized for on-policy RL algorithms and enable rapid training of complex robotic tasks. A key challenge lies in computing state-action log-likelihoods under diffusion policies, which is straightforward for Gaussian policies but intractable for flow-based models due to irreversible forward-reverse processes and discretization errors (e.g., Euler-Maruyama approximations). To bridge this gap, we propose GenPO, a generative policy optimization framework that leverages exact diffusion inversion to construct invertible action mappings. GenPO introduces a novel doubled dummy action mechanism that enables invertibility via alternating updates, resolving log-likelihood computation barriers. Furthermore, we also use the action log-likelihood for unbiased entropy and KL divergence estimation, enabling KL-adaptive learning rates and entropy regularization in on-policy updates. Extensive experiments on eight IsaacLab benchmarks, including legged locomotion (Ant, Humanoid, Anymal-D, Unitree H1, Go2), dexterous manipulation (Shadow Hand), aerial control (Quadcopter), and robotic arm tasks (Franka), demonstrate GenPO’s superiority over existing RL baselines. Notably, GenPO is the first method to successfully integrate diffusion policies into on-policy RL, unlocking their potential for large-scale parallelized training and real-world robotic deployment. Shutong Ding, Haoyang Luo, Weinan Zhang 0001, Jingya Wang 0001, Ye Shi 0001 |
NeurIPS | 1 |
| 2024 | Guidance with Spherical Gaussian Constraint for Conditional DiffusionabstractRecent advances in diffusion models attempt to handle conditional generative tasks by utilizing a differentiable loss function for guidance without the need for additional training. While these methods achieved certain success, they often compromise on sample quality and require small guidance step sizes, leading to longer sampling processes. This paper reveals that the fundamental issue lies in the manifold deviation during the sampling process when loss guidance is employed. We theoretically show the existence of manifold deviation by establishing a certain lower bound for the estimation error of the loss guidance. To mitigate this problem, we propose Diffusion with Spherical Gaussian constraint (DSG), drawing inspiration from the concentration phenomenon in high-dimensional Gaussian distributions. DSG effectively constrains the guidance step within the intermediate data manifold through optimization and enables the use of larger guidance steps. Furthermore, we present a closed-form solution for DSG denoising with the Spherical Gaussian constraint. Notably, DSG can seamlessly integrate as a plugin module within existing training-free conditional diffusion methods. Implementing DSG merely involves a few lines of additional code with almost no extra computational overhead, yet it leads to significant performance improvements. Comprehensive experimental results in various conditional generation tasks validate the superiority and adaptability of DSG in terms of both sample quality and time efficiency. Lingxiao Yang, Shutong Ding, Jingyi Yu 0001, Jingya Wang 0001, Ye Shi 0001 |
ICML | 2 |
| 2024 | Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationabstractDiffusion models have garnered widespread attention in Reinforcement Learning (RL) for their powerful expressiveness and multimodality. It has been verified that utilizing diffusion policies can significantly improve the performance of RL algorithms in continuous control tasks by overcoming the limitations of unimodal policies, such as Gaussian policies. Furthermore, the multimodality of diffusion policies also shows the potential of providing the agent with enhanced exploration capabilities. However, existing works mainly focus on applying diffusion policies in offline RL, while their incorporation into online RL has been less investigated. The diffusion model's training objective, known as the variational lower bound, cannot be applied directly in online RL due to the unavailability of 'good' samples (actions). To harmonize the diffusion model with online RL, we propose a novel model-free diffusion-based online RL algorithm named Q-weighted Variational Policy Optimization (QVPO). Specifically, we introduce the Q-weighted variational loss and its approximate implementation in practice. Notably, this loss is shown to be a tight lower bound of the policy objective. To further enhance the exploration capability of the diffusion policy, we design a special entropy regularization term. Unlike Gaussian policies, the log-likelihood in diffusion policies is inaccessible; thus this entropy term is nontrivial. Moreover, to reduce the large variance of diffusion policies, we also develop an efficient behavior policy through action selection. This can further improve its sample efficiency during online interaction. Consequently, the QVPO algorithm leverages the exploration capabilities and multimodality of diffusion policies, preventing the RL agent from converging to a sub-optimal policy. To verify the effectiveness of QVPO, we conduct comprehensive experiments on MuJoCo continuous control benchmarks. The final results demonstrate that QVPO achieves state-of-the-art performance in terms of both cumulative reward and sample efficiency. Shutong Ding, Kan Ren, Weinan Zhang 0001, Jingyi Yu 0001, Jingya Wang 0001, Ye Shi 0001 |
NeurIPS | 1 |
| 2023 | Two Sides of The Same Coin: Bridging Deep Equilibrium Models and Neural ODEs via Homotopy ContinuationabstractDeep Equilibrium Models (DEQs) and Neural Ordinary Differential Equations (Neural ODEs) are two branches of implicit models that have achieved remarkable success owing to their superior performance and low memory consumption. While both are implicit models, DEQs and Neural ODEs are derived from different mathematical formulations. Inspired by homotopy continuation, we establish a connection between these two models and illustrate that they are actually two sides of the same coin. Homotopy continuation is a classical method of solving nonlinear equations based on a corresponding ODE. Given this connection, we proposed a new implicit model called HomoODE that inherits the property of high accuracy from DEQs and the property of stability from Neural ODEs. Unlike DEQs, which explicitly solve an equilibrium-point-finding problem via Newton's methods in the forward pass, HomoODE solves the equilibrium-point-finding problem implicitly using a modified Neural ODE via homotopy continuation. Further, we developed an acceleration method for HomoODE with a shared learnable initial point. It is worth noting that our model also provides a better understanding of why Augmented Neural ODEs work as long as the augmented part is regarded as the equilibrium point to find. Comprehensive experiments with several image classification tasks demonstrate that HomoODE surpasses existing implicit models in terms of both accuracy and memory consumption. Shutong Ding, Tianyu Cui, Jingya Wang 0001, Ye Shi 0001 |
NeurIPS | 1 |
| 2023 | Reduced Policy Optimization for Continuous Control with Hard ConstraintsabstractRecent advances in constrained reinforcement learning (RL) have endowed reinforcement learning with certain safety guarantees. However, deploying existing constrained RL algorithms in continuous control tasks with general hard constraints remains challenging, particularly in those situations with non-convex hard constraints. Inspired by the generalized reduced gradient (GRG) algorithm, a classical constrained optimization technique, we propose a reduced policy optimization (RPO) algorithm that combines RL with GRG to address general hard constraints. RPO partitions actions into basic actions and nonbasic actions following the GRG method and outputs the basic actions via a policy network. Subsequently, RPO calculates the nonbasic actions by solving equations based on equality constraints using the obtained basic actions. The policy network is then updated by implicitly differentiating nonbasic actions with respect to basic actions. Additionally, we introduce an action projection procedure based on the reduced gradient and apply a modified Lagrangian relaxation technique to ensure inequality constraints are satisfied. To the best of our knowledge, RPO is the first attempt that introduces GRG to RL as a way of efficiently handling both equality and inequality hard constraints. It is worth noting that there is currently a lack of RL environments with complex hard constraints, which motivates us to develop three new benchmarks: two robotics manipulation tasks and a smart grid operation control task. With these benchmarks, RPO achieves better performance than previous constrained RL algorithms in terms of both cumulative reward and constraint violation. We believe RPO, along with the new benchmarks, will open up new opportunities for applying RL to real-world problems with complex constraints. Shutong Ding, Jingya Wang 0001, Yali Du 0001, Ye Shi 0001 |
NeurIPS | 1 |