Donghao Ying

dblp:304/3590 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0001-7329-5917ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 86% Optimization for machine learning · 10% Generative modeling · 4%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 100%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 75% Mathematical optimization · 25%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine rescheduling
1.922026
Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026
Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025
Machine learning › Reinforcement learning
constrained reinforcement learning
0.912025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems
0.912025
Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025
Machine learning › Reinforcement learning
safe reinforcement learning
0.912025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.912025
Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.712023
Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes · AAAI 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023
Machine learning › Optimization for machine learning
primal-dual methods
0.712023
Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023
Machine learning › Reinforcement learning › multi-agent reinforcement learning
safe multi-agent reinforcement learning
0.712023
Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023
Algorithmic game theory and mechanism design
dynamic pricing
0.712023
No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023
Algorithmic game theory and mechanism design
equilibrium computation
0.712023
No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023
Mathematical optimization
primal-dual method
0.712023
Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes · AAAI 2023
Algorithmic game theory and mechanism design
regret minimization
0.712023
No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.312026
Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026
Machine learning › Generative modeling
diffusion model
0.312025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025
Machine learning › Reinforcement learning › policy learning
policy extraction
0.312025
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

deep reinforcement learning · 1.7combinatorial optimization · 1.7projected subgradient descent · 1.3policy gradient · 1.3two-stage decision-making · 1.0risk-aware evaluation · 1.0reinforcement learning · 1.0gradient manipulation · 0.9diffusion model · 0.9shadow reward · 0.7online projected gradient ascent · 0.7neighbor truncation · 0.7multinomial logit model · 0.7correlation decay · 0.7
YearPublicationVenuePosition
2026 Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers
abstract
Managing a vast number of virtual machines (VMs) efficiently is a critical challenge in modern large-scale data centers. The continuous creation and termination of VMs lead to resource fragmentation across physical machines (PMs), necessitating periodic VM rescheduling to optimize resource utilization. Despite its significance, VM rescheduling has received limited attention in the literature. A key challenge is that, unlike conventional combinatorial optimization problems, the efficiency of rescheduling algorithms is heavily impacted by inference time, as VM states evolve dynamically during execution. This scalability bottleneck hampers existing methods. To address this, we propose VMR$^{2}$L, a reinforcement learning framework tailored for VM rescheduling. VMR$^{2}$L integrates a two-stage decision-making process to accommodate complex operational constraints, a feature extraction mechanism that captures critical relational information for rescheduling, and a risk-aware evaluation strategy that enables users to balance execution speed and rescheduling accuracy. Extensive experiments using real-world data from a production-scale data center demonstrate that VMR$^{2}$L achieves near-optimal performance while reducing inference time to a matter of seconds. To facilitate reproducibility, we provide access to our implementation and datasets.
Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du
IEEE Trans. Parallel Distributed Syst.4
2025 Towards VM Rescheduling Optimization Through Deep Reinforcement Learning
abstract
Modern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM rescheduling. Despite the increasing importance of VM rescheduling as data centers grow in size, the problem remains understudied. We first show that, unlike most combinatorial optimization tasks, the inference time of VM rescheduling algorithms significantly influences their performance, due to dynamic VM state changes during this period. This causes existing methods to scale poorly. Therefore, we develop a reinforcement learning system for VM rescheduling, VMR2L, which incorporates a set of customized techniques, such as a two-stage framework that accommodates diverse constraints and workload conditions, a feature extraction module that captures relational information specific to rescheduling, as well as a risk-seeking evaluation enabling users to optimize the trade-off between latency and accuracy. We conduct extensive experiments with data from an industry-scale data center. Our results show that VMR2L can achieve a performance comparable to the optimal solution but with a running time of seconds. Code12 and datasets3 are open-sourced.
Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du
EuroSys4
2025 Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
abstract
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent learns from a fixed dataset—a common requirement in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constrained Offline Reinforcement Learning (DRCORL), which first uses a diffusion model to capture the behavioral policy from offline data and then extracts a simplified policy to enable efficient inference. We further apply gradient manipulation for safety adaptation, balancing the reward objective and constraint satisfaction. This approach leverages high-quality offline data while incorporating safety requirements. Empirical results show that DRCORL achieves reliable safety performance, fast inference, and strong reward outcomes across robot learning tasks. Compared to existing safe offline RL methods, it consistently meets cost limits and performs well with the same hyperparameters, indicating practical applicability in real-world scenarios. We open-source our implementation at https://github.com/JamesJunyuGuo/DRCORL.
Donghao Ying, Ming Jin 0002, Shangding Gu, Costas J. Spanos, Javad Lavaei
NeurIPS3
2025 Subsampled Ensemble Can Improve Generalization Tail Exponentially
abstract
Ensemble learning is a popular technique to improve the accuracy of machine learning models. It traditionally hinges on the rationale that aggregating multiple weak models can lead to better models with lower variance and hence higher stability, especially for discontinuous base learners. In this paper, we provide a new perspective on ensembling. By selecting the most frequently generated model from the base learner when repeatedly applied to subsamples, we can attain exponentially decaying tails for the excess risk, even if the base learner suffers from slow (i.e., polynomial) decay rates. This tail enhancement power of ensembling applies to base learners that have reasonable predictive power to begin with and is stronger than variance reduction in the sense of exhibiting rate improvement. We demonstrate how our ensemble methods can substantially improve out-of-sample performances in a range of numerical examples involving heavy-tailed data or intrinsically slow rates.
Huajie Qian, Donghao Ying, Henry Lam, Wotao Yin
NeurIPS2
2025 Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction
abstract
We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an O(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments.
Donghao Ying, Mengzi Guo, Hyunin Lee, Yuhao Ding, Javad Lavaei, Zuo-Jun Max Shen
J. Artif. Intell. Res.1
2023 Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes
abstract
We study convex Constrained Markov Decision Processes (CMDPs) in which the objective is concave and the constraints are convex in the state-action occupancy measure. We propose a policy-based primal-dual algorithm that updates the primal variable via policy gradient ascent and updates the dual variable via projected sub-gradient descent. Despite the loss of additivity structure and the nonconvex nature, we establish the global convergence of the proposed algorithm by leveraging a hidden convexity in the problem, and prove the O(T^-1/3) convergence rate in terms of both optimality gap and constraint violation. When the objective is strongly concave in the occupancy measure, we prove an improved convergence rate of O(T^-1/2). By introducing a pessimistic term to the constraint, we further show that a zero constraint violation can be achieved while preserving the same convergence rate for the optimality gap. This work is the first one in the literature that establishes non-asymptotic convergence guarantees for policy-based primal-dual methods for solving infinite-horizon discounted convex CMDPs.
Donghao Ying, Mengzi Guo, Yuhao Ding, Javad Lavaei, Zuo-Jun Max Shen
AAAI1
2023 No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand
abstract
This work is dedicated to the algorithm design in a competitive framework, with the primary goal of learning a stable equilibrium. We consider the dynamic price competition between two firms operating within an opaque marketplace, where each firm lacks information about its competitor. The demand follows the multinomial logit (MNL) choice model, which depends on the consumers' observed price and their reference price, and consecutive periods in the repeated games are connected by reference price updates. We use the notion of stationary Nash equilibrium (SNE), defined as the fixed point of the equilibrium pricing policy for the single-period game, to simultaneously capture the long-run market equilibrium and stability. We propose the online projected gradient ascent algorithm (OPGA), where the firms adjust prices using the first-order derivatives of their log-revenues that can be obtained from the market feedback mechanism. Despite the absence of typical properties required for the convergence of online games, such as strong monotonicity and variational stability, we demonstrate that under diminishing step-sizes, the price and reference price paths generated by OPGA converge to the unique SNE, thereby achieving the no-regret learning and a stable market. Moreover, with appropriate step-sizes, we prove that this convergence exhibits a rate of $\mathcal{O}(1/t)$.
Mengzi Guo, Donghao Ying, Javad Lavaei, Zuo-Jun Max Shen
NeurIPS2
2023 Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities
abstract
We investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints. The objective and constraints are described by general utilities, i.e., nonlinear functions of the long-term state-action occupancy measure, which encompass broader decision-making goals such as risk, exploration, or imitations. The exponential growth of the state-action space size with the number of agents presents challenges for global observability, further exacerbated by the global coupling arising from agents' safety constraints. To tackle this issue, we propose a primal-dual method utilizing shadow reward and $\kappa$-hop neighbor truncation under a form of correlation decay property, where $\kappa$ is the communication radius. In the exact setting, our algorithm converges to a first-order stationary point (FOSP) at the rate of $\mathcal{O}\left(T^{-2/3}\right)$. In the sample-based setting, we demonstrate that, with high probability, our algorithm requires $\widetilde{\mathcal{O}}\left(\epsilon^{-3.5}\right)$ samples to achieve an $\epsilon$-FOSP with an approximation error of $\mathcal{O}(\phi_0^{2\kappa})$, where $\phi_0\in (0,1)$. Finally, we demonstrate the effectiveness of our model through extensive numerical experiments.
Donghao Ying, Yunkai Zhang 0002, Yuhao Ding, Alec Koppel, Javad Lavaei
NeurIPS1
2022 A Dual Approach to Constrained Markov Decision Processes with Entropy Regularization
abstract
We study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total utility. By leveraging the entropy regularization, our theoretical analysis shows that its Lagrangian dual function is smooth and the Lagrangian duality gap can be decomposed into the primal optimality gap and the constraint violation. Furthermore, we propose an accelerated dual-descent method for entropy-regularized CMDPs. We prove that our method achieves the global convergence rate $\widetilde{\mathcal{O}}(1/T)$ for both the optimality gap and the constraint violation for entropy-regularized CMDPs. A discussion about a linear convergence rate for CMDPs with a single constraint is also provided.
Donghao Ying, Yuhao Ding, Javad Lavaei
AISTATS1