VLDB 2026 Research / reviewers in the wild / expert
Donghao Ying
dblp:304/3590
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0001-7329-5917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 86% Optimization for machine learning · 10% Generative modeling · 4% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 100% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 75% Mathematical optimization · 25% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine rescheduling |
1.9 | 2 | 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026 Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.9 | 1 | 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025 |
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.9 | 1 | 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.7 | 1 | 2023 | Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes · AAAI 2023 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Machine learning › Optimization for machine learning
primal-dual methods |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
safe multi-agent reinforcement learning |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Algorithmic game theory and mechanism design
dynamic pricing |
0.7 | 1 | 2023 | No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023 |
Algorithmic game theory and mechanism design
equilibrium computation |
0.7 | 1 | 2023 | No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023 |
Mathematical optimization
primal-dual method |
0.7 | 1 | 2023 | Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision Processes · AAAI 2023 |
Algorithmic game theory and mechanism design
regret minimization |
0.7 | 1 | 2023 | No-Regret Learning in Dynamic Competition with Reference Effects Under Logit Demand · NeurIPS 2023 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.3 | 1 | 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025 |
Machine learning › Reinforcement learning › policy learning
policy extraction |
0.3 | 1 | 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
deep reinforcement learning · 1.7combinatorial optimization · 1.7projected subgradient descent · 1.3policy gradient · 1.3two-stage decision-making · 1.0risk-aware evaluation · 1.0reinforcement learning · 1.0gradient manipulation · 0.9diffusion model · 0.9shadow reward · 0.7online projected gradient ascent · 0.7neighbor truncation · 0.7multinomial logit model · 0.7correlation decay · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data CentersabstractManaging a vast number of virtual machines (VMs) efficiently is a critical challenge in modern large-scale data centers. The continuous creation and termination of VMs lead to resource fragmentation across physical machines (PMs), necessitating periodic VM rescheduling to optimize resource utilization. Despite its significance, VM rescheduling has received limited attention in the literature. A key challenge is that, unlike conventional combinatorial optimization problems, the efficiency of rescheduling algorithms is heavily impacted by inference time, as VM states evolve dynamically during execution. This scalability bottleneck hampers existing methods. To address this, we propose VMR$^{2}$L, a reinforcement learning framework tailored for VM rescheduling. VMR$^{2}$L integrates a two-stage decision-making process to accommodate complex operational constraints, a feature extraction mechanism that captures critical relational information for rescheduling, and a risk-aware evaluation strategy that enables users to balance execution speed and rescheduling accuracy. Extensive experiments using real-world data from a production-scale data center demonstrate that VMR$^{2}$L achieves near-optimal performance while reducing inference time to a matter of seconds. To facilitate reproducibility, we provide access to our implementation and datasets. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement LearningabstractModern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM rescheduling. Despite the increasing importance of VM rescheduling as data centers grow in size, the problem remains understudied. We first show that, unlike most combinatorial optimization tasks, the inference time of VM rescheduling algorithms significantly influences their performance, due to dynamic VM state changes during this period. This causes existing methods to scale poorly. Therefore, we develop a reinforcement learning system for VM rescheduling, VMR2L, which incorporates a set of customized techniques, such as a two-stage framework that accommodates diverse constraints and workload conditions, a feature extraction module that captures relational information specific to rescheduling, as well as a risk-seeking evaluation enabling users to optimize the trade-off between latency and accuracy. We conduct extensive experiments with data from an industry-scale data center. Our results show that VMR2L can achieve a performance comparable to the optimal solution but with a running time of seconds. Code12 and datasets3 are open-sourced. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
EuroSys | 4 |
| 2025 | Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RLabstractConstrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent learns from a fixed dataset—a common requirement in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constrained Offline Reinforcement Learning (DRCORL), which first uses a diffusion model to capture the behavioral policy from offline data and then extracts a simplified policy to enable efficient inference. We further apply gradient manipulation for safety adaptation, balancing the reward objective and constraint satisfaction. This approach leverages high-quality offline data while incorporating safety requirements. Empirical results show that DRCORL achieves reliable safety performance, fast inference, and strong reward outcomes across robot learning tasks. Compared to existing safe offline RL methods, it consistently meets cost limits and performs well with the same hyperparameters, indicating practical applicability in real-world scenarios. We open-source our implementation at https://github.com/JamesJunyuGuo/DRCORL. Donghao Ying, Ming Jin 0002, Shangding Gu, Costas J. Spanos, Javad Lavaei |
NeurIPS | 3 |
| 2025 | Subsampled Ensemble Can Improve Generalization Tail ExponentiallyabstractEnsemble learning is a popular technique to improve the accuracy of machine learning models. It traditionally hinges on the rationale that aggregating multiple weak models can lead to better models with lower variance and hence higher stability, especially for discontinuous base learners. In this paper, we provide a new perspective on ensembling. By selecting the most frequently generated model from the base learner when repeatedly applied to subsamples, we can attain exponentially decaying tails for the excess risk, even if the base learner suffers from slow (i.e., polynomial) decay rates. This tail enhancement power of ensembling applies to base learners that have reasonable predictive power to begin with and is stronger than variance reduction in the sense of exhibiting rate improvement. We demonstrate how our ensemble methods can substantially improve out-of-sample performances in a range of numerical examples involving heavy-tailed data or intrinsically slow rates. Huajie Qian, Donghao Ying, Henry Lam, Wotao Yin |
NeurIPS | 2 |
| 2025 | Policy-based Primal-Dual Methods for Concave CMDP with Variance ReductionabstractWe study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an O(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments. Donghao Ying, Mengzi Guo, Hyunin Lee, Yuhao Ding, Javad Lavaei, Zuo-Jun Max Shen |
J. Artif. Intell. Res. | 1 |
| 2023 | Policy-Based Primal-Dual Methods for Convex Constrained Markov Decision ProcessesabstractWe study convex Constrained Markov Decision Processes (CMDPs) in which the objective is concave and the constraints are convex in the state-action occupancy measure. We propose a policy-based primal-dual algorithm that updates the primal variable via policy gradient ascent and updates the dual variable via projected sub-gradient descent. Despite the loss of additivity structure and the nonconvex nature, we establish the global convergence of the proposed algorithm by leveraging a hidden convexity in the problem, and prove the O(T^-1/3) convergence rate in terms of both optimality gap and constraint violation. When the objective is strongly concave in the occupancy measure, we prove an improved convergence rate of O(T^-1/2). By introducing a pessimistic term to the constraint, we further show that a zero constraint violation can be achieved while preserving the same convergence rate for the optimality gap. This work is the first one in the literature that establishes non-asymptotic convergence guarantees for policy-based primal-dual methods for solving infinite-horizon discounted convex CMDPs. Donghao Ying, Mengzi Guo, Yuhao Ding, Javad Lavaei, Zuo-Jun Max Shen |
AAAI | 1 |
| 2023 | No-Regret Learning in Dynamic Competition with Reference Effects Under Logit DemandabstractThis work is dedicated to the algorithm design in a competitive framework, with the primary goal of learning a stable equilibrium. We consider the dynamic price competition between two firms operating within an opaque marketplace, where each firm lacks information about its competitor. The demand follows the multinomial logit (MNL) choice model, which depends on the consumers' observed price and their reference price, and consecutive periods in the repeated games are connected by reference price updates. We use the notion of stationary Nash equilibrium (SNE), defined as the fixed point of the equilibrium pricing policy for the single-period game, to simultaneously capture the long-run market equilibrium and stability. We propose the online projected gradient ascent algorithm (OPGA), where the firms adjust prices using the first-order derivatives of their log-revenues that can be obtained from the market feedback mechanism. Despite the absence of typical properties required for the convergence of online games, such as strong monotonicity and variational stability, we demonstrate that under diminishing step-sizes, the price and reference price paths generated by OPGA converge to the unique SNE, thereby achieving the no-regret learning and a stable market. Moreover, with appropriate step-sizes, we prove that this convergence exhibits a rate of $\mathcal{O}(1/t)$. Mengzi Guo, Donghao Ying, Javad Lavaei, Zuo-Jun Max Shen |
NeurIPS | 2 |
| 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General UtilitiesabstractWe investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints. The objective and constraints are described by general utilities, i.e., nonlinear functions of the long-term state-action occupancy measure, which encompass broader decision-making goals such as risk, exploration, or imitations. The exponential growth of the state-action space size with the number of agents presents challenges for global observability, further exacerbated by the global coupling arising from agents' safety constraints. To tackle this issue, we propose a primal-dual method utilizing shadow reward and $\kappa$-hop neighbor truncation under a form of correlation decay property, where $\kappa$ is the communication radius. In the exact setting, our algorithm converges to a first-order stationary point (FOSP) at the rate of $\mathcal{O}\left(T^{-2/3}\right)$. In the sample-based setting, we demonstrate that, with high probability, our algorithm requires $\widetilde{\mathcal{O}}\left(\epsilon^{-3.5}\right)$ samples to achieve an $\epsilon$-FOSP with an approximation error of $\mathcal{O}(\phi_0^{2\kappa})$, where $\phi_0\in (0,1)$. Finally, we demonstrate the effectiveness of our model through extensive numerical experiments. Donghao Ying, Yunkai Zhang 0002, Yuhao Ding, Alec Koppel, Javad Lavaei |
NeurIPS | 1 |
| 2022 | A Dual Approach to Constrained Markov Decision Processes with Entropy RegularizationabstractWe study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total utility. By leveraging the entropy regularization, our theoretical analysis shows that its Lagrangian dual function is smooth and the Lagrangian duality gap can be decomposed into the primal optimality gap and the constraint violation. Furthermore, we propose an accelerated dual-descent method for entropy-regularized CMDPs. We prove that our method achieves the global convergence rate $\widetilde{\mathcal{O}}(1/T)$ for both the optimality gap and the constraint violation for entropy-regularized CMDPs. A discussion about a linear convergence rate for CMDPs with a single constraint is also provided. Donghao Ying, Yuhao Ding, Javad Lavaei |
AISTATS | 1 |