EDBT 2026 Demo / reviewers in the wild / expert
Yunkai Zhang 0002
dblp:34/3300-2
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-7203-2883ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 100% | |
| Artificial intelligence
4 papers |
Reinforcement learning · 56% Optimization for machine learning · 17% Question answering and dialogue systems · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine rescheduling |
1.9 | 2 | 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026 Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Cloud and datacenter computing
cluster resource management and scheduling |
1.1 | 2 | 2026 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026 |
Recommender systems
generative recommendation |
1.0 | 1 | 2026 | Guiding Generative Recommender Systems with Structured Human Priors via Multi-head Decoding · WWW 2026 |
Recommender systems › beyond-accuracy recommendation
novelty and diversity |
1.0 | 1 | 2026 | Guiding Generative Recommender Systems with Structured Human Priors via Multi-head Decoding · WWW 2026 |
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
container placement |
0.8 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Machine learning › Optimization for machine learning
primal-dual methods |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
safe multi-agent reinforcement learning |
0.7 | 1 | 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General Utilities · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis
fact-checking |
0.4 | 1 | 2020 | TabFact: A Large-scale Dataset for Table-based Fact Verification · ICLR 2020 |
Natural language and speech › Question answering and dialogue systems
table-based fact verification |
0.4 | 1 | 2020 | TabFact: A Large-scale Dataset for Table-based Fact Verification · ICLR 2020 |
Machine learning › Graph learning
graph neural network |
0.2 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.2 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Methods — techniques the papers use, named apart from their topics
heuristic migration planning · 2.3graph neural network · 2.3column generation · 2.3deep reinforcement learning · 1.7combinatorial optimization · 1.7mixed integer programming · 1.5two-stage decision-making · 1.0risk-aware evaluation · 1.0reinforcement learning · 1.0multi-head decoding · 1.0adapter tuning · 1.0mixed-integer programming · 0.8shadow reward · 0.7neighbor truncation · 0.7correlation decay · 0.7natural language inference · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Guiding Generative Recommender Systems with Structured Human Priors via Multi-head DecodingabstractOptimizing recommender systems for objectives beyond accuracy, such as diversity, novelty, and personalization, is crucial for long-term user satisfaction. To this end, industrial practitioners have accumulated vast amounts of structured domain knowledge, which we term human priors (e.g., item taxonomies, temporal patterns). This knowledge is typically applied through post-hoc adjustments during ranking or post-ranking. However, this approach remains decoupled from the core model learning, which is particularly undesirable as the industry shifts to end-to-end generative recommendation foundation models. On the other hand, many methods targeting these beyond-accuracy objectives often require architecture-specific modifications and discard these valuable human priors by learning user intent in a fully unsupervised manner. Instead of discarding the human priors accumulated over years of practice, we introduce a backbone-agnostic framework that seamlessly integrates these human priors directly into the end-to-end training of generative recommenders. With lightweight, prior-conditioned adapter heads inspired by efficient LLM decoding strategies, our approach guides the model to disentangle user intent along human-understandable axes (e.g., interaction types, long- vs. short-term interests). We also introduce a hierarchical composition strategy for modeling complex interactions across different prior types. Extensive experiments on three large-scale datasets demonstrate that our method significantly enhances both accuracy and beyond-accuracy objectives. We also show that human priors allow the backbone model to more effectively leverage longer context lengths and larger model sizes. Yunkai Zhang 0002, Diji Yang, Ryan Lin, Ruizhong Qiu, Benyu Zhang, Hanchao Yu, Yinglong Xia, Zhuokai Zhao, Lizhu Zhang, Xiangjun Fan, Zhuoran Yu, Zeyu Zheng 0002 |
WWW | 1 |
| 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data CentersabstractManaging a vast number of virtual machines (VMs) efficiently is a critical challenge in modern large-scale data centers. The continuous creation and termination of VMs lead to resource fragmentation across physical machines (PMs), necessitating periodic VM rescheduling to optimize resource utilization. Despite its significance, VM rescheduling has received limited attention in the literature. A key challenge is that, unlike conventional combinatorial optimization problems, the efficiency of rescheduling algorithms is heavily impacted by inference time, as VM states evolve dynamically during execution. This scalability bottleneck hampers existing methods. To address this, we propose VMR$^{2}$L, a reinforcement learning framework tailored for VM rescheduling. VMR$^{2}$L integrates a two-stage decision-making process to accommodate complex operational constraints, a feature extraction mechanism that captures critical relational information for rescheduling, and a risk-aware evaluation strategy that enables users to balance execution speed and rescheduling accuracy. Extensive experiments using real-world data from a production-scale data center demonstrate that VMR$^{2}$L achieves near-optimal performance while reducing inference time to a matter of seconds. To facilitate reproducibility, we provide access to our implementation and datasets. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement LearningabstractModern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM rescheduling. Despite the increasing importance of VM rescheduling as data centers grow in size, the problem remains understudied. We first show that, unlike most combinatorial optimization tasks, the inference time of VM rescheduling algorithms significantly influences their performance, due to dynamic VM state changes during this period. This causes existing methods to scale poorly. Therefore, we develop a reinforcement learning system for VM rescheduling, VMR2L, which incorporates a set of customized techniques, such as a two-stage framework that accommodates diverse constraints and workload conditions, a feature extraction module that captures relational information specific to rescheduling, as well as a risk-seeking evaluation enabling users to optimize the trade-off between latency and accuracy. We conduct extensive experiments with data from an industry-scale data center. Our results show that VMR2L can achieve a performance comparable to the optimal solution but with a running time of seconds. Code12 and datasets3 are open-sourced. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
EuroSys | 2 |
| 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud EnvironmentsabstractContainerization has garnered substantial favor among cloud service providers. Nevertheless, the notable network overhead incurred between containers has prompted concerns within the community. In cloud resource scheduling, collocating service containers that frequently communicate to the same machine - termed “service affinity” - is instrumental in enhancing application performance. In response to this concern, we present a solution that harnesses service affinity and collocates containers to enhance the overall system performance and stability. To maximize the benefits of collocating containers, it is necessary to calculate a new schedule that optimally and efficiently maximizes service affinity, especially within the expansive domain of industry-scale cloud environments. In pursuit of this, we leverage the skewness property of affinity and machine learning to fuse solver-based algorithms, thereby assuring both quality and efficiency for problems at scale. Our methodology encompasses the partitioning of a given task into discrete subproblems, with a keen focus on resolving the most critical ones. Via a graph neural network classifier, we assign each subproblem to be solved independently using methods based on off-the-shelf solvers in our algorithm pool - namely, MIP-based, or column generation. This strategic approach enables the efficient computation of a schedule for a cloud cluster that fully optimizes the overall service affinity. We further propose a heuristic algorithm to compute executable container migration plans for practical use, facilitating the transition to the new placement where service affinity is well optimized. Our solution has been deployed in our large-scale production environment, covering over a million cores within ByteDance. Through the successful real-world production deployment, our approach exhibits an average improvement in end-to-end latency by 23.75% and a reduction in request error rates by 24.09% compared to the original system. Zuzhi Chen, Fuxin Jiang, Binbin Chen 0005, Yu Li 0003, Yunkai Zhang 0002, Jianjun Chen 0001, Wu Xiang, Guozhu Cheng, Wei Zhang 0172, Tieying Zhang |
ICDE | 5 |
| 2023 | Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General UtilitiesabstractWe investigate safe multi-agent reinforcement learning, where agents seek to collectively maximize an aggregate sum of local objectives while satisfying their own safety constraints. The objective and constraints are described by general utilities, i.e., nonlinear functions of the long-term state-action occupancy measure, which encompass broader decision-making goals such as risk, exploration, or imitations. The exponential growth of the state-action space size with the number of agents presents challenges for global observability, further exacerbated by the global coupling arising from agents' safety constraints. To tackle this issue, we propose a primal-dual method utilizing shadow reward and $\kappa$-hop neighbor truncation under a form of correlation decay property, where $\kappa$ is the communication radius. In the exact setting, our algorithm converges to a first-order stationary point (FOSP) at the rate of $\mathcal{O}\left(T^{-2/3}\right)$. In the sample-based setting, we demonstrate that, with high probability, our algorithm requires $\widetilde{\mathcal{O}}\left(\epsilon^{-3.5}\right)$ samples to achieve an $\epsilon$-FOSP with an approximation error of $\mathcal{O}(\phi_0^{2\kappa})$, where $\phi_0\in (0,1)$. Finally, we demonstrate the effectiveness of our model through extensive numerical experiments. Donghao Ying, Yunkai Zhang 0002, Yuhao Ding, Alec Koppel, Javad Lavaei |
NeurIPS | 2 |
| 2020 | TabFact: A Large-scale Dataset for Table-based Fact Verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 0002, Hong Wang 0023, Xiyou Zhou, William Yang Wang |
ICLR | 4 |