EDBT 2026 Demo / reviewers in the wild / expert
Binbin Chen 0005
dblp:38/8396-5
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-8598-2442ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 89% Reconfigurable computing and FPGAs · 7% Performance modeling and evaluation · 2% | |
| Artificial intelligence
3 papers |
Reinforcement learning · 95% Graph learning · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Energy systems and smart grids · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
cluster resource management and scheduling |
2.6 | 4 | 2026 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job Scheduling · Proc. VLDB Endow. 2024 Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024 Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine rescheduling |
1.9 | 2 | 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data Centers · IEEE Trans. Parallel Distributed Syst. 2026 Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
implicit communication |
0.9 | 1 | 2025 | Learning to Communicate Through Implicit Communication Channels · ICLR 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
learned communication protocol |
0.9 | 1 | 2025 | Learning to Communicate Through Implicit Communication Channels · ICLR 2025 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
multi-agent communication |
0.9 | 1 | 2025 | Learning to Communicate Through Implicit Communication Channels · ICLR 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | Learning to Communicate Through Implicit Communication Channels · ICLR 2025 |
Machine learning › Reinforcement learning › online decision making
reinforcement learning for systems |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.9 | 1 | 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement Learning · EuroSys 2025 |
Cloud and datacenter computing › configuration tuning
configuration auto-tuning |
0.8 | 1 | 2024 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing › cluster resource management and scheduling
container placement |
0.8 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Cloud and datacenter computing › cluster resource management and scheduling
geo-distributed scheduling |
0.8 | 1 | 2024 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job Scheduling · Proc. VLDB Endow. 2024 |
Reconfigurable computing and FPGAs › FPGA resource optimization
resource utilization balancing |
0.8 | 1 | 2024 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job Scheduling · Proc. VLDB Endow. 2024 |
Machine learning › Graph learning
graph neural network |
0.2 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Distributed systems
fault tolerance |
0.2 | 1 | 2024 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job Scheduling · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing › big data platform
shuffle service |
0.2 | 1 | 2024 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2024 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.2 | 1 | 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud Environments · ICDE 2024 |
Methods — techniques the papers use, named apart from their topics
graph neural network · 2.3column generation · 2.3deep reinforcement learning · 1.7combinatorial optimization · 1.7heuristic migration planning · 1.5two-stage decision-making · 1.0risk-aware evaluation · 1.0risk heuristics · 1.0reinforcement learning · 1.0precedent case retrieval · 1.0knowledge base construction · 1.0theory of mind · 0.9implicit channel protocol · 0.9mixed-integer programming · 0.8mixed integer programming · 0.8security assessment · 0.6frequency-domain modeling · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proactive Change Risk Detection in Production Cloud Systems: ByteDance's ExperienceabstractModern cloud services rely on a high volume of changes for rapid innovation, yet these changes are a primary cause of production incidents. To manage this risk, industry practice employs tiered change management pipelines that concentrate static rules and human reviews. However, our analysis of ByteDance's cloud platform (Volcano Engine) reveals a critical long tail problem: while high-risk changes undergo extensive scrutiny, 78.1% of change-induced incidents originate from the sheer volume of low-risk changes that receive minimal review. At this scale, exhaustive manual review is fundamentally infeasible. To address this, we present Aegis, a novel knowledge-driven system that provides interpretable change risk assessment. Aegis automatically constructs a knowledge base from historical operational data, distilling it into generalized risk heuristics and identifying relevant precedent cases. When a new change is proposed, Aegis generates human-readable warnings grounded in this historical evidence, explaining why a change is risky. We have deployed Aegis in Volcano Engine's production environment for five months, where it processed tens of thousands of change requests. Its risk escalations achieved a 75% acceptance rate from production engineers and successfully prevented multiple potential incidents. Jinyang Liu 0002, Yichen Li 0003, Tieying Zhang, Binbin Chen 0005, Xiao He 0008, Yi Li 0098 |
EuroSys | 4 |
| 2026 | Scalable and Efficient Reinforcement Learning for Virtual Machine Rescheduling in Cloud Data CentersabstractManaging a vast number of virtual machines (VMs) efficiently is a critical challenge in modern large-scale data centers. The continuous creation and termination of VMs lead to resource fragmentation across physical machines (PMs), necessitating periodic VM rescheduling to optimize resource utilization. Despite its significance, VM rescheduling has received limited attention in the literature. A key challenge is that, unlike conventional combinatorial optimization problems, the efficiency of rescheduling algorithms is heavily impacted by inference time, as VM states evolve dynamically during execution. This scalability bottleneck hampers existing methods. To address this, we propose VMR$^{2}$L, a reinforcement learning framework tailored for VM rescheduling. VMR$^{2}$L integrates a two-stage decision-making process to accommodate complex operational constraints, a feature extraction mechanism that captures critical relational information for rescheduling, and a risk-aware evaluation strategy that enables users to balance execution speed and rescheduling accuracy. Extensive experiments using real-world data from a production-scale data center demonstrate that VMR$^{2}$L achieves near-optimal performance while reducing inference time to a matter of seconds. To facilitate reproducibility, we provide access to our implementation and datasets. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Towards VM Rescheduling Optimization Through Deep Reinforcement LearningabstractModern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM rescheduling. Despite the increasing importance of VM rescheduling as data centers grow in size, the problem remains understudied. We first show that, unlike most combinatorial optimization tasks, the inference time of VM rescheduling algorithms significantly influences their performance, due to dynamic VM state changes during this period. This causes existing methods to scale poorly. Therefore, we develop a reinforcement learning system for VM rescheduling, VMR2L, which incorporates a set of customized techniques, such as a two-stage framework that accommodates diverse constraints and workload conditions, a feature extraction module that captures relational information specific to rescheduling, as well as a risk-seeking evaluation enabling users to optimize the trade-off between latency and accuracy. We conduct extensive experiments with data from an industry-scale data center. Our results show that VMR2L can achieve a performance comparable to the optimal solution but with a running time of seconds. Code12 and datasets3 are open-sourced. Xianzhong Ding, Yunkai Zhang 0002, Binbin Chen 0005, Donghao Ying, Tieying Zhang, Jianjun Chen 0001, Lei Zhang 0213, Alberto Cerpa, Wan Du |
EuroSys | 3 |
| 2025 | Learning to Communicate Through Implicit Communication ChannelsabstractEffective communication is an essential component in collaborative multi-agent systems. Situations where explicit messaging is not feasible have been common in human society throughout history, which motivate the study of implicit communication. Previous works on learning implicit communication mostly rely on theory of mind (ToM), where agents infer the mental states and intentions of others by interpreting their actions. However, ToM-based methods become less effective in making accurate inferences in complex tasks. In this work, we propose the Implicit Channel Protocol (ICP) framework, which allows agents to communicate through implicit communication channels similar to the explicit ones. ICP leverages a subset of actions, denoted as the scouting actions, and a mapping between information and these scouting actions that encodes and decodes the messages. We propose training algorithms for agents to message and act, including learning with a randomly initialized information map and with a delayed information map. The efficacy of ICP has been tested on the tasks of Guessing Numbers, Revealing Goals, and Hanabi, where ICP significantly outperforms baseline methods through more efficient information transmission. Binbin Chen 0005, Tieying Zhang, Baoxiang Wang 0001 |
ICLR | 2 |
| 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud EnvironmentsabstractContainerization has garnered substantial favor among cloud service providers. Nevertheless, the notable network overhead incurred between containers has prompted concerns within the community. In cloud resource scheduling, collocating service containers that frequently communicate to the same machine - termed “service affinity” - is instrumental in enhancing application performance. In response to this concern, we present a solution that harnesses service affinity and collocates containers to enhance the overall system performance and stability. To maximize the benefits of collocating containers, it is necessary to calculate a new schedule that optimally and efficiently maximizes service affinity, especially within the expansive domain of industry-scale cloud environments. In pursuit of this, we leverage the skewness property of affinity and machine learning to fuse solver-based algorithms, thereby assuring both quality and efficiency for problems at scale. Our methodology encompasses the partitioning of a given task into discrete subproblems, with a keen focus on resolving the most critical ones. Via a graph neural network classifier, we assign each subproblem to be solved independently using methods based on off-the-shelf solvers in our algorithm pool - namely, MIP-based, or column generation. This strategic approach enables the efficient computation of a schedule for a cloud cluster that fully optimizes the overall service affinity. We further propose a heuristic algorithm to compute executable container migration plans for practical use, facilitating the transition to the new placement where service affinity is well optimized. Our solution has been deployed in our large-scale production environment, covering over a million cores within ByteDance. Through the successful real-world production deployment, our approach exhibits an average improvement in end-to-end latency by 23.75% and a reduction in request error rates by 24.09% compared to the original system. Zuzhi Chen, Fuxin Jiang, Binbin Chen 0005, Yu Li 0003, Yunkai Zhang 0002, Jianjun Chen 0001, Wu Xiang, Guozhu Cheng, Wei Zhang 0172, Tieying Zhang |
ICDE | 3 |
| 2024 | Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDanceabstractAt ByteDance, where we execute over a million Spark jobs and handle 500PB of shuffled data daily, ensuring resource efficiency is paramount for cost savings. However, achieving optimization of resource efficiency in large-scale production environments poses significant challenges. Drawing from our practical experiences, we have identified three key issues critical to addressing resource efficiency in real-world production settings: 1 slow I/Os leading to excessive CPU and memory idleness, 2 coarse-grained resource control causing wastage, and 3 sub-optimal job configurations resulting in low utilization. To tackle these issues, we propose a resource efficiency governance framework for Spark workloads. Specifically, 1 we devise the multi-mechanism shuffle services, including Enhanced External Shuffle Service (ESS) and Cloud Shuffle Service (CSS), where CSS employs a push-based approach to enhance I/O efficiency through sequential reading. 2 We modify the Spark configuration parameter protocol, allowing for fine-grained resource control by introducing several new parameters such as milliCores and memoryBurst, as well as supporting operators with additional spill modes. 3 We design a two-stage configuration autotuning method, comprising rule-based and algorithm-based tuning, providing more reliable Spark configuration optimizations. By deploying these techniques on millions of Spark jobs in production over the last two years, we have achieved over 22% CPU utilization increase, 5% memory utilization increase, and 10% shuffle block time ratio decrease, effectively saving millions of CPU cores and petabytes of memory daily. Xiuqi Huang, Wei Zhongjia, Hang Cheng, Chaohui Xin, Zuzhi Chen, Binbin Chen 0005, Yufei Wu 0014, Hao Wang 0210, Tieying Zhang, Xiaofeng Gao 0001, Yuming Liang, Pengwei Zhao, Guihai Chen |
Proc. VLDB Endow. | 7 |
| 2024 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job SchedulingabstractAt internet scale companies like ByteDance, data is generated and consumed at enormously high speed by many different applications. Achieving low latency on such big data jobs is an important problem. However, the naive approach of aggregating all the data required by a job to a single location is not always feasible in a geo-distributed environment. Similarly, existing approaches in geo-distributed job scheduling often try to minimize WAN usage, which may come at the cost of latency. Another crucial element to ensure low latency is resource load balancing among DCs, which enables flexibility in job scheduling and avoids resource bottlenecks. Therefore, to minimize latency, optimizing job completion time (JCT) while maintaining resource utilization balance is important. To this end, we propose ResLake , a global scheduling platform for data-intensive workloads. ResLake aims to reduce JCT of geo-distributed applications while balancing the compute (CPU/Memory) and storage (Disk) usages across DCs and efficiently using WAN interconnections. We have deployed ResLake in ByteDance's production for over 1.5 years. ResLake has scheduled billions of jobs since its deployment. We find that ResLake improves JCT of jobs by at least 20%, and can improve resource utilization balance across DCs by up to 53%. Xin-Chun Zhang, Aqsa Kashaf, Yihan Zou, Wei Zhang 0172, Weibo Liao, Song Haoxiang, Jintao Ye, Binbin Chen 0005, Zuzhi Chen, Tieying Zhang, Yongping Tang |
Proc. VLDB Endow. | 12 |
| 2022 | Energy-Circuit-Based Integrated Energy Management System: Theory, Implementation, and ApplicationabstractIntegrated energy systems (IESs), in which various energy flows are interconnected and coordinated to release potential flexibility for more efficient and secure operation, have drawn increasing attention in recent years. In this article, an integrated energy management system (IEMS) that performs online analysis and optimization on coupling energy flows in an IES is comprehensively introduced. From the theory perspective, an energy circuit method (ECM) that models natural gas networks and heating networks in the frequency domain is discussed. This method extends the electric circuit modeling of power systems to IESs and enables the IEMS to manage large-scale IESs. From the implementation perspective, the architecture design and function development of the IEMS are presented. Tutorial examples with illustrative case studies are provided to demonstrate its functions of dynamic state estimation, energy flow analysis, security assessment and control, and optimal energy flow. From the application perspective, real-world engineering demonstrations that apply IEMSs in managing building-, park-, and city-scale IESs are reported. The economic and environmental benefits obtained in these demonstration projects indicate that the IEMS has broad application prospects for a low/zero-carbon future energy system. Binbin Chen 0005, Qinglai Guo, Guanxiong Yin, Bin Wang 0092, Zhaoguang Pan, Yuwei Chen 0008, WenChuan Wu 0001, Hongbin Sun 0002 |
Proc. IEEE | 1 |