VLDB 2026 Research / reviewers in the wild / expert
Decang Sun
dblp:303/4263
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-3442-4656ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learn-to-Probe: Achieving Signal Distinguishability in Learning-based Congestion ControlabstractInternet congestion control remains a fundamental challenge, and recent learning-based congestion control algorithms (CCAs) have shown potential in optimizing network performance. However, their reliance on heuristically chosen input signals often leads to suboptimal behavior across diverse network conditions. In this paper, we identify the root cause as the lack of signal distinguishability-the ability of signals to reflect meaningful differences in network states. To address this, we propose Learn-to-Probe (LTP), a novel signal engineering paradigm that actively generates distinguishable network signals to improve the learning process. LTP (i) employs Bayesian filtering to accurately estimate network states from historical signals, and (ii) introduces an intrinsic reinforcement learning reward that encourages the flow to probe the network, inducing signal sequences that minimize the estimation uncertainty in (i). This probing behavior naturally enhances signal distinguishability, enabling the learning model to make more informed decisions. Extensive evaluations show that LTP consistently achieves high link utilization, low queuing delay, and stable convergence across diverse environments. Our results underscore the importance of signal distinguishability and offer a new direction for robust, adaptive congestion control. Han Tian, Junxue Zhang 0001, Xudong Liao, Decang Sun, Bin Huang 0024, Wenxue Li 0004, Yong Wang 0046, Kai Chen 0005 |
EuroSys | 5 |
| 2026 | PolicyCache: Intra-flow Learning in Congestion Control
Han Tian, Xudong Liao, Decang Sun, Wenxue Li 0004, Bin Huang 0024, Senbo Fu, Junxue Zhang 0001, Dian Shen, Kai Chen 0005 |
NSDI | 5 |
| 2025 | Design and Operation of Shared Machine Learning Clusters on CampusabstractThe rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems. Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005 |
ASPLOS (1) | 2 |
| 2025 | Achieving Fairness Generalizability for Learning-based Congestion Control with JuryabstractInternet congestion control (CC) has long posed a challenging control problem in networking systems, with recent approaches increasingly incorporating deep reinforcement learning (DRL) to enhance adaptability and performance. Despite promising, DRL-based CC schemes often suffer from poor fairness, particularly when applied to network environments unseen during training. This paper introduces Jury, a novel DRL-based CC scheme designed to achieve fairness generalizability. At its heart, Jury decouples the fairness control from the principal DRL model with two design elements: i) By transforming network signals, it provides a universal view of network environments among competing flows, and ii) It adopts a post-processing phase to dynamically module the sending rate based on flow bandwidth occupancy estimation, ensuring large flows behave more conservatively and smaller flows more aggressively, thus achieving a fair and balanced bandwidth allocation. We have fully implemented Jury, and extensive evaluations demonstrate its robust convergence properties and high performance across a broad spectrum of both emulated and real-world network conditions. Han Tian, Xudong Liao, Decang Sun, Chaoliang Zeng, Yilun Jin, Junxue Zhang 0001, Xinchen Wan, Zilong Wang 0007, Yong Wang 0046, Kai Chen 0005 |
EuroSys | 3 |
| 2025 | Enabling In-Network Acceleration Over the Cloud
Hao Wang 0116, Decang Sun, Jinbin Hu 0001, Kai Chen 0005 |
INFOCOM | 2 |
| 2025 | GREEN: Carbon-efficient Resource Scheduling for Machine Learning Clusters
Kaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 0001, Kai Chen 0005 |
NSDI | 2 |
| 2024 | Efficient DRL-Based Congestion Control With Ultra-Low OverheadabstractPrevious congestion control (CC) algorithms based on deep reinforcement learning (DRL) directly adjust flow sending rate to respond to dynamic bandwidth change, resulting in high inference overhead. Such overhead may consume considerable CPU resources and hurt the datapath performance. In this paper, we present, a hierarchical congestion control algorithm that fully utilizes the performance gain from deep reinforcement learning but with ultra-low overhead. At its heart, decouples the congestion control task into two subtasks in different timescales and handles them with different components: 1) lightweight CC executor that performs fine-grained control responding to dynamic bandwidth changes; and 2) RL agent that works at a coarse-grained level that generates control sub-policies for the CC executor. Such two-level control architecture can provide fine-grained DRL-based control with a low model inference overhead. Real-world experiments and emulations show that achieves consistent high performance across various network conditions with an ultra-low control overhead reduced by at least 80% compared to its DRL-based counterparts, similar to classic CC schemes such as Cubic. Han Tian, Xudong Liao, Chaoliang Zeng, Decang Sun, Junxue Zhang 0001, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Egeria: Efficient DNN Training with Knowledge-Guided Layer FreezingabstractTraining deep neural networks (DNNs) is time-consuming. While most existing solutions try to overlap/schedule computation and communication for efficient training, this paper goes one step further by skipping computing and communication through DNN layer freezing. Our key insight is that the training progress of internal DNN layers differs significantly, and front layers often become well-trained much earlier than deep layers. To explore this, we first introduce the notion of training plasticity to quantify the training progress of internal DNN layers. Then we design Egeria, a knowledge-guided DNN training system that employs semantic knowledge from a reference model to accurately evaluate individual layers' training plasticity and safely freeze the converged ones, saving their corresponding backward computation and communication. Our reference model is generated on the fly using quantization techniques and runs forward operations asynchronously on available CPUs to minimize the overhead. In addition, Egeria caches the intermediate outputs of the frozen layers with prefetching to further skip the forward computation. Our implementation and testbed experiments with popular vision and language models show that Egeria achieves 19%-43% training speedup w.r.t. the state-of-the-art without sacrificing accuracy. Decang Sun, Kai Chen 0005, Fan Lai 0001, Mosharaf Chowdhury |
EuroSys | 2 |