EDBT 2026 Demo / reviewers in the wild / expert
Chenyang Hei
dblp:370/3985
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-5010-1529ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DistSRE: Synthesizing Runtime Executions with Automated Co-Optimization for Distributed LLM Training
Xiuzhu Sha, Chenyang Hei, Fuliang Li, Chengxi Gao, Rongfei Zeng, Xingwei Wang 0001 |
IWQoS | 2 |
| 2026 | UPServe: Backend Agnostic Proxy for Black-box Heterogeneous LLM Scheduling
Haorui Wan, Chenyang Hei, Fuliang Li, Chengxi Gao, Yuhan Jia, Tongrui Liu, Xingwei Wang 0001 |
IWQoS | 2 |
| 2026 | HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001 |
NSDI | 1 |
| 2025 | Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU ClustersabstractState-of-the-art deep learning models rely on large GPU clusters and various parallelism strategies, which in turn depend on collective communication (CC) operators to synchronize data. While vendor libraries (e.g., NCCL, RCCL) provide standard CC algorithms, they often suffer from bandwidth bottlenecks in imbalanced topologies. Recent synthesis-based methods improve performance but face three key limitations: poor scalability due to the combinatorial explosion of scheduling space, lack of support for multistage execution, and suboptimal communication throughput. We propose Canvas, a scalable and near-optimal CC scheduling framework that addresses these challenges. Canvas introduces: (1) Hierarchical synthesis to decompose the global scheduling problem into tractable subproblems for scalability. (2) Collective decomposition to enable structured, multi-stage algorithm generation. (3) Cross-micro-batch pipeline scheduling to parallelize communication across micro-batches and maximize link utilization. Evaluations show that Canvas achieves up to 1.98× bandwidth speedup over TACCL and 3.56× over TE-CCL, and synthesizes algorithms for 512-GPU topologies within 1.77 hours, whereas TACCL fails to produce results within 24 hours. Chenyang Hei, Fuliang Li, Chengxi Gao, Tongrui Liu, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 1 |
| 2025 | TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication LibraryabstractModern distributed training systems face escalating communication bottlenecks as GPU clusters scale to accommodate large models. While collective communication libraries and automated synthesizers address algorithmic efficiency, they suffer from three critical limitations including labor-intensive manual intervention requirement, overreliance on predefined input optimization, and suboptimal isolated configuration optimization. To solve these problems, we present TuCCL, a systematic framework that co-optimizes communication algorithms and runtime parameters through three innovations: Topology-Aware Sketch Generation that automatically produces high-performance primitives, Hierarchical Configuration Optimization modeling nonlinear parameter-performance relationships, and Multi-phase Resource-Aware Configuration Optimization enabling joint configuration tuning with adaptive search space pruning. Evaluations demonstrate TuCCL’s superiority over state-of-the-art systems with 1.75x–11.49x bandwidth improvements for AllGather/AllReduce on NVIDIA V100/A100 clusters, 90.2% faster configuration search than grid methods, and 1.22x-2.52x end-to-end training speedups across diverse model scales. Chenyang Hei, Fuliang Li, Tongrui Liu, Chengxi Gao, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 2 |
| 2025 | MAZ3: Memory-Assisted ZeRO-3 for Efficient Collective CommunicationabstractLarge Language Models (LLMs) have advanced rapidly, but their growing parameter scales and memory demands pose critical challenges for distributed training. Although GPU memory capacity improves steadily, model sizes expand much faster, causing frequent Out-of-Memory (OOM) errors and rising training costs. Existing memory optimization approaches, such as ZeRO-3 and offloading, alleviate per-GPU memory pressure but introduce excessive collective communication, limited computation–communication overlap, and degraded scalability. We present MAZ3, a distributed training framework that mitigates these limitations through three key techniques: (1) Collaborative CPU–GPU memory management, storing full parameters in CPU memory and broadcasting them within nodes to reduce global synchronization; (2) Fine-grained communication–computation overlap, aligning collective operations with model computation to hide latency; (3) Hierarchical aggregation operators, leveraging intra-node NVLink and inter-node NIC channels concurrently to minimize communication overhead. We implement MAZ3 on a multi-GPU cluster and evaluate it with large-scale models. Results show that MAZ3 reduces inter-node model communication (gradients and parameters) by 33%, improves training efficiency by 40.3%, and increases throughput by 67.9% compared with ZeRO-3. Moreover, MAZ3 retains the memory efficiency of ZeRO-3 while approaching the training efficiency and throughput of ZeRO-2 Offload, achieving a balanced trade-off between memory optimization and performance. Chenyang Hei, Fuliang Li, Chengxi Gao, Xingwei Wang 0001 |
ICNP | 2 |
| 2025 | SmartTC: A Real-Time ML-Based Traffic Classification with SmartnicabstractReal-time network traffic classification plays a crucial role in ensuring Quality of Service and network security, and machine learning (ML) based methods achieve high classification accuracy but induce significant computational overhead. While SmartNIC solutions can offload classification tasks thus reducing CPU burdens, they still suffer from various limitations, manifested in limited computing capabilities, insufficiency in dynamic load handling and high latency from heterogeneous computing architectures. To solve these problems, we propose SmartTC, with three key designs: (1) SmartTC employs hardware-software co-design to optimize SmartNIC processing power, (2) SmartTC adopts a trafficaware dynamic batch submission strategy that adjusts submission policies based on real-time network load, and (3) SmartTC proposes parallel pipeline scheduling that ensures efficient task execution while minimizing communication overhead. Finally, we implement SmartTC on the BlueField-3 DPU and conduct extensive experiments for evaluations, and comparison results demonstrate that SmartTC significantly outperforms existing solutions. For example, it reduces average traffic classification time by up to$\mathbf{1 6. 8 \%}$under low loads and$\mathbf{9 0. 9 \%}$under high loads. Besides, SmartTC does not affect Bluefield-3 network services, and saves host CPU usage by at least two cores. Lingxiang Hu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 2 |
| 2025 | ResCCL: Resource-Efficient Scheduling for Collective CommunicationabstractAs distributed deep learning training (DLT) systems scale, collective communication has become a significant performance bottleneck. While current approaches optimize bandwidth utilization and task completion time, existing communication libraries (CCLs) backends fail to efficiently manage GPU resources during algorithm execution, limiting the performance of advanced algorithms. This paper proposes ResCCL, a novel CCL backend designed for Resource-Efficient Scheduling to address key limitations in current systems. ResCCL enhances execution efficiency by optimizing scheduling at the primitive level (e.g., send and recvReduceCopy), enabling flexible thread block (TB) allocation, and generating lightweight communication kernels to minimize runtime overhead. Our approach tackles the global scheduling problem, reduces idle TB resources, and enhances communication bandwidth. Evaluation results demonstrate that ResCCL achieves up to 2.5× improvement in bandwidth performance compared to both NCCL and MSCCL. It reduces SM resource overhead by 77.8% and increases TB utilization by 41.6% while running the same algorithms. In end-to-end DLT, ResCCL boosts Megatron's throughput by up to 39%. Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Ennan Zhai, Xingwei Wang 0001 |
SIGCOMM | 2 |
| 2024 | Human-Intent-Driven Cellular Configuration Generation Using Program SynthesisabstractCellular networks are vital for emerging applications like the Metaverse, which impose demanding quality and quantity requirements. This necessitates frequent reconfiguration of both new and existing base stations to balance network service quality (e.g., ultra-low latency and high bandwidth) and resource consumption. Existing data-driven configuration methods learn from historical data, but have two key limitations. First, they yield only approximate solutions, lacking precision. Second, poor bootstrapping for new base stations with previously unobserved attributes. In this paper, we pioneer intent-driven configuration synthesis by designing an intent language and utilizing satisfiability modulo theory (SMT) for cellular networks to enable exact and precise solutions. We formulate synthesis as an SMT problem, permitting verification of precision. First, we cast configuration generation as a program synthesis problem via novel modeling to bridge the intent-configuration gap. Second, we extend SMT synthesis to scale to large networks. However, vanilla SMT approaches have poor scalability. Hence, we propose an optimization using sampling for constraint verification instead of exhaustive forward solving. We also design a domain-specific optimization to prune the sample space and improve efficiency. Experiments on various network scales demonstrate the effectiveness of our proposed SMT-based cellular network configuration synthesis. Fuliang Li, Chenyang Hei, Jiaxing Shen, Qing Li 0006, Xingwei Wang 0001 |
IEEE J. Sel. Areas Commun. | 2 |