EDBT 2026 Demo / reviewers in the wild / expert
Kun Qian 0021
dblp:77/2062-21
· DBLP profile ↗
14ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0001-9882-9279ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai |
NSDI | 2 |
| 2025 | Evolution of Aegis: Fault Diagnosis for AI Model Training Service in Production
Jianbo Dong, Kun Qian 0021, Zhilong Zheng, Liang Chen 0001, Yichi Xu, Yikai Zhu, Xue Li 0024, Zhihui Ren, Yang Liu 0245, Yu Guan 0005, Chaojie Yang, Yang Zhang 0102, Man Yuan, Yong Li 0008, Xianlong Zeng, Zhiping Yao, Binzhang Fu, Ennan Zhai, Wei Lin 0016, Dennis Cai |
NSDI | 2 |
| 2025 | SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Dan Li 0001, Li Chen 0008, Heyang Zhou, Linkang Zheng, Yikai Zhu, Yang Liu 0245, Kun Qian 0021, Kunling He, Ennan Zhai, Dennis Cai, Binzhang Fu |
NSDI | 13 |
| 2025 | SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingabstractThe flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting. Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai |
SIGCOMM | 2 |
| 2025 | SyCCL: Exploiting Symmetry for Efficient Collective Communication SchedulingabstractThe performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. State-of-the-art collective schedule synthesizers (such as TECCL and TACCL) utilize Mixed Integer Linear Program for modeling but encounter search space explosion and scalability challenges. In this paper, we propose SyCCL, a scalable collective schedule synthesizer that aims to synthesize near-optimal schedules in tens of minutes for production-scale machine-learning jobs. SyCCL leverages collective and topology symmetries to decompose the original collective communication demand into smaller sub-demands within smaller topology subsets. SyCCL proposes efficient search strategies to quickly explore potential sub-demands, synthesizes corresponding sub-schedules, and integrates these sub-schedules into complete schedules. Our 32-A100 testbed and production-scale simulation experiments show that SyCCL improves collective performance by up to 127% while reducing synthesis time by 2 to 4 orders of magnitude compared to state-of-the-art efforts. Jiamin Cao, Shangfeng Shi, Weisen Liu, Yifan Yang 0009, Yichi Xu, Zhilong Zheng, Yu Guan 0005, Kun Qian 0021, Ying Liu 0024, Mingwei Xu 0001, Ning Wang 0001, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 9 |
| 2025 | Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market
Yuxing Xiang, Xue Li 0024, Kun Qian 0021, Diwen Zhu, Wenyuan Yu, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008, Jingren Zhou 0001 |
SOSP | 3 |
| 2024 | Near-Lossless Gradient Compression for Data-Parallel Distributed DNN TrainingabstractData parallelism has become a cornerstone in scaling up the training of deep neural networks (DNNs). However, the communication overhead associated with synchronizing gradients across multiple nodes has emerged as a significant bottleneck, adversely affecting training efficiency and leading to a surge in large-scale distributed model training costs. By leveraging insights into the statistical characteristics of gradients, we present GComp, a near-lossless gradient compression scheme designed to reduce the communication burden during data-parallel training significantly. GComp develops an optimized Huffman encoding/decoding strategy to compress gradient exponents effectively. Additionally, it introduces an innovative multi-level quantization method for mantissa, complemented by a pruning strategy that eliminates zero-valued gradients. These integrated approaches significantly reduce the volume of data for synchronization, while virtually not affecting the DNN model's training accuracy. We conduct comprehensive evaluations of GComp, demonstrating that our method can decrease the communication volume by as much as 67.1%, and enhance training speed by up to 1.9×. Xue Li 0024, Cheng Guo 0007, Kun Qian 0021, Menghao Zhang 0001, Mengyu Yang, Mingwei Xu 0001 |
SoCC | 3 |
| 2024 | Burstable Cloud Block Storage with Data Processing Units
Junyi Shu, Kun Qian 0021, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
OSDI | 2 |
| 2024 | Crux: GPU-Efficient Communication Scheduling for Deep Learning TrainingabstractDeep learning training (DLT), e.g., large language model (LLM) training, has become one of the most important services in multitenant cloud computing. By deeply studying in-production DLT jobs, we observed that communication contention among different DLT jobs seriously influences the overall GPU computation utilization, resulting in the low efficiency of the training cluster. In this paper, we present Crux, a communication scheduler that aims to maximize GPU computation utilization by mitigating the communication contention among DLT jobs. Maximizing GPU computation utilization for DLT, nevertheless, is NP-Complete; thus, we formulate and prove a novel theorem to approach this goal by GPU intensity-aware communication scheduling. Then, we propose an approach that prioritizes the DLT flows with high GPU computation intensity, reducing potential communication contention. Our 96-GPU testbed experiments show that Crux improves 8.3% to 14.8% GPU computation utilization. The large-scale production trace-based simulation further shows that Crux increases GPU computation utilization by up to 23% compared with alternatives including Sincronia, TACCL, and CASSINI. Jiamin Cao, Yu Guan 0005, Kun Qian 0021, Wencong Xiao, Jianbo Dong, Binzhang Fu, Dennis Cai, Ennan Zhai |
SIGCOMM | 3 |
| 2024 | Alibaba HPN: A Data Center Network for Large Language Model TrainingabstractThis paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production. Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai |
SIGCOMM | 1 |
| 2023 | XRON: A Hybrid Elastic Cloud Overlay Network for Video Conferencing at Planetary ScaleabstractQuality and cost are two key considerations for video conferencing services. Service providers face a dilemma when selecting network tiers to build their infrastructure---relying on Internet links has poor quality, while using premium links brings excessive cost. Bingyang Wu, Kun Qian 0021, Bo Li 0061, Dennis Cai, Ennan Zhai, Xuanzhe Liu, Xin Jin 0008 |
SIGCOMM | 2 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 3 |
| 2022 | From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudabstractThis paper presents the two generations of storage network stacks that reduced the average I/O latency of Alibaba Cloud's EBS service by 72% in the last five years: Luna, a user-space TCP stack that corresponds the latency of network to the speed of SSD; and Solar, a storage-oriented UDP stack that enables both storage and network hardware accelerations. Rui Miao 0001, Lingjun Zhu, Kun Qian 0021, Shujun Zhuang, Bo Li 0061, Shuguang Cheng, Binzhang Fu, Jiaji Zhu, Jiesheng Wu, Dennis Cai, Hongqiang Harry Liu |
SIGCOMM | 4 |