Qianyu Zhang 0001

dblp:139/5941-1 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-3805-4480ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
abstract
As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions.
Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao
PPoPP5
2024 ALEPH: Accelerating Distributed Training With eBPF-Based Hierarchical Gradient Aggregation
abstract
Distributed training includes two important operations: gradient transmission and gradient aggregation, which will consume massive bandwidth and computing resources. To achieve efficient distributed training, one must overcome two critical challenges: heterogeneity of bandwidth resources and limitation of computing resources among compute nodes. Existing architectures based on Parameter Server (PS) and All-Reduce (AR) fail to cope with these challenges because the PS will aggregate gradients from all workers and suffers from bandwidth bottlenecks, while AR intends to alleviate bandwidth bottlenecks at the PS, but the workers need to process many gradient packets thus can be overloaded. To address these shortcomings, we design a new distributed training system called ALEPH. In the control plane, ALEPH uses an efficient algorithm to group workers into clusters with different sizes so as to fully utilize heterogeneous bandwidth. We show that the proposed algorithm can achieve a good approximation performance. In the data plane, ALEPH leverages, for the first time, extended Berkeley Packet Filter (eBPF) programs to aggregate and forward gradient packets to reduce computation overhead. We show how to overcome several hurdles in using eBPF for distributed training. We implement ALEPH and evaluate its performance on a small-scale testbed and large-scale simulations. Experimental results show that ALEPH reduces training time by 20%-31% and increases bandwidth utilization by 88% compared with state-of-the-art frameworks.
Peng Yang 0022, Hongli Xu 0001, Gongming Zhao, Qianyu Zhang 0001, Jiawei Liu 0007, Chunming Qiao
IEEE/ACM Trans. Netw.4
2024 XAgg: Accelerating Heterogeneous Distributed Training Through XDP-Based Gradient Aggregation
abstract
With the growth of model/dataset/system size for distributed model training in datacenters, the widely used Parameter Server (PS) architecture suffers from communication bottleneck of gradient transmission. Recent works attempt to utilize programmable switches to implement in-network gradient aggregation and alleviate communication bottlenecks on PSs. Due to the limited on-chip memory of programmable switches, gradient transmission requires strict synchronization to achieve ideal aggregation performance. However, the distributed training system is usually heterogeneous in datacenters (e.g., computation and bandwidth heterogeneity), and the gradient will reach the aggregation nodes asynchronously, thereby seriously affecting the aggregation performance. To solve the above issue, we propose XAgg, which accelerates heterogeneous gradient aggregation by deploying the eXpress Data Path (XDP) based aggregator on servers. Specifically, the abundant idle memory on servers can cache the entire gradient, so as to effectively deal with asynchronous gradient transmission in heterogeneous scenarios. Moreover, XDP can provide high-performance and low-latency gradient aggregation. We conduct microbenchmark and testbed with real-world DNN models and datasets. Experimental results show that XAgg improves the gradient aggregation throughput by 3.3$\times$compared with TCP-based aggregation, reaching 100 Gbps with 10 CPU cores. In addition, XAgg reduces communication time by 49%-82% compared with state-of-the-art solutions.
Qianyu Zhang 0001, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022
IEEE/ACM Trans. Netw.1
2023 Scalable and Robust East-West Forwarding Framework for Hyperscale Clouds
abstract
With the broad deployment of distributed applications on clouds, east-west traffic is now dominating the majority of cloud networks. The existing communication solutions are tightly coupled with either the control plane (e.g., preprogrammed model) or the location of compute nodes (e.g., conventional gateway model). As a result, it is difficult to flexibly respond to the rapidly expanding networks and frequent abnormal events (e.g., burst traffic and device failures). Accordingly, they may not provide high-performance east-west forwarding while ensuring scalability and robustness. To address this issue, we design Zeta, a scalable and robust east-west forwarding framework with gateway clusters for hyperscale clouds. Zeta abstracts the traffic forwarding capability as a Gateway Cluster Layer, decoupled from the logic of control plane and the location of compute nodes. Specifically, Zeta adopts gateway clusters to support large-scale networks and cope with burst traffic. Moreover, a transparent Multi IPs Migration is proposed for fast recovery from unpredictable failures. We implement Zeta based on eXpress Data Path (XDP) and evaluate its scalability and robustness through comprehensive experiments with up to 100k container instances. Our evaluation shows that Zeta reduces the 99% RTT by$5.1 {\times }$in burst video traffic, and reduces the gateway pure recovery delay by$10.8 {\times }$compared with the state-of-the-art solutions.
Qianyu Zhang 0001, Gongming Zhao, Liguang Xie, Hongli Xu 0001, Zhuolong Yu, Yangming Zhao, Chunming Qiao, Liusheng Huang
IEEE/ACM Trans. Netw.1
2022 OXDP: Offloading XDP to SmartNIC for Accelerating Packet Processing
abstract
Traditional kernel network processing suffers from high delay and overhead, which has become the bottleneck of high-speed networks. A natural method to accelerate packet processing is to bypass the kernel network stack and process packets in user space directly, $e.g$., DPDK. However, due to many network functions are implemented in the kernel network stack, bypassing the stack means that we need to redesign the required functions elsewhere, leading to poor compatibility. One promising technology to address this problem is called eXpress Data Path (XDP), which can support high-performance packet processing while preserving the kernel stack. However, existing solutions mainly run XDP in software mode, resulting in relatively poor packet processing performance. Fortunately, with the development of programmable hardware, running XDP in hardware mode is a more promising approach. Thus, in this paper, we design and implement OXDP, the first-of-its-kind work on accelerating packet processing by offloading XDP to SmartNICs. Since today’s SmartNICs are still subject to some limitations regarding the rigid runtime environment, it is nontrivial to offload XDP to SmartNICs. To address this issue, OXDP performs best-effort offloading based on the primitive packet operations, thus maximizing the use of SmartNIC’s resources. Specifically, OXDP splits the forwarding function into two parts, one part offloading on SmartNIC with hardware XDP and the other part deploying on host. We evaluate the efficiency of OXDP with comprehensive experiments. Evaluation results show that the forwarding rate of OXDP can reach 18.7 Mpps, which improves $30 \times$ compared with the single-core performance of software XDP.
Gongming Zhao, Qianyu Zhang 0001, Hongli Xu 0001, Liguang Xie
ICPADS3
2022 Zeta: A Scalable and Robust East-West Communication Framework in Large-Scale Clouds
Qianyu Zhang 0001, Gongming Zhao, Hongli Xu 0001, Zhuolong Yu, Liguang Xie, Yangming Zhao, Chunming Qiao, Liusheng Huang
NSDI1
2022 A Robust Service Mapping Scheme for Multi-Tenant Clouds
abstract
In a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants’ traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants’ traffic are two critical challenges that affect the tenants’ QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach:service node assignmentandtenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+$\epsilon $, where$\epsilon $is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives.
Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Yutong Zhai, Qianyu Zhang 0001, He Huang 0001, Yongqiang Yang
IEEE/ACM Trans. Netw.5