Peng Yang 0022

dblp:57/5443-22 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4400-2289ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 2 first-author · 9 since 2021
YearPublicationVenuePosition
2026 NAST: In-Network Aggregation with Worker Selection for Accelerating Distributed Training
Jianfeng Bao, Peng Yang 0022, Gongming Zhao, Huihui Tang, Hongli Xu 0001, Qianpiao Ma
IWQoS2
2026 Achieving Efficient and Robust Multi-Job Resource Scheduling in Deep Learning Clusters
Jianfeng Bao, Wentao Fan 0002, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022, Xiaohu Xu
IEEE Trans. Netw.5
2025 Network Resource-Aware Multi-Job Deployment in Deep Learning Clusters
Ai Zhong, Gongming Zhao, Hongli Xu 0001, Jiawei Liu 0007, Peng Yang 0022
ICCCN6
2025 CROP: Efficient and Robust Multi-Job Placement in Deep Learning Clusters
abstract
Deep learning (DL) has seen a growing dataset, an expanding model scale, and increasing applications in recent years. There is a notable trend of shifting DL training jobs from local computing units to powerful DL clusters built by cloud providers. These clusters allocate physical training nodes to DL jobs through a process referred to as multi-job placement. Existing multi-job placement strategies fail to achieve high efficiency in resource utilization, DL training, and robustness simultaneously, resulting in poor performance when resources are limited or when abnormalities occur in some devices. To tackle these challenges, we present CROP, an approach that performs efficient and robust multi-job placement in DL clusters. We formulate the efficient and robust multi-job placement problem as a non-linear program and prove its NP-hardness. To solve this problem, we present an effective submodular-based algorithm with a tight approximation factor of ($1-1/e$). We evaluate CROP on a small-scale testbed consisting of 8 physical GPUs and a large-scale simulation employing real-world job traces. Experimental results demonstrate that CROP achieves nearoptimal communication overhead while improving the training throughput of the DL cluster by up to$57.5\%$compared to state-of-the-art solutions.
Peng Yang 0022, Gongming Zhao, Hongli Xu 0001, Haibo Wang 0004, Wentao Fan 0002, Xiaohu Xu
IWQoS1
2024 HifiCNet: High-Fidelity Cloud Network Validation Platform at Scale by Hybrid Architecture
abstract
Ensuring reliable operation of cloud networks is critical for cloud service providers to guarantee quality of service for tenants. A promising solution is to design a high-fidelity cloud network validation platform that proactively validates the correctness of all operations before implementing changes to the production network. However, the tight coupling between physical and virtual networks in the cloud poses challenges to achieving high-fidelity cloud network validation. Existing network validation platforms focus primarily on traditional physical networks, while ignoring virtual network validation. Regrettably, neglecting the combined validation of physical and virtual networks will result in inaccurate evaluations. To bridge this gap, we present HifiCNet, a high-fidelity platform that concurrently validates both physical and virtual networks. HifiCNet designs an orchestrator to elegantly coordinate the interaction between physical and virtual networks in the cloud and innovatively adopts an emulator-simulator hybrid architecture to ensure high fidelity and scalability for cloud network validation. Through extensive evaluation based on real topologies and traffic traces, we show that HifiCNet enables high-fidelity validation of cloud network configurations, services, and exceptions. Notably, HifiCNet can use 38 servers to establish a physical network comprising 10k hosts, and a virtual network consisting of 200 k virtual machines.
Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Baoqing Wang, Peng Yang 0022, Chun-Jen Chung, Min Chen 0033, Xuwei Yang
ICNP5
2024 InGo: In-Network Aggregation Routing with Batch Size Adjustment for Distributed Training
abstract
Distributed training has emerged as a critical application in clusters due to the widespread adoption of AI technology across various domains. However, as distributed training continues to advance, it has become increasingly time-consuming. To address this challenge, researchers have explored leveraging In-Network Aggregation (INA) to expedite distributed model training. Specifically, by harnessing programmable hardware, such as Intel Tofino switches, INA can aggregate gradients within the network, thereby reducing the amount of gradient transmission and accelerating distributed training. However, previous works assume fixed routing selection and batch size, ignoring their impact on model convergence and resulting in extended completion time. To bridge this gap, we propose InGo, a pioneering approach that considers both in-network aggregation routing and batch size adjustment, and provide the rigorous convergence analysis. Then, we formally define the problem of in-network aggregation routing with batch size adjustment, and present an efficient algorithm with bounded approximation factors to solve this problem. Through extensive experiments on both physical platforms and simulated environments, we demonstrate that InGo significantly reduces the completion time by 25.2%-74.7% compared to state-of-the-art solutions.
Jianfeng Bao, Gongming Zhao, Hongli Xu 0001, Haibo Wang 0004, Peng Yang 0022
IWQoS5
2024 Toward a Service Availability-Guaranteed Cloud Through VM Placement
abstract
In a multi-tenant cloud, the cloud service provider (CSP) leases physical resources to tenants in the form of virtual machines (VMs) with an agreed service level agreement (SLA). As the most important indicator of SLA, we should guarantee the service availability of tenants when placing the VMs. However, previous works about VM placement mainly concentrate on optimizing the cloud resource utilization, but only a few works consider the service availability by measuring the hardware availability. In fact, abnormal tenants can make the corresponding service unavailable by launching network attacks. That is, both the hardware availability and the tenant uncertainty will affect the service availability of VMs on physical machines (PMs). Without considering this factor, the CSP may fail to meet the tenant’s SLA requirements, leading to a reduction in revenue. To solve such a problem, this paper considers the service availability in terms of both the hardware availability and the tenant uncertainty, and studies the service availability-guaranteed VM placement in multi-tenant clouds (SAG-VMP) problem. This problem is very challenging since the service availability actually changes with the tenants served on the PM. To address this issue, we propose a two-phase approach: PM assignment and VM placement. The first phase determines the availability of each PM through a long-term tenant-PM mapping algorithm and the second phase places each VM on a PM that meets the service availability requirement based on a primal-dual online algorithm. Two algorithms with bounded approximation factors are proposed for these two phases, respectively. Both small-scale experiment results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives.
Jiawei Liu 0007, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022, Baoqing Wang, Chunming Qiao
IEEE/ACM Trans. Netw.4
2024 ALEPH: Accelerating Distributed Training With eBPF-Based Hierarchical Gradient Aggregation
abstract
Distributed training includes two important operations: gradient transmission and gradient aggregation, which will consume massive bandwidth and computing resources. To achieve efficient distributed training, one must overcome two critical challenges: heterogeneity of bandwidth resources and limitation of computing resources among compute nodes. Existing architectures based on Parameter Server (PS) and All-Reduce (AR) fail to cope with these challenges because the PS will aggregate gradients from all workers and suffers from bandwidth bottlenecks, while AR intends to alleviate bandwidth bottlenecks at the PS, but the workers need to process many gradient packets thus can be overloaded. To address these shortcomings, we design a new distributed training system called ALEPH. In the control plane, ALEPH uses an efficient algorithm to group workers into clusters with different sizes so as to fully utilize heterogeneous bandwidth. We show that the proposed algorithm can achieve a good approximation performance. In the data plane, ALEPH leverages, for the first time, extended Berkeley Packet Filter (eBPF) programs to aggregate and forward gradient packets to reduce computation overhead. We show how to overcome several hurdles in using eBPF for distributed training. We implement ALEPH and evaluate its performance on a small-scale testbed and large-scale simulations. Experimental results show that ALEPH reduces training time by 20%-31% and increases bandwidth utilization by 88% compared with state-of-the-art frameworks.
Peng Yang 0022, Hongli Xu 0001, Gongming Zhao, Qianyu Zhang 0001, Jiawei Liu 0007, Chunming Qiao
IEEE/ACM Trans. Netw.1
2024 XAgg: Accelerating Heterogeneous Distributed Training Through XDP-Based Gradient Aggregation
abstract
With the growth of model/dataset/system size for distributed model training in datacenters, the widely used Parameter Server (PS) architecture suffers from communication bottleneck of gradient transmission. Recent works attempt to utilize programmable switches to implement in-network gradient aggregation and alleviate communication bottlenecks on PSs. Due to the limited on-chip memory of programmable switches, gradient transmission requires strict synchronization to achieve ideal aggregation performance. However, the distributed training system is usually heterogeneous in datacenters (e.g., computation and bandwidth heterogeneity), and the gradient will reach the aggregation nodes asynchronously, thereby seriously affecting the aggregation performance. To solve the above issue, we propose XAgg, which accelerates heterogeneous gradient aggregation by deploying the eXpress Data Path (XDP) based aggregator on servers. Specifically, the abundant idle memory on servers can cache the entire gradient, so as to effectively deal with asynchronous gradient transmission in heterogeneous scenarios. Moreover, XDP can provide high-performance and low-latency gradient aggregation. We conduct microbenchmark and testbed with real-world DNN models and datasets. Experimental results show that XAgg improves the gradient aggregation throughput by 3.3$\times$compared with TCP-based aggregation, reaching 100 Gbps with 10 CPU cores. In addition, XAgg reduces communication time by 49%-82% compared with state-of-the-art solutions.
Qianyu Zhang 0001, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022
IEEE/ACM Trans. Netw.4