Yijun Li 0002

dblp:52/6049-2 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0003-4335-8742ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 16 · 3 first-author · 16 since 2021Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FAR: Fast and Accurate Rate Control for Lossless Datacenter Networks
abstract
In recent years, end-to-end congestion control algorithms or flow pausing mechanisms are proposed to achieve high throughput and low latency in datacenter networks. However, prior end-to-end congestion control works without complex signals fail to achieve fast convergence to a stable equilibrium state and effectively handle the transient congestion, while existing flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient states and incomplete queue elimination in equilibrium states. To address these issues, we present FAR, a rate control protocol that combines the advantages of flow pausing and congestion control. At its heart, FAR couples the bandwidth-estimation-based congestion control and the end-to-end flow pausing mechanisms. After flow pausing, FAR quickly explore the available bandwidth with a binary-search probe to achieve high throughput and low latency. Meanwhile, FAR employs a probe staggering mechanism to address the queue oscillation issue in high-concurrency scenarios. We implement the prototype of FAR using DPDK. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67% compared with the state-of-the-art designs.
Jingling Liu, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0001, Ping Zhong 0002, Jiawei Huang 0001
IEEE Trans. Netw.4
2026 SIM: Accelerating Distributed DNN Training by Exploring Gradient Similarity
abstract
Synchronous stochastic gradient descent (SSGD) has been widely used in distributed deep learning. However, since the local gradients need to be shared among workers at every iteration, SSGD performance is significantly influenced by network bottlenecks caused by either heterogeneous environment or bandwidth contention. To solve this problem, asynchronous parallel (ASP) strategy allows each worker to update parameters independently without synchronization, while suffering from accuracy loss and convergence inefficiency. In this paper, we propose a novel similarity-based synchronization scheme called SIM, which mitigates the impact of network bottlenecks and ensures convergence efficiency. Specifically, SIM reduces the number of aggregation workers based on the gradient similarity between global and local gradients, therefore shrinking the waiting time for the stragglers. We provide a theoretical analysis of convergence efficiency and conduct large-scale testbed experiments on CIFAR-10 and SQUAD dataset. The experimental results show that SIM reduces the convergence time of four classical deep learning models by up to 40%.
Jin Ye 0003, Yijun Li 0002, Xiaojuan Lu, Qichen Su, Jiawei Huang 0001, Jianxin Wang 0001
IEEE Trans. Netw.2
2025 ORC: Online Reinforcement Learning for Congestion Control with Fast Convergence
Yijun Li 0002, Jiawei Huang 0001, Chuliang Wu, Jianxin Wang 0001
APNet1
2025 Elastic Scheduling for Mix-Flow in Time-Sensitive Networking
abstract
Time-Sensitive Networking (TSN) is the most promising network infrastructure for various time-critical applications in Industry 4.0. However, industry applications generate a mix of time-triggered (TT) and event-triggered (ET) flows. Scheduling such mix-flows is a key challenge for TSN. Though current TSN scheduling mechanisms commonly provide deterministic transmission for TT flows with stringent latency requirements, they cannot flexibly accommodate ET flows, which are usually generated by emergency events. In this paper, we propose Elastic Backoff (EBO), a systematic solution for scheduling mix-flows in an elastic way. Our key insight is that the network resources should be reasonably allocated for ET flows while minimally impacting TT flows. To this end, we incorporate the elasticity into the TSN scheduling to make resource reservations for ET flows without hurting TT flows. We conduct extensive experiments on both testbeds and simulations. The evaluation results show that, compared with the state-of-the-art designs, EBO improves the schedulability of ET flows by up to 7.5×, while still ensuring the deterministic transmission of TT flows.
Jiawei Huang 0001, Shengwen Zhou, Hui Li 0120, Yijun Li 0002, Qile Wang, Jishu Tian, Kengchang Chen
ICDCS7
2025 Aion: A Memory-Efficient Approach for Long-Term Periodic Flow Detection
abstract
Sketch-based measurement approaches have recently become a promising solution for detecting periodic flows. However, current sketch approaches struggle to achieve accurate detection of periodic flows due to their short-sighted record of the flow arrival information. Recording the long-term information of flow arrivals could mitigate this issue, while the large memory consumption will hurt the detection accuracy. Consequently, achieving a satisfactory trade-off between memory efficiency and detection accuracy remains a tough challenge. To address this issue, we propose Aion for periodic flow detection. Specifically, Aion uses the Sidon sequence to compress historical flow arrival information in multiple time windows into very small size. Based on the periodicity information from successive time windows, Aion updates the estimated frequencies of periodic flows and promptly evicts non-periodic ones to enable accurate detection of frequent periodic flows. We implement Aion on a P4-based testbed and demonstrate that it achieves superior resource efficiency compared to state-of-the-art approaches. Trace-driven evaluations show that Aion improves F1-Score by up to 9.88×, particularly under small memory conditions.
Jiawei Huang 0001, Xianshi Su, Yijun Li 0002, Sitan Li
ICNP5
2025 DACC: Data Augmentation for Learning-based Congestion Control
Jiawei Huang 0001, Yijun Li 0002, Shengwen Zhou, Hui Li 0120, Weihe Li, Jingling Liu, Wanchun Jiang
INFOCOM5
2025 SwitchTop-k: Scaling Top-k Compression on Programmable Switches
abstract
Distributed deep learning has been widely deployed in data centers to provide various services such as image classification and speech recognition. To reduce the training time, Top-k compression has become one of the most popular solutions used to shrink the data volume of gradients. Nevertheless, we observe that existing Top-k compression solutions are inefficient when used for large-scale distributed training due to gradient build-up, missing of Top-k gradients, and high compression overhead at the end hosts. To address these problems, we propose SwitchTop-k, which improves the accuracy of selecting Top-k values while ensuring a high compression rate and zero compression overhead. Specifically, SwitchTop-k offloads the Top-k compression from the end hosts to the programmable switches, thus alleviating the gradient build-up and compression overhead. Meanwhile, we propose a sketch-based solution to achieve high accuracy in selecting global Top-k gradients. We also co-design switch logic and end host logic to improve communication efficiency of uncompressed traffic. Finally, we implement SwitchTop-k on Intel Tofino switches and integrate it with Pytorch. The test results show that SwitchTop-k reduces iteration time by up to 91% compared with existing compression algorithms.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
KDD (2)1
2025 Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and Aggregation
Jiawei Huang 0001, Yijun Li 0002, Jingling Liu, Junxue Zhang 0001, Hui Li 0120, Shengwen Zhou, Xiaojuan Lu, Qichen Su, Jianxin Wang 0001, Chee-Wei Tan 0001, Yong Cui 0001, Kai Chen 0005
USENIX ATC3
2025 Asynchronous Control Based Aggregation Transport Protocol for Distributed Deep Learning
abstract
With the rapid growth scale of dataset and model, the training of deep neural networks (DNN) tends to be deployed in a distributed manner. In the large-scale distributed training, the bottlenecks have gradually moved from computational resources to communication process. Recent researches adopt in-network aggregation (INA) that offloads the gradient aggregation process to programmable switches, thereby reducing network traffic amount and transmission latency. Unfortunately, due to the bandwidth competition in shared training clusters, the straggler will slow down the training efficiency of INA. To address this issue, we propose an Asynchronous Control based Aggregation Transport Protocol (AC-ATP), which makes full use uncongested links to transmit gradients and the switch memory to cache gradients from the fast workers to accelerate the gradient aggregation. Meanwhile, AC-ATP performs congestion control according to the transmission progress of worker and the remaining completion time of the job. The evaluation results of real testbed and large-scale simulations show that AC-ATP reduces the aggregate time by up to 68% and speeds up training in real-world benchmark models.
Jin Ye 0003, Yajun Peng, Yijun Li 0002, Jiawei Huang 0001
IEEE Trans. Computers3
2025 Proactive Transport With High Link Utilization Using Opportunistic Packets in Cloud Data Centers
abstract
To meet the stringent demanding low latency and high throughput of cloud datacenter applications, recent receiver-driven transport protocols transmit only one packet once receiving each credit packet from the receiver to achieve ultra-low queueing delay. However, the round-trip time variation and the highly dynamic background traffic significantly deteriorate the performance of receiver-driven transport protocols, resulting in under-utilized bandwidth. This paper designs a simple yet effective solution called RPO, which retains the advantages of receiver-driven transmission while efficiently utilizing the available bandwidth. Specifically, RPO rationally uses low-priority opportunistic packets to ensure high network utilization without increasing the queueing delay of high-priority normal packets. Furthermore, to tackle the queueing buildup due to line-rate transmission in the first RTT, we design a selective dropping mechanism called SDM to help the majority of small flows complete within only one RTT by prioritizing the first-RTT bursty packets over the packets triggered by grants. We implement RPO in Linux hosts with DPDK. The experimental results show that RPO significantly improves the network utilization by up to 35% over the state-of-the-art schemes, without introducing additional queueing delay. Moreover, RPO integrated with SDM reduces the AFCT of small flows by up to 45% compared with RPO integrated with Aeolus.
Jinbin Hu 0001, Jiawei Huang 0001, Yijun Li 0002, Shuying Rao, Wenchao Jiang, Kai Chen 0005, Jianxin Wang 0001, Tian He 0001
IEEE Trans. Mob. Comput.4
2025 Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine Learning
abstract
Large-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols.
Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
IEEE Trans. Netw.6
2024 Proactive Buffer Management of Shared-Memory Switches for Distributed Deep Learning
abstract
Each output port in a shared memory switch can compete for shared memory pool resources. The allocation strategy of the shared buffer directly affects the ability of each output port to absorb network traffic. Due to the unpredictability of traditional network traffic, existing switch buffer management strategies take a passive approach, allocating buffers to each port only after traffic arrives. This passive response has the problem of untimely buffer allocation and cannot effectively absorb burst traffic. Distributed deep learning follows a specific training pattern, and network traffic exhibits obvious periodic characteristics during transmission. Thus, we propose a Proactive Dynamic Threshold (PDT) strategy, which realizes the pre-adjustment of switch port threshold by detecting the traffic characteristics of distributed training.
Jin Ye 0003, Yajun Peng, Yijun Li 0002, Jiawei Huang 0001
APNet3
2024 A Conditional Diffusion-based Data Augmentation for Anomaly Detection in AIOps
abstract
Data augmentation plays a crucial role in AIOps for enhancing the performance of classification models in scenarios with limited supervision. However, current methods used for generating pseudo-anomaly samples may fail in AIOps: existing data augmentation methods suffer from poor sample quality due to class imbalance, high dimensionality, and high diversity. Inspired by the conditional DDPM, we address the problem by generating realistic anomaly samples between normal and abnormal ones. Unfortunately, due to the lack of pre-trained encoders and the difficulty of determining conditional information, it is hard to directly use conditional DDPM. In this work, we present C-Aug which combines sample mixing and conditional diffusion to overcome the above issues. C-Aug respectively achieves F1-Scores of 0.76, 0.98, and 0.90 on three public datasets, which significantly outperforms the other five baselines.
Jiawei Huang 0001, Hanyu Deng, Yijun Li 0002, Jingling Liu, Qichen Su
CSCWD4
2024 D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center Switches
abstract
Nowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively.
Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001
ICDCS8
2024 Achieving Efficient Scheduling based on Accurate Measurement of Small Flows in Data Center
abstract
In modern data centers, many flow scheduling schemes are proposed to accelerate data transfer and improve user experience. However, these schemes assume ideally the prior knowledge of the flow size information, which, unfortunately, is hard to obtain without modifying data center applications. The sketch-based approaches measure the flow size at switch with a compact memory structure, high throughput, and acceptable accuracy loss. However, existing sketches commonly focus on large or specific flows, while most flows in data center networks are small, resulting in missing or overestimated size information about small flows. We propose Strainer Sketch, which enables accurate and fast measurement of small flows with small memory and flexible deployment in a variety of scheduling algorithms. Specifically, Strainer Sketch uses the hierarchical structure to mitigate hash collisions between large and small flows, and the probabilistic counting algorithm to mitigate overestimation due to hash collisions between small flows. Furthermore, we propose a packet scheduling algorithm SW-PIFO, which provides the flow discrimination for a huge number of small flows by using a limited number of queues. Through the testbed experiments and simulations of typical data center applications, we show that our scheme reduces the small flow completion time (FCT) by up to 56.7 <?TeX $\%$?> Math 1 compared with flow scheduling using classic sketches.
Jiawei Huang 0001, Qile Wang, Yijun Li 0002, Sitan Li, Jingling Liu, Min Zhan, Jianxin Wang 0001
ICPP4
2024 Coupling Congestion Control and Flow Pausing in Data Center Network
abstract
To achieve high throughput and low latency for data center applications, there are two broad lines of work: end-to-end congestion control algorithms and flow pausing mechanisms. It is challenging for end-to-end congestion control algorithms without complex signals to achieve fast convergence to a stable equilibrium state while effectively handling the transient congestion. Additionally, flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient state and incomplete queue elimination in equilibrium state. We propose a transport protocol that combines the advantages of flow pausing and congestion control, called FAR. The key idea is coupling the bandwidth-estimation based congestion control and the end-to-end flow pausing mechanisms. FAR quickly explores the available bandwidth with binary-search based packet train probe to achieve high throughput and low latency. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67 <?TeX $\%$?> Math 1 compared with the state-of-the-art designs.
Jiawei Huang 0001, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0007, Ping Zhong 0002
ICPP4
2024 Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep Learning
abstract
Distributed deep learning has been widely employed to train deep neural network over large-scale dataset. However, the commonly used parameter server architecture suffers from long synchronization time in data-parallel training. Although the existing solutions are proposed to reduce synchronization overhead by breaking the synchronization barriers or limiting the staleness bound, they inevitably experience low convergence efficiency and long synchronization waiting. To address these problems, we propose Gsyn to reduce both synchronization overhead and staleness. Specifically, Gsyn divides workers into multiple groups. The workers in the same group coordinate with each other using the bulk synchronous parallel scheme to achieve high convergence efficiency, and each group communicates with parameter server asynchronously to reduce the synchronization waiting time, consequently increasing the convergence efficiency. Furthermore, we theoretically analyze the optimal number of groups to achieve a good tradeoff between staleness and synchronization waiting. The evaluation test in the realistic cluster with multiple training tasks demonstrates that Gsyn is beneficial and accelerates distributed training by up to 27% over the state-of-the-art solutions.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001
INFOCOM1
2024 SGC: Similarity-Guided Gradient Compression for Distributed Deep Learning
abstract
The collective communication has become the bottleneck of large-scale distributed deep learning due to the huge volume of gradients aggregated during the training process. Despite much recent progress in reducing traffic volume by compressing the stochastic gradients inside each training worker, how to share the inter-worker data redundancy to alleviate communication overhead has remained elusive. In this paper, we reveal that most gradients have a great similarity with close value among training workers. From this hypothesis, we propose a Similarity-guided Gradient Compression framework named SGC which skips aggregating the similar gradients among each worker which utilizes local one rather than average value to save communication expenses. Each worker utilizes local SGC firstly quantifies the similarity of gradients among workers, and then elaborately adjusts the aggregation frequency of similar gradients without hurting DNN model accuracy. Meanwhile, we theoretically analyze the convergency accuracy of SGC. The comprehensive evaluation demonstrates that SGC outperforms the state-of-the-art schemes by up to 47% in convergence time.
Jingling Liu, Jiawei Huang 0001, Yijun Li 0002, Wenjun Lyu, Wenchao Jiang, Jianxin Wang 0001
IWQoS3
2024 Straggler-Aware Gradient Aggregation for Large-Scale Distributed Deep Learning System
abstract
Deep Neural Network (DNN) is a critical component of a wide range of applications. However, with the rapid growth of the training dataset and model size, communication becomes the bottleneck, resulting in low utilization of computing resources. To accelerate communication, recent works propose to aggregate gradients from multiple workers in the programmable switch to reduce the volume of exchanged data. Unfortunately, since using synchronization transmission to aggregate data, current in-network aggregation designs suffer from the straggler problem, which often occurs in shared clusters due to resource contention. To address this issue, we propose a straggler-aware aggregation transport protocol (SA-ATP), which enables the leading worker to leverage the spare computing and storage resources to help the straggling worker. We implement SA-ATP atop clusters using P4-programmable switches. The evaluation results show that SA-ATP reduces the iteration time by up to 57% and accelerates training by up to$1.8\times $in real-world benchmark models.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001
IEEE/ACM Trans. Netw.1
2023 A2TP: Aggregator-aware In-network Aggregation for Multi-tenant Learning
abstract
Distributed Machine Learning (DML) techniques are widely used to accelerate the training of large-scale machine learning models. However, during training iterations, gradients need to be frequently aggregated across multiple workers, resulting in communication bottleneck. To reduce the communication overhead of DML, several In-Network Aggregation (INA) protocols are proposed to reduce the volume of aggregation traffic by offloading aggregation functions into switches, thus alleviating network bottlenecks. Nevertheless, these protocols couple the congestion control of in-switch aggregator resources and link bandwidth resources, together with the straggler-oblivious manner in aggregator allocation, leading to low aggregation efficiency.
Jiawei Huang 0001, Yijun Li 0002, Aikun Xu, Shengwen Zhou, Jingling Liu, Jianxin Wang 0001
EuroSys3
2023 PA-ATP: Progress-Aware Transmission Protocol for In-Network Aggregation
abstract
Large-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols.
Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
ICNP6
2023 Traffic-aware rate control for mix-flow in datacenter
abstract
Abstract Datacenter applications generate diverse flows, including deadline flows and non‐deadline flows. The deadline flows require to complete within strict deadline, while non‐deadline flows seek a shorter flow completion time. The state‐of‐the‐art deadline‐aware methods either transmit deadline flows with best‐effort at high priority, resulting in the starvation of non‐deadline flows, or blindly restrict the sending rates of deadline flows, leading to a high deadline missing ratio. To meet the different requirements of mix‐flows, a novel traffic‐aware rate control (TRC) method is proposed. TRC dynamically adjusts the sending rates of deadline flows according to their deadlines and the predicted future traffic patterns. If the intense competition is predicted among deadline flows, TRC will adopt a more aggressive manner to transmit the current deadline flows to avoid bandwidth contention in the future, reducing the deadline missing ratio. Otherwise, TRC will conservatively transmit deadline flows and complete these flows near their respective deadlines, relinquishing the excess bandwidth to non‐deadline flows. Meanwhile, TRC schedules non‐deadline flows in accordance with their sizes, minimizing the average FCT. The performance of TRC in large‐scale scenarios is evaluated through NS2 simulations. The test results show that TRC reduces the deadline missing ratio of deadline flows and the FCT of non‐deadline flows by up to 69.5% and 78.7% compared to the state‐of‐the‐art deadline‐aware schemes, respectively.
Jiawei Huang 0001, Yijun Li 0002, Jianxin Wang 0001
IET Commun.3
2023 ChainSketch: An Efficient and Accurate Sketch for Heavy Flow Detection
abstract
Identifying heavy flows is essential for network management. However, it is challenging to detect heavy flow quickly and accurately under the highly dynamic traffic and rapid growth of network capacity. Existing heavy flow detection schemes can make a trade-off in efficiency, accuracy and speed. However, these schemes still require memory large enough to obtain acceptable performance. To address this issue, we propose ChainSketch, which has the advantages of good memory efficiency, high accuracy and fast detection. Specifically, ChainSketch uses the selective replacement strategy to mitigate the over-estimation issue. Meanwhile, ChainSketch utilizes the hash chain and compact structure to improve memory efficiency. We implement the ChainSketch on OVS platform, P4-based testbed and large-scale simulations to process heavy hitter and heavy changer detection. The results of trace-driven tests show that, ChainSketch greatly improves the F1-score by up to$3.43\times $compared with the state-of-the-art solutions especially for small memory.
Jiawei Huang 0001, Wenlu Zhang, Yijun Li 0002, Jin Ye 0003, Jianxin Wang 0001
IEEE/ACM Trans. Netw.3
2022 HSP: Hybrid Synchronous Parallelism for Fast Distributed Deep Learning
abstract
In the parameter-server-based distributed deep learning system, the workers simultaneously communicate with the parameter server to refine model parameters, easily resulting in severe network contention. To solve this problem, Asynchronous Parallel (ASP) strategy enables each worker to update the parameter independently without synchronization. However, due to the inconsistency of parameters among workers, ASP experiences accuracy loss and slow convergence. In this paper, we propose Hybrid Synchronous Parallelism (HSP), which mitigates the communication contention without excessive degradation of convergence speed. Specifically, the parameter server sequentially pulls gradients from workers to eliminate network congestion and synchronizes all up-to-date parameters after each iteration. Meanwhile, HSP cautiously lets idle workers to compute with out-of-date weights to maximize the utilizations of computing resources. We provide theoretical analysis of convergence efficiency and implement HSP on popular deep learning (DL) framework. The test results show that HSP improves the convergence speedup of three classical deep learning models by up to 67%.
Yijun Li 0002, Jiawei Huang 0001, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001
ICPP1
2022 UA-Sketch: An Accurate Approach to Detect Heavy Flow based on Uninterrupted Arrival
abstract
Heavy flow detection in enormous network traffic is a critical task for network measurement. Due to the limited memory size and high link capacity, accurate detection of heavy flows becomes challenging in large-scale networks. Almost all existing approaches of detecting heavy flows use single-dimension statistics of flow size to make flow-replacement decisions. However, under the mass number of small flows, the heavy flows are prone to be frequently and mistakenly replaced, resulting in unsatisfactory accuracy. To solve this problem, we reveal that the number of uninterrupted arrival packets is a useful metric in identifying flow types. We further propose UA-Sketch that expels small flows and protects heavy ones according to the multiple-dimension statistics including both estimated flow size and number of uninterrupted arrival packets. The test results of trace-driven simulations and OVS experiments show that, even under small memory, UA-Sketch achieves higher accuracy than the existing works, with the F1 Score by up to 2.1 ×.
Jin Ye 0003, Wenlu Zhang, Guihao Chen, Yuanchao Shan, Yijun Li 0002, Weihe Li, Jiawei Huang 0001
ICPP6
2022 Achieving Per-Flow Fairness and High Utilization With Limited Priority Queues in Data Center
abstract
Modern data centers often host multiple applications with diverse network demands. To provide fair bandwidth allocation to several thousand traversing flows, Approximate Fair Queueing (AFQ) utilizes multiple priority queues in switch to approximate ideal fair queueing. However, due to limited number of queues in programmable switches, AFQ easily experiences high packet loss and low link utilization. In this paper, we propose Elastic Fair Queueing (EFQ), which leverages limited priority queues to flexibly achieve both high network utilization and fair bandwidth allocation. EFQ dynamically assigns the free buffer space in priority queues for each packet to obtain high utilization without sacrificing flow-level fairness. The results of simulation experiments and real implementations show that EFQ reduces the average flow completion time by up to 82% over the state-of-the-art fair bandwidth allocation mechanisms.
Jingling Liu, Jiawei Huang 0001, Yijun Li 0002, Jianxin Wang 0001, Tian He 0001
IEEE/ACM Trans. Netw.4
2021 RPO: Receiver-driven Transport Protocol Using Opportunistic Transmission in Data Center
abstract
Modern datacenter applications bring fundamental challenges to transport protocols as they simultaneously require low latency and high throughput. Recent receiver-driven trans-port protocols transmit only one data packet once receiving each grant or credit packet from the receiver to achieve ultra-low queueing delay and zero packet loss. However, the round-trip time variation and the highly dynamic background traffic significantly deteriorate the performance of receiver-driven transport protocols, resulting in under-utilized bandwidth. This paper designs a simple yet effective solution called RPO that retains the advantages of receiver-driven transmission while efficiently utilizing the available bandwidth. Specifically, RPO rationally uses low-priority opportunistic packets to ensure high network utilization without increasing the queueing delay of high-priority normal packets. In addition, since RPO only uses Explicit Congestion Notification (ECN) marking function and priority queues, RPO is ready to deploy on switches. We implement RPO in Linux hosts with DPDK. Our small-scale testbed experiments and large-scale simulations show that RPO significantly improves the network utilization by up to 35% under high workload over the state-of-the-art receiver-driven transmission schemes, without introducing additional queueing delay.
Jinbin Hu 0001, Jiawei Huang 0001, Yijun Li 0002, Wenchao Jiang, Kai Chen 0005, Jianxin Wang 0001, Tian He 0001
ICNP4