Jingling Liu

dblp:160/7597 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0001-8743-0270ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 17 · 7 first-author · 15 since 2021Systems, architecture and hardware · 12 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Achieving accurate and stateless multicast with customized hash function in data center
Zhidong He, Jiawei Huang 0001, Jingling Liu
Comput. Networks4
2026 FAR: Fast and Accurate Rate Control for Lossless Datacenter Networks
abstract
In recent years, end-to-end congestion control algorithms or flow pausing mechanisms are proposed to achieve high throughput and low latency in datacenter networks. However, prior end-to-end congestion control works without complex signals fail to achieve fast convergence to a stable equilibrium state and effectively handle the transient congestion, while existing flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient states and incomplete queue elimination in equilibrium states. To address these issues, we present FAR, a rate control protocol that combines the advantages of flow pausing and congestion control. At its heart, FAR couples the bandwidth-estimation-based congestion control and the end-to-end flow pausing mechanisms. After flow pausing, FAR quickly explore the available bandwidth with a binary-search probe to achieve high throughput and low latency. Meanwhile, FAR employs a probe staggering mechanism to address the queue oscillation issue in high-concurrency scenarios. We implement the prototype of FAR using DPDK. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67% compared with the state-of-the-art designs.
Jingling Liu, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0001, Ping Zhong 0002, Jiawei Huang 0001
IEEE Trans. Netw.1
2025 SOLB: Synchronization-Objective Load Balancing for Distributed DNN Training
abstract
Distributed training is the most common way to scale out and accelerate Deep Neural Network (DNN) training. Distributed DNN training requires synchronization of gradient aggregation among all workers through collective communication operations before proceeding to the next training round. Any long-tail delay would be an obstacle to model training performance and convergence. This paper proposes a load balancing scheme SOLB for distributed DNN training acceleration, which utilizes a synchronization-objective load balancing to ensure packets with the same gradient indices from the same training job arrive at the receiver back to back, enabling optimal synchronous communication. Furthermore, leveraging the independence of parameter updates, we design an order-agnostic transmission protocol to avoid the overhead of packet reordering without affecting the training accuracy. Through both testbed and simulation experiments, our scheme significantly reduces gradient aggregation time by up to 82.5% and accelerates overall model training by up to 84.4% compared to the state-of-the-art load balancing schemes.
Jingling Liu, Zhong He, Rui Cui
ICDCS1
2025 DACC: Data Augmentation for Learning-based Congestion Control
Jiawei Huang 0001, Yijun Li 0002, Shengwen Zhou, Hui Li 0120, Weihe Li, Jingling Liu, Wanchun Jiang
INFOCOM9
2025 Sadra: Size-Aware Demotion Rate Adjustment for Flow Scheduling in Data Center
abstract
Most existing flow scheduling schemes aim to minimize the flow completion time in data center network. However, these schemes either lack fine-grained flow differentiation, sacrificing performance for deployability or require precise flow information and significant hardware modifications for nearoptimal transmission. Thus, we present Sadra, a novel flow scheduling solution leveraging multiple priority queues available in existing commodity switches to minimize FCT by assigning different initial priorities and demotion thresholds to flows with various sizes. Our intuition is that the average queuing length of small flows increases as the number of large flows grows, yet reducing queuing length improves latency of small flows with negligible impact on throughput of large flows. Sadra consists of two key parts: i) A novel scheduling scheme that utilizes imprecise predictions to assign initial priorities to flows with different sizes, enabling differentiation from the outset of transmission; ii) A size-aware demotion threshold strategy that assigns different demotion thresholds to further differentiate flows with the same initial priority, thereby reducing the blocking of small flows by large ones. We have implemented a Sadra prototype and evaluated Sadra through both testbed experiments and NS3 simulations. Simulation and testbed evaluations show that Sadra significantly reduces the average and tail FCT of small flows by up to 61% and 89% compared with the state-of-the-art schemes, respectively.
Jingling Liu, Rui Cui, Zhong He, Wenjun Lyu
IWQoS1
2025 SwitchTop-k: Scaling Top-k Compression on Programmable Switches
abstract
Distributed deep learning has been widely deployed in data centers to provide various services such as image classification and speech recognition. To reduce the training time, Top-k compression has become one of the most popular solutions used to shrink the data volume of gradients. Nevertheless, we observe that existing Top-k compression solutions are inefficient when used for large-scale distributed training due to gradient build-up, missing of Top-k gradients, and high compression overhead at the end hosts. To address these problems, we propose SwitchTop-k, which improves the accuracy of selecting Top-k values while ensuring a high compression rate and zero compression overhead. Specifically, SwitchTop-k offloads the Top-k compression from the end hosts to the programmable switches, thus alleviating the gradient build-up and compression overhead. Meanwhile, we propose a sketch-based solution to achieve high accuracy in selecting global Top-k gradients. We also co-design switch logic and end host logic to improve communication efficiency of uncompressed traffic. Finally, we implement SwitchTop-k on Intel Tofino switches and integrate it with Pytorch. The test results show that SwitchTop-k reduces iteration time by up to 91% compared with existing compression algorithms.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
KDD (2)3
2025 On the Stability and Generalization of Meta-Learning: the Impact of Inner-Levels
abstract
Meta-learning has achieved significant advancements, with generalization emerging as a key metric for evaluating meta-learning algorithms. While recent studies have mainly focused on training strategies, data-split methods, and tightening generalization bounds, they often ignore the impact of inner-levels on generalization. To bridge this gap, this paper focuses on several prominent meta-learning algorithms and establishes two generalization analytical frameworks for them based on their inner-processes: the Gradient Descent Framework (GDF) and the Proximal Descent Framework (PDF). Within these frameworks, we introduce two novel algorithmic stability definitions and derive the corresponding generalization bounds. Our findings reveal a trade-off of inner-levels under GDF, whereas PDF exhibits a beneficial relationship. Moreover, we highlight the critical role of the meta-objective function in minimizing generalization error. Inspired by this, we propose a new, simplified meta-objective function definition to enhance generalization performance. Many real-world experiments support our findings and show the improvement of the new meta-objective function.
Wenjun Ding, Jingling Liu, Lixing Chen, Xiu Su
NeurIPS2
2025 Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and Aggregation
Jiawei Huang 0001, Yijun Li 0002, Jingling Liu, Junxue Zhang 0001, Hui Li 0120, Shengwen Zhou, Xiaojuan Lu, Qichen Su, Jianxin Wang 0001, Chee-Wei Tan 0001, Yong Cui 0001, Kai Chen 0005
USENIX ATC4
2025 Tile-size aware bitrate allocation for adaptive 360$^{\circ }$ video streaming
Jiawei Huang 0001, Jingling Liu, Feng Gao 0001, Weihe Li, Jianxin Wang 0001
Multim. Tools Appl.3
2025 Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine Learning
abstract
Large-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols.
Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
IEEE Trans. Netw.7
2025 Automatic Dual Threshold Tuning for Switch Buffer Sharing in Datacenter Networking
abstract
For the widely deployed on-chip shared buffer, efficient buffer management is the key to absorbing bursts and avoiding packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. What’s more, we introduce D2T${}^{*}$which combines D2T with advanced DRL techniques to move toward mastering buffer management for further improving performance across various scenarios. We implement D2T at a P4-programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively.
Jingling Liu, Hui Li 0120, Jiawei Huang 0001, Ping Zhong 0002, Boyan Huang, Pingping Dong, Wensheng Tang, Wanchun Jiang, Jianxin Wang 0001, Yong Cui 0001
IEEE Trans. Netw.1
2024 A Conditional Diffusion-based Data Augmentation for Anomaly Detection in AIOps
abstract
Data augmentation plays a crucial role in AIOps for enhancing the performance of classification models in scenarios with limited supervision. However, current methods used for generating pseudo-anomaly samples may fail in AIOps: existing data augmentation methods suffer from poor sample quality due to class imbalance, high dimensionality, and high diversity. Inspired by the conditional DDPM, we address the problem by generating realistic anomaly samples between normal and abnormal ones. Unfortunately, due to the lack of pre-trained encoders and the difficulty of determining conditional information, it is hard to directly use conditional DDPM. In this work, we present C-Aug which combines sample mixing and conditional diffusion to overcome the above issues. C-Aug respectively achieves F1-Scores of 0.76, 0.98, and 0.90 on three public datasets, which significantly outperforms the other five baselines.
Jiawei Huang 0001, Hanyu Deng, Yijun Li 0002, Jingling Liu, Qichen Su
CSCWD5
2024 D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center Switches
abstract
Nowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively.
Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001
ICDCS5
2024 Achieving Efficient Scheduling based on Accurate Measurement of Small Flows in Data Center
abstract
In modern data centers, many flow scheduling schemes are proposed to accelerate data transfer and improve user experience. However, these schemes assume ideally the prior knowledge of the flow size information, which, unfortunately, is hard to obtain without modifying data center applications. The sketch-based approaches measure the flow size at switch with a compact memory structure, high throughput, and acceptable accuracy loss. However, existing sketches commonly focus on large or specific flows, while most flows in data center networks are small, resulting in missing or overestimated size information about small flows. We propose Strainer Sketch, which enables accurate and fast measurement of small flows with small memory and flexible deployment in a variety of scheduling algorithms. Specifically, Strainer Sketch uses the hierarchical structure to mitigate hash collisions between large and small flows, and the probabilistic counting algorithm to mitigate overestimation due to hash collisions between small flows. Furthermore, we propose a packet scheduling algorithm SW-PIFO, which provides the flow discrimination for a huge number of small flows by using a limited number of queues. Through the testbed experiments and simulations of typical data center applications, we show that our scheme reduces the small flow completion time (FCT) by up to 56.7 <?TeX $\%$?> Math 1 compared with flow scheduling using classic sketches.
Jiawei Huang 0001, Qile Wang, Yijun Li 0002, Sitan Li, Jingling Liu, Min Zhan, Jianxin Wang 0001
ICPP8
2024 Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep Learning
abstract
Distributed deep learning has been widely employed to train deep neural network over large-scale dataset. However, the commonly used parameter server architecture suffers from long synchronization time in data-parallel training. Although the existing solutions are proposed to reduce synchronization overhead by breaking the synchronization barriers or limiting the staleness bound, they inevitably experience low convergence efficiency and long synchronization waiting. To address these problems, we propose Gsyn to reduce both synchronization overhead and staleness. Specifically, Gsyn divides workers into multiple groups. The workers in the same group coordinate with each other using the bulk synchronous parallel scheme to achieve high convergence efficiency, and each group communicates with parameter server asynchronously to reduce the synchronization waiting time, consequently increasing the convergence efficiency. Furthermore, we theoretically analyze the optimal number of groups to achieve a good tradeoff between staleness and synchronization waiting. The evaluation test in the realistic cluster with multiple training tasks demonstrates that Gsyn is beneficial and accelerates distributed training by up to 27% over the state-of-the-art solutions.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001
INFOCOM4
2024 SGC: Similarity-Guided Gradient Compression for Distributed Deep Learning
abstract
The collective communication has become the bottleneck of large-scale distributed deep learning due to the huge volume of gradients aggregated during the training process. Despite much recent progress in reducing traffic volume by compressing the stochastic gradients inside each training worker, how to share the inter-worker data redundancy to alleviate communication overhead has remained elusive. In this paper, we reveal that most gradients have a great similarity with close value among training workers. From this hypothesis, we propose a Similarity-guided Gradient Compression framework named SGC which skips aggregating the similar gradients among each worker which utilizes local one rather than average value to save communication expenses. Each worker utilizes local SGC firstly quantifies the similarity of gradients among workers, and then elaborately adjusts the aggregation frequency of similar gradients without hurting DNN model accuracy. Meanwhile, we theoretically analyze the convergency accuracy of SGC. The comprehensive evaluation demonstrates that SGC outperforms the state-of-the-art schemes by up to 47% in convergence time.
Jingling Liu, Jiawei Huang 0001, Yijun Li 0002, Wenjun Lyu, Wenchao Jiang, Jianxin Wang 0001
IWQoS1
2024 Learning Audio and Video Bitrate Selection Strategies via Explicit Requirements
abstract
Mobile video streaming dominates today's network traffic, and adaptive bitrate (ABR) algorithms have been routinely adopted for transmitting media content across dynamic mobile networks. State-of-the-art ABR algorithms mainly alter video bitrate without considering audio bitrate as they consider the impact on the video negligible due to their small size. However, to bring users an immersive experience, recent content providers have applied high-quality audio with large sizes, like stereophonic sound. Therefore, improper audio bitrate selection will adversely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, and vice versa) and frequent playback interruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bitrate selections. By learning from explicit goals, SPA can match the actual requirements and attain good performance. By conducting trace-driven and testbed-based experiments, we observe SPA's considerable superiority compared to existing approaches, including reducing the undesirable combinations by up to 34.17× and achieving zero stall time across 88.57% of traces. We also invite 35 volunteers to join a subjective test, and the result shows that 33/35 people consider SPA provides them with a satisfactory viewing experience.
Weihe Li, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
IEEE Trans. Mob. Comput.3
2024 Optimizing Video Streaming in Dynamic Networks: An Intelligent Adaptive Bitrate Solution Considering Scene Intricacy and Data Budget
abstract
Adaptive Bitrate (ABR) algorithms have become increasingly important for delivering high-quality video content over fluctuating networks. Considering the complexity of video scenes, video chunks can be separated into two categories: those with intricate scenes and those with simple scenes. In practice, it has been observed that improving the quality of intricate chunks yields more substantial improvements in Quality of Experience (QoE) compared with focusing solely on simple chunks. However, the current ABR schemes either treat all chunks equally or rely on fixed linear-based reward functions, which limits their ability to meet real-world requirements. To tackle these limitations, this paper introduces a novel ABR approach called CAST (Complex-scene Aware bitrate algorithm via Self-play reinforcemenT learning), which considers the scene complexity and formulates the bitrate adaptation task as an explicit objective. Leveraging the power of parallel computing with multiple agents, CAST trains a neural network to achieve superior video playback quality for intricate scenes while minimizing playback freezing time. Moreover, we also introduce a new variant of our proposed approach called CAST-DU, to address the critical issue of efficiently managing users' limited cellular data budgets while ensuring a satisfactory viewing experience. Furthermore, we present CAST-Live, tailored for live streaming scenarios with constrained playback buffers and considerations for energy costs. Extensive trace-driven evaluations and subjective tests demonstrate that CAST, CAST-DU, and CAST-Live outperform existing off-the-shelf schemes, delivering a superior video streaming experience over fluctuating networks while efficiently utilizing data resources. Moreover, CAST-Live demonstrates effectiveness even under limited buffer size constraints while incurring minimal energy costs.
Weihe Li, Jiawei Huang 0001, Qichen Su, Jingling Liu, Wenjun Lyu, Jianxin Wang 0001
IEEE Trans. Mob. Comput.5
2024 Straggler-Aware Gradient Aggregation for Large-Scale Distributed Deep Learning System
abstract
Deep Neural Network (DNN) is a critical component of a wide range of applications. However, with the rapid growth of the training dataset and model size, communication becomes the bottleneck, resulting in low utilization of computing resources. To accelerate communication, recent works propose to aggregate gradients from multiple workers in the programmable switch to reduce the volume of exchanged data. Unfortunately, since using synchronization transmission to aggregate data, current in-network aggregation designs suffer from the straggler problem, which often occurs in shared clusters due to resource contention. To address this issue, we propose a straggler-aware aggregation transport protocol (SA-ATP), which enables the leading worker to leverage the spare computing and storage resources to help the straggling worker. We implement SA-ATP atop clusters using P4-programmable switches. The evaluation results show that SA-ATP reduces the iteration time by up to 57% and accelerates training by up to$1.8\times $in real-world benchmark models.
Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001
IEEE/ACM Trans. Netw.4
2023 MEB: an Efficient and Accurate Multicast using Bloom Filter with Customized Hash Function
abstract
Multicast is widely used to support a huge range of applications with one-to-many or many-to-many communication patterns. However, multicast systems do not scale due to considerable state and communication overheads. Some stateful multicast approaches require maintaining the state of each multicast session at switches, thus incurring large memory overhead. Some stateless ones utilize Bloom filter (BF) to encode multicast tree into the packet header to minimize communication overhead, but potentially suffer from the substantial false positive due to the probabilistic nature of Bloom filter. In this paper, we propose a stateless multicast scheme MEB, which uses Bloom filter to achieve large-scale multicast communication with low error, small overhead and high scalability. Specifically, to control the rate of false positive, MEB elaborately selects the hash functions for Bloom filters when constructing the packet header at the sender side, and makes forwarding decision according to packet header at the switch with negligible overhead. We compare MEB against the state-of-the-art multicast system in large-scale simulations. The test results show that MEB reduces the traffic overhead by up to 70% with small error rate.
Jiawei Huang 0001, Qile Wang, Jingling Liu, Shengwen Zhou, Zhidong He
APNet4
2023 A2TP: Aggregator-aware In-network Aggregation for Multi-tenant Learning
abstract
Distributed Machine Learning (DML) techniques are widely used to accelerate the training of large-scale machine learning models. However, during training iterations, gradients need to be frequently aggregated across multiple workers, resulting in communication bottleneck. To reduce the communication overhead of DML, several In-Network Aggregation (INA) protocols are proposed to reduce the volume of aggregation traffic by offloading aggregation functions into switches, thus alleviating network bottlenecks. Nevertheless, these protocols couple the congestion control of in-switch aggregator resources and link bandwidth resources, together with the straggler-oblivious manner in aggregator allocation, leading to low aggregation efficiency.
Jiawei Huang 0001, Yijun Li 0002, Aikun Xu, Shengwen Zhou, Jingling Liu, Jianxin Wang 0001
EuroSys6
2023 CAST: An Intricate-Scene Aware Adaptive Bitrate Approach for Video Streaming via Parallel Training
Weihe Li, Jiawei Huang 0001, Jingling Liu, Wenlu Zhang, Wenjun Lyu, Jianxin Wang 0001
ICA3PP (4)4
2023 PA-ATP: Progress-Aware Transmission Protocol for In-Network Aggregation
abstract
Large-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols.
Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001
ICNP7
2023 Achieving Fast Convergence and High Efficiency using Differential Explicit Feedback in Data Center
abstract
Since most flows are short-lived in data center networks, fast convergence becomes very important to help the short flows effectively utilize high bandwidth. Though current explicit feedback-based transport control protocols (TCPs) provide fast convergence via fine-grained congestion information from customized switches, they unavoidably incur large traffic overhead for widely existing small packets in data center applications, resulting in suboptimal network efficiency. To solve this issue, we propose a datacenter TCP based onDifferentialExplicitCongestionNotification, called DECN, to achieve fast convergence without any traffic overhead. Specifically, DECN feeds rate difference between the target and current rate back to the source by using multiple consecutive packets. Besides, we propose an enhanced version DECN* which obtains the optimal number of consecutive packets according to the packet loss rate. The experimental results of NS2 simulation and testbed implementation show that DECN and its enhanced version DECN* achieve comparable fast convergence as XCP without incurring any extra feedback overhead. Compared with the state-of-the-art explicit feedback-based TCPs, they reduce the flow completion time by up to 34% in typical data center applications.
Jiawei Huang 0001, Jingling Liu, Sen Liu 0002, Jinbin Hu 0001, Jianxin Wang 0001
IEEE Trans. Cloud Comput.2
2023 REN: Receiver-Driven Congestion Control Using Explicit Notification for Data Center
abstract
In recent years, receiver-driven transport protocols have been proposed to use proactive congestion control to meet the stringent latency requirements of large-scale applications in data center. However, the receiver-driven proposals face the challenges brought by network dynamic. First, when the bursty flows start, the aggressive and blind line-rate transmission in the first RTT easily leads to persistent queue backlog. Second, when some flows finish transmissions, the remaining ones cannot increase their sending rates to seize the available bandwidth. To address these problems, this article presents a new receiver-driven congestion control design, called REN, which uses the under- and over-utilization notifications from switch to handle the dynamic traffic. With the aid of explicit feedback, REN alleviates the traffic burstiness due to aggressive start, mitigates the conservativeness in utilizing available bandwidth, and still retains the receiver-driven feature to achieve ultra-low latency. We implement the prototype of REN using DPDK. The experimental results of real testbed and large-scale NS2 simulation show that REN effectively reduces the average flow completion time (AFCT) by up to 68% over the state-of-the-art receiver-driven transmission schemes.
Jiawei Huang 0001, Jinbin Hu 0001, Weihe Li, Tao Zhang 0019, Jingling Liu, Jianxin Wang 0001, Tian He 0001
IEEE Trans. Cloud Comput.6
2023 Asymmetry-Aware Load Balancing With Adaptive Switching Granularity in Data Center
abstract
Datacenter networks provide large bisection bandwidth by load balancing traffic over rich parallel paths in multi-rooted tree topologies. Nevertheless, production datacenters operate under various path diversities caused by traffic dynamics, hardware failures and heterogeneous switching equipment. Therefore, the load balancing schemes in data center should be resilient to network asymmetry. Prior fine-grained schemes such as RPS and Presto are prone to experience packet reordering problem under asymmetric topology since they split flows into small units which are spread across all parallel paths. The coarse-grained solutions such as ECMP and LetFlow effectively avoid packet reordering, but easily leading to under-utilization of multiple paths. To solve these problems, we propose a load balancing mechanism called AG, which adaptively adjusts switching granularity according to the asymmetric degree of multiple paths. AG increases switching granularity to alleviate packet reordering under large degrees of topology asymmetry, while reducing switching granularity to obtain high link utilization under small degrees of topology asymmetry. Moreover, we design a switch-based scheme which measures the difference of one-way delay of multiple paths to obtain accurate state of topology asymmetry with low overhead. AG is a practical switch-based solution without modification at end hosts. The experimental results of NS2 simulations and real implementation show that AG reduces the average and$99^{th}$flow completion time by up to 54% and 65% compared with the state-of-the-art load balancing schemes, respectively.
Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001, Tian He 0001
IEEE/ACM Trans. Netw.1
2022 Synthesizing Audio and Video Bitrate Selections via Learning from Actual Requirements
abstract
Adaptive bitrate (ABR) algorithms are routinely adopted for transmitting media contents across dynamic networks. State-of-the-art ABR algorithms only adapt to video bitrate without considering audio bitrate adaption as they consider the im-pact on the video to be negligible due to the small size of the audio. However, to bring users an immersive experience, more and more content providers have applied high-quality audio with large sizes, like stereophonic and surround (Dolby Atmos). Therefore, improper audio bitrate selection will ad-versely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, vice versa) and frequent playback inter-ruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bi-trate selections. Experimental results demonstrate SPA's con-siderable superiority as compared with existing approaches.
Weihe Li, Jiawei Huang 0001, Jingling Liu, Feng Gao 0001
ICME4
2022 APS: Adaptive Packet Spraying to Isolate Mix-Flows in Data Center Network
abstract
Modern data centers host diverse applications, which generate a mix of short flows with stringent latency requirement and long flows requiring large sustained throughput. To solve the problem of resource competition between the mixed flows, we propose an adaptive traffic isolation scheme APS. Based on the packet spraying scheme in the multipath transmission, APS dynamically separates long flows from short ones on different paths to provide the low latency for the short flows. Meanwhile, to resolve the out-of-order problem, APS limits the long flows to a few paths with Equal Cost Multi Path (ECMP). Experimental results of NS2 simulation and testbed implementation show that, APS reduces the average completion time for short flows by up to 60 percent and increases the throughputs for long flows by about 1.68x over the state-of-the-art multipath transmission schemes.
Jingling Liu, Jiawei Huang 0001, Wenjun Lv, Jianxin Wang 0001
IEEE Trans. Cloud Comput.1
2022 Achieving Per-Flow Fairness and High Utilization With Limited Priority Queues in Data Center
abstract
Modern data centers often host multiple applications with diverse network demands. To provide fair bandwidth allocation to several thousand traversing flows, Approximate Fair Queueing (AFQ) utilizes multiple priority queues in switch to approximate ideal fair queueing. However, due to limited number of queues in programmable switches, AFQ easily experiences high packet loss and low link utilization. In this paper, we propose Elastic Fair Queueing (EFQ), which leverages limited priority queues to flexibly achieve both high network utilization and fair bandwidth allocation. EFQ dynamically assigns the free buffer space in priority queues for each packet to obtain high utilization without sacrificing flow-level fairness. The results of simulation experiments and real implementations show that EFQ reduces the average flow completion time by up to 82% over the state-of-the-art fair bandwidth allocation mechanisms.
Jingling Liu, Jiawei Huang 0001, Yijun Li 0002, Jianxin Wang 0001, Tian He 0001
IEEE/ACM Trans. Netw.1
2021 Mitigating Port Starvation for Shallow-buffered Switches in Datacenter Networks
abstract
Explicit Congestion Notification (ECN) is widely utilized in modern data centers to achieve low latency and high throughput for various applications. In recent years, however, even with the sustainable growth of link bandwidth in data centers, the switch buffer size does not increase remarkably. Consequently, the standard per-port ECN scheme suffers from excessive packet loss. Though the shared-buffer ECN scheme alleviates the packet loss, we observe that it leads to severe unfairness, which we term as the Port Starvation problem. When flows destined for some ports have aggressively occupied the shared buffer, the later-arrival flows destined for other ports will be ECN-marked unfairly and obtain significantly lower throughput. To address the port starvation problem, we design a buffer-aware fair ECN-marking (BFEM) scheme for shallow-buffered switch. BFEM leverages the shared buffer to reduce packet loss and meanwhile punishes aggressive flows by ECN marking. We evaluate BFEM with both 40Gbps P4 testbed implementation and large-scale NS2 simulation. The test results show that, by improving fairness between egress ports, BFEM increases total link utilization and reduces the average flow completion time by up to 40% compared with the state-of-the-art per-port and shared-buffer ECN marking schemes.
Wenjun Lyu, Jiawei Huang 0001, Jingling Liu, Shaojun Zou, Weihe Li, Jianxin Wang 0001, Desheng Zhang 0002
ICDCS3
2021 GTCP: Hybrid Congestion Control for Cross-Datacenter Networks
abstract
To improve the quality of experience for worldwide users, an increasing number of service providers deploy their services on geographically dispersed data centers, which are connected by wide area network (WAN). In the cross-datacenter networks, however, the intra- and inter-datacenter parts have different characteristics, including switch buffer depth, round-trip time and bandwidth. Besides, most of intra-DC flows belong to interactive services that require low delay while inter-DC flows typically need to achieve high throughput. Unfortunately, existing sender-based and receiver-driven transport protocols do not consider the network heterogeneity between inter- and intra- DC networks so that they fail to simultaneously achieve low latency for intra-DC flows and high throughput for inter-DC flows. This paper proposes a general hybrid congestion control mechanism called GTCP to address this problem. When the inter-DC flow detects congestion inside data center, it switches to the receiver-driven mode to avoid the impact on intra-DC flows. Otherwise, it switches back to the sender-based mode to proactively explore the available bandwidth. Besides, the intra-DC flow leverages the pausing mechanism to eliminate the queue build-up. Through a series of testbed experiments and large-scale NS2 simulations, we demonstrate that GTCP reduces flow completion time by up to 79.3% compared with existing protocols.
Shaojun Zou, Jiawei Huang 0001, Jingling Liu, Tao Zhang 0019, Jianxin Wang 0001
ICDCS3
2020 Achieving High Utilization for Approximate Fair Queueing in Data Center
abstract
Modern data centers often host multiple applications with diverse network demands. To provide fair bandwidth allocation to several thousand traversing flows, Approximate Fair Queueing (AFQ) utilizes multiple priority queues in switch to approximate ideal fair queueing. However, due to limited number of queues in commodity switches, AFQ easily experiences high packet loss and low link utilization. In this paper, we propose Elastic Fair Queueing (EFQ), which leverages limited priority queues to flexibly achieve both high network utilization and fair bandwidth allocation. EFQ dynamically assigns the free buffer space in priority queues for each packet to obtain high utilization without sacrificing flow-level fairness. The results of simulation experiments and real implementations show that EFQ reduces the average flow completion time by up to 82% over the state-of-the-art fair bandwidth allocation mechanisms.
Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001
ICDCS1
2019 DDT: Mitigating the Competitiveness Difference of Data Center TCPs
abstract
To achieve better network performance, the cloud service providers are widely deploying the ECN-based transport protocols (i.e., DCTCP) in their data center networks (DCN). In multi-tenant environment, however, the newly introduced ECN-enabled TCP greatly impairs the performance of applications with out-dated and miscon figured TCP stacks. The reason is that the ECN-enabled datacenter switch fails to treat the mixed TCP traffic fairly, causing the distinguished performance gap between the ECN-enabled and ECN-disabled TCPs. This paper proposes DDT (Dual Dynamic Thresholds), an active queue management algorithm (AQM) that aims to achieve the flow-level fairness when the heterogeneous TCP traffic coexists. DDT monitors the switch queue in real time, and dynamically tunes the distance between ECN-marking and packet-dropping thresholds to mitigate the competitiveness difference between the ECN-enabled and ECN-disabled TCP. Our preliminary real implementations and testing results show that DDT elegantly fills the competitiveness gap of heterogeneous TCP traffic without disturbing their own control loops, while only introducing acceptable deployment overhead at the switch.
Tao Zhang 0019, Jiawei Huang 0001, Shaojun Zou, Sen Liu 0002, Jinbin Hu 0001, Jingling Liu, Chang Ruan, Jianxin Wang 0001, Geyong Min
APNet6
2019 AG: Adaptive Switching Granularity for Load Balancing with Asymmetric Topology in Data Center Network
abstract
Modern data center topologies often take the form of a multi-rooted tree with rich parallel paths to provide high bandwidth. However, various path diversities caused by traffic dynamics, link failures and heterogeneous switching equipments widely exist in production datacenter network. Therefore, the multi-path load balancer in data center should be robust to these diversities. Although prior fine-grained schemes such as RPS and Presto make full use of available paths, they are prone to experience packet reordering problem under asymmetric topology. The coarse-grained solutions such as ECMP and LetFlow effectively avoid packet reordering, but easily lead to under-utilization of multiple paths. To cope with these inefficiencies, we propose a load balancing mechanism called AG, which adaptively adjusts switching granularity according to the asymmetric degree of multiple paths. AG increases switching granularity to alleviate packet reordering under large degrees of topology asymmetry, while reducing switching granularity to obtain high link utilization under small degrees of topology asymmetry. AG is deployed on the switches with negligible overhead, while making no modification on end-hosts. We evaluate AG through both Mininet testbed and large-scale NS2 simulations. The experimental results show that AG reduces the average and 99thflow completion time by up to 51% and 56% over the state-of-the-art load balancing schemes, respectively.
Jingling Liu, Jiawei Huang 0001, Weihe Li, Jianxin Wang 0001
ICNP1