VLDB 2026 Research / reviewers in the wild / expert
Wanchun Jiang
dblp:05/8330
· DBLP profile ↗
76ranked-venue papers
30as first author
50since 2021 · last 2026
0000-0001-5067-321XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 38 · 13 first-author · 27 since 2021Systems, architecture and hardware · 28 · 12 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCC: Re-architecting Delay-based Congestion Control in Datacenter Networks
Wanchun Jiang, Haoyang Li 0006, Danfeng Shan, Fengyuan Ren, Jiawei Huang 0001, Jianxin Wang 0001 |
NSDI | 1 |
| 2026 | Odin: Rethinking Congestion Control under All-to-All Traffic
Wanchun Jiang, Jianxi Ye, Huichen Dai |
SIGCOMM | 1 |
| 2026 | Cardinality is Not Enough: Super Host Detection via Segmented Cardinality EstimationabstractAccurately detecting super host that establishes connections to a large number of distinct peers is significant for mitigating web attacks and ensuring high quality of web service. Existing sketch-based approaches estimate the number of distinct connections called flow cardinality according to full IP addresses, while ignoring the fact that a malicious or victim super host often communicates with hosts within the same subnet, resulting in high false positive rates and low accuracy. Though hierarchical-structure based approaches could capture flow cardinality in subnet, they inherently suffer from high memory usage. To address these limitations, we propose SegSketch, a segmented cardinality estimation approach that employs a lightweight halved-segment hashing strategy to infer common prefix lengths of IP addresses, and estimates cardinality within subnet to enhance detection accuracy under constrained memory size. Experiments driven by real-world traces demonstrate that, SegSketch improves F1-Score by up to 8.04× compared to state-of-the-art solutions, particularly under small memory budgets. Jiawei Huang 0001, Xianshi Su, Weihe Li, Qichen Su, Jin Ye 0003, Wanchun Jiang, Jianxin Wang 0001 |
WWW | 10 |
| 2026 | Enhance energy efficient ethernet with reinforcement learning based periodic strategy
Wanchun Jiang, Renfu Yao, Jiawei Huang 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Comprehensive Multistep Prefetching Strategy for Short Video Streaming
Wanchun Jiang, Yongxin Shan, Jin-Tian Hu, Qichen Su, Jia-Wei Huang |
J. Comput. Sci. Technol. | 1 |
| 2026 | FAR: Fast and Accurate Rate Control for Lossless Datacenter NetworksabstractIn recent years, end-to-end congestion control algorithms or flow pausing mechanisms are proposed to achieve high throughput and low latency in datacenter networks. However, prior end-to-end congestion control works without complex signals fail to achieve fast convergence to a stable equilibrium state and effectively handle the transient congestion, while existing flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient states and incomplete queue elimination in equilibrium states. To address these issues, we present FAR, a rate control protocol that combines the advantages of flow pausing and congestion control. At its heart, FAR couples the bandwidth-estimation-based congestion control and the end-to-end flow pausing mechanisms. After flow pausing, FAR quickly explore the available bandwidth with a binary-search probe to achieve high throughput and low latency. Meanwhile, FAR employs a probe staggering mechanism to address the queue oscillation issue in high-concurrency scenarios. We implement the prototype of FAR using DPDK. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67% compared with the state-of-the-art designs. Jingling Liu, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0001, Ping Zhong 0002, Jiawei Huang 0001 |
IEEE Trans. Netw. | 9 |
| 2026 | Fast and Accurate Software Traffic Shaping With Inter-Flow Batching
Danfeng Shan, Shihao Hu, Hao Li 0011, Yazhe Tang, Peng Zhang 0011, Wanchun Jiang, Fengyuan Ren |
IEEE Trans. Netw. | 7 |
| 2026 | Toward QoE-Fairness for Video Streaming Over Heterogeneous Networks: An Innovative Bandwidth Allocation MechanismabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streams, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streams but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairnessawarebandwidthallocationmechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video stream to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. In addition, we propose a Deep Neural Network (DNN)-based multi-step mapping model aimed at balancing the performance and overhead of Fabam, thereby enhancing its deployment potential in practical applications. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 44.01% in QoE fairness and an improvement of 36.39% in QoE efficiency. Meanwhile, Fabam-DNN maintains satisfactory QoE fairness while supporting multiple users at a low cost. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | Introspective Congestion Control for Consistent High PerformanceabstractThe congestion control (CC) algorithm is expected to achieve consistent high performance under different network environments. Traditionally, classic CCs are designed with the methodology of inferring path conditions to guide the rate adjustment. However, this methodology suffers from wrong path condition inferences in certain cases, which mislead the rate adjustment and lead to performance degradation. To avoid wrong path condition inferences, we develop the projection-based introspective method and design the introspective congestion control (ICC) algorithm in this paper. Specifically, the rate adjustment rules are designed to possess a specialized profile such that the projection of the profile can be distinguished under unchanged path conditions. In this way, the projection, which can be distinguished from the time series of delay signals in the frequency domain, facilitates ICC to extract more information for path condition inferences. Consequently, with the introspection on the projection, ICC can avoid being misled by wrong path condition inferences and thus achieve consistent high performance under different conditions. The advantages of ICC are confirmed through extensive experiments conducted on various locally emulated scenarios, global testbeds over the Internet, and the Alipay platform. Wanchun Jiang, Haoyang Li 0006, Jia Wu 0011, Fengyuan Ren, Jianxin Wang 0001 |
EuroSys | 1 |
| 2025 | Occamy: A Preemptive Buffer Management for On-chip Shared-memory SwitchesabstractToday's high-speed switches employ an on-chip shared packet buffer. The buffer is becoming increasingly insufficient as it cannot scale with the growing switching capacity. Nonetheless, the buffer needs to face highly intense bursts and meet stringent performance requirements for datacenter applications. This imposes rigorous demand on the Buffer Management (BM) scheme, which dynamically allocates the buffer across queues. However, the de facto BM scheme, designed over two decades ago, is ill-suited to meet the requirements of today's network. Danfeng Shan, Yunguang Li, Jinchao Ma, Xinyu Wen, Hao Li 0011, Wanchun Jiang, Nan Li 0047, Fengyuan Ren |
EuroSys | 8 |
| 2025 | Scalable Bolt: Taming the Burst Queue via Scalable Token ManagementabstractEfficient congestion control (CC) is critical to maintaining high performance in modern datacenter networks. Recently, Bolt, proposed in NSDI 23, builds a sub-RTT control loop that enables ultra-low-latency reactions to congestion. Combined with an effective ramp-up mechanism driven by the proactive tokens, Bolt achieves superior performance compared to other existing CC schemes, especially in the face of workloads dominated by short flows. However, our study reveals that Bolt suffers from a severe burst bottleneck queue in heterogeneous topologies. This issue is particularly concerning because such heterogeneity is common in widely deployed Fat-tree and Leaf-spine topologies, and is expected to increase as data centers continue to scale out. To solve this problem, Scalable Bolt (S-Bolt) is proposed. By scaling the generation and consumption of tokens, S-Bolt eliminates the burst bottleneck queue while retaining the advantages of fast responding to congestion and spare bandwidth. Moreover, S-Bolt is friendly to the implementation on switches. Evaluating results show that S-Bolt works well under the current workloads and heterogeneous topologies. Specifically, S-Bolt reduces the burst queue by 70.6% and cuts down the tail flow completion time by 12.5%, compared with Bolt. Haoyang Li 0006, Peile Chen, Ang Jiang, Wanchun Jiang |
ICPADS | 5 |
| 2025 | DACC: Data Augmentation for Learning-based Congestion Control
Jiawei Huang 0001, Yijun Li 0002, Shengwen Zhou, Hui Li 0120, Weihe Li, Jingling Liu, Wanchun Jiang |
INFOCOM | 10 |
| 2025 | Cut the Response Time of Key-Value Stores by the SDN-Based SchedulerabstractAs the foundational components of large-scale applications, distributed key-value stores must respond to user requests quickly. However, a user request typically comprises multiple key-value access operations, which are processed in parallel across different servers, and the response time is determined by the slowest operation. To reduce the mean response time of requests, existing approaches schedule the sequence of key-value access operations across different servers so that all operations of a request complete at approximately the same time. Nevertheless, all of these approaches operate in a distributive manner, and their theoretical performance boundaries are unknown. To address these issues, we designed SDN-KVS (Key-Value Scheduler based on Software Defined Network), which migrates the waiting queue of key-value access operations from overloaded servers to the SDN controller. In this way, SDN-KVS centrally schedules the requests from different clients to overloaded servers without extra latency overhead. The scheduling result is proven to be$(1+2 \eta)$-approximation, i.e., the mean response time of requests is smaller than ($1+2 \eta$) times of the optimal value, where$\eta$is the parameter to make a trade-off between mean and tail response time. Simulation results confirm the excellent performance of the SDN-KVS algorithm. Specifically, SDN-KVS outperforms existing algorithms up to 37.6% and 79.8% in terms of mean and tail response time, respectively. Wanchun Jiang, Haoyang Li 0006, Chengke Wen, Rongfei Zeng, Jiawei Huang 0001, Jianxin Wang 0001 |
IWQoS | 1 |
| 2025 | SwitchTop-k: Scaling Top-k Compression on Programmable SwitchesabstractDistributed deep learning has been widely deployed in data centers to provide various services such as image classification and speech recognition. To reduce the training time, Top-k compression has become one of the most popular solutions used to shrink the data volume of gradients. Nevertheless, we observe that existing Top-k compression solutions are inefficient when used for large-scale distributed training due to gradient build-up, missing of Top-k gradients, and high compression overhead at the end hosts. To address these problems, we propose SwitchTop-k, which improves the accuracy of selecting Top-k values while ensuring a high compression rate and zero compression overhead. Specifically, SwitchTop-k offloads the Top-k compression from the end hosts to the programmable switches, thus alleviating the gradient build-up and compression overhead. Meanwhile, we propose a sketch-based solution to achieve high accuracy in selecting global Top-k gradients. We also co-design switch logic and end host logic to improve communication efficiency of uncompressed traffic. Finally, we implement SwitchTop-k on Intel Tofino switches and integrate it with Pytorch. The test results show that SwitchTop-k reduces iteration time by up to 91% compared with existing compression algorithms. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
KDD (2) | 5 |
| 2025 | Bandwidth Prediction Within Playback-Time of Buffer Duration Based on Frame-Level Stable Segment for Live Streaming
Xianshi Su, Wanchun Jiang, Danfeng Shan |
WASA (2) | 3 |
| 2025 | Teaching to Fish Rather Than Giving a Fish: The Concentrator Method of Teaching Classic Congestion Control With Learning-Based moduleabstractNowadays, Congestion Control (CC) algorithms are expected to satisfy the diverse demands of applications running over diverse networks. To achieve this goal, the combinations, which are expected to inherit both the advantages of classic CC in terms of convergence, overhead, and explainability, and the advantages of learning-based CC on adapting to diverse networks and demands, become a hot topic. In this paper, we reveal the existing combination works are eithergiving a fishorteaching to fish. Based on the insight of their essential issues, we develop the Concentrator method ofteaching to fish. According to this method, we propose Seagull as a step further. Specifically, Seagull captures the network characteristics and application demands in a coarse-grained manner via an online learning module. Moreover, the online learning module guides the customization of the rate adjustment rules of the classic CC module for fine-grained system evolution. Replacing the assumption on networks by the captured characteristics, the classic CC module of Seagull can fulfill the specified application demands. Real-world experimental results show Seagull respectively outperforms Orca, PCC-Vivace, and CUBIC by$49.3\%,\ 30.4\%$, and 24.9% in terms of throughput over the Internet, and improves the video quality of experience (QoE) by$12.9\sim 33.5\%$compared to CUBIC over cellular links. Haoyang Li 0006, Wanchun Jiang, Jie Wang 0067, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine LearningabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 8 |
| 2025 | Automatic Dual Threshold Tuning for Switch Buffer Sharing in Datacenter NetworkingabstractFor the widely deployed on-chip shared buffer, efficient buffer management is the key to absorbing bursts and avoiding packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. What’s more, we introduce D2T${}^{*}$which combines D2T with advanced DRL techniques to move toward mastering buffer management for further improving performance across various scenarios. We implement D2T at a P4-programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively. Jingling Liu, Hui Li 0120, Jiawei Huang 0001, Ping Zhong 0002, Boyan Huang, Pingping Dong, Wensheng Tang, Wanchun Jiang, Jianxin Wang 0001, Yong Cui 0001 |
IEEE Trans. Netw. | 9 |
| 2024 | Cooperative Simulation of RDMA-based Network and StorageabstractNowadays, the NVMe SSDs and the NVMe-oF (NVMe over Fabric) protocols are popular, and RDMA (Remote Direct Memory Access) has become one of the de-facto network fabric. Meanwhile, as the network bandwidth rapidly expands, potential bottlenecks can occur in both the RDMA-based network and the SSDs of the host. Accordingly, traffic management has become a hot research topic. However, the RDMA-based network and SSDs cannot be simulated at the same time, although simulators play an important role in research. To address this issue, we developed RSNet, a simulator for both the RDMA-based network and storage. Specifically, RSNet implements the RDMA and NVMe-oF protocols and integrates the open-source SSD simulator MQSim into the NS3 platform. The results validate that RSNet can simultaneously treat both network and SSD, as well as realistically simulate the collaboration between network and SSD. Moreover, RSNet is also capable of replicating certain issues within RDMA-based networks and storage. In total, we believe that RSNet would facilitate further research on traffic management in RDMA-based networks and storage. Wanchun Jiang, Xingping Zhang |
CSCWD | 1 |
| 2024 | D2T: Dynamic Dual Threshold Policy of Shared-Memory in Data Center SwitchesabstractNowadays the data center switches employ the on-chip shared buffer to absorb bursts and avoid packet loss during transient congestion. However, as the buffer-per-port-per-Gbps in production data centers decreases, it becomes more challenging to provide efficient buffer management to meet the requirements of heterogeneous traffic. We observe that typical shared buffer management policies have two steps: first, they identify short flows arriving at ports and then allocate more buffer room for these ports. Unfortunately, the lack of isolation between long and short flows leads to increased queue buildup and even packet loss of short flows. To address this limitation, we propose D2T, which uses different queue length thresholds for long and short flows. Specifically, we first design a compact data structure to distinguish between long and short flows. Then when two kinds of flows coexist at the same port, the threshold of long flows will decrease to absorb the bursty short flows. We implement D2T at a P4- programmable switch and large-scale simulations. The results demonstrate that D2T reduces both average and tail flow completion times (FCT) of short flows by up to 29% and 62% compared with the state-of-the-art policies, respectively. Jiawei Huang 0001, Hui Li 0120, Jingling Liu, Wenlu Zhang, Yijun Li 0002, Sitan Li, Shengwen Zhou, Ping Zhong 0002, Jianxin Wang 0001, Wanchun Jiang, Yong Cui 0001 |
ICDCS | 13 |
| 2024 | Achieving High Efficiency for Datacenter Multicast using Skewed Bloom FilterabstractMulticast serves as an important approach for one-to-many communication in data center networks. To reduce overhead and improve scalability, bloom filters are employed in current multicast approaches to store forwarding ports of switches. However, the well-known false positive issue of bloom filter incurs wrong forwarding behaviors and redundant traffic in multicast tree, degrading transmission efficiency and increasing the risk of data leakage. Inspired by the fact that, given the same false positive ratio, the switch in the upper layers of multicast tree generates more redundant traffic, we propose RSBF, a fine-grained and resource-aware multicast approach using skewed bloom filters. Specifically, RSBF maintains multiple bloom filters corresponding to different layers of multicast tree, and allocates more ample space to the bloom filter of the upper layer switches, thereby reducing the overall redundant traffic. The test results of large-scale simulation demonstrate that RSBF reduces both redundant traffic and header overhead by up to 64% and 49% compared with the state-of-the-art approaches, respectively. Jiawei Huang 0001, Hui Li 0120, Qile Wang, Sitan Li, Zhidong He, Wanchun Jiang |
ICPP | 9 |
| 2024 | Coupling Congestion Control and Flow Pausing in Data Center NetworkabstractTo achieve high throughput and low latency for data center applications, there are two broad lines of work: end-to-end congestion control algorithms and flow pausing mechanisms. It is challenging for end-to-end congestion control algorithms without complex signals to achieve fast convergence to a stable equilibrium state while effectively handling the transient congestion. Additionally, flow pausing mechanisms are decoupled from congestion control, which leads to long convergence time after transient state and incomplete queue elimination in equilibrium state. We propose a transport protocol that combines the advantages of flow pausing and congestion control, called FAR. The key idea is coupling the bandwidth-estimation based congestion control and the end-to-end flow pausing mechanisms. FAR quickly explores the available bandwidth with binary-search based packet train probe to achieve high throughput and low latency. Extensive evaluation results demonstrate that our protocol achieves accurate bandwidth estimation and reduces the tail flow completion time (FCT) by up to 67 <?TeX $\%$?> Math 1 compared with the state-of-the-art designs. Jiawei Huang 0001, Shengwen Zhou, Yijun Li 0002, Sitan Li, Wanchun Jiang, Jianxin Wang 0007, Ping Zhong 0002 |
ICPP | 9 |
| 2024 | Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningabstractDistributed deep learning has been widely employed to train deep neural network over large-scale dataset. However, the commonly used parameter server architecture suffers from long synchronization time in data-parallel training. Although the existing solutions are proposed to reduce synchronization overhead by breaking the synchronization barriers or limiting the staleness bound, they inevitably experience low convergence efficiency and long synchronization waiting. To address these problems, we propose Gsyn to reduce both synchronization overhead and staleness. Specifically, Gsyn divides workers into multiple groups. The workers in the same group coordinate with each other using the bulk synchronous parallel scheme to achieve high convergence efficiency, and each group communicates with parameter server asynchronously to reduce the synchronization waiting time, consequently increasing the convergence efficiency. Furthermore, we theoretically analyze the optimal number of groups to achieve a good tradeoff between staleness and synchronization waiting. The evaluation test in the realistic cluster with multiple training tasks demonstrates that Gsyn is beneficial and accelerates distributed training by up to 27% over the state-of-the-art solutions. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001 |
INFOCOM | 6 |
| 2024 | RLPS: Reinforcement Learning based Periodic Strategy for 40/100Gbps Energy Efficient EthernetabstractThe strategy for 40/100Gbps Energy Efficient Ethernet (EEE) determines when to enter and leave the power-saving modes. Accordingly, it directly decides both the energy savings and the incurred latency of frames in the EEE. The performance of the EEE strategy is greatly influenced by the network traffic, and thus existing EEE strategies need either proper parameter configuration under certain traffic loads or parameter adaptation mechanisms based on the traffic prediction under the assumptions of certain distribution. Consequently, these EEE strategies hardly keep consistent high performance under variable traffic in reality. To address this issue, we bring the reinforcement learning method into the design of the EEE strategy and propose the reinforcement learning based periodic strategy (RLPS) in this paper. Specifically, RLPS transmits existing frames at first and then stays in the selected power-saving mode for the rest time in each cycle. Moreover, RLPS learns the time length of each cycle online to reflect the impacts of traffic, instead of directly outputting power-saving mode transition decisions. In this way, the power consumption in each cycle is optimal with the help of learned information, and the overhead of online learning is reduced. Simulations driven by both synthetic traffic and real traces confirm that RLPS outperforms existing strategies, i.e., can achieve consistent high performance regardless of the traffic loads and distributions. Wanchun Jiang, Zhuang Tian, Renfu Yao, Xunyong Tan, Jiawei Huang 0001, Jianxin Wang 0001 |
ISPA | 1 |
| 2024 | Modeling and Analyzing the Shared Receive Queue of RDMAabstractNowadays, the RDMA (Remote Direct Memory Access) technology has been broadly employed in data centers. The Shared Receive Queue (SRQ) is an embedded mechanism in RDMA protocol, which reduces the memory cost of queue pairs sharing the same receiver. However, the configurations of SRQ are often heuristic and empirical nowadays. Consequently, the Receiver Not Ready (RNR) signal would be easily triggered, leading to utilization loss in the face of dynamic traffic. In other words, configuring SRQ reasonably is the key to the performance of RDMA and remains a challenge due to the variable traffic and environment. To address this issue, we propose a theoretical model for SRQ to guide its configuration. Simulations demonstrate that the system utilization is significantly improved and the triggering of RNR signals is reduced with the proper SRQ configuration guided by the theoretical model. Zhuang Tian, Wanchun Jiang |
ISPA | 3 |
| 2024 | Analysis and Improvement of PowerTCPabstractNowadays, Congestion Control (CC) algorithms based on In-Network Telemetry (INT) are popular in datacenter networks, because INT is supported by many commercial devices and provides comprehensive information about network congestion. In this work, we analyze the recent INT-based CC algorithm PowerTCP and reveal its fairness and large delay issues under the conditions of a large number of flows. Inspired by the analytical results, we propose Trident, which enhances PowerTCP by accelerating the speed of converging to fairness and maintaining a low queuing delay with a small equilibrium point. Simulations based on the open source codes confirm that Trident outperforms existing INT-based CC algorithms HPCC, PowerTCP and Poseidon by 14.3%, 14.8%, and 63.5% in terms of FCT of short flows. Moreover, Trident outperforms existing CC algorithms such as Timely, DCQCN and DCTCP, benefiting from the advantages inherited from PowerTCP. Wanchun Jiang, Haoyang Li 0006, Jiawei Huang 0001, Jianxin Wang 0001 |
IWQoS | 1 |
| 2024 | Achieving QoE Fairness in Video Streaming over Heterogeneous Congestion Control ProtocolsabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streaming, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streaming but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairness aware bandwidth allocation mechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video streaming to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 24.48% in QoE fairness and an improvement of 16.63% in QoE efficiency. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IWQoS | 6 |
| 2024 | A learning-based approach for video streaming over fluctuating networks with limited playback buffers
Weihe Li, Jiawei Huang 0001, Qichen Su, Wanchun Jiang, Jianxin Wang 0001 |
Comput. Commun. | 4 |
| 2024 | Improve video QoE by practical bandwidth allocation
Wanchun Jiang, Pan Ning, Zhicheng Ren, Jintian Hu, Jianxin Wang 0001 |
Multim. Syst. | 1 |
| 2024 | Learning Audio and Video Bitrate Selection Strategies via Explicit RequirementsabstractMobile video streaming dominates today's network traffic, and adaptive bitrate (ABR) algorithms have been routinely adopted for transmitting media content across dynamic mobile networks. State-of-the-art ABR algorithms mainly alter video bitrate without considering audio bitrate as they consider the impact on the video negligible due to their small size. However, to bring users an immersive experience, recent content providers have applied high-quality audio with large sizes, like stereophonic sound. Therefore, improper audio bitrate selection will adversely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, and vice versa) and frequent playback interruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bitrate selections. By learning from explicit goals, SPA can match the actual requirements and attain good performance. By conducting trace-driven and testbed-based experiments, we observe SPA's considerable superiority compared to existing approaches, including reducing the undesirable combinations by up to 34.17× and achieving zero stall time across 88.57% of traces. We also invite 35 volunteers to join a subjective test, and the result shows that 33/35 people consider SPA provides them with a satisfactory viewing experience. Weihe Li, Jiawei Huang 0001, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | VASE: Enhancing Adaptive Bitrate Selection for VBR-Encoded Audio and Video Content With Deep Reinforcement LearningabstractAdaptive BitRate (ABR) algorithms have become increasingly prevalent in modern streaming platforms, offering users significant improvements in the Quality of Experience (QoE). With streaming providers like YouTube and Netflix shifting to high-fidelity audio formats such as stereophonic sound and Dolby Atoms, ensuring proper audio and video adaptation has become a critical aspect of modern streaming platforms. Additionally, Variable Bitrate (VBR) encoding has gained great popularity in encoding audio and video content, given its higher quality-to-bits ratio. However, the considerable variability in network bandwidth, in combination with VBR features such as significantly fluctuating audio/video chunk sizes and diverse content complexity, makes existing ABR schemes formidable to make optimal bitrate selection due to their overlook of audio adaptation or oblivious to VBR features. In this paper, we introduce a new ABR approach forVBR-basedAudio-aware videoStrEaming named VASE, which harnesses deep reinforcement learning (DRL) and exploits parallel computing with multiple agents to swiftly and adeptly manage fluctuations in video/audio chunk sizes, network bandwidth, and varying content complexity, all while operating without any assumptions. Besides, two variants are proposed to mitigate the download energy cost and handle audio and video content in finer granularity. Extensive trace-driven, testbed, and subjective evaluations show that our scheme surpasses existing advanced adaptation schemes regarding the overall QoE, effectively demonstrating its superiority. Weihe Li, Jiawei Huang 0001, Qichen Su, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Improvement of Copa: Behaviors and Friendliness of Delay-Based Congestion Control AlgorithmabstractDelay-based congestion control has drawn a lot of attention in both academics and industry recently. Specifically, the Copa algorithm proposed in NSDI can achieve consistent high performance under various network environments and has already been deployed on Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to achieve its expected behaviors, i.e., clear the bottleneck buffer occupancy periodically. Moreover, we also reveal that the pathological competitive mode of Copa fails to guarantee friendliness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy and become friendly to Cubic in the competitive mode. As a result, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by real-world experiments and simulations. Specifically, Copa+ has the highest average throughput over different Internet links among different cloud nodes, compared to Cubic, BBR, PCC Vivace, Remy, and Indigo. Meanwhile, Copa+ has an 8.1% increase in throughput and similar low queuing delay compared to Copa. Moreover, Copa+ achieves 14.6% lower queuing delay and 2.4% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Jia Wu 0002, Zheyuan Liu 0008, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Straggler-Aware Gradient Aggregation for Large-Scale Distributed Deep Learning SystemabstractDeep Neural Network (DNN) is a critical component of a wide range of applications. However, with the rapid growth of the training dataset and model size, communication becomes the bottleneck, resulting in low utilization of computing resources. To accelerate communication, recent works propose to aggregate gradients from multiple workers in the programmable switch to reduce the volume of exchanged data. Unfortunately, since using synchronization transmission to aggregate data, current in-network aggregation designs suffer from the straggler problem, which often occurs in shared clusters due to resource contention. To address this issue, we propose a straggler-aware aggregation transport protocol (SA-ATP), which enables the leading worker to leverage the spare computing and storage resources to help the straggling worker. We implement SA-ATP atop clusters using P4-programmable switches. The evaluation results show that SA-ATP reduces the iteration time by up to 57% and accelerates training by up to$1.8\times $in real-world benchmark models. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2024 | Enforcing Fairness in the Traffic Policer Among Heterogeneous Congestion Control AlgorithmsabstractTraffic policing is widely used by ISPs to limit their customers’ traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms: flows using new CC algorithms can easily occupy the majority of bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without a buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient. When being policed, they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for competing TCP flows. Simply adding a buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards a token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. To avoid bandwidth waste when flows come and go, FairPolicer puts all available tokens in a global bucket and maintains the amount of residual bucket space rather than the number of available tokens. To scale to massive concurrent flows, FairPolicer uses a Count-Min Sketch structure to maintain per-flow data with a small memory footprint. Testbed experiments show that FairPolicer can allocate bandwidth in a max-min fair manner and achieve much lower latency than other kinds of rate limiters. Danfeng Shan, Linbing Jiang, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Yazhe Tang, Fengyuan Ren |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Reinforcement-Learning Based Preload Strategy for Short Video
Zhicheng Ren, Yongxin Shan, Wanchun Jiang, Yijing Shan, Danfeng Shan, Jianxin Wang 0001 |
ICIC (5) | 3 |
| 2023 | PA-ATP: Progress-Aware Transmission Protocol for In-Network AggregationabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
ICNP | 8 |
| 2023 | weBurst can be Harmless: Achieving Line-rate Software Traffic Shaping by Inter-flow BatchingabstractTraffic shaping is a common function at end hosts. Compared with hardware ones, software shapers are more flexible to be developed and deployed, and thus are very attractive. Nevertheless, software approaches are still unsatisfactory as they struggle to saturate 40Gbps and higher speed.While much effort has been made to reduce the intrinsic overhead of software traffic shaping, we find that it is the extrinsic overhead, such as PCIe communications and interrupts, that hinders shaping from achieving 40Gbps - 100Gbps speed. Batching is an effective way to amortize these overheads. However, blindly batching can degrade the network performance, as it introduces bursts into the network. Diving into the dilemma, we find that intra-flow burst is to blame for harming the network performance, while inter-flow burst, consisting of packets from different flows, can be naturally demultiplexed in the network.Based on the insight, we present FlowBundler, which can achieve efficient traffic shaping by inter-flow batching. Testbed experiments show that FlowBundler can achieve an accurate shaping of 98Gbps with a single CPU core, which is 2.6× better than state-of-the-art approaches. Large-scale simulations show that FlowBundler can batch packet transmissions without harming the network performance. Danfeng Shan, Shihao Hu, Wanchun Jiang, Hao Li 0011, Peng Zhang 0011, Yazhe Tang, Huanzhao Wang, Fengyuan Ren |
INFOCOM | 4 |
| 2023 | Practical periodic strategy for 40/100 Gbps Energy Efficient Ethernet
Wanchun Jiang, Renfu Yao, Kaiqin Liao, Yulong Yan, Jiawei Huang 0001, Weiping Wang 0003, Jianxin Wang 0001 |
Comput. Networks | 1 |
| 2023 | AMS: Adaptive Multiget Scheduling Algorithm for Distributed Key-Value StoresabstractDistributed key-value stores provide the Multiget API, where many key-value operations are batched together, to meet the parallel requirement of applications. Correspondingly, reducing the latency of Multigets is crucial for the responsiveness of the distributed key-value stores. The latency of a Multiget depends on both which replica server its key-value operations are scheduled to, i.e., the replica selection for each key-value operation, and when these key-value operations are served, i.e., the scheduling of the service sequence at different replica servers. Existing solutions solely focus on either one of them and accordingly lead to suboptimal latency for Multigets. To address these issues, this paper proposes an Adaptive Multiget Scheduling (AMS) algorithm in this paper, and specifically, our AMS re-architectures the framework to remove the conflict between replica selection and service sequence scheduling in the existing solution Rein. Based on the new framework, a sophisticated replica selection method is designed. Furthermore, AMS guides both replica selection and service sequence scheduling by the piggybacked information of replica servers, being adaptive to the heterogeneous time-varying server performance. Consequently, AMS can respectively reduce the median,$95^{th}$, and$99^{th}$percentile latencies of Multigets by a factor of 4, 3.1, and 1.86 compared to the default FIFO algorithm and significantly outperforms Rein. Wanchun Jiang, Yujia Qiu, Fa Ji, Yongjia Zhang, Xiangqian Zhou, Jianxin Wang 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | RAV: Learning-Based Adaptive Streaming to Coordinate the Audio and Video Bitrate SelectionsabstractMost commercial players adopt adaptive bitrate (ABR) algorithms to dynamically decide each chunk's bitrate based on the perceived network bandwidth and buffer occupancy. However, current ABR algorithms are agnostic of audio bitrate selection since they deem it has negligible influence on video bitrate selection due to small size of audio chunks. Nevertheless, with the development of audio technologies, the bitrate of audio content increases dramatically in recent years. Thus, inappropriate audio selection can significantly affect video selection and deteriorate the viewing experience. To tackle these inefficiencies, we propose a deepReinforcement learning-based ABR algorithm that takesAudio andVideo quality into account (RAV) to circumvent a series of suboptimal performances, like low playback quality, frequent playback interruptions, poor playback smoothness, and undesirable combinations of video and audio chunks. Furthermore, RAV trains a neural network model that automatically outputs the bitrates for future audio and video chunks without relying on any presumptions about the environment, achieving good robustness to a broad spectrum of conditions. By conducting trace-driven and real-world experiments, we demonstrate that RAV significantly ameliorates the average overall viewing quality by 37.96%-118.20% over the state-of-the-art ABR algorithms. In addition, we also conduct subjective experiments by inviting 32 volunteers, and 27/32 users strongly agree that RAV provides them a better viewing experience than existing ABR solutions. Weihe Li, Jiawei Huang 0001, Wenjun Lyu, Baoshen Guo, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Consistent Low Latency Scheduler for Distributed Key-Value StoresabstractNowadays, the distributed key-value stores have become the basic building block for large-scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for a single end-user request. Accordingly, the completion time of an end-user request is determined by the last completed key-value access operation. Scheduling the order of serving key-value access operations can effectively reduce the completion times of end requests, thereby improving the user experience. However, existing scheduling algorithms hardly achieve consistent low latency due to the following challenges: the large overhead of cooperating clients and servers, the time-varying load and performance of servers, the traffic distribution can be either heavy-tailed or light-tailed and both the mean and the tail completion time are expected to be low. In this paper, we formalize the problem of scheduling key-value access operations and show it is NP-hard. Furthermore, we heuristically design the distributed adaptive scheduler (DAS), which distributively combines the largest remaining processing time last and the shortest remaining process time first algorithms. Theoretical analysis shows that DAS is adaptive to the time-varying traffic and server performance and can achieve consistent low mean and tail latency regardless of traffic distributions. Extensive simulations show that DAS reduces the mean request completion time by$17 \! \sim \! 50\%$with heavy-tailed traffic and$2 \! \sim 26 \! \%$with light-tailed traffic, while keeping the smallest tail completion time, compared to the default first come first served algorithm. Moreover, DAS outperforms the existing Rein-SBF algorithm under various scenarios. Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jiawei Huang 0001, Jianxin Wang 0001, Tong Zhang 0018 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Accelerated Information Dissemination for Replica Selection in Distributed Key-Value Store SystemsabstractIn distributed key-value stores, multiple replica servers are always available for each key-value access operation when the eventual consistency model is employed. Accordingly, the completion times of the key-value access operations generated by an end-user request at different servers may be of great difference, especially when the replica servers are heterogeneous and have time-varying performance. Accordingly, the replica selection algorithm is crucial to cut the response time of end-user requests. The main challenge of making replica selection for each light-weighted key-value access operation is to timely know the status of replica serves. Recently, the adaptive replica selection algorithm C3 suggests guiding the replica selection with the piggybacked information of replica server in the returned “value”. Although C3 has good performance, the poor timeliness of feedback information makes a large performance gap between C3 and the ideal replica selection algorithm. To narrow this gap, the Accelerated Information Dissemination (AID) mechanism is proposed in this paper. Specifically, AID removes the bottleneck of information dissemination at the “slow” servers by letting both “client” and “replica server” store the records about the status of replica servers and both “key” and “value” piggyback multiple records. AID is implemented in Cassandra and evaluated by experiments and large scale simulations. The results show AID can significantly improve the timeliness of feedback information, especially when the number of nodes is large. Accordingly, AID helps C3 to greatly reduce the latency. Wanchun Jiang, Yujia Qiu, Fa Ji, HaiMing Xie, Xiangqian Zhou, Jiawei Huang 0001, Jianxin Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | HSP: Hybrid Synchronous Parallelism for Fast Distributed Deep LearningabstractIn the parameter-server-based distributed deep learning system, the workers simultaneously communicate with the parameter server to refine model parameters, easily resulting in severe network contention. To solve this problem, Asynchronous Parallel (ASP) strategy enables each worker to update the parameter independently without synchronization. However, due to the inconsistency of parameters among workers, ASP experiences accuracy loss and slow convergence. In this paper, we propose Hybrid Synchronous Parallelism (HSP), which mitigates the communication contention without excessive degradation of convergence speed. Specifically, the parameter server sequentially pulls gradients from workers to eliminate network congestion and synchronizes all up-to-date parameters after each iteration. Meanwhile, HSP cautiously lets idle workers to compute with out-of-date weights to maximize the utilizations of computing resources. We provide theoretical analysis of convergence efficiency and implement HSP on popular deep learning (DL) framework. The test results show that HSP improves the convergence speedup of three classical deep learning models by up to 67%. Yijun Li 0002, Jiawei Huang 0001, Shengwen Zhou, Wanchun Jiang, Jianxin Wang 0001 |
ICPP | 5 |
| 2022 | Copa+: Analysis and Improvement of the Delay-based Congestion Control Algorithm CopaabstractCopa is a delay-based congestion control algorithm proposed in NSDI recently. It can achieve consistent high performance under various network environments and has already been deployed in Facebook. In this paper, we theoretically analyze Copa and reveal its large queuing delay and poor fairness issue under certain conditions. The root cause is that Copa fails to clear the bottleneck buffer occupancy periodically as expected. Accordingly, Copa may get a wrong base RTT estimation and enter its competitive mode by mistake, leading to large delay and unfairness. To address these issues, we propose Copa+, which enhances Copa with a parameter adaptation mechanism and an optimized competitive mode entrance criterion. Designed based on our theoretical analysis, Copa+ can adaptively clear the bottleneck buffer occupancy for correct estimation of base RTT. Consequently, Copa+ inherits the advantages of Copa but achieves lower queuing delay and better fairness under different environments, as confirmed by the real-world experiments and simulations. Specifically, Copa+ has the highest throughput similar to Copa but 11.9% lower queuing delay over different Internet links among different cloud nodes, and achieves 39.4% lower queuing delay and 8.9% higher throughput compared to Sprout over emulated cellular links. Wanchun Jiang, Haoyang Li 0006, Zheyuan Liu 0008, Jia Wu 0002, Jiawei Huang 0001, Danfeng Shan, Jianxin Wang 0001 |
INFOCOM | 1 |
| 2022 | Opportunistic Transmission for Video Streaming over Wild InternetabstractThe video streaming system employs adaptive bitrate (ABR) algorithms to optimize a user’s quality of experience. However, it is hard for ABR algorithms to choose the right bitrate consistently under highly dynamic bandwidth fluctuations in wild Internet. In this article, we propose a building block on the client side named Opportunistic Chunk Replacement Mechanism (OCRM) to help existing ABR algorithms make full use of the available bandwidth to improve the network utilization and viewing experience of users. Specifically, the servers take advantages of the spare bandwidth to opportunistically transmit high-quality chunks (called opportunistic chunks ) with low priority to the client, without incurring any extra delay. Then, the client player replaces the low-quality chunks with the opportunistic ones that have high quality. We compare OCRM with state-of-the-art ABR algorithms by using trace-driven experiments spanning a wide variety of quality of experience metrics and network conditions. The test results show that OCRM effectively achieves high network utilization and improves the user’s viewing experience by up to 35%. Jiawei Huang 0001, Qichen Su, Weihe Li, Zhuoran Liu 0003, Tao Zhang 0019, Sen Liu 0002, Ping Zhong 0002, Wanchun Jiang, Jianxin Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2021 | Cutting the Request Completion Time in Key-value Stores with Distributed Adaptive SchedulerabstractNowadays, the distributed key-value stores have become the basic building block for large scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for the data required by a single end-user request. Hence, the completion time of the end request is determined by the last completed key-value access operation. Accordingly, scheduling the order of key-value access operations of different end requests can effectively reduce their completion time, improving the user experience. However, existing algorithms are either hard to employ in distributed key-value stores due to the relatively large cooperation overhead for centralized information or unable to adapt to the time-varying load and server performance under different traffic patterns. In this paper, we first formalize the scheduling problem for small mean request completion time. As a step further, because of the NP-hardness of this problem, we heuristically design the distributed adaptive scheduler (DAS) for distributed key-value stores. DAS reduces the average request completion time by a distributed combination of the largest remaining processing time last and shortest remaining process time first algorithms. Moreover, DAS is adaptive to the time-varying server load and performance. Extensive simulations show that DAS reduces the mean request completion time by more than 15 ~ 50% compared to the default first come first served algorithm and outperforms the existing Rein-SBF algorithm under various scenarios. Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jianxin Wang 0001, Tong Zhang 0018 |
ICDCS | 1 |
| 2021 | Towards the Fairness of Traffic PolicerabstractTraffic policing is widely used by ISPs to limit their customers' traffic rates. It has long been believed that a well-tuned traffic policer offers a satisfactory performance for TCP. However, we find this belief breaks with the emergence of new congestion control (CC) algorithms like BBR: flows using these new CC algorithms can easily occupy the majority of the bandwidth, starving traditional TCP flows. We confirm this problem with experiments and reveal its root cause as follows. Without buffer in traffic policers, congestion only causes packet losses, while new CC algorithms are loss-resilient, i.e. they adjust the sending rate based on other network feedback like delay. Thus, when being policed they will not reduce the sending rate until an unacceptable loss ratio for TCP is reached, resulting in low throughput for TCP. Simply adding buffer to the traffic policer improves fairness but incurs high latency. To this end, we propose FairPolicer, which can achieve fair bandwidth allocation without sacrificing latency. FairPolicer regards token as a basic unit of bandwidth and fairly allocates tokens to active flows in a round-robin manner. Testbed experiments show that FairPolicer can significantly improve the fairness and achieve much lower latency than other kinds of rate-limiters. Danfeng Shan, Peng Zhang 0011, Wanchun Jiang, Hao Li 0011, Fengyuan Ren |
INFOCOM | 3 |
| 2021 | Practical Bandwidth Allocation for Video QoE Fairness
Wanchun Jiang, Pan Ning, Jintian Hu, Zhicheng Ren, Jianxin Wang 0001 |
WASA (1) | 1 |
| 2021 | Analysis and improvement of the latency-based congestion control algorithm DX
Wanchun Jiang, Haoyang Li 0006, Lijuan Peng, Jia Wu 0011, Chang Ruan, Jianxin Wang 0001 |
Future Gener. Comput. Syst. | 1 |
| 2021 | Optimizing the Response Time of Memcached Systems via Model and Quantitative AnalysisabstractMemcached is a widely used in-memory caching solution in large-scale searching scenarios. The most crucial metric of Memcached systems is the response time, which is affected by various factors such as workload, service rate, unbalanced load distribution, and cache miss ratio. This article aims to quantify the influence of each factor on the response time of Memcached systems. First, we establish a theoretical model for Memcached systems that captures their main features, including burst and concurrent key arrival, unbalanced load distribution, and cache miss process. By solving this model using queuing and stochastic theories, we obtain an estimate of the response time in Memcached systems. Intensive experiments based on real-world components demonstrate that the estimate always matches perfectly with the actual value. Furthermore, we obtain a comprehensive and quantitative understanding of all factors. The main insights are threefold. 1) There exists an optimum range of utilization at Memcached servers in which the response time is kept at a low level with a small penalty. 2) The influence of the cache miss ratio on the response time is logarithmic rather than linear. 3) The number of keys generated from an end-user request has the greatest impact in Memcached systems. Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018 |
IEEE Trans. Computers | 3 |
| 2020 | Information Dissemination for the Adaptive Replica Selection algorithm in Key-Value StoresabstractIn distributed key-value stores, multiple replica servers are always available for each key-value access operation when the eventually consistency model is employed. Accordingly, the replica selection algorithm is crucial to the tail latency of the key-value access operations generated by end-user requests, especially under the environment of heterogeneous replica servers. The main challenge of making replica selection decision is to know the status of replica serves with acceptable overhead for each lightweight key-value access operation. Recently, aware of the time-varying performance and load of replica servers, the adaptive replica selection algorithm C3 suggests piggybacking the information of replica server via the returned value” to guide the replica selection. Although the good performance of C3 has been verified by experiments and simulations, the poor timeliness of feedback information makes a large performance gap between C3 and the ideal replica selection algorithm. To narrow this gap, we propose the information dissemination mechanism for C3, which lets both “client” and “server” store the records about the status of replica servers, and both key” and value” piggyback multiple records. In this way, the information dissemination bottleneck at “slow” server would be removed with the help of multiple clients and servers. As confirmed by simulation results, the information dissemination mechanism can significantly improve the timeliness of feedback information with acceptable overhead and accordingly helps C3 to greatly reduce the tail latency (by about 35% under the default scenario). Wanchun Jiang, Fa Ji, HaiMing Xie, Xiangqian Zhou, Jianxin Wang 0001 |
ICC | 1 |
| 2020 | PS: Periodic Strategy for the 40-100Gbps Energy Efficient EthernetabstractThe 40-100Gbps Energy Efficient Ethernet (EEE) standardized by IEEE 802.3bj contains not only the DeepSleep mode, which is 10% of the normal power consumption, but also the FastWake mode, which has much less state transition time and 70% of the normal power consumption. Correspondingly, the strategy of determining the state transitions is crucial to both the power consumption and the incurred latency of frames. As EEE applications always place tail latency constraint on frame transmission, existing strategies for the 40-100Gbps EEE are rarely configured with proper parameters for good power saving under the given tail latency constraint, especially in face of the variable traffic load. To address this problem, we design the Periodic Strategy (PS) for the 40-100Gbps EEE. Specifically, PS works periodically and optimizes the power saving independently in each cycle. In this way, PS decouples the latency constraint and energy saving. 1) It limits the largest incurred latency by the length of the cycle to satisfy the given latency constraint. 2) It achieves good power saving by selecting the proper low power state based on the prediction of incoming frames and entering the selected state just once in each cycle. What’s more, the good performance of PS is insensitive to the traffic pattern. Extensive simulations confirm that PS achieves better energy saving than all the existing strategies under the given latency constraint, under several kinds of different traffic patterns. Wanchun Jiang, Kaiqin Liao, Yulong Yan, Jianxin Wang 0001 |
ICPP | 1 |
| 2020 | Polo: Receiver-Driven Congestion Control for Low Latency over Commodity Network FabricabstractRecently, numerous novel transport protocols are proposed for the low latency of applications deployed in data center networks, e.g., web search and retail recommendation system. The state-of-art receiver-driven protocols, e.g., Homa and NDP, show the superior performance for achieving the lowest possible latency. However, Homa assumes that the core layer in data center network has no congestion, which limits its application for the existing over-subscribed networks. NDP requires to modify the switch hardware since it trims packets to headers when the packets cause the switch buffer to overflow, resulting in high deployment cost. In this paper, we present Polo to realize low latency for flows over commodity network fabric relying on Explicit Congestion Notification (ECN) and priority queues. According to packets with ECN marking, the Polo receiver obtains the congestion information such that it dynamically adjusts the number of data packets in network for maintaining the small switch queue. The adjustment is carried out periodically. The time interval is determined by keeping the extra high priority packet always in flight nor by a fine-grained timer. Further, Polo designs the packet recovery mechanisms to retransmit the lost packets as soon as possible. Simulation experiment results show that Polo outperforms the state-of-art receiver-driven protocols in a wide range of scenarios including incast. Chang Ruan, Jianxin Wang 0001, Wanchun Jiang, Tao Zhang 0019 |
ICPP | 3 |
| 2020 | Re-architecting Congestion Management in Lossless Ethernet
Wenxue Cheng, Kun Qian 0017, Wanchun Jiang, Tong Zhang 0018, Fengyuan Ren |
NSDI | 3 |
| 2020 | PTCP: A priority-based transport control protocol for timeout mitigation in commodity data center
Chang Ruan, Jianxin Wang 0001, Wanchun Jiang, Geyong Min, Yi Pan 0001 |
Future Gener. Comput. Syst. | 3 |
| 2020 | Achieving high utilization of flowlet-based load balancing in data center networks
Shaojun Zou, Jiawei Huang 0001, Wanchun Jiang, Jianxin Wang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2019 | Modeling and Analysis of the Latency-Based Congestion Control Algorithm DX
Wanchun Jiang, Lijuan Peng, Chang Ruan, Jia Wu 0011, Jianxin Wang 0001 |
NPC | 1 |
| 2019 | TAP: Timeliness-aware predication-based replica selection algorithm for key-value storesabstractSummary In current large‐scale distributed key‐value stores, a single end‐user request may lead to key‐value access across tens or hundreds of servers. The tail latency of these key‐value accesses is crucial to user experience and greatly impacts the revenue. To cut the tail latency, it is crucial for clients to choose the best replica server as much as possible for the service of each key‐value access operation. Aware of the challenges on the time‐varying performance across servers and the herd behaviors, an adaptive replica selection scheme C3 has been proposed recently. In C3, feedback from individual servers is brought into replica ranking to reflect the time‐varying performance of servers, and the distributed rate control and backpressure mechanisms are invented. Despite C3's good performance, we reveal the timeliness issue of C3, which has large impacts on both the replica ranking and the rate control. To address this issue, we propose the TAP (timeliness‐aware predication‐based) replica selection algorithm, which predicts the queue size of replica servers under the poor timeliness condition, instead of utilizing the exponentially weighted moving average of the piggybacked queue sizes in history as in C3. Consequently, compared with C3, TAP can obtain more accurate queue‐size estimation to guide the replica selection at clients. Simulation results also confirm the advantage of TAP over C3 in terms of cutting the tail latency. Xianqian Zhou, Liyuan Fang, HaiMing Xie, Wanchun Jiang |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Understanding and improvement of the selection of replica servers in key-value stores
Wanchun Jiang, HaiMing Xie, Xiangqian Zhou, Liyuan Fang, Jianxin Wang 0001 |
Inf. Syst. | 1 |
| 2019 | Haste makes waste: The On-Off algorithm for replica selection in key-value stores
Wanchun Jiang, HaiMing Xie, Xiangqian Zhou, Liyuan Fang, Jianxin Wang 0001 |
J. Parallel Distributed Comput. | 1 |
| 2018 | A fine-grained rule partition algorithm in cloud data centers
Wei Jiang 0042, Wanchun Jiang, Weiping Wang 0003, Yi Pan 0001, Jianxin Wang 0001 |
J. Netw. Comput. Appl. | 2 |
| 2017 | Performance Analysis and Improvement of Replica Selection Algorithms for Key-Value StoresabstractIn current large-scale distributed key-value stores for cloud computing, the tail latency of the hundreds of key-value accesses generated by an end-user request determines the response time of this request. Replica selection algorithms, which select the best replica server for each key-value access as much as possible, is crucial to reduce the tail latency. This paper summarizes current replica selection algorithms and classifies them into three categories: information-agnostic, client-independence and feedback, according to their demanded information. Furthermore, simulation-based performance analysis of these algorithms is conducted. Based on the insights obtained from performance analysis, we design the L2 algorithm by assembling the basic ideas of the Least OSK algorithm and the Least RPT algorithm. The L2 algorithm has similar best performance with the recently proposed C3 algorithm, but is much simpler than C3. Wanchun Jiang, HaiMing Xie, Xiangqian Zhou, Liyuan Fang, Jianxin Wang 0001 |
CLOUD | 1 |
| 2017 | Tars: Timeliness-Aware Adaptive Replica Selection for Key-Value StoresabstractIn current large-scale distributed key-value stores, a single end-user request may lead to key-value access across tens or hundreds of servers. The tail latency of these key-value accesses is crucial to the user experience and greatly impacts the revenue. To cut the tail latency, it is crucial for clients to choose the best replica server as much as possible for the service of each key-value access. Aware of the challenges on the time-varying performance across servers and the herd behaviors, an adaptive replica selection scheme C3 is proposed recently. In C3, feedback from individual servers is brought into replica ranking to reflect the time-varying performance of servers, and the distributed rate control and backpressure mechanism is invented. Despite of C3's good performance, we reveal the timeliness issue of C3, which has large impacts on both the replica ranking and the rate control, and propose the Tars (timeliness-aware adaptive replica selection) scheme. Following the same framework as C3, Tars improves the replica ranking by taking the timeliness of the feedback information into consideration, as well as revises the rate control of C3. Simulation results confirm that Tars outperforms C3. Wanchun Jiang, Liyuan Fang, HaiMing Xie, Xiangqian Zhou, Jianxin Wang 0001 |
ICCCN | 1 |
| 2017 | Modeling and Analyzing Latency in the Memcached systemabstractMemcached is a widely used in-memory caching solution in large-scale searching scenarios. The most pivotal performance metric in Memcached is latency, which is affected by various factors including the workload pattern, the service rate, the unbalanced load distribution and the cache miss ratio. To quantitate the impact of each factor on latency, we establish a theoretical model for the Memcached system. Specially, we formulate the unbalanced load distribution among Memcached servers by a set of probabilities, capture the burst and concurrent key arrivals at Memcached servers in form of batching blocks, and add a cache miss processing stage. Based on this model, algebraic derivations are conducted to estimate latency in Memcached. The latency estimation is validated by intensive experiments. Moreover, we obtain a quantitative understanding of how much improvement of latency performance can be achieved by optimizing each factor and provide several useful recommendations to optimal latency in Memcached. Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018 |
ICDCS | 3 |
| 2017 | Congestion control in Converged Ethernet with heterogeneous and time-varying delaysabstractCongestion control is an indispensable mechanism in the new trend of enhanced Ethernet as a unified fabric for traditional LAN, SAN, and high-performance computing networks. A congestion management framework for Converged Ethernet (CE) networks has been standardized by IEEE 802.1 Qau work group, and QCN is recommended as the congestion control scheme in the standard draft. QCN is heuristically designed for 1/10Gbps Ethernet without considering the impact of delays. Recent work find that QCN will encounter stability issues with feedback delays, and these issues will be more serious as Ethernet extends to 40/100Gbps and the delays become heterogeneous and time-varying. This work aims to mitigate the negative impact of delays on congestion control scheme in CE. Specially, considering the delays are heterogeneous and time-varying, we build a model for Converged Ethernet with the standard congestion management framework. The model provides a new congestion detector to estimate the real congestion status under the impact of delays and regards the heterogeneous and time-varying feature as disturbances. Leveraging the new congestion detector and tolerating the disturbance through the sliding mode control method, we design the Delay-tolerant Sliding Mode (DSM) congestion control scheme. Extensive simulations show that DSM outperforms other congestion control schemes when the Ethernet ranges from 1Gbps to 100Gbps and the delays are heterogeneous and time-varying. Wenxue Cheng, Wanchun Jiang, Tong Zhang 0018, Bo Wang 0066, Kun Qian 0017, Fengyuan Ren |
IWQoS | 2 |
| 2017 | FSQCN: Fast and simple quantized congestion notification in data center ethernet
Chang Ruan, Jianxin Wang 0001, Wanchun Jiang, Jiawei Huang 0001, Geyong Min, Yi Pan 0001 |
J. Netw. Comput. Appl. | 3 |
| 2017 | Analyzing and Enhancing Dynamic Threshold Policy of Data Center SwitchesabstractToday's data center switches usually employ on-chip shared memory; buffer management policy in them is essential to ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) policy is widely used by switch vendors. Meanwhile, in data centers, distributed applications such as MapReduce often introduce micro-burst traffic into network and the packet dropping caused by micro-burst usually leads to serious performance degradation. When micro-burst traffic arrives at switches, DT is unable to fully utilize the buffer to absorb it. Therefore, in this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and quantitatively estimate the free buffer size when packets are dropped. The results show that the free buffer size can be very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose the Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT. Danfeng Shan, Wanchun Jiang, Fengyuan Ren |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | FSQCN: Fast and Simple Quantized Congestion Notification in Data Center EthernetabstractCurrently, Quantized Congestion Notification (QCN) has been accepted as the standard layer 2 congestion control protocol in Data Center Ethernet. Although the good performance of QCN has been validated in many experiments, we find that QCN has two drawbacks. First, the incomplete binary search in QCN fails to find the proper sending rate, leading to complicate supplement mechanisms. Second, in face of unknown network environment, the rate setting of QCN is inconsistent. Consequently, QCN spends much time on acquiring the spare bandwidth. To address theses problems, we propose the Fast and Simple QCN (FSQCN), following the same framework as QCN. FSQCN complements the binary search and removes other complicate supplement mechanisms in QCN. Thus, FSQCN is much simpler than QCN. Moreover, FSQCN resets the sending rate to the link rate explicitly when the switch detects spare bandwidth. Extensive simulations validate that FSQCN controls the queue length well like QCN and responds faster than QCN. Chang Ruan, Jianxin Wang 0001, Wanchun Jiang |
ICDCS | 3 |
| 2015 | Absorbing micro-burst traffic by enhancing dynamic threshold policy of data center switchesabstractIn data center networks, micro-burst is a common traffic pattern and the packet dropping caused by it usually leads to serious performance degradation. Meanwhile, most of the current commodity switches employ on-chip shared memory, and the buffer management policies of them ensure fair sharing of memory among all ports. Among various polices, Dynamic Threshold (DT) is widely used by switch vendors. However, because DT needs to reserve a fraction of switch buffer, there is free buffer space while packets from micro-burst traffic are dropped. In this paper, we theoretically deduce the sufficient conditions for packet dropping caused by micro-burst traffic, and estimate the corresponding free buffer size. The results show that the free buffer size is very large when the number of overloaded ports is small. What's worse, to ensure fair sharing of memory among output ports, packets from micro-burst traffic may be dropped even when the traffic size is much smaller than the buffer size. In light of these results, we propose Enhanced Dynamic Threshold (EDT) policy, which can alleviate packet dropping caused by micro-burst traffic through fully utilizing the switch buffer and temporarily relaxing the fairness constraint. The simulation results show that EDT can absorb more micro-burst traffic than DT. Danfeng Shan, Wanchun Jiang, Fengyuan Ren |
INFOCOM | 2 |
| 2015 | Sliding Mode Congestion Control for Data Center Ethernet NetworksabstractRecently, Ethernet is enhanced as the unified switch fabric of data centers, called data center Ethernet. One of the indispensable enhancements is end-to-end congestion management, and currently quantized congestion notification (QCN) has been ratified as the corresponding standard. However, our experiments show that QCN suffers from large oscillations of the queue length at the bottleneck link such that the buffer is emptied frequently and accordingly the link utilization degrades, with certain system parameters and network configurations. This phenomenon is corresponding to our theoretical analysis result that QCN fails to enter into the sliding mode motion (SMM) pattern with certain system parameters and network configurations. Knowing the drawbacks of QCN and realizing the advantage that congestion management system is insensitive to the changes of parameters and network configurations in the SMM pattern, we present sliding mode congestion control (SMCC), which can enter into the SMM pattern under any conditions. SMCC is simple, stable, fair, has short response time, and can be easily used to replace QCN because both of them follow the framework developed by the IEEE 802.1Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially when traffic pattern and network states are variable. Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Yongwei Wu 0001, Chuang Lin 0002 |
IEEE Trans. Computers | 1 |
| 2015 | Phase Plane Analysis of Quantized Congestion Notification for Data Center EthernetabstractCurrently, Ethernet is being enhanced to become the unified switch fabric in data centers. With the unified switch fabric, the cost on redundant devices is reduced, while the design and management of data center networks are simplified. Congestion management is one of the indispensable enhancements on Ethernet, and Quantized Congestion Notification (QCN) has just been ratified as the formal standard. Though QCN has been investigated for several years, there exist few in-depth theoretical analyses on QCN. The most possible reason is that QCN is heuristically designed and involves the property of variable structure. The classic linear analysis method is incapable of handling the segmented nonlinearity of the variable structure system. In this paper, we use the phase plane method, which is suitable for systems of segmented nonlinearity, to analyze the QCN system. The overall dynamic behaviors of the QCN system are presented, and the sufficient conditions for the stable QCN system are deduced. These sufficient conditions serve as guidelines toward proper parameters setting. Moreover, we find that the stability of QCN is mainly promised by the sliding mode motion, which is the underlying reason for QCN's stable queue shown in numerous simulations and experiments. Experiments on the NetFPGA platform verify that the analytical results can explain the complex behaviors of QCN. Wanchun Jiang, Fengyuan Ren, Chuang Lin 0002 |
IEEE/ACM Trans. Netw. | 1 |
| 2014 | Analysis of Backward Congestion Notification with Delay for Enhanced Ethernet NetworksabstractAt present, companies and standards organizations are enhancing Ethernet as the unified switch fabric for all of the TCP/IP traffic, the storage traffic and the high performance computing traffic in data centers. Backward congestion notification (BCN) is the basic mechanism for the end-to-end congestion management enhancement of Ethernet. To fulfill the special requirements of the unified switch fabric, i.e., losslessness and low transmission delay, BCN should hold the buffer occupancy around a target point tightly. Thus, the stability of the control loop and the buffer size are critical to BCN. Currently, the impacts of delay on the performance of BCN are unidentified. When the speed of Ethernet increases to 40 Gbps or 100 Gbps in the near future, the number of on-the-fly packets becomes the same order with the buffer size of switch. Accordingly, the impacts of delay will become significant. In this paper, we analyze BCN, paying special attention on the delay. We model the BCN system with a set of segmented delayed differential equations, and then deduce sufficient condition for the uniformly asymptotic stability of BCN. Subsequently, the bounds of buffer occupancy are estimated, which provides direct guidelines on setting buffer size. Finally, numerical analysis and experiments on the NetFPGA platform verify our theoretical analysis. Wanchun Jiang, Fengyuan Ren, Yongwei Wu 0001, Chuang Lin 0002, Ivan Stojmenovic |
IEEE Trans. Computers | 1 |
| 2013 | Modeling and understanding burst transmission algorithms for energy efficient ethernetabstractRecently, the energy consumption of Ethernet has become one of the hottest topics focused by both academic committee and industry, especially with the increase of the link speed from 1Gbps to 10Gbps nowadays or even 40/100/200Gbps in the near future. To save the energy consumed by the Ethernet, the Energy Efficient Ethernet (EEE) is developed and standardized by the IEEE 802.3az work group. When there is no incoming traffic, the EEE can saves 90 % of its energy consumption by entering into the Low Power Idle (LPI) mode. To maximize the energy saving of Ethernet, the Burst TRansmission (BTR) algorithm, which defines a new way to utilize the LPI mode, is developed as a policy for EEE. Prior work theoretically shows that the BTR algorithm makes a tradeoff between the energy saving and the queuing delay. However, the traffic pattern, on which the performance of EEE greatly depends, is assumed to be deterministic in their analyses. Besides, their models made estimation for many situations. In this paper, assuming that the arrival time of packets can be modeled by Poisson process, we build Markov model for EEE with the BTR algorithm and provide analytical understanding on the BTR algorithm. We propose two actual models: one focuses on the buffer size limit, the other concentrates on tolerable packet delay additionally. We draw some guidelines of parameter selection and policy design for EEE from combination of theory conclusions and simulation results. The results show that the saved energy can be constrained by link occupancy even though the buffer size is variational. The other policy buffer full triggered wake-up can achieve ideal ratio of energy consumption and arrival rate within the scope of the buffer as well. However, the tolerable delay can not be guaranteed by any policies. The buffer size is even fixed, which affects the flexibility of demanded delay for different business. The policy considering tolerable delay is supposed to be a little better than the other policy, with a little more complicated design. Thus we design an adaptive policy: detect the load utilization, apply the buffer full triggered wake-up policy for higher load utilization link, while applying the buffer full and timeout triggered wakeup policy for the delay sensitive business and tiny arrival rate. Jinli Meng, Fengyuan Ren, Wanchun Jiang, Chuang Lin 0002 |
IWQoS | 3 |
| 2012 | Analysis of backward congestion notification with delay for enhanced ethernet networksabstractRecently, companies and standards organizations are enhancing Ethernet as the unified switch fabric for all of the TCP/IP traffic, the storage traffic and the interprocess communication(IPC) traffic in Data Center Networks(DCNs). Backward Congestion Notification(BCN) is the basic mechanism for the end-to-end congestion management enhancement. To fulfill the special requirements of the unified switch fabric that being lossless and of extremely low latency, BCN should hold the queue length around a target point tightly. Thus, the stability of the control loop and the buffer size are critical to BCN. Currently, the impacts of delay on the performance of BCN are unidentified. When the link capacity increases to 40Gbps or 100Gbps in the near future, the number of on-the-fly packets becomes the same order with the shallow buffer size of switches. Thus, the impacts of delay on the performance of BCN will become significant. In this paper, we analyze BCN, paying special attention on the delay. Firstly, we model the BCN system with a set of segmented delayed differential equations. Then, the sufficient condition for the uniformly asymptotic stability of the BCN system is deduced. Subsequently, the bound of buffer occupancy under this sufficient condition are estimated, which provides guidelines on setting buffer size. Finally, the numerical analysis and the experiments on the NetFPGA platform verify the theoretical analysis. Wanchun Jiang, Fengyuan Ren, Chuang Lin 0002, Ivan Stojmenovic |
INFOCOM | 1 |
| 2012 | Sliding Mode Congestion Control for data center Ethernet networksabstractRecently, Ethernet is being enhanced as the unified switch fabric of data centers, called Data Center Ethernet. The end-to-end congestion management is one of the indispensable enhancements, and Quantized Congestion Notification (QCN) has been ratified to be the standard. Our experiments show that QCN suffers from the oscillation of the queue at the bottleneck link. With the changes of system parameters and network configurations, the oscillation may become so serious that the queue is emptied frequently. As a result, the utilization of the bottleneck link degrades. Theoretical analysis shows that QCN approaches to the equilibrium point mainly through the sliding mode motion. But whether QCN enters into the sliding mode motion also depends on both system parameters and network configurations. Hence, we present the Sliding Mode Congestion Control (SMCC) scheme, which can drive the system into the sliding mode motion under any conditions. SMCC benefits from the advantage that the sliding mode motion is insensitive to system parameters and external disturbances. Moreover, SMCC is simple, stable and has short response time. QCN can be replaced by SMCC easily since both of them follow the framework developed by the IEEE 802.1 Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially in the condition that the traffic pattern and the network state are variable. Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002 |
INFOCOM | 1 |
| 2010 | Phase Plane Analysis of Congestion Control in Data Center Ethernet NetworksabstractEthernet has some attractive properties for network consolidation in the data center, but needs further enhancement to satisfy the additional requirements of unified network fabrics. Congestion management is introduced in Ethernet networks to avoid dropping packets due to congestion. The BCN (Backward Congestion Notification) mechanism is a basic element of several standard drafts, and its stability underlies normal network operations. Because the linear stability analysis method is incapable of handling the nonlinearity of the variable structure control employed by the BCN mechanism, some particular phenomena are unexposed and the insights are insufficient. In this paper, we propose the concept of strong stability of the queuing system to satisfy the requirements of no dropped packets in the data center, and build a fluid-flow analytical model for the BCN congestion control system. Considering the nonlinearity involved in the rate regulation laws, we classify the system into different categories according to the shapes of phase trajectories, and conduct a nonlinear stability analysis using phase plane analysis techniques on a case by case basis. The analysis details can provide a comprehensive understanding about the behaviors of the overall congestion control system. Finally, we also deduce an explicit stability criterion presenting the parameters constraints for the strongly stable BCN system, which can provide straightforward guidelines for proper parameter settings. Fengyuan Ren, Wanchun Jiang |
ICDCS | 2 |