VLDB 2026 Research / reviewers in the wild / expert
Jinbin Hu 0001
dblp:116/0938-1
· DBLP profile ↗
59ranked-venue papers
36as first author
53since 2021 · last 2026
0000-0001-8216-9683ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 30 · 18 first-author · 26 since 2021Systems, architecture and hardware · 22 · 15 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaLB: Gap-Aware Load Balancing for High-Performance AI Training in RDMA Networks
Jinbin Hu 0001, Zijing Zhong, Rui Zhi, Jin Wang 0001 |
IWQoS | 1 |
| 2026 | RTDSeg: Hard example sampling driven Real-Time Concrete Structural Damage Segmentation network
Jing Wang 0209, Haizhou Yao, Jinbin Hu 0001, Jin Wang 0001, Yafei Ma |
Adv. Eng. Informatics | 4 |
| 2026 | A High-Performance Sketch With Dynamic Memory Allocation for Priority-Oriented Data Stream ProcessingabstractSketch is widely used in many traffic estimation tasks due to its good balance among accuracy, speed, and memory usage. In scenarios with priority flows, priority-aware sketch, as an emerging method, provides differentiated detection accuracy for flows of different priorities, optimizing resource allocation and improving the detection accuracy of high-priority flows. However, existing priority-aware sketches methods struggle to effectively handle the dynamic changes in flow priority distribution in realworld detection environments, leading to wasted or insufficient storage space. To address this issue, this paper proposes a new priority-aware sketch with Dynamic Memory Allocation called DMA-Sketch. It dynamically adjusts the detection framework based on flow priority distribution information and adaptively allocates appropriate memory space to each storage region. The experimental results show that DMA-Sketch improves the overall priority accuracy, high-priority accuracy and throughput by up to 1.33×, 16.39× and 1.88×, respectively, under the scenarios with changing flow priority distribution over the state-of-the-art schemes. Jinbin Hu 0001, Houqiang Shen, Jiawei Huang 0001, Robert Simon Sherratt, Jin Wang 0001 |
IEEE Trans. Computers | 1 |
| 2026 | Reordering-Resilient Multipath Transport for RDMA-Enabled Cloud DatacentersabstractRemote direct memory access (RDMA) is widely deployed in production data centers to enable low-latency transmission. The current multipath RDMA transmission protocols effectively improve link utilization by allocating traffic to equal-cost parallel paths. To address packet reordering, they struggle to control the level of out-of-order packets by using bitmaps. However, under asymmetric path status and highly dynamic traffic scenarios, a large number of out-of-order packets easily cause bitmap overflow and frequent unnecessary retransmission, resulting in goodput far below throughput. Motivated by this, we present MPTR, an efficient multipath transport with robust reordering for RDMA networks. At its core, MPTR continuously monitors the multipath congestion status at the receiver and distributes the traffic in a congestion-aware manner to proactively reduce the degree of out-of-order and avoid triggering retransmission due to bitmap cache overflow. The NS-3 simulation results show that MPTR effectively reduces unnecessary retransmission and improves goodput under realistic workloads by up to 34%, 49%, and 51% compared to multi-path remote direct memory access (MP-RDMA), ConWeave, and data center quantized congestion notification (DCQCN), respectively. Jin Wang 0001, Ruiqian Li, Jalel Ben-Othman, Jinbin Hu 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2026 | SALB: Security-Aware Load Balancing for Large Language Model Training in Datacenter NetworksabstractTo meet the massive compute and high-speed communication demands of Large Language Model (LLM) training, modern datacenters typically adopt multipath topologies such as Fat-Tree and Clos to host parallel jobs across hundreds to thousands of GPUs. However, LLM training exhibits periodic, high-bandwidth communication patterns. Existing load-balancing schemes become misaligned under dynamic congestion and anomalous surges: they struggle to promptly mitigate iteration-peak congestion and lack effective isolation of anomalous traffic. To address this, we propose Security-Aware Load Balancing (SALB) for LLM training. SALB leverages a Deep Reinforcement Learning (DRL) controller with queue and delay signals for packet-level multipath load balancing and employs path binding to confine suspicious flows. By integrating data security into load balancing, SALB simultaneously achieves high throughput and robust traffic isolation. NS-3 simulation results show that, compared with CONGA, Hermes, and ConWeave, SALB reduces the 99th-percentile flow completion time (FCT) of short flows by an average of 65% and increases the throughput of long flows by an average of 54%. It further outperforms the baselines in aggregate throughput, path utilization, and packet loss rate, thereby significantly enhancing system stability, robustness, and data security. Wangqing Luo, Jinbin Hu 0001, Pradip Kumar Sharma, Jin Wang 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2026 | Toward Fine-Grained Load Balancing With Congested-Flow Isolation in Lossless DatacentersabstractRemote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) cooperating with Priority Flow Control (PFC) has been widely deployed in production datacenters to enable low latency, lossless transmission. At the same time, modern datacenters typically offer parallel transmission paths between any pair of end-hosts, underscoring the importance of load balancing. However, the well-studied load balancing mechanisms designed for lossy datacenter networks (DCNs) are ill-suited for such lossless environments. Through extensive experiments, we are among the first to comprehensively inspect the interactions between PFC and load balancing, and uncover that existing fine-grained rerouting schemes can be counterproductive to spread the congested flows among more paths, further aggravating PFC’s head-of-line (HoL) blocking. Motivated by this, we present FLB, a Fine-grained Load Balancing scheme for lossless DCNs. At its core, FLB employs threshold-free rerouting to effectively balance traffic load and improve link utilization during normal conditions and leverages timely congested flow isolation to eliminate HoL blocking on non-congested flows when congestion occurs. To handle complex multi-bottleneck scenarios, we further introduce FLB*, which incorporates an enhanced congestion-point-aware isolation mechanism using Congestion Point Identifiers (CPI) to eliminate HoL blocking among different congested flows.We have fully implemented a FLB prototype, and our evaluation results show that FLB reduces PFC PAUSE rate by up to 96% and avoids HoL blocking, translating to up to 45% improvement in goodput over CONGA+DCQCN and 40%, 36%, 29% and 18% reduction in average flow completion time (FCT) over LetFlow+Swift, MP-RDMA, Proteus+DCQCN and LetFlow+PCN, respectively. Jinbin Hu 0001, Siyao Li, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Mengyu Ma, Jin Wang 0001, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 1 |
| 2026 | DSA: Efficient Data-Plane Memory Scheduler for In-Network Aggregation to Accelerate Distributed TrainingabstractTo reduce the traffic volume and accelerate communication in distributed training (DT) jobs, recent works introduce In-Network Aggregation (INA) to move the gradient summation into network programmable switches. However, switch memory is a scarce resource, unable to support massive DT jobs in data centers, and existing INA solutions have not utilized switch memory to the best extent. We propose DSA, an Efficient Data-Plane switch memory Scheduler for in-network Aggregation. DSA introduces preemption to the switch memory management for INA jobs. Furthermore, under packet preemption scenarios, DSA optimizes the selective retransmission mechanism to reduce redundant retransimtting packets to alleviate congestion. In the data plane, DSA allows gradient tensors with high priority to preempt the switch aggregators (basic computation unit in INA) from tensors with low priority, which avoids an aggregator wasting time in idle. In the control plane, DSA devises a priority policy which assigns high priority to gradient tensors that benefit overall job efficiency more, e.g., communication-intensive jobs. We implement the prototype of DSA. The experimental results show that DSA can improve the average job completion time (JCT) by up to 1.35x compared with baseline solutions. Jinbin Hu 0001, Xinming Xu, Hao Wang 0116, Jin Wang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 1 |
| 2026 | Towards Optimal Communication Scheduling With Automatic Configuration for Distributed DNN TrainingabstractByteScheduler partitions and rearranges tensor transmissions to improve the communication efficiency of distributed Deep Neural Network (DNN) training. The configuration of hyper-parameters (i.e., the partition size and the credit size) is critical to the effectiveness of partitioning and rearrangement. Currently ByteScheduler adopts Bayesian Optimization (BO) to find the optimal configuration for the hyper-parameters beforehand. In practice, however, various runtime factors (such as worker node status and network conditions) change over time, making the statically-determined one-shot configuration result suboptimal for real-world DNN training. To address this problem, in this paper we present a realtime configuration method (called AutoByte) that automatically and timely searches the optimal hyper-parameters as the training systems dynamically change. AutoByte extends the ByteScheduler framework with a meta network, which takes the systems’ runtime statistics as its input, dynamically adjusts the triggering threshold based on system environment characteristics, and outputs predictions for speedups under specific configurations. Evaluation results on various DNN models show that AutoByte can dynamically tune the hyper-parameters with low resource usage, and deliver up to 33.2% higher performance than the best static configuration method on the ByteScheduler framework. Jinbin Hu 0001, Xinming Xu, Hao Wang 0116, Yiqing Ma, Yiming Zhang 0003, Jin Wang 0001, Kai Chen 0005 |
IEEE Trans. Netw. | 1 |
| 2025 | Hierarchical-Caching-Driven Distributed Architecture for Accelerating Model Training
Jinbin Hu 0001, Wenda Tang, Jin Wang 0001 |
ICA3PP (8) | 1 |
| 2025 | HaLB: Heterogeneous Traffic-Aware Load Balancing for Minimizing Deadline Misses in AI-Centric Datacenter Networks
Jinbin Hu 0001, Rui Zhi, Jin Wang 0001 |
ICA3PP (8) | 1 |
| 2025 | Towards Efficient Multi-path Transport with Robust Reordering in RDMA Datacenter Networks
Jin Wang 0001, Jinbin Hu 0001 |
ICA3PP (8) | 3 |
| 2025 | A High-Accuracy Sketch for Measuring Low-Entropy Flows in Distributed AI TrainingabstractDistributed AI training generates unique low-entropy flow patterns with predictable, singular and repetitive flows that differ fundamentally from traditional network flow with heavy-tailed distributions. While sketch-based methods are widely used for network measurement, existing approaches fail to exploit these distinctive characteristics, resulting in poor measurement accuracy. To address this issue, this paper proposes FP-Sketch, a high-accuracy sketch for measuring low-entropy flows with Flow Prediction. FP-Sketch utilizes a staging queue to predict and classify flows of different sizes, thereby leveraging the singular, repetitive, and predictable nature of low-entropy AI flows. Combined with hierarchical storage, our method achieves superior measurement precision for distributed AI workloads. We establish rigorous error bounds for FP-Sketch through theoretical analysis. The experimental results show that FP-Sketch reduces flow estimation error by 38.6% and improves insertion throughput by 49.4% compared to the state-of-the-art alternatives. Jin Wang 0001, Chenye Zhu, Jinbin Hu 0001 |
ICPP | 3 |
| 2025 | Enabling In-Network Acceleration Over the Cloud
Hao Wang 0116, Decang Sun, Jinbin Hu 0001, Kai Chen 0005 |
INFOCOM | 3 |
| 2025 | DMA-Sketch: A Fast and Accurate Sketch for Priority-Oriented Data Stream ProcessingabstractSketch is widely used in many traffic estimation tasks due to its good balance among accuracy, speed, and memory usage. In scenarios with priority flows, priority-aware sketch, as an emerging method, provides differentiated detection accuracy for flows of different priorities, optimizing resource allocation and improving the detection accuracy of high-priority flows. However, existing priority-aware sketches methods struggle to effectively handle the dynamic changes in flow priority skew in realworld detection environments, leading to wasted or insufficient storage space. To address this issue, this paper proposes a new priority-aware sketch with Dynamic Memory Allocation called DMA-Sketch. It dynamically adjusts the detection framework based on flow priority skew information and adaptively allocates appropriate memory space to each storage region. The experimental results show that DMA-Sketch improves the overall priority accuracy, high-priority accuracy and throughput by up to$1.33 \times, 16.39 \times$and$1.88 \times$, respectively, under the scenarios with changing flow priority skew over the state-of-theart schemes. Jinbin Hu 0001, Houqiang Sheng, Ying Liu 0064, Jin Wang 0001 |
IWQoS | 1 |
| 2025 | FLB: Fine-grained Load Balancing for Lossless Datacenter Networks
Jinbin Hu 0001, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005 |
USENIX ATC | 1 |
| 2025 | Deadline-aware load balancing for coflow in datacenter networks
Zhichen Wang, Jinbin Hu 0001, Jin Wang 0001, Fayez Alqahtani 0001, Amr Tolba |
Comput. Networks | 3 |
| 2025 | Traffic-Aware Load Balancing Based on Deep Reinforcement Learning in Cloud-Based Industrial Data CentersabstractModern industrial datacenter networks employ multirooted tree topologies to accommodate a diverse range of cloud applications, which generate heterogeneous traffic with low-latency short flows and high-throughput long flows. Recently, the proposed learning-based load balancing mechanisms are resilient to dynamic network, but they are agnostic to heterogeneous traffic, resulting in large tail delay. In this article, we propose a new deep reinforcement learning (DRL) based load balancing called DRLB, which uses DRL with the distributed distributional deterministic policy gradients algorithm to make (re)routing for long flows, and adopts the weighted cost multipathing mechanism for short flows. Furthermore, this article introduces a traffic feature-based dynamic training cycle mechanism to adaptively adjust the training cycles. The experimental results show DRLB reduces the flow completion time of short flows by up to 58% and improves the throughput of long flows by 38% compared to the state-of-the-art load balancing mechanisms. Jinbin Hu 0001, Wangqing Luo, Amr Tolba, Jin Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | SRCC: Sub-RTT Congestion Control for Lossless Datacenter NetworksabstractTo meet the stringent requirements of industrial applications, modern Ethernet datacenter networks widely deployed with remote direct memory access (RDMA) technology and priority-based flow control (PFC) scheme aim at providing low latency and high throughput transmission performance. However, the existing end-to-end congestion control cannot handle the transient congestion timely due to the round-trip-time (RTT) level control loop, inevitably resulting in PFC triggering. In this article, we propose a Sub-RTT congestion control mechanism called SRCC to alleviate bursty congestion timely. Specifically, SRCC identifies the congested flows accurately, notifies congestion directly from the hotspot to the corresponding source at the sub-RTT control loop and adjusts the sending rate to avoid PFC's head-of-line blocking. Compared to the state-of-the-art end-to-end transmission protocols, the evaluation results show that SRCC effectively reduces the average flow completion time (FCT) by up to 61%, 52%, 40%, and 24% over datacenter quantized congestion notification (DCQCN), Swift, high precision congestion control (HPCC), and photonic congestion notification (PCN), respectively. Jinbin Hu 0001, Shuying Rao, Jiawei Huang 0001, Jianxin Wang 0001, Jin Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | ARS: Adaptive Routing System for Heterogeneous Traffic in Industrial Data CentersabstractModern industrial datacenter networks carry latency-sensitive and throughput-oriented applications with diverse requirements. Recent load balancing mechanisms effectively reduce latency and improve throughput for heterogeneous traffic. However, deadline-sensitive flows still often miss deadlines due to being blocked. In this article, we introduce an adaptive routing system (ARS) to avoid missing deadlines. Specifically, an ARS computes a heuristic function using three influence factors, derives the probability of choosing the next node, and finds optimal (re)routing path. The experimental results show that an ARS enhances the throughput for long flows and decreases the average flow completion time and the deadline miss rate by 24% and 55%, respectively, compared to state-of-the-art load balancing schemes. Jinbin Hu 0001, Rui Zhi, Jin Wang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Proactive Transport With High Link Utilization Using Opportunistic Packets in Cloud Data CentersabstractTo meet the stringent demanding low latency and high throughput of cloud datacenter applications, recent receiver-driven transport protocols transmit only one packet once receiving each credit packet from the receiver to achieve ultra-low queueing delay. However, the round-trip time variation and the highly dynamic background traffic significantly deteriorate the performance of receiver-driven transport protocols, resulting in under-utilized bandwidth. This paper designs a simple yet effective solution called RPO, which retains the advantages of receiver-driven transmission while efficiently utilizing the available bandwidth. Specifically, RPO rationally uses low-priority opportunistic packets to ensure high network utilization without increasing the queueing delay of high-priority normal packets. Furthermore, to tackle the queueing buildup due to line-rate transmission in the first RTT, we design a selective dropping mechanism called SDM to help the majority of small flows complete within only one RTT by prioritizing the first-RTT bursty packets over the packets triggered by grants. We implement RPO in Linux hosts with DPDK. The experimental results show that RPO significantly improves the network utilization by up to 35% over the state-of-the-art schemes, without introducing additional queueing delay. Moreover, RPO integrated with SDM reduces the AFCT of small flows by up to 45% compared with RPO integrated with Aeolus. Jinbin Hu 0001, Jiawei Huang 0001, Yijun Li 0002, Shuying Rao, Wenchao Jiang, Kai Chen 0005, Jianxin Wang 0001, Tian He 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Hierarchical Adaptive Learning-Based Congestion Control With Low Training Overhead for Datacenter NetworksabstractMost congestion control mechanisms perform well in specific datacenter networks, but none can consistently deliver good performance across varying scenarios. Recently proposed frameworks based on reinforcement learning can flexibly select congestion control algorithms to adapt to dynamic network. However, frequently altering the congestion control mechanisms during relatively stable periods of the network actually leads to instability and unnecessary computational overhead. In this paper, we propose a lightweight and hierarchical adaptive congestion control algorithm (LACC) to be resilient to the varying network. LACC dynamically selects the appropriate congestion control mechanism only when the current congestion control algorithm is not suitable for the current network state, rather than changing the congestion control scheme every training cycle to ensure network stability. The simulation results show that LACC significantly reduces the average overhead by 31% and improves throughput by up to 47%, 35%, 23% and 15% compared to Cubic, Reno, BBR and Antelope, respectively. Jinbin Hu 0001, Zikai Zhou, Jing Wang 0209 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2025 | $R^{3}$R3: A Building Block for Disordering-Tolerant Load Balancing in Data Center NetworksabstractPacket-level load balancing has shown its massive potential for long in utilizing super high bisection bandwidth of data center network (DCN). This kind of potential, however, has still not been completely transformed into huge performance enhancement of data transmission. The fundamental reason is that packet-level load balancing can fully utilize the parallel paths of underlying physical network, but suffer from the problem of packet disordering transmission, which greatly impairs the flow-level transmission performance of DCN. This paper explores the root cause of performance impairment generated by packet disordering transmission, and proposes$R^{3}$, a solution focusing on “recognizably releasing redundant acknowledgements” as a building block for data center packet-level load balancer. In$R^{3}$'s heart, the source leaf switch perceives the global packet loss information and selectively intercepts the redundant acknowledgement packets, thus avoiding the TCP-driven end-host from experiencing frequent window reductions and unnecessary packet retransmissions. Experimental results of numerous simulation tests and real implementations show that, after integrating$R^{3}$into the representative data center packet-level load balancing schemes, the transmission performances of both delay-sensitive and throughput-oriented data center flows are significantly improved. Furthermore,$R^{3}$is merely implemented by switch, leaving the end hosts and the deployed load balancing scheme totally unchanged. Tao Zhang 0019, Yuanzhen Hu, Jinbin Hu 0001, Haotian Jing, Yangfan Li 0001, Xidao Luan |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | AutoPipe: Automatic Configuration of Pipeline Parallelism in Shared GPU ClusterabstractAs training Deep Neural Network (DNN) is time-consuming, people resort to parallelization across multiple accelerators. A plethora of solutions adopt data/model parallelization, but they suffer from frequent weight synchronization overhead or resource under-utilization. Recent work introduces pipeline parallelism to improve the utilization of accelerators, however, most existing pipeline parallelism approaches take a one-shot configuration, while ignoring the fluctuation of available resources, e.g., bandwidth and GPUs. Moreover, the heuristic work partition methods oversimplify the computation and communication process, leading to sub-optimal results. To address this challenge, we present AutoPipe, a self-adaptive pipeline parallelism optimization solution. At its core, AutoPipe introduces a reinforcement learning (RL) based work partitioning model, which takes into account both exact communication procedure and dynamic state switching. To mitigate the stalls on state switching, AutoPipe adopts layer-by-layer computation under switching. We have implemented an AutoPipe prototype and evaluated it via testbed experiments. Our results show that the AutoPipe-enhanced PipeDream can find better work partitioning and benefit from dynamic configuration, outperforming the vanilla solutions by up to 89% for exclusive tasks and 143% in dynamic workloads. Furthermore, we show that AutoPipe can also work well with other pipeline parallelism schemes and achieve considerable performance gains. Jinbin Hu 0001, Ying Liu 0064, Hao Wang 0116, Jin Wang 0001 |
ICPP | 1 |
| 2024 | Improving Availability and Scalability for RDMA Load Balancing with In-network ReorderingabstractRemote Direct Memory Access (RDMA) is widely deployed in datacenter networks (DCNs) due to its ultra-low latency, high throughput, and low CPU overhead. Since RDMA is sensitive to out-of-order packets, the previous load balancing schemes designed based on TCP do not work well in RDMA networks. Recently proposed load balancing schemes focus on solving packet reordering within the network. However, the existing solutions cannot be extended in practice because the required queues far exceed common switch capabilities. In this paper, we propose a scalable and efficient load balancing called SELB to improve availability and scalability. SELB employs a clustering algorithm to categorize equal-cost paths and then reroutes traffic to the same cluster parallel paths to reduce the degree of out-of-order and improve queue utilization. The NS-3 simulation results demonstrate that SELB reduces the average flow completion time (FCT) and the 99th percentile FCT by up to 33% and 21%, respectively, compared to the state-of-the-art load balancing schemes. Jinbin Hu 0001, Ruiqian Li, Shuying Rao, Jin Wang 0001 |
ISPA | 1 |
| 2024 | DAR: Deadline-Aware Rerouting for Mix-flows in Datacenter NetworksabstractIn modern datacenter networks (DCNs), the booming online data-intensive applications generate mix-flows with or without deadlines. Balancing these heterogenous flows among parallel equal-cost paths to meet the tight deadlines is crucial. However, due to the unaware of deadlines, the existing load balancing mechanisms cannot choose suitable (re)routing path for mix-flows to meet their respective stringent requirements. In this paper, we propose a deadline-aware rerouting scheme called DAR, which applies different routing strategies for mix-flows. Specifically, DAR first perceives the deadline flows and then categorizes them based on the urgency of the deadline, and employs different (re)routing strategies to ensure that flows with more urgent deadlines are completed earlier. The NS-3 simulation results show that DAR effectively balances mix-flows. For example, compared to the state-of-the-art load balancing schemes, DAR reduces the deadline miss rate and the average flow completion time (AFCT) by up to 38% and 35.5%, respectively. Jinbin Hu 0001, Rui Zhi, Shuying Rao, Ying Liu 0064, Jin Wang 0001 |
ISPA | 1 |
| 2024 | Learning-Based Hierarchical Adaptive Congestion Control with Low Training OverheadabstractMost congestion control mechanisms perform well in specific network environments, but none can consistently deliver good performance across all scenarios. Recently proposed frameworks based on reinforcement learning can flexibly select congestion control algorithms to adapt to dynamic changes in network conditions. However, frequently altering the congestion control mechanisms during relatively stable periods of the network actually leads to instability and unnecessary computational overhead. In this paper, we propose a hierarchical adaptive congestion control algorithm (HACC) to be resilient to the varying network. HACC dynamically selects the appropriate congestion control mechanism only when the current congestion control algorithm is not suitable for the current network state, rather than changing the congestion control scheme every training cycle to ensure network stability. The simulation results show that under different realistic workloads, HACC significantly reduces the computational overhead and improves throughput. Specifically, HACC reduces average overhead by 31% and improves throughput by up to 47%, 35%, 23%, and 15% compared to Cubic, Reno, BBR, and Antelope, respectively. Jinbin Hu 0001, Zikai Zhou, Shuying Rao, Yujie Peng, Bowen Bao, Chang Ruan |
ISPA | 1 |
| 2024 | TaLB: Tensor-aware Load Balancing for Distributed DNN Training AccelerationabstractIncreasingly large-scale models and rich data sets make communication overhead a key bottleneck for distributed Deep Neural Network (DNN) training, constantly attracting the attention of academia and industry. Despite continuous efforts, prior solutions such as pipelining computation/communication and in-network gradient compression/scheduling do not focus on how to accelerate DNN training through load balancing in datacenter networks (DCNs). However, the existing load balancing mechanisms are unaware of tensor integrity and priority for gradient parameter synchronization during the DNN training iterations, resulting in severe tensor tail latency and slow model convergence speed. In this paper, we present a Tensor-aware Load Balancing (TaLB) scheme to accelerate DNN training. Specifically, TaLB identifies the different priority tensors and makes (re)routing decisions based on the tensor-level granularity to cut the high-priority tensors tail delay. The testbed implementation and large-scale NS-3 simulation results show that TaLB effectively accelerates DNN training speed. For example, TaLB significantly reduces the average flow completion time (FCT) by up to 55%, and accelerates the model training speed up to 2.37× on VGG19, ResNet50 and AlexNet models. Jinbin Hu 0001, Yi He 0017, Wangqing Luo, Jiawei Huang 0001, Jianxin Wang 0001, Jin Wang 0001 |
IWQoS | 1 |
| 2024 | Deployment optimization in wireless sensor networks using advanced artificial bee colony algorithm
Jueyu Zhu, Jifang Rong, Ying Liu 0064, Fayez Alqahtani 0001, Amr Tolba, Jinbin Hu 0001 |
Peer Peer Netw. Appl. | 8 |
| 2024 | HG: Leveraging Hybrid Switching Granularity to Balance Heterogeneous Data Center Traffic Load for Cloud-Based Industrial ApplicationsabstractNowadays, the deluge of heterogeneous data generated by various cloud-based industrial applications often has to be delivered to the data center for analysis and storage. To speed up data processing thus facilitating application performance, the modern data center network offers rich parallel paths and super high bisection bandwidth for data communications between servers, expecting to provide good transmission performance for the heterogeneous data traffic caused by cloud-based industrial applications. Due to high path diversities, however, balancing the heterogeneous traffic load across multiple parallel paths for fully utilizing the offered super high bisection bandwidth is full of challenges (i.e., how to achieve high path utilization without incurring adverse impact). Although prior studies demonstrate that the flowlet-based solutions are promising to fill the bill, we argue that their rerouting operations are still inappropriate in timing and manner. This article presents HG, a load balancing scheme adopting hybrid switching granularity to make traffic rerouting. HG embeds the flow-fragment-based and flowcell-based path switching into the flowlet-based path switching, and employs state-weighted path measurement to choose paths for newly appeared flow fragments, flowcells, and flowlets. The results of numerous NS2 tests show that, compared with the state-of-the-art data center load balancing schemes, HG significantly reduces the average and tail-flow completion times for delay-sensitive flows, and the throughput of throughput-oriented flows is always maintained at high level. Tao Zhang 0019, Shengli He, Ku Jin, Yuanzhen Hu, Chang Ruan, Shaojun Zou, Jinbin Hu 0001, Fangmin Li |
IEEE Trans. Ind. Informatics | 9 |
| 2024 | Lightweight Automatic ECN Tuning Based on Deep Reinforcement Learning With Ultra-Low Overhead in Datacenter NetworksabstractIn modern datacenter networks (DCNs), mainstream congestion control (CC) mechanisms essentially rely on Explicit Congestion Notification (ECN) to reflect congestion. The traditional static ECN threshold performs poorly under dynamic scenarios, and setting a proper ECN threshold under various traffic patterns is challenging and time-consuming. The recently proposed reinforcement learning (RL) based ECN Tuning algorithm (ACC) consumes a large number of computational resources, making it difficult to deploy on switches. In this paper, we present a lightweight and hierarchical automated ECN tuning algorithm called LAECN, which can fully exploit the performance benefits of deep reinforcement learning with ultra-low overhead. The simulation results show that LAECN improves performance significantly by reducing latency and increasing throughput in stable network conditions, and also shows consistent high performance in small flows network environments. For example, LAECN effectively improves throughput by up to 47%, 34%, 32% and 24% over DCQCN, TIMELY, HPCC and ACC, respectively. Jinbin Hu 0001, Zikai Zhou, Jin Zhang 0018 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2024 | Enhancing Load Balancing With In-Network Recirculation to Prevent Packet Reordering in Lossless Data CentersabstractMany existing load balancing mechanisms work effectively in lossy datacenter networks (DCNs), but they suffer from serious packet reordering in lossless Ethernet DCNs deployed with the hop-by-hop Priority-based Flow Control (PFC). The key reason is that the prior solutions are not able to perceive PFC triggering correctly and in a timely manner when making load balancing decisions. Once the forwarding path pauses transmission due to PFC triggering, the packets allocated on it are blocked, inevitably leading to out-of-order packets and retransmission. In this paper, we present an Reordering-robust Load Balancing (RLB) scheme with PFC prediction in lossless DCNs. At its heart, RLB leverages the derivative of ingress queue length to predict PFC triggering and proactively notifies the upstream switches to choose an appropriate rerouting path or perform packet recirculation to avoid reordering. Furthermore, under switch failure scenarios, RLB adjusts the recirculation threshold adaptively to mitigate the risk of packets over-recirculation. We have implemented RLB in the hardware programmable switch. As a building block for existing load balancing mechanisms, we have integrated RLB into Presto, LetFlow, Hermes and DRILL. The evaluation results show that the RLB-enhanced solutions deliver significant performance by avoiding packet reordering. For example, it reduces the$99^{th}$percentile flow completion time (FCT) by up to 72%, 67%, 58% and 54% over DRILL, Presto, LetFlow and Hermes, respectively. Jinbin Hu 0001, Yi He 0017, Wangqing Luo, Jiawei Huang 0001, Jin Wang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Load Balancing With Multi-Level Signals for Lossless Datacenter NetworksabstractVarious datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-level signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35%, 31%, 28%, 22% and 46%, 42%, 34%, 29% better average FCT and$99^{th}$percentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively. Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | FlowSail: Fine-Grained and Practical Flow Control for Datacenter NetworksabstractAs datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effectiveness due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce FlowSail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of FlowSail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FlowSail using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that FlowSail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FlowSail achieves$4.3\times $lower tail latency under the same number of queues, matches existing FC schemes with$4\times $fewer queues, and holds robust performance with a minimum of 2 queues. Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 3 |
| 2024 | High-Performance Hardware Acceleration Architecture for Cross-Silo Federated LearningabstractCross-silo federated learning (FL) adopts various cryptographic operations to preserve data privacy, which introduces significant performance overhead. In this paper, we identify nine widely-used cryptographic operations and design an efficient hardware architecture to accelerate them. However, directly offloading them on hardware statically leads to (1) inadequate hardware acceleration due to the limited resources allocated to each operation; (2) insufficient resource utilization, since different operations are used at different times. To address these challenges, we propose FLASH, a high-performance hardware acceleration architecture for cross-silo FL systems. At its heart, FLASH extracts two basic operators—modular exponentiation and multiplication—behind the nine cryptographic operations and implements them as highly-performant engines to achieve adequate acceleration. Furthermore, it leverages a dataflow scheduling scheme to dynamically compose different cryptographic operations based on these basic engines to obtain sufficient resource utilization. We have implemented a fully-functional FLASH prototype with Xilinx VU13P FPGA and integrated it with FATE, the most widely-adopted cross-silo FL framework. Experimental results show that, for the nine cryptographic operations, FLASH achieves up to$14.0\times$and$3.4\times$acceleration over CPU and GPU, translating to up to$6.8\times$and$2.0\times$speedup for realistic FL applications, respectively. We finally evaluate the FLASH design as an ASIC, and it achieves$23.6\times$performance improvement upon the FPGA prototype. Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Han Tian, Kai Chen 0005 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Scaling Switch-driven Flow Control with AquariusabstractAs datacenter networks support more diverse applications and faster link speeds, effective end-to-end congestion control becomes increasingly challenging due to the inherent feedback delay. To address this issue, switch-driven per-hop flow control (FC) has gained popularity due to its natural flow isolation, timely control loop, and ability to handle transient congestion. However, the ideal FC requires impractical hardware resources, and the state-of-the-art approximation approach still demands a large number of queues that exceeds common switch capabilities, limiting scalability in practice. Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005 |
APNet | 3 |
| 2023 | Adaptive Routing for Datacenter Networks Using Ant Colony Optimization
Jinbin Hu 0001, Man He, Shuying Rao, Jing Wang 0209, Shiming He |
ICA3PP (3) | 1 |
| 2023 | Deep Reinforcement Learning Based Load Balancing for Heterogeneous Traffic in Datacenter Networks
Jinbin Hu 0001, Wangqing Luo, Yi He 0017, Jin Wang 0001, Dengyong Zhang |
ICA3PP (3) | 1 |
| 2023 | Enabling Traffic-Differentiated Load Balancing for Datacenter Networks
Jinbin Hu 0001, Ying Liu 0064, Shuying Rao, Jing Wang 0209, Dengyong Zhang |
ICA3PP (3) | 1 |
| 2023 | HAECN: Hierarchical Automatic ECN Tuning with Ultra-Low Overhead in Datacenter Networks
Jinbin Hu 0001, Youyang Wang, Zikai Zhou, Shuying Rao, Rundong Xin, Jing Wang 0209, Shiming He |
ICA3PP (3) | 1 |
| 2023 | Image Inpainting Forensics Algorithm Based on Dual-Domain Encoder-Decoder Network
Dengyong Zhang, En Tan, Feng Li 0065, Jing Wang 0209, Jinbin Hu 0001 |
ICA3PP (5) | 6 |
| 2023 | Enabling Load Balancing for Lossless DatacentersabstractVarious datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-Ievel signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35 %, 31 %, 28%, 22% and 46 %, 42 %, 34 %, 29 % better average FCT and 99thpercentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively. Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005 |
ICNP | 1 |
| 2023 | Towards Fine-Grained and Practical Flow Control for Datacenter NetworksabstractAs datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effective due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce Flowsail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of Flowsail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FLOWSAIL using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that Flowsail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FLOWSAIL achieves 4.3 x lower tail latency under the same number of queues, matches existing FC schemes with 4 x fewer queues, and holds robust performance with a minimum of 2 queues. Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005 |
ICNP | 3 |
| 2023 | RLB: Reordering-Robust Load Balancing in Lossless Datacenter NetworksabstractMany existing load balancing mechanisms work effectively in lossy datacenter networks (DCNs), but they suffer from serious packet reordering in lossless Ethernet DCNs deployed with the hop-by-hop Priority-based Flow Control (PFC). The key reason is that the prior solutions are not able to correctly and timely perceive PFC triggering when making load balancing decisions. Once the forwarding path pauses transmission due to PFC triggering, the packets allocated on it are blocked, inevitably leading to out-of-order packets and retransmission. In this paper, we present a Reordering-robust Load Balancing (RLB) scheme with PFC prediction in lossless DCNs. At its heart, RLB leverages the derivative of ingress queue length to predict PFC triggering and proactively notifies the upstream switches to choose an appropriate rerouting path or perform packet recirculation to avoid reordering. As a building block for existing load balancing mechanisms, we have integrated RLB into Presto, LetFlow, Hermes and DRILL. The test results show that the RLB-enhanced solutions deliver significant performance by avoiding packet reordering. For example, it reduces the 99th percentile flow completion time (FCT) by up to 58%, 67%, 72% and 54% over Presto, LetFlow, Hermes and DRILL, respectively. Jinbin Hu 0001, Yi He 0017, Jin Wang 0001, Wangqing Luo, Jiawei Huang 0001 |
ICPP | 1 |
| 2023 | FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated Learning
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Kai Chen 0005 |
NSDI | 5 |
| 2023 | A novel self-adaptive multi-strategy artificial bee colony algorithm for coverage optimization in wireless sensor networks
Jin Wang 0001, Ying Liu 0064, Shuying Rao, Xinyu Zhou 0002, Jinbin Hu 0001 |
Ad Hoc Networks | 5 |
| 2023 | Load balancing for heterogeneous traffic in datacenter networks
Jin Wang 0001, Shuying Rao, Ying Liu 0064, Pradip Kumar Sharma, Jinbin Hu 0001 |
J. Netw. Comput. Appl. | 5 |
| 2023 | Reducing tail latency with coding-based packet spraying in edge datacentersabstractModern cloud computing applications have stringent low-latency and high-throughput requirements to meet the increasingly diverse demands from customers. In edge datacenters, the popular load balancing scheme based on packet spraying, i.e., random packet spraying (RPS), makes good use of multiple equal-cost paths to ensure high link utilization and efficient transmission. However, RPS performs poorly under the bursty traffic scenario due to packet loss or even timeout. Therefore, to reduce the tail latency caused by retransmission, we design a coding-based random packet spraying named CRPS. Specifically, the source host transmits forward error correction (FEC) encoded packets and dynamically adjusts the data redundancy based on the packet loss rate. Then the switch randomly sprays the encoded packets across all equal-cost multiple paths to implement parallel transmission. In this way, once enough encoded packets from any parallel paths arrive at the destination host, the original packets can be decoded immediately to address the adverse impact of packet loss. The NS-2 simulation results show that CRPS effectively reduces the probability of timeout and significantly improves the tail flow completion time (FCT) by up to 72% compared with the state-of-the-art multipath transmission schemes. Jing Wang 0209, Man He, Jinbin Hu 0001, Naixue Xiong |
J. Syst. Archit. | 4 |
| 2023 | Achieving Fast Convergence and High Efficiency using Differential Explicit Feedback in Data CenterabstractSince most flows are short-lived in data center networks, fast convergence becomes very important to help the short flows effectively utilize high bandwidth. Though current explicit feedback-based transport control protocols (TCPs) provide fast convergence via fine-grained congestion information from customized switches, they unavoidably incur large traffic overhead for widely existing small packets in data center applications, resulting in suboptimal network efficiency. To solve this issue, we propose a datacenter TCP based onDifferentialExplicitCongestionNotification, called DECN, to achieve fast convergence without any traffic overhead. Specifically, DECN feeds rate difference between the target and current rate back to the source by using multiple consecutive packets. Besides, we propose an enhanced version DECN* which obtains the optimal number of consecutive packets according to the packet loss rate. The experimental results of NS2 simulation and testbed implementation show that DECN and its enhanced version DECN* achieve comparable fast convergence as XCP without incurring any extra feedback overhead. Compared with the state-of-the-art explicit feedback-based TCPs, they reduce the flow completion time by up to 34% in typical data center applications. Jiawei Huang 0001, Jingling Liu, Sen Liu 0002, Jinbin Hu 0001, Jianxin Wang 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | REN: Receiver-Driven Congestion Control Using Explicit Notification for Data CenterabstractIn recent years, receiver-driven transport protocols have been proposed to use proactive congestion control to meet the stringent latency requirements of large-scale applications in data center. However, the receiver-driven proposals face the challenges brought by network dynamic. First, when the bursty flows start, the aggressive and blind line-rate transmission in the first RTT easily leads to persistent queue backlog. Second, when some flows finish transmissions, the remaining ones cannot increase their sending rates to seize the available bandwidth. To address these problems, this article presents a new receiver-driven congestion control design, called REN, which uses the under- and over-utilization notifications from switch to handle the dynamic traffic. With the aid of explicit feedback, REN alleviates the traffic burstiness due to aggressive start, mitigates the conservativeness in utilizing available bandwidth, and still retains the receiver-driven feature to achieve ultra-low latency. We implement the prototype of REN using DPDK. The experimental results of real testbed and large-scale NS2 simulation show that REN effectively reduces the average flow completion time (AFCT) by up to 68% over the state-of-the-art receiver-driven transmission schemes. Jiawei Huang 0001, Jinbin Hu 0001, Weihe Li, Tao Zhang 0019, Jingling Liu, Jianxin Wang 0001, Tian He 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | A Receiver-Driven Transport Protocol With High Link Utilization Using Anti-ECN Marking in Data Center NetworksabstractExisting reactive or proactive congestion control protocols are hard to simultaneously achieve ultra-low latency and high link utilization across all workloads ranging from delay-sensitive flows to bandwidth-hungry ones in datacenter networks. We present an Anti-ECN (Explicit Congestion Notification) Marking Receiver-driven Transport protocol called AMRT, which achieves both near-zero queueing delay and full link utilization by reasonably increasing sending rate in the case of under-utilization. Specifically, switches mark the ECN bit of data packets once detecting spare bandwidth. When receiving the anti-ECN marked packet, the receiver generates the corresponding marked grant to trigger more data packets. The testbed and simulation experiments show that AMRT effectively reduces the average flow completion time (AFCT) by up to 42% and improves the link utilization by up to 38% over the state-of-the-art receiver-driven transmission schemes. Jinbin Hu 0001, Jiawei Huang 0001, Jianxin Wang 0001, Tian He 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | Load Balancing in PFC-Enabled Datacenter NetworksabstractIn Priority Flow Control (PFC) enabled datacenter networks (DCNs), PFC is inevitably triggered due to bursty traffic even with end-to-end congestion control. Load balancing as a complementary mechanism to transport protocols can make rerouting decisions in time to alleviate PFC’s head-of-line (HoL) blocking problem. However, prior solutions designed for lossy DCNs do not work well in PFC-enabled networks, because the unreliable rerouting signals such as separate local queue length, round-trip time (RTT), explicit congestion notification (ECN), and link load cannot timely and correctly reflect PFC pausing. Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005 |
APNet | 1 |
| 2021 | RPO: Receiver-driven Transport Protocol Using Opportunistic Transmission in Data CenterabstractModern datacenter applications bring fundamental challenges to transport protocols as they simultaneously require low latency and high throughput. Recent receiver-driven trans-port protocols transmit only one data packet once receiving each grant or credit packet from the receiver to achieve ultra-low queueing delay and zero packet loss. However, the round-trip time variation and the highly dynamic background traffic significantly deteriorate the performance of receiver-driven transport protocols, resulting in under-utilized bandwidth. This paper designs a simple yet effective solution called RPO that retains the advantages of receiver-driven transmission while efficiently utilizing the available bandwidth. Specifically, RPO rationally uses low-priority opportunistic packets to ensure high network utilization without increasing the queueing delay of high-priority normal packets. In addition, since RPO only uses Explicit Congestion Notification (ECN) marking function and priority queues, RPO is ready to deploy on switches. We implement RPO in Linux hosts with DPDK. Our small-scale testbed experiments and large-scale simulations show that RPO significantly improves the network utilization by up to 35% under high workload over the state-of-the-art receiver-driven transmission schemes, without introducing additional queueing delay. Jinbin Hu 0001, Jiawei Huang 0001, Yijun Li 0002, Wenchao Jiang, Kai Chen 0005, Jianxin Wang 0001, Tian He 0001 |
ICNP | 1 |
| 2021 | Adjusting Switching Granularity of Load Balancing for Heterogeneous Datacenter TrafficabstractThe state-of-the-art datacenter load balancing designs commonly optimize bisection bandwidth with homogeneous switching granularity. Their performances surprisingly degrade under mixed traffic containing both short and long flows. Specifically, the short flows suffer from long-tailed delay, while the throughputs of long flows also degrade dramatically due to low link utilization and packet reordering. To solve these problems, we design a traffic-aware load balancing (TLB) scheme to adaptively adjust the switching granularity of long flows according to the load strength of short ones. Under the heavy load of short flows, the long flows use large switching granularity to help short ones obtain more opportunities in choosing short queues to complete quickly. On the contrary, the long flows reroute flexibly with small switching granularity to achieve high throughput. Furthermore, under extremely bursty scenario, we utilize the packet slicing scheme for long flows to release bandwidth for short ones. The experimental results of NS2 simulation and testbed implementation show that TLB significantly reduces the average flow completion time of short flows by 16%-67% over the state-of-the-art load balancers and achieves the high throughput for long flows. Moreover, for extreme bursty case, at the acceptable throughput degradation of long flows, TLB with packet slicing reduces the deadline missing ratio of bursty short flows by up to 80%. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lyu, Weihe Li, Wenchao Jiang, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | Achieving Fast Convergence and High Efficiency using Differential Explicit Feedback in Data CenterabstractSince most flows are short-lived in data center networks, fast convergence becomes very important to help the short flows effectively utilize high bandwidth. Though current feedback-based transport control protocols (TCPs) provide fast convergence via fine-grained explicit congestion information from customized switches, they unavoidably incur large traffic overhead for widely existing small packets in data center applications, resulting in suboptimal network efficiency. To solve this issue, we propose a datacenter TCP based on differential feedbacks, called DECN, to achieve fast convergence without any traffic overhead. Specifically, DECN feeds rate difference between the target and current rate back to the source by using multiple consecutive packets. The experimental results of NS2 simulation and testbed implementation show that DECN achieves comparable fast convergence as XCP without incurring any extra feedback overhead. Compared with the state-of-the-art feedback-based TCPs, DECN reduces the flow completion time by up to 34.1% in typical data center applications. Jiawei Huang 0001, Sen Liu 0002, Jinbin Hu 0001, Jianxin Wang 0001 |
ICC | 4 |
| 2020 | AMRT: Anti-ECN Marking to Improve Utilization of Receiver-driven Transmission in Data CenterabstractCloud applications generate a variety of workloads ranging from delay-sensitive flows to bandwidth-hungry ones in data centers. Existing reactive or proactive congestion control protocols are hard to simultaneously achieve ultra-low latency and high link utilization across all workloads in data center networks. We present a new receiver-driven transport scheme using anti-ECN (Explicit Congestion Notification) marking to achieve both near-zero queueing delay and full link utilization by reasonably increasing sending rate in the case of under-utilization. Specifically, switches mark the ECN bit of data packets once detecting spare bandwidth. When receiving the anti-ECN marked packet, the receiver generates the corresponding marked grant to trigger more data packets. The experimental results of small-scale testbed implementation and large-scale NS2 simulation show that AMRT effectively reduces the average flow completion time (AFCT) by up to 40.8% and improves the link utilization by up to 36.8% under high workload over the state-of-the-art receiver-driven transmission schemes. Jinbin Hu 0001, Jiawei Huang 0001, Jianxin Wang 0001, Tian He 0001 |
ICPP | 1 |
| 2019 | DDT: Mitigating the Competitiveness Difference of Data Center TCPsabstractTo achieve better network performance, the cloud service providers are widely deploying the ECN-based transport protocols (i.e., DCTCP) in their data center networks (DCN). In multi-tenant environment, however, the newly introduced ECN-enabled TCP greatly impairs the performance of applications with out-dated and miscon figured TCP stacks. The reason is that the ECN-enabled datacenter switch fails to treat the mixed TCP traffic fairly, causing the distinguished performance gap between the ECN-enabled and ECN-disabled TCPs. This paper proposes DDT (Dual Dynamic Thresholds), an active queue management algorithm (AQM) that aims to achieve the flow-level fairness when the heterogeneous TCP traffic coexists. DDT monitors the switch queue in real time, and dynamically tunes the distance between ECN-marking and packet-dropping thresholds to mitigate the competitiveness difference between the ECN-enabled and ECN-disabled TCP. Our preliminary real implementations and testing results show that DDT elegantly fills the competitiveness gap of heterogeneous TCP traffic without disturbing their own control loops, while only introducing acceptable deployment overhead at the switch. Tao Zhang 0019, Jiawei Huang 0001, Shaojun Zou, Sen Liu 0002, Jinbin Hu 0001, Jingling Liu, Chang Ruan, Jianxin Wang 0001, Geyong Min |
APNet | 5 |
| 2019 | TLB: Traffic-aware Load Balancing with Adaptive Granularity in Data Center NetworksabstractModern datacenter topologies typically are multi-rooted trees consisting of multiple paths between any given pair of hosts. Recent load balancing designs focus on making full use of available parallel paths to provide high bisection bandwidth. However, they are agnostic to the mixed traffic generated by diverse applications in data centers and respectively use the same granularity in rerouting flows regardless of the flow type. Therefore, the short flows suffer the long-tailed queueing delay and reordering problems, while the throughputs of long flows are also degraded dramatically due to low link utilization and packet reordering under the non-adaptive granularity. To solve these problems, we design a traffic-aware load balancing (TLB) scheme to adopt different rerouting granularities for two kinds of flows. Specifically, TLB adaptively adjusts the switching granularity of long flows according to the load strength of short ones. Under the heavy load of short flows, the long flows use large switching granularity to help short ones obtain more opportunities in choosing short queues to complete quickly. When the load strength of short flows is low, the long flows switch paths more flexibly with small switching granularity to achieve high throughput. TLB is deployed at the switch, without any modifications on the end-hosts. The experimental results of NS2 simulations and Mininet implementation show that TLB significantly reduces the average flow completion time (AFCT) of short flows by ~15%-40% over the state-of-the-art load balancing schemes and achieves the high throughput for long flows. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lv, Weihe Li, Jianxin Wang 0001, Tian He 0001 |
ICPP | 1 |
| 2019 | CAPS: Coding-Based Adaptive Packet Spraying to Reduce Flow Completion Time in Data CenterabstractModern data-center applications generate a diverse mix of short and long flows with different performance requirements and weaknesses. The short flows are typically delay-sensitive but to suffer the head-of-line blocking and out-of-order problems. Recent solutions prioritize the short flows to meet their latency requirements, while damaging the throughput-sensitive long flows. To solve these problems, we design a Coding-based Adaptive Packet Spraying (CAPS) that effectively mitigates the negative impact of short and long flows on each other. To exploit the availability of multiple paths and avoid the head-of-line blocking, CAPS spreads the packets of short flows to all paths, while the long flows are limited to a few paths with Equal Cost Multi Path (ECMP). Meanwhile, to resolve the out-of-order problem with low overhead, CAPS encodes the short flows using forward error correction (FEC) technology and adjusts the coding redundancy according to the blocking probability. Moreover, since the coding efficiency decreases when the coding unit is too small or large, we demonstrate how to obtain the optimal size of coding unit. The coding layer is deployed between the TCP and IP layers, without any modifications on the existing TCP/IP protocols. The test results of NS2 simulation and small-scale testbed experiments show that CAPS significantly reduces the average flow completion time of short flows by ~30%-70% over the state-of-the-art multipath transmission schemes and achieves the high throughput for long flows with negligible traffic overhead. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lv, Yutao Zhou, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | CAPS: Coding-based Adaptive Packet Spraying to Reduce Flow Completion Time in Data CenterabstractModern data-center applications generate a diverse mix of short and long flows with different performance requirements and weaknesses. The short flows are typically delay-sensitive but to suffer the head-of-line blocking and out-of-order problems. Recent solutions prioritize the short flows to meet their latency requirements, while damaging the throughput-sensitive long flows. To solve these problems, we design a Coding-based Adaptive Packet Spraying (CAPS) that effectively mitigates the negative impact of short and long flows on each other. To exploit the availability of multiple paths and avoid the head-of-line blocking, CAPS spreads the packets of short flows to all paths, while the long flows are limited to a few paths with Equal Cost Multi Path (ECMP). Meanwhile, to resolve the out-of-order problem with low overhead, CAPS encodes the short flows using forward error correction (FEC) technology and adjusts the coding redundancy according to the blocking probability. The coding layer is deployed between the TCP and IP layers, without any modifications on the existing TCP/IP protocols. The experimental results of NS2 simulation and Mininet implementation show that CAPS significantly reduces the average flow completion time of short flows by ~30% -70% over the state-of-the-art multipath transmission schemes and achieves the high throughput for long flows with negligible traffic overhead. Jinbin Hu 0001, Jiawei Huang 0001, Wenjun Lv, Yutao Zhou, Jianxin Wang 0001, Tian He 0001 |
INFOCOM | 1 |