VLDB 2026 Research / reviewers in the wild / expert
Jiao Zhang 0002
dblp:04/527-2
· DBLP profile ↗
97ranked-venue papers
21as first author
60since 2021 · last 2026
0000-0001-5614-3420ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 86 · 17 first-author · 56 since 2021Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Argus: Scalable and Deterministic Network Fault Localization for AI Training ClustersabstractNetwork failures in Artificial Intelligence (AI) training clusters can degrade entire jobs, making fast and accurate fault localization critical. Existing active probing systems suffer from two fundamental limitations: probabilistic path coverage that cannot guarantee complete link observability, and binary anomaly detection that fails to distinguish concurrent failures or localize gray failures. Equal-Cost Multi-Path (ECMP) routing is deterministic given the same 5-tuple, and ECMP configurations are accessible in operator-controlled clusters. We exploit this property to derive exact probe paths through offline hash computation without network measurement. Based on this approach, we design a hash-aware probing system that constructs a deterministic coverage matrix to select minimal probes guaranteeing complete link coverage. We introduce edge signatures to ensure fault distinguishability and a tiered diagnosis approach where lightweight iterative localization handles hard failures while sparse regression localizes gray failures. Preliminary evaluations on fat-tree topologies with up to 10,240 hosts show that our system achieves 100% link coverage with 39× fewer probes than R-Pingmesh, F1 score of 0.75–0.92 for multi-link failures, and 0.67 F1 for gray failures where existing methods fail entirely. Yuxiang Wang 0011, Jiao Zhang 0002, Xianyu Huang, Yubo Ruan, Yingjie Duan, Shoushou Ren, Xianjun He, Tao Huang 0005 |
APNet | 2 |
| 2026 | Lossless-SR: Towards Non-Disruptive Source Routing for Topology-Varying LEO Satellite Networks
Tian Pan 0001, Guohao Ruan, Zijia Xu, Yuehui Tan, Jiao Zhang 0002, Tao Huang 0005 |
ICC | 8 |
| 2026 | Symphony: Enhancing RDMA Connection Scalability through Sender-Receiver Coordination
Jiao Zhang 0002, Dexuan Liao, Xianyu Huang |
INFOCOM | 2 |
| 2026 | Cadence: Scalable RDMA Queue Pair Multiplexing through Fine-grained WQE Scheduling
Dexuan Liao, Jiao Zhang 0002, Kaitai Zhang |
INFOCOM | 2 |
| 2026 | Tlaloc: A Generic Multipath Load Balancing for RoCE
Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
INFOCOM | 2 |
| 2026 | CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
NSDI | 24 |
| 2026 | HyperEdge: An Edge CDN Infrastructure for Cost Efficient Video Streaming
Dehui Wei, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Jialin Li 0001 |
NSDI | 2 |
| 2026 | Euler: An Out-of-Order-Aware Load Balancing with Adaptive Granularity for AI Clusters
Jiafeng Jiang, Jiao Zhang 0002, Huimin Luo, Shuo Wang 0006, Tao Huang 0005 |
WCNC | 2 |
| 2026 | Mercury: Multipath Spraying for Joint Congestion and Reordering Control in RDMAabstractDue to the low entropy traffic characteristics of LLM (Large Language Model) training, existing load balancing mechanisms such as Equal-Cost Multi-Path (ECMP) fail to fully utilize the redundant bandwidth between computing nodes in RDMA over Converged Ethernet (RoCE). Packet spraying mechanism has become a typical solution to the load balancing problem in RoCEs. However, it has a negative effect on congestion control mechanisms and suffers severe out-of-order problems. In this paper, we propose Mercury, a host-driven spraying scheme that synergizes congestion feedback and reordering control. Mercury selects paths by leveraging ECN, RTT, and reordering metrics, adjusts rates via multi-metric window. It also employs receiver-side buffers with priority-based dropping to mitigate out-of-order penalties. Evaluations in ns-3 under AllReduce and All-to-All traffic show that Mercury consistently outperforms the ECMP-based baselines, including DCQCN, TIMELY, HPCC, SWIFT, and BOLT, with the largest reduction in Max FCT reaching 63%. Under multi-path load balancing, Mercury delivers the lowest Max FCT for large messages in AllReduce and for most message sizes in All-to-All. It outperforms STRACK and MP-RDMA by up to 28% and 35% in AllReduce, and by up to 25% and 30% in All-to-All. Yuxiang Wang 0011, Jiao Zhang 0002, Leixin Cai, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2026 | HierCC: Taming Traffic Uncertainty in RDMA Data Centers With Hierarchical Congestion ControlabstractExisting congestion control schemes for RDMA resolve the dilemma of guaranteeing high throughput and ultra-low latency to some extent from a variety of perspectives. However, they are inefficient in addressing transient large queue build-up and under-utilized bandwidth caused by frequent traffic bursts. In this paper, we argue that traffic uncertainty is the fundamental challenge that limits these schemes from addressing the aforementioned dilemma. Inspired by the investigation that aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack to ease traffic uncertainty and further provides hierarchically control within the first-hop ToR and between racks. Specifically, the inter-rack rates of aggregate flows are controlled by a credit-based mechanism. Then the bandwidth obtained by the aggregated flow is allocated to the corresponding intra-rack individual flows promptly and accurately. We implement HierCC in a testbed that consists of DPDK-based end-hosts and P4-based Tofino switches. The performance of HierCC is evaluated by comprehensive testbed experiments and SystemC/NS3 simulations. Results indicate that, compared with state-of-the-art, HierCC can mitigate buffer usage by up to$10\times $and reduce the average and 99th percentile FCT by up to 84% and 80%, respectively. Zirui Wan, Jiao Zhang 0002, Xiaolong Zhong, Zixuan Guan, Haoyu Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 2 |
| 2026 | Weir: Scalable RDMA With Delay-Based RNIC Cache Control Software Middleware for Data Center NetworksabstractRemote Direct Memory Access (RDMA) is widely used in distributed services in Data Center Networks (DCNs) due to its high performance. As DCNs expand in scale, RDMA faces scalability issues. The reason is that the high concurrency Queue Pairs (QPs) lead to cache misses on RDMA Network Interface Card (RNIC) and frequent evictions, and the behaviour of fetching the cache via PCIe leads to performance degradation of RDMA. In this paper, we model the behaviour of Work Queue Element (WQE) on RNIC as a producer-consumer model and investigate that the root cause of WQE cache misses is the mismatch between the production rate of the CPU and the consumption rate of the RNIC. We design Weir from the perspective of WQE cache control to avoid cache misses and improve throughput under high concurrent QPs. Weir determines the cache occupancy on the RNIC by monitoring the number of active QPs and the increase/decrease in the life cycle of WQEs, and calculates the production rate and pacing by credit. The implementation of Weir exhibits minimal CPU overhead. Evaluation results show that Weir can maintain 97Gbps throughput without degradation even with up to 16K concurrent QPs, and effectively reduces various observable cache misses by$5\times $to$10\times $compared to commercial RNICs. Additionally, experiments show that Weir has better connection scalability than XRC and DCT. Jiao Zhang 0002, Yongchen Pan, Dexuan Liao, Huimin Luo, Tao Huang 0005, Haipeng Yao |
IEEE Trans. Netw. | 1 |
| 2025 | Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Yang Song 0031, Tian Pan 0001, Zhigang Zong, Bengbeng Xue, Xionglie Wei, Yisong Qiao, Donglin Lai, Baohai Hu, Jin Ke 0005, Enge Song, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
APNet | 18 |
| 2025 | Accelerating Distributed Training on Parameter Server Architecture With Path-Aware MulticastabstractIt is observed that the bottleneck in distributed training has shifted from computation to communication due to contention in concurrent transmissions and substantial redundant traffic. In the Parameter Server (PS) architecture, the server aggregates gradients from multiple workers and then distributes updated model parameters back to the workers in a one-to-many manner. Currently, model parameters are distributed via unicast, sending multiple identical copies of the data, which leads to significant bandwidth waste. Although multicast can save bandwidth, current approaches have two main drawbacks: on one hand, many protocols require maintaining excessive multicast state inside the network; on the other hand, the lack of coordination among multiple multicast trees can still lead to path conflicts. In this work, we propose path-aware multicast, which includes innetwork multicast tree reservation and per-hop control multicast. Specifically, before each round of model parameter distribution, the server queries the network for a multicast tree that satisfies the bandwidth requirement. The calculated multicast tree is then returned with bandwidth reserved at its tree nodes. Next, model parameters are forwarded with hop-by-hop control along the multicast tree. After the multicast is completed, the reserved network resources are released. Our evaluation shows that in an$8 \times 8$spine-leaf topology, path-aware multicast improves link load balancing by 32.6 % compared to random multicast and accelerates model parameter distribution by up to nearly$N \times$compared to unicast, where$N$is the number of workers. Chuanying Yuan, Tian Pan 0001, Guohao Ruan, Hao Li 0011, Yan Zou, Jiao Zhang 0002, Tao Huang 0005 |
ICC | 9 |
| 2025 | Achieving Adaptive Multi-Path Routing and Order-Preserving Time Slot Planning in TSNabstractWith the rise of autonomous driving, the performance requirements for In-vehicular networks are continuously increasing. Existing research leverages Frame Replication and Elimination for Reliability (FRER) and Time-Aware Shaper (TAS) mechanisms in Time-Sensitive Networking (TSN) to achieve deterministic transmission. FRER requires transmitting flows over multiple disjoint paths. However, FRER lacks redundancy degree selection strategies, and the delay differences between redundant paths lead to packet disorder, which increases network resource overhead and compromises traffic QoS. In this paper, we propose a Bandwidth-Aware with Frame Replication and Elimination for Reliability (BA-FRER) algorithm dynamically selects redundancy degree based on the current network bandwidth resources, delay, and reliability utility function. Additionally, we propose a Redundant-Aware Order Preservation (RAOP) algorithm configures time slots for each flow based on the TAS mechanism to align the delays of redundant paths. The evaluation results show that the BA-FRER algorithm improves the utilization of network bandwidth resources, the flow access rate, and the reliability, while the RAOP algorithm reduces the probability of packet disorder. Yanke Li, Shuo Wang 0006, Guoyu Peng, Guizhen Li, Jiao Zhang 0002, Tao Huang 0005 |
ICCCN | 6 |
| 2025 | Mercury: A Dynamic Multi-path Packet Spraying Scheme for RDMA NetworksabstractDue to the low entropy traffic characteristics of LLM (Large Language Model) training, existing load balancing mechanisms such as Equal-Cost Multi-Path (ECMP) fail to fully utilize the redundant bandwidth between computing nodes in RDMA over Converged Ethernet (RoCE). Packet spraying mechanism has become a typical solution to the load balancing problem in RoCEs. However, it has a negative effect on congestion control mechanisms and suffers severe out-of-order problems.In this paper, we propose Mercury, an host-driven spraying scheme that synergizes congestion feedback and reordering control. Mercury selects paths by leveraging ECN, RTT, and reordering metrics, adjusts rates via multi-metric window. It also employs receiver-side buffers with priority-based dropping to mitigate out-of-order penalties. Evaluations in ns-3 under AllReduce/All-to-All traffic show Mercury reduces maximum flow completion time (Max FCT) by 40%-63% compared to ECMP-based DCQCN/TIMELY/HPCC. It also achieves at least 10%-20% improvement against switch-based spraying. Yuxiang Wang 0011, Jiao Zhang 0002, Zirui Wan, Leixin Cai, Shuo Wang 0006, Tao Huang 0005 |
ICCCN | 2 |
| 2025 | Valve: Scalable RDMA with Gap-based Cache Control Middleware for Data Center NetworksabstractRemote Direct Memory Access (RDMA) has become a cornerstone technology in Data Center Networks (DCNs). However, DCNs have expanded substantially, leading to severe connection scalability issues for RDMA. The critical reason behind these issues stems from frequent cache misses on RDMA NICs (RNICs) when handling numerous Queue Pair (QP) connections. Cache misses require time-consuming retrievals from the host via PCIe, resulting in a degradation in RDMA performance. Existing software solutions primarily aim to alleviate QP Context (QPC) cache pressure, while hardware solutions incur prohibitive costs. In this paper, we identify Work Queue Element (WQE) cache, rather than QPC cache, as the fundamental bottleneck limiting connection scalability. Hence, we propose a software middleware, Valve, designed to mitigate cache misses through WQE cache control. Valve regulates WQE posting to control WQE cache by monitoring RNIC cache usage and adaptively adjusting the gap of WQE posting. Valve boasts ease of deployment, requiring the addition of approximately 1000 lines of code, and incurs low CPU overhead. Valve maintains peak performance of RNIC throughput regardless of the number of QPs and considerably reduces observable cache misses (such as ICM, MTT, and MPT cache misses) by 2.8× to 3.1× compared to XRC and DCT. Jiao Zhang 0002, Dexuan Liao, Yongchen Pan, Tao Huang 0005 |
ICNP | 2 |
| 2025 | Weir: Delay-based RNIC Cache Control Software Middleware for Scalable RDMA Networks
Yongchen Pan, Jiao Zhang 0002, Zirui Wan, Baohong Lin, Junliang Wang, Huimin Luo |
INFOCOM | 2 |
| 2025 | ACC: Addressing Performance Limitations in Datacenters with Atomic Congestion Control
Zirui Wan, Jiao Zhang 0002, Tian Pan 0001, Pingping Lin, Tao Huang 0005 |
INFOCOM | 2 |
| 2025 | HELDR: Packet Loss Detection and Retransmission for Live Streaming Hyper-Edge NetworkabstractLive streaming platforms like Douyin have developed the Live Streaming Hyper-Edge Delivery Network (LSHEDN) to reduce bandwidth cost. In LS-HEDN, the Content Delivery Network (CDN) splits the live streaming into multiple substreams by randomly assigning each frame to them. Hyperedge devices like set-top boxes with cheap and idle bandwidth resources forward a substream from CDN to multiple users. A protocol based on User Datagram Protocol (UDP) is adopted between devices and users, with users detecting packet loss and requesting retransmissions via Negative Acknowledgment (NACK). Given the demand for lower latency and the inherent fluctuations in public network, existing receiver-side packet loss detection and retransmission methods fall short in achieving both timeliness and accuracy simultaneously. This is manifested as frequent rebuffering and excessive redundancy. Notably, when head-of-line blocking(HOL blocking) occurs in the upstream link of the device, these issues become even more pronounced. To address this, we propose Hyper-Edge Loss Detection and Retransmission (HELDR) algorithm. It features a loss detection algorithm tailored to the transmission characteristics in LSHEDN, which improves detection accuracy. Its immediate retransmission mechanism and the backup devices retransmission mechanism enhance timeliness. Large-scale online A/B tests results show that HELDR reduces the average rebuffering rate by 41.2%, reduces the average redundancy rate by 15.7%. Peisheng Guo, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Tao Huang 0005, Ruili Fang, Zhenpeng Zhu, Dehui Wei |
IWQoS | 2 |
| 2025 | Achilles: an Enhanced Scheme for Reactive Transport in Datacenters
Zirui Wan, Jiao Zhang 0002, Haoyu Pan, Tao Huang 0005 |
IWQoS | 2 |
| 2025 | Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event NotificationabstractLayer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation. Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu |
SIGCOMM | 12 |
| 2025 | ByteTracker: An Agentless and Real-time Path-aware Network Probing SystemabstractAs the number of data center servers grows into the millions and due to the demand for more accurate, rapid and powerful network fault detection and location, the existing Pingmesh-centric monitoring and diagnostic system is not efficient enough. In this paper, we propose ByteTracker, the first agentless probing and diagnostic system for large-scale data center networks. It does not need to deploy probe processes or make any configurations on end hosts, and all probes are launched by a small number of centralized Probers. ByteTracker achieves accurate, real-time probe path tracking with packet mirroring on switches. By reducing end-host probe noise, precisely identifying network timeout probes, accurately tracking probe paths, and marking the failed switch with multiple network timeout probes, ByteTracker can locate network failures with nearly 100% accuracy. We have deployed ByteTracker in all of our data centers for over half a year. During deployment, ByteTracker can detect almost all network anomalies and locate them within 5 seconds with 100% accuracy. Shixian Guo, Kefei Liu 0004, Yulin Lai, Yangyang Bai, Jianghang Ning, Yongbin Dong, Sisi Wen, Jiale Feng, Chengcai Yao, Zhuo Jiang, Jiao Zhang 0002, Tao Huang 0005 |
SIGCOMM | 22 |
| 2025 | Torrent: Re-Architecting End-to-End Transmission for Cross-Datacenter RDMA NetworksabstractSustainability is becoming increasingly challenging in today's data centers with limited space, power and connectivity. Large cloud service providers interconnect geographically distributed datacenters for better scalability and availability. Applications running on cross-datacenter network impose great challenges in transport design. In this paper, we identify two inherent limitations of extending the existing transport technology, RDMA, and its Ethernet derivative, RoCE, to long-haul transmission. First, the on-chip resources of commodity RDMA NICs are insufficient for long-haul transmission. Second, applying existing traffic control schemes to inter-datacenter environment exhibits poor performance. Motivated by this, we propose Torrent, a switch-driven transport framework which partitions end-to-end control into three sub-control loops. To achieve the combined goals of fairness and high performance in cross-datacenter scenarios, Torrent employs fast acknowledgment and near-end congestion control on datacenter interconnection (DCI) switches. We implement Torrent prototypes on commodity programmable switches and evaluate it through real-world testbed experiments. Our results show that Torrent can achieve high link utilization over ultra-long distances and quickly converge congested flows to steady rates. Haoyu Pan, Zirui Wan, Jiao Zhang 0002, Tao Huang 0005 |
WCNC | 4 |
| 2025 | SeqBalance: Congestion-Aware Load Balancing With No Reordering in Data Center NetworksabstractWith the rapid development of the Internet of Things (IoT), an increasing amount of sensor data generated by IoT applications has been transferred to data center networks for storage and data analysis. Remote Direct Memory Access (RDMA) is widely used in data center networks because of its high performance. However, due to the characteristics of RDMA’s retransmission strategy, current load balancing schemes for data center networks are unsuitable for RDMA. In this paper, we propose SeqBalance, a load balancing framework designed for RDMA. SeqBalance implements fine-grained load balancing for RDMA through a reasonable design and does not cause reordering problems. SeqBalance detects link congestion at the switch by sensing ECN signals and link utilization, and guides routing decisions accordingly. SeqBalance’s designs are all based on existing commercial RNICs and commercial programmable switches, so they are compatible with existing data center networks. We have implemented SeqBalance Shaper for fine-grained sub-flow splitting in Mellanox CX-6 RNIC and implemented routing decisions in Intel Tofino P4 programmable switch. The results of hardware testbed experiments and large-scale simulations show that compared with existing load balancing schemes, SeqBalance improves 24.7% and 15.9% on average FCT and 99th-percentile FCT. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Internet Things J. | 2 |
| 2025 | QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNsabstractVideo content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions. Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | RoCELet: Host-Based Flowlet Load Balancing for RoCEabstractRemote Direct Memory Access (RDMA) is becoming a popular high-speed networking technology. It uses kernel bypass and zero copy to achieve high throughput and low latency with little CPU overhead. However, standard RoCE transmission uses Equal Cost Multipath (ECMP) for load balancing, which can result in lower transmission performance due to hash conflicts. Meanwhile, it has been verified that, unlike TCP, the unique retransmission mode and flow characteristics of RoCE make previous load balancing algorithms not well applied to RoCE. In this paper, we introduce RoCELet, a load balancing algorithm for RoCE. It achieves fine-grained RoCE load balancing by actively generating flowlets, effectively utilizing the rich end-to-end paths in the data center. We implement a prototype based on DPDK and evaluate it through small-scale testbed experiments and large-scale simulations. Our results show that compared to state-of-the-art load balancing algorithms, RoCELet optimizes 48.2% and 16.4% in average FCT and$99^{th}$-ile FCT, respectively. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Jiafeng Jiang, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 2 |
| 2025 | RHCC: Revisiting Intra-Host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with CPU-to-memory traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to revisit the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines CPU-to-memory traffic congestion avoidance with sub-RTT granularity and proactive RNIC traffic adjustment. RHCC ensures fast congestion avoidance and can work with different inter-host congestion control methods. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2$\times$and 1.4$\times$, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Yongchen Pan, Tao Huang 0005 |
IEEE Trans. Netw. | 2 |
| 2025 | Re-Architecting Traffic Control in Cross-Datacenter RDMA NetworksabstractThe network-intensive applications, like machine learning and cloud storage, are increasingly driving two critical trends:1)RDMA has been widely deployed to provide high-speed networks;2)applications are distributively deployed across multiple regional datacenters to satisfy demands for content providers and customers. To fully utilize the benefits of RDMA, we desire to extend it to support cross-datacenter networks. However, the long-haul transport suffers a considerably long control loop, and thus the hybrid of long-haul and intra-datacenter traffic can easily cause severe congestion. We revisit existing traffic control methods and find they are insufficient to resolve this hybrid traffic congestion. Generally, regional datacenters are connected using dedicated long-haul optical fiber and datacenter interconnection (DCI) switches. In this paper, we propose Approach Traffic Control (ATC), a novel solution focusing on two-side DCI-switches (i.e., the approach point for datacenters) to separately alleviate the hybrid traffic congestion in the local and distal datacenters, as a building block for host-driven control methods. This design principle helps ATC shorten the control loop to a single datacenter scale while aggregating congestion information of the whole datacenter range with minor deployment complexity. We implement ATC on P4-based switches and conduct evaluations using real-world testbeds and large-scale NS3 simulations. The results show that ATC ensures fast congestion avoidance and delivers significant performance. For example, ATC reduces the FCT of intra-datacenter and long-haul traffic by up to 88% and 52%, respectively. Zirui Wan, Jiao Zhang 0002, Yuzhen Su, Haoyu Pan, Mingxuan Yu, Tao Huang 0005 |
IEEE Trans. Netw. | 2 |
| 2024 | Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE ClustersabstractRoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary. Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
APNet | 2 |
| 2024 | FCC : A Fast-Converging Low-Latency Congestion Control Algorithm for Datacenter RDMA NetworkabstractCongestion control plays a crucial role in ensuring the performance of data center networks. However, mainstream RDMA congestion control algorithms still face challenges such as slow congestion response and poor deployability. In this paper, we propose a novel fast-convergence congestion control algorithm, FCC, to address these shortcomings. FCC leverages Explicit Congestion Notification (ECN) and Round-Trip Time (RTT) signals, utilizing the gradient of RTT to enhance response speed and employing a Sigmoid curve for rate increment. Through experiments, we demonstrate that compared to existing state-of-the-art algorithms, FCC achieves superior performance in terms of convergence speed, fairness, and small flow latency metrics. Biyao Che, Yuxiang Wang 0011, Zirui Wan, Zixiao Wang 0003, Yuan Tian 0038, Jizhuang Zhao, Shuo Wang 0006, Jiao Zhang 0002 |
APNet | 9 |
| 2024 | Rethinking Intra-host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with intra-host traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to rethink the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines sub-RTT granularity intra-host traffic congestion avoidance and proactive RNIC traffic adjustment. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2 × and 1.4 ×, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Tao Huang 0005 |
APNet | 2 |
| 2024 | D-Router: Decoupled Content Routers with Remote Content StoreabstractNamed Data Networking (NDN) enables efficient content distribution through in-network caching. However, the additional states of network intermediary nodes make NDN forwarding more burdensome, and the unpredictability of cache hits during forwarding leads to uncertain content retrieval latency. To overcome performance bottlenecks at the router's data plane and enhance network determinism, we propose the decoupled content router with remote content store (D-Router). This novel architecture decouples the local content store (CS) from routers and introduces the remote CS device for pooling important content. When Interest packets arrive at a router whose CS is overloaded, we ensure determinism by forwarding them to the remote CS for processing if the requested content is cached there, preventing blocking before the local CS of routers and potential random cache hits along the forwarding path. The dual-path bypass forwarding is supported through the design of routers and a dual-path routing protocol. D-Router is compatible with traditional NDN. Experiments show notable enhancements in data plane performance, including a 30% reduction in round-trip time (RTT), a 25% increase in throughput, improved determinism, and reduced network jitter. Additionally, the decoupling of CS makes it easier for network administrators to deploy network upgrades. Tian Pan 0001, Chunyang Wu, Guohao Ruan, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 5 |
| 2024 | Accelerating Mega-Scale Satellite Network Simulation in NS-3 via MPI-based ParallelizationabstractDue to the high costs of low Earth orbit (LEO) satellite manufacturing and launch, as well as the complexity of in-orbit network protocol debugging, simulating and verifying satellite network protocols on the ground before satellite launch holds significant importance. Compared to the expensive emulation with one-to-one replication, simulation (e.g., using ns-3) can achieve discrete event processing at a relatively lower cost by extending the wall clock time. However, very few studies have used ns-3 for LEO satellite network simulation, facing challenges such as faithfully simulating the on/off state switching of inter-satellite links (ISLs) and achieving simulation performance scalability for high-density satellite constellations. In this work, we propose a system to accelerate mega-scale LEO satellite network simulation in ns-3 via MPI-based parallelization. Specifically, we simulate ISLs based on ns-3's P2P channels/P2P remote channels and achieve runtime link connection/disconnection by implementing stateful traffic dropping inside the network interface. Then, we conduct concurrent simulation with ns-3's parallel and distributed simulation capability and partition the satellite constellation into multiple simulation processes through a hierarchical clustering algorithm and automated scripts, considering satellite locality and inter-process workload balance. Our evaluation shows significant speed improvements via parallelization, e.g., a 373% speedup with 12 processes for LEO-192, and a 156% speedup with 3 processes for LEO-3072. Haibin Song, Tian Pan 0001, Guohao Ruan, Ying Wan 0001, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2024 | BiCC: Bilateral Congestion Control in Cross-datacenter RDMA NetworksabstractWith the development of network-intensive applications like machine learning and cloud storage, there are two growing trends: (i) RDMA has been widely deployed to enhance underlying high-speed networks; (ii) applications are deployed on geographically distributed datacenters to meet customer demands (e.g., low access latency to services or regular data backups). To fully utilize the benefits of RDMA, we desire to support long-haul RDMA transport for cross-datacenter applications. Different from common intra-datacenter communications, the hybrid of long-haul and intra-datacenter traffic complicates the congestion state, and the considerably long control loop makes it more severe. We revisit existing congestion control methods and find they are insufficient to address the hybrid traffic congestion.Note that regional datacenters are connected by dedicated long-haul optical fiber and datacenter interconnection (DCI) switches directly. In this paper, we propose Bilateral Congestion Control (BiCC), a novel solution relying on two-side DCI-switches to bilaterally alleviate the hybrid traffic congestion in the sender-side and receiver-side datacenter while serving as a building block for existing host-driven methods. BiCC can shorten the control loop to a single datacenter scale and aggregate congestion information across the whole datacenter. We implement BiCC on commodity P4-based switches and conduct evaluations using both testbed experiments and NS3 simulations. The extensive evaluation results show that BiCC ensures fast congestion avoidance. Thus, BiCC reduces the average FCT for intra-datacenter and inter-datacenter traffic by up to 53% and 51%, respectively, in large-scale simulations. Zirui Wan, Jiao Zhang 0002, Mingxuan Yu, Xinghua Zhao, Tao Huang 0005 |
INFOCOM | 2 |
| 2024 | SARO: Intelligent Data-Driven Routing Optimization for LEO Satellite NetworksabstractThe rapid development of satellite networks has precipitated an increasing demand for high bandwidth. However, compared to terrestrial networks, bandwidth resources in satellite networks are often more constrained. Hence, reasonable traffic scheduling is crucial. Because satellite networks present complex network structures, uneven traffic patterns, and strong dynamics, traditional traffic scheduling algorithms often fail in accurate analysis and modeling. Numerous intelligent routing schemes based on learning have been proposed and validated. However, due to the limitations of network modeling and generalization, it is difficult for them to quickly adapt to changing network conditions. In this paper, we propose SARO, an intelligent routing algorithm that integrates Deep Reinforcement Learning (DRL), Graph Neural Networks (GNN), and In-band Network Telemetry (INT). Compared to traditional deep learning algorithms, SARO not only overcomes the challenge of acquiring network features but also addresses the issue of limited generalization performance. Experiments show that regardless of changes in topology, SARO’s maximum link utilization is reduced by 7.6% to 15.6% compared to baseline algorithms, demonstrating SARO performs excellently in terms of load balancing and generalization. Jiao Zhang 0002, Tian Pan 0001, Tao Huang 0005 |
ISCC | 2 |
| 2024 | FTA-detector: Troubleshooting Gray Link Failures Based on Fault Tree AnalysisabstractDetecting link failures is critical to ensuring the operation of data center networks (DCNs). However, some gray link failures may go undetected by switches, leading to silent packet drops. In this paper, we propose FTA-detector, a gray link failure detection and localization approach leveraging Fault Tree Analysis (FTA), a technique previously applied in the field of reliability engineering. On the data plane, we collect fine-grained hop-by-hop information through In-band Network Telemetry (INT), detect the bidirectional connectivity of end-to-end paths through a novel aging mechanism, and implement fast reroute in response to gray link failures. On the control plane, we introduce a faulty link localization algorithm based on FTA to recommend the most likely faulty links. Specifically, we use Top K and progressive failure repair to discover and repair link faults as early as possible during failure inference, significantly reducing the overall computation complexity of sequential root cause analysis. For large-scale network topology, we propose a divide and conquer optimization scheme for scalability. To verify the efficiency of our system, we build a virtual network test platform with P4 switch software and Redis database. The test results show that FTA-detector can troubleshoot multi-point failures in DCNs in a very short time with high accuracy. Yan Zou, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Qingqiang Yi, Ying Wan 0001, Jiao Zhang 0002, Tao Huang 0005 |
NOMS | 7 |
| 2024 | LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu |
NSDI | 22 |
| 2024 | R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic SystemabstractRoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related. Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
SIGCOMM | 3 |
| 2024 | Blaze: Delay-Aware Cloud-Edge Collaborative Service Function Chain Deployment with Network CalculusabstractWith the rapid development of Internet of the Things (IoT) technology, IoT services have higher and higher requirements for latency. In the IoT environment, virtual network functions (VNFs) are deployed on general-purpose hardware and are sequentially connected to form service function chain (SFC) to provide network services for IoT devices. However, the high latency of the link between the cloud center and the edge nodes and the resource capacity limitation of the edge nodes pose challenges to the deployment of SFCs in IoT devices. In this paper, we study the cloud-edge collaborative SFC deployment problem. We applied the network calculus theory to the cloud-edge collaborative SFC deployment for the first time, aiming to provide the end-to-end delay guarantee for the deployed SFC. We model the SFC deployment problem as Mixed Integer Nonlinear Programming (MINLP). Then we propose a heuristic algorithm (Blaze) to solve this problem. Blaze is proven to complete the deployment of SFCs in polynomial time. Finally, the algorithm is evaluated by experimental simulation. The experimental results show that compared with the existing state-of-the-art corresponding algorithms, the proposed algorithm achieves better performance in terms of the number of VNFs deployed in the cloud, resource consumption of edge nodes, and SFC request acceptance rate. Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
WCNC | 2 |
| 2024 | Breaking the Inertial Thinking: Non-Blocking Multipath Congestion Control Based on the Single-Subflow Reinforcement Learning ModelabstractThe Multipath TCP (MPTCP) protocol has received more attention due to the increasing number of terminals with multiple network interfaces. To meet the higher network performance demand of terminal services, many researches leverage reinforcement learning (RL) for MPTCP congestion control (CC) algorithms to improve the performance of MPTCP. However, we observe two limitations of existing RL-based mechanisms that make them impractical: 1) Fail to break the restriction of the input and output dimensions of RL, making the mechanisms unadaptable to the varying number of subflows. 2) Frequent model decisions block packet transmission, leading to under-utilization of bandwidth. This paper breaks the inertial thinking By “inertial thinking” here, we are referring to the initial reaction of others when dealing with CC in MPTCP. Given the interdependence between MPTCP subflows, scholars have traditionally opted for coupled CC. However, we have challenged this conventional thinking by independently handling the CC of different subflows in a single MPTCP flow and ensuring fairness. to overcome the above limitations and proposes Maggey, a non-blocking CC mechanism that applies the single-subflow model to multipath transmission. To this end, Maggey employs loosely coupled design principles and a unique reward function to ensure the fairness of the algorithm. Additionally, Maggey introduces iterative training to ensure the accuracy of training of the single-subflow model. Furthermore, a mode transition framework is artfully designed to avoid blocking, preserving the flexibility of RL-based CCs. These two features enhance the practicability of Maggey and the paper analyze the stability of Maggey. We implement Maggey in the Linux kernel and evaluate the performance of Maggey through extensive emulation and live experiments. The evaluation results show that Maggey boosts 26% throughput over DRL-CC at high bandwidth and improves 2%-60% throughput over traditional algorithms under different network conditions. Besides, Maggey maintains fairness in different scenarios. Dehui Wei, Jiao Zhang 0002, Yuanjie Liu, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Diagnosing End-Host Network Bottlenecks in RDMA ServersabstractIn RDMA (Remote Direct Memory Access) networks, end-host networks, including intra-host networks and RNICs (RDMA NIC), were considered robust and have received little attention. However, as the RNIC line rate rapidly increases to multi-hundred gigabits, the intra-host network becomes a potential performance bottleneck for network applications. Intra-host network bottlenecks can result in degraded intra-host bandwidth and increased intra-host latency. In addition, RNIC network problems can result in connection failures and packet drops. Host network problems can severely degrade network performance. However, when host network problems occur, they can hardly be noticed due to the lack of a monitoring system. Furthermore, existing diagnostic mechanisms cannot efficiently diagnose host network problems. In this paper, we analyze the symptom of host network problems based on our long-term troubleshooting experience and propose Hostping, the first monitoring and diagnostic system dedicated to host networks. The core idea of Hostping is to conduct 1) loopback tests between RNICs and endpoints within the host to measure intra-host latency and bandwidth, and 2) mutual probing between RNICs on a host to measure RNIC connectivity. We have deployed Hostping on thousands of servers in our distributed machine learning system. Not only can Hostping detect and diagnose host network problems we already knew in minutes, but it also reveals eight problems we did not notice before. Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | PACC: A Proactive CNP Generation Scheme for Datacenter NetworksabstractThe rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a proactive and accurate switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability, convergence and key parameter settings of PACC. Then, we implement PACC in a testbed consisting of DPDK-based end-hosts and Tofino P4 switches. In our evaluation, PACC achieves better fairness, fast reaction, high throughput, and 6$\sim$69% lower FCT (Flow Completion Time) than DCQCN, TIMELY, HPCC and RoCC. Jiao Zhang 0002, Xiaolong Zhong, Mingxuan Yu, Haoyu Pan, Zixuan Guan, Biyao Che, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | INT-Label: Lightweight In-Band Network-Wide Telemetry via Distributed LabelingabstractIn-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for maintaining and troubleshooting data center networks. To achievenetwork-widetelemetry coverage, orchestration on top of the INT primitive is required. A straightforward solution would flood the network with INT probe packets for maximum measurement coverage, which leads to a huge bandwidth overhead. A refined solution leverages the SDN controller to collect the network topology information and carry out centralized probing path planning, which, however, is inefficient in reacting to topology changes. To tackle the above problems, we proposeINT-label, a lightweight In-band Network-Wide Telemetry architecture via the distributed labeling approach. INT-label periodically labels the sampled packets with device-internal states. It is cost-effective with a minor bandwidth overhead and able to seamlessly adapt to topology changes. In order to reduce the number of labeled packets, we introduce a times-based probabilistic labeling algorithm, which allows fewer packets to carry more INT information than the interval-based algorithm. In addition, to counteract the degradation of telemetry resolution due to loss of labeled packets, we design a feedback mechanism which can adaptively change the instant labeling frequency. We provide theoretical proof that INT-label can achieve network-wide telemetry. We analyze the impact of transmission delay on coverage rate and labeling times distribution under the INT-label architecture. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under the labeling frequency of 20 times per second. With the adaptive labeling enabled, even if 60% of the packets are lost, the coverage can still reach 92%. Enge Song, Tian Pan 0001, Haoyu Song 0001, Qiang Fu 0011, Yingjiang Liu, Chenhao Jia, Chuanying Yuan, Minglan Gao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2023 | Amphis: Rearchitecturing Congestion Control for Capturing Internet Application VarietyabstractTCP was designed to provide stream-oriented communication service for bulk data transfer applications (e.g., FTP and Email). With four-decade development, Internet applications have undergone significant changes, which now involve highly dynamic traffic pattern and message-oriented communication paradigm. However, the impact of this substantial evolution on congestion control (CC) has not been fully studied. Most of the network transports today still make the long-held assumption about application traffic, i.e., a byte stream with an unlimited data arrival rate. Tian Pan 0001, Shuihai Hu, Guangyu An, Xincai Fei, Fanzhao Wang, Yueke Chi, Minglan Gao, Hao Wu 0023, Jiao Zhang 0002, Tao Huang 0005, Jingbin Zhou |
APNet | 9 |
| 2023 | Performance Modeling and Analysis of Distributed Deep Neural Network Training with Parameter ServerabstractWith the growth of dataset size and the development of hardware accelerators, the application of deep neural networks (DNN) in various fields has made great breakthroughs. In order to improve the training speed of DNN, distributed training has been widely used. However, the imbalance between computation and communication makes distributed training difficult to achieve maximum efficiency. Therefore there is a need to detect the bottleneck state and verify the effect of some optimization schemes. Testing on a physical cluster incurs additional time and cost overhead. This paper builds a DNN-specific performance model that is used for bottleneck detection and tuning at a low cost. We build this model through detailed analysis and reasonable assumptions. We also focus on fine-grained modeling of scalability and network components, which are key factors affecting performance. Then we verify the performance model with an average error of 5% on testbed and emulator. Finally, we provide use cases of the performance model. Jiao Zhang 0002, Dehui Wei, Tian Pan 0001, Tao Huang 0005 |
GLOBECOM | 2 |
| 2023 | Leopard: A Pragmatic Learning-Based Multipath Congestion Control for Rapid Adaptation to New Network ConditionsabstractMultipath TCP is a multipath transport protocol deployed on end devices, and many learning-based multipath congestion control schemes have been proposed and verified. However, these schemes cannot adapt rapidly to new network conditions because of their convergence problems and generalization issues. To rapidly adapt to new network conditions, we propose Leopard, a learning-based multi-path congestion control frame-work that uses reinforcement learning to combine offline learning with online fine-tuning. The extensive experiments in emulated network conditions and the real world demonstrate that Leopard converges quickly and maintains consistent high performance in new network conditions, which avoids long retraining when the network environment changes. Leopard improves throughput by 13% compared with DRL-CC and reduces the convergence time by 20% compared with MPCC in new network conditions. Yuanjie Liu, Jiao Zhang 0002, Dehui Wei |
ICC | 2 |
| 2023 | Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA Servers
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005 |
NSDI | 3 |
| 2023 | SLIT: Achieving Fast Bandwidth Isolation Across Virtual MachinesabstractNetwork performance guarantee in the cloud is one of the hottest research topics recently. However, current approaches either lack scalability or fail to achieve high performance. To avoid the hard tradeoff between the scalability and performance, we adopt a hybrid approach to get the best of both worlds. In this paper, we follow the stateless design principle and propose SLIT, which provides fast VM-level network performance isolation and maintains core-stateless characteristics. Specifically, at the end host, we develop a novel network-adaptive labeling mechanism and it computes the correct scheduling priority for each hop. At the switch, packets are scheduled based on their labels and we develop a novel straggler detection algorithm to find the new incoming flow. Our evaluation results show that SLIT can approximate the optimal performance of Weighted Fair Queuing (WFQ) effectively. Chengyuan Huang, Jiao Zhang 0002, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | RCC: Enabling Receiver-Driven RDMA Congestion Control With Congestion Divide-and-Conquer in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and in-network congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional in-network support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is$4{\sim }79\%$better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Lightweight Route Flooding via Flooding Topology Pruning for LEO Satellite NetworksabstractWith the low latency and high coverage, the low earth orbit (LEO) satellite systems are attracting more and more venture capitals as well as research attentions. Due to their highly dynamic constellation topologies, routing protocols on the ground have to be tailored to efficiently adapt to the regular topology changes. However, for irregular topology changes caused by exceptional link failure/recovery, network-wide route flooding is still necessary for route convergence. But, this will cause significant traffic flooding redundancy due to the high density of constellation topologies. For larger-scale constellations, the redundancy issue will be exacerbated. To lessen the redundancy, this work proposes a lightweight route flooding mechanism by generating a sparse flooding topology that prunes the original full-mesh topology, and only flooding the route information on the sparse topology. By considering the maximum flooding hop as well as the robustness of the flooding topology, we design an algorithm to calculate the optimal topology instead of just applying the minimum spanning tree. The evaluation shows that, for the LEO-96 constellation, the new flooding topology has a 28.1% reduction in the inter-satellite links (ISLs) compared with the original topology, and the new flooding mechanism has a 37.52% reduction in the traffic flooded and a 10.03% reduction in the route convergence time compared with OSPF. Such improvements will be amplified on larger-scale constellations. Guohao Ruan, Tian Pan 0001, Chengcheng Lu, Zhengjie Luo, Houtian Wang, Jiao Zhang 0002, Yushi Shen, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2022 | MIMIC: SmartNIC-aided Flow Backpressure for CPU Overloading Protection in Multi-Tenant CloudsabstractIn multi-tenant clouds, off-the-shelf x86 boxes are widely deployed as middleboxes. With the rapid growth of cloud traffic and the migration to NFV deployment in recent years, CPU overloading at middleboxes becomes more of an issue. From our data centers, we observed that the CPU overloading was caused by heavy hitters. To address this issue, we propose MIMIC, a cloud-scale flow backpressure system, implemented onto our existing SmartNIC with FPGA acceleration. MIMIC rate-limits the selected heavy hitters through a new per-flow backpressure protocol and a new heavy-hitter detection system, to protect the other tenants. The detection system is based on hierarchical memory design, leveraging on-chip SRAM and off-chip DRAM, which can handle highly concurrent cloud traffic without the losses of flow information. We extend the design by adding a pre-filtering procedure for rapid detection. To avoid CPU being flooded by FPGA through frequent heavy-hitter reporting, due to their performance disparity, the CPU queries the FPGA on demand. The backpressure protocol is non-invasive to protect tenant privacy and allows controllable rate-limiting through the novel use of ECN and meter tables. The SmartNIC acts as a man in the middle to facilitate heavy-hitter detection and per-flow backpressuring. In a production setting, we observe that MIMIC can react quickly and bring down CPU load to the normal level within 10ms without packet losses. Enge Song, Nianbing Yu, Tian Pan 0001, Qiang Fu 0011, Xionglie Wei, Yisong Qiao, Jianyuan Lu, Yijian Dong, Mingxu Xie, Jinkui Mao, Zhengjie Luo, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Shunmin Zhu |
ICNP | 15 |
| 2022 | WebQMon.ai: Gateway-Based Web QoE Assessment Using Lightweight Neural Networks
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICSOC | 5 |
| 2022 | PACC: Proactive and Accurate Congestion Feedback for RDMA Congestion ControlabstractThe rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability and key parameter settings of PACC. Then, we conduct both micro-benchmark and large-scale simulations to evaluate the performance of PACC. The results show that PACC achieves fairness, fast reaction, high throughput, and 6~69% lower FCT (Flow Completion Time) than DCQCN, TIMELY and HPCC. Xiaolong Zhong, Jiao Zhang 0002, Zixuan Guan, Zirui Wan |
INFOCOM | 2 |
| 2021 | HierCC: Hierarchical RDMA Congestion ControlabstractRDMA has been increasingly deployed in data centers to decrease latency and CPU utilization. However, existing RDMA congestion control schemes fail to address instantaneous large queue build-up or bandwidth under-utilization associated with frequent traffic bursty. In this paper, we argue that traffic uncertainty is the essential reason that constrains data center congestion control from simultaneously achieving high throughput and deterministic latency. Since aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack and hierarchically controls the rate of flows. The rate of aggregate flows between racks is controlled by a credit-based congestion control mechanism. Then the bandwidth obtained by an aggregate flow in a rack is allocated to the corresponding individual flows from that rack promptly and accurately. We evaluate HierCC using SystemC and large-scale NS3 simulations. Results indicate that HierCC can significantly mitigate buffer usage and reduce the 99th percentile FCT by up to 20% and 40% compared with HPCC and DCQCN under a realistic workload, respectively. Jiao Zhang 0002, Zixuan Guan, Zirui Wan, Yinben Xia, Tian Pan 0001, Tao Huang 0005, Dezhi Tang |
APNet | 1 |
| 2021 | INT-probe: Lightweight In-band Network-Wide Telemetry with Stationary ProbesabstractVisibility is essential for operating and troubleshooting intricate networks. In-band Network Telemetry (INT) has been embedded in the latest merchant silicons to offer high-precision device and traffic state visibility. INT is actually an underlying technique and each INT instance covers only one monitoring path. The network-wide measurement coverage therefore requires a high-level orchestration to provision multiple INT paths. An optimal path planning is expected to produce a minimum number of paths with a minimum number of overlapping links. Eulerian trail has been used to solve the general problem. However, in production networks, the vantage points where one can deploy probes to start and terminate INT paths are constrained. In this work, we propose an optimal path planning algorithm, INT-probe, which achieves the network-wide telemetry coverage under the constraint of stationary probes. INT-probe formulates the constrained path planning into an extended multi-depot k-Chinese postman problem (MDCPP-set) and then reduces it to a solvable minimum weight perfect matching problem. We analyze algorithm's theoretical bound and the complexity. Extensive evaluation on both wide area networks and data center networks with different scales and topologies are conducted. We show INT-probe is efficient, high-performance, and practical for real-world deployment. For a large-scale data center networks with 1125 switches, INT-probe can generate 112 monitoring paths (reduced by 50.4 %) by allowing only 1.79% increase of the total path length, promptly resolving link failures within 744.71ms. Tian Pan 0001, Xingchen Lin, Haoyu Song 0001, Enge Song, Zizheng Bian, Hao Li 0011, Jiao Zhang 0002, Fuliang Li, Tao Huang 0005, Chenhao Jia, Bin Liu 0001 |
ICDCS | 7 |
| 2021 | Loom: Switch-based Cloud Load Balancer with Compressed StatesabstractLayer-4 load balancers play a critical role in large-scale data centers. Recently, load balancers implemented on programmable switches have attracted much attention since they overcome the inflexibility of dedicated load balancers and high latency of software load balancers. However, keeping per-connection state easily leads to storage exhaustion, especially under resource exhaustion attacks. Although several stateless load balancers are proposed to address this issue, the state management burden is offloaded to backend servers, causing high deployment and running costs. In this paper, a load balancer called Loom with compressed states is proposed for large-scale data centers. Firstly, we propose a novel classifier-based load balancer idea to avoid directly maintaining per-connection state. Then, a circulating Bloom filter structure is proposed that can efficiently classify connections as well as be implemented on existing programmable switches. Theoretical analysis shows that Loom can maintain 11 ~ 30x more concurrent connections than those directly storing the 5-tuple of connections. Loom is implemented in hardware P4 switches and experimental results indicate that 11 ~ 29x more concurrent connections can be maintained in Loom, which is close to the theoretical results. Besides, Loom is resistant to resource exhaustion attacks and reduces the percentage of broken connections by up to 57% with an SYN flood. Jiao Zhang 0002, Shubo Wen, Tian Pan 0001, Tao Huang 0005 |
ICNP | 1 |
| 2021 | Receiver-Driven RDMA Congestion Control by Differentiating Congestion Types in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and innetwork congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional innetwork support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is 4~79% better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Jiaming Shi, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
ICNP | 1 |
| 2021 | INT-label: Lightweight In-band Network-Wide Telemetry via Interval-based Distributed LabellingabstractThe In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for reliably maintaining and troubleshooting data center networks. For achieving network-wide telemetry, orchestration on top of the INT primitive is further required. One straightforward solution is to flood the INT probe packets into the network topology for maximum measurement coverage, which, however, leads to huge bandwidth overhead. A refined solution is to leverage the SDN controller to collect the topology and carry out centralized probing path planning, which, however, cannot seamlessly adapt to occasional topology changes. To tackle the above problems, in this work, we propose INT-label, a lightweight In-band Network-Wide Telemetry architecture via interval-based distributed labelling. INT-label periodically labels device-internal states onto sampled packets, which is cost-effective with minor bandwidth overhead and able to seamlessly adapt to topology changes. Furthermore, to avoid telemetry resolution degradation due to loss of labelled packets, we also design a feedback mechanism to adaptively change the instant label frequency. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under a label frequency of 20 times per second. With adaptive labelling enabled, the coverage can still reach 92% even if 60% of the packets are lost in the data plane. Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
INFOCOM | 5 |
| 2021 | Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesabstractThe cloud gateway is essential in the public cloud as the central hub of cloud traffic. We show that horizontal scaling of software gateways, once sustainable for years, is no longer future-proof facing the massive scale and rapid growth of today's cloud. The root cause is the stagnant performance of the CPU core, which is prone to be overloaded by heavy hitters as traffic growth goes far beyond Moore's law. To address this, we propose \emph{Sailfish}, a cloud-scale multi-tenant multi-service gateway accelerated by programmable switches. The new challenge is that large forwarding tables due to multi-tenancy cannot be fit into the limited on-chip memories. To this end, we devise a multi-pronged approach with (1) hardware/software co-design for table sharing, (2) horizontal table splitting among gateway clusters, (3) pipeline-aware table compression for a single node. Compared with the x86 gateway of a similar price, Sailfish reduces latency by 95% (2μs), improves throughput by more than 20x in bps (3.2Tbps) and 71x in pps (1.8Gpps) with packet length < 256B. Sailfish has been deployed in Alibaba Cloud for more than two years. It is the first P4-based cloud gateway in the industry, of which a single cluster carries dozens of Tbps traffic, withstanding peak-hour traffic in large online shopping festivals. Tian Pan 0001, Nianbing Yu, Chenhao Jia, Jianwen Pi, Yisong Qiao, Jianyuan Lu, Enge Song, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
SIGCOMM | 12 |
| 2021 | NB-Cache: Non-Blocking In-Network Caching for High-Performance Content RoutersabstractInformation-Centric Networking (ICN) provides scalable and efficient content distribution at the Internet scale due to in-network caching and native multicast. To support these features, a content router needs high performance at its data plane, which consists of three forwarding steps: checking the Content Store (CS), then the Pending Interest Table (PIT), and finally the Forwarding Information Base (FIB). In this work, we build an analytical model of the router and identify that CS is the actual bottleneck. Then, we propose a novel mechanism called “NB-Cache” to address CS’s performance issue from a network-wide point of view. In NB-Cache, when packets arrive at a router whose CS is fully loaded, instead of being blocked and waiting for the CS, these packets are forwarded to the next-hop router, whose CS may not be fully loaded. This approach essentially utilizes Content Stores of all the routers along the forwarding path in parallel rather than checking each CS sequentially. NB-Cache follows a design pattern of on-demand load balancing and can be formulated into a non-trivial N-queue bypass model. We use the Markov chain to establish its theoretical base and find an algorithm for automated transition rate matrix generation. Experiments show significant improvement of data plane performance: 70% reduction in round-trip time (RTT) and 130% increase in throughput. NB-Cache decouples the fast packet forwarding from the slower content retrieval thus substantially reducing CS’s heavy dependency on fast but expensive memory. Tian Pan 0001, Xingchen Lin, Enge Song, Jiao Zhang 0002, Hao Li 0011, Jianhui Lv, Tao Huang 0005, Bin Liu 0001, Beichuan Zhang 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2020 | PLB: Adaptive Partial Congestion-aware Load Balancing for Datacenter NetworksabstractIn order to accommodate ever-increasing new tenants and applications, datacenter networks (DCNs) require an efficient load balancing scheme to fully utilize their bisection bandwidth. Equal-cost MultiPath routing (ECMP) is a widely used load-balancing mechanism in the DCN. However, ECMP blindly hashes traffic to parallel paths and results in imbalance and collisions. Motivated by ECMP's shortcomings, some recent schemes provide more visibility into networks via active probing. They could be broadly classified as probing all the paths or a fixed number of paths (e.g., 3 paths) each probe interval. However, they all suffer from some limitations. Probing all paths introduces high probing overhead while probing a fixed number of paths is suboptimal when the network topology and traffic load change. To our best knowledge, none of the existing schemes adapt the number of paths being probed to the network conditions. Enlightened by the defects of previous work, we introduce PLB, an adaptive partial congestion-aware load-balancing mechanism. At its heart, PLB randomly probes partial paths each probe interval and the number of them changes according to the network topology and the traffic load. Besides, PLB splits flow into flowlets and makes careful routing/rerouting decisions for them. Through analysis, we formulate the correlations between the number of paths being probed and the network conditions. Furthermore, simulations with realistic workloads validate our conclusions and show that PLB reduces overall flow completion times compared to the state-of-the-art load balancing schemes both in symmetric and asymmetric topologies. Kefei Liu 0004, Jiao Zhang 0002, Dehui Wei, Tao Huang 0005 |
GLOBECOM | 2 |
| 2020 | INT-filter: Mitigating Data Collection Overhead for High-Resolution In-band Network TelemetryabstractIn-band Network Telemetry (INT) enables fine-grained network monitoring to ease the management of large-scale networks, which, however, relies on the real-time collection of a huge amount of telemetry data through the southbound interface. For example, the INT telemetry data upload rate of a 28-pod FatTree topology reaches 3Tbps under a probe frequency of 100 times/s, which is rather unacceptable since the controller-switch link bandwidth is limited. To mitigate the telemetry data collection overhead, in this work, we propose INT-filter, a novel measurement architecture that deploys the same prediction algorithm on both the data plane and the control plane to predict the traffic state in the near future instead of uploading all the telemetry data. Such prediction-based approach leverages the observation that there is considerable redundancy in the telemetry data sequence. In addition, we design an integration mechanism that conducts predictions using multiple methods simultaneously and uploads the predicted result from the least-error method to further decrease the upload volume. Extensive evaluation suggests that INT-filter can achieve at least 33.6% data collection decrease under a 10ms probe interval. With prediction integration, the upload reduction can further reach 58.5%. Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 5 |
| 2020 | Data-driven Routing Optimization based on Programmable Data PlaneabstractTo meet the growing demand for high bandwidth of Multimedia network, IP Network Providers spend millions of dollars overprovisioning bandwidth of their network. However, due to the lack of reasonable traffic scheduling, the over-provisioning network still has a severe issue of utilization imbalance. Traffic Engineering (TE) is proposed to solve this problem. Network measurement and routing optimization strategies are two key components of TE. Effective real-time network measurement provides the basis for the generation of route optimization strategies, which makes the network congestion-aware. Existing out-band network telemetry that transmits extra probes to measure network status has the problem of inaccurate measurement information in the network. Besides, the relationship between complex network status and routing optimization strategy is difficult to describe with an exact mathematical model. Therefore, we propose a novel TE approach, which is called DPRO. It combines In-band Network Telemetry based on programmable language P4 with Reinforcement Learning to minimize network max-link-utilization. Extensive experiments show that our approach significantly outperforms several widely-used baseline methods in terms of max-link-utilization. Qian Li 0006, Jiao Zhang 0002, Tian Pan 0001, Tao Huang 0005, Yunjie Liu 0001 |
ICCCN | 2 |
| 2020 | Hieff: Enabling Efficient VNF Clusters by Coordinating VNF Scaling and Flow SchedulingabstractA cluster of Virtual Network Functions (VNF) can serve massive fluctuating traffic by managing VNF instances and distributing flows. However, how to schedule the flows and manage VNF scaling efficiently in a VNF cluster is still an open question. Existing solutions such as hash based schemes encounter imbalance and passive flow remapping obstacles while flow table-based scheme suffers from high processing latency and flow entries overflow challenges. In this paper, we design and present Hieff, an efficient NFV system that coordinates VNF scaling and flow scheduling within a VNF cluster. The key idea of Hieff is to precisely manage the heavy flows with a flow table while simply allowing light flows be distributed by hash. Though this idea has been explored in previous work, we are the first to apply it with the VNF scaling process. We mathematically model the Hieff system and propose a heuristic algorithm to determine the optimized VNF scaling and flow scheduling strategies. We implement Hieff based on BESS and Click and use real-world tracing to evaluate the system. Results show that Hieff can handle co-existing massive flows efficiently with low latency while balancing the load of VNF instances at low cost. Jiao Zhang 0002, Tao Huang 0005 |
IPCCC | 2 |
| 2020 | Fast Switch-Based Load Balancer Considering Application Server StatesabstractLarge-scale services are generally hosted on multiple application servers to scale out in today's data centers. Load balancers distribute users' requests across these servers. Software load balancer and switch-based load balancer are two typical classes of load balancers. However, most of the existing mechanisms either exhibit high processing latency at load balancers or likely lead to unbalanced requests distribution without considering the disparity of the application servers. In this paper, we study how the disparity of application servers significantly impacts the response time of requests. A fast switch-based Load Balancer considering Application Server states (LBAS) then is proposed to minimize the processing latency at both load balancers and application servers. The data plane of LBAS is well designed to store millions of connections in limited storage capacity without violating per-connection consistency. Besides, a partial dynamic weighting algorithm based on the Ridge Regression theory is designed and implemented to decrease the processing latency at application servers. We implement LBAS using the P4 programming language and conduct a series of extensive experiments to evaluate the performance. The results demonstrate that the proposed LBAS mechanism significantly reduces the response time of requests compared with Uniform random, Static weight, and Spotlight in various scenarios. Jiao Zhang 0002, Shubo Wen, Jinsheng Zhang, Tian Pan 0001, Tao Huang 0005, Linquan Zhang, Yunjie Liu 0001, F. Richard Yu |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | INT-path: Towards Optimal Path Planning for In-band Network-Wide TelemetryabstractWith the ever-increasing complexity of networks, fine-grained network monitoring enables better network reliability and timely feedback control. The In-band Network Telemetry (INT) allows cost-effective network monitoring by encapsulating device-internal states into probe packets. However, INT only specifies an underlying device-level primitive while how to achieve network-wide traffic monitoring remains undefined. In this work, we propose INT-path, a network-wide telemetry framework, by decoupling the system into a routing mechanism and a routing path generation policy. Specifically, we embed source routing into INT probes to allow specifying the route the probe packet takes through the network. Above the mechanism, we develop an Euler trail-based path planning policy to generate non-overlapped INT paths that cover the entire network with a minimum path number. Besides, an exhaustive analysis of algorithm's run-time complexity is also provided. INT-path can “encode” the network-wide traffic status into a series of “bitmap images”, transforming network troubleshooting into pattern recognition problems. INT-path is very suitable for deployment in data center networks thanks to their symmetric network topologies. Tian Pan 0001, Enge Song, Zizheng Bian, Xingchen Lin, Xiaoyu Peng, Jiao Zhang 0002, Tao Huang 0005, Bin Liu 0001, Yunjie Liu 0001 |
INFOCOM | 6 |
| 2019 | RABA: Resource-Aware Backup Allocation For A Chain of Virtual Network FunctionsabstractNetwork Function Virtualization (NFV) turns a sequence of network functions on hardwares into a service chain of virtual network functions (VNFs) provisioned on virtual machines or containers. However, the chain of VNFs may suffer from interruption as long as one VNF fails due to software faults or hardware malfunctions. A common approach to ensuring high availability is to provide backup nodes for primary VNFs. However, existing work on allocating backup nodes have not considered the heterogeneous resource demands of different VNFs. In this paper, we formalize the resource-aware backup allocation problem, which aims to minimize the backup resource consumption while meeting the overall availability demand. To this end, we prove the NP-hardness of this problem and propose the RABA-CDDE algorithm based on differential evolution to solve it. Besides, to reduce the computation overhead of RABA-CDDE, a greedy algorithm is proposed. Our extensive evaluation shows that the proposed algorithms can reduce the resource consumption by about 15% and 35% respectively compared to the state-of-art solutions in dedicated and shared protection scenarios. Jiao Zhang 0002, Chunyi Peng 0001, Linquan Zhang, Tao Huang 0005, Yunjie Liu 0001 |
INFOCOM | 1 |
| 2019 | NB-cache: non-blocking in-network caching for high-speed content routersabstractInformation-Centric Networking (ICN) provides scalable and efficient content distribution at the Internet scale due to its in-network caching and native multicast capabilities. To support these features, a content router needs high performance at its data plane, which consists of three forwarding steps: checking the Content Store (CS), then the Pending Interest Table (PIT), and finally the Forwarding Information Base (FIB). While prior works focus on performance optimization of a single step, we build an analytical model of content router's entire data plane and identify that CS is the actual bottleneck in the pipeline. Compared with PIT and FIB, CS is more challenging because it has more data to read/write, may have more entries in its table to store and lookup, and needs to organize content objects to sustain frequent cache replacement. Then, we propose a novel mechanism called "NB-Cache" to address CS's performance issue from a network-wide point of view rather than a single router's. In NB-Cache, when packets arrive at a router whose CS is fully loaded, instead of being blocked and waiting for the CS, these packets are forwarded to the next-hop router, whose CS may not be fully loaded. This approach essentially utilizes Content Stores of all the routers along the forwarding path in parallel rather than checking each CS sequentially. Our experiments show significant improvement of data plane performance: 70% reduction in round-trip time (RTT) and 130% increase in throughput. Tian Pan 0001, Xingchen Lin, Jiao Zhang 0002, Hao Li 0011, Jianhui Lv, Tao Huang 0005, Bin Liu 0001, Beichuan Zhang 0001 |
IWQoS | 3 |
| 2019 | Future Internet: trends and challengesabstractTraditional networks face many challenges due to the diversity of applications, such as cloud computing, Internet of Things, and the industrial Internet. Future Internet needs to address these challenges to improve network scalability, security, mobility, and quality of service. In this work, we survey the recently proposed architectures and the emerging technologies that meet these new demands. Some cases for these architectures and technologies are also presented. We propose an integrated framework called the service customized network which combines the strength of current architectures, and discuss some of the open challenges and opportunities for future Internet. We hope that this work can help readers quickly understand the problems and challenges in the current research and serves as a guide and motivation for future network research. Jiao Zhang 0002, Tao Huang 0005, Shuo Wang 0006, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2019 | Service Function Chain Composition, Placement, and Assignment in Data CentersabstractWith the development of network function virtualization (NFV), service function chains (SFCs) are deployed via virtual network functions (VNFs). In general, the SFCs are served via composition and then deployed into data center infrastructures. However, most of the existing works neglect SFC composition. Furthermore, they consider that VNF instances are independently deployed for each SFC, which may underutilize the computational power of servers. We consider, for each required VNF in the chain, the operator can either place it on a new instance or assign it to an established instance if the residual resource of that instance is sufficient. Such a deployment scheme can leverage resources more efficiently and we define it as SFC placement and assignment. In this paper, we first combine SFC composition, placement and assignment together to enhance resource allocation. We present the system model and formulate the problem as 0-1 integer programming. We aim to improve the VNF instance utilization as well as reduce the link consumption. A heuristic approach called Jcap is developed to solve the problem in two stages. The simulations show that Jcap achieves competitive performance with the optimal results obtained from mathematical model. Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2018 | Analysing and improving convergence of quantized congestion notification in Data Center Ethernet
Ran Shu 0001, Fengyuan Ren, Jiao Zhang 0002, Tong Zhang 0018, Chuang Lin 0002 |
Comput. Networks | 3 |
| 2018 | Multi-Attributes-Based Coflow Scheduling Without Prior Knowledge
Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2017 | Low Latency Software Rate Limiters for Cloud NetworksabstractA lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency. Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022 |
APNet | 8 |
| 2017 | Leveraging multiple coflow attributes for information-agnostic coflow schedulingabstractRecently, designing information-agnostic coflow scheduling mechanisms attracts much attention since by leveraging priority queues, they could reduce coflow completion time in data-parallel clusters without a priori knowledge, such as flow size, coflow size. However, existing information-agnostic mechanisms generally schedule coflows only according to the sent data size of different coflows and ignore other useful coflow-level attributes like width, length and communication patterns. In this paper, we investigate that the coflow completion time could be further decreased by jointly leveraging multiple coflow-level attributes. Based on this investigation, we present a Multiple-attributes-based Coflow Scheduling (MCS) mechanism to reduce the coflow completion time. In MCS, a Shortest and Narrowest Coflow First (SNCF) algorithm is designed to separate coflows based on their widths and estimated lengths at the start of a coflow. During the transmission of coflows, one type of demotion thresholds employed in previous coflow scheduling mechanisms is too crude for various coflows. Therefore, we proposed a double-threshold scheme to adjust the priorities of narrow (small coflow width) and wide (large coflow width) coflows according to different thresholds. Trace-driven simulations with production workloads show that MCS outperforms the previous information-agnostic scheduler Aalo, and reduces the coflow completion time of small coflows. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
ICC | 2 |
| 2017 | Adaptively adjusting ECN marking thresholds for datacenter networksabstractECN thresholds have limited operational range and very strict scope. Lower thresholds exacerbate the queue underflow while higher thresholds increase the queueing delays. In this paper, an Adaptive ECN (A-ECN) marking scheme is proposed to enhance the performance of ECN. A-ECN can adaptively adjust ECN marking thresholds in different scenarios to achieve good generality. Therefore, network operators can directly deploy A-ECN in various environments regardless of underlying queue types and bandwidth. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
ICNP | 2 |
| 2017 | A clustering-based approach for Virtual Network Function Mapping and AssigningabstractNetwork Function Virtualization has attracted attention from both academia and industry as it can help the service provider to obtain agility and flexibility in network service deployment. In general, the enterprises require their flows to pass through a specific sequence of virtual network function (VNF) that varies from service to service. In addition, for each VNF required in the coming service demands, the operator can either launch a new instance for it or assign it to an established instance. This makes the network service deployment tasks even more complicated. In this paper, we first propose a method based on min-K-cut to cluster the VNFs. With clustering results as guidance, we determine whether to launch or reuse the instance to improve utilization rate of the VNF instance. Furthermore, for purpose of decreasing link bandwidth occupation, we aggregate the instances that are deployed with VNFs from the same cluster into the same server or rack. We evaluate our approach considering the average link bandwidth occupied by every accepted demand, the instance utilization rate and the total number of served demands. The simulation shows that our approach reduces link occupation effectively, and, meanwhile, guarantees the VNF instance utilization rate advantageously. Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IWQoS | 2 |
| 2017 | Skipping congestion-links for coflow schedulingabstractData transfer duration accounts for a great proportion of job completion time in big-data systems. To reduce the time spent on data transfer, some traffic scheduling mechanisms at coflow-level are proposed recently. Most of them abstract datacenter networks as an ideal non-blocking big-switch, and the bottleneck is located at egress or ingress ports of end-hosts instead of in networks. Thus, they mainly focus on how to allocate port capacities of end-hosts to jobs without considering innetwork congestion. However, link congestion frequently occurs in datacenter networks due to network oversubscription and load imbalance. When link congestion occurs, bottleneck locations will move from the ports of end-hosts to network links. In this paper, we design and implement SkipL, a congestionaware coflow scheduler which could detect congestion and schedules coflows at end-hosts to effectively reduce coflow completion time. In addition, to be easily deployed in cloud environments, SkipL does not require to control flow routes. SkipL prototype system is implemented in Linux. The results of experiments conducted in a real small testbed and simulations conducted in the flow-level simulator show that SkipL reduces the average Coflow Completion Time(CCT) compared to the per-flow fair sharing scheduling method and Varys. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IWQoS | 2 |
| 2017 | Flow distribution-aware load balancing for the datacenter
Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
Comput. Commun. | 2 |
| 2017 | FlowTrace: measuring round-trip time and tracing path in software-defined networking with low communication overheadabstractIn today’s networks, load balancing and priority queues in switches are used to support various quality-of-service (QoS) features and provide preferential treatment to certain types of traffic. Traditionally, network operators use ‘traceroute’ and ‘ping’ to troubleshoot load balancing and QoS problems. However, these tools are not supported by the common OpenFlow-based switches in software-defined networking (SDN). In addition, traceroute and ping have potential problems. Because load balancing mechanisms balance flows to different paths, it is impossible for these tools to send a single type of probe packet to find the forwarding paths of flows and measure latencies. Therefore, tracing flows’ real forwarding paths is needed before measuring their latencies, and path tracing and latency measurement should be jointly considered. To this end, FlowTrace is proposed to find arbitrary flow paths and measure flow latencies in OpenFlow networks. FlowTrace collects all flow entries and calculates flow paths according to the collected flow entries. However, polling flow entries from switches will induce high overhead in the control plane of SDN. Therefore, a passive flow table collecting method with zero control plane overhead is proposed to address this problem. After finding flows’ real forwarding paths, FlowTrace uses a new measurement method to measure the latencies of different flows. Results of experiments conducted in Mininet indicate that FlowTrace can correctly find flow paths and accurately measure the latencies of flows in different priority classes. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001, F. Richard Yu |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2016 | TFC: token flow control in data center networksabstractServices in modern data center networks pose growing performance demands. However, the widely existed special traffic patterns, such as micro-burst, highly concurrent flows, on-off pattern of flow transmission, exacerbate the performance of transport protocols. In this work, an clean-slate explicit transport control mechanism, called Token Flow Control (TFC), is proposed for data center networks to achieve high link utilization, ultra-low latency, fast convergence, and rare packets dropping. TFC uses tokens to represent the link bandwidth resource and define the concept of effective flows to stand for consumers. The total tokens will be explicitly allocated to each consumer every time slot. TFC excludes in-network buffer space from the flow pipeline and thus achieves zero-queueing. Besides, a packet delay function is added at switches to prevent packets dropping with highly concurrent flows. The performance of TFC is evaluated using both experiments on a small real testbed and large-scale simulations. The results show that TFC achieves high throughput, fast convergence, near zero-queuing and rare packets loss in various scenarios. Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Peng Cheng 0005 |
EuroSys | 1 |
| 2016 | Deadline-aware bandwidth sharing by allocating switch buffer in data center networksabstractMost of today's data center applications are sensitive to latency. In this paper, a Deadline-aware bandwidth Sharing mechanism by Allocating switch Buffer (DSAB) is proposed to satisfy the deadline requirements of flows. As the existing work SAB does, DSAB also leverages the feature that the buffer size of a switch port is usually much larger than the product of bandwidth and round trip delay in data center networks. The basic idea of DSAB is to allocate switch buffer to deadline-sensitive flows in priority and the remaining buffer is fairly allocated to background flows. However, this method will possibly lead to bandwidth wastage in topologies with multiple bottlenecks. Therefore, enlightened by the pre-authorization method in credit card systems, we propose a Pre-Authorization (PA) algorithm to address this problem. The PA algorithm allows switches take back the extra allocated bandwidth by adding a bandwidth confirmation phase after the bandwidth request phase. End hosts and switch functions of DSAB is implemented in Linux kernel and NetFPGA platform, respectively. The results of experiments conducted in a real small testbed indicate that DSAB can indeed satisfy the deadline requirements of deadline-sensitive flows by allocating switch buffer to them in priority. Jiao Zhang 0002 |
INFOCOM | 1 |
| 2016 | FDALB: Flow distribution aware load balancing for datacenter networksabstractWe present FDALB, a flow distribution aware load balancing mechanism aimed at reducing flow collisions and achieving high scalability. FDALB, like the most of centralized methods, uses a centralized controller to get the view of networks and congestion information. However, FDALB classifies flows into short flows and long flows. The paths of short flows and long flows are controlled by distributed switches and the centralized controller respectively. Thus, the controller handles only a small part of flows to achieve high scalability. To further reduce the controller's overhead, FDALB leverages end-hosts to tag long flows, thus switches can easily determine long flows by inspecting the tag. Besides, FDALB can adaptively adjust the threshold at each end-host to keep up with the flow distribution dynamics. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IWQoS | 2 |
| 2016 | Guaranteeing Delay of Live Virtual Machine Migration by Determining and Provisioning Appropriate BandwidthabstractThe proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. Most existing literature on live VM migration focus on designing migration mechanisms to shorten the two metrics or making a tradeoff between them. Few of them can be applied to applications with delay requirements, such as a VM backup process that needs to be done in a specific time. This will negatively impact the user experiences and reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration, and then propose a novel transport control mechanism to guarantee the computed bandwidth. The experimental results demonstrate that the bandwidth obtained from the proposed reciprocal-based model guarantees the expected total migration time and downtime, and the proposed transport control mechanism ensures that the live VM migration flow obtains the expected bandwidth even if there are background flows. Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Computers | 1 |
| 2015 | Congestion-aware adaptive forwarding in datacenter networks
Jiao Zhang 0002, Fengyuan Ren, Tao Huang 0005, Yunjie Liu 0001 |
Comput. Commun. | 1 |
| 2015 | Dynamic Routing for Data Integrity and Delay Differentiated Services in Wireless Sensor NetworksabstractApplications running on the same Wireless Sensor Network (WSN) platform usually have different Quality of Service (QoS) requirements. Two basic requirements are low delay and high data integrity. However, in most situations, these two requirements cannot be satisfied simultaneously. In this paper, based on the concept ofpotentialin physics, we propose IDDR, a multi-path dynamic routing algorithm, to resolve this conflict. By constructing a virtual hybrid potential field, IDDR separates packets of applications with different QoS requirements according to the weight assigned to each packet, and routes them towards the sink through different paths to improve the data fidelity for integrity-sensitive applications as well as reduce the end-to-end delay for delay-sensitive ones. Using the Lyapunov drift technique, we prove that IDDR is stable. Simulation results demonstrate that IDDR provides data integrity and delay differentiated services. Jiao Zhang 0002, Fengyuan Ren, Hongkun Yang, Chuang Lin 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2015 | Modeling and Solving TCP Incast Problem in Data Center NetworksabstractTCP Incast problem attracts much attention due to the catastrophic goodput drop. In this paper, a goodput model of the problem is built to understand why goodput collapse occurs and a solution to the problem based on the theoretical analysis is proposed. We found that the TCP Incast goodput deterioration is mainly caused by two types of timeouts, one happens at the tail of data blocks and dominates the goodput when the number of senders is small, while the other one at the head of data blocks and governs the goodput when the number of senders is large. The proposed model describes the relationship between these two types of timeouts and the Incast communication pattern, block size, bottleneck buffer size, and so on. The simulation results indicate that the model well characterizes the features of the TCP Incast problem. Enlightened by the analysis, a PRiority-based solution to the TCP INcast problem (PRIN) is proposed, which avoids timeouts at the head of blocks by reducing TCP send window and prevents timeouts at the tail of blocks by leveraging priority technology. The experimental results show that PRIN solves the TCP Incast problem. Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Delay guaranteed live migration of Virtual MachinesabstractThe proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration, which can be employed to facilitate load balancing, fault management and server maintenance etc. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. The existing literature on live VM migration mainly focus on designing migration mechanisms to shorten these two metrics or making a tradeoff between them. Few of them can be applied to the applications with delay requirements, such as, delay-sensitive web services or a VM backup process that needs to be done in a specific time. This will not only negatively impact the user experiences, but also reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration. We first propose a deterministic-based model as a simple example, then assume that the dirtying frequency of each page obeys the bernoulli distribution. At last, we analyze the statistic features of the typical workload running in a VM and build a reciprocal-based workload model, and theoretically give the required bandwidth value to satisfy the performance metrics of a live VM migration. The experimental results demonstrate that the bandwidth obtained from the reciprocal-based model can guarantee the expected total migration time and downtime. Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002 |
INFOCOM | 1 |
| 2014 | Analysing convergence of Quantized Congestion Notification in Data Center EthernetabstractEnhancing Ethernet as the unified data center fabric to concurrently handle the traffic of Local Area Network (LAN), Storage Area Network (SAN), and High Performance Computing (HPC) has attracted much attention. Congestion management is one critical enhancement to fill the performance gap between traditional Ethernet and the unified data center fabric. Currently, Quantized Congestion Notification (QCN) has been approved as the standard congestion management mechanism. However, lots of work pointed out that QCN suffers from the problem of unfairness among different flows. In this paper, we found that QCN could achieve fairness, merely the convergence time to fairness is quite long. Thus, we build a convergence time model to investigate the reasons of the slow convergence process of QCN. The model indicates that the convergence time of QCN can be decreased if RPs have the same rate increase probability or the rate increase step becomes larger at steady state. We validate the precise of our model by comparing with experimental data on the NetFPGA platform. The results show that it well characterizes the convergence time to fairness of QCN. Based on the proposed model, the impact of QCN parameters, network parameters, and QCN variants on the convergence time is analysed. Finally, enlightened by the analysis, we proposed a mechanism, called QCN-T, which replaces the Byte Counter and Timer at sources with a single modified Timer, to reduce the convergence time of QCN. Ran Shu 0001, Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002 |
IWQoS | 2 |
| 2014 | Sharing Bandwidth by Allocating Switch Buffer in Data Center NetworksabstractIn today's data centers, the round trip propagation delay is quite small. Therefore, switch buffer sizes are much larger than the Bandwidth Delay Product (BDP). Based on this observation, in this paper we introduce a new transport protocol which provides bandwidth Sharing by Allocating switch Buffer (SAB) for data centers. SAB sets the congestion windows for flows based on the buffer size of the switches along the path. On one hand, as long as the total buffer allocated to all the flows is larger than the BDP, the network bandwidth can be fully utilized. On the other hand, since SAB only allocates the buffer space to flows, the totally injected traffic will not exceed the network capacity. Thus, SAB rarely loses packets. SAB also reduces flow completion time by allowing flows to reach their fair share of bandwidth quickly. The results of a series of experiments and simulations demonstrate that SAB has the features of fast convergence and rare packet loss. It reduces the latency of short flows and solves theTCP Incast and TCP Outcast problems. Jiao Zhang 0002, Fengyuan Ren, Xin Yue, Ran Shu 0001, Chuang Lin 0002 |
IEEE J. Sel. Areas Commun. | 1 |
| 2013 | Taming TCP incast throughput collapse in data center networksabstractThe TCP incast problem attracts a lot of attention due to its wide existence in cloud services and catastrophic performance degradation. Some effort has been made to solve it. However, the industry is still struggling with it, such as Facebook. Based on the investigation that the TCP incast problem is mainly caused by the TimeOuts (TOs) occurring at the boundary of the stripe units, this paper presents a simple and effective TCP enhanced mechanism, called GIP (Guarantee Important Packets), for the applications with the TCP incast problem. The main idea is making TCP aware of the boundaries of the stripe units, and reducing the congestion window of each flow at the start of each stripe unit as well as redundantly transmitting the last packet of each stripe unit. GIP modifies TCP a little at the end hosts, thus it can be easily implemented. Also, it poses no impact on the other TCP-based applications. The results of both experiments on our testbed and simulations on the ns-2 platform demonstrate that TCP with GIP can avoid almost all of the TOs and achieve high goodput for applications with the incast communication pattern. Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002 |
ICNP | 1 |
| 2013 | Frequency Domain Packet Scheduling with Stability Analysis for 3GPP LTE UplinkabstractIn this paper, we investigate the Frequency Domain Packet Scheduling (FDPS) problem for 3GPP Long Term Evolution (LTE) Uplink (UL). Instead of studying a specific scheduling policy, we provide a unified approach to tackle this issue. First, we formalize a general LTE UL FDPS problem, which is suitable for various scheduling policies. Then, we prove that the problem is MAX SNP-hard, which implies that approximation algorithms with constant approximation ratios are the best that we can hope for. Therefore, we design two approximation algorithms, both of which have polynomial runtime. The first algorithm is based on a simple greedy method. The second one is based on the Local Ratio (L-R) technique and it can approximately solve the LTE UL FDPS problem with an approximation ratio of 2. To further analyze the stability of the 2-approximation L-R algorithm, we derive a specific FDPS problem, which incorporates the queue length and channel quality information. We utilize the Lyapunov Drift to prove the L-R algorithm is stable for any $((\omega_0, \epsilon_0))$-admissible LTE UL systems. The simulation results indicate good performance of the L-R scheduler. Fengyuan Ren, Yinsheng Xu, Hongkun Yang, Jiao Zhang 0002, Chuang Lin 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2013 | Attribute-Aware Data Aggregation Using Potential-Based Dynamic Routing in Wireless Sensor NetworksabstractThe resources especially energy in wireless sensor networks (WSNs) are quite limited. Since sensor nodes are usually much dense, data sampled by sensor nodes have much redundancy, data aggregation becomes an effective method to eliminate redundancy, minimize the number of transmission, and then to save energy. Many applications can be deployed in WSNs and various sensors are embedded in nodes, the packets generated by heterogenous sensors or different applications have different attributes. The packets from different applications cannot be aggregated. Otherwise, most data aggregation schemes employ static routing protocols, which cannot dynamically or intentionally forward packets according to network state or packet types. The spatial isolation caused by static routing protocol is unfavorable to data aggregation. To make data aggregation more efficient, in this paper, we introduce the concept of packet attribute, defined as the identifier of the data sampled by different kinds of sensors or applications, and then propose an attribute-aware data aggregation (ADA) scheme consisting of a packet-driven timing algorithm and a special dynamic routing protocol. Inspired by the concept of potential in physics and pheromone in ant colony, a potential-based dynamic routing is elaborated to support an ADA strategy. The performance evaluation results in series of scenarios verify that the ADA scheme can make the packets with the same attribute spatially convergent as much as possible and therefore improve the efficiency of data aggregation. Furthermore, the ADA scheme also offers other properties, such as scalable with respect to network size and adaptable for tracking mobile events. Fengyuan Ren, Jiao Zhang 0002, Yongwei Wu 0001, Tao He 0008, Canfeng Chen, Chuang Lin 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | Modeling and understanding TCP incast in data center networksabstractRecently, TCP incast problem attracts increasing attention since the receiver suffers drastic goodput drop when it simultaneously strips data over multiple servers. Lots of attempts have been made to address the problem through experiments and simulations. However, to the best of our knowledge, few solutions can solve it fundamentally at low cost. In this paper, a goodput model of TCP incast is built to understand why goodput collapse occurs. We conclude that TCP incast goodput deterioration is mainly caused by two types of timeouts, one happens at the tail of a data block and dominates the goodput when the number of senders is small, while the other one at the head of a data block and governs the goodput when the number of senders is large. The proposed model describes the causes of these two types of timeouts which are related to the incast communication pattern, block size, bottleneck buffer and so on. We validate the proposed model by comparing with simulation data, finding that it can well characterize the features of TCP incast. We also discuss the impact of most parameters on the goodput of TCP incast. Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002 |
INFOCOM | 1 |
| 2011 | EBRP: Energy-Balanced Routing Protocol for Data Gathering in Wireless Sensor NetworksabstractEnergy is an extremely critical resource for battery-powered wireless sensor networks (WSN), thus making energy-efficient protocol design a key challenging problem. Most of the existing energy-efficient routing protocols always forward packets along the minimum energy path to the sink to merely minimize energy consumption, which causes an unbalanced distribution of residual energy among sensor nodes, and eventually results in a network partition. In this paper, with the help of the concept of potential in physics, we design an Energy-Balanced Routing Protocol (EBRP) by constructing a mixed virtual potential field in terms of depth, energy density, and residual energy. The goal of this basic approach is to force packets to move toward the sink through the dense energy area so as to protect the nodes with relatively low residual energy. To address the routing loop problem emerging in this basic algorithm, enhanced mechanisms are proposed to detect and eliminate loops. The basic algorithm and loop elimination mechanism are first validated through extensive simulation experiments. Finally, the integrated performance of the full potential-based energy-balanced routing algorithm is evaluated through numerous simulations in a random deployed network running event-driven applications, the impact of the parameters on the performance is examined and guidelines for parameter settings are summarized. Our experimental results show that there are significant improvements in energy balance, network lifetime, coverage ratio, and throughput as compared to the commonly used energy-efficient routing algorithm. Fengyuan Ren, Jiao Zhang 0002, Tao He 0008, Chuang Lin 0002, Sajal K. Das 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | Effective Data Aggregation Supported by Dynamic Routing in Wireless Sensor NetworksabstractData aggregation is an main method to conserve energy in wireless sensor network (WSN). Prior work on data aggregation protocols are generally based on static routing schemes, such as tree-based, cluster-based or chain-based routing schemes. Although they can save energy to some extent, in dynamic scenarios where the source nodes are changing frequently, they will not only incur high overhead to continuously reconstruct the routing but also can not reduce the communication overhead effectively. Our work aims to design an effective data aggregation mechanism supported by dynamic routing (DASDR) which can adapt to different scenarios without incurring much overhead. Enlightened by the concept of potential field in the discipline of physics, the dynamic routing in DASDR is designed based on two potential fields: depth potential field which guarantees packets reaching the sink at last and queue potential field which makes packets more spatially convergent and thus data aggregation will be more efficient. Simulation results show that DASDR is more effective in energy savings as well as scales well with regard to the network size. Jiao Zhang 0002, Qian Wu 0001, Fengyuan Ren, Tao He 0008, Chuang Lin 0002 |
ICC | 1 |
| 2010 | Frequency-Domain Packet Scheduling for 3GPP LTE UplinkabstractIn this paper, we investigate the frequency-domain packet scheduling (FDPS) problem for 3GPP LTE Uplink (UL). Instead of studying a specific scheduling policy, we provide a unified approach to tackle this issue. First we formalize a general LTE UL FDPS problem which is suitable for various scheduling policies. Then we prove that the problem is MAX SNP-hard, which implies that approximation algorithms with constant approximation ratios are the best that we can hope for. Therefore we design two approximation algorithms, both of which have polynomial runtime. Subsequently, we analyze the two algorithms and find their approximation ratios. The first algorithm is easy to follow, since it is based on a simple greedy method. The second one is based on the local ratio technique and it can approximately solve the LTE UL FDPS problem with a approximation ratio of 2. Hongkun Yang, Fengyuan Ren, Chuang Lin 0002, Jiao Zhang 0002 |
INFOCOM | 4 |
| 2010 | Attribute-aware data aggregation using dynamic routing in wireless sensor networksabstractData aggregation has been widely recognized as an efficient method to reduce energy consumption in wireless sensor networks, which can support a wide range of applications such as monitoring temperature, humidity, level, speed etc. The data sampled by the same kind of sensors have much redundancy since the sensor nodes are usually quite dense in wireless sensor networks. To make data aggregation more efficient, the packets with the same attribute, defined as the identifier of different data sampled by different sensors such as temperature sensors, humidity sensors, etc., should be gathered together. However, to the best of our knowledge, present data aggregation mechanisms did not take packet attribute into consideration. In this paper, we take the lead in introducing packet attribute into data aggregation and propose an Attribute-aware Data Aggregation mechanism using Dynamic Routing (ADADR) which can make packets with the same attribute convergent as much as possible and therefore improve the efficiency of data aggregation. This goal cannot be achieved by present static routing schemes employed in most of data aggregation mechanisms since they construct routes before transmitting the sampled data and thus can not dynamically forward packets in response to the variation of packets at intermediate nodes. Hence, we present a potential-based dynamic routing scheme which employs the concept of potential in physics and pheromone in ant colony to achieve our goal. The results of simulations in series of scenarios show that ADADR indeed conserve energy by reducing the average number of transmissions each packet needs to reach the sink and is scalable with regard to the network size. Jiao Zhang 0002, Fengyuan Ren, Tao He 0008, Chuang Lin 0002 |
WOWMOM | 1 |