EDBT 2026 Demo / reviewers in the wild / expert
Zirui Wan
dblp:309/5644
· DBLP profile ↗
18ranked-venue papers
8as first author
18since 2021 · last 2026
0009-0000-1673-8104ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 17 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RiCE: Precise Remote In-Network Congestion Elimination in Inter-Datacenter RDMA Networks
Chengyuan Huang, Guangyu Zhao, Lu Lu 0016, Zirui Wan, Jiaqing Dong, Zhuo Tang, Guihai Chen, Chen Tian 0001 |
IWQoS | 6 |
| 2026 | HierCC: Taming Traffic Uncertainty in RDMA Data Centers With Hierarchical Congestion ControlabstractExisting congestion control schemes for RDMA resolve the dilemma of guaranteeing high throughput and ultra-low latency to some extent from a variety of perspectives. However, they are inefficient in addressing transient large queue build-up and under-utilized bandwidth caused by frequent traffic bursts. In this paper, we argue that traffic uncertainty is the fundamental challenge that limits these schemes from addressing the aforementioned dilemma. Inspired by the investigation that aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack to ease traffic uncertainty and further provides hierarchically control within the first-hop ToR and between racks. Specifically, the inter-rack rates of aggregate flows are controlled by a credit-based mechanism. Then the bandwidth obtained by the aggregated flow is allocated to the corresponding intra-rack individual flows promptly and accurately. We implement HierCC in a testbed that consists of DPDK-based end-hosts and P4-based Tofino switches. The performance of HierCC is evaluated by comprehensive testbed experiments and SystemC/NS3 simulations. Results indicate that, compared with state-of-the-art, HierCC can mitigate buffer usage by up to$10\times $and reduce the average and 99th percentile FCT by up to 84% and 80%, respectively. Zirui Wan, Jiao Zhang 0002, Xiaolong Zhong, Zixuan Guan, Haoyu Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 1 |
| 2025 | Mercury: A Dynamic Multi-path Packet Spraying Scheme for RDMA NetworksabstractDue to the low entropy traffic characteristics of LLM (Large Language Model) training, existing load balancing mechanisms such as Equal-Cost Multi-Path (ECMP) fail to fully utilize the redundant bandwidth between computing nodes in RDMA over Converged Ethernet (RoCE). Packet spraying mechanism has become a typical solution to the load balancing problem in RoCEs. However, it has a negative effect on congestion control mechanisms and suffers severe out-of-order problems.In this paper, we propose Mercury, an host-driven spraying scheme that synergizes congestion feedback and reordering control. Mercury selects paths by leveraging ECN, RTT, and reordering metrics, adjusts rates via multi-metric window. It also employs receiver-side buffers with priority-based dropping to mitigate out-of-order penalties. Evaluations in ns-3 under AllReduce/All-to-All traffic show Mercury reduces maximum flow completion time (Max FCT) by 40%-63% compared to ECMP-based DCQCN/TIMELY/HPCC. It also achieves at least 10%-20% improvement against switch-based spraying. Yuxiang Wang 0011, Jiao Zhang 0002, Zirui Wan, Leixin Cai, Shuo Wang 0006, Tao Huang 0005 |
ICCCN | 3 |
| 2025 | Weir: Delay-based RNIC Cache Control Software Middleware for Scalable RDMA Networks
Yongchen Pan, Jiao Zhang 0002, Zirui Wan, Baohong Lin, Junliang Wang, Huimin Luo |
INFOCOM | 4 |
| 2025 | ACC: Addressing Performance Limitations in Datacenters with Atomic Congestion Control
Zirui Wan, Jiao Zhang 0002, Tian Pan 0001, Pingping Lin, Tao Huang 0005 |
INFOCOM | 1 |
| 2025 | Achilles: an Enhanced Scheme for Reactive Transport in Datacenters
Zirui Wan, Jiao Zhang 0002, Haoyu Pan, Tao Huang 0005 |
IWQoS | 1 |
| 2025 | Torrent: Re-Architecting End-to-End Transmission for Cross-Datacenter RDMA NetworksabstractSustainability is becoming increasingly challenging in today's data centers with limited space, power and connectivity. Large cloud service providers interconnect geographically distributed datacenters for better scalability and availability. Applications running on cross-datacenter network impose great challenges in transport design. In this paper, we identify two inherent limitations of extending the existing transport technology, RDMA, and its Ethernet derivative, RoCE, to long-haul transmission. First, the on-chip resources of commodity RDMA NICs are insufficient for long-haul transmission. Second, applying existing traffic control schemes to inter-datacenter environment exhibits poor performance. Motivated by this, we propose Torrent, a switch-driven transport framework which partitions end-to-end control into three sub-control loops. To achieve the combined goals of fairness and high performance in cross-datacenter scenarios, Torrent employs fast acknowledgment and near-end congestion control on datacenter interconnection (DCI) switches. We implement Torrent prototypes on commodity programmable switches and evaluate it through real-world testbed experiments. Our results show that Torrent can achieve high link utilization over ultra-long distances and quickly converge congested flows to steady rates. Haoyu Pan, Zirui Wan, Jiao Zhang 0002, Tao Huang 0005 |
WCNC | 3 |
| 2025 | RHCC: Revisiting Intra-Host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with CPU-to-memory traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to revisit the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines CPU-to-memory traffic congestion avoidance with sub-RTT granularity and proactive RNIC traffic adjustment. RHCC ensures fast congestion avoidance and can work with different inter-host congestion control methods. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2$\times$and 1.4$\times$, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Yongchen Pan, Tao Huang 0005 |
IEEE Trans. Netw. | 1 |
| 2025 | Re-Architecting Traffic Control in Cross-Datacenter RDMA NetworksabstractThe network-intensive applications, like machine learning and cloud storage, are increasingly driving two critical trends:1)RDMA has been widely deployed to provide high-speed networks;2)applications are distributively deployed across multiple regional datacenters to satisfy demands for content providers and customers. To fully utilize the benefits of RDMA, we desire to extend it to support cross-datacenter networks. However, the long-haul transport suffers a considerably long control loop, and thus the hybrid of long-haul and intra-datacenter traffic can easily cause severe congestion. We revisit existing traffic control methods and find they are insufficient to resolve this hybrid traffic congestion. Generally, regional datacenters are connected using dedicated long-haul optical fiber and datacenter interconnection (DCI) switches. In this paper, we propose Approach Traffic Control (ATC), a novel solution focusing on two-side DCI-switches (i.e., the approach point for datacenters) to separately alleviate the hybrid traffic congestion in the local and distal datacenters, as a building block for host-driven control methods. This design principle helps ATC shorten the control loop to a single datacenter scale while aggregating congestion information of the whole datacenter range with minor deployment complexity. We implement ATC on P4-based switches and conduct evaluations using real-world testbeds and large-scale NS3 simulations. The results show that ATC ensures fast congestion avoidance and delivers significant performance. For example, ATC reduces the FCT of intra-datacenter and long-haul traffic by up to 88% and 52%, respectively. Zirui Wan, Jiao Zhang 0002, Yuzhen Su, Haoyu Pan, Mingxuan Yu, Tao Huang 0005 |
IEEE Trans. Netw. | 1 |
| 2024 | FCC : A Fast-Converging Low-Latency Congestion Control Algorithm for Datacenter RDMA NetworkabstractCongestion control plays a crucial role in ensuring the performance of data center networks. However, mainstream RDMA congestion control algorithms still face challenges such as slow congestion response and poor deployability. In this paper, we propose a novel fast-convergence congestion control algorithm, FCC, to address these shortcomings. FCC leverages Explicit Congestion Notification (ECN) and Round-Trip Time (RTT) signals, utilizing the gradient of RTT to enhance response speed and employing a Sigmoid curve for rate increment. Through experiments, we demonstrate that compared to existing state-of-the-art algorithms, FCC achieves superior performance in terms of convergence speed, fairness, and small flow latency metrics. Biyao Che, Yuxiang Wang 0011, Zirui Wan, Zixiao Wang 0003, Yuan Tian 0038, Jizhuang Zhao, Shuo Wang 0006, Jiao Zhang 0002 |
APNet | 3 |
| 2024 | Rethinking Intra-host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with intra-host traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to rethink the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines sub-RTT granularity intra-host traffic congestion avoidance and proactive RNIC traffic adjustment. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2 × and 1.4 ×, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Tao Huang 0005 |
APNet | 1 |
| 2024 | Crowd Modeling and Control Via Cooperative Adaptive FilteringabstractThis paper introduces a crowd modeling and motion control approach that employs diffusion adaptation within an adaptive network. In the network, nodes collaboratively address specific estimation problems while simultaneously moving as agents governed by certain motion control mechanisms. Our research delves into the behaviors of agents when they encounter spatial constraints. Within this framework, the agents pursue several objectives, such as target tracking, coherent motion, and obstacle evasion. Throughout their navigation, they demonstrate a nature of self-organization and self-adjustment that drives them to maintain certain social distances from each other and adaptively adjust their behaviors in response to environmental changes. Zirui Wan, Saeid Sanei |
ICASSP | 1 |
| 2024 | BiCC: Bilateral Congestion Control in Cross-datacenter RDMA NetworksabstractWith the development of network-intensive applications like machine learning and cloud storage, there are two growing trends: (i) RDMA has been widely deployed to enhance underlying high-speed networks; (ii) applications are deployed on geographically distributed datacenters to meet customer demands (e.g., low access latency to services or regular data backups). To fully utilize the benefits of RDMA, we desire to support long-haul RDMA transport for cross-datacenter applications. Different from common intra-datacenter communications, the hybrid of long-haul and intra-datacenter traffic complicates the congestion state, and the considerably long control loop makes it more severe. We revisit existing congestion control methods and find they are insufficient to address the hybrid traffic congestion.Note that regional datacenters are connected by dedicated long-haul optical fiber and datacenter interconnection (DCI) switches directly. In this paper, we propose Bilateral Congestion Control (BiCC), a novel solution relying on two-side DCI-switches to bilaterally alleviate the hybrid traffic congestion in the sender-side and receiver-side datacenter while serving as a building block for existing host-driven methods. BiCC can shorten the control loop to a single datacenter scale and aggregate congestion information across the whole datacenter. We implement BiCC on commodity P4-based switches and conduct evaluations using both testbed experiments and NS3 simulations. The extensive evaluation results show that BiCC ensures fast congestion avoidance. Thus, BiCC reduces the average FCT for intra-datacenter and inter-datacenter traffic by up to 53% and 51%, respectively, in large-scale simulations. Zirui Wan, Jiao Zhang 0002, Mingxuan Yu, Xinghua Zhao, Tao Huang 0005 |
INFOCOM | 1 |
| 2024 | PACC: A Proactive CNP Generation Scheme for Datacenter NetworksabstractThe rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a proactive and accurate switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability, convergence and key parameter settings of PACC. Then, we implement PACC in a testbed consisting of DPDK-based end-hosts and Tofino P4 switches. In our evaluation, PACC achieves better fairness, fast reaction, high throughput, and 6$\sim$69% lower FCT (Flow Completion Time) than DCQCN, TIMELY, HPCC and RoCC. Jiao Zhang 0002, Xiaolong Zhong, Mingxuan Yu, Haoyu Pan, Zixuan Guan, Biyao Che, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 9 |
| 2023 | RCC: Enabling Receiver-Driven RDMA Congestion Control With Congestion Divide-and-Conquer in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and in-network congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional in-network support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is$4{\sim }79\%$better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | PACC: Proactive and Accurate Congestion Feedback for RDMA Congestion ControlabstractThe rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability and key parameter settings of PACC. Then, we conduct both micro-benchmark and large-scale simulations to evaluate the performance of PACC. The results show that PACC achieves fairness, fast reaction, high throughput, and 6~69% lower FCT (Flow Completion Time) than DCQCN, TIMELY and HPCC. Xiaolong Zhong, Jiao Zhang 0002, Zixuan Guan, Zirui Wan |
INFOCOM | 5 |
| 2021 | HierCC: Hierarchical RDMA Congestion ControlabstractRDMA has been increasingly deployed in data centers to decrease latency and CPU utilization. However, existing RDMA congestion control schemes fail to address instantaneous large queue build-up or bandwidth under-utilization associated with frequent traffic bursty. In this paper, we argue that traffic uncertainty is the essential reason that constrains data center congestion control from simultaneously achieving high throughput and deterministic latency. Since aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack and hierarchically controls the rate of flows. The rate of aggregate flows between racks is controlled by a credit-based congestion control mechanism. Then the bandwidth obtained by an aggregate flow in a rack is allocated to the corresponding individual flows from that rack promptly and accurately. We evaluate HierCC using SystemC and large-scale NS3 simulations. Results indicate that HierCC can significantly mitigate buffer usage and reduce the 99th percentile FCT by up to 20% and 40% compared with HPCC and DCQCN under a realistic workload, respectively. Jiao Zhang 0002, Zixuan Guan, Zirui Wan, Yinben Xia, Tian Pan 0001, Tao Huang 0005, Dezhi Tang |
APNet | 4 |
| 2021 | Receiver-Driven RDMA Congestion Control by Differentiating Congestion Types in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and innetwork congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional innetwork support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is 4~79% better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Jiaming Shi, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
ICNP | 4 |