VLDB 2026 Research / reviewers in the wild / expert
Rui Zhuang
dblp:148/1297
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-2993-9316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center NetworksabstractHigh-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption. Junye Zhang, Zhigang Ji, Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng, Zixuan Guan |
SIGCOMM | 6 |
| 2026 | Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI ClustersabstractLarge-scale AI clusters impose higher requirements on network scalability, cost, and communication efficiency. The traditional Clos topology suffers from superlinear cost growth when scaling to over 100k GPUs, while the more cost-effective Dragonfly+ introduces "down-up" detours, deadlock risks, and complex routing design. This paper presents Dragonfly-Ultra, a scalable, low-cost network architecture for high-performance AI clusters. Dragonfly-Ultra can scale to over 260k GPUs with only 82% cost and 81% power consumption of a 3-layer Clos architecture. Dragonfly-Ultra optimizes inter-group connectivity to eliminate intra-group detours entirely. Beyond the topological benefits, Dragonfly-Ultra incorporates three key mechanisms to further improve network performance and optimize collective communication, including lightweight dual-waterline adaptive routing for fast congestion mitigation, virtual-link-based deadlock avoidance with lower hardware overhead, and uniform affinity-aware rank placement for balanced inter-group traffic across all phases. Simulation results on a 4k-node cluster show that, compared to Clos, Dragonfly-Ultra achieves up to 18.8% and 39.2% lower completion time for AllReduce and AlltoAll, respectively. Compared to Dragonfly+, the reductions are up to 27.9% and 62.1%, outperforming current mainstream topologies. Rui Zhuang, Junye Zhang, Kefei Liu, Weiqiang Cheng, Zixuan Guan, Shengnan Yue, Ruixue Wang, Tong Yang 0003 |
SIGCOMM | 1 |
| 2024 | ProactMP: A Proactive Multipath Transport Protocol for Low-Latency DatacentersabstractWith the development of datacenter networks (DCNs) towards high bandwidth and low latency, the demands of high-level datacenter applications are heading towards high performance and high reliability, which makes traffic congestion one of the most notable problems in DCNs and brings new challenges to transport protocols. Proactive transport protocols are gaining prevalence due to their ability to provide accurate feedback and precise end-to-end control, while multipath transmission is having a broader application space in the multi-path topology of large-scale DCNs. However, these advanced transport protocols aim to improve their performance by addressing some specific congestion problems, but fail to handle multiple congestion problems caused by incast, high workload and load imbalance. Their performance in terms of flow completion time (FCT), delay, robustness, and balance still has room for further improvement. In this paper, we propose ProactMP, a novel proactive multipath transport protocol for further improvement of datacenter communications. ProactMP utilizes the rich resources of parallel paths in modern DCN and spreads the load across available network paths to improve network efficiency. ProactMP deploys a credit-based bandwidth allocation strategy to achieve low delay and zero packet loss, and overcommits receiver downlinks to ensure high link utilization. We have implemented ProactMP in the Linux system. Our testbed experiments show that ProactMP outperforms the TCP variants, MPTCP variants and a leading proactive transport protocol in FCT, link utilization, fairness and latency. Rui Zhuang, Jiangping Han, Kaiping Xue, Jian Li 0031, Qibin Sun, Jun Lu 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | Early Marking for Controllable Maximum Queue Length in Data Center NetworksabstractIn data center networks (DCNs), numerous congestion control schemes utilize explicit congestion notification (ECN) to achieve low average queue delay. Such schemes generally mark packets based on the current queue length exceeding a marking threshold. However, due to the delay of ECN feedback, the queue length may further increase before the congestion notification is delivered to senders, which may lead to uncontrollable maximum queue length when bursts occur. In this paper, we propose an early ECN marking scheme based on prediction, E-ECN, to control the maximum queue length in DCNs. E-ECN uses predicted queue length rather than the current to indicate congestion with an advance time which offsets the hysteresis of ECN. We theoretically and experimentally demonstrate that early marking does not impact the throughput with appropriate selection of the advance time, and we provide guidelines for the selection in DCNs. Our simulation results show that E-ECN achieves shorter average queue delay and controllable maximum queue length in general with a bandwidth utilization guarantee. E-ECN greatly reduces queue overflow and improves the robustness of DCNs. Jiangping Han, Rui Zhuang, Kaiping Xue, Qibin Sun, Jun Lu 0001 |
ICCCN | 3 |
| 2023 | Achieving Flexible and Lightweight Multipath Congestion Control Through Online LearningabstractThe upgrade of network devices to be equipped with multiple network interfaces makes it possible to improve network throughput performance through multipath transmission protocols, especially multipath TCP (MPTCP). However, so far the mostly used MPTCP protocols have a common limitation, namely the rigid and conservative method. They have been designed with little consideration of the fact that real networks are dynamic and the network status changes frequently, thus leading to the poor performance of current MPTCP in many realistic scenarios. In this paper, we propose a lightweight multipath congestion control algorithm based on online learning, named MP-OL. MP-OL models congestion control as a multi-armed bandit problem, and adjusts the sending rate of each subflow flexibly and adaptively through online learning. Therefore, MP-OL possesses the capability of suiting various network scenarios, and can achieve fairness and high performance in dynamic network environment. It can also flexibly switch between online learning and traditional method, which reduces the computational complexity while ensuring the learning efficiency, thus making MP-OL easy to deploy and use. As the experimental results demonstrated, compared with the leading MPTCP variants, MP-OL achieves significant improvements in fairness and link utilization, and shows better resilience to non-congestion loss and better adaptability to unstable network conditions. In real networks, MP-OL also obtains better throughput performance. Rui Zhuang, Jiangping Han, Kaiping Xue, Jian Li 0031, David S. L. Wei, Ruidong Li 0001, Qibin Sun, Jun Lu 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2023 | EdAR: An Experience-Driven Multipath Scheduler for Seamless Handoff in Mobile NetworksabstractMultipath TCP (MPTCP) improves the bandwidth utilization in wireless network scenarios, since it can simultaneously utilize multiple interfaces for data transmission. However, with the fast growth of mobile devices and applications, link interruptions caused by handoffs still lead to drastic performance degradation in such scenarios. Typically, a series of packet losses on part of the links will block the transmission of the entire connection when handoff occurs. This paper proposes an Experience-driven Adaptive Redundant packet scheduler (EdAR) for MPTCP, aiming at achieving seamless handoffs in mobile networks. EdAR enables flexibly scheduling redundant packets with an experience-driven learning-based approach in the face of drastic network environment changes for multipath performance enhancement. To enable accurate learning and prediction, both the network environment and the best course of actions are jointly learned via a Deep Reinforcement Learning (DRL) agent, which we design with a hybrid structure to deal with the complexity of system states. Furthermore, both offline and online learning are utilized to allow the agent to adapt to different and changing network environments. Evaluation results show that EdAR outperforms the state-of-the-art MPTCP schedulers in most network scenarios. Specifically in mobile networks with frequent handoffs, EdAR brings$2\times $improvement in terms of the overall goodput. Jiangping Han, Kaiping Xue, Jian Li 0031, Rui Zhuang, Ruidong Li 0001, Ruozhou Yu, Guoliang Xue, Qibin Sun |
IEEE Trans. Wirel. Commun. | 4 |
| 2021 | Low Priority Congestion Control for Multipath TCPabstractMany applications are bandwidth consuming but may tolerate longer flow completion times. Multipath protocols, such as multipath TCP (MPTCP), can offer bandwidth aggregation and resilience to link failures for such applications, and low priority congestion control (LPCC) mechanisms can make these applications yield to other time-sensitive ones. Properly combining the above two can improve the overall user experience. However, the existing LPCC mechanisms are not adequate for MPTCP. They do not take into account the characteristics of multiple network paths, and cannot ensure fairness among the same priority flows. Therefore, we propose a multipath LPCC mechanism, i.e., Dynamic Coupled Low Extra Delay Background Transport, named DC-LEDBAT. Our scheme is designed based on a standardized LPCC mechanism LEDBAT. To avoid unfairness among the same priority flows, DC-LEDBAT trades little throughput for precisely measuring the minimum delay. Moreover, to be friendly to single-path LEDBAT, our scheme leverages the correlation of the queuing delay to detect whether multiple paths go through a shared bottleneck. Then, DC-LEDBAT couples the congestion window at shared bottlenecks to control the sending rate. We implement DC-LEDBAT in a Linux kernel and experimental results show that DC-LEDBAT can not only utilize the excess bandwidth of MPTCP but also ensure fairness among the same priority flows. Jian Li 0031, Yitao Xing, Rui Zhuang, Kaiping Xue |
GLOBECOM | 5 |
| 2021 | MP-VR: An MPTCP-Based Adaptive Streaming Framework for 360-degree Virtual Reality Videosabstract360-degree virtual reality videos greatly improve the video experience by providing users with a more immersive and interactive environment than standard streaming video. However, 360-degree videos suffer from bandwidth limits. Existing bandwidth-efficient solutions mainly focus on spatially cutting 360-degree video into tiles, and only provide video content in the Field-of-View (FoV) of users with high quality to reduce bandwidth consumption. Although existing tile-based schemes can reduce the bandwidth consumption, the bandwidth and transmission delay provided by a single-path TCP may still not meet the high requirements of 360-degree videos. Multipath TCP (MPTCP) allows a TCP connection to operate across multiple paths simultaneously and becomes highly attractive to support the mobile devices with various radio interfaces to aggregate multipath bandwidth and improve the throughput. In this paper, by taking the advantage of MPTCP, we propose an MPTCP-based adaptive streaming framework for 360-degree Virtual Reality videos, named MP-VR. MP-VR dynamically selects the appropriate tile bitrate according to the bandwidth and transmission delay of different subflows. Then it schedules the video segments to subflows to improve QoE of users. We conduct experiments on a testbed in our lab and simulations on NS-3. Evaluation results show that MP-VR outperforms existing tile-based strategies when network fluctuations or errors in FoV predictions occur. Wenjia Wei, Jiangping Han, Yitao Xing, Kaiping Xue, Jianqing Liu, Rui Zhuang |
ICC | 6 |
| 2021 | wCompound: Enhancing Performance of Multipath Transmission in High-speed and Long Distance NetworksabstractAs the user demand for data transmission over high-speed and long distance (hereafter abbreviated as HSLD) networks increases significantly, multipath TCP (MPTCP) shows a great potential to further improve the utilization of HSLD network resources than traditional TCP, and provides better quality of service (QoS). It has been reported that TCP causes serious waste of bandwidth in HSLD networks, while MPTCP can transmit data by using multiple network paths simultaneously between two distant hosts, thus provides better resource utilization, higher throughput and smoother failure recovery for applications. However, the existing multipath congestion control algorithms cannot perfectly meet the efficiency requirements of HSLD network, since they mainly emphasize fairness rather than other critical indicators of QoS such as throughput, but still encounter fairness issues when coexist with various TCP variants. To solve these problems, we develop weighted Compound (wCompound), a loss-and-delay-based compound multipath congestion control algorithm which is originated from Compound TCP, and is applicable to HSLD networks. Different from the traditional methods of setting an empirical value as the threshold, wCompound innovatively adopts a dynamic threshold and have the flexibility to adjust the sending window of each subflow based on current network state, so as to effectively couple all subflows and fully utilize the network capacity. Moreover, with the cooperation of delay-based and loss-based methods, wCompound also ensures good fairness to different types of TCP variants. We implement wCompound in the Linux kernel, then carry out sufficient experiments on our testbed. The results show that wCompound achieves higher utilization of network resources and can always maintain an appropriate throughput no matter competing with loss-based or delay-based network traffic. Rui Zhuang, Yitao Xing, Wenjia Wei, Kaiping Xue |
IWQoS | 1 |
| 2014 | Compiling Abstract Specifications into Concrete Systems - Bringing Order to the Cloud
Ian Unruh, Alexandru G. Bardas, Rui Zhuang, Xinming Ou, Scott A. DeLoach |
LISA | 3 |