VLDB 2026 Research / reviewers in the wild / expert
Shuihai Hu
dblp:143/8245
· DBLP profile ↗
25ranked-venue papers
10as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 22 · 9 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniDMA: Scalable RDMA Transport over WAN
Jinyang Li 0009, Shuihai Hu, Zhenyu Li 0001, Gaogang Xie, Jingbin Zhou |
APNet | 6 |
| 2025 | FlexSpark: Robust and Efficient Multi-Device Collaborative Inference over Wireless Network
Yiyang Shao, Shuihai Hu, Xinle Du, Jingbin Zhou |
APNet | 3 |
| 2024 | Revisiting Congestion Control for WiFi NetworksabstractWiFi networks, widely utilized by wireless devices, have become increasingly complex and congested environments, leading to noticeable delays, jitter, and throughput degradation for end-to-end network flows in today’s Internet. Through detailed experimental observations, we identified the TCP victim problem in WiFi, where TCP erroneously detects congestion and interacts with WiFi routers, resulting in a significant throughput decrease for certain hosts. In this paper, we introduce Cupid, a novel congestion control algorithm that relies on receiver-side WiFi physical layer measurements. Cupid accurately assesses congestion by measuring parameters such as airtime utilization, concurrency, and rates, allowing for precise rate adjustments. Our results demonstrate that whether employed alone or alongside other congestion controls, Cupid effectively mitigates the TCP victim problem. Furthermore, when used independently, Cupid can also reduce latency. Xinle Du, Yiyang Shao, Wei Wang 0505, Shuihai Hu, Jingbin Zhou, Kun Tan 0003 |
APNet | 5 |
| 2024 | Accelerating Privacy-Preserving Machine Learning With GeniBatchabstractCross-silo privacy-preserving machine learning (PPML) adopt; Partial Homomorphic Encryption (PHE) for secure data combination and high-quality model training across multiple organizations (e.g., medical and financial). However, PHE introduces significant computation and communication overheads due to data inflation. Batch optimization is an encouraging direction to mitigate the problem by compressing multiple data into a single ciphertext. While promising, it is impractical for a large number of cross-silo PPML applications due to the limited vector operations support and severe data corruption. Junxue Zhang 0001, Xiaodian Cheng, Hong Zhang 0025, Yilun Jin, Shuihai Hu, Han Tian, Kai Chen 0005 |
EuroSys | 6 |
| 2024 | Panorama: Optimizing Internet-scale Users' Routes from End to End
Shuihai Hu |
USENIX ATC | 2 |
| 2024 | LiteFlow: Toward High-Performance Adaptive Neural Networks for Kernel DatapathabstractAdaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution. This paper presents LiteFlow, a hybrid solution to build high-performance adaptive NNs for kernel datapath. At its core, LiteFlow decouples the control path of adaptive NNs into: 1) a kernel-space fast path for efficient model inference; and 2) a userspace slow path for effective model tuning. We have implemented LiteFlow with Linux kernel datapath and evaluated it with three popular datapath functions including congestion control, flow scheduling, and load balancing. Compared to prior works, LiteFlow achieves 44.4% better goodput for congestion control, and improves the completion time for long flows by 33.7% and 56.7% for flow scheduling and load balancing, respectively. Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Amphis: Rearchitecturing Congestion Control for Capturing Internet Application VarietyabstractTCP was designed to provide stream-oriented communication service for bulk data transfer applications (e.g., FTP and Email). With four-decade development, Internet applications have undergone significant changes, which now involve highly dynamic traffic pattern and message-oriented communication paradigm. However, the impact of this substantial evolution on congestion control (CC) has not been fully studied. Most of the network transports today still make the long-held assumption about application traffic, i.e., a byte stream with an unlimited data arrival rate. Tian Pan 0001, Shuihai Hu, Guangyu An, Xincai Fei, Fanzhao Wang, Yueke Chi, Minglan Gao, Hao Wu 0023, Jiao Zhang 0002, Tao Huang 0005, Jingbin Zhou |
APNet | 2 |
| 2022 | ComboTE: Scalable Mixed-link based Traffic Engineering for Hybrid WANsabstractIn recent years, enterprises are increasingly moving from private WANs to hybrid WANs, for the purpose of cost saving, better scalability and improved user experience. To utilize network resource of both private and public networks, existing traffic engineering (TE) solutions for hybrid WANs perform traffic classification at the application level, i.e., sending traffic of high-priority applications to private WANs while letting other traffic go to the Internet. Unfortunately, such a strategy falls short of achieving optimal performance and has inherent limitation in making desired tradeoff between cost and performance. In this paper, we present comboTE, a mixed-link based TE framework to optimize both performance and cost over hybrid WAN. Specifically, comboTE designs a fine-grained TE strategy that considers link(s) from both types of networks at every hop when deciding the routing path for each traffic demand. To address the scalability issues, we leverage Lagrangian relaxation, a decomposition technique for solving large scale integer linear programs, combined with a novel augmented-graph-based approach to derive the most cost-efficient link compositions within a segment. Our comboTE achieves the near-optimal solution with a theoretically proven gap that is better than linear programming relaxation. The experimental results on realistic topologies and traffic matrices show that comboTE scales well and achieves solutions of high quality. Compared to prior TE algorithms, comboTE achieves a solution with tighter gap in limited time, and takes 2.2× less execution time without a time budget. Xincai Fei, Hao Wu 0006, Shuihai Hu |
ICNP | 4 |
| 2022 | LiteFlow: towards high-performance adaptive neural networks for kernel datapathabstractAdaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution. Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005 |
SIGCOMM | 4 |
| 2022 | Aeolus: A Building Block for Proactive Transport in Datacenter NetworksabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth isproactivelyallocated as “credits” to senders who then can send “scheduled packets” at a right rate to ensure high link utilization, low latency, and zero packet loss. Consequently, proactive solutions such as ExpressPass, NDP, Homa, etc., have been proposed recently. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT “pre-credit” phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on “pre-credit” packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. Aeolus is compatible with all existing proactive solutions and readily implementable with commodity switches. We have integrated Aeolus into ExpressPass, NDP and Homa, and shown, via both implementation and simulations, that the Aeolus-enhanced solutions deliver significant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by$20\times $, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Gaoxiong Zeng, Wei Bai 0001, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | Providing Bandwidth Guarantees, Work Conservation and Low Latency Simultaneously in the CloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity-a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | One More Config is Enough: Saving (DC)TCP for High-Speed Extremely Shallow-Buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1 Gbps to 40 Gbps or even 100 Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4 MB at 1 Gbps to 16 MB at 100 Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC,1a simple yet effective solution that requires only one more ECN configuration (i.e., shared buffer ECN/RED) at commodity switches. BCC operates upon real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; When it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing a small amount of throughput. BCC is readily deployable with existing commodity switches. We validate BCC's efficacy in a 100G testbed and evaluate its performance using extensive simulations. Our results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while only degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 2 |
| 2020 | RAT - Resilient Allreduce Tree for Distributed Machine LearningabstractParameter/gradient exchange plays an important role in large-scale distributed machine learning (DML). However, prior solutions such as parameter server (PS) or ring-allreduce (Ring) fall short since they are not resilient to issues or uncertainties like oversubscription, congestion or failures that may occur in datacenter networks (DCN). Xinchen Wan, Hong Zhang 0025, Hao Wang 0116, Shuihai Hu, Junxue Zhang 0001, Kai Chen 0005 |
APNet | 4 |
| 2020 | One More Config is Enough: Saving (DC)TCP for High-speed Extremely Shallow-buffered DatacentersabstractThe link speed in production datacenters is growing fast, from 1Gbps to 40Gbps or even 100Gbps. However, the buffer size of commodity switches increases slowly, e.g., from 4MB at 1Gbps to 16MB at 100Gbps, thus significantly outpaced by the link speed. In such extremely shallow-buffered networks, today's TCP/ECN solutions, such as DCTCP, suffer from either excessive packet loss or substantial throughput degradation.To this end, we present BCC1, a simple yet effective solution that requires just one more ECN config (i.e., shared buffer ECN/RED) over prior solutions. BCC operates based on real-time global shared buffer utilization. When available buffer space suffices, BCC delivers both high throughput and low packet loss rate as prior work; Once it gets insufficient, BCC automatically triggers the shared buffer ECN to prevent packet loss at the cost of sacrificing little throughput. BCC is readily deployable with existing commodity switches. We validate BCC's hardware feasibility in a small 100G testbed and evaluate its performance using large-scale simulations. Our results show that BCC maintains low packet loss rate while slightly degrading throughput when the available buffer becomes insufficient. For example, compared to current practice, BCC achieves up to 94.4% lower 99th percentile flow completion time (FCT) for small flows while degrading average FCT for large flows by up to 3%. Wei Bai 0001, Shuihai Hu, Kai Chen 0005, Kun Tan 0002, Yongqiang Xiong |
INFOCOM | 2 |
| 2020 | Aeolus: A Building Block for Proactive Transport in DatacentersabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth is proactively allocated as "credits" to senders who then can send "scheduled packets" at a right rate to ensure high link utilization, low latency, and zero packet loss. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT "pre-credit" phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on "pre-credit" packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. We have integrated Aeolus into ExpressPass[14], NDP[18] and Homa[29], and shown, through both implementation and simulations, that the Aeolus-enhanced solutions deliver signiicant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by 20x, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Wei Bai 0001, Gaoxiong Zeng, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
SIGCOMM | 1 |
| 2019 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote direct memory access over converged Ethernet deployments is vulnerable to deadlocks induced by priority flow control. Prior solutions for deadlock prevention either require significant changes to routing protocols or require excessive buffers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol and needs only modest buffers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock, and implement it efficiently on commodity hardware. Shuihai Hu, Yibo Zhu 0001, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | Augmenting Proactive Congestion Control with AeolusabstractRecently, proactive congestion control solutions have drawn great attention in the community. By explicitly scheduling data transmissions based on the availability of network bandwidth, proactive solutions offer a lossless, near-zero queueing network for serving network transfers. Despite the advantages, proactive solutions require an extra RTT to allocate the ideal sending rate for new arrival flows. To resolve this, current solutions let new flows blindly transmit unscheduled packets in the first RTT, and assign these packets with high priority in the network. The unscheduled packets, however, can cause serious network congestion, resulting in large queue buildups and excessive packet losses. Shuihai Hu, Wei Bai 0001, Baochen Qiao, Kai Chen 0005, Kun Tan 0002 |
APNet | 1 |
| 2018 | Enabling Work-Conserving Bandwidth Guarantees for Multi-Tenant Datacenters via Dynamic Tenant-Queue BindingabstractToday's cloud networks are shared among many tenants. Bandwidth guarantees and work conservation are two key properties to ensure predictable performance for tenant applications and high network utilization for providers. Despite significant efforts, very little prior work can really achieve both properties simultaneously even some of them claimed so. In this paper, we present QShare, a comprehensive in-network solution to achieve bandwidth guarantees and work conservation simultaneously. QShare leverages weighted fair queuing on commodity switches to slice network bandwidth for tenants, and solves the challenge of queue scarcity through balanced tenant placement and dynamic tenant-queue binding. We have implemented a QShare prototype and evaluated it extensively via both testbed experiments and simulations. Our results show that QShare ensures bandwidth guarantees while driving network utilization to over 91% even under unpredictable traffic demands. Zhuotao Liu, Kai Chen 0005, Shuihai Hu, Yih-Chun Hu, Yi Wang 0004, Gong Zhang 0001 |
INFOCOM | 4 |
| 2017 | Congestion Control for High-speed Extremely Shallow-buffered Datacenter NetworksabstractThe link speed in datacenters is growing fast, from 1Gbps to 100Gbps. However, the buffer size of commodity switches increases slowly, thus significantly outpaced by the link speed. In such extremely shallow-buffered datacenter networks, prior TCP/ECN solutions suffer from either excessive packet losses or significant throughput degradation. Motivated by this, we introduce BCC, a simple yet effective solution with only one more configuration (shared buffer ECN/RED) at commodity switches. BCC operates based on real-time shared buffer utilization. When the buffer is abundant, BCC delivers both high throughput and low packet loss rate. When it becomes scarce, BCC triggers shared buffer ECN/RED to prevent packet losses at the cost of sacrificing a small amount of throughput. Our preliminary results show that BCC maintains low packet loss rate persistently while only slightly degrading throughput when the buffer becomes insufficient. Compared to current practice, BCC achieves up to 94.4% lower 99th percentile completion time for small flows while only degrading large flows by up to 2.8%. Wei Bai 0001, Kai Chen 0005, Shuihai Hu, Kun Tan 0002, Yongqiang Xiong |
APNet | 3 |
| 2017 | Tagger: Practical PFC Deadlock Prevention in Data Center NetworksabstractRemote Direct Memory Access over Converged Ethernet (RoCE) deployments are vulnerable to deadlocks induced by Priority Flow Control (PFC). Prior solutions for deadlock prevention either require signi.cant changes to routing protocols, or require excessive bu.ers in the switches. In this paper, we propose Tagger, a scheme for deadlock prevention. It does not require any changes to the routing protocol, and needs only modest bu.ers. Tagger is based on the insight that given a set of expected lossless routes, a simple tagging scheme can be developed to ensure that no deadlock will occur under any failure conditions. Packets that do not travel on these lossless routes may be dropped under extreme conditions. We design such a scheme, prove that it prevents deadlock and implement it e.ciently on commodity hardware. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
CoNEXT | 1 |
| 2016 | Deadlocks in Datacenter Networks: Why Do They Form, and How to Avoid ThemabstractDriven by the need for ultra-low latency, high throughput and low CPU overhead, Remote Direct Memory Access (RDMA) is being deployed by many cloud providers. To deploy RDMA in Ethernet networks, Priority-based Flow Control (PFC) must be used. PFC, however, makes Ethernet networks prone to deadlocks. Prior work on deadlock avoidance has focused on {\em necessary} condition for deadlock formation, which leads to rather onerous and expensive solutions for deadlock avoidance. In this paper, we investigate {\em sufficient} conditions for deadlock formation, conjecturing that avoiding {\em sufficient} conditions might be less onerous. Shuihai Hu, Peng Cheng 0005, Chuanxiong Guo, Jitendra Padhye, Kai Chen 0005 |
HotNets | 1 |
| 2016 | Providing bandwidth guarantees, work conservation and low latency simultaneously in the cloudabstractToday's cloud is shared among multiple tenants running different applications, and a desirable multi-tenant datacenter network infrastructure should provide bandwidth guarantees for throughput-intensive applications, low latency for latency-sensitive short messages, as well as work conservation to fully utilize the network bandwidth. Despite significant efforts in recent years, none of them can achieve these three properties simultaneously. In this paper, we identify the key deficiency of prior solutions and use this insight to motivate our design of Trinity - a simple, practical yet effective solution that achieves bandwidth guarantees, work conservation and low latency simultaneously in the cloud. We implement Trinity using existing commodity hardwares and demonstrate its superior performance over prior solutions using testbed experiments. Shuihai Hu, Wei Bai 0001, Kai Chen 0005, Chen Tian 0001, Ying Zhang 0022 |
INFOCOM | 1 |
| 2016 | Explicit Path Control in Commodity Data Centers: Design and ApplicationsabstractMany data center network DCN applications require explicit routing path control over the underlying topologies. In this paper, we present XPath, a simple, practical and readily-deployable way to implement explicit path control, using existing commodity switches. At its core, XPath explicitly identifies an end-to-end path with a path ID and leverages a two-step compression algorithm to pre-install all the desired paths into IP TCAM tables of commodity switches. Our evaluation and implementation show that XPath scales to large DCNs and is readily-deployable. Furthermore, on our testbed, we integrate XPath into four applications to showcase its utility. Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Explicit Path Control in Commodity Data Centers: Design and Applications
Shuihai Hu, Kai Chen 0005, Wei Bai 0001, Chang Lan, Hao Wang 0022, Chuanxiong Guo |
NSDI | 1 |
| 2013 | Towards minimal-delay deadline-driven data center TCPabstractThis paper presents MCP, a novel distributed and reactive transport protocol for data center networks (DCNs) to achieve minimal per-packet delay while providing guaranteed transmission rates to meet flow deadlines. To design MCP, we first formulate a stochastic packet delay minimization problem with constraints on deadline completion and network stability. By solving this problem, we derive an optimal congestion window update function which establishes the theoretical foundation for MCP. To be incrementally deployable with existing switch hardware, MCP leverages functionality available on commodity switch, i.e., ECN, to approximate the optimal window update function. Our preliminary results show that MCP holds great promise in terms of deadline miss rate and goodput. Lei Chen 0002, Shuihai Hu, Kai Chen 0005, Danny H. K. Tsang |
HotNets | 2 |