VLDB 2026 Research / reviewers in the wild / expert
Chuhao Chen 0001
dblp:251/4860-1
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0003-6286-8964ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI CloudsabstractAI training and inference are driving cloud networks toward terabit-per-second (Tbps) bandwidth per server, challenging the scalability and efficiency of today's cloud network architectures. A prevalent design scales bandwidth by stacking monolithic Data Processing Units (DPUs), but this approach tightly couples control and data plane resources, leading to excessive cost, power consumption, and operational complexity. We identify a fundamental control-data plane divergence in AI clouds: while data plane bandwidth demand grows rapidly, control plane demand remains largely flat due to the dominance of elephant flows. As a result, monolithic DPUs become systematically over-provisioned when used as bandwidth scaling primitives. Lizhou Gao, Yuanyi Zhu, Chao Pei, Chuhao Chen 0001, Zijian Li 0003, Jian Zhao 0006, Dongbo Gu, Hongchen Ren, Jiyuan Chen, Yunpeng Guan, Jianye Yuan, Yibo Huang 0005, Yang Xu 0010 |
SIGCOMM | 7 |
| 2025 | Empowering Flowlet Load Balancing in RDMA with Host-Based Flowlet Fine-TuningabstractFlowlet-level load balancing has not demonstrated the expected robust capability in RDMA networks due to insufficient flowlets and the adverse effects of PFC. To delve deeper, we conduct measurements at end hosts and perform a detailed analysis of time gaps between packets. Our investigation reveals that in RDMA networks, the number of time gaps exceeding the flowlet timeout is considerably lower than the number in TCP networks. We also identify a stepwise time gap pattern that predicts the occurrence of PFC. Based on these observations, we propose$\text{HF}^{2} \mathrm{T}$, a host-based time gap adjustment method to improve the effectiveness of flowlet-level load balancing in RDMA networks. The core idea involves delaying a minimal number of specific packets at the host, actively extending the time gaps between them, thereby fostering the generation of sufficient flowlets at the switch and enhancing the utilization of equal-cost links. Incorporating an identification algorithm for the time gap pattern that predicts PFC,$\text{HF}^{2} \mathrm{T}$also leverages the time gap extension to reroute traffic away from potential PFC paths in advance, thus mitigating PFC occurrences. The minor cost of delaying a few packets is vastly offset by the benefits of generating flowlets and reducing PFC. We use DPDK to implement a prototype of$\text{HF}^{2} \mathrm{T}$, and through testbed experiments, we demonstrate that$\text{HF}^{2} \mathrm{T}$, serving as a building block for flowlet load balancing, can enhance the throughput of CONGA by 16.82%. The simulation results also show that$\text{HF}^{2} \mathrm{T}$can reduce the average FCT by 16.38% and the 99-percentile FCT by 21.13% compared to the state-of-the-art RDMA load balancing ConWeave. Chuhao Chen 0001, Deli Huang, Zerui Tian, Ruyi Yao, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 1 |
| 2024 | HF^2T: Host-Based Flowlet Fine-Tuning for RDMA Load BalancingabstractIn modern data center networks, RDMA is widely applied in scenarios such as high-performance computing, distributed storage and machine learning. In recent studies, it has been observed that flowlet switching load balancers cannot fully unleash their robust capabilities due to an insufficient number of flowlets in RDMA networks. In this paper, we scrutinize the traffic pattern at the end hosts and meticulously analyze time gaps between packets. Our findings reveal that in RDMA, the proportion of time gaps between packets larger than the flowlet threshold is notably scarce, constituting only a fraction of those in TCP, averaging 1/300. Based on this observation, we propose HF2T, a host-based method to improve the effectiveness of flowlet-level load balancing in RDMA. The core idea is to postpone a minimal number of specific packets at the host, actively elongating the time gaps between them, and promoting flowlet generation at the switch. The cost of postponing a minimal number of packets is far outweighed by the benefits of flowlets generation at the switch, improving the network performance. Simulation experiments confirm that HF2T, when deployed in conjunction with the flowlet load balancing, achieves an average reduction of 37.32% in Medium FCT and an average reduction of 28.75% in 99-percentile FCT, compared to deploying the same flowlet load balancing scheme solely at switches. Chuhao Chen 0001, Jiarui Ye, Yongbo Gao, Sen Liu 0002, Yang Xu 0010 |
APNet | 1 |
| 2023 | CoLUE: Collaborative TCAM Update in SDN Switches
Ruyi Yao, Chuhao Chen 0001, Wenjun Li 0004, Ying Wan 0001, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
INFOCOM | 4 |
| 2023 | Rusen: Rule Semantics Enabler toward Fast TCAM Update for Commodity SDN SwitchesabstractTernary Content Addressable Memory (TCAM) is widely used in Software-Defined Networking (SDN) switches due to its impressive throughput. But its unique circuit design results in long and inconsistent update delays. To overcome this challenge, many TCAM update algorithms based on rule semantics have been proposed. These algorithms eliminate unnecessary order restrictions, thus reducing update delays in theory. However, most commodity switches are semantic-unaware, which maintain rules in strict priority order. These algorithms are therefore not available for practical use. To address this issue, this paper proposes Rusen, a framework that enables the use of many semantic-based algorithms on Semantic-unaware commodity switches. Working as a transparent middle layer, the core idea of Rusen is to express the update scheme derived by semantic-based algorithms as messages that the Semantic-unaware switches can execute. In addition, Rusen optimizes the update scheme based on the specific characteristics of each switch, leading to improved performance of these algorithms. We evaluate the performance of Rusen by enabling several state-of-the-art semantic-based algorithms on commodity SDN switches. Results show that the average update delay can be significantly reduced by 23%∼94% on OpenFlow switches and 39%∼84% on a P4 switch. Ruoshi Sun, Ruyi Yao, Chuhao Chen 0001, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 4 |
| 2022 | BubbleTCAM: Bubble Reservation in SDN Switches for Fast TCAM UpdateabstractThe unique hardware structure of Ternary Content-Addressable Memory (TCAM) enables its unparalleled lookup throughput but also causes slow update due to the Priority Order Constraint (POC). With the increase of application demands, TCAM update has become a bottleneck in the network. This paper proposes a new TCAM management mechanism named BubbleTCAM to enable fast TCAM update, in which available empty entries are defined as bubbles. The core idea of Bub-bleTCAM is to uniformly distribute bubbles and dependency chains in TCAM, which is beneficial to updates. BubbleTCAM consists of two components: bubble management and rule insertion. Bubble management enables TCAM to have uniformly distributed bubbles at all times through three key procedures: bubble lock reservation, bubble lock release and bubble generation. Rule insertion ensures that dependency chains of rules are uniformly stretched and distributed in TCAM. In addition, BubbleTCAM avoids the reorder problem by pre-sorting. Our evaluation based on the rulesets generated by ClassBench shows that BubbleTCAM effectively reduces the average cost and worst cost (in units of rule movements) during rule updates by at least 48% and 50%, respectively. Especially for the worst cost, the performance can be improved by up to 196x. Chuhao Chen 0001, Ruyi Yao, Ying Wan 0001, Wenjun Li 0004, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
IWQoS | 2 |