VLDB 2026 Research / reviewers in the wild / expert
Chengyuan Huang
dblp:202/5480
· DBLP profile ↗
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-6079-6579ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 6 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RiCE: Precise Remote In-Network Congestion Elimination in Inter-Datacenter RDMA Networks
Chengyuan Huang, Guangyu Zhao, Lu Lu 0016, Zirui Wan, Jiaqing Dong, Zhuo Tang, Guihai Chen, Chen Tian 0001 |
IWQoS | 2 |
| 2026 | Revisiting Flow Control in Node-Centric Datacenter NetworksabstractNode-centric Data Centers (NDCs) are highly flexible, cost-efficient, and failure-resilient, and have gained growing popularity in recent years. However, RDMA technology used in NDC still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems. Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of NDC. In this paper, we propose a per-port flow control named PortFC for NDC. PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a server-aware queue scheduling method. Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks. PortFC achieves 1.7-8.0 times higher throughput and reduces latency by 11.7%-87.7% compared to the state-of-the-art lossy RDMA based on IRN and the lossless RDMA method based on PFC. In particular, PortFC still demonstrates good performance in a Rail-only NDC with heterogeneous bandwidth domains. Peirui Cao, Rui Ning, Guangyu Zhao, Zhaochen Zhang, Chang Liu 0001, Yunzhuo Liu, Rui Li 0020, Chengyuan Huang, Tao Sun 0010, Guihai Chen, Baochun Li, Chen Tian 0001 |
IEEE Trans. Netw. | 8 |
| 2026 | UDMP: Unified Delay-Driven Multipath Protocol for AI ClustersabstractDistributed AI model training generates bursty, low-entropy elephant flows that challenge existing single-path transport protocols in multi-stage Clos networks, leading to congestion and inefficiency. Multipath transport emerges as a promising solution, leveraging multiple paths to balance traffic and enhance resilience. However, current multipath RDMA solutions suffer from scalability, congestion control, and load-balancing inefficiencies. This paper introduces Unified Delay-driven Multipath Protocol (UDMP), a novel approach that co-designs congestion control and load balancing using network delay as a unified signal. UDMP employs delay-gradient-based congestion control to precisely resolve unavoidable congestion. Moreover, UDMP leverages delay-assisted load balancing to shift traffic across paths with minimal latency adaptively, maintaining throughput when encountering avoidable congestion. A novel Token Pool design integrates these components, eliminating per-path state overhead while achieving fine-grained traffic distribution. Implementations on DPDK and NS3 demonstrate that UDMP achieves up to 2x higher throughput and reduces flow completion times by up to 30% compared to state-of-the-art methods like MPRDMA and QP-Scaling. These results highlight UDMP’s effectiveness in meeting the stringent performance requirements of modern distributed AI training workloads. Chengyuan Huang, Zhengqi Cui, Jun Xu 0037, Zhaochen Zhang, Li Wang 0110, Peirui Cao, Zhongming Ji, Jilei Chen, Shengju Zhang, Lingkun Meng, Ahmed M. Abdelmoniem, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Keqiang He, Chen Tian 0001 |
IEEE Trans. Netw. | 1 |
| 2026 | Rail: ReArranging Inter-GPU Links for GPU-Centric ClustersabstractIn modern GPU-centric clusters, large-scale AI training relies on two distinct communication domains: a high-bandwidth intra-node domain using proprietary interconnects (e.g., NVLink), and a scale-out inter-node network domain (e.g., RDMA). We observe that the widely-used ring algorithm, often create a significant load imbalance across these domains. This leads to the counter-intuitive scenario where the expensive, high-bandwidth intra-node domain becomes a performance bottleneck, while the inter-node network remains underutilized. This inefficiency is further exacerbated by the disparity in bandwidth provisioning: inter-node network bandwidth is generally more cost-effective and accessible, whereas intra-node bandwidth is often proprietary and more costly to scale. To address this fundamental imbalance, we propose RAIL, aimed at resolving the intra-node bottleneck by strategically rearranging inter-GPU communication paths. This rebalancing ensures that traffic loads are appropriately matched with the distinct transmission capabilities of each domain, thereby maximizing overall communication performance. RAIL incorporates a Load Distributing Strategy (LDS) that can accurately partition physical nodes into logical nodes based on the a transmission capabilities of both domains, shifting excess traffic from the overloaded intra-node domain to the underutilized network domain. Additionally, the Intra-Rail Strategy (IRS) leverages topological characteristics to ensure optimal communication paths through the network domain between logical nodes. Our evaluation demonstrates that RAIL effectively mitigates congestion and achieves a 30.7% average increase in collective communication bus bandwidth compared to the widely-used NCCL solution. Haixin Nan, Jun Xu 0037, Peirui Cao, Zhaochen Zhang, Yizhi Wang 0004, Zhehao Lin, Yuhang Li 0002, Chengyuan Huang, Xiaohu Xu, Zhongming Ji, Shengju Zhang, Lingkun Meng, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 8 |
| 2025 | PortFC: Designing High-performance Deadlock-free BCube NetworksabstractBCube is a modular data center network.Compared with other topologies, BCube has natural advantages, such as lower deployment costs and stronger failure recovery capabilities.However, RDMA technology used in BCube still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems.Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of BCube.In this paper, we propose a per-port flow control named PortFC for BCube.PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a serveraware queue scheduling method.Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks.PortFC achieves 1.7-8.0times higher throughput and reduces latency by 11.7%-87.7%compared to the state-of-the-art Peirui Cao, Rui Ning, Zhaochen Zhang, Chang Liu 0001, Rui Li 0020, Yongqi Yang, Yunzhuo Liu, Chengyuan Huang, Tao Sun 0010, Xiaodong Duan, Guihai Chen, Chen Tian 0001 |
ICS | 9 |
| 2025 | Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleabstractThe flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers. Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001 |
SIGCOMM | 5 |
| 2025 | Reunion: Receiver-driven network load balancing mechanism in AI training clusters
Mingyao Wang, Keqiang He, Peirui Cao, Jiong Duan, Dongliang Lv, Chengyuan Huang, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
Comput. Networks | 8 |
| 2025 | Troubleshooting Programmable Data Planes via Real-Time Table Information RecordingabstractWhile the flexibility of programmable switches brings opportunities, it also introduces security risks. Hence, it is vital to conduct effective troubleshooting in the programmable switch to mitigate frequent network failures. However, troubleshooting programmable switch failures is challenging due to their enhanced flexibility and functionality compared to regular switches, posing increased difficulty in debugging, particularly with limited debugging tools and information. To address this problem, we propose an efficient troubleshooting method that records real-time information about packets in the data plane, including the tables involved in packet processing. Unfortunately, due to hardware limitations, it is infeasible to record all tables’ information in the data plane. Thus, the key is to find the table set reflecting the execution path a packet goes through while minimizing the resource overhead. We first represent P4 programs as a probabilistic transition directed acyclic graph (DAG) and employ information entropy to quantify the information within a set of tracked tables. Then, we adopt a two-step approach and design algorithms to find both optimal and approximately optimal table record plans. The evaluation results show the efficacy of the proposed method, including achieving the same path recovery rate as the related works with less than one-third of the resource consumption. Chengyuan Huang, Yibo Xiao, Tianfan Zhang, Bingheng Yan, Ahmed M. Abdelmoniem, Gianni Antichi, Xiaoliang Wang 0001, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 1 |
| 2024 | μMon: Empowering Microsecond-level Network Monitoring with WaveletsabstractNetwork monitoring is essential for network management and optimization. In modern data centers, fluctuations in flow rates and network congestion events (e.g., microbursts) typically manifest on a microsecond timescale. However, the time granularity of network monitoring systems has not been refined correspondingly to efficiently capture these behaviors. Attaining the monitoring granularity at the microsecond scale can greatly facilitate network performance analysis and management, but poses considerable challenges regarding memory, bandwidth, and deployment costs. We propose μMon, a novel microsecond-level network monitoring system for data centers. The key of μMon is WaveSketch, an innovative algorithm that measures and compresses flow rate curves using in-dataplane wavelet transform. WaveSketch allows for more accurate characterization of application traffic patterns and aids in profiling transport algorithms. Furthermore, by combining the fine-grained flow rate measurements with network-collected congestion information, μMon can 'replay' congestion events to analyze their cause and impact. We evaluate μMon through testbed deployment and simulations at a granularity of 8.192 μs. The evaluation results demonstrate that μMon can achieve a 90% accuracy in microsecond-level rate measurements with an average of 5 Mbps bandwidth overhead per host. Additionally, it can capture 99% heavy congestion events with 31--82 Mbps bandwidth overhead per switch. Chengyuan Huang, Xiangyu Han, Jiaqi Zheng 0001, Xiaoliang Wang 0001, Chen Tian 0001, Wan-Chun Dou, Guihai Chen |
SIGCOMM | 2 |
| 2024 | MpScope: Enabling multi-pipeline monitoring inside a switch
Chengyuan Huang, Tianfan Zhang, Li Wang 0110, Yibo Xiao, Chen Tian 0001, Xiaoliang Wang 0001, Bingheng Yan, Ahmed M. Abdelmoniem, Wan-Chun Dou, Guihai Chen |
Comput. Networks | 1 |
| 2024 | Minimizing Buffer Utilization for Lossless Inter-DC LinksabstractRDMA over Converged Ethernet (RoCEv2) has been widely deployed to data centers (DCs) for its better compatibility with Ethernet/IP than Infiniband (IB). As cross-DC applications emerge, they also demand high throughput, low latency, and lossless network for cross-DC data transmission. However, RoCEv2’s underlying lossless mechanism Priority-based Flow Control (PFC) cannot fit into the long-haul transmission scenario and degrades the performance of RoCEv2. PFC is myopic and only considers queue length to pause upstream senders, which leads to large queueing delay. This paper proposes Bifrost, a downstream-driven lossless flow control that supports long distance cross-DC data transmission. Bifrost uses virtual incoming packets, which indicates the upper bound of in-flight packets, together with buffered packets to control the flow rate. It minimizes the buffer space requirement to one-hop bandwidth delay product (BDP) and achieves low one-way latency. Moreover, we extend Bifrost and propose BifrostX, to accommodate the multi-priority queue of the current switch implementation. BifrostX enables flow control for each queue separately while maintaining low buffer reservation, no throughput loss, and no packet loss. Real-world experiments are conducted with prototype switches and 80 kilometers cables. Evaluations demonstrate that compared to PFC, Bifrost reduces average/tail flow completion time (FCT) of inter-DC flows by up to 22.5%/42.0%, respectively. Bifrost is compatible with existing infrastructure and can support distance of thousands of kilometers. Chengyuan Huang, Feiyang Xue, Xiaoliang Wang 0001, Tao Wu 0011, Zifa Han, Xiangyu Gong, Chen Tian 0001, Wan-Chun Dou, Guihai Chen |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | MC-RDMA: Improving Replication Performance of RDMA-based Distributed Systems with Reliable Multicast SupportabstractRemote Direct Memory Access has been widely adopted in distributed storage systems. However, it only supports unicast operations, which degrades the performance significantly for data replication because of bandwidth waste and CPU overhead. To address the problem, we propose MC-RDMA, a distributed and reliable multicast RDMA. It is compatible with existing unicast RDMA but supports lazy packet replication with reliable RDMA multicasting. The key idea of MC-RDMA is utilizing in-network programmable switches to build a NIC-transparent reliable multicast protocol for RDMA. MC-RDMA combines the address information of the IP and RoCEv2 into a sender-initialized multicast routing protocol. Besides, it synchronizes the hardware transmission states of multiple receivers by merging ACKs and NAKs. To verify the effectiveness of MC-RDMA, we implement it with Mellanox ConnectX-6 commodity RNICs and Intel Tofino P4 programmable switches. Experimental results show that MC-RDMA can double the sender bandwidth utilization and reduce the CPU overhead significantly compared to unicast-based RDMA replications. Moreover, it reduces the storage request latency by -30% with realistic workloads and decreases the training time by -50% in the distributed training system. Chengyuan Huang, Yixiao Gao, Duoxing Li, Yibo Xiao, Ruyi Zhang 0005, Chen Tian 0001, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Fu Xiao 0001 |
ICNP | 1 |
| 2023 | Achieving Zero-copy Serialization for Datacenter RPCabstractRemote Procedure Call (RPC) is widely used in distributed systems and it usually needs to serialize data before transmission. Serialization accounts for a large proportion of the overhead in RPC and becomes a bottleneck for RPC communications. Because the size of the output serialized message cannot be predicted in advance, there could be multiple memory reallocations and copies in typical serialization libraries (e.g., FlatBuffers), which dominates the overhead. We propose the novel serialization library, zFlatBuffers, to eliminate these avoidable copies during the serialization process and realize zero copy during communication. Unlike the typical serialization library, FlatBuffers, the message generated by zFlatBuffers consists of multiple non-contiguous buffers due to its zero-copy nature. Moreover, we integrate zFlatBuffers with RDMA-based RPC systems. For RDMA Unreliable Datagram, we modify the message buffer of eRPC to enable it to transmit messages composed of multiple buffers. We also build the zRPC system based on RDMA Reliable Connection, which transmits the zFlatBuffers message by the scatter/gather function. Compared to the original FlatBuffers, zFlatBuffers improves the throughput of eRPC and zRPC by 11.2%-33.7% and 5.8%-53.6%, respectively. Tianfan Zhang, Huaping Zhou, Chengyuan Huang, Chen Tian 0001, Xiaoliang Wang 0001, Ahmed M. Abdelmoniem, Matthew Tan, Wan-Chun Dou, Guihai Chen |
IPCCC | 3 |
| 2023 | SLIT: Achieving Fast Bandwidth Isolation Across Virtual MachinesabstractNetwork performance guarantee in the cloud is one of the hottest research topics recently. However, current approaches either lack scalability or fail to achieve high performance. To avoid the hard tradeoff between the scalability and performance, we adopt a hybrid approach to get the best of both worlds. In this paper, we follow the stateless design principle and propose SLIT, which provides fast VM-level network performance isolation and maintains core-stateless characteristics. Specifically, at the end host, we develop a novel network-adaptive labeling mechanism and it computes the correct scheduling priority for each hop. At the switch, packets are scheduled based on their labels and we develop a novel straggler detection algorithm to find the new incoming flow. Our evaluation results show that SLIT can approximate the optimal performance of Weighted Fair Queuing (WFQ) effectively. Chengyuan Huang, Jiao Zhang 0002, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 1 |