EDBT 2026 Demo / reviewers in the wild / expert
Gaoxiong Zeng
dblp:204/5313
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-1876-0329ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 5 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reliable RDMA Over Lossy Fabrics via Data-Control Partitioning
Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Bingyang Liu, Kai Chen 0005 |
IEEE Trans. Netw. | 7 |
| 2026 | High-Performance RoCE-Capable Multicast for Commodity RDMA DatacentersabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
IEEE Trans. Netw. | 4 |
| 2025 | Sub-RTT Congestion Control for Inter-Datacenter NetworksabstractWith the explosive growth in the scale and complexity of large language models (LLMs), there is an urgent need to extend training and inference workloads from within a single data center to across multiple data centers. However, this also introduces new challenges for network transport protocols. To address these issues, we propose SRCC (Sub-RTT Congestion Control), a method designed for inter-datacenter networks. Specifically, SRCC introduces a flowset-based mechanism along with shared node tables, enabling Datacenter Interconnect (DCI) switches to be aware of the path status of each flow. By leveraging information shared among different flows, SRCC can accurately adjust the sending rate at a sub-RTT timescale, thereby significantly improving network performance. Building on this approach, we design detailed mechanisms to address the following challenges: (1) applying INT technology in wide-area networks; (2) acquiring INT information with low overhead; and (3) achieving precise congestion window adjustments under sub-RTT perception.We conducted large-scale simulations using NS3, and the experimental results show that our scheme reduces the average FCT slowdown by 44.17% and 53.86% compared to HPCC and DCTCP, respectively. Jun Wang 0178, Yuchao Zhang 0004, Gaoxiong Zeng, Chenyue Zheng, Wendong Wang 0003, Haipeng Yao |
ICNP | 3 |
| 2025 | Revisiting RDMA Reliability for Lossy FabricsabstractDue to the high operational complexity and limited deployment scale of lossless RDMA networks, the community has been exploring efficient RDMA communication over lossy fabrics. State-of-the-art (SOTA) lossy RDMA solutions implement a simplified selective repeat mechanism in RDMA NICs (RNICs) to enhance loss recovery efficiency. However, these solutions still face performance challenges, such as unavoidable ECMP hash collisions and excessive retransmission timeouts (RTOs). In this paper, we revisit RDMA reliability with the goals of being independent of PFC, compatible with packet-level load balancing, free from RTO, and friendly to hardware offloading. To this end, we propose DCP, a transport architecture that co-designs both the switch and RNICs, fully meeting the design goals. At its core, DCP-Switch introduces a simple yet effective lossless control plane, which is leveraged by DCP-RNIC to enhance reliability support for high-speed lossy fabrics, primarily including header-only-based retransmission and bitmap-free packet tracking. We prototype DCP-Switch using P4 switch and DCP-RNIC using FPGA. Extensive experiments demonstrate that DCP achieves 1.6× and 2.1× performance improvements, compared to SOTA lossless and lossy RDMA solutions, respectively. Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Kai Chen 0005, Bingyang Liu |
SIGCOMM | 7 |
| 2024 | Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable MulticastabstractModern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage. Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005 |
HPCA | 4 |
| 2024 | PAS: Towards Accurate and Efficient Federated Learning with Parameter-Adaptive SynchronizationabstractFederated Learning (FL) is a distributed paradigm that supports collaborated model training while preserving data privacy, where clients periodically synchronize their local gradients once after multiple local iterations. Due to non-uniform data distribution and poor network condition, FL processes often suffer degraded training accuracy and efficiency. In this work, we analyze the microscopic parameter variation behaviors in FL, and find that an effective method to improve FL accuracy is to switch to more frequent synchronization at proper moments. Moreover, such moments can be detected from gradient characteristics, and are heterogeneous across different parameters. Motivated by such observations, we propose Parameter-Adaptive Synchronization (PAS), a FL scheme that adaptively tunes the synchronization period for each scalar parameter. The benefits of PAS are two-fold: By switching to more frequent synchronization when necessary, we can improve the FL training accuracy; by synchronizing different parameters independently, we can enable communication-computation overlapping and enhance the network utilization. We implemented PAS atop PyTorch, and extensive experiments show that it can substantially improve FL performance in both accuracy and communication efficiency. Zuo Gan, Chen Chen 0067, Jiayi Zhang 0006, Gaoxiong Zeng, Yifei Zhu 0001, Jieru Zhao, Quan Chen 0002, Minyi Guo |
IWQoS | 4 |
| 2024 | Towards Domain-Specific Network Transport for Distributed DNN Training
Hao Wang 0116, Han Tian, Jingrong Chen 0004, Xinchen Wan, Jiacheng Xia, Gaoxiong Zeng, Wei Bai 0001, Junchen Jiang, Yong Wang 0046, Kai Chen 0005 |
NSDI | 6 |
| 2024 | FLAIR: A Fast and Low-Redundancy Failure Recovery Framework for Inter Data Center NetworkabstractDue to the fast developments of 5G and IoT technologies, Inter-Datacenter (Inter-DC) networks are facing unprecedented pressure to duplicate large volumes of geographically distributed user data in a real-time manner. Meanwhile, with the expansion of Inter-DC networks scale, link/node failures also become increasingly frequent, negatively affecting the data transmission efficiency. Therefore, link failure recovery methods become of utmost importance. Many works investigated fast failure recovery, yet none of them consider the deployment overhead of such recovery schemes. While in this paper, we found that the side-effect of deploying recovery strategies and the future availability of the recovered transmissions are also crucial for fast recovery. So we propose a fast and low-redundancy failure recovery framework, FLAIR, which consists of a fast recovery strategy FRAVaR and a redundancy removal algorithm ROSE. FRAVaR takes full consideration of deployment overhead by minimizing shuffle traffic. On its base, ROSE regularly eliminates the cumulative rerouting redundancy by removing unnecessary routing updates. The experiment results on 4 realistic network topologies show that FLAIR successfully reduces up to 48.2% deployment overhead compared with the state-of-the-art solutions, and thus reduces up to 70.2% recovery speed and improves up to 36% network utilization. Yuchao Zhang 0004, Haoqiang Huang, Ahmed M. Abdelmoniem, Gaoxiong Zeng, Chenyue Zheng, Xirong Que, Wendong Wang 0003, Ke Xu 0002 |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Cutting Tail Latency in Commodity Datacenters with CloudburstabstractLong tail latency of short flows (or messages) greatly affects user-facing applications in datacenters. Prior solutions to the problem introduce significant implementation complexities, such as global state monitoring, complex network control, or non-trivial switch modifications. While promising superior performance, they are hard to implement in practice.This paper presents Cloudburst, a simple, effective yet readily deployable solution achieving similar or even better results without introducing the above complexities. At its core, Cloudburst explores forward error correction (FEC) over multipath — it proactively spreads FEC-coded packets generated from messages over multipath in parallel, and recovers them with the first few arriving ones. As a result, Cloudburst is able to obliviously exploit underutilized paths, thus achieving low tail latency. We have implemented Cloudburst as a user-space library, and deployed it on a testbed with commodity switches. Our testbed and simulation experiments show the superior performance of Cloudburst. For example, Cloudburst achieves 63.69% and 60.06% reduction in 99th percentile message/flow completion time (FCT) compared to DCTCP and PIAS, respectively. Gaoxiong Zeng, Li Chen 0008, Bairen Yi, Kai Chen 0005 |
INFOCOM | 1 |
| 2022 | Aeolus: A Building Block for Proactive Transport in Datacenter NetworksabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth isproactivelyallocated as “credits” to senders who then can send “scheduled packets” at a right rate to ensure high link utilization, low latency, and zero packet loss. Consequently, proactive solutions such as ExpressPass, NDP, Homa, etc., have been proposed recently. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT “pre-credit” phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on “pre-credit” packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. Aeolus is compatible with all existing proactive solutions and readily implementable with commodity switches. We have integrated Aeolus into ExpressPass, NDP and Homa, and shown, via both implementation and simulations, that the Aeolus-enhanced solutions deliver significant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by$20\times $, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Gaoxiong Zeng, Wei Bai 0001, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run oncross-datacenter networkthat consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network poses significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., buffer depths, RTTs). In this paper, we find that existing DCN or WAN transport reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself only, can simultaneously capture the location and degree of congestion, mainly due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low buffer occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31%, 76% and 2% reduction of small flow average completion times, and up to 34%, 39%, 9% and 58% reduction of large flow average completion times compared to TCP Cubic, DCTCP, BBR and TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | FlashPass: Proactive Congestion Control for Shallow-buffered WANabstractIn recent years, large enterprises (e.g., Google, Alibaba, etc.) have been building and deploying their wide-area routers based on shallow-buffered switching chips. However, with legacy reactive transport (e.g., TCP Cubic), shallow buffer can easily get overwhelmed by large BDP wide-area traffic, leading to high packet losses and degraded throughput. To address it, we ask: can we design a transport to simultaneously achieve high throughput and low loss for shallow-buffered WAN?We answer this question affirmatively by employing proactive congestion control (PCC). However, two issues exist for existing PCC to work on WAN. Firstly, wide-area traffics have diverse RTTs, leading to what we called imperfect scheduling issue (e.g., data crash in time). Secondly, there is one RTT delay for credits to trigger data sending, which may degrade network performance. Therefore, we propose a novel PCC design - FlashPass. To address the first issue, FlashPass adopts sender-driven emulation process with send time calibration to avoid the data packet crash. To address the second issue, FLASHPASS enables early data transmission in the starting phase, and incorporates an over-provisioning with selective dropping mechanism for efficient credit allocation in the finishing phase. Our evaluation with production workload demonstrates that FlashPass reduces the overall flow completion times of TCP Cubic and ExpressPass by up to 32% and 11.4%, and the 99-th tail completion times of small flows by up to 49.5% and 38%, respectively. Gaoxiong Zeng, Jianxin Qiu, Hongqiang Liu, Kai Chen 0005 |
ICNP | 1 |
| 2020 | Aeolus: A Building Block for Proactive Transport in DatacentersabstractAs datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth is proactively allocated as "credits" to senders who then can send "scheduled packets" at a right rate to ensure high link utilization, low latency, and zero packet loss. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT "pre-credit" phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on "pre-credit" packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. We have integrated Aeolus into ExpressPass[14], NDP[18] and Homa[29], and shown, through both implementation and simulations, that the Aeolus-enhanced solutions deliver signiicant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by 20x, while achieving similar performance as NDP without switch modifications. Shuihai Hu, Wei Bai 0001, Gaoxiong Zeng, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004 |
SIGCOMM | 3 |
| 2019 | Rethinking Transport Layer Design for Distributed Machine LearningabstractMotivated by the increasing scale of data, we see a growing need of high performance distributed machine learning systems. Many research works are being proposed to improve distributed machine learning performance. Jiacheng Xia, Gaoxiong Zeng, Junxue Zhang 0001, Weiyan Wang, Wei Bai 0001, Junchen Jiang, Kai Chen 0005 |
APNet | 2 |
| 2019 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run on cross-datacenter network that consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network imposes significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., butter depths, RTTs). In this paper, we find that existing DCN or WAN transports reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself, can simultaneously capture the location and degree of congestion. This is due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low butter occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31% and 76% reduction of small flow average completion times compared to TCP Cubic, DCTCP and BBR; and up to 58% reduction of large flow average completion times compared to TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
ICNP | 1 |
| 2017 | Combining ECN and RTT for Datacenter TransportabstractDatacenter transports should provide low average and tail flow completion times (FCT) to achieve desired application performance. While most prior datacenter transports take either ECN or RTT as congestion signal, this paper makes a case that both signals are indispensable: ECN, as a per-hop signal, is more effective to prevent packet loss; while RTT, as an end-to-end signal, controls end-to-end queueing delay better. As persistent low flow completion times imply low queueing delay and near zero packet loss, we introduce EAR, a new datacenter transport that hears and reacts to both ECN and RTT. Our preliminary results show that: 1) compared to delay-based DCTCP, EAR achieves up to 91% lower packet losses and 93% fewer timeouts; 2) compared to ECN-based DCTCP, EAR reduces RTT by up to 32% for cross-rack traffic in a 4-level fattree. As a result, EAR delivers persistent low average and tail completion times under various scenarios in large scale simulations. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
APNet | 1 |