Chaoliang Zeng

dblp:257/6290 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-5151-0997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 3 first-author · 18 since 2021Systems, architecture and hardware · 5 · 4 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 High-Performance RoCE-Capable Multicast for Commodity RDMA Datacenters
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
IEEE Trans. Netw.6
2026 Towards Fair and Efficient Congestion Control Through Multi-Agent Deep Reinforcement Learning
abstract
Recent years have witnessed a plethora of learning-based solutions for congestion control (CC) that demonstrate better performance over traditional TCP schemes. However, they fail to provide consistently good convergence properties, includingfairness, fast convergenceandstability, due to the mismatch between their objective functions and these properties. Despite being intuitive, integrating these properties into existing learning-based CC is challenging, because: 1) their training environments are designed for the performance optimization of single flow but incapable of cooperative multi-flow optimization, and 2) there is no directly measurable metric to represent these properties into the training objective function. We present Astraea, a new learning-based congestion control that ensures fast convergence to fairness with stability. At the heart of Astraea is a multi-agent deep reinforcement learning framework that explicitly optimizes these convergence properties during the training process by enabling the learning of interactive policy between multiple competing flows, while maintaining high performance. We further build a faithful multi-flow environment that emulates the competing behaviors of concurrent flows, explicitly expressing convergence properties to enable their optimization during training. We have fully implemented Astraea and our comprehensive experiments show that Astraea can quickly converge to fairness point and exhibit better stability than its counterparts. For example, Astraea achieves near-optimal bandwidth sharing (i.e., fairness) when multiple flows compete for the same bottleneck, delivers up to 8.4× faster convergence speed and 2.8× smaller throughput deviation, while achieving comparable or even better performance over prior solutions.
Han Tian, Xudong Liao, Chaoliang Zeng, Xinchen Wan, Junxue Zhang 0001, Kai Chen 0005
IEEE Trans. Netw.3
2025 Achieving Fairness Generalizability for Learning-based Congestion Control with Jury
abstract
Internet congestion control (CC) has long posed a challenging control problem in networking systems, with recent approaches increasingly incorporating deep reinforcement learning (DRL) to enhance adaptability and performance. Despite promising, DRL-based CC schemes often suffer from poor fairness, particularly when applied to network environments unseen during training. This paper introduces Jury, a novel DRL-based CC scheme designed to achieve fairness generalizability. At its heart, Jury decouples the fairness control from the principal DRL model with two design elements: i) By transforming network signals, it provides a universal view of network environments among competing flows, and ii) It adopts a post-processing phase to dynamically module the sending rate based on flow bandwidth occupancy estimation, ensuring large flows behave more conservatively and smaller flows more aggressively, thus achieving a fair and balanced bandwidth allocation. We have fully implemented Jury, and extensive evaluations demonstrate its robust convergence properties and high performance across a broad spectrum of both emulated and real-world network conditions.
Han Tian, Xudong Liao, Decang Sun, Chaoliang Zeng, Yilun Jin, Junxue Zhang 0001, Xinchen Wan, Zilong Wang 0007, Yong Wang 0046, Kai Chen 0005
EuroSys4
2025 A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Han Tian, Xudong Liao, Chaoliang Zeng, Zilong Wang 0007, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen 0005
INFOCOM6
2025 Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload Shaping
Xudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng, Hao Wang 0116, Junxue Zhang 0001, Mengyu Ma, Guyue Liu, Kai Chen 0005
USENIX ATC4
2024 Astraea: Towards Fair and Efficient Learning-based Congestion Control
abstract
Recent years have witnessed a plethora of learning-based solutions for congestion control (CC) that demonstrate better performance over traditional TCP schemes. However, they fail to provide consistently good convergence properties, including fairness, fast convergence and stability, due to the mismatch between their objective functions and these properties. Despite being intuitive, integrating these properties into existing learning-based CC is challenging, because: 1) their training environments are designed for the performance optimization of single flow but incapable of cooperative multi-flow optimization, and 2) there is no directly measurable metric to represent these properties into the training objective function.
Xudong Liao, Han Tian, Chaoliang Zeng, Xinchen Wan, Kai Chen 0005
EuroSys3
2024 Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable Multicast
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
HPCA6
2024 Accelerating Neural Recommendation Training with Embedding Scheduling
Chaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian, Xinchen Wan, Hao Wang 0116, Kai Chen 0005
NSDI1
2024 Accelerating Secure Collaborative Machine Learning with Protocol-Aware RDMA
Zhenghang Ren, Mingxuan Fan, Zilong Wang 0007, Junxue Zhang 0001, Chaoliang Zeng, Cheng Hong 0001, Kai Chen 0005
USENIX Security Symposium5
2024 Load Balancing With Multi-Level Signals for Lossless Datacenter Networks
abstract
Various datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-level signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35%, 31%, 28%, 22% and 46%, 42%, 34%, 29% better average FCT and$99^{th}$percentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
IEEE/ACM Trans. Netw.2
2024 FlowSail: Fine-Grained and Practical Flow Control for Datacenter Networks
abstract
As datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effectiveness due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce FlowSail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of FlowSail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FlowSail using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that FlowSail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FlowSail achieves$4.3\times $lower tail latency under the same number of queues, matches existing FC schemes with$4\times $fewer queues, and holds robust performance with a minimum of 2 queues.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
IEEE/ACM Trans. Netw.2
2024 Efficient DRL-Based Congestion Control With Ultra-Low Overhead
abstract
Previous congestion control (CC) algorithms based on deep reinforcement learning (DRL) directly adjust flow sending rate to respond to dynamic bandwidth change, resulting in high inference overhead. Such overhead may consume considerable CPU resources and hurt the datapath performance. In this paper, we present, a hierarchical congestion control algorithm that fully utilizes the performance gain from deep reinforcement learning but with ultra-low overhead. At its heart, decouples the congestion control task into two subtasks in different timescales and handles them with different components: 1) lightweight CC executor that performs fine-grained control responding to dynamic bandwidth changes; and 2) RL agent that works at a coarse-grained level that generates control sub-policies for the CC executor. Such two-level control architecture can provide fine-grained DRL-based control with a low model inference overhead. Real-world experiments and emulations show that achieves consistent high performance across various network conditions with an ultra-low control overhead reduced by at least 80% compared to its DRL-based counterparts, similar to classic CC schemes such as Cubic.
Han Tian, Xudong Liao, Chaoliang Zeng, Decang Sun, Junxue Zhang 0001, Kai Chen 0005
IEEE/ACM Trans. Netw.3
2024 LiteFlow: Toward High-Performance Adaptive Neural Networks for Kernel Datapath
abstract
Adaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution. This paper presents LiteFlow, a hybrid solution to build high-performance adaptive NNs for kernel datapath. At its core, LiteFlow decouples the control path of adaptive NNs into: 1) a kernel-space fast path for efficient model inference; and 2) a userspace slow path for effective model tuning. We have implemented LiteFlow with Linux kernel datapath and evaluated it with three popular datapath functions including congestion control, flow scheduling, and load balancing. Compared to prior works, LiteFlow achieves 44.4% better goodput for congestion control, and improves the completion time for long flows by 33.7% and 56.7% for flow scheduling and load balancing, respectively.
Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005
IEEE/ACM Trans. Netw.2
2023 Scaling Switch-driven Flow Control with Aquarius
abstract
As datacenter networks support more diverse applications and faster link speeds, effective end-to-end congestion control becomes increasingly challenging due to the inherent feedback delay. To address this issue, switch-driven per-hop flow control (FC) has gained popularity due to its natural flow isolation, timely control loop, and ability to handle transient congestion. However, the ideal FC requires impractical hardware resources, and the state-of-the-art approximation approach still demands a large number of queues that exceeds common switch capabilities, limiting scalability in practice.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
APNet2
2023 Accurate and Scalable Rate Limiter for RDMA NICs
abstract
Rate limiter is required by RDMA NIC (RNIC) to enforce the rate limits calculated by congestion control. RNIC expects the rate limiter to be accurate and scalable: to precisely shape the traffic for numerous flows with minimized resource consumption, thereby mitigating the incasts and congestions and improving the network performance. Previous works, however, fail to meet the performance requirements of RNIC while achieving accuracy and scalability.
Zilong Wang 0007, Xinchen Wan, Chaoliang Zeng, Kai Chen 0005
APNet3
2023 Enabling Load Balancing for Lossless Datacenters
abstract
Various datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-Ievel signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35 %, 31 %, 28%, 22% and 46 %, 42 %, 34 %, 29 % better average FCT and 99thpercentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
ICNP2
2023 Towards Fine-Grained and Practical Flow Control for Datacenter Networks
abstract
As datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effective due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce Flowsail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of Flowsail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FLOWSAIL using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that Flowsail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FLOWSAIL achieves 4.3 x lower tail latency under the same number of queues, matches existing FC schemes with 4 x fewer queues, and holds robust performance with a minimum of 2 queues.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
ICNP2
2023 SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo
NSDI4
2022 Load Balancing in PFC-Enabled Datacenter Networks
abstract
In Priority Flow Control (PFC) enabled datacenter networks (DCNs), PFC is inevitably triggered due to bursty traffic even with end-to-end congestion control. Load balancing as a complementary mechanism to transport protocols can make rerouting decisions in time to alleviate PFC’s head-of-line (HoL) blocking problem. However, prior solutions designed for lossy DCNs do not work well in PFC-enabled networks, because the unreliable rerouting signals such as separate local queue length, round-trip time (RTT), explicit congestion notification (ECN), and link load cannot timely and correctly reflect PFC pausing.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
APNet2
2022 Herald: An Embedding Scheduler for Distributed Embedding Model Training
abstract
Given the ability to represent categorical features, embedding models have gained great success on many internet services. State-of-the-art training frameworks enable embedding cache in GPU workers to benefit from hardware acceleration while supporting massive category representations (embeddings) in the limited-capacity GPU device memory. However, based on our measurements, naively adopting a cache system in embedding model training leads to non-negligible communications overhead between caches and the global parameter server. We observe that many such communications are avoidable, given the predictability and sparsity natures of embedding cache accesses in distributed training.
Chaoliang Zeng, Xiaodian Cheng, Han Tian, Hao Wang 0116, Kai Chen 0005
APNet1
2022 Spine: an efficient DRL-based congestion control with ultra-low overhead
abstract
Previous congestion control (CC) algorithms based on deep reinforcement learning (DRL) directly adjust flow sending rate to respond to dynamic bandwidth change, resulting in high inference overhead. Such overhead may consume considerable CPU resources and hurt the datapath performance. In this paper, we present Spine, a hierarchical congestion control algorithm that fully utilizes the performance gain from deep reinforcement learning but with ultra-low overhead. At its heart, Spine decouples the congestion control task into two subtasks in different timescales and handles them with different components: i) a lightweight CC executor that performs fine-grained control responding to dynamic bandwidth changes, and ii) an RL agent that works at a coarse-grained level that generates control sub-policies for the CC executor. Such two-level control architecture can provide fine-grained DRL-based control with a low model inference overhead. Real-world experiments and emulations show that Spine achieves consistent high performance across various network conditions with an ultra-low control overhead reduced by at least 80% compared to its DRL-based counterparts, similar to classic CC schemes such as Cubic.
Han Tian, Xudong Liao, Chaoliang Zeng, Junxue Zhang 0001, Kai Chen 0005
CoNEXT3
2022 Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load Balancing
Chaoliang Zeng, Layong Luo, Zilong Wang 0007, Wenchen Han, Lebing Wan, Zhipeng Ding, Xiongfei Geng, Feng Ning, Kai Chen 0005, Chuanxiong Guo
NSDI1
2022 FAERY: An FPGA-accelerated Embedding-based Retrieval System
Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Ding Tang, Zilong Wang 0007, Kai Chen 0005, Chuanxiong Guo
OSDI1
2022 LiteFlow: towards high-performance adaptive neural networks for kernel datapath
abstract
Adaptive neural networks (NN) have been used to optimize OS kernel datapath functions because they can achieve superior performance under changing environments. However, how to deploy these NNs remains a challenge. One approach is to deploy these adaptive NNs in the userspace. However, such userspace deployments suffer from either high cross-space communication overhead or low responsiveness, significantly compromising the function performance. On the other hand, pure kernel-space deployments also incur a large performance degradation because the computation logic of model tuning algorithm is typically complex, interfering with the performance of normal datapath execution.
Junxue Zhang 0001, Chaoliang Zeng, Hong Zhang 0025, Shuihai Hu, Kai Chen 0005
SIGCOMM2
2022 Sphinx: Enabling Privacy-Preserving Online Learning over the Cloud
abstract
With the growing complexity of deep learning applications, users have started to delegate their data and models to the cloud. Among these applications, online learning services, which involve both training and inference procedures, are widely deployed. To ensure privacy guarantee on the public cloud, researchers have proposed a plethora of privacy-preserving deep learning algorithms with different techniques, ranging from obfuscation mechanisms to cryptographic tools. However, none of them is applicable to online learning services. They either focus only on inference or training procedure while ignoring the other, or require non-colluding or trusted third parties. In this paper, we present Sphinx, an efficient and privacy-preserving online deep learning system without any trusted third parties. Sphinx strikes a balance between model performance, computational efficiency, and privacy preservation with systematical optimizations on both private inference and training protocols. At its core, Sphinx synthesizes homomorphic encryption and differential privacy reciprocally to maintain the model by keeping most of its parameters as plaintexts, enabling fast training and inference protocol designs. Meanwhile, by refining the homomorphic operation behaviors, Sphinx avoids most of the heavyweight homomorphic operations and minimizes the communication cost. As a result, Sphinx is able to reduce the training time significantly while achieving real-time inference without exposing user privacy. In our experiments, we find that compared to the pure homomorphic encryption solution, Sphinx is $35 \times$ faster for training and 4 orders of magnitude faster for inference, providing real-time inference response (0.05 seconds for MNIST and 0.08 seconds for CIFAR-10). Our experiments also demonstrate that Sphinx achieves promising model accuracy under a tight privacy budget (96% accuracy under $\epsilon=2, \delta=10^{-5}$ for MNIST) without a trusted data aggregator, and is more robust against practical reconstruction attacks.
Han Tian, Chaoliang Zeng, Zhenghang Ren, Di Chai, Junxue Zhang 0001, Kai Chen 0005, Qiang Yang 0001
SP2
2019 Joint Heterogeneous Server Placement and Application Configuration in Edge Computing
abstract
The rapid development of the Internet of Things (IoT) has brought profound changes in the cloud computing paradigm. One promising computing model in IoT-related applications is edge computing, which can decrease the request response time by deploying edge servers close to IoT devices. Two fundamental problems in edge computing include how to place a limited number of edge servers on the candidate locations (e.g., the Access Points) and how to configure the applications on each server. In this paper, we jointly study the edge server placement and the application configuration to minimize the weighted sum of the service cost and edge server opening cost. We propose a local-search based algorithm, named SPAC, to solve the problem efficiently with its approximation ratio analyzed. Extensive simulations on Google data traces demonstrate SPAC reduces the total cost by up to 60% compared with state-of-the-art methods, and outperforms the baselines consistently in different parameter settings.
Jiaying Meng, Chaoliang Zeng, Haisheng Tan, Zining Li, Bojie Li, Xiang-Yang Li 0001
ICPADS2