Zilong Wang 0007

dblp:42/898-7 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-3184-4081ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 3 first-author · 11 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 High-Performance RoCE-Capable Multicast for Commodity RDMA Datacenters
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
IEEE Trans. Netw.5
2025 Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud
abstract
FPGAs are gaining popularity in the cloud as accelerators for various applications. To make FPGAs more accessible for users and streamline system management, cloud providers have widely adopted the shell-role architecture on their homogeneous FPGA servers. However, the increasing heterogeneity of cloud FPGAs poses new challenges for this architecture. Previous studies either focus on homogeneous FPGAs or only partially address the portability issues for roles, while still requiring laborious shell development for providers and ad-hoc software modifications for users.
Xinchen Wan, Zilong Wang 0007, Qian Zhao 0001, Feng Ning, Qingsong Ning, Shideng Zhang, Zhenyu Li 0001, Layong Luo, Gaogang Xie
ASPLOS (2)5
2025 Design and Operation of Shared Machine Learning Clusters on Campus
abstract
The rapid advancement of large machine learning (ML) models has driven universities worldwide to invest heavily in GPU clusters. Effectively sharing these resources among multiple users is essential for maximizing both utilization and accessibility. However, managing shared GPU clusters presents significant challenges, ranging from system configuration to fair resource allocation among users. This paper introduces SING, a full-stack solution tailored to simplify shared GPU cluster management. Aimed at addressing the pressing need for efficient resource sharing with limited staffing, SING enhances operational efficiency by reducing maintenance costs and optimizing resource utilization. We provide a comprehensive overview of its four extensible architectural layers, explore the features of each layer, and share insights from real-world deployment, including usage patterns and incident management strategies. As part of our commitment to advancing shared ML cluster management, we open-source SING's resources to support the development and operation of similar systems.
Kaiqiang Xu, Decang Sun, Hao Wang 0116, Zhenghang Ren, Xinchen Wan, Xudong Liao, Zilong Wang 0007, Junxue Zhang 0001, Kai Chen 0005
ASPLOS (1)7
2025 Achieving Fairness Generalizability for Learning-based Congestion Control with Jury
abstract
Internet congestion control (CC) has long posed a challenging control problem in networking systems, with recent approaches increasingly incorporating deep reinforcement learning (DRL) to enhance adaptability and performance. Despite promising, DRL-based CC schemes often suffer from poor fairness, particularly when applied to network environments unseen during training. This paper introduces Jury, a novel DRL-based CC scheme designed to achieve fairness generalizability. At its heart, Jury decouples the fairness control from the principal DRL model with two design elements: i) By transforming network signals, it provides a universal view of network environments among competing flows, and ii) It adopts a post-processing phase to dynamically module the sending rate based on flow bandwidth occupancy estimation, ensuring large flows behave more conservatively and smaller flows more aggressively, thus achieving a fair and balanced bandwidth allocation. We have fully implemented Jury, and extensive evaluations demonstrate its robust convergence properties and high performance across a broad spectrum of both emulated and real-world network conditions.
Han Tian, Xudong Liao, Decang Sun, Chaoliang Zeng, Yilun Jin, Junxue Zhang 0001, Xinchen Wan, Zilong Wang 0007, Yong Wang 0046, Kai Chen 0005
EuroSys8
2025 A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Han Tian, Xudong Liao, Chaoliang Zeng, Zilong Wang 0007, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen 0005
INFOCOM7
2025 Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005
OSDI3
2025 MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
abstract
Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively.
Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005
SIGCOMM6
2024 Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable Multicast
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
HPCA5
2024 Fast, Scalable, and Accurate Rate Limiter for RDMA NICs
abstract
RDMA NICs desire a rate limiter that is accurate, scalable, and fast: to precisely enforce the policies such as congestion control and traffic isolation, to support a large number of flows, and to sustain high packet rates. Prior works such as SENIC and PIEO can achieve accuracy and scalability, but they are not fast enough, thus fail to fulfill the performance requirement of RNICs, due primarily to their monolithic design and one-packet-per-sorting transmission. We present Tassel, a hierarchical rate limiter for RDMA NICs that can deliver high packet rates by enabling multiple-packet-per-sorting transmission, while preserving accuracy and scalability. At its heart, Tassel renovates the workflow of the rate limiter hierarchically: by first applying scalable rate limiting to the flows to be scheduled, followed by accurate rate limiting to the packets to be transmitted, while leveraging adaptive batching and packet filtering to improve the performance of these two steps. We integrate Tassel into the RNIC architecture by replacing the original QP scheduler module and implement the prototype of Tassel using FPGA. Experimental results show that Tassel delivers 125 Mpps packet rate, outperforming SENIC and PIEO by 3.6×, while supporting 16 K flows with low resource usage, 7.5% - 25.6% as compared to SENIC and PIEO, and preserving high accuracy, precisely enforcing rate limits from 100 Kbps to 100 Gbps.
Zilong Wang 0007, Xinchen Wan, Yijun Sun, Qingsong Ning, Junxue Zhang 0001, Kai Chen 0005
SIGCOMM1
2024 Accelerating Secure Collaborative Machine Learning with Protocol-Aware RDMA
Zhenghang Ren, Mingxuan Fan, Zilong Wang 0007, Junxue Zhang 0001, Chaoliang Zeng, Cheng Hong 0001, Kai Chen 0005
USENIX Security Symposium3
2024 Load Balancing With Multi-Level Signals for Lossless Datacenter Networks
abstract
Various datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-level signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35%, 31%, 28%, 22% and 46%, 42%, 34%, 29% better average FCT and$99^{th}$percentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
IEEE/ACM Trans. Netw.3
2023 Accurate and Scalable Rate Limiter for RDMA NICs
abstract
Rate limiter is required by RDMA NIC (RNIC) to enforce the rate limits calculated by congestion control. RNIC expects the rate limiter to be accurate and scalable: to precisely shape the traffic for numerous flows with minimized resource consumption, thereby mitigating the incasts and congestions and improving the network performance. Previous works, however, fail to meet the performance requirements of RNIC while achieving accuracy and scalability.
Zilong Wang 0007, Xinchen Wan, Chaoliang Zeng, Kai Chen 0005
APNet1
2023 Enabling Load Balancing for Lossless Datacenters
abstract
Various datacenter network (DCN) load balancing schemes have been proposed in the past decade. Unfortunately, most of these solutions designed for lossy DCNs do not work well for Priority Flow Control (PFC) enabled lossless DCNs, primarily due to the reason that the individual congestion signals used in these solutions, e.g., link load, queue length, Round Trip Time (RTT) and Explicit Congestion Notification (ECN), may not be able to correctly or timely reflect the hop-by-hop PFC pausing. This paper first reveals the above problems via extensive experiments, and then based on the insights learned, we present Proteus, a PFC-aware load balancing scheme that is resilient to PFC pausing by exploring a combination of multi-level congestion signals. At its heart, Proteus leverages RTT-Ievel signals (i.e., RTT and link utilization) to detect path status for initial routing decision, and exploits sub-RTT level signal (i.e., cumulative sojourn time) to reflect instantaneous PFC pausing and make timely rerouting choices based on the idea of better-late-than-never. We have implemented Proteus in the hardware programmable switch. Our testbed experiments as well as large-scale simulations show that Proteus can effectively handle PFC pausing under realistic workloads and achieve up to 35 %, 31 %, 28%, 22% and 46 %, 42 %, 34 %, 29 % better average FCT and 99thpercentile FCT than CONGA, DRILL, Hermes and MP-RDMA, respectively.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Junxue Zhang 0001, Kun Guo 0003, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
ICNP3
2023 SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo
NSDI1
2022 Load Balancing in PFC-Enabled Datacenter Networks
abstract
In Priority Flow Control (PFC) enabled datacenter networks (DCNs), PFC is inevitably triggered due to bursty traffic even with end-to-end congestion control. Load balancing as a complementary mechanism to transport protocols can make rerouting decisions in time to alleviate PFC’s head-of-line (HoL) blocking problem. However, prior solutions designed for lossy DCNs do not work well in PFC-enabled networks, because the unreliable rerouting signals such as separate local queue length, round-trip time (RTT), explicit congestion notification (ECN), and link load cannot timely and correctly reflect PFC pausing.
Jinbin Hu 0001, Chaoliang Zeng, Zilong Wang 0007, Hong Xu 0001, Jiawei Huang 0001, Kai Chen 0005
APNet3
2022 Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load Balancing
Chaoliang Zeng, Layong Luo, Zilong Wang 0007, Wenchen Han, Lebing Wan, Zhipeng Ding, Xiongfei Geng, Feng Ning, Kai Chen 0005, Chuanxiong Guo
NSDI4
2022 FAERY: An FPGA-accelerated Embedding-based Retrieval System
Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Ding Tang, Zilong Wang 0007, Kai Chen 0005, Chuanxiong Guo
OSDI7
2022 Aeolus: A Building Block for Proactive Transport in Datacenter Networks
abstract
As datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth isproactivelyallocated as “credits” to senders who then can send “scheduled packets” at a right rate to ensure high link utilization, low latency, and zero packet loss. Consequently, proactive solutions such as ExpressPass, NDP, Homa, etc., have been proposed recently. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT “pre-credit” phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on “pre-credit” packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. Aeolus is compatible with all existing proactive solutions and readily implementable with commodity switches. We have integrated Aeolus into ExpressPass, NDP and Homa, and shown, via both implementation and simulations, that the Aeolus-enhanced solutions deliver significant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by$20\times $, while achieving similar performance as NDP without switch modifications.
Shuihai Hu, Gaoxiong Zeng, Wei Bai 0001, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004
IEEE/ACM Trans. Netw.4
2020 Aeolus: A Building Block for Proactive Transport in Datacenters
abstract
As datacenter network bandwidth keeps growing, proactive transport becomes attractive, where bandwidth is proactively allocated as "credits" to senders who then can send "scheduled packets" at a right rate to ensure high link utilization, low latency, and zero packet loss. While promising, a fundamental challenge is that proactive transport requires at least one-RTT for credits to be computed and delivered. In this paper, we show such one-RTT "pre-credit" phase could carry a substantial amount of flows at high link-speeds, but none of existing proactive solutions treats it appropriately. We present Aeolus, a solution focusing on "pre-credit" packet transmission as a building block for proactive transports. Aeolus contains unconventional design principles such as scheduled-packet-first (SPF) that de-prioritizes the first-RTT packets, instead of prioritizing them as prior work. It further exploits the preserved, deterministic nature of proactive transport as a means to recover lost first-RTT packets efficiently. We have integrated Aeolus into ExpressPass[14], NDP[18] and Homa[29], and shown, through both implementation and simulations, that the Aeolus-enhanced solutions deliver signiicant performance or deployability advantages. For example, it improves the average FCT of ExpressPass by 56%, cuts the tail FCT of Homa by 20x, while achieving similar performance as NDP without switch modifications.
Shuihai Hu, Wei Bai 0001, Gaoxiong Zeng, Zilong Wang 0007, Baochen Qiao, Kai Chen 0005, Kun Tan 0002, Yi Wang 0004
SIGCOMM4