Wenxue Li 0004

dblp:12/8919-4 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-8228-2552ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 7 first-author · 14 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enabling Packet Spraying over Commodity RNICs with In-Network Support
abstract
AI training workloads exhibit unique traffic patterns that mismatch the ECMP load balancing of RDMA networks, leading to severe throughput degradation. While packet-level load balancing (e.g., random packet spraying, adaptive routing, etc.) offers a promising alternative to ECMP by providing fine-grained traffic distribution, it introduces out-of-order (OOO) packet arrivals. The reliable transport mechanism of current commodity RDMA NICs (RNICs) misinterprets these OOO arrivals as packet loss, causing spurious retransmissions and unnecessary slow starts.
Xiangzhou Liu, Wenxue Li 0004, Kai Chen 0005
EuroSys2
2026 Learn-to-Probe: Achieving Signal Distinguishability in Learning-based Congestion Control
abstract
Internet congestion control remains a fundamental challenge, and recent learning-based congestion control algorithms (CCAs) have shown potential in optimizing network performance. However, their reliance on heuristically chosen input signals often leads to suboptimal behavior across diverse network conditions. In this paper, we identify the root cause as the lack of signal distinguishability-the ability of signals to reflect meaningful differences in network states. To address this, we propose Learn-to-Probe (LTP), a novel signal engineering paradigm that actively generates distinguishable network signals to improve the learning process. LTP (i) employs Bayesian filtering to accurately estimate network states from historical signals, and (ii) introduces an intrinsic reinforcement learning reward that encourages the flow to probe the network, inducing signal sequences that minimize the estimation uncertainty in (i). This probing behavior naturally enhances signal distinguishability, enabling the learning model to make more informed decisions. Extensive evaluations show that LTP consistently achieves high link utilization, low queuing delay, and stable convergence across diverse environments. Our results underscore the importance of signal distinguishability and offer a new direction for robust, adaptive congestion control.
Han Tian, Junxue Zhang 0001, Xudong Liao, Decang Sun, Bin Huang 0024, Wenxue Li 0004, Yong Wang 0046, Kai Chen 0005
EuroSys8
2026 MFS: An Efficient Model Family Serving System for LLMs
abstract
LLM serving providers typically offer a suite of structurally similar models, known as model families, such as the open-source Llama2 series featuring 7B, 13B, and 70B models. While numerous optimizations for LLM serving have been proposed, the potential for leveraging synergies between models within the same family has not been thoroughly explored. This paper introduces MFS, an innovative multi-tiered LLM model family serving system to exploit the structural similarities and parameter redundancies across different scales of models within a family. By utilizing a novel fine-tuning technique called Knowledge Precipitation, MFS restructures the largest model in a family to encapsulate smaller models within its architecture, enabling a unified multi-tiered serving pipeline. Based on the multi-tiered model, MFS realizes a highly parallelized tiered-level batching approach, significantly enhancing system efficiency. It also enables the sharing of intermediate features and KV-cache between models and facilitates multi-level sampling techniques during the inference phase. Experimental results demonstrate that MFS achieves substantial improvements over existing methods, including a 56.1% reduction in end-to-end token generation latency and a 47.8% decrease in GPU memory footprint without compromising the quality of generated content.
Yunxuan Zhang, Hao Wang 0116, Han Tian, Liu Yang 0008, Xudong Liao, Wenxue Li 0004, Ping Yin, Bowen Liu 0002, Kai Chen 0005
EuroSys6
2026 PolicyCache: Intra-flow Learning in Congestion Control
Han Tian, Xudong Liao, Decang Sun, Wenxue Li 0004, Bin Huang 0024, Senbo Fu, Junxue Zhang 0001, Dian Shen, Kai Chen 0005
NSDI6
2026 Toward Fine-Grained Load Balancing With Congested-Flow Isolation in Lossless Datacenters
abstract
Remote Direct Memory Access (RDMA) over Converged Ethernet (RoCE) cooperating with Priority Flow Control (PFC) has been widely deployed in production datacenters to enable low latency, lossless transmission. At the same time, modern datacenters typically offer parallel transmission paths between any pair of end-hosts, underscoring the importance of load balancing. However, the well-studied load balancing mechanisms designed for lossy datacenter networks (DCNs) are ill-suited for such lossless environments. Through extensive experiments, we are among the first to comprehensively inspect the interactions between PFC and load balancing, and uncover that existing fine-grained rerouting schemes can be counterproductive to spread the congested flows among more paths, further aggravating PFC’s head-of-line (HoL) blocking. Motivated by this, we present FLB, a Fine-grained Load Balancing scheme for lossless DCNs. At its core, FLB employs threshold-free rerouting to effectively balance traffic load and improve link utilization during normal conditions and leverages timely congested flow isolation to eliminate HoL blocking on non-congested flows when congestion occurs. To handle complex multi-bottleneck scenarios, we further introduce FLB*, which incorporates an enhanced congestion-point-aware isolation mechanism using Congestion Point Identifiers (CPI) to eliminate HoL blocking among different congested flows.We have fully implemented a FLB prototype, and our evaluation results show that FLB reduces PFC PAUSE rate by up to 96% and avoids HoL blocking, translating to up to 45% improvement in goodput over CONGA+DCQCN and 40%, 36%, 29% and 18% reduction in average flow completion time (FCT) over LetFlow+Swift, MP-RDMA, Proteus+DCQCN and LetFlow+PCN, respectively.
Jinbin Hu 0001, Siyao Li, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Mengyu Ma, Jin Wang 0001, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005
IEEE Trans. Netw.3
2026 Reliable RDMA Over Lossy Fabrics via Data-Control Partitioning
Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Bingyang Liu, Kai Chen 0005
IEEE Trans. Netw.1
2026 High-Performance RoCE-Capable Multicast for Commodity RDMA Datacenters
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g.,$5.2\times $faster multicast communication and$2.7\times $higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
IEEE Trans. Netw.1
2025 Congestion Control for AI Workloads with Message-Level Signaling
Zhenghang Ren, Wenxue Li 0004, Xiangzhou Liu, Kai Chen 0005
APNet3
2025 Enabling Packet Spraying over Commodity RNICs with In-Network Support
Xiangzhou Liu, Wenxue Li 0004, Kai Chen 0005
APNet2
2025 Enabling Efficient GPU Communication over Multiple NICs with FuseLink
Zhenghang Ren, Zilong Wang 0007, Wenxue Li 0004, Kaiqiang Xu, Xudong Liao, Yijun Sun, Bowen Liu 0002, Han Tian, Junxue Zhang 0001, Mingfei Wang, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005
OSDI5
2025 CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data Paths
abstract
Efficient Input/Output (I/O) data path between NICs and CPUs/DRAMs is critical for supporting datacenter applications with high-performance network transmission, especially as link speed scales to 100Gbps and beyond. Traditional I/O acceleration strategies, such as Data Direct I/O (DDIO) and Remote Direct Memory Access (RDMA), perform suboptimally due to the inefficient utilization of the Last-Level Cache (LLC). This paper presents CEIO, a novel cache-efficient network I/O architecture that employs proactive rate control and elastic buffering to achieve zero LLC misses in the I/O data path while ensuring the effectiveness of DDIO and RDMA under various network conditions. We have implemented CEIO on commodity SmartNICs and incorporated it into widely-used DPDK and RDMA libraries. Experiments with well-optimized RPC framework and distributed file system under realistic workloads demonstrate that CEIO achieves up to 2.9× higher throughput and 1.9× lower P99.9 latency over prior work.
Bowen Liu 0002, Qijing Li, Zhuobin Huang, Yijun Sun, Wenxue Li 0004, Junxue Zhang 0001, Ping Yin, Kai Chen 0005
SIGCOMM6
2025 Revisiting RDMA Reliability for Lossy Fabrics
abstract
Due to the high operational complexity and limited deployment scale of lossless RDMA networks, the community has been exploring efficient RDMA communication over lossy fabrics. State-of-the-art (SOTA) lossy RDMA solutions implement a simplified selective repeat mechanism in RDMA NICs (RNICs) to enhance loss recovery efficiency. However, these solutions still face performance challenges, such as unavoidable ECMP hash collisions and excessive retransmission timeouts (RTOs). In this paper, we revisit RDMA reliability with the goals of being independent of PFC, compatible with packet-level load balancing, free from RTO, and friendly to hardware offloading. To this end, we propose DCP, a transport architecture that co-designs both the switch and RNICs, fully meeting the design goals. At its core, DCP-Switch introduces a simple yet effective lossless control plane, which is leveraged by DCP-RNIC to enhance reliability support for high-speed lossy fabrics, primarily including header-only-based retransmission and bitmap-free packet tracking. We prototype DCP-Switch using P4 switch and DCP-RNIC using FPGA. Extensive experiments demonstrate that DCP achieves 1.6× and 2.1× performance improvements, compared to SOTA lossless and lossy RDMA solutions, respectively.
Wenxue Li 0004, Xiangzhou Liu, Yunxuan Zhang, Gaoxiong Zeng, Shoushou Ren, Zhenghang Ren, Bowen Liu 0002, Junxue Zhang 0001, Kai Chen 0005, Bingyang Liu
SIGCOMM1
2025 MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
abstract
Mixture-of-Expert (MoE) models outperform conventional models by selectively activating different subnets, named experts, on a per-token basis. This gated computation generates dynamic communications that cannot be determined beforehand, challenging the existing GPU interconnects that remain static during distributed training. In this paper, we advocate for a first-of-its-kind system, called MixNet, that unlocks topology reconfiguration during distributed MoE training. Towards this vision, we first perform a production measurement study and show that the MoE dynamic communication pattern has strong locality, alleviating the need for global reconfiguration. Based on this, we design and implement a regionally reconfigurable high-bandwidth domain that augments existing electrical interconnects using optical circuit switching (OCS), achieving scalability while maintaining rapid adaptability. We build a fully functional MixNet prototype with commodity hardware and a customized collective communication runtime. Our prototype trains state-of-the-art MoE models with in-training topology reconfiguration across 32 A100 GPUs. Large-scale packet-level simulations show that MixNet achieves performance comparable to a non-blocking fat-tree fabric while boosting the networking cost efficiency (e.g., performance per dollar) of four representative MoE models by 1.2×–1.5× and 1.9×–2.3× at 100 Gbps and 400 Gbps link bandwidths, respectively.
Xudong Liao, Yijun Sun, Han Tian, Xinchen Wan, Yilun Jin, Zilong Wang 0007, Zhenghang Ren, Wenxue Li 0004, Kin Fai Tse, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Xiaofeng Ye, Yiming Zhang 0003, Kai Chen 0005
SIGCOMM9
2025 FLB: Fine-grained Load Balancing for Lossless Datacenter Networks
Jinbin Hu 0001, Wenxue Li 0004, Xiangzhou Liu, Bowen Liu 0002, Ping Yin, Jianxin Wang 0001, Jiawei Huang 0001, Kai Chen 0005
USENIX ATC2
2024 Understanding Communication Characteristics of Distributed Training
abstract
Communication is pivotal in distributed training and a thorough understanding of its characteristics is essential for future optimizations. However, prior works are limited, either focusing on customized optimizations or conducting incomplete explorations on communication characteristics. In this work, we systematically analyze the communication characteristics of distributed training, considering two key aspects of communication: pattern and overhead, and assessing a broad spectrum of determinant factors. In particular, we extensively investigate the features of communication patterns, such as predictability, and comprehensively evaluate the impact of various factors on communication overhead. Additionally, we develop and validate an analytical formulation to estimate communication overhead, providing a mathematical understanding of models with predictability.
Wenxue Li 0004, Xiangzhou Liu, Yilun Jin, Han Tian, Zhizhen Zhong, Guyue Liu, Ying Zhang 0022, Kai Chen 0005
APNet1
2024 Cepheus: Accelerating Datacenter Applications with High-Performance RoCE-Capable Multicast
abstract
Modern datacenter applications widely exhibit multicast communication patterns. Meanwhile, RDMA is emerging as the de-facto networking architecture to meet the stringent performance requirement of applications. However, existing multicast approaches fail to efficiently collaborate multicast with commodity RDMA transport, either causing inefficient multicast traffic transmission or being trapped in the insufficient end-host transport protocol. In this paper, we propose Cepheus, which delivers performance gains from both multicast (i.e., traffic reduction and transmission hop minimization) and RDMA transport (i.e., ultra-low latency, high throughput and low CPU overhead). Cepheus reuses RoCE as its transport layer and provides a RoCE-capable multicast primitive via in-network assistance. At its core, Cepheus builds on and goes beyond the native multicast architecture by exploiting more switch functionalities to tackle the incompatibilities between multicast flow structure and RoCE processing logic. We prototype Cepheus on an FPGA board, as a building block attached to an Ethernet switch. Extensive experiments demonstrate Cepheus inter-operates with commodity RoCE protocol and outperforms existing RDMA multicast schemes, e.g., 5.2 × faster multicast communication and 2.7 × higher replication throughput for distributed storage.
Wenxue Li 0004, Junyi Zhang 0005, Gaoxiong Zeng, Zilong Wang 0007, Chaoliang Zeng, Pengpeng Zhou, Qiaoling Wang, Kai Chen 0005
HPCA1
2024 FlowSail: Fine-Grained and Practical Flow Control for Datacenter Networks
abstract
As datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effectiveness due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce FlowSail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of FlowSail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FlowSail using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that FlowSail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FlowSail achieves$4.3\times $lower tail latency under the same number of queues, matches existing FC schemes with$4\times $fewer queues, and holds robust performance with a minimum of 2 queues.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
IEEE/ACM Trans. Netw.1
2023 Scaling Switch-driven Flow Control with Aquarius
abstract
As datacenter networks support more diverse applications and faster link speeds, effective end-to-end congestion control becomes increasingly challenging due to the inherent feedback delay. To address this issue, switch-driven per-hop flow control (FC) has gained popularity due to its natural flow isolation, timely control loop, and ability to handle transient congestion. However, the ideal FC requires impractical hardware resources, and the state-of-the-art approximation approach still demands a large number of queues that exceeds common switch capabilities, limiting scalability in practice.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
APNet1
2023 Towards Fine-Grained and Practical Flow Control for Datacenter Networks
abstract
As datacenter networks continue to support a wider range of applications and faster link speeds, they face the challenge of managing bursty traffic and transient congestion. End-to-end congestion controls (CCs) find it increasingly difficult to maintain effective due to the inherent feedback delay. To address this issue, per-hop flow control (FC) has gained popularity due to its ability to react promptly to transient congestion. However, existing FC mechanisms either lack fine-grained (i.e., per-flow granularity) control or require an impractical number of queues that exceeds the capabilities of commodity switches. In this paper, we introduce Flowsail, an innovative FC scheme that enables fine-grained control at the per-flow level while requiring a practical number of switch queues, theoretically as few as two. The core of Flowsail is an effective approximation of ideal FC by three key design components: dynamic flow-to-queue mapping, hierarchical congested flow identification, and on-demand isolation. We have implemented a prototype of FLOWSAIL using the programmable P4 switch and conducted extensive testbed experiments and simulations. The results indicate that Flowsail effectively sustains performance with significantly fewer queues compared to existing FC schemes. For instance, FLOWSAIL achieves 4.3 x lower tail latency under the same number of queues, matches existing FC schemes with 4 x fewer queues, and holds robust performance with a minimum of 2 queues.
Wenxue Li 0004, Chaoliang Zeng, Jinbin Hu 0001, Kai Chen 0005
ICNP1
2023 SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo
NSDI5