Yuanwei Lu

dblp:182/6479 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0006-2554-6490ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 3 first-author · 7 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Accelerating Hardware/Software Combined Traffic Processing With Fast and Efficient Asynchronous Flow Offloading
abstract
eHardware/software (hw/sw) combined systems are necessary to meet modern clouds’ requirements for processing huge amounts of network traffic by efficiently offloading large flows to hardware. However, existing hw/sw flow offloading systems typically perform traffic statistics collection and large flow selection within a time window—a time-window-based approach. Their offloading decision of large flows issynchronizedin the unit of a time window, which is mismatched to the asynchronous and dynamic nature of each flow’s sending rate. Additionally, the flow measurement and selection for large flows are decoupled in these solutions, leading to memory and CPU inefficiency. In this paper, we introduce TAO, a novel solution to the hw/sw combined flow offloading problem byasynchronouslyselecting and offloading flows based on flow table entries. TAO can reactfasterto the rapid dynamics of flows by taking actions at each table entry and ismore efficientby coupling flow measurement and selection into the entry. We have implemented a full-fledged TAO prototype based on the P4 switch and DPDK. Testbed results demonstrate that TAO can offload ∼16% more traffic to hardware, outperforming existing solutions by achieving 42× lower memory overhead. Meanwhile, it reduces software CPU utilization by 66.7% and cuts tail forwarding latency by 95.59% compared to state-of-the-art methods.
Xijin Yin, Yuanwei Lu, Xin Zhang 0117, Xingtong Lin, Shengli Zheng, Bangwen Deng, Xianneng Zou, Yachen Wang, Guo Chen 0001
IEEE Trans. Netw.2
2025 Fast and Scalable Selective Retransmission for RDMA
Peihao Huang, Guo Chen 0001, Xin Zhang 0117, Huijun Shen, Ying Bian, Yuanwei Lu, Zhenyuan Ruan, Bojie Li, Jiansong Zhang 0001, Yongfeng Liu, Zhigang Chen 0001
INFOCOM8
2025 InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
abstract
Scaling Large Language Model (LLM) training relies on multidimensional parallelism, where High-Bandwidth Domains (HBDs) are critical for communication-intensive parallelism like Tensor Parallelism. However, existing HBD architectures face fundamental limitations in scalability, cost, and fault resiliency: switch-centric HBDs (e.g., NVL-72) incur prohibitive scaling costs, while GPU-centric HBDs (e.g., TPUv3/Dojo) suffer from severe fault propagation. Switch-GPU hybrid HBDs (e.g., TPUv4) take a middle-ground approach, but the fault explosion radius remains large.
Chenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng, Yu Zhou 0008, Wenqing Lv, Yelong Xu, Yuanwei Lu, Yanbo Yu, Yichen Shen 0001, Yibo Zhu 0001, Daxin Jiang
SIGCOMM9
2023 Fast, Scalable and Robust Centralized Routing for Data Center Networks
abstract
This paper presents a fast and robust centralized data center network (DCN) routing solution, called . For fast routing calculation, uses centralized controllers to collect/disseminate the network’s link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch’s routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate with extensive experiments on Linux-machine controllers and white-box switches. provides$\sim$1200x and$\sim$100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively. Furthermore, Primus maintains good routing controllability/manageability thanks to its centralized architecture, which enables us to build several advanced routing features in our testbed, including routing failure visualization and weighted-cost-multi-path routing.
Fusheng Lin, Guo Chen 0001, Guihua Zhou, Dehui Wei, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001
IEEE/ACM Trans. Netw.8
2022 Elixir: A High-performance and Low-cost Approach to Managing Hardware/Software Hybrid Flow Tables Considering Flow Burstiness
Yanshu Wang, Dan Li 0001, Yuanwei Lu
NSDI3
2021 Primus: Fast and Robust Centralized Routing for Large-scale Data Center Networks
abstract
This paper presents a fast and robust centralized data center network (DCN) routing solution called Primus. For fast routing calculation, Primus uses centralized controller to collect/disseminates the network's link-states (LS), and offload the actual routing calculation onto each switch. Observing that the routing changes can be classified into a few fixed patterns in DCNs which have regular topologies, we simplify each switch's routing calculation into a table-lookup manner, i.e., comparing LS changes with pre-installed base topology and updating routing paths according to predefined rules. As such, the routing calculation time at each switch only needs 10s of us even in a large network topology containing 10K+ switches. For efficient controller fault-tolerance, Primus purposely uses reporter switch to ensure the LS updates successfully delivered to all affected switches. As such, Primus can use multiple stateless controllers and little redundant traffic to tolerate failures, which incurs little overhead under normal case, and keeps 10s of ms fast routing reaction time even under complex data-/control-plane failures. We design, implement and evaluate Primus with extensive experiments on Linux-machine controllers and white-box switches. Primus provides ~1200x and ~100x shorter convergence time than current distributed protocol BGP and the state-of-the-art centralized routing solution, respectively.
Guihua Zhou, Guo Chen 0001, Fusheng Lin, Dehui Wei, Jianbing Wu, Li Chen 0008, Yuanwei Lu, Andrew Qu, Hongbo Jiang 0001
INFOCOM8
2021 Accessing Cloud with Disaggregated Software-Defined Router
Xiaoliang Wang 0001, Yuanwei Lu, Yanbo Yu, Shengli Zheng, Youjian Zhao
NSDI3
2019 MP-RDMA: Enabling RDMA With Multi-Path Transport in Datacenters
abstract
RDMA is becoming prevalent because of its low latency, high throughput and low CPU overhead. However, in current datacenters, RDMA remains a single path transport which is prone to failures and falls short to utilize the rich parallel network paths. Unlike previous multi-path approaches, which mainly focus on TCP, this paper presents a multi-path transport for RDMA, i.e. MP-RDMA, which efficiently utilizes the rich network paths in datacenters. MP-RDMA employs three novel techniques to address the challenge of limited RDMA NICs on-chip memory size: 1) a multi-path ACK-clocking mechanism to distribute traffic in a congestion-aware manner without incurring per-path states; 2) an out-of-order aware path selection mechanism to control the level of out-of-order delivered packets, thus minimizes the meta data required to them; 3) a synchronise mechanism to ensure in-order memory update whenever needed. With all these techniques, MP-RDMA only adds 66B to each connection state compared to single-path RDMA. Our evaluation with an FPGA-based prototype demonstrates that compared with single-path RDMA, MP-RDMA can significantly improve the robustness under failures ( $2\times \sim 4\times $ higher throughput under 0.5%~10% link loss ratio) and improve the overall network utilization by up to 47%.
Guo Chen 0001, Yuanwei Lu, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Thomas Moscibroda
IEEE/ACM Trans. Netw.2
2018 Multi-Path Transport for RDMA in Datacenters
Yuanwei Lu, Guo Chen 0001, Bojie Li, Kun Tan 0002, Yongqiang Xiong, Peng Cheng 0005, Jiansong Zhang 0001, Enhong Chen, Thomas Moscibroda
NSDI1
2018 FUSO: Fast Multi-Path Loss Recovery for Data Center Networks
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
IEEE/ACM Trans. Netw.2
2017 Memory Efficient Loss Recovery for Hardware-based Transport in Datacenter
abstract
Limited by the small on-chip memory, hardware-based transport typically implements go-back-N loss recovery mechanism, which costs very few memory but is well-known to perform inferior even under small packet loss ratio. We present MELO, an efficient selective retransmission mechanism for hardware-based transport, which consumes only a constant small memory regardless of the number of concurrent connections. Specifically, MELO employs an architectural separation between data and meta data storage and uses a shared bits pool allocation mechanism to reduce meta data on-chip memory footprint. By only adding in average 23B extra on-chip states for each connection, MELO achieves up to 14.02x throughput while reduces 99% tail FCT by 3.11x compared with go-back-N under certain loss ratio.
Yuanwei Lu, Guo Chen 0001, Zhenyuan Ruan, Wencong Xiao, Bojie Li, Jiansong Zhang 0001, Yongqiang Xiong, Peng Cheng 0005, Enhong Chen
APNet1
2017 One more queue is enough: Minimizing flow completion time with explicit priority notification
abstract
Ideally, minimizing the flow completion time (FCT) requires millions of priorities supported by the underlying network so that each flow has its unique priority. However, in production datacenters, the available switch priority queues for flow scheduling are very limited (merely 2 or 3). This practical constraint seriously degrades the performance of previous approaches. In this paper, we introduce Explicit Priority Notification (EPN), a novel scheduling mechanism which emulates fine-grained priorities (i.e., desired priorities or DP) using only two switch priority queues. EPN can support various flow scheduling disciplines with or without flow size information. We have implemented EPN on commodity switches and evaluated its performance with both testbed experiments and extensive simulations. Our results show that, with flow size information, EPN achieves comparable FCT as pFabric that requires clean-slate switch hardware. And EPN also outperforms TCP by up to 60.5% if it bins the traffic into two priority queues according to flow size. In information-agnostic setting, EPN outperforms PIAS with two priority queues by up to 37.7%. To the best of our knowledge, EPN is the first system that provides millions of priorities for flow scheduling with commodity switches.
Yuanwei Lu, Guo Chen 0001, Larry Luo, Kun Tan 0002, Yongqiang Xiong, Xiaoliang Wang 0001, Enhong Chen
INFOCOM1
2017 KV-Direct: High-Performance In-Memory Key-Value Store with Programmable NIC
abstract
Performance of in-memory key-value store (KVS) continues to be of great importance as modern KVS goes beyond the traditional object-caching workload and becomes a key infrastructure to support distributed main-memory computation in data centers. Recent years have witnessed a rapid increase of network bandwidth in data centers, shifting the bottleneck of most KVS from the network to the CPU. RDMA-capable NIC partly alleviates the problem, but the primitives provided by RDMA abstraction are rather limited. Meanwhile, programmable NICs become available in data centers, enabling in-network processing. In this paper, we present KV-Direct, a high performance KVS that leverages programmable NIC to extend RDMA primitives and enable remote direct key-value access to the main host memory.
Bojie Li, Zhenyuan Ruan, Wencong Xiao, Yuanwei Lu, Yongqiang Xiong, Andrew Putnam, Enhong Chen
SOSP4
2016 Fast and Cautious: Leveraging Multi-path Diversity for Transport Loss Recovery in Data Centers
Guo Chen 0001, Yuanwei Lu, Yuan Meng 0002, Bojie Li, Kun Tan 0002, Dan Pei, Peng Cheng 0005, Layong Luo, Yongqiang Xiong, Xiaoliang Wang 0001, Youjian Zhao
USENIX ATC2