VLDB 2026 Research / reviewers in the wild / expert
Xuandong Liu
dblp:266/7436
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0002-6521-6200ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource PoolingabstractThe increasing popularity of large models and datasets has highlighted the significance of distributed training networks. As gradient synchronization generates substantial traffic, in-network aggregation (INA) has emerged as a solution to offload aggregation onto the switch, alleviating network congestion and accelerating distributed training. However, the limited memory capacity of the INA switch becomes a potential bottleneck as computation shifts into the network, especially in multi-tenant scenarios. To address this bottleneck and enhance network throughput, we propose the Aggregation with Innetwork Resource Pooling (AIRP) framework. Unlike existing approaches that optimize individual switches in a localized manner, AIRP takes a holistic view and efficiently pools switch memory resources across the entire network, allocating them to multiple tenants. Evaluation using the ns-3 simulator and P4 testbed demonstrates that AIRP can accelerate the training of various models, including computer vision and language models. The experimental results show that AIRP outperforms existing INA approaches by up to 7 times in terms of network throughput in multi-tenant scenarios, while also achieving great flexibility and efficiency in deployment. Huifeng Xing, Hao Wang 0231, Yang Chen 0001, Yinfan Hu, Xuandong Liu, Zijian Li 0003, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
IEEE Trans. Netw. | 6 |
| 2024 | Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model TrainingabstractParameter Server (PS) and Ring-AllReduce (RAR) are two widely utilized synchronization architectures in multiworker Deep Learning (DL), also referred to as Distributed Deep Learning (DDL). However, PS encounters challenges with the “incast” issue, while RAR struggles with problems caused by the long dependency chain. The emerging In-network Aggregation (INA) has been proposed to integrate with PS to mitigate its incast issue. However, such PS-based INA has poor incremental deployment abilities as it requires replacing all the switches to show significant performance improvement, which is not costeffective. In this study, we present the incorporation of INA capabilities into RAR, called RAR with In-Network Aggregation (Rina), to tackle both the problems above. Rina features its agent-worker mechanism. When an INA-capable ToR switch is deployed, all workers in this rack run as one abstracted worker with the help of the agent, resulting in both excellent incremental deployment capabilities and better throughput. We conducted extensive testbed and simulation evaluations to substantiate the throughput advantages of Rina over existing DDL training synchronization structures. Compared with the state-of-the-art PS-based INA methods ATP, Rina can achieve more than$\mathbf{5 0 \%}$throughput with the same hardware cost. Xuandong Liu, Minglin Li, Yinfan Hu, Huifeng Xing, Hao Wang 0231, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
ICNP | 2 |
| 2023 | OSP: Boosting Distributed Model Training with 2-stage SynchronizationabstractDistributed deep learning (DDL) is a promising research area, which aims to increase the efficiency of training deep learning tasks with large size of datasets and models. As the computation capability of DDL nodes continues to increase, the network connection between nodes is becoming a major bottleneck. Various methods of gradient compression and improved model synchronization have been proposed to address this bottleneck in Parameter-Server-based DDL. However, these two types of methods can result in accuracy loss due to discarded gradients and have limited enhancement on the throughput of model synchronization, respectively. To address these challenges, we propose a new model synchronization method named Overlapped Synchronization Parallel (OSP), which achieves efficient communication with a 2-stage synchronization approach and uses Local-Gradient-based Parameter correction (LGP) to avoid accuracy loss caused by stale parameters. The prototype of OSP has been implemented using PyTorch and evaluated on commonly used deep learning models and datasets with a 9-node testbed. Evaluation results show that OSP can achieve up to 50% improvement in throughput without accuracy loss compared to popular synchronization models. Lei Shi 0031, Xuandong Liu, Sen Liu 0002, Yang Xu 0010 |
ICPP | 3 |
| 2023 | Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission ProtocolabstractDistributed Machine Learning (DML) systems are utilized to enhance the speed of model training in data centers (DCs) and edge nodes. The Parameter Server (PS) communication architecture is commonly employed, but it faces severe long-tail latency caused by many-to-one “incast” traffic patterns, negatively impacting training throughput. To address this challenge, we design the Loss-tolerant Transmission Protocol (LTP), which permits partial loss of gradients during synchronization to avoid unneeded retransmission and contributes to faster synchronization per iteration. LTP implements loss-tolerant transmission through out-of-order transmission and out-of-order Acknowledges (ACKs). LTP employs Early Close to adjust the loss-tolerant threshold based on network conditions and bubble-filling for data correction to maintain training accuracy. LTP is implemented by C++ and integrated into PyTorch. Evaluations on a testbed of 8 worker nodes and one PS node demonstrate that LTP can significantly improve DML training task throughput by up to 30x compared to traditional TCP congestion controls, with no sacrifice to final accuracy. Lei Shi 0031, Xuandong Liu, Xin Ai 0008, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 3 |
| 2021 | MagicTCAM: A Multiple-TCAM Scheme for Fast TCAM UpdateabstractTernary Content-Addressable Memory (TCAM) is a popular solution for high-speed flow table lookup in Software-Defined Networking (SDN). Rule insertion in TCAM is a time-consuming operation. To ensure semantic correctness, rules overlapped must be stored in TCAM with decreasing priority order and many rule movements may be needed to make space for a single inserted rule. When a rule insertion is in progress, the regular flow table lookup will be suspended, which could lead to a degraded user experience for SDN applications. In this paper, we propose a multiple-TCAM framework named MagicTCAM to reduce the rule movements during a rule insertion. The core of MagicTCAM lies in three operations: layering, partitioning and rotating. By layering, rules with the least overlapping will be grouped (i.e., layered) into a sub-ruleset. The number of rule movements is therefore greatly reduced as most of rules in a sub-ruleset are non-overlapped. To achieve balanced load in TCAMs, rules in each sub-ruleset are further partitioned and dispatched into different TCAMs in a rotating manner. In addition, an inter-TCAM movement algorithm is proposed to allow rules to be moved between TCAMs for reduced rule movement. Experiment results show that with two half-sized TCAMs, MagicTCAM reduces the rule movements by 39% on average compared with the state-of-the-art work while the computation time is shortened by half as well. Ruyi Yao, Xuandong Liu, Ying Wan 0001, Bin Liu 0001, Wenjun Li 0004, Yang Xu 0010 |
ICNP | 3 |