Shuo Liu 0002

dblp:07/6773-2 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-3546-4070ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2024 Training Job Placement in Clusters with Statistical In-Network Aggregation
abstract
In-Network Aggregation (INA) offloads the gradient aggregation in distributed training (DT) onto programmable switches, where the switch memory could be allocated to jobs in either synchronous or statistical multiplexing mode. Statistical INA has advantages in switch memory utilization, control-plane simplicity, and management safety, but it faces the problem of cross-layer resource efficiency in job placement. This paper presents a job placement system NetPack for clusters with statistical INA, which aims to maximize the utilization of both computation and network resources. NetPack periodically batches and places jobs into the cluster. When placing a job, NetPack runs a steady state estimation algorithm to acquire the available resources in the cluster, heuristically values each server according to its available resources (GPU and bandwidth), and runs a dynamic programming algorithm to efficiently search for servers with the highest value for the job. Our prototype of NetPack and the experiments demonstrate that NetPack outperforms prior job placement methods by 45% in terms of average job completion time on production traces.
Bohan Zhao, Wei Xu 0005, Shuo Liu 0002, Yang Tian 0012, Qiaoling Wang, Wenfei Wu
ASPLOS (1)3
2023 In-Network Aggregation with Transport Transparency for Distributed Training
abstract
Recent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation.
Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He
ASPLOS (3)1
2021 Scalable Fully Pipelined Hardware Architecture for In-Network Aggregated AllReduce Communication
abstract
The Ring-AllReduce framework is currently the most popular solution to deploy industry-level distributed machine learning tasks. However, only about half of the maximum bandwidth can be achieved in the optimal condition. In recent years, several in-network aggregation frameworks have been proposed to overcome the drawback, but limited hardware information have been disclosed. In this paper, we propose a scalable fully-pipelined architecture that handles tasks like forwarding, aggregation and retransmission with no bandwidth loss. The architecture is implemented on a Xilinx Ultrascale FPGA that connects to 8 working servers with 10 Gb/s network adapters, and it is able to scale to more complicated scenarios involving more workers. Compared with Ring-AllReduce, using AllReduce-Switch improves the efficient bandwidth of AllReduce communication with a ratio of$1.75\times $. In image training tasks, the proposed hardware architecture helps to achieve up to$1.67\times $speedup to the training process. For computing-intensive models, the speedup from communication may be partially hidden by computing. In particular, for ResNet-50, AllReduce-Switch improves the training process with MPI and NCCL by$1.30\times $and$1.04\times $respectively.
Yao Liu 0006, Junyi Zhang 0005, Shuo Liu 0002, Qiaoling Wang, Wangchen Dai, Ray C. C. Cheung
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 Adaptive Coherent Sampling for Network Delay Measurement
abstract
End-to-end network delay, as a metric to indicate the QoS (Quality-of-Service), plays an important role in distributed services. Unfortunately, it is infeasible in practice to know all node-pair delay information due to the quadratic growth of overhead by active probing. In this paper, we leverage the stateof-the-art matrix completion technology for better network delay estimation from limited measurements. Although the number of samples required for exact matrix completion is theoretically bounded, it is practically less helpful as the number cannot be specified. This motivates us to propose an adaptive coherent sampling algorithm to select the elements with larger leverage scores to maintain the characteristic of important rows or columns in the delay matrix. The number of samples is adaptively determined by a proposed stopping criterion. Simulation results based on real-world network delay datasets indicate that our proposed algorithm is capable of providing better performance (improves estimation error by 16.9% and convergence stress by 28.9%) at less cost (reduces number of samples by 3.9% and processing time by 78.6%) than traditionally used algorithms.
Shuo Liu 0002, Qiaoling Wang
ICC1
2019 Poster: online adaptive sampling for network delay measurement via matrix completion
abstract
End-to-end network delay plays an important role in distributed services. Unfortunately, it is infeasible to know all node-pair delay information in practice due to the quadratic growth of overhead by active probing. In this paper, we leverage the state-of-the-art matrix completion technology for better network delay estimation from limited measurements. Specifically, we formulate the matrix completion problem as a nuclear norm minimization problem which can be solved via convex optimization. We propose an online adaptive sampling strategy for network delay measurement. The key idea is to sample the elements with larger leverage scores to maintain characteristic of important rows or columns of the matrix. The number of samples is adaptively determined by a proposed stopping criterion. A preliminary simulation result based on real-world network delay datasets indicates that our proposed sampling algorithm is capable of providing better performance (smaller estimation error and less convergence stress) at less cost (fewer samples and shorter processing time) than other traditionally used algorithms.
Shuo Liu 0002, Qiaoling Wang
Networking1
2014 Improved indoor tracking based on generalized t-distribution noise model
abstract
The use of wireless sensor networks for indoor localization application has emerged as a significant area of interest over the last decade, primarily motivated by its low cost and convenient deployment. The weighted centroid localization algorithm is a suitable positioning technique in a wireless sensor network due to its easy implementation. However, the performance of this method is easily affected by outliers and interference in the measurement of radio signal strength. In order to overcome this limitation, a more robust ARMA filter using generalized t-distribution noise model based on influence function approach is proposed. A hardware prototype was implemented to demonstrate that the ARMA filter could improve system performance, especially when dealing with the case of measurement outliers.
Shuo Liu 0002, Le Yin, Weng Khuen Ho, Keck Voon Ling
ICARCV1