Shiwei Zhang 0002

dblp:15/1308-2 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-0838-1883ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 DyOrc: Efficient Serving of Dynamic Machine Learning Workflows
abstract
The landscape of machine learning applications has shifted from monolithic end-to-end models to compositions of pretrained large foundation models. For instance, multi-modal chatbots are often built by composition of a large language model and modality-specific encoder models. Such applications often feature dynamic workflows, with models conditionally evoked according to different inputs and intermediate processing results. Conditional model execution prevents conventional request batching and hinders efficient hardware utilization, due to dynamic, diverging execution paths across requests. Separately deploying models as dedicated services and invoking them on the go during dynamic workflow executions can potentially allow service-wise request batching, boosting resource efficiency. However, generic workflow orchestrators are proven inefficient for machine learning applications, due to schedulers that do not exploit batching, communication methods that are suboptimal for GPU tensors, and the considerable cold-start delays associated with model deployment.
Shiwei Zhang 0002, Lansong Diao, Zisheng Meng, Siyu Wang 0006, Wei Lin 0016, Chuan Wu 0001
SoCC1
2024 HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
abstract
Single-Program-Multiple-Data (SPMD) parallelism has recently been adopted to train large deep neural networks (DNNs). Few studies have explored its applicability on heterogeneous clusters, to fully exploit available resources for large model learning. This paper presents HAP, an automated system designed to expedite SPMD DNN training on heterogeneous clusters. HAP jointly optimizes the tensor sharding strategy, sharding ratios across heterogeneous devices and the communication methods for tensor exchanges for optimized distributed training with SPMD parallelism. We novelly formulate model partitioning as a program synthesis problem, in which we generate a distributed program from scratch on a distributed instruction set that semantically resembles the program designed for a single device, and systematically explore the solution space with an A-based search algorithm. We derive the optimal tensor sharding ratios by formulating it as a linear programming problem. Additionally, HAP explores tensor communication optimization in a heterogeneous cluster and integrates it as part of the program synthesis process, for automatically choosing optimal collective communication primitives and applying sufficient factor broadcasting technique. Extensive experiments on representative workloads demonstrate that HAP achieves up to 2.41x speed-up on heterogeneous clusters.
Shiwei Zhang 0002, Lansong Diao, Chuan Wu 0001, Zongyan Cao, Siyu Wang 0006, Wei Lin 0016
EuroSys1
2023 Expediting Distributed DNN Training With Device Topology-Aware Graph Deployment
abstract
This paper presents TAG, an automatic system to derive optimized DNN training graph and its deployment onto any device topology, for expedited training in device- and topology- heterogeneous ML clusters. We novelly combine both the DNN computation graph and the device topology graph as input to a graph neural network (GNN), and join the GNN with a search-based method to quickly identify optimized distributed training strategies. To reduce communication in a heterogeneous cluster, we further explore a lossless gradient compression technique and solve a combinatorial optimization problem to automatically apply the technique for training time minimization. We evaluate TAG with various representative DNN models and device topologies, showing that it can achieve up to 4.56x training speed-up as compared to existing schemes. TAG can produce efficient deployment strategies for both unseen DNN models and unseen device topologies, without heavy fine-tuning.
Shiwei Zhang 0002, Xiaodong Yi 0001, Lansong Diao, Chuan Wu 0001, Siyu Wang 0006, Wei Lin 0016
IEEE Trans. Parallel Distributed Syst.1
2022 Accelerating large-scale distributed neural network training with SPMD parallelism
abstract
Deep neural networks (DNNs) with trillions of parameters have emerged, e.g., Mixture-of-Experts (MoE) models. Training models of this scale requires sophisticated parallelization strategies like the newly proposed SPMD parallelism, that shards each tensor along different dimensions. A common problem using SPMD is that computation stalls during communication due to data dependencies, resulting in low GPU utilization and long training time. We present a general technique to accelerate SPMD-based DNN training by maximizing computation-communication overlap and automatic SPMD strategy search. The key idea is to duplicate the DNN model into two copies that have no dependency, and interleave their execution such that computation of one copy overlaps with communication of the other. We propose a dynamic programming algorithm to automatically identify optimized sharding strategies that minimize model training time by maximally enabling computation-communication overlap. Experiments show that our designs achieve up to 61% training speed-up as compared to existing frameworks.
Shiwei Zhang 0002, Lansong Diao, Chuan Wu 0001, Siyu Wang 0006, Wei Lin 0016
SoCC1
2022 Optimizing DNN Compilation for Distributed Training With Joint OP and Tensor Fusion
abstract
This article proposesDisCo, an automatic deep learning compilation module for data-parallel distributed training. Unlike most deep learning compilers that focus on training or inference on a single device,DisCooptimizes a DNN model for distributed training over multiple GPU machines. Existing single-device compilation strategies do not work well in distributed training, due mainly to communication inefficiency that they incur.DisCogenerates optimized, joint computation operator and communication tensor fusion strategies to enable highly efficient distributed training. A GNN-based simulator is built to effectively estimate per-iteration training time achieved by operator/tensor fusion candidates. A backtracking search algorithm is driven by the simulator, navigating efficiently in the large strategy space to identify good operator/tensor fusion strategies that minimize distributed training time. We compareDisCowith existing DL fusion schemes and show that it achieves good training speed-up close to the ideal, full computation-communication overlap case.
Xiaodong Yi 0001, Shiwei Zhang 0002, Lansong Diao, Chuan Wu 0001, Zhen Zheng, Shiqing Fan, Siyu Wang 0006, Jun Yang 0052, Wei Lin 0016
IEEE Trans. Parallel Distributed Syst.2
2020 Optimizing distributed training deployment in heterogeneous GPU clusters
abstract
This paper proposes HeteroG, an automatic module to accelerate deep neural network training in heterogeneous GPU clusters. To train a deep learning model with large amounts of data, distributed training using data or model parallelism has been widely adopted, mostly over homogeneous devices (GPUs, network bandwidth). Heterogeneous training environments may often exist in shared clusters with GPUs of different models purchased in different batches and network connections of different bandwidth availability (e.g., due to contention). Classic data parallelism does not work well in a heterogeneous cluster, while model-parallel training is hard to plan. HeteroG enables highly-efficient distributed training over heterogeneous devices, by automatically converting a single-GPU training model to a distributed one according to the deep learning graph and available resources. HeteroG embraces operation-level hybrid parallelism, communication architecture selection and execution scheduling, based on a carefully designed strategy framework exploiting both GNN-based learning and combinatorial optimization. We compare HeteroG with existing parallelism schemes and show that it achieves up-to 222% training speed-up. HeteroG also enables efficient training of large models over a set of heterogeneous devices where simple parallelism is infeasible.
Xiaodong Yi 0001, Shiwei Zhang 0002, Ziyue Luo, Guoping Long, Lansong Diao, Chuan Wu 0001, Zhen Zheng, Jun Yang 0052, Wei Lin 0016
CoNEXT2