Zhen Zhang 0063

dblp:19/5112-63 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0002-0164-0849ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Slapo: A Schedule Language for Progressive Optimization of Large Deep Learning Model Training
abstract
Recent years have seen an increase in the development of large deep learning (DL) models, which makes training efficiency crucial. Common practice is struggling with the trade-off between usability and performance. On one hand, DL frameworks such as PyTorch use dynamic graphs to facilitate model developers at a price of sub-optimal model training performance. On the other hand, practitioners propose various approaches to improving the training efficiency by sacrificing some of the flexibility, ranging from making the graph static for more thorough optimization (e.g., XLA) to customizing optimization towards large-scale distributed training (e.g., DeepSpeed and Megatron-LM).
Hongzheng Chen, Cody Hao Yu, Shuai Zheng 0004, Zhen Zhang 0063, Zhiru Zhang, Yida Wang 0003
ASPLOS (2)4
2024 Distributed Training of Large Language Models on AWS Trainium
abstract
Large language models (LLMs) are ubiquitously powerful but prohibitively expensive to train, often requiring thousands of compute devices, typically GPUs. To reduce the cost of training LLMs for customers, Amazon Web Services (AWS) launched the Amazon EC2 trn1 instances, powered by AWS Trainium, an Amazon's homegrown deep learning accelerator, as an alternative to distributed LLM training. The trn1 instances provide a high-performance LLM training solution at a lower cost compared to their GPU-based counterpart, the p4d instances, which are powered by Nvidia A100 GPUs. This paper describes the design and development of the Neuron Distributed Training Library, a component of the AWS Neuron SDK, which enables distributed training of large language models on AWS Trainium. Neuron Distributed Training Library supports a variety of existing distributed training techniques with unified interfaces, and provides features to address trn1-specific challenges as well. Our evaluation shows that trn1 instances, specifically the trn1.32xlarge, achieve better or comparable performance (up to 24.6% improvement) while offering significant lower costs (up to 46.3% cost saving) in selected workloads when compared to p4d.24xlarge instances. As a result, AWS Trainium has been adopted for training numerous external and internal models, showcasing its high-performance and cost-effectiveness. Several supported open-source LLMs are accessible via HuggingFace Optimum Neuron.
Xinwei Fu, Zhen Zhang 0063, Haozheng Fan, Guangtai Huang, Mohammad El-Shabani, Randy Huang, Rahul Solanki, Ron Diamant, Yida Wang 0003
SoCC2
2024 DISTMM: Accelerating Distributed Multimodal Model Training
Zhen Zhang 0063, Shuai Zheng 0004, Yida Wang 0003
NSDI2
2024 SDCC: software-defined collective communication for distributed training
Xin Jin 0008, Zhen Zhang 0063, Yunshan Jia, Yun Ma 0002, Xuanzhe Liu
Sci. China Inf. Sci.2
2024 DistMind: Efficient Resource Disaggregation for Deep Learning Workloads
abstract
Deep learning (DL) systems suffer from low resource utilization due to 1) monolithic server model that tightly couples compute and memory; and 2) limited sharing between different inference applications, and across inference and training, because of strict service level objectives (SLOs). To address this problem, we present, a disaggregated DL system that enables efficient multiplexing of DL applications with near-optimal resource utilization. decouples compute from host memory, and exposes the abstractions of a GPU pool and a memory pool, each of which can be independently provisioned. The key challenge is to dynamically allocate GPU resources to different applications based on their real-time demands while meeting strict SLOs. We tackle this challenge by exploiting the power of high-speed 100 Gbps networks, and design three-stage pipelining, cache-aware load balancing, and DNN-aware sharding mechanisms based on the characteristics of DL workloads, to achieve millisecond-scale application loading overhead and improve system efficiency. We have implemented a prototype of and integrated it with PyTorch. Experimental results on AWS EC2 show that achieves near 100% resource utilization, and compared with NVIDIA MPS and Ray, improves the throughput by up to 279% and reduces the inference latency by up to 94%.
Xin Jin 0008, Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Yinmin Zhong, Xuanzhe Liu
IEEE/ACM Trans. Netw.3
2023 Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
abstract
Oobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set of heterogeneous pipeline templates and instantiates at least f + 1 logically equivalent pipeline replicas to tolerate any f simultaneous failures. During execution, it relies on already-replicated model states across the replicas to provide fast recovery. Oobleck provably guarantees that some combination of the initially created pipeline templates can be used to cover all available resources after f or fewer simultaneous failures, thereby avoiding resource idling at all times. Evaluation on large DNN models with billions of parameters shows that Oobleck provides consistently high throughput, and it outperforms state-of-the-art fault tolerance solutions like Bamboo and Varuna by up to 13.9×.
Insu Jang, Zhenning Yang, Zhen Zhang 0063, Xin Jin 0008, Mosharaf Chowdhury
SOSP3
2023 GEMINI: Fast Failure Recovery in Distributed Training with In-Memory Checkpoints
abstract
Large deep learning models have recently garnered substantial attention from both academia and industry. Nonetheless, frequent failures are observed during large model training due to large-scale resources involved and extended training time. Existing solutions have significant failure recovery costs due to the severe restriction imposed by the bandwidth of remote storage in which they store checkpoints.
Zhen Jia 0001, Shuai Zheng 0004, Zhen Zhang 0063, Xinwei Fu, T. S. Eugene Ng, Yida Wang 0003
SOSP4
2022 MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
abstract
Existing general purpose frameworks for gigantic model training, i.e., dense models with billions of parameters, cannot scale efficiently on cloud environment with various networking conditions due to large communication overheads. In this paper, we propose MiCS, which Minimizes the Communication Scale to bring down communication overhead. Specifically, by decreasing the number of participants in a communication collective, MiCS can utilize heterogeneous network bandwidth, reduce network traffic over slower links, reduce the latency of communications for maintaining high network bandwidth utilization, and amortize expensive global gradient synchronization overhead. Our evaluation on AWS shows that the system throughput of MiCS is up to 2.89× that of the state-of-the-art large model training systems. MiCS achieves near-linear scaling efficiency, which is up to 1.27× that of DeepSpeed. MiCS allows us to train a proprietary model with 100 billion parameters on 512 GPUs with 99.4% weak-scaling efficiency, and it is able to saturate over 54.5% theoretical computation power of each GPU on a public cloud with less GPU memory and more restricted networks than DGX-A100 clusters.
Zhen Zhang 0063, Shuai Zheng 0004, Yida Wang 0003, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li 0003, Xin Jin 0008
Proc. VLDB Endow.1
2020 PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications
Zhihao Bai, Zhen Zhang 0063, Yibo Zhu 0001, Xin Jin 0008
OSDI2