Zongyan Cao

dblp:69/10075 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0003-8278-1612ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 52% Distributed systems · 17% High-performance computing · 17%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning
distributed training
0.812024
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis · EuroSys 2024
High-performance computing › cluster computing
heterogeneous clusters
0.812024
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis · EuroSys 2024
Parallel and multicore computing › parallel programming models
SPMD parallelism
0.812024
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis · EuroSys 2024
Machine learning › Efficient and distributed learning
distributed training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing
data parallelism
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing
pipeline parallelism
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training
0.212024
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis · EuroSys 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.212024
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis · EuroSys 2024

Methods — techniques the papers use, named apart from their topics

pipeline parallelism · 1.0parallelization strategy planning · 1.0data parallelism · 1.0tensor sharding · 0.8program synthesis · 0.8linear programming · 0.8a* search · 0.8
YearPublicationVenuePosition
2024 HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
abstract
Single-Program-Multiple-Data (SPMD) parallelism has recently been adopted to train large deep neural networks (DNNs). Few studies have explored its applicability on heterogeneous clusters, to fully exploit available resources for large model learning. This paper presents HAP, an automated system designed to expedite SPMD DNN training on heterogeneous clusters. HAP jointly optimizes the tensor sharding strategy, sharding ratios across heterogeneous devices and the communication methods for tensor exchanges for optimized distributed training with SPMD parallelism. We novelly formulate model partitioning as a program synthesis problem, in which we generate a distributed program from scratch on a distributed instruction set that semantically resembles the program designed for a single device, and systematically explore the solution space with an A-based search algorithm. We derive the optimal tensor sharding ratios by formulating it as a linear programming problem. Additionally, HAP explores tensor communication optimization in a heterogeneous cluster and integrates it as part of the program synthesis process, for automatically choosing optimal collective communication primitives and applying sufficient factor broadcasting technique. Extensive experiments on representative workloads demonstrate that HAP achieves up to 2.41x speed-up on heterogeneous clusters.
Shiwei Zhang 0002, Lansong Diao, Chuan Wu 0001, Zongyan Cao, Siyu Wang 0006, Wei Lin 0016
EuroSys4
2021 DAPPLE: a pipelined data parallel approach for training large models
abstract
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an effective approach for improving device utilization. However, there are still several tricky issues to address: improving computing efficiency while ensuring convergence, and reducing memory usage without incurring additional computing costs. We propose DAPPLE, a synchronous training framework which combines data parallelism and pipeline parallelism for large DNN models. It features a novel parallelization strategy planner to solve the partition and placement problems, and explores the optimal hybrid strategies of data and pipeline parallelism. We also propose a new runtime scheduling algorithm to reduce device memory usage, which is orthogonal to re-computation approach and does not come at the expense of training throughput. Experiments show that DAPPLE planner consistently outperforms strategies generated by PipeDream's planner by up to 3.23× speedup under synchronous training scenarios, and DAPPLE runtime outperforms GPipe by 1.6× speedup of training throughput and saves 12% of memory consumption at the same time.
Shiqing Fan, Zongyan Cao, Siyu Wang 0006, Zhen Zheng, Chuan Wu 0001, Guoping Long, Jun Yang 0052, Lixue Xia, Lansong Diao, Wei Lin 0016
PPoPP4
2015 Parallel simulation of high-dimensional American option pricing based on CPU versus MIC
abstract
Summary American option pricing is a high‐dimensional problem, and its computational challenges have attracted significant attention. We examine this problem using a stochastic mesh method enhanced with bias reduction within the classic Black–Scholes framework. We present Many Integrated Core (MIC) parallelization and acceleration techniques, which result in significant numerical acceleration for large‐scale simulations. In particular, we observe speed‐ups of 21‐fold and 28‐fold for CPU and MIC, respectively, over conventional means. Convergence performance is also examined. Copyright © 2014 John Wiley & Sons, Ltd.
Yonghong Hu, Zongyan Cao, Jue Wang 0013
Concurr. Comput. Pract. Exp.3
2009 USGPA: A User-Centric and Secure Grid Portal Architecture for High-Performance Computing
abstract
A grid portal is one of the most important ways to access grid systems. The disadvantages existing in current grid portals are that they are designed for specific applications, difficult to add new applications. Besides, the security issues are not considered fully in these portals. In this paper, we proposed a user-centric and secure grid portal architecture-USGPA based on portlet. In USGPA, a security portlet model is proposed in which security issues for sensitive data and the portal server are solved. Based on the model, the application portlet model is proposed with which new applications can be encapsulated into portal quickly. Then all portlets for high-performance computing-HPC are designed based on these two models. With these portlets, users can not only manage jobs via portal but also can custom the portal by selecting and managing the applications they required. Last but important, a prototype is implemented to evaluate the scalability, security and efficiency of the architecture.
Rongqiang Cao, Xuebin Chi, Zongyan Cao, Zhihui Dai, Haili Xiao
ISPA3