Wumo Yan

dblp:254/1176 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0002-4506-795XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 61% Performance modeling and evaluation · 30% Cloud and datacenter computing · 9%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
distributed machine learning
0.612022
Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022
Distributed systems › distributed machine learning
distributed stochastic gradient descent
0.612022
Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022
Performance modeling and evaluation
performance prediction
0.612022
Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

profiling · 0.6analytical performance modeling · 0.6
YearPublicationVenuePosition
2022 Predicting Throughput of Distributed Stochastic Gradient Descent
abstract
Training jobs of deep neural networks (DNNs) can be accelerated through distributed variants of stochastic gradient descent (SGD), where multiple nodes process training examples and exchange updates. The total throughput of the nodes depends not only on their computing power, but also on their networking speeds and coordination mechanism (synchronous or asynchronous, centralized or decentralized), since communication bottlenecks and stragglers can result in sublinear scaling when additional nodes are provisioned. In this paper, we propose two classes of performance models to predict throughput of distributed SGD:fine-grained models, representing many elementary computation/communication operations and their dependencies; andcoarse-grained models, where SGD steps at each node are represented as a sequence of high-level phases without parallelism between computation and communication. Using a PyTorch implementation, real-world DNN models and different cloud environments, our experimental evaluation illustrates that, while fine-grained models are more accurate and can be easily adapted to new variants of distributed SGD, coarse-grained models can provide similarly accurate predictions when augmented with ad hoc heuristics, and their parameters can be estimated with profiling information that is easier to collect.
Zhuojin Li, Marco Paolieri, Leana Golubchik, Sung-Han Lin, Wumo Yan
IEEE Trans. Parallel Distributed Syst.5
2020 Throughput Prediction of Asynchronous SGD in TensorFlow
abstract
Modern machine learning frameworks can train neural networks using multiple nodes in parallel, each computing parameter updates with stochastic gradient descent (SGD) and sharing them asynchronously through a central parameter server. Due to communication overhead and bottlenecks, the total throughput of SGD updates in a cluster scales sublinearly, saturating as the number of nodes increases. In this paper, we present a solution to predicting training throughput from profiling traces collected from a single-node configuration. Our approach is able to model the interaction of multiple nodes and the scheduling of concurrent transmissions between the parameter server and each node. By accounting for the dependencies between received parts and pending computations, we predict overlaps between computation and communication and generate synthetic execution traces for configurations with multiple nodes. We validate our approach on TensorFlow training jobs for popular image classification neural networks, on AWS and on our in-house cluster, using nodes equipped with GPUs or only with CPUs. We also investigate the effects of data transmission policies used in TensorFlow and the accuracy of our approach when combined with optimizations of the transmission schedule.
Zhuojin Li, Wumo Yan, Marco Paolieri, Leana Golubchik
ICPE2