VLDB 2026 Research / reviewers in the wild / expert
Wumo Yan
dblp:254/1176
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0002-4506-795XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 61% Performance modeling and evaluation · 30% Cloud and datacenter computing · 9% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
distributed machine learning |
0.6 | 1 | 2022 | Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022 |
Distributed systems › distributed machine learning
distributed stochastic gradient descent |
0.6 | 1 | 2022 | Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022 |
Performance modeling and evaluation
performance prediction |
0.6 | 1 | 2022 | Predicting Throughput of Distributed Stochastic Gradient Descent · IEEE Trans. Parallel Distributed Syst. 2022 |
Methods — techniques the papers use, named apart from their topics
profiling · 0.6analytical performance modeling · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Predicting Throughput of Distributed Stochastic Gradient DescentabstractTraining jobs of deep neural networks (DNNs) can be accelerated through distributed variants of stochastic gradient descent (SGD), where multiple nodes process training examples and exchange updates. The total throughput of the nodes depends not only on their computing power, but also on their networking speeds and coordination mechanism (synchronous or asynchronous, centralized or decentralized), since communication bottlenecks and stragglers can result in sublinear scaling when additional nodes are provisioned. In this paper, we propose two classes of performance models to predict throughput of distributed SGD:fine-grained models, representing many elementary computation/communication operations and their dependencies; andcoarse-grained models, where SGD steps at each node are represented as a sequence of high-level phases without parallelism between computation and communication. Using a PyTorch implementation, real-world DNN models and different cloud environments, our experimental evaluation illustrates that, while fine-grained models are more accurate and can be easily adapted to new variants of distributed SGD, coarse-grained models can provide similarly accurate predictions when augmented with ad hoc heuristics, and their parameters can be estimated with profiling information that is easier to collect. Zhuojin Li, Marco Paolieri, Leana Golubchik, Sung-Han Lin, Wumo Yan |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Throughput Prediction of Asynchronous SGD in TensorFlowabstractModern machine learning frameworks can train neural networks using multiple nodes in parallel, each computing parameter updates with stochastic gradient descent (SGD) and sharing them asynchronously through a central parameter server. Due to communication overhead and bottlenecks, the total throughput of SGD updates in a cluster scales sublinearly, saturating as the number of nodes increases. In this paper, we present a solution to predicting training throughput from profiling traces collected from a single-node configuration. Our approach is able to model the interaction of multiple nodes and the scheduling of concurrent transmissions between the parameter server and each node. By accounting for the dependencies between received parts and pending computations, we predict overlaps between computation and communication and generate synthetic execution traces for configurations with multiple nodes. We validate our approach on TensorFlow training jobs for popular image classification neural networks, on AWS and on our in-house cluster, using nodes equipped with GPUs or only with CPUs. We also investigate the effects of data transmission policies used in TensorFlow and the accuracy of our approach when combined with optimizations of the transmission schedule. Zhuojin Li, Wumo Yan, Marco Paolieri, Leana Golubchik |
ICPE | 2 |