Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jo Sanghoon

dblp:343/5602 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 67% Parallel and multicore computing · 33%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › collective communication
all-reduce
0.712023
Logical/Physical Topology-Aware Collective Communication in Deep Learning Training · HPCA 2023
High-performance computing
collective communication
0.712023
Logical/Physical Topology-Aware Collective Communication in Deep Learning Training · HPCA 2023
Parallel and multicore computing
data parallelism
0.712023
Logical/Physical Topology-Aware Collective Communication in Deep Learning Training · HPCA 2023
Machine learning › Efficient and distributed learning
distributed training
0.212023
Logical/Physical Topology-Aware Collective Communication in Deep Learning Training · HPCA 2023
Machine learning › Efficient and distributed learning › distributed training › parallelization › parallel training
multi-GPU training
0.212023
Logical/Physical Topology-Aware Collective Communication in Deep Learning Training · HPCA 2023

Methods — techniques the papers use, named apart from their topics

topology-aware communication · 1.3gradient queuing · 1.3
YearPublicationVenuePosition
2023 Logical/Physical Topology-Aware Collective Communication in Deep Learning Training
abstract
Training is an important aspect of deep learning to enable network models to be deployed. To scale training, multiple GPUs are commonly used with data parallelism to exploit the additional GPU compute and memory capacity. However, one challenge in scalability is the collective communication between GPUs. In this work, we propose to accelerate the AllReduce collective. AllReduce communication is often based on a logical topology (e.g., ring or tree algorithms) that is mapped to a physical topology or the physical connectivity between the nodes. In this work, we propose a logical/physical topology-aware collective communication that we refer to as C-Cube architecture – Chaining Collective Communication with Computation. C-Cube exploits the opportunity to overlap or chain different phases of collective communication as well as forward computation in a tree algorithm AllReduce. We exploit the communication pattern in a logical tree topology to overlap the different phases of communication. Since ordering is maintained in the tree collective algorithm, we propose gradient queuing to enable chaining of communication with forward computation to accelerate overall performance while having no impact on training accuracy. We also exploit the physical topology characteristics to further improve the performance, including proposing detour connections for collective communication while leveraging the additional connectivity to enable a double-tree C-Cube implementation. We implement a C-Cube proof-of-concept on a real system (8-GPU NVIDIA DGX-1) and show C-Cube results in performance improvement in communication performance compared to non-overlapped tree algorithms as well as overall performance.
Jo Sanghoon, Hyojun Son, John Kim 0001
HPCA1