Daegun Yoon

dblp:264/1445 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-7520-1144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2024 Preserving Near-Optimal Gradient Sparsification Cost for Scalable Distributed Deep Learning
abstract
Communication overhead is a major obstacle to scaling distributed training systems. Gradient sparsification is a potential optimization approach to reduce the communication volume without significant loss of model fidelity. However, existing gradient sparsification methods have low scalability owing to inefficient design of their algorithms, which raises the communication overhead significantly. In particular, gradient build-up and inadequate sparsity control methods degrade the sparsification performance considerably. Moreover, communication traffic increases drastically owing to workload imbalance of gradient selection between workers.To address these challenges, we propose a novel gradient sparsification scheme called ExDyna. In ExDyna, the gradient tensor of the model comprises fined-grained blocks, and contiguous blocks are grouped into non-overlapping partitions. Each worker selects gradients in its exclusively allocated partition so that gradient build-up never occurs. To balance the workload of gradient selection between workers, ExDyna adjusts the topology of partitions by comparing the workloads of adjacent partitions. In addition, ExDyna supports online threshold scaling, which estimates the accurate threshold of gradient selection on-the-fly. Accordingly, ExDyna can satisfy the user-required sparsity level during a training period regardless of models and datasets. Therefore, ExDyna can enhance the scalability of distributed training systems by preserving near-optimal gradient sparsification cost. In experiments, ExDyna outperformed state-of-the-art sparsifiers in terms of training speed and sparsification performance while achieving high accuracy.
Daegun Yoon, Sangyoon Oh 0001
CCGrid1
2023 MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
abstract
Gradient sparsification is a communication optimisation technique for scaling and accelerating distributed deep neural network (DNN) training. It reduces the increasing communication traffic for gradient aggregation. However, existing sparsifiers have poor scalability because of the high computational cost of gradient selection and/or increase in communication traffic. In particular, an increase in communication traffic is caused by gradient build-up and inappropriate threshold for gradient selection. To address these challenges, we propose a novel gradient sparsification method called MiCRO. In MiCRO, the gradient vector is partitioned, and each partition is assigned to the corresponding worker. Each worker then selects gradients from its partition, and the aggregated gradients are free from gradient build-up. Moreover, MiCRO estimates the accurate threshold to maintain the communication traffic as per user requirement by minimising the compression ratio error. MiCRO enables near-zero cost gradient sparsification by solving existing problems that hinder the scalability and acceleration of distributed DNN training. In our extensive experiments, MiCRO outperformed state-of-the-art sparsifiers with an outstanding convergence rate.
Daegun Yoon, Sangyoon Oh 0001
HiPC1
2023 DEFT: Exploiting Gradient Norm Difference between Model Layers for Scalable Gradient Sparsification
abstract
Gradient sparsification is a widely adopted solution for reducing the excessive communication traffic in distributed deep learning. However, most existing gradient sparsifiers have relatively poor scalability because of considerable computational cost of gradient selection and/or increased communication traffic owing to gradient build-up. To address these challenges, we propose a novel gradient sparsification scheme, DEFT, that partitions the gradient selection task into sub tasks and distributes them to workers. DEFT differs from existing sparsifiers, wherein every worker selects gradients among all gradients. Consequently, the computational cost can be reduced as the number of workers increases. Moreover, gradient build-up can be eliminated because DEFT allows workers to select gradients in partitions that are non-intersecting (between workers). Therefore, even if the number of workers increases, the communication traffic can be maintained as per user requirement.
Daegun Yoon, Sangyoon Oh 0001
ICPP1
2023 WAVE: designing a heuristics-based three-way breadth-first search on GPUs
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001
J. Supercomput.1
2023 SAGE: toward on-the-fly gradient compression ratio scaling
Daegun Yoon, Minjoong Jeong, Sangyoon Oh 0001
J. Supercomput.1
2022 AMBLE: Adjusting mini-batch and local epoch for federated learning with heterogeneous devices
Juwon Park, Daegun Yoon, Sangho Yeo, Sangyoon Oh 0001
J. Parallel Distributed Comput.2
2021 Exploring a system architecture of content-based publish/subscribe system for efficient on-the-fly data dissemination
abstract
Summary In a cloud‐scale publish/subscribe messaging system, it is difficult to partition subscription data among several servers. Without a sophisticated scheme and a system architecture, the messaging system would either waste resources or fail to deliver messages on time. In this study, we propose DRDA, a dynamic replication degree adjustment technology, for efficient message delivery. The technology calculates and maintains the number of subscription replications at a reasonable level by monitoring the statuses of servers, based on the number of subscription replications and the frequency of event dissemination. To verify the effectiveness of our proposed scheme and system architecture, we build a prototype of a content‐based publish/subscribe system that dynamically adjusts the number of replications among brokers. Furthermore, we compare the load balance, resource overhead, and performance of a publish/subscribe system with DRDA with a publish/subscribe system without DRDA. The experimental results show that DRDA outperforms other approaches under various parameter configurations. We have added the prototype code to a GitHub repository to make it publicly available.
Daegun Yoon, Gyudong Park, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.1
2021 Balanced content space partitioning for pub/sub: a study on impact of varying partitioning granularity
Daegun Yoon, Zhetao Li, Sangyoon Oh 0001
J. Supercomput.1