Oh-Kyoung Kwon

dblp:16/4583 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-9734-9257ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 When HPC Scheduling Meets Active Learning: Maximizing The Performance with Minimal Data
Jiheon Choi, Minsol Choo, Oh-Kyoung Kwon, Sangyoon Oh 0001
HPC Asia5
2025 HYPERF: End-to-End Autotuning Framework for High-Performance Computing
Juseong Park, Yongwon Shin, Oh-Kyoung Kwon, Hyojin Sung
HPDC6
2024 GPUTucker: Large-Scale GPU-Based Tucker Decomposition Using Tensor Partitioning
Donghyoung Han, Oh-Kyoung Kwon, Kang-Wook Chon, Min-Soo Kim 0002
Expert Syst. Appl.3
2023 Addressing Client Heterogeneity in Synchronous Federated Learning: The CHAFL Approach
abstract
Federated learning (FL) is proposed to address the security vulnerabilities of conventional distributed deep learning. Since the capabilities of participating FL clients are highly variable in terms of both statistical and system aspects, FL training will face diminished convergence accuracy and speed. Hence, we propose CHAFL (client heterogeneity aware federated learning) to address client heterogeneity (i.e., statistical and system heterogeneity) in synchronous FL. CHAFL selects clients based on global loss with contribution, defined as local loss, enabling higher round-to-accuracy for the global model than previous studies. Additionally, to handle system heterogeneity, it proposes a lightweight algorithm that eliminates the profiling process previously employed to calculate adaptive local epochs in existing studies, thereby improving time-to-accuracy. In order to verify the effectiveness of CHAFL, we exploit three benchmark datasets on non-IID and system heterogeneous setting for empirical evaluation. Compared to the baseline, CHAFL achieves an accuracy improvement of 3.1 to 5.7%, along with 1.04 to 1.73x higher round-to-accuracy and 1.03 to 1.98x higher time-to-accuracy.
Miri Yu, Oh-Kyoung Kwon, Sangyoon Oh 0001
ICPADS2
2023 Crossover-SGD: A gossip-based communication in distributed deep learning for alleviating large mini-batch problem and enhancing scalability
abstract
Summary Distributed deep learning is an effective way to reduce the training time for large datasets as well as complex models. However, the limited scalability caused by network‐overheads makes it difficult to synchronize the parameters of all workers and gossip‐based methods that demonstrate stable scalability regardless of the number of workers have been proposed. However, to use gossip‐based methods in general cases, the validation accuracy for a large mini‐batch needs to be verified. For this, we first empirically study the characteristics of gossip methods in a large mini‐batch problem and observe that gossip methods preserve higher validation accuracy than AllReduce‐SGD (stochastic gradient descent) when the number of batch sizes is increased, and the number of workers is fixed. However, the delayed parameter propagation of the gossip‐based models decreases validation accuracy in large node scales. To cope with this problem, we propose Crossover‐SGD that alleviates the delay propagation of weight parameters via segment‐wise communication and random network topology with fair peer selection. We also adapt hierarchical communication to limit the number of workers in gossip‐based communication methods. To validate the effectiveness of our method, we conduct empirical experiments and observe that our Crossover‐SGD shows higher node scalability than stochastic gradient push.
Sangho Yeo, Minho Bae, Minjoong Jeong, Oh-Kyoung Kwon, Sangyoon Oh 0001
Concurr. Comput. Pract. Exp.4
2022 Efficient Task-Mapping of Parallel Applications Using a Space-Filling Curve
abstract
Improving the communication performance of parallel programs is an important but difficult problem in a large-scale distributed memory-based cluster. Efforts to improve parallel scalability often face severe huddles in managing communication overheads. This paper proposes a framework of a space-filing curve(SFC)-based task-remapping for communication intensive parallel applications. An SFC-based mapping, when applied for task-mapping of parallel applications preserves locality in terms of communications and produce a less fragmented task-mapping, reducing communication overheads. The framework also provides tools for performance analysis to see if the proposed task-mapping is appropriate for a given application running on a target system. It further develops a binary classifier as a predictor to decide whether or not to apply the proposed mapping before run-time. We evaluate the framework with three communication intensive applications in Cartesian coordinates: P3DFFT solver and Channel code using 2D domain decomposition model, and Poisson solver using 3D domain decomposition. The evaluation is conducted on a large-scale cluster system of fat-tree topology with up to 1,024 compute nodes. The proposed task-mapping achieves the overall performance improvement ranging from ~30% to ~66% over the baseline approach depending on the workloads. Also, when used in combination with the binary classifier-based predictor, it achieves the expected performance gains from 4% to 8%.
Oh-Kyoung Kwon, Ji Hoon Kang 0002, Seungchul Lee, Wonjung Kim 0002, Junehwa Song
PACT1
2004 GRASP: A Grid Resource Allocation System based on OGSA
Oh-Kyoung Kwon, Jaegyoon Hahm, Jong-Suk Ruth Lee
HPDC1