Gyudong Kim

dblp:97/7003 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Deep learning architectures and training · 67% Efficient and distributed learning · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.912025
First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training · NeurIPS 2025
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.912025
First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

tensor parallelism · 0.9all-reduce elimination · 0.9
YearPublicationVenuePosition
2026 Exploring Heterogeneity-Aware Optimizations for Resource Efficient Edge Recommendation
abstract
Recommendation systems are widely deployed on edge devices to enable personalized user experience. While recommendation inference has traditionally been performed on centralized servers, recent advances in mobile SoCs have motivated a shift toward on-device execution. However, achieving efficient recommendation inference on edge devices remains challenging due to edge-specific execution characteristics and heterogeneity. In this paper, we characterize resource inefficiencies under realistic edge constraints and propose optimization strategies.
Yerin Lee, Gyudong Kim, Eunjin Lee, Jeff Zhang 0001, Young-Ho Gong, Carole-Jean Wu
DATE2
2025 First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training
abstract
As training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), where each block’s MHA–MLP connection requires an all-reduce communication. Through our investigation, we show that the MHA-MLP connections can be bypassed for efficiency, while the attention output of the first layer can serve as an alternative signal for the bypassed connection. Motivated by the observations, we propose FAL (First Attentions Last), an efficient transformer architecture that redirects the first MHA output to the MLP inputs of the following layers, eliminating the per-block MHA-MLP connections. This removes the all-reduce communication and enables parallel execution of MHA and MLP on a single GPU. We also introduce FAL+, which adds the normalized first attention output to the MHA outputs of the following layers to augment the MLP input for the model quality. Our evaluation shows that FAL reduces multi-GPU training time by up to 44%, improves single-GPU throughput by up to 1.18×, and achieves better perplexity compared to the baseline GPT. FAL+ achieves even lower perplexity without increasing the training time than the baseline. Codes are available at: https://casl-ku.github.io/FAL/
Gyudong Kim, Hyukju Na, Jin Kyu Kim, Hyunsung Jang, Jaegi Hwang, Namkoo Ha, Seungryong Kim, Younggeun Kim 0001
NeurIPS1