Yunki Han

dblp:309/4348 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-0432-9324ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector Quantization
abstract
Enabling real-time GNN inference services requires low end-to-end latency to meet service level agreements. However, intensive preparation steps and the neighborhood explosion problem pose significant challenges to efficient GNN inference serving. In this paper, we propose FLAG, an FPGA-based GNN inference serving system using vector quantization. To reduce preparation overhead, we introduce offline preprocessing to precompute and compress hidden embeddings for serving. A dedicated FPGA accelerator leverages the precomputed data to enable lightweight aggregation. As a result, FLAG achieves average speedups of $154 \times 176 \times$, and $333 \times$ on three GNN models compared to the baseline system.
Yunki Han, Taehwan Kim 0011, Seohye Ha, Lee-Sup Kim
DAC1
2025 SAFF: Scalable Acceleration of GNN-based Machine Learning Force Fields using Tensor-Aware Hardware for Molecular Simulation
abstract
Machine Learning Force Fields (MLFFs) have emerged as a key technique in molecular dynamics (MD) simulations, enabling accurate prediction of molecular energies and forces with a level of precision comparable to that of Density Functional Theory (DFT), while significantly reducing computational cost. However, applying MLFFs to real-world scenarios, such as semiconductor process simulations involving massive atomic scales and long simulation times, exposes severe performance limitations, resulting in current MLFF models being inadequate for practical use and emphasizing the necessity of significant computational acceleration. Graph Neural Networks (GNNs) are frequently utilized to construct MLFFs, with tensor product operations constituting a substantial portion of the computational workload. These operations involve high-dimensional tensors whose sizes are determined by the number of channels and the order of rotation. This results in frequent memory accesses and low data reuse. This leads to memory bottlenecks that limit the efficient use of GPU computational resources. In this work, we propose a tensor-based function fusion technique to improve data continuity between functions, thereby increasing on-chip data reuse and reducing off-chip memory access. Furthermore, we identified inefficient utilization of hardware resources and addressed this by restructuring the computation flow and removing redundant operations to enhance computational performance. To optimize performance, we have also designed a dedicated tensor product hardware architecture that is optimized for GNN-based MLFFs. The experimental results demonstrate that the proposed system achieves an average speedup of 6.26× and a 424× improvement in energy efficiency compared to GPU-based execution, while maintaining the original model accuracy.
Seohye Ha, Yunki Han, Taehwan Kim 0011, Gunhee Park, Lee-Sup Kim
ICCAD2
2025 EOD: Enabling Low Latency GNN Inference via Near-Memory Concatenate Aggregation
abstract
As online services based on graph databases increasingly integrate with machine learning, serving low-latency Graph Neural Network (GNN) inference for individual requests has become a critical challenge.Real-time GNN inference services operate in an inductive setup, which can handle newly added, previously unseen nodes and their edges.In this setup, the system must prepare the computational graph and input node features for target nodes, followed by GNN inference using the given input data.However, the workflow of a GNN serving system presents two key challenges that hinder low-latency inference.The first challenge arises from the extensive preparation step, which involves heavy memory access in host memory and significant data I/O to devices, constituting the largest portion of end-to-end inference latency.The second challenge is the well-known neighborhood explosion problem in GNN research.As the receptive field for target nodes increases exponentially with the number of layers, this issue exacerbates overall latency.To address these challenges, we propose a Near-Memory Processing (NMP) based low-latency GNN inference serving system named EOD.To ensure low-latency real-time GNN service, we co-design the algorithm and hardware to tackle the aforementioned issues.First, to mitigate the neighborhood explosion problem, we propose a precomputation method for the training node set, reducing memory access and computational complexity from exponential to linear growth.Additionally, we introduce a concatenated ZVC compression method to minimize the overhead of storing precomputed hidden features.Finally, to alleviate heavy host-side memory access and data I/O, we design an NMP architecture that enables efficient aggregation on concatenated ZVC-compressed data.As a result, EOD achieves a geometric mean of 981.1× and 912.0× aggregation speedup over the baseline and the existing architecture for GNN aggregation.Additionally, EOD achieves a geometric mean of 17.9×, and up to 74.3× end-to-end latency speedup over the GPU baseline.
Taehwan Kim 0011, Yunki Han, Seohye Ha, Lee-Sup Kim
ISCA2
2025 AToM: Adaptive Token Merging for Efficient Acceleration of Vision Transformer
abstract
Recently, Vision Transformers (ViTs) have set a new standard in computer vision (CV), showing unparalleled image processing performance. However, their substantial computational requirements hinder practical deployment, especially on resource-limited devices common in CV applications. Token merging has emerged as a solution, condensing tokens with similar features to cut computational and memory demands. Yet, existing applications on ViTs often miss the mark in token compression, with rigid merging strategies and a lack of in-depth analysis of ViT merging characteristics. To overcome these issues, this paper introduces Adaptive Token Merging (AToM), a comprehensive algorithm-architecture co-design for accelerating ViTs. The AToM algorithm employs an image-adaptive, fine-grained merging strategy, significantly boosting computational efficiency. We also optimize the merging and unmerging processes to minimize overhead, employing techniques like First-Come-First-Merge mapping and Linear Distance Calculation. On the hardware side, the AToM architecture is tailor-made to exploit the AToM algorithm's benefits, with specialized engines for efficient merge and unmerge operations. Our pipeline architecture ensures end-to-end ViT processing, minimizing latency and memory overhead from the AToM algorithm. Across various hardware platforms including CPU, EdgeGPU, and GPU, AToM achieves average end-to-end speedups of 10.9$\boldsymbol{\times}$, 7.7$\boldsymbol{\times}$, and 5.4$\boldsymbol{\times}$, alongside energy savings of 24.9$\boldsymbol{\times}$, 1.8$\boldsymbol{\times}$, and 16.7$\boldsymbol{\times}$. Moreover, AToM offers 1.2$\boldsymbol{\times}$1.9$\boldsymbol{\times}$higher effective throughput compared to existing transformer accelerators.
Jaekang Shin, Myeonggu Kang, Yunki Han, Lee-Sup Kim
IEEE Trans. Computers3
2024 Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
abstract
The attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency.
Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Jaekang Shin, Lee-Sup Kim
DAC3
2024 CoCoA: Algorithm-Hardware Co-Design for Large-Scale GNN Training using Compressed Graph
abstract
Scaling Graph Neural Network (GNN) training on large-scale graph data poses a critical challenge for implementing GNN applications on real-world giant graphs. The size of real-world graphs often exceeds the memory capacity of accelerator devices, necessitating the use of multiple devices or host memory for training. While expanding memory space alleviates the out-of-memory problem, this approach introduces another bottleneck through heavy communication via low-bandwidth interconnection. Therefore, achieving efficient, scalable GNN training on large-scale graphs requires addressing both capacity and communication issues.
Yunki Han, Jaekang Shin, Gunhee Park, Lee-Sup Kim
ICCAD1
2023 OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die Processing
abstract
Training deep neural network (DNN) models is a resource-intensive, iterative process. For this reason, nowadays, complex optimizers like Adam are widely adopted as it increases the speed and efficiency of training. These optimizers, however, employ additional variables and raise the memory demand 2× to 3× of model parameters, worsening the memory capacity bottleneck. Moreover, as the size of DNN models is projected to grow even further, it is not practical to assume that the future models will fit in accelerator memory. This has triggered various efforts to offload models to flash-based storage. However, when the model, especially the optimizer, is offloaded to flash, the limited I/O bandwidth severely slows down the overall training process. To this end, we present OptimStore, a solid-state drive (SSD) system with on-die processing (ODP) architectures for gradient descent-based machine learning models. OptimStore accelerates the training process of such large-scale models by processing model optimization in the storage device, specifically inside the flash dies. ODP capability of OptimStore eliminates the heavy data movement over external interconnect and internal flash channels. Overall, OptimStore achieves, on average, a 2.8× speedup and a 3.6× improved energy efficiency in the weight update stage over baseline SSD offloading.
Junkyum Kim, Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Lee-Sup Kim
HPCA3
2022 EGCN: An Efficient GCN Accelerator for Minimizing Off-Chip Memory Access
abstract
As Graph Convolutional Networks (GCNs) have emerged as a promising solution for graph representation learning, designing specialized GCN accelerators has become an important challenge. An analysis of GCN workloads shows that the main bottleneck of GCN processing is not computation but the memory latency of intensive off-chip data transfer. Therefore, minimizing off-chip data transfer is the primary challenge for designing an efficient GCN accelerator. To address this challenge, optimization is initialized by considering GCNs as tiled matrix multiplication. In this paper, we optimize off-chip memory access from both the in- and out-of-tile perspectives. From the out-of-tile perspective, we find optimal tile configurations of given datasets and on-chip buffer capacity, then observe the dataflow across phases and layers. Inter-layer phase fusion dataflow with optimal tile configuration reduces data transfer of intermediate outputs. From the in-tile perspective, due to the sparsity of tiles, tiles have redundant data which does not participate in computation. Redundant data load is eliminated with hardware support. Finally, we introduce an efficient GCN inference accelerator, EGCN, specialized for minimizing off-chip memory access. EGCN achieves 41.9% off-chip DRAM access reduction, 1.49× speedup, and 1.95× energy efficiency improvement on average over the state-of-the-art accelerators.
Yunki Han, Kangkyu Park, Youngbeom Jung, Lee-Sup Kim
IEEE Trans. Computers1
2021 Deferred Dropout: An Algorithm-Hardware Co-Design DNN Training Method Provisioning Consistent High Activation Sparsity
abstract
This paper proposes a deep neural network training method that provisions consistent high activation sparsity and the ability to adjust the sparsity. To improve training performance, prior work reduces the memory footprint for training by exploiting input activation sparsity which is observed due to the ReLU function. However, the previous approach relies solely on the inherent sparsity caused by the function, and thus the footprint reduction is not guaranteed. In particular, models for natural language processing tasks like BERT do not use the function, so the models have almost zero activation sparsity and the previous approach loses its efficiency. In this paper, a new training method, Deferred Dropout, and its hardware architecture are proposed. With the proposed method, input activations are dropped out after the conventional forward-pass computation. In contrast to the conventional dropout where activations are zeroed before forward-pass computation, the dropping timing is deferred until the completion of the computation. Then, the sparsified activations are compressed and stashed in memory. This approach is based on our observation that networks preserve training quality even if only a few high magnitude activations are used in the backward pass. The hardware architecture enables designers to exploit the tradeoff between training quality and activation sparsity. Evaluation results demonstrate that the proposed method achieves 1.21-3.60 × memory footprint reduction and 1.06-1.43 x speedup on the TPUv3 architecture, compared to the prior work.
Kangkyu Park, Yunki Han, Lee-Sup Kim
ICCAD2