Kwangrae Kim

dblp:247/3453 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TiMM-ViT: Token-Importance-Aware Token Merging with Multi-Exit Training and an FPGA-based Vision Transformer Accelerator for Image Classification
abstract
Token merging (ToMe) reduces the computational cost of Vision Transformers (ViTs) by merging similar tokens. However, ToMe may inadvertently merge critical tokens, leading to accuracy degradation. Furthermore, the sequential execution of ViT and ToMe operations incurs increased latency on edge GPUs. To address these challenges, we propose TiMM, a token-importance-aware merging method that preserves essential tokens by considering both token importance and similarity. We also present a dedicated FPGA-based accelerator designed to exploit the concurrent execution between ViT and TiMM. Experimental results across DeiT-Tiny, Small, and Base demonstrate that TiMM achieves accuracy improvements of 0.23%-0.73% over ToMe, respectively, within a 1% accuracy drop margin relative to the baseline. When implemented on the Xilinx ZCU102, it demonstrates speedups of 1.45×-10×, along with energy efficiency gains of 1.03×-3.1× compared to the NVIDIA Jetson TX2.
Soomin Rho, Sangki Park, Kwangrae Kim, Min-Gwon Song, Ki-Seok Chung
ISLPED4
2024 FloatMax: An Efficient Accelerator for Transformer-Based Models Exploiting Tensor-Wise Adaptive Floating-Point Quantization
abstract
The rapid growth of the Transformer model size results in significant computational resources and memory requirements. To mitigate this complexity, quantization to reduce the bit width to represent numbers is being actively studied. However, prior quantization methods that use integers of 8 bits or less suffer from accuracy loss due to lower resolution because tensors of the Transformer model often have a non-uniform distribution containing outliers. In this paper, we propose a novel quantization method that utilizes tensor-wise adaptive floating-point quantization. Two key strategies address the issue of accuracy degradation. First, we leverage the characteristics of the floating-point data type, which provides higher precision for normal values and lower precision for outlier values. Second, since the degree of non-uniformity and outliers in each tensor vary from layer to layer, we adaptively assign different bit widths to the exponent and the mantissa of the floating-point representation for each tensor, quantizing them with different levels of precision. This approach allows high accuracy across the tensor distribution while efficiently managing outliers. We design the FloatMax processing element and decoders to efficiently carry out these floating-point computations. In addition, FloatMax is integrated into a systolic array to accelerate the linear layer and attention mechanism. Evaluation results show that FloatMax achieves 0.75×, 0.85×, and 0.79× area reduction and 1.70×, 1.84×, and 1.89× average performance improvement over Olive, ANT, and AdaFloat, respectively, without accuracy loss.
Seoho Chung, Kwangrae Kim, Soomin Rho, Chanhoon Kim, Ki-Seok Chung
ICCD2
2023 Dynamic Partitioning Method for Near-Memory Parallel Processing of Sparse Matrix-Vector Multiplication
abstract
Near-memory processing (NMP), which places lightweight processing units near the DRAM memory, has been actively studied to speed up the execution of memory-intensive applications by reducing the amount of data traffic between the DRAM and the CPU. Sparse matrix-vector multiplication (SpMV) is a representative memory-bound kernel used in various applications such as graph analytics, scientific computing, and machine learning. There are prior works to accelerate SpMV by NMP employing a fixed partitioning scheme that hides random access of SpMV using a parallel NMP core. However, due to the various distributions of the matrix, the fixed partitioning of prior works causes a load imbalance in which a sparse matrix is unevenly allocated to processing units for NMP. To resolve this, dynamic partitioning methods to distribute matrices and vectors to the NMP processing units can be effective. In this paper, we propose a dynamic partitioning algorithm (DPA) that analyzes the distribution of non-zero elements in a sparse matrix to classify it into three types (even distribution, skewed distribution, and power-law distribution) and partitions the matrix according to each distribution. Our proposed distribution scheme alleviates load imbalance by up to 73% when compared to static distribution schemes, and such improvement achieves an average speed-up 1.37x (up to 1.84x) over the NMP architecture with static distribution schemes.
Dae-Eun Wi, Kwangrae Kim, Ki-Seok Chung
IECON2
2021 HammerFilter: Robust Protection and Low Hardware Overhead Method for RowHammer
abstract
The continuous scaling-down of the dynamic random access memory (DRAM) manufacturing process has made it possible to improve DRAM density. However, it makes small DRAM cells susceptible to electromagnetic interference between nearby cells. Unless DRAM cells are adequately isolated from each other, the frequent switching access of some cells may lead to unintended bit flips in adjacent cells. This phenomenon is commonly referred to as RowHammer. It is often considered a security issue because unusually frequent accesses to a small set of rows generated by malicious attacks can cause bit flips. Such bit flips may also be caused by general applications. Although several solutions have been proposed, most approaches either incur excessive area overhead or exhibit limited prevention capabilities against maliciously crafted attack patterns. Therefore, the goals of this study are (1) to mitigate RowHammer, even when the number of aggressor rows increases and attack patterns become complicated, and (2) to implement the method with a low area overhead.We propose a robust hardware-based protection method for RowHammer attacks with a low hardware cost called HammerFilter, which employs a modified version of the counting bloom filter. It tracks all attacking rows efficiently by leveraging the fact that the counting bloom filter is a space-efficient data structure, and we add an operation, HALF-DELETE, to mitigate the energy overhead. According to our experimental results, the proposed method can completely prevent bit flips when facing artificially crafted attack patterns (five patterns in our experiments), whereas state-of-the-art probabilistic solutions can only mitigate less than 56% of bit flips on average. Furthermore, the proposed method has a much lower area cost compared to existing counter-based solutions (40.6× better than TWiCe and 2.3× better than Graphene).
Kwangrae Kim, Jeonghyun Woo, Junsu Kim 0004, Ki-Seok Chung
ICCD1