VLDB 2026 Research / reviewers in the wild / expert
Myeonggu Kang
dblp:217/8569
· DBLP profile ↗
15ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-3557-8526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AToM: Adaptive Token Merging for Efficient Acceleration of Vision TransformerabstractRecently, Vision Transformers (ViTs) have set a new standard in computer vision (CV), showing unparalleled image processing performance. However, their substantial computational requirements hinder practical deployment, especially on resource-limited devices common in CV applications. Token merging has emerged as a solution, condensing tokens with similar features to cut computational and memory demands. Yet, existing applications on ViTs often miss the mark in token compression, with rigid merging strategies and a lack of in-depth analysis of ViT merging characteristics. To overcome these issues, this paper introduces Adaptive Token Merging (AToM), a comprehensive algorithm-architecture co-design for accelerating ViTs. The AToM algorithm employs an image-adaptive, fine-grained merging strategy, significantly boosting computational efficiency. We also optimize the merging and unmerging processes to minimize overhead, employing techniques like First-Come-First-Merge mapping and Linear Distance Calculation. On the hardware side, the AToM architecture is tailor-made to exploit the AToM algorithm's benefits, with specialized engines for efficient merge and unmerge operations. Our pipeline architecture ensures end-to-end ViT processing, minimizing latency and memory overhead from the AToM algorithm. Across various hardware platforms including CPU, EdgeGPU, and GPU, AToM achieves average end-to-end speedups of 10.9$\boldsymbol{\times}$, 7.7$\boldsymbol{\times}$, and 5.4$\boldsymbol{\times}$, alongside energy savings of 24.9$\boldsymbol{\times}$, 1.8$\boldsymbol{\times}$, and 16.7$\boldsymbol{\times}$. Moreover, AToM offers 1.2$\boldsymbol{\times}$1.9$\boldsymbol{\times}$higher effective throughput compared to existing transformer accelerators. Jaekang Shin, Myeonggu Kang, Yunki Han, Lee-Sup Kim |
IEEE Trans. Computers | 2 |
| 2024 | Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability EstimationabstractThe attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency. Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Jaekang Shin, Lee-Sup Kim |
DAC | 2 |
| 2024 | ToEx: Accelerating Generation Stage of Transformer-Based Language Models via Token-Adaptive Early ExitabstractTransformer-based language models have recently gained popularity in numerous natural language processing (NLP) applications due to their superior performance compared to traditional algorithms. These models involve two execution stages: summarization and generation. The generation stage accounts for a significant portion of the total execution time due to its auto-regressive property, which necessitates considerable and repetitive off-chip accesses. Consequently, our objective is to minimize off-chip accesses during the generation stage to expedite transformer execution. To achieve the goal, we propose a token-adaptive early exit (ToEx) that generates output tokens using fewer decoders, thereby reducing off-chip accesses for loading weight parameters. Although our approach has the potential to minimize data communication, it brings two challenges: 1) inaccurate self-attention computation, and 2) significant overhead for exit decision. To overcome these challenges, we introduce a methodology that facilitates accurate self-attention by lazily performing computations for previously exited tokens. Moreover, we mitigate the overhead of exit decision by incorporating a lightweight output embedding layer. We also present a hardware design to efficiently support the proposed work. Evaluation results demonstrate that our work can reduce the number of decoders by 2.6× on average. Accordingly, it achieves 3.2× speedup on average compared to transformer execution without our work. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 1 |
| 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die ProcessingabstractTraining deep neural network (DNN) models is a resource-intensive, iterative process. For this reason, nowadays, complex optimizers like Adam are widely adopted as it increases the speed and efficiency of training. These optimizers, however, employ additional variables and raise the memory demand 2× to 3× of model parameters, worsening the memory capacity bottleneck. Moreover, as the size of DNN models is projected to grow even further, it is not practical to assume that the future models will fit in accelerator memory. This has triggered various efforts to offload models to flash-based storage. However, when the model, especially the optimizer, is offloaded to flash, the limited I/O bandwidth severely slows down the overall training process. To this end, we present OptimStore, a solid-state drive (SSD) system with on-die processing (ODP) architectures for gradient descent-based machine learning models. OptimStore accelerates the training process of such large-scale models by processing model optimization in the storage device, specifically inside the flash dies. ODP capability of OptimStore eliminates the heavy data movement over external interconnect and internal flash channels. Overall, OptimStore achieves, on average, a 2.8× speedup and a 3.6× improved energy efficiency in the weight update stage over baseline SSD offloading. Junkyum Kim, Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Lee-Sup Kim |
HPCA | 2 |
| 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERTabstractRecently, multiple transformer models, such as BERT, have been utilized together to support multiple natural language processing (NLP) tasks in a system, also known as multi-task BERT. Multi-task BERT with very high weight parameters increases the area requirement of a processing in resistive memory (ReRAM) architecture, and several works have attempted to address this model size issue. Despite the reduced parameters, the number of multi-task BERT computations remains the same, leading to massive energy consumption in ReRAM-based deep neural network (DNN) accelerators. Therefore, we suggest a framework for better energy efficiency during the ReRAM acceleration of multi-task BERT. First, we analyze the inherent redundancies of multi-task BERT and the computational properties of the ReRAM-based DNN accelerator, after which we propose what is termed the model generator, which produces optimal BERT models supporting multiple tasks. The model generator reduces multi-task BERT computations while maintaining the algorithmic performance. Furthermore, we present task scheduler, which adjusts the execution order of multiple tasks, to run the produced models efficiently. As a result, the proposed framework achieves maximally 4.4× higher energy efficiency over the baseline, and it can also be combined with the previous multi-task BERT works to achieve both a smaller area and higher energy efficiency. Myeonggu Kang, Hyein Shin, Junkyum Kim, Lee-Sup Kim |
IEEE Trans. Computers | 1 |
| 2023 | Fault-Free: A Framework for Analysis and Mitigation of Stuck-at-Fault on Realistic ReRAM-Based DNN AcceleratorsabstractResistive RAM(ReRAM) is gaining attention as a suitable memory platform for accelerating deep neural networks(DNNs) in an energy-efficient way. However, energy-efficient ReRAM-based DNN accelerators suffer from serious Stuck-At-Fault(SAF) issues that significantly degrade the inference accuracy. SAF is a device-level non-ideality, and the problems of SAF worsen in the realistic ReRAM with low cell resolution. To address the problem in the realistic ReRAM, we present a framework for mitigating SAF on ReRAM-based accelerators(Fault-free). We first analyze the impact of SAF on low-resolution cells. Based on the analysis, we present offline compilation, which drastically reduces the impact of SAF on inference accuracy. At the first stage, we extract indices of distorted weights due to SAF. For extracted weights, fault-aware weight decomposition and closest value mapping are applied to minimize the error of weights. In the online phase, the target DNN model is executed on the ReRAM-based accelerator along with lightweight compensation units. The online compensation is selectively performed for a small portion of weights to reduce the hardware overhead. With the proposed framework, the ReRAM-based accelerator successfully ensures the inference accuracy of various DNN models with an average area of 5% and an energy overhead of 0.8% from an ideal ReRAM-based accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
IEEE Trans. Computers | 2 |
| 2022 | Re2fresh: A Framework for Mitigating Read Disturbance in ReRAM-Based DNN AcceleratorsabstractA severe read disturbance problem degrades the inference accuracy of a resistive RAM (ReRAM) based deep neural network (DNN) accelerator. Refresh, which reprograms the ReRAM cells, is the most obvious solution for the problem, but programming ReRAM consumes huge energy. To address the issue, we first analyze the resistance drift pattern of each conductance state and the actual read stress applied to the ReRAM array by considering the characteristics of ReRAM-based DNN accelerators. Based on the analysis, we cluster ReRAM cells into a few groups for each layer of DNN and generate a proper refresh cycle for each group in the offline phase. The individual refresh cycles reduce energy consumption by reducing the number of unnecessary refresh operations. In the online phase, the refresh controller selectively launches refresh operations according to the generated refresh cycles. ReRAM cells are selectively refreshed by minimally modifying the conventional structure of the ReRAM-based DNN accelerator. The proposed work successfully resolves the read disturbance problem by reducing 97% of the energy consumption for the refresh operation while preserving inference accuracy. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 2 |
| 2022 | S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level SparsityabstractThe processing in-memory (PIM) approach that combines memory and processor appears to solve the memory wall problem. NAND flash memory, which is widely adopted in edge devices, is one of the promising platforms for PIM with its high-density property and the intrinsic ability for analog vector-matrix multiplication. Despite its potential, the domain conversion process, which converts an analog current to a digital value, accounts for most energy consumption on the NAND flash-based accelerator. It restricts the NAND flash memory usage for PIM compared to the other platforms. In this paper, we propose a NAND flash-based DNN accelerator to achieve both large memory density and energy efficiency among various platforms. As the NAND flash memory already shows higher memory density than other memory platforms, we aim to enhance energy efficiency by reducing the domain conversion process burden. Firstly, we optimize the bit width of partial multiplication by considering the analog-to-digital converter (ADC) resource. For further optimization, we propose a methodology to exploit many zero partial multiplication results for enhancing both energy efficiency and throughput. The proposed work successfully exploits the bit-level sparsity of DNN, which results in achieving up to 8.6/8.2 larger energy efficiency/throughput over the provisioned baseline. Myeonggu Kang, Hyeonuk Kim, Hyein Shin, Jaehyeong Sim, Kyeonghan Kim, Lee-Sup Kim |
IEEE Trans. Computers | 1 |
| 2022 | A Framework for Accelerating Transformer-Based Language Model on ReRAM-Based ArchitectureabstractTransformer-based language models have become thede-factostandard model for various natural language processing (NLP) applications given the superior algorithmic performances. Processing a transformer-based language model on a conventional accelerator induces the memory wall problem, and the ReRAM-based accelerator is a promising solution to this problem. However, due to the characteristics of the self-attention mechanism and the ReRAM-based accelerator, the pipeline hazard arises when processing the transformer-based language model on the ReRAM-based accelerator. This hazard issue greatly increases the overall execution time. In this article, we propose a framework to resolve the hazard issue. First, we propose the concept of window self-attention to reduce the attention computation scope by analyzing the properties of the self-attention mechanism. After that, we present a window-size search algorithm, which finds an optimal window size set according to the target application/algorithmic performance. We also suggest a hardware design that exploits the advantages of the proposed algorithm optimization on the general ReRAM-based accelerator. The proposed work successfully alleviates the hazard issue while maintaining the algorithmic performance, leading to a$5.8\times $speedup over the provisioned baseline. It also delivers up to$39.2\times /643.2\times $speedup/higher energy efficiency over GPU, respectively. Myeonggu Kang, Hyein Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Fault-free: A Fault-resilient Deep Neural Network Accelerator based on Realistic ReRAM DevicesabstractEnergy-efficient Resistive RAM (ReRAM) based deep neural network (DNN) accelerator suffers from severe Stuck-At-Fault (SAF) problem that drastically degrades the inference accuracy. The SAF problem gets even worse in realistic ReRAM devices with low cell resolution. To address the issue, we propose a fault-resilient DNN accelerator based on realistic ReRAM devices. We first analyze the SAF problem in a realistic ReRAM device and propose a 3-stage offline fault-resilient compilation and lightweight online compensation. The proposed work enables the reliable execution of DNN with only 5% area and 0.8% energy overhead from the ideal ReRAM-based DNN accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
DAC | 2 |
| 2021 | Optimizing ADC Utilization through Value-Aware Bypass in ReRAM-based DNN AcceleratorabstractReRAM-based Processing-In-Memory (PIM) has been widely studied as a promising approach for Deep Neural Networks (DNN) accelerator with its energy-efficient analog operations. However, the domain conversion process for the analog operation requires frequent accesses to power-hungry Analog-to-Digital Converter (ADC), hindering the overall energy efficiency. Although previous research has been suggested to address this problem, the ADC cost has not been sufficiently reduced because of its unsuitable approach for ReRAM. In this paper, we propose mixed-signal-based value-aware bypass techniques to optimize the ADC utilization of the ReRAM-based PIM. By utilizing the property of bit-line (BL) level value distribution, the proposed work bypasses the redundant ADC operations depending on the magnitude of value. Evaluation results show that our techniques successfully reduce ADC access and improve overall energy efficiency by 2.48 × -3.07 × compared to ISAAC. HanCheon Yun, Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
DAC | 3 |
| 2021 | A Framework for Area-efficient Multi-task BERT Execution on ReRAM-based AcceleratorsabstractWith the superior algorithmic performances, BERT has become the de-facto standard model for various NLP tasks. Accordingly, multiple BERT models have been adopted on a single system, which is also called multi-task BERT. Although the ReRAM-based accelerator shows the sufficient potential to execute a single BERT model by adopting in-memory computation, processing multi-task BERT on the ReRAM-based accelerator extremely increases the overall area due to multiple fine-tuned models. In this paper, we propose a framework for area-efficient multi-task BERT execution on the ReRAM-based accelerator. Firstly, we decompose the fine-tuned model of each task by utilizing the base-model. After that, we propose a two-stage weight compressor, which shrinks the decomposed models by analyzing the properties of the ReRAM-based accelerator. We also present a profiler to generate hyper-parameters for the proposed compressor. By sharing the base-model and compressing the decomposed models, the proposed framework successfully reduces the total area of the ReRAM-based accelerator without an additional training procedure. It achieves a 0.26 x area than baseline while maintaining the algorithmic performances. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
ICCAD | 1 |
| 2020 | A Thermal-aware Optimization Framework for ReRAM-based Deep Neural Network AccelerationabstractResistive RAM (ReRAM) is widely regarded as a promising platform for deep neural network (DNN) acceleration. However, the ReRAM device suffers from severe thermal problems that degrade the lifetime and inference accuracy of the ReRAM-based DNN accelerator. To address the issues, we propose a thermal-aware optimization framework for accelerating DNN on ReRAM (TOPAR). TOPAR includes 3-stage offline thermal optimization and online thermal-aware error compensation. Offline thermal optimization consists of thermal-aware weight decomposition, thermal-aware column reordering, and fine-grained weight adjustment to reduce the temperature of the ReRAM-based DNN accelerator. For online thermal-aware error compensation, we compensate conductance change according to the temperature variation. With TOPAR, the endurance degradation due to temperature rise improves up to 2.39×, and inference accuracy is preserved without harming the performance of the ReRAM-based DNN accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 2 |
| 2019 | An Energy-efficient Processing-in-memory Architecture for Long Short Term Memory in Spin Orbit Torque MRAMabstractMany recent studies have focused on Processing-in-memory (PIM) architectures for neural networks to resolve the memory bottleneck problem. Especially, an increased interest in Spin Orbit Torque (SOT)-MRAMs has emerged due to its low latency, high energy efficiency, and non-volatility. However, the previous work added extra computing circuits to support complicated computations, which results in large energy overheads. In this work, we propose a new PIM architecture with relatively small peripheral circuit, which produces the highest energy efficiency for processing a Long Short Term Memory (LSTM) among the PIM architectures. We improve the efficiency with a new computing method for logical operations, which exploits characteristics of SOT-MRAMs. We reduce the number of word lines (WLs) activated concurrently to one from two in the previous works. As a result, the energy for driving WLs is saved, and the sensing current for computation is reduced. Moreover, we propose efficient methods for additions, multiplications and non-linear activation functions in memory to process an LSTM. Accordingly, we achieve 1.26x energy efficiency with the proposed computing method for logical operations compared to the previous study based on SOT-MRAMs and up to 5.54x energy efficiency over the previous PIM architectures based on other memories. Kyeonghan Kim, Hyein Shin, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 4 |
| 2018 | TrainWare: A Memory Optimized Weight Update Architecture for On-Device Convolutional Neural Network TrainingabstractTraining convolutional neural network on device has become essential where it allows applications to consider user's individual environment. Meanwhile, the weight update operation from the training process is the primary factor of high energy consumption due to its substantial memory accesses. We propose a dedicated weight update architecture with two key features: (1) a specialized local buffer for the DRAM access deduction (2) a novel dataflow and its suitable processing element array structure for weight gradient computation to optimize the energy consumed by internal memories. Our scheme achieves 14.3%-30.2% total energy reduction by drastically eliminating the memory accesses. Seungkyu Choi, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim |
ISLPED | 3 |