EDBT 2026 Demo / reviewers in the wild / expert
Jaekang Shin
dblp:241/4263
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-5943-1599ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AToM: Adaptive Token Merging for Efficient Acceleration of Vision TransformerabstractRecently, Vision Transformers (ViTs) have set a new standard in computer vision (CV), showing unparalleled image processing performance. However, their substantial computational requirements hinder practical deployment, especially on resource-limited devices common in CV applications. Token merging has emerged as a solution, condensing tokens with similar features to cut computational and memory demands. Yet, existing applications on ViTs often miss the mark in token compression, with rigid merging strategies and a lack of in-depth analysis of ViT merging characteristics. To overcome these issues, this paper introduces Adaptive Token Merging (AToM), a comprehensive algorithm-architecture co-design for accelerating ViTs. The AToM algorithm employs an image-adaptive, fine-grained merging strategy, significantly boosting computational efficiency. We also optimize the merging and unmerging processes to minimize overhead, employing techniques like First-Come-First-Merge mapping and Linear Distance Calculation. On the hardware side, the AToM architecture is tailor-made to exploit the AToM algorithm's benefits, with specialized engines for efficient merge and unmerge operations. Our pipeline architecture ensures end-to-end ViT processing, minimizing latency and memory overhead from the AToM algorithm. Across various hardware platforms including CPU, EdgeGPU, and GPU, AToM achieves average end-to-end speedups of 10.9$\boldsymbol{\times}$, 7.7$\boldsymbol{\times}$, and 5.4$\boldsymbol{\times}$, alongside energy savings of 24.9$\boldsymbol{\times}$, 1.8$\boldsymbol{\times}$, and 16.7$\boldsymbol{\times}$. Moreover, AToM offers 1.2$\boldsymbol{\times}$1.9$\boldsymbol{\times}$higher effective throughput compared to existing transformer accelerators. Jaekang Shin, Myeonggu Kang, Yunki Han, Lee-Sup Kim |
IEEE Trans. Computers | 1 |
| 2024 | Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability EstimationabstractThe attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency. Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Jaekang Shin, Lee-Sup Kim |
DAC | 5 |
| 2024 | CoCoA: Algorithm-Hardware Co-Design for Large-Scale GNN Training using Compressed GraphabstractScaling Graph Neural Network (GNN) training on large-scale graph data poses a critical challenge for implementing GNN applications on real-world giant graphs. The size of real-world graphs often exceeds the memory capacity of accelerator devices, necessitating the use of multiple devices or host memory for training. While expanding memory space alleviates the out-of-memory problem, this approach introduces another bottleneck through heavy communication via low-bandwidth interconnection. Therefore, achieving efficient, scalable GNN training on large-scale graphs requires addressing both capacity and communication issues. Yunki Han, Jaekang Shin, Gunhee Park, Lee-Sup Kim |
ICCAD | 2 |
| 2024 | ToEx: Accelerating Generation Stage of Transformer-Based Language Models via Token-Adaptive Early ExitabstractTransformer-based language models have recently gained popularity in numerous natural language processing (NLP) applications due to their superior performance compared to traditional algorithms. These models involve two execution stages: summarization and generation. The generation stage accounts for a significant portion of the total execution time due to its auto-regressive property, which necessitates considerable and repetitive off-chip accesses. Consequently, our objective is to minimize off-chip accesses during the generation stage to expedite transformer execution. To achieve the goal, we propose a token-adaptive early exit (ToEx) that generates output tokens using fewer decoders, thereby reducing off-chip accesses for loading weight parameters. Although our approach has the potential to minimize data communication, it brings two challenges: 1) inaccurate self-attention computation, and 2) significant overhead for exit decision. To overcome these challenges, we introduce a methodology that facilitates accurate self-attention by lazily performing computations for previously exited tokens. Moreover, we mitigate the overhead of exit decision by incorporating a lightweight output embedding layer. We also present a hardware design to efficiently support the proposed work. Evaluation results demonstrate that our work can reduce the number of decoders by 2.6× on average. Accordingly, it achieves 3.2× speedup on average compared to transformer execution without our work. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 4 |
| 2023 | Accelerating On-Device DNN Training Workloads via Runtime Convergence MonitorabstractWith the growing demand for processing deep learning applications on edge devices, on-device DNN training has become a major workload to execute a variety of vision tasks suited for users. Therefore, architectures employed with algorithm co-design to accelerate the training process have been steadily studied. However, previous solutions are mostly supported by extended versions of the inference studies, such as sparsity, data flow, quantization, etc. Moreover, most works examine their schemes on from-the-scratch training that cannot tolerate inaccurate computing. Accordingly, there are still factors that hinder the overall speed of the DNN training process that has not been addressed in practical workloads. In this work, we propose a runtime convergence monitor to achieve massive computational savings in the practical on-device training workloads (i.e., transfer-learning-based task adaptation). By monitoring the network output data, we determine the training intensity of incoming tasks and adaptively detect the convergence in iteration intervals for training diverse datasets. Furthermore, we enable the computation skip of converged images determined by the monitored prediction probability to enhance the training speed within an iteration. As a result, we perform an accurate but fast convergence in model training for the task adaptation with minimal overhead. Unlike the previous approximation methods, our monitoring system enables runtime optimization and can be easily applicable to any type of accelerator attaining significant speedup. Evaluation results on various datasets show geomean of$2.2\times $speedup when applied in any systolic architectures and further enhancement of$3.6\times $when applied in accelerators dedicated for on-device training. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Energy-Efficient CNN Personalized Training by Adaptive Data ReformationabstractTo adopt deep neural networks in resource-constrained edge devices, various energy- and memory-efficient embedded accelerators have been proposed. However, most off-the-shelf networks are well trained with vast amounts of data, but unexplored users’ data or accelerator’s constraints can lead to unexpected accuracy loss. Therefore, a network adaptation suitable for each user and device is essential to make a high confidence prediction in given environment. We propose simple but efficient data reformation methods that can effectively reduce the communication cost with off-chip memory during the adaptation. Our proposal utilizes the data’s zero-centered distribution and spatial correlation to concentrate the sporadically spread bit-level zeros to the units of value. Consequently, we reduced the communication volume by up to 55.6% per task with an area overhead of 0.79% during the personalization training. Youngbeom Jung, Hyeonuk Kim, Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Algorithm/architecture co-design for energy-efficient acceleration of multi-task DNNabstractReal-world AI applications, such as augmented reality or autonomous driving, require processing multiple CV tasks simultaneously. However, the enormous data size and the memory footprint have been a crucial hurdle for deep neural networks to be applied in resource-constrained devices. To solve the problem, we propose an algorithm/architecture co-design. The proposed algorithmic scheme, named SqueeD, reduces per-task weight and activation size by 21.9x and 2.1x, respectively, by sharing those data between tasks. Moreover, we design architecture and dataflow to minimize DRAM access by fully utilizing benefits from SqueeD. As a result, the proposed architecture reduces the DRAM access increment and energy consumption increment per task by 2.2x and 1.3x, respectively. Jaekang Shin, Seungkyu Choi, Jongwoo Ra, Lee-Sup Kim |
DAC | 1 |
| 2022 | A Deep Neural Network Training Architecture With Inference-Aware Heterogeneous Data-TypeabstractAs deep learning applications often encounter accuracy degradation due to the distorted inputs from a variety of environmental conditions, training with personal data has become essential for the edge devices. Hence, ‘training on edge’ by supporting a trainable deep learning accelerator has been actively studied. Nevertheless, previous research does not consider the fundamental datapath for training and the importance of retaining the high performance for inference tasks. In this work, we propose NeuroFlix, a deep neural network training accelerator supporting heterogeneous data-type of floating- and fixed-point for input operands. From two perspectives: 1) separate precision decision for each input data, 2) maintenance of high performance on inference, we configure the data with low-bit fixed-point of activation/weight and floating-point based error gradient securing up to half-precision. A novel MAC architecture is designed to compute low- and high-precision modes for the different input combinations. By substituting a high-cost floating-point based addition to brick-level separate accumulations, we realize both area-efficient architecture and high throughput for low-precision computation. Consequently, NeuroFlix outperforms the previous architectures of state-of-the-art configurations proving its high efficiency in both training and inference. By also comparing with the off-the-shelf bfloat16-based accelerator, it achieves 1.2 ×/2.0 × of speedup/energy-efficiency at training and further enhancement of 3.6 ×/4.5 × at inference. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 2 |
| 2021 | A Convergence Monitoring Method for DNN Training of On-Device Task AdaptationabstractDNN training has become a major workload in on-device situations to execute various vision tasks with high performance. Accordingly, training architectures accompanying approximate computing have been steadily studied for efficient acceleration. However, most of the works examine their scheme on from-the-scratch training where inaccurate computing is not tolerable. Moreover, previous solutions are mostly provided as an extended version of the inference works, e.g., sparsity/pruning, quantization, dataflow, etc. Therefore, unresolved issues in practical workloads that hinder the total speed of the DNN training process remain still. In this work, with targeting the transfer learning-based task adaptation of the practical on-device training workload, we propose a convergence monitoring method to resolve the redundancy in massive training iterations. By utilizing the network's output value, we detect the training intensity of incoming tasks and monitor the prediction convergence with the given intensity to provide early-exits in the scheduled training iteration. As a result, an accurate approximation over various tasks is performed with minimal overhead. Unlike the sparsity-driven approximation, our method enables runtime optimization and can be easily applicable to off-the-shelf accelerators achieving significant speedup. Evaluation results on various datasets show a geomean of$2.2\times$speedup over baseline and$1.8\times$speedup over the latest convergence-related training method. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
ICCAD | 2 |
| 2021 | A Framework for Area-efficient Multi-task BERT Execution on ReRAM-based AcceleratorsabstractWith the superior algorithmic performances, BERT has become the de-facto standard model for various NLP tasks. Accordingly, multiple BERT models have been adopted on a single system, which is also called multi-task BERT. Although the ReRAM-based accelerator shows the sufficient potential to execute a single BERT model by adopting in-memory computation, processing multi-task BERT on the ReRAM-based accelerator extremely increases the overall area due to multiple fine-tuned models. In this paper, we propose a framework for area-efficient multi-task BERT execution on the ReRAM-based accelerator. Firstly, we decompose the fine-tuned model of each task by utilizing the base-model. After that, we propose a two-stage weight compressor, which shrinks the decomposed models by analyzing the properties of the ReRAM-based accelerator. We also present a profiler to generate hyper-parameters for the proposed compressor. By sharing the base-model and compressing the decomposed models, the proposed framework successfully reduces the total area of the ReRAM-based accelerator without an additional training procedure. It achieves a 0.26 x area than baseline while maintaining the algorithmic performances. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
ICCAD | 3 |
| 2020 | A Pragmatic Approach to On-device Incremental Learning System with Selective Weight UpdatesabstractIncremental learning is drawing attention to widen capabilities of device-AI. Previous works have researched to reduce numerous computations and memory accesses required for the training process of IL, but they could not show a noticeable improvement in the weight gradient computation (WGC) phase. Therefore, we propose a selective weight update technique that searches for critical weights to be updated by applying the IL algorithm that training per-task binary masks. Also, we introduce a novel dataflow for the implementation of selective WGC on typical NPUs with minimum overheads. On average, our system shows a 2.9× speed up and 2.5× energy efficiency in WGC without degrading training quality. Jaekang Shin, Seungkyu Choi, Yeongjae Choi, Lee-Sup Kim |
DAC | 1 |
| 2019 | An Optimized Design Technique of Low-bit Neural Network Training for Personalization on IoT DevicesabstractPersonalization by incremental learning has become essential for IoT devices to enhance the performance of the deep learning models trained with global datasets. To avoid massive transmission traffic in the network, exploiting on-device learning is necessary. We propose a software/hardware co-design technique that builds an energy-efficient low-bit trainable system: (1) software optimizations by local low-bit quantization and computation freezing to minimize the on-chip storage requirement and computational complexity, (2) hardware design of a bit-flexible multiply-and-accumulate (MAC) array sharing the same resources in inference and training. Our scheme saves 99.2% on on-chip buffer storage and achieves 12.8x higher peak energy efficiency compared to previous trainable accelerators. Seungkyu Choi, Jaekang Shin, Yeongjae Choi, Lee-Sup Kim |
DAC | 2 |