EDBT 2026 Demo / reviewers in the wild / expert
Eunji Kwon
dblp:261/0407
· DBLP profile ↗
14ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-5313-8976ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Down-sampling in Hybrid Neural Networks using Adversarial AutoencodersabstractEarly convolutional layers in hybrid neural networks enable efficient down-sampling but pose a significant burden on inference latency and energy consumption. We propose a method to replace the conventional down-sampling block with lightweight autoencoders to enhance applicability in edge devices. We introduce an adversarial training strategy to align the autoencoder’s latent features with the original stem, ensuring compatibility with succeeding layers. Applying our method to MobileViTV2-050 yields a 1.23x speedup and a 47% reduction in Energy-Delay Product (EDP) with only a 1.0% accuracy drop on ImageNet-1K. Jonghyeon Nam, JoonSeok Kim, Eunji Kwon, Seokhyeong Kang |
DATE | 3 |
| 2026 | Autonomous Model Quantization Framework for Hybrid Vision Transformers Based on Reinforcement LearningabstractExisting quantization approaches often suffer from significant accuracy degradation when compressing hybrid convolution and transformer models with low bit-width. This paper presents RL-PTQv2, an extension of our previous RL-PTQ framework [1], which introduces a new reinforcement learning (RL)-based post-training quantization (PTQ) method. RL-PTQv2 introduces two key advances: (i) hardware (HW)-aware PTQ (optional), where RL is guided by real latency and energy feedback from an in-loop PIM simulator, enabling deployable designs that jointly optimize accuracy, latency, and energy, and (ii) improved quantization techniques, supporting symmetric/ asymmetric quantization and mixed adaptive rounding to better balance precision and efficiency. Across various hybrid vision transformer families, including MobileViTv1 and v2 [2], [3], EfficientFormerv1 and v2 [4], [5], and MobileFormer [6], RL-PTQv2 achieves state-of-the-art quantized accuracy compared to previous PTQ methods [7], [8], [9], [10]. Furthermore, our quantized model showed an improvement in energy efficiency of 10.1× on TransPIM [11] and 22.6× on the Titan RTX GPU compared to the baseline model, specifically when deployed on HViT-PIM, a dedicated processing framework for efficiently executing MobileViT models. HViT-PIM was developed primarily to explore the potential of HW-aware PTQ. However, the RL-PTQv2 is not limited to processing-in-memory (PIM). It can also be seamlessly integrated with a variety of bit-serial accelerators, enabling automatic quantization tailored to the underlying HW. Eunji Kwon, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | RL-PTQ: RL-based Mixed Precision Quantization for Hybrid Vision TransformersabstractExisting quantization approaches incur significant accuracy loss when compressing hybrid convolution and transformer models with low bit-width. This paper presents RL-PTQ, a novel post-training quantization (PTQ) framework utilizing reinforcement learning (RL). Our focus is on determining the most effective bit-width and observer for quantization configurations tailored for mixed precision by grouping layers and addressing the challenges of quantization of hybrid transformers. We achieved the highest quantized accuracy for MobileViTs compared to the previous PTQ methods [5--7]. Furthermore, our quantized model on Processing In Memory (PIM) architecture exhibited an energy efficiency enhancement of 10.1× and 22.6× compared to the baseline model, on the state-of-the-art PIM accelerator [15] and GPU, respectively. Eunji Kwon, Minxuan Zhou, Tajana Rosing, Seokhyeong Kang |
DAC | 1 |
| 2024 | ViT- ToGo: Vision Transformer Accelerator with Grouped Token PruningabstractVision Transformer ($V$iT) has gained prominence for its performance in various vision tasks but comes with considerable computational and memory demands, posing a challenge when deploying it on resource-constrained edge devices. To address this limitation, various token pruning methods have been proposed to reduce the computation. However, the majority of token pruning techniques do not account for practical use in actual embedded devices, which demand a significant reduction in computational load. In this paper, we introduce ViT-ToGo, a$V$iT accelerator with grouped token pruning. This enables the parallel execution of the$V$iT models and the token pruning process. We implement grouped token pruning with a head-wise importance estimator which simplifies the process need for token pruning, including sorting and reordering. Our proposed method achieves up to 66 % reduction in the number of tokens, resulting in up to 36% reduction in GFLOPs, with only a minimal accuracy drop of around 1 %. Furthermore, the hardware implementation incurs a marginal resource overhead of 1.13% in average. Seungju Lee, Kyumin Cho, Eunji Kwon, Sejin Park 0001, Seojeong Kim, Seokhyeong Kang |
DATE | 3 |
| 2024 | Mobile Transformer Accelerator Exploiting Various Line Sparsity and Tile-Based Dynamic QuantizationabstractTransformer models are difficult to employ in mobile devices due to their memory-and computation-intensive properties. Accordingly, there is ongoing research on various methods for compressing transformer models, such as pruning and quantization. However, general computing platforms such as central processing units (CPUs) and graphics processing units (GPUs) are not energy-efficient to accelerate the pruned model because the unstructured sparsity they exhibit causes degradation of parallelism. In this paper, we propose a low-power accelerator for transformers that can handle various levels of structured sparsity induced by line pruning with different granularity. Our approach accelerates pruned transformers in a head-wise and line-wise manner. We present a head reorganization and shuffling method that supports head-wise skip operations and resolves the load imbalance problem among processing engines (PEs) caused by the varying number of operations in each head. Furthermore, we implemented a sparse quantized general matrix-to-matrix multiplication (SQ-GEMM) module that supports line-wise skipping and on-the-fly tile-based dynamic quantization of activations. As a result, compared to mobile GPU and CPU, the proposed accelerator improved the energy efficiency by 2.9× and 12.3× for the detection transformer (DETR), and 3.0× and 12.4× for the vision transformer (ViT) models, respectively. In addition, our proposed mobile accelerator achieved the highest energy efficiency among the current state-of-the-art FPGA-based transformer accelerators. Eunji Kwon, Jongho Yoon 0001, Seokhyeong Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Mobile Accelerator Exploiting Sparsity of Multi-Heads, Lines, and Blocks in Transformers in Computer VisionabstractIt is difficult to employ transformer models for computer vision in mobile devices due to their memory- and computation-intensive properties. Accordingly, there is ongoing research on various methods for compressing transformer models, such as pruning. However, general computing platforms such as central processing units (CPUs) and graphics processing units (GPUs) are not energy-efficient to accelerate the pruned model due to their structured sparsity. This paper proposes a low-power accelerator for transformers with various sizes of structured sparsity induced by pruning with different granularity. In this study, we can accelerate a transformer that has been pruned in a head-wise, line-wise, or block-wise manner. We developed a head scheduling algorithm to support head-wise skip operations and resolve the processing engine (PE) load imbalance problem caused by different number of operations in one head. Moreover, we implemented a sparse general matrix-to-matrix multiplication (sparse GEMM) module that supports line-wise and block-wise skipping. As a result, when compared with a mobile GPU and mobile CPU respectively, our proposed accelerator achieved$6.1\times$and$13.6\times$improvements in energy efficiency for the detection transformer (DETR) model and achieved approximately$2.6\times$and$7.9\times$improvements in the energy efficiency on average for the vision transformer (ViT) models. Eunji Kwon, Haena Song, Seokhyeong Kang |
DATE | 1 |
| 2023 | FPGA-Based Accelerator for Rank-Enhanced and Highly-Pruned Block-Circulant Neural NetworksabstractNumerous network compression methods have been proposed to deploy deep neural networks in a resource-constrained embedded system. Among them, block-circulant matrix (BCM) compression is one of the promising hardware-friendly methods for both acceleration and compression. However, it has several limitations; (i) limited representation due to the structural characteristic of circulant matrix, (ii) limitation of the compression parameter, (iii) need to specialize the dataflow for BCM-compressed network accelerators. In this paper, rank-enhanced and highly-pruned block-circulant matrices compression (RP-BCM) framework is proposed to overcome these limitations. RP-BCM comprises two stages: Hadamard-BCM and BCM-wise pruning. Moreover, a dedicated skip scheme is introduced to processing element design for exploiting high-parallelism with BCM-wise sparsity. Furthermore, we propose specialized dataflow for a BCM-compressed network on a resource-constrained FPGA. As a result, the proposed method achieves parameter reduction and FLOPs reduction for ResNet-50 in ImageNet by 92.4% and 77.3%, respectively. Moreover, the proposed hardware design achieves$3.1\times$improvement in energy efficiency on the Xilinx PYNQ-Z2 FPGA board for ResNet-18 on ImageNet compared to the GPU. Haena Song, Jongho Yoon 0001, Eunji Kwon, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 4 |
| 2022 | Adaptive FSP: Adaptive Architecture Search with Filter Shape Pruning
Aeri Kim, Seungju Lee, Eunji Kwon, Seokhyeong Kang |
ACCV (1) | 3 |
| 2021 | Approach to Improve the Performance Using Bit-level Sparsity in Neural NetworksabstractThis paper presents a convolutional neural network (CNN) accelerator that can skip zero weights and handle outliers, which are few but have a significant impact on the accuracy of CNNs, to achieve speedup and increase the energy efficiency of CNN. We propose an offline weight-scheduling algorithm which can skip zero weights and combine two non-outlier weights simultaneously using bit-level sparsity of CNNs. We use a reconfigurable multiplier-and-accumulator (MAC) unit for two purposes; usually used to compute combined two non-outliers and sometimes to compute outliers. We further improve the speedup of our accelerator by clipping some of the outliers with negligible accuracy loss. Compared to DaDianNao [7] and Bit-Tactical [16] architectures, our CNN accelerator can improve the speed by 3.34 and 2.31 times higher and reduce energy consumption by 29.3% and 30.2%, respectively. Yesung Kang, Eunji Kwon, Seunggyu Lee, Younghoon Byun, Youngjoo Lee 0002, Seokhyeong Kang |
DATE | 2 |
| 2021 | MDARTS: Multi-objective Differentiable Neural Architecture SearchabstractIn this work, we present a differentiable neural architecture search (NAS) method that takes into account two competing objectives, quality of result (QoR) and quality of service (QoS) with hardware design constraints. NAS research has recently received a lot of attention due to its ability to automatically find architecture candidates that can outperform handcrafted ones. However, the NAS approach which complies with actual HW design constraints has been under-explored. A naive NAS approach for this would be to optimize a combination of two criteria of QoR and QoS, but the simple extension of the prior art often yields degenerated architectures, and suffers from a sensitive hyperparameter tuning. In this work, we propose a multi-objective differential neural architecture search, called MDARTS. MDARTS has an affordable search time and can find Pareto frontier of QoR versus QoS. We also identify the problematic gap between all the existing differentiable NAS results and those final post-processed architectures, where soft connections are binarized. This gap leads to performance degradation when the model is deployed. To mitigate this gap, we propose a separation loss that discourages indefinite connections of components by implicitly minimizing entropy. Hyun-jeong Kwon, Eunji Kwon, Youngchang Choi, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 3 |
| 2021 | Reinforcement Learning-Based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that utilizes reinforcement learning to increase the power efficiency of mobile device systems based on a multiprocessor system-on-a-chip (MPSoC). The proposed policy predicts a system’s characteristics and learns power management controls to adapt to the variations in the system. We consider the behavioral characteristics of systems that run on mobile devices under diverse scenarios. Therefore, the policy can flexibly manage the system power regardless of the application scenario and achieve lower energy consumption without compromising the user satisfaction. The average energy per unit quality of service (QoS) of the proposed policy is lower than that of the previous six dynamic voltage/frequency scaling governors by 31.66%. Furthermore, we reduce the runtime overhead by implementing the proposed policy as hardware. We implemented the policy on the field programmable gate array (FPGA) and construct a communication interface between the central processing units (CPUs) and the hardware of the proposed policy. Decision-making by the hardware-implemented policy is 3.92 times faster than by the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Jongho Yoon 0001, Seokhyeong Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Late Breaking Results: Reinforcement Learning-based Power Management Policy for Mobile Device SystemsabstractThis paper presents a power management policy that exploits reinforcement learning to increase power efficiency of mobile device systems. Our Q-learning-based policy predicts a system’s characteristics and learns power management controls to adapt to the system’s variations. Therefore, we can flexibly manage the system power regardless of the application scenario and can achieve lower energy per QoS compared to previous dynamic voltage/frequency scaling governors. To minimize the process overhead, we implemented our power management policy as hardware; the hardware-implemented policy reduced the average latency up to 40× compared to the software-implemented policy. Eunji Kwon, Sodam Han, Yoonho Park, Young Hwan Kim, Seokhyeong Kang |
DAC | 1 |
| 2020 | Analysis and Solution of CNN Accuracy Reduction over Channel Loop TilingabstractOwing to the growth of the size of convolutional neural networks (CNNs), quantization and loop tiling (also called loop breaking) are mandatory to implement CNN on an embedded system. However, channel loop tiling of quantized CNNs induces unexpected errors. We explain why channel loop tiling of quantized CNNs induces the unexpected errors, and how the errors affect the accuracy of state-of-the-art CNNs. We also propose a method to recover accuracy under channel tiling by compressing and decompressing the most-significant bits of partial sums. Using the proposed method, we can recover accuracy by 12.3% with only 1% circuit area overhead and an additional 2% of power consumption. Yesung Kang, Yoonho Park, Eunji Kwon, Taeho Lim, Sangyun Oh, Mingyu Woo, Seokhyeong Kang |
DATE | 4 |
| 2020 | GRLC: grid-based run-length compression for energy-efficient CNN acceleratorabstractConvolutional neural networks (CNNs) require a huge amount of off-chip DRAM access, which accounts for most of its energy consumption. Compression of feature maps can reduce the energy consumption of DRAM access. However, previous compression methods show poor compression ratio if the feature maps are either extremely sparse or dense. To improve the compression ratio efficiently, we have exploited the spatial correlation and the distribution of non-zero activations in output feature maps. In this work, we propose a grid-based run-length compression (GRLC) and have implemented a hardware for the GRLC. Compared with a previous compression method [1], GRLC reduces 11% of the DRAM access and 5% of the energy consumption on average in VGG-16, ExtractionNet and ResNet-18. Yoonho Park, Yesung Kang, Eunji Kwon, Seokhyeong Kang |
ISLPED | 4 |