Dahoon Park

dblp:305/4358 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MX-SAFE: Versatile Inference-and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
abstract
As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8_E2M5) and a subnormal FP mode (FP5_E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8_E2M5 and MXFP8_E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.
Dahoon Park, Jahyun Koo 0002, Sangwoo Hwang, Jaeha Kung 0001
DATE1
2026 A Survey on Binary and Ternary Neural Networks and Their Realization in Compute-in-Memory for Edge Intelligence
abstract
Deep learning has achieved remarkable success across a wide range of applications, such as language modeling, computer vision, recommendation systems, and robotics. However, the growing size of models and their increasing computational demands pose significant challenges, particularly for resource-constrained devices. A promising approach to address these challenges is extreme quantization, exemplified by binary and ternary neural networks. These techniques significantly reduce model size by quantizing weights and activations to 1 bit or 1.58 bits, while simplifying computation, making them well-suited for efficient deployment in resource-limited environments. This paper presents a comprehensive review of extreme quantization techniques, organized into three key areas: (1) a comparative analysis of quantizing only the weights (e.g., binary weight networks, ternary weight networks) versus quantizing both weights and activations (e.g., binary neural networks, ternary neural networks), along with a discussion of the progress and trade-offs of their approaches; (2) an examination of how extreme quantization, initially applied to convolutional neural networks, has been extended to Transformer architectures; and (3) an overview of compute-in-memory architectures optimized for binarization and ternarization, including designs based on advanced bit-cell technologies.
Dahoon Park, Hyungdong Park, Inguk Yeo, Suhak Lee, Hyunseob Shin, Sung-Il Pae, Deliang Fan, Jaeha Kung 0001, Kon-Woo Kwon
IEEE Internet Things J.1
2026 A Hybrid Digital-Analog Compute-in-Memory Using Content-Addressable Memory With Flexible Multi-Bit Slicing
abstract
Compute-in-memory (CIM) reduces data movement and enhances compute parallelism, making it suitable for AI applications. However, analog CIMs, yet energy-efficient, are vulnerable to PVT variations, while digital CIMs offer robustness but limited efficiency due to their bit-wise computation overhead. To address these challenges, we propose a hybrid CIM architecture that integrates content-addressable memory (CAM) and cluster-based CIM, named CAM-CIM, fabricated in 65nm CMOS technology. The proposed CAM-CIM flexibly slices multi-bit weights, assigning MSBs to CAM and LSBs to CIM, enabling dynamic accuracy-efficiency trade-offs across various bit precisions. A two-stage 8:3 compressor-based adder tree improves CAM efficiency and a reference voltage search algorithm ensures accurate CIM computation with low-bit ADCs. Our CAM-CIM supports 1-8b inputs/weights with reconfigurable compute modes, leveraging ternary-CAM based selective columns and cluster-wise CIM processing to produce multiple trade-off points even in the same bit precision. A prototype chip with a RISC-V controller and custom instructions is demonstrated that shows energy efficiencies of 32.4TOPS/W (8b/8b) and 76.0-354.9TOPS/W (4b/4b) with 0.66% accuracy loss, on average, across a wide range of DNN benchmarks including CNNs and vision transformers on CIFAR and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 RISC-V Driven Orchestration of Vector Processing Units and eFlash Compute-in-Memory Arrays for Fast and Accurate Keyword Spotting
abstract
In this paper, we propose a computationally efficient keyword spotting (KWS) model, named hybrid reparameterized FSMN (HRepFSMN), by carefully examining the impact of binarization on the accuracy. In particular, we found that binarizing depthwise convolution (DW-Conv) within the previous binarized KWS model, i.e., BiFSMNv2, does not lead to a significant reduction in FLOPs. Therefore, we allow floating-point (FP) operations on less computation-intensive DW-Conv layers while the remaining layers are computed in a binary fashion (hybrid data type). In addition, we remove skip connections, which require data fetching in full precision, by applying a reparameterization technique. More importantly, to efficiently compute the proposed HRepFSMN, we present a RISC-V controlled hardware accelerator that consists of reconfigurable vector processing units for FP operations and eFlash compute-in-memory arrays for binary operations. We extend RISC-V instructions so that the core can efficiently manage both computing fabrics. As a result, our HRepFSMN improves accuracy by 2.57%/4.98% with 24.02×/3.66× speed-up compared to BiFSMNv2/BiFSMNv2_small. By shrinking down our HRepFSMN, we achieve 0.95% higher accuracy with 20.87× speed-up compared to BiFSMNv2_small.
Gunil Kang, Dahoon Park, Sangwoo Jung 0001, Jung Gyu Min, Youngjoo Lee 0002, Jaeha Kung 0001
ASP-DAC2
2025 CAM-CIM: A Hybrid Compute-in-Memory Using Content-Addressable Memory with Subword Split Mapping for Reduced ADC Resolution
abstract
Recently, compute-in-memory (CIM) has become a promising architecture for data-intensive applications such as deep learning. However, analog or digital CIM (ACIM or DCIM) faces some design challenges. ACIMs inherently have non-idealities, which lead to significant accuracy degradation. In addition, a substantial amount of power is consumed by analog-to-digital converters (ADC). On the other hand, DCIMs show an exponential increase in power consumption and computing cycles as the operand bit-width increases, particularly due to an accumulation stage. In this paper, to overcome these challenges, we propose a hybrid DCIM-ACIM architecture that consists of a content addressable memory (CAM) as DCIM and a cluster-based multi-cycle ACIM, called CAM-CIM. As a weight mapping strategy, we present a subword split mapping that assigns some MSBs to DCIM for improved accuracy and the remaining LSBs to ACIM for reduced ADC resolution. The accuracy of using the proposed CAM-CIM array is evaluated on various deep learning benchmarks from CNNs to Swin-Tiny. A 65nm CAM-CIM macro with either 3-bit or 4-bit ADCs shows 10.3 × and 5.4 × improvement in energy efficiency, on average, compared to CAM- and CIM-only architectures, respectively. Compared to recent CIM architectures, CAM-CIM demonstrates 1.4 × higher energy efficiency.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
ISLPED5
2025 RIMIX: RISC-V Core with MIXed-Precision SIMD Instruction Extensions Supported by Oracle-Assisted Sub-Network Search for Efficient TinyML
abstract
As the size of the deep learning model increases, mixed-precision quantization has become an efficient compression technique. However, the lack of MCU support for mixed-precision computation limits its performance in running tinyML tasks. To address this issue, we propose a RISC-V core, named RIMIX, designed to support various bit combinations with minimal hardware overhead. The RIMIX features an optimized bit packing mechanism, an extended ISA tailored for mixed-precision arithmetic, and a neural unit capable of multi-precision computations, achieving up to 28.6 × speedup over an Ibex core. To maximize the quality of tinyML processing with RIMIX, we also present an oracle-based neural architecture search to discover optimized models under a target constraint. To speed up the search process, we propose a novel two-stage approach that decouples model topology search and mixed-precision training. First, we search for an oracle network using training-free NAS, i.e., a high-bit optimized network that serves as a basis for mixed-precision training. Once the oracle architecture is identified, we distill the model with weight sharing so that it performs well on any bit combinations. We further propose a sub-network selection strategy from the oracle network, considering actual RIMIX instruction cycles to better meet the target constraints. Our sub-network selection method outperforms conventional BOPs-based search methods. Finally, the proposed SW/HW co-design method enables 2.0× faster execution in running TinyML tasks while keeping the accuracy drop by less than 2% compared to the existing state-of-the-art method on the Artix-7 FPGA board.
Dahoon Park, Yeeun Hong
ISLPED2
2024 OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
abstract
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while lacking research on quantizing activations. In this paper, we present a hardware-software co-design method that results in an energy-efficient LLM accelerator, named OPAL, for generation tasks. First of all, a novel activation quantization method that leverages the microscaling data format while preserving several outliers per subtensor block (e.g., four out of 128 elements) is proposed. Second, on top of preserving outliers, mixed precision is utilized that sets 5-bit for inputs to sensitive layers in the decoder block of an LLM, while keeping inputs to less sensitive layers to 3-bit. Finally, we present the OPAL hardware architecture that consists of FP units for handling outliers and vectorized INT multipliers for dominant non-outlier related operations. In addition, OPAL uses log2-based approximation on softmax operations that only requires shift and subtraction to maximize power efficiency. As a result, we are able to improve the energy efficiency by 1.6~2.2×, and reduce the area by 2.4~3.1× with negligible accuracy loss, i.e., <1 perplexity increase.
Jahyun Koo 0002, Dahoon Park, Sangwoo Jung 0001, Jaeha Kung 0001
DAC2
2024 SpikedAttention: Training-Free and Fully Spike-Driven Transformer-to-SNN Conversion with Winner-Oriented Spike Shift for Softmax Operation
abstract
Event-driven spiking neural networks(SNNs) are promising neural networks that reduce the energy consumption of continuously growing AI models. Recently, keeping pace with the development of transformers, transformer-based SNNs were presented. Due to the incompatibility of self-attention with spikes, however, existing transformer-based SNNs limit themselves by either restructuring self-attention architecture or conforming to non-spike computations. In this work, we propose a novel transformer-to-SNN conversion method that outputs an end-to-end spike-based transformer, named SpikedAttention. Our method directly converts the well-trained transformer without modifying its attention architecture. For the vision task, the proposed method converts Swin Transformer into an SNN without post-training or conversion-aware training, achieving state-of-the-art SNN accuracy on ImageNet dataset, i.e., 80.0\% with 28.7M parameters. Considering weight accumulation, neuron potential update, and on-chip data movement, SpikedAttention reduces energy consumption by 42\% compared to the baseline ANN, i.e., Swin-T. Furthermore, for the first time, we demonstrate that SpikedAttention successfully converts a BERT model to an SNN with only 0.3\% accuracy loss on average consuming 58\% less energy on GLUE benchmark. Our code is available at Github ( https://github.com/sangwoohwang/SpikedAttention ).
Sangwoo Hwang, Dahoon Park
NeurIPS3
2024 A Dual-Precision and Low-Power CNN Inference Engine Using a Heterogeneous Processing-in-Memory Architecture
abstract
In this article, we present an energy-scalable CNN model that can adapt to different hardware resource constraints. Specifically, we propose a dual-precision network, named DualNet, that leverages two independent bit-precision paths (INT4 and ternary-binary). DualNet achieves both high accuracy and low complexity by balancing the ratio between two paths. We also present an evolutionary algorithm that allows the automatic search of the optimal ratios. In addition to the novel CNN architecture design, we develop a heterogeneous processing-in-memory (PIM) hardware that integrates SRAM-and eDRAM-based PIMs to efficiently compute two precision paths in parallel. To verify the energy efficiency of DualNet computed on the heterogeneous PIM, we prototyped a test chip in 28nm CMOS technology. To maximize the hardware efficiency, we utilize an improved data mapping scheme achieving the most effective deployment of DualNets on multiple PIM arrays. With the proposed SW-HW co-optimization, we can obtain the most energy-efficient DualNet model operating on the actual PIM hardware. Compared to the other quantized networks with a single bit-precision, DualNet reduces the energy consumption, memory footprint, and latency by 29.0%, 49.5%, 47.3% on average, respectively, for CIFAR-10/100 and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Youngjoo Lee 0002, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 LightNorm: Area and Energy-Efficient Batch Normalization Hardware for On-Device DNN Training
abstract
When training early-stage deep neural networks (DNNs), generating intermediate features via convolution or linear layers occupied most of the execution time. Accordingly, extensive research has been done to reduce the computational burden of the convolution or linear layers. In recent mobile-friendly DNNs, however, the relative number of operations involved in processing these layers has significantly reduced. As a result, the proportion of the execution time of other layers, such as batch normalization layers, has increased. Thus, in this work, we conduct a detailed analysis of the batch normalization layer to efficiently reduce the runtime overhead in the batch normalization process. Backed up by the thorough analysis, we present an extremely efficient batch normalization, named LightNorm, and its associated hardware module. In more detail, we fuse three approximation techniques that are i) low bit-precision, ii) range batch normalization, and iii) block floating point. All these approximate techniques are carefully utilized not only to maintain the statistics of intermediate feature maps, but also to minimize the off-chip memory accesses. By using the proposed LightNorm hardware, we can achieve significant area and energy savings during the DNN training without hurting the training accuracy. This makes the proposed hardware a great candidate for the on-device training.
Seock-Hwan Noh, Junsang Park, Dahoon Park, Jahyun Koo 0002, Jeik Choi, Jaeha Kung 0001
ICCD3
2021 ZeBRA: Precisely Destroying Neural Networks with Zero-Data Based Repeated Bit Flip Attack
Dahoon Park, Kon-Woo Kwon, Sunghoon Im 0001
BMVC1