Jahyun Koo 0002

dblp:315/9569 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-5652-5306ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MX-SAFE: Versatile Inference-and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
abstract
As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8_E2M5) and a subnormal FP mode (FP5_E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8_E2M5 and MXFP8_E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.
Dahoon Park, Jahyun Koo 0002, Sangwoo Hwang, Jaeha Kung 0001
DATE2
2026 GustavSNN: Unleashing the Power of Gustavson's Algorithm on SNN Acceleration with Column-Parallel Tick-Batch Dataflow
abstract
Spiking neural networks (SNNs) require sequential computation over long timesteps, introducing substantial memory and energy overheads due to frequent updates of neuron membrane potentials. Previous SNN accelerators address this by employing tick-batch techniques, which process all timesteps within a layer before moving on to the next. However, existing approaches rely on neuron-centric scheduling, limiting their ability to exploit temporal sparsity. In this work, we propose a novel scheduling approach along with its hardware architecture for Gustavson product (GP)-based SNN acceleration. We introduce a column-parallel tick-batch (CPTB) dataflow that partitions the spike matrix into multiple submatrices and processes each submatrix of a timestep in parallel while maintaining tickbatch semantics. To support this, we present the first GP-based SNN accelerator, named GustavSNN, which avoids accessing the global membrane potential memory by updating neuron states directly in local registers. In addition, we propose a non-zero row vector (NRV) spike format that enables fine-grained skipping of inactive spike rows. As a result, our proposed architecture achieves up to 11.8× higher energy efficiency (GOPS/W) than naïve GP-based accelerator and 1.43× higher energy efficiency compared to state-of-the-art SNN accelerators.
Sangwoo Hwang, Jahyun Koo 0002, Jaeha Kung 0001
HPCA3
2024 OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
abstract
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while lacking research on quantizing activations. In this paper, we present a hardware-software co-design method that results in an energy-efficient LLM accelerator, named OPAL, for generation tasks. First of all, a novel activation quantization method that leverages the microscaling data format while preserving several outliers per subtensor block (e.g., four out of 128 elements) is proposed. Second, on top of preserving outliers, mixed precision is utilized that sets 5-bit for inputs to sensitive layers in the decoder block of an LLM, while keeping inputs to less sensitive layers to 3-bit. Finally, we present the OPAL hardware architecture that consists of FP units for handling outliers and vectorized INT multipliers for dominant non-outlier related operations. In addition, OPAL uses log2-based approximation on softmax operations that only requires shift and subtraction to maximize power efficiency. As a result, we are able to improve the energy efficiency by 1.6~2.2×, and reduce the area by 2.4~3.1× with negligible accuracy loss, i.e., <1 perplexity increase.
Jahyun Koo 0002, Dahoon Park, Sangwoo Jung 0001, Jaeha Kung 0001
DAC1
2023 DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions
abstract
Over the past decade, it has been found that deep neural networks (DNNs) perform better on visual perception and language understanding tasks as their size increases. However, this comes at the cost of high energy consumption and large memory requirement to train such large models. As the training DNNs necessitates a wide dynamic range in representing tensors, floating point formats are normally used. In this work, we utilize a block floating point (BFP) format that significantly reduces the size of tensors and the power consumption of arithmetic units. Unfortunately, prior work on BFP-based DNN training empirically selects the block size and the precision that maintain the training accuracy. To make the BFP-based training more feasible, we propose dynamic block size and precision scaling (DBPS) for highly efficient DNN training. We also present a hardware accelerator, called DBPS core, which supports the DBPS control by configuring arithmetic units with custom instructions extended in a RISC-V processor. As a result, the training time and energy consumption reduce by 67.1% and 72.0%, respectively, without hurting the training accuracy.
Jeik Choi, Seock-Hwan Noh, Jahyun Koo 0002, Jaeha Kung 0001
DAC4
2023 FlexBlock: A Flexible DNN Training Accelerator With Multi-Mode Block Floating Point Support
abstract
When training deep neural networks (DNNs), expensive floating point arithmetic units are used in GPUs or custom neural processing units (NPUs). To reduce the burden of floating point arithmetic, community has started exploring the use of more efficient data representations, e.g., block floating point (BFP). The BFP format allows a group of values to share an exponent, which effectively reduces the memory footprint and enables cheaper fixed point arithmetic for multiply-accumulate (MAC) operations. However, existing BFP-based DNN accelerators are targeted for a specific precision, making them less versatile. In this paper, we present FlexBlock, a DNN training accelerator with three BFP modes, possibly different among activation, weight, and gradient tensors. By configuring FlexBlock to a lower BFP precision, the number of MACs handled by the core increases by up to 4× in 8-bit mode or 16× in 4-bit mode compared to 16-bit mode. To reach this theoretical upper bound, FlexBlock maximizes the core utilization at various precision levels or layer types, and allows dynamic precision control to keep throughput at its peak without sacrificing training accuracy. We evaluate the effectiveness of FlexBlock using representative DNNs on CIFAR, ImageNet and WMT14 datasets. As a result, training in FlexBlock significantly improves training speed by 1.5$\sim 5.3\times$and energy efficiency by 2.4$\sim 7.0\times$compared to other training accelerators.
Seock-Hwan Noh, Jahyun Koo 0002, Jongse Park, Jaeha Kung 0001
IEEE Trans. Computers2
2022 LightNorm: Area and Energy-Efficient Batch Normalization Hardware for On-Device DNN Training
abstract
When training early-stage deep neural networks (DNNs), generating intermediate features via convolution or linear layers occupied most of the execution time. Accordingly, extensive research has been done to reduce the computational burden of the convolution or linear layers. In recent mobile-friendly DNNs, however, the relative number of operations involved in processing these layers has significantly reduced. As a result, the proportion of the execution time of other layers, such as batch normalization layers, has increased. Thus, in this work, we conduct a detailed analysis of the batch normalization layer to efficiently reduce the runtime overhead in the batch normalization process. Backed up by the thorough analysis, we present an extremely efficient batch normalization, named LightNorm, and its associated hardware module. In more detail, we fuse three approximation techniques that are i) low bit-precision, ii) range batch normalization, and iii) block floating point. All these approximate techniques are carefully utilized not only to maintain the statistics of intermediate feature maps, but also to minimize the off-chip memory accesses. By using the proposed LightNorm hardware, we can achieve significant area and energy savings during the DNN training without hurting the training accuracy. This makes the proposed hardware a great candidate for the on-device training.
Seock-Hwan Noh, Junsang Park, Dahoon Park, Jahyun Koo 0002, Jeik Choi, Jaeha Kung 0001
ICCD4