Jaeha Kung 0001

dblp:12/10064-1 · DBLP profile ↗
← Back
36ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-2027-8531ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 6 first-author · 20 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MX-SAFE: Versatile Inference-and Training-Proof Microscaling Format with On-the-Fly Exponent and Mantissa Bit Allocation
abstract
As the demand for deep learning grows, cost reduction through quantization has become essential for both training and inference. In 2022, the Open Compute Project (OCP) consortium standardized narrow precision formats for deep learning, called the microscaling (MX) format. The MX format is a hardware-friendly dynamic quantization scheme that effectively reduces the data size by sharing an 8-bit exponent across multiple operands. The MX format can be categorized into two types with their own strengths: (i) MXINT which focuses on a high precision consisting only of mantissa bits and (ii) MXFP which focuses on a wider dynamic range by allowing local exponent bits. In this work, we present a versatile MXFP format, called MX-SAFE (MXSF in short), that adaptively uses two modes, i.e., a wider mantissa mode (FP8_E2M5) and a subnormal FP mode (FP5_E3M2), to support both training and direct-cast inference. Furthermore, we propose a tile-based block design to increase hardware efficiency by reducing the burden of re-quantization process during the training with the MXSF format. Owing to the use of the proposed MXSF format, 0.05%/11.1% and 3.55%/3.57% improvements in accuracy, on average, for inference/full-training compared to MXFP8_E2M5 and MXFP8_E4M3 are observed, respectively. Moreover, we present a training-inference accelerator that supports the MXSF format and it achieves similar accuracy to the BF16 baseline while using 24.9% less total energy consumption.
Dahoon Park, Jahyun Koo 0002, Sangwoo Hwang, Jaeha Kung 0001
DATE4
2026 GustavSNN: Unleashing the Power of Gustavson's Algorithm on SNN Acceleration with Column-Parallel Tick-Batch Dataflow
abstract
Spiking neural networks (SNNs) require sequential computation over long timesteps, introducing substantial memory and energy overheads due to frequent updates of neuron membrane potentials. Previous SNN accelerators address this by employing tick-batch techniques, which process all timesteps within a layer before moving on to the next. However, existing approaches rely on neuron-centric scheduling, limiting their ability to exploit temporal sparsity. In this work, we propose a novel scheduling approach along with its hardware architecture for Gustavson product (GP)-based SNN acceleration. We introduce a column-parallel tick-batch (CPTB) dataflow that partitions the spike matrix into multiple submatrices and processes each submatrix of a timestep in parallel while maintaining tickbatch semantics. To support this, we present the first GP-based SNN accelerator, named GustavSNN, which avoids accessing the global membrane potential memory by updating neuron states directly in local registers. In addition, we propose a non-zero row vector (NRV) spike format that enables fine-grained skipping of inactive spike rows. As a result, our proposed architecture achieves up to 11.8× higher energy efficiency (GOPS/W) than naïve GP-based accelerator and 1.43× higher energy efficiency compared to state-of-the-art SNN accelerators.
Sangwoo Hwang, Jahyun Koo 0002, Jaeha Kung 0001
HPCA4
2026 A Survey on Binary and Ternary Neural Networks and Their Realization in Compute-in-Memory for Edge Intelligence
abstract
Deep learning has achieved remarkable success across a wide range of applications, such as language modeling, computer vision, recommendation systems, and robotics. However, the growing size of models and their increasing computational demands pose significant challenges, particularly for resource-constrained devices. A promising approach to address these challenges is extreme quantization, exemplified by binary and ternary neural networks. These techniques significantly reduce model size by quantizing weights and activations to 1 bit or 1.58 bits, while simplifying computation, making them well-suited for efficient deployment in resource-limited environments. This paper presents a comprehensive review of extreme quantization techniques, organized into three key areas: (1) a comparative analysis of quantizing only the weights (e.g., binary weight networks, ternary weight networks) versus quantizing both weights and activations (e.g., binary neural networks, ternary neural networks), along with a discussion of the progress and trade-offs of their approaches; (2) an examination of how extreme quantization, initially applied to convolutional neural networks, has been extended to Transformer architectures; and (3) an overview of compute-in-memory architectures optimized for binarization and ternarization, including designs based on advanced bit-cell technologies.
Dahoon Park, Hyungdong Park, Inguk Yeo, Suhak Lee, Hyunseob Shin, Sung-Il Pae, Deliang Fan, Jaeha Kung 0001, Kon-Woo Kwon
IEEE Internet Things J.10
2026 A Hybrid Digital-Analog Compute-in-Memory Using Content-Addressable Memory With Flexible Multi-Bit Slicing
abstract
Compute-in-memory (CIM) reduces data movement and enhances compute parallelism, making it suitable for AI applications. However, analog CIMs, yet energy-efficient, are vulnerable to PVT variations, while digital CIMs offer robustness but limited efficiency due to their bit-wise computation overhead. To address these challenges, we propose a hybrid CIM architecture that integrates content-addressable memory (CAM) and cluster-based CIM, named CAM-CIM, fabricated in 65nm CMOS technology. The proposed CAM-CIM flexibly slices multi-bit weights, assigning MSBs to CAM and LSBs to CIM, enabling dynamic accuracy-efficiency trade-offs across various bit precisions. A two-stage 8:3 compressor-based adder tree improves CAM efficiency and a reference voltage search algorithm ensures accurate CIM computation with low-bit ADCs. Our CAM-CIM supports 1-8b inputs/weights with reconfigurable compute modes, leveraging ternary-CAM based selective columns and cluster-wise CIM processing to produce multiple trade-off points even in the same bit precision. A prototype chip with a RISC-V controller and custom instructions is demonstrated that shows energy efficiencies of 32.4TOPS/W (8b/8b) and 76.0-354.9TOPS/W (4b/4b) with 0.66% accuracy loss, on average, across a wide range of DNN benchmarks including CNNs and vision transformers on CIFAR and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 Fused-STA: Automated Design Space Exploration of a Fused Systolic Tensor Array for Universal Deep Learning Acceleration
abstract
Systolic array architectures have become the dominant solution for accelerating deep neural network (DNN) computations, yet designing optimal configurations remains challenging due to the vast design space spanning array dimensions, dataflow strategies, SRAM allocations, and tiling sizes. Existing tool-aided optimization frameworks constrain this design space by treating SRAM sizes or tiling sizes as fixed input parameters, while focusing only on conventional systolic arrays, limiting their ability to discover a truly optimal design. This article presents a comprehensive framework for systolic tensor array (STA) design space exploration that jointly optimizes array configurations, SRAM sizes, tiling sizes, and dataflow types. We introduce TensorSim, a cycle-accurate simulator with an ML-based synthesis prediction model, and TensorOptimizer, an automated multiobjective optimization framework. Evaluation with TensorOptimizer across 12 representative DNNs—from edge-level convolutional neural networks (CNNs) to billion-parameter large language models (LLMs)—reveals that output stationary (OS) dataflows outperform weight stationary (WS) dataflows by$1.53\times $on edge workloads, while WS achieves$2.13\times $higher efficiency over OS on server-scale computations. Based on these insights, we propose Fused-STA, a reconfigurable architecture that dynamically switches between OS and WS modes within a unified hardware substrate. Fused-STA achieves over 91% of the oracle single-dataflow performance, specifically designed for a single network, providing$2.34\times $and$1.92\times $improvements over a tensor processing unit (TPU)-like baseline on edge-level and server-level benchmarks, respectively.
Jooyeon Lee, Sangwoo Jung 0001, Jaeha Kung 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2025 RISC-V Driven Orchestration of Vector Processing Units and eFlash Compute-in-Memory Arrays for Fast and Accurate Keyword Spotting
abstract
In this paper, we propose a computationally efficient keyword spotting (KWS) model, named hybrid reparameterized FSMN (HRepFSMN), by carefully examining the impact of binarization on the accuracy. In particular, we found that binarizing depthwise convolution (DW-Conv) within the previous binarized KWS model, i.e., BiFSMNv2, does not lead to a significant reduction in FLOPs. Therefore, we allow floating-point (FP) operations on less computation-intensive DW-Conv layers while the remaining layers are computed in a binary fashion (hybrid data type). In addition, we remove skip connections, which require data fetching in full precision, by applying a reparameterization technique. More importantly, to efficiently compute the proposed HRepFSMN, we present a RISC-V controlled hardware accelerator that consists of reconfigurable vector processing units for FP operations and eFlash compute-in-memory arrays for binary operations. We extend RISC-V instructions so that the core can efficiently manage both computing fabrics. As a result, our HRepFSMN improves accuracy by 2.57%/4.98% with 24.02×/3.66× speed-up compared to BiFSMNv2/BiFSMNv2_small. By shrinking down our HRepFSMN, we achieve 0.95% higher accuracy with 20.87× speed-up compared to BiFSMNv2_small.
Gunil Kang, Dahoon Park, Sangwoo Jung 0001, Jung Gyu Min, Youngjoo Lee 0002, Jaeha Kung 0001
ASP-DAC8
2025 Dissecting and Re-Architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMS
abstract
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional hardware is challenging due to limited DRAM capacity and high GPU costs. Thus, in this work, we propose offloading the single-batch token generation to a 3D NAND flash processing-in-memory (PIM) device, leveraging its high storage density to overcome the DRAM capacity wall. We explore 3D NAND flash configurations and present a re-architected PIM array with an H-tree network for optimal latency and cell density. Along with the well-chosen PIM array size, we develop operation tiling and mapping methods for LLM layers, achieving a$2.4 \times$speedup over four RTX4090 with vLLM and comparable performance to four A100 with only 4.9% latency overhead. Our detailed area analysis reveals that the proposed 3D NAND flash PIM architecture can be integrated within a$4.98 ~\text{mm}^{2}$die area under the memory array, without extra area overhead.
Yongjoo Jang, Sangwoo Hwang, Sangwoo Jung 0001, Wonbo Shim, Jaeha Kung 0001
ICCD7
2025 CAM-CIM: A Hybrid Compute-in-Memory Using Content-Addressable Memory with Subword Split Mapping for Reduced ADC Resolution
abstract
Recently, compute-in-memory (CIM) has become a promising architecture for data-intensive applications such as deep learning. However, analog or digital CIM (ACIM or DCIM) faces some design challenges. ACIMs inherently have non-idealities, which lead to significant accuracy degradation. In addition, a substantial amount of power is consumed by analog-to-digital converters (ADC). On the other hand, DCIMs show an exponential increase in power consumption and computing cycles as the operand bit-width increases, particularly due to an accumulation stage. In this paper, to overcome these challenges, we propose a hybrid DCIM-ACIM architecture that consists of a content addressable memory (CAM) as DCIM and a cluster-based multi-cycle ACIM, called CAM-CIM. As a weight mapping strategy, we present a subword split mapping that assigns some MSBs to DCIM for improved accuracy and the remaining LSBs to ACIM for reduced ADC resolution. The accuracy of using the proposed CAM-CIM array is evaluated on various deep learning benchmarks from CNNs to Swin-Tiny. A 65nm CAM-CIM macro with either 3-bit or 4-bit ADCs shows 10.3 × and 5.4 × improvement in energy efficiency, on average, compared to CAM- and CIM-only architectures, respectively. Compared to recent CIM architectures, CAM-CIM demonstrates 1.4 × higher energy efficiency.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
ISLPED8
2025 Jack Unit: An Area- and Energy-Efficient Multiply-Accumulate (MAC) Unit Supporting Diverse Data Formats
abstract
In this work, we introduce an area- and energy-efficient multiply-accumulate (MAC) unit, named Jack Unit, that is a jack-of-all-trades, supporting various data formats such as integer (INT), floating point (FP), and microscaling data format (MX). It provides bit-level flexibility and enhances hardware efficiency by i) replacing the carry-save multiplier (CSM) in the FP multiplier with a precision-scalable CSM, ii) performing the adjustment of significands based on the exponent differences within the CSM, and iii) utilizing 2D sub-word parallelism. To assess effectiveness, we implemented the layout of the Jack unit and three baseline MAC units. Additionally, we designed an AI accelerator equipped with our Jack units to compare with a state-of-the-art AI accelerator supporting various data formats. The proposed MAC unit achieves an area reduction of 14.53∼50.25% and a power reduction of 4.76 ∼ 45.65% compared to the baseline MAC units. On five AI benchmarks, the accelerator de-signed with our Jack units improves energy efficiency by 1.32 ∼ 5.41× over the baseline across various data formats.
Seock-Hwan Noh, Sungju Kim, Daehoon Kim 0001, Jaeha Kung 0001, Yeseong Kim
ISLPED5
2025 All-Rounder: A Flexible AI Accelerator With Diverse Data Format Support and Morphable Structure for Multi-DNN Processing
abstract
Recognizing the explosive increase in the use of artificial intelligence (AI)-based applications, several industrial companies developed custom application-specific integrated circuits (ASICs) (e.g., Google TPU, IBM RaPiD, and Intel NNP-I/NNP-T) and constructed a hyperscale cloud infrastructure with them. These ASICs perform operations of the inference or training process of AI models which are requested by users. Since the AI models have different data formats and types of operations, the ASICs need to support diverse data formats and various operation shapes. However, the previous ASIC solutions do not or less fulfill these requirements. To overcome these limitations, we first present an area-efficient multiplier, named all-in-one multiplier, which supports multiple bit-widths for both integer (INT) and floating-point (FP) data types. Then, we build a multiply-and-accumulation (MAC) array equipped with these multipliers with multiformat support. In addition, the MAC array can be partitioned into multiple blocks that can be flexibly fused to support various deep neural network (DNN) operation types. We evaluate the practical effectiveness of the proposed MAC array by making an accelerator out of it, named All-rounder. According to our evaluation, the proposed all-in-one multiplier occupies$1.49\times $smaller area compared to the baselines with dedicated multipliers for each data format. Then, we compare the performance and energy efficiency of the proposed All-rounder with three different accelerators showing consistent speedup and higher efficiency across various AI benchmarks from vision to large language model (LLM)-based language tasks.
Seock-Hwan Noh, Seungpyo Lee, Banseok Shin, Sehun Park, Yongjoo Jang, Jaeha Kung 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2024 NDPipe: Exploiting Near-data Processing for Scalable Inference and Continuous Training in Photo Storage
abstract
This paper proposes a novel photo storage system called NDPipe, which accelerates the performance of training and inference for image data by leveraging near-data processing in photo storage servers. NDPipe distributes storage servers with inexpensive commodity GPUs in a data center and uses their collective intelligence to perform inference and training near image data. By efficiently partitioning deep neural network (DNN) models and exploiting the data parallelism of many storage servers, NDPipe can achieve high training throughput with low synchronization costs. NDPipe optimizes the near-data processing engine to maximally utilize system components in each storage server. Our results show that, given the same energy budget, NDPipe exhibits 1.39× higher inference throughput and 2.64× faster training speed than typical photo storage systems.
Jungwoo Kim 0004, Seonggyun Oh, Jaeha Kung 0001, Yeseong Kim, Sungjin Lee 0001
ASPLOS (3)3
2024 OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
abstract
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while lacking research on quantizing activations. In this paper, we present a hardware-software co-design method that results in an energy-efficient LLM accelerator, named OPAL, for generation tasks. First of all, a novel activation quantization method that leverages the microscaling data format while preserving several outliers per subtensor block (e.g., four out of 128 elements) is proposed. Second, on top of preserving outliers, mixed precision is utilized that sets 5-bit for inputs to sensitive layers in the decoder block of an LLM, while keeping inputs to less sensitive layers to 3-bit. Finally, we present the OPAL hardware architecture that consists of FP units for handling outliers and vectorized INT multipliers for dominant non-outlier related operations. In addition, OPAL uses log2-based approximation on softmax operations that only requires shift and subtraction to maximize power efficiency. As a result, we are able to improve the energy efficiency by 1.6~2.2×, and reduce the area by 2.4~3.1× with negligible accuracy loss, i.e., <1 perplexity increase.
Jahyun Koo 0002, Dahoon Park, Sangwoo Jung 0001, Jaeha Kung 0001
DAC4
2024 A Ready-to-Use RTL Generator for Systolic Tensor Arrays and Analysis Using Open-Source EDA Tools
abstract
From simple image classifiers to complex and large language models, generalized matrix multiplication (GEMM) is the fundamental and the most time-consuming operation among all mathematical operations involved in them. To accelerate the computation of matrix multiplication in deep learning, many off-the-shelf neural processors utilize systolic arrays as dedicated hardware for the GEMM operations. Recently, more generalized form of the systolic array, i.e., a systolic tensor array (STA) which includes vectorized MAC units within a single processing unit, has been proposed. However, the optimal selection of STA configuration on a given deep learning model is difficult due to large configuration search space. To help select the optimal STA configuration in many deep learning models, in this work, we present a ready-to-use and open-source RTL generator for various STA configurations. The power consumption and post-layout area of several STAs are analyzed by using open-source EDA tools.
Jooyeon Lee, Jaeha Kung 0001
ISCAS3
2024 A Dual-Precision and Low-Power CNN Inference Engine Using a Heterogeneous Processing-in-Memory Architecture
abstract
In this article, we present an energy-scalable CNN model that can adapt to different hardware resource constraints. Specifically, we propose a dual-precision network, named DualNet, that leverages two independent bit-precision paths (INT4 and ternary-binary). DualNet achieves both high accuracy and low complexity by balancing the ratio between two paths. We also present an evolutionary algorithm that allows the automatic search of the optimal ratios. In addition to the novel CNN architecture design, we develop a heterogeneous processing-in-memory (PIM) hardware that integrates SRAM-and eDRAM-based PIMs to efficiently compute two precision paths in parallel. To verify the energy efficiency of DualNet computed on the heterogeneous PIM, we prototyped a test chip in 28nm CMOS technology. To maximize the hardware efficiency, we utilize an improved data mapping scheme achieving the most effective deployment of DualNets on multiple PIM arrays. With the proposed SW-HW co-optimization, we can obtain the most energy-efficient DualNet model operating on the actual PIM hardware. Compared to the other quantized networks with a single bit-precision, DualNet reduces the energy consumption, memory footprint, and latency by 29.0%, 49.5%, 47.3% on average, respectively, for CIFAR-10/100 and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Youngjoo Lee 0002, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 DBPS: Dynamic Block Size and Precision Scaling for Efficient DNN Training Supported by RISC-V ISA Extensions
abstract
Over the past decade, it has been found that deep neural networks (DNNs) perform better on visual perception and language understanding tasks as their size increases. However, this comes at the cost of high energy consumption and large memory requirement to train such large models. As the training DNNs necessitates a wide dynamic range in representing tensors, floating point formats are normally used. In this work, we utilize a block floating point (BFP) format that significantly reduces the size of tensors and the power consumption of arithmetic units. Unfortunately, prior work on BFP-based DNN training empirically selects the block size and the precision that maintain the training accuracy. To make the BFP-based training more feasible, we propose dynamic block size and precision scaling (DBPS) for highly efficient DNN training. We also present a hardware accelerator, called DBPS core, which supports the DBPS control by configuring arithmetic units with custom instructions extended in a RISC-V processor. As a result, the training time and energy consumption reduce by 67.1% and 72.0%, respectively, without hurting the training accuracy.
Jeik Choi, Seock-Hwan Noh, Jahyun Koo 0002, Jaeha Kung 0001
DAC5
2023 FlexBlock: A Flexible DNN Training Accelerator With Multi-Mode Block Floating Point Support
abstract
When training deep neural networks (DNNs), expensive floating point arithmetic units are used in GPUs or custom neural processing units (NPUs). To reduce the burden of floating point arithmetic, community has started exploring the use of more efficient data representations, e.g., block floating point (BFP). The BFP format allows a group of values to share an exponent, which effectively reduces the memory footprint and enables cheaper fixed point arithmetic for multiply-accumulate (MAC) operations. However, existing BFP-based DNN accelerators are targeted for a specific precision, making them less versatile. In this paper, we present FlexBlock, a DNN training accelerator with three BFP modes, possibly different among activation, weight, and gradient tensors. By configuring FlexBlock to a lower BFP precision, the number of MACs handled by the core increases by up to 4× in 8-bit mode or 16× in 4-bit mode compared to 16-bit mode. To reach this theoretical upper bound, FlexBlock maximizes the core utilization at various precision levels or layer types, and allows dynamic precision control to keep throughput at its peak without sacrificing training accuracy. We evaluate the effectiveness of FlexBlock using representative DNNs on CIFAR, ImageNet and WMT14 datasets. As a result, training in FlexBlock significantly improves training speed by 1.5$\sim 5.3\times$and energy efficiency by 2.4$\sim 7.0\times$compared to other training accelerators.
Seock-Hwan Noh, Jahyun Koo 0002, Jongse Park, Jaeha Kung 0001
IEEE Trans. Computers5
2022 LightNorm: Area and Energy-Efficient Batch Normalization Hardware for On-Device DNN Training
abstract
When training early-stage deep neural networks (DNNs), generating intermediate features via convolution or linear layers occupied most of the execution time. Accordingly, extensive research has been done to reduce the computational burden of the convolution or linear layers. In recent mobile-friendly DNNs, however, the relative number of operations involved in processing these layers has significantly reduced. As a result, the proportion of the execution time of other layers, such as batch normalization layers, has increased. Thus, in this work, we conduct a detailed analysis of the batch normalization layer to efficiently reduce the runtime overhead in the batch normalization process. Backed up by the thorough analysis, we present an extremely efficient batch normalization, named LightNorm, and its associated hardware module. In more detail, we fuse three approximation techniques that are i) low bit-precision, ii) range batch normalization, and iii) block floating point. All these approximate techniques are carefully utilized not only to maintain the statistics of intermediate feature maps, but also to minimize the off-chip memory accesses. By using the proposed LightNorm hardware, we can achieve significant area and energy savings during the DNN training without hurting the training accuracy. This makes the proposed hardware a great candidate for the on-device training.
Seock-Hwan Noh, Junsang Park, Dahoon Park, Jahyun Koo 0002, Jeik Choi, Jaeha Kung 0001
ICCD6
2022 A 46-nF/10-MΩ Range 114-aF/0.37-Ω Resolution Parasitic- and Temperature-Insensitive Reconfigurable RC-to-Digital Converter in 0.18-μm CMOS
abstract
This paper presents a 46 nF/10$\text{M}\Omega $-range, digital-intensive, reconfigurable RC-to-digital converter (R2CDC) that can readout multiple C and R sensors in a time-interleaved fashion. Ratio-metric conversion using swing-boosted period-modulation (SB-PM) front-end by the R2CDC results in 114 aFrms/$0.37 \Omega _{\text {rms}}$resolutions and a worst-case temperature-drift of 64.2 ppm/°C over −40 to 125°C. Femto-farad capacitances can be sensed with a relative code-deviation less than 0.16 % even when parasitics vary 30 times the baseline. Implemented in a$0.18 ~\mu \text{m}$standard CMOS process, the R2CDC consumes 140$\mu \text{A}$from a 1 V supply, occupying an active area of$0.175 ~\mu \text{m} ^{\mathrm{ 2}}$.
Arup K. George, Wooyoon Shim, Jaeha Kung 0001, Ji-Hoon Kim 0003, Minkyu Je, Junghyup Lee
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Implication of Optimizing NPU Dataflows on Neural Architecture Search for Mobile Devices
abstract
Recent advances in deep learning have made it possible to implement artificial intelligence in mobile devices. Many studies have put a lot of effort into developing lightweight deep learning models optimized for mobile devices. To overcome the performance limitations of manually designed deep learning models, an automated search algorithm, called neural architecture search ( NAS ), has been proposed. However, studies on the effect of hardware architecture of the mobile device on the performance of NAS have been less explored. In this article, we show the importance of optimizing a hardware architecture, namely, NPU dataflow, when searching for a more accurate yet fast deep learning model. To do so, we first implement an optimization framework, named FlowOptimizer, for generating a best possible NPU dataflow for a given deep learning operator. Then, we utilize this framework during the latency-aware NAS to find the model with the highest accuracy satisfying the latency constraint. As a result, we show that the searched model with FlowOptimizer outperforms the performance by 87.1% and 92.3% on average compared to the searched model with NVDLA and Eyeriss, respectively, with better accuracy on a proxy dataset. We also show that the searched model can be transferred to a larger model to classify a more complex image dataset, i.e., ImageNet, achieving 0.2%/5.4% higher Top-1/Top-5 accuracy compared to MobileNetV2-1.0 with 3.6 \( \times \) lower latency.
Jooyeon Lee, Junsang Park, Jaeha Kung 0001
ACM Trans. Design Autom. Electr. Syst.4
2021 Design and Analysis of Approximate Compressors for Balanced Error Accumulation in MAC Operator
abstract
In this paper, we present a novel approximate computing scheme suitable for realizing the energy-efficient multiply-accumulate (MAC) processing. In contrast to the prior works that suffer from the error accumulation limiting the approximate range, we utilize different approximate multipliers in an interleaved way to compensate errors in the opposite direction during accumulate operations. For the balanced error accumulation, we first design the approximate 4-2 compressors generating errors in the opposite direction while minimizing the computational costs. Based on the probabilistic analysis, positive and negative multipliers are then carefully developed to provide a similar error distance. Simulation results on various practical applications reveal that the proposed MAC processing offers the energy-efficient computing scenario by extending the range of approximate parts. Even compared to the state-of-the-art solutions, for example, the proposed interleaving scheme relaxes the core-level energy consumption of the recent CNN accelerator by more than 35% without degrading the recognition accuracy.
Gunho Park, Jaeha Kung 0001, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 High-throughput Near-Memory Processing on CNNs with 3D HBM-like Memory
abstract
This article discusses the high-performance near-memory neural network (NN) accelerator architecture utilizing the logic die in three-dimensional (3D) High Bandwidth Memory– (HBM) like memory. As most of the previously reported 3D memory-based near-memory NN accelerator designs used the Hybrid Memory Cube (HMC) memory, we first focus on identifying the key differences between HBM and HMC in terms of near-memory NN accelerator design. One of the major differences between the two 3D memories is that HBM has the centralized through- silicon-via (TSV) channels while HMC has distributed TSV channels for separate vaults. Based on the observation, we introduce the Round-Robin Data Fetching and Groupwise Broadcast schemes to exploit the centralized TSV channels for improvement of the data feeding rate for the processing elements. Using synthesized designs in a 28-nm CMOS technology, performance and energy consumption of the proposed architectures with various dataflow models are evaluated. Experimental results show that the proposed schemes reduce the runtime by 16.4–39.3% on average and the energy consumption by 2.1–5.1% on average compared to conventional data fetching schemes.
Naebeom Park, Sungju Ryu, Jaeha Kung 0001, Jae-Joon Kim
ACM Trans. Design Autom. Electr. Syst.3
2020 Balancing Computation Loads and Optimizing Input Vector Loading in LSTM Accelerators
abstract
The long short-term memory (LSTM) is a widely used neural network model for dealing with time-varying data. To reduce the memory requirement, pruning is often applied to the weight matrix of the LSTM, which makes the matrix sparse. In this paper, we present a new sparse matrix format, named rearranged compressed sparse column (RCSC), to maximize the inference speed of the LSTM hardware accelerator. The RCSC format speeds up the inference by: 1) evenly distributing the computation loads to processing elements (PEs) and 2) reducing the input vector load miss within the local buffer. We also propose a hardware architecture adopting hierarchical input buffer to further reduce the pipeline stalls which cannot be handled by the RCSC format alone. The simulation results for various datasets show that combined use of the RSCS format and the proposed hardware requires 2× smaller inference runtime on average compared to the previous work.
Junki Park, Wooseok Yi, Daehyun Ahn, Jaeha Kung 0001, Jae-Joon Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Peregrine: A Flexible Hardware Accelerator for LSTM with Limited Synaptic Connection Patterns
abstract
In this paper, we present an integrated solution to design a high-performance LSTM accelerator. We propose a fast and flexible hardware architecture, named Peregrine, supported by a stack of innovations from algorithm to hardware design. Peregrine first minimizes the memory footprint by limiting the synaptic connection patterns within the LSTM network. Also, Peregrine provides parallel Huffman decoders with adaptive clocking to provide flexibility in dealing with a wide range of sparsity levels in the weight matrices. All these features are incorporated in a novel hardware architecture to maximize energy-efficiency. As a result, Peregrine improves performance by ~38% and energy-efficiency by ~33% in speech recognition compared to the state-of-the-art LSTM accelerator.
Jaeha Kung 0001, Junki Park, Sehun Park, Jae-Joon Kim
DAC1
2019 Similarity-Based LSTM Architecture for Energy-Efficient Edge-Level Speech Recognition
abstract
Targeting the resource-limited edge devices, we present a novel processing architecture of long short-term memory (LSTM) networks for low-power speech recognition. The proposed scheme newly defines the similarity score between two inputs of adjacent LSTM cells, and then the processing mode of the current LSTM cell is dynamically determined to reduce the energy while providing the accurate recognition. If the similarity is high, more precisely, the current cell is disabled and the outputs are directly copied from the prior vectors, totally eliminating complex LSTM operations. To maximize the skipping ratio without degrading the accuracy, for the first time, we analyze the effects of skipping the consecutive cells and set the upper limit of the number of consecutive skips. When two adjacent inputs are weakly similar, in addition, we modify the concept of the previous delta-computing, which approximately activate the LSTM cell with low computational resolution, further reducing the energy consumption. Compared to the previous state-of-the-art solutions, as a result, the proposed LSTM architecture remarkably saves the energy consumed for the accurate speech recognition, which is suitable to the resource-limited embedded edges.
Junseo Jo, Jaeha Kung 0001, Sunggu Lee, Youngjoo Lee 0002
ISLPED2
2019 WMixNet: An Energy-Scalable and Computationally Lightweight Deep Learning Accelerator
abstract
In this paper, we present a lightweight CNN model named MixNet which is easily scalable to different energy requirements in embedded platforms. The MixNet model uses two extreme bit-precisions that efficiently balances the model accuracy and energy consumption. The energy consumption in processing MixNet is managed by controlling the ratio between high-precision (16bit) and low-precision (1bit) paths. Since only two bit-precisions are required in designing a hardware accelerator, the control logic becomes simpler compared to other multi-precision accelerators. In addition, a reconfigurable multiplier is proposed to enable highly parallel MixNet computations for faster prediction and/or training. Overall, the energy efficiency in terms of run-time per unit power improves by 1.75 ~1.94 × over the recently proposed reduced-precision CNN model.
Sangwoo Jung 0001, Seungsik Moon, Youngjoo Lee 0002, Jaeha Kung 0001
ISLPED4
2018 The CAMEL approach to stacked sensor smart cameras
abstract
Stacked image sensor systems combine an image sensor, memory, and processors using 3D technology. Stacking camera components that have traditionally been packaged separately provides several benefits: very high bandwidth out of the image sensor, allowing for higher frame rates; very low latency, providing opportunities for image processing and computer vision algorithms which can adapt at very high rates; and lower power consumption. This paper will review the characteristics of stacked image sensor systems and discuss novel algorithmic and systems concepts that are made possible by these stacked sensors.
Saibal Mukhopadhyay, Marilyn Wolf, Mohammed Faisal Amir, Evan Gebhardt, Jong Hwan Ko, Jaeha Kung 0001, Burhan Ahmad Mudassar
DATE6
2018 Maximizing system performance by balancing computation loads in LSTM accelerators
abstract
The LSTM is a popular neural network model for modeling or analyzing the time-varying data. The main operation of LSTM is a matrix-vector multiplication and it becomes sparse (spMxV) due to the widely-accepted weight pruning in deep learning. This paper presents a new sparse matrix format, named CBSR, to maximize the inference speed of the LSTM accelerator. In the CBSR format, speed-up is achieved by balancing out the computation loads over PEs. Along with the new format, we present a simple network transformation to completely remove the hardware overhead incurred when using the CBSR format. Also, the detailed analysis on the impact of network size or the number of PEs is performed, which lacks in the prior work. The simulation results show 16~38% improvement in the system performance compared to the well-known CSC/CSR format. The power analysis is also performed in 65nm CMOS technology to show 9~22% energy savings.
Junki Park, Jaeha Kung 0001, Wooseok Yi, Jae-Joon Kim
DATE2
2018 Adaptive Precision Cellular Nonlinear Network
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Adaptive weight compression for memory-efficient neural networks
abstract
Neural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents an application of JPEG image encoding to compress the weights by exploiting the spatial locality and smoothness of the weight matrix. To minimize the loss of accuracy due to JPEG encoding, we propose to adaptively control the quantization factor of the JPEG algorithm depending on the error-sensitivity (gradient) of each weight. With the adaptive compression technique, the weight blocks with higher sensitivity are compressed less for higher accuracy. The adaptive compression reduces memory requirement, which in turn results in higher performance and lower energy of neural network hardware. The simulation for inference hardware for multilayer perceptron with the MNIST dataset shows up to 42X compression with less than 1% loss of recognition accuracy, resulting in 3X higher effective memory bandwidth and ~19X lower system energy.
Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Jaeha Kung 0001, Saibal Mukhopadhyay
DATE4
2017 On-chip training of recurrent neural networks with limited numerical precision
abstract
Training of neural network can be accelerated by limited numerical precision together with specialized low-precision hardware. This paper studies how low precision can impact on entire training of RNNs. We emulate low precision training for recently proposed gated recurrent unit (GRU) and use dynamic fixed point as a target numeric format. We first show that batch normalization on input sequences can help speed up training with low precision as well as high precision. We also show that the overflow rate should be carefully controlled for dynamic fixed point. We study low precision training with various rounding options including bit truncation, round to nearest, and stochastic rounding. Stochastic rounding shows superior results than the other options. The effect of fully low precision training is also analyzed by comparing partial low precision training. We show that the piecewise linear activation function with stochastic rounding can achieve comparable training results with floating point precision. Low precision multiplier and accumulator (MAC) with linear-feedback shift register (LFSR) is implemented with 28nm Synopsys PDK for energy and performance analysis. Implementation results show low precision hardware is 4.7× faster, and energy per task is up to 4.55× lower than that of floating point hardware.
Taesik Na, Jong Hwan Ko, Jaeha Kung 0001, Saibal Mukhopadhyay
IJCNN3
2017 A Programmable Hardware Accelerator for Simulating Dynamical Systems
abstract
The fast and energy-efficient simulation of dynamical systems defined by coupled ordinary/partial differential equations has emerged as an important problem. The accelerated simulation of coupled ODE/PDE is critical for analysis of physical systems as well as computing with dynamical systems. This paper presents a fast and programmable accelerator for simulating dynamical systems. The computing model of the proposed platform is based on multilayer cellular nonlinear network (CeNN) augmented with nonlinear function evaluation engines. The platform can be programmed to accelerate wide classes of ODEs/PDEs by modulating the connectivity within the multilayer CeNN engine. An innovative hardware architecture including data reuse, memory hierarchy, and near-memory processing is designed to accelerate the augmented multilayer CeNN. A dataflow model is presented which is supported by optimized memory hierarchy for efficient function evaluation. The proposed solver is designed and synthesized in 15nm technology for the hardware analysis. The performance is evaluated and compared to GPU nodes when solving wide classes of differential equations and the power consumption is analyzed to show orders of magnitude improvement in energy efficiency.
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay
ISCA1
2016 ReRAM Crossbar based Recurrent Neural Network for human activity detection
abstract
We present a programmable high-efficient Recurrent Neural Network (RNN) with Synapses design using Resistive Random Access Memory (ReRAM). The presented ReRAM-RNN employs crossbar ReRAM arrays as synapses. A fast synapses programming methodology is realized by CMOS-based neuron with in-built programming circuitry. The simulations are performed using experimentally verified physical resistive switching model, instead of only functional models, providing better estimate of system speed and power efficiency. Simulation results show that ReRAM-RNN can provide higher computation efficiency and/or more compact design than software realization of RNN, and dedicated CMOS based digital-and analog-RNN. We show that the efficiency improvement of ReRAM-based neural network design is more significant in feedback networks than in feedforward networks.
Eui Min Jung, Jaeha Kung 0001, Saibal Mukhopadhyay
IJCNN3
2016 Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory
abstract
This paper presents a programmable and scalable digital neuromorphic architecture based on 3D high-density memory integrated with logic tier for efficient neural computing. The proposed architecture consists of clusters of processing engines, connected by 2D mesh network as a processing tier, which is integrated in 3D with multiple tiers of DRAM. The PE clusters access multiple memory channels (vaults) in parallel. The operating principle, referred to as the memory centric computing, embeds specialized state-machines within the vault controllers of HMC to drive data into the PE clusters. The paper presents the basic architecture of the Neurocube and an analysis of the logic tier synthesized in 28nm and 15nm process technologies. The performance of the Neurocube is evaluated and illustrated through the mapping of a Convolutional Neural Network and estimating the subsequent power and performance for both training and inference.
Duckhwan Kim 0001, Jaeha Kung 0001, Sek M. Chai, Sudhakar Yalamanchili, Saibal Mukhopadhyay
ISCA2
2016 Dynamic Approximation with Feedback Control for Energy-Efficient Recurrent Neural Network Hardware
abstract
This paper presents methodology of feedback-controlled dynamic approximation to enable energy-accuracy trade-off in digital recurrent neural network (RNN). A low-power digital RNN engine is presented that employs the proposed dynamic approximation. The on-chip feedback controller is realized by utilizing hysteretic or proportional controller. The dynamic adaptation of bit-precisions during the RNN computation is selected as approximation approach. Considering various applications, the digital RNN engine designed in 28nm CMOS shows ~36% average energy saving compared to the baseline case, with only ~4% of accuracy degradation on average.
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay
ISLPED1
2015 A power-aware digital feedforward neural network platform with backpropagation driven approximate synapses
abstract
This paper proposes a power-aware digital feedforward neural network platform that utilizes the backpropagation algorithm during training to enable energy-quality trade-off. Given a quality constraint, the proposed approach identifies a set of synaptic weights for approximation in a neural network. The approach selects synapses with small impact on output error, estimated by the backpropagation algorithm, for approximation. The approximations are achieved by a coupled software (reduced bit-width) and hardware (approximate multiplication in the processing engine) based design approaches. The full-chip design in 130nm CMOS shows, compared to a baseline accurate design, the proposed approach reduces system power by ~38% with 0.4% lower recognition accuracy in a classification problem.
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay
ISLPED1
2015 On the Impact of Energy-Accuracy Tradeoff in a Digital Cellular Neural Network for Image Processing
abstract
This paper studies the opportunities of energy-accuracy tradeoff in cellular neural network (CNN). Algorithmic characteristics of CNN is coupled with hardware-induced error distribution of a digital CNN cell to evaluate energy-accuracy tradeoff for simple image processing tasks as well as a complex application. The analysis shows that errors modulate the cell dynamics and propagate through the network degrading the output quality and increasing the convergence time. The error propagation is determined by the task being performed by the CNN, specifically, the strength of the feedback template. Controlling precision is observed to be a more effective approach for energy-accuracy tradeoff in CNN than voltage over scaling.
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1