Sangwoo Jung 0001

dblp:248/3490-1 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0008-1150-9888ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Hybrid Digital-Analog Compute-in-Memory Using Content-Addressable Memory With Flexible Multi-Bit Slicing
abstract
Compute-in-memory (CIM) reduces data movement and enhances compute parallelism, making it suitable for AI applications. However, analog CIMs, yet energy-efficient, are vulnerable to PVT variations, while digital CIMs offer robustness but limited efficiency due to their bit-wise computation overhead. To address these challenges, we propose a hybrid CIM architecture that integrates content-addressable memory (CAM) and cluster-based CIM, named CAM-CIM, fabricated in 65nm CMOS technology. The proposed CAM-CIM flexibly slices multi-bit weights, assigning MSBs to CAM and LSBs to CIM, enabling dynamic accuracy-efficiency trade-offs across various bit precisions. A two-stage 8:3 compressor-based adder tree improves CAM efficiency and a reference voltage search algorithm ensures accurate CIM computation with low-bit ADCs. Our CAM-CIM supports 1-8b inputs/weights with reconfigurable compute modes, leveraging ternary-CAM based selective columns and cluster-wise CIM processing to produce multiple trade-off points even in the same bit precision. A prototype chip with a RISC-V controller and custom instructions is demonstrated that shows energy efficiencies of 32.4TOPS/W (8b/8b) and 76.0-354.9TOPS/W (4b/4b) with 0.66% accuracy loss, on average, across a wide range of DNN benchmarks including CNNs and vision transformers on CIFAR and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2026 Fused-STA: Automated Design Space Exploration of a Fused Systolic Tensor Array for Universal Deep Learning Acceleration
abstract
Systolic array architectures have become the dominant solution for accelerating deep neural network (DNN) computations, yet designing optimal configurations remains challenging due to the vast design space spanning array dimensions, dataflow strategies, SRAM allocations, and tiling sizes. Existing tool-aided optimization frameworks constrain this design space by treating SRAM sizes or tiling sizes as fixed input parameters, while focusing only on conventional systolic arrays, limiting their ability to discover a truly optimal design. This article presents a comprehensive framework for systolic tensor array (STA) design space exploration that jointly optimizes array configurations, SRAM sizes, tiling sizes, and dataflow types. We introduce TensorSim, a cycle-accurate simulator with an ML-based synthesis prediction model, and TensorOptimizer, an automated multiobjective optimization framework. Evaluation with TensorOptimizer across 12 representative DNNs—from edge-level convolutional neural networks (CNNs) to billion-parameter large language models (LLMs)—reveals that output stationary (OS) dataflows outperform weight stationary (WS) dataflows by$1.53\times $on edge workloads, while WS achieves$2.13\times $higher efficiency over OS on server-scale computations. Based on these insights, we propose Fused-STA, a reconfigurable architecture that dynamically switches between OS and WS modes within a unified hardware substrate. Fused-STA achieves over 91% of the oracle single-dataflow performance, specifically designed for a single network, providing$2.34\times $and$1.92\times $improvements over a tensor processing unit (TPU)-like baseline on edge-level and server-level benchmarks, respectively.
Jooyeon Lee, Sangwoo Jung 0001, Jaeha Kung 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 RISC-V Driven Orchestration of Vector Processing Units and eFlash Compute-in-Memory Arrays for Fast and Accurate Keyword Spotting
abstract
In this paper, we propose a computationally efficient keyword spotting (KWS) model, named hybrid reparameterized FSMN (HRepFSMN), by carefully examining the impact of binarization on the accuracy. In particular, we found that binarizing depthwise convolution (DW-Conv) within the previous binarized KWS model, i.e., BiFSMNv2, does not lead to a significant reduction in FLOPs. Therefore, we allow floating-point (FP) operations on less computation-intensive DW-Conv layers while the remaining layers are computed in a binary fashion (hybrid data type). In addition, we remove skip connections, which require data fetching in full precision, by applying a reparameterization technique. More importantly, to efficiently compute the proposed HRepFSMN, we present a RISC-V controlled hardware accelerator that consists of reconfigurable vector processing units for FP operations and eFlash compute-in-memory arrays for binary operations. We extend RISC-V instructions so that the core can efficiently manage both computing fabrics. As a result, our HRepFSMN improves accuracy by 2.57%/4.98% with 24.02×/3.66× speed-up compared to BiFSMNv2/BiFSMNv2_small. By shrinking down our HRepFSMN, we achieve 0.95% higher accuracy with 20.87× speed-up compared to BiFSMNv2_small.
Gunil Kang, Dahoon Park, Sangwoo Jung 0001, Jung Gyu Min, Youngjoo Lee 0002, Jaeha Kung 0001
ASP-DAC4
2025 Dissecting and Re-Architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMS
abstract
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional hardware is challenging due to limited DRAM capacity and high GPU costs. Thus, in this work, we propose offloading the single-batch token generation to a 3D NAND flash processing-in-memory (PIM) device, leveraging its high storage density to overcome the DRAM capacity wall. We explore 3D NAND flash configurations and present a re-architected PIM array with an H-tree network for optimal latency and cell density. Along with the well-chosen PIM array size, we develop operation tiling and mapping methods for LLM layers, achieving a$2.4 \times$speedup over four RTX4090 with vLLM and comparable performance to four A100 with only 4.9% latency overhead. Our detailed area analysis reveals that the proposed 3D NAND flash PIM architecture can be integrated within a$4.98 ~\text{mm}^{2}$die area under the memory array, without extra area overhead.
Yongjoo Jang, Sangwoo Hwang, Sangwoo Jung 0001, Wonbo Shim, Jaeha Kung 0001
ICCD4
2025 CAM-CIM: A Hybrid Compute-in-Memory Using Content-Addressable Memory with Subword Split Mapping for Reduced ADC Resolution
abstract
Recently, compute-in-memory (CIM) has become a promising architecture for data-intensive applications such as deep learning. However, analog or digital CIM (ACIM or DCIM) faces some design challenges. ACIMs inherently have non-idealities, which lead to significant accuracy degradation. In addition, a substantial amount of power is consumed by analog-to-digital converters (ADC). On the other hand, DCIMs show an exponential increase in power consumption and computing cycles as the operand bit-width increases, particularly due to an accumulation stage. In this paper, to overcome these challenges, we propose a hybrid DCIM-ACIM architecture that consists of a content addressable memory (CAM) as DCIM and a cluster-based multi-cycle ACIM, called CAM-CIM. As a weight mapping strategy, we present a subword split mapping that assigns some MSBs to DCIM for improved accuracy and the remaining LSBs to ACIM for reduced ADC resolution. The accuracy of using the proposed CAM-CIM array is evaluated on various deep learning benchmarks from CNNs to Swin-Tiny. A 65nm CAM-CIM macro with either 3-bit or 4-bit ADCs shows 10.3 × and 5.4 × improvement in energy efficiency, on average, compared to CAM- and CIM-only architectures, respectively. Compared to recent CIM architectures, CAM-CIM demonstrates 1.4 × higher energy efficiency.
Sangwoo Jung 0001, Dahoon Park, Hyunseob Shin, Jong-Hyeok Yoon, Jaeha Kung 0001
ISLPED1
2024 OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
abstract
To overcome the burden on the memory size and bandwidth due to ever-increasing size of large language models (LLMs), aggressive weight quantization has been recently studied, while lacking research on quantizing activations. In this paper, we present a hardware-software co-design method that results in an energy-efficient LLM accelerator, named OPAL, for generation tasks. First of all, a novel activation quantization method that leverages the microscaling data format while preserving several outliers per subtensor block (e.g., four out of 128 elements) is proposed. Second, on top of preserving outliers, mixed precision is utilized that sets 5-bit for inputs to sensitive layers in the decoder block of an LLM, while keeping inputs to less sensitive layers to 3-bit. Finally, we present the OPAL hardware architecture that consists of FP units for handling outliers and vectorized INT multipliers for dominant non-outlier related operations. In addition, OPAL uses log2-based approximation on softmax operations that only requires shift and subtraction to maximize power efficiency. As a result, we are able to improve the energy efficiency by 1.6~2.2×, and reduce the area by 2.4~3.1× with negligible accuracy loss, i.e., <1 perplexity increase.
Jahyun Koo 0002, Dahoon Park, Sangwoo Jung 0001, Jaeha Kung 0001
DAC3
2024 A Dual-Precision and Low-Power CNN Inference Engine Using a Heterogeneous Processing-in-Memory Architecture
abstract
In this article, we present an energy-scalable CNN model that can adapt to different hardware resource constraints. Specifically, we propose a dual-precision network, named DualNet, that leverages two independent bit-precision paths (INT4 and ternary-binary). DualNet achieves both high accuracy and low complexity by balancing the ratio between two paths. We also present an evolutionary algorithm that allows the automatic search of the optimal ratios. In addition to the novel CNN architecture design, we develop a heterogeneous processing-in-memory (PIM) hardware that integrates SRAM-and eDRAM-based PIMs to efficiently compute two precision paths in parallel. To verify the energy efficiency of DualNet computed on the heterogeneous PIM, we prototyped a test chip in 28nm CMOS technology. To maximize the hardware efficiency, we utilize an improved data mapping scheme achieving the most effective deployment of DualNets on multiple PIM arrays. With the proposed SW-HW co-optimization, we can obtain the most energy-efficient DualNet model operating on the actual PIM hardware. Compared to the other quantized networks with a single bit-precision, DualNet reduces the energy consumption, memory footprint, and latency by 29.0%, 49.5%, 47.3% on average, respectively, for CIFAR-10/100 and ImageNet datasets.
Sangwoo Jung 0001, Dahoon Park, Youngjoo Lee 0002, Jong-Hyeok Yoon, Jaeha Kung 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2019 WMixNet: An Energy-Scalable and Computationally Lightweight Deep Learning Accelerator
abstract
In this paper, we present a lightweight CNN model named MixNet which is easily scalable to different energy requirements in embedded platforms. The MixNet model uses two extreme bit-precisions that efficiently balances the model accuracy and energy consumption. The energy consumption in processing MixNet is managed by controlling the ratio between high-precision (16bit) and low-precision (1bit) paths. Since only two bit-precisions are required in designing a hardware accelerator, the control logic becomes simpler compared to other multi-precision accelerators. In addition, a reconfigurable multiplier is proposed to enable highly parallel MixNet computations for faster prediction and/or training. Overall, the energy efficiency in terms of run-time per unit power improves by 1.75 ~1.94 × over the recently proposed reduced-precision CNN model.
Sangwoo Jung 0001, Seungsik Moon, Youngjoo Lee 0002, Jaeha Kung 0001
ISLPED1