Soyeon Um

dblp:264/0270 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0002-8526-2047ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 9 since 2021
YearPublicationVenuePosition
2025 A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization
abstract
Non-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference.
Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
ISCAS4
2025 A 32.65µm2 Spin/Area Large Scale Ising CIM with Progressive Circular Dataflow and Bi-directional eDRAM Cell Array
abstract
This paper presents a large-scale Ising computing-in-memory (CIM) for a real-world combinatorial optimization problem (COP). While the CIM approach shows promising performance improvement compared to digital-based Ising machines, it can be only used for simple COP due to limited CIM connectivity, bit precision, and graph size. To overcome limitations, the proposed processor supports reconfigurable high-bit Chimera graph topology achieving a small cell area and high energy efficiency through three key features : 1) Progressive circular dataflow reduce area by 58.8% and improved energy efficiency by 44.7% thanks to fully reuse spin and coefficient. 2) Bi-directional eDRAM cell array support bi-directional ising computation for circular dataflow with 3T-2C eDRAM cell. Due to the signed operation with compact coefficient storing, the area was reduced by 36.1%. 3) Reconfigurable spin exchange link and reconfigurable C-2C ladder support various graph sizes and various bit precision of coefficients. In conclusion, the proposed large-scale Ising CIM achieves 32.65µm2spin area and 2.09µW effective spin power which is 3.21× and 5.17× smaller than previous state-of-the art.
Jingu Lee, Sangwoo Ha, Sunjoo Whang, Soyeon Um, Wooyoung Jo, Hoi-Jun Yoo
ISCAS6
2025 A 62.8 TOPS/W FP-INT Digital Computing-in Memory Processor with Bit-Reordered Adder Tree and Low Active Hierarchical Accumulator
abstract
This paper introduces an energy-efficient DRAM-based FP-INT Digital Computing-in-Memory (CIM) processor for data-intensive AI workloads, addressing challenges in mixed precision, excess power consumption, and area overhead. The processor features three key innovations: 1) Bit-reordered adder tree (BRAT) that reduces adder bit width through exponent-aware reordering, achieving power and area reductions by 23% and 56%, respectively; 2) Low active hierarchical accumulator (LAHA) that conditionally updates the FP accumulator, minimizing accumulator activity by 57%; and 3) Speculative input block skipping (SIBS) to avoid unnecessary computations, reducing energy consumption by 15%. Implemented in 28nm CMOS technology, the processor achieves 62.8 TOPS/W in FP8 activation and INT4 weight precision for GPT2-Small inference. Our results show an overall energy savings of 31%, making this architecture highly efficient for modern large language models.
Sunjoo Whang, Sangwoo Ha, Soyeon Um, Hoi-Jun Yoo
ISCAS3
2024 Two-Step Spike Encoding Scheme and Architecture for Highly Sparse Spiking-Neural-Network
abstract
This paper proposes a two-step spike encoding, which consists of the source encoding and process encoding for energy-efficient spiking-neural-network (SNN) acceleration. The eigen-train generation and its superposition generate spike trains which show high accuracy with low spike ratio. Sparsity boosting (SB) and spike generation skipping (SGS) reduce the number of operations for SNN. Time shrinking multi-level encoding (TS-MLE) compresses the number of spikes in a train along time axis, and spike-level clock skipping (SLCS) decreases the processing time. Eigen-train generation achieves 90.3% accuracy, the same accuracy as CNN, under the condition of 4.18% spike ratio for CIFAR-10 classification. SB reduces spike ratio by 0.49× with only 0.1% accuracy loss, and the SGS reduces the spike ratio by 20.9% with 0.5% accuracy loss. TS- MLE and SLCS increase the throughput of SNN by 2.8× while decreasing the hardware resource for spike generator by 75% compared with previous generators.
Sangyeob Kim, Soyeon Um, Hoi-Jun Yoo
ISCAS3
2023 A 332 TOPS/W Input/Weight-Parallel Computing-in-Memory Processor with Voltage-Capacitance-Ratio Cell and Time-Based ADC
abstract
Recent computing-in-memory (CIM) achieves high energy efficiency with charge-domain computation and multi-bit input driving. However, the previous works still require high power consumption and trade computation signal-to-noise ratio (SNR) for energy efficiency. This work proposes an energy-efficient and accurate multi-bit input/weight-parallel CIM processor with four key features: 1) a 10T2C sign-magnitude cell with voltage-capacitance-ratio (VCR) decoding for 5-bit analog inputs with only 2-level supply voltages, 2) a computation word line (CWL) charge reuse method for input driver power reduction, 3) a signal-amplifying noise canceling voltage-to-time converter (SANC-VTC) for SNR improvement, and 4) a distribution-aware time-to-digital converter (DA-TDC) for ADC power reduction. The proposed CIM processor is simulated in 28 nm CMOS technology with 1.25 mm2area. As a result, it achieves 4.44 mW power consumption and 332 TOPS/W energy efficiency with 72.43% benchmark accuracy (@ ImageNet, ResNet50, 5-bit input/5-bit weight).
Seongyon Hong, Soyeon Um, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo
ISCAS2
2023 A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level Efficiency
abstract
Spiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA.
Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo
ISCAS3
2023 A 5.99 TFLOPS/W Heterogeneous CIM-NPU Architecture for an Energy Efficient Floating-Point DNN Acceleration
abstract
This work presents an energy-efficient digital-based computing-in-memory (CIM) processor to support floating-point (FP) deep neural network (DNN) acceleration. Previous FP-CIM processors have two limitations. Processors with post-alignment shows low throughput due to serial operation, and the other processor with pre-alignment incurs truncation error. To resolve these problems, we focus on the statistics that outlier exists according to shift amount in pre-alignment-based FP operation. As those outlier decreases energy efficiency due to long operation cycles, it needs to be processed separately. The proposed Hetero-FP-CIM integrates both CIM arrays and shared NPU, so they compute both dense inlier and sparse outlier respectively. It also includes efficient weight caching system to avoid entire weight copy in shared NPU. The proposed Hetero-FP-CIM is simulated in 28 nm CMOS technology and occupies 2.7 mm2. As a result, it achieves 5.99 TOPS/W at ImageNet (ResNet50) with bfloat16 representation.
Wonhoon Park, Junha Ryu, Soyeon Um, Wooyoung Jo, Sangyoeb Kim, Hoi-Jun Yoo
ISCAS4
2022 Neuro-CIM: A 310.4 TOPS/W Neuromorphic Computing-in-Memory Processor with Low WL/BL activity and Digital-Analog Mixed-mode Neuron Firing
abstract
Multi WLs Driving ➔ Low Energy Efficiency by ADC (<100 TOPS/W)
Sangyeob Kim, Soyeon Um, Kwantae Kim, Hoi-Jun Yoo
HCS3
2022 A 161.6 TOPS/W Mixed-mode Computing-in-Memory Processor for Energy-Efficient Mixed-Precision Deep Neural Networks
abstract
A Mixed-mode Computing-in memory (CIM) processor for the mixed-precision Deep Neural Network (DNN) processing is proposed. Due to the bit-serial processing for the multi-bit data, the previous CIM processors could not exploit the energy-efficient computation of mixed-precision DNNs. This paper proposes an energy-efficient mixed-mode CIM processor with two key features: 1) Mixed-Mode Mixed-precision CIM (M3-CIM) which achieves 55.46% energy efficiency improvement. 2) Digital-CIM for In-memory MAC for the increased throughput of M3-CIM. The proposed CIM processor was simulated in 28nm CMOS technology and occupies 1.96 mm2. It achieves a state-of-the-art energy efficiency of 161.6 TOPS/W with 72.8% accuracy at ImageNet (ResNet50).
Wooyoung Jo, Juhyeong Lee, Soyeon Um, Zhiyong Li 0016, Hoi-Jun Yoo
ISCAS4