Sangwoo Ha

dblp:320/0524 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-9191-6326ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2025 A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform Quantization
abstract
Non-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference.
Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo
ISCAS3
2025 A 32.65µm2 Spin/Area Large Scale Ising CIM with Progressive Circular Dataflow and Bi-directional eDRAM Cell Array
abstract
This paper presents a large-scale Ising computing-in-memory (CIM) for a real-world combinatorial optimization problem (COP). While the CIM approach shows promising performance improvement compared to digital-based Ising machines, it can be only used for simple COP due to limited CIM connectivity, bit precision, and graph size. To overcome limitations, the proposed processor supports reconfigurable high-bit Chimera graph topology achieving a small cell area and high energy efficiency through three key features : 1) Progressive circular dataflow reduce area by 58.8% and improved energy efficiency by 44.7% thanks to fully reuse spin and coefficient. 2) Bi-directional eDRAM cell array support bi-directional ising computation for circular dataflow with 3T-2C eDRAM cell. Due to the signed operation with compact coefficient storing, the area was reduced by 36.1%. 3) Reconfigurable spin exchange link and reconfigurable C-2C ladder support various graph sizes and various bit precision of coefficients. In conclusion, the proposed large-scale Ising CIM achieves 32.65µm2spin area and 2.09µW effective spin power which is 3.21× and 5.17× smaller than previous state-of-the art.
Jingu Lee, Sangwoo Ha, Sunjoo Whang, Soyeon Um, Wooyoung Jo, Hoi-Jun Yoo
ISCAS2
2025 A 2.67 mJ/frame Video Mamba Accelerator with Importance-aware Redundancy Elimination and SSM Computing Reformulation
abstract
An energy-efficient video understanding processor, SLYTHERIN, is proposed to accelerate the new AI model, mamba, efficiently on edge devices. Mamba, a state-of-the-art model for in-context learning, is designed to replace the transformer, whose computational complexity increases significantly in video applications. However, the acceleration of video mamba on edge devices presents two main challenges: 1) slow inference due to the iterative operation phase of mamba and 2) the large energy consumption caused by external memory access (EMA). To address these challenges, the SLYTHERIN is proposed with 3 key building blocks. 1) A 6-stage pipelined task allocator minimizes computational complexity by dynamically managing redundant computations with importance-aware prediction. 2) Reformed computing SSM engine increases core efficiency by tackling the overheads in mamba’s iterative stages with reordering and distributed L1 cache. 3) Patch data management unit addresses the large EMA with difference-based bit-sliced data compression. It finally achieves 2.67-mJ/frame system energy efficiency with 274 FPS.
Youngjin Moon, Sangwoo Ha, Junha Ryu, Hoi-Jun Yoo, Donghyeon Han
ISCAS2
2025 A 62.8 TOPS/W FP-INT Digital Computing-in Memory Processor with Bit-Reordered Adder Tree and Low Active Hierarchical Accumulator
abstract
This paper introduces an energy-efficient DRAM-based FP-INT Digital Computing-in-Memory (CIM) processor for data-intensive AI workloads, addressing challenges in mixed precision, excess power consumption, and area overhead. The processor features three key innovations: 1) Bit-reordered adder tree (BRAT) that reduces adder bit width through exponent-aware reordering, achieving power and area reductions by 23% and 56%, respectively; 2) Low active hierarchical accumulator (LAHA) that conditionally updates the FP accumulator, minimizing accumulator activity by 57%; and 3) Speculative input block skipping (SIBS) to avoid unnecessary computations, reducing energy consumption by 15%. Implemented in 28nm CMOS technology, the processor achieves 62.8 TOPS/W in FP8 activation and INT4 weight precision for GPT2-Small inference. Our results show an overall energy savings of 31%, making this architecture highly efficient for modern large language models.
Sunjoo Whang, Sangwoo Ha, Soyeon Um, Hoi-Jun Yoo
ISCAS2