EDBT 2026 Demo / reviewers in the wild / expert
Sangwoo Ha
dblp:320/0524
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-9191-6326ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A 4.21 TFLOPS/W Memory-Efficient LLM Inference Accelerator with Bit-Layered Non-Uniform QuantizationabstractNon-uniform Quantization (NUQ) is widely used in LLM accelerators due to its high accuracy. However, employing NUQ with models of varying sizes can substantially increase storage requirements on mobile devices. This paper presents a bit-layered NUQ accelerator architecture that supports multiple bit-width configurations while minimizing memory usage. Key features include Reconfigurable Condensed Look-up Accumulator (RCLA), Dual-Sign Path Accumulation (DSPA), and MSB-Sparse Encoding Compression (MSEC). RCLA enables the use of multiple weight precisions within a uniform PE array. In particular, it optimizes PE utilization in high-bit-width NUQ modes, reducing accumulation cycles by 63.2 %. DSPA facilitates energy-efficient computation, resulting in an average power reduction of 40.7 % across various weight modes. MSEC enhances weight compression, reducing the data storage size of each bit plane by up to 47.7 %. The proposed design supports models of different sizes, improves energy efficiency, and reduces memory capacity requirements, rendering it ideal for mobile LLM inference. Byeongcheol Kim, Sangwoo Ha, Soyeon Um, Kyomin Sohn, Hoi-Jun Yoo |
ISCAS | 3 |
| 2025 | A 32.65µm2 Spin/Area Large Scale Ising CIM with Progressive Circular Dataflow and Bi-directional eDRAM Cell ArrayabstractThis paper presents a large-scale Ising computing-in-memory (CIM) for a real-world combinatorial optimization problem (COP). While the CIM approach shows promising performance improvement compared to digital-based Ising machines, it can be only used for simple COP due to limited CIM connectivity, bit precision, and graph size. To overcome limitations, the proposed processor supports reconfigurable high-bit Chimera graph topology achieving a small cell area and high energy efficiency through three key features : 1) Progressive circular dataflow reduce area by 58.8% and improved energy efficiency by 44.7% thanks to fully reuse spin and coefficient. 2) Bi-directional eDRAM cell array support bi-directional ising computation for circular dataflow with 3T-2C eDRAM cell. Due to the signed operation with compact coefficient storing, the area was reduced by 36.1%. 3) Reconfigurable spin exchange link and reconfigurable C-2C ladder support various graph sizes and various bit precision of coefficients. In conclusion, the proposed large-scale Ising CIM achieves 32.65µm2spin area and 2.09µW effective spin power which is 3.21× and 5.17× smaller than previous state-of-the art. Jingu Lee, Sangwoo Ha, Sunjoo Whang, Soyeon Um, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 2 |
| 2025 | A 2.67 mJ/frame Video Mamba Accelerator with Importance-aware Redundancy Elimination and SSM Computing ReformulationabstractAn energy-efficient video understanding processor, SLYTHERIN, is proposed to accelerate the new AI model, mamba, efficiently on edge devices. Mamba, a state-of-the-art model for in-context learning, is designed to replace the transformer, whose computational complexity increases significantly in video applications. However, the acceleration of video mamba on edge devices presents two main challenges: 1) slow inference due to the iterative operation phase of mamba and 2) the large energy consumption caused by external memory access (EMA). To address these challenges, the SLYTHERIN is proposed with 3 key building blocks. 1) A 6-stage pipelined task allocator minimizes computational complexity by dynamically managing redundant computations with importance-aware prediction. 2) Reformed computing SSM engine increases core efficiency by tackling the overheads in mamba’s iterative stages with reordering and distributed L1 cache. 3) Patch data management unit addresses the large EMA with difference-based bit-sliced data compression. It finally achieves 2.67-mJ/frame system energy efficiency with 274 FPS. Youngjin Moon, Sangwoo Ha, Junha Ryu, Hoi-Jun Yoo, Donghyeon Han |
ISCAS | 2 |
| 2025 | A 62.8 TOPS/W FP-INT Digital Computing-in Memory Processor with Bit-Reordered Adder Tree and Low Active Hierarchical AccumulatorabstractThis paper introduces an energy-efficient DRAM-based FP-INT Digital Computing-in-Memory (CIM) processor for data-intensive AI workloads, addressing challenges in mixed precision, excess power consumption, and area overhead. The processor features three key innovations: 1) Bit-reordered adder tree (BRAT) that reduces adder bit width through exponent-aware reordering, achieving power and area reductions by 23% and 56%, respectively; 2) Low active hierarchical accumulator (LAHA) that conditionally updates the FP accumulator, minimizing accumulator activity by 57%; and 3) Speculative input block skipping (SIBS) to avoid unnecessary computations, reducing energy consumption by 15%. Implemented in 28nm CMOS technology, the processor achieves 62.8 TOPS/W in FP8 activation and INT4 weight precision for GPT2-Small inference. Our results show an overall energy savings of 31%, making this architecture highly efficient for modern large language models. Sunjoo Whang, Sangwoo Ha, Soyeon Um, Hoi-Jun Yoo |
ISCAS | 2 |