Wonbo Shim

dblp:271/1999 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-9669-7310ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference Computing
abstract
Transformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead.
Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu
IEEE Trans. Very Large Scale Integr. Syst.4
2025 Dissecting and Re-Architecting 3D NAND Flash PIM Arrays for Efficient Single-Batch Token Generation in LLMS
abstract
The advancement of large language models has led to models with billions of parameters, significantly increasing memory and compute demands. Serving such models on conventional hardware is challenging due to limited DRAM capacity and high GPU costs. Thus, in this work, we propose offloading the single-batch token generation to a 3D NAND flash processing-in-memory (PIM) device, leveraging its high storage density to overcome the DRAM capacity wall. We explore 3D NAND flash configurations and present a re-architected PIM array with an H-tree network for optimal latency and cell density. Along with the well-chosen PIM array size, we develop operation tiling and mapping methods for LLM layers, achieving a$2.4 \times$speedup over four RTX4090 with vLLM and comparable performance to four A100 with only 4.9% latency overhead. Our detailed area analysis reveals that the proposed 3D NAND flash PIM architecture can be integrated within a$4.98 ~\text{mm}^{2}$die area under the memory array, without extra area overhead.
Yongjoo Jang, Sangwoo Hwang, Sangwoo Jung 0001, Wonbo Shim, Jaeha Kung 0001
ICCD6
2025 E-Flash: Energy-Efficient DNN Mapping on NAND Flash Memory with State-Switching Algorithm
abstract
Deep neural network (DNN) has been widely adopted in various applications. Ranging from image classification to text generation, Transformer-based models have demonstrated unprecedented performance. However, they suffer from a significant computational complexity and a large memory footprint, leading to memory-bound issues. While previous research on NAND flash-based neural network computation has been performed, these studies often encounter accuracy problems, as analog processing-in memory (PIM) operations typically lead to inaccurate results. Moreover, studies on NAND flash using single-level cell (SLC) are unable to fully leverage the advantages of efficient storage density on the multi-level cell (MLC) memory. We propose an E-Flash hardware architecture with an energy-efficient DNN mapping method. E-Flash introduces a state-switching algorithm to perform data movements between flash memory and host device in an energy-efficient manner. By reallocating the data in triple-level cell (TLC) NAND flash memory, we reduce the energy consumption during the data read operation. Experimental results demonstrate that E-Flash achieves improved energy consumption compared to baseline under significantly small area overhead for quantized BERT and Llama 2 weights by 37.73% and 16.74%, respectively.
Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu
ISLPED4
2021 RRAM for Compute-in-Memory: From Inference to Training
abstract
To efficiently deploy machine learning applications to the edge, compute-in-memory (CIM) based hardware accelerator is a promising solution with improved throughput and energy efficiency. Instant-on inference is further enabled by emerging non-volatile memory technologies such as resistive random access memory (RRAM). This paper reviews the recent progresses of the RRAM based CIM accelerator design. First, the multilevel states RRAM characteristics are measured from a test vehicle to examine the key device properties for inference. Second, a benchmark is performed to study the scalability of the RRAM CIM inference engine and the feasibility towards monolithic 3D integration that stacks RRAM arrays on top of advanced logic process node. Third, grand challenges associated with in-situ training are presented. To support accurate and fast in-situ training and enable subsequent inference in an integrated platform, a hybrid precision synapse that combines RRAM with volatile memory (e.g. capacitor) is designed and evaluated at system-level. Prospects and future research needs are discussed.
Shimeng Yu, Wonbo Shim, Xiaochen Peng, Yandong Luo
IEEE Trans. Circuits Syst. I Regul. Pap.2