Sanghun Shin

dblp:290/3702 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 NITRO: 3D NAND Flash-Based In-Storage LLM Computing with Enhanced Activation Dataflow
abstract
In-storage computing (ISC) has emerged as a next-generation memory architecture to relieve the data movement bottleneck between host processors and memory systems. While recent NAND flash-based processing-in-memory works leverage the high density of 3D NAND flash for deep neural networks, they primarily focus on optimizing computation inside the NAND array. Consequently, these approaches often fail to address the critical latency overhead associated with managing intermediate activation data. To overcome such a limitation, we propose a heterogeneous NAND flash-based ISC architecture with enhanced activation buffering. By buffering intermediate values in a DRAM subsystem rather than programming them into the NAND flash array, our approach effectively mitigates the high programming latency penalties. We also introduce a distributed dataflow scheme that maximizes computational parallelism through optimized plane- and bank-level data mapping. The results show that our proposed architecture achieves performance improvements, reducing inference latency by up to 86% compared to the baseline.
Sanghun Shin, Gisan Ji, Sungju Ryu
DATE1
2026 E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference Computing
abstract
Transformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead.
Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu
IEEE Trans. Very Large Scale Integr. Syst.2
2025 E-Flash: Energy-Efficient DNN Mapping on NAND Flash Memory with State-Switching Algorithm
abstract
Deep neural network (DNN) has been widely adopted in various applications. Ranging from image classification to text generation, Transformer-based models have demonstrated unprecedented performance. However, they suffer from a significant computational complexity and a large memory footprint, leading to memory-bound issues. While previous research on NAND flash-based neural network computation has been performed, these studies often encounter accuracy problems, as analog processing-in memory (PIM) operations typically lead to inaccurate results. Moreover, studies on NAND flash using single-level cell (SLC) are unable to fully leverage the advantages of efficient storage density on the multi-level cell (MLC) memory. We propose an E-Flash hardware architecture with an energy-efficient DNN mapping method. E-Flash introduces a state-switching algorithm to perform data movements between flash memory and host device in an energy-efficient manner. By reallocating the data in triple-level cell (TLC) NAND flash memory, we reduce the energy consumption during the data read operation. Experimental results demonstrate that E-Flash achieves improved energy consumption compared to baseline under significantly small area overhead for quantized BERT and Llama 2 weights by 37.73% and 16.74%, respectively.
Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu
ISLPED2
2023 Decoupling thermal effects in GaN photodetectors for accurate measurement of ultraviolet intensity using deep neural network
Keuntae Baek, Sanghun Shin, Hongyun So
Eng. Appl. Artif. Intell.2