EDBT 2026 Demo / reviewers in the wild / expert
Gisan Ji
dblp:415/5281
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0009-0061-433XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NITRO: 3D NAND Flash-Based In-Storage LLM Computing with Enhanced Activation DataflowabstractIn-storage computing (ISC) has emerged as a next-generation memory architecture to relieve the data movement bottleneck between host processors and memory systems. While recent NAND flash-based processing-in-memory works leverage the high density of 3D NAND flash for deep neural networks, they primarily focus on optimizing computation inside the NAND array. Consequently, these approaches often fail to address the critical latency overhead associated with managing intermediate activation data. To overcome such a limitation, we propose a heterogeneous NAND flash-based ISC architecture with enhanced activation buffering. By buffering intermediate values in a DRAM subsystem rather than programming them into the NAND flash array, our approach effectively mitigates the high programming latency penalties. We also introduce a distributed dataflow scheme that maximizes computational parallelism through optimized plane- and bank-level data mapping. The results show that our proposed architecture achieves performance improvements, reducing inference latency by up to 86% compared to the baseline. Sanghun Shin, Gisan Ji, Sungju Ryu |
DATE | 2 |
| 2026 | E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference ComputingabstractTransformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead. Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | RADiT: Redundancy-Aware Diffusion Transformer Acceleration Leveraging Timestep SimilarityabstractDiffusion Transformers (DiTs) have demonstrated unprecedented performance across various generative tasks including image and video generation. However, a large amount of computations on the inference process and iterative sampling steps in the DiT models result in high computational costs, leading to substantial latency and energy consumption challenges. To address these issues, we propose a redundancy-aware DiT (RADiT), a novel software-hardware co-optimization accelerator for DiTs that minimizes redundant operations in the iterative sampling stages. We identify data redundancy by evaluating blockwise input features and skip redundant computations by reusing results from consecutive timesteps. Furthermore, to minimize accuracy degradation and maximize computational efficiency, the Dynamic Threshold Scaling Module (DTSM) and Compress and Compare Unit (CCU) are employed in the redundancy detection process. This approach enables DiTs to achieve up to $1.8 \times$ and $1.7 \times$ faster speeds for image and video generation, respectively, without compromising quality, along with 41% and 45.5% reductions in energy consumption. Our RADiT scheme improves throughput by $1.67 \times$ and $1.76 \times$ for image and video generation tasks, respectively, while maintaining output quality and significantly reducing energy consumption. Youngjun Park, Yeonggeon Kim, Gisan Ji, Sungju Ryu |
DAC | 4 |
| 2025 | OptiRange: An Efficient ReRAM-Based PIM Accelerator with ADC Resolution OptimizationabstractProcessing-in-memory (PIM) techniques have attracted significant attention in computer architecture research due to their ability to mitigate the memory wall bottleneck for deep neural network (DNN) applications. In-memory computing (IMC) using ReRAM crossbar arrays is promising for energy efficiency but suffers from high-cost peripheral circuits, particularly analog-to-digital converters (ADCs). There are existing methods to reduce ADC energy consumption, such as lowering resolution or sharing an ADC among multiple columns. Unfortunately, these techniques reduce throughput. In this work, we propose OptiRange for PIM accelerators that enhances both energy efficiency and throughput while maintaining accuracy and model flexibility. OptiRange proposes a split-the-burden (STB) algorithm that manipulates crossbar cell values without changing the weights. STB reduces large cell values, shifts the column-sum distribution center closer to zero, and minimizes the number of low-resistance state (LRS) cells, significantly lowering the hardware-level fixed ADC resolution and improving accuracy. Furthermore, OptiRange utilizes dynamic ADC range (DAR) adaptation, determining the optimal ADC operating range per column at the software level and enabling the skipping of redundant ADC conversion cycles at runtime. An associated ADC range-based grouping (AG) strategy leverages the resulting dynamic latencies to manage synchronization and further boost system throughput. By combining weight-preserving cell manipulation with adaptive ADC operation and synchronization, OptiRange dynamically optimizes the ADC workload. Experimental results indicate that our method significantly enhances overall performance compared to conventional ReRAM-based accelerators. Sangkyu Jeon, Gisan Ji, Yeonggeon Kim, Youngjun Park, Sungju Ryu |
ICCAD | 2 |
| 2025 | E-Flash: Energy-Efficient DNN Mapping on NAND Flash Memory with State-Switching AlgorithmabstractDeep neural network (DNN) has been widely adopted in various applications. Ranging from image classification to text generation, Transformer-based models have demonstrated unprecedented performance. However, they suffer from a significant computational complexity and a large memory footprint, leading to memory-bound issues. While previous research on NAND flash-based neural network computation has been performed, these studies often encounter accuracy problems, as analog processing-in memory (PIM) operations typically lead to inaccurate results. Moreover, studies on NAND flash using single-level cell (SLC) are unable to fully leverage the advantages of efficient storage density on the multi-level cell (MLC) memory. We propose an E-Flash hardware architecture with an energy-efficient DNN mapping method. E-Flash introduces a state-switching algorithm to perform data movements between flash memory and host device in an energy-efficient manner. By reallocating the data in triple-level cell (TLC) NAND flash memory, we reduce the energy consumption during the data read operation. Experimental results demonstrate that E-Flash achieves improved energy consumption compared to baseline under significantly small area overhead for quantized BERT and Llama 2 weights by 37.73% and 16.74%, respectively. Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu |
ISLPED | 1 |