Jae Yong Lee 0004

dblp:72/6418-4 · also Jaeyong Lee 0004 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-2724-8888ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 STRAW: Stress-Aware WL-Based Read Disturbance Management for High-Density NAND Flash Memory
abstract
While NAND flash memory has continuously increased its storage density over decades, this progress has exacerbated the read-disturbance problem. In this work, we identify two fundamental limitations of existing read-disturbance management techniques that trigger read reclaim (RR) at the block granularity: (i) they overlook the heterogeneous reliability impact of read disturbance across individual wordlines (WLs), leading to unnecessary RR in many cases; and (ii) they address read disturbance only after disturbance-induced errors have already accumulated, which forces substantial RR-induced copy overheads in read-disturbance-prone modern NAND flash memory. To address these limitations, we propose STRAW (STRess-Aware Wordline-based read-disturbance management), a new technique that minimizes RR overheads through two key ideas: (i) stress-aware WL-based read reclaim, which monitors the accumulated read-disturbance effect on each WL and reclaims only heavily disturbed WLs, and (ii) stress-reduced read, which mitigates disturbance on valid WLs during each read operation by scaling pass-through voltages based on WL validity. Our experimental results using a modern SSD emulator show that STRAW reduces RR-induced page-copy overhead by 88.6% on average compared with the state-of-the-art technique.
Myoungjun Chun, Jae Yong Lee 0004, Inhyuk Choi, Jisung Park 0001, Myungsuk Kim, Jihong Kim 0001
ASPLOS (2)2
2026 Processing-in-Memory Architecture for In-Storage Information Retrieval on 3-D NAND Flash Solid-State Drives
abstract
Recent advances in retrieval-augmented generation (RAG) highlight retrieval latency, dominated by storage access, and large embedding-index size as the major bottlenecks. Although processing-in-memory (PIM) based on 3-Dnandflash and binary passage retriever (BPR) algorithm offer potential solutions for reducing latency and memory usage, naively integrating BPR algorithm into an SSD based on prior in-memory search (IMS) architectures results in negligible performance gain with substantial hardware overhead. This stems from the fact that prior IMS architectures are optimized only for search, leaving candidate readout misaligned with thenandpage direction and requiring excessive readcycles, while BPR's original top-$L$selection further requires high-precision ADCs and sorting logic. To address these limitations, we propose an algorithm–hardware co-designed in-storage information retrieval architecture. We introduce a hardware-efficient optimized BPR algorithm with a threshold-based selection and two-stage in-storage processing approach to minimize data transfer overheads. We further present an IMS architecture with page-aligned encoding that efficiently supports both search and candidate read operations, eliminating the read cycle penalty. We also present a lightweight in situ error monitoring scheme that bypasses error-correcting code (ECC) decoding while preserving error-aware refresh triggering. Experiments show that our design reduces the amount of data read fromnandflash by 7.94–$1330.40\times $compared with host-side processing baselines and further reduces controller-to-host data transfer, while maintaining the retrieval accuracy. It also reduces total operationcycles by up to 96.93% compared with prior IMS architectures, and lowers error-detection power by 94.71% and total read energy by 17.22% compared with a conventional ECC decoder for error monitoring.
Huiwon Yun, Inho Jeong, Kyoyun Lee, Jae Yong Lee 0004, Myoungjun Chun, Jihong Kim 0001, Dongsuk Jeon
IEEE Trans. Very Large Scale Integr. Syst.4
2025 AiF: Accelerating On-Device LLM Inference Using In-Flash Processing
abstract
While large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint.
Jae Yong Lee 0004, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001
ISCA1
2025 DEAR: Improving Performance and Lifetime of SSDs Using Dynamic Error-Aware Refresh
Jae Yong Lee 0004, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001
MICRO1
2024 RiF: Improving Read Performance of Modern SSDs Using an On-Die Early-Retry Engine
abstract
Modern high-performance SSDs have multiple flash channels operating in parallel to achieve their high I/O bandwidth. However, when the effective bandwidth of these flash channels declines, the SSD's overall bandwidth is substantially impacted. In contemporary SSDs featuring high-density 3D NAND flash memory, frequent invocations of a read-retry procedure pose a significant challenge to fully utilizing the maximum I/O bandwidth of a flash channel. In this paper, we propose a novel read-retry optimization scheme, Retry-in-Flash (RiF), which proactively minimizes the amount of time wasted in conventional read-retry procedures. Unlike existing read-retry solutions that focus on identifying an optimal read-reference voltage for a sensed page, the RiF scheme focuses on determining early on whether a read-retry will be required for the sensed data. To know if a read-retry is needed or not at the earliest possible time, we propose a RiF-enabled flash chip with an on-die early-retry (ODEAR) engine. When the ODEAR engine determines that a sensed page requires a read-retry, a readreference voltage is immediately adjusted and the same page is re-read while ignoring the previously sensed page. By performing the key steps of a read-retry procedure inside a RiF flash chip without transferring the sensed uncorrectable page to an offchip controller, the RiF scheme prevents the read bandwidth of a flash channel from being wasted due to failed read data. To evaluate the RiF scheme, we developed a prototype RiF-enabled flash chip and constructed a RiF-aware SSD simulator using RiF flash chips. Our evaluation results show that the proposed RiF scheme improves the effective SSD bandwidth by 72.1% on average over a state-of-the-art read-retry solution at 2K P/E cycles with negligible power and area overheads.
Myoungjun Chun, Jae Yong Lee 0004, Myungsuk Kim, Jisung Park 0001, Jihong Kim 0001
HPCA2
2022 TailCut: improving performance and lifetime of SSDs using pattern-aware state encoding
abstract
Although lateral charge spreading is considered as a dominant error source in 3D NAND flash memory, little is known about its detailed characteristics at the storage system level. From a device characterization study, we observed that lateral charge spreading strongly depends on vertically adjacent state patterns and a few specific patterns are responsible for a large portion of bit errors from lateral charge spreading. We propose a new state encoding scheme, called TailCut, which removes vulnerable state patterns by modifying encoded states. By removing vulnerable patterns, TailCut can improve the SSD lifetime and read latency by 80% and 25%, respectively.
Jae Yong Lee 0004, Myungsuk Kim, Wonil Choi, Sanggu Lee, Jihong Kim 0001
DAC1
2022 PiF: in-flash acceleration for data-intensive applications
abstract
To minimize unnecessary data movements from storage to a host, processing-in-storage (PiS) techniques, which move a compute unit to storage, have been proposed. In this position paper, we propose an extreme version of PiS solutions, called a processing-in-flash (PiF) scheme, that moves computation inside flash chips where data are physically present. As a key building block of a PiF solution, we present a novel flash chip architecture, CoX. Using a prototype PiF SSD based on CoX chips, we demonstrate that PiF-based SSDs are promising in accelerating data-intensive applications.
Myoungjun Chun, Jae Yong Lee 0004, Sanggu Lee, Myungsuk Kim, Jihong Kim 0001
HotStorage2