EDBT 2026 Demo / reviewers in the wild / expert
Myoungjun Chun
dblp:228/0270
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-8188-4324ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STRAW: Stress-Aware WL-Based Read Disturbance Management for High-Density NAND Flash MemoryabstractWhile NAND flash memory has continuously increased its storage density over decades, this progress has exacerbated the read-disturbance problem. In this work, we identify two fundamental limitations of existing read-disturbance management techniques that trigger read reclaim (RR) at the block granularity: (i) they overlook the heterogeneous reliability impact of read disturbance across individual wordlines (WLs), leading to unnecessary RR in many cases; and (ii) they address read disturbance only after disturbance-induced errors have already accumulated, which forces substantial RR-induced copy overheads in read-disturbance-prone modern NAND flash memory. To address these limitations, we propose STRAW (STRess-Aware Wordline-based read-disturbance management), a new technique that minimizes RR overheads through two key ideas: (i) stress-aware WL-based read reclaim, which monitors the accumulated read-disturbance effect on each WL and reclaims only heavily disturbed WLs, and (ii) stress-reduced read, which mitigates disturbance on valid WLs during each read operation by scaling pass-through voltages based on WL validity. Our experimental results using a modern SSD emulator show that STRAW reduces RR-induced page-copy overhead by 88.6% on average compared with the state-of-the-art technique. Myoungjun Chun, Jae Yong Lee 0004, Inhyuk Choi, Jisung Park 0001, Myungsuk Kim, Jihong Kim 0001 |
ASPLOS (2) | 1 |
| 2026 | Processing-in-Memory Architecture for In-Storage Information Retrieval on 3-D NAND Flash Solid-State DrivesabstractRecent advances in retrieval-augmented generation (RAG) highlight retrieval latency, dominated by storage access, and large embedding-index size as the major bottlenecks. Although processing-in-memory (PIM) based on 3-Dnandflash and binary passage retriever (BPR) algorithm offer potential solutions for reducing latency and memory usage, naively integrating BPR algorithm into an SSD based on prior in-memory search (IMS) architectures results in negligible performance gain with substantial hardware overhead. This stems from the fact that prior IMS architectures are optimized only for search, leaving candidate readout misaligned with thenandpage direction and requiring excessive readcycles, while BPR's original top-$L$selection further requires high-precision ADCs and sorting logic. To address these limitations, we propose an algorithm–hardware co-designed in-storage information retrieval architecture. We introduce a hardware-efficient optimized BPR algorithm with a threshold-based selection and two-stage in-storage processing approach to minimize data transfer overheads. We further present an IMS architecture with page-aligned encoding that efficiently supports both search and candidate read operations, eliminating the read cycle penalty. We also present a lightweight in situ error monitoring scheme that bypasses error-correcting code (ECC) decoding while preserving error-aware refresh triggering. Experiments show that our design reduces the amount of data read fromnandflash by 7.94–$1330.40\times $compared with host-side processing baselines and further reduces controller-to-host data transfer, while maintaining the retrieval accuracy. It also reduces total operationcycles by up to 96.93% compared with prior IMS architectures, and lowers error-detection power by 94.71% and total read energy by 17.22% compared with a conventional ECC decoder for error monitoring. Huiwon Yun, Inho Jeong, Kyoyun Lee, Jae Yong Lee 0004, Myoungjun Chun, Jihong Kim 0001, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | AiF: Accelerating On-Device LLM Inference Using In-Flash ProcessingabstractWhile large language models (LLMs) achieve remarkable performance across diverse application domains, their substantial memory demands present challenges, especially on personal devices with limited DRAM capacity.Recent LLM inference engines have introduced SSD offloading for model parameters to reduce memory footprint.However, the highly memory-bound nature of on-device LLMs makes inference speed heavily dependent on read bandwidth, leading to significant performance degradation due to the limited bandwidth of SSDs.In this paper, we propose an in-flash processing solution for on-device LLM, called Accelerator-in-Flash (AiF), which integrates matrix-vector multiplication (GEMV) operations directly into flash chips.By enabling in-flash GEMV operations, AiF leverages the high internal bandwidth of flash chips without being constrained by the limited external bandwidth.Building on this core structure, AiF employs two novel flash read techniques that were specifically optimized for reading LLM parameters stored in flash memory.AiF achieves a 4x boost in internal read bandwidth during inference with minimal implementation overhead, thanks to its streamlined error correction process.Evaluations on eight real-world LLMs reveal that AiF provides a 14.6x throughput improvement compared to baseline SSD offloading schemes.Furthermore, AiF surpasses in-memory inference, delivering 1.4x higher throughput with a significantly reduced memory footprint. Jae Yong Lee 0004, Hyeunjoo Kim, Sanghun Oh, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001 |
ISCA | 4 |
| 2025 | DEAR: Improving Performance and Lifetime of SSDs Using Dynamic Error-Aware Refresh
Jae Yong Lee 0004, Myoungjun Chun, Myungsuk Kim, Jihong Kim 0001 |
MICRO | 3 |
| 2024 | RiF: Improving Read Performance of Modern SSDs Using an On-Die Early-Retry EngineabstractModern high-performance SSDs have multiple flash channels operating in parallel to achieve their high I/O bandwidth. However, when the effective bandwidth of these flash channels declines, the SSD's overall bandwidth is substantially impacted. In contemporary SSDs featuring high-density 3D NAND flash memory, frequent invocations of a read-retry procedure pose a significant challenge to fully utilizing the maximum I/O bandwidth of a flash channel. In this paper, we propose a novel read-retry optimization scheme, Retry-in-Flash (RiF), which proactively minimizes the amount of time wasted in conventional read-retry procedures. Unlike existing read-retry solutions that focus on identifying an optimal read-reference voltage for a sensed page, the RiF scheme focuses on determining early on whether a read-retry will be required for the sensed data. To know if a read-retry is needed or not at the earliest possible time, we propose a RiF-enabled flash chip with an on-die early-retry (ODEAR) engine. When the ODEAR engine determines that a sensed page requires a read-retry, a readreference voltage is immediately adjusted and the same page is re-read while ignoring the previously sensed page. By performing the key steps of a read-retry procedure inside a RiF flash chip without transferring the sensed uncorrectable page to an offchip controller, the RiF scheme prevents the read bandwidth of a flash channel from being wasted due to failed read data. To evaluate the RiF scheme, we developed a prototype RiF-enabled flash chip and constructed a RiF-aware SSD simulator using RiF flash chips. Our evaluation results show that the proposed RiF scheme improves the effective SSD bandwidth by 72.1% on average over a state-of-the-art read-retry solution at 2K P/E cycles with negligible power and area overheads. Myoungjun Chun, Jae Yong Lee 0004, Myungsuk Kim, Jisung Park 0001, Jihong Kim 0001 |
HPCA | 1 |
| 2024 | ReadGuard: Integrated SSD Management for Priority-Aware Read Performance DifferentiationabstractWhen multiple apps with different I/O priorities share a high-performance SSD, it is important to differentiate the I/O QoS level based on the I/O priority of each app. In this paper, we study how a modern flash-based SSD should be designed to support priority-aware read performance differentiation. From an in-depth evaluation study using 3D TLC SSDs, we observed that existing FTLs have several weaknesses that need to be improved for better read performance differentiation. In order to overcome the existing FTL weaknesses, we propose ReadGuard , a novel priority-aware SSD management technique that enables an FTL to manage its blocks in a fully read-latency-aware fashion. ReadGuard leverages a new read-latency-centric block quality marker that can accurately distinguish the read latency of a block and ensures that higher-quality blocks are used for higher-priority apps. ReadGuard extends an existing suspend/resume technique to handle collisions among reads. Our experimental results show that a ReadGuard -enabled SSD is effective in supporting differentiated read performance in modern 3D flash SSDs. Myoungjun Chun, Myungsuk Kim, Dusol Lee, Jisung Park 0001, Jihong Kim 0001 |
ACM Trans. Storage | 1 |
| 2022 | PiF: in-flash acceleration for data-intensive applicationsabstractTo minimize unnecessary data movements from storage to a host, processing-in-storage (PiS) techniques, which move a compute unit to storage, have been proposed. In this position paper, we propose an extreme version of PiS solutions, called a processing-in-flash (PiF) scheme, that moves computation inside flash chips where data are physically present. As a key building block of a PiF solution, we present a novel flash chip architecture, CoX. Using a prototype PiF SSD based on CoX chips, we demonstrate that PiF-based SSDs are promising in accelerating data-intensive applications. Myoungjun Chun, Jae Yong Lee 0004, Sanggu Lee, Myungsuk Kim, Jihong Kim 0001 |
HotStorage | 1 |
| 2021 | Reducing solid-state drive read latency by optimizing read-retryabstract3D NAND flash memory with advanced multi-level cell techniques provides high storage density, but suffers from significant performance degradation due to a large number of read-retry operations. Although the read-retry mechanism is essential to ensuring the reliability of modern NAND flash memory, it can significantly in-crease the read latency of an SSD by introducing multiple retry steps that read the target page again with adjusted read-reference voltage values. Through a detailed analysis of the read mechanism and rigorous characterization of 160 real 3D NAND flash memory chips, we find new opportunities to reduce the read-retry latency by exploiting two advanced features widely adopted in modern NAND flash-based SSDs: 1) the CACHE READ command and 2) strong ECC engine. First, we can reduce the read-retry latency using the advanced CACHE READ command that allows a NAND flash chip to perform consecutive reads in a pipelined manner. Second, there exists a large ECC-capability margin in the final retry step that can be used for reducing the chip-level read latency. Based on our new findings, we develop two new techniques that effectively reduce the read-retry latency: 1) Pipelined Read-Retry (PR²) and 2) Adaptive Read-Retry (AR²). PR² reduces the latency of a read-retry operation by pipelining consecutive retry steps using the CACHE READ command. AR² shortens the latency of each retry step by dynamically reducing the chip-level read latency depending on the current operating conditions that determine the ECC-capability margin. Our evaluation using twelve real-world workloads shows that our proposal improves SSD response time by up to 31.5% (17% on average)over a state-of-the-art baseline with only small changes to the SSD controller. Jisung Park 0001, Myungsuk Kim, Myoungjun Chun, Lois Orosa 0001, Jihong Kim 0001, Onur Mutlu |
ASPLOS | 3 |
| 2021 | RealWear: Improving performance and lifetime of SSDs using a NAND aging marker
Myungsuk Kim, Myoungjun Chun, Duwon Hong, Yoona Kim, Geonhee Cho, Dusol Lee, Jihong Kim 0001 |
Perform. Evaluation | 2 |
| 2021 | Reparo: A Fast RAID Recovery Scheme for Ultra-large SSDsabstractA recent ultra-large SSD (e.g., a 32-TB SSD) provides many benefits in building cost-efficient enterprise storage systems. Owing to its large capacity, however, when such SSDs fail in a RAID storage system, a long rebuild overhead is inevitable for RAID reconstruction that requires a huge amount of data copies among SSDs. Motivated by modern SSD failure characteristics, we propose a new recovery scheme, called reparo , for a RAID storage system with ultra-large SSDs. Unlike existing RAID recovery schemes, reparo repairs a failed SSD at the NAND die granularity without replacing it with a new SSD, thus avoiding most of the inter-SSD data copies during a RAID recovery step. When a NAND die of an SSD fails, reparo exploits a multi-core processor of the SSD controller in identifying failed LBAs from the failed NAND die and recovering data from the failed LBAs. Furthermore, reparo ensures no negative post-recovery impact on the performance and lifetime of the repaired SSD. Experimental results using 32-TB enterprise SSDs show that reparo can recover from a NAND die failure about 57 times faster than the existing rebuild method while little degradation on the SSD performance and lifetime is observed after recovery. Duwon Hong, Keonsoo Ha, Minseok Ko, Myoungjun Chun, Yoona Kim, Sungjin Lee 0001, Jihong Kim 0001 |
ACM Trans. Storage | 4 |
| 2019 | Fully Automatic Stream Management for Multi-Streamed SSDs Using Program Contexts
Duwon Hong, Sangwook Shane Hahn, Myoungjun Chun, Sungjin Lee 0001, Joo Young Hwang, Jongyoul Lee, Jihong Kim 0001 |
FAST | 4 |
| 2019 | Exploiting Process Similarity of 3D Flash Memory for High Performance SSDsabstract3D NAND flash memory exhibits two contrasting process characteristics from its manufacturing process. While process variability between different horizontal layers are well known, little has been systematically investigated about strong process similarity (PS) within the horizontal layer. In this paper, based on an extensive characterization study using real 3D flash chips, we show that 3D NAND flash memory possesses very strong process similarity within a 3D flash block: the word lines (WLs) on the same horizontal layer of the 3D flash block exhibit virtually equivalent reliability characteristics. This strong process similarity, which was not previously utilized, opens simple but effective new optimization opportunities for 3D flash memory. In this paper, we focus on exploiting the process similarity for improving the I/O latency. By carefully reusing various flash operating parameters monitored from accessing the leading WL, the remaining WLs on the same horizontal layer can be quickly accessed, avoiding unnecessary redundant steps for subsequent program and read operations. We also propose a new program sequence, called mixed order scheme (MOS), for 3D NAND flash memory which can further reduce the program latency. We have implemented a PS-aware FTL, called cubeFTL, which takes advantage of the proposed techniques. Our evaluation results show that cubeFTL can improve the IOPS by up to 48% over an existing PS-unaware FTL. Youngseop Shim, Myungsuk Kim, Myoungjun Chun, Jisung Park 0001, Yoona Kim, Jihong Kim 0001 |
MICRO | 3 |
| 2018 | FlashShare: Punching Through Server Storage Stack from Kernel to Firmware for Ultra-Low Latency SSDs
Jie Zhang 0048, Miryeong Kwon, Donghyun Gouk, Sungjoon Koh, Changlim Lee, Mohammad Alian, Myoungjun Chun, Mahmut T. Kandemir, Nam Sung Kim, Jihong Kim 0001, Myoungsoo Jung |
OSDI | 7 |