Jinhua Cui 0001

dblp:156/2428-1 · DBLP profile ↗
← Back
22ranked-venue papers
16as first author
12since 2021 · last 2026
0000-0002-3252-7937ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 15 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author
YearPublicationVenuePosition
2026 WOM-FTL: An Efficient FTL for High-Density Flash Memory Through WOM-v Codes
abstract
High-density NAND flash memory, such as quadruple-level cell (QLC) flash, has has been widely adopted in emerging storage systems. However, its limited endurance and performance challenges necessitate novel solutions. Voltage-based write-once memory (WOM-v) codes have demonstrated their effectiveness in extending flash memory lifespan by reducing the erase count of flash blocks. Concurrently, secure deletion is essential to ensure data privacy in flash-based storage systems. Existing secure deletion approaches—encryption-based, erasure-based, and scrubbing-based—often face limitations such as susceptibility to deciphering or significant performance overheads. Additionally, the inherent “big block problem” in high-density flash memory complicates garbage collection (GC), further degrading system performance. To address these challenges, this paper proposes WOM-FTL, a flash translation layer (FTL) that integrates secure deletion and garbage collection (GC) with WOM-v codes to enhance both security and performance. WOM-FTL classifies request data into four categories based on access frequency and privacy requirements: hot-secure (HS), cold-secure (CS), hot-unsecure (HU) and cold-unsecure (CU). Additionally, WOM-FTL divides each block into several equal-sized sub-blocks and further classifies them into top sub-blocks and bottom sub-blocks according to their data storage characteristics. When a secure data deletion command is issued to the storage device, WOM-FTL leverages unsecure data (HU and CU) to overwrite the secure data (HS and CS). Furthermore, WOM-FTL allocates different types of request data to the corresponding sub-blocks, creating a data allocation pattern that is both scrubbing-friendly and GC-friendly. Experimental results demonstrate that WOM-FTL improves the I/O performance of storage systems by 60.91% compared to state-of-the-art solutions, providing a significant advancement in secure and efficient management of high-density flash memory.
Jinhua Cui 0001, Canghao Wen, Shiqiang Nie, Debin Liu, Yaliang Zhao, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 SecureIMR: A Plausibly Deniable Storage System for Interlaced Magnetic Recording
Jinhua Cui 0001, Canghao Wen, Shiqiang Nie
ICA3PP (1)1
2025 Z-STREAM: A Transparent Reordering Engine for High-Performance Zoned Namespace SSDs
Canghao Wen, Jinhua Cui 0001, Xuanxuan Fu, Laurence T. Yang
ICA3PP (1)2
2024 SmartNetSSD: Exploiting Path Resources for Read Performance Improvement in Network-Based SSDs
abstract
With the bit density improvement and the three-dimensional NAND flash techniques, solid-state drives (SSDs) dramatically increase the storage capacity and performance. However, the incorporation of multiple flash chips within a single flash channel structure in SSDs introduces the access path conflicts when servicing multiple I/O requests accessing flash chips on the same channel. To meet the increasing performance demands of modern applications, network SSDs that employ the interconnection network of flash chips, has a high potential to address the access path conflicts by fundamentally increasing the number of the available access paths. In this paper, we propose SmartNetSSD, a new scheme that utilizes the path diversity in network SSD to dramatically mitigate access path conflicts. SmartNetSSD employs the following key techniques: 1) A multi-path routing algorithm, MProuting, identifies the multiple access paths for read operations. 2) A network-based read-retry technique, NetRR, pipelines the con-secutive read-retry steps for a read operation across the multiple access paths. The experimental results show that SmartNetSSD improves I/O performance by an average of 46.29% over the state-of-the-art approach.
Jinhua Cui 0001, Shiqiang Nie, Laurence T. Yang
ICCD1
2024 Optimizing Secure Deletion in Interlaced Magnetic Recording With Move-on-Cover Approach
abstract
Secure deletion plays a crucial role in safeguarding privacy and maintaining confidentiality. Historically, secure deletion was achieved on magnetic storage devices by data overwriting. However, with the rapid advancements in magnetic storage technology, existing secure deletion strategies are still oblivious to the unique characteristics of these storage devices. As a result, there is currently a lack of research effectively addressing the implementation of secure deletion for the emerging interlaced magnetic recording (IMR) technology. Regrettably, applying traditional read-modify-write (RMW) processes to enable secure deletion on IMR technology leads to significant challenges such as severe write amplification (WA) and performance degradation. In this paper, we propose move-on-cover (MOC), a novel secure deletion strategy for securely deleting on IMR, to optimize the performance of secure deletion. Specifically, when securely deleting a bottom track, the MOC first moves one of affected top track to a free top track, and then covers the to-be-deleted bottom track with another top track, minimizing redundant data backups and hiding secure deletion latency of bottom tracks. A series of evaluation experiments show that MOC is highly effective in reducing extra I/Os and improving the performance of secure deletion.
Zhimin Zeng, Jinhua Cui 0001, Lizhao Wan, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 A Fast Secure Deletion Strategy for High-Density Flash Memory through WOM-v Codes
abstract
High bit-density NAND flash memory, such as quadruple-level cell (QLC) flash, has been widely adopted in emerging storage systems. As voltage-based write-once-memory (WOM-v) codes reduce the erase count of a flash block before storing new data, WOM-v codes are proven feasible to extend the lifetime of high-density flash memory significantly. To ensure data security, secure deletion is widely employed in flash-based storage system.In this paper, we propose FSD, a fast secure deletion strategy for high-density flash memory that cooperates secure deletion with WOM-v codes. FSD classifies request data into secure, and unsecure data by considering data privacy. When secure data deletion command issues to the storage device, FSD cooperates unsecure data program operation to cover secure data. The results show that FSD improves the I/O performance of storage system by 70.37% over the state-of-the-art.
Jinhua Cui 0001, Laurence T. Yang
DAC1
2023 Improving 3-D NAND SSD Read Performance by Parallelizing Read-Retry
abstract
With the bit density improvement and the 3-D flash techniques, modern NAND flash-memory-based solid-state disks (SSDs) dramatically increase the flash storage capacity. However, in high-density SSDs, the long read latency overheads due to massive read-retry steps have become a serious performance concern to develop flash memory in storage devices. In this article, we proposed a parallel read-retry scheme (PaRR) to utilize the read-retry characteristics among flash memory cells. Specifically, when reading the multiple data pages simultaneously, PaRR reduces the read-response time by parallelizing read-retry operations. PaRR reduces the number of read-retry operations by parallelizing read-retry steps with the different aggressive read reference voltages if simultaneously reading the data pages that exhibit virtually equivalent reliability characteristic. Our evaluation shows that PaRR improves the I/O performance by 33.77% on average compared with the state-of-the-art scheme.
Jinhua Cui 0001, Zhimin Zeng, Jianhang Huang, Weiqi Yuan, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A Space-Efficient Fair Cache Scheme Based on Machine Learning for NVMe SSDs
abstract
Non-volatile memory express (NVMe) solid-state drives (SSDs) have been widely adopted in multi-tenant cloud computing environments or multi-programming systems. The on-board DRAM cache inside NVMe SSDs can efficiently reduce the disk accesses and extend the lifetime of SSDs. Current SSD cache management research either improves cache hit ratio while ignoring fairness, or improves fairness while sacrificing overall performance. In this paper, we present MLCache, a space-efficient shared cache management scheme for NVMe SSDs. By learning the impact of reuse distance on cache allocation, a workload-generic neural network model is built. At runtime, MLCache continuously monitors the reuse distance distribution for the neural network module to obtain space-efficient allocation decisions. MLCache also proposes an efficient parallel writing back strategy based on hit ratio and response time, to improve fairness. Experimental results show MLCache improves the write hit ratio when compared to baseline, and MLCache strongly safeguards the fairness of SSDs with parallel write-back and maintains a low level of degradation.
Jinhua Cui 0001, Laurence T. Yang
IEEE Trans. Parallel Distributed Syst.2
2022 IMRSim: A Disk Simulator for Interlaced Magnetic Recording Technology
Zhimin Zeng, Laurence T. Yang, Jinhua Cui 0001
NPC4
2022 ADS: Leveraging Approximate Data for Efficient Data Sanitization in SSDs
abstract
NAND flash memory has been widely adopted in emerging storage systems. To ensure data security, the support of data sanitization in NAND flash memory-based storage systems is widely employed. Although some existing studies made efforts in employing encryption-based, erasure-based, and scrubbing-based secure deletion approaches to achieve the security requirement, they suffer from the risk of being deciphered, the severe performance, and the scrubbing disturbance problems. Meanwhile, 3-D NAND flash technology, which stacks flash cells in vertical direction, is gaining traction in the modern systems. This made the problems more severe because of the increased number of scrubbing disturbance directions in 3-D NAND flash memory. To address the above issue, this work proposes an approximate-data-aware data sanitization scheme (ADS) with the assistance of the error-resilient data of modern applications, which guarantees the highest degree of security for security-sensitive data sanitization (i.e., storage systems do not keep any old version of secure data once secure data are updated). ADS classifies request data into approx-secure (AS), precise-secure (PS), approx-unsecure (AU), and precise-unsecure (PU) data by considering two factors, including data error resilience and data privacy. Then, a novel data allocation strategy is proposed to selectively interleave secure data and approximate data within the flash blocks, which creates the scrubbing friendly data patterns to minimize the overhead of secure deletion. Our experimental results show that ADS reduces the average secure deletion latency by 58.93% over the state of the art.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Exploiting Uncorrectable Data Reuse for Performance Improvement of Flash Memory
abstract
With the increased density and the technology scaling, flash memory is more vulnerable to noise effects, overwhelmingly lowering the data retention time. The periodic data refresh technique is commonly used to retain the long-term data integrity. However, the frequent refresh requests introduce the increased access conflict and severe write amplification, leading to suboptimal performance and lifetime improvement of flash memory-based storage systems. In this article, we propose ApproxRefresh, which enables the uncorrectable data reuse with the assistance of approximate-computing applications to reduce the refresh costs. Specifically, we design a lightweight remapping-free refresh technique, called RFR, to periodically correct, compute an enhanced ECC, and only remap new ECC parity bits, which dramatically reduces the refresh costs. Then, with approximate-read and precise-read hotness awareness, ApproxRefresh selectively adopts RFR or the traditional remapping-based refresh technique to reduce the refresh costs and the read disturbance. Besides, ApproxRefresh further improves flash access performance based on data hotness and process variations. In addition, a corresponding ApproxRefresh-aware garbage collection algorithm is proposed to complete the design. Evaluations show that ApproxRefresh reduces the refresh latency by 36.19% over the state of the art.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 SmartHeating: On the Performance and Lifetime Improvement of Self-Healing SSDs
abstract
In NAND flash memory-based solid-state drives (SSDs), during the idle time between the consecutive program/erase cycles (dwell time), the dielectric damage of flash cell can be partially repaired, also known as the self-recovery effect. As the effectiveness of the self-recovery effect can be improved under high temperature, self-healing SSDs are proven feasible to extend the flash endurance significantly. However, current self-healing SSDs perform the heating operations on all the worn-out blocks without considering the data retention requirement, and measures the lifetime of flash memory based on the worst-case self-recovery effect, leading to some unnecessary heating operations and the degraded performance. We propose SmartHeating, a smart heating scheme that exploits the dwell time variation and the write hotness variation to improve the I/O performance and the lifetime of self-healing SSDs. SmartHeating tracks the dwell time of all worn-out flash blocks, predicts their self-recovery effect and reliability, and avoids performing heating operations on the worn-out flash blocks that still have strong flash reliability. In addition, by exploiting the data hotness variation, SmartHeating only heats the worn-out flash blocks that store write-cold data, while allocating write-hot data to a small portion of worn-out flash blocks with negligible refresh overhead. The experimental results show that SmartHeating reduces the number of heating operations by 12.5% on average, boosts I/O performance of flash storage systems by 21.0%, and improves the lifetime of flash memory by $1.20\times $ compared with conventional heating scheme.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Exploiting Disturbance-Aware Read Redirection for Performance Improvement in 3D Flash Memory
abstract
Three-dimensional (3D) NAND flash memory has been widely adopted in emerging storage systems, ranging from mobile devices to cloud systems, due to its performance benefits and scalability. While 3D flash improves capacity through stacking more cells in the vertical direction, the read disturbance has become a serious reliability concern to develop 3D flash memory in storage devices, especially for read-intensive applications. To address the above challenge, we propose a novel disturbance-aware read redirection scheme (DRR). By exploiting parity-based redundancy and data duplication to regenerate data content, DRR redirects reads from strong disturbed flash blocks to weak ones without any data replication. DRR makes the redirection decision based on the accumulative read disturbance, data redundancy, data duplication, and the current I/O traffic. Results show that DRR provides as much as 24.6% performance improvement in the average read/write response time, and extends the endurance of 3D flash devices by 4.7%.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
ACM Great Lakes Symposium on VLSI1
2020 MLCache: A Space-Efficient Cache Scheme based on Reuse Distance and Machine Learning for NVMe SSDs
abstract
Non-volatile memory express (NVMe) solid-state drives (SSDs) have been widely adopted in emerging storage systems, which can provide multiple I/O queues and high-speed bus to maximize high data transfer rate. NVMe SSD use streams (also called "Multi-Queue") to store related data in associated locations or for other performance enhancements. The on-board DRAM cache inside NVMe SSDs can efficiently reduce the disk accesses and extend the lifetime of SSDs, thus improving the overall efficiency of the storage systems. However, in previous studies, such SSD cache has been only used as a shared cache for all streams or a statically partitioned cache for each stream, which may seriously degrade the performance-per-stream and underutilize the valuable cache resources.
Jinhua Cui 0001, Laurence T. Yang
ICCAD2
2020 ApproxRefresh: Enabling Uncorrectable Data Reuse on Flash Memory with Approximate Read Awareness
abstract
With the increased density and the technology scaling, flash memory is more vulnerable to noise effects, overwhelmingly lowering the data retention time. The periodic data refresh technique is commonly used to retain the long-term data integrity. However, the frequent refresh requests also introduce the increased access conflict and severe write amplification, leading to sub-optimal performance and lifetime improvement.
Jinhua Cui 0001, Jianhang Huang, Laurence T. Yang
LCTES1
2020 Leveraging partial-refresh for performance and lifetime improvement of 3D NAND flash memory in cyber-physical systems
Jinhua Cui 0001, Youtao Zhang, Liang Shi 0001, Chun Jason Xue, Jun Yang 0002, Laurence T. Yang
J. Syst. Archit.1
2018 ShadowGC: Cooperative garbage collection with multi-level buffer for performance improvement in NAND flash-based SSDs
abstract
Garbage collection, an essential background activity in NAND flash based SSDs, often introduces large runtime overhead. Recent studies showed that it is beneficial to separate the flash pages that have dirty copies in the write buffers from those that do not. However, the existing schemes exploring this observation have limitations, which prevent them from maximizing the performance improvement. In this paper, we address the above challenge through ShadowGC, a novel GC design that exploits the pages in both host-side and device-side write buffers and adopts different read and write strategies to minimize the GC overhead. When garbage collecting flash pages that have dirty copies in the device-side write buffer, ShadowGC reads data from the write buffer. When garbage collecting flash pages that have dirty copies in the host-side write buffer, ShadowGC moves them to dedicated blocks and speeds up the movement with fast-write operations. Our experimental results show that, on average, ShadowGC reduces the write amplification by 16.2% and the GC latency by 20.5% over the state-of-the-art.
Jinhua Cui 0001, Youtao Zhang, Jianhang Huang, Weiguo Wu, Jun Yang 0002
DATE1
2018 ApproxFTL: On the Performance and Lifetime Improvement of 3-D NAND Flash-Based SSDs
abstract
3-D NAND flash is one of the most prospective advances in flash memory industry. While 3-D flash improves cell density and reduces lithography cost through die stacking, it suffers from severe program disturbance, which leads to significant performance and lifetime degradation for 3-D flash-based SSDs. To address the above challenge, we propose ApproxFTL, an approximate-write aware flash translation layer design, that uses approximate-write operations to store error-resilient data of modern applications. By reducing the maximal threshold voltage and tightening the guard bands between multilevel cell states, approximate write operations not only finish early but also exhibit large disturbance reduction, which can be exploited to alleviate disturbance in physical blocks that save both precise and approximate data. ApproxFTL maximizes the disturbance mitigation through approximate-write aware data placement, wear leveling, and garbage collection enhancements. Our experimental results show that ApproxFTL, while preserving high data quality, improves the read and write response time of flash accesses by 41.38% and 45.64% on average, respectively, and extends the lifetime of 3-D flash-based SSDs by 5.75% when comparing to the state-of-the-art.
Jinhua Cui 0001, Youtao Zhang, Liang Shi 0001, Chun Jason Xue, Weiguo Wu, Jun Yang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 DLV: Exploiting Device Level Latency Variations for Performance Improvement on Flash Memory Storage Systems
abstract
NAND flash has been widely adopted in storage systems due to its better read and write performance and lower power consumption over traditional mechanical hard drives. To meet the increasing performance demand of modern applications, recent studies speed up flash accesses by exploiting access latency variations at the device level. Unfortunately, existing flash access schedulers are still oblivious to such variations, leading to suboptimal I/O performance improvements. In this paper, we propose DLV, a novel flash access scheduler for exploring scheduling opportunities due to device level access latency variations. DLV improves flash access speeds based on process variations and data retention time difference across flash blocks. More importantly, DLV integrates access speed optimization with access scheduling such that the average access response time can be effectively reduced on flash memory storage systems. Our experimental results show that DLV achieves an average of 41.5% performance improvement over the state-of-the-art.
Jinhua Cui 0001, Youtao Zhang, Weiguo Wu, Jun Yang 0002, Yinfeng Wang, Jianhang Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Exploiting latency variation for access conflict reduction of NAND flash memory
abstract
NAND flash memory has been widely used in storage systems by offering greater read/write performance and lower power consumption than mechanical hard drives. Recently, the tradeoff between endurance, write speed, and read speed has been exploited from many ways for I/O performance improvement, which also induce the read/write latency variation. In this paper, the latency variation is exploited in I/O scheduling for access characteristic guided read and write latency minimization. First, with the understanding of the relationship among read latency, write latency and raw bit error rates (RBER), different ways to exploit the relationship for read and write latency reduction is discussed. Then, an I/O scheduling scheme is proposed by using hotness and retention age of accessed data to determine the speed of writes or reads, giving scheduling priority to fast writes and fast reads for conflict reduction. Experiments with various traces reveal that the proposed technique achieves significant read and write performance improvements.
Jinhua Cui 0001, Weiguo Wu, Xingjun Zhang, Jianhang Huang, Yinfeng Wang
MSST1
2016 VIOS: A Variation-Aware I/O Scheduler for Flash-Based Storage Systems
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC1
2016 Exploiting Cross-Layer Hotness Identification to Improve Flash Memory System Performance
Jinhua Cui 0001, Weiguo Wu, Shiqiang Nie, Jianhang Huang, Zhuang Hu, Nianjun Zou, Yinfeng Wang
NPC1