Zhigang Cai

dblp:20/9950 · DBLP profile ↗
← Back
40ranked-venue papers
1as first author
36since 2021 · last 2026
0000-0002-8406-8461ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 1 first-author · 33 since 2021Software engineering, systems software and programming languages · 8 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Multilevel SLC Cache Architecture for SLC-TLC Hybrid Storage
abstract
While high-density NAND flash memory, such as triple-level-cell (TLC) memory achieves enhanced storage capacity and cost efficiency by storing three bits per cell compared to the single-level-cell (SLC) flash memory, it exhibits compromised I/O throughput and reduced endurance. To mitigate the intrinsic limitations of TLC flash memory, contemporary designs employ hybrid architectures that reconfigure designated TLC chips to operate in the SLC mode, forming an efficient cache to address latency and endurance constraints inherent to conventional TLC implementations. However, it is challenging to optimize the SLC cache, such as the granularity of cached data and cold/hot data separation. In this paper, we propose supporting a two-level hierarchy (i.e. L1 and L2) of SLC cache stores based on the varying granularity of cached, and we present a mathematical model to direct the segmentation of the L1 and the L2 cache in the SLC region, by considering the garbage collection counts in both level caches, and the write size characteristics of user applications. Besides, we introduce a scheme to select the garbage collection (GC) victim block for retaining the hot data pages in the SLC cache and then improving I/O responsiveness. The evaluation results show that our proposal can improve I/O performance by between 5.4% and 63.3%, in contrast to existing cache management schemes for SLC-TLC hybrid storage.
Zhibing Sha, Jun Li 0062, Huanhuan Tian, Zhigang Cai, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 SWEEP: Gathering Hot Read/Write Data Together to Minimize Read Reclaims in SSDs
abstract
Read disturb is a circuit-level noise in high-density solid-state drives (SSDs), which may corrupt existing data in SSD blocks and then cause high read error rates. The approach of read reclaim (RR) is commonly used to avoid read disturb errors by migrating the valid data pages to other free blocks for resetting the negative effects of read disturbs, but it affects both I/O responsiveness and SSD lifetime. This article proposes SWEEP to minimize the number of RR operations. Specifically, it gathers hot read and hot write data pages together, and saves them in a portion of designated blocks. After that, the hot write data pages will be invalidated and the hot read data pages in such blocks are likely to be migrated to other blocks through garbage collection (GC). Consequently, the read disturbs on these hot read data pages will be reset, without additional RR operations. Trace-driven simulation experiments show that our proposal can significantly reduce the number of RR operations by between 1.6 % and 74.4 %, which contributes to a maximum reduction of 29.7 % on total erase operations, indicating a better lifetime of SSD devices. In addition, it can reduce the overall I/O response time by up to 52.9 %, compared to existing optimization schemes for SSDs.
Shiyu Zhong, Jiaxu Wu, Zhigang Cai, Jianwei Liao 0001
ACM Trans. Design Autom. Electr. Syst.6
2026 Prefetching Mapping Table Entries to Speed Up Address Translation in DRAM-Less SSDs
abstract
Certain NAND flash-based storage devices are not equipped with dynamic random access memory (DRAM) for holding the whole mapping table, due to the constraints of chip size and the cost, such as secure digital cards. Such DRAM-less flash memory can load only a subset of frequently accessed mapping table entries into a limited-capacity, on-board static random access memory (SRAM) cache, to expedite address translation. To improve the use efficiency of the SRAM cache, this article proposes to prefetch mapping table entries into the cache according to their locality . Then, it can minimize the number of translation page reads at the flash array caused by loading the required entries, thus improving I/O performance. Specifically, we use the indicator of runs test to reflect the locality of mapping table entries on the same translation page. When processing a missed mapping entry, it determines whether adjacent mapping entries accompanying the missed one should be loaded into the cache or not, on the basis of the runs of the target translation page. Consequently, subsequent requests requiring access to these mapping entries can be quickly responded to with the cached ones, instead of reading the target translation pages. In addition, we support cache management based on the runs test of the mapping entries to further improve the cache hit ratio. Experimental results show that our proposal can increase the hit ratio of mapping table entries by 39.4 % and reduce overall I/O latency by 23.4 % on average, in contrast to state-of-the-art schemes.
Zhibing Sha, Jun Li 0062, Zhigang Cai, Yuanquan Shi, Jianwei Liao 0001
ACM Trans. Storage4
2026 Cache Partition Management for Improving Fairness and I/O Responsiveness in NVMe SSDs
abstract
NVMe SSDs have become mainstream storage devices thanks to their compact size and ultra-low latency. It has been observed that the impact of interference among all concurrently running streams (i.e., I/O workloads) on their overall responsiveness differs significantly, thus leading to unfairness. The intensity and access locality of streams are the primary factors contributing to interference. A small-sized data cache is commonly equipped in the front-end of SSDs to improve I/O performance and extend the device's lifetime. The degree of parallelism at this level, however, is limited compared to that of the SSD back end, which consists of multiple channels, chips, and planes. Therefore, the impact of interference can be more significant at the data cache level. In this paper, we propose a cache division management scheme that not only contributes to fairness but also boosts I/O responsiveness across all workloads in NVMe SSDs. Specifically, our proposal supports long-term data cache partitioning and short-term cache adjustment with global sharing, ensuring better fairness and further enhancing cache utilization efficiency in multi-stream scenarios. Trace-driven simulation experiments show that our proposal improves fairness by an average of66.0% and reduces overall I/O response time by between3.8% and18.0%, compared to existing cache management schemes for NVMe SSDs.
Fan Yang 0110, Zhibing Sha, Zhigang Cai, Balazs Gerofi, Yuanquan Shi, Jianwei Liao 0001
IEEE Trans. Parallel Distributed Syst.5
2025 A Two-level SLC Cache Hierarchy for Hybrid SSDs
abstract
Although high-density NAND flash memory, such as triple-level-cell (TLC) flash memory can offer high density, its lower write performance and endurance compared to single-level-cell (SLC) flash memory are impediments to the proliferation of TLC products. To overcome such disadvantages of TLC flash memory, hybrid architectures, which integrate a portion of SLC chips and employ them as a write cache, are widely adopted in commercial solid-state disks (SSDs). However, it is challenging to optimize the SLC cache, such as the granularity of cached data and the cold/hot data separation. In this paper, we propose supporting two-level hierarchy (i.e. L1 and L2) of SLC cache stores based on varying granularity of cached data. Moreover, we support the segmentation of the L1 and the L2 cache in the SLC region with a dynamic manner, by considering the write size characteristics of user applications. The evaluation results show that our proposal can improve I/O performance by between 12.6% and 25.1%, in contrast to existing cache management schemes for SLC-TLC hybrid storage.
Zhibing Sha, Jun Li 0062, Huanhuan Tian, Zhigang Cai, Jianwei Liao 0001
DATE6
2025 SetMP: Set Associative Mapping Management for Multi-plane Optimization in SSDs
abstract
Modern solid state drives (SSDs) are composed of a four-level parallel structure, including channels, chips, dies, and planes, to enhance SSD performance with maximum access parallelism. Because the planes within the same die share the same set of control units and peripheral circuits, it generally has to open multiple aligned blocks to enable access parallelism through multi-plane (MP) operations. Such a passive method, however, cannot effectively exploit plane level parallelism, since MP operations can only be triggered when the accessed data pages have the same offset address across the planes. In addition, it will worsen the block open time issue, as multiple aligned blocks are opened to enable MP operations for simultaneous data writing. This, in turn, increases the error rate when reading data from blocks with long open time. This paper introduces SetMP, a novel approach that proactively aggregates requests to exploit plane level parallelism through set associative management. By increasing the frequency of MP operations, SetMP enhances I/O responsiveness while reducing the open time of block associated with maintaining multiple open blocks for MP operations. Evaluation results demonstrate that SetMP achieves an average reduction in I/O latency of 16.9%, without significantly increasing the open time of block, outperforming existing optimization schemes.
Aobo Yang, Huanhuan Tian, Yuyang He, Jiaxu Wu, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001
LCTES7
2025 DeferRefresh: Deferred retention-aware refresh for archival management on high-density SSD-based RAID systems
Huanhuan Tian, Zhibing Sha, Fan Yang 0110, Aobo Yang, Zhigang Cai, Jianwei Liao 0001
J. Syst. Archit.7
2025 Balancing I/O and wear-out distribution inside SSDs with optimized cache management
Jiaxu Wu, Aobo Yang, Fan Yang 0110, Zhigang Cai, Jianwei Liao 0001
J. Syst. Archit.5
2025 Set associative address mapping to improve data throughput and reduce tail latency in SSDs
Aobo Yang, Jiaxu Wu, Fan Yang 0110, Zhibing Sha, Shiyu Zhong, Zhigang Cai, Jianwei Liao 0001
J. Syst. Archit.7
2025 Supports of Data Cache Division for Computational Solid-state Drives
abstract
The computational SSD ( CompSSD ), with high computing capabilities, can function not only as a storage device but also as a computing node. The data cache of the CompSSD device stores both the output data from host-side tasks and the input data for tasks executed on the CompSSD . However, current cache management strategies are optimized for traditional SSDs and are incompatible with the unique requirements of CompSSD . To address the issue of cache management for CompSSD , this article proposes a novel cache division scheme, to dynamically divide the cache into two parts, for separately buffering output data from host-side tasks and input data used by CompSSD -side tasks. To this end, we construct a mathematical model that periodically estimate an optimal cache division ratio, by considering the factors of the ratios of read/write data amount, the cache hits, and the overhead of data transfer between the storage device and the host. Besides, we propose a scheme of proactive data flushing to write the output data to the underlying flash arrays, without impacts on I/O responsiveness. The trace-driven experiments show that our scheme can improve the overall I/O latency by 35.4% on average, in contrast to existing cache management schemes for CompSSD devices.
Zhibing Sha, Shuaiwen Yu, Chengyong Tang, Zhigang Cai, Min Huang 0018, Jun Li 0062, Jianwei Liao 0001
ACM Trans. Archit. Code Optim.4
2025 Improving I/O Performance and Fairness in NVMe SSDs With Pooling Portions of Cache Partitions
abstract
Nonvolatile memory express (NVMe) solid-state drives (SSDs) have become mainstream storage devices in today’s computing systems, due to their high throughput and ultralow latency. It has been observed that the impact of interference among all concurrently running streams (i.e., I/O workloads) on their overall responsiveness differs significantly in multistream SSDs, resulting in unfairness. This article proposes a cache division management scheme built on top of the evenly partition scheme for NVMe SSDs, to enhance I/O responsiveness without consciously sacrificing fairness. To this end, we first build a mathematical model to directly cut portions from the Local cache partitions allocated to concurrently running streams, considering their run-time performance measures. Then, our approach pools these portions together for the use of all streams. As a result, each stream has its corresponding Local cache space for ensuring fairness, meanwhile the pooled Global cache space is shared by all streams for enhancing I/O responsiveness. Trace-driven simulation experiments demonstrate that our proposal reduces the overall I/O latency by up to24.4%, and improve the measure of fairness by$\mathtt{2.5}\times $on average, in contrast to existing cache management schemes for NVMe SSDs.
Zhigang Cai, Fengxiang Zhang, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Adaptive DRAM Cache Division for Computational Solid-state Drives
abstract
High computational capabilities enable modern solid-state drives (SSDs) to be computing nodes, not just faster storage devices, and the SSD having such capability is generally called as the computational SSD (CompSSD). Then, the DRAM data cache of CompSSD should hold not only the output data of the tasks running at the host side, but also the input data of the tasks executed at the SSD side. To boost the use efficiency of the cache inside CompSSD, this paper proposes an adaptive cache division scheme, to dynamically split the cache space for separately buffering the output data running at the host and the input data running at the CompSSD. Specifically, we construct a mathematical model running at flash translation layer of CompSSD, to periodically determine the cache proportion of the workloads running at the host side and the CompSSD side, by considering the factors of the ratios of read/write data amount, the cache hits, and the overhead of data transfer between the storage device and the host. Then, both the output data and the input data can be buffered in their own private cache parts, so that the overall I/O performance can be enhanced. Trace-driven simulation experiments show that our proposal can reduce the overall I/O latency by 27.5 % on average, in contrast to existing cache management schemes.
Shuaiwen Yu, Zhibing Sha, Chengyong Tang, Zhigang Cai, Min Huang 0018, Jun Li 0062, Jianwei Liao 0001
DATE4
2024 Page Type-Aware Full-Sequence Program Scheduling via Reinforcement Learning in High Density SSDs
abstract
Full-sequence program (FSP) can program multiple bits simultaneously, and thus complete a multiple-page write at one time for naturally enhancing write performance of high density 3-D solid-state drives (SSDs). This article proposes an FSP scheduling approach for the 3-D quad-level cell (QLC) SSDs, to further boost their read responsiveness. Considering each FSP operation in QLC SSDs spansfourdifferent types of QLC pages having dissimilar read latency, we introduce matching four pages of application data to the suited QLC pages and flush them together with the one-shot program of FSP. To this end, we employ reinforcement learning to classify the (cached) application data intofourcategories on the basis of their historical access frequency and the associating request size. Thus, the frequently read data can be mapped to the QLC pages having less access latency, meanwhile the other data can be flushed onto the slow QLC pages. Then, we can group four different categories of data pages and flush them together into a four-page unit of 3-D QLC SSDs with an FSP operation. In addition, a proactive rewrite method is also triggered for grouping the hot read data with the cached data to form an FSP unit. Through a series of emulation tests on several realistic disk traces, we show that the proposed mechanisms yields notable performance improvement on the read responsiveness.
Jun Li 0062, Zhigang Cai, Balazs Gerofi, Yutaka Ishikawa, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Fast Online Reconstruction for SSD-Based RAID-5 Storage Systems
abstract
NAND based solid state drives (SSDs) are almost ubiquitously used in safety-critical systems, and recent advances have demonstrated RAID implementations that built on the top of SSDs can effectively enhance the data integrity and reliability. RAID can restore the lost data chunks in case of failures of RAID components (i.e., SSDs in the context), through a process of RAID reconstruction. Specially, online RAID reconstruction allows the RAID system to continue fulfilling user I/O requests during reconstruction. Servicing user I/O requests, however, significantly affects the performance of reconstruction due to contention for the shared SSD bandwidth. This paper proposes a fast online reconstruction method for SSD-based RAID systems, that preferably restores the lost chunks if the replaced SSD device is idle to reduce the reconstruction time, thus minimizing the probability of a second disk failure in the RAID system during reconstruction. Furthermore, it schedules the tasks of restoring data/parity chunks according to the their impacts on other working SSDs in the RAID system, for the purpose of reducing the overall I/O latency. Through a series of experiments based on the selected disk traces of real-world applications, we show that the proposed reconstruction scheme can reduce the reconstruction time by up to 45.6%, and meanwhile cut down the I/O latency by 9.8% on average compared to state-of-the-art methods.
Haodong Lin, Junhao Luo, Jun Li 0062, Zhibing Sha, Zhigang Cai, Yuanquan Shi, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Modeling Retention Errors of 3D NAND Flash for Optimizing Data Placement
abstract
Considering 3D NAND flash has a new property of process variation (PV) , which causes different raw bit error rates (RBER) among different layers of the flash block. This article builds a mathematical model for estimating the retention errors of flash cells, by considering the factor of layer-to-layer PV in 3D NAND flash memory, as well as the factors of program/erase (P/E) cycle and retention time of data. Then, it proposes classifying the layers of flash block in 3D NAND flash memory into profitable and unprofitable categories, according to the error correction overhead. After understanding the retention error variation of different layers in 3D NAND flash, we design a mechanism of data placement, which maps the write data onto a suitable layer of flash block, according to the data hotness and the error correction overhead of layers, to boost read performance of 3D NAND flash. The experimental results demonstrate that our proposed retention error estimation model can yield a R 2 value of 0.966 on average, verifying the accuracy of the model. Based on the estimated retention error rates of layers, the proposed data placement mechanism can noticeably reduce the read latency by 29.8 % on average, compared with state-of-the-art methods against retention errors for 3D NAND flash memory.
Jiewen Tang, Jun Li 0062, Zhibing Sha, Fan Yang 0110, Zhigang Cai, Jianwei Liao 0001
ACM Trans. Design Autom. Electr. Syst.6
2024 Polling Sanitization to Balance I/O Latency and Data Security of High-density SSDs
abstract
Sanitization is an effective approach for ensuring data security through scrubbing invalid but sensitive data pages, with the cost of impacts on storage performance due to moving out valid pages from the sanitization-required wordline, which is a logical read/write unit and consists of multiple pages in high-density SSDs. To minimize the impacts on I/O latency and data security, this article proposes a polling-based scheduling approach for data sanitization in high-density SSDs. Our method polls a specific SSD channel for completing data sanitization at the block granularity, meanwhile other channels can still service I/O requests. Furthermore, our method assigns a low priority to the blocks that are more likely to have future adjacent page invalidations inside sanitization-required wordlines, while selecting the sanitization block, to minimize the negative impacts of moving valid pages. Through a series of emulation experiments on several disk traces of real-world applications, we show that our proposal can decrease the negative effects of data sanitization in terms of the risk-performance index, which is a united time metric of I/O responsiveness and the unsafe time interval, by 16.34% , on average, compared to related sanitization methods.
Zhigang Cai, Fan Yang 0110, Jun Li 0062, François Trahay, Zheng Yang 0001, Chao Wang 0069, Jianwei Liao 0001
ACM Trans. Storage2
2023 Out-of-channel data placement for balancing wear-out and I/O workloads in RAID-enabled SSDs
abstract
Channel-level RAID implementation SSDs can fight against channel failures inside SSDs, but greatly suffer from imbalanced wear-out (i.e. erase) and I/O workloads across all SSD channels, due to the nature of in-channel updates on data/parity chunks of data stripes. This paper proposes exchanging channel locations of data/parity chunks belonging to the same stripe when satisfying update (write) requests, termed as out-of-channel data placement. Consequently, it can smooth wear-out and I/O workloads across SSD channels, thus reducing I/O response time. Through a series of emulation experiments on several realistic disk traces, we show that our proposal can greatly improve I/O performance, as well as noticeably balance the wear-out and I/O workloads, in contrast to related methods.
Fan Yang 0110, Chengqi Xiao, Jun Li 0062, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001
DATE5
2023 Re-aligning Across-page Requests for Flash-based Solid-state Drives
abstract
In flash-based solid-state drives (SSDs), certain small unaligned I/O requests span two logical pages though their size is not larger than the basic write/read unit of SSDs (i.e. an SSD page), and we term them as across-page requests. Servicing such across-page requests triggers two separated I/O operations on different SSD pages, and thus impacts the I/O performance and the endurance of SSDs. For mitigating negative effects caused by across-page requests, this paper proposes a novel flash translation layer (FTL) scheme for SSDs to separately re-align such requests via remapping them onto a single SSD page. Consequently, both read and write requests on the across-page data can be completed with one page-level I/O operation. Through a series of experiments based on the selected disk traces of real-world applications, we demonstrate that the proposed realigning method at FTL of SSD devices, can noticeably reduce the I/O latency by between 4.6% and 11.6%, and the erase number (i.e. the indicator of SSD endurance) by between 6.4% and 19.11%, compared to state-of-the-art methods.
Zhigang Cai, Chengyong Tang, Minjun Li, François Trahay, Jun Li 0062, Zhibing Sha, Fan Yang 0110, Jianwei Liao 0001
ICPP1
2023 Modeling Retention Errors on Modern 3D-Flash Products
abstract
The innovative stacking architecture of 3D NAND flash offers a promising solution to increase the capacity of flash memory and further cut down the per-unit price. Such compact architectures, however, cause varied kinds of errors because of the hardware nature. Specially, retention errors are primary causes of read retries in high-density flash memory, which are induced by charge leakage over time. This paper proposes an empirical mathematical model to estimate the error rate caused by retention errors on the granularity of block, which is the basic program/erase (P/E) unit in 3D NAND flash memory. Specifically, we build the generalized model by considering the factors of layer-to-layer interference, early retention loss, the P/E cycle and the retention time of data of 3D NAND flash memory. Then, we validate the model on four commercial 3D NAND flash products, and the experimental results verify the accuracy of our proposed estimation model. At last, we apply our model in the existing I/O optimization of write scheduling for 3D NAND flash memory, for showing its contributions to better I/O performance.
Jianwei Liao 0001, Jiewen Tang, Jun Li 0062, Junhao Luo, Chenqi Xiao, Zhigang Cai, Lei Chen 0002
ISCAS6
2023 Rep-RAID: An Integrated Approach to Optimizing Data Replication and Garbage Collection in RAID-Enabled SSDs
abstract
Redundant Array of Independent Disks (RAID) technology has been recently introduced to flash memory based SSDs to enhance their data reliability. Although RAID increases reliability, it doubles the number of write operations and requires additional parity computation as every write operation on a data chunk leads to another update on the corresponding parity chunk. Data replication has been proposed to mitigate the overhead of write requests in RAID enabled SSDs, however, replication increases the cost of garbage collection (GC), which in turn limits the improvement of I/O performance compared to the baseline RAID implementation. This paper introduces Rep-RAID, an improved data replication management scheme accompanied with optimized GC for RAID-enabled SSDs. Guided by a mathematical model, Rep-RAID only replicates frequently updated data chunks. Furthermore, Rep-RAID reorganizes new data stripes during the GC process by utilizing replicated data to replace invalid data chunks caused by data replication in old stripes. As a result, it decreases I/O latency for both read and write requests and significantly reduces the GC overhead induced by data movement. Experimental results show that the proposed scheme can improve I/O performance by 16.7%, and reduce tail latency by up to 17.9% at the 99.99th percentile, when compared to the state-of-the-art RAID-enabled SSDs.
Jun Li 0062, Balazs Gerofi, François Trahay, Zhigang Cai, Jianwei Liao 0001
LCTES4
2023 Cache eviction for SSD-HDD hybrid storage based on sequential packing
Chengyong Tang, Zhibing Sha, Jun Li 0062, Haodong Lin, Lei Chen 0002, Zhigang Cai, Jianwei Liao 0001
J. Syst. Archit.6
2023 Adaptive Management With Request Granularity for DRAM Cache Inside nand-Based SSDs
abstract
Most flash-based solid-state drives (SSDs) adopt an onboard dynamic random access memory (DRAM) to buffer hot write data. Then, the write or overwrite operations can be absorbed by the DRAM cache, given that there is sufficient locality in the applications’ I/O access pattern, to consequently avoid flushing the write data onto underlying SSD cells. After analyzing typical real-world workloads over SSDs, we observed that the buffered data of small-size requests are more likely to be reaccessed than those of large write requests. To efficiently utilize the limited space of DRAM cache, this article proposes an adaptive request granularity-based cache management scheme for SSDs. First, we introduce a request block corresponding to a write request, as the cache management granularity, and propose a dynamic manner for classifying small and large request blocks. Next, we design three-level linked lists for supporting different routines of upgradation for small and large request blocks, once their data have been hit in the cache. Finally, we present a scheme of evicting the request blocks having the minimum cost in cache replacement, by taking both factors of access hotness and time discounting into account. Experimental results show that our proposal can yield improvements on cache hits and the overall I/O latency by21.8% and14.7% on average, compared to state-of-the-art cache management schemes inside SSDs.
Haodong Lin, Jun Li 0062, Zhibing Sha, Zhigang Cai, Yuanquan Shi, Balazs Gerofi, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Proactive Stripe Reconstruction to Improve Cache Use Efficiency of SSD-Based RAID Systems
abstract
Solid-State Drives (SSDs) exhibit different failure characteristics compared to conventional hard disk drives. In particular, the Bit Error Rate (BER) of an SSD increases as it bears more writes. Then, Parity-based Redundant Array of Inexpensive Disks (RAID) arrays composed from SSDs are introduced to address correlated failures. In the RAID-5 implementation, specifically, the process of parity generation (or update) associating with a data stripe, consists of read and write operations to the SSDs. Whenever a new update request comes to the RAID system, the related parity must be also updated and flushed onto the RAID component of SSD. Such frequent parity updates result in poor RAID performance and shorten the life-time of the SSDs. Consequently, a DRAM cache is commonly equipped accompanying with the RAID controller, called the parity cache, and used to buffer the parity chunks that are most frequently updated data, for boosting I/O performance. To better improve the use efficiency of the parity cache, this paper proposes a stripe reconstruction approach to minimize the number of parity updates on SSDs, thus boosting I/O performance of the SSD RAID system. When the currently updated stripe has both cold and hot updated data chunks, it will proactively carry out stripe reconstruction if we can find another matched stripe that also includes cold and hot update data chunks on the complementary RAID components. In the reconstruction process, we first group the cold data chunks of two matched stripes, to build a new stripe and flush the parity chunk on the RAID component. After that, the hot data chunks are organized as a new stripe as well, and its parity chunk is buffered in the parity cache. This results in better cache use efficiency, as it can reduce the number of parity updates on RAID components of SSDs, as well as proactively free up cache space for quickly absorbing subsequent write requests. In addition, the proposed method adjusts the target SSD of write requests based on stripe reconstructions through considering the I/O workload balance of all SSDs. Experimental results show that our proposal can reduce the number of parity chunk updates in SSDs by 2.3% and overall I/O latency by 12.2% on average, compared to state-of-the-art parity cache management techniques.
Zhibing Sha, Jun Li 0062, Balazs Gerofi, Zhigang Cai, Jianwei Liao 0001
ACM Trans. Embed. Comput. Syst.5
2023 Visibility Graph-based Cache Management for DRAM Buffer Inside Solid-state Drives
abstract
Most solid-state drives (SSDs) adopt an on-board Dynamic Random Access Memory (DRAM) to buffer the write data, which can significantly reduce the amount of write operations committed to the flash array of SSD if data exhibits locality in write operations. This article focuses on efficiently managing the small amount of DRAM cache inside SSDs. The basic idea is to employ the visibility graph technique to unify both temporal and spatial locality of references of I/O accesses, for directing cache management in SSDs. Specifically, we propose to adaptively generate the visibility graph of cached data pages and then support batch adjustment of adjacent or nearby (hot) cached data pages by referring to the connection situations in the visibility graph. In addition, we propose to evict the buffered data pages in batches by also referring to the connection situations, to maximize the internal flushing parallelism of SSD devices without worsening I/O congestion. The trace-driven simulation experiments show that our proposal can yield improvements on cache hits by between 0.8 % and 19.8 %, and the overall I/O latency by 25.6 % on average, compared to state-of-the-art cache management schemes inside SSDs.
Zhibing Sha, Jun Li 0062, Fengxiang Zhang, Min Huang 0018, Zhigang Cai, François Trahay, Jianwei Liao 0001
ACM Trans. Storage5
2022 Unifying Temporal and Spatial Locality for Cache Management inside SSDs
abstract
To ensure better I/O performance of solid-state drivers (SSDs), a dynamic random access memory (DRAM) is commonly equipped as a cache to absorb overwrites or writes, instead of directly flushing them onto underlying SSD cells. This paper focuses on the management of the small amount cache inside SSDs. First, we propose to unify both factors of temporal and spatial locality of user applications by employing the visibility graph technique, for directing cache management. Next, we propose to support batch adjustment of adjacent or nearby (hot) cached data pages by referring to the connection situations in the visibility graph of all cached pages. At last, we propose to evict the buffered data pages in batches, to maximize the internal flushing parallelism of SSD devices, without worsening I/O congestion. The trace-driven simulation experiments show that our proposal can yield improvements on cache hits by more than 2.8%, and the overall I/O latency by 20.2% on average, in contrast to conventional cache schemes inside SSDs.
Zhibing Sha, Zhigang Cai, François Trahay, Jianwei Liao 0001
DATE2
2022 Multi-stream Information-Based Neural Network for Mammogram Mass Segmentation
Zijian Deng, Yu Gui, Zhigang Cai, Jianwei Liao 0001
ICANN (1)5
2022 DRAM Cache Management with Request Granularity for NAND-based SSDs
abstract
Most flash-based solid-state drives (SSDs) employ an on-board Dynamic Random Access Memory (DRAM) to cache hot data at the SSD page granularity. This can significantly reduce the number of flush operations to the underlying arrays of SSDs given that there is sufficient locality in the applications’ I/O access pattern. We observe, however, that in most I/O workloads over SSDs the buffered data of small sized requests are more likely to be re-accessed than those of larger requests, which also require more DRAM space for caching their data.
Haodong Lin, Zhibing Sha, Jun Li 0062, Zhigang Cai, Balazs Gerofi, Yuanquan Shi, Jianwei Liao 0001
ICPP4
2022 Adaptive Switch on Wear Leveling for Enhancing I/O Latency and Lifetime of High-Density SSDs
abstract
NAND flash-based high-density solid-state drives (SSDs) commonly have limited write endurance, as a single block becomes unavailable after a finite number of program/erase cycles. The (static) wear leveling algorithms migrate cold data to more worn blocks for yielding an even distribution of wears in SSDs. However, migrating cold data must delay normal I/O processing, and unnecessary wear leveling operations may damage SSD endurance as each operation causes a program/erase cycle. This article proposes an adaptive switch on wear leveling (called ASWL), to boost I/O performance and lifetime for high-density SSDs. The basic idea of ASWL is to adaptively switch the threshold function for intelligently triggering wear leveling operations at different stages of SSD lifetime. To this end, we build a mathematical model determining the wear leveling threshold condition with a dynamic manner, by considering the real-time factors of I/O intensity and the wear difference of SSD blocks. Through a series of emulation experiments on several realistic disk traces, we show that the proposed ASWL mechanism can not only greatly improve I/O performance but also noticeably extend the lifetime of high-density SSDs, in contrast to existing wear leveling methods.
Jun Li 0062, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Read Refresh Scheduling and Data Reallocation against Read Disturb in SSDs
abstract
Read disturb is a circuit-level noise in flash-based Solid-State Drives (SSDs), induced by intensive read requests, which may result in unexpected read errors. The approach of read refresh (RR) is commonly adopted to mitigate its negative effects by unconditionally migrating all valid data pages in the RR block to another new block. However, routine RR operations greatly impact the I/O responsiveness of SSDs, because the processing on normal I/O requests must be blocked at the same time. To further reduce the negative effects of read refresh, this article proposes a read refresh scheduling and data reallocation method to deal with two primary issues with respect to an RR operation, including where to place data pages and when to trigger page migrations. Specifically, we first construct a data reallocation model to match the data pages in the RR block and the destination blocks for addressing the issue of where to place the data. The model considers not only the read hotness of pages in the RR block, but also the accumulated read counts of the destination blocks. Moreover, for addressing the issue of when to trigger data migrations, we build a timing decision model to determine the time points for completing page migrations by considering the factors of the intensity of I/Os and the disturb situation on the RR block. Through a series of simulation experiments based on several realistic disk traces, we illustrate that the proposed RR scheduling and data reallocation mechanism can noticeably reduce the read errors by more than 10.3% , on average, and the long-tail latency by between 43.9% and 64.0% at the 99.99th percentile, in contrast to state-of-the-art methods.
Jianwei Liao 0001, Jun Li 0062, Mingwang Zhao, Zhibing Sha, Zhigang Cai
ACM Trans. Embed. Comput. Syst.5
2022 Degraded Mode-benefited I/O Scheduling to Ensure I/O Responsiveness in RAID-enabled SSDs
abstract
RAID-enabled SSDs commonly have unbalanced I/O workloads on their components (e.g., SSD channels), as the data/parity chunks in the same stripe may have varied access frequency, which greatly impacts I/O responsiveness. This article proposes a I/O scheduling scheme by resorting to the degraded read mode and the read-modify-write mode to reduce the long-tail latency of I/O requests in RAID-enabled SSDs. The basic idea is to avoid scheduling read or update requests to the heavily congested but targeted RAID components. Such requests are satisfied by accessing other relevant RAID components by certain XOR computations (we call the degraded modes ). Specially, we build a queuing overhead assessment model on the top of factors of data redundancy and the current blocked I/O traffics on SSD channels to precisely dispatch incoming I/O requests to be fulfilled with the degraded mode or not. The trace-driven experiments illustrate that the proposed scheme can reduce the long-tail latency of read requests by 23.1% on average at the 99.99th percentile, in contrast to state-of-the-art scheduling methods.
Zhibing Sha, Jun Li 0062, Zhigang Cai, Min Huang 0018, Jianwei Liao 0001, François Trahay
ACM Trans. Design Autom. Electr. Syst.3
2022 Pattern-Based Prefetching with Adaptive Cache Management Inside of Solid-State Drives
abstract
This article proposes a pattern-based prefetching scheme with the support of adaptive cache management, at the flash translation layer of solid-state drives ( SSDs ). It works inside of SSDs and has features of OS dependence and uses transparency. Specifically, it first mines frequent block access patterns that reflect the correlation among the occurred I/O requests. Then, it compares the requests in the current time window with the identified patterns to direct prefetching data into the cache of SSDs. More importantly, to maximize the cache use efficiency, we build a mathematical model to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the written data. Experimental results show that our proposal can yield improvements on average read latency by 1.8 %– 36.5 % without noticeably increasing the write latency, in contrast to conventional SSD-inside prefetching schemes.
Jun Li 0062, Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Kenli Li 0001, Balazs Gerofi, Yutaka Ishikawa
ACM Trans. Storage3
2021 Block Attribute-aware Data Reallocation to Alleviate Read Disturb in SSDs
abstract
This paper proposes a data reallocation method in RR processes, by taking account of block attributes including P/E cycles and read counts. Specifically, it can distribute the data in the RR block onto a number of available SSD blocks, to allow different blocks keeping their own near optimal (tolerable) read count. Then, it can reduce the number of read fresh operations and the number of raw bit errors. Through a series of simulation experiments based on several realistic disk traces, we demonstrate that the proposed method can decrease the read response time by between 15.11% and 24.96% and the raw bit error rate by 4.14% on average, in contrast to state-of-the-art approaches.
Mingwang Zhao, Jun Li 0062, Zhigang Cai, Jianwei Liao 0001, Yuanquan Shi
DATE3
2021 Intra-page Cache Update in SLC-mode with Partial Programming in High Density SSDs
abstract
Modern high density SSDs commonly designate a part of their capacity as a cache using an Single-level Cell (SLC)-mode region. Partial programming is then adopted for reducing space fragmentation in the SLC-mode pages, but it exacerbates program disturb including both in-page disturb and neighbouring page disturb. This paper proposes a partial programming scheme (called intra-page update) by updating hot, small size data inside a given page to minimize the negative impact induced by program disturb. Moreover, we introduce a novel data movement principle to separate hot and cold write data in the SLC-mode cache when updating the data or carrying out garbage collection. As a result, the hot updated data can be kept in the SLC-mode cache and the cold data will be flushed onto the high density SSD region. Simulation tests on several realistic disk traces show that our proposal improves bit error rate by 9.2%, and I/O performance by 9.3% on average, compared to state-of-the-art methods, without a noticeable decrease in total endurance.
Jun Li 0062, Minjun Li, Zhigang Cai, François Trahay, Mohamed Wahib, Balazs Gerofi, Zhiming Liu 0001, Min Huang 0018, Jianwei Liao 0001
ICPP3
2021 A Novel CFLRU-Based Cache Management Approach for NAND-Based SSDs
Haodong Lin, Jun Li 0062, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001, Yuanquan Shi
NPC4
2021 Low I/O Intensity-aware Partial GC Scheduling to Reduce Long-tail Latency in SSDs
abstract
This article proposes a low I/O intensity-aware scheduling scheme on garbage collection (GC) in SSDs for minimizing the I/O long-tail latency to ensure I/O responsiveness. The basic idea is to assemble partial GC operations by referring to several determinable factors (e.g., I/O characteristics) and dispatch them to be processed together in idle time slots of I/O processing. To this end, it first makes use of Fourier transform to explore the time slots having relative sparse I/O requests for conducting time-consuming GC operations, as the number of affected I/O requests can be limited. After that, it constructs a mathematical model to further figure out the types and quantities of partial GC operations, which are supposed to be dealt with in the explored idle time slots, by taking the factors of I/O intensity, read/write ratio, and the SSD use state into consideration. Through a series of simulation experiments based on several realistic disk traces, we illustrate that the proposed GC scheduling mechanism can noticeably reduce the long-tail latency by between 5.5% and 232.3% at the 99.99th percentile, in contrast to state-of-the-art methods.
Zhibing Sha, Jun Li 0062, Lihao Song, Jiewen Tang, Min Huang 0018, Zhigang Cai, Lianju Qian, Jianwei Liao 0001, Zhiming Liu 0001
ACM Trans. Archit. Code Optim.6
2021 Mitigating Negative Impacts of Read Disturb in SSDs
abstract
Read disturb is a circuit-level noise in solid-state drives (SSDs), which may corrupt existing data in SSD blocks and then cause high read error rate and longer read latency. The approach of read refresh is commonly used to avoid read disturb errors by periodically migrating the hot read data to other free blocks, but it places considerable negative impacts on I/O (Input/Output) responsiveness. This article proposes scheduling approaches on write data and read refresh operations, to mitigate the negative effects caused by read disturb. To be specific, we first construct a model to classify SSD blocks into two categories according to the estimated read error rate by referring to the factors of block’s P/E (Program/Erase) cycle and the accumulated read count to the block. Then, the data being intensively read will be redirected to the block having a small read error rate, as it is not sensitive to read disturb even though the data will be heavily requested. Moreover, we take advantage of reinforcement learning to predict the idle interval between two I/O requests for purposely conducting (partial) read refresh operations. As a result, it is able to minimize negative impacts toward subsequent incoming I/O requests and to ensure I/O responsiveness. Through a series of emulation tests on several realistic disk traces, we demonstrate that the proposed mechanisms can noticeably yield performance improvements on the metrics of read error rate and I/O latency.
Jun Li 0062, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001, Balazs Gerofi, Yutaka Ishikawa
ACM Trans. Design Autom. Electr. Syst.4
2020 Frequent Access Pattern-based Prefetching Inside of Solid-State Drives
abstract
This paper proposes an SSD-inside data prefetching scheme, which has features of OS-dependence and use transparency. To be specific, it first mines frequent block access patterns that reflect the correlation among the occurred requests. Then it compares the requests in the current time window with the identified patterns, to direct fetching data in advance. Furthermore, to maximize the cache use efficiency, we construct a method to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the write data. Experimental results demonstrate that our proposal can yield improvements on average read latency by 6.3% to 9.3% without noticeably increasing write latency, in contrast to conventional SSD-inside prefetching schemes.
Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Yutaka Ishikawa
DATE2
2020 Patch-Based Data Management for Dual-Copy Buffers in RAID-Enabled SSDs
abstract
While a dual-copy redundant array of independent disk (RAID) buffering can remove single point of failure in the buffer of RAID-enabled solid-state drives (SSDs), it loses a half of buffer space efficiency as a result. This article introduces a patch-based data management scheme for dual-copy buffers in RAID-enabled SSDs, for better improving buffer use efficiency. Different from conventional methods that caches data/parity chunks in the RAID buffer, it caches the modified portions of data chunks (called data patches) and then employs a cost-evaluation model to direct patch data replacement. Besides, in order to correctly respond to applications' read/write requests, we design the patch-based read/write mode by taking the original dirty data chunks and relevant up-to-date patches into account. Through a series of simulation tests on several realistic disk traces, we illustrate that patch-based data management can work for RAID-enabled SSDs to enhance buffer use efficiency, and then yield better I/O performance. Especially, our proposal can noticeably reduce the I/O latency by 39.1% and the number of block erases by 3.9% on average, in contrast to state-of-the-art approaches.
Jun Li 0062, Zhibing Sha, Zhigang Cai, François Trahay, Jianwei Liao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Adaptive Process Migrations in Coupled Applications for Exchanging Data in Local File Cache
abstract
Many problems in science and engineering are usually emulated as a set of mutually interacting models, resulting in a coupled or multiphysics application. These component models show challenges originating from their interdisciplinary nature and from their computational and algorithmic complexities. In general, these models are independently developed and maintained, so that they commonly employ the global file system for exchanging their data in the coupled application. To effectively use the local file cache on the compute node for exchanging the data among the processes of such applications, and consequently boosting I/O performance, this article presents a novel mechanism to migrate a process from one compute node to another node on the basis of block I/O dependency. In this newly proposed mechanism, the block I/O dependency between two involved processes running on the different nodes is profiled as block access similarity by taking advantage of the Cohen’s kappa statistic . Then, the process is supposed to be dynamically migrated from its source node to the destination node, on which there is another process having heavy block I/O dependency. As a result, both processes can exchange their data by utilizing the local file cache instead of the global file system to reduce I/O time. The experimental results demonstrate that the I/O performance can be significantly improved, and the time required for executing the application can be resultantly decreased, as expected.
Jianwei Liao 0001, Zhigang Cai, François Trahay, Guoqiang Xiao 0001
ACM Trans. Auton. Adapt. Syst.2
2006 A Matching-Based Automatic Registration for Remotely Sensed Imagery
abstract
How to register and rectify the multi-resource images of remote sensing is becomes one of the urgent problems in the field of earth observation and information acquirement recently. To meet this requirement of multi-source image application, an automatic imagery registration algorithm based on image matching was proposed and described, and an automatic registration workflow of multi-source remote sensing imagery was established and implemented in this paper. Many remote sensing images of SPOT5 (5 m), SPOT4 (10 m) and TM(30 m) which cover Hangzhou area were employed for the experimentation. It could be clearly indicated by the result that the algorithm and workflow designed here are precise, prompt, and practical for geometry registration.
Dengrong Zhang, Le Yu 0002, Zhigang Cai
IGARSS3