Congming Gao

dblp:152/3982 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0003-2611-2652ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 12 first-author · 23 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Nemo: A Low-Write-Amplification Cache for Tiny Objects on Log-Structured Flash Devices
Xufeng Yang, Jingxin Hu, Congming Gao, Tianyang Jiang, Linbo Long, Yina Lv, Jiwu Shu
ASPLOS (2)4
2026 Tetris: Lightweight Hyperparameter Auto-Tuning for Mitigating Performance Spikes in LSM-KVS
Yina Lv, Qiao Li 0001, Quanqing Xu, Congming Gao, Chuanhui Yang, Xiaoli Wang 0002, Chun Jason Xue
ICDE5
2026 LOONG: Utilizing Long-Stride Reprogramming to Enhance the Performance of SSDs
Congming Gao, Jiancong Zheng, Xufeng Yang, Qiao Li 0001, Yina Lv, Chun Jason Xue, Jiwu Shu
ISCA1
2026 Revisiting the RAID Performance Bottleneck in the SSD Era: A Roofline Model Perspective
abstract
RAID is widely employed in modern storage systems due to its high bandwidth and reliability. While traditional RAID controllers utilize large DRAM caches to mitigate HDD latency, the advent of SSDs has shifted potential bottlenecks to the controller architecture itself. In this paper, we introduce a Roofline model adapted for RAID systems, defining I/O Processing Intensity (IOPI) as the ratio of I/O requests to total bytes transferred. We demonstrate that write performance is constrained by DRAM bandwidth under low IOPI and by computational overhead under high IOPI. To mitigate the DRAM bottleneck, we propose a high-bandwidth SRAM-based architecture, bringing parity-based RAID write performance closer to theoretical limits.
Xiaoyang Wang 0006, Congming Gao, Jiwu Shu
ISPASS4
2026 AtRS: Auto-tuning RAID system with GAN
Fangzheng Wang, Congming Gao, Bohong Zhu, Jiwu Shu
Future Gener. Comput. Syst.2
2025 MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Device
Chao Wu 0006, Cheng Ji 0002, Li-Pin Chang, Zongwei Zhu, Congming Gao, Weichao Guo, Yanzhi Wang 0001
FAST5
2025 Overlapping Aware Data Placement Optimizations for LSM Tree-Based Store on ZNS SSDs
abstract
Solid State Drives (SSDs) based on the NVMe Zoned Namespaces (ZNS) interface can notably reduce the costs of address mapping, garbage collection, and over-provisioning by dividing the storage space into multiple zones for sequential writes and random reads. The Log-Structured Merge (LSM) tree, which is extensively used in key-value storage systems, converts random writes to sequential writes, hence a suitable scenario to utilize ZNS SSDs. However, LSM tree associated data significantly varies in lifetime due to the levels and merging mechanisms of the LSM tree. Therefore, without an accurate method to estimate data lifetime, data with disparate lifetimes may be placed in the same zone, thus causing low space utilization and high write amplification within the SSD. To address these issues, the article proposes two data overlapping aware optimizations to realize intelligent data placement: a zone allocation scheme and a garbage collection scheme. The key technique of these optimizations is an accurate data-lifetime estimation by considering both the associated tree level of the data and the data overlapping ratio between the data and those in the neighboring level. Using the estimation technique, the zone allocation optimization can place data with similar lifetimes in the same zone. Besides, the garbage collection optimization can reclaim zones in an adaptive manner based on overlapping ratios to reduce the amount of data migration. Experimental results demonstrate that the optimization schemes effectively reduce garbage collection-incurred data copy by average factors of 2.11× and 1.50× in comparison to a conventional work and a state-of-the-art work, respectively. Consequently, the proposed work successfully alleviates the write amplification effect by 18% and 6%, compared to the conventional work and the state-of-the-art work, respectively.
Jingcheng Shen, Linbo Long, Zhenhua Tan, Congming Gao, Kan Zhong, Masao Okita, Fumihiko Ino
ACM Trans. Archit. Code Optim.5
2024 Overlapping Aware Zone Allocation for LSM Tree-Based Store on ZNS SSDs
abstract
NVMe Zoned Namespace (ZNS) devices partition the storage space into sequential-write zones, notably reducing the costs of address mapping, garbage collection (GC), and overprovisioning. Log-Structured Merge (LSM) tree-based databases convert random writes into sequential writes and can thus be efficiently handled by ZNS devices. Efficient zone-allocation methods play a pivotal role in maximizing the performance of LSM tree-based store running on ZNS devices. However, existing zone-allocation methods encounter high write-amplification factors due to inaccurate lifetime estimation solely based on the LSM-tree levels. To address this, this paper proposes an overlapping-aware zone-allocation method, termed OAZA, which efficiently selects suitable zones to place data. First, OAZA estimates the data lifetime by considering both the LSM-tree level of the data and the relative data hotness within the same tree level. Secondly, OAZA intelligently selects an appropriate zone to store the data based on the estimated lifetime. Experimental results demonstrate that OAZA outperforms two zone-allocation methods that correlate data lifetime merely to the tree level. Specially, OAZA reduces the amount of GC-induced data copy by average factors of 2.7 × and 1.7× in comparison to the two methods, respectively. Additionally, OAZA achieves an impressively low write-amplification factor of 1.1 ×, outperforming the factors of 1.2× and 1.3× achieved by the two compared methods, respectively.
Jingcheng Shen, Linbo Long, Renping Liu 0002, Zhenhua Tan, Congming Gao
ASPDAC6
2024 Para-ZNS: Improving Small-Zone ZNS SSDs Parallelism Through Dynamic Zone Mapping
abstract
The emerging Zoned Namespace (ZNS) interface helps flash-based SSDs achieve high performance by dividing the logical space into fixed-size zones. Typically, a zone is mapped to blocks across multiple dies to achieve I/O parallelism. Small zones can make better use of space and are therefore widely studied. However, a small zone fails to be mapped to blocks residing on all dies, causing underutilized die-level parallelism. Meanwhile, a fine-grained (i.e., plane-level) parallelism is rarely exploited for ZNS SSDs due to a strict limitation mandating that only the same type of operation can be simultaneously performed on the same address across different planes within a die. To address these issues, this paper proposes a novel small-zone ZNS-SSD design with dynamic zone mapping, named Para-ZNS. First, a new parallel block grouping module is devised to group blocks across all planes from multiple dies as a basic unit to be mapped to a zone. Such a basic mapping unit achieves parallelism among multiple dies and plane-level parallelism. Then, a die-parallelism identification module is implemented to locate idle dies. Subsequently, to fully exploit the die-level parallelism, a dynamic zone mapping scheme is employed to intelligently map the basic mapping units on the identified idle dies to open zones. The evaluation results based on a widely-used I/O tester (FIO) demonstrate that Para-ZNS improves the bandwidth by 3.42× on average in comparison to state-of-the-art work.
Zhenhua Tan, Linbo Long, Jingcheng Shen, Congming Gao, Renping Liu 0002
DATE4
2024 Midas Touch: Invalid-Data Assisted Reliability and Performance Boost for 3d High-Density Flash
abstract
High-density 3D NAND flash like QLC (Quad-Level Cell) is prevailing in providing large capacities for data-intensive applications. Because of the structure limitation, a two-step programming with a specific sequence is adopted in 3D QLC flash, where data could become invalid between the two programming steps. This is called invalid programming, as the second-step programming is conducted on partially-invalid wordlines (WLs). By exploiting this phenomenon, this work proposes invalid-data assisted strategies for performance and reliability boosting of valid data in 3D QLC-based flash storage systems. We first propose a high-efficiency re-programming (RP) scheme to reprogram the valid data and a high-reliability not-programming (NP) scheme to program data on the partially-invalid WLs. An adaptive data allocation (ADA) strategy for data management between the SLC and QLC regions is further introduced to reduce the occurrence of invalid programming. The simulator-based experiments show the proposed RP scheme combined with ADA can reduce the execution time for programming by 13.51%, on average. Through real-device evaluations, we present that the NP scheme can reduce the bit error rate of NP-programmed data by 32.8% of the worst page type, thus improving overall reliability, which translates to 12% reduction in refreshing overheads and 30% lifetime extension, on average. Besides, the NP scheme with ADA averagely reduces energy consumption by 4.8%.
Qiao Li 0001, Hongyang Dang, Congming Gao, Jie Zhang 0048, Tei-Wei Kuo, Chun Jason Xue
HPCA4
2024 Ares-Flash: Efficient Parallel Integer Arithmetic Operations Using NAND Flash Memory
abstract
In-Flash Processing (IFP) has been proposed in recent years to realize computation ability inside NAND flash memory. Distinguished from processing-in-memory (PIM) and in-storage-processing (ISP), IFP reduces the data movement starting from the most bottom flash memory medium. It is especially beneficial to those applications that require high processing parallelism (e.g., huge databases, large-scale image processing, etc.). However, IFP works often require expensive extra hardware modification, which brings lots of energy dissipation and area overhead. Recent works, Parabit and Flash-Cosmos, try to implement calculations using flash memory with minimum hardware modification. But their accomplishments still remain at the level of simple bitwise operations, and thus it prevents their work from being widely adopted. In this work, we propose Ares-Flash, a new technology using flash memory to support more complex integer arithmetic operations (e.g., addition, accumulation, and multiplication). Ares leverages the designed page buffer to perform basic mechanisms: full-adder logic and bit-shift. Moreover, we construct computational operations using dedicated control sequences in the page buffer upon two basic mechanisms. Our experimental results indicate that Ares is highly efficient and significantly mitigates data movement from storage to memory or computing units (e.g., CPUs, GPUs). Quantitatively, Ares averagely improves performance and energy efficiency by 8.58×/9.89× and 97.9×/14× compared to the out-storage-processing(OSP)/instorage-processing(ISP) under accumulation tasks with real workloads. It also improves 8.2×/4.5× and 89×/13.2× when performing vector-vector multiplication on real-world workloads.
Congming Gao, Youyou Lu, Yuhao Zhang 0006, Jiwu Shu
MICRO2
2024 In-place Switch: Reprogramming based SLC Cache Design for Hybrid 3D SSDs
abstract
To increase SSD capacity, high bit-density cells, such as Triple-Level Cell (TLC), are utilized within 3D SSDs. However, due to the inferior performance of TLC, SLC/TLC hybrid 3D SSD is designed to use a portion of the TLC space as an SLC cache to achieve high SSD performance by writing host data at the SLC speed. However, our preliminary studies indicate that the SLC cache can lead to a performance cliff if filled rapidly and cause significant write amplification when data migration occurs during idle times. In this work, we propose leveraging a reprogram operation to address these challenges. Specifically, when the SLC cache is full or during idle periods, a reprogram operation is performed to switch used SLC pages to TLC pages in place (termed In-place Switch, IPS). Subsequently, other free TLC space is allocated as the new SLC cache. IPS can continuously provide sufficient SLC cache within SSDs, significantly improving write performance and reducing write amplification. Experimental results demonstrate that IPS can reduce write latency and write amplification by up to 0.75 times and 0.53 times, respectively, compared to state-of-the-art SLC cache technologies.
Xufeng Yang, Jiancong Zheng, Cheng Ji 0002, Congming Gao
NAS4
2024 Space-efficient and high-performance inline deduplication for emerging hybrid storage system with Libra+
Renhui Chen, Tianmeng Zhang, Zijing Li, Congming Gao, Youtao Zhang, Qiao Li 0001, Jun Yang 0002, Jiwu Shu
J. Syst. Archit.4
2024 WA-Zone: Wear-Aware Zone Management Optimization for LSM-Tree on ZNS SSDs
abstract
ZNS SSDs divide the storage space into sequential-write zones, reducing costs of DRAM utilization, garbage collection, and over-provisioning. The sequential-write feature of zones is well-suited for LSM-based databases, where random writes are organized into sequential writes to improve performance. However, the current compaction mechanism of LSM-tree results in widely varying access frequencies (i.e., hotness) of data and thus incurs an extreme imbalance in the distribution of erasure counts across zones. The imbalance significantly limits the lifetime of SSDs. Moreover, the current zone-reset method involves a large number of unnecessary erase operations on unused blocks, further shortening the SSD lifetime. Considering the access pattern of LSM-tree, this article proposes a wear-aware zone-management technique, termed WA-Zone , to effectively balance inter- and intra-zone wear in ZNS SSDs. In WA-Zone, a wear-aware zone allocator is first proposed to dynamically allocate data with different hotness to zones with corresponding lifetimes, enabling an even distribution of the erasure counts across zones. Then, a partial-erase-based zone-reset method is presented to avoid unnecessary erase operations. Furthermore, because the novel zone-reset method might lead to an unbalanced distribution of erasure counts across blocks in a zone, a wear-aware block allocator is proposed. Experimental results based on the FEMU emulator demonstrate the proposed WA-Zone enhances the ZNS-SSD lifetime by 5.23×, compared with the baseline scheme.
Linbo Long, Shuiyong He, Jingcheng Shen, Renping Liu 0002, Zhenhua Tan, Congming Gao, Duo Liu 0002, Kan Zhong
ACM Trans. Archit. Code Optim.6
2024 Optimizing Garbage Collection for ZNS SSDs via In-storage Data Migration and Address Remapping
abstract
The NVMe Zoned Namespace (ZNS) is a high-performance interface for flash-based solid-state drives (SSDs), which divides the logical address space into fixed-size and sequential-write zones. Meanwhile, ZNS SSDs eliminate in-device garbage collection (GC) by shifting the responsibility of GC to the host. However, the host-side GC of ZNS SSDs is not efficient. On the one hand, data migration during GC first moves data to the host buffer and then writes back the transferred data to the new location in the SSD, resulting in an unnecessary end-to-end transfer overhead. On the other hand, due to the pre-configured mapping between zones and blocks, GC incurs a large block-to-block rewrite overhead, i.e., even if most of the data in a block of the victim zone is valid, the valid data will still be rewritten to another block in the target zone. To address these issues, this article proposes a novel ZNS SSD design that features dynamic zone mapping, termed Brick-ZNS . Brick-ZNS implements two key functionalities: in-storage data migration and address remapping. New ZNS commands are first designed to realize in-storage data migration to avoid the end-to-end transfer overhead of GC while ensuring performance predictability. Then, a remapping strategy exploiting parallel physical blocks is proposed to reduce the large block-to-block rewrite overhead while ensuring zone-level access parallelism. The basic idea of the strategy is to directly remap the parallel physical blocks with a sufficient amount of valid data in the victim zone to the target zone, hence avoiding the large block-to-block rewrite overhead. Based on a full-stack SSD emulator, the evaluation results show that Brick-ZNS improves write throughput by 25% and SSD lifetime by 1.41×.
Zhenhua Tan, Linbo Long, Jingcheng Shen, Renping Liu 0002, Congming Gao, Kan Zhong
ACM Trans. Archit. Code Optim.5
2024 Extremely-Compressed SSDs with I/O Behavior Prediction
abstract
As the data volume continues to grow exponentially, there is an increasing demand for large storage system capacity. Data compression techniques effectively reduce the volume of written data, enhancing space efficiency. As a result, many modern SSDs have already incorporated data compression capabilities. However, data compression introduces additional processing overhead in critical I/O paths, potentially affecting system performance. Currently, most compression solutions in flash-based storage systems employ fixed compression algorithms for all incoming data without leveraging differences among various data access patterns. This leads to sub-optimal compression efficiency. This article proposes a data-type-aware Flash Translation Layer (DAFTL) scheme to maximize space efficiency without compromising system performance. First, we propose an I/O behavior prediction method to forecast future access on specific data. Then, DAFTL matches data types with distinct I/O behaviors to compression algorithms of varying intensities, achieving an optimal balance between performance and space efficiency. Specifically, it employs higher-intensity compression algorithms for less frequently accessed data to maximize space efficiency. For frequently accessed data, it utilizes lower-intensity but faster compression algorithms to maintain system performance. Finally, an improved compact compression method is proposed to effectively eliminate page fragmentation and further enhance space efficiency. Extensive evaluations using a variety of real-world workloads, as well as the workloads with real data we collected on our platforms, demonstrate that DAFTL achieves more data reductions than other approaches. When compared to the state-of-the-art compression schemes, DAFTL reduces the total number of pages written to the SSD by an average of 8%, 21.3%, and 25.6% for data with high, medium, and low compressibility, respectively. In the case of workloads with real data, DAFTL achieves an average reduction of 10.4% in the total number of pages written to SSD. Furthermore, DAFTL exhibits comparable or even improved read and write performance compared to other solutions.
Xiangyu Yao, Qiao Li 0001, Kaihuan Lin, Xinbiao Gan, Jie Zhang 0048, Congming Gao, Zhirong Shen, Quanqing Xu, Chuanhui Yang, Chun Jason Xue
ACM Trans. Storage6
2024 Design and Implementation of Deduplication on F2FS
abstract
Data deduplication technology has gained popularity in modern file systems due to its ability to eliminate redundant writes and improve storage space efficiency. In recent years, the flash-friendly file system (F2FS) has been widely adopted in flash memory-based storage devices, including smartphones, fast-speed servers, and Internet of Things. In this article, we propose F2DFS (deduplication-based F2FS), which introduces three main design contributions. First, F2DFS integrates inline and offline hybrid deduplication. Inline deduplication eliminates redundant writes and enhances flash device endurance, while offline deduplication mitigates the negative I/O performance impact and saves more storage space. Second, F2DFS follows the file system coupling design principle, effectively leveraging the potentials and benefits of both deduplication and native F2FS. Also, with the aid of this principle, F2DFS achieves high-performance and space-efficient incremental deduplication. Third, F2DFS adopts virtual indexing to mitigate deduplication-induced many-to-one mapping updates during the segment cleaning. We conducted comprehensive experimental comparisons between F2DFS, native F2FS, and other state-of-the-art deduplication schemes, using both synthetic and real-world workloads. For inline deduplication, F2DFS outperforms SmartDedup, Dmdedup, and ZFS, in terms of both I/O bandwidth performance and deduplication rates. And for offline deduplication, compared to SmartDedup, XFS, and BtrFS, F2DFS shows higher execution efficiency, lower resource usage, and greater storage space savings. Moreover, F2DFS demonstrates more efficient segment cleanings than native F2FS.
Tiangmeng Zhang, Renhui Chen, Zijing Li, Congming Gao, Chengke Wang, Jiwu Shu
ACM Trans. Storage4
2023 RLAlloc: A Deep Reinforcement Learning-Assisted Resource Allocation Framework for Enhanced Both I/O Throughput and QoS Performance of Multi-Streamed SSDs
abstract
Multi-streamed Solid-State Disks (SSDs) have attracted increasing adoption in modern flash storage devices. Despite their excellent promise, effective flash resource allocation is still limiting both their achievable I/O performance and practical implementation. To this end, we develop the first-of-its-kind framework dubbed RLAlloc, which for the first time demonstrates deep Reinforcement Learning-assisted resource Allocation for boosting both I/O throughput and QoS performance of multi-streamed SSDs. Extensive experiments consistently validate the effectiveness of RLAlloc, improving up to 39.9% on I/O throughput and 44.0% on QoS performance over the state-of-the-art competitors.
Mengquan Li, Chao Wu 0006, Congming Gao, Cheng Ji 0002, Kenli Li 0001
DAC3
2023 Optimizing Data Migration for Garbage Collection in ZNS SSDs
abstract
ZNS SSDs shift the responsibility of garbage collection (GC) to the host. However, data migration in GC needs to move data to the host's buffer first and write back to the new location, resulting in an unnecessary end-to-end transfer overhead. Moreover, due to the pre-configured mapping between zones and blocks, GC needs to perform a large number of unnecessary block-to-block data migrations between zones. To address these issues, this paper proposes a simple and efficient data migration method, called IS-AR, with in-storage data migration and address remapping. Based on a full-stack SSD emulator, our evaluation shows that IS-AR reduces GC latency by 6.78× and improves SSD lifetime by 1.17× on average.
Zhenhua Tan, Linbo Long, Renping Liu 0002, Congming Gao
DATE4
2023 MGC: Multiple-Gray-Code for 3D NAND Flash based High-Density SSDs
abstract
QLC (4-bit-per-cell) and more-bit-per-cell 3D NAND flash memories are increasingly adopted in large storage systems. While achieving significant cost reduction, these memories face degraded performance and reliability issues. The industry has adopted two-step programming (TSP), rather than one-step programming, to perform fine-granularity program control and choose gray-code encoding, as well as LDPC (Low-Density Parity-Check Code) for error correction. Different flash manufacturers often integrate different gray-codes in their products, which exhibit different performance and reliability characteristics. Unfortunately, a fixed gray-code encoding design lacks the ability to meet the dynamic read and program performance requirements at both application and device levels.In this paper, we propose MGC, a multiple-gray-code encoding strategy, that adaptively chooses the best gray-code to meet the optimization goals at runtime. In particular, MGC first extracts the performance and reliability requirements based on application-level access patterns and detects the reliability degree of SSD. It then determines the appropriate gray-code to encode the data, either from host/user application or due to garbage collection, before writing the pages to the flash memory. MGC is integrated in FTL (flash translation layer) and enhances the flash controller to enable runtime gray-code arbitration. We evaluate the proposed MGC scheme. The results show that MGC achieves better performance and lifetime guarantee compared with state-of-the-arts and introduces little overhead.
Yina Lv, Liang Shi 0001, Qiao Li 0001, Congming Gao, Yunpeng Song, Longfei Luo, Youtao Zhang
HPCA4
2023 ADAR: Application-Specific Data Allocation and Reprogramming Optimization for 3-D TLC Flash Memory
abstract
High bit-density flash memories, such as triple-level cell (TLC) and quad-level cell (QLC), have been widely used in flash memory-based storage systems, offering significantly high capacity. However, these high bit-density flash memories suffer from asymmetric access performance on the different pages that sharing the same physical cells. Meanwhile, 3-D flash memory adopts stacking technology to increase capacity and reduce cost per bit. The flash unit can be reprogrammed many times as long as the voltage increases. The reprogramming technology is also an effective solution for further increasing the 3-D flash capacity, allowing multiple program operations in an erase cycle. Considering the restrictions of reprogram operations, solid-state drives (SSDs) should capture the access pattern to perform more reprogramming operations to realize the joint optimization of read and write performance. In this work, we propose an application-specific data allocation and reprogramming technique named ADAR to enhance the read and write performance of 3-D TLC flash memory-based SSDs. The core idea is to allocate low-latency least significant bit (LSB) and central significant bit pages to frequently updated write data (termed hot write data) to improve the write performance, and reprogram the pages from high-latency pages (e.g., most significant bit page) to low-latency pages (e.g., LSB page mode) to enhance the read performance while initially storing frequently read data (termed hot read data) in high-latency pages. We explored data access patterns and designed an effective hotness identification method to present a new data allocation and reprogramming technique for 3-D TLC flash memory. Based on a modified 3-D TLC SSD simulator with typical workloads, our evaluation showed that our technique achieved 35.36% and 25.72% performance improvements in read and write latencies, respectively.
Linbo Long, Jinpeng Huang, Congming Gao, Duo Liu 0002, Renping Liu 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Stop unnecessary refreshing: extending 3D NAND flash lifetime with ORBER
Qiao Li 0001, Congming Gao, Shun Deng, Tei-Wei Kuo, Chun Jason Xue
CCF Trans. High Perform. Comput.3
2022 Reprogramming 3D TLC Flash Memory based Solid State Drives
abstract
NAND flash memory-based SSDs have been widely adopted. The scaling of SSD has evolved from plannar (2D) to 3D stacking. For reliability and other reasons, the technology node in 3D NAND SSD is larger than in 2D, but data density can be increased via increasing bit-per-cell. In this work, we develop a novel reprogramming scheme for TLCs in 3D NAND SSD, such that a cell can be programmed and reprogrammed several times before it is erased. Such reprogramming can improve the endurance of a cell and the speed of programming, and increase the amount of bits written in a cell per program/erase cycle, i.e., effective capacity. Our work is the first to perform a real 3D NAND SSD test to validate the feasibility of the reprogram operation. From the collected data, we derive the restrictions of performing reprogramming due to reliability challenges. Furthermore, a reprogrammable SSD (ReSSD) is designed to structure reprogram operations. ReSSD is evaluated in a case study in RAID 5 system (RSS-RAID). Experimental results show that RSS-RAID can improve the endurance by 35.7%, boost write performance by 15.9%, and increase effective capacity by 7.71%, with negligible overhead compared with conventional 3D SSD-based RAID 5 system.
Congming Gao, Chun Jason Xue, Youtao Zhang, Liang Shi 0001, Jiwu Shu, Jun Yang 0002
ACM Trans. Storage1
2021 Pattern-Guided File Compression with User-Experience Enhancement for Log-Structured File System on Mobile Devices
Cheng Ji 0002, Li-Pin Chang, Riwei Pan, Chao Wu 0006, Congming Gao, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue
FAST5
2021 ParaBit: Processing Parallel Bitwise Operations in NAND Flash Memory based SSDs
abstract
Processing-in-memory (PIM) and in-storage-computing (ISC) architectures have been constructed to implement computation inside memory and near storage, respectively. While effectively mitigating the overhead of data movement from memory and storage to the processor, due to the limited bandwidth of existing systems, these architectures still suffer from the large data movement overhead between storage and memory, in particular, if the amount of required data is large. It has become a major constraint for further improving the computation efficiency in PIM and ISC architectures.
Congming Gao, Xin Xin 0008, Youyou Lu, Youtao Zhang, Jun Yang 0002, Jiwu Shu
MICRO1
2020 Maximizing I/O Throughput and Minimizing Performance Variation via Reinforcement Learning Based I/O Merging for SSDs
abstract
Merging technique is widely adopted by I/O schedulers to maximize system I/O throughput. However, I/O merging could increase the latency of individual I/O, thus incurring prolonged I/O latencies and enlarged performance variations. Even with better system throughput, higher worst-case latency experienced by some requests could block the SSD storage system, which violates the QoS (Quality of Service) requirement. In order to improve QoS performance while providing higher I/O throughput, this paper proposes a reinforcement learning based I/O merging approach. Through learning the characteristic of various I/O patterns, the proposed approach makes merging decisions adaptively based on different I/O workloads. Evaluation results show that the proposed scheme is capable of reducing the standard deviation of I/O latency by 19.1 percent on average, worst-case latency by 7.3-60.9 percent at the 99.9th percentile compared with the latest I/O merging scheme, while maximizing system throughput.
Chao Wu 0006, Cheng Ji 0002, Qiao Li 0001, Congming Gao, Riwei Pan, Chenchen Fu, Liang Shi 0001, Chun Jason Xue
IEEE Trans. Computers4
2020 Aging Capacitor Supported Cache Management Scheme for Solid-State Drives
abstract
Solid-state drives (SSDs) have been widely adopted in embedded systems, data centers, and cloud storage due to its well-identified advantages. Inside SSD, random access memory (RAM) is adopted as the built-in cache for achieving better performance. However, due to the volatility characteristic of RAM, data loss may happen when sudden power interrupts. In order to solve this issue, a capacitor has been equipped inside emerging SSDs as an interim power supplier. But due to the capacitor aging issue, which will result in capacitance decreases over time, there still may exist data loss when power interruption occurs. Once the remaining capacitance drops to the threshold value where all dirty pages in the cache can not be written back to flash memory, data loss happens. To solve the above issue, an efficient cache management scheme for capacitor equipped SSDs is proposed in this article. The basic idea of this scheme is to bound the number of dirty pages in a cache within the capability of the equipped capacitor. The proposed scheme includes three steps: 1) a periodical dirty page budget detection (DPBD) scheme is proposed to acquire the maximal number of dirty pages that can be written back within current capability of equipped capacitor; 2) a smart dirty page synchronizing scheme is proposed during normal run time to bound the number of dirty pages in the cache; and 3) when power supply interrupts, an efficient writing back method is applied to further reduce the capacitance consumption of capacitor. The simulation results show that the proposed scheme achieves encouraging improvement on lifetime and performance while power interruption induced data loss is avoided.
Congming Gao, Liang Shi 0001, Qiao Li 0001, Kai Liu 0001, Chun Jason Xue, Jun Yang 0002, Youtao Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Boosting the Performance of SSDs via Fully Exploiting the Plane Level Parallelism
abstract
Solid state drives (SSDs) are constructed with multiple level parallel organization, including channels, chips, dies, and planes. Among these parallel levels, plane level parallelism, which is the last level parallelism of SSDs, has the most strict restrictions. Only the same type of operations that access the same address in different planes can be processed in parallel. In order to maximize the access performance, several previous works have been proposed to exploit the plane level parallelism for host accesses and internal operations of SSDs. However, our preliminary studies show that the plane level parallelism is farfrom well utilized and should be further improved. The reason is that the strict restrictions of plane level parallelism are hard to be satisfied. In this article, a from plane to die parallel optimization framework is proposed to exploit the plane level parallelism through smartly satisfying the strict restrictions all the time. In order to achieve the objective, there are at least two challenges. First, due to that host access patterns are always complex, receiving multiple same-type requests to different planes at the same time is uncommon. Second, there are many internal activities, such as garbage collection (GC), which may destroy the restrictions. In order to solve above challenges, two schemes are proposed in the SSD controller: First, a die level write construction scheme is designed to make sure there are always N pages of data written by each write operation. Second, in a further step, a die level GC scheme is proposed to activate GC in the unit of all planes in the same die. Combing the die level write and die level GC, write accesses from both host write operations and GC induced valid page movements can be processed in parallel at all time. To further improve the performance of SSDs, host write operations blocked by GCs are suggested to be processed in parallel with GC induced valid page movements, bringing lesser waiting time cost of host write operations. As a result, the GC cost and average write latency can be significantly reduced. Experiment results show that the proposed framework is able to significantly improve the write performance without read performance impact.
Congming Gao, Liang Shi 0001, Kai Liu 0001, Chun Jason Xue, Jun Yang 0002, Youtao Zhang
IEEE Trans. Parallel Distributed Syst.1
2020 Process Variation Aware Read Performance Improvement for LDPC-Based nand Flash Memory
abstract
With the rapid development of technology scaling and cell density improvement for capacity increase and cost reduction, nand flash memory is confronted with degraded reliability. On one hand, while low-density parity-check (LDPC) codes have been deployed in today's nand flash memories to enhance reliability, flash read latency has still been a performance bottleneck with the increased raw bit error rates (RBER). On the other hand, significant process variations (PV) have been found on existing nand flash memories, which introduce great reliability variations among different flash blocks. Recent studies have proposed to exploit PV to improve endurance by better wear leveling or to improve write performance. These approaches are prone to allocate read data to blocks with low reliability, which further degrades read performance. This paper proposes to enhance read performance of LDPC-equipped nand flash memory by exploiting the reliability variations from PV. The paper consists of three parts. First, a block grouping approach is presented to categorize flash blocks according to their reliability. Second, according to the grouping scheme, a data placement scheme is proposed, which allocates read-hot data to flash blocks with high reliability. At the same time, the read-cold data is moved to blocks with low reliability. As a result, the read performance is enhanced. However, allocating high reliable blocks for read-hot data collides with previous PV-based wear leveling methods. To address the issue, the third part is a grouping partition scheme which limits the amount of high reliable blocks occupied by read-hot data. Therefore, read performance enhancement can be achieved and the wear leveling schemes will be impacted slightly. Experiment results present that, the proposed approach can provide significant read performance improvement on LDPC-equipped nand flash memory and is compatible with the previous PV-based wear leveling.
Qiao Li 0001, Liang Shi 0001, Yejia Di, Congming Gao, Cheng Ji 0002, Yu Liang 0004, Chun Jason Xue
IEEE Trans. Reliab.4
2019 Constructing Large, Durable and Fast SSD System via Reprogramming 3D TLC Flash Memory
abstract
NAND flash memory based SSDs have been widely studied and adopted. The scaling of SSD has evolved from plannar (2D) to 3D stacking. Compared with 2D SSD, 3D SSD stacks more layers into one block, constructing one block with more flash pages. For reliability and other reasons, technology node in 3D NAND SSD is larger than in 2D, but data density can be increased via increasing bit-per-cell. However, representing multiple bits per cell encounters additional challenges such as endurance and access latency. In this work, we develop a novel reprogramming scheme for TLCs in 3D NAND SSD, such that a cell can be programmed and reprogrammed several times before it is erased. Such reprogramming can reduce the frequency of erases which determines the endurance of a cell, improve the speed of programming, and increase the amount of bits written in a cell per program/erase cycle, i.e., effective capacity. Our work is the first to perform real 3D NAND SSD test to validate the feasibility of the reprogram operation. From the collected data, we derive the restrictions of performing reprogramming due to reliability challenges. Further, a reprogrammable SSD (ReSSD) is designed to structure reprogram operations, and when they should be applied. ReSSD is evaluated in a case study in 3D TLC SSD based RAID 5 system (RSS-RAID). Experimental results show that RSS-RAID can improve the endurance by 30.3%, boost write performance by 16.7%, and increase effective capacity by 7.71%, with negligible overhead compared with conventional 3D SSD based RAID 5 system.
Congming Gao, Qiao Li 0001, Chun Jason Xue, Youtao Zhang, Liang Shi 0001, Jun Yang 0002
MICRO1
2019 Parallel all the time: Plane Level Parallelism Exploration for High Performance SSDs
abstract
Solid state drives (SSDs) are constructed with multiple level parallel organization, including channels, chips, dies and planes. Among these parallel levels, plane level parallelism, which is the last level parallelism of SSDs, has the most strict restrictions. Only the same type of operations which access the same address in different planes can be processed in parallel. In order to maximize the access performance, several previous works have been proposed to exploit the plane level parallelism for host accesses and internal operations of SSDs. However, our preliminary studies show that the plane level parallelism is far from well utilized and should be further improved. The reason is that the strict restrictions of plane level parallelism are hard to be satisfied. In this work, a from plane to die parallel optimization framework is proposed to exploit the plane level parallelism through smartly satisfying the strict restrictions all the time. In order to achieve the objective, there are at least two challenges. First, due to that host access patterns are always complex, receiving multiple same-type requests to different planes at the same time is uncommon. Second, there are many internal activities, such as garbage collection (GC), which may destroy the restrictions. In order to solve above challenges, two schemes are proposed in the SSD controller: First, a die level write construction scheme is designed to make sure there are always N pages of data written by each write operation. Second, in a further step, a die level GC scheme is proposed to activate GC in the unit of all planes in the same die. Combing the die level write and die level GC, write accesses from both host write operations and GC induced valid page movements can be processed in parallel at all time. As a result, the GC cost and average write latency can be significantly reduced. Experiment results show that the proposed framework is able to significantly improve the write performance without read performance impact.
Congming Gao, Liang Shi 0001, Chun Jason Xue, Cheng Ji 0002, Jun Yang 0002, Youtao Zhang
MSST1
2019 Optimizing Tail Latency of LDPC based Flash Memory Storage Systems Via Smart Refresh
abstract
Flash memory has been developed with bit density improvement, technology scaling, and 3D stacking. With this trend, its reliability has been degraded significantly. Error correction code, low density parity code (LDPC), which has strong error correction capability, has been employed to solve this issue. However, one of the critical issues of LDPC is that it would introduce a long decoding latency on devices with low reliability. In this case, tail latency would happen, which will significantly impact the quality of service (QoS). In this work, a set of smart refresh schemes is proposed to optimize the tail latency. The basic idea of the work is to refresh data when the accessed data has a long decoding latency. Two smart refresh schemes are proposed for this work: The first refresh scheme is designed to refresh long access latency data when it is accessed several times for access performance optimization; The second refresh scheme is designed to periodical detecting data with extremely long access latency and refreshing them for tail latency optimization. Experiment results show that the proposed schemes are able to significantly improve the tail latency and access performance with little overhead.
Yina Lv, Liang Shi 0001, Qiao Li 0001, Congming Gao, Chun Jason Xue, Edwin H.-M. Sha
NAS4
2019 Minimizing Retention Induced Refresh Through Exploiting Process Variation of Flash Memory
abstract
Refresh schemes have been the default approach in NAND flash memory to avoid data losses. The critical issue of the refresh schemes is that they introduce additional costs on lifetime and performance. Recent work proposed to minimize the refresh costs by using uniform refresh frequencies based on the number of program/erase (P/E) cycles. However, from our investigation, we find that the refresh costs still have a high burden on the lifetime performance. In this paper, a novel refresh minimization scheme is proposed by exploiting the process variation (PV) of flash memory. State-of-the-art flash memory always has significant PV, which introduces large variations on the retention time of flash blocks. In order to reduce the refresh costs, we first propose a new refresh frequency determination scheme by detecting the supported retention time of flash blocks. If the detected retention time is large, a low refresh frequency can be applied to minimize the refresh costs. Second, considering that the retention time requirements of data are varied with each others, we further propose a data hotness and refresh frequency matching scheme. The matching scheme is designed to allocate data to blocks with right higher supported retention time. Through simulation studies, the lifetime and performance are significantly improved compared with state-of-the-art refresh schemes.
Yejia Di, Liang Shi 0001, Congming Gao, Qiao Li 0001, Chun Jason Xue, Kaijie Wu 0001
IEEE Trans. Computers3
2018 Loss is Gain: Shortening Data for Lifetime Improvement on Low-Cost ECC Enabled Consumer-Level Flash Memory
abstract
Reliability has been a challenge in the development of NAND flash memory, due to its technology size scaling and bit density improvement. To ensure the data integrity, error correction codes (ECC) with high error correction capability have been suggested. However, much higher costs will be introduced which cannot be supported for cost-limited consumer-level flash memory. Thus, low-cost ECCs are usually applied. In this work, a reliability improvement scheme is proposed for low-cost ECC enabled consumer-level flash memory. The scheme is motivated by the finding that low-cost ECC is able to protect shortened encoded data with improved reliability. This is because that the less the encoded data are, the less the errors will be occurred. With this motivation, a design is proposed to construct the shortened data case for a low-cost ECC when it cannot be able to provide the reliability requirement. Second, two relaxation approaches are proposed to relax the space reduction as it has bad effects on flash memory. A model guided evaluation is finally presented, and the results show that the lifetime can be significantly improved with little space reduction.
Yejia Di, Liang Shi 0001, Congming Gao, Qiao Li 0001, Kaijie Wu 0001, Chun Jason Xue
ACM Great Lakes Symposium on VLSI3
2018 An Efficient Cache Management Scheme for Capacitor Equipped Solid State Drives
abstract
Within SSDs, random access memory (RAM) has been adopted as cache inside controller for achieving better performance. However, due to the volatility characteristic of RAM, data loss may happen when sudden power interrupts. To solve this issue, capacitor has been equipped inside emerging SSDs as interim supplier. However, the aging issue of capacitor will result in capacitance decreases over time. Once the remaining capacitance is not able to write all dirty pages in the cache back to flash memory, data loss may happen. In order to solve the above issue, an efficient cache management scheme for capacitor equipped SSDs is proposed in this work. The basic idea of the scheme is to bound the number of dirty pages in cache within the capability of the capacitor. Simulation results show that the proposed scheme achieves encourage improvement on lifetime and performance while power interruption induced data loss is avoided.
Congming Gao, Liang Shi 0001, Yejia Di, Qiao Li 0001, Chun Jason Xue, Edwin H.-M. Sha
ACM Great Lakes Symposium on VLSI1
2018 Access Characteristic Guided Read and Write Regulation on Flash Based Storage Systems
abstract
NAND flash memory is now used in various storage systems, such as embedded systems, personal computers, and web servers. The developments in bit density and technology scaling have reduced its price, but worsen the reliability, leading to shortened lifetime and degraded access performance. This paper proposes to exploit access characteristics of workloads to improve flash performance and lifetime. The basic idea is to regulate the read and write operations based on the identified access characteristics. First, an access cost model is presented, which indicates a tradeoff between read and write time cost on NAND flash memory. Based on the access characteristics of workloads, read-only pages will be written with high cost so that they can be read with low cost, and write-only pages will be written with low cost. Second, the tradeoff between read cost and flash wearing is exploited for lifetime improvement. The write requests on write-only data are processed with reduced wearing by regulating the program threshold voltage. Finally, as these approaches apply different write operations on write-only data for performance and lifetime improvement respectively, a combined approach is proposed to satisfy both goals. Simulation results show that the proposed approaches can improve performance and lifetime significantly with negligible overhead.
Qiao Li 0001, Liang Shi 0001, Congming Gao, Yejia Di, Chun Jason Xue
IEEE Trans. Computers3
2018 Exploiting Parallelism for Access Conflict Minimization in Flash-Based Solid State Drives
abstract
Solid state drives (SSDs) have been widely deployed in personal computers, data centers, and cloud storages. In order to improve performance, SSDs are usually constructed with a number of channels with each channel connecting to a number of nand flash chips, each flash chip consisting of multiple dies and each die containing multiple planes. Based on this parallel architecture, I/O requests are potentially able to access parallel units simultaneously. Despite the rich parallelism offered by the parallel architecture, recent studies show that the utilization of flash parallel units is seriously low. This paper shows that the low parallel unit utilization is highly caused by the access conflict among I/O requests. In this paper, we propose parallel issue queueing (PIQ), a novel I/O scheduler at the host systems. PIQ groups I/O requests without conflicts into the same batch and I/O requests with conflicts into different batches. Hence, the multiple I/O requests in one batch can be fulfilled simultaneously by exploiting the rich parallelism of SSDs. Extensive experimental results show that PIQ delivers significant performance improvement especially for the applications which have heavy access conflicts.
Congming Gao, Liang Shi 0001, Cheng Ji 0002, Yejia Di, Kaijie Wu 0001, Chun Jason Xue, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Exploiting Chip Idleness for Minimizing Garbage Collection - Induced Chip Access Conflict on SSDs
abstract
Solid state drives (SSDs) are normally constructed with a number of parallel-accessible flash chips, where host I/O requests are processed in parallel. In addition, there are many internal activities in SSDs, such as garbage collection and wear leveling induced read, write, and erase operations, to solve the issues of inability of in-place updates and limited lifetime. When internal activities are triggered on a chip, the chip will be blocked. Our preliminary studies on several workloads show that when internal activities are frequently triggered, the host I/O performance will be significantly impacted because of the access conflict between them. In this work, in order to improve the access conflict induced performance degradation, a novel access conflict minimization scheme is proposed. The basic idea of the scheme is motivated by an interesting observation in SSDs: several chips are idle when other chips are busy with internal activities and host I/O requests. Based on this observation, we propose to schedule internal activities induced operations for minimized access conflict by exploiting the idleness of the multiple chips of SSDs. This approach is realized by two steps: First, read internal activities accessed data to the controller; second, by exploiting the idle chips during internal activities, write internal activities accessed data back to these idle chips. With this scheme, the internal activities can be processed with minimized access conflict to the host requests. Simulation results show that the proposed approach significantly reduces the access conflict, and in turn leads to a significant performance improvement of SSDs.
Congming Gao, Liang Shi 0001, Yejia Di, Qiao Li 0001, Chun Jason Xue, Kaijie Wu 0001, Edwin H.-M. Sha
ACM Trans. Design Autom. Electr. Syst.1
2017 Asymmetric Error Rates of Cell States Exploration for Performance Improvement on Flash Memory Based Storage Systems
abstract
Recent studies show that a multilevel cell flash cell in different states suffers from diverse error patterns in varying degrees. That is, the error rates of each page are highly dependent on the data content. Consequently, pages with different data will exhibit quite different error rates. However, existing technologies equipped with one uniform error correction code (ECC) scheme for all pages in a flash memory do not take the different error rates of pages into consideration. In this paper, we propose to exploit the asymmetric error rates of flash memory exhibited by the flash pages with different data for performance improvement. Before a page is programmed, its specific error rates, called content-dependent bit error rates (CDBERs), are estimated according to the content of the page. The margin between the CDBER of a page and the maximal error rates correctable by the uniform ECC code is exploited for performance improvement. On one hand, a faster and suitable write operation is selected to speed up the progress of programming while the increased speed induced CDBER does not exceed the maximal correctable error rates. On the other hand, a light-weight ECC scheme can be chosen for a faster read operation since the page decoding process of a light-weight ECC scheme incurs less time overhead. Finally, a state mapping scheme, which further reduces the CDBER through mapping high error rate states to the low error rate states of a page, is proposed. Simulation results show that the proposed approaches lead to significant write and read performance improvement.
Edwin H.-M. Sha, Congming Gao, Liang Shi 0001, Kaijie Wu 0001, Mengying Zhao, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Lightweight Data Compression for Mobile Flash Storage
abstract
Data compression is beneficial to flash storage lifespan. However, because the design of mobile flash storage is highly cost-sensitive, hardware compression becomes a less attractive option. This study investigates the feasibility of data compression on mobile flash storage. It first characterizes data compressibility based on mobile apps, and the analysis shows that write traffic bound for mobile storage volumes is highly compressible. Based on this finding, a lightweight approach is introduced for firmware-based data compression in mobile flash storage. The controller and flash module work in a pipelined fashion to hide the data compression overhead. Together with this pipelined design, the proposed approach selectively compresses incoming data of high compressibility, while leaving data of low compressibility to a compression-aware garbage collector. Experimental results show that our approach greatly reduced the frequency of block erase by 50.5% compared to uncompressed flash storage. Compared to unconditional data compression, our approach improved the write latency by 10.4% at a marginal cost of 4% more block erase operations.
Cheng Ji 0002, Li-Pin Chang, Liang Shi 0001, Congming Gao, Chao Wu 0006, Yuangang Wang, Chun Jason Xue
ACM Trans. Embed. Comput. Syst.4
2015 Maximizing IO performance via conflict reduction for flash memory storage systems
Qiao Li 0001, Liang Shi 0001, Congming Gao, Kaijie Wu 0001, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha
DATE3
2014 Exploit asymmetric error rates of cell states to improve the performance of flash memory storage systems
abstract
The reliability of flash memory is getting worse with the introduction of Multiple Level Cell (MLC) and Triple Level Cell (TLC) technologies. To account for possible errors, each page in a flash memory is equipped with an Error Correction Code (ECC) module. An ECC scheme is chosen according to the worst-case error occurrences across all pages in the flash memory. Recent studies show that an MLC flash cell in different states exhibits diverse error rates and the difference is dramatic. Consequently, pages with different data will exhibit quite different error rates. Existing technologies that use one uniform ECC scheme for all pages in a flash memory is far from optimal. This paper exploits the asymmetric error rates exhibited by the pages with different data for write performance improvement. Before a page is programmed, its specific error rate, called Content-Dependent Bit Error Rate (CDBER), is estimated according to the content of the page. The margin between the CDBER of a page and the maximal error rate correctable by the uniform ECC code is exploited for write performance improvement. Simulation results show that the proposed approach leads to significant write performance improvement.
Congming Gao, Liang Shi 0001, Kaijie Wu 0001, Chun Jason Xue, Edwin H.-M. Sha
ICCD1
2014 Exploiting parallelism in I/O scheduling for access conflict minimization in flash-based solid state drives
abstract
Solid state drives (SSDs) have been widely deployed in personal computers, data centers, and cloud storages. In order to improve performance, SSDs are usually constructed with a number of channels with each channel connecting to a number of NAND flash chips. Despite the rich parallelism offered by multiple channels and multiple chips per channel, recent studies show that the utilization of flash chips (i.e. the number of flash chips being accessed simultaneously) is seriously low. Our study shows that the low chip utilization is caused by the access conflict among I/O requests. In this work, we propose Parallel Issue Queuing (PIQ), a novel I/O scheduler at the host system, to minimize the access conflicts between I/O requests. The proposed PIQ schedules I/O requests without conflicts into the same batch and I/O requests with conflicts into different batches. Hence the multiple I/O requests in one batch can be fulfilled simultaneously by exploiting the rich parallelism of SSD. And because PIQ is implemented at the host side, it can take advantage of rich resource at host system such as main memory and CPU, which makes the overhead negligible. Extensive experimental results show that PIQ delivers significant performance improvement to the applications that have heavy access conflicts.
Congming Gao, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Kaijie Wu 0001, Edwin H.-M. Sha
MSST1