You Zhou 0009

dblp:20/2165-9 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 6 first-author · 21 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 APT: Securing Against DRAM Read Disturbance via Adaptive Probabilistic In-DRAM Trackers
abstract
With exacerbated DRAM read disturbance, transparent in-DRAM defenses require space (to track the aggressor rows) and time (to perform more mitigations). To reduce storage overhead, recent works have developed probabilistic row-sampling techniques with several entries. To address the time issue, these techniques employ Refresh Management (RFM) commands introduced in DDR5. However, probabilistic defenses with RFM face two critical challenges: (i) fixed-probability sampling under dynamic activation patterns can cause row-sampling misses, allowing attacks to evade mitigation, and (ii) RFM causes timing variations that can be exploited for side and covert channels to leak sensitive information. The goal of this paper is to design a low-cost and secure in-DRAM defense that overcomes these challenges.
Runjin Wu, Meng Zhang 0014, You Zhou 0009, Changsheng Xie 0001, Fei Wu 0005
ASPLOS (2)3
2026 CEMU: Enabling Full-System Emulation of Computational Storage Beyond Hardware Limits
abstract
Computational storage drives (CSDs) present a promising approach to improve system performance through near data processing in SSDs. However, current research platforms are fragmented and inadequate to explore the full design space of CSD systems. Existing hardware and emulator platforms are constrained by physical compute resources, while simulators lack full-system fidelity. To address the problems, we introduce CEMU, a new software-based CSD emulation platform that enables full-system research. It consists of a CSD device emulator and a CSD-oriented software stack. Through a novel virtual machine freezing mechanism, CSD emulation achieves high configurability. While the CSD can utilize the host CPU to physically perform computation to preserve full-system behaviors, the computational delay can be modeled separately to emulate CSDs with CPU-unbounded high computing power. The software stack is designed with two principles, adhering to recent industry CSD standards and being compatible with the existing I/O stack, which is achieved via a newly developed file system FDMFS. We verify CEMU's emulation fidelity across a range of applications by benchmarking against actual CSD hardware, demonstrating average end-to-end performance accuracy of 95% or higher. We also use two case studies on large language model training and LevelDB to demonstrate that CEMU is effective in exploring CSD system research and can uncover insights that have not been discovered in previous research platforms.
Jiapin Wang, You Zhou 0009, Kai Lu 0002, Jiguang Wan 0001, Fei Wu 0005, Tao Lu 0014
ASPLOS (2)3
2026 Crash-Consistent SSD Array With Hardware-Guaranteed Transactional Atomicity
abstract
Crash consistency is a critical challenge that storage systems need to address carefully in multiple layers, including the database, file system, and block-layer RAID. Software approaches typically resort to logging to ensure transactional atomicity, which causes significant performance overhead and write amplification over underlying high-speed SSDs. Motivated by the out-of-place update feature of SSD-internalflash translation layer (FTL), previous studies propose crash-consistent SSDs to offload transactional atomicity guarantee and demonstrate their effectiveness in eliminating software-based logging overheads. However, existing hardware offloading approaches only consider single-SSD systems and would fail in an SSD array. This paper presents a crash-consistent SSD array, using FTLs and coordinating multiple SSDs to provide transactional atomicity across the array. The key is to design anarray-wide transaction commit (ARC)protocol, which resolves the multi-SSD coordination challenge and tolerates disk failures. We implement an ARC array manager and ARC SSDs to verify the design. Two case studies are conducted, where the ARC array is utilized to address the transaction logging overhead in the SQLite database and the stripe write-hole problem in the RAID subsystem. Experimental results demonstrate that the ARC SSD array can improve system performance by 32% to 93% and reduce write amplification by 39% to 49%, on average, compared with software-based logging approaches.
Zeyu Niu, Xiang Chen 0028, You Zhou 0009, Zibin Sun, Zhihu Tan, Changsheng Xie 0001, Fei Wu 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 StreamCSD: SSD-Autonomous Stream Management via In-Storage Content Learning
abstract
Write amplification (WA) from migrating valid pages during garbage collection (GC) degrades SSD performance and lifespan. Although stream management based on high-level software semantics reduces WA, existing solutions require host modifications, hindering their adoption. We introduce StreamCSD, an SSD-autonomous stream management approach using in-storage content learning, eliminating host-side changes. Leveraging compression ratios from embedded compressors in computational storage drives (CSDs), StreamCSD employs a streaming Kmeans algorithm to cost-efficiently cluster data into streams. Evaluations show that StreamCSD reduces WA from 1.7 to 1.06 under multimodal generative AI workloads, matching state-of-the-art methods with minimal impact on bandwidth. StreamCSD operates without host modifications, promoting broader adoption of multi-stream SSDs.
Xiang Chen 0028, Yelin Shan, Jiapin Wang, Yunxin Huang, Yafei Yang, Tao Lu 0014, You Zhou 0009, Fei Wu 0005
DAC8
2024 Eliminating Storage Management Overhead of Deduplication over SSD Arrays Through a Hardware/Software Co-Design
abstract
This paper presents a hardware/software co-design solution to efficiently implement block-layer deduplication over SSD arrays. By introducing complex and varying dependency over the entire storage space, deduplication is infamously subject to high storage management overheads in terms of CPU/memory resource usage and I/O performance degradation. To fundamentally address this problem, one intuitive idea is to offload deduplication storage management from host into SSDs, which is motivated by the redundant dual address mapping in host-side deduplication layer and intra-SSD flash translation layer (FTL). The practical implementation of this idea is nevertheless challenging because of the array-wide deduplication vs. per-SSD FTL management scope mismatch. Aiming to tackle this challenge, this paper presents a solution, called ARM-Dedup, that makes SSD FTL deduplication-oriented and array-aware and accordingly re-architects deduplication software to achieve lightweight and high-performance deduplication over an SSD array. We implemented an ARM-Dedup prototype based on the Linux Dmdedup engine and mdraid software RAID over FEMU SSD emulators. Experimental results show that ARM-Dedup has good scalability and can improve system performance significantly, such as by up to 272% and 127% higher IOPS in synthetic and real-world workloads, respectively.
Yuhong Wen, You Zhou 0009, Tong Zhang 0002, Shangjun Yang, Changsheng Xie 0001, Fei Wu 0005
ASPLOS (2)3
2024 Balloon-ZNS: Constructing High-Capacity and Low-Cost ZNS SSDs with Built-in Compression
abstract
ZNS SSDs are emerging storage devices promising low cost, high performance, and software definability. This paper proposes Balloon-ZNS that enables transparent compression in ZNS SSDs to enhance cost efficiency. ZNS SSDs, unlike traditional SSDs, require data pages to be stored in logical zones and flash blocks with aligned offsets, conflicting with the management of compressed, variable-length pages. Motivated by a key observation that compressibility locality widely exists in data streams, Balloon-ZNS employs a compressibility-adaptive, slot-aligned storage management scheme to address the intractable conflict. Evaluation with RocksDB shows Balloon-ZNS can reap more than 80% of the compression gain while achieving 7.3% lower to 14.4% higher throughput than a vanilla ZNS SSD, on average, when data compressibility is not poor.
Yu Wang 0168, Zibin Sun, You Zhou 0009, Tao Lu 0014, Changsheng Xie 0001, Fei Wu 0005
DAC3
2024 LaVA: An Effective Layer Variation Aware Bad Block Management for 3D CT NAND Flash
abstract
3D NAND flash with charge trap (CT) technology has been developed by stacking multiple layers vertically to boost storage capacity while ensuring reliability and scalability. One of its critical characteristics is the large endurance variation among and inside blocks and layers. With this feature, traditional bad block management (BBM), which determines block lifetime by the page with worst endurance, results in underutilization of solid state drive (SSD) usage. In this paper, a layer variation aware and fault-tolerant bad block management, named LaVA, is proposed to prolong the lifetime of 3D NAND flash storage. The relevant layer, instead of the entire flash block, is discarded at a finer granularity when a page failure is encountered. Experimental results based on real-world workloads show that LaVA can significantly extend the endurance of 3D CT NAND flash (30.6%-62.2%) with a small performance degradation (less than 10% increase of tail I/Oresponse time), compared to the conventional technique.
Shuhan Bai, You Zhou 0009, Fei Wu 0005, Changsheng Xie 0001, Tei-Wei Kuo, Chun Jason Xue
DATE2
2024 HA-CSD: Host and SSD Coordinated Compression for Capacity and Performance
abstract
Integrating data compression capability into SSDs has demonstrated great potential to improve the utilization and lifetime of the storage device and also the performance of the entire system. It is advocated to add a hardware engine into the SSD for low-latency compression and decompression. However, this requires a new and long hardware product development cycle, which would prevent current storage systems from reaping the benefits of in-SSD compression. In this paper, we explore a software-based in-SSD compression solution, which can be delivered to users quickly through a simple SSD firmware update. The most critical challenge is the severe performance bottleneck caused by compression and decompression, as the in-SSD embedded CPU has quite limited computing power. To tackle this challenge, we propose a host-assisted computational storage device, called HA-CSD. It employs an offline, data hotness- and compressibility-aware compression strategy to remove compression from the critical write I/O path. A novel decompression architecture is devised to utilize the powerful host CPU for fast decompression. We implement HA-CSD in a commercial enterprise SSD with a code change of more than 25K lines in the host NVMe driver and SSD firmware. Experimental results show that HA-CSD achieves 2.1GB/s and 5.2GB/s read and write bandwidth. Compared with RocksDB built-in compression, HA-CSD can increase the YCSB benchmark throughput by up to 5.7×, and improve the host CPU efficiency significantly.
Xiang Chen 0028, Tao Lu 0014, Jiapin Wang, Guangchun Xie, Xueming Cao, Yuanpeng Ma, Bing Si, Yunxin Huang, Yafei Yang, You Zhou 0009, Fei Wu 0005
IPDPS13
2024 TieredHM: Hotspot-Optimized Hash Indexing for Memory-Semantic SSD-Based Hybrid Memory
abstract
Memory semantic Solid State Drives (MS-SSDs) provide a promising opportunity to enable the hybrid memory architecture (HMA). The memory semantic interface enables the CPUs to directly access structured data in SSDs and eliminate bulk data copy/swap between the memory and storage devices. However, existing hash indexings issue many random writes, resulting in two problems when directly deployed on MS-SSD-based HMA: 1) Highly random traffic persisted to the underlying NAND flash of MS-SSDs incurs significant garbage collection (GC) overhead. 2) Placing frequently updated memory pages of hash indexings in persistent memories (PMs) is anticipated to reduce write latency, failing to work effectively due to the lack of skewness. To address the above problems, we propose a novel MS-SSD-friendly hash indexing scheme called TieredHM. It employs a multi-layer structure and opportunistic data movement (ODM) to construct skewed writes. Hence, the MS-SSD can transform the writes into multi-streamed writes, separating data with different update frequencies to reduce GC overhead. Besides, since the top layer is updated much more frequently (more skewed) than other layers, placing the top layer of TieredHM into persistent memory can significantly reduce write latency. TieredHM further leverages a prefetch mechanism based on the internal parallelism of NAND flash to reduce search overhead incurred by ODM. Experimental results show that TieredHM reduces the average write latency and GC overhead by up to 8.3X and 20.0X compared to state-of-the-art hash indexings without sacrificing read performance.
Weizhou Huang, Jian Zhou 0004, You Zhou 0009, Feng Zhu 0024, Kun Wang 0029, Fei Wu 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 FlexZNS: Building High-Performance ZNS SSDs with Size-Flexible and Parity-Protected Zones
abstract
NVMe zoned namespace (ZNS) SSDs present a new class of storage devices with attractive features including low cost, software definability, and stable performance. However, one primary culprit that hinders the adoption of ZNS is the high garbage collection (GC) overhead it brings to host software. The ZNS interface divides the logical address space into size-fixed zones that must be written sequentially. Despite being friendly to flash memory, ZNS requires host software to perform out-of-place updates and GC on individual zones. Current ZNS SSDs typically employ a large zone size (e.g., of GBs) to be conducive to die-level RAID protection on flash memory. This impedes flexible data placement, such as mixing data with different lifetimes in the same zone, and incurs sizable data migrations during zone GC. To address this problem, we propose FlexZNS, a novel ZNS SSD design that provides reliable zoned storage allowing host software to configure the zone size flexibly as well as multiple zone sizes. The size variability of zones poses two interrelated challenges, one for the SSD controller to establish per-zone RAID protection, and the other for host software to manage variable zone capacity loss caused by parity storage. To tackle the challenges, FlexZNS decouples the storage of parity from individual zones on flash memory and hides the zone capacity loss from the host software. We verify FlexZNS on a ZNS-compatible file system F2FS and a popular key-value store RocksDB. Extensive experiments demonstrate that FlexZNS can significantly improve the system performance and reduce GC-induced write amplification, compared with a conventional ZNS SSD with large-sized zones.
Yu Wang 0168, You Zhou 0009, Zhonghai Lu, Kun Wang 0029, Feng Zhu 0024, Changsheng Xie 0001, Fei Wu 0005
ICCD2
2023 ADT-FSE: A New Encoder for SZ
abstract
SZ is a lossy floating-point data compressor that excels in compression ratio and throughput for high-performance computing (HPC), time series databases, and deep learning applications. However, SZ performs poorly for small chunks and has slow decompression. We pinpoint the Huffman tree in the quantization factor encoder as the bottleneck of SZ. In this paper, we propose ADT-FSE, a new quantization factor encoder for SZ. Based on the Gaussian distribution of quantization factors, we design an adaptive data transcoding (ADT) scheme to map quantization factors to codes for better compressibility, and then use finite state entropy (FSE) to compress the codes. Experiments show that ADT-FSE improves the quantization factor compression ratio, compression and decompression throughput by up to 5×, 2× and 8×, respectively, over the original SZ Huffman encoder. On average, SZ_ADT is over 2× faster than ZFP in decompression. Case studies of the TDengine time series database and HDF5 file store confirm that SZ_ADT significantly boosts user-perceived application performance. In addition, ADT-FSE makes the compression ratio prediction of SZ_ADT easy and accurate, and has the potential to dramatically reduce the area size of SZ hardware implementation.
Tao Lu 0014, Zibin Sun, Xiang Chen 0028, You Zhou 0009, Fei Wu 0005, Yunxin Huang, Yafei Yang
SC5
2023 Holistic and Opportunistic Scheduling of Background I/Os in Flash-Based SSDs
abstract
Background (BG)tasks are maintained indispensably in multiple layers of storage systems, from applications to flash-based SSDs. They launch a large amount of I/Os, causing significant interference withforeground (FG)I/O performance. Our key insight is that, to mitigate such interference, holistic scheduling of system-wide, multi-source BG I/Os is required and can only be realized at the underlying SSD layer. Only the SSD has a global view of all FG and BG I/Os as well as direct information and control about flash storage resources. We are thus inspired to propose a novel I/O scheduling architecture, calledHuFu. It provides a framework for host software to register BG tasks and offload their I/O scheduling into the SSD. Then, the SSD-internal I/O scheduler prioritizes FG I/O processing, while BG I/Os are scheduled opportunistically by utilizing flash parallelism and idleness. To verifyHuFu, we perform case studies on RocksDB and compares it with several state-of-the-art host-side I/O scheduling schemes. Experimental results show thatHuFucan significantly alleviate performance interference caused by BG I/Os and improve SSD bandwidth utilization, thus improving the FG throughput, average and tail latencies (e.g., by about 18% in a write-heavy workload).
Yu Wang 0168, You Zhou 0009, Fei Wu 0005, Jian Zhou 0004, Zhonghai Lu, Zhengyong Wang, Changsheng Xie 0001
IEEE Trans. Computers2
2022 Tiered Hashing: Revamping Hash Indexing under a Unified Memory-Storage Hierarchy
abstract
NAND flash-based Solid State Drives (SSDs) provide a promising opportunity to enable the unified memory-storage hierarchy (UMH). The UMH renders a single memory address space for heterogeneous memories. Thus, the CPUs can directly access structured data in SSDs and eliminate bulk data copy/swap between the memory and storage devices. However, applying traditional indexing structures directly on SSDs may lead to poor performance. Particularly, the popular hash indexing generates highly randomized write traffic, incurring significant garbage collection overhead in SSDs. To address this problem, we propose a novel SSD-friendly hash indexing scheme called Tiered Hashing. It employs a multi-layer structure and opportunistic data movement (ODM) to construct skewed writes. Hence, the SSD can transform the writes into multi-streamed writes, where hot and cold data are separated to reduce GC overhead. Experimental results show Tiered Hashing reduces the average write latency and GC overhead by up to 94.98% and 90.71% compared to state-of-the-art hash indexings, without sacrificing read performance.
Jian Zhou 0004, Weizhou Huang, You Zhou 0009, Fei Wu 0005, Liu Shi, Kun Wang 0029, Feng Zhu 0024
PACT4
2022 PACA: A Page Type Aware Read Cache Scheme in QLC Flash-based SSDs
abstract
QLC flash-based SSDs are gaining increasing attention and are expected to be widely used in read-intensive application scenarios, since they provide high density and low cost but suffer from poor write endurance and performance. QLC flash has four types of pages, between which read latency variation is as large as 1.6 to 4.8 times. This raises a critical concern for QLC SSDs to provide adequate and stable read performance. Notice that the SSD-internal cache (built with DRAM or non-volatile RAM) has long been utilized to improve write performance and lifetime. In this paper, we argue that the cache also plays an important role in read performance optimization of QLC SSDs. We design a novel flash page type aware read cache scheme, called PACA. It exploits read latency variation of QLC pages to prioritize caching data stored in high-latency QLC pages in a workload-adaptive manner. We verified PACA in FEMU, a popular SSD emulator. Experimental results show that PACA can reduce the average SSD read latency by up to 44.5%, compared with a baseline read cache scheme being unaware of flash page types.
Qihui Chen, You Zhou 0009, Fei Wu 0005, Zhengyong Wang, Changsheng Xie 0001
ICCD3
2022 WA-OPShare: Workload-Adaptive Over-Provisioning Space Allocation for Multi-Tenant SSDs
abstract
Sharing a flash-based solid-state drive (SSD) among multiple tenants has become a common practice to improve storage utilization and cost efficiency. Meanwhile, how to allocate limited storage resources, especially the over-provisioning space (OPS) resources, among competitive tenants has emerged as a critical problem. The OPS refers to additional user-invisible storage space, whose size influences garbage collection (GC) efficiency. Due to unawareness of workload characteristics of different tenants, prior studies on multitenant OPS allocation lead to suboptimal SSD performance. In this article, we propose a novel workload-adaptive OPS allocation scheme for multitenant SSDs, called WA-OPShare. It targets an OPS sharing scheme that dynamically allocates the OPS among tenants to improve overall SSD performance. Two models are developed to identify underutilized storage space and predict the OPS-induced performance benefit of each tenant, respectively. Guided by the models, WA-OPShare regularly releases the underutilized storage space and then reallocates it to the tenant who can benefit the most. Experimental results show that compared to the traditional Partition and Sharing schemes, WA-OPShare improves the performance by up to 40.3% and 31.2%, and reduces the write amplification by up to 37.0% and 17.5%, respectively.
Yuhong Wen, You Zhou 0009, Fei Wu 0005, Zhenghong Wang, Changsheng Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Understanding and Exploiting the Full Potential of SSD Address Remapping
abstract
Duplicate writes are prevalent in storage systems, originating from data duplication, journaling, and data relocations, etc. As flash-based solid state drives (SSDs) have been widely deployed, duplicate writes can significantly degrade their performance and lifetime. Prior studies have proposed innovative approaches that exploit the address remapping utility inside an SSD to eliminate duplicate writes. However, remap operations modify the logical-to-physical (L2P) address mapping table while the physical-to-logical (P2L) mappings persisted on flash memory remain unchanged. Such inconsistency between L2P and P2L mappings may cause data corruption and has long been a major obstacle to utilize SSD address remapping. In this article, we propose a novel SSD design, called Remap-SSD-LH, that realizes the full potential of SSD address remapping. It provides a remap primitive, which allows the host software and SSD firmware to perform logical writes of duplicate data at almost zero cost. To ensure mapping consistency as well as fast mapping lookups, Remap-SSD-LH employs a local log scheme based on hybrid storage. A local log is maintained for each flash garbage collection unit to record relevant P2L mapping changes induced by remap operations. The logs are stored in small nonvolatile RAM (NVRAM), e.g., capacitor-protected DRAM, and can be destaged to flash memory if NVRAM is full. We verify Remap-SSD-LH on a software SSD emulator with three case studies: 1) intra-SSD deduplication; 2) SQLite journaling; and 3) F2FS cleaning. The experimental results show that Remap-SSD-LH can maximally and efficiently exploit address remapping to improve SSD performance and lifetime.
Qiulin Wu, You Zhou 0009, Fei Wu 0005, Hong Jiang 0001, Jian Zhou 0004, Changsheng Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 SW-WAL: Leveraging Address Remapping of SSDs to Achieve Single-Write Write-Ahead Logging
abstract
Write-ahead logging (WAL) has been widely used to provide transactional atomicity in databases, such as SQLite and MySQL/InnoDB. However, the WAL introduces duplicate writes, where changes are recorded in the WAL file and then written to the database file, called checkpointing writes. On the other hand, NAND flash-based SSDs, which have an inherent indirection software layer, called flash translation layer (FTL), become commonplace in modern storage systems. Innovative SSD designs have been proposed to eliminate the WAL overheads by exploiting the FTL, such as providing an atomic write interface or utilizing its address remapping. However, these designs introduce significant performance overheads of maintaining and persisting extra transactional information to guarantee the transactional atomicity or mapping consistency. In this paper, we propose single-write WAL (SW-WAL), a novel cross-layer design, to eliminate WAL-induced duplicate writes on SSDs with minimal overheads. The SSD exposes an address remapping interface to the host, through which the checkpointing writes can be completed without conducting real data writes. To ensure the transactional atomicity and mapping consistency, we make the SSD aware of the transactional writes to the WAL file. Specifically, when transactional data are written to the WAL file, both transactional and mapping semantics are delivered from the host to the SSD and persisted in relevant flash pages as housekeeping metadata without any extra overheads. We implement a prototype of SW-WAL, which runs a popular database SQLite on an emulated NVMe SSD. Experimental results show that SW-WAL improves the database performance by up to 62% compared with original SQLite that bears the WAL overheads and up to 32% compared with the state-of-the-art design that eliminates the WAL overheads.
Qiulin Wu, You Zhou 0009, Fei Wu 0005, Jiguang Wan 0001, Changsheng Xie 0001
DATE2
2021 Remap-SSD: Safely and Efficiently Exploiting SSD Address Remapping to Eliminate Duplicate Writes
You Zhou 0009, Qiulin Wu, Fei Wu 0005, Hong Jiang 0001, Jian Zhou 0004, Changsheng Xie 0001
FAST1
2021 Seer-SSD: Bridging Semantic Gap between Log-Structured File Systems and SSDs to Reduce SSD Write Amplification
abstract
Log-structured file systems (LS-FSs) sequentialize writes, so they are expected to perform well on flash-based SSDs. However, we observe a semantic gap between the LS- FS and SSD that causes a stale-LBA problem. When data are updated, the LS-FS allocates new logical block addresses (LBAs). The relevant stale LBAs are invalidated and then trimmed or reused with a delay by the LS-FS. During the time interval, stale LBAs are regarded temporarily as valid and migrated unnecessarily by garbage collection in the SSD. Our experimental study of real-world traces reveals that stale-LBA migrations amount to 59%-150% of host data writes. To solve this serious problem, we propose Seer-SSD to deliver stale-LBA metadata along with written data from the LS-FS to the SSD. Then, stale LBAs are invalidated actively and selectively in the SSD without compromising file system consistency. Seer-SSD can be implemented easily based on existing block interfaces and maintain compatibility with non-LS-FSs. We perform a case study on an emulated NVMe SSD hosting F2FS (a state-of-the- art LS-FS). Experimental results with popular databases show that Seer-SSD improves the throughput by 99.8% and reduces the write amplification by 53.6%, on average, compared to a traditional SSD unaware of stale LBAs.
You Zhou 0009, Fei Wu 0005, Changsheng Xie 0001
ICCD1
2021 An Efficient Data Migration Scheme to Optimize Garbage Collection in SSDs
abstract
Garbage collection (GC) is time consuming and frequently executed all over the lifetime of solid-state drives (SSDs), which has a significant impact on system performance. Manufactures provide the copyback that directly transfers data within the same plane to accelerate data migration in GC. However, the introduction of copyback leads to two issues: 1) high detection overhead of copyback feasibility (whether data are carried out via copyback with guaranteed reliability) and 2) interplane unbalanced wear distribution. In this article, we first explore copyback error characteristics on the real NAND flash chip, then propose a fast GC scheme called FastGC. It utilizes copyback error characteristics to efficiently detect the copyback feasibility of data instead of transferring out all valid data for detecting. FastGC further utilizes a data migration leveler which aims at relieving migration overhead per GC to realize the wear leveling. Regarding data migrated via external data move (EDM), FastGC takes data coldness and erase counts of planes into consideration to even out the number of migrating data per plane and prolong the lifetime of SSDs. SSDsim, a validate simulation is used to implement FastGC and comprehensive experiments are carried out with various enterprise workloads to evaluate the system performance and the wear difference of SSDs. The experimental results in the SSDsim show the FastGC greatly promotes system performance and the wear leveling up to 46.68% and 12X, respectively, compared to the traditional copyback-based GC.
Shunzhuo Wang, You Zhou 0009, Jiaona Zhou, Fei Wu 0005, Changsheng Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 LiveSSD: A Low-Interference RAID Scheme for Hardware Virtualized SSDs
abstract
Hardware virtualization has been increasingly used to provide performance isolation between multiple tenants sharing an SSD. It exploits the SSD's highly parallel architecture by allocating dedicated flash dies to each tenant. On the other hand, intra-SSD RAID, which stripes data and parity across flash dies, is essential to enhance storage reliability, such as protecting data against die failures and read errors. However, parity updates introduce I/O interference, degrading tenants' performance significantly, and violating performance isolation. To solve this problem, we propose a low-interference RAID scheme for hardware virtualized SSDs, called LiveSSD. Flash pages with the same offset across dies constitute a stripe in a RAID-4 manner. High-speed NVRAM is employed as parity storage. Thus, LiveSSD allows each tenant to read/write its flash die(s) independently and avoids parity updates being a performance bottleneck. Nonetheless, parity updates introduce I/O interference during garbage collection, i.e., extra reads of invalid flash pages. LiveSSD actively conducts parity updates in advance by utilizing both page access feature of flash memory and idle time in workloads. Extensive simulation results show that LiveSSD enables RAID protection in a hardware-virtualized SSD with minimum I/O interference caused by parity updates.
You Zhou 0009, Fei Wu 0005, Weizhou Huang, Changsheng Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 SCORE: A Novel Scheme to Efficiently Cache Overlong ECCs in NAND Flash Memory
abstract
Technology scaling and program/erase cycling result in an increasing bit error rate in NAND flash storage. Some solid state drives (SSDs) adopt overlong error correction codes (ECCs) , whose redundancy size exceeds the spare area limit of flash pages, to protect user data for improved reliability and lifetime. However, the read performance is significantly degraded, because a logical data page and its ECC redundancy are stored in two flash pages. In this article, we find that caching ECCs has a large potential to reduce flash reads by achieving higher hit rates, compared to caching data. Then, we propose a novel scheme to efficiently cache overlong ECCs, called SCORE , to improve the SSD performance. Exceeding ECC redundancy (called ECC residues ) of logically consecutive data pages are grouped into ECC pages . SCORE partitions RAM to cache both data pages and ECC pages in a workload-adaptive manner. Finally, we verify SCORE using extensive trace-driven simulations. The results show that SCORE obtains high ECC hit rates without sacrificing data hit rates, thus improving the read performance by an average of 22% under various workloads, compared to the state-of-the-art schemes.
You Zhou 0009, Fei Wu 0005, Zhonghai Lu, Xubin He, Ping Huang 0001, Changsheng Xie 0001
ACM Trans. Archit. Code Optim.1
2018 OSPADA: One-Shot Programming Aware Data Allocation Policy to Improve 3D NAND Flash Read Performance
abstract
Charge trap (CT) based 3D NAND flash is predominating the flash storage market due to higher density, better performance and endurance than planar flash. CT-based 3D flash programs multiple pages in a word line at a time, called one-shot programming, unlike planar flash which programs one page at a time. Solid state drives (SSDs) utilize the internal parallelism to improve the performance, but one-shot programming is likely to program logically sequential data into one parallel unit (i.e., a plane) and thus degrades the read parallelism. In this paper, we propose a one-shot programming aware data allocation policy, called OSPADA, to improve the read performance of CT flash based SSDs by enhancing read parallelism. OSPADA reorders written data to distribute logically sequential data into different parallel units using the distance aware round-robin strategy. Experimental results show that OSPADA improves the read performance by up to 22.8% compared with traditional dynamic data allocation policies.
Fei Wu 0005, Zuo Lu, You Zhou 0009, Xubin He, Zhihu Tan, Changsheng Xie 0001
ICCD3
2018 Characterizing 3D Charge Trap NAND Flash: Observations, Analyses and Applications
abstract
In the 3D era, the Charge Trap (CT) NAND flash is employed by mainstream products, thus having a deep understanding of its characteristics is becoming increasingly crucial for designing flash-based systems. In this paper, to enable such understanding, we implement comprehensive experiments on advanced 3D CT NAND flash chips by developing an ARM-and FPGA-based evaluation platform. Based on the experimental results, we first make distinct observations on the characteristics of 3D CT NAND flash, including its performance and reliability features. Then we give analyses of the observations from physical and circuit aspects. Finally, based on the unique characteristics of 3D CT NAND flash, suggestions to optimize the flash management algorithms in real applications are presented.
Fei Wu 0005, Qin Xiong, Zhonghai Lu, You Zhou 0009, Weizhen Kong, Changsheng Xie 0001
ICCD5
2018 Exploiting Minipage-Level Mapping to Improve Write Efficiency of NAND Flash
abstract
Pushing NAND flash memory to higher density, manufacturers are aggressively enlarging the flash page size. However, the sizes of I/O requests in a wide range of scenarios do not grow accordingly. Since a page is the unit of flash read/write operations, traditional flash translation layers (FTLs) maintain the page mapping regularity. Hence, small random write requests become common, leading to extensive partial logical page writes. This write inefficiency significantly degrades the performance and increases the write amplification of flash storage. In this paper, we first propose a configurable mapping layer, called minipage, whose size is set to match I/O request sizes. The minipage-level mapping provides better flexibility in handling small writes at the cost of sequential read performance degradation and a larger mapping table. Then, we propose a new FTL, called PM-FTL, that exploits the minipage-level mapping to improve write efficiency and utilizes the page-level mapping to reduce the costs caused by the minipage-level mapping. Finally, trace-driven simulation results show that compared to traditional FTLs, PM-FTL reduces the write amplification and flash storage response time by an average of 33.4% and 19.1%, up to 57.7% and 34%, respectively, under 16KB flash pages and 4KB minipages.
You Zhou 0009, Fei Wu 0005, Weijun Xiao, Xubin He, Zhonghai Lu, Changsheng Xie 0001
NAS2
2018 Characterizing 3D Floating Gate NAND Flash: Observations, Analyses, and Implications
abstract
As both NAND flash memory manufacturers and users are turning their attentions from planar architecture towards three-dimensional (3D) architecture, it becomes critical and urgent to understand the characteristics of 3D NAND flash memory. These characteristics, especially those different from planar NAND flash, can significantly affect design choices of flash management techniques. In this article, we present a characterization study on the state-of-the-art 3D floating gate (FG) NAND flash memory through comprehensive experiments on an FPGA-based 3D NAND flash evaluation platform. We make distinct observations on its performance and reliability, such as operation latencies and various error patterns, followed by careful analyses from physical and circuit-level perspectives. Although 3D FG NAND flash provides much higher storage densities than planar NAND flash, it faces new performance challenges of garbage collection overhead and program performance variations and more complicated reliability issues due to, e.g., distinct location dependence and value dependence of errors. We also summarize the differences between 3D FG NAND flash and planar NAND flash and discuss implications on the designs of NAND flash management techniques brought by the architecture innovation. We believe that our work will facilitate developing novel 3D FG NAND flash-oriented designs to achieve better performance and reliability.
Qin Xiong, Fei Wu 0005, Zhonghai Lu, You Zhou 0009, Yibing Chu, Changsheng Xie 0001, Ping Huang 0001
ACM Trans. Storage5
2017 Lifetime adaptive ECC in NAND flash page management
abstract
NAND flash memory has decreasing storage reliability, as the density or program/erase (P/E) cycle increases. To ensure data integrity, error correction codes (ECCs) are widely employed and typically stored in the out-of-band area (OOB) of flash pages. However, the worst-case oriented ECC is largely under-utilized in the early stage (small P/E cycles), and the required ECC redundancy may be too large to fit in OOB in the late stage (high P/E cycles). In this paper, we propose LAE-FTL, which employs a lifetime-adaptive ECC scheme, to improve the performance and lifetime of NAND flash memory. LAE-FTL uses weak ECCs in the early stage and strong ECCs in the late stage to guarantee the storage reliability. Since OOB is large enough to store weak ECCs in the early stage, small and size-incremental codewords are adaptively used to improve data transfer and decoding parallelism. In the late stage, strong ECCs have to be employed and the ECC redundancies become too large to be stored in OOB. Thus, LAE-FTL stores the exceeding ECC redundancies in the data space of flash pages and stores user data in a cross-page fashion. Finally, our trace-driven simulation results show that LAE-FTL improves the read performance by up to 63.42%, compared to the worst-case oriented ECC scheme in the early stage, and significantly improve the storage reliability at low cost in the late stage.
Shunzhuo Wang, Fei Wu 0005, Zhonghai Lu, You Zhou 0009, Qin Xiong, Meng Zhang 0014, Changsheng Xie 0001
DATE4
2017 Understanding and Alleviating the Impact of the Flash Address Translation on Solid State Devices
abstract
Flash-based solid state devices (SSDs) have been widely employed in consumer and enterprise storage systems. However, the increasing SSD capacity imposes great pressure on performing efficient logical to physical address translation in a page-level flash translation layer (FTL). Existing schemes usually employ a built-in RAM to store mapping information, called mapping cache , to speed up the address translation. Since only a fraction of the mapping table can be cached due to limited cache space, a large number of extra flash accesses are required for cache management and garbage collection, degrading the performance and lifetime of an SSD. In this paper, we first apply analytical models to investigate the key factors that incur extra flash accesses during address translation. Then, we propose a novel page-level FTL with an efficient translation page-level caching mechanism, named TPFTL , to minimize the extra flash accesses. TPFTL employs a two-level least recently used (LRU) list with space-efficient optimizations to organize cached mapping entries. Inspired by the models, we further design a workload-adaptive loading policy combined with an efficient replacement policy to increase the cache hit rate and reduce the writebacks of replaced dirty entries. Finally, we evaluate TPFTL using extensive trace-driven simulations. Our evaluation results show that compared to the state-of-the-art FTLs, TPFTL significantly reduces the extra operations caused by address translation, achieving reductions on system response time and write amplification by up to 27.1% and 32.2%, respectively.
You Zhou 0009, Fei Wu 0005, Ping Huang 0001, Xubin He, Changsheng Xie 0001, Jian Zhou 0004
ACM Trans. Storage1
2015 An efficient page-level FTL to optimize address translation in flash memory
abstract
Flash-based solid state disks (SSDs) have been very popular in consumer and enterprise storage markets due to their high performance, low energy, shock resistance, and compact sizes. However, the increasing SSD capacity imposes great pressure on performing efficient logical to physical address translation in a page-level flash translation layer (FTL). Existing schemes usually employ a built-in RAM cache for storing mapping information, called the mapping cache, to speed up the address translation. Since only a fraction of the mapping table can be cached due to limited cache space, a large number of extra operations to flash memory are required for cache management and garbage collection, degrading the performance and lifetime of an SSD. In this paper, we first apply analytical models to investigate the key factors that incur extra operations. Then, we propose an efficient page-level FTL, named TPFTL, which employs two-level LRU lists to organize cached mapping entries to minimize the extra operations. Inspired by the models, we further design a workload-adaptive loading policy combined with an efficient replacement policy to increase the cache hit ratio and reduce the writebacks of replaced dirty entries. Finally, we evaluate TPFTL using extensive trace-driven simulations. Our evaluation results show that compared to the state-of-the-art FTLs, TPFTL reduces random writes caused by address translation by an average of 62% and improves the response time by up to 24%.
You Zhou 0009, Fei Wu 0005, Ping Huang 0001, Xubin He, Changsheng Xie 0001, Jian Zhou 0004
EuroSys1
2015 A novel optimization algorithm for Chien search of BCH Codes in NAND flash memory devices
abstract
As NAND flash memory chips become denser, they are more vulnerable to random errors caused by ageing, read or write interference, and erase operations. These errors compromise both the data integrity and lifetime of flash memory so that error correction codes (ECC) are employed by the flash controller to strengthen the fault tolerance. The BCH (Bose Chaudhuri Hochquenghem) code is a widely used ECC technique in flash-based storage devices due to its strong error correction capability and high performance. The third step of decoding a BCH code is the Chien search process, which locates the errors in the received codeword. To increase the decoding throughput, parallel Chien search algorithms are used, but existing algorithms occupy more than 60% area of the total decoding logic, increasing the hardware complexity and energy consumption. To reduce the hardware complexity and overhead, in this paper, we propose a plane optimization algorithm to reduce the redundant XOR gates used in the Chien search process. Our study based on intensive experiments shows that for a (2047,1926, 11) BCH code with the parallel factor of 32, the proposed optimization algorithm reduces the number of XOR gates used in the Chien search process by 79%, 46% and 13%, respectively, compared to the straightforward implementation, the GMA approach and the strength-reduced architecture.
Meng Zhang 0014, Fei Wu 0005, Changsheng Xie 0001, You Zhou 0009
NAS4