VLDB 2026 Research / reviewers in the wild / expert
Feng Zhu 0024
dblp:71/2791-24
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0008-1119-3422ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shiro: Efficient and Accurate In-Storage Data Lifetime Separation for nand Flash SSDsabstractThe log-structured nature of NAND flash storage necessitates garbage collection in SSDs. Garbage collection (GC) is a major source of runtime write amplification (WA), leading to faster device wear out and interference with host I/Os. The key to mitigating this problem is separating data by lifetime so that data in the same flash block are invalidated within temporal proximity. For higher lifetime prediction accuracy and adaptibility, prior works proposed using machine learning algorithms for data separation. However, existing learning-based solutions perform data lifetime prediction at the host side, leading to several drawbacks. First, host-side prediction does not have knowledge of the internal data movement inside the SSD during GC, and thus fails to leverage the opportunity to further separate GC writes, resulting in suboptimal WA reduction in the long term. Second, performing prediction at the host significantly prolongs the I/O critical path and consumes host resources that could otherwise be used for serving user applications. We present Shiro, a holistic FTL design that performs instorage data separation for both user writes and GC writes for maximal long-term WA reduction. For user writes, Shiro uses a sequence model to accurately predict data lifetime by learning lifetime distribution from long historical access patterns. For GC writes, Shiro incorporates a reinforcement learning-assisted page migration strategy that takes direct feedback from longterm WA to further improve data separation efficacy. To address the challenges posed by performing fine-grained and real-time machine learning decisions inside the resource-constrained SSD, we propose a suite of enabling techniques to keep computation and storage overhead low. Extensive evaluation of Shiro on real-world traces shows that Shiro can deliver 29 WA compared with conventional FTL and state-of-the-art instorage data separation schemes. Furthermore, thanks to lower data migration overhead during GC, Shiro achieves significantly higher steady-state I/O performance. Penghao Sun, Shengan Zheng, Litong You, Wanru Zhang, Ruoyan Ma, Feng Zhu 0024, Linpeng Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | TxISC: Transactional File Processing in Computational SSDsabstractComputational SSDs implement the in-storage computing (ISC) paradigm and benefit applications by taking over I/O-intensive tasks from the host. Existing works have proposed various frameworks aiming at easy access to ISC functionalities, and among them generic frameworks with file-based abstractions offer better usability. However, since intermediate output by ISC tasks may leave files in a dirty state, concurrent access to and the integrity of file data should be properly managed, which has not been fully addressed. In this paper, we present TxISC, a generic ISC framework that coordinates the host kernel and device firmware to offer a versatile file-based programming model. Under the hood, TxISC turns each invocation of an ISC task into a transaction with full ACID guarantee, fully covering concurrency control and data protection. TxISC implements transactions at low cost by leveraging the out-of-place write characteristic of NAND flash. Evaluation on full-stack hardware shows that transactions incur almost no runtime performance penalty compared with existing ISC architectures. Application case studies demonstrate that the programming model of TxISC can be used to offload complex logic and deliver significant speedup over host-only solutions. Penghao Sun, Shengan Zheng, Kaijiang Deng, Jin Pu, Maojun Yuan, Feng Zhu 0024, Linpeng Huang |
DATE | 8 |
| 2025 | An Efficient Independent Read Scheme for Contemporary QLC SSDsabstractQLC solid-state disks (SSDs) are increasingly deployed in large-scale storage systems. While achieving remarkable storage density and cost-effectiveness, QLC NAND exhibits degraded performance. To alleviate the issue, Independent Multi-Plane (IMP) read has been proposed to leverage the plane-level parallelism under random read workloads. However, compared to the previous generations of chips, the variation in read latency of QLC chip has widened significantly, and the number of planes in a QLC chip has increased. As a result, the idle time in IMP commands has escalated dramatically. Moreover, the conventional layout of data and parities in redundant array of independent NAND (RAIN) increases the probability of high latency read occureneces, exacerbating the contribution to idle time. Consequently, the incorporation of IMP is inherently inefficient in contemporary QLC NAND flash chips. In this paper, we propose an efficient independent read (EFFIR) scheme to tackle this challenge. EFFIR features a read latency variation aware transaction service that properly combines read transactions to minimize idle time and proactively transfers read data to eliminate unnecessary delays. Moreover, EFFIR incorporates a read latency variation aware RAIN that reorganizes the layout of data and parity to mitigate the impact of high-latency data access on idle time. Our comprehensive experimental results demonstrate and elucidate how EFFIR significantly enhances SSD responsiveness while consistently delivering favorable performance across a diverse range of read-intensive workloads. Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001, Feng Zhu 0024, Maojun Yuan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2024 | TieredHM: Hotspot-Optimized Hash Indexing for Memory-Semantic SSD-Based Hybrid MemoryabstractMemory semantic Solid State Drives (MS-SSDs) provide a promising opportunity to enable the hybrid memory architecture (HMA). The memory semantic interface enables the CPUs to directly access structured data in SSDs and eliminate bulk data copy/swap between the memory and storage devices. However, existing hash indexings issue many random writes, resulting in two problems when directly deployed on MS-SSD-based HMA: 1) Highly random traffic persisted to the underlying NAND flash of MS-SSDs incurs significant garbage collection (GC) overhead. 2) Placing frequently updated memory pages of hash indexings in persistent memories (PMs) is anticipated to reduce write latency, failing to work effectively due to the lack of skewness. To address the above problems, we propose a novel MS-SSD-friendly hash indexing scheme called TieredHM. It employs a multi-layer structure and opportunistic data movement (ODM) to construct skewed writes. Hence, the MS-SSD can transform the writes into multi-streamed writes, separating data with different update frequencies to reduce GC overhead. Besides, since the top layer is updated much more frequently (more skewed) than other layers, placing the top layer of TieredHM into persistent memory can significantly reduce write latency. TieredHM further leverages a prefetch mechanism based on the internal parallelism of NAND flash to reduce search overhead incurred by ODM. Experimental results show that TieredHM reduces the average write latency and GC overhead by up to 8.3X and 20.0X compared to state-of-the-art hash indexings without sacrificing read performance. Weizhou Huang, Jian Zhou 0004, You Zhou 0009, Feng Zhu 0024, Kun Wang 0029, Fei Wu 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Learning-based Data Separation for Write Amplification Reduction in Solid State DrivesabstractGarbage collection in SSDs causes write amplification. The key to mitigating this problem is separating data by lifetime. Prior works proposed using machine learning to accurately predict data lifetime but prediction is performed at the host side, burdening the host storage stack. We present PHFTL, a practical, holistic FTL design with device-side learning-based data separation. The machine learning model in PHFTL accurately and adaptively predicts the lifetime of every written page. A suite of enabling techniques are introduced to keep computation and storage overhead low. Extensive evaluation of PHFTL demonstrates superiority over state-of-the-art and feasibility on real hardware. Penghao Sun, Litong You, Shengan Zheng, Wanru Zhang, Ruoyan Ma, Guanzhong Wang, Feng Zhu 0024, Linpeng Huang |
DAC | 8 |
| 2023 | FlexZNS: Building High-Performance ZNS SSDs with Size-Flexible and Parity-Protected ZonesabstractNVMe zoned namespace (ZNS) SSDs present a new class of storage devices with attractive features including low cost, software definability, and stable performance. However, one primary culprit that hinders the adoption of ZNS is the high garbage collection (GC) overhead it brings to host software. The ZNS interface divides the logical address space into size-fixed zones that must be written sequentially. Despite being friendly to flash memory, ZNS requires host software to perform out-of-place updates and GC on individual zones. Current ZNS SSDs typically employ a large zone size (e.g., of GBs) to be conducive to die-level RAID protection on flash memory. This impedes flexible data placement, such as mixing data with different lifetimes in the same zone, and incurs sizable data migrations during zone GC. To address this problem, we propose FlexZNS, a novel ZNS SSD design that provides reliable zoned storage allowing host software to configure the zone size flexibly as well as multiple zone sizes. The size variability of zones poses two interrelated challenges, one for the SSD controller to establish per-zone RAID protection, and the other for host software to manage variable zone capacity loss caused by parity storage. To tackle the challenges, FlexZNS decouples the storage of parity from individual zones on flash memory and hides the zone capacity loss from the host software. We verify FlexZNS on a ZNS-compatible file system F2FS and a popular key-value store RocksDB. Extensive experiments demonstrate that FlexZNS can significantly improve the system performance and reduce GC-induced write amplification, compared with a conventional ZNS SSD with large-sized zones. Yu Wang 0168, You Zhou 0009, Zhonghai Lu, Kun Wang 0029, Feng Zhu 0024, Changsheng Xie 0001, Fei Wu 0005 |
ICCD | 6 |
| 2022 | Tiered Hashing: Revamping Hash Indexing under a Unified Memory-Storage HierarchyabstractNAND flash-based Solid State Drives (SSDs) provide a promising opportunity to enable the unified memory-storage hierarchy (UMH). The UMH renders a single memory address space for heterogeneous memories. Thus, the CPUs can directly access structured data in SSDs and eliminate bulk data copy/swap between the memory and storage devices. However, applying traditional indexing structures directly on SSDs may lead to poor performance. Particularly, the popular hash indexing generates highly randomized write traffic, incurring significant garbage collection overhead in SSDs. To address this problem, we propose a novel SSD-friendly hash indexing scheme called Tiered Hashing. It employs a multi-layer structure and opportunistic data movement (ODM) to construct skewed writes. Hence, the SSD can transform the writes into multi-streamed writes, where hot and cold data are separated to reduce GC overhead. Experimental results show Tiered Hashing reduces the average write latency and GC overhead by up to 94.98% and 90.71% compared to state-of-the-art hash indexings, without sacrificing read performance. Jian Zhou 0004, Weizhou Huang, You Zhou 0009, Fei Wu 0005, Liu Shi, Kun Wang 0029, Feng Zhu 0024 |
PACT | 9 |
| 2021 | Optimizing Performance for Open-Channel SSDs in Cloud Storage SystemabstractIn large-scale cloud storage systems, Solid-State Drive (SSD) has been broadly used as the mainstream storage device because it has the advantages of low access latency and high throughput. However, conventional SSD is a black-box system to host softwares, thus failing to fully exploit the benefits of NAND flash and provide high quality of service (QoS). On the other hand, Open-Channel SSD (OCSSD) which exposes its internal information to the host software, has the potential to solve this problem. However, existing OCSSD fails to achieve anticipated performance under heavy workloads. To this end, we propose an advanced OCSSD-based driver developed with the novel data placement policy, redefined garbage collection (GC) with copyback technique, efficient prefetch read scheme, and fast live upgrade method. Our work describes the consistent efforts to pursue high performance and QoS in OCSSDs with different approaches. The evaluation results show that our novel Open-Channel SSD is able to provide high I/O throughputs and predictable I/O latencies. For example, our Open-Channel SSD can improve I/O throughputs by 103% and reduce the 99th percentile latency by 62.9% on average compared with the state-of-the-art NVMe SSDs. Feng Zhu 0024, Kun Wang 0029, Dengcai Xu |
IPDPS | 2 |