EDBT 2026 Demo / reviewers in the wild / expert
Bo Ding 0002
dblp:37/516-2
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-7588-0140ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SuperCopyback: Revisiting Copyback on Modern High-Performance NAND Flash-based SSDsabstractNAND flash-based SSDs have emerged as a critical storage solution due to their exceptional performance and cost-effectiveness. However, the sequential write limitation of NAND flash blocks necessitates garbage collection (GC) to reclaim space occupied by stale data. Nevertheless, the extensive data migration involved in GC significantly impacts performance and Quality of Service (QoS) of SSDs. To mitigate this issue, copyback has been proposed as a means to accelerate GC by eliminating off-chip data movements. Specifically, copyback reads data into on-plane latches and immidiately re-writes it into another page on the same plane. However, in the case of modern high-performance SSDs, copyback is rarely utilized due to the following challenges: (1) Copyback operates at the page-level and thus fails to effectively reclaim invalid data within the context of subpage-level mapping; (2) Copyback eliminates off-chip data movements, preventing data pages from being read out for Redundant Array of Independent NAND (RAIN) parity computation, thereby compromising SSD reliability. In this study, we introduce SuperCopyback as a solution that efficiently addresses these issues for modern SSDs. Firstly, we propose a Multiple-Read-One-Write (MROW) copyback approach through lightweight latch circuit modifications to enable subpage-level copyback implementation; additionally, we propose an orchestrated GC method to effectively utilize MROW copyback. Furthermore, we present a novel copyback-based RAIN scheme that conceals data pages readout latency in write operations and relocates parities to support efficient copyback. The experimental results on both synthetic and real traces demonstrate that SuperCopyback achieves a performance comparable to an ideal scenario where data movement of $\mathbf{4 K B}$ takes only 1ns. Bo Ding 0002, Wei Tong 0001, Dan Feng 0001 |
DAC | 2 |
| 2025 | MPFS: A Scalable User-Space Persistent Memory File System for Multiple Processesabstract11This work was supported by the Young Scientists Fund of the National Natural Science Foundation of China under Grant 62302182.Persistent memory (PM) leveraging memory-mapped I/O(MMIO) delivers superior I/O performance, leading to the development of user-space PM file systems based on MMIO. While effective in single-process scenarios, these systems encounter challenges in multi-process environments, such as performance degradation due to repeated page faults and cross-process synchronizations, as well as a large memory footprint from duplicated paging structures. To address these problems, we propose a Multi-process PM File System (MPFS). MPFS builds a shareable page table and shares it among processes, avoiding building duplicate paging structures for distinct processes, thereby significantly reducing the software overhead and memory footprint caused by repeated page faults. MPFS further proposes a PGD-aligned (512GB) mapping method to accelerate page table sharing. Furthermore, MPFS provides a cross-process memory protection mechanism based on the PGD-aligned mapping, ensuring multi-process data reliability with negligible overheads. The experimental results show that MPFS outperforms existing user-space PM file systems by 1560% in multi-process scenarios. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Yuchong Hu, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
DATE | 1 |
| 2025 | An Efficient Independent Read Scheme for Contemporary QLC SSDsabstractQLC solid-state disks (SSDs) are increasingly deployed in large-scale storage systems. While achieving remarkable storage density and cost-effectiveness, QLC NAND exhibits degraded performance. To alleviate the issue, Independent Multi-Plane (IMP) read has been proposed to leverage the plane-level parallelism under random read workloads. However, compared to the previous generations of chips, the variation in read latency of QLC chip has widened significantly, and the number of planes in a QLC chip has increased. As a result, the idle time in IMP commands has escalated dramatically. Moreover, the conventional layout of data and parities in redundant array of independent NAND (RAIN) increases the probability of high latency read occureneces, exacerbating the contribution to idle time. Consequently, the incorporation of IMP is inherently inefficient in contemporary QLC NAND flash chips. In this paper, we propose an efficient independent read (EFFIR) scheme to tackle this challenge. EFFIR features a read latency variation aware transaction service that properly combines read transactions to minimize idle time and proactively transfers read data to eliminate unnecessary delays. Moreover, EFFIR incorporates a read latency variation aware RAIN that reorganizes the layout of data and parity to mitigate the impact of high-latency data access on idle time. Our comprehensive experimental results demonstrate and elucidate how EFFIR significantly enhances SSD responsiveness while consistently delivering favorable performance across a diverse range of read-intensive workloads. Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001, Feng Zhu 0024, Maojun Yuan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | A Read Latency Variation Aware Independent Read Scheme for QLC SSDsabstractQLC flash-based SSDs has attracted growing interest and is expected to fit in read-intensive scenarios owing to its higher cost-effectiveness and shorter write endurance compared with the Triple-Level Cell (TLC) SSDs. Recently, new commands supporting independent reads such as Single Operation Multiple Locations (SOML) and Independent multi-plane (IMP) read are proposed to improve read performance. Unfortuantely, while independent read exhibits a significant performance improvement, we show in this paper that the exisiting approach fails to fully exploit its potential due to the larger read latency variation and more planes per die for current SSD architecture. Through a set of experiments, we demonstrate that the lack of a read latency variation aware machanism leads to low performance of independent read among a wide variety of workloads. To alleviate this issue, we propose LITA, which key idea is to combine read transactions with similar latency into one command. LITA includes (1) LIT, a latency variation aware transaction combination, (2) and TASAP, a latency variation aware transaction completion service. The experimental results show LITA can reduce read latency by 20.4% and 9.6% on average for 4-planes QLC SSDs under IMP and SOML, respectively. Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001 |
DATE | 4 |
| 2024 | Enabling Reliable Memory-Mapped I/O With Auto-Snapshot for Persistent Memory SystemsabstractPersistent memory (PM) is promising to be the next-generation storage device with better I/O performance. Since the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of being corrupted by scribbles and non-atomic I/O interfaces, causing serious reliability issues. To address these issues, we propose RMMIO, an efficient user-space library that provides reliable MMIO for PM systems. RMMIO provides atomic I/O interfaces and lightweight snapshots to ensure the reliability of MMIO. Compared with existing schemes, RMMIO mitigates additional writes and extra software overheads caused by reliability guarantees, thus achieving MMIO-like performance. In addition, we also propose an automatic snapshot with efficient memory management for RMMIO to minimize data loss incurred by reliability issues. The experimental results of microbenchmarks show that RMMIO achieves 8.49x and 2.31x higher throughput than ext4-DAX and the state-of-the-art MMIO-based scheme, respectively, while ensuring data reliability. The real-world application accelerated by RMMIO achieves at most 7.06x higher throughput than that of ext4-DAX. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
IEEE Trans. Computers | 1 |
| 2023 | Lock-Free High-performance Hashing for Persistent Memory via PM-aware Holistic OptimizationabstractPersistent memory (PM) provides large-scale non-volatile memory (NVM) with DRAM-comparable performance. The non-volatility and other unique characteristics of PM architecture bring new opportunities and challenges for the efficient storage system design. For example, some recent crash-consistent and write-friendly hashing schemes are proposed to provide fast queries for PM systems. However, existing PM hashing indexes suffer from the concurrency bottleneck due to the blocking resizing and expensive lock-based concurrency control for queries. Moreover, the lack of PM awareness and systematical design further increases the query latency. To address the concurrency bottleneck of lock contention in PM hashing, we propose clevel hashing, a lock-free concurrent level hashing scheme that provides non-blocking resizing via background threads and lock-free search/insertion/update/deletion using atomic primitives to enable high concurrency for PM hashing. By exploiting the PM characteristics, we present a holistic approach to building clevel hashing for high throughput and low tail latency via the PM-aware index/allocator co-design. The proposed volatile announcement array with a helping mechanism coordinates lock-free insertions and guarantees a strong consistency model. Our experiments using real-world YCSB workloads on Intel Optane DC PMM show that clevel hashing, respectively, achieves up to 5.7× and 1.6× higher throughput than state-of-the-art P-CLHT and Dash while guaranteeing low tail latency, e.g., 1.9×–7.2× speedup for the p99 latency with the insert-only workload. Zhangyu Chen, Yu Hua 0001, Luochangqi Ding, Bo Ding 0002, Pengfei Zuo, Xue (Steve) Liu |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | SplitZNS: Towards an Efficient LSM-Tree on Zoned Namespace SSDsabstractThe Zoned Namespace (ZNS) Solid State Drive (SSD) is a nascent form of storage device that offers novel prospects for the Log Structured Merge Tree (LSM-tree). ZNS exposes erase blocks in SSD as append-only zones, enabling the LSM-tree to gain awareness of the physical layout of data. Nevertheless, LSM-tree on ZNS SSDs necessitates Garbage Collection (GC) owing to the mismatch between the gigantic zones and relatively small Sorted String Tables (SSTables). Through extensive experiments, we observe that a smaller zone size can reduce data migration in GC at the cost of a significant performance decline owing to inadequate parallelism exploitation. In this article, we present SplitZNS, which introduces small zones by tweaking the zone-to-chip mapping to maximize GC efficiency for LSM-tree on ZNS SSDs. Following the multi-level peculiarity of LSM-tree and the inherent parallel architecture of ZNS SSDs, we propose a number of techniques to leverage and accelerate small zones to alleviate the performance impact due to underutilized parallelism. (1) First, we use small zones selectively to prevent exacerbating write slowdowns and stalls due to their suboptimal performance. (2) Second, to enhance parallelism utilization, we propose SubZone Ring, which employs a per-chip FIFO buffer to imitate a large zone writing style; (3) Read Prefetcher, which prefetches data concurrently through multiple chips during compactions; (4) and Read Scheduler, which assigns query requests the highest priority. We build a prototype integrated with SplitZNS to validate its efficiency and efficacy. Experimental results demonstrate that SplitZNS achieves up to 2.77× performance and reduces data migration considerably compared to the lifetime-based data placement. 1 Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | RMMIO: Enabling Reliable Memory-Mapped I/O for Persistent Memory SystemsabstractThe byte-addressable persistent memory (PM) is coming to be the next-generation storage device for better I/O performance. As the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works have proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of user-space scribbles and non-atomic I/O interfaces, termed reliability issues. Since existing reliability schemes cause significant extra overheads, we propose RMMIO, an efficient user-space library that provides reliable memory-mapped I/O interfaces for PM systems. RMMIO achieves a good balance between efficiency and reliability by introducing a memory-mapped cache layer upon kernel file systems. The cache layer accelerates I/O requests and carries the file system’s responsibility for data reliability by data isolation. In addition, RMMIO further employs lightweight snapshots and efficient atomic I/O interfaces to guarantee the integrity and consistency of the data in the cache layer at low costs. The experimental results show that RMMIO achieves 8.49x higher throughput than ext4-DAX and 2.31x higher throughput than state-of-the-art MMIO-based schemes for PM while ensuring data reliability. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
ICCD | 1 |
| 2020 | Lock-free Concurrent Level Hashing for Persistent Memory
Zhangyu Chen, Yu Hua 0001, Bo Ding 0002, Pengfei Zuo |
USENIX ATC | 3 |