Bohong Zhu

dblp:247/9449 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0004-5763-1740ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AtRS: Auto-tuning RAID system with GAN
Fangzheng Wang, Congming Gao, Bohong Zhu, Jiwu Shu
Future Gener. Comput. Syst.3
2024 PireSPM: Efficient and Recoverable Secure Persistent Memory for Multi-cores
abstract
Secure persistent memory (PM) systems have to guarantee the crash consistency for secure data, which requires atomically persisting both data and its security metadata (e.g., the counter and Merkle tree nodes). Existing work introduces persistent registers and uses the redo-style method to ensure the atomic durability of the data and security metadata for a single-core system. Unfortunately, their performance decreases sharply in multi-core secure PM systems because the atomic durability guarantee procedure for each write request is serialized.In this work, we propose PireSPM, a pipeline-based and recoverable scheme for secure persistent memory for multi-core architecture. PireSPM supports processing multiple atomic durability requests at a time without incurring recoverability problems to the secure PM. We further optimize PireSPM by combining Merkle tree updates and reducing the latency incurred by the atomic durability guarantee without breaking the system’s recoverability. Experimental results show that PireSPM with all these optimizations achieve up to 1.49× speedup and 4.97× speedup respectively over the state-of-the-art method in 1-core and 16-core systems.
Bohong Zhu, Jiwu Shu, Zhengyong Wang
CCGrid2
2024 ZUFS: Enhancing Stability and Endurance in Mobile Devices with Integrated Zoned Namespaces in Universal Flash Storage
abstract
This paper presents Zoned UFS (ZUFS), an innovative approach to address the inherent challenges in traditional Universal Flash Storage (UFS) systems in mobile devices. ZUFS integrates Zoned Namespaces (ZNS) with the UFS framework, aiming to optimize storage efficiency, reduce latency variations, and enhance endurance. We introduced key innovations, including a dual data path design and a host-based Flash Translation Layer (FTL), specifically designed to reduce overprovisioning and address the Write Amplification Factor (WAF) issues common in Triple-Level Cell (TLC) NAND technology. Through comprehensive experiments, ZUFS demonstrates significant improvements in storage performance, including reduced tail latency and better capacity utilization, compared to conventional UFS-based systems. The findings indicate that ZUFS not only offers a viable solution to current storage limitations in mobile platforms but also sets a new direction for future advancements in mobile storage technology.
Pengbo Yan 0002, Bohong Zhu, Zhirong Shen, Jiwu Shu, Jiadong Yang
CCGrid2
2024 Exploring the Asynchrony of Slow Memory Filesystem with EasyIO
abstract
We introduce EasyIO, a new approach to explore asynchronous I/O on filesystems designed for (disaggregated) nonvolatile memories to improve CPU efficiency. EasyIO offloads expensive memory movement operations to the on-chip DMA engine and harvests the unleashed CPU cycles by transparently interleaving asynchronous I/Os with fine-grained application tasks. We further adopt a completion buffer-centric design to improve EasyIO's efficiency and schedulability; internally, orderless file operation and two-level locking are incorporated to break the serial order between file metadata and data, thus accelerating read and write operations and defusing deadlock risks. EasyIO also introduces a traffic-aware channel manager to fulfill the diverse performance goals of applications. Extensive experimental results show that, compared to conventional synchronous filesystems, EasyIO significantly reduces CPU consumption (using less than 88% of cores at most) while achieving comparable peak bandwidth; EasyIO also achieves 1.03-2.3× speedups across eight real-world applications. When achieving these goals, EasyIO exhibits higher but tolerable latencies for read operations due to the task interleaving.
Bohong Zhu, Youmin Chen, Jiwu Shu
EuroSys1
2021 Scalable Persistent Memory File System with Kernel-Userspace Collaboration
Youmin Chen, Youyou Lu, Bohong Zhu, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau, Jiwu Shu
FAST3
2021 Octopus+: An RDMA-Enabled Distributed Persistent Memory File System
abstract
Non-volatile memory and remote direct memory access (RDMA) provide extremely high performance in storage and network hardware. However, existing distributed file systems strictly isolate file system and network layers, and the heavy layered software designs leave high-speed hardware under-exploited. In this article, we propose an RDMA-enabled distributed persistent memory file system, Octopus + , to redesign file system internal mechanisms by closely coupling non-volatile memory and RDMA features. For data operations, Octopus + directly accesses a shared persistent memory pool to reduce memory copying overhead, and actively fetches and pushes data all in clients to rebalance the load between the server and network. For metadata operations, Octopus + introduces self-identified remote procedure calls for immediate notification between file systems and networking, and an efficient distributed transaction mechanism for consistency. Octopus + is enabled with replication feature to provide better availability. Evaluations on Intel Optane DC Persistent Memory Modules show that Octopus + achieves nearly the raw bandwidth for large I/Os and orders of magnitude better performance than existing distributed file systems.
Bohong Zhu, Youmin Chen, Qing Wang 0031, Youyou Lu, Jiwu Shu
ACM Trans. Storage1
2020 TH-DPMS: Design and Implementation of an RDMA-enabled Distributed Persistent Memory Storage System
abstract
The rapidly increasing data in recent years requires the datacenter infrastructure to store and process data with extremely high throughput and low latency. Fortunately, persistent memory (PM) and RDMA technologies bring new opportunities towards this goal. Both of them are capable of delivering more than 10 GB/s of bandwidth and sub-microsecond latency. However, our past experiences and recent studies show that it is non-trivial to build an efficient and distributed storage system with such new hardware. In this article, we design and implement TH-DPMS (TsingHua Distributed Persistent Memory System) based on persistent memory and RDMA, which unifies the memory, file system, and key-value interface in a single system. TH-DPMS is designed based on a unified distributed persistent memory abstract, pDSM. pDSM acts as a generic layer to connect the PMs of different storage nodes via high-speed RDMA network and organizes them into a global shared address space. It provides the fundamental functionalities, including global address management, space management, fault tolerance, and crash consistency guarantees. Applications are enabled to access pDSM with a group of flexible and easy-to-use APIs by using either raw read/write interfaces or the transactional ones with ACID guarantees. Based on pDSM, we implement a distributed file system and a key-value store named pDFS and pDKVS, respectively. Together, they uphold TH-DPMS with high-performance, low-latency, and fault-tolerant data storage. We evaluate TH-DPMS with both micro-benchmarks and real-world memory-intensive workloads. Experimental results show that TH-DPMS is capable of delivering an aggregated bandwidth of 120 GB/s with 6 nodes. When processing memory-intensive workloads such as YCSB and Graph500, TH-DPMS improves the performance by one order of magnitude compared to existing systems and keeps consistent high efficiency when the workload size grows to multiple terabytes.
Jiwu Shu, Youmin Chen, Qing Wang 0031, Bohong Zhu, Youyou Lu
ACM Trans. Storage4