VLDB 2026 Research / reviewers in the wild / expert
Zhangyu Chen
dblp:241/5080
· DBLP profile ↗
15ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-9020-3693ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MPFS: A Scalable User-Space Persistent Memory File System for Multiple Processesabstract11This work was supported by the Young Scientists Fund of the National Natural Science Foundation of China under Grant 62302182.Persistent memory (PM) leveraging memory-mapped I/O(MMIO) delivers superior I/O performance, leading to the development of user-space PM file systems based on MMIO. While effective in single-process scenarios, these systems encounter challenges in multi-process environments, such as performance degradation due to repeated page faults and cross-process synchronizations, as well as a large memory footprint from duplicated paging structures. To address these problems, we propose a Multi-process PM File System (MPFS). MPFS builds a shareable page table and shares it among processes, avoiding building duplicate paging structures for distinct processes, thereby significantly reducing the software overhead and memory footprint caused by repeated page faults. MPFS further proposes a PGD-aligned (512GB) mapping method to accelerate page table sharing. Furthermore, MPFS provides a cross-process memory protection mechanism based on the PGD-aligned mapping, ensuring multi-process data reliability with negligible overheads. The experimental results show that MPFS outperforms existing user-space PM file systems by 1560% in multi-process scenarios. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Yuchong Hu, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
DATE | 7 |
| 2025 | GPHash: An Efficient Hash Index for GPU with Byte-Granularity Persistent Memory
Menglei Chen, Yu Hua 0001, Zhangyu Chen, Gen Dong |
FAST | 3 |
| 2025 | Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud Systems
Gen Dong, Yu Hua 0001, Yongle Zhang 0007, Zhangyu Chen, Menglei Chen |
USENIX ATC | 4 |
| 2024 | Approximate Similarity-Aware Compression for Non-Volatile Main Memory
Zhangyu Chen, Yu Hua 0001, Pengfei Zuo, Yuan-Yuan Sun, Yuncheng Guo |
J. Comput. Sci. Technol. | 1 |
| 2024 | Enabling Reliable Memory-Mapped I/O With Auto-Snapshot for Persistent Memory SystemsabstractPersistent memory (PM) is promising to be the next-generation storage device with better I/O performance. Since the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of being corrupted by scribbles and non-atomic I/O interfaces, causing serious reliability issues. To address these issues, we propose RMMIO, an efficient user-space library that provides reliable MMIO for PM systems. RMMIO provides atomic I/O interfaces and lightweight snapshots to ensure the reliability of MMIO. Compared with existing schemes, RMMIO mitigates additional writes and extra software overheads caused by reliability guarantees, thus achieving MMIO-like performance. In addition, we also propose an automatic snapshot with efficient memory management for RMMIO to minimize data loss incurred by reliability issues. The experimental results of microbenchmarks show that RMMIO achieves 8.49x and 2.31x higher throughput than ext4-DAX and the state-of-the-art MMIO-based scheme, respectively, while ensuring data reliability. The real-world application accelerated by RMMIO achieves at most 7.06x higher throughput than that of ext4-DAX. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
IEEE Trans. Computers | 4 |
| 2023 | ROLEX: A Scalable RDMA-oriented Learned Key-Value Store for Disaggregated Memory Systems
Yu Hua 0001, Pengfei Zuo, Zhangyu Chen, Jiajie Sheng |
FAST | 4 |
| 2023 | Lock-Free High-performance Hashing for Persistent Memory via PM-aware Holistic OptimizationabstractPersistent memory (PM) provides large-scale non-volatile memory (NVM) with DRAM-comparable performance. The non-volatility and other unique characteristics of PM architecture bring new opportunities and challenges for the efficient storage system design. For example, some recent crash-consistent and write-friendly hashing schemes are proposed to provide fast queries for PM systems. However, existing PM hashing indexes suffer from the concurrency bottleneck due to the blocking resizing and expensive lock-based concurrency control for queries. Moreover, the lack of PM awareness and systematical design further increases the query latency. To address the concurrency bottleneck of lock contention in PM hashing, we propose clevel hashing, a lock-free concurrent level hashing scheme that provides non-blocking resizing via background threads and lock-free search/insertion/update/deletion using atomic primitives to enable high concurrency for PM hashing. By exploiting the PM characteristics, we present a holistic approach to building clevel hashing for high throughput and low tail latency via the PM-aware index/allocator co-design. The proposed volatile announcement array with a helping mechanism coordinates lock-free insertions and guarantees a strong consistency model. Our experiments using real-world YCSB workloads on Intel Optane DC PMM show that clevel hashing, respectively, achieves up to 5.7× and 1.6× higher throughput than state-of-the-art P-CLHT and Dash while guaranteeing low tail latency, e.g., 1.9×–7.2× speedup for the p99 latency with the insert-only workload. Zhangyu Chen, Yu Hua 0001, Luochangqi Ding, Bo Ding 0002, Pengfei Zuo, Xue (Steve) Liu |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | APPcache+: An STT-MRAM-Based Approximate Cache System With Low Power and Long LifetimeabstractDue to high static power and low scalability, the traditional SRAM-based cache is not a good solution for image processing applications. Emerging spin transfer torque magnetic RAM (STT-MRAM) is a promising candidate for cache due to its low leakage power and high density. However, STT-MRAM suffers from high write energy. Therefore, by making use of the ability of tolerating minor errors in image processing applications, this work presents an STT-MRAM-basedAPProximatecachearchitecture (APPcache+) to write/read approximate data, which can largely reduce the cache energy and improve the STT-MRAM lifetime. APPcache+ includes three main designs. First, we find that there are many similar elements (e.g., pixels in images) in cache lines. Therefore, APPcache+ presents several lightweight similarity-based encoding techniques to remove redundant elements, thus, shortening the data size and reducing the energy of STT-MRAM cache. Second, we design a partial read scheme to reduce the read energy of the STT-MRAM cache. In the traditional decompression process, the whole line is fetched into the decompressor, leading to unnecessary read energy. The partial read scheme can largely reduce read energy while keeping the overhead low. Third, we observe the encoding schemes may lead to bit write imbalance. Therefore, we propose a lightweight Ping-Pong intraline wear-leveling scheme to improve the lifetime. Compared with the baseline, extensive evaluation results show that our APPcache+ can largely reduce the overall energy by 32.58%, improve lifetime by 40.7% with only 2.2% performance degradation, and 1.86% output quality loss. Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Zhangyu Chen, Bing Wu 0001, Chengning Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | A High-performance RDMA-oriented Learned Key-value Store for Disaggregated Memory SystemsabstractDisaggregated memory systems separate monolithic servers into different components, including compute and memory nodes, to enjoy the benefits of high resource utilization, flexible hardware scalability, and efficient data sharing. By exploiting the high-performance RDMA (Remote Direct Memory Access), the compute nodes directly access the remote memory pool without involving remote CPUs. Hence, the ordered key-value (KV) stores (e.g., B-trees and learned indexes) keep all data sorted to provide range query services via the high-performance network. However, existing ordered KVs fail to work well on the disaggregated memory systems, due to either consuming multiple network roundtrips to search the remote data or heavily relying on the memory nodes equipped with insufficient computing resources to process data modifications. In this article, we propose a scalable RDMA-oriented KV store with learned indexes, called ROLEX, to coalesce the ordered KV store in the disaggregated systems for efficient data storage and retrieval. ROLEX leverages a retraining-decoupled learned index scheme to dissociate the model retraining from data modification operations via adding a bias and some data movement constraints to learned models. Based on the operation decoupling, data modifications are directly executed in compute nodes via one-sided RDMA verbs with high scalability. The model retraining is hence removed from the critical path of data modification and asynchronously executed in memory nodes by using dedicated computing resources. ROLEX efficiently alleviates the fragmentation and garbage collection issues, due to allocating and reclaiming space via fixed-size leaves that are accessed via the atomic-size leaf numbers. Our experimental results on YCSB and real-world workloads demonstrate that ROLEX achieves competitive performance on the static workloads, as well as significantly improving the performance on dynamic workloads by up to 2.2× over state-of-the-art schemes on the disaggregated memory systems. We have released the open-source codes for public use in GitHub. Yu Hua 0001, Pengfei Zuo, Zhangyu Chen, Jiajie Sheng |
ACM Trans. Storage | 4 |
| 2022 | Efficiently detecting concurrency bugs in persistent memory programsabstractDue to the salient DRAM-comparable performance, TB-scale capacity, and non-volatility, persistent memory (PM) provides new opportunities for large-scale in-memory computing with instant crash recovery. However, programming PM systems is error-prone due to the existence of crash-consistency bugs, which are challenging to diagnose especially with concurrent programming widely adopted in PM applications to exploit hardware parallelism. Existing bug detection tools for DRAM-based concurrency issues cannot detect PM crash-consistency bugs because they are oblivious to PM operations and PM consistency. On the other hand, existing PM-specific debugging tools only focus on sequential PM programs and cannot effectively detect crash-consistency issues hidden in concurrent executions. Zhangyu Chen, Yu Hua 0001, Yongle Zhang 0007, Luochangqi Ding |
ASPLOS | 1 |
| 2022 | RMMIO: Enabling Reliable Memory-Mapped I/O for Persistent Memory SystemsabstractThe byte-addressable persistent memory (PM) is coming to be the next-generation storage device for better I/O performance. As the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works have proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of user-space scribbles and non-atomic I/O interfaces, termed reliability issues. Since existing reliability schemes cause significant extra overheads, we propose RMMIO, an efficient user-space library that provides reliable memory-mapped I/O interfaces for PM systems. RMMIO achieves a good balance between efficiency and reliability by introducing a memory-mapped cache layer upon kernel file systems. The cache layer accelerates I/O requests and carries the file system’s responsibility for data reliability by data isolation. In addition, RMMIO further employs lightweight snapshots and efficient atomic I/O interfaces to guarantee the integrity and consistency of the data in the cache layer at low costs. The experimental results show that RMMIO achieves 8.49x higher throughput than ext4-DAX and 2.31x higher throughput than state-of-the-art MMIO-based schemes for PM while ensuring data reliability. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
ICCD | 4 |
| 2021 | Improving the energy efficiency of STT-MRAM based approximate cacheabstractApproximate computing applications lead to large energy consumption and performance demand for the memory system. However, traditional SRAM based cache cannot satisfy these demands due to high leakage power and limited density. Spin Transfer Torque Magnetic RAM (STT-MRAM) is a promising candidate of cache due to low leakage power and high density. However, STT-MRAM suffers from high write energy. To leverage the ability of tolerating acceptable quality loss via approximations to data, we propose an STT-MRAM based APProximate cache architecture (APPcache) to write/read approximate data thus largely reducing energy. We find many similar elements (e.g. pixels in images) existing in cache lines while running approximate computing applications. Therefore, APPcache uses several lightweight similarity-based encoding schemes to eliminate the similar elements to reduce the data size thus reducing the write energy of STT-MRAM based cache. Besides, we design a software interface to manually control the output quality. APPcache can significantly eliminate similar elements, thus improving energy efficiency. Experimental results show that our scheme can reduce write energy and improve the image raw data compression ratio by 21.9% and 38.0% compared with the state-of-the-art scheme with 1 % error rate, respectively. As for the output quality, the losses of all benchmarks are within 5% with 1 % error rate. Wei Zhao 0034, Wei Tong 0001, Dan Feng 0001, Jingning Liu, Zhangyu Chen, Jie Xu 0013, Bing Wu 0001, Chengning Wang, Bo Liu 0057 |
DATE | 5 |
| 2020 | Reducing Bit Writes in Non-volatile Main Memory by Similarity-aware CompressionabstractVarious applications use image bitmaps (data containing pixels) in main memory for fast accesses, thereby leading to lots of memory consumption. Unlike legacy DRAM, nonvolatile memories (NVMs) have larger capacity. However, NVM writes consume higher energy and latency compared with reads. Existing data compression schemes leverage precise general-purpose data patterns or precision scaling to reduce data sizes, which suffer from limited compression performance for bitmaps due to large variance or serious quality loss. By exploiting the pixel-level similarity due to the analogous contents in adjacent pixels, we propose SimCom, an approximate Similarity-aware Compression scheme, to compress the write accesses to bitmaps in NVMs, thus efficiently improving the memory performance for image/video applications. SimCom reduces the data size by compressing data into base words and runs. The storage costs for small runs are further mitigated by reusing the least significant bits of base words. The adaptive compression scheme handles various data formats without user annotations on data types. Our experimental results with real-world image/video workloads demonstrate the efficacy and efficiency of SimCom. Zhangyu Chen, Yu Hua 0001, Pengfei Zuo, Yuncheng Guo |
DAC | 1 |
| 2020 | Lock-free Concurrent Level Hashing for Persistent Memory
Zhangyu Chen, Yu Hua 0001, Bo Ding 0002, Pengfei Zuo |
USENIX ATC | 1 |
| 2019 | Mitigating Asymmetric Read and Write Costs in Cuckoo Hashing for Storage Systems
Yu Hua 0001, Zhangyu Chen, Yuncheng Guo |
USENIX ATC | 3 |