Zhe Yang 0012

dblp:181/2876-12 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-2031-5937ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Switch$\Delta$: Asynchronous Metadata Updating for Distributed Storage with in-Network Data Visibility
abstract
Distributed storage systems typically maintain strong consistency between data nodes and metadata nodes by adopting ordered writes: 1) first installing data; 2) then updating metadata to make data visible.We propose SwitchDelta to accelerate ordered writes by moving metadata updates out of the critical path. It buffers in-flight metadata updates in programmable switches to enable data visibility in the network and retain strong consistency. SwitchDelta uses a best-effort data plane design to overcome the resource limitation of switches and designs a novel metadata update protocol to exploit the benefits of in-network data visibility. We evaluate SwitchDelta in three distributed in-memory storage systems: log-structured key-value stores, file systems, and secondary indexes. The evaluation shows that SwitchDelta reduces the latency of write operations by up to 52.4% and boosts the throughput by up to 126.9% under write-heavy workloads.
Qing Wang 0031, Zhe Yang 0012, Jiwu Shu, Youyou Lu
ICDE3
2025 Efficiently Enlarging RDMA-Attached Memory with SSD
abstract
RDMA-based in-memory storage systems offer high performance but are restricted by the capacity of physical memory. In this article, we propose TeRM to extend RDMA-attached memory with SSD. TeRM achieves fast remote access on the SSD-extended memory by eliminating page faults of RDMA NIC and CPU from the critical path. We also introduce a set of techniques to reduce the consumption of CPU and network resources. Evaluation shows that TeRM performs close to the performance of the ideal upper bound where all pages are pinned in the physical memory. Compared with existing approaches, TeRM significantly improves the performance of unmodified RDMA-based storage systems, including a file system and a key-value system.
Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu
ACM Trans. Storage1
2024 TeRM: Extending RDMA-Attached Memory with SSD
Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu
FAST1
2023 RIO: Order-Preserving and CPU-Efficient Remote Storage Access
abstract
Modern NVMe SSDs and RDMA networks provide dramatically higher bandwidth and concurrency. Existing networked storage systems (e.g., NVMe over Fabrics) fail to fully exploit these new devices due to inefficient storage ordering guarantees. Severe synchronous execution for storage order in these systems stalls the CPU and I/O devices and lowers the CPU and I/O performance efficiency of the storage system.
Xiaojian Liao, Zhe Yang 0012, Jiwu Shu
EuroSys2
2023 λ-IO: A Unified IO Stack for Computational Storage
Zhe Yang 0012, Youyou Lu, Xiaojian Liao, Youmin Chen, Siyu He, Jiwu Shu
FAST1
2023 Revisiting Swapping in User-Space With Lightweight Threading
abstract
Memory-intensive applications, such as in-memory databases, caching systems, and key-value stores, are increasingly demanding larger main memory to fit their working sets. Conventional swapping can enlarge the memory capacity by paging out inactive pages to backend stores. However, existing swapping solutions suffer several performance and compatibility issues, making them unsuitable for high-concurrency and memory-intensive applications. In this article, we redesign the swapping system and propose Lightswap, a high-performance user-space swapping solution that supports paging with both local SSDs and remote memories. First, to avoid kernel involvement, we propose to leverage the extended Berkeley packet filter (eBPF) for handling page faults (PFs) in user space and further eliminate the heavy I/O stack with the help of user-space I/O drivers. Then, we co-design the PF handling with lightweight thread (LWT) scheduling to improve system throughput and reduce the end-to-end PF latency. Finally, we propose a try-catch framework in Lightswap to deal with swap-in errors which have been exacerbated by the scaling in process technology. We implement Lightswap in our production-level system and evaluate it with various benchmarks. Results show that Lightswap achieves scalable PF notification latency ($4 \mu \text{s}$under 128 LWTs), reduces the PF handling latency by 3–5 times, and improves the throughput of memcached by more than 40% compared with the state-of-art swapping systems.
Kan Zhong, Wenlin Cui, Qiao Li 0001, Zhe Yang 0012, Youyou Lu, Xiaodan Yan, Siwei Luo, Qizhao Yuan, Keji Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Efficient Crash Consistency for NVMe over PCIe and RDMA
abstract
This article presents crash-consistent Non-Volatile Memory Express (ccNVMe), a novel extension of the NVMe that defines how host software communicates with the non-volatile memory (e.g., solid-state drive) across a PCI Express bus and RDMA-capable networks with both crash consistency and performance efficiency. Existing storage systems pay a huge tax on crash consistency, and thus cannot fully exploit the multi-queue parallelism and low latency of the NVMe and RDMA interfaces. ccNVMe alleviates this major bottleneck by coupling the crash consistency to the data dissemination. This new idea allows the storage system to achieve crash consistency by taking the free rides of the data dissemination mechanism of NVMe, using only two lightweight memory-mapped I/Os (MMIOs), unlike traditional systems that use complex update protocol and synchronized block I/Os. ccNVMe introduces a series of techniques including transaction-aware MMIO/doorbell and I/O command coalescing to reduce the PCIe traffic as well as to provide atomicity. We present how to build a high-performance and crash-consistent file system named MQFS atop ccNVMe. We experimentally show that MQFS increases the IOPS of RocksDB by 36% and 28% compared to a state-of-the-art file system and Ext4 without journaling, respectively.
Xiaojian Liao, Youyou Lu, Zhe Yang 0012, Jiwu Shu
ACM Trans. Storage3
2022 AlNiCo: SmartNIC-accelerated Contention-aware Request Scheduling for Transaction Processing
Youyou Lu, Qing Wang 0031, Jiazhen Lin, Zhe Yang 0012, Jiwu Shu
USENIX ATC5
2021 Crash Consistent Non-Volatile Memory Express
abstract
This paper presents crash consistent Non-Volatile Memory Express (ccNVMe), a novel extension of the NVMe that defines how host software communicates with the non-volatile memory (e.g., solid-state drive) across a PCI Express bus with both crash consistency and performance efficiency. Existing storage systems pay a huge tax on crash consistency, and thus can not fully exploit the multi-queue parallelism and low latency of the NVMe interface. ccNVMe alleviates this major bottleneck by coupling the crash consistency to the data dissemination. This new idea allows the storage system to achieve crash consistency by taking the free rides of the data dissemination mechanism of NVMe, using only two lightweight memory-mapped I/Os (MMIO), unlike traditional systems that use complex update protocol and heavyweight block I/Os. ccNVMe introduces transaction-aware MMIO and doorbell to reduce the PCIe traffic as well as to provide atomicity. We present how to build a high-performance and crash-consistent file system namely MQFS atop ccNVMe. We experimentally show that MQFS increases the IOPS of RocksDB by 36% and 28% compared to a state-of-the-art file system and Ext4 without journaling, respectively.
Xiaojian Liao, Youyou Lu, Zhe Yang 0012, Jiwu Shu
SOSP3
2020 CoinPurse: A Device-Assisted File System with Dual Interfaces
abstract
Block I/O serves as a classic interface for accessing storage devices with portability. But it can also cause extra overhead by enforcing transferring data in the unit of blocks. In this paper, we present CoinPurse, a device-assisted file system with dual interfaces. By leveraging non-volatile memory (NVM) in SSD, CoinPurse manages to adaptively persist writes through both the block I/O and a byte-addressable partial update interface. In addition, we also develop a set of techniques to overcome hardware limitations and resolve possible consistency conflicts. Evaluation shows that CoinPurse outperforms F2FS, a popular flash-optimized file system, by up to 33.2%.
Zhe Yang 0012, Youyou Lu, Erci Xu, Jiwu Shu
DAC1
2020 OCVM: Optimizing the Isolation of Virtual Machines with Open-Channel SSDs
Xiaojian Liao, Zhe Yang 0012, Youyou Lu, Jiwu Shu
ICA3PP (1)4
2020 SineKV: Decoupled Secondary Indexing for LSM-based Key-Value Stores
abstract
Secondary indexing is highly demanded for key-value stores by many applications to accelerate query performance. Current secondary indices on key-value stores are typically built on top of the primary index. In a secondary key query, the primary index has to be accessed to fetch the records, with the retrieved primary keys from the secondary index. The record fetching process invokes lots of point lookups in the primary index and exacerbates the read amplification. In this paper, we present SineKV, a decoupled Secondary indexing Key-Value store, aiming to avoid fetching records from the primary index and improve the secondary key query performance. Firstly, SineKV separates the records from the indices and keeps each index pointing to the record values independently. Secondly, SineKV proposes a mapping-based lazy index maintenance strategy to ensure the consistency of secondary indices. Finally, SineKV leverages the CMB feature of the underlying NVMe SSDs to guarantee crash consistency. We implement and evaluate SineKV against LevelDB and Wisc-Key based designs. The evaluations show SineKV outperforms LevelDB and WiscKey based systems by up to 6.12× and 2.78× under microbenchmark and mixed workloads.
Youyou Lu, Zhe Yang 0012, Jiwu Shu
ICDCS3
2019 OCStore: Accelerating Distributed Object Storage with Open-Channel SSDs
abstract
SSDs are getting widely used in data centers. It is a critical issue to design efficient software for exploiting the benefits of fast SSD hardware. In this paper, we propose OCStore, an object store based on open-channel SSDs for distributed object storage system. OCStore manages the objects directly on raw flash memory, mitigating redundant functions across the object store, the file system, and the FTL layers. It provides streamed transactional update, which not only ensures the multi-page atomicity leveraging the non-overwrite flash write features, but also provides isolation for independent I/O streams while enabling parallel accesses to different channels. OCStore also coordinates different channels to enable transaction-aware scheduling, so as to reduce transaction-level latency and provide low response time to distributed storage. We implement OCStore in Linux kernel on the real open-channel SSDs, and evaluate them as OSDs in Ceph. Evaluations show that OCStore outperforms state-of-the-art object stores by 1.5× to 3.0×, while providing much lower and stable latencies, and decreases up to 70% write traffic under heavy workloads.
Youyou Lu, Zhe Yang 0012, Liyang Pan, Jiwu Shu
ICDCS3
2019 Cognitive SSD: A Deep Learning Engine for In-Storage Data Retrieval
Shengwen Liang, Ying Wang 0001, Youyou Lu, Zhe Yang 0012, Huawei Li 0001, Xiaowei Li 0001
USENIX ATC4