VLDB 2026 Research / reviewers in the wild / expert
Xiaomin Zou
dblp:292/5854
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-5867-847XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RED-ANNS: A RDMA-Enabled Distributed Framework for Graph-Based Approximate Nearest Neighbor Search
Yue Chen 0029, Kai Zhang 0006, Sipeng Chen, Shihai Xiao, Xiaomin Zou, Yinan Jing, Xiaoyang Sean Wang, Mingxiang Wan |
Proc. VLDB Endow. | 5 |
| 2024 | A Scalable and Write-Optimized Disaggregated B+-Tree With Adaptive Cache AssistanceabstractDisaggregated memory (DM) architecture separates CPU and DRAM into computing/memory resource pools and interconnects them with high-speed networks. Storage systems on DM locate data by distributed index. However, existing distributed indexes either suffer from prohibitive synchronization overhead of write operation or sacrifice the performance of read operation, resulting in low throughput, high tail latency, and challenging trade-off. In this paper, we present Marlin+, a scalable and optimized B+-tree on DM. Marlin+ provides atomic granularity synchronization between write operations via three strategies: 1) a concurrent algorithm that is friendly to IDU operations (Insert, Delete, and Update), enabling different clients to concurrently operate on the same leaf node, 2) shared-exclusive leaf node lock, effectively preventing conflicts between index structure modification operation (SMO) and IDU operations, and 3) critical path compression of write to reduce latency of write operation. Moreover, Marlin+ proposes an adaptive remote address cache to accelerate the access of hot data. Compared to the state-of-the-art schemes based on DM, Marlin achieves 2.21× higher throughput and 83.4% lower P99 latency under YCSB hybrid workloads. Compared to Marlin, Marlin+ improves the throughput by up to 1.58× and reduces the P50 latency by up to 50.5% under YCSB read-intensive workloads. Hang An, Fang Wang 0001, Dan Feng 0001, Xiaomin Zou, Zefeng Liu, Jianshun Zhang |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | Marlin: A Concurrent and Write-Optimized B+-tree Index on Disaggregated MemoryabstractMemory disaggregation architecture can achieve higher resource utilization, independent scaling of CPUs and memory. Disaggregated memory systems manage memory resources and locate data by distributed index. However, existing distributed indexes suffer from high synchronization overhead of write, due to naive node lock, thus resulting in high tail latency, low throughput, and poor scalability. Hang An, Fang Wang 0001, Dan Feng 0001, Xiaomin Zou, Zefeng Liu, Jianshun Zhang |
ICPP | 4 |
| 2023 | Fast One-Sided RDMA-Based State Machine Replication for Disaggregated MemoryabstractDisaggregated memory architecture has risen in popularity for large datacenters with the advantage of improved resource utilization, failure isolation, and elasticity. Replicated state machines (RSMs) have been extensively used for reliability and consistency. In traditional RSM protocols, each replica stores replicated data and has the computing power to participate in some part of the protocols. However, traditional RSM protocols fail to work in the disaggregated memory architecture due to asymmetric resources on CPU nodes and memory nodes. This article proposes ECHO, a fast one-sided RDMA-based RSM protocol with lightweight log replication and remote applying, efficient linearizability guarantee, and fast coordinator failure recovery. ECHO enables all operations in the protocol to be efficiently executed using only one-sided RDMA, without the participation of any computing resource in the memory pool. To provide lightweight log replication and remote applying, ECHO couples the replicated log and the state machine to avoid dual-copy and performs remote applying by updating pointers. To enable efficient remote log state management, ECHO leverages a hitchhiked log state updating scheme to eliminate extra network round trips. To provide efficient linearizability guarantee, ECHO performs immediate remote applying after log replication and leverages the local locks at the coordinator to ensure linear consistency. Moreover, ECHO adopts a commit-aware log cache to make data visible immediately after being committed. To achieve fast failure recovery, ECHO leverages a commit point identification scheme to reduce the overhead of log consistency recovery. Experimental results demonstrate that ECHO outperforms the state-of-the-art RSM protocol (namely Sift) in multiple scenarios. For example, ECHO achieves 27%–52% higher throughput on typical write-intensive workloads. Moreover, ECHO reduces the consistency recovery time by three orders of magnitude for coordinator failure. Jingwen Du, Fang Wang 0001, Dan Feng 0001, Changchen Gan, Yuchao Cao, Xiaomin Zou |
ACM Trans. Archit. Code Optim. | 6 |
| 2022 | ROWE-tree: A Read-Optimized and Write-Efficient B+-tree for Persistent MemoryabstractPersistent memory (PM) exhibits a huge potential to provide B+-tree indexes with high performance, efficient persistence, and instant recovery. A large number of PM-optimized B+-tree indexes have been proposed, but most of them fail to provide high performance for both read and write operations because: (1) their designs of search optimization and insert improvement are often traded off against each other, and (2) they overlook the read/write interference problem of PM which incurs unpredictable performance degradation. In this paper, we propose ROWE-tree, a read-optimized and write-efficient B+-tree for PM. The designs of our ROWE-tree consist of three key points. First, we propose two techniques to make a good trade-off between write and read performance: self-verifying insertion, which reduces consistency overhead by using the key itself as a persist mark instead of additional metadata, and semi-sorted leaf nodes, which use append-only insertion to avoid the shifting overhead of sorting nodes but keep intra-cache-line items sorted to accelerate the lookup. Second, based on the observation that data accesses are highly skewed in real-world workloads, we build an in-DRAM cache of hot items to outsource accesses to hot items to DRAM. By doing so, we can alleviate the read/write interference of PM and significantly improve overall performance. Third, to cope with the dynamic changes of hot items, we exploit a lightweight mechanism to track such changes at run-time. Using Intel Optane DCPMM, our evaluations show that ROWE-tree obtains up to 3.86 × higher performance than the state-of-the-art PM B+-tree indexes under YCSB workloads. Xiaomin Zou, Fang Wang 0001, Dan Feng 0001, Tianjin Guan |
ICPP | 1 |
| 2022 | A write-optimal and concurrent persistent dynamic hashing with radix tree assistance
Xiaomin Zou, Fang Wang 0001, Dan Feng 0001, Renzhi Xiao |
J. Syst. Archit. | 1 |
| 2022 | SPHT: A scalable and high-performance hashing scheme for persistent memoryabstractAbstract The evolution of persistent memory (PM) has significantly affected the design of today's indexing structures. Hashing‐based structures are widely used in storage systems to achieve fast query responses. Recently, several concurrent and failure‐atomic hashing schemes for PM have been proposed to improve the scalability. However, these works still suffer from limited scalability, especially under write‐intensive workloads or at a high number of threads. Our empirical study concludes three issues harm the scalability of PM hashing schemes: the NUMA effects, the resizing operations, and the inter‐thread interference overhead in PM. Based on the above scalability issues, we present SPHT, a scalable and persistent hashing scheme for hybrid DRAM‐PM memory. To eliminate the NUMA effects, SPHT maintains a hash subtable for every NUMA node and stores the key‐value items in the designated nodes. By doing so, all operations can be executed in the local memory without cross‐node communication. The hash subtable consists of two components: the search layer in DRAM for fast accesses and the data layer in PM for efficient persistence. The data layer is organized in a log‐structured way and its log chunks own separate PM space, thus avoiding the resizing operations in PM. Meanwhile, the compacted log structure also supports batching multiple small items, which effectively reduces the persistence overhead. Furthermore, to maximize concurrency, we assign threads to different partitions in the data layer to reduce the inter‐thread interference overhead. On Intel Optane DCPMM, our evaluations show that SPHT scales well and achieves up to 2.7 higher performance than state‐of‐the‐art PM hashing schemes under YCSB workloads. Xiaomin Zou, Fang Wang 0001, Dan Feng 0001, Feiyu Yang, Mengya Lei, Chaojie Liu |
Softw. Pract. Exp. | 1 |
| 2022 | SecNVM: An Efficient and Write-Friendly Metadata Crash Consistency Scheme for Secure NVMabstractData security is an indispensable part of non-volatile memory (NVM) systems. However, implementing data security efficiently on NVM is challenging, since we have to guarantee the consistency of user data and the related security metadata. Existing consistency schemes ignore the recoverability of the SGX style integrity tree (SIT) and the access correlation between metadata blocks, thereby generating unnecessary NVM write traffic. In this article, we propose SecNVM, an efficient and write-friendly metadata crash consistency scheme for secure NVM. SecNVM utilizes the observation that for a lazily updated SIT, the lost tree nodes after a crash can be recovered by the corresponding child nodes in NVM. It reduces the SIT persistency overhead through a restrained write-back metadata cache and exploits the SIT inter-layer dependency for recovery. Next, leveraging the strong access correlation between the counter and DMAC, SecNVM improves the efficiency of security metadata access through a novel collaborative counter-DMAC scheme. In addition, it adopts a lightweight address tracker to reduce the cost of address tracking for fast recovery. Experiments show that compared to the state-of-the-art schemes, SecNVM improves the performance and decreases write traffic a lot, and achieves an acceptable recovery time. Mengya Lei, Fang Wang 0001, Dan Feng 0001, Xiaomin Zou, Renzhi Xiao |
ACM Trans. Archit. Code Optim. | 5 |
| 2021 | HDNH: a read-efficient and write-optimized hashing scheme for hybrid DRAM-NVM memoryabstractWith high memory density, non-volatility and DRAM-scale latency, non-volatile memory (NVM) brings evolution to storage systems and durable data structures. And Intel Optane DC persistent memory module (AEP), the first commercial product of NVM, shows some features that are different from previous assumptions: higher read latency, lower bandwidth and block access granularity compared with DRAM. It is reasonable to build up hybrid memory to give full play to the complementary advantages of DRAM and NVM. In this paper, we present a read-efficient and write-optimized hashing scheme for hybrid DRAM-NVM memory, named HDNH (Hybrid DRAM-NVM Hashing). Our design can be summarized into three key points. First, we decouple the storage for data and metadata by placing key-value items in non-volatile table for persistence while placing metadata in Optimistic Compression Filter (OCF) to reduce excessive NVM accesses. Second, we design hot table in DRAM to speed up search requests and propose an efficient replacement strategy called RAFL. Third, we develop a fine-grained optimistic concurrency mechanism to enable high-performance concurrent accesses on multi-core systems. Experimental results on the AEP platform show that HDNH outperforms its counterparts by up to 2.9x under various YCSB workloads. Kaixin Huang, Xiaomin Zou, Nuo Xu 0001, Liang Fang 0008 |
ICPP | 3 |