EDBT 2026 Demo / reviewers in the wild / expert
Yuanhui Zhou
dblp:05/3711
· DBLP profile ↗
9ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0003-3685-7033ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeapKV: Enabling Efficient Garbage Collection for KV-Separated LSM Stores on Modern SSDsabstractKey-value (KV) separation has emerged as a pivotal solution to tackle write amplification in LSM-tree-based KV stores (LSM stores). However, the garbage collection (GC) mechanism essential for reclaiming obsolete values introduces substantial overheads: (i) Additional LSM-tree I/O operations degrade foreground performance and cause data inconsistency issues; (ii) Exacerbated write amplification arises from excessive valid data migration in update-intensive workloads. Moreover, existing GC optimization schemes fundamentally struggle to balance space overhead, write amplification, and system performance. In this article, we propose HeapKV, a high-performance KV-separated LSM store that improves GC efficiency through three key technologies: (i) A lightweight two-level index and a global garbage view decouple GC operations of value storage from the LSM-tree, eliminating additional I/O operations; (ii) A novel valid data migration scheme mitigates write amplification during space reclamation by in-place overwrites and logical data copying; (iii) SSD-conscious I/O optimizations featuring asynchronous value flushing, fast read paths and concurrent prefetching for range queries. Extensive experiments demonstrate that HeapKV achieves 40%–7.4× higher throughput under diverse workloads with lower write/space amplification, compared to other state-of-the-art KV-separated LSM stores. Kai Lu 0002, Yuanhui Zhou, Nengjie Wang, Jiguang Wan 0001, Bisheng Huang |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | A Contract-aware and Cost-effective LSM Store for Cloud Storage with Low Latency SpikesabstractCloud storage is gaining popularity because features such as pay-as-you-go significantly reduce storage costs. However, the community has not sufficiently explored its contract model and latency characteristics. As LSM-Tree-based key-value stores (LSM stores) become the building block for numerous cloud applications, how cloud storage would impact the performance of key-value accesses is vital. This study reveals the significant latency variances of Amazon Elastic Block Store (EBS) under various I/O pressures, which challenges LSM store read performance on cloud storage. To reduce the corresponding tail latency, we propose Calcspar, a contract-aware LSM store for cloud storage, which efficiently addresses the challenges by regulating the rate of I/O requests to cloud storage and absorbing surplus I/O requests with the data cache. We specifically developed a fluctuation-aware cache to lower the high latency brought on by workload fluctuations. Additionally, we build a congestion-aware IOPS allocator to reduce the impact of LSM store internal operations on read latency. We evaluated Calcspar on EBS with different real-world workloads and compared it to the cutting-edge LSM stores. The results show that Calcspar can significantly reduce tail latency while maintaining regular read and write performance, keeping the 99 th percentile latency under 550μs and reducing average latency by 66%. In addition, Calcspar has lower write prices and average latency compared to Cloud NoSQL services offered by cloud vendors. Yuanhui Zhou, Jian Zhou 0004, Kai Lu 0002, Shuning Chen, Jiguang Wan 0001 |
ACM Trans. Storage | 1 |
| 2023 | Calcspar: A Contract-Aware LSM Store for Cloud Storage with Low Latency Spikes
Yuanhui Zhou, Jian Zhou 0004, Shuning Chen, Yanguang Wang, Jiguang Wan 0001 |
USENIX ATC | 1 |
| 2022 | Building a Fast and Efficient LSM-tree Store by Integrating Local Storage with Cloud StorageabstractThe explosive growth of modern web-scale applications has made cost-effectiveness a primary design goal for their underlying databases. As a backbone of modern databases, LSM-tree based key–value stores (LSM store) face limited storage options. They are either designed for local storage that is relatively small, expensive, and fast or for cloud storage that offers larger capacities at reduced costs but slower. Designing an LSM store by integrating local storage with cloud storage services is a promising way to balance the cost and performance. However, such design faces challenges such as data reorganization, metadata overhead, and reliability issues. In this article, we propose RocksMash , a fast and efficient LSM store that uses local storage to store frequently accessed data and metadata while using cloud to hold the rest of the data to achieve cost-effectiveness. To improve metadata space-efficiency and read performance, RocksMash uses an LSM-aware persistent cache that stores metadata in a space-efficient way and stores popular data blocks by using compaction-aware layouts. Moreover, RocksMash uses an extended write-ahead log for fast parallel data recovery. We implemented RocksMash by embedding these designs into RocksDB. The evaluation results show that RocksMash improves the performance by up to 1.7 \( \times \) compared to the state-of-the-art schemes and delivers high reliability, cost-effectiveness, and fast recovery. Jiguang Wan 0001, Shuning Chen, Yuanhui Zhou, Hadeel Albahar, Zhihu Tan |
ACM Trans. Archit. Code Optim. | 6 |
| 2021 | Building A Fast and Efficient LSM-tree Store by Integrating Local Storage with Cloud StorageabstractThe explosive growth of modern web-scale applications has made cost-effectiveness a primary design goal for their underlying databases. As a backbone of modern databases, LSM-tree based key-value stores (LSM store) face limited storage options. They are either designed for local storage that is relatively small, expensive, and fast or for cloud storage that offers larger capacities at reduced costs but slower. Designing an LSM store by integrating local storage with cloud storage services is a promising way to balance the cost and performance. However, such design faces challenges such as data reorganization and metadata overhead issues. In this paper, we propose ROCKSMASH, a fast and efficient LSM store that uses local storage to store frequently accessed data and metadata while using cloud to hold the rest of the data to achieve cost-effectiveness. To improve metadata space-efficiency and read performance, ROCKSMASH uses an LSM-aware persistent cache that stores metadata in a space-efficient way and stores popular data blocks by using compaction-aware layouts. We implemented ROCKSMASH by embedding these designs into RocksDB. The evaluation results show that ROCKSMASH improves the performance by up to 1.7 × compared to the state-of-the-art schemes and delivers higher reliability and cost-effectiveness. Jiguang Wan 0001, Shuning Chen, Yuanhui Zhou, Hadeel Albahar, Changsheng Xie 0001 |
CLUSTER | 6 |
| 1999 | A Hybrid Lazy-Eager Approach to Reducing the Computation and Memory Requirements of Local Parametric Learning Algorithms
Yuanhui Zhou, Carla E. Brodley |
ICML | 1 |
| 1998 | Mining Classification Rules in Multistrategy Learning ApproachabstractClassification, which involves finding rules that partition a given dataset into disjoint groups, is one class of data mining problems. Approaches proposed so far for mining classification rules from databases are mainly decision tree based on symbolic learning methods. In this paper, we combine artificial neural network and genetic algorithm to mine classification rules. Some experiments have demonstrated that our method generates rules of better performance than the decision tree approach and the number of extracted rules is fewer than that of C4.5. Yuanhui Zhou, Yuchang Lu, Chunyi Shi |
Intell. Data Anal. | 1 |
| 1997 | A Connectionist Approach to Extracting Knowledge from Databases
Yuanhui Zhou, Yuchang Lu, Chunyi Shi |
IDA | 1 |
| 1997 | Using Neural Network to Extract Knowledge from Database
Yuanhui Zhou, Yuchang Lu, Chunyi Shi |
PKDD | 1 |