Chenlei Tang

dblp:241/1119 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2022 Accelerating range queries of primary and secondary indices for key-value separation
abstract
Primary and secondary indices in LSM-tree-based key-value (KV) stores play significant roles for real-world applications, but they suffer severe I/O amplification due to compaction operations. Prior works show that KV separation can mitigate the I/O amplification under various workloads for either primary or secondary indices. However, range queries of primary and secondary indices only achieve suboptimal efficiency for two reasons: (1) KV separation improves insert/update performance by sacrificing the performance of range queries, (2) range queries of primary and secondary indices may conflict with each other.
Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li
SoCC1
2022 RepKV: A Replicated Key-Value Store to Boost Multiple Indices for Key-Value Separation
abstract
Primary and secondary indices are demanded in real-world applications. Recent works show that key-value(KV) separation is efficient to improve queries of multiple secondary indices in LSM-based KV stores. It stores the value in a separate value log and only stores keys and the value address in primary and secondary indices. However, this share-value scheme on multiple indices results in suboptimal efficiency: (1) queries on secondary indices and range queries on all indices cannot fully exploit the bandwidth of SSD devices simultaneously (2) the put operation is inefficient to update secondary indices.To address the above inefficiency, we propose RepKV, a replicated KV store aiming to boost operations of multiple in-dices. Firstly, RepKV uses a primary-backup replication scheme. Each replication stores the same KV pairs but adopts different organizations to exploit the benefit of SSD devices. Secondly, RepKV proposes a lightweight replication scheme to mitigate the extra KV pairs synchronized in replications. Thirdly, RepKV uses a parallel parsing policy to boost the put operation. Experimental results show that RepKV can improve the query performance on secondary indices by up to 22.24%, the range query performance of all indices by up to 31.38%, and the put performance by up to 13.8%. Besides, RepKV can reduce the I/O amplification of replication by up to 3.05x via the lightweight replication scheme.
Chenlei Tang, Jiguang Wan 0001, Zhihu Tan, Guokuan Li
ICCD1
2022 FenceKV: Enabling Efficient Range Query for Key-Value Separation
abstract
LSM-tree is widely used in key-value stores for big data storage, but it suffers from write amplification brought by frequent compaction operations. An effective solution for this problem is key-value separation, which decouples values from the LSM-tree and stores them in a separate value log. However, existing key-value separation schemes achieve poor range query performance, especially for small key-value pairs, because they focus on mitigating write amplification but neglect access characteristics of the SSD. In this article, we propose FenceKV, which aims to achieve better range query performance while maintaining reasonable update performance for update-intensive workloads. FenceKV employs a new partition method to map values to the storage space based on the key-range to achieve efficient update and range query. Moreover, it adopts a key-range garbage collection policy to mitigate the garbage collection overhead and maintain sequential access for range queries. We compare FenceKV with modern key-value stores with various workloads, and results show that FenceKV can improve the range query performance significantly, while maintaining reasonable update performance compared to the existing designs of key-value separation.
Chenlei Tang, Jiguang Wan 0001, Changsheng Xie 0001
IEEE Trans. Parallel Distributed Syst.1
2019 RAFS: A RAID-Aware File System to Reduce the Parity Update Overhead for SSD RAID
abstract
In a parity-based SSD RAID, small write requests not only accelerate the wear-out of SSDs due to extra writes for updating parities but also deteriorate performance due to associated expensive garbage collection. To mitigate the problem of small writes, a buffer is often added at the RAID controller to absorb overwrites and writes performed to the same stripe. However, this approach achieves only suboptimal efficiency because file layout information is invisible at the block level.This paper proposes RAFS, a RAID-aware file system, which utilizes a RAID-friendly data layout to improve the reliability and performance of SSD-based RAID 5. By leveraging delayed allocation of modern file systems, RAFS employs a stripe-aware buffer policy to coalesce writes to the same file. To reduce parity updates, RAFS compacts buffered updates and flushes back in stripe units to mitigate the parity update overhead. RAFS adopts a stripe-granularity allocation scheme to align writes to stripe boundaries. Experimental results show that RAFS can improve throughput by up to 90%, compared to Ext4.
Chenlei Tang, Jiguang Wan 0001, Fei Wu 0005, Changsheng Xie 0001
DATE1
2019 An Active Method to Mitigate the Long Latencies for Host-Aware Shingle Magnetic Recording Drives
abstract
Shingled Magnetic Recording (SMR) is one of the most promising techniques that satisfy the ever-growing storage volume demands. By overlapping tracks, SMR enormously improves the storage area density, which in turn brings higher storage volumes. However, SMR sacrifices the random write performance for better storage volumes. Current SMR drives propose to remedy this problem by employing an in-drive persistent cache to temporally store incoming writes and migrate them to their disk destinations later on. Unfortunately, cleaning processes for the persistent cache takes up to tens of seconds, and the unpredictable timing of these time-consuming operations chokes normal requests and drastically degrades SMR drive performance. In this paper, we propose to remedy this issue by proactively freeing the persistent cache space so that keeping these lengthy processes transparent with regards to normal requests, therefore reducing the long tails and delivering steady and predictable performance for SMR drive-based storage systems. We prototype our design as a Host-Aware SMR drive aware userspace file system, AM FS, and evaluate it on the real HA-SMR drive with libzbc. Evaluations results show that AM FS reduces the long tails of HA-SMR drives.
Jiguang Wan 0001, Ping Huang 0001, Bihua Shu, Chenlei Tang, Changsheng Xie 0001
ICPADS5