Xiangxiang Jiang

dblp:203/2033 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0000-6747-214XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GDH+: A GPGPU-Empowered Gradient Data Hierarchy and Key-Value Separation for Optimizing LSM-Tree-Based KV Stores
abstract
The rapid growth of unstructured data has driven the widespread adoption of LSM-tree-based key-value stores (KV stores). The write amplification resulting from compaction in LSM-trees causes a performance bottleneck. Existing solutions attempt to address this issue through key-value separation strategies. However, these studies fail to optimize the memory components of LSM-trees or provide efficient garbage collection (GC) strategies that achieve high performance while minimizing CPU overhead. These limitations motivate us to propose a GPGPU-empowered gradient data hierarchy and key-value separation for optimizing KV stores, named GDH+ . We utilize GPGPU acceleration for sorting and flushing operations, optimizing the memory components of the LSM-tree. Additionally, we enhance read performance with an LRU-based memory component that distinguishes between hot and cold data, in combination with an adaptive migration strategy. Furthermore, we propose an in-place GC strategy to reduce CPU overhead while maintaining high performance. GDH+ achieves a 2×, 1.5×, and 40% improvement in write performance, read performance, and CPU utilization, respectively, compared to state-of-the-art KV stores.
Hui Sun 0002, Xiangxiang Jiang, Jinfeng Xu 0004, Enhui Wang, Yinliang Yue, Xiao Qin 0001
ACM Trans. Archit. Code Optim.2
2025 gParaKV: A GPGPU-accelerated Key-Value Separation-based KV Store with Optimized Compaction and Garbage Collection
abstract
LSM-tree-based key-value stores or KV stores are widely deployed in modern cloud storage systems thanks to high data storage efficiency and retrieval capabilities. The compaction process in the LSM-tree, however, results in severe performance bottlenecks, especially in scenarios involving large volumes of data. While key-value separation methods mitigate the performance bottlenecks caused by compaction, the existing methods do not fully address merge-sorting during compaction and expensive garbage collection (GC). We propose gParaKV, a GPGPU-empowered KV store with a KV separation mechanism, leveraging the GPGPU parallel technology to accelerate merge-sorting in compaction and GC. gParaKV embraces unique features like a GPGPU bitmap structure, parallel data marking, and a parallel GC mechanism. These critical components effectively curtail the overhead of merge-sorting and GC operations by virtue of parallel computing. We compare it with state-of-the-art KV stores (e.g., RocksDB, BlobDB, Wisckey, DiffKV, UniKV, and HPDK) under various workloads. The experimental results show that gParaKV can improve the write performance and GC efficiency compared to the existing key-value separation-based KV stores.
Hui Sun 0002, Xiangxiang Jiang, Xiao Qin 0001, Song Jiang 0001, Enhui Wang
SC2
2025 RGKV: A GPGPU-Empowered Compaction Framework for LSM-Tree-Based KV Stores With Optimized Data Transfer and Parallel Processing
abstract
The Log-structured merge-tree (LSM-tree), widely adopted in key-value stores (KV stores), is esteemed for its efficient write performance and superb scalability amid large-scale data processing. The compaction process of LSM-trees consumes significant computational resources, thereby becoming a bottleneck for system performance. Traditionally, compaction is handled by CPUs, but CPU processing capacity often falls short of increasing demands with the surge in data volumes. To address this challenge, existing solutions attempt to accelerate compaction using GPGPUs. Due to low GPGPU parallelism and data transfer delay in prior studies, the anticipated performance improvements have not yet been fully realized. In this paper, we bring forth RGKV – a comprehensive optimization approach to overcoming the limitations of current GPGPU-empowered KV stores. RGKV features the GPGPU-adapted contiguous memory allocation and GPGPU-optimized key-value block architecture to furnish high-efficient GPGPU parallel encoding and decoding catering to the needs of KV stores. To enhance the computational efficiency and overall performance of KV stores, RGKV employs a parallel merge-sorting algorithm to maximize the parallel processing capabilities of the GPGPU. Moreover, RGKV incorporates a data transfer module anchored on the GPUDirect storage technology – designed for KV stores – and designs an efficient data structure to substantially curtail data transfer latency between an SSD and a GPGPU, boosting data transfer speed and alleviating CPU load. The experimental results demonstrate that RGKV achieves a remarkable 4$\times$improvement in overall throughput and a 7$\times$improvement in compaction throughput compared to the state-of-the-art KV stores, while also reducing average write latency by 70.6%.
Hui Sun 0002, Xiangxiang Jiang, Yinliang Yue, Xiao Qin 0001
IEEE Trans. Computers2
2024 gLSM: Using GPGPU to Accelerate Compactions in LSM-tree-based Key-value Stores
abstract
Log-structured-merge tree or LSM-tree is a technological underpinning in key-value (KV) stores to support a wide range of performance-critical applications. By conducting data re-organization in the background by virtue of compaction operations, the KV stores have the potential to swiftly service write requests with sequential batched disk writes and read requests for KV items constantly sorted by the compaction. Compaction demands high I/O bandwidth and CPU speed to facilitate quality service to user read/write requests. With the emergence of high-speed SSDs, CPUs are increasingly becoming a performance bottleneck. To mitigate the bottleneck limiting the KV-store’s performance and that of the applications supported by the store, we propose a system - gLSM - to leverage GPGPU to remarkably accelerate the compaction operations. gLSM fully utilizes the parallelism and computational capability inside GPGPUs to improve the compaction performance. We design a driver framework to parallelize compaction operations handled between a pair of CPU and GPGPU. We employ data independence and GPGPU-orient radix-sorting algorithm to concurrently conduct compaction. A key-value separation method is devised to slash the transfer of data volume from CPU-side memory to the GPGPU counterpart. The results reveal that gLSM improves the throughput and compaction bandwidth by up to a factor of 2.9 and 26.0, respectively, compared with the four state-of-the-art KV stores. gLSM also reduces the write latency by 73.3%. gLSM exhibits a performance improvement by up to 45% compared against its variant where there are no KV separation and collaboration sort modules.
Hui Sun 0002, Jinfeng Xu 0004, Xiangxiang Jiang, Yinliang Yue, Xiao Qin 0001
ACM Trans. Storage3
2022 A smart file-level continuous data protection scheme based on security baseline
abstract
There is a rapidly growing interest in securing big data due to the rapid development of cloud computing, big data, and other information technologies according to the fourth industrial revolution. Continuous data protection (CDP) is an effective method to deal with huge loss caused by data loss. More optimal design methods are available and studies on the establishment of the knowledge base for an efficient backup data management in the CDP field. In this paper, a knowledge-based smart file-level CDP scheme is suggested. The user's foundation database and context information are applied to machine learning technology, enabling a large amount of files' log data and context information accumulated continuously to be stored in the knowledge base using B + tree structure. This enables high performance and flexibility in the data protection management system. The result of comparative evaluation with different security risk levels for verifying the validity shows that the suggested method presented a higher performance in write/query operations and storage overhead.
Xiao Yu 0005, Xiangxiang Jiang, Zhizhuo Sun, Liu Cong
Int. J. Intell. Syst.2