VLDB 2026 Research / reviewers in the wild / expert
Ruicheng Liu
dblp:01/4595
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-7786-9896ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GaussDB-Vector: A Large-Scale Persistent Real-Time Vector Database for LLM ApplicationsabstractVector databases are widely used as a fundamental tool for addressing the weaknesses of large language model (LLM) applications, specifically hallucinations and the high cost of inference. However, existing vector databases either cater to niche applications with low-latency in-memory search, or offer sophisticated data management capabilities but at the cost of low performance. To address these limitations, we propose GaussDB-Vector, a high-performance, real-time persistent vector database that excels in low-latency scalable search, real-time inserts and deletes, high availability, large-scale distributed search, and hybrid scalar-vector filtered search capabilities. These features are primarily achieved through an innovative storage architecture designed for a graph-based vector index, optimized for I/O operations and adaptable across various dataset sizes and dimensions, complemented by novel buffering strategies to further reduce I/O burdens. GaussDB-Vector supports product quantization, parallel search, and hardware acceleration via SIMD, GPUs, and NPUs in order to further accelerate queries. Experimental results show that GaussDB-Vector outperforms competitive baselines by a factor of 1 to 5 times. Guoliang Li 0001, Ji Sun 0001, James Pan, Yongqing Xie, Ruicheng Liu, Wen Nie |
Proc. VLDB Endow. | 6 |
| 2025 | A Topology-Aware Localized Update Strategy for Graph-Based ANN Index
Song Yu 0004, Shengyuan Lin, Shufeng Gong 0001, Yongqing Xie, Ruicheng Liu, Ji Sun 0001, Yanfeng Zhang 0001, Guoliang Li 0001, Ge Yu 0001 |
Proc. VLDB Endow. | 5 |
| 2024 | Online Nonstop Task Management for Storm-Based Distributed Stream Processing Engines
Zhou Zhang 0006, Peiquan Jin, Xike Xie, Ruicheng Liu, Shouhong Wan |
J. Comput. Sci. Technol. | 5 |
| 2023 | Closing the Performance Gap between Leveling and Tiering Compaction via Bundle CompactionabstractSo far, most LSM-tree-based storage engines adopt either leveling or tiering compaction. We note that while leveling compaction can deliver high search performance and low space amplification, it has a high rate of write amplification (therefore delivering poor write performance). On the other hand, tiering compaction has a low rate of write amplification (therefore delivering good write performance) but has poor search performance and high space amplification. Aiming to close the performance gap between leveling and tiering databases, this paper proposes a new storage engine called B+LSM. The novel ideas of B+LSM lie in two aspects: (1) B+LSM replaces the underlying level structure of LSM-tree with a B+-tree-like tree, and each tree node is defined as a Bundle Compaction Unit (BCU), whose size is allowed to be dynamically changed with workload statistics to balance read and write performance. (2) B+LSM proposes a new node-grained compaction scheme called Bundle Compaction. Bundle compaction is always triggered to merge all the data within a BCU node, partition them into bundles, and then send bundles to the children. Such a compaction scheme can take advantage of leveling and tiering compaction by auto-tuning the size of BCU nodes. We implemented B+LSM and compared it with LevelDB, RocksDB, PebblesDB, and L2SM on the YCSB workloads. The results show that B+LSM can achieve high time performance and reduce space amplification on both static and dynamic workloads. Ruicheng Liu, Peiquan Jin, Yongping Luo, Zhaole Chu, Yigui Yuan |
HPDC | 1 |
| 2023 | A Multi-task Learning Model for Gold-two-mention Co-reference ResolutionabstractThe task of resolving repeated objects in natural languages is known as co-reference resolution. It is an important part of modern natural language processing and semantic cognition as these implicit relationships are particularly difficult in natural language understanding in downstream tasks. Mention identification and mention linking are the two sub-tasks in the general co-reference resolution research community. Gold-two-mention style co-reference resolution is a special type of co-reference resolution that focuses on linking the ambiguous pronoun to one of the two candidate antecedents. In this paper, we proposed a joint learning model that learns mention identification and mention linking tasks together, because we find that the learning of mention identification can provide supportive dependent information for the learning of mention linking. As far as we know, we propose the first model that introduces a multi-task learning framework to the gold-two-mention co-reference resolution task. We find that our proposed model outperforms state-of-the-art baselines and a single-task learning model on three gold-two-mention co-reference resolution datasets. By comparing the errors made by either the single-task learning model or the multi-task learning model, our error analysis also yields interesting findings about in which way our multi-task learning model makes fewer resolution errors. Ruicheng Liu, Guanyi Chen, Rui Mao 0010, Erik Cambria |
IJCNN | 1 |
| 2022 | Design Considerations of A Novel Distributed Key-Value Store for New StorageabstractThe emergence of new storage like persistent memory (PM) and zoned namespaces SSDs (ZNS-SSDs) introduces new challenges and opportunities for distributed key-value stores. Since LSM-tree has been widely adopted in distributed key-value stores, such as RocksDB and HBase, it is necessary to revisit the LSM-tree to make it adapt to new storage. In this paper, we first analyze the challenges of adapting the LSM-tree for new storage. Then, we propose a high-level architecture for a new-storage-aware LSM-tree-based key-value store called Hybrid-LSM. We explain the key structural issues of different storage layers in Hybrid-LSM and present some preliminary design ideas. Ruicheng Liu, Peiquan Jin, Yongping Luo, Zhaole Chu |
ICDCS | 1 |
| 2022 | ZonedStore: A Concurrent ZNS-Aware Cache System for Cloud Data StorageabstractCloud data storage relies on efficient cache systems to offer high performance for intensive reads/writes on big data. Due to the large data volume of cloud data storage and the limited capacity of DRAM, current cloud vendors prefer to use SSDs (Solid State Drives) but not DRAM to build the cache system. However, traditional SSDs have a serious over-provisioning problem and a high cost in garbage collection. Thus, the performance of SSD-based cache systems will drop quickly when the usage of SSDs increases. Recently, Zoned Namespaces (ZNS) SSDs have emerged as a hot topic in both academics and industries. Compared to conventional SSDs, ZNS SSDs have the advantages of less overhead of garbage collection and lower over-provisioning costs. Therefore, ZNS SSDs have been a better candidate for the cache system for cloud storage. However, ZNS SSDs only accept sequential writes, and the zones inside ZNS SSDs need to be carefully managed to maximize the advantages of ZNS SSDs. Therefore, making the cache system adapt to ZNS SSDs is becoming a challenging issue. In this paper, we demonstrate ZonedStore, a novel ZNS-aware cache system for cloud data storage. After a brief introduction to the architecture of ZonedStore, we present the key designs of ZonedStore, including a Zone Manager to control the space allocation and operations on ZNS SSDs, a Multi-Layer Buffer Manager, and an In-Memory Concurrent Index to accelerate accesses. Finally, we present a case study to demonstrate the working process and performance of ZonedStore. Yanqi Lv 0001, Peiquan Jin, Ruicheng Liu, Yuanjin Lin, Kuankuan Guo |
ICDCS | 4 |
| 2022 | MMH-index: Enhancing Apache Lucene with High-Performance Multi-Modal Indexing and SearchingabstractData diversity is one of the main characteristics of big data, which makes a growing number of multi-modal data. For example, in e-commerce applications, a product is often described simultaneously through text, image, video, etc. However, traditional text search engines are mainly oriented to text data and cannot offer high-performance indexing and search on multi-modal data. In this paper, we propose a high-performance hybrid index structure named MMH-index to enhance Apache Lucene, an open-source Java library providing powerful indexing and search features, with multi-modal indexing and searching. The MMH-index optimizes the classical inverted index in Lucene with an innovative design called modality bitmap. By adding modality bitmaps to the dictionary of the inverted index, MMH-index can reduce the redundant data in the index and the average length of the posting lists, thus decreasing the space consumption and the time cost of multi-modal queries. We implement MMH-index in Lucene and evaluate its performance on two real multi-modal datasets. The experimental results show that compared with the traditional inverted index in Lucene, MMH-index achieves significant improvement in space consumption and multi-modal query performance. Ruicheng Liu, Jialing Liang, Peiquan Jin |
ACM Multimedia | 1 |
| 2021 | RoBF: An Auto-Tuning Bloom Filter for Mixed Queries on LSM-TreeabstractBloom filter is an efficient technique to improve query performance in LSM-tree-based databases, such as RocksDB, HBase, and Cassandra.However, the original Bloom filter uses a fixed false positive rate (FPR), which makes it inefficient for mixed queries that involve both point and range queries.To solve this problem, in this paper, we present an improved Bloom filter called RoBF (Range-Query-Oriented Bloom Filter), which uses a mixture of Bloom filters and can process mixed queries on LSM-tree efficiently.We design an efficient algorithm for generating the solution based on the query distribution.We compare our proposal with the trie-based filter and find out that each has its own advantages for various scenarios.Therefore, we propose to use different filters with varied sizes for different levels on LSM-tree.Following this idea, we present an algorithm to generate specific filters with a specific size for different levels on LSM-tree to optimize the performance of mixed queries under limited memory space.We conduct comparative experiments and compare the proposed RoBF with various competitors, and the results show that RoBF can improve the performance of evaluating mixed queries by up to 6x to 30x, compared to the original Bloom filter in RocksDB. Ruicheng Liu, Peiquan Jin, Shouhong Wan, Bei Hua |
SEKE | 1 |
| 2021 | Water-Wheel: Real-Time Storage with High Throughput and Scalability for Big Data Streams
Yanqi Lv 0001, Ruicheng Liu, Peiquan Jin |
SEKE | 2 |
| 2005 | Aspect-Oriented Real-Time System Modeling Method Based on UMLabstractReal-time systems could be modeled using AOP based on UML. Timing requirements could be separated from the system according the separation of concerns techniques, expressed as a time-aspect independence of the system, and designed occurrence. So a timing model could be created to describe the time of the system. Finally the timing model could be woven into the system to compose a real-time system based on the AOP technology only when needed for a particular application. Also UML had been extended to express AOP and the time model. The real-time systems could be modeled from the static structure, dynamic behaviors and weaving of the time-aspect, and an elevator case had been given as an example. Ruicheng Liu |
RTCSA | 2 |