Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rongfeng He

dblp:320/7021 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-0798-966XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Database system architecture and tuning · 60% Query processing and optimization · 23% Indexing and storage engines · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Interconnection networks and networks-on-chip · 67% Memory systems · 33%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning › database design › physical database design
index selection
1.012026
BISLearner: Block-Aware Index Selection using Attention-Based Reinforcement Learning for Data Analytics · ACM Trans. Database Syst. 2026
Database system architecture and tuning
index tuning
1.012026
BISLearner: Block-Aware Index Selection using Attention-Based Reinforcement Learning for Data Analytics · ACM Trans. Database Syst. 2026
Database system architecture and tuning › index tuning
learned index tuning
1.012026
BISLearner: Block-Aware Index Selection using Attention-Based Reinforcement Learning for Data Analytics · ACM Trans. Database Syst. 2026
Query processing and optimization › runtime optimization
data skipping
0.712023
Sieve: A Learned Data-Skipping Index for Data Analytics · Proc. VLDB Endow. 2023
Indexing and storage engines
learned index
0.712023
Sieve: A Learned Data-Skipping Index for Data Analytics · Proc. VLDB Endow. 2023
Memory systems
memory disaggregation
0.612022
RACE: One-sided RDMA-conscious Extendible Hashing · ACM Trans. Storage 2022
Interconnection networks and networks-on-chip › remote direct memory access
one-sided RDMA
0.612022
RACE: One-sided RDMA-conscious Extendible Hashing · ACM Trans. Storage 2022
Interconnection networks and networks-on-chip
remote direct memory access
0.612022
RACE: One-sided RDMA-conscious Extendible Hashing · ACM Trans. Storage 2022
Information retrieval › hashing › hash table design
extendible hashing
0.212022
RACE: One-sided RDMA-conscious Extendible Hashing · ACM Trans. Storage 2022

Methods — techniques the papers use, named apart from their topics

lock-free concurrency control · 1.1extendible remote resizing · 1.1invalid action masking · 1.0histogram-based block summaries · 1.0attention-based reinforcement learning · 1.0piecewise linear function · 0.7learned index · 0.7
YearPublicationVenuePosition
2026 BISLearner: Block-Aware Index Selection using Attention-Based Reinforcement Learning for Data Analytics
abstract
The development of data analytics services has fueled many optimizations in data scans, and indexes are one of the most important techniques to improve scan efficiency. Meanwhile, block-based data organization has become standard practice in these services, providing an opportunity for more fine-grained index selection at the block level. However, today’s systems ignore data distribution differences among blocks and usually tune indexes over the entire database table, leading to unnecessary storage costs and potential degradation in query performance. To bridge this gap, we propose BISLearner, a fast, block-aware index selecting approach based on reinforcement learning. One major challenge lies in differentiating the data distribution among data blocks. To solve this problem, BISLearner maintains simplified histograms that represent the data distribution of each block. When a query is issued, BISLearner leverages the query predicate and histogram-based block summaries to generate a specific workload representation for each block. However, such block-aware workload representation leads to an excessive number of input features, resulting in a slow or even incorrect convergence of neural networks. Inspired by the human-learning process, where more attention is devoted to the important parts of data, we design an attention-based neural model to efficiently handle the high volume of input features caused by table partitioning and select the best-suited index combinations at the block level. Additionally, to handle the expensive search space caused by attribute combinations and data partitioning, we employ heuristic-based invalid action masking at the block level to accelerate the training process. Our evaluation using PostgreSQL and Greenplum database systems demonstrates BISLearner is able to reduce job completion time by up to 28.45% compared to its best counterparts.
Yulai Tong, Hua Wang 0008, Ke Zhou 0001, JiaLe Miao, Rongfeng He
ACM Trans. Database Syst.8
2024 DAG-aware harmonizing job scheduling and data caching for disaggregated analytics frameworks
Yulai Tong, Hua Wang 0008, Ke Zhou 0001, Rongfeng He
Future Gener. Comput. Syst.6
2023 Sieve: A Learned Data-Skipping Index for Data Analytics
abstract
Modern data analytics services are coupled with external data storage services, making I/O from remote cloud storage one of the dominant costs for query processing. Techniques such as columnar block-based data organization and compression have become standard practices for these services to save storage and processing cost. However, the problem of effectively skipping irrelevant blocks at low overhead is still open. Existing data-skipping efforts maintain lightweight summaries (e.g., min/max, histograms) for each block to filter irrelevant data. However, such techniques ignore patterns in real-world data, enabling ineffective use of the storage budget and may cause serious false positives. This paper presents Sieve, a learning-enhanced index designed to efficiently filter out irrelevant blocks by capturing data patterns. Specifically, Sieve utilizes piece-wise linear functions to capture block distribution trends over the key space. Based on the captured trends, Sieve trades off storage consumption and false positives by grouping neighboring keys with similar block distributions into a single region. We have evaluated Sieve using Presto, and experiments on real-world datasets demonstrate that Sieve achieves up to 80% reduction in blocks accessed and 42% reduction in query times compared to its counterparts.
Yulai Tong, Hua Wang 0008, Ke Zhou 0001, Rongfeng He
Proc. VLDB Endow.5
2022 DSDP: Dual Stream Data Prefetcher
abstract
Hardware prefetching is an important DRAM latency hiding technology. Designing prefetchers to maximize system performance often requires a delicate balance between coverage and accuracy. As the number of cores increases, the accuracy of the prefetching algorithm becomes more important. Separating streams based on memory access instructions is an effective way to improve accuracy. However, this technique may lose prefetch opportunities by losing cross-PC relationships, and even reduce algorithm coverage.
Hua Wang 0008, Ke Zhou 0001, Kaichao Cui, Huabing Yan, Rongfeng He
PACT7
2022 RACE: One-sided RDMA-conscious Extendible Hashing
abstract
Memory disaggregation is a promising technique in datacenters with the benefit of improving resource utilization, failure isolation, and elasticity. Hashing indexes have been widely used to provide fast lookup services in distributed memory systems. However, traditional hashing indexes become inefficient for disaggregated memory, since the computing power in the memory pool is too weak to execute complex index requests. To provide efficient indexing services in disaggregated memory scenarios, this article proposes RACE hashing, a one-sided RDMA-Conscious Extendible hashing index with lock-free remote concurrency control and efficient remote resizing. RACE hashing enables all index operations to be efficiently executed by using only one-sided RDMA verbs without involving any compute resource in the memory pool. To support remote concurrent access with high performance, RACE hashing leverages a lock-free remote concurrency control scheme to enable different clients to concurrently operate the same hashing index in the memory pool in a lock-free manner. To resize the hash table with low overheads, RACE hashing leverages an extendible remote resizing scheme to reduce extra RDMA accesses caused by extendible resizing and allow concurrent request execution during resizing. Extensive experimental results demonstrate that RACE hashing outperforms state-of-the-art distributed in-memory hashing indexes by 1.4–13.7× in YCSB hybrid workloads.
Pengfei Zuo, Qihui Zhou, Jiazhao Sun, Shuangwu Zhang, Yu Hua 0001, James Cheng, Rongfeng He, Huabing Yan
ACM Trans. Storage8