Rain Jiang

dblp:300/4242 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 83% Distributed and cloud data management · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › search engines
full-text search
1.722025
LogCloud: Fast Search of Compressed Logs on Object Storage · Proc. VLDB Endow. 2025
Rottnest: Indexing Data Lakes for Search · ICDE 2025
Distributed and cloud data management
data lake
0.912025
Rottnest: Indexing Data Lakes for Search · ICDE 2025
Information retrieval
indexing
0.912025
Rottnest: Indexing Data Lakes for Search · ICDE 2025
Information retrieval › similarity search
nearest neighbor search
0.912025
Rottnest: Indexing Data Lakes for Search · ICDE 2025
Information retrieval
search engines
0.912025
LogCloud: Fast Search of Compressed Logs on Object Storage · Proc. VLDB Endow. 2025
Storage systems
log management
0.912025
LogCloud: Fast Search of Compressed Logs on Object Storage · Proc. VLDB Endow. 2025
Storage systems
object storage
0.312025
LogCloud: Fast Search of Compressed Logs on Object Storage · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

inverted index · 1.7FM-index · 1.7
YearPublicationVenuePosition
2025 Rottnest: Indexing Data Lakes for Search
abstract
Data lakes have become widely popular in managing enterprise data. Their widespread integration with query engines has allowed them to displace specialized data warehouses as the single source of truth for enterprise data. While the columnar storage format and block min-max indices allow query engines to achieve competitive performance on relational data analytics queries, they are not yet suitable for other search-oriented queries like full text and vector nearest neighbor search. We present Rottnest, a general system that builds additional lightweight indices on top of data lakes. We show that our system is more cost efficient compared to un-indexed data lakes or specialized databases across several orders of magnitude of total query loads and operating time horizons.
Sasha Krassovsky, Conor Kennedy, Alex Aiken, Weston Pace, Rain Jiang, Huayi Zhang
ICDE6
2025 LogCloud: Fast Search of Compressed Logs on Object Storage
abstract
Large organizations emit terabytes of logs every day in their cloud environment. Efficient data science on these logs via text search is crucial for gleaning operational insights and debugging production outages. Current log management systems either perform full-text indexing on a cluster of dedicated servers to provide efficient search at the expense of high storage cost, or store unindexed compressed logs on object storage at the expense of high search cost. We propose LogCloud, a new object-storage based log management system that supports both cheap compressed log storage and efficient search. LogCloud constructs inverted indices on compressed logs using a novel FM-index implementation that supports efficient querying from object storage directly, removing the need for dedicated indexing servers. Experiments on five public and five production log datasets show that LogCloud can achieve both cheap storage and search, scaling to TB-scale datasets.
Junyu Wei, Alex Aiken, Guangyan Zhang, Jacob Odgård Tørring, Rain Jiang
Proc. VLDB Endow.6