VLDB 2026 Research / reviewers in the wild / expert
Seiyeon Cho
dblp:372/4579
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Indexing and storage engines · 50% Machine learning and data management · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Reconfigurable computing and FPGAs · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning and data management › continual learning
incremental learning |
0.8 | 1 | 2024 | Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024 |
Indexing and storage engines
learned index |
0.8 | 1 | 2024 | Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.8 | 1 | 2024 | Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024 |
Methods — techniques the papers use, named apart from their topics
memoization · 1.5matrix decomposition · 1.5QR factorization · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Accelerating String-key Learned Index Structures via Memoization-based Incremental TrainingabstractLearned indexes use machine learning models to learn the mappings between keys and their corresponding positions in key-value indexes. These indexes use the mapping information as training data. Learned indexes require frequent retrainings of their models to incorporate the changes introduced by update queries. To efficiently retrain the models, existing learned index systems often harness a linear algebraic QR factorization technique that performs matrix decomposition. This factorization approach processes all key-position pairs during each retraining, resulting in compute operations that grow linearly with the total number of keys and their lengths. Consequently, the retrainings create a severe performance bottleneck, especially for variable-length string keys, while the retrainings are crucial for maintaining high prediction accuracy and in turn, ensuring low query service latency. To address this performance problem, we develop an algorithm-hardware co-designed string-key learned index system, dubbed SIA. In designing SIA, we leverage a unique algorithmic property of the matrix decomposition-based training method. Exploiting the property, we develop a memoization-based incremental training scheme, which only requires computation over updated keys, while decomposition results of non-updated keys from previous computations can be reused. We further enhance SIA to offload a portion of this training process to an FPGA accelerator to not only relieve CPU resources for serving index queries (i.e., inference), but also accelerate the training itself. Our evaluation shows that compared to ALEX, LIPP, and SIndex, a state-of-the-art learned index systems, SIA-accelerated learned indexes offer 2.6× and 3.4× higher throughput on the two real-world benchmark suites, YCSB and Twitter cache trace, respectively. Minsu Kim 0004, Jinwoo Hwang, Guseul Heo, Seiyeon Cho, Divya Mahajan 0001, Jongse Park |
Proc. VLDB Endow. | 4 |