VLDB 2026 Research / reviewers in the wild / expert
Robin Nittka
dblp:49/10329
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.9 | 1 | 2025 | Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe · NeurIPS 2025 |
Information retrieval › document retrieval › structure-aware retrieval
hierarchical retrieval |
0.9 | 1 | 2025 | Hierarchical Retrieval: The Geometry and a Pretrain-Finetune Recipe · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
dual encoder · 1.7pretrain-finetune · 0.9pre-train/fine-tune · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeabstractDual encoder (DE) models, where a pair of matching query and document are embedded into similar vector representations, are widely used in information retrieval due to their simplicity and scalability. However, the Euclidean geometry of the embedding space limits the expressive power of DEs, which may compromise their quality. This paper investigates such limitations in the context of hierarchical retrieval (HR),
where the document set has a hierarchical structure and the matching documents for a query are all of its ancestors. We first prove that DEs are feasible for HR as long as the embedding dimension is linear in the depth of the hierarchy and logarithmic in the number of documents.
Then we study the problem of learning such embeddings in a standard retrieval setup where DEs are trained on samples of matching query and document pairs. Our experiments reveal a lost-in-the-long-distance phenomenon, where retrieval accuracy degrades for documents further away in the hierarchy. To address this, we introduce a pretrain-finetune recipe that significantly improves long-distance retrieval without sacrificing performance on closer documents. We experiment on a realistic hierarchy from WordNet for retrieving documents at various levels of abstraction, and show that pretrain-finetune boosts the recall on long-distance pairs from 19% to 76%. Finally, we demonstrate that our method improves retrieval of relevant products on a shopping queries dataset. Chong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka, Felix X. Yu, Sanjiv Kumar |
NeurIPS | 4 |