VLDB 2026 Research / reviewers in the wild / expert
Maroua Maachou
dblp:309/5846
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
token pruning |
0.6 | 1 | 2022 | Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022 |
Information retrieval › retrieval models › neural retrieval
late interaction retrieval |
0.2 | 1 | 2022 | Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022 |
Information retrieval
retrieval models |
0.2 | 1 | 2022 | Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022 |
Methods — techniques the papers use, named apart from their topics
attention mechanism · 0.6ColBERT · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT)abstractBERT-based rankers have been shown very effective as rerankers in information retrieval tasks. In order to extend these models to full-ranking scenarios, the ColBERT model has been recently proposed, which adopts a late interaction mechanism. This mechanism allows for the representation of documents to be precomputed in advance. However, the late-interaction mechanism leads to large index size, as one needs to save a representation for each token of every document. In this work, we focus on token pruning techniques in order to mitigate this problem. We test four methods, ranging from simpler ones to the use of a single layer of attention mechanism to select the tokens to keep at indexing time. Our experiments show that for the MS MARCO-passages collection, indexes can be pruned up to 70% of their original size, without a significant drop in performance. We also evaluate on the MS MARCO-documents collection and the BEIR benchmark, which reveals some challenges for the proposed mechanism. Carlos Eduardo Rosar Kós Lassance, Maroua Maachou, Joohee Park, Stéphane Clinchant |
SIGIR | 2 |