Maroua Maachou

dblp:309/5846 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
token pruning
0.612022
Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022
Information retrieval › retrieval models › neural retrieval
late interaction retrieval
0.212022
Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022
Information retrieval
retrieval models
0.212022
Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT) · SIGIR 2022

Methods — techniques the papers use, named apart from their topics

attention mechanism · 0.6ColBERT · 0.6
YearPublicationVenuePosition
2022 Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT)
abstract
BERT-based rankers have been shown very effective as rerankers in information retrieval tasks. In order to extend these models to full-ranking scenarios, the ColBERT model has been recently proposed, which adopts a late interaction mechanism. This mechanism allows for the representation of documents to be precomputed in advance. However, the late-interaction mechanism leads to large index size, as one needs to save a representation for each token of every document. In this work, we focus on token pruning techniques in order to mitigate this problem. We test four methods, ranging from simpler ones to the use of a single layer of attention mechanism to select the tokens to keep at indexing time. Our experiments show that for the MS MARCO-passages collection, indexes can be pruned up to 70% of their original size, without a significant drop in performance. We also evaluate on the MS MARCO-documents collection and the BEIR benchmark, which reveals some challenges for the proposed mechanism.
Carlos Eduardo Rosar Kós Lassance, Maroua Maachou, Joohee Park, Stéphane Clinchant
SIGIR2