Yash Kankanampati

dblp:280/0847 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0001-1451-4593ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › retrieval models › neural retrieval
dense retrieval
1.012026
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026
Information retrieval › indexing
index compression
1.012026
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026
Information retrieval › retrieval models › neural retrieval
late interaction retrieval
1.012026
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026
Information retrieval
token pruning
1.012026
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026
Information retrieval
embedding space geometry
0.312026
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models · SIGIR 2026

Methods — techniques the papers use, named apart from their topics

voronoi cell estimation · 1.0hyperspace geometry · 1.0
YearPublicationVenuePosition
2026 Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha 0002, Timothee Mickus, Raúl Vázquez, Patanjali Bhamidipati, Claudio Savelli, Ahana Chattopadhyay, Laura A. Zanella, Yash Kankanampati, Binesh Arakkal Remesh, Aryan Ashok Chandramania, Chuyuan Li, Ioana Buhnila, Radhika Mamidi
LREC9
2026 A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
abstract
Late-interaction models such as ColBERT offer competitive performance across various retrieval tasks but require storing a dense embedding for each document token, leading to a substantial index storage overhead. Past works address this by attempting to prune low-importance token embeddings based on statistical and empirical measures, but they often either lack formal grounding or are ineffective. To address these shortcomings, we introduce a framework grounded in hyperspace geometry and cast token pruning as a Voronoi cell estimation problem in the embedding space. By interpreting each token's influence as a measure of its Voronoi region, our approach enables principled pruning that retains retrieval quality while reducing index size. Through our experiments, we demonstrate that this approach serves not only as a competitive pruning strategy but also as a valuable tool for improving and interpreting token-level behavior within dense retrieval systems.
Yash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski, Joseph Le Roux
SIGIR1
2020 Multitask Easy-First Dependency Parsing: Exploiting Complementarities of Different Dependency Representations
abstract
We present a parsing model for projective dependency trees which takes advantage of the existence of complementary dependency annotations for a language.This is the case for Arabic with the availability of CATiB and UD treebanks.Our system performs syntactic parsing according to both annotation types jointly as a sequence of arc-creating operations following the Easy-First approach, and partially created trees for one annotation type are also available to the other as features for the score function.This method gives error reduction of 9.9% on CATiB and 6.1% on UD compared to a single-task baseline, and ablation tests show that the main contribution of this reduction is given by sharing tree representation between tasks, and not simply sharing BiLSTM layers as is usually performed in NLP multitask systems.
Yash Kankanampati, Joseph Le Roux, Nadi Tomeh, Dima Taji, Nizar Habash
COLING1