Rohan Jha

dblp:214/8071 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0001-5008-4949ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Vector Index Compression in Any Modality
Hanxiang Qin, Alexander Martin 0006, Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
SIGIR3
2026 ColBERTSaR: Sparsified ColBERT Index via Product Quantization
abstract
While ColBERT is an effective neural retrieval architecture, it requires a heavy index structure to support candidate set retrieval based on approximated token embeddings, gathering and decompressing document token embeddings, and applying the MaxSim operation. Indexes in PLAID and similar ColBERT implementations require five to ten times the disk storage of the original raw text, which limits their scalability. Furthermore, prior work has identified that the gathering and decompression stages are the primary inefficiencies at query time. Limiting the number of document tokens that must be gathered by thresholding and score approximation does not eliminate the need for the entire index to support ad hoc queries. In this work, we propose an embedding quantization approach that turns a ColBERT index into a true inverted index. We show that, theoretically, ColBERT with embedding quantization is equivalent to learned-sparse retrieval except for the scoring mechanism. Empirically, we demonstrate that our index is 50-70% smaller than a one-bit PLAID index while retaining retrieval effectiveness.
Eugene Yang 0001, Andrew Yates, Dawn J. Lawrie, James Mayfield, Saron Samuel, Rohan Jha
SIGIR6
2024 Generalizable Tip-of-the-Tongue Retrieval with LLM Re-ranking
abstract
Tip-of-the-Tongue (ToT) retrieval is challenging for search engines because the queries are usually natural-language, verbose, and contain uncertain and inaccurate information. This paper studies the generalization capabilities of existing retrieval methods with ToT queries in multiple domains. We curate a multi-domain dataset and evaluate the effectiveness of recall-oriented first-stage retrieval methods across the different domains, considering in-domain, out-of-domain, and multi-domain training settings. We further explore the use of a Large Language Model (LLM), i.e. GPT-4, for zero-shot re-ranking in various ToT domains, relying solely on the item titles. Results show that multi-domain training enhances recall, and that LLMs are strong zero-shot re-rankers, especially for popular items, outperforming direct GPT-4 prompting without first-stage retrieval. Datasets and code can be found on GitHub https://github.com/LuisPB7/TipTongue
Luís Borges, Rohan Jha, Jamie Callan, Bruno Martins 0001
SIGIR2
2023 COILcr: Efficient Semantic Matching in Contextualized Exact Match Retrieval
Zhen Fan 0003, Luyu Gao, Rohan Jha, Jamie Callan
ECIR (1)3
2021 Predicting Inductive Biases of Pre-Trained Models
Charles Lovering, Rohan Jha, Tal Linzen, Ellie Pavlick
ICLR2