VLDB 2026 Research / reviewers in the wild / expert
Suraj Nair 0001
dblp:52/5152-1 · also Suraj Rajappan Nair
· DBLP profile ↗
7ranked-venue papers in the field
3as first author
6since 2021 · last 2025
0000-0003-2283-7672ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 7 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Token Pruning Optimization for Efficient Multi-vector Dense Retrieval
Shanxiu He, Mutasem Al-Darabsah, Suraj Nair 0001, Jonathan May, Tarun Agarwal, Tao Yang 0009, Choon Hui Teo |
ECIR (1) | 3 |
| 2023 | HC3: A Suite of Test Collections for CLIR Evaluation over Informal TextabstractWhile there are many test collections for Cross-Language Information Retrieval (CLIR), none of the large public test collections focus on short informal text documents. This paper introduces a new pair of CLIR test collections with millions of Chinese or Persian Tweets or Tweet threads as documents, sixty event-motivated topics written both in English and in each of the two document languages, and three-point graded relevance judgments constructed using interactive search and active learning. The design and construction of these new test collections are described, and baseline results are presented that demonstrate the utility of the collections for system evaluation. Shallow pooling is used to assess the efficacy of active learning to select documents for judgment. Dawn J. Lawrie, James Mayfield, Douglas W. Oard, Eugene Yang 0001, Suraj Nair 0001, Petra Galuscáková |
SIGIR | 5 |
| 2023 | BLADE: Combining Vocabulary Pruning and Intermediate Pretraining for Scaleable Neural CLIRabstractLearning sparse representations using pretrained language models enhances the monolingual ranking effectiveness. Such representations are sparse vectors in the vocabulary of a language model projected from document terms. Extending such approaches to Cross-Language Information Retrieval (CLIR) using multilingual pretrained language models poses two challenges. First, the larger vocabularies of multilingual models affect both training and inference efficiency. Second, the representations of terms from different languages with similar meanings might not be sufficiently similar. To address these issues, we propose a learned sparse representation model, BLADE, combining vocabulary pruning with intermediate pre-training based on cross-language supervision. Our experiments reveal BLADE significantly reduces indexing time compared to its monolingual counterpart, SPLADE, on machine-translated documents, and it generates rankings with strengths complementary to those of other efficient CLIR methods. Suraj Nair 0001, Eugene Yang 0001, Dawn J. Lawrie, James Mayfield, Douglas W. Oard |
SIGIR | 1 |
| 2023 | Neural Methods for Cross-Language Information RetrievalabstractThis half day tutorial introduces the participant to the basic concepts underlying neural Cross-Language Information Retrieval (CLIR). It discusses the most common algorithmic approaches to CLIR, focusing on modern neural methods; the history of CLIR; where to find and how to use CLIR training collections, test collections and baseline systems; how CLIR training and test collections are constructed; and open research questions in CLIR. Eugene Yang 0001, Dawn J. Lawrie, James Mayfield, Suraj Nair 0001, Douglas W. Oard |
SIGIR | 4 |
| 2022 | Transfer Learning Approaches for Building Cross-Language Dense Retrieval Models
Suraj Nair 0001, Eugene Yang 0001, Dawn J. Lawrie, Kevin Duh, Paul McNamee, Kenton Murray, James Mayfield, Douglas W. Oard |
ECIR (1) | 1 |
| 2022 | C3: Continued Pretraining with Contrastive Weak Supervision for Cross Language Ad-Hoc RetrievalabstractPretrained language models have improved effectiveness on numerous tasks, including ad-hoc retrieval. Recent work has shown that continuing to pretrain a language model with auxiliary objectives before fine-tuning on the retrieval task can further improve retrieval effectiveness. Unlike monolingual retrieval, designing an appropriate auxiliary task for cross-language mappings is challenging. To address this challenge, we use comparable Wikipedia articles in different languages to further pretrain off-the-shelf multilingual pretrained models before fine-tuning on the retrieval task. We show that our approach yields improvements in retrieval effectiveness. Eugene Yang 0001, Suraj Nair 0001, Ramraj Chandradevan, Rebecca Iglesias-Flores, Douglas W. Oard |
SIGIR | 2 |
| 2020 | Combining Contextualized and Non-contextualized Query Translations to Improve CLIRabstractIn cross-language information retrieval using probabilistic structured queries (PSQ), translation probabilities from statistical machine translation act as a bridge between the query and document vocabulary. These translation probabilities are typically estimated from a sentence-aligned corpus on a word to word basis without taking into account the context. Neural methods, by contrast, can learn to translate using the context around the words, and this can be used as a basis for estimating context-dependent translation probabilities. However, sparsity limits the accuracy of context-specific translation probabilities for rare words, which can be important in retrieval applications. This paper presents evidence that combining such context-dependent translation probabilities with context-independent translation probabilities learned from the same parallel corpus can yield improvements in the effectiveness of cross-language ranked retrieval. Suraj Nair 0001, Petra Galuscáková, Douglas W. Oard |
SIGIR | 1 |