Dwaipayan Roy 0001

dblp:151/2098 · DBLP profile ↗
← Back
24ranked-venue papers in the field
7as first author
14since 2021 · last 2026
0000-0002-5962-5983ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 24 (7 first)
YearPublicationVenuePosition
2026 Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR
Suchana Datta, Dwaipayan Roy 0001, Derek Greene, Gerardine Meaney, Karen Wade, Philipp Mayr 0001
ECIR (3)2
2026 AgriIR: A Scalable Framework for Domain-Specific Knowledge Retrieval
Shuvam Banerji Seal, Aheli Poddar, Dwaipayan Roy 0001
ECIR (3)4
2026 When RAG Disagrees: Detecting Latent Epistemic Conflict via Logit Interactions
Saisab Sadhu, Dwaipayan Roy 0001, Tanmay Basu
SIGIR2
2026 MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval
abstract
Users increasingly expect modern search systems to offer a unified interface that seamlessly retrieves information from diverse data sources and formats. However, current information retrieval (IR) evaluation benchmarks have not kept pace with this development, primarily due to the lack of test collections that represent the diversity of contemporary search domains. We address this critical gap with MIRA, a novel benchmark based on a large-scale social science search platform. MIRA is designed for category-aware ranking across heterogeneous categories – Publications, Research Data, Variables, and Instruments & Tools – within a single, unified evaluation framework. The proposed collection is distinctive in several ways: (1) it is built upon real user queries, providing a more realistic basis for evaluation; (2) it covers scholarly items from four distinct categories, enabling multi-faceted evaluation; and (3) it leverages a Large Language Model to generate topic descriptions and narratives, as well as for relevance assessment with respect to these topics, substantially reducing the labor and cost of test collection generation. We release this resource to benefit the community by providing a foundational testbed for the research on multi-faceted, category-aware, integrated, or cross-category information retrieval.
Mehmet Deniz Türkmen, Suchana Datta, Dwaipayan Roy 0001, Daniel Hienert, Philipp Mayr 0001, Derek Greene
SIGIR3
2025 Tales and Truths: Exploring the Linguistic Journey of 19th Century Literature and Non-fiction
Suchana Datta, Dwaipayan Roy 0001, Derek Greene, Gerardine Meaney
ECIR (4)2
2025 Combining Query Performance Predictors: A Reproducibility Study
Sourav Saha 0003, Suchana Datta, Dwaipayan Roy 0001, Mandar Mitra, Derek Greene
ECIR (4)3
2025 BAAF: A Framework for Media Bias Detection
Soumyadeep Sar, Subinay Adhikary, Dwaipayan Roy 0001
ECIR (3)3
2025 FACTors: A New Dataset for Studying the Fact-checking Ecosystem
abstract
Our fight against false information is spearheaded by fact-checkers. They investigate the veracity of claims and document their findings as fact-checking reports. With the rapid increase in the amount of false information circulating online, the use of automation in fact-checking processes aims to strengthen this ecosystem by enhancing scalability. Datasets containing fact-checked claims play a key role in developing such automated solutions. However, to the best of our knowledge, there is no fact-checking dataset at the ecosystem level, covering claims from a sufficiently long period of time and sourced from a wide range of actors reflecting the entire ecosystem that admittedly follows widely-accepted codes and principles of fact-checking.
Enes Altuncu, Can Baskent, Sanjay Bhattacherjee, Shujun Li 0001, Dwaipayan Roy 0001
SIGIR5
2024 Exploring the Nexus Between Retrievability and Query Generation Strategies
Aman Sinha 0003, Priyanshu Raj Mall, Dwaipayan Roy 0001
ECIR (4)3
2023 Findability: A Novel Measure of Information Accessibility
abstract
The overwhelming volume of data generated and indexed by search engines poses a significant challenge in retrieving documents from the index efficiently and effectively. Even with a well-crafted query, several relevant documents often get buried among a multitude of competing documents, resulting in reduced accessibility or "findability" of the desired document. Consequently, it is crucial to develop a robust methodology for assessing this dimension of Information Retrieval (IR) system performance. While previous studies have focused on measuring document accessibility disregarding user queries and document relevance, there exists no metric to quantify the findability of a document within a given IR system without resorting to manual labor. This paper aims to address this gap by defining and deriving a metric to evaluate the findability of documents as perceived by end-users. Through experiments, we demonstrate the varying impact of different retrieval models and collections on the findability of documents. Furthermore, we establish the findability measure as an independent metric distinct from retrievability, an accessibility measure introduced in prior literature.
Aman Sinha 0003, Priyanshu Raj Mall, Dwaipayan Roy 0001
CIKM3
2023 Weakly supervised deep metric learning on discrete metric spaces for privacy-preserved clustering
Chandan Biswas, Debasis Ganguly, Dwaipayan Roy 0001, Ujjwal Bhattacharya
Inf. Process. Manag.3
2022 Measuring and Comparing the Consistency of IR Models for Query Pairs with Similar and Different Information Needs
abstract
A widespread use of supervised ranking models has necessitated an investigation on how consistent their outputs align with user expectations. While a match between the user expectations and system outputs can be sought at different levels of granularity, we study this alignment for search intent transformation across a pair of queries. Specifically, we propose a consistency metric, which for a given pair of queries - one reformulated from the other with at least one term in common, measures if the change in the set of the top-retrieved documents induced by this reformulation is as per a user's expectation. Our experiments led to a number of observations, such as DRMM (an early interaction based IR model) exhibits better alignment with set-level user expectations, whereas transformer-based neural models (e.g., MonoBERT) agree more consistently with the content and rank-based expectations of overlap.
Procheta Sen, Sourav Saha 0003, Debasis Ganguly, Manisha Verma, Dwaipayan Roy 0001
CIKM5
2022 Information asymmetry in Wikipedia across different languages: A statistical analysis
abstract
Abstract Wikipedia is the largest web‐based open encyclopedia covering more than 300 languages. Different language editions of Wikipedia differ significantly in terms of their information coverage. In this article, we compare the information coverage in English Wikipedia (most exhaustive) and Wikipedias in 8 other widely spoken languages, namely Arabic, German, Hindi, Korean, Portuguese, Russian, Spanish, and Turkish. We analyze variations in different language editions of Wikipedia in terms of the number of topics covered as well as the amount of information discussed about different topics. Further, as a step towards bridging the information gap, we present WikiCompare—a browser plugin that allows Wikipedia readers to have a comprehensive overview of topics by incorporating missing information from Wikipedia page in other language.
Dwaipayan Roy 0001, Sumit Bhatia
J. Assoc. Inf. Sci. Technol.1
2021 Tag embedding based personalized point of interest recommendation system
Suraj Agrawal, Dwaipayan Roy 0001, Mandar Mitra
Inf. Process. Manag.2
2020 Characteristics of Dataset Retrieval Sessions: Experiences from a Real-Life Digital Library
Zeljko Carevic, Dwaipayan Roy 0001, Philipp Mayr 0001
TPDL2
2020 Retrieving Potential Causes from a Query Event
abstract
Different to traditional IR, which retrieves a set of topically relevant documents given a user query, we investigate causal retrieval, which involves retrieving a set of documents that describe a set of potential causes leading to an effect specified in the query. We argue that the nature of causal relevance should be different to that of traditional topical relevance. This is because although the causally relevant documents would have partial term overlap with the ones that are topically relevant for a query, yet it is expected that a majority of these documents would use a different set of terms to describe a number of causes possibly leading to their effects. To address this, we propose a feedback model to estimate a distribution of terms which are relatively infrequent but associated with high weights in the topically relevant distribution, leading to potential causal relevance. Our experiments demonstrate that such a feedback model turns out to be substantially more effective than traditional IR models and a number of other causality heuristic baselines.
Suchana Datta, Debasis Ganguly, Dwaipayan Roy 0001, Francesca Bonin, Charles Jochim, Mandar Mitra
SIGIR3
2019 Privacy Preserving Approximate K-means Clustering
abstract
Privacy preserving computation is of utmost importance in a cloud computing environment where a client often requires to send sensitive data to servers offering computing services over untrusted networks. Eavesdropping over the network or malware at the server may lead to leaking sensitive information from the data. To prevent this, we propose to encode the input data in such a way that, firstly, it should be difficult to decode it back to the true data, and secondly, the computational results obtained with the encoded data should not be substantially different from those obtained with the true data. Specifically, the computational activity that we focus on is the K-means clustering, which is widely used for many data mining tasks. Our proposed variant of the K-means algorithm is capable of privacy preservation in the sense that it requires as input only binary encoded data, and is not allowed to access the true data vectors at any stage of the computation. During intermediate stages of K-means computation, our algorithm is able to effectively process the inputs with incomplete information seeking to yield outputs relatively close to the complete information (non-encoded) case. Evaluation on real datasets show that the proposed methods yields comparable clustering effectiveness in comparison to the standard K-means algorithm on image clustering (MNIST-8M dataset), and in fact outperforms the standard K-means on text clustering (ODPtweets dataset).
Chandan Biswas, Debasis Ganguly, Dwaipayan Roy 0001, Ujjwal Bhattacharya
CIKM3
2019 I-REX: A Lucene Plugin for EXplainable IR
abstract
Providing high-level, intuitive explanations of the performance of IR systems is generally difficult due to their complexity, and the various low-level implementation details involved. We present I-REX, a tool built on top of Lucene, that is intended to provide a systematic view into the inner workings of retrieval models and methods (specifically query expansion). This should help researchers study, compare, understand and explain the performance of these models and methods. I-REX can be run either as a Web service accessible through a browser, or as a terminal-based tool with a shell-like interactive interface. In this article, we describe a session that illustrates how I-REX can be used to explain the observed difference in the performance of two variants of the Language Model.
Dwaipayan Roy 0001, Sourav Saha 0003, Mandar Mitra, Bihan Sen, Debasis Ganguly
CIKM1
2019 Selecting Discriminative Terms for Relevance Model
abstract
Pseudo-relevance feedback based on the relevance model does not take into account the inverse document frequency of candidate terms when selecting expansion terms. As a result, common terms are often included in the expanded query constructed by this model. We propose three possible extensions of the relevance model that address this drawback. Our proposed extensions are simple to compute and are independent of the base retrieval model. Experiments on several TREC news and web collections show that the proposed modifications yield significantly better MAP, precision, NDCG, and recall values than the original relevance model as well as its two recently proposed state-of-the-art variants.
Dwaipayan Roy 0001, Sumit Bhatia, Mandar Mitra
SIGIR1
2019 Estimating Gaussian mixture models in the local neighbourhood of embedded word vectors for query performance prediction
Dwaipayan Roy 0001, Debasis Ganguly, Mandar Mitra, Gareth J. F. Jones
Inf. Process. Manag.1
2018 Using Word Embeddings for Information Retrieval: How Collection and Term Normalization Choices Affect Performance
abstract
Neural word embedding approaches, due to their ability to capture semantic meanings of vocabulary terms, have recently gained attention of the information retrieval (IR) community and have shown promising results in improving ad hoc retrieval performance. It has been observed that these approaches are sensitive to various choices made during the learning of word embeddings and their usage, often leading to poor reproducibility. We study the effect of varying following two parameters, viz., i) the term normalization and ii) the choice of training collection, on ad hoc retrieval performance with word2vec and fastText embeddings. We present quantitative estimates of similarity of word vectors obtained under different settings, and use embeddings based query expansion task to understand the effects of these parameters on IR effectiveness.
Dwaipayan Roy 0001, Debasis Ganguly, Sumit Bhatia, Srikanta J. Bedathur, Mandar Mitra
CIKM1
2017 An Improved Test Collection and Baselines for Bibliographic Citation Recommendation
abstract
The problem of recommending bibliographic citations to an author who is writing an article has been well-studied. However, different researchers have used different datasets to evaluate proposed techniques, and have sometimes reported contradictory findings regarding the relative effectiveness of various approaches. In addition, these datasets are problematic in one way or another (e.g., in terms of size or availability), precluding the possibility of adopting one (or some) of them as standard benchmarks. A recently created test collection that makes use of data from CiteSeerx is large, heterogenous, and publicly available, but has certain other limitations. In this paper, we propose a way to modify this test collection to address these limitations. We also use the improved test collection to establish a set of baseline results using elementary content-based techniques, as well as reference directed indexing.
Dwaipayan Roy 0001
CIKM1
2016 Word Vector Compositionality based Relevance Feedback using Kernel Density Estimation
abstract
A limitation of standard information retrieval (IR) models is that the notion of term composionality is restricted to pre-defined phrases and term proximity. Standard text based IR models provide no easy way of representing semantic relations between terms that are not necessarily phrases, such as the equivalence relationship between `osteoporosis' and the terms `bone' and `decay'. To alleviate this limitation, we introduce a relevance feedback (RF) method which makes use of word embedded vectors. We leverage the fact that the vector addition of word embeddings leads to a semantic composition of the corresponding terms, e.g. addition of the vectors for `bone' and `decay' yields a vector that is likely to be close to the vector for the word `osteoporosis'. Our proposed RF model enables incorporation of semantic relations by exploiting term compositionality with embedded word vectors. We develop our model for RF as a generalization of the relevance model (RLM). Our experiments demonstrate that our word embedding based RF model significantly outperforms the RLM model on standard TREC test collections, namely the TREC 6,7,8 and Robust ad-hoc and the TREC 9 and 10 WT10G test collections.
Dwaipayan Roy 0001, Debasis Ganguly, Mandar Mitra, Gareth J. F. Jones
CIKM1
2015 Word Embedding based Generalized Language Model for Information Retrieval
abstract
Word2vec, a state-of-the-art word embedding technique has gained a lot of interest in the NLP community. The embedding of the word vectors helps to retrieve a list of words that are used in similar contexts with respect to a given word. In this paper, we focus on using the word embeddings for enhancing retrieval effectiveness. In particular, we construct a generalized language model, where the mutual independence between a pair of words (say t and t') no longer holds. Instead, we make use of the vector embeddings of the words to derive the transformation probabilities between words. Specifically, the event of observing a term t in the query from a document d is modeled by two distinct events, that of generating a different term t', either from the document itself or from the collection, respectively, and then eventually transforming it to the observed query term t. The first event of generating an intermediate term from the document intends to capture how well does a term contextually fit within a document, whereas the second one of generating it from the collection aims to address the vocabulary mismatch problem by taking into account other related terms in the collection. Our experiments, conducted on the standard TREC collection, show that our proposed method yields significant improvements over LM and LDA-smoothed LM baselines.
Debasis Ganguly, Dwaipayan Roy 0001, Mandar Mitra, Gareth J. F. Jones
SIGIR2