Sumit Bhatia

dblp:52/7536 · DBLP profile ↗
← Back
32ranked-venue papers in the field
12as first author
12since 2021 · last 2026
0000-0002-8146-4100ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 24 (8 first)Data Mining & Knowledge Discovery · 3 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)Database Systems & Data Management · 1 (1 first)Business Process & Enterprise Data · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 The 9th International Workshop on Narrative Extraction from Text: Text2Story 2026
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (3)4
2026 Enhancing Enterprise Assistant Responses with Rich Multimodal Artifacts from Product Documentation
Sohan Patnaik, Sai Sree Harsha, H. S. V. N. S. Kowndinya Renduchintala, Milan Aggarwal, Sumit Bhatia, Yunyao Li 0001
SIGIR5
2025 The 8th International Workshop on Narrative Extraction from Texts: Text2Story 2025
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (5)4
2025 Exploring the Role of Diversity in Example Selection for In-Context Learning
abstract
In-Context Learning (ICL) has gained prominence due to its ability to perform tasks without requiring extensive training data and its robustness to noisy labels. A typical ICL workflow involves selecting localized examples relevant to a given input using sparse or dense embedding-based similarity functions. However, relying solely on similarity-based selection may introduce topical biases in the retrieved contexts, potentially leading to suboptimal downstream performance. We posit that reranking the retrieved context to enhance topical diversity can improve downstream task performance. To achieve this, we leverage maximum marginal relevance (MMR) which balances topical similarity with inter-example diversity. Our experimental results demonstrate that diversifying the selected examples leads to consistent improvements in downstream performance across various context sizes and similarity functions. The implementation of our approach is made available at https://github.com/janak11111/Diverse-ICL.
Janak Kapuriya, Manit Kaushik, Debasis Ganguly, Sumit Bhatia
SIGIR4
2024 The 7th International Workshop on Narrative Extraction from Texts: Text2Story 2024
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (5)4
2024 GenACT: An Ontology-Based Temporal Web Data Generator
Gunjan Singh, Udit Arora, Shashikant Kumar, Riccardo Tommasini 0001, Pieter Bonte, Sumit Bhatia, Raghava Mutharaju
ER6
2023 The 6th International Workshop on Narrative Extraction from Texts: Text2Story 2023
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (3)4
2023 Explain Like I am BM25: Interpreting a Dense Model's Ranked-List with a Sparse Approximation
abstract
Neural retrieval models (NRMs) have been shown to outperform their statistical counterparts owing to their ability to capture semantic meaning via dense document representations. These models, however, suffer from poor interpretability as they do not rely on explicit term matching. As a form of local per-query explanations, we introduce the notion of equivalent queries that are generated by maximizing the similarity between the NRM's results and the result set of a sparse retrieval system with the equivalent query. We then compare this approach with existing methods such as RM3-based query expansion and contrast differences in retrieval effectiveness and in the terms generated by each approach.
Michael Llordes, Debasis Ganguly, Sumit Bhatia, Chirag Agarwal
SIGIR3
2022 Why Did You Not Compare with That? Identifying Papers for Use as Baselines
Manjot Bedi, Tanisha Pandey, Sumit Bhatia, Tanmoy Chakraborty 0002
ECIR (1)3
2022 The 5th International Workshop on Narrative Extraction from Texts: Text2Story 2022
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Marina Litvak
ECIR (2)4
2022 Information asymmetry in Wikipedia across different languages: A statistical analysis
abstract
Abstract Wikipedia is the largest web‐based open encyclopedia covering more than 300 languages. Different language editions of Wikipedia differ significantly in terms of their information coverage. In this article, we compare the information coverage in English Wikipedia (most exhaustive) and Wikipedias in 8 other widely spoken languages, namely Arabic, German, Hindi, Korean, Portuguese, Russian, Spanish, and Turkish. We analyze variations in different language editions of Wikipedia in terms of the number of topics covered as well as the amount of information discussed about different topics. Further, as a step towards bridging the information gap, we present WikiCompare—a browser plugin that allows Wikipedia readers to have a comprehensive overview of topics by incorporating missing information from Wikipedia page in other language.
Dwaipayan Roy 0001, Sumit Bhatia
J. Assoc. Inf. Sci. Technol.2
2021 The 4th International Workshop on Narrative Extraction from Texts: Text2Story 2021
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia, Mark A. Finlayson
ECIR (2)4
2020 The 3rd International Workshop on Narrative Extraction from Texts: Text2Story 2020
Ricardo Campos 0001, Alípio Mário Jorge, Adam Jatowt, Sumit Bhatia
ECIR (2)4
2020 OWL2Bench: A Benchmark for OWL 2 Reasoners
Gunjan Singh, Sumit Bhatia, Raghava Mutharaju
ISWC (2)2
2019 The 2nd International Workshop on Narrative Extraction from Text: Text2Story 2019
Alípio Mário Jorge, Ricardo Campos 0001, Adam Jatowt, Sumit Bhatia
ECIR (2)4
2019 Selecting Discriminative Terms for Relevance Model
abstract
Pseudo-relevance feedback based on the relevance model does not take into account the inverse document frequency of candidate terms when selecting expansion terms. As a result, common terms are often included in the expanded query constructed by this model. We propose three possible extensions of the relevance model that address this drawback. Our proposed extensions are simple to compute and are independent of the base retrieval model. Experiments on several TREC news and web collections show that the proposed modifications yield significantly better MAP, precision, NDCG, and recall values than the original relevance model as well as its two recently proposed state-of-the-art variants.
Dwaipayan Roy 0001, Sumit Bhatia, Mandar Mitra
SIGIR2
2018 Using Word Embeddings for Information Retrieval: How Collection and Term Normalization Choices Affect Performance
abstract
Neural word embedding approaches, due to their ability to capture semantic meanings of vocabulary terms, have recently gained attention of the information retrieval (IR) community and have shown promising results in improving ad hoc retrieval performance. It has been observed that these approaches are sensitive to various choices made during the learning of word embeddings and their usage, often leading to poor reproducibility. We study the effect of varying following two parameters, viz., i) the term normalization and ii) the choice of training collection, on ad hoc retrieval performance with word2vec and fastText embeddings. We present quantitative estimates of similarity of word vectors obtained under different settings, and use embeddings based query expansion task to understand the effects of these parameters on IR effectiveness.
Dwaipayan Roy 0001, Debasis Ganguly, Sumit Bhatia, Srikanta J. Bedathur, Mandar Mitra
CIKM3
2018 That's Interesting, Tell Me More! Finding Descriptive Support Passages for Knowledge Graph Relationships
Sumit Bhatia, Purusharth Dwivedi
ISWC (1)1
2017 Tools and Infrastructure for Supporting Enterprise Knowledge Graphs
Sumit Bhatia, Nidhi Rajshree, Anshu N. Jain, Nitish Aggarwal
ADMA1
2016 Proactive Information Retrieval: Anticipating Users' Information Need
Sumit Bhatia, Debapriyo Majumdar, Nitish Aggarwal
ECIR1
2016 Identifying the role of individual user messages in an online discussion and its use in thread retrieval
abstract
Online discussion forums have become a popular medium for users to discuss with and seek information from other users having similar interests. A typical discussion thread consists of a sequence of posts posted by multiple users. Each post in a thread serves a different purpose providing different types of information and, thus, may not be equally useful for all applications. Identifying the purpose and nature of each post in a discussion thread is thus an interesting research problem as it can help in improving information extraction and intelligent assistance techniques. We study the problem of classifying a given post as per its purpose in the discussion thread and employ features based on the post's content, structure of the thread, behavior of the participating users, and sentiment analysis of the post's content. We evaluate our approach on two forum data sets belonging to different genres and achieve strong classification performance. We also analyze the relative importance of different features used for the post classification task. Next, as a use case, we describe how the post class information can help in thread retrieval by incorporating this information in a state‐of‐the‐art thread retrieval model.
Sumit Bhatia, Prakhar Biyani, Prasenjit Mitra 0001
J. Assoc. Inf. Sci. Technol.1
2015 Using Subjectivity Analysis to Improve Thread Retrieval in Online Forums
Prakhar Biyani, Sumit Bhatia, Cornelia Caragea, Prasenjit Mitra 0001
ECIR2
2015 Predicting Future Scientific Discoveries Based on a Networked Analysis of the Past Literature
abstract
We present KnIT, the Knowledge Integration Toolkit, a system for accelerating scientific discovery and predicting previously unknown protein-protein interactions. Such predictions enrich biological research and are pertinent to drug discovery and the understanding of disease. Unlike a prior study, KnIT is now fully automated and demonstrably scalable. It extracts information from the scientific literature, automatically identifying direct and indirect references to protein interactions, which is knowledge that can be represented in network form. It then reasons over this network with techniques such as matrix factorization and graph diffusion to predict new, previously unknown interactions. The accuracy and scope of KnIT's knowledge extractions are validated using comparisons to structured, manually curated data sources as well as by performing retrospective studies that predict subsequent literature discoveries using literature available prior to a given date. The KnIT methodology is a step towards automated hypothesis generation from text, with potential application to other scientific domains.
Meena Nagarajan, Angela D. Wilkins, Benjamin J. Bachman, Ilya B. Novikov, Shenghua Bao, Peter J. Haas, María E. Terrón-Díaz, Sumit Bhatia, Anbu K. Adikesavan, Jacques J. Labrie, Sam Regenbogen, Christie M. Buchovecky, Curtis R. Pickering, Linda Kato, Andreas Martin Lisewski, Ana Lelescu, Houyin Zhang, Stephen Boyer, Griff Weber, Ying Chen 0001, Lawrence A. Donehower, W. Scott Spangler, Olivier Lichtarge
KDD8
2013 Monitoring and analyzing customer feedback through social media platforms for identifying and remedying customer problems
abstract
The tremendous growth and popularity of social media platforms like Twitter, Facebook, etc. provides business organizations an opportunity to monitor the feedback from its customers, identify their problems and take corrective measures. In this paper, we describe a system to automatically monitor and analyze customer feedback through various social media platforms like Facebook, Twitter, etc. and detect issues faced by the customers. Business organizations can use this system to engage with their customers and help alleviate the problems faced by them. The system uses statistical event detection techniques for identifying various customer issues. The system offers a batch version as well as real time version of event detection algorithm depending upon the client's requirements. We also describe a few case studies illustrating the utility of our proposed system for business organizations in identifying issues faced by their customers through social media channels.
Sumit Bhatia, Wei Peng 0001, Tong Sun 0001
ASONAM1
2013 Automatic Detection of Pseudocodes in Scholarly Documents Using Machine Learning
abstract
A significant number of scholarly articles in computer science and other disciplines contain algorithms that provide concise descriptions for solving a wide variety of computational problems. For example, Dijkstra's algorithm describes how to find the shortest paths between two nodes in a graph. Automatic identification and extraction of these algorithms from scholarly digital documents would enable automatic algorithm indexing, searching, analysis and discovery. An algorithm search engine, which identifies pseudocodes in scholarly documents and makes them searchable, has been implemented as a part of the CiteSeerX suite. Here, we illustrate the limitations of start-of-the-art rule based pseudocode detection approach, and present a novel set of machine learning based techniques that extend previous methods.
Suppawong Tuarob, Sumit Bhatia, Prasenjit Mitra 0001, C. Lee Giles
ICDAR2
2012 A scalable approach for performing proximal search for verbose patent search queries
abstract
Even though queries received by traditional information retrieval systems are quite short, there are many application scenarios where long natural language queries are more effective. Further, incorporating term position information can help improve results of long queries. However, the techniques for incorporating term position information have been developed for terse queries and hence, can not be directly applied to long queries. Though there exist some methods for performing proximal search for long queries, they are not scalable due to long query response times. We describe an intuitive and simple, yet effective technique that implicitly incorporates term position information for long queries in a scalable manner. Our proposed approach achieves more than 700% faster query response times while maintaining the quality of retrieved results when compared with a state-of-the-art method for performing proximal search for very long queries.
Sumit Bhatia, Bin He 0001, Qi He 0002, W. Scott Spangler
CIKM1
2012 Classifying User Messages For Managing Web Forum Data
Sumit Bhatia, Prakhar Biyani, Prasenjit Mitra 0001
WebDB1
2012 Summarizing figures, tables, and algorithms in scientific publications to augment search results
abstract
Increasingly, special-purpose search engines are being built to enable the retrieval of document-elements like tables, figures, and algorithms [Bhatia et al. 2010; Liu et al. 2007; Hearst et al. 2007]. These search engines present a thumbnail view of document-elements, some document metadata such as the title of the papers and their authors, and the caption of the document-element. While some authors in some disciplines write carefully tailored captions, generally, the author of a document assumes that the caption will be read in the context of the text in the document. When the caption is presented out of context as in a document-element-search-engine result, it may not contain enough information to help the end-user understand what the content of the document-element is. Consequently, end-users examining document-element search results would want a short “synopsis” of this information presented along with the document-element. Having access to the synopsis allows the end-user to quickly understand the content of the document-element without having to download and read the entire document as examining the synopsis takes a shorter time than finding information about a document element by downloading, opening and reading the file. Furthermore, it may allow the end-user to examine more results than they would otherwise. In this paper, we present the first set of methods to extract this useful information (synopsis) related to document-elements automatically. We use Naïve Bayes and support vector machine classifiers to identify relevant sentences from the document text based on the similarity and the proximity of the sentences with the caption and the sentences in the document text that refer to the document-element. We compare the two classification methods and study the effects of different features used. We also investigate the problem of choosing the optimum synopsis-size that strikes a balance between the information content and the size of the generated synopses. A user study is also performed to measure how the synopses generated by our proposed method compare with other state-of-the-art approaches.
Sumit Bhatia, Prasenjit Mitra 0001
ACM Trans. Inf. Syst.1
2011 Multidimensional search result diversification: diverse search results for diverse users
abstract
Hundreds of millions of people today rely on Web based Search Engines to satisfy their information needs. In order to meet the expectations of this vast and diverse user population, the search engine should present a list of results such that the probability of satisfying the average user is maximized. This leads us to the problem of Search Result Diversification. Given a user submitted query, the search engine should include results that are relevant to the user query and at the same time, diverse enough to meet the expectations of diverse user populations. However, it is not clear in what respect the results should be diversified.
Sumit Bhatia
SIGIR1
2011 Query suggestions in the absence of query logs
abstract
After an end-user has partially input a query, intelligent search engines can suggest possible completions of the partial query to help end-users quickly express their information needs. All major web-search engines and most proposed methods that suggest queries rely on search engine query logs to determine possible query suggestions. However, for customized search systems in the enterprise domain, intranet search, or personalized search such as email or desktop search or for infrequent queries, query logs are either not available or the user base and the number of past user queries is too small to learn appropriate models. We propose a probabilistic mechanism for generating query suggestions from the corpus without using query logs. We utilize the document corpus to extract a set of candidate phrases. As soon as a user starts typing a query, phrases that are highly correlated with the partial user query are selected as completions of the partial query and are offered as query suggestions. Our proposed approach is tested on a variety of datasets and is compared with state-of-the-art approaches. The experimental results clearly demonstrate the effectiveness of our approach in suggesting queries with higher quality.
Sumit Bhatia, Debapriyo Majumdar, Prasenjit Mitra 0001
SIGIR1
2010 Finding algorithms in scientific articles
abstract
Algorithms are an integral part of computer science literature. However, none of the current search engines offer specialized algorithm search facility. We describe a vertical search engine that identifies the algorithms present in documents and extracts and indexes the related metadata and textual description of the identified algorithms. This algorithm specific information is then utilized for algorithm ranking in response to user queries. Experimental results show the superiority of our system on other popular search engines.
Sumit Bhatia, Prasenjit Mitra 0001, C. Lee Giles
WWW1
2009 Generating synopses for document-element search
abstract
Scientists often search for document-elements like tables, figures, or algorithm pseudo-codes. Domain scientists and researchers report important data, results and algorithms using these document-elements; readers want to compare the reported results with their findings. Some document-element search engines have been proposed (especially to search for tables and figures) to make this task easier. While searching for document-elements today, the end-user is presented with the caption of the document-element and a sentence in the document text that refers to the document-element. Oftentimes, the caption and the reference text do not contain enough information to interpret the document-element. In this paper, we present the first set of methods to extract this useful information (synopsis) related to document-elements automatically. We also investigate the problem of choosing the optimum synopsis-size that strikes a balance between information content and size of the generated synopses.
Sumit Bhatia, Shibamouli Lahiri, Prasenjit Mitra 0001
CIKM1