VLDB 2026 Research / reviewers in the wild / expert
Bishal Santra
dblp:191/6050
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-0380-689XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 67% Data mining · 33% | |
| Artificial intelligence
4 papers |
Language models and text generation · 51% Information extraction and text analysis · 29% Knowledge representation and reasoning · 21% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization |
0.9 | 1 | 2025 | SCULPT: Systematic Tuning of Long Prompts · ACL (1) 2025 |
Information retrieval › retrieval models › neural retrieval › dense retrieval
bi-encoder retrieval |
0.9 | 1 | 2025 | Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval · EMNLP 2025 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.9 | 1 | 2025 | Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for Retrieval · EMNLP 2025 |
Data mining › predictive modeling › classification › multi-label classification
extreme classification |
0.9 | 1 | 2025 | On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification · KDD (1) 2025 |
Data mining › text mining › text classification › weakly supervised classification
partial label learning |
0.9 | 1 | 2025 | On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification · KDD (1) 2025 |
Information retrieval › retrieval models
query-document relevance |
0.9 | 1 | 2025 | On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme Classification · KDD (1) 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
domain knowledge integration |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › text generation › surface realization
linearization |
0.4 | 1 | 2019 | Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource Languages · ACL (1) 2019 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge Graphs · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › syntactic parsing
word ordering |
0.4 | 1 | 2019 | Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource Languages · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging |
0.3 | 1 | 2018 | Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
word segmentation |
0.3 | 1 | 2018 | Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in Sanskrit · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
small language models · 0.9propensity weighting · 0.9large language model data augmentation · 0.9large language model · 0.9evolutionary search · 0.9data imputation · 0.9black-box optimization · 0.9token embedding · 0.4seq2seq · 0.4pre-training · 0.4near-infrared imaging · 0.4knowledge graph embedding · 0.4convolutional neural network · 0.4energy-based model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SCULPT: Systematic Tuning of Long PromptsabstractShanu Kumar, Akhila Yesantarao Venkata, Shubhanshu Khandelwal, Bishal Santra, Parag Agrawal, Manish Gupta. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shanu Kumar, Akhila Yesantarao Venkata, Shubhanshu Khandelwal, Bishal Santra, Parag Agrawal, Manish Gupta 0001 |
ACL (1) | 4 |
| 2025 | Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for RetrievalabstractCompact dual-encoder models are widely used for retrieval owing to their efficiency and scalability.However, such models often underperform compared to their Large Language Model (LLM)-based retrieval counterparts, likely due to their limited world knowledge.While LLMbased data augmentation has been proposed as a strategy to bridge this performance gap, there is insufficient understanding of its effectiveness and scalability to real-world retrieval problems.Existing research does not systematically explore key factors such as the optimal augmentation scale, the necessity of using large augmentation models, and whether diverse augmentations improve generalization, particularly in out-of-distribution (OOD) settings.This work presents a comprehensive study of the effectiveness of LLM augmentation for retrieval, comprising over 100 distinct experimental settings of retrieval models, augmentation models and augmentation strategies.We find that, while augmentation enhances retrieval performance, its benefits diminish beyond a certain augmentation scale, even with diverse augmentation strategies.Surprisingly, we observe that augmentation with smaller LLMs can achieve performance competitive with larger augmentation models.Moreover, we examine how augmentation effectiveness varies with retrieval model pre-training, revealing that augmentation provides the most benefit to models which are not well pre-trained.Our insights pave the way for more judicious and efficient augmentation strategies, thus enabling informed decisions and maximising retrieval performance while being more cost-effective. Pranjal A. Chitale, Bishal Santra, Yashoteja Prabhu, Amit Sharma 0007 |
EMNLP | 2 |
| 2025 | On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme ClassificationabstractExtreme Classification (XC) aims to map a query to the most relevant documents from a very large document set. XC algorithms used in real-world applications typically learn this mapping from datasets curated from implicit feedback, such as user clicks. However, these datasets often suffer from missing labels. In this work, we observe that systematic missing labels lead to missing knowledge, which is critical for modelling relevance between queries and documents. We formally show that this absence of knowledge is hard to recover using existing methods such as propensity weighting and data imputation strategies that solely rely on the training dataset. While Large Language Models (LLMs) provide an attractive solution to augment the missing knowledge, leveraging them in applications with low latency requirements and large document sets is challenging. To mitigate missing knowledge at scale, we propose SKIM (Scalable Knowledge Infusion for Missing Labels), an algorithm that leverages a combination of Small Language Models or SLMs, e.g., Llama2-7b, and abundant unstructured meta-data to effectively address the missing label problem. We show the efficacy of our method on large-scale public datasets through a combination of unbiased evaluation strategies, such as exhaustive human annotations and simulation-based evaluation benchmarks. SKIM outperforms existing methods on Recall@100 by more than 10 absolute points. Additionally, SKIM scales to proprietary query-ad retrieval datasets containing 10 million documents, outperforming baseline methods by 12% in offline evaluations and increasing ad click-yield by 1.23% in an online A/B test conducted on Bing Search. We release the code and trained models at: github.com/bicycleman15/skim Jatin Prakash, Anirudh Buvanesh, Bishal Santra, Deepak Saini, Sachin Yadav 0002, Jian Jiao 0007, Yashoteja Prabhu, Amit Sharma 0007, Manik Varma |
KDD (1) | 3 |
| 2022 | Representation Learning for Conversational Data using Discourse Mutual Information MaximizationabstractBishal Santra, Sumegh Roychowdhury, Aishik Mandal, Vasu Gurram, Atharva Naik, Manish Gupta, Pawan Goyal. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Bishal Santra, Sumegh Roychowdhury, Aishik Mandal, Vasu Gurram, Atharva Naik, Manish Gupta 0001, Pawan Goyal 0002 |
NAACL-HLT | 1 |
| 2021 | Hierarchical Transformer for Task Oriented Dialog SystemsabstractGenerative models for dialog systems have gained much interest because of the recent success of RNN and Transformer based models in tasks like question answering and summarization.Although the task of dialog response generation is generally seen as a sequence to sequence (Seq2Seq) problem, researchers in the past have found it challenging to train dialog systems using the standard Seq2Seq models.Therefore, to help the model learn meaningful utterance and conversation level features, Sordoni et al. (2015b); Serban et al. ( 2016) proposed Hierarchical RNN architecture, which was later adopted by several other RNN based dialog systems.With the transformer-based models dominating the seq2seq problems lately, the natural question to ask is the applicability of the notion of hierarchy in transformer based dialog systems.In this paper, we propose a generalized framework for Hierarchical Transformer Encoders and show how a standard transformer can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings.We demonstrate that Hierarchical Encoding helps achieve better natural language understanding of the contexts in transformer-based models for task-oriented dialog systems through a wide range of experiments.The code and data for all experiments in this paper has been open-sourced 1 2 . Bishal Santra, Potnuru Anusha, Pawan Goyal 0002 |
NAACL-HLT | 1 |
| 2020 | A Graph-Based Framework for Structured Prediction Tasks in SanskritabstractWe propose a framework using energy-based models for multiple structured prediction tasks in Sanskrit. Ours is an arc-factored model, similar to the graph-based parsing approaches, and we consider the tasks of word segmentation, morphological parsing, dependency parsing, syntactic linearization, and prosodification, a “prosody-level” task we introduce in this work. Ours is a search-based structured prediction framework, which expects a graph as input, where relevant linguistic information is encoded in the nodes, and the edges are then used to indicate the association between these nodes. Typically, the state-of-the-art models for morphosyntactic tasks in morphologically rich languages still rely on hand-crafted features for their performance. But here, we automate the learning of the feature function. The feature function so learned, along with the search space we construct, encode relevant linguistic information for the tasks we consider. This enables us to substantially reduce the training data requirements to as low as 10%, as compared to the data requirements for the neural state-of-the-art models. Our experiments in Czech and Sanskrit show the language-agnostic nature of the framework, where we train highly competitive models for both the languages. Moreover, our framework enables us to incorporate language-specific constraints to prune the search space and to filter the candidates during inference. We obtain significant improvements in morphosyntactic tasks for Sanskrit by incorporating language-specific constraints into the model. In all the tasks we discuss for Sanskrit, we either achieve state-of-the-art results or ours is the only data-driven solution for those tasks. Amrith Krishna, Bishal Santra, Ashim Gupta, Pavankumar Satuluri, Pawan Goyal 0002 |
Comput. Linguistics | 2 |
| 2019 | VPDS: An AI-Based Automated Vehicle Occupancy and Violation Detection SystemabstractHigh Occupancy Vehicle/High Occupancy Tolling (HOV/HOT) lanes are operated based on voluntary HOV declarations by drivers. A majority of these declarations are wrong to leverage faster HOV lane speeds illegally. It is a herculean task to manually regulate HOV lanes and identify these violators. Therefore, an automated way of counting the number of people in a car is prudent for fair tolling and for violator detection.In this paper, we propose a Vehicle Passenger Detection System (VPDS) which works by capturing images through Near Infrared (NIR) cameras on the toll lanes and processing them using deep Convolutional Neural Networks (CNN) models. Our system has been deployed in 3 cities over a span of two years and has served roughly 30 million vehicles with an accuracy of 97% which is a remarkable improvement over manual review which is 37% accurate. Our system can generate an accurate report of HOV lane usage which helps policy makers pave the way towards de-congestion. Abhinav Kumar 0004, Aishwarya Gupta 0001, Bishal Santra, Lalitha K. S., Manasa Kolla, Mayank Gupta 0002, Rishabh Singh |
AAAI | 3 |
| 2019 | Poetry to Prose Conversion in Sanskrit as a Linearisation Task: A Case for Low-Resource LanguagesabstractThe word ordering in a Sanskrit verse is often not aligned with its corresponding prose order.Conversion of the verse to its corresponding prose helps in better comprehension of the construction.Owing to the resource constraints, we formulate this task as a word ordering (linearisation) task.In doing so, we completely ignore the word arrangement at the verse side.kāvya guru, the approach we propose, essentially consists of a pipeline of two pretraining steps followed by a seq2seq model.The first pretraining step learns task specific token embeddings from pretrained embeddings.In the next step, we generate multiple hypotheses for possible word arrangements of the input (Wang et al., 2018).We then use them as inputs to a neural seq2seq model for the final prediction.We empirically show that the hypotheses generated by our pretraining step result in predictions that consistently outperform predictions based on the original order in the verse.Overall, kāvya guru outperforms current state of the art models in linearisation for the poetry to prose conversion task in Sanskrit. Amrith Krishna, Vishnu Dutt Sharma, Bishal Santra, Aishik Chakraborty, Pavankumar Satuluri, Pawan Goyal 0002 |
ACL (1) | 3 |
| 2019 | Incorporating Domain Knowledge into Medical NLI using Knowledge GraphsabstractSoumya Sharma, Bishal Santra, Abhik Jana, Santosh Tokala, Niloy Ganguly, Pawan Goyal. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Soumya Sharma, Bishal Santra, Abhik Jana, Santosh Tokala, Niloy Ganguly, Pawan Goyal 0002 |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Free as in Free Word Order: An Energy Based Model for Word Segmentation and Morphological Tagging in SanskritabstractAmrith Krishna, Bishal Santra, Sasi Prasanth Bandaru, Gaurav Sahu, Vishnu Dutt Sharma, Pavankumar Satuluri, Pawan Goyal. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Amrith Krishna, Bishal Santra, Sasi Prasanth Bandaru, Gaurav Sahu, Vishnu Dutt Sharma, Pavankumar Satuluri, Pawan Goyal 0002 |
EMNLP | 2 |
| 2016 | Word Segmentation in Sanskrit Using Path Constrained Random WalksabstractIn Sanskrit, the phonemes at the word boundaries undergo changes to form new phonemes through a process called as sandhi. A fused sentence can be segmented into multiple possible segmentations. We propose a word segmentation approach that predicts the most semantically valid segmentation for a given sentence. We treat the problem as a query expansion problem and use the path-constrained random walks framework to predict the correct segments. Amrith Krishna, Bishal Santra, Pavankumar Satuluri, Sasi Prasanth Bandaru, Bhumi Faldu, Yajuvendra Singh, Pawan Goyal 0002 |
COLING | 2 |