VLDB 2026 Research / reviewers in the wild / expert
Rohit Saluja
dblp:213/8378
· DBLP profile ↗
16ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-0773-3480ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiFOLD: A Multimodal Framework to Correct OCR Lapses in Cluttered Documents
Rajat Verma, Vriti Sharma, Manikandan Ravikiran, Rohit Saluja |
ICDAR (2) | 4 |
| 2025 | GISA: Gradual Information Selection Attention for MCQ Difficulty Estimation
Manikandan Ravikiran, Arnav Bhavsar, Rohit Saluja |
AIED (1) | 4 |
| 2025 | Can Language Models Verify Indian Classical Music Note Sequences for Early Learners?
Radhika Grover, Ankit Maurya, Manikandan Ravikiran, Rohit Saluja |
IEEE Big Data | 4 |
| 2025 | Prompted to Fly: Translating Free-Form Instructions into Schema-Constrained Mission Generation for UAVs Using LLMs
Manikandan Ravikiran, Sohom Chakrabarty, Rohit Saluja, Mrunmayee Limaye, Harshal Kolhe, Samiron |
IEEE Big Data | 4 |
| 2025 | TEEMIL : Towards Educational MCQ Difficulty Estimation in Indic LanguagesabstractDifficulty estimation of multiple-choice questions (MCQs) is crucial for creating effective educational assessments, yet remains underexplored in Indic languages like Hindi and Kannada due to the lack of comprehensive datasets. This paper addresses this gap by introducing two datasets, TEEMIL-H and TEEMIL-K, containing 4689 and 4215 MCQs, respectively, with manually annotated difficulty labels. We benchmark these datasets using state-of-the-art multilingual models and conduct ablation studies to analyze the effect of context, the impact of options, and the presence of the None of the Above (NOTA) option on difficulty estimation. Our findings establish baselines for difficulty estimation in Hindi and Kannada, offering valuable insights into improving model performance and guiding future research in MCQ difficulty estimation . Manikandan Ravikiran, Siddharth Vohra, Rajat Verma, Rohit Saluja, Arnav Bhavsar |
COLING | 4 |
| 2025 | Towards Scene Text Recognition in Rainy Weather Conditions
Anandita Jamwal, Lalithya Koneti, Manikandan Ravikiran, Dinesh Singh 0001, Rohit Saluja |
ICDAR (5) | 5 |
| 2024 | IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured TrafficabstractIntelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle’s driving behavior for safe and efficient navigation. This is particularly critical in developing countries where traffic situations are often dense and unstructured with heterogeneous road occupants. Existing datasets, predominantly geared towards structured and sparse traffic scenarios, fall short of capturing the complexity of driving in such environments. To fill this gap, we present IDD-X, a large-scale dual-view driving video dataset. With 697K bounding boxes, 9K important object tracks, and 1-12 objects per video, IDD-X offers comprehensive ego-relative annotations for multiple important road objects covering 10 categories and 19 explanation label categories. The dataset also incorporates rearview information to provide a more complete representation of the driving environment. We also introduce custom-designed deep networks aimed at multiple important object localization and per-object explanation prediction. Overall, our dataset and introduced prediction models form the foundation for studying how road conditions and surrounding entities affect driving behavior in complex traffic situations. Chirag Parikh, Rohit Saluja, C. V. Jawahar, Ravi Kiran Sarvadevabhatla |
ICRA | 2 |
| 2023 | CueCAn: Cue-driven Contextual Attention for Identifying Missing Traffic Signs on Unconstrained RoadsabstractUnconstrained Asian roads often involve poor infrastructure, affecting overall road safety. Missing traffic signs are a regular part of such roads. Missing or non-existing object detection has been studied for locating missing curbs and estimating reasonable regions for pedestrians on road scene images. Such methods involve analyzing task-specific single object cues. In this paper, we present the first and most challenging video dataset for missing objects, with multiple types of traffic signs for which the cues are visible without the signs in the scenes. We refer to it as the Missing Traffic Signs Video Dataset (MTSVD). MTSVD is challenging compared to the previous works in two aspects i) The traffic signs are generally not present in the vicinity of their cues, ii) The traffic signs' cues are diverse and unique. Also, MTSVD is the first publicly available missing object dataset. To train the models for identifying missing signs, we complement our dataset with 10K traffic sign tracks, with 40% of the traffic signs having cues visible in the scenes. For identifying missing signs, we propose the Cue-driven Contextual Attention units (CueCAn), which we incorporate in our model's encoder. We first train the encoder to classify the presence of traffic sign cues and then train the entire segmentation model end-to-end to localize missing traffic signs. Quantitative and qualitative analysis shows that CueCAn significantly improves the performance of base models. Varun Gupta 0007, Anbumani Subramanian, C. V. Jawahar, Rohit Saluja |
ICRA | 4 |
| 2022 | New Objects on the Road? No Problem, We'll Learn Them TooabstractObject detection plays an essential role in providing localization, path planning, and decision making capabilities in autonomous navigation systems. However, existing object detection models are trained and tested on a fixed number of known classes. This setting makes the object detection model difficult to generalize well in real-world road scenarios while encountering an unknown object. We address this problem by introducing our framework that handles the issue of unknown object detection and updates the model when unknown object labels are available. Next, our solution includes three major components that address the inherent problems present in the road scene datasets. The novel components are a) Feature-Mix that improves the unknown object detection by widening the gap between known and unknown classes in latent feature space, b) Focal regression loss handling the problem of improving small object detection and intra-class scale variation, and c) Curriculum learning further enhances the detection of small objects. We use Indian Driving Dataset (IDD) and Berkeley Deep Drive (BDD) dataset for evaluation. Our solution provides state-of-the-art performance on open-world evaluation metrics. We hope this work will create new directions for open-world object detection for road scenes, making it more reliable and robust autonomous systems. Shyam Nandan Rai, K. J. Joseph, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar |
IROS | 4 |
| 2022 | Multi-Domain Incremental Learning for Semantic SegmentationabstractRecent efforts in multi-domain learning for semantic segmentation attempt to learn multiple geographical datasets in a universal, joint model. A simple fine-tuning experiment performed sequentially on three popular road scene segmentation datasets demonstrates that existing segmentation frameworks fail at incrementally learning on a series of visually disparate geographical domains. When learning a new domain, the model catastrophically forgets previously learned knowledge. In this work, we pose the problem of multi-domain incremental learning for semantic segmentation. Given a model trained on a particular geographical domain, the goal is to (i) incrementally learn a new geographical domain, (ii) while retaining performance on the old domain, (iii) given that the previous domain’s dataset is not accessible. We propose a dynamic architecture that assigns universally shared, domain-invariant parameters to capture homogeneous semantic features present in all domains, while dedicated domain-specific parameters learn the statistics of each domain. Our novel optimization strategy helps achieve a good balance between retention of old knowledge (stability) and acquiring new knowledge (plasticity). We demonstrate the effectiveness of our proposed solution on domain incremental settings pertaining to real-world driving scenes from roads of Germany (Cityscapes), the United States (BDD100k), and India (IDD).1 Prachi Garg, Rohit Saluja, Vineeth N. Balasubramanian, Chetan Arora 0001, Anbumani Subramanian, C. V. Jawahar |
WACV | 2 |
| 2022 | To miss-attend is to misalign! Residual Self-Attentive Feature Alignment for Adapting Object DetectorsabstractAdvancements in adaptive object detection can lead to tremendous improvements in applications like autonomous navigation, as they alleviate the distributional shifts along the detection pipeline. Prior works adopt adversarial learning to align image features at global and local levels, yet the instance-specific misalignment persists. Also, adaptive object detection remains challenging due to visual diversity in background scenes and intricate combinations of objects. Motivated by structural importance, we aim to attend prominent instance-specific regions, overcoming the feature misalignment issue. We propose a novel resIduaL seLf-attentive featUre alignMEnt (ILLUME) method for adaptive object detection. ILLUME comprises Self-Attention Feature Map (SAFM) module that enhances structural attention to object-related regions and thereby generates domain invariant features. Our approach significantly reduces the domain distance with the improved feature alignment of the instances. Qualitative results demonstrate the ability of ILLUME to attend important object instances required for alignment. Experimental results on several benchmark datasets show that our method outperforms the existing state-of-the-art approaches. Vaishnavi Khindkar, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, Rohit Saluja, C. V. Jawahar |
WACV | 5 |
| 2022 | FLUID: Few-Shot Self-Supervised Image DerainingabstractSelf-supervised methods have shown promising results in denoising and dehazing tasks, where the collection of the paired dataset is challenging and expensive. However, we find that these methods fail to remove the rain streaks when applied for image deraining tasks. The method’s poor performance is due to the explicit assumptions: (i) the distribution of noise or haze is uniform and (ii) the value of a noisy or hazy pixel is independent of its neighbors. The rainy pixels are non-uniformly distributed, and it is not necessarily dependant on its neighboring pixels. Hence, we conclude that the self-supervised method needs to have some prior knowledge about rain distribution to perform the deraining task. To provide this knowledge, we hypothesize a network trained with minimal supervision to estimate the likelihood of rainy pixels. This leads us to our proposed method called FLUID: Few Shot Sel f-Supervised Image Deraining.We perform extensive experiments and comparisons with existing image deraining and few-shot image-to-image translation methods on Rain 100L and DDN-SIRR datasets containing real and synthetic rainy images. In addition, we use the Rainy Cityscapes dataset to show that our method trained in a few-shot setting can improve semantic segmentation and object detection in rainy conditions. Our approach obtains a mIoU gain of 51.20 over the current best-performing deraining method. [Project Page] Shyam Nandan Rai, Rohit Saluja, Chetan Arora 0001, Vineeth N. Balasubramanian, Anbumani Subramanian, C. V. Jawahar |
WACV | 2 |
| 2020 | Leaf Counting in Rice (Oryza Sativa L.) Using Object Detection: A Deep Learning ApproachabstractLeaf count is one of the crucial tasks in plant phenotyping, and leaves are the basic unit of plant architecture involved in photosynthesis, growth, and yield of a plant. Therefore, the total number of leaves per plant is considered as one of the essential physio-morphological plant traits for phenotyping. The current work proposes to estimate the total number of leaves of a rice plant by detecting their leaves tips. A rice plant has a single tip for a single leaf. Hence, this proposed framework counts the total number of leaves by counting the number of leaves tips equal to the number of leaves. You Only Look Once (YOLO) algorithm is used for the detection of the leaves tips as an object. This hypothesis builds a basis for counting the total number of leaves in a plant like rice, and similar field crops such as wheat (Triticum aestivum L), maize (Zea mays L.), sorghum (Sorghum bicolor), barley (Hordeum vulgare L.). The model detected leaves of a rice plant (RGB images) by detecting corresponding leaves tips with YOLO having average accuracy up to 82% and IOU around 0.53-0.60 and estimates the number of leaves in a plant by counting predicted bounding boxes around tips. The model also performed well with the wheat crop. Mukesh Kumar Vishal, Biplab Banerjee, Rohit Saluja, Raju Dhandapani, Viswanathan Chinnusamy, Sudhir Kumar 0003, Rabi N. Sahoo, J. Adinarayana |
IGARSS | 3 |
| 2019 | OCR On-the-Go: Robust End-to-End Systems for Reading License Plates & Street SignsabstractWe work on the problem of recognizing license plates and street signs automatically in challenging conditions such as chaotic traffic. We leverage state-of-the-art text spotters to generate a large amount of noisy labeled training data. The data is filtered using a pattern derived from domain knowledge. We augment training and testing data with interpolated boxes and annotations that makes our training and testing robust. We further use synthetic data during training to increase the coverage of the training data. We train two different models for recognition. Our baseline is a conventional Convolution Neural Network (CNN) encoder followed by a Recurrent Neural Network (RNN) decoder. As our first contribution, we bypass the detection phase by augmenting the baseline with an Attention mechanism in the RNN decoder. Next, we build in the capability of training the model end-to-end on scenes containing license plates by incorporating inception based CNN encoder that makes the model robust to multiple scales. We achieve improvements of 3.75% at the sequence level, over the baseline model. We present the first results of using multi-headed attention models on text recognition in images and illustrate the advantages of using multiple-heads over a single head. We observe gains as large as 7.18% by incorporating multi-headed attention. We also experiment with multi-headed attention models on French Street Name Signs dataset (FSNS) and a new Indian Street dataset that we release for experiments. We observe that such models with multiple attention masks perform better than the model with single-headed attention on three different datasets with varying complexities. Our models outperform state-of-the-art methods on FSNS and IIIT-ILST Devanagari datasets by 1.1% and 8.19% respectively. Rohit Saluja, Ayush Maheshwari, Ganesh Ramakrishnan, Parag Chaudhuri, Mark J. Carman |
ICDAR | 1 |
| 2019 | Sub-Word Embeddings for OCR Corrections in Highly Fusional Indic LanguagesabstractTexts in Indic Languages contain a large proportion of out-of-vocabulary (OOV) words due to frequent fusion using conjoining rules (of which there are around 4000 in Sanskrit). OCR errors further accentuate this complexity for the error correction systems. Variations of sub-word units such as n-grams, possibly encapsulating the context, can be extracted from the OCR text as well as the language text individually. Some of the sub-word units that are derived from the texts in such languages highly correlate to the word conjoining rules. Signals such as frequency values (on a corpus) associated with such sub-word units have been used previously with log-linear classifiers for detecting errors in Indic OCR texts. We explore two different encodings to capture such signals and augment the input to Long Short Term Memory (LSTM) based OCR correction models, that have proven useful in the past for jointly learning the language as well as OCR-specific confusions. The first type of encoding makes direct use of sub-word unit frequency values, derived from the training data. The formulation results in faster convergence and better accuracy values of the error correction model on four different languages with varying complexities. The second type of encoding makes use of trainable sub-word embeddings. We introduce a new procedure for training fastText embeddings on the sub-word units and further observe a large gain in F-Scores, as well as word-level accuracy values. Rohit Saluja, Mayur Punjabi, Mark J. Carman, Ganesh Ramakrishnan, Parag Chaudhuri |
ICDAR | 1 |
| 2017 | Error Detection and Corrections in Indic OCR Using LSTMsabstractConventional approaches to spell checking suggest spelling corrections using proximity-based matches to a known vocabulary. For highly inflectional Indian languages, any off-the-shelf vocabulary is significantly incomplete, since a large fraction of words in Indic documents are generated using word conjoining rules. Therefore, a tremendous manual effort is needed in spell-correcting words in Indic OCR documents. Moreover, in a spell checking system, a vocabulary may suggest multiple alternatives to the incorrect word. The ranking of these corrective suggestions is improved using language models. Owing to corpus resource scarcity, however, Indian languages lack reliable language models. Thus, learning the character (or n-gram) confusions or error patterns of the OCR system can be helpful in correcting the Out of Vocabulary (OOV) words in OCR documents. We adopt a Long Short-Term Memory (LSTM) based character level language model with a fixed delay for discriminative language modeling in the context of OCR errors for jointly addressing the problems of error detection and correction in Indic OCR. For words that need not be corrected in the OCR output, our model simply abstains from suggesting any changes. We present extensive results to validate the performance of our model on four Indian languages with different inflectional complexities. We achieve F-Scores above 92.4% and decreases in Word Error Rates (WER) of at least 26.7% across the four languages. Rohit Saluja, Devaraj Adiga, Parag Chaudhuri, Ganesh Ramakrishnan, Mark J. Carman |
ICDAR | 1 |