Vinod K. Kurmi

dblp:222/1567 · also Vinod Kumar Kurmi · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0003-0332-5647ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dense Frame Annotations for Low-Resource ISL Fingerspelling Recognition
Kirandevraj R, Vinod K. Kurmi, Vinay P. Namboodiri, C. V. Jawahar
ICPR (14)2
2026 Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization
abstract
We present HAQAGen, a unified generative model for resolution-invariant NIR-to-RGB colorization that balances chromatic realism with structural fidelity. The proposed model introduces (i) a combined loss term aligning the global color statistics through differentiable histogram matching, perceptual image quality measure, and feature-based similarity to preserve texture information, (ii) local hue–saturation priors injected via Spatially Adaptive Denormalization (SPADE) to stabilize chromatic reconstruction, and (iii) texture-aware supervision within a Mamba backbone to preserve fine details. We introduce an adaptive-resolution inference engine that further enables high-resolution translation without sacrificing quality. Our proposed NIR-to-RGB translation model simultaneously enforces global color statistics and local chromatic consistency, while scaling to native resolutions without compromising texture fidelity or generalization. Extensive evaluations on FANVID, OMSIV, VCIP2020, and RGB2NIR using different evaluation metrics demonstrate consistent improvements over state-of-the-art baseline methods. HAQAGen produces images with sharper textures, natural colors, attaining significant gains as per perceptual metrics. These results position HAQAGen as a scalable and effective solution for NIR-to-RGB translation across diverse imaging scenarios.
Abhinav Attri, Rajeev Ranjan Dwivedi, Samiran Das, Vinod K. Kurmi
WACV4
2025 Multi-Attribute Bias Mitigation via Representation Learning
abstract
Real-world images frequently exhibit multiple overlapping biases, including textures, watermarks, gendered makeup, scene-object pairings, etc. These biases collectively impair the performance of modern vision models, undermining both their robustness and fairness. Addressing these biases individually proves inadequate, as mitigating one bias often permits or intensifies others. We tackle this multi-bias problem with Generalized Multi-Bias Mitigation (GMBM), a lean two-stage framework that needs group labels only while training and minimizes bias at test time. First, Adaptive Bias-Integrated Learning (ABIL) deliberately identifies the influence of known shortcuts by training encoders for each attribute and integrating them with the main backbone, compelling the classifier to explicitly recognize these biases. Then Gradient-Suppression Fine-Tuning prunes those very bias directions from the backbone’s gradients, leaving a single compact network that ignores all the shortcuts it just learned to recognize. Moreover we find that existing bias metrics break under subgroup imbalance and train–test distribution shifts, so we introduce Scaled Bias Amplification (SBA): a test-time measure that disentangles model-induced bias amplification from distributional differences. We validate GMBM on FB-CMNIST, CelebA, and COCO, where we boost worst-group accuracy, halve multi-attribute bias amplification, and set a new low in SBA—even as bias complexity and distribution shifts intensify—making GMBM the first practical, end-to-end multi-bias solution for visual recognition. Project page: https://visdomlab.github.io/GMBM/
Rajeev Ranjan Dwivedi, Vinod K. Kurmi
ECAI3
2025 Multilingual Query-by-Example KWS for Indian Languages using Transliteration
Kirandevraj R, Vinod K. Kurmi, Vinay P. Namboodiri, C. V. Jawahar
INTERSPEECH2
2025 Spectral-Temporal Attention for Robust Change Detection
abstract
Change detection has long been used for various tasks. With advancements in robotic systems and computer vision, change detection techniques can be further explored for diverse applications. Current state-of-the-art methods primarily use either satellite images or street-level images to detect changes. However, the techniques used for these two types of images differ substantially, though their core objective remains identical.We introduce a spectral-temporal attention network capable of detecting changes in both satellite and street-level images across multiple temporal instances. Additionally, we present an indoor environmental dataset featuring significantly more frequent changes. We analyze the impact of temporal and spatial domain shifts on the performance of various methods and demonstrate that performing attention in the spectral domain not only enhances overall performance but also increases robustness against spatial domain shifts.
Mayank Thakur, Radhe Shyam Sharma, Vinod K. Kurmi, Raj Samant, Badri Narayana Patro
IROS3
2025 Strategic Base Representation Learning via Feature Augmentations for Few-Shot Class Incremental Learning
abstract
Few-shot class incremental learning implies the model to learn new classes while retaining knowledge of previously learned classes with a small number of training instances. Existing frameworks typically freeze the parameters of the previously learned classes during the incorporation of new classes. However, this approach often results in suboptimal class separation of previously learned classes, leading to overlap between old and new classes. Consequently, the performance of old classes degrades on new classes. To address these challenges, we propose a novel feature augmentation driven contrastive learning framework designed to enhance the separation of previously learned classes to accommodate new classes. Our approach involves augmenting feature vectors and assigning proxy labels to these vectors. This strategy expands the feature space, ensuring seamless integration of new classes within the expanded space. Additionally, we employ a self-supervised contrastive loss to improve the separation between previous classes. We validate our framework through experiments on three FSCIL benchmark datasets: CIFAR100, miniImageNet, and CUB200. The results demonstrate that our Feature Augmentation driven Contrastive Learning framework significantly outperforms other approaches, achieving state-of-the-art performance. Code is available at https://visdomlab.github.io/FeatAugFSCIL/.
Parinita Nema, Vinod K. Kurmi
WACV2
2025 Label Calibration in Source Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) utilizes a pretrained source model with unlabeled target data. Self-supervised SFDA techniques generate pseudolabels from the pretrained source model, but these pseudolabels often contain noise due to domain discrepancies between the source and target domains. Traditional self-supervised SFDA techniques rely on deterministic model predictions using the softmax function, leading to unreliable pseudolabels. In this work, we propose to introduce predictive uncertainty and softmax calibration for pseudolabel refinement using evidential deep learning. The Dirichlet prior is placed over the output of the target network to capture uncertainty using evidence with a single forward pass. Furthermore, softmax calibration solves the translation invariance problem to assist in learning with noisy labels. We incorporate a combination of evidential deep learning loss and information maximization loss with calibrated softmax in both prior and non-prior target knowledge SFDA settings. Extensive experimental analysis shows that our method outperforms other state-of-the-art methods on benchmark datasets. The code is available at https://visdomlab.github.io/EKS/.
Shivangi Rai, Rini Smita Thakur, Kunal Jangid, Vinod K. Kurmi
WACV4
2025 Uncertainty and Energy based Loss Guided Semi-Supervised Semantic Segmentation
abstract
Semi-supervised (SS) semantic segmentation exploits both labeled and unlabeled images to overcome tedious and costly pixel-level annotation problems. Pseudolabel supervision is one of the core approaches of training net-works with both pseudo labels and ground-truth labels. This work uses aleatoric or data uncertainty and energy based modeling in intersection-union pseudo supervised network. The aleatoric uncertainty is modeling the inherent noise variations of the data in a network with two predic-tive branches. The per-pixel variance parameter obtained from the network gives a quantitative idea about the data uncertainty. Moreover, energy-based loss realizes the potential of generative modeling on the downstream SS seg-mentation task. The aleatoric and energy loss are applied in conjunction with pseudo-intersection labels, pseudo-union labels, and ground-truth on the respective network branch. The comparative analysis with state-of-the-art methods has shown improvement in performance metrics. The code is availaible at https://visdomlab.github.io/DUEB/.
Rini Smita Thakur, Vinod K. Kurmi
WACV2
2024 CosFairNet: A Parameter-Space based Approach for Bias Free Learning
Rajeev Ranjan Dwivedi, Priyadarshini Kumari, Vinod K. Kurmi
BMVC3
2024 Quantifying Uncertainty in Neural Networks through Residuals
Dalavai Udbhav Mallanna, Rini Smita Thakur, Rajeev Ranjan Dwivedi, Vinod K. Kurmi
CIKM4
2024 Towards Robust Few-shot Class Incremental Learning in Audio Classification using Contrastive Representation
Riyansha Singh, Parinita Nema, Vinod K. Kurmi
INTERSPEECH3
2023 Distilling knowledge for occlusion robust monocular 3D face reconstruction
Hitika Tiwari, Vinod K. Kurmi, K. S. Venkatesh, Yong-Sheng Chen
Image Vis. Comput.2
2022 Gradient Based Activations for Accurate Bias-Free Learning
abstract
Bias mitigation in machine learning models is imperative, yet challenging. While several approaches have been proposed, one view towards mitigating bias is through adversarial learning. A discriminator is used to identify the bias attributes such as gender, age or race in question. This discriminator is used adversarially to ensure that it cannot distinguish the bias attributes. The main drawback in such a model is that it directly introduces a trade-off with accuracy as the features that the discriminator deems to be sensitive for discrimination of bias could be correlated with classification. In this work we solve the problem. We show that a biased discriminator can actually be used to improve this bias-accuracy tradeoff. Specifically, this is achieved by using a feature masking approach using the discriminator's gradients. We ensure that the features favoured for the bias discrimination are de-emphasized and the unbiased features are enhanced during classification. We show that this simple approach works well to reduce bias as well as improve accuracy significantly. We evaluate the proposed model on standard benchmarks. We improve the accuracy of the adversarial methods while maintaining or even improving the unbiasness and also outperform several other recent methods.
Vinod K. Kurmi, Yash Vardhan Sharma, Vinay P. Namboodiri
AAAI1
2022 Generalized Keyword Spotting using ASR embeddings
Kirandevraj R, Vinod K. Kurmi, Vinay P. Namboodiri, C. V. Jawahar
INTERSPEECH2
2022 Occlusion Resistant Network for 3D Face Reconstruction
abstract
3D face reconstruction from a monocular face image is a mathematically ill-posed problem. Recently, we observed a surge of interest in deep learning-based approaches to address the issue. These methods possess extreme sensitivity towards occlusions. Thus, in this paper, we present a novel context-learning-based distillation approach to tackle the occlusions in the face images. Our training pipeline focuses on distilling the knowledge from a pre-trained occlusion-sensitive deep network. The proposed model learns the context of the target occluded face image. Hence our approach uses a weak model (unsuitable for occluded face images) to train a highly robust network towards partially and fully-occluded face images. We obtain a landmark accuracy of 0.77 against 5.84 of recent state-of-the-art-method for real-life challenging facial occlusions. Also, we propose a novel end-to-end training pipeline to reconstruct 3D faces from multiple variations of the target image per identity to emphasize the significance of visible facial features during learning. For this purpose, we leverage a novel composite multi-occlusion loss function. Our multi-occlusion per identity model shows a dip in the landmark error by a large margin of 6.67 in comparison to a recent state-of-the-art method. We deploy the occluded variations of the CelebA validation dataset and AFLW2000-3D face dataset: naturally-occluded and artificially occluded, for the comparisons. We comprehensively compare our results with the other approaches concerning the accuracy of the reconstructed 3D face mesh for occluded face images.
Hitika Tiwari, Vinod K. Kurmi, K. S. Venkatesh, Yong-Sheng Chen
WACV2
2021 Collaborative Learning to Generate Audio-Video Jointly
abstract
There have been a number of techniques that have demonstrated the generation of multimedia data for one modality at a time using GANs, such as the ability to generate images, videos, and audio. However, so far, the task of multi-modal generation of data, specifically for audio and videos both, has not been sufficiently well-explored. Towards this, we propose a method that demonstrates that we are able to generate naturalistic samples of video and audio data by the joint correlated generation of audio and video modalities. The proposed method uses multiple discriminators to ensure that the audio, video, and the joint output are also indistinguishable from real-world samples. We present a dataset for this task and show that we are able to generate realistic samples. This method is validated using various standard metrics such as Inception Score, Frechet Inception Distance (FID) and through human evaluation.
Vinod K. Kurmi, Vipul Bajaj, Badri Narayana Patro, K. S. Venkatesh, Vinay P. Namboodiri, Preethi Jyothi
ICASSP1
2021 Data Uncertainty Guided Noise-aware Preprocessing Of Fingerprints
abstract
The effectiveness of fingerprint-based authentication systems on good quality fingerprints is established long back. However, the performance of standard fingerprint matching systems on noisy and poor quality fingerprints is far from satisfactory. Towards this, we propose a data uncertainty-based framework which enables the state-of-the-art fingerprint pre-processing models to quantify noise present in the input image and identify fingerprint regions with background noise and poor ridge clarity. Quantification of noise helps the model two folds: firstly, it makes the objective function adaptive to the noise in a particular input fingerprint and consequently, helps to achieve robust performance on noisy and distorted fingerprint regions. Secondly, it provides a noise variance map which indicates noisy pixels in the input fingerprint image. The predicted noise variance map enables the end-users to understand erroneous predictions due to noise present in the input image. Extensive experimental evaluation on 13 publicly available fingerprint databases, across different architectural choices and two fingerprint processing tasks demonstrate effectiveness of the proposed framework.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN4
2021 Sensor-invariant Fingerprint ROI Segmentation Using Recurrent Adversarial Learning
abstract
A fingerprint region of interest (roi) segmentation algorithm is designed to separate the foreground fingerprint from the background noise. All the learning based state-of-the-art fingerprint roi segmentation algorithms proposed in the literature are benchmarked on scenarios when both training and testing databases consist of fingerprint images acquired from the same sensors. However, when testing is conducted on a different sensor, the segmentation performance obtained is often unsatisfactory. As a result, every time a new fingerprint sensor is used for testing, the fingerprint roi segmentation model needs to be re-trained with the fingerprint image acquired from the new sensor and its corresponding manually marked ROI. Manually marking fingerprint ROI is expensive because firstly, it is time consuming and more importantly, requires domain expertise. In order to save the human effort in generating annotations required by state-of-the-art, we propose a fingerprint roi segmentation model which aligns the features of fingerprint images derived from the unseen sensor such that they are similar to the ones obtained from the fingerprints whose ground truth roi masks are available for training. Specifically, we propose a recurrent adversarial learning based feature alignment network that helps the fingerprint roi segmentation model to learn sensor-invariant features. Consequently, sensor-invariant features learnt by the proposed roi segmentation model help it to achieve improved segmentation performance on fingerprints acquired from the new sensor. Experiments on publicly available FVC databases demonstrate the efficacy of the proposed work.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN4
2021 Do not Forget to Attend to Uncertainty while Mitigating Catastrophic Forgetting
abstract
One of the major limitations of deep learning models is that they face catastrophic forgetting in an incremental learning scenario. There have been several approaches proposed to tackle the problem of incremental learning. Most of these methods are based on knowledge distillation and do not adequately utilize the information provided by older task models, such as uncertainty estimation in pre-dictions. The predictive uncertainty provides the distributional information can be applied to mitigate catastrophic forgetting in a deep learning framework. In the proposed work, we consider a Bayesian formulation to obtain the data and model uncertainties. We also incorporate self-attention framework to address the incremental learning problem. We define distillation losses in terms of aleatoric uncertainty and self-attention. In the proposed work, we investigate different ablation analyses on these losses. Furthermore, we are able to obtain better results in terms of accuracy on standard benchmarks.
Vinod K. Kurmi, Badri Narayana Patro, K. S. Venkatesh, Vinay P. Namboodiri
WACV1
2021 Domain Impression: A Source Data Free Domain Adaptation Method
abstract
Unsupervised Domain adaptation methods solve the adaptation problem for an unlabeled target set, assuming that the source dataset is available with all labels. However, the availability of actual source samples is not always possible in practical cases. It could be due to memory constraints, privacy concerns, and challenges in sharing data. This practical scenario creates a bottleneck in the domain adaptation problem. This paper addresses this challenging scenario by proposing a domain adaptation technique that does not need any source data. Instead of the source data, we are only provided with a classifier that is trained on the source data. Our proposed approach is based on a generative framework, where the trained classifier is used for generating samples from the source classes. We learn the joint distribution of data by using the energy-based modeling of the trained classifier. At the same time, a new classifier is also adapted for the target domain. We perform various ablation analysis under different experimental setups and demonstrate that the proposed approach achieves better results than the baseline models in this extremely novel scenario.
Vinod K. Kurmi, K. S. Venkatesh, Vinay P. Namboodiri
WACV1
2021 Exploring dropout discriminator for domain adaptation
Vinod K. Kurmi, K. S. Venkatesh, Vinay P. Namboodiri
Neurocomputing1
2021 Revisiting paraphrase question generator using pairwise discriminator
Badri Narayana Patro, Dev Chauhan, Vinod K. Kurmi, Vinay P. Namboodiri
Neurocomputing3
2021 Informative discriminator for domain adaptation
Vinod K. Kurmi, K. S. Venkatesh, Vinay P. Namboodiri
Image Vis. Comput.1
2021 MUMC: Minimizing uncertainty of mixture of cues
Badri Narayana Patro, Vinod K. Kurmi, Vinay P. Namboodiri
Image Vis. Comput.2
2020 Deep Bayesian Network for Visual Question Generation
abstract
Generating natural questions from an image is a semantic task that requires using vision and language modalities to learn multimodal representations. Images can have multiple visual and language cues such as places, captions, and tags. In this paper, we propose a principled deep Bayesian learning framework that combines these cues to produce natural questions. We observe that with the addition of more cues and by minimizing uncertainty in the among cues, the Bayesian network becomes more confident. We propose a Minimizing Uncertainty of Mixture of Cues (MUMC), that minimizes uncertainty present in a mixture of cues experts for generating probabilistic questions. This is a Bayesian framework and the results show a remarkable similarity to natural questions as validated by a human study. We observe that with the addition of more cues and by minimizing uncertainty among the cues, the Bayesian framework becomes more confident. Ablation studies of our model indicate that a subset of cues is inferior at this task and hence the principled fusion of cues is preferred. Further, we observe that the proposed approach substantially improves over state-of-the-art benchmarks on the quantitative metrics (BLEU-n, METEOR, ROUGE, and CIDEr). Here we provide project link for Deep Bayesian VQG https: //delta-lab-iitk.github.io/BVQG/.
Badri Narayana Patro, Vinod K. Kurmi, Vinay P. Namboodiri
WACV2
2019 Curriculum based Dropout Discriminator for Domain Adaptation
Vinod K. Kurmi, Vipul Bajaj, K. S. Venkatesh, Vinay P. Namboodiri
BMVC1
2019 Attending to Discriminative Certainty for Domain Adaptation
abstract
In this paper, we aim to solve for unsupervised domain adaptation of classifiers where we have access to label information for the source domain while these are not available for a target domain. While various methods have been proposed for solving these including adversarial discriminator based methods, most approaches have focused on the entire image based domain adaptation. In an image, there would be regions that can be adapted better, for instance, the foreground object may be similar in nature. To obtain such regions, we propose methods that consider the probabilistic certainty estimate of various regions and specific focus on these during classification for adaptation. We observe that just by incorporating the probabilistic certainty of the discriminator while training the classifier, we are able to obtain state of the art results on various datasets as compared against all the recent methods. We provide a thorough empirical analysis of the method by providing ablation analysis, statistical significance test, and visualization of the attention maps and t-SNE embeddings. These evaluations convincingly demonstrate the effectiveness of the proposed approach.
Vinod K. Kurmi, Shanu Kumar, Vinay P. Namboodiri
CVPR1
2019 Looking back at Labels: A Class based Domain Adaptation Technique
abstract
In this paper, we solve the problem of adapting classifiers across domains. We consider the problem of domain adaptation for multi-class classification where we are provided a labeled set of examples in a source dataset and we are provided a target dataset with no supervision. In this setting, we propose an adversarial discriminator based approach. While the approach based on adversarial discriminator has been previously proposed; in this paper, we present an informed adversarial discriminator. Our observation relies on the analysis that shows that if the discriminator has access to all the information available including the class structure present in the source dataset, then it can guide the transformation of features of the target set of classes to a more structure adapted space. Using this formulation, we obtain state-of-the-art results for the standard evaluation on benchmark datasets. We further provide detailed analysis which shows that using all the labeled information results in an improved domain adaptation.
Vinod K. Kurmi, Vinay P. Namboodiri
IJCNN1
2018 Learning Semantic Sentence Embeddings using Sequential Pair-wise Discriminator
Badri Narayana Patro, Vinod K. Kurmi, Vinay P. Namboodiri
COLING2
2018 Multimodal Differential Network for Visual Question Generation
abstract
Generating natural questions from an image is a semantic task that requires using visual and language modality to learn multimodal representations. Images can have multiple visual and language contexts that are relevant for generating questions namely places, captions, and tags. In this paper, we propose the use of exemplars for obtaining the relevant context. We obtain this by using a Multimodal Differential Network to produce natural and engaging questions. The generated questions show a remarkable similarity to the natural questions as validated by a human study. Further, we observe that the proposed approach substantially improves over state-of-the-art benchmarks on the quantitative metrics (BLEU, METEOR, ROUGE, and CIDEr).
Badri Narayana Patro, Vinod K. Kurmi, Vinay P. Namboodiri
EMNLP3