VLDB 2026 Research / reviewers in the wild / expert
Pankaj Wasnik
dblp:184/0627 · also Pankaj Shivdayal Wasnik
· DBLP profile ↗
18ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Listen like a Teacher: Mitigating Whisper Hallucinations Using Adaptive Layer Attention and Knowledge DistillationabstractThe Whisper model, an open-source automatic speech recognition system, is widely adopted for its strong performance across multilingual and zero-shot settings. However, it frequently suffers from hallucination errors, especially under noisy acoustic conditions. Previous works to reduce hallucinations in Whisper-style ASR systems have primarily focused on audio preprocessing or post-processing of transcriptions to filter out erroneous content. However, modifications to the Whisper model itself remain largely unexplored to mitigate hallucinations directly. To address this challenge, we present a two-stage architecture that first enhances encoder robustness through Adaptive Layer Attention (ALA) and further suppresses hallucinations using a multi-objective knowledge distillation (KD) framework. In the first stage, ALA groups encoder layers into semantically coherent blocks via inter-layer correlation analysis. A learnable multi-head attention module then fuses these block representations, enabling the model to jointly exploit low- and high-level features for more robust encoding. In the second stage, our KD framework trains the student model on noisy audio to align its semantic and attention distributions with a teacher model processing clean inputs. Our experiments on noisy speech benchmarks show notable reductions in hallucinations and word error rates, while preserving performance on clean speech. Together, ALA and KD offer a principled strategy to improve Whisper’s reliability under real-world noisy conditions. Kumud Tripathi, Aditya Srinivas Menon, Aman Gaurav, Raj Prakash Gohil, Pankaj Wasnik |
AAAI | 5 |
| 2025 | EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice ConversionabstractThe Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion intensity in the diffusion-based EVC framework to generate precise speech of the target emotion. Traditional approaches control the intensity of an emotional state in the utterance via emotion class probabilities or intensity labels that often lead to inept style manipulations and degradations in quality. On the contrary, we aim to regulate emotion intensity using self-supervised learning-based feature representations and unsupervised directional latent vector modeling (DVM) in the emotional embedding space within a diffusion-based framework. These emotion embeddings can be modified based on the given target emotion intensity and the corresponding direction vector. Furthermore, the updated embeddings can be fused in the reverse diffusion process to generate the speech with the desired emotion and intensity. In summary, this paper aims to achieve high-quality emotional intensity regularization in the diffusion-based EVC framework, which is the first of its kind work. The effectiveness of the proposed method has been shown across state-of-the-art (SOTA) baselines in terms of subjective and objective evaluations for the English and Hindi languages. Ashishkumar Prabhakar Gudmalwar, Ishan D. Biyani, Nirmesh J. Shah, Pankaj Wasnik, Rajiv Ratn Shah |
AAAI | 4 |
| 2025 | Enhancing Entertainment Translation for Indian Languages Using Adaptive Context, Style and LLMsabstractWe address the challenging task of neural machine translation (NMT) in the entertainment domain, where the objective is to automatically translate a given dialogue from a source language content to a target language. This task has various applications, particularly in automatic dubbing, subtitling, and other content localization tasks, enabling source content to reach a wider audience. Traditional NMT systems typically translate individual sentences in isolation, without facilitating knowledge transfer of crucial elements such as the context and style from previously encountered sentences. In this work, we emphasize the significance of these fundamental aspects in producing pertinent and captivating translations. We demonstrate their significance through several examples and propose a novel framework for entertainment translation, which, to our knowledge, is the first of its kind. Furthermore, we introduce an algorithm to estimate the context and style of the current session and use these estimations to generate a prompt that guides a Large Language Model (LLM) to generate high-quality translations. Our method is both language and LLM-agnostic, making it a general-purpose tool. We demonstrate the effectiveness of our algorithm through various numerical studies and observe significant improvement in the COMET scores over various state-of-the-art LLMs. Moreover, our proposed method consistently outperforms baseline LLMs in terms of win-ratio. Pratik Rakesh Singh, Mohammadi Zaki, Pankaj Wasnik |
AAAI | 3 |
| 2025 | Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class ImbalanceabstractPrecise Event Spotting (PES) aims to identify events and their class from long, untrimmed videos, particularly in sports. The main objective of PES is to detect the event at the exact moment it occurs. Existing methods mainly rely on features from a large pre-trained network, which may not be ideal for the task. Furthermore, these methods overlook the issue of imbalanced event class distribution present in the data, negatively impacting performance in challenging scenarios. This paper demonstrates that an appropriately designed network, trained end-to-end, can outperform state-of-the-art (SOTA) methods. Particularly, we propose a network with a convolutional spatial-temporal feature extractor enhanced with our proposed Adaptive Spatio-Temporal Refinement Module (ASTRM) and a long-range temporal module. The ASTRM enhances the features with spatio-temporal information. Meanwhile, the long-range temporal module helps extract global context from the data by modeling long-range dependencies. To address the class imbalance issue, we introduce the Soft Instance Contrastive (Soft-Ic) loss that promotes feature compactness and class separation. Extensive experiments show that the proposed method is efficient and outperforms the SOTA methods, specifically in more challenging settings. Sanchayan Santra, Vishal M. Chudasama, Pankaj Wasnik, Vineeth N. Balasubramanian |
CVPR | 3 |
| 2025 | Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and TokenizationabstractAutomatic speech recognition has recently seen a significant advancement with large foundational models such as Whisper. However, these models often struggle to perform well in low-resource languages, such as Indian languages. This paper explores two novel approaches to enhance Whisper’s multilingual speech recognition performance in Indian languages. First, we propose prompt-tuning with language family information, which enhances Whisper’s accuracy in linguistically similar languages. Second, we introduce a novel tokenizer that reduces the number of generated tokens, thereby accelerating Whisper’s inference speed. Our extensive experiments demonstrate that the tokenizer significantly reduces inference time, while prompt-tuning enhances accuracy across various Whisper model sizes, including Small, Medium, and Large. Together, these techniques achieve a balance between optimal WER and inference speed. Kumud Tripathi, Raj Gothi, Pankaj Wasnik |
ICASSP | 3 |
| 2025 | DuET: Dual Incremental Object Detection via Exemplar-Free Task ArithmeticabstractReal-world object detection systems, such as those in autonomous driving and surveillance, must continuously learn new object categories and simultaneously adapt to changing environmental conditions. Existing approaches, Class Incremental Object Detection (CIOD) and Domain Incremental Object Detection (DIOD) only address one aspect of this challenge. CIOD struggles in unseen domains, while DIOD suffers from catastrophic forgetting when learning new classes, limiting their real-world applicability. To overcome these limitations, we introduce Dual Incremental Object Detection (DuIOD), a more practical setting that simultaneously handles class and domain shifts in an exemplar-free manner. We propose DuET, a Task Arithmetic-based model merging framework that enables stable incremental learning while mitigating sign conflicts through a novel Directional Consistency Loss. Unlike prior methods, DuET is detector-agnostic, allowing models like YOLO11 and RT-DETR to function as real-time incremental object detectors. To comprehensively evaluate both retention and adaptation, we introduce the Retention-Adaptability Index (RAI), which combines the Average Retention Index (Avg RI) for catastrophic forgetting and the Average Generalization Index for domain adaptability into a common ground. Extensive experiments on the Pascal Series and Diverse Weather Series demonstrate DuET's effectiveness, achieving a +13.12% RAI improvement while preserving 89.3% Avg RI on the Pascal Series (4 tasks), as well as a +11.39% RAI improvement with 88.57% Avg RI on the Diverse Weather Series (3 tasks), outperforming existing methods. Munish Monga, Vishal M. Chudasama, Pankaj Wasnik, Biplab Banerjee |
ICCV | 3 |
| 2025 | REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice ConversionabstractSpeech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. In this paper, we propose leveraging speaker representations learned from time reversed speech as an augmentation strategy to enhance speaker representation. Notably, speaker and language disentanglement in voice conversion (VC) is essential to accurately preserve a speaker's unique vocal traits while minimizing interference from linguistic content. The effectiveness of the proposed approach is evaluated in the context of state-of-the-art diffusion-based VC models. Experimental results indicate that the proposed approach significantly improves speaker similarity-related scores while maintaining high speech quality. Ishan D. Biyani, Nirmesh J. Shah, Ashishkumar Prabhakar Gudmalwar, Pankaj Wasnik, Rajiv Ratn Shah |
INTERSPEECH | 4 |
| 2025 | LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
Aditya Srinivas Menon, Raj Prakash Gohil, Kumud Tripathi, Pankaj Wasnik |
INTERSPEECH | 4 |
| 2025 | Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
Kumud Tripathi, Chowdam Venkata Kumar, Pankaj Wasnik |
INTERSPEECH | 3 |
| 2025 | AdaPrefix++: Integrating Adapters, Prefixes and Hypernetwork for Continual Learning
Sayanta Adhikari, Dupati Srikar Chandra, P. K. Srijith, Pankaj Wasnik, Naoyuki Onoe |
WACV | 4 |
| 2024 | VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
Ashishkumar Gudmalwar, Nirmesh J. Shah, Sai Akarsh, Pankaj Wasnik, Rajiv Ratn Shah |
INTERSPEECH | 4 |
| 2024 | DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
Neha Sahipjohn, Ashishkumar Gudmalwar, Nirmesh J. Shah, Pankaj Wasnik, Rajiv Ratn Shah |
INTERSPEECH | 4 |
| 2024 | Open-Set Object Detection By Aligning Known Class RepresentationsabstractOpen-Set Object Detection (OSOD) has emerged as a contemporary research direction to address the detection of unknown objects. Recently, few works have achieved remarkable performance in the OSOD task by employing contrastive clustering to separate unknown classes. In contrast, we propose a new semantic clustering-based approach to facilitate a meaningful alignment of clusters in semantic space and introduce a class decorrelation module to enhance inter-cluster separation. Our approach further incorporates an object focus module to predict objectness scores, which enhances the detection of unknown objects. Further, we employ i) an evaluation technique that penalizes low-confidence outputs to mitigate the risk of misclassification of the unknown objects and ii) a new metric called HMP that combines known and unknown precision using harmonic mean. Our extensive experiments demonstrate that the proposed model achieves significant improvement on the MS-COCO & PASCAL VOC dataset for the OSOD task. Hiran Sarkar, Vishal M. Chudasama, Naoyuki Onoe, Pankaj Wasnik, Vineeth N. Balasubramanian |
WACV | 4 |
| 2023 | Fiducial Focus Augmentation for Facial Landmark Detection
Purbayan Kar, Vishal M. Chudasama, Naoyuki Onoe, Pankaj Wasnik, Vineeth N. Balasubramanian |
BMVC | 4 |
| 2018 | Fusion of Multi-Scale Local Phase Quantization Features for Face Presentation Attack DetectionabstractFace recognition systems are widely known for their vulnerability against presentation attacks or spoofing attacks. The exponential deployment of face recognition systems has been further challenged even by the simple and low-cost face artefacts generated using conventional printers. In this paper, we present a novel scheme to detect face presentation attacks posed by high-quality print attacks which are relatively difficult to detect. The proposed scheme leverages the phase information extracted from the spatial-frequency representation of the given image. We also present a new face presentation attack database collected using the iPhone 6S. The new database is comprised of 100 subjects collected in two different sessions that have resulted in a total of 31228 samples (or images). Extensive experiments are carried out on the newly constructed database and the obtained results show the improved performance of the proposed scheme when compared aaainst six different state-of-the-art methods. Ramachandra Raghavendra, Sushma Venkatesh, Kiran B. Raja, Pankaj Wasnik, Martin Stokkenes, Christoph Busch 0001 |
FUSION | 4 |
| 2018 | Subjective Logic Based Score Level Fusion: Combining Faces and FingerprintsabstractBiometric systems are prone to random and systematic errors which are typically attributed to the variations in terms of inter-session data capture and intra-session variability. Furthermore, these errors cannot be defined and modeled mathematically in many cases, but we can associate them with uncertainty based on certain conditions. In such cases, one of the possible approach to improve biometric system performance is to employ multi-biometric fusion by incorporating the uncertainties. In the literature, researchers have proposed many fusion techniques, but most of these techniques do not take uncertainty into account while performing fusion. Since the decision made by uni-modal biometric comparators do not consider the uncertainty involved in such decisions, it is essential first to model the uncertainty before combining the decision from multiple uni-modal biometric systems efficiently. To this end, we propose a score level multi-biometric fusion scheme using Subjective Logic which incorporates the uncertainty of the system's information channels while fusing the scores. Extensive experiments are carried out on the multi-biometric NIST BSSR1, and the proposed scheme has indicated a superior performance with a genuine match rate of 99.02 % at a false match rate fixed to 0.01 %. Pankaj Wasnik, Ramachandra Raghavendra, Kiran B. Raja, Christoph Busch 0001 |
FUSION | 1 |
| 2017 | Robust face presentation attack detection on smartphones : An approach based on variable focusabstractSmartphone based facial biometric systems have been well used in many of the security applications starting from simple phone unlocking to secure banking applications. This work presents a new approach of exploring the intrinsic characteristics of the smartphone camera to capture a number of stack images in the depth-of-field. With the set of stack images obtained, we present a new feature-free and classifier-free approach to provide the presentation attack resistant face biometric system. With the entire system implemented on the smartphone, we demonstrate the applicability of the proposed scheme in obtaining a stack of images with varying focus to effectively determine the presentation attacks. We create a new database of 13250 images at different focal length to present a detailed analysis of vulnerability together with the evaluation of proposed scheme. An extensive evaluation of the newly created database comprising of 5 different Presentation Attack Instruments (PAI) has demonstrated an outstanding performance on all 5 PAI through proposed approach. With the set ofcomplementary benefits of proposed approach illustrated in this work, we deduce the robustness towards unseen 2D attacks. Kiran B. Raja, Pankaj Wasnik, Ramachandra Raghavendra, Christoph Busch 0001 |
IJCB | 2 |
| 2016 | Eye region based multibiometric fusion to mitigate the effects of body weight variations in face recognition
Pankaj Wasnik, Kiran B. Raja, Ramachandra Raghavendra, Christoph Busch 0001 |
FUSION | 1 |