VLDB 2026 Research / reviewers in the wild / expert
Bhavin Jawade
dblp:310/1267
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2025
0009-0008-6059-5364ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio-Visual Representation Learning For Lip-Sync Estimation Through Ranking Augmented Contrastive TrainingabstractIn many applications, particularly in media production and content localization, it is crucial to detect and evaluate varying degrees of audio-visual synchronization, such as selecting high-quality dubbed audio over poorly synchronized tracks. Traditional contrastively pre-trained LipSync models are designed to distinguish perfectly synced audio from unsynced audio. However, these models fall short when it comes to detecting partial synchronization, such as in dubbed audio, because their training objective is focused on pulling synced lip-motion and audio closer together while pushing everything else apart. This approach limits their ability to accurately gauge varying levels of sync, leading to challenges in scenarios that require a more nuanced understanding of synchronization quality. To address this limitation, we propose a novel deep metric learning approach, the Ranking Supervised Multi-Similarity (RSMS) loss formulation, which introduces a ranking prior as a supervision signal. Our method integrates hard-sample mining to enforce this ranking, allowing the model to better differentiate between partial-syncs and completely unsynced audios. Furthermore, we demonstrate the effectiveness of using “Dubbed Audio” as a train-time example of partial-syncs, leading to improved performance in lip-sync models. Bhavin Jawade, Ravi Gadde, Christophe Bejjani, Yinghong Lan |
ICASSP | 1 |
| 2025 | Ridgeformer: Mutli-Stage Contrastive Training for Fine-Grained Cross-Domain Fingerprint RecognitionabstractThe increasing demand for hygienic and portable biometric systems has underscored the critical need for advancements in contactless fingerprint recognition. Despite its potential, this technology faces notable challenges, including out-of-focus image acquisition, reduced contrast between fingerprint ridges and valleys, variations in finger positioning, and perspective distortion. These factors significantly hinder the accuracy and reliability of contactless fingerprint matching. To address these issues, we propose a novel multi-stage transformer-based contactless fingerprint matching approach that first captures global spatial features and subsequently refines localized feature alignment across fingerprint samples. By employing a hierarchical feature extraction and matching pipeline, our method ensures fine-grained, cross-sample alignment while maintaining the robustness of global feature representation. We perform extensive evaluations on publicly available datasets such as HKPolyU and RidgeBase under different evaluation protocols, such as contactless-to-contact matching and contactless-to-contactless matching and demonstrate that our proposed approach outperforms existing methods, including COTS solutions. Our codebase is available at https://github.com/KNITPhoenix/Ridgeformer Shubham Pandey, Bhavin Jawade, Srirangaraj Setlur |
ICIP | 2 |
| 2025 | SCOT: Self-Supervised Contrastive Pretraining for Zero-Shot Compositional RetrievalabstractCompositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains including product retrieval (e-commerce) and web search. Existing methods primarily focus on fully-supervised learning, wherein models are trained on datasets of labeled triplets such as FashionIQ and CIRR. This poses two significant challenges: (i) curating such triplet datasets is labor intensive; and (ii) models lack generalization to un-seen objects and domains. In this work, we propose SCOT (Self-supervised COmpositional Training), a novel zero-shot compositional pretraining strategy that combines existing large image-text pair datasets with the generative capabilities of large language models to contrastively train an embedding composition network. Specifically, we show that the text embedding from a large-scale contrastively-pretrained vision-language model can be utilized as proxy target supervision during compositional pretraining, replacing the target image embedding. In zero-shot settings, this strategy surpasses SOTA zero-shot compositional re-trieval methods as well as many fully-supervised methods on standard benchmarks such as FashionIQ and CIRR. Our code and models are available at https://github.com/yahoo/SCOT. Bhavin Jawade, João V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, Venu Govindaraju |
WACV | 1 |
| 2024 | GestSpoof: Gesture Based Spatio-Temporal Representation Learning for Robust Fingerprint Presentation Attack DetectionabstractFingerprint spoof attacks represent one of the most prevalent forms of biometric presentation attacks. While significant progress has been made in framing fingerprint spoof detection as a general image classification problem, limited attention has been given to treating it as a temporal learning problem. The distinctions in the elastic properties between authentic and synthetically created counterfeit fingerprints can be more accurately captured under motion-induced gestures during acquisition. In this study, we introduce a novel method for detecting fake fingerprints by deliberately introducing distortions through sliding and twisting motions during acquisition. As widely used spoof datasets such as those from LivDet 2009 to 2021 or MSU FPAD lack the temporal information essential for this investigation, we assembled a new dataset focused on distortion-based fake and real fingerprints, encompassing various types of spoof materials and diverse distortions. This gesture-equipped dataset comprises more than 3680 videos gathered from 184 unique fingers. Additionally, we present a novel spatial-temporal multi-modal network for detecting fingerprint spoofs using intentional-distortion. Our proposed approach yields significantly improved results compared to traditional static classification-based methods for spoof detection, across various metrics and for both known and unknown (generalization) scenarios, thereby highlighting the substantial impact that introducing gestures can have on enhancing fingerprint spoof detection. The dataset can be downloaded from here: https://www.buffalo.edu/cubs/research/datasets/gestspoof-dataset.html Bhavin Jawade, Shreeram Subramanya, Atharv Dabhade, Srirangaraj Setlur, Venu Govindaraju |
FG | 1 |
| 2024 | DIOR: Dataset for Indoor-Outdoor Reidentification Long Range 3D/2D Skeleton Gait Collection Pipeline, Semi-Automated Gait Keypoint Labeling and Baseline Evaluation MethodsabstractRecently, there has been growing interest in the identification and re-identification of individuals from long distances using rooftop cameras, UAV cameras, street cams, and similar devices. This type of recognition extends beyond facial recognition, utilizing whole-body markers such as gait. However, datasets to train and test such recognition algorithms are scarce and often lack labeling. This paper introduces DIOR—a comprehensive framework for data collection, semi-automated annotation, and a dataset comprising 1.649 million RGB frames labeled with 3D/2D skeleton gait markers across 14 subjects. This dataset includes 200,000 RGB frames captured from long-range cam-eras. Our approach employs advanced 3D computer vision techniques to achieve pixel-level accuracy in indoor environments using motion capture systems. For outdoor, long-range environments, we eliminate the reliance on motion capture systems and implement a cost-effective, hybrid 3D computer vision and learning pipeline using only four inexpensive RGB cameras. This method successfully achieves precise skeleton labeling of distant subjects, even when their visual size is as small as 20-25 pixels within an RGB frame. We benchmark models trained on existing datasets such as CASIA-B, on our proposed dataset for the task of Gait recognition. Our pipeline and the accompanying dataset will be made publicly available following acceptance. Praveen Raj Masilamani, Bhavin Jawade, Srirangaraj Setlur, Karthik Dantu |
IJCB | 3 |
| 2024 | ProxyFusion: Face Feature Aggregation Through Sparse ExpertsabstractFace feature fusion is indispensable for robust face recognition, particularly in scenarios involving long-range, low-resolution media (unconstrained environments) where not all frames or features are equally informative. Existing methods often rely on large intermediate feature maps or face metadata information, making them incompatible with legacy biometric template databases that store pre-computed features. Additionally, real-time inference and generalization to large probe sets remains challenging.
To address these limitations, we introduce a linear time O(N) proxy based sparse expert selection and pooling approach for context driven feature-set attention. Our approach is order invariant on the feature-set, generalizes to large sets, is compatible with legacy template stores, and utilizes significantly less parameters making it suitable real-time inference and edge use-cases. Through qualitative experiments, we demonstrate that ProxyFusion learns discriminative information for importance weighting of face features without relying on intermediate features. Quantitative evaluations on challenging low-resolution face verification datasets such as IARPA BTS3.1 and DroneSURF show the superiority of ProxyFusion in unconstrained long-range face recognition setting.
Our code and pretrained models are available at: https://github.com/bhavinjawade/ProxyFusion Bhavin Jawade, Alexander Stone, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
NeurIPS | 1 |
| 2023 | CoNAN: Conditional Neural Aggregation Network For Unconstrained Face Feature FusionabstractFace recognition from image sets acquired under unregulated and uncontrolled settings, such as at large distances, low resolutions, varying viewpoints, illumination, pose, and atmospheric conditions, is challenging. Face feature aggregation, which involves aggregating a set of N feature representations present in a template into a single global representation, plays a pivotal role in such recognition systems. Existing works in traditional face feature aggregation either utilize metadata or high-dimensional intermediate feature representations to estimate feature quality for aggregation. However, generating high-quality metadata or style information is not feasible for extremely low-resolution faces captured in long-range and high altitude settings. To overcome these limitations, we propose a feature distribution conditioning approach called CoNAN for template aggregation. Specifically, our method aims to learn a context vector conditioned over the distribution information of the incoming feature set, which is utilized to weigh the features based on their estimated informativeness. The proposed method produces state-of-the-art results on long-range unconstrained face recognition datasets such as BTS, and DroneSURF, validating the advantages of such an aggregation strategy. Bhavin Jawade, Deen Dayal Mohan, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
IJCB | 1 |
| 2023 | Liveness Detection Competition - Noncontact-based Fingerprint Algorithms and Systems (LivDet-2023 Noncontact Fingerprint)abstractLiveness Detection (LivDet) is an international competition series open to academia and industry with the objective to assess and report state-of-the-art in Presentation Attack Detection (PAD). LivDet-2023 Noncontact Fingerprint is the first edition of the noncontact fingerprint-based PAD competition for algorithms and systems. The competition serves as an important benchmark in noncontact-based fingerprint PAD, offering (a) independent assessment of the state-of-the-art in noncontact-based fingerprint PAD for algorithms and systems, and (b) common evaluation protocol, which includes finger photos of a variety of Presentation Attack Instruments (PAIs) and live fingers to the biometric research community (c) provides standard algorithm and system evaluation protocols, along with the comparative analysis of state-of-the-art algorithms from academia and industry with both old and new android smartphones. The winning algorithm achieved an APCER of 11.35% averaged over all PAIs and a BPCER of 0.62%. The winning system achieved an APCER of 13.0.4%, averaged over all PAIs tested over all the smartphones, and a BPCER of 1.68% over all smartphones tested. Four-finger systems that make individual finger-based PAD decisions were also tested. The dataset used for competition will be available1, to all researchers as per data share protocol.1https://noncontactfingerprint2023.1ivdet.org/index.php Sandip Purnapatra, Humaira Rezaie, Bhavin Jawade, Yu Liu 0069, Luke Brosell, Mst Rumana Sumi, Lambert Igene, Alden Dimarco, Srirangaraj Setlur, Soumyabrata Dey, Stephanie Schuckers, Marco Huber, Jan Niklas Kolf, Meiling Fang, Naser Damer, Banafsheh Adami, Raul Chitic, Karsten Seelert, Vishesh Mistry, Rahul Parthe, Umit Kacar |
IJCB | 3 |
| 2023 | RealCQA: Scientific Chart Question Answering as a Test-Bed for First-Order Logic
Saleem Ahmed, Bhavin Jawade, Shubham Pandey, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (3) | 2 |
| 2023 | Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source LocalizationabstractLearning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to localize the source of the sound. In a video, oftentimes, the objects exhibiting movement are the ones generating the sound. In this work, we capture this characteristic by modeling the optical flow in a video as a prior to better aid in localizing the sound source. We further demonstrate that the addition of flow-based attention substantially improves visual sound source localization. Finally, we benchmark our method on standard sound source localization datasets and achieve state-of-the-art performance on the Soundnet Flickr and VGG Sound Source datasets. Code: https://github.com/denfed/heartheflow. Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju |
WACV | 3 |
| 2023 | NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal EmbeddingsabstractCross-modal retrieval is a fundamental vision-language task with a broad range of practical applications. Text-to-image matching is the most common form of cross-modal retrieval where, given a large database of images and a textual query, the task is to retrieve the most relevant set of images. Existing methods utilize dual encoders with an attention mechanism and a ranking loss for learning embeddings that can be used for retrieval based on cosine similarity. Despite the fact that these methods attempt to perform semantic alignment across visual regions and textual words using tailored attention mechanisms, there is no explicit supervision from the training objective to enforce such alignment. To address this, we propose NAPReg, a novel regularization formulation that projects high-level semantic entities i.e Nouns into the embedding space as shared learnable proxies. We show that using such a formulation allows the attention mechanism to learn better word-region alignment while also utilizing region information from other samples to build a more generalized latent representation for semantic concepts. Experiments on three benchmark datasets i.e. MS-COCO, Flickr30k and Flickr8k demonstrate that our method achieves state-of-the-art results in cross-modal metric learning for text-image and image-text retrieval tasks. Code: https://github.com/bhavinjawade/NAPReq Bhavin Jawade, Deen Dayal Mohan, Naji Mohamed Ali, Srirangaraj Setlur, Venu Govindaraju |
WACV | 1 |
| 2022 | Attribute De-biased Vision Transformer (AD-ViT) for Long-Term Person Re-identificationabstractPerson re-identification (re-ID) aims to retrieve images of the same identity from a gallery of person images across cameras and viewpoints. However, most works in person re-ID assume a short-term setting characterized by invariance in appearance. In contrast, a high visual variance can be frequently seen in a long-term setting due to changes in apparel and accessories, which makes the task more challenging. Therefore, learning identity-specific features agnostic of temporally variant features is crucial for robust long-term person Re-ID. To this end, we propose an Attribute De-biased Vision Transformer (AD-ViT) to provide direct supervision to learn identity-specific features. Specifically, we produce attribute labels for person instances and utilize them to guide our model to focus on identity features through gradient reversal. Our experiments on two long-term re-ID datasets - LTCC and NKUP show that the proposed work consistently outperforms current state-of-the-art methods. Kyung Won Lee, Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
AVSS | 2 |
| 2022 | RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint DatasetabstractContactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks. However, development of practical and robust contactless fingerprint matching techniques is constrained by the limited availablity of large scale real-world datasets. To motivate further advances in contactless fingerprint matching across sensors, we introduce the RidgeBase benchmark dataset. RidgeBase consists of more than 15,000 contactless and contact-based fingerprint image pairs acquired from 88 individuals under different background and lighting conditions using two smartphone cameras and one flatbed contact sensor. Unlike existing datasets, RidgeBase is designed to promote research under different matching scenarios that include Single Finger Matching and Multi-Finger Matching for both contactless-to-contactless (CL2CL) and contact-to-contactless (C2CL) verification and identification. Furthermore, due to the high intra-sample variance in contactless fingerprints belonging to the same finger, we propose a set-based matching protocol inspired by the advances in facial recognition datasets. This protocol is specifically designed for pragmatic contactless fingerprint matching that can account for variances in focus, polarity and finger-angles. We report qualitative and quantitative baseline results for different protocols using a COTS fingerprint matcher (Verifinger) and a Deep CNN based approach on the RidgeBase dataset. The dataset can be downloaded here: https://www.buffalo.edu/cubs/research/datasets/ridgebase-benchmark-dataset.html Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Nalini K. Ratha, Venu Govindaraju |
IJCB | 1 |