VLDB 2026 Research / reviewers in the wild / expert
Karthik Nandakumar
dblp:83/3874
· DBLP profile ↗
53ranked-venue papers
4as first author
39since 2021 · last 2026
0000-0002-6274-9725ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 33 since 2021Artificial intelligence and machine learning · 27 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Security and privacy · 8 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VFace: A Training-Free Approach for Diffusion-Based Video Face SwappingabstractWe present a training-free, plug-and-play method, namely VFace, for high-quality face swapping in videos. It can be seamlessly integrated with image-based face swapping approaches built on diffusion models. First, we introduce a Frequency Spectrum Attention Interpolation technique to facilitate generation and intact key identity characteristics. Second, we achieve Target Structure Guidance via plug-and-play attention injection to better align the structural features from the target frame to the generation. Third, we present a Flow-Guided Attention Temporal Smoothening mechanism that enforces spatiotemporal coherence without modifying the underlying diffusion model to reduce temporal inconsistencies typically encountered in frame-wise generation. Our method requires no additional training or video-specific fine-tuning. Extensive experiments show that our method significantly enhances temporal consistency and visual fidelity, offering a practical and modular solution for video-based face swapping. Our code is available at VFace. Sanoojan Baliah, Yohan Abeysinghe, Rusiru Thushara, Khan Muhammad 0001, Abhinav Dhall, Karthik Nandakumar, Muhammad Haris Khan |
WACV | 6 |
| 2026 | GenMix: Effective data augmentation with generative diffusion model image editingabstractData augmentation is widely used to enhance generalization in visual classification tasks. However, traditional methods struggle when source and target domains differ, as in domain adaptation, due to their inability to address domain gaps. This paper introduces GenMix, a generalizable prompt-guided generative data augmentation approach that enhances both in-domain and cross-domain image classification. Our technique leverages image editing to generate augmented images based on custom conditional prompts, designed specifically for each problem type. By blending portions of the input image with its edited generative counterpart and incorporating fractal patterns, our approach mitigates unrealistic images and label ambiguity, improving the performance and adversarial robustness of the resulting models. Efficacy of our method is established with extensive experiments on eight public datasets for general and fine-grained classification, in both in-domain and cross-domain settings. Additionally, we demonstrate performance improvements for self-supervised learning, learning with data scarcity, and adversarial robustness. As compared to the existing state-of-the-art methods, our technique achieves stronger performance across the board. Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar, Naveed Akhtar |
Expert Syst. Appl. | 4 |
| 2025 | STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion ModelsabstractThe rapid proliferation of large-scale text-to-image diffusion (T2ID) models has raised serious concerns about their potential misuse in generating harmful content. Although numerous methods have been proposed for erasing undesired concepts from T2ID models, they often provide a false sense of security; concept-erased models (CEMs) can still be manipulated via adversarial attacks to regenerate the erased concept. While a few robust concept erasure methods based on adversarial training have emerged recently, they compromise on utility (generation quality for benign concepts) to achieve robustness and/or remain vulnerable to advanced embedding space attacks. These limitations stem from the failure of robust CEMs to thoroughly search for "blind spots" in the embedding space. To bridge this gap, we propose STEREO, a novel two-stage framework that employs adversarial training as a first step rather than the only step for robust concept erasure. In the first stage, STEREO employs adversarial training as a vulnerability identification mechanism to search thoroughly enough. In the second robustly erase once stage, STEREO introduces an anchor-concept-based compositional objective to robustly erase the target concept in a single fine-tuning stage, while minimizing the degradation of model utility. We benchmark STEREO against seven state-of-the-art concept erasure methods, demonstrating its superior robustness to both white-box and black-box attacks, while largely preserving utility. Koushik Srivatsan, Fahad Shamshad, Muzammal Naseer, Vishal M. Patel, Karthik Nandakumar |
CVPR | 5 |
| 2025 | TrojanWave: Exploiting Prompt Learning for Stealthy Backdoor Attacks on Large Audio-Language ModelsabstractPrompt learning has emerged as an efficient alternative to full fine-tuning for adapting large audio-language models (ALMs) to downstream tasks.While this paradigm enables scalable deployment via Prompt-as-a-Service frameworks, it also introduces a critical yet underexplored security risk of backdoor attacks.In this work, we present TrojanWave, the first backdoor attack tailored to the prompt-learning setting in frozen ALMs.Unlike prior audio backdoor methods that require training from scratch on full datasets, TrojanWave injects backdoors solely through learnable prompts, making it highly scalable and effective in few-shot settings.TrojanWave injects imperceptible audio triggers in both time and spectral domains to effectively induce targeted misclassification during inference.To mitigate this threat, we further propose TrojanWave-Defense, a lightweight prompt purification method that neutralizes malicious prompts without hampering the clean performance.Extensive experiments across 11 diverse audio classification benchmarks demonstrate the robustness and practicality of both the attack and defense.Our code is publicly available at Github † . Asif Hanif, Maha Tufail Agro, Fahad Shamshad, Karthik Nandakumar |
EMNLP | 4 |
| 2025 | FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space MixingabstractAdvancements in face recognition (FR) technologies have amplified privacy concerns, necessitating methods that protect identity while maintaining recognition utility. Existing face anonymization methods typically focus on obscuring identity but fail to meet the requirements of biometric template protection, including revocability, unlinkability, and irreversibility. We propose FaceAnonyMixer, a cancelable face generation framework that leverages the latent space of a pre-trained generative model to synthesize privacy-preserving face images. The core idea of FaceAnonyMixer is to irreversibly mix the latent code of a real face image with a synthetic code derived from a revocable key. The mixed latent code is further refined through a carefully designed multi-objective loss to satisfy all cancelable biometric requirements. FaceAnonyMixer is capable of generating high-quality cancelable faces that can be directly matched using existing FR systems without requiring any modifications. Extensive experiments on benchmark datasets demonstrate that FaceAnonyMixer delivers superior recognition accuracy while providing significantly stronger privacy protection, achieving over an 11% absolute gain on commercial API compared to recent cancelable biometric methods. Code is available at: https://github.com/talha-alam/faceanonymixer Mohammed Talha Alam, Fahad Shamshad, Fakhri Karray, Karthik Nandakumar |
IJCB | 4 |
| 2025 | A Framework for Double-Blind Federated Adaptation of Foundation ModelsabstractFoundation models (FMs) excel in zero-shot tasks but benefit from task-specific adaptation. However, privacy concerns prevent data sharing among multiple data owners, and proprietary restrictions prevent the learning service provider (LSP) from sharing the FM. In this work, we propose BlindFed, a framework enabling collaborative FM adaptation while protecting both parties: data owners do not access the FM or each other's data, and the LSP does not see sensitive task data. BlindFed relies on fully homomorphic encryption (FHE) and consists of three key innovations: (i) FHE-friendly architectural modifications via polynomial approximations and low-rank adapters, (ii) a two-stage split learning approach combining offline knowledge distillation and online encrypted inference for adapter training without backpropagation through the FM, and (iii) a privacy-boosting scheme using sample permutations and stochastic block sampling to mitigate model extraction attacks. Empirical results on four image classification datasets demonstrate the practical feasibility of the BlindFed framework, albeit at a high communication cost and large computational complexity for the LSP. Nurbek Tastan, Karthik Nandakumar |
ICCV | 2 |
| 2025 | Aequa: Fair Model Rewards in Collaborative Learning via Slimmable NetworksabstractCollaborative learning enables multiple participants to learn a single global model by exchanging focused updates instead of sharing data. One of the core challenges in collaborative learning is ensuring that participants are rewarded fairly for their contributions, which entails two key sub-problems: contribution assessment and reward allocation. This work focuses on fair reward allocation, where the participants are incentivized through model rewards - differentiated final models whose performance is commensurate with the contribution. In this work, we leverage the concept of slimmable neural networks to collaboratively learn a shared global model whose performance degrades gracefully with a reduction in model width. We also propose a post-training fair allocation algorithm that determines the model width for each participant based on their contributions. We theoretically study the convergence of our proposed approach and empirically validate it using extensive experiments on different datasets and architectures. We also extend our approach to enable training-time model reward allocation. Nurbek Tastan, Samuel Horváth, Karthik Nandakumar |
ICML | 3 |
| 2025 | Forget-MI: Machine Unlearning for Forgetting Multimodal Information in Healthcare Settings
Shahad Hardan, Darya Taratynova, Abdelmajid Essofi, Karthik Nandakumar, Mohammad Yaqub |
MICCAI (3) | 4 |
| 2025 | Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language ModelsabstractThe conventional modus operandi for adapting pre-trained vision-language models (VLMs) during test-time involves tuning learnable prompts, i.e., test-time prompt tuning. This paper introduces Test-Time Low-rank adaptation (TTL) as an alternative to prompt tuning for zero-shot generalization of large-scale VLMs. Taking inspiration from recent advancements in efficiently fine-tuning large language models, TTL offers a test-time parameter-efficient adaptation approach that updates the attention weights of the transformer encoder by maximizing prediction confidence. The self-supervised confidence maximization objective is specified using a weighted entropy loss that enforces consistency among predictions of augmented samples. TTL introduces only a small amount of trainable parameters for low-rank adapters in the model space while keeping the prompts and backbone frozen. Extensive experiments on a variety of natural distribution and cross-domain tasks show that TTL can outperform other techniques for test-time optimization of VLMs in strict zero-shot settings. Specifically, TTL outperforms test-time prompt tuning baselines with a significant improvement on average. Our code is available at https://github.com/Razaimam45/TTLTest-Time-Low-Rank-Adaptation. Raza Imam, Hanan Gani, Muhammad Huzaifa, Karthik Nandakumar |
WACV | 4 |
| 2024 | Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New BaselineabstractUnsupervised (US) video anomaly detection (VAD) in surveillance applications is gaining more popularity recently due to its practical real-world applications. As surveillance videos are privacy sensitive and the availability of large-scale video data may enable better US- VAD systems, collaborative learning can be highly rewarding in this setting. However, due to the extremely challenging nature of the US- VAD task, where learning is carried out without any annotations, privacy-preserving collaborative learning of us- VAD systems has not been studied yet. In this paper, we propose a new baseline for anomaly detection capable of localizing anomalous events in complex surveil-lance videos in a fully unsupervised fashion without any labels on a privacy-preserving participant-based distributed training configuration. Additionally, we propose three new evaluation protocols to benchmark anomaly detection approaches on various scenarios of collaborations and data availability. Based on these protocols, we modify existing VAD datasets to extensively evaluate our approach as well as existing US SOTA methods on two large-scale datasets including UCF-Crime and XD- Violence. All proposed evaluation protocols, dataset splits, and codes are available here: https://github.com/AnasEmadllICLAP. Anas Al-lahham, Muhammad Zaigham Zaheer, Nurbek Tastan, Karthik Nandakumar |
CVPR | 4 |
| 2024 | Attack To Defend: Exploiting Adversarial Attacks for Detecting Poisoned ModelsabstractPoisoning (trojan/backdoor) attacks enable an adversary to train and deploy a corrupted machine learning (ML) model, which typically works well and achieves good ac-curacy on clean input samples but behaves maliciously on poisoned samples containing specific trigger patterns. Using such poisoned ML models as the foundation to build real-world systems can compromise application safety. Hence, there is a critical need for algorithms that detect whether a given target model has been poisoned. This work proposes a novel approach for detecting poisoned models called Attack To Defend (A2D), which is based on the observation that poisoned models are more sensitive to adversarial pertur-bations compared to benign models. We propose a metric called sensitivity to adversarial perturbations (SAP) to mea-sure the sensitivity of a ML model to adversarial attacks at a specific perturbation bound. We then generate strong ad-versarial attacks against an unrelated reference model and estimate the SAP value of the target model by transferring the generated attacks. The target model is deemed to be a trojan if its SAP value exceeds a decision threshold. The A2D framework requires only black-box access to the target model and a small clean set, while being computationally efficient. The A2D approach has been evaluated on four standard image datasets and its effectiveness under various types of poisoning attacks has been demonstrated. Samar Fares, Karthik Nandakumar |
CVPR | 2 |
| 2024 | Diffusemix: Label-Preserving Data Augmentation with Diffusion ModelsabstractRecently, a number of image-mixing-based augmentation techniques have been introduced to improve the gen-eralization of deep neural networks. In these techniques, two or more randomly selected natural images are mixed together to generate an augmented image. Such methods may not only omit important portions of the input images but also introduce label ambiguities by mixing images across labels resulting in misleading supervisory signals. To address these limitations, we propose Diffusemix, a novel data augmentation technique that leverages a diffusion model to reshape training images, supervised by our bespoke conditional prompts. First, concatenation of a partial natural image and its generated counterpart is ob-tained which helps in avoiding the generation of unrealistic images or label ambiguities. Then, to enhance resilience against adversarial attacks and improves safety measures, a randomly selected structural pattern from a set of frac-tal images is blended into the concatenated image to form the final augmented image for training. Our empirical results on seven different datasets reveal that Diffusemix achieves superior performance compared to existing state-of-the-art methods on tasks including general classification, fine- grained classification, fine-tuning, data scarcity, and adversarial robustness. Augmented datasets and codes are available here: https://diffusemix.github.io/ Khawar Islam, Muhammad Zaigham Zaheer, Arif Mahmood, Karthik Nandakumar |
CVPR | 4 |
| 2024 | Feature Map Purification for Enhancing Adversarial Robustness of Deep Timeseries ClassifiersabstractDeep-learning based timeseries classifiers are known to be susceptible to adversarial attacks, where the adversary adds imperceptible perturbations to the input sample to cause mis-classification. While principled adversarial defense mechanisms such as adversarial training and certified robustness have been proposed for image classifiers, they have seldom been studied in the timeseries domain. Existing defenses for timeseries classifiers are primarily centered around adversarial sample detection, but have mixed performance and fail to generalize well across attacks. This work proposes an alternative approach based on purifying the intermediate representations within a deep convolutional timeseries (DCT) classifier. We design a learnable sub-network with residual connections that filters the feature maps in multiple wavelet basis spaces to suppress the adversarial perturbations. Given any pretrained non-robust DCT classifier, the proposed feature map purification (FeMPure) module can be trained in isolation without affecting the given classifier and can be seamlessly plugged back in to enhance the adversarial robustness of the original classifier. Experiments based on 2 well-known architectures for DCT classifiers, 6 adversarial attacks, and 80 public-domain datasets demonstrate that the proposed FeMPure approach can provide good adversarial robustness, irrespective of whether the adversary is unaware or has full knowledge of the defense mechanism. With minor modifications, the FeMPure approach can also be employed for adversarial sample detection or for enhancing certified robustness. Mubarak G. Abdu-Aguye, Muhammad Zaigham Zaheer, Karthik Nandakumar |
ICDM | 3 |
| 2024 | Multi-Attribute Vision Transformers are Efficient and Robust LearnersabstractSince their inception, Vision Transformers (ViTs) have emerged as a compelling alternative to Convolutional Neural Networks (CNNs) across a wide spectrum of tasks. ViTs exhibit notable characteristics, including global attention, resilience against occlusions, and adaptability to distribution shifts. One underexplored aspect of ViTs is their potential for multi-attribute learning, referring to their ability to simultaneously grasp multiple attribute-related tasks. In this paper, we delve into the multi-attribute learning capability of ViTs, presenting a straightforward yet effective strategy for training various attributes through a single ViT network as distinct tasks. We assess the resilience of multi-attribute ViTs against adversarial attacks and compare their performance against ViTs designed for single attributes. Moreover, we further evaluate the robustness of multi-attribute ViTs against a recent transformer based attack called Patch-Fool. Our empirical findings on the CelebA dataset provide validation for our assertions. Hanan Gani, Nada Saadi, Noor Hussein, Karthik Nandakumar |
ICIP | 4 |
| 2024 | Dirichlet-based Uncertainty Quantification for Personalized Federated Learning with Improved Posterior Networks
Nikita Kotelevskii, Samuel Horváth, Karthik Nandakumar, Martin Takác 0001, Maxim Panov |
IJCAI | 3 |
| 2024 | Redefining Contributions: Shapley-Driven Federated Learning
Nurbek Tastan, Samar Fares, Toluwani Aremu, Samuel Horváth, Karthik Nandakumar |
IJCAI | 5 |
| 2024 | BAPLe: Backdoor Attacks on Medical Foundational Models Using Prompt Learning
Asif Hanif, Fahad Shamshad, Muhammad Awais Hassan, Muzammal Naseer, Fahad Shahbaz Khan, Karthik Nandakumar, Salman Khan 0001, Rao Muhammad Anwer |
MICCAI (12) | 6 |
| 2024 | PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
Noor Hussein, Fahad Shamshad, Muzammal Naseer, Karthik Nandakumar |
MICCAI (12) | 4 |
| 2024 | PEMMA: Parameter-Efficient Multi-Modal Adaptation for Medical Image Segmentation
Nada Saadi, Numan Saeed, Mohammad Yaqub, Karthik Nandakumar |
MICCAI (12) | 4 |
| 2024 | SurvRNC: Learning Ordered Representations for Survival Prediction Using Rank-N-Contrast
Numan Saeed, Muhammad Ridzuan, Fadillah A. Maani, Hussain Alasmawi, Karthik Nandakumar, Mohammad Yaqub |
MICCAI (5) | 5 |
| 2024 | A Synopsis of FAME 2024 Challenge: Associating Faces with Voices in Multilingual Environments
Muhammad Saad Saeed, Shah Nawaz, Marta Moscati, Rohan Kumar Das, Muhammad Salman Tahir, Muhammad Zaigham Zaheer, Muhammad Irzam Liaqat, Muhammad Haris Khan, Karthik Nandakumar, Muhammad Haroon Yousaf, Markus Schedl |
ACM Multimedia | 9 |
| 2024 | A Coarse-to-Fine Pseudo-Labeling (C2FPL) Framework for Unsupervised Video Anomaly DetectionabstractDetection of anomalous events in videos is an important problem in applications such as surveillance. Video anomaly detection (VAD) is well-studied in the one-class classification (OCC) and weakly supervised (WS) settings. However, fully unsupervised (US) video anomaly detection methods, which learn a complete system without any annotation or human supervision, have not been explored in depth. This is because the lack of any ground truth annotations significantly increases the magnitude of the VAD challenge. To address this challenge, we propose a simple-but-effective two-stage pseudo-label generation framework that produces segment-level (normal/anomaly) pseudo-labels, which can be further used to train a segment-level anomaly detector in a supervised manner. The proposed coarse-to-fine pseudo-label (C2FPL) generator employs carefully-designed hierarchical divisive clustering and statistical hypothesis testing to identify anomalous video segments from a set of completely unlabeled videos. The trained anomaly detector can be directly applied on segments of an unseen test video to obtain segment-level, and subsequently, frame-level anomaly predictions. Extensive studies on two large-scale public-domain datasets, UCF-Crime and XD-Violence, demonstrate that the proposed unsupervised approach achieves superior performance compared to all existing OCC and US methods, while yielding comparable performance to the state-of-the-art WS methods. Anas Al-lahham, Nurbek Tastan, Muhammad Zaigham Zaheer, Karthik Nandakumar |
WACV | 4 |
| 2023 | CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent SearchabstractThe success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to generate “naturalistic” images that can protect facial privacy without compromising user experience. We propose a novel two-step approach for facial privacy protection that relies on finding adversarial latent codes in the low- dimensional manifold of a pretrained generative model. The first step inverts the given face image into the latent space and finetunes the generative model to achieve an accurate reconstruction of the given image from its latent code. This step produces a good initialization, aiding the generation of high-quality faces that resemble the given identity. Subsequently, user-defined makeup text prompts and identity- preserving regularization are used to guide the search for adversarial codes in the latent space. Extensive experiments demonstrate that faces generated by our approach have stronger black-box transferability with an absolute gain of 12.06% over the state-of-the-art facial privacy protection approach under the face verification task. Finally, we demonstrate the effectiveness of the proposed approach for commercial face recognition systems. Our code is available at https://github.com/fahadshamshad/Clip2Protect. Fahad Shamshad, Muzammal Naseer, Karthik Nandakumar |
CVPR | 3 |
| 2023 | Evading Forensic Classifiers with Attribute-Conditioned Adversarial FacesabstractThe ability of generative models to produce highly realistic synthetic face images has raised security and ethical concerns. As a first line of defense against such fake faces, deep learning based forensic classifiers have been developed. While these forensic models can detect whether a face image is synthetic or real with high accuracy, they are also vulnerable to adversarial attacks. Although such attacks can be highly successful in evading detection by forensic classifiers, they introduce visible noise patterns that are detectable through careful human scrutiny. Additionally, these attacks assume access to the target model(s) which may not always be true. Attempts have been made to directly perturb the latent space of GANs to produce adversarial fake faces that can circumvent forensic classifiers. In this work, we go one step further and show that it is possible to successfully generate adversarial fake faces with a specified set of attributes (e.g., hair color, eye size, race, gender, etc.). To achieve this goal, we leverage the state-of-the-art generative model StyleGAN with disentangled representations, which enables a range of modifications without leaving the manifold of natural images. We propose a framework to search for adversarial latent codes within the feature space of StyleGAN, where the search can be guided either by a text prompt or a reference image. We also propose a meta-learning based optimization strategy to achieve transferable performance on unknown target models. Extensive experiments demonstrate that the proposed approach can produce semantically manipulated adversarial fake faces, which are true to the specified attribute set and can successfully fool forensic face classifiers, while remaining undetectable by humans. Code: https://github.com/koushiksrivats/face_attribute_attack. Fahad Shamshad, Koushik Srivatsan, Karthik Nandakumar |
CVPR | 3 |
| 2023 | CaPriDe Learning: Confidential and Private Decentralized Learning Based on Encryption-Friendly Distillation LossabstractLarge volumes of data required to train accurate deep neural networks (DNNs) are seldom available with any single entity. Often, privacy concerns prevent entities from sharing data with each other or with a third-party learning service provider. While crosssilo federated learning (FL) allows collaborative learning of large DNNs without sharing the data itself, most existing cross-silo FL algorithms have an unacceptable utility-privacy trade-off. In this work, we propose a framework called Confidential and Private Decentralized (CaPriDe) learning, which optimally leverages the power of fully homomorphic encryption (FHE) to enable collaborative learning without compromising on the confidentiality and privacy of data. In CaPridDe learning, participating entities release their private data in an encrypted form allowing other participants to perform inference in the encrypted domain. The crux of CaPriDe learning is mutual knowledge distillation between multiple local models through a novel distillation loss, which is an approximation of the Kullback-Leibler (KL) divergence between the local predictions and encrypted inferences of other participants on the same data that can be computed in the encrypted domain. Extensive experiments on three datasets show that CaPriDe learning can improve the accuracy of local models without any central coordination, provide strong guarantees of data confidentiality and privacy, and has the ability to handle statistical heterogeneity. Constraints on the model architecture (arising from the need to be FHE-friendly), limited scalability, and computational complexity of encrypted domain inference are the main limitations of the proposed approach. The code can be found at https://github.com/tnurbek/capride-learning. Nurbek Tastan, Karthik Nandakumar |
CVPR | 2 |
| 2023 | Towards Building Text-to-Speech Systems for the Next Billion UsersabstractDeep learning based text-to-speech (TTS) systems have been evolving rapidly with advances in model architectures, training methodologies, and generalization across speakers and languages. However, these advances have not been thoroughly investigated for Indian language speech synthesis. Such investigation is computationally expensive given the number and diversity of Indian languages, relatively lower resource availability, and the diverse set of advances in neural TTS that remain untested. In this paper, we evaluate the choice of acoustic models, vocoders, supplementary loss functions, training schedules, and speaker and language diversity for Dravidian and Indo-Aryan languages. Based on this, we identify monolingual models with FastPitch and HiFi-GAN V1, trained jointly on male and female speakers to perform the best. With this setup, we train and evaluate TTS models for 13 languages and find our models to significantly improve upon existing models in all languages as measured by mean opinion scores. We open-source all models on the Bhashini platform. Gokul Karthik Kumar, Praveen S. V, Mitesh M. Khapra, Karthik Nandakumar |
ICASSP | 5 |
| 2023 | Single-branch Network for Multimodal TrainingabstractWith the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval, matching, and verification. Existing works use separate networks to extract embeddings of each modality to bridge the gap between them. The modular structure of their branched networks is fundamental in creating numerous multimodal applications and has become a defacto standard to handle multiple modalities. In contrast, we propose a novel single-branch network capable of learning discriminative representation of unimodal as well as multimodal tasks without changing the network. An important feature of our single-branch network is that it can be trained either using single or multiple modalities without sacrificing performance. We evaluated our proposed single-branch network on the challenging multimodal problem (face-voice association) for cross-modal verification and matching tasks with various loss formulations. Experimental results demonstrate the superiority of our proposed single-branch network over the existing methods in a wide range of experiments. Code: https://github.com/msaadsaeed/SBNet Muhammad Saad Saeed, Shah Nawaz, Muhammad Haris Khan, Muhammad Zaigham Zaheer, Karthik Nandakumar, Muhammad Haroon Yousaf, Arif Mahmood |
ICASSP | 5 |
| 2023 | FedSIS: Federated Split Learning with Intermediate Representation Sampling for Privacy-preserving Generalized Face Presentation Attack DetectionabstractLack of generalization to unseen domains/attacks is the Achilles heel of most face presentation attack detection (FacePAD) algorithms. Existing attempts to enhance the generalizability of FacePAD solutions assume that data from multiple source domains are available with a single entity to enable centralized training. In practice, data from different source domains may be collected by diverse entities, who are often unable to share their data due to legal and privacy constraints. While collaborative learning paradigms such as federated learning (FL) can overcome this problem, standard FL methods are ill-suited for domain generalization because they struggle to surmount the twin challenges of handling non-iid client data distributions during training and generalizing to unseen domains during inference. In this work, a novel framework called Federated Split learning with Intermediate representation Sampling (FedSIS) is introduced for privacy-preserving domain generalization. In FedSIS, a hybrid Vision Transformer (ViT) architecture is learned using a combination of FL and split learning to achieve robustness against statistical heterogeneity in the client data distributions without any sharing of raw data (thereby preserving privacy). To further improve generalization to unseen domains, a novel feature augmentation strategy called intermediate representation sampling is employed, and discriminative information from intermediate blocks of a ViT is distilled using a shared adapter network. The FedSIS approach has been evaluated on two well-known benchmarks for cross-domain FacePAD to demonstrate that it is possible to achieve state-of-the-art generalization performance without data sharing. Code: https://github.com/Naiftt/FedSIS Naif Alkhunaizi, Koushik Srivatsan, Faris Almalik, Ibrahim Almakky, Karthik Nandakumar |
IJCB | 5 |
| 2023 | On Self-Supervised Learning and Prompt Tuning of Vision Transformers for Cross-sensor Fingerprint Presentation Attack DetectionabstractPresentation attacks pose a serious threat to the integrity of fingerprint-based biometric systems. Existing methods for fingerprint presentation attack detection (FpPAD) suffer from a lack of generalizability across different sensors and attack instruments, especially those that are not encountered during training. Recently, deep neural networks based on the Vision Transformer (ViT) architecture have demonstrated impressive generalization performance across many image recognition tasks due to their ability to effectively model long-range dependencies between image patches through the self-attention mechanism. While ViT models have been considered for FpPAD, many practical intricacies involved in learning generalizable ViTs for the FpPAD task have not been explored in depth. These include: (i) what is the best way to pre-process a fingerprint image to generate patches required by a ViT?, (ii) how to pre-train the ViT backbone to be used in FpPAD?, (iii) how to finetune the pre-trained ViT backbone for the FpPAD task?, and (iv) what is the most effective classifier design for a ViT-based FpPAD system? In this study, we undertake a thorough empirical study based on two public-domain datasets (LivDet 2015 and MSU-FPAD) in search of answers to the above questions. The key findings of this study are as follows: (i) Using minutia-aligned local patches provides the best PAD performance compared to partitioning the image into fixed number of non-overlapping patches. (ii) Self-supervised pre-training based on Masked Image Modeling (MIM) leads to better generalization performance than multimodal approaches such as image-text alignment. (iii) Visual prompt learning is a more effective way to adapt a pre-trained ViT model for FpPAD compared to full-tuning. (iv) Learning a linear classification head together with the visual prompts provides superior performance compared to linear probing and alignment with fixed text prompts. We hope that the above findings will be useful to the biometrics community and accelerate the deployment of ViT models for practical FpPAD systems. Maryam Nadeem, Karthik Nandakumar |
IJCB | 2 |
| 2023 | FLIP: Cross-domain Face Anti-spoofing with Language GuidanceabstractFace anti-spoofing (FAS) or presentation attack detection is an essential component of face recognition systems deployed in security-critical applications. Existing FAS methods have poor generalizability to unseen spoof types, camera sensors, and environmental conditions. Recently, vision transformer (ViT) models have been shown to be effective for the FAS task due to their ability to capture long-range dependencies among image patches. However, adaptive modules or auxiliary loss functions are often required to adapt pre-trained ViT weights learned on large-scale datasets such as ImageNet. In this work, we first show that initializing ViTs with multimodal (e.g., CLIP) pre-trained weights improves generalizability for the FAS task, which is in line with the zero-shot transfer capabilities of vision-language pre-trained (VLP) models. We then propose a novel approach for robust cross-domain FAS by grounding visual representations with the help of natural language. Specifically, we show that aligning the image representation with an ensemble of class descriptions (based on natural language semantics) improves FAS generalizability in low-data regimes. Finally, we propose a multimodal contrastive learning strategy to boost feature generalization further and bridge the gap between source and target domains. Extensive experiments on three standard protocols demonstrate that our method significantly outperforms the state-of-the-art methods, achieving better zero-shot transfer performance than five-shot transfer of "adaptive ViTs". Code: https://github.com/koushiksrivats/FLIP Koushik Srivatsan, Muzammal Naseer, Karthik Nandakumar |
ICCV | 3 |
| 2023 | FeSViBS: Federated Split Learning of Vision Transformer with Block Sampling
Faris Almalik, Naif Alkhunaizi, Ibrahim Almakky, Karthik Nandakumar |
MICCAI (2) | 4 |
| 2023 | DCTM: Dilated Convolutional Transformer Model for Multimodal Engagement Estimation in ConversationabstractConversational engagement estimation is posed as a regression problem, entailing the identification of the favorable attention and involvement of the participants in the conversation. This task arises as a crucial pursuit to gain insights into human's interaction dynamics and behavior patterns within a conversation. In this research, we introduce a dilated convolutional Transformer for modeling and estimating human engagement in the MULTIMEDIATE 2023 competition. Our proposed system surpasses the baseline models, exhibiting a noteworthy 7% improvement on test set and 4% on validation set. Moreover, we employ different modality fusion mechanism and show that for this type of data, a simple concatenated method with self-attention fusion gains the best performance. Vu Ngoc Tu, Van Thong Huynh, Soo-Hyung Kim, Shah Nawaz, Karthik Nandakumar, Muhammad Zaigham Zaheer |
ACM Multimedia | 6 |
| 2023 | Byzantine-Tolerant Methods for Distributed Variational InequalitiesabstractRobustness to Byzantine attacks is a necessity for various distributed training scenarios. When the training reduces to the process of solving a minimization problem, Byzantine robustness is relatively well-understood. However, other problem formulations, such as min-max problems or, more generally, variational inequalities, arise in many modern machine learning and, in particular, distributed learning tasks. These problems significantly differ from the standard minimization ones and, therefore, require separate consideration. Nevertheless, only one work [Abidi et al., 2022] addresses this important question in the context of Byzantine robustness. Our work makes a further step in this direction by providing several (provably) Byzantine-robust methods for distributed variational inequality, thoroughly studying their theoretical convergence, removing the limitations of the previous work, and providing numerical comparisons supporting the theoretical findings. Nazarii Tupitsa, Abdulla Jasem Almansoori, Yanlin Wu, Martin Takác 0001, Karthik Nandakumar, Samuel Horváth, Eduard Gorbunov |
NeurIPS | 5 |
| 2022 | On the Importance of Image Encoding in Automated Chest X-Ray Report Generation
Otabek Nazarov, Mohammad Yaqub, Karthik Nandakumar |
BMVC | 3 |
| 2022 | On Demographic Bias in Fingerprint RecognitionabstractFingerprint recognition systems have been deployed globally in numerous applications including personal devices, forensics, law enforcement, banking, and national identity systems. For these systems to be socially acceptable and trustworthy, it is critical that they perform equally well across different demographic groups. In this work, we propose a formal statistical framework to test for the existence of bias (demographic differentials) in fingerprint recognition across four major demographic groups (white male, white female, black male, and black female) for two state-of-the-art (SOTA) fingerprint matchers operating in verification and identification modes. Experiments on two different fingerprint databases (with 15,468 and 1,014 subjects) show that demographic differentials in SOTA fingerprint recognition systems decrease as the matcher accuracy increases and any small bias that may be evident is likely due to certain outlier, low-quality fingerprint images. Akash Godbole, Steven A. Grosz, Karthik Nandakumar, Anil K. Jain 0001 |
IJCB | 3 |
| 2022 | Suppressing Poisoning Attacks on Federated Learning for Medical Imaging
Naif Alkhunaizi, Dmitry Kamzolov, Martin Takác 0001, Karthik Nandakumar |
MICCAI (8) | 4 |
| 2022 | Self-Ensembling Vision Transformer (SEViT) for Robust Medical Image Classification
Faris Almalik, Mohammad Yaqub, Karthik Nandakumar |
MICCAI (3) | 3 |
| 2022 | How to Democratise and Protect AI: Fair and Differentially Private Decentralised Deep LearningabstractThis article first considers the research problem of fairness in collaborative deep learning, while ensuring privacy. A novel reputation system is proposed through digital tokens and local credibility to ensure fairness, in combination with differential privacy to guarantee privacy. In particular, we build a fair and differentially private decentralised deep learning framework called FDPDDL, which enables parties to derive more accurate local models in a fair and private manner by using our developed two-stage scheme: during the initialisation stage, artificial samples generated by Differentially Private Generative Adversarial Network (DPGAN) are used to mutually benchmark the local credibility of each party and generate initial tokens; during the update stage, Differentially Private SGD (DPSGD) is used to facilitate collaborative privacy-preserving deep learning, and local credibility and tokens of each party are updated according to the quality and quantity of individually released gradients. Experimental results on benchmark datasets under three realistic settings demonstrate that FDPDDL achieves high fairness, yields comparable accuracy to the centralised and distributed frameworks, and delivers better accuracy than the standalone framework. Lingjuan Lyu, Yitong Li 0002, Karthik Nandakumar, Jiangshan Yu, Xingjun Ma |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | Optimizing Homomorphic Encryption based Secure Image AnalyticsabstractData privacy is a growing concern as more cloud-based solutions for extracting insights from images are available. Fully Homomorphic Encryption (FHE) is one of the state-of-the-art techniques to enable privacy preserving machine learning. However, encrypted deep neural network-based inference methods face the fundamental challenge of optimizing computational depth and complexity to achieve an acceptable trade-off between accuracy and practical feasibility. Existing works only report high level implementations without providing rigorous analysis of the intricacies involved in the FHE implementation of generic primitive operators of a convolutional neural network (CNN). In this paper, we use the CKKS encryption scheme available in the open-source HElib library to run encrypted inference experiments on the MNIST dataset. The experiments indicate that efficient ciphertext packing schemes, model optimization and multi-threading strategies play a critical role in determining the throughput and latency of the inference process. We also show that operational parameters of the chosen FHE scheme such as the degree of the cyclotomic polynomial, depth limitations of the underlying leveled HE scheme, and the computational precision parameters result in significant trade-offs between accuracy, security level and computational time of the machine learning model. The key contribution of the paper is the analysis and recommendation of optimization techniques for efficient encrypted CNN inference. Nayna Jain, Karthik Nandakumar, Nalini K. Ratha, Sharath Pankanti, Uttam Kumar 0001 |
MMSP | 2 |
| 2020 | Cancelable Biometrics Vault: A Secure Key-Binding Biometric Cryptosystem based on Chaffing and WinnowingabstractExisting key-binding biometric cryptosystems, such as the Fuzzy Vault Scheme (FVS) and Fuzzy Commitment Scheme (FCS), employ Error Correcting Codes (ECC) to handle intra-user variations in biometric data. As a result, a trade-off exists between the key length and matching accuracy. Moreover, these systems are vulnerable to privacy leakage, i.e., it is trivial to recover the original biometric template given the secure sketch and its associated cryptographic key. In this work, we propose a novel key-binding biometric cryptosystem framework, referred to as Cancelable Biometrics Vault (CBV), to address the above two limitations. The CBV framework is inspired by the cryptographic principle of chaffing and winnowing. It utilizes the concept of cancelable biometrics (CB) to generate secure biometric templates, which in turn are used to encode bits in a cryptographic key. While the CBV framework is generic and does not rely on a specific biometric representation, it does assume the availability of a suitable (satisfying the requirements of accuracy preservation, non-invertibility, and non-linkability) CB scheme for the given representation. To demonstrate the usefulness of the proposed CBV framework, we implement this approach using an extended BioEncoding scheme, which is a CB scheme appropriate for bit strings such as iris-codes. Unlike the baseline BioEncoding scheme, the extended version proposed in this work fulfills all the three requirements of a CB construct. Experiments show that the decoding accuracy of the proposed CBV framework is comparable to the recognition accuracy of the underlying CB construct, namely, the extended BioEncoding scheme, regardless of the cryptographic key size. Osama Ouda, Karthik Nandakumar, Arun Ross |
ICPR | 2 |
| 2020 | Towards Fair and Privacy-Preserving Federated Deep ModelsabstractThe current standalone deep learning framework tends to result in overfitting and low utility. This problem can be addressed by either a centralized framework that deploys a central server to train a global model on the joint data from all parties, or a distributed framework that leverages a parameter server to aggregate local model updates. Server-based solutions are prone to the problem of a single-point-of-failure. In this respect, collaborative learning frameworks, such as federated learning (FL), are more robust. Existing federated learning frameworks overlook an important aspect of participation: fairness. All parties are given the same final model without regard to their contributions. To address these issues, we propose a decentralized Fair and Privacy-Preserving Deep Learning (FPPDL) framework to incorporate fairness into federated deep learning models. In particular, we design a local credibility mutual evaluation mechanism to guarantee fairness, and a three-layer onion-style encryption scheme to guarantee both accuracy and privacy. Different from existing FL paradigm, under FPPDL, each participant receives a different version of the FL model with performance commensurate with his contributions. Experiments on benchmark datasets demonstrate that FPPDL balances fairness, privacy and accuracy. It enables federated learning ecosystems to detect and isolate low-contribution parties, thereby promoting responsible participation. Lingjuan Lyu, Jiangshan Yu, Karthik Nandakumar, Yitong Li 0002, Xingjun Ma, Jiong Jin, Han Yu 0001, Kee Siong Ng |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Proving Multimedia Integrity using Sanitizable Signatures Recorded on BlockchainabstractWhile significant advancements have been made in the field of multimedia forensics to detect altered content, existing techniques mostly focus on enabling the content recipient to verify the content integrity without any inputs from the content creator. In many application scenarios, the creator has a strong incentive to establish the provenance and integrity of the multimedia data created and released by him. Hence, there is a strong need for mechanisms that allow the content creator to prove the authenticity of the released content. Since blockchain technology provides an immutable distributed database, it is an ideal solution for reliably time-stamping content with its creation time and storing an irrefutable signature of the content at the time of its creation. However, a simple digital signature scheme does not allow modification of the content after the initial commitment. Authorized multimedia content alteration by its creator is often necessary (e.g., redaction of faces to protect the privacy of individuals in a video, redaction of sensitive fields in a text document) before the content is distributed. The main contributions of this paper are: (i) a novel sanitizable signature scheme that enables the content creator to prove the integrity of the redacted content, while preventing the recipients from reconstructing the redacted segments based on the published commitment, and (ii) a blockchain-based solution for securely managing the sanitizable signature. The proposed solution employs a robust hashing scheme using chameleon hash function and Merkle tree to generate the initial signature, which is stored on the blockchain. The auxiliary data required for the integrity verification step is retained by the content creator and only a signature of this auxiliary data is stored on the blockchain. Any modifications to the multimedia content requires only updating the signature of the auxiliary data, which is securely recorded on the blockchain. We demonstrate that the proposed approach enables verification of integrity of redacted multimedia content without compromising the content privacy requirements. Karthik Nandakumar, Nalini K. Ratha, Sharath Pankanti |
IH&MMSec | 1 |
| 2018 | Double-Blind Consent-Driven Data Sharing on BlockchainabstractBlockchains are designed for trustworthy and transparent execution of transactions involving multiple parties. An important class of applications requires data to be shared selectively among mutually anonymous transacting peers while retaining the tamper-resistant evidentiary and validation features of a blockchain. KYC validations of corporate customers by banks is one example, where both banks and customers benefit from sharing process and data on a blockchain network. However, sharing of confidential KYC data must be authorized by customers, and a bank-customer relationship must be kept secret from other banks in the network. In this paper, we describe the design and implementation of a smart contract for consent-driven and double-blind data sharing on the Hyperledger Fabric blockchain platform. We show how a KYC application was built around this model to address the needs of the banks while meeting regulatory requirements. Kumar Bhaskaran, Peter Ilfrich, Dain Liffman, Christian Vecchiola, Praveen Jayachandran, Apurva Kumar, Fabian Lim, Karthik Nandakumar, Zhengquan Qin, Venkatraman Ramakrishna, Ernie G. S. Teo, Chun Hui Suen |
IC2E | 8 |
| 2018 | PPFA: Privacy Preserving Fog-Enabled Aggregation in Smart GridabstractFor constrained end devices in Internet of Things, such as smart meters (SMs), data transmission is an energy-consuming operation. To address this problem, we propose an efficient and privacy-preserving aggregation system with the aid of Fog computing architecture, named PPFA, which enables the intermediate Fog nodes to periodically collect data from nearby SMs and accurately derive aggregate statistics as the fine-grained Fog level aggregation. The Cloud/utility supplier computes overall aggregate statistics by aggregating Fog level aggregation. To minimize the privacy leakage and mitigate the utility loss, we use more efficient and concentrated Gaussian mechanism to distribute noise generation among parties, thus offering provable differential privacy guarantees of the aggregate statistic on both Fog level and Cloud level. In addition, to ensure aggregator obliviousness and system robustness, we put forward a two-layer encryption scheme: the first layer applies OTP to encrypt individual noisy measurement to achieve aggregator obliviousness, while the second layer uses public-key cryptography for authentication purpose. Our scheme is simple, efficient, and practical, it requires only one round of data exchange among a SM, its connected Fog node and the Cloud if there are no node failures, otherwise, one extra round is needed between a meter, its connected Fog node, and the trusted third party. Lingjuan Lyu, Karthik Nandakumar, Benjamin I. P. Rubinstein, Jiong Jin, Justin Bedo, Marimuthu Palaniswami |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | 50 years of biometric research: Accomplishments, challenges, and opportunities
Anil K. Jain 0001, Karthik Nandakumar, Arun Ross |
Pattern Recognit. Lett. | 2 |
| 2015 | Robust Representation and Recognition of Facial Emotions Using Extreme Sparse LearningabstractRecognition of natural emotions from human faces is an interesting topic with a wide range of potential applications, such as human-computer interaction, automated tutoring systems, image and video retrieval, smart environments, and driver warning systems. Traditionally, facial emotion recognition systems have been evaluated on laboratory controlled data, which is not representative of the environment faced in real-world applications. To robustly recognize the facial emotions in real-world natural situations, this paper proposes an approach called extreme sparse learning, which has the ability to jointly learn a dictionary (set of basis) and a nonlinear classification model. The proposed approach combines the discriminative power of extreme learning machine with the reconstruction property of sparse representation to enable accurate classification when presented with noisy signals and imperfect data recorded in natural settings. In addition, this paper presents a new local spatio-temporal descriptor that is distinctive and pose-invariant. The proposed framework is able to achieve the state-of-the-art recognition accuracy on both acted and spontaneous facial emotion databases. Seyedehsamaneh Shojaeilangari, Weiyun Yau, Karthik Nandakumar, Jun Li 0005, Eam Khwang Teoh |
IEEE Trans. Image Process. | 3 |
| 2013 | A multi-modal gesture recognition system using audio, video, and skeletal joint dataabstractThis paper describes the gesture recognition system developed by the Institute for Infocomm Research (I2R) for the 2013 ICMI CHALEARN Multi-modal Gesture Recognition Challenge. The proposed system adopts a multi-modal approach for detecting as well as recognizing the gestures. Automated gesture detection is performed using both audio signals and information about hand joints obtained from the Kinect sensor to segment a sample into individual gestures. Once the gestures are detected and segmented, features extracted from three different modalities, namely, audio, 2-dimensional video (RGB), and skeletal joints (Kinect) are used to classify a given sequence of frames into one of the 20 known gestures or an unrecognized gesture. Mel frequency cepstral coefficients (MFCC) are extracted from the audio signals and a Hidden Markov Model (HMM) is used for classification. While Space-Time Interest Points (STIP) are used to represent the RGB modality, a covariance descriptor is extracted from the skeletal joint data. In the case of both RGB and Kinect modalities, Support Vector Machines (SVM) are used for gesture classification. Finally, a fusion scheme is applied to accumulate evidence from all the three modalities and predict the sequence of gestures in each test sample. The proposed gesture recognition system is able to achieve an average edit distance of 0.2074 over the 275 test samples containing 2,742 unlabeled gestures. While the proposed system is able to recognize the known gestures with high accuracy, most of the errors are caused due to insertion, which occurs when an unrecognized gesture is misclassified as one of the 20 known gestures. Karthik Nandakumar, Kong-Wah Wan, Siu Man Alice Chan, Wen Zheng Terence Ng, Jian-Gang Wang 0001, Weiyun Yau |
ICMI | 1 |
| 2012 | Multibiometric Cryptosystems Based on Feature-Level FusionabstractMultibiometric systems are being increasingly de- ployed in many large-scale biometric applications (e.g., FBI-IAFIS, UIDAI system in India) because they have several advantages such as lower error rates and larger population coverage compared to unibiometric systems. However, multibiometric systems require storage of multiple biometric templates (e.g., fingerprint, iris, and face) for each user, which results in increased risk to user privacy and system security. One method to protect individual templates is to store only the secure sketch generated from the corresponding template using a biometric cryptosystem. This requires storage of multiple sketches. In this paper, we propose a feature-level fusion framework to simultaneously protect multiple templates of a user as a single secure sketch. Our main contributions include: (1) practical implementation of the proposed feature-level fusion framework using two well-known biometric cryptosystems, namery,fuzzy vault and fuzzy commitment, and (2) detailed analysis of the trade-off between matching accuracy and security in the proposed multibiometric cryptosystems based on two different databases (one real and one virtual multimodal database), each containing the three most popular biometric modalities, namely, fingerprint, iris, and face. Experimental results show that both the multibiometric cryptosystems proposed here have higher security and matching performance compared to their unibiometric counterparts. Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2010 | A hybrid biometric cryptosystem for securing fingerprint minutiae templates
Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001 |
Pattern Recognit. Lett. | 2 |
| 2008 | Securing fingerprint template: Fuzzy vault with minutiae descriptorsabstractFuzzy vault has been shown to be an effective technique for securing fingerprint minutiae templates. Its security depends on the difficulty in identifying the set of genuine minutiae points among a mixture of genuine and chaff points and reconstructing the secure polynomial using the evaluations (ordinate values) available for each point in the vault. We show that the security of fuzzy vault can be improved by ldquoencryptingrdquo these polynomial evaluations using a fuzzy commitment scheme. This encryption makes it difficult for an adversary to decode the vault even if the correct set of minutiae is selected. We use minutiae descriptors, which capture orientation and ridge frequency information in a minutiapsilas neighborhood, for securing the polynomial evaluations. This modification leads to a significant increase in both the security (number of tries an adversary has to make in order to guess the secure key) and matching accuracy of the vault. We validate our results on FVC2002 DB2 and show that false accept rate (FAR) is reduced from 0.7% to 0.01% at a genuine accept rate (GAR) of 95%. At the same time, vault security as measured in terms of min-entropy, is increased from 31 bits to 47 bits in case a perfect code is used. Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001 |
ICPR | 2 |
| 2008 | Likelihood Ratio-Based Biometric Score FusionabstractMultibiometric systems fuse information from different sources to compensate for the limitations in performance of individual matchers. We propose a framework for optimal combination of match scores that is based on the likelihood ratio test. The distributions of genuine and impostor match scores are modeled as finite Gaussian mixture model. The proposed fusion approach is general in its ability to handle (i) discrete values in biometric match score distributions, (ii) arbitrary scales and distributions of match scores, (iii) correlation between the scores of multiple matchers and (iv) sample quality of multiple biometric sources. Experiments on three multibiometric databases indicate that the proposed fusion framework achieves consistently high performance compared to commonly used score fusion techniques based on score transformation and classification. Karthik Nandakumar, Yi Chen 0015, Sarat C. Dass, Anil K. Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Fingerprint-Based Fuzzy Vault: Implementation and PerformanceabstractReliable information security mechanisms are required to combat the rising magnitude of identity theft in our society. While cryptography is a powerful tool to achieve information security, one of the main challenges in cryptosystems is to maintain the secrecy of the cryptographic keys. Though biometric authentication can be used to ensure that only the legitimate user has access to the secret keys, a biometric system itself is vulnerable to a number of threats. A critical issue in biometric systems is to protect the template of a user which is typically stored in a database or a smart card. The fuzzy vault construct is a biometric cryptosystem that secures both the secret key and the biometric template by binding them within a cryptographic framework. We present a fully automatic implementation of the fuzzy vault scheme based on fingerprint minutiae. Since the fuzzy vault stores only a transformed version of the template, aligning the query fingerprint with the template is a challenging task. We extract high curvature points derived from the fingerprint orientation field and use them as helper data to align the template and query minutiae. The helper data itself do not leak any information about the minutiae template, yet contain sufficient information to align the template and query fingerprints accurately. Further, we apply a minutiae matcher during decoding to account for nonlinear distortion and this leads to significant improvement in the genuine accept rate. We demonstrate the performance of the vault implementation on two different fingerprint databases. We also show that performance improvement can be achieved by using multiple fingerprint impressions during enrollment and verification. Karthik Nandakumar, Anil K. Jain 0001, Sharath Pankanti |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2005 | Score normalization in multimodal biometric systems
Anil K. Jain 0001, Karthik Nandakumar, Arun Ross |
Pattern Recognit. | 2 |