Akshay Agarwal 0001

dblp:152/3672-1 · DBLP profile ↗
← Back
50ranked-venue papers
18as first author
40since 2021 · last 2026
0000-0001-7362-4752ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 12 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 12 first-author · 31 since 2021Security and privacy · 8 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 WingBeats and Snapshots: Fusing Sound and Vision for Mosquito Monitoring (Student Abstract)
Ahana Chanda, Akshay Agarwal 0001
AAAI2
2026 Semantic-Guided Sketch-to-RGB Image Generation via Controlled Diffusion for Improved Sketch Recognition (Student Abstract)
abstract
Although deep networks excel on RGB images, their performance degrades sharply under severe domain shifts—such as sketch recognition, where color and texture cues are missing. In this work, we propose a novel pipeline that leverages semantic cues extracted from sketches to guide the synthesis of photorealistic RGB images using diffusion-based generative models. Our framework operates by extracting two crucial cues from the input sketch: semantic captions via the BLIP model and structural outlines via Canny edge detection. These cues are then integrated using ControlNet to guide a Stable Diffusion model, ensuring the synthesized RGB image is both semantically consistent with the content and structurally faithful to the original sketch. We evaluated our synthesized images by benchmarking classification performance. We trained standard architectures (from convolutional to transformer-based) on Tiny-ImageNet subsets and tested them on sketches, their synthesized counterparts, and the original RGB images. Experimental results demonstrate that our approach produces realistic, identity-preserving images, which significantly improve classification accuracy and effectively bridge the semantic gap. While BLIP-based captioning and ControlNet-guided diffusion are established methods, our contribution lies in their integration into a unified, caption-guided pipeline that enhances sketch-to-RGB translation with improved semantic consistency. The proposed method generalizes well across architectures, providing a scalable and cost-efficient solution for sketch-based image synthesis.
Ritika Jain, Akshay Agarwal 0001
AAAI3
2026 Q-MoFusion: A Quantum Classifier for Masquito Species Classification (Student Abstract)
abstract
Automated mosquito species identification is critical for combating vector-borne diseases. We introduce Q-MoFusion, a novel hybrid quantum-classical framework that fuses deep features from pre-trained Audio Spectrogram Transformer (AST) and Whisper models using a Variational Quantum Circuit (VQC). Our approach significantly outperforms individual backbones and prior state-of-the-art benchmarks, demonstrating superior accuracy and robustness, particularly on imbalanced classes. Q-MoFusion demonstrates the potential of hybrid quantum computing to enhance bioacoustic surveillance for addressing critical public health challenges.
Vishesh Kumar, Ahana Chanda, Poulomi Bhattacharya, Akshay Agarwal 0001
AAAI4
2026 Improving CAPTCHA Robustness via Controlled Image Corruptions (Student Abstract)
abstract
The Completely Automated Public Turing test to Tell Computers and Humans Apart (CAPTCHA) is widely deployed on the web as a security mechanism to distinguish humans from automated bots. However, their robustness is being challenged by the rapid advancements in AI, with models capable of near-human level character recognition rendering CAPTCHA obsolete. This research aims to systematically study the effect of multiple image corruptions, including elastic transformations, blur, noise, and occlusions, on human readability and automated solvers in text-based CAPTCHA recognition. We conduct experiments on multimodal large language models (MLLMs), a traditional deep learning-based optical character recognition (OCR) system, and human subjects. Using an existing CAPTCHA dataset and artificially corrupted versions, we analyze the recognition performance of AI models and humans, identifying vulnerabilities and patterns of robustness. The findings will contribute to a better understanding of CAPTCHA vulnerabilities and explore potential methods to increase the robustness of CAPTCHA in the era of advanced AI models.
Suchetan G. Uppur, Akshay Agarwal 0001
AAAI3
2026 Guarding Digital Identity: Attention-Guided Fusion for Detecting Forged ID Documents (Student Abstract)
abstract
Government verification systems are increasingly relying on internet-based platforms, where users authenticate their identities by uploading images captured with ordinary mobile devices. However, the rapid advancements in generative algorithms have enabled the creation of highly realistic forged ID cards that can easily bypass such verification pipelines. These forgeries are not restricted to a single modality; they may target facial imagery, textual content, or both, posing significant challenges to existing detection approaches. We present a framework that analyzes visual features for ID forgery detection by integrating feature fusion with attention mechanisms, leveraging both convolutional neural network (CNN) architectures, such as ResNet-50 and EfficientNet, and transformer-based models, including ViT-16 and Swin Transformer. This study emphasises the significance of feature fusion and attention-driven representation learning in developing robust and trustworthy ID forgery detection systems for real-world deployment.
Gargi S. Yeole, Poulomi Bhattacharya, Akshay Agarwal 0001
AAAI3
2025 A Unified, Resilient, and Explainable Adversarial Patch Detector
abstract
Deep Neural Networks (DNNs), backbone architecture in ‘almost’ every computer vision task, are vulnerable to adversarial attacks, particularly physical out-of-distribution (OOD) adversarial patches. Existing defense models often struggle with interpreting these attacks in ways that align with human visual perception. Our proposed AdvPatchXAI approach introduces a generalized, robust, and explainable defense algorithm designed to defend DNNs against physical adversarial threats. AdvPatchXAI employs a novel patch decorrelation loss that reduces feature redundancy and enhances the distinctiveness of patch representations, enabling better generalization across unseen adversarial scenarios. It learns prototypical parts self-supervised, enhancing interpretability and correlation with human vision. The model utilizes a sparse linear layer for classification, making the decision process globally interpretable through a set of learned prototypes and locally explainable by pinpointing relevant prototypes within an image. Our comprehensive evaluation shows that AdvPatchXAI closes the "semantic" gap between latent space and pixel space and effectively handles unseen adversarial patches even perturbed with unseen corruptions, thereby significantly advancing DNN robustness in practical settings1.
Vishesh Kumar, Akshay Agarwal 0001
CVPR2
2025 Your Face, Your Privacy: Combating Unauthorized Usage
abstract
The high performance of current deep face recognition systems and their unauthorized usage have raised a severe concern for privacy in the physical, adversarial, and digital domains. To protect privacy, users are exploring several ways, and one such method that recently gained attention is individuals deliberately obscuring their faces with their hands, presumably to avoid facial recognition technology. Since deep face recognition algorithms can handle partial tampering of faces, this raises a critical question of whether these deliberate attempts can protect privacy. In the literature, no evaluation exists that showcases that this type of hiding can bypass the face recognition algorithms. Therefore, in this first-ever study, we have performed extensive research by first developing multiple nose and mouth occlusion datasets using synthetic patches and real-life objects. Our extensive experimentation reveals several interesting observations reflecting the fact that even when a patch is a face patch extracted from an unseen subject, it can fool the face recognition networks. Further, not only face recognition networks, but also it is observed that the proposed patches are effective in deceiving the soft biometric classifier, i.e., the classifier detecting the gender and ethnicity of individuals.
Akshay Agarwal 0001, Nalini K. Ratha
FG2
2025 Gesture Recognition for Emergencies: Dataset and Cross-Condition Analysis
abstract
Gestures serve as a powerful medium of nonverbal communication, facilitating the transmission of a wide range of feelings. Hand gestures, in particular, represent a type of nonverbal communication utilized in various areas, including communication among deaf individuals, home automation, and medical applications. This makes gesture recognition an essential focus of research and development in computer vision, robotics, and human-computer interaction (HCI). The primary limitations of the existing research are the limited exploration of Indian ethnicity and the gesture captured in nighttime uncontrolled settings. In response, we have curated a large-scale hand gesture recognition dataset comprising signs frequently used in emergency scenarios. A benchmark study using the proposed dataset has been performed using state-of-the-art hand detection and gesture classification networks through in-the-wild evaluation settings.
Jiya Sinha, Poulomi Bhattacharya, Akshay Agarwal 0001
FG3
2025 Family Resemblance or Fraud? Face Morphing Attacks on Kinship Verification
abstract
Kinship verification using facial images is widely applied in forensic analysis, immigration, and child trafficking prevention. However, deep learning models for kinship verification are vulnerable to morphed images, where the facial features of two individuals are blended to create realistic but fake images. This work investigates the influence of different morphing ratios (95% child + 5% random parent to 50% child + 50% random parent) on kinship verification algorithms. Developing a morphed dataset allows us to experiment with deep learning and kinship-specific models on original and morphed child images to determine the threshold beyond which non-kin morphs are labeled kin. The experiments show a continuous rise in misclassification rates with the increasing percentage of parental features in morphed images, underscoring the difficulties encountered by current kinship verification systems. It is to be noted that the current study is the first to present significant insights into the vulnerability of existing kinship verification models against different morph ratios. It highlights the necessity for more effective verification methods to counter the risks associated with facial morphing in real-world applications.
Gargi S. Yeole, S. Aarthi, Shalvika Srivastav, Akshay Agarwal 0001
FG4
2025 Unmasking the Audio Illusion: A Survey on Spoofing and Deepfake Detection
abstract
With breakthroughs in deep learning algorithms, the practice of manipulating audio to produce believable fakes is expanding rapidly. This survey paper provides a comprehensive overview of the current state of deepfake audio research, encompassing generation methods, online platforms to generate fake audio, the latest detection techniques, human perception of fake audio, and the underlying security concerns. We examine different methods for speech synthesis, audio splicing, and voice cloning, pointing out their advantages and disadvantages. Furthermore, we investigate various detection algorithms, encompassing supervised, unsupervised, and hybrid techniques, and assess their effectiveness in detecting audio manipulation. We review deepfake audio’s impacts, including possible adverse effects on reputation, fraud, and misinformation. We present a concise analysis of AI versus human detection of deepfake audio, drawing insights from existing literature and validating them through our experiments. Finally, we highlight future research directions and recommendations for mitigating the societal risks associated with this powerful technology.
S. Aarthi, Akshay Agarwal 0001
IJCB2
2025 On Which Data Distribution (Synthetic or Real) We Should Rely for Soft Biometric Classification
abstract
Identification of gender is critical not only for human-computer interaction but also for scrutinizing the search space in which an identity needs to be determined. Traditionally, “real” facial images are employed for gender identification by computer vision algorithms. Due to the tremendous rise of privacy and advancement in generative networks, synthetic face images are heavily developed and can be used for several face-related studies including gender classification. However, their effectiveness compared to real images is still unexplored for gender classification. In response, this study explores the effectiveness of gender classification networks trained on real and synthetic face images, offering novel insights into the effectiveness of these two data distributions. For that, we implemented several state-of-the-art gender classification architectures covering convolutional neural networks (CNNs) and vision transformers (ViT). Our research builds on the rigorous evaluation of 8 Deep Neural Networks (DNNs) across 4 diverse datasets and 6 types of image corruptions. To make the research interpretable, we have also used several explainable mechanisms, including Grad-CAM and t-SNE visualizations. In brief, the impact of the proposed research is multifold: (i) understand the effectiveness of real vs. synthetic data distributions in network training and (ii) whether the synthetic models reflect the true physical world distribution to ensure that the models trained on them are resilient against image perturbations.
Manju R. A, Akshay Agarwal 0001
WACV3
2025 Detection of identity swapping attacks in low-resolution image settings
Akshay Agarwal 0001, Nalini K. Ratha
J. Inf. Secur. Appl.1
2025 Robustness Benchmarking of Convolutional and Transformer Architectures for Image Classification
abstract
Amidst the burgeoning landscape of thousands of deep neural networks (DNNs), selecting a robust architecture for image classification poses a formidable challenge. The prime reason is the vulnerability of these DNNs to image corruption. Surprisingly, the literature still does not understand which network is sensitive to which kind of corruption and to what extent. Our study rigorously analyzes DNNs across the classification spectrum, from pure convolutional neural networks (CNNs: without attention layers) to state-of-the-art Vision Transformers (ViTs). To reach a robust conclusion, we have performed extensive experiments using 18 DNNs, 5 diverse datasets, and 15 corruption types of varying severity. Our analysis uncovers insightful and surprising findings concerning the robustness of different DNNs across corruption types. For example, it is observed that the ViT trained on CIFAR10 is found to be highly robust in handling noise corruptions, even of considerable severity (say, severity 4), but is found vulnerable to environmental corruptions of the same severity. At the same time, while the performance of pure CNNs is lower than that of ViT, they can handle environmental factors better than noise corruption. Another interesting observation showcased that these corruptions can even fool one of the popular network explainability algorithms, Grad-CAM. The heat map highlights the same region of interest on the noisy image, similar to clean images, but the class label differs. Shedding light on the vulnerabilities of deep learning models and revealing their vulnerability against particular corruption ensures the deployment of a correct network in the real world that deals with specific corruption frequently.
Vishesh Kumar, Shivam Shukla, Akshay Agarwal 0001
IEEE Trans. Big Data3
2024 Deepfake: Classifiers, Fairness, and Demographically Robust Algorithm
abstract
Deepfake detection research has seen tremendous success and has achieved remarkably high performance on a few existing datasets. However, the significant drawback of the existing works is the generalizability of the detection algorithms under cross-datasets and cross-attack/manipulation settings. On top of that, another critical bottleneck of deepfake detection literature is the understanding of the fairness quotient of these algorithms. One big reason for such a less explored domain is the unavailability of deep fake datasets covering multiple ethnicities and genders with proper annotations. For example, the popular deepfake detection datasets such as FaceForensics++ and Celeb-DF are highly biased toward Caucasian ethnicity. Recently, a multi-ethnicity multi-modal dataset namely FakeAVCeleb has been released which can fulfill this gap. Henceforth by utilizing the potential of this dataset, we have performed the fairness study of deepfake detection algorithms. For that, several image classifiers are selected which range from deep convolutional neural networks to handcrafted image feature extraction to vision transformers. The experiments performed using such a wide variety of classifiers reveal that the deepfake detectors are not fair and can detect one ethnicity with high accuracy but fail miserably on others. For instance, the performance of one of the popular deepfake detection networks namely XceptionNet shows a reduction of more than 30% when dealing with different ethnicities and genders. Not only ethnicity or gender but also the type of classifiers have a huge impact on the performance. We assert that the proposed study can help in building a fair, robust, and accurate deepfake classifier utilizing insightful findings that can help in the selection of an effective and robust backbone architecture.
Akshay Agarwal 0001, Nalini K. Ratha
FG1
2024 Benchmarking In-the-wild Soft Biometric Attribute Identification
abstract
The active involvement of different demographic entities, whether positive or negative, demands effective identification of diverse populations. One way to quickly perform this identification is by segregating individuals based on their soft biometric attributes such as race and gender. Further, the soft biometric recognition technology has the potential to enhance individual privacy by providing a less invasive identification mode than conventional biometric methods. For example, a specific disbursement of government benefits might not require the exact identity of a person but only soft biometric attributes for its actual distribution. Secondly, the correct identification of race and gender can restrict the identity search space in which we need to look for an imposter. However, despite tremendous literature on identity recognition, limited work has been done to identify soft biometric attributes in challenging evaluation settings accurately. Therefore, in this research, for the first time, we have performed extensive experiments to benchmark the effectiveness of pure convolutional networks and attention networks to identify gender and ethnicity in the in-distribution and out-of-distribution (OOD) settings. We employed diverse face datasets to benchmark our methods, including UTKFace and FairFace. In contrast to the general understanding that OOD images lead to poor performance, we observe significant performance in identifying soft biometric attributes, including race and gender classification.
Manju R. A, Akshay Agarwal 0001
IJCB2
2024 Enhancing Drug Abuse Face Recognition: A Study on Image Corruption and Restoration
abstract
Drug abuse poses a pressing societal issue and limits the success of face recognition algorithms. Further, when the faces are acquired in an unconstrained environment, they get affected by several environmental factors including snow, noise, and blur. One prominent limitation of the existing research is that no work has been done to mitigate the impact of drug abuse on deep face recognition. Further, the problem became complicated in the presence of input corruption. Through this study, we aim to fill the gap by investigating the effects of drug-induced facial alterations coupled with input corruptions on deep face recognition. Our study indicates that drug alteration drastically alters facial features and on top of that the corruption further degrades the image quality and hence leads to poor performance by deep face recognition networks. Therefore, the one question we asked inspired by the success of generative networks is whether these models can help mitigate the impact of drug abuse and image corruption to boost performance. Through experiments simulating real-world scenarios, we demonstrate several interesting observations about the success of generative models including diffusion and generative architectures. The efficacy of these models is not only observed for face recognition but also for the improvement of gender classification performance. To the best of our knowledge mitigation of illicit drug abuse is not explored in the literature and hence, this research contributes to enhancing the reliability of face recognition systems amidst the challenges posed by drug abuse, thus advancing both understanding and application in practical settings.
Hruturaj Dhake, Akshay Agarwal 0001
IJCB2
2024 Is Face Super Resolution Truly Pushing the Boundaries of Face Recognition?
abstract
With the improving efficacy of generative algorithms, the performance of face super-resolution algorithms is also increasing towards generating high-quality facial data. However, are these images useful for face recognition? This paper investigates whether these enhanced super-resolved facial images only improve the visual quality or they also aid in improving the recognizability of these images, thus contributing towards addressing the challenge of low-resolution face recognition. We conduct a comprehensive empirical and statistical analysis of human perception and face recognition tasks. Extensive experiments are performed using multiple state-of-the-art generative and face recognition models across six publicly available face datasets to assess whether face super-resolution algorithms are effective in recognizing individuals in low-resolution conditions. The results and supporting analysis indicate that the ability of super-resolution images to improve recognizability is limited, and further research is required to design generative AI algorithms that improve both visual appearance and recognizability of low-resolution images.
Muskan Dosi, Udaybhan Rathore, Chiranjeev Chiranjeev, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
IJCB4
2024 Are Object Recognition Models Effective and Unbiased for Biometric Recognition?
abstract
Can the general-purpose pre-trained image classifier be effective for biometric recognition? Due to the prevalence of generic features in these models, we assert that they can be used to extract the discriminating feature of the biometrics images which help encode the object images. The utilization of features from the pre-trained model can significantly limit the training requirements and can be deployed on computationally limited devices. In this research, we have conducted an extensive study to perform biometric recognition using existing state-of-the-art pre-trained deep neural networks. The experiments are not limited to any modality or imaging spectrum to ensure the findings are trustworthy and thorough. The experiments conducted on various modalities including multi-modal biometric recognition as well showcase that the pre-trained deep networks can be a suitable option for cost-effective biometric recognition. Further, we have also evaluated the fusion of deep neural networks to see if the performance can further be improved without significantly increasing the computational cost.
Vishesh Kumar, Akshay Agarwal 0001
IJCB2
2024 Face Morphing Detection in Social Media Content
abstract
Face being an active medium of communication is a significant part of our social media life; however, faces are vulnerable to manipulations. Among various manipulations, face morphing is a well-known tampering technique that aims to generate images containing information from more than one identity. Morphed images are heavily used for various malicious purposes including sarcasm, money laundering, and pornography. For many of the above harmful purposes, these manipulated images are uploaded on social media platforms where they can further go through tampering using social-media filters. Interestingly, the existing morph attack detection works have not addressed social media’s impact on deceiving face morph detectors. In this research, for the first time, we have generated authentic (or real) and face-morphed images impacted by one of the premium features of social media platforms known as filtering. We have used 13 Instagram filters and performed an extensive study on the proposed social-media morphed dataset. It is demonstrated that these filters can radically reduce the morph detection performances of several popular deep-learning classifiers. Therefore, to effectively address the concerns of face morphing and social media filtering, we propose a robust ViT-CNN architecture to advance the morph image detection performance.
Akshay Agarwal 0001, Nalini K. Ratha
ICIP1
2024 On the Effectiveness of a Hybrid Model for Volatility Prediction
abstract
Volatility prediction of the stock market has been a hot topic in economics and finance. However, it is challenging to forecast the volatility accurately due to the complex and dynamic nature of the stock market. It is observed that deep neural network architectures are universal approximators and effective for financial tasks; moreover, a single model is generally not found effective in solving the volatile nature of the input. Therefore, we propose a hybrid model that uniquely combines the strength of two state-of-the-art deep network models to estimate stock volatility precisely. The proposed hybrid model is evaluated on several real-world datasets covering two stock indexes and top volatile bank stock. The proposed algorithm surpasses the existing and state-of-the-art uni models and showcases its strength in effectively capturing the volatility of different stocks.
SaiAsrith EVNM Baddepudi, Akshay Agarwal 0001
ICMLA2
2024 Robustness of Classifiers for AI-Generated Text Detectors for Copyright and Privacy Protected Society
Akshay Agarwal 0001, Mohammed Uzair
ICPR (20)1
2024 Restoring Noisy Images Using Dual-Tail Encoder-Decoder Signal Separation Network
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
ICPR (1)1
2024 Supervised Mixup: Protecting the Likely Classes for Adversarial Robustness
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
ICPR (5)1
2024 A Multi-modal Framework to Counter Hate Speeches
Kirtilekha Bhesra, Akshay Agarwal 0001
ICPR (31)2
2024 An Unconstrained Dataset for Face Recognition Across Distance, Pose, and Resolution
Udaybhan Rathore, Akshay Agarwal 0001
ICPR (14)2
2024 Neural Encoding of Odors: Translating Odors into Unique Digital Representation with EEG Signals
Archana Yadav, Vishakha Pareek, Akshay Agarwal 0001, Santanu Chaudhury
ICPR (7)3
2024 Corruption depth: Analysis of DNN depth for misclassification
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
Neural Networks1
2024 Indian Traffic Sign Detection and Classification Through a Unified Framework
abstract
Traffic sign boards are vital in facilitating smart transportation systems. More than 90% of accidents happen due to drivers’ inattentiveness over these boards. Hence, relaying traffic sign board information automatically to drivers becomes crucial to avoid such accidents. While numerous traffic sign detection and classification systems exist, it is important to note that these automated systems have not been adequately assessed within the context of Indian settings. The task of traffic sign detection and classification presents unique challenges in the Indian context due to the presence of multiple variations of signboards for a single action (e.g., left turn and right turn). This paper proposes a novel deep neural network architecture for detecting and classifying traffic signboards simultaneously. In the proposed architecture, several state-of-the-art convolutional neural networks such as AlexNet, VGG-19, ResNet-50, and EfficientNet v2 are used as a backbone to solve the task of bounding box regression and sign classification. Along with the proposed method, we collected the Indian traffic signs and information boards dataset. The collected dataset consists of 4257 raw images without any augmentation. We comprehensively evaluated the proposed algorithm on the collected dataset, resulting in the detection and classification accuracies of 85.5% and 98.5%, respectively. The proposed algorithm has also been evaluated on the existing traffic sign recognition dataset, and a comparison with state-of-the-art algorithms is done for both traffic sign detection and classification, demonstrating the effectiveness of the proposed multi-task network.
Rishabh Uikey, Haroon R. Lone, Akshay Agarwal 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Misclassifications of Contact Lens Iris PAD Algorithms: Is it Gender Bias or Environmental Conditions?
abstract
One of the critical steps in biometrics pipeline is detection of presentation attacks, a physical adversary. Several presentation (adversary) attack detection (PAD) algorithms, including iris PAD, have been proposed and have shown superlative performance. However, a recent study, on a small-scale database, has highlighted that iris PAD may have gender biases. In this research, we present a rigorous study on gender bias in iris presentation attack detection algorithms using a large-scale and gender-balanced database. The paper provides several interesting observations which can help in building future presentation attack detection algorithms with aim of fair treatment of each demography. In addition, we also present a robust iris presentation attack detection algorithm by combining gender-covariate based classifiers. The proposed robust classifier not only reduces the difference in accuracy between different genders but also improves the overall performance of the PAD system.
Akshay Agarwal 0001, Nalini K. Ratha, Afzel Noore, Richa Singh 0001, Mayank Vatsa
WACV1
2022 Robust IRIS Presentation Attack Detection Through Stochastic Filter Noise
abstract
The vulnerability of iris recognition algorithms against presentation attacks demands a robust defense mechanism. Much research has been done in the literature to create a robust attack detection algorithm; however, most of the algorithms suffer from generalizability, such as inter database testing or unseen attack type. The problem of attack detection can further be exacerbated if the images contain noise such as Gaussian or Salt-Pepper noise. In this research, we propose a multi-task deep learning model with a denoising convolutional skip autoencoder and a classifier to inbuilt robustness against noisy images. The Gaussian noise layer is introduced as a dropout between the encoder network’s hidden layers, which helps the model learn generalized features that are robust to data noise. The proposed algorithm is evaluated on multiple presentation attack databases and extensive experiments across different noise types and a comparison with other deep learning models show the generalizability and efficacy of the proposed model.
Vishi Jain, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha
ICPR2
2022 Enhanced iris presentation attack detection via contraction-expansion CNN
Akshay Agarwal 0001, Afzel Noore, Mayank Vatsa, Richa Singh 0001
Pattern Recognit. Lett.1
2022 Crafting Adversarial Perturbations via Transformed Image Component Swapping
abstract
Adversarial attacks have been demonstrated to fool the deep classification networks. There are two key characteristics of these attacks: firstly, these perturbations are mostly additive noises carefully crafted from the deep neural network itself. Secondly, the noises are added to the whole image, not considering them as the combination of multiple components from which they are made. Motivated by these observations, in this research, we first study the role of various image components and the impact of these components on the classification of the images. These manipulations do not require the knowledge of the networks and external noise to function effectively and hence have the potential to be one of the most practical options for real-world attacks. Based on the significance of the particular image components, we also propose a transferable adversarial attack against unseen deep networks. The proposed attack utilizes the projected gradient descent strategy to add the adversarial perturbation to the manipulated component image. The experiments are conducted on a wide range of networks and four databases including ImageNet and CIFAR-100. The experiments show that the proposed attack achieved better transferability and hence gives an upper hand to an attacker. On the ImageNet database, the success rate of the proposed attack is up to 88.5%, while the current state-of-the-art attack success rate on the database is 53.8%. We have further tested the resiliency of the attack against one of the most successful defenses namely adversarial training to measure its strength. The comparison with several challenging attacks shows that: (i) the proposed attack has a higher transferability rate against multiple unseen networks and (ii) it is hard to mitigate its impact. We claim that based on the understanding of the image components, the proposed research has been able to identify a newer adversarial attack unseen so far and unsolvable using the current defense mechanisms.
Akshay Agarwal 0001, Nalini K. Ratha, Mayank Vatsa, Richa Singh 0001
IEEE Trans. Image Process.1
2022 DAMAD: Database, Attack, and Model Agnostic Adversarial Perturbation Detector
abstract
Adversarial perturbations have demonstrated the vulnerabilities of deep learning algorithms to adversarial attacks. Existing adversary detection algorithms attempt to detect the singularities; however, they are in general, loss-function, database, or model dependent. To mitigate this limitation, we propose DAMAD-a generalized perturbation detection algorithm which is agnostic to model architecture, training data set, and loss function used during training. The proposed adversarial perturbation detection algorithm is based on the fusion of autoencoder embedding and statistical texture features extracted from convolutional neural networks. The performance of DAMAD is evaluated on the challenging scenarios of cross-database, cross-attack, and cross-architecture training and testing along with traditional evaluation of testing on the same database with known attack and model. Comparison with state-of-the-art perturbation detection algorithms showcase the effectiveness of the proposed algorithm on six databases: ImageNet, CIFAR-10, Multi-PIE, MEDS, point and shoot challenge (PaSC), and MNIST. Performance evaluation with nearly a quarter of a million adversarial and original images and comparison with recent algorithms show the effectiveness of the proposed algorithm.
Akshay Agarwal 0001, Gaurav Goswami, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
IEEE Trans. Neural Networks Learn. Syst.1
2021 Role of Optimizer on Network Fine-tuning for Adversarial Robustness (Student Abstract)
abstract
The solutions proposed in the literature for adversarial robustness are either not effective against the challenging gradient-based attacks or are computationally demanding, such as adversarial training. Adversarial training or network training based data augmentation shows the potential to increase the adversarial robustness. While the training seems compelling, it is not feasible for resource-constrained institutions, especially academia, to train the network from scratch multiple times. The two fold contributions are: (i) providing an effective solution against white-box adversarial attacks via network fine-tuning steps and (ii) observing the role of different optimizers towards robustness. Extensive experiments are performed on a range of databases, including Fashion-MNIST and a subset of ImageNet. It is found that the few steps of network fine-tuning effectively increases the robustness of both shallow and deep architectures. To know other interesting observations, especially regarding the role of the optimizer, refer to the paper.
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001
AAAI1
2021 Detection of Digital Manipulation in Facial Images (Student Abstract)
abstract
Advances in deep learning have enabled the creation of photo-realistic DeepFakes by switching the identity or expression of individuals. Such technology in the wrong hands can seed chaos through blackmail, extortion, and forging false statements of influential individuals. This work proposes a novel approach to detect forged videos by magnifying their temporal inconsistencies. A study is also conducted to understand role of ethnicity bias due to skewed datasets on deepfake detection. A new dataset comprising forged videos of Indian ethnicity individuals is presented to facilitate this study.
Aman Mehra, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001
AAAI2
2021 MD-CSDNetwork: Multi-Domain Cross Stitched Network for Deepfake Detection
abstract
The rapid progress in the ease of creating and spreading ultra-realistic media over social platforms calls for an urgent need to develop a generalizable deepfake detection technique. It has been observed that current deepfake generation methods leave discriminative artifacts in the frequency spectrum of fake images and videos. Inspired by this observation, in this paper, we present a novel approach, termed as MD-CSDNetwork, for combining the features in the spatial and frequency domains to extract a shared discriminative representation for classifying deepfakes. MD-CSDNetwork is a novel cross-stitched network with two parallel branches carrying spatial and frequency information, respectively. We hypothesize that these multi-domain input data streams can be considered as related supervisory signals and can ensure better performance and generalization. Further, the concept of cross-stitch connections is utilized where they are inserted between the two branches to learn an optimal combination of domain-specific and shared representations from other domains automatically. Extensive experiments are conducted on the popular benchmark datasets. We report improvements over all the manipulation types in the FaceForensics++ dataset and comparable results with state-of-the-art methods for cross-database evaluation on the Celeb-DF dataset and the Deepfake Detection Dataset.
Aayushi Agarwal 0001, Akshay Agarwal 0001, Sayan Sinha, Mayank Vatsa, Richa Singh 0001
FG2
2021 When Sketch Face Recognition Meets Mask Obfuscation: Database and Benchmark
abstract
During this unprecedented time of the COVID19 pandemic, wearing face masks has become a necessity. While these masks aim to secure an individual from getting infected by any kind of viruses including COVID-19; they significantly obfuscate the identity. The situation becomes even worse when an attacker performs a crime and the place does not have any surveillance cameras. The identification of criminals in such conditions highly depends on the witnesses and generation of sketches based on their description. To the best of our knowledge, in the literature, no work has been performed for matching sketch images with masks. In this research, we have first created the mask sketch face database using more than 50 identities. The sketch images are generated using a different variant of pencils, which can be seen as different sketch artists. The recognition experiments are performed using state-of-the-art face embedding networks including ArcFace and DeepID which show that the recognition performance degrades significantly when the sketch mask images are used for identification. In another set of experiments, it is observed that the recognition algorithm is robust in handling the digital face mask images. However, the ineffectiveness in handling the variations that occurred due to sketches is a serious concern and needs attention.
Akshay Agarwal 0001, Nalini K. Ratha, Mayank Vatsa, Richa Singh 0001
FG1
2021 Intelligent and Adaptive Mixup Technique for Adversarial Robustness
abstract
Deep neural networks are generally trained using large amounts of data to achieve state-of-the-art accuracy in many possible computer vision and image analysis applications ranging from object recognition to natural language processing. It is also claimed that these networks can memorize the data which can be extracted from the network parameters such as weights and gradient information. The adversarial vulnerability of the deep networks is usually evaluated on the unseen test set of the databases. If the network is memorizing the data, then the small perturbation in the training image data should not drastically change its performance. Based on this assumption, we first evaluate the robustness of deep neural networks on small perturbations added in the training images used for learning the parameters of the network. It is observed that, even if the network has seen the images it is still vulnerable to these small perturbations. Further, we propose a novel data augmentation technique to increase the robustness of deep neural networks to such perturbations.
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
ICIP1
2021 Cognitive data augmentation for adversarial defense via pixel masking
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
Pattern Recognit. Lett.1
2021 Image Transformation-Based Defense Against Adversarial Perturbation on Deep Learning Models
abstract
Deep learning algorithms provide state-of-the-art results on a multitude of applications. However, it is also well established that they are highly vulnerable to adversarial perturbations. It is often believed that the solution to this vulnerability of deep learning systems must come from deep networks only. Contrary to this common understanding, in this article, we propose a non-deep learning approach that searches over a set of well-known image transforms such as Discrete Wavelet Transform and Discrete Sine Transform, and classifying the features with a support vector machine-based classifier. Existing deep networks-based defense have been proven ineffective against sophisticated adversaries, whereas image transformation-based solution makes a strong defense because of the non-differential nature, multiscale, and orientation filtering. The proposed approach, which combines the outputs of two transforms, efficiently generalizes across databases as well as different unseen attacks and combinations of both (i.e., cross-database and unseen noise generation CNN model). The proposed algorithm is evaluated on large scale databases, including object database (validation set of ImageNet) and face recognition (MBGC) database. The proposed detection algorithm yields at-least 84.2% and 80.1% detection accuracy under seen and unseen database test settings, respectively. Besides, we also show how the impact of the adversarial perturbation can be neutralized using a wavelet decomposition-based filtering method of denoising. The mitigation results with different perturbation methods on several image databases demonstrate the effectiveness of the proposed method.
Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha
IEEE Trans. Dependable Secur. Comput.1
2020 On the Robustness of Face Recognition Algorithms Against Attacks and Bias
abstract
Face recognition algorithms have demonstrated very high recognition performance, suggesting suitability for real world applications. Despite the enhanced accuracies, robustness of these algorithms against attacks and bias has been challenged. This paper summarizes different ways in which the robustness of a face recognition algorithm is challenged, which can severely affect its intended working. Different types of attacks such as physical presentation attacks, disguise/makeup, digital adversarial attacks, and morphing/tampering using GANs have been discussed. We also present a discussion on the effect of bias on face recognition models and showcase that factors such as age and gender variations affect the performance of modern algorithms. The paper also presents the potential reasons for these challenges and some of the future research directions for increasing the robustness of face recognition models.
Richa Singh 0001, Akshay Agarwal 0001, Maneet Singh, Shruti Nagpal, Mayank Vatsa
AAAI2
2020 Attack Agnostic Adversarial Defense via Visual Imperceptible Bound
abstract
The high susceptibility of deep learning algorithms against structured and unstructured perturbations has motivated the development of efficient adversarial defense algorithms. However, the lack of generalizability of existing defense algorithms and the high variability in the performance of the attack algorithms for different databases raises several questions on the effectiveness of the defense algorithms. In this research, we aim to design a defense model that is robust within a certain bound against both seen and unseen adversarial attacks. This bound is related to the visual appearance of an image, and we termed it as Visual Imperceptible Bound (VIB). To compute this bound, we propose a novel method that uses the database characteristics. The VIB is further used to measure the effectiveness of attack algorithms. The performance of the proposed defense model is evaluated on the MNIST, CIFAR-10, and Tiny ImageNet databases on multiple attacks that include C&W ( l2) and DeepFool. The proposed defense model is not only able to increase the robustness against several attacks but also retain or improve the classification accuracy on an original clean test set. The proposed algorithm is attack agnostic, i.e. it does not require any knowledge of the attack algorithm.
Saheb Chhabra, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
ICPR2
2020 Generalized Iris Presentation Attack Detection Algorithm under Cross-Database Settings
abstract
Presentation attacks are posing major challenges to most of the biometric modalities. Iris recognition, which is considered as one of the most accurate biometric modality for person identification, has also been shown to be vulnerable to advanced presentation attacks such as 3D contact lenses and textured lens. While in the literature, several presentation attack detection (PAD) algorithms are presented; a significant limitation is the generalizability against an unseen database, unseen sensor, and different imaging environment. To address this challenge, we propose a generalized deep learning-based PAD network, MVANet, which utilizes multiple representation layers. It is inspired by the simplicity and success of hybrid algorithm or fusion of multiple detection networks. The computational complexity is an essential factor in training deep neural networks; therefore, to reduce the computational complexity while learning multiple feature representation layers, a fixed base model has been used. The performance of the proposed network is demonstrated on multiple databases such as IIITD-WVU MUIPA and IIITD-CLI databases under cross-database training-testing settings, to assess the generalizability of the proposed algorithm.
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001
ICPR3
2020 MixNet for Generalized Face Presentation Attack Detection
abstract
The non-intrusive nature and high accuracy of face recognition algorithms have led to their successful deployment across multiple applications ranging from border access to mobile unlocking and digital payments. However, their vulnerability against sophisticated and cost-effective presentation attack mediums raises essential questions regarding its reliability. In the literature, several presentation attack detection algorithms are presented; however, they are still far behind from reality. The major problem with existing work is the generalizability against multiple attacks both in the seen and unseen setting. The algorithms which are useful for one kind of attack (such as print) perform unsatisfactorily for another type of attack (such as silicone masks). In this research, we have proposed a deep learning-based network termed as MixNet to detect presentation attacks in cross-database and unseen attack settings. The proposed algorithm utilizes state-of-the-art convolutional neural network architectures and learns the feature mapping for each attack category. Experiments are performed using multiple challenging face presentation attack databases such as SMAD and Spoof In the Wild (SiW-M) databases. Extensive experiments and comparison with existing state of the art algorithms show the effectiveness of the proposed algorithm.
Nilay Sanghvi, Sushant Kumar Singh, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001
ICPR3
2019 Detecting and Mitigating Adversarial Perturbations for Robust Face Recognition
Gaurav Goswami, Akshay Agarwal 0001, Nalini K. Ratha, Richa Singh 0001, Mayank Vatsa
Int. J. Comput. Vis.2
2018 Unravelling Robustness of Deep Learning Based Face Recognition Against Adversarial Attacks
abstract
Deep neural network (DNN) architecture based models have high expressive power and learning capacity. However, they are essentially a black box method since it is not easy to mathematically formulate the functions that are learned within its many layers of representation. Realizing this, many researchers have started to design methods to exploit the drawbacks of deep learning based algorithms questioning their robustness and exposing their singularities. In this paper, we attempt to unravel three aspects related to the robustness of DNNs for face recognition: (i) assessing the impact of deep architectures for face recognition in terms of vulnerabilities to attacks inspired by commonly observed distortions in the real world that are well handled by shallow learning methods along with learning based adversaries; (ii) detecting the singularities by characterizing abnormal filter response behavior in the hidden layers of deep networks; and (iii) making corrections to the processing pipeline to alleviate the problem. Our experimental evaluation using multiple open-source DNN-based face recognition networks, including OpenFace and VGG-Face, and two publicly available databases (MEDS and PaSC) demonstrates that the performance of deep learning based face recognition algorithms can suffer greatly in the presence of such distortions. The proposed method is also compared with existing detection algorithms and the results show that it is able to detect the attacks with very high accuracy by suitably designing a classifier using the response of the hidden layers in the network. Finally, we present several effective countermeasures to mitigate the impact of adversarial attacks and improve the overall robustness of DNN-based face recognition.
Gaurav Goswami, Nalini K. Ratha, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
AAAI3
2017 SWAPPED! Digital face presentation attack detection via weighted local magnitude pattern
abstract
Advancements in smartphone applications have empowered even non-technical users to perform sophisticated operations such as morphing in faces as few tap operations. While such enablements have positive effects, as a negative side, now anyone can digitally attack face (biometric) recognition systems. For example, face swapping application of Snapchat can easily create “swapped” identities and circumvent face recognition system. This research presents a novel database, termed as SWAPPED - Digital Attack Video Face Database, prepared using Snap chat's application which swaps/stitches two faces and creates videos. The database contains bonafide face videos and face swapped videos of multiple subjects. Baseline face recognition experiments using commercial system shows over 90% rank-1 accuracy when attack videos are used as probe. As a second contribution, this research also presents a novel Weighted Local Magnitude Pattern feature descriptor based presentation attack detection algorithm which outperforms several existing approaches.
Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Afzel Noore
IJCB1
2016 Mobile periocular matching with pre-post cataract surgery
abstract
Ocular recognition algorithms, including iris matching, have been used in several applications including large scale national ID projects such as India's Aadaar. Deployment of large-scale biometric systems is expected to rely on using multiple devices including mobile devices to ensure widespread adoption of biometric recognition systems. Ocular images captured using mobile devices may have challenges such as uncontrolled illumination, complex background, and geometric distortions. Further, among many enrollees of large scale biometrics program, some may have ocular diseases. One of the most common ocular disease in elderly is cataract. While it is established that iris recognition may be challenging due to ocular diseases, this paper investigates periocular recognition with pre and post cataract surgery images. In this research, we present a mobile periocular database of 145 subjects1. Baseline results also include a framework that achieves over 69% rank-10 accuracy and around 24% genuine accept rate at 1% false accept rate in inter-session experiments.
Rohit Keshari, Soumyadeep Ghosh, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
ICIP3
2016 Fingerprint sensor classification via Mélange of handcrafted features
abstract
Large scale biometrics projects rely on capturing images/signal from multiple sensors. For example, in India's Aadhaar project, multiple fingerprint sensors of different make and model are used for data collection. Similarly, in law enforcement applications, different agencies use different fingerprint sensors. These scenarios cause two potential problems: (i) sensor inter-operability and (ii) protecting/recording chain of evidence. While sensor inter-operability in fingerprints is a well studied problem, automatically recording chain of evidence is a relatively less explored research problem. For both the problems, one potential approach includes automatically identifying sensors based on the input image. This paper presents a novel fingerprint sensor identification algorithm based on multiple features such as Haralick, entropy, statistical and image quality features. The proposed algorithm is evaluated on a large database with 30,000 images with 15 fingerprint sensor classes. The proposed algorithm achieves an accuracy of 96% and computationally requires less than 10 milliseconds for an image.
Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
ICPR1
2016 Face anti-spoofing with multifeature videolet aggregation
abstract
Biometric systems can be attacked in several ways and the most common being spoofing the input sensor. Therefore, anti-spoofing is one of the most essential prerequisite against attacks on biometric systems. For face recognition it is even more vulnerable as the image capture is non-contact based. Several anti-spoofing methods have been proposed in the literature for both contact and non-contact based biometric modalities often using video to study the temporal characteristics of a real vs. spoofed biometric signal. This paper presents a novel multi-feature evidence aggregation method for face spoofing detection. The proposed method fuses evidence from features encoding of both texture and motion (liveness) properties in the face and also the surrounding scene regions. The feature extraction algorithms are based on a configuration of local binary pattern and motion estimation using histogram of oriented optical flow. Furthermore, the multi-feature windowed videolet aggregation of these orthogonal features coupled with support vector machine-based classification provides robustness to different attacks. We demonstrate the efficacy of the proposed approach by evaluating on three standard public databases: CASIA-FASD, 3DMAD and MSU-MFSD with equal error rate of 3.14%, 0%, and 0%, respectively.
Talha Ahmad Siddiqui, Samarth Bharadwaj, Tejas I. Dhamecha, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha
ICPR4