EDBT 2026 Demo / reviewers in the wild / expert
Richa Singh 0001
dblp:75/3512
· DBLP profile ↗
195ranked-venue papers
10as first author
84since 2021 · last 2026
0000-0003-4060-4573ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 137 · 6 first-author · 69 since 2021Graphics, computer vision, multimedia, augmented reality and games · 133 · 4 first-author · 63 since 2021Security and privacy · 43 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 36 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACID Test: A Benchmark for Cultural Safety and Alignment in LALMsabstractLarge Audio Language Models (LALMs) are transforming AI by processing and generating human language directly from audio. As these models proliferate in real-world applications, it becomes critical to evaluate their performance to ensure equitable and safe use across diverse linguistic and cultural contexts. We present the first comprehensive study of cultural bias in LALMs, extending text-based harm frameworks to the audio modality to analyze how linguistic diversity influences model behavior and uncover challenges in interpreting audio nuances. To address this, we introduce the Audio Cultural Intelligence Dataset (ACID), a multilingual audio–text benchmark spanning 1,315 hours across diverse languages and cultural contexts, and we conduct a systematic evaluation of 10 open-source and two closed-source models. Our results reveal substantial performance disparities across languages and cultural settings and show that biases manifest distinctly when models process audio inputs. These findings highlight the need to evaluate LALMs not only for technical accuracy but also for fair and culturally sensitive behavior, motivating the development of inclusive datasets and culturally aware training practices for safer and more equitable audio language models. Bikash Dutta, Adit Jain, Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
AAAI | 5 |
| 2026 | NutriScreener: Retrieval Augmented Multi-Pose Graph Attention Network for Malnourishment ScreeningabstractChild malnutrition remains a global crisis, yet existing screening methods are laborious and poorly scalable, hindering early intervention. In this work, we present NutriScreener, a retrieval-augmented, multi-pose graph attention network that combines CLIP-based visual embeddings, class-boosted knowledge retrieval, and context awareness to enable robust malnutrition detection and anthropometric prediction from children's images, simultaneously addressing generalizability and class-imbalance. In a clinical study, doctors rated it 4.3/5 for accuracy and 4.6/5 for efficiency, confirming its deployment readiness in low-resource settings. NutriScreener was trained and tested on 2,141 children from AnthroVision and additionally evaluated on diverse cross-continent populations, including ARAN and an in-house collected CampusPose dataset, achieving 0.79 recall, 0.82 AUC, and significantly lower anthropometric RMSEs, demonstrating reliable measurement in unconstrained, pediatric settings. Cross-dataset results show up to 25\% recall gain and up to 2.3 cm reduction in head circumference RMSE using demographically matched knowledge bases. NutriScreener offers a scalable and accurate solution for early malnutrition detection in low-resource environments. Misaal Khan, Mayank Vatsa, Richa Singh 0001 |
AAAI | 4 |
| 2026 | Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image GenerationabstractThe architectural blueprint of today’s leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown across three core primitives—negation, counting, and spatial relations. Our analysis reveals a dramatic performance collapse: models that are accurate on single primitives fail precipitously when these are combined, exposing severe interference. We trace this failure to three key factors. First, training data show a near-total absence of explicit negations. Second, continuous attention architectures are fundamentally unsuitable for discrete logic. Third, evaluation metrics reward visual plausibility over constraint satisfaction. By analyzing recent benchmarks and methods, we show that current solutions and simple scaling cannot bridge this gap. Achieving genuine compositionality, we conclude, will require fundamental advances in representation and reasoning rather than incremental adjustments to existing architectures. Mayank Vatsa, Aparna Bharati, Richa Singh 0001 |
AAAI | 3 |
| 2026 | HumanBench: Two Heads, No Legs, But Mostly Human, the State of Generative Capabilities in T2I ModelsabstractDespite rapid advances, text-to-image (T2I) models still struggle to generate anatomically coherent and semantically grounded humans. We present HumanBench, a large-scale benchmark of 36K images designed to evaluate T2I models across four axes: template consistency, spatial reasoning, action understanding, and texture recognition. To quantify alignment, we propose two metrics—Agreement and Distinction—that capture both fidelity to prompts and separation from counterfactual or negated variants. HumanBench introduces a formal measure of prompt complexity based on slot counts and compositional bindings, enabling controlled analysis of model performance. Evaluating three publicly accessible models (Midjourney, Stable Diffusion, FLUX) comprehensively and three additional models on a subset, we find persistent failures, including disfigurements, species leakage, and counting errors. Prompt–image consistency declines monotonically with complexity, a trend observed across the models. A human study of 300 samples further corroborates these findings, showing a strong correlation between Agreement scores and human ratings. By unifying automated metrics, controlled prompt design, and human validation, HumanBench provides a robust foundation for auditing the compositional and anatomical capabilities of modern T2I models. Dataset is available here. Anubhooti Jain, Mayank Vatsa, Richa Singh 0001 |
WACV | 3 |
| 2026 | Introduction to the Special Issue on Transformers
Feng Xia 0001, Tyler Derr, Anh Tuan Luu, Richa Singh 0001, Aline Villavicencio |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | AQUAFace: Age-Invariant Quality Adaptive Face Recognition for Unconstrained Selfie vs ID VerificationabstractFace recognition in the presence of age and quality variations poses a formidable challenge. While recent margin-based loss functions have shown promise in addressing these variations individually, real-world scenarios such as selfie versus ID face matching often involve simultaneous variations of both age and quality. In response, we propose a comprehensive framework aimed at mitigating the impact of these variations while preserving vital identity-related information crucial for accurate face recognition. The proposed adaptive margin-based loss function AQUAFace exhibits adaptiveness towards hard samples characterized by significant age and quality variations. This loss function is meticulously designed to prioritize the preservation of identity-related features while simultaneously mitigating the adverse effects of age and quality variations on face recognition accuracy. To validate the effectiveness of our approach, we focus on the specific task of selfie versus ID document matching. Our results demonstrate that AQUAFace effectively handles age and quality differences, leading to enhanced recognition performance. Additionally, we explore the benefits of fine-tuning the recognition model with synthetic data, further boosting performance. As a result, our proposed model, AQUAFace, achieves state-of-the-art performance on six benchmark datasets (CALFW, CPLFW, CFP-FP, AgeDB, IJB-C, and TinyFace), each exhibiting diverse age and quality variations. Shivang Agarwal, Jyoti Chaudhary, Sadiq Siraj Ebrahim, Mayank Vatsa, Richa Singh 0001, Shyam Prasad Adhikari, Sangeeth Reddy Battu |
AAAI | 5 |
| 2025 | Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation ModelsabstractExisting unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts—a challenge known as adjacency. To address this, we propose FADE (Fine-Grained Attenuation for Diffusion Erasure), introducing adjacency-aware unlearning in diffusion models. FADE comprises two components: (1) the Concept Neighborhood, which identifies an adjacency set of related concepts, and (2) Mesh Modules, employing a structured combination of Expungement, Adjacency, and Guidance loss components. These enable precise erasure of target concepts while preserving fidelity across related and unrelated concepts. Evaluated on datasets like Stanford Dogs, Oxford Flowers, CUB, I2P, Imagenette, and ImageNet-1k, FADE effectively removes target concepts with minimal impact on correlated concepts, achieving at least a 12% improvement in retention performance over state-of-the-art methods. Our code and models are available on the project page: iab-rubric/unlearning/FG-Un. Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh 0001 |
CVPR | 5 |
| 2025 | Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?abstractThe proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi – across tasks like Automatic Speech Recognition (ASR), Audio Question Answering (AQA), and audio classification tasks in Hindi and code-mixed Hindi-English (aka Hinglish). We also explore the potential of Retrieval-Augmented Generation (RAG) to boost LALM performance in these low-resource settings. Our findings highlight significant performance discrepancies, with LALMs performing well in audio classification but struggling with ASR and AQA. While RAG shows potential, especially for audio classification, its impact is inconsistent across tasks. This work offers critical insights into the challenges of using LALMs for low-resource languages and provides a foundation for developing more inclusive and adaptable AI systems for complex multilingual tasks. Bikash Dutta, Rishabh Ranjan, Akshat Jain, Richa Singh 0001, Mayank Vatsa |
ICASSP | 4 |
| 2025 | Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice BiometricsabstractThe widespread adoption of voice-enabled authentication and audio biometric systems have significantly increased privacy vulnerabilities associated with sensitive speech data. Compliance with privacy regulations such as GDPR’s right to be forgotten and India’s DPDP Act necessitates targeted and efficient erasure of individual-specific voice signatures from already-trained biometric models. Existing unlearning methods designed for visual data inadequately handle the sequential, temporal, and high-dimensional nature of audio signals, leading to ineffective or incomplete speaker and accent erasure. To address this, we introduce QPAudioEraser, a quantum-inspired audio unlearning framework. Our four-phase approach involves: (1) weight initialization using destructive interference to nullify target features, (2) superposition-based label transformations that obscure class identity, (3) an uncertainty-maximizing quantum loss function, and (4) entanglement-inspired mixing of correlated weights to retain model knowledge. Comprehensive evaluations with ResNet18, ViT, and CNN architectures across AudioM-NIST, Speech Commands, LibriSpeech, and Speech Accent Archive datasets validate QPAudioEraser’s superior performance. The framework achieves complete erasure of target data (0% Forget Accuracy) while incurring minimal impact on model utility, with a performance degradation on retained data as low as 0.05%. QPAudioEraser consistently surpasses conventional baselines across single-class, multi-class, sequential, and accent-level erasure scenarios, establishing the proposed approach as a robust privacy-preserving solution. Shreyansh Pathak, Sonu Sreshtha, Richa Singh 0001, Mayank Vatsa |
IJCB | 3 |
| 2025 | Erasing Shadows: Residual-Guided Watermark Removal Via Reverse DiffusionabstractInvisible watermarking is essential for protecting the authenticity and ownership of generative content. However, sophisticated removal techniques pose significant threats. This paper introduces a novel attack framework that exploits diffusion model dynamics through uniform noise injection, residual-based watermark extraction, and reverse diffusion reconstruction. The key insight is that watermarks manifest as structured artifacts within the residuals of consecutive denoising steps. By analyzing these residuals over multiple diffusion steps, the proposed method isolates and removes watermark traces. Extensive experiments demonstrate effective watermark removal from images embedded using classical and state-of-the-art algorithms, including DWT-DCT, RivaGAN, and Stable Signature. Evaluation on diverse datasets, CelebA, FFHQ, and synthetic images generated using SDXL and DALL•E 3, shows watermark removal rates of up to 99%, while maintaining high perceptual fidelity, validated by SSIM and FID metrics. These findings highlight critical vulnerabilities in existing watermarking strategies and establish a benchmark for assessing resilience in protecting generative AI content. Nidhi Soni, Ravi Kumar Saxena, Mayank Vatsa, Richa Singh 0001 |
IJCB | 4 |
| 2025 | ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetabstractThe proliferation of deepfakes and AI-generated content has led to a surge in media forgeries and misinformation, necessitating robust detection systems. However, current datasets lack diversity across modalities, languages, and real-world scenarios. To address this gap, we present ILLUSION (Integration of Life-Like Unique Synthetic Identities and Objects from Neural Networks), a large-scale, multi-modal
deepfake dataset comprising 1.3 million samples spanning audio-visual forgeries, 26 languages, challenging noisy environments, and various manipulation protocols. Generated using 28 state-of-the-art generative techniques, ILLUSION includes
faceswaps, audio spoofing, synchronized audio-video manipulations, and synthetic media while ensuring a balanced representation of gender and skin tone for unbiased evaluation. Using Jaccard Index and UpSet plot analysis, we demonstrate ILLUSION’s distinctiveness and minimal overlap with existing datasets, emphasizing its novel generative coverage. We benchmarked image, audio, video, and multi-modal detection models, revealing key challenges such as performance degradation in multilingual and multi-modal contexts, vulnerability to real-world distortions, and limited generalization to zero-day attacks. By bridging synthetic and real-world complexities, ILLUSION provides a challenging yet essential platform for advancing deepfake detection research. The dataset is publicly available at https://www.iab-rubric.org/illusion-database. Kartik Thakral, Rishabh Ranjan, Akshat Jain, Mayank Vatsa, Richa Singh 0001 |
ICLR | 6 |
| 2025 | Harmonizing Geometry and Uncertainty: Diffusion with HyperspheresabstractDo contemporary diffusion models preserve the class geometry of hyperspherical data? Standard diffusion models rely on isotropic Gaussian noise in the forward process, inherently favoring Euclidean spaces. However, many real-world problems involve non-Euclidean distributions, such as hyperspherical manifolds, where class-specific patterns are governed by angular geometry within hypercones. When modeled in Euclidean space, these angular subtleties are lost, leading to suboptimal generative performance. To address this limitation, we introduce \textbf{HyperSphereDiff} to align hyperspherical structures with directional noise, preserving class geometry and effectively capturing angular uncertainty. We demonstrate both theoretically and empirically that this approach aligns the generative process with the intrinsic geometry of hyperspherical data, resulting in more accurate and geometry-aware generative models. We evaluate our framework on four object datasets and two face datasets, showing that incorporating angular uncertainty better preserves the underlying hyperspherical manifold. Muskan Dosi, Chiranjeev Chiranjeev, Kartik Thakral, Mayank Vatsa, Richa Singh 0001 |
ICML | 5 |
| 2025 | SHIELD: A Self-supervised, Silicosis-focused Hierarchical Imaging Framework for Occupational Lung Disease DiagnosisabstractSilicosis is an irreversible lung disease caused by silica dust exposure in industrial settings. Early detection is crucial, but automatic diagnostic methods are hindered by limited data availability. We propose SHIELD - a self-supervised, Silicosis-focused Hierarchical Imaging framework for early occupational Lung disease Diagnosis. Our method leverages a multi-resolution jigsaw puzzle pretext task on CXR images to extract and preserve features for lung region analysis. By employing a pyramidal strategy to generate pretrained models at various resolutions, followed by fine-tuning and a two-level ensembling across diverse deep learning architectures, SHIELD achieves enhanced diagnostic accuracy. We validate our approach on a publicly collected CXR dataset of 3044 samples from public health centers in India. SHIELD achieves 72% accuracy, demonstrating up to 20% improvement over baseline approaches. This work advances medical image analysis and supports UN Sustainable Development Goal 3 by providing cost-effective early screening in resource-limited settings. Yasmeena Akhter, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa |
IJCAI | 3 |
| 2025 | Words Over Pixels? Rethinking Vision in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) promise seamless integration of vision and language understanding. However, despite their strong performance, recent studies reveal that MLLMs often fail to effectively utilize visual information, frequently relying on textual cues instead. This survey provides a comprehensive analysis of the vision component in MLLMs, covering both application-level and architectural aspects. We investigate critical challenges such as weak spatial reasoning, poor fine-grained visual perception, and suboptimal fusion of visual and textual modalities. Additionally, we explore limitations in current vision encoders, benchmark inconsistencies, and their implications for downstream tasks. By synthesizing recent advancements, we highlight key research opportunities to enhance visual understanding, improve cross-modal alignment, and develop more robust and efficient MLLMs. Our observations emphasize the urgent need to elevate vision to an equal footing with language, paving the path for more reliable and perceptually aware multimodal models. Anubhooti Jain, Mayank Vatsa, Richa Singh 0001 |
IJCAI | 3 |
| 2025 | Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
Bikash Dutta, Rishabh Ranjan, Shyam Sathvik, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 5 |
| 2025 | LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security
Nidheesh Gorthi, Kartik Thakral, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa |
INTERSPEECH | 4 |
| 2025 | Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
Rishabh Ranjan, Ayinala Likhith, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 4 |
| 2025 | SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
Rishabh Ranjan, Kishan Pipariya, Mayank Vatsa, Richa Singh 0001 |
INTERSPEECH | 4 |
| 2025 | Non-invasive TB Detection Using Acoustic and Semantic Features from Cough Sounds
Yasmeena Akhter, Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
MICCAI (1) | 5 |
| 2025 | Learning the Power of "No": Foundation Models with NegationsabstractNegation is a fundamental aspect of natural language reasoning, yet foundational vision-language models (VLMs) like CLIP face significant challenges in accurately interpreting it. These models often process text prompts holistically, making it difficult to isolate and understand the role of negated terms. To overcome this limitation, we present CC-Neg: a novel dataset consisting of 228,246 images, each paired with both true captions and their corresponding negated versions. CC-Neg provides a critical benchmark to assess and improve foundational VLMs' ability to process negations, focusing specifically on how the presence of terms like ‘not’ alters the semantic relationship between images and their textual descriptions. To illustrate the effectiveness of the CC-Neg dataset in enhancing negation understanding, we introduce the CoN-CLIP framework, which incorporates targeted modifications to CLIP's contrastive loss function. When trained with CC-Neg, CoN-CLIP achieves a 3.85% average improvement in top-1 accuracy for zero-shot image classification across eight datasets, and a 4.4% performance boost on challenging compositionality benchmarks such as SugarCREPE. These results highlight CoN-CLIP's enhanced understanding of the nuanced semantic relationships involving negation. Our code and the CC-Neg benchmark are available at: https://github.com/jaisidhsingh/CoN-CLIP. Jaisidh Singh, Ishaan Shrivastava, Mayank Vatsa, Richa Singh 0001, Aparna Bharati |
WACV | 4 |
| 2025 | Beyond shadows and light: Odyssey of face recognition for social good
Chiranjeev Chiranjeev, Muskan Dosi, Shivang Agarwal, Jyoti Chaudhary, Pranav Pant, Mayank Vatsa, Richa Singh 0001 |
Comput. Vis. Image Underst. | 7 |
| 2025 | On learning discriminative embeddings for optimized top-k matching
Soumyadeep Ghosh, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
Pattern Recognit. | 3 |
| 2024 | BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock AttributesabstractAutomatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually rather than in species-rich flocks. We present Bird-Collect, a comprehensive benchmark dataset for monitoring dense bird flock attributes. It includes a unique collection of more than 6,000 high-resolution images of Demoiselle Cranes (Anthropoides virgo) feeding and nesting in the vicinity of Khichan region of Rajasthan. Particularly, each image contains an average of 190 individual birds, illustrating the complex dynamics of densely populated bird flocks on a scale that has not previously been studied. In addition, a total of 433 distinct pictures captured at Keoladeo National Park, Bharatpur provide a comprehensive representation of 34 distinct bird species belonging to various taxonomic groups. These images offer details into the diversity and the behaviour of birds in vital natural ecosystem along the migratory flyways. Additionally, we provide a set of 2,500 point-annotated samples which serve as ground truth for benchmarking various computer vision tasks like crowd counting, density estimation, segmentation, and species classification. The benchmark performance for these tasks highlight the need for tailored approaches for specific wildlife applications, which include varied conditions including views, illumination, and resolutions. With around 46.2 GBs in size encompassing data collected from two distinct nesting ground sets, it is the largest birds dataset containing detailed annotations, showcasing a substantial leap in bird research possibilities. We intend to publicly release the dataset to the research community. The database is available at: https://iab-rubric.org/resources/wildlife-dataset/birdcollect Kshitiz, Sonu Sreshtha, Bikash Dutta, Muskan Dosi, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
AAAI | 6 |
| 2024 | Adventures of Trustworthy Vision-Language Models: A SurveyabstractRecently, transformers have become incredibly popular in computer vision and vision-language tasks. This notable rise in their usage can be primarily attributed to the capabilities offered by attention mechanisms and the outstanding ability of transformers to adapt and apply themselves to a variety of tasks and domains. Their versatility and state-of-the-art performance have established them as indispensable tools for a wide array of applications. However, in the constantly changing landscape of machine learning, the assurance of the trustworthiness of transformers holds utmost importance. This paper conducts a thorough examination of vision-language transformers, employing three fundamental principles of responsible AI: Bias, Robustness, and Interpretability. The primary objective of this paper is to delve into the intricacies and complexities associated with the practical use of transformers, with the overarching goal of advancing our comprehension of how to enhance their reliability and accountability. Mayank Vatsa, Anubhooti Jain, Richa Singh 0001 |
AAAI | 3 |
| 2024 | ToonerGAN: Reinforcing GANs for Obfuscating Automated Facial IndexingabstractThe rapid evolution of automatic facial indexing technologies increases the risk of compromising personal and sensitive information. To mitigate the issue, we propose creating cartoon avatars, or ‘toon avatars', designed to effectively obscure identity features. The primary objective is to deceive current AI systems, preventing them from accurately identifying individuals while making minimal modifications to their facial features. Moreover, we aim to ensure that a human observer can still recognize the person depicted in these altered avatar images. To achieve this, we introduce ‘ToonerGAN’, a novel approach that utilizes Generative Adversarial Networks (GANs) to craft personalized cartoon avatars. The ToonerGAN framework consists of a style and a de-identification module that work together to produce high-resolution, realistic cartoon images. For the efficient training of our network, we have developed ‘ToonSet’ dataset, consisting of around 23,000 facial images and their cartoon renditions. Through comprehensive experiments and benchmarking against existing datasets, including CelebA-HQ, our method demonstrates superior performance in obfuscating identity while preserving the utility of data. Additionally, a user-centric study exploring the effectiveness of ToonerGAN has yielded compelling observations. Kartik Thakral, Shashikant Prasad, Stuti Aswani, Mayank Vatsa, Richa Singh 0001 |
CVPR | 5 |
| 2024 | HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
Chiranjeev Chiranjeev, Muskan Dosi, Kartik Thakral, Mayank Vatsa, Richa Singh 0001 |
ECCV (88) | 5 |
| 2024 | Navigating Text-to-Image Generative Bias Across Indic Languages
Surbhi Mittal, Arnav Sudan, Mayank Vatsa, Richa Singh 0001, Tamar Glaser, Tal Hassner |
ECCV (88) | 4 |
| 2024 | Is Face Super Resolution Truly Pushing the Boundaries of Face Recognition?abstractWith the improving efficacy of generative algorithms, the performance of face super-resolution algorithms is also increasing towards generating high-quality facial data. However, are these images useful for face recognition? This paper investigates whether these enhanced super-resolved facial images only improve the visual quality or they also aid in improving the recognizability of these images, thus contributing towards addressing the challenge of low-resolution face recognition. We conduct a comprehensive empirical and statistical analysis of human perception and face recognition tasks. Extensive experiments are performed using multiple state-of-the-art generative and face recognition models across six publicly available face datasets to assess whether face super-resolution algorithms are effective in recognizing individuals in low-resolution conditions. The results and supporting analysis indicate that the ability of super-resolution images to improve recognizability is limited, and further research is required to design generative AI algorithms that improve both visual appearance and recognizability of low-resolution images. Muskan Dosi, Udaybhan Rathore, Chiranjeev Chiranjeev, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa |
IJCB | 5 |
| 2024 | Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional NoisesabstractDeep learning models, such as those used for face recognition and attribute prediction, are susceptible to manipulations like adversarial noise and unintentional noise, including Gaussian and impulse noise. This paper introduces CIAI, a Class-Independent Adversarial Intent detection network built on a modified vision transformer with detection layers. CIAI employs a novel loss function that combines Maximum Mean Discrepancy and Center Loss to detect both intentional (adversarial attacks) and unintentional noise, regardless of the image class. It is trained in a multistep fashion. We also introduce the aspect of intent during detection that can act as an added layer of security. We further showcase the performance of our proposed detector on CelebA, CelebA-HQ, LFW, AgeDB, and CIFAR-10 datasets. Our detector is able to detect both intentional (like FGSM, PGD, and DeepFool) and unintentional (like Gaussian and Salt & Pepper noises) perturbations. Anubhooti Jain, Susim Mukul Roy, Kwanit Gupta, Mayank Vatsa, Richa Singh 0001 |
IJCB | 5 |
| 2024 | Faking Fluent: Unveiling the Achilles' Heel of Multilingual Deepfake DetectionabstractWith the rapid advancement of deep learning techniques, the generation of audio deepfakes has achieved remarkable realism across various languages and accents. However, the effectiveness of audio deepfake detection models in diverse linguistic environments remains a crucial area of investigation. This paper presents the first empirical study on the robustness of current audio deepfake detection algorithms across different languages and accents. We evaluate whether these models maintain their effectiveness across varied linguistic domains or perform better in specific language contexts. Our comprehensive analysis examines state-of-the-art audio deepfake detection models trained on the ASVspoof 2019 and BhashaBluff datasets, assessing their performance across four diverse datasets: three representing similar-language variations (Speech Accent Archive, Svarah, and the UK English Accent Dataset) and one representing a different language (Vaani). Our results and supporting analysis indicate that while current models perform well on benchmark datasets, their ability to generalize across diverse linguistic conditions is limited. We identify potential vulnerabilities in existing models when faced with unfamiliar languages or accents, highlighting the need for more inclusive and adaptable detection systems. Our results highlight the need to enhance the robustness of audio deepfake detection across the global linguistic spectrum and emphasize the importance of developing models capable of effectively identifying synthetic speech, regardless of language or accent. Rishabh Ranjan, Bikash Dutta, Mayank Vatsa, Richa Singh 0001 |
IJCB | 4 |
| 2024 | Context Encoded Multi-Modal Attention Network for Detecting Audio SpoofingabstractHumans interpret speech through both acoustic and textual signals. Recognizing the importance of incorporating contextual information into speech processing, especially for spoofing detection, we introduce a multimodal representation framework that integrates text with audio signals to enhance spoof detection effectiveness. This research advances in two significant ways: firstly, through the creation of a new multimodal spoof detection model called Raw-BERT, and secondly, by developing an innovative multi-headed multimodal attention network that synergistically merges text and audio data for enhanced detection capabilities. We have extensively evaluated the proposed framework across diverse datasets in various languages, demonstrating that the addition of textual context significantly boosts model performance. Overall, the proposed model demonstrates superior results, consistently outperforming existing spoof detection algorithms in multiple evaluations on benchmark audio spoofing datasets. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 3 |
| 2024 | Restoring Noisy Images Using Dual-Tail Encoder-Decoder Signal Separation Network
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
ICPR (1) | 3 |
| 2024 | Supervised Mixup: Protecting the Likely Classes for Adversarial Robustness
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
ICPR (5) | 3 |
| 2024 | PSIVUS: Atherosclerotic Plaque Segmentation in Intravascular Ultrasound Images via Active Learning
Anuradha Mahato, Paromita Banerjee, Rutvik Narendrabhai Jethava, Bhanu Duggal, Angshuman Paul, Mayank Vatsa, Richa Singh 0001 |
ICPR (28) | 7 |
| 2024 | DomainAdapt: Leveraging Multitask Learning and Domain Insights for Children's Nutritional Status Assessment
Misaal Khan, Richa Singh 0001, Mayank Vatsa |
MICCAI (3) | 2 |
| 2024 | SynthProv: Interpretable Framework for Profiling Identity LeakageabstractGenerative Adversarial Networks (GANs) can generate hyperrealistic face images of synthetic identities based on a latent understanding of real images from a large training set. Despite their proficiency, the term "synthetic identity" remains ambiguous, and the uniqueness of the faces GANs produce is rarely assessed. Recent studies have found that identities from the training data can unintentionally appear in the faces generated by StyleGAN2, but the cause of this phenomenon is unclear. In this work, we propose a novel framework, SynthProv, that utilizes the improved interpolation ability of StyleGAN2 latent space and employs image composition to analyze leakage. This is the first method that goes beyond detection and traces the source or provenance of constituent identity signals in the generated image. Experiments show that SynthProv succeeds in both detection and provenance tasks using multiple matching strategies. We identify identities from FFHQ and CelebA-HQ training datasets with the highest leakage into the latent space as "leaking reals". Analyzing latent space behavior to evaluate generative model privacy via leakage is an important research direction, as undetected leaking reals pose a significant threat to training data privacy. Our code is available at https://github.com/jaisidhsingh/SynthProv. Jaisidh Singh, Harshil Bhatia, Mayank Vatsa, Richa Singh 0001, Aparna Bharati |
WACV | 4 |
| 2024 | Corruption depth: Analysis of DNN depth for misclassification
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
Neural Networks | 3 |
| 2024 | Multi-Surface Multi-Technique (MUST) Latent Fingerprint DatabaseabstractLatent fingerprint recognition involves acquisition and comparison of latent fingerprints with an exemplar gallery of fingerprints. The diversity in the type of surface leads to different procedures to recover the latent fingerprint. The appearance of latent fingerprints vary significantly due to the development techniques, leading to large intra-class variation. Due to lack of large datasets acquired using multiple mechanisms and surfaces, existing algorithms for latent fingerprint enhancement and comparison may perform poorly. In this study, we propose a Multi-Surface Multi-Technique (MUST) Latent Fingerprint Database. The database consists of more than 16,000 latent fingerprint impressions from 120 unique classes (120 fingers from 12 participants). Including corresponding exemplar fingerprints (livescan and rolled) and extended gallery, the dataset has nearly 21,000 impressions. It has latent fingerprints acquired under 35 different scenarios and additional four subsets of exemplar prints captured using live scan sensor and inked-rolled prints. With 39 different subsets, the database illustrates intra-class variations in latent fingerprints. The database has a potential usage towards building robust algorithms for latent fingerprint enhancement, segmentation, comparison, and multi-task learning. We also provide annotations for manually marked minutiae, acquisition Pixel Per Inch (PPI), and semantic segmentation masks. We also present the experimental protocol and the baseline results for the proposed dataset. The availability of the proposed database can encourage research in handling intra-class variation in latent fingerprint recognition. Aakarsh Malhotra, Mayank Vatsa, Richa Singh 0001, Keith B. Morris, Afzel Noore |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | IdProv: Identity-Based Provenance for Synthetic Image Generation (Student Abstract)abstractRecent advancements in Generative Adversarial Networks (GANs) have made it possible to obtain high-quality face images of synthetic identities. These networks see large amounts of real faces in order to learn to generate realistic looking synthetic images. However, the concept of a synthetic identity for these images is not very well-defined. In this work, we verify identity leakage from the training set containing real images into the latent space and propose a novel method, IdProv, that uses image composition to trace the source of identity signals in the generated image. Harshil Bhatia, Jaisidh Singh, Gaurav Sangwan, Aparna Bharati, Richa Singh 0001, Mayank Vatsa |
AAAI | 5 |
| 2023 | DF-Platter: Multi-Face Heterogeneous Deepfake DatasetabstractDeepfake detection is gaining significant importance in the research community. While most of the research efforts are focused towards high-quality images and videos with controlled appearance of individuals, deepfake generation algorithms now have the capability to generate deep-fakes with low-resolution, occlusion, and manipulation of multiple subjects. In this research, we emulate the real-world scenario of deepfake generation and propose the DF-Platter dataset, which contains (i) both low-resolution and high-resolution deepfakes generated using multiple generation techniques and (ii) single-subject and multiple-subject deepfakes, with face images of Indian ethnicity. Faces in the dataset are annotated for various attributes such as gender, age, skin tone, and occlusion. The dataset is prepared in 116 days with continuous usage of 32 GPUs accounting to 1,800 GB cumulative memory. With over 500 GBs in size, the dataset contains a total of 133,260 videos encompassing three sets. To the best of our knowledge, this is one of the largest datasets containing vast variability and multiple challenges. We also provide benchmark results under multiple evaluation settings using popular and state-of-the-art deepfake detection models, for c0 images and videos along with c23 and c40 compression variants. The results demonstrate a significant performance reduction in the deepfake detection task on low-resolution deep-fakes. Furthermore, existing techniques yield declined detection accuracy on multiple-subject deepfakes. It is our assertion that this database will improve the state-of-the-art by extending the capabilities of deepfake detection algorithms to real-world scenarios. The database is available at: http://iab-rubric.org/df-platter-database. Kartik Narayan, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
CVPR | 6 |
| 2023 | Are Face Detection Models Biased?abstractThe presence of bias in deep models leads to unfair outcomes for certain demographic subgroups. Research in bias focuses primarily on facial recognition and attribute prediction with scarce emphasis on face detection. Existing studies consider face detection as binary classification into ‘face’ and ‘non-face’ classes. In this work, we investigate possible bias in the domain of face detection through facial region localization which is currently unexplored. Since facial region localization is an essential task for all face recognition pipelines, it is imperative to analyze the presence of such bias in popular deep models. Most existing face detection datasets lack suitable annotation for such analysis. Therefore, we web-curate the Fair Face Localization with Attributes (F2LA) dataset and manually annotate more than 10 attributes per face, including facial localization information. Utilizing the extensive annotations from F2LA, an experimental setup is designed to study the performance of four pre-trained face detectors. We observe (i) a high disparity in detection accuracies across gender and skin-tone, and (ii) interplay of confounding factors beyond demography. The F2LA data and associated annotations can be accessed at http://iab-rubric.org/index.php/F2LA. Surbhi Mittal, Kartik Thakral, Puspita Majumdar, Mayank Vatsa, Richa Singh 0001 |
FG | 5 |
| 2023 | PhygitalNet: Unified Face Presentation Attack Detection via One-Class Isolation LearningabstractFace biometric systems are shown to be vulnerable to various kinds of presentation attacks including physical and digital attacks. Existing research generally focuses on individual attacks and very few focus on generalizability across digital and physical attacks. In this research, we propose PhygitalNet model that generalizes to both physical and digital presentation attacks on face biometric systems. The proposed model is based on novel one-class iSOLatiOn Learning (SOLO Learning) which is a two-step training process aimed at reducing of the covariate shift between the bonafide samples of the physical as well as digital attack dataset in the pre-training step. In the downstream step, the algorithm introduces a novel single-class iSOLatiOn loss (SOLO loss) function that isolates the samples belonging to the bonafide class away from the samples of the attacked class for both the attack methods. Experimental results show that PhygitalNet achieves a significant performance gain when compared with the baseline techniques, evaluated on a combination of MLFP, MSU-MFSD dataset (for physical attack) and FaceForensics++ (for digital attack) datasets. Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
FG | 4 |
| 2023 | Leveraging Synthetic Data and Hard Pair Mining for Selfie vs ID Face VerificationabstractThis paper delves into the challenging task of selfie vs ID face verification which involves matching high-resolution selfies with low-resolution faces extracted from scanned ID documents. Existing face verification models often face performance degradation when confronted with this task, mainly due to disparities in data distributions, such as age-difference, degradation due to scanning, and difference in appearance. To address this issue and enhance performance, the paper explores the implementation of facial quality assessment and hard-pair mining techniques. In addition, the paper investigates the potential of synthetic data for training face verification models tailored for this specific task. The integration of synthetic data as an alternative training source is explored to improve robustness and overcome legal and privacy concerns arising from authentic datasets. By combining hard pair mining, facial quality assessment, and the utilization of synthetic data, this paper presents a comprehensive framework that aims to achieve improved face verification results in the complex scenario of selfie vs ID matching. The goal is to optimize the models’ performance and enhance their ability to accurately match selfies with the corresponding ID images, even under challenging conditions. Shivang Agarwal, Jyoti Chaudhary, Hard Savani, Mayank Vatsa, Richa Singh 0001, Shyam Prasad Adhikari, Sangeeth Reddy, Kshitij Agrawal, Hemant Misra |
IJCB | 6 |
| 2023 | UG-LDFace: Unified and Generalized Framework for Long-Range Disguised Face RecognitionabstractLong-range, low-resolution videos have widespread applications in active monitoring, crowd counting, traffic analysis, and person verification/identification. The problem of analyzing faces in such an environment is exacerbated by the presence of disguise and occlusion. This research presents UG-LDFace, a novel face recognition model to address this arduous challenge. A single-stage unified framework is proposed that integrates two feature refinement techniques: feature enhancement for low-resolution data and feature selection for disguised faces. The proposed model also comprises a revised distribution technique to generalize UG-LDFace on unseen data. The proposed approach shows its efficacy on five different datasets, DroneSURF, SCface, D-LORD, DSIMF, and LFW, containing various levels of occlusion and low-resolution data. Muskan Dosi, Chiranjeev Chiranjeev, Richa Singh 0001, Mayank Vatsa |
IJCB | 3 |
| 2023 | SV-DeiT: Speaker Verification with DeiTCap Spoofing DetectionabstractAs advancements in automatic speech generation continue to progress, the ability to distinguish between real and fake samples has diminished. In addition, current spoofing detection algorithms struggle to perform well on new and unseen test distributions. To address these challenges, this paper presents two contributions. First, inspired by the success of transformer and capsule networks in high representation capabilities, we propose the DeiTCap spoof detection network on spectrogram audio features. This framework utilizes multi-head attention, sub-entities (capsules) in the audio domain and a modified routing algorithm to identify capsule agreement. The proposed spoof detection algorithm is integrated into the spoofing aware speaker recognition framework SV-DeiT. Second, we introduce a novel text-to-speech dataset TRADIF created with cutting-edge transformers and diffusion models to evaluate the generalizability of countermeasure systems. Our proposed DeiT-Cap achieves an EER of 1.08% on the evaluation set of the ASVSpoof2019 LA dataset. Moreover, the proposed network demonstrates strength in cross-domain training-testing with two different datasets, highlighting its robustness and versatility. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 3 |
| 2023 | On AI-Assisted Pneumoconiosis Detection from Chest X-raysabstractAccording to theWorld Health Organization, Pneumoconiosis affects millions of workers globally, with an estimated 260,000 deaths annually. The burden of Pneumoconiosis is particularly high in low-income countries, where occupational safety standards are often inadequate, and the prevalence of the disease is increasing rapidly. The reduced availability of expert medical care in rural areas, where these diseases are more prevalent, further adds to the delayed screening and unfavourable outcomes of the disease. This paper aims to highlight the urgent need for early screening and detection of Pneumoconiosis, given its significant impact on affected individuals, their families, and societies as a whole. With the help of low-cost machine learning models, early screening, detection, and prevention of Pneumoconiosis can help reduce healthcare costs, particularly in low-income countries. In this direction, this research focuses on designing AI solutions for detecting different kinds of Pneumoconiosis from chest X-ray data. This will contribute to the Sustainable Development Goal 3 of ensuring healthy lives and promoting well-being for all at all ages, and present the framework for data collection and algorithm for detecting Pneumoconiosis for early screening. The baseline results show that the existing algorithms are unable to address this challenge. Therefore, it is our assertion that this research will improve state-of-the-art algorithms of segmentation, semantic segmentation, and classification not only for this disease but in general medical image analysis literature. Yasmeena Akhter, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa, Santanu Chaudhury |
IJCAI | 3 |
| 2023 | NutriAI: AI-Powered Child Malnutrition Assessment in Low-Resource EnvironmentsabstractMalnutrition among infants and young children is a pervasive public health concern, particularly in developing countries where resources are limited. Millions of children globally suffer from malnourishment and its complications1. Despite the best efforts of governments and organizations, malnourishment persists and remains a leading cause of morbidity and mortality among children under five. Physical measurements, such as weight, height, middle-upper-arm-circumference (muac), and head circumference are commonly used to assess the nutritional status of children. However, this approach can be resource-intensive and challenging to carry out on a large scale. In this research, we are developing NutriAI, a low-cost solution that leverages small sample size classification approach to detect malnutrition by analyzing 2D images of the subjects in multiple poses. The proposed solution will not only reduce the workload of health workers but also provide a more efficient means of monitoring the nutritional status of children. On the dataset prepared as part of this research, the baseline results highlight that the modern deep learning approaches can facilitate malnutrition detection via anthropometric indicators in the presence of diversity with respect to age, gender, physical characteristics, and accessories including clothing. Misaal Khan, Shivang Agarwal, Mayank Vatsa, Richa Singh 0001 |
IJCAI | 4 |
| 2023 | Long-term Monitoring of Bird Flocks in the WildabstractMonitoring and analysis of wildlife are key to conservation planning and conflict management. The widespread use of camera traps coupled with AI-based analysis tools serves as an excellent example of successful and non-invasive use of technology for design, planning, and evaluation of conservation policies. As opposed to the typical use of camera traps that capture still images or short videos, in this project, we propose to analyze longer term videos monitoring a large flock of birds. This project, which is part of the NSF-TIH Indo-US joint R&D partnership, focuses on solving challenges associated with the analysis of long-term videos captured at feeding grounds and nesting sites, among other such locations that host large flocks of migratory birds. We foresee that the objectives of this project would lead to datasets and benchmarking tools as well as novel algorithms that would be instrumental in developing automated video analysis tools that could in turn help understand individual and social behavior of birds. The first of the key outcomes of this research will include the curation of challenging, real-world datasets for benchmarking various image and video analytics algorithms for tasks such as counting, detection, segmentation, and tracking. Our recent efforts towards this outcome is a curated dataset of 812 high-resolution, point-annotated, images (4K - 32MP) of a flock of Demoiselle cranes (Anthropoides virgo) taken from their feeding site at Khichan, Rajasthan, India. The average number of birds in each image is about 207, with a maximum count of 1500. The benchmark experiments show that state-of-the-art vision techniques struggle with tasks such as segmentation, detection, localization, and density estimation for the proposed dataset. Over the execution of this open science research, we will be scaling this dataset for segmentation and tracking in videos, as well as developing novel techniques for video analytics for wildlife monitoring. Kshitiz, Sonu Sreshtha, Ramy Mounir, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar |
IJCAI | 5 |
| 2023 | Uncovering the Deceptions: An Analysis on Audio Spoofing Detection and Future ProspectsabstractAudio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications including person authentication to banking to virtual assistants. Research has shown that these systems are also susceptible to spoofing and attacks. Therefore, protecting audio processing systems against fraudulent activities such as identity theft, financial fraud, and spreading misinformation, is of paramount importance. This paper reviews the current state-of-the-art techniques for detecting audio spoofing and discusses the current challenges along with open research problems. The paper further highlights the importance of considering the ethical and privacy implications of audio spoofing detection systems. Lastly, the work aims to accentuate the need for building more robust and generalizable methods, the integration of automatic speaker verification and countermeasure systems, and better evaluation protocols. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCAI | 3 |
| 2023 | Misclassifications of Contact Lens Iris PAD Algorithms: Is it Gender Bias or Environmental Conditions?abstractOne of the critical steps in biometrics pipeline is detection of presentation attacks, a physical adversary. Several presentation (adversary) attack detection (PAD) algorithms, including iris PAD, have been proposed and have shown superlative performance. However, a recent study, on a small-scale database, has highlighted that iris PAD may have gender biases. In this research, we present a rigorous study on gender bias in iris presentation attack detection algorithms using a large-scale and gender-balanced database. The paper provides several interesting observations which can help in building future presentation attack detection algorithms with aim of fair treatment of each demography. In addition, we also present a robust iris presentation attack detection algorithm by combining gender-covariate based classifiers. The proposed robust classifier not only reduces the difference in accuracy between different genders but also improves the overall performance of the PAD system. Akshay Agarwal 0001, Nalini K. Ratha, Afzel Noore, Richa Singh 0001, Mayank Vatsa |
WACV | 4 |
| 2023 | Uniform misclassification loss for unbiased model predictionabstractDeep learning algorithms have achieved tremendous success over the past few years. However, the biased behavior of deep models, where the models favor/disfavor certain demographic subgroups, is a major concern in the deep learning community. Several adverse consequences of biased predictions have been observed in the past. One solution to alleviate the problem is to train deep models for fair outcomes. Therefore, in this research, we propose a novel loss function, termed as Uniform Misclassification Loss (UML) to train deep models for unbiased outcomes. The proposed UML function penalizes the model for the worst-performing subgroup for mitigating bias and enhancing the overall model performance. The proposed loss function is also effective while training with imbalanced data as well. Further, a metric, Joint Performance Disparity Measure (JPD) is introduced to jointly measure the overall model performance and the bias in model prediction. Multiple experiments have been performed on four publicly available datasets for facial attribute prediction and comparisons are performed with existing bias mitigation algorithms. Experimental results are reported using performance and bias evaluation metrics . The proposed loss function outperforms existing bias mitigation algorithms that showcase its effectiveness in obtaining unbiased outcomes and improved performance. Puspita Majumdar, Mayank Vatsa, Richa Singh 0001 |
Pattern Recognit. | 3 |
| 2022 | Anatomizing Bias in Facial AnalysisabstractExisting facial analysis systems have been shown to yield biased results against certain demographic subgroups. Due to its impact on society, it has become imperative to ensure that these systems do not discriminate based on gender, identity, or skin tone of individuals. This has led to research in the identification and mitigation of bias in AI systems. In this paper, we encapsulate bias detection/estimation and mitigation algorithms for facial analysis. Our main contributions include a systematic review of algorithms proposed for understanding bias, along with a taxonomy and extensive overview of existing bias mitigation algorithms. We also discuss open challenges in the field of biased facial analysis. Richa Singh 0001, Puspita Majumdar, Surbhi Mittal, Mayank Vatsa |
AAAI | 1 |
| 2022 | Mannet: A Large-Scale Manipulated Image Detection Dataset And Baseline EvaluationsabstractThe sharing of fake content on social media platforms has become a major concern. In many cases, the same content with small variations is shared multiple times on different social media platforms. This leads to the circulation of manipulated content on the web. With the rapid advancement in deep learning algorithms, the generation of manipulated images with small variations in original images has become an easy task. These contents raise serious concerns when used for malicious activities. Therefore, detection of manipulated contents is of paramount importance. However, no large-scale dataset having manipulated images generated using both handcrafted and deep learning algorithms is available. Therefore, in this research, we have proposed a large dataset with more than 5.5 million images, termed as ManNet dataset. Additionally, we have benchmarked the performance of existing algorithms for manipulated image detection. The experimental results highlight that inter-set (disjoint training testing) evaluations are the major challenge of manipulated image detection. Saheb Chhabra, Puspita Majumdar, Richa Singh 0001, Mayank Vatsa |
ICASSP | 4 |
| 2022 | In-group and Out-group Performance Bias in Facial Retouching DetectionabstractAccuracy alone is not sufficient to establish the efficacy of an AI algorithm-issues of demographic bias are an important area of concern. Demographic bias in face recognition algorithms has attracted more attention from the re-search community to date, but bias can also be a problem for face image analysis algorithms, such as detection of manipulated face images. In this paper, we investigate performance of humans and algorithms at detecting retouched face images of subjects from different origin (America, India, China) and gender groups. To be representative of the state of retouching detection, we use eight different algorithms from the literature. In addition to overall human accuracy, differences across origin and gender of the human performing the task are analyzed. We observe different bias patterns, such as algorithms show higher in-group accuracy than out-group, while the extent of retouching and familiarity drives differences in detection accuracy for humans. This is the first work to analyze and compare bias exhibited by humans and algorithms in similar tasks of detecting retouched face images. Aparna Bharati, Emma Connors, Mayank Vatsa, Richa Singh 0001, Kevin W. Bowyer |
IJCB | 4 |
| 2022 | DeePhy: On Deepfake PhylogenyabstractDeepfake refers to tailored and synthetically generated videos which are now prevalent and spreading on a large scale, threatening the trustworthiness of the information available online. While existing datasets contain different kinds of deepfakes which vary in their generation technique, they do not consider progression of deepfakes in a “phylogenetic” manner. It is possible that an existing deepfake face is swapped with another face. This process of face swapping can be performed multiple times and the resultant deepfake can be evolved to confuse the deepfake detection algorithms. Further, many databases do not provide the employed generative model as target labels. Model attribution helps in enhancing the explainability of the detection results by providing information on the generative model employed. In order to enable the research community to address these questions, this paper proposes DeePhy, a novel Deepfake Phylogeny dataset which consists of 5040 deep-fake videos generated using three different generation techniques. There are 840 videos of one-time swapped deep-fakes, 2520 videos of two-times swapped deepfakes and 1680 videos of three-times swapped deepfakes. With over 30 GBs in size, the database is prepared in over 1100 hours using 18 GPUs of 1,352 GB cumulative memory. We also present the benchmark on DeePhy dataset using six deep-fake detection algorithms. The results highlight the need to evolve the research of model attribution of deepfakes and generalize the process over a variety of deepfake generation techniques. The database is available at: http://iab-rubric.org/deephy-database Kartik Narayan, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
IJCB | 6 |
| 2022 | STATNet: Spectral and Temporal features based Multi-Task Network for Audio Spoofing DetectionabstractWith the rise in mobile phone users and VoIP, voice has emerged as an easy and accessible biometric modality for identification or verification tasks. Given the increasing usage of voice biometrics, the security of these systems is also of paramount importance. Researchers have demon-strated that Automatic Speaker Verification (ASV) systems are prone to spoofing attacks like synthetic speech or fake speech, which can be used maliciously for a variety of tasks such as impersonation, fake news spreading, and opinion formation. This research proposes a deep convolution-based multi-task network which performs both spoof detection and source identification for synthetic speech. The pro-posed model is evaluated on three datasets ASVspoof2019 LA, FOR-Norm and In-the- Wild Audio Deepfake dataset. The results demonstrate the EER of 2.456%, 0.814%, and 0.199% on the ASVspoof2019 LA, FOR-Norm, and In-the-Wild Audio Deepfake datasets. In addition, we have also demonstrated results for cross-dataset evaluation and speech source identification. Rishabh Ranjan, Mayank Vatsa, Richa Singh 0001 |
IJCB | 3 |
| 2022 | Robust IRIS Presentation Attack Detection Through Stochastic Filter NoiseabstractThe vulnerability of iris recognition algorithms against presentation attacks demands a robust defense mechanism. Much research has been done in the literature to create a robust attack detection algorithm; however, most of the algorithms suffer from generalizability, such as inter database testing or unseen attack type. The problem of attack detection can further be exacerbated if the images contain noise such as Gaussian or Salt-Pepper noise. In this research, we propose a multi-task deep learning model with a denoising convolutional skip autoencoder and a classifier to inbuilt robustness against noisy images. The Gaussian noise layer is introduced as a dropout between the encoder network’s hidden layers, which helps the model learn generalized features that are robust to data noise. The proposed algorithm is evaluated on multiple presentation attack databases and extensive experiments across different noise types and a comparison with other deep learning models show the generalizability and efficacy of the proposed model. Vishi Jain, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha |
ICPR | 3 |
| 2022 | MTCD: Cataract detection via near infrared eye images
Pavani Tripathi, Yasmeena Akhter, Mahapara Khurshid, Aditya Lakra, Rohit Keshari, Mayank Vatsa, Richa Singh 0001 |
Comput. Vis. Image Underst. | 7 |
| 2022 | DeriveNet for (Very) Low Resolution Image ClassificationabstractImages captured from a distance often result in (very) low resolution (VLR/LR) region of interest, requiring automated identification. VLR/LR images (or regions of interest) often contain less information content, rendering ineffective feature extraction and classification. To this effect, this research proposes a novel DeriveNet model for VLR/LR classification, which focuses on learning effective class boundaries by utilizing the class-specific domain knowledge. DeriveNet model is jointly trained via two losses: (i) proposed Derived-Margin softmax loss and (ii) the proposed Reconstruction-Center (ReCent) loss. The Derived-Margin softmax loss focuses on learning an effective VLR classifier while explicitly modeling the inter-class variations. The ReCent loss incorporates domain information by learning a HR reconstruction space for approximating the class variations for the VLR/LR samples. It is utilized to derive inter-class margins for the Derived-Margin softmax loss. The DeriveNet model has been trained with a novel Multi-resolution Pyramid based data augmentation which enables the model to learn from varying resolutions during training. Experiments and analysis have been performed on multiple datasets for (i) VLR/LR face recognition, (ii) VLR digit classification, and (iii) VLR/LR face recognition from drone-shot videos. The DeriveNet model achieves state-of-the-art performance across different datasets, thus promoting its utility for several VLR/LR classification tasks. Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Multi-task driven explainable diagnosis of COVID-19 using chest X-ray images
Aakarsh Malhotra, Surbhi Mittal, Puspita Majumdar, Saheb Chhabra, Kartik Thakral, Mayank Vatsa, Richa Singh 0001, Santanu Chaudhury, Ashwin Pudrod, Anjali Agrawal |
Pattern Recognit. | 7 |
| 2022 | Enhanced iris presentation attack detection via contraction-expansion CNN
Akshay Agarwal 0001, Afzel Noore, Mayank Vatsa, Richa Singh 0001 |
Pattern Recognit. Lett. | 4 |
| 2022 | Disguise Resilient Face VerificationabstractWith increasing usage of face recognition algorithms, it is well established that external artifacts and makeup accessories can be applied to different facial features such as eyes, nose, mouth, and cheek, to obfuscate one’s identity or to impersonate someone else’s identity. Recognizing faces in the presence of these artifacts comprises the problem of disguised face recognition, which is one of the most arduous covariates of face recognition. The challenge becomes exacerbated when disguised faces are captured in real-time environment, with low resolution images. To address the challenge of disguised face recognition, this paper first proposes a novel multi-objective encoder-decoder network, termed as DED-Net. DED-Net attempts to learn the class variations in the feature space generated by both disguised as well non-disguised images, using a combination of Mahalanobis and Cosine distance metrics, along with Mutual Information based supervision. The DED-Net is then extended to learn from the local and global features of both disguised and non-disguised face images for efficient face recognition, and the complete framework is termed as Disguise Resilient (D-Res) framework. The efficacy of the proposed framework has been demonstrated on two real-world benchmark datasets: Disguised Faces in the Wild (DFW) 2018 and DFW2019 competition datasets. In addition, this research also emphasizes on the importance of recognizing disguised faces in low resolution settings and proposes three experimental protocols to simulate the real-world surveillance scenario. To this effect, benchmark results have been shown on seven protocols for three low resolution settings ($32\times 32$,$24\times 24$, and$16\times 16$) of the two DFW benchmark datasets. The results demonstrate superior performance of the D-Res framework, in comparison with benchmark algorithms. For example, an improvement of around 3% is observed on the Overall protocol of the DFW2019 dataset, where the D-Res framework achieves 96.3%. Experiments have also been performed on benchmark face verification datasets (LFW, YTF, and IJB-B), where the D-Res framework achieves improved verification accuracy. Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Crafting Adversarial Perturbations via Transformed Image Component SwappingabstractAdversarial attacks have been demonstrated to fool the deep classification networks. There are two key characteristics of these attacks: firstly, these perturbations are mostly additive noises carefully crafted from the deep neural network itself. Secondly, the noises are added to the whole image, not considering them as the combination of multiple components from which they are made. Motivated by these observations, in this research, we first study the role of various image components and the impact of these components on the classification of the images. These manipulations do not require the knowledge of the networks and external noise to function effectively and hence have the potential to be one of the most practical options for real-world attacks. Based on the significance of the particular image components, we also propose a transferable adversarial attack against unseen deep networks. The proposed attack utilizes the projected gradient descent strategy to add the adversarial perturbation to the manipulated component image. The experiments are conducted on a wide range of networks and four databases including ImageNet and CIFAR-100. The experiments show that the proposed attack achieved better transferability and hence gives an upper hand to an attacker. On the ImageNet database, the success rate of the proposed attack is up to 88.5%, while the current state-of-the-art attack success rate on the database is 53.8%. We have further tested the resiliency of the attack against one of the most successful defenses namely adversarial training to measure its strength. The comparison with several challenging attacks shows that: (i) the proposed attack has a higher transferability rate against multiple unseen networks and (ii) it is hard to mitigate its impact. We claim that based on the understanding of the image components, the proposed research has been able to identify a newer adversarial attack unseen so far and unsolvable using the current defense mechanisms. Akshay Agarwal 0001, Nalini K. Ratha, Mayank Vatsa, Richa Singh 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | DAMAD: Database, Attack, and Model Agnostic Adversarial Perturbation DetectorabstractAdversarial perturbations have demonstrated the vulnerabilities of deep learning algorithms to adversarial attacks. Existing adversary detection algorithms attempt to detect the singularities; however, they are in general, loss-function, database, or model dependent. To mitigate this limitation, we propose DAMAD-a generalized perturbation detection algorithm which is agnostic to model architecture, training data set, and loss function used during training. The proposed adversarial perturbation detection algorithm is based on the fusion of autoencoder embedding and statistical texture features extracted from convolutional neural networks. The performance of DAMAD is evaluated on the challenging scenarios of cross-database, cross-attack, and cross-architecture training and testing along with traditional evaluation of testing on the same database with known attack and model. Comparison with state-of-the-art perturbation detection algorithms showcase the effectiveness of the proposed algorithm on six databases: ImageNet, CIFAR-10, Multi-PIE, MEDS, point and shoot challenge (PaSC), and MNIST. Performance evaluation with nearly a quarter of a million adversarial and original images and comparison with recent algorithms show the effectiveness of the proposed algorithm. Akshay Agarwal 0001, Gaurav Goswami, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Role of Optimizer on Network Fine-tuning for Adversarial Robustness (Student Abstract)abstractThe solutions proposed in the literature for adversarial robustness are either not effective against the challenging gradient-based attacks or are computationally demanding, such as adversarial training. Adversarial training or network training based data augmentation shows the potential to increase the adversarial robustness. While the training seems compelling, it is not feasible for resource-constrained institutions, especially academia, to train the network from scratch multiple times. The two fold contributions are: (i) providing an effective solution against white-box adversarial attacks via network fine-tuning steps and (ii) observing the role of different optimizers towards robustness. Extensive experiments are performed on a range of databases, including Fashion-MNIST and a subset of ImageNet. It is found that the few steps of network fine-tuning effectively increases the robustness of both shallow and deep architectures. To know other interesting observations, especially regarding the role of the optimizer, refer to the paper. Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001 |
AAAI | 3 |
| 2021 | NEAP-F: Network Epoch Accuracy Prediction Framework (Student Abstract)abstractRecent work in neural architecture search has spawned interest in algorithms that can predict the performance of convolutional neural networks using minimum time and computation resources. We propose a new framework, Network Epoch Accuracy Prediction Framework (NEAP-F) which can predict the testing accuracy achieved by a convolutional neural network in one or more epochs. We introduce a novel approach to generate vector representations for networks, and encode ``ease" of classifying image datasets into a vector. For vector representations of networks, we focus on the layer parameters and connections between the network layers. A network achieves different accuracies on different image datasets; therefore, we use the image dataset characteristics to create a vector signifying the ``ease" of classifying the image dataset. After generating these vectors, the prediction models are trained with architectures having skip connections seen in current state-of-the-art architectures. The framework predicts accuracies in order of milliseconds, demonstrating its computational efficiency. It can be easily applied to neural architecture search methods to predict the performance of candidate networks and can work on unseen datasets as well. Arushi Chauhan, Mayank Vatsa, Richa Singh 0001 |
AAAI | 3 |
| 2021 | On Learning Deep Models with Imbalanced Data DistributionabstractThe availability of large training data has led to the development of sophisticated deep learning algorithms to achieve state-of-the-art performance on various tasks and several applications have been benefited immensely. Despite the unparalleled success, the performance of deep learning algorithms depends significantly on the training data distribution. An imbalance in training data distribution affects the performance of deep models. Our research focuses on designing and developing solutions for different real-world problems, specifically related to facial analytic tasks, with imbalanced data distribution. These problems include injured face recognition, fake image detection, and estimation and mitigation of bias in model prediction. Puspita Majumdar, Richa Singh 0001, Mayank Vatsa |
AAAI | 2 |
| 2021 | Detection of Digital Manipulation in Facial Images (Student Abstract)abstractAdvances in deep learning have enabled the creation of photo-realistic DeepFakes by switching the identity or expression of individuals. Such technology in the wrong hands can seed chaos through blackmail, extortion, and forging false statements of influential individuals. This work proposes a novel approach to detect forged videos by magnifying their temporal inconsistencies. A study is also conducted to understand role of ethnicity bias due to skewed datasets on deepfake detection. A new dataset comprising forged videos of Indian ethnicity individuals is presented to facilitate this study. Aman Mehra, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001 |
AAAI | 4 |
| 2021 | Semi-Supervised Learning via Triplet Network Based Active Learning (Student Abstract)abstractIn recent years, deep learning models have pushed state-of-the-art accuracies for several machine learning tasks. However, such models require a large amount of (supervised) data for training. While unlabelled data is available in abundance, manually labeling them is very costly. Active learning techniques helps in utilizing unlabelled data which may result in an improved classification model. In this research, we present an active learning algorithm which can help in increasing performance of deep learning models by using large amount of unlabelled data. A novel active learning algorithm, Triplet AL is proposed which uses a triplet network to select samples from an unlabelled data set. Previous active learning methods rely on classification model's final prediction scores as a measure of confidence for an unlabelled sample. We propose a more reliable confidence measure, termed as Top-Two-Margin which is given by the Triplet Network. The proposed algorithm shows improved performance compared to other active learning approaches. Divyanshu Sundriyal, Soumyadeep Ghosh, Mayank Vatsa, Richa Singh 0001 |
AAAI | 4 |
| 2021 | MD-CSDNetwork: Multi-Domain Cross Stitched Network for Deepfake DetectionabstractThe rapid progress in the ease of creating and spreading ultra-realistic media over social platforms calls for an urgent need to develop a generalizable deepfake detection technique. It has been observed that current deepfake generation methods leave discriminative artifacts in the frequency spectrum of fake images and videos. Inspired by this observation, in this paper, we present a novel approach, termed as MD-CSDNetwork, for combining the features in the spatial and frequency domains to extract a shared discriminative representation for classifying deepfakes. MD-CSDNetwork is a novel cross-stitched network with two parallel branches carrying spatial and frequency information, respectively. We hypothesize that these multi-domain input data streams can be considered as related supervisory signals and can ensure better performance and generalization. Further, the concept of cross-stitch connections is utilized where they are inserted between the two branches to learn an optimal combination of domain-specific and shared representations from other domains automatically. Extensive experiments are conducted on the popular benchmark datasets. We report improvements over all the manipulation types in the FaceForensics++ dataset and comparable results with state-of-the-art methods for cross-database evaluation on the Celeb-DF dataset and the Deepfake Detection Dataset. Aayushi Agarwal 0001, Akshay Agarwal 0001, Sayan Sinha, Mayank Vatsa, Richa Singh 0001 |
FG | 5 |
| 2021 | When Sketch Face Recognition Meets Mask Obfuscation: Database and BenchmarkabstractDuring this unprecedented time of the COVID19 pandemic, wearing face masks has become a necessity. While these masks aim to secure an individual from getting infected by any kind of viruses including COVID-19; they significantly obfuscate the identity. The situation becomes even worse when an attacker performs a crime and the place does not have any surveillance cameras. The identification of criminals in such conditions highly depends on the witnesses and generation of sketches based on their description. To the best of our knowledge, in the literature, no work has been performed for matching sketch images with masks. In this research, we have first created the mask sketch face database using more than 50 identities. The sketch images are generated using a different variant of pencils, which can be seen as different sketch artists. The recognition experiments are performed using state-of-the-art face embedding networks including ArcFace and DeepID which show that the recognition performance degrades significantly when the sketch mask images are used for identification. In another set of experiments, it is observed that the recognition algorithm is robust in handling the digital face mask images. However, the ineffectiveness in handling the variations that occurred due to sketches is a serious concern and needs attention. Akshay Agarwal 0001, Nalini K. Ratha, Mayank Vatsa, Richa Singh 0001 |
FG | 4 |
| 2021 | AECNet: Attentive EfficientNet For Crowd CountingabstractIn the COVID pandemic situation, crowd counting became one of the tools to monitor if the social-distancing norms are being followed or not. However, in designing crowd counting algorithm, there are several challenges such as background noise, camera-to-objects distance, occlusion, and variations due to illumination, scale, and viewpoint. In this research, we propose a novel pipeline for density estimation in crowd counting. The proposed pipeline makes use of an encoder-decoder-based architecture in which we explore the family of EfficientN ets for the encoder architecture. For the decoder, we propose a deeper attention network to assist the model in a better distinction between foreground and background pixels. We empirically show that for a crowd counting dataset, the use of average pooling operation for any backbone architecture of encoder gives a significant improvement in performance. In terms of Mean Absolute Error, the proposed pipeline outperforms existing state-of-the-art techniques by a large margin on large-scale and small-scale counting datasets, UCF-QNRF and UCF _CC_50 dataset. We also achieve state-of-the-art results on the ShanghaiTech and Mall datasets. We additionally propose a crowd counting dataset captured using drones. We perform benchmark experiments on this dataset with existing and the proposed methods. The proposed dataset can be found at http://www.iab-rubric.org/resources/CrowdUAV.html. Muskan Dosi, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
FG | 5 |
| 2021 | RGB-D Face Recognition using Reconstruction based Shared RepresentationabstractLow cost time-of-flight based depth sensors such as Kinect have opened new avenues for their usage in video surveillance scenarios. RGB-D images obtained from such sensors have shown their utility in improved face recognition capabilities. Generally, existing RGB-D face recognition methods fuse the depth information with RGB information which results in enhanced recognition performance. However, in the real world surveillance scenarios, cameras are placed at a distance too large for low cost depth sensors to capture good quality depth information. Such poor quality depth information may not contribute significantly to face recognition. In this paper, we present a novel representation learning algorithm by learning shared representation of RGB and depth information using a reconstruction based deep neural network. The proposed network, once trained in the offline mode, can generate a shared representation of RGB and depth data using only the RGB image. This feature rich representation is then utilized for face identification. This allows the framework to be used in scenarios where low quality or no depth image is captured. Experiments on multiple real world databases show the effectiveness of the proposed approach. Soumyadeep Ghosh, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
FG | 2 |
| 2021 | Dual Sensor Indian Masked Face DatasetabstractWith the advancements in deep learning technologies, real-world applications like face detection, gender prediction, and face recognition have achieved human-level performance. However, the emergence of the COVID-19 pandemic brought new challenges to existing deep learning algorithms. People are forced to wear a mask to limit the spread of COVID-19. These face masks occlude a significant portion of the face, thereby posing multiple challenges to existing algorithms. Images captured using surveillance cameras have a low resolution which hinders the model performance. Along with this, skin tone, ethnicity and attire also play a significant role in detection and recognition performance. India is a large country with huge diversity in skin tone and attire of the people. To address the challenges due to masks in the Indian context, we propose a novel Dual Sensor Indian Masked Face (DS- IMF) dataset, which contains images captured in constrained environmental settings with a variety of masks and degrees of occlusion. Multiple experiments are performed on the DS- IMF dataset at different resolutions. Experimental results demonstrate the limitations of existing algorithms on low-resolution masked face images. The proposed dataset can be found at http://www.iab-rubric.org/resources/dsimf.html. Shiksha Mishra, Puspita Majumdar, Muskan Dosi, Mayank Vatsa, Richa Singh 0001 |
FG | 5 |
| 2021 | Intelligent and Adaptive Mixup Technique for Adversarial RobustnessabstractDeep neural networks are generally trained using large amounts of data to achieve state-of-the-art accuracy in many possible computer vision and image analysis applications ranging from object recognition to natural language processing. It is also claimed that these networks can memorize the data which can be extracted from the network parameters such as weights and gradient information. The adversarial vulnerability of the deep networks is usually evaluated on the unseen test set of the databases. If the network is memorizing the data, then the small perturbation in the training image data should not drastically change its performance. Based on this assumption, we first evaluate the robustness of deep neural networks on small perturbations added in the training images used for learning the parameters of the network. It is observed that, even if the network has seen the images it is still vulnerable to these small perturbations. Further, we propose a novel data augmentation technique to increase the robustness of deep neural networks to such perturbations. Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
ICIP | 3 |
| 2021 | Indian Masked Faces in the Wild DatasetabstractDue to the COVID-19 pandemic, wearing face masks has become a mandate in public places worldwide. Face masks occlude a significant portion of the facial region. Additionally, people wear different types of masks, from simple ones to ones with graphics and prints. These pose new challenges to face recognition algorithms. Researchers have recently proposed a few masked face datasets for designing algorithms to overcome the challenges of masked face recognition. However, existing datasets lack the cultural diversity and collection in the unrestricted settings. Country like India with attire diversity, people are not limited to wearing traditional masks but also clothing like a thin cotton printed towel (locally called as “gamcha”), “stoles”, and “handkerchiefs” to cover their faces. In this paper, we present a novel Indian Masked Faces in the Wild (IMFW) dataset which contains images with variations in pose, illumination, resolution, and the variety of masks worn by the subjects. We have also benchmarked the performance of existing face recognition models on the proposed IMFW dataset. Experimental results demonstrate the limitations of existing algorithms in presence of diverse conditions. Shiksha Mishra, Puspita Majumdar, Richa Singh 0001, Mayank Vatsa |
ICIP | 3 |
| 2021 | Class Equilibrium using Coulomb's LawabstractProjection algorithms learn a transformation function to project the data from input space to the feature space, with the objective of increasing the inter-class distance. However, increasing the inter-class distance can affect the intra-class distance. Maintaining an optimal inter-class separation among the classes without affecting the intra-class distance of the data distribution is a challenging task. In this paper, inspired by the Coulomb's law of Electrostatics, we propose a new algorithm to compute the equilibrium space of any data distribution where the separation among the classes is optimal. The algorithm further learns the transformation between the input space and equilibrium space to perform classification in the equilibrium space. The performance of the proposed algorithm is evaluated on four publicly available datasets at three different resolutions. It is observed that the proposed algorithm performs well for lowresolution images. Saheb Chhabra, Puspita Majumdar, Mayank Vatsa, Richa Singh 0001 |
IJCNN | 4 |
| 2021 | Enhancing Fine-Grained Classification for Low Resolution ImagesabstractLow resolution fine-grained classification has widespread applicability for applications where data is captured at a distance such as surveillance and mobile photography. While fine-grained classification with high resolution images has received significant attention, limited attention has been given to low resolution images. These images suffer from the inherent challenge of limited information content and the absence of fine details useful for sub-category classification. This results in low inter-class variations across samples of visually similar classes. In order to address these challenges, this research proposes a novel attribute-assisted loss, which utilizes ancillary information to learn discriminative features for classification. The proposed loss function enables a model to learn class-specific discriminative features, while incorporating attribute-level separability. Evaluation is performed on multiple datasets with different models, for four resolutions varying from$32\times 32$to$224\times 224$. Different experiments demonstrate the efficacy of the proposed attribute-assisted loss for low resolution fine-grained classification. Maneet Singh, Shruti Nagpal, Mayank Vatsa, Richa Singh 0001 |
IJCNN | 4 |
| 2021 | Understanding Neural Responses to Face Verification of Cross-Domain RepresentationsabstractFace verification involves identifying whether two faces belong to the same person or not. It relies heavily upon face perception, processing, and the decision making of an individual. This research studies cross-domain face verification, where one face image belongs to a controlled, well-illuminated environment, while the other is of a varying representation having differences in image type or quality. Specifically, two cross-domain face verification tasks are analyzed: controlled-low resolution and controlled-sketch face verification. functional Magnetic Resonance Imaging (fMRI) data has been collected for 23 participants of two ethnic groups while performing face verification. Statistical comparisons were performed with same-domain controlled face verification for both the tasks. Our findings reveal regions of Right Frontal Gyrus, Bilateral Insula, and Right Middle Cingulate Cortex demonstrating higher activation for controlled-sketch face verification, as compared to controlled face verification. Similar analysis were performed for controlled-low resolution face verification, where regions responsible for higher visual load and difficult tasks result in higher activation. Further, stimuli ethnicity differences influence activations for low-resolution face verification but do not affect sketch face verification. Regions of Right Middle Occipital Gyrus and Right Fusiform Gyrus present higher activity, suggesting increased face processing effort for within ethnicity low resolution face verification. We believe the findings of this research will help enable further development in the field of brain-inspired facial recognition algorithms. Maneet Singh, Shruti Nagpal, Daksha Yadav, Naman Kohli, Prateekshit Pandey, Gokulraj Prabhakaran, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Julie Brefczynski-Lewis, Harsh Mahajan |
IJCNN | 7 |
| 2021 | Guest Editorial: Adversarial Deep Learning in Biometrics & Forensics
Rama Chellappa, Diego Gragnaniello, Chang-Tsun Li, Francesco Marra, Richa Singh 0001 |
Comput. Vis. Image Underst. | 5 |
| 2021 | Discriminative shared transform learning for sketch to image matching
Shruti Nagpal, Maneet Singh, Richa Singh 0001, Mayank Vatsa |
Pattern Recognit. | 3 |
| 2021 | Cognitive data augmentation for adversarial defense via pixel masking
Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
Pattern Recognit. Lett. | 3 |
| 2021 | Improving face recognition performance using TeCS2 dictionary
Saksham Suri, Anush Sankaran, Mayank Vatsa, Richa Singh 0001 |
Pattern Recognit. Lett. | 4 |
| 2021 | Image Transformation-Based Defense Against Adversarial Perturbation on Deep Learning ModelsabstractDeep learning algorithms provide state-of-the-art results on a multitude of applications. However, it is also well established that they are highly vulnerable to adversarial perturbations. It is often believed that the solution to this vulnerability of deep learning systems must come from deep networks only. Contrary to this common understanding, in this article, we propose a non-deep learning approach that searches over a set of well-known image transforms such as Discrete Wavelet Transform and Discrete Sine Transform, and classifying the features with a support vector machine-based classifier. Existing deep networks-based defense have been proven ineffective against sophisticated adversaries, whereas image transformation-based solution makes a strong defense because of the non-differential nature, multiscale, and orientation filtering. The proposed approach, which combines the outputs of two transforms, efficiently generalizes across databases as well as different unseen attacks and combinations of both (i.e., cross-database and unseen noise generation CNN model). The proposed algorithm is evaluated on large scale databases, including object database (validation set of ImageNet) and face recognition (MBGC) database. The proposed detection algorithm yields at-least 84.2% and 80.1% detection accuracy under seen and unseen database test settings, respectively. Besides, we also show how the impact of the adversarial perturbation can be neutralized using a wavelet decomposition-based filtering method of denoising. The mitigation results with different perturbation methods on several image databases demonstrate the effectiveness of the proposed method. Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | On the Robustness of Face Recognition Algorithms Against Attacks and BiasabstractFace recognition algorithms have demonstrated very high recognition performance, suggesting suitability for real world applications. Despite the enhanced accuracies, robustness of these algorithms against attacks and bias has been challenged. This paper summarizes different ways in which the robustness of a face recognition algorithm is challenged, which can severely affect its intended working. Different types of attacks such as physical presentation attacks, disguise/makeup, digital adversarial attacks, and morphing/tampering using GANs have been discussed. We also present a discussion on the effect of bias on face recognition models and showcase that factors such as age and gender variations affect the performance of modern algorithms. The paper also presents the potential reasons for these challenges and some of the future research directions for increasing the robustness of face recognition models. Richa Singh 0001, Akshay Agarwal 0001, Maneet Singh, Shruti Nagpal, Mayank Vatsa |
AAAI | 1 |
| 2020 | Generalized Zero-Shot Learning via Over-Complete DistributionabstractA well trained and generalized deep neural network (DNN) should be robust to both seen and unseen classes. However, the performance of most of the existing supervised DNN algorithms degrade for classes which are unseen in the training set. To learn a discriminative classifier which yields good performance in Zero-Shot Learning (ZSL) settings, we propose to generate an Over-Complete Distribution (OCD) using Conditional Variational Autoencoder (CVAE) of both seen and unseen classes. In order to enforce the separability between classes and reduce the class scatter, we propose the use of Online Batch Triplet Loss (OBTL) and Center Loss (CL) on the generated OCD. The effectiveness of the framework is evaluated using both Zero-Shot Learning and Generalized Zero-Shot Learning protocols on three publicly available benchmark databases, SUN, CUB and AWA2. The results show that generating over-complete distributions and enforcing the classifier to learn a transform function from overlapping to non-overlapping distributions can improve the performance on both seen and unseen classes. Rohit Keshari, Richa Singh 0001, Mayank Vatsa |
CVPR | 2 |
| 2020 | Diversity Blocks for De-biasing Classification ModelsabstractRecent studies have highlighted a major caveat in various high performing automated systems for tasks such as facial analysis (e.g. gender prediction), object classification, and image to caption generation. Several of the existing systems have been shown to yield biased results towards or against a particular subgroup. The biased behavior exhibited by these models when deployed and used in a real world scenario presents with the challenge of automated systems being unfair. In this research, we propose a novel technique, diversity block, for de-biasing existing models without re-training them. The proposed technique requires small amount of training data and can be incorporated with an existing model for addressing the challenge of biased predictions. This is done by adding a diversity block and computing the prediction based on the scores of the original model and the diversity block in order to get a more confident and de-biased prediction. The efficacy of the proposed technique has been demonstrated on the task of gender prediction, along with an auxiliary case study on object classification. Shruti Nagpal, Maneet Singh, Richa Singh 0001, Mayank Vatsa |
IJCB | 3 |
| 2020 | Attack Agnostic Adversarial Defense via Visual Imperceptible BoundabstractThe high susceptibility of deep learning algorithms against structured and unstructured perturbations has motivated the development of efficient adversarial defense algorithms. However, the lack of generalizability of existing defense algorithms and the high variability in the performance of the attack algorithms for different databases raises several questions on the effectiveness of the defense algorithms. In this research, we aim to design a defense model that is robust within a certain bound against both seen and unseen adversarial attacks. This bound is related to the visual appearance of an image, and we termed it as Visual Imperceptible Bound (VIB). To compute this bound, we propose a novel method that uses the database characteristics. The VIB is further used to measure the effectiveness of attack algorithms. The performance of the proposed defense model is evaluated on the MNIST, CIFAR-10, and Tiny ImageNet databases on multiple attacks that include C&W ( l2) and DeepFool. The proposed defense model is not only able to increase the robustness against several attacks but also retain or improve the classification accuracy on an original clean test set. The proposed algorithm is attack agnostic, i.e. it does not require any knowledge of the attack algorithm. Saheb Chhabra, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa |
ICPR | 3 |
| 2020 | Generalized Iris Presentation Attack Detection Algorithm under Cross-Database SettingsabstractPresentation attacks are posing major challenges to most of the biometric modalities. Iris recognition, which is considered as one of the most accurate biometric modality for person identification, has also been shown to be vulnerable to advanced presentation attacks such as 3D contact lenses and textured lens. While in the literature, several presentation attack detection (PAD) algorithms are presented; a significant limitation is the generalizability against an unseen database, unseen sensor, and different imaging environment. To address this challenge, we propose a generalized deep learning-based PAD network, MVANet, which utilizes multiple representation layers. It is inspired by the simplicity and success of hybrid algorithm or fusion of multiple detection networks. The computational complexity is an essential factor in training deep neural networks; therefore, to reduce the computational complexity while learning multiple feature representation layers, a fixed base model has been used. The performance of the proposed network is demonstrated on multiple databases such as IIITD-WVU MUIPA and IIITD-CLI databases under cross-database training-testing settings, to assess the generalizability of the proposed algorithm. Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001 |
ICPR | 5 |
| 2020 | MixNet for Generalized Face Presentation Attack DetectionabstractThe non-intrusive nature and high accuracy of face recognition algorithms have led to their successful deployment across multiple applications ranging from border access to mobile unlocking and digital payments. However, their vulnerability against sophisticated and cost-effective presentation attack mediums raises essential questions regarding its reliability. In the literature, several presentation attack detection algorithms are presented; however, they are still far behind from reality. The major problem with existing work is the generalizability against multiple attacks both in the seen and unseen setting. The algorithms which are useful for one kind of attack (such as print) perform unsatisfactorily for another type of attack (such as silicone masks). In this research, we have proposed a deep learning-based network termed as MixNet to detect presentation attacks in cross-database and unseen attack settings. The proposed algorithm utilizes state-of-the-art convolutional neural network architectures and learns the feature mapping for each attack category. Experiments are performed using multiple challenging face presentation attack databases such as SMAD and Spoof In the Wild (SiW-M) databases. Extensive experiments and comparison with existing state of the art algorithms show the effectiveness of the proposed algorithm. Nilay Sanghvi, Sushant Kumar Singh, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001 |
ICPR | 5 |
| 2020 | Age Gap Reducer-GAN for Recognizing Age-Separated FacesabstractIn this paper, we propose a novel algorithm for matching faces with temporal variations caused due to age progression. The proposed generative adversarial network algorithm is a unified framework that combines facial age estimation and age-separated face verification. The key idea of this approach is to learn the age variations across time by conditioning the input image on the subject's gender and the target age group to which the face needs to be progressed. The loss function accounts for reducing the age gap between the original image and generated face image as well as preserving the identity. Both visual fidelity and quantitative evaluations demonstrate the efficacy of the proposed architecture on different facial age databases for age-separated face recognition. Daksha Yadav, Naman Kohli, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
ICPR | 4 |
| 2020 | Detecting Face2Face Facial Reenactment in VideosabstractVisual content has become the primary source of information, as evident in the billions of images and videos, shared and uploaded on the Internet every single day. This has led to an increase in alterations in images and videos to make them more informative and eye-catching for the viewers worldwide. Some of these alterations are simple, like copy-move, and are easily detectable, while other sophisticated alterations like reenactment based DeepFakes are hard to detect. Reenactment alterations allow the source to change the target expressions and create photo-realistic images and videos. While the technology can be potentially used for several applications, the malicious usage of automatic reenactment has a very large social implication. It is therefore important to develop detection techniques to distinguish real images and videos with the altered ones. This research proposes a learning-based algorithm for detecting reenactment based alterations. The proposed algorithm uses a multi-stream network that learns regional artifacts and provides a robust performance at various compression levels. We also propose a loss function for the balanced learning of the streams for the proposed network. The performance is evaluated on the publicly available FaceForen- sics dataset. The results show state-of-the-art classification accuracy of 99.96%, 99.10%, and 91.20% for no, easy, and hard compression factors, respectively. Prabhat Kumar 0005, Mayank Vatsa, Richa Singh 0001 |
WACV | 3 |
| 2019 | Data Fine-TuningabstractIn real-world applications, commercial off-the-shelf systems are utilized for performing automated facial analysis including face recognition, emotion recognition, and attribute prediction. However, a majority of these commercial systems act as black boxes due to the inaccessibility of the model parameters which makes it challenging to fine-tune the models for specific applications. Stimulated by the advances in adversarial perturbations, this research proposes the concept of Data Fine-tuning to improve the classification accuracy of a given model without changing the parameters of the model. This is accomplished by modeling it as data (image) perturbation problem. A small amount of “noise” is added to the input with the objective of minimizing the classification loss without affecting the (visual) appearance. Experiments performed on three publicly available datasets LFW, CelebA, and MUCT, demonstrate the effectiveness of the proposed concept. Saheb Chhabra, Puspita Majumdar, Mayank Vatsa, Richa Singh 0001 |
AAAI | 4 |
| 2019 | Guided DropoutabstractDropout is often used in deep neural networks to prevent over-fitting. Conventionally, dropout training invokes random drop of nodes from the hidden layers of a Neural Network. It is our hypothesis that a guided selection of nodes for intelligent dropout can lead to better generalization as compared to the traditional dropout. In this research, we propose “guided dropout” for training deep neural network which drop nodes by measuring the strength of each node. We also demonstrate that conventional dropout is a specific case of the proposed guided dropout. Experimental evaluation on multiple datasets including MNIST, CIFAR10, CIFAR100, SVHN, and Tiny ImageNet demonstrate the efficacy of the proposed guided dropout. Rohit Keshari, Richa Singh 0001, Mayank Vatsa |
AAAI | 2 |
| 2019 | On Learning Density Aware EmbeddingsabstractDeep metric learning algorithms have been utilized to learn discriminative and generalizable models which are effective for classifying unseen classes. In this paper, a novel noise tolerant deep metric learning algorithm is proposed. The proposed method, termed as Density Aware Metric Learning, enforces the model to learn embeddings that are pulled towards the most dense region of the clusters for each class. It is achieved by iteratively shifting the estimate of the center towards the dense region of the cluster thereby leading to faster convergence and higher generalizability. In addition to this, the approach is robust to noisy samples in the training data, often present as outliers. Detailed experiments and analysis on two challenging cross-modal face recognition databases and two popular object recognition databases exhibit the efficacy of the proposed approach. It has superior convergence, requires lesser training time, and yields better accuracies than several popular deep metric learning methods. Soumyadeep Ghosh, Richa Singh 0001, Mayank Vatsa |
CVPR | 2 |
| 2019 | FaceSurv: A Benchmark Video Dataset for Face Detection and Recognition Across Spectra and ResolutionsabstractExisting face recognition algorithms achieve high recognition performance for frontal face images with good illumination and close proximity to the imaging device. However, most of the existing algorithms fail to perform equally well in surveillance scenarios, where videos are captured across varying resolutions and spectra. In surveillance settings, cameras are usually placed far away from the subjects, thereby resulting in variations across pose, illumination, occlusion, and resolution. Current video datasets used for face recognition are often captured in constrained environments, and thus fail to simulate the real world scenarios. In this paper, we present the FaceSurv database featuring 252 subjects in 460 videos. The proposed dataset contains over 142K face images, spread across videos captured in both visible and near-infrared spectra. Each video contains a group of individuals walking from 36ft towards the imaging device, offering a plethora of challenges common to surveillance settings. Benchmark experimental protocol and baseline results have been reported with state-of-the-art algorithms for face detection and recognition. It is our assertion that the availability of such a challenging database will facilitate the development of robust face recognition systems relevant to real world surveillance scenarios. Sanchit Gupta, Nikita Gupta, Soumyadeep Ghosh, Maneet Singh, Shruti Nagpal, Mayank Vatsa, Richa Singh 0001 |
FG | 7 |
| 2019 | DroneSURF: Benchmark Dataset for Drone-based Face RecognitionabstractUnmanned Aerial Vehicles (UAVs) or drones are often used to reach remote areas or regions which are inaccessible to humans. Equipped with a large field of view, compact size, and remote control abilities, drones are deemed suitable for monitoring crowded or disaster-hit areas, and performing aerial surveillance. While research has focused on area monitoring, object detection and tracking, limited attention has been given to person identification, especially face recognition, using drones. This research presents a novel large-scale drone dataset, DroneSURF: Drone Surveillance of Faces, in order to facilitate research for face recognition. The dataset contains 200 videos of 58 subjects, captured across 411K frames, having over 786K face annotations. The proposed dataset demonstrates variations across two surveillance use cases: (i) active and (ii) passive, two locations, and two acquisition times. DroneSURF encapsulates challenges due to the effect of motion, variations in pose, illumination, background, altitude, and resolution, especially due to the large and varying distance between the drone and the subjects. This research presents a detailed description of the proposed DroneSURF dataset, along with information regarding the data distribution, protocols for evaluation, and baseline results. Isha Kalra, Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa, P. B. Sujit |
FG | 4 |
| 2019 | Dual Directed Capsule Network for Very Low Resolution Image RecognitionabstractVery low resolution (VLR) image recognition corresponds to classifying images with resolution 16×16 or less. Though it has widespread applicability when objects are captured at a very large stand-off distance (e.g. surveillance scenario) or from wide angle mobile cameras, it has received limited attention. This research presents a novel Dual Directed Capsule Network model, termed as DirectCapsNet, for addressing VLR digit and face recognition. The proposed architecture utilizes a combination of capsule and convolutional layers for learning an effective VLR recognition model. The architecture also incorporates two novel loss functions: (i) the proposed HR-anchor loss and (ii) the proposed targeted reconstruction loss, in order to overcome the challenges of limited information content in VLR images. The proposed losses use high resolution images as auxiliary data during training to "direct" discriminative feature learning. Multiple experiments for VLR digit classification and VLR face recognition are performed along with comparisons with state-of-the-art algorithms. The proposed DirectCapsNet consistently showcases state-of-the-art results; for example, on the UCCS face database, it shows over 95% face recognition accuracy when 16×16 images are matched with 80×80 images. Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa |
ICCV | 3 |
| 2019 | Triplet Transform Learning for Automated Primate Face RecognitionabstractAutomated primate face recognition has enormous potential in effective conservation of species facing endangerment or extinction. The task is characterized by lack of training data, low inter-class variations, and large intra-class differences. Owing to the challenging nature of the problem, limited research has been performed to automate the process of primate face recognition. In this research, we propose a novel Triplet Transform Learning (TTL) model for learning discriminative representations of primate faces. The proposed model reduces the intra-class variations and increases the inter-class variations to obtain robust sparse representations for the primate faces. It is utilized to present a novel framework for primate face recognition, which is evaluated on the primate dataset, comprising of 80 identities including monkeys, gorillas, and chimpanzees. Experimental results demonstrate the efficacy of the proposed approach, where it outperforms the existing approaches and attains state-of-the-art performance on the primates database. Mohit Agarwal 0007, Sanchit Sinha, Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa |
ICIP | 5 |
| 2019 | AUTO-G: Gesture Recognition in the Crowd for Autonomous VehiclabstractAutonomous driving is an active area of research. An important aspect of this problem is recognizing the gestures made by humans, both inside and outside the vehicle. In this paper, we present the Auto-G database that comprises different hand gestures for autonomous driving. The database encompasses several challenges such as occlusion, low resolution, motion blur, illumination variation, extreme pose variations, along with the presence of multiple gestures within a frame. We also propose an end-to-end pipeline for hand detection and gesture recognition. The proposed pipeline achieves a frame gesture recognition accuracy of 90.23% on the proposed Auto-G database. Pavani Tripathi, Rohit Keshari, Soumyadeep Ghosh, Mayank Vatsa, Richa Singh 0001 |
ICIP | 5 |
| 2019 | Siamese Deep Dictionary LearningabstractResearchers have explored the importance of Siamese networks in deep learning. With recent developments in deep learning and the effectiveness of deep dictionary learning, this research proposes the architecture of Siamese Deep Dictionary Learning. We first propose the architecture followed by solving the optimization problem. The experimental effectiveness is demonstrated on five different image databases pertaining to two classification problems: face verification and kinship verification. The experiments show that the proposed Siamese Deep Dictionary Learning yields comparable results compared to state-of-the-art algorithms on all five databases. Vanika Singhal, Angshul Majumdar, Mayank Vatsa, Richa Singh 0001 |
IJCNN | 4 |
| 2019 | Latent Fingerprint Enhancement Using Generative Adversarial NetworksabstractLatent fingerprints recognition is very useful in law enforcement and forensics applications. However, automated matching of latent fingerprints with a gallery of live scan images is very challenging due to several compounding factors such as noisy background, poor ridge structure, and overlapping unstructured noise. In order to efficiently match latent fingerprints, an effective enhancement module is a necessity so that it can facilitate correct minutiae extraction. In this research, we propose a Generative Adversarial Network based latent fingerprint enhancement algorithm to enhance the poor quality ridges and predict the ridge information. Experiments on two publicly available datasets, IIITD-MOLF and IIITD-MSLFD show that the proposed enhancement algorithm improves the fingerprints quality while preserving the ridge structure. It helps the standard feature extraction and matching algorithms to boost latent fingerprints matching performance. Indu Joshi, Adithya Anand, Mayank Vatsa, Richa Singh 0001, Sumantra Dutta Roy, Prem Kumar Kalra |
WACV | 4 |
| 2019 | Detecting and Mitigating Adversarial Perturbations for Robust Face Recognition
Gaurav Goswami, Akshay Agarwal 0001, Nalini K. Ratha, Richa Singh 0001, Mayank Vatsa |
Int. J. Comput. Vis. | 4 |
| 2019 | Between-subclass piece-wise linear solutions in large scale kernel SVM learning
Tejas I. Dhamecha, Afzel Noore, Richa Singh 0001, Mayank Vatsa |
Pattern Recognit. | 3 |
| 2019 | Residual Codean Autoencoder for Facial Attribute Analysis
Akshay Sethi, Maneet Singh, Richa Singh 0001, Mayank Vatsa |
Pattern Recognit. Lett. | 3 |
| 2019 | Are you eligible? Predicting adulthood from face images via Class Specific Mean Autoencoder
Maneet Singh, Shruti Nagpal, Mayank Vatsa, Richa Singh 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Supervised Mixed Norm Autoencoder for Kinship Verification in Unconstrained VideosabstractIdentifying kinship relations has garnered interest due to several applications such as organizing and tagging the enormous amount of videos being uploaded on the Internet. Existing research in kinship verification primarily focuses on kinship prediction with image pairs. In this research, we propose a new deep learning framework for kinship verification in unconstrained videos using a novel Supervised Mixed Norm regularization Autoencoder (SMNAE). This new autoencoder formulation introduces class-specific sparsity in the weight matrix. The proposed three-stage SMNAE based kinship verification framework utilizes the learned spatio-temporal representation in the video frames for verifying kinship in a pair of videos. A new kinship video (KIVI) database of more than 500 individuals with variations due to illumination, pose, occlusion, ethnicity, and expression is collected for this research. It comprises a total of 355 true kin video pairs with over 250,000 still frames. The effectiveness of the proposed framework is demonstrated on the KIVI database and six existing kinship databases. On the KIVI database, SMNAE yields video-based kinship verification accuracy of 83.18% which is at least 3.2% better than existing algorithms. The algorithm is also evaluated on six publicly available kinship databases and compared with best reported results. It is observed that the proposed SMNAE consistently yields best results on all the databases. Naman Kohli, Daksha Yadav, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IEEE Trans. Image Process. | 4 |
| 2018 | Unravelling Robustness of Deep Learning Based Face Recognition Against Adversarial AttacksabstractDeep neural network (DNN) architecture based models have high expressive power and learning capacity. However, they are essentially a black box method since it is not easy to mathematically formulate the functions that are learned within its many layers of representation. Realizing this, many researchers have started to design methods to exploit the drawbacks of deep learning based algorithms questioning their robustness and exposing their singularities. In this paper, we attempt to unravel three aspects related to the robustness of DNNs for face recognition: (i) assessing the impact of deep architectures for face recognition in terms of vulnerabilities to attacks inspired by commonly observed distortions in the real world that are well handled by shallow learning methods along with learning based adversaries; (ii) detecting the singularities by characterizing abnormal filter response behavior in the hidden layers of deep networks; and (iii) making corrections to the processing pipeline to alleviate the problem. Our experimental evaluation using multiple open-source DNN-based face recognition networks, including OpenFace and VGG-Face, and two publicly available databases (MEDS and PaSC) demonstrates that the performance of deep learning based face recognition algorithms can suffer greatly in the presence of such distortions. The proposed method is also compared with existing detection algorithms and the results show that it is able to detect the attacks with very high accuracy by suitably designing a classifier using the response of the hidden layers in the network. Finally, we present several effective countermeasures to mitigate the impact of adversarial attacks and improve the overall robustness of DNN-based face recognition. Gaurav Goswami, Nalini K. Ratha, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa |
AAAI | 4 |
| 2018 | Learning Structure and Strength of CNN Filters for Small Sample Size TrainingabstractConvolutional Neural Networks have provided state-of-the-art results in several computer vision problems. However, due to a large number of parameters in CNNs, they require a large number of training samples which is a limiting factor for small sample size problems. To address this limitation, we propose SSF-CNN which focuses on learning the "structure" and "strength" of filters. The structure of the filter is initialized using a dictionary based filter learning algorithm and the strength of the filter is learned using the small sample training data. The architecture provides the flexibility of training with both small and large training databases, and yields good accuracies even with small size training data. The effectiveness of the algorithm is first demonstrated on MNIST, CIFAR10, and NORB databases, with varying number of training samples. The results show that SSF-CNN significantly reduces the number of parameters required for training while providing high accuracies on the test databases. On small sample size problems such as newborn face recognition and Omniglot, it yields state-of-the-art results. Specifically, on the IIITD Newborn Face Database, the results demonstrate improvement in rank-1 identification accuracy by at least 10%. Rohit Keshari, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
CVPR | 3 |
| 2018 | Scattering Transform for Matching Surgically Altered Face ImagesabstractThe use of face as a biometric feature has been widely accepted and used in security and surveillance systems. Recent studies have made significant advancements to address various challenges in face recognition such as illumination, age, pose and disguise. Another important covariate is recognizing faces with pre-and-post facial plastic surgery. Facial plastic surgeries alter the geometry and texture of facial regions, the extent of which is dependent on both the number, and the type of surgeries performed. The increasing reach of plastic surgery and its expanding user base present an indispensable challenge that must be dealt with while devising robust face recognition systems. In this paper, we present Invariant Scattering transform based feature extraction to compute translation invariant representation at local and global levels that is stable against plastic surgery variations. The identification accuracy achieved by the proposed algorithm is over 97% at rank-10 on the IIITD plastic surgery face database. Ishita Gupta, Ikshu Bhalla, Richa Singh 0001, Mayank Vatsa |
ICPR | 3 |
| 2018 | SegDenseNet: Iris Segmentation for Pre-and-Post Cataract SurgeryabstractCataract is one of the major ophthalmic diseases worldwide which can potentially affect the performance of iris-based biometric systems. While existing research has shown that cataract does not have a major impact on iris recognition, our observations suggest that iris segmentation algorithms are not well equipped to handle cataract or post cataract surgery cases, thereby affecting the overall iris recognition performance. This paper presents an efficient iris segmentation algorithm with variations due to cataract and post cataract surgery. The proposed algorithm, termed as SegDenseNet, is a deep learning algorithm based on DenseNet. The experiments on the IIITD Cataract Surgery Database show that improving iris segmentation enhances the recognition performance by up to 25% across different sensors and matchers. Aditya Lakra, Pavani Tripathi, Rohit Keshari, Mayank Vatsa, Richa Singh 0001 |
ICPR | 5 |
| 2018 | Face Recognition for Newborns, Toddlers, and Pre-School Children: A Deep Learning ApproachabstractBiometric recognition of newborns, toddlers, and pre-school children aims is an important research challenge with applications in identifying newborn swapping, missing kids, and disbursing benefits. In this research, we propose a representation learning algorithm to extract unique and invariant features from face images of newborns and toddlers, to design an efficient face recognition algorithm. Specifically, we propose a deep learning model which applies class-based penalties while learning the filters of a convolutional neural network. The proposed CNN architecture achieves a rank-1 identification accuracy of 62.7% for single gallery newborn face recognition and 85.1% for single gallery toddler face recognition, forming state-of-the-results for both the databases. Comparison with several existing algorithms also showcases the effectiveness of the proposed algorithm on both the databases. Sahar Siddiqui, Mayank Vatsa, Richa Singh 0001 |
ICPR | 3 |
| 2018 | Anonymizing k Facial Attributes via Adversarial PerturbationsabstractA face image not only provides details about the identity of a subject but also reveals several attributes such as gender, race, sexual orientation, and age. Advancements in machine learning algorithms and popularity of sharing images on the World Wide Web, including social media websites, have increased the scope of data analytics and information profiling from photo collections. This poses a serious privacy threat for individuals who do not want to be profiled. This research presents a novel algorithm for anonymizing selective attributes which an individual does not want to share without affecting the visual quality of images. Using the proposed algorithm, a user can select single or multiple attributes to be surpassed while preserving identity information and visual content. The proposed adversarial perturbation based algorithm embeds imperceptible noise in an image such that attribute prediction algorithm for the selected attribute yields incorrect classification result, thereby preserving the information according to user's choice. Experiments on three popular databases i.e. MUCT, LFWcrop, and CelebA show that the proposed algorithm not only anonymizes \textit{k}-attributes, but also preserves image quality and identity information. Saheb Chhabra, Richa Singh 0001, Mayank Vatsa |
IJCAI | 2 |
| 2018 | Person Authentication Using Head ImagesabstractIn many surveillance applications, the cameras are placed at overhead heights for human identification. In such real-world scenarios, the person of interest might be walking away from the camera and the only information available is "image of the person's head". In this research, we investigate the usage of head images for person recognition and propose it as a soft-biometric modality. With its viability for human recognition, application of head images can also be extended with other face recognition algorithms for surveillance. We propose a head image database pertaining to 103 subjects with more than 600 images. In addition to the database, we propose a framework for head image-based person verification. As a pre-processing stage, the framework includes evaluation of two segmentation algorithms. We also perform benchmarking evaluations of various texture, key-point, and learning-based representation algorithms and establish the baseline results. The experiments suggest that head images can be effectively used to ascertain human identity and the availability of this database could pave further research in this field. Aakarsh Malhotra, Richa Singh 0001, Mayank Vatsa, Vishal M. Patel |
WACV | 2 |
| 2018 | Iris Presentation Attack via Textured Contact Lens in Unconstrained EnvironmentabstractThe widespread use of smartphones has spurred the research in mobile iris devices. Due to their convenience, these mobile devices are also utilized in unconstrained outdoor scenarios. This has necessitated the development of reliable iris recognition algorithms for such uncontrolled environment. At the same time, iris presentation attacks pose a major challenge to current iris recognition systems. It has been shown that print attacks and textured contact lens may significantly degrade the iris recognition performance. Motivated by these factors, we present a novel Mobile Uncontrolled Iris Presentation Attack Database (MUIPAD). The database contains more than 10,000 iris images that are acquired with and without textured contact lenses in indoor and outdoor environments using a mobile sensor. We also investigate the efficacy of textured contact lens in identity impersonation and obfuscation. Moreover, we demonstrate the effectiveness of deep learning based features for iris presentation attack detection on the proposed database. Daksha Yadav, Naman Kohli, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
WACV | 4 |
| 2018 | CrowdFaceDB: Database and benchmarking for face verification in crowd
Tejas I. Dhamecha, Mahek Shah, Mayank Vatsa, Richa Singh 0001 |
Pattern Recognit. Lett. | 5 |
| 2017 | SWAPPED! Digital face presentation attack detection via weighted local magnitude patternabstractAdvancements in smartphone applications have empowered even non-technical users to perform sophisticated operations such as morphing in faces as few tap operations. While such enablements have positive effects, as a negative side, now anyone can digitally attack face (biometric) recognition systems. For example, face swapping application of Snapchat can easily create “swapped” identities and circumvent face recognition system. This research presents a novel database, termed as SWAPPED - Digital Attack Video Face Database, prepared using Snap chat's application which swaps/stitches two faces and creates videos. The database contains bonafide face videos and face swapped videos of multiple subjects. Baseline face recognition experiments using commercial system shows over 90% rank-1 accuracy when attack videos are used as probe. As a second contribution, this research also presents a novel Weighted Local Magnitude Pattern feature descriptor based presentation attack detection algorithm which outperforms several existing approaches. Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
IJCB | 2 |
| 2017 | Multimodal biometric recognition for toddlers and pre-school childrenabstractIn many applications such as law enforcement, attendance systems, and medical services, biometrics is utilized for identifying individuals. However, current systems, in general, do not enroll all possible age groups, particularly, toddlers and pre-school children. This research is the first of its kind attempt to prepare a multimodal biometric database for such potential users of biometric systems. In the proposed database, face, fingerprint, and iris modalities of over 100 children (age range of 18 months to 4 years) are captured in two different sessions, months apart. We also perform benchmarking evaluation of existing tools and algorithms to establish the baseline results for different unimodal and multimodal scenarios. Our experience and results suggest that while iris is highly accurate, it requires constant adult supervision to attain cooperation from children. On the other hand, face is the most easy-to-capture modality but yields very low verification performance. We assert that the availability of this database can instigate research in this important research problem. Pratichi Basak, Saurabh De, Mallika Agarwal, Aakarsh Malhotra, Mayank Vatsa, Richa Singh 0001 |
IJCB | 6 |
| 2017 | Demography-based facial retouching detection using subclass supervised sparse autoencoderabstractDigital retouching of face images is becoming more widespread due to the introduction of software packages that automate the task. Several researchers have introduced algorithms to detect whether a face image is original or retouched. However, previous work on this topic has not considered whether or how accuracy of retouching detection varies with the demography of face images. In this paper, we introduce a new Multi-Demographic Retouched Faces (MDRF) dataset, which contains images belonging to two genders, male and female, and three ethnicities, Indian, Chinese, and Caucasian. Further, retouched images are created using two different retouching software packages. The second major contribution of this research is a novel semi-supervised autoencoder incorporating “sub-class” information to improve classification. The proposed approach outperforms existing state-of-the-art detection algorithms for the task of generalized retouching detection. Experiments conducted with multiple combinations of ethnicities show that accuracy of retouching detection can vary greatly based on the demographics of the training and testing images. Aparna Bharati, Mayank Vatsa, Richa Singh 0001, Kevin W. Bowyer |
IJCB | 3 |
| 2017 | Synthetic iris presentation attack using iDCGANabstractReliability and accuracy of iris biometric modality has prompted its large-scale deployment for critical applications such as border control and national ID projects. The extensive growth of iris recognition systems has raised apprehensions about susceptibility of these systems to various attacks. In the past, researchers have examined the impact of various iris presentation attacks such as textured contact lenses and print attacks. In this research, we present a novel presentation attack using deep learning based synthetic iris generation. Utilizing the generative capability of deep con-volutional generative adversarial networks and iris quality metrics, we propose a new framework, named as iDCGAN (iris deep convolutional generative adversarial network) for generating realistic appearing synthetic iris images. We demonstrate the effect of these synthetically generated iris images as presentation attack on iris recognition by using a commercial system. The state-of-the-art presentation attack detection framework, DESIST is utilized to analyze if it can discriminate these synthetically generated iris images from real images. The experimental results illustrate that mitigating the proposed synthetic presentation attack is of paramount importance. Naman Kohli, Daksha Yadav, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IJCB | 4 |
| 2017 | On matching skulls to digital face images: A preliminary approachabstractForensic application of automatically matching skull with face images is an important research area linking biometrics with practical applications inforensics. It is an opportunity for biometrics and face recognition researchers to help the law enforcement and forensic experts in giving an identity to unidentified human skulls. It is an extremely challenging problem which is further exacerbated due to lack of any publicly available database related to this problem. This is the first research in this direction with a twofold contribution: (i) introducing the first of its kind skull-face image pair database, Identify Me, and (ii) presenting a preliminary approach using the proposed semi-supervised formulation of transform learning. The experimental results and comparison with existing algorithms showcase the challenging nature of the problem. We assert that the availability of the database will inspire researchers to build sophisticated skull-to-face matching algorithms. Shruti Nagpal, Maneet Singh, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
IJCB | 4 |
| 2017 | Gender and ethnicity classification of Iris images using deep class-encoderabstractSoft biometric modalities have shown their utility in different applications including reducing the search space significantly. This leads to improved recognition performance, reduced computation time, and faster processing of test samples. Some common soft biometric modalities are ethnicity, gender, age, hair color, iris color, presence of facial hair or moles, and markers. This research focuses on performing ethnicity and gender classification on iris images. We present a novel supervised auto-encoder based approach, Deep Class-Encoder, which uses class labels to learn discriminative representation for the given sample by mapping the learned feature vector to its label. The proposed model is evaluated on two datasets each for ethnicity and gender classification. The results obtained using the proposed Deep Class-Encoder demonstrate its effectiveness in comparison to existing approaches and state-of-the-art methods. Maneet Singh, Shruti Nagpal, Mayank Vatsa, Richa Singh 0001, Afzel Noore, Angshul Majumdar |
IJCB | 4 |
| 2017 | Unconstrained visible spectrum iris with textured contact lens variations: Database and benchmarkingabstractIris recognition in visible spectrum has developed into an active area of research. This has elevated the importance of efficient presentation attack detection algorithms, particularly in security based critical applications. In this paper, we present the first detailed analysis of the effect of textured contact lenses on iris recognition in visible spectrum. We introduce the first contact lens database in visible spectrum, Unconstrained Visible Contact Lens Iris (UVCLI) Database, containing samples from 70 classes with subjects wearing textured contact lenses in indoor and outdoor environments across multiple sessions. We observe that textured contact lenses degrade the visible spectrum iris recognition performance by over 25% and thus, may be utilized intentionally or unintentionally to attack existing iris recognition systems. Next, three iris presentation attack detection (PAD) algorithms are evaluated on the proposed database and highest PAD accuracy of 82.85%c is observed. This illustrates that there is a significant scope of improvement in developing efficient PAD algorithms for detection of textured contact lenses in unconstrained visible spectrum iris images. Daksha Yadav, Naman Kohli, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IJCB | 4 |
| 2017 | LivDet iris 2017 - Iris liveness detection competition 2017abstractPresentation attacks such as using a contact lens with a printed pattern or printouts of an iris can be utilized to bypass a biometric security system. The first international iris liveness competition was launched in 2013 in order to assess the performance of presentation attack detection (PAD) algorithms, with a second competition in 2015. This paper presents results of the third competition, LivDet-Iris 2017. Three software-based approaches to Presentation Attack Detection were submitted. Four datasets of live and spoof images were tested with an additional cross-sensor test. New datasets and novel situations of data have resulted in this competition being of a higher difficulty than previous competitions. Anonymous received the best results with a rate of rejected live samples of 3.36% and rate of accepted spoof samples of 14.71%. The results show that even with advances, printed iris attacks as well as patterned contacts lenses are still difficult for software-based systems to detect. Printed iris images were easier to be differentiated from live images in comparison to patterned contact lenses as was also seen in previous competitions. David Yambay, Benedict Becker, Naman Kohli, Daksha Yadav, Adam Czajka, Kevin W. Bowyer, Stephanie Schuckers, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Diego Gragnaniello, Carlo Sansone, Luisa Verdoliva, Lingxiao He, Yiwei Ru, Nianfeng Liu, Zhenan Sun, Tieniu Tan |
IJCB | 8 |
| 2017 | Face Sketch Matching via Coupled Deep Transform LearningabstractFace sketch to digital image matching is an important challenge of face recognition that involves matching across different domains. Current research efforts have primarily focused on extracting domain invariant representations or learning a mapping from one domain to the other. In this research, we propose a novel transform learning based approach termed as DeepTransformer, which learns a transformation and mapping function between the features of two domains. The proposed formulation is independent of the input information and can be applied with any existing learned or hand-crafted feature. Since the mapping function is directional in nature, we propose two variants of DeepTransformer: (i) semi-coupled and (ii) symmetrically-coupled deep transform learning. This research also uses a novel IIIT-D Composite Sketch with Age (CSA) variations database which contains sketch images of 150 subjects along with age-separated digital photos. The performance of the proposed models is evaluated on a novel application of sketch-to-sketch matching, along with sketch-to-digital photo matching. Experimental results demonstrate the robustness of the proposed models in comparison to existing state-of-the-art sketch matching algorithms and a commercial face recognition system. Shruti Nagpal, Maneet Singh, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Angshul Majumdar |
ICCV | 3 |
| 2017 | Kernel group sparse representation based classifier for multimodal biometricsabstractClassification is an important pattern recognition paradigm with a multitude of applications in popular research problems. Utilizing multiple data representations to improve the accuracy of classification has been explored in literature. However, approaches such as combining classifiers using majority voting and score level fusion do not utilize the underlying structure of the data which is available at the representation stage itself. In this paper, we propose a kernelization based extension to the group sparse representation classifier which can utilize multiple representations of input data to improve classification performance. By using a kernel, these representations are processed in a higher dimensional space where they are more separable, without substantially increasing computational costs. The proposed algorithm selects the ideal kernel to use along with its parameters automatically as part of the training process. We evaluate the proposed algorithm on three challenging biometric problems namely, cross distance face recognition, RGB-D face recognition, and multimodal biometrics to showcase its efficacy. Experimentally, we observe that the proposed algorithm can efficiently combine multiple data representations to further improve classification performance. Gaurav Goswami, Richa Singh 0001, Mayank Vatsa, Angshul Majumdar |
IJCNN | 2 |
| 2017 | Class representative autoencoder for low resolution multi-spectral gender classificationabstractGender is one of the most common attributes used to describe an individual. It is used in multiple domains such as human computer interaction, marketing, security, and demographic reports. Research has been performed to automate the task of gender recognition in constrained environment using face images, however, limited attention has been given to gender classification in unconstrained scenarios. This work attempts to address the challenging problem of gender classification in multi-spectral low resolution face images. We propose a robust Class Representative Autoencoder model, termed as AutoGen for the same. The proposed model aims to minimize the intra-class variations while maximizing the inter-class variations for the learned feature representations. Results on visible as well as near infrared spectrum data for different resolutions and multiple databases depict the efficacy of the proposed model. Comparative results with existing approaches and two commercial off-the-shelf systems further motivate the use of class representative features for classification. Maneet Singh, Shruti Nagpal, Richa Singh 0001, Mayank Vatsa |
IJCNN | 3 |
| 2017 | Region-specific fMRI dictionary for decoding face verification in humansabstractThis paper focuses on decoding the process of face verification in the human brain using fMRI responses. 2400 fMRI responses are collected from different participants while they perform face verification on genuine and imposter stimuli face pairs. The first part of the paper analyzes the responses covering both cognitive and fMRI neuro-imaging results. With an average verification accuracy of 64.79% by human participants, the results of the cognitive analysis depict that the performance of female participants is significantly higher than the male participants with respect to imposter pairs. The results of the neuro-imaging analysis identifies regions of the brain such as the left fusiform gyrus, caudate nucleus, and superior frontal gyrus that are activated when participants perform face verification tasks. The second part of the paper proposes a novel two-level fMRI dictionary learning approach to predict if the stimuli observed is genuine or imposter using the brain activation data for selected regions. A comparative analysis with existing machine learning techniques illustrates that the proposed approach yields at least 4.5% higher classification accuracy than other algorithms. It is envisioned that the result of this study is the first step in designing brain-inspired automatic face verification algorithms. Daksha Yadav, Naman Kohli, Shruti Nagpal, Maneet Singh, Prateekshit Pandey, Mayank Vatsa, Richa Singh 0001, Afzel Noore, Gokulraj Prabhakaran, Harsh Mahajan |
IJCNN | 7 |
| 2017 | Group sparse autoencoder
Anush Sankaran, Mayank Vatsa, Richa Singh 0001, Angshul Majumdar |
Image Vis. Comput. | 3 |
| 2017 | Face Verification via Class Sparsity Based Supervised EncodingabstractAutoencoders are deep learning architectures that learn feature representation by minimizing the reconstruction error. Using an autoencoder as baseline, this paper presents a novel formulation for a class sparsity based supervised encoder, termed as CSSE. We postulate that features from the same class will have a common sparsity pattern/support in the latent space. Therefore, in the formulation of the autoencoder, a supervision penalty is introduced as a jointsparsity promoting l2;1-norm. The formulation of CSSE is derived for a single hidden layer and it is applied for multiple hidden layers using a greedy layer-bylayer learning approach. The proposed CSSE approach is applied for learning face representation and verification experiments are performed on the LFW and PaSC face databases. The experiments show that the proposed approach yields improved results compared to autoencoders and comparable results with state-ofthe-art face recognition algorithms. Angshul Majumdar, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Class sparsity signature based Restricted Boltzmann Machine
Anush Sankaran, Gaurav Goswami, Mayank Vatsa, Richa Singh 0001, Angshul Majumdar |
Pattern Recognit. | 4 |
| 2017 | Face Verification via Learned Representation on Feature-Rich Video FramesabstractAbundance and availability of video capture devices, such as mobile phones and surveillance cameras, have instigated research in video face recognition, which is highly pertinent in law enforcement applications. While the current approaches have reported high accuracies at equal error rates, performance at lower false accept rates requires significant improvement. In this paper, we propose a novel face verification algorithm, which starts with selecting feature-rich frames from a video sequence using discrete wavelet transform and entropy computation. Frame selection is followed by representation learning-based feature extraction, where three contributions are presented: 1) deep learning architecture, which is a combination of stacked denoising sparse autoencoder (SDAE) and deep Boltzmann machine (DBM); 2) formulation for joint representation in an autoencoder; and 3) updating the loss function of DBM by including sparse and low rank regularization. Finally, a multilayer neural network is used as the classifier to obtain the verification decision. The results are demonstrated on two publicly available databases, YouTube Faces and Point and Shoot Challenge. Experimental analysis suggests that: 1) the proposed feature-richness-based frame selection offers noticeable and consistent performance improvement compared with frontal only frames, random frames, or frame selection using perceptual no-reference image quality measures and 2) joint feature learning in SDAE and sparse and low rank regularization in DBM helps in improving face verification performance. On the benchmark Point and Shoot Challenge database, the algorithm yields the verification accuracy of over 97% at 1% false accept rate whereas, on the YouTube Faces database, over 95% verification accuracy is observed at equal error rate. Gaurav Goswami, Mayank Vatsa, Richa Singh 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Detecting Silicone Mask-Based Presentation Attack via Deep Dictionary LearningabstractIn movies, film stars portray another identity or obfuscate their identity with the help of silicone/latex masks. Such realistic masks are now easily available and are used for entertainment purposes. However, their usage in criminal activities to deceive law enforcement and automatic face recognition systems is also plausible. Therefore, it is important to guard biometrics systems against such realistic presentation attacks. This paper introduces the first-of-its-kind silicone mask attack database which contains 130 real and attacked videos to facilitate research in developing presentation attack detection algorithms for this challenging scenario. Along with silicone mask, there are several other presentation attack instruments that are explored in literature. The next contribution of this research is a novel multilevel deep dictionary learning-based presentation attack detection algorithm that can discern different kinds of attacks. An efficient greedy layer by layer training approach is formulated to learn the deep dictionaries followed by SVM to classify an input sample as genuine or attacked. Experimental are performed on the proposed SMAD database, some samples with real world silicone mask attacks, and four existing presentation attack databases, namely, replay-attack, CASIA-FASD, 3DMAD, and UVAD. The results show that the proposed algorithm yields better performance compared with state-ofthe-art algorithms, in both intra-database and cross-database experiments. Ishan Manjani, Snigdha Tariyal, Mayank Vatsa, Richa Singh 0001, Angshul Majumdar |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2017 | Hierarchical Representation Learning for Kinship VerificationabstractKinship verification has a number of applications such as organizing large collections of images and recognizing resemblances among humans. In this paper, first, a human study is conducted to understand the capabilities of human mind and to identify the discriminatory areas of a face that facilitate kinship-cues. The visual stimuli presented to the participants determine their ability to recognize kin relationship using the whole face as well as specific facial regions. The effect of participant gender and age and kin-relation pair of the stimulus is analyzed using quantitative measures such as accuracy, discriminability index d' , and perceptual information entropy. Utilizing the information obtained from the human study, a hierarchical kinship verification via representation learning (KVRL) framework is utilized to learn the representation of different face regions in an unsupervised manner. We propose a novel approach for feature representation termed as filtered contractive deep belief networks (fcDBN). The proposed feature representation encodes relational information present in images using filters and contractive regularization penalty. A compact representation of facial images of kin is extracted as an output from the learned model and a multi-layer neural network is utilized to verify the kin accurately. A new WVU kinship database is created, which consists of multiple images per subject to facilitate kinship verification. The results show that the proposed deep learning framework (KVRL-fcDBN) yields the state-of-the-art kinship verification accuracy on the WVU kinship database and on four existing benchmark data sets. Furthermore, kinship information is used as a soft biometric modality to boost the performance of face verification via product of likelihood ratio and support vector machine based approaches. Using the proposed KVRL-fcDBN framework, an improvement of over 20% is observed in the performance of face verification. Naman Kohli, Mayank Vatsa, Richa Singh 0001, Afzel Noore, Angshul Majumdar |
IEEE Trans. Image Process. | 3 |
| 2016 | Face identification from low resolution near-infrared imagesabstractFace identification from low quality and low resolution Near-Infrared (NIR) face images is a challenging problem. Since surveillance cameras typically acquire images at a large standoff distance, the effective resolution of the face is not large enough to identify the individuals. Moreover for a 24-hour surveillance footage, images in low light and at nighttime are acquired in NIR mode which makes the identification problem even more challenging. We propose an effective method using both hand-crafted and learned features for face identification of low resolution NIR images. We show that learned features contribute considerably to the performance of identification algorithm, and that using both feature level and score level fusion in a hierarchal approach gives good performance. The results demonstrate the effectiveness of the proposed approach on images which are of low quality, low resolution and acquired under challenging illumination conditions in near-infrared mode by surveillance cameras. Soumyadeep Ghosh, Rohit Keshari, Richa Singh 0001, Mayank Vatsa |
ICIP | 3 |
| 2016 | Mobile periocular matching with pre-post cataract surgeryabstractOcular recognition algorithms, including iris matching, have been used in several applications including large scale national ID projects such as India's Aadaar. Deployment of large-scale biometric systems is expected to rely on using multiple devices including mobile devices to ensure widespread adoption of biometric recognition systems. Ocular images captured using mobile devices may have challenges such as uncontrolled illumination, complex background, and geometric distortions. Further, among many enrollees of large scale biometrics program, some may have ocular diseases. One of the most common ocular disease in elderly is cataract. While it is established that iris recognition may be challenging due to ocular diseases, this paper investigates periocular recognition with pre and post cataract surgery images. In this research, we present a mobile periocular database of 145 subjects1. Baseline results also include a framework that achieves over 69% rank-10 accuracy and around 24% genuine accept rate at 1% false accept rate in inter-session experiments. Rohit Keshari, Soumyadeep Ghosh, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa |
ICIP | 4 |
| 2016 | At-a-distance person recognition via combining ocular featuresabstractPerson recognition is a challenging research problem particularly if the images are captured at a distance and only ocular region is present. In this research, we present a framework that extracts multiple features from iris and periocular regions from near infrared images captured at a distance of 2 meters or more. Using these features and random decision forest, fusion and classification is performed and verification results are reported. On CASIA V4-at-a-distance and FOCS databases, the proposed algorithm yields state-of-the-art results; particularly achieving over 61% genuine accept rate at 0.1% false accept rate on complete CASIA V4-at-a-distance database. Shalini Verma, Paritosh Mittal, Mayank Vatsa, Richa Singh 0001 |
ICIP | 4 |
| 2016 | Low rank group sparse representation based classifier for pose variationabstractFace recognition under uncontrolled environment persists to be an unresolved problem having challenges such as varying pose, illumination, occlusion etc. In this research, we propose an algorithm for identification of faces with pose and illumination variations. An adaptive dictionary learning framework built upon group sparse representation classifier is presented in order to learn dictionary parameters and pose invariant sparse codes for given images. Low rank regularization is utilized for dictionary learning, to address the noise present in training samples that can hinder the discriminative power of the learnt dictionary. Experimental results illustrate state-of-the-art performance on the CMU Multi-PIE dataset. Shivangi Yadav, Maneet Singh, Mayank Vatsa, Richa Singh 0001, Angshul Majumdar |
ICIP | 4 |
| 2016 | Fingerprint sensor classification via Mélange of handcrafted featuresabstractLarge scale biometrics projects rely on capturing images/signal from multiple sensors. For example, in India's Aadhaar project, multiple fingerprint sensors of different make and model are used for data collection. Similarly, in law enforcement applications, different agencies use different fingerprint sensors. These scenarios cause two potential problems: (i) sensor inter-operability and (ii) protecting/recording chain of evidence. While sensor inter-operability in fingerprints is a well studied problem, automatically recording chain of evidence is a relatively less explored research problem. For both the problems, one potential approach includes automatically identifying sensors based on the input image. This paper presents a novel fingerprint sensor identification algorithm based on multiple features such as Haralick, entropy, statistical and image quality features. The proposed algorithm is evaluated on a large database with 30,000 images with 15 fingerprint sensor classes. The proposed algorithm achieves an accuracy of 96% and computationally requires less than 10 milliseconds for an image. Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa |
ICPR | 2 |
| 2016 | Improving classifier fusion via Pool Adjacent Violators normalizationabstractClassifier fusion is a well-studied problem in which decisions from multiple classifiers are combined at the score, rank, or decision level to obtain better results than a single classifier. Subsequently, various techniques for combining classifiers at each of these levels have been proposed in the literature. Many popular methods entail scaling and normalizing the scores obtained by each classifier to a common numerical range before combining the normalized scores using the sum rule or another classifier. In this research, we explore an alternative method to combine classifiers at the score level. The Pool Adjacent Violators (PAV) algorithm has traditionally been utilized to convert classifier match scores to confidence values that model posterior probabilities for test data. The PAV algorithm and other score normalization techniques have studied the same problem without being aware of each other. In this first ever study to combine the two, we propose the PAV algorithm for classifier fusion on publicly available NIST multi-modal biometrics score dataset. We observe that it provides several advantages over existing techniques and find that the interpretation learned by the PAV algorithm is more robust than the scaling learned by other popular normalization algorithms such as min-max. Moreover, the PAV algorithm enables the combined score to be interpreted as confidence and is able to further improve the results obtained by other approaches. We also observe that utilizing traditional normalization techniques first for individual classifiers and then normalizing the fused score using PAV offers a performance boost compared to only using the PAV algorithm. Gaurav Goswami, Nalini K. Ratha, Richa Singh 0001, Mayank Vatsa |
ICPR | 3 |
| 2016 | Face anti-spoofing with multifeature videolet aggregationabstractBiometric systems can be attacked in several ways and the most common being spoofing the input sensor. Therefore, anti-spoofing is one of the most essential prerequisite against attacks on biometric systems. For face recognition it is even more vulnerable as the image capture is non-contact based. Several anti-spoofing methods have been proposed in the literature for both contact and non-contact based biometric modalities often using video to study the temporal characteristics of a real vs. spoofed biometric signal. This paper presents a novel multi-feature evidence aggregation method for face spoofing detection. The proposed method fuses evidence from features encoding of both texture and motion (liveness) properties in the face and also the surrounding scene regions. The feature extraction algorithms are based on a configuration of local binary pattern and motion estimation using histogram of oriented optical flow. Furthermore, the multi-feature windowed videolet aggregation of these orthogonal features coupled with support vector machine-based classification provides robustness to different attacks. We demonstrate the efficacy of the proposed approach by evaluating on three standard public databases: CASIA-FASD, 3DMAD and MSU-MFSD with equal error rate of 3.14%, 0%, and 0%, respectively. Talha Ahmad Siddiqui, Samarth Bharadwaj, Tejas I. Dhamecha, Akshay Agarwal 0001, Mayank Vatsa, Richa Singh 0001, Nalini K. Ratha |
ICPR | 6 |
| 2016 | Discriminative FaceTopics for face recognition via latent Dirichlet allocationabstractLatent Dirichlet Allocation is a widely used approach for topic modeling and it has been successfully applied in several information retrieval applications. In this paper, we introduce this modeling technique for face recognition, by making an analogy between the two domains. We utilize latent Dirichlet allocation to represent facial regions in terms of FaceTopics. Further, linear discriminant analysis is utilized to obtain discriminative FaceTopics which are more suitable for classification tasks. The performance of the proposed approach is evaluated on the CMU-MultiPIE dataset under illumination and expression variations. The evaluation on over more than 50k images shows the effectiveness of the proposed approach. Further, the proposed approach shows improved identification results on e-PRIP dataset for matching composite sketches to photos. Tejas I. Dhamecha, Praneet Sharma, Richa Singh 0001, Mayank Vatsa |
WACV | 3 |
| 2016 | Effect of illicit drug abuse on face recognitionabstractOver the years, significant research has been undertaken to improve the performance of face recognition in the presence of covariates such as variations in pose, illumination, expressions, aging, and use of disguises. This paper highlights the effect of illicit drug abuse on facial features. An Illicit Drug Abuse Face (IDAF) database of 105 subjects has been created to study the performance on two commercial face recognition systems and popular face recognition algorithms. The experimental results show the decreased performance of current face recognition algorithms on drug abuse face images. This paper also proposes projective Dictionary learning based illicit Drug Abuse face Classification (DDAC) framework to effectively detect and separate faces affected by drug abuse from normal faces. This important pre-processing step stimulates researchers to develop a new class of face recognition algorithms specifically designed to improve the face recognition performance on faces affected by drug abuse. The highest classification accuracy of 88.81% is observed to detect such faces by the proposed DDAC framework on a combined database of illicit drug abuse and regular faces. Daksha Yadav, Naman Kohli, Prateekshit Pandey, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
WACV | 4 |
| 2016 | Sketch Recognition: What Lies Ahead?
Shruti Nagpal, Mayank Vatsa, Richa Singh 0001 |
Image Vis. Comput. | 3 |
| 2016 | On incremental semi-supervised discriminant analysis
Tejas I. Dhamecha, Richa Singh 0001, Mayank Vatsa |
Pattern Recognit. | 2 |
| 2016 | Incremental granular relevance vector machine: A case study in multimodal biometrics
Hunny Mehrotra, Richa Singh 0001, Mayank Vatsa, Banshidhar Majhi |
Pattern Recognit. | 2 |
| 2016 | Domain Specific Learning for Newborn Face RecognitionabstractBiometric recognition of newborn babies is an opportunity for the realization of several useful applications, such as improved security against swapping and abduction, accurate census, and effective drug delivery. This paper explores the possibility of using face recognition toward an affordable and friendly biometric modality for newborns. The paper proposes an autoencoder-based feature representation followed by problem specific distance metric learning via one-shot similarity with one class-online support vector machine. The largest publicly available database of newborns collected from various sources to study face recognition is introduced. Several existing face recognition approaches and commercial systems are also evaluated on a common benchmark protocol. The efficacy of the proposed algorithm is evaluated under both verification and identification settings. With multiple galleries, rank-1 identification accuracy of 78.5% and verification accuracy of 63.4% at 0.1% false accept rate are achieved. Samarth Bharadwaj, Himanshu S. Bhatt, Mayank Vatsa, Richa Singh 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2016 | Detecting Facial Retouching Using Supervised Deep LearningabstractDigitally altering, or retouching, face images is a common practice for images on social media, photo sharing websites, and even identification cards when the standards are not strictly enforced. This research demonstrates the effect of digital alterations on the performance of automatic face recognition, and also introduces an algorithm to classify face images as original or retouched with high accuracy. We first introduce two face image databases with unaltered and retouched images. Face recognition experiments performed on these databases show that when a retouched image is matched with its original image or an unaltered gallery image, the identification performance is considerably degraded, with a drop in matching accuracy of up to 25%. However, when images are retouched with the same style, the matching accuracy can be misleadingly high in comparison with matching original images. To detect retouching in face images, a novel supervised deep Boltzmann machine algorithm is proposed. It uses facial parts to learn discriminative features to classify face images as original or retouched. The proposed approach for classifying images as original or retouched yields an accuracy of over 87% on the data sets introduced in this paper and over 99% on three other makeup data sets used by previous researchers. This is a substantial increase in accuracy over the previous state-of-the-art algorithm, which has shown <;50% accuracy in classifying original and retouched images from the ND-IIITD retouched faces database. Aparna Bharati, Richa Singh 0001, Mayank Vatsa, Kevin W. Bowyer |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | QFuse: Online learning framework for adaptive biometric system
Samarth Bharadwaj, Himanshu S. Bhatt, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
Pattern Recognit. | 3 |
| 2014 | Aiding face recognition with social context association rule based re-rankingabstractHumans are very efficient at recognizing familiar face images even in challenging conditions. One reason for such capabilities is the ability to understand social context between individuals. Sometimes the identity of the person in a photo can be inferred based on the identity of other persons in the same photo, when some social context between them is known. This research presents an algorithm to utilize co-occurrence of individuals as the social context to improve face recognition. Association rule mining is utilized to infer multi-level social context among subjects from a large repository of social transactions. The results are demonstrated on the G-album and on the SN-collection pertaining to 4675 identities prepared by the authors from a social networking Web site. The results show that association rules extracted from social context can be used to augment face recognition and improve the identification performance. Samarth Bharadwaj, Mayank Vatsa, Richa Singh 0001 |
IJCB | 3 |
| 2014 | MDLFace: Memorability augmented deep learning for video face recognitionabstractVideos have ample amount of information in the form of frames that can be utilized for feature extraction and matching. However, face images in not all of the frames are “memorable” and useful. Therefore, utilizing all the frames available in a video for recognition does not necessarily improve the performance but significantly increases the computation time. In this research, we present a memorability based frame selection algorithm that enables automatic selection of memorable frames for facial feature extraction and matching. A deep learning algorithm is then proposed that utilizes a stack of denoising autoencoders and deep Boltzmann machines to perform face recognition using the most memorable frames. The proposed algorithm, termed as MDLFace, is evaluated on two publicly available video face databases, Youtube Faces and Point and Shoot Challenge. The results show that the proposed algorithm achieves state-of-the-art performance at low false accept rates. Gaurav Goswami, Romil Bhardwaj, Richa Singh 0001, Mayank Vatsa |
IJCB | 3 |
| 2014 | Recognizing composite sketches with digital face images via SSD dictionaryabstractSketch recognition has important law enforcement applications in detecting and apprehending suspects. Compared to hand drawn sketches, software generated composite sketches are faster to create and require lesser skill sets as well as bring consistency in sketch generation. While sketch generation is one side of the problem, recognizing composite sketches with digital images is another side. This paper presents an algorithm to address the second problem, i.e. matching composite sketches with digital images. The proposed algorithm utilizes a SSD based dictionary generated via 50,000 images from the CMU Multi-PIE database. The gallery-probe feature vectors created using SSD dictionary are matched using GentleBoostKO classifier. The results on extended PRIP composite sketch database show the effectiveness of the proposed algorithm. Paritosh Mittal, Aishwarya Jain, Gaurav Goswami, Richa Singh 0001, Mayank Vatsa |
IJCB | 4 |
| 2014 | On latent fingerprint minutiae extraction using stacked denoising sparse AutoEncodersabstractLatent fingerprint identification is of critical importance in criminal investigation. FBI's Next Generation Identification program demands latent fingerprint identification to be performed in lights-out mode, with very little or no human intervention. However, the performance of an automated latent fingerprint identification is limited due to imprecise automated feature (minutiae) extraction, specifically due to noisy ridge pattern and presence of background noise. In this paper, we propose a novel descriptor based minutiae detection algorithm for latent fingerprints. Minutia and non-minutia descriptors are learnt from a large number of tenprint fingerprint patches using stacked denoising sparse autoencoders. Latent fingerprint minutiae extraction is then posed as a binary classification problem to classify patches as minutia or non-minutia patch. Experiments performed on the NIST SD-27 database shows promising results on latent fingerprint matching. Anush Sankaran, Prateekshit Pandey, Mayank Vatsa, Richa Singh 0001 |
IJCB | 4 |
| 2014 | Leap signature recognition using HOOF and HOT featuresabstractWith the growing need for secure authentication, there is an increasing interest in establishing newer biometric modalities that are verifiable in a fast manner with as few associated complexities as possible. In this research, we propose a new biometric modality using a Leap Motion device. The Leap signature is created by an individual in three-dimensional space in absence of any feedback from objects or surfaces. The proposed framework combines an adaptation of 3D Histogram of Oriented Optical Flow and a new feature descriptor, termed as Histogram of Oriented Trajectories. Experiments are performed on the IIITD Leap Signature Database, which consists of 900 samples from 60 subjects. The results are combined with a four-patch local binary pattern based face verification algorithm. An accuracy of over 91% is achieved on this database, with rate of successful spoofing attempts being approximately 1.4%. Ishan Nigam, Mayank Vatsa, Richa Singh 0001 |
ICIP | 3 |
| 2014 | On cross spectral periocular recognitionabstractThis paper introduces the challenge of cross spectral periocular matching. The proposed algorithm utilizes neural network for learning the variabilities caused by two different spectrums. Two neural networks are first trained on each spectrum individually and then combined such that, by using the cross spectral training data, they jointly learn the cross spectral variability. To evaluate the performance, a cross spectral periocular database is prepared that contains images pertaining to visible night vision and near infrared spectrums. The proposed combined neural network architecture, on the cross spectral database, shows improved performance compared to existing feature descriptors and cross domain algorithms. Shalini Verma, Mayank Vatsa, Richa Singh 0001 |
ICIP | 4 |
| 2014 | On Effectiveness of Histogram of Oriented Gradient Features for Visible to Near Infrared Face MatchingabstractThe advent of near infrared imagery and it's applications in face recognition has instigated research in cross spectral (visible to near infrared) matching. Existing research has focused on extracting textural features including variants of histogram of oriented gradients. This paper focuses on studying the effectiveness of these features for cross spectral face recognition. On NIR-VIS-2.0 cross spectral face database, three HOG variants are analyzed along with dimensionality reduction approaches and linear discriminant analysis. The results demonstrate that DSIFT with subspace LDA outperforms a commercial matcher and other HOG variants by at least 15%. We also observe that histogram of oriented gradient features are able to encode similar facial features across spectrums. Tejas I. Dhamecha, Praneet Sharma, Richa Singh 0001, Mayank Vatsa |
ICPR | 3 |
| 2014 | On Iris Spoofing Using Print AttackabstractHuman iris contains rich textural information which serves as the key information for biometric identifications. It is very unique and one of the most accurate biometric modalities. However, spoofing techniques can be used to obfuscate or impersonate identities and increase the risk of false acceptance or false rejection. This paper revisits iris recognition with spoofing attacks and analyzes their effect on the recognition performance. Specifically, print attack with contact lens variations is used as the spoofing mechanism. It is observed that print attack and contact lens, individually and in conjunction, can significantly change the inter-personal and intra-personal distributions and thereby increase the possibility to deceive the iris recognition systems. The paper also presents the IIITD iris spoofing database, which contains over 4800 iris images pertaining to over 100 individuals with variations due to contact lens, sensor, and print attack. Finally, the paper also shows that cost effective descriptor approaches may help in counter-measuring spooking attacks. Priyanshu Gupta, Shipra Behera, Mayank Vatsa, Richa Singh 0001 |
ICPR | 4 |
| 2014 | FaceDCAPTCHA: Face detection based color image CAPTCHA
Gaurav Goswami, Brian M. Powell, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
Future Gener. Comput. Syst. | 4 |
| 2014 | Saliency based mass detection from screening mammograms
Praful Agrawal, Mayank Vatsa, Richa Singh 0001 |
Signal Process. | 3 |
| 2014 | On Recognizing Faces in Videos Using Clustering-Based Re-Ranking and FusionabstractDue to widespread applications, availability of large intra-personal variations in video and limited information content in still images, video-based face recognition has gained significant attention. Unlike still face images, videos provide abundant information that can be leveraged to address variations in pose, illumination, and expression as well as enhance the face recognition performance. This paper presents a video-based face recognition algorithm that computes a discriminative video signature as an ordered list of still face images from a large dictionary. A three-stage approach is proposed for optimizing ranked lists across multiple video frames and fusing them into a single composite ordered list to compute the video signature. This signature embeds diverse intra-personal variations and facilitates in matching two videos with large variations. For matching two videos, a discounted cumulative gain measure is utilized, which uses the ranking of images in the video signature as well as the usefulness of images in characterizing the individual in the video. The efficacy of the proposed algorithm is evaluated under different video-based face recognition scenarios such as matching still face images with videos and matching videos with videos. The efficacy of the proposed algorithm is demonstrated on the YouTube faces database and the MBGC v2 video challenge database that comprise different types of video-based face recognition challenges such as matching still face images with videos and matching videos with videos. Performance comparison with the benchmark results on both the databases and a commercial face recognition system shows the efficiency of the proposed algorithm for video-based face recognition. Himanshu S. Bhatt, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | RGB-D Face Recognition With Texture and Attribute FeaturesabstractFace recognition algorithms generally utilize 2D images for feature extraction and matching. To achieve higher resilience toward covariates, such as expression, illumination, and pose, 3D face recognition algorithms are developed. While it is challenging to use specialized 3D sensors due to high cost, RGB-D images can be captured by low-cost sensors such as Kinect. This research introduces a novel face recognition algorithm using RGB-D images. The proposed algorithm computes a descriptor based on the entropy of RGB-D faces along with the saliency feature obtained from a 2D face. Geometric facial attributes are also extracted from the depth image and face recognition is performed by fusing both the descriptor and attribute match scores. The experimental results indicate that the proposed algorithm achieves high face recognition accuracy on RGB-D images obtained using Kinect compared with existing 2D and 3D approaches. Gaurav Goswami, Mayank Vatsa, Richa Singh 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Unraveling the Effect of Textured Contact Lenses on Iris RecognitionabstractThe presence of a contact lens, particularly a textured cosmetic lens, poses a challenge to iris recognition as it obfuscates the natural iris patterns. The main contribution of this paper is to present an in-depth analysis of the effect of contact lenses on iris recognition. Two databases, namely, the IIIT-D Iris Contact Lens database and the ND-Contact Lens database, are prepared to analyze the variations caused due to contact lenses. We also present a novel lens detection algorithm that can be used to reduce the effect of contact lenses. The proposed approach outperforms other lens detection algorithms on the two databases and shows improved iris recognition performance. Daksha Yadav, Naman Kohli, James S. Doyle Jr., Richa Singh 0001, Mayank Vatsa, Kevin W. Bowyer |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Improving Cross-Resolution Face Matching Using Ensemble-Based Co-Transfer LearningabstractFace recognition algorithms are generally trained for matching high-resolution images and they perform well for similar resolution test data. However, the performance of such systems degrades when a low-resolution face image captured in unconstrained settings, such as videos from cameras in a surveillance scenario, are matched with high-resolution gallery images. The primary challenge, here, is to extract discriminating features from limited biometric content in low-resolution images and match it to information rich high-resolution face images. The problem of cross-resolution face matching is further alleviated when there is limited labeled positive data for training face recognition algorithms. In this paper, the problem of cross-resolution face matching is addressed where low-resolution images are matched with high-resolution gallery. A co-transfer learning framework is proposed, which is a cross-pollination of transfer learning and co-training paradigms and is applied for cross-resolution face matching. The transfer learning component transfers the knowledge that is learnt while matching high-resolution face images during training to match low-resolution probe images with high-resolution gallery during testing. On the other hand, co-training component facilitates this transfer of knowledge by assigning pseudolabels to unlabeled probe instances in the target domain. Amalgamation of these two paradigms in the proposed ensemble framework enhances the performance of cross-resolution face recognition. Experiments on multiple face databases show the efficacy of the proposed algorithm and compare with some existing algorithms and a commercial system. In addition, several high profile real-world cases have been used to demonstrate the usefulness of the proposed approach in addressing the tough challenges. Himanshu S. Bhatt, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha |
IEEE Trans. Image Process. | 2 |
| 2013 | Can holistic representations be used for face biometric quality assessment?abstractA face quality metric must quantitatively measure the usability of an image as a biometric sample. Though it is well established that quality measures are an integral part of robust face recognition systems, automatic measurement of bio-metric quality in face is still challenging. Inspired by scene recognition research, this paper investigates the use of holistic super-ordinate representations, namely, Gist and sparsely pooled Histogram of Orientated Gradient (HOG), in classifying images into different quality categories that are derived from matching performance. The experiments on the CAS-PEAL and SCFace databases containing covariates such as illumination, expression, pose, low-resolution and occlusion by accessories, suggest that the proposed algorithm can efficiently classify input face image into relevant quality categories and be utilized in face recognition systems. Samarth Bharadwaj, Mayank Vatsa, Richa Singh 0001 |
ICIP | 3 |
| 2013 | On rank aggregation for face recognition from videosabstractFace recognition from still face images suffers due to intrapersonal variations caused by pose, illumination, and expression that degrade the performance. On the other hand, videos provide abundant information that can be leveraged to compensate the limitations of still face images and enhance face recognition performance. This paper presents a video based face recognition algorithm that computes a discriminative video signature as an ordered list of still face images. The video signature embeds diverse intra-personal and temporal variations across multiple frames, thus facilitates matching two videos with large variations. Two videos are matched by comparing their discriminative signatures using the Kendall tau similarity distance measure. Performance comparison with the benchmark results and a commercial face recognition system on the publicly available YouTube faces database show the efficacy of the proposed video based face recognition algorithm. Himanshu S. Bhatt, Richa Singh 0001, Mayank Vatsa |
ICIP | 2 |
| 2013 | Boosting local descriptors for matching composite and digital face imagesabstractSketch recognition is one of the most challenging applications of face recognition. Due to the incorrectness of features in the witness description, standard face recognition algorithms are generally not applicable to matching sketches with digital face images. This research designs a patch based face recognition algorithm that generates patches around fiducial features and extracts local information from these patches using Daisy descriptor. The information extracted from these patches are then efficiently matched using GentleBoostKO algorithm. The experiments performed on the PRIP composite face image database show that the proposed algorithm yields promising results and outperforms existing state-of-the-art algorithms and a commercial system. Paritosh Mittal, Aishwarya Jain, Richa Singh 0001, Mayank Vatsa |
ICIP | 3 |
| 2013 | Recognizing Surgically Altered Face Images Using Multiobjective Evolutionary AlgorithmabstractWidespread acceptability and use of biometrics for person authentication has instigated several techniques for evading identification. One such technique is altering facial appearance using surgical procedures that has raised a challenge for face recognition algorithms. Increasing popularity of plastic surgery and its effect on automatic face recognition has attracted attention from the research community. However, the nonlinear variations introduced by plastic surgery remain difficult to be modeled by existing face recognition systems. In this research, a multiobjective evolutionary granular algorithm is proposed to match face images before and after plastic surgery. The algorithm first generates non-disjoint face granules at multiple levels of granularity. The granular information is assimilated using a multiobjective genetic approach that simultaneously optimizes the selection of feature extractor for each face granule along with the weights of individual granules. On the plastic surgery face database, the proposed algorithm yields high identification accuracy as compared to existing algorithms and a commercial face recognition system. Himanshu S. Bhatt, Samarth Bharadwaj, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2012 | Work in progress: On entrance test criteria for CS and IT UG programsabstractIndian universities broadly follow an entrance process for undergraduate engineering process in which physics, chemistry and mathematics based questions are asked. Either the individual scores or some combination of these scores is used to set the admission criterion for different engineering streams. However, it has not been analyzed whether this is a good predictor of undergraduate performance or not, especially in Indian context. The broad goal of this research is to understand the relationship between entrance criterion and undergraduate Computer Science/Information Technology performance. The analysis includes data collected from the students who appeared in the entrance exams and took admission in IIIT-Delhi in between 2008 to 2010. The scores obtained in entrance exams and cumulative grade point average earned during the first year of their studies are analyzed. Preliminary analysis suggests that for computer science undergraduate studies, mathematics or numerical understanding is an important aspect. Further, logic and aptitude based tests are better correlated in context of computer science education. Richa Singh 0001, Mayank Pundir |
FIE | 1 |
| 2012 | Work in progress: A quantitative study of effectiveness in group learningabstractIt is generally assumed that group studies are more effective for students than individual studies. The objective of this work in progress is to quantitatively evaluate and analyze the effect of collaborative studies on individual students performance. This effort would help the student stimulate interest in group learning and collaboration along with exposing them towards multiple problem solving approaches while working individually or in groups. This way the students are challenged to use their existing knowledge and approach, and augment it further with the knowledge and approach provided by group partners. While there are several efforts that focus on developing new group learning techniques, we intend to study the efficacy of previously proposed techniques under various test settings for EE and CS courses without significantly diverting from the course framework. Saket Srivastava, Richa Singh 0001 |
FIE | 2 |
| 2012 | Matching cross-resolution face images using co-transfer learningabstractFace recognition systems, trained in controlled environment, often fail to efficiently match low resolution images with high resolution images. In this research, a co-transfer learning framework is proposed in which knowledge learnt in controlled high resolution environment is transferred for matching low resolution probe images with high resolution gallery. The proposed framework seamlessly combines transfer learning and co-training to perform knowledge transfer by updating classifier's decision boundary with low resolution probe instances. Experiments are performed on the CMU-Multi-PIE and SCface database with gallery images of size 72 × 72 and size of probe images varying from 48 × 48 to 16 × 16. The results show that, in terms of rank-1 identification accuracy, the proposed algorithm outperforms existing approaches by at least 5%. Himanshu S. Bhatt, Richa Singh 0001, Mayank Vatsa, Nalini K. Ratha |
ICIP | 2 |
| 2012 | Incremental subclass discriminant analysis: A case study in face recognitionabstractSubclass discriminant analysis is found to be applicable under various scenarios. However, it is computationally expensive to update the between-class and within-class scatter matrices in batch mode. This research presents an incremental subclass discriminant analysis algorithm to update SDA in incremental manner with increasing number of samples per class. The effectiveness of the proposed algorithm is demonstrated using face recognition in terms of identification accuracy and training time. Experiments are performed on the AR face database and compared with other subspace based incremental and batch learning algorithms. The results illustrate that, compared to SDA, incremental SDA yields significant reduction in time along with comparable accuracy. Hemank Lamba, Tejas I. Dhamecha, Mayank Vatsa, Richa Singh 0001 |
ICIP | 4 |
| 2012 | Memetically Optimized MCWLD for Matching Sketches With Digital Face ImagesabstractOne of the important cues in solving crimes and apprehending criminals is matching sketches with digital face images. This paper presents an automated algorithm to extract discriminating information from local regions of both sketches and digital face images. Structural information along with minute details present in local facial regions are encoded using multiscale circular Weber's local descriptor. Further, an evolutionary memetic optimization algorithm is proposed to assign optimal weight to every local facial region to boost the identification performance. Since forensic sketches or digital face images can be of poor quality, a preprocessing technique is used to enhance the quality of images and improve the identification performance. Comprehensive experimental evaluation on different sketch databases show that the proposed algorithm yields better identification performance compared to existing face recognition algorithms and two commercial face recognition systems. Himanshu S. Bhatt, Samarth Bharadwaj, Richa Singh 0001, Mayank Vatsa |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Evolutionary granular approach for recognizing faces altered due to plastic surgeryabstractRecognizing faces with altered appearances is a challenging task and is only now beginning to be addressed by researchers. The paper presents an evolutionary granular approach for matching face images that have been altered by plastic surgery procedures. The algorithm extracts discriminating information from non-disjoint face granules obtained at different levels of granularity. At the first level of granularity, both pre and post-surgery face images are processed by Gaussian and Laplacian operators to obtain face granules at varying resolutions. The second level of granularity divides face image into horizontal and vertical face granules of varying size and information content. At the third level of granularity, face image is tessellated into non-overlapping local facial regions. An evolutionary approach is proposed using genetic algorithm to simultaneously optimize the selection of feature extractor for each face granule along with finding optimal weights corresponding to each face granule for matching. Experiments on pre and post-plastic surgery face images show that the proposed algorithm provides at least 15% better identification performance as compared to other face recognition algorithms. Himanshu S. Bhatt, Samarth Bharadwaj, Richa Singh 0001, Mayank Vatsa, Afzel Noore |
FG | 3 |
| 2011 | On co-training online biometric classifiersabstractIn an operational biometric verification system, changes in biometric data over a period of time can affect the classification accuracy. Online learning has been used for updating the classifier decision boundary. However, this requires labeled data that is only available during new enrolments. This paper presents a biometric classifier update algorithm in which the classifier decision boundary is updated using both labeled enrolment instances and unlabeled probe in- stances. The proposed co-training online classifier update algorithm is presented as a semi-supervised learning task and is applied to a face verification application. Experiments indicate that the proposed algorithm improves the performance both in terms of classification accuracy and computational time. Himanshu S. Bhatt, Samarth Bharadwaj, Richa Singh 0001, Mayank Vatsa, Afzel Noore, Arun Ross |
IJCB | 3 |
| 2011 | A framework for quality-based biometric classifier selectionabstractMultibiometric systems fuse the evidence (e.g., match scores) pertaining to multiple biometric modalities or classifiers. Most score-level fusion schemes discussed in the literature require the processing (i.e., feature extraction and matching) of every modality prior to invoking the fusion scheme. This paper presents a framework for dynamic classifier selection and fusion based on the quality of the gallery and probe images associated with each modality with multiple classifiers. The quality assessment algorithm for each biometric modality computes a quality vector for the gallery and probe images that is used for classifier selection. These vectors are used to train Support Vector Machines (SVMs) for decision making. In the proposed framework, the bio- metric modalities are arranged sequentially such that the stronger biometric modality has higher priority for being processed. Since fusion is required only when all unimodal classifiers are rejected by the SVM classifiers, the average computational time of the proposed framework is significantly reduced. Experimental results on different multi-modal databases involving face and fingerprint show that the proposed quality-based classifier selection framework yields good performance even when the quality of the bio- metric sample is sub-optimal. Himanshu S. Bhatt, Samarth Bharadwaj, Mayank Vatsa, Richa Singh 0001, Arun Ross, Afzel Noore |
IJCB | 4 |
| 2011 | Is gender classification across ethnicity feasible using discriminant functions?abstractOver the years, automatic gender recognition has been used in many applications. However, limited research has been done on analyzing gender recognition across ethnicity scenario. This research aims at studying the performance of discriminant functions including Principal Component Analysis, Linear Discriminant Analysis and Subclass Discriminant Analysis with the availability of limited training database and unseen ethnicity variations. The experiments are performed on a heterogeneous database of 8112 images that includes variations in illumination, expression, minor pose and ethnicity. Contrary to existing literature, the results show that PCA provides comparable but slightly better performance compared to PCA+LDA, PCA+SDA and PCA+SVM. The results also suggest that linear discriminant functions provide good generalization capability even with limited number of training samples, principal components and with cross-ethnicity variations. Tejas I. Dhamecha, Anush Sankaran, Richa Singh 0001, Mayank Vatsa |
IJCB | 3 |
| 2011 | Face recognition for look-alikes: A preliminary studyabstractOne of the major challenges efface recognition is to design a feature extractor and matcher that reduces the intra class variations and increases the inter-class variations. The feature extraction algorithm has to be robust enough to extract similar features for a particular subject despite variations in quality, pose, illumination, expression, aging, and disguise. The problem is exacerbated when there are two individuals with lower inter-class variations, i.e., look alikes. In such cases, the intra-class similarity is higher than the inter-class variation for these two individuals. This research explores the problem of look-alike faces and their effect on human performance and automatic face recognition algorithms. There is three fold contribution in this re search: firstly, we analyze the human recognition capabilities for look-alike appearances. Secondly, we compare human recognition performance with ten existing face recognition algorithms, and finally, proposed an algorithm to improve the face verification accuracy. The analysis shows that neither humans nor automatic face recognition algorithms are efficient in recognizing look-alikes. Hemank Lamba, Ankit Sarkar, Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IJCB | 4 |
| 2011 | On matching latent to latent fingerprintsabstractThis research presents a forensics application of matching two latent fingerprints. In crime scene settings, it is often required to match multiple latent fingerprints. Unlike matching latent with inked or live fingerprints, this research problem is very challenging and requires proper analysis and attention. The contribution of this paper is three fold: (i) a comparative analysis of existing algorithms is presented for this application, (ii) fusion and context switching frameworks are presented to improve the identification performance, and (Hi) a multi-latent fingerprint database is prepared. The experiments highlight the need for improved feature extraction and processing methods and exhibit large scope of improvement in this important research problem. Anush Sankaran, Tejas I. Dhamecha, Mayank Vatsa, Richa Singh 0001 |
IJCB | 4 |
| 2010 | Quality-Based Fusion for Multichannel Iris RecognitionabstractWe propose a quality-based fusion scheme for improving the recognition accuracy using color iris images characterized by three spectral channels - Red, Green and Blue. In the proposed method, quality scores are employed to select two channels of a color iris image which are fused at the image level using a Redundant Discrete Wavelet Transform (RDWT). The fused image is then used in a score-level fusion framework along with the remaining channel to improve recognition accuracy. Experimental results on a heterogenous color iris database demonstrate the efficacy of the technique when compared against other score-level and image-level fusion methods. The proposed method can potentially benefit the use of color iris images in conjunction with their NIR counterparts. Mayank Vatsa, Richa Singh 0001, Arun Ross, Afzel Noore |
ICPR | 2 |
| 2010 | Biometric classifier update using online learning: A case study in near infrared face verification
Richa Singh 0001, Mayank Vatsa, Arun Ross, Afzel Noore |
Image Vis. Comput. | 1 |
| 2010 | Plastic surgery: a new dimension to face recognitionabstractAdvancement and affordability is leading to the popularity of plastic surgery procedures. Facial plastic surgery can be reconstructive to correct facial feature anomalies or cosmetic to improve the appearance. Both corrective as well as cosmetic surgeries alter the original facial information to a large extent thereby posing a great challenge for face recognition algorithms. The contribution of this research is 1) preparing a face database of 900 individuals for plastic surgery, and 2) providing an analytical and experimental underpinning of the effect of plastic surgery on face recognition algorithms. The results on the plastic surgery database suggest that it is an arduous research challenge and the current state-of-art face recognition algorithms are unable to provide acceptable levels of identification performance. Therefore, it is imperative to initiate a research effort so that future face recognition systems will be able to address this important problem. Richa Singh 0001, Mayank Vatsa, Himanshu S. Bhatt, Samarth Bharadwaj, Afzel Noore, Shahin S. Nooreyezdan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2010 | On the dynamic selection of biometric fusion algorithmsabstractBiometric fusion consolidates the output of multiple biometric classifiers to render a decision about the identity of an individual. We consider the problem of designing a fusion scheme when 1) the number of training samples is limited, thereby affecting the use of a purely density-based scheme and the likelihood ratio test statistic; 2) the output of multiple matchers yields conflicting results; and 3) the use of a single fusion rule may not be practical due to the diversity of scenarios encountered in the probe dataset. To address these issues, a dynamic reconciliation scheme for fusion rule selection is proposed. In this regard, the contribution of this paper is two-fold: 1) the design of a sequential fusion technique that uses the likelihood ratio test-statistic in conjunction with a support vector machine classifier to account for errors in the former; and 2) the design of a dynamic selection algorithm that unifies the constituent classifiers and fusion schemes in order to optimize both verification accuracy and computational cost. The case study in multiclassifier face recognition suggests that the proposed algorithm can address the issues listed above. Indeed, it is observed that the proposed method performs well even in the presence of confounding covariate factors thereby indicating its potential for large-scale face recognition. Mayank Vatsa, Richa Singh 0001, Afzel Noore, Arun Ross |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Quality-augmented fusion of level-2 and level-3 fingerprint information using DSm theory
Mayank Vatsa, Richa Singh 0001, Afzel Noore, Max M. Houck |
Int. J. Approx. Reason. | 2 |
| 2009 | Face recognition with disguise and single gallery images
Richa Singh 0001, Mayank Vatsa, Afzel Noore |
Image Vis. Comput. | 1 |
| 2009 | Feature based RDWT watermarking for multimodal biometric system
Mayank Vatsa, Richa Singh 0001, Afzel Noore |
Image Vis. Comput. | 2 |
| 2009 | Combining pores and ridges with minutiae for improved fingerprint verification
Mayank Vatsa, Richa Singh 0001, Afzel Noore, Sanjay Kumar Singh 0001 |
Signal Process. | 2 |
| 2009 | Unification of Evidence-Theoretic Fusion Algorithms: A Case Study in Level-2 and Level-3 Fingerprint FeaturesabstractThis paper formulates an evidence-theoretic multimodal unification approach using belief functions that take into account the variability in biometric image characteristics. While processing nonideal images, the variation in the quality of features at different levels of abstraction may cause individual classifiers to generate conflicting genuine-impostor decisions. Existing fusion approaches are nonadaptive and do not always guarantee optimum performance improvements. We propose a contextual unification framework to dynamically select the most appropriate evidence-theoretic fusion algorithm for a given scenario. In the first approach, the unification framework uses deterministic rules to select the most appropriate fusion algorithm; while in the second approach, the framework intelligently learns from the input evidences using a 2nu-granular support vector machine. The effectiveness of the unification approach is experimentally validated by fusing match scores from level-2 and level-3 fingerprint features. Compared to existing fusion algorithms, the proposed unification approach is computationally efficient, and the verification accuracy is not compromised even when conflicting decisions are encountered. Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2008 | Multiclass mv-granular soft support vector machine: A case study in dynamic classifier selection for multispectral face recognitionabstractThis paper presents a novel formulation of multiclass support vector machine by integrating the concepts of soft labels and granular computing. The proposed multiclass mv-granular soft support vector machine uses soft labels to address the issues due to noisy and incorrectly labeled data, and granular computing to make it adaptable to data distributions both globally and locally. The proposed multiclass classifier is used for dynamic selection in a multispectral face recognition application. Specifically, for the given probe face images, mv-GSSVM is used to optimally choose one of the four options: visible spectrum face recognition, short-wave infrared face recognition, multispectral face image fusion, and multispectral match score fusion. Experimental results on a multispectral face database show that the proposed algorithm improves the verification accuracy and also decreases the computational time. Richa Singh 0001, Mayank Vatsa, Afzel Noore |
ICPR | 1 |
| 2008 | Integrated multilevel image fusion and match score fusion of visible and infrared face images for robust face recognition
Richa Singh 0001, Mayank Vatsa, Afzel Noore |
Pattern Recognit. | 1 |
| 2008 | Improving Iris Recognition Performance Using Segmentation, Quality Enhancement, Match Score Fusion, and IndexingabstractThis paper proposes algorithms for iris segmentation, quality enhancement, match score fusion, and indexing to improve both the accuracy and the speed of iris recognition. A curve evolution approach is proposed to effectively segment a nonideal iris image using the modified Mumford-Shah functional. Different enhancement algorithms are concurrently applied on the segmented iris image to produce multiple enhanced versions of the iris image. A support-vector-machine-based learning algorithm selects locally enhanced regions from each globally enhanced image and combines these good-quality regions to create a single high-quality iris image. Two distinct features are extracted from the high-quality iris image. The global textural feature is extracted using the 1-D log polar Gabor transform, and the local topological feature is extracted using Euler numbers. An intelligent fusion algorithm combines the textural and topological matching scores to further improve the iris recognition performance and reduce the false rejection rate, whereas an indexing algorithm enables fast and accurate iris identification. The verification and identification performance of the proposed algorithms is validated and compared with other algorithms using the CASIA Version 3, ICE 2005, and UBIRIS iris databases. Mayank Vatsa, Richa Singh 0001, Afzel Noore |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2007 | Integrating Image Quality in 2nu-SVM Biometric Match Score FusionabstractThis paper proposes an intelligent 2nu-support vector machine based match score fusion algorithm to improve the performance of face and iris recognition by integrating the quality of images. The proposed algorithm applies redundant discrete wavelet transform to evaluate the underlying linear and non-linear features present in the image. A composite quality score is computed to determine the extent of smoothness, sharpness, noise, and other pertinent features present in each subband of the image. The match score and the corresponding quality score of an image are fused using 2nu-support vector machine to improve the verification performance. The proposed algorithm is experimentally validated using the FERET face database and the CASIA iris database. The verification performance and statistical evaluation show that the proposed algorithm outperforms existing fusion algorithms. Mayank Vatsa, Richa Singh 0001, Afzel Noore |
Int. J. Neural Syst. | 2 |
| 2007 | Improving verification accuracy by synthesis of locally enhanced biometric images and deformable model
Richa Singh 0001, Mayank Vatsa, Afzel Noore |
Signal Process. | 1 |
| 2007 | A Mosaicing Scheme for Pose-Invariant Face RecognitionabstractMosaicing entails the consolidation of information represented by multiple images through the application of a registration and blending procedure. We describe a face mosaicing scheme that generates a composite face image during enrollment based on the evidence provided by frontal and semiprofile face images of an individual. Face mosaicing obviates the need to store multiple face templates representing multiple poses of a user's face image. In the proposed scheme, the side profile images are aligned with the frontal image using a hierarchical registration algorithm that exploits neighborhood properties to determine the transformation relating the two images. Multiresolution splining is then used to blend the side profiles with the frontal image, thereby generating a composite face image of the user. A texture-based face recognition technique that is a slightly modified version of the C2 algorithm proposed by Serre et al. is used to compare a probe face image with the gallery face mosaic. Experiments conducted on three different databases indicate that face mosaicing, as described in this paper, offers significant benefits by accounting for the pose variations that are commonly observed in face images. Richa Singh 0001, Mayank Vatsa, Arun Ross, Afzel Noore |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Signature Verification Using Static and Dynamic Features
Mayank Vatsa, Richa Singh 0001, Pabitra Mitra, Afzel Noore |
ICONIP | 2 |
| 2002 | A comparison of face recognition algorithms neural network based & line based approachesabstractOne of the most successful applications of image analysis and understanding, face recognition has received significant attention. There are at least two reasons for the trend: the first is the wide range of commercial and law enforcement applications and the second is the availability of feasible technologies. In general, few methods of face recognition are in practice: feature based face recognition methods, eigen face based, line based, elastic bunch graph method and neural network based methods. All have their possibilities and features. In the neural network approach automatic detection of eyes and mouth is followed by a spatial normalization of the images. The classification of the normalized images is carried out by hybrid neural network that combines unsupervised and supervised methods for finding structures and reducing classification errors respectively. The line-based is a type of image-based approach. It does not use any detailed biometric knowledge of the human face. These techniques use either the pixel-based bi-dimensional array representation of the entire face image or a set of transformed images or template sub-images of facial features as the image representation. An image-based metric such as correlation is then used to match the resulting image with the set of model images. In the context of image-based techniques, two approaches are there namely template-based and neural networks. In the template-based approach, the face is represented as a set of templates of the major facial features, which are then matched with the prototypical model face templates. Sanjay K. Singh, Mayank Vatsa, Richa Singh 0001, D. S. Chauhan |
SMC | 3 |