Anil K. Jain 0001

dblp:j/AnilKJain · also Anil Kumar Jain · DBLP profile ↗
← Back
483ranked-venue papers
80as first author
41since 2021 · last 2026
0000-0002-6369-6995ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 370 · 64 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 188 · 23 first-author · 22 since 2021Security and privacy · 67 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 32 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 32 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorSystems, architecture and hardware · 5Theory of computation · 1
YearPublicationVenuePosition
2026 50 Years of Automated Face Recognition
abstract
Over the past five decades, automated face recognition (FR) has progressed from handcrafted geometric and statistical approaches to advanced deep learning architectures that now approach, and in many cases exceed, human performance. This paper traces the historical and technological evolution of FR, encompassing early algorithmic paradigms through to contemporary neural systems trained on extensive real and synthetically generated datasets. We examine pivotal innovations that have driven this progression, including advances in dataset construction, loss function formulation, network architecture design, and feature fusion strategies. Furthermore, we analyze the relationship between data scale, diversity, and model generalization, highlighting how dataset expansion correlates with benchmark performance gains. Recent systems have achieved near-perfect large-scale identification accuracy, with the leading algorithm in the latest NIST FRTE 1:N benchmark reporting a False Negative Identification Rate (FNIR) of 0.15 percent at False Positive Identification Rate (FPIR) of 0.001 on a gallery of over 10 million identities. Larger galleries increase false positive rates and deployments at greater scales will see higher error rates. We delineate key open problems and emerging directions, including scalable training, multi-modal fusion, synthetic data, and interpretable recognition frameworks.
Anil K. Jain 0001, Xiaoming Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Adapting to the Wild: From Human Face to Animal Face Recognition
Maria De Marsico, Anil K. Jain 0001, Michele Miranda, Alessio Orlando
CAIP (1)2
2025 Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
abstract
Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of their face data with service providers. In this work, we introduce a novel method in which the face anti-spoofing model can be adapted by the client itself to a target domain at test time using only a small sample of data while keeping model parameters and training data inaccessible to the client. Specifically, we develop a prototype-based base model and an optimal transport-guided adaptor that enables adaptation in either a lightweight training or training-free fashion, without updating base model’s parameters. Furthermore, we propose geodesic mixup, an optimal transport-based synthesis method that generates augmented training data along the geodesic path between source prototypes and target data distribution. This allows training a lightweight classifier to effectively adapt to target-specific characteristics while retaining essential knowledge learned from the source domain. In cross-domain and cross-attack settings, compared with recent methods, our method achieves average relative improvements of 19.17% in HTER and 8.58% in AUC, respectively.
Zhuowei Li 0002, Tianchen Zhao, Xuanbai Chen, Alessandro Bergamo, Anil K. Jain 0001, Yifan Xing
CVPR9
2025 A Quality-Guided Mixture of Score-Fusion Experts Framework for Human Recognition
abstract
Whole-body biometric recognition is a challenging multimodal task that integrates various biometric modalities, including face, gait, and body. This integration is essential for overcoming the limitations of unimodal systems. Traditionally, whole-body recognition involves deploying different models to process multiple modalities, achieving the final outcome by score-fusion (e.g., weighted averaging of similarity matrices from each model). However, these conventional methods may overlook the variations in score distributions of individual modalities, making it challenging to improve final performance. In this work, we present \textbf{Q}uality-guided \textbf{M}ixture of score-fusion \textbf{E}xperts (QME), a novel framework designed for improving whole-body biometric recognition performance through a learnable score-fusion strategy using a Mixture of Experts (MoE). We introduce a novel pseudo-quality loss for quality estimation with a modality-specific Quality Estimator (QE), and a score triplet loss to improve the metric performance. Extensive experiments on multiple whole-body biometric datasets demonstrate the effectiveness of our proposed approach, achieving state-of-the-art results across various metrics compared to baseline methods. Our method is effective for multimodal and multi-model, addressing key challenges such as model misalignment in the similarity score domain and variability in data quality.
Yiyang Su, Anil K. Jain 0001, Xiaoming Liu 0002
ICCV4
2025 Universal Fingerprint Generation: Controllable Diffusion Model With Multimodal Conditions
abstract
The utilization of synthetic data for fingerprint recognition has garnered increased attention due to its potential to alleviate privacy concerns surrounding sensitive biometric data. However, current methods for generating fingerprints have limitations in creating impressions of the same finger with useful intra-class variations. To tackle this challenge, we present GenPrint, a framework to produce fingerprint images of various types while maintaining identity and offering humanly understandable control over different appearance factors, such as fingerprint class, acquisition type, sensor device, and quality level. Unlike previous fingerprint generation approaches, GenPrint is not confined to replicating style characteristics from the training dataset alone: it enables the generation of novel styles from unseen devices without requiring additional fine-tuning. To accomplish these objectives, we developed GenPrint using latent diffusion models with multimodal conditions (text and image) for consistent generation of style and identity. Our experiments leverage a variety of publicly available datasets for training and evaluation. Results demonstrate the benefits of GenPrint in terms of identity preservation, explainable control, and universality of generated images. Importantly, the GenPrint-generated images yield comparable or even superior accuracy to models trained solely on real data and further enhances performance when augmenting the diversity of existing real fingerprint datasets.
Steven A. Grosz, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Open-Set Biometrics: Beyond Good Closed-Set Models
Yiyang Su, Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002
ECCV (62)4
2024 GenPalm: Contactless Palmprint Generation with Diffusion Models
abstract
The scarcity of large-scale palmprint databases poses a significant bottleneck to advancements in contactless palmprint recognition. To address this, researchers have turned to synthetic data generation. While Generative Adversarial Networks (GANs) have been widely used, they suffer from instability and mode collapse. Recently, diffusion probabilistic models have emerged as a promising alternative, offering stable training and better distribution coverage. This paper introduces a novel palmprint generation method using diffusion probabilistic models, develops an end-to-end framework for synthesizing multiple palm identities, and validates the realism and utility of the generated palmprints. Experimental results demonstrate the effectiveness of our approach in generating palmprint images which enhance contactless palmprint recognition performance across several test databases utilizing challenging cross-database and time-separated evaluation protocols.
Steven A. Grosz, Anil K. Jain 0001
IJCB2
2024 CLIP4Sketch: Enhancing Sketch to Mugshot Matching through Dataset Augmentation using Diffusion Models
abstract
Forensic sketch-to-mugshot matching is a challenging task in face recognition, primarily hindered by the scarcity of annotated forensic sketches and the modality gap between sketches and photographs. To address this, we propose CLIP4Sketch, a novel approach that leverages diffusion models to generate a large and diverse set of sketch images, which helps in enhancing the performance of face recognition systems in sketch-to-mugshot matching. Our method utilizes Denoising Diffusion Probabilistic Models (DDPMs) to generate sketches with explicit control over identity and style. We combine CLIP and Adaface embeddings of a reference mugshot, along with textual descriptions of style, as the conditions to the diffusion model. We demonstrate the efficacy of our approach by generating a comprehensive dataset of sketches corresponding to mugshots and training a face recognition model on our synthetic data. Our results show significant improvements in sketch-to-mugshot matching accuracy over training on an existing, limited amount of real face sketch data, validating the potential of diffusion models in enhancing the performance of face recognition systems across modalities. We also compare our dataset with datasets generated using GAN-based methods to show its superiority.
Kushal Kumar Jain, Steven A. Grosz, Anoop M. Namboodiri, Anil K. Jain 0001
IJCB4
2024 Learning a Robust Minutiae Extractor via an Ensemble of Expert Models
abstract
Recognizing the limitations of manual methods in creating ground truth minutiae sets for fingerprint recognition, we propose an ensemble method that leverages the strengths of two well-known SDKs: Innovatrics ANSI&ISO v2.4.10 and Verifinger v12.4, referred to as experts. By combining the predictions of these two expert systems, our method aims to learn a robust ground truth minutiae set, serving as a foundation for training fingerprint-matching models. We introduce four ensemble configurations using union and intersection operations to capture a comprehensive set of minutiae, addressing the limitations inherent in any single extractor. Using the ensemble minutiae ground truth, we trained fingerprint recognition models using two distinct architectures: U-Net and Vision Transformers (ViT). These were evaluated using authentication scores and statistics were computed against a hand-marked ground truth dataset. Our experiment results are encouraging across both architectures, particularly for difficult test sets like NIST SD 302. Training with a ViT architecture led to a 1.59% improvement in matching accuracy, increasing from 92.79% for the model trained on Innovatrics ground truth minutiae to 94.38% for the model trained on the ensemble ground truth. The U-Net architecture performance improvement was 1.74%, from 92.00% to 93.74%. These are both significant improvements from the results obtained from the Innovatrics and Verifinger SDKs. These results demonstrate the effectiveness of ensemble ground truth minutiae in enhancing the performance of fingerprint recognition systems.
Arhan A. Mulay, Steven A. Grosz, Anil K. Jain 0001
IJCB3
2024 Data Pruning via Separability, Integrity, and Model Uncertainty-Aware Importance Sampling
Steven A. Grosz, Manoj Aggarwal, Gérard G. Medioni, Anil K. Jain 0001
ICPR (2)7
2024 Mobile Contactless Palmprint Recognition: Use of Multiscale, Multimodel Embeddings
abstract
Contactless palmprints are comprised of both global and local discriminative features. Most prior work focuses on extracting global features or local features alone for palmprint matching, whereas this research introduces a novel framework that combines global and local features for enhanced palmprint matching accuracy. Leveraging recent advancements in deep learning, this study integrates a vision transformer (ViT) and a convolutional neural network (CNN) to extract complementary local and global features. Next, a mobile-based, end-to-end palmprint recognition system is developed, referred to as Palm-ID. On top of the ViT and CNN features, Palm-ID incorporates a palmprint enhancement module and efficient dimensionality reduction (for faster matching). Palm-ID balances the trade-off between accuracy and latency, requiring just 18ms to extract a template of size 516 bytes, which can be efficiently searched against a 10,000 palmprint gallery in 0.33ms on an AMD EPYC 7543 32-Core CPU utilizing 128-threads. Cross-database matching protocols and evaluations on large-scale operational datasets demonstrate the robustness of the proposed method, achieving a TAR of 98.06% at FAR=0.01% on a newly collected, time-separated dataset. To show a practical deployment of the end-to-end system, the entire recognition pipeline is embedded within a mobile device for enhanced user privacy and security.
Steven A. Grosz, Akash Godbole, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2023 DCFace: Synthetic Face Generation with Dual Condition Diffusion Model
abstract
Generating synthetic datasets for training face recognition models is challenging because dataset generation entails more than creating high fidelity images. It involves generating multiple images of same subjects under different factors (e.g., variations in pose, illumination, expression, aging and occlusion) which follows the real image conditional distribution. Previous works have studied the generation of synthetic datasets using GAN or 3D models. In this work, we approach the problem from the aspect of combining subject appearance (ID) and external factor (style) conditions. These two conditions provide a direct way to control the inter-class and intra-class variations. To this end, we propose a Dual Condition Face Generator (DCFace) based on a diffusion model. Our novel Patch-wise style extractor and Time-step dependent ID loss enables DCFace to consistently produce face images of the same subject under different styles with precise control. Face recognition models trained on synthetic images from the proposed DCFace provide higher verification accuracies compared to previous works by 6.11% on average in 4 out of 5 test datasets, LFW, CFP-FP, CPLFW, AgeDB and CALFW. Code Link
Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002
CVPR3
2023 FaceGuard: A Self-Supervised Defense Against Adversarial Face Images
abstract
Prevailing defense schemes against adversarial face images tend to overfit to the perturbations in the training set and fail to generalize to unseen adversarial attacks. We propose a new self-supervised adversarial defense framework, namely FaceGuard, that can automatically detect, localize, and purify a wide variety of adversarial faces without utilizing pre-computed adversarial training samples. During training, FaceGuard automatically synthesizes challenging and diverse adversarial attacks, enabling a classifier to learn to distinguish them from real faces. Concurrently, a purifier attempts to remove the adversarial perturbations in the image space. Experimental results on LFW, Celeb-A, and FFHQ datasets show that FaceGuard can achieve 99.81%, 98.73%, and 99.35% detection accuracies, respectively, on six unseen adversarial attack types. In addition, the proposed method can enhance the face recognition performance of ArcFace from 34.27% TAR @ 0.1% FAR under no defense to 77.46% TAR @ 0.1% FAR. Code, pre-trained models and dataset will be publicly available.
Debayan Deb, Xiaoming Liu 0002, Anil K. Jain 0001
FG3
2023 Unified Detection of Digital and Physical Face Attacks
abstract
State-of-the-art defense mechanisms against face attacks achieve near perfect accuracies within one of three attack categories, namely adversarial, digital manipulation, or physical spoofs, however, they fail to generalize well when tested across all three categories. Poor generalization can be attributed to learning incoherent attacks jointly. To over-come this shortcoming, we propose a unified attack detection framework, namely UniFAD, that can automatically cluster 25 coherent attack types belonging to the three categories. Using a multi-task learning framework along with k-means clustering, UniFAD learns joint representations for coherent attacks, while uncorrelated attack types are learned separately. Proposed UniFAD outperforms prevailing defense methods and their fusion with an overall TDR = 94.73% @ 0.2% FDR on a large fake face dataset consisting of 341K bona fide images and 448K attack images of 25 types across all 3 categories. Proposed method can detect an attack within 3 milliseconds on a Nvidia 2080Ti. UniFAD can also identify the attack categories with 97.37% accuracy. Code and dataset will be publicly available.
Debayan Deb, Xiaoming Liu 0002, Anil K. Jain 0001
FG3
2023 ViT Unified: Joint Fingerprint Recognition and Presentation Attack Detection
abstract
A secure fingerprint recognition system must contain both a presentation attack (i.e., spoof) detection and recognition module in order to protect users against unwanted access by malicious users. Traditionally, these tasks would be carried out by two independent systems; however, recent studies have demonstrated the potential to have one unified system architecture in order to reduce the computational burdens on the system, while maintaining high accuracy. In this work, we leverage a vision transformer architecture for joint spoof detection and matching and report competitive results with state-of-the-art (SOTA) models for both a sequential system (two ViT models operating independently) and a unified architecture (a single ViT model for both tasks). ViT models are particularly well suited for this task as the ViT’s global embedding encodes features useful for recognition, whereas the individual, local embeddings are useful for spoof detection. We demonstrate the capability of our unified model to achieve an average integrated matching (IM) accuracy of 98.87% across LivDet 2013 and 2015 CrossMatch sensors. This is comparable to IM accuracy of 98.95% of our sequential dual-ViT system, but with $\sim 50\%$ of the parameters and $\sim 58\%$ of the latency.
Steven A. Grosz, Kanishka P. Wijewardena, Anil K. Jain 0001
IJCB3
2023 AdvGen: Physical Adversarial Attack on Face Presentation Attack Detection Systems
abstract
Evaluating the risk level of adversarial images is essential for safely deploying face authentication models in the real world. Popular approaches for physical-world attacks, such as print or replay attacks, suffer from some limitations, like including physical and geometrical artifacts. Recently adversarial attacks have gained attraction, which try to digitally deceive the learning strategy of a recognition system using slight modifications to the captured image. While most previous research assumes that the adversarial image could be digitally fed into the authentication systems, this is not always the case for systems deployed in the real world. This paper demonstrates the vulnerability of face authentication systems to adversarial images in physical world scenarios. We propose AdvGen, an automated Generative Adversarial Network, to simulate print and replay attacks and generate adversarial images that can fool state-of-the-art PADs in a physical domain attack setting. Using this attack strategy, the attack success rate goes up to 82.01%. We test AdvGen extensively on four datasets and ten state-of-the-art PADs. We also demonstrate the effectiveness of our attack by conducting experiments in a realistic, physical environment.
Sai Amrit Patnaik, Shivali Chansoriya, Anoop M. Namboodiri, Anil K. Jain 0001
IJCB4
2023 How does the Memorization of Neural Networks Impact Adversarial Robust Models?
abstract
Recent studies suggest that "memorization" is one necessary factor for overparameterized deep neural networks (DNNs) to achieve optimal performance. Specifically, the perfectly fitted DNNs can memorize the labels of many atypical samples, generalize their memorization to correctly classify test atypical samples and enjoy better test performance. While, DNNs which are optimized via adversarial training algorithms can also achieve perfect training performance by memorizing the labels of atypical samples, as well as the adversarially perturbed atypical samples. However, adversarially trained models always suffer from poor generalization, with both relatively low clean accuracy and robustness on the test set. In this work, we study the effect of memorization in adversarial trained DNNs and disclose two important findings: (a) Memorizing atypical samples is only effective to improve DNN's accuracy on clean atypical samples, but hardly improve their adversarial robustness and (b) Memorizing certain atypical samples will even hurt the DNN's performance on typical samples. Based on these two findings, we propose Benign Adversarial Training (BAT) which can facilitate adversarial training to avoid fitting "harmful" atypical samples and fit as more "benign" atypical samples as possible. In our experiments, we validate the effectiveness of BAT, and show that it can achieve better clean accuracy vs. robustness trade-off than baseline methods, in benchmark datasets for image classification.
Han Xu 0002, Wentao Wang 0006, Zitao Liu 0001, Anil K. Jain 0001, Jiliang Tang
KDD5
2023 Synthetic Latent Fingerprint Generator
abstract
Given a full fingerprint image (rolled or slap), we present CycleGAN models to generate multiple latent impressions of the same identity as the full print. Our models can control the degree of distortion, noise, blurriness and occlusion in the generated latent print images to obtain Good, Bad and Ugly latent image categories as introduced in the NIST SD27 latent database. The contributions of our work are twofold: (i) demonstrate the similarity of synthetically generated latent fingerprint images to crime scene latents in NIST SD27 and MSP databases as evaluated by the NIST NFIQ 2 quality measure and recognition accuracies obtained by a SOTA fingerprint matcher, and (ii) use of synthetic latents to augment small-size latent training databases in the public domain to improve the performance of DeepPrint, a SOTA fingerprint matcher designed for rolled to rolled fingerprint matching on three latent databases (NIST SD27, NIST SD302, and IIITD-SLF). As an example, with synthetic latent data augmentation, the Rank-1 retrieval performance of DeepPrint is improved from 15.50% to 29.07% on challenging NIST SD27 latent database. Our approach for generating synthetic latent fingerprints can be used to improve the recognition performance of any latent matcher and its individual components (e.g., enhancement, segmentation and feature extraction). https://prip-lab.github.io/Synthetic-Latent-Fingerprint-Generator/
André Brasil Vieira Wyzykowski, Anil K. Jain 0001
WACV2
2023 PrintsGAN: Synthetic Fingerprint Generator
abstract
A major impediment to researchers working in the area of fingerprint recognition is the lack of publicly available, large-scale, fingerprint datasets. The publicly available datasets that do exist contain very few identities and impressions per finger. This limits research on a number of topics, including e.g., using deep networks to learn fixed length fingerprint embeddings. Therefore, we propose PrintsGAN, a synthetic fingerprint generator capable of generating unique fingerprints along with multiple impressions for a given fingerprint. Using PrintsGAN, we synthesize a database of 525k fingerprints (35K distinct fingers, each with 15 impressions). Next, we show the utility of the PrintsGAN generated dataset by training a deep network to extract a fixed-length embedding from a fingerprint. In particular, an embedding model trained on our synthetic fingerprints and fine-tuned on a small number of publicly available real fingerprints (25K prints from NIST SD 302) obtains a TAR of 87.03% @ FAR=0.01% on the NIST SD4 database (a boost from TAR=73.37% when only trained on NIST SD 302). Prevailing synthetic fingerprint generation methods do not enable such performance gains due to i) lack of realism or ii) inability to generate multiple impressions per finger. Our dataset is released to the public: https://biometrics.cse.msu.edu/Publications/Databases/MSU_PrintsGAN/.
Joshua J. Engelsma, Steven A. Grosz, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Transfer Learning in Deep Reinforcement Learning: A Survey
abstract
Reinforcement learning is a learning paradigm for solving sequential decision-making problems. Recent years have witnessed remarkable progress in reinforcement learning upon the fast development of deep neural networks. Along with the promising prospects of reinforcement learning in numerous domains such as robotics and game-playing, transfer learning has arisen to tackle various challenges faced by reinforcement learning, by transferring knowledge from external expertise to facilitate the efficiency and effectiveness of the learning process. In this survey, we systematically investigate the recent progress of transfer learning approaches in the context of deep reinforcement learning. Specifically, we provide a framework for categorizing the state-of-the-art transfer learning approaches, under which we analyze their goals, methodologies, compatible reinforcement learning backbones, and practical applications. We also draw connections between transfer learning and other relevant topics from the reinforcement learning perspective and explore their potential challenges that await future research progress.
Zhuangdi Zhu, Kaixiang Lin, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 SpoofGAN: Synthetic Fingerprint Spoof Images
abstract
A major limitation to advances in fingerprint presentation attack detection (PAD) is the lack of publicly available, large-scale datasets, a problem which has been compounded by increased concerns surrounding privacy and security of biometric data. Furthermore, most state-of-the-art PAD algorithms rely on deep networks which perform best in the presence of a large amount of training data. This work aims to demonstrate the utility of synthetic (both bona fide and PA style) fingerprints in supplying these algorithms with sufficient data to improve the performance of fingerprint PAD algorithms beyond the capabilities when training on a limited amount of publicly available “real” datasets. First, we provide details of our approach in modifying a state-of-the-art generative architecture to synthesize high quality bona fide and PA fingerprints. Then, we provide quantitative and qualitative analysis to verify the quality of our synthetic fingerprints in mimicking the distribution of real data samples. We showcase the utility of our synthetic bona fide and PA fingerprints in training a deep network for fingerprint PAD, which dramatically boosts the performance across three different evaluation datasets compared to an identical model trained on real data alone. Finally, we demonstrate that only 25% of the original (real) dataset is required to obtain similar detection performance when augmenting the training dataset with synthetic data. We make our synthetic dataset and model publicly available to encourage further research on this topic:https://github.com/groszste/SpoofGAN.
Steven A. Grosz, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2023 Latent Fingerprint Recognition: Fusion of Local and Global Embeddings
abstract
One of the most challenging problems in fingerprint recognition continues to be establishing the identity of a suspect associated with partial and smudgy fingerprints left at a crime scene (i.e., latent prints or fingermarks). Despite the success of fixed-length embeddings for rolled and slap fingerprint recognition, the features learned for latent fingerprint matching have mostly been limited to local minutiae-based embeddings and have not directly leveraged global representations for matching. In this paper, we combine global embeddings with local embeddings for state-of-the-art latent to rolled matching accuracy with high throughput. The combination of both local and global representations leads to improved recognition accuracy across NIST SD 27, NIST SD 302, MSP, MOLF DB1/DB4, and MOLF DB2/DB4 latent fingerprint datasets for both closed-set (84.11%, 54.36%, 84.35%, 70.43%, 62.86% rank-1 retrieval rate, respectively) and open-set (0.50, 0.74, 0.44, 0.60, 0.68 FNIR at FPIR=0.02, respectively) identification scenarios on a gallery of 100K rolled fingerprints. Not only do we fuse the complimentary representations, we also use the local features to guide the global representations to focus on discriminatory regions in two fingerprint images to be compared. This leads to a multi-stage matching paradigm in which subsets of the retrieved candidate lists for each probe image are passed to subsequent stages for further processing, resulting in a considerable reduction in latency (requiring just 0.068 ms per latent to rolled comparison on an AMD EPYC 7543 32-Core Processor, roughly 15K comparisons per second). Finally, we show the generalizability of the fused representations for improving authentication accuracy across several rolled, plain, and contactless fingerprint datasets.
Steven A. Grosz, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2023 Fingerprint Template Invertibility: Minutiae vs. Deep Templates
abstract
Much of the success of fingerprint recognition is attributed to minutiae-based fingerprint representation. It was believed that minutiae templates could not be inverted to obtain a high fidelity fingerprint image, but this assumption has been shown to be false. The success of deep learning has resulted in alternative fingerprint representations (embeddings), in the hope that they might offer better recognition accuracy as well as non-invertibility of deep network-based templates. We evaluate whether deep fingerprint templates suffer from the same reconstruction attacks as the minutiae templates. We show that while a deep template can be inverted to produce a fingerprint image that could be matched to its source image, deep templates are more resistant to reconstruction attacks than minutiae templates. In particular, reconstructed fingerprint images from minutiae templates yield a TAR of about 100.0% (98.3%) @ FAR of 0.01% for type-I (type-II) attacks using a state-of-the-art commercial fingerprint matcher, when tested on NIST SD4. The corresponding attack performance for reconstructed fingerprint images from deep templates using the same commercial matcher yields a TAR of less than 1% for both type-I and type-II attacks; however, when the reconstructed images are matched using the same deep network, they achieve a TAR of 85.95% (68.10%) for type-I (type-II) attacks. Furthermore, what is missing from previous fingerprint template inversion studies is an evaluation of the black-box attack performance, which we perform using 3 different state-of-the-art fingerprint matchers. We conclude that fingerprint images generated by inverting minutiae templates are highly susceptible to both white-box and black-box attack evaluations, while fingerprint images generated by deep templates are resistant to black-box evaluations and comparatively less susceptible to white-box evaluations.
Kanishka P. Wijewardena, Steven A. Grosz, Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2023 Trustworthy AI: A Computational Perspective
abstract
In the past few decades, artificial intelligence (AI) technology has experienced swift developments, changing everyone’s daily life and profoundly altering the course of human society. The intention behind developing AI was and is to benefit humans by reducing labor, increasing everyday conveniences, and promoting social good. However, recent research and AI applications indicate that AI can cause unintentional harm to humans by, for example, making unreliable decisions in safety-critical scenarios or undermining fairness by inadvertently discriminating against a group or groups. Consequently, trustworthy AI has recently garnered increased attention regarding the need to avoid the adverse effects that AI could bring to people, so people can fully trust and live in harmony with AI technologies. A tremendous amount of research on trustworthy AI has been conducted and witnessed in recent years. In this survey, we present a comprehensive appraisal of trustworthy AI from a computational perspective to help readers understand the latest technologies for achieving trustworthy AI. Trustworthy AI is a large and complex subject, involving various dimensions. In this work, we focus on six of the most crucial dimensions in achieving trustworthy AI: (i) Safety & Robustness, (ii) Nondiscrimination & Fairness, (iii) Explainability, (iv) Privacy, (v) Accountability & Auditability, and (vi) Environmental Well-being. For each dimension, we review the recent related technologies according to a taxonomy and summarize their applications in real-world systems. We also discuss the accordant and conflicting interactions among different dimensions and discuss potential aspects for trustworthy AI to investigate in the future.
Yiqi Wang 0001, Wenqi Fan, Yaxin Li 0001, Shaili Jain, Yunhao Liu 0001, Anil K. Jain 0001, Jiliang Tang
ACM Trans. Intell. Syst. Technol.8
2022 AdaFace: Quality Adaptive Margin for Face Recognition
abstract
Recognition in low quality face datasets is challenging because facial attributes are obscured and degraded. Advances in margin-based loss functions have resulted in enhanced discriminability of faces in the embedding space. Further, previous studies have studied the effect of adaptive losses to assign more importance to misclassified (hard) examples. In this work, we introduce another aspect of adaptiveness in the loss function, namely the image quality. We argue that the strategy to emphasize misclassified samples should be adjusted according to their image quality. Specifically, the relative importance of easy or hard samples should be based on the sample's image quality. We propose a new loss function that emphasizes samples of different difficulties based on their image quality. Our method achieves this in the form of an adaptive margin function by approximating the image quality with feature norms. Extensive experiments show that our method, AdaFace, improves the face recognition performance over the state-of-the-art (SoTA) on four datasets (IJB-B, IJB-C, IJB-S and TinyFace). Code and models are released in Supp.
Anil K. Jain 0001, Xiaoming Liu 0002
CVPR2
2022 Multi-domain Learning for Updating Face Anti-spoofing Models
Yaojie Liu, Anil K. Jain 0001, Xiaoming Liu 0002
ECCV (13)3
2022 Controllable and Guided Face Synthesis for Unconstrained Face Recognition
Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002
ECCV (12)3
2022 On Demographic Bias in Fingerprint Recognition
abstract
Fingerprint recognition systems have been deployed globally in numerous applications including personal devices, forensics, law enforcement, banking, and national identity systems. For these systems to be socially acceptable and trustworthy, it is critical that they perform equally well across different demographic groups. In this work, we propose a formal statistical framework to test for the existence of bias (demographic differentials) in fingerprint recognition across four major demographic groups (white male, white female, black male, and black female) for two state-of-the-art (SOTA) fingerprint matchers operating in verification and identification modes. Experiments on two different fingerprint databases (with 15,468 and 1,014 subjects) show that demographic differentials in SOTA fingerprint recognition systems decrease as the matcher accuracy increases and any small bias that may be evident is likely due to certain outlier, low-quality fingerprint images.
Akash Godbole, Steven A. Grosz, Karthik Nandakumar, Anil K. Jain 0001
IJCB4
2022 Robust Unsupervised Domain Adaptation from A Corrupted Source
abstract
Unsupervised Domain Adaptation (UDA) provides a promising solution for learning without supervision, which transfers knowledge from relevant source domains with accessible labeled training data. Existing UDA solutions hinge on clean training data with a short-tail distribution from the source domain, which can be fragile when the source domain data is corrupted either inherently or via adversarial attacks. In this work, we propose an effective framework to address the challenges of UDA from corrupted source domains in a principled manner. Specifically, we perform knowledge ensemble from multiple domain-invariant models that are learned on random partitions of training data. To further address the distribution shift from the source to the target domain, we refine each of the learned models via mutual information maximization, which adaptively obtains the predictive information of the target domain with high confidence. Extensive empirical studies demonstrate that the proposed approach is robust against various types of poisoned data attacks while achieving high asymptotic performance on the target domain.
Shuyang Yu, Zhuangdi Zhu, Anil K. Jain 0001
ICDM4
2022 On Missing Scores in Evolving Multibiometric Systems
abstract
The use of multiple modalities (e.g., face and fingerprint) or multiple algorithms (e.g., three face comparators) has shown to improve the recognition accuracy of an operational biometric system. Over time a biometric system may evolve to add new modalities, retire old modalities, or be merged with other biometric systems. This can lead to scenarios where there are missing scores corresponding to the input probe set. Previous work on this topic has focused on either the verification or identification tasks, but not both. Further, the proportion of missing data considered has been less than 50%. In this work, we study the impact of missing score data for both the verification and identification tasks. We show that the application of various score imputation methods along with simple sum fusion can improve recognition accuracy, even when the proportion of missing scores increases to 90%. Experiments show that fusion after score imputation outperforms fusion with no imputation. Specifically, iterative imputation with K nearest neighbors consistently surpasses other imputation methods in both the verification and identification tasks, regardless of the amount of scores missing, and provides imputed values that are consistent with the ground truth complete dataset.
Melissa R. Dale, Anil K. Jain 0001, Arun Ross
ICPR2
2022 Cluster and Aggregate: Face Recognition with Large Probe Set
abstract
Feature fusion plays a crucial role in unconstrained face recognition where inputs (probes) comprise of a set of $N$ low quality images whose individual qualities vary. Advances in attention and recurrent modules have led to feature fusion that can model the relationship among the images in the input set. However, attention mechanisms cannot scale to large $N$ due to their quadratic complexity and recurrent modules suffer from input order sensitivity. We propose a two-stage feature fusion paradigm, Cluster and Aggregate, that can both scale to large $N$ and maintain the ability to perform sequential inference with order invariance. Specifically, Cluster stage is a linear assignment of $N$ inputs to $M$ global cluster centers, and Aggregation stage is a fusion over $M$ clustered features. The clustered features play an integral role when the inputs are sequential as they can serve as a summarization of past features. By leveraging the order-invariance of incremental averaging operation, we design an update rule that achieves batch-order invariance, which guarantees that the contributions of early image in the sequence do not diminish as time steps increase. Experiments on IJB-B and IJB-S benchmark datasets show the superiority of the proposed two-stage paradigm in unconstrained face recognition.
Feng Liu 0037, Anil K. Jain 0001, Xiaoming Liu 0002
NeurIPS3
2022 Infant-ID: Fingerprints for Global Good
abstract
In many of the least developed and developing countries, a multitude of infants continue to suffer and die from vaccine-preventable diseases and malnutrition. Lamentably, the lack of official identification documentation makes it exceedingly difficult to track which infants have been vaccinated and which infants have received nutritional supplements. Answering these questions could prevent this infant suffering and premature death around the world. To that end, we propose Infant-Prints, an end-to-end, low-cost, infant fingerprint recognition system. Infant-Prints is comprised of our (i) custom built, compact, low-cost (85 USD), high-resolution (1,900 ppi), ergonomic fingerprint reader, and (ii) high-resolution infant fingerprint matcher. To evaluate the efficacy of Infant-Prints, we collected a longitudinal infant fingerprint database captured in 4 different sessions over a 12-month time span (December 2018 to January 2020), from 315 infants at the Saran Ashram Hospital, a charitable hospital in Dayalbagh, Agra, India. Our experimental results demonstrate, for the first time, that Infant-Prints can deliver accurate and reliable recognition (over time) of infants enrolled between the ages of 2-3 months, in time for effective delivery of vaccinations, healthcare, and nutritional supplements (TAR=95.2% @ FAR = 1.0% for infants aged 8-16 weeks at enrollment and authenticated 3 months later).
Joshua J. Engelsma, Debayan Deb, Kai Cao 0001, Anjoo Bhatnagar, Prem Sewak Sudhish, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 C2CL: Contact to Contactless Fingerprint Matching
abstract
Matching contactless fingerprints or finger photos to contact-based fingerprint impressions has received increased attention in the wake of COVID-19 due to the superior hygiene of the contactless acquisition and the widespread availability of low cost mobile phones capable of capturing photos of fingerprints with sufficient resolution for verification purposes. This paper presents an end-to-end automated system, called C2CL, comprised of a mobile finger photo capture app, preprocessing, and matching algorithms to handle the challenges inhibiting previous cross-matching methods; namely i) low ridge-valley contrast of contactless fingerprints, ii) varying roll, pitch, yaw, and distance of the finger to the camera, iii) non-linear distortion of contact-based fingerprints, and vi) different image qualities of smartphone cameras. Our preprocessing algorithm segments, enhances, scales, and unwarps contactless fingerprints, while our matching algorithm extracts both minutiae and texture representations. A sequestered dataset of 9, 888 contactless 2D fingerprints and corresponding contact-based fingerprints from 206 subjects (2 thumbs and 2 index fingers for each subject) acquired using our mobile capture app is used to evaluate the cross-database performance of our proposed algorithm. Furthermore, additional experimental results on 3 publicly available datasets show substantial improvement in the state-of-the-art for contact to contactless fingerprint matching (TAR in the range of 96.67% to 98.30% at FAR=0.01%).
Steven A. Grosz, Joshua J. Engelsma, Eryun Liu, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2021 Mitigating Face Recognition Bias via Group Adaptive Classifier
abstract
Face recognition is known to exhibit bias - subjects in a certain demographic group can be better recognized than other groups. This work aims to learn a fair face representation, where faces of every group could be more equally represented. Our proposed group adaptive classifier mitigates bias by using adaptive convolution kernels and attention mechanisms on faces based on their demographic attributes. The adaptive module comprises kernel masks and channel-wise attention maps for each demographic group so as to activate different facial regions for identification, leading to more discriminative features pertinent to their demographics. Our introduced automated adaptation strategy determines whether to apply adaptation to a certain layer by iteratively computing the dissimilarity among demographic-adaptive parameters. A new de-biasing loss function is proposed to mitigate the gap of average intra-class distance between demographic groups. Experiments on face benchmarks (RFW, LFW, IJB-A, and IJB-C) show that our work is able to mitigate face recognition bias across demographic groups while maintaining the competitive accuracy.
Sixue Gong, Xiaoming Liu 0002, Anil K. Jain 0001
CVPR3
2021 Lifting 2D StyleGAN for 3D-Aware Face Generation
abstract
We propose a framework, called LiftedGAN, that disentangles and lifts a pre-trained StyleGAN2 for 3D-aware face generation. Our model is "3D-aware" in the sense that it is able to (1) disentangle the latent space of StyleGAN2 into texture, shape, viewpoint, lighting and (2) generate 3D components for rendering synthetic images. Unlike most previous methods, our method is completely self-supervised, i.e. it neither requires any manual annotation nor 3DMM model for training. Instead, it learns to generate images as well as their 3D components by distilling the prior knowledge in StyleGAN2 with a differentiable renderer. The proposed model is able to output both the 3D shape and texture, allowing explicit pose and lighting control over generated images. Qualitative and quantitative results show the superiority of our approach over existing methods on 3D-controllable GANs in content controllability while generating realistic high quality images.
Yichun Shi, Divyansh Aggarwal, Anil K. Jain 0001
CVPR3
2021 FedFace: Collaborative Learning of Face Recognition Model
abstract
DNN-based face recognition models require large centrally aggregated face datasets for training. However, due to the growing data privacy concerns and legal restrictions, accessing and sharing face datasets has become exceedingly difficult. We propose FedFace, a federated learning (FL) framework for collaborative learning of face recognition models in a privacy aware manner. FedFace utilizes the face images available on multiple clients to learn an accurate and generalizable face recognition model where the face images stored at each client are neither shared with other clients nor the central host and each client is a mobile device containing face images pertaining to only the owner of the device (one identity per client). Our experiments show the effectiveness of FedFace in enhancing the verification performance of pre-trained face recognition system on standard face verification benchmarks namely LFW, IJB-A and IJB-C.
Divyansh Aggarwal, Anil K. Jain 0001
IJCB3
2021 To be Robust or to be Fair: Towards Fairness in Adversarial Training
abstract
Adversarial training algorithms have been proved to be reliable to improve machine learning models’ robustness against adversarial examples. However, we find that adversarial training algorithms tend to introduce severe disparity of accuracy and robustness between different groups of data. For instance, PGD adversarially trained ResNet18 model on CIFAR-10 has 93% clean accuracy and 67% PGD l_infty-8 adversarial accuracy on the class ”automobile” but only 65% and 17% on class ”cat”. This phenomenon happens in balanced datasets and does not exist in naturally trained models when only using clean samples. In this work, we empirically and theoretically show that this phenomenon can generally happen under adversarial training algorithms which minimize DNN models’ robust errors. Motivated by these findings, we propose a Fair-Robust-Learning (FRL) framework to mitigate this unfairness problem when doing adversarial defenses and experimental results validate the effectiveness of FRL.
Han Xu 0002, Yaxin Li 0001, Anil K. Jain 0001, Jiliang Tang
ICML4
2021 Learning a Fixed-Length Fingerprint Representation
abstract
We present DeepPrint, a deep network, which learns to extract fixed-length fingerprint representations of only 200 bytes. DeepPrint incorporates fingerprint domain knowledge, including alignment and minutiae detection, into the deep network architecture to maximize the discriminative power of its representation. The compact, DeepPrint representation has several advantages over the prevailing variable length minutiae representation which (i) requires computationally expensive graph matching techniques, (ii) is difficult to secure using strong encryption schemes (e.g., homomorphic encryption), and (iii) has low discriminative power in poor quality fingerprints where minutiae extraction is unreliable. We benchmark DeepPrint against two top performing COTS SDKs (Verifinger and Innovatrics) from the NIST and FVC evaluations. Coupled with a re-ranking scheme, the DeepPrint rank-1 search accuracy on the NIST SD4 dataset against a gallery of 1.1 million fingerprints is comparable to the top COTS matcher, but it is significantly faster (DeepPrint: 98.80% in 0.3 seconds vs. COTS A: 98.85% in 27 seconds). To the best of our knowledge, the DeepPrint representation is the most compact and discriminative fixed-length fingerprint representation reported in the academic literature.
Joshua J. Engelsma, Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Learning Continuous Face Age Progression: A Pyramid of GANs
abstract
The two underlying requirements of face age progression, i.e., aging accuracy and identity permanence, are not well studied in the literature. This paper presents a novel generative adversarial network based approach to address the issues in a coupled manner. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while keeping personalized properties stable. To render photo-realistic facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer way. Further, an adversarial learning scheme is introduced to simultaneously train a single generator and multiple parallel discriminators, resulting in smooth continuous face aging sequences. The proposed method is applicable even in the presence of variations in pose, expression, makeup, etc., achieving remarkably vivid aging effects. Quantitative evaluations by a COTS face recognition system demonstrate that the target age distributions are accurately recovered, and 99.88 and 99.98 percent age progressed faces can be correctly verified at 0.001 percent FAR after age transformations of approximately 28 and 23 years elapsed time on the MORPH and CACD databases, respectively. Both visual and quantitative assessments show that the approach advances the state-of-the-art.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Fingerprint Spoof Detector Generalization
abstract
We present a style-transfer based wrapper, called Universal Material Generator (UMG), to improve the generalization performance of any fingerprint spoof (presentation attack) detector against spoofs made from materials not seen during training. Specifically, we transfer the style (texture) characteristics between fingerprint images of known materials with the goal of synthesizing fingerprint images corresponding to unknown materials, that may occupy the space between the known materials in the deep feature space. Synthetic live fingerprint images are also added to the training dataset to supervise the CNN to learn generative-noise invariant features which discriminate between lives and spoofs. The proposed approach is shown to improve the generalization performance of two state-of-the-art spoof detectors, namely Fingerprint Spoof Buster and Slim-ResCNN, winner of the LivDet 2017 spoof detection competition. Specifically, the performance is improved from TDR of 75.24% and 73.09% to TDR of 91.78% and 90.63% @ FDR = 0.2% for Spoof Buster and Slim-ResCNN, respectively. These results are based on a large-scale dataset of 5,743 live and 4,912 spoof images fabricated using 12 different materials. In addition to generalization across different spoof materials, the proposed approach is also shown to improve the average cross-sensor spoof detection performance from 67.60% and 64.62% to 80.63% and 77.59%, for Fingerprint Spoof Buster and Slim-ResCNN, respectively, when tested on the LivDet 2017 dataset.
Tarang Chugh, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2021 Look Locally Infer Globally: A Generalizable Face Anti-Spoofing Approach
abstract
State-of-the-art presentation attack detection approaches tend to overfit to the presentation attack instruments seen during training and fail to generalize to unknown presentation attack instruments. Given that face presentation attack detection is inherently a local task, we propose a face presentation attack detection framework, namely Self-Supervised Regional Fully Convolutional Network (SSR-FCN), that is trained to learn local discriminative cues from a face image in a self-supervised manner. The proposed framework (i) improves generalizability while maintaining the computational efficiency of holistic face presentation attack detection approaches (<; 4 ms on a Nvidia GTX 1080Ti GPU), and (ii) is more interpretable since it localizes the parts of the face that are labeled as presentation attacks. Experimental results show that SSR-FCN can achieve TDR = 65% @ 2.0% FDR when evaluated on a dataset, SiW-M, comprising of 13 different presentation attack instruments under unknown attacks while achieving competitive performances under standard benchmark datasets (Oulu-NPU, CASIA-MFSD, and Replay-Attack).
Debayan Deb, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2020 On the Detection of Digital Face Manipulation
abstract
Detecting manipulated facial images and videos is an increasingly important topic in digital media forensics. As advanced face synthesis and manipulation methods are made available, new types of fake face representations are being created which have raised significant concerns for their use in social media. Hence, it is crucial to detect manipulated face images and localize manipulated regions. Instead of simply using multi-task learning to simultaneously detect manipulated images and predict the manipulated mask (regions), we propose to utilize an attention mechanism to process and improve the feature maps for the classification task. The learned attention maps highlight the informative regions to further improve the binary classification (genuine face v. fake face), and also visualize the manipulated regions. To enable our study of manipulated face detection and localization, we collect a large-scale database that contains numerous types of facial forgeries. With this dataset, we perform a thorough analysis of data-driven fake face detection. We show that the use of an attention mechanism improves facial forgery detection and manipulated region localization.
Hao Dang, Feng Liu 0037, Joel Stehouwer, Xiaoming Liu 0002, Anil K. Jain 0001
CVPR5
2020 Towards Universal Representation Learning for Deep Face Recognition
abstract
Recognizing wild faces is extremely hard as they appear with all kinds of variations. Traditional methods either train with specifically annotated variation data from target domains, or by introducing unlabeled target variation data to adapt from the training data. Instead, we propose a universal representation learning framework that can deal with larger variation unseen in the given training data without leveraging target domain knowledge. We firstly synthesize training data alongside some semantically meaningful variations, such as low resolution, occlusion and head pose. However, directly feeding the augmented data for training will not converge well as the newly introduced samples are mostly hard examples. We propose to split the feature embedding into multiple sub-embeddings, and associate different confidence values for each sub-embedding to smooth the training procedure. The sub-embeddings are further decorrelated by regularizing variation classification loss and variation adversarial loss on different partitions of them. Experiments show that our method achieves top performance on general face recognition datasets such as LFW and MegaFace, while significantly better on extreme benchmarks such as TinyFace and IJB-S.
Yichun Shi, Xiang Yu 0002, Kihyuk Sohn, Manmohan Krishna Chandraker, Anil K. Jain 0001
CVPR5
2020 Jointly De-Biasing Face Recognition and Demographic Attribute Estimation
Sixue Gong, Xiaoming Liu 0002, Anil K. Jain 0001
ECCV (29)3
2020 Fingerprint Spoof Detection: Temporal Analysis of Image Sequence
abstract
We utilize the dynamics involved in the imaging of a fingerprint on a touch-based fingerprint reader, such as perspiration, changes in skin color (blanching), and skin distortion, to differentiate real fingers from spoof (fake) fingers. Specifically, we utilize a deep learning-based architecture (CNN-LSTM) trained end-to-end using sequences of minutiae-centered local patches extracted from ten color frames captured on a COTS fingerprint reader. A time-distributed CNN (MobileNet-v1) extracts spatial features from each local patch, while a bi-directional LSTM layer learns the temporal relationship between the patches in the sequence. Experimental results on a database of 26, 650 live frames from 685 subjects (1,333 unique fingers), and 32,910 spoof frames of 7 spoof materials (with a total of 14 material variants), show that the proposed approach exceeds the state-of-the-art performance in both known-material and cross-material (generalization) scenarios. For instance, the proposed approach improves the state-of-the-art cross-material performance from TDR of 81.65% to 86.20% @ FDR = 0.2%.
Tarang Chugh, Anil K. Jain 0001
IJCB2
2020 AdvFaces: Adversarial Face Synthesis
abstract
Face recognition systems have been shown to be vulnerable to adversarial faces resulting from adding small perturbations to probe images. Such adversarial images can lead state-of-the-art face matchers to falsely reject a genuine subject (obfuscation attack) or falsely match to an impostor (impersonation attack). Current approaches to crafting adversarial faces lack perceptual quality and take an unreasonable amount of time to generate them. We propose, AdvFaces, an automated adversarial face synthesis method that learns to generate minimal perturbations in the salient facial regions via Generative Adversarial Networks. Once AdvFaces is trained, a hacker can automatically generate imperceptible face perturbations that can evade four black-box state-of-the-art face matchers with attack success rates as high as 97.22% and 24.30% at 0.1 % False Accept Rate, for obfuscation and impersonation attacks, respectively.
Debayan Deb, Jianbang Zhang, Anil K. Jain 0001
IJCB3
2020 Fingerprint Presentation Attack Detection: A Sensor and Material Agnostic Approach
abstract
The vulnerability of automated fingerprint recognition systems to presentation attacks (PAs), i.e., spoof or altered fingers, has been a growing concern, warranting the development of accurate and efficient presentation attack detection (PAD) methods. However, one major limitation of the existing PAD solutions is their poor generalization to new PA materials and fingerprint sensors, not used in training. In this study, we propose a robust PAD solution with improved cross-material and cross-sensor generalization. Specifically, we build on top of any CNN-based architecture trained for fingerprint spoof detection combined with cross-material spoof generalization using a style transfer network wrapper. We also incorporate adversarial representation learning (ARL) in deep neural networks (DNN) to learn sensor and material invariant representations for PAD. Experimental results on LivDet 2015 and 2017 public domain datasets exhibit the effectiveness of the proposed approach.
Steven A. Grosz, Tarang Chugh, Anil K. Jain 0001
IJCB3
2020 White-Box Evaluation of Fingerprint Matchers: Robustness to Minutiae Perturbations
abstract
Prevailing evaluations of fingerprint recognition systems have been performed as end-to-end black-box tests of fingerprint identification or authentication accuracy. However, performance of the end-to-end system is subject to errors arising in any of its constituent modules, including: fingerprint scanning, preprocessing, feature extraction, and matching. Conversely, white-box evaluations provide a more granular evaluation by studying the individual subcomponents of a system. While a few studies have conducted stand-alone evaluations of the fingerprint reader and feature extraction modules of fingerprint recognition systems, little work has been devoted towards white-box evaluations of the fingerprint matching module. We report results of a controlled, white-box evaluation of one open-source and two commercial-off-the-shelf (COTS) minutiae-based matchers in terms of their robustness against controlled perturbations (random noise and non-linear distortions) introduced into the input minutiae feature sets. Our white-box evaluations reveal that the performance of fingerprint minutiae matchers are more susceptible to non-linear distortion and missing minutiae than spurious minutiae and small positional displacements of the minutiae locations.
Steven A. Grosz, Joshua J. Engelsma, Nicholas G. Paulter Jr., Anil K. Jain 0001
IJCB4
2020 Fingerprint Synthesis: Search with 100 Million Prints
abstract
Evaluation of large-scale fingerprint search algorithms has been limited due to lack of publicly available datasets. To address this problem, we utilize a Generative Adversarial Network (GAN) to synthesize a fingerprint dataset consisting of 100 million fingerprint images. In contrast to existing fingerprint synthesis algorithms, we incorporate an identity loss which guides the generator to synthesize fingerprints corresponding to more distinct identities. The characteristics of our synthesized fingerprints are shown to be more similar to real fingerprints than existing meth- ods via eight different metrics (minutiae count - block and template, minutiae direction - block and template, minutiae convex hull area, minutiae spatial distribution, block minutiae quality distribution, and NFIQ 2.0 scores). Additionally, the synthetic fingerprints based on our approach are shown to be more distinct than synthetic fingerprints based on published methods through search results and imposter distribution statistics. Finally, we report for the first time in open literature, search accuracy against a gallery of 1 00 million fingerprints (NIST SD4 Rank-1 accuracy of 89.7%).
Vishesh Mistry, Joshua J. Engelsma, Anil K. Jain 0001
IJCB3
2020 Identifying Missing Children: Face Age-Progression via Deep Feature Aging
abstract
Given a face image of a recovered child at age ageprobe, we search a gallery of missing children with known identities and age agegallery at which they were either lost or stolen in an attempt to unite the recovered child with his family. We propose a feature aging module that can age-progress deep face features output by a face matcher to improve the recognition accuracy of age-separated child face images. In addition, the feature aging module guides age-progression in the image space such that synthesized aged gallery faces can be utilized to further enhance cross-age face matching accuracy of any commodity face matcher. For time lapses larger than 10 years (the missing child is recovered after 10 or more years), the proposed age-progression module improves the rank-1 open-set identification accuracy of CosFace from 22.91 % to 25.04% on a child celebrity dataset, namely ITWCC. The proposed method also outperforms state-of-the-art approaches with a rank-1 identification rate of 95.91 %, compared to 94.91 %, on a public aging dataset, FG-NET, and 99.58%, compared to 99.50%, on CACD-VS. These results suggest that aging face features enhances the ability to identify young children who are possible victims of child trafficking or abduction.
Debayan Deb, Divyansh Aggarwal, Anil K. Jain 0001
ICPR3
2020 3D face reconstruction from mugshots: Application to arbitrary view face recognition
Huan Tu, Feng Liu 0037, Qijun Zhao, Anil K. Jain 0001
Neurocomputing5
2020 End-to-End Latent Fingerprint Search
abstract
Latent fingerprints are one of the most important and widely used sources of evidence in law enforcement and forensic agencies. Yet the performance of the state-of-the-art latent recognition systems is far from satisfactory, and they often require manual markups to boost the latent search performance. Further, the COTS systems are proprietary and do not output the true comparison scores between a latent and reference prints to conduct quantitative evidential analysis. We present an end-to-end latent fingerprint search system, including automated region of interest (ROI) cropping, latent image preprocessing, feature extraction, feature comparison, and outputs a candidate list. Two separate minutiae extraction models provide complementary minutiae templates. To compensate for the small number of minutiae in small ridge area and poor quality latents, a virtual minutiae set is generated to construct a texture template. A 96-dimensional descriptor is extracted for each minutia from its neighborhood. For computational efficiency, the descriptor length for virtual minutiae is further reduced to 16 using product quantization. Our end-to-end system is evaluated on four latent databases: NIST SD27 (258 latents); MSP (1200 latents), WVU (449 latents), and N2N (10 000 latents) against a background set of 100K rolled prints, which includes the true rolled mates of the latents with rank-1 retrieval rates of 65.7%, 69.4%, 65.5%, and 7.6%, respectively. A multi-core solution implemented on 24 cores obtains 1-ms per latent to rolled comparison.
Kai Cao 0001, Dinh-Luan Nguyen, Cori Tymoszek, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2019 On the Intrinsic Dimensionality of Image Representations
abstract
This paper addresses the following questions pertaining to the intrinsic dimensionality of any given image representation: (i) estimate its intrinsic dimensionality, (ii) develop a deep neural network based non-linear mapping, dubbed DeepMDS, that transforms the ambient representation to the minimal intrinsic space, and (iii) validate the veracity of the mapping through image matching in the intrinsic space. Experiments on benchmark image datasets (LFW, IJB-C and ImageNet-100) reveal that the intrinsic dimensionality of deep neural network representations is significantly lower than the dimensionality of the ambient features. For instance, SphereFace's 512-dim face representation and ResNet's 512-dim image representation have an intrinsic dimensionality of 16 and 19 respectively. Further, the DeepMDS mapping is able to obtain a representation of significantly lower dimensionality while maintaining discriminative ability to a large extent, 59.75% TAR @ 0.1% FAR in 16-dim vs 71.26% TAR in 512-dim on IJB-C and a Top-1 accuracy of 77.0% at 19-dim vs 83.4% at 512-dim on ImageNet-100.
Sixue Gong, Vishnu Naresh Boddeti, Anil K. Jain 0001
CVPR3
2019 WarpGAN: Automatic Caricature Generation
abstract
We propose, WarpGAN, a fully automatic network that can generate caricatures given an input face photo. Besides transferring rich texture styles, WarpGAN learns to automatically predict a set of control points that can warp the photo into a caricature, while preserving identity. We introduce an identity-preserving adversarial loss that aids the discriminator to distinguish between different subjects. Moreover, WarpGAN allows customization of the generated caricatures by controlling the exaggeration extent and the visual styles. Experimental results on a public domain dataset, WebCaricature, show that WarpGAN is capable of generating caricatures that not only preserve the identities but also outputs a diverse set of caricatures for each input photo. Five caricature experts suggest that caricatures generated by WarpGAN are visually similar to hand-drawn ones and only prominent facial features are exaggerated.
Yichun Shi, Debayan Deb, Anil K. Jain 0001
CVPR3
2019 Probabilistic Face Embeddings
abstract
Embedding methods have achieved success in face recognition by comparing facial features in a latent semantic space. However, in a fully unconstrained face setting, the facial features learned by the embedding model could be ambiguous or may not even be present in the input face, leading to noisy representations. We propose Probabilistic Face Embeddings (PFEs), which represent each face image as a Gaussian distribution in the latent space. The mean of the distribution estimates the most likely feature values while the variance shows the uncertainty in the feature values. Probabilistic solutions can then be naturally derived for matching and fusing PFEs using the uncertainty information. Empirical evaluation on different baseline models, training datasets and benchmarks show that the proposed method can improve the face recognition performance of deterministic embeddings by converting them into PFEs. The uncertainties estimated by PFEs also serve as good indicators of the potential matching accuracy, which are important for a risk-controlled recognition system.
Yichun Shi, Anil K. Jain 0001
ICCV2
2019 Automated Latent Fingerprint Recognition
abstract
Latent fingerprints are one of the most important and widely used evidence in law enforcement and forensic agencies worldwide. Yet, NIST evaluations show that the performance of state-of-the-art latent recognition systems is far from satisfactory. An automated latent fingerprint recognition system with high accuracy is essential to compare latents found at crime scenes to a large collection of reference prints to generate a candidate list of possible mates. In this paper, we propose an automated latent fingerprint recognition algorithm that utilizes Convolutional Neural Networks (ConvNets) for ridge flow estimation and minutiae descriptor extraction, and extract complementary templates (two minutiae templates and one texture template) to represent the latent. The comparison scores between the latent and a reference print based on the three templates are fused to retrieve a short candidate list from the reference database. Experimental results show that the rank-1 identification accuracies (query latent is matched with its true mate in the reference database) are 64.7 percent for the NIST SD27 and 75.3 percent for the WVU latent databases, against a reference database of 100K rolled prints. These results are the best among published papers on latent recognition and competitive with the performance (66.7 and 70.8 percent rank-1 accuracies on NIST SD27 and WVU DB, respectively) of a leading COTS latent Automated Fingerprint Identification System (AFIS). By score-level (rank-level) fusion of our system with the commercial off-the-shelf (COTS) latent AFIS, the overall rank-1 identification performance can be improved from 64.7 and 75.3 to 73.3 percent (74.4 percent) and 76.6 percent (78.4 percent) on NIST SD27 and WVU latent databases, respectively.
Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 RaspiReader: Open Source Fingerprint Reader
abstract
We open source an easy to assemble, spoof resistant, high resolution, optical fingerprint reader, called RaspiReader, using ubiquitous components. By using our open source STL files and software, RaspiReader can be built in under one hour for only US $175. As such, RaspiReader provides the fingerprint research community a seamless and simple method for quickly prototyping new ideas involving fingerprint reader hardware. In particular, we posit that this open source fingerprint reader will facilitate the exploration of novel fingerprint spoof detection techniques involving both hardware and software. We demonstrate one such spoof detection technique by specially customizing RaspiReader with two cameras for fingerprint image acquisition. One camera provides high contrast, frustrated total internal reflection (FTIR) fingerprint images, and the other outputs direct images of the finger in contact with the platen. Using both of these image streams, we extract complementary information which, when fused together and used for spoof detection, results in marked performance improvement over previous methods relying only on grayscale FTIR images provided by COTS optical readers. Finally, fingerprint matching experiments between images acquired from the FTIR output of RaspiReader and images acquired from a COTS reader verify the interoperability of the RaspiReader with existing COTS optical readers.
Joshua J. Engelsma, Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Tattoo Image Search at Scale: Joint Detection and Compact Representation Learning
abstract
The explosive growth of digital images in video surveillance and social media has led to the significant need for efficient search of persons of interest in law enforcement and forensic applications. Despite tremendous progress in primary biometric traits (e.g., face and fingerprint) based person identification, a single biometric trait alone can not meet the desired recognition accuracy in forensic scenarios. Tattoos, as one of the important soft biometric traits, have been found to be valuable for assisting in person identification. However, tattoo search in a large collection of unconstrained images remains a difficult problem, and existing tattoo search methods mainly focus on matching cropped tattoos, which is different from real application scenarios. To close the gap, we propose an efficient tattoo search approach that is able to learn tattoo detection and compact representation jointly in a single convolutional neural network (CNN) via multi-task learning. While the features in the backbone network are shared by both tattoo detection and compact representation learning, individual latent layers of each sub-network optimize the shared features toward the detection and feature learning tasks, respectively. We resolve the small batch size issue inside the joint tattoo detection and compact representation learning network via random image stitch and preceding feature buffering. We evaluate the proposed tattoo search system using multiple public-domain tattoo benchmarks, and a gallery set with about 300K distracter tattoo images compiled from these datasets and images from the Internet. In addition, we also introduce a tattoo sketch dataset containing 300 tattoos for sketch-based tattoo search. Experimental results show that the proposed approach has superior performance in tattoo detection and tattoo search at scale compared to several state-of-the-art tattoo retrieval algorithms.
Hu Han 0001, Anil K. Jain 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 On the Reconstruction of Face Images from Deep Face Templates
abstract
State-of-the-art face recognition systems are based on deep (convolutional) neural networks. Therefore, it is imperative to determine to what extent face templates derived from deep networks can be inverted to obtain the original face image. In this paper, we study the vulnerabilities of a state-of-the-art face recognition system based on template reconstruction attack. We propose a neighborly de-convolutional neural network (NbNet) to reconstruct face images from their deep templates. In our experiments, we assumed that no knowledge about the target subject and the deep network are available. To train the NbNet reconstruction models, we augmented two benchmark face datasets (VGG-Face and Multi-PIE) with a large collection of images synthesized using a face generator. The proposed reconstruction was evaluated using type-I (comparing the reconstructed images against the original face images used to generate the deep template) and type-II (comparing the reconstructed images against a different face image of the same subject) attacks. Given the images reconstructed from NbNets, we show that for verification, we achieve TAR of 95.20 percent (58.05 percent) on LFW under type-I (type-II) attacks @ FAR of 0.1 percent. Besides, 96.58 percent (92.84 percent) of the images reconstructed from templates of partition fa (fb) can be identified from partition fa in color FERET. Our study demonstrates the need to secure deep templates in face recognition systems.
Guangcan Mai, Kai Cao 0001, Pong C. Yuen, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Learning Face Age Progression: A Pyramid Architecture of GANs
abstract
The two underlying requirements of face age progression, i.e. aging accuracy and identity permanence, are not well studied in the literature. In this paper, we present a novel generative adversarial network based approach. It separately models the constraints for the intrinsic subject-specific characteristics and the age-specific facial changes with respect to the elapsed time, ensuring that the generated faces present desired aging effects while simultaneously keeping personalized properties stable. Further, to generate more lifelike facial details, high-level age-specific features conveyed by the synthesized face are estimated by a pyramidal adversarial discriminator at multiple scales, which simulates the aging effects in a finer manner. The proposed method is applicable to diverse face samples in the presence of variations in pose, expression, makeup, etc., and remarkably vivid aging effects are achieved. Both visual fidelity and quantitative evaluations show that the approach advances the state-of-the-art.
Hongyu Yang 0001, Di Huang 0001, Yunhong Wang 0001, Anil K. Jain 0001
CVPR4
2018 Heterogeneous Hyper-Network Embedding
abstract
Heterogeneous hyper-networks is used to represent multi-modal and composite interactions between data points. In such networks, several different types of nodes form a hyperedge. Heterogeneous hyper-network embedding learns a distributed node representation under such complex interactions while preserving the network structure. However, this is a challenging task due to the multiple modalities and composite interactions. In this study, a deep approach is proposed to embed heterogeneous attributed hyper-networks with complicated and non-linear node relationships. In particular, a fully-connected and graph convolutional layers are designed to project different types of nodes into a common low-dimensional space, a tuple-wise similarity function is proposed to preserve the network structure, and a ranking based loss function is used to improve the similarity scores of hyperedges in the embedding space. The proposed approach is evaluated on synthetic and real world datasets and a better performance is obtained compared with baselines.
Inci M. Baytas, Cao Xiao, Fei Wang 0001, Anil K. Jain 0001
ICDM4
2018 On Mugshot-based Arbitrary View Face Recognition
abstract
Despite the wide usage of mugshot images in forensic applications, they are underutilized in existing automated face recognition systems. In this paper, we propose a novel mugshot-based arbitrary view face recognition method. Our approach reconstructs full 3D faces via cascaded regression in shape space with efficient seamless texture recovery. Unlike existing methods, it makes full use of the frontal and profile views available in mugshot images, and thus generates accurate and realistic 3D faces. Multi-view face images are synthesized from the reconstructed 3D faces to enlarge the gallery so that arbitrary view faces can be better recognized. Evaluation experiments were conducted on BFM and Multi-PIE databases by using state-of-the-art deep learning (DL) based face matchers. The results demonstrate the effectiveness of our proposed method and show that DL-based face matchers can benefit from mugshot images and the reconstructed 3D faces, especially for recognizing large off-angle faces.
Feng Liu 0037, Huan Tu, Qijun Zhao, Anil K. Jain 0001
ICPR5
2018 Longitudinal Study of Automatic Face Recognition
abstract
The two underlying premises of automatic face recognition are uniqueness and permanence. This paper investigates the permanence property by addressing the following: Does face recognition ability of state-of-the-art systems degrade with elapsed time between enrolled and query face images? If so, what is the rate of decline w.r.t. the elapsed time? While previous studies have reported degradations in accuracy, no formal statistical analysis of large-scale longitudinal data has been conducted. We conduct such an analysis on two mugshot databases, which are the largest facial aging databases studied to date in terms of number of subjects, images per subject, and elapsed times. Mixed-effects regression models are applied to genuine similarity scores from state-of-the-art COTS face matchers to quantify the population-mean rate of change in genuine scores over time, subject-specific variability, and the influence of age, sex, race, and face image quality. Longitudinal analysis shows that despite decreasing genuine scores, 99% of subjects can still be recognized at 0.01% FAR up to approximately 6 years elapsed time, and that age, sex, and race only marginally influence these trends. The methodology presented here should be periodically repeated to determine age-invariant properties of face recognition as state-of-the-art evolves to better address facial aging.
Lacey Best-Rowden, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Heterogeneous Face Attribute Estimation: A Deep Multi-Task Learning Approach
abstract
Face attribute estimation has many potential applications in video surveillance, face retrieval, and social media. While a number of methods have been proposed for face attribute estimation, most of them did not explicitly consider the attribute correlation and heterogeneity (e.g., ordinal versus nominal and holistic versus local) during feature representation learning. In this paper, we present a Deep Multi-Task Learning (DMTL) approach to jointly estimate multiple heterogeneous attributes from a single face image. In DMTL, we tackle attribute correlation and heterogeneity with convolutional neural networks (CNNs) consisting of shared feature learning for all the attributes, and category-specific feature learning for heterogeneous attributes. We also introduce an unconstrained face database (LFW+), an extension of public-domain LFW, with heterogeneous demographic attributes (age, gender, and race) obtained via crowdsourcing. Experimental results on benchmarks with multiple face attributes (MORPH II, LFW+, CelebA, LFWA, and FotW) show that the proposed approach has superior performance compared to state of the art. Finally, evaluations on a public-domain face database (LAP) with a single attribute show that the proposed approach has excellent generalization ability.
Hu Han 0001, Anil K. Jain 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Clustering Millions of Faces by Identity
abstract
Given a large collection of unlabeled face images, we address the problem of clustering faces into an unknown number of identities. This problem is of interest in social media, law enforcement, and other applications, where the number of faces can be of the order of hundreds of million, while the number of identities (clusters) can range from a few thousand to millions. To address the challenges of run-time complexity and cluster quality, we present an approximate Rank-Order clustering algorithm that performs better than popular clustering algorithms (k-Means and Spectral). Our experiments include clustering up to 123 million face images into over 10 million clusters. Clustering results are analyzed in terms of external (known face labels) and internal (unknown face labels) quality measures, and run-time. Our algorithm achieves an F-measure of 0.87 on the LFW benchmark (13 K faces of 5,749 individuals), which drops to 0.27 on the largest dataset considered (13 K faces in LFW + 123M distractor images). Additionally, we show that frames in the YouTube benchmark can be clustered with an F-measure of 0.71. An internal per-cluster quality measure is developed to rank individual clusters for manual exploration of high quality clusters that are compact and isolated.
Charles Otto, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Learning Face Image Quality From Human Assessments
abstract
Face image quality can be defined as a measure of the utility of a face image to automatic face recognition. In this paper, we propose (and compare) two methods for learning face image quality based on target face quality values from: 1) human assessments of face image quality (matcher-independent) and 2) quality values computed from similarity scores (matcher-dependent). A support vector regression model trained on face features extracted using a deep convolutional neural network (ConvNet) is used to predict the quality of a face image. The proposed methods are evaluated on two unconstrained face image databases, Labeled Faces in the Wild and IARPA Janus Benchmark-A (IJB-A), which both contain facial variations encompassing a multitude of quality factors. Evaluation of the proposed automatic face image quality measures shows we are able to reduce the false non-match rate at 1% false match rate by at least 13% for two face matchers (a commercial off-the-shelf matcher and a ConvNet matcher) by using the proposed face quality to select subsets of face images and video frames for matching templates (i.e., multiple faces per subject) in the IJB-A protocol. To the best of our knowledge, this is the first work to utilize human assessments of face image quality in designing a predictor of unconstrained face quality that is shown to be effective in cross-database evaluation.
Lacey Best-Rowden, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2018 Fingerprint Spoof Buster: Use of Minutiae-Centered Patches
abstract
The primary purpose of a fingerprint recognition system is to ensure a reliable and accurate user authentication, but the security of the recognition system itself can be jeopardized by spoof attacks. This paper addresses the problem of developing accurate, generalizable, and efficient algorithms for detecting fingerprint spoof attacks. Specifically, we propose a deep convolutional neural network-based approach utilizing local patches centered and aligned using fingerprint minutiae. Experimental results on three public-domain LivDet datasets (2011, 2013, and 2015) show that the proposed approach provides the state-of-the-art accuracies in fingerprint spoof detection for intra-sensor, cross-material, cross-sensor, as well as cross-dataset testing scenarios. For example, in LivDet 2015, the proposed approach achieves 99.03% average accuracy over all sensors compared with 95.51% achieved by the LivDet 2015 competition winners. In addition, two new fingerprint presentation attack datasets containing more than 20,000 images, using two different fingerprint readers, and over 12 different spoof fabrication materials are collected. We also present a graphical user interface, called Fingerprint Spoof Buster, that allows the operator to visually examine the local regions of the fingerprint highlighted as live or spoof, instead of relying on only a single score as output by the traditional approaches.
Tarang Chugh, Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2018 Latent Fingerprint Value Prediction: Crowd-Based Learning
abstract
Latent fingerprints are one of the most crucial sources of evidence in forensic investigations. As such, development of automatic latent fingerprint recognition systems to quickly and accurately identify the suspects is one of the most pressing problems facing fingerprint researchers. One of the first steps in manual latent processing is for a fingerprint examiner to perform a triage by assigning one of the following three values to a query latent: Value for Individualization (VID), Value for Exclusion Only (VEO), or No Value (NV). However, latent value determination by examiners is known to be subjective, resulting in large intra-examiner and inter-examiner variations. Furthermore, in spite of the guidelines available, the underlying bases that examiners implicitly use for value determination are unknown. In this paper, we propose a crowdsourcing based framework for understanding the underlying bases of value assignment by fingerprint examiners, and use it to learn a predictor for quantitative latent value assignment. Experimental results are reported using four latent fingerprint databases, two from forensic casework (NIST SD27 and MSP) and two collected in laboratory settings (WVU and IIITD), and a state-of-the-art latent automated fingerprint identification system (AFIS). The main conclusions of this paper are as follows: 1) crowdsourced latent value is more robust than prevailing value determination (VID, VEO, and NV) and latent fingerprint image quality for predicting AFIS performance; 2) two bases can explain expert value assignments, which can be interpreted in terms of latent features; and 3) our value predictor can rank a collection of latents from most informative to least informative.
Tarang Chugh, Kai Cao 0001, Elham Tabassi, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.5
2018 Universal 3D Wearable Fingerprint Targets: Advancing Fingerprint Reader Evaluations
abstract
We present the design and manufacturing of high-fidelity universal 3D fingerprint targets, which can be imaged on a variety of fingerprint sensing technologies, namely, capacitive, contact optical, and contactless optical. Universal 3D fingerprint targets enable, for the first time, not only a repeatable and controlled evaluation of fingerprint readers but also the ability to conduct fingerprint reader interoperability studies. Fingerprint reader interoperability refers to how robust fingerprint recognition systems are to variations in the images acquired by different types of fingerprint readers. To build universal 3D fingerprint targets, we adopt a molding and casting framework consisting of: 1) digital mapping of fingerprint images to a negative mold; 2) CAD modeling a scaffolding system to hold the negative mold; 3) fabricating the mold and scaffolding system with a high resolution 3D printer; 4) producing or mixing a material with similar electrical, optical, and mechanical properties to that of the human finger; and 5) fabricating a 3D fingerprint target using controlled casting. Our experiments conducted with personal identity verification and Appendix F certified optical (contact and contactless) and capacitive fingerprint readers demonstrate the usefulness of universal 3D fingerprint targets for controlled and repeatable fingerprint reader evaluations and also fingerprint reader interoperability studies.
Joshua J. Engelsma, Sunpreet S. Arora, Anil K. Jain 0001, Nicholas G. Paulter Jr.
IEEE Trans. Inf. Forensics Secur.3
2018 Face Clustering: Representation and Pairwise Constraints
abstract
Clustering face images according to their latent identity has two important applications: 1) grouping a collection of face images when no external labels are associated with images, and 2) indexing for efficient large scale face retrieval. The clustering problem is composed of two key parts: representation and similarity metric for face images, and choice of the partition algorithm. We first propose a representation based on ResNet, which has been shown to perform very well in image classification problems. Given this representation, we design a clustering algorithm, Conditional Pairwise Clustering (ConPaC), which directly estimates the adjacency matrix only based on the similarities between face images. This allows a dynamic selection of number of clusters and retains pairwise similarities between faces. ConPaC formulates the clustering problem as a Conditional Random Field model and uses Loopy Belief Propagation to find an approximate solution for maximizing the posterior probability of the adjacency matrix. Experimental results on two benchmark face datasets (LFW and IJB-B) show that ConPaC outperforms well known clustering algorithms such as k-means, spectral clustering, and approximate Rank-order. Additionally, our algorithm can naturally incorporate pairwise constraints to work in a semi-supervised way that leads to improved clustering performance. We also propose a k-NN variant of ConPaC, which has a linear time complexity given a k-NN graph, suitable for large datasets.
Yichun Shi, Charles Otto, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2017 Fingerprint indexing and matching: An integrated approach
abstract
Large scale fingerprint recognition systems have been deployed worldwide not only in law enforcement but also in many civilian applications. Thus, it is of great value o identify a query fingerprint in a large background finger-print database both effectively and efficiently based on indexing strategies. The published indexing algorithms do not meet the requirements, especially at low penetrate rates, because of the difficulty in extracting reliable minutiae and other features in low quality fingerprint images. We propose a Convolutional Neural Network (ConvNet) based fingerprint indexing algorithm. An orientation field dictionary is learned to align fingerprints in a unified coordinate system and a large longitudinal fingerprint database, where each finger has multiple impressions over time, is used to train the ConvNet. Experimental results on NIST SD4 and NIST SD14 show that the proposed approach outperforms state-of-the-art fingerprint indexing techniques reported in the literature. Further indexing results on an augmented gallery set of 250K rolled prints demonstrate the scalability of the proposed algorithm. At a penetrate rate of 1%, a score-level fusion of the proposed indexing and a state-of-the-art COTS SDK provides 97.8% rank-1 identification accuracy with a 100-fold reduction in the search space.
Kai Cao 0001, Anil K. Jain 0001
IJCB2
2017 Fingerprint spoof detection using minutiae-based local patches
abstract
The individuality of fingerprints is being leveraged for a plethora of day-to-day applications, ranging from unlocking a smartphone to international border security. While the primary purpose of a fingerprint recognition system is to ensure a reliable and accurate user authentication, the security of the recognition system itself can be jeopardized by spoof attacks. This study addresses the problem of developing accurate and generalizable algorithms for detecting fingerprint spoof attacks. We propose a deep convolutional neural network based approach utilizing local patches extracted around fingerprint minutiae. Experimental results on three public-domain LivDet datasets (2011, 2013, and 2015) show that the proposed approach provides state of the art accuracies in fingerprint spoof detection for intra-sensor, cross-material, cross-sensor, as well as cross-dataset testing scenarios. For example, the proposed approach achieves a 69% reduction in average classification error for spoof detection under both known material and cross-material scenarios on LivDet 2015 datasets.
Tarang Chugh, Kai Cao 0001, Anil K. Jain 0001
IJCB3
2017 Patient Subtyping via Time-Aware LSTM Networks
abstract
In the study of various diseases, heterogeneity among patients usually leads to different progression patterns and may require different types of therapeutic intervention. Therefore, it is important to study patient subtyping, which is grouping of patients into disease characterizing subtypes. Subtyping from complex patient data is challenging because of the information heterogeneity and temporal dynamics. Long-Short Term Memory (LSTM) has been successfully used in many domains for processing sequential data, and recently applied for analyzing longitudinal patient records. The LSTM units are designed to handle data with constant elapsed times between consecutive elements of a sequence. Given that time lapse between successive elements in patient records can vary from days to months, the design of traditional LSTM may lead to suboptimal performance. In this paper, we propose a novel LSTM unit called Time-Aware LSTM (T-LSTM) to handle irregular time intervals in longitudinal patient records. We learn a subspace decomposition of the cell memory which enables time decay to discount the memory content according to the elapsed time. We propose a patient subtyping model that leverages the proposed T-LSTM in an auto-encoder to learn a powerful single representation for sequential records of patients, which are then used to cluster patients into clinical subtypes. Experiments on synthetic and real world datasets show that the proposed T-LSTM architecture captures the underlying structures in the sequences with time irregularities.
Inci M. Baytas, Cao Xiao, Fei Wang 0001, Anil K. Jain 0001
KDD5
2017 Face Search at Scale
abstract
Given the prevalence of social media websites, one challenge facing computer vision researchers is to devise methods to search for persons of interest among the billions of shared photos on these websites. Despite significant progress in face recognition, searching a large collection of unconstrained face images remains a difficult problem. To address this challenge, we propose a face search system which combines a fast search procedure, coupled with a state-of-the-art commercial off the shelf (COTS) matcher, in a cascaded framework. Given a probe face, we first filter the large gallery of photos to find the top- k most similar faces using features learned by a convolutional neural network. The k retrieved candidates are re-ranked by combining similarities based on deep features and those output by the COTS matcher. We evaluate the proposed face search system on a gallery containing 80 million web-downloaded face images. Experimental results demonstrate that while the deep features perform worse than the COTS matcher on a mugshot dataset (93.7 percent versus 98.6 percent TAR@FAR of 0.01 percent), fusing the deep features with the COTS matcher improves the overall performance ( 99.5 percent TAR@FAR of 0.01 percent). This shows that the learned deep features provide complementary information over representations used in state-of-the-art face matchers. On the unconstrained face image benchmarks, the performance of the learned deep features is competitive with reported accuracies. LFW database: 98.20 percent accuracy under the standard protocol and 88.03 percent TAR@FAR of 0.1 percent under the BLUFR protocol; IJB-A benchmark: 51.0 percent TAR@FAR of 0.1 percent (verification), rank 1 retrieval of 82.2 percent (closed-set search), 61.5 percent FNIR@FAR of 1 percent (open-set search). The proposed face search system offers an excellent trade-off between accuracy and scalability on galleries with millions of images. Additionally, in a face search experiment involving photos of the Tsarnaev brothers, convicted of the Boston Marathon bombing, the proposed cascade face search system could find the younger brother's (Dzhokhar Tsarnaev) photo at rank 1 in 1 second on a 5 M gallery and at rank 8 in 7 seconds on an 80 M gallery.
Charles Otto, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Editorial: Special issue on ubiquitous biometrics
Ran He 0001, Brian C. Lovell, Rama Chellappa, Anil K. Jain 0001, Zhenan Sun
Pattern Recognit.4
2017 Gold Fingers: 3D Targets for Evaluating Capacitive Readers
abstract
With capacitive fingerprint readers being increasingly used for access control as well as for smartphone unlock and payments, there is a growing interest among metrology agencies (e.g., the National Institute of Standards and Technology) to develop standard artifacts (targets) and procedures for repeatable evaluation of capacitive readers. We present our design and fabrication procedures to create conductive 3D targets (gold fingers) for capacitive readers. Wearable 3D targets with known feature markings (e.g., fingerprint ridge flow and ridge spacing) are first fabricated using a high-resolution 3D printer. A sputter coating process is subsequently used to deposit a thin layer (~300 nm) of conductive materials (titanium and gold) on 3D printed targets. The wearable gold finger targets are used to evaluate a PIV-certified single-finger capacitive reader as well as small-area capacitive readers embedded in smartphones and access control terminals. In additional, we show that a simple procedure to create 3D printed spoofs with conductive carbon coating is able to successfully spoof a PIV-certified single-finger capacitive reader as well as a capacitive reader embedded in an access control terminal.
Sunpreet S. Arora, Anil K. Jain 0001, Nicholas G. Paulter Jr.
IEEE Trans. Inf. Forensics Secur.2
2017 Fingerprint Recognition of Young Children
abstract
In 1899, Galton first captured ink-on-paper fingerprints of a single child from birth until the age of 4.5 years, manually compared the prints, and concluded that “the print of a child at the age of 2.5 years would serve to identify him ever after.” Since then, ink-on-paper fingerprinting and manual comparison methods have been superseded by digital capture and automatic fingerprint comparison techniques, but only a few feasibility studies on child fingerprint recognition have been conducted. Here, we present the first systematic and rigorous longitudinal study that addresses the following questions: (1) Do fingerprints of young children possess the salient features required to uniquely recognize a child? (2) If so, at what age can a child's fingerprints be captured with sufficient fidelity for recognition? (3) Can a child's fingerprints be used to reliably recognize the child as he ages? For this paper, we collected fingerprints of 309 children (0-5 years old) four different times over a one year period. We show, for the first time, that fingerprints acquired from a child as young as 6-h old exhibit distinguishing features necessary for recognition, and that state-of-the-art fingerprint technology achieves high recognition accuracy (98.9% true accept rate at 0.1% false accept rate) for children older than six months. In addition, we use mixed-effects statistical models to study the persistence of child fingerprint recognition accuracy and show that the recognition accuracy is not significantly affected over the one year time lapse in our data. Given rapidly growing requirements to recognize children for vaccination tracking, delivery of supplementary food, and national identification documents, this paper demonstrates that fingerprint recognition of young children (six months and older) is a viable solution based on available capture and recognition technology.
Anil K. Jain 0001, Sunpreet S. Arora, Kai Cao 0001, Lacey Best-Rowden, Anjoo Bhatnagar
IEEE Trans. Inf. Forensics Secur.1
2016 Asynchronous Multi-task Learning
abstract
Many real-world machine learning applications involve several learning tasks which are inter-related. For example, in healthcare domain, we need to learn a predictive model of a certain disease for many hospitals. The models for each hospital may be different because of the inherent differences in the distributions of the patient populations. However, the models are also closely related because of the nature of the learning tasks modeling the same disease. By simultaneously learning all the tasks, multi-task learning (MTL) paradigm performs inductive knowledge transfer among tasks to improve the generalization performance. When datasets for the learning tasks are stored at different locations, it may not always be feasible to transfer the data to provide a data centralized computing environment due to various practical issues such as high data volume and privacy. In this paper, we propose a principled MTL framework for distributed and asynchronous optimization to address the aforementioned challenges. In our framework, gradient update does not wait for collecting the gradient information from all the tasks. Therefore, the proposed method is very efficient when the communication delay is too high for some task nodes. We show that many regularized MTL formulations can benefit from this framework, including the low-rank MTL for shared subspace learning. Empirical studies on both synthetic and real-world datasets demonstrate the efficiency and effectiveness of the proposed framework.
Inci M. Baytas, Anil K. Jain 0001
ICDM3
2016 Giving Infants an Identity: Fingerprint Sensing and Recognition
abstract
There is a growing demand for biometrics-based recognition of children for a number of applications, particularly in developing countries where children do not have any form of identification. These applications include tracking child vaccination schedules, identifying missing children, preventing fraud in food subsidies, and preventing newborn baby swaps in hospitals. Our objective is to develop a fingerprint-based identification system for infants (age range: 0-12 months)1. Our ongoing research has addressed the following issues: (i) design of a compact, comfortable, high-resolution (>1,000 ppi) fingerprint reader; (ii) image enhancement algorithms to improve quality of infant fingerprint images; and (iii) collection of longitudinal infant fingerprint data to evaluate identification accuracy over time. This collaboration between Michigan State University, Dayalbagh Educational Institute, Saran Ashram Hospital, Agra, India and NEC Corporation, has demonstrated the feasibility of recognizing infants older than 4 weeks using fingerprints.
Anil K. Jain 0001, Sunpreet S. Arora, Lacey Best-Rowden, Kai Cao 0001, Prem Sewak Sudhish, Anjoo Bhatnagar, Yoshinori Koda
ICTD1
2016 A Fast and Accurate Unconstrained Face Detector
abstract
We propose a method to address challenges in unconstrained face detection, such as arbitrary pose variations and occlusions. First, a new image feature called Normalized Pixel Difference (NPD) is proposed. NPD feature is computed as the difference to sum ratio between two pixel values, inspired by the Weber Fraction in experimental psychology. The new feature is scale invariant, bounded, and is able to reconstruct the original image. Second, we propose a deep quadratic tree to learn the optimal subset of NPD features and their combinations, so that complex face manifolds can be partitioned by the learned rules. This way, only a single soft-cascade classifier is needed to handle unconstrained face detection. Furthermore, we show that the NPD features can be efficiently obtained from a look up table, and the detection template can be easily scaled, making the proposed face detector very fast. Experimental results on three public face datasets (FDDB, GENKI, and CMU-MIT) show that the proposed method achieves state-of-the-art performance in detecting unconstrained faces with arbitrary pose variations and occlusions in cluttered scenes.
Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2016 50 years of biometric research: Accomplishments, challenges, and opportunities
Anil K. Jain 0001, Karthik Nandakumar, Arun Ross
Pattern Recognit. Lett.1
2016 Adaptive fusion of biometric and biographic information for identity de-duplication
Prem Sewak Sudhish, Anil K. Jain 0001, Kai Cao 0001
Pattern Recognit. Lett.2
2016 Design and Fabrication of 3D Fingerprint Targets
abstract
Standard targets are typically used for structural (white-box) evaluation of fingerprint readers, e.g., for calibrating imaging components of a reader. However, there is no standard method for behavioral (black-box) evaluation of fingerprint readers in operational settings where variations in finger placement by the user are encountered. The goal of this research is to design and fabricate 3D targets for repeatable behavioral evaluation of fingerprint readers. 2D calibration patterns with known characteristics (e.g., sinusoidal gratings of pre-specified orientation and frequency, and fingerprints with known singular points and minutiae) are projected onto a generic 3D finger surface to create electronic 3D targets. A state-of-the-art 3D printer (Stratasys Objet350 Connex) is used to fabricate wearable 3D targets with materials similar in hardness and elasticity to the human finger skin. The 3D printed targets are cleaned using 2M NaOH solution to obtain evaluation-ready 3D targets. Our experimental results show that: 1) features present in the 2D calibration pattern are preserved during the creation of the electronic 3D target; 2) features engraved on the electronic 3D target are preserved during the physical 3D target fabrication; and 3) intra-class variability between multiple impressions of the physical 3D target is small. We also demonstrate that the generated 3D targets are suitable for behavioral evaluation of three different (500/1000 ppi) PIV/Appendix F certified optical fingerprint readers in the operational settings.
Sunpreet S. Arora, Kai Cao 0001, Anil K. Jain 0001, Nicholas G. Paulter Jr.
IEEE Trans. Inf. Forensics Secur.3
2016 Secure Face Unlock: Spoof Detection on Smartphones
abstract
With the wide deployment of the face recognition systems in applications from deduplication to mobile device unlocking, security against the face spoofing attacks requires increased attention; such attacks can be easily launched via printed photos, video replays, and 3D masks of a face. We address the problem of face spoof detection against the print (photo) and replay (photo or video) attacks based on the analysis of image distortion (e.g., surface reflection, moiré pattern, color distortion, and shape deformation) in spoof face images (or video frames). The application domain of interest is smartphone unlock, given that the growing number of smartphones have the face unlock and mobile payment capabilities. We build an unconstrained smartphone spoof attack database (MSU USSA) containing more than 1000 subjects. Both the print and replay attacks are captured using the front and rear cameras of a Nexus 5 smartphone. We analyze the image distortion of the print and replay attacks using different: 1) intensity channels (R, G, B, and grayscale); 2) image regions (entire image, detected face, and facial component between nose and chin); and 3) feature descriptors. We develop an efficient face spoof detection system on an Android smartphone. Experimental results on the public-domain Idiap Replay-Attack, CASIA FASD, and MSU-MFSD databases, and the MSU USSA database show that the proposed approach is effective in face spoof detection for both the cross-database and intra-database testing scenarios. User studies of our Android face spoof detection system involving 20 participants show that the proposed approach works very well in real application scenarios.
Keyurkumar Patel, Hu Han 0001, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2016 PhenoTree: Interactive Visual Analytics for Hierarchical Phenotyping From Large-Scale Electronic Health Records
abstract
Electronic health records (EHRs) capture comprehensive patient information in digital form from a variety of sources. Increasing availability of EHRs has facilitated development of data and visual analytic tools for healthcare analytics, such as clinical decision support and patient care management systems. Many healthcare analytic tools are used to investigate fundamental problems, such as study of patient population, exploring complicated interactions among patients and their medical histories, and extracting structured phenotypes characterizing the patient population. In this paper, we propose PHENOTREE, a novel data-driven, hierarchical, and interactive phenotyping tool, that enables physicians and medical researchers to participate in the phenotyping process of large-scale EHR cohorts. The proposed visual analytic tool allows users to interactively explore EHR cohorts, and generate, interpret, evaluate, and refine phenotypes by building and navigating a phenotype hierarchy. Specifically, given a cohort or subcohort, PHENOTREE employs sparse principal component analysis (SPCA) to identify key clinical features that characterize the population. The clinical features provide a natural way to generate deeper phenotypes at finer granularities by expanding the phenotype hierarchy. To facilitate the intensive computation required for interactive analytics, we design an efficient SPCA solver based on a variance reduced stochastic gradient technique. The benefits of our method are demonstrated by analyzing two different EHR patient cohorts, a public and a private dataset containing EHRs of 101 767 and 223 076 patients, respectively. Our evaluations show that PHENOTREE can detect clinically meaningful hierarchical phenotypes.
Inci M. Baytas, Kaixiang Lin, Fei Wang 0001, Anil K. Jain 0001
IEEE Trans. Multim.4
2015 Pushing the frontiers of unconstrained face detection and recognition: IARPA Janus Benchmark A
abstract
Rapid progress in unconstrained face recognition has resulted in a saturation in recognition accuracy for current benchmark datasets. While important for early progress, a chief limitation in most benchmark datasets is the use of a commodity face detector to select face imagery. The implication of this strategy is restricted variations in face pose and other confounding factors. This paper introduces the IARPA Janus Benchmark A (IJB-A), a publicly available media in the wild dataset containing 500 subjects with manually localized face images. Key features of the IJB-A dataset are: (i) full pose variation, (ii) joint use for face recognition and face detection benchmarking, (iii) a mix of images and videos, (iv) wider geographic variation of subjects, (v) protocols supporting both open-set identification (1:N search) and verification (1:1 comparison), (vi) an optional protocol that allows modeling of gallery subjects, and (vii) ground truth eye and nose locations. The dataset has been developed using 1,501,267 million crowd sourced annotations. Baseline accuracies for both face detection and face recognition from commercial and open source algorithms demonstrate the challenge offered by this new unconstrained benchmark.
Brendan Klare, Benjamin Klein, Emma Taborsky, Austin Blanton, Jordan Cheney, Kristen Allen, Patrick Grother, Alan Mah, Mark James Burge, Anil K. Jain 0001
CVPR10
2015 Demographic Estimation from Face Images: Human vs. Machine Performance
abstract
Demographic estimation entails automatic estimation of age, gender and race of a person from his face image, which has many potential applications ranging from forensics to social media. Automatic demographic estimation, particularly age estimation, remains a challenging problem because persons belonging to the same demographic group can be vastly different in their facial appearances due to intrinsic and extrinsic factors. In this paper, we present a generic framework for automatic demographic (age, gender and race) estimation. Given a face image, we first extract demographic informative features via a boosting algorithm, and then employ a hierarchical approach consisting of between-group classification, and within-group regression. Quality assessment is also developed to identify low-quality face images that are difficult to obtain reliable demographic estimates. Experimental results on a diverse set of face image databases, FG-NET (1K images), FERET (3K images), MORPH II (75K images), PCSO (100K images), and a subset of LFW (4K images), show that the proposed approach has superior performance compared to the state of the art. Finally, we use crowdsourcing to study the human perception ability of estimating demographics from face images. A side-by-side comparison of the demographic estimates from crowdsourced data and the proposed algorithm provides a number of insights into this challenging problem.
Hu Han 0001, Charles Otto, Xiaoming Liu 0002, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2015 Learning Fingerprint Reconstruction: From Minutiae to Image
abstract
The set of minutia points is considered to be the most distinctive feature for fingerprint representation and is widely used in fingerprint matching. It was believed that the minutiae set does not contain sufficient information to reconstruct the original fingerprint image from which minutiae were extracted. However, recent studies have shown that it is indeed possible to reconstruct fingerprint images from their minutiae representations. Reconstruction techniques demonstrate the need for securing fingerprint templates, improving the template interoperability, and improving fingerprint synthesis. But, there is still a large gap between the matching performance obtained from original fingerprint images and their corresponding reconstructed fingerprint images. In this paper, the prior knowledge about fingerprint ridge structures is encoded in terms of orientation patch and continuous phase patch dictionaries to improve the fingerprint reconstruction. The orientation patch dictionary is used to reconstruct the orientation field from minutiae, while the continuous phase patch dictionary is used to reconstruct the ridge pattern. Experimental results on three public domain databases (FVC2002 DB1_A, FVC2002 DB2_A, and NIST SD4) demonstrate that the proposed reconstruction algorithm outperforms the state-of-the-art reconstruction algorithms in terms of both: 1) spurious minutiae and 2) matching performance with respect to type-I attack (matching the reconstructed fingerprint against the same impression from which minutiae set was extracted) and type-II attack (matching the reconstructed fingerprint against a different impression of the same finger).
Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2015 Face Spoof Detection With Image Distortion Analysis
abstract
Automatic face recognition is now widely used in applications ranging from deduplication of identity to authentication of mobile payment. This popularity of face recognition has raised concerns about face spoof attacks (also known as biometric sensor presentation attacks), where a photo or video of an authorized person's face could be used to gain access to facilities or services. While a number of face spoof detection techniques have been proposed, their generalization ability has not been adequately addressed. We propose an efficient and rather robust face spoof detection algorithm based on image distortion analysis (IDA). Four different features (specular reflection, blurriness, chromatic moment, and color diversity) are extracted to form the IDA feature vector. An ensemble classifier, consisting of multiple SVM classifiers trained for different face spoof attacks (e.g., printed photo and replayed video), is used to distinguish between genuine (live) and spoof faces. The proposed approach is extended to multiframe face spoof detection in videos using a voting-based scheme. We also collect a face spoof database, MSU mobile face spoofing database (MSU MFSD), using two mobile devices (Google Nexus 5 and MacBook Air) with three types of spoof attacks (printed photo, replayed video with iPhone 5S, and replayed video with iPad Air). Experimental results on two public-domain face spoof databases (Idiap REPLAY-ATTACK and CASIA FASD), and the MSU MFSD database show that the proposed approach outperforms the state-of-the-art methods in spoof detection. Our results also highlight the difficulty in separating genuine and spoof faces, especially in cross-database and cross-device scenarios.
Di Wen 0001, Hu Han 0001, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2014 Image Tag Completion by Noisy Matrix Recovery
Zheyun Feng, Songhe Feng, Rong Jin 0001, Anil K. Jain 0001
ECCV (7)4
2014 Unconstrained face recognition: Establishing baseline human performance via crowdsourcing
abstract
Research focus in face recognition has shifted towards recognition of faces “in the wild” for both still images and videos which are captured in unconstrained imaging environments and without user cooperation. Due to confounding factors of pose, illumination, and expression, as well as occlusion and low resolution, current face recognition systems deployed in forensic and security applications operate in a semi-automatic manner; an operator typically reviews the top results from the face recognition system to manually determine the final match. For this reason, it is important to analyze the accuracies achieved by both the matching algorithms (machines) and humans on unconstrained face recognition tasks. In this paper, we report human accuracy on unconstrained faces in still images and videos via crowd-sourcing on Amazon Mechanical Turk. In particular, we report the first human performance on the YouTube Faces database and show that humans are superior to machines, especially when videos contain contextual cues in addition to the face image. We investigate the accuracy of humans from two different countries (United States and India) and find that humans from the United States are more accurate, possibly due to their familiarity with the faces of the public figures in the YouTube Faces database. A fusion of recognitions made by humans and a commercial-off-the-shelf face matcher improves performance over humans alone.
Lacey Best-Rowden, Shiwani Bisht, Joshua C. Klontz, Anil K. Jain 0001
IJCB4
2014 Recognizing infants and toddlers using fingerprints: Increasing the vaccination coverage
abstract
One of the major goals of most national, international and non-governmental health organizations is to eradicate the occurrence of vaccine-preventable childhood diseases (e.g., polio). Without a high vaccination coverage in a country or a geographical region, these deadly diseases take a heavy toll on children. Therefore, it is important for an effective immunization program to keep track of children who have been immunized and those who have received the required booster shots during the first 4 years of life to improve the vaccination coverage. Given that children, as well as the adults, in low income countries typically do not have any form of identification documents which can be used for this purpose, we address the following question: can fingerprints be effectively used to recognize children from birth to 4 years? We have collected 1,600 fingerprint images (500 ppi) of 20 infants and toddlers captured over a 30-day period in East Lansing, Michigan and 420 fingerprints of 70 infants and toddlers at two different health clinics in Benin, West Africa. We devised the following strategies to improve the fingerprint recognition accuracy when comparing the acquired fingerprints against an extended gallery database of 32,768 infant fingerprints collected by VaxTrac in Benin: (i) upsample the acquired fingerprint image to facilitate minutiae extraction, (ii) match the query print against templates created from each enrollment impression and fuse the match scores, (iii) fuse the match scores of the thumb and index finger, and (iv) update the gallery with fingerprints acquired over multiple sessions. A rank-1 (rank-10) identification accuracy of 83.8% (89.6%) on the East Lansing data, and 40.00% (48.57%) on the Benin data is obtained after incorporating these strategies when matching infant and toddler fingerprints using a commercial fingerprint SDK. This is an improvement of about 38% and 20%, respectively, on the two datasets without using the proposed strategies. A state-of-the-art latent finger-print SDK achieves an even higher rank-1 (rank-10) identification accuracy of 98.97% (99.39%) and 67.14% (71.43%) on the two datasets, respectively, using these strategies; an improvement of about 23% and 24%, respectively, on the two datasets without using the proposed strategies.
Anil K. Jain 0001, Kai Cao 0001, Sunpreet S. Arora
IJCB1
2014 Suspect identification based on descriptive facial attributes
abstract
We present a method for using human describable face attributes to perform face identification in criminal investigations. To enable this approach, a set of 46 facial attributes were carefully defined with the goal of capturing all describable and persistent facial features. Using crowd sourced labor, a large corpus of face images were manually annotated with the proposed attributes. In turn, we train an automated attribute extraction algorithm to encode target repositories with the attribute information. Attribute extraction is performed using localized face components to improve the extraction accuracy. Experiments are conducted to compare the use of attribute feature information, derived from crowd workers, to face sketch information, drawn by expert artists. In addition to removing the dependence on expert artists, the proposed method complements sketchbased face recognition by allowing investigators to immediately search face repositories without the time delay that is incurred due to sketch generation.
Brendan Klare, Scott Klum, Joshua C. Klontz, Emma Taborsky, Tayfun Akgül, Anil K. Jain 0001
IJCB6
2014 A Single-Pass Algorithm for Efficiently Recovering Sparse Cluster Centers of High-dimensional Data
abstract
Learning a statistical model for high-dimensional data is an important topic in machine learning. Although this problem has been well studied in the supervised setting, little is known about its unsupervised counterpart. In this work, we focus on the problem of clustering high-dimensional data with sparse centers. In particular, we address the following open question in unsupervised learning: “is it possible to reliably cluster high-dimensional data when the number of samples is smaller than the data dimensionality?" We develop an efficient clustering algorithm that is able to estimate sparse cluster centers with a single pass over the data. Our theoretical analysis shows that the proposed algorithm is able to accurately recover cluster centers with only O(s\log d) number of samples (data points), provided all the cluster centers are s-sparse vectors in a d dimensional space. Experimental results verify both the effectiveness and efficiency of the proposed clustering algorithm compared to the state-of-the-art algorithms on several benchmark datasets.
Jinfeng Yi, Lijun Zhang 0005, Jun Wang 0006, Rong Jin 0001, Anil K. Jain 0001
ICML5
2014 3D Fingerprint Phantoms
abstract
One of the critical factors prior to deployment of any large scale biometric system is to have a realistic estimate of its matching performance. In practice, evaluations are conducted on the operational data to set an appropriate threshold on match scores before the actual deployment. These performance estimates, though, are restricted by the amount of available test data. To overcome this limitation, use of a large number of 2D synthetic fingerprints for evaluating fingerprint systems had been proposed. However, the utility of 2D synthetic fingerprints is limited in the context of testing end-to-end fingerprint systems which involve the entire matching process, from image acquisition to feature extraction and matching. For a comprehensive evaluation of fingerprint systems, we propose creating 3D fingerprint phantoms (phantoms or imaging phantoms are specially designed objects with known properties scanned or imaged to evaluate, analyze, and tune the performance of various imaging devices) with known characteristics (e.g., type, singular points and minutiae) by (i) projecting 2D synthetic fingerprints with known characteristics onto a generic 3D finger surface and (ii) printing the 3D fingerprint phantoms using a commodity 3D printer. Preliminary experimental results show that the captured images of the 3D fingerprint phantoms can be successfully matched to the 2D synthetic fingerprint images (from which the phantoms were generated) using a commercial fingerprint matcher. This demonstrates that our method preserves the ridges and valleys during the 3D fingerprint phantom creation process ensuring that the synthesized 3D phantoms can be utilized for comprehensive evaluations of fingerprint systems.
Sunpreet S. Arora, Kai Cao 0001, Anil K. Jain 0001, Nicholas G. Paulter Jr.
ICPR3
2014 Segmentation and Enhancement of Latent Fingerprints: A Coarse to Fine RidgeStructure Dictionary
abstract
Latent fingerprint matching has played a critical role in identifying suspects and criminals. However, compared to rolled and plain fingerprint matching, latent identification accuracy is significantly lower due to complex background noise, poor ridge quality and overlapping structured noise in latent images. Accordingly, manual markup of various features (e.g., region of interest, singular points and minutiae) is typically necessary to extract reliable features from latents. To reduce this markup cost and to improve the consistency in feature markup, fully automatic and highly accurate ("lights-out" capability) latent matching algorithms are needed. In this paper, a dictionary-based approach is proposed for automatic latent segmentation and enhancement towards the goal of achieving "lights-out" latent identification systems. Given a latent fingerprint image, a total variation (TV) decomposition model with L1 fidelity regularization is used to remove piecewise-smooth background noise. The texture component image obtained from the decomposition of latent image is divided into overlapping patches. Ridge structure dictionary, which is learnt from a set of high quality ridge patches, is then used to restore ridge structure in these latent patches. The ridge quality of a patch, which is used for latent segmentation, is defined as the structural similarity between the patch and its reconstruction. Orientation and frequency fields, which are used for latent enhancement, are then extracted from the reconstructed patch. To balance robustness and accuracy, a coarse to fine strategy is proposed. Experimental results on two latent fingerprint databases (i.e., NIST SD27 and WVU DB) show that the proposed algorithm outperforms the state-of-the-art segmentation and enhancement algorithms and boosts the performance of a state-of-the-art commercial latent matcher.
Kai Cao 0001, Eryun Liu, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Latent Fingerprint Matching: Performance Gain via Feedback from Exemplar Prints
abstract
Latent fingerprints serve as an important source of forensic evidence in a court of law. Automatic matching of latent fingerprints to rolled/plain (exemplar) fingerprints with high accuracy is quite vital for such applications. However, latent impressions are typically of poor quality with complex background noise which makes feature extraction and matching of latents a significantly challenging problem. We propose incorporating top-down information or feedback from an exemplar to refine the features extracted from a latent for improving latent matching accuracy. The refined latent features (e.g. ridge orientation and frequency), after feedback, are used to re-match the latent to the top K candidate exemplars returned by the baseline matcher and resort the candidate list. The contributions of this research include: (i) devising systemic ways to use information in exemplars for latent feature refinement, (ii) developing a feedback paradigm which can be wrapped around any latent matcher for improving its matching performance, and (iii) determining when feedback is actually necessary to improve latent matching accuracy. Experimental results show that integrating the proposed feedback paradigm with a state-of-the-art latent matcher improves its identification accuracy by 0.5-3.5 percent for NIST SD27 and WVU latent databases against a background database of 100k exemplars.
Sunpreet S. Arora, Eryun Liu, Kai Cao 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Multiple Kernel Learning for Visual Object Recognition: A Review
abstract
Multiple kernel learning (MKL) is a principled approach for selecting and combining kernels for a given recognition task. A number of studies have shown that MKL is a useful tool for object recognition, where each image is represented by multiple sets of features and MKL is applied to combine different feature sets. We review the state-of-the-art for MKL, including different formulations and algorithms for solving the related optimization problems, with the focus on their applications to object recognition. One dilemma faced by practitioners interested in using MKL for object recognition is that different studies often provide conflicting results about the effectiveness and efficiency of MKL. To resolve this, we conduct extensive experiments on standard datasets to evaluate various approaches to MKL for object recognition. We argue that the seemingly contradictory conclusions offered by studies are due to different experimental setups. The conclusions of our study are: (i) given a sufficient number of training examples and feature/kernel types, MKL is more effective for object recognition than simple kernel combination (e.g., choosing the best performing kernel or average of kernels); and (ii) among the various approaches proposed for MKL, the sequential minimal optimization, semi-infinite programming, and level method based ones are computationally most efficient.
Serhat Selcuk Bucak, Rong Jin 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2014 Nighttime face recognition at large standoff: Cross-distance and cross-spectral matching
Dongoh Kang, Hu Han 0001, Anil K. Jain 0001, Seong-Whan Lee
Pattern Recognit.3
2014 Robust Keypoint Detection Using Higher-Order Scale Space Derivatives: Application to Image Retrieval
abstract
Image retrieval has been extensively studied over the last two decades due to the increasing demands for the effective use of multimedia data. Among various approaches to image retrieval, scale space representation and local keypoint descriptors have been shown to be a promising approach. Even though the concept of scale space representation has been known for a long time, it has now gained prominence as a powerful method for image retrieval mostly due to the invention of the Scale Invariant Feature Transform (SIFT). We will review the characteristics of the scale space operation and provide an extended method of scale space operation that significantly improves the image matching accuracy in the context of image retrieval. We use an operational tattoo image database containing 1,000 near duplicate images to show the superior retrieval performance of the proposed method compared to SIFT keypoints.
Unsang Park, Jongseung Park, Anil K. Jain 0001
IEEE Signal Process. Lett.3
2014 Unconstrained Face Recognition: Identifying a Person of Interest From a Media Collection
abstract
As face recognition applications progress from constrained sensing and cooperative subjects scenarios (e.g., driver's license and passport photos) to unconstrained scenarios with uncooperative subjects (e.g., video surveillance), new challenges are encountered. These challenges are due to variations in ambient illumination, image resolution, background clutter, facial pose, expression, and occlusion. In forensic investigations where the goal is to identify a person of interest, often based on low quality face images and videos, we need to utilize whatever source of information is available about the person. This could include one or more video tracks, multiple still images captured by bystanders (using, for example, their mobile phones), 3-D face models constructed from image(s) and video(s), and verbal descriptions of the subject provided by witnesses. These verbal descriptions can be used to generate a face sketch and provide ancillary information about the person of interest (e.g., gender, race, and age). While traditional face matching methods generally take a single media (i.e., a still face image, video track, or face sketch) as input, this paper considers using the entire gamut of media as a probe to generate a single candidate list for the person of interest. We show that the proposed approach boosts the likelihood of correctly identifying the person of interest through the use of different fusion schemes, 3-D face models, and incorporation of quality measures for fusion and video frame selection.
Lacey Best-Rowden, Hu Han 0001, Charles Otto, Brendan Klare, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.5
2014 The FaceSketchID System: Matching Facial Composites to Mugshots
abstract
Facial composites are widely used by law enforcement agencies to assist in the identification and apprehension of suspects involved in criminal activities. These composites, generated from witness descriptions, are posted in public places and media with the hope that some viewers will provide tips about the identity of the suspect. This method of identifying suspects is slow, tedious, and may not lead to the timely apprehension of a suspect. Hence, there is a need for a method that can automatically and efficiently match facial composites to large police mugshot databases. Because of this requirement, facial composite recognition is an important topic for biometrics researchers. While substantial progress has been made in nonforensic facial composite (or viewed composite) recognition over the past decade, very little work has been done using operational composites relevant to law enforcement agencies. Furthermore, no facial composite to mugshot matching systems have been documented that are readily deployable as standalone software. Thus, the contributions of this paper include: 1) an exploration of composite recognition use cases involving multiple forms of facial composites; 2) the FaceSketchID System, a scalable, and operationally deployable software system that achieves state-of-the-art matching accuracy on facial composites using two algorithms (holistic and component based); and 3) a study of the effects of training data on algorithm performance. We present experimental results using a large mugshot gallery that is representative of a law enforcement agency’s mugshot database. All results are compared against three state-of-the-art commercial-off-the-shelf face recognition systems.
Scott Klum, Hu Han 0001, Brendan Klare, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2014 Prostate Cancer Grading: Use of Graph Cut and Spatial Arrangement of Nuclei
abstract
Tissue image grading is one of the most important steps in prostate cancer diagnosis, where the pathologist relies on the gland structure to assign a Gleason grade to the tissue image. In this grading scheme, the discrimination between grade 3 and grade 4 is the most difficult, and receives the most attention from researchers. In this study, we propose a novel method (called nuclei-based method) that 1) utilizes graph theory techniques to segment glands and 2) computes a gland-score (based on the spatial arrangement of nuclei) to estimate how similar a segmented region is to a gland. Next, we create a fusion method by combining this nuclei-based method with the lumen-based method presented in our previous work to improve the performance of grade 3 versus grade 4 classification problem (the accuracy is now improved to 87.3% compared to 81.1% of the lumen-based method alone). To segment glands, we build a graph of nuclei and lumina in the image, and use the normalized cut method to partition the graph into different components, each corresponding to a gland. Unlike most state-of-the-art lumen-based gland segmentation method, the nuclei-based method is able to segment glands without lumen or glands with multiple lumina. Moreover, another important contribution in this research is the development of a set of measures to exploit the difference in nuclei spatial arrangement between grade 3 images (where nuclei form closed chain structure on the gland boundary) and grade 4 image (where nuclei distribute more randomly in the gland). These measures are combined to generate a single gland-score value, which estimates how similar a segmented region (which is a set of nuclei and lumina) is to a gland.
Kien Nguyen 0005, Anindya Sarkar, Anil K. Jain 0001
IEEE Trans. Medical Imaging3
2013 Inferring Users' Preferences from Crowdsourced Pairwise Comparisons: A Matrix Completion Approach
abstract
Inferring user preferences over a set of items is an important problem that has found numerous applications. This work focuses on the scenario where the explicit feature representation of items is unavailable, a setup that is similar to collaborative filtering. In order to learn a user's preferences from his/her response to only a small number of pairwise comparisons, we propose to leverage the pairwise comparisons made by many crowd users, a problem we refer to as crowdranking. The proposed crowdranking framework is based on the theory of matrix completion, and we present efficient algorithms for solving the related optimization problem. Our theoretical analysis shows that, on average, only O(r log m) pairwise queries are needed to accurately recover the ranking list of m items for the target user, where r is the rank of the unknown rating matrix, r << m. Our empirical study with two real-world benchmark datasets for collaborative filtering and one crowdranking dataset we collected via Amazon Mechanical Turk shows the promising performance of the proposed algorithm compared to the state-of-the-art approaches.
Jinfeng Yi, Rong Jin 0001, Shaili Jain, Anil K. Jain 0001
HCOMP4
2013 Large-Scale Image Annotation by Efficient and Robust Kernel Metric Learning
abstract
One of the key challenges in search-based image annotation models is to define an appropriate similarity measure between images. Many kernel distance metric learning (KML) algorithms have been developed in order to capture the nonlinear relationships between visual features and semantics of the images. One fundamental limitation in applying KML to image annotation is that it requires converting image annotations into binary constraints, leading to a significant information loss. In addition, most KML algorithms suffer from high computational cost due to the requirement that the learned matrix has to be positive semi-definitive (PSD). In this paper, we propose a robust kernel metric learning (RKML) algorithm based on the regression technique that is able to directly utilize image annotations. The proposed method is also computationally more efficient because PSD property is automatically ensured by regression. We provide the theoretical guarantee for the proposed algorithm, and verify its efficiency and effectiveness for image annotation by comparing it to state-of-the-art approaches for both distance metric learning and image annotation.
Zheyun Feng, Rong Jin 0001, Anil K. Jain 0001
ICCV3
2013 Semi-supervised Clustering by Input Pattern Assisted Pairwise Similarity Matrix Completion
abstract
Many semi-supervised clustering algorithms have been proposed to improve the clustering accuracy by effectively exploring the available side information that is usually in the form of pairwise constraints. Despite the progress, there are two main shortcomings of the existing semi-supervised clustering algorithms. First, they have to deal with non-convex optimization problems, leading to clustering results that are sensitive to the initialization. Second, none of these algorithms is equipped with theoretical guarantee regarding the clustering performance. We address these limitations by developing a framework for semi-supervised clustering based on \it input pattern assisted matrix completion. The key idea is to cast clustering into a matrix completion problem, and solve it efficiently by exploiting the correlation between input patterns and cluster assignments. Our analysis shows that under appropriate conditions, only O(\log n) pairwise constraints are needed to accurately recover the true cluster partition. We verify the effectiveness of the proposed algorithm by comparing it to the state-of-the-art semi-supervised clustering algorithms on several benchmark datasets.
Jinfeng Yi, Lijun Zhang 0005, Rong Jin 0001, Qi Qian 0001, Anil K. Jain 0001
ICML (3)5
2013 Orientation Field Estimation for Latent Fingerprint Enhancement
abstract
Identifying latent fingerprints is of vital importance for law enforcement agencies to apprehend criminals and terrorists. Compared to live-scan and inked fingerprints, the image quality of latent fingerprints is much lower, with complex image background, unclear ridge structure, and even overlapping patterns. A robust orientation field estimation algorithm is indispensable for enhancing and recognizing poor quality latents. However, conventional orientation field estimation algorithms, which can satisfactorily process most live-scan and inked fingerprints, do not provide acceptable results for most latents. We believe that a major limitation of conventional algorithms is that they do not utilize prior knowledge of the ridge structure in fingerprints. Inspired by spelling correction techniques in natural language processing, we propose a novel fingerprint orientation field estimation algorithm based on prior knowledge of fingerprint structure. We represent prior knowledge of fingerprints using a dictionary of reference orientation patches. which is constructed using a set of true orientation fields, and the compatibility constraint between neighboring orientation patches. Orientation field estimation for latents is posed as an energy minimization problem, which is solved by loopy belief propagation. Experimental results on the challenging NIST SD27 latent fingerprint database and an overlapped latent fingerprint database demonstrate the advantages of the proposed orientation field estimation algorithm over conventional algorithms.
Jianjiang Feng, Jie Zhou 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Heterogeneous Face Recognition Using Kernel Prototype Similarities
abstract
Heterogeneous face recognition (HFR) involves matching two face images from alternate imaging modalities, such as an infrared image to a photograph or a sketch to a photograph. Accurate HFR systems are of great value in various applications (e.g., forensics and surveillance), where the gallery databases are populated with photographs (e.g., mug shot or passport photographs) but the probe images are often limited to some alternate modality. A generic HFR framework is proposed in which both probe and gallery images are represented in terms of nonlinear similarities to a collection of prototype face images. The prototype subjects (i.e., the training set) have an image in each modality (probe and gallery), and the similarity of an image is measured against the prototype images from the corresponding modality. The accuracy of this nonlinear prototype representation is improved by projecting the features into a linear discriminant subspace. Random sampling is introduced into the HFR framework to better handle challenges arising from the small sample size problem. The merits of the proposed approach, called prototype random subspace (P-RS), are demonstrated on four different heterogeneous scenarios: 1) near infrared (NIR) to photograph, 2) thermal to photograph, 3) viewed sketch to photograph, and 4) forensic sketch to photograph.
Brendan Klare, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Partial Face Recognition: Alignment-Free Approach
abstract
Numerous methods have been developed for holistic face recognition with impressive performance. However, few studies have tackled how to recognize an arbitrary patch of a face image. Partial faces frequently appear in unconstrained scenarios, with images captured by surveillance cameras or handheld devices (e.g., mobile phones) in particular. In this paper, we propose a general partial face recognition approach that does not require face alignment by eye coordinates or any other fiducial points. We develop an alignment-free face representation method based on Multi-Keypoint Descriptors (MKD), where the descriptor size of a face is determined by the actual content of the image. In this way, any probe face image, holistic or partial, can be sparsely represented by a large dictionary of gallery descriptors. A new keypoint descriptor called Gabor Ternary Pattern (GTP) is also developed for robust and discriminative face recognition. Experimental results are reported on four public domain face databases (FRGCv2.0, AR, LFW, and PubFig) under both the open-set identification and verification scenarios. Comparisons with two leading commercial face recognition SDKs (PittPatt and FaceVACS) and two baseline algorithms (PCA+LDA and LBP) show that the proposed method, overall, is superior in recognizing both holistic and partial faces without requiring alignment.
Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 A Coarse to Fine Minutiae-Based Latent Palmprint Matching
abstract
With the availability of live-scan palmprint technology, high resolution palmprint recognition has started to receive significant attention in forensics and law enforcement. In forensic applications, latent palmprints provide critical evidence as it is estimated that about 30 percent of the latents recovered at crime scenes are those of palms. Most of the available high-resolution palmprint matching algorithms essentially follow the minutiae-based fingerprint matching strategy. Considering the large number of minutiae (about 1,000 minutiae in a full palmprint compared to about 100 minutiae in a rolled fingerprint) and large area of foreground region in full palmprints, novel strategies need to be developed for efficient and robust latent palmprint matching. In this paper, a coarse to fine matching strategy based on minutiae clustering and minutiae match propagation is designed specifically for palmprint matching. To deal with the large number of minutiae, a local feature-based minutiae clustering algorithm is designed to cluster minutiae into several groups such that minutiae belonging to the same group have similar local characteristics. The coarse matching is then performed within each cluster to establish initial minutiae correspondences between two palmprints. Starting with each initial correspondence, a minutiae match propagation algorithm searches for mated minutiae in the full palmprint. The proposed palmprint matching algorithm has been evaluated on a latent-to-full palmprint database consisting of 446 latents and 12,489 background full prints. The matching results show a rank-1 identification accuracy of 79.4 percent, which is significantly higher than the 60.8 percent identification accuracy of a state-of-the-art latent palmprint matching algorithm on the same latent database. The average computation time of our algorithm for a single latent-to-full match is about 141 ms for genuine match and 50 ms for impostor match, on a Windows XP desktop system with 2.2-GHz CPU and 1.00-GB RAM. The computation time of our algorithm is an order of magnitude faster than a previously published state-of-the-art-algorithm.
Eryun Liu, Anil K. Jain 0001, Jie Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Tag Completion for Image Retrieval
abstract
Many social image search engines are based on keyword/tag matching. This is because tag-based image retrieval (TBIR) is not only efficient but also effective. The performance of TBIR is highly dependent on the availability and quality of manual tags. Recent studies have shown that manual tags are often unreliable and inconsistent. In addition, since many users tend to choose general and ambiguous tags in order to minimize their efforts in choosing appropriate words, tags that are specific to the visual content of images tend to be missing or noisy, leading to a limited performance of TBIR. To address this challenge, we study the problem of tag completion, where the goal is to automatically fill in the missing tags as well as correct noisy tags for given images. We represent the image-tag relation by a tag matrix, and search for the optimal tag matrix consistent with both the observed tags and the visual similarity. We propose a new algorithm for solving this optimization problem. Extensive empirical studies show that the proposed algorithm is significantly more effective than the state-of-the-art algorithms. Our studies also verify that the proposed algorithm is computationally efficient and scales well to large databases.
Lei Wu 0017, Rong Jin 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Component-Based Representation in Automated Face Recognition
abstract
This paper presents a framework for component-based face alignment and representation that demonstrates improvements in matching performance over the more common holistic approach to face alignment and representation. This work is motivated by recent evidence from the cognitive science community demonstrating the efficacy of component-based facial representations. The component-based framework presented in this paper consists of the following major steps: 1) landmark extraction using Active Shape Models (ASM), 2) alignment and cropping of components using Procrustes Analysis, 3) representation of components with Multiscale Local Binary Patterns (MLBP), 4) per-component measurement of facial similarity, and 5) fusion of per-component similarities. We demonstrate on three public datasets and an operational dataset consisting of face images of 8000 subjects, that the proposed component-based representation provides higher recognition accuracies over holistic-based representations. Additionally, we show that the proposed component-based representations: 1) are more robust to changes in facial pose, and 2) improve recognition accuracy on occluded face images in forensic scenarios.
Kathryn Bonnen, Brendan Klare, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2013 Matching Composite Sketches to Face Photos: A Component-Based Approach
abstract
The problem of automatically matching composite sketches to facial photographs is addressed in this paper. Previous research on sketch recognition focused on matching sketches drawn by professional artists who either looked directly at the subjects (viewed sketches) or used a verbal description of the subject's appearance as provided by an eyewitness (forensic sketches). Unlike sketches hand drawn by artists, composite sketches are synthesized using one of the several facial composite software systems available to law enforcement agencies. We propose a component-based representation (CBR) approach to measure the similarity between a composite sketch and mugshot photograph. Specifically, we first automatically detect facial landmarks in composite sketches and face photos using an active shape model (ASM). Features are then extracted for each facial component using multiscale local binary patterns (MLBPs), and per component similarity is calculated. Finally, the similarity scores obtained from individual facial components are fused together, yielding a similarity score between a composite sketch and a face photo. Matching performance is further improved by filtering the large gallery of mugshot images using gender information. Experimental results on matching 123 composite sketches against two galleries with 10,123 and 1,316 mugshots show that the proposed method achieves promising performance (rank-100 accuracies of 77.2% and 89.4%, respectively) compared to a leading commercial face recognition system (rank-100 accuracies of 22.8% and 52.0%) and densely sampled MLBP on holistic faces (rank-100 accuracies of 27.6% and 10.6%). We believe our prototype system will be of great value to law enforcement agencies in apprehending suspects in a timely fashion.
Hu Han 0001, Brendan Klare, Kathryn Bonnen, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2013 Face Tracking and Recognition at a Distance: A Coaxial and Concentric PTZ Camera System
abstract
Face recognition has been regarded as an effective method for subject identification at a distance because of its covert and remote sensing capability. However, face images have a low resolution when they are captured at a distance (say, larger than 5 meters) thereby degrading the face matching performance. To address this problem, we propose an imaging system consisting of static and pan-tilt-zoom (PTZ) cameras to acquire high resolution face images up to a distance of 12 meters. We propose a novel coaxial-concentric camera configuration between the static and PTZ cameras to achieve the distance invariance property using a simple calibration scheme. We also use a linear prediction model and camera motion control to mitigate delays in image processing and mechanical camera motion. Our imaging system was used to track 50 different subjects and their faces at distances ranging from 6 to 12 meters. The matching scenario consisted of these 50 subjects as probe and additional 10 000 subjects as gallery. Rank-1 identification accuracy of 91.5% was achieved compared to 0% rank-1 accuracy of the conventional camera system using a state-of-the-art matcher. The proposed camera system can operate at a larger distance (up to 50 meters) by replacing the static camera with a PTZ camera to detect a subject at a larger distance and control the second PTZ camera to capture the high-resolution face image.
Unsang Park, Hyun-Cheol Choi, Anil K. Jain 0001, Seong-Whan Lee
IEEE Trans. Inf. Forensics Secur.3
2013 Latent Fingerprint Matching Using Descriptor-Based Hough Transform
abstract
Identifying suspects based on impressions of fingers lifted from crime scenes (latent prints) is a routine procedure that is extremely important to forensics and law enforcement agencies. Latents are partial fingerprints that are usually smudgy, with small area and containing large distortion. Due to these characteristics, latents have a significantly smaller number of minutiae points compared to full (rolled or plain) fingerprints. The small number of minutiae and the noise characteristic of latents make it extremely difficult to automatically match latents to their mated full prints that are stored in law enforcement databases. Although a number of algorithms for matching full-to-full fingerprints have been published in the literature, they do not perform well on the latent-to-full matching problem. Further, they often rely on features that are not easy to extract from poor quality latents. In this paper, we propose a new fingerprint matching algorithm which is especially designed for matching latents. The proposed algorithm uses a robust alignment algorithm (descriptor-based Hough transform) to align fingerprints and measures similarity between fingerprints by considering both minutiae and orientation field information. To be consistent with the common practice in latent matching (i.e., only minutiae are marked by latent examiners), the orientation field is reconstructed from minutiae. Since the proposed algorithm relies only on manually marked minutiae, it can be easily used in law enforcement applications. Experimental results on two different latent databases (NIST SD27 and WVU latent databases) show that the proposed algorithm outperforms two well optimized commercial fingerprint matchers. Further, a fusion of the proposed algorithm and commercial fingerprint matchers leads to improved matching accuracy.
Alessandra A. Paulino, Jianjiang Feng, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2012 Nighttime Face Recognition at Long Distance: Cross-Distance and Cross-Spectral Matching
Hyunju Maeng, Shengcai Liao, Dongoh Kang, Seong-Whan Lee, Anil K. Jain 0001
ACCV (2)5
2012 Efficient Kernel Clustering Using Random Fourier Features
abstract
Kernel clustering algorithms have the ability to capture the non-linear structure inherent in many real world data sets and thereby, achieve better clustering performance than Euclidean distance based clustering algorithms. However, their quadratic computational complexity renders them non-scalable to large data sets. In this paper, we employ random Fourier maps, originally proposed for large scale classification, to accelerate kernel clustering. The key idea behind the use of random Fourier maps for clustering is to project the data into a low-dimensional space where the inner product of the transformed data points approximates the kernel similarity between them. An efficient linear clustering algorithm can then be applied to the points in the transformed space. We also propose an improved scheme which uses the top singular vectors of the transformed data matrix to perform clustering, and yields a better approximation of kernel clustering under appropriate conditions. Our empirical studies demonstrate that the proposed schemes can be efficiently applied to large data sets containing millions of data points, while achieving accuracy similar to that achieved by state-of-the-art kernel clustering algorithms.
Radha Chitta, Rong Jin 0001, Anil K. Jain 0001
ICDM3
2012 Robust Ensemble Clustering by Matrix Completion
abstract
Data clustering is an important task and has found applications in numerous real-world problems. Since no single clustering algorithm is able to identify all different types of cluster shapes and structures, ensemble clustering was proposed to combine different partitions of the same data generated by multiple clustering algorithms. The key idea of most ensemble clustering algorithms is to find a partition that is consistent with most of the available partitions of the input data. One problem with these algorithms is their inability to handle uncertain data pairs, i.e. data pairs for which about half of the partitions put them into the same cluster and the other half do the opposite. When the number of uncertain data pairs is large, they can mislead the ensemble clustering algorithm in generating the final partition. To overcome this limitation, we propose an ensemble clustering approach based on the technique of matrix completion. The proposed algorithm constructs a partially observed similarity matrix based on the data pairs whose cluster memberships are agreed upon by most of the clustering algorithms in the ensemble. It then deploys the matrix completion algorithm to complete the similarity matrix. The final data partition is computed by applying an efficient spectral clustering algorithm to the completed matrix. Our empirical studies with multiple real-world datasets show that the proposed algorithm performs significantly better than the state-of-the-art algorithms for ensemble clustering.
Jinfeng Yi, Tianbao Yang, Rong Jin 0001, Anil K. Jain 0001, Mehrdad Mahdavi
ICDM4
2012 Face Recognition in the Virtual World: Recognizing Avatar Faces
abstract
Criminal activity in virtual worlds is becoming a major problem for law enforcement agencies. Forensic investigators are becoming interested in being able to accurately and automatically track people in virtual communities. In this paper a set of algorithms capable of verification and recognition of avatar faces with high degree of accuracy are described. Results of experiments aimed at within-virtual-world avatar authentication and inter-reality-based scenarios of tracking a person between real and virtual worlds are reported. In the FERET-to-Avatar face dataset, where an avatar face was generated from every photo in the FERET database, a COTS FR algorithm achieved a near perfect 99.58% accuracy on 725 subjects. On a dataset of avatars from Second Life, the proposed avatar-to-avatar matching algorithm (which uses a fusion of local structural and appearance descriptors) achieved average true accept rates of (i) 96.33% using manual eye detection, and (ii) 86.5% in a fully automated mode at a false accept rate of 1.0%. A combination of the proposed face matcher and a state-of-the art commercial matcher (FaceVACS) resulted in further improvement on the inter-reality-based scenario.
Roman V. Yampolskiy, Brendan Klare, Anil K. Jain 0001
ICMLA (1)3
2012 Structure and Context in Prostatic Gland Segmentation and Classification
Kien Nguyen 0005, Anindya Sarkar, Anil K. Jain 0001
MICCAI (1)3
2012 Semi-Crowdsourced Clustering: Generalizing Crowd Labeling by Robust Distance Metric Learning
abstract
One of the main challenges in data clustering is to define an appropriate similarity measure between two objects. Crowdclustering addresses this challenge by defining the pairwise similarity based on the manual annotations obtained through crowdsourcing. Despite its encouraging results, a key limitation of crowdclustering is that it can only cluster objects when their manual annotations are available. To address this limitation, we propose a new approach for clustering, called \textit{semi-crowdsourced clustering} that effectively combines the low-level features of objects with the manual annotations of a subset of the objects obtained via crowdsourcing. The key idea is to learn an appropriate similarity measure, based on the low-level features of objects, from the manual annotations of only a small portion of the data to be clustered. One difficulty in learning the pairwise similarity measure is that there is a significant amount of noise and inter-worker variations in the manual annotations obtained via crowdsourcing. We address this difficulty by developing a metric learning algorithm based on the matrix completion method. Our empirical study with two real-world image data sets shows that the proposed algorithm outperforms state-of-the-art distance metric learning algorithms in both clustering accuracy and computational efficiency.
Jinfeng Yi, Rong Jin 0001, Anil K. Jain 0001, Shaili Jain, Tianbao Yang
NIPS3
2012 Simultaneous classification and community detection on heterogeneous network data
Prakash Mandayam Comar, Pang-Ning Tan, Anil K. Jain 0001
Data Min. Knowl. Discov.3
2012 A framework for joint community detection across multiple related networks
Prakash Mandayam Comar, Pang-Ning Tan, Anil K. Jain 0001
Neurocomputing3
2012 Altered Fingerprints: Analysis and Detection
abstract
The widespread deployment of Automated Fingerprint Identification Systems (AFIS) in law enforcement and border control applications has heightened the need for ensuring that these systems are not compromised. While several issues related to fingerprint system security have been investigated, including the use of fake fingerprints for masquerading identity, the problem of fingerprint alteration or obfuscation has received very little attention. Fingerprint obfuscation refers to the deliberate alteration of the fingerprint pattern by an individual for the purpose of masking his identity. Several cases of fingerprint obfuscation have been reported in the press. Fingerprint image quality assessment software (e.g., NFIQ) cannot always detect altered fingerprints since the implicit image quality due to alteration may not change significantly. The main contributions of this paper are: 1) compiling case studies of incidents where individuals were found to have altered their fingerprints for circumventing AFIS, 2) investigating the impact of fingerprint alteration on the accuracy of a commercial fingerprint matcher, 3) classifying the alterations into three major categories and suggesting possible countermeasures, 4) developing a technique to automatically detect altered fingerprints based on analyzing orientation field and minutiae distribution, and 5) evaluating the proposed technique and the NFIQ algorithm on a large database of altered fingerprints provided by a law enforcement agency. Experimental results show the feasibility of the proposed approach in detecting altered fingerprints and highlight the need to further pursue this problem.
Soweon Yoon, Jianjiang Feng, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 Special Issue on Awards from ICPR 2010
Kim Boyer, Müjdat Çetin, Anil K. Jain 0001, Seong-Whan Lee
Pattern Recognit. Lett.3
2012 Pill-ID: Matching and retrieval of drug pill images
Young-Beom Lee, Unsang Park, Anil K. Jain 0001, Seong-Whan Lee
Pattern Recognit. Lett.3
2012 Prostate cancer grading: Gland segmentation and structural features
Kien Nguyen 0005, Bikash Sabata, Anil K. Jain 0001
Pattern Recognit. Lett.3
2012 Face Recognition Performance: Role of Demographic Information
abstract
This paper studies the influence of demographics on the performance of face recognition algorithms. The recognition accuracies of six different face recognition algorithms (three commercial, two nontrainable, and one trainable) are computed on a large scale gallery that is partitioned so that each partition consists entirely of specific demographic cohorts. Eight total cohorts are isolated based on gender (male and female), race/ethnicity (Black, White, and Hispanic), and age group (18–30, 30–50, and 50–70 years old). Experimental results demonstrate that both commercial and the nontrainable algorithms consistently have lower matching accuracies on the same cohorts (females, Blacks, and age group 18–30) than the remaining cohorts within their demographic. Additional experiments investigate the impact of the demographic distribution in the training set on the performance of a trainable face recognition algorithm. We show that the matching accuracy for race/ethnicity and age cohorts can be improved by training exclusively on that specific cohort. Operationally, this leads to a scenario, called dynamic face matcher selection, where multiple face recognition algorithms (each trained on different demographic cohorts) are available for a biometric system operator to select based on the demographic information extracted from a probe image. This procedure should lead to improved face recognition accuracy in many intelligence and law enforcement face recognition scenarios. Finally, we show that an alternative to dynamic face matcher selection is to train face recognition algorithms on datasets that are evenly distributed across demographics, as this approach offers consistently high accuracy across all cohorts.
Brendan Klare, Mark James Burge, Joshua C. Klontz, Richard W. Vorder Bruegge, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.5
2012 Coupled Discriminant Analysis for Heterogeneous Face Recognition
abstract
Coupled space learning is an effective framework for heterogeneous face recognition. In this paper, we propose a novel coupled discriminant analysis method to improve the heterogeneous face recognition performance. There are two main advantages of the proposed method. First, all samples from different modalities are used to represent the coupled projections, so that sufficient discriminative information could be extracted. Second, the locality information in kernel space is incorporated into the coupled discriminant analysis as a constraint to improve the generalization ability. In particular, two implementations of locality constraint in kernel space (LCKS)-based coupled discriminant analysis methods, namely LCKS-coupled discriminant analysis (LCKS-CDA) and LCKS-coupled spectral regression (LCKS-CSR), are presented. Extensive experiments on three cases of heterogeneous face matching (high versus low image resolution, digital photo versus video image, and visible light versus near infrared) validate the efficacy of the proposed method.
Zhen Lei 0001, Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.3
2012 Evidential Value of Automated Latent Fingerprint Comparison: An Empirical Approach
abstract
Latent prints are routinely recovered from crime scenes and are compared with available databases of known fingerprints for identifying criminals. However, current procedures to compare latent prints to large databases of exemplar (rolled or plain) prints are prone to errors. This suggests caution in making conclusions about a suspect's identity based on a latent fingerprint comparison. A number of attempts have been made to statistically model the utility of a fingerprint comparison in making a correct accept/reject decision or its evidential value. These approaches, however, either make unrealistic assumptions about the model or they lack simple interpretation. We argue that the posterior probability of two fingerprints belonging to different fingers given their match score, referred to as the nonmatch probability (NMP), effectively captures any implicating evidence of the comparison. NMP is computed using state-of-the-art matchers and is easy to interpret. To incorporate the effect of image quality, number of minutiae, and size of the latent on NMP value, we compute the NMP vs. match score plots separately for image pairs (latent and exemplar prints) with different characteristics. Given the paucity of latent fingerprint databases in public domain, we simulate latent prints using two exemplar print databases (NIST SD-14 and Michigan State Police) by cropping regions of three different sizes. We appropriately validate this simulation using four latent databases (NIST SD-27 and three proprietary latent databases) and two state-of-the-art fingerprint matchers to compute their respective match scores. We also discuss a practical scenario where a latent examiner uses the proposed framework to compute the evidential value of a latent-exemplar print pair comparison.
Abhishek Nagar, Heeseung Choi, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2012 Multibiometric Cryptosystems Based on Feature-Level Fusion
abstract
Multibiometric systems are being increasingly de- ployed in many large-scale biometric applications (e.g., FBI-IAFIS, UIDAI system in India) because they have several advantages such as lower error rates and larger population coverage compared to unibiometric systems. However, multibiometric systems require storage of multiple biometric templates (e.g., fingerprint, iris, and face) for each user, which results in increased risk to user privacy and system security. One method to protect individual templates is to store only the secure sketch generated from the corresponding template using a biometric cryptosystem. This requires storage of multiple sketches. In this paper, we propose a feature-level fusion framework to simultaneously protect multiple templates of a user as a single secure sketch. Our main contributions include: (1) practical implementation of the proposed feature-level fusion framework using two well-known biometric cryptosystems, namery,fuzzy vault and fuzzy commitment, and (2) detailed analysis of the trade-off between matching accuracy and security in the proposed multibiometric cryptosystems based on two different databases (one real and one virtual multimodal database), each containing the three most popular biometric modalities, namely, fingerprint, iris, and face. Experimental results show that both the multibiometric cryptosystems proposed here have higher security and matching performance compared to their unibiometric counterparts.
Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2012 Model Based Separation of Overlapping Latent Fingerprints
abstract
Latent fingerprints lifted from crime scenes often contain overlapping prints, which are difficult to separate and match by state-of-the-art fingerprint matchers. A few methods have been proposed to separate overlapping fingerprints to enable fingerprint matchers to successfully match the component fingerprints. These methods are limited by the accuracy of the estimated orientation field, which is not reliable for poor quality overlapping latent fingerprints. In this paper, we improve the robustness of overlapping fingerprints separation, particularly for low quality images. Our algorithm reconstructs the orientation fields of component prints by modeling fingerprint orientation fields. In order to facilitate this, we utilize the orientation cues of component fingerprints, which are manually marked by fingerprint examiners. This additional markup is acceptable in forensics, where the first priority is to improve the latent matching accuracy. The effectiveness of the proposed method has been evaluated not only on simulated overlapping prints, but also on real overlapped latent fingerprint images. Compared with available methods, the proposed algorithm is more effective in separating poor quality overlapping fingerprints and enhancing the matching accuracy of overlapping fingerprints.
Qijun Zhao, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2011 Multi-label learning with incomplete class assignments
abstract
We consider a special type of multi-label learning where class assignments of training examples are incomplete. As an example, an instance whose true class assignment is (c1, c2, c3) is only assigned to class c1when it is used as a training sample. We refer to this problem as multi-label learning with incomplete class assignment. Incompletely labeled data is frequently encountered when the number of classes is very large (hundreds as in MIR Flickr dataset) or when there is a large ambiguity between classes (e.g., jet vs plane). In both cases, it is difficult for users to provide complete class assignments for objects. We propose a ranking based multi-label learning framework that explicitly addresses the challenge of learning from incompletely labeled data by exploiting the group lasso technique to combine the ranking errors. We present a learning algorithm that is empirically shown to be efficient for solving the related optimization problem. Our empirical study shows that the proposed framework is more effective than the state-of-the-art algorithms for multi-label learning in dealing with incompletely labeled data.
Serhat Selcuk Bucak, Rong Jin 0001, Anil K. Jain 0001
CVPR3
2011 Face recognition: Some challenges in forensics
abstract
Face recognition has become a valuable and routine forensic tool used by criminal investigators. Compared to automated face recognition, forensic face recognition is more demanding because it must be able to handle facial images captured under non-ideal conditions and it has high liability for following legal procedures. This paper discusses recent developments in automated face recognition that impact the forensic face recognition community. Improvements in forensic face recognition through research in facial aging, facial marks, forensic sketch recognition, face recognition in video, near-infrared face recognition, and use of soft biometrics will be discussed. Finally, current limitations and future research directions for face recognition in forensics are suggested.
Anil K. Jain 0001, Brendan Klare, Unsang Park
FG1
2011 Speedup of fuzzy and possibilistic kernel c-means for large-scale clustering
abstract
The ubiquity of personal computing technology has produced an abundance of staggeringly large data sets-the Library of Congress has stored over 160 terabytes of web data and it is estimated that Facebook alone logs over 25 terabytes of data per day. There is a great need for systems by which one can elucidate the similarity and dissimilarity among and between groups in these data sets. Clustering is one way to find these groups. In this paper, we propose an approximation method for the fuzzy and possibilistic kernel c-means clustering algorithms. Our approximation constrains the cluster centers to be linear combinations of a size m randomly selected subset of the n input objects, where m ≪ n. The proposed algorithm only requires an m × n rectangular portion of the full n × n kernel matrix and the n diagonal values, resulting in significant memory savings. Furthermore, the computational complexity of the c-means algorithm is substantially reduced. We demonstrate that up to 3 orders of magnitude of speedup are possible while achieving almost the same performance as the original kernel c-means algorithm.
Timothy C. Havens, Radha Chitta, Anil K. Jain 0001, Rong Jin 0001
FUZZ-IEEE3
2011 On the evidential value of fingerprints
abstract
Fingerprint evidence is routinely used by forensics and law enforcement agencies worldwide to apprehend and convict criminals, a practice in use for over 100 years. The use of fingerprints has been accepted as an infallible proof of identity based on two premises: (i) permanence or persistence, and (ii) uniqueness or individuality. However, in the absence of any theoretical results that establish the unique ness or individuality of fingerprints, the use of fingerprints in various court proceedings is being questioned. This has raised awareness in the forensics community about the need to quantify the evidential value of fingerprint matching. A few studies that have studied this problem estimate this evidential value in one of two ways: (i) feature modeling, where a statistical (generative) model for fingerprint features, primarily minutiae, is developed which is then used to estimate the matching error and (ii) match score modeling, where a set of match scores obtained over a database is used to estimate the matching error rates. Our focus here is on match score modeling and we develop metrics to evaluate the effectiveness and reliability of the proposed evidential measure. Compared to previous approaches, the proposed measure allows explicit utilization of prior odds. Further, we also incorporate fingerprint image quality to improve the reliability of the estimated evidential value.
Heeseung Choi, Abhishek Nagar, Anil K. Jain 0001
IJCB3
2011 Face recognition across time lapse: On learning feature subspaces
abstract
There is a growing interest in understanding the impact of aging on face recognition performance, as well as de- signing recognition algorithms that are mostly invariant to temporal changes. While some success has been made on this front, a fundamental questions has yet to be answered: do face recognition systems that compensate for the effects of aging compromise recognition performance for faces that have not undergone any aging? The studies in this paper help confirm that age invariant systems do seem to decrease performance in non-aging scenarios. This is demonstrated by performing training experiments on the largest face aging dataset studied in the literature to date (over 200,000 images from roughly 64,000 subjects). Further experiments conducted in this research help demonstrate the impact of aging on two leading commercial face recognition systems. We also determine the regions of the face that remain the most stable over time.
Brendan Klare, Anil K. Jain 0001
IJCB2
2011 Analysis of facial features in identical twins
abstract
A study of the distinctiveness of different facial features (MLBP, SIFT, and facial marks) with respect to distinguishing identical twins is presented. The accuracy of distinguishing between identical twin pairs is measured using the entire face, as well as each facial component (eyes, eye brows, nose, and mouth). The impact of discriminant learning methods on twin face recognition is investigated. Experimental results indicate that features that perform well in distinguishing identical twins are not always consistent with the features that best distinguish two non-twin faces.
Brendan Klare, Alessandra A. Paulino, Anil K. Jain 0001
IJCB3
2011 Biometric recognition of newborns: Identification using palmprints
abstract
We present some results on newborn identification through high-resolution images of palmar surfaces. To our knowledge, there is no biometric system currently available that can be effectively used for newborn identification. The manual procedure of capturing inked footprints in practice for this purpose is limited for use inside hospitals and is not an effective solution for identification purposes. The use of friction ridge patterns on the hands of newborns is challenging due to both the small size of newborn's papillary ridges, which are, on average, 2.5 to 3 times smaller than the ridges in adult fingerprints, and their fragility, making them amenable to deformation. The proposed palmprint based automatic system for newborn identification is relatively easy to use and shows the feasibility of this approach. Experiments were performed on images collected from 250 newborns at the University Hospital (Universidade Federal do Parana). An image acquisition protocol was developed in order to collect suitable images. When considering the good quality palmar images, the results show that the pro- posed approach is promising.
Rubisley de P. Lemes, Olga R. P. Bellon, Luciano Silva, Anil K. Jain 0001
IJCB4
2011 Partial face recognition: An alignment free approach
abstract
Many approaches have been developed for holistic face recognition with impressive performance. However, few studies have addressed the question of how to recognize an arbitrary image patch of a holistic face. In this paper we ad- dress this problem of partial face recognition. Partial faces frequently appear in unconstrained image capture environments, particularly when faces are captured by surveillance cameras or handheld devices (e.g. mobile phones). The pro- posed approach adopts a variable-size description which represents each face with a set of keypoint descriptors. In this way, we argue that a probe face image, holistic or partial, can be sparsely represented by a large dictionary of gallery descriptors. The proposed method is alignment free and we address large-scale face recognition problems by a fast filtering strategy. Experimental results on three public domain face databases (FRGCv2.0, AR, and LFW) show that the proposed method achieves promising results in recognizing both holistic and partial faces.
Shengcai Liao, Anil K. Jain 0001
IJCB2
2011 NFRAD: Near-Infrared Face Recognition at a Distance
abstract
Face recognition at a distance is gaining wide attention in order to augment the surveillance systems with face recognition capability. However, face recognition at a distance in nighttime has not yet received adequate attention considering the increased security threats at nighttime. We introduce a new face image database, called Near-Infrared Face Recognition at a Distance Database (NFRAD-DB). Images in NFRAD-DB are collected at a distance of up to 60 meters with 50 different subjects using a near-infrared camera, a telescope, and near-infrared illuminator. We provide face recognition performance using FaceVACS, DoG-SIFT, and DoG-MLBP representations. The face recognition test consisted of NIR images of these 50 subjects at 60 meters as probe and visible images at 1 meter with additional mug shot images of 10,000 subjects as gallery. Rank-1 identification accuracy of 28 percent was achieved from the proposed method compared to 18 percent rank-1 accuracy of a state of the art face recognition system, FaceVACS. These recognition results are encouraging given this challenging matching problem due to the illumination pattern and insufficient brightness in NFRAD images.
Hyunju Maeng, Hyun-Cheol Choi, Unsang Park, Seong-Whan Lee, Anil K. Jain 0001
IJCB5
2011 Latent fingerprint matching using descriptor-based hough transform
abstract
Identifying suspects based on impressions of fingers lifted from crime scenes (latent prints) is extremely important to law enforcement agencies. Latents are usually partial fingerprints with small area, contain nonlinear distortion, and are usually smudgy and blurred. Due to some of these characteristics, they have a significantly smaller number of minutiae points (one of the most important features in fingerprint matching) and therefore it can be extremely difficult to automatically match latents to plain or rolled fingerprints that are stored in law enforcement databases. Our goal is to develop a latent matching algorithm that uses only minutiae information. The proposed approach consists of following three modules: (i) align two sets of minutiae by using a descriptor-based Hough Transform; (ii) establish the correspondences between minutiae; and (iii) compute a similarity score. Experimental results on NIST SD27 show that the proposed algorithm outperforms a commercial fingerprint matcher.
Alessandra A. Paulino, Jianjiang Feng, Anil K. Jain 0001
IJCB3
2011 Latent fingerprint enhancement via robust orientation field estimation
abstract
Latent fingerprints, or simply latents, have been considered as cardinal evidence for identifying and convicting criminals. The amount of information available for identification from latents is often limited due to their poor quality, unclear ridge structure and occlusion with complex back ground or even other latent prints. We propose a latent fingerprint enhancement algorithm, which expects manually marked region of interest (ROI) and singular points. The core of the proposed algorithm is a robust orientation field estimation algorithm for latents. Short-time Fourier transform is used to obtain multiple orientation elements in each image block. This is followed by a hypothesize-and test paradigm based on randomized RANSAC, which generates a set of hypothesized orientation fields. Experimental results on NIST SD27 latent fingerprint database show that the matching performance of a commercial matcher is significantly improved by utilizing the enhanced latent finger prints produced by the proposed algorithm.
Soweon Yoon, Jianjiang Feng, Anil K. Jain 0001
IJCB3
2011 3D to 2D fingerprints: Unrolling and distortion correction
abstract
Touchless 3D fingerprint sensors can capture both 3D depth information and albedo images of the finger surface. Compared with 2D fingerprint images acquired by traditional contact-based fingerprint sensors, the 3D fingerprints are generally free from the distortion caused by non-uniform pressure and undesirable motion of the finger. Several unrolling algorithms have been proposed for virtual rolling of 3D fingerprints to obtain 2D equivalent fingerprints, so that they can be matched with the legacy 2D fingerprint databases. However, available unrolling algorithms do not consider the impact of distortion that is typically present in the legacy 2D fingerprint images. In this paper, we conduct a comparative study of representative unrolling algorithms and propose an effective approach to incorporate distortion into the unrolling process. The 3D fingerprint database was acquired by using a 3D fingerprint sensor being developed by the General Electric Global Research. By matching the 2D equivalent fingerprints with the corresponding 2D fingerprints collected with a commercial contact-based fingerprint sensor, we show that the compatibility between the 2D unrolled fingerprints and the traditional contact-based 2D fingerprints is improved after incorporating the distortion into the unrolling process.
Qijun Zhao, Anil K. Jain 0001, Gil Abramovich
IJCB2
2011 LinkBoost: A Novel Cost-Sensitive Boosting Framework for Community-Level Network Link Prediction
abstract
Link prediction is a challenging task due to the inherent skew ness of network data. Typical link prediction methods can be categorized as either local or global. Local methods consider the link structure in the immediate neighborhood of a node pair to determine the presence or absence of a link, whereas global methods utilize information from the whole network. This paper presents a community (cluster) level link prediction method without the need to explicitly identify the communities in a network. Specifically, a variable-cost loss function is defined to address the data skew ness problem. We provide theoretical proof that shows the equivalence between maximizing the well-known modularity measure used in community detection and minimizing a special case of the proposed loss function. As a result, any link prediction method designed to optimize the loss function would result in more links being predicted within a community than between communities. We design a boosting algorithm to minimize the loss function and present an approach to scale-up the algorithm by decomposing the network into smaller partitions and aggregating the weak learners constructed from each partition. Experimental results show that our proposed Link Boost algorithm consistently performs as good as or better than many existing methods when evaluated on 4 real-world network datasets.
Prakash Mandayam Comar, Pang-Ning Tan, Anil K. Jain 0001
ICDM3
2011 Approximate kernel k-means: solution to large scale kernel clustering
abstract
Digital data explosion mandates the development of scalable tools to organize the data in a meaningful and easily accessible form. Clustering is a commonly used tool for data organization. However, many clustering algorithms designed to handle large data sets assume linear separability of data and hence do not perform well on real world data sets. While kernel-based clustering algorithms can capture the non-linear structure in data, they do not scale well in terms of speed and memory requirements when the number of objects to be clustered exceeds tens of thousands. We propose an approximation scheme for kernel k-means, termed approximate kernel k-means, that reduces both the computational complexity and the memory requirements by employing a randomized approach. We show both analytically and empirically that the performance of approximate kernel k-means is similar to that of the kernel k-means algorithm, but with dramatically reduced run-time complexity and memory requirements.
Radha Chitta, Rong Jin 0001, Timothy C. Havens, Anil K. Jain 0001
KDD4
2011 A kernel density based approach for large scale image retrieval
abstract
Local image features, such as SIFT descriptors, have been shown to be effective for content-based image retrieval (CBIR). In order to achieve efficient image retrieval using local features, most existing approaches represent an image by a bag-of-words model in which every local feature is quantized into a visual word. Given the bag-of-words representation for images, a text search engine is then used to efficiently find the matched images for a given query. The main drawback with these approaches is that the two key steps, i.e., key point quantization and image matching, are separated, leading to sub-optimal performance in image retrieval. In this work, we present a statistical framework for large-scale image retrieval that unifies key point quantization and image matching by introducing kernel density function. The key ideas of the proposed framework are (a) each image is represented by a kernel density function from which the observed key points are sampled, and (b) the similarity of a gallery image to a query image is estimated as the likelihood of generating the key points in the query image by the kernel density function of the gallery image. We present efficient algorithms for kernel density estimation as well as for effective image matching. Experiments with large-scale image retrieval confirm that the proposed method is not only more effective but also more efficient than the state-of-the-art approaches in identifying visually similar images for given queries from large image databases.
Tianbao Yang, Rong Jin 0001, Anil K. Jain 0001
ICMR5
2011 Fingerprint Reconstruction: From Minutiae to Phase
abstract
Fingerprint matching systems generally use four types of representation schemes: grayscale image, phase image, skeleton image, and minutiae, among which minutiae-based representation is the most widely adopted one. The compactness of minutiae representation has created an impression that the minutiae template does not contain sufficient information to allow the reconstruction of the original grayscale fingerprint image. This belief has now been shown to be false; several algorithms have been proposed that can reconstruct fingerprint images from minutiae templates. These techniques try to either reconstruct the skeleton image, which is then converted into the grayscale image, or reconstruct the grayscale image directly from the minutiae template. However, they have a common drawback: Many spurious minutiae not included in the original minutiae template are generated in the reconstructed image. Moreover, some of these reconstruction techniques can only generate a partial fingerprint. In this paper, a novel fingerprint reconstruction algorithm is proposed to reconstruct the phase image, which is then converted into the grayscale image. The proposed reconstruction algorithm not only gives the whole fingerprint, but the reconstructed fingerprint contains very few spurious minutiae. Specifically, a fingerprint image is represented as a phase image which consists of the continuous phase and the spiral phase (which corresponds to minutiae). An algorithm is proposed to reconstruct the continuous phase from minutiae. The proposed reconstruction algorithm has been evaluated with respect to the success rates of type-I attack (match the reconstructed fingerprint against the original fingerprint) and type-II attack (match the reconstructed fingerprint against different impressions of the original fingerprint) using a commercial fingerprint recognition system. Given the reconstructed image from our algorithm, we show that both types of attacks can be successfully launched against a fingerprint recognition system.
Jianjiang Feng, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Latent Fingerprint Matching
abstract
Latent fingerprint identification is of critical importance to law enforcement agencies in identifying suspects: Latent fingerprints are inadvertent impressions left by fingers on surfaces of objects. While tremendous progress has been made in plain and rolled fingerprint matching, latent fingerprint matching continues to be a difficult problem. Poor quality of ridge impressions, small finger area, and large nonlinear distortion are the main difficulties in latent fingerprint matching compared to plain or rolled fingerprint matching. We propose a system for matching latent fingerprints found at crime scenes to rolled fingerprints enrolled in law enforcement databases. In addition to minutiae, we also use extended features, including singularity, ridge quality map, ridge flow map, ridge wavelength map, and skeleton. We tested our system by matching 258 latents in the NIST SD27 database against a background database of 29,257 rolled fingerprints obtained by combining the NIST SD4, SD14, and SD27 databases. The minutiae-based baseline rank-1 identification rate of 34.9 percent was improved to 74 percent when extended features were used. In order to evaluate the relative importance of each extended feature, these features were incrementally used in the order of their cost in marking by latent experts. The experimental results indicate that singularity, ridge quality map, and ridge flow map are the most effective features in improving the matching accuracy.
Anil K. Jain 0001, Jianjiang Feng
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Matching Forensic Sketches to Mug Shot Photos
abstract
The problem of matching a forensic sketch to a gallery of mug shot images is addressed in this paper. Previous research in sketch matching only offered solutions to matching highly accurate sketches that were drawn while looking at the subject (viewed sketches). Forensic sketches differ from viewed sketches in that they are drawn by a police sketch artist using the description of the subject provided by an eyewitness. To identify forensic sketches, we present a framework called local feature-based discriminant analysis (LFDA). In LFDA, we individually represent both sketches and photos using SIFT feature descriptors and multiscale local binary patterns (MLBP). Multiple discriminant projections are then used on partitioned vectors of the feature-based representation for minimum distance matching. We apply this method to match a data set of 159 forensic sketches against a mug shot gallery containing 10,159 images. Compared to a leading commercial face recognition system, LFDA offers substantial improvements in matching forensic sketches to the corresponding face images. We were able to further improve the matching performance using race and gender information to reduce the target gallery size. Additional experiments demonstrate that the proposed framework leads to state-of-the-art accuracys when matching viewed sketches.
Brendan Klare, Zhifeng Li 0001, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 A Network of Dynamic Probabilistic Models for Human Interaction Analysis
abstract
We propose a novel method of analyzing human interactions based on the walking trajectories of human subjects, which provide elementary and necessary components for understanding and interpretation of complex human interactions in visual surveillance tasks. Our principal assumption is that an interaction episode is composed of meaningful small unit interactions, which we call “sub-interactions”. We model each sub-interaction by a dynamic probabilistic model and propose a modified factorial hidden Markov model (HMM) with factored observations. The complete interaction is represented with a network of dynamic probabilistic models (DPMs) by an ordered concatenation of sub-interaction models. The rationale for this approach is that it is more effective in utilizing common components, i.e., sub-interaction models, to describe complex interaction patterns. By assembling these sub-interaction models in a network, possibly with a mixture of different types of DPMs, such as standard HMMs, variants of HMMs, dynamic Bayesian networks, and so on, we can design a robust model for the analysis of human interactions. We show the feasibility and effectiveness of the proposed method by analyzing the structure of network of DPMs and its success on four different databases: a self-collected dataset, Tsinghua University's dataset, the public domain CAVIAR dataset, and the Edinburgh Informatics Forum Pedestrian dataset.
Heung-Il Suk, Anil K. Jain 0001, Seong-Whan Lee
IEEE Trans. Circuits Syst. Video Technol.2
2011 Restoring Degraded Face Images: A Case Study in Matching Faxed, Printed, and Scanned Photos
abstract
We study the problem of restoring severely degraded face images such as images scanned from passport photos or images subjected to fax compression, downscaling, and printing. The purpose of this paper is to illustrate the complexity of face recognition in such realistic scenarios and to provide a viable solution to it. The contributions of this work are two-fold. First, a database of face images is assembled and used to illustrate the challenges associated with matching severely degraded face images. Second, a preprocessing scheme with low computational complexity is developed in order to eliminate the noise present in degraded images and restore their quality. An extensive experimental study is performed to establish that the proposed restoration scheme improves the quality of the ensuing face images while simultaneously improving the performance of face matching.
Thirimachos Bourlai, Arun Ross, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2011 Separating Overlapped Fingerprints
abstract
Fingerprint images generally contain either a single fingerprint (e.g., rolled images) or a set of nonoverlapped fingerprints (e.g., slap fingerprints). However, there are situations where several fingerprints overlap on top of each other. Such situations are frequently encountered when latent (partial) fingerprints are lifted from crime scenes or residue fingerprints are left on fingerprint sensors. Overlapped fingerprints constitute a serious challenge to existing fingerprint recognition algorithms, since these algorithms are designed under the assumption that fingerprints have been properly segmented. In this paper, a novel algorithm is proposed to separate overlapped fingerprints into component or individual fingerprints. The basic idea is to first estimate the orientation field of the given image with overlapped fingerprints and then separate it into component orientation fields using a relaxation labeling technique. We also propose an algorithm to utilize fingerprint singularity information to further improve the separation performance. Experimental results indicate that the algorithm leads to good separation of overlapped fingerprints that leads to a significant improvement in the matching accuracy.
Fanglin Chen 0001, Jianjiang Feng, Anil K. Jain 0001, Jie Zhou 0001
IEEE Trans. Inf. Forensics Secur.3
2011 A Discriminative Model for Age Invariant Face Recognition
abstract
Aging variation poses a serious problem to automatic face recognition systems. Most of the face recognition studies that have addressed the aging problem are focused on age estimation or aging simulation. Designing an appropriate feature representation and an effective matching framework for age invariant face recognition remains an open problem. In this paper, we propose a discriminative model to address face matching in the presence of age variation. In this framework, we first represent each face by designing a densely sampled local feature description scheme, in which scale invariant feature transform (SIFT) and multi-scale local binary patterns (MLBP) serve as the local descriptors. By densely sampling the two kinds of local descriptors from the entire facial image, sufficient discriminatory information, including the distribution of the edge direction in the face image (that is expected to be age invariant) can be extracted for further analysis. Since both SIFT-based local features and MLBP-based local features span a high-dimensional feature space, to avoid the overfitting problem, we develop an algorithm, called multi-feature discriminant analysis (MFDA) to process these two local feature spaces in a unified framework. The MFDA is an extension and improvement of the LDA using multiple features combined with two different random sampling methods in feature and sample space. By random sampling the training set as well as the feature space, multiple LDA-based classifiers are constructed and then combined to generate a robust decision via a fusion rule. Experimental results show that our approach outperforms a state-of-the-art commercial face recognition engine on two public domain face aging data sets: MORPH and FG-NET. We also compare the performance of the proposed discriminative model with a generative aging model. A fusion of discriminative and generative models further improves the face matching accuracy in the presence of aging.
Zhifeng Li 0001, Unsang Park, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2011 Erratum to "Soft Biometric Traits for Continuous User Authentication"
abstract
In the Acknowledgment for the above paper (ibid., vol. 5, no. 4, pp. 771-780, Dec 2010), due to a production error, the corresponding author's name was spelled incorrectly. The correct spelling is Unsang Park.
Koichiro Niinuma, Unsang Park, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2011 Periocular Biometrics in the Visible Spectrum
abstract
The term periocular refers to the facial region in the immediate vicinity of the eye. Acquisition of the periocular biometric is expected to require less subject cooperation while permitting a larger depth of field compared to traditional ocular biometric traits (viz., iris, retina, and sclera). In this work, we study the feasibility of using the periocular region as a biometric trait. Global and local information are extracted from the periocular region using texture and point operators resulting in a feature set for representing and matching this region. A number of aspects are studied in this work, including the 1) effectiveness of incorporating the eyebrows, 2) use of side information (left or right) in matching, 3) manual versus automatic segmentation schemes, 4) local versus global feature extraction schemes, 5) fusion of face and periocular biometrics, 6) use of the periocular biometric in partially occluded face images, 7) effect of disguising the eyebrows, 8) effect of pose variation and occlusion, 9) effect of masking the iris and eye region, and 10) effect of template aging on matching performance. Experimental results show a rank-one recognition accuracy of 87.32% using 1136 probe and 1136 gallery periocular images taken from 568 different subjects (2 images/subject) in the Face Recognition Grand Challenge (version 2.0) database with the fusion of three different matchers.
Unsang Park, Raghavender R. Jillela, Arun Ross, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.4
2010 Multi task learning on multiple related networks
abstract
With the rapid proliferation of online social networks, the need for newer class of learning algorithm to simultaneously deal with multiple related networks has become increasingly important. This paper proposes an approach for multi-task learning in multiple related networks, where in we perform different tasks such as classification on one network and clustering on the other. We show that the framework can be extended to incorporate prior information about the correspondences between the clusters and classes in different networks. We have performed experiments on real-world data sets to demonstrate the effectiveness of the proposed framework.
Prakash Mandayam Comar, Pang-Ning Tan, Anil K. Jain 0001
CIKM3
2010 Online visual vocabulary pruning using pairwise constraints
abstract
Given a pair of images represented using bag-of-visual-words and a label corresponding to whether the images are “related”(must-link constraint) or “unrelated” (cannot-link constraint), we address the problem of selecting a subset of visual words that are salient in explaining the relation between the image pair. In particular, a subset of features is selected such that the distance computed using these features satisfies the given pairwise constraints. An efficient online feature selection algorithm is presented based on the dual-gradient descent approach. Side information in the form of pair-wise constraints is incorporated into the feature selection stage, providing the user with flexibility to use an unsupervised or semi-supervised algorithm at a later stage. Correlated subsets of visual words, usually resulting from hierarchical quantization process (called groups), are exploited to select a significantly smaller vocabulary. A group-LASSO regularizer is used to drive as many feature weights to zero as possible. We evaluate the quality of the pruned vocabulary by clustering the data using the resulting feature subset. Experiments on PASCAL VOC 2007 dataset using 5000 visual keywords, resulted in around 80% reduction in the number of keywords, with little or no loss in performance.
Pavan Kumar Mallapragada, Rong Jin 0001, Anil K. Jain 0001
CVPR3
2010 Learning from Noisy Side Information by Generalized Maximum Entropy Model
Tianbao Yang, Rong Jin 0001, Anil K. Jain 0001
ICML3
2010 Detecting Altered Fingerprints
abstract
The widespread deployment of Automated Fingerprint Identification Systems (AFIS) in law enforcement and border control applications has prompted some individuals with criminal background to evade identification by purposely altering their fingerprints. Available fingerprint quality assessment software cannot detect most of the altered fingerprints since the implicit image quality does not always degrade due to alteration. In this paper, we classify the alterations observed in an operational database into three categories and propose an algorithm to detect altered fingerprints. Experiments were conducted on both real-world altered fingerprints and synthetically generated altered fingerprints. At a false alarm rate of 7%, the proposed algorithm detected 92% of the altered fingerprints, while a well-known fingerprint quality software, NFIQ, only detected 20% of the altered fingerprints.
Jianjiang Feng, Anil K. Jain 0001, Arun Ross
ICPR2
2010 Heterogeneous Face Recognition: Matching NIR to Visible Light Images
abstract
Matching near-infrared (NIR) face images to visible light (VIS) face images offers a robust approach to face recognition with unconstrained illumination. In this paper we propose a novel method of heterogeneous face recognition that uses a common feature-based representation for both NIR images as well as VIS images. Linear discriminant analysis is performed on a collection of random subspaces to learn discriminative projections. NIR and VIS images are matched (i) directly using the random subspace projections, and (ii) using sparse representation classification. Experimental results demonstrate the effectiveness of the proposed approach for matching NIR and VIS face images.
Brendan Klare, Anil K. Jain 0001
ICPR2
2010 Clustering Face Carvings: Exploring the Devatas of Angkor Wat
abstract
We propose a framework for clustering and visualization of images of face carvings at archaeological sites. The pairwise similarities among face carvings are computed by performing Procrustes analysis on local facial features (eyes, nose, mouth, etc.). The distance between corresponding face features is computed using point distribution models; the final pairwise similarity is the weighted sum of feature similarities. A web-based interface is provided to allow domain experts to interactively assign different weights to each face feature, and display hierarchical clustering results in 2D or 3D projections obtained by multidimensional scaling. The proposed framework has been successfully applied to the devata goddesses depicted in the ancient Angkor Wat temple. The resulting clusterings and visualization will enable a systematic anthropological, ethnological and artistic analysis of nearly 1,800 stone portraits of devatas of Angkor Wat.
Brendan Klare, Pavan Kumar Mallapragada, Anil K. Jain 0001, Kent Davis
ICPR3
2010 Unsupervised Ensemble Ranking: Application to Large-Scale Image Retrieval
abstract
The continued explosion in the growth of image and video databases makes automatic image search and retrieval an extremely important problem. Among the various approaches to Content-based Image Retrieval (CBIR), image similarity based on local point descriptors has shown promising performance. However, this approach suffers from the scalability problem. Although bag-of-words model resolves the scalability problem, it suffers from loss in retrieval accuracy. We circumvent this performance loss by an ensemble ranking approach in which rankings from multiple bag-of-words models are combined to obtain more accurate retrieval results. An unsupervised algorithm is developed to learn the weights for fusing the rankings from multiple bag-of-words models. Experimental results on a database of 100,000 images show that this approach is both efficient and effective in finding visually similar images.
Rong Jin 0001, Anil K. Jain 0001
ICPR3
2010 PILL-ID: Matching and Retrieval of Drug Pill Imprint Images
abstract
Automatic illicit drug pill matching and retrieval is becoming an important problem due to an increase in the number of tablet type illicit drugs being circulated in our society. We propose an automatic method to match drug pill images based on the imprints appearing on the tablet. This will help identify the source and manufacturer of the illicit drugs. The feature vector extracted from tablet images is based on edge localization and invariant moments. Instead of storing a single template for each pill type, we generate multiple templates during the edge detection process. This circumvents the difficulties during matching due to variations in illumination and viewpoint. Experimental results using a set of real drug pill images (822 illicit drug pill images and 1,294 legal drug pill images) showed 76.74% (93.02%) rank one (rank-20) matching accuracy.
Young-Beom Lee, Unsang Park, Anil K. Jain 0001
ICPR3
2010 On the Scalability of Evidence Accumulation Clustering
abstract
This work focuses on the scalability of the Evidence Accumulation Clustering (EAC) method. We first address the space complexity of the co-association matrix. The sparseness of the matrix is related to the construction of the clustering ensemble. Using a split and merge strategy combined with a sparse matrix representation, we empirically show that a linear space complexity is achievable in this framework, leading to the scalability of EAC method to clustering large data-sets.
André Lourenço, Ana Fred, Anil K. Jain 0001
ICPR3
2010 Automated Gland Segmentation and Classification for Gleason Grading of Prostate Tissue Images
abstract
The well-known Gleason grading method for an H&E prostatic carcinoma tissue image uses morphological features of histology patterns within a tissue slide to classify it into 5 grades. We have developed an automated gland segmentation and classification method that will be used for automated Gleason grading of a prostatic carcinoma tissue image. We demonstrate the performance of the proposed classification system for a three-class classification problem (benign, grade 3 carcinoma and grade 4 carcinoma) on a dataset containing 78 tissue images and achieve a classification accuracy of 88.84%. In comparison to the other segmentation-based methods, our approach combines the similarity of morphological patterns associated with a grade with the domain knowledge such as the appearance of nuclei and blue mucin for the grading task.
Kien Nguyen 0005, Anil K. Jain 0001, Ronald L. Allen
ICPR2
2010 Unsupervised transfer classification: application to text categorization
abstract
We study the problem of building the classification model for a target class in the absence of any labeled training example for that class. To address this difficult learning problem, we extend the idea of transfer learning by assuming that the following side information is available: (i) a collection of labeled examples belonging to other classes in the problem domain, called the auxiliary classes; (ii) the class information including the prior of the target class and the correlation between the target class and the auxiliary classes. Our goal is to construct the classification model for the target class by leveraging the above data and information. We refer to this learning problem as unsupervised transfer classification. Our framework is based on the generalized maximum entropy model that is effective in transferring the label information of the auxiliary classes to the target class. A theoretical analysis shows that under certain assumption, the classification model obtained by the proposed approach converges to the optimal model when it is learned from the labeled examples for the target class. Empirical study on text categorization over four different data sets verifies the effectiveness of the proposed approach.
Tianbao Yang, Rong Jin 0001, Anil K. Jain 0001, Yang Zhou 0033
KDD3
2010 Multi-label Multiple Kernel Learning by Stochastic Approximation: Application to Visual Object Recognition
abstract
Recent studies have shown that multiple kernel learning is very effective for object recognition, leading to the popularity of kernel learning in computer vision problems. In this work, we develop an efficient algorithm for multi-label multiple kernel learning (ML-MKL). We assume that all the classes under consideration share the same combination of kernel functions, and the objective is to find the optimal kernel combination that benefits all the classes. Although several algorithms have been developed for ML-MKL, their computational cost is linear in the number of classes, making them unscalable when the number of classes is large, a challenge frequently encountered in visual object recognition. We address this computational challenge by developing a framework for ML-MKL that combines the worst-case analysis with stochastic approximation. Our analysis shows that the complexity of our algorithm is $O(m^{1/3}\sqrt{ln m})$, where $m$ is the number of classes. Empirical studies with object recognition show that while achieving similar classification accuracy, the proposed method is significantly more efficient than the state-of-the-art algorithms for ML-MKL.
Serhat Selcuk Bucak, Rong Jin 0001, Anil K. Jain 0001
NIPS3
2010 Identifying Cohesive Subgroups and Their Correspondences in Multiple Related Networks
abstract
Identifying cohesive subgroups in networks, also known as clustering is an active area of research in link mining with many practical applications. However, most of the early work in this area has focused on partitioning a single network or a bipartite graph into clusters/communities. This paper presents a framework that simultaneously clusters nodes from multiple related networks and learns the correspondences between subgroups in different networks. The framework also allows the incorporation of prior information about potential relationships between the subgroups. We have performed extensive experiments on both synthetic and real-life data sets to evaluate the effectiveness of our framework. Our results show superior performance of simultaneous clustering over independent clustering of individual networks.
Prakash Mandayam Comar, Pang-Ning Tan, Anil K. Jain 0001
Web Intelligence3
2010 Age-Invariant Face Recognition
abstract
One of the challenges in automatic face recognition is to achieve temporal invariance. In other words, the goal is to come up with a representation and matching scheme that is robust to changes due to facial aging. Facial aging is a complex process that affects both the 3D shape of the face and its texture (e.g., wrinkles). These shape and texture changes degrade the performance of automatic face recognition systems. However, facial aging has not received substantial attention compared to other facial variations due to pose, lighting, and expression. We propose a 3D aging modeling technique and show how it can be used to compensate for the age variations to improve the face recognition performance. The aging modeling technique adapts view-invariant 3D face models to the given 2D face aging database. The proposed approach is evaluated on three different databases (i.g., FG-NET, MORPH, and BROWNS) using FaceVACS, a state-of-the-art commercial face recognition engine.
Unsang Park, Yiying Tong, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Data clustering: 50 years beyond K-means
Anil K. Jain 0001
Pattern Recognit. Lett.1
2010 A hybrid biometric cryptosystem for securing fingerprint minutiae templates
Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001
Pattern Recognit. Lett.3
2010 A hybrid approach for generating secure and discriminating face template
abstract
Biometric template protection is one of the most important issues in deploying a practical biometric system. To tackle this problem, many algorithms, that do not store the template in its original form, have been reported in recent years. They can be categorized into two approaches, namely biometric cryptosystem and transform-based. However, most (if not all) algorithms in both approaches offer a trade-off between the template security and matching performance. Moreover, we believe that no single template protection method is capable of satisfying the security and performance simultaneously. In this paper, we propose a hybrid approach which takes advantage of both the biometric cryptosystem approach and the transform-based approach. A three-step hybrid algorithm is designed and developed based on random projection, discriminability-preserving (DP) transform, and fuzzy commitment scheme. The proposed algorithm not only provides good security, but also enhances the performance through the DP transform. Three publicly available face databases, namely FERET, CMU-PIE, and FRGC, are used for evaluation. The security strength of the binary templates generated from FERET, CMU-PIE, and FRGC databases are 206.3, 203.5, and 347.3 bits, respectively. Moreover, noninvertibility analysis and discussion on data leakage of the proposed hybrid algorithm are also reported. Experimental results show that, using Fisherface to construct the input facial feature vector (face template), the proposed hybrid method can improve the recognition accuracy by 4%, 11%, and 15% on the FERET, CMU-PIE, and FRGC databases, respectively. A comparison with the recently developed random multispace quantization biohashing algorithm is also reported.
Yi C. Feng, Pong C. Yuen, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2010 Soft Biometric Traits for Continuous User Authentication
abstract
Most existing computer and network systems authenticate a user only at the initial login session. This could be a critical security weakness, especially for high-security systems because it enables an impostor to access the system resources until the initial user logs out. This situation is encountered when the logged in user takes a short break without logging out or an impostor coerces the valid user to allow access to the system. To address this security flaw, we propose a continuous authentication scheme that continuously monitors and authenticates the logged in user. Previous methods for continuous authentication primarily used hard biometric traits, specifically fingerprint and face to continuously authenticate the initial logged in user. However, the use of these biometric traits is not only inconvenient to the user, but is also not always feasible due to the user's posture in front of the sensor. To mitigate this problem, we propose a new framework for continuous user authentication that primarily uses soft biometric traits (e.g., color of user's clothing and facial skin). The proposed framework automatically registers (enrolls) soft biometric traits every time the user logs in and fuses soft biometric matching with the conventional authentication schemes, namely password and face biometric. The proposed scheme has high tolerance to the user's posture in front of the computer system. Experimental results show the effectiveness of the proposed method for continuous user authentication.
Koichiro Niinuma, Unsang Park, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2010 Face matching and retrieval using soft biometrics
abstract
Soft biometric traits embedded in a face (e.g., gender and facial marks) are ancillary information and are not fully distinctive by themselves in face-recognition tasks. However, this information can be explicitly combined with face matching score to improve the overall face-recognition accuracy. Moreover, in certain application domains, e.g., visual surveillance, where a face image is occluded or is captured in off-frontal pose, soft biometric traits can provide even more valuable information for face matching or retrieval. Facial marks can also be useful to differentiate identical twins whose global facial appearances are very similar. The similarities found from soft biometrics can also be useful as a source of evidence in courts of law because they are more descriptive than the numerical matching scores generated by a traditional face matcher. We propose to utilize demographic information (e.g., gender and ethnicity) and facial marks (e.g., scars, moles, and freckles) for improving face image matching and retrieval performance. An automatic facial mark detection method has been developed that uses (1) the active appearance model for locating primary facial features (e.g., eyes, nose, and mouth), (2) the Laplacian-of-Gaussian blob detection, and (3) morphological operators. Experimental results based on the FERET database (426 images of 213 subjects) and two mugshot databases from the forensic domain (1225 images of 671 subjects and 10 000 images of 10 000 subjects, respectively) show that the use of soft biometric traits is able to improve the face-recognition performance of a state-of-the-art commercial matcher.
Unsang Park, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.2
2009 A co-classification framework for detecting web spam and spammers in social media web sites
abstract
Social media are becoming increasingly popular and have attracted considerable attention from spammers. Using a sample of more than ninety thousand known spam Web sites, we found between 7% to 18% of their URLs are posted on two popular social media Web sites, digg.com and delicious.com. In this paper, we present a co-classification framework to detect Web spam and the spammers who are responsible for posting them on the social media Web sites. The rationale for our approach is that since both detection tasks are related, it would be advantageous to train them simultaneously to make use of the labeled examples in the Web spam and spammer training data. We have evaluated the effectiveness of our algorithm on the delicious.com data set. Our experimental results showed that the proposed co-classification algorithm significantly outperforms classifiers that learn each detection task independently.
Pang-Ning Tan, Anil K. Jain 0001
CIKM3
2009 Efficient multi-label ranking for multi-class learning: Application to object recognition
abstract
Multi-label learning is useful in visual object recognition when several objects are present in an image. Conventional approaches implement multi-label learning as a set of binary classification problems, but they suffer from imbalanced data distributions when the number of classes is large. In this paper, we address multi-label learning with many classes via a ranking approach, termed multi-label ranking. Given a test image, the proposed scheme aims to order all the object classes such that the relevant classes are ranked higher than the irrelevant ones. We present an efficient algorithm for multi-label ranking based on the idea of block coordinate descent. The proposed algorithm is applied to visual object recognition. Empirical results on the PASCAL VOC 2006 and 2007 data sets show promising results in comparison to the state-of-the-art algorithms for multi-label learning.
Serhat Selcuk Bucak, Pavan Kumar Mallapragada, Rong Jin 0001, Anil K. Jain 0001
ICCV4
2009 Content-based image retrieval: An application to tattoo images
abstract
Tattoo images on human body have been routinely collected and used in law enforcement to assist in suspect and victim identification. However, the current practice of matching tattoos is based on keywords. Assigning keywords to individual tattoo images is both tedious and subjective. We have developed a content-based image retrieval system for a tattoo image database. The system automatically extracts image features based on the Scale Invariant Feature Transform (SIFT). Side information, i.e., body location of tattoos and tattoo classes, is utilized to improve the retrieval time and retrieval accuracy. Geometrical constraints are also introduced in SIFT keypoint matching to reduce false retrievals. Experimental results on 1,000 queries against an operational database of 63,593 tattoo images show a rank-20 accuracy of 94.2%; the average matching time per query is 2.9 sec. on Intel Core 2, 2.66 GHz, 3 GB RAM processor.
Anil K. Jain 0001, Rong Jin 0001, Nicholas Gregg
ICIP1
2009 Facial marks: Soft biometric for face recognition
abstract
We propose to utilize micro features, namely facial marks (e.g., freckles, moles, and scars) to improve face recognition and retrieval performance. Facial marks can be used in three ways: i) to supplement the features in an existing face matcher, ii) to enable fast retrieval from a large database using facial mark based queries, and iii) to enable matching or retrieval from a partial or profile face image with marks. We use Active Appearance Model (AAM) to locate and segment primary facial features (e.g., eyes, nose, and mouth). Then, Laplacian-of-Gaussian (LoG) and morphological operators are used to detect facial marks. Experimental results based on FERET (426 images, 213 subjects) and Mugshot (1,225 images, 671 subjects) databases show that the use of facial marks improves the rank-1 identification accuracy of a state-of-the-art face recognition system from 92.96% to 93.90% and from 91.88% to 93.14%, respectively.
Anil K. Jain 0001, Unsang Park
ICIP1
2009 Latent Palmprint Matching
abstract
The evidential value of palmprints in forensic applications is clear as about 30 percent of the latents recovered from crime scenes are from palms. While biometric systems for palmprint-based personal authentication in access control type of applications have been developed, they mostly deal with low-resolution (about 100 ppi) palmprints and only perform full-to-full palmprint matching. We propose a latent-to-full palmprint matching system that is needed in forensic applications. Our system deals with palmprints captured at 500 ppi (the current standard in forensic applications) or higher resolution and uses minutiae as features to be compatible with the methodology used by latent experts. Latent palmprint matching is a challenging problem because latent prints lifted at crime scenes are of poor image quality, cover only a small area of the palm, and have a complex background. Other difficulties include a large number of minutiae in full prints (about 10 times as many as fingerprints), and the presence of many creases in latents and full prints. A robust algorithm to reliably estimate the local ridge direction and frequency in palmprints is developed. This facilitates the extraction of ridge and minutiae features even in poor quality palmprints. A fixed-length minutia descriptor, MinutiaCode, is utilized to capture distinctive information around each minutia and an alignment-based minutiae matching algorithm is used to match two palmprints. Two sets of partial palmprints (150 live-scan partial palmprints and 100 latent palmprints) are matched to a background database of 10,200 full palmprints to test the proposed system. Despite the inherent difficulty of latent-to-full palmprint matching, rank-1 recognition rates of 78.7 and 69 percent, respectively, were achieved in searching live-scan partial palmprints and latent palmprints against the background database.
Anil K. Jain 0001, Jianjiang Feng
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 SemiBoost: Boosting for Semi-Supervised Learning
abstract
Semi-supervised learning has attracted a significant amount of attention in pattern recognition and machine learning. Most previous studies have focused on designing special algorithms to effectively exploit the unlabeled data in conjunction with labeled data. Our goal is to improve the classification accuracy of any given supervised learning algorithm by using the available unlabeled examples. We call this as the Semi-supervised improvement problem, to distinguish the proposed approach from the existing approaches. We design a metasemi-supervised learning algorithm that wraps around the underlying supervised algorithm and improves its performance using unlabeled data. This problem is particularly important when we need to train a supervised learning algorithm with a limited number of labeled examples and a multitude of unlabeled examples. We present a boosting framework for semi-supervised learning, termed as SemiBoost. The key advantages of the proposed semi-supervised learning approach are: 1) performance improvement of any supervised learning algorithm with a multitude of unlabeled data, 2) efficient computation by the iterative boosting algorithm, and 3) exploiting both manifold and cluster assumption in training classification models. An empirical study on 16 different data sets and text categorization demonstrates that the proposed framework improves the performance of several commonly used supervised learning algorithms, given a large number of unlabeled examples. We also show that the performance of the proposed algorithm, SemiBoost, is comparable to the state-of-the-art semi-supervised learning algorithms.
Pavan Kumar Mallapragada, Rong Jin 0001, Anil K. Jain 0001, Yi Liu 0054
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Rank-based distance metric learning: An application to image retrieval
abstract
We present a novel approach to learn distance metric for information retrieval. Learning distance metric from a number of queries with side information, i.e., relevance judgements, has been studied widely, for example pairwise constraint-based distance metric learning. However, the capacity of existing algorithms is limited, because they usually assume that the distance between two similar objects is smaller than the distance between two dissimilar objects. This assumption may not hold, especially in the case of information retrieval when the input space is heterogeneous. To address this problem explicitly, we propose rank-based distance metric learning. Our approach overcomes the drawback of existing algorithms by comparing the distances only among the relevant and irrelevant objects for a given query. To avoid over-fitting, a regularizer based on the Burg matrix divergence is also introduced. We apply the proposed framework to tattoo image retrieval in forensics and law enforcement application domain. The goal of the application is to retrieve tattoo images from a gallery database that are visually similar to a tattoo found on a suspect or a victim. The experimental results show encouraging results in comparison to the standard approaches for distance metric learning.
Rong Jin 0001, Anil K. Jain 0001
CVPR3
2008 Face recognition with temporal invariance: A 3D aging model
abstract
The variation caused by aging has not received adequate attention compared with pose, lighting, and expression variations. Aging is a complex process that affects both the 3D shape of the face and its texture (e.g., wrinkles). While the facial age modeling has been widely studied in computer graphics community, only a few studies have been reported in computer vision literature on age-invariant face recognition. We propose an automatic aging simulation technique that can assist any existing face recognition engine for aging-invariant face recognition. We learn the aging patterns of shape and the corresponding texture in 3D domain by adapting a 3D morphable model to the 2D aging database (public domain FG-NET). At recognition time, each probe and all gallery images are modified to compensate for the age-induced variation using an intermediate 3D model deformation and a texture modification, prior to matching. The proposed approach is evaluated on a set of age-separated probe and gallery data using a state-of-the-art commercial face recognition engine, FaceVACS. Use of 3D aging model improves the rank-1 matching accuracy on FG-NET database from 28.0% to 37.8%, on average.
Unsang Park, Yiying Tong, Anil K. Jain 0001
FG3
2008 Filtering large fingerprint database for latent matching
abstract
Latent fingerprint identification is of critical importance to law enforcement agencies in apprehending criminals. Considering the huge size of fingerprint databases maintained by law enforcement agencies, exhaustive one-to-one matching is impractical and a database filtering technique is necessary to reduce the search space. Due to low image quality and small finger area of latent fingerprints, it is necessary to use several features for an efficient and reliable filtering system. A multi-stage filtering system is proposed, which utilizes pattern type, singular points and orientation field. We have tested our system by searching 258 latent fingerprints in NIST SD27 against a background database containing 10,258 rolled fingerprints (obtained by combining 2,000 in NIST SD4, 8,000 in SD14 and 258 in SD27). Although latent fingerprints contain very limited information, the filtering system not only improved the matching speed by three fold but also improved the rank-1 matching accuracy from 70.9% to 73.3%.
Jianjiang Feng, Anil K. Jain 0001
ICPR2
2008 Cluster validation using a probabilistic attributed graph
abstract
We propose a new cluster validity index. A data partition is described by a set of disjoint sub-graphs, each corresponding to the minimum spanning tree of a cluster, taking as edge weight the dissimilarity between linked objects. Based on the assumption that each cluster has a characteristic parametric distribution of dissimilarity increments, graph probabilities are estimated. The validity index is defined as the minimum description length for both estimated model parameters and data partition, according to this probabilistic model. This new index can be used to evaluate various partitions of a given data set obtained by: (i) a single clustering algorithm, (ii) different clustering algorithms, or (iii) cluster ensemble methods. Experimental evaluation of the proposed index on synthetic and real data taken from the UCI repository confirms the usefulness of the method in selecting good clustering solutions.
Ana Fred, Anil K. Jain 0001
ICPR2
2008 Active query selection for semi-supervised clustering
abstract
Semi-supervised clustering allows a user to specify available prior knowledge about the data to improve the clustering performance. A common way to express this information is in the form of pair-wise constraints. A number of studies have shown that, in general, these constraints improve the resulting data partition. However, the choice of constraints is critical since improperly chosen constraints might actually degrade the clustering performance. We focus on constraint (also known as query) selection for improving the performance of semi-supervised clustering algorithms. We present an active query selection mechanism, where the queries are selected using a min-max criterion. Experimental results on a variety of datasets, using MPCK-means as the underlying semi-clustering algorithm, demonstrate the superior performance of the proposed query selection procedure.
Pavan Kumar Mallapragada, Rong Jin 0001, Anil K. Jain 0001
ICPR3
2008 Securing fingerprint template: Fuzzy vault with minutiae descriptors
abstract
Fuzzy vault has been shown to be an effective technique for securing fingerprint minutiae templates. Its security depends on the difficulty in identifying the set of genuine minutiae points among a mixture of genuine and chaff points and reconstructing the secure polynomial using the evaluations (ordinate values) available for each point in the vault. We show that the security of fuzzy vault can be improved by ldquoencryptingrdquo these polynomial evaluations using a fuzzy commitment scheme. This encryption makes it difficult for an adversary to decode the vault even if the correct set of minutiae is selected. We use minutiae descriptors, which capture orientation and ridge frequency information in a minutiapsilas neighborhood, for securing the polynomial evaluations. This modification leads to a significant increase in both the security (number of tries an adversary has to make in order to guess the secure key) and matching accuracy of the vault. We validate our results on FVC2002 DB2 and show that false accept rate (FAR) is reduced from 0.7% to 0.01% at a genuine accept rate (GAR) of 95%. At the same time, vault security as measured in terms of min-entropy, is increased from 31 bits to 47 bits in case a perfect code is used.
Abhishek Nagar, Karthik Nandakumar, Anil K. Jain 0001
ICPR3
2008 Data Clustering: 50 Years Beyond K-means
Anil K. Jain 0001
ECML/PKDD (1)1
2008 Semi-Supervised Boosting for Multi-Class Classification
Hamed Valizadegan, Rong Jin 0001, Anil K. Jain 0001
ECML/PKDD (2)3
2008 Deformation Modeling for Robust 3D Face Matching
abstract
Face recognition based on 3D surface matching is promising for overcoming some of the limitations of current 2D image-based face recognition systems. The 3D shape is generally invariant to the pose and lighting changes, but not invariant to the nonrigid facial movement such as expressions. Collecting and storing multiple templates to account for various expressions for each subject in a large database is not practical. We propose a facial surface modeling and matching scheme to match 2.5D facial scans in the presence of both nonrigid deformations and pose changes (multiview) to a stored 3D face model with neutral expression. A hierarchical geodesic-based resampling approach is applied to extract landmarks for modeling facial surface deformations. We are able to synthesize the deformation learned from a small group of subjects (control group) onto a 3D neutral model (not in the control group), resulting in a deformed template. A user-specific (3D) deformable model is built for each subject in the gallery with respect to the control group by combining the templates with synthesized deformations. By fitting this generative deformable model to a test scan, the proposed approach is able to handle expressions and pose changes simultaneously. A fully automatic and prototypic deformable model based 3D face matching system has been developed. Experimental results demonstrate that the proposed deformation modeling scheme increases the 3D face matching accuracy in comparison to matching with 3D neutral models by 7 and 10 percentage points, respectively, on a subset of the FRGC v2.0 3D benchmark and the MSU multiview 3D face database with expression variations.
Xiaoguang Lu, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Likelihood Ratio-Based Biometric Score Fusion
abstract
Multibiometric systems fuse information from different sources to compensate for the limitations in performance of individual matchers. We propose a framework for optimal combination of match scores that is based on the likelihood ratio test. The distributions of genuine and impostor match scores are modeled as finite Gaussian mixture model. The proposed fusion approach is general in its ability to handle (i) discrete values in biometric match score distributions, (ii) arbitrary scales and distributions of match scores, (iii) correlation between the scores of multiple matchers and (iv) sample quality of multiple biometric sources. Experiments on three multibiometric databases indicate that the proposed fusion framework achieves consistently high performance compared to commonly used score fusion techniques based on score transformation and classification.
Karthik Nandakumar, Yi Chen 0015, Sarat C. Dass, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Biometric Recognition: Overview and Recent Advances
Anil K. Jain 0001
CIARP1
2007 Face Recognition in Video: Adaptive Fusion of Multiple Matchers
abstract
Face recognition in video is being actively studied as a covert method of human identification in surveillance systems. Identifying human faces in video is a difficult problem due to the presence of large variations in facial pose and lighting, and poor image resolution. However, by taking advantage of the diversity of the information contained in video, the performance of a face recognition system can be enhanced. In this work we explore (a) the adaptive use of multiple face matchers in order to enhance the performance of face recognition in video, and (b) the possibility of appropriately populating the database (gallery) in order to succinctly capture intra class variations. To extract the dynamic information in video, the facial poses in various frames are explicitly estimated using active appearance model (AAM) and a factorization based 3D face reconstruction technique. We also estimate the motion blur using discrete cosine transformation (DCT). Our experimental results on 204 subjects in CMU's face-in-action (FIA) database show that the proposed recognition method provides consistent improvements in the matching performance using three different face matchers.
Unsang Park, Anil K. Jain 0001, Arun Ross
CVPR2
2007 BoostCluster: boosting clustering by pairwise constraints
abstract
Data clustering is an important task in many disciplines. A large number of studies have attempted to improve clustering by using the side information that is often encoded as pairwise constraints. However, these studies focus on designing special clustering algorithms that can effectively exploit the pairwise constraints. We present a boosting framework for data clustering,termed as BoostCluster, that is able to iteratively improve the accuracy of any given clustering algorithm by exploiting the pairwise constraints. The key challenge in designing a boosting framework for data clustering is how to influence an arbitrary clustering algorithm with the side information since clustering algorithms by definition are unsupervised. The proposed framework addresses this problem by dynamically generating new data representations at each iteration that are, on the one hand, adapted to the clustering results at previous iterations by the given algorithm, and on the other hand consistent with the given side information. Our empirical study shows that the proposed boosting framework is effective in improving the performance of a number of popular clustering algorithms (K-means, partitional SingleLink, spectral clustering), and its performance is comparable to the state-of-the-art algorithms for data clustering with side information.
Yi Liu 0054, Rong Jin 0001, Anil K. Jain 0001
KDD3
2007 Pores and Ridges: High-Resolution Fingerprint Matching Using Level 3 Features
abstract
Fingerprint friction ridge details are generally described in a hierarchical order at three different levels, namely, Level 1 (pattern), Level 2 (minutia points), and Level 3 (pores and ridge contours). Although latent print examiners frequently take advantage of Level 3 features to assist in identification, Automated Fingerprint Identification Systems (AFIS) currently rely only on Level 1 and Level 2 features. In fact, the Federal Bureau of Investigation's (FBI) standard of fingerprint resolution for AFIS is 500 pixels per inch (ppi), which is inadequate for capturing Level 3 features, such as pores. With the advances in fingerprint sensing technology, many sensors are now equipped with dual resolution (500 ppi/1,000 ppi) scanning capability. However, increasing the scan resolution alone does not necessarily provide any performance improvement in fingerprint matching, unless an extended feature set is utilized. As a result, a systematic study to determine how much performance gain one can achieve by introducing Level 3 features in AFIS is highly desired. We propose a hierarchical matching system that utilizes features at all the three levels extracted from 1,000 ppi fingerprint scans. Level 3 features, including pores and ridge contours, are automatically extracted using Gabor filters and wavelet transform and are locally matched using the Iterative Closest Point (ICP) algorithm. Our experiments show that Level 3 features carry significant discriminatory information. There is a relative reduction of 20 percent in the equal error rate (EER) of the matching system when Level 3 features are employed in combination with Level 1 and 2 features. This significant performance gain is consistently observed across various quality fingerprint images.
Anil K. Jain 0001, Yi Chen 0015, Meltem Demirkus
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 From Template to Image: Reconstructing Fingerprints from Minutiae Points
abstract
Most fingerprint-based biometric systems store the minutiae template of a user in the database. It has been traditionally assumed that the minutiae template of a user does not reveal any information about the original fingerprint. In this paper, we challenge this notion and show that three levels of information about the parent fingerprint can be elicited from the minutiae template alone, viz., 1) the orientation field information, 2) the class or type information, and 3) the friction ridge structure. The orientation estimation algorithm determines the direction of local ridges using the evidence of minutiae triplets. The estimated orientation field, along with the given minutiae distribution, is then used to predict the class of the fingerprint. Finally, the ridge structure of the parent fingerprint is generated using streamlines that are based on the estimated orientation field. Line Integral Convolution is used to impart texture to the ensuing ridges, resulting in a ridge map resembling the parent fingerprint. The salient feature of this noniterative method to generate ridges is its ability to preserve the minutiae at specified locations in the reconstructed ridge map. Experiments using a commercial fingerprint matcher suggest that the reconstructed ridge structure bears close resemblance to the parent fingerprint.
Arun Ross, Jidnya Shah, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Fingerprint-Based Fuzzy Vault: Implementation and Performance
abstract
Reliable information security mechanisms are required to combat the rising magnitude of identity theft in our society. While cryptography is a powerful tool to achieve information security, one of the main challenges in cryptosystems is to maintain the secrecy of the cryptographic keys. Though biometric authentication can be used to ensure that only the legitimate user has access to the secret keys, a biometric system itself is vulnerable to a number of threats. A critical issue in biometric systems is to protect the template of a user which is typically stored in a database or a smart card. The fuzzy vault construct is a biometric cryptosystem that secures both the secret key and the biometric template by binding them within a cryptographic framework. We present a fully automatic implementation of the fuzzy vault scheme based on fingerprint minutiae. Since the fuzzy vault stores only a transformed version of the template, aligning the query fingerprint with the template is a challenging task. We extract high curvature points derived from the fingerprint orientation field and use them as helper data to align the template and query minutiae. The helper data itself do not leak any information about the minutiae template, yet contain sufficient information to align the template and query fingerprints accurately. Further, we apply a minutiae matcher during decoding to account for nonlinear distortion and this leads to significant improvement in the genuine accept rate. We demonstrate the performance of the vault implementation on two different fingerprint databases. We also show that performance improvement can be achieved by using multiple fingerprint impressions during enrollment and verification.
Karthik Nandakumar, Anil K. Jain 0001, Sharath Pankanti
IEEE Trans. Inf. Forensics Secur.2
2007 Statistical Models for Assessing the Individuality of Fingerprints
abstract
Following the Daubert ruling in 1993, forensic evidence based on fingerprints was first challenged in the 1999 case of the U.S. versus Byron C. Mitchell and, subsequently, in 20 other cases involving fingerprint evidence. The main concern with the admissibility of fingerprint evidence is the problem of individualization, namely, that the fundamental premise for asserting the uniqueness of fingerprints has not been objectively tested and matching error rates are unknown. In order to assess the error rates, we require quantifying the variability of fingerprint features, namely, minutiae in the target population. A family of finite mixture models has been developed in this paper to represent the distribution of minutiae in fingerprint images, including minutiae clustering tendencies and dependencies in different regions of the fingerprint image domain. A mathematical model that computes the probability of a random correspondence (PRC) is derived based on the mixture models. A PRC of 2.25$\,\times 10^{-6}$corresponding to 12 minutiae matches was computed for the NIST4 Special Database, when the numbers of query and template minutiae both equal 46. This is also the estimate of the PRC for a target population with a similar composition as that of NIST4.
Yongfang Zhu, Sarat C. Dass, Anil K. Jain 0001
IEEE Trans. Inf. Forensics Secur.3
2006 Deformation Modeling for Robust 3D Face Matching
abstract
Human face recognition based on 3D surface matching is promising for overcoming the limitations of current 2D image-based face recognition systems. The 3D shape is invariant to the pose and lighting changes, but not invariant to the non-rigid facial movement, such as expressions. Collecting and storing multiple templates for each subject in a large database (associated with various expressions) is not practical. We present a facial surface modeling and matching scheme to match 2.5D test scans in the presence of both non-rigid deformations and large pose changes (multiview) to a neutral expression 3D face model. A geodesic-based resampling approach is applied to extract landmarks for modeling facial surface deformations. We are able to synthesize the deformation learned from a small group of subjects (control group) onto a 3D neutral model (not in the control group), resulting in a deformed template. A personspecific (3D) deformable model is built for each subject in the gallery w.r.t. the control group by combining the templates with synthesized deformations. By fitting this generative deformable model to a test scan, the proposed approach is able to handle expressions and large pose changes simultaneously. Experimental results demonstrate that the proposed matching scheme based on deformation modeling improves the matching accuracy.
Xiaoguang Lu, Anil K. Jain 0001
CVPR (2)2
2006 Performance Evaluation of Fingerprint Verification Systems
abstract
This paper is concerned with the performance evaluation of fingerprint verification systems. After an initial classification of biometric testing initiatives, we explore both the theoretical and practical issues related to performance evaluation by presenting the outcome of the recent Fingerprint Verification Competition (FVC2004). FVC2004 was organized by the authors of this work for the purpose of assessing the state-of-the-art in this challenging pattern recognition application and making available a new common benchmark for an unambiguous comparison of fingerprint-based biometric systems. FVC2004 is an independent, strongly supervised evaluation performed at the evaluators' site on evaluators' hardware. This allowed the test to be completely controlled and the computation times of different algorithms to be fairly compared. The experience and feedback received from previous, similar competitions (FVC2000 and FVC2002) allowed us to improve the organization and methodology of FVC2004 and to capture the attention of a significantly higher number of academic and commercial organizations (67 algorithms were submitted for FVC2004). A new, "Light" competition category was included to estimate the loss of matching performance caused by imposing computational constraints. This paper discusses data collection and testing protocols, and includes a detailed analysis of the results. We introduce a simple but effective method for comparing algorithms at the score level, allowing us to isolate difficult cases (images) and to study error correlations and algorithm "fusion." The huge amount of information obtained, including a structured classification of the submitted algorithms on the basis of their features, makes it possible to better understand how current fingerprint recognition systems work and to delineate useful research directions for the future.
Raffaele Cappelli, Dario Maio, Davide Maltoni, James L. Wayman, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2006 Validating a Biometric Authentication System: Sample Size Requirements
abstract
Authentication systems based on biometric features (e.g., fingerprint impressions, iris scans, human face images, etc.) are increasingly gaining widespread use and popularity. Often, vendors and owners of these commercial biometric systems claim impressive performance that is estimated based on some proprietary data. In such situations, there is a need to independently validate the claimed performance levels. System performance is typically evaluated by collecting biometric templates from n different subjects, and for convenience, acquiring multiple instances of the biometric for each of the n subjects. Very little work has been done in 1) constructing confidence regions based on the ROC curve for validating the claimed performance levels and 2) determining the required number of biometric samples needed to establish confidence regions of prespecified width for the ROC curve. To simplify the analysis that address these two problems, several previous studies have assumed that multiple acquisitions of the biometric entity are statistically independent. This assumption is too restrictive and is generally not valid. We have developed a validation technique based on multivariate copula models for correlated biometric acquisitions. Based on the same model, we also determine the minimum number of samples required to achieve confidence bands of desired width for the ROC curve. We illustrate the estimation of the confidence bands as well as the required number of biometric samples using a fingerprint matching system that is applied on samples collected from a small population.
Sarat C. Dass, Yongfang Zhu, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Incremental Nonlinear Dimensionality Reduction by Manifold Learning
abstract
Understanding the structure of multidimensional patterns, especially in unsupervised cases, is of fundamental importance in data mining, pattern recognition, and machine learning. Several algorithms have been proposed to analyze the structure of high-dimensional data based on the notion of manifold learning. These algorithms have been used to extract the intrinsic characteristics of different types of high-dimensional data by performing nonlinear dimensionality reduction. Most of these algorithms operate in a "batch" mode and cannot be efficiently applied when data are collected sequentially. In this paper, we describe an incremental version of ISOMAP, one of the key manifold learning algorithms. Our experiments on synthetic data as well as real world images demonstrate that our modified algorithm can maintain an accurate low-dimensional representation of the data in an efficient manner.
Martin H. C. Law, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Matching 2.5D Face Scans to 3D Models
abstract
The performance of face recognition systems that use two-dimensional images depends on factors such as lighting and subject's pose. We are developing a face recognition system that utilizes three-dimensional shape information to make the system more robust to arbitrary pose and lighting. For each subject, a 3D face model is constructed by integrating several 2.5D face scans which are captured from different views. 2.5D is a simplified 3D (x, y, z) surface representation that contains at most one depth value (z direction) for every point in the (x, y) plane. Two different modalities provided by the facial scan, namely, shape and texture, are utilized and integrated for face matching. The recognition engine consists of two components, surface matching and appearance-based matching. The surface matching component is based on a modified Iterative Closest Point (ICP) algorithm. The candidate list from the gallery used for appearance matching is dynamically generated based on the output of the surface matching component, which reduces the complexity of the appearance-based matching stage. Three-dimensional models in the gallery are used to synthesize new appearance samples with pose and illumination variations and the synthesized face images are used in discriminant subspace analysis. The weighted sum rule is applied to combine the scores given by the two matching components. Experimental results are given for matching a database of 200 3D face models with 598 2.5D independent test scans acquired under different pose and some lighting and expression changes. These results show the feasibility of the proposed matching scheme.
Xiaoguang Lu, Anil K. Jain 0001, Dirk Colbry
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Fingerprint Warping Using Ridge Curve Correspondences
abstract
The performance of a fingerprint matching system is affected by the nonlinear deformation introduced in the fingerprint impression during image acquisition. This nonlinear deformation causes fingerprint features such as minutiae points and ridge curves to be distorted in a complex manner. A technique is presented to estimate the nonlinear distortion in fingerprint pairs based on ridge curve correspondences. The nonlinear distortion, represented using the thin-plate spline (TPS) function, aids in the estimation of an "average" deformation model for a specific finger when several impressions of that finger are available. The estimated average deformation is then utilized to distort the template fingerprint prior to matching it with an input fingerprint. The proposed deformation model based on ridge curves leads to a better alignment of two fingerprint images compared to a deformation model based on minutiae patterns. An index of deformation is proposed for selecting the "optimal" deformation model arising from multiple impressions associated with a finger. Results based on experimental data consisting of 1,600 fingerprints corresponding to 50 different fingers collected over a period of two weeks show that incorporating the proposed deformation model results in an improvement in the matching performance.
Arun Ross, Sarat C. Dass, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Personal authentication using hand images
Ajay Kumar 0001, David C. M. Wong, Helen C. Shen, Anil K. Jain 0001
Pattern Recognit. Lett.4
2006 Biometrics: a tool for information security
abstract
Establishing identity is becoming critical in our vastly interconnected society. Questions such as "Is she really who she claims to be?," "Is this person authorized to use this facility?," or "Is he in the watchlist posted by the government?" are routinely being posed in a variety of scenarios ranging from issuing a driver's license to gaining entry into a country. The need for reliable user authentication techniques has increased in the wake of heightened concerns about security and rapid advancements in networking, communication, and mobility. Biometrics, described as the science of recognizing an individual based on his or her physical or behavioral traits, is beginning to gain acceptance as a legitimate method for determining an individual's identity. Biometric systems have now been deployed in various commercial, civilian, and forensic applications as a means of establishing identity. In this paper, we provide an overview of biometrics and discuss some of the salient research issues that need to be addressed for making biometric technology an effective tool for providing information security. The primary contribution of this overview includes: 1) examining applications where biometric scan solve issues pertaining to information security; 2) enumerating the fundamental challenges encountered by biometric systems in real-world applications; and 3) discussing solutions to address the problems of scalability and security in large-scale authentication systems.
Anil K. Jain 0001, Arun Ross, Sharath Pankanti
IEEE Trans. Inf. Forensics Secur.1
2005 Learning with Constrained and Unlabelled Data
abstract
Classification problems abundantly arise in many computer vision tasks eing of supervised, semi-supervised or unsupervised nature. Even when class labels are not available, a user still might favor certain grouping solutions over others. This bias can be expressed either by providing a clustering criterion or cost function and, in addition to that, by specifying pairwise constraints on the assignment of objects to classes. In this work, we discuss a unifying formulation for labelled and unlabelled data that can incorporate constrained data for model fitting. Our approach models the constraint information by the maximum entropy principle. This modeling strategy allows us (i) to handle constraint violations and soft constraints, and, at the same time, (ii) to speed up the optimization process. Experimental results on face classification and image segmentation indicates that the proposed algorithm is computationally efficient and generates superior groupings when compared with alternative techniques.
Tilman Lange, Martin H. C. Law, Anil K. Jain 0001, Joachim M. Buhmann
CVPR (1)3
2005 Model-based Clustering With Probabilistic Constraints
abstract
The problem of clustering with constraints is receiving increasing attention. Many existing algorithms assume the specified constraints are correct and consistent. We take a new approach and model the uncertainty of constraints in a principled manner by treating the constraints as random variables. The effect of specified constraints on a subset of points is propagated to other data points by biasing the search for cluster boundaries. By combining the a posteriori enforcement of constraints with the log-likelihood, we obtain a new objective function. An EM-type algorithm derived by variational method is used for efficient parameter estimation. Experimental results demonstrate the usefulness of the proposed algorithm. In particular, our approach can identify the desired clusters even when only a small portion of data participates in constraints.
Martin H. C. Law, Alexander P. Topchy, Anil K. Jain 0001
SDM3
2005 Dental Biometrics: Alignment and Matching of Dental Radiographs
abstract
Dental biometrics utilizes dental radiographs for human identification. The dental radiographs provide information about teeth, including tooth contours, elative positions of neighboring teeth, and shapes of the dental work (e.g., crowns, fillings, and bridges). The proposed system has two main stages: feature extraction and matching. The feature extraction stage uses anisotropic diffusion to enhance the images and a Mixture of Gaussians model to segment the dental work. The matching stage has three sequential steps: tooth-level matching, computation of image distances, and subject identification. In the tooth-level matching step, tooth contours are matched using a shape registration method, and the dental work is matched on overlapping areas. The distance between the tooth contours and the distance between the dental work are then combined using posterior probabilities. In the second step, the tooth correspondences between the given query (postmortem) radiograph and the database (antemortem) radiograph are established. A distance based on the corresponding teeth is then used to measure the similarity between the two radiographs. Finally, all the distances between the given postmortem radiographs and the antemortem radiographs that provide candidate identities are combined to establish the identity of the subject associated with the postmortem radiographs.
Hong Chen 0005, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Combining Multiple Clusterings Using Evidence Accumulation
abstract
We explore the idea of evidence accumulation (EAC) for combining the results of multiple clusterings. First, a clustering ensemble--a set of object partitions, is produced. Given a data set (n objects or patterns in d dimensions), different ways of producing data partitions are: 1) applying different clustering algorithms and 2) applying the same clustering algorithm with different values of parameters or initializations. Further, combinations of different data representations (feature spaces) and clustering algorithms can also provide a multitude of significantly different data partitionings. We propose a simple framework for extracting a consistent clustering, given the various partitions in a clustering ensemble. According to the EAC concept, each partition is viewed as an independent evidence of data organization, individual data partitions being combined, based on a voting mechanism, to generate a new n x n, similarity matrix between the n patterns. The final data partition of the n patterns is obtained by applying a hierarchical agglomerative clustering algorithm on this matrix. We have developed a theoretical framework for the analysis of the proposed clustering combination strategy and its evaluation, based on the concept of mutual information between data partitions. Stability of the results is evaluated using bootstrapping techniques. A detailed discussion of an evidence accumulation-based clustering algorithm, using a split and merge strategy based on the K-means clustering algorithm, is presented. Experimental results of the proposed method on several synthetic and real data sets are compared with other combination strategies, and with individual clustering results produced by well-known clustering algorithms.
Ana Fred, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Large-Scale Evaluation of Multimodal Biometric Authentication Using State-of-the-Art Systems
abstract
We examine the performance of multimodal biometric authentication systems using state-of-the-art Commercial Off-the-Shelf (COTS) fingerprint and face biometric systems on a population approaching 1,000 individuals. The majority of prior studies of multimodal biometrics have been limited to relatively low accuracy non-COTS systems and populations of a few hundred users. Our work is the first to demonstrate that multimodal fingerprint and face biometric systems can achieve significant accuracy gains over either biometric alone, even when using highly accurate COTS systems on a relatively large-scale population. In addition to examining well-known multimodal methods, we introduce new methods of normalization and fusion that further improve the accuracy.
Robert Snelick, Umut Uludag, Alan Mink, Mike Indovina, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2005 Clustering Ensembles: Models of Consensus and Weak Partitions
abstract
Clustering ensembles have emerged as a powerful method for improving both the robustness as well as the stability of unsupervised classification solutions. However, finding a consensus clustering from multiple partitions is a difficult problem that can be approached from graph-based, combinatorial, or statistical perspectives. This study extends previous research on clustering ensembles in several respects. First, we introduce a unified representation for multiple clusterings and formulate the corresponding categorical clustering problem. Second, we propose a probabilistic model of consensus using a finite mixture of multinomial distributions in a space of clusterings. A combined partition is found as a solution to the corresponding maximum-likelihood problem using the EM algorithm. Third, we define a new consensus function that is related to the classical intraclass variance criterion using the generalized mutual information definition. Finally, we demonstrate the efficacy of combining partitions generated by weak clustering algorithms that use data projections and random data splits. A simple explanatory model is offered for the behavior of combinations of such weak clustering components. Combination accuracy is analyzed as a function of several parameters that control the power and resolution of component partitions as well as the number of partitions. We also analyze clustering ensembles with incomplete information and the effect of missing cluster labels on the quality of overall consensus. Experimental results demonstrate the effectiveness of the proposed methods on several real-world data sets.
Alexander P. Topchy, Anil K. Jain 0001, William F. Punch
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Score normalization in multimodal biometric systems
Anil K. Jain 0001, Karthik Nandakumar, Arun Ross
Pattern Recognit.1
2005 A deformable model for fingerprint matching
Arun Ross, Sarat C. Dass, Anil K. Jain 0001
Pattern Recognit.3
2005 A wrapper-based approach to image segmentation and classification
abstract
The traditional processing flow of segmentation followed by classification in computer vision assumes that the segmentation is able to successfully extract the object of interest from the background image. It is extremely difficult to obtain a reliable segmentation without any prior knowledge about the object that is being extracted from the scene. This is further complicated by the lack of any clearly defined metrics for evaluating the quality of segmentation or for comparing segmentation algorithms. We propose a method of segmentation that addresses both of these issues, by using the object classification subsystem as an integral part of the segmentation. This will provide contextual information regarding the objects to be segmented, as well as allow us to use the probability of correct classification as a metric to determine the quality of the segmentation. We view traditional segmentation as a filter operating on the image that is independent of the classifier, much like the filter methods for feature selection. We propose a new paradigm for segmentation and classification that follows the wrapper methods of feature selection. Our method wraps the segmentation and classification together, and uses the classification accuracy as the metric to determine the best segmentation. By using shape as the classification feature, we are able to develop a segmentation algorithm that relaxes the requirement that the object of interest to be segmented must be homogeneous in some low-level image parameter, such as texture, color, or grayscale. This represents an improvement over other segmentation methods that have used classification information only to modify the segmenter parameters, since these algorithms still require an underlying homogeneity in some parameter space. Rather than considering our method as, yet, another segmentation algorithm, we propose that our wrapper method can be considered as an image segmentation framework, within which existing image segmentation algorithms may be executed. We show the performance of our proposed wrapper-based segmenter on real-world and complex images of automotive vehicle occupants for the purpose of recognizing infants on the passenger seat and disabling the vehicle airbag. This is an interesting application for testing the robustness of our approach, due to the complexity of the images, and, consequently, we believe the algorithm will be suitable for many other real-world applications.
Michael E. Farmer, Anil K. Jain 0001
IEEE Trans. Image Process.2
2004 Multiobjective Data Clustering
Martin H. C. Law, Alexander P. Topchy, Anil K. Jain 0001
CVPR (2)3
2004 Analysis of Consensus Partition in Cluster Ensemble
abstract
In combination of multiple partitions, one is usually interested in deriving a consensus solution with a quality better than that of given partitions. Several recent studies have empirically demonstrated improved accuracy of clustering ensembles on a number of artificial and real-world data sets. Unlike certain multiple supervised classifier systems, convergence properties of unsupervised clustering ensembles remain unknown for conventional combination schemes. In this paper, we present formal arguments on the effectiveness of cluster ensemble from two perspectives. The first is based on a stochastic partition generation model related to re-labeling and consensus function with plurality voting. The second is to study the property of the "mean" partition of an ensemble with respect to a metric on the space of all possible partitions. In both the cases, the consensus solution can be shown to converge to a true underlying clustering solution as the number of partitions in the ensemble increases. This paper provides a rigorous justification for the use of cluster ensemble.
Alexander P. Topchy, Martin H. C. Law, Anil K. Jain 0001, Ana Fred
ICDM3
2004 Robust motion-based image segmentation using fusion
abstract
To support real-time tracking of objects in video sequences, there has been considerable effort directed at developing optical flow and general motion-based image segmentation algorithms. The goal is to segment multiple moving objects in the image based on their relative motion. This task can be complicated by the presence of lighting variations. Furthermore, a combination of multiple motions and complex lighting effects can lead to dramatic image variations that may not be adequately accounted for by a single motion-based segmentation algorithm. We propose to fuse the results of multiple motion segmentation algorithms to improve the system robustness. Our approach uses the expectation maximization (EM) algorithm as a fusion engine. It also uses principal components analysis (PCA) to perform dimensionality reduction to improve the performance of the EM algorithm and reduce the processing burden. The performance of the proposed fusion algorithm has been demonstrated in the "smart airbag" application of monitoring occupants in a moving automobile to determine if they are too close to the instrument panel (airbag). Through fusion of the outputs of multiple algorithms we are able to reduce the percentage of pixels missed on the target by 35.
Michael E. Farmer, Xiaoguang Lu, Hong Chen 0005, Anil K. Jain 0001
ICIP4
2004 Nonlinear Manifold Learning for Data Stream
abstract
There has been a renewed interest in understanding the structure of high dimensional data set based on manifold learning. Examples include ISOMAP [25], LLE [20] and Laplacian Eigenmap [2] algorithms. Most of these algorithms operate in a “batch” mode and cannot be applied efficiently for a data stream. We propose an incremental version of ISOMAP. Our experiments not only demonstrate the accuracy and efficiency of the proposed algorithm, but also reveal interesting behavior of the ISOMAP as the size of available data increases.
Martin H. C. Law, Nan Zhang 0002, Anil K. Jain 0001
SDM3
2004 A Mixture Model for Clustering Ensembles
abstract
Clustering ensembles have emerged as a powerful method for improving both the robustness and the stability of unsupervised classification solutions. However, finding a consensus clustering from multiple partitions is a difficult problem that can be approached from graph-based, combinatorial or statistical perspectives. We offer a probabilistic model of consensus using a finite mixture of multinomial distributions in a space of clusterings. A combined partition is found as a solution to the corresponding maximum likelihood problem using the EM algorithm. The excellent scalability of this algorithm and comprehensible underlying model are particularly important for clustering of large datasets. This study compares the performance of the EM consensus algorithm with other fusion approaches for clustering ensembles. We also analyze clustering ensembles with incomplete information and the effect of missing cluster labels on the quality of overall consensus. Experimental results demonstrate the effectiveness of the proposed method on large real-world datasets.
Alexander P. Topchy, Anil K. Jain 0001, William F. Punch
SDM2
2004 Simultaneous Feature Selection and Clustering Using Mixture Models
abstract
Clustering is a common unsupervised learning technique used to discover group structure in a set of data. While there exist many algorithms for clustering, the important issue of feature selection, that is, what attributes of the data should be used by the clustering algorithms, is rarely touched upon. Feature selection for clustering is difficult because, unlike in supervised learning, there are no class labels for the data and, thus, no obvious criteria to guide the search. Another important problem in clustering is the determination of the number of clusters, which clearly impacts and is influenced by the feature selection issue. In this paper, we propose the concept of feature saliency and introduce an expectation-maximization (EM) algorithm to estimate it, in the context of mixture-based clustering. Due to the introduction of a minimum message length model selection criterion, the saliency of irrelevant features is driven toward zero, which corresponds to performing feature selection. The criterion and algorithm are then extended to simultaneously estimate the feature saliencies and the number of clusters.
Martin H. C. Law, Mário A. T. Figueiredo, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Online Handwritten Script Recognition
Anoop M. Namboodiri, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Biometric Cryptosystems: Issues and Challenges
abstract
In traditional cryptosystems, user authentication is based on possession of secret keys; the method falls apart if the keys are not kept secret (i.e., shared with non-legitimate users). Further, keys can be forgotten, lost, or stolen and, thus, cannot provide non-repudiation. Current authentication systems based on physiological and behavioral characteristics of persons (known as biometrics), such as fingerprints, inherently provide solutions to many of these problems and may replace the authentication component of traditional cryptosystems. We present various methods that monolithically bind a cryptographic key with the biometric template of a user stored in the database in such a way that the key cannot be revealed without a successful biometric authentication. We assess the performance of one of these biometric key binding/generation algorithms using the fingerprint biometric. We illustrate the challenges involved in biometric key generation primarily due to drastic acquisition variations in the representation of a biometric identifier and the imperfect nature of biometric feature extraction and matching algorithms. We elaborate on the suitability of these algorithms for digital rights management systems.
Umut Uludag, Sharath Pankanti, Salil Prabhakar, Anil K. Jain 0001
Proc. IEEE4
2004 Matching of dental X-ray images for human identification
Anil K. Jain 0001, Hong Chen 0005
Pattern Recognit.1
2004 Text information extraction in images and video: a survey
Keechul Jung, Kwang In Kim, Anil K. Jain 0001
Pattern Recognit.3
2004 Biometric template selection and update: a case study in fingerprints
Umut Uludag, Arun Ross, Anil K. Jain 0001
Pattern Recognit.3
2004 An introduction to biometric recognition
abstract
A wide variety of systems requires reliable personal recognition schemes to either confirm or determine the identity of an individual requesting their services. The purpose of such schemes is to ensure that the rendered services are accessed only by a legitimate user and no one else. Examples of such applications include secure access to buildings, computer systems, laptops, cellular phones, and ATMs. In the absence of robust personal recognition schemes, these systems are vulnerable to the wiles of an impostor. Biometric recognition, or, simply, biometrics, refers to the automatic recognition of individuals based on their physiological and/or behavioral characteristics. By using biometrics, it is possible to confirm or establish an individual's identity based on "who she is", rather than by "what she possesses" (e.g., an ID card) or "what she remembers" (e.g., a password). We give a brief overview of the field of biometrics and summarize some of its advantages, disadvantages, strengths, limitations, and related privacy concerns.
Anil K. Jain 0001, Arun Ross, Salil Prabhakar
IEEE Trans. Circuits Syst. Video Technol.1
2003 Occupant Classification System for Automotive Airbag Suppression
abstract
The introduction of airbags into automobiles has significantly improved the safety of the occupants. Unfortunately, airbags can also cause fatal injuries if the occupant is a child smaller (in weight) than a typical 6 year old. In response to this, The National Highway Transportation and Safety Administration (NHTSA) has mandated that starting in the 2006 model year all automobiles be equipped with an automatic suppression system to detect the presence of a child or infant and suppress the airbag. The classification problem we address is a four-class problem with the classes being rear-facing infant seat, child, adult, and empty seat. We describe a machine vision-based occupant classification system using a single grayscale camera and a digital signal processor that can perform this function in "real time" (< 5 seconds). The system has been extensively tested on a database of over 21,000 real-world images collected over a period of 4 months in moderate lighting conditions with a wide variety of passengers in eight different vehicles. We have achieved a classification accuracy of /spl sim/ 95%. We believe this system serves the need for a low-cost, high reliability embedded real-time airbag suppression system. Additional testing and improvements of the classification system are currently underway.
Michael E. Farmer, Anil K. Jain 0001
CVPR (1)2
2003 Robust Data Clustering
abstract
We address the problem of robust clustering by combining data partitions (forming a clustering ensemble) produced by multiple clusterings. We formulate robust clustering under an information-theoretical framework; mutual information is the underlying concept used in the definition of quantitative measures of agreement or consistency between data partitions. Robustness is assessed by variance of the cluster membership, based on bootstrapping. We propose and analyze a voting mechanism on pairwise associations of patterns for combining data partitions. We show that the proposed technique attempts to optimize the mutual information based criteria, although the optimality is not ensured in all situations. This evidence accumulation method is demonstrated by combining the well-known K-means algorithm to produce clustering ensembles. Experimental results show the ability of the technique to identify clusters with arbitrary shapes and sizes.
Ana Fred, Anil K. Jain 0001
CVPR (2)2
2003 Indexing and Retrieval of On-line Handwritten Documents
abstract
Recent advances in on-line data capturing technologiesand its widespread deployment in devices like PDAsand notebook PCs is creating large amounts of handwrittendata that need to be archived and retrieved efficiently.Word-spotting, which is based on a direct comparison ofa handwritten keyword to words in the document, is commonlyused for indexing and retrieval. We propose a stringmatching-based method for word-spotting in on-line documents.The retrieval algorithm achieves a precision of92.3% at a recall rate of 90% on a database of 6,672 wordswritten by 10 different writers. Indexing experiments showan accuracy of 87.5% using a database of 3,872 on-linewords.
Anil K. Jain 0001, Anoop M. Namboodiri
ICDAR1
2003 Combining Multiple Weak Clusterings
abstract
A data set can be clustered in many ways depending on the clustering algorithm employed, parameter settings used and other factors. Can multiple clusterings be combined so that the final partitioning of data provides better clustering? The answer depends on the quality of clusterings to be combined as well as the properties of the fusion method. First, we introduce a unified representation for multiple clusterings and formulate the corresponding categorical clustering problem. As a result, we show that the consensus function is related to the classical intra-class variance criterion using the generalized mutual information definition. Second, we show the efficacy of combining partitions generated by weak clustering algorithms that use data projections and random data splits. A simple explanatory model is offered for the behavior of combinations of such weak clustering components. We analyze the combination accuracy as a function of parameters controlling the power and resolution of component partitions as well as the learning dynamics vs. the number of clusterings involved. Finally, some empirical studies compare the effectiveness of several consensus functions.
Alexander P. Topchy, Anil K. Jain 0001, William F. Punch
ICDM2
2003 Integrated segmentation and classification for automotive airbag suppression
abstract
The use of airbags into automobiles has significantly improved the safety of the occupants. Unfortunately, when airbags are deployed in the case of a crash, they can also cause fatal injuries if the occupant is a child smaller (in weight) than a typical 6 year old. In response to this, The National Highway Transportation and Safety Administration (NHTSA) has mandated that, starting in the 2006 model year, all automobiles be equipped with an automatic suppression system. These systems are supposed to suppress the airbag if a child or an infant is occupying the front passenger seat. We are investigating the use of machine vision to classify front-seat passenger occupants into four classes: (i) adult, (ii) empty, (iii) RFIS (rear facing infant seat), and child. The design and integration of such a vision system into automobiles is very difficult due to (i) occupant variability (e.g., different types of infant seats and children clothing), and (ii) extreme lighting variability (e.g., bright sunny days, and night time operation). Our approach is based on recognition-driven segmentation. Preliminary results show that by integrating the segmentation and classification stages of processing, we are able to more reliably recognize the various occupant classes.
Michael E. Farmer, Anil K. Jain 0001
ICIP (3)2
2003 Combining classifiers for face recognition
abstract
Current two-dimensional face recognition approaches can obtain a good performance only under constrained environments. However, in the real applications, face appearance changes significantly due to different illumination, pose, and expression. Face recognizers based on different representations of the input face images have different sensitivity to these variations. Therefore, a combination of different face classifiers which can integrate the complementary information should lead to improved classification accuracy. We use the sum rule and RBF-based integration strategies to combine three commonly used face classifiers based on PCA, ICA and LDA representations. Experiments conducted on a face database containing 206 subjects (2,060 face images) show that the proposed classifier combination approaches outperform individual classifiers.
Xiaoguang Lu, Yunhong Wang 0001, Anil K. Jain 0001
ICME3
2003 Multimedia content protection via biometrics-based encryption
abstract
We propose a multimedia content protection framework that is based on biometric data of the users and a layered encryption/decryption scheme. Password-only encryption schemes are vulnerable to illegal key exchange problems. By using biometric data along with hardware identifiers as keys, it is possible to alleviate fraudulent usage of protected content. A combination of symmetric and asymmetric key systems is utilized for this purpose. The computational requirements and applicability of the proposed method are addressed. The results of encryption and decryption experiments related to time measurements are included. Watermarking systems can be used to complement the proposed method to permit novel uses of protected multimedia data.
Umut Uludag, Anil K. Jain 0001
ICME2
2003 Multimodal user interfaces: who's the user?
abstract
A wide variety of systems require reliable personal recognition schemes to either confirm or determine the identity of an individual requesting their services. The purpose of such schemes is to ensure that only a legitimate user, and not anyone else, accesses the rendered services. Examples of such applications include secure access to buildings, computer systems, laptops, cellular phones and ATMs. Biometric recognition, or simply biometrics, refers to the automatic recognition of individuals based on their physiological and/or behavioral characteristics. By using biometrics it is possible to confirm or establish an individual's identity based on "who she is", rather than by "what she possesses" (e.g., an ID card) or "what she remembers" (e.g., a password). Current biometric systems make use of fingerprints, hand geometry, iris, face, voice, etc. to establish a person's identity. Biometric systems also introduce an aspect of user convenience. For example, they alleviate the need for a user to remember multiple passwords associated with different applications. A biometric system that uses a single biometric trait for recognition has to contend with problems related to non-universality of the trait, spoof attacks, limited degrees of freedom, large intra-class variability, and noisy data. Some of these problems can be addressed by integrating the evidence presented by multiple biometric traits of a user (e.g., face and iris). Such systems, known as multimodal biometric systems, demonstrate substantial improvement in recognition performance. In this talk, we will present various applications of biometrics, challenges associated in designing biometric systems, and various fusion strategies available to implement a multimodal biometric system.
Anil K. Jain 0001
ICMI1
2003 Generating Discriminating Cartoon Faces Using Interacting Snakes
abstract
As a computational bridge between the high-level a priori knowledge of object shape and the low-level image data, active contours (or snakes) are useful models for the extraction of deformable objects. We propose an approach for manipulating multiple snakes iteratively, called interacting snakes, that minimizes the attraction energy functionals on both contours and enclosed regions of individual snakes and the repulsion energy functionals among multiple snakes that interact with each other. We implement the interacting snakes through explicit curve (parametric active contours) representation in the domain of face recognition. We represent human faces semantically via facial components such as eyes, mouth, face outline, and the hair outline. Each facial component is encoded by a closed (or open) snake that is drawn from a 3D generic face model. A collection of semantic facial components form a hypergraph, called semantic face graph, which employs interacting snakes to align the general facial topology onto the sensed face images. Experimental results show that a successful interaction among multiple snakes associated with facial components makes the semantic face graph a useful model for face representation, including cartoon faces and caricatures, and recognition.
Rein-Lien Hsu, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Hiding Biometric Data
abstract
With the wide spread utilization of biometric identification systems, establishing the authenticity of biometric data itself has emerged as an important research issue. The fact that biometric data is not replaceable and is not secret, combined with the existence of several types of attacks that are possible in a biometric system, make the issue of security/integrity of biometric data extremely critical. We introduce two applications of an amplitude modulation-based watermarking method, in which we hide a user's biometric data in a variety of images. This method has the ability to increase the security of both the hidden biometric data (e.g., eigen-face coefficients) and host images (e.g., fingerprints). Image adaptive data embedding methods used in our scheme lead to low visibility of the embedded signal. Feature analysis of host images guarantees high verification accuracy on watermarked (e.g., fingerprint) images.
Anil K. Jain 0001, Umut Uludag
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Learning fingerprint minutiae location and type
Salil Prabhakar, Anil K. Jain 0001, Sharath Pankanti
Pattern Recognit.2
2003 A hybrid fingerprint matcher
Arun Ross, Anil K. Jain 0001, James Reisman
Pattern Recognit.2
2003 Information fusion in biometrics
Arun Ross, Anil K. Jain 0001
Pattern Recognit. Lett.2
2002 Fingerprint mosaicking
abstract
It has been observed that the reduced contact area offered by solid-state fingerprint sensors does not provide sufficient information (e.g., number of minutiae) for high accuracy user verification. Further, multiple impressions of the same finger acquired by these sensors, may have only a small region of overlap thereby degrading the matching performance of the verification system. To deal with this problem, we have developed a fingerprint mosaicking scheme that constructs a composite fingerprint template using multiple impressions. A composite template reduces storage, improves matching time and alleviates the problem of template selection. In the proposed algorithm, two impressions (templates) of a finger are initially aligned using the corresponding minutiae points. This alignment is used by a modified version of the well-known iterative closest point algorithm (ICP) to compute a transformation matrix that defines the spatial relationship between the two impressions. The resulting transformation matrix is used in two ways: (a) the two templates are stitched together to generate a composite image. Minutiae points are then detected in this composite image; (b) the minutia maps obtained from each of the individual impressions are integrated to create a larger minutia map. Our experiments show that a composite template improves the performance of the fingerprint matching system by ∼ 4%.
Anil K. Jain 0001, Arun Ross
ICASSP1
2002 Learning user-specific parameters in a multibiometric system
abstract
Biometric systems that use a single biometric trait have to contend with noisy data, restricted degrees of freedom, failure-to-enroll problems, spoof attacks, and unacceptable error rates. Multibiometric systems that use multiple traits of an individual for authentication, alleviate some of these problems while improving verification performance. We demonstrate that the performance of multibiometric systems can be further improved by learning user-specific parameters. Two types of parameters are considered here. (i) Thresholds that are used to decide if a matching score indicates a genuine user or an impostor, and (ii) weights that are used to indicate the importance of matching scores output by each biometric trait. User-specific thresholds are computed using the cumulative histogram of impostor matching scores corresponding to each user. The user-specific weights associated with each biometric are estimated by searching for that set of weights which minimizes the total verification error. The tests were conducted on a database of 50 users who provided fingerprint, face and hand geometry data, with 10 of these users providing data over a period of two months. We observed that user-specific thresholds improved system performance by /spl sim/ 2%, while user-specific weights improved performance by /spl sim/ 3%.
Anil K. Jain 0001, Arun Ross
ICIP (1)1
2002 Semantic face matching
abstract
The need for efficient methods for archiving and retrieving personal digital photo collections arises due to a significant increase in the number of digital images and videos that people have to manage. We propose a semantic face matching approach for managing consumer photographs based on semantic face attributes. These attributes are organized as a semantic face graph (derived from a 3D generic face model) containing facial components such as eyes and mouth in the spatial domain. We align the semantic facial components in the semantic face graph with the extracted facial features in a given image. Aligned facial components are transformed to a feature space spanned by Fourier descriptors of facial components for face matching. The semantic face graph allows face matching based on selected facial components. Our experimental results demonstrate that the proposed semantic representation of the face is useful for face matching and visualization (e.g., generating facial caricatures).
Rein-Lien Hsu, Anil K. Jain 0001
ICME (2)2
2002 Feature Selection in Mixture-Based Clustering
abstract
There exist many approaches to clustering, but the important issue of feature selection, i.e., selecting the data attributes that are relevant for clustering, is rarely addressed. Feature selection for clustering is difficult due to the absence of class labels. We propose two approaches to feature selection in the context of Gaussian mixture-based clustering. In the first one, instead of making hard selections, we estimate feature saliencies. An expectation-maximization (EM) algorithm is derived for this task. The second approach extends Koller and Sahami’s mutual-information- based feature relevance criterion to the unsupervised case. Feature selec- tion is then carried out by a backward search scheme. This scheme can be classified as a “wrapper”, since it wraps mixture estimation in an outer layer that performs feature selection. Experimental results on synthetic and real data show that both methods have promising performance.
Martin H. C. Law, Anil K. Jain 0001, Mário A. T. Figueiredo
NIPS2
2002 Writer Adaptation for Online Handwriting Recognition
abstract
Writer-adaptation is the process of converting a writer-independent handwriting recognition system into a writer-dependent system. It can greatly increasing recognition accuracy, given adequate writer models. The limited amount of data a writer provides during training constrains the models' complexity. We show how appropriate use of writer-independent models is important for the adaptation. Our approach uses writer-independent writing style models (lexemes) to identify the styles present in a particular writer's training data. These models are then updated using the writer's data. Lexemes in the writer's data for which an inadequate number of training examples is available are replaced with the writer-independent models. We demonstrate the feasibility of this approach on both isolated handwritten character recognition and unconstrained word recognition tasks. Our results show an average reduction in error rate of 16.3 percent for lowercase characters as compared against representing each of the writer's character classes with a single model. In addition, an average error rate reduction of 9.2 percent is shown on handwritten words using only a small amount of data for adaptation.
Scott D. Connell, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Unsupervised Learning of Finite Mixture Models
abstract
This paper proposes an unsupervised algorithm for learning a finite mixture model from multivariate data. The adjective "unsupervised" is justified by two properties of the algorithm: 1) it is capable of selecting the number of components and 2) unlike the standard expectation-maximization (EM) algorithm, it does not require careful initialization. The proposed method also avoids another drawback of EM for mixture fitting: the possibility of convergence toward a singular estimate at the boundary of the parameter space. The novelty of our approach is that we do not use a model selection criterion to choose one among a set of preestimated candidate models; instead, we seamlessly integrate estimation and model selection in a single algorithm. Our technique can be applied to any type of parametric mixture model for which it is possible to write an EM algorithm; in this paper, we illustrate it with experiments involving Gaussian mixtures. These experiments testify for the good performance of our approach.
Mário A. T. Figueiredo, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Face Detection in Color Images
abstract
Human face detection plays an important role in applications such as video surveillance, human computer interface, face recognition, and face image database management. We propose a face detection algorithm for color images in the presence of varying lighting conditions as well as complex backgrounds. Based on a novel lighting compensation technique and a nonlinear color transformation, our method detects skin regions over the entire image and then generates face candidates based on the spatial arrangement of these skin patches. The algorithm constructs eye, mouth, and boundary maps for verifying each face candidate. Experimental results demonstrate successful face detection over a wide range of facial variations in color, position, scale, orientation, 3D pose, and expression in images from several photo collections (both indoors and outdoors).
Rein-Lien Hsu, Mohamed Abdel-Mottaleb, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 FVC2000: Fingerprint Verification Competition
abstract
Reliable and accurate fingerprint recognition is a challenging pattern recognition problem, requiring algorithms robust in many contexts. FVC2000 competition attempted to establish the first common benchmark, allowing companies and academic institutions to unambiguously compare performance and track improvements in their fingerprint recognition algorithms. Three databases were created using different state-of-the-art sensors and a fourth database was artificially generated; 11 algorithms were extensively tested on the four data sets. We believe that FVC2000 protocol, databases, and results will be useful to all practitioners in the field not only as a benchmark for improving methods, but also for enabling an unbiased evaluation of algorithms.
Dario Maio, Davide Maltoni, Raffaele Cappelli, James L. Wayman, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2002 On the Individuality of Fingerprints
abstract
Fingerprint identification is based on two basic premises: (1) persistence and (2) individuality. We address the problem of fingerprint individuality by quantifying the amount of information available in minutiae features to establish a correspondence between two fingerprint images. We derive an expression which estimates the probability of a false correspondence between minutiae-based representations from two arbitrary fingerprints belonging to different fingers. Our results show that (1) contrary to the popular belief, fingerprint matching is not infallible and leads to some false associations, (2) while there is an overwhelming amount of discriminatory information present in the fingerprints, the strength of the evidence degrades drastically with noise in the sensed fingerprint images, (3) the performance of the state-of-the-art automatic fingerprint matchers is not even close to the theoretical limit, and (4) because automatic fingerprint verification systems based on minutia use only a part of the discriminatory information present in the fingerprints, it may be desirable to explore additional complementary representations of fingerprints for automatic matching.
Sharath Pankanti, Salil Prabhakar, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 Pattern recognition in information systems
Ana Fred, Anil K. Jain 0001
Pattern Recognit.2
2002 On-line signature verification,
Anil K. Jain 0001, Friederike D. Griess, Scott D. Connell
Pattern Recognit.1
2002 On the similarity of identical twin fingerprints
Anil K. Jain 0001, Salil Prabhakar, Sharath Pankanti
Pattern Recognit.1
2002 Decision-level fusion in fingerprint verification
Salil Prabhakar, Anil K. Jain 0001
Pattern Recognit.2
2002 Matching of palmprints
Nicolae Duta, Anil K. Jain 0001, Kanti V. Mardia
Pattern Recognit. Lett.2
2002 Automatic image orientation detection
abstract
We present an algorithm for automatic image orientation estimation using a Bayesian learning framework. We demonstrate that a small codebook (the optimal size of codebook is selected using a modified MDL criterion) extracted from a learning vector quantizer (LVQ) can be used to estimate the class-conditional densities of the observed features needed for the Bayesian methodology. We further show how principal component analysis (PCA) and linear discriminant analysis (LDA) can be used as a feature extraction mechanism to remove redundancies in the high-dimensional feature vectors used for classification. The proposed method is compared with four different commonly used classifiers, namely k-nearest neighbor, support vector machine (SVM), a mixture of Gaussians, and hierarchical discriminating regression (HDR) tree. Experiments on a database of 16 344 images have shown that our proposed algorithm achieves an accuracy of approximately 98% on the training set and over 97% on an independent test set. A slight improvement in classification accuracy is achieved by employing classifier combination techniques.
Aditya Vailaya, HongJiang Zhang, Changjiang Yang, Feng-I Liu, Anil K. Jain 0001
IEEE Trans. Image Process.5
2002 Learning similarity measure for natural image retrieval with relevance feedback
abstract
A new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two clusters. Images inside the boundary are ranked by their Euclidean distances to the query. The scheme is called constrained similarity measure (CSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance of the Euclidean distance measure. Two techniques, support vector machine (SVM) and AdaBoost from machine learning, are utilized to learn the boundary. They are compared to see their differences in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The CSM metric is evaluated in a large database of 10009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval.
Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang
IEEE Trans. Neural Networks2
2001 Bayesian Learning of Sparse Classifiers
abstract
Bayesian approaches to supervised learning use priors on the classifier parameters. However, few priors aim at achieving "sparse" classifiers, where irrelevant/redundant parameters are automatically set to zero. Two well-known ways of obtaining sparse classifiers are: use a zero-mean Laplacian prior on the parameters, and the "support vector machine" (SVM). Whether one uses a Laplacian prior or an SVM, one still needs to specify/estimate the parameters that control the degree of sparseness of the resulting classifiers. We propose a Bayesian approach to learning sparse classifiers which does not involve any parameters controlling the degree of sparseness. This is achieved by a hierarchical-Bayes interpretation of the Laplacian prior, followed by the adoption of a Jeffreys' non-informative hyper-prior Implementation is carried out by an EM algorithm. Experimental evaluation of the proposed method shows that it performs competitively with (often better than) the best classification techniques available.
Mário A. T. Figueiredo, Anil K. Jain 0001
CVPR (1)2
2001 Learning Similarity Measure for Natural Image Retrieval with Relevance Feedback
abstract
A new scheme of learning similarity measure is proposed for content-based image retrieval (CBIR). It learns a boundary that separates the images in the database into two parts. Images on the positive side of the boundary are ranked by their Euclidean distances to the query. The scheme is called restricted similarity measure (RSM), which not only takes into consideration the perceptual similarity between images, but also significantly improves the retrieval performance based on the Euclidean distance measure. Two techniques, support vector machine and AdaBoost, are utilized to learn the boundary, and compared with respect to their performance in boundary learning. The positive and negative examples used to learn the boundary are provided by the user with relevance feedback. The RSM metric is evaluated on a large database of 10,009 natural images with an accurate ground truth. Experimental results demonstrate the usefulness and effectiveness of the proposed similarity measure for image retrieval.
Guodong Guo, Anil K. Jain 0001, Wei-Ying Ma, HongJiang Zhang
CVPR (1)2
2001 On the Individuality of Fingerprints
abstract
Fingerprint identification is based on two basic premises: (i) persistence: the basic characteristics of fingerprints do not change with time; and (ii) individuality: the fingerprint is unique to an individual. The validity of the first premise has been established. While the second premise is generally accepted to be true, the underlying scientific basis of fingerprint individuality has not been formally tested. We address the problem of fingerprint individuality by quantifying the amount of information available in minutiae points to establish a correspondence between two fingerprint images. We derive an expression which estimates the probability of falsely associating minutiae-based representations from two arbitrary fingerprints. For example, the probability that a fingerprint with 36 minutiae points will share 12 minutiae points with another arbitrarily chosen fingerprint with 36 minutiae points is 6.10/spl times/10/sup -8/. These probability estimates are compared with typical fingerprint matcher accuracy results. Our results show that: (i) fingerprint matching is not infallible and leads to some false associations, (ii) the performance of automatic fingerprint matcher does not even come close to the theoretical performance, and (iii) due to the limited information content of the minutiae-based representation, automatic system designers should explore the use of non-minutiae-based information present in the fingerprints.
Sharath Pankanti, Salil Prabhakar, Anil K. Jain 0001
CVPR (1)3
2001 Markov Face Models
abstract
The spatial distribution of gray level intensities in an image can be naturally modeled using Markov random field (MRF) models. We develop and investigate the performance of face detection algorithms derived from MRF considerations. For enhanced detection, the MRF models are defined for every permutation of site indices (pixels) in the image. We find the optimal permutation that provides maximum discriminatory power to identify faces from nonfaces. The methodology presented here is a generalization of the face detection algorithm described previously where a most discriminating Markov chain model was used. The MRF models successfully detect faces in a number of test images.
Sarat C. Dass, Anil K. Jain 0001
ICCV2
2001 A Background Model Initialization Algorithm for Video Surveillance
Daniel Gutchess, Miroslav Trajkovic, Eric Cohen-Solal, Damian M. Lyons, Anil K. Jain 0001
ICCV5
2001 Structure in On-line Documents
abstract
We present a hierarchical approach for extracting homogeneous regions in on-line documents. The problem of identifying and processing ruled and unruled tables, text and drawings is addressed. The on-line document is first segmented into regions with only text strokes and regions with both text and non-text strokes. The text region is further classified as unruled table or plain text. Stroke clustering is used to segment the non-text regions. Each nontext segment is then classified as drawing, ruled table or underlined keyword using stroke properties. The individual regions are processed and the results are assembled to identify the structure of the on-line document.
Anil K. Jain 0001, Anoop M. Namboodiri, Jayashree Subrahmonia
ICDAR1
2001 Face detection in color images
abstract
Human face detection is often the first step in applications such as video surveillance, human computer interface, face recognition, and image database management. We propose a face detection algorithm for color images in the presence of varying lighting conditions as well as complex backgrounds. Our method detects skin regions over the entire image, and then generates face candidates based on the spatial arrangement of these skin patches. The algorithm constructs eye, mouth, and boundary maps for verifying each face candidate. Experimental results demonstrate successful detection over a wide variety of facial variations in color, position, scale, rotation, pose, and expression from several photo collections.
Rein-Lien Hsu, Mohamed Abdel-Mottaleb, Anil K. Jain 0001
ICIP (1)3
2001 Face modeling for recognition
abstract
3D human face models have been widely used in applications such as facial animation, video compression/coding, augmented reality, head tracking, facial expression recognition, human action recognition, and face recognition. Modeling human faces provides a potential solution to identifying faces with variations in illumination, pose, and facial expression. We propose a method of modeling human faces based on a generic face model (a triangular mesh model) and individual facial measurements containing both shape and texture information. The modeling method adapts a generic face model to the given facial features, extracted from registered range and color images, in a global-to-local fashion. It iteratively moves the vertices of the mesh model to smoothen the non-feature areas, and uses the 2.5D active contours to refine feature boundaries. The resultant face model has been shown to be visually similar to the true face. Initial results show that the constructed model is quite useful for recognizing nonfrontal views.
Rein-Lien Hsu, Anil K. Jain 0001
ICIP (2)2
2001 Fingerprint matching using minutiae and texture features
abstract
The advent of solid-state fingerprint sensors presents a fresh challenge to traditional fingerprint matching algorithms. These sensors provide a small contact area (/spl ap/0.6"/spl times/0.6") for the fingertip and, therefore, sense only a limited portion of the fingerprint. Thus multiple impressions of the same fingerprint may have only a small region of overlap. Minutiae-based matching algorithms, which consider ridge activity only in the vicinity of minutiae points, are not likely to perform well on these images due to the insufficient number of corresponding points in the input and template images. We present a hybrid matching algorithm that uses both minutiae (point) information and texture (region) information for matching the fingerprints. Results obtained on the MSU-VERIDICOM database shows that a combination of the texture-based and minutiae-based matching scores leads to a substantial improvement in the overall matching performance.
Anil K. Jain 0001, Arun Ross, Salil Prabhakar
ICIP (3)1
2001 Automatic Construction of 2D Shape Models
abstract
A procedure for automated 2D shape model design is presented. The system is given a set of training example shapes defined by contour point coordinates. The shapes are automatically aligned using Procrustes analysis and clustered to obtain cluster prototypes (typical objects) and statistical information about intracluster shape variation. One difference from previous methods is that the training set is first automatically clustered and shapes considered to be outliers are discarded. In this way, cluster prototypes are not distorted by outliers. A second difference is in the manner in which registered sets of points are extracted from each shape contour. We propose a flexible point matching technique that takes into account both pose/scale differences and nonlinear shape differences. The matching method is independent of the objects' initial relative position/scale and does not require any manually tuned parameters. Our shape model design method was used to learn 11 different shapes from contours that were manually traced in MR brain images. The resulting model was then employed to segment several MR brain images that were not included in the shape-training set. A quantitative analysis of our shape registration approach, within the main cluster of each structure, demonstrated results that compare very well to those achieved by manual registration; achieving an average registration error of about 1 pixel. Our approach can serve as a fully automated substitute to the tedious and time-consuming manual 2D shape registration and analysis.
Nicolae Duta, Anil K. Jain 0001, Marie-Pierre Jolly
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Aircraft Detection: A Case Study in Using Human Similarity Measure
abstract
After the most prominent signal in an infrared image of the sky is extracted, the question is whether the signal corresponds to an aircraft. We present a new approach that avoids metric similarity measures and the use of thresholds, and instead attempts to learn similarity measures like those used by humans. In the absence of sufficient real data, the approach allows one to specifically generate an arbitrarily large number of training exemplars projecting near the classification boundary. Once trained on such a training set, the performance of our neural network-based system is comparable to that of a human expert and far better than a network trained only on the available real data. Furthermore, the results obtained are considerably better than those obtained using an Euclidean discriminator.
Behrooz Kamgar-Parsi, Behzad Kamgar-Parsi, Anil K. Jain 0001, Judith E. Dayhoff
IEEE Trans. Pattern Anal. Mach. Intell.3
2001 Template-based online character recognition
Scott D. Connell, Anil K. Jain 0001
Pattern Recognit.2
2001 Image classification for content-based indexing
abstract
Grouping images into (semantically) meaningful categories using low-level visual features is a challenging and important problem in content-based image retrieval. Using binary Bayesian classifiers, we attempt to capture high-level concepts from low-level image features under the constraint that the test image does belong to one of the classes. Specifically, we consider the hierarchical classification of vacation images; at the highest level, images are classified as indoor or outdoor; outdoor images are further classified as city or landscape; finally, a subset of landscape images is classified into sunset, forest, and mountain classes. We demonstrate that a small vector quantizer (whose optimal size is selected using a modified MDL criterion) can be used to model the class-conditional densities of the features, required by the Bayesian methodology. The classifiers have been designed and evaluated on a database of 6931 vacation photographs. Our system achieved a classification accuracy of 90.5% for indoor/outdoor, 95.3% for city/landscape, 96.6% for sunset/forest and mountain, and 96% for forest/mountain classification problems. We further develop a learning method to incrementally train the classifiers as additional data become available. We also show preliminary results for feature reduction using clustering techniques. Our goal is to combine multiple two-class classifiers into a single hierarchical classifier.
Aditya Vailaya, Mário A. T. Figueiredo, Anil K. Jain 0001, HongJiang Zhang
IEEE Trans. Image Process.3
2000 Recognition of Unconstrained On-Line Devanagari Characters
abstract
Devanagari is a script used for several major languages such as Hindi, Sanskrit, Marathi and Nepali, and is used by more than 500 million people. Unconstrained Devanagari writing is more complex than English cursive due to the possible variations in the order, number, directional and shape of the constituent strokes. An online pen computing environment has numerous application in providing an easy human interface for a complex script like Devanagari. A Devanagari character recognition experiment with 20 different writers with each writer writing 5 samples of each character in a totally unconstrained way, has been conducted. An accuracy of 86.5% with no rejects is achieved through the combination of multiple classifiers that focus on either local online properties, or global off-line properties. Further improvements in performance are expected by using word-level contextual information. We also explore the use of writer dependent models to improve the recognition accuracy.
Scott D. Connell, R. Mahesh K. Sinha, Anil K. Jain 0001
ICPR3
2000 Unsupervised Selection and Estimation of Finite Mixture Models
abstract
We describe a method for fitting mixture models to multivariate data which performs component selection and does not require external initialization. The novelty of our approach includes: an MML-like (minimum message length) model selection criterion; inclusion of the criterion into the expectation-maximization (EM) algorithm (increasing its ability to escape from local maxima); an initialization strategy supported on the interpretation of EM as a self-annealing algorithm.
Mário A. T. Figueiredo, Anil K. Jain 0001
ICPR2
2000 Minutia Verification and Classification for Fingerprint Matching
abstract
We propose a feedback path for the feature extraction stage, followed by a feature refinement stage for improving the matching performance. This performance improvement is illustrated in the context of a minutia-based fingerprint verification system. We show that a minutia verification stage based on re-examining the gray-scale profile in a detected minutia's spatial neighborhood in the sensed image can improve the matching performance by /spl sim/4% on our database. Further, we show that a feature refinement stage which assigns a class label to each detected minutia (ridge ending and ridge bifurcation) before matching can also improve the matching performance by /spl sim/3%. A combination of feedback (minutia verification) in the feature extraction phase and feature refinement (minutia classification) improves the overall performance of the fingerprint verification system by /spl sim/8%.
Salil Prabhakar, Anil K. Jain 0001, Sharath Pankanti, Ruud M. Bolle
ICPR2
2000 Reject Option for VQ-Based Bayesian Classification
abstract
We have developed a reject option for VQ-based supervised Bayesian classification to improve classification accuracy by sieving out patterns that are classified with a low confidence value. A small codebook extracted from a learning vector quantizer (LVQ) is used to estimate the class-conditional densities of the feature vector. We adapt the two commonly used rejection criteria, outlier rejection and ambiguity rejection, for the VQ-based Bayesian classifiers. Using three high-level image classification problems, we demonstrate how local rejection criteria can improve the error vs. reject characteristics of our classifier over a global rejection method.
Aditya Vailaya, Anil K. Jain 0001
ICPR2
2000 Statistical Pattern Recognition: A Review
abstract
The primary goal of pattern recognition is supervised or unsupervised classification. Among the various frameworks in which pattern recognition has been traditionally formulated, the statistical approach has been most intensively studied and used in practice. More recently, neural network techniques and methods imported from statistical learning theory have been receiving increasing attention. The design of a recognition system requires careful attention to the following issues: definition of pattern classes, sensing environment, pattern representation, feature extraction and selection, cluster analysis, classifier design and learning, selection of training and test samples, and performance evaluation. In spite of almost 50 years of research and development in this field, the general problem of recognizing complex patterns with arbitrary orientation, location, and scale remains unsolved. New and emerging applications, such as data mining, web searching, retrieval of multimedia data, face recognition, and cursive handwriting recognition, require robust and efficient pattern recognition techniques. The objective of this review paper is to summarize and compare some of the well-known methods used in various stages of a pattern recognition system and identify research topics and applications which are at the forefront of this exciting and challenging field.
Anil K. Jain 0001, Robert P. W. Duin, Jianchang Mao
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 Object Tracking Using Deformable Templates
abstract
We propose a method for object tracking using prototype-based deformable template models. To track an object in an image sequence, we use a criterion which combines two terms: the frame-to-frame deviations of the object shape and the fidelity of the modeled shape to the input image. The deformable template model utilizes the prior shape information which is extracted from the previous frames along with a systematic shape deformation scheme to model the object shape in a new frame. The following image information is used in the tracking process: 1) edge and gradient information: the object boundary consists of pixels with large image gradient, 2) region consistency: the same object region possesses consistent color and texture throughout the sequence, and 3) interframe motion: the boundary of a moving object is characterized by large interframe motion. The tracking proceeds by optimizing an objective function which combines both the shape deformation and the fidelity of the modeled shape to the current image (in terms of gradient, texture, and interframe motion). The inherent structure in the deformable template, together with region, motion, and image gradient cues, makes the proposed algorithm relatively insensitive to the adverse effects of weak image features and moderate amounts of occlusion.
Yu Zhong 0001, Anil K. Jain 0001, Marie-Pierre Jolly
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Automatic Caption Localization in Compressed Video
abstract
We present a method to automatically localize captions in JPEG compressed images and the I-frames of MPEG compressed videos. Caption text regions are segmented from background images using their distinguishing texture characteristics. Unlike previously published methods which fully decompress the video sequence before extracting the text regions, this method locates candidate caption text regions directly in the DCT compressed domain using the intensity variation information encoded in the DCT domain. Therefore, only a very small amount of decoding is required. The proposed algorithm takes about 0.006 second to process a 240/spl times/350 image and achieves a recall rate of 99.17 percent while falsely accepting about 1.87 percent nontext DCT blocks on a variety of MPEG compressed videos containing more than 2,300 I-frames.
Yu Zhong 0001, HongJiang Zhang, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2000 Image-based form document retrieval
Anil K. Jain 0001
Pattern Recognit.2
2000 Object localization using color, texture and shape
Yu Zhong 0001, Anil K. Jain 0001
Pattern Recognit.2
2000 Dimensionality reduction using genetic algorithms
abstract
Pattern recognition generally requires that objects be described in terms of a set of measurable features. The selection and quality of the features representing each pattern affect the success of subsequent classification. Feature extraction is the process of deriving new features from original features to reduce the cost of feature measurement, increase classifier efficiency, and allow higher accuracy. Many feature extraction techniques involve linear transformations of the original pattern vectors to new vectors of lower dimensionality. While this is useful for data visualization and classification efficiency, it does not necessarily reduce the number of features to be measured since each new feature may be a linear combination of all of the features in the original pattern vector. Here, we present a new approach to feature extraction in which feature selection and extraction and classifier training are performed simultaneously using a genetic algorithm. The genetic algorithm optimizes a feature weight vector used to scale the individual features in the original pattern vectors. A masking vector is also employed for simultaneous selection of a feature subset. We employ this technique in combination with the k nearest neighbor classification rule, and compare the results with classical feature selection and extraction techniques, including sequential floating forward feature selection, and linear discriminant analysis. We also present results for the identification of favorable water-binding sites on protein surfaces.
Michael L. Raymer, William F. Punch, Erik D. Goodman, Leslie A. Kuhn, Anil K. Jain 0001
IEEE Trans. Evol. Comput.5
2000 Unsupervised contour representation and estimation using B-splines and a minimum description length criterion
abstract
This paper describes a new approach to adaptive estimation of parametric deformable contours based on B-spline representations. The problem is formulated in a statistical framework with the likelihood function being derived from a region-based image model. The parameters of the image model, the contour parameters, and the B-spline parameterization order (i.e., the number of control points) are all considered unknown. The parameterization order is estimated via a minimum description length (MDL) type criterion. A deterministic iterative algorithm is developed to implement the derived contour estimation criterion, the result is an unsupervised parametric deformable contour: it adapts its degree of smoothness/complexity (number of control points) and it also estimates the observation (image) model parameters. The experiments reported in the paper, performed on synthetic and real (medical) images, confirm the adequate and good performance of the approach.
Mário A. T. Figueiredo, José M. N. Leitão, Anil K. Jain 0001
IEEE Trans. Image Process.3
2000 Filterbank-based fingerprint matching
abstract
With identity fraud in our society reaching unprecedented proportions and with an increasing emphasis on the emerging automatic personal identification applications, biometrics-based verification, especially fingerprint-based identification, is receiving a lot of attention. There are two major shortcomings of the traditional approaches to fingerprint representation. For a considerable fraction of population, the representations based on explicit detection of complete ridge structures in the fingerprint are difficult to extract automatically. The widely used minutiae-based representation does not utilize a significant component of the rich discriminatory information available in the fingerprints. Local ridge structures cannot be completely characterized by minutiae. Further, minutiae-based matching has difficulty in quickly matching two fingerprint images containing a different number of unregistered minutiae points. The proposed filter-based algorithm uses a bank of Gabor filters to capture both local and global details in a fingerprint as a compact fixed length FingerCode. The fingerprint matching is based on the Euclidean distance between the two corresponding FingerCodes and hence is extremely fast. We are able to achieve a verification accuracy which is only marginally inferior to the best results of minutiae-based algorithms published in the open literature. Our system performs better than a state-of-the-art minutiae-based system when the performance requirement of the application system does not demand a very low false acceptance rate. Finally, we show that the matching performance can be improved by combining the decisions of the matchers based on complementary (minutiae-based and filter-based) fingerprint information.
Anil K. Jain 0001, Salil Prabhakar, Lin Hong, Sharath Pankanti
IEEE Trans. Image Process.1
1999 Learning 2D Shape Models
abstract
A new fully automated shape learning method is presented. It is based on clustering a set of training shapes in the original shape space (defined by the coordinates of the contour points) and performing a Procrustes analysis on each cluster to obtain cluster prototypes and information about shape variation. The main difference from previously reported methods is that the training set is first automatically clustered and those shapes considered to be outliers are discarded. The second difference is in the manner in which registered sets of points are extracted from each shape contour. As a direct application of our shape learning method, an 11-structure shape model of brain substructures was extracted from MR image data, an eigen-shape model was automatically trained, and employed to segment several MR brain images not present in the shape-training set. A quantitative analysis of our shape registration approach, within the main cluster of each structure, shows that our results compare very well to those achieved by manual registration; achieving an average rms error of about 1 pixel. Our approach can serve as a fully automated substitute to the tedious and time-consuming manual shape registration and analysis.
Nicolae Duta, Anil K. Jain 0001, Marie-Pierre Jolly
CVPR2
1999 FingerCode: A Filterbank for Fingerprint Representation and Matching
abstract
With the identity fraud in our society reaching unprecedented proportions and with an increasing emphasis on the emerging automatic positive personal identification applications, biometrics-based identification, especially fingerprint-based identification, is receiving a lot of attention. There are two major shortcomings of the traditional approaches to fingerprint representation. For a significant fraction of population, the representations based on explicit detection of complete ridge structures in the fingerprint are difficult to extract automatically. The widely used minutiae-based representation does not utilize a significant component of the rich discriminatory information, available in the fingerprints. The proposed filter-based algorithm uses a bank of Gabor filters to capture both the local and the global details in a fingerprint as a compact 640-byte fixed length FingerCode. The fingerprint matching is based on the Euclidean distance between the two corresponding FingerCodes and hence is extremely fast. Our initial results show identification accuracies comparable to the best results of minutiae-based algorithms published in the open literature. Finally, we show that the matching performance can be improved by combining the decisions of the matchers based on complementary fingerprint information.
Anil K. Jain 0001, Salil Prabhakar, Lin Hong, Sharath Pankanti
CVPR1
1999 Automatic Aircraft Recognition: Toward Using Human Similarity Measure in a Recognition System
abstract
The problem of screening images of the skies to determine whether they contain aircraft or not is both of theoretical and practical interest. After the most prominent visual signal in the infrared image of the sky is extracted, the question is whether the signal is a correct match of an aircraft. Common approaches calculate the degree of similarity of the shape of the signal with a model aircraft using a similarity measure such as Euclidean distance, and make a decision based on whether the degree of similarity exceeds a (pre-specified) threshold. Our approach avoids metric similarity measures and the use of thresholds as it attempts to employ similarity measures used by humans. In the absence of sufficient real data, the approach allows to specifically generate an arbitrarily large number of training exemplars projecting near classification boundary. Once trained on such a training set, the performance of the neural network was comparable to that of a human expert, and far better than a network trained only on the available real data. Furthermore, the results were considerably better than those obtained using a Euclidean discriminator.
Behrooz Kamgar-Parsi, Behzad Kamgar-Parsi, Anil K. Jain 0001
CVPR3
1999 Model-Guided Segmentation of Corpus Callosum in MR Images
abstract
Magnetic resonance imaging (MRI) of the brain, followed by automated segmentation of the corpus callosum (CC) in midsagittal sections has important applications in neurology and neurocognitive research since the size and shape of the CC are shown to be correlated to sex, age, neurodegenerative diseases and various lateralized behavior in man. Moreover, whole head, multispectral 3D MRI recordings enable voxel-based tissue classification and estimation of total brain volumes, in addition to CC morphometric parameters. We propose a new algorithm that uses both multispectral MRI measurements (intensity values) and prior information about shape (CC template) to segment CC in midsagittal slices with very little user interaction. The algorithm has been successfully tested on a sample of 10 subjects scanned with multispectral 3D MRI, collected for a study of dyslexia. We conclude that the proposed method for CC segmentation is promising for clinical use when multispectral MR images are recorded.
Arvid Lundervold, Torfinn Taxt, Nicolae Duta, Anil K. Jain 0001
CVPR4
1999 Learning-based Object Detection in Cardiac MR Images
abstract
An automated method for left ventricle detection in MR cardiac images is presented. Ventricle detection is the first step in a fully automated segmentation system used to compute volumetric information about the heart. Our method is based on learning the gray level appearance of the ventricle by maximizing the discrimination between positive and negative examples in a training set. The main differences from previously reported methods are feature definition and solution to the optimization problem involved in the learning process. Our method was trained on a set of 1,350 MR cardiac images from which 101,250 positive examples and 123,096 negative examples were generated. The detection results on a test set of 887 different images demonstrate an excellent performance: 98% detection rate, a false alarm rate of 0.05% of the number of windows analyzed (10 false alarms per image) and a detection time of 2 seconds per 256/spl times/256 image on a Sun Ultra 10 for an 8-scale search. The false alarms ore eventually eliminated by a position/scale consistency check along all the images that represent the same anatomical slice.
Nicolae Duta, Anil K. Jain 0001, Marie-Pierre Jolly
ICCV2
1999 Writer Adaptation of Online Handwriting Models
abstract
Writer adaptation is the process of converting a writer-independent handwriting recognition system, which models the characteristics of a large group of writers, into a writer-dependent system, which is tuned for a particular writer. Adaptation has the potential of increasing recognition accuracies, provided adequate models can be constructed for a particular writer. The limited amount of data that a writer typically provides makes the role of writer-independent models crucial in the adaptation process. Our approach to writer-adaptation makes use of writer-independent writing style models (called lexemes), to identify the styles present in a particular writer's training data. These models are then retrained using the writer's data. We demonstrate the feasibility of this approach using hidden Markov models trained on a combination of discretely and cursively written lower case characters. Our results show an average reduction in error rate of 16.3% for lower case characters as compared against representing each of the writer's character classes with a single model.
Scott D. Connell, Anil K. Jain 0001
ICDAR2
1999 Local Weight Selection for Two-Dimensional Phase Unwrapping
abstract
Local weights in phase unwrapping play an important role in guiding the flow of the phase integration. Unwrapping algorithms can proceed in a path-following or a non-path-following fashion. In path-following phase unwrapping algorithms, local weights guide the selection of unwrapping paths. In nonpath-following algorithms, the introduction of predetermined local weights is necessary to accommodate phase inconsistencies in practice. In current unwrapping algorithms, these local weights are extracted either from the quality information (correlation in interferometric synthetic aperture radar and modulation in phase shifting interferometry) or from the information contained in the wrapped phase difference alone. In other words, the crucial issue is how a phase unwrapping algorithm should select these location-associated weights in order to prevent the propagation of phase errors. This paper presents a quantitative scheme for selecting local weights for phase unwrapping opt the basis of both these types of information (the quality map and the wrapped phase). Real objects as well as synthetic ones have been investigated for various unwrapping methods in experiments. Integrated 3D free-form object models demonstrate the validity of the phasor-weighted phase unwrapping approach.
Rein-Lien Hsu, Shaoyun Chen, Anil K. Jain 0001, Carolyn R. Mercer
ICIP (3)3
1999 Deformable Matching of Hand Shapes for User Verification
abstract
We present a method for personal authentication based on deformable matching of hand shapes. Authentication systems are already employed in domains that require some sort of user verification. Unlike previous methods on hand shape based verification, our method aligns the hand shapes before extracting a feature set. We also base the verification decision on the shape distance which is automatically computed during the alignment stage. The shape distance proves to be a more reliable classification criterion than the handcrafted feature sets used by previous systems. Our verification system attained a high level of accuracy: 96.5% genuine accept rate vs. false accept rate. This performance is further improved by learning an enrolment template shape for each user.
Anil K. Jain 0001, Nicolae Duta
ICIP (2)1
1999 A Case Study in Using Human Similarity Measure for Automated Object Recognition
abstract
Image understanding often involves object recognition, where a basic question is how to decide whether a match is correct. Typically the best match (among a set of prestored objects) is assumed to be the correct match. This may work well in controlled environments (closed world). But, in uncontrolled environments (open world), the test object may not belong to the prestored object classes. In uncontrolled environments, a metric similarity measure (e.g. Euclidean) in conjunction with a threshold is used. However, based on psychophysical studies this is very different from, and far inferior to, human capabilities. To accept or reject a match, we introduce an approach that avoids metric similarity measures and the use of thresholds as it attempts to employ similarity measures used by humans. In the absence of sufficient real data, the approach allows to specifically generate an arbitrarily large number of training exemplars projecting near classification boundary. For aircraft detection, the performance of a neural network trained on such a training set, was comparable to that of a human expert, and far better than a network trained only on the available real data. Furthermore, the results were considerably better than those obtained using a Euclidean discriminator.
Behrooz Kamgar-Parsi, Behzad Kamgar-Parsi, Anil K. Jain 0001
ICIP (1)3
1999 Incremental Learning for Bayesian Classification of Images
abstract
Grouping images into (semantically) meaningful categories using low-level visual features is a challenging and important problem in content-based image retrieval. In this paper, we develop an incremental learning paradigm for Bayesian classification of images. Under the Bayesian paradigm, the class-conditional densities are represented in terms of codebook vectors. Learning is thus incrementally updating these codebook vectors as new training data become available. The proposed learning scheme estimates the already learnt training samples from the existing codebook vectors and augments these to the new training set for re-training the classifier. The above paradigm is shown to yield good results on three complex image classification problems. A classifier trained incrementally has comparable accuracies to the one which is trained using the true training samples.
Aditya Vailaya, Anil K. Jain 0001
ICIP (2)2
1999 Automatic Image Orientation Detection
abstract
We present an algorithm for automatic image orientation estimation using a Bayesian learning framework. We demonstrate that a small codebook (the optimal size of codebook is selected using a modified MDL criterion) extracted from a vector quantizer can be used to estimate the class-conditional densities of the observed features needed for the Bayesian methodology. We further show how feature clustering can be used as a feature selection mechanism to remove redundancies in the high-dimensional feature vectors used for classification. Experiments on a database of 17,901 images have shown that our proposed algorithm achieves an accuracy of approximately 97% on the training set and over 89% on an independent test set.
Aditya Vailaya, HongJiang Zhang, Anil K. Jain 0001
ICIP (2)3
1999 Automatic Caption Localization in Compressed Video
abstract
We present a method to automatically locate captions in MPEG video. Caption text regions are segmented from the background using their distinguishing texture characteristics. This method first locates candidate text regions directly in the DCT compressed domain, and then reconstructs the candidate regions for further refinement in the spatial domain. Therefore, only a small amount of decoding is required. The proposed algorithm achieves about 4.0% false reject rate and less than 5.7% false positive rate on a variety of MPEG compressed video containing more than 42,000 frames.
Yu Zhong 0001, HongJiang Zhang, Anil K. Jain 0001
ICIP (2)3
1999 A filterbank-based representation for classification and matching of fingerprints
abstract
We view fingerprints as oriented texture patterns, textures which exhibit an inherent and well-defined sense of directionality. Given a fingerprint image, we demonstrate that reliable translation and rotation invariant representations can be built based entirely on the inherent properties of the underlying fingerprint texture. We also illustrate that the representations thus derived are useful for robust discrimination of the fingerprints. In particular, we present the application of our novel representation scheme for solving the problems of fingerprint classification and matching on large datasets of fingerprint images acquired in real situations. The proposed scheme of generic representation for oriented textures relies on extracting one or more invariant frames of reference of the oriented texture based on an analysis of its orientation field. The fingerprint classification is based on a two-stage classifier which uses a K-NN classifier in its first stage and a set of neural network classifiers in its second stage.
Anil K. Jain 0001, Salil Prabhakar, Sharath Pankanti
IJCNN1
1999 Query by Video Clip
Anil K. Jain 0001, Aditya Vailaya
Multim. Syst.1
1999 A Multichannel Approach to Fingerprint Classification
abstract
Fingerprint classification provides an important indexing mechanism in a fingerprint database. An accurate and consistent classification can greatly reduce fingerprint matching time for a large database. We present a fingerprint classification algorithm which is able to achieve an accuracy better than previously reported in the literature. We classify fingerprints into five categories: whorl, right loop, left loop, arch, and tented arch. The algorithm uses a novel representation (FingerCode) and is based on a two-stage classifier to make a classification. It has been tested on 4000 images in the NIST-4 database. For the five-class problem, a classification accuracy of 90 percent is achieved (with a 1.8 percent rejection during the feature extraction phase). For the four-class problem (arch and tented arch combined into one class), we are able to achieve a classification accuracy of 94.8 percent (with 1.8 percent rejection). By incorporating a reject option at the classifier, the classification accuracy can be increased to 96 percent for the five-class classification task, and to 97.8 percent for the four-class classification task after a total of 32.5 percent of the images are rejected.
Anil K. Jain 0001, Salil Prabhakar, Lin Hong
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Automatic fruit recognition: a survey and new results using Range/Attenuation images
Antonio Ramón Jiménez, Anil K. Jain 0001, Ramón Ceres Ruíz, José Luis Pons Rovira
Pattern Recognit.2
1999 Combining multiple matchers for a high security fingerprint verification system
Anil K. Jain 0001, Salil Prabhakar, Shaoyun Chen
Pattern Recognit. Lett.1
1999 Two-dimensional phase unwrapping using a block least-squares method
abstract
We present a block least-squares (BLS) method for two-dimensional (2-D) phase unwrapping. The method works by tessellating the input image into small square blocks with only one phase wrap. These blocks are unwrapped using a simple procedure, and the unwrapped blocks are merged together using one of two proposed block merging algorithms. By specifying a suitable mask, the method can easily handle objects of any shape. This approach is compared with the Ghiglia-Romero method and the Marroquin-Rivera method. On synthetic images with different noise levels, the BLS method is shown to be superior, both with respect to the resulting gray values in the unwrapped image as well as visual inspection. The method is also shown to successfully unwrap synthetic and real images with shears, fiber-optic interferometry images, and medical magnetic resonance images. We believe the new method has the potential to improve the present quality of phase unwrapped images of several different image modalities.
Jarle Strand, Torfinn Taxt, Anil K. Jain 0001
IEEE Trans. Image Process.3
1999 Computer Vision Algorithms on Reconfigurable Logic Arrays
abstract
Computer vision algorithms are natural candidates for high performance computing systems. Algorithms in computer vision are characterized by complex and repetitive operations on large amounts of data involving a variety of data interactions (e.g., point operations, neighborhood operations, global operations). In this paper, we describe the use of the custom computing approach to meet the computation and communication needs of computer vision algorithms. By customizing hardware architecture at the instruction level for every application, the optimal grain size needed for the problem at hand and the instruction granularity can be matched. A custom computing approach can also reuse the same hardware by reconfiguring at the software level for different levels of the computer vision application. We demonstrate the advantages of our approach using Splash 2-a Xilinx 4010-based custom computer.
Nalini K. Ratha, Anil K. Jain 0001
IEEE Trans. Parallel Distributed Syst.2
1998 Integrating Faces and Fingerprints for Personal Identification
Lin Hong, Anil K. Jain 0001
ACCV (1)2
1998 Object Tracking Using Deformable Templates
abstract
We propose a novel method for object tracking using prototype-based deformable template models. To track an object in an image sequence, we use a criterion which combines two terms: the deviation of the object shape from its shape in the previous frame, and the fidelity of the detected shape to the input image. Shape and gradient information are used to track the object. We have also used the consistency between corresponding object regions throughout the sequence to help in trading the object of interest. Inter-frame motion is also used to track the boundary of moving objects. We have applied the algorithm to a number of image sequences from different sources. The inherent structure in the deformable template, together with region, motion, and image gradient cues, make the algorithm relatively insensitive to the adverse effects of weak image features and moderate partial occlusion.
Yu Zhong 0001, Anil K. Jain 0001, Marie-Pierre Jolly
ICCV2
1998 Learning prototypes for online handwritten digits
abstract
A writer independent handwriting recognition system must be able to recognize a wide variety of handwriting styles, while attempting to obtain a high degree of accuracy when recognizing data from any one of those styles. As the number of writing styles increases, so does the variability of the data's distribution. We then have an optimization problem: how to best model the data, while keeping the representation as simple as possible? If we can identify N different styles of writing individual characters (referred to as lexemes), these can then be modeled as N relatively simple independent distributions. We describe here a template-based system using a string-matching distance measure for the recognition of online handwriting which takes advantage of lexemes to reduce the number of templates that must be stored. A method of identifying lexemes and lexeme representatives is shown, and experimental results are given for a set of handwritten digits taken from 21 different writers. The use of lexeme representatives reduces classification time by 90.2% while retaining approximately 98% of the recognition accuracy.
Scott D. Connell, Anil K. Jain 0001
ICPR2
1998 Learning the human face concept in black and white images
abstract
Presents a learning approach for the face detection problem. The problem can be stated as follows: given an arbitrary black and white, still image, find the location and size of every human face it contains. Numerous applications of automatic face detection have attracted considerable interest in this problem, but no present face detection system is completely satisfactory from the point of view of detection rate, false alarm rate and detection time. We describe an inductive learning-based detection method that produces a maximally specific hypothesis consistent with the training data. Three different sets of features were considered for defining the concept of a human face. The performance achieved is as follows: 85% detection rate, a false alarm rate of 0.04% of the number of windows analyzed and 1 minute detection table for a 320/spl times/240 image on a Sun Ultrasparc 1.
Nicolae Duta, Anil K. Jain 0001
ICPR2
1998 F2ID: a personal identification system using faces and fingerprints
abstract
A real-time automatic personal identification system should meet the conflicting dual requirements of accuracy and response time. In addition, it also should be user-friendly. We introduce a medium-size realtime automatic personal identification system, F2ID, which integrates faces and fingerprints to make a personal identification. F2ID overcomes some of the limitations of face recognition systems and fingerprint verification systems and can achieve a desirable identification accuracy with a tolerable response time. We have tested our system on a limited set of face and fingerprint images collected in a laboratory environment. Experimental results show that that our system meets both the identification accuracy as well as the speed requirements.
Anil K. Jain 0001, Lin Hong, Yatin Kulkarni
ICPR1
1998 Query by video clip
abstract
Typical video search is based on queries involving a single shot. We generalize this problem by allowing queries that involve a video clip. We propose two schemes for query by video clip. In the first scheme, retrieval based on key frames where the database and query video are segmented into shots which are represented via key frames. For every query key frame, a similarity value (using color, texture, and motion) is associated with the key frames in the database video clip. Boundaries marking highly similar consecutive shots are then used to generate the set of retrieved video sub-clips. In the second scheme, in retrieval using sub-sampled frames, we uniformly sub-sample the query clip as well as the database video. The retrieval is based on matching color and texture features of the sub-sampled frames. Initial experiments on two video databases show promising results.
Anil K. Jain 0001, Aditya Vailaya
ICPR1
1998 Automatic text location in images and video frames
abstract
Automatic text location (without character recognition capabilities) deals with extracting image regions that contain text only. The images of these regions can then be fed to an optical character recognition module or highlighted for users. This is very useful in a number of applications such as database indexing and converting paper documents to their electronic versions. The performance of our automatic text location algorithm is shown in several applications. Compared with some traditional text location methods, our method has the following advantages: 1) low computational cost; 2) robust to font size; and 3) high accuracy.
Anil K. Jain 0001, Bin Yu 0002
ICPR1
1998 Classification of text documents
abstract
We investigate four different classification methods for document classification: the naive Bayes classifier, nearest neighbor classifier, decision tree classifier, and subspace method. The classifiers were applied to seven-class Yahoo newsgroups individually and in combination. We study three classifier combination approaches: simple voting, dynamic classifier selection, and adaptive classifier combination. Our experimental results indicate that the naive Bayes classifier and the subspace method outperform the other two classification methods on our data sets. Combinations of multiple classifiers did not always improve classification accuracy. Among the three different combination approaches, the adaptive classifier combination method proposed here performed the best.
Anil K. Jain 0001
ICPR2
1998 Image-based form document retrieval
abstract
We address the problem of image-based form document retrieval. The essential element of this problem is the definition of a similarity measure that is applicable in real situations, where query images are allowed to differ from the database images. Based on the definition of form signature, we have proposed a similarity measure that is insensitive to translation, scaling, moderate skew (<5/spl deg/) and variations in the geometrical proportions of the form layout. This similarity measure also has a good tolerance to line detection errors. We have developed a prototype form retrieval system which has been tested on a database containing 100 different kinds of forms.
Anil K. Jain 0001
ICPR2
1998 Classification of Text Documents
abstract
The exponential growth of the internet has led to a great deal of interest in developing useful and efficient tools and software to assist users in searching the Web. Document retrieval, categorization, routing and filtering can all be formulated as classification problems. However, the complexity of natural languages and the extremely high dimensionality of the feature space of documents have made this classification problem very difficult. We investigate four different methods for document classification: the naive Bayes classifier, the nearest neighbour classifier, decision trees and a subspace method. These were applied to seven-class Yahoo news groups (business, entertainment, health, international, politics, sports and technology) individually and in combination. We studied three classifier combination approaches: simple voting, dynamic classifier selection and adaptive classifier combination. Our experimental results indicate that the naive Bayes classifier and the subspace method outperform the other two classifiers on our data sets. Combinations of multiple classifiers did not always improve the classification accuracy compared to the best individual classifier. Among the three different combination approaches, our adaptive classifier combination method introduced here performed the best. The best classification accuracy that we are able to achieve on this seven-class problem is approximately 83%, which is comparable to the performance of other similar studies. However, the classification problem considered here is more difficult because the pattern classes used in our experiments have a large overlap of words in their corresponding documents.
Anil K. Jain 0001
Comput. J.2
1998 Extended Attributed String Matching for Shape Recognition
Sei-Wang Chen, S. T. Tung, Chiung-Yao Fang, Shen Cherng, Anil K. Jain 0001
Comput. Vis. Image Underst.5
1998 Registration and Integration of Multiple Object Views for 3D Model Construction
abstract
Automatic 3D object model construction is important in applications ranging from manufacturing to entertainment, since CAD models of existing objects may be either unavailable or unusable. We describe a prototype system for automatically registering and integrating multiple views of objects from range data. The results can then be used to construct geometric models of the objects. New techniques for handling key problems such as robust estimation of transformations relating multiple views and seamless integration of registered data to form an unbroken surface have been proposed and implemented in the system. Experimental results on real surface data acquired using a digital interferometric sensor as well as a laser range scanner demonstrate the good performance of our system.
Chitra Dorai, Anil K. Jain 0001, Carolyn R. Mercer
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 Integrating Faces and Fingerprints for Personal Identification
abstract
An automatic personal identification system based solely on fingerprints or faces is often not able to meet the system performance requirements. We have developed a prototype biometrics system which integrates faces and fingerprints. The system overcomes the limitations of face recognition systems as well as fingerprint verification systems. The integrated prototype system operates in the identification mode with an admissible response time. The identity established by the system is more reliable than the identity established by a face recognition system. In addition, the proposed decision fusion scheme enables performance improvement by integrating multiple cues with different confidence measures. Experimental results demonstrate that our system performs very well. It meets the response time as well as the accuracy requirements.
Lin Hong, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Fingerprint Image Enhancement: Algorithm and Performance Evaluation
abstract
In order to ensure that the performance of an automatic fingerprint identification/verification system will be robust with respect to the quality of input fingerprint images, it is essential to incorporate a fingerprint enhancement algorithm in the minutiae extraction module. We present a fast fingerprint enhancement algorithm, which can adaptively improve the clarity of ridge and valley structures of input fingerprint images based on the estimated local ridge orientation and frequency. We have evaluated the performance of the image enhancement algorithm using the goodness index of the extracted minutiae and the accuracy of an online fingerprint verification system. Experimental results show that incorporating the enhancement algorithm improves both the goodness index and the verification accuracy.
Lin Hong, Yifei Wan, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 Document Representation and Its Application to Page Decomposition
abstract
Transforming a paper document to its electronic version in a form suitable for efficient storage, retrieval, and interpretation continues to be a challenging problem. An efficient representation scheme for document images is necessary to solve this problem. Document representation involves techniques of thresholding, skew detection, geometric layout analysis, and logical layout analysis. The derived representation can then be used in document storage and retrieval. Page segmentation is an important stage in representing document images obtained by scanning journal pages. The performance of a document understanding system greatly depends on the correctness of page segmentation and labeling of different regions such as text, tables, images, drawings, and rulers. We use the traditional bottom-up approach based on the connected component extraction to efficiently implement page segmentation and region identification. A new document model which preserves top-down generation information is proposed based on which a document is logically represented for interactive editing, storage, retrieval, transfer, and logical analysis. Our algorithm has a high accuracy and takes approximately 1.4 seconds on a SGI Indy workstation for model creation, including orientation estimation, segmentation, and labeling (text, table, image, drawing, and ruler) for a 2550/spl times/3300 image of a typical journal page scanned at 300 dpi. This method is applicable to documents from various technical journals and can accommodate moderate amounts of skew and noise.
Anil K. Jain 0001, Bin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Large-Scale Parallel Data Clustering
abstract
Algorithmic enhancements are described that enable large computational reduction in mean square-error data clustering. These improvements are incorporated into a parallel data-clustering tool, P-CLUSTER, designed to execute on a network of workstations. Experiments involving the unsupervised segmentation of standard texture images were performed. For some data sets, a 96 percent reduction in computation was achieved.
Dan Judd, Philip K. McKinley, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1998 Shape-Based Retrieval: A Case Study With Trademark Image Databases
Anil K. Jain 0001, Aditya Vailaya
Pattern Recognit.1
1998 Automatic text location in images and video frames
Anil K. Jain 0001, Bin Yu 0002
Pattern Recognit.1
1998 On image classification: city images vs. landscapes
Aditya Vailaya, Anil K. Jain 0001, HongJiang Zhang
Pattern Recognit.2
1998 Deformable template models: A review
Anil K. Jain 0001, Yu Zhong 0001, Marie-Pierre Jolly
Signal Process.1
1997 Adaptive B-Splines and Boundary Estimation
abstract
This paper describes a boundary estimation scheme based on a new adaptive approach to B-spline curve fitting. The number of control points of the spline, their locations, and the observation parameters, are all considered unknown. The optimal number of control points is estimated via a new minimum description length (MDL) type criterion. The result is an adaptive parametrically deformable contour which also estimates the observation model parameters. Experiments on synthetic and real (medical) images confirm the adequacy and good performance of the approach.
Mário A. T. Figueiredo, José M. N. Leitão, Anil K. Jain 0001
CVPR3
1997 Page Segmentation Using Document Model
abstract
Transforming a paper document to its electronic version in a form suitable for efficient storage, retrieval and interpretation continues to be a challenging problem. An efficient document model is necessary to solve this problem. Document modeling involves techniques of thresholding, skew detection, geometric layout analysis and logical layout analysis. The derived model can then be used in document storage and retrieval. We use the traditional bottom-up approach based on the connected component extraction to efficiently implement page segmentation and region identification. A new document model which preserves top-down generation information is proposed based on which a document is logically represented for interactive editing, storage, retrieval, transfer and logical analysis.
Anil K. Jain 0001, Bin Yu 0002
ICDAR1
1997 Lane Boundary Detection Using a Multiresolution Hough Transform
abstract
Lane boundary detection is the problem of estimating the geometric structure of the lane boundaries of a road based on the images grabbed by a camera on board a vehicle. We use the Hough transform to detect lane boundaries with a parabolic model under a variety of road pavement types, lane structures and weather conditions. In the three-dimensional Hough space, a parabolic curve is represented as a straight line. To simplify the computation, the parametric space can be divided into (i) a two-dimensional space measured by the parameters which are shared by all the lane edges, and (ii) a one-dimensional space of the parameter which makes a distinction among different edges in an image. A multiresolution strategy is used to improve both the speed and accuracy of the Hough transform. Experimental results show that the proposed method is relatively less prone to the image noise and is computationally tractable.
Bin Yu 0002, Anil K. Jain 0001
ICIP (2)2
1997 Panel report: the potential of geons for generic 3-D object recognition
Sven J. Dickinson, Robert Bergevin, Irving Biederman, Jan-Olof Eklundh, Roger Munck-Fairwood, Anil K. Jain 0001, Alex Pentland
Image Vis. Comput.6
1997 COSMOS - A Representation Scheme for 3D Free-Form Objects
abstract
We address the problem of representing and recognizing 3D free-form objects when (1) the object viewpoint is arbitrary, (2) the objects may vary in shape and complexity, and (3) no restrictive assumptions are made about the types of surfaces on the object. We assume that a range image of a scene is available, containing a view of a rigid 3D object without occlusion. We propose a new and general surface representation scheme for recognizing objects with free-form (sculpted) surfaces. In this scheme, an object is described concisely in terms of maximal surface patches of constant shape index. The maximal patches that represent the object are mapped onto the unit sphere via their orientations, and aggregated via shape spectral functions. Properties such as surface area, curvedness, and connectivity, which are required to capture local and global information, are also built into the representation. The scheme yields a meaningful and rich description useful for object recognition. A novel concept, the shape spectrum of an object is also introduced within the framework of COSMOS for object view grouping and matching. We demonstrate the generality and the effectiveness of our scheme using real range images of complex objects.
Chitra Dorai, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Shape Spectrum Based View Grouping and Matching of 3D Free-Form Objects
abstract
We address the problem of constructing view aspects of 3D free-form objects for efficient matching during recognition. We introduce a novel view representation based on "shape spectrum" features, and propose a general and powerful technique for organizing multiple views of objects of complex shape and geometry into compact and homogeneous clusters. Our view grouping technique obviates the need for surface segmentation and edge detection. Experiments on 6,400 synthetically generated views of 20 free-form objects and 100 real range images of 10 sculpted objects demonstrate the good performance of our shape spectrum based model view selection technique.
Chitra Dorai, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Optimal Registration of Object Views Using Range Data
abstract
This paper deals with robust registration of object views in the presence of uncertainties and noise in depth data. Errors in registration of multiple views of a 3D object severely affect view integration during automatic construction of object models. We derive a minimum variance estimator (MVE) for computing the view transformation parameters accurately from range data of two views of a 3D object. The results of our experiments show that view transformation estimates obtained using MVE are significantly more accurate than those computed with an unweighted error criterion for registration.
Chitra Dorai, Juyang Weng, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 lamda-tau-Space Representation of Images and Generalized Edge Detector
abstract
An image and surface representation based on regularization theory is introduced in this paper. This representation is based on a hybrid model derived from the physical membrane and plate models. The representation, called the /spl lambda//spl tau/-representation, has two dimensions; one dimension represents smoothness or scale while the other represents the continuity of the image or surface. It contains images/surfaces sampled both in scale space and the weighted Sobolev space of continuous functions. Thus, this new representation can be viewed as an extension of the well-known scale space representation. We have experimentally shown that the proposed hybrid model results in improved results compared to the two extreme constituent models, i.e., the membrane and the plate models. Based on this hybrid model, a generalized edge detector (GED) which encompasses most of the well-known edge detectors under a common framework is developed. The existing edge detectors can be obtained from the generalized edge detector by simply specifying the values of two parameters, one of which controls the shape of the filter (/spl tau/) and the other controls the scale of the filter (/spl lambda/). By sweeping the values of these two parameters continuously, one can generate an edge representation in the /spl lambda//spl tau/ space, which is very useful for developing a goal-directed edge detection scheme for a specific task. The proposed representation and the edge detector have been evaluated qualitatively and quantitatively on several different types of image data such as intensity, range, and stereo images.
Muhittin Gökmen, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 On-Line Fingerprint Verification
abstract
Fingerprint verification is one of the most reliable personal identification methods. However, manual fingerprint verification is incapable of meeting today's increasing performance requirements. An automatic fingerprint identification system (AFIS) is needed. This paper describes the design and implementation of an online fingerprint verification system which operates in two stages: minutia extraction and minutia matching. An improved version of the minutia extraction algorithm proposed by Ratha et al. (1995), which is much faster and more reliable, is implemented for extracting features from an input fingerprint image captured with an online inkless scanner. For minutia matching, an alignment-based elastic matching algorithm has been developed. This algorithm is capable of finding the correspondences between minutiae in the input image and the stored template without resorting to exhaustive search and has the ability of adaptively compensating for the nonlinear deformations and inexact pose transformations between fingerprints. The system has been tested on two sets of fingerprint images captured with inkless scanners. The verification accuracy is found to be acceptable. Typically, a complete fingerprint verification procedure takes, on an average, about eight seconds on a SPARC 20 workstation. These experimental results show that our system meets the response time requirements of online verification with high accuracy.
Anil K. Jain 0001, Lin Hong, Ruud M. Bolle
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Representation and Recognition of Handwritten Digits Using Deformable Templates
abstract
We investigate the application of deformable templates to recognition of handprinted digits. Two characters are matched by deforming the contour of one to fit the edge strengths of the other, and a dissimilarity measure is derived from the amount of deformation needed, the goodness of fit of the edges, and the interior overlap between the deformed shapes. Classification using the minimum dissimilarity results in recognition rates up to 99.25 percent on a 2,000 character subset of NIST Special Database 1. Additional experiments on an independent test data were done to demonstrate the robustness of this method. Multidimensional scaling is also applied to the 2,000/spl times/2,000 proximity matrix, using the dissimilarity measure as a distance, to embed the patterns as points in low-dimensional spaces. A nearest neighbor classifier is applied to the resulting pattern matrices. The classification accuracies obtained in the derived feature space demonstrate that there does exist a good low-dimensional representation space. Methods to reduce the computational requirements, the primary limiting factor of this method, are discussed.
Anil K. Jain 0001, Douglas E. Zongker
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Feature Selection: Evaluation, Application, and Small Sample Performance
abstract
A large number of algorithms have been proposed for feature subset selection. Our experimental results show that the sequential forward floating selection algorithm, proposed by Pudil et al. (1994), dominates the other algorithms tested. We study the problem of choosing an optimal feature set for land use classification based on SAR satellite images using four different texture models. Pooling features derived from different texture models, followed by a feature selection results in a substantial improvement in the classification accuracy. We also illustrate the dangers of using feature selection in small sample size situations.
Anil K. Jain 0001, Douglas E. Zongker
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 Obituary: Pierre Devijver
Josef Kittler, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1997 Recognition of Digits in Hydrographic Maps: Binary Versus Topographic Analysis
abstract
Compares the performance of topographic analysis and binary analysis for recognition of digits in hydrographic maps. The performance of each method was measured by the correct classification rate of the final symbol recognition step when processing a complete hydrographic map of size 0.45/spl times/0.6 m/sup 2/ with about 35000 digits. The experimental results indicated that binary analysis had a better performance than topographic analysis. Overall, the performance of the binary analysis was acceptable.
Øivind Due Trier, Torfinn Taxt, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1997 An identity-authentication system using fingerprints
abstract
Fingerprint verification is an important biometric technique for personal identification. We describe the design and implementation of a prototype automatic identity-authentication system that uses fingerprints to authenticate the identity of an individual. We have developed an improved minutiae-extraction algorithm that is faster and more accurate than our earlier algorithm (1995). An alignment-based minutiae-matching algorithm has been proposed. This algorithm is capable of finding the correspondences between input minutiae and the stored template without resorting to exhaustive search and has the ability to compensate adaptively for the nonlinear deformations and inexact transformations between an input and a template. To establish an objective assessment of our system, both the Michigan State University and the National Institute of Standards and Technology NIST 9 fingerprint data bases have been used to estimate the performance numbers. The experimental results reveal that our system can achieve a good performance on these data bases. We also have demonstrated that our system satisfies the response-time requirement. A complete authentication procedure, on average, takes about 1.4 seconds on a Sun ULTRA I workstation (it is expected to run as fast or faster on a 200 HMz Pentium).
Anil K. Jain 0001, Lin Hong, Sharath Pankanti, Ruud M. Bolle
Proc. IEEE1
1997 Mobile robot localization in indoor environment
Hansye S. Dulimart, Anil K. Jain 0001
Pattern Recognit.2
1997 Practicing vision: Integration, evaluation and applications
Anil K. Jain 0001, Chitra Dorai
Pattern Recognit.1
1997 Object detection using gabor filters
Anil K. Jain 0001, Nalini K. Ratha, Sridhar Lakshmanan
Pattern Recognit.1
1997 Editorial
John Chung-Mong Lee, Anil K. Jain 0001
Pattern Recognit.2
1997 Texture fusion and feature selection applied to SAR imagery
abstract
The discrimination ability of four different methods for texture computation in ERS SAR imagery is examined and compared. Feature selection methodology and discriminant analysis are applied to find the optimal combination of texture features. By combining features derived from different texture models, the classification accuracy increased significantly.
Anne H. Schistad Solberg, Anil K. Jain 0001
IEEE Trans. Geosci. Remote. Sens.2
1997 Guest Editorial Special Issue on Artificial Neural Networks and Statistical Pattern Recognition
Anil K. Jain 0001, Jianchang Mao
IEEE Trans. Neural Networks1
1997 Authors' Reply
Jianchang Mao, Anil K. Jain 0001
IEEE Trans. Neural Networks2
1997 Comments on "A self-organizing network for hyperellipsoidal clustering (HEC)" [and reply]
abstract
In the above paper by Mao-Jain (ibid., vol.7 (1996)), the Mahalanobis distance is used instead of Euclidean distance as the distance measure in order to acquire the hyperellipsoidal clustering. We prove that the clustering cost function is a constant under this condition, so hyperellipsoidal clustering cannot be realized. We also explains why the clustering algorithm developed in the above paper can get some good hyperellipsoidal clustering results. In reply, Mao-Jain state that the Wang-Xia failed to point out that their HEC clustering algorithm used a regularized Mahalanobis distance instead of the standard Mahalanobis distance. It is the regularized Mahalanobis distance which plays an important role in realizing hyperellipsoidal clusters. In conclusion, the comments made by Wang-Xia together with this response provide some new insights into the behavior of their HEC clustering algorithm. It further confirms that the HEC algorithm is a useful tool for understanding the structure of multidimensional data.
Wang Song, Shaowei Xia, Jianchang Mao, Anil K. Jain 0001, Danil V. Prokhorov, Donald C. Wunsch II
IEEE Trans. Neural Networks4
1996 λτ-space representation of images and generalized edge detector
abstract
An image clad surface representation based on regularization theory is introduced in this paper. This representation is based on a hybrid model derived from the physical membrane and plate models. The representation, called the /spl lambda//spl tau/-representation, has two dimensions; one dimension represents smoothness or scale while the other represents the continuity of the image or surface. It contains images/surfaces sampled both in scale space and the weighted Sobolev space of continuous functions. Thus, this new representation can be viewed as an extension of the well-known scale space representation. We have experimentally shown that the proposed hybrid model results in improved results compared to the two extreme constituent model, i.e., the membrane and the plate models. Based on this hybrid model, a generalized edge detector (GED) which encompasses most of the well-known edge detectors under a common framework is developed. The existing edge detectors can be obtained from the generalized edge detector by simply specifying the valves of two parameters, one of which controls the shape of the filter (/spl tau/) and the other controls the scale of the filter (/spl lambda/). By sweeping the valves of these two parameters continuously, one can generate an edge representation in the /spl lambda//spl tau/ space, which is very useful for developing a goal-directed edge detection scheme for a specific task. The proposed representation and the edge detector have been evaluated qualitatively and quantitatively on several different types of image data such as intensity, range and stereo images.
Muhittin Gökmen, Anil K. Jain 0001
CVPR2
1996 FPGA-based high performance page layout segmentation
abstract
A page layout segmentation algorithm for locating text, background and halftone areas is presented. The algorithm has been implemented on Splash 2-an FPGA-based array processor. The speed as determined by the Xilinx synthesis tools projects an application speed of 5 MHz. For documents of size 1,024/spl times/1,024 pixels, a significant speedup of two orders of magnitude compared to a SparcStation 20 has been achieved.
Nalini K. Ratha, Anil K. Jain 0001, Diane T. Rover
Great Lakes Symposium on VLSI2
1996 Compression of fingerprint images using hybrid image model
abstract
We present an efficient model-based fingerprint image compression scheme based on a hybrid image model. Our model is based on extracting ridge and valley contours and then reconstructing a hybrid surface by using the gray values on these contours. The hybrid model we utilized is the convex combination of the membrane and plate functionals used for surface reconstruction by regularization. Two parameters of this model is determined to obtain a good approximation of the original fingerprint image given the sparse data, on ridges and valleys. In this compression scheme, the ridge contours are coded efficiently by using a differential chain code, while the differences between consecutive gray values along the chains are encoded using Huffman coding. Also included in the compressed image are the mean value of each valley segment and two parameters of the hybrid model. One advantage of our approach as compared to transform based algorithms and wavelet-based algorithms is that features such as delta and core points, end points, bifurcation points can be extracted directly from compressed image even for very high values of the compression ratio. The algorithm has been applied to various fingerprint images, and high compression ratios like 45:1 have been obtained while keeping all the important features in the images.
Muhittin Gökmen, Ilker Ersoy, Anil K. Jain 0001
ICIP (3)3
1996 Recognition of 3D free-form objects
abstract
We address the problem of recognizing 3D rigid free-form objects using dense range data when the objects can be imaged from arbitrary viewpoints and the objects vary in shape and complexity. We propose a multi-level matching strategy that employs shape spectral analysis and features derived from the COSMOS representations of free-form objects for fast and efficient recognition. We demonstrate that with a large model database of object views, a small set of ranked candidate matches can be selected quickly using shape spectrum based matching for further verification. We propose a graph-based matching scheme for view hypothesis verification using COSMOS representations of object views to establish the correct identity and the pose of the sensed object. Preliminary experimental results on a database containing views often different objects are shown to demonstrate the effectiveness of the COSMOS-based 3D object recognition system.
Chitra Dorai, Anil K. Jain 0001
ICPR2
1996 From images to models: automatic 3D object model construction from multiple views
abstract
Automatic 3D object model construction is important in applications ranging from manufacturing to entertainment industry, since CAD models of existing objects may be either unavailable or unusable. We describe a prototype system for automatically registering and integrating multiple views of objects using surface depth data to construct their geometric models. New techniques for handling key issues such as robust estimation of view transformations relating multiple views and seamless integration of registered data to form an unbroken surface have been proposed and implemented in our system. Our system has been tested on depth data that were obtained using digital interferometry techniques as well as a laser range scanner. Experimental results on real object surfaces demonstrate the good performance of our system.
Chitra Dorai, Anil K. Jain 0001, Carolyn R. Mercer
ICPR3
1996 On-line fingerprint verification
abstract
We describe the design and implementation of an online fingerprint verification system which operates in two stages: (i) minutia extraction and (ii) minutia matching. An improved minutia extraction algorithm that is much faster and more accurate than our earlier algorithm has been implemented. For minutia matching, an alignment-based elastic matching algorithm has been developed. This algorithm is capable of finding the correspondences between input minutiae and the stored template without resorting to exhaustive search and has the ability to adaptively compensate for the nonlinear deformations and inexact pose transformations between finger prints. The system has been tested on two sets of finger print images captured with inkless scanners. The verification accuracy is found to be over 99% with a 15% reject rate. Typically, a complete fingerprint verification procedure takes, on an average, about 8 seconds on a SPARC 20 workstation. It meets the response time requirements of on-line verification with high accuracy.
Anil K. Jain 0001, Lin Hong
ICPR1
1996 Large-scale parallel data clustering
abstract
Algorithmic enhancements are described that allow large reduction (for some data sets, over 95 percent) in the number of floating point operations in mean square error data clustering. These improvements are incorporated into a parallel data clustering tool, P-CLUSTER, developed in an earlier study. Experiments on segmenting standard texture images show that the proposed enhancements enable clustering of an entire 512/spl times/512 image at approximately the same computational cost as that of previous methods applied to only 5 percent of the image pixels.
Dan Judd, Philip K. McKinley, Anil K. Jain 0001
ICPR3
1996 Is there any texture in the image?
abstract
Texture analysis methods have been used in various image processing tasks, such as image segmentation, recognition, shape analysis, texture synthesis, and image compression. When applying any of these methods, we assume that the input image has some textural characteristics. This paper addresses the problem of deciding whether an image has texture; in other words, whether texture-based methods are suitable for processing the image. We define a texture to have a spatially uniform distribution of local gray-value variations. A fast algorithm for detecting regions that have texture according to this definition is presented. The performance of the method is demonstrated on several synthetic and natural images.
Kalle Karu, Anil K. Jain 0001, Ruud M. Bolle
ICPR2
1996 Image databases: a case study in Norwegian silver authentication
abstract
Establishing the identity of the craftsmen of old Norwegian silver objects can be accomplished through the recognition of master-marks. Manually searching through the 6000 master-marks and 250 city-marks is tedious and unreliable. Our-goal is to develop an efficient retrieval of images of master- and city-marks from large databases based on the shape content of the image. We are building an automatic retrieval system to browse through the entire database. The first step is to extract primitive visual features from the images and to retrieve images on the basis of these features. We have developed a prototype image database of master-marks from a catalog by scanning 171 of these marks which appear on silver tankards. We have successfully extracted shape features from these marks and matched master-marks in the database against the unknown query images. Initial experiments on master-marks are encouraging. The same shape-based matching techniques are used to match city-marks. Our future work will involve building an expert system for a stylistic analysis of the silver tankards.
Britt Kroepelien, Aditya Vailaya, Anil K. Jain 0001
ICPR3
1996 Gray scale processing of hydrographic maps
abstract
This paper investigates how gray scale information can be used in a hydrographic map understanding system to improve the system performance. To process gray scale scanned map images, we have implemented a topographic analysis method and a binary analysis method. In addition, deconvolution of the gray scale map image was used as an optional preprocessing step for both the methods. Both the methods process the input image by extracting binary print components, recognizing long lines, splitting touching digits and recognizing the digits. The topographic analysis extracts the information by computing topographic labels for each pixel, while the binary analysis is based on locally adaptive thresholding of the gray scale image. The performance of each method was evaluated by measuring the recognition performance of the digit recognition module. Experimental results indicate that the computationally intensive deconvolution and topographic analysis does not improve system performance. The same high performance is achieved by binary analysis, provided a high quality locally adaptive binary method is used.
Øivind Due Trier, Torfinn Taxt, Anil K. Jain 0001
ICPR3
1996 A hierarchical system for efficient image retrieval
abstract
Retrieval efficiency and accuracy are two important issues in designing a content-based database retrieval system. We propose a new image database retrieval method based on shape information. This system achieves both the desired efficiency and accuracy using a two-stage hierarchy: in the first stage, simple and easily computable statistical shape features are used to quickly browse through the database to generate a moderate number of plausible retrievals; in the second stage, the outputs from the first stage are screened using a deformable template matching process to discard spurious matches. We have tested the algorithm using hand drawn queries on a trademark database containing 1,100 images. Each retrieval takes a reasonable amount of computation time. The top most retrieved image from the system agrees with that obtained by human subjects, but there are significant differences between the top 10 retrieved images by our system and that provided by human subjects. This demonstrates the need for developing shape features that are better able to capture human perceptual similarity of shapes.
Aditya Vailaya, Yu Zhong 0001, Anil K. Jain 0001
ICPR3
1996 A form dropout system
abstract
This paper describes a system for form dropout when the filled-in characters or symbols are either touching or crossing the form frames and the form model is unknown. Since some of the character strokes are either touching or crossing the form frames, we need to address the following three issues: (i) localization of form frames; (ii) separation between characters and form frames, and (ii) reconstruction of broken strokes introduced during separation. The form frame is automatically located by finding long straight lines based on a data structure, called block adjacency graph. Form frame removal and character reconstruction are implemented in this graph. When the same process is applied to a blank form, followed by the procedure of connected component extraction and clustering, a form structure-based template is automatically generated which includes form model, skew angle and preprinted data areas. Given the form template, our system can extract both handwritten and machine-typed filled-in data. Experimental results on three different types of forms demonstrate the performance of our system.
Bin Yu 0002, Anil K. Jain 0001
ICPR2
1996 Algorithms for feature selection: An evaluation
abstract
A large number of algorithms have been proposed for doing feature subset selection. The goal of this paper is to evaluate the quality of feature subsets generated by the various algorithms, and also compare their computational requirements. Our results show that the sequential forward floating selection (SFFS) algorithm, proposed by Pudil et al. (1994), dominates the other algorithms tested. This paper also illustrates the dangers of using feature selection in small sample size situations. It gives the results of applying feature selection to land use classification of SAR satellite images using four different texture models. Pooling features derived from different texture models, followed by a feature selection results in a substantial improvement in the classification accuracy. Application of feature selection to classification of handprinted characters illustrates the value of feature selection in reducing the number of features needed for classifier design.
Douglas E. Zongker, Anil K. Jain 0001
ICPR2
1996 Reconstruction and Boundary Detection of Range and Intensity Images Using Multiscale MRF Representations
Bilge Günsel, Anil K. Jain 0001, Erdal Panayirci
Comput. Vis. Image Underst.2
1996 Vehicle Segmentation and Classification Using Deformable Templates
abstract
This paper proposes a segmentation algorithm using deformable template models to segment a vehicle of interest both from the stationary complex background and other moving vehicles in an image sequence. We define a polygonal template to characterize a general model of a vehicle and derive a prior probability density function to constrain the template to be deformed within a set of allowed shapes. We propose a likelihood probability density function which combines motion information and edge directionality to ensure that the deformable template is contained within the moving areas in the image and its boundary coincides with strong edges with the same orientation in the image. The segmentation problem is reduced to a minimization problem and solved by the Metropolis algorithm. The system was successfully tested on 405 image sequences containing multiple moving vehicles on a highway.
Marie-Pierre Jolly, Sridhar Lakshmanan, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
1996 Learning Texture Discrimination Masks
abstract
A neural network texture classification method is proposed in this paper. The approach is introduced as a generalization of the multichannel filtering method. Instead of using a general filter bank, a neural network is trained to find a minimal set of specific filters, so that both the feature extraction and classification tasks are performed by the same unified network. The authors compute the error rates for different network parameters, and show the convergence speed of training and node pruning algorithms. The proposed method is demonstrated in several texture classification experiments. It is successfully applied in the tasks of locating barcodes in the images and segmenting a printed page into text, graphics, and background. Compared with the traditional multichannel filtering method, the neural network approach allows one to perform the same texture classification or segmentation task more efficiently. Extensions of the method, as well as its limitations, are discussed in the paper.
Anil K. Jain 0001, Kalle Karu
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Object Matching Using Deformable Templates
abstract
We propose a general object localization and retrieval scheme based on object shape using deformable templates. Prior knowledge of an object shape is described by a prototype template which consists of the representative contour/edges, and a set of probabilistic deformation transformations on the template. A Bayesian scheme, which is based on this prior knowledge and the edge information in the input image, is employed to find a match between the deformed template and objects in the image. Computational efficiency is achieved via a coarse-to-fine implementation of the matching algorithm. Our method has been applied to retrieve objects with a variety of shapes from images with complex background. The proposed scheme is invariant to location, rotation, and moderate scale changes of the template.
Anil K. Jain 0001, Yu Zhong 0001, Sridhar Lakshmanan
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Parameter Estimation in Markov Random Field Contextual Models Using Geometric Models of Objects
abstract
We present a new scheme for the estimation of Markov random field line process parameters which uses geometric CAD models of the objects in the scene. The models are used to generate synthetic images of the objects from random view points. The edge maps computed from the synthesized images are used as training samples to estimate the line process parameters using a least squares method. We show that this parameter estimation method is useful for detecting edges in range as well as intensity edges. The main contributions of the paper are: 1) use of CAD models to obtain true edge labels which are otherwise not available; and 2) use of canonical Markov random field representation to reduce the number of parameters.
Sateesha G. Nadabar, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 A Real-Time Matching System for Large Fingerprint Databases
abstract
With the current rapid growth in multimedia technology, there is an imminent need for efficient techniques to search and query large image databases. Because of their unique and peculiar needs, image databases cannot be treated in a similar fashion to other types of digital libraries. The contextual dependencies present in images, and the complex nature of two-dimensional image data make the representation issues more difficult for image databases. An invariant representation of an image is still an open research issue. For these reasons, it is difficult to find a universal content-based retrieval technique. Current approaches based on shape, texture, and color for indexing image databases have met with limited success. Further, these techniques have not been adequately tested in the presence of noise and distortions. A given application domain offers stronger constraints for improving the retrieval performance. Fingerprint databases are characterized by their large size as well as noisy and distorted query images. Distortions are very common in fingerprint images due to elasticity of the skin. In this paper, a method of indexing large fingerprint image databases is presented. The approach integrates a number of domain-specific high-level features such as pattern class and ridge density at higher levels of the search. At the lowest level, it incorporates elastic structural feature-based matching for indexing the database. With a multilevel indexing approach, we have been able to reduce the search space. The search engine has also been implemented on Splash 2-a field programmable gate array (FPGA)-based array processor to obtain near-ASIC level speed of matching. Our approach has been tested on a locally collected test data and on NIST-9, a large fingerprint database available in the public domain.
Nalini K. Ratha, Kalle Karu, Shaoyun Chen, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
1996 A Generic System for Form Dropout
abstract
Recent advances in intelligent character recognition are enabling us to address many challenging problems in document image analysis. One of them is intelligent form analysis. This paper describes a generic system for form dropout when the filled-in characters or symbols are either touching or crossing the form frames. We propose a method to separate these characters from form frames whose locations are unknown. Since some of the character strokes are either touching or crossing the form frames, we need to address the following three issues: 1) localization of form frames; 2) separation of characters and form frames; and 3) reconstruction of broken strokes introduced during separation. The form frame is automatically located by finding long straight lines based on the block adjacency graph. Form frame separation and character reconstruction are implemented by means of this graph. The proposed system includes form structure learning and form dropout. First, a form structure-based template is automatically generated from a blank form which includes form frames, preprinted data areas and skew angle. With this form template, our system can then extract both handwritten and machine-typed filled-in data. Experimental results on three different types of forms show the performance of our system. Further, the proposed method is robust to noise and skew that is introduced during scanning.
Bin Yu 0002, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1996 A client/server control architecture for robot navigation
Hansye S. Dulimart, Anil K. Jain 0001
Pattern Recognit.2
1996 On retrieving textured images from an image database
Georgy L. Gimel'farb, Anil K. Jain 0001
Pattern Recognit.2
1996 Image retrieval using color and shape
Anil K. Jain 0001, Aditya Vailaya
Pattern Recognit.1
1996 Page segmentation using tecture analysis
Anil K. Jain 0001, Yu Zhong 0001
Pattern Recognit.1
1996 Fingerprint classification
Kalle Karu, Anil K. Jain 0001
Pattern Recognit.2
1996 Is there any texture in the image?
Kalle Karu, Anil K. Jain 0001, Ruud M. Bolle
Pattern Recognit.2
1996 Feature extraction methods for character recognition-A survey
Øivind Due Trier, Anil K. Jain 0001, Torfinn Taxt
Pattern Recognit.2
1996 A robust and fast skew detection algorithm for generic documents
Bin Yu 0002, Anil K. Jain 0001
Pattern Recognit.2
1996 Pre/post-filter for performance improvement of transform coding
Chung J. Kuo, John R. Deller Jr., Anil K. Jain 0001
Signal Process. Image Commun.3
1996 A Markov random field model for classification of multisource satellite imagery
abstract
A general model for multisource classification of remotely sensed data based on Markov random fields (MRF) is proposed. A specific model for fusion of optical images, synthetic aperture radar (SAR) images, and GIS (geographic information systems) ground cover data is presented in detail and tested. The MRF model exploits spatial class dependencies (spatial context) between neighboring pixels in an image, and temporal class dependencies between different images of the same scene. By including the temporal aspect of the data, the proposed model is suitable for detection of class changes between the acquisition dates of different images. The performance of the proposed model is investigated by fusing Landsat TM images, multitemporal ERS-1 SAR images, and GIS ground-cover maps for land-use classification, and on agricultural crop classification based on Landsat TM images, multipolarization SAR images, and GIS crop field border maps. The performance of the MRF model is compared to a simpler reference fusion model. On an average, the MRF model results in slightly higher (2%) classification accuracy when the same data is used as input to the two models. When GIS field border data is included in the MRF model, the classification accuracy of the MRF model improves by 8%. For change detection in agricultural areas, 75% of the actual class changes are detected by the MRF model, compared to 62% for the reference model. Based on the well-founded theoretical basis of Markov random field models for classification tasks and the encouraging experimental results in our small-scale study, the authors conclude that the proposed MRF model is useful for classification of multisource satellite imagery.
Anne H. Schistad Solberg, Torfinn Taxt, Anil K. Jain 0001
IEEE Trans. Geosci. Remote. Sens.3
1996 A self-organizing network for hyperellipsoidal clustering (HEC)
abstract
We propose a self-organizing network for hyperellipsoidal clustering (HEC). It consists of two layers. The first employs a number of principal component analysis subnetworks to estimate the hyperellipsoidal shapes of currently formed clusters. The second performs competitive learning using the cluster shape information from the first. The network performs partitional clustering using the proposed regularized Mahalanobis distance, which was designed to deal with the problems in estimating the Mahalanobis distance when the number of patterns in a cluster is less than or not considerably larger than the dimensionality of the feature space during clustering. This distance also achieves a tradeoff between hyperspherical and hyperellipsoidal cluster shapes so as to prevent the HEC network from producing unusually large or small clusters. The significance level of the Kolmogorov-Smirnov test on the distribution of the Mahalanobis distances of patterns in a cluster to the cluster center under the Gaussian cluster assumption is used as a compactness measure. The HEC network has been tested on a number of artificial data sets and real data sets, We also apply the HEC network to texture segmentation problems. Experiments show that the HEC network leads to a significant improvement in the clustering results over the K-means algorithm with Euclidean distance. Our results on real data sets also indicate that hyperellipsoidal shaped clusters are often encountered in practice.
Jianchang Mao, Anil K. Jain 0001
IEEE Trans. Neural Networks2
1995 Convolution on Splash 2
abstract
Convolution is a fundamental operation in many signal and image processing applications. Since the computation and communication pattern in a convolution operation is regular, a number of special architectures have been designed and implemented for this operator. The Von Neumann architectures cannot meet the real-time requirements of applications that use convolution as an intermediate step. We combine the advantages of systolic algorithms with the low cost of developing application specific designs using field programmable gate arrays (FPGAs) to build a scalable convolver for use in computer vision systems. The performance of the systolic algorithm of (Kung et al., 1981) is compared theoretically and experimentally with many other convolution algorithms reported in the literature. The implementation of a convolution operation on Splash 2, an attached processor based on Xilinx 4010 FPGAs, is reported with impressive performance gains.
Nalini K. Ratha, Anil K. Jain 0001, Diane T. Rover
FCCM2
1995 COSMOS-A Representation Scheme for Free-Form Surfaces
abstract
We address the problem of representing and recognizing arbitrarily curved 3D rigid objects when: the objects may vary in shape and complexity, and no restrictive assumptions are made about the types of surfaces on the object. We propose a new and general surface representation scheme for recognizing objects with free form (sculpted) surfaces from range data. In this scheme, an object is described concisely in terms of maximal surface patches of constant shape index. These maximal patches are mapped onto the unit sphere via their orientations, and aggregated via shape spectral functions. Properties such as surface area, curvedness and connectivity that capture local and global information are also built into the representation. The scheme yields not only a meaningful and rich surface description useful for the recoverability of the object, but also a set of powerful indexing primitives for object matching. We demonstrate the generality and the effectiveness of our scheme using real range images of complex objects. We also present results on the categorization of object views based on a novel shape spectral matching technique.>
Chitra Dorai, Anil K. Jain 0001
ICCV2
1995 Data capture from maps based on gray scale topographic analysis
abstract
There is a large number of documents, including hand-printed maps, where useful information is lost if binarization is performed on the scanned image before further processing. For such documents, methods which utilize the gray scale values must be used in order to extract as much of the available information as possible. Topographic analysis has been used in the literature to recognize characters directly in the gray scale images. The authors extend the topographic analysis method, so that characters and lines can be extracted from gray scale map images where methods using only the information in the binary image fail.
Øivind Due Trier, Torfinn Taxt, Anil K. Jain 0001
ICDAR3
1995 Locating text in complex color images
abstract
There is a substantial interest in retrieving images from a large database using the textual information contained in the images. An algorithm which will automatically locate the textual regions in the input image will facilitate this task; the optical character recognizer can then be applied to only those regions of the image which contain text. We present a method for automatically locating text in complex color images. The algorithm first finds the approximate locations of text lines using horizontal spatial variance, and then extracts text components in these boxes using color segmentation. The proposed method has been used to locate text in compact disc (CD) and book cover images, as well as in the images of traffic scenes captured by a video camera. Initial results are encouraging and suggest that these algorithms can be used in image retrieval applications.
Yu Zhong 0001, Kalle Karu, Anil K. Jain 0001
ICDAR3
1995 Shape spectra based view grouping for free-form objects
abstract
The concept of "shape spectrum" of an object view is introduced for constructing view aspects of free-form objects. The shape spectrum captures the shape characteristics of an object by aggregating the areas of the surfaces of the object at each value of the shape index. We derive a novel view representation based on shape spectral features and propose a general and powerful technique for organizing multiple views of objects of complex shape and geometry into compact and homogeneous clusters. Our view-grouping technique obviates the need for surface segmentation and edge detection. We demonstrate with experimental results on a database of 3,200 views of 10 objects that view aspects can be determined for sculpted objects easily and effectively.
Chitra Dorai, Anil K. Jain 0001
ICIP (3)2
1995 Page segmentation using texture discrimination masks
abstract
We propose a new texture-based page segmentation algorithm which automatically extracts the text, halftone, and line-drawing regions from input greyscale document images. This approach utilizes a neural network to train a set of masks which is optimal for discriminating the three main texture classes in the page segmentation problem: halftone, background, and text and line-drawing regions. The test and line-drawing regions are further discriminated based on connectivity analysis. We have applied the algorithm to successfully segment English and Chinese document images. We also demonstrate that the masks can perform language separation (English/Chinese) when appropriately trained.
Anil K. Jain 0001, Yu Zhong 0001
ICIP (3)1
1995 Detecting straight edges in millimeter-wave images
abstract
This paper presents two new methods for detecting edges in millimeter-wave radar images. The first method is based on a deformable template model of edge shapes and a random field model of the millimeter-wave imaging process. The second method is similar to the first, except that the imaging model component is replaced by one that is based on the magnitude and direction of the image gradient. Experimental results are shown to illustrate the advantages of using these methods over traditional edge detectors.
Sridhar Lakshmanan, Anil K. Jain 0001, Yu Zhong 0001
ICIP2
1995 Integration of Multiple Feature Groups and Multiple Views into a 3D Object Recognition System
Jianchang Mao, Patrick J. Flynn, Anil K. Jain 0001
Comput. Vis. Image Underst.3
1995 A Survey of Automated Visual Inspection
abstract
In this paper, we survey the automated visual inspection systems and techniques that have been reported in the literature from 1988 to 1993. Several earlier systems are also discussed. In the survey, a taxonomy of inspection systems based on their sensory input is presented. The general benefits and feasibility of automated visual inspection are also discussed. We present common approaches to visual inspection and also consider the specification and analysis of dimensional tolerances and their influence on the inspection task(s). One of the recent developments in automated visual inspection, namely the expanded role of computer-aided design (CAD) data in many systems, is examined in detail.
Timothy S. Newman, Anil K. Jain 0001
Comput. Vis. Image Underst.2
1995 Contour extraction of moving objects in complex outdoor scenes
Marie-Pierre Dubuisson, Anil K. Jain 0001
Int. J. Comput. Vis.2
1995 Integrating Vision Modules: Stereo, Shading, Grouping, and Line Labeling
abstract
It is generally agreed that individual visual cues are fallible and often ambiguous. This has generated a lot of interest in design of integrated vision systems which are expected to give a reliable performance in practical situations. The design of such systems is challenging since each vision module works under a different and possibly conflicting set of assumptions. We have proposed and implemented a multiresolution system which integrates perceptual organization (grouping), segmentation, stereo, shape from shading, and line labeling modules. We demonstrate the efficacy of our approach using images of several different realistic scenes. The output of the integrated system is shown to be insensitive to the constraints imposed by the individual modules. The numerical accuracy of the recovered depth is assessed in case of synthetically generated data. Finally, we have qualitatively evaluated our approach by reconstructing geons from the depth data obtained from the integrated system. These results indicate that integrated vision systems are likely to produce better reconstruction of the input scene than the individual modules.>
Sharath Pankanti, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Goal-Directed Evaluation of Binarization Methods
abstract
This paper presents a methodology for evaluation of low-level image analysis methods, using binarization (two-level thresholding) as an example. Binarization of scanned gray scale images is the first step in most document image analysis systems. Selection of an appropriate binarization method for an input image domain is a difficult problem. Typically, a human expert evaluates the binarized images according to his/her visual criteria. However, to conduct an objective evaluation, one needs to investigate how well the subsequent image analysis steps will perform on the binarized image. We call this approach goal-directed evaluation, and it can be used to evaluate other low-level image processing methods as well. Our evaluation of binarization methods is in the context of digit recognition, so we define the performance of the character recognition module as the objective measure. Eleven different locally adaptive binarization methods were evaluated, and Niblack's method gave the best performance.
Øivind Due Trier, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Knowledge-based clustering scheme for collection management and retrieval of library books
M. Narasimha Murty, Anil K. Jain 0001
Pattern Recognit.2
1995 Fusion of range and intensity images on a connection machine (CM-2)
Sateesha G. Nadabar, Anil K. Jain 0001
Pattern Recognit.2
1995 A system for 3D CAD-based inspection using range images
Timothy S. Newman, Anil K. Jain 0001
Pattern Recognit.2
1995 Adaptive flow orientation-based feature extraction in fingerprint images
Nalini K. Ratha, Shaoyun Chen, Anil K. Jain 0001
Pattern Recognit.3
1995 Locating text in complex color images
Yu Zhong 0001, Kalle Karu, Anil K. Jain 0001
Pattern Recognit.3
1995 A nonlinear projection method based on Kohonen's topology preserving maps
abstract
A nonlinear projection method is presented to visualize high-dimensional data as a 2D image. The proposed method is based on the topology preserving mapping algorithm of Kohonen. The topology preserving mapping algorithm is used to train a 2D network structure. Then the interpoint distances in the feature space between the units in the network are graphically displayed to show the underlying structure of the data. Furthermore, we present and discuss a new method to quantify how well a topology preserving mapping algorithm maps the high-dimensional input data onto the network structure. This is used to compare our projection method with a well-known method of Sammon (1969). Experiments indicate that the performance of the Kohonen projection method is comparable or better than Sammon's method for the purpose of classifying clustered data. Its time-complexity only depends on the resolution of the output image, and not on the size of the dataset. A disadvantage, however, is the large amount of CPU time required.
Martin A. Kraaijveld, Jianchang Mao, Anil K. Jain 0001
IEEE Trans. Neural Networks3
1995 Artificial neural networks for feature extraction and multivariate data projection
abstract
Classical feature extraction and data projection methods have been well studied in the pattern recognition and exploratory data analysis literature. We propose a number of networks and learning algorithms which provide new or alternative tools for feature extraction and data projection. These networks include a network (SAMANN) for J.W. Sammon's (1969) nonlinear projection, a linear discriminant analysis (LDA) network, a nonlinear discriminant analysis (NDA) network, and a network for nonlinear projection (NP-SOM) based on Kohonen's self-organizing map. A common attribute of these networks is that they all employ adaptive learning algorithms which makes them suitable in some environments where the distribution of patterns in feature space changes with respect to time. The availability of these networks also facilitates hardware implementation of well-known classical feature extraction and projection approaches. Moreover, the SAMANN network offers the generalization ability of projecting new data, which is not present in the original Sammon's projection algorithm; the NDA method and NP-SOM network provide new powerful approaches for visualizing high dimensional data. We evaluate five representative neural networks for feature extraction and data projection based on a visual judgement of the two-dimensional projection maps and three quantitative criteria on eight data sets with various properties.
Jianchang Mao, Anil K. Jain 0001
IEEE Trans. Neural Networks2
1994 2D matching of 3D moving objects in color outdoor scenes
abstract
This paper describes an object matching system which is able to extract objects of interest from outdoor scenes and match them. Our application (in the domain of IVHS) involves measuring the average travel time in a road network. The extraction of the object of interest is performed by fusing multiple cues including motion, color, edges, and model information. Two objects extracted from images captured by two independent cameras at different times are then matched to evaluate their similarity. Color indexing based on histogram matching is used to avoid matching all possible pairs of objects. To resolve ambiguities, further matching is done by measuring the Hausdorff distance between two sets of edge points. The object matching system was given 2 sets of 40 vehicles. It was able to identify 23 of the 30 correct matches and all the false matches were rejected. Color indexing reduced the number of candidates for a match from 40 to 2. This matching accuracy is adequate to obtain a reliable estimate of the average travel time.>
Marie-Pierre Dubuisson, Anil K. Jain 0001
CVPR2
1994 On integration of vision modules
abstract
Individual cues from visual modules are fallible and often ambiguous. As a result, only integrated vision systems can be expected to give a reliable performance in practice. The design of such systems is challenging since each vision module works under different and possibly conflicting sets of assumptions. We have proposed and implemented a multiresolution system which integrates perceptual grouping, segmentation, stereo, shape from shading, and line labelling modules. The output of the integrated system is shown to be relatively insensitive to the constraints imposed by the individual modules.>
Sharath Pankanti, Anil K. Jain 0001, Mihran Tüceryan
CVPR2
1994 Fusing Color and Edge Information for Object Matching
abstract
This paper illustrates the advantages of using multiple cues for object matching. Given two sets of people entering and leaving a room, the goal is to identify the matching pairs assuming that the viewing aspects of the people in the two scenes are similar. Color information or 2D shape information alone is not enough to find all the matching pairs, but all the matching pairs are correctly identified when both the features are combined.>
Marie-Pierre Jolly, Anil K. Jain 0001
ICIP (3)2
1994 Multi-resolution Image Representation using Markov Random Fields
abstract
This paper presents a new method for representing the spatial information present in digital grey-tone images. The method is based on using multi-resolution decompositions (MRDs) and Markov random fields (MRFs) concurrently. A given image is represented by a MRD of it, along with an optimally estimated set of Gaussian MRF (GMRF) parameters. Since the GMRF parameters are very small in number, this addition to the usual MRD results in only a small increase in the number of bits in the representation. It is shown, however, that such a minor addition helps when reconstructing the (given) original image from its MRD. Experimental results are presented to illustrate the usefulness of this new method.>
Sridhar Lakshmanan, Anil K. Jain 0001, Yu Zhong 0001
ICIP (1)2
1994 Automatic filter design for texture discrimination
abstract
Multichannel filtering has been shown by many researchers to provide good features for texture segmentation and classification. In this paper the authors exploit neural networks to construct optimal filters and to combine the outputs of these filters for the classification of known textures. The authors use the neural network training together with node pruning, so that both the classification error and the number of filters or, equivalently, the number of features, are minimized. The performance of the neural network classifier is demonstrated an several experiments involving classification of natural textures. The authors study the effects of using different sized filters with different network configurations. The authors show that the number of filters, and, therefore, the processing time, can be greatly reduced while preserving the classification accuracy, using the proposed scheme compared to using a general set of filters (e.g., Gabor filters).
Anil K. Jain 0001, Kalle Karu
ICPR (1)1
1994 Optimal registration of multiple range views
abstract
Errors in registration of multiple views of an object based on estimated transformations between views can affect surface classification. The authors derive a minimum variance estimator (MVE) for computing the transformation parameters accurately from range data of two different views of a 3D object. The results of the authors' experiments show that the solution obtained using MVE is significantly more reliable than the estimate obtained with an unweighted distance criterion for registration.
Chitra Dorai, John Weng, Anil K. Jain 0001
ICPR (1)3
1994 A modified Hausdorff distance for object matching
abstract
The purpose of object matching is to decide the similarity between two objects. This paper introduces 24 possible distance measures based on the Hausdorff distance between two point sets. These measures can be used to match two sets of edge points extracted from any two objects. Based on experiments on synthetic images containing various levels of noise, the authors determined that one of these distance measures, called the modified Hausdorff distance (MHD) has the best performance for object matching. The advantages of MHD ever other distances are also demonstrated on several edge snaps of objects extracted from real images.
Marie-Pierre Dubuisson, Anil K. Jain 0001
ICPR (1)2
1994 Boundary detection using multiscale Markov random fields
abstract
The basic difficulty encountered in filtering-based multiscale boundary detection methods is the elimination of noise and insignificant edges while preserving positional accuracy at the image discontinuities. In this paper, a nonlinear multiscale boundary detection method which prevents the conflict between the detection and localization goals is introduced. The method uses multiscale representations of coupled Markov random fields and applies a stochastic regularization scheme based on the Bayesian approach. This allows the integration of boundary information extracted at multiple scales simultaneously resulting in robust integration of the information at a variety of spatial scales. The scheme is applicable to intensity images as well as to range images and eliminates the dependency on edge operator size which is the main difficulty in filtering-based multiscale techniques.
Bilge Günsel, Erdal Panayirci, Anil K. Jain 0001
ICPR (2)3