Patrick J. Flynn

dblp:45/2338 · DBLP profile ↗
← Back
122ranked-venue papers
12as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 11 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 59 · 5 first-author · 10 since 2021Security and privacy · 24 · 5 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1
YearPublicationVenuePosition
2026 PaW-ViT: A Patch-based Warping Vision Transformer for Robust Ear Verification
Deeksha Arun, Kevin W. Bowyer, Patrick J. Flynn
FG3
2026 BARS: A Blockchain-Anchored System for rPPG-Based Biometric Verification
Lu Niu, Patrick J. Flynn
ICBC2
2023 Non-Contrastive Unsupervised Learning of Physiological Signals from Video
abstract
Subtle periodic signals such as blood volume pulse and respiration can be extracted from RGB video, enabling non-contact health monitoring at low cost. Advancements in remote pulse estimation - or remote photoplethysmography (rPPG) - are currently driven by deep learning solutions. However, modern approaches are trained and evaluated on benchmark datasets with ground truth from contact-P PG sensors. We present the first non-contrastive unsuper-vised learning framework for signal regression to mitigate the need for labelled video data. With minimal assumptions of periodicity and finite bandwidth, our approach discovers the blood volume pulse directly from unlabelled videos. We find that encouraging sparse power spectra within normal physiological bandlimits and variance over batches of power spectra is sufficient for learning visual features of periodic signals. We perform the first experiments utilizing unlabelled video data not specifically created for rPPG to train robust pulse rate estimators. Given the limited inductive biases and impressive empirical results, the approach is theoretically capable of discovering other periodic signals from video, enabling multiple physiological measurements without the need for ground truth signals.
Jeremy Speth, Nathan Vance, Patrick J. Flynn, Adam Czajka
CVPR3
2023 Analyzing the Impact of Shape & Context on the Face Recognition Performance of Deep Networks
abstract
In this article, we analyze how changing the underlying 3D shape of the base identity in face images can distort their overall appearance, especially from the perspective of deep face recognition. As done in popular training data augmentation schemes, we graphically render real and synthetic face images with randomly chosen or best-fitting 3D face models to generate novel views of the base identity. We compare deep features generated from these images to assess the perturbation these renderings introduce into the original identity. We perform this analysis at various degrees of facial yaw with the base identities varying in gender and ethnicity. Additionally, we investigate if adding some form of context and background pixels in these rendered images, when used as training data, further improves the downstream performance of a face recognition model. Our experiments demonstrate the significance of facial shape in accurate face matching and underpin the importance of contextual data for network training.
Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn
FG4
2023 Iris Liveness Detection Competition (LivDet-Iris) - The 2023 Edition
abstract
This paper describes the results of the 2023 edition of the “LivDet” series of iris presentation attack detection (PAD) competitions. New elements in this fifth competition include (1) GAN-generated iris images as a category of presentation attack instruments (PAI), and (2) an evaluation of human accuracy at detecting PAI as a reference benchmark. Clarkson University and the University of Notre Dame contributed image datasets for the competition, composed of samples representing seven different PAI categories, as well as baseline PAD algorithms. Fraunhofer IGD, Beijing University of Civil Engineering and Architecture, and Hochschule Darmstadt contributed results for a total of eight PAD algorithms to the competition. Accuracy results are analyzed by different PAI types, and compared to human accuracy. Overall, the Fraunhofer IGD algorithm, using an attention-based pixel-wise binary supervision network, showed the best-weighted accuracy results (average classification error rate of 37.31%), while the Beijing University of Civil Engineering and Architecture’s algorithm won when equal weights for each PAI were given (average classification rate of 22.15%). These results suggest that iris PAD is still a challenging problem.
Patrick Tinsley, Sandip Purnapatra, Mahsa Mitcheff, Aidan Boyd, Colton R. Crum, Kevin W. Bowyer, Patrick J. Flynn, Stephanie Schuckers, Adam Czajka, Meiling Fang, Naser Damer, Caiyong Wang, Xianyun Sun, Zhaohua Chang, Guangzhe Zhao, Juan E. Tapia, Christoph Busch 0001, Carlos M. Aravena, Daniel Schulz
IJCB7
2022 Haven't I Seen You Before? Assessing Identity Leakage in Synthetic Irises
abstract
Generative Adversarial Networks (GANs) have proven to be a preferred method of synthesizing fake images of ob-jects, such as faces, animals, and automobiles. It is not surprising these models can also generate ISO-compliant, yet synthetic iris images, which can be used to augment training data for iris matchers and liveness detectors. In this work, we trained one of the most recent GAN mod-els (StyleGAN3 [15]) to generate fake iris images with two primary goals: (i) to understand the GAN's ability to produce “never-before-seen” irises, and (ii) to investigate the phenomenon of identity leakage as a function of the GAN's training time. Previous work has shown that personal biometric data can inadvertently flow from training data into synthetic samples, raising a privacy concern for subjects who accidentally appear in the training dataset. This paper presents analysis for three different iris matchers at varying points in the GAN training process to diagnose where and when authentic training samples are in jeopardy of leaking through the generative process. Our results show that while most synthetic samples do not show signs of identity leak-age, a handful of generated samples match authentic (training) samples nearly perfectly, with consensus across all matchers. In order to prioritize privacy, security, and trust in the machine learning model development process, the re-search community must strike a delicate balance between the benefits of using synthetic data and the corresponding threats against privacy from potential identity leakage.
Patrick Tinsley, Adam Czajka, Patrick J. Flynn
IJCB3
2022 On Low-Resolution Face Re-identification with High-Resolution-Mapping
Loreto Prieto, Sebastian A. Pulgar, Patrick J. Flynn, Domingo Mery
PSIVT3
2022 Digital and Physical-World Attacks on Remote Pulse Detection
abstract
Remote photoplethysmography (rPPG) is a technique for estimating blood volume changes from reflected light without the need for a contact sensor. We present the first examples of presentation attacks in the digital and physical domains on rPPG from face video. Digital attacks are easily performed by adding imperceptible periodic noise to the input videos. Physical attacks are performed with illumination from visible spectrum LEDs placed in close proximity to the face, while still being difficult to perceive with the human eye. We also show that our attacks extend beyond medical applications, since the method can effectively generate a strong periodic pulse on 3D-printed face masks, which presents difficulties for pulse-based face presentation attack detection (PAD). The paper concludes with ideas for using this work to improve robustness of rPPG methods and pulse-based face PAD.
Jeremy Speth, Nathan Vance, Patrick J. Flynn, Kevin W. Bowyer, Adam Czajka
WACV3
2021 Deception Detection and Remote Physiological Monitoring: A Dataset and Baseline Experimental Results
abstract
We present the Deception Detection and Physiological Monitoring (DDPM) dataset and initial baseline results on this dataset. Our application context is an interview scenario in which the interviewee attempts to deceive the interviewer on selected responses. The interviewee is recorded in RGB, near-infrared, and long-wave infrared, along with cardiac pulse, blood oxygenation, and audio. After collection, data were annotated for interviewer/interviewee, curated, ground-truthed, and organized into train / test parts for a set of canonical deception detection experiments. Baseline experiments found random accuracy for micro-expressions as an indicator of deception, but that saccades can give a statistically significant response. We also estimated subject heart rates from face videos (remotely) with a mean absolute error as low as 3.16 bpm. The database contains almost 13 hours of recordings of 70 subjects, and over 8 million visible-light, near-infrared, and thermal video frames, along with appropriate meta, audio and pulse oximeter data. To our knowledge, this is the only collection offering recordings of five modalities in an interview scenario that can be used in both deception detection and remote photoplethysmography research.
Jeremy Speth, Nathan Vance, Adam Czajka, Kevin W. Bowyer, Diane Wright, Patrick J. Flynn
IJCB6
2021 This Face Does Not Exist... But It Might Be Yours! Identity Leakage in Generative Models
abstract
Generative adversarial networks (GANs) are able to generate high resolution photo-realistic images of objects that "do not exist." These synthetic images are rather difficult to detect as fake. However, the manner in which these generative models are trained hints at a potential for information leakage from the supplied training data, especially in the context of synthetic faces. This paper presents experiments suggesting that identity information in face images can flow from the training corpus into synthetic samples without any adversarial actions when building or using the existing model. This raises privacy-related questions, but also stimulates discussions of (a) the face manifold's characteristics in the feature space and (b) how to create generative models that do not inadvertently reveal identity information of real subjects whose images were used for training. We used five different face matchers (face_recognition, FaceNet, ArcFace, SphereFace and Neurotechnology MegaMatcher) and the StyleGAN2 synthesis model, and show that this identity leakage does exist for some, but not all methods. So, can we say that these synthetically generated faces truly do not exist? Databases of real and synthetically generated faces are made available with this paper to allow full replicability of the results discussed in this work.
Patrick Tinsley, Adam Czajka, Patrick J. Flynn
WACV3
2021 Unifying frame rate and temporal dilations for improved remote pulse detection
Jeremy Speth, Nathan Vance, Patrick J. Flynn, Kevin W. Bowyer, Adam Czajka
Comput. Vis. Image Underst.3
2021 Transformation-Aware Embeddings for Image Provenance
abstract
A dramatic rise in the flow of manipulated image content on the Internet has led to a prompt response from the media forensics research community. New mitigation efforts leverage cutting-edge data-driven strategies and increasingly incorporate usage of techniques from computer vision and machine learning to detect and profile the space of image manipulations. This paper addresses Image Provenance Analysis, which aims at discovering relationships among different manipulated image versions that share content. One important task in provenance analysis, like most visual understanding problems, is establishing a visual description and dissimilarity computation method that connects images that share full or partial content. But the existing handcrafted or learned descriptors - generally appropriate for tasks such as object recognition - may not sufficiently encode the subtle differences between near-duplicate image variants, which significantly characterize the provenance of any image. This paper introduces a novel data-driven learning-based approach that provides the context for ordering images that have been generated from a single image source through various transformations. Our approach learns transformation-aware embeddings using weak supervision via composited transformations and a rank-based Edit Sequence Loss. To establish the effectiveness of the proposed approach, comparisons are made with state-of-the-art handcrafted and deep-learning-based descriptors, as well as image matching approaches. Further experimentation validates the proposed approach in the context of image provenance analysis and improves upon existing approaches.
Aparna Bharati, Daniel Moreira, Patrick J. Flynn, Anderson Rocha 0001, Kevin W. Bowyer, Walter J. Scheirer
IEEE Trans. Inf. Forensics Secur.3
2021 Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image Retrieval
abstract
Images from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates - and to do so with enough efficiency at large scale. We propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit.
Joel Brogan, Aparna Bharati, Daniel Moreira, Anderson Rocha 0001, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer
IEEE Trans. Image Process.6
2020 On Hallucinating Context and Background Pixels from a Face Mask using Multi-scale GANs
abstract
We propose a multi-scale GAN model to hallucinate realistic context (forehead, hair, neck, clothes) and background pixels automatically from a single input face mask, without any user supervision. Instead of swapping a face on to an existing picture, our model directly generates realistic context and background pixels based on the features of the provided face mask. Unlike facial inpainting algorithms, it can generate realistic hallucinations even for a large number of missing pixels. Our model is composed of a cascaded network of GAN blocks, each tasked with hallucination of missing pixels at a particular resolution while guiding the synthesis process of the next GAN block. The hallucinated full face image is made photo-realistic by using a combination of reconstruction, perceptual, adversarial and identity preserving losses at each block of the network. With a set of extensive experiments, we demonstrate the effectiveness of our model in hallucinating context and background pixels from face masks varying in facial pose, expression and lighting, collected from multiple datasets subject disjoint with our training data. We also compare our method with popular face inpainting and face swapping models in terms of visual quality, realism and identity preservation. Additionally, we analyze our cascaded pipeline and compare it with the progressive growing of GANs, and explore its usage as a data augmentation module for training CNNs.
Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn
WACV4
2019 Fast Face Image Synthesis With Minimal Training
abstract
We propose an algorithm to generate realistic face images of both real and synthetic identities (people who do not exist) with different facial yaw, shape and resolution. The synthesized images can be used to augment datasets to train CNNs or as massive distractor sets for biometric verification experiments without any privacy concerns. Additionally, law enforcement can make use of this technique to train forensic experts to recognize faces. Our method samples face components from a pool of multiple face images of real identities to generate the synthetic texture. Then, a real 3D head model compatible to the generated texture is used to render it under different facial yaw transformations. We perform multiple quantitative experiments to assess the effectiveness of our synthesis procedure in CNN training and its potential use to generate distractor face images. Additionally, we compare our method with popular GAN models in terms of visual quality and execution time.
Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn
WACV4
2019 Beyond Pixels: Image Provenance Analysis Leveraging Metadata
abstract
Creative works, whether paintings or memes, follow unique journeys that result in their final form. Understanding these journeys, a process known as "provenance analysis," provides rich insights into the use, motivation, and authenticity underlying any given work. The application of this type of study to the expanse of unregulated content on the Internet is what we consider in this paper. Provenance analysis provides a snapshot of the chronology and validity of content as it is uploaded, re-uploaded, and modified over time. Although still in its infancy, automated provenance analysis for online multimedia is already being applied to different types of content. Most current works seek to build provenance graphs based on the shared content between images or videos. This can be a computationally expensive task, especially when considering the vast influx of content that the Internet sees every day. Utilizing non-content-based information, such as timestamps, geotags, and camera IDs can help provide important insights into the path a particular image or video has traveled during its time on the Internet without large computational overhead. This paper tests the scope and applicability of metadata-based inferences for provenance graph construction in two different scenarios: digital image forensics and cultural analytics.
Aparna Bharati, Daniel Moreira, Joel Brogan, Patricia Hale, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
WACV6
2019 "Keep Me In, Coach!": A Computer Vision Perspective on Assessing ACL Injury Risk in Female Athletes
abstract
We present and share a foundational dataset of multi-angle video recordings of scripted athletic movements to enable the development of computer vision research applications that evaluate and identify lower-body injury risk. The focus of the dataset is female athletes, who are at a substantially increased risk of anterior cruciate ligament (ACL) injury and are therefore a top priority for sports science. In our study, varsity and club sport athletes perform two assessment movements (the countermovement jump and the drop jump). These jump tasks are used ubiquitously in sports medicine research to characterize athleticism and to identify risk factors that indicate ACL injury propensity. The novelty of the dataset centers on (i) the type of movement data (purposeful, evaluative movements that need to be tracked with a high degree of precision), (ii) our generalized collection method that can be replicated with ease by non-experts, and (iii) the amount of data collected (we collected data from 55 division one (D1) female athletes performing 3-5 iterations of each jumps, for a total of 480 jumps). Data from each camera was manually aligned and a fully automated pipeline was built to extract knee information from athletes. Ideally, any athlete or researcher will be able to easily replicate our setup and assemble a compatible and complementary dataset to propel the development and assessment of injury propensity models.
Nathaniel Blanchard, Kyle Skinner, Aden Kemp, Walter J. Scheirer, Patrick J. Flynn
WACV5
2019 Domain-Specific Human-Inspired Binarized Statistical Image Features for Iris Recognition
abstract
Binarized statistical image features (BSIF) have been successfully used for texture analysis in many computer vision tasks, including iris recognition and biometric presentation attack detection. One important point is that all applications of BSIF in iris recognition have used the original BSIF filters, which were trained on image patches extracted from natural images. This paper tests the question of whether domain-specific BSIF can give better performance than the default BSIF. The second important point is in the selection of image patches to use in training for BSIF. Can image patches derived from eye-tracking experiments, in which humans perform an iris recognition task, give better performance than random patches? Our results say that (1) domain-specific BSIF features can out-perform the default BSIF features, and (2) selecting image patches in a task-specific manner guided by human performance can out-perform selecting random patches. These results are important because BSIF is often regarded as a generic texture tool that does not need any domain adaptation, and human-task-guided selection of patches for training has never (to our knowledge) been done. This paper follows the reproducible research requirements, and the new iris-domain-specific BSIF filters, the patches used in filter training, the database used in testing and the source codes of the designed iris recognition method are made available along with this paper to facilitate applications of this concept.
Adam Czajka, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn
WACV4
2019 Performance of Humans in Iris Recognition: The Impact of Iris Condition and Annotation-Driven Verification
abstract
This paper advances the state of the art in human examination of iris images by (1) assessing the impact of different iris conditions in identity verification, and (2) introducing an annotation step that improves the accuracy of people's decisions. In a first experimental session, 114 subjects were asked to decide if pairs of iris images depict the same eye (genuine pairs) or two distinct eyes (impostor pairs). The image pairs sampled six conditions: (1) easy for algorithms to classify, (2) difficult for algorithms to classify, (3) large difference in pupil dilation, (4) disease-affected eyes, (5) identical twins, and (6) post-mortem samples. In a second session, 85 of the 114 subjects were asked to annotate matching and non-matching regions that supported their decisions. Subjects were allowed to change their initial classification as a result of the annotation process. Results suggest that: (a) people improve their identity verification accuracy when asked to annotate matching and non-matching regions between the pair of images, (b) images depicting the same eye with large difference in pupil dilation were the most challenging to subjects, but benefited well from the annotation-driven classification, (c) humans performed better than iris recognition algorithms when verifying genuine pairs of post-mortem and disease-affected eyes (i.e., samples showing deformations that go beyond the distortions of a healthy iris due to pupil dilation), and (d) annotation does not improve accuracy of analyzing images from identical twins, which remain confusing for people.
Daniel Moreira, Mateusz Trokielewicz, Adam Czajka, Kevin W. Bowyer, Patrick J. Flynn
WACV5
2019 On Low-Resolution Face Recognition in the Wild: Comparisons and New Techniques
abstract
Although face recognition systems have achieved impressive performance in recent years, the low-resolution face recognition task remains challenging, especially when the low-resolution faces are captured under non-ideal conditions, which is widely prevalent in surveillance-based applications. Faces captured in such conditions are often contaminated by blur, non-uniform lighting, and non-frontal face pose. In this paper, we analyze the face recognition techniques using data captured under low-quality conditions in the wild. We provide a comprehensive analysis of the experimental results for two of the most important applications in real surveillance applications, and demonstrate practical approaches to handle both cases that show promising performance. The following three contributions are made: (i) we conduct experiments to evaluate super-resolution methods for low-resolution face recognition; (ii) we study face re-identification on various public face datasets, including real surveillance and low-resolution subsets of large-scale datasets, presenting a baseline result for several deep learning-based approaches, and improve them by introducing a generative adversarial network pre-training approach and fully convolutional architecture; and (iii) we explore the low-resolution face identification by employing a state-of-the-art supervised discriminative learning approach. The evaluations are conducted on challenging portions of the SCface and UCCSface datasets.
Loreto Prieto, Domingo Mery, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.4
2019 Assessing the Impact of Corneal Refraction and Iris Tissue Non-Planarity on Iris Recognition
abstract
Assumptions regarding eye morphology are implicit in iris recognition algorithms. The cornea is assumed to have little to no effect on the view of the iris texture when the eye gaze is non-frontal, the iris is assumed flat, and the eye is assumed to be imaged via an orthographic projection. If these assumptions hold, affine transformations may be used to rectify non-frontally posed images to a frontal view and the rubber sheet model may be accurately used to normalize the iris annulus to a rectangular image. This paper examines how iris recognition performance degrades when the first two assumptions are violated. Using a computer renderable eye model, a data set is created varying the presence of the cornea and a parameterized non-planarity of the iris shape across a large range of eye gaze angles. Matching scores are created using a commercial matcher. When comparing the relative impact of each assumption violation, it is observed that iris non-planarity presents a more significant problem than corneal refractive distortion with regard to iris recognition accuracy in non-frontal images.
Joseph Thompson, Patrick J. Flynn, Chris Boehnen, Hector J. Santos-Villalobos
IEEE Trans. Inf. Forensics Secur.2
2018 To Frontalize or Not to Frontalize: Do We Really Need Elaborate Pre-processing to Improve Face Recognition?
abstract
Face recognition performance has improved remarkably in the last decade. Much of this success can be attributed to the development of deep learning techniques such as convolutional neural networks (CNNs). While CNNs have pushed the state-of-the-art forward, their training process requires a large amount of clean and correctly labelled training data. If a CNN is intended to tolerate facial pose, then we face an important question: should this training data be diverse in its pose distribution, or should face images be normalized to a single pose in a pre-processing step? To address this question, we evaluate a number of facial landmarking algorithms and a popular frontalization method to understand their effect on facial recognition performance. Additionally, we introduce a new, automatic, single-image frontalization scheme that exceeds the performance of the reference frontalization algorithm for video-to-video face matching on the Point and Shoot Challenge (PaSC) dataset. Additionally, we investigate failure modes of each frontalization method on different facial yaw using the CMU Multi-PIE dataset. We assert that the subsequent recognition and verification performance serves to quantify the effectiveness of each pose correction scheme.
Sandipan Banerjee, Joel Brogan, Janez Krizaj, Aparna Bharati, Brandon RichardWebster, Vitomir Struc, Patrick J. Flynn, Walter J. Scheirer
WACV7
2018 Image Provenance Analysis at Scale
abstract
Prior art has shown it is possible to estimate, through image processing and computer vision techniques, the types and parameters of transformations that have been applied to the content of individual images to obtain new images. Given a large corpus of images and a query image, an interesting further step is to retrieve the set of original images whose content is present in the query image, as well as the detailed sequences of transformations that yield the query image given the original images. This is a problem that recently has received the name of image provenance analysis. In these times of public media manipulation (e.g., fake news and meme sharing), obtaining the history of image transformations is relevant for fact checking and authorship verification, among many other applications. This article presents an end-to-end processing pipeline for image provenance analysis, which works at real-world scale. It employs a cutting-edge image filtering solution that is custom-tailored for the problem at hand, as well as novel techniques for obtaining the provenance graph that expresses how the images, as nodes, are ancestrally connected. A comprehensive set of experiments for each stage of the pipeline is provided, comparing the proposed solution with state-of-the-art results, employing previously published datasets. In addition, this work introduces a new dataset of real-world provenance cases from the social media site Reddit, along with baseline results.
Aparna Bharati, Joel Brogan, Allan Pinto, Michael Parowski, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
IEEE Trans. Image Process.6
2017 SREFI: Synthesis of realistic example face images
abstract
In this paper, we propose a novel face synthesis approach that can generate an arbitrarily large number of synthetic images of both real and synthetic identities. Thus a face image dataset can be expanded in terms of the number of identities represented and the number of images per identity using this approach, without the identity-labeling and privacy complications that come from downloading images from the web. To measure the visual fidelity and uniqueness of the synthetic face images and identities, we conducted face matching experiments with both human participants and a CNN pre-trained on a dataset of 2.6M real face images. To evaluate the stability of these synthetic faces, we trained a CNN model with an augmented dataset containing close to 200,000 synthetic faces. We used a snapshot of this trained CNN to recognize extremely challenging frontal (real) face images. Experiments showed training with the augmented faces boosted the face recognition performance of the CNN.
Sandipan Banerjee, John S. Bernhard, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn
IJCB5
2017 Learning face similarity for re-identification from real surveillance video: A deep metric solution
abstract
Person re-identification (ReID) is the task of automatically matching persons across surveillance cameras with location or time differences. Nearly all proposed ReID approaches exploit body features. Even if successfully captured in the scene, faces are often assumed to be unhelpful to the ReID process[3]. As cameras and surveillance systems improve, ‘Facial ReID’ approaches deserve attention. The following contributions are made in this work: 1) We describe a high-quality dataset for person re-identification featuring faces. This dataset was collected from a real surveillance network in a municipal rapid transit system, and includes the same people appearing in multiple sites at multiple times wearing different attire. 2) We employ new DNN architectures and patch matching techniques to handle face misalignment in quality regimes where landmarking fails. We further boost the performance by adopting the fully convolutional structure and spatial pyramid pooling (SPP).
Loreto Prieto, Patrick J. Flynn, Domingo Mery
IJCB3
2017 U-Phylogeny: Undirected provenance graph construction in the wild
abstract
Deriving relationships between images and tracing back their history of modifications are at the core of Multimedia Phylogeny solutions, which aim to combat misinformation through doctored visual media. Nonetheless, most recent image phylogeny solutions cannot properly address cases of forged composite images with multiple donors, an area known as multiple parenting phylogeny (MPP). This paper presents a preliminary undirected graph construction solution for MPP, without any strict assumptions. The algorithm is underpinned by robust image representative keypoints and different geometric consistency checks among matching regions in both images to provide regions of interest for direct comparison. The paper introduces a novel technique to geometrically filter the most promising matches as well as to aid in the shared region localization task. The strength of the approach is corroborated by experiments with real-world cases, with and without image distractors (unrelated cases).
Aparna Bharati, Daniel Moreira, Allan Pinto, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001
ICIP6
2017 Spotting the difference: Context retrieval and analysis for improved forgery detection and localization
abstract
As image tampering becomes ever more sophisticated and commonplace, the need for image forensics algorithms that can accurately and quickly detect forgeries grows. In this paper, we revisit the ideas of image querying and retrieval to provide clues to better localize forgeries. We propose a method to perform large-scale image forensics on the order of one million images using the help of an image search algorithm and database to gather contextual clues as to where tampering may have taken place. In this vein, we introduce five new strongly invariant image comparison methods and test their effectiveness under heavy noise, rotation, and color space changes. Lastly, we show the effectiveness of these methods compared to passive image forensics using Nimble [1], a new, state-of-the-art dataset from the National Institute of Standards and Technology (NIST).
Joel Brogan, Paolo Bestagini, Aparna Bharati, Allan Pinto, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
ICIP7
2017 Provenance filtering for multimedia phylogeny
abstract
Departing from traditional digital forensics modeling, which seeks to analyze single objects in isolation, multimedia phylogeny analyzes the evolutionary processes that influence digital objects and collections over time. One of its integral pieces is provenance filtering, which consists of searching a potentially large pool of objects for the most related ones with respect to a given query, in terms of possible ancestors (donors or contributors) and descendants. In this paper, we propose a two-tiered provenance filtering approach to find all the potential images that might have contributed to the creation process of a given query q. In our solution, the first (coarse) tier aims to find the most likely “host” images - the major donor or background - contributing to a composite/doctored image. The search is then refined in the second tier, in which we search for more specific (potentially small) parts of the query that might have been extracted from other images and spliced into the query image. Experimental results with a dataset containing more than a million images show that the two-tiered solution underpinned by the context of the query is highly useful for solving this difficult task.
Allan Pinto, Daniel Moreira, Aparna Bharati, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001
ICIP6
2017 Lessons from collecting a million biometric samples
P. Jonathon Phillips, Patrick J. Flynn, Kevin W. Bowyer
Image Vis. Comput.2
2017 Special issue on Best of Biometrics 2015
Massimo Tistarelli, J. Ross Beveridge, Patrick J. Flynn, Michele Nappi
Image Vis. Comput.3
2017 Crowd Scene Understanding from Video: A Survey
abstract
Crowd video analysis has applications in crowd management, public space design, and visual surveillance. Example tasks potentially aided by automated analysis include anomaly detection (such as a person walking against the grain of traffic or rapid assembly/dispersion of groups of people), population and density measurements, and interactions between groups of people. This survey explores crowd analysis as it relates to two primary research areas: crowd statistics and behavior understanding. First, we survey methods for counting individuals and approximating the density of the crowd. Second, we showcase research efforts on behavior understanding as related to crowds. These works focus on identifying groups, interactions within small groups, and abnormal activity detection such as riots and bottlenecks in large crowds. Works presented in this section also focus on tracking groups of individuals, either as a single entity or a subset of individuals within the frame of reference. Finally, a summary of datasets available for crowd activity video research is provided.
Jason M. Grant, Patrick J. Flynn
ACM Trans. Multim. Comput. Commun. Appl.2
2016 Iris Recognition Based on Human-Interpretable Features
abstract
The iris is a stable biometric trait that has been widely used for human recognition in various applications. However, deployment of iris recognition in forensic applications has not been reported. A primary reason is the lack of human-friendly techniques for iris comparison. To further promote the use of iris recognition in forensics, the similarity between irises should be made visualizable and interpretable. Recently, a human-in-the-loop iris recognition system was developed, based on detecting and matching iris crypts. Building on this framework, we propose a new approach for detecting and matching iris crypts automatically. Our detection method is able to capture iris crypts of various sizes. Our matching scheme is designed to handle potential topological changes in the detection of the same crypt in different images. Our approach outperforms the known visible-feature-based iris recognition method on three different data sets. In particular, our approach achieves over 22% higher rank one hit rate in identification, and over 51% lower equal error rate in verification. In addition, the benefit of our approach on multi-enrollment is experimentally demonstrated.
Jianxu Chen 0001, Danny Ziyi Chen, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.4
2015 Strong, Neutral, or Weak: Exploring the Impostor Score Distribution
abstract
The strong, neutral, or weak (SNoW) face impostor pairs problem is intended to explore the causes and impact of impostor face pairs that are inherently strong (easily recognized as nonmatches) or weak (possible false matches). The SNoW technique develops three partitions within the impostor score distribution of a given data set. Results provide evidence that varying degrees of impostor scores impact the overall performance of a face recognition system. This paper extends our earlier work to incorporate improvements regarding outlier detection for partitioning, explores the SNoW concept for the additional modalities of fingerprint and iris, and presents methods for how to begin to reveal the causes of weak impostor pairs. We also show a clear operational difference between strong and weak comparisons as well as identify partition stability across multiple algorithms.
Amanda Sgroi, Patrick J. Flynn, Kevin W. Bowyer, P. Jonathon Phillips
IEEE Trans. Inf. Forensics Secur.2
2014 The IJCB 2014 PaSC video face and person recognition competition
abstract
The Point-and-Shoot Face Recognition Challenge (PaSC) is a performance evaluation challenge including 1401 videos of 265 people acquired with handheld cameras and depicting people engaged in activities with non-frontal head pose. This report summarizes the results from a competition using this challenge problem. In the Video-to-video Experiment a person in a query video is recognized by comparing the query video to a set of target videos. Both target and query videos are drawn from the same pool of 1401 videos. In the Still-to-video Experiment the person in a query video is to be recognized by comparing the query video to a larger target set consisting of still images. Algorithm performance is characterized by verification rate at a false accept rate of 0.01 and associated receiver operating characteristic (ROC) curves. Participants were provided eye coordinates for video frames. Results were submitted by 4 institutions: (i) Advanced Digital Science Center, Singapore; (ii) CPqD, Brasil; (iii) Stevens Institute of Technology, USA; and (iv) University of Ljubljana, Slovenia. Most competitors demonstrated video face recognition performance superior to the baseline provided with PaSC. The results represent the best performance to date on the handheld video portion of the PaSC.
J. Ross Beveridge, Hao Zhang 0013, Patrick J. Flynn, Yooyoung Lee, Venice Erin Liong, Jiwen Lu, Marcus A. Angeloni, Tiago de Freitas Pereira, Gang Hua 0001, Vitomir Struc, Janez Krizaj, P. Jonathon Phillips
IJCB3
2014 An optimal strategy for dilation based iris image enrollment
abstract
The progression of research into understanding and mitigating the effects of pupil dilation on iris biometrics is at a point where a formalization of the problem is necessary to tie together several research directions and results. Past research has shown that differences in dilation in a (probe, gallery) pair lead to an increase in false non-match rates. Additionally, analysis continues to show that there is at least an approximate linear relationship between increase in dilation difference and degradation in match scores. Lastly, dilation-aware based enrollment techniques have shown to be a promising approach to addressing matching errors due to pupil dilation difference. This paper establishes a framework based on an assumed linear relationship between match scores and dilation difference and shows that the optimal image to enroll based on pupil dilation is the image which has a dilation value near the mean or median depending on the measure of dilation difference.
Estefan Ortiz, Kevin W. Bowyer, Patrick J. Flynn
IJCB3
2014 The effectiveness of face detection algorithms in unconstrained crowd scenes
abstract
The 2013 Boston Marathon bombing represents a case where automatic facial biometrics tools could have proven invaluable to law enforcement officials, yet the lack of robustness of current tools in unstructured environments limited their utility. In this work, we focus on complications that confound face detection algorithms. We first present a simple multi-pose generalization of the Viola-Jones algorithm. Our results on the Face Detection Data set and Benchmark (FDDB) show that it makes a significant improvement over the state of the art for published algorithms. Conversely, our experiments demonstrate that the improvements attained by accommodating multiple poses can be negligible compared to the gains yielded by normalizing scores and using the most appropriate classifier for uncontrolled data. We conclude with a qualitative evaluation of the proposed algorithm on publicly available images of the Boston Marathon crowds. Although the results of our evaluations are encouraging, they confirm that there is still room for improvement in terms of robustness to out-of-plane rotation, blur and occlusion.
Jeremiah R. Barr, Kevin W. Bowyer, Patrick J. Flynn
WACV3
2014 Active Clustering with Ensembles for Social structure extraction
abstract
We introduce a method for extracting the social network structure for the persons appearing in a set of video clips. Individuals are unknown, and are not matched against known enrollments. An identity cluster representing an individual is formed by grouping similar-appearing faces from different videos. Each identity cluster is represented by a node in the social network. Two nodes are linked if the faces from their clusters appeared together in one or more video frames. Our approach incorporates a novel active clustering technique to create more accurate identity clusters based on feedback from the user about ambiguously matched faces. The final output consists of one or more network structures that represent the social group(s), and a list of persons who potentially connect multiple social groups. Our results demonstrate the efficacy of the proposed clustering algorithm and network analysis techniques.
Jeremiah R. Barr, Leonardo A. Cament, Kevin W. Bowyer, Patrick J. Flynn
WACV4
2014 Finger-knuckle-print verification based on vector consistency of corresponding interest points
abstract
This paper proposes a novel finger-knuckle-print (FKP) verification method based on vector consistency among corresponding interest points (CIPs) detected from aligned finger images. We used two different approaches for reliable detection of CIPs; one method employs SIFT features and captures gradient directionality, and the other method employs phase correlation to represent the intensity field surrounding an interest point. The consistency of interframe displacements between pairs of matching CIPs in a match pair is used as a matching score. Such displacements will show consistency in a genuine match but not in an impostor match. Experimental results show that the proposed approach is effective in FKP verification.
Min-Ki Kim, Patrick J. Flynn
WACV2
2014 Iris crypts: Multi-scale detection and shape-based matching
abstract
This paper presents an improved framework for iris crypt detection and matching that outperforms both previous methods and manual annotations. The system uses a multi-scale pyramid architecture to detect feature candidates before they are further examined and optimized by heuristic-based methods. The dissimilarity between irises are measured by a two-stage matcher in the simple to complex order. The first stage estimates the global dissimilarity and rejects the majority of unmatching candidates. The surviving pairs are matched by local dissimilarities between each crypt pair using shape descriptors. The proposed framework showed significant performance improvement in both identification and verification context.
Patrick J. Flynn
WACV2
2014 Framework for Active Clustering With Ensembles
abstract
Clustering approaches can alleviate the burden of tagging face identities in ad hoc video and image collections. We introduce a novel semisupervised framework for clustering face patterns into identity groups using minimal human interaction. This technique combines concepts from ensemble clustering and active learning to improve clustering accuracy. The framework actively queries the user for a soft link constraint between each pair of neighboring faces that are ambiguously matched according to the ensemble. We demonstrate the efficacy of our approach with the broadest evaluation of active face clustering algorithms to date. Our evaluations focus on data that is appropriate for human-in-the-loop face recognition, including blurry point-and-shoot videos, images of women seen before and after the application of makeup, and photographs of twins. The results indicate that ensemble-based constrained clustering algorithms are generally more robust to noise than alternative approaches. Finally, we show that the proposed clustering algorithm is more accurate and parsimonious than the current state-of-the-art.
Jeremiah R. Barr, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.3
2014 Double Trouble: Differentiating Identical Twins by Face Recognition
abstract
Facial recognition algorithms should be able to operate even when similar-looking individuals are encountered, or even in the extreme case of identical twins. An experimental data set comprised of 17486 images from 126 pairs of identical twins (252 subjects) collected on the same day and 6864 images from 120 pairs of identical twins (240 subjects) with images taken a year later was used to measure the performance on seven different face recognition algorithms. Performance is reported for variations in illumination, expression, gender, and age for both the same day and cross-year image sets. Regardless of the conditions of image acquisition, distinguishing identical twins are significantly harder than distinguishing subjects who are not identical twins for all algorithms.
Jeffrey R. Paone, Patrick J. Flynn, P. Jonathon Phillips, Kevin W. Bowyer, Richard W. Vorder Bruegge, Patrick Grother, George W. Quinn, Matthew Pruitt, Jason M. Grant
IEEE Trans. Inf. Forensics Secur.2
2013 Using isolated vowel sounds for classification of Mild Traumatic Brain Injury
abstract
Concussions are Mild Traumatic Brain Injuries (mTBI) that are common in contact sports and are often difficult to diagnose due to the delayed appearance of symptoms. This paper explores the feasibility of using speech analysis for detecting mTBI. Recordings are taken on a mobile device from athletes participating in a boxing tournament following each match. Vowel sounds are isolated from the recordings and acoustic features are extracted and used to train several one-class machine learning algorithms in order to predict whether an athlete is concussed. Prediction results are verified against the diagnoses made by a ringside medical team at the time of recording and performance evaluation shows prediction accuracies of up to 98%.
Michael Falcone, Nikhil Yadav, Christian Poellabauer, Patrick J. Flynn
ICASSP4
2013 Pose-Robust Recognition of Low-Resolution Face Images
abstract
Face images captured by surveillance cameras usually have poor resolution in addition to uncontrolled poses and illumination conditions, all of which adversely affect the performance of face matching algorithms. In this paper, we develop a completely automatic, novel approach for matching surveillance quality facial images to high-resolution images in frontal pose, which are often available during enrollment. The proposed approach uses multidimensional scaling to simultaneously transform the features from the poor quality probe images and the high-quality gallery images in such a manner that the distances between them approximate the distances had the probe images been captured in the same conditions as the gallery images. Tensor analysis is used for facial landmark localization in the low-resolution uncontrolled probe images for computing the features. Thorough evaluation on the Multi-PIE dataset and comparisons with state-of-the-art super-resolution and classifier-based approaches are performed to illustrate the usefulness of the proposed approach. Experiments on surveillance imagery further signify the applicability of the framework. We also show the usefulness of the proposed approach for the application of tracking and recognition in surveillance videos.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn, Kevin W. Bowyer
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 A sparse representation approach to face matching across plastic surgery
abstract
Plastic surgery procedures can significantly alter facial appearance, thereby posing a serious challenge even to the state-of-the-art face matching algorithms. In this paper, we propose a novel approach to address the challenges involved in automatic matching of faces across plastic surgery variations. In the proposed formulation, part-wise facial characterization is combined with the recently popular sparse representation approach to address these challenges. The sparse representation approach requires several images per subject in the gallery to function effectively which is often not available in several use-cases, as in the problem we address in this work. The proposed formulation utilizes images from sequestered non-gallery subjects with similar local facial characteristics to fulfill this requirement. Extensive experiments conducted on a recently introduced plastic surgery database [17] consisting of 900 subjects highlight the effectiveness of the proposed approach.
Gaurav Aggarwal, Soma Biswas, Patrick J. Flynn, Kevin W. Bowyer
WACV3
2012 Predicting good, bad and ugly match Pairs
abstract
Several sources of variation in facial appearance that affect face matching performance have long been investigated. The recently introduced GBU challenge problem [1] indicates that there can be significant variation in performance across different partitions of the data, even when the impact of most known factors is eliminated or significantly reduced by the data collection and experimentation protocol. The GBU challenge problem consists of three partitions which are called the Good (easy to match), the Bad (average matching difficulty) and the Ugly (difficult to match). In this paper, we investigate various image and facial characteristics that can account for the observed significant difference in performance across these partitions. Given a match pair, we aim to predict the partition it belongs to. Partial Least Squares (PLS)-based regression is used to perform the prediction task. Our analysis indicates that the match pairs from the three partitions differ from each other in terms of simple but often ignored factors like image sharpness, hue, saturation and extent of facial expressions.
Gaurav Aggarwal, Soma Biswas, Patrick J. Flynn, Kevin W. Bowyer
WACV3
2012 ROARS: a robust object archival system for data intensive scientific computing
Hoang Bui, Peter Bui, Patrick J. Flynn, Douglas Thain
Distributed Parallel Databases3
2012 Face Recognition from Video: a Review
abstract
Driven by key law enforcement and commercial applications, research on face recognition from video sources has intensified in recent years. The ensuing results have demonstrated that videos possess unique properties that allow both humans and automated systems to perform recognition accurately in difficult viewing conditions. However, significant research challenges remain as most video-based applications do not allow for controlled recordings. In this survey, we categorize the research in this area and present a broad and deep review of recently proposed methods for overcoming the difficulties encountered in unconstrained settings. We also draw connections between the ways in which humans and current algorithms recognize faces. An overview of the most popular and difficult publicly available face video databases is provided to complement these discussions. Finally, we cover key research challenges and opportunities that lie ahead for the field as a whole.
Jeremiah R. Barr, Kevin W. Bowyer, Patrick J. Flynn, Soma Biswas
Int. J. Pattern Recognit. Artif. Intell.3
2012 Multidimensional Scaling for Matching Low-Resolution Face Images
abstract
Face recognition performance degrades considerably when the input images are of Low Resolution (LR), as is often the case for images taken by surveillance cameras or from a large distance. In this paper, we propose a novel approach for matching low-resolution probe images with higher resolution gallery images, which are often available during enrollment, using Multidimensional Scaling (MDS). The ideal scenario is when both the probe and gallery images are of high enough resolution to discriminate across different subjects. The proposed method simultaneously embeds the low-resolution probe images and the high-resolution gallery images in a common space such that the distance between them in the transformed space approximates the distance had both the images been of high resolution. The two mappings are learned simultaneously from high-resolution training images using an iterative majorization algorithm. Extensive evaluation of the proposed approach on the Multi-PIE data set with probe image resolution as low as 8 6 pixels illustrates the usefulness of the method. We show that the proposed approach improves the matching performance significantly as compared to performing matching in the low-resolution domain or using super-resolution techniques to obtain a higher resolution test image prior to recognition. Experiments on low-resolution surveillance images from the Surveillance Cameras Face Database further highlight the effectiveness of the approach.
Soma Biswas, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.3
2012 A Multialgorithm Analysis of Three Iris Biometric Sensors
abstract
The issue of interoperability between iris sensors is an important topic in large-scale and long-term applications of iris biometric systems. This work compares three commercially available iris sensors and three iris matching systems and investigates the impact of cross-sensor matching on system performance in comparison to single-sensor performance. Several factors which may impact single-sensor and cross-sensor performance are analyzed, including changes in the acquisition environment and differences in dilation ratio between iris images. The sensors are evaluated using three different iris matching algorithms, and conclusions are drawn regarding the interaction between the sensors and the matching algorithm in both the cross-sensor and single-sensor scenarios. Finally, the relative performances of the three sensors are compared.
Ryan Connaughton, Amanda Sgroi, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.4
2012 Human and Machine Performance on Periocular Biometrics Under Near-Infrared Light and Visible Light
abstract
Periocular biometrics is the recognition of individuals based on the appearance of the region around the eye. Periocular recognition may be useful in applications where it is difficult to obtain a clear picture of an iris for iris biometrics, or a complete picture of a face for face biometrics. Previous periocular research has used either visible-light (VL) or near-infrared (NIR) light images, but no prior research has directly compared the two illuminations using images with similar resolution. We conducted an experiment in which volunteers were asked to compare pairs of periocular images. Some pairs showed images taken in VL, and some showed images taken in NIR light. Participants labeled each pair as belonging to the same person or to different people. Untrained participants with limited viewing times correctly classified VL image pairs with 88% accuracy, and NIR image pairs with 79% accuracy. For comparison, we presented pairs of iris images from the same subjects. In addition, we investigated differences between performance on light and dark eyes and relative helpfulness of various features in the periocular region under different illuminations. We calculated performance of three computer algorithms on the periocular images. Performance for humans and computers was similar.
Karen Hollingsworth, Shelby Solomon Darnell, Philip E. Miller, Damon L. Woodard, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.6
2012 Analysis of Facial Marks to Distinguish Between Identical Twins
abstract
Identical twin face recognition is a challenging task due to the existence of a high degree of correlation in overall facial appearance. Commercial face recognition systems exhibit poor performance in differentiating between identical twins under practical conditions. In this paper, we study the usability of facial marks as biometric signatures to distinguish between identical twins. We propose a multiscale automatic facial mark detector based on a gradient-based operator known as the fast radial symmetry transform. The transform detects bright or dark regions with high radial symmetry at different scales. Next, the detections are tracked across scales to determine the prominence of facial marks. Extensive experiments are performed both on manually annotated and on automatically detected facial marks to evaluate the usefulness of facial marks as biometric signatures. Experiment results are based on identical twin images acquired at the 2009 Twins Days Festival in Twinsburg, Ohio. The results of our analysis signify the usefulness of the distribution of facial marks as a biometric signature. In addition, our results indicate the existence of some degree of correlation between geometric distribution of facial marks across identical twins.
Nisha Srinivas, Gaurav Aggarwal, Patrick J. Flynn, Richard W. Vorder Bruegge
IEEE Trans. Inf. Forensics Secur.3
2011 Pose-robust recognition of low-resolution face images
abstract
Face images captured by surveillance cameras usually have poor resolution in addition to uncontrolled poses and illumination conditions which adversely affect performance of face matching algorithms. In this paper, we develop a novel approach for matching surveillance quality facial images to high resolution images in frontal pose which are often available during enrollment. The proposed approach uses Multidimensional Scaling to simultaneously transform the features from the poor quality probe images and the high quality gallery images in such a manner that the distances between them approximate the distances had the probe images been captured in the same conditions as the gallery images. Thorough evaluation on the Multi-PIE dataset and comparisons with state-of-the-art super-resolution and classifier based approaches are performed to illustrate the usefulness of the proposed approach. Experiments on real surveillance images further signify the applicability of the framework.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
CVPR3
2011 Distinguishing identical twins by face recognition
abstract
The paper measures the ability of face recognition algorithms to distinguish between identical twin siblings. The experimental dataset consists of images taken of 126 pairs of identical twins (252 people) collected on the same day and 24 pairs of identical twins (48 people) with images collected one year apart. In terms of both the number of paris of twins and lapsed time between acquisitions, this is the most extensive investigation of face recognition performance on twins to date. Recognition experiments are conducted using three of the top submissions to the Multiple Biometric Evaluation (MBE) 2010 Still Face Track [1]. Performance results are reported for both same day and cross year matching. Performance results are broken out by lighting conditions (studio and outside); expression (neutral and smiling); gender and age. Confidence intervals were generated by a bootstrap method. This is the most detailed covariate analysis of face recognition of twins to date.
P. Jonathon Phillips, Patrick J. Flynn, Kevin W. Bowyer, Richard W. Vorder Bruegge, Patrick Grother, George W. Quinn, Matthew Pruitt
FG2
2011 Face recognition in low-resolution videos using learning-based likelihood measurement model
abstract
Low-resolution surveillance videos with uncontrolled pose and illumination present a significant challenge to both face tracking and recognition algorithms. Considerable appearance difference between the probe videos and high-resolution controlled images in the gallery acquired during enrollment makes the problem even harden In this paper, we extend the simultaneous tracking and recognition framework [22] to address the problem of matching high-resolution gallery images with surveillance quality probe videos. We propose using a learning-based likelihood measurement model to handle the large appearance and resolution difference between the gallery images and probe videos. The measurement model consists of a mapping which transforms the gallery and probe features to a space in which their inter-Euclidean distances approximate the distances that would have been obtained had all the descriptors been computed from good quality frontal images. Experimental results on real surveillance quality videos and comparisons with related approaches show the effectiveness of the proposed framework.
Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
IJCB3
2011 Difficult imaging covariates or difficult subjects? - An empirical investigation
abstract
The performance of face recognition algorithms is affected both by external factors and internal subject characteristics [I]. Reliably identifying these factors and understanding their behavior on performance can potentially serve two important goals to predict the performance of the algorithms at novel deployment sites and to design appropriate acquisition environments at prospective sites to optimize performance. There have been a few recent efforts in this direction that focus on identifying factors that affect face recognition performance but there has been no extensive study regarding the consistency of the effects various factors have on algorithms when other covariates vary. To give an example, a smiling target image has been reported to be better than a neutral expression image, but is this true across all possible illumination conditions, head poses, gender, etc.? In this paper, we perform rigorous experiments to provide answers to such questions. Our investigation indicates that controlled lighting and smiling expression are the most favorable conditions that consistently give superior performance even when other factors are allowed to vary. We also observe that internal subject characterization using biometric menagerie-based classification shows very weak consistency when external conditions are allowed to vary.
Jeffrey R. Paone, Soma Biswas, Gaurav Aggarwal, Patrick J. Flynn
IJCB4
2011 Facial recognition of identical twins
abstract
Biometric identification systems must be able to distinguish between individuals even in situations where the bio metric signature may be similar, such as in the case of identical twins. This paper presents experiments done in facial recognition using data from a set of images of twins. This work establishes the current state of facial recognition in regards to twins and the accuracy of current state-of-the art programs in distinguishing between identical twins using three commercial face matchers, Cognitec 8.3.2.0, VeriLook 4.0, and PittPatt 4.2.1 and a baseline matcher employing Local Region PCA. Overall, it was observed that Cognitec had the best performance. All matchers, how ever, saw degradation in performance compared to an experiment where the ability to distinguish unrelated persons was assessed. In particular, lighting and expression seemed to have affected performance the most.
Matthew Pruitt, Jason M. Grant, Jeffrey R. Paone, Patrick J. Flynn, Richard W. Vorder Bruegge
IJCB4
2011 Twins 3D face recognition challenge
abstract
Existing 3D face recognition algorithms have achieved high enough performances against public datasets like FRGC v2, that it is difficult to achieve further significant increases in recognition performance. However, the 3D TEC dataset is a more challenging dataset which consists of 3D scans of 107 pairs of twins that were acquired in a single session, with each subject having a scan of a neutral expression and a smiling expression. The combination of factors related to the facial similarity of identical twins and the variation in facial expression makes this a challenging dataset. We conduct experiments using state of the art face recognition algorithms and present the results. Our results indicate that 3D face recognition of identical twins in the presence of varying facial expressions is far from a solved problem, but that good performance is possible.
Vipin Vijayan, Kevin W. Bowyer, Patrick J. Flynn, Di Huang 0001, Liming Chen 0002, Mark Hansen, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris
IJCB3
2011 Textured mesh generation of extracted regions from urban range-scanned LIDAR data
abstract
LIDAR range scanners are a popular tool for data acquisition in the field of urban modeling, archaeological preservation, city planning, and more. However, range scanners output a series of discrete distance samples as disconnected points, providing a fundamentally incomplete representation of the underlying structure in a scene. Data from a single LIDAR scan of a region can be trivially triangulated, but when multiple scanners are involved, or when the acquisition platform is mobile, those inter-point relationships are lost. We propose a technique to triangulate such data by identifying logical surfaces within the data and triangulating those surfaces individually. The resulting triangulations are simplified dramatically by using information about the shape of the regions, and texture is applied from camera imagery on the scan vehicle. The result is a high fidelity representation of a scene which is more efficiently rendered on modern hardware than point sets, which occupies less space in memory, and which brings us closer to a solid representation of the true scanned scene.
Alexandri Zavodny, Patrick J. Flynn
ICME2
2011 Detecting questionable observers using face track clustering
abstract
We introduce the questionable observer detection problem: Given a collection of videos of crowds, determine which individuals appear unusually often across the set of videos. The algorithm proposed here detects these individuals by clustering sequences of face images. To provide robustness to sensor noise, facial expression and resolution variations, blur, and intermittent occlusions, we merge similar face image sequences from the same video and discard outlying face patterns prior to clustering. We present experiments on a challenging video dataset. The results show that the proposed method can surpass the performance of a clustering algorithm based on the VeriLook face recognition software by Neurotechnology both in terms of the detection rate and the false detection frequency.
Jeremiah R. Barr, Kevin W. Bowyer, Patrick J. Flynn
WACV3
2011 Information fusion in low-resolution iris videos using Principal Components Transform
abstract
The focus of this work is on improving the recognition performance of low-resolution iris video frames acquired under varying illumination. To facilitate this, an image-level fusion scheme with modest computational requirements is proposed. The proposed algorithm uses the evidence of multiple image frames of the same iris to extract discriminatory information via the Principal Components Transform (PCT). Experimental results on a subset of the MBGC NIR iris database demonstrate the utility of this scheme to achieve improved recognition accuracy when low-resolution probe images are compared against high-resolution gallery images.
Raghavender R. Jillela, Arun Ross, Patrick J. Flynn
WACV3
2011 Genetically identical irises have texture similarity that is not detected by iris biometrics
Karen Hollingsworth, Kevin W. Bowyer, Stephen Lagree, Samuel P. Fenker, Patrick J. Flynn
Comput. Vis. Image Underst.5
2011 Useful features for human verification in near-infrared periocular images
Karen Hollingsworth, Kevin W. Bowyer, Patrick J. Flynn
Image Vis. Comput.3
2011 Improved Iris Recognition through Fusion of Hamming Distance and Fragile Bit Distance
abstract
The most common iris biometric algorithm represents the texture of an iris using a binary iris code. Not all bits in an iris code are equally consistent. A bit is deemed fragile if its value changes across iris codes created from different images of the same iris. Previous research has shown that iris recognition performance can be improved by masking these fragile bits. Rather than ignoring fragile bits completely, we consider what beneficial information can be obtained from the fragile bits. We find that the locations of fragile bits tend to be consistent across different iris codes of the same eye. We present a metric, called the fragile bit distance, which quantitatively measures the coincidence of the fragile bit patterns in two iris codes. We find that score fusion of fragile bit distance and Hamming distance works better for recognition than Hamming distance alone. To our knowledge, this is the first and only work to use the coincidence of fragile bit locations to improve the accuracy of matches.
Karen Hollingsworth, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 ROARS: a scalable repository for data intensive scientific computing
abstract
As scientific research becomes more data intensive, there is an increasing need for scalable, reliable, and high performance storage systems. Such data repositories must provide both data archival services and rich metadata, and cleanly integrate with large scale computing resources. ROARS is a hybrid approach to distributed storage that provides both large, robust, scalable storage and efficient rich metadata queries for scientific applications. In this paper, we demonstrate that ROARS is capable of importing and exporting large quantities of data, migrating data to new storage nodes, providing robust fault tolerance, and generating materialized views based on metadata queries. Our experimental results demonstrate that ROARS' aggregate throughput scales with the number of concurrent clients while providing fault-tolerant data access. ROARS is currently being used to store 5.1TB of data in our local biometrics repository.
Hoang Bui, Peter Bui, Patrick J. Flynn, Douglas Thain
HPDC3
2010 Towards long term data quality in a large scale biometrics experiment
abstract
Quality of data plays a very important role in any scientific research. In this paper we present some of the challenges that we face in managing and maintaining data quality for a terabyte scale biometrics repository. We have developed a step by step model to capture, ingest, validate, and prepare data for biometrics research. During these processes, there are many hidden errors which can be introduced into the data. Those errors can affect the overall quality of data, and thus can skew the results of biometrics research. We discuss necessary steps we have taken to reduce and eliminate the errors. Steps such as data replication, automated data validation, and logging metadata changes are both necessary and crucial to improve the quality and reliability of our data.
Hoang Bui, Diane Wright, Clarence Helm, Rachel Witty, Patrick J. Flynn, Douglas Thain
HPDC5
2010 Degradation of iris recognition performance due to non-cosmetic prescription contact lenses
Sarah E. Baker, Amanda Hentz, Kevin W. Bowyer, Patrick J. Flynn
Comput. Vis. Image Underst.4
2010 FRVT 2006 and ICE 2006 Large-Scale Experimental Results
abstract
This paper describes the large-scale experimental results from the Face Recognition Vendor Test (FRVT) 2006 and the Iris Challenge Evaluation (ICE) 2006. The FRVT 2006 looked at recognition from high-resolution still frontal face images and 3D face images, and measured performance for still frontal face images taken under controlled and uncontrolled illumination. The ICE 2006 evaluation reported verification performance for both left and right irises. The images in the ICE 2006 intentionally represent a broader range of quality than the ICE 2006 sensor would normally acquire. This includes images that did not pass the quality control software embedded in the sensor. The FRVT 2006 results from controlled still and 3D images document at least an order-of-magnitude improvement in recognition performance over the FRVT 2002. The FRVT 2006 and the ICE 2006 compared recognition performance from high-resolution still frontal face images, 3D face images, and the single-iris images. On the FRVT 2006 and the ICE 2006 data sets, recognition performance was comparable for high-resolution frontal face, 3D face, and the iris images. In an experiment comparing human and algorithms on matching face identity across changes in illumination on frontal face images, the best performing algorithms were more accurate than humans on unfamiliar faces.
P. Jonathon Phillips, W. Todd Scruggs, Alice J. O'Toole, Patrick J. Flynn, Kevin W. Bowyer, Cathy L. Schott, Matthew Sharpe
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 All-Pairs: An Abstraction for Data-Intensive Computing on Campus Grids
abstract
Today, campus grids provide users with easy access to thousands of CPUs. However, it is not always easy for nonexpert users to harness these systems effectively. A large workload composed in what seems to be the obvious way by a naive user may accidentally abuse shared resources and achieve very poor performance. To address this problem, we argue that campus grids should provide end users with high-level abstractions that allow for the easy expression and efficient execution of data-intensive workloads. We present one example of an abstraction—All-Pairs—that fits the needs of several applications in biometrics, bioinformatics, and data mining. We demonstrate that an optimized All-Pairs abstraction is both easier to use than the underlying system, achieve performance orders of magnitude better than the obvious but naive approach, and is both faster and more efficient than a tuned conventional approach. This abstraction has been in production use for one year on a 500 CPU campus grid at the University of Notre Dame and has been used to carry out a groundbreaking analysis of biometric data.
Christopher Moretti, Hoang Bui, Karen Hollingsworth, Brandon Rich, Patrick J. Flynn, Douglas Thain
IEEE Trans. Parallel Distributed Syst.5
2009 Pupil dilation degrades iris biometric performance
Karen Hollingsworth, Kevin W. Bowyer, Patrick J. Flynn
Comput. Vis. Image Underst.3
2009 Iris recognition using signal-level fusion of frames from video
abstract
We take advantage of the temporal continuity in an iris video to improve matching performance using signal-level fusion. From multiple frames of a frontal iris video, we create a single average image. For comparison, we reimplement three score-level fusion methods (Ma, Krichen, and Schmid). We find that our signal-level fusion ofNimages performs better than Ma's or Krichen's score-level fusion methods ofNHamming distance scores. Our signal-level fusion performs comparably to Schmid's log-likelihood method of score-level fusion, and our method achieves this performance using less computation time. We compare our signal fusion method with another new method: a multigallery, multiprobe method involving score-level fusion ofN2Hamming distances. The multigallery, multiprobe score fusion has slightly better recognition performance, while the signal fusion has significant advantages in memory and computation requirements. No published prior work has shown any advantage of the use of video over still images in iris biometrics.
Karen Hollingsworth, Tanya Peters, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.4
2008 BXGrid: A Data Repository and Workflow Abstraction for Biometrics Research
abstract
Research in the field of biometrics depends on the effective management of large amounts of data and computation. Current research projects in biometrics acquire many terabytes of images and video of subjects in many different modes and situations, annotated with detailed metadata. To study the effectiveness of new algorithms for identifying people, researchers must exhaustively compare large numbers of measurements with a variety of custom functions. The quality of the end results is often dependent upon the sheer amount of data marshalled to support it.To address these challenges, we are constructing BXGrid, an end-to-end computing system for conducting biometrics research. BXGrid assists with the entire research process from data acquisition all the way to generating results for publication. Because the entire chain of research is kept consistently within one system, multiple users may easily share tools and results, building off of each other's work. BXGrid also helps to ensure scientific integrity by automating a variety of consistency checks, external data audits, and reproduction of existing results.
Hoang Bui, Deborah Thomas, Michael Kelly, Christopher Lyon, Douglas Thain, Patrick J. Flynn
eScience6
2008 Using Small Abstractions to Program Large Distributed Systems
abstract
Distributed systems such as clusters, clouds and grids remain a difficult platform for executing large data intensive workloads. Even sophisticated users struggle to shape complex workloads into the simple 'assembly language' of file transfer and job submission. To address this, problem, we advocate the use of abstractions, which are simple frameworks for expressing large structured problems. In this talk, we will discuss three examples of abstractions developed at the University of Notre Dame for scientific applications. In each case, we have been able to scale up workloads one to wto orders of magnitude larger than we previously feasible. Through each example, we will address some persistent obstacles in the field of distributed computing.
Douglas Thain, Christopher Moretti, Hoang Bui, Nitesh V. Chawla, Patrick J. Flynn
eScience6
2008 Rotated Profile Signatures for robust 3D feature detection
abstract
While recent years have seen progress in face recognition from 3D images, nonfrontal head pose is still a challenge to existing techniques. We introduce a new system for 3D face recognition that is robust to facial pose variation. Large degrees of facial pose variation may lead to a significant fraction of the features visible in frontal images being occluded. High accuracy automatic feature and pose detection is performed by a new technique called rotated profile signatures (RPS). Experiments are performed on the largest available database of 3D faces acquired under varying pose. This database contains over 7,300 total images of 406 unique subjects gathered at the University of Notre Dame. Experimental results show that the RPS detection algorithm is capable of performing nose detection with greater than 96.5% accuracy across the pose variation represented in the data set used.
Timothy C. Faltemier, Kevin W. Bowyer, Patrick J. Flynn
FG3
2008 All-pairs: An abstraction for data-intensive cloud computing
abstract
Although modern parallel and distributed computing systems provide easy access to large amounts of computing power, it is not always easy for non-expert users to harness these large systems effectively. A large workload composed in what seems to be the obvious way by a naive user may accidentally abuse shared resources and achieve very poor performance. To address this problem, we propose that production systems should provide end users with high-level abstractions that allow for the easy expression and efficient execution of data intensive workloads. We present one example of an abstraction - all-pairs - that fits the needs of several data-intensive scientific applications. We demonstrate that an optimized all-pairs abstraction is both easier to use than the underlying system, and achieves performance orders of magnitude better than the obvious but naive approach, and twice as fast as a hand-optimized conventional approach.
Christopher Moretti, Jared Bulosan, Douglas Thain, Patrick J. Flynn
IPDPS4
2008 Image understanding for iris biometrics: A survey
Kevin W. Bowyer, Karen Hollingsworth, Patrick J. Flynn
Comput. Vis. Image Underst.3
2008 Using multi-instance enrollment to improve performance of 3D face recognition
Timothy C. Faltemier, Kevin W. Bowyer, Patrick J. Flynn
Comput. Vis. Image Underst.3
2008 A Region Ensemble for 3-D Face Recognition
abstract
In this paper, we introduce a new system for 3D face recognition based on the fusion of results from a committee of regions that have been independently matched. Experimental results demonstrate that using 28 small regions on the face allow for the highest level of 3D face recognition. Score-based fusion is performed on the individual region match scores and experimental results show that the Borda count and consensus voting methods yield higher performance than the standard sum, product, and min fusion rules. In addition, results are reported that demonstrate the robustness of our algorithm by simulating large holes and artifacts in images. To our knowledge, no other work has been published that uses a large number of 3D face regions for high-performance face matching. Rank one recognition rates of 97.2% and verification rates of 93.2% at a 0.1% false accept rate are reported and compared to other methods published on the face recognition grand challenge v2 data set.
Timothy C. Faltemier, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Inf. Forensics Secur.3
2007 Challenges in Executing Data Intensive Biometric Workloads on a Desktop Grid
abstract
Desktop grids have traditionally focused on executing computation intensive workloads. Can they also be used to execute data-intensive workloads? To answer this question, we present a case study of a data intensive biometric application which is infeasible to process on a single machine. We evaluate the capacity of a desktop grid to store and deliver the data need to execute the workload, and compare several general techniques for data deployment. Selecting the most scalable technique, we execute and evaluate five large production workloads on a 350-CPU desktop grid. We observe that this technique is sensitive to many parameters, and propose that an ideal system should be responsible for choosing the proper decomposition of a workload.
Christopher Moretti, Timothy C. Faltemier, Douglas Thain, Patrick J. Flynn
IPDPS4
2007 Comments on the CASIA version 1.0 Iris Data Set
abstract
We note that the images in the CASIA version 1.0 iris dataset have been edited so that the pupil area is replaced by a circular region of uniform intensity. We recommend that this dataset is no longer used in iris biometrics research, unless there this a compelling reason that takes into account the nature of the images. In addition, based on our experience with the Iris Challenge Evaluation (ICE) 2005 technology development project, we make recommendations for reporting results of iris recognition experiments.
P. Jonathon Phillips, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 A survey of approaches and challenges in 3D and multi-modal 3D + 2D face recognition
Kevin W. Bowyer, Kyong I. Chang, Patrick J. Flynn
Comput. Vis. Image Underst.3
2006 Multiple Nose Region Matching for 3D Face Recognition under Varying Facial Expression
abstract
An algorithm is proposed for 3D face recognition in the presence of varied facial expressions. It is based on combining the match scores from matching multiple overlapping regions around the nose. Experimental results are presented using the largest database employed to date in 3D face recognition studies, over 4,000 scans of 449 subjects. Results show substantial improvement over matching the shape of a single larger frontal face region. This is the first approach to use multiple overlapping regions around the nose to handle the problem of expression variation.
Kyong I. Chang, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Face Recognition Using 2-D, 3-D, and Infrared: Is Multimodal Better Than Multisample?
abstract
This work examines face recognition using normal intensity images, infrared images, three-dimensional shape, and combinations of these. We compare the performance improvement obtained by combining three-dimensional or infrared with normal intensity images (a "multimodal" approach) to the performance improvement obtained by using multiple intensity images (a "multisample" approach). Combining results from different types of imagery gives significantly higher recognition rates than are obtained by using a single intensity image. However, significantly higher recognition rates are also obtained by combining results from multiple intensity images. Overall, initial results indicate that, using an "eigen-face" recognition algorithm and weighted score fusion, multisample techniques can result in a performance increase comparable to that of multimodal techniques
Kevin W. Bowyer, Kyong I. Chang, Patrick J. Flynn
Proc. IEEE3
2005 Overview of the Face Recognition Grand Challenge
abstract
Over the last couple of years, face recognition researchers have been developing new techniques. These developments are being fueled by advances in computer vision techniques, computer design, sensor design, and interest in fielding face recognition systems. Such advances hold the promise of reducing the error rate in face recognition systems by an order of magnitude over Face Recognition Vendor Test (FRVT) 2002 results. The face recognition grand challenge (FRGC) is designed to achieve this performance goal by presenting to researchers a six-experiment challenge problem along with data corpus of 50,000 images. The data consists of 3D scans and high resolution still imagery taken under controlled and uncontrolled conditions. This paper describes the challenge problem, data corpus, and presents baseline performance and preliminary results on natural statistics of facial imagery.
P. Jonathon Phillips, Patrick J. Flynn, W. Todd Scruggs, Kevin W. Bowyer, Jin Chang, Kevin Hoffman, Joe Marques, Jaesik Min, William J. Worek
CVPR (1)2
2005 Personal Identification Utilizing Finger Surface Features
abstract
In this paper we present a novel approach for personal identification, which utilizes finger surface features as a biometric identifier. Using dense range data images of the hand, we calculate the curvature-based surface representation, shape index, for the index, middle, and ring fingers. This representation is used for comparisons to determine subject similarity. Our experiments involve the use of a large data set of range images collected over time. We examine the performance of individual finger surfaces as a biometric identifier as well as the performance when using the three finger surfaces in conjunction. The results of our experiments are presented, which indicate that this approach performs well for a first-of-its-kind biometric technique.
Damon L. Woodard, Patrick J. Flynn
CVPR (2)2
2005 IR and visible light face recognition
Patrick J. Flynn, Kevin W. Bowyer
Comput. Vis. Image Underst.2
2005 Finger surface as a biometric identifier
Damon L. Woodard, Patrick J. Flynn
Comput. Vis. Image Underst.2
2005 An Evaluation of Multimodal 2D+3D Face Biometrics
abstract
We report on the largest experimental study to date in multimodal 2D+3D face recognition, involving 198 persons in the gallery and either 198 or 670 time-lapse probe images. PCA-based methods are used separately for each modality and match scores in the separate face spaces are combined for multimodal recognition. Major conclusions are: 1) 2D and 3D have similar recognition performance when considered individually, 2) combining 2D and 3D results using a simple weighting scheme outperforms either 2D or 3D alone, 3) combining results from two or more 2D images using a similar weighting scheme also outperforms a single 2D image, and 4) combined 2D+3D outperforms the multiimage 2D result. This is the first (so far, only) work to present such an experimental control to substantiate multimodal performance improvement.
Kyong I. Chang, Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.3
2003 Free-Form 3D Object Recognition In Range Data Using Weak Correspondence Between Local Features
abstract
Model-Based 3D object recognition systems have a variety of potential applications, but widespread use of such systems has not occurred, due to a number of factors including the representational limitations of models. One historical limitation is the discriminatory representation of free-form objects. The system described in this paper recognizes free-form objects in dense range data acquired by a structured light rangefinder. Images and object models are represented as a network of salient segments which are then brought into correspondence until a reliable pose estimate is available. Experiments with a database of images and object models highlight the contributions of this system.
Richard J. Campbell, Patrick J. Flynn
Int. J. Pattern Recognit. Artif. Intell.2
2003 Aggressive region growing for speckle reduction in ultrasound images
Ruming Yin, Patrick J. Flynn, Shira L. Broschat
Pattern Recognit. Lett.3
2002 Saliency Sequential Surface Organization for Free-Form Object Recognition
Kim L. Boyer, Ravi Srikantiah, Patrick J. Flynn
Comput. Vis. Image Underst.3
2002 Pair-Wise Range Image Registration: A Study in Outlier Classification
Gerald Dalley, Patrick J. Flynn
Comput. Vis. Image Underst.2
2001 Edge-Based Artifact Mitigation in a Wavelet Transform Coding Framework
Anand Kalyanaraman, Patrick J. Flynn
Data Compression Conference2
2001 A Survey Of Free-Form Object Representation and Recognition Techniques
Richard J. Campbell, Patrick J. Flynn
Comput. Vis. Image Underst.2
2001 Special Issue on Empirical Evaluation of Computer Vision Algorithms
Patrick J. Flynn, Adam W. Hoover, P. Jonathon Phillips
Comput. Vis. Image Underst.1
2000 A 20th Anniversary Survey: Introduction to 'Content-Based Image Retrieval at the End of the Early Years'
Kevin W. Bowyer, Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 The 20th Anniversary of the IEEE Transactions on Pattern Analysis and Machine Intelligence
Kevin W. Bowyer, Patrick J. Flynn, Rangachar Kasturi
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Eigenshapes for 3D Object Recognition in Range Data
abstract
Much of the recent research in object recognition has adopted an appearance-based scheme, wherein objects to be recognized are represented as a collection of prototypes in a multidimensional space spanned by a number of characteristic vectors (eigen-images) obtained from training views. In this paper, we extend the appearance-based recognition scheme to handle range (shape) data. The result of training is a set of 'eigensurfaces' that capture the gross shape of the objects. These techniques are used to form a system that recognizes objects under an arbitrary rotational pose transformation. The system has been tested on a 20 object database including free-form objects and a 54 object database of manufactured parts. Experiments with the system point out advantages and also highlight challenges that must be studied in future research.
Richard J. Campbell, Patrick J. Flynn
CVPR2
1998 On Approximating Rough Curves with Fractal Functions
Wayne O. Cochran, John C. Hart, Patrick J. Flynn
Graphics Interface3
1998 Recent Progress in CAD-Based Computer Vision: An Introduction to the Special Issue
Octavia I. Camps, Patrick J. Flynn, George C. Stockman
Comput. Vis. Image Underst.2
1996 New subband geometries for image texture segmentation
abstract
The usefulness of multichannel filtering for image texture segmentation has been demonstrated in the literature. Based on psychophysical studies, it has been largely believed that Gabor filters are optimal for segmentation. In this paper, the application of maximally decimated, non-separable, perfect reconstruction filter banks (subband systems) are investigated for texture segmentation. The subband decomposition used in this paper forms a maximally decimated filter bank and therefore computation time and complexity is much less than the full rate case of the Gabor filters. The directional selectivity achieved by the non-separable, multirate filter bank results in performance similar to that achievable with the Gabor filters, but with a drastically reduced computational load.
Dhiraj Kacker, Roberto H. Bamberger, Patrick J. Flynn
ICIP (3)3
1996 Realistic range rendering for object hypothesis verification
Patrick J. Flynn
Image Vis. Comput.1
1996 An Experimental Comparison of Range Image Segmentation Algorithms
abstract
A methodology for evaluating range image segmentation algorithms is proposed. This methodology involves (1) a common set of 40 laser range finder images and 40 structured light scanner images that have manually specified ground truth and (2) a set of defined performance metrics for instances of correctly segmented, missed, and noise regions, over- and under-segmentation, and accuracy of the recovered geometry. A tool is used to objectively compare a machine generated segmentation against the specified ground truth. Four research groups have contributed to evaluate their own algorithm for segmenting a range image into planar patches.
Adam W. Hoover, Gillian Jean-Baptiste, Xiaoyi Jiang 0001, Patrick J. Flynn, Horst Bunke, Dmitry B. Goldgof, Kevin W. Bowyer, David W. Eggert, Andrew W. Fitzgibbon, Robert B. Fisher
IEEE Trans. Pattern Anal. Mach. Intell.4
1996 Fractal Volume Compression
abstract
This research explores the principles, implementation, and optimization of a competitive volume compression system based on fractal image compression. The extension of fractal image compression to volumetric data is trivial in theory. However, the simple addition of a dimension to existing fractal image compression algorithms results in infeasible compression times and noncompetitive volume compression results. This paper extends several fractal image compression enhancements to perform properly and efficiently on volumetric data, and introduces a new 3D edge classification scheme based on principal component analysis. Numerous experiments over the many parameters of fractal volume compression suggest aggressive settings of its system parameters. At this peak efficiency, fractal volume compression surpasses vector quantization and approaches within 1 dB PSNR of the discrete cosine transform. When compared to the DCT, fractal volume compression represents surfaces in volumes exceptionally well at high compression rates, and the artifacts of its compression error appear as noise instead of deceptive smoothing or distracting ringing.
Wayne O. Cochran, John C. Hart, Patrick J. Flynn
IEEE Trans. Vis. Comput. Graph.3
1995 Recent Progress in CAD-Based Vision
Katsushi Ikeuchi, Patrick J. Flynn
Comput. Vis. Image Underst.2
1995 Integration of Multiple Feature Groups and Multiple Views into a 3D Object Recognition System
Jianchang Mao, Patrick J. Flynn, Anil K. Jain 0001
Comput. Vis. Image Underst.2
1994 Realistic range rendering
abstract
In many model-based object recognition systems, a synthesize-and-verify technique is used to evaluate the quality of hypotheses. This technique synthesises images of hypothesized objects in hypothesized poses, and compares them against the input imagery, producing a matching score. In this paper, we examine the image synthesis process in the context of triangulation-based range finding. We motivate the use of synthetically shadowed range data, for verification, present a simple and efficient algorithm for generation of shadowed range imagery, and demonstrate its usefulness in a set of real imagery.>
Patrick J. Flynn
CVPR1
1994 A Robust System for Lineament Analysis of Aero-magnetic Imagery using Orientation Analysis and Edge Linking
abstract
This paper presents a robust system for the detection, enhancement, and symbolic description of linear features in imagery. The system consists of two separate sub-systems: an orientation analysis engine based on the steerable filters, and an edge linking scheme which utilizes local orientation information.>
Jianxin Hou, Roberto H. Bamberger, Patrick J. Flynn
ICIP (1)3
1994 Guaranteed geometric hashing
abstract
Geometric hashing is an invariant feature-driven approach to model-based object recognition. Previous interest has focused on its ability to accommodate sensor error. This paper presents an enhancement of the geometric hashing technique which guarantees, under only a few constraints, that models will not be missed due to sensor noise. The authors' geometric hashing algorithm enters model affine invariants into hash table regions defined by an exact error model, brings together known optimizations (table symmetry and the use of more than 3 model-scene point correspondences) and uses novel data organization. Experimental results (on both synthetic and real data) suggest that the authors' modifications to a geometric hashing recognition scheme effectively overcome sensor noise.
Matthew P. Howell, Patrick J. Flynn
ICPR (1)2
1994 3-D Object Recognition with Symmetric Models: Symmetry Extraction and Encoding
abstract
Object recognition systems which employ solid models and range data have been a topic of interest for several years. Model databases have the potential to become large in some environments. This paper proposes a pair of techniques for incorporating knowledge of the symmetries of object models into the recognition process. The effects of symmetric models on the speed of an object recognition system is examined in the context of an implemented system employing invariant feature indexing as a correspondence-building mechanism. Groups of model surfaces are enumerated and examined to yield a list of segment label permutations which summarize the model's symmetry. This symmetry extraction process is followed by a symmetry encoding procedure which replaces groups of features which are indistinguishable because of symmetry with a single prototype feature group. Experiments with a large model database demonstrate the utility of these symmetry extraction and encoding techniques.>
Patrick J. Flynn
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Ground state texture patterns for the second-order Ising model
abstract
Ground state (energy-minimizing) texture patterns in Markov random field (MRF) image models, in particular, the second-order Ising model, are addressed. Ground state texture patterns are obtained for arbitrary parameter values. A small number of texture classes can be generated from the binary Ising model if the global minimum field energy criterion is used to terminate the sampler. Implications for texture synthesis are discussed.>
Hongjiu Lu, Patrick J. Flynn
CVPR2
1993 Parallel hypothesis verification
John Moody, Patrick J. Flynn, David L. Cohn
Pattern Recognit.2
1992 Saliencies and symmetries: toward 3D object recognition from large model databases
abstract
The construction of interpretation tables from database models is introduced, and a recognition procedure using scene feature groups is discussed. Techniques for extraction of feature group equivalence classes and computation of feature group saliency are discussed. Two methods to reduce the computational burdens associated with a large model database are proposed and tested on polyhedral objects. The first method reduces the population of protohypotheses in the interpretation tables consulted during recognition by excluding redundant feature groups produced from object symmetries. The second method assigns a population-based numerical measure of saliency to each feature group retrieved from the scene; this measure allows only the most salient feature groups to be used in object recognition.>
Patrick J. Flynn
CVPR1
1992 Parallel hypothesis verification
abstract
The verification of identifying and pose hypotheses in model-based 3D object recognition systems can involve a time-consuming image rendering operation followed by pixel-level comparison of the input and rendered images. In situations where many such hypotheses need to be verified, exploitation of inherent data-parallelism between hypotheses and their corresponding object models can increase the efficiency of the object recognition system. The authors describe a prototype system for distribution of hypotheses and the accompanying rendering tasks to individual processors in a loosely-coupled computing environment, and demonstrate excellent performance improvements over a single-processor implementation.>
John Moody, Patrick J. Flynn, David L. Cohn
ICPR (4)2
1992 3D object recognition using invariant feature indexing of interpretation tables
Patrick J. Flynn, Anil K. Jain 0001
CVGIP Image Underst.1
1991 CAD-Based Computer Vision: From CAD Models to Relational Graphs
abstract
The topic of model-building for 3-D objects is examined. Most 3-D object recognition systems construct models either manually or by training. Neither approach has been very satisfactory, particularly in designing object recognition systems which can handle a large number of objects. Recent interest in integrating mechanical CAD systems and vision systems has led to a third type of model building for vision: adaptation of preexisting CAD models of objects for recognition. If a solid model of an object to be recognized is already available in a manufacturing database, then it should be possible to infer automatically a model appropriate for vision tasks from the manufacturing model. Such a system has been developed. It uses 3-D object descriptions created on a commercial CAD system and expressed in both the industry-standard IGES form and a polyhedral approximation and performs geometric inferencing to obtain a relational graph representation of the object which can be stored in a database of models for object recognition. Relational graph models contain both view-independent information extracted from the IGES description and view-dependent information (patch areas) extracted from synthetic views of the object. It is argued that such a system is needed to efficiently create a large database (more than 100 objects) of 3-D models to evaluate matching strategies.>
Patrick J. Flynn, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1991 BONSAI: 3D Object Recognition Using Constrained Search
abstract
BONSAI, a model-based 3D object recognition system, is described. It identifies and localizes 3D objects in range images of one or more parts that have been designed on a computer-aided-design (CAD) system. Recognition is performed via constrained search of the interpretation tree, using unary and binary constraints (derived automatically from the CAD models) to prune the search space. Attention is focused on the recognition procedure, but the model-building, image acquisition, and segmentation procedures are also outlined. Experiments with over 200 images demonstrate that the constrained search approach to 3D object recognition has an accuracy comparable to that of previous systems.>
Patrick J. Flynn, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 BONSAI: 3D object recognition using constrained search
abstract
A description is presented of BONSAI, a model-based 3-D object recognition system, which identifies and localizes 3-D objects in range images of one or more parts which have been designed on a CAD system. Recognition is performed via constrained search of the interpretation tree, using unary and binary constraints (derived automatically from the CAD models) to prune the search space. Experiments with over 200 images of 20 different parts demonstrate that the constrained search approach to 3-D object recognition has comparable accuracy to other existing systems.>
Patrick J. Flynn, Anil K. Jain 0001
ICCV1
1989 On reliable curvature estimation
abstract
An empirical study of the accuracy of five different curvature estimation techniques, using synthetic range images and images obtained from three range sensors, is presented. The results obtained highlight the problems inherent in accurate estimation of curvatures, which are second-order quantities, and thus highly sensitive to noise contamination. The numerical curvature estimation methods are found to perform about as accurately as the analytic techniques, although ensemble estimates of overall surface curvature such as averages are unreliable unless trimmed estimates are used. The median proved to be the best estimator of location. As an exception, it is shown theoretically that zero curvature can be fairly reliably detected, with appropriate selection of threshold values.>
Patrick J. Flynn, Anil K. Jain 0001
CVPR1
1989 CAD-based computer vision: from CAD models to relational graphs
abstract
The authors outline their approach for automatic translation of geometric entities produced by a CAD system into a relational graph structure. They have developed a system which uses 3-D object descriptions created on a commercial CAD system and expressed in the industry-standard IGES form, and performs geometric inferencing to object a relational graph representation of the object which can be stored in a database of models of object recognition. Details of the IGES standard, the geometric inference engine, and some formal properties of 3-D models are discussed. In addition to the process of translation from one data format to another, the interference engine extracts higher-level information from the CAD model and stores it explicitly in the new data structure. The higher-level features will allow the search space explored during the object recognition stage to be pruned early.>
Patrick J. Flynn, Anil K. Jain 0001
SMC1
1989 Segmentation of document images
abstract
Several methods for segmentation of document images are explored. The authors pose the segmentation operation as a statistical classification task with two pattern classes: print and background. A number of classification strategies are available. All require some prior information about the distribution of gray levels for the two classes. Learning (either supervised or unsupervised) and automatic updating of the class-conditional densities are performed within image subregions to adapt global density estimates to the local area. After local densities have been obtained, each pixel within the window is classified; several techniques for this are considered. Results on four test images indicate that the commonly used contextual models are not suitable to all document images.>
Torfinn Taxt, Patrick J. Flynn, Anil K. Jain 0001
SMC2
1989 Segmentation of Document Images
abstract
Several methods for segmentation of document images (maps, drawings, etc.) are explored. The segmentation operation is posed as a statistical classification task with two pattern classes: print and background. A number of classification strategies are available. All require some prior information about the distribution of gray levels for the two classes. Training (either supervised or unsupervised) is employed to form these initial density estimates. Automatic updating of the class-conditional densities is performed within subregions in the image to adapt these global density estimates to the local image area. After local class-conditional densities have been obtained, each pixel is classified within the window using several techniques: a noncontextual Bayes classifier, Besag's classifier, relaxation, Owen and Switzer's classifier, and Haslett's classifier. Four test images were processed. In two of these, the relaxation method performed best, and in the other two, the noncontextual method performed best. Automatic updating improved the results for both classifiers.>
Torfinn Taxt, Patrick J. Flynn, Anil K. Jain 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
1988 Surface classification: hypothesis testing and parameter estimation
abstract
A 3-D surface classification method based on the quadric surface model is described. This technique does not require the points from the surface to lie on a grid. A sample of surface points is classified as planar or nonplanar through two hypothesis tests. If the sample is nonplanar, curvature features are evaluated at each point to classify the sample as spherical, cylindrical, or conical. A nonlinear optimization technique is then used to refine the parameters (e.g. radius, orientation) of the resulting surface type.>
Patrick J. Flynn, Anil K. Jain 0001
CVPR1