EDBT 2026 Demo / reviewers in the wild / expert
Alice J. O'Toole
dblp:35/3887
· DBLP profile ↗
32ranked-venue papers
13as first author
11since 2021 · last 2025
0000-0001-7981-1508ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 11 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dissecting Human Body Representations in Deep Networks Trained for Person IdentificationabstractLong-term body identification algorithms have emerged recently with the increased availability of high-quality training data. We seek to fill knowledge gaps about these models by analyzing body image embeddings from four body identification networks trained with 1.9 million images across 4,788 identities and 9 databases. By analyzing a diverse range of architectures (ViT, SWIN-ViT, CNN, and linguistically primed CNN), we first show that the face contributes to the accuracy of body identification algorithms and that these algorithms can identify faces to some extent-with no explicit face training. Second, we show that representations (embeddings) generated by body identification algorithms encode information about gender, as well as image-based information including view (yaw) and even the dataset from which the image originated. Third, we demonstrate that identification accuracy can be improved without additional training by operating directly and selectively on the learned embedding space. Leveraging principal component analysis (PCA), identity comparisons were consistently more accurate in subspaces that eliminated dimensions that explained large amounts of variance. These three findings were surprisingly consistent across architectures and test datasets. This work represents the first analysis of body representations produced by long-term re-identification networks trained on challenging unconstrained datasets. Thomas M. Metz, Matthew Q. Hill, Blake A. Myers, Veda Nandan Gandi, Rahul Chilakapati, Alice J. O'Toole |
FG | 6 |
| 2025 | Unconstrained Body Recognition at Altitude and Range: Comparing Four ApproachesabstractThis study presents an investigation of four distinct approaches to long-term person identification using body shape. Unlike short-term re-identification systems that rely on temporary features (e.g., clothing), we focus on learning persistent body shape characteristics that remain stable over time. We introduce a body identification model based on a Vision Transformer (ViT) (Body Identification from Diverse Datasets, BIDDS) and on a Swin-ViT model (Swin-BIDDS). We also expand on previous approaches [26] based on the Linguistic and Non-linguistic Core ResNet Identity Models (LCRIM and NLCRIM), but with improved training. All models are trained on a large and diverse dataset of over 1.9 million images of approximately 5k identities across 9 databases. Performance was evaluated on standard re-identification benchmark datasets (MARS [42], MSMT17 [34], Outdoor Gait [31], DeepChange [35]) and on an unconstrained dataset [3] that includes images at a distance (from close-range to $\mathbf{1 0 0 0 m}$), at altitude (from an unmanned aerial vehicle, UAV), and with clothing change. A comparative analysis across these models provides insights into how different backbone architectures and input image sizes impact long-term body identification performance across realworld conditions. Blake A. Myers, Matthew Q. Hill, Veda Nandan Gandi, Thomas M. Metz, Alice J. O'Toole |
FG | 5 |
| 2024 | GAMMA-FACE: GAussian Mixture Models Amend Diffusion Models for Bias Mitigation in Face Images
Basudha Pal, Arunkumar Kannan, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa |
ECCV (67) | 4 |
| 2024 | Designing Cross-Race Tests for Forensic Facial Examiners, Super-recognizers, and Face Recognition AlgorithmsabstractHumans and machines vary in the accuracy with which they recognize faces of different races. This can impact the fairness of face identification in security and forensic settings. We introduce a protocol for designing a cross-race face identification test for evaluating people (e.g., forensic facial examiners, super-recognizers) and machines with superior face-identification ability. We followed this protocol to create a cross-race test and report the test's benchmarks on untrained human participants and two state-of-the-art face recognition algorithms. The goal of the protocol is to select a relatively small number of challenging test items (facial image comparisons) of two races, with approximately equally challenging items of both races. Item selection consisted of pre-screening with an open-source face recognition algorithm, followed by a second round of prescreening using the performance of untrained human participants. We sampled face-images (Black and White identities) from a large biometric data set and applied the protocol to assemble face comparisons. The protocol yielded a cross-race test with 20 comparison pairs portraying Black and White identities (10 same-identity; 10 different-identity). Untrained participants (54 Black; 51 White) judged whether face-image pairs showed the same or different identities using a 7-point scale. By design, the test proved challenging for untrained participants, with performance comparable across Black and White image pairs for both Black and White participants. Two top-performing face recognition systems from the Face Recognition Vendor Test-ongoing [6] scored perfectly (no errors) on both Black and White face-image pairs from the Cross-Race Test. The human and machine benchmarks established here make this test ideal for evaluating cross-race face recognition bias in people with high levels of skill and training. Géraldine Jeckeln, Selin Yavuzcan, Kate A. Marquis, Prajay Sandipkumar Mehta, Amy N. Yates, P. Jonathon Phillips, Alice J. O'Toole |
FG | 7 |
| 2024 | DiversiNet: Mitigating Bias in Deep Classification Networks across Sensitive Attributes through Diffusion-Generated DataabstractDeep learning models trained on sensitive data often show biases towards certain demographics, posing fairness challenges, especially with limited datasets. Diffusion generated data effectively supplement the underrepresented dataset, serving as a regularization technique to enhance feature learning. In addition to the original balanced dataset, we incorporate synthetic data generated by the diffusion model to train classifiers and subsequently assess their performance. Experimental results demonstrate a reduction in bias across all target attributes along with an increase in overall accuracy. For instance, for gender classification in the FFHQ dataset, the overall accuracy rises to 94.44% from 93.92% after including data generated from a diffusion model. Simultaneously, the bias, measured as the absolute difference between the true positive rates of young and old individuals, decreases from 0.0340 to 0.0204 (reduction of 40%). Moreover, we extend our analysis to multi-attribute scenarios, successfully mitigating bias with respect to multiple sensitive attributes simultaneously in sensitive attribute classification as well as in other downstream tasks. To the best of our knowledge, this study introduces a novel approach to bias mitigation, highlighting the versatility of diffusion-based data augmentation in addressing biases concerning age, gender, and race. Basudha Pal, Aniket Roy, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa |
IJCB | 4 |
| 2024 | Combining Images and Words in Deep Networks that Identify People from Body Shape
Alice J. O'Toole |
ICPRAM | 1 |
| 2024 | The Influence of the Other-Race Effect on Susceptibility to Face Morphing AttacksabstractFacial morphs created between two identities resemble both of the faces used to create the morph. Consequently, humans and machines are prone to mistake morphs made from two identities for either of the faces used to create the morph. This vulnerability has been exploited in "morph attacks" in security scenarios. Here, we asked whether the "other-race effect" (ORE)-the human advantage for identifying own- vs. other-race faces-exacerbates morph attack susceptibility for humans. We also asked whether face-identification performance in a deep convolutional neural network (DCNN) is affected by the race of morphed faces. Caucasian (CA) and East-Asian (EA) participants performed a face-identity matching task on pairs of CA and EA face images in two conditions. In the morph condition, different-identity pairs consisted of an image of identity "A" and a 50/50 morph between images of identity "A" and "B". In the baseline condition, morphs of different identities never appeared. As expected, morphs were identified mistakenly more often than original face images. Of primary interest, morph identification was substantially worse for cross-race faces than for own-race faces. Similar to humans, the DCNN performed more accurately for original face images than for morphed image pairs. Notably, the deep network proved substantially more accurate than humans in both cases. The results point to the possibility that DCNNs might be useful for improving face identification accuracy when morphed faces are presented. They also indicate the significance of the race of a face in morph attack susceptibility in applied settings. Snipta Mallick, Géraldine Jeckeln, Connor J. Parde, Carlos Domingo Castillo, Alice J. O'Toole |
ACM Trans. Appl. Percept. | 5 |
| 2023 | Recognizing People by Body Shape Using Deep Networks of Images and WordsabstractCommon and important applications of person identification occur at distances and viewpoints in which the face is not visible or is not sufficiently resolved to be useful. We examine body shape as a biometric across distance and viewpoint variation. We propose an approach that combines standard object classification networks with representations based on linguistic (word-based) descriptions of bodies. Algorithms with and without linguistic training were compared on their ability to identify people from body shape in images captured across a large range of distances/views (close-range, 100m, 200m, 270m, 300m, 370m, 400m, 490m, 500m, 600m, and at elevated pitch in images taken by an unmanned aerial vehicle [UAV]). Accuracy, as measured by identity-match ranking and false accept errors in an open-set test, was surprisingly good. For identity-ranking, linguistic models were more accurate for close-range images, whereas non-linguistic models fared better at intermediary distances. Fusion of the linguistic and non-linguistic embeddings improved performance at all but the farthest distance. Although the non-linguistic model yielded fewer false accepts at most distances, fusion of the linguistic and non-linguistic models decreased false accepts for all but the UAV images. We conclude that linguistic and non-linguistic representations of body shape can offer complementary identity information for bodies that can improve identification in applications of interest. Blake A. Myers, Lucas Jaggernauth, Thomas M. Metz, Matthew Q. Hill, Veda Nandan Gandi, Carlos Domingo Castillo, Alice J. O'Toole |
IJCB | 7 |
| 2023 | Twin Identification over Viewpoint Change: A Deep Convolutional Neural Network Surpasses HumansabstractDeep convolutional neural networks (DCNNs) have achieved human-level accuracy in face identification (Phillips et al., 2018), though it is unclear how accurately they discriminate highly-similar faces. Here, humans and a DCNN performed a challenging face-identity matching task that included identical twins. Participants ( N = 87) viewed pairs of face images of three types: same-identity, general imposters (different identities from similar demographic groups), and twin imposters (identical twin siblings). The task was to determine whether the pairs showed the same person or different people. Identity comparisons were tested in three viewpoint-disparity conditions: frontal to frontal, frontal to 45° profile, and frontal to 90° profile. Accuracy for discriminating matched-identity pairs from twin-imposter pairs and general-imposter pairs was assessed in each viewpoint-disparity condition. Humans were more accurate for general-imposter pairs than twin-imposter pairs, and accuracy declined with increased viewpoint disparity between the images in a pair. A DCNN trained for face identification (Ranjan et al., 2018) was tested on the same image pairs presented to humans. Machine performance mirrored the pattern of human accuracy, but with performance at or above all humans in all but one condition. Human and machine similarity scores were compared across all image-pair types. This item-level analysis showed that human and machine similarity ratings correlated significantly in six of nine image-pair types [range r = 0.38 to r = 0.63], suggesting general accord between the perception of face similarity by humans and the DCNN. These findings also contribute to our understanding of DCNN performance for discriminating high-resemblance faces, demonstrate that the DCNN performs at a level at or above humans, and suggest a degree of parity between the features used by humans and the DCNN. Connor J. Parde, Virginia E. Strehle, Vivekjyoti Banerjee, Jacqueline G. Cavazos, Carlos Domingo Castillo, Alice J. O'Toole |
ACM Trans. Appl. Percept. | 7 |
| 2021 | A Study of the Human Perception of Synthetic FacesabstractAdvances in face synthesis have raised alarms about the deceptive use of synthetic faces. Can synthetic identities be effectively used to fool human observers? In this paper, we introduce a study of the human perception of synthetic faces generated using different strategies including a state-of-the-art deep learning-based GAN model. This is the first rigorous study of the effectiveness of synthetic face generation techniques grounded in experimental techniques from psychology. We answer important questions such as how often do GAN-based and more traditional image processing-based techniques confuse human observers, and are there subtle cues within a synthetic face image that cause humans to perceive it as a fake without having to search for obvious clues? To answer these questions, we conducted a series of large-scale crowdsourced behavioral experiments with different sources of face imagery. Results show that humans are unable to distinguish synthetic faces from real faces under several different circumstances. This finding has serious implications for many different applications where face images are presented to human users. Bingyu Shen 0001, Brandon RichardWebster, Alice J. O'Toole, Kevin W. Bowyer, Walter J. Scheirer |
FG | 3 |
| 2021 | Evaluating Automated Face Identity-Masking Methods with Human Perception and a Deep Convolutional Neural NetworkabstractFace de-identification (or “masking”) algorithms have been developed in response to the prevalent use of video recordings in public places. We evaluated the success of face identity masking for human perceivers and a deep convolutional neural network (DCNN). Eight de-identification algorithms were applied to videos of drivers’ faces, while they actively operated a motor vehicle. These masks were pre-selected to be applicable to low-quality video and to maintain coarse information about facial actions. Humans studied high-resolution images to learn driver identities and were tested on their recognition of active drivers in low-resolution videos. Faces in the videos were either unmasked or were masked by one of the eight algorithms. When participants were tested immediately after learning (Experiment 1), all masks reduced identification, with six of eight masks reducing identification to extremely poor performance. In a second experiment, two of the most effective masks were tested after a delay of 7 or 28 days. The delay did not further reduce identification of the masked faces. In all masked conditions, participants maintained stringent decision criteria, with low confidence in recognition, further indicating the effectiveness of the masks. Next, the DCNN performed an identity-matching task between high-resolution images and masked videos—a task analogous to that done by humans. The pattern of accuracy for the DCNN mirrored some, but not all, aspects of human performance, highlighting the need to test the effectiveness of identity masking for both humans and machines. The DCNN was also tested on its ability to match identity between masked and unmasked versions of the same video, based only on the face. DCNN performance for the eight masks offers insight into the nature of the information in faces that is coded in these networks. Kimberley D. Orsten-Hooge, Asal Baragchizadeh, Thomas P. Karnowski, David S. Bolme, Regina K. Ferrell, Parisa R. Jesudasen, Carlos Domingo Castillo, Alice J. O'Toole |
ACM Trans. Appl. Percept. | 8 |
| 2017 | Evaluation of Automated Identity Masking Method (AIM) in Naturalistic Driving Study (NDS)abstractIdentity masking methods have been developed in recent years for use in multiple applications aimed at protecting privacy. There is only limited work, however, targeted at evaluating effectiveness of methods-with only a handful of studies testing identity masking effectiveness for human perceivers. Here, we employed human participants to evaluate identity masking algorithms on video data of drivers, which contains subtle movements of the face and head. We evaluated the effectiveness of the “personalized supervised bilinear regression method for Facial Action Transfer (FAT)” de-identification algorithm. We also evaluated an edge-detection filter, as an alternate “fill-in” method when face tracking failed due to abrupt or fast head motions. Our primary goal was to develop methods for humanbased evaluation of the effectiveness of identity masking. To this end, we designed and conducted two experiments to address the effectiveness of masking in preventing recognition and in preserving action perception. 1- How effective is an identity masking algorithm?We conducted a face recognition experiment and employed Signal Detection Theory (SDT) to measure human accuracy and decision bias. The accuracy results show that both masks (FAT mask and edgedetection) are effective, but that neither completely eliminated recognition. However, the decision bias data suggest that both masks altered the participants' response strategy and made them less likely to affirm identity. 2- How effectively does the algorithm preserve actions? We conducted two experiments on facial behavior annotation. Results showed that masking had a negative effect on annotation accuracy for the majority of actions, with differences across action types. Notably, the FAT mask preserved actions better than the edge-detection mask. To our knowledge, this is the first study to evaluate a deidentification method aimed at preserving facial actions employing human evaluators in a laboratory setting. Asal Baragchizadeh, Thomas P. Karnowski, David S. Bolme, Alice J. O'Toole |
FG | 4 |
| 2017 | Five Principles for Crowd-Source Experiments in Face RecognitionabstractThe past few years have seen impressive gains in long standing and difficult problems in face recognition. These gains have come about through the use of deep learning algorithms that consist of multi-layered neural networks. In part,the success of these algorithms is due to the easy availability of extremely large datasets of faces that are annotated and labelled by humans. The reliance on crowd-sourced data for machine learning and algorithm evaluation raises methodological issues that are not widely appreciated in computer vision. Several of these issues have come to light in recent work using crowd sourcing to benchmark human face identification on large databases that are used to test face recognition algorithms. Wedefine and discuss these issues using face recognition as a case study. We focus on: a.) the characteristics of the human participants;b.) the difference between aggregate and fused measures of human accuracy; and c.) the lack of standard methods for controlling critical characteristics of the “imposter” distribution in large and variably diverse data sets. We will show that estimates of human accuracy can vary widely depending on howthese factors combine in any given evaluation.We conclude with recommendations on best practices in mitigating this variability and arriving at stable estimates of ground truth acquired by crowd-sourcing. Alice J. O'Toole, P. Jonathon Phillips |
FG | 1 |
| 2017 | Face and Image Representation in Deep CNN FeaturesabstractFace recognition algorithms based on deep convolutional neural networks (DCNNs) have made progress on the task of recognizing faces in unconstrained viewing conditions. These networks operate with compact feature-based face representations derived from learning a very large number of face images. Although the learned feature sets produced by DCNNs can be highly robust to changes in viewpoint, illumination, and appearance, little is known about the nature of the face code that emerges at the top level of these networks. We analyzed the DCNN features produced by two recent face recognition algorithms. In the first set of experiments, we used the top-level features from the DCNNs as input into linear classifiers aimed at predicting metadata about the images. The results showed that the DCNN features contained surprisingly accurate information about the yaw and pitch of a face, and about whether the input face came from a still image or a video frame. In the second set of experiments, we measured the extent to which individual DCNN features operated in a view-dependent or view-invariant manner for different identities. We found that view-dependent coding was a characteristic of the identities rather than the DCNN features– with some identities coded consistently in a view-dependent way and others in a view-independent way. In our third analysis, we visualized the DCNN feature space for 24,000+ images of 500 identities. Images in the center of the space were uniformly of low quality (e.g., extreme views, face occlusion, poor contrast, low resolution). Image quality increased monotonically as a function of distance from the origin. This result suggests that image quality information is available in the DCNN features, such that consistently average feature values reflect coding failures that reliably indicate poor or unusable images. Combined, the results offer insight into the coding mechanisms that support robust representation of faces in DCNNs. Connor J. Parde, Carlos Domingo Castillo, Matthew Q. Hill, Y. Ivette Colon, Swami Sankaranarayanan, Jun-Cheng Chen, Alice J. O'Toole |
FG | 7 |
| 2016 | Body talk: crowdshaping realistic 3D avatars with wordsabstractRealistic, metrically accurate, 3D human avatars are useful for games, shopping, virtual reality, and health applications. Such avatars are not in wide use because solutions for creating them from high-end scanners, low-cost range cameras, and tailoring measurements all have limitations. Here we propose a simple solution and show that it is surprisingly accurate. We use crowdsourcing to generate attribute ratings of 3D body shapes corresponding to standard linguistic descriptions of 3D shape. We then learn a linear function relating these ratings to 3D human shape parameters. Given an image of a new body, we again turn to the crowd for ratings of the body shape. The collection of linguistic ratings of a photograph provides remarkably strong constraints on the metric 3D shape. We call the process crowdshaping and show that our Body Talk system produces shapes that are perceptually indistinguishable from bodies created from high-resolution scans and that the metric accuracy is sufficient for many tasks. This makes body "scanning" practical without a scanner, opening up new applications including database search, visualization, and extracting avatars from books. Stephan Streuber, Maria Alejandra Quiros-Ramirez, Matthew Q. Hill, Carina A. Hahn, Silvia Zuffi, Alice J. O'Toole, Michael J. Black |
ACM Trans. Graph. | 6 |
| 2014 | Comparison of human and computer performance across face recognition experiments
P. Jonathon Phillips, Alice J. O'Toole |
Image Vis. Comput. | 2 |
| 2012 | Demographic effects on estimates of automatic face recognition performance
Alice J. O'Toole, P. Jonathon Phillips, Xiaobo An, Joseph P. Dunlop |
Image Vis. Comput. | 1 |
| 2012 | The Good, the Bad, and the Ugly Face Challenge Problem
P. Jonathon Phillips, J. Ross Beveridge, Bruce A. Draper, Geof H. Givens, Alice J. O'Toole, David S. Bolme, Joseph P. Dunlop, Yui Man Lui, Hassan Sahibzada, Samuel Weimer |
Image Vis. Comput. | 5 |
| 2012 | Comparing face recognition algorithms to humans on challenging tasksabstractWe compared face identification by humans and machines using images taken under a variety of uncontrolled illumination conditions in both indoor and outdoor settings. Natural variations in a person's day-to-day appearance (e.g., hair style, facial expression, hats, glasses, etc.) contributed to the difficulty of the task. Both humans and machines matched the identity of people (same or different) in pairs of frontal view face images. The degree of difficulty introduced by photometric and appearance-based variability was estimated using a face recognition algorithm created by fusing three top-performing algorithms from a recent international competition. The algorithm computed similarity scores for a constant set of same-identity and different-identity pairings from multiple images. Image pairs were assigned to good , moderate , and poor accuracy groups by ranking the similarity scores for each identity pairing, and dividing these rankings into three strata. This procedure isolated the role of photometric variables from the effects of the distinctiveness of particular identities. Algorithm performance for these constant identity pairings varied dramatically across the groups. In a series of experiments, humans matched image pairs from the good, moderate, and poor conditions, rating the likelihood that the images were of the same person (1: sure same - 5: sure different). Algorithms were more accurate than humans in the good and moderate conditions, but were comparable to humans in the poor accuracy condition. To date, these are the most variable illumination- and appearance-based recognition conditions on which humans and machines have been compared. The finding that machines were never less accurate than humans on these challenging frontal images suggests that face recognition systems may be ready for applications with comparable difficulty. We speculate that the superiority of algorithms over humans in the less challenging conditions may be due to the algorithms' use of detailed, view-specific identity information. Humans may consider this information less important due to its limited potential for robust generalization in suboptimal viewing conditions. Alice J. O'Toole, Xiaobo An, Joseph P. Dunlop, Vaidehi S. Natu, P. Jonathon Phillips |
ACM Trans. Appl. Percept. | 1 |
| 2011 | Demographic effects on estimates of automatic face recognition performanceabstractThe intended applications of automatic face recognition systems include venues that vary widely in demographic diversity. Formal evaluations of algorithms do not commonly consider the effects of population diversity on performance. We document the effects of racial and gender demographics on the accuracy of algorithms that match identity in pairs of face images. In particular, we focus on the effects of the “background” population distribution of non-matched identities against which identity matches are compared. The algorithm we tested was created by fusing three of the top performers from a recent US Government competition. First, we demonstrate the variability of algorithm performance estimates when the non-matched identities were demographically “yoked” by race and/or gender (i.e., “yoking” constrains non-matched pairs to be of the same race or gender). We also found a shift in the match threshold required to maintain a stable false positive rate when demographic control scenarios varied. These results were verified with two independent data sets that differed in demographic characteristics. In a second experiment, we explored the effects of progressive increases in population diversity on algorithm performance. We found systematic, but non-general, effects when the balance between majority and minority populations of non-matched identities shifted. Finally, we show that identity match accuracy differs substantially when the non-match identity population varied by race. The results indicate the importance of the demographic composition and modeling of the background population in predicting the accuracy of face recognition algorithms. Alice J. O'Toole, Xiaobo An, P. Jonathon Phillips, Joseph P. Dunlop |
FG | 1 |
| 2011 | An introduction to the good, the bad, & the ugly face recognition challenge problemabstractThe Good, the Bad, & the Ugly Face Challenge Problem was created to encourage the development of algorithms that are robust to recognition across changes that occur in still frontal faces. The Good, the Bad, & the Ugly consists of three partitions. The Good partition contains pairs of images that are considered easy to recognize. On the Good partition, the base verification rate (VR) is 0.98 at a false accept rate (FAR) of 0.001. The Bad partition contains pairs of images of average difficulty to recognize. For the Bad partition, the VR is 0.80 at a FAR of 0.001. The Ugly partition contains pairs of images considered difficult to recognize, with a VR of 0.15 at a FAR of 0.001. The base performance is from fusing the output of three of the top performers in the FRVT 2006. The design of the Good, the Bad, & the Ugly controls for pose variation, subject aging, and subject “recognizability.” Subject recognizability is controlled by having the same number of images of each subject in every partition. This implies that the differences in performance among the partitions are result of how a face is presented in each image. P. Jonathon Phillips, J. Ross Beveridge, Bruce A. Draper, Geof H. Givens, Alice J. O'Toole, David S. Bolme, Joseph P. Dunlop, Yui Man Lui, Hassan Sahibzada, Samuel Weimer |
FG | 5 |
| 2011 | An other-race effect for face recognition algorithmsabstractPsychological research indicates that humans recognize faces of their own race more accurately than faces of other races. This “other-race effect” occurs for algorithms tested in a recent international competition for state-of-the-art face recognition algorithms. We report results for a Western algorithm made by fusing eight algorithms from Western countries and an East Asian algorithm made by fusing five algorithms from East Asian countries. At the low false accept rates required for most security applications, the Western algorithm recognized Caucasian faces more accurately than East Asian faces and the East Asian algorithm recognized East Asian faces more accurately than Caucasian faces. Next, using a test that spanned all false alarm rates, we compared the algorithms with humans of Caucasian and East Asian descent matching face identity in an identical stimulus set. In this case, both algorithms performed better on the Caucasian faces—the “majority” race in the database. The Caucasian face advantage, however, was far larger for the Western algorithm than for the East Asian algorithm. Humans showed the standard other-race effect for these faces, but showed more stable performance than the algorithms over changes in the race of the test faces. State-of-the-art face recognition algorithms, like humans, struggle with “other-race face” recognition. P. Jonathon Phillips, Abhijit Narvekar, Julianne H. Ayyad, Alice J. O'Toole |
ACM Trans. Appl. Percept. | 5 |
| 2010 | FRVT 2006 and ICE 2006 Large-Scale Experimental ResultsabstractThis paper describes the large-scale experimental results from the Face Recognition Vendor Test (FRVT) 2006 and the Iris Challenge Evaluation (ICE) 2006. The FRVT 2006 looked at recognition from high-resolution still frontal face images and 3D face images, and measured performance for still frontal face images taken under controlled and uncontrolled illumination. The ICE 2006 evaluation reported verification performance for both left and right irises. The images in the ICE 2006 intentionally represent a broader range of quality than the ICE 2006 sensor would normally acquire. This includes images that did not pass the quality control software embedded in the sensor. The FRVT 2006 results from controlled still and 3D images document at least an order-of-magnitude improvement in recognition performance over the FRVT 2002. The FRVT 2006 and the ICE 2006 compared recognition performance from high-resolution still frontal face images, 3D face images, and the single-iris images. On the FRVT 2006 and the ICE 2006 data sets, recognition performance was comparable for high-resolution frontal face, 3D face, and the iris images. In an experiment comparing human and algorithms on matching face identity across changes in illumination on frontal face images, the best performing algorithms were more accurate than humans on unfamiliar faces. P. Jonathon Phillips, W. Todd Scruggs, Alice J. O'Toole, Patrick J. Flynn, Kevin W. Bowyer, Cathy L. Schott, Matthew Sharpe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Humans versus algorithms: Comparisons from the Face Recognition Vendor Test 2006abstractWe present a synopsis of results comparing the performance of humans with face recognition algorithms tested in the face recognition vendor test (FRVT) 2006 and face recognition grand challenge (FRGC). Algorithms and humans matched face identity in images taken under controlled and uncontrolled illumination. The human-machine comparisons include accuracy benchmarks, an error pattern analysis, and a test of human and machine performance stability across data sets varying in image quality. The results indicate that: (1.) machines can compete quantitatively with humans matching face identity across changes in illumination; (2.) qualitative differences between humans and machines can be exploited to improve identification by fusing human and machine match scores; and (3.) recognition skills for humans and machines are comparably stable across changes in image quality. Combined the results suggest that face recognition algorithms may be ready for applications with task constraints similar to those evaluated in the FRVT 2006. Alice J. O'Toole, P. Jonathon Phillips, Abhijit Narvekar |
FG | 1 |
| 2007 | Face Recognition Algorithms Surpass Humans Matching Faces Over Changes in IlluminationabstractThere has been significant progress in improving the performance of computer-based face recognition algorithms over the last decade. Although algorithms have been tested and compared extensively with each other, there has been remarkably little work comparing the accuracy of computer-based face recognition systems with humans. We compared seven state-of-the-art face recognition algorithms with humans on a facematching task. Humans and algorithms determined whether pairs of face images, taken under different illumination conditions, were pictures of the same person or of different people. Three algorithms surpassed human performance matching face pairs prescreened to be "difficult" and six algorithms surpassed humans on "easy" face pairs. Although illumination variation continues to challenge face recognition algorithms, current algorithms compete favorably with humans. The superior performance of the best algorithms over humans, in light of the absolute performance levels of the algorithms, underscores the need to compare algorithms with the best current control--humans. Alice J. O'Toole, P. Jonathon Phillips, Janet H. Ayyad, Nils Penard, Hervé Abdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Fusing Face-Verification Algorithms and HumansabstractIt has been demonstrated recently that state-of-the-art face-recognition algorithms can surpass human accuracy at matching faces over changes in illumination. The ranking of algorithms and humans by accuracy, however, does not provide information about whether algorithms and humans perform the task comparably or whether algorithms and humans can be fused to improve performance. In this paper, we fused humans and algorithms using partial least square regression (PLSR). In the first experiment, we applied PLSR to face-pair similarity scores generated by seven algorithms participating in the Face Recognition Grand Challenge. The PLSR produced an optimal weighting of the similarity scores, which we tested for generality with a jackknife procedure. Fusing the algorithms' similarity scores using the optimal weights produced a twofold reduction of error rate over the most accurate algorithm. Next, human-subject-generated similarity scores were added to the PLSR analysis. Fusing humans and algorithms increased the performance to near-perfect classification accuracy. These results are discussed in terms of maximizing face-verification accuracy with hybrid systems consisting of multiple algorithms and humans. Alice J. O'Toole, Hervé Abdi, P. Jonathon Phillips |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2005 | A Video Database of Moving Faces and PeopleabstractWe describe a database of static images and video clips of human faces and people that is useful for testing algorithms for face and person recognition, head/eye tracking, and computer graphics modeling of natural human motions. For each person there are nine static "facial mug shots" and a series of video streams. The videos include a "moving facial mug shot," a facial speech clip, one or more dynamic facial expression clips, two gait videos, and a conversation video taken at a moderate distance from the camera. Complete data sets are available for 284 subjects and duplicate data sets, taken subsequent to the original set, are available for 229 subjects. Alice J. O'Toole, Joshua Harms, Sarah L. Snow, Dawn R. Hurst, Matthew R. Pappas, Janet H. Ayyad, Hervé Abdi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Face Recognition Algorithms as Models of Human Face ProcessingabstractWe evaluated the adequacy of computational algorithms as models of human face processing by looking at how the algorithms and humans process individual faces. By comparing model- and human-generated measures of the similarity between pairs of faces, we were able to assess the accord between several automatic face recognition algorithms and human perceivers. Multidimensional scaling (MDS) was used to create a spatial representation of the subject response patterns. Next, the model response patterns were projected into this space. The results revealed a common bimodal structure for both the subjects and for most of the models. The bimodal subject structure reflected strategy differences in making similarity decisions. For the models, the bimodal structure was related to combined aspects of the representations and the distance metrics used in the implementations. Alice J. O'Toole, Yi D. Cheng, Brendan Ross, Heather A. Wild, P. Jonathon Phillips |
FG | 1 |
| 1999 | 3D shape and 2D surface textures of human faces: the role of "averages" in attractiveness and age
Alice J. O'Toole, Theodore J. Price, Thomas Vetter, James C. Bartlett, Volker Blanz |
Image Vis. Comput. | 1 |
| 1996 | Face distinctiveness in recognition across viewpoint: An analysis of the statistical structure of face spacesabstractWe present an analysis of the effects of face distinctiveness on the performance of a computational model of recognition over viewpoint change. In the first stage of the model, the face stimulus is normalized by being mapped to an arbitrary standard view. In the second stage, the normalized stimulus is mapped into a "face space" spanned by a number of reference faces, and is classified as familiar or unfamiliar We carried out experiments employing a parametrically generated family of face stimuli that vary in distinctiveness. The experiments show that while the "view-mapping" process operates more accurately for typical versus distinctive faces, the base level distinctiveness of the faces is preserved in the face space coding. These data provide insight into how the psychophysically well-established inverse relationship between the typically and recognizability of faces might operate for recognition across changes in viewpoint. Alice J. O'Toole, Shimon Edelman |
FG | 1 |
| 1994 | Connectionist models of face processing: A survey
Dominique Valentin, Hervé Abdi, Alice J. O'Toole, Garrison W. Cottrell |
Pattern Recognit. | 3 |
| 1988 | A physical system approach to recognition memory for spatially transformed faces
Alice J. O'Toole, Richard B. Millward, James A. Anderson |
Neural Networks | 1 |