VLDB 2026 Research / reviewers in the wild / expert
Michael C. King
dblp:214/3280
· DBLP profile ↗
16ranked-venue papers
0as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Security and privacy · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hue Are You? Can Skin Depigmentation Affect Face Recognition Performance?abstractVitiligo causes localized loss of skin pigmentation, producing visible tone variations that may affect face recognition accuracy. While skin tone is a known covariate, the specific impact of vitiligo has not been studied. This work addresses two key questions: (1) does the presence of vitiligo present a challenge for face recognition systems, and (2) does the severity of vitiligo influence performance? First, we introduce Vitiligo Faces (VF), a new dataset of individuals with vitiligo, and compare matcher performance on VF against CFPW, a benchmark composed of difficult image pairs. Two of the three evaluated matchers perform worse on VF than CFPW, with greater score overlap and higher false non-match rates—indicating that vitiligo presents a unique and substantial challenge. Second, we develop a computational pipeline to quantify severity and find that increasing severity is associated with declining performance, as measured by d-prime (d’) separability. These findings underscore the need to evaluate recognition systems under underrepresented, clinically relevant conditions. Vitiligo introduces real-world variation that current models are not well equipped to handle, highlighting the importance of inclusive training and evaluation. Joyce Annan, Xavier Merino, Michael C. King |
FG | 3 |
| 2025 | The AgeDB-30M Dataset: Melanated Faces for Age-Invariant Face RecognitionabstractFor the task of evaluating face recognition algorithms, the research community has adopted a set of de facto standard datasets. These datasets tend to emphasize “difficult pairs” – paired images chosen for differences in factors like age (as in AgeDB-30 and CALFW) and pose (as in CPLFW and CFP-FP). Difficult pairs allow for a more robust evaluation of algorithms, offering granular insight into the factors that are problematic for specific algorithms. However, the existing datasets ignore one factor that has historically proven highly challenging for many face recognition algorithms: race. The faces in these datasets are overwhelmingly White.In this work, we address this demographic gap with the curation of an all-Black dataset for evaluation: AgeDB-30M, where “M” indicates “melanated”. It is the first publicly-available dataset of difficult cross-age image pairs solely from the Black demographic. We hope that AgeDB-30M is a valuable tool for the research community, supporting continued efforts toward more robust algorithmic evaluation, particularly with respect to issues of bias and fairness. Audison Beaubrun, Joyce Annan, Haiyu Wu, Xavier Merino, Kevin W. Bowyer, Michael C. King |
FG | 6 |
| 2025 | Deep CNN Face Matchers Inherently Support Revocable Biometric TemplatesabstractOne common critique of biometric authentication is that if an individual’s biometric is compromised, then the individual has no recourse. The concept of revocable biometrics was developed to address this concern. A biometric scheme is revocable if an individual can have their current enrollment in the scheme revoked, so that the compromised biometric template becomes worthless, and the individual can re-enroll with a new template that has similar recognition power. We show that modern deep CNN face matchers inherently allow for a robust revocable biometric scheme. For a given state-of-the-art deep CNN backbone and training set, it is possible to generate an unlimited number of distinct face matcher models that have both (1) equivalent recognition power, and (2) strongly incompatible biometric templates. The equivalent recognition power extends to the point of generating impostor and genuine distributions that have the same shape and placement on the similarity dimension, meaning that the models can share a similarity threshold for a 1-in-10,000 false match rate. The biometric templates from different model instances are so strongly incompatible that the cross-instance similarity score for images of the same person is typically lower than the sameinstance similarity score for images of different persons. That is, a stolen biometric template that is revoked is of less value in attempting to match the re-enrolled identity than the average impostor template. We also explore the feasibility of using a Vision Transformer (ViT) backbone-based face matcher in the revocable biometric system proposed in this work and demonstrate that it is less suitable compared to typical ResNet-based deep CNN backbones. Aman Bhatta, Michael C. King, Kevin W. Bowyer |
FG | 2 |
| 2025 | Tiny Faces, Big Trouble: Evaluating Super-Resolution for Face RecognitionabstractLow-resolution imagery presents a critical challenge for face recognition (FR), particularly in use cases such as law enforcement and surveillance, where real-world conditions are unconstrained. Despite the development of FR systems tailored for low-resolution input and the availability of super-resolution (SR) techniques, there is no evidence that such enhancements are used in operational deployments. This work evaluates the effectiveness of six SR methods in enhancing low-resolution face images prior to recognition. We simulate low-resolution probes at interpupillary distances (IPD) of 5-30px and upscale them using SR methods, while keeping gallery images fixed at high resolution (~100px IPD). Our analysis proceeds in two stages. First, we assess whether SR methods preserve image fidelity using standard image quality assessment (IQA) metrics and 1:1 “self-matching” scores. Second, we measure their impact on biometric performance by performing 1:1 and 1:N matching. Results show that although SR techniques improve perceptual quality, they do not fully recover identity-relevant features, especially at lower resolutions. These findings highlight the limitations of current SR methods in restoring biometric utility and underscore the need for resolution-aware FR pipelines in real-world applications. Xavier Merino, Gabriella Pangelinan, Samuel Langborgh, Michael C. King |
FG | 4 |
| 2025 | Peepers & Pixels: Human Recognition Accuracy on Low Resolution FacesabstractAutomated one-to-many ($1: \mathrm{N}$) face recognition is a powerful investigative tool commonly used by law enforcement agencies. In this context, potential matches resulting from automated 1:N recognition are reviewed by human examiners prior to possible use as investigative leads. While automated 1:N recognition can achieve near-perfect accuracy under ideal imaging conditions, operational scenarios may necessitate the use of surveillance imagery, which is often degraded in various quality dimensions. One important quality dimension is image resolution, typically quantified by the number of pixels on the face. The common metric for this is inter-pupillary distance (IPD), which measures the number of pixels between the pupils. Low IPD is known to degrade the accuracy of automated face recognition. However, the threshold IPD for reliability in human face recognition remains undefined. This study aims to explore the boundaries of human recognition accuracy by systematically testing accuracy across a range of IPD values. We find that at low IPDs ($10 \mathrm{px}, 5 \mathrm{px}$), human accuracy is at or below chance levels ($50.7 \%, 35.9 \%$), even as confidence in decision-making remains relatively high ($77 \%, 70.7 \%$). Our findings indicate that, for low IPD images, human recognition ability could be a limiting factor to overall system accuracy. Xavier Merino, Gabriella Pangelinan, Samuel Langborgh, Michael C. King, Kevin W. Bowyer |
FG | 4 |
| 2025 | Testing Peepers on Pixels: A Demo of Human Recognition Accuracy for Low Resolution FacesabstractHow well can humans recognize faces at extremely low resolution? We conducted a controlled study with 100 participants to evaluate this question—and now FG2025 attendees can try it for themselves. Our interactive demo challenges attendees to match heavily degraded probe images to high-quality reference images, simulating conditions common in operational face recognition contexts. In doing so, it highlights the perceptual limits of human recognition and the risk of misidentification in high-stakes settings. The demo runs offline on standard laptops, collects no personal data, and takes about three minutes to complete. Xavier Merino, Gabriella Pangelinan, Samuel Langborgh, Michael C. King, Kevin W. Bowyer |
FG | 4 |
| 2025 | Impact of Sunglasses on One-to-Many Facial Identification AccuracyabstractOne-to-many facial identification is documented to achieve high accuracy in the case where both the probe and the gallery are ‘mugshot quality’ images. However, an increasing number of documented instances of wrongful arrest following one-to-many facial identification have raised questions about its accuracy. Probe images used in one-to-many facial identification are often cropped from frames of surveillance video and deviate from ‘mugshot quality’ in various ways. This paper systematically explores how the accuracy of one-to-many facial identification is degraded by the person in the probe image choosing to wear dark sunglasses. We show that sunglasses degrade accuracy for mugshot-quality images by an amount similar to strong blur or noticeably lower resolution. Further, we demonstrate that the combination of sunglasses with blur or lower resolution results in even more pronounced loss in accuracy. These results have important implications for developing objective criteria to qualify a probe image for the level of accuracy to be expected if it used for one-to-many identification. To ameliorate the accuracy degradation caused by dark sunglasses, we show that it is possible to recover about 38% of the lost accuracy by synthetically adding sunglasses to all the gallery images, without model re-training. We also show that the frequency of wearing-sunglasses images is very low in existing training sets, and that increasing the representation of wearing-sunglasses images can greatly reduce the error rate. The image set assembled for this research is available at https://cvrl.nd.edu/projects/data/ to support replication and further research. Sicong Tian, Haiyu Wu, Michael C. King, Kevin W. Bowyer |
FG | 3 |
| 2025 | One Face, Many Views: Cross-View Consistency of Facial Action Unit Analysis in Multi-Camera SettingsabstractFacial Action Units (AUs) represent individual facial muscle movements and are the building blocks for recognizing expressions and emotions. As such, AU detection is fundamental in facial affect analysis (FAA). While FAA systems are increasingly deployed in real-world applications, most AU detection models are trained and evaluated on frontfacing, well-framed images, which overlook the variability of different camera angles. To address this gap, we introduce MultiFace7, a large-scale, multi-view video corpus designed to evaluate the robustness of AU detection across camera views. Using synchronized recordings, we benchmark open-source and commercial FAA tools by extracting AU features and comparing their consistency across views using a statistical correlation analysis. Our analysis reveals substantial inconsistencies in AU intensity detection depending on the camera angle. These findings highlight limitations in current FAA systems and raise concerns about their trustworthiness in unconstrained environments where optimal camera positioning cannot be guaranteed. While a few view-invariant models and multi-view datasets exist, they are limited in scope, often rely on still images or synthetic views, and lack evidence of use in realworld FAA applications. Our work underscores the need for updated, video-based multi-view benchmarks and more robust, operationally viable AU detection models. Kushal Vangara, Xavier Merino, Gabriella Pangelinan, Michael C. King |
FG | 4 |
| 2022 | Face Regions Impact Recognition Accuracy Differently Across DemographicsabstractVariation in face recognition accuracy across demographic groups has attracted attention from news media, civil liberties advocates and academic researchers. The problem is challenging, in that both the impostor distribution (matches across different people) and the genuine distribution (matches across same people) may vary across demographic groups. Simple answers such as balancing the number of subjects and images in the training data do not have a substantial impact on demographic accuracy disparities. We present the first investigation into whether parts of the face - such as eyes, nose, mouth - show the same accuracy differences across demographic groups as are seen with matching the whole face. We show that matching focused on different parts of the face may result in opposite accuracy differences across demographics. For example, using the eye region for face matching results in Caucasian males having a better impostor distribution (lower similarity scores) than Caucasian females, but using the nose regionfor face matching results in Caucasian females having a better impostor distribution. We also show that it is possible to select face region(s) that effectively minimize the difference in the impostor or genuine distributions across at least some demographics. Our results suggest that a new pathway to reducing accuracy disparity across demographic groups may be to weight the parts of the face differently in matching. Vitor Albiero, Kevin W. Bowyer, Michael C. King |
IJCB | 3 |
| 2022 | Gendered Differences in Face Recognition Accuracy Explained by Hairstyles, Makeup, and Facial MorphologyabstractMedia reports have accused face recognition of being “biased”, “sexist” and “racist”. There is consensus in the research literature that face recognition accuracy is lower for females, who often have both a higher false match rate and a higher false non-match rate. However, there is little published research aimed at identifying the cause of lower accuracy for females. For instance, the 2019 Face Recognition Vendor Test that documents lower female accuracy across a broad range of algorithms and datasets also lists “Analyze cause and effect” under the heading “What we did not do”. We present the first experimental analysis to identify major causes of lower face recognition accuracy for females on datasets where previous research has observed this result. Controlling for equal amount of visible face in the test images mitigates the apparent higher false non-match rate for females. Additional analysis shows that makeup-balanced datasets further improves females to achieve lower false non-match rates. Finally, a clustering experiment suggests that images of two different females are inherently more similar than of two different males, potentially accounting for a difference in false match rates. Vitor Albiero, Kai Zhang 0052, Michael C. King, Kevin W. Bowyer |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Mitigating Attacks on Fake News Detection Systems using Genetic-Based Adversarial TrainingabstractThe study of adversarial effects on AI systems is not a new concept, but much of the research has been devoted to deep learning. In this paper we explore the effects of adversarial examples on 4 machine learning classifiers and measure the effectiveness of adversarial training. Additionally, we present a novel method for selecting adversarial training examples that lead to a more robust machine learning system. Our results suggest that adversarial examples can significantly hinder the classification performance and that adversarial training is an effective defensive counter-measure. Marcellus Smith, Brandon Brown, Gerry V. Dozier, Michael C. King |
CEC | 4 |
| 2021 | Does Face Recognition Error Echo Gender Classification Error?abstractThis paper is the first to explore the question of whether images that are classified incorrectly by a face analytics algorithm (e.g., gender classification) are any more or less likely to participate in an image pair that results in a face recognition error. We analyze results from three different gender classification algorithms (one open-source and two commercial), and two face recognition algorithms (one open-source and one commercial), on image sets representing four demographic groups (African-American female and male, Caucasian female and male). For impostor image pairs, our results show that pairs in which one image has a gender classification error have a better impostor distribution than pairs in which both images have correct gender classification, and so are less likely to generate a false match error. For genuine image pairs, our results show that individuals whose images have a mix of correct and incorrect gender classification have a worse genuine distribution (increased false non-match rate) compared to individuals whose images consistently have correct gender classification. Thus, compared to images that generate correct gender classification, images with gender classification error have a lower false match rate and a higher false non-match rate. Vitor Albiero, Michael C. King, Kevin W. Bowyer |
IJCB | 3 |
| 2020 | Harvesting Faces from Social Media Photos for Biometric AnalysisabstractThe accuracy of automatic face recognition has increased significantly over the last decade. Technology developers actively try to improve their tools and algorithms; for this to occur, there is a need for high-quality datasets with a large number of images to test and develop new techniques. Online social networks provide a vast digital media resource, given the volume of traffic that goes through its infrastructure. The content within it varies but is predominately flooded by images. In this era where selfies are the norm, we examine a collection method employed to harvest face data from the subject's images via the web. We then show how it can be processed and organized so that it is useful for biometric applications. In addition, this experiment demonstrates how restrictions put in place by social media platforms are inadequate in the protection of their user's data. Giordano Benitez Torres, Michael C. King |
ISTAS | 2 |
| 2020 | 'Uh-oh Spaghetti-oh': When Successful Genetic and Evolutionary Feature Selection Makes You More Susceptible to Adversarial Authorship AttacksabstractFeature selection is a technique used to reduce an original set of features to a subset containing the most salient features. Reducing the feature set to the most significant subset of features typically results in an increase in the overall accuracy of a system. It has been shown that in some cases, the use of feature selection can make an underlying system susceptible to adversarial attacks. In this paper, we investigate the susceptibility of a feature selection-based Authorship Attribution System (AAS) to adversarial authorship attacks. The AAS studied is an instance of a Linear Support Vector Machine (LSVM). The feature selection algorithm used is an instance of Genetic & Evolutionary Feature Selection (GEFeS)In order to evaluate the GEFeS+LSVM-based AAS, we use three adversarial authorship masking techniques to generate adversarial texts to attack the AAS. Our results show that in some cases the GEFeS+LSVM-based AAS is more susceptible to adversarial authorship attacks. We provide a simple measurement to determine whether the use of GEFeS is beneficial or detrimental to a LSVM-based AAS. Alexicia Richardson, Gerry V. Dozier, Michael C. King, Richard Chapman 0001 |
SMC | 3 |
| 2020 | Does Face Recognition Accuracy Get Better With Age? Deep Face Matchers Say NoabstractPrevious studies generally agree that face recognition accuracy is higher for older persons than for younger persons. But most previous studies were before the wave of deep learning matchers, and most considered accuracy only in terms of the verification rate for genuine pairs. This paper investigates accuracy for age groups 16-29, 30-49 and 50-70, using three modern deep CNN matchers, and considers differences in the impostor and genuine distributions as well as verification rates and ROC curves. We find that accuracy is lower for older persons and higher for younger persons. In contrast, a pre deep learning matcher on the same dataset shows the traditional result ofhigher accuracy for older persons, although its overall accuracy is much lower than that of the deep learning matchers. Comparing the impostor and genuine distributions, we conclude that impostor scores have a larger effect than genuine scores in causing lower accuracy for the older age group. We also investigate the effects of training data across the age groups. Our results show that fine-tuning the deep CNN models on additional images ofolder persons actually lowers accuracy for the older age group. Also, we fine-tune and train from scratch two models using age-balanced training datasets, and these results also show lower accuracy for older age group. These results argue that the lower accuracy for the older age group is not due to imbalance in the original training data. Vitor Albiero, Kevin W. Bowyer, Kushal Vangara, Michael C. King |
WACV | 4 |
| 2019 | Cyber Influence of Human Behavior: Personal and National Security, Privacy, and Fraud Awareness to Prevent HarmabstractThe Internet and connected technology platforms have enabled an increase of cyber influence activity. These actions target a range of personal to national level security and privacy attributes related to cybercrime, behavior, and identities. These emerging threats call for new indicators for improved awareness, decisions, and action. This research proposes a cyber-physical-human spectrum of identification with a prototyped classification method. Classifier goals are to aid in awareness of activity and potential harmful intent such as detection of identity feature acquisition, fraudulent identities and entities, and targeting or influential behavior. Emerging malicious influence actors prey on human social demographic groups and trends using the Internet infrastructure with social network platform access to large target populations as their attack surface. The methodology discusses how this problem could benefit from a combined human-technical approach to understand indicators of influencing human perception that persuade someone perform a desired action. This method is designed to aid in rapid influence awareness and introduce a counter-influence concept. A prototyped experiment trial demonstrates how awareness may be beneficial to balancing national security with personal privacy. Mary C. Kay Michel, Michael C. King |
ISTAS | 2 |