VLDB 2026 Research / reviewers in the wild / expert
Walter J. Scheirer
dblp:80/4900
· DBLP profile ↗
83ranked-venue papers
12as first author
33since 2021 · last 2026
0000-0001-9649-8074ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 45 · 8 first-author · 18 since 2021Security and privacy · 11 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do We Need Subsidiarity in Software?abstractSubsidiarity is a principle of social organization that promotes human dignity and resists over-centralization by balancing personal autonomy with intervention from higher authorities only when necessary. Thus it is a relevant, but not previously explored, critical lens for discerning the tradeoffs between complete user control of software and surrendering control to "big tech" for convenience, as is common in surveillance capitalism. Our study explores data privacy through the lens of subsidiarity: we employ a multi-method approach of data flow monitoring and user interviews to determine the level of control different everyday technologies currently operate at, and the level of control everyday computer users think is necessary. We found that chat platforms like Slack and Discord violate subsidiarity the most. Our work provides insight into when users are willing to surrender privacy for convenience and demonstrates how subsidiarity can inform designs that promote human flourishing. Louisa Conwill, Megan Levis Scheirer, Walter J. Scheirer |
CHI | 3 |
| 2026 | FaceMINT: A library for gaining insights into biometric face recognition via mechanistic interpretabilityabstractDeep-learning models, including those used in biometric recognition, have achieved remarkable performance on benchmark datasets as well as real-world recognition tasks. However, a major drawback of these models is their lack of transparency in decision-making. Mechanistic interpretability has emerged as a promising research field intended to help us gain insights into such models, but its application to biometric data remains limited. In this work, we bridge this gap by introducing the FaceMINT library, a publicly available Python library (build on top of Pytorch) that enables biometric researchers to inspect their models through mechanistic interpretability. It provides a plug-and-play solution that allows researchers to seamlessly switch between the analyzed biometric models, evaluate state-of-the-art sparse autoencoders, select from various image parametrizations, and fine-tune hyperparameters. Using a large scale Glint360K dataset, we demonstrate the usability of FaceMINT by applying its functionality to two state-of-the-art (deep-learning) face recognition models: AdaFace, based on Convolutional Neural Networks (CNN), and SwinFace, based on transformers. The proposed library implements various sparse auto-encoders (SAEs), including vanilla SAE, Gated SAE, JumpReLU SAE, and TopK SAE, which have achieved state-of-the-art results in the mechanistic interpretability of large language models. Our study highlights the promise of mechanistic interpretability in the biometric field, providing new avenues for researchers to explore model transparency and refine biometric recognition systems. The library is publicly available at www.gitlab.com/peterrot/facemint . Peter Rot, Robert Jutresa, Peter Peer, Vitomir Struc, Walter J. Scheirer, Klemen Grm |
Image Vis. Comput. | 5 |
| 2025 | Design Patterns for the Common Good: Building Better Technologies Using the Wisdom of Virtue Ethics
Louisa Conwill, Megan K. Levis, Karla A. Badillo-Urquiola, Walter J. Scheirer |
CHI | 4 |
| 2025 | Towards Fair and Robust Face Parsing for Generative AI: A Multi-Objective ApproachabstractFace parsing is a fundamental task in computer vision, enabling applications such as identity verification, facial editing, and controllable image synthesis. However, existing face parsing models often lack fairness and robustness, leading to biased segmentation across demographic groups and errors under occlusions, noise, and domain shifts. These limitations affect downstream face synthesis, where segmentation biases can degrade generative model outputs. We propose a multi-objective learning framework that optimizes accuracy, fairness, and robustness in face parsing. Our approach introduces a homotopy-based loss function that dynamically adjusts the importance of these objectives during training. To evaluate its impact, we compare multi-objective and single-objective U-Net models in a GAN-based face synthesis pipeline (Pix2PixHD). Our results show that fairness-aware and robust segmentation improves photorealism and consistency in face generation. Additionally, we conduct preliminary experiments using ControlNet, a structured conditioning model for diffusion-based synthesis, to explore how segmentation quality influences guided image generation. Our findings demonstrate that multi-objective face parsing improves demographic consistency and robustness, leading to higher-quality GAN-based synthesis.11Source code and trained model weights are available on GitHub: https://github.com/sabraha2/Towards-Fair-and-Robust-Face-Parsing-for-Generative-AI-A-Multi-Objective-Approach. Sophia J. Abraham, Jonathan D. Hauenstein, Walter J. Scheirer |
FG | 3 |
| 2025 | A Comprehensive Evaluation Framework for the Study of the Effects of Facial Filters on Face Recognition AccuracyabstractFacial filters are now commonplace for social media users around the world. Previous work has demonstrated that facial filters can negatively impact automated face recognition performance. However, these studies focus on small numbers of hand-picked filters in particular styles. In order to more effectively incorporate the wide ranges of filters present on various social media applications, we introduce a framework that allows for larger-scale study of the impact of facial filters on automated recognition. This framework includes a controlled dataset of face images, a principled filter selection process that selects a representative range of filters for experimentation, and a set of experiments to evaluate the filters’ impact on recognition. We demonstrate our framework with a case study of filters from the American applications Instagram and Snapchat and the Chinese applications Meitu and Pitu to uncover cross-cultural differences. Finally, we show how the filtering effect in a face embedding space can easily be detected and restored to improve face recognition performance. Kagan Öztürk, Louisa Conwill, Jacob Gutierrez, Kevin W. Bowyer, Walter J. Scheirer |
IJCB | 5 |
| 2025 | COSTARR: Consolidated Open Set Technique with Attenuation for Robust RecognitionabstractHandling novelty remains a key challenge in visual recognition systems. Existing open-set recognition (OSR) methods rely on the familiarity hypothesis, detecting novelty by the absence of familiar features. We propose a novel attenuation hypothesis: small weights learned during training attenuate features and serve a dual role-differentiating known classes while discarding information useful for distinguishing known from unknown classes. To leverage this overlooked information, we present COSTARR, a novel approach that combines both the requirement of familiar features and the lack of unfamiliar ones. We provide a probabilistic interpretation of the COSTARR score, linking it to the likelihood of correct classification and belonging in a known class. To determine the individual contributions of the pre- and post-attenuated features to COSTARR's performance, we conduct ablation studies that show both pre-attenuated deep features and the underutilized post-attenuated Hadamard product features are essential for improving OSR. Also, we evaluate COSTARR in a large-scale setting using ImageNet2012-1K as known data and NINCO, iNaturalist, OpenImage-O, and other datasets as unknowns, across multiple modern pre-trained architectures (ViTs, ConvNeXts, and ResNet). The experiments demonstrate that COSTARR generalizes effectively across various architectures and significantly outperforms prior state-of-the-art methods by incorporating previously discarded attenuation information, advancing open-set recognition capabilities. Ryan Rabinowitz, Steve Cruz, Walter J. Scheirer, Terrance E. Boult |
ICCV | 3 |
| 2025 | Human Activity Recognition in an Open World (Abstract Reprint)abstractManaging novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released. Derek S. Prijatelj, Samuel Grieggs, Dawei Du, Ameya Shringi, Christopher Funk, Adam Kaufman, Eric Robertson 0001, Walter J. Scheirer |
IJCAI | 9 |
| 2025 | Psych-Occlusion: Using Visual Psychophysics for Aerial Detection of Occluded Persons During Search and RescueabstractThe success of Emergency Response (ER) scenarios, such as search and rescue, is often dependent upon the prompt location of a lost or injured person. With the increasing use of small Unmanned Aerial Systems (sUAS) as “eyes in the sky” during ER scenarios, efficient detection of persons from aerial views plays a crucial role in achieving a successful mission outcome. Fatigue of human operators during prolonged ER missions, coupled with limited human resources, highlights the need for sUAS equipped with Computer Vision (CV) capabilities to aid in finding the person from aerial views. However, the performance of CV models onboard sUAS substantially degrades under real-life rigorous conditions of a typical ER scenario, where person search is hampered by occlusion and low target resolution. To address these challenges, we extracted images from the NOMAD dataset and performed a crowdsource experiment to collect behavioural measurements when humans were asked to “find the person in the picture”. We exemplify the use of our behavioral dataset, Psych-ER, by using its human accuracy data to adapt the loss function of a detection model. We tested our loss adaptation on a RetinaNet model evaluated on NOMAD against increasing distance and occlusion, with our psychophysical loss adaptation showing improvements over the baseline at higher distances across different levels of occlusion, without degrading performance at closer distances. To the best of our knowledge, our work is the first human-guided approach to address the location task of a detection model, while addressing real-world challenges of aerial search and rescue. All datasets and code can be found at: https://github.com/ArtRuss/NOMad. Arturo Miguel Russell Bernal, Jane Cleland-Huang, Walter J. Scheirer |
WACV | 3 |
| 2024 | This Probably Looks Exactly Like That: An Invertible Prototypical Network
Zachariah Carmichael, Timothy Redgrave, Daniel Gonzalez 0001, Walter J. Scheirer |
ECCV (37) | 4 |
| 2024 | NOMAD: A Natural, Occluded, Multi-scale Aerial Dataset, for Emergency Response ScenariosabstractWith the increasing reliance on small Unmanned Aerial Systems (sUAS) for Emergency Response Scenarios, such as Search and Rescue, the integration of computer vision capabilities has become a key factor in mission success. Nevertheless, computer vision performance for detecting humans severely degrades when shifting from ground to aerial views. Several aerial datasets have been created to mitigate this problem, however, none of them has specifically addressed the issue of occlusion, a critical component in Emergency Response Scenarios. Natural Occluded Multi-scale Aerial Dataset (NOMAD) presents a benchmark for human detection under occluded aerial views, with five different aerial distances and rich imagery variance. NOMAD is composed of 100 different Actors, all performing sequences of walking, laying and hiding. It includes 42,825 frames, extracted from 5.4k resolution videos, and manually annotated with a bounding box and a label describing 10 different visibility levels, categorized according to the percentage of the human body visible inside the bounding box. This allows computer vision models to be evaluated on their detection performance across different ranges of occlusion. NOMAD is designed to improve the effectiveness of aerial search and rescue and to enhance collaboration between sUAS and humans, by providing a new benchmark dataset for human detection under occluded aerial views. Arturo Miguel Russell Bernal, Walter J. Scheirer, Jane Cleland-Huang |
WACV | 2 |
| 2024 | Pixel-Grounded Prototypical Part NetworksabstractPrototypical part neural networks (ProtoPartNNs), namely ProtoPNet and its derivatives, are an intrinsically interpretable approach to machine learning. Their prototype learning scheme enables intuitive explanations of the form, this (prototype) looks like that (testing image patch). But, does this actually look like that? In this work, we delve into why object part localization and associated heat maps in past work are misleading. Rather than localizing to object parts, existing ProtoPartNNs localize to the entire image, contrary to generated explanatory visualizations. We argue that detraction from these underlying issues is due to the alluring nature of visualizations and an over-reliance on intuition. To alleviate these issues, we devise new receptive field-based architectural constraints for meaningful localization and a principled pixel space mapping for ProtoPartNNs. To improve interpretability, we propose additional architectural improvements, including a simplified classification head. We also make additional corrections to ProtoPNet and its derivatives, such as the use of a validation set, rather than a test set, to evaluate generalization during training. Our approach, PixPNet (Pixel-grounded Prototypical part Network), is the only ProtoPartNN that truly learns and localizes to prototypical object parts. We demonstrate that PixPNet achieves quantifiably improved interpretability without sacrificing accuracy1. Zachariah Carmichael, Suhas Lohit, Anoop Cherian, Michael J. Jones 0001, Walter J. Scheirer |
WACV | 5 |
| 2024 | The Paleographer's Eye ex machina: Using Computer Vision to Assist Humanists in Scribal Hand IdentificationabstractThe steady digitization of medieval manuscripts is rapidly changing the field of paleography, challenging existing assumptions about handwriting and book production. This development has identified historically important centers for the production of scribal texts, and even individual scribes themselves. For example, scholars of late medieval English literature have identified the copyists of a number of literary manuscripts, and the important role of London government clerks in shaping literary culture. However, traditional paleography has no agreed-upon methodology or fixed criteria for the attribution of handwriting to a particular community, period, or scribe. The approach taken by paleographers is inherently qualitative and subject to personal bias. Even those wielding the mighty "paleographer’s eye" cannot claim objectivity. Computer vision offers solutions with spectacular performance on writer identification and retrieval benchmarks, but these have not been widely adopted by the paleography community because they tend not to hold up in practice. In this work, we attempt to bridge the divide with a software package designed not to automate paleography, but to augment the paleographer’s eye. We introduce automated handwriting identification tools for which the results can be quickly visually understood and assessed, and used as one feature among many by expert paleographers when attributing previously unknown scribal hands. We also demonstrate a use case for our software by analyzing several items believed to be written by Thomas Hoccleve, a highly productive clerk of the Privy Seal who is also an important fifteenth-century English poet. Samuel Grieggs, C. E. M. Henderson, Sebastian Sobecki, Alexandra Gillespie, Walter J. Scheirer |
WACV | 5 |
| 2024 | C-CLIP: Contrastive Image-Text Encoders to Close the Descriptive-Commentative GapabstractThe interplay between the image and comment on a social media post is one of high importance for understanding its overall message. Recent strides in multimodal embedding models, namely CLIP, have provided an avenue forward in relating image and text. However the current training regime for CLIP models is insufficient for matching content found on social media, regardless of site or language. Current CLIP training data is based on what we call "descriptive" text: text in which an image is merely described. This is something rarely seen on social media, where the vast majority of text content is "commentative" in nature. The captions provide commentary and broader context related to the image, rather than describing what is in it. Current CLIP models perform poorly on retrieval tasks where image-caption pairs display a commentative relationship. Closing this gap would be beneficial for several important application areas related to social media. For instance, it would allow groups focused on Open-Source Intelligence Operations (OSINT) to further aid efforts during disaster events, such as the ongoing Russian invasion of Ukraine, by easily exposing data to non-technical users for discovery and analysis. In order to close this gap we demonstrate that training contrastive image-text encoders on explicitly commentative pairs results in large improvements in retrieval results, with the results extending across a variety of nonEnglish languages. William Theisen, Walter J. Scheirer |
WACV | 2 |
| 2024 | Human Activity Recognition in an Open WorldabstractManaging novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released. Derek S. Prijatelj, Samuel Grieggs, Dawei Du, Ameya Shringi, Christopher Funk, Adam Kaufman, Eric Robertson 0001, Walter J. Scheirer |
J. Artif. Intell. Res. | 9 |
| 2024 | Informing Machine Perception With PsychophysicsabstractGustav Fechner’s 1860 delineation of psychophysics, the measurement of sensation in relation to its stimulus, is widely considered to be the advent of modern psychological science. In psychophysics, a researcher parametrically varies some aspects of a stimulus and measures the resulting changes in a human subject’s experience of that stimulus; doing so gives insight into the determining relationship between a sensation and the physical input that evoked it. This approach is used heavily in perceptual domains, including signal detection, threshold measurement, and ideal observer analysis. Scientific fields, such as vision science, have always leaned heavily on the methods and procedures of psychophysics, but there is now growing appreciation of them by machine learning researchers, sparked by widening overlap between biological and artificial perception[1],[2],[3],[4],[5]. Machine perception that is guided by behavioral measurements, as opposed to guidance restricted to arbitrarily assigned human labels, has significant potential to fuel further progress in artificial intelligence (AI). Justin Dulay, Sonia Poltoratski, Till S. Hartmann, Samuel E. Anthony, Walter J. Scheirer |
Proc. IEEE | 5 |
| 2023 | Unfooling Perturbation-Based Post Hoc ExplainersabstractMonumental advancements in artificial intelligence (AI) have lured the interest of doctors, lenders, judges, and other professionals. While these high-stakes decision-makers are optimistic about the technology, those familiar with AI systems are wary about the lack of transparency of its decision-making processes. Perturbation-based post hoc explainers offer a model agnostic means of interpreting these systems while only requiring query-level access. However, recent work demonstrates that these explainers can be fooled adversarially. This discovery has adverse implications for auditors, regulators, and other sentinels. With this in mind, several natural questions arise - how can we audit these black box systems? And how can we ascertain that the auditee is complying with the audit in good faith? In this work, we rigorously formalize this problem and devise a defense against adversarial attacks on perturbation-based explainers. We propose algorithms for the detection (CAD-Detect) and defense (CAD-Defend) of these attacks, which are aided by our novel conditional anomaly detection approach, KNN-CAD. We demonstrate that our approach successfully detects whether a black box system adversarially conceals its decision-making process and mitigates the adversarial attack on real-world data for the prevalent explainers, LIME and SHAP. The code for this work is available at https://github.com/craymichael/unfooling. Zachariah Carmichael, Walter J. Scheirer |
AAAI | 2 |
| 2023 | Psychophysical-Score: A Behavioral Measure for Assessing the Biological Plausibility of Visual Recognition Models
Brandon RichardWebster, Justin Dulay, Anthony DiFalco, Walter J. Scheirer |
CogSci | 4 |
| 2023 | Analyzing the Impact of Shape & Context on the Face Recognition Performance of Deep NetworksabstractIn this article, we analyze how changing the underlying 3D shape of the base identity in face images can distort their overall appearance, especially from the perspective of deep face recognition. As done in popular training data augmentation schemes, we graphically render real and synthetic face images with randomly chosen or best-fitting 3D face models to generate novel views of the base identity. We compare deep features generated from these images to assess the perturbation these renderings introduce into the original identity. We perform this analysis at various degrees of facial yaw with the base identities varying in gender and ethnicity. Additionally, we investigate if adding some form of context and background pixels in these rendered images, when used as training data, further improves the downstream performance of a face recognition model. Our experiments demonstrate the significance of facial shape in accurate face matching and underpin the importance of contextual data for network training. Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn |
FG | 2 |
| 2023 | Photoshop FantasiesabstractThe possibility of an altered photo revising history in a convincing way highlights a salient threat of imaging technology. Afterall, seeing is believing. Or is it? The examples history has preserved make it clear that the observer is more often than not meant to understand that something has changed. Surprisingly, the objectives of photographic manipulation have remained largely the same since the camera first appeared in the 19th century. The old battleworn techniques have simply evolved to keep pace with technological developments. In this talk, we will learn about the history of photographic manipulation, from the invention of the camera to the present day. Importantly, we will consider the reception of photo editing and its relationship to the notion of reality, which is more significant than the technologies themselves. Surprisingly, we will discover that creative mythmaking has found a new medium to embed itself in. Walter J. Scheirer |
IH&MMSec | 1 |
| 2023 | Motif Mining: Finding and Summarizing Remixed Image ContentabstractOn the Internet, images are no longer static; they have become dynamic content. Thanks to the availability of smartphones with cameras and easy-to-use editing software, images can be remixed (i.e., redacted, edited, and re-combined with other content) on-the-fly, allowing a world-wide audience to repeat the process many times. From digital art to memes, the evolution of images through time is now an important topic of study for digital humanists, social scientists, and media forensics specialists. However, because typical data sets in computer vision are composed of static content, there has been limited development of automated algorithms for analyzing remixed content. In this paper, we propose the idea of Motif Mining: the process of finding and summarizing remixed image content in large collections of unlabeled and unsorted data. For the first time, this idea is formalized and a reference implementation grounded in that formalism is introduced. We conduct experiments on three meme-style data sets, including a newly collected set associated with the Russo-Ukrainian conflict. The proposed motif mining approach is able to identify related remixed content that, when compared to similar approaches, more closely aligns with the preferences and expectations of human observers. William Theisen, Daniel Gonzalez 0001, Zachariah Carmichael, Daniel Moreira, Tim Weninger, Walter J. Scheirer |
WACV | 6 |
| 2023 | Measuring Human Perception to Improve Open Set RecognitionabstractThe human ability to recognize when an object belongs or does not belong to a particular vision task outperforms all open set recognition algorithms. Human perception as measured by the methods and procedures of visual psychophysics from psychology provides an additional data stream for algorithms that need to manage novelty. For instance, measured reaction time from human subjects can offer insight as to whether a class sample is prone to be confused with a different class - known or novel. In this work, we designed and performed a large-scale behavioral experiment that collected over 200,000 human reaction time measurements associated with object recognition. The data collected indicated reaction time varies meaningfully across objects at the sample-level. We therefore designed a new psychophysical loss function that enforces consistency with human behavior in deep networks which exhibit variable reaction time for different images. As in biological vision, this approach allows us to achieve good open set recognition performance in regimes with limited labeled training data. Through experiments using data from ImageNet, significant improvement is observed when training Multi-Scale DenseNets with this new formulation: it significantly improved top-1 validation accuracy by 6.02%, top-1 test accuracy on known samples by 9.81%, and top-1 test accuracy on unknown samples by 33.18%. We compared our method to 10 open set recognition methods from the literature, which were all outperformed on multiple metrics. Derek S. Prijatelj, Justin Dulay, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Measuring Human Perception to Improve Handwritten Document TranscriptionabstractIn this paper, we consider how to incorporate psychophysical measurements of human visual perception into the loss function of a deep neural network being trained for a recognition task, under the assumption that such information can reduce errors. As a case study to assess the viability of this approach, we look at the problem of handwritten document transcription. While good progress has been made towards automatically transcribing modern handwriting, significant challenges remain in transcribing historical documents. Here we describe a general enhancement strategy, underpinned by the new loss formulation, which can be applied to the training regime of any deep learning-based document transcription system. Through experimentation, reliable performance improvement is demonstrated for the standard IAM and RIMES datasets for three different network architectures. Further, we go on to show feasibility for our approach on a new dataset of digitized Latin manuscripts, originally produced by scribes in the Cloister of St. Gall in the the 9th century. Samuel Grieggs, Bingyu Shen 0001, Greta Rauch, David Chiang 0001, Brian L. Price, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | A Bayesian evaluation framework for subjectively annotated visual recognition tasks
Derek S. Prijatelj, Mel McCurrie, Samuel E. Anthony, Walter J. Scheirer |
Pattern Recognit. | 4 |
| 2021 | Towards a Unifying Framework for Formal Theories of NoveltyabstractManaging inputs that are novel, unknown, or out-of-distribution is critical as an agent moves from the lab to the open world. Novelty-related problems include being tolerant to novel perturbations of the normal input, detecting when the input includes novel items, and adapting to novel inputs. While significant research has been undertaken in these areas, a noticeable gap exists in the lack of a formalized definition of novelty that transcends problem domains. As a team of researchers spanning multiple research groups and different domains, we have seen, first hand, the difficulties that arise from ill-specified novelty problems, as well as inconsistent definitions and terminology. Therefore, we present the first unified framework for formal theories of novelty and use the framework to formally define a family of novelty types. Our framework can be applied across a wide range of domains, from symbolic AI to reinforcement learning, and beyond to open world image recognition. Thus, it can be used to help kick-start new research efforts and accelerate ongoing work on these important novelty-related problems. Terrance E. Boult, Przemyslaw A. Grabowicz, Derek S. Prijatelj, Roni Stern, Lawrence B. Holder, Joshua Alspector, Mohsen Jafarzadeh, Touqeer Ahmad, Akshay Raj Dhamija, Chunchun Li, Steve Cruz, Abhinav Shrivastava, Carl Vondrick, Walter J. Scheirer |
AAAI | 14 |
| 2021 | A Study of the Human Perception of Synthetic FacesabstractAdvances in face synthesis have raised alarms about the deceptive use of synthetic faces. Can synthetic identities be effectively used to fool human observers? In this paper, we introduce a study of the human perception of synthetic faces generated using different strategies including a state-of-the-art deep learning-based GAN model. This is the first rigorous study of the effectiveness of synthetic face generation techniques grounded in experimental techniques from psychology. We answer important questions such as how often do GAN-based and more traditional image processing-based techniques confuse human observers, and are there subtle cues within a synthetic face image that cause humans to perceive it as a fake without having to search for obvious clues? To answer these questions, we conducted a series of large-scale crowdsourced behavioral experiments with different sources of face imagery. Results show that humans are unable to distinguish synthetic faces from real faces under several different circumstances. This finding has serious implications for many different applications where face images are presented to human users. Bingyu Shen 0001, Brandon RichardWebster, Alice J. O'Toole, Kevin W. Bowyer, Walter J. Scheirer |
FG | 5 |
| 2021 | Handwriting Recognition with Novelty
Derek S. Prijatelj, Samuel Grieggs, Futoshi Yumoto, Eric Robertson 0001, Walter J. Scheirer |
ICDAR (4) | 5 |
| 2021 | Automatic Discovery of Political Meme Genres with Diverse Appearances
William Theisen, Joel Brogan, Pamela Bilo Thomas, Daniel Moreira, Pascal Phoa, Tim Weninger, Walter J. Scheirer |
ICWSM | 7 |
| 2021 | Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed SequencesabstractUnderstanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- or even hour-long video data, but annotating such data is very time consuming and can not easily be automated or scaled. To address this problem, this paper proposes an approach for the unsupervised learning of actions in untrimmed video sequences based on a joint visual-temporal embedding space. To this end, we combine a visual embedding based on a predictive U-Net architecture with a temporal continuous function. The resulting representation space allows detecting relevant action clusters based on their visual as well as their temporal appearance. The proposed method is evaluated on three standard benchmark datasets, Breakfast Actions, INRIA YouTube Instructional Videos, and 50 Salads. We show that the proposed approach is able to provide a meaningful visual and temporal embedding out of the visual cues present in contiguous video frames and is suitable for the task of unsupervised temporal segmentation of actions. Rosaura G. VidalMata, Walter J. Scheirer, Anna Kukleva, David D. Cox, Hilde Kuehne |
WACV | 2 |
| 2021 | Report on UG2+ challenge Track 1: Assessing algorithms to improve video object detection and classification from unconstrained mobility platforms
Sreya Banerjee, Rosaura G. VidalMata, Zhangyang Wang, Walter J. Scheirer |
Comput. Vis. Image Underst. | 4 |
| 2021 | Bridging the Gap Between Computational Photography and Visual RecognitionabstractWhat is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area. Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 24 |
| 2021 | Transformation-Aware Embeddings for Image ProvenanceabstractA dramatic rise in the flow of manipulated image content on the Internet has led to a prompt response from the media forensics research community. New mitigation efforts leverage cutting-edge data-driven strategies and increasingly incorporate usage of techniques from computer vision and machine learning to detect and profile the space of image manipulations. This paper addresses Image Provenance Analysis, which aims at discovering relationships among different manipulated image versions that share content. One important task in provenance analysis, like most visual understanding problems, is establishing a visual description and dissimilarity computation method that connects images that share full or partial content. But the existing handcrafted or learned descriptors - generally appropriate for tasks such as object recognition - may not sufficiently encode the subtle differences between near-duplicate image variants, which significantly characterize the provenance of any image. This paper introduces a novel data-driven learning-based approach that provides the context for ordering images that have been generated from a single image source through various transformations. Our approach learns transformation-aware embeddings using weak supervision via composited transformations and a rank-based Edit Sequence Loss. To establish the effectiveness of the proposed approach, comparisons are made with state-of-the-art handcrafted and deep-learning-based descriptors, as well as image matching approaches. Further experimentation validates the proposed approach in the context of image provenance analysis and improves upon existing approaches. Aparna Bharati, Daniel Moreira, Patrick J. Flynn, Anderson Rocha 0001, Kevin W. Bowyer, Walter J. Scheirer |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Privacy-Enhancing Face Biometrics: A Comprehensive SurveyabstractBiometric recognition technology has made significant advances over the last decade and is now used across a number of services and applications. However, this widespread deployment has also resulted in privacy concerns and evolving societal expectations about the appropriate use of the technology. For example, the ability to automatically extract age, gender, race, and health cues from biometric data has heightened concerns about privacy leakage. Face recognition technology, in particular, has been in the spotlight, and is now seen by many as posing a considerable risk to personal privacy. In response to these and similar concerns, researchers have intensified efforts towards developing techniques and computational models capable of ensuring privacy to individuals, while still facilitating the utility of face recognition technology in several application scenarios. These efforts have resulted in a multitude of privacy-enhancing techniques that aim at addressing privacy risks originating from biometric systems and providing technological solutions for legislative requirements set forth in privacy laws and regulations, such as GDPR. The goal of this overview paper is to provide a comprehensive introduction into privacy-related research in the area of biometrics and review existing work on Biometric Privacy-Enhancing Techniques (B-PETs) applied to face biometrics. To make this work useful for as wide of an audience as possible, several key topics are covered as well, including evaluation strategies used with B-PETs, existing datasets, relevant standards, and regulations and critical open issues that will have to be addressed in the future. Blaz Meden, Peter Rot, Philipp Terhörst, Naser Damer, Arjan Kuijper, Walter J. Scheirer, Arun Ross, Peter Peer, Vitomir Struc |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image RetrievalabstractImages from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates - and to do so with enough efficiency at large scale. We propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit. Joel Brogan, Aparna Bharati, Daniel Moreira, Anderson Rocha 0001, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer |
IEEE Trans. Image Process. | 7 |
| 2020 | The Next Generation of Human-Drone Partnerships: Co-Designing an Emergency Response SystemabstractThe use of semi-autonomous Unmanned Aerial Vehicles (UAV) to support emergency response scenarios, such as fire surveillance and search and rescue, offers the potential for huge societal benefits. However, designing an effective solution in this complex domain represents a "wicked design" problem, requiring a careful balance between trade-offs associated with drone autonomy versus human control, mission functionality versus safety, and the diverse needs of different stakeholders. This paper focuses on designing for situational awareness (SA) using a scenario-driven, participatory design process. We developed SA cards describing six common design-problems, known as SA demons, and three new demons of importance to our domain. We then used these SA cards to equip domain experts with SA knowledge so that they could more fully engage in the design process. We designed a potentially reusable solution for achieving SA in multi-stakeholder, multi-UAV, emergency response applications. Ankit Agrawal 0002, Sophia J. Abraham, Benjamin Burger, Chichi Christine, Luke Fraser, John M. Hoeksema, Sarah Hwang, Elizabeth Travnik, Shreya Kumar, Walter J. Scheirer, Jane Cleland-Huang, Michael Vierhauser, Ryan Bauer, Steve Cox 0002 |
CHI | 10 |
| 2020 | Backdooring Convolutional Neural Networks via Targeted Weight PerturbationsabstractWe present a new white-box backdoor attack that exploits a vulnerability of convolutional neural networks (CNNs). In particular, we examine the application of facial recognition. Deep learning techniques are at the top of the game for facial recognition, which means they have now been implemented in many production-level systems. Alarmingly, unlike other commercial technologies such as operating systems and network devices, deep learning-based facial recognition algorithms are not presently designed with security requirements or audited for security vulnerabilities before deployment. Given how young the technology is and how abstract many of the internal workings of these algorithms are, neural network-based facial recognition systems are prime targets for security breaches. As more and more of our personal information begins to be guarded by facial recognition (e.g., the iPhone X), exploring the security vulnerabilities of these systems from a penetration testing standpoint is crucial. Along these lines, we describe a general methodology for backdooring CNNs via targeted weight perturbations. Using a five-layer CNN and ResNet-50 as case studies, we show that an attacker is able to significantly increase the chance that inputs they supply will be falsely accepted by a CNN while simultaneously preserving the error rates for legitimate enrolled classes. Jacob Dumford, Walter J. Scheirer |
IJCB | 2 |
| 2020 | Modeling Score Distributions and Continuous Covariates: A Bayesian ApproachabstractComputer Vision practitioners must thoroughly understand their model's performance, but conditional evaluation is complex and error-prone. In biometric verification, model performance over continuous covariates - known, real-number attributes of images that affect performance - is particularly challenging to study. We develop a generative model of the match and non-match score distributions over continuous covariates and perform inference with modern Bayesian methods. We use mixture models to capture arbitrary distributions and local basis functions to capture non-linear, multivariate trends. Three experiments demonstrate the accuracy and effectiveness of our approach. First, we study the relationship between age and face verification performance and find previous methods may overstate performance and confidence. Second, we study preprocessing for CNNs and find a highly non-linear, multivariate surface of model performance. Our method is accurate and data efficient when evaluated against previous synthetic methods. Third, we demonstrate the novel application of our method to pedestrian tracking and calculate variable thresholds and expected performance while controlling for multiple covariates. Mel McCurrie, Hamish Nicholson, Walter J. Scheirer, Samuel E. Anthony |
IJCB | 3 |
| 2020 | On Hallucinating Context and Background Pixels from a Face Mask using Multi-scale GANsabstractWe propose a multi-scale GAN model to hallucinate realistic context (forehead, hair, neck, clothes) and background pixels automatically from a single input face mask, without any user supervision. Instead of swapping a face on to an existing picture, our model directly generates realistic context and background pixels based on the features of the provided face mask. Unlike facial inpainting algorithms, it can generate realistic hallucinations even for a large number of missing pixels. Our model is composed of a cascaded network of GAN blocks, each tasked with hallucination of missing pixels at a particular resolution while guiding the synthesis process of the next GAN block. The hallucinated full face image is made photo-realistic by using a combination of reconstruction, perceptual, adversarial and identity preserving losses at each block of the network. With a set of extensive experiments, we demonstrate the effectiveness of our model in hallucinating context and background pixels from face masks varying in facial pose, expression and lighting, collected from multiple datasets subject disjoint with our training data. We also compare our method with popular face inpainting and face swapping models in terms of visual quality, realism and identity preservation. Additionally, we analyze our cascaded pipeline and compare it with the progressive growing of GANs, and explore its usage as a data augmentation module for training CNNs. Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn |
WACV | 2 |
| 2020 | Face Hallucination Using Cascaded Super-Resolution and Identity PriorsabstractIn this paper we address the problem of hallucinating high-resolution facial images from low-resolution inputs at high magnification factors. We approach this task with convolutional neural networks (CNNs) and propose a novel (deep) face hallucination model that incorporates identity priors into the learning procedure. The model consists of two main parts: i) a cascaded super-resolution network that upscales the low-resolution facial images, and ii) an ensemble of face recognition models that act as identity priors for the super-resolution network during training. Different from most competing super-resolution techniques that rely on a single model for upscaling (even with large magnification factors), our network uses a cascade of multiple SR models that progressively upscale the low-resolution images using steps of 2× . This characteristic allows us to apply supervision signals (target appearances) at different resolutions and incorporate identity constraints at multiple-scales. The proposed C-SRIP model (Cascaded Super Resolution with Identity Priors) is able to upscale (tiny) low-resolution images captured in unconstrained conditions and produce visually convincing results for diverse low-resolution inputs. We rigorously evaluate the proposed model on the Labeled Faces in the Wild (LFW), Helen and CelebA datasets and report superior performance compared to the existing state-of-the-art. Klemen Grm, Walter J. Scheirer, Vitomir Struc |
IEEE Trans. Image Process. | 2 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 5 |
| 2019 | Learning and the Unknown: Surveying Steps toward Open World RecognitionabstractAs science attempts to close the gap between man and machine by building systems capable of learning, we must embrace the importance of the unknown. The ability to differentiate between known and unknown can be considered a critical element of any intelligent self-learning system. The ability to reject uncertain inputs has a very long history in machine learning, as does including a background or garbage class to account for inputs that are not of interest. This paper explains why neither of these is genuinely sufficient for handling unknown inputs – uncertain is not unknown, and unknowns need not appear to be uncertain to a learning system. The past decade has seen the formalization and development of many open set algorithms, which provably bound the risk from unknown classes. We summarize the state of the art, core ideas, and results and explain why, despite the efforts to date, the current techniques are genuinely insufficient for handling unknown inputs, especially for deep networks. Terrance E. Boult, Steve Cruz, Akshay Raj Dhamija, Manuel Günther, James Henrydoss, Walter J. Scheirer |
AAAI | 6 |
| 2019 | A Neurobiological Evaluation Metric for Neural Network Model SearchabstractNeuroscience theory posits that the brain's visual system coarsely identifies broad object categories via neural activation patterns, with similar objects producing similar neural responses. Artificial neural networks also have internal activation behavior in response to stimuli. We hypothesize that networks exhibiting brain-like activation behavior will demonstrate brain-like characteristics, e.g., stronger generalization capabilities. In this paper we introduce a human-model similarity (HMS) metric, which quantifies the similarity of human fMRI and network activation behavior. To calculate HMS, representational dissimilarity matrices (RDMs) are created as abstractions of activation behavior, measured by the correlations of activations to stimulus pairs. HMS is then the correlation between the fMRI RDM and the neural network RDM across all stimulus pairs. We test the metric on unsupervised predictive coding networks, which specifically model visual perception, and assess the metric for statistical significance over a large range of hyperparameters. Our experiments show that networks with increased human-model similarity are correlated with better performance on two computer vision tasks: next frame prediction and object matching accuracy. Further, HMS identifies networks with high performance on both tasks. An unexpected secondary finding is that the metric can be employed during training as an early-stopping mechanism. Nathaniel Blanchard, Jeffery Kinnison, Brandon RichardWebster, Pouya Bashivan, Walter J. Scheirer |
CVPR | 5 |
| 2019 | Fast Face Image Synthesis With Minimal TrainingabstractWe propose an algorithm to generate realistic face images of both real and synthetic identities (people who do not exist) with different facial yaw, shape and resolution. The synthesized images can be used to augment datasets to train CNNs or as massive distractor sets for biometric verification experiments without any privacy concerns. Additionally, law enforcement can make use of this technique to train forensic experts to recognize faces. Our method samples face components from a pool of multiple face images of real identities to generate the synthetic texture. Then, a real 3D head model compatible to the generated texture is used to render it under different facial yaw transformations. We perform multiple quantitative experiments to assess the effectiveness of our synthesis procedure in CNN training and its potential use to generate distractor face images. Additionally, we compare our method with popular GAN models in terms of visual quality and execution time. Sandipan Banerjee, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn |
WACV | 2 |
| 2019 | Beyond Pixels: Image Provenance Analysis Leveraging MetadataabstractCreative works, whether paintings or memes, follow unique journeys that result in their final form. Understanding these journeys, a process known as "provenance analysis," provides rich insights into the use, motivation, and authenticity underlying any given work. The application of this type of study to the expanse of unregulated content on the Internet is what we consider in this paper. Provenance analysis provides a snapshot of the chronology and validity of content as it is uploaded, re-uploaded, and modified over time. Although still in its infancy, automated provenance analysis for online multimedia is already being applied to different types of content. Most current works seek to build provenance graphs based on the shared content between images or videos. This can be a computationally expensive task, especially when considering the vast influx of content that the Internet sees every day. Utilizing non-content-based information, such as timestamps, geotags, and camera IDs can help provide important insights into the path a particular image or video has traveled during its time on the Internet without large computational overhead. This paper tests the scope and applicability of metadata-based inferences for provenance graph construction in two different scenarios: digital image forensics and cultural analytics. Aparna Bharati, Daniel Moreira, Joel Brogan, Patricia Hale, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
WACV | 8 |
| 2019 | "Keep Me In, Coach!": A Computer Vision Perspective on Assessing ACL Injury Risk in Female AthletesabstractWe present and share a foundational dataset of multi-angle video recordings of scripted athletic movements to enable the development of computer vision research applications that evaluate and identify lower-body injury risk. The focus of the dataset is female athletes, who are at a substantially increased risk of anterior cruciate ligament (ACL) injury and are therefore a top priority for sports science. In our study, varsity and club sport athletes perform two assessment movements (the countermovement jump and the drop jump). These jump tasks are used ubiquitously in sports medicine research to characterize athleticism and to identify risk factors that indicate ACL injury propensity. The novelty of the dataset centers on (i) the type of movement data (purposeful, evaluative movements that need to be tracked with a high degree of precision), (ii) our generalized collection method that can be replicated with ease by non-experts, and (iii) the amount of data collected (we collected data from 55 division one (D1) female athletes performing 3-5 iterations of each jumps, for a total of 480 jumps). Data from each camera was manually aligned and a fully automated pipeline was built to extract knee information from athletes. Ideally, any athlete or researcher will be able to easily replicate our setup and assemble a compatible and complementary dataset to propel the development and assessment of injury propensity models. Nathaniel Blanchard, Kyle Skinner, Aden Kemp, Walter J. Scheirer, Patrick J. Flynn |
WACV | 4 |
| 2019 | PsyPhy: A Psychophysics Driven Evaluation Framework for Visual RecognitionabstractBy providing substantial amounts of data and standardized evaluation protocols, datasets in computer vision have helped fuel advances across all areas of visual recognition. But even in light of breakthrough results on recent benchmarks, it is still fair to ask if our recognition algorithms are doing as well as we think they are. The vision sciences at large make use of a very different evaluation regime known as Visual Psychophysics to study visual perception. Psychophysics is the quantitative examination of the relationships between controlled stimuli and the behavioral responses they elicit in experimental test subjects. Instead of using summary statistics to gauge performance, psychophysics directs us to construct item-response curves made up of individual stimulus responses to find perceptual thresholds, thus allowing one to identify the exact point at which a subject can no longer reliably recognize the stimulus class. In this article, we introduce a comprehensive evaluation framework for visual recognition models that is underpinned by this methodology. Over millions of procedurally rendered 3D scenes and 2D images, we compare the performance of well-known convolutional neural networks. Our results bring into question recent claims of human-like performance, and provide a path forward for correcting newly surfaced algorithmic deficiencies. Brandon RichardWebster, Samuel E. Anthony, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Visual Psychophysics for Making Face Recognition Algorithms More Explainable
Brandon RichardWebster, So Yon Kwon, Christopher Clarizio, Samuel E. Anthony, Walter J. Scheirer |
ECCV (15) | 5 |
| 2018 | Coupling Story to Visualization: Using Textual Analysis as a Bridge Between Data and InterpretationabstractOnline writers and journalism media are increasingly combining visualization (and other multimedia content) with narrative text to create narrative visualizations. Often, however, the two elements are presented independently of one another. We propose an approach to automatically integrate text and visualization elements. We begin with a writer»s narrative that presumably can be supported with visual data evidence. We leverage natural language processing, quantitative narrative analysis, and information visualization to (1) automatically extract narrative components (who, what, when, where) from data-rich stories, and (2) integrate the supporting data evidence with the text to develop a narrative visualization. We also employ bidirectional interaction from text to visualization and visualization to text to support reader exploration in both directions. We demonstrate the approach with a case study in the data-rich field of sports journalism. Ronald A. Metoyer, Qiyu Zhi, Bart Janczuk, Walter J. Scheirer |
IUI | 4 |
| 2018 | To Frontalize or Not to Frontalize: Do We Really Need Elaborate Pre-processing to Improve Face Recognition?abstractFace recognition performance has improved remarkably in the last decade. Much of this success can be attributed to the development of deep learning techniques such as convolutional neural networks (CNNs). While CNNs have pushed the state-of-the-art forward, their training process requires a large amount of clean and correctly labelled training data. If a CNN is intended to tolerate facial pose, then we face an important question: should this training data be diverse in its pose distribution, or should face images be normalized to a single pose in a pre-processing step? To address this question, we evaluate a number of facial landmarking algorithms and a popular frontalization method to understand their effect on facial recognition performance. Additionally, we introduce a new, automatic, single-image frontalization scheme that exceeds the performance of the reference frontalization algorithm for video-to-video face matching on the Point and Shoot Challenge (PaSC) dataset. Additionally, we investigate failure modes of each frontalization method on different facial yaw using the CMU Multi-PIE dataset. We assert that the subsequent recognition and verification performance serves to quantify the effectiveness of each pose correction scheme. Sandipan Banerjee, Joel Brogan, Janez Krizaj, Aparna Bharati, Brandon RichardWebster, Vitomir Struc, Patrick J. Flynn, Walter J. Scheirer |
WACV | 8 |
| 2018 | SHADHO: Massively Scalable Hardware-Aware Distributed Hyperparameter OptimizationabstractComputer vision is experiencing an AI renaissance, in which machine learning models are expediting important breakthroughs in academic research and commercial applications. Effectively training these models, however, is not trivial due in part to hyperparameters: user-configured values that control a model's ability to learn from data. Existing hyperparameter optimization methods are highly parallel but make no effort to balance the search across heterogeneous hardware or to prioritize searching high-impact spaces. In this paper, we introduce a framework for massively Scalable Hardware-Aware Distributed Hyperparameter Optimization (SHADHO). Our framework calculates the relative complexity of each search space and monitors performance on the learning task over all trials. These metrics are then used as heuristics to assign hyperparameters to distributed workers based on their hardware. We first demonstrate that our framework achieves double the throughput of a standard distributed hyperparameter optimization framework by optimizing SVM for MNIST using 150 distributed workers. We then conduct model search with SHADHO over the course of one week using 74 GPUs across two compute clusters to optimize U-Net for a cell segmentation task, discovering 515 models that achieve a lower validation loss than standard U-Net. Jeffery Kinnison, Nathaniel Kremer-Herman, Douglas Thain, Walter J. Scheirer |
WACV | 4 |
| 2018 | UG^2: A Video Benchmark for Assessing the Impact of Image Restoration and Enhancement on Automatic Visual RecognitionabstractAdvances in image restoration and enhancement techniques have led to discussion about how such algorithms can be applied as a pre-processing step to improve automatic visual recognition. In principle, techniques like deblurring and super-resolution should yield improvements by de-emphasizing noise and increasing signal in an input image. But the historically divergent goals of computational photography and visual recognition communities have created a significant need for more work in this direction. To facilitate new research, we introduce a new benchmark dataset called UG2, which contains three difficult real-world scenarios: uncontrolled videos taken by UAVs and manned gliders, as well as controlled videos taken on the ground. Over 150,000 annotated frames for hundreds of ImageNet classes are available, which are used for baseline experiments that assess the impact of known and unknown image artifacts and other conditions on common deep learning-based object classification approaches. Further, current image restoration and enhancement techniques are evaluated by determining whether or not they improve baseline classification performance. Results show that there is plenty of room for algorithmic innovation, making this dataset a useful tool going forward. Rosaura G. VidalMata, Sreya Banerjee, Klemen Grm, Vitomir Struc, Walter J. Scheirer |
WACV | 5 |
| 2018 | Convolutional Neural Networks for Subjective Face Attributes
Mel McCurrie, Fernando Beletti, Lucas Parzianello, Allen Westendorp, Samuel E. Anthony, Walter J. Scheirer |
Image Vis. Comput. | 6 |
| 2018 | The Extreme Value MachineabstractIt is often desirable to be able to recognize when inputs to a recognition function learned in a supervised manner correspond to classes unseen at training time. With this ability, new class labels could be assigned to these inputs by a human operator, allowing them to be incorporated into the recognition function-ideally under an efficient incremental update mechanism. While good algorithms that assume inputs from a fixed set of classes exist, e.g. , artificial neural networks and kernel machines, it is not immediately obvious how to extend them to perform incremental learning in the presence of unknown query classes. Existing algorithms take little to no distributional information into account when learning recognition functions and lack a strong theoretical foundation. We address this gap by formulating a novel, theoretically sound classifier-the Extreme Value Machine (EVM). The EVM has a well-grounded interpretation derived from statistical Extreme Value Theory (EVT), and is the first classifier to be able to perform nonlinear kernel-free variable bandwidth incremental learning. Compared to other classifiers in the same deep network derived feature space, the EVM is accurate and efficient on an established benchmark partition of the ImageNet dataset. Ethan M. Rudd, Lalit P. Jain, Walter J. Scheirer, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Image Provenance Analysis at ScaleabstractPrior art has shown it is possible to estimate, through image processing and computer vision techniques, the types and parameters of transformations that have been applied to the content of individual images to obtain new images. Given a large corpus of images and a query image, an interesting further step is to retrieve the set of original images whose content is present in the query image, as well as the detailed sequences of transformations that yield the query image given the original images. This is a problem that recently has received the name of image provenance analysis. In these times of public media manipulation (e.g., fake news and meme sharing), obtaining the history of image transformations is relevant for fact checking and authorship verification, among many other applications. This article presents an end-to-end processing pipeline for image provenance analysis, which works at real-world scale. It employs a cutting-edge image filtering solution that is custom-tailored for the problem at hand, as well as novel techniques for obtaining the provenance graph that expresses how the images, as nodes, are ancestrally connected. A comprehensive set of experiments for each stage of the pipeline is provided, comparing the proposed solution with state-of-the-art results, employing previously published datasets. In addition, this work introduces a new dataset of real-world provenance cases from the social media site Reddit, along with baseline results. Aparna Bharati, Joel Brogan, Allan Pinto, Michael Parowski, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
IEEE Trans. Image Process. | 8 |
| 2017 | Predicting First Impressions With Deep LearningabstractDescribable visual facial attributes are now commonplacein human biometrics and affective computing, withexisting algorithms even reaching a sufficient point of maturityfor placement into commercial products. These algorithmsmodel objective facets of facial appearance, such as hair andeye color, expression, and aspects of the geometry of theface. A natural extension, which has not been studied toany great extent thus far, is the ability to model subjectiveattributes that are assigned to a face or body based purelyon visual judgements. The fundamental question: how do wecreate models for problems where there is no ground truth,only measurable behavior? In this demo, we show a newconvolutional neural network-based regression framework thatallows us to train predictive models of crowd behavior for socialattribute assignment. Mel McCurrie, Samuel E. Anthony, Walter J. Scheirer |
FG | 3 |
| 2017 | Predicting First Impressions with Deep LearningabstractDescribable visual facial attributes are now commonplace in human biometrics and affective computing, with existing algorithms even reaching a sufficient point of maturityfor placement into commercial products. These algorithms model objective facets of facial appearance, such as hair and eye color, expression, and aspects of the geometry of the face. A natural extension, which has not been studied to any great extent thus far, is the ability to model subjective attributes that are assigned to a face based purely on visual judgements. For instance, with just a glance, our first impression of a face may lead us to believe that a person is smart, worthy of our trust, and perhaps even our admiration - regardless of the underlying truth behind such attributes. Psychologists believe that these judgements are based on a variety of factors such as emotional states, personality traits, and other physiognomic cues. But work in this direction leads to an interesting question: how do we create models for problems where there is no ground truth, only measurable behavior? In this paper, we introduce a convolutional neural network-based regression framework that allows us to train predictive models of crowd behavior for social attribute assignment. Over images from the AFLW face database, these models demonstrate strong correlations with human crowd ratings. Mel McCurrie, Fernando Beletti, Lucas Parzianello, Allen Westendorp, Samuel E. Anthony, Walter J. Scheirer |
FG | 6 |
| 2017 | SREFI: Synthesis of realistic example face imagesabstractIn this paper, we propose a novel face synthesis approach that can generate an arbitrarily large number of synthetic images of both real and synthetic identities. Thus a face image dataset can be expanded in terms of the number of identities represented and the number of images per identity using this approach, without the identity-labeling and privacy complications that come from downloading images from the web. To measure the visual fidelity and uniqueness of the synthetic face images and identities, we conducted face matching experiments with both human participants and a CNN pre-trained on a dataset of 2.6M real face images. To evaluate the stability of these synthetic faces, we trained a CNN model with an augmented dataset containing close to 200,000 synthetic faces. We used a snapshot of this trained CNN to recognize extremely challenging frontal (real) face images. Experiments showed training with the augmented faces boosted the face recognition performance of the CNN. Sandipan Banerjee, John S. Bernhard, Walter J. Scheirer, Kevin W. Bowyer, Patrick J. Flynn |
IJCB | 3 |
| 2017 | U-Phylogeny: Undirected provenance graph construction in the wildabstractDeriving relationships between images and tracing back their history of modifications are at the core of Multimedia Phylogeny solutions, which aim to combat misinformation through doctored visual media. Nonetheless, most recent image phylogeny solutions cannot properly address cases of forged composite images with multiple donors, an area known as multiple parenting phylogeny (MPP). This paper presents a preliminary undirected graph construction solution for MPP, without any strict assumptions. The algorithm is underpinned by robust image representative keypoints and different geometric consistency checks among matching regions in both images to provide regions of interest for direct comparison. The paper introduces a novel technique to geometrically filter the most promising matches as well as to aid in the shared region localization task. The strength of the approach is corroborated by experiments with real-world cases, with and without image distractors (unrelated cases). Aparna Bharati, Daniel Moreira, Allan Pinto, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001 |
ICIP | 7 |
| 2017 | Spotting the difference: Context retrieval and analysis for improved forgery detection and localizationabstractAs image tampering becomes ever more sophisticated and commonplace, the need for image forensics algorithms that can accurately and quickly detect forgeries grows. In this paper, we revisit the ideas of image querying and retrieval to provide clues to better localize forgeries. We propose a method to perform large-scale image forensics on the order of one million images using the help of an image search algorithm and database to gather contextual clues as to where tampering may have taken place. In this vein, we introduce five new strongly invariant image comparison methods and test their effectiveness under heavy noise, rotation, and color space changes. Lastly, we show the effectiveness of these methods compared to passive image forensics using Nimble [1], a new, state-of-the-art dataset from the National Institute of Standards and Technology (NIST). Joel Brogan, Paolo Bestagini, Aparna Bharati, Allan Pinto, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer |
ICIP | 9 |
| 2017 | Provenance filtering for multimedia phylogenyabstractDeparting from traditional digital forensics modeling, which seeks to analyze single objects in isolation, multimedia phylogeny analyzes the evolutionary processes that influence digital objects and collections over time. One of its integral pieces is provenance filtering, which consists of searching a potentially large pool of objects for the most related ones with respect to a given query, in terms of possible ancestors (donors or contributors) and descendants. In this paper, we propose a two-tiered provenance filtering approach to find all the potential images that might have contributed to the creation process of a given query q. In our solution, the first (coarse) tier aims to find the most likely “host” images - the major donor or background - contributing to a composite/doctored image. The search is then refined in the second tier, in which we search for more specific (potentially small) parts of the query that might have been extracted from other images and spliced into the query image. Experimental results with a dataset containing more than a million images show that the two-tiered solution underpinned by the context of the query is highly useful for solving this difficult task. Allan Pinto, Daniel Moreira, Aparna Bharati, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001 |
ICIP | 7 |
| 2017 | Déjà vu: Scalable place recognition using mutually supportive feature frequenciesabstractLearning and recognition is a fundamental process performed in many robot operations such as mapping and localization. The majority of approaches share some common characteristics, such as attempting to extract salient features, landmarks or signatures, and growth in data storage and computational requirements as the size of the environment increases. In biological systems, spatial encoding in the brain is definitively known to be performed using a fixed-size neural encoding framework - the place, head-direction and grid cells found in the mammalian hippocampus and entorhinal cortex. Particularly paradoxically, one of the main encoding centers - the grid cells - represents the world using a highly aliased, repetitive encoding structure where one neuron represents an unbounded number of places in the world. Inspired by this system, in this paper we invert the normal approach used in forming mapping and localization algorithms, by developing a novel place recognition algorithm that seeks out and leverages repetitive, mutually complementary landmark frequencies in the world. The combinatorial encoding capacity of multiple different frequencies enables not only the ability to achieve efficient data storage, but also the potential for sub-linear storage growth in a learning and recall system. Using both ground-based and aerial camera datasets, we demonstrate the system finding and utilizing these frequencies to achieve successful place recognition, and discuss how this approach might scale to arbitrarily large global datasets and dimensions. Adam Jacobson, Walter J. Scheirer, Michael Milford |
IROS | 2 |
| 2017 | Neuron Segmentation Using Deep Complete Bipartite Networks
Jianxu Chen 0001, Sreya Banerjee, Abhinav Grama, Walter J. Scheirer, Danny Ziyi Chen |
MICCAI (2) | 4 |
| 2017 | Authorship Attribution for Social Media ForensicsabstractThe veil of anonymity provided by smartphones with pre-paid SIM cards, public Wi-Fi hotspots, and distributed networks like Tor has drastically complicated the task of identifying users of social media during forensic investigations. In some cases, the text of a single posted message will be the only clue to an author's identity. How can we accurately predict who that author might be when the message may never exceed 140 characters on a service like Twitter? For the past 50 years, linguists, computer scientists, and scholars of the humanities have been jointly developing automated methods to identify authors based on the style of their writing. All authors possess peculiarities of habit that influence the form and content of their written works. These characteristics can often be quantified and measured using machine learning algorithms. In this paper, we provide a comprehensive review of the methods of authorship attribution that can be applied to the problem of social media forensics. Furthermore, we examine emerging supervised learning-based methods that are effective for small sample sizes, and provide step-by-step explanations for several scalable approaches as instructional case studies for newcomers to the field. We argue that there is a significant need in forensics for new authorship attribution algorithms that can exploit context, can process multi-modal data, and are tolerant to incomplete knowledge of the space of all possible authors at training time. Anderson Rocha 0001, Walter J. Scheirer, Christopher W. Forstall, Thiago Cavalcante, Antonio Theophilo, Bingyu Shen 0001, Ariadne Carvalho, Efstathios Stamatatos |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | One-class slab support vector machineabstractThis work introduces the one-class slab SVM (OCSSVM), a one-class classifier that aims at improving the performance of the one-class SVM. The proposed strategy reduces the false positive rate and increases the accuracy of detecting instances from novel classes. To this end, it uses two parallel hyperplanes to learn the normal region of the decision scores of the target class. OCSSVM extends one-class SVM since it can scale and learn non-linear decision functions via kernel methods. The experiments on two publicly available datasets show that OCSSVM can consistently outperform the one-class SVM and perform comparable to or better than other state-of-the-art one-class classifiers. Victor Fragoso, Walter J. Scheirer, João Pedro Hespanha, Matthew Turk 0001 |
ICPR | 2 |
| 2015 | Improving Optimum-Path Forest Classification Using Confidence Measures
Silas Evandro Nachif Fernandes, Walter J. Scheirer, David D. Cox, João Paulo Papa |
CIARP | 2 |
| 2015 | Fine-Tuning Convolutional Neural Networks Using Harmony Search
Gustavo H. Rosa, João Paulo Papa, Aparecido Nilceu Marana, Walter J. Scheirer, David D. Cox |
CIARP | 4 |
| 2015 | Open Set Fingerprint Spoof Detection Across Novel Fabrication MaterialsabstractA fingerprint spoof detector is a pattern classifier that is used to distinguish a live finger from a fake (spoof) one in the context of an automated fingerprint recognition system. Most spoof detectors are learning-based and rely on a set of training images. Consequently, the performance of any such spoof detector significantly degrades when encountering spoofs fabricated using novel materials not found in the training set. In real-world applications, the problem of fingerprint spoof detection must be treated as an open set recognition problem where incomplete knowledge of the fabrication materials used to generate spoofs is present at training time, and novel materials may be encountered during system deployment. To mitigate the security risk posed by novel spoofs, this paper introduces: 1) the use of the Weibull-calibrated SVM (W-SVM), which is relatively robust for open set recognition, as a novel-material detector and a spoof detector and 2) a scheme for the automatic adaptation of the W-SVM-based spoof detector to new spoof materials that leverages interoperability across classifiers. Experiments conducted on new partitions of the LivDet 2011 database designed for open set evaluation suggest: 1) a 97% increase in the error rate of the existing spoof detectors when tested using new spoof materials and 2) up to 44% improvement in spoof detection performance across spoof materials when the proposed adaptive approach is used. Ajita Rattani, Walter J. Scheirer, Arun Ross |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Multi-class Open Set Recognition Using Probability of Inclusion
Lalit P. Jain, Walter J. Scheirer, Terrance E. Boult |
ECCV (3) | 2 |
| 2014 | Condition-invariant, top-down visual place recognitionabstractIn this paper we present a novel, condition-invariant place recognition algorithm inspired by recent discoveries in human visual neuroscience. The algorithm combines intolerant but fast low resolution whole image matching with highly tolerant, sub-image patch matching processes. The approach does not require prior training and works on single images, alleviating the need for either a velocity signal or image sequence, differentiating it from current state of the art methods. We conduct an exhaustive set of experiments evaluating the relationship between place recognition performance and computational resources using part of the challenging Alderley sunny day - rainy night dataset, which has only been previously solved by integrating over 320 frame long image sequences. We achieve recall rates of up to 51% at 100% precision, matching places that have undergone drastic perceptual change while rejecting match hypotheses between highly aliased images of different places. Human trials demonstrate the performance is approaching human capability. The results provide a new benchmark for single image, condition-invariant place recognition. Michael Milford, Walter J. Scheirer, Eleonora Vig, Arren Glover, Oliver Baumann, Jason B. Mattingley, David D. Cox |
ICRA | 2 |
| 2014 | Perceptual Annotation: Measuring Human Vision to Improve Computer VisionabstractFor many problems in computer vision, human learners are considerably better than machines. Humans possess highly accurate internal recognition and learning mechanisms that are not yet understood, and they frequently have access to more extensive training data through a lifetime of unbiased experience with the visual world. We propose to use visual psychophysics to directly leverage the abilities of human subjects to build better machine learning systems. First, we use an advanced online psychometric testing platform to make new kinds of annotation data available for learning. Second, we develop a technique for harnessing these new kinds of information-"perceptual annotations"-for support vector machines. A key intuition for this approach is that while it may remain infeasible to dramatically increase the amount of data and high-quality labels available for the training of a given system, measuring the exemplar-by-exemplar difficulty and pattern of errors of human annotators can provide important information for regularizing the solution of the system at hand. A case study for the problem face detection demonstrates that this approach yields state-of-the-art results on the challenging FDDB data set. Walter J. Scheirer, Samuel E. Anthony, Ken Nakayama, David D. Cox |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Probability Models for Open Set RecognitionabstractReal-world tasks in computer vision often touch upon open set recognition: multi-class recognition with incomplete knowledge of the world and many unknown inputs. Recent work on this problem has proposed a model incorporating an open space risk term to account for the space beyond the reasonable support of known classes. This paper extends the general idea of open space risk limiting classification to accommodate non-linear classifiers in a multiclass setting. We introduce a new open set recognition model called compact abating probability (CAP), where the probability of class membership decreases in value (abates) as points move from known data toward open space. We show that CAP models improve open set recognition for multiple algorithms. Leveraging the CAP formulation, we go on to describe the novel Weibull-calibrated SVM (W-SVM) algorithm, which combines the useful properties of statistical extreme value theory for score calibration with one-class and binary support vector machines. Our experiments show that the W-SVM is significantly better for open set object detection and OCR problems when compared to the state-of-the-art for the same tasks. Walter J. Scheirer, Lalit P. Jain, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Good recognition is non-metric
Walter J. Scheirer, Kimberly Wilber, Michael Eckmann, Terrance E. Boult |
Pattern Recognit. | 1 |
| 2014 | Open set source camera attribution and device linking
Filipe de Oliveira Costa, Ewerton Silva, Michael Eckmann, Walter J. Scheirer, Anderson Rocha 0001 |
Pattern Recognit. Lett. | 4 |
| 2013 | Animal recognition in the Mojave Desert: Vision tools for field biologistsabstractThe outreach of computer vision to non-traditional areas has enormous potential to enable new ways of solving real world problems. One such problem is how to incorporate technology in the effort to protect endangered and threatened species in the wild. This paper presents a snapshot of our interdisciplinary team's ongoing work in the Mojave Desert to build vision tools for field biologists to study the currently threatened Desert Tortoise and Mohave Ground Squirrel. Animal population studies in natural habitats present new recognition challenges for computer vision, where open set testing and access to just limited computing resources lead us to algorithms that diverge from common practices. We introduce a novel algorithm for animal classification that addresses the open set nature of this problem and is suitable for implementation on a smartphone. Further, we look at a simple model for object recognition applied to the problem of individual species identification. A thorough experimental analysis is provided for real field data collected in the Mojave desert. Kimberly Wilber, Walter J. Scheirer, Phil Leitner, Brian Heflin, James Zott, Daniel Reinke, David K. Delaney, Terrance E. Boult |
WACV | 2 |
| 2013 | Toward Open Set RecognitionabstractTo date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This paper explores the nature of open set recognition and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step toward a solution, we introduce a novel "1-vs-set machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale cross-dataset experiments performed over the Caltech 256 and ImageNet sets, as well as face matching experiments performed over the Labeled Faces in the Wild set. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks. Walter J. Scheirer, Anderson Rocha 0001, Archana Sapkota, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Multi-attribute spaces: Calibration for attribute fusion and similarity searchabstractRecent work has shown that visual attributes are a powerful approach for applications such as recognition, image description and retrieval. However, fusing multiple attribute scores - as required during multi-attribute queries or similarity searches - presents a significant challenge. Scores from different attribute classifiers cannot be combined in a simple way; the same score for different attributes can mean different things. In this work, we show how to construct normalized “multi-attribute spaces” from raw classifier outputs, using techniques based on the statistical Extreme Value Theory. Our method calibrates each raw score to a probability that the given attribute is present in the image. We describe how these probabilities can be fused in a simple way to perform more accurate multiattribute searches, as well as enable attribute-based similarity searches. A significant advantage of our approach is that the normalization is done after-the-fact, requiring neither modification to the attribute classification system nor ground truth attribute annotations. We demonstrate results on a large data set of nearly 2 million face images and show significant improvements over prior work. We also show that perceptual similarity of search results increases by using contextual attributes. Walter J. Scheirer, Neeraj Kumar 0006, Peter N. Belhumeur, Terrance E. Boult |
CVPR | 1 |
| 2012 | For your eyes onlyabstractIn this paper, we take a look at an enhanced approach for eye detection under difficult acquisition circumstances such as low-light, distance, pose variation, and blur. We present a novel correlation filter based eye detection pipeline that is specifically designed to reduce face alignment errors, thereby increasing eye localization accuracy and ultimately face recognition accuracy. The accuracy of our eye detector is validated using data derived from the Labeled Faces in the Wild (LFW) and the Face Detection on Hard Datasets Competition 2011 (FDHD) sets. The results on the LFW dataset also show that the proposed algorithm exhibits enhanced performance, compared to another correlation filter based detector, and that a considerable increase in face recognition accuracy may be achieved by focusing more effort on the eye localization stage of the face recognition process. Our results on the FDHD dataset show that our eye detector exhibits superior performance, compared to 11 different state-of-the-art algorithms, on the entire set of difficult data without any per set modifications to our detection or preprocessing algorithms. The immediate application of eye detection is automatic face recognition, though many good applications exist in other areas, including medical research, training simulators, communication systems for the disabled, and automotive engineering. Brian Heflin, Walter J. Scheirer, Terrance E. Boult |
WACV | 2 |
| 2012 | Learning for Meta-RecognitionabstractIn this paper, we consider meta-recognition, an approach for postrecognition score analysis, whereby a prediction of matching accuracy is made from an examination of the tail of the scores produced by a recognition algorithm. This is a general approach that can be applied to any recognition algorithm producing distance or similarity scores. In practice, meta-recognition can be implemented in two different ways: a statistical fitting algorithm based on the extreme value theory, and a machine learning algorithm utilizing features computed from the raw scores. While the statistical algorithm establishes a strong theoretical basis for meta-recognition, the machine learning algorithm is more accurate in its predictions in all of our assessments. In this paper, we present a study of the machine learning algorithm and its associated features for the purpose of building a highly accurate meta-recognition system for security and surveillance applications. Through the use of feature- and decision-level fusion, we achieve levels of accuracy well beyond those of the statistical algorithm, as well as the popular “cohort” model for postrecognition score analysis. In addition, we also explore the theoretical question of why machine learning-based algorithms tend to outperform statistical meta-recognition and provide a partial explanation. We show that our proposed methods are effective for a variety of different recognition applications across security and forensics-oriented computer vision, including biometrics, object recognition, and content-based image retrieval. Walter J. Scheirer, Anderson Rocha 0001, Jonathan Parris, Terrance E. Boult |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Fusing with context: A Bayesian approach to combining descriptive attributesabstractFor identity related problems, descriptive attributes can take the form of any information that helps represent an individual, including age data, describable visual attributes, and contextual data. With a rich set of descriptive at- tributes, it is possible to enhance the base matching accuracy of a traditional face identification system through intelligent score weighting. If we can factor any attribute differences between people into our match score calculation, we can deemphasize incorrect results, and ideally lift the correct matching record to a higher rank position. Naturally, the presence of all descriptive attributes during a match instance cannot be expected, especially when considering non-biometric context. Thus, in this paper, we examine the application of Bayesian Attribute Networks to combine descriptive attributes and produce accurate weighting factors to apply to match scores from face recognition systems based on incomplete observations made at match time. We also examine the pragmatic concerns of attribute network creation, and introduce a Noisy-OR formulation for stream- lined truth value assignment and more accurate weighting. Experimental results show that incorporating descriptive attributes into the matching process significantly enhances face identification over the baseline by up to 32.8%. Walter J. Scheirer, Neeraj Kumar 0006, Karl Ricanek, Peter N. Belhumeur, Terrance E. Boult |
IJCB | 1 |
| 2011 | Meta-Recognition: The Theory and Practice of Recognition Score AnalysisabstractIn this paper, we define meta-recognition, a performance prediction method for recognition algorithms, and examine the theoretical basis for its postrecognition score analysis form through the use of the statistical extreme value theory (EVT). The ability to predict the performance of a recognition system based on its outputs for each match instance is desirable for a number of important reasons, including automatic threshold selection for determining matches and nonmatches, and automatic algorithm selection or weighting for multi-algorithm fusion. The emerging body of literature on postrecognition score analysis has been largely constrained to biometrics, where the analysis has been shown to successfully complement or replace image quality metrics as a predictor. We develop a new statistical predictor based upon the Weibull distribution, which produces accurate results on a per instance recognition basis across different recognition problems. Experimental results are provided for two different face recognition algorithms, a fingerprint recognition algorithm, a SIFT-based object recognition system, and a content-based image retrieval system. Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Robust Fusion: Extreme Value Theory for Recognition Score Normalization
Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
ECCV (3) | 1 |
| 2008 | INSPEC2T: Inexpensive Spectrometer Color Camera TechnologyabstractModern spectrometer equipment tends to be expensive, thus increasing the cost of emerging systems that take advantage of spectral properties as part of their operation. This paper introduces a novel technique that exploits the spectral response characteristics of a traditional sensor (i.e. CMOS or CCD) to utilize it as a low-cost spectrometer. Using the raw Bayer pattern data from a sensor, we estimate the brightness and wavelength of the measured light at a particular point. We use this information to support wide dynamic range, high noise tolerance, and, if sampling takes place on a slope, sub-pixel resolution. Experimental results are provided for both simulation and real data. Further, we investigate the potential of this low-cost technology for spoof detection in biometric systems. Lastly, an actual hardware systhesis is conducted to show the ease with which this algorithm can be implemented onto an FPGA. Walter J. Scheirer, S. R. Kirkbride, Terrance E. Boult |
WACV | 1 |
| 2007 | Revocable Fingerprint Biotokens: Accuracy and Security AnalysisabstractThis paper reviews the biometric dilemma, the pending threat that may limit the long-term value of biometrics in security applications. Unlike passwords, if a biometric database is ever compromised or improperly shared, the underlying biometric data cannot be changed. The concept of revocable or cancelable biometric-based identity tokens (biotokens), if properly implemented, can provide significant enhancements in both privacy and security and address the biometric dilemma. The key to effective revocable biotokens is the need to support the highly accurate approximate matching needed in any biometric system as well as protecting privacy/security of the underlying data. We briefly review prior work and show why it is insufficient in both accuracy and security. This paper adapts a recently introduced approach that separates each datum into two fields, one of which is encoded and one which is left to support the approximate matching. Previously applied to faces, this paper uses this approach to enhance an existing fingerprint system. Unlike previous work in privacy-enhanced biometrics, our approach improves the accuracy of the underlying svstem! The security analysis of these biotokens includes addressing the critical issue of protection of small fields. The resulting algorithm is tested on three different fingerprint verification challenge datasets and shows an average decrease in the Equal Error Rate of over 30% - providing improved security and improved privacy. Terrance E. Boult, Walter J. Scheirer, Robert Woodworth |
CVPR | 2 |
| 2006 | Network intrusion detection with semantics-aware capabilityabstractMalicious network traffic, including widespread worm activity, is a growing threat to Internet-connected networks and hosts. In this paper, we propose a network intrusion detection system (NIDS) with semantics-aware capability. Our NIDS segregates suspicious traffic from the regular traffic flow, extracts binary code from the suspicious traffic, and performs semantic analysis on it to identify potential threats. Our contributions in this work are threefold: (a) we believe our prototype is the first NIDS that provides semantics-aware capability, (b) our implementation is more efficient than what is reported in (M. Christodorescu et al., 2005) (c) our designed templates can capture polymorphic shellcodes with added sequences of stack and mathematic operations. Walter J. Scheirer, Mooi Choo Chuah |
IPDPS | 1 |