VLDB 2026 Research / reviewers in the wild / expert
Terrance E. Boult
dblp:08/6458
· DBLP profile ↗
123ranked-venue papers
17as first author
16since 2021 · last 2025
0000-0001-5007-2529ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 10 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 14 · 4 since 2021Security and privacy · 12 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Computer networks · 4Theory of computation · 3 · 3 first-authorSystems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GHOST: Gaussian Hypothesis Open-Set TechniqueabstractEvaluations of large-scale recognition methods typically focus on overall performance. While this approach is common, it often fails to provide insights into performance across individual classes, which can lead to fairness issues and misrepresentation. Addressing these gaps is crucial for accurately assessing how well methods handle novel or unseen classes and ensuring a fair evaluation. To address fairness in Open-Set Recognition (OSR), we demonstrate that per-class performance can vary dramatically. We introduce Gaussian Hypothesis Open Set Technique (GHOST), a novel hyperparameter-free algorithm that models deep features using class-wise multivariate Gaussian distributions with diagonal covariance matrices. We apply Z-score normalization to logits to mitigate the impact of feature magnitudes that deviate from the model’s expectations, thereby reducing the likelihood of the network assigning a high score to an unknown sample. We evaluate GHOST across multiple ImageNet-1K pre-trained deep networks and test it with four different unknown datasets. Using standard metrics such as AUOSCR, AUROC and FPR95, we achieve statistically significant improvements, advancing the state-of-the-art in large-scale OSR. Source code is provided online. Ryan Rabinowitz, Steve Cruz, Manuel Günther, Terrance E. Boult |
AAAI | 4 |
| 2025 | COSTARR: Consolidated Open Set Technique with Attenuation for Robust RecognitionabstractHandling novelty remains a key challenge in visual recognition systems. Existing open-set recognition (OSR) methods rely on the familiarity hypothesis, detecting novelty by the absence of familiar features. We propose a novel attenuation hypothesis: small weights learned during training attenuate features and serve a dual role-differentiating known classes while discarding information useful for distinguishing known from unknown classes. To leverage this overlooked information, we present COSTARR, a novel approach that combines both the requirement of familiar features and the lack of unfamiliar ones. We provide a probabilistic interpretation of the COSTARR score, linking it to the likelihood of correct classification and belonging in a known class. To determine the individual contributions of the pre- and post-attenuated features to COSTARR's performance, we conduct ablation studies that show both pre-attenuated deep features and the underutilized post-attenuated Hadamard product features are essential for improving OSR. Also, we evaluate COSTARR in a large-scale setting using ImageNet2012-1K as known data and NINCO, iNaturalist, OpenImage-O, and other datasets as unknowns, across multiple modern pre-trained architectures (ViTs, ConvNeXts, and ResNet). The experiments demonstrate that COSTARR generalizes effectively across various architectures and significantly outperforms prior state-of-the-art methods by incorporating previously discarded attenuation information, advancing open-set recognition capabilities. Ryan Rabinowitz, Steve Cruz, Walter J. Scheirer, Terrance E. Boult |
ICCV | 4 |
| 2025 | SMART-vision: survey of modern action recognition techniques in vision
Ali K. AlShami, Ryan Rabinowitz, Khang Nhut Lam, Yousra Shleibik, Melkamu Mersha, Terrance E. Boult, Jugal K. Kalita |
Multim. Tools Appl. | 6 |
| 2024 | Operational Open-Set Recognition and PostMax Refinement
Steve Cruz, Ryan Rabinowitz, Manuel Günther, Terrance E. Boult |
ECCV (6) | 4 |
| 2024 | Watchlist Challenge: 3rd Open-set Face Detection and IdentificationabstractIn the current landscape of biometrics and surveillance, the ability to accurately recognize faces in uncontrolled settings is paramount. The Watchlist Challenge addresses this critical need by focusing on face detection and open-set identification in real-world surveillance scenarios. This paper presents a comprehensive evaluation of participating algorithms, using the enhanced UnConstrained College Students (UCCS) dataset with new evaluation protocols. In total, four participants submitted four face detection and nine open-set face recognition systems. The evaluation demonstrates that while detection capabilities are generally robust, closed-set identification performance varies significantly, with models pre-trained on large-scale datasets showing superior performance. However, open-set scenarios require further improvement, especially at higher true positive identification rates, i.e., lower thresholds. Furkan Kasim, Terrance E. Boult, Rensso Mora Colque, Bernardo Biesseck, Rafael O. Ribeiro, Jan Schlüter, Tomás Repák, Rafael Henrique Vareto, David Menotti, William Robson Schwartz, Manuel Günther |
IJCB | 2 |
| 2024 | Comparative study on chromatin loop callers using Hi-C data reveals their effectivenessabstractAbstract Background Chromosome is one of the most fundamental part of cell biology where DNA holds the hierarchical information. DNA compacts its size by forming loops, and these regions house various protein particles, including CTCF, SMC3, H3 histone. Numerous sequencing methods, such as Hi-C, ChIP-seq, and Micro-C, have been developed to investigate these properties. Utilizing these data, scientists have developed a variety of loop prediction techniques that have greatly improved their methods for characterizing loop prediction and related aspects. Results In this study, we categorized 22 loop calling methods and conducted a comprehensive study of 11 of them. Additionally, we have provided detailed insights into the methodologies underlying these algorithms for loop detection, categorizing them into five distinct groups based on their fundamental approaches. Furthermore, we have included critical information such as resolution, input and output formats, and parameters. For this analysis, we utilized the GM12878 Hi-C datasets at 5 KB, 10 KB, 100 KB and 250 KB resolutions. Our evaluation criteria encompassed various factors, including memory usages, running time, sequencing depth, and recovery of protein-specific sites such as CTCF, H3K27ac, and RNAPII. Conclusion This analysis offers insights into the loop detection processes of each method, along with the strengths and weaknesses of each, enabling readers to effectively choose suitable methods for their datasets. We evaluate the capabilities of these tools and introduce a novel Biological, Consistency, and Computational robustness score ( $$BCC_{score}$$ B C C score ) to measure their overall robustness ensuring a comprehensive evaluation of their performance. H. M. A. Mohit Chowdhury, Terrance E. Boult, Oluwatosin Oluwadare |
BMC Bioinform. | 2 |
| 2024 | Open-set face recognition with maximal entropy and Objectosphere loss
Rafael Henrique Vareto, Yu Linghu, Terrance E. Boult, William Robson Schwartz, Manuel Günther |
Image Vis. Comput. | 3 |
| 2023 | DOERS: Distant Observation Enhancement and Recognition SystemabstractIn order to recognize people across long distances and from elevated viewpoints, biometric systems must handle the challenges of imaging through atmospheric turbulence and non-frontal presentations, in addition to the traditional A-PIE challenges of aging, pose, illumination, and expression. While individual biometric modalities such as facial appearance, gait, and whole body appearance each have a role to play, no single modality can address all of these challenges. This paper describes a novel multi-modal biometric recognition system that addresses the challenges of atmospheric turbulence, occlusions, and elevated viewpoints by combining these modalities. We demonstrate our system on both $R G B$ video-based identity verification and both open and closed-world search. Dawei Du, Cole Hill, Gabriel Bertocco, Maurício Pamplona Segundo, Wes Robbins, Brandon RichardWebster, Roderic Collins, Sudeep Sarkar, Terrance E. Boult, Scott McCloskey |
IJCB | 9 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 16 |
| 2023 | CAST: Conditional Attribute Subsampling Toolkit for Fine-grained EvaluationabstractThorough evaluation is critical for developing models that are fair and robust. In this work, we describe the Conditional Attribute Subsampling Toolkit (CAST) for selecting data subsets for fine-grained scientific evaluations. Our toolkit efficiently filters data given an arbitrary number of conditions for metadata attributes. The purpose of the toolkit is to allow researchers to easily to evaluate models on targeted test distributions. The functionality of CAST is demonstrated on the WebFace42M face Recognition dataset. We calculate over 50 attributes for this dataset including race, image quality, facial features, and accessories. Using our toolkit, we create over a hundred test sets conditioned on one or multiple attributes. Results are presented for subsets of various demographics and image quality ranges. Using eleven different subsets, we build a face recognition 1:1 verification benchmark called C11 that exclusively contains pairs that are near the decision threshold. Evaluation on C11 with state-of-the-art methods demonstrates the suitability of the proposed benchmark. The toolkit is publicly available at https://github.com/WesRobbins/CAST. Wes Robbins, Steven Zhou, Aman Bhatta, Chad Mello, Vitor Albiero, Kevin W. Bowyer, Terrance E. Boult |
WACV | 7 |
| 2023 | Pose2Trajectory: Using transformers on body pose to predict tennis player's trajectory
Ali K. AlShami, Terrance E. Boult, Jugal K. Kalita |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Enhanced Performance of Pre-Trained Networks by Matched Augmentation DistributionsabstractThere exists a distribution discrepancy between training and testing, in the way images are fed to modern CNNs. Recent work tried to bridge this gap either by fine-tuning or re-training the network at different resolutions. However retraining a network is rarely cheap and not always viable. To this end, we propose a simple solution to address the train-test distributional shift and enhance the performance of pretrained models - which commonly ship as a package with deep learning platforms e.g., PyTorch. Specifically, we demonstrate that running inference on the center crop of an image is not always the best as important discriminatory information may be cropped-off. Instead we propose to combine results for multiple random crops for a test image. This not only matches the train time augmentation but also provides the full coverage of the input image. We explore combining representation of random crops through averaging at different levels i.e., deep feature level, logit level, and softmax level. We demonstrate that, for various families of modern deep networks, such averaging results in better validation accuracy compared to using a single central crop per image. The softmax averaging results in the best performance for various pre-trained networks without requiring any re-training or fine-tuning whatsoever. On modern GPUs with batch processing, the paper's approach to inference of pre-trained networks, is essentially free as all images in a batch can all be processed at once. Our code is available at: https://github.com/TouqeerAhmad/MID Touqeer Ahmad, Mohsen Jafarzadeh, Akshay Raj Dhamija, Ryan Rabinowitz, Steve Cruz, Chunchun Li, Terrance E. Boult |
IJCNN | 7 |
| 2022 | Open-Set Support Vector MachinesabstractOften, when dealing with real-world recognition problems, we do not need, and often cannot have, knowledge of the entire set of possible classes that might appear during operational testing. In such cases, we need to think of robust classification methods able to deal with the “unknown” and properly reject samples belonging to classes never seen during training. Notwithstanding, existing classifiers to date were mostly developed for the closed-set scenario, i.e., the classification setup in which it is assumed that all test samples belong to one of the classes with which the classifier was trained. In the open-set scenario, however, a test sample can belong to none of the known classes and the classifier must properly reject it by classifying it as unknown. In this work, we extend upon the well-known support vector machines (SVMs) classifier and introduce the open-set SVMs (OSSVMs), which is suitable for recognition in open-set setups. OSSVM balances the empirical risk and the risk of the unknown and ensures that the region of the feature space in which a test sample would be classified as known (one of the known classes) is always bounded, ensuring a finite risk of the unknown. In this work, we also highlight the properties of the SVM classifier related to the open-set scenario, and provide necessary and sufficient conditions for an RBF SVM to have bounded open-space risk. Pedro Ribeiro Mendes Júnior, Terrance E. Boult, Jacques Wainer, Anderson Rocha 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2021 | Towards a Unifying Framework for Formal Theories of NoveltyabstractManaging inputs that are novel, unknown, or out-of-distribution is critical as an agent moves from the lab to the open world. Novelty-related problems include being tolerant to novel perturbations of the normal input, detecting when the input includes novel items, and adapting to novel inputs. While significant research has been undertaken in these areas, a noticeable gap exists in the lack of a formalized definition of novelty that transcends problem domains. As a team of researchers spanning multiple research groups and different domains, we have seen, first hand, the difficulties that arise from ill-specified novelty problems, as well as inconsistent definitions and terminology. Therefore, we present the first unified framework for formal theories of novelty and use the framework to formally define a family of novelty types. Our framework can be applied across a wide range of domains, from symbolic AI to reinforcement learning, and beyond to open world image recognition. Thus, it can be used to help kick-start new research efforts and accelerate ongoing work on these important novelty-related problems. Terrance E. Boult, Przemyslaw A. Grabowicz, Derek S. Prijatelj, Roni Stern, Lawrence B. Holder, Joshua Alspector, Mohsen Jafarzadeh, Touqeer Ahmad, Akshay Raj Dhamija, Chunchun Li, Steve Cruz, Abhinav Shrivastava, Carl Vondrick, Walter J. Scheirer |
AAAI | 1 |
| 2021 | ComFu: Improving Visual Clustering by Commonality FusionabstractClustering has a long history in the computer vision community with a myriad of applications. Clustering is a family of unsupervised machine learning techniques that group samples based on similarity. Multiple ad hoc techniques have been developed to combine or fuse clustering algorithms with dozens of different clustering techniques. This paper presents a new formalization of clustering fusion and introduces the novel Commonality Fusion (ComFu) technique to combine the advantages of different clustering algorithms by fusing their results on datasets. ComFu builds a pairwise commonality matrix of samples by computing how many clustering algorithms group each pair together. Using this matrix, ComFu builds initial clusters of points with high commonality and then assigns points with low commonality to clusters with the highest average commonality to those points with an automatic distance measure selection process. We start experiments by comparing ComFu with the prior state-of-the-art cluster fusion algorithms on eight UCI datasets. We then evaluate ComFu on practical vision clustering problems, advancing the state-of-the-art on a wide range of applications including clustering faces in the IJB-B dataset. We apply ComFu to fuse FINCH, the state-of-the-art ”parameter-free” approach, which returns multiple partitions and can use multiple distance metrics, and show that ComFu improves their result by fusing over metrics and partitions. Chunchun Li, Manuel Günther, Terrance E. Boult |
ICMLA | 3 |
| 2021 | Automatic Open-World Reliability AssessmentabstractImage classification in the open-world must handle out-of-distribution (OOD) images. Systems should ideally reject OOD images, or they will map atop of known classes and reduce reliability. Using open-set classifiers that can reject OOD inputs can help. However, optimal accuracy of open-set classifiers depend on the frequency of OOD data. Thus, for either standard or open-set classifiers, it is important to be able to determine when the world changes and increasing OOD inputs will result in reduced system reliability. However, during operations, we cannot directly assess accuracy as there are no labels. Thus, the reliability assessment of these classifiers must be done by human operators, made more complex because networks are not 100% accurate, so some failures are to be expected. To automate this process, herein, we formalize the open-world recognition reliability problem and propose multiple automatic reliability assessment policies to address this new problem using only the distribution of reported scores/probability data. The distributional algorithms can be applied to both classic classifiers with SoftMax as well as the open-world Extreme Value Machine (EVM) to provide automated reliability assessment. We show that all of the new algorithms significantly outperform detection using the mean of SoftMax. Mohsen Jafarzadeh, Touqeer Ahmad, Akshay Raj Dhamija, Chunchun Li, Steve Cruz, Terrance E. Boult |
WACV | 6 |
| 2020 | Enhancing Open-Set Recognition using Clustering-based Extreme Value Machine (C-EVM)abstractIn real-world deployments, machine learning applications find challenges when accessing ever-increasing volumes of data - the real world is open and often presents data from classes not seen in training. Open-set recognition is a growing area of machine learning addressing such problems. This research work advances the state-of-the-art in open-set recognition, the Extreme Value Machine (EVM), with a novel clustering-based extension (C-EVM) during training to improve the end-to-end prediction performance. The C-EVM combines Density-based spatial clustering of applications with noise (DBSCAN)-based clustering with a novel Nearby Clusters (NC) algorithm during model fitting to reduce computation while improving accuracy. Our experiments show a statistically significant improvement of 5-10% in macro F1-score over the state-of-the-art EVM on open-set testing using the KDD CUP-99 data set. Past work on open set recognition often traded improved open-set robustness for a decrease in closed-set accuracy, whereas C-EVM outperforms the EVM in both closed-set and open-set recognition. Testing on subsets of ImageNet-2012 with varying numbers of classes, the C-EVM statistically significantly out performs EVM when using deep features. A parameterless Hierarchical DBSCAN (HDBSCAN)-based C-EVM variant is introduced as part of this work that scales well for large data sets. Finally, both EVM and C-EVM can operate as kernel-free incremental learners, enabling these open-set multi-class classifiers to be useful for streaming and big data applications. James Henrydoss, Steve Cruz, Chunchun Li, Manuel Günther, Terrance E. Boult |
IEEE BigData | 5 |
| 2020 | Infoprint: Information Theoretic Digital Image ForensicsabstractTampered images pose a serious predicament since digitized media is a ubiquitous part of our lives. These are facilitated by the availability of image editing software and recent advances in deep Generative Adversarial Networks (GANs). We propose an innovative method to formulate the problem of 10-calizing manipulated regions in fake images as a deep representation learning problem using the Information Bottleneck (IB) principle. We devise a convolutional neural net-based architecture, InfoPrint (IP), that uses variational inference to approximate the IB formulation. Testing on three standard datasets, we demonstrate that InfoPrint outperforms the state-of-the-art by 3% points or more. Additionally, we demonstrate that it has the ability to to detect alterations made by inpainting GANs. Aurobrata Ghosh, Steve Cruz, Subbu Veeravasarapu, Maneesh Kumar Singh 0001, Terrance E. Boult |
ICIP | 6 |
| 2020 | The Overlooked Elephant of Object Detection: Open SetabstractEven though object detection is a popular area of research that has found considerable applications in the real world, it has some fundamental aspects that have never been formally discussed and experimented. One of the core aspects of evaluating object detectors has been the ability to avoid false detections. While major datasets like PASCAL VOC or MSCOCO extensively test the detectors on their ability to avoid false positives, they do not differentiate between their closed-set and open-set performance. Despite systems being trained to reject everything other than the classes of interest, unknown objects from the open world end up being incorrectly detected as known objects, often with very high confidence. This paper is the first to formalize the problem of open-set object detection and propose the first open-set object detection protocol. Moreover, the paper provides a new evaluation metric to analyze the performance of some state-of-the-art detectors and discusses their performance differences. Akshay Raj Dhamija, Manuel Günther, Jonathan Ventura, Terrance E. Boult |
WACV | 4 |
| 2020 | I-MOVE: Independent Moving Objects for Velocity Estimation
Jonathan Schwan, Akshay Raj Dhamija, Terrance E. Boult |
WACV | 3 |
| 2019 | Learning and the Unknown: Surveying Steps toward Open World RecognitionabstractAs science attempts to close the gap between man and machine by building systems capable of learning, we must embrace the importance of the unknown. The ability to differentiate between known and unknown can be considered a critical element of any intelligent self-learning system. The ability to reject uncertain inputs has a very long history in machine learning, as does including a background or garbage class to account for inputs that are not of interest. This paper explains why neither of these is genuinely sufficient for handling unknown inputs – uncertain is not unknown, and unknowns need not appear to be uncertain to a learning system. The past decade has seen the formalization and development of many open set algorithms, which provably bound the risk from unknown classes. We summarize the state of the art, core ideas, and results and explain why, despite the efforts to date, the current techniques are genuinely insufficient for handling unknown inputs, especially for deep networks. Terrance E. Boult, Steve Cruz, Akshay Raj Dhamija, Manuel Günther, James Henrydoss, Walter J. Scheirer |
AAAI | 1 |
| 2019 | Facial attributes: Accuracy and adversarial robustness
Andras Rozsa, Manuel Günther, Ethan M. Rudd, Terrance E. Boult |
Pattern Recognit. Lett. | 4 |
| 2018 | A Multimodal Approach for Predicting Changes in PTSD Symptom SeverityabstractThe rising prevalence of mental illnesses is increasing the demand for new digital tools to support mental wellbeing. Numerous collaborations spanning the fields of psychology, machine learning and health are building such tools. Machine-learning models that estimate effects of mental health interventions currently rely on either user self-reports or measurements of user physiology. In this paper, we present a multimodal approach that combines self-reports from questionnaires and skin conductance physiology in a web-based trauma-recovery regime. We evaluate our models on the EASE multimodal dataset and create PTSD symptom severity change estimators at both total and cluster-level. We demonstrate that modeling the PTSD symptom severity change at the total-level with self-reports can be statistically significantly improved by the combination of physiology and self-reports or just skin conductance measurements. Our experiments show that PTSD symptom cluster severity changes using our novel multimodal approach are significantly better modeled than using self-reports and skin conductance alone when extracting skin conductance features from triggers modules for avoidance, negative alterations in cognition & mood and alterations in arousal & reactivity symptoms, while it performs statistically similar for intrusion symptom. Adria Mallol-Ragolta, Svati Dhamija, Terrance E. Boult |
ICMI | 3 |
| 2018 | Reducing Network AgnostophobiaabstractAgnostophobia, the fear of the unknown, can be experienced by deep learning engineers while applying their networks to real-world applications. Unfortunately, network behavior is not well defined for inputs far from a networks training set. In an uncontrolled environment, networks face many instances that are not of interest to them and have to be rejected in order to avoid a false positive. This problem has previously been tackled by researchers by either a) thresholding softmax, which by construction cannot return "none of the known classes", or b) using an additional background or garbage class. In this paper, we show that both of these approaches help, but are generally insufficient when previously unseen classes are encountered. We also introduce a new evaluation metric that focuses on comparing the performance of multiple approaches in scenarios where such unseen classes or unknowns are encountered. Our major contributions are simple yet effective Entropic Open-Set and Objectosphere losses that train networks using negative samples from some classes. These novel losses are designed to maximize entropy for unknown inputs while increasing separation in deep feature space by modifying magnitudes of known and unknown samples. Experiments on networks trained to classify classes from MNIST and CIFAR-10 show that our novel loss functions are significantly better at dealing with unknown inputs from datasets such as Devanagari, NotMNIST, CIFAR-100 and SVHN. Akshay Raj Dhamija, Manuel Günther, Terrance E. Boult |
NeurIPS | 3 |
| 2018 | Chainlets: A New Descriptor for Detection and RecognitionabstractDetecting and recognizing objects in images is one of the most challenging tasks in computer vision, as it seeks to detect subtle objects while ignoring massive numbers of negatives. While deep networks have led to advances in many problems, new representations and approaches are needed for applications without millions of training samples or where explanations are required. This paper focuses on a new representation that can be used for detection/recognition in many applications of computer vision, and demonstrates it on two very different applications: pedestrian detection and ear recognition. This paper proposes the use of Chainlets, ordered oriented data computed from deep contourbased edge detection, as a novel object descriptor. Chainlets address the problem with Histograms of Oriented Gradients, in that HOG does not model edge connectedness. We extend HOG using Histograms of Chain Codes, which improve object descriptiveness and can even provide orientation invariance. These descriptors significantly outperform existing feature sets, including both existing hand-crafted and deep features for human ear recognition, and are near state of the art on pedestrian detection. Results from our Chainlets algorithm underwent independent testing as part of the new Unconstrained Ear Recognition Challenge dataset, where the competition's evaluation showed Chainlets yielded a significant improvement over other state of the art approaches. To show further generality, we performed an evaluation on the INRIA person detection dataset with results that are near state-of-the-art deep network and boosted classifier results. Overall, the experimental results show that the novel Chainlets representation is competitive with, or better than, state-of-the-art algorithms on both pedestrian detection and ear recognition applications. Adil M. Ahmad, Daniel Lemmond, Terrance E. Boult |
WACV | 3 |
| 2018 | Automated Action Units Vs. Expert Raters: Face offabstractUser engagement is an essential component of any application design. Finding reliable methods to forecast continuos engagement can aid in creating adaptive applications like web-based interventions, intelligent student tutoring, the creation of socially intelligent human-robots, etc. In this paper, we compare observational estimates from expert raters to vision-based learning, for estimating user engagement. The vision-based approach uses automated computation of Action Units combined with an RNN. Several data collection techniques have been explored in the past that capture different modalities for engagement from obtaining self-reports and gathering external observations via crowd-sourcing or even trained expert raters. Traditional machine learning approaches discard annotations from inconsistent raters, use rater averages or apply raterspecific weighting schemes. Such approaches often end up throwing away expensive annotations. We introduce a novel approach that exploits the inherent confusion and disagreement in raters annotations to build a scalable engagement estimation model that learns to appropriately weigh subjective behavioral cues. We show that actively modeling the uncertainty, either explicitly from expert raters or from automated estimation with AU, significantly improves prediction over prediction from just the average engagement ratings. Our approach performs significantly better or on par with experts in predicting engagement for a trauma-recovery application. Svati Dhamija, Terrance E. Boult |
WACV | 2 |
| 2018 | ECLIPSE: Ensembles of Centroids Leveraging Iteratively Processed Spatial Eclipse ClusteringabstractClustering is an unsupervised technique for machine learning and data analysis. Different clustering methods such as centroid, connectivity, density, or distribution-based clustering have been applied as a step in many vision applications. Recently, face clustering has become an important task in the face recognition field, and evaluation benchmarks on the LFW and IJB-B datasets have been created. In this paper, we present the Ensembles of Centroids Leveraging Iteratively Processed Spatial Eclipse (ECLIPSE) clustering algorithm, where we combine the advantages of centroid, density, and connectivity-based clustering algorithms. We show that ECLIPSE can work with most kinds of distance measures such as Euclidean, Cosine, and Bray-Curtis distance. We present the Alignment-Free Facial Feature Extraction (AFFFE) network to extract deep features for the LFW and IJB-B datasets. Using these features, our experimental results show that ECLIPSE can estimate the true number of clusters better than related algorithms and delivers state-of-the-art clustering results, especially for large datasets. Using only the hint in the IJB-B protocol, AFFFE and ECLIPSE significantly advance the state of the art. Chunchun Li, Manuel Günther, Terrance E. Boult |
WACV | 3 |
| 2018 | Towards Robust Deep Neural Networks with BANGabstractMachine learning models, including state-of-the-art deep neural networks, are vulnerable to small perturbations that cause unexpected classification errors. This unexpected lack of robustness raises fundamental questions about their generalization properties and poses a serious concern for practical deployments. As such perturbations can remain imperceptible - the formed adversarial examples demonstrate an inherent inconsistency between vulnerable machine learning models and human perception - some prior work casts this problem as a security issue. Despite the significance of the discovered instabilities and ensuing research, their cause is not well understood and no effective method has been developed to address the problem. In this paper, we present a novel theory to explain why this unpleasant phenomenon exists in deep neural networks. Based on that theory, we introduce a simple, efficient, and effective training approach, Batch Adjusted Network Gradients (BANG), which significantly improves the robustness of machine learning models. While the BANG technique does not rely on any form of data augmentation or the utilization of adversarial images for training, the resultant classifiers are more resistant to adversarial perturbations while maintaining or even enhancing the overall classification performance. Andras Rozsa, Manuel Günther, Terrance E. Boult |
WACV | 3 |
| 2018 | The Extreme Value MachineabstractIt is often desirable to be able to recognize when inputs to a recognition function learned in a supervised manner correspond to classes unseen at training time. With this ability, new class labels could be assigned to these inputs by a human operator, allowing them to be incorporated into the recognition function-ideally under an efficient incremental update mechanism. While good algorithms that assume inputs from a fixed set of classes exist, e.g. , artificial neural networks and kernel machines, it is not immediately obvious how to extend them to perform incremental learning in the presence of unknown query classes. Existing algorithms take little to no distributional information into account when learning recognition functions and lack a strong theoretical foundation. We address this gap by formulating a novel, theoretically sound classifier-the Extreme Value Machine (EVM). The EVM has a well-grounded interpretation derived from statistical Extreme Value Theory (EVT), and is the first classifier to be able to perform nonlinear kernel-free variable bandwidth incremental learning. Compared to other classifiers in the same deep network derived feature space, the EVM is accurate and efficient on an established benchmark partition of the ImageNet dataset. Ethan M. Rudd, Lalit P. Jain, Walter J. Scheirer, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Automated mood-aware engagement predictionabstractDeveloping intelligent machines that recognize facial expressions, detect spontaneous emotions and infer affective states of an individual are all challenging problems. While significant amount of work in recent years has focussed on advancing machine learning techniques for affect recognition and affect classification, the prediction of mood from facial analysis and the usage of mood data have received less attention. Questionnaires for psychometric measurement of mood-states are common, but using them during interventions that target psychological well-being of people are arduous and may burden an already troubled population. In this work, we present mood prediction as a sequence learning problem that uses facial Action Units (AUs) as inputs to a Long Short-Term Memory (LSTM) machine. We create two separate automated LSTM models - a total mood disturbance predictor and a mood sub-scale predictor, and then use them to aid behavioral assessments of engagement. Our mood-aware engagement predictor uses total mood disturbance score, and our analysis compares both mood sub-scale predictors and an overall mood disturbance predictor for engagement prediction. We evaluate our mood models on a large scale dataset consisting of 8M+ frames from multiple videos collected from 110 subjects during a web-intervention for trauma recovery. Our experiments show that mood-aware engagement predictor using our novel visual analysis approach performs significantly better or on par with using self-reports. Svati Dhamija, Terrance E. Boult |
ACII | 2 |
| 2017 | Adversarial Robustness: Softmax versus Openmax
Andras Rozsa, Manuel Günther, Terrance E. Boult |
BMVC | 3 |
| 2017 | Cross-Modal Facial Attribute Recognition with Geometric FeaturesabstractWe propose a purely geometric approach to facial attribute recognition which has better cross-modal performance than a state-of-the-art appearance-based method. While labeled color imagery is plentiful for facial attribute learning, labeled imagery in other modalities such as infrared is comparatively rare. Because face appearance is significantly altered in infrared imagery, standard attribute recognition methods trained on color imagery may not transfer well. To address this problem, we propose attribute recognition based purely on geometric information, i.e. geometric relationships derived from a facial landmark detector. We show that our method outperforms a state-of-the-art appearance-based method in attribute recognition when both are trained on color images and tested on infrared images. Chloe Bradley, Terrance E. Boult, Jonathan Ventura |
FG | 2 |
| 2017 | The unconstrained ear recognition challengeabstractIn this paper we present the results of the Unconstrained Ear Recognition Challenge (UERC), a group benchmarking effort centered around the problem of person recognition from ear images captured in uncontrolled conditions. The goal of the challenge was to assess the performance of existing ear recognition techniques on a challenging large-scale dataset and identify open problems that need to be addressed in the future. Five groups from three continents participated in the challenge and contributed six ear recognition techniques for the evaluation, while multiple baselines were made available for the challenge by the UERC organizers. A comprehensive analysis was conducted with all participating approaches addressing essential research questions pertaining to the sensitivity of the technology to head rotation, flipping, gallery size, large-scale recognition and others. The top performer of the UERC was found to ensure robust performance on a smaller part of the dataset (with 180 subjects) regardless of image characteristics, but still exhibited a significant performance drop when the entire dataset comprising 3,704 subjects was used for testing. Ziga Emersic, Dejan Stepec, Vitomir Struc, Peter Peer, Anjith George, Adil M. Ahmad, Elshibani Omar, Terrance E. Boult, Reza Safdari, Stefanos Zafeiriou, Doggucan Yaman, Fevziye Irem Eyiokur, Hazim Kemal Ekenel |
IJCB | 8 |
| 2017 | Unconstrained Face Detection and Open-Set Face Recognition ChallengeabstractFace detection and recognition benchmarks have shifted toward more difficult environments. The challenge presented in this paper addresses the next step in the direction of automatic detection and identification of people from outdoor surveillance cameras. While face detection has shown remarkable success in images collected from the web, surveillance cameras include more diverse occlusions, poses, weather conditions and image blur. Although face verification or closed-set face identification have surpassed human capabilities on some datasets, open-set identification is much more complex as it needs to reject both unknown identities and false accepts from the face detector. We show that unconstrained face detection can approach high detection rates albeit with moderate false accept rates. By contrast, open-set face recognition is currently weak and requires much more attention. Manuel Günther, Peiyun Hu, Christian Herrmann 0001, Chi-Ho Chan, Min Jiang 0003, Shufan Yang, Akshay Raj Dhamija, Deva Ramanan, Jürgen Beyerer, Josef Kittler, Mohamad Al Jazaery, Mohammad Iqbal Nouyed, Guodong Guo, Cezary Stankiewicz, Terrance E. Boult |
IJCB | 15 |
| 2017 | AFFACT: Alignment-free facial attribute classification techniqueabstractFacial attributes are soft-biometrics that allow limiting the search space, e.g., by rejecting identities with non-matching facial characteristics such as nose sizes or eyebrow shapes. In this paper, we investigate how the latest versions of deep convolutional neural networks, ResNets, perform on the facial attribute classification task. We test two loss functions: the sigmoid cross-entropy loss and the Euclidean loss, and find that for classification performance there is little difference between these two. Using an ensemble of three ResNets, we obtain the new state-of-the-art facial attribute classification error of 8.00 % on the aligned images of the CelebA dataset. More significantly, we introduce the Alignment-Free Facial Attribute Classification Technique (AFFACT), a data augmentation technique that allows a network to classify facial attributes without requiring alignment beyond detected face bounding boxes. To our best knowledge, we are the first to report similar accuracy when using only the detected bounding boxes - rather than requiring alignment based on automatically detected facial landmarks - and who can improve classification accuracy with rotating and scaling test images. We show that this approach outperforms the CelebA baseline on unaligned images with a relative improvement of 36.8 %. Manuel Günther, Andras Rozsa, Terrance E. Boult |
IJCB | 3 |
| 2017 | LOTS about attacking deep featuresabstractDeep neural networks provide state-of-the-art performance on various tasks and are, therefore, widely used in real world applications. DNNs are becoming frequently utilized in biometrics for extracting deep features, which can be used in recognition systems for enrolling and recognizing new individuals. It was revealed that deep neural networks suffer from a fundamental problem, namely, they can unexpectedly misclassify examples formed by slightly perturbing correctly recognized inputs. Various approaches have been developed for generating these so-called adversarial examples, but they aim at attacking end-to-end networks. For biometrics, it is natural to ask whether systems using deep features are immune to or, at least, more resilient to attacks than end-to-end networks. In this paper, we introduce a general technique called the layerwise origin-target synthesis (LOTS) that can be efficiently used to form adversarial examples that mimic the deep features of the target. We analyze and compare the adversarial robustness of the end-to-end VGG Face network with systems that use Euclidean or cosine distance between gallery templates and extracted deep features. We demonstrate that iterative LOTS is very effective and show that systems utilizing deep features are easier to attack than the end-to-end network. Andras Rozsa, Manuel Günther, Terrance E. Boult |
IJCB | 3 |
| 2017 | Incremental Open Set Intrusion Recognition Using Extreme Value MachineabstractTypically, most network intrusion detection systems use supervised learning techniques to identify network anomalies. A problem exists when identifying the unknowns and automatically updating a classifier with new query classes. This is defined as an open set incremental learning problem and we propose to extend a recently introduced method, the Extreme Value Machine (EVM) to address the issue of identifying new classes during query time. The EVM is derived from the statistical extreme value theory and is the first classifier that can perform kernel-free, nonlinear, variable bandwidth outlier detection combined with incremental learning. In this paper, we utilize the EVM for intrusion detection and measure the open set recognition performance of identifying known and unknown classes. Additionally, we evaluate the performance on the KDDCUP'99 dataset and compare the results with the state-of-the-art Weibull-SVM (W-SVM). Our findings demonstrate that the EVM mirrors the performance of the W-SVM classifier, while it supports incremental learning. James Henrydoss, Steve Cruz, Ethan M. Rudd, Manuel Günther, Terrance E. Boult |
ICMLA | 5 |
| 2016 | Towards Application-centric Fairness in Multi-tenant Clouds with Adaptive CPU Sharing ModelabstractThe performance of cloud application is often quite disappointing due to unmanaged consolidation. Therefore, efforts are required to reduce co-tenants interference and provide predictable application performance in multi-tenant cloud environments. In this paper, we examined the complex interplay among cloud tenants as they compete for CPU time, and shared hardware resources. We propose Adaptive CPU Sharing (ACS) approach that reduces co-tenants interference and provides predictable application performance. Our approach is to monitor the progress of submitted applications at runtime, tracks the slowdown of individual application and applies adjustment until convergence. Thus, when an application suffered more slowdown, we allocate more CPU to reduce unfairness. In establishing system support for fine-grained profiling, we report system level activities at sub-second granularity. We predicted application performance degradation by creating a mathematical relationship between high-level application performance and low-level machine events (i.e., CPU steal time and L2 caches miss rate). We validate the added value of our approach by comparing application performance slowdowns (average) with various datasets. Based on our experimental results, our approach helps mitigate co-tenant interference and reduces unfairness by minimizing the overall application slowdowns. Anthony O. Ayodele, Jia Rao, Terrance E. Boult |
CLOUD | 3 |
| 2016 | Automated big security text pruning and classificationabstractMany security related big data problems, including document, traffic, and system log analysis require analysis of unstructured text. Consider the task of analyzing company documents for secure storage. Some might be too sensitive to put on a public cloud and require private storage with associated backup overhead, some may safe on the cloud in encrypted form, and some may be sufficiently non-sensitive to be stored on the cloud in plain-text without encryption and decryption overhead. Being able to make such categorizations autonomously can significantly strengthen data security, organization, and storage efficiency. In this paper, we analyze several base machine learning based security risk assessment algorithms and develop techniques to improve upon standard algorithms. In particular, we examine labeling document sensitivity, labeling each paragraph in the document with one of three levels of security risk. For evaluation, we use real sensitive texts, from documents leaked by the WikiLeaks organization. We improve upon the base models using probabilistic topic modeling via Latent Dirichlet Analysis to identify samples from impure subtopics in the training set, prior to training a logistic regression classifier. Khudran Alzhrani, Ethan M. Rudd, C. Edward Chow, Terrance E. Boult |
IEEE BigData | 4 |
| 2016 | Towards Open Set Deep NetworksabstractDeep networks have produced significant gains for various visual recognition problems, leading to high impact academic and commercial applications. Recent work in deep networks highlighted that it is easy to generate images that humans would never classify as a particular object class, yet networks classify such images high confidence as that given class - deep network are easily fooled with images humans do not consider meaningful. The closed set nature of deep networks forces them to choose from one of the known classes leading to such artifacts. Recognition in the real world is open set, i.e. the recognition system should reject unknown/unseen classes at test time. We present a methodology to adapt deep networks for open set recognition, by introducing a new model layer, OpenMax, which estimates the probability of an input being from an unknown class. A key element of estimating the unknown probability is adapting Meta-Recognition concepts to the activation patterns in the penultimate layer of the network. Open-Max allows rejection of "fooling" and unrelated open set images presented to the system, OpenMax greatly reduces the number of obvious errors made by a deep network. We prove that the OpenMax concept provides bounded open space risk, thereby formally providing an open set recognition solution. We evaluate the resulting open set deep networks using pre-trained networks from the Caffe Model-zoo on ImageNet 2012 validation data, and thousands of fooling and open set images. The proposed OpenMax model significantly outperforms open set recognition accuracy of basic deep networks as well as deep networks with thresholding of SoftMax probabilities. Abhijit Bendale, Terrance E. Boult |
CVPR | 2 |
| 2016 | MOON: A Mixed Objective Optimization Network for the Recognition of Facial Attributes
Ethan M. Rudd, Manuel Günther, Terrance E. Boult |
ECCV (5) | 3 |
| 2016 | Assessing Threat of Adversarial Examples on Deep Neural NetworksabstractDeep neural networks are facing a potential security threat from adversarial examples, inputs that look normal but cause an incorrect classification by the deep neural network. For example, the proposed threat could result in hand-written digits on a scanned check being incorrectly classified but looking normal when humans see them. This research assesses the extent to which adversarial examples pose a security threat, when one considers the normal image acquisition process. This process is mimicked by simulating the transformations that normally occur in of acquiring the image in a real world application, such as using a scanner to acquire digits for a check amount or using a camera in an autonomous car. These small transformations negate the effect of the carefully crafted perturbations of adversarial examples, resulting in a correct classification by the deep neural network. Thus just acquiring the image decreases the potential impact of the proposed security threat. We also show that the already widely used process of averaging over multiple crops neutralizes most adversarial examples. Normal preprocessing, such as text binarization, almost completely neutralizes adversarial examples. This is the first paper to show that for text driven classification, adversarial examples are an academic curiosity, not a security threat. Abigail Graese, Andras Rozsa, Terrance E. Boult |
ICMLA | 3 |
| 2016 | Are Accuracy and Robustness CorrelatedabstractMachine learning models are vulnerable to adversarial examples formed by applying small carefully chosen perturbations to inputs that cause unexpected classification errors. In this paper, we perform experiments on various adversarial example generation approaches with multiple deep convolutional neural networks including Residual Networks, the best performing models on ImageNet Large-Scale Visual Recognition Challenge 2015. We compare the adversarial example generation techniques with respect to the quality of the produced images, and measure the robustness of the tested machine learning models to adversarial examples. Finally, we conduct large-scale experiments on cross-model adversarial portability. We find that adversarial examples are mostly transferable across similar network topologies, and we demonstrate that better machine learning models are less vulnerable to adversarial examples. Andras Rozsa, Manuel Günther, Terrance E. Boult |
ICMLA | 3 |
| 2016 | Are facial attributes adversarially robust?abstractFacial attributes are emerging soft biometrics that have the potential to reject non-matches, for example, based on mismatching gender. To be usable in stand-alone systems, facial attributes must be extracted from images automatically and reliably. In this paper, we propose a simple yet effective solution for automatic facial attribute extraction by training a deep convolutional neural network (DCNN) for each facial attribute separately, without using any pre-training or dataset augmentation, and we obtain new state-of-the-art facial attribute classification results on the CelebA benchmark. To test the stability of the networks, we generated adversarial images - formed by adding imperceptible non-random perturbations to original inputs which result in classification errors - via a novel fast flipping attribute (FFA) technique. We show that FFA generates more adversarial examples than other related algorithms, and that DCNNs for certain attributes are generally robust to adversarial inputs, while DCNNs for other attributes are not. This result is surprising because no DCNNs tested to date have exhibited robustness to adversarial images without explicit augmentation in the training procedure to account for adversarial examples. Finally, we introduce the concept of natural adversarial samples, i.e., images that are misclassified but can be easily turned into correctly classified images by applying small perturbations. We demonstrate that natural adversarial samples commonly occur, even within the training set, and show that many of these images remain misclassified even with additional training epochs. This phenomenon is surprising because correcting the misclassification, particularly when guided by training data, should require only a small adjustment to the DCNN parameters. Andras Rozsa, Manuel Günther, Ethan M. Rudd, Terrance E. Boult |
ICPR | 4 |
| 2016 | Automated big text security classificationabstractIn recent years, traditional cybersecurity safeguards have proven ineffective against insider threats. Famous cases of sensitive information leaks caused by insiders, including the WikiLeaks release of diplomatic cables and the Edward Snowden incident, have greatly harmed the U.S. government's relationship with other governments and with its own citizens. Data Leak Prevention (DLP) is a solution for detecting and preventing information leaks from within an organization's network. However, state-of-art DLP detection models are only able to detect very limited types of sensitive information, and research in the field has been hindered due to the lack of available sensitive texts. Many researchers have focused on document-based detection with artificially labeled “confidential documents” for which security labels are assigned to the entire document, when in reality only a portion of the document is sensitive. This type of whole-document based security labeling increases the chances of preventing authorized users from accessing non-sensitive information within sensitive documents. In this paper, we introduce Automated Classification Enabled by Security Similarity (ACESS), a new and innovative detection model that penetrates the complexity of big text security classification/detection. To analyze the ACESS system, we constructed a novel dataset, containing formerly classified paragraphs from diplomatic cables made public by the WikiLeaks organization. To our knowledge this paper is the first to analyze a dataset that contains actual formerly sensitive information annotated at paragraph granularity. Khudran Alzhrani, Ethan M. Rudd, Terrance E. Boult, C. Edward Chow |
ISI | 3 |
| 2016 | Furthering fingerprint-based authentication: Introducing the true-neighbor templateabstractThis paper introduces the True-Neighbor Template (TNT), a novel, minutiae-only, fingerprint representation and matching approach for authentication. The TNT representation approach overcomes consistency-limitations affecting many existing approaches that stem from relative distortion and spurious minutiae. The TNT matching approach maximizes exploitation of captured fingerprint complexity. A standard benchmark experiment, using the FVC protocol and FVC2006 and FVC2002 databases, is used for evaluation. TNT demonstrates generally-superior neighbor-selection-consistency regarding several established approaches, including: fixed radius, k-nearest neighbors, fixed sectors, and Voronoi diagram. TNT demonstrates generally-superior authentication performance regarding several well-known approaches, including: Bozorth, K-plet, and Minutia Cylinder-Code (MCC). Fawaz E. Alsaadi, Terrance E. Boult |
WACV | 2 |
| 2015 | Performance Measurement and Interference Profiling in Multi-tenant CloudsabstractThe ongoing rush for cloud-based services by small, medium, and large-scale organizations to reduce operational cost and to have more flexibility in the deployment and management of business applications cannot be overemphasized. However, the performance of in-cloud applications is often quite disappointing and unpredictable. Cloud users often perceive the sub optimal and unpredictable performance as anomalies as it is hard to conduct capacity planning based on such unreliable measurements. Performance interference due to resource sharing has been well studied in literature. Representative work includes the study of shared CPU caches, memory bandwidth, hard disks, network bandwidth, and the fair allocation CPU time. There lacks a comprehensive understanding of the complex interplay for shared resource under contention such as CPU. In this research work, we focus on predicting application performance by establishing a mathematical relationship between the high-level performance and the low-level CPU multiplexing. We design a synthetic workload with controllable CPU demands to emulate interference workloads in the cloud. We begin our measurements in a controlled environment to study the impact of CPU allocation on application performance. Based on the results from our experiments, we established the interdependency between CPU steal time, and application performance, and confirms that the percentage of CPU steal time influence application performance, even when workloads of equal parameters were submitted for processing on the same system platform. Our experimental results were evaluated against similar experimental results we performed on Amazon EC2 m3.medium model instance. We confirmed that the workload runtime duration on Amazon EC2 m3. medium model instance are significantly been impacted by high CPU steal time percentage due to interference from co-tenants. Therefore, the workload runtime slowdown percentage on submitted workloads on Amazon EC2.medium model instance is the hidden cost incurred by Cloud subscribers in term of time lost. Cloud service providers should pay close attention to CPU steal time percentage as part of system optimization efforts on Xen based cloud platform. We present Multi-tenant Performance Measurement and Interference Profiling system, a Xen based multi-tenant cloud environment designed for performance measurement and profiling. Anthony O. Ayodele, Jia Rao, Terrance E. Boult |
CLOUD | 3 |
| 2015 | Towards Open World RecognitionabstractWith the of advent rich classification models and high computational power visual recognition systems have found many operational applications. Recognition in the real world poses multiple challenges that are not apparent in controlled lab environments. The datasets are dynamic and novel categories must be continuously detected and then added. At prediction time, a trained system has to deal with myriad unseen categories. Operational systems require minimal downtime, even to learn. To handle these operational issues, we present the problem of Open World Recognition and formally define it. We prove that thresholding sums of monotonically decreasing functions of distances in linearly transformed feature space can balance “open space risk” and empirical risk. Our theory extends existing algorithms for open world recognition. We present a protocol for evaluation of open world recognition systems. We present the Nearest Non-Outlier (NNO) algorithm that evolves model efficiently, adding object categories incrementally while detecting outliers and managing open space risk. We perform experiments on the ImageNet dataset with 1.2M+ images to validate the effectiveness of our method on large scale visual recognition tasks. NNO consistently yields superior results on open world recognition. Abhijit Bendale, Terrance E. Boult |
CVPR | 2 |
| 2014 | Multi-class Open Set Recognition Using Probability of Inclusion
Lalit P. Jain, Walter J. Scheirer, Terrance E. Boult |
ECCV (3) | 3 |
| 2014 | Exemplar codes for facial attributes and tattoo recognitionabstractWhen implementing real-world computer vision systems, researchers can use mid-level representations as a tool to adjust the trade-off between accuracy and efficiency. Unfortunately, existing mid-level representations that improve accuracy tend to decrease efficiency, or are specifically tailored to work well within one pipeline or vision problem at the exclusion of others. We introduce a novel, efficient mid-level representation that improves classification efficiency without sacrificing accuracy. Our Exemplar Codes are based on linear classifiers and probability normalization from extreme value theory. We apply Exemplar Codes to two problems: facial attribute extraction and tattoo classification. In these settings, our Exemplar Codes are competitive with the state of the art and offer efficiency benefits, making it possible to achieve high accuracy even on commodity hardware with a low computational budget. Kimberly Wilber, Ethan M. Rudd, Brian Heflin, Yui-Man Lui, Terrance E. Boult |
WACV | 5 |
| 2014 | Probability Models for Open Set RecognitionabstractReal-world tasks in computer vision often touch upon open set recognition: multi-class recognition with incomplete knowledge of the world and many unknown inputs. Recent work on this problem has proposed a model incorporating an open space risk term to account for the space beyond the reasonable support of known classes. This paper extends the general idea of open space risk limiting classification to accommodate non-linear classifiers in a multiclass setting. We introduce a new open set recognition model called compact abating probability (CAP), where the probability of class membership decreases in value (abates) as points move from known data toward open space. We show that CAP models improve open set recognition for multiple algorithms. Leveraging the CAP formulation, we go on to describe the novel Weibull-calibrated SVM (W-SVM) algorithm, which combines the useful properties of statistical extreme value theory for score calibration with one-class and binary support vector machines. Our experiments show that the W-SVM is significantly better for open set object detection and OCR problems when compared to the state-of-the-art for the same tasks. Walter J. Scheirer, Lalit P. Jain, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Good recognition is non-metric
Walter J. Scheirer, Kimberly Wilber, Michael Eckmann, Terrance E. Boult |
Pattern Recognit. | 4 |
| 2013 | Animal recognition in the Mojave Desert: Vision tools for field biologistsabstractThe outreach of computer vision to non-traditional areas has enormous potential to enable new ways of solving real world problems. One such problem is how to incorporate technology in the effort to protect endangered and threatened species in the wild. This paper presents a snapshot of our interdisciplinary team's ongoing work in the Mojave Desert to build vision tools for field biologists to study the currently threatened Desert Tortoise and Mohave Ground Squirrel. Animal population studies in natural habitats present new recognition challenges for computer vision, where open set testing and access to just limited computing resources lead us to algorithms that diverge from common practices. We introduce a novel algorithm for animal classification that addresses the open set nature of this problem and is suitable for implementation on a smartphone. Further, we look at a simple model for object recognition applied to the problem of individual species identification. A thorough experimental analysis is provided for real field data collected in the Mojave desert. Kimberly Wilber, Walter J. Scheirer, Phil Leitner, Brian Heflin, James Zott, Daniel Reinke, David K. Delaney, Terrance E. Boult |
WACV | 8 |
| 2013 | TPAMI CVPR Special SectionabstractThe articles in this special issue include papers from the CVPR'11 conference which was held in Colorado Spring, CO, June 2011. Pedro F. Felzenszwalb, David A. Forsyth, Pascal Fua, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Toward Open Set RecognitionabstractTo date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This paper explores the nature of open set recognition and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step toward a solution, we introduce a novel "1-vs-set machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale cross-dataset experiments performed over the Caltech 256 and ImageNet sets, as well as face matching experiments performed over the Labeled Faces in the Wild set. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks. Walter J. Scheirer, Anderson Rocha 0001, Archana Sapkota, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Multi-attribute spaces: Calibration for attribute fusion and similarity searchabstractRecent work has shown that visual attributes are a powerful approach for applications such as recognition, image description and retrieval. However, fusing multiple attribute scores - as required during multi-attribute queries or similarity searches - presents a significant challenge. Scores from different attribute classifiers cannot be combined in a simple way; the same score for different attributes can mean different things. In this work, we show how to construct normalized “multi-attribute spaces” from raw classifier outputs, using techniques based on the statistical Extreme Value Theory. Our method calibrates each raw score to a probability that the given attribute is present in the image. We describe how these probabilities can be fused in a simple way to perform more accurate multiattribute searches, as well as enable attribute-based similarity searches. A significant advantage of our approach is that the normalization is done after-the-fact, requiring neither modification to the attribute classification system nor ground truth attribute annotations. We demonstrate results on a large data set of nearly 2 million face images and show significant improvements over prior work. We also show that perceptual similarity of search results increases by using contextual attributes. Walter J. Scheirer, Neeraj Kumar 0006, Peter N. Belhumeur, Terrance E. Boult |
CVPR | 4 |
| 2012 | For your eyes onlyabstractIn this paper, we take a look at an enhanced approach for eye detection under difficult acquisition circumstances such as low-light, distance, pose variation, and blur. We present a novel correlation filter based eye detection pipeline that is specifically designed to reduce face alignment errors, thereby increasing eye localization accuracy and ultimately face recognition accuracy. The accuracy of our eye detector is validated using data derived from the Labeled Faces in the Wild (LFW) and the Face Detection on Hard Datasets Competition 2011 (FDHD) sets. The results on the LFW dataset also show that the proposed algorithm exhibits enhanced performance, compared to another correlation filter based detector, and that a considerable increase in face recognition accuracy may be achieved by focusing more effort on the eye localization stage of the face recognition process. Our results on the FDHD dataset show that our eye detector exhibits superior performance, compared to 11 different state-of-the-art algorithms, on the entire set of difficult data without any per set modifications to our detection or preprocessing algorithms. The immediate application of eye detection is automatic face recognition, though many good applications exist in other areas, including medical research, training simulators, communication systems for the disabled, and automotive engineering. Brian Heflin, Walter J. Scheirer, Terrance E. Boult |
WACV | 3 |
| 2012 | Secure remote matching with privacy: Scrambled support vector vaulted verification (S2V3)abstractAs biometric authentication systems become common in everyday use, researchers are beginning to address privacy issues in biometric recognition. With the growing use of mobile devices, it is important to develop approaches that support remote mobile verification. This paper outlines the need for a mobile/remote SVM-based authentication system that does not compromise the privacy of the subject being recognized. We discuss limitations of earlier privacy-preserving authentication systems and present necessary privacy and security requirements that make a system attractive from both the server's security point of view and from the client's privacy-centric point of view. We then present a novel protocol we call “Vaulted Verification” that allows a server to remotely authenticate a client's biometric in a privacy preserving way. We conclude with a small evaluation of performance, discussion of security implications, and ideas for future work. Kimberly Wilber, Terrance E. Boult |
WACV | 2 |
| 2012 | Learning for Meta-RecognitionabstractIn this paper, we consider meta-recognition, an approach for postrecognition score analysis, whereby a prediction of matching accuracy is made from an examination of the tail of the scores produced by a recognition algorithm. This is a general approach that can be applied to any recognition algorithm producing distance or similarity scores. In practice, meta-recognition can be implemented in two different ways: a statistical fitting algorithm based on the extreme value theory, and a machine learning algorithm utilizing features computed from the raw scores. While the statistical algorithm establishes a strong theoretical basis for meta-recognition, the machine learning algorithm is more accurate in its predictions in all of our assessments. In this paper, we present a study of the machine learning algorithm and its associated features for the purpose of building a highly accurate meta-recognition system for security and surveillance applications. Through the use of feature- and decision-level fusion, we achieve levels of accuracy well beyond those of the statistical algorithm, as well as the popular “cohort” model for postrecognition score analysis. In addition, we also explore the theoretical question of why machine learning-based algorithms tend to outperform statistical meta-recognition and provide a partial explanation. We show that our proposed methods are effective for a variety of different recognition applications across security and forensics-oriented computer vision, including biometrics, object recognition, and content-based image retrieval. Walter J. Scheirer, Anderson Rocha 0001, Jonathan Parris, Terrance E. Boult |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2011 | Face and eye detection on hard datasetsabstractFace and eye detection algorithms are deployed in a wide variety of applications. Unfortunately, there has been no quantitative comparison of how these detectors perform under difficult circumstances. We created a dataset of low light and long distance images which possess some of the problems encountered by face and eye detectors solving real world problems. The dataset we created is composed of reimaged images (photohead) and semi-synthetic heads imaged under varying conditions of low light, atmospheric blur, and distances of 3m, 50m, 80m, and 200m. This paper analyzes the detection and localization performance of the participating face and eye algorithms compared with the Viola Jones detector and four leading commercial face detectors. Performance is characterized under the different conditions and parameterized by per-image brightness and contrast. In localization accuracy for eyes, the groups/companies focusing on long-range face detection outperform leading commercial applications. Jonathan Parris, Kimberly Wilber, Brian Heflin, Ham M. Rara, Ahmed El-Barkouky, Aly A. Farag, Javier R. Movellan, Modesto Castrillón-Santana, Javier Lorenzo-Navarro, Mohammad Nayeem Teli, Sébastien Marcel, Cosmin Atanasoaei, Terrance E. Boult |
IJCB | 14 |
| 2011 | Fusing with context: A Bayesian approach to combining descriptive attributesabstractFor identity related problems, descriptive attributes can take the form of any information that helps represent an individual, including age data, describable visual attributes, and contextual data. With a rich set of descriptive at- tributes, it is possible to enhance the base matching accuracy of a traditional face identification system through intelligent score weighting. If we can factor any attribute differences between people into our match score calculation, we can deemphasize incorrect results, and ideally lift the correct matching record to a higher rank position. Naturally, the presence of all descriptive attributes during a match instance cannot be expected, especially when considering non-biometric context. Thus, in this paper, we examine the application of Bayesian Attribute Networks to combine descriptive attributes and produce accurate weighting factors to apply to match scores from face recognition systems based on incomplete observations made at match time. We also examine the pragmatic concerns of attribute network creation, and introduce a Noisy-OR formulation for stream- lined truth value assignment and more accurate weighting. Experimental results show that incorporating descriptive attributes into the matching process significantly enhances face identification over the baseline by up to 32.8%. Walter J. Scheirer, Neeraj Kumar 0006, Karl Ricanek, Peter N. Belhumeur, Terrance E. Boult |
IJCB | 5 |
| 2011 | Realistic stereo error models and finite optimal stereo baselinesabstractStereo reconstruction is an important research and application area, both for general 3D reconstruction and for operations like robotic navigation and remote sensing. This paper addresses the determination of parameters for a stereo system to optimize/minimize 3D reconstruction errors. Previous work on error analysis in stereo reconstruction optimized error in disparity space which led to the erroneous conclusion that, ignoring matching errors, errors decrease when the baseline goes to infinity. In this paper, we derive the first formal error model based on the more realistic “point-of-closest-approach” ray model used in modern stereo systems. We then show this results in finite optimal baseline that minimizes reconstruction errors in all three world directions. We also show why previous oversimplified error analysis results in infinite baselines. We derive the mathematical relationship between the error variances and the stereo system parameters. In our analysis, we consider the situations where errors exist in only one camera as well as errors in both cameras. We have derived the results for both parallel and verged systems, though only the simpler models are presented algebraically herein. The paper includes simulations to highlight the results and validate the approximations in the error propagation. The results should allow stereo system designers, or those using motion-stereo, to improve their system. Terrance E. Boult |
WACV | 2 |
| 2011 | Meta-Recognition: The Theory and Practice of Recognition Score AnalysisabstractIn this paper, we define meta-recognition, a performance prediction method for recognition algorithms, and examine the theoretical basis for its postrecognition score analysis form through the use of the statistical extreme value theory (EVT). The ability to predict the performance of a recognition system based on its outputs for each match instance is desirable for a number of important reasons, including automatic threshold selection for determining matches and nonmatches, and automatic algorithm selection or weighting for multi-algorithm fusion. The emerging body of literature on postrecognition score analysis has been largely constrained to biometrics, where the analysis has been shown to successfully complement or replace image quality metrics as a predictor. We develop a new statistical predictor based upon the Weibull distribution, which produces accurate results on a per instance recognition basis across different recognition problems. Experimental results are provided for two different face recognition algorithms, a fingerprint recognition algorithm, a SIFT-based object recognition system, and a content-based image retrieval system. Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Robust Fusion: Extreme Value Theory for Recognition Score Normalization
Walter J. Scheirer, Anderson Rocha 0001, Ross J. Micheals, Terrance E. Boult |
ECCV (3) | 4 |
| 2008 | INSPEC2T: Inexpensive Spectrometer Color Camera TechnologyabstractModern spectrometer equipment tends to be expensive, thus increasing the cost of emerging systems that take advantage of spectral properties as part of their operation. This paper introduces a novel technique that exploits the spectral response characteristics of a traditional sensor (i.e. CMOS or CCD) to utilize it as a low-cost spectrometer. Using the raw Bayer pattern data from a sensor, we estimate the brightness and wavelength of the measured light at a particular point. We use this information to support wide dynamic range, high noise tolerance, and, if sampling takes place on a slope, sub-pixel resolution. Experimental results are provided for both simulation and real data. Further, we investigate the potential of this low-cost technology for spoof detection in biometric systems. Lastly, an actual hardware systhesis is conducted to show the ease with which this algorithm can be implemented onto an FPGA. Walter J. Scheirer, S. R. Kirkbride, Terrance E. Boult |
WACV | 3 |
| 2007 | Revocable Fingerprint Biotokens: Accuracy and Security AnalysisabstractThis paper reviews the biometric dilemma, the pending threat that may limit the long-term value of biometrics in security applications. Unlike passwords, if a biometric database is ever compromised or improperly shared, the underlying biometric data cannot be changed. The concept of revocable or cancelable biometric-based identity tokens (biotokens), if properly implemented, can provide significant enhancements in both privacy and security and address the biometric dilemma. The key to effective revocable biotokens is the need to support the highly accurate approximate matching needed in any biometric system as well as protecting privacy/security of the underlying data. We briefly review prior work and show why it is insufficient in both accuracy and security. This paper adapts a recently introduced approach that separates each datum into two fields, one of which is encoded and one which is left to support the approximate matching. Previously applied to faces, this paper uses this approach to enhance an existing fingerprint system. Unlike previous work in privacy-enhanced biometrics, our approach improves the accuracy of the underlying svstem! The security analysis of these biotokens includes addressing the critical issue of protection of small fields. The resulting algorithm is tested on three different fingerprint verification challenge datasets and shows an average decrease in the Equal Error Rate of over 30% - providing improved security and improved privacy. Terrance E. Boult, Walter J. Scheirer, Robert Woodworth |
CVPR | 1 |
| 2007 | PrivacyCam: a Privacy Preserving Camera Using uCLinux on the Blackfin DSPabstractConsiderable research work has been done in the area of surveillance and biometrics, where the goals have always been high performance, robustness in security and cost optimization. With the emergence of more intelligent and complex video surveillance mechanisms, the issue of "privacy invasion" has been looming large. Very little investment or effort has gone into looking after this issue in an efficient and cost-effective way. The process of PICO (privacy through invertible cryptographic obscuration) is a way of using cryptographic techniques and combining them with image processing and video surveillance to provide a practical solution to the critical issue of "privacy invasion". This paper presents the idea and example of a realtime embedded application of the PICO technique, using uCLinux on the tiny Blackfin DSP architecture, along with a small Omnivision camera. It demonstrates how the practical problem of "privacy invasion" can be successfully addressed through DSP hardware in terms of smallness in size and cost optimization. After review of previous applications of "privacy protection", and system components, we discuss the "embedded jpeg-space" detection of regions of interest and the real time application of encryption techniques to improve privacy while allowing general surveillance to continue. The resulting approach permits full access (violation of privacy) only by access to the private-key to recover the decryption key, thereby striking a fine trade-off among privacy, security, cost and space. Ankur Chattopadhyay, Terrance E. Boult |
CVPR | 2 |
| 2007 | Improving Variance Estimation in Biometric SystemsabstractMeasuring system performance seems conceptually straightforward. However, the interpretation of the results and predicting future performance remain as exceptional challenges in system evaluation. Robust experimental design is critical in evaluation, but there have been very few techniques to check designs for either overlooked associations or weak assumptions. For biometric & vision system evaluation, the complexity of the systems make a thorough exploration of the problem space impossible - this lack of verifiability in experimental design is a serious issue. In this paper, we present a new evaluation methodology that improves the accuracy of variance estimator via the discovery of false assumptions about the homogeneity of cofactors - i.e., when the data is not ''well mixed". The new methodology is then applied in the context of a biometric system evaluation with highly influential cofactors. Ross J. Micheals, Terrance E. Boult |
CVPR | 2 |
| 2007 | Systems issues in distributed multi-modal surveillanceabstractTo be viable commercial multi-modal surveillance systems, the systems need to be reliable, robust and must be able to work at night (maybe the most critical time). They must handle small and non-distinctive targets that are as far away as possible. Like other commercial applications, end users of the systems must be able to operate them in a proper way. In this paper, we focus on three significant inherent limitations of current surveillance systems: the effective accuracy at relevant distances, the ability to define and visualize the events on a large scale, and the usability of the system. Terrance E. Boult |
CVPR | 2 |
| 2007 | Two thresholds are better than oneabstractThe concept of the Bayesian optimal single threshold is a well established and widely used classification technique. In this paper, we prove that when spatial cohesion is assumed for targets, a better classification result than the "optimal" single threshold classification can be achieved. Under the assumption of spatial cohesion and certain prior knowledge about the target and background, the method can be further simplified as dual threshold classification. In core-dual threshold classification, spatial cohesion within the target core allows "continuation" linking values to fall between the two thresholds to the target core; classical Bayesian classification is employed beyond the dual thresholds. The core-dual threshold algorithm can be built into a Markov random field model (MRF). From this MRF model, the dual thresholds can be obtained and optimal classification can be achieved. In some practical applications, a simple method called symmetric subtraction may be employed to determine effective dual thresholds in real time. Given dual thresholds, the quasi-connected component algorithm is shown to be a deterministic implementation of the MRF core-dual threshold model combining the dual thresholds, extended neighborhoods and efficient connected component computation. Terrance E. Boult, RC Johnson |
CVPR | 2 |
| 2007 | On Channel Reliability Measure Training for Multi-Camera Face RecognitionabstractSingle-camera face recognition has severe limitations when the subject is not cooperative, or there are pose changes and different illumination conditions. Face recognition using multiple synchronized cameras is proposed to overcome the limitations. We introduce a reliability measure trained from examples to evaluate the inherent quality of channel recognition. The recognition from the channel predicted to be the most reliable is selected as the final recognition results. In this paper, we enhance Adaboost to improve the component based face detector running in each channel as well as the channel reliability measure training. Effective features are designed to train the channel reliability measure using data from both face detection and recognition. The recognition rate is far better than that of either single channel, and consistently better than common classifier fusion rules Binglong Xie, Visvanathan Ramesh, Ying Zhu 0006, Terrance E. Boult |
WACV | 4 |
| 2007 | Hop-count based probabilistic packet dropping: Congestion mitigation with loss rate differentiation
Xiaobo Zhou 0002, Dennis Ippoliti, Terrance E. Boult |
Comput. Commun. | 3 |
| 2007 | Application of Projective Invariants in Hand Geometry BiometricsabstractOur research focuses on finding mathematical representations of biometric features that are not only distinctive, but also invariant to projective transformations. We have chosen hand geometry technology to work with, because it has wide public awareness and acceptance and most important, large space for improvement. Unlike the traditional hand geometry technologies, the hand descriptor in our hand geometry system is constructed using projective-invariant features. Hand identification can be accomplished by a single view of a hand regardless of the viewing angles. The noise immunity and the discriminability possessed by a hand feature vector using different types of projective invariants are studied. We have found an appropriate symmetric polynomial representation of the hand features with which both noise immunity and discrimminability of a hand feature vector are considerably improved. Experimental results show that the system achieves an equal error rate (EER) of 2.1% by a 5-D feature vector on a database of 52 hand images. The EER reduces to 0.00% when the feature vector dimension increases to 18. In this paper, we extend the concept of hand geometry from a geometrical size-based technique that requires physical hand constraints to a projective invariant-based technique that allows free hand motion. Chia-Jiu Wang, Terrance E. Boult |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2006 | HPPD: A Hop-count Probabilistic Packet DropperabstractNetwork applications and users have very diverse service expectations and requirements, demanding for provisioning of different levels of quality of service on the Internet. Packet loss rate differentiation has been an active research topic. However, the existing packet dropping schemes for loss rate differentiation have not considered an important issue, that is, the retransmission overhead of dropped packets. In this paper, we design a hop-count based probabilistic packet dropper (HPPD) for congestion mitigation and loss rate differentiation. HPPD aims to meet a two-fold objective by two-dimensional loss rate differentiation: one is the congestion mitigation that aims to reduce congestion in the first place by dropping intra-class packets differently based on their maturity levels to reduce retransmission cost; the other is inter-class proportional loss rate differentiation. The maturity level of a packet, the number of hops travelled, is inferred from its time-to-live value in the IP header. We propose an intra-class nth-root proportional dropping scheme, where n is a controllable parameter trading off dropping fairness for congestion mitigation. Simulation results show that HPPD can significantly mitigate the congestion by reducing the retransmission overhead of dropped packets and achieve the proportional loss rate differentiation at the same time. Xiaobo Zhou 0002, Dennis Ippoliti, Terrance E. Boult |
ICC | 3 |
| 2006 | Programmable Imaging: Towards a Flexible Camera
Shree K. Nayar, Vlad Branzoi, Terrance E. Boult |
Int. J. Comput. Vis. | 3 |
| 2005 | On the Small Sample Performance of Boosted ClassifiersabstractBoosting algorithms have been widely applied in the machine vision systems. Two fundamental issues that have to be solved in these systems are how much training data and how many Boosting rounds are needed to achieve a desired performance. We view the Boosting algorithm as a nonlinear estimation scheme that estimates a strong classifier from a given training sample set (that is generated by sampling a true unknown distribution), the weak classifiers, and the number of Boosting rounds T. The performance characterization of this estimator involves the derivation of the classification error statistics of the trained strong classifier as a function of the training set and the collection of the weak classifiers. Although the convergence and the error bounds for the training error and generalization error of the algorithms have been studied for several years, the estimated bounds are still loose bounds that are only meaningful for large training sets. With no effective tools for determining the error bounds, users are now collecting training samples with as much data as they can afford, with no good way to know if they are sufficient. In this paper, we characterize the classification error statistics of the trained strong classifier as a function of the true distributions of classes, the collection of the weak classifiers, and the size of the training set. We show that the statistics can be numerically computed and the results are more accurate than previous bounds in the literature. Theoretical results are verified through the simulations. Face detection is used as a case study to illustrate the application of the theory on real data. Weiliang Li, Ying Zhu 0006, Visvanathan Ramesh, Terrance E. Boult |
CVPR (2) | 5 |
| 2004 | Programmable Imaging Using a Digital Micromirror Array
Shree K. Nayar, Vlad Branzoi, Terrance E. Boult |
CVPR (1) | 3 |
| 2004 | Omni-directional visual surveillance
Terrance E. Boult, Ross J. Micheals, Michael Eckmann |
Image Vis. Comput. | 1 |
| 2004 | Sudden illumination change detection using order consistency
Binglong Xie, Visvanathan Ramesh, Terrance E. Boult |
Image Vis. Comput. | 3 |
| 2004 | Lighting sensitive displayabstractAlthough display devices have been used for decades, they have functioned without taking into account the illumination of their environment. We present the concept of a lighting sensitive display (LSD)---a display that measures the incident illumination and modifies its content accordingly. An ideal LSD would be able to measure the 4D illumination field incident upon it and generate a 4D light field in response to the illumination. However, current sensing and display technologies do not allow for such an ideal implementation. Our initial LSD prototype uses a 2D measurement of the illumination field and produces a 2D image in response to it. In particular, it renders a 3D scene such that it always appears to be lit by the real environment that the display resides in. The current system is designed to perform best when the light sources in the environment are distant from the display, and a single user in a known location views the display. The displayed scene is represented by compressing a very large set of images (acquired or rendered) of the scene that correspond to different lighting conditions. The compression algorithm is a lossy one that exploits not only image correlations over the illumination dimensions but also coherences over the spatial dimensions of the image. This results in a highly compressed representation of the original image set. This representation enables us to achieve high quality relighting of the scene in real time. Our prototype LSD can render 640 × 480 images of scenes under complex and varying illuminations at 15 frames per second using a 2 GHz processor. We conclude with a discussion on the limitations of the current implementation and potential areas for future research. Shree K. Nayar, Peter N. Belhumeur, Terrance E. Boult |
ACM Trans. Graph. | 3 |
| 2002 | Statistical Characterization of Morphological Operator Sequences
Visvanathan Ramesh, Terrance E. Boult |
ECCV (4) | 3 |
| 2002 | Dynamic home agent reassignment in Mobile IPabstractUnder conventional Mobile IP, all traffic to a MN is tunneled to the current foreign agent (FA) from the remote HA, which can result in poor efficiency. A dynamic home agent (HA) reassignment model is proposed in this paper to reduce both the signaling traffic and the data traffic to the home network. Simultaneous bindings with two or more HA are supported in our model to provide seamless HA handover. Furthermore, the whole dynamic home agent/home address reassignment procedure is completed in a single registration signaling cycle. Hence it minimizes the delay of the (regional) foreign agent (FA) handoff. Finally, the network access identifier (NAI) or full qualified domain name (FQDN) is used instead of the home address to uniquely identify a MN since the home address is no longer permanent. Terrance E. Boult |
WCNC | 2 |
| 2002 | Guest Editorial: Stereo and Multi-Baseline Vision
Gary R. Bradski, Terrance E. Boult |
Int. J. Comput. Vis. | 2 |
| 2002 | Introduction to the Special Issue on Innovative Applications of Computer Vision
Bir Bhanu, Terrance E. Boult, Alok Gupta, David Michael |
Mach. Vis. Appl. | 2 |
| 2001 | Efficient Evaluation of Classification and Recognition SystemsabstractIn this paper, a new framework for evaluating a variety of computer vision systems and components is introduced. This framework is particularly well suited for domains such as classification or recognition systems, where blind application of the i.i.d. assumption would reduce an evaluation's accuracy, such as with classification or recognition systems. With few exceptions, most previous work on vision system evaluation does not include confidence intervals, since they are difficult to calculate, and are often coupled with strict requirements. We show how a set of previously overlooked replicate statistics tools can be used to obtain tighter confidence intervals of evaluation estimates while simultaneously reducing the amount of data and computation required to reach such sound evaluatory conclusions. In the included application of the new methodology, the well-known FERET face recognition system evaluation is extended to incorporate standard errors and confidence intervals. Ross J. Micheals, Terrance E. Boult |
CVPR (1) | 2 |
| 2001 | Into the woods: visual surveillance of noncooperative and camouflaged targets in complex outdoor settingsabstractAutonomous video surveillance and monitoring of human subjects in video has a rich history. Many deployed systems are able to reliably track human motion in indoor and controlled outdoor environments, e.g., parking lots and university campuses. A challenging domain of vital military importance is the surveillance of noncooperative and camouflaged targets within cluttered outdoor settings. These situations require both sensitivity and a very wide field of view and, therefore, are a natural application of omnidirectional video. Fundamentally, target finding is a change detection problem. Detection of camouflaged and adversarial targets implies the need for extreme sensitivity. Unfortunately, blind change detection in woods and fields may lead to a high fraction of false alarms, since natural scene motion and lighting changes produce highly dynamic scenes. Naturally, this desire for high sensitivity leads to a direct tradeoff between miss detections and false alarms. This paper discusses the current state of the art in video-based target detection, including an analysis of background adaptation techniques. The primary focus of the paper is the Lehigh Omnidirectional Tracking System (LOTS) and its components. This includes adaptive multibackground modeling, quasi-connected components (a novel approach to spatio-temporal grouping), background subtraction analyses, and an overall system evaluation. Terrance E. Boult, Ross J. Micheals, Michael Eckmann |
Proc. IEEE | 1 |
| 2000 | Error Analysis of Background AdaptionabstractBackground modeling is a common component in video surveillance systems and is used to quickly identify regions of interest. To increase the robustness of background subtraction techniques, researchers have developed techniques to update the background model and also developed probabilistic/statistical approaches for thresholding the difference. This paper presents an error analysis of this type of background modeling and pixel labeling, providing both theoretical analysis and experimental validation. Evaluation is centered around the tradeoff of probability of false alarm and probability of miss detection, and this paper shows how to efficiently compute these probabilities front simpler values that are more easily measured. It includes an analysis for both static and dynamic background modeling. The paper also examines the assumptions of Gaussian and mixture of Gaussian models for a pixel. Terrance E. Boult, Frans Coetzee, Visvanathan Ramesh |
CVPR | 2 |
| 2000 | Physical Panoramic Pyramid and Noise Sensitivity in PyramidsabstractMulti-resolution techniques have been used in a wide range of vision applications. Unfortunately, the costly operation of building a proper pyramid strongly reduces its value as a tool for reducing computational cost. A new approach, physical panoramic pyramid, is introduced in this paper. Physical panoramic pyramid measures multiple resolutions simultaneously resulting in multi-resolution panoramic images. No computation is needed to construct these image pyramids. We also analyze general noise sensitivity in image pyramids, including the interaction of the loss of resolution, random background noise and aliasing noise. The paper also discusses the issue of indexing between the neighboring layer the viewpoint variation and the applications of the physical panoramic pyramid. Weihong Yin, Terrance E. Boult |
CVPR | 2 |
| 2000 | Efficient super-resolution via image warpingabstractThis paper introduces a new algorithm for enhancing image resolution from an image sequence. The approach we propose herein uses the integrating resampler for warping. The method is a direct computation, which is fundamentally different from the iterative back-projection approaches proposed in previous work. This paper shows that image-warping techniques may have a strong impact on the quality of image resolution enhancement. By coupling the degradation model of the imaging system directly into the integrating resampler, we can better approximate the warping characteristics of real sensors, which also significantly improve the quality of super-resolution images. Examples of super-resolutions are given for gray-scale images. Evaluations are made visually by comparing the resulting images and those using bi-linear resampling and back-projection and quantitatively using OCR as a fundamental measure. The paper shows that even when the images are qualitatively similar, quantitative differences appear in machine processing. Ming-Chao Chiang, Terrance E. Boult |
Image Vis. Comput. | 2 |
| 1998 | Remote Reality DemonstrationabstractRemote Reality is an approach to providing an immersive environment via omni-directional imaging. The system can use a live video-feed from a remote location or can use recorded data and be remote in both space and time. While less interactive than traditional VR, remote reality has an important advantage: there is little to no need for model building. In addition, the objects, the textures and the motions are not just realistic, they are remote views of reality. Terrance E. Boult |
CVPR | 1 |
| 1998 | Applications of omnidirectional imaging: multi-body tracking and remote realityabstractRecently, S. Nayar (1997) introduced a parabolic imaging system that has a field of view of a full hemisphere or more. When used with a video camera the result is an omni-directional video stream that captures everything going around it. In the VAST Lab at Lehigh, we have been experimenting with these cameras, developing new variants, and developing omni-directional vision applications. We present an overview of omni-directional imaging and then two of our applications which we will be demonstrating. The first application is a frame-rate multi-body tracking system. The system uses an omni-directional imager and a standard PC to track multiple moving objects in all directions. The system is designed to provide perspective views of the most significant targets, either locally or over a network. The second application is something we call Remote Reality, which provides an immersive environment via omnidirectional imaging. It can use live or pre-recorded video from a remote location. While less interactive than traditional VR, remote reality has important advantages: there is no need for "model building" and the objects, textures and motions are not graphical approximations. Terrance E. Boult, Weihong Yin, Ali Erkin, Peter Lewis, Chris Power, Ross J. Micheals |
WACV | 1 |
| 1998 | Catadioptric video sensorsabstractConventional video cameras have limited fields of view which make them restrictive in a variety of applications. A catadioptric sensor uses a combination of lenses and mirrors placed in a carefully arranged configuration to capture a much wider field of view. At Columbia University, we have developed a wide range of catadioptric sensors. Some of these sensors have been designed to produce unusually large fields of view. Others have been constructed for the purpose of depth computation. All our sensors perform in real time using just a PC. Shree K. Nayar, Joshua Gluckman, Rahul Swaminathan, Simon Lok, Terrance E. Boult |
WACV | 5 |
| 1997 | Local Blur Estimation and Super-ResolutionabstractUntil now, all super-resolution algorithms have presumed that the images were taken under the same illumination conditions. This paper introduces a new approach to super-resolution, based on edge models and a local blur estimate, which circumvents these difficulties. The paper presents the theory and the experimental results using the new approach. Ming-Chao Chiang, Terrance E. Boult |
CVPR | 2 |
| 1997 | Separation of Reflection Components Using Color and Polarization
Shree K. Nayar, Xi-Sheng Fang, Terrance E. Boult |
Int. J. Comput. Vis. | 3 |
| 1996 | Global Models with Parametric Offsets as Applied to Cardiac Motion RecoveryabstractWe introduce a new solid shape model formulation that includes built-in offsets from a base global component (e.g. an ellipsoid) which are functions of the global component's parameters. The offsets provide two features. First, they help to form an expected model shape which facilitates appropriate model data correspondences. Second, they scale with the base global model to maintain the expected shape even in the presence of large global deformations. We apply this model formulation to the recovery of 3-D cardiac motion from a volunteer dataset of tagged-MR images. The model instance is a variation of the hybrid volumetric ventriculoid (HVV), a deformable thick-walled ellipsoid model resembling the left ventricle (LV) of the heart. A unique aspect of of implementation is the employment of constant volume constraints when recovering the cardiac motion. In addition, we present a novel geodesic-like prismoidal tessellation of the model which provides for more stable fits. Thomas O'Donnell, Terrance E. Boult, Alok Gupta |
CVPR | 2 |
| 1996 | Efficient image warping and super-resolutionabstractThis paper introduces a new algorithm for enhancing image resolution from an image sequence. The approach we propose herein uses the integrating resampler proposed by M. Chiang and T. Boult (1996) as the underlying resampling algorithm. Moreover, it is a direct method, which is fundamentally different from the iterative, back-projection approaches proposed in previous work. We show that image warping techniques may have a strong impact on the quality of image resolution enhancement. By coupling the degradation model of the imaging system directly into the integrating resampler, we can better approximate the warping characteristics of real sensors, which also highly improve the quality of super-resolution images. Examples of super-resolutions are given for gray-scale images. Evaluations are made by comparing the resulting images and those using bi-linear resampling and back-projection. Results from our experiments show that integrating resampler outperforms traditional bi-linear resampling. Ming-Chao Chiang, Terrance E. Boult |
WACV | 2 |
| 1996 | Recovery of SHGCs From a Single Intensity ViewabstractGeneralized cylinders are a flexible, loosely-defined class of parametric shapes capable of modeling many real-world objects. Straight homogeneous generalized cylinders are an important subclass of generalized cylinders, whose cross-sections are scaled versions of a reference curve. Although there has been considerable research into recovering the shape of SHGCs from their contour, this work has almost exclusively involved methods that couple contour and heuristic constraints. A rigorous approach to the problem of recovering solid parametric shape from a single intensity view should involve at least two stages: (1) deriving the contour constraints, and (2) determining if additional image constraints, e.g., intensity, can be used to uniquely determine the 3D object shape. In this paper, the authors follow the approach just described. This methodology is also important for the recovery of object classes like tubes, where contour and heuristic constraints are shown to be insufficient for shape recovery. First, the authors prove that SHGC contours generated under orthography have exactly two degrees of freedom. Next, the authors show that the remaining free parameters can be resolved using reflectance-based constraints, without knowledge of the number of light sources, their positions, intensities, the amount of ambient light; or the surface albedo. Finally, the reflectance-based recovery algorithm is demonstrated on both synthetic and real SHGC images. Ari D. Gross, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | The extruded generalized cylinder: a deformable model for object recoveryabstractThere is increasing interest in the recovery of generalized cylinders (GCs) with curved spines. However, existing formulations of such GCs, for example those based an the Frenet-Serret frame or the tube model, suffer serious drawbacks: discontinuities, a lack of expressive power, "narrowing" in the plane normal to the spine, non-intuitive twisting behavior, and/or off-axis nonorthogonality of their local coordinate systems. We discuss some of the problems associated with the non-orthogonality of the coordinate system based on the Frenet-Serret frame. This non-orthogonality is induced by torsion effects and we show how to correct for it. We then introduce a new model, the extruded GC (EGC) model, which overcomes all the problems mentioned above. For complex axes, the EGC model is also simpler to understand and use than existing models. The EGC model is further extended by including local surface deformations. Recovery of the deformable EGC via a physically-motivated paradigm is demonstrated on pre-segmented data from a human carotid artery.> Thomas O'Donnell, Terrance E. Boult, Xi-Sheng Fang, Alok Gupta |
CVPR | 2 |
| 1994 | A periodic generalized cylinder model with local deformations for tracking closed contours exhibiting repeating motionabstractPeriodic data is often most appropriately described with a periodic model. This allows for the subsumption of tracking into model fitting and a natural method for the interpolation of data over time. We present a deformable periodic generalized cylinder (DPGC) for simultaneously modeling and tracking closed contours which exhibit a repeating motion. Results of the recovery of the projection of a beating left ventricle (BLV) from segmented ultrasound data as well as the segmentation of a BLV cross section from series of MR image slices are presented. Thomas O'Donnell, Alok Gupta, Terrance E. Boult |
ICPR (1) | 3 |
| 1994 | A Study of Upper and Lower Bounds for Minimum Congestion Routing in Lightwave NetworksabstractThis paper considers the combined problem of finding allocation of wavelengths to the stations (configuration) and finding the associated routing of the traffic to minimize congestion (the amount of maximum flow on any link). This work presents efficient algorithms for computing both upper and lower bounds on the congestion. The upper bounds-to obtain approximate solutions of this problem-are based modification of two heuristics (i) variable depth local search, and (ii) simulated annealing. A more significant contribution is a lower bound computation based on building flow trees to find a lower bound on the total flow, and then distributing the total flow over the links provide a lower bound on the congestion. This technique yields a tool which can be used in evaluating the quality of heuristic algorithms, and determining a termination criteria during minimization. This technique can be applied to other problems with flow-based objectives. Performance of the heuristics is analysed, and compared via simulation studies. It is shown that the heuristics perform on the average, within 20% of the computed lower bound, and 15% better than the previous methods to solve this problem.> Bülent Yener, Terrance E. Boult |
INFOCOM | 2 |
| 1994 | Analyzing skewed symmetries
Ari D. Gross, Terrance E. Boult |
Int. J. Comput. Vis. | 2 |
| 1993 | Removal of specularities using color and polarizationabstractAn algorithm for separating the specular and diffuse components of reflection from images is presented. The method uses color and polarization simultaneously to obtain strong constraints on the reflection components at each image point. Polarization is used to locally determine the color of the specular component, constraining the diffuse color at a pixel to a one-dimensional linear subspace. This subspace is used to find neighboring pixels whose color is consistent with the pixel. Diffuse color information from consistent neighbors is used to determine the diffuse color of the pixel. In contrast to previous separation algorithms, the proposed method can handle highlights that have a varying diffuse component, as well as highlights that include regions with different reflectance and material properties. Experimental results obtained by applying the algorithm to complex scenes with textured objects and strong interreflections are presented.> Shree K. Nayar, Xi-Sheng Fang, Terrance E. Boult |
CVPR | 3 |
| 1993 | Local Image Reconstruction and Subpixel Restoration AlgorithmsabstractThis paper introduces a new class of reconstruction algorithms that are fundamentally different from traditional approaches. We deviate from the standard practice that treats images as point samples. In this work, image values are treated as area samples generated by nonoverlapping integrators. This is consistent with the image formation process, particularly for CCD and CID cameras. We show that superior results are obtained by formulating reconstruction as a two-stage process: image restoration followed by application of the point spread function (PSF) of the imaging sensor. By coupling the PSF to the reconstruction process, we satisfy a more intuitive fidelity measure of accuracy that is based on the physical limitations of the sensor. Efficient local techniques for image restoration are derived to invert the effects of the PSF and estimate the underlying image that passed through the sensor. The reconstruction algorithms derived herein are local methods that compare favorably to cubic convolution, a well-known local technique, and they even rival global algorithms such as interpolating cubic splines. Evaluations are made by comparing their passband and stopband performances in the frequency domain, as well as by direct inspection of the resulting images in the spatial domain. A secondary advantage of the algorithms derived with this approach is that they satisfy an imaging-consistency property. This means that they exactly reconstruct the image for some function in the given class of functions. Their error can be shown to be at most twice that of the "optimal" algorithm for a wide range of optimality constraints. Terrance E. Boult, George Wolberg |
CVGIP Graph. Model. Image Process. | 1 |
| 1992 | Correcting chromatic aberrations using image warpingabstractChromatic aberration is due to refraction affecting each color channel differently. This paper addresses the use of image warping to reduce the impact of these aberrations in vision applications. The warp is determined using edge displacements which are fit with cubic splines. A new image reconstruction algorithm is used for nonlinear resampling. The main contribution of this work is to analyze the quality of the warping approach by comparing it with active lens control. Two different imaging systems are tested. 1 Introduction In an imaging system, refraction causes each color channel to focus differently. This phenomenon is called chromatic aberration. Chromatic aberration (hereafter CA) is generally broken up into two categories: axial chromatic aberrations (ACA) and lateral chromatic aberrations (LCA), e.g. see [7]. ACA manifests itself as blurring; LCA as geometric distortions. Often these sources of degradation cause measurable differences in color images, e.g., a simple CCTV le... Terrance E. Boult, George Wolberg |
CVPR | 1 |
| 1992 | The image understanding environment programabstractThe history of the image understanding environment (IUE) project, a five-year program to develop a common software environment for the development of algorithms and application systems, is reviewed. An overview of some of the data structures that are currently evolving as a specification for the IUE is provided. The ultimate goal of the project is to provide the basic data structures and algorithms that are required to carry state-of-the-art research in image understanding.> Joseph L. Mundy, Thomas O. Binford, Terrance E. Boult, Allen R. Hanson, J. Ross Beveridge, Robert M. Haralick, Visvanathan Ramesh, Charles A. Kohl, Daryl T. Lawton, Doug Morgan, Keith Price, Tom Strat |
CVPR | 3 |
| 1991 | Physically-based edge labelingabstractThe authors present a physically based approach, using polarization, to distinguish three types of image edges; limb edges, specular edges, and albedo/physical edges. Assuming general imaging conditions and smooth dielectric surfaces, a labeling scheme which enables one to distinguish among these edge types has been developed. The method is demonstrated on laboratory images.> Terrance E. Boult, Lawrence B. Wolff |
CVPR | 1 |
| 1991 | SYMAN: a symmetry analyzerabstractA description is given of the construction of a symmetry analyzer. Examples using SYMAN on both real and synthetic images are shown. SYMAN's combination of both global and local methods is discussed. The derivation of a global analytic solution for the skew axes when the degree of skew symmetry is known is described. A local tangent-based algorithm which has advantages over previous methods is presented.> Ari D. Gross, Terrance E. Boult |
CVPR | 2 |
| 1991 | Constraining Object Features Using a Polarization Reflectance ModelabstractThe authors present a polarization reflectance model that uses the Fresnel reflection coefficients. This reflectance model accurately predicts the magnitudes of polarization components of reflected light, and all the polarization-based methods presented follow from this model. The authors demonstrate the capability of polarization-based methods to segment material surfaces according to varying levels of relative electrical conductivity, in particular distinguishing dielectrics, which are nonconducting, and metals, which are highly conductive. Polarization-based methods can provide cues for distinguishing different intensity-edge types arising from intrinsic light-dark or color variations, intensity edges caused by specularities, and intensity edges caused by occluding contours where the viewing direction becomes nearly orthogonal to surface normals. Analysis of reflected polarization components is also shown to enable the separation of diffuse and specular components of reflection, unobscuring intrinsic surface detail saturated by specular glare. Polarization-based methods used for constraining surface normals are discussed.> Lawrence B. Wolff, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1990 | Energy-based segmentation of very sparse range surfacesabstractA segmentation technique for very sparse surfaces is described. It is based on minimizing the energy of the surfaces in the scene. While it could be used in almost any system as part of surface reconstruction/model recovery, the algorithm is designed to be usable when the depth information is scattered and very sparse, as is generally the case with depth generated by stereo algorithms. Results from a sequential algorithm are presented, and a working prototype that executes on the massively parallel Connection Machine is discussed. The technique presented models the surfaces with reproducing kernel-based splines which can be shown to solve a regularized surface reconstruction problem. From the functional form of these splines the authors derive computable upper and lower bounds on the energy of a surface over a given finite region. The computation of the spline, and the corresponding surface representation are quite efficient for very sparse data.> Terrance E. Boult, Mark Lerner |
ICRA | 1 |
| 1990 | An algorithm to recover generalized cylinders from a single intensity viewabstractA general method is presented for recovering straight homogeneous generalized cylinders from monocular intensity images. In this method, it is assumed that the generalized cylinder being recovered has purely diffuse reflectance and that the diffuse reflectance coefficient is constant. It is demonstrated that contour information alone is insufficient to recover a straight homogeneous generalized cylinder uniquely. It is shown that the sign and magnitude of the Gaussian curvature at a point vary among members of a contour-equivalent class. The contour image fails to constrain two parameters of the underlying generalized cylinder, the 3D axis tilt and translation. A method for ruling straight homogeneous generalized cylinder images is described. Once the rulings of the image have been recovered, all parameters derivable from contour alone can be recovered, all parameters derivable from contour alone can be recovered. To recover the two remaining parameters (modulo scale) not constrained by image contour, additional information must be incorporated into the recovery process, e.g. intensity information. A method for recovering the tilt of the object using the ruled contour image and intensity values along extremal cross-section curves is derived, along with a method for recovering the location of the object's 3D axis from intensity values along meridians of the surface. The methods outlined constitute an algorithm for recovering all the shape parameters (modulo scale).> Ari D. Gross, Terrance E. Boult |
ICRA | 2 |
| 1990 | Pruning bayesian networks for efficient computation
Michelle Baker, Terrance E. Boult |
UAI | 2 |
| 1990 | Dynamic digital distance maps in two dimensionsabstractAn efficient means for dealing with obstacles in motion is provided to extend the usefulness of digital distance maps. An algorithm is presented that allows one to compute what portions of a map will probably be affected by an obstacle's motion. The algorithm is based on an analysis of the distance transform as a problem in wave propagation. The regions that must be checked for possible updates when an obstacle moves are those that are in its or in the shadow of obstacles that are partially in the shadow of the moving obstacle. The technique can handle multiple fixed goals, multiple obstacles moving and interacting in an arbitrary fashion, and it is independent of the technique used for calculation of the distance map. The algorithm is demonstrated on a number of synthetic two-dimensional examples, and example timing results are reported.> Terrance E. Boult |
IEEE Trans. Robotics Autom. | 1 |
| 1989 | Polarization/radiometric based material classificationabstractA technique for identifying the material properties of objects in an image using multiple images taken through a polarizing lens at various rotations in front of a stationary camera (only the filter moves). Using these images, it is possible to obtain the classification of material surfaces at all points on a spectacular highlight. The algorithm is demonstrated on laboratory images. The authors assume a point source, the theory can only be applied at points where specular reflection dominates. Extensions of the theory to deal with extended light sources, which greatly increase the portion of the image giving rise to specular reflection, are also considered.> Lawrence B. Wolff, Terrance E. Boult |
CVPR | 2 |
| 1989 | Using Line Correspondence Stereo to Measure Surface Orientation
Lawrence B. Wolff, Terrance E. Boult |
IJCAI | 2 |
| 1989 | Separable image warping with spatial lookup tablesabstractImage warping refers to the 2-D resampling of a source image onto a target image. In the general case, this requires costly 2-D filtering operations. Simplifications are possible when the warp can be expressed as a cascade of orthogonal 1-D transformations. In these cases, separable transformations have been introduced to realize large performance gains. The central ideas in this area were formulated in the 2-pass algorithm by Catmull and Smith. Although that method applies over an important class of transformations, there are intrinsic problems which limit its usefulness.The goal of this work is to extend the 2-pass approach to handle arbitrary spatial mapping functions. We address the difficulties intrinsic to 2-pass scanline algorithms: bottlenecking, foldovers, and the lack of closed-form inverse solutions. These problems are shown to be resolved in a general, efficient, separable technique, with graceful degradation for transformations of increasing complexity. George Wolberg, Terrance E. Boult |
SIGGRAPH | 2 |
| 1989 | Straight homogeneous generalized cylinders: analysis of reflectance properties and a necessary condition for class membershipabstractConsideration is given to two membership tests for straight homogeneous generalized cylinders to determine if an object in the image is a member of the shape class. It is shown that contour information alone is insufficient to recover a straight homogeneous generalized cylinder uniquely. It is then shown that the sign and magnitude of the Gaussian curvature at a point vary among members of a contour-equivalent class. Next, a method of ruling straight homogeneous generalized cylinder images is developed. This ruling of the surface serves two functions. First, the ruling algorithm provides a heuristic test of whether or not the image is consistent with that of a straight homogeneous generalized cylinder. Secondly, the ruling makes explicit certain parameters of the underlying straight homogeneous generalized cylinder that the authors use in their second membership test. The second membership test is an intensity-based method which assumes that the surface has geodesics and that the albedo is constant. The method compares intensity values at corresponding meridian points along cross-sectional geodesics.> Ari D. Gross, Terrance E. Boult |
SMC | 2 |
| 1988 | Analysis of two new stereo algorithmsabstractThe authors present two algorithms for stereo matching that make use of simultaneous matching and surface reconstruction. By integrating matching and reconstruction, which are traditionally separated temporally, the algorithms can make use of the current surface approximation to help disambiguate the remaining matches. The result is a consistent smoothness assumption in both matching and surface reconstruction. The two methods differ in the surfaces reconstructed; one uses world surfaces and the other disparity surfaces. The authors present an initial experimental analysis of each algorithm and discuss the limitations of the approach and future work.> Terrance E. Boult, Liang-Hua Chen |
CVPR | 1 |
| 1988 | The integration of information from stereo and multiple shape-from-texture cuesabstractAn approach is described that integrates multiple visual sensing methodologies that yield three-dimensional information. The current system integrates feature-based stereo algorithms with various shape-from-texture algorithms. Unlike most systems for multisensor integration, that fuse all the information at one conceptual level, e.g., the surface level, the system under development uses two levels of data fusion: intraprocess integration and interprocess integration. Intraprocess integration techniques for feature-based stereo and shape-from-texture algorithms are briefly discussed. The authors also discuss an interprocess integration technique for fusing feature-based stereo and shape-from-texture based on smooth models of surfaces. Examples are presented using camera-acquired images.> Mark L. Moerdler, Terrance E. Boult |
CVPR | 2 |
| 1988 | Synergistic Smooth Surface StereoabstractThis paper presents a new algorithm for stereo matching. The algo- rithm combines what are generally three processes, feature matching, surface reconstruction, and segmentation of world surfaces, in a consis- tent and synergistic way. By integrating these phases, which are usually sequential, the algorithm can make use of the current surface approxi- mation to disambiguate potential matches. This results in higher data densities, a consistency of interpretation, and greater system flexibility. Examples of the algorithm are presented on real and synthetic images, including a scene with a transparent surface. Terrance E. Boult, Liang-Hua Chen |
ICCV | 1 |
| 1988 | Error Of Fit Measures For Recovering Parametric SolidsabstractParametric models of objects are becoming increasingly more impor- tant in computer vision. In the past few years, a number of researchers have investigated the recovery of a class of parametric models by the minimization of an error of fit measure. The measures used have typ- ically been chosen in an ad hoc fashion. This paper looks at how these measures affect the performance of a recovery system. This research can be divided into two parts. The first studies the biases of the po- tential error-of-fit measures with respect to the parameters recovered and examines the cross-sectional shape of their respective error of fit surfaces. This study is done in simulation by holding all but one pa- rameter constant. The second part of the research compares two of the better error of fit measures by using them in a recovery system. Both the number of iterations and the quality of the reconstruction are considered. Ari D. Gross, Terrance E. Boult |
ICCV | 2 |
| 1988 | Can we approximate zeros of functions with nonzero topological degree?abstractThe bisection method provides an affirmative answer for scalar functions. We show that the answer is negative for bivariate functions. This means, in particular, that an arbitrary continuation method cannot approximate a zero of every smooth bivariate function with non-zero topological degree. Terrance E. Boult, Christopher A. Sikorski |
J. Complex. | 1 |
| 1987 | Optimal algorithms: Tools for mathematical modelingabstractIn this paper we discuss the use of optimal error algorithms as tools to aid the process of mathematical modeling. Often a model cannot directly be tested, but computational experiments are performed and the models are evaluated based on the performance of algorithms which embody the model. In general the only conclusion that should be drawn from such comparisons is which algorithm, not which model, is best. If, however, we compare optimal error algorithms which embody different models, we can draw conclusions about the appropriateness of the different models. After a general discussion of the use of optimal algorithms in modeling, we present an example from the modeling of human reconstruction of surfaces from sparse visual depth data. We then discuss the interplay of contaminated data and modeling. We end with a short discussion of the interplay between optimal error algorithms, algorithm complexity, and modeling. Terrance E. Boult |
J. Complex. | 1 |
| 1986 | Complexity of computing topological degree of lipschitz functions in n dimensionsabstractWe find lower and upper bounds on the complexity, comp(deg), of computing the topological degree of functions defined on the n -dimensional unit cube C n , f : C n → R n , n ≥ 2, which satisfy a Lipschitz condition with constant K and whose infinity norm at each point on the boundary of C n is at least d , d > 0, and such that K 8d ≥ 1 . A lower bound, comp low ≅ 2n( K 8d ) n−1 (c + n) is obtained for comp(deg), assuming that each function evaluation costs c and elementary arithmetic operations and comparisons cost unity. We prove that the topological degree can be computed using A = (⌊ K 2d + 1⌋ + 1) n − (⌊ K 2d + 1⌋ − 1) n function evaluations. It can be done by an algorithm ϕ ∗ due to Kearfott, with cost given by comp (ϕ ∗ ) ≅ A (c + ( n 2 2 )(n − 1)!) . Thus for small n, say n ≤ 5, and small K 2d , say K 2d ≤ 9 , the degree can be computed in time at most 10 5 ( c + 300). For large n and/or large K 2d the problem is intractable. Terrance E. Boult, Christopher A. Sikorski |
J. Complex. | 1 |