Michael J. Jones 0001

dblp:56/4313-2 · also Michael Jones 0002 · DBLP profile ↗
← Back
40ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0001-5215-2346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
abstract
Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost of annotating training data with the class labels of all these events, along with their start and end times, imposes constraints on the scalability of AVVP techniques unless they can be trained in a weakly-supervised setting, where only modality-agnostic, video-level labels are available in the training data. To this end, recently proposed approaches seek to generate segment-level pseudo-labels to better guide model training. However, the absence of inter-segment dependencies when generating these pseudo-labels and the general bias towards predicting labels that are absent in a segment limit their performance. This work proposes a novel approach towards overcoming these weaknesses called Uncertainty-Weighted Weakly-Supervised Audio-Visual Video Parsing (UWAV). Additionally, our innovative approach factors in the uncertainty associated with these estimated pseudo-labels and incorporates a feature mixup based training regularization for improved training. Empirical results show that UWAV outperforms state-of-the-art methods for the AVVP task on multiple metrics, across two different datasets, attesting to its effectiveness and generalizability.1
Yung-Hsuan Lai, Janek Ebbers, Yu-Chiang Frank Wang, François G. Germain, Michael J. Jones 0001, Moitreya Chatterjee
CVPR5
2024 Equivariant Spatio-temporal Self-supervision for LiDAR Object Detection
Deepti Hegde, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones 0001, Vishal M. Patel
ECCV (26)4
2024 Pixel-Grounded Prototypical Part Networks
abstract
Prototypical part neural networks (ProtoPartNNs), namely ProtoPNet and its derivatives, are an intrinsically interpretable approach to machine learning. Their prototype learning scheme enables intuitive explanations of the form, this (prototype) looks like that (testing image patch). But, does this actually look like that? In this work, we delve into why object part localization and associated heat maps in past work are misleading. Rather than localizing to object parts, existing ProtoPartNNs localize to the entire image, contrary to generated explanatory visualizations. We argue that detraction from these underlying issues is due to the alluring nature of visualizations and an over-reliance on intuition. To alleviate these issues, we devise new receptive field-based architectural constraints for meaningful localization and a principled pixel space mapping for ProtoPartNNs. To improve interpretability, we propose additional architectural improvements, including a simplified classification head. We also make additional corrections to ProtoPNet and its derivatives, such as the use of a validation set, rather than a test set, to evaluate generalization during training. Our approach, PixPNet (Pixel-grounded Prototypical part Network), is the only ProtoPartNN that truly learns and localizes to prototypical object parts. We demonstrate that PixPNet achieves quantifiably improved interpretability without sacrificing accuracy1.
Zachariah Carmichael, Suhas Lohit, Anoop Cherian, Michael J. Jones 0001, Walter J. Scheirer
WACV4
2023 EVAL: Explainable Video Anomaly Localization
abstract
We develop a novel framework for single-scene video anomaly localization that allows for human-understandable reasons for the decisions the system makes. We first learn general representations of objects and their motions (using deep networks) and then use these representations to build a high-level, location-dependent model of any particular scene. This model can be used to detect anomalies in new videos of the same scene. Importantly, our approach is explainable - our high-level appearance and motion features can provide human-understandable reasons for why any part of a video is classified as normal or anomalous. We conduct experiments on standard video anomaly detection datasets (Street Scene, CUHK Avenue, ShanghaiTech and UCSD Ped1, Ped2) and show significant improvements over the previous state-of-the-art. All of our code and extra datasets will be made publicly available.
Michael J. Jones 0001, Erik G. Learned-Miller
CVPR2
2023 Robust Frame-to-Frame Camera Rotation Estimation in Crowded Scenes
abstract
We present an approach to estimating camera rotation in crowded, real-world scenes from handheld monocular video. While camera rotation estimation is a well-studied problem, no previous methods exhibit both high accuracy and acceptable speed in this setting. Because the setting is not addressed well by other datasets, we provide a new dataset and benchmark, with high-accuracy, rigorously verified ground truth, on 17 video sequences. Methods developed for wide baseline stereo (e.g., 5-point methods) perform poorly on monocular video. On the other hand, methods used in autonomous driving (e.g., SLAM) leverage specific sensor setups, specific motion models, or local optimization strategies (lagging batch processing) and do not generalize well to handheld video. Finally, for dynamic scenes, commonly used robustification techniques like RANSAC require large numbers of iterations, and become prohibitively slow. We introduce a novel generalization of the Hough transform on SO(3) to efficiently and robustly find the camera rotation most compatible with optical flow. Among comparably fast methods, ours reduces error by almost 50% over the next best, and is more accurate than any method, irrespective of speed. This represents a strong new performance point for crowded scenes, an important setting for computer vision. The code and the dataset are available at https://fabiendelattre.com/robustrotation-estimation.
Fabien Delattre, David Dirnfeld, Phat Nguyen, Stephen Scarano, Michael J. Jones 0001, Pedro Miraldo, Erik G. Learned-Miller
ICCV5
2022 Cross-Modal Knowledge Transfer Without Task-Relevant Source Data
Sk Miraj Ahmed, Suhas Lohit, Kuan-Chuan Peng, Michael J. Jones 0001, Amit K. Roy-Chowdhury
ECCV (34)4
2022 An Empirical Analysis of Boosting Deep Networks
abstract
Boosting is a method for finding a highly accurate classifier by linearly combining many “weak” classifiers, each of which may be only moderately accurate. Thus, boosting is a method for learning an ensemble of classifiers. While boosting has been shown to be very effective for decision trees, its impact on neural networks has not been extensively studied. Using standard object recognition datasets, we verify experimentally the well-known result that a boosted ensemble of decision trees usually generalizes much better on testing data than a single decision tree with the same number of parameters. In contrast, using the same datasets and boosting algorithms, our experiments show the opposite to be true when using neural networks (both convolutional neural networks (CNNs) and multilayer perceptrons (MLPs)). We find that a single neural network usually generalizes better than a boosted ensemble of smaller neural networks with the same total number of parameters. While this is an experimental investigation, more theoretical research is warranted to understand the role of boosting in deep learning-based classifiers.
Sai Saketh Rambhatla, Michael J. Jones 0001, Rama Chellappa
IJCNN2
2022 What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical Perspective
abstract
Knowledge distillation (KD) is a general neural network training approach that uses a teacher model to guide the student model. Existing works mainly study KD from the network output side (e.g., trying to design a better KD loss function), while few have attempted to understand it from the input side. Especially, its interplay with data augmentation (DA) has not been well understood. In this paper, we ask: Why do some DA schemes (e.g., CutMix) inherently perform much better than others in KD? What makes a "good" DA in KD? Our investigation from a statistical perspective suggests that a good DA scheme should reduce the covariance of the teacher-student cross-entropy. A practical metric, the stddev of teacher’s mean probability (T. stddev), is further presented and well justified empirically. Besides the theoretical understanding, we also introduce a new entropy-based data-mixing DA scheme, CutMixPick, to further enhance CutMix. Extensive empirical studies support our claims and demonstrate how we can harvest considerable performance gains simply by using a better DA scheme in knowledge distillation. Code: https://github.com/MingSun-Tse/Good-DA-in-KD.
Huan Wang 0014, Suhas Lohit, Michael J. Jones 0001, Yun Fu 0001
NeurIPS3
2022 Model Compression Using Optimal Transport
abstract
Model compression methods are important to allow for easier deployment of deep learning models in compute, memory and energy-constrained environments such as mobile phones. Knowledge distillation is a class of model compression algorithms where knowledge from a large teacher network is transferred to a smaller student network thereby improving the student’s performance. In this paper, we show how optimal transport-based loss functions can be used for training a student network which encourages learning student network parameters that help bring the distribution of student features closer to that of the teacher features. We present image classification results on CIFAR-100, SVHN and ImageNet and show that the proposed optimal transport loss functions perform comparably to or better than other loss functions.
Suhas Lohit, Michael J. Jones 0001
WACV2
2022 A Survey of Single-Scene Video Anomaly Detection
abstract
This article summarizes research trends on the topic of anomaly detection in video feeds of a single scene. We discuss the various problem formulations, publicly available datasets and evaluation criteria. We categorize and situate past research into an intuitive taxonomy and provide a comprehensive comparison of the accuracy of many algorithms on standard test sets. Finally, we also provide best practices and suggest some possible directions for future research.
Bharathkumar Ramachandra, Michael J. Jones 0001, Ranga Raju Vatsavai
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Perceptual metric learning for video anomaly detection
Bharathkumar Ramachandra, Michael J. Jones 0001, Ranga Raju Vatsavai
Mach. Vis. Appl.2
2020 LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood
abstract
Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained using our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmarks is visible, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on mulitple standard face alignment datasets. Our method's estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks.
Abhinav Kumar 0004, Tim K. Marks, Wenxuan Mou, Ye Wang 0001, Michael J. Jones 0001, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu 0002, Chen Feng 0002
CVPR5
2020 Street Scene: A new dataset and evaluation protocol for video anomaly detection
abstract
Progress in video anomaly detection research is currently slowed by small datasets that lack a wide variety of activities as well as flawed evaluation criteria. This paper aims to help move this research effort forward by introducing a large and varied new dataset called Street Scene, as well as two new evaluation criteria that provide a better estimate of how an algorithm will perform in practice. In addition to the new dataset and evaluation criteria, we present two variations of a novel baseline video anomaly detection algorithm and show they are much more accurate on Street Scene than two well known algorithms from the literature.
Bharathkumar Ramachandra, Michael J. Jones 0001
WACV2
2020 Learning a distance function with a Siamese network to localize anomalies in videos
abstract
This work introduces a new approach to localize anomalies in surveillance video. The main novelty is the idea of using a Siamese convolutional neural network (CNN) to learn a distance function between a pair of video patches (spatio-temporal regions of video). The learned distance function, which is not specific to the target video, is used to measure the distance between each video patch in the testing video and the video patches found in normal training video. If a testing video patch is not similar to any normal video patch then it must be anomalous. We compare our approach to previously published algorithms using 4 evaluation measures and 3 challenging target benchmark datasets. Experiments show that our approach either surpasses or performs comparably to current state-of-the-art methods.
Bharathkumar Ramachandra, Michael J. Jones 0001, Ranga Raju Vatsavai
WACV2
2019 Body Part Alignment and Temporal Attention Pooling for Video-Based Person Re-Identification
Sai Saketh Rambhatla, Michael J. Jones 0001
BMVC2
2016 A Multi-stream Bi-directional Recurrent Neural Network for Fine-Grained Action Detection
abstract
We present a multi-stream bi-directional recurrent neural network for fine-grained action detection. Recently, twostream convolutional neural networks (CNNs) trained on stacked optical flow and image frames have been successful for action recognition in videos. Our system uses a tracking algorithm to locate a bounding box around the person, which provides a frame of reference for appearance and motion and also suppresses background noise that is not within the bounding box. We train two additional streams on motion and appearance cropped to the tracked bounding box, along with full-frame streams. Our motion streams use pixel trajectories of a frame as raw features, in which the displacement values corresponding to a moving scene point are at the same spatial position across several frames. To model long-term temporal dynamics within and between actions, the multi-stream CNN is followed by a bi-directional Long Short-Term Memory (LSTM) layer. We show that our bi-directional LSTM network utilizes about 8 seconds of the video sequence to predict an action label. We test on two action detection datasets: the MPII Cooking 2 Dataset, and a new MERL Shopping Dataset that we introduce and make available to the community with this paper. The results demonstrate that our method significantly outperforms state-of-the-art action detection methods on both datasets.
Tim K. Marks, Michael J. Jones 0001, Oncel Tuzel, Ming Shao
CVPR3
2016 Exemplar learning for extremely efficient anomaly detection in real-valued time series
Michael J. Jones 0001, Daniel Nikovski, Makoto Imamura, Takahisa Hirata
Data Min. Knowl. Discov.1
2015 An improved deep learning architecture for person re-identification
abstract
In this work, we propose a method for simultaneously learning features and a corresponding similarity metric for person re-identification. We present a deep convolutional architecture with layers specially designed to address the problem of re-identification. Given a pair of images as input, our network outputs a similarity value indicating whether the two input images depict the same person. Novel elements of our architecture include a layer that computes cross-input neighborhood differences, which capture local relationships between the two input images based on mid-level features from each input image. A high-level summary of the outputs of this layer is computed by a layer of patch summary features, which are then spatially integrated in subsequent layers. Our method significantly outperforms the state of the art on both a large data set (CUHK03) and a medium-sized data set (CUHK01), and is resistant to over-fitting. We also demonstrate that by initially training on an unrelated large data set before fine-tuning on a small target data set, our network can achieve results comparable to the state of the art even on a small data set (VIPeR).
Ejaz Ahmed 0002, Michael J. Jones 0001, Tim K. Marks
CVPR2
2015 Real-time 3D head pose and facial landmark estimation from depth images using triangular surface patch features
abstract
We present a real-time system for 3D head pose estimation and facial landmark localization using a commodity depth sensor. We introduce a novel triangular surface patch (TSP) descriptor, which encodes the shape of the 3D surface of the face within a triangular area. The proposed descriptor is viewpoint invariant, and it is robust to noise and to variations in the data resolution. Using a fast nearest neighbor lookup, TSP descriptors from an input depth map are matched to the most similar ones that were computed from synthetic head models in a training phase. The matched triangular surface patches in the training set are used to compute estimates of the 3D head pose and facial landmark positions in the input depth map. By sampling many TSP descriptors, many votes for pose and landmark positions are generated which together yield robust final estimates. We evaluate our approach on the publicly available Biwi Kinect Head Pose Database to compare it against state-of-the-art methods. Our results show a significant improvement in the accuracy of both pose and landmark location estimates while maintaining real-time speed.
Chavdar Papazov, Tim K. Marks, Michael J. Jones 0001
CVPR3
2011 Pose Normalization via Learned 2D Warping for Fully Automatic Face Recognition
abstract
We present a novel approach to pose-invariant face recognition that handles continuous pose variations, is not database-specific, and achieves high accuracy without any manual intervention. Our method uses multidimensional Gaussian process regression to learn a nonlinear mapping function from the 2D shapes of faces at any non-frontal pose to the corresponding 2D frontal face shapes. We use this mapping to take an input image of a new face at an arbitrary pose and pose-normalize it, generating a synthetic frontal image of the face that is then used for recognition. Our fully automatic system for face recognition includes automatic methods for extracting 2D facial feature points and accurately estimating 3D head pose, and this information is used as input to the 2D pose-normalization algorithm. The current system can handle pose variation up to 45 degrees to the left or right (yaw angle) and up to 30 degrees up or down (pitch angle). The system demonstrates high accuracy in recognition experiments on the CMU-PIE, USF 3D, and Multi-PIE databases, showing excellent generalization across databases and convincingly outperforming other automatic methods.
Akshay Asthana, Michael J. Jones 0001, Tim K. Marks, Kinh H. Tieu, Roland Göcke
BMVC2
2011 Fully automatic pose-invariant face recognition via 3D pose normalization
abstract
An ideal approach to the problem of pose-invariant face recognition would handle continuous pose variations, would not be database specific, and would achieve high accuracy without any manual intervention. Most of the existing approaches fail to match one or more of these goals. In this paper, we present a fully automatic system for pose-invariant face recognition that not only meets these requirements but also outperforms other comparable methods. We propose a 3D pose normalization method that is completely automatic and leverages the accurate 2D facial feature points found by the system. The current system can handle 3D pose variation up to ±45° in yaw and ±30° in pitch angles. Recognition experiments were conducted on the USF 3D, Multi-PIE, CMU-PIE, FERET, and FacePix databases. Our system not only shows excellent generalization by achieving high accuracy on all 5 databases but also outperforms other methods convincingly.
Akshay Asthana, Tim K. Marks, Michael J. Jones 0001, Kinh H. Tieu, M. V. Rohith
ICCV3
2010 Morphable Reflectance Fields for enhancing face recognition
abstract
In this paper, we present a novel framework to address the confounding effects of illumination variation in face recognition. By augmenting the gallery set with realistically relit images, we enhance recognition performance in a classifier-independent way. We describe a novel method for single-image relighting, Morphable Reflectance Fields (MoRF), which does not require manual intervention and provides relighting superior to that of existing automatic methods. We test our framework through face recognition experiments using various state-of-the-art classifiers and popular benchmark datasets: CMU PIE, Multi-PIE, and MERL Dome. We demonstrate that our MoRF relighting and gallery augmentation framework achieves improvements in terms of both rank-1 recognition rates and ROC curves. We also compare our model with other automatic relighting methods to confirm its advantage. Finally, we show that the recognition rates achieved using our framework exceed those of state-of-the-art recognizers on the aforementioned databases.
Ritwik Kumar, Michael J. Jones 0001, Tim K. Marks
CVPR2
2008 Pedestrian detection using boosted features over many frames
abstract
A scanning window type pedestrian detector is presented that uses both appearance and motion information to find walking people in surveillance video. We extend the work of Viola, Jones and Snow [10] to use many more frames as input to the detector thus allowing a much more detailed analysis of motion. The resulting detector is about an order of magnitude more accurate than the detector of Viola, Jones and Snow. It is also computationally efficient, processing frames at the rate of 5 Hz on a 3 GHz Pentium processor.
Michael J. Jones 0001, Daniel Snow
ICPR1
2008 Iris Extraction Based on Intensity Gradient and Texture Difference
abstract
Biometrics has become more and more important in security applications. In comparison with many other biometric features, iris recognition has very high recognition accuracy. Successful iris recognition depends largely on correct iris localization, however, the performance of current techniques for iris localization still leaves room for improvement. To improve the iris localization performance, we propose a novel method that optimally utilizes both the intensity gradient and texture difference. Experimental results demonstrate that our new approach gives much better results than previous approaches. In order to make the iris boundary more accurate, we present a new issue called model selection and propose a method to choose between ellipse/circle and circle/circle models. Furthermore, we propose a dome model to compute mask images and remove eyelid occlusions in the unwrapped images rather than in the original eye images with a least commitment strategy.
Guodong Guo, Michael J. Jones 0001
WACV2
2005 Detecting Pedestrians Using Patterns of Motion and Appearance
Paul A. Viola, Michael J. Jones 0001, Daniel Snow
Int. J. Comput. Vis.2
2004 Robust Real-Time Face Detection
Paul A. Viola, Michael J. Jones 0001
Int. J. Comput. Vis.2
2003 Detecting Pedestrians Using Patterns of Motion and Appearance
abstract
This paper describes a pedestrian detection system that integrates image intensity information with motion information. We use a detection style algorithm that scans a detector over two consecutive frames of a video sequence. The detector is trained (using AdaBoost) to take advantage of both motion and appearance information to detect a walking person. Past approaches have built detectors based on appearance information, but ours is the first to combine both sources of information in a single detector. The implementation described runs at about 4 frames/second, detects pedestrians at very small scales (as small as 20/spl times/15 pixels), and has a very low false positive rate. Our approach builds on the detection work of Viola and Jones. Novel contributions of this paper include: i) development of a representation of image motion which is extremely efficient, and ii) implementation of a state of the art pedestrian detection system which operates on low resolution images under difficult conditions (such as rain and snow).
Paul A. Viola, Michael J. Jones 0001, Daniel Snow
ICCV2
2002 Statistical Color Models with Application to Skin Detection
Michael J. Jones 0001, James M. Rehg
Int. J. Comput. Vis.1
2001 Rapid Object Detection using a Boosted Cascade of Simple Features
abstract
This paper describes a machine learning approach for visual object detection which is capable of processing images extremely rapidly and achieving high detection rates. This work is distinguished by three key contributions. The first is the introduction of a new image representation called the "integral image" which allows the features used by our detector to be computed very quickly. The second is a learning algorithm, based on AdaBoost, which selects a small number of critical visual features from a larger set and yields extremely efficient classifiers. The third contribution is a method for combining increasingly more complex classifiers in a "cascade" which allows background regions of the image to be quickly discarded while spending more computation on promising object-like regions. The cascade can be viewed as an object specific focus-of-attention mechanism which unlike previous approaches provides statistical guarantees that discarded regions are unlikely to contain the object of interest. In the domain of face detection the system yields detection rates comparable to the best previous systems. Used in real-time applications, the detector runs at 15 frames per second without resorting to image differencing or skin color detection.
Paul A. Viola, Michael J. Jones 0001
CVPR (1)2
2001 Robust Real-Time Face Detection
abstract
We have constructed a frontal face detection system which achieves detection and false positive rates which are equivalent to the best published results [7, 5, 6, 4, 1]. This face detection system is most clearly distinguished from previous approaches in its ability to detect faces extremely rapidly. Operating on 384 by 288 pixel images, faces are detected at 15 frames per second on a conventional 700 MHz Intel Pentium III. In other face detection systems, auxiliary information, such as image differences in video sequences, or pixel color in color images, have been used to achieve high frame rates. Our system achieves high frame rates working only with the information present in a single grey scale image. These alternative sources of information can also be integrated with our system to achieve even higher frame rates.
Paul A. Viola, Michael J. Jones 0001
ICCV2
2001 Fast and Robust Classification using Asymmetric AdaBoost and a Detector Cascade
abstract
This paper develops a new approach for extremely fast detection in do- mains where the distribution of positive and negative examples is highly skewed (e.g. face detection or database retrieval). In such domains a cascade of simple classifiers each trained to achieve high detection rates and modest false positive rates can yield a final detector with many desir- able features: including high detection rates, very low false positive rates, and fast performance. Achieving extremely high detection rates, rather than low error, is not a task typically addressed by machine learning al- gorithms. We propose a new variant of AdaBoost as a mechanism for training the simple classifiers used in the cascade. Experimental results in the domain of face detection show the training algorithm yields sig- nificant improvements in performance over conventional AdaBoost. The final face detection system can process 15 frames per second, achieves over 90% detection, and a false positive rate of 1 in a 1,000,000.
Paul A. Viola, Michael J. Jones 0001
NIPS2
1999 Statistical Color Models with Application to Skin Detection
abstract
The existence of large image datasets such as photos on the World Wide Web make it possible to build powerful generic models for low-level image attributes like color using simple histogram learning techniques. We describe the construction of color models for skin and non-skin classes from a dataset of nearly 1 billion labeled pixels. These classes exhibit a surprising degree of separability which we exploit by building a skin pixel detector that achieves an equal error rate of 88%. We compare the performance of histogram and mixture models in skin detection and find histogram models to be superior in accuracy and computational cost. Using aggregate features computed from the skin detector we build a remarkably effective detector for naked people. We believe this work is the most comprehensive and detailed exploration of skin color models to date.
Michael J. Jones 0001, James M. Rehg
CVPR1
1999 A Cluster-based Statistical Model for Object Detection
abstract
This paper presents an approach to object detection which is based on recent work in statistical models for texture synthesis and recognition. Our method follows the texture recognition work of De Bonet and Viola (1998). We use feature vectors which capture the joint occurrence of local features at multiple resolutions. The distribution of feature vectors for a set of training images of an object class is estimated by clustering the data and then forming a mixture of Gaussian models. The mixture model is further refined by determining which clusters are the most discriminative for the class and retaining only those clusters. After the model is learned, test images are classified by computing the likelihood of their feature vectors with respect to the model. We present promising results in applying our technique to face detection and car detection.
Thomas D. Rikert, Michael J. Jones 0001, Paul A. Viola
ICCV2
1998 Hierarchical Morphable Models
abstract
This paper presents a new technique for modelling object classes (such as faces) and matching the model to novel images from the object class. The technique can be used for a variety of image analysis applications including face recognition, object verification and facial expression analysis. The model, called a hierarchical morphable model, is "learned" from example images (partioned into components) and their correspondences. This is an extension to the work on morphable models described in previous papers. Hierarchical morphable models are shown to find good matches to novel face images and are also robust to partial occlusion.
Michael J. Jones 0001, Tomaso A. Poggio
CVPR1
1998 Gaze Estimation Using Morphable Models
Thomas D. Rikert, Michael J. Jones 0001
FG2
1998 Multidimensional Morphable Models
abstract
We describe a flexible model for representing images of objects of a certain class, known a priori, such as faces, and introduce a new algorithm for matching it to a novel image and thereby performing image analysis. We call this model a multidimensional morphable model or just a, morphable model. The morphable model is learned from example images (called prototypes) of objects of a class. In this paper we introduce an effective stochastic gradient descent algorithm that automaticaIly matches a model to a novel image by finding the parameters that minimize the error between the image generated by the model and the novel image. Two examples demonstrate the robustness and the broad range of applicability of the matching algorithm and the underlying morphable model. Our approach can provide novel solutions to several vision tasks, including the computation of image correspondence, object verification, image synthesis and image compression.
Michael J. Jones 0001, Tomaso A. Poggio
ICCV1
1998 Multidimensional Morphable Models: A Framework for Representing and Matching Object Classes
Michael J. Jones 0001, Tomaso A. Poggio
Int. J. Comput. Vis.1
1997 A bootstrapping algorithm for learning linear models of object classes
abstract
Flexible models of object classes, based on linear combinations of prototypical images, are capable of matching novel images of the same class and have been shown to be a powerful tool to solve several fundamental vision tasks such as recognition, synthesis and correspondence. The key problem in creating a specific flexible model is the computation of pixelwise correspondence between the prototypes, a task done until now in a semiautomatic way. In this paper we describe an algorithm that automatically bootstraps the correspondence between the prototypes. The algorithm -which can be used for 2D images as well as for 3D models-is shown to synthesize successfully a flexible model of frontal face images and a flexible model of handwritten digits.
Thomas Vetter, Michael J. Jones 0001, Tomaso A. Poggio
CVPR2
1995 Model-Based Matching of Line Drawings by Linear Combinations of Prototypes
abstract
We describe a technique for finding pixelwise correspondences between two images by using models of objects of the same class to guide the search. The object models are "learned" from example images (also called prototypes) of an object class. The models consist of a linear combination of prototypes. The flow fields giving pixelwise correspondences between a base prototype and each of the other prototypes must be given. A novel image of an object of the same class is matched to a model by minimizing an error between the novel image and the current guess for the closest model image. Currently, the algorithm applies to line drawings of objects. An extension to real grey level images is discussed.>
Michael J. Jones 0001, Tomaso A. Poggio
ICCV1
1995 Regularization Theory and Neural Networks Architectures
abstract
We had previously shown that regularization principles lead to approximation schemes that are equivalent to networks with one layer of hidden units, called regularization networks. In particular, standard smoothness functionals lead to a subclass of regularization networks, the well known radial basis functions approximation schemes. This paper shows that regularization networks encompass a much broader range of approximation schemes, including many of the popular general additive models and some of the neural networks. In particular, we introduce new classes of smoothness functionals that lead to different classes of basis functions. Additive splines as well as some tensor product splines can be obtained from appropriate classes of smoothness functionals. Furthermore, the same generalization that extends radial basis functions (RBF) to hyper basis functions (HBF) also leads from additive models to ridge approximation models, containing as special cases Breiman's hinge functions, some forms of projection pursuit regression, and several types of neural networks. We propose to use the term generalized regularization networks for this broad class of approximation schemes that follow from an extension of regularization. In the probabilistic interpretation of regularization, the different classes of basis functions correspond to different classes of prior probabilities on the approximating function spaces, and therefore to different types of smoothness assumptions. In summary, different multilayer networks with one hidden layer, which we collectively call generalized regularization networks, correspond to different classes of priors and associated smoothness functionals in a classical regularization principle. Three broad classes are (1) radial basis functions that can be generalized to hyper basis functions, (2) some tensor product splines, and (3) additive splines that can be generalized to schemes of the type of ridge approximation, hinge functions, and several perceptron-like neural networks with one hidden layer.
Federico Girosi, Michael J. Jones 0001, Tomaso A. Poggio
Neural Comput.2