VLDB 2026 Research / reviewers in the wild / expert
Simon Prince
dblp:p/SimonPrince · also Simon J. D. Prince
· DBLP profile ↗
34ranked-venue papers
10as first author
2since 2021 · last 2021
0000-0002-5545-3344ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 7 first-authorArtificial intelligence and machine learning · 21 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Deep learning architectures and training · 20% Generative modeling · 20% Face, body and person analysis · 14% | |
| Computer graphics and multimedia
10 papers |
Image and video processing · 62% Visual content generation and editing · 16% Virtual and augmented reality · 13% |
Topics — the 30 heaviest of 56, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
0.5 | 1 | 2021 | Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021 |
Machine learning › Generative modeling
normalizing flow |
0.5 | 1 | 2021 | Normalizing Flows: An Introduction and Review of Current Methods · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Machine learning › Deep learning architectures and training › neural network training
training on small datasets |
0.5 | 1 | 2021 | Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021 |
Machine learning › Deep learning architectures and training
transformer |
0.5 | 1 | 2021 | Optimizing Deeper Transformers on Small Datasets · ACL/IJCNLP (1) 2021 |
Computer vision › Face, body and person analysis
face recognition |
0.4 | 5 | 2012 | Probabilistic Models for Inference about Identity · IEEE Trans. Pattern Anal. Mach. Intell. 2012 Joint and implicit registration for face recognition · CVPR 2009 Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008 |
Image and video processing › image restoration
image inpainting |
0.2 | 2 | 2015 | Modeling object appearance using Context-Conditioned Component Analysis · CVPR 2015 Visio-lization: generating novel facial images · ACM Trans. Graph. 2009 |
Machine learning › Representation and self-supervised learning
component analysis |
0.2 | 1 | 2015 | Modeling object appearance using Context-Conditioned Component Analysis · CVPR 2015 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning |
0.2 | 1 | 2015 | Modeling object appearance using Context-Conditioned Component Analysis · CVPR 2015 |
Visual content generation and editing
appearance transfer |
0.2 | 1 | 2015 | Modeling object appearance using Context-Conditioned Component Analysis · CVPR 2015 |
Image and video processing
image segmentation |
0.2 | 2 | 2010 | "Lattice Cut" - Constructing superpixels using layer constraints · CVPR 2010 Superpixel lattices · CVPR 2008 |
Image and video processing › image segmentation › superpixel segmentation
superpixel lattice |
0.2 | 2 | 2010 | "Lattice Cut" - Constructing superpixels using layer constraints · CVPR 2010 Superpixel lattices · CVPR 2008 |
Image and video processing › image segmentation
superpixel segmentation |
0.2 | 2 | 2010 | "Lattice Cut" - Constructing superpixels using layer constraints · CVPR 2010 Superpixel lattices · CVPR 2008 |
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection |
0.1 | 2 | 2007 | Pre-Attentive and Attentive Detection of Humans in Wide-Field Scenes · Int. J. Comput. Vis. 2007 Statistical Cue Integration for Foveated Wide-Field Surveillance · CVPR (2) 2005 |
Machine learning › Generative modeling › face synthesis
generative face model |
0.1 | 2 | 2012 | Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008 Probabilistic Models for Inference about Identity · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Computer vision › Image recognition and object detection
attribute recognition |
0.1 | 1 | 2009 | Patch-based within-object classification · ICCV 2009 |
Computer vision › Face, body and person analysis
face manipulation |
0.1 | 1 | 2009 | Visio-lization: generating novel facial images · ACM Trans. Graph. 2009 |
Machine learning › Generative modeling
face synthesis |
0.1 | 1 | 2009 | Visio-lization: generating novel facial images · ACM Trans. Graph. 2009 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2009 | Epitomized priors for multi-labeling problems · CVPR 2009 |
Machine learning › Generative modeling › diffusion model
image editing |
0.1 | 1 | 2009 | Visio-lization: generating novel facial images · ACM Trans. Graph. 2009 |
Machine learning › Generative modeling
image generation |
0.1 | 1 | 2009 | Visio-lization: generating novel facial images · ACM Trans. Graph. 2009 |
Computer vision › 3D vision › low-level vision › feature detection
keypoint detection |
0.1 | 1 | 2009 | Joint and implicit registration for face recognition · CVPR 2009 |
Computer vision › Segmentation and scene understanding › semantic segmentation
multi-label segmentation |
0.1 | 1 | 2009 | Epitomized priors for multi-labeling problems · CVPR 2009 |
Computer vision › Image recognition and object detection › image classification
patch-based classification |
0.1 | 1 | 2009 | Patch-based within-object classification · ICCV 2009 |
Computer vision › Segmentation and scene understanding
scene parsing |
0.1 | 1 | 2009 | Epitomized priors for multi-labeling problems · CVPR 2009 |
Computer vision › Segmentation and scene understanding › image segmentation
segmentation evaluation |
0.1 | 1 | 2009 | Scene shape priors for superpixel segmentation · ICCV 2009 |
Computer vision › Segmentation and scene understanding › image segmentation › region-based segmentation
superpixel segmentation |
0.1 | 1 | 2009 | Scene shape priors for superpixel segmentation · ICCV 2009 |
Computer vision › 3D vision
3d reconstruction |
0.1 | 2 | 2005 | Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005 3D Live: Real Time Captured Content for Mixed Reality · ISMAR 2002 |
Computer vision › 3D vision › 3d reconstruction
shape from silhouette |
0.1 | 2 | 2005 | Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005 3D Live: Real Time Captured Content for Mixed Reality · ISMAR 2002 |
Virtual and augmented reality
augmented reality |
0.1 | 2 | 2005 | Live three-dimensional content for augmented reality · IEEE Trans. Multim. 2005 3-D live: real time interaction for mixed reality · CSCW 2002 |
Computer vision › Face, body and person analysis › face recognition › robust face recognition
pose-invariant face recognition |
0.1 | 1 | 2008 | Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008 |
Methods — techniques the papers use, named apart from their topics
normalizing flow · 0.5density estimation · 0.5subspace modeling · 0.4context-conditioned component analysis · 0.4EM algorithm · 0.3tied factor analysis · 0.2graph cuts · 0.2non-linear manifold modeling · 0.1generative modeling · 0.1shape-from-silhouette · 0.1tree-structured belief network · 0.1probabilistic model · 0.1parametric global model · 0.1non-parametric texture model · 0.1epitome · 0.1bayesian marginalization · 0.1fiducial marker tracking · 0.1view synthesis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Optimizing Deeper Transformers on Small DatasetsabstractPeng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Simon J.D. Prince, Yanshuai Cao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dhruv Kumar 0005, Wei Yang 0017, Wenjie Zi, Keyi Tang, Chenyang Huang 0001, Jackie Chi Kit Cheung, Simon Prince, Yanshuai Cao |
ACL/IJCNLP (1) | 8 |
| 2021 | Normalizing Flows: An Introduction and Review of Current MethodsabstractNormalizing Flows are generative models which produce tractable distributions where both sampling and density evaluation can be efficient and exact. The goal of this survey article is to give a coherent and comprehensive review of the literature around the construction and use of Normalizing Flows for distribution learning. We aim to provide context and explanation of the models, review current state-of-the-art literature, and identify open questions and promising future directions. Ivan Kobyzev, Simon Prince, Marcus A. Brubaker |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Modeling object appearance using Context-Conditioned Component AnalysisabstractSubspace models have been very successful at modeling the appearance of structured image datasets when the visual objects have been aligned in the images (e.g., faces). Even with extensions that allow for global transformations or dense warps of the image, the set of visual objects whose appearance may be modeled by such methods is limited. They are unable to account for visual objects where occlusion leads to changing visibility of different object parts (without a strict layered structure) and where a one-to-one mapping between parts is not preserved. For example bunches of bananas contain different numbers of bananas but each individual banana shares an appearance subspace. In this work we remove the image space alignment limitations of existing subspace models by conditioning the models on a shape dependent context that allows for the complex, non-linear structure of the appearance of the visual object to be captured and shared. This allows us to exploit the advantages of subspace appearance models with non-rigid, deformable objects whilst also dealing with complex occlusions and varying numbers of parts. We demonstrate the effectiveness of our new model with examples of structured inpainting and appearance transfer. Daniyar Turmukhambetov, Neill D. F. Campbell, Simon Prince, Jan Kautz |
CVPR | 3 |
| 2012 | Probabilistic Models for Inference about IdentityabstractMany face recognition algorithms use "distance-based" methods: Feature vectors are extracted from each face and distances in feature space are compared to determine matches. In this paper, we argue for a fundamentally different approach. We consider each image as having been generated from several underlying causes, some of which are due to identity (latent identity variables, or LIVs) and some of which are not. In recognition, we evaluate the probability that two faces have the same underlying identity cause. We make these ideas concrete by developing a series of novel generative models which incorporate both within-individual and between-individual variation. We consider both the linear case, where signal and noise are represented by a subspace, and the nonlinear case, where an arbitrary face manifold can be described and noise is position-dependent. We also develop a "tied" version of the algorithm that allows explicit comparison of faces across quite different viewing conditions. We demonstrate that our model produces results that are comparable to or better than the state of the art for both frontal face recognition and face recognition under varying pose. Simon Prince, Yun Fu 0002, Umar Mohammed, James H. Elder |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Interactive Lesion Segmentation with Shape Priors From Offline and Online LearningabstractIn medical image segmentation, tumors and other lesions demand the highest levels of accuracy but still call for the highest levels of manual delineation. One factor holding back automatic segmentation is the exemption of pathological regions from shape modelling techniques that rely on high-level shape information not offered by lesions. This paper introduces two new statistical shape models (SSMs) that combine radial shape parameterization with machine learning techniques from the field of nonlinear time series analysis. We then develop two dynamic contour models (DCMs) using the new SSMs as shape priors for tumor and lesion segmentation. From training data, the SSMs learn the lower level shape information of boundary fluctuations, which we prove to be nevertheless highly discriminant. One of the new DCMs also uses online learning to refine the shape prior for the lesion of interest based on user interactions. Classification experiments reveal superior sensitivity and specificity of the new shape priors over those previously used to constrain DCMs. User trials with the new interactive algorithms show that the shape priors are directly responsible for improvements in accuracy and reductions in user demand. Tony Shepherd, Simon Prince, Daniel C. Alexander |
IEEE Trans. Medical Imaging | 2 |
| 2010 | Context-based additive logistic model for facial keypoint localizationabstractFacial keypoint localization is an important step for face recognition. The “Average of Synthetic Exact Filter (ASEF)” approach [2] finds a correlation filter for each training image and averages them together. The resulting classifier is efficient as the filtering can be implemented in the Fourier domain and performance is good for frontal images. However, it cannot cope with a range of poses. In this paper, we generalize this approach to find keypoints using a technique that (i) combines together information from training images in a more principled way than averaging, (ii) can be extended to form non-linear combinations of filters and (iii) can adapt based on context (e.g. pose). These innovations are presented within a greedy boosting-style probabilistic framework. We demonstrate state of the art performance of these algorithms using a challenging data set. Jonathan Warrell, Jania Aghajanian, Simon Prince |
BMVC | 4 |
| 2010 | StyP-Boost: A Bilinear Boosting Algorithm for Learning Style-Parameterized ClassifiersabstractWe introduce a novel bilinear boosting algorithm, which extends the multi-class boosting framework of JointBoost to optimize a bilinear objective function. This allows style parameters to be introduced to aid classification, where style is any factor which the classes vary with systematically, modeled by a vector quantity. The algorithm allows learning to take place across different styles. We apply this Style Parameterized Boosting framework (StyP-Boost) to two object class segmentation tasks: road surface segmentation and general scene parsing. In the former the style parameters represent global surface appearance, and in the latter the probability of belonging to a scene-class. We show how our framework improves on 1) learning without style, and 2) learning independent classifiers within each style. Further, we achieve state-of-the-art results on the Corel database for scene parsing. Jonathan Warrell, Philip Torr 0001, Simon Prince |
BMVC | 3 |
| 2010 | "Lattice Cut" - Constructing superpixels using layer constraintsabstractUnsupervised over-segmentation of an image into super-pixels is a common preprocessing step for image parsing algorithms. Superpixels are used as both regions of support for feature vectors and as a starting point for the final segmentation. Recent algorithms that construct superpixels that conform to a regular grid (or superpixel lattice) have used greedy solutions. In this paper we show that we can construct a globally optimal solution in either the horizontal or vertical direction using a single graph cut. The solution takes into account both edges in the image, and the coherence of the resulting superpixel regions. We show that our method outperforms existing algorithms for computing superpixel lattices. Additionally, we show that performance can be comparable or better than other contemporary segmentation algorithms which are not constrained to produce a lattice. Alastair Philip Moore, Simon Prince, Jonathan Warrell |
CVPR | 2 |
| 2010 | CUDA Implementation of Deformable Pattern Recognition and its Application to MNIST Handwritten Digit DatabaseabstractIn this study we propose a deformable pattern recognition method with CUDA implementation. In order to achieve the proper correspondence between foreground pixels of input and prototype images, a pair of distance maps are generated from input and prototype images, whose pixel values are given based on the distance to the nearest foreground pixel. Then a regularization technique computes the horizontal and vertical displacements based on these distance maps. The dissimilarity is measured based on the eight-directional derivative of input and prototype images in order to leverage characteristic information on the curvature of line segments that might be lost after the deformation. The prototype-parallel displacement computation on CUDA and the gradual prototype elimination technique are employed for reducing the computational time without sacrificing the accuracy. A simulation shows that the proposed method with the k-nearest neighbor classifier gives the error rate of 0.57% for the MNIST handwritten digit database. Yoshiki Mizukami, Katsumi Tadamura, Jonathan Warrell, Simon Prince |
ICPR | 5 |
| 2009 | Face Pose Estimation in Uncontrolled EnvironmentsabstractAutomatic estimation of head pose faciliates human facial analysis. It has widespread applications such as, gaze direction detection, video teleconferencing and human computer interaction (HCI). It can also be integrated in a multi-view face detection and recognition system. Most current methods estimate pose in a limited range or treat pose as a classification problem by assigning the face to one of many discrete poses [1,2]. Mainly tested on images taken in controlled environments e.g. the FacePix dataset [3] (Fig. 1a). Jania Aghajanian, Simon Prince |
BMVC | 2 |
| 2009 | Vistas: Hierarchial Boundary priors using Multiscale Conditional Random FieldsabstractBoundary detection is a fundamental problem in computer vision. However, bound-ary detection is difficult as it involves integrating multiple cues (intensity, color, texture) as well as trying to incorporate object class or scene level descriptions to mitigate the am-biguity of the local signal. In this paper we investigate incorporating a priori information into boundary detection. We learn a probabilistic model that describes a prior for object boundaries over small patches of the image. We then incorporate this boundary model into a mixture of multiscale conditional random fields, where the mixture components represent different contexts formed by clustering overall spatial distributions of bound-aries across images and image regions (vistas). We demonstrate this approach using challenging real-world road scenes. Importantly, we show that recent spectral methods that have been used in state-of-the-art boundary detection algorithms do not generalize well to these complex scenes. We show that our algorithm successfully learns these boundary distributions and can exploit this knowledge to improve state-of-the-art bound-ary detectors. 1 Jonathan Warrell, Alastair Philip Moore, Simon Prince |
BMVC | 3 |
| 2009 | Joint and implicit registration for face recognitionabstractContemporary face recognition algorithms rely on precise localization of keypoints (corner of eye, nose etc.). Unfortunately, finding keypoints reliably and accurately remains a hard problem. In this paper we pose two questions. First, is it possible to exploit the gallery image in order to find keypoints in the probe image? For instance, consider finding the left eye in the probe image. Rather than using a generic eye model, we use a model that is informed by the appearance of the eye in the gallery image. To this end we develop a probabilistic model which combines recognition and keypoint localization. Second, is it necessary to localize keypoints? Alternatively we can consider keypoint position as a hidden variable which we marginalize over in a Bayesian manner. We demonstrate that both of these innovations improve performance relative to conventional methods in both frontal and cross-pose face recognition. Simon Prince |
CVPR | 2 |
| 2009 | Epitomized priors for multi-labeling problemsabstractImage parsing remains difficult due to the need to combine local and contextual information when labeling a scene. We approach this problem by using the epitome as a prior over label configurations. Several properties make it suited to this task. First, it allows a condensed patch-based representation. Second, efficient E-M based learning and inference algorithms can be used. Third, non-stationarity is easily incorporated. We consider three existing priors, and show how each can be extended using the epitome. The simplest prior assumes patches of labels are drawn independently from either a mixture model or an epitome. Next we investigate a `conditional epitome' model, which substitutes an epitome for a conditional mixture model. Finally, we develop an `epitome tree' model, which combines the epitome with a tree structured belief network prior. Each model is combined with a per-pixel classifier to perform segmentation. In each case, the epitomized form of the prior provides superior segmentation performance, with the epitome tree performing best overall. We also apply the same models to denoising binary images, with similar results. Jonathan Warrell, Simon Prince, Alastair Philip Moore |
CVPR | 2 |
| 2009 | Patch-based within-object classificationabstractAdvances in object detection have made it possible to collect large databases of certain objects. In this paper we exploit these datasets for within-object classification. For example, we classify gender in face images, pose in pedestrian images and phenotype in cell images. Previous work has mainly targeted the above tasks individually using object specific representations. Here, we propose a general Bayesian framework for within-object classification. Images are represented as a regular grid of non-overlapping patches. In training, these patches are approximated by a predefined library. In inference, the choice of approximating patch determines the classification decision. We propose a Bayesian framework in which we marginalize over the patch frequency parameters to provide a posterior probability for the class. We test our algorithm on several challenging “real world” databases. Jania Aghajanian, Jonathan Warrell, Simon Prince, Jennifer L. Rohn, Buzz Baum |
ICCV | 3 |
| 2009 | Scene shape priors for superpixel segmentationabstractUnsupervised over-segmentation of an image into super-pixels is a common preprocessing step for image parsing algorithms. Superpixels are used as both regions of support for feature vectors and as a starting point for the final segmentation. In this paper we investigate incorporating a priori information into superpixel segmentations. We learn a probabilistic model that describes the spatial density of the object boundaries in the image. We then describe an over-segmentation algorithm that partitions this density roughly equally between superpixels whilst still attempting to capture local object boundaries. We demonstrate this approach using road scenes where objects in the center of the image tend to be more distant and smaller than those at the edge. We show that our algorithm successfully learns this foveated spatial distribution and can exploit this knowledge to improve the segmentation. Lastly, we introduce a new metric for evaluating vision labeling problems. We measure performance on a challenging real-world dataset and illustrate the limitations of conventional evaluation metrics. Alastair Philip Moore, Simon Prince, Jonathan Warrell, Umar Mohammed, Graham Jones |
ICCV | 2 |
| 2009 | Gender classification in uncontrolled settings using additive logistic modelsabstractMany previous studies have investigated gender classification in well-lit frontal images. In this paper we consider images where the pose, expression and lighting are relatively unconstrained. We localize faces using a standard sliding-window detector. We preprocess the facial region by convolving with Gabor filters at at four scales and four orientations. We sample these responses and concatenate them to form a feature vector. We develop a classifier based on an additive sum of non-linear functions of one-dimensional projections of the data. In particular we investigate arc tangent and weighted sums of Gaussians. We describe a training method based on increasing the binomial log likelihood. We demonstrate that our system on two databases and show that it performs well relative to the state of the art. Simon Prince, Jania Aghajanian |
ICIP | 1 |
| 2009 | Labelfaces: Parsing facial features by multiclass labeling with an epitome priorabstractWe consider the problem of parsing facial features from an image labeling perspective. We learn a per-pixel unary classifier, and a prior over expected label configurations, allowing us to estimate a dense labeling of facial images by part (e.g. hair, mouth, moustache, hat). This approach deals naturally with large variations in shape and appearance characteristic of unconstrained facial images, and also the problem of detecting classes that may be present or absent. We use an Adaboost-based unary classifier, and develop a family of priors based on `epitomes' which are shown to be particularly effective in capturing the non-stationary aspects of face label distributions. Jonathan Warrell, Simon Prince |
ICIP | 2 |
| 2009 | Visio-lization: generating novel facial imagesabstractOur goal is to generate novel realistic images of faces using a model trained from real examples. This model consists of two components: First we consider face images as samples from a texture with spatially varying statistics and describe this texture with a local non-parametric model. Second, we learn a parametric global model of all of the pixel values. To generate realistic faces, we combine the strengths of both approaches and condition the local non-parametric model on the global parametric model. We demonstrate that with appropriate choice of local and global models it is possible to reliably generate new realistic face images that do not correspond to any individual in the training data. We extend the model to cope with considerable intra-class variation (pose and illumination). Finally, we apply our model to editing real facial images: we demonstrate image in-painting, interactive techniques for improving synthesized images and modifying facial expressions. Umar Mohammed, Simon Prince, Jan Kautz |
ACM Trans. Graph. | 2 |
| 2008 | Superpixel latticesabstractUnsupervised over-segmentation of an image into superpixels is a common preprocessing step for image parsing algorithms. Ideally, every pixel within each superpixel region will belong to the same real-world object. Existing algorithms generate superpixels that forfeit many useful properties of the regular topology of the original pixels: for example, the nth superpixel has no consistent position or relationship with its neighbors. We propose a novel algorithm that produces superpixels that are forced to conform to a grid (a regular superpixel lattice). Despite this added topological constraint, our algorithm is comparable in terms of speed and accuracy to alternative segmentation approaches. To demonstrate this, we use evaluation metrics based on (i) image reconstruction (ii) comparison to human-segmented images and (iii) stability of segmentation over subsequent frames of video sequences. Alastair Philip Moore, Simon Prince, Jonathan Warrell, Umar Mohammed, Graham Jones |
CVPR | 2 |
| 2008 | Mosaicfaces: a discrete representation for face recognitionabstractMost face recognition algorithms use a "distance- based" approach: gallery and probe images are projected into a low dimensional feature space and decisions about matching are based on distance in this space. In this paper we use a very different representation, where each face is approximated by a regular grid of patches (a mosaicface). Each of these patches is chosen from a library. Faces are now represented as a list of indices to this library. Since there is no obvious way to measure distance between two such lists, we use a probabilistic approach in which the observed face data is explained by a generative model. There are two phases: (i) Learning - we estimate library contents and associated variability (noise), (ii) Recognition - we evaluate the probability that probe and gallery images were generated from the same library patches. Our method performs significantly better than contemporary approaches in the presence of large illumination changes. Variation in viewing conditions is handled by extending this model to learn equivalences between multiple patch appearances. We demonstrate that our method provides a major improvement on the lighting subset of the XM2VTS database compared to "distance-based" methods. Jania Aghajanian, Simon Prince |
WACV | 2 |
| 2008 | Tied Factor Analysis for Face Recognition across Large Pose DifferencesabstractFace recognition algorithms perform very unreliably when the pose of the probe face is different from the gallery face: typical feature vectors vary more with pose than with identity. We propose a generative model that creates a one-to-many mapping from an idealized "identity" space to the observed data space. In identity space, the representation for each individual does not vary with pose. We model the measured feature vector as being generated by a pose-contingent linear transformation of the identity variable in the presence of Gaussian noise. We term this model "tied" factor analysis. The choice of linear transformation (factors) depends on the pose, but the loadings are constant (tied) for a given individual. We use the EM algorithm to estimate the linear transformations and the noise parameters from training data. We propose a probabilistic distance metric which allows a full posterior over possible matches to be established. We introduce a novel feature extraction process and investigate recognition performance using the FERET, XM2VTS and PIE databases. Recognition performance compares favourably to contemporary approaches. Simon Prince, James H. Elder, Jonathan Warrell, Fatima M. Felisberti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Probabilistic Linear Discriminant Analysis for Inferences About IdentityabstractMany current face recognition algorithms perform badly when the lighting or pose of the probe and gallery images differ. In this paper we present a novel algorithm designed for these conditions. We describe face data as resulting from a generative model which incorporates both within-individual and between-individual variation. In recognition we calculate the likelihood that the differences between face images are entirely due to within-individual variability. We extend this to the non-linear case where an arbitrary face manifold can be described and noise is position-dependent. We also develop a "tied" version of the algorithm that allows explicit comparison across quite different viewing conditions. We demonstrate that our model produces state of the art results for (i) frontal face recognition (ii) face recognition under varying pose. Simon Prince, James H. Elder |
ICCV | 1 |
| 2007 | Pre-Attentive and Attentive Detection of Humans in Wide-Field Scenes
James H. Elder, Simon Prince, Yuqian Hou, Mikhail Sizintsev, E. Olevskiy |
Int. J. Comput. Vis. | 2 |
| 2006 | Tied Factor Analysis for Face Recognition Across Large Pose ChangesabstractAbstract—Face recognition algorithms perform very unreliably when the pose of the probe face is different from the gallery face: typical feature vectors vary more with pose than with identity. We propose a generative model that creates a one-to-many mapping from an idealized “identity ” space to the observed data space. In identity space, the representation for each individual does not vary with pose. We model the measured feature vector as being generated by a pose-contingent linear transformation of the identity variable in the presence of Gaussian noise. We term this model “tied ” factor analysis. The choice of linear transformation (factors) depends on the pose, but the loadings are constant (tied) for a given individual. We use the EM algorithm to estimate the linear transformations and the noise parameters from training data. We propose a probabilistic distance metric that allows a full posterior over possible matches to be established. We introduce a novel feature extraction process and investigate recognition performance by using the FERET, XM2VTS, and PIE databases. Recognition performance compares favorably with contemporary approaches. Index Terms—Computing methodologies, pattern recognition, applications, face and gesture recognition. Ç 1 Simon Prince, James H. Elder |
BMVC | 1 |
| 2005 | Creating Invariance to "Nuisance Parameters" in Face RecognitionabstractA major goal for face recognition is to identify faces where the pose of the probe is different from the stored face. Typical feature vectors vary more with pose than with identity, leading to very poor recognition performance. We propose a non-linear many-to-one mapping from a conventional feature space to a new space constructed so that each individual has a unique feature vector regardless of pose. Training data is used to implicitly parameterize the position of the multi-dimensional face manifold by pose. We introduce a co-ordinate transform, which depends on the position on the manifold. This transform is chosen so that different poses of the same face are mapped to the same feature vector. The same approach is applied to illumination changes. We investigate different methods for creating features, which are invariant to both pose and illumination. We provide a metric to assess the discriminability of the resulting features. Our technique increases the discriminability of faces under unknown pose and lighting compared to contemporary methods. Simon Prince, James H. Elder |
CVPR (2) | 1 |
| 2005 | Statistical Cue Integration for Foveated Wide-Field SurveillanceabstractReliable wide-field detection of human activity is an unsolved problem. The main difficulty is that low resolution and the unconstrained nature of realistic environments and human behaviour make form cues unreliable. Here we argue that reliability in far- or wide-field detection can still be achieved by probabilistic combination of multiple weak but complementary visual cues that do not depend on detailed form analysis. To demonstrate, we describe a real-time Bayesian algorithm for localizing human activity in relatively unconstrained scenes, using motion, background subtraction and skin colour cues. Fast sampling of scale space is achieved using integral images and a flexible norm that can handle sparse cues without loss of statistical power. We show that the probabilistic approach far outperforms a representative logical approach in which skin and background subtraction classifiers are combined conjunctively. Our method is currently used in a pre-attentive human activity sensor, generating saccadic targets for an attentive foveated vision system that reliably fixates faces over a 130 deg field of view, allowing high-resolution capture of facial images over a large dynamic scene. Simon Prince, James H. Elder, Yuqian Hou, Mikhail Sizintsev, Yevgen Olevskiy |
CVPR (2) | 1 |
| 2005 | Live three-dimensional content for augmented realityabstractWe describe an augmented reality system for superimposing three-dimensional (3-D) live content onto two-dimensional fiducial markers in the scene. In each frame, the Euclidean transformation between the marker and the camera is estimated. The equivalent virtual view of the live model is then generated and rendered into the scene at interactive speeds. The 3-D structure of the model is calculated using a fast shape-from-silhouette algorithm based on the outputs of 15 cameras surrounding the subject. The novel view is generated by projecting rays through each pixel of the desired image and intersecting them with the 3-D structure. Pixel color is estimated by taking a weighted sum of the colors of the projections of this 3-D point in nearby real camera images. Using this system, we capture live human models and present them via the augmented reality interface at a remote location. We can generate 384/spl times/288 pixel images of the models at 25 fps, with a latency of <100 ms. The result gives the strong impression that the model is a real 3-D part of the scene. Farzam Farbiz, Adrian David Cheok, Wei Liu 0009, Zhiying Zhou, Ke Xu 0004, Simon Prince, Mark Billinghurst, Hirokazu Kato 0001 |
IEEE Trans. Multim. | 6 |
| 2003 | Visual registration for unprepared augmented reality environments
Ke Xu 0004, Simon Prince, Adrian David Cheok, Krishnamoorthy Ganesh Kumar |
Pers. Ubiquitous Comput. | 2 |
| 2002 | 3-D live: real time interaction for mixed realityabstractWe describe a real-time 3-D augmented reality video- conferencing system. With this technology, an observer sees the real world from his viewpoint, but modified so that the image of a remote collaborator is rendered into the scene. We register the image of the collaborator with the world by estimating the 3-D transformation between the camera and a fiducial marker. We describe a novel shape- from-silhouette algorithm, which generates the appropriate view of the collaborator and the associated depth map at 30 fps. When this view is superimposed upon the real world, it gives the strong impression that the collaborator is a real part of the scene. We also demonstrate interaction in virtual environments with a live fully 3-D collaborator. Finally, we consider interaction between users in the real world and collaborators in a virtual space, using a tangible AR interface. Simon Prince, Adrian David Cheok, Farzam Farbiz, Todd Williamson, Nikolas Johnson, Mark Billinghurst, Hirokazu Kato 0001 |
CSCW | 1 |
| 2002 | Interactive Theatre Experience in Embodied + Wearable Mixed Reality SpaceabstractThis paper presents an interactive theatre based on an embodied mixed reality space and wearable computers. Embodied computing mixed reality spaces integrate ubiquitous computing, tangible interaction and social computing within a mixed reality space, which enables intuitive interaction with physical world and virtual world. We believe it has potential advantages to support novel interactive theatre experiences. Therefore, we explored the novel interactive theatre experience supported in the embodied mixed reality space, and implemented live 3D characters to interact with user in such a system. Adrian David Cheok, Xubo Yang, Simon Prince, Fong Siew Wan, Mark Billinghurst, Hirokazu Kato 0001 |
ISMAR | 4 |
| 2002 | Interactive Theatre Experience in Embodied + Wearable Mixed Reality Space
Adrian David Cheok, Xubo Yang, Simon Prince, Fong Siew Wan, Mark Billinghurst, Hirokazu Kato 0001 |
ISMAR | 4 |
| 2002 | Online 6 DOF Augmented Reality Registration from Natural FeaturesabstractWe present a complete scalable system for 6 DOF camera tracking based on natural features. Crucially, the calculation is based only on pre-captured reference images and previous estimates of the camera pose and is hence suitable for online applications. We match natural features in the current frame to two spatially separated reference images. We overcome the wide baseline matching problem by matching to the previous frame and transferring point positions to the reference images. We then minimize deviations from the two-view and three-view constraints between the reference images and the current frame as a function of camera position parameters. We stabilize this calculation using a recursive form of temporal regularization that is similar in spirit to the Kalman filter. We can track camera pose over hundreds of frames and realistically integrate virtual objects with only slight jitter. Kar Wee Chia, Adrian David Cheok, Simon Prince |
ISMAR | 3 |
| 2002 | 3D Live: Real Time Captured Content for Mixed RealityabstractWe present a complete system for live capture of 3D content and simultaneous presentation in augmented reality. The user sees the real world from his viewpoint, but modified so that the image of a remote collaborator is rendered into the scene. Fifteen cameras surround the collaborator, and the resulting video streams are used to construct a three-dimensional model of the subject using a shape-from-silhouette algorithm. Users view a two-dimensional fiducial marker using a video-see-through augmented reality interface. The geometric relationship between the marker and head-mounted camera is calculated, and the equivalent view of the subject is computed and drawn into the scene. Our system can generate 384 /spl times/ 288 pixel images of the models at 25 fps, with a latency of < 100 ms. The result gives the strong impression that the subject is a real part of the 3D scene. We demonstrate applications of this system in 3D videoconferencing and entertainment. Simon Prince, Adrian David Cheok, Farzam Farbiz, Todd Williamson, Nikolas Johnson, Mark Billinghurst, Hirokazu Kato 0001 |
ISMAR | 1 |
| 2002 | 3D Live: Real Time Captured Content for Mixed Reality
Simon Prince, Adrian David Cheok, Farzam Farbiz, Todd Williamson, Nikolas Johnson, Mark Billinghurst, Hirokazu Kato 0001 |
ISMAR | 1 |