John M. Winn

dblp:w/JohnMWinn · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
28 papers
Image recognition and object detection · 25% Probabilistic and Bayesian machine learning · 25% Segmentation and scene understanding · 18%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 100%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 62% Multimedia analysis and retrieval · 29% Rendering · 6%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 54% Medical and health informatics · 46%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.662015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
3D LayoutCRF for Multi-View Object Class Recognition and Segmentation · CVPR 2007
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.452013
Learning to Pass Expectation Propagation Messages · NIPS 2013
Gates · NIPS 2008
Hybrid learning of large jigsaws · CVPR 2007
Machine learning › Probabilistic and Bayesian machine learning
boltzmann machine
0.322014
The Shape Boltzmann Machine: A Strong Model of Object Shape · Int. J. Comput. Vis. 2014
The Shape Boltzmann Machine: A strong model of object shape · CVPR 2012
Computer vision › Image recognition and object detection › object detection › object detection evaluation
object detection benchmark
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Performance modeling and evaluation
benchmarking
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Performance modeling and evaluation › benchmarking › machine learning benchmarking
vision benchmark
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Computer vision › Segmentation and scene understanding
video segmentation
0.332011
Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Tree-based Classifiers for Bilayer Video Segmentation · CVPR 2007
Escaping local minima through hierarchical model selection: Automatic object discovery, segmentation, and tracking in video · CVPR (1) 2006
Computer vision › Image recognition and object detection
object recognition
0.242009
Hybrid learning of large jigsaws · CVPR 2007
Incorporating On-demand Stereo for Real Time Recognition · CVPR 2007
TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation · ECCV (1) 2006
Computer vision › 3D vision › 3d shape modeling
object shape modeling
0.222014
The Shape Boltzmann Machine: A strong model of object shape · CVPR 2012
The Shape Boltzmann Machine: A Strong Model of Object Shape · Int. J. Comput. Vis. 2014
Computer vision › 3D vision › 3d scene modeling › scene representation
epitome model
0.222009
Epitomic Location Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Epitomic location recognition · CVPR 2008
Computer vision › 3D vision › visual localization
location recognition
0.222009
Epitomic Location Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Epitomic location recognition · CVPR 2008
Computer vision › 3D vision
visual localization
0.222009
Epitomic Location Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Epitomic location recognition · CVPR 2008
Machine learning › Learning theory
classification
0.212013
Decision Jungles: Compact and Rich Models for Classification · NIPS 2013
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.212013
Learning to Pass Expectation Propagation Messages · NIPS 2013
Machine learning › Efficient and distributed learning
model compression
0.212013
Decision Jungles: Compact and Rich Models for Classification · NIPS 2013
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.212013
Decision Jungles: Compact and Rich Models for Classification · NIPS 2013
Machine learning › Kernel, tree and ensemble methods › ensemble learning
tree ensembles
0.212013
Decision Jungles: Compact and Rich Models for Classification · NIPS 2013
Medical and health informatics › clinical data analysis
phenotyping
0.112012
ShapePheno: unsupervised extraction of shape phenotypes from biological image collections · Bioinform. 2012
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
variational message passing
0.122008
Gates · NIPS 2008
Variational Message Passing · J. Mach. Learn. Res. 2005
Computer vision › Segmentation and scene understanding
object segmentation
0.122007
3D LayoutCRF for Multi-View Object Class Recognition and Segmentation · CVPR 2007
The Layout Consistent Random Field for Recognizing and Segmenting Partially Occluded Objects · CVPR (1) 2006
Computer vision › Image recognition and object detection › image classification
object classification
0.122007
3D LayoutCRF for Multi-View Object Class Recognition and Segmentation · CVPR 2007
Object Categorization by Learned Universal Visual Dictionary · ICCV 2005
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
bi-layer segmentation
0.112011
Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.112011
Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › Video understanding and tracking
motion segmentation
0.112011
Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.122008
Gates · NIPS 2008
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112009
TextonBoost for Image Understanding: Multi-Class Object Recognition and Segmentation by Jointly Modeling Texture, Layout, and Context · Int. J. Comput. Vis. 2009
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.122005
Variational Message Passing · J. Mach. Learn. Res. 2005
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.122005
Variational Message Passing · J. Mach. Learn. Res. 2005
VIBES: A Variational Inference Engine for Bayesian Networks · NIPS 2002
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
factor graphs
0.112008
Gates · NIPS 2008
Computer vision › 3D vision
feature matching
0.112008
Epitomic location recognition · CVPR 2008

Methods — techniques the papers use, named apart from their topics

random forest · 0.4conditional random field · 0.3shape boltzmann machine · 0.2just-in-time learning · 0.2generative model · 0.2epitomic image analysis · 0.2neural network · 0.2entropy minimization · 0.2discriminative learning · 0.2decision directed acyclic graph · 0.2probabilistic machine learning · 0.1deformable template model · 0.1image segmentation · 0.1image blending · 0.13d object retrieval · 0.1visual dictionary · 0.1gaussian mixture model · 0.1variational inference · 0.0
YearPublicationVenuePosition
2020 Learning Direct Optimization for scene understanding
Lukasz Romaszko, Christopher K. I. Williams, John M. Winn
Pattern Recognit.3
2015 Consensus Message Passing for Layered Graphical Models
abstract
Generative models provide a powerful framework for probabilistic reasoning. However, in many domains their use has been hampered by the practical difficulties of inference. This is particularly the case in computer vision, where models of the imaging process tend to be large, loopy and layered. For this reason bottom-up conditional models have traditionally dominated in such domains. We find that widely-used, general-purpose message passing inference algorithms such as Expectation Propagation (EP) and Variational Message Passing (VMP) fail on the simplest of vision models. With these models in mind, we introduce a modification to message passing that learns to exploit their layered structure by passing ’consensus’ messages that guide inference towards good solutions. Experiments on a variety of problems show that the proposed technique leads to significantly more accurate inference results, not only when compared to standard EP and VMP, but also when compared to competitive bottom-up conditional models.
Varun Jampani, S. M. Ali Eslami, Daniel Tarlow, Pushmeet Kohli, John M. Winn
AISTATS5
2015 The Pascal Visual Object Classes Challenge: A Retrospective
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John M. Winn, Andrew Zisserman
Int. J. Comput. Vis.5
2014 Just-In-Time Learning for Fast and Flexible Inference
S. M. Ali Eslami, Daniel Tarlow, Pushmeet Kohli, John M. Winn
NIPS4
2014 The Shape Boltzmann Machine: A Strong Model of Object Shape
S. M. Ali Eslami, Nicolas Heess, Christopher K. I. Williams, John M. Winn
Int. J. Comput. Vis.4
2013 Structural Expectation Propagation (SEP): Bayesian structure learning for networks with latent variables
abstract
Learning the structure of discrete Bayesian networks has been the subject of extensive research in machine learning, with most Bayesian approaches focusing on fully observed networks. One of few the methods that can handle networks with latent variables is the "structural EM algorithm" which interleaves greedy structure search with the estimation of latent variables and parameters, maintaining a single best network at each step. We introduce Structural Expectation Propagation (SEP), an extension of EP which can infer the structure of Bayesian networks having latent variables and missing data. SEP performs variational inference in a joint model of structure, latent variables, and parameters, offering two advantages: (i) it accounts for uncertainty in structure and parameter values when making local distribution updates (ii) it returns a variational distribution over network structures rather than a single network. We demonstrate the performance of SEP both on synthetic problems and on real-world clinical data.
Nevena Lazic, Christopher M. Bishop, John M. Winn
AISTATS3
2013 Learning to Pass Expectation Propagation Messages
abstract
Expectation Propagation (EP) is a popular approximate posterior inference algorithm that often provides a fast and accurate alternative to sampling-based methods. However, while the EP framework in theory allows for complex non-Gaussian factors, there is still a significant practical barrier to using them within EP, because doing so requires the implementation of message update operators, which can be difficult and require hand-crafted approximations. In this work, we study the question of whether it is possible to automatically derive fast and accurate EP updates by learning a discriminative model e.g., a neural network or random forest) to map EP message inputs to EP message outputs. We address the practical concerns that arise in the process, and we provide empirical analysis on several challenging and diverse factors, indicating that there is a space of factors where this approach appears promising.
Nicolas Heess, Daniel Tarlow, John M. Winn
NIPS3
2013 Decision Jungles: Compact and Rich Models for Classification
abstract
Randomized decision trees and forests have a rich history in machine learning and have seen considerable success in application, perhaps particularly so for computer vision. However, they face a fundamental limitation: given enough data, the number of nodes in decision trees will grow exponentially with depth. For certain applications, for example on mobile or embedded processors, memory is a limited resource, and so the exponential growth of trees limits their depth, and thus their potential accuracy. This paper proposes decision jungles, revisiting the idea of ensembles of rooted decision directed acyclic graphs (DAGs), and shows these to be compact and powerful discriminative models for classification. Unlike conventional decision trees that only allow one path to every node, a DAG in a decision jungle allows multiple paths from the root to each leaf. We present and compare two new node merging algorithms that jointly optimize both the features and the structure of the DAGs efficiently. During training, node splitting and node merging are driven by the minimization of exactly the same objective function, here the weighted sum of entropies at the leaves. Results on varied datasets show that, compared to decision forests and several other baselines, decision jungles require dramatically less memory while considerably improving generalization.
Jamie Shotton, Toby Sharp, Pushmeet Kohli, Sebastian Nowozin, John M. Winn, Antonio Criminisi
NIPS5
2012 The Shape Boltzmann Machine: A strong model of object shape
abstract
A good model of object shape is essential in applications such as segmentation, object detection, inpainting and graphics. For example, when performing segmentation, local constraints on the shape can help where the object boundary is noisy or unclear, and global constraints can resolve ambiguities where background clutter looks similar to part of the object. In general, the stronger the model of shape, the more performance is improved. In this paper, we use a type of Deep Boltzmann Machine [22] that we call a Shape Boltzmann Machine (ShapeBM) for the task of modeling binary shape images. We show that the ShapeBM characterizes a strong model of shape, in that samples from the model look realistic and it can generalize to generate samples that differ from training examples. We find that the ShapeBM learns distributions that are qualitatively and quantitatively better than existing models for this task.
S. M. Ali Eslami, Nicolas Heess, John M. Winn
CVPR3
2012 ShapePheno: unsupervised extraction of shape phenotypes from biological image collections
abstract
MOTIVATION: Accurate large-scale phenotyping has recently gained considerable importance in biology. For example, in genome-wide association studies technological advances have rendered genotyping cheap, leaving phenotype acquisition as the major bottleneck. Automatic image analysis is one major strategy to phenotype individuals in large numbers. Current approaches for visual phenotyping focus predominantly on summarizing statistics and geometric measures, such as height and width of an individual, or color histograms and patterns. However, more subtle, but biologically informative phenotypes, such as the local deformation of the shape of an individual with respect to the population mean cannot be automatically extracted and quantified by current techniques. RESULTS: We propose a probabilistic machine learning model that allows for the extraction of deformation phenotypes from biological images, making them available as quantitative traits for downstream analysis. Our approach jointly models a collection of images using a learned common template that is mapped onto each image through a deformable smooth transformation. In a case study, we analyze the shape deformations of 388 guppy fish (Poecilia reticulata). We find that the flexible shape phenotypes our model extracts are complementary to basic geometric measures. Moreover, these quantitative traits assort the observations into distinct groups and can be mapped to polymorphic genetic loci of the sample set. AVAILABILITY: Code is available under: http://bioweb.me/GEBI CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Theofanis Karaletsos, Oliver Stegle, Christine Dreyer, John M. Winn, Karsten M. Borgwardt
Bioinform.4
2012 In Memoriam: Mark Everingham
abstract
Recounts the career and contributions pf Mark Everingham.
Andrew Zisserman, John M. Winn, Andrew W. Fitzgibbon, Luc Van Gool, Josef Sivic, Christopher K. I. Williams, David C. Hogg
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Weakly Supervised Learning of Foreground-Background Segmentation Using Masked RBMs
Nicolas Heess, Nicolas Le Roux, John M. Winn
ICANN (2)3
2011 Learning a Generative Model of Images by Factoring Appearance and Shape
abstract
Computer vision has grown tremendously in the past two decades. Despite all efforts, existing attempts at matching parts of the human visual system's extraordinary ability to understand visual scenes lack either scope or power. By combining the advantages of general low-level generative models and powerful layer-based and hierarchical models, this work aims at being a first step toward richer, more flexible models of images. After comparing various types of restricted Boltzmann machines (RBMs) able to model continuous-valued data, we introduce our basic model, the masked RBM, which explicitly models occlusion boundaries in image patches by factoring the appearance of any patch region from its shape. We then propose a generative model of larger images using a field of such RBMs. Finally, we discuss how masked RBMs could be stacked to form a deep model able to generate more complicated structures and suitable for various tasks such as segmentation or object recognition.
Nicolas Le Roux, Nicolas Heess, Jamie Shotton, John M. Winn
Neural Comput.4
2011 Bilayer Segmentation of Webcam Videos Using Tree-Based Classifiers
abstract
This paper presents an automatic segmentation algorithm for video frames captured by a (monocular) webcam that closely approximates depth segmentation from a stereo camera. The frames are segmented into foreground and background layers that comprise a subject (participant) and other objects and individuals. The algorithm produces correct segmentations even in the presence of large background motion with a nearly stationary foreground. This research makes three key contributions: First, we introduce a novel motion representation, referred to as "motons," inspired by research in object recognition. Second, we propose estimating the segmentation likelihood from the spatial context of motion. The estimation is efficiently learned by random forests. Third, we introduce a general taxonomy of tree-based classifiers that facilitates both theoretical and experimental comparisons of several known classification algorithms and generates new ones. In our bilayer segmentation algorithm, diverse visual cues such as motion, motion context, color, contrast, and spatial priors are fused by means of a conditional random field (CRF) model. Segmentation is then achieved by binary min-cut. Experiments on many sequences of our videochat application demonstrate that our algorithm, which requires no initialization, is effective in a variety of scenes, and the segmentation results are comparable to those obtained by stereo systems.
Pei Yin, Antonio Criminisi, John M. Winn, Irfan A. Essa
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 The Pascal Visual Object Classes (VOC) Challenge
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, Andrew Zisserman
Int. J. Comput. Vis.4
2010 A Bayesian Framework to Account for Complex Non-Genetic Factors in Gene Expression Levels Greatly Increases Power in eQTL Studies
abstract
Gene expression measurements are influenced by a wide range of factors, such as the state of the cell, experimental conditions and variants in the sequence of regulatory regions. To understand the effect of a variable of interest, such as the genotype of a locus, it is important to account for variation that is due to confounding causes. Here, we present VBQTL, a probabilistic approach for mapping expression quantitative trait loci (eQTLs) that jointly models contributions from genotype as well as known and hidden confounding factors. VBQTL is implemented within an efficient and flexible inference framework, making it fast and tractable on large-scale problems. We compare the performance of VBQTL with alternative methods for dealing with confounding variability on eQTL mapping datasets from simulations, yeast, mouse, and human. Employing Bayesian complexity control and joint modelling is shown to result in more precise estimates of the contribution of different confounding factors resulting in additional associations to measured transcript levels compared to alternative approaches. We present a threefold larger collection of cis eQTLs than previously found in a whole-genome eQTL scan of an outbred human population. Altogether, 27% of the tested probes show a significant genetic association in cis, and we validate that the additional eQTLs are likely to be real by replicating them in different sets of individuals. Our method is the next step in the analysis of high-dimensional phenotype data, and its application has revealed insights into genetic regulation of gene expression by demonstrating more abundant cis-acting eQTLs in human than previously shown. Our software is freely available online at http://www.sanger.ac.uk/resources/software/peer/.
Oliver Stegle, Leopold Parts, Richard Durbin, John M. Winn
PLoS Comput. Biol.4
2009 TextonBoost for Image Understanding: Multi-Class Object Recognition and Segmentation by Jointly Modeling Texture, Layout, and Context
Jamie Shotton, John M. Winn, Carsten Rother, Antonio Criminisi
Int. J. Comput. Vis.2
2009 Epitomic Location Recognition
abstract
This paper presents a novel method for location recognition, which exploits an epitomic representation to achieve both high efficiency and good generalization. A generative model based on epitomic image analysis captures the appearance and geometric structure of an environment while allowing for variations due to motion, occlusions, and non-Lambertian effects. The ability to model translation and scale invariance together with the fusion of diverse visual features yields enhanced generalization with economical training. Experiments on both existing and new labeled image databases result in recognition accuracy superior to state of the art with real-time computational performance.
Anitha Kannan, Antonio Criminisi, John M. Winn
IEEE Trans. Pattern Anal. Mach. Intell.4
2008 Epitomic location recognition
abstract
This paper presents a novel method for location recognition, which exploits an epitomic representation to achieve both high efficiency and good generalization. A generative model based on epitomic image analysis captures the appearance and geometric structure of an environment while allowing for variations due to motion, occlusions and non-Lambertian effects. The ability to model translation and scale invariance together with the fusion of diverse visual features yield enhanced generalization with economical training. Experiments on both existing and new labelled image databases result in recognition accuracy superior to state of the art with real-time computational performance.
Anitha Kannan, Antonio Criminisi, John M. Winn
CVPR4
2008 Immune System Modeling with Infer.NET
abstract
Graphical models allow scientific prior knowledge to be incorporated into the statistical analysis of data, whilst also providing a vivid way to represent and communicate this knowledge. In this paper we develop a graphical model of the immune system as a means of analyzing immunological data from the Manchester asthma and allergy study (MAAS). The analysis is achieved using the Infer.NET tool which allows Bayesian inference to be applied automatically to a specified graphical model.Our immune system model consists firstly of a hidden Markov model representing how allergen-specific skin prick tests (SPTs) and serum-specific IgE tests (SITs) change over time. By introducing a latent multinomial variable, we also cluster the children in an unsupervised manner into different sensitization classes. For 2 sensitization classes, the children who are vulnerable to allergies and have a high probability of having asthma (22%) are identified. For 5 sensitization classes, children in the first cluster, those who are vulnerable to allergies, have an even higher probability of having asthma (42%). The second part of the model involves using the inferred sensitization class as a label and 8 exposure variables in a Bayes point machine. Using multiple permutation tests, we conclude that the level of endotoxins and gender have a significant effect on a child's vulnerability to allergies.
Vincent Y. F. Tan, John M. Winn, Angela Simpson, Adnan Custovic
eScience2
2008 Gates
abstract
Gates are a new notation for representing mixture models and context-sensitive independence in factor graphs. Factor graphs provide a natural representation for message-passing algorithms, such as expectation propagation. However, message passing in mixture models is not well captured by factor graphs unless the entire mixture is represented by one factor, because the message equations have a containment structure. Gates capture this containment structure graphically, allowing both the independences and the message-passing equations for a model to be readily visualized. Different variational approximations for mixture models can be understood as different ways of drawing the gates in a model. We present general equations for expectation propagation and variational message passing in the presence of gates.
Tom Minka, John M. Winn
NIPS2
2008 Accounting for Non-genetic Factors Improves the Power of eQTL Studies
Oliver Stegle, Anitha Kannan, Richard Durbin, John M. Winn
RECOMB4
2007 Incorporating On-demand Stereo for Real Time Recognition
abstract
A new method for localising and recognising hand poses and objects in real-time is presented. This problem is important in vision-driven applications where it is natural for a user to combine hand gestures and real objects when interacting with a machine. Examples include using a real eraser to remove words from a document displayed on an electronic surface. In this paper the task of simultaneously recognising object classes, hand gestures and detecting touch events is cast as a single classification problem. A random forest algorithm is employed which adaptively selects and combines a minimal set of appearance, shape and stereo features to achieve maximum class discrimination for a given image. This minimal set leads to both efficiency at run time and good generalisation. Unlike previous stereo works which explicitly construct disparity maps, here the stereo matching costs are used directly as visual cue and only computed on-demand, i.e. only for pixels where they are necessary for recognition. This leads to improved efficiency. The proposed method is assessed on a database of a variety of objects and hand poses selected for interacting on a flat surface in an office environment.
Thomas Deselaers, Antonio Criminisi, John M. Winn, Ankur Agarwal
CVPR3
2007 3D LayoutCRF for Multi-View Object Class Recognition and Segmentation
abstract
We introduce an approach to accurately detect and segment partially occluded objects in various viewpoints and scales. Our main contribution is a novel framework for combining object-level descriptions (such as position, shape, and color) with pixel-level appearance, boundary, and occlusion reasoning. In training, we exploit a rough 3D object model to learn physically localized part appearances. To find and segment objects in an image, we generate proposals based on the appearance and layout of local parts. The proposals are then refined after incorporating object-level information, and overlapping objects compete for pixels to produce a final description and segmentation of objects in the scene. A further contribution is a novel instance penalty, which is handled very efficiently during inference. We experimentally validate our approach on the challenging PASCAL'06 car database.
Derek Hoiem, Carsten Rother, John M. Winn
CVPR3
2007 Hybrid learning of large jigsaws
abstract
A jigsaw is a recently proposed generative model that describes an image as a composition of non-overlapping patches of varying shape, extracted from a latent image. By learning the latent jigsaw image which best explains a set of images, it is possible to discover the shape, size and appearance of repeated structures in the images. A challenge when learning this model is the very large space of possible jigsaw pixels which can potentially be used to explain each image pixel. The previous method of inference for this model scales linearly with the number of jigsaw pixels, making it unusable for learning the large jigsaws needed for many practical applications. In this paper, we make three contributions that enable the learning of large jigsaws -a novel sparse belief propagation algorithm, a hybrid method which significantly improves the sparseness of this algorithm, and a method that uses these techniques to make learning of large jigsaws feasible. We provide detailed analysis of how our hybrid inference method leads to significant savings in memory and computation time. To demonstrate the success of our method, we present experimental results applying large jigsaws to an object recognition task.
Julia A. Lasserre, Anitha Kannan, John M. Winn
CVPR3
2007 Tree-based Classifiers for Bilayer Video Segmentation
abstract
This paper presents an algorithm for the automatic segmentation of monocular videos into foreground and background layers. Correct segmentations are produced even in the presence of large background motion with nearly stationary foreground. There are three key contributions. The first is the introduction of a novel motion representation, "motons", inspired by research in object recognition. Second, we propose learning the segmentation likelihood from the spatial context of motion. The learning is efficiently performed by Random Forests. The third contribution is a general taxonomy of tree-based classifiers, which facilitates theoretical and experimental comparisons of several known classification algorithms, as well as spawning new ones. Diverse visual cues such as motion, motion context, colour, contrast and spatial priors are fused together by means of a conditional random field (CRF) model. Segmentation is then achieved by binary min-cut. Our algorithm requires no initialization. Experiments on many video-chat type sequences demonstrate the effectiveness of our algorithm in a variety of scenes. The segmentation results are comparable to those obtained by stereo systems.
Pei Yin, Antonio Criminisi, John M. Winn, Irfan A. Essa
CVPR3
2007 Photo clip art
abstract
We present a system for inserting new objects into existing photographs by querying a vast image-based object library, pre-computed using a publicly available Internet object database. The central goal is to shield the user from all of the arduous tasks typically involved in image compositing. The user is only asked to do two simple things: 1) pick a 3D location in the scene to place a new object; 2) select an object to insert using a hierarchical menu. We pose the problem of object insertion as a data-driven, 3D-based, context-sensitive object retrieval task. Instead of trying to manipulate the object to change its orientation, color distribution, etc. to fit the new image, we simply retrieve an object of a specified class that has all the required properties (camera pose, lighting, resolution, etc) from our large object library. We present new automatic algorithms for improving object segmentation and blending, estimating true 3D object size and orientation, and estimating scene lighting conditions. We also present an intuitive user interface that makes object insertion fast and simple even for the artistically challenged.
Jean-François Lalonde, Derek Hoiem, Alexei A. Efros, Carsten Rother, John M. Winn, Antonio Criminisi
ACM Trans. Graph.5
2006 Escaping local minima through hierarchical model selection: Automatic object discovery, segmentation, and tracking in video
abstract
Recently, the generative modeling approach to video segmentation has been gaining popularity in the computer vision community. For example, the flexible sprites framework has been studied in, among other references, [11,13,14,24]. In general, detailed generative models are vulnerable to intractability of inference and local minima problems when approximations are made (see, e.g., [25]). Recent approaches to dealing with these problems focused on inference techniques for increasingly more expressive models. Simpler models, on the other hand, while less precise, are often not just faster, but less prone to local minima. In addition, while many different models may be based on similar hidden variables, some models may be more amenable to inference of some of the shared variables, while other models lead to efficient and accurate inference of other components of the hierarchical data description. In this paper, we empirically illustrate that forcing multiple models to share the posterior distribution leads to inference less prone to local minima. We define a set of key hidden variables that describe aspects of the data that we care about. The relationships among these key variables are defined through multiple conditional distribution models on the same pairs of variables, controlled by switch variables. The posterior distribution over the key hidden variables is shared, and inference of the switch variables serves as a mechanism for combinatorial model selection. The key observation here is that while the most expressive model often ends up a winner by the end of the iterative learning of model parameters, early iterations are dominated by simpler model components, and upon convergence, the free energy is lower than the ones reached by switching on all the most complex components from the beginning of the learning. We illustrate the performance of this approach on the unsupervised video segmentation task.
Nebojsa Jojic, John M. Winn, C. Lawrence Zitnick
CVPR (1)2
2006 Discriminative Object Class Models of Appearance and Shape by Correlatons
abstract
This paper presents a new model of object classes which incorporates appearance and shape information jointly. Modeling objects appearance by distributions of visual words has recently proven successful. Here appearancebased models are augmented by capturing the spatial arrangement of visual words. Compact spatial modeling without loss of discrimination is achieved through the introduction of adaptive vector quantized correlograms, which we call correlatons. Efficiency is further improved by means of integral images. The robustness of our new models to geometric transformations, severe occlusions and missing information is also demonstrated. The accuracy of discrimination of the proposed models is assessed with respect to existing databases with large numbers of object classes viewed under general conditions, and shown to outperform appearance-only models.
Silvio Savarese, John M. Winn, Antonio Criminisi
CVPR (2)2
2006 The Layout Consistent Random Field for Recognizing and Segmenting Partially Occluded Objects
abstract
This paper addresses the problem of detecting and segmenting partially occluded objects of a known category. We first define a part labelling which densely covers the object. Our Layout Consistent Random Field (LayoutCRF) model then imposes asymmetric local spatial constraints on these labels to ensure the consistent layout of parts whilst allowing for object deformation. Arbitrary occlusions of the object are handled by avoiding the assumption that the whole object is visible. The resulting system is both efficient to train and to apply to novel images, due to a novel annealed layout-consistent expansion move algorithm paired with a randomised decision tree classifier. We apply our technique to images of cars and faces and demonstrate state-of-the-art detection and segmentation performance even in the presence of partial occlusion.
John M. Winn, Jamie Shotton
CVPR (1)1
2006 Located Hidden Random Fields: Learning Discriminative Parts for Object Detection
Ashish Kapoor, John M. Winn
ECCV (3)2
2006 TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation
Jamie Shotton, John M. Winn, Carsten Rother, Antonio Criminisi
ECCV (1)2
2006 Clustering appearance and shape by learning jigsaws
abstract
Patch-based appearance models are used in a wide range of computer vision ap- plications. To learn such models it has previously been necessary to specify a suitable set of patch sizes and shapes by hand. In the jigsaw model presented here, the shape, size and appearance of patches are learned automatically from the repeated structures in a set of training images. By learning such irregularly shaped ‘jigsaw pieces’, we are able to discover both the shape and the appearance of object parts without supervision. When applied to face images, for example, the learned jigsaw pieces are surprisingly strongly associated with face parts of different shapes and scales such as eyes, noses, eyebrows and cheeks, to name a few. We conclude that learning the shape of the patch not only improves the accuracy of appearance-based part detection but also allows for shape-based part detection. This enables parts of similar appearance but different shapes to be dis- tinguished; for example, while foreheads and cheeks are both skin colored, they have markedly different shapes.
Anitha Kannan, John M. Winn, Carsten Rother
NIPS2
2005 Object Categorization by Learned Universal Visual Dictionary
abstract
This paper presents a new algorithm for the automatic recognition of object classes from images (categorization). Compact and yet discriminative appearance-based object class models are automatically learned from a set of training images. The method is simple and extremely fast, making it suitable for many applications such as semantic image retrieval, Web search, and interactive image editing. It classifies a region according to the proportions of different visual words (clusters in feature space). The specific visual words and the typical proportions in each object are learned from a segmented training set. The main contribution of this paper is twofold: i) an optimally compact visual dictionary is learned by pair-wise merging of visual words from an initially large dictionary. The final visual words are described by GMMs. ii) A novel statistical measure of discrimination is proposed which is optimized by each merge operation. High classification accuracy is demonstrated for nine object classes on photographs of real objects viewed under general lighting conditions, poses and viewpoints. The set of test images used for validation comprise: i) photographs acquired by us, ii) images from the Web and iii) images from the recently released Pascal dataset. The proposed algorithm performs well on both texture-rich objects (e.g. grass, sky, trees) and structure-rich ones (e.g. cars, bikes, planes)
John M. Winn, Antonio Criminisi, Tom Minka
ICCV1
2005 LOCUS: Learning Object Classes with Unsupervised Segmentation
abstract
We address the problem of learning object class models and object segmentations from unannotated images. We introduce LOCUS (learning object classes with unsupervised segmentation) which uses a generative probabilistic model to combine bottom-up cues of color and edge with top-down cues of shape and pose. A key aspect of this model is that the object appearance is allowed to vary from image to image, allowing for significant within-class variation. By iteratively updating the belief in the object's position, size, segmentation and pose, LOCUS avoids making hard decisions about any of these quantities and so allows for each to be refined at any stage. We show that LOCUS successfully learns an object class model from unlabeled images, whilst also giving segmentation accuracies that rival existing supervised methods. Finally, we demonstrate simultaneous recognition and segmentation in novel images using the learned models for a number of object classes, as well as unsupervised object discovery and tracking in video.
John M. Winn, Nebojsa Jojic
ICCV1
2005 Variational Message Passing
abstract
Bayesian inference is now widely established as one of the principal foundations for machine learning. In practice, exact inference is rarely possible, and so a variety of approximation techniques have been developed, one of the most widely used being a deterministic framework called variational inference. In this paper we introduce Variational Message Passing (VMP), a general purpose algorithm for applying variational inference to Bayesian Networks. Like belief propagation, VMP proceeds by sending messages between nodes in the network and updating posterior beliefs using local operations at each node. Each such update increases a lower bound on the log evidence (unless already at a local maximum). In contrast to belief propagation, VMP can be applied to a very general class of conjugate-exponential models because it uses a factorised variational approximation. Furthermore, by introducing additional variational parameters, VMP can be applied to models containing non-conjugate distributions. The VMP framework also allows the lower bound to be evaluated, and this can be used both for model comparison and for detection of convergence. Variational message passing has been implemented in the form of a general purpose inference engine called VIBES ('Variational Inference for BayEsian networkS') which allows models to be specified graphically and then solved variationally without recourse to coding.
John M. Winn, Christopher M. Bishop
J. Mach. Learn. Res.1
2004 Generative Affine Localisation and Tracking
abstract
We present an extension to the Jojic and Frey (2001) layered sprite model which allows for layers to undergo affine transformations. This extension allows for affine object pose to be inferred whilst simultaneously learn- ing the object shape and appearance. Learning is carried out by applying an augmented variational inference algorithm which includes a global search over a discretised transform space followed by a local optimisa- tion. To aid correct convergence, we use bottom-up cues to restrict the space of possible affine transformations. We present results on a number of video sequences and show how the model can be extended to track an object whose appearance changes throughout the sequence.
John M. Winn, Andrew Blake 0001
NIPS1
2002 VIBES: A Variational Inference Engine for Bayesian Networks
abstract
In recent years variational methods have become a popular tool for approximate inference and learning in a wide variety of proba- bilistic models. For each new application, however, it is currently necessary (cid:12)rst to derive the variational update equations, and then to implement them in application-speci(cid:12)c code. Each of these steps is both time consuming and error prone. In this paper we describe a general purpose inference engine called VIBES (‘Variational Infer- ence for Bayesian Networks’) which allows a wide variety of proba- bilistic models to be implemented and solved variationally without recourse to coding. New models are speci(cid:12)ed either through a simple script or via a graphical interface analogous to a drawing package. VIBES then automatically generates and solves the vari- ational equations. We illustrate the power and (cid:13)exibility of VIBES using examples from Bayesian mixture modelling.
Christopher M. Bishop, David J. Spiegelhalter, John M. Winn
NIPS3
2000 Non-linear Bayesian Image Modelling
Christopher M. Bishop, John M. Winn
ECCV (1)2