Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Paul Sturgess

dblp:37/8610 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-authorSystems, architecture and hardware · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Segmentation and scene understanding · 36% 3D vision · 20% Kernel, tree and ensemble methods · 16%
Computer graphics and multimedia
1 paper
Multimedia analysis and retrieval · 100%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
semantic segmentation
0.632015
Semantic octree: Unifying recognition, reconstruction and representation via an octree constrained higher order MRF · ICRA 2015
Dense Semantic Image Segmentation with Objects and Attributes · CVPR 2014
Joint Optimization for Object Class Segmentation and Dense Stereo Reconstruction · Int. J. Comput. Vis. 2012
Computer vision › 3D vision
3d reconstruction
0.212015
Semantic octree: Unifying recognition, reconstruction and representation via an octree constrained higher order MRF · ICRA 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation
0.212014
Dense Semantic Image Segmentation with Objects and Attributes · CVPR 2014
Multimedia analysis and retrieval › image analysis › image understanding
image parsing
0.212014
ImageSpirit: Verbal Guided Image Parsing · ACM Trans. Graph. 2014
Human-AI interaction › large language model interaction › language-based interaction
verbal interaction
0.212014
ImageSpirit: Verbal Guided Image Parsing · ACM Trans. Graph. 2014
Machine learning › Learning theory › classification
classifier design
0.112012
Efficient discriminative learning of parametric nearest neighbor classifiers · CVPR 2012
Computer vision › 3D vision › stereo vision
dense stereo reconstruction
0.112012
Joint Optimization for Object Class Segmentation and Dense Stereo Reconstruction · Int. J. Comput. Vis. 2012
Machine learning › Kernel, tree and ensemble methods
locally linear classifier
0.112012
Efficient discriminative learning of parametric nearest neighbor classifiers · CVPR 2012
Machine learning › Kernel, tree and ensemble methods › nearest neighbor methods
nearest neighbor classifier
0.112012
Efficient discriminative learning of parametric nearest neighbor classifiers · CVPR 2012
Computer vision › Image recognition and object detection
object detection
0.112010
What, Where and How Many? Combining Object Detectors and CRFs · ECCV (4) 2010
Computer vision › Segmentation and scene understanding
scene understanding
0.112010
What, Where and How Many? Combining Object Detectors and CRFs · ECCV (4) 2010
Machine learning › Optimization for machine learning › optimization
joint optimization
0.012012
Joint Optimization for Object Class Segmentation and Dense Stereo Reconstruction · Int. J. Comput. Vis. 2012
Machine learning › Optimization for machine learning › regularized risk minimization
max-margin learning
0.012012
Efficient discriminative learning of parametric nearest neighbor classifiers · CVPR 2012

Methods — techniques the papers use, named apart from their topics

joint estimation · 0.4interactive refinement · 0.4octree representation · 0.2markov random field · 0.2mean-field approximation · 0.2hierarchical model · 0.2boosting-based piecewise training · 0.2weighted squared euclidean distance · 0.1stereo matching · 0.1max-margin learning · 0.1joint optimization · 0.1ensemble learning · 0.1
YearPublicationVenuePosition
2015 Semantic octree: Unifying recognition, reconstruction and representation via an octree constrained higher order MRF
abstract
On the one hand, mainly within the computer vision community, multi-resolution image labelling problems with pixel, super-pixel and object levels, have made great progress towards the modelling of holistic scene understanding. On the other hand, mainly within the robotics and graphics communities, multi-resolution 3Drepresentations of the world have matured to be efficient and accurate. In this paper we bring together the two hands and move towards the new direction of unified recognition, reconstruction and representation. We tackle the problem by embedding an octree into a hierarchical robust PNMarkov Random Field. This allows us to jointly infer the multi-resolution 3Dvolume along with the object-class labels, all within the constraints of an octree data-structure. The octree representation is chosen as this data-structure is efficient for further processing such as dynamic updates, data compression, and surface reconstruction. We perform experiments in inferring our semantic octree on the The kitti Vision Benchmark Suite in order to demonstrate its efficacy.
Sunando Sengupta, Paul Sturgess
ICRA2
2014 Dense Semantic Image Segmentation with Objects and Attributes
abstract
The concepts of objects and attributes are both important for describing images precisely, since verbal descriptions often contain both adjectives and nouns (e.g. 'I see a shiny red chair'). In this paper, we formulate the problem of joint visual attribute and object class image segmentation as a dense multi-labelling problem, where each pixel in an image can be associated with both an object-class and a set of visual attributes labels. In order to learn the label correlations, we adopt a boosting-based piecewise training approach with respect to the visual appearance and co-occurrence cues. We use a filtering-based mean-field approximation approach for efficient joint inference. Further, we develop a hierarchical model to incorporate region-level object and attribute information. Experiments on the aPASCAL, CORE and attribute augmented NYU indoor scenes datasets show that the proposed approach is able to achieve state-of-the-art results.
Shuai Zheng 0001, Ming-Ming Cheng, Jonathan Warrell, Paul Sturgess, Vibhav Vineet, Carsten Rother, Philip Torr 0001
CVPR4
2014 ImageSpirit: Verbal Guided Image Parsing
abstract
Humans describe images in terms of nouns and adjectives while algorithms operate on images represented as sets of pixels. Bridging this gap between how humans would like to access images versus their typical representation is the goal of image parsing, which involves assigning object and attribute labels to pixels. In this article we propose treating nouns as object labels and adjectives as visual attribute labels. This allows us to formulate the image parsing problem as one of jointly estimating per-pixel object and attribute labels from a set of training images. We propose an efficient (interactive time) solution. Using the extracted labels as handles, our system empowers a user to verbally refine the results. This enables hands-free parsing of an image into pixel-wise object/attribute labels that correspond to human semantics. Verbally selecting objects of interest enables a novel and natural interaction modality that can possibly be used to interact with new generation devices (e.g., smartphones, Google Glass, livingroom devices). We demonstrate our system on a large number of real-world images with varying complexity. To help understand the trade-offs compared to traditional mouse-based interactions, results are reported for both a large-scale quantitative evaluation and a user study.
Ming-Ming Cheng, Shuai Zheng 0001, Wen-Yan Lin, Vibhav Vineet, Paul Sturgess, Nigel T. Crook, Niloy J. Mitra, Philip Torr 0001
ACM Trans. Graph.5
2012 Scalable Cascade Inference for Semantic Image Segmentation
abstract
Semantic image segmentation is a problem of simultaneous segmentation and recognition of an input image into regions and their associated categorical labels, such as person, car or cow. A popular way to achieve this goal is to assign a label to every pixel in the input image and impose simple structural constraints on the output label space. Efficient approximation algorithms for solving this labelling problem such as α-expansion have, at best, linear runtime complexity with respect to the number of labels, making them practical only when working in a specific domain that has few classes-of-interest. However when working in a more general setting where the number of classes could easily reach tens of thousands, sub-linear complexity is desired. In this paper we propose meeting this requirement by performing cascaded inference that wraps around the α-expansion algorithm. The cascade both divides the large label set into smaller more manageable ones by way of a hierarchy, and dynamically subdivides the image into smaller and smaller regions during inference. We test our method on the SUN09 dataset with 107 accurately hand labelled classes.
Paul Sturgess, Lubor Ladicky, Nigel T. Crook, Philip Torr 0001
BMVC1
2012 Improved Initialization and Gaussian Mixture Pairwise Terms for Dense Random Fields with Mean-field Inference
abstract
Recently, Krahenbuhl and Koltun proposed an efficient inference method for densely connected pairwise random fields using the mean-field approximation for a Conditional Random Field (CRF). However, they restrict their pairwise weights to take the form of a weighted combination of Gaussian kernels where each Gaussian component is allowed to take only zero mean, and can only be rescaled by a single value for each label pair. Further, their method is sensitive to initialization. In this paper, we propose methods to alleviate these issues. First, we propose a hierarchical mean-field approach where labelling from the coarser level is propagated to the finer level for better initialisation. Further, we use SIFT-flow based label transfer to provide a good initial condition at the coarsest level. Second, we allow our approach to take general Gaussian pairwise weights, where we learn the mean, the co-variance matrix, and the mixing co-efficient for every mixture component. We propose a variation of Expectation Maximization (EM) for piecewise learning of the parameters of the mixture model determined by the maximum likelihood function. Finally, we demonstrate the efficiency and accuracy offered by our method for object class segmentation problems on two challenging datasets: PascalVOC-10 segmentation and CamVid datasets. We show that we are able to achieve state of the art performance on the CamVid dataset, and an almost 3% improvement on the PascalVOC10 dataset compared to baseline graph-cut and mean-field methods, while also reducing the inference time by almost a factor of 3 compared to graph-cuts based methods.
Vibhav Vineet, Jonathan Warrell, Paul Sturgess, Philip Torr 0001
BMVC3
2012 Efficient discriminative learning of parametric nearest neighbor classifiers
abstract
Linear SVMs are efficient in both training and testing, however the data in real applications is rarely linearly separable. Non-linear kernel SVMs are too computationally intensive for applications with large-scale data sets. Recently locally linear classifiers have gained popularity due to their efficiency whilst remaining competitive with kernel methods. The vanilla nearest neighbor algorithm is one of the simplest locally linear classifiers, but it lacks robustness due to the noise often present in real-world data. In this paper, we introduce a novel local classifier, Parametric Nearest Neighbor (P-NN) and its extension Ensemble of P-NN (EP-NN). We parameterize the nearest neighbor algorithm based on the minimum weighted squared Euclidean distances between the data points and the prototypes, where a prototype is represented by a locally linear combination of some data points. Meanwhile, our method attempts to jointly learn both the prototypes and the classifier parameters discriminatively via max-margin. This makes our classifiers suitable to approximate the classification decision boundaries locally based on nonlinear functions. During testing, the computational complexity of both classifiers is linear in the product of the dimension of data and the number of prototypes. Our classification results on MNIST, USPS, LETTER, and Chars 74K are comparable and in some cases are better than many other methods such as the state-of-the-art locally linear classifiers.
Paul Sturgess, Sunando Sengupta, Nigel T. Crook, Philip Torr 0001
CVPR2
2012 Automatic dense visual semantic mapping from street-level imagery
abstract
This paper describes a method for producing a semantic map from multi-view street-level imagery. We define a semantic map as an overhead, or bird's eye view of a region with associated semantic object labels, such as car, road and pavement. We formulate the problem using two conditional random fields. The first is used to model the semantic image segmentation of the street view imagery treating each image independently. The outputs of this stage are then aggregated over many images to form the input for our semantic map that is a second random field defined over a ground plane. Each image is related by a simple, yet effective, geometrical function that back projects a region from the street view image into the overhead ground plane map. We introduce, and make publicly available, a new dataset created from real world data. Our qualitative evaluation is performed on this data consisting of a 14.8 km track, and we also quantify our results on a representative subset.
Sunando Sengupta, Paul Sturgess, Lubor Ladicky, Philip Torr 0001
IROS2
2012 Joint Optimization for Object Class Segmentation and Dense Stereo Reconstruction
Lubor Ladicky, Paul Sturgess, Chris Russell 0001, Sunando Sengupta, Yalin Bastanlar, W. F. Clocksin, Philip Torr 0001
Int. J. Comput. Vis.2
2010 Joint Optimisation for Object Class Segmentation and Dense Stereo Reconstruction
abstract
This work is supported by EPSRC research grants, HMGCC, TUBITAK researcher exchange grant, the IST Programme of the European Community, under the PASCAL2 Network of Excellence, IST-2007-216886.
Lubor Ladicky, Paul Sturgess, Chris Russell 0001, Sunando Sengupta, Yalin Bastanlar, W. F. Clocksin, Philip Torr 0001
BMVC2
2010 What, Where and How Many? Combining Object Detectors and CRFs
Lubor Ladicky, Paul Sturgess, Karteek Alahari, Chris Russell 0001, Philip Torr 0001
ECCV (4)2
2009 Combining Appearance and Structure from Motion Features for Road Scene Understanding
abstract
In this paper we present a framework for pixel-wise object segmentation of road scenes that combines motion and appearance features. It is designed to handle street-level imagery such as that on Google Street View and Microsoft Bing Maps. We formulate the problem in a CRF framework in order to probabilistically model the label likelihoods and the a priori knowledge. An extended set of appearance-based features is used, which consists of textons, colour, location and HOG descriptors. A novel boosting approach is then applied to combine the motion and appearance-based features. We also incorporate higher order potentials in our CRF model, which produce segmentations with precise object boundaries. We evaluate our method both quantitatively and qualitatively on the challenging Cambridge-driving Labeled Video dataset. Our approach shows an overall recognition accuracy of 84 % compared to the state-of-the-art accuracy of 69%.
Paul Sturgess, Karteek Alahari, Lubor Ladicky, Philip Torr 0001
BMVC1