Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mark Everingham

dblp:63/4673 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Image recognition and object detection · 44% Face, body and person analysis · 28% Video understanding and tracking · 15%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 100%
Computer graphics and multimedia
2 papers
Image and video processing · 100%

Topics — the 26 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.542015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
Shared parts for deformable part-based models · CVPR 2011
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Computer vision › Image recognition and object detection › object detection › object detection evaluation
object detection benchmark
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Performance modeling and evaluation
benchmarking
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Performance modeling and evaluation › benchmarking › machine learning benchmarking
vision benchmark
0.322015
The Pascal Visual Object Classes Challenge: A Retrospective · Int. J. Comput. Vis. 2015
The Pascal Visual Object Classes (VOC) Challenge · Int. J. Comput. Vis. 2010
Computer vision › Face, body and person analysis
human pose estimation
0.322014
Automatic and Efficient Human Pose Estimation for Sign Language Videos · Int. J. Comput. Vis. 2014
Learning effective human pose estimation from inaccurate annotation · CVPR 2011
Computer vision › Video understanding and tracking
sign language video analysis
0.212014
Automatic and Efficient Human Pose Estimation for Sign Language Videos · Int. J. Comput. Vis. 2014
Computer vision › Face, body and person analysis
face recognition
0.122009
"Who are you?" - Learning person specific classifiers from video · CVPR 2009
Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models · ICCV 2005
Computer vision › Image recognition and object detection › object detection › part-based object detection
deformable part model
0.112011
Shared parts for deformable part-based models · CVPR 2011
Computer vision › Image recognition and object detection
object localization
0.112011
Shared parts for deformable part-based models · CVPR 2011
Image and video processing › motion analysis
human motion analysis
0.112011
Upper Body Detection and Tracking in Extended Signing Sequences · Int. J. Comput. Vis. 2011
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.112009
Implicit color segmentation features for pedestrian and object detection · ICCV 2009
Machine learning › Learning paradigms
multiple instance learning
0.112009
Learning sign language by watching TV (using weakly aligned subtitles) · CVPR 2009
Computer vision › Image recognition and object detection
pedestrian detection
0.112009
Implicit color segmentation features for pedestrian and object detection · ICCV 2009
Computer vision › Video understanding and tracking
sign language recognition
0.112009
Learning sign language by watching TV (using weakly aligned subtitles) · CVPR 2009
Machine learning › Learning paradigms
weakly supervised learning
0.112009
Learning sign language by watching TV (using weakly aligned subtitles) · CVPR 2009
Computer vision › Face, body and person analysis
face detection
0.112005
Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models · ICCV 2005
Computer vision › Face, body and person analysis
head pose estimation
0.112005
Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models · ICCV 2005
Computer vision › Face, body and person analysis › face recognition
video-based face recognition
0.112005
Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models · ICCV 2005
Empirical software engineering › data annotation
annotation agreement
0.012011
Learning effective human pose estimation from inaccurate annotation · CVPR 2011
Empirical software engineering
crowdsourcing
0.012011
Learning effective human pose estimation from inaccurate annotation · CVPR 2011
Image and video processing
image segmentation
0.012002
Evaluating Image Segmentation Algorithms Using the Pareto Front · ECCV (4) 2002
Image and video processing › image segmentation
segmentation evaluation
0.012002
Evaluating Image Segmentation Algorithms Using the Pareto Front · ECCV (4) 2002
Machine learning › Representation and self-supervised learning › representation learning
feature extraction
0.012009
Implicit color segmentation features for pedestrian and object detection · ICCV 2009
Computer vision › 3D vision › 3d shape modeling
3d head modeling
0.012005
Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models · ICCV 2005
Mathematical optimization
multi-objective optimization
0.012002
Evaluating Image Segmentation Algorithms Using the Pareto Front · ECCV (4) 2002
Mathematical optimization › multi-objective optimization
pareto front
0.012002
Evaluating Image Segmentation Algorithms Using the Pareto Front · ECCV (4) 2002

Methods — techniques the papers use, named apart from their topics

part sharing · 0.2nonlinear classifier · 0.2latent annotation update · 0.2energy function · 0.2coupled training · 0.2amazon mechanical turk · 0.2pose estimation · 0.2deep learning · 0.2multiple instance learning · 0.1distance function · 0.1pareto front analysis · 0.1
YearPublicationVenuePosition
2016 Harvesting Training Images for Fine-Grained Object Categories Using Visual Descriptions
Josiah Wang, Katja Markert, Mark Everingham
ECIR3
2015 The Pascal Visual Object Classes Challenge: A Retrospective
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John M. Winn, Andrew Zisserman
Int. J. Comput. Vis.1
2014 Automatic and Efficient Human Pose Estimation for Sign Language Videos
James Charles, Tomas Pfister, Mark Everingham, Andrew Zisserman
Int. J. Comput. Vis.3
2012 Automatic and Efficient Long Term Arm and Hand Tracking for Continuous Sign Language TV Broadcasts
abstract
We present a fully automatic arm and hand tracker that detects joint positions over continuous sign language video sequences of more than an hour in length. Our framework replicates the state-of-the-art long term tracker by Buehler et al. (IJCV 2011), but does not require the manual annotation and, after automatic initialisation, performs tracking in real-time. We cast the problem as a generic frame-by-frame random forest regressor without a strong spatial model. Our contributions are (i) a co-segmentation algorithm that automatically separates the signer from any signed TV broadcast using a generative layered model; (ii) a method of predicting joint positions given only the segmentation and a colour model using a random forest regressor; and (iii) demonstrating that the random forest can be trained from an existing semi-automatic, but computationally expensive, tracker. The method is applied to signing footage with changing background, challenging imaging conditions, and for different signers. We achieve superior joint localisation results to those obtained using the method of Buehler et al.
Tomas Pfister, James Charles, Mark Everingham, Andrew Zisserman
BMVC3
2011 Learning effective human pose estimation from inaccurate annotation
abstract
The task of 2-D articulated human pose estimation in natural images is extremely challenging due to the high level of variation in human appearance. These variations arise from different clothing, anatomy, imaging conditions and the large number of poses it is possible for a human body to take. Recent work has shown state-of-the-art results by partitioning the pose space and using strong nonlinear classifiers such that the pose dependence and multi-modal nature of body part appearance can be captured. We propose to extend these methods to handle much larger quantities of training data, an order of magnitude larger than current datasets, and show how to utilize Amazon Mechanical Turk and a latent annotation update scheme to achieve high quality annotations at low cost. We demonstrate a significant increase in pose estimation accuracy, while simultaneously reducing computational expense by a factor of 10, and contribute a dataset of 10,000 highly articulated poses.
Sam Johnson, Mark Everingham
CVPR2
2011 Shared parts for deformable part-based models
abstract
The deformable part-based model (DPM) proposed by Felzenszwalb et al. has demonstrated state-of-the-art results in object localization. The model offers a high degree of learnt invariance by utilizing viewpoint-dependent mixture components and movable parts in each mixture component. One might hope to increase the accuracy of the DPM by increasing the number of mixture components and parts to give a more faithful model, but limited training data prevents this from being effective. We propose an extension to the DPM which allows for sharing of object part models among multiple mixture components as well as object classes. This results in more compact models and allows training examples to be shared by multiple components, ameliorating the effect of a limited size training set. We (i) reformulate the DPM to incorporate part sharing, and (ii) propose a novel energy function allowing for coupled training of mixture components and object classes. We report state-of-the-art results on the PASCAL VOC dataset.
Patrick Ott, Mark Everingham
CVPR2
2011 Upper Body Detection and Tracking in Extended Signing Sequences
Patrick Buehler, Mark Everingham, Daniel P. Huttenlocher, Andrew Zisserman
Int. J. Comput. Vis.2
2010 Clustered Pose and Nonlinear Appearance Models for Human Pose Estimation
abstract
We investigate the task of 2D articulated human pose estimation in unconstrained still images. This is extremely challenging because of variation in pose, anatomy, clothing, and imaging conditions. Current methods use simple models of body part appearance and plausible configurations due to limitations of available training data and constraints on computational expense. We show that such models severely limit accuracy. Building on the successful pictorial structure model (PSM) we propose richer models of both appearance and pose, using state-of-the-art discriminative classifiers without introducing unacceptable computational expense. We introduce a new annotated database of challenging consumer images, an order of magnitude larger than currently available datasets, and demonstrate over 50 % relative improvement in pose estimation accuracy over a stateof-the-art method.
Sam Johnson, Mark Everingham
BMVC2
2010 The Pascal Visual Object Classes (VOC) Challenge
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, Andrew Zisserman
Int. J. Comput. Vis.1
2009 Learning Models for Object Recognition from Natural Language Descriptions
abstract
We investigate the task of learning models for visual object recognition from natural language descriptions alone. The approach contributes to the recognition of fine-grain object categories, such as animal and plant species, where it may be difficult to collect many images for training, but where textual descriptions of visual attributes are readily available. As an example we tackle recognition of butterfly species, learning models from descriptions in an online nature guide. We propose natural language processing methods for extracting salient visual attributes from these descriptions to use as ‘templates ’ for the object categories, and apply vision methods to extract corresponding attributes from test images. A generative model is used to connect textual terms in the learnt templates to visual attributes. We report experiments comparing the performance of humans and the proposed method on a dataset of ten butterfly categories. 1
Josiah Wang, Katja Markert, Mark Everingham
BMVC3
2009 Learning sign language by watching TV (using weakly aligned subtitles)
abstract
The goal of this work is to automatically learn a large number of British sign language (BSL) signs from TV broadcasts. We achieve this by using the supervisory information available from subtitles broadcast simultaneously with the signing. This supervision is both weak and noisy: it is weak due to the correspondence problem since temporal distance between sign and subtitle is unknown and signing does not follow the text order; it is noisy because subtitles can be signed in different ways, and because the occurrence of a subtitle word does not imply the presence of the corresponding sign. The contributions are: (i) we propose a distance function to match signing sequences which includes the trajectory of both hands, the hand shape and orientation, and properly models the case of hands touching; (ii) we show that by optimizing a scoring function based on multiple instance learning, we are able to extract the sign of interest from hours of signing footage, despite the very weak and noisy supervision. The method is automatic given the English target word of the sign to be learnt. Results are presented for 210 words including nouns, verbs and adjectives.
Patrick Buehler, Andrew Zisserman, Mark Everingham
CVPR3
2009 "Who are you?" - Learning person specific classifiers from video
abstract
We investigate the problem of automatically labelling faces of characters in TV or movie material with their names, using only weak supervision from automatically-aligned subtitle and script text. Our previous work (Everingham et al. [8]) demonstrated promising results on the task, but the coverage of the method (proportion of video labelled) and generalization was limited by a restriction to frontal faces and nearest neighbour classification. In this paper we build on that method, extending the coverage greatly by the detection and recognition of characters in profile views. In addition, we make the following contributions: (i) seamless tracking, integration and recognition of profile and frontal detections, and (ii) a character specific multiple kernel classifier which is able to learn the features best able to discriminate between the characters. We report results on seven episodes of the TV series "Buffy the Vampire Slayer", demonstrating significantly increased coverage and performance with respect to previous methods on this material.
Josef Sivic, Mark Everingham, Andrew Zisserman
CVPR2
2009 Implicit color segmentation features for pedestrian and object detection
abstract
We investigate the problem of pedestrian detection in still images. Sliding window classifiers, notably using the Histogram-of-Gradient (HOG) features proposed by Dalal and Triggs are the state-of-the-art for this task, and we base our method on this approach. We propose a novel feature extraction scheme which computes implicit `soft segmentations' of image regions into foreground/background. The method yields stronger object/background edges than gray-scale gradient alone, suppresses textural and shading variations, and captures local coherence of object appearance. The main contributions of our work are: (i) incorporation of segmentation cues into object detection; (ii) integration with classifier learning cf. a post-processing filter; (iii) high computational efficiency. We report results on the INRIA person detection dataset, achieving state-of-the-art results considerably exceeding those of the original HOG detector. Preliminary results for generic object detection on the PASCAL VOC2006 dataset also show substantial improvements in accuracy.
Patrick Ott, Mark Everingham
ICCV2
2009 Taking the bite out of automated naming of characters in TV video
Mark Everingham, Josef Sivic, Andrew Zisserman
Image Vis. Comput.1
2008 Long Term Arm and Hand Tracking for Continuous Sign Language TV Broadcasts
abstract
The goal of this work is to detect hand and arm positions over continuous sign language video sequences of more than one hour in length. We cast the problem as inference in a generative model of the image. Un-der this model, limb detection is expensive due to the very large number of possible configurations each part can assume. We make the following con-tributions to reduce this cost: (i) using efficient sampling from a pictorial structure proposal distribution to obtain reasonable configurations; (ii) iden-tifying a large set of frames where correct configurations can be inferred, and using temporal tracking elsewhere. Results are reported for signing footage with changing background, chal-lenging image conditions, and different signers; and we show that the method is able to identify the true arm and hand locations. The results exceed the state-of-the-art for the length and stability of continuous limb tracking. 1
Patrick Buehler, Mark Everingham, Daniel P. Huttenlocher, Andrew Zisserman
BMVC2
2006 Hello! My name is... Buffy'' -- Automatic Naming of Characters in TV Video
abstract
We investigate the problem of automatically labelling appearances of characters in TV or film material. This is tremendously challenging due to the huge variation in imaged appearance of each character and the weakness and ambiguity of available annotation. However, we demonstrate that high precision can be achieved by combining multiple sources of information, both visual and textual. The principal novelties that we introduce are: (i) automatic generation of time stamped character annotation by aligning subtitles and transcripts; (ii) strengthening the supervisory information by identifying when characters are speaking; (iii) using complementary cues of face matching and clothing matching to propose common annotations for face tracks. Results are presented on episodes of the TV series “Buffy the Vampire Slayer”. 1
Mark Everingham, Josef Sivic, Andrew Zisserman
BMVC1
2005 Identifying Individuals in Video by Combining "Generative" and Discriminative Head Models
abstract
The objective of this work is automatic detection and identification of individuals in unconstrained consumer video, given a minimal number of labelled faces as training data. Whilst much work has been done on (mainly frontal) face detection and recognition, current methods are not sufficiently robust to deal with the wide variations in pose and appearance found in such video. These include variations in scale, illumination, expression, partial occlusion, motion blur, etc. We describe two areas of innovation: the first is to capture the 3-D appearance of the entire head, rather than just the face region, so that visual features such as the hairline can be exploited. The second is to combine discriminative and 'generative' approaches for detection and recognition. Images rendered using the head model are used to train a discriminative tree-structured classifier giving efficient detection and pose estimates over a very wide pose range with three degrees of freedom. Subsequent verification of the identity is obtained using the head model in a 'generative' framework. We demonstrate excellent performance in detecting and identifying three characters and their poses in a TV situation comedy
Mark Everingham, Andrew Zisserman
ICCV1
2003 Wearable Mobility Aid for Low Vision Using Scene Classification in a Markov Random Field Model Framework
abstract
This article describes work on a novel approach to vision enhancement for people with severe visual impairments. This approach utilizes computer vision techniques to classify scene content so that visual enhancement of the scene can identify semantically important concepts. The mediated view of a scene presented to the user is in the form of a highly-saturated color image in which distinct colors represent important object types in the scene. The effectiveness of this scheme was demonstrated in a pilot study participated in by people with a range of visual impairments. The scene classification technique uses an artificial neural network classifier within the framework of a Markov random field model, and the accuracy and robustness of this technique using low quality video images from a hand-held camera is demonstrated.
Mark Everingham, Barry T. Thomas, Tom Troscianko
Int. J. Hum. Comput. Interact.1
2002 Evaluating Image Segmentation Algorithms Using the Pareto Front
Mark Everingham, Henk L. Muller, Barry T. Thomas
ECCV (4)1
2001 Evaluating image segmentation algorithms using monotonic hulls in fitness/cost space
abstract
Image segmentation is the first stage of processing in many practical computer vision systems. While development of particular segmentation algorithms has attracted considerable research interest, relatively little work has been published on the subject of their evaluation. In this paper we propose a framework for quantitative evaluation of segmentation algorithms which we believe addresses shortcomings of previous approaches, and use this framework to compare several state-of-the-art algorithms.
Mark Everingham, Henk L. Muller, Barry T. Thomas
BMVC1
2001 Supervised segmentation and tracking of nonrigid objects using a "mixture of histograms" model
abstract
Segmentation and tracking of objects in video sequences is important for a number of applications. In the supervised variant, segmentation can be achieved by modelling the probability density of image observations taken from an object for use in a Bayesian classifier, and Gaussian mixture models have been applied to this task by several researchers. Motivated by practical difficulties we have experienced with these models we propose a novel and simple alternative approach which combines a strong shape model with histograms of image features and gives good empirical results on test sequences requiring flexible models.
Mark Everingham, Barry T. Thomas
ICIP (1)1