James H. Elder

dblp:21/2977 · DBLP profile ↗
← Back
38ranked-venue papers
12as first author
5since 2021 · last 2024
0000-0003-3880-4808ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 11 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 4 since 2021Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
15 papers
Computational photography and imaging · 33% Image and video processing · 30% Geometric modeling and processing · 14%
Artificial intelligence
9 papers
Face, body and person analysis · 42% Generative modeling · 17% Image recognition and object detection · 11%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational photography and imaging
camera calibration
0.822022
A Reliable Online Method for Joint Estimation of Focal Length and Camera Rotation · ECCV (1) 2022
Automatic Single-View Calibration and Rectification from Parallel Planar Curves · ECCV (4) 2014
Computer vision › Face, body and person analysis
face recognition
0.442012
Probabilistic Models for Inference about Identity · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Probabilistic Linear Discriminant Analysis for Inferences About Identity · ICCV 2007
Image and video processing › pattern detection › curve detection
line segment detection
0.312017
MCMLSD: A Dynamic Programming Approach to Line Segment Detection · CVPR 2017
Image and video processing › image warping
image rectification
0.212014
Automatic Single-View Calibration and Rectification from Parallel Planar Curves · ECCV (4) 2014
Virtual and augmented reality › tracking
camera tracking
0.212022
A Reliable Online Method for Joint Estimation of Focal Length and Camera Rotation · ECCV (1) 2022
Virtual and augmented reality › tracking
camera pose estimation
0.112012
3DTown: The automatic urban awareness project · VR 2012
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection
0.122007
Pre-Attentive and Attentive Detection of Humans in Wide-Field Scenes · Int. J. Comput. Vis. 2007
Statistical Cue Integration for Foveated Wide-Field Surveillance · CVPR (2) 2005
Machine learning › Generative modeling › face synthesis
generative face model
0.122012
Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Probabilistic Models for Inference about Identity · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Geometric modeling and processing
shape deformation
0.112010
On growth and formlets: Sparse multi-scale coding of planar shape · CVPR 2010
Geometric modeling and processing
shape representation
0.112010
On growth and formlets: Sparse multi-scale coding of planar shape · CVPR 2010
Image and video processing
edge detection
0.162001
Image Editing in the Contour Domain · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Local Scale Control for Edge Detection and Blur Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Local Scale Control for Edge Detection and Blur Estimation · ECCV (2) 1996
Algorithms and data structures
dynamic programming
0.112017
MCMLSD: A Dynamic Programming Approach to Line Segment Detection · CVPR 2017
Computer vision › 3D vision › multi-view geometry › camera geometry
manhattan frame estimation
0.112008
Efficient Edge-Based Methods for Estimating Manhattan Frames in Urban Imagery · ECCV (2) 2008
Computer vision › Face, body and person analysis › face recognition › robust face recognition
pose-invariant face recognition
0.112008
Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Machine learning › Generative modeling
generative model
0.112007
Probabilistic Linear Discriminant Analysis for Inferences About Identity · ICCV 2007
Natural language and speech › Speech recognition and synthesis › speaker recognition
probabilistic linear discriminant analysis
0.112007
Probabilistic Linear Discriminant Analysis for Inferences About Identity · ICCV 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112005
Statistical Cue Integration for Foveated Wide-Field Surveillance · CVPR (2) 2005
Computer vision › Video understanding and tracking › activity recognition
human activity localization
0.112005
Statistical Cue Integration for Foveated Wide-Field Surveillance · CVPR (2) 2005
Computer vision › Face, body and person analysis › face recognition › robust face recognition
illumination-invariant face recognition
0.112005
Creating Invariance to "Nuisance Parameters" in Face Recognition · CVPR (2) 2005
Biometric security
face recognition
0.112005
Creating Invariance to "Nuisance Parameters" in Face Recognition · CVPR (2) 2005
Visual content generation and editing
image editing
0.122001
Image Editing in the Contour Domain · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Image Editing in the Contour Domain · CVPR 1998
Internet of things and sensor networks › camera sensor networks
camera networks
0.012012
3DTown: The automatic urban awareness project · VR 2012
Image and video processing › image segmentation
contour grouping
0.012003
Contour Grouping with Prior Models · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Image and video processing
image segmentation
0.012003
Contour Grouping with Prior Models · IEEE Trans. Pattern Anal. Mach. Intell. 2003
Computational photography and imaging
omnidirectional imaging
0.012002
Image Registration for Foveated Omnidirectional Sensing · ECCV (4) 2002
Geometric modeling and processing
shape analysis
0.012010
On growth and formlets: Sparse multi-scale coding of planar shape · CVPR 2010
Computer vision › Segmentation and scene understanding › perceptual grouping
contour grouping
0.012001
Contour Grouping with Strong Prior Models · CVPR (2) 2001
Multimedia analysis and retrieval › image analysis › image blur analysis
blur estimation
0.031998
Local Scale Control for Edge Detection and Blur Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 1998
Local Scale Control for Edge Detection and Blur Estimation · ECCV (2) 1996
Space Scale Localization, Blur, and Contour-Based Image Coding · CVPR 1996
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
factor analysis
0.012008
Tied Factor Analysis for Face Recognition across Large Pose Differences · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Computer vision › Segmentation and scene understanding
edge detection
0.011999
Are Edges Incomplete? · Int. J. Comput. Vis. 1999

Methods — techniques the papers use, named apart from their topics

probabilistic hough transform · 0.6online estimation · 0.6markov chain · 0.6joint estimation · 0.6dynamic programming · 0.6manhattan structure estimation · 0.3PTZ camera tracking · 0.3tied factor analysis · 0.2edge detection · 0.2non-linear manifold modeling · 0.1generative modeling · 0.1matching pursuit · 0.1greedy coarse-to-fine pursuit · 0.1probabilistic distance metric · 0.1EM algorithm · 0.1tied model · 0.1probabilistic linear discriminant analysis · 0.1pre-attentive and attentive detection · 0.1
YearPublicationVenuePosition
2024 Class-conditional domain adaptation for semantic segmentation
abstract
Semantic segmentation is an important sub-task for many applications. However, pixel-level ground-truth labeling is costly, and there is a tendency to overfit to training data, thereby limiting the generalization ability. Unsupervised domain adaptation can potentially address these problems by allowing systems trained on labelled datasets from the source domain (including less expensive synthetic domain) to be adapted to a novel target domain. The conventional approach involves automatic extraction and alignment of the representations of source and target domains globally. One limitation of this approach is that it tends to neglect the differences between classes: representations of certain classes can be more easily extracted and aligned between the source and target domains than others, limiting the adaptation over all classes. Here, we address this problem by introducing a Class-Conditional Domain Adaptation (CCDA) method. This incorporates a class-conditional multi-scale discriminator and class-conditional losses for both segmentation and adaptation. Together, they measure the segmentation, shift the domain in a class-conditional manner, and equalize the loss over classes. Experimental results demonstrate that the performance of our CCDA method matches, and in some cases, surpasses that of state-of-the-art methods.
Yue Wang 0038, James H. Elder, Runmin Wu, Huchuan Lu
Comput. Vis. Media3
2023 A uniform transformer-based structure for feature fusion and enhancement for RGB-D saliency detection
Yue Wang 0038, Xu Jia 0012, Lu Zhang 0053, James H. Elder, Huchuan Lu
Pattern Recognit.5
2022 Blind Image Super-Resolution with Degradation-Aware Adaptation
Yue Wang 0038, Jiawen Ming, Xu Jia 0012, James H. Elder, Huchuan Lu
ACCV (3)4
2022 A Reliable Online Method for Joint Estimation of Focal Length and Camera Rotation
Yiming Qian, James H. Elder
ECCV (1)2
2022 VCSeg: Virtual Camera Adaptation for Road Segmentation
abstract
Domain shift limits generalization in many problem domains. For road segmentation, one of the principal causes of domain shift is variation in the geometric camera parameters, which results in misregistration of scene structure between images. To address this issue, we decompose the shift into two components: Between-camera shift and within-camera shift. To handle between-camera shift, we assume that average camera parameters are known or can be estimated and use this knowledge to rectify both source and target domain images to a standard virtual camera model. To handle within-camera shift, we use estimates of road vanishing points to correct for shifts in camera pan and tilt. While this approach improves alignment, it produces gaps in the virtual image that complicates network training. To solve this problem, we introduce a novel projective image completion method that fills these gaps in a plausible way. Using five diverse and challenging road segmentation datasets, we demonstrate that our virtual camera method dramatically improves road segmentation performance when generalizing across cameras, and propose that this be integrated as a standard component of road segmentation systems to improve generalization.
Gong Cheng 0007, James H. Elder
WACV2
2020 Synergistic Saliency and Depth Prediction for RGB-D Saliency Detection
Yue Wang 0038, James H. Elder, Runmin Wu, Huchuan Lu, Lu Zhang 0053
ACCV (2)3
2019 Keep Your Eye on the Puck: Automatic Hockey Videography
abstract
While hockey involves a large playing surface, instantaneous play is typically localized to a smaller region of the ice. Live spectators thus attentively shift their gaze to follow play, and professional sports videographers pan and tilt their cameras to mimic this process. Unfortunately, manual videography is economically prohibitive below the elite level. Here we propose a system for automatically tracking play, allowing a high-definition video feed to be dynamically cropped and retargeted to a spectator's display device. We employ the puck as an objective surrogate for the location of play, and develop a novel method for ground-truthing puck location from high-definition video. This allows us to train a deep network regressor that uses the video imagery, optic flow, estimated player positions and team affiliation to predict the location of play. We show that our algorithm outperforms a simple 'follow the herd' strategy and results in a practical system for delivering high-quality curated video of amateur-level hockey games to remote spectators.
Hemanth Pidaparthy, James H. Elder
WACV2
2018 LS3D: Single-View Gestalt 3D Surface Reconstruction from Manhattan Line Segments
Yiming Qian, Srikumar Ramalingam, James H. Elder
ACCV (4)3
2017 MCMLSD: A Dynamic Programming Approach to Line Segment Detection
abstract
Prior approaches to line segment detection typically involve perceptual grouping in the image domain or global accumulation in the Hough domain. Here we propose a probabilistic algorithm that merges the advantages of both approaches. In a first stage lines are detected using a global probabilistic Hough approach. In the second stage each detected line is analyzed in the image domain to localize the line segments that generated the peak in the Hough map. By limiting search to a line, the distribution of segments over the sequence of points on the line can be modeled as a Markov chain, and a probabilistically optimal labelling can be computed exactly using a standard dynamic programming algorithm, in linear time. The Markov assumption also leads to an intuitive ranking method that uses the local marginal posterior probabilities to estimate the expected number of correctly labelled points on a segment. To assess the resulting Markov Chain Marginal Line Segment Detector (MCMLSD) we develop and apply a novel quantitative evaluation methodology that controls for under-and over-segmentation. Evaluation on the YorkUrbanDB dataset shows that the proposed MCMLSD method outperforms the state-of-the-art by a substantial margin.
Emilio J. Almazán, Ron Tal, Yiming Qian, James H. Elder
CVPR4
2016 Unsupervised Crowd Counting
Nada Elassal, James H. Elder
ACCV (5)2
2016 Evaluating features and classifiers for road weather condition analysis
abstract
Weather-dependent road conditions are a major factor in many automobile incidents; computer vision algorithms for automatic classification of road conditions can thus be of great benefit. This paper presents a system for classification of road conditions using still-frames taken from an uncalibrated dashboard camera. The problem is challenging due to variability in camera placement, road layout, weather and illumination conditions. The system uses a prior distribution of road pixel locations learned from training data then fuses normalized luminance and texture features probabilistically to categorize the segmented road surface. We attain an accuracy of 80% for binary classification (bare vs. snow/ice-covered) and 68% for 3 classes (dry vs. wet vs. snow/ice-covered) on a challenging dataset, suggesting that a useful system may be viable.
Yiming Qian, Emilio J. Almazán, James H. Elder
ICIP3
2014 Automatic Single-View Calibration and Rectification from Parallel Planar Curves
Eduardo R. Corral-Soto, James H. Elder
ECCV (4)2
2014 Effects of Specular Highlights on Perceived Surface Convexity
abstract
Shading is known to produce vivid perceptions of depth. However, the influence of specular highlights on perceived shape is unclear: some studies have shown that highlights improve quantitative shape perception while others have shown no effect. Here we ask how specular highlights combine with Lambertian shading cues to determine perceived surface curvature, and to what degree this is based upon a coherent model of the scene geometry. Observers viewed ambiguous convex/concave shaded surfaces, with or without highlights. We show that the presence/absence of specular highlights has an effect on qualitative shape, their presence biasing perception toward convex interpretations of ambiguous shaded objects. We also find that the alignment of a highlight with the Lambertian shading modulates its effect on perceived shape; misaligned highlights are less likely to be perceived as specularities, and thus have less effect on shape perception. Increasing the depth of the surface or the slant of the illuminant also modulated the effect of the highlight, increasing the bias toward convexity. The effect of highlights on perceived shape can be understood probabilistically in terms of scene geometry: for deeper objects and/or highly slanted illuminants, highlights will occur on convex but not concave surfaces, due to occlusion of the illuminant. Given uncertainty about the exact object depth and illuminant direction, the presence of a highlight increases the probability that the surface is convex.
Wendy J. Adams, James H. Elder
PLoS Comput. Biol.2
2013 Combining Local and Global Cues for Closed Contour Extraction
abstract
Algorithms for computing closed contours are generally based upon local Gestalt cues relating pairs of oriented elements, and a Markov assumption to then group these elements into chains. Without additional global constraints, these algorithms generally do not perform well on general natural scenes. Such global cues could include symmetry, shape priors or global colour appearance. A key challenge is to combine these local and global cues in a statistically optimal way. Here we propose a novel, effective method for rigorously combining local and global cues, both at the stage of forming new closed contour hypotheses, and at the stage of evaluating and ranking these hypotheses. We also demonstrate the importance of promoting the diversity of hypotheses. We evaluate our results on a standard public dataset, and demonstrate a substantial performance improvement over prior methods.
Vida Movahedi, James H. Elder
BMVC2
2013 On growth and formlets: Sparse multi-scale coding of planar shape
James H. Elder, Timothy D. Oleskiw, Alex Yakubovich, Gabriel Peyré
Image Vis. Comput.1
2012 3DTown: The automatic urban awareness project
abstract
In this work the goal is to develop a distributed system for sensing, interpreting and visualizing the real-time dynamics of urban life within the 3D context of a city focusing on typical, useful dynamic information such as walking pedestrians and moving vehicles captured by pan-tilt-zoom (PTZ) video cameras. Three-dimensionalization of the data extracted from video cameras is achieved by an algorithm that uses the Manhattan structure of the urban scene to automatically estimate the camera pose. Thus, if the pose of the video camera changes, our system will automatically update the corresponding projection matrix to maintain accurate geo-location of the scene dynamics.
Eduardo R. Corral-Soto, Ron Tal, Larry Wang, Ravi Ancil Persad, Chan Solomon, Bob Hou, Gunho Sohn, James H. Elder
VR9
2012 Probabilistic Models for Inference about Identity
abstract
Many face recognition algorithms use "distance-based" methods: Feature vectors are extracted from each face and distances in feature space are compared to determine matches. In this paper, we argue for a fundamentally different approach. We consider each image as having been generated from several underlying causes, some of which are due to identity (latent identity variables, or LIVs) and some of which are not. In recognition, we evaluate the probability that two faces have the same underlying identity cause. We make these ideas concrete by developing a series of novel generative models which incorporate both within-individual and between-individual variation. We consider both the linear case, where signal and noise are represented by a subspace, and the nonlinear case, where an arbitrary face manifold can be described and noise is position-dependent. We also develop a "tied" version of the algorithm that allows explicit comparison of faces across quite different viewing conditions. We demonstrate that our model produces results that are comparable to or better than the state of the art for both frontal face recognition and face recognition under varying pose.
Simon Prince, Yun Fu 0002, Umar Mohammed, James H. Elder
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 Image registration for foveated panoramic sensing
abstract
This article addresses the problem of registering high-resolution, small field-of-view images with low-resolution panoramic images provided by a panoramic catadioptric video sensor. Such systems may find application in surveillance and telepresence systems that require a large field of view and high resolution at selected locations. Although image registration has been studied in more conventional applications, the problem of registering panoramic and conventional video has not previously been addressed, and this problem presents unique challenges due to (i) the extreme differences in resolution between the sensors (more than a 16:1 linear resolution ratio in our application), and (ii) the resolution inhomogeneity of panoramic images. The main contributions of this article are as follows. First, we introduce our foveated panoramic sensor design. Second, we show how a coarse registration can be computed from the raw images using parametric template matching techniques. Third, we propose two refinement methods allowing automatic and near real-time registration between the two image streams. The first registration method is based on matching extracted interest points using a closed form method. The second registration method is featureless and based on minimizing the intensity discrepancy allowing the direct recovery of both the geometric and the photometric transforms. Fourth, a comparison between the two registration methods is carried out, which shows that the featureless method is superior in accuracy. Registration examples using the developed methods are presented.
Fadi Dornaika, James H. Elder
ACM Trans. Multim. Comput. Commun. Appl.2
2010 On growth and formlets: Sparse multi-scale coding of planar shape
abstract
This paper presents a sparse representation of 2D planar shape through the composition of warping functions, termed formlets, localized in scale and space. Each formlet subjects the 2D space in which the shape is embedded to a localized isotropic radial deformation. By constraining these localized warping transformations to be diffeomorphisms, the topology of shape is preserved, and the set of simple closed curves is closed under any sequence of these warpings. A generative model based on a composition of formlets applied to an embryonic shape, e.g., an ellipse, has the advantage of synthesizing only those shapes that could correspond to the boundaries of physical objects. To compute the set of formlets that represent a given boundary, we demonstrate a greedy coarse-to-fine formlet pursuit algorithm that serves as a non-commutative generalization of matching pursuit for sparse approximations. We evaluate our method by pursuing partially occluded shapes, comparing performance against a contour-based sparse shape coding framework.
Timothy D. Oleskiw, James H. Elder, Gabriel Peyré
CVPR2
2009 Hierarchical appearance-based classifiers for qualitative spatial localization
abstract
This paper presents a novel appearance-based technique for qualitative spatial localization. A vocabulary of visual words is built automatically, representing local features that repeatedly occur in the set of training images. An information maximization technique is then applied to build a hierarchical classifier for each environment by learning informative visual words. Child nodes in this hierarchy encode information redundant with information coded by their parents. In localization, hierarchical classifiers are used in a top-down manner, where top-level visual words are examined first, and for each top-level visual word which does not respond as expected, its lower-level visual words are examined. This allows inference to recover from missing features encoded by higher-level visual words. Several experiments on a challenging localization database demonstrate the advantages of our hierarchical framework and show a significant improvement over the traditional bag-of-features approaches.
Ehsan Fazl Ersi, James H. Elder, John K. Tsotsos
IROS2
2008 Efficient Edge-Based Methods for Estimating Manhattan Frames in Urban Imagery
Patrick Denis, James H. Elder, Francisco J. Estrada
ECCV (2)2
2008 Tied Factor Analysis for Face Recognition across Large Pose Differences
abstract
Face recognition algorithms perform very unreliably when the pose of the probe face is different from the gallery face: typical feature vectors vary more with pose than with identity. We propose a generative model that creates a one-to-many mapping from an idealized "identity" space to the observed data space. In identity space, the representation for each individual does not vary with pose. We model the measured feature vector as being generated by a pose-contingent linear transformation of the identity variable in the presence of Gaussian noise. We term this model "tied" factor analysis. The choice of linear transformation (factors) depends on the pose, but the loadings are constant (tied) for a given individual. We use the EM algorithm to estimate the linear transformations and the noise parameters from training data. We propose a probabilistic distance metric which allows a full posterior over possible matches to be established. We introduce a novel feature extraction process and investigate recognition performance using the FERET, XM2VTS and PIE databases. Recognition performance compares favourably to contemporary approaches.
Simon Prince, James H. Elder, Jonathan Warrell, Fatima M. Felisberti
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Probabilistic Linear Discriminant Analysis for Inferences About Identity
abstract
Many current face recognition algorithms perform badly when the lighting or pose of the probe and gallery images differ. In this paper we present a novel algorithm designed for these conditions. We describe face data as resulting from a generative model which incorporates both within-individual and between-individual variation. In recognition we calculate the likelihood that the differences between face images are entirely due to within-individual variability. We extend this to the non-linear case where an arbitrary face manifold can be described and noise is position-dependent. We also develop a "tied" version of the algorithm that allows explicit comparison across quite different viewing conditions. We demonstrate that our model produces state of the art results for (i) frontal face recognition (ii) face recognition under varying pose.
Simon Prince, James H. Elder
ICCV2
2007 Pre-Attentive and Attentive Detection of Humans in Wide-Field Scenes
James H. Elder, Simon Prince, Yuqian Hou, Mikhail Sizintsev, E. Olevskiy
Int. J. Comput. Vis.1
2006 Tied Factor Analysis for Face Recognition Across Large Pose Changes
abstract
Abstract—Face recognition algorithms perform very unreliably when the pose of the probe face is different from the gallery face: typical feature vectors vary more with pose than with identity. We propose a generative model that creates a one-to-many mapping from an idealized “identity ” space to the observed data space. In identity space, the representation for each individual does not vary with pose. We model the measured feature vector as being generated by a pose-contingent linear transformation of the identity variable in the presence of Gaussian noise. We term this model “tied ” factor analysis. The choice of linear transformation (factors) depends on the pose, but the loadings are constant (tied) for a given individual. We use the EM algorithm to estimate the linear transformations and the noise parameters from training data. We propose a probabilistic distance metric that allows a full posterior over possible matches to be established. We introduce a novel feature extraction process and investigate recognition performance by using the FERET, XM2VTS, and PIE databases. Recognition performance compares favorably with contemporary approaches. Index Terms—Computing methodologies, pattern recognition, applications, face and gesture recognition. Ç 1
Simon Prince, James H. Elder
BMVC2
2005 Creating Invariance to "Nuisance Parameters" in Face Recognition
abstract
A major goal for face recognition is to identify faces where the pose of the probe is different from the stored face. Typical feature vectors vary more with pose than with identity, leading to very poor recognition performance. We propose a non-linear many-to-one mapping from a conventional feature space to a new space constructed so that each individual has a unique feature vector regardless of pose. Training data is used to implicitly parameterize the position of the multi-dimensional face manifold by pose. We introduce a co-ordinate transform, which depends on the position on the manifold. This transform is chosen so that different poses of the same face are mapped to the same feature vector. The same approach is applied to illumination changes. We investigate different methods for creating features, which are invariant to both pose and illumination. We provide a metric to assess the discriminability of the resulting features. Our technique increases the discriminability of faces under unknown pose and lighting compared to contemporary methods.
Simon Prince, James H. Elder
CVPR (2)2
2005 Statistical Cue Integration for Foveated Wide-Field Surveillance
abstract
Reliable wide-field detection of human activity is an unsolved problem. The main difficulty is that low resolution and the unconstrained nature of realistic environments and human behaviour make form cues unreliable. Here we argue that reliability in far- or wide-field detection can still be achieved by probabilistic combination of multiple weak but complementary visual cues that do not depend on detailed form analysis. To demonstrate, we describe a real-time Bayesian algorithm for localizing human activity in relatively unconstrained scenes, using motion, background subtraction and skin colour cues. Fast sampling of scale space is achieved using integral images and a flexible norm that can handle sparse cues without loss of statistical power. We show that the probabilistic approach far outperforms a representative logical approach in which skin and background subtraction classifiers are combined conjunctively. Our method is currently used in a pre-attentive human activity sensor, generating saccadic targets for an attentive foveated vision system that reliably fixates faces over a 130 deg field of view, allowing high-resolution capture of facial images over a large dynamic scene.
Simon Prince, James H. Elder, Yuqian Hou, Mikhail Sizintsev, Yevgen Olevskiy
CVPR (2)2
2003 Contour Grouping with Prior Models
abstract
Conventional approaches to perceptual grouping assume little specific knowledge about the object(s) of interest. However, there are many applications in which such knowledge is available and useful. Here, we address the problem of finding the bounding contour of an object in an image when some prior knowledge about the object is available. We introduce a framework for combining prior probabilistic knowledge of the appearance of the object with probabilistic models for contour grouping. A constructive search technique is used to compute candidate closed object boundaries, which are then evaluated by combining figure, ground, and prior probabilities to compute the maximum a posteriori estimate. A significant advantage of our formulation is that it rigorously combines probabilistic local cues with important global constraints such as simplicity (no self-intersections), closure, completeness, and nontrivial scale priors. We apply this approach to the problem of computing exact lake boundaries from satellite imagery, given approximate prior knowledge from an existing digital database. We quantitatively evaluate the performance of our algorithm and find that it exceeds the performance of human mapping experts and a competing active contour approach, even with relatively weak prior knowledge. While the priors may be task-specific, the approach is general, as we demonstrate by applying it to a completely different problem: the computation of human skin boundaries in natural imagery.
James H. Elder, Amnon Krupnik, Leigh A. Johnston
IEEE Trans. Pattern Anal. Mach. Intell.1
2002 Image Registration for Foveated Omnidirectional Sensing
Fadi Dornaika, James H. Elder
ECCV (4)2
2001 Contour Grouping with Strong Prior Models
abstract
Conventional approaches to perceptual grouping assume little specific knowledge about the object(s) of interest. However, there are many applications in which such knowledge is available and useful. We address the problem of finding the bounding contour of an object in an image when some prior knowledge about the object is available. We introduce a framework for combining prior probabilistic knowledge of the appearance of the object with probabilistic models for contour grouping. While prior probabilistic approaches have employed shortest-path algorithms to compute contours, this approach is limited in that many global properties cannot easily be incorporated in the computation. We propose as an alternative an approximate, constructive search technique, which finds a good (not necessarily optimal) solution, and which can accommodate important global cues and constraints. We apply this approach to the problem of computing exact lake boundaries from satellite imagery, given approximate prior models from an existing digital database. Our algorithm improves the accuracy of the prior GIS lake models by an average of 41%.
James H. Elder, Amnon Krupnik
CVPR (2)1
2001 Image Editing in the Contour Domain
abstract
We propose a novel method for image editing in which the primitive working unit is not a pixel but an edge. The feasibility of this proposal is suggested by the recent work of Elder et al. (1998) showing that a gray-scale image can be accurately represented by its edge map if a suitable edge model and scale selection method are employed. In particular, an efficient algorithm has been reported by Elder et al. (1996) and Elder (1999) to invert such an edge representation to yield a high-fidelity reconstruction of the original image. We combined these algorithms together with an efficient method for contour grouping and an intuitive user interface to allow users to perform image editing operations directly in the contour domain. Experimental results suggest that this novel combination of vision algorithms may increase the efficiency of certain classes of image editing operations.
James H. Elder, Richard M. Goldberg
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Are Edges Incomplete?
James H. Elder
Int. J. Comput. Vis.1
1998 Image Editing in the Contour Domain
abstract
Image editing systems are essentially pixel-based. In this paper we propose a novel method for image editing in which the primitive working unit is not a pixel but an edge. The feasibility of this proposal is suggested by recent work showing that a grey-scale image can be accurately represented by its edge map if a suitable edge model and scale selection method are employed. In particular, an efficient algorithm has been reported to invert such an edge representation to yield a high-fidelity reconstruction of the original image. We have combined these algorithms together with an efficient method for contour grouping and an intuitive user interface to allow users to perform image editing operations directly in the contour domain. Experimental results suggest that this novel combination of vision algorithms may lead to substantial improvements in the efficiency of certain classes of image editing operations.
James H. Elder, Richard M. Goldberg
CVPR1
1998 Local Scale Control for Edge Detection and Blur Estimation
abstract
We show that knowledge of sensor properties and operator norms can be exploited to define a unique, locally computable minimum reliable scale for local estimation at each point in the image. This method for local scale control is applied to the problem of detecting and localizing edges in images with shallow depth of field and shadows. We show that edges spanning a broad range of blur scales and contrasts can be recovered accurately by a single system with no input parameters other than the second moment of the sensor noise. A natural dividend of this approach is a measure of the thickness of contours which can be used to estimate focal and penumbral blur. Local scale control is shown to be important for the estimation of blur in complex images, where the potential for interference between nearby edges of very different blur scale requires that estimates be made at the minimum reliable scale.
James H. Elder, Steven W. Zucker
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Space Scale Localization, Blur, and Contour-Based Image Coding
abstract
We have recently proposed a scale-adaptive algorithm for reliable edge detection and blur estimation. The algorithm produces a contour code which consists of estimates of position, brightness, contrast and blur for each edge point in the image. Here we address two questions: 1. Can scale adaptation be used to achieve precise localization of blurred edges? 2. How much of the perceptual content of an image is carried by the 1-D contour code? We report an efficient algorithm for subpixel localization, and show that local scale control allows excellent precision even for highly blurred edges. We further show how local scale control can quantitatively account for human visual acuity of blurred edge stimuli. To address the question of perceptual content, we report an algorithm for inverting the contour code to reconstruct an estimate of the original image. While reconstruction based on edge brightness and contrast alone introduces significant artifact, restitution of the local blur signal is shown to produce perceptually accurate reconstructions.
James H. Elder, Steven W. Zucker
CVPR1
1996 Computing Contour Closure
James H. Elder, Steven W. Zucker
ECCV (1)1
1996 Local Scale Control for Edge Detection and Blur Estimation
James H. Elder, Steven W. Zucker
ECCV (2)1
1995 Shadows, Defocus and Reliable Estimation
James H. Elder, Steven W. Zucker
CAIP1