Robert T. Collins

dblp:76/2806 · DBLP profile ↗
← Back
75ranked-venue papers
18as first author
1since 2021 · last 2025
0000-0001-9062-4252ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 16 first-authorGraphics, computer vision, multimedia, augmented reality and games · 58 · 13 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
43 papers
Video understanding and tracking · 32% 3D vision · 20% Image recognition and object detection · 12%
Computer graphics and multimedia
10 papers
Geometric modeling and processing · 46% Image and video processing · 29% Multimedia analysis and retrieval · 22%

Topics — the 30 heaviest of 102, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
multi-object tracking
0.652013
Tracking Sports Players with Context-Conditioned Motion Models · CVPR 2013
Multi-target Tracking by Lagrangian Relaxation to Min-cost Network Flow · CVPR 2013
Vision-Based Analysis of Small Groups in Pedestrian Crowds · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Computer vision › Video understanding and tracking
object tracking
0.582009
Shape constrained figure-ground segmentation and tracking · CVPR 2009
Object tracking and detection after occlusion via numerical hybrid local and global mode-seeking · CVPR 2008
On-the-fly Object Modeling while Tracking · CVPR 2007
Computer vision › Face, body and person analysis
human pose analysis
0.412020
From Image to Stability: Learning Dynamics from Human Pose · ECCV (23) 2020
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.322013
Multi-target Tracking by Lagrangian Relaxation to Min-cost Network Flow · CVPR 2013
Multitarget data association with higher-order motion models · CVPR 2012
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation
0.332010
Data driven mean-shift belief propagation for non-gaussian MRFs · CVPR 2010
Deformed Lattice Detection in Real-World Images Using Mean-Shift Belief Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Efficient mean shift belief propagation for vision tracking · CVPR 2008
Computer vision › 3D vision
feature matching
0.322016
Regularity-Driven Building Facade Matching between Aerial and Street Views · CVPR 2016
A Space-Sweep Approach to True Multi-Image Matching · CVPR 1996
Computer vision › 3D vision
multi-view geometry
0.212016
Regularity-Driven Building Facade Matching between Aerial and Street Views · CVPR 2016
Computer vision › Video understanding and tracking
crowd analysis
0.222012
Vision-Based Analysis of Small Groups in Pedestrian Crowds · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Marked point processes for crowd counting · CVPR 2009
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference
0.222010
Data driven mean-shift belief propagation for non-gaussian MRFs · CVPR 2010
Efficient mean shift belief propagation for vision tracking · CVPR 2008
Machine learning › Learning theory › statistical estimation
nonparametric estimation
0.222010
Data driven mean-shift belief propagation for non-gaussian MRFs · CVPR 2010
Efficient mean shift belief propagation for vision tracking · CVPR 2008
Robotics › Motion planning and robot control
trajectory optimization
0.212014
Hybrid Stochastic / Deterministic Optimization for Tracking Sports Players and Pedestrians · ECCV (2) 2014
Geometric modeling and processing › feature recognition
lattice detection
0.222009
Deformed Lattice Detection in Real-World Images Using Mean-Shift Belief Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Deformed Lattice Discovery Via Efficient Mean-Shift Belief Propagation · ECCV (2) 2008
Computer vision › Video understanding and tracking › motion analysis › motion modeling
motion model
0.212013
Tracking Sports Players with Context-Conditioned Motion Models · CVPR 2013
Computer vision › Image recognition and object detection
object detection
0.212013
Optimized Pedestrian Detection for Multiple and Occluded People · CVPR 2013
Computer vision › Image recognition and object detection › pedestrian detection
occluded pedestrian detection
0.212013
Optimized Pedestrian Detection for Multiple and Occluded People · CVPR 2013
Computer vision › Image recognition and object detection
pedestrian detection
0.212013
Optimized Pedestrian Detection for Multiple and Occluded People · CVPR 2013
Computer vision › Video understanding and tracking › multi-object tracking
player tracking
0.212013
Tracking Sports Players with Context-Conditioned Motion Models · CVPR 2013
Computer vision › Segmentation and scene understanding › image segmentation › model-based segmentation
shape-prior segmentation
0.122009
Shape constrained figure-ground segmentation and tracking · CVPR 2009
Corrected Laplacians: Closer Cuts and Segmentation with Shape Priors · CVPR (2) 2005
Computer vision › 3D vision › human mesh recovery
human body shape estimation
0.112012
A Generative Model for Simultaneous Estimation of Human Body Shape and Pixel-Level Segmentation · ECCV (5) 2012
Computer vision › Segmentation and scene understanding › image segmentation
pixel-level segmentation
0.112012
A Generative Model for Simultaneous Estimation of Human Body Shape and Pixel-Level Segmentation · ECCV (5) 2012
Computer vision › Image recognition and object detection › object detection › multi-object detection
crowd detection
0.112010
Crowd Detection with a Multiview Sampler · ECCV (5) 2010
Multimedia analysis and retrieval
geotagging
0.112010
Beyond GPS: determining the camera viewing direction of a geotagged image · ACM Multimedia 2010
Geometric modeling and processing › shape analysis
symmetry detection
0.132004
A Computational Model for Periodic Pattern Perception Based on Frieze and Wallpaper Groups · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Skewed Symmetry Groups · CVPR (1) 2001
A Computational Model for Repeated Pattern Perception Using Frieze and Wallpaper Groups · CVPR 2000
Computer vision › Image recognition and object detection › object counting
crowd counting
0.112009
Marked point processes for crowd counting · CVPR 2009
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation
0.112009
Shape constrained figure-ground segmentation and tracking · CVPR 2009
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
marked point processes
0.112009
Marked point processes for crowd counting · CVPR 2009
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.112009
Deformed Lattice Detection in Real-World Images Using Mean-Shift Belief Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.112009
Marked point processes for crowd counting · CVPR 2009
Multimedia analysis and retrieval
image analysis
0.112009
Deformed Lattice Detection in Real-World Images Using Mean-Shift Belief Propagation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Computer vision › 3D vision
3d reconstruction
0.152002
Three-Dimensional Reconstruction of Points and Lines with Unknown Correspondence · Int. J. Comput. Vis. 2002
Three-Dimensional Reconstruction of Points and Lines with Unknown Correspondence across Images · Int. J. Comput. Vis. 2001
Site Model Acquisition and Extension from Aerial Images · ICCV 1995

Methods — techniques the papers use, named apart from their topics

belief propagation · 0.6pose estimation · 0.4dynamics learning · 0.4mean shift · 0.3shape context · 0.2self-similarity descriptors · 0.2SIFT · 0.2stochastic optimization · 0.2deterministic optimization · 0.2lagrangian relaxation · 0.2linear assignment · 0.1iterated conditional modes · 0.1visual matching · 0.1scale space theory · 0.1near-orthogonal view matching · 0.1mean-shift · 0.1thin-plate spline warping · 0.1simulated annealing · 0.1
YearPublicationVenuePosition
2025 Recurrence-Based Vanishing Point Detection
Skanda Bharadwaj, Robert T. Collins, Yanxi Liu 0001
WACV2
2020 From Image to Stability: Learning Dynamics from Human Pose
Jesse Scott, Bharadwaj Ravichandran, Christopher Funk, Robert T. Collins, Yanxi Liu 0001
ECCV (23)4
2016 Regularity-Driven Building Facade Matching between Aerial and Street Views
abstract
We present an approach for detecting and matching building facades between aerial view and street-view images. We exploit the regularity of urban scene facades as captured by their lattice structures and deduced from median-tiles' shape context, color, texture and spatial similarities. Our experimental results demonstrate effective matching of oblique and partially-occluded facades between aerial and ground views. Quantitative comparisons for automated urban scene facade matching from three cities show superior performance of our method over baseline SIFT, Root-SIFT and the more sophisticated Scale-Selective Self-Similarity and Binary Coherent Edge descriptors. We also illustrate regularity-based applications of occlusion removal from street views and higher-resolution texture-replacement in aerial views.
Mark Wolff, Robert T. Collins, Yanxi Liu 0001
CVPR2
2016 Assessing tracking performance in complex scenarios using mean time between failures
abstract
Existing measures for evaluating the performance of tracking algorithms are difficult to interpret, which makes it hard to identify the best approach for a particular situation. As we show, a dummy algorithm which does not actually track scores well under most existing measures. Although some measures characterize specific error sources quite well, combining them into a single aggregate measure for comparing approaches or tuning parameters is not straightforward. In this work we propose `mean time between failures' as a viable summary of solution quality - especially when the goal is to follow objects for as long as possible. In addition to being sensitive to all tracking errors, the performance numbers are directly interpretable: how long can an algorithm operate before a mistake has likely occurred (the object is lost, its identity is confused, etc.)? We illustrate the merits of this measure by assessing solutions from different algorithms on a challenging dataset.
Peter Carr 0001, Robert T. Collins
WACV2
2014 Hybrid Stochastic / Deterministic Optimization for Tracking Sports Players and Pedestrians
Robert T. Collins, Peter Carr 0001
ECCV (2)1
2014 Estimating the camera direction of a geotagged image using reference images
Jiebo Luo 0001, Robert T. Collins, Yanxi Liu 0001
Pattern Recognit.3
2013 Multi-target Tracking by Lagrangian Relaxation to Min-cost Network Flow
abstract
We propose a method for global multi-target tracking that can incorporate higher-order track smoothness constraints such as constant velocity. Our problem formulation readily lends itself to path estimation in a trellis graph, but unlike previous methods, each node in our network represents a candidate pair of matching observations between consecutive frames. Extra constraints on binary flow variables in the graph result in a problem that can no longer be solved by min-cost network flow. We therefore propose an iterative solution method that relaxes these extra constraints using Lagrangian relaxation, resulting in a series of problems that ARE solvable by min-cost flow, and that progressively improve towards a high-quality solution to our original optimization problem. We present experimental results showing that our method outperforms the standard network-flow formulation as well as other recent algorithms that attempt to incorporate higher-order smoothness constraints.
Asad A. Butt, Robert T. Collins
CVPR2
2013 Tracking Sports Players with Context-Conditioned Motion Models
abstract
We employ hierarchical data association to track players in team sports. Player movements are often complex and highly correlated with both nearby and distant players. A single model would require many degrees of freedom to represent the full motion diversity and could be difficult to use in practice. Instead, we introduce a set of Game Context Features extracted from noisy detections to describe the current state of the match, such as how the players are spatially distributed. Our assumption is that players react to the current situation in only a finite number of ways. As a result, we are able to select an appropriate simplified affinity model for each player and time instant using a random decision forest based on current track and game context features. Our context-conditioned motion models implicitly incorporate complex inter-object correlations while remaining tractable. We demonstrate significant performance improvements over existing multi-target tracking algorithms on basketball and field hockey sequences several minutes in duration and containing 10 and 20 players respectively.
Jingchen Liu, Peter Carr 0001, Robert T. Collins, Yanxi Liu 0001
CVPR3
2013 Optimized Pedestrian Detection for Multiple and Occluded People
abstract
We present a quadratic unconstrained binary optimization (QUBO) framework for reasoning about multiple object detections with spatial overlaps. The method maximizes an objective function composed of unary detection confidence scores and pairwise overlap constraints to determine which overlapping detections should be suppressed, and which should be kept. The framework is flexible enough to handle the problem of detecting objects as a shape covering of a foreground mask, and to handle the problem of filtering confidence weighted detections produced by a traditional sliding window object detector. In our experiments, we show that our method outperforms two existing state-of-the-art pedestrian detectors.
Sitapa Watcharapinchai, Robert T. Collins
CVPR2
2013 Robust autocalibration for a surveillance camera network
abstract
We propose a novel approach for multi-camera autocalibration by observing multiview surveillance video of pedestrians walking through the scene. Unlike existing methods, we do NOT require tracking or explicit correspondences of the same person across time/views. Instead, we take noisy foreground blobs as the only input and rely on a joint optimization framework with robust statistics to achieve accurate calibration under challenging scenarios. First, each individual camera is roughly calibrated into its local World Coordinate System (lWCS) based on analysis of relative 3D pedestrian height distribution. Then, all lWCSs are iteratively registered with respect to a shared global World Coordinate System (gWCS) by incorporating robust matching with a partial Direct Linear Transform (pDLT). As demonstrated by extensive evaluation, our algorithm achieves satisfactory results in various camera settings with up to moderate crowd densities with a large proportion of foreground outliers.
Jingchen Liu, Robert T. Collins, Yanxi Liu 0001
WACV2
2012 Multiple Target Tracking Using Frame Triplets
Asad A. Butt, Robert T. Collins
ACCV (3)2
2012 Multitarget data association with higher-order motion models
abstract
We present an iterative approximate solution to the multidimensional assignment problem under general cost functions. The method maintains a feasible solution at every step, and is guaranteed to converge. It is similar to the iterated conditional modes (ICM) algorithm, but applied at each step to a block of variables representing correspondences between two adjacent frames, with the optimal conditional mode being calculated exactly as the solution to a two-frame linear assignment problem. Experiments with ground-truthed trajectory data show that the method outperforms both network-flow data association and greedy recursive filtering using a constant velocity motion model.
Robert T. Collins
CVPR1
2012 A Generative Model for Simultaneous Estimation of Human Body Shape and Pixel-Level Segmentation
Ingmar Rauschert, Robert T. Collins
ECCV (5)2
2012 Vision-Based Analysis of Small Groups in Pedestrian Crowds
abstract
Building upon state-of-the-art algorithms for pedestrian detection and multi-object tracking, and inspired by sociological models of human collective behavior, we automatically detect small groups of individuals who are traveling together. These groups are discovered by bottom-up hierarchical clustering using a generalized, symmetric Hausdorff distance defined with respect to pairwise proximity and velocity. We validate our results quantitatively and qualitatively on videos of real-world pedestrian scenes. Where human-coded ground truth is available, we find substantial statistical agreement between our results and the human-perceived small group structure of the crowd. Results from our automated crowd analysis also reveal interesting patterns governing the shape of pedestrian groups. These discoveries complement current research in crowd dynamics, and may provide insights to improve evacuation planning and real-time situation awareness during public disturbances.
Weina Ge, Robert T. Collins, Barry Ruback
IEEE Trans. Pattern Anal. Mach. Intell.2
2011 Automatic Surveillance Camera Calibration without Pedestrian Tracking
Jingchen Liu, Robert T. Collins, Yanxi Liu 0001
BMVC2
2010 Translation-Symmetry-Based Perceptual Grouping with Applications to Urban Scenes
Kyle Brocklehurst, Robert T. Collins, Yanxi Liu 0001
ACCV (3)3
2010 Image De-fencing Revisited
Kyle Brocklehurst, Robert T. Collins, Yanxi Liu 0001
ACCV (4)3
2010 Data driven mean-shift belief propagation for non-gaussian MRFs
abstract
We introduce a novel data-driven mean-shift belief propagation (DDMSBP) method for non-Gaussian MRFs, which often arise in computer vision applications. With the aid of scale space theory, optimization of non-Gaussian, multimodal MRF models using DDMSBP becomes less sensitive to local maxima. This is a significant improvement over standard BP inference, and extends the range of methods that are computationally tractable. In particular, when pair-wise potentials are Gaussians, the time complexity of DDMSBP becomes bilinear in the numbers of states and nodes in the MRF. Experimental results from simulation and non-rigid deformable neuroimage registration demonstrate that our method is faster and more accurate than state-of-the-art inference algorithms.
Somesh Kashyap, Robert T. Collins, Yanxi Liu 0001
CVPR3
2010 A GPU based implementation of Center-Surround Distribution Distance for feature extraction and matching
abstract
The release of general purpose GPU programming environments has garnered universal access to computing performance that was once only available to super-computers. The availability of such computational power has fostered the creation and re-deployment of algorithms, new and old, creating entirely new classes of applications. In this paper, a GPU implementation of the Center-Surround Distribution Distance (CSDD) algorithm for detecting features within images and video is presented. While an optimized CPU implementation requires anywhere from several seconds to tens of minutes to perform analysis of an image, the GPU based approach has the potential to improve upon this by up to 28X, with no loss in accuracy.
Aditi Rathi, Michael DeBole, Weina Ge, Robert T. Collins, Narayanan Vijaykrishnan
DATE4
2010 Crowd Detection with a Multiview Sampler
Weina Ge, Robert T. Collins
ECCV (5)2
2010 RISN: An Efficient, Dynamically Tasked and Interoperable Sensor Network Overlay
abstract
Sensor network has given rise to a comprehensive view of various environments through their harnessed data. Usage of the data is, however, limited since the current sensor network paradigm promotes isolated networks that are statically tasked. In recent years, users have become mobile entities that require constant access to data for efficient processing. Under the current limitations of sensor networks, users would be restricted to using only a subset of the vast amount of data being collected, depending on the networks they are able to access. Through reliance on isolated networks, proliferation of sensor nodes can easily occur in any area that has high appeals to users. Furthermore, support for dynamic tasking of nodes and efficient processing of data is contrary to the general view of sensor networks as subject to severe resource constraints. Addressing the aforementioned challenges requires the deployment of a system that allows users to take full advantage of data collected in the area of interest to their tasks. In light of these observations, we introduce a hardware-overlay system designed to allow users to efficiently collect and utilize data from various heterogeneous sensor networks. The hardware-overlay takes advantage of FPGA devices and the mobile agent paradigm in order to efficiently collect and process data from cooperating networks. The computational and power efficiency of the system are herein demonstrated.
Evens Jean, Robert T. Collins, Ali R. Hurson, Sahra Sedigh Sarvestani
EUC2
2010 Beyond GPS: determining the camera viewing direction of a geotagged image
abstract
Increasingly, geographic information is being associated with personal photos. Recent research results have shown that the additional global positioning system (GPS) information helps visual recognition for geotagged photos by providing location context. However, the current GPS data only identifies the camera location, leaving the viewing direction uncertain. To produce more precise location information, i.e. the viewing direction for geotagged photos, we utilize both Google Street View and Google Earth satellite images. Our proposed system is two-pronged: 1) visual matching between a user photo and any available street views in the vicinity determine the viewing direction, and 2) when only an overhead satellite view is available, near-orthogonal view matching between the user photo and satellite imagery computes the viewing direction. Experimental results have shown the promise of the proposed framework.
Jiebo Luo 0001, Robert T. Collins, Yanxi Liu 0001
ACM Multimedia3
2009 Marked point processes for crowd counting
abstract
A Bayesian marked point process (MPP) model is developed to detect and count people in crowded scenes. The model couples a spatial stochastic process governing number and placement of individuals with a conditional mark process for selecting body shape. We automatically learn the mark (shape) process from training video by estimating a mixture of Bernoulli shape prototypes along with an extrinsic shape distribution describing the orientation and scaling of these shapes for any given image location. The reversible jump Markov Chain Monte Carlo framework is used to efficiently search for the maximum a posteriori configuration of shapes, leading to an estimate of the count, location and pose of each person in the scene. Quantitative results of crowd counting are presented for two publicly available datasets with known ground truth.
Weina Ge, Robert T. Collins
CVPR2
2009 Shape constrained figure-ground segmentation and tracking
abstract
Global shape information is an effective top-down complement to bottom-up figure-ground segmentation as well as a useful constraint to avoid drift during adaptive tracking. We propose a novel method to embed global shape information into local graph links in a Conditional Random Field (CRF) framework. Given object shapes from several key frames, we automatically collect a shape dataset on-the-fly and perform statistical analysis to build a collection of deformable shape templates representing global object shape. In new frames, simulated annealing and local voting align the deformable template with the image to yield a global shape probability map. The global shape probability is combined with a region-based probability of object boundary map and the pixel-level intensity gradient to determine each link cost in the graph. The CRF energy is minimized by min-cut, followed by Random Walk on the uncertain boundary region to get a soft segmentation result. Experiments on both medical and natural images with deformable object shapes are demonstrated.
Zhaozheng Yin, Robert T. Collins
CVPR2
2009 Automatically detecting the small group structure of a crowd
abstract
Recent work on computer vision analysis of crowds tends to focus on robustly tracking individuals through the crowd or on analyzing the overall pattern of flow. Our work seeks a deeper analysis of social behavior by identifying the small group structure of crowds, forming the basis for mid-level activity analysis at the granularity of human social groups. Building upon state-of-the-art algorithms for pedestrian detection and multi-object tracking, and inspired by social science models of human collective behavior, we automatically detect small groups of individuals who are traveling together. These groups are discovered using a bottom-up hierarchical clustering approach that compares sets of individuals based on a generalized, symmetric Hausdorff distance defined with respect to pairwise proximity and velocity. We validate our results quantitatively and qualitatively on videos of real-world pedestrian scenes. Where human-coded ground truth is available, we find substantial statistical agreement between our results and the human-perceived small group structure of the crowd.
Weina Ge, Robert T. Collins, Barry Ruback
WACV2
2009 Deformed Lattice Detection in Real-World Images Using Mean-Shift Belief Propagation
abstract
We propose a novel and robust computational framework for automatic detection of deformed 2D wallpaper patterns in real-world images. The theory of 2D crystallographic groups provides a sound and natural correspondence between the underlying lattice of a deformed wallpaper pattern and a degree-4 graphical model. We start the discovery process with unsupervised clustering of interest points and voting for consistent lattice unit proposals. The proposed lattice basis vectors and pattern element contribute to the pairwise compatibility and joint compatibility (observation model) functions in a Markov Random Field (MRF). Thus, we formulate the 2D lattice detection as a spatial, multitarget tracking problem, solved within an MRF framework using a novel and efficient Mean-Shift Belief Propagation (MSBP) method. Iterative detection and growth of the deformed lattice are interleaved with regularized thin-plate spline (TPS) warping, which rectifies the current deformed lattice into a regular one to ensure stability of the MRF model in the next round of lattice recovery. We provide quantitative comparisons of our proposed method with existing algorithms on a diverse set of 261 real-world photos to demonstrate significant advances in accuracy and speed over the state of the art in automatic discovery of regularity in real images.
Kyle Brocklehurst, Robert T. Collins, Yanxi Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Multi-target Data Association by Tracklets with Unsupervised Parameter Estimation
abstract
We consider multi-target tracking via probabilistic data association among tracklets (trajectory fragments), a mid-level representation that provides good spatio-temporal context for efficient tracking. Model parameter estimation and the search for the best association among tracklets are unified naturally within a Markov Chain Monte Carlo sampling procedure. The proposed approach is able to infer the optimal model parameters for different tracking scenarios in an unsupervised manner. 1
Weina Ge, Robert T. Collins
BMVC2
2008 Online Figure-ground Segmentation with Edge Pixel Classification
abstract
The need for figure-ground segmentation in video arises in many vision problems like tracker initialization, accurate object shape representation and drift-free appearance model adaptation. This paper uses a 3D spatio-temporal Conditional Random Field (CRF) to combine different segmentation cues while enforcing temporal coherence. Without supervised parameter training, the weighting factors for different data potential functions in the CRF model are adapted online to reflect changes in object appearance and environment. To get an accurate boundary based on the 3D CRF segmentation result, edge pixels are classified into three classes: foreground, background and boundary. The final foreground region bitmask is constructed from the foreground and boundary edge pixels. The effectiveness of our approach is demonstrated on several airborne videos with large appearance change and heavy occlusion. 1
Zhaozheng Yin, Robert T. Collins
BMVC2
2008 Rotation symmetry group detection via frequency analysis of frieze-expansions
abstract
We present a novel and effective algorithm for rotation symmetry group detection from real-world images. We propose a frieze-expansion method that transforms rotation symmetry group detection into a simple translation symmetry detection problem. We define and construct a dense symmetry strength map from a given image, and search for potential rotational symmetry centers automatically. Frequency analysis, using discrete Fourier transform (DFT), is applied to the frieze-expansion patterns to uncover the types and the cardinality of multiple rotation symmetry groups in an image, concentric or otherwise. Furthermore, our detection algorithm can discriminate discrete versus continuous and cyclic versus dihedral symmetry groups, and identify the corresponding supporting regions in the image. Experimental results on over 80 synthetic and natural images demonstrate superior performance of our rotation detection algorithm in accuracy and in speed over the state of the art rotation detection algorithms.
Seungkyu Lee 0001, Robert T. Collins, Yanxi Liu 0001
CVPR2
2008 Efficient mean shift belief propagation for vision tracking
abstract
A mechanism for efficient mean-shift belief propagation (MSBP) is introduced. The novelty of our work is to use mean-shift to perform nonparametric mode-seeking on belief surfaces generated within the belief propagation framework. Belief propagation (BP) is a powerful solution for performing inference in graphical models. However, there is a quadratic increase in the cost of computation with respect to the size of the hidden variable space. While the recently proposed nonparametric belief propagation (NBP) has better performance in terms of speed, even for continuous hidden variable spaces, computation is still slow due to the particle filter sampling process. Our MSBP method only needs to compute a local grid of samples of the belief surface during each iteration. This approach needs a significantly smaller number of samples than NBP, reducing computation time, yet it also yields more accurate and stable solutions. The efficiency and robustness of MSBP is compared against other variants of BP on applications in multi-target tracking and 2D articulated body tracking.
Yanxi Liu 0001, Robert T. Collins
CVPR3
2008 Object tracking and detection after occlusion via numerical hybrid local and global mode-seeking
abstract
Given an object model and a black-box measure of similarity between the model and candidate targets, we consider visual object tracking as a numerical optimization problem. During normal tracking conditions when the object is visible from frame to frame, local optimization is used to track the local mode of the similarity measure in a parameter space of translation, rotation and scale. However, when the object becomes partially or totally occluded, such local tracking is prone to failure, especially when common prediction techniques like the Kalman filter do not provide a good estimate of object parameters in future frames. To recover from these inevitable tracking failures, we consider object detection as a global optimization problem and solve it via Adaptive Simulated Annealing (ASA), a method that avoids becoming trapped at local modes and is much faster than exhaustive search. As a Monte Carlo approach, ASA stochastically samples the parameter space, in contrast to local deterministic search. We apply cluster analysis on the sampled parameter space to redetect the object and renew the local tracker. Our numerical hybrid local and global mode-seeking tracker is validated on challenging airborne videos with heavy occlusion and large camera motions. Our approach outperforms state-of-the-art trackers on the VIVID benchmark datasets.
Zhaozheng Yin, Robert T. Collins
CVPR2
2008 CSDD Features: Center-Surround Distribution Distance for Feature Extraction and Matching
Robert T. Collins, Weina Ge
ECCV (3)1
2008 Deformed Lattice Discovery Via Efficient Mean-Shift Belief Propagation
Robert T. Collins, Yanxi Liu 0001
ECCV (2)2
2008 Likelihood Map Fusion for Visual Object Tracking
abstract
Visual object tracking can be considered as a figure-ground classification task. In this paper, different features are used to generate a set of likelihood maps for each pixel indicating the probability of that pixel belonging to foreground object or scene background. For example, intensity, texture, motion, saliency and template matching can all be used to generate likelihood maps. We propose a generic likelihood map fusion framework to combine these heterogeneous features into a fused soft segmentation suitable for mean-shift tracking. All the component likelihood maps contribute to the segmentation based on their classification confidence scores (weights) learned from the previous frame. The evidence combination framework dynamically updates the weights such that, in the fused likelihood map, discriminative foreground/background information is preserved while ambiguous information is suppressed. The framework is applied here to track ground vehicles from thermal airborne video, and is also compared to other state-of-the-art algorithms.
Zhaozheng Yin, Fatih Porikli, Robert T. Collins
WACV3
2007 Shape Variation-Based Frieze Pattern for Robust Gait Recognition
abstract
Gait is an attractive biometric for vision-based human identification. Previous work on existing public data sets has shown that shape cues yield improved recognition rates compared to pure motion cues. However, shape cues are fragile to gross appearance variations of an individual, for example, walking while carrying a ball or a backpack. We introduce a novel, spatiotemporal shape variation-based frieze pattern (SVB frieze pattern) representation for gait, which captures motion information over time. The SVB frieze pattern represents normalized frame difference over gait cycles. Rows/columns of the vertical/horizontal SVB frieze pattern contain motion variation information augmented by key frame information with body shape. A temporal symmetry map of gait patterns is also constructed and combined with vertical/horizontal SVB frieze patterns for measuring the dissimilarity between gait sequences. Experimental results show that our algorithm improves gait recognition performance on sequences with and without gross differences in silhouette shape. We demonstrate superior performance of this computational framework over previous algorithms using shape cues alone on both CMU MoBo and UoS HumanID gait databases.
Seungkyu Lee 0001, Yanxi Liu 0001, Robert T. Collins
CVPR3
2007 Belief Propagation in a 3D Spatio-temporal MRF for Moving Object Detection
abstract
Previous pixel-level change detection methods either contain a background updating step that is costly for moving cameras (background subtraction) or can not locate object position and shape accurately (frame differencing). In this paper we present a belief propagation approach for moving object detection using a 3D Markov random field (MRF) model. Each hidden state in the 3D MRF model represents a pixel's motion likelihood and is estimated using message passing in a 6-connected spatio-temporal neighborhood. This approach deals effectively with difficult moving object detection problems like objects camouflaged by similar appearance to the background, or objects with uniform color that frame difference methods can only partially detect. Three examples are presented where moving objects are detected and tracked successfully while handling appearance change, shape change, varied moving speed/direction, scale change and occlusion/clutter.
Zhaozheng Yin, Robert T. Collins
CVPR2
2007 On-the-fly Object Modeling while Tracking
abstract
To implement a persistent tracker, we build a set of view-dependent object appearance models adoptively and automatically while tracking an object under different viewing angles. This collection of acquired models is indexed with respect to the view sphere. The acquired models aid recovery from tracking failure due to occlusion and changing view angle. In this paper, view-dependent object appearance is represented by intensity patches around detected Harris corners. The intensity patches from a model are matched to the current frame by solving a bipartite linear assignment problem with outlier exclusion and missed inlier recovery. Based on these reliable matches, the change in object rotation, translation and scale is estimated between consecutive frames using Procrustes analysis. The experimental results show good performance using a collection of view-specific patch-based models for detection and tracking of vehicles in low-resolution airborne video.
Zhaozheng Yin, Robert T. Collins
CVPR2
2007 Background estimation under rapid gain change in thermal imagery
Hulya Yalcin, Robert T. Collins, Martial Hebert
Comput. Vis. Image Underst.2
2006 Spatial Divide and Conquer with Motion Cues for Tracking through Clutter
abstract
Tracking can be considered a two-class classification problem between the foreground object and its surrounding background. Feature selection to better discriminate object from background is thus a critical step to ensure tracking robustness. In this paper, a spatial divide and conquer approach is used to subdivide foreground and background into smaller regions, with different features being selected to distinguish between different pairs of object and background regions. Temporal cues are incorporated into the process using foreground motion prediction and motion segmentation. Appearance weight maps tailored to each spatial region are merged and combined with the motion information to form a joint weight image suitable for mean-shift tracking. Examples are presented to illustrate that divide and conquer feature selection combined with motion cues handles spatial background clutter and camouflage well.
Zhaozheng Yin, Robert T. Collins
CVPR (1)2
2006 Body Localization in Still Images Using Hierarchical Models and Hybrid Search
abstract
We present a 3-level hierarchical model for localizing human bodies in still images from arbitrary viewpoints. We first fit a simple tree-structured model defined on a small landmark set along the body contours by Dynamic Programming (DP). The output is a series of proposal maps that encode the probabilities of partial body configurations. Next, we fit a mixture of view-dependent models by Sequential Monte Carlo (SMC), which handles self-occlusion, anthropometric constraints, and large viewpoint changes. DP and SMC are designed to search in opposite directions such that the DP proposals are utilized effectively to initialize and guide the SMC inference. This hybrid strategy of combining deterministic and stochastic search ensures both the robustness and efficiency of DP, and the accuracy of SMC. Finally, we fit an expanded mixture model with increased landmark density through local optimization. The model hierarchy is trained on a large number of gait images. Extensive tests on cluttered images with varying poses including walking, dancing and various types of sports activities demonstrate the feasibility of the proposed approach.
Jiayong Zhang, Jiebo Luo 0001, Robert T. Collins, Yanxi Liu 0001
CVPR (2)3
2005 Unsupervised Learning of Object Features from Video Sequences
abstract
We develop an efficient algorithm for unsupervised learning of object models as constellations of features, from low resolution video sequences. The input images typically contain single or multiple objects that change in pose, scale and degree of occlusion. Also, the objects can move significantly between consecutive frames. The content of an input sequence is unlabeled so the learner has to cluster the data based on the data's implicit coherence over time and space. Our approach takes advantage of the dependent pairwise co-occurrences of objects' features within local neighborhoods vs. the independent behavior of unrelated features. We couple or decouple pairs of features based on a probabilistic interpretation of their pairwise statistics and then extract objects as connected components of features.
Marius Leordeanu, Robert T. Collins
CVPR (1)2
2005 Corrected Laplacians: Closer Cuts and Segmentation with Shape Priors
abstract
We optimize over the set of corrected Laplacians (CL) associated with a weighted graph to improve the average case normalized cut (NCut) of a graph. Unlike edge-relaxation SDPs, optimizing over the set CL naturally exploits the matrix sparsity by operating solely on the diagonal. This structure is critical to image segmentation applications because the number of vertices is generally proportional to the number of pixels in the image. CL optimization provides a guiding principle for improving the combinatorial solution over the spectral relaxation, which is important because small improvements in the cut cost often result in significant improvements in the perceptual relevance of the segmentation. We develop an optimization procedure to accommodate prior information in the form of statistical shape models, resulting in a segmentation method that produces foreground regions which are consistent with a parameterized family of shapes. We validate our technique with ground truth on MRI medical images, providing a quantitative comparison against results produced by current spectral relaxation approaches to graph partitioning.
David Tolliver, Gary L. Miller, Robert T. Collins
CVPR (2)3
2005 A Flow-Based Approach to Vehicle Detection and Background Mosaicking in Airborne Video
abstract
In this work, we address the detection of vehicles in a video stream obtained from a moving airborne platform. We propose a Bayesian framework for estimating dense optical flow over time that explicitly estimates a persistent model of background appearance. The approach assumes that the scene can be described by background and occlusion layers, estimated within an expectation-maximization framework. The mathematical formulation of the paper is an extension of the work in (H. Yalcin et al., 2005) where motion and appearance models for foreground and background layers are estimated simultaneously in a Bayesian framework.
Hulya Yalcin, Martial Hebert, Robert T. Collins, Michael J. Black
CVPR (2)3
2005 Bayesian Body Localization Using Mixture of Nonlinear Shape Models
abstract
We present a 2D model-based approach to localizing human body in images viewed from arbitrary and unknown angles. The central component is a statistical shape representation of the nonrigid and articulated body contours, where a nonlinear deformation is decomposed based on the concept of parts. Several image cues are combined to relate the body configuration to the observed image, with self occlusion explicitly treated. To accommodate large viewpoint changes, a mixture of view-dependent models is employed. Inference is done by direct sampling of the posterior mixture, using Sequential Monte Carlo (SMC) simulation enhanced with annealing and kernel move. The fitting method is independent of the number of mixture components, and does not require the preselection of a "correct" viewpoint. The models were trained on a large number of interactively labeled gait images. Preliminary tests demonstrated the feasibility of the proposed approach.
Jiayong Zhang, Robert T. Collins, Yanxi Liu 0001
ICCV2
2005 Online Selection of Discriminative Tracking Features
abstract
This paper presents an online feature selection mechanism for evaluating multiple features while tracking and adjusting the set of features used to improve tracking performance. Our hypothesis is that the features that best discriminate between object and background are also best for tracking the object. Given a set of seed features, we compute log likelihood ratios of class conditional sample densities from object and background to form a new set of candidate features tailored to the local object/background discrimination task. The two-class variance ratio is used to rank these new features according to how well they separate sample distributions of object and background pixels. This feature evaluation mechanism is embedded in a mean-shift tracking system that adaptively selects the top-ranked discriminative features for tracking. Examples are presented that demonstrate how this method adapts to changing appearances of both tracked object and scene background. We note susceptibility of the variance ratio feature selection method to distraction by spatially correlated background clutter and develop an additional approach that seeks to minimize the likelihood of distraction.
Robert T. Collins, Yanxi Liu 0001, Marius Leordeanu
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Three-Dimensional Scene Flow
abstract
Just as optical flow is the two-dimensional motion of points in an image, scene flow is the three-dimensional motion of points in the world. The fundamental difficulty with optical flow is that only the normal flow can be computed directly from the image measurements, without some form of smoothing or regularization. In this paper, we begin by showing that the same fundamental limitation applies to scene flow; however, many cameras are used to image the scene. There are then two choices when computing scene flow: 1) perform the regularization in the images or 2) perform the regularization on the surface of the object in the scene. In this paper, we choose to compute scene flow using regularization in the images. We describe three algorithms, the first two for computing scene flow from optical flows and the third for constraining scene tructure from the inconsistencies in multiple optical flows.
Sundar Vedula, Simon Baker, Peter Rander, Robert T. Collins, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.4
2004 Representation and Matching of Articulated Shapes
Jiayong Zhang, Robert T. Collins, Yanxi Liu 0001
CVPR (2)2
2004 A Computational Model for Periodic Pattern Perception Based on Frieze and Wallpaper Groups
abstract
We present a computational model for periodic pattern perception based on the mathematical theory of crystallographic groups. In each N-dimensional Euclidean space, a finite number of symmetry groups can characterize the structures of an infinite variety of periodic patterns. In 2D space, there are seven frieze groups describing monochrome patterns that repeat along one direction and 17 wallpaper groups for patterns that repeat along two linearly independent directions to tile the plane. We develop a set of computer algorithms that "understand" a given periodic pattern by automatically finding its underlying lattice, identifying its symmetry group, and extracting its representative motifs. We also extend this computational model for near-periodic patterns using geometric AIC. Applications of such a computational model include pattern indexing, texture synthesis, image compression, and gait analysis.
Yanxi Liu 0001, Robert T. Collins, Yanghai Tsin
IEEE Trans. Pattern Anal. Mach. Intell.2
2003 Mean-shift Blob Tracking through Scale Space
abstract
The mean-shift algorithm is an efficient technique for tracking 2D blobs through an image. Although the scale of the mean-shift kernel is a crucial parameter, there is presently no clean mechanism for choosing or updating scale while tracking blobs that are changing in size. We adapt Lindeberg's (1998) theory of feature scale selection based on local maxima of differential scale-space filters to the problem of selecting kernel scale for mean-shift blob tracking. We show that a difference of Gaussian (DOG) mean-shift kernel enables efficient tracking of blobs through scale space. Using this kernel requires generalizing the mean-shift algorithm to handle images that contain negative sample weights.
Robert T. Collins
CVPR (2)1
2003 On-Line Selection of Discriminative Tracking Features
abstract
We present a method for evaluating multiple feature spaces while tracking, and for adjusting the set of features used to improve tracking performance. Our hypothesis is that the features that best discriminate between object and background are also best for tracking the object. We develop an online feature selection mechanism based on the two-class variance ratio measure, applied to log likelihood distributions computed with respect to a given feature from samples of object and background pixels. This feature selection mechanism is embedded in a tracking system that adaptively selects the top-ranked discriminative features for tracking. Examples are presented to illustrate how the method adapts to changing appearances of both tracked object and scene background.
Robert T. Collins, Yanxi Liu 0001
ICCV1
2002 Gait Sequence Analysis Using Frieze Patterns
Yanxi Liu 0001, Robert T. Collins, Yanghai Tsin
ECCV (2)2
2002 An active camera system for acquiring multi-view video
abstract
A system is described for acquiring multi-view video of a person moving through the environment. A real-time tracking algorithm adjusts the pan, tilt, zoom and focus parameters of multiple active cameras to keep the moving person centered in each view. The output of the system is a set of synchronized, time-stamped video streams, showing the person simultaneously from several viewpoints.
Robert T. Collins, Omead Amidi, Takeo Kanade
ICIP (1)1
2002 Three-Dimensional Reconstruction of Points and Lines with Unknown Correspondence
Yong-Qing Cheng, Robert T. Collins, Xiaoguang Wang 0013, Edward M. Riseman, Allen R. Hanson
Int. J. Comput. Vis.2
2001 Skewed Symmetry Groups
abstract
We introduce the term skewed symmetry groups and provide a complete theoretical treatment for 2D wallpaper groups under affine transformations. For the first time, a given periodic pattern can be classified not simply by its Euclidean symmetry group but,by its highest "potential" symmetry group under affine deformation. A concise wallpaper group migration map is constructed that separates the 17 affinely deformed wallpaper groups into small, distinct orbits. The practical value of this result includes a novel indexing and retrieval scheme for regular patterns, and a maximal-symmetry-based method for estimating shape and orientation from texture under unknown views.
Yanxi Liu 0001, Robert T. Collins
CVPR (1)2
2001 Bayesian Color Constancy for Outdoor Object Recognition
abstract
Outdoor scene classification is challenging due to irregular geometry, uncontrolled illumination, and noisy reflectance distributions. This paper discusses a Bayesian approach to classifying a color image of an outdoor scene. A likelihood model factors in the physics of the image formation process, sensor noise distribution, and prior distributions over geometry, material types, and illuminant spectrum parameters. These prior distributions are learned through a training process that uses color observations of planar scene patches over time. An iterative linear algorithm estimates the maximum likelihood reflectance, spectrum, geometry, and object class labels for a new image. Experiments on images taken by outdoor surveillance cameras classify known material types and shadow regions correctly, and flag as outliers material types that were not seen previously.
Yanghai Tsin, Robert T. Collins, Visvanathan Ramesh, Takeo Kanade
CVPR (1)2
2001 Three-Dimensional Reconstruction of Points and Lines with Unknown Correspondence across Images
Yong-Qing Cheng, Xiaoguang Wang 0013, Robert T. Collins, Edward M. Riseman, Allen R. Hanson
Int. J. Comput. Vis.3
2001 Algorithms for cooperative multisensor surveillance
abstract
The Video Surveillance and Monitoring (VSAM) team at Carnegie Mellon University (CMU) has developed an end-to-end, multicamera surveillance system that allows a single human operator to monitor activities in a cluttered environment using a distributed network of active video sensors. Video understanding algorithms have been developed to automatically detect people and vehicles, seamlessly track them using a network of cooperating active sensors, determine their three-dimensional locations with respect to a geospatial site model, and present this information to a human operator who controls the system through a graphical user interface. The goal is to automatically collect and disseminate real-time information to improve the situational awareness of security providers and decision makers. The feasibility of real-time video surveillance has been demonstrated within a multicamera testbed system developed on the campus of CMU. This paper presents an overview of the issues and algorithms involved in creating this semiautonomous, multicamera surveillance system.
Robert T. Collins, Alan J. Lipton, Hironobu Fujiyoshi, Takeo Kanade
Proc. IEEE1
2001 Robust Midsagittal Plane Extraction from Normal and Pathological 3D Neuroradiology Images
abstract
This paper focuses on extracting the ideal midsagittal plane (iMSP) from three-dimensional (3-D) normal and pathological neuroimages. The main challenges in this work are the structural asymmetry that may exist in pathological brains, and the anisotropic, unevenly sampled image data that is common in clinical practice. We present an edge-based, cross-correlation approach that decomposes the plane fitting problem into discovery of two-dimensional symmetry axes on each slice, followed by a robust estimation of plane parameters. The algorithm's tolerance to brain asymmetries, input image offsets and image noise is quantitatively evaluated. We find that the algorithm can extract the iMSP from input 3-D images with 1) large asymmetrical lesions; 2) arbitrary initial rotation offsets; 3) low signal-to-noise ratio or high bias field. The iMSP algorithm is compared with an approach based on maximization of mutual information registration, and is found to exhibit superior performance under adverse conditions. Finally, no statistically significant difference is found between the midsagittal plane computed by the iMSP algorithm and that estimated by two trained neuroradiologists.
Yanxi Liu 0001, Robert T. Collins, William E. Rothfus
IEEE Trans. Medical Imaging2
2000 A Computational Model for Repeated Pattern Perception Using Frieze and Wallpaper Groups
abstract
Humans have an innate ability to perceive symmetry, but it is not obvious how to automate this powerful insight. In this paper the mathematical theory of frieze and wallpaper groups is used to extract visually meaningful building blocks (motifs) from a repeated pattern. A novel peak detection algorithm based on "regions of dominance" is used to automatically detect the underlying translational lattice of a repeated pattern. Following automatic classification of the pattern's symmetry group, knowledge of the interplay between rotation, reflection, glide-reflection and translation in that group leads to a small set of candidate motifs that exhibit local symmetry consistent with the global symmetry of the entire pattern. Although other work has addressed detection of the translational lattice of a repeated pattern, ours is the first to seek a principled method for determining a representative motif. Experiments show that the-resulting pattern motifs conform well with human perception.
Yanxi Liu 0001, Robert T. Collins
CVPR2
2000 Robust Midsagittal Plane Extraction from Coarse, Pathological 3D Images
Yanxi Liu 0001, Robert T. Collins, William E. Rothfus
MICCAI2
2000 Introduction to the Special Section on Video Surveillance
abstract
UTOMATED video surveillance addresses real-time observation of people and vehicles within a busy environment, leading to a description of their actions and interactions. The technical issues include moving object detection and tracking, object classification, human motion analysis, and activity understanding, touching on many of the core topics of computer vision, pattern analysis, and aritificial intelligence. Video surveillance has spawned large research projects in the United States, Europe, and Japan, and has been the topic of several international conferences and workshops in recent years. There are immediate needs for automated surveillance systems in commercial, law enforcement, and military applications. Mounting video cameras is cheap, but finding available human resources to observe the output is expensive. Although surveillance cameras are already prevalent in banks, stores, and parking lots, video data currently is used only “after the fact” as a forensic tool, thus losing its primary benefit as an active, real-time medium. What is needed is continuous 24-hour monitoring of surveillance video to alert security officers to a burglary in progress or to a suspicious individual loitering in the parking lot, while there is still time to prevent the crime. In addition to the obvious security applications, video surveillance technology has been proposed to measure traffic flow, detect accidents on highways, monitor pedestrian congestion in public spaces, compile consumer demographics in shopping malls and amusement parks, log routine maintainence tasks at nuclear facilities, and count endangered species. The numerous military applications include patrolling national borders, measuring the flow of refugees in troubled areas, monitoring peace treaties, and providing secure perimeters around bases and embassies. The 11 papers in this special section illustrate topics and techniques at the forefront of video surveillance research. These papers can be loosely organized into three categories. Detection and tracking involves real-time extraction of moving objects from video and continuous tracking over time to form persistent object trajectories. C. Stauffer and W.E.L. Grimson introduce unsupervised statistical learning techniques to cluster object trajectories produced by adaptive background subtraction into descriptions of normal scene activity. Viewpoint-specific trajectory descriptions from multiple cameras are combined into a common scene coordinate system using a calibration technique described by L. Lee, R. Romano, and G. Stein, who automatically determine the relative exterior orientation of overlapping camera views by observing a sparse set of moving objects on flat terrain. Two papers address the accumulation of noisy motion evidence over time. R. Pless, T. Brodský, and Y. Aloimonos detect and track small objects in aerial video sequences by first compensating for the self-motion of the aircraft, then accumulating residual normal flow to acquire evidence of independent object motion. L. Wixson notes that motion in the image does not always signify purposeful travel by an independently moving object (examples of such “motion clutter” are wind-blown tree branches and sun reflections off rippling water) and devises a flow-based salience measure to highlight objects that tend to move in a consistent direction over time. Human motion analysis is concerned with detecting periodic motion signifying a human gait and acquiring descriptions of human body pose over time. R. Cutler and L.S. Davis plot an object’s self-similarity across all pairs of frames to form distinctive patterns that classify bipedal, quadripedal, and rigid object motion. Y. Ricquebourg and P. Bouthemy track apparent contours in XT slices of an XYT sequence volume to robustly delineate and track articulated human body structure. I. Haritaoglu, D. Harwood, and L.S. Davis present W4, a surveillance system specialized to the task of looking at people. The W4 system can locate people and segment their body parts, build simple appearance models for tracking, disambiguate between and separately track multiple individuals in a group, and detect carried objects such as boxes and backpacks. Activity analysis deals with parsing temporal sequences of object observations to produce high-level descriptions of agent actions and multiagent interactions. In our opinion, this will be the most important area of future research in video surveillance. N.M. Oliver, B. Rosario, and A.P. Pentland introduce Coupled Hidden Markov Models (CHMMs) to detect and classify interactions consisting of two interleaved agent action streams and present a training method based on synthetic agents to address the problem of parameter estimation from limited real-world training examples. M. Brand and V. Kettnaker present an entropyminimization approach to estimating HMM topology and
Robert T. Collins, Christophe Biernacki, Gilles Celeux, Alan J. Lipton, Gérard Govaert, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Calibration of an Outdoor Active Camera System
abstract
A parametric camera model and calibration procedures are developed for an outdoor active camera system with pan, tilt and zoom control. Unlike traditional methods, active camera motion plays a key role in the calibration process, and no special laboratory setups are required. Intrinsic parameters are estimated automatically by fitting parametric models to the optic flow induced by rotating and zooming. No knowledge of 3D scene structure is needed. Extrinsic parameters are calculated by actively rotating the camera to sight a sparse set of surveyed landmarks over a virtual hemispherical field of view, yielding a well-conditioned pose estimation problem.
Robert T. Collins, Yanghai Tsin
CVPR1
1999 Three-Dimensional Scene Flow
abstract
Scene flow is the three-dimensional motion field of points in the world, just as optical flow is the two-dimensional motion field of points in an image. Any optical flow is simply the projection of the scene flow onto the image plane of a camera. We present a framework for the computation of dense, non-rigid scene flow from optical flow. Our approach leads to straightforward linear algorithms and a classification of the task into three major scenarios: complete instantaneous knowledge of the scene structure; knowledge only of correspondence information; and no knowledge of the scene structure. We also show that multiple estimates of the normal flow cannot be used to estimate dense scene flow directly without some form of smoothing or regularization.
Sundar Vedula, Simon Baker, Peter Rander, Robert T. Collins, Takeo Kanade
ICCV4
1998 The Ascender System: Automated Site Modeling from Multiple Aerial Images
Robert T. Collins, Christopher O. Jaynes, Yong-Qing Cheng, Xiaoguang Wang 0013, Frank Stolle, Edward M. Riseman, Allen R. Hanson
Comput. Vis. Image Underst.1
1997 Multi-Image Focus of Attention for Rapid Site Model Construction
abstract
A multi-image focus of attention mechanism has been developed that can quickly distinguish raised objects like buildings from structured background clutter typical to many aerial image scenarios. The underlying approach is the space-sweep stereo method, in which features from multiple images are backprojected onto a virtual, horizontal plane that is methodically swept through the scene. Back-projected gradient orientations from multiple images are highly correlated when they come from scene locations containing structural edges that are roughly horizontal, like building roofs and terrain; otherwise, they tend to be uniformly distributed. These observations are used to define a structural salience measure that can determine whether a given volume of space contains a statistically significant number of structural edges, without first performing precise reconstruction of those edges. The utility of structural salience for computing focus of attention regions is illustrated on sample data from Ft. Hood, Texas.
Robert T. Collins
CVPR1
1997 Reply to Pizlo, Rosenfeld, and Weiss
Robert T. Collins
Comput. Vis. Image Underst.1
1996 A Space-Sweep Approach to True Multi-Image Matching
abstract
The problem of determining feature correspondences across multiple views is considered. The term "true multi-image" matching is introduced to describe techniques that make full and efficient use of the geometric relationships between multiple images and the scene. A true multi-image technique must generalize to any number of images, be of linear algorithmic complexity in the number of images, and use all the images in an equal manner. A new space-sweep approach to true multi-image matching is presented that simultaneously determines 2D feature correspondences and the 3D positions of feature points in the scene. The method is illustrated on a seven-image matching example from the aerial image domain.
Robert T. Collins
CVPR1
1996 Determining Correspondences and Rigid Motion of 3-D Point Sets with Missing Data
abstract
This paper addresses the general 3-D rigid motion problem, where the point correspondences and the motion parameters between two sets of 3-D points are to be recovered. The existence of missing points in the two sets is the most difficult problem. We first show a mathematical symmetry in the solutions of rotation parameters and point correspondences. A closed-form solution based on the correlation matrix eigenstructure decomposition is proposed for correspondence recovery with no missing points. Using a heuristic measure of point pair affinity derived from the eigenstructure, a weighted bipartite matching algorithm is developed to determine the correspondences in general cases where missing points occur. The use of the affinity heuristic also leads to a fast outlier removal algorithm, which can be run iteratively to refine the correspondence recovery. Simulation results and experiments on real images are shown in both ideal and general cases.
Xiaoguang Wang 0013, Yong-Qing Cheng, Robert T. Collins, Allen R. Hanson
CVPR3
1995 Site Model Acquisition and Extension from Aerial Images
abstract
A system has been developed to acquire, extend and refine 3D geometric site models from aerial imagery. This system hypothesizes potential building roofs in an image, automatically locates supporting geometric evidence in other images, and determines the precise shape and position of the new buildings via multiimage triangulation. Model-to-image registration techniques are applied to align new, incoming images against the site model. Model extension and refinement procedures are then performed to add previously unseen buildings and to improve the geometric accuracy of the existing 3D building models.>
Robert T. Collins, Yong-Qing Cheng, Christopher O. Jaynes, Frank Stolle, Xiaoguang Wang 0013, Allen R. Hanson, Edward M. Riseman
ICCV1
1994 Task driven perceptual organization for extraction of rooftop polygons
abstract
A new method for extracting planar polygonal rooftops in monocular aerial imagery is proposed. Structural features are extracted and hierarchically related using perceptual grouping techniques. Top-down feature verification is used so that features, and links between the features, are verified with local information in the image and weighed in a graph. Cycles in the graph correspond to possible building rooftop hypotheses. Virtual features are hypothesized for the perceptual completion of partially occluded rooftops. Extraction of the "best" grouping features into a building rooftop hypothesis is posed as a graph search problem. The maximally weighted, independent set of cycles in the graph is extracted as the final set of roof boundaries.>
Christopher O. Jaynes, Frank Stolle, Robert T. Collins
WACV3
1993 Matching perspective views of coplanar structures using projective unwarping and similarity matching
abstract
We consider the problem of matching perspective views of coplanar structures composed of line segments. Both model-to-image and image-to-image correspondence matching are given a consistent treatment. These matching scenarios generally require discovery of an eight parameter projective mapping. However, when the horizon line of the object plane can be found in the image, done here using vanishing point analysis, these problems reduce to a simpler six parameter affine matching problem. When the intrinsic lens parameters of the camera are known, the problem further reduces to four parameter affine similarity matching.
Robert T. Collins, J. Ross Beveridge
CVPR1
1992 Single plane model extension using projective transformations and data fusion
abstract
A priori knowledge of the relative positions of four or more coplanar points or lines is used to derive the positions of other points and lines on the same plane in a manner invariant to camera location and intrinsic camera parameters. A framework for data fusion in the projective plane is presented to merge the position estimates of coplanar points and lines derived in this way.>
Robert T. Collins
CVPR1
1990 Vanishing point calculation as a statistical inference on the unit sphere
abstract
An examination is made of vanishing point calculation as a statistical estimation problem. It is assumed that image line segments have been previously clustered into groups of convergent lines. For each group, the vanishing point location is estimated as the polar axis of an equatorial distribution on the unit sphere, and the statistical error of the estimate is determined. The sensitivity of the estimates to the number of lines in a convergent cluster is studied.>
Robert T. Collins, Richard Weiss 0001
ICCV1
1989 The schema system
Bruce A. Draper, Robert T. Collins, John Brolio, Allen R. Hanson, Edward M. Riseman
Int. J. Comput. Vis.2
1988 Image interpretation by distributed cooperative processes
abstract
The VISIONS schema system provides a framework for building a general interpretation system as a distributed network of many small special-purpose interpretation systems. Each scheme is an 'expert' at recognizing one type of object. A discussion is presented of the problems of knowledge representation in a distributed A1 environments and the scheme system approach to those problems. A series of interpretation experiments have been performed on nine images from two natural scene domains; results from three images are presented.>
Bruce A. Draper, John Brolio, Robert T. Collins, Allen R. Hanson, Edward M. Riseman
CVPR3