Sven J. Dickinson

dblp:d/SvenJDickinson · DBLP profile ↗
← Back
90ranked-venue papers
21as first author
9since 2021 · last 2025
0000-0001-6566-9635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 20 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Probabilistic Directed Distance Fields for Ray-Based Shape Representations
abstract
In modern computer vision, the optimal representation of 3D shape remains task-dependent. One fundamental operation applied to such representations is differentiable rendering, which enables learning-based inverse graphics approaches. Standard explicit representations are often easily rendered, but can suffer from limited geometric fidelity, among other issues. On the other hand, implicit representations generally preserve greater fidelity, but suffer from difficulties with rendering, limiting scalability. In this work, we devise Directed Distance Fields (DDFs), which map a ray or oriented point (position and direction) to surface visibility and depth. This enables efficient differentiable rendering, obtaining depth with a single forward pass per pixel, as well as higher-order geometry with only additional backward passes. Using probabilistic DDFs (PDDFs), we can model the inherent discontinuities in the underlying field. We then apply DDFs to single-shape fitting, generative modelling, and 3D reconstruction, showcasing strong performance with simple architectural components via the versatility of our representation. Finally, since the dimensionality of DDFs permits view-dependent geometric artifacts, we conduct a theoretical investigation of the constraints necessary for view consistency. We find a small set of field properties that are sufficient to guarantee a DDF is consistent, without knowing which shape the field is expressing.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Shape-Based Measures Improve Scene Categorization
abstract
Converging evidence indicates that deep neural network models that are trained on large datasets are biased toward color and texture information. Humans, on the other hand, can easily recognize objects and scenes from images as well as from bounding contours. Mid-level vision is characterized by the recombination and organization of simple primary features into more complex ones by a set of so-called Gestalt grouping rules. While described qualitatively in the human literature, a computational implementation of these perceptual grouping rules is so far missing. In this article, we contribute a novel set of algorithms for the detection of contour-based cues in complex scenes. We use the medial axis transform (MAT) to locally score contours according to these grouping rules. We demonstrate the benefit of these cues for scene categorization in two ways: (i) Both human observers and CNN models categorize scenes most accurately when perceptual grouping information is emphasized. (ii) Weighting the contours with these measures boosts performance of a CNN model significantly compared to the use of unweighted contours. Our work suggests that, even though these measures are computed directly from contours in the image, current CNN models do not appear to extract or utilize these grouping cues.
Morteza Rezanejad, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable Simulation
abstract
Multi-finger grasping relies on high quality training data, which is hard to obtain: human data is hard to transfer and synthetic data relies on simplifying assumptions that reduce grasp quality. By making grasp simulation differentiable, and contact dynamics amenable to gradient-based optimization, we accelerate the search for high-quality grasps with fewer limiting assumptions. We present Grasp'D-1M: a large-scale dataset for multi-finger robotic grasping, synthesized with Fast-Grasp'D, a novel differentiable grasping simulator. Grasp'D-1M contains one million training examples for three robotic hands (three, four and five-fingered), each with multimodal visual inputs (RGB+depth+segmentation, available in mono and stereo). Grasp synthesis with Fast-Grasp'D is 10x faster than GraspIt! [1] and 20x faster than the prior Grasp'D differentiable simulator [2]. Generated grasps are more stable and contact-rich than GraspIt! grasps, regardless of the distance threshold used for contact generation. We validate the usefulness of our dataset by retraining an existing vision-based grasping pipeline [3] on Grasp'D-1M, and showing a dramatic increase in model performance, predicting grasps with 30% more contact, a 33% higher epsilon metric, and 35% lower simulated displacement. Additional details at fast-graspd.github.io.
Dylan Turpin, Tao Zhong 0003, Shutong Zhang, Guanglei Zhu, Eric Heiden, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg
ICRA8
2023 Disentangling Geometric Deformation Spaces in Generative Latent Shape Models
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
Int. J. Comput. Vis.3
2022 Representing 3D Shapes with Probabilistic Directed Distance Fields
abstract
Differentiable rendering is an essential operation in modern vision, allowing inverse graphics approaches to 3D understanding to be utilized in modern machine learning frameworks. Explicit shape representations (voxels, point clouds, or meshes), while relatively easily rendered, often suffer from limited geometric fidelity or topological con-straints. On the other hand, implicit representations (occu-pancy, distance, or radiance fields) preserve greater fidelity, but suffer from complex or inefficient rendering processes, limiting scalability. In this work, we endeavour to address both shortcomings with a novel shape representation that allows fast differentiable rendering within an implicit ar-chitecture. Building on implicit distance representations, we define Directed Distance Fields (DDFs), which map an oriented point (position and direction) to surface visibility and depth. Such a field can render a depth map with a single forward pass per pixel, enable differential surface geometry extraction (e.g., surface normals and curvatures) via network derivatives, be easily composed, and permit extraction of classical unsigned distance fields. Using probabilistic DDFs (PDDFs), we show how to model inherent discontinuities in the underlying field. Finally, we apply our method to fitting single shapes, unpaired 3D-aware generative image modelling, and single-image 3D reconstruction tasks, showcasing strong performance with simple architectural components via the versatility of our representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson
CVPR3
2022 Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands
Dylan Turpin, Liquan Wang, Eric Heiden, Yun-Chun Chen, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg
ECCV (6)7
2022 State of the Journal Editorial
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 DeepFlux for Skeleton Detection in the Wild
Yongchao Xu, Yukang Wang, Stavros Tsogkas, Jianqiang Wan, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
Int. J. Comput. Vis.6
2021 State of the Journal Editorial
abstract
Presents the state of the journal review for this issue of the publication.
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 Appearance Shock Grammar for Fast Medial Axis Extraction From Real Images
abstract
We combine ideas from shock graph theory with more recent appearance-based methods for medial axis extraction from complex natural scenes, improving upon the present best unsupervised method, in terms of efficiency and performance. We make the following specific contributions: i) we extend the shock graph representation to the domain of real images, by generalizing the shock type definitions using local, appearance-based criteria; ii) we then use the rules of a Shock Grammar to guide our search for medial points, drastically reducing run time when compared to other methods, which exhaustively consider all points in the input image; iii) we remove the need for typical post-processing steps including thinning, non-maximum suppression, and grouping, by adhering to the Shock Grammar rules while deriving the medial axis solution; iv) finally, we raise some fundamental concerns with the evaluation scheme used in previous work and propose a more appropriate alternative for assessing the performance of medial axis extraction from scenes. Our experiments on the BMAX500 and SK-LARGE datasets demonstrate the effectiveness of our approach. We outperform the present state-of-the-art, excelling particularly in the high-precision regime, while running an order of magnitude faster and requiring no post-processing.
Charles-Olivier Dufresne Camaro, Morteza Rezanejad, Stavros Tsogkas, Kaleem Siddiqi, Sven J. Dickinson
CVPR5
2020 State of the Journal Editorial
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 Scene Categorization From Contours: Medial Axis Based Salience Measures
Morteza Rezanejad, Gabriel Downs, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi
CVPR6
2019 DeepFlux for Skeletons in the Wild
abstract
Computing object skeletons in natural images is challenging, owing to large variations in object appearance and scale, and the complexity of handling background clutter. Many recent methods frame object skeleton detection as a binary pixel classification problem, which is similar in spirit to learning-based edge detection, as well as to semantic segmentation methods. In the present article, we depart from this strategy by training a CNN to predict a two-dimensional vector field, which maps each scene point to a candidate skeleton pixel, in the spirit of flux-based skeletonization algorithms. This ``image context flux'' representation has two major advantages over previous approaches. First, it explicitly encodes the relative position of skeletal pixels to semantically meaningful entities, such as the image points in their spatial context, and hence also the implied object boundaries. Second, since the skeleton detection context is a region-based vector field, it is better able to cope with object parts of large width. We evaluate the proposed method on three benchmark datasets for skeleton detection and two for symmetry detection, achieving consistently superior performance over state-of-the-art methods.
Yukang Wang, Yongchao Xu, Stavros Tsogkas, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi
CVPR5
2019 Geometric Disentanglement for Generative Latent Shape Models
abstract
Representing 3D shapes is a fundamental problem in artificial intelligence, which has numerous applications within computer vision and graphics. One avenue that has recently begun to be explored is the use of latent representations of generative models. However, it remains an open problem to learn a generative model of shapes that is interpretable and easily manipulated, particularly in the absence of supervised labels. In this paper, we propose an unsupervised approach to partitioning the latent space of a variational autoencoder for 3D point clouds in a natural way, using only geometric information, that builds upon prior work utilizing generative adversarial models of point sets. Our method makes use of tools from spectral geometry to separate intrinsic and extrinsic shape information, and then considers several hierarchical disentanglement penalties for dividing the latent space in this manner. We also propose a novel disentanglement penalty that penalizes the predicted change in the latent representation of the output,with respect to the latent variables of the initial shape. We show that the resulting latent representation exhibits intuitive and interpretable behaviour, enabling tasks such as pose transfer that cannot easily be performed by models with an entangled representation.
Tristan Aumentado-Armstrong, Stavros Tsogkas, Allan Douglas Jepson, Sven J. Dickinson
ICCV4
2019 State of the Journal
abstract
Presents the current state of the journal and discusses future directions.
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 State of the Journal
abstract
Presents an editorial on the current state of the IEEE Transactions on Pattern Analysis and Machine Intelligence.
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 AMAT: Medial Axis Transform for Natural Images
abstract
We introduce Appearance-MAT (AMAT), a generalization of the medial axis transform for natural images, that is framed as a weighted geometric set cover problem. We make the following contributions: i) we extend previous medial point detection methods for color images, by associating each medial point with a local scale; ii) inspired by the invertibility property of the binary MAT, we also associate each medial point with a local encoding that allows us to invert the AMAT, reconstructing the input image; iii) we describe a clustering scheme that takes advantage of the additional scale and appearance information to group individual points into medial branches, providing a shape decomposition of the underlying image regions. In our experiments, we show state-of-the-art performance in medial point detection on Berkeley Medial AXes (BMAX500), a new dataset of medial axes based on the BSDS500 database, and good generalization on the SK506 and WH-SYMMAX datasets. We also measure the quality of reconstructed images from BMAX500, obtained by inverting their computed AMAT. Our approach delivers significantly better reconstruction quality w.r.t. to three baselines, using just 10% of the image pixels. Our code and annotations are available at https://github.com/tsogkas/amat.
Stavros Tsogkas, Sven J. Dickinson
ICCV2
2017 Incoming EIC Editorial
abstract
Presents the incoming editorial by the new Editor-In-Chief.
Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Learning to Combine Mid-Level Cues for Object Proposal Generation
abstract
In recent years, region proposals have replaced sliding windows in support of object recognition, offering more discriminating shape and appearance information through improved localization. One powerful approach for generating region proposals is based on minimizing parametric energy functions with parametric maxflow. In this paper, we introduce Parametric Min-Loss (PML), a novel structured learning framework for parametric energy functions. While PML is generally applicable to different domains, we use it in the context of region proposals to learn to combine a set of mid-level grouping cues to yield a small set of object region proposals with high recall. Our learning framework accounts for multiple diverse outputs, and is complemented by diversification seeds based on image location and color. This approach casts perceptual grouping and cue combination in a novel structured learning framework which yields baseline improvements on VOC 2012 and COCO 2014.
Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson
ICCV3
2014 Multi-cue Mid-level Grouping
Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson
ACCV (3)3
2013 Recognize Human Activities from Partially Observed Videos
abstract
Recognizing human activities in partially observed videos is a challenging problem and has many practical applications. When the unobserved subsequence is at the end of the video, the problem is reduced to activity prediction from unfinished activity streaming, which has been studied by many researchers. However, in the general case, an unobserved subsequence may occur at any time by yielding a temporal gap in the video. In this paper, we propose a new method that can recognize human activities from partially observed videos in the general case. Specifically, we formulate the problem into a probabilistic framework: 1) dividing each activity into multiple ordered temporal segments, 2) using spatiotemporal features of the training video samples in each segment as bases and applying sparse coding (SC) to derive the activity likelihood of the test video sample at each segment, and 3) finally combining the likelihood at each segment to achieve a global posterior for the activities. We further extend the proposed method to include more bases that correspond to a mixture of segments with different temporal lengths (MSSC), which can better represent the activities with large intra-class variations. We evaluate the proposed methods (SC and MSSC) on various real videos. We also evaluate the proposed methods on two special cases: 1) activity prediction where the unobserved subsequence is at the end of the video, and 2) human activity recognition on fully observed videos. Experimental results show that the proposed methods outperform existing state-of-the-art comparison methods.
Yu Cao 0003, Daniel Paul Barrett, Andrei Barbu, N. Siddharth 0001, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002
CVPR8
2013 Detecting Curved Symmetric Parts Using a Deformable Disc Model
abstract
Symmetry is a powerful shape regularity that's been exploited by perceptual grouping researchers in both human and computer vision to recover part structure from an image without a priori knowledge of scene content. Drawing on the concept of a medial axis, defined as the locus of centers of maximal inscribed discs that sweep out a symmetric part, we model part recovery as the search for a sequence of deformable maximal inscribed disc hypotheses generated from a multiscale super pixel segmentation, a framework proposed by LEV09. However, we learn affinities between adjacent super pixels in a space that's invariant to bending and tapering along the symmetry axis, enabling us to capture a wider class of symmetric parts. Moreover, we introduce a global cost that perceptually integrates the hypothesis space by combining a pair wise and a higher-level smoothing term, which we minimize globally using dynamic programming. The new framework is demonstrated on two datasets, and is shown to significantly outperform the baseline LEV09.
Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson
ICCV3
2013 Multiscale Symmetric Part Detection and Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson
Int. J. Comput. Vis.3
2012 Superedge grouping for object localization by combining appearance and shape information
abstract
Both appearance and shape play important roles in object localization and object detection. In this paper, we propose a new superedge grouping method for object localization by incorporating both boundary shape and appearance information of objects. Compared with the previous edge grouping methods, the proposed method does not subdivide detected edges into short edgels before grouping. Such long, unsubdivided superedges not only facilitate the incorporation of object shape information into localization, but also increase the robustness against image noise and reduce computation. We identify and address several important problems in achieving the proposed superedge grouping, including gap filling for connecting superedges, accurate encoding of region-based information into individual edges, and the incorporation of object-shape information into object localization. In this paper, we use the bag of visual words technique to quantify the region-based appearance features of the object of interest. We find that the proposed method, by integrating both boundary and region information, can produce better localization performance than previous subwindow search and edge grouping methods on most of the 20 object categories from the VOC 2007 database. Experiments also show that the proposed method is roughly 50 times faster than the previous edge grouping method.
Sanja Fidler, Jarrell W. Waggoner, Yu Cao 0003, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002
CVPR5
2012 Detecting Reduplication in Videos of American Sign Language
Zoya Gavrilov, Stan Sclaroff, Carol Neidle, Sven J. Dickinson
LREC4
2012 3D Object Detection and Viewpoint Estimation with a Deformable 3D Cuboid Model
abstract
This paper addresses the problem of category-level 3D object detection. Given a monocular image, our aim is to localize the objects in 3D by enclosing them with tight oriented 3D bounding boxes. We propose a novel approach that extends the well-acclaimed deformable part-based model[Felz.] to reason in 3D. Our model represents an object class as a deformable 3D cuboid composed of faces and parts, which are both allowed to deform with respect to their anchors on the 3D box. We model the appearance of each face in fronto-parallel coordinates, thus effectively factoring out the appearance variation induced by viewpoint. Our model reasons about face visibility patters called aspects. We train the cuboid model jointly and discriminatively and share weights across all aspects to attain efficiency. Inference then entails sliding and rotating the box in 3D and scoring object hypotheses. While for inference we discretize the search space, the variables are continuous in our model. We demonstrate the effectiveness of our approach in indoor and outdoor scenarios, and show that our approach outperforms the state-of-the-art in both 2D[Felz09] and 3D object detection[Hedau12].
Sanja Fidler, Sven J. Dickinson, Raquel Urtasun
NIPS2
2012 Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven J. Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, N. Siddharth 0001, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell W. Waggoner, Song Wang 0002, Jinlian Wei
UAI5
2012 Discovering hierarchical object models from captioned images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson
Comput. Vis. Image Underst.5
2012 Optimal Image and Video Closure by Superpixel Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson
Int. J. Comput. Vis.3
2011 Efficient many-to-many feature matching under the l1 norm
M. Fatih Demirci, Yusuf Osmanlioglu, Ali Shokoufandeh, Sven J. Dickinson
Comput. Vis. Image Underst.4
2011 Bone graphs: Medial shape parsing and abstraction
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi
Comput. Vis. Image Underst.2
2011 Object categorization using bone graphs
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi
Comput. Vis. Image Underst.2
2010 Spatiotemporal Closure
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson
ACCV (1)3
2010 Spatiotemporal Contour Grouping Using Abstract Part Models
Pablo Sala, Diego Macrini, Sven J. Dickinson
ACCV (4)3
2010 Discovering Multipart Appearance Models from Captioned Images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson
ECCV (5)5
2010 Optimal Contour Closure by Superpixel Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson
ECCV (2)3
2010 Contour Grouping and Abstraction Using Simple Part Models
Pablo Sala, Sven J. Dickinson
ECCV (5)2
2010 Using Language to Learn Structured Appearance Models for Image Annotation
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to simultaneously learn the names and appearances of the objects. Only a small fraction of local features within any given image are associated with a particular caption word, and captions may contain irrelevant words not associated with any image object. We propose a novel algorithm that uses the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to learn meaningful feature configurations (representing named objects). We also introduce a graph-based appearance model that captures some of the structure of an object by encoding the spatial relationships among the local visual features. In an iterative procedure, we use language (the words) to drive a perceptual grouping process that assembles an appearance model for a named object. Results of applying our method to three data sets in a variety of conditions demonstrate that, from complex, cluttered, real-world scenes with noisy captions, we can learn both the names and appearances of objects, resulting in a set of models invariant to translation, scale, orientation, occlusion, and minor changes in viewpoint or articulation. These named models, in turn, are used to automatically annotate new, uncaptioned images, thereby facilitating keyword-based image retrieval.
Michael Jamieson, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson, Sven Wachsmuth
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 Multiscale symmetric part detection and grouping
abstract
Skeletonization algorithms typically decompose an object's silhouette into a set of symmetric parts, offering a powerful representation for shape categorization. However, having access to an object's silhouette assumes correct figure-ground segmentation, leading to a disconnect with the mainstream categorization community, which attempts to recognize objects from cluttered images. In this paper, we present a novel approach to recovering and grouping the symmetric parts of an object from a cluttered scene. We begin by using a multiresolution superpixel segmentation to generate medial point hypotheses, and use a learned affinity function to perceptually group nearby medial points likely to belong to the same medial branch. In the next stage, we learn higher granularity affinity functions to group the resulting medial branches likely to belong to the same object. The resulting framework yields a skeletal approximation that's free of many of the instabilities plaguing traditional skeletons. More importantly, it doesn't require a closed contour, enabling the application of skeleton-based categorization systems to more realistic imagery
Alex Levinshtein, Sven J. Dickinson, Cristian Sminchisescu
ICCV2
2009 Skeletal Shape Abstraction from Examples
abstract
Learning a class prototype from a set of exemplars is an important challenge facing researchers in object categorization. Although the problem is receiving growing interest, most approaches assume a one-to-one correspondence among local features, restricting their ability to learn true abstractions of a shape. In this paper, we present a new technique for learning an abstract shape prototype from a set of exemplars whose features are in many-to-many correspondence. Focusing on the domain of 2D shape, we represent a silhouette as a medial axis graph whose nodes correspond to "parts" defined by medial branches and whose edges connect adjacent parts. Given a pair of medial axis graphs, we establish a many-to-many correspondence between their nodes to find correspondences among articulating parts. Based on these correspondences, we recover the abstracted medial axis graph along with the positional and radial attributes associated with its nodes. We evaluate the abstracted prototypes in the context of a recognition task.
M. Fatih Demirci, Ali Shokoufandeh, Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 TurboPixels: Fast Superpixels Using Geometric Flows
abstract
We describe a geometric-flow-based algorithm for computing a dense oversegmentation of an image, often referred to as superpixels. It produces segments that, on one hand, respect local image boundaries, while, on the other hand, limiting undersegmentation through a compactness constraint. It is very fast, with complexity that is approximately linear in image size, and can be applied to megapixel sized images with high superpixel densities in a matter of minutes. We show qualitative demonstrations of high-quality results on several complex images. The Berkeley database is used to quantitatively compare its performance to a number of oversegmentation algorithms, showing that it yields less undersegmentation than algorithms that lack a compactness constraint while offering a significant speedup over N-cuts, which does enforce compactness.
Alex Levinshtein, Adrian Stere, Kiriakos N. Kutulakos, David J. Fleet, Sven J. Dickinson, Kaleem Siddiqi
IEEE Trans. Pattern Anal. Mach. Intell.5
2008 From skeletons to bone graphs: Medial abstraction for object recognition
abstract
Medial descriptions, such as shock graphs, have gained significant momentum in the shape-based object recognition community due to their invariance to translation, rotation, scale and articulation and their ability to cope with moderate amounts of within-class deformation. While they attempt to decompose a shape into a set of parts, this decomposition can suffer from ligature-induced instability. In particular, the addition of even a small part can have a dramatic impact on the representation in the vicinity of its attachment. We present an algorithm for identifying and representing the ligature structure, and restoring the non-ligature structures that remain. This leads to a bone graph, a new medial shape abstraction that captures a more intuitive notion of an objectpsilas parts than a skeleton or a shock graph, and offers improved stability and within-class deformation invariance. We demonstrate these advantages by comparing the use of bone graphs to shock graphs in a set of view-based object recognition and pose estimation trials.
Diego Macrini, Kaleem Siddiqi, Sven J. Dickinson
CVPR3
2008 Retrieving articulated 3-D models using medial surfaces
Kaleem Siddiqi, Diego Macrini, Ali Shokoufandeh, Sylvain Bouix, Sven J. Dickinson
Mach. Vis. Appl.6
2008 A generalized family of fixed-radius distribution-based distance measures for content-based fMRI image retrieval
John Novatnack, Nicu D. Cornea, Ali Shokoufandeh, Deborah Silver, Sven J. Dickinson, Paul B. Kantor
Pattern Recognit. Lett.5
2007 Learning Structured Appearance Models from Captioned Images of Cluttered Scenes
abstract
Given an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to learn both the names and appearances of the objects. Only a small number of local features within any given image are associated with a particular caption word. We describe a connected graph appearance model where vertices represent local features and edges encode spatial relationships. We use the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to guide the search for meaningful feature configurations. We demonstrate improved results on a dataset to which an unstructured object model was previously applied. We also apply the new method to a more challenging collection of captioned images from the Web, detecting and annotating objects within highly cluttered realistic scenes.
Michael Jamieson, Afsaneh Fazly, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth
ICCV3
2006 Using Language to Drive the Perceptual Grouping of Local Image Features
abstract
We address the problem of learning both the semantics (names) and the visual features (SIFT collections) of objects appearing in a training set of unstructured, captioned images of cluttered scenes. Prior work in applying machine translation models to learn the associations between image features and caption nouns has assumed a one-toone correspondence between features and nouns. However, each training image may contain thousands of SIFT features belonging to multiple objects. Our challenge is two-fold: 1) grouping the SIFT features into meaningful collections, and 2) learning the object names associated with those collections. Since better collections tend to have stronger associations with object names, we offer an integrated solution that uses the caption words to drive the feature grouping process. The result is a more general model acquisition framework that does not assume words correspond to individual features and does not require training images with isolated objects or unambiguous labels. The model that is learned performs well at labeling cluttered scenes in a set of test images.
Michael Jamieson, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth
CVPR (2)2
2006 The representation and matching of categorical shape
Ali Shokoufandeh, Lars Bretzner, Diego Macrini, M. Fatih Demirci, Clas Jönsson, Sven J. Dickinson
Comput. Vis. Image Underst.6
2006 Object Recognition as Many-to-Many Feature Matching
M. Fatih Demirci, Ali Shokoufandeh, Yakov Keselman, Lars Bretzner, Sven J. Dickinson
Int. J. Comput. Vis.5
2006 Integrating region and boundary information for spatiallycoherent object tracking
Desmond Chung, W. James MacLean, Sven J. Dickinson
Image Vis. Comput.3
2006 Landmark Selection for Vision-Based Navigation
abstract
Recent work in the object recognition community has yielded a class of interest-point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions, such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem, and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data. Finally, we use the decompositions from the real image data to measure the localization performance versus the undecomposed map
Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson
IEEE Trans. Robotics4
2005 3D Object Retrieval using Many-to-many Matching of Curve Skeletons
abstract
We present a 3D matching framework based on a many-to-many matching algorithm that works with skeletal representations of 3D volumetric objects. We demonstrate the performance of this approach on a large database of 3D objects containing more than 1000 exemplars. The method is especially suited to matching objects with distinct part structure and is invariant to part articulation. Skeletal matching has an intuitive quality that helps in defining the search and visualizing the results. In particular, the matching algorithm produces a direct correspondence between two skeletons and their parts, which can be used for registration and juxtaposition.
Nicu D. Cornea, M. Fatih Demirci, Deborah Silver, Ali Shokoufandeh, Sven J. Dickinson, Paul B. Kantor
SMI5
2005 A Visualization Tool for fMRI Data Mining
Nicu D. Cornea, Ulukbek Ibraev, Deborah Silver, Paul B. Kantor, Ali Shokoufandeh, Jeff Abrahamson, Sven J. Dickinson
IEEE Visualization7
2005 Generic Model Abstraction from Examples
abstract
The recognition community has typically avoided bridging the representational gap between traditional, low-level image features and generic models. Instead, the gap has been artificially eliminated by either bringing the image closer to the models using simple scenes containing idealized, textureless objects or by bringing the models closer to the images using 3D CAD model templates or 2D appearance model templates. In this paper, we attempt to bridge the representational gap for the domain of model acquisition. Specifically, we address the problem of automatically acquiring a generic 2D view-based class model from a set of images, each containing an exemplar object belonging to that class. We introduce a novel graph-theoretical formulation of the problem in which we search for the lowest common abstraction among a set of lattices, each representing the space of all possible region groupings in a region adjacency graph representation of an input image. The problem is intractable and we present a shortest path-based approximation algorithm to yield an efficient solution. We demonstrate the approach on real imagery.
Yakov Keselman, Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Indexing Hierarchical Structures Using Graph Spectra
abstract
Hierarchical image structures are abundant in computer vision and have been used to encode part structure, scale spaces, and a variety of multiresolution features. In this paper, we describe a framework for indexing such representations that embeds the topological structure of a directed acyclic graph (DAG) into a low-dimensional vector space. Based on a novel spectral characterization of a DAG, this topological signature allows us to efficiently retrieve a promising set of candidates from a database of models using a simple nearest-neighbor search. We establish the insensitivity of the signature to minor perturbation of graph structure due to noise, occlusion, or node split/merge. To accommodate large-scale occlusion, the DAG rooted at each nonleaf node of the query "votes" for model objects that share that "part," effectively accumulating local evidence in a model DAG's topological subspaces. We demonstrate the approach with a series of indexing experiments in the domain of view-based 3D object recognition using shock graphs.
Ali Shokoufandeh, Diego Macrini, Sven J. Dickinson, Kaleem Siddiqi, Steven W. Zucker
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Incremental Model-Based Estimation Using Geometric Constraints
abstract
We present a model-based framework for incremental, adaptive object shape estimation and tracking in monocular image sequences. Parametric structure and motion estimation methods usually assume a fixed class of shape representation (splines, deformable superquadrics, etc.) that is initialized prior to tracking. Since the model shape coverage is fixed a priori, the incremental recovery of structure is decoupled from tracking, thereby limiting both processes in their scope and robustness. In this work, we describe a model-based framework that supports the automatic detection and integration of low-level geometric primitives (lines) incrementally. Such primitives are not explicitly captured in the initial model, but are moving consistently with its image motion. The consistency tests used to identify new structure are based on trinocular constraints between geometric primitives. The method allows not only an increase in the model scope, but also improves tracking accuracy by including the newly recovered features in its state estimation. The formulation is a step toward automatic model building, since it allows both weaker assumptions on the availability of a prior shape representation and on the number of features that would otherwise be necessary for entirely bottom-up reconstruction. We demonstrate the proposed approach on two separate image-based tracking domains, each involving complex 3D object structure and motion.
Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Many-to-Many Feature Matching Using Spherical Coding of Directed Graphs
M. Fatih Demirci, Ali Shokoufandeh, Sven J. Dickinson, Yakov Keselman, Lars Bretzner
ECCV (1)3
2004 Landmark selection for vision-based navigation
abstract
Recent work in the object recognition community has yielded a class of interest point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data.
Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson
IROS4
2003 Many-to-Many Graph Matching via Metric Embedding
abstract
Graph matching is an important component in many object recognition algorithms. Although most graph matching algorithms seek a one-to-one correspondence between nodes, it is often the case that a more meaningful correspondence exists between a cluster of nodes in one graph and a cluster of nodes in the other. We present a matching algorithm that establishes many-to-many correspondences between nodes of noisy, vertex-labeled weighted graphs. The algorithm is based on recent developments in efficient low-distortion metric embedding of graphs into normed vector spaces. By embedding weighted graphs into normed vector spaces, we reduce the problem of many-to-many graph matching to that of computing a distribution-based distance measure between graph embeddings. We use a specific measure, the earth mover's distance, to compute distances between sets of weighted vectors. Empirical evaluation of the algorithm on an extensive set of recognition trials demonstrates both the robustness and efficiency of the overall approach.
Yakov Keselman, Ali Shokoufandeh, M. Fatih Demirci, Sven J. Dickinson
CVPR (1)4
2003 Skeleton Based Shape Matching and Retrieval
abstract
We describe a novel method for searching and comparing 3D objects. The method encodes the geometric and topological information in the form of a skeletal graph and uses graph matching techniques to match the skeletons and to compare them. The skeletal graphs can be manually annotated to refine or restructure the search. This helps in choosing between a topological similarity and a geometric (shape) similarity. A feature of skeletal matching is the ability to perform part-matching, and its inherent intuitiveness, which helps in defining the search and in visualizing the results. Also, the matching results, which are presented in a per-node basis can be used for driving a number of registration algorithms, most of which require a good initial guess to perform registration. We also describe a visualization tool to aid in the selection and specification of the matched objects.
H. Sundar, Deborah Silver, Nikhil Gagvani, Sven J. Dickinson
Shape Modeling International4
2002 On the Representation and Matching of Qualitative Shape at Multiple Scales
Ali Shokoufandeh, Sven J. Dickinson, Clas Jönsson, Lars Bretzner, Tony Lindeberg
ECCV (3)2
2001 Generic Model Abstraction from Examples
abstract
The recognition community has long avoided bridging the representational gap between traditional, low-level image features and generic models. Instead, the gap has been artificially eliminated by either bringing the image closer to the models, using simple scenes containing idealized, textureless objects,,or by bringing the models closer to the images, using 3-D CAD model templates or 2-D appearance model templates. In this paper, we attempt to bridge the representational gap for the domain of model acquisition. Specifically, we address the problem of automatically acquiring a generic 2-D view-based class model from a set of images, each containing an exemplar object belonging to that class. We introduce a novel graph-theoretical formulation of the problem, and demonstrate the approach on real imagery.
Yakov Keselman, Sven J. Dickinson
CVPR (1)2
2001 Improving the Scope of Deformable Model Shape and Motion Estimation
abstract
Previous approaches to deformable model shape estimation and tracking have assumed a fixed class of shapes representation (e.g., deformable superquadrics), initialized prior to tracking. Since the shape coverage of the model is fixed, such approaches do not directly accommodate incremental representation discovery during tracking. As a result, model shape coverage is decoupled from tracking, thereby limiting both processes in terms of scope and robustness. We present a novel deformable model framework that accommodates the incremental incorporation during tracking of new geometric primitives (lines, in addition to points) that are not explicitly captured in the initial deformable model but that are moving consistently with its image motion. As these new features are detected via consistency checks, they are added to the model, providing incremental soft constraints on the estimation of its rigid parameters. The consistency checks are based on trilinear relationships between geometric primitives. Consequently, we not only increase both model scope and, ultimately, its higher-level shape coverage, but improve tracking robustness and accuracy, by directly employing the new features in both forward prediction and reconstruction. Our new formulation is a step towards automating model shape estimation and tracking, since it requires significantly reduced initial model hand-crafting. We demonstrate our approach on two separate image-based tracking domains, each involving complex 3D object shape and motion.
Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson
CVPR (1)3
2001 Introduction to the Special Section on Graph Algorithms in Computer Vision
abstract
N a letter to C. Huygens of 1679, G.W. Leibniz expressed his dissatisfaction with the standard coordinate treatment of geometric figures and maintained that we need yet another kind of analysis, geometric or linear, which deals directly with position, as algebra deals with magnitude (1). In fact, Leibniz initiated the study of the so-called geometry of positions (geometria situs) which, as L. Euler clearly put it in his famous 1736 Konigsberg bridges paper which had to mark the beginning of graph theory, concerned only with the determination of position, and its properties; it does not involve measurements nor calculations made with them (2). After about two centuries, this study developed into two of the richest branches of modern mathematics: graph theory and combinatorial topology. Mutatis mutandis, an analogous discontent is nowadays being felt among many researchers working in computer vision, a field that is currently dominated by purely geometric methods, who are increasingly making use of sophisticated graph-theoretic concepts, results, and algorithms. Indeed, graphs have long been an important tool in computer vision, especially because of their representational power and flexibility. However, there is now a renewed and growing interest toward explicitly formulating computer vision problems as graph problems. This is particularly advanta- geous because it allows vision problems to be cast in a pure, abstract setting with solid theoretical underpinnings and also permits access to the full arsenal of graph algorithms developed in computer science and operations research. Graph-theoretic problems which have proven to be relevant to computer vision include maximum flow, minimum spanning tree, maximum clique, shortest path, maximal common subtree/subgraph, etc. In addition, a number of fundamental techniques that were designed in the graph algorithms community have recently been applied to computer vision problems. Examples include spectral
Sven J. Dickinson, Marcello Pelillo, Ramin Zabih
IEEE Trans. Pattern Anal. Mach. Intell.1
1999 Indexing using a Spectral Encoding of Topological Structure
abstract
In an object recognition system, if the extracted image features are multilevel or multiscale, the indexing structure may take the form of a tree. Such structures are not only common in computer vision, but also appear in linguistics, graphics, computational biology, and a wide range of other domains. In this paper, we develop an indexing mechanism that maps the topological structure of a tree into a low-dimensional vector space. Based on a novel eigenvalue characterization of a tree, this topological signature allows us to efficiently retrieve a small set of candidates from a database of models. To accommodate occlusion and local deformation, local evidence is accumulated in each of the tree's topological subspaces. We demonstrate the approach with a series of indexing experiments in the domain of 2-D object recognition.
Ali Shokoufandeh, Sven J. Dickinson, Kaleem Siddiqi, Steven W. Zucker
CVPR2
1999 Shock Graphs and Shape Matching
Kaleem Siddiqi, Ali Shokoufandeh, Sven J. Dickinson, Steven W. Zucker
Int. J. Comput. Vis.3
1999 View-based object recognition using saliency maps
Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson
Image Vis. Comput.3
1999 A Computational Model of View Degeneracy
abstract
We quantify the observation by Kender and Freudenstein (1987) that degenerate views occupy a significant fraction of the viewing sphere surrounding an object. For a perspective camera geometry, we introduce a computational model that can be used to estimate the probability that a view degeneracy will occur in a random view of a polyhedral object. For a typical recognition system parameterization, view degeneracies typically occur with probabilities of 20 percent and, depending on the parameterization, as high as 50 percent. We discuss the impact of view degeneracy on the problem of object recognition and, for a particular recognition framework, relate the cost of object disambiguation to the probability of view degeneracy. To reduce this cost, we incorporate our model of view degeneracy in an active focal length control paradigm that balances the probability of view degeneracy with the camera field of view. In order to validate both our view degeneracy model as well as our active focal length control model, a set of experiments are reported using a real recognition system operating on real images.
Sven J. Dickinson, David Wilkes, John K. Tsotsos
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 View-Based Object Matching
abstract
We introduce a novel view-based object representation, called the saliency map graph (SMG), which captures the salient regions of an object view at multiple scales using a wavelet transform. This compact representation is highly invariant to translation, rotation (image and depth), and scaling, and offers the locality of representation required for occluded object recognition. To compare two saliency map graphs, we introduce two graph similarity algorithms. The first computes the topological similarity between two SMG's, providing a coarse-level matching of two graphs. The second computes the geometrical similarity between two SMG's, providing a fine-level matching of two graphs. We test and compare these two algorithms on a large database of model object views.
Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson
ICCV3
1998 Shock Graphs and Shape Matching
abstract
We have been developing a theory for the generic representation of 2-D shape, where structural descriptions are derived from the shocks (singularities) of a curve evolution process, acting on bounding contours. We now apply the theory to the problem of shape matching. The shocks are organized into a directed, acyclic shock graph, and complexity is managed by attending to the most significant (central) shape components first. The space of all such graphs is highly structured and can be characterized by the rules of a shock graph grammar. The grammar permits a reduction of a shockgraph to a unique rooted shock tree. We introduce a novel tree matching algorithm which finds the best set of corresponding nodes between two shock trees in polynomial time. Using a diverse database of shapes, we demonstrate our system's performance under articulation, occlusion, and changes in viewpoint.
Kaleem Siddiqi, Ali Shokoufandeh, Sven J. Dickinson, Steven W. Zucker
ICCV3
1998 PLAYBOT A visually-guided robot for physically disabled children
John K. Tsotsos, Gilbert Verghese, Sven J. Dickinson, Michael R. M. Jenkin, Allan Douglas Jepson, Evangelos E. Milios, Fernando Nuflo, Suzanne Stevenson, Michael J. Black, Dimitris N. Metaxas
Image Vis. Comput.3
1997 Active Object Recognition Integrating Attention and Viewpoint Control
Sven J. Dickinson, Henrik I. Christensen, John K. Tsotsos, Göran Olofsson
Comput. Vis. Image Underst.1
1997 Using Aspection Graphs to Control The Recovery Tracking of Deformable Models
abstract
Active or deformable models have emerged as a popular modeling paradigm in computer vision. These models have the flexibility to adapt themselves to the image data, offering the potential for both generic object recognition and non-rigid object tracking. Because these active models are underconstrained, however, deformable shape recovery often requires manual segmentation or good model initialization, while active contour trackers have been able to track only an object's translation in the image. In this paper, we report our current progress in using a part-based aspect graph representation of an object14 to provide the missing constraints on data-driven deformable model recovery and tracking processes.
Sven J. Dickinson
Int. J. Pattern Recognit. Artif. Intell.1
1997 Panel report: the potential of geons for generic 3-D object recognition
Sven J. Dickinson, Robert Bergevin, Irving Biederman, Jan-Olof Eklundh, Roger Munck-Fairwood, Anil K. Jain 0001, Alex Pentland
Image Vis. Comput.1
1997 The Role of Model-Based Segmentation in the Recovery of Volumetric Parts From Range Data
abstract
We present a method for segmenting and estimating the shape of 3D objects from range data. The technique uses model views, or aspects, to constrain the fitting of deformable models to range data. Based on an initial region segmentation of a range image, regions are grouped into aspects corresponding to the volumetric parts that make up an object. The qualitative segmentation of the range image into a set of volumetric parts not only captures the coarse shape of the parts, but qualitatively encodes the orientation of each part through its aspect. Knowledge of a part's coarse shape, its orientation, as well as the mapping between the faces in its aspect and the surfaces on the part provides strong constraints on the fitting of a deformable model (supporting both global and local deformations) to the data. Unlike previous work in physics-based deformable model recovery from range data, the technique does not require presegmented data. Furthermore, occlusion is handled at segmentation time and does not complicate the fitting process, as only 3D points known to belong to a part participate in the fitting of a model to the part. We present the approach in detail and apply it to the recovery of objects from range data.
Sven J. Dickinson, Dimitris N. Metaxas, Alex Pentland
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 A Quantitative Analysis of View Degeneracy and its use for Active Focal Length control
abstract
We quantify the observation by Kender and Freudenstein (1987) that degenerate views occupy a significant fraction of the viewing sphere surrounding an object. This demonstrates that systems for recognition must explicitly account for the possibility of view degeneracy. We show that view degeneracy cannot be detected from a single camera viewpoint. As a result, systems designed to recognize objects from a single arbitrary viewpoint must be able to function in spite of possible undetected degeneracies, or else operate with imaging parameters that cause acceptably low probabilities of degeneracy. To address this need, we give a prescription for active control of focal length that allows a principled tradeoff between the camera field of view and probability of view degeneracy.>
David Wilkes, Sven J. Dickinson, John K. Tsotsos
ICCV2
1995 Recognition by Functional Parts
Ehud Rivlin, Sven J. Dickinson, Azriel Rosenfeld
Comput. Vis. Image Underst.2
1994 A New Approach to Tracking 3D Objects in 2D Image Sequences
Michael Chan 0001, Dimitris N. Metaxas, Sven J. Dickinson
AAAI3
1994 Qualitative tracking of 3-D objects using active contour networks
abstract
In this paper, we track changes in the appearance of the object as it moves from one frame to the next. At a symbolic level, an aspect graph clusters all the views of an object into a set of topologically distinct classes in terms of which surfaces of an object are visible from a given viewpoint (Koenderink and van Doom (1979). Two nodes (or aspects) in the aspect graph are connected by an arc if it is possible to directly move from a viewpoint in which the first aspect is visible to a viewpoint in which the second aspect is visible. Qualitatively, we can envision a tracking strategy which simply tracks an object as it moves from one node to another in the object's aspect graph. Although it does not provide us with accurate pose of the object, it does qualitatively describe the motion of the object without the need for a CAD representation of the object.>
Sven J. Dickinson, Piotr Jasiobedzki, Göran Olofsson, Henrik I. Christensen
CVPR1
1994 Recognition by functional parts [function-based object recognition]
abstract
We present an approach to function-based object recognition that reasons about the functionality of an object's initiative parts. We extend the popular "recognition by parts" shape recognition framework to support "recognition, by functional parts", by combining a set of functional primitives and their relations with a set of abstract volumetric shape primitives and their relations. Previous approaches have relied on more global object features, often ignoring the problem of object segmentation, and thereby restricting themselves to range images of unoccluded scenes. We show how these shape primitives and relations can be easily recovered from superquadric ellipsoids which, in turn, can be recovered from either range or intensity images of occluded scenes. Furthermore, the proposed framework supports both unexpected (bottom-up) object recognition and expected (top-down) object recognition. We demonstrate the approach on, a simple domain by recognizing a restricted class of hand-tools from 2-D images.>
Ehud Rivlin, Sven J. Dickinson, Azriel Rosenfeld
CVPR2
1994 Active Object Recognition Integrating Attention and Viewpoint Control
Sven J. Dickinson, Henrik I. Christensen, John K. Tsotsos, Göran Olofsson
ECCV (2)1
1994 Physics-based tracking of 3D objects in 2D image sequences
abstract
We present a new technique for tracking 3D objects in 2D image sequences. We assume that objects are constructed from a class of volumetric part primitives. The models are initially recovered using a qualitative shape recovery process. We subsequently track the objects using local forces computed from image potentials. Therefore we avoid the expensive computation of image features. By integrating measurements from stereo images, 3D positions (as well as other model parameters) of the objects can be continuously updated using an extended Kalman filter. Our model-based approach can handle occlusions in scenes with multiple moving objects by predicting their occurrences. To handle severe or unexpected occlusion we use a feedback mechanism between the quantitative and qualitative shape estimation systems. We demonstrate our technique in experiments involving image sequences from complex motions of objects.
Michael Chan 0001, Dimitris N. Metaxas, Sven J. Dickinson
ICPR (1)3
1994 Navigation based on a network of 2D images
abstract
This paper describes the integration of 2D stimulus-driven robot localization and positioning with a token-based correspondence method in a practical robot navigation system. The approach allows for modular acquisition and update of world knowledge for navigation, and robustness of navigation to low-level errors. No special marking of the world is necessary, so the robot may operate in quite general environments. Tests in a real industrial environment confirm the potential of the method.
David Wilkes, Sven J. Dickinson, Ehud Rivlin, Ronen Basri
ICPR (1)2
1994 Integrating qualitative and quantitative shape recovery
Sven J. Dickinson, Dimitris N. Metaxas
Int. J. Comput. Vis.1
1993 Integration of quantitative and qualitative techniques for deformable model fitting from orthographic, perspective, and stereo projections
abstract
The authors synthesize a new approach to 3-D object shape recovery by integrating qualitative shape recovery techniques and quantitative physics-based shape estimation techniques. They first use qualitative shape recovery and recognition techniques to provide strong fitting constraints on physics-based deformable model recovery techniques. Previously developed techniques of fitting deformable models to occluding image contours are then extended to the case of image data captured under general orthographic, perspective, and stereo projections. Experimental results are presented to illustrate the shape recovery approach.>
Dimitris N. Metaxas, Sven J. Dickinson
ICCV2
1993 The Use of Geons for Generic 3D Object Recognition
Sven J. Dickinson, Robert Bergevin, Irving Biederman, Jan-Olof Eklundh, Roger Munck-Fairwood, Alex Pentland
IJCAI1
1992 From volumes to views: An approach to 3-D object recognition
Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld
CVGIP Image Underst.1
1992 3-D Shape Recovery Using Distributed Aspect Matching
abstract
An approach to the recovery of 3-D volumetric primitives from a single 2-D image is presented. The approach first takes a set of 3-D volumetric modeling primitives and generates a hierarchical aspect representation based on the projected surfaces of the primitives; conditional probabilities capture the ambiguity of mappings between levels of the hierarchy. From a region segmentation of the input image, the authors present a formulation of the recovery problem based on the grouping of the regions into aspects. No domain-independent heuristics are used; only the probabilities inherent in the aspect hierarchy are exploited. Once the aspects are recovered, the aspect hierarchy is used to infer a set of volumetric primitives and their connectivity. As a front end to an object recognition system, the approach provides the indexing power of complex 3-D object-centered primitives while exploiting the convenience of 2-D viewer-centered aspect matching; aspects are used to represent a finite vocabulary of 3-D parts from which objects can be constructed.>
Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 Qualitative 3-D shape reconstruction using distributed aspect graph matching
abstract
An approach is presented to 3-D primitive reconstruction that is independent of the selection of volumetric primitives used to model objects. The approach first takes an arbitrary set of 3-D volumetric primitives and generates a hierarchical aspect representation based on the projected surfaces of the primitives; conditional probabilities capture the ambiguity of mappings between levels of the hierarchy. The integration of object-centered and viewer-centered representations provides the indexing power of 3-D volumetric primitives, while supporting a 2-D matching paradigm for primitive reconstruction. Formulation of the problem based on grouping the image regions according to aspect is presented. No domain dependent heuristics are used; the authors exploit only the probabilities inherent in the aspect hierarchy. For a given selection of primitives, the success of the heuristic depends on the likelihood of the various aspects; best results are achieved when certain aspects are more likely, and fewer primitives project to a given aspect.>
Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld
ICCV1
1990 A flexible tool for prototyping ALV road following algorithms
abstract
A production system model of problem-solving is applied to the design of a vision system by which an autonomous land vehicle (ALV) navigates roads. The ALV vision task consists of hypothesizing objects in a scene model and verifying these hypotheses using the vehicle's sensors. Object hypothesis generation is based on the local navigation task, an a priori road map, and the contents of the scene model. Verification of an object hypothesis involves directing the sensors towards the expected location of the object, collecting evidence in support of the object, and reasoning about the evidence. Constructing the scene model consists of building a semantic network of object frames exhibiting component, spatial, and inheritance relationships. The control structure is provided by a set of communicating production systems implementing a structured blackboard; each production system contains rules for defining the attributes of a particular class of object frame. The combination of production system and object-oriented programming techniques results in a flexible control structure able to accommodate new object classes, reasoning strategies, vehicle sensors, and image analysis techniques.>
Sven J. Dickinson, Larry Davis 0001
IEEE Trans. Robotics Autom.1
1988 An expert vision system for autonomous land vehicle road following
abstract
A production-system model of problem solving is applied to the design of a vision system by which an autonomous land vehicle (ALV) navigates roads. The ALV vision task consists of hypothesizing objects in a scene model and verifying these hypotheses using the vehicles sensors. Object hypothesis generation is based on the local navigation task, and a priori road map, and the contents of the scene model. Verification of an object hypothesis involves directing the sensors toward the expected location of the object, collecting evidence in support of the object, and reasoning about the evidence. Constructing the scene model consists of building a semantic network of object frames exhibiting component, spatial, and inheritance relationships. The control structure is provided by a set of communicating production systems implementing a structured blackboard; each production system contains the rules for defining the attributes of a particular class of object frame.>
Sven J. Dickinson, Larry Davis 0001
CVPR1