VLDB 2026 Research / reviewers in the wild / expert
Sven J. Dickinson
dblp:d/SvenJDickinson
· DBLP profile ↗
90ranked-venue papers
21as first author
9since 2021 · last 2025
0000-0001-6566-9635ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 85 · 20 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 5 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Probabilistic Directed Distance Fields for Ray-Based Shape RepresentationsabstractIn modern computer vision, the optimal representation of 3D shape remains task-dependent. One fundamental operation applied to such representations is differentiable rendering, which enables learning-based inverse graphics approaches. Standard explicit representations are often easily rendered, but can suffer from limited geometric fidelity, among other issues. On the other hand, implicit representations generally preserve greater fidelity, but suffer from difficulties with rendering, limiting scalability. In this work, we devise Directed Distance Fields (DDFs), which map a ray or oriented point (position and direction) to surface visibility and depth. This enables efficient differentiable rendering, obtaining depth with a single forward pass per pixel, as well as higher-order geometry with only additional backward passes. Using probabilistic DDFs (PDDFs), we can model the inherent discontinuities in the underlying field. We then apply DDFs to single-shape fitting, generative modelling, and 3D reconstruction, showcasing strong performance with simple architectural components via the versatility of our representation. Finally, since the dimensionality of DDFs permits view-dependent geometric artifacts, we conduct a theoretical investigation of the constraints necessary for view consistency. We find a small set of field properties that are sufficient to guarantee a DDF is consistent, without knowing which shape the field is expressing. Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Shape-Based Measures Improve Scene CategorizationabstractConverging evidence indicates that deep neural network models that are trained on large datasets are biased toward color and texture information. Humans, on the other hand, can easily recognize objects and scenes from images as well as from bounding contours. Mid-level vision is characterized by the recombination and organization of simple primary features into more complex ones by a set of so-called Gestalt grouping rules. While described qualitatively in the human literature, a computational implementation of these perceptual grouping rules is so far missing. In this article, we contribute a novel set of algorithms for the detection of contour-based cues in complex scenes. We use the medial axis transform (MAT) to locally score contours according to these grouping rules. We demonstrate the benefit of these cues for scene categorization in two ways: (i) Both human observers and CNN models categorize scenes most accurately when perceptual grouping information is emphasized. (ii) Weighting the contours with these measures boosts performance of a CNN model significantly compared to the use of unweighted contours. Our work suggests that, even though these measures are computed directly from contours in the image, current CNN models do not appear to extract or utilize these grouping cues. Morteza Rezanejad, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Fast-Grasp'D: Dexterous Multi-finger Grasp Generation Through Differentiable SimulationabstractMulti-finger grasping relies on high quality training data, which is hard to obtain: human data is hard to transfer and synthetic data relies on simplifying assumptions that reduce grasp quality. By making grasp simulation differentiable, and contact dynamics amenable to gradient-based optimization, we accelerate the search for high-quality grasps with fewer limiting assumptions. We present Grasp'D-1M: a large-scale dataset for multi-finger robotic grasping, synthesized with Fast-Grasp'D, a novel differentiable grasping simulator. Grasp'D-1M contains one million training examples for three robotic hands (three, four and five-fingered), each with multimodal visual inputs (RGB+depth+segmentation, available in mono and stereo). Grasp synthesis with Fast-Grasp'D is 10x faster than GraspIt! [1] and 20x faster than the prior Grasp'D differentiable simulator [2]. Generated grasps are more stable and contact-rich than GraspIt! grasps, regardless of the distance threshold used for contact generation. We validate the usefulness of our dataset by retraining an existing vision-based grasping pipeline [3] on Grasp'D-1M, and showing a dramatic increase in model performance, predicting grasps with 30% more contact, a 33% higher epsilon metric, and 35% lower simulated displacement. Additional details at fast-graspd.github.io. Dylan Turpin, Tao Zhong 0003, Shutong Zhang, Guanglei Zhu, Eric Heiden, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg |
ICRA | 8 |
| 2023 | Disentangling Geometric Deformation Spaces in Generative Latent Shape Models
Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson |
Int. J. Comput. Vis. | 3 |
| 2022 | Representing 3D Shapes with Probabilistic Directed Distance FieldsabstractDifferentiable rendering is an essential operation in modern vision, allowing inverse graphics approaches to 3D understanding to be utilized in modern machine learning frameworks. Explicit shape representations (voxels, point clouds, or meshes), while relatively easily rendered, often suffer from limited geometric fidelity or topological con-straints. On the other hand, implicit representations (occu-pancy, distance, or radiance fields) preserve greater fidelity, but suffer from complex or inefficient rendering processes, limiting scalability. In this work, we endeavour to address both shortcomings with a novel shape representation that allows fast differentiable rendering within an implicit ar-chitecture. Building on implicit distance representations, we define Directed Distance Fields (DDFs), which map an oriented point (position and direction) to surface visibility and depth. Such a field can render a depth map with a single forward pass per pixel, enable differential surface geometry extraction (e.g., surface normals and curvatures) via network derivatives, be easily composed, and permit extraction of classical unsigned distance fields. Using probabilistic DDFs (PDDFs), we show how to model inherent discontinuities in the underlying field. Finally, we apply our method to fitting single shapes, unpaired 3D-aware generative image modelling, and single-image 3D reconstruction tasks, showcasing strong performance with simple architectural components via the versatility of our representation. Tristan Aumentado-Armstrong, Stavros Tsogkas, Sven J. Dickinson, Allan Douglas Jepson |
CVPR | 3 |
| 2022 | Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands
Dylan Turpin, Liquan Wang, Eric Heiden, Yun-Chun Chen, Miles Macklin, Stavros Tsogkas, Sven J. Dickinson, Animesh Garg |
ECCV (6) | 7 |
| 2022 | State of the Journal Editorial
Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | DeepFlux for Skeleton Detection in the Wild
Yongchao Xu, Yukang Wang, Stavros Tsogkas, Jianqiang Wan, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi |
Int. J. Comput. Vis. | 6 |
| 2021 | State of the Journal EditorialabstractPresents the state of the journal review for this issue of the publication. Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Appearance Shock Grammar for Fast Medial Axis Extraction From Real ImagesabstractWe combine ideas from shock graph theory with more recent appearance-based methods for medial axis extraction from complex natural scenes, improving upon the present best unsupervised method, in terms of efficiency and performance. We make the following specific contributions: i) we extend the shock graph representation to the domain of real images, by generalizing the shock type definitions using local, appearance-based criteria; ii) we then use the rules of a Shock Grammar to guide our search for medial points, drastically reducing run time when compared to other methods, which exhaustively consider all points in the input image; iii) we remove the need for typical post-processing steps including thinning, non-maximum suppression, and grouping, by adhering to the Shock Grammar rules while deriving the medial axis solution; iv) finally, we raise some fundamental concerns with the evaluation scheme used in previous work and propose a more appropriate alternative for assessing the performance of medial axis extraction from scenes. Our experiments on the BMAX500 and SK-LARGE datasets demonstrate the effectiveness of our approach. We outperform the present state-of-the-art, excelling particularly in the high-precision regime, while running an order of magnitude faster and requiring no post-processing. Charles-Olivier Dufresne Camaro, Morteza Rezanejad, Stavros Tsogkas, Kaleem Siddiqi, Sven J. Dickinson |
CVPR | 5 |
| 2020 | State of the Journal Editorial
Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Scene Categorization From Contours: Medial Axis Based Salience Measures
Morteza Rezanejad, Gabriel Downs, John Wilder, Dirk Bernhardt-Walther, Allan Douglas Jepson, Sven J. Dickinson, Kaleem Siddiqi |
CVPR | 6 |
| 2019 | DeepFlux for Skeletons in the WildabstractComputing object skeletons in natural images is challenging, owing to large variations in object appearance and scale, and the complexity of handling background clutter. Many recent methods frame object skeleton detection as a binary pixel classification problem, which is similar in spirit to learning-based edge detection, as well as to semantic segmentation methods. In the present article, we depart from this strategy by training a CNN to predict a two-dimensional vector field, which maps each scene point to a candidate skeleton pixel, in the spirit of flux-based skeletonization algorithms. This ``image context flux'' representation has two major advantages over previous approaches. First, it explicitly encodes the relative position of skeletal pixels to semantically meaningful entities, such as the image points in their spatial context, and hence also the implied object boundaries. Second, since the skeleton detection context is a region-based vector field, it is better able to cope with object parts of large width. We evaluate the proposed method on three benchmark datasets for skeleton detection and two for symmetry detection, achieving consistently superior performance over state-of-the-art methods. Yukang Wang, Yongchao Xu, Stavros Tsogkas, Xiang Bai, Sven J. Dickinson, Kaleem Siddiqi |
CVPR | 5 |
| 2019 | Geometric Disentanglement for Generative Latent Shape ModelsabstractRepresenting 3D shapes is a fundamental problem in artificial intelligence, which has numerous applications within computer vision and graphics. One avenue that has recently begun to be explored is the use of latent representations of generative models. However, it remains an open problem to learn a generative model of shapes that is interpretable and easily manipulated, particularly in the absence of supervised labels. In this paper, we propose an unsupervised approach to partitioning the latent space of a variational autoencoder for 3D point clouds in a natural way, using only geometric information, that builds upon prior work utilizing generative adversarial models of point sets. Our method makes use of tools from spectral geometry to separate intrinsic and extrinsic shape information, and then considers several hierarchical disentanglement penalties for dividing the latent space in this manner. We also propose a novel disentanglement penalty that penalizes the predicted change in the latent representation of the output,with respect to the latent variables of the initial shape. We show that the resulting latent representation exhibits intuitive and interpretable behaviour, enabling tasks such as pose transfer that cannot easily be performed by models with an entangled representation. Tristan Aumentado-Armstrong, Stavros Tsogkas, Allan Douglas Jepson, Sven J. Dickinson |
ICCV | 4 |
| 2019 | State of the JournalabstractPresents the current state of the journal and discusses future directions. Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | State of the JournalabstractPresents an editorial on the current state of the IEEE Transactions on Pattern Analysis and Machine Intelligence. Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | AMAT: Medial Axis Transform for Natural ImagesabstractWe introduce Appearance-MAT (AMAT), a generalization of the medial axis transform for natural images, that is framed as a weighted geometric set cover problem. We make the following contributions: i) we extend previous medial point detection methods for color images, by associating each medial point with a local scale; ii) inspired by the invertibility property of the binary MAT, we also associate each medial point with a local encoding that allows us to invert the AMAT, reconstructing the input image; iii) we describe a clustering scheme that takes advantage of the additional scale and appearance information to group individual points into medial branches, providing a shape decomposition of the underlying image regions. In our experiments, we show state-of-the-art performance in medial point detection on Berkeley Medial AXes (BMAX500), a new dataset of medial axes based on the BSDS500 database, and good generalization on the SK506 and WH-SYMMAX datasets. We also measure the quality of reconstructed images from BMAX500, obtained by inverting their computed AMAT. Our approach delivers significantly better reconstruction quality w.r.t. to three baselines, using just 10% of the image pixels. Our code and annotations are available at https://github.com/tsogkas/amat. Stavros Tsogkas, Sven J. Dickinson |
ICCV | 2 |
| 2017 | Incoming EIC EditorialabstractPresents the incoming editorial by the new Editor-In-Chief. Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Learning to Combine Mid-Level Cues for Object Proposal GenerationabstractIn recent years, region proposals have replaced sliding windows in support of object recognition, offering more discriminating shape and appearance information through improved localization. One powerful approach for generating region proposals is based on minimizing parametric energy functions with parametric maxflow. In this paper, we introduce Parametric Min-Loss (PML), a novel structured learning framework for parametric energy functions. While PML is generally applicable to different domains, we use it in the context of region proposals to learn to combine a set of mid-level grouping cues to yield a small set of object region proposals with high recall. Our learning framework accounts for multiple diverse outputs, and is complemented by diversification seeds based on image location and color. This approach casts perceptual grouping and cue combination in a novel structured learning framework which yields baseline improvements on VOC 2012 and COCO 2014. Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson |
ICCV | 3 |
| 2014 | Multi-cue Mid-level Grouping
Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson |
ACCV (3) | 3 |
| 2013 | Recognize Human Activities from Partially Observed VideosabstractRecognizing human activities in partially observed videos is a challenging problem and has many practical applications. When the unobserved subsequence is at the end of the video, the problem is reduced to activity prediction from unfinished activity streaming, which has been studied by many researchers. However, in the general case, an unobserved subsequence may occur at any time by yielding a temporal gap in the video. In this paper, we propose a new method that can recognize human activities from partially observed videos in the general case. Specifically, we formulate the problem into a probabilistic framework: 1) dividing each activity into multiple ordered temporal segments, 2) using spatiotemporal features of the training video samples in each segment as bases and applying sparse coding (SC) to derive the activity likelihood of the test video sample at each segment, and 3) finally combining the likelihood at each segment to achieve a global posterior for the activities. We further extend the proposed method to include more bases that correspond to a mixture of segments with different temporal lengths (MSSC), which can better represent the activities with large intra-class variations. We evaluate the proposed methods (SC and MSSC) on various real videos. We also evaluate the proposed methods on two special cases: 1) activity prediction where the unobserved subsequence is at the end of the video, and 2) human activity recognition on fully observed videos. Experimental results show that the proposed methods outperform existing state-of-the-art comparison methods. Yu Cao 0003, Daniel Paul Barrett, Andrei Barbu, N. Siddharth 0001, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002 |
CVPR | 8 |
| 2013 | Detecting Curved Symmetric Parts Using a Deformable Disc ModelabstractSymmetry is a powerful shape regularity that's been exploited by perceptual grouping researchers in both human and computer vision to recover part structure from an image without a priori knowledge of scene content. Drawing on the concept of a medial axis, defined as the locus of centers of maximal inscribed discs that sweep out a symmetric part, we model part recovery as the search for a sequence of deformable maximal inscribed disc hypotheses generated from a multiscale super pixel segmentation, a framework proposed by LEV09. However, we learn affinities between adjacent super pixels in a space that's invariant to bending and tapering along the symmetry axis, enabling us to capture a wider class of symmetric parts. Moreover, we introduce a global cost that perceptually integrates the hypothesis space by combining a pair wise and a higher-level smoothing term, which we minimize globally using dynamic programming. The new framework is demonstrated on two datasets, and is shown to significantly outperform the baseline LEV09. Tom Sie Ho Lee, Sanja Fidler, Sven J. Dickinson |
ICCV | 3 |
| 2013 | Multiscale Symmetric Part Detection and Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson |
Int. J. Comput. Vis. | 3 |
| 2012 | Superedge grouping for object localization by combining appearance and shape informationabstractBoth appearance and shape play important roles in object localization and object detection. In this paper, we propose a new superedge grouping method for object localization by incorporating both boundary shape and appearance information of objects. Compared with the previous edge grouping methods, the proposed method does not subdivide detected edges into short edgels before grouping. Such long, unsubdivided superedges not only facilitate the incorporation of object shape information into localization, but also increase the robustness against image noise and reduce computation. We identify and address several important problems in achieving the proposed superedge grouping, including gap filling for connecting superedges, accurate encoding of region-based information into individual edges, and the incorporation of object-shape information into object localization. In this paper, we use the bag of visual words technique to quantify the region-based appearance features of the object of interest. We find that the proposed method, by integrating both boundary and region information, can produce better localization performance than previous subwindow search and edge grouping methods on most of the 20 object categories from the VOC 2007 database. Experiments also show that the proposed method is roughly 50 times faster than the previous edge grouping method. Sanja Fidler, Jarrell W. Waggoner, Yu Cao 0003, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002 |
CVPR | 5 |
| 2012 | Detecting Reduplication in Videos of American Sign Language
Zoya Gavrilov, Stan Sclaroff, Carol Neidle, Sven J. Dickinson |
LREC | 4 |
| 2012 | 3D Object Detection and Viewpoint Estimation with a Deformable 3D Cuboid ModelabstractThis paper addresses the problem of category-level 3D object detection. Given a monocular image, our aim is to localize the objects in 3D by enclosing them with tight oriented 3D bounding boxes. We propose a novel approach that extends the well-acclaimed deformable part-based model[Felz.] to reason in 3D. Our model represents an object class as a deformable 3D cuboid composed of faces and parts, which are both allowed to deform with respect to their anchors on the 3D box. We model the appearance of each face in fronto-parallel coordinates, thus effectively factoring out the appearance variation induced by viewpoint. Our model reasons about face visibility patters called aspects. We train the cuboid model jointly and discriminatively and share weights across all aspects to attain efficiency. Inference then entails sliding and rotating the box in 3D and scoring object hypotheses. While for inference we discretize the search space, the variables are continuous in our model. We demonstrate the effectiveness of our approach in indoor and outdoor scenarios, and show that our approach outperforms the state-of-the-art in both 2D[Felz09] and 3D object detection[Hedau12]. Sanja Fidler, Sven J. Dickinson, Raquel Urtasun |
NIPS | 2 |
| 2012 | Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven J. Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, N. Siddharth 0001, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell W. Waggoner, Song Wang 0002, Jinlian Wei |
UAI | 5 |
| 2012 | Discovering hierarchical object models from captioned images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson |
Comput. Vis. Image Underst. | 5 |
| 2012 | Optimal Image and Video Closure by Superpixel Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson |
Int. J. Comput. Vis. | 3 |
| 2011 | Efficient many-to-many feature matching under the l1 norm
M. Fatih Demirci, Yusuf Osmanlioglu, Ali Shokoufandeh, Sven J. Dickinson |
Comput. Vis. Image Underst. | 4 |
| 2011 | Bone graphs: Medial shape parsing and abstraction
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi |
Comput. Vis. Image Underst. | 2 |
| 2011 | Object categorization using bone graphs
Diego Macrini, Sven J. Dickinson, David J. Fleet, Kaleem Siddiqi |
Comput. Vis. Image Underst. | 2 |
| 2010 | Spatiotemporal Closure
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson |
ACCV (1) | 3 |
| 2010 | Spatiotemporal Contour Grouping Using Abstract Part Models
Pablo Sala, Diego Macrini, Sven J. Dickinson |
ACCV (4) | 3 |
| 2010 | Discovering Multipart Appearance Models from Captioned Images
Michael Jamieson, Yulia Eskin, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson |
ECCV (5) | 5 |
| 2010 | Optimal Contour Closure by Superpixel Grouping
Alex Levinshtein, Cristian Sminchisescu, Sven J. Dickinson |
ECCV (2) | 3 |
| 2010 | Contour Grouping and Abstraction Using Simple Part Models
Pablo Sala, Sven J. Dickinson |
ECCV (5) | 2 |
| 2010 | Using Language to Learn Structured Appearance Models for Image AnnotationabstractGiven an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to simultaneously learn the names and appearances of the objects. Only a small fraction of local features within any given image are associated with a particular caption word, and captions may contain irrelevant words not associated with any image object. We propose a novel algorithm that uses the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to learn meaningful feature configurations (representing named objects). We also introduce a graph-based appearance model that captures some of the structure of an object by encoding the spatial relationships among the local visual features. In an iterative procedure, we use language (the words) to drive a perceptual grouping process that assembles an appearance model for a named object. Results of applying our method to three data sets in a variety of conditions demonstrate that, from complex, cluttered, real-world scenes with noisy captions, we can learn both the names and appearances of objects, resulting in a set of models invariant to translation, scale, orientation, occlusion, and minor changes in viewpoint or articulation. These named models, in turn, are used to automatically annotate new, uncaptioned images, thereby facilitating keyword-based image retrieval. Michael Jamieson, Afsaneh Fazly, Suzanne Stevenson, Sven J. Dickinson, Sven Wachsmuth |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2009 | Multiscale symmetric part detection and groupingabstractSkeletonization algorithms typically decompose an object's silhouette into a set of symmetric parts, offering a powerful representation for shape categorization. However, having access to an object's silhouette assumes correct figure-ground segmentation, leading to a disconnect with the mainstream categorization community, which attempts to recognize objects from cluttered images. In this paper, we present a novel approach to recovering and grouping the symmetric parts of an object from a cluttered scene. We begin by using a multiresolution superpixel segmentation to generate medial point hypotheses, and use a learned affinity function to perceptually group nearby medial points likely to belong to the same medial branch. In the next stage, we learn higher granularity affinity functions to group the resulting medial branches likely to belong to the same object. The resulting framework yields a skeletal approximation that's free of many of the instabilities plaguing traditional skeletons. More importantly, it doesn't require a closed contour, enabling the application of skeleton-based categorization systems to more realistic imagery Alex Levinshtein, Sven J. Dickinson, Cristian Sminchisescu |
ICCV | 2 |
| 2009 | Skeletal Shape Abstraction from ExamplesabstractLearning a class prototype from a set of exemplars is an important challenge facing researchers in object categorization. Although the problem is receiving growing interest, most approaches assume a one-to-one correspondence among local features, restricting their ability to learn true abstractions of a shape. In this paper, we present a new technique for learning an abstract shape prototype from a set of exemplars whose features are in many-to-many correspondence. Focusing on the domain of 2D shape, we represent a silhouette as a medial axis graph whose nodes correspond to "parts" defined by medial branches and whose edges connect adjacent parts. Given a pair of medial axis graphs, we establish a many-to-many correspondence between their nodes to find correspondences among articulating parts. Based on these correspondences, we recover the abstracted medial axis graph along with the positional and radial attributes associated with its nodes. We evaluate the abstracted prototypes in the context of a recognition task. M. Fatih Demirci, Ali Shokoufandeh, Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | TurboPixels: Fast Superpixels Using Geometric FlowsabstractWe describe a geometric-flow-based algorithm for computing a dense oversegmentation of an image, often referred to as superpixels. It produces segments that, on one hand, respect local image boundaries, while, on the other hand, limiting undersegmentation through a compactness constraint. It is very fast, with complexity that is approximately linear in image size, and can be applied to megapixel sized images with high superpixel densities in a matter of minutes. We show qualitative demonstrations of high-quality results on several complex images. The Berkeley database is used to quantitatively compare its performance to a number of oversegmentation algorithms, showing that it yields less undersegmentation than algorithms that lack a compactness constraint while offering a significant speedup over N-cuts, which does enforce compactness. Alex Levinshtein, Adrian Stere, Kiriakos N. Kutulakos, David J. Fleet, Sven J. Dickinson, Kaleem Siddiqi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2008 | From skeletons to bone graphs: Medial abstraction for object recognitionabstractMedial descriptions, such as shock graphs, have gained significant momentum in the shape-based object recognition community due to their invariance to translation, rotation, scale and articulation and their ability to cope with moderate amounts of within-class deformation. While they attempt to decompose a shape into a set of parts, this decomposition can suffer from ligature-induced instability. In particular, the addition of even a small part can have a dramatic impact on the representation in the vicinity of its attachment. We present an algorithm for identifying and representing the ligature structure, and restoring the non-ligature structures that remain. This leads to a bone graph, a new medial shape abstraction that captures a more intuitive notion of an objectpsilas parts than a skeleton or a shock graph, and offers improved stability and within-class deformation invariance. We demonstrate these advantages by comparing the use of bone graphs to shock graphs in a set of view-based object recognition and pose estimation trials. Diego Macrini, Kaleem Siddiqi, Sven J. Dickinson |
CVPR | 3 |
| 2008 | Retrieving articulated 3-D models using medial surfaces
Kaleem Siddiqi, Diego Macrini, Ali Shokoufandeh, Sylvain Bouix, Sven J. Dickinson |
Mach. Vis. Appl. | 6 |
| 2008 | A generalized family of fixed-radius distribution-based distance measures for content-based fMRI image retrieval
John Novatnack, Nicu D. Cornea, Ali Shokoufandeh, Deborah Silver, Sven J. Dickinson, Paul B. Kantor |
Pattern Recognit. Lett. | 5 |
| 2007 | Learning Structured Appearance Models from Captioned Images of Cluttered ScenesabstractGiven an unstructured collection of captioned images of cluttered scenes featuring a variety of objects, our goal is to learn both the names and appearances of the objects. Only a small number of local features within any given image are associated with a particular caption word. We describe a connected graph appearance model where vertices represent local features and edges encode spatial relationships. We use the repetition of feature neighborhoods across training images and a measure of correspondence with caption words to guide the search for meaningful feature configurations. We demonstrate improved results on a dataset to which an unstructured object model was previously applied. We also apply the new method to a more challenging collection of captioned images from the Web, detecting and annotating objects within highly cluttered realistic scenes. Michael Jamieson, Afsaneh Fazly, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth |
ICCV | 3 |
| 2006 | Using Language to Drive the Perceptual Grouping of Local Image FeaturesabstractWe address the problem of learning both the semantics (names) and the visual features (SIFT collections) of objects appearing in a training set of unstructured, captioned images of cluttered scenes. Prior work in applying machine translation models to learn the associations between image features and caption nouns has assumed a one-toone correspondence between features and nouns. However, each training image may contain thousands of SIFT features belonging to multiple objects. Our challenge is two-fold: 1) grouping the SIFT features into meaningful collections, and 2) learning the object names associated with those collections. Since better collections tend to have stronger associations with object names, we offer an integrated solution that uses the caption words to drive the feature grouping process. The result is a more general model acquisition framework that does not assume words correspond to individual features and does not require training images with isolated objects or unambiguous labels. The model that is learned performs well at labeling cluttered scenes in a set of test images. Michael Jamieson, Sven J. Dickinson, Suzanne Stevenson, Sven Wachsmuth |
CVPR (2) | 2 |
| 2006 | The representation and matching of categorical shape
Ali Shokoufandeh, Lars Bretzner, Diego Macrini, M. Fatih Demirci, Clas Jönsson, Sven J. Dickinson |
Comput. Vis. Image Underst. | 6 |
| 2006 | Object Recognition as Many-to-Many Feature Matching
M. Fatih Demirci, Ali Shokoufandeh, Yakov Keselman, Lars Bretzner, Sven J. Dickinson |
Int. J. Comput. Vis. | 5 |
| 2006 | Integrating region and boundary information for spatiallycoherent object tracking
Desmond Chung, W. James MacLean, Sven J. Dickinson |
Image Vis. Comput. | 3 |
| 2006 | Landmark Selection for Vision-Based NavigationabstractRecent work in the object recognition community has yielded a class of interest-point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions, such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem, and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data. Finally, we use the decompositions from the real image data to measure the localization performance versus the undecomposed map Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson |
IEEE Trans. Robotics | 4 |
| 2005 | 3D Object Retrieval using Many-to-many Matching of Curve SkeletonsabstractWe present a 3D matching framework based on a many-to-many matching algorithm that works with skeletal representations of 3D volumetric objects. We demonstrate the performance of this approach on a large database of 3D objects containing more than 1000 exemplars. The method is especially suited to matching objects with distinct part structure and is invariant to part articulation. Skeletal matching has an intuitive quality that helps in defining the search and visualizing the results. In particular, the matching algorithm produces a direct correspondence between two skeletons and their parts, which can be used for registration and juxtaposition. Nicu D. Cornea, M. Fatih Demirci, Deborah Silver, Ali Shokoufandeh, Sven J. Dickinson, Paul B. Kantor |
SMI | 5 |
| 2005 | A Visualization Tool for fMRI Data Mining
Nicu D. Cornea, Ulukbek Ibraev, Deborah Silver, Paul B. Kantor, Ali Shokoufandeh, Jeff Abrahamson, Sven J. Dickinson |
IEEE Visualization | 7 |
| 2005 | Generic Model Abstraction from ExamplesabstractThe recognition community has typically avoided bridging the representational gap between traditional, low-level image features and generic models. Instead, the gap has been artificially eliminated by either bringing the image closer to the models using simple scenes containing idealized, textureless objects or by bringing the models closer to the images using 3D CAD model templates or 2D appearance model templates. In this paper, we attempt to bridge the representational gap for the domain of model acquisition. Specifically, we address the problem of automatically acquiring a generic 2D view-based class model from a set of images, each containing an exemplar object belonging to that class. We introduce a novel graph-theoretical formulation of the problem in which we search for the lowest common abstraction among a set of lattices, each representing the space of all possible region groupings in a region adjacency graph representation of an input image. The problem is intractable and we present a shortest path-based approximation algorithm to yield an efficient solution. We demonstrate the approach on real imagery. Yakov Keselman, Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Indexing Hierarchical Structures Using Graph SpectraabstractHierarchical image structures are abundant in computer vision and have been used to encode part structure, scale spaces, and a variety of multiresolution features. In this paper, we describe a framework for indexing such representations that embeds the topological structure of a directed acyclic graph (DAG) into a low-dimensional vector space. Based on a novel spectral characterization of a DAG, this topological signature allows us to efficiently retrieve a promising set of candidates from a database of models using a simple nearest-neighbor search. We establish the insensitivity of the signature to minor perturbation of graph structure due to noise, occlusion, or node split/merge. To accommodate large-scale occlusion, the DAG rooted at each nonleaf node of the query "votes" for model objects that share that "part," effectively accumulating local evidence in a model DAG's topological subspaces. We demonstrate the approach with a series of indexing experiments in the domain of view-based 3D object recognition using shock graphs. Ali Shokoufandeh, Diego Macrini, Sven J. Dickinson, Kaleem Siddiqi, Steven W. Zucker |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Incremental Model-Based Estimation Using Geometric ConstraintsabstractWe present a model-based framework for incremental, adaptive object shape estimation and tracking in monocular image sequences. Parametric structure and motion estimation methods usually assume a fixed class of shape representation (splines, deformable superquadrics, etc.) that is initialized prior to tracking. Since the model shape coverage is fixed a priori, the incremental recovery of structure is decoupled from tracking, thereby limiting both processes in their scope and robustness. In this work, we describe a model-based framework that supports the automatic detection and integration of low-level geometric primitives (lines) incrementally. Such primitives are not explicitly captured in the initial model, but are moving consistently with its image motion. The consistency tests used to identify new structure are based on trinocular constraints between geometric primitives. The method allows not only an increase in the model scope, but also improves tracking accuracy by including the newly recovered features in its state estimation. The formulation is a step toward automatic model building, since it allows both weaker assumptions on the availability of a prior shape representation and on the number of features that would otherwise be necessary for entirely bottom-up reconstruction. We demonstrate the proposed approach on two separate image-based tracking domains, each involving complex 3D object structure and motion. Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Many-to-Many Feature Matching Using Spherical Coding of Directed Graphs
M. Fatih Demirci, Ali Shokoufandeh, Sven J. Dickinson, Yakov Keselman, Lars Bretzner |
ECCV (1) | 3 |
| 2004 | Landmark selection for vision-based navigationabstractRecent work in the object recognition community has yielded a class of interest point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data. Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson |
IROS | 4 |
| 2003 | Many-to-Many Graph Matching via Metric EmbeddingabstractGraph matching is an important component in many object recognition algorithms. Although most graph matching algorithms seek a one-to-one correspondence between nodes, it is often the case that a more meaningful correspondence exists between a cluster of nodes in one graph and a cluster of nodes in the other. We present a matching algorithm that establishes many-to-many correspondences between nodes of noisy, vertex-labeled weighted graphs. The algorithm is based on recent developments in efficient low-distortion metric embedding of graphs into normed vector spaces. By embedding weighted graphs into normed vector spaces, we reduce the problem of many-to-many graph matching to that of computing a distribution-based distance measure between graph embeddings. We use a specific measure, the earth mover's distance, to compute distances between sets of weighted vectors. Empirical evaluation of the algorithm on an extensive set of recognition trials demonstrates both the robustness and efficiency of the overall approach. Yakov Keselman, Ali Shokoufandeh, M. Fatih Demirci, Sven J. Dickinson |
CVPR (1) | 4 |
| 2003 | Skeleton Based Shape Matching and RetrievalabstractWe describe a novel method for searching and comparing 3D objects. The method encodes the geometric and topological information in the form of a skeletal graph and uses graph matching techniques to match the skeletons and to compare them. The skeletal graphs can be manually annotated to refine or restructure the search. This helps in choosing between a topological similarity and a geometric (shape) similarity. A feature of skeletal matching is the ability to perform part-matching, and its inherent intuitiveness, which helps in defining the search and in visualizing the results. Also, the matching results, which are presented in a per-node basis can be used for driving a number of registration algorithms, most of which require a good initial guess to perform registration. We also describe a visualization tool to aid in the selection and specification of the matched objects. H. Sundar, Deborah Silver, Nikhil Gagvani, Sven J. Dickinson |
Shape Modeling International | 4 |
| 2002 | On the Representation and Matching of Qualitative Shape at Multiple Scales
Ali Shokoufandeh, Sven J. Dickinson, Clas Jönsson, Lars Bretzner, Tony Lindeberg |
ECCV (3) | 2 |
| 2001 | Generic Model Abstraction from ExamplesabstractThe recognition community has long avoided bridging the representational gap between traditional, low-level image features and generic models. Instead, the gap has been artificially eliminated by either bringing the image closer to the models, using simple scenes containing idealized, textureless objects,,or by bringing the models closer to the images, using 3-D CAD model templates or 2-D appearance model templates. In this paper, we attempt to bridge the representational gap for the domain of model acquisition. Specifically, we address the problem of automatically acquiring a generic 2-D view-based class model from a set of images, each containing an exemplar object belonging to that class. We introduce a novel graph-theoretical formulation of the problem, and demonstrate the approach on real imagery. Yakov Keselman, Sven J. Dickinson |
CVPR (1) | 2 |
| 2001 | Improving the Scope of Deformable Model Shape and Motion EstimationabstractPrevious approaches to deformable model shape estimation and tracking have assumed a fixed class of shapes representation (e.g., deformable superquadrics), initialized prior to tracking. Since the shape coverage of the model is fixed, such approaches do not directly accommodate incremental representation discovery during tracking. As a result, model shape coverage is decoupled from tracking, thereby limiting both processes in terms of scope and robustness. We present a novel deformable model framework that accommodates the incremental incorporation during tracking of new geometric primitives (lines, in addition to points) that are not explicitly captured in the initial deformable model but that are moving consistently with its image motion. As these new features are detected via consistency checks, they are added to the model, providing incremental soft constraints on the estimation of its rigid parameters. The consistency checks are based on trilinear relationships between geometric primitives. Consequently, we not only increase both model scope and, ultimately, its higher-level shape coverage, but improve tracking robustness and accuracy, by directly employing the new features in both forward prediction and reconstruction. Our new formulation is a step towards automating model shape estimation and tracking, since it requires significantly reduced initial model hand-crafting. We demonstrate our approach on two separate image-based tracking domains, each involving complex 3D object shape and motion. Cristian Sminchisescu, Dimitris N. Metaxas, Sven J. Dickinson |
CVPR (1) | 3 |
| 2001 | Introduction to the Special Section on Graph Algorithms in Computer VisionabstractN a letter to C. Huygens of 1679, G.W. Leibniz expressed his dissatisfaction with the standard coordinate treatment of geometric figures and maintained that we need yet another kind of analysis, geometric or linear, which deals directly with position, as algebra deals with magnitude (1). In fact, Leibniz initiated the study of the so-called geometry of positions (geometria situs) which, as L. Euler clearly put it in his famous 1736 Konigsberg bridges paper which had to mark the beginning of graph theory, concerned only with the determination of position, and its properties; it does not involve measurements nor calculations made with them (2). After about two centuries, this study developed into two of the richest branches of modern mathematics: graph theory and combinatorial topology. Mutatis mutandis, an analogous discontent is nowadays being felt among many researchers working in computer vision, a field that is currently dominated by purely geometric methods, who are increasingly making use of sophisticated graph-theoretic concepts, results, and algorithms. Indeed, graphs have long been an important tool in computer vision, especially because of their representational power and flexibility. However, there is now a renewed and growing interest toward explicitly formulating computer vision problems as graph problems. This is particularly advanta- geous because it allows vision problems to be cast in a pure, abstract setting with solid theoretical underpinnings and also permits access to the full arsenal of graph algorithms developed in computer science and operations research. Graph-theoretic problems which have proven to be relevant to computer vision include maximum flow, minimum spanning tree, maximum clique, shortest path, maximal common subtree/subgraph, etc. In addition, a number of fundamental techniques that were designed in the graph algorithms community have recently been applied to computer vision problems. Examples include spectral Sven J. Dickinson, Marcello Pelillo, Ramin Zabih |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Indexing using a Spectral Encoding of Topological StructureabstractIn an object recognition system, if the extracted image features are multilevel or multiscale, the indexing structure may take the form of a tree. Such structures are not only common in computer vision, but also appear in linguistics, graphics, computational biology, and a wide range of other domains. In this paper, we develop an indexing mechanism that maps the topological structure of a tree into a low-dimensional vector space. Based on a novel eigenvalue characterization of a tree, this topological signature allows us to efficiently retrieve a small set of candidates from a database of models. To accommodate occlusion and local deformation, local evidence is accumulated in each of the tree's topological subspaces. We demonstrate the approach with a series of indexing experiments in the domain of 2-D object recognition. Ali Shokoufandeh, Sven J. Dickinson, Kaleem Siddiqi, Steven W. Zucker |
CVPR | 2 |
| 1999 | Shock Graphs and Shape Matching
Kaleem Siddiqi, Ali Shokoufandeh, Sven J. Dickinson, Steven W. Zucker |
Int. J. Comput. Vis. | 3 |
| 1999 | View-based object recognition using saliency maps
Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson |
Image Vis. Comput. | 3 |
| 1999 | A Computational Model of View DegeneracyabstractWe quantify the observation by Kender and Freudenstein (1987) that degenerate views occupy a significant fraction of the viewing sphere surrounding an object. For a perspective camera geometry, we introduce a computational model that can be used to estimate the probability that a view degeneracy will occur in a random view of a polyhedral object. For a typical recognition system parameterization, view degeneracies typically occur with probabilities of 20 percent and, depending on the parameterization, as high as 50 percent. We discuss the impact of view degeneracy on the problem of object recognition and, for a particular recognition framework, relate the cost of object disambiguation to the probability of view degeneracy. To reduce this cost, we incorporate our model of view degeneracy in an active focal length control paradigm that balances the probability of view degeneracy with the camera field of view. In order to validate both our view degeneracy model as well as our active focal length control model, a set of experiments are reported using a real recognition system operating on real images. Sven J. Dickinson, David Wilkes, John K. Tsotsos |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | View-Based Object MatchingabstractWe introduce a novel view-based object representation, called the saliency map graph (SMG), which captures the salient regions of an object view at multiple scales using a wavelet transform. This compact representation is highly invariant to translation, rotation (image and depth), and scaling, and offers the locality of representation required for occluded object recognition. To compare two saliency map graphs, we introduce two graph similarity algorithms. The first computes the topological similarity between two SMG's, providing a coarse-level matching of two graphs. The second computes the geometrical similarity between two SMG's, providing a fine-level matching of two graphs. We test and compare these two algorithms on a large database of model object views. Ali Shokoufandeh, Ivan Marsic, Sven J. Dickinson |
ICCV | 3 |
| 1998 | Shock Graphs and Shape MatchingabstractWe have been developing a theory for the generic representation of 2-D shape, where structural descriptions are derived from the shocks (singularities) of a curve evolution process, acting on bounding contours. We now apply the theory to the problem of shape matching. The shocks are organized into a directed, acyclic shock graph, and complexity is managed by attending to the most significant (central) shape components first. The space of all such graphs is highly structured and can be characterized by the rules of a shock graph grammar. The grammar permits a reduction of a shockgraph to a unique rooted shock tree. We introduce a novel tree matching algorithm which finds the best set of corresponding nodes between two shock trees in polynomial time. Using a diverse database of shapes, we demonstrate our system's performance under articulation, occlusion, and changes in viewpoint. Kaleem Siddiqi, Ali Shokoufandeh, Sven J. Dickinson, Steven W. Zucker |
ICCV | 3 |
| 1998 | PLAYBOT A visually-guided robot for physically disabled children
John K. Tsotsos, Gilbert Verghese, Sven J. Dickinson, Michael R. M. Jenkin, Allan Douglas Jepson, Evangelos E. Milios, Fernando Nuflo, Suzanne Stevenson, Michael J. Black, Dimitris N. Metaxas |
Image Vis. Comput. | 3 |
| 1997 | Active Object Recognition Integrating Attention and Viewpoint Control
Sven J. Dickinson, Henrik I. Christensen, John K. Tsotsos, Göran Olofsson |
Comput. Vis. Image Underst. | 1 |
| 1997 | Using Aspection Graphs to Control The Recovery Tracking of Deformable ModelsabstractActive or deformable models have emerged as a popular modeling paradigm in computer vision. These models have the flexibility to adapt themselves to the image data, offering the potential for both generic object recognition and non-rigid object tracking. Because these active models are underconstrained, however, deformable shape recovery often requires manual segmentation or good model initialization, while active contour trackers have been able to track only an object's translation in the image. In this paper, we report our current progress in using a part-based aspect graph representation of an object14 to provide the missing constraints on data-driven deformable model recovery and tracking processes. Sven J. Dickinson |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1997 | Panel report: the potential of geons for generic 3-D object recognition
Sven J. Dickinson, Robert Bergevin, Irving Biederman, Jan-Olof Eklundh, Roger Munck-Fairwood, Anil K. Jain 0001, Alex Pentland |
Image Vis. Comput. | 1 |
| 1997 | The Role of Model-Based Segmentation in the Recovery of Volumetric Parts From Range DataabstractWe present a method for segmenting and estimating the shape of 3D objects from range data. The technique uses model views, or aspects, to constrain the fitting of deformable models to range data. Based on an initial region segmentation of a range image, regions are grouped into aspects corresponding to the volumetric parts that make up an object. The qualitative segmentation of the range image into a set of volumetric parts not only captures the coarse shape of the parts, but qualitatively encodes the orientation of each part through its aspect. Knowledge of a part's coarse shape, its orientation, as well as the mapping between the faces in its aspect and the surfaces on the part provides strong constraints on the fitting of a deformable model (supporting both global and local deformations) to the data. Unlike previous work in physics-based deformable model recovery from range data, the technique does not require presegmented data. Furthermore, occlusion is handled at segmentation time and does not complicate the fitting process, as only 3D points known to belong to a part participate in the fitting of a model to the part. We present the approach in detail and apply it to the recovery of objects from range data. Sven J. Dickinson, Dimitris N. Metaxas, Alex Pentland |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | A Quantitative Analysis of View Degeneracy and its use for Active Focal Length controlabstractWe quantify the observation by Kender and Freudenstein (1987) that degenerate views occupy a significant fraction of the viewing sphere surrounding an object. This demonstrates that systems for recognition must explicitly account for the possibility of view degeneracy. We show that view degeneracy cannot be detected from a single camera viewpoint. As a result, systems designed to recognize objects from a single arbitrary viewpoint must be able to function in spite of possible undetected degeneracies, or else operate with imaging parameters that cause acceptably low probabilities of degeneracy. To address this need, we give a prescription for active control of focal length that allows a principled tradeoff between the camera field of view and probability of view degeneracy.> David Wilkes, Sven J. Dickinson, John K. Tsotsos |
ICCV | 2 |
| 1995 | Recognition by Functional Parts
Ehud Rivlin, Sven J. Dickinson, Azriel Rosenfeld |
Comput. Vis. Image Underst. | 2 |
| 1994 | A New Approach to Tracking 3D Objects in 2D Image Sequences
Michael Chan 0001, Dimitris N. Metaxas, Sven J. Dickinson |
AAAI | 3 |
| 1994 | Qualitative tracking of 3-D objects using active contour networksabstractIn this paper, we track changes in the appearance of the object as it moves from one frame to the next. At a symbolic level, an aspect graph clusters all the views of an object into a set of topologically distinct classes in terms of which surfaces of an object are visible from a given viewpoint (Koenderink and van Doom (1979). Two nodes (or aspects) in the aspect graph are connected by an arc if it is possible to directly move from a viewpoint in which the first aspect is visible to a viewpoint in which the second aspect is visible. Qualitatively, we can envision a tracking strategy which simply tracks an object as it moves from one node to another in the object's aspect graph. Although it does not provide us with accurate pose of the object, it does qualitatively describe the motion of the object without the need for a CAD representation of the object.> Sven J. Dickinson, Piotr Jasiobedzki, Göran Olofsson, Henrik I. Christensen |
CVPR | 1 |
| 1994 | Recognition by functional parts [function-based object recognition]abstractWe present an approach to function-based object recognition that reasons about the functionality of an object's initiative parts. We extend the popular "recognition by parts" shape recognition framework to support "recognition, by functional parts", by combining a set of functional primitives and their relations with a set of abstract volumetric shape primitives and their relations. Previous approaches have relied on more global object features, often ignoring the problem of object segmentation, and thereby restricting themselves to range images of unoccluded scenes. We show how these shape primitives and relations can be easily recovered from superquadric ellipsoids which, in turn, can be recovered from either range or intensity images of occluded scenes. Furthermore, the proposed framework supports both unexpected (bottom-up) object recognition and expected (top-down) object recognition. We demonstrate the approach on, a simple domain by recognizing a restricted class of hand-tools from 2-D images.> Ehud Rivlin, Sven J. Dickinson, Azriel Rosenfeld |
CVPR | 2 |
| 1994 | Active Object Recognition Integrating Attention and Viewpoint Control
Sven J. Dickinson, Henrik I. Christensen, John K. Tsotsos, Göran Olofsson |
ECCV (2) | 1 |
| 1994 | Physics-based tracking of 3D objects in 2D image sequencesabstractWe present a new technique for tracking 3D objects in 2D image sequences. We assume that objects are constructed from a class of volumetric part primitives. The models are initially recovered using a qualitative shape recovery process. We subsequently track the objects using local forces computed from image potentials. Therefore we avoid the expensive computation of image features. By integrating measurements from stereo images, 3D positions (as well as other model parameters) of the objects can be continuously updated using an extended Kalman filter. Our model-based approach can handle occlusions in scenes with multiple moving objects by predicting their occurrences. To handle severe or unexpected occlusion we use a feedback mechanism between the quantitative and qualitative shape estimation systems. We demonstrate our technique in experiments involving image sequences from complex motions of objects. Michael Chan 0001, Dimitris N. Metaxas, Sven J. Dickinson |
ICPR (1) | 3 |
| 1994 | Navigation based on a network of 2D imagesabstractThis paper describes the integration of 2D stimulus-driven robot localization and positioning with a token-based correspondence method in a practical robot navigation system. The approach allows for modular acquisition and update of world knowledge for navigation, and robustness of navigation to low-level errors. No special marking of the world is necessary, so the robot may operate in quite general environments. Tests in a real industrial environment confirm the potential of the method. David Wilkes, Sven J. Dickinson, Ehud Rivlin, Ronen Basri |
ICPR (1) | 2 |
| 1994 | Integrating qualitative and quantitative shape recovery
Sven J. Dickinson, Dimitris N. Metaxas |
Int. J. Comput. Vis. | 1 |
| 1993 | Integration of quantitative and qualitative techniques for deformable model fitting from orthographic, perspective, and stereo projectionsabstractThe authors synthesize a new approach to 3-D object shape recovery by integrating qualitative shape recovery techniques and quantitative physics-based shape estimation techniques. They first use qualitative shape recovery and recognition techniques to provide strong fitting constraints on physics-based deformable model recovery techniques. Previously developed techniques of fitting deformable models to occluding image contours are then extended to the case of image data captured under general orthographic, perspective, and stereo projections. Experimental results are presented to illustrate the shape recovery approach.> Dimitris N. Metaxas, Sven J. Dickinson |
ICCV | 2 |
| 1993 | The Use of Geons for Generic 3D Object Recognition
Sven J. Dickinson, Robert Bergevin, Irving Biederman, Jan-Olof Eklundh, Roger Munck-Fairwood, Alex Pentland |
IJCAI | 1 |
| 1992 | From volumes to views: An approach to 3-D object recognition
Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld |
CVGIP Image Underst. | 1 |
| 1992 | 3-D Shape Recovery Using Distributed Aspect MatchingabstractAn approach to the recovery of 3-D volumetric primitives from a single 2-D image is presented. The approach first takes a set of 3-D volumetric modeling primitives and generates a hierarchical aspect representation based on the projected surfaces of the primitives; conditional probabilities capture the ambiguity of mappings between levels of the hierarchy. From a region segmentation of the input image, the authors present a formulation of the recovery problem based on the grouping of the regions into aspects. No domain-independent heuristics are used; only the probabilities inherent in the aspect hierarchy are exploited. Once the aspects are recovered, the aspect hierarchy is used to infer a set of volumetric primitives and their connectivity. As a front end to an object recognition system, the approach provides the indexing power of complex 3-D object-centered primitives while exploiting the convenience of 2-D viewer-centered aspect matching; aspects are used to represent a finite vocabulary of 3-D parts from which objects can be constructed.> Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Qualitative 3-D shape reconstruction using distributed aspect graph matchingabstractAn approach is presented to 3-D primitive reconstruction that is independent of the selection of volumetric primitives used to model objects. The approach first takes an arbitrary set of 3-D volumetric primitives and generates a hierarchical aspect representation based on the projected surfaces of the primitives; conditional probabilities capture the ambiguity of mappings between levels of the hierarchy. The integration of object-centered and viewer-centered representations provides the indexing power of 3-D volumetric primitives, while supporting a 2-D matching paradigm for primitive reconstruction. Formulation of the problem based on grouping the image regions according to aspect is presented. No domain dependent heuristics are used; the authors exploit only the probabilities inherent in the aspect hierarchy. For a given selection of primitives, the success of the heuristic depends on the likelihood of the various aspects; best results are achieved when certain aspects are more likely, and fewer primitives project to a given aspect.> Sven J. Dickinson, Alex Pentland, Azriel Rosenfeld |
ICCV | 1 |
| 1990 | A flexible tool for prototyping ALV road following algorithmsabstractA production system model of problem-solving is applied to the design of a vision system by which an autonomous land vehicle (ALV) navigates roads. The ALV vision task consists of hypothesizing objects in a scene model and verifying these hypotheses using the vehicle's sensors. Object hypothesis generation is based on the local navigation task, an a priori road map, and the contents of the scene model. Verification of an object hypothesis involves directing the sensors towards the expected location of the object, collecting evidence in support of the object, and reasoning about the evidence. Constructing the scene model consists of building a semantic network of object frames exhibiting component, spatial, and inheritance relationships. The control structure is provided by a set of communicating production systems implementing a structured blackboard; each production system contains rules for defining the attributes of a particular class of object frame. The combination of production system and object-oriented programming techniques results in a flexible control structure able to accommodate new object classes, reasoning strategies, vehicle sensors, and image analysis techniques.> Sven J. Dickinson, Larry Davis 0001 |
IEEE Trans. Robotics Autom. | 1 |
| 1988 | An expert vision system for autonomous land vehicle road followingabstractA production-system model of problem solving is applied to the design of a vision system by which an autonomous land vehicle (ALV) navigates roads. The ALV vision task consists of hypothesizing objects in a scene model and verifying these hypotheses using the vehicles sensors. Object hypothesis generation is based on the local navigation task, and a priori road map, and the contents of the scene model. Verification of an object hypothesis involves directing the sensors toward the expected location of the object, collecting evidence in support of the object, and reasoning about the evidence. Constructing the scene model consists of building a semantic network of object frames exhibiting component, spatial, and inheritance relationships. The control structure is provided by a set of communicating production systems implementing a structured blackboard; each production system contains the rules for defining the attributes of a particular class of object frame.> Sven J. Dickinson, Larry Davis 0001 |
CVPR | 1 |