Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daniel P. Huttenlocher

dblp:h/DPHuttenlocher · also Dan Huttenlocher · DBLP profile ↗
← Back
81ranked-venue papers
24as first author
0since 2021 · last 2015
0009-0005-2026-3064ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 15 first-authorGraphics, computer vision, multimedia, augmented reality and games · 44 · 15 first-authorDatabases, data management, data science and information retrieval · 13 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 10Theory of computation · 6 · 5 first-authorHuman-computer interaction and ubiquitous computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
45 papers
3D vision · 37% Image recognition and object detection · 19% Probabilistic and Bayesian machine learning · 13%
Databases, data mining, and information retrieval
10 papers
Web and social media mining · 88% Knowledge graphs · 7% Information retrieval · 3%
Computer graphics and multimedia
14 papers
Image and video processing · 61% Multimedia analysis and retrieval · 25% Computational photography and imaging · 12%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Computational social science and digital humanities · 74% Computing education · 24% Bioinformatics and computational biology · 2%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 33% Games and playful interaction · 33% Human-AI interaction · 33%

Topics — the 30 heaviest of 139, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › structure from motion
bundle adjustment
0.322013
SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from Motion · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Discrete-continuous optimization for large-scale structure from motion · CVPR 2011
Machine learning › Optimization for machine learning
discrete-continuous optimization
0.322013
SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from Motion · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Discrete-continuous optimization for large-scale structure from motion · CVPR 2011
Computer vision › 3D vision
structure from motion
0.322013
SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from Motion · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Discrete-continuous optimization for large-scale structure from motion · CVPR 2011
Computer vision › Image recognition and object detection
object recognition
0.3182006
Weakly Supervised Learning of Part-Based Spatial Models for Visual Object Recognition · ECCV (1) 2006
Pictorial Structures for Object Recognition · Int. J. Comput. Vis. 2005
A New Bayesian Framework for Object Recognition · CVPR 1999
Web and social media mining
social media analysis
0.332012
Discovering value from community activity on focused question answering sites: a case study of stack overflow · KDD 2012
Feedback effects between similarity and social influence in online communities · KDD 2008
Mapping the world's photos · WWW 2009
Web and social media mining
social network analysis
0.332010
Predicting positive and negative links in online social networks · WWW 2010
Feedback effects between similarity and social influence in online communities · KDD 2008
Group formation in large social networks: membership, growth, and evolution · KDD 2006
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.232013
SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from Motion · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Efficient Belief Propagation with Learned Higher-Order Markov Random Fields · ECCV (2) 2006
A New Bayesian Framework for Object Recognition · CVPR 1999
Computational social science and digital humanities › social network analysis
information diffusion
0.212015
Global Diffusion via Cascading Invitations: Structure, Growth, and Homophily · WWW 2015
Computational social science and digital humanities
social network analysis
0.212015
Global Diffusion via Cascading Invitations: Structure, Growth, and Homophily · WWW 2015
Web and social media mining › social network analysis
structural balance theory
0.222010
Predicting positive and negative links in online social networks · WWW 2010
Signed networks in social media · CHI 2010
Computing education › online education
massive open online courses
0.212014
Engaging with massive online courses · WWW 2014
Computational social science and digital humanities
social media analysis
0.222012
Effects of user similarity in social media · WSDM 2012
Signed networks in social media · CHI 2010
Games and playful interaction
gamification
0.212013
Steering user behavior with badges · WWW 2013
Collaborative and social computing
online communities
0.212013
Steering user behavior with badges · WWW 2013
Computer vision › 3D vision
camera pose estimation
0.222012
Worldwide Pose Estimation Using 3D Point Clouds · ECCV (1) 2012
A study of affine matching with bounded sensor error · Int. J. Comput. Vis. 1994
Computational social science and digital humanities › social computing
online community analysis
0.112012
Discovering value from community activity on focused question answering sites: a case study of stack overflow · KDD 2012
Web and social media mining › trust assessment
reputation system
0.112012
Discovering value from community activity on focused question answering sites: a case study of stack overflow · KDD 2012
Image and video processing
image restoration
0.122008
Sparse Long-Range Random Field and Its Application to Image Denoising · ECCV (3) 2008
Efficient Belief Propagation for Early Vision · CVPR (1) 2004
Computer vision › Image recognition and object detection › image classification
object classification
0.122007
Composite Models of Objects and Scenes for Category Recognition · CVPR 2007
Spatial Priors for Part-Based Recognition Using Statistical Models · CVPR (1) 2005
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.112011
Efficient Unbiased Tracking of Multiple Dynamic Obstacles Under Large Viewpoint Changes · IEEE Trans. Robotics 2011
Computer vision › 3D vision › structure from motion
large-scale reconstruction
0.112011
Discrete-continuous optimization for large-scale structure from motion · CVPR 2011
Computer vision › Video understanding and tracking
multi-object tracking
0.112011
Efficient Unbiased Tracking of Multiple Dynamic Obstacles Under Large Viewpoint Changes · IEEE Trans. Robotics 2011
Image and video processing › motion analysis
human motion analysis
0.112011
Upper Body Detection and Tracking in Extended Signing Sequences · Int. J. Comput. Vis. 2011
Computer vision › 3D vision › visual localization
location recognition
0.112010
Location Recognition Using Prioritized Feature Matching · ECCV (2) 2010
Knowledge graphs › link prediction
signed link prediction
0.112010
Predicting positive and negative links in online social networks · WWW 2010
Web and social media mining › social network analysis
signed social networks
0.112010
Signed networks in social media · CHI 2010
Image and video processing › image restoration
image deblurring
0.112010
Generating sharp panoramas from motion-blurred videos · CVPR 2010
Image and video processing › image restoration › image deblurring
multi-image deblurring
0.112010
Generating sharp panoramas from motion-blurred videos · CVPR 2010
Computational photography and imaging › panoramic imaging
panorama generation
0.112010
Generating sharp panoramas from motion-blurred videos · CVPR 2010
Computer vision › 3D vision › stereo vision
stereo matching
0.122008
Learning for stereo vision using the structured support vector machine · CVPR 2008
Computing Visual Correspondence: Incorporating the Probability of a False Match · ICCV 1995

Methods — techniques the papers use, named apart from their topics

homophily measurement · 0.4cascade analysis · 0.4badge systems · 0.3markov random field · 0.3social psychology theory · 0.2large-scale dataset analysis · 0.2temporal features · 0.2structural analysis · 0.2content analysis · 0.2classification · 0.2survey · 0.2qualitative analysis · 0.2levenberg-marquardt · 0.2discrete-continuous optimization · 0.2graphical model · 0.1rao-blackwellized particle filter · 0.1parametric filters · 0.1levenberg-marquardt refinement · 0.1
YearPublicationVenuePosition
2015 Coordination and Efficiency in Decentralized Collaboration
Daniel M. Romero, Daniel P. Huttenlocher, Jon M. Kleinberg
ICWSM2
2015 Global Diffusion via Cascading Invitations: Structure, Growth, and Homophily
abstract
Many of the world's most popular websites catalyze their growth through invitations from existing members. New members can then in turn issue invitations, and so on, creating cascades of member signups that can spread on a global scale. Although these diffusive invitation processes are critical to the popularity and growth of many websites, they have rarely been studied, and their properties remain elusive. For instance, it is not known how viral these cascades structures are, how cascades grow over time, or how diffusive growth affects the resulting distribution of member characteristics present on the site. In this paper, we study the diffusion of LinkedIn, an online professional network comprising over 332 million members, a large fraction of whom joined the site as part of a signup cascade. First we analyze the structural patterns of these signup cascades, and find them to be qualitatively different from previously studied information diffusion cascades. We also examine how signup cascades grow over time, and observe that diffusion via invitations on LinkedIn occurs over much longer timescales than are typically associated with other types of online diffusion. Finally, we connect the cascade structures with rich individual-level attribute data to investigate the interplay between the two. Using novel techniques to study the role of homophily in diffusion, we find striking differences between the local, edge-wise homophily and the global, cascade-level homophily we observe in our data, suggesting that signup cascades form surprisingly coherent groups of members.
Ashton Anderson, Daniel P. Huttenlocher, Jon M. Kleinberg, Jure Leskovec, Mitul Tiwari
WWW2
2014 Engaging with massive online courses
abstract
The Web has enabled one of the most visible recent developments in education---the deployment of massive open online courses. With their global reach and often staggering enrollments, MOOCs have the potential to become a major new mechanism for learning. Despite this early promise, however, MOOCs are still relatively unexplored and poorly understood.
Ashton Anderson, Daniel P. Huttenlocher, Jon M. Kleinberg, Jure Leskovec
WWW2
2013 Steering user behavior with badges
abstract
An increasingly common feature of online communities and social media sites is a mechanism for rewarding user achievements based on a system of badges. Badges are given to users for particular contributions to a site, such as performing a certain number of actions of a given type. They have been employed in many domains, including news sites like the Huffington Post, educational sites like Khan Academy, and knowledge-creation sites like Wikipedia and Stack Overflow. At the most basic level, badges serve as a summary of a user's key accomplishments; however, experience with these sites also shows that users will put in non-trivial amounts of work to achieve particular badges, and as such, badges can act as powerful incentives. Thus far, however, the incentive structures created by badges have not been well understood, making it difficult to deploy badges with an eye toward the incentives they are likely to create.
Ashton Anderson, Daniel P. Huttenlocher, Jon M. Kleinberg, Jure Leskovec
WWW2
2013 SfM with MRFs: Discrete-Continuous Optimization for Large-Scale Structure from Motion
abstract
Recent work in structure from motion (SfM) has built 3D models from large collections of images downloaded from the Internet. Many approaches to this problem use incremental algorithms that solve progressively larger bundle adjustment problems. These incremental techniques scale poorly as the image collection grows, and can suffer from drift or local minima. We present an alternative framework for SfM based on finding a coarse initial solution using hybrid discrete-continuous optimization and then improving that solution using bundle adjustment. The initial optimization step uses a discrete Markov random field (MRF) formulation, coupled with a continuous Levenberg-Marquardt refinement. The formulation naturally incorporates various sources of information about both the cameras and points, including noisy geotags and vanishing point (VP) estimates. We test our method on several large-scale photo collections, including one with measured camera positions, and show that it produces models that are similar to or better than those produced by incremental bundle adjustment, but more robustly and in a fraction of the time.
David Crandall, Andrew Owens, Noah Snavely, Daniel P. Huttenlocher
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Worldwide Pose Estimation Using 3D Point Clouds
Yunpeng Li 0002, Noah Snavely, Daniel P. Huttenlocher, Pascal Fua
ECCV (1)3
2012 Discovering value from community activity on focused question answering sites: a case study of stack overflow
abstract
Question answering (Q&A) websites are now large repositories of valuable knowledge. While most Q&A sites were initially aimed at providing useful answers to the question asker, there has been a marked shift towards question answering as a community-driven knowledge creation process whose end product can be of enduring value to a broad audience. As part of this shift, specific expertise and deep knowledge of the subject at hand have become increasingly important, and many Q&A sites employ voting and reputation mechanisms as centerpieces of their design to help users identify the trustworthiness and accuracy of the content.
Ashton Anderson, Daniel P. Huttenlocher, Jon M. Kleinberg, Jure Leskovec
KDD2
2012 Effects of user similarity in social media
abstract
There are many settings in which users of a social media application provide evaluations of one another. In a variety of domains, mechanisms for evaluation allow one user to say whether he or she trusts another user, or likes the content they produced, or wants to confer special levels of authority or responsibility on them. Earlier work has studied how the relative status between two users - that is, their comparative levels of status in the group - affects the types of evaluations that one user gives to another.
Ashton Anderson, Daniel P. Huttenlocher, Jon M. Kleinberg, Jure Leskovec
WSDM2
2011 Discrete-continuous optimization for large-scale structure from motion
abstract
Recent work in structure from motion (SfM) has successfully built 3D models from large unstructured collections of images downloaded from the Internet. Most approaches use incremental algorithms that solve progressively larger bundle adjustment problems. These incremental techniques scale poorly as the number of images grows, and can drift or fall into bad local minima. We present an alternative formulation for SfM based on finding a coarse initial solution using a hybrid discrete-continuous optimization, and then improving that solution using bundle adjustment. The initial optimization step uses a discrete Markov random field (MRF) formulation, coupled with a continuous Levenberg-Marquardt refinement. The formulation naturally incorporates various sources of information about both the cameras and the points, including noisy geotags and vanishing point estimates. We test our method on several large-scale photo collections, including one with measured camera positions, and show that it can produce models that are similar to or better than those produced with incremental bundle adjustment, but more robustly and in a fraction of the time.
David Crandall, Andrew Owens, Noah Snavely, Daniel P. Huttenlocher
CVPR4
2011 Upper Body Detection and Tracking in Extended Signing Sequences
Patrick Buehler, Mark Everingham, Daniel P. Huttenlocher, Andrew Zisserman
Int. J. Comput. Vis.3
2011 Efficient Unbiased Tracking of Multiple Dynamic Obstacles Under Large Viewpoint Changes
abstract
A novel-tracking algorithm is presented as a computationally feasible, real-time solution to the joint estimation problem of data assignment and dynamic obstacle tracking from a potentially moving robotic platform. The algorithm implements a Rao-Blackwellized particle filter (RBPF) to factorize the joint estimation problem into 1) a data assignment problem solved via particle filter and 2) a multiple dynamic obstacle-tracking problem solved with efficient parametric filters. The parametric filters make use of a new target representation and stable features developed specifically for tracking full-size vehicles in a dense traffic environment. The algorithm is validated in real time, both in controlled experiments with full-size robotic vehicles and on data collected at the 2007 Defense Advanced Research Projects Agency (DARPA) Urban Challenge.
Isaac Miller, Mark E. Campbell, Daniel P. Huttenlocher
IEEE Trans. Robotics3
2010 Signed networks in social media
abstract
Relations between users on social media sites often reflect a mixture of positive (friendly) and negative (antagonistic) interactions. In contrast to the bulk of research on social networks that has focused almost exclusively on positive interpretations of links between people, we study how the interplay between positive and negative relationships affects the structure of on-line social networks. We connect our analyses to theories of signed networks from social psychology. We find that the classical theory of structural balance tends to capture certain common patterns of interaction, but that it is also at odds with some of the fundamental phenomena we observe --- particularly related to the evolving, directed nature of these on-line networks. We then develop an alternate theory of status that better explains the observed edge signs and provides insights into the underlying social mechanisms. Our work provides one of the first large-scale evaluations of theories of signed networks using on-line datasets, as well as providing a perspective for reasoning about social media sites.
Jure Leskovec, Daniel P. Huttenlocher, Jon M. Kleinberg
CHI2
2010 Generating sharp panoramas from motion-blurred videos
abstract
In this paper, we show how to generate a sharp panorama from a set of motion-blurred video frames. Our technique is based on joint global motion estimation and multi-frame deblurring. It also automatically computes the duty cycle of the video, namely the percentage of time between frames that is actually exposure time. The duty cycle is necessary for allowing the blur kernels to be accurately extracted and then removed. We demonstrate our technique on a number of videos.
Yunpeng Li 0002, Sing Bing Kang, Neel Joshi, Steven M. Seitz, Daniel P. Huttenlocher
CVPR5
2010 Location Recognition Using Prioritized Feature Matching
Yunpeng Li 0002, Noah Snavely, Daniel P. Huttenlocher
ECCV (2)3
2010 Sequential Influence Models in Social Networks
Dan Cosley, Daniel P. Huttenlocher, Jon M. Kleinberg, Xiangyang Lan, Siddharth Suri
ICWSM2
2010 Governance in Social Media: A Case Study of the Wikipedia Promotion Process
Jure Leskovec, Daniel P. Huttenlocher, Jon M. Kleinberg
ICWSM2
2010 Predicting positive and negative links in online social networks
abstract
We study online social networks in which relationships can be either positive (indicating relations such as friendship) or negative (indicating relations such as opposition or antagonism). Such a mix of positive and negative links arise in a variety of online settings; we study datasets from Epinions, Slashdot and Wikipedia. We find that the signs of links in the underlying social networks can be predicted with high accuracy, using models that generalize across this diverse range of sites. These models provide insight into some of the fundamental principles that drive the formation of signed links in networks, shedding light on theories of balance and status from social psychology; they also suggest social computing applications by which the attitude of one user toward another can be estimated from evidence provided by their relationships with other members of the surrounding social network.
Jure Leskovec, Daniel P. Huttenlocher, Jon M. Kleinberg
WWW2
2009 Landmark classification in large-scale image collections
abstract
With the rise of photo-sharing websites such as Facebook and Flickr has come dramatic growth in the number of photographs online. Recent research in object recognition has used such sites as a source of image data, but the test images have been selected and labeled by hand, yielding relatively small validation sets. In this paper we study image classification on a much larger dataset of 30 million images, including nearly 2 million of which have been labeled into one of 500 categories. The dataset and categories are formed automatically from geotagged photos from Flickr, by looking for peaks in the spatial geotag distribution corresponding to frequently-photographed landmarks. We learn models for these landmarks with a multiclass support vector machine, using vector-quantized interest point descriptors as features. We also explore the non-visual information available on modern photo-sharing sites, showing that using textual tags and temporal constraints leads to significant improvements in classification rate. We find that in some cases image features alone yield comparable classification accuracy to using text tags as well as to the performance of human observers.
Yunpeng Li 0002, David Crandall, Daniel P. Huttenlocher
ICCV3
2009 Mapping the world's photos
abstract
We investigate how to organize a large collection of geotagged photos, working with a dataset of about 35 million images collected from Flickr. Our approach combines content analysis based on text tags and image data with structural analysis based on geospatial data. We use the spatial distribution of where people take photos to define a relational structure between the photos that are taken at popular places. We then study the interplay between this structure and the content, using classification methods for predicting such locations from visual, textual and temporal features of the photos. We find that visual and temporal features improve the ability to estimate the location of a photo, compared to using just textual features. We illustrate using these techniques to organize a large photo collection, while also revealing various interesting properties about popular cities and landmarks at a global scale.
David Crandall, Lars Backstrom, Daniel P. Huttenlocher, Jon M. Kleinberg
WWW3
2008 Long Term Arm and Hand Tracking for Continuous Sign Language TV Broadcasts
abstract
The goal of this work is to detect hand and arm positions over continuous sign language video sequences of more than one hour in length. We cast the problem as inference in a generative model of the image. Un-der this model, limb detection is expensive due to the very large number of possible configurations each part can assume. We make the following con-tributions to reduce this cost: (i) using efficient sampling from a pictorial structure proposal distribution to obtain reasonable configurations; (ii) iden-tifying a large set of frames where correct configurations can be inferred, and using temporal tracking elsewhere. Results are reported for signing footage with changing background, chal-lenging image conditions, and different signers; and we show that the method is able to identify the true arm and hand locations. The results exceed the state-of-the-art for the length and stability of continuous limb tracking. 1
Patrick Buehler, Mark Everingham, Daniel P. Huttenlocher, Andrew Zisserman
BMVC3
2008 Learning for stereo vision using the structured support vector machine
abstract
We present a random field based model for stereo vision with explicit occlusion labeling in a probabilistic framework. The model employs non-parametric cost functions that can be learnt automatically using the structured support vector machine. The learning algorithm enables the training of models that are steered towards optimizing for a particular desired loss function, such as the metric used to evaluate the quality of the stereo labeling. Experimental results demonstrate that the performance of our method surpasses that of previous learning approaches and is comparable to the state-of-the-art for pixel-based stereo. Moreover, our method achieves good results even when trained on different image sets, in contrast with the common practice of hand tuning to specific benchmark images. In addition, we investigate the impact of graph structure on model performance. Our study shows that random field models with longer-range edges generally outperform the 4-connected grid and that this advantage is especially pronounced for noisy images.
Yunpeng Li 0002, Daniel P. Huttenlocher
CVPR2
2008 Learning for Optical Flow Using Stochastic Optimization
Yunpeng Li 0002, Daniel P. Huttenlocher
ECCV (2)2
2008 Sparse Long-Range Random Field and Its Application to Image Denoising
Yunpeng Li 0002, Daniel P. Huttenlocher
ECCV (3)2
2008 Feedback effects between similarity and social influence in online communities
abstract
A fundamental open question in the analysis of social networks is to understand the interplay between similarity and social ties. People are similar to their neighbors in a social network for two distinct reasons: first, they grow to resemble their current friends due to social influence; and second, they tend to form new links to others who are already like them, a process often termed selection by sociologists. While both factors are present in everyday social processes, they are in tension: social influence can push systems toward uniformity of behavior, while selection can lead to fragmentation. As such, it is important to understand the relative effects of these forces, and this has been a challenge due to the difficulty of isolating and quantifying them in real settings.
David Crandall, Dan Cosley, Daniel P. Huttenlocher, Jon M. Kleinberg, Siddharth Suri
KDD3
2007 Composite Models of Objects and Scenes for Category Recognition
abstract
This paper presents a method of learning and recognizing generic object categories using part-based spatial models. The models are multiscale, with a scene component that specifies relationships between the object and surrounding scene context, and an object component that specifies relationships between parts of the object. The underlying graphical model forms a tree structure, with a star topology for both the contextual and object components. A partially supervised paradigm is used for learning the models, where each training image is labeled with bounding boxes indicating the overall location of object instances, but parts or regions of the objects and scene are not specified. The parts, regions and spatial relationships are learned automatically. We demonstrate the method on the detection task on the PASCAL 2006 Visual Object Classes Challenge dataset, where objects must be correctly localized. Our results demonstrate better overall performance than those of previously reported techniques, in terms of the average precision measure used in the PASCAL detection evaluation. Our results also show that incorporating scene context into the models improves performance in comparison with not using such contextual information.
David Crandall, Daniel P. Huttenlocher
CVPR2
2006 Weakly Supervised Learning of Part-Based Spatial Models for Visual Object Recognition
David Crandall, Daniel P. Huttenlocher
ECCV (1)2
2006 Efficient Belief Propagation with Learned Higher-Order Markov Random Fields
Xiangyang Lan, Stefan Roth 0001, Daniel P. Huttenlocher, Michael J. Black
ECCV (2)3
2006 Group formation in large social networks: membership, growth, and evolution
abstract
The processes by which communities come together, attract new members, and develop over time is a central research issue in the social sciences - political movements, professional organizations, and religious denominations all provide fundamental examples of such communities. In the digital domain, on-line groups are becoming increasingly prominent due to the growth of community and social networking sites such as MySpace and LiveJournal. However, the challenge of collecting and analyzing large-scale time-resolved data on social groups and communities has left most basic questions about the evolution of such groups largely unresolved: what are the structural features that influence whether individuals will join communities, which communities will grow rapidly, and how do the overlaps among pairs of communities change over time.Here we address these questions using two large sources of data: friendship links and community membership on LiveJournal, and co-authorship and conference publications in DBLP. Both of these datasets provide explicit user-defined communities, where conferences serve as proxies for communities in DBLP. We study how the evolution of these communities relates to properties such as the structure of the underlying social networks. We find that the propensity of individuals to join communities, and of communities to grow rapidly, depends in subtle ways on the underlying network structure. For example, the tendency of an individual to join a community is influenced not just by the number of friends he or she has within the community, but also crucially by how those friends are connected to one another. We use decision-tree techniques to identify the most significant structural determinants of these properties. We also develop a novel methodology for measuring movement of individuals between communities, and show how such movements are closely aligned with changes in the topics of interest within the communities.
Lars Backstrom, Daniel P. Huttenlocher, Jon M. Kleinberg, Xiangyang Lan
KDD2
2006 Efficient Belief Propagation for Early Vision
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
Int. J. Comput. Vis.2
2005 Spatial Priors for Part-Based Recognition Using Statistical Models
abstract
We present a class of statistical models for part-based object recognition that are explicitly parameterized according to the degree of spatial structure they can represent. These models provide a way of relating different spatial priors that have been used for recognizing generic classes of objects, including joint Gaussian models and tree-structured models. By providing explicit control over the degree of spatial structure, our models make it possible to study the extent to which additional spatial constraints among parts are actually helpful in detection and localization, and to consider the tradeoff in representational power and computational cost. We consider these questions for object classes that have substantial geometric structure, such as airplanes, faces and motorbikes, using datasets employed by other researchers to facilitate evaluation. We find that for these classes of objects, a relatively small amount of spatial structure in the model can provide statistically indistinguishable recognition performance from more powerful models, and at a substantially lower computational cost.
David Crandall, Pedro F. Felzenszwalb, Daniel P. Huttenlocher
CVPR (1)3
2005 Beyond Trees: Common-Factor Models for 2D Human Pose Recovery
abstract
Tree structured models have been widely used for determining the pose of a human body, from either 2D or 3D data. While such models can effectively represent the kinematic constraints of the skeletal structure, they do not capture additional constraints such as coordination of the limbs. Tree structured models thus miss an important source of information about human body pose, as limb coordination is necessary for balance while standing, walking, or running, as well as being evident in other activities such as dancing and throwing. In this paper, we consider the use of undirected graphical models that augment a tree structure with latent variables in order to account for coordination between limbs. We refer to these as common-factor models, since they are constructed by using factor analysis to identify additional correlations in limb position that are not accounted for by the kinematic tree structure. These common-factor models have an underlying tree structure and thus a variant of the standard Viterbi algorithm for a tree can be applied for efficient estimation. We present some experimental results contrasting common-factor models with tree models, and quantify the improvement in pose estimation for 2D image data.
Xiangyang Lan, Daniel P. Huttenlocher
ICCV2
2005 Pictorial Structures for Object Recognition
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
Int. J. Comput. Vis.2
2004 Efficient Belief Propagation for Early Vision
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
CVPR (1)2
2004 A Unified Spatio-Temporal Articulated Model for Tracking
Xiangyang Lan, Daniel P. Huttenlocher
CVPR (1)2
2004 Efficient Graph-Based Image Segmentation
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
Int. J. Comput. Vis.2
2003 Fast Algorithms for Large-State-Space HMMs with Applications to Web Usage Analysis
abstract
In applying Hidden Markov Models to the analysis of massive data streams, it is often necessary to use an arti(cid:12)cially reduced set of states; this is due in large part to the fact that the basic HMM estimation algorithms have a quadratic dependence on the size of the state set. We present algorithms that reduce this computational bottleneck to linear or near-linear time, when the states can be embedded in an underlying grid of parameters. This type of state representation arises in many domains; in particular, we show an application to tra(cid:14)c analysis at a high-volume Web site.
Pedro F. Felzenszwalb, Daniel P. Huttenlocher, Jon M. Kleinberg
NIPS2
2000 Adaptive Bayesian Recognition in Tracking Rigid Objects
abstract
We present a framework for tracking rigid objects based on an adaptive Bayesian recognition technique that incorporates dependencies between object features. At each frame we find a maximum a posteriori (MAP) estimate of the object parameters that include positioning and configuration of non-occluded features. This estimate may be rejected based on its quality. Our careful selection of data points in each frame allows temporal fusion via Kalman filtering. Despite "unimodality" of our tracking scheme, we demonstrate fairly robust results in highly cluttered aerial scenes. Our technique forms a natural feedback loop between the recognition method and the filter that helps to explain such robustness. We study this loop and derive a number of interesting properties. First, the effective threshold for recognition in each frame is adaptive. It depends on the current level of noise in the system. This allows the system to identify partially occluded or distorted objects as long as the predicted locations are accurate. But requires a very good match if there is uncertainty as to the object location. Second, the search area for the recognition method is automatically pruned based on the current system uncertainty, yielding an efficient overall method.
Yuri Boykov, Daniel P. Huttenlocher
CVPR2
2000 Efficient Matching of Pictorial Structures
abstract
A pictorial structure is a collection of parts arranged in a deformable configuration. Each part is represented using a simple appearance model and the deformable configuration is represented by spring-like connections between pairs of parts. While pictorial structures were introduced a number of years ago, they have not been broadly applied to matching and recognition problems. This has been due in part to the computational difficulty of matching pictorial structures to images. In this paper we present an efficient algorithm for finding the best global match of a pictorial stucture to an image. With this improved algorithm, pictorial structures provide a practical and powerful framework for quantitative descriptions of objects and scenes, and are suitable for many generic image recognition problems. We illustrate the approach using simple models of a person and a car.
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
CVPR2
2000 Integrating Color, Texture, and Geometry for Image Retrieval
abstract
This paper examines the problem of image retrieval from large, heterogeneous image databases. We present a technique that fulfils several needs identified by surveying recent research in the field. This technique fairly integrates a diverse and expandable set of image properties (for example, color, texture, and location) in a retrieval framework, and allows end-users substantial control over their use. We propose a novel set of evaluation methods in addition to applying established tests for image retrieval; our technique proves competitive with state-of-the-art methods in these tests and does better on certain tasks. Furthermore, it improves on many standard image retrieval algorithms by supporting queries based on subsections of images. For certain queries this capability significantly increases the relevance of the images retrieved, and further expands the user's control over the retrieval process.
Nicholas R. Howe, Daniel P. Huttenlocher
CVPR2
2000 Scene Modeling for Wide Area Surveillance and Image Synthesis
abstract
We present a method for modeling a scene that is observed by a moving camera, where only a portion of the scene is visible at any time. This method uses mixture models to represent pixels in a panoramic view, and to construct a "background image" that contains only static (non-moving) parts of the scene. The method can be used to reliably detect moving objects in a video sequence, detect patterns of activity over a wide field of view, and remove moving objects from a video or panoramic mosaic. The method also yields improved results in detecting moving objects and in constructing mosaics in the presence of moving objects, when compared with techniques that are not based on scene modeling. We present examples illustrating the results.
Anurag Mittal, Daniel P. Huttenlocher
CVPR2
1999 A New Bayesian Framework for Object Recognition
abstract
We introduce an approach to feature-based object recognition, using maximum a posteriori (MAP) estimation under a Markov random field (MRF) model. This approach provides an efficienct solution for a wide class of priors that explicitly model dependencies between individual features of an object. These priors capture phenomena such as the fact that unmatched features due to partial occlusion are generally spatially correlated rather than independent. The main focus of this paper is a special case of the framework that yields a particularly efficient approximation method. We call this special case spatially coherent matching (SCM), as it reflects the spatial correlation among neighboring features of an object. The SCM method operates directly on the image feature map, rather than relying on the graph-based methods used in the general framework. We present some Monte Carlo experiments showing that SCM yields substantial improvements over Hausdorff matching for cluttered scenes and partially occluded objects.
Yuri Boykov, Daniel P. Huttenlocher
CVPR2
1999 Digipaper: A Versatile Color Document Image Representation
abstract
We describe a segmentation method and associated file format for storing images of color documents. We separate each page of the document into three layers, containing the background (usually one or more photographic images), the text, and the color of the text. Each of these layers has different properties, making it desirable to use different compression methods to represent the three layers. The background layers are compressed using any method designed for photographic images, the text layers are compressed using a token-based representation, and the text color layers are compressed by augmenting the representation used for the text layers. We also describe an algorithm for segmenting images into these three layers. This representation and algorithm can produce very highly-compressed document files that nonetheless retain excellent image quality.
Daniel P. Huttenlocher, Pedro F. Felzenszwalb, William Rucklidge
ICIP (1)1
1999 Fast detection of common geometric substructure in proteins
abstract
We consider the problem of identifying common threedimensional substructures between proteins.Our method is based on comparing the shape of the a-carbon backbone structures of the proteins in order to find 3D rigid motions that bring portions of the geometric structures into correspondence.We propose a geometric representation of protein backbone chains that is compact yet allows for similarity measures that are robust against noise and outliers.We represent the structure of the backbone as a sequence of unit vectors, defined by each adjacent pair of a-carbons; we then define a measure of the similarity of two protein structures baaed on the RMS (root mean squared) distance between corresponding orientation vectors of the two proteins.Our measure has several advantages over measures that are commonly used for comparing protein shapes, such as the minimum RMS distance between the 3D positions of corresponding atoms in two proteins.This
L. Paul Chew, Daniel P. Huttenlocher, Klara Kedem, Jon M. Kleinberg
RECOMB2
1999 View-Based Recognition Using an Eigenspace Approximation to the Hausdorff Measure
abstract
View-based recognition methods, such as those using eigenspace techniques, have been successful for a number of recognition tasks. Such approaches, however, are somewhat limited in their ability to recognize objects that are partly hidden from view or occur against cluttered backgrounds. In order to address these limitations, we have developed a view matching technique based on an eigenspace approximation to the generalized Hausdorff measure. This method achieves compact storage and fast indexing that are the main advantages of eigenspace view matching techniques, while also being tolerant of partial occlusion and background clutter. The method applies to binary feature maps, such as intensity edges, rather than directly to intensity images.
Daniel P. Huttenlocher, Ryan H. Lilien, Clark F. Olson
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Image Segmentation Using Local Variation
abstract
We present a new graph-theoretic approach to the problem of image segmentation. Our method uses local criteria and yet produces results that reflect global properties of the image. We develop a framework that provides specific definitions of what it means for an image to be under- or over-segmented. We then present an efficient algorithm for computing a segmentation that is neither under- nor over-segmented according to these definitions. Our segmentation criterion is based on intensity differences between neighboring pixels. An important characteristic of the approach is that it is able to preserve detail in low-variability regions while ignoring detail in high-variability regions, which we illustrate with several examples on both real and synthetic images.
Pedro F. Felzenszwalb, Daniel P. Huttenlocher
CVPR2
1997 Geometric Pattern Matching Under Euclidean Motion
abstract
Given two planar sets A and B, we examine the problem of determining the smallest ϵ such that there is a Euclidean motion (rotation and translation) of A that brings each member of A within distance ϵ of some member of B. We establish upper bounds on the combinatorial complexity of this subproblem in model-based computer vision, when the sets A and B contain points, line segments, or (filled-in) polygons. We also show how to use our methods to substantially improve on existing algorithms for finding the minimum Hausdorff distance under Euclidean motion.
L. Paul Chew, Michael T. Goodrich, Daniel P. Huttenlocher, Klara Kedem, Jon M. Kleinberg, Dina Kravets
Comput. Geom.3
1997 Automatic target recognition by matching oriented edge pixels
abstract
This paper describes techniques to perform efficient and accurate target recognition in difficult domains. In order to accurately model small, irregularly shaped targets, the target objects and images are represented by their edge maps, with a local orientation associated with each edge pixel. Three dimensional objects are modeled by a set of two-dimensional (2-D) views of the object. Translation, rotation, and scaling of the views are allowed to approximate full three-dimensional (3-D) motion of the object. A version of the Hausdorff measure that incorporates both location and orientation information is used to determine which positions of each object model are reported as possible target locations. These positions are determined efficiently through the examination of a hierarchical cell decomposition of the transformation space. This allows large volumes of the space to be pruned quickly. Additional techniques are used to decrease the computation time required by the method when matching is performed against a catalog of object models. The probability that this measure will yield a false alarm and efficient methods for estimating this probability at run time are considered in detail. This information can be used to maintain a low false alarm rate or to rank competing hypotheses based on their likelihood of being a false alarm. Finally, results of the system recognizing objects in infrared and intensity images are given.
Clark F. Olson, Daniel P. Huttenlocher
IEEE Trans. Image Process.2
1996 Recognizing Three-Dimensional Objects by Comparing Two-Dimensional Images
abstract
In this paper we address the problem of recognizing an object from a novel viewpoint, given a single "model" view of that object. As is common in model-based recognition, objects and images are represented as sets of feature points. We present an efficient algorithm for determining whether two sets of image points (in the plane) could be projections of a common object (a three-dimensional point set). The method relies on the fact that two sets of points in the plane are orthographic projections of the same three-dimensional point set exactly when they have a common projection onto a line. This is a form of the well-known epipolar constraint used in stereopsis. Our algorithm can be used to recognize an object by comparing a stored two-dimensional view of the object against an unknown view, without requiring the correspondence between points in the views to be known a priori. We provide some examples illustrating the approach.
Daniel P. Huttenlocher, Liana M. Lorigo
CVPR1
1996 Object Recognition Using Subspace Methods
Daniel P. Huttenlocher, Ryan H. Lilien, Clark F. Olson
ECCV (1)1
1995 Computing Visual Correspondence: Incorporating the Probability of a False Match
abstract
We describe a method for computing visual correspondence which employs a formal model of the probability of a false match. This model estimates the chance that the best match for each point could have occurred at random. The model is effective at identifying points in one image for which there is no corresponding point in the other image, as occurs at depth boundaries in stereo and at motion boundaries in optical flow. More generally, the model can be used to identify points where the best match is of poor quality, as occurs in regions of uniform texture. We describe the similarity measure used in the method and present the formal model of a false match. We also show examples of using the method to compute stereo disparity.>
Daniel P. Huttenlocher, Eric W. Jaquith
ICCV1
1994 Visually-guided navigation by comparing two-dimensional edge images
abstract
We present a method for navigating a robot from an initial position to a specified landmark in its visual field, using a sequence of monocular images. The location of the landmark with respect to the robot is determined using the change in size and position of the landmark in the image as the robot moves. The landmark location is estimated after the first three images are taken, and this estimate is refined after each motion. One of the novel aspects of this method is that it uses no explicit three-dimensional information.>
Daniel P. Huttenlocher, Michael E. Leventon, William Rucklidge
CVPR1
1994 Comparing Point Sets Under Projection
Daniel P. Huttenlocher, Jon M. Kleinberg
SODA1
1994 A study of affine matching with bounded sensor error
W. Eric L. Grimson, Daniel P. Huttenlocher, David Jacobs 0001
Int. J. Comput. Vis.2
1993 A multi-resolution technique for comparing images using the Hausdorff distance
abstract
The Hausdorff distance measures the extent to which each point of a model set lies near some point of an image set and vice versa. An efficient method of computing this distance is developed, based on a multi-resolution tessellation of the space is possible transformations of the model set. One of the key ideas is that entire cells in this tessellation can be ruled out quickly, without actually computing the Hausdorff distance for many of them. Emphasis is placed on the case in which the model is allowed to translate and scale (independently in x and y) with respect to the image. This four-dimensional transformation space is searched rapidly while guaranteeing that no match will be missed. Some examples of identifying an object in a cluttered scene are presented, including cases where the object is partially hidden from view.>
Daniel P. Huttenlocher, William Rucklidge
CVPR1
1993 Tracking non-rigid objects in complex scenes
abstract
The authors describe a model-based method for tracking nonrigid objects moving in a complex scene. The method operates by extracting two-dimensional models of an object from a sequence of images. The basic idea underlying the technique is to decompose the image of a solid object moving in space into two components: a two-dimensional motion and a two-dimensional shape change. The motion component is factored out and the shape change is represented explicitly by a sequence of two-dimensional models, one corresponding to each image frame. The major assumption underlying the method is that the two-dimensional shape of an object will change slowly from one frame to the next. There is no assumption, however, that the two-dimensional image motion between successive frames will be small.>
Daniel P. Huttenlocher, Jae J. Noh, William Rucklidge
ICCV1
1993 The Upper Envelope of voronoi Surfaces and Its Applications
Daniel P. Huttenlocher, Klara Kedem, Micha Sharir
Discret. Comput. Geom.1
1993 Comparing Images Using the Hausdorff Distance
abstract
The Hausdorff distance measures the extent to which each point of a model set lies near some point of an image set and vice versa. Thus, this distance can be used to determine the degree of resemblance between two objects that are superimposed on one another. Efficient algorithms for computing the Hausdorff distance between all possible relative positions of a binary image and a model are presented. The focus is primarily on the case in which the model is only allowed to translate with respect to the image. The techniques are extended to rigid motion. The Hausdorff distance computation differs from many other shape comparison methods in that no correspondence between the model and the image is derived. The method is quite tolerant of small position errors such as those that occur with edge detectors and other feature extraction methods. It is shown that the method extends naturally to the problem of comparing a portion of a model against an image.>
Daniel P. Huttenlocher, Gregory A. Klanderman, William Rucklidge
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 On Dynamic Voronoi Diagrams and the Minimum Hausdorff Distance for Point Sets Under Euclidean Motion in the Plane
abstract
We show that the dynamic Voronoi diagram of k sets of points in the plane, where each set consists of m points moving rigidly, has complexity O(n2k2λs(k)) for some fixed s, where λs(n) is the maximum length of a (n, s) Davenport-Schinzel sequence. This improves the result of Aonuma et al., who show an upper bound of O(n3k4 log* k) for the complexity of such Voronoi diagrams. We then apply this result to the problem of finding the minimum Hausdorff distance between two point sets in the plane under Euclidean motion. We show that this distance can be computed in time O((m + n)6 log (mn)), where the two sets contain m and n points respectively.
Daniel P. Huttenlocher, Klara Kedem, Jon M. Kleinberg
SCG1
1992 Recognizing 3D objects from 2D images: an error analysis
abstract
Object recognition systems that use a small number of pairings of data and model features to compute the 3D transformation from model to sensor coordinates are considered. The effects of 2D sensor uncertainty on such computations are examined. The uncertainty in transformation parameters is bounded, and the effect of this uncertainty on false positive recognition rates is analyzed.>
W. Eric L. Grimson, Daniel P. Huttenlocher, Tao Daniel Alter
CVPR2
1992 Comparing images using the Hausdorff distance under translation
abstract
Efficient algorithms are provided for computing the Hausdorff distance between a binary image and all possible relative positions (translations) of a model, or a portion of that model. The computation is in many ways similar to binary correlation. However, it is more tolerant of perturbations in the locations of points because it measures proximity rather than exact superposition.>
Daniel P. Huttenlocher, William Rucklidge, Gregory A. Klanderman
CVPR1
1992 A Study of Affine Matching With Bounded Sensor Error
W. Eric L. Grimson, Daniel P. Huttenlocher, David Jacobs 0001
ECCV2
1992 Measuring the Quality of Hypotheses in Model-Based Recognition
Daniel P. Huttenlocher, Todd A. Cass
ECCV1
1992 Finding convex edge groupings in an image
Daniel P. Huttenlocher, Peter C. Wayner
Int. J. Comput. Vis.1
1992 Voronoi Diagrams of Rigidly Moving Sets of Points
abstract
Consider k sets each consisting of n points in the plane, with each set allowed to move rigidly according to some continuous function of time. A paper by Aonuma, Imai, Imai, and Tokuyama shows an upper bound of O(n3k4log∗n) on the number of combinatorial changes to the Voronoi diagram of the kn points over all time. We present a bound of O(n2k2λs(k)) for s fixed s, thus improving their result by slightly more than a factor of kn.
Daniel P. Huttenlocher, Klara Kedem, Jon M. Kleinberg
Inf. Process. Lett.1
1992 Introduction to the Special Issue on Interpretation of 3-D Scenes
W. Eric L. Grimson, Daniel P. Huttenlocher
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 The Upper Envelope of Voronoi Surfaces and Its Applications
abstract
Given a set S of sources (points or segments), we consider the surface that is the graph of the functien d(z) = minPcS p(z, p), for some metric p.This surface is closely related to the Voronoi diagram, Vor(S), of S under the metric p.The upper envelope of a set of these Voronoi surfaces, each defined for a differ-
Daniel P. Huttenlocher, Klara Kedem, Micha Sharir
SCG1
1991 Fast affine point matching: an output-sensitive method
abstract
A model-based recognition method that runs in time proportional to the actual number of instances of a model that are found in an image is presented. The key idea is to filter out many of the possible matches without having to explicitly consider each one. This contrasts with the hypothesize-and-test paradigm, commonly used in model-based recognition, where each possible match is tested and either accepted or rejected. For most recognition problems the number of possible matches is very large, whereas the number of actual matches is quite small, making output-sensitive methods such as this one very attractive. The method is based on an affine invariant representation of an object that uses distance ratios defined by quadruples of feature points. A central property of this representation is that it can be recovered from an image using only pairs of feature points.>
Daniel P. Huttenlocher
CVPR1
1991 Finding convex edge groupings in an image
abstract
A method for identifying groups of intensity edges in an image that are likely to result from the same convex object in a scene is described. A key property of the method is that its output is no more complex than the original image. The method uses a triangulation of linear edge segments to define a local neighborhood that is scale invariant. From this local neighborhood a local convexity graph that encodes which neighboring image edges could be part of a convex group of image edges is constructed. A path in the graph corresponds to a convex polygonal chain in the image, such as a convex polygon or a spiral. Examples are presented to illustrate that the technique find intuitively salient groups.>
Daniel P. Huttenlocher, Peter C. Wayner
CVPR1
1991 An Efficiently Computable Metric for Comparing Polygonal Shapes
abstract
A method for comparing polygons that is a metric, invariant under translation, rotation, and change of scale, reasonably easy to compute, and intuitive is presented. The method is based on the L/sub 2/ distance between the turning functions of the two polygons. It works for both convex and nonconvex polygons and runs in time O(mn log mn), where m is the number of vertices in one polygon and n is the number of vertices in the other. Some examples showing that the method produces answers that are intuitively reasonable are presented.>
Esther M. Arkin, L. Paul Chew, Daniel P. Huttenlocher, Klara Kedem, Joseph S. B. Mitchell
IEEE Trans. Pattern Anal. Mach. Intell.3
1991 Introduction to the Special Issue on Interpretation of 3-D Scenes-Part I
W. Eric L. Grimson, Daniel P. Huttenlocher
IEEE Trans. Pattern Anal. Mach. Intell.2
1991 On the Verification of Hypothesized Matches in Model-Based Recognition
abstract
Model-based recognition methods generally use ad hoc techniques to decide whether or not a model of an object matches a given scene. The most common such technique is to set an empirically determined threshold on the fraction of model features that must be matched to data features. Conditions under which to accept a match as correct are rigorously derived. The analysis is based on modeling the recognition process as a statistical occupancy problem. This model makes the assumption that pairings of object and data features can be characterized as a random process with a uniform distribution. The authors present a number of examples illustrating that real image data are well approximated by such a random process. Using a statistical occupancy model, they derive an expression for the probability that a randomly occurring match will account for a given fraction of the features of a particular object. This expression is a function of the number of model features, the number of data features, and bounds on the degree of sensor noise. It provides a means of setting a threshold such that the probability of a random match is very small.>
W. Eric L. Grimson, Daniel P. Huttenlocher
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Computing the Minimum Hausdorff Distance for Point Sets Under Translation
abstract
We consider the problem of computing a translation that minimizes the Hausdorff distance between two sets of points. For points in @@@@1 in the worst case there are ⊖(mn) translations at which the Hausdorff distance is a local minimum, where m is the number of points in one set and n is the number in the other. For points in @@@@2 there are ⊖(mn(m + n)) such local minima. We show how to compute the minimal Hausdorff distance in time Ο(mn log mn) for points in @@@@1 and in time Ο(m2n2α(mn)) for points in @@@@2. The results for the one-dimensional case are applied to the problem of comparing polygons under general affine transformations, where we extend the recent results of Arkin et al on polygon resemblance under rigid body motion. The two-dimensional case is closely related to the problem of finding an approximate congruence between two points sets under translation in the plane, as considered by Alt et al.
Daniel P. Huttenlocher, Klara Kedem
SCG1
1990 On the Verification of Hypthesized Matches in Model-Based Recognition
W. Eric L. Grimson, Daniel P. Huttenlocher
ECCV2
1990 On the sensitivity of geometric hashing
abstract
A formal model is presented for analyzing how the method performs in the presence of sensory uncertainty. The method performs well for simple images or for exact data. However, the performance degrades rapidly for cluttered scenes or in the presence of even moderate sensor error (3-5 pixels).>
W. Eric L. Grimson, Daniel P. Huttenlocher
ICCV2
1990 An Efficiently Computable Metric for Comparing Polygonal Shapes
Esther M. Arkin, L. Paul Chew, Daniel P. Huttenlocher, Klara Kedem, Joseph S. B. Mitchell
SODA3
1990 Recognizing solid objects by alignment with an image
Daniel P. Huttenlocher, Shimon Ullman
Int. J. Comput. Vis.1
1990 On the Sensitivity of the Hough Transform for Object Recognition
abstract
Object recognition from sensory data involves, in part, determining the pose of a model with respect to a scene. A common method for finding an object's pose is the generalized Hough transform, which accumulates evidence for possible coordinate transformations in a parameter space whose axes are the quantized transformation parameters. Large clusters of similar transformations in that space are taken as evidence of a correct match. A theoretical analysis of the behavior of such methods is presented. The authors derive bounds on the set of transformations consistent with each pairing of data and model features, in the presence of noise and occlusion in the image. Bounds are provided on the likelihood of false peaks in the parameter space, as a function of noise, occlusion, and tessellation effects. It is argued that haphazardly applying such methods to complex recognition tasks is risky, as the probability of false positives can be very high.>
W. Eric L. Grimson, Daniel P. Huttenlocher
IEEE Trans. Pattern Anal. Mach. Intell.2
1988 On The Sensitivity Of The Hough Transform For Object Recognition
abstract
A common method for finding an object's pose is the generalized Hough transform, which accumulates evidence for possible coordinate transformations in a parameter space and takes large clusters of similar transformations as evidence of a correct solution. We analyze this approach by deriving theoretical bounds on the set of transformations consistent with each data-model feature pairing, and by deriving bounds on the likelihood of false peaks in the parameter space, as a function of noise, occlusion, and tessellation effects. We argue that blithely applying such methods to complex recognition tasks is a risky proposition, as the probability of false positives can be very high.
W. Eric L. Grimson, Daniel P. Huttenlocher
ICCV2
1986 A broad phonetic classifier
abstract
It has been shown that broad phonetic sequences partition a large lexicon into small equivalence classes of words sharing the same sequence. While these results illustrate the power of broad phonetic constraints for differentiating words from one another, they do not suggest how to exploit sequential constraints in recognition. This paper presents a method for decoupling sequential phonetic constraints from a lexicon, by representing allowable broad phonetic sequences in terms of n-th order Markov models. A simple frame-based broad phonetic classifier is used to evaluate the effectiveness of these models in recognition. Tests on 300 sentences from 30 male speakers demonstrate that the addition of sequential constraints improves the classifier's performance.
Daniel P. Huttenlocher
ICASSP1
1984 A model of lexical access from partial phonetic information
abstract
Current approaches to isolated word recognition rely on classical pattern recognition techniques which utilize little or no speech specific knowledge. While the performance of these systems is quite good, they are not readily extensible to tasks involving very large vocabularies and many different speakers. This paper presents a model of lexical access using partial phonetic information. Rather than performing detailed phonetic analysis, a word is characterized in terms of broad phonetic and prosodic information. This partial description is then used to retrieve a small set of words from a large lexicon. The broad class representation used in the model is both relatively insensitive to variability in the speech signal, and very powerful in differentiating among the words in a large lexicon. In order to evaluate the use of this model, we have implemented a word hypothesizer which uses partial phonetic information in lexical access. The system performs a broad phonetic categorization of the acoustic signal. This broad classification is used to return a small set of word candidates from a 20,000 word lexicon. The system is not trained to a specific speaker or vocabulary.
Daniel P. Huttenlocher, Victor Zue
ICASSP1
1983 Phonotactic and Lexical Constraints in Speech Recognition
Daniel P. Huttenlocher, Victor W. Sue
AAAI1