VLDB 2026 Research / reviewers in the wild / expert
Erik B. Sudderth
dblp:22/3923
· DBLP profile ↗
60ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-0595-9726ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
45 papers |
Probabilistic and Bayesian machine learning · 44% 3D vision · 18% Graph learning · 7% | |
| Computer graphics and multimedia
5 papers |
Image and video processing · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Accessibility and assistive technology · 100% |
Topics — the 30 heaviest of 125, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.4 | 7 | 2021 | Marginalized Stochastic Natural Gradients for Black-Box Variational Inference · ICML 2021 Scalable Adaptation of State Complexity for Nonparametric Hidden Markov Models · NIPS 2015 Efficient Online Inference for Bayesian Nonparametric Relational Models · NIPS 2013 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model |
1.4 | 12 | 2015 | Scalable Adaptation of State Complexity for Nonparametric Hidden Markov Models · NIPS 2015 Efficient Online Inference for Bayesian Nonparametric Relational Models · NIPS 2013 Memoized Online Variational Inference for Dirichlet Process Mixture Models · NIPS 2013 |
Computer vision › 3D vision
3d object detection |
1.0 | 3 | 2020 | Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts · IEEE Trans. Pattern Anal. Mach. Intell. 2020 3D Object Detection With Latent Support Surfaces · CVPR 2018 Three-Dimensional Object Detection and Layout Prediction Using Clouds of Oriented Gradients · CVPR 2016 |
Accessibility and assistive technology › assistive technology for visual impairment
assistive technology for blind and low-vision users |
1.0 | 1 | 2026 | "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models · CHI 2026 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.8 | 4 | 2023 | Unbiased learning of deep generative models with structured discrete representations · NeurIPS 2023 Loop Series and Bethe Variational Bounds in Attractive Graphical Models · NIPS 2007 Distributed Occlusion Reasoning for Tracking with Nonparametric Belief Propagation · NIPS 2004 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle smoother |
0.8 | 1 | 2024 | Learning to be Smooth: An End-to-End Differentiable Particle Smoother · NeurIPS 2024 |
Robotics › Robot navigation and mapping › localization
vehicle localization |
0.8 | 1 | 2024 | Learning to be Smooth: An End-to-End Differentiable Particle Smoother · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle filtering |
0.7 | 2 | 2023 | Differentiable and Stable Long-Range Tracking of Multiple Posterior Modes · NeurIPS 2023 Nonparametric Belief Propagation · CVPR (1) 2003 |
Machine learning › Representation and self-supervised learning
discrete representation |
0.7 | 1 | 2023 | Unbiased learning of deep generative models with structured discrete representations · NeurIPS 2023 |
Machine learning › Generative modeling › variational autoencoder
structured variational autoencoder |
0.7 | 1 | 2023 | Unbiased learning of deep generative models with structured discrete representations · NeurIPS 2023 |
Machine learning › Generative modeling
variational autoencoder |
0.7 | 1 | 2023 | Unbiased learning of deep generative models with structured discrete representations · NeurIPS 2023 |
Computer vision › 3D vision
3d scene understanding |
0.6 | 3 | 2018 | 3D Object Detection With Latent Support Surfaces · CVPR 2018 Three-Dimensional Object Detection and Layout Prediction Using Clouds of Oriented Gradients · CVPR 2016 Depth from Familiar Objects: A Hierarchical Model for 3D Scenes · CVPR (2) 2006 |
Machine learning › Graph learning
graph clustering |
0.6 | 2 | 2022 | Thinned random measures for sparse graphs with overlapping communities · NeurIPS 2022 Efficient Online Inference for Bayesian Nonparametric Relational Models · NIPS 2013 |
Machine learning › Graph learning
stochastic block model |
0.6 | 2 | 2022 | Thinned random measures for sparse graphs with overlapping communities · NeurIPS 2022 Efficient Online Inference for Bayesian Nonparametric Relational Models · NIPS 2013 |
Computer vision › 3D vision › 3d object detection › multimodal 3d object detection
RGB-D 3D object detection |
0.6 | 2 | 2018 | 3D Object Detection With Latent Support Surfaces · CVPR 2018 Three-Dimensional Object Detection and Layout Prediction Using Clouds of Oriented Gradients · CVPR 2016 |
Machine learning › Graph learning › graph clustering
overlapping community detection |
0.6 | 1 | 2022 | Thinned random measures for sparse graphs with overlapping communities · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › gradient-based variational inference
black-box variational inference |
0.5 | 1 | 2021 | Marginalized Stochastic Natural Gradients for Black-Box Variational Inference · ICML 2021 |
Machine learning › Trustworthy machine learning › fairness
fair classification |
0.5 | 1 | 2021 | Scalable and Stable Surrogates for Flexible Classifiers with Fairness Constraints · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
fairness |
0.5 | 1 | 2021 | Scalable and Stable Surrogates for Flexible Classifiers with Fairness Constraints · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › fairness › algorithmic fairness
fairness constraints |
0.5 | 1 | 2021 | Scalable and Stable Surrogates for Flexible Classifiers with Fairness Constraints · NeurIPS 2021 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent |
0.5 | 1 | 2021 | Marginalized Stochastic Natural Gradients for Black-Box Variational Inference · ICML 2021 |
Machine learning › Optimization for machine learning › gradient estimation
stochastic gradient estimation |
0.5 | 1 | 2021 | Marginalized Stochastic Natural Gradients for Black-Box Variational Inference · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
hierarchical dirichlet process |
0.5 | 4 | 2013 | Efficient Online Inference for Bayesian Nonparametric Relational Models · NIPS 2013 Truly Nonparametric Online Variational Inference for Hierarchical Dirichlet Processes · NIPS 2012 Nonparametric Bayesian Learning of Switching Linear Dynamical Systems · NIPS 2008 |
Image and video processing › motion estimation
optical flow |
0.5 | 3 | 2015 | Layered RGBD scene flow estimation · CVPR 2015 Layered segmentation and optical flow estimation over time · CVPR 2012 Layered image motion with explicit occlusions, temporal consistency, and depth ordering · NIPS 2010 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference |
0.4 | 3 | 2015 | Proteins, Particles, and Pseudo-Max-Marginals: A Submodular Approach · ICML 2015 Preserving Modes and Messages via Diverse Particle Selection · ICML 2014 Nonparametric Belief Propagation · CVPR (1) 2003 |
Computer vision › 3D vision
depth estimation |
0.4 | 2 | 2019 | 3D Scene Reconstruction With Multi-Layer Depth and Epipolar Transformers · ICCV 2019 Depth from Familiar Objects: A Hierarchical Model for 3D Scenes · CVPR (2) 2006 |
Computer vision › 3D vision › 3d scene understanding › scene structure
3d scene layout prediction |
0.4 | 1 | 2020 | Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Computer vision › Segmentation and scene understanding › scene understanding
indoor scene understanding |
0.4 | 1 | 2020 | Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Computer vision › 3D vision
3d scene reconstruction |
0.4 | 1 | 2019 | 3D Scene Reconstruction With Multi-Layer Depth and Epipolar Transformers · ICCV 2019 |
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction |
0.4 | 1 | 2019 | 3D Scene Reconstruction With Multi-Layer Depth and Epipolar Transformers · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
survey · 2.0model evaluation · 2.0dataset annotation · 2.0two-filter smoother · 0.8stratification · 0.8importance sampling · 0.8natural gradient · 0.7importance-sampling gradient estimator · 0.7implicit differentiation · 0.7deep neural network encoder · 0.7generative model · 0.4dirichlet process mixture · 0.4graph-cut optimization · 0.3discrete-continuous optimization · 0.3MRF · 0.3variational inference · 0.3scalable learning · 0.3online inference · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsabstractVision-Language Models (VLMs) are increasingly used by blind and low-vision (BLV) people to identify and understand products in their everyday lives, such as food, personal care items, and household goods. Despite their prevalence, we lack an empirical understanding of how common image quality issues—such as blur, misframing, and rotation—affect the accuracy of VLM-generated captions and whether the resulting captions meet BLV people’s information needs. Based on a survey of 86 BLV participants, we develop an annotated dataset of 1,859 product images from BLV people to systematically evaluate how image quality issues affect VLM-generated captions. While the best VLM achieves 98% accuracy on images with no quality issues, accuracy drops to 75% overall when quality issues are present, worsening considerably as issues compound. We discuss the need for model evaluations that center on disabled people’s experiences throughout the process and offer concrete recommendations for HCI and ML researchers to make VLMs more reliable for BLV people. Kapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan, Darren Gergle, Erik B. Sudderth, Anne Marie Piper |
CHI | 6 |
| 2024 | Learning to be Smooth: An End-to-End Differentiable Particle SmootherabstractFor challenging state estimation problems arising in domains like vision and robotics, particle-based representations attractively enable temporal reasoning about multiple posterior modes. Particle smoothers offer the potential for more accurate offline data analysis by propagating information both forward and backward in time, but have classically required human-engineered dynamics and observation models. Extending recent advances in discriminative training of particle filters, we develop a framework for low-variance propagation of gradients across long time sequences when training particle smoothers. Our "two-filter" smoother integrates particle streams that are propagated forward and backward in time, while incorporating stratification and importance weights in the resampling step to provide low-variance gradient estimates for neural network dynamics and observation models. The resulting mixture density particle smoother is substantially more accurate than state-of-the-art particle filters, as well as search-based baselines, for city-scale global vehicle localization from real-world videos and maps. Ali Younis, Erik B. Sudderth |
NeurIPS | 2 |
| 2023 | Unbiased learning of deep generative models with structured discrete representationsabstractBy composing graphical models with deep learning architectures, we learn generative models with the strengths of both frameworks. The structured variational autoencoder (SVAE) inherits structure and interpretability from graphical models, and flexible likelihoods for high-dimensional data from deep learning, but poses substantial optimization challenges. We propose novel algorithms for learning SVAEs, and are the first to demonstrate the SVAE's ability to handle multimodal uncertainty when data is missing by incorporating discrete latent variables. Our memory-efficient implicit differentiation scheme makes the SVAE tractable to learn via gradient descent, while demonstrating robustness to incomplete optimization. To more rapidly learn accurate graphical model parameters, we derive a method for computing natural gradients without manual derivations, which avoids biases found in prior work. These optimization innovations enable the first comparisons of the SVAE to state-of-the-art time series models, where the SVAE performs competitively while learning interpretable and structured discrete data representations. Henry C. Bendekgey, Gabe Hope, Erik B. Sudderth |
NeurIPS | 3 |
| 2023 | Differentiable and Stable Long-Range Tracking of Multiple Posterior ModesabstractParticle filters flexibly represent multiple posterior modes nonparametrically, via a collection of weighted samples, but have classically been applied to tracking problems with known dynamics and observation likelihoods. Such generative models may be inaccurate or unavailable for high-dimensional observations like images. We instead leverage training data to discriminatively learn particle-based representations of uncertainty in latent object states, conditioned on arbitrary observations via deep neural network encoders. While prior discriminative particle filters have used heuristic relaxations of discrete particle resampling, or biased learning by truncating gradients at resampling steps, we achieve unbiased and low-variance gradient estimates by representing posteriors as continuous mixture densities. Our theory and experiments expose dramatic failures of existing reparameterization-based estimators for mixture gradients, an issue we address via an importance-sampling gradient estimator. Unlike standard recurrent neural networks, our mixture density particle filter represents multimodal uncertainty in continuous latent states, improving accuracy and robustness. On a range of challenging tracking and robot localization problems, our approach achieves dramatic improvements in accuracy, will also showing much greater stability across multiple training runs. Ali Younis, Erik B. Sudderth |
NeurIPS | 2 |
| 2023 | A decoder suffices for query-adaptive variational inferenceabstractDeep generative models like variational autoencoders (VAEs) are widely used for density estimation and dimensionality reduction, but infer latent representations via amortized inference algorithms, which require that all data dimensions are observed. VAEs thus lack a key strength of probabilistic graphical models: the ability to infer posteriors for test queries with arbitrary structure. We demonstrate that many prior methods for imputation with VAEs are costly and ineffective, and achieve superior performance via query-adaptive variational inference (QAVI) algorithms based directly on the generative decoder. By analytically marginalizing arbitrary sets of missing features, and optimizing expressive posteriors including mixtures and density flows, our non-amortized QAVI algorithms achieve excellent performance while avoiding expensive model retraining. On standard image and tabular datasets, our approach substantially outperforms prior methods in the plausibility and diversity of imputations. We also show that QAVI effectively generalizes to recent hierarchical VAE models for high-dimensional images. Sakshi Agarwal, Gabriel Hope, Ali Younis, Erik B. Sudderth |
UAI | 4 |
| 2022 | Thinned random measures for sparse graphs with overlapping communitiesabstractNetwork models for exchangeable arrays, including most stochastic block models, generate dense graphs with a limited ability to capture many characteristics of real-world social and biological networks. A class of models based on completely random measures like the generalized gamma process (GGP) have recently addressed some of these limitations. We propose a framework for thinning edges from realizations of GGP random graphs that models observed links via nodes' overall propensity to interact, as well as the similarity of node memberships within a large set of latent communities. Our formulation allows us to learn the number of communities from data, and enables efficient Monte Carlo methods that scale linearly with the number of observed edges, and thus (unlike dense block models) sub-quadratically with the number of entities or nodes. We compare to alternative models for both dense and sparse networks, and demonstrate effective recovery of latent community structure for real-world networks with thousands of nodes. Federica Zoe Ricci, Michele Guindani, Erik B. Sudderth |
NeurIPS | 3 |
| 2021 | Marginalized Stochastic Natural Gradients for Black-Box Variational InferenceabstractBlack-box variational inference algorithms use stochastic sampling to analyze diverse statistical models, like those expressed in probabilistic programming languages, without model-specific derivations. While the popular score-function estimator computes unbiased gradient estimates, its variance is often unacceptably large, especially in models with discrete latent variables. We propose a stochastic natural gradient estimator that is as broadly applicable and unbiased, but improves efficiency by exploiting the curvature of the variational bound, and provably reduces variance by marginalizing discrete latent variables. Our marginalized stochastic natural gradients have intriguing connections to classic coordinate ascent variational inference, but allow parallel updates of variational parameters, and provide superior convergence guarantees relative to naive Monte Carlo approximations. We integrate our method with the probabilistic programming language Pyro and evaluate real-world models of documents, images, networks, and crowd-sourcing. Compared to score-function estimators, we require far fewer Monte Carlo samples and consistently convergence orders of magnitude faster. Geng Ji 0001, Debora Sujono, Erik B. Sudderth |
ICML | 3 |
| 2021 | Scalable and Stable Surrogates for Flexible Classifiers with Fairness ConstraintsabstractWe investigate how fairness relaxations scale to flexible classifiers like deep neural networks for images and text. We analyze an easy-to-use and robust way of imposing fairness constraints when training, and through this framework prove that some prior fairness surrogates exhibit degeneracies for non-convex models. We resolve these problems via three new surrogates: an adaptive data re-weighting, and two smooth upper-bounds that are provably more robust than some previous methods. Our surrogates perform comparably to the state-of-the-art on low-dimensional fairness benchmarks, while achieving superior accuracy and stability for more complex computer vision and natural language processing tasks. Harry Bendekgey, Erik B. Sudderth |
NeurIPS | 2 |
| 2020 | Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene LayoutsabstractWe develop new representations and algorithms for three-dimensional (3D) object detection and spatial layout prediction in cluttered indoor scenes. We first propose a clouds of oriented gradient (COG) descriptor that links the 2D appearance and 3D pose of object categories, and thus accurately models how perspective projection affects perceived image gradients. To better represent the 3D visual styles of large objects and provide contextual cues to improve the detection of small objects, we introduce latent support surfaces. We then propose a "Manhattan voxel" representation which better captures the 3D room layout geometry of common indoor environments. Effective classification rules are learned via a latent structured prediction framework. Contextual relationships among categories and layout are captured via a cascade of classifiers, leading to holistic scene hypotheses that exceed the state-of-the-art on the SUN RGB-D database. Zhile Ren, Erik B. Sudderth |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | 3D Scene Reconstruction With Multi-Layer Depth and Epipolar TransformersabstractWe tackle the problem of automatically reconstructing a complete 3D model of a scene from a single RGB image. This challenging task requires inferring the shape of both visible and occluded surfaces. Our approach utilizes viewer-centered, multi-layer representation of scene geometry adapted from recent methods for single object shape completion. To improve the accuracy of view-centered representations for complex scenes, we introduce a novel "Epipolar Feature Transformer" that transfers convolutional network features from an input view to other virtual camera viewpoints, and thus better covers the 3D scene geometry. Unlike existing approaches that first detect and localize objects in 3D, and then infer object shape using category-specific models, our approach is fully convolutional, end-to-end differentiable, and avoids the resolution and memory limitations of voxel representations. We demonstrate the advantages of multi-layer depth representations and epipolar feature transformers on the reconstruction of a large database of indoor scenes. Daeyun Shin, Zhile Ren, Erik B. Sudderth, Charless C. Fowlkes |
ICCV | 3 |
| 2019 | Variational Training for Large-Scale Noisy-OR Bayesian Networks
Geng Ji 0001, Dehua Cheng, Huazhong Ning, Changhe Yuan, Hanning Zhou, Liang Xiong, Erik B. Sudderth |
UAI | 7 |
| 2019 | A Fusion Approach for Multi-Frame Optical Flow EstimationabstractTo date, top-performing optical flow estimation methods only take pairs of consecutive frames into account. While elegant and appealing, the idea of using more than two frames has not yet produced state-of-the-art results. We present a simple, yet effective fusion approach for multi-frame optical flow that benefits from longer-term temporal cues. Our method first warps the optical flow from previous frames to the current, thereby yielding multiple plausible estimates. It then fuses the complementary information carried by these estimates into a new optical flow field. At the time of writing, our method ranks first among published results in the MPI Sintel and KITTI 2015 benchmarks. Our models will be available on https://github.com/NVlabs/PWC-Net. Zhile Ren, Orazio Gallo, Deqing Sun, Ming-Hsuan Yang 0001, Erik B. Sudderth, Jan Kautz |
WACV | 5 |
| 2018 | Semi-Supervised Prediction-Constrained Topic ModelsabstractSupervisory signals can help topic models discover low-dimensional data representations which are useful for a specific prediction task. We propose a framework for training supervised latent Dirichlet allocation that balances two goals: faithful generative explanations of high-dimensional data and accurate prediction of associated class labels. Existing approaches fail to balance these goals by not properly handling a fundamental asymmetry: the intended application is always predicting labels from data, not data from labels. Our new prediction-constrained objective for training generative models coherently integrates supervisory signals even when only a small fraction of training examples are labeled. We demonstrate improved prediction quality compared to previous supervised topic models, achieving results competitive with high-dimensional logistic regression on text analysis and electronic health records tasks while simultaneously learning interpretable topics. Michael C. Hughes, Gabriel Hope, Leah Weiner, Thomas H. McCoy Jr., Roy H. Perlis, Erik B. Sudderth, Finale Doshi-Velez |
AISTATS | 6 |
| 2018 | 3D Object Detection With Latent Support SurfacesabstractWe develop a 3D object detection algorithm that uses latent support surfaces to capture contextual relationships in indoor scenes. Existing 3D representations for RGB-D images capture the local shape and appearance of object categories, but have limited power to represent objects with different visual styles. The detection of small objects is also challenging because the search space is very large in 3D scenes. However, we observe that much of the shape variation within 3D object categories can be explained by the location of a latent support surface, and smaller objects are often supported by larger objects. Therefore, we explicitly use latent support surfaces to better represent the 3D appearance of large objects, and provide contextual cues to improve the detection of small objects. We evaluate our model with 19 object categories from the SUN RGB-D database, and demonstrate state-of-the-art performance. Zhile Ren, Erik B. Sudderth |
CVPR | 2 |
| 2017 | Cascaded Scene Flow Prediction Using Semantic SegmentationabstractGiven two consecutive frames from a pair of stereo cameras, 3D scene flow methods simultaneously estimate the 3D geometry and motion of the observed scene. Many existing approaches use superpixels for regularization, but may predict inconsistent shapes and motions inside rigidly moving objects. We instead assume that scenes consist of foreground objects rigidly moving in front of a static background, and use semantic cues to produce pixel-accurate scene flow estimates. Our cascaded classification framework accurately models 3D scenes by iteratively refining semantic segmentation masks, stereo correspondences, 3D rigid motion estimates, and optical flow fields. We evaluate our method on the challenging KITTI autonomous driving benchmark, and show that accounting for the motion of segmented vehicles leads to state-of-the-art performance. Zhile Ren, Deqing Sun, Jan Kautz, Erik B. Sudderth |
3DV | 4 |
| 2017 | From Patches to Images: A Nonparametric Generative ModelabstractWe propose a hierarchical generative model that captures the self-similar structure of image regions as well as how this structure is shared across image collections. Our model is based on a novel, variational interpretation of the popular expected patch log-likelihood (EPLL) method as a model for randomly positioned grids of image patches. While previous EPLL methods modeled image patches with finite Gaussian mixtures, we use nonparametric Dirichlet process (DP) mixtures to create models whose complexity grows as additional images are observed. An extension based on the hierarchical DP then captures repetitive and self-similar structure via image-specific variations in cluster frequencies. We derive a structured variational inference algorithm that adaptively creates new patch clusters to more accurately model novel image textures. Our denoising performance on standard benchmarks is superior to EPLL and comparable to the state-of-the-art, and provides novel statistical justifications for common image processing heuristics. We also show accurate image inpainting results. Geng Ji 0001, Michael C. Hughes, Erik B. Sudderth |
ICML | 3 |
| 2017 | Multiscale Semi-Markov Dynamics for Intracortical Brain-Computer InterfacesabstractIntracortical brain-computer interfaces (iBCIs) have allowed people with tetraplegia to control a computer cursor by imagining the movement of their paralyzed arm or hand. State-of-the-art decoders deployed in human iBCIs are derived from a Kalman filter that assumes Markov dynamics on the angle of intended movement, and a unimodal dependence on intended angle for each channel of neural activity. Due to errors made in the decoding of noisy neural data, as a user attempts to move the cursor to a goal, the angle between cursor and goal positions may change rapidly. We propose a dynamic Bayesian network that includes the on-screen goal position as part of its latent state, and thus allows the person’s intended angle of movement to be aggregated over a much longer history of neural activity. This multiscale model explicitly captures the relationship between instantaneous angles of motion and long-term goals, and incorporates semi-Markov dynamics for motion trajectories. We also introduce a multimodal likelihood model for recordings of neural populations which can be rapidly calibrated for clinical applications. In offline experiments with recorded neural data, we demonstrate significantly improved prediction of motion directions compared to the Kalman filter. We derive an efficient online inference algorithm, enabling a clinical trial participant with tetraplegia to control a computer cursor with neural activity in real time. The observed kinematics of cursor movement are objectively straighter and smoother than prior iBCI decoding models without loss of responsiveness. Daniel Milstein, Jason L. Pacheco, Leigh J. Hochberg, John D. Simeral, Beata Jarosiewicz, Erik B. Sudderth |
NIPS | 6 |
| 2017 | Refinery: An Open Source Topic Modeling Web PlatformabstractWe introduce Refinery, an open source platform for exploring large text document collections with topic models. Refinery is a standalone web application driven by a graphical interface, so it is usable by those without machine learning or programming expertise. Users can interactively organize articles by topic and also refine this organization with phrase-level analysis. Under the hood, we train Bayesian nonparametric topic models that can adapt model complexity to the provided data with scalable learning algorithms. The project website contains Python code and further documentation. Dae Il Kim, Benjamin F. Swanson, Michael C. Hughes, Erik B. Sudderth |
J. Mach. Learn. Res. | 4 |
| 2016 | Three-Dimensional Object Detection and Layout Prediction Using Clouds of Oriented GradientsabstractWe develop new representations and algorithms for three-dimensional (3D) object detection and spatial layout prediction in cluttered indoor scenes. RGB-D images are traditionally described by local geometric features of the 3D point cloud. We propose a cloud of oriented gradient (COG) descriptor that links the 2D appearance and 3D pose of object categories, and thus accurately models how perspective projection affects perceived image boundaries. We also propose a "Manhattan voxel" representation which better captures the 3D room layout geometry of common indoor environments. Effective classification rules are learned via a structured prediction framework that accounts for the intersection-over-union overlap of hypothesized 3D cuboids with human annotations, as well as orientation estimation errors. Contextual relationships among categories and layout are captured via a cascade of classifiers, leading to holistic scene hypotheses with improved accuracy. Our model is learned solely from annotated RGB-D images, without the benefit of CAD models, but nevertheless its performance substantially exceeds the state-of-the-art on the SUN RGB-D database. Avoiding CAD models allows easier learning of detectors for many object categories. Zhile Ren, Erik B. Sudderth |
CVPR | 2 |
| 2015 | Reliable and Scalable Variational Inference for the Hierarchical Dirichlet ProcessabstractWe introduce a new variational inference objective for hierarchical Dirichlet process admixture models. Our approach provides novel and scalable algorithms for learning nonparametric topic models of text documents and Gaussian admixture models of image patches. Improving on the point estimates of topic probabilities used in previous work, we define full variational posteriors for all latent variables and optimize parameters via a novel surrogate likelihood bound. We show that this approach has crucial advantages for data-driven learning of the number of topics. Via merge and delete moves that remove redundant or irrelevant topics, we learn compact and interpretable models with less computation. Scaling to millions of documents is possible using stochastic or memoized variational updates. Michael C. Hughes, Dae Il Kim, Erik B. Sudderth |
AISTATS | 3 |
| 2015 | Layered RGBD scene flow estimationabstractAs consumer depth sensors become widely available, estimating scene flow from RGBD sequences has received increasing attention. Although the depth information allows the recovery of 3D motion from a single view, it poses new challenges. In particular, depth boundaries are not well-aligned with RGB image edges and therefore not reliable cues to localize 2D motion boundaries. In addition, methods that extend the 2D optical flow formulation to 3D still produce large errors in occlusion regions. To better use depth for occlusion reasoning, we propose a layered RGBD scene flow method that jointly solves for the scene segmentation and the motion. Our key observation is that the noisy depth is sufficient to decide the depth ordering of layers, thereby avoiding a computational bottleneck for RGB layered methods. Furthermore, the depth enables us to estimate a per-layer 3D rigid motion to constrain the motion of each layer. Experimental results on both the Middlebury and real-world sequences demonstrate the effectiveness of the layered approach for RGBD scene flow estimation. Deqing Sun, Erik B. Sudderth, Hanspeter Pfister |
CVPR | 2 |
| 2015 | Proteins, Particles, and Pseudo-Max-Marginals: A Submodular ApproachabstractVariants of max-product (MP) belief propagation effectively find modes of many complex graphical models, but are limited to discrete distributions. Diverse particle max-product (D-PMP) robustly approximates max-product updates in continuous MRFs using stochastically sampled particles, but previous work was specialized to tree-structured models. Motivated by the challenging problem of protein side chain prediction, we extend D-PMP in several key ways to create a generic MAP inference algorithm for loopy models. We define a modified diverse particle selection objective that is provably submodular, leading to an efficient greedy algorithm with rigorous optimality guarantees, and corresponding max-marginal error bounds. We further incorporate tree-reweighted variants of the MP algorithm to allow provable verification of global MAP recovery in many models. Our general-purpose Matlab library is applicable to a wide range of pairwise graphical models, and we validate our approach using optical flow benchmarks. We further demonstrate superior side chain prediction accuracy compared to baseline algorithms from the state-of-the-art Rosetta package. Jason L. Pacheco, Erik B. Sudderth |
ICML | 2 |
| 2015 | Scalable Adaptation of State Complexity for Nonparametric Hidden Markov ModelsabstractBayesian nonparametric hidden Markov models are typically learned via fixed truncations of the infinite state space or local Monte Carlo proposals that make small changes to the state space. We develop an inference algorithm for the sticky hierarchical Dirichlet process hidden Markov model that scales to big datasets by processing a few sequences at a time yet allows rapid adaptation of the state space cardinality. Unlike previous point-estimate methods, our novel variational bound penalizes redundant or irrelevant states and thus enables optimization of the state space. Our birth proposals use observed data statistics to create useful new states that escape local optima. Merge and delete proposals remove ineffective states to yield simpler models with more affordable future computations. Experiments on speaker diarization, motion capture, and epigenetic chromatin datasets discover models that are more compact, more interpretable, and better aligned to ground truth segmentations than competitors. We have released an open-source Python implementation which can parallelize local inference steps across sequences. Michael C. Hughes, William T. Stephenson, Erik B. Sudderth |
NIPS | 3 |
| 2015 | Guest Editors' Introduction to the Special Issue on Bayesian NonparametricsabstractThe articles in this special issue discuss the applications supported by Bayesian nonparametric modeling. These probabilistic models defined over infinite-dimensional parameter spaces. For Gaussian process models of regression and classification functions, the parameter space consists of a set of continuous functions. For the Dirichlet process mixture models used in density estimation and clustering, the parameter space is dense in the space of probability measures. Bayesian nonparametric models provide a flexible framework for modeling complex data and a promising alternative to classical model selection methods. Due to recent computational advances, these approaches have received increasing attention in machine learning, statistics, probability, and related application domains. Ryan P. Adams, Emily B. Fox, Erik B. Sudderth, Yee Whye Teh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Preserving Modes and Messages via Diverse Particle SelectionabstractIn applications of graphical models arising in domains such as computer vision and signal processing, we often seek the most likely configurations of high-dimensional, continuous variables. We develop a particle-based max-product algorithm which maintains a diverse set of posterior mode hypotheses, and is robust to initialization. At each iteration, the set of hypotheses at each node is augmented via stochastic proposals, and then reduced via an efficient selection algorithm. The integer program underlying our optimization-based particle selection minimizes errors in subsequent max-product message updates. This objective automatically encourages diversity in the maintained hypotheses, without requiring tuning of application-specific distances among hypotheses. By avoiding the stochastic resampling steps underlying particle sum-product algorithms, we also avoid common degeneracies where particles collapse onto a single hypothesis. Our approach significantly outperforms previous particle-based algorithms in experiments focusing on the estimation of human pose from single images. Jason L. Pacheco, Silvia Zuffi, Michael J. Black, Erik B. Sudderth |
ICML | 4 |
| 2014 | Nonparametric Clustering with Distance Dependent Hierarchies
Soumya Ghosh, Michalis Raptis, Leonid Sigal, Erik B. Sudderth |
UAI | 4 |
| 2013 | A Fully-Connected Layered Model of Foreground and Background FlowabstractLayered models allow scene segmentation and motion estimation to be formulated together and to inform one another. Traditional layered motion methods, however, employ fairly weak models of scene structure, relying on locally connected Ising/Potts models which have limited ability to capture long-range correlations in natural scenes. To address this, we formulate a fully-connected layered model that enables global reasoning about the complicated segmentations of real objects. Optimization with fully-connected graphical models is challenging, and our inference algorithm leverages recent work on efficient mean field updates for fully-connected conditional random fields. These methods can be implemented efficiently using high-dimensional Gaussian filtering. We combine these ideas with a layered flow model, and find that the long-range connections greatly improve segmentation into figure-ground layers when compared with locally connected MRF models. Experiments on several benchmark datasets show that the method can recover fine structures and large occlusion regions, with good flow accuracy and much lower computational cost than previous locally-connected layered models. Deqing Sun, Jonas Wulff, Erik B. Sudderth, Hanspeter Pfister, Michael J. Black |
CVPR | 3 |
| 2013 | Memoized Online Variational Inference for Dirichlet Process Mixture ModelsabstractVariational inference algorithms provide the most effective framework for large-scale training of Bayesian nonparametric models. Stochastic online approaches are promising, but are sensitive to the chosen learning rate and often converge to poor local optima. We present a new algorithm, memoized online variational inference, which scales to very large (yet finite) datasets while avoiding the complexities of stochastic gradient. Our algorithm maintains finite-dimensional sufficient statistics from batches of the full dataset, requiring some additional memory but still scaling to millions of examples. Exploiting nested families of variational bounds for infinite nonparametric models, we develop principled birth and merge moves allowing non-local optimization. Births adaptively add components to the model to escape local optima, while merges remove redundancy and improve speed. Using Dirichlet process mixture models for image clustering and denoising, we demonstrate major improvements in robustness and accuracy. Michael C. Hughes, Erik B. Sudderth |
NIPS | 2 |
| 2013 | Efficient Online Inference for Bayesian Nonparametric Relational ModelsabstractStochastic block models characterize observed network relationships via latent community memberships. In large social networks, we expect entities to participate in multiple communities, and the number of communities to grow with the network size. We introduce a new model for these phenomena, the hierarchical Dirichlet process relational model, which allows nodes to have mixed membership in an unbounded set of communities. To allow scalable learning, we derive an online stochastic variational inference algorithm. Focusing on assortative models of undirected networks, we also propose an efficient structured mean field variational bound, and online methods for automatically pruning unused communities. Compared to state-of-the-art online learning methods for parametric relational models, we show significantly improved perplexity and link prediction accuracy for sparse networks with tens of thousands of nodes. We also showcase an analysis of LittleSis, a large network of who-knows-who at the heights of business and government. Dae Il Kim, Prem Gopalan, David M. Blei, Erik B. Sudderth |
NIPS | 4 |
| 2012 | Nonparametric learning for layered segmentation of natural imagesabstractWe explore recently proposed Bayesian nonparametric models of image partitions, based on spatially dependent Pitman-Yor processes. These models are attractive because they adapt to images of varying complexity, successfully modeling uncertainty in the structure and scale of human segmentations of natural scenes. By developing substantially improved inference and learning algorithms, we achieve performance comparable to state-of-the-art methods. For learning, we show how the Gaussian process (GP) covariance functions underlying these models can be calibrated to accurately match the statistics of example human segmentations. For inference, we develop a stochastic search-based algorithm which is substantially less susceptible to local optima than conventional variational methods. Our approach utilizes the expectation propagation algorithm to approximately marginalize latent GPs, and a low rank covariance representation to improve computational efficiency. Experiments with two benchmark datasets show that our learning and inference innovations substantially improve segmentation accuracy. By hypothesizing multiple partitions for each image, we also take steps towards capturing the variability of human scene interpretations. Soumya Ghosh, Erik B. Sudderth |
CVPR | 2 |
| 2012 | Layered segmentation and optical flow estimation over timeabstractLayered models provide a compelling approach for estimating image motion and segmenting moving scenes. Previous methods, however, have failed to capture the structure of complex scenes, provide precise object boundaries, effectively estimate the number of layers in a scene, or robustly determine the depth order of the layers. Furthermore, previous methods have focused on optical flow between pairs of frames rather than longer sequences. We show that image sequences with more frames are needed to resolve ambiguities in depth ordering at occlusion boundaries; temporal layer constancy makes this feasible. Our generative model of image sequences is rich but difficult to optimize with traditional gradient descent methods. We propose a novel discrete approximation of the continuous objective in terms of a sequence of depth-ordered MRFs and extend graph-cut optimization methods with new “moves” that make joint layer segmentation and motion estimation feasible. Our optimizer, which mixes discrete and continuous optimization, automatically determines the number of layers and reasons about their depth ordering. We demonstrate the value of layered models, our optimization strategy, and the use of more than two frames on both the Middlebury optical flow benchmark and the MIT layer segmentation benchmark. Deqing Sun, Erik B. Sudderth, Michael J. Black |
CVPR | 2 |
| 2012 | The Nonparametric Metadata Dependent Relational Model
Dae Il Kim, Michael C. Hughes, Erik B. Sudderth |
ICML | 3 |
| 2012 | Truly Nonparametric Online Variational Inference for Hierarchical Dirichlet ProcessesabstractVariational methods provide a computationally scalable alternative to Monte Carlo methods for large-scale, Bayesian nonparametric learning. In practice, however, conventional batch and online variational methods quickly become trapped in local optima. In this paper, we consider a nonparametric topic model based on the hierarchical Dirichlet process (HDP), and develop a novel online variational inference algorithm based on split-merge topic updates. We derive a simpler and faster variational approximation of the HDP, and show that by intelligently splitting and merging components of the variational posterior, we can achieve substantially better predictions of test data than conventional online and batch variational algorithms. For streaming analysis of large datasets where batch analysis is infeasible, we show that our split-merge updates better capture the nonparametric properties of the underlying model, allowing continual learning of new topics. Michael Bryant 0002, Erik B. Sudderth |
NIPS | 2 |
| 2012 | From Deformations to Parts: Motion-based Segmentation of 3D ObjectsabstractWe develop a method for discovering the parts of an articulated object from aligned meshes capturing various three-dimensional (3D) poses. We adapt the distance dependent Chinese restaurant process (ddCRP) to allow nonparametric discovery of a potentially unbounded number of parts, while simultaneously guaranteeing a spatially connected segmentation. To allow analysis of datasets in which object instances have varying shapes, we model part variability across poses via affine transformations. By placing a matrix normal-inverse-Wishart prior on these affine transformations, we develop a ddCRP Gibbs sampler which tractably marginalizes over transformation uncertainty. Analyzing a dataset of humans captured in dozens of poses, we infer parts which provide quantitatively better motion predictions than conventional clustering methods. Soumya Ghosh, Erik B. Sudderth, Matthew Loper, Michael J. Black |
NIPS | 2 |
| 2012 | Effective Split-Merge Monte Carlo Methods for Nonparametric Models of Sequential DataabstractApplications of Bayesian nonparametric methods require learning and inference algorithms which efficiently explore models of unbounded complexity. We develop new Markov chain Monte Carlo methods for the beta process hidden Markov model (BP-HMM), enabling discovery of shared activity patterns in large video and motion capture databases. By introducing split-merge moves based on sequential allocation, we allow large global changes in the shared feature structure. We also develop data-driven reversible jump moves which more reliably discover rare or unique behaviors. Our proposals apply to any choice of conjugate likelihood for observed data, and we show success with multinomial, Gaussian, and autoregressive emission models. Together, these innovations allow tractable analysis of hundreds of time series, where previous inference required clever initialization and at least ten thousand burn-in iterations for just six sequences. Michael C. Hughes, Emily B. Fox, Erik B. Sudderth |
NIPS | 3 |
| 2012 | Minimization of Continuous Bethe Approximations: A Positive VariationabstractWe develop convergent minimization algorithms for Bethe variational approximations which explicitly constrain marginal estimates to families of valid distributions. While existing message passing algorithms define fixed point iterations corresponding to stationary points of the Bethe free energy, their greedy dynamics do not distinguish between local minima and maxima, and can fail to converge. For continuous estimation problems, this instability is linked to the creation of invalid marginal estimates, such as Gaussians with negative variance. Conversely, our approach leverages multiplier methods with well-understood convergence properties, and uses bound projection methods to ensure that marginal approximations are valid at all iterations. We derive general algorithms for discrete and Gaussian pairwise Markov random fields, showing improvements over standard loopy belief propagation. We also apply our method to a hybrid model with both discrete and continuous variables, showing improvements over expectation propagation. Jason L. Pacheco, Erik B. Sudderth |
NIPS | 2 |
| 2011 | Global Seismic Monitoring: A Bayesian ApproachabstractThe automated processing of multiple seismic signals to detect and localize seismic events is a central tool in both geophysics and nuclear treaty verification. This paper reports on a project, begun in 2009, to reformulate this problem in a Bayesian framework. A Bayesian seismic monitoring system, NET-VISA, has been built comprising a spatial event prior and generative models of event transmission and detection, as well as an inference algorithm. Applied in the context of the International Monitoring System (IMS), a global sensor network developed for the Comprehensive Nuclear-Test-Ban Treaty (CTBT), NET-VISA achieves a reduction of around 50% in the number of missed events compared to the currently deployed system. It also finds events that are missed even by the human analysts who post-process the IMS output. Nimar S. Arora, Stuart Russell 0001, Paul Kidwell, Erik B. Sudderth |
AAAI | 4 |
| 2011 | Spatial distance dependent Chinese restaurant processes for image segmentationabstractThe distance dependent Chinese restaurant process (ddCRP) was recently introduced to accommodate random partitions of non-exchangeable data. The ddCRP clusters data in a biased way: each data point is more likely to be clustered with other data that are near it in an external sense. This paper examines the ddCRP in a spatial setting with the goal of natural image segmentation. We explore the biases of the spatial ddCRP model and propose a novel hierarchical extension better suited for producing "human-like" segmentations. We then study the sensitivity of the models to various distance and appearance hyperparameters, and provide the first rigorous comparison of nonparametric Bayesian models in the image segmentation domain. On unsupervised image segmentation, we demonstrate that similar performance to existing nonparametric Bayesian models is possible with substantially simpler models and algorithms. Soumya Ghosh, Andrei B. Ungureanu, Erik B. Sudderth, David M. Blei |
NIPS | 3 |
| 2011 | The Doubly Correlated Nonparametric Topic ModelabstractTopic models are learned via a statistical model of variation within document collections, but designed to extract meaningful semantic structure. Desirable traits include the ability to incorporate annotations or metadata associated with documents; the discovery of correlated patterns of topic usage; and the avoidance of parametric assumptions, such as manual specification of the number of topics. We propose a doubly correlated nonparametric topic (DCNT) model, the first model to simultaneously capture all three of these properties. The DCNT models metadata via a flexible, Gaussian regression on arbitrary input features; correlations via a scalable square-root covariance representation; and nonparametric selection from an unbounded series of potential topics via a stick-breaking construction. We validate the semantic structure and predictive performance of the DCNT using a corpus of NIPS documents annotated by various metadata. Dae Il Kim, Erik B. Sudderth |
NIPS | 2 |
| 2010 | Global seismic monitoring as probabilistic inferenceabstractThe International Monitoring System (IMS) is a global network of sensors whose purpose is to identify potential violations of the Comprehensive Nuclear-Test-Ban Treaty (CTBT), primarily through detection and localization of seismic events. We report on the first stage of a project to improve on the current automated software system with a Bayesian inference system that computes the most likely global event history given the record of local sensor data. The new system, VISA (Vertically Integrated Seismological Analysis), is based on empirically calibrated, generative models of event occurrence, signal propagation, and signal detection. VISA exhibits significantly improved precision and recall compared to the current operational system and is able to detect events that are missed even by the human analysts who post-process the IMS output. Nimar S. Arora, Stuart Russell 0001, Paul Kidwell, Erik B. Sudderth |
NIPS | 4 |
| 2010 | Layered image motion with explicit occlusions, temporal consistency, and depth orderingabstractLayered models are a powerful way of describing natural scenes containing smooth surfaces that may overlap and occlude each other. For image motion estimation, such models have a long history but have not achieved the wide use or accuracy of non-layered methods. We present a new probabilistic model of optical flow in layers that addresses many of the shortcomings of previous approaches. In particular, we define a probabilistic graphical model that explicitly captures: 1) occlusions and disocclusions; 2) depth ordering of the layers; 3) temporal consistency of the layer segmentation. Additionally the optical flow in each layer is modeled by a combination of a parametric model and a smooth deviation based on an MRF with a robust spatial prior; the resulting model allows roughness in layers. Finally, a key contribution is the formulation of the layers using an image-dependent hidden field prior based on recent models for static scene segmentation. The method achieves state-of-the-art results on the Middlebury benchmark and produces meaningful scene segmentations as well as detected occlusion regions. Deqing Sun, Erik B. Sudderth, Michael J. Black |
NIPS | 2 |
| 2010 | Gibbs Sampling in Open-Universe Stochastic Languages
Nimar S. Arora, Rodrigo de Salvo Braz, Erik B. Sudderth, Stuart Russell 0001 |
UAI | 3 |
| 2009 | Nonparametric belief propagation for distributed tracking of robot networks with noisy inter-distance measurementsabstractWe consider the problem of tracking multiple moving robots using noisy sensing of inter-robot and inter-beacon distances. Sensing is local: there are three fixed beacons at known locations, so distance and position estimates propagate across multiple robots. We show that the technique of Nonparametric Belief Propagation (NBP), a graph-based generalization of particle filtering, can address this problem and model multi-modal and ring-shaped uncertainty distributions. NBP provides the basis for distributed algorithms in which messages are exchanged between local neighbors. Generalizing previous approaches to localization in static sensor networks, we improve efficiency and accuracy by using a dynamics model for temporal tracking. We compare the NBP dynamic tracking algorithm with SMCL+R, a sequential Monte Carlo algorithm. Whereas NBP currently requires more computation, it converges in more cases and provides estimates that are 3 to 4 times more accurate. NBP also facilitates probabilistic models of sensor accuracy and network connectivity. Jeremy Schiff, Erik B. Sudderth, Kenneth Y. Goldberg |
IROS | 2 |
| 2009 | Sharing Features among Dynamical Systems with Beta ProcessesabstractWe propose a Bayesian nonparametric approach to relating multiple time series via a set of latent, dynamical behaviors. Using a beta process prior, we allow data-driven selection of the size of this set, as well as the pattern with which behaviors are shared among time series. Via the Indian buffet process representation of the beta process predictive distributions, we develop an exact Markov chain Monte Carlo inference method. In particular, our approach uses the sum-product algorithm to efficiently compute Metropolis-Hastings acceptance probabilities, and explores new dynamical behaviors via birth/death proposals. We validate our sampling algorithm using several synthetic datasets, and also demonstrate promising unsupervised segmentation of visual motion capture data. Emily B. Fox, Erik B. Sudderth, Michael I. Jordan, Alan S. Willsky |
NIPS | 2 |
| 2009 | Guest Editors' Introduction to the Special Section on Probabilistic Graphical ModelsabstractThe ten papers in this special section focus on applications of probabilistic graphical models in all areas of computer vision. Jiebo Luo 0001, Dimitris N. Metaxas, Antonio Torralba 0001, Thomas S. Huang, Erik B. Sudderth |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2008 | An HDP-HMM for systems with state persistenceabstractThe hierarchical Dirichlet process hidden Markov model (HDP-HMM) is a flexible, nonparametric model which allows state spaces of unknown size to be learned from data. We demonstrate some limitations of the original HDP-HMM formulation (Teh et al., 2006), and propose a sticky extension which allows more robust learning of smoothly varying dynamics. Using DP mixtures, this formulation also allows learning of more complex, multimodal emission distributions. We further develop a sampling algorithm that employs a truncated approximation of the DP to jointly resample the full state sequence, greatly improving mixing rates. Via extensive experiments with synthetic data and the NIST speaker diarization database, we demonstrate the advantages of our sticky extension, and the utility of the HDP-HMM in real-world applications. Emily B. Fox, Erik B. Sudderth, Michael I. Jordan, Alan S. Willsky |
ICML | 2 |
| 2008 | Nonparametric Bayesian Learning of Switching Linear Dynamical SystemsabstractMany nonlinear dynamical phenomena can be effectively modeled by a system that switches among a set of conditionally linear dynamical modes. We consider two such models: the switching linear dynamical system (SLDS) and the switching vector autoregressive (VAR) process. In this paper, we present a nonparametric approach to the learning of an unknown number of persistent, smooth dynamical modes by utilizing a hierarchical Dirichlet process prior. We develop a sampling algorithm that combines a truncated approximation to the Dirichlet process with an efficient joint sampling of the mode and state sequences. The utility and flexibility of our model are demonstrated on synthetic data, sequences of dancing honey bees, and the IBOVESPA stock index. Emily B. Fox, Erik B. Sudderth, Michael I. Jordan, Alan S. Willsky |
NIPS | 2 |
| 2008 | Shared Segmentation of Natural Scenes Using Dependent Pitman-Yor ProcessesabstractWe develop a statistical framework for the simultaneous, unsupervised segmentation and discovery of visual object categories from image databases. Examining a large set of manually segmented scenes, we use chi--square tests to show that object frequencies and segment sizes both follow power law distributions, which are well modeled by the Pitman--Yor (PY) process. This nonparametric prior distribution leads to learning algorithms which discover an unknown set of objects, and segmentation methods which automatically adapt their resolution to each image. Generalizing previous applications of PY processes, we use Gaussian processes to discover spatially contiguous segments which respect image boundaries. Using a novel family of variational approximations, our approach produces segmentations which compare favorably to state--of--the--art methods, while simultaneously discovering categories shared among natural scenes. Erik B. Sudderth, Michael I. Jordan |
NIPS | 1 |
| 2008 | Describing Visual Scenes Using Transformed Objects and Parts
Erik B. Sudderth, Antonio Torralba 0001, William T. Freeman, Alan S. Willsky |
Int. J. Comput. Vis. | 1 |
| 2007 | Hierarchical Dirichlet processes for tracking maneuvering targetsabstractWe consider the problem of state estimation for a dynamic system driven by unobserved, correlated inputs. We model these inputs via an uncertain set of temporally correlated dynamic models, where this uncertainty includes the number of modes, their associated statistics, and the rate of mode transitions. The dynamic system is formulated via two interacting graphs: a hidden Markov model (HMM) and a linear-Gaussian state space model. The HMM's state space indexes system modes, while its outputs are the unobserved inputs to the linear dynamical system. This Markovian structure accounts for temporal persistence of input regimes, but avoids rigid assumptions about their detailed dynamics. Via a hierarchical Dirichlet process (HDP) prior, the complexity of our infinite state space robustly adapts to new observations. We present a learning algorithm and computational results that demonstrate the utility of the HDP for tracking, and show that it efficiently learns typical dynamics from noisy data. Emily B. Fox, Erik B. Sudderth, Alan S. Willsky |
FUSION | 2 |
| 2007 | Learning Multiscale Representations of Natural Scenes Using Dirichlet ProcessesabstractWe develop nonparametric Bayesian models for multiscale representations of images depicting natural scene categories. Individual features or wavelet coefficients are marginally described by Dirichlet process (DP) mixtures, yielding the heavy-tailed marginal distributions characteristic of natural images. Dependencies between features are then captured with a hidden Markov tree, and Markov chain Monte Carlo methods used to learn models whose latent state space grows in complexity as more images are observed. By truncating the potentially infinite set of hidden states, we are able to exploit efficient belief propagation methods when learning these hierarchical Dirichlet process hidden Markov trees (HDP-HMTs) from data. We show that our generative models capture interesting qualitative structure in natural scenes, and more accurately categorize novel images than models which ignore spatial relationships among features. Jyri J. Kivinen, Erik B. Sudderth, Michael I. Jordan |
ICCV | 2 |
| 2007 | Image Denoising with Nonparametric Hidden Markov TreesabstractWe develop a hierarchical, nonparametric statistical model for wavelet representations of natural images. Extending previous work on Gaussian scale mixtures, wavelet coefficients are marginally distributed according to infinite, Dirichlet process mixtures. A hidden Markov tree is then used to couple the mixture assignments at neighboring nodes. Via a Monte Carlo learning algorithm, the resulting hierarchical Dirichlet process hidden Markov tree (HDP-HMT) model automatically adapts to the complexity of different images and wavelet bases. Image denoising results demonstrate the effectiveness of this learning process. Jyri J. Kivinen, Erik B. Sudderth, Michael I. Jordan |
ICIP (3) | 2 |
| 2007 | Loop Series and Bethe Variational Bounds in Attractive Graphical ModelsabstractVariational methods are frequently used to approximate or bound the partition or likelihood function of a Markov random field. Methods based on mean field theory are guaranteed to provide lower bounds, whereas certain types of convex relaxations provide upper bounds. In general, loopy belief propagation (BP) provides (often accurate) approximations, but not bounds. We prove that for a class of attractive binary models, the value specified by any fixed point of loopy BP always provides a lower bound on the true likelihood. Empirically, this bound is much better than the naive mean field bound, and requires no further work than running BP. We establish these lower bounds using a loop series expansion due to Chertkov and Chernyak, which we show can be derived as a consequence of the tree reparameterization characterization of BP fixed points. Erik B. Sudderth, Martin J. Wainwright, Alan S. Willsky |
NIPS | 1 |
| 2006 | Depth from Familiar Objects: A Hierarchical Model for 3D ScenesabstractWe develop an integrated, probabilistic model for the appearance and three-dimensional geometry of cluttered scenes. Object categories are modeled via distributions over the 3D location and appearance of visual features. Uncertainty in the number of object instances depicted in a particular image is then achieved via a transformed Dirichlet process. In contrast with image-based approaches to object recognition, we model scale variations as the perspective projection of objects in different 3D poses. To calibrate the underlying geometry, we incorporate binocular stereo images into the training process. A robust likelihood model accounts for outliers in matched stereo features, allowing effective learning of 3D object structure from partial 2D segmentations. Applied to a dataset of office scenes, our model detects objects at multiple scales via a coarse reconstruction of the corresponding 3D geometry. Erik B. Sudderth, Antonio Torralba 0001, William T. Freeman, Alan S. Willsky |
CVPR (2) | 1 |
| 2005 | Learning Hierarchical Models of Scenes, Objects, and PartsabstractWe describe a hierarchical probabilistic model for the detection and recognition of objects in cluttered, natural scenes. The model is based on a set of parts which describe the expected appearance and position, in an object centered coordinate frame, of features detected by a low-level interest operator. Each object category then has its own distribution over these parts, which are shared between objects. We learn the parameters of this model via a Gibbs sampler which uses the graphical model's structure to analytically average over many parameters. Applied to a database of images of isolated objects, the sharing of parts among objects improves detection accuracy when few training examples are available. We also extend this hierarchical framework to scenes containing multiple objects Erik B. Sudderth, Antonio Torralba 0001, William T. Freeman, Alan S. Willsky |
ICCV | 1 |
| 2005 | Describing Visual Scenes using Transformed Dirichlet ProcessesabstractMotivated by the problem of learning to detect and recognize objects with minimal supervision, we develop a hierarchical probabilistic model for the spatial structure of visual scenes. In contrast with most existing models, our approach explicitly captures uncertainty in the number of object instances depicted in a given image. Our scene model is based on the transformed Dirichlet process (TDP), a novel extension of the hierarchical DP in which a set of stochastically transformed mixture components are shared between multiple groups of data. For visual scenes, mixture components describe the spatial structure of visual features in an objectcentered coordinate frame, while transformations model the object positions in a particular image. Learning and inference in the TDP, which has many potential applications beyond computer vision, is based on an empirically effective Gibbs sampler. Applied to a dataset of partially labeled street scenes, we show that the TDP's inclusion of spatial structure improves detection performance, flexibly exploiting partially labeled training images. Erik B. Sudderth, Antonio Torralba 0001, William T. Freeman, Alan S. Willsky |
NIPS | 1 |
| 2004 | Distributed Occlusion Reasoning for Tracking with Nonparametric Belief PropagationabstractWe describe a threedimensional geometric hand model suitable for vi- sual tracking applications. The kinematic constraints implied by the model's joints have a probabilistic structure which is well described by a graphical model. Inference in this model is complicated by the hand's many degrees of freedom, as well as multimodal likelihoods caused by ambiguous image measurements. We use nonparametric belief propaga- tion (NBP) to develop a tracking algorithm which exploits the graph's structure to control complexity, while avoiding costly discretization. While kinematic constraints naturally have a local structure, self occlusions created by the imaging process lead to complex interpenden- cies in color and edgebased likelihood functions. However, we show that local structure may be recovered by introducing binary hidden vari- ables describing the occlusion state of each pixel. We augment the NBP algorithm to infer these occlusion variables in a distributed fashion, and then analytically marginalize over them to produce hand position esti- mates which properly account for occlusion events. We provide simula- tions showing that NBP may be used to refine inaccurate model initializa- tions, as well as track hand motion through extended image sequences. 1 Introduction Accurate visual detection and tracking of threedimensional articulated objects is a chal- lenging problem with applications in humancomputer interfaces, motion capture, and scene understanding [1]. In this paper, we develop a probabilistic method for tracking a geometric hand model from monocular image sequences. Because articulated hand mod- els have many (roughly 26) degrees of freedom, exact representation of the posterior dis- tribution over model configurations is intractable. Trackers based on extended and un- scented Kalman filters [2, 3] have difficulties with the multimodal uncertainties produced by ambiguous image evidence. This has motived many researchers to consider nonparamet- ric representations, including particle filters [4, 5] and deterministic multiscale discretiza- tions [6]. However, the hand's high dimensionality can cause these trackers to suffer catas- trophic failures, requiring the use of models which limit the hand's motion [4] or sophisti- cated prior models of hand configurations and dynamics [5, 6]. An alternative way to address the high dimensionality of articulated tracking problems is to describe the posterior distribution's statistical structure using a graphical model. Graph- Figure 1: Projected edges (left block) and silhouettes (right block) for a configuration of the 3D structural hand model matching the given image. To aid visualization, the model is also projected following rotations by 35 (center) and 70 (right) about the vertical axis. ical models have been used to track viewbased human body representations [7], con- tour models of restricted hand configurations [8], viewbased 2.5D "cardboard" models of hands and people [9], and a full 3D kinematic human body model [10]. Because the variables in these graphical models are continuous, and discretization is intractable for threedimensional models, most traditional graphical inference algorithms are inapplica- ble. Instead, these trackers are based on recently proposed extensions of particle filters to general graphs: mean field Monte Carlo in [9], and nonparametric belief propagation (NBP) [11, 12] in [10]. In this paper, we show that NBP may be used to track a threedimensional geometric model of the hand. To derive a graphical model for the tracking problem, we consider a redun- dant local representation in which each hand component is described by its own three dimensional position and orientation. We show that the model's kinematic constraints, including selfintersection constraints not captured by joint angle representations, take a simple form in this local representation. We also provide a local decomposition of the likelihood function which properly handles occlusion in a distributed fashion, a significant improvement over our earlier tracking results [13]. We conclude with simulations demon- strating our algorithm's robustness to occlusions. 2 Geometric Hand Modeling Structurally, the hand is composed of sixteen approximately rigid components: three pha- langes or links for each finger and thumb, as well as the palm [1]. As proposed by [2, 3], we model each rigid body by one or more truncated quadrics (ellipsoids, cones, and cylin- ders) of fixed size. These geometric primitives are well matched to the true geometry of the hand, allow tracking from arbitrary orientations (in contrast to 2.5D "cardboard" mod- els [5, 9]), and permit efficient computation of projected boundaries and silhouettes [3]. Figure 1 shows the edges and silhouettes corresponding to a sample hand model configu- ration. Note that only a coarse model of the hand's geometry is necessary for tracking. 2.1 Kinematic Representation and Constraints The kinematic constraints between different hand model components are well described by revolute joints [1]. Figure 2(a) shows a graph describing this kinematic structure, in which nodes correspond to rigid bodies and edges to joints. The two joints connecting the phalanges of each finger and thumb have a single rotational degree of freedom, while the joints connecting the base of each finger to the palm have two degrees of freedom (cor- responding to grasping and spreading motions). These twenty angles, combined with the palm's global position and orientation, provide 26 degrees of freedom. Forward kinematic transformations may be used to determine the finger positions corresponding to a given set of joint angles. While most modelbased hand trackers use this joint angle parameteriza- tion, we instead explore a redundant representation in which the ith rigid body is described by its position qi and orientation ri (a unit quaternion). Let xi = (qi, ri) denote this local description of each component, and x = {x1, . . . , x16} the overall hand configuration. Clearly, there are dependencies among the elements of x implied by the kinematic con- (a) (b) (c) (d) Figure 2: Graphs describing the hand model's constraints. (a) Kinematic constraints (EK ) de- rived from revolute joints. (b) Structural constraints (ES) preventing 3D component intersections. (c) Dynamics relating two consecutive time steps. (d) Occlusion consistency constraints (EO). straints. Let EK be the set of all pairs of rigid bodies which are connected by joints, or equivalently the edges in the kinematic graph of Fig. 2(a). For each joint (i, j) EK , define an indicator function K (x i,j i, xj ) which is equal to one if the pair (xi, xj ) are valid rigid body configurations associated with some setting of the angles of joint (i, j), and zero otherwise. Viewing the component configurations xi as random variables, the following prior explicitly enforces all constraints implied by the original joint angle representation: pK(x) K (x i,j i, xj ) (1) (i,j)EK Equation (1) shows that pK (x) is an undirected graphical model, whose Markov structure is described by the graph representing the hand's kinematic structure (Fig. 2(a)). 2.2 Structural and Temporal Constraints In reality, the hand's joint angles are coupled because different fingers can never occupy the same physical volume. This constraint is complex in a joint angle parameterization, but simple in our local representation: the position and orientation of every pair of rigid bodies must be such that their component quadric surfaces do not intersect. We approximate this ideal constraint in two ways. First, we only explicitly constrain those pairs of rigid bodies which are most likely to intersect, corresponding to the edges ES of the graph in Fig. 2(b). Furthermore, because the relative orientations of each finger's quadrics are implicitly constrained by the kinematic prior pK (x), we may detect most intersections based on the distance between object centroids. The structural prior is then given by 1 ||q p i - qj || > i,j S (x) S (x (x i,j i, xj ) S i,j i, xj ) = (2) 0 otherwise (i,j)ES where i,j is determined from the quadrics composing rigid bodies i and j. Empirically, we find that this constraint helps prevent different fingers from tracking the same image data. In order to track hand motion, we must model the hand's dynamics. Let xt denote the i position and orientation of the ith hand component at time t, and xt = {xt1, . . . , xt16}. For each component at time t, our dynamical model adds a Gaussian potential connecting it to the corresponding component at the previous time step (see Fig. 2(c)): 16 pT xt | xt-1 = N xt - xt-1; 0, i i i (3) i=1 Although this temporal model is factorized, the kinematic constraints at the following time step implicitly couple the corresponding random walks. These dynamics can be justified as the maximum entropy model given observations of the nodes' marginal variances i. 3 Observation Model Skin colored pixels have predictable statistics, which we model using a histogram distribu- tion pskin estimated from training patches [14]. Images without people were used to create a histogram model pbkgd of nonskin pixels. Let (x) denote the silhouette of projected hand configuration x. Then, assuming pixels are independent, an image y has likelihood p p skin(u) C (y | x) = pskin(u) pbkgd(v) (4) pbkgd(u) u(x) v(x) u(x) The final expression neglects the proportionality constant p v bkgd(v), which is inde- pendent of x, and thereby limits computation to the silhouette region [8]. 3.1 Distributed Occlusion Reasoning In configurations where there is no selfocclusion, pC (y | x) decomposes as a product of local likelihood terms involving the projections (xi) of individual hand components [13]. To allow a similar decomposition (and hence distributed inference) when there is occlu- sion, we augment the configuration xi of each node with a set of binary hidden variables zi = {zi } = 0 if pixel u in the projection of rigid body i is occluded (u) u. Letting zi(u) by any other body, and 1 otherwise, the color likelihood (eq. (4)) may be rewritten as 16 p z 16 i(u) p skin(u) C (y | x, z) = = p p C (y | xi, zi) (5) bkgd(u) i=1 u(xi) i=1 Assuming they are set consistently with the hand configuration x, the hidden occlusion variables z ensure that the likelihood of each pixel in (x) is counted exactly once. We may enforce consistency of the occlusion variables using the following function: 0 if x = 1 (x j occludes xi, u (xj ), and zi(u) j , zi ; x (6) (u) i) = 1 otherwise Note that because our rigid bodies are convex and nonintersecting, they can never take mutually occluding configurations. The constraint (xj, zi ; x (u) i) is zero precisely when pixel u in the projection of xi should be occluded by xj, but zi is in the unoccluded state. (u) The following potential encodes all of the occlusion relationships between nodes i and j: O (x (x ; x ; x i,j i, zi, xj , zj ) = j , zi(u) i) (xi, zj(u) j ) (7) u These occlusion constraints exist between all pairs of nodes. As with the structural prior, we enforce only those pairs EO (see Fig. 2(d)) most prone to xj occlusion: pO(x, z) O (x i,j i, zi, xj , zj ) (8) (i,j)EO z y i(u) xi Figure 3 shows a factor graph for the occlusion relationships between xi and its neighbors, as well as the observation potential pC (y | xi, zi). x u k The occlusion potential (xj, zi ; x (u) i) has a very Figure 3: Factor graph showing weak dependence on xi, depending only on p(y | xi, zi), and the occlusion con- whether xi is behind xj relative to the camera. straints placed on xi by xj , xk. Dashed lines denote weak dependencies. The 3.2 Modeling Edge Filter Responses plate is replicated once per pixel. Edges provide another important hand tracking cue. Using boundaries labeled in training images, we estimated a histogram pon of the response of a derivative of Gaussian filter steered to the edge's orientation [8, 10]. A similar histogram poff was estimated for filter outputs at randomly chosen locations. Let (x) denote the oriented edges in the projection of model configuration x. Then, again assuming pixel independence, image y has edge likelihood p 16 p z 16 i(u) p on(u) on(u) E (y | x, z) = = p p E (y | xi, zi) (9) off (u) poff(u) u(x) i=1 u(xi) i=1 where we have used the same occlusion variables z to allow a local decomposition. 4 Nonparametric Belief Propagation Over the previous sections, we have shown that a redundant, local representation of the geometric hand model's configuration xt allows p (xt | yt), the posterior distribution of the hand model at time t given image observations yt, to be written as 16 p xt | yt pK(xt)pS(xt)pO(xt, zt) pC(yt | xt, zt)p , zt) i i E (yt | xti i (10) zt i=1 The summation marginalizes over the hidden occlusion variables zt, which were needed to locally decompose the edge and color likelihoods. When video frames are observed, the overall posterior distribution is given by p (x | y) p xt | yt pT (xt | xt-1) (11) t=1 Excluding the potentials involving occlusion variables, which we discuss in detail in Sec. 4.2, eq. (11) is an example of a pairwise Markov random field: p (x | y) i,j (xi, xj) i (xi, y) (12) (i,j)E iV Hand tracking can thus be posed as inference in a graphical model, a problem we propose to solve using belief propagation (BP) [15]. At each BP iteration, some node i V calculates a message m (x ij j ) to be sent to a neighbor j (i) {j | (i, j) E}: mn (x mn-1 (x ij j ) j,i (xj , xi) i (xi, y) ki i) dxi (13) xi k(i)\j At any iteration, each node can produce an approximation ^ p(xi | y) to the marginal distri- bution p (xi | y) by combining the incoming messages with the local observation: ^ pn(xi | y) i (xi, yi) mn (x ji i) (14) j(i) For treestructured graphs, the beliefs ^ pn(xi | y) will converge to the true marginals p (xi | y). On graphs with cycles, BP is approximate but often highly accurate [15]. 4.1 Nonparametric Representations For the hand tracking problem, the rigid body configurations xi are sixdimensional con- tinuous variables, making accurate discretization intractable. Instead, we employ nonpara- metric, particlebased approximations to these messages using the nonparametric belief propagation (NBP) algorithm [11, 12]. In NBP, each message is represented using either a samplebased density estimate (a mixture of Gaussians) or an analytic function. Both types of messages are needed for hand tracking, as we discuss below. Each NBP message update involves two stages: sampling from the estimated marginal, followed by Monte Carlo ap- proximation of the outgoing message. For the general form of these updates, see [11]; the following sections focus on the details of the hand tracking implementation. The hand tracking application is complicated by the fact that the orientation component ri of xi = (qi, ri) is an element of the rotation group SO(3). Following [10], we represent orientations as unit quaternions, and use a linearized approximation when constructing den- sity estimates, projecting samples back to the unit sphere as necessary. This approximation is most appropriate for densities with tightly concentrated rotational components. 4.2 Marginal Computation BP's estimate of the belief ^ p(xi | y) is equal to the product of the incoming messages from neighboring nodes with the local observation potential (see eq. (14)). NBP approximates this product using importance sampling, as detailed in [13] for cases where there is no selfocclusion. First, M samples are drawn from the product of the incoming kinematic and temporal messages, which are Gaussian mixtures. We use a recently proposed multi- scale Gibbs sampler [16] to efficiently draw accurate (albeit approximate) samples, while avoiding the exponential cost associated with direct sampling (a product of d M Gaussian mixtures contains M d Gaussians). Following normalization of the rotational component, each sample is assigned a weight equal to the product of the color and edge likelihoods with any structural messages. Finally, the computationally efficient "rule of thumb" heuris- tic [17] is used to set the bandwidth of Gaussian kernels placed around each sample. To derive BP updates for the occlusion masks zi, we first cluster (xi, zi) for each hand component so that p (xt, zt | yt) has a pairwise form (as in eq. (12)). In principle, NBP could manage occlusion constraints by sampling candidate occlusion masks zi along with rigid body configurations xi. However, due to the exponentially large number of possible occlusion masks, we employ a more efficient analytic approximation. Consider the BP message sent from xj to (zi, xi), calculated by applying eq. (13) to the occlusion potential (x ; x u j , zi(u) i). We assume that ^ p(xj | y) is well separated from any candidate xi, a situation typically ensured by the kinematic and structural constraints. The occlusion constraint's weak dependence on xi (see Fig. 3) then separates the message computation into two cases. If xi lies in front of typical xj configurations, the BP message j,i(u)(zi ) is uninformative. If x (u) i is occluded, the message approximately equals j,i(u)(zi = 0) = 1 = 1) = 1 - Pr [u (x (u) j,i(u)(zi(u) j )] (15) where we have neglected correlations among pixel occlusion states, and where the prob- ability is computed with respect to ^ p(xj | y). By taking the product of these messages k,i(u)(zi ) from all potential occluders x (u) k and normalizing, we may determine an ap- proximation to the marginal occlusion probability i Pr[z = 0]. (u) i(u) Because the color likelihood pC (y | xi, zi) factorizes across pixels u, the BP approximation to pC (y | xi) may be written in terms of these marginal occlusion probabilites: p p skin(u) C (y | xi) i + (1 - ) (16) (u) i(u) pbkgd(u) u(xi) Intuitively, this equation downweights the color evidence at pixel u as the probability of that pixel's occlusion increases. The edge likelihood pE(y | xi) averages over zi similarly. The NBP estimate of ^ p(xi | y) is determined by sampling configurations of xi as before, and reweighting them using these occlusionsensitive likelihood functions. 4.3 Message Propagation To derive the propagation rule for nonocclusion edges, as suggested by [18] we rewrite the message update equation (13) in terms of the marginal distribution ^ p(xi | y): ^ pn-1(x mn (x i | y) dx ij j ) = j,i (xj , xi) i (17) x mn-1 (x i ji i) Our explicit use of the current marginal estimate ^ pn-1(xi | y) helps focus the Monte Carlo approximation on the most important regions of the state space. Note that messages sent 1 2 1 2 Figure 4: Refinement of a coarse initialization following one and two NBP iterations, both without (left) and with (right) occlusion reasoning. Each plot shows the projection of the five most significant modes of the estimated marginal distributions. Note the difference in middle finger estimates. along kinematic, structural, and temporal edges depend only on the belief ^ p(xi | y) follow- ing marginalization over occlusion variables zi. Details and pseudocode for the message propagation step are provided in [13]. For kine- matic constraints, we sample uniformly among permissable joint angles, and then use forward kinematics to propagate samples from ^ pn-1(xi | y) /mn-1 (x ji i) to hypothesized configurations of xj. Following [12], temporal messages are determined by adjusting the bandwidths of the current marginal estimate ^ p(xi | y) to match the temporal covariance i. Because structural potentials (eq. (2)) equal one for all state configurations outside some ball, the ideal structural messages are not finitely integrable. We therefore approximate the structural message m (x ij j ) as an analytic function equal to the weights of all kernels in ^ p(xi | y) outside a ball centered at qj, the position of xj. Erik B. Sudderth, Michael I. Mandel, William T. Freeman, Alan S. Willsky |
NIPS | 1 |
| 2003 | Nonparametric Belief PropagationabstractIn many applications of graphical models arising in computer vision, the hidden variables of interest are most naturally specified by continuous, non-Gaussian distributions. There exist inference algorithms for discrete approximations to these continuous distributions, but for the high-dimensional variables typically of interest, discrete inference becomes infeasible. Stochastic methods such as particle filters provide an appealing alternative. However, existing techniques fail to exploit the rich structure of the graphical models describing many vision problems. Drawing on ideas from regularized particle filters and belief propagation (BP), this paper develops a nonparametric belief propagation (NBP) algorithm applicable to general graphs. Each NBP iteration uses an efficient sampling procedure to update kernel-based approximations to the true, continuous likelihoods. The algorithm can accommodate an extremely broad class of potential functions, including nonparametric representations. Thus, NBP extends particle filtering methods to the more general vision problems that graphical models can describe. We apply the NBP algorithm to infer component interrelationships in a parts-based face model, allowing location and reconstruction of occluded features. Erik B. Sudderth, Alexander Ihler, William T. Freeman, Alan S. Willsky |
CVPR (1) | 1 |
| 2003 | Efficient Multiscale Sampling from Products of Gaussian MixturesabstractThe problem of approximating the product of several Gaussian mixture distributions arises in a number of contexts, including the nonparametric belief propagation (NBP) inference algorithm and the training of prod- uct of experts models. This paper develops two multiscale algorithms for sampling from a product of Gaussian mixtures, and compares their performance to existing methods. The first is a multiscale variant of pre- viously proposed Monte Carlo techniques, with comparable theoretical guarantees but improved empirical convergence rates. The second makes use of approximate kernel density evaluation methods to construct a fast approximate sampler, which is guaranteed to sample points to within a tunable parameter (cid:15) of their true probability. We compare both multi- scale samplers on a set of computational examples motivated by NBP, demonstrating significant improvements over existing methods. Alexander Ihler, Erik B. Sudderth, William T. Freeman, Alan S. Willsky |
NIPS | 2 |
| 2000 | Tree-Based Modeling and Estimation of Gaussian Processes on Graphs with CyclesabstractWe present the embedded trees algorithm, an iterative technique for estimation of Gaussian processes defined on arbitrary graphs. By exactly solving a series of modified problems on embedded span(cid:173) ning trees, it computes the conditional means with an efficiency comparable to or better than other techniques. Unlike other meth(cid:173) ods, the embedded trees algorithm also computes exact error co(cid:173) variances. The error covariance computation is most efficient for graphs in which removing a small number of edges reveals an em(cid:173) bedded tree. In this context, we demonstrate that sparse loopy graphs can provide a significant increase in modeling power rela(cid:173) tive to trees, with only a minor increase in estimation complexity. Martin J. Wainwright, Erik B. Sudderth, Alan S. Willsky |
NIPS | 2 |