John W. Fisher III

dblp:23/954 · DBLP profile ↗
← Back
89ranked-venue papers
8as first author
0since 2021 · last 2020
0000-0003-4844-3495ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 50 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 first-authorSystems, architecture and hardware · 4Computer networks · 2Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
41 papers
Probabilistic and Bayesian machine learning · 32% 3D vision · 25% Planning, search and constraint satisfaction · 12%
Computer graphics and multimedia
14 papers
Image and video processing · 62% Geometric modeling and processing · 15% Computational photography and imaging · 8%
Theoretical computer science
7 papers
Mathematical optimization · 50% Algorithmic game theory and mechanism design · 23% Information theory · 16%

Topics — the 30 heaviest of 118, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.852020
Streaming, Distributed Variational Inference for Bayesian Nonparametrics · NIPS 2015
Bayesian Nonparametric Intrinsic Image Decomposition · ECCV (4) 2014
Coupling Nonparametric Mixtures via Latent Dirichlet Processes · NIPS 2012
Robotics › Robot navigation and mapping
sensor planning
0.622018
Task-Specific Sensor Planning for Robotic Assembly Tasks · ICRA 2018
Information-Driven Adaptive Structured-Light Scanners · CVPR 2016
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.542014
Parallel Sampling of HDPs using Sub-Cluster Splits · NIPS 2014
Parallel Sampling of DP Mixture Models using Sub-Cluster Splits · NIPS 2013
Efficient MCMC sampling with implicit shape representations · CVPR 2011
Computer vision › 3D vision › 3d motion analysis
articulated motion analysis
0.522020
Nonparametric Object and Parts Modeling With Lie Group Dynamics · CVPR 2020
Recovering Articulated Model Topology from Observed Rigid Motion · NIPS 2002
Machine learning › Reinforcement learning › hierarchical reinforcement learning
macro-action discovery
0.412020
Belief-Dependent Macro-Action Discovery in POMDPs using the Value of Information · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty
0.412020
Belief-Dependent Macro-Action Discovery in POMDPs using the Value of Information · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
POMDP planning
0.412020
Belief-Dependent Macro-Action Discovery in POMDPs using the Value of Information · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information
0.412020
Belief-Dependent Macro-Action Discovery in POMDPs using the Value of Information · NeurIPS 2020
Algorithmic game theory and mechanism design
resource allocation
0.412020
Sequential Bayesian Experimental Design with Variable Cost Structure · NeurIPS 2020
Image and video processing
image registration
0.432017
Highly-Expressive Spaces of Well-Behaved Transformations: Keeping it Simple · ICCV 2015
Automatic registration of LIDAR and optical images of urban scenes · CVPR 2009
Transformations Based on Continuous Piecewise-Affine Velocity Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Machine learning › Generative modeling › diffusion model
parallel sampling
0.422014
Parallel Sampling of HDPs using Sub-Cluster Splits · NIPS 2014
Parallel Sampling of DP Mixture Models using Sub-Cluster Splits · NIPS 2013
Image and video processing
image segmentation
0.332013
A Video Representation Using Temporal Superpixels · CVPR 2013
Low level vision via switchable Markov random fields · CVPR 2012
Analysis of orientation and scale in smoothly varying textures · ICCV 2009
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation
0.342015
Adaptive Belief Propagation · ICML 2015
Loopy Belief Propagation: Convergence and Effects of Message Errors · J. Mach. Learn. Res. 2005
Message Errors in Belief Propagation · NIPS 2004
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty › information gathering › informative planning
information-theoretic planning
0.312018
A Robust Approach to Sequential Information Theoretic Planning · ICML 2018
Knowledge, reasoning and agents › Multi-agent systems
multi-robot systems
0.312018
Task-Specific Sensor Planning for Robotic Assembly Tasks · ICRA 2018
Knowledge, reasoning and agents › Multi-agent systems › multi-robot coordination
multi-robot task planning
0.312018
Task-Specific Sensor Planning for Robotic Assembly Tasks · ICRA 2018
Machine learning › Representation and self-supervised learning › mutual information
mutual information estimation
0.312018
A Robust Approach to Sequential Information Theoretic Planning · ICML 2018
Machine learning › Reinforcement learning
sequential experimental design
0.312018
A Robust Approach to Sequential Information Theoretic Planning · ICML 2018
Computer vision › 3D vision
surface normal estimation
0.312018
The Manhattan Frame Model - Manhattan World Inference in the Space of Surface Normals · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
0.312018
Task-Specific Sensor Planning for Robotic Assembly Tasks · ICRA 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.322013
Parallel Sampling of DP Mixture Models using Sub-Cluster Splits · NIPS 2013
Coupling Nonparametric Mixtures via Latent Dirichlet Processes · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.322012
Low level vision via switchable Markov random fields · CVPR 2012
Manifold guided composite of Markov random fields for image modeling · CVPR 2012
Computer vision › 3D vision
point cloud registration
0.312017
Efficient Global Point Cloud Alignment Using Bayesian Nonparametric Mixtures · CVPR 2017
Image and video processing › image restoration
image denoising
0.322012
Low level vision via switchable Markov random fields · CVPR 2012
Manifold guided composite of Markov random fields for image modeling · CVPR 2012
Image and video processing › image restoration
image inpainting
0.322012
Low level vision via switchable Markov random fields · CVPR 2012
Manifold guided composite of Markov random fields for image modeling · CVPR 2012
Computer vision › 3D vision
3d reconstruction
0.322014
Aerial Reconstructions via Probabilistic Data Fusion · CVPR 2014
Automatic registration of LIDAR and optical images of urban scenes · CVPR 2009
Computational photography and imaging
intrinsic image decomposition
0.322014
Bayesian Nonparametric Intrinsic Image Decomposition · ECCV (4) 2014
Analysis of orientation and scale in smoothly varying textures · ICCV 2009
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model
0.322012
Coupling Nonparametric Mixtures via Latent Dirichlet Processes · NIPS 2012
Construction of Dependent Dirichlet Processes based on Poisson Processes · NIPS 2010
Robotics › Robot navigation and mapping › active perception
active sensing
0.212016
Information-Driven Adaptive Structured-Light Scanners · CVPR 2016
Computer vision › 3D vision › range sensing
depth sensing
0.212016
Information-Driven Adaptive Structured-Light Scanners · CVPR 2016

Methods — techniques the papers use, named apart from their topics

markov chain monte carlo · 1.0gibbs sampling · 0.8probabilistic generative model · 0.8particle filtering · 0.7metropolis-hastings · 0.5value of information · 0.4two-sided bounds · 0.4regret bounds · 0.4lie group dynamics · 0.4knapsack optimization · 0.4best-arm identification · 0.4coreset · 0.4variational inference · 0.4split-merge proposal · 0.4tetrahedral tessellation · 0.3diffeomorphic integration · 0.3branch-and-bound · 0.3bayesian nonparametric mixture · 0.3
YearPublicationVenuePosition
2020 Nonparametric Object and Parts Modeling With Lie Group Dynamics
abstract
Articulated motion analysis often utilizes strong prior knowledge such as a known or trained parts model for humans. Yet, the world contains a variety of articulating objects--mammals, insects, mechanized structures--where the number and configuration of parts for a particular object is unknown in advance. Here, we relax such strong assumptions via an unsupervised, Bayesian nonparametric parts model that infers an unknown number of parts with motions coupled by a body dynamic and parameterized by SE(D), the Lie group of rigid transformations. We derive an inference procedure that utilizes short observation sequences (image, depth, point cloud or mesh) of an object in motion without need for markers or learned body models. Efficient Gibbs decompositions for inference over distributions on SE(D) demonstrate robust part decompositions of moving objects under both 3D and 2D observation models. The inferred representation permits novel analysis, such as object segmentation by relative part motion, and transfers to new observations of the same object type.
David S. Hayden, Jason L. Pacheco, John W. Fisher III
CVPR3
2020 Belief-Dependent Macro-Action Discovery in POMDPs using the Value of Information
abstract
This work introduces macro-action discovery using value-of-information (VoI) for robust and efficient planning in partially observable Markov decision processes (POMDPs). POMDPs are a powerful framework for planning under uncertainty. Previous approaches have used high-level macro-actions within POMDP policies to reduce planning complexity. However, macro-action design is often heuristic and rarely comes with performance guarantees. Here, we present a method for extracting belief-dependent, variable-length macro-actions directly from a low-level POMDP model. We construct macro-actions by chaining sequences of open-loop actions together when the task-specific value of information (VoI) --- the change in expected task performance caused by observations in the current planning iteration --- is low. Importantly, we provide performance guarantees on the resulting VoI macro-action policies in the form of bounded regret relative to the optimal policy. In simulated tracking experiments, we achieve higher reward than both closed-loop and hand-coded macro-action baselines, selectively using VoI macro-actions to reduce planning complexity while maintaining near-optimal task performance.
Genevieve Flaspohler, Nicholas Roy, John W. Fisher III
NeurIPS3
2020 Sequential Bayesian Experimental Design with Variable Cost Structure
abstract
Mutual information (MI) is a commonly adopted utility function in Bayesian optimal experimental design (BOED). While theoretically appealing, MI evaluation poses a significant computational burden for most real world applications. As a result, many algorithms utilize MI bounds as proxies that lack regret-style guarantees. Here, we utilize two-sided bounds to provide such guarantees. Bounds are successively refined/tightened through additional computation until a desired guarantee is achieved. We consider the problem of adaptively allocating computational resources in BOED. Our approach achieves the same guarantee as existing methods, but with fewer evaluations of the costly MI reward. We adapt knapsack optimization of best arm identification problems, with important differences that impact overall algorithm design and performance. First, observations of MI rewards are biased. Second, evaluating experiments incurs shared costs amongst all experiments (posterior sampling) in addition to per experiment costs that may vary with increasing evaluation. We propose and demonstrate an algorithm that accounts for these variable costs in the refinement decision.
Sue Zheng, David S. Hayden, Jason L. Pacheco, John W. Fisher III
NeurIPS4
2019 Variational Information Planning for Sequential Decision Making
abstract
We consider the setting of sequential decision making where, at each stage, potential actions are evaluated based on expected reduction in posterior uncertainty, given by mutual information (MI). As MI typically lacks a closed form, we propose an approach which maintains variational approximations of, both, the posterior and MI utility. Our planning objective extends an established variational bound on MI to the setting of sequential planning. The result, variational information planning (VIP), is an efficient method for sequential decision making. We further establish convexity of the variational planning objective and, under conditional exponential family approximations, we show that the optimal MI bound arises from a relaxation of the well-known exponential family moment matching property. We demonstrate VIP for sensor selection, experiment design, and active learning, where it meets or exceeds methods requiring more computation, or those specialized to the task.
Jason L. Pacheco, John W. Fisher III
AISTATS2
2019 Distributed MCMC Inference in Dirichlet Process Mixture Models Using Julia
abstract
Due to the increasing availability of large data sets, the need for general-purpose massively-parallel analysis tools become ever greater. In unsupervised learning, Bayesian nonparametric mixture models, exemplified by the Dirichlet-Process Mixture Model (DPMM), provide a principled Bayesian approach to adapt model complexity to the data. Despite their potential, however, DPMMs have yet to become a popular tool. This is partly due to the lack of friendly software tools that can handle large datasets efficiently. Here we show how, using Julia, one can achieve efficient and easily-modifiable implementation of distributed inference in DPMMs. Particularly, we show how a recent parallel MCMC inference algorithm - originally implemented in C++ for a single multi-core machine - can be distributed efficiently across multiple multi-core machines using a distributed-memory model. This leads to speedups, alleviates memory and storage limitations, and lets us learn DPMMs from significantly larger datasets and of higher dimensionality. It also turned out that even on a single machine the proposed Julia implementation handles higher dimensions more gracefully (at least for Gaussians) than the original C++ implementation. Finally, we use the proposed implementation to learn a model of image patches and apply the learned model for image denoising. While we speculate that a highly-optimized distributed implementation in, say, C++ could have been faster than the proposed implementation in Julia, from our perspective as machine-learning researchers (as opposed to HPC researchers), the latter also offers a practical and monetary value due to the ease of development and abstraction level. Our code is publicly available at https://github.com/dinarior/dpmm subclusters.jl
Or Dinari, Angel Yu, Oren Freifeld, John W. Fisher III
CCGRID4
2018 A Robust Approach to Sequential Information Theoretic Planning
abstract
In many sequential planning applications a natural approach to generating high quality plans is to maximize an information reward such as mutual information (MI). Unfortunately, MI lacks a closed form in all but trivial models, and so must be estimated. In applications where the cost of plan execution is expensive, one desires planning estimates which admit theoretical guarantees. Through the use of robust M-estimators we obtain bounds on absolute deviation of estimated MI. Moreover, we propose a sequential algorithm which integrates inference and planning by maximally reusing particles in each stage. We validate the utility of using robust estimators in the sequential approach on a Gaussian Markov Random Field wherein information measures have a closed form. Lastly, we demonstrate the benefits of our integrated approach in the context of sequential experiment design for inferring causal regulatory networks from gene expression levels. Our method shows improvements over a recent method which selects intervention experiments based on the same MI objective.
Sue Zheng, Jason L. Pacheco, John W. Fisher III
ICML3
2018 Task-Specific Sensor Planning for Robotic Assembly Tasks
abstract
When performing multi-robot tasks, sensory feedback is crucial in reducing uncertainty for correct execution. Yet the utilization of sensors should be planned as an integral part of the task planning, taken into account several factors such as the tolerance of different inferred properties of the scene and interaction with different agents. In this paper we handle this complex problem in a principled, yet efficient way. We use surrogate predictors based on open-loop simulation to estimate and bound the probability of success for specific tasks. We reason about such task-specific uncertainty approximants and their effectiveness. We show how they can be incorporated into a multi-robot planner, and demonstrate results with a team of robots performing assembly tasks.
Guy Rosman, Changhyun Choi, Mehmet Remzi Dogar, John W. Fisher III, Daniela Rus
ICRA4
2018 The Manhattan Frame Model - Manhattan World Inference in the Space of Surface Normals
abstract
Objects and structures within man-made environments typically exhibit a high degree of organization in the form of orthogonal and parallel planes. Traditional approaches utilize these regularities via the restrictive, and rather local, Manhattan World (MW) assumption which posits that every plane is perpendicular to one of the axes of a single coordinate system. The aforementioned regularities are especially evident in the surface normal distribution of a scene where they manifest as orthogonally-coupled clusters. This motivates the introduction of the Manhattan-Frame (MF) model which captures the notion of an MW in the surface normals space, the unit sphere, and two probabilistic MF models over this space. First, for a single MF we propose novel real-time MAP inference algorithms, evaluate their performance and their use in drift-free rotation estimation. Second, to capture the complexity of real-world scenes at a global scale, we extend the MF model to a probabilistic mixture of Manhattan Frames (MMF). For MMF inference we propose a simple MAP inference algorithm and an adaptive Markov-Chain Monte-Carlo sampling algorithm with Metropolis-Hastings split/merge moves that let us infer the unknown number of mixture components. We demonstrate the versatility of the MMF model and inference algorithm across several scales of man-made environments.
Julian Straub, Oren Freifeld, Guy Rosman, John J. Leonard, John W. Fisher III
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Efficient Global Point Cloud Alignment Using Bayesian Nonparametric Mixtures
abstract
Point cloud alignment is a common problem in computer vision and robotics, with applications ranging from 3D object recognition to reconstruction. We propose a novel approach to the alignment problem that utilizes Bayesian nonparametrics to describe the point cloud and surface normal densities, and branch and bound (BB) optimization to recover the relative transformation. BB uses a novel, refinable, near-uniform tessellation of rotation space using 4D tetrahedra, leading to more efficient optimization compared to the common axis-angle tessellation. We provide objective function bounds for pruning given the proposed tessellation, and prove that BB converges to the optimum of the cost function along with providing its computational complexity. Finally, we empirically demonstrate the efficiency of the proposed approach as well as its robustness to real-world conditions such as missing data and partial overlap.
Julian Straub, Trevor Campbell, Jonathan P. How, John W. Fisher III
CVPR4
2017 Transformations Based on Continuous Piecewise-Affine Velocity Fields
abstract
We propose novel finite-dimensional spaces of well-behaved transformations. The latter are obtained by (fast and highly-accurate) integration of continuous piecewise-affine velocity fields. The proposed method is simple yet highly expressive, effortlessly handles optional constraints (e.g., volume preservation and/or boundary conditions), and supports convenient modeling choices such as smoothing priors and coarse-to-fine analysis. Importantly, the proposed approach, partly due to its rapid likelihood evaluations and partly due to its other properties, facilitates tractable inference over rich transformation spaces, including using Markov-Chain Monte-Carlo methods. Its applications include, but are not limited to: monotonic regression (more generally, optimization over monotonic functions); modeling cumulative distribution functions or histograms; time-warping; image warping; image registration; real-time diffeomorphic image editing; data augmentation for image classifiers. Our GPU-based code is publicly available.
Oren Freifeld, Søren Hauberg, Kayhan Batmanghelich, John W. Fisher III
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Dreaming More Data: Class-dependent Distributions over Diffeomorphisms for Learned Data Augmentation
abstract
Data augmentation is a key element in training high-dimensional models. In this approach, one synthesizes new observations by applying pre-specified transformations to the original training data; e.g. new images are formed by rotating old ones. Current augmentation schemes, however, rely on manual specification of the applied transformations, making data augmentation an implicit form of feature engineering. With an eye towards true end-to-end learning, we suggest learning the applied transformations on a per-class basis. Particularly, we align image pairs within each class under the assumption that the spatial transformation between images belongs to a large class of diffeomorphisms. We then learn a class-specific probabilistic generative models of the transformations in a Riemannian submanifold of the Lie group of diffeomorphisms. We demonstrate significant performance improvements in training deep neural nets over manually-specified augmentation schemes. Our code and augmented datasets are available online.
Søren Hauberg, Oren Freifeld, Anders Boesen Lindbo Larsen, John W. Fisher III, Lars Kai Hansen
AISTATS4
2016 Information-Driven Adaptive Structured-Light Scanners
abstract
Sensor planning and active sensing, long studied in robotics, adapt sensor parameters to maximize a utility function while constraining resource expenditures. Here we consider information gain as the utility function. While these concepts are often used to reason about 3D sensors, these are usually treated as a predefined, black-box, component. In this paper we show how the same principles can be used as part of the 3D sensor. We describe the relevant generative model for structured-light 3D scanning and show how adaptive pattern selection can maximize information gain in an open-loop-feedback manner. We then demonstrate how different choices of relevant variable sets (corresponding to the subproblems of locatization and mapping) lead to different criteria for pattern selection and can be computed in an online fashion. We show results for both subproblems with several pattern dictionary choices and demonstrate their usefulness for pose estimation and depth acquisition.
Guy Rosman, Daniela Rus, John W. Fisher III
CVPR3
2016 Efficient Observation Selection in Probabilistic Graphical Models Using Bayesian Lower Bounds
Dilin Wang, John W. Fisher III, Qiang Liu 0001
UAI2
2015 A Theoretical Analysis of Optimization by Gaussian Continuation
abstract
Optimization via continuation method is a widely used approach for solving nonconvex minimization problems. While this method generally does not provide a global minimum, empirically it often achieves a superior local minimum compared to alternative approaches such as gradient descent. However, theoretical analysis of this method is largely unavailable. Here, we provide a theoretical analysis that provides a bound on the endpoint solution of the continuation method. The derived bound depends on a problem specific characteristic that we refer to as optimization complexity. We show that this characteristic can be analytically computed when the objective function is expressed in some suitable basis functions. Our analysis combines elements of scale-space theory, regularization and differential equations.
Hossein Mobahi, John W. Fisher III
AAAI2
2015 A Dirichlet Process Mixture Model for Spherical Data
abstract
Directional data, naturally represented as points on the unit sphere, appear in many applications. However, unlike the case of Euclidean data, flexible mixture models on the sphere that can capture correlations, handle an unknown number of components and extend readily to high-dimensional data have yet to be suggested. For this purpose we propose a Dirichlet process mixture model of Gaussian distributions in distinct tangent spaces (DP-TGMM) to the sphere. Importantly, the formulation of the proposed model allows the extension of recent advances in efficient inference for Bayesian nonparametric models to the spherical domain. Experiments on synthetic data as well as real-world 3D surface normal and 20-dimensional semantic word vector data confirm the expressiveness and applicability of the DP-TGMM.
Julian Straub, Jason Chang 0001, Oren Freifeld, John W. Fisher III
AISTATS4
2015 Small-variance nonparametric clustering on the hypersphere
abstract
Structural regularities in man-made environments reflect in the distribution of their surface normals. Describing these surface normal distributions is important in many computer vision applications, such as scene understanding, plane segmentation, and regularization of 3D reconstructions. Based on the small-variance limit of Bayesian nonparametric von-Mises-Fisher (vMF) mixture distributions, we propose two new flexible and efficient k-means-like clustering algorithms for directional data such as surface normals. The first, DP-vMF-means, is a batch clustering algorithm derived from the Dirichlet process (DP) vMF mixture. Recognizing the sequential nature of data collection in many applications, we extend this algorithm to DDP-vMF-means, which infers temporally evolving cluster structure from streaming data. Both algorithms naturally respect the geometry of directional data, which lies on the unit sphere. We demonstrate their performance on synthetic directional data and real 3D surface normals from RGB-D sensors. While our experiments focus on 3D data, both algorithms generalize to high dimensional directional data such as protein backbone configurations and semantic word vectors.
Julian Straub, Trevor Campbell, Jonathan P. How, John W. Fisher III
CVPR4
2015 Boosting crowdsourcing with expert labels: Local vs. global effects
Qiang Liu 0001, Alexander Ihler, John W. Fisher III
FUSION3
2015 Efficient information planning in Gaussian MRFs
Georgios Papachristoudis, John W. Fisher III
FUSION2
2015 On the complexity of information planning in Gaussian models
abstract
We analyze the complexity of evaluating information rewards for measurement selection in sparse graphical models under the assumption that measurements are drawn from a limited number of nodes subject to a finite budget. Previous analyses [1, 2, 3] exploit the submodular property of conditional mutual information to demonstrate that greedy measurement selection come with near-optimal guarantees As noted in [4] typical formulations assume oracle value models. However, [1, 2, 5] allude to a more significant source of complexity, namely computing the measurement reward. Here, we focus on Gaussian models and show that by exploiting sparsity in the measurement model, the complexity of planning is substantially reduced. We also demonstrate that by utilizing the information form additional significant reductions in complexity may be realized.
Georgios Papachristoudis, John W. Fisher III
ICASSP2
2015 Semantically-Aware Aerial Reconstruction from Multi-modal Data
abstract
We consider a methodology for integrating multiple sensors along with semantic information to enhance scene representations. We propose a probabilistic generative model for inferring semantically-informed aerial reconstructions from multi-modal data within a consistent mathematical framework. The approach, called Semantically-Aware Aerial Reconstruction (SAAR), not only exploits inferred scene geometry, appearance, and semantic observations to obtain a meaningful categorization of the data, but also extends previously proposed methods by imposing structure on the prior over geometry, appearance, and semantic labels. This leads to more accurate reconstructions and the ability to fill in missing contextual labels via joint sensor and semantic information. We introduce a new multi-modal synthetic dataset in order to provide quantitative performance analysis. Additionally, we apply the model to real-world data and exploit OpenStreetMap as a source of semantic observations. We show quantitative improvements in reconstruction accuracy of large-scale urban scenes from the combination of LiDAR, aerial photography, and semantic data. Furthermore, we demonstrate the model's ability to fill in for missing sensed data, leading to more interpretable reconstructions.
Randi Cabezas, Julian Straub, John W. Fisher III
ICCV3
2015 Highly-Expressive Spaces of Well-Behaved Transformations: Keeping it Simple
abstract
We propose novel finite-dimensional spaces of Rn→ Rntransformations, n ∈ {1, 2, 3}, derived from (continuously-defined) parametric stationary velocity fields. Particularly, we obtain these transformations, which are diffeomorphisms, by fast and highly-accurate integration of continuous piecewise-affine velocity fields, we also provide an exact solution for n = 1. The simple-yet-highly-expressive proposed representation handles optional constraints (e.g., volume preservation) easily and supports convenient modeling choices and rapid likelihood evaluations (facilitating tractable inference over latent transformations). Its applications include, but are not limited to: unconstrained optimization over monotonic functions, modeling cumulative distribution functions or histograms, time warping, image registration, landmark-based warping, real-time diffeomorphic image editing. Our code is available at https://github.com/freifeld/cpabDiffeo.
Oren Freifeld, Søren Hauberg, Kayhan Batmanghelich, John W. Fisher III
ICCV4
2015 A fast method for inferring high-quality simply-connected superpixels
abstract
Superpixel segmentation is a key step in many image processing and vision tasks. Our recently-proposed connectivity-constrained probabilistic model [1] yields high-quality super-pixels. Seemingly, however, connectivity constraints preclude parallelized inference. As such, the implementation from [1] is serial. The contributions of this work are as follows. First, we demonstrate that effective parallelization is possible via a fast GPU implementation that scales gracefully with both the number of pixels and number of superpixels. Second, we show that the superpixels are improved by replacing the fixed and restricted spatial covariances from [1] with a flexible Bayesian prior. Quantitative evaluation on public benchmarks shows the proposed method outperforms the state-of-the-art. We make our implementation publicly available.
Oren Freifeld, John W. Fisher III
ICIP3
2015 Adaptive Belief Propagation
abstract
Graphical models are widely used in inference problems. In practice, one may construct a single large-scale model to explain a phenomenon of interest, which may be utilized in a variety of settings. The latent variables of interest, which can differ in each setting, may only represent a small subset of all variables. The marginals of variables of interest may change after the addition of measurements at different time points. In such adaptive settings, naive algorithms, such as standard belief propagation (BP), may utilize many unnecessary computations by propagating messages over the entire graph. Here, we formulate an efficient inference procedure, termed adaptive BP (AdaBP), suitable for adaptive inference settings. We show that it gives exact results for trees in discrete and Gaussian Markov Random Fields (MRFs), and provide an extension to Gaussian loopy graphs. We also provide extensions on finding the most likely sequence of the entire latent graph. Lastly, we compare the proposed method to standard BP and to that of (Sumer et al., 2011), which tackles the same problem. We show in synthetic and real experiments that it outperforms standard BP by orders of magnitude and explore the settings that it is advantageous over (Sumer et al., 2011).
Georgios Papachristoudis, John W. Fisher III
ICML2
2015 Coresets for visual summarization with applications to loop closure
abstract
In continuously operating robotic systems, efficient representation of the previously seen camera feed is crucial. Using a highly efficient compression coreset method, we formulate a new method for hierarchical retrieval of frames from large video streams collected online by a moving robot. We demonstrate how to utilize the resulting structure for efficient loop-closure by a novel sampling approach that is adaptive to the structure of the video. The same structure also allows us to create a highly-effective search tool for large-scale videos, which we demonstrate in this paper. We show the efficiency of proposed approaches for retrieval and loop closure on standard datasets, and on a large-scale video from a mobile camera.
Mikhail Volkov 0002, Guy Rosman, Dan Feldman, John W. Fisher III, Daniela Rus
ICRA4
2015 Real-time manhattan world rotation estimation in 3D
abstract
Drift of the rotation estimate is a well known problem in visual odometry systems as it is the main source of positioning inaccuracy. We propose three novel algorithms to estimate the full 3D rotation to the surrounding Manhattan World (MW) in as short as 20 ms using surface-normals derived from the depth channel of a RGB-D camera. Importantly, this rotation estimate acts as a structure compass which can be used to estimate the bias of an odometry system, such as an inertial measurement unit (IMU), and thus remove its angular drift. We evaluate the run-time as well as the accuracy of the proposed algorithms on groundtruth data. They achieve zerodrift rotation estimation with RMSEs below 3.4° by themselves and below 2.8° when integrated with an IMU in a standard extended Kalman filter (EKF). Additional qualitative results show the accuracy in a large scale indoor environment as well as the ability to handle fast motion. Selected segmentations of scenes from the NYU depth dataset demonstrate the robustness of the inference algorithms to clutter and hint at the usefulness of the segmentation for further processing.
Julian Straub, Nishchal Bhandari, John J. Leonard, John W. Fisher III
IROS4
2015 Streaming, Distributed Variational Inference for Bayesian Nonparametrics
abstract
This paper presents a methodology for creating streaming, distributed inference algorithms for Bayesian nonparametric (BNP) models. In the proposed framework, processing nodes receive a sequence of data minibatches, compute a variational posterior for each, and make asynchronous streaming updates to a central model. In contrast to previous algorithms, the proposed framework is truly streaming, distributed, asynchronous, learning-rate-free, and truncation-free. The key challenge in developing the framework, arising from fact that BNP models do not impose an inherent ordering on their components, is finding the correspondence between minibatch and central BNP posterior components before performing each update. To address this, the paper develops a combinatorial optimization problem over component correspondences, and provides an efficient solution technique. The paper concludes with an application of the methodology to the DP mixture model, with experimental results demonstrating its practical scalability and performance.
Trevor Campbell, Julian Straub, John W. Fisher III, Jonathan P. How
NIPS3
2015 Probabilistic Variational Bounds for Graphical Models
abstract
Variational algorithms such as tree-reweighted belief propagation can provide deterministic bounds on the partition function, but are often loose and difficult to use in an any-time'' fashion, expending more computation for tighter bounds. On the other hand, Monte Carlo estimators such as importance sampling have excellent any-time behavior, but depend critically on the proposal distribution. We propose a simple Monte Carlo based inference method that augments convex variational bounds by adding importance sampling (IS). We argue that convex variational methods naturally provide good IS proposals thatcover the probability of the target distribution, and reinterpret the variational optimization as designing a proposal to minimizes an upper bound on the variance of our IS estimator. This both provides an accurate estimator and enables the construction of any-time probabilistic bounds that improve quickly and directly on state of-the-art variational bounds, which provide certificates of accuracy given enough samples relative to the error in the initial bound.
Qiang Liu 0001, John W. Fisher III, Alexander Ihler
NIPS2
2015 Estimating the Partition Function by Discriminance Sampling
Qiang Liu 0001, Jian Peng 0001, Alexander Ihler, John W. Fisher III
UAI4
2014 Bayesian Switching Interaction Analysis Under Uncertainty
abstract
We introduce a Bayesian discrete-time framework for switching-interaction analysis under uncertainty, in which latent interactions, switching pattern and signal states and dynamics are inferred from noisy (and possibly missing) observations of these signals. We propose reasoning over full posterior distribution of these latent variables as a means of combating and characterizing uncertainty. This approach also allows for answering a variety of questions probabilistically, which is suitable for exploratory pattern discovery and post-analysis by human experts. This framework is based on a fully-Bayesian learning of the structure of a switching dynamic Bayesian network (DBN) and utilizes a state-space approach to allow for noisy observations and missing data. It generalizes the autoregressive switching interaction model of Siracusa et al., which does not allow observation noise, and the switching linear dynamic system model of Fox et al., which does not infer interactions among signals. Posterior samples are obtained via a Gibbs sampling procedure, which is particularly efficient in the case of linear Gaussian dynamics and observation models. We demonstrate the utility of our framework on a controlled human-generated data, and a real-world climate data.
Zoran Dzunic, John W. Fisher III
AISTATS2
2014 Aerial Reconstructions via Probabilistic Data Fusion
abstract
We propose an integrated probabilistic model for multi-modal fusion of aerial imagery, LiDAR data, and (optional) GPS measurements. The model allows for analysis and dense reconstruction (in terms of both geometry and appearance) of large 3D scenes. An advantage of the approach is that it explicitly models uncertainty and allows for missing data. As compared with image-based methods, dense reconstructions of complex urban scenes are feasible with fewer observations. Moreover, the proposed model allows one to estimate absolute scale and orientation and reason about other aspects of the scene, e.g., detection of moving objects. As formulated, the model lends itself to massively-parallel computing. We exploit this in an efficient inference scheme that utilizes both general purpose and domain-specific hardware components. We demonstrate results on large-scale reconstruction of urban terrain from LiDAR and aerial photography data.
Randi Cabezas, Oren Freifeld, Guy Rosman, John W. Fisher III
CVPR4
2014 A Mixture of Manhattan Frames: Beyond the Manhattan World
abstract
Objects and structures within man-made environments typically exhibit a high degree of organization in the form of orthogonal and parallel planes. Traditional approaches to scene representation exploit this phenomenon via the somewhat restrictive assumption that every plane is perpendicular to one of the axes of a single coordinate system. Known as the Manhattan-World model, this assumption is widely used in computer vision and robotics. The complexity of many real-world scenes, however, necessitates a more flexible model. We propose a novel probabilistic model that describes the world as a mixture of Manhattan frames: each frame defines a different orthogonal coordinate system. This results in a more expressive model that still exploits the orthogonality constraints. We propose an adaptive Markov-Chain Monte-Carlo sampling algorithm with Metropolis-Hastings split/merge moves that utilizes the geometry of the unit sphere. We demonstrate the versatility of our Mixture-of-Manhattan-Frames model by describing complex scenes using depth images of indoor scenes as well as aerial-LiDAR measurements of an urban center. Additionally, we show that the model lends itself to focal-length calibration of depth cameras and to plane segmentation.
Julian Straub, Guy Rosman, Oren Freifeld, John J. Leonard, John W. Fisher III
CVPR5
2014 Bayesian Nonparametric Intrinsic Image Decomposition
Jason Chang 0001, Randi Cabezas, John W. Fisher III
ECCV (4)3
2014 Bayesian nonparametric modeling of driver behavior
abstract
Modern vehicles are equipped with increasingly complex sensors. These sensors generate large volumes of data that provide opportunities for modeling and analysis. Here, we are interested in exploiting this data to learn aspects of behaviors and the road network associated with individual drivers. Our dataset is collected on a standard vehicle used to commute to work and for personal trips. A Hidden Markov Model (HMM) trained on the GPS position and orientation data is utilized to compress the large amount of position information into a small amount of road segment states. Each state has a set of observations, i.e. car signals, associated with it that are quantized and modeled as draws from a Hierarchical Dirichlet Process (HDP). The inference for the topic distributions is carried out using an online variational inference algorithm. The topic distributions over joint quantized car signals characterize the driving situation in the respective road state. In a novel manner, we demonstrate how the sparsity of the personal road network of a driver in conjunction with a hierarchical topic model allows data driven predictions about destinations as well as likely road conditions.
Julian Straub, Sue Zheng, John W. Fisher III
Intelligent Vehicles Symposium3
2014 Parallel Sampling of HDPs using Sub-Cluster Splits
Jason Chang 0001, John W. Fisher III
NIPS2
2014 Coresets for k-Segmentation of Streaming Data
Guy Rosman, Mikhail Volkov 0002, Dan Feldman, John W. Fisher III, Daniela Rus
NIPS4
2013 A Video Representation Using Temporal Superpixels
abstract
We develop a generative probabilistic model for temporally consistent super pixels in video sequences. In contrast to supermodel methods, object parts in different frames are tracked by the same temporal super pixel. We explicitly model flow between frames with a bilateral Gaussian process and use this information to propagate super pixels in an online fashion. We consider four novel metrics to quantify performance of a temporal super pixel representation and demonstrate superior performance when compared to supermodel methods.
Jason Chang 0001, Donglai Wei 0001, John W. Fisher III
CVPR3
2013 Topology-Constrained Layered Tracking with Latent Flow
abstract
We present an integrated probabilistic model for layered object tracking that combines dynamics on implicit shape representations, topological shape constraints, adaptive appearance models, and layered flow. The generative model combines the evolution of appearances and layer shapes with a Gaussian process flow and explicit layer ordering. Efficient MCMC sampling algorithms are developed to enable a particle filtering approach while reasoning about the distribution of object boundaries in video. We demonstrate the utility of the proposed tracking algorithm on a wide variety of video sources while achieving state-of-the-art results on a boundary-accurate tracking dataset.
Jason Chang 0001, John W. Fisher III
ICCV2
2013 Parallel Sampling of DP Mixture Models using Sub-Cluster Splits
abstract
We present a novel MCMC sampler for Dirichlet process mixture models that can be used for conjugate or non-conjugate prior distributions. The proposed sampler can be massively parallelized to achieve significant computational gains. A non-ergodic restricted Gibbs iteration is mixed with split/merge proposals to produce a valid sampler. Each regular cluster is augmented with two sub-clusters to construct likely split moves. Unlike many previous parallel samplers, the proposed sampler accurately enforces the correct stationary distribution of the Markov chain without the need for approximate models. Empirical results illustrate that the new sampler exhibits better convergence properties than current methods.
Jason Chang 0001, John W. Fisher III
NIPS2
2012 Manifold guided composite of Markov random fields for image modeling
abstract
We present a new generative image model, integrating techniques arising from two different domains: manifold modeling and Markov random fields. First, we develop a probabilistic model with a mixture of hyperplanes to approximate the manifold of orientable image patches, and demonstrate that it is more effective than the field of experts in expressing local texture patterns. Next, we develop a construction that yields an MRF for coherent image generation, given a configuration of local patch models, and thereby establish a prior distribution over an MRF space. Taking advantage of the model structure, we derive a variational inference algorithm, and apply it to low-level vision. In contrast to previous methods that rely on a single MRF, the method infers an approximate posterior distribution of MRFs, and recovers the underlying images by combining the predictions in a Bayesian fashion. Experiments quantitatively demonstrate superior performance as compared to state-of-the-art methods on image denoising and inpainting.
Dahua Lin, John W. Fisher III
CVPR2
2012 Low level vision via switchable Markov random fields
abstract
Markov random fields play a central role in solving a variety of low level vision problems, including denoising, in-painting, segmentation, and motion estimation. Much previous work was based on MRFs with hand-crafted networks, yet the underlying graphical structure is rarely explored. In this paper, we show that if appropriately estimated, the MRF's graphical structure, which captures significant information about appearance and motion, can provide crucial guidance to low level vision tasks. Motivated by this observation, we propose a principled framework to solve low level vision tasks via an exponential family of MRFs with variable structures, which we call Switchable MRFs. The approach explicitly seeks a structure that optimally adapts to the image or video along the pursuit of task-specific goals. Through theoretical analysis and experimental study, we demonstrate that the proposed method addresses a number of drawbacks suffered by previous methods, including failure to capture heavy-tail statistics, computational difficulties, and lack of generality.
Dahua Lin, John W. Fisher III
CVPR2
2012 Learning Deformations with Parallel Transport
Donglai Wei 0001, Dahua Lin, John W. Fisher III
ECCV (2)3
2012 Efficient topology-controlled sampling of implicit shapes
abstract
Sampling from distributions of implicitly defined shapes enables analysis of various energy functionals used for image segmentation. Recent work [1] describes a computationally efficient Metropolis- Hastings method for accomplishing this task. Here, we extend that framework so that samples are accepted at every iteration of the sampler, achieving an order of magnitude speed up in convergence. Additionally, we show how to incorporate topological constraints.
Jason Chang 0001, John W. Fisher III
ICIP2
2012 Coupling Nonparametric Mixtures via Latent Dirichlet Processes
abstract
Mixture distributions are often used to model complex data. In this paper, we develop a new method that jointly estimates mixture models over multiple data sets by exploiting the statistical dependencies between them. Specifically, we introduce a set of latent Dirichlet processes as sources of component models (atoms), and for each data set, we construct a nonparametric mixture model by combining sub-sampled versions of the latent DPs. Each mixture model may acquire atoms from different latent DPs, while each atom may be shared by multiple mixtures. This multi-to-multi association distinguishes the proposed method from prior constructions that rely on tree or chain structures, allowing mixture models to be coupled more flexibly. In addition, we derive a sampling algorithm that jointly infers the model parameters and present experiments on both document analysis and image modeling.
Dahua Lin, John W. Fisher III
NIPS2
2011 Efficient MCMC sampling with implicit shape representations
abstract
We present a method for sampling from the posterior distribution of implicitly defined segmentations conditioned on the observed image. Segmentation is often formulated as an energy minimization or statistical inference problem in which either the optimal or most probable configuration is the goal. Exponentiating the negative energy functional provides a Bayesian interpretation in which the solutions are equivalent. Sampling methods enable evaluation of distribution properties that characterize the solution space via the computation of marginal event probabilities. We develop a Metropolis-Hastings sampling algorithm over level-sets which improves upon previous methods by allowing for topological changes while simultaneously decreasing computational times by orders of magnitude. An M-ary extension to the method is provided.
Jason Chang 0001, John W. Fisher III
CVPR2
2010 Modeling and estimating persistent motion with geometric flows
abstract
We propose a principled framework to model persistent motion in dynamic scenes. In contrast to previous efforts on object tracking and optical flow estimation that focus on local motion, we primarily aim at inferring a global model of persistent and collective dynamics. With this in mind, we first introduce the concept of geometric flow that describes motion simultaneously over space and time, and derive a vector space representation based on Lie algebra. We then extend it to model complex motion by combining multiple flows in a geometrically consistent manner. Taking advantage of the linear nature of this representation, we formulate a stochastic flow model, and incorporate a Gaussian process to capture the spatial coherence more effectively. This model leads to an efficient and robust algorithm that can integrate both point pairs and frame differences in motion estimation. We conducted experiments on different types of videos. The results clearly demonstrate that the proposed approach is effective in modeling persistent motion.
Dahua Lin, W. Eric L. Grimson, John W. Fisher III
CVPR3
2010 Construction of Dependent Dirichlet Processes based on Poisson Processes
abstract
We present a novel method for constructing dependent Dirichlet processes. The approach exploits the intrinsic relationship between Dirichlet and Poisson pro- cesses in order to create a Markov chain of Dirichlet processes suitable for use as a prior over evolving mixture models. The method allows for the creation, re- moval, and location variation of component models over time while maintaining the property that the random measures are marginally DP distributed. Addition- ally, we derive a Gibbs sampling algorithm for model inference and test it on both synthetic and real data. Empirical results demonstrate that the approach is effec- tive in estimating dynamically varying mixture models.
Dahua Lin, W. Eric L. Grimson, John W. Fisher III
NIPS3
2009 Learning visual flows: A Lie algebraic approach
abstract
We present a novel method for modeling dynamic visual phenomena, which consists of two key aspects. First, the integral motion of constituent elements in a dynamic scene is captured by a common underlying geometric transform process. Second, a Lie algebraic representation of the transform process is introduced, which maps the transformation group to a vector space, and thus overcomes the difficulties due to the group structure. Consequently, the statistical learning techniques based on vector spaces can be readily applied. Moreover, we discuss the intrinsic connections between the Lie algebra and the Linear dynamical processes, showing that our model induces spatially varying fields that can be estimated from local motions without continuous tracking. Following this, we further develop a statistical framework to robustly learn the flow models from noisy and partially corrupted observations. The proposed methodology is demonstrated on real world phenomenon, inferring common motion patterns from surveillance videos of crowded scenes and satellite data of weather evolution.
Dahua Lin, W. Eric L. Grimson, John W. Fisher III
CVPR3
2009 Automatic registration of LIDAR and optical images of urban scenes
abstract
Fusion of 3D laser radar (LIDAR) imagery and aerial optical imagery is an efficient method for constructing 3D virtual reality models. One difficult aspect of creating such models is registering the optical image with the LIDAR point cloud, which is characterized as a camera pose estimation problem. We propose a novel application of mutual information registration methods, which exploits the statistical dependency in urban scenes of optical appearance with measured LIDAR elevation. We utilize the well known downhill simplex optimization to infer camera pose parameters. We discuss three methods for measuring mutual information between LIDAR imagery and optical imagery. Utilization of OpenGL and graphics hardware in the optimization process yields registration times dramatically lower than previous methods. Using an initial registration comparable to GPS/INS accuracy, we demonstrate the utility of our algorithm with a collection of urban images and present 3D models created with the fused imagery.
Andrew Mastin, Jeremy Kepner, John W. Fisher III
CVPR3
2009 Analysis of orientation and scale in smoothly varying textures
abstract
We present a novel representation for modeling textured regions subject to smooth variations in orientation and scale. Utilizing the steerable pyramid of Simoncelli and Freeman as a basis, we decompose textured regions of natural images into explicit local attributes of contrast, bias, scale, and orientation. Additionally, we impose smoothness on these attributes via Markov random fields. The combination allows for demonstrable improvements in common scene analysis applications including unsupervised segmentation, reflectance and shading estimation, and estimation of the radiometric response function from a single image.
Jason Chang 0001, John W. Fisher III
ICCV2
2008 Learning max-weight discriminative forests
abstract
We present a method for sequential learning of increasingly complex graphical models for discriminating between two hypotheses. We generate forests for each hypothesis, each with no more edges than a spanning tree, which optimize an information-theoretic criteria. The method relies on a straightforward extension of the efficient max-weight spanning tree (MWST) algorithm by incorporating multivalued edge-weights. Each iteration produces nested forests with increasing number of edges; each provably optimal as compared to alternative forests. Empirical results demonstrate superior probability of error as compared to generative approaches.
Vincent Y. F. Tan, John W. Fisher III, Alan S. Willsky
ICASSP2
2007 Dynamic Dependency Tests for Audio-Visual Speaker Association
abstract
We formulate the problem of audio-visual speaker association as a dynamic dependency test. That is, given an audio stream and multiple video streams, we wish to determine their dependency structure as it evolves over time. To this end, we propose the use of a hidden factorization Markov model in which the hidden state encodes a finite number of possible dependency structures. Each dependency structure has an explicit semantic meaning, namely "who is speaking". This model takes advantage of both structural and parametric changes associated with changes in speaker. This is contrasted with standard sliding window based dependence analysis. Using this model we obtain state-of-the-art performance on an audio-visual association task without benefit of training data.
Michael Siracusa, John W. Fisher III
ICASSP (2)2
2007 Performance Guarantees for Information Theoretic Sensor Resource Management
abstract
Many estimation problems involve sensors which can be actively controlled to alter the information received and utilized in the underlying inference task. In this paper, we discuss performance guarantees for heuristic algorithms for adaptive sensor control in sequential estimation problems, where the inference criterion is mutual information. We also demonstrate the performance of our tighter online computable performance guarantees through computational simulations. The guarantees may be applied to other estimation criteria including the Cramer-Rao bound.
Jason Williams 0002, John W. Fisher III, Alan S. Willsky
ICASSP (3)2
2007 Detecting Cortical Surface Regions in Structural MR Data
abstract
We present a novel level-set method for evolving open surfaces embedded in three-dimensional volumes. We adapt the method for statistical detection and segmentation of cytoarchitectonic regions of the cortical ribbon (e.g., Brodmann areas). In addition, we incorporate an explicit interface appearance model which is oriented normal to the open surface, allowing one to model characteristics beyond voxel intensities and high gradients. We show that such models are well suited to detecting embedded cortical structures. Appearance models of the interface are used in two ways: firstly, to evolve an open surface in the normal direction for the purpose of detecting the location of the surface, and secondly, to evolve the boundary of the surface in a direction tangential to the surface in order to delineate the extent of a specific Brodmann area within the cortical ribbon. The utility of the method is demonstrated on a challenging ex-vivo structural MR dataset for detection of Brodmann area 17.
Biswajit Bose, John W. Fisher III, Bruce Fischl, Oliver Hinds, W. Eric L. Grimson
ICCV2
2007 MCMC Curve Sampling for Image Segmentation
Ayres C. Fan, John W. Fisher III, William M. Wells III, James J. Levitt, Alan S. Willsky
MICCAI (2)2
2007 Combining object and feature dynamics in probabilistic tracking
Leonid Taycher, John W. Fisher III, Trevor Darrell
Comput. Vis. Image Underst.2
2007 Using the logarithm of odds to define a vector space on probabilistic atlases
Kilian M. Pohl, John W. Fisher III, Sylvain Bouix, Martha Elizabeth Shenton, Robert W. McCarley, W. Eric L. Grimson, Ron Kikinis, William M. Wells III
Medical Image Anal.2
2006 Detection and Localization of Material Releases with Sparse Sensor Configurations
abstract
We consider the problem of detecting and localizing a material release utilizing sparse sensor measurements. We formulate the problem as one of abrupt change detection. The problem is challenging because of the sparse sensor deployment and complex system dynamics. We restrict ourselves to propagation models consisting of diffusion plus transport according to a Gaussian puff model. We derive optimal inference algorithms, provided the model parametrization is known precisely, within a hybrid detection-localization hypothesis testing framework with linear growth in the hypothesis space. The primary assumptions are that the mean wind field is deterministically known and that the Gaussian puff model is valid. Under these assumptions, we characterize the change in performance of detection, time-to-detection and localization as a function of the number of sensors. We then examine some performance impacts when the underlying dynamical model deviates from the assumed model
Emily B. Fox, Jason Williams 0002, John W. Fisher III, Alan S. Willsky
ICASSP (4)3
2006 Logarithm Odds Maps for Shape Representation
Kilian M. Pohl, John W. Fisher III, Martha Elizabeth Shenton, Robert W. McCarley, W. Eric L. Grimson, Ron Kikinis, William M. Wells III
MICCAI (2)2
2005 Combining Object and Feature Dynamics in Probabilistic Tracking
abstract
Objects can exhibit different dynamics at different scales, and this is often exploited by visual tracking algorithms. A local dynamic model is typically used to extract image features that are then used as input to a system for tracking the entire object using a global dynamic model. Approximate local dynamics may be brittle - point trackers drift due to image noise and adaptive background models adapt to foreground objects that become stationary - but constraints from the global model can make them more robust. We propose a probabilistic framework for incorporating global dynamics knowledge into the local feature extraction processes. A global tracking algorithm can be formulated as a generative model and used to predict feature values that are incorporated into an observation process of the feature extractor. We combine such models in a multichain graphical model framework. We show the utility of our framework for improving feature tracking and thus shape and motion estimates in a batch factorization algorithm. We also propose an approximate filtering algorithm appropriate for online applications, and demonstrate its application to background subtraction.
Leonid Taycher, John W. Fisher III, Trevor Darrell
CVPR (2)2
2005 Estimating dependency and significance for high-dimensional data
abstract
Understanding the dependency structure of a set of variables is a key component in various signal processing applications which involve data association. The simple task of detecting whether any dependency exists is particularly difficult when models of the data are unknown or difficult to characterize because of high-dimensional measurements. We review the use of nonparametric tests for characterizing dependency and how to carry out these tests with high-dimensional observations. In addition we present a method to assess the significance of the tests.
Michael Siracusa, Kinh Tieu, Alexander Ihler, John W. Fisher III, Alan S. Willsky
ICASSP (5)4
2005 Optimization approaches to dynamic routing of measurements and models in a sensor network object tracking problem
abstract
Inter-sensor communication often comprises a significant portion of energy expenditures in a sensor network as compared to sensing and computation. We discuss an integrated approach to dynamically routing measurements and models in a sensor network. Specifically, we examine the problem of tracking objects within a region wherein the responsibility for combining measurements and updating a posterior state distribution is assigned to a single sensor at any given time step. The so called leader node may change over time. Sensor nodes communicate for two reasons: firstly, measurements of target state are transmitted from sensors to the current leader node for incorporation into the state estimate model; secondly, the state model is transmitted between sensors when the leader node changes. The trade-off between these two types of communication is of primary importance to dynamic selection of the leader node. We propose an algorithm based on a dynamic programming roll-out formulation of the minimum cost problem. We obtain a cost function which can be efficiently minimized by simplifying the problem to that of an open loop feedback controller which is an upper bound to the optimal cost. We present empirical results which compare methods previously proposed in the literature to the algorithm presented here.
Jason Williams 0002, John W. Fisher III, Alan S. Willsky
ICASSP (5)2
2005 A Unifying Approach to Registration, Segmentation, and Intensity Correction
Kilian M. Pohl, John W. Fisher III, James J. Levitt, Martha Elizabeth Shenton, Ron Kikinis, W. Eric L. Grimson, William M. Wells III
MICCAI2
2005 Loopy Belief Propagation: Convergence and Effects of Message Errors
abstract
Belief propagation (BP) is an increasingly popular method of performing approximate inference on arbitrary graphical models. At times, even further approximations are required, whether due to quantization of the messages or model parameters, from other simplified message or model representations, or from stochastic approximation methods. The introduction of such errors into the BP message computations has the potential to affect the solution obtained adversely. We analyze the effect resulting from message approximation under two particular measures of error, and show bounds on the accumulation of errors in the system. This analysis leads to convergence conditions for traditional BP message passing, and both strict bounds and estimates of the resulting error in systems of approximate BP message passing.
Alexander Ihler, John W. Fisher III, Alan S. Willsky
J. Mach. Learn. Res.2
2005 Nonparametric belief propagation for self-localization of sensor networks
abstract
Automatic self-localization is a critical need for the effective use of ad hoc sensor networks in military or civilian applications. In general, self-localization involves the combination of absolute location information (e.g., from a global positioning system) with relative calibration information (e.g., distance measurements between sensors) over regions of the network. Furthermore, it is generally desirable to distribute the computational burden across the network and minimize the amount of intersensor communication. We demonstrate that the information used for sensor localization is fundamentally local with regard to the network topology and use this observation to reformulate the problem within a graphical model framework. We then present and demonstrate the utility of nonparametric belief propagation (NBP), a recent generalization of particle filtering, for both estimating sensor locations and representing location uncertainties. NBP has the advantage that it is easily implemented in a distributed fashion, admits a wide variety of statistical models, and can represent multimodal uncertainty. Using simulations of small to moderately sized sensor networks, we show that NBP may be made robust to outlier measurement errors by a simple model augmentation, and that judicious message construction can result in better estimates. Furthermore, we provide an analysis of NBP's communications requirements, showing that typically only a few messages per sensor are required, and that even low bit-rate approximations of these messages can be used with little or no performance impact.
Alexander Ihler, John W. Fisher III, Randolph L. Moses, Alan S. Willsky
IEEE J. Sel. Areas Commun.2
2005 A Nonparametric Statistical Method for Image Segmentation Using Information Theory and Curve Evolution
abstract
In this paper, we present a new information-theoretic approach to image segmentation. We cast the segmentation problem as the maximization of the mutual information between the region labels and the image pixel intensities, subject to a constraint on the total length of the region boundaries. We assume that the probability densities associated with the image pixel intensities within each region are completely unknown a priori, and we formulate the problem based on nonparametric density estimates. Due to the nonparametric structure, our method does not require the image regions to have a particular type of probability distribution and does not require the extraction and use of a particular statistic. We solve the information-theoretic optimization problem by deriving the associated gradient flows and applying curve evolution techniques. We use level-set methods to implement the resulting evolution. The experimental results based on both synthetic and real images demonstrate that the proposed technique can solve a variety of challenging image segmentation problems. Futhermore, our method, which does not require any training, performs as good as methods based on training.
Junmo Kim 0004, John W. Fisher III, Anthony J. Yezzi, Müjdat Çetin, Alan S. Willsky
IEEE Trans. Image Process.2
2004 Nonparametric belief propagation for sensor self-calibration
abstract
Automatic self-calibration of ad-hoc sensor networks is a critical need for their use in military or civilian applications. In general, self-calibration involves the combination of absolute location information (e.g. GPS) with relative calibration information (e.g. estimated distance between sensors) over regions of the network. We formulate the self-calibration problem as a graphical model, enabling the application of nonparametric belief propagation (NBP), a recent generalization of particle filtering, for both estimating sensor locations and representing location uncertainties. NBP has the advantage that it is easily implemented in a distributed fashion, can represent multi-modal uncertainty, and admits a wide variety of statistical models. This last point is particularly appealing in that it can be used to provide robustness against occasional high-variance (outlier) noise. We illustrate the performance of NBP using Monte Carlo analysis on an example network.
Alexander Ihler, John W. Fisher III, Randolph L. Moses, Alan S. Willsky
ICASSP (3)2
2004 Nonparametric belief propagation for self-calibration in sensor networks
abstract
Automatic self-calibration of ad-hoc sensor networks is a critical need for their use in military or civilian applications. In general, self-calibration involves the combination of absolute location information (e.g. GPS) with relative calibration information (e.g. time delay or received signal strength between sensors) over regions of the network. Furthermore, it is generally desirable to distribute the computational burden across the network and minimize the amount of inter-sensor communication. We demonstrate that the information used for sensor calibration is fundamentally local with regard to the network topology and use this observation to reformulate the problem within a graphical model framework. We then demonstrate the utility of nonparametric belief propagation (NBP), a recent generalization of particle filtering, for both estimating sensor locations and representing location uncertainties. NBP has the advantage that it is easily implemented in a distributed fashion, admits a wide variety of statistical models, and can represent multi-modal uncertainty. We illustrate the performance of NBP on several example networks while comparing to a previously published nonlinear least squares method.
Alexander Ihler, John W. Fisher III, Randolph L. Moses, Alan S. Willsky
IPSN2
2004 Exact MAP Activity Detection in f MRI Using a GLM with an Ising Spatial Prior
Eric R. Cosman Jr., John W. Fisher III, William M. Wells III
MICCAI (2)2
2004 Message Errors in Belief Propagation
abstract
Belief propagation (BP) is an increasingly popular method of perform- ing approximate inference on arbitrary graphical models. At times, even further approximations are required, whether from quantization or other simplified message representations or from stochastic approxima- tion methods. Introducing such errors into the BP message computations has the potential to adversely affect the solution obtained. We analyze this effect with respect to a particular measure of message error, and show bounds on the accumulation of errors in the system. This leads both to convergence conditions and error bounds in traditional and approximate BP message passing.
Alexander Ihler, John W. Fisher III, Alan S. Willsky
NIPS2
2004 Speaker association with signal-level audiovisual fusion
abstract
Audio and visual signals arriving from a common source are detected using a signal-level fusion technique. A probabilistic multimodal generation model is introduced and used to derive an information theoretic measure of cross-modal correspondence. Nonparametric statistical density modeling techniques can characterize the mutual information between signals from different domains. By comparing the mutual information between different pairs of signals, it is possible to identify which person is speaking a given utterance and discount errant motion or audio from other utterances or nonspeech events.
John W. Fisher III, Trevor Darrell
IEEE Trans. Multim.1
2003 Incorporating complex statistical information in active contour-based image segmentation
abstract
An information-theoretic method for multiphase image segmentation, in an active contour-based framework is proposed. Our approach is based on nonparametric density estimates, and is able to solve problems involving arbitrary probability densities for the region intensities. This is achieved by maximizing the mutual information between the region labels and the image pixel intensities, in order to segment up to 2/sup m/ regions using m curves. The method does not require any prior training regarding the regions of interest, but rather learns the probability densities during the evolution process. We present some illustrative experimental results, demonstrating the power of the proposed segmentation approach.
Junmo Kim 0004, John W. Fisher III, Müjdat Çetin, Anthony J. Yezzi, Alan S. Willsky
ICIP (2)2
2003 Learning cross-modal appearance models with application to tracking
abstract
Objects of interest are rarely silent or invisible. Analysis of multi-modal signal generation from a single object represents a rich and challenging area for smart sensor arrays. We consider the problem of simultaneously learning and audio and visual appearance model of a moving subject. We present a method which successfully learns such a model without benefit of hand initialization using only the associated audio signal to "decide" which object to model and track. We are interested in particular in modeling joint audio and video variation, such as produced by a speaking face. We present an algorithm and experimental results of a human speaker moving in a scene.
John W. Fisher III, Trevor Darrell
ICME1
2003 A multi-modal approach for determining speaker location and focus
abstract
This paper presents a multi-modal approach to locate a speaker in a scene and determine to whom he or she is speaking. We present a simple probabilistic framework that combines multiple cues derived from both audio and video information. A purely visual cue is obtained using a head tracker to identify possible speakers in a scene and provide both their 3-D positions and orientation. In addition, estimates of the audio signal's direction of arrival are obtained with the help of a two-element microphone array. A third cue measures the association between the audio and the tracked regions in the video. Integrating these cues provides a more robust solution than using any single cue alone. The usefulness of our approach is shown in our results for video sequences with two or more people in a prototype interactive kiosk environment.
Michael Siracusa, Louis-Philippe Morency, Kevin W. Wilson, John W. Fisher III, Trevor Darrell
ICMI4
2003 ICA Using Spacings Estimates of Entropy
Erik G. Learned-Miller, John W. Fisher III
J. Mach. Learn. Res.2
2002 Probabalistic Models and Informative Subspaces for Audiovisual Correspondence
John W. Fisher III, Trevor Darrell
ECCV (3)1
2002 Face Recognition from Long-Term Observations
Gregory Shakhnarovich, John W. Fisher III, Trevor Darrell
ECCV (3)2
2002 Informative subspaces for audio-visual processing: High-level function from low-level fusion
abstract
We propose a new probabilistic model of single source multi-modal generation, and show algorithms for maximizing mutual information which find correspondences between signal components. We show a nonparametric method for finding informative subspaces that captures complex statistical relationships between different modalities. We extend a previous subspace method to include new priors on the projection weights, yielding more robust results. Applied to human speakers, our model finds a relationship between audio speech and video of facial motion, and partially segments background events in both channels. We present new results on the problem of audio-visual verification, and show how the audio and video of a speaker can be matched without a prior model of the speaker's voice or appearance.
John W. Fisher III, Trevor Darrell
ICASSP1
2002 Nonparametric methods for image segmentation using information theory and curve evolution
abstract
We present a novel information theoretic approach to image segmentation. We cast the segmentation problem as the maximization of the mutual information between the region labels and the image pixel intensities, subject to a constraint on the total length of the region boundaries. We assume that the probability densities associated with the image pixel intensities within each region are completely unknown a priori, and we formulate the problem based on nonparametric density estimates. Due to the nonparametric structure, our method does not require the image regions to have a particular type of probability distribution, and does not require the extraction and use of a particular statistic. We solve the information-theoretic optimization problem by deriving the associated gradient flows and applying curve evolution techniques. We use fast level set methods to implement the resulting evolution The evolution equations are based on nonparametric statistics, and have an intuitive appeal. The experimental results based on both synthetic and real images demonstrate that the proposed technique can solve a variety of challenging image segmentation problems.
Junmo Kim 0004, John W. Fisher III, Anthony J. Yezzi, Müjdat Çetin, Alan S. Willsky
ICIP (3)2
2002 Recovering Articulated Model Topology from Observed Rigid Motion
abstract
Accurate representation of articulated motion is a challenging problem for machine perception. Several successful tracking algorithms have been developed that model human body as an articulated tree. We pro- pose a learning-based method for creating such articulated models from observations of multiple rigid motions. This paper is concerned with recovering topology of the articulated model, when the rigid motion of constituent segments is known. Our approach is based on finding the Maximum Likelihood tree shaped factorization of the joint probability density function (PDF) of rigid segment motions. The topology of graph- ical model formed from this factorization corresponds to topology of the underlying articulated body. We demonstrate the performance of our al- gorithm on both synthetic and real motion capture data.
Leonid Taycher, John W. Fisher III, Trevor Darrell
NIPS2
2001 Nonparametric estimators for online signature authentication
abstract
We present extensions to our previous work in modelling dynamical processes. The approach uses an information theoretic criterion for searching over subspaces of the past observations, combined with a nonparametric density characterizing its relation to one-step-ahead prediction and uncertainty. We use this methodology to model handwriting stroke data, specifically signatures, as a dynamical system and show that it is possible to learn a model capturing their dynamics for use either in synthesizing realistic signatures and in discriminating between signatures and forgeries even though no forgeries have been used in constructing the model. This novel approach yields promising results even for small training sets.
Alexander Ihler, John W. Fisher III, Alan S. Willsky
ICASSP2
2001 Adaptive Entropy Rates for f MRI Time-Series Analysis
John W. Fisher III, Eric R. Cosman Jr., Cindy Wible, William M. Wells III
MICCAI1
2000 Ausio-visual Segmentation and "The Cocktail Party Effect"
Trevor Darrell, John W. Fisher III, Paul A. Viola, William T. Freeman
ICMI2
2000 Incorporating Spatial Priors into an Information Theoretic Approach for fMRI Data Analysis
Junmo Kim 0004, John W. Fisher III, Andy Tsai, Cindy Wible, Alan S. Willsky, William M. Wells III
MICCAI2
2000 Learning Joint Statistical Models for Audio-Visual Fusion and Segregation
abstract
People can understand complex auditory and visual information, often using one to disambiguate the other. Automated analysis, even at a low(cid:173) level, faces severe challenges, including the lack of accurate statistical models for the signals, and their high-dimensionality and varied sam(cid:173) pling rates. Previous approaches [6] assumed simple parametric models for the joint distribution which, while tractable, cannot capture the com(cid:173) plex signal relationships. We learn the joint distribution of the visual and auditory signals using a non-parametric approach. First, we project the data into a maximally informative, low-dimensional subspace, suitable for density estimation. We then model the complicated stochastic rela(cid:173) tionships between the signals using a nonparametric density estimator. These learned densities allow processing across signal modalities. We demonstrate, on synthetic and real signals, localization in video of the face that is speaking in audio, and, conversely, audio enhancement of a particular speaker selected from the video.
John W. Fisher III, Trevor Darrell, William T. Freeman, Paul A. Viola
NIPS1
1999 Analysis of Functional MRI Data Using Mutual Information
Andy Tsai, John W. Fisher III, Cindy Wible, William M. Wells III, Junmo Kim 0004, Alan S. Willsky
MICCAI2
1999 Learning Informative Statistics: A Nonparametnic Approach
John W. Fisher III, Alexander Ihler, Paul A. Viola
NIPS1
1998 A novel measure for independent component analysis (ICA)
abstract
Measures of independence (and dependence) are fundamental in many areas of engineering and signal processing. Shannon introduced the idea of information entropy which has a sound theoretical foundation but sometimes is not easy to implement in engineering applications. In this paper, Renyi's entropy is used and a novel independence measure is proposed. When integrated with a nonparametric estimator of the probability density function (Parzen Window), the measure can be related to the "potential energy of the samples" which is easy to understand and implement. The experimental results on blind source separation confirm the theory. Although the work is preliminary, the "potential energy" method is rather general and will have many applications.
Dongxin Xu, José C. Príncipe, John W. Fisher III, Hsiao-Chun Wu
ICASSP3
1998 Target discrimination in synthetic aperture radar using artificial neural networks
abstract
This paper addresses target discrimination in synthetic aperture radar (SAR) imagery using linear and nonlinear adaptive networks. Neural networks are extensively used for pattern classification but here the goal is discrimination. We show that the two applications require different cost functions. We start by analyzing with a pattern recognition perspective the two-parameter constant false alarm rate (CFAR) detector which is widely utilized as a target detector in SAR. Then we generalize its principle to construct the quadratic gamma discriminator (QGD), a nonparametrically trained classifier based on local image intensity. The linear processing element of the QCD is further extended with nonlinearities yielding a multilayer perceptron (MLP) which we call the NL-QGD (nonlinear QGD). MLPs are normally trained based on the L(2) norm. We experimentally show that the L(2) norm is not recommended to train MLPs for discriminating targets in SAR. Inspired by the Neyman-Pearson criterion, we create a cost function based on a mixed norm to weight the false alarms and the missed detections differently. Mixed norms can easily be incorporated into the backpropagation algorithm, and lead to better performance. Several other norms (L(8), cross-entropy) are applied to train the NL-QGD and all outperformed the L(2) norm when validated by receiver operating characteristics (ROC) curves. The data sets are constructed from TABILS 24 ISAR targets embedded in 7 km(2) of SAR imagery (MIT/LL mission 90).
José C. Príncipe, Munchurl Kim, John W. Fisher III
IEEE Trans. Image Process.3
1995 A nonlinear extension of the MACE filter
John W. Fisher III, José C. Príncipe
Neural Networks1