Matthew Hoffman 0001

dblp:07/4433 · also Matthew D. Hoffman, Matthew Douglas Hoffman · DBLP profile ↗
← Back
43ranked-venue papers
13as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 12 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
24 papers
Probabilistic and Bayesian machine learning · 60% Generative modeling · 11% 3D vision · 7%
Databases, data mining, and information retrieval
3 papers
Recommender systems · 70% Data mining · 30%

Topics — the 30 heaviest of 60, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
2.192020
Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics · ICML 2020
Automatic Reparameterisation of Probabilistic Programs · ICML 2020
A Variational Analysis of Stochastic Gradient Algorithms · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming
1.442020
Automatic Reparameterisation of Probabilistic Programs · ICML 2020
Simple, Distributed, and Accelerated Probabilistic Programming · NeurIPS 2018
Autoconj: Recognizing and Exploiting Conjugacy Without a Domain-Specific Language · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
1.352020
Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics · ICML 2020
Generalizing Hamiltonian Monte Carlo with Neural Networks · ICLR (Poster) 2018
Learning Deep Latent Gaussian Models with Markov Chain Monte Carlo · ICML 2017
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo
1.032021
What Are Bayesian Neural Network Posteriors Really Like? · ICML 2021
Generalizing Hamiltonian Monte Carlo with Neural Networks · ICLR (Poster) 2018
The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo · J. Mach. Learn. Res. 2014
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.822021
What Are Bayesian Neural Network Posteriors Really Like? · ICML 2021
Stochastic Gradient Descent as Approximate Bayesian Inference · J. Mach. Learn. Res. 2017
Machine learning › Probabilistic and Bayesian machine learning
statistical inference
0.822020
Automatic Reparameterisation of Probabilistic Programs · ICML 2020
Autoconj: Recognizing and Exploiting Conjugacy Without a Domain-Specific Language · NeurIPS 2018
Computer vision › 3D vision
3d scene reconstruction
0.812024
Robust Inverse Graphics via Probabilistic Inference · ICML 2024
Machine learning › Generative modeling
diffusion model
0.812024
Robust Inverse Graphics via Probabilistic Inference · ICML 2024
Machine learning › Generative modeling › diffusion model
diffusion model conditioning
0.812024
Robust Inverse Graphics via Probabilistic Inference · ICML 2024
Computer vision › 3D vision
inverse rendering
0.812024
Robust Inverse Graphics via Probabilistic Inference · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.812024
Robust Inverse Graphics via Probabilistic Inference · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.712023
Sequential Monte Carlo Learning for Time Series Structure Discovery · ICML 2023
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting
0.712023
Training Chain-of-Thought via Latent-Variable Inference · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.712023
Training Chain-of-Thought via Latent-Variable Inference · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
0.712023
Sequential Monte Carlo Learning for Time Series Structure Discovery · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.712023
Sequential Monte Carlo Learning for Time Series Structure Discovery · ICML 2023
Machine learning › Trustworthy machine learning
robustness
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Trustworthy machine learning › robustness
underspecification
0.612022
Underspecification Presents Challenges for Credibility in Modern Machine Learning · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning
stochastic gradient descent
0.522017
Stochastic Gradient Descent as Approximate Bayesian Inference · J. Mach. Learn. Res. 2017
A Variational Analysis of Stochastic Gradient Algorithms · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
stochastic variational inference
0.532015
A trust-region method for stochastic variational inference with applications to streaming data · ICML 2015
Stochastic variational inference · J. Mach. Learn. Res. 2013
Sparse stochastic inference for latent Dirichlet allocation · ICML 2012
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.512021
What Are Bayesian Neural Network Posteriors Really Like? · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › gradient-based variational inference
black-box variational inference
0.412020
Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
fokker-planck equation
0.412020
Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
langevin dynamics
0.412020
Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics · ICML 2020
Machine learning › Deep learning architectures and training
reparameterization
0.412020
Automatic Reparameterisation of Probabilistic Programs · ICML 2020
Machine learning › Generative modeling
autoregressive model
0.412019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Audio and music processing
music generation
0.412019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Machine learning › Generative modeling › diffusion model
conditional generation
0.312018
Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models · ICLR (Poster) 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models · ICLR (Poster) 2018
Recommender systems › collaborative filtering
autoencoder-based collaborative filtering
0.312018
Variational Autoencoders for Collaborative Filtering · WWW 2018

Methods — techniques the papers use, named apart from their topics

variational inference · 2.5normalizing flow · 1.5markov chain monte carlo · 1.4bayesian inference · 1.1neural radiance field · 0.8diffusion model · 0.8sequential monte carlo · 0.7involutive MCMC · 0.7control variates · 0.7bayesian nonparametric prior · 0.7pattern pruning · 0.6pattern mining · 0.6coordinated exploration · 0.6self-attention · 0.4relative positional encoding · 0.4variational autoencoder · 0.3annealing · 0.3deep neural network · 0.3
YearPublicationVenuePosition
2024 Robust Inverse Graphics via Probabilistic Inference
abstract
How do we infer a 3D scene from a single image in the presence of corruptions like rain, snow or fog? Straightforward domain randomization relies on knowing the family of corruptions ahead of time. Here, we propose a Bayesian approach—dubbed robust inverse graphics (RIG)—that relies on a strong scene prior and an uninformative uniform corruption prior, making it applicable to a wide range of corruptions. Given a single image, RIG performs posterior inference jointly over the scene and the corruption. We demonstrate this idea by training a neural radiance field (NeRF) scene prior and using a secondary NeRF to represent the corruptions over which we place an uninformative prior. RIG, trained only on clean data, outperforms depth estimators and alternative NeRF approaches that perform point estimation instead of full inference. The results hold for a number of scene prior architectures based on normalizing flows and diffusion models. For the latter, we develop reconstruction-guidance with auxiliary latents (ReGAL)—a diffusion conditioning algorithm that is applicable in the presence of auxiliary latent variables such as the corruption. RIG demonstrates how scene priors can be used beyond generation tasks.
Pavel Sountsov, Matthew Hoffman 0001, Ben Lee, Brian Patton, Rif A. Saurous
ICML3
2023 ProbNeRF: Uncertainty-Aware Inference of 3D Shapes from 2D Images
abstract
The problem of inferring object shape from a single 2D image is underconstrained. Prior knowledge about what objects are plausible can help, but even given such prior knowledge there may still be uncertainty about the shapes of occluded parts of objects. Recently, conditional neural radiance field (NeRF) models have been developed that can learn to infer good point estimates of 3D models from single 2D images. The problem of inferring uncertainty estimates for these models has received less attention. In this work, we propose probabilistic NeRF (ProbNeRF), a model and inference strategy for learning probabilistic generative models of 3D objects’ shapes and appearances, and for doing posterior inference to recover those properties from 2D images. ProbNeRF is trained as a variational autoencoder, but at test time we use Hamiltonian Monte Carlo (HMC) for inference. Given one or a few 2D images of an object (which may be partially occluded), ProbNeRF is able not only to accurately model the parts it sees, but also to propose realistic and diverse hypotheses about the parts it does not see. We show that key to the success of ProbNeRF are (i) a deterministic rendering scheme, (ii) an annealed-HMC strategy, (iii) a hypernetwork-based decoder architecture, and (iv) doing inference over a full set of NeRF weights, rather than just a low-dimensional code. Videos and code are available at https://probnerf.github.io.
Matthew Hoffman 0001, Pavel Sountsov, Christopher Suter, Ben Lee, Vikash Mansinghka 0001, Rif A. Saurous
AISTATS1
2023 Sequential Monte Carlo Learning for Time Series Structure Discovery
abstract
This paper presents a new approach to automatically discovering accurate models of complex time series data. Working within a Bayesian nonparametric prior over a symbolic space of Gaussian process time series models, we present a novel structure learning algorithm that integrates sequential Monte Carlo (SMC) and involutive MCMC for highly effective posterior inference. Our method can be used both in "online” settings, where new data is incorporated sequentially in time, and in “offline” settings, by using nested subsets of historical data to anneal the posterior. Empirical measurements on real-world time series show that our method can deliver 10x–100x runtime speedups over previous MCMC and greedy-search structure learning algorithms targeting the same model family. We use our method to perform the first large-scale evaluation of Gaussian process time series structure learning on a prominent benchmark of 1,428 econometric datasets. The results show that our method discovers sensible models that deliver more accurate point forecasts and interval forecasts over multiple horizons as compared to widely used statistical and neural baselines that struggle on this challenging data.
Feras Saad, Brian Patton, Matthew Hoffman 0001, Rif A. Saurous, Vikash Mansinghka 0001
ICML3
2023 Training Chain-of-Thought via Latent-Variable Inference
abstract
Large language models (LLMs) solve problems more accurately and interpretably when instructed to work out the answer step by step using a "chain-of-thought" (CoT) prompt. One can also improve LLMs' performance on a specific task by supervised fine-tuning, i.e., by using gradient ascent on some tunable parameters to maximize the average log-likelihood of correct answers from a labeled training set. Naively combining CoT with supervised tuning requires supervision not just of the correct answers, but also of detailed rationales that lead to those answers; these rationales are expensive to produce by hand. Instead, we propose a fine-tuning strategy that tries to maximize the \emph{marginal} log-likelihood of generating a correct answer using CoT prompting, approximately averaging over all possible rationales. The core challenge is sampling from the posterior over rationales conditioned on the correct answer; we address it using a simple Markov-chain Monte Carlo (MCMC) expectation-maximization (EM) algorithm inspired by the self-taught reasoner (STaR), memoized wake-sleep, Markovian score climbing, and persistent contrastive divergence. This algorithm also admits a novel control-variate technique that drives the variance of our gradient estimates to zero as the model improves. Applying our technique to GSM8K and the tasks in BIG-Bench Hard, we find that this MCMC-EM fine-tuning technique typically improves the model's accuracy on held-out examples more than STaR or prompt-tuning with or without CoT.
Matthew Hoffman 0001, Du Phan, David Dohan, Sholto Douglas, Aaron Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, Rif A. Saurous
NeurIPS1
2022 Tuning-Free Generalized Hamiltonian Monte Carlo
abstract
Hamiltonian Monte Carlo (HMC) has become a go-to family of Markov chain Monte Carlo (MCMC) algorithms for Bayesian inference problems, in part because we have good procedures for automatically tuning its parameters. Much less attention has been paid to automatic tuning of generalized HMC (GHMC), in which the auxiliary momentum vector is partially updated frequently instead of being completely resampled infrequently. Since GHMC spreads progress over many iterations, it is not straightforward to tune GHMC based on quantities typically used to tune HMC such as average acceptance rate and squared jumped distance. In this work, we propose an ensemble-chain adaptation (ECA) algorithm for GHMC that automatically selects values for all of GHMC’s tunable parameters each iteration based on statistics collected from a population of many chains. This algorithm is designed to make good use of SIMD hardware accelerators such as GPUs, allowing most chains to be updated in parallel each iteration. Unlike typical adaptive-MCMC algorithms, our ECA algorithm does not perturb the chain’s stationary distribution, and therefore does not need to be “frozen” after warmup. Empirically, we find that the proposed algorithm quickly converges to its stationary distribution, producing accurate estimates of posterior expectations with relatively few gradient evaluations per chain.
Matthew Hoffman 0001, Pavel Sountsov
AISTATS1
2022 Underspecification Presents Challenges for Credibility in Modern Machine Learning
abstract
Machine learning (ML) systems often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification in ML pipelines as a key reason for these failures. An ML pipeline is the full procedure followed to train and validate a predictor. Such a pipeline is underspecified when it can return many distinct predictors with equivalently strong test performance. Underspecification is common in modern ML pipelines that primarily validate predictors on held-out data that follow the same distribution as the training data. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment domains. This ambiguity can lead to instability and poor model behavior in practice, and is a distinct failure mode from previously identified issues arising from structural mismatch between training and deployment domains. We provide evidence that underspecfication has substantive implications for practical ML pipelines, using examples from computer vision, medical imaging, natural language processing, clinical risk prediction based on electronic health records, and medical genomics. Our results show the need to explicitly account for underspecification in modeling pipelines that are intended for real-world deployment in any domain.
Alexander D'Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew Hoffman 0001, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman 0003, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin G. Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang 0002, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, D. Sculley
J. Mach. Learn. Res.10
2021 An Adaptive-MCMC Scheme for Setting Trajectory Lengths in Hamiltonian Monte Carlo
abstract
Hamiltonian Monte Carlo (HMC) is a powerful MCMC algorithm based on simulating Hamiltonian dynamics. Its performance depends strongly on choosing appropriate values for two parameters: the step size used in the simulation, and how long the simulation runs for. The step-size parameter can be tuned using standard adaptive-MCMC strategies, but it is less obvious how to tune the simulation-length parameter. The no-U-turn sampler (NUTS) eliminates this problematic simulation-length parameter, but NUTS’s relatively complex control flow makes it difficult to efficiently run many parallel chains on accelerators such as GPUs. NUTS also spends some extra gradient evaluations relative to HMC in order to decide how long to run each iteration without violating detailed balance. We propose ChEES-HMC, a simple adaptive-MCMC scheme for automatically tuning HMC’s simulation-length parameter, which minimizes a proxy for the autocorrelation of the state’s second moments. We evaluate ChEES-HMC and NUTS on many tasks, and find that ChEES-HMC typically yields larger effective sample sizes per gradient evaluation than NUTS does. When running many chains on a GPU, ChEES-HMC can also run significantly more gradient evaluations per second than NUTS, allowing it to quickly provide accurate estimates of posterior expectations.
Matthew Hoffman 0001, Alexey Radul, Pavel Sountsov
AISTATS1
2021 What Are Bayesian Neural Network Posteriors Really Like?
abstract
The posterior over Bayesian neural network (BNN) parameters is extremely high-dimensional and non-convex. For computational reasons, researchers approximate this posterior using inexpensive mini-batch methods such as mean-field variational inference or stochastic-gradient Markov chain Monte Carlo (SGMCMC). To investigate foundational questions in Bayesian deep learning, we instead use full batch Hamiltonian Monte Carlo (HMC) on modern architectures. We show that (1) BNNs can achieve significant performance gains over standard training and deep ensembles; (2) a single long HMC chain can provide a comparable representation of the posterior to multiple shorter chains; (3) in contrast to recent studies, we find posterior tempering is not needed for near-optimal performance, with little evidence for a “cold posterior” effect, which we show is largely an artifact of data augmentation; (4) BMA performance is robust to the choice of prior scale, and relatively similar for diagonal Gaussian, mixture of Gaussian, and logistic priors; (5) Bayesian neural networks show surprisingly poor generalization under domain shift; (6) while cheaper alternatives such as deep ensembles and SGMCMC can provide good generalization, their predictive distributions are distinct from HMC. Notably, deep ensemble predictive distributions are similarly close to HMC as standard SGLD, and closer than standard variational inference.
Pavel Izmailov, Sharad Vikram, Matthew Hoffman 0001, Andrew Gordon Wilson
ICML3
2020 Hamiltonian Monte Carlo Swindles
abstract
Hamiltonian Monte Carlo (HMC) is a powerful Markov chain Monte Carlo (MCMC) algorithm for estimating expectations with respect to continuous un-normalized probability distributions. MCMC estimators typically have higher variance than classical Monte Carlo with i.i.d. samples due to autocorrelations; most MCMC research tries to reduce these autocorrelations. In this work, we explore a complementary approach to variance reduction based on two classical Monte Carlo ’swindles’: first, running an auxiliary coupled chain targeting a tractable approximation to the target distribution, and using the auxiliary samples as control variates; and second, generating anti-correlated ("antithetic") samples by running two chains with flipped randomness. Both ideas have been explored previously in the context of Gibbs samplers and random-walk Metropolis algorithms, but we argue that they are ripe for adaptation to HMC in light of recent coupling results from the HMC theory literature. For many posterior distributions, we find that these swindles generate effective sample sizes orders of magnitude larger than plain HMC, as well as being more efficient than analogous swindles for Metropolis-adjusted Langevin algorithm and random-walk Metropolis.
Dan Piponi, Matthew Hoffman 0001, Pavel Sountsov
AISTATS2
2020 Automatic Reparameterisation of Probabilistic Programs
abstract
Probabilistic programming has emerged as a powerful paradigm in statistics, applied science, and machine learning: by decoupling modelling from inference, it promises to allow modellers to directly reason about the processes generating data. However, the performance of inference algorithms can be dramatically affected by the parameterisation used to express a model, requiring users to transform their programs in non-intuitive ways. We argue for automating these transformations, and demonstrate that mechanisms available in recent modelling frameworks can implement non-centring and related reparameterisations. This enables new inference algorithms, and we propose two: a simple approach using interleaved sampling and a novel variational formulation that searches over a continuous space of parameterisations. We show that these approaches enable robust inference across a range of models, and can yield more efficient samplers than the best fixed parameterisation.
Maria I. Gorinova 0001, Dave Moore, Matthew Hoffman 0001
ICML3
2020 Black-Box Variational Inference as a Parametric Approximation to Langevin Dynamics
abstract
Variational inference (VI) and Markov chain Monte Carlo (MCMC) are approximate posterior inference algorithms that are often said to have complementary strengths, with VI being fast but biased and MCMC being slower but asymptotically unbiased. In this paper, we analyze gradient-based MCMC and VI procedures and find theoretical and empirical evidence that these procedures are not as different as one might think. In particular, a close examination of the Fokker-Planck equation that governs the Langevin dynamics (LD) MCMC procedure reveals that LD implicitly follows a gradient flow that corresponds to a variational inference procedure based on optimizing a nonparametric normalizing flow. This result suggests that the transient bias of LD (due to the Markov chain not having burned in) may track that of VI (due to the optimizer not having converged), up to differences due to VI’s asymptotic bias and parameterization. Empirically, we find that the transient biases of these algorithms (and their momentum-accelerated counterparts) do evolve similarly. This suggests that practitioners with a limited time budget may get more accurate results by running an MCMC procedure (even if it’s far from burned in) than a VI procedure, as long as the variance of the MCMC estimator can be dealt with (e.g., by running many parallel chains).
Matthew Hoffman 0001, Yi-An Ma
ICML1
2019 The LORACs Prior for VAEs: Letting the Trees Speak for the Data
abstract
In variational autoencoders, the prior on the latent codes $z$ is often treated as an afterthought, but the prior shapes the kind of latent representation that the model learns. If the goal is to learn a representation that is interpretable and useful, then the prior should reflect the ways in which the high-level factors that describe the data vary. The “default” prior is a standard normal, but if the natural factors of variation in the dataset exhibit discrete structure or are not independent, then the isotropic-normal prior will actually encourage learning representations that \emph{mask} this structure. To alleviate this problem, we propose using a flexible Bayesian nonparametric hierarchical clustering prior based on the time-marginalized coalescent (TMC). To scale learning to large datasets, we develop a new inducing-point approximation and inference algorithm. We then apply the method without supervision to several datasets and examine the interpretability and practical performance of the inferred hierarchies and learned latent space.
Sharad Vikram, Matthew Hoffman 0001, Matthew J. Johnson 0002
AISTATS2
2019 Music Transformer: Generating Music with Long-Term Structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew Hoffman 0001, Monica Dinculescu, Douglas Eck
ICLR (Poster)8
2018 On the challenges of learning with inference networks on sparse, high-dimensional data
abstract
We study parameter estimation in Nonlinear Factor Analysis (NFA) where the generative model is parameterized by a deep neural network. Recent work has focused on learning such models using inference (or recognition) networks; we identify a crucial problem when modeling large, sparse, high-dimensional datasets – underfitting. We study the extent of underfitting, highlighting that its severity increases with the sparsity of the data. We propose methods to tackle it via iterative optimization inspired by stochastic variational inference (Hoffman et al., 2013) and improvements in the data representation used for inference. The proposed techniques drastically improve the ability of these powerful models to fit sparse data, achieving state-of-the-art results on a benchmark text-count dataset and excellent results on the task of top-N recommendation.
Rahul G. Krishnan, Dawen Liang, Matthew Hoffman 0001
AISTATS3
2018 Multimodal Prediction and Personalization of Photo Edits with Deep Generative Models
abstract
Professional-grade software applications are powerful but complicated – expert users can achieve impressive results, but novices often struggle to complete even basic tasks. Photo editing is a prime example: after loading a photo, the user is confronted with an array of cryptic sliders like "clarity", "temp", and "highlights". An automatically generated suggestion could help, but there is no single "correct" edit for a given image – different experts may make very different aesthetic decisions when faced with the same image, and a single expert may make different choices depending on the intended use of the image (or on a whim). We therefore want a system that can propose multiple diverse, high-quality edits while also learning from and adapting to a user’s aesthetic preferences. In this work, we develop a statistical model that meets these objectives. Our model builds on recent advances in neural network generative modeling and scalable inference, and uses hierarchical structure to learn editing patterns across many diverse users. Empirically, we find that our model outperforms other approaches on this challenging multimodal prediction task.
Ardavan Saeedi, Matthew Hoffman 0001, Stephen DiVerdi, Asma Ghandeharioun, Matthew J. Johnson 0002, Ryan P. Adams
AISTATS2
2018 Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models
Jesse H. Engel, Matthew Hoffman 0001, Adam Roberts
ICLR (Poster)2
2018 Generalizing Hamiltonian Monte Carlo with Neural Networks
Daniel Levy 0002, Matthew Hoffman 0001, Jascha Sohl-Dickstein
ICLR (Poster)2
2018 Autoconj: Recognizing and Exploiting Conjugacy Without a Domain-Specific Language
abstract
Deriving conditional and marginal distributions using conjugacy relationships can be time consuming and error prone. In this paper, we propose a strategy for automating such derivations. Unlike previous systems which focus on relationships between pairs of random variables, our system (which we call Autoconj) operates directly on Python functions that compute log-joint distribution functions. Autoconj provides support for conjugacy-exploiting algorithms in any Python-embedded PPL. This paves the way for accelerating development of novel inference algorithms and structure-exploiting modeling strategies. The package can be downloaded at https://github.com/google-research/autoconj.
Matthew Hoffman 0001
NeurIPS1
2018 Simple, Distributed, and Accelerated Probabilistic Programming
abstract
We describe a simple, low-level approach for embedding probabilistic programming in a deep learning ecosystem. In particular, we distill probabilistic programming down to a single abstraction—the random variable. Our lightweight implementation in TensorFlow enables numerous applications: a model-parallel variational auto-encoder (VAE) with 2nd-generation tensor processing units (TPUv2s); a data-parallel autoregressive model (Image Transformer) with TPUv2s; and multi-GPU No-U-Turn Sampler (NUTS). For both a state-of-the-art VAE on 64x64 ImageNet and Image Transformer on 256x256 CelebA-HQ, our approach achieves an optimal linear speedup from 1 to 256 TPUv2 chips. With NUTS, we see a 100x speedup on GPUs over Stan and 37x over PyMC3.
Dustin Tran, Matthew Hoffman 0001, Dave Moore, Christopher Suter, Srinivas Vasudevan, Alexey Radul
NeurIPS2
2018 Variational Autoencoders for Collaborative Filtering
abstract
We extend variational autoencoders (VAEs) to collaborative filtering for implicit feedback. This non-linear probabilistic model enables us to go beyond the limited modeling capacity of linear factor models which still largely dominate collaborative filtering research.We introduce a generative model with multinomial likelihood and use Bayesian inference for parameter estimation. Despite widespread use in language modeling and economics, the multinomial likelihood receives less attention in the recommender systems literature. We introduce a different regularization parameter for the learning objective, which proves to be crucial for achieving competitive performance. Remarkably, there is an efficient way to tune the parameter using annealing. The resulting model and learning algorithm has information-theoretic connections to maximum entropy discrimination and the information bottleneck principle. Empirically, we show that the proposed approach significantly outperforms several state-of-the-art baselines, including two recently-proposed neural network approaches, on several real-world datasets. We also provide extended experiments comparing the multinomial likelihood with other commonly used likelihood functions in the latent factor collaborative filtering literature and show favorable results. Finally, we identify the pros and cons of employing a principled Bayesian inference approach and characterize settings where it provides the most significant improvements.
Dawen Liang, Rahul G. Krishnan, Matthew Hoffman 0001, Tony Jebara
WWW3
2018 Characterizing User Skills from Application Usage Traces with Hierarchical Attention Recurrent Networks
abstract
Predicting users’ proficiencies is a critical component of AI-powered personal assistants. This article introduces a novel approach for the prediction based on users’ diverse, noisy, and passively generated application usage histories. We propose a novel bi-directional recurrent neural network with hierarchical attention mechanism to extract sequential patterns and distinguish informative traces from noise. Our model is able to attend to the most discriminative actions and sessions to make more accurate and directly interpretable predictions while requiring 50× less training data than the state-of-the-art sequential learning approach. We evaluate our model with two large scale datasets collected from 68K Photoshop users: a digital design skill dataset where the user skill is determined by the quality of the end products and a software skill dataset where users self-disclose their software usage skill levels. The empirical results demonstrate our model’s superior performance compared to existing user representation learning techniques that leverage action frequencies and sequential patterns. In addition, we qualitatively illustrate the model’s significant interpretative power. The proposed approach is broadly relevant to applications that generate user time-series analytics.
Longqi Yang 0001, Hailin Jin, Matthew Hoffman 0001, Deborah Estrin
ACM Trans. Intell. Syst. Technol.4
2017 Deep Probabilistic Programming
Dustin Tran, Matthew Hoffman 0001, Rif A. Saurous, Eugene Brevdo, Kevin Murphy 0002, David M. Blei
ICLR (Poster)2
2017 Learning Deep Latent Gaussian Models with Markov Chain Monte Carlo
abstract
Deep latent Gaussian models are powerful and popular probabilistic models of high-dimensional data. These models are almost always fit using variational expectation-maximization, an approximation to true maximum-marginal-likelihood estimation. In this paper, we propose a different approach: rather than use a variational approximation (which produces biased gradient signals), we use Markov chain Monte Carlo (MCMC, which allows us to trade bias for computation). We find that our MCMC-based approach has several advantages: it yields higher held-out likelihoods, produces sharper images, and does not suffer from the variational overpruning effect. MCMC’s additional computational overhead proves to be significant, but not prohibitive.
Matthew Hoffman 0001
ICML1
2017 CoreFlow: Extracting and Visualizing Branching Patterns from Event Sequences
abstract
Abstract Event sequence datasets with high event cardinality and long sequences are difficult to visualize and analyze. In particular, it is hard to generate a high level visual summary of paths and volume of flow. Existing approaches of mining and visualizing frequent sequential patterns look promising, but have limitations in terms of scalability, interpretability and utility. We propose CoreFlow, a technique that automatically extracts and visualizes branching patterns in event sequences. CoreFlow constructs a tree by recursively applying a three‐step procedure: rank events, divide sequences into groups, and trim sequences by the chosen event. The resulting tree contains key events as nodes, and links represent aggregated flows between key events. Based on CoreFlow, we have developed an interactive system for event sequence analysis. Our approach can compute branching patterns for millions of events in a few seconds, with improved interpretability of extracted patterns compared to previous work. We also present case studies of using the system in three different domains and discuss success and failure cases of applying CoreFlow to real‐world analytic problems. These case studies call forth future research on metrics and models to evaluate the quality of visual summaries of event sequences.
Zhicheng Liu 0001, Bernard Kerr, Mira Dontcheva, Justin Grover, Matthew Hoffman 0001, Alan Wilson 0004
Comput. Graph. Forum5
2017 Stochastic Gradient Descent as Approximate Bayesian Inference
abstract
Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show that constant SGD can be used as an approximate Bayesian posterior inference algorithm. Specifically, we show how to adjust the tuning parameters of constant SGD to best match the stationary distribution to a posterior, minimizing the Kullback-Leibler divergence between these two distributions. (2) We demonstrate that constant SGD gives rise to a new variational EM algorithm that optimizes hyperparameters in complex probabilistic models. (3) We also show how to tune SGD with momentum for approximate sampling. (4) We analyze stochastic-gradient MCMC algorithms. For Stochastic- Gradient Langevin Dynamics and Stochastic-Gradient Fisher Scoring, we quantify the approximation errors due to finite learning rates. Finally (5), we use the stochastic process perspective to give a short proof of why Polyak averaging is optimal. Based on this idea, we propose a scalable approximate MCMC algorithm, the Averaged Stochastic Gradient Sampler.
Stephan Mandt, Matthew Hoffman 0001, David M. Blei
J. Mach. Learn. Res.2
2017 Patterns and Sequences: Interactive Exploration of Clickstreams to Understand Common Visitor Paths
abstract
Modern web clickstream data consists of long, high-dimensional sequences of multivariate events, making it difficult to analyze. Following the overarching principle that the visual interface should provide information about the dataset at multiple levels of granularity and allow users to easily navigate across these levels, we identify four levels of granularity in clickstream analysis: patterns, segments, sequences and events. We present an analytic pipeline consisting of three stages: pattern mining, pattern pruning and coordinated exploration between patterns and sequences. Based on this approach, we discuss properties of maximal sequential patterns, propose methods to reduce the number of patterns and describe design considerations for visualizing the extracted sequential patterns and the corresponding raw sequences. We demonstrate the viability of our approach through an analysis scenario and discuss the strengths and limitations of the methods based on user feedback.
Zhicheng Liu 0001, Mira Dontcheva, Matthew Hoffman 0001, Seth Walker, Alan Wilson 0004
IEEE Trans. Vis. Comput. Graph.4
2016 Fast and easy crowdsourced perceptual audio evaluation
abstract
Automated objective methods of audio evaluation are fast, cheap, and require little effort by the investigator. However, objective evaluation methods do not exist for the output of all audio processing algorithms, often have output that correlates poorly with human quality assessments, and require ground truth data in their calculation. Subjective human ratings of audio quality are the gold standard for many tasks, but are expensive, slow, and require a great deal of effort to recruit subjects and run listening tests. Moving listening tests from the lab to the micro-task labor market of Amazon Mechanical Turk speeds data collection and reduces investigator effort. However, it also reduces the amount of control investigators have over the testing environment, adding new variability and potential biases to the data. In this work, we compare multiple stimulus listening tests performed in a lab environment to multiple stimulus listening tests performed in web environment on a population drawn from Mechanical Turk.
Mark Cartwright, Bryan Pardo, Gautham J. Mysore, Matthew Hoffman 0001
ICASSP4
2016 A Variational Analysis of Stochastic Gradient Algorithms
abstract
Stochastic Gradient Descent (SGD) is an important algorithm in machine learning. With constant learning rates, it is a stochastic process that, after an initial phase of convergence, generates samples from a stationary distribution. We show that SGD with constant rates can be effectively used as an approximate posterior inference algorithm for probabilistic modeling. Specifically, we show how to adjust the tuning parameters of SGD such as to match the resulting stationary distribution to the posterior. This analysis rests on interpreting SGD as a continuous-time stochastic process and then minimizing the Kullback-Leibler divergence between its stationary distribution and the target posterior. (This is in the spirit of variational inference.) In more detail, we model SGD as a multivariate Ornstein-Uhlenbeck process and then use properties of this process to derive the optimal parameters. This theoretical framework also connects SGD to modern scalable inference algorithms; we analyze the recently proposed stochastic gradient Fisher scoring under this perspective. We demonstrate that SGD with properly chosen constant rates gives a new way to optimize hyperparameters in probabilistic models.
Stephan Mandt, Matthew Hoffman 0001, David M. Blei
ICML2
2016 The Segmented iHMM: A Simple, Efficient Hierarchical Infinite HMM
abstract
We propose the segmented iHMM (siHMM), a hierarchical infinite hidden Markov model (iHMM) that supports a simple, efficient inference scheme. The siHMM is well suited to segmentation problems, where the goal is to identify points at which a time series transitions from one relatively stable regime to a new regime. Conventional iHMMs often struggle with such problems, since they have no mechanism for distinguishing between high-and low-level dynamics. Hierarchical HMMs (HHMMs) can do better, but they require much more complex and expensive inference algorithms. The siHMM retains the simplicity and efficiency of the iHMM, but outperforms it on a variety of segmentation problems, achieving performance that matches or exceeds that of a more complicated HHMM.
Ardavan Saeedi, Matthew Hoffman 0001, Matthew J. Johnson 0002, Ryan P. Adams
ICML2
2016 Scalable Nonparametric Bayesian Multilevel Clustering
Viet Huynh, Dinh Q. Phung, Svetha Venkatesh, XuanLong Nguyen, Matthew Hoffman 0001, Hung Hai Bui
UAI5
2015 Stochastic Structured Variational Inference
abstract
Stochastic variational inference makes it possible to approximate posterior distributions induced by large datasets quickly using stochastic optimization. The algorithm relies on the use of fully factorized variational distributions. However, this “mean-field” independence approximation limits the fidelity of the posterior approximation, and introduces local optima. We show how to relax the mean-field approximation to allow arbitrary dependencies between global parameters and local hidden variables, producing better parameter estimates by reducing bias, sensitivity to local optima, and sensitivity to hyperparameters.
Matthew Hoffman 0001, David M. Blei
AISTATS1
2015 Speech dereverberation using a learned speech model
abstract
We present a general single-channel speech dereverberation method based on an explicit generative model of reverberant and noisy speech. To regularize the model, we use a pre-learned speech model of clean and dry speech as a prior and perform posterior inference over the latent clean speech. The reverberation kernel and additive noise are estimated under the maximum-likelihood framework. Our model assumes no prior knowledge about specific speakers or rooms, and consequently our method can automatically adapt to various reverberant and noisy conditions. We evaluate the proposed model with both simulated data and real recordings from the REVERB Challenge1in the task of speech enhancement and obtain results comparable to or better than the state-of-the-art.
Dawen Liang, Matthew Hoffman 0001, Gautham J. Mysore
ICASSP2
2015 Celeste: Variational inference for a generative model of astronomical images
abstract
We present a new, fully generative model of optical telescope image sets, along with a variational procedure for inference. Each pixel intensity is treated as a Poisson random variable, with a rate parameter dependent on latent properties of stars and galaxies. Key latent properties are themselves random, with scientific prior distributions constructed from large ancillary data sets. We check our approach on synthetic images. We also run it on images from a major sky survey, where it exceeds the performance of the current state-of-the-art method for locating celestial bodies and measuring their colors.
Jeffrey Regier, Andrew C. Miller, Jon D. McAuliffe, Ryan P. Adams, Matthew Hoffman 0001, Dustin Lang, David Schlegel, Prabhat
ICML5
2015 A trust-region method for stochastic variational inference with applications to streaming data
abstract
Stochastic variational inference allows for fast posterior inference in complex Bayesian models. However, the algorithm is prone to local optima which can make the quality of the posterior approximation sensitive to the choice of hyperparameters and initialization. We address this problem by replacing the natural gradient step of stochastic varitional inference with a trust-region update. We show that this leads to generally better results and reduced sensitivity to hyperparameters. We also describe a new strategy for variational inference on streaming data and show that here our trust-region method is crucial for getting good performance.
Lucas Theis, Matthew Hoffman 0001
ICML2
2014 Exploiting long-term temporal dependencies in NMF using recurrent neural networks with application to source separation
abstract
This paper seeks to exploit high-level temporal information during feature extraction from audio signals via non-negative matrix factorization. Contrary to existing approaches that impose local temporal constraints, we train powerful recurrent neural network models to capture long-term temporal dependencies and event co-occurrence in the data. This gives our method the ability to “fill in the blanks” in a smart way during feature extraction from complex audio mixtures, an ability very useful for a number of audio applications. We apply these ideas to source separation problems.
Nicolas Boulanger-Lewandowski, Gautham J. Mysore, Matthew Hoffman 0001
ICASSP3
2014 Speech decoloration based on the product-of-filters model
abstract
We present a single-channel speech decoloration method based on a recently proposed generative product-of-filters (PoF) model. We take a spectral approach and attempt to learn the magnitude response of the actual coloration filter, given only the degraded speech signal. Experiments on synthetic data demonstrate that the proposed method effectively captures both coarse and fine structure of the coloration filter. On real recordings, we find that simply subtracting the learned coloration filter from the log-spectra yields promising decoloration results.
Dawen Liang, Daniel P. W. Ellis, Matthew Hoffman 0001, Gautham J. Mysore
ICASSP3
2014 The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew Hoffman 0001, Andrew Gelman
J. Mach. Learn. Res.1
2013 Stochastic variational inference
Matthew Hoffman 0001, David M. Blei, Chong Wang 0002, John W. Paisley
J. Mach. Learn. Res.1
2012 Poisson-uniform nonnegative matrix factorization
abstract
Probabilistic models of audio spectrograms used in audio source separation often rely on Poisson or multinomial noise models corresponding to the generalized Kullback-Leibler (GKL) divergence popular in methods using Nonnegative Matrix Factorization (NMF). This noise model works well in practice, but it is difficult to justify since these distributions are technically only applicable to discrete counts data. This issue is particularly problematic in hierarchical and non-parametric Bayesian models where estimates of uncertainty depend strongly on the likelihood model. In this paper, we present a hierarchical Bayesian model that retains the flavor of the Poisson likelihood model but yields a coherent generative process for continuous spectrogram data. This model allows for more principled, accurate, and effective Bayesian inference in probabilistic NMF models based on GKL.
Matthew Hoffman 0001
ICASSP1
2012 Nonparametric variational inference
Samuel Gershman, Matthew Hoffman 0001, David M. Blei
ICML2
2012 Sparse stochastic inference for latent Dirichlet allocation
David M. Mimno, Matthew Hoffman 0001, David M. Blei
ICML2
2010 Bayesian Nonparametric Matrix Factorization for Recorded Music
Matthew Hoffman 0001, David M. Blei, Perry R. Cook
ICML1
2010 Online Learning for Latent Dirichlet Allocation
abstract
We develop an online variational Bayes (VB) algorithm for Latent Dirichlet Allocation (LDA). Online LDA is based on online stochastic optimization with a natural gradient step, which we show converges to a local optimum of the VB objective function. It can handily analyze massive document collections, including those arriving in a stream. We study the performance of online LDA in several ways, including by fitting a 100-topic topic model to 3.3M articles from Wikipedia in a single pass. We demonstrate that online LDA finds topic models as good or better than those found with batch VB, and in a fraction of the time.
Matthew Hoffman 0001, David M. Blei, Francis R. Bach
NIPS1