Andrew Gelman

dblp:60/2648 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0002-6975-2601ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Probabilistic and Bayesian machine learning · 89% Efficient and distributed learning · 11% Question answering and dialogue systems · 0%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
2.262023
Federated Learning as Variational Inference: A Scalable Expectation Propagation Approach · ICLR 2023
Pathfinder: Parallel quasi-Newton variational inference · J. Mach. Learn. Res. 2022
Yes, but Did It Work?: Evaluating Variational Inference · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.932022
Pathfinder: Parallel quasi-Newton variational inference · J. Mach. Learn. Res. 2022
The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo · J. Mach. Learn. Res. 2014
Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo
0.822022
Pathfinder: Parallel quasi-Newton variational inference · J. Mach. Learn. Res. 2022
The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo · J. Mach. Learn. Res. 2014
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling
0.812024
Pareto Smoothed Importance Sampling · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
monte carlo integration
0.812024
Pareto Smoothed Importance Sampling · J. Mach. Learn. Res. 2024
Machine learning › Efficient and distributed learning
federated learning
0.712023
Federated Learning as Variational Inference: A Scalable Expectation Propagation Approach · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection › model averaging
bayesian model combination
0.612022
Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › sampling › posterior sampling
multimodal posterior sampling
0.612022
Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.612022
Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors · J. Mach. Learn. Res. 2022
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming
0.522017
Automatic Differentiation Variational Inference · J. Mach. Learn. Res. 2017
Automatic Variational Inference in Stan · NIPS 2015
Machine learning › Efficient and distributed learning › distributed inference
distributed bayesian inference
0.412020
Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned Data · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.412020
Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned Data · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
posterior approximation evaluation
0.312018
Yes, but Did It Work?: Evaluating Variational Inference · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
approximate bayesian inference
0.212015
Automatic Variational Inference in Stan · NIPS 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
automated variational inference
0.212015
Automatic Variational Inference in Stan · NIPS 2015
Natural language and speech › Question answering and dialogue systems › conversational agents
intelligent assistants
0.011988
FRM: An Intelligent Assistant for Financial Resource Management · AAAI 1988

Methods — techniques the papers use, named apart from their topics

variational inference · 1.5expectation propagation · 1.1automatic differentiation · 1.1importance weighting · 0.8generalized pareto distribution · 0.8quasi-newton optimization · 0.6importance resampling · 0.6cross-validation · 0.6bayesian stacking · 0.6MCMC · 0.6simulation · 0.0intelligent control · 0.0
YearPublicationVenuePosition
2024 Pareto Smoothed Importance Sampling
abstract
Importance weighting is a general way to adjust Monte Carlo integration to account for draws from the wrong distribution, but the resulting estimate can be highly variable when the importance ratios have a heavy right tail. This routinely occurs when there are aspects of the target distribution that are not well captured by the approximating distribution, in which case more stable estimates can be obtained by modifying extreme importance ratios. We present a new method for stabilizing importance weights using a generalized Pareto distribution fit to the upper tail of the distribution of the simulated importance ratios. The method, which empirically performs better than existing methods for stabilizing importance sampling estimates, includes stabilized effective sample size estimates, Monte Carlo error estimates, and convergence diagnostics. The presented Pareto $\hat{k}$ finite sample convergence rate diagnostic is useful for any Monte Carlo estimator.
Aki Vehtari, Daniel Simpson, Andrew Gelman, Yuling Yao, Jonah Gabry
J. Mach. Learn. Res.3
2024 Bayesian workflow for time-varying transmission in stratified compartmental infectious disease transmission models
abstract
Compartmental models that describe infectious disease transmission across subpopulations are central for assessing the impact of non-pharmaceutical interventions, behavioral changes and seasonal effects on the spread of respiratory infections. We present a Bayesian workflow for such models, including four features: (1) an adjustment for incomplete case ascertainment, (2) an adequate sampling distribution of laboratory-confirmed cases, (3) a flexible, time-varying transmission rate, and (4) a stratification by age group. Within the workflow, we benchmarked the performance of various implementations of two of these features (2 and 3). For the second feature, we used SARS-CoV-2 data from the canton of Geneva (Switzerland) and found that a quasi-Poisson distribution is the most suitable sampling distribution for describing the overdispersion in the observed laboratory-confirmed cases. For the third feature, we implemented three methods: Brownian motion, B-splines, and approximate Gaussian processes (aGP). We compared their performance in terms of the number of effective samples per second, and the error and sharpness in estimating the time-varying transmission rate over a selection of ordinary differential equation solvers and tuning parameters, using simulated seroprevalence and laboratory-confirmed case data. Even though all methods could recover the time-varying dynamics in the transmission rate accurately, we found that B-splines perform up to four and ten times faster than Brownian motion and aGPs, respectively. We validated the B-spline model with simulated age-stratified data. We applied this model to 2020 laboratory-confirmed SARS-CoV-2 cases and two seroprevalence studies from the canton of Geneva. This resulted in detailed estimates of the transmission rate over time and the case ascertainment. Our results illustrate the potential of the presented workflow including stratified transmission to estimate age-specific epidemiological parameters. The workflow is freely available in the R package HETTMO, and can be easily adapted and applied to other infectious diseases.
Judith A. Bouman, Anthony Hauser, Simon L. Grimm, Martin Wohlfender, Samir Bhatt, Elizaveta Semenova, Andrew Gelman, Christian L. Althaus, Julien Riou
PLoS Comput. Biol.7
2023 Federated Learning as Variational Inference: A Scalable Expectation Propagation Approach
Philip Greengard, Hongyi Wang 0001, Andrew Gelman, Eric P. Xing
ICLR4
2023 Bayesian spatial modelling of localised SARS-CoV-2 transmission through mobility networks across England
abstract
In the early phases of growth, resurgent epidemic waves of SARS-CoV-2 incidence have been characterised by localised outbreaks. Therefore, understanding the geographic dispersion of emerging variants at the start of an outbreak is key for situational public health awareness. Using telecoms data, we derived mobility networks describing the movement patterns between local authorities in England, which we have used to inform the spatial structure of a Bayesian BYM2 model. Surge testing interventions can result in spatio-temporal sampling bias, and we account for this by extending the BYM2 model to include a random effect for each timepoint in a given area. Simulated-scenario modelling and real-world analyses of each variant that became dominant in England were conducted using our BYM2 model at local authority level in England. Simulated datasets were created using a stochastic metapopulation model, with the transmission rates between different areas parameterised using telecoms mobility data. Different scenarios were constructed to reproduce real-world spatial dispersion patterns that could prove challenging to inference, and we used these scenarios to understand the performance characteristics of the BYM2 model. The model performed better than unadjusted test positivity in all the simulation-scenarios, and in particular when sample sizes were small, or data was missing for geographical areas. Through the analyses of emerging variant transmission across England, we found a reduction in the early growth phase geographic clustering of later dominant variants as England became more interconnected from early 2022 and public health interventions were reduced. We have also shown the recent increased geographic spread and dominance of variants with similar mutations in the receptor binding domain, which may be indicative of convergent evolution of SARS-CoV-2 variants.
Mitzi Morris, Andrew Gelman, Bob Carpenter, William Ferguson, Christopher E. Overton, Martyn Fyles
PLoS Comput. Biol.3
2022 The Worst of Both Worlds: A Comparative Analysis of Errors in Learning from Data in Psychology and Machine Learning
abstract
Arguments that machine learning (ML) is facing a reproducibility and replication crisis suggest that some published claims in research cannot be taken at face value. Concerns inspire analogies to the replication crisis affecting the social and medical sciences. A deeper understanding of what reproducibility concerns in supervised ML research have in common with the replication crisis in experimental science puts the new concerns in perspective, and helps researchers avoid "the worst of both worlds," where ML researchers begin borrowing methodologies from explanatory modeling without understanding their limitations and vice versa. We contribute a comparative analysis of concerns about inductive learning that arise in causal attribution as exemplified in psychology versus predictive modeling as exemplified in ML. We identify common themes in reform discussions, like overreliance on asymptotic theory and non-credible beliefs about real-world data generating processes. We argue that in both fields, claims from learning are implied to generalize outside the specific environment studied (e.g., the input dataset or subject sample, modeling implementation, etc.) but are often difficult to refute due to underspecification of key parts of the learning pipeline. We conclude by discussing risks that arise when sources of errors are misdiagnosed and the need to acknowledge the role of human inductive biases in learning and reform.
Jessica Hullman, Sayash Kapoor, Priyanka Nanayakkara, Andrew Gelman, Arvind Narayanan
AIES4
2022 Stacking for Non-mixing Bayesian Computations: The Curse and Blessing of Multimodal Posteriors
abstract
When working with multimodal Bayesian posterior distributions, Markov chain Monte Carlo (MCMC) algorithms have difficulty moving between modes, and default variational or mode-based approximate inferences will understate posterior uncertainty. And, even if the most important modes can be found, it is difficult to evaluate their relative weights in the posterior. Here we propose an approach using parallel runs of MCMC, variational, or mode-based inference to hit as many modes or separated regions as possible and then combine these using Bayesian stacking, a scalable method for constructing a weighted average of distributions. The result from stacking efficiently samples from multimodal posterior distribution, minimizes cross validation prediction error, and represents the posterior uncertainty better than variational inference, but it is not necessarily equivalent, even asymptotically, to fully Bayesian inference. We present theoretical consistency with an example where the stacked inference approximates the true data generating process from the misspecified model and a non-mixing sampler, from which the predictive performance is better than full Bayesian inference, hence the multimodality can be considered a blessing rather than a curse under model misspecification. We demonstrate practical implementation in several model families: latent Dirichlet allocation, Gaussian process regression, hierarchical regression, horseshoe variable selection, and neural networks.
Yuling Yao, Aki Vehtari, Andrew Gelman
J. Mach. Learn. Res.3
2022 Pathfinder: Parallel quasi-Newton variational inference
abstract
We propose Pathfinder, a variational method for approximately sampling from differentiable probability densities. Starting from a random initialization, Pathfinder locates normal approximations to the target density along a quasi-Newton optimization path, with local covariance estimated using the inverse Hessian estimates produced by the optimizer. Pathfinder returns draws from the approximation with the lowest estimated Kullback-Leibler (KL) divergence to the target distribution. We evaluate Pathfinder on a wide range of posterior distributions, demonstrating that its approximate draws are better than those from automatic differentiation variational inference (ADVI) and comparable to those produced by short chains of dynamic Hamiltonian Monte Carlo (HMC), as measured by 1-Wasserstein distance. Compared to ADVI and short dynamic HMC runs, Pathfinder requires one to two orders of magnitude fewer log density and gradient evaluations, with greater reductions for more challenging posteriors. Importance resampling over multiple runs of Pathfinder improves the diversity of approximate draws, reducing 1-Wasserstein distance further and providing a measure of robustness to optimization failures on plateaus, saddle points, or in minor modes. The Monte Carlo KL divergence estimates are embarrassingly parallelizable in the core Pathfinder algorithm, as are multiple runs in the resampling version, further increasing Pathfinder's speed advantage with multiple cores.
Bob Carpenter, Andrew Gelman, Aki Vehtari
J. Mach. Learn. Res.3
2020 Expectation Propagation as a Way of Life: A Framework for Bayesian Inference on Partitioned Data
abstract
A common divide-and-conquer approach for Bayesian computation with big data is to partition the data, perform local inference for each piece separately, and combine the results to obtain a global posterior approximation. While being conceptually and computationally appealing, this method involves the problematic need to also split the prior for the local inferences; these weakened priors may not provide enough regularization for each separate computation, thus eliminating one of the key advantages of Bayesian methods. To resolve this dilemma while still retaining the generalizability of the underlying local inference method, we apply the idea of expectation propagation (EP) as a framework for distributed Bayesian inference. The central idea is to iteratively update approximations to the local likelihoods given the state of the other approximations and the prior. The present paper has two roles: we review the steps that are needed to keep EP algorithms numerically stable, and we suggest a general approach, inspired by EP, for approaching data partitioning problems in a way that achieves the computational benefits of parallelism while allowing each local update to make use of relevant information from the other sites. In addition, we demonstrate how the method can be applied in a hierarchical context to make use of partitioning of both data and parameters. The paper describes a general algorithmic framework, rather than a specific algorithm, and presents an example implementation for it.
Aki Vehtari, Andrew Gelman, Tuomas Sivula, Pasi Jylänki, Dustin Tran, Swupnil Sahai, Paul Blomstedt, John P. Cunningham, David Schiminovich, Christian P. Robert
J. Mach. Learn. Res.2
2018 Yes, but Did It Work?: Evaluating Variational Inference
abstract
While it’s always possible to compute a variational approximation to a posterior distribution, it can be difficult to discover problems with this approximation. We propose two diagnostic algorithms to alleviate this problem. The Pareto-smoothed importance sampling (PSIS) diagnostic gives a goodness of fit measurement for joint distributions, while simultaneously improving the error in the estimate. The variational simulation-based calibration (VSBC) assesses the average performance of point estimates.
Yuling Yao, Aki Vehtari, Daniel Simpson, Andrew Gelman
ICML4
2017 The statistical significance filter leads to overconfident expectations of replicability
Shravan Vasishth, Andrew Gelman
CogSci2
2017 A Safe Depth Forecasting Model for Insuring Tubewell Installations Against Arsenic Risk in Bangladesh
Matilde Trevisani, Alexander van Geen, Andrew Gelman, Shuky Ehrenberg, John Immel
ICCSA (5)4
2017 Automatic Differentiation Variational Inference
abstract
Probabilistic modeling is iterative. A scientist posits a simple model, fits it to her data, refines it according to her analysis, and repeats. However, fitting complex models to large data is a bottleneck in this process. Deriving algorithms for new models can be both mathematically and computationally challenging, which makes it difficult to efficiently cycle through the steps. To this end, we develop ADVI. Using our method, the scientist only provides a probabilistic model and a dataset, nothing else. ADVI automatically derives an efficient variational inference algorithm, freeing the scientist to refine and explore many models. ADVI supports a broad class of models ---no conjugacy assumptions are required. We study ADVI across ten modern probabilistic models and apply it to a dataset with millions of observations. We deploy ADVI as part of Stan, a probabilistic programming system.
Alp Kucukelbir, Dustin Tran, Rajesh Ranganath, Andrew Gelman, David M. Blei
J. Mach. Learn. Res.4
2015 Automatic Variational Inference in Stan
abstract
Variational inference is a scalable technique for approximate Bayesian inference. Deriving variational inference algorithms requires tedious model-specific calculations; this makes it difficult for non-experts to use. We propose an automatic variational inference algorithm, automatic differentiation variational inference (ADVI); we implement it in Stan (code available), a probabilistic programming system. In ADVI the user provides a Bayesian model and a dataset, nothing else. We make no conjugacy assumptions and support a broad class of models. The algorithm automatically determines an appropriate variational family and optimizes the variational objective. We compare ADVI to MCMC sampling across hierarchical generalized linear models, nonconjugate matrix factorization, and a mixture model. We train the mixture model on a quarter million images. With ADVI we can use variational inference on any model we write in Stan.
Alp Kucukelbir, Rajesh Ranganath, Andrew Gelman, David M. Blei
NIPS3
2014 The No-U-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew Hoffman 0001, Andrew Gelman
J. Mach. Learn. Res.2
1988 FRM: An Intelligent Assistant for Financial Resource Management
Andrew Gelman, Susan Altman, Matt Pallakoff, Ketan Doshi, Catherine Manago, Thomas C. Rindfleisch, Bruce G. Buchanan
AAAI1