Jon D. McAuliffe

dblp:80/5859 · also Jon McAuliffe · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0003-2626-7320ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 61% Optimization for machine learning · 22% Reinforcement learning · 6%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 56% Computational science and engineering · 44%

Topics — the 29 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.232023
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference
0.712023
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Computational science and engineering
astronomy
0.632023
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
rao-blackwellization
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Optimization for machine learning › gradient estimation
stochastic gradient estimation
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Optimization for machine learning
variance reduction
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › gradient-based variational inference
black-box variational inference
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Optimization for machine learning
stochastic optimization
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Reinforcement learning › policy optimization
trust region methods
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212015
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Natural language and speech › Language models and text generation
generative inference
0.212015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Bioinformatics and computational biology
cancer genomics
0.212015
GLAD: a mixed-membership model for heterogeneous tumor subtype classification · Bioinform. 2015
Bioinformatics and computational biology › cancer genomics › cancer subtype analysis
cancer subtype classification
0.212015
GLAD: a mixed-membership model for heterogeneous tumor subtype classification · Bioinform. 2015
Machine learning › Probabilistic and Bayesian machine learning
discrete distribution
0.112019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning
reparameterization trick
0.112017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Bioinformatics and computational biology › population genetics › coalescent theory
coalescent model
0.112008
A spatially varying two-sample recombinant coalescent, with applications to HIV escape response · NIPS 2008
Bioinformatics and computational biology
phylogenetics
0.112008
A spatially varying two-sample recombinant coalescent, with applications to HIV escape response · NIPS 2008
Bioinformatics and computational biology › population genetics
recombination
0.112008
A spatially varying two-sample recombinant coalescent, with applications to HIV escape response · NIPS 2008
Bioinformatics and computational biology › molecular evolution
viral evolution
0.112008
A spatially varying two-sample recombinant coalescent, with applications to HIV escape response · NIPS 2008
Natural language and speech › Information extraction and text analysis › topic model
supervised latent dirichlet allocation
0.112007
Supervised Topic Models · NIPS 2007
Natural language and speech › Information extraction and text analysis
text classification
0.112007
Supervised Topic Models · NIPS 2007
Data mining › text mining › topic modeling
supervised topic modeling
0.112007
Supervised Topic Models · NIPS 2007
Data mining › text mining
topic model
0.112007
Supervised Topic Models · NIPS 2007
Bioinformatics and computational biology › genome annotation
functional annotation
0.012004
Multiple-sequence functional annotation and the generalized hidden Markov phylogeny · Bioinform. 2004
Bioinformatics and computational biology › genome annotation
gene structure prediction
0.012004
Multiple-sequence functional annotation and the generalized hidden Markov phylogeny · Bioinform. 2004
Machine learning › Learning theory › loss function › surrogate loss
convex surrogate loss
0.012003
Large Margin Classifiers: Convex Loss, Low Noise, and Convergence Rates · NIPS 2003
Machine learning › Learning theory
statistical learning theory
0.012003
Large Margin Classifiers: Convex Loss, Low Noise, and Convergence Rates · NIPS 2003
Bioinformatics and computational biology
comparative genomics
0.012004
Multiple-sequence functional annotation and the generalized hidden Markov phylogeny · Bioinform. 2004
Machine learning › Kernel, tree and ensemble methods › large margin methods
maximum margin classifiers
0.012003
Large Margin Classifiers: Convex Loss, Low Noise, and Convergence Rates · NIPS 2003

Methods — techniques the papers use, named apart from their topics

variational inference · 1.9forward KL divergence · 1.3MCMC · 1.3poisson model · 0.4bayesian inference · 0.4unbiased estimator · 0.4rao-blackwellization · 0.4trust region optimization · 0.3stochastic second-order optimization · 0.3reparameterization trick · 0.3sparse biomarker signature learning · 0.2mixed-membership modeling · 0.2latent variable model · 0.2gaussian process · 0.2hierarchical bayesian model · 0.1maximum likelihood estimation · 0.1latent dirichlet allocation · 0.1
YearPublicationVenuePosition
2023 Variational Inference for Deblending Crowded Starfields
abstract
In images collected by astronomical surveys, stars and galaxies often overlap visually. Deblending is the task of distinguishing and characterizing individual light sources in survey images. We propose StarNet, a Bayesian method to deblend sources in astronomical images of crowded star fields. StarNet leverages recent advances in variational inference, including amortized variational distributions and an optimization objective targeting an expectation of the forward KL divergence. In our experiments with SDSS images of the M2 globular cluster, StarNet is substantially more accurate than two competing methods: Probabilistic Cataloging (PCAT), a method that uses MCMC for inference, and DAOPHOT, a software pipeline employed by SDSS for deblending. In addition, the amortized approach to inference gives StarNet the scaling characteristics necessary to perform Bayesian inference on modern astronomical surveys.
Runjing Liu, Jon D. McAuliffe, Jeffrey Regier
J. Mach. Learn. Res.2
2019 Rao-Blackwellized Stochastic Gradients for Discrete Distributions
abstract
We wish to compute the gradient of an expectation over a finite or countably infinite sample space having K $\leq$ $\infty$ categories. When K is indeed infinite, or finite but very large, the relevant summation is intractable. Accordingly, various stochastic gradient estimators have been proposed. In this paper, we describe a technique that can be applied to reduce the variance of any such estimator, without changing its bias{—}in particular, unbiasedness is retained. We show that our technique is an instance of Rao-Blackwellization, and we demonstrate the improvement it yields on a semi-supervised classification problem and a pixel attention task.
Runjing Liu, Jeffrey Regier, Nilesh Tripuraneni, Michael I. Jordan, Jon D. McAuliffe
ICML5
2019 Cataloging the visible universe through Bayesian inference in Julia at petascale
Jeffrey Regier, Keno Fischer, Kiran Pamnany, Andreas Noack 0001, Jarrett Revels, Maximilian Lam, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
J. Parallel Distributed Comput.10
2018 Cataloging the Visible Universe Through Bayesian Inference at Petascale
abstract
Astronomical catalogs derived from wide-field imaging surveys are an important tool for understanding the Universe. We construct an astronomical catalog from 55 TB of imaging data using Celeste, a Bayesian variational inference code written entirely in the high-productivity programming language Julia. Using over 1.3 million threads on 650,000 Intel Xeon Phi cores of the Cori Phase II supercomputer, Celeste achieves a peak rate of 1.54 DP PFLOP/s. Celeste is able to jointly optimize parameters for 188M stars and galaxies, loading and processing 178 TB across 8192 nodes in 14.6 minutes. To achieve this, Celeste exploits parallelism at multiple levels (cluster, node, and thread) and accelerates I/O through Cori's Burst Buffer. Julia's native performance enables Celeste to employ high-level constructs without resorting to hand-written or generated low-level code (C/C++/Fortran), and yet achieve petascale performance.
Jeffrey Regier, Kiran Pamnany, Keno Fischer, Andreas Noack 0001, Maximilian Lam, Jarrett Revels, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
IPDPS10
2017 Fast Black-box Variational Inference through Stochastic Trust-Region Optimization
abstract
We introduce TrustVI, a fast second-order algorithm for black-box variational inference based on trust-region optimization and the reparameterization trick. At each iteration, TrustVI proposes and assesses a step based on minibatches of draws from the variational distribution. The algorithm provably converges to a stationary point. We implemented TrustVI in the Stan framework and compared it to two alternatives: Automatic Differentiation Variational Inference (ADVI) and Hessian-free Stochastic Gradient Variational Inference (HFSGVI). The former is based on stochastic first-order optimization. The latter uses second-order information, but lacks convergence guarantees. TrustVI typically converged at least one order of magnitude faster than ADVI, demonstrating the value of stochastic second-order information. TrustVI often found substantially better variational distributions than HFSGVI, demonstrating that our convergence theory can matter in practice.
Jeffrey Regier, Michael I. Jordan, Jon D. McAuliffe
NIPS3
2015 Celeste: Variational inference for a generative model of astronomical images
abstract
We present a new, fully generative model of optical telescope image sets, along with a variational procedure for inference. Each pixel intensity is treated as a Poisson random variable, with a rate parameter dependent on latent properties of stars and galaxies. Key latent properties are themselves random, with scientific prior distributions constructed from large ancillary data sets. We check our approach on synthetic images. We also run it on images from a major sky survey, where it exceeds the performance of the current state-of-the-art method for locating celestial bodies and measuring their colors.
Jeffrey Regier, Andrew C. Miller, Jon D. McAuliffe, Ryan P. Adams, Matthew Hoffman 0001, Dustin Lang, David Schlegel, Prabhat
ICML3
2015 A Gaussian Process Model of Quasar Spectral Energy Distributions
abstract
We propose a method for combining two sources of astronomical data, spectroscopy and photometry, that carry information about sources of light (e.g., stars, galaxies, and quasars) at extremely different spectral resolutions. Our model treats the spectral energy distribution (SED) of the radiation from a source as a latent variable that jointly explains both photometric and spectroscopic observations. We place a flexible, nonparametric prior over the SED of a light source that admits a physically interpretable decomposition, and allows us to tractably perform inference. We use our model to predict the distribution of the redshift of a quasar from five-band (low spectral resolution) photometric data, the so called ``photo-z'' problem. Our method shows that tools from machine learning and Bayesian statistics allow us to leverage multiple resolutions of information to make accurate predictions with well-characterized uncertainties.
Andrew C. Miller, Albert Wu, Jeffrey Regier, Jon D. McAuliffe, Dustin Lang, Prabhat, David Schlegel, Ryan P. Adams
NIPS4
2015 GLAD: a mixed-membership model for heterogeneous tumor subtype classification
abstract
MOTIVATION: Genomic analyses of many solid cancers have demonstrated extensive genetic heterogeneity between as well as within individual tumors. However, statistical methods for classifying tumors by subtype based on genomic biomarkers generally entail an all-or-none decision, which may be misleading for clinical samples containing a mixture of subtypes and/or normal cell contamination. RESULTS: We have developed a mixed-membership classification model, called glad, that simultaneously learns a sparse biomarker signature for each subtype as well as a distribution over subtypes for each sample. We demonstrate the accuracy of this model on simulated data, in-vitro mixture experiments, and clinical samples from the Cancer Genome Atlas (TCGA) project. We show that many TCGA samples are likely a mixture of multiple subtypes. AVAILABILITY: A python module implementing our algorithm is available from http://genomics.wpi.edu/glad/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hachem Saddiki, Jon D. McAuliffe, Patrick Flaherty
Bioinform.2
2008 A spatially varying two-sample recombinant coalescent, with applications to HIV escape response
abstract
Statistical evolutionary models provide an important mechanism for describing and understanding the escape response of a viral population under a particular therapy. We present a new hierarchical model that incorporates spatially varying mutation and recombination rates at the nucleotide level. It also maintains sep- arate parameters for treatment and control groups, which allows us to estimate treatment effects explicitly. We use the model to investigate the sequence evolu- tion of HIV populations exposed to a recently developed antisense gene therapy, as well as a more conventional drug therapy. The detection of biologically rele- vant and plausible signals in both therapy studies demonstrates the effectiveness of the method.
Alexander Braunstein, Zhi Wei 0001, Shane T. Jensen, Jon D. McAuliffe
NIPS4
2007 Supervised Topic Models
abstract
We introduce supervised latent Dirichlet allocation (sLDA), a statistical model of labelled documents. The model accommodates a variety of response types. We derive a maximum-likelihood procedure for parameter estimation, which relies on variational approximations to handle intractable posterior expectations. Prediction problems motivate this research: we use the fitted model to predict response values for new documents. We test sLDA on two real-world problems: movie ratings predicted from reviews, and web page popularity predicted from text descriptions. We illustrate the benefits of sLDA versus modern regularized regression, as well as versus an unsupervised LDA analysis followed by a separate regression.
David M. Blei, Jon D. McAuliffe
NIPS2
2004 Multiple-sequence functional annotation and the generalized hidden Markov phylogeny
abstract
MOTIVATION: Phylogenetic shadowing is a comparative genomics principle that allows for the discovery of conserved regions in sequences from multiple closely related organisms. We develop a formal probabilistic framework for combining phylogenetic shadowing with feature-based functional annotation methods. The resulting model, a generalized hidden Markov phylogeny (GHMP), applies to a variety of situations where functional regions are to be inferred from evolutionary constraints. RESULTS: We show how GHMPs can be used to predict complete shared gene structures in multiple primate sequences. We also describe shadower, our implementation of such a prediction system. We find that shadower outperforms previously reported ab initio gene finders, including comparative human-mouse approaches, on a small sample of diverse exonic regions. Finally, we report on an empirical analysis of shadower's performance which reveals that as few as five well-chosen species may suffice to attain maximal sensitivity and specificity in exon demarcation. AVAILABILITY: A Web server is available at http://bonaire.lbl.gov/shadower
Jon D. McAuliffe, Lior Pachter, Michael I. Jordan
Bioinform.1
2003 Large Margin Classifiers: Convex Loss, Low Noise, and Convergence Rates
abstract
Many classification algorithms, including the support vector machine, boosting and logistic regression, can be viewed as minimum contrast methods that minimize a convex surrogate of the 0-1 loss function. We characterize the statistical consequences of using such a surrogate by pro- viding a general quantitative relationship between the risk as assessed us- ing the 0-1 loss and the risk as assessed using any nonnegative surrogate loss function. We show that this relationship gives nontrivial bounds un- der the weakest possible condition on the loss function—that it satisfy a pointwise form of Fisher consistency for classification. The relationship is based on a variational transformation of the loss function that is easy to compute in many applications. We also present a refined version of this result in the case of low noise. Finally, we present applications of our results to the estimation of convergence rates in the general setting of function classes that are scaled hulls of a finite-dimensional base class.
Peter L. Bartlett, Michael I. Jordan, Jon D. McAuliffe
NIPS3