Jeffrey Regier

dblp:164/7281 · also Jeffrey C. Regier · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-1472-5235ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Probabilistic and Bayesian machine learning · 68% Optimization for machine learning · 9% Learning theory · 7%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Computational science and engineering · 88% Bioinformatics and computational biology · 12%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 22 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
3.162024
Globally Convergent Variational Inference · NeurIPS 2024
Variational Inference with Coverage Guarantees in Simulation-Based Inference · ICML 2024
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
amortized variational inference
1.422024
Variational Inference with Coverage Guarantees in Simulation-Based Inference · ICML 2024
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference › simulation-based inference
neural posterior estimation
0.812024
Globally Convergent Variational Inference · NeurIPS 2024
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel analysis
0.812024
Globally Convergent Variational Inference · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference
0.812024
Variational Inference with Coverage Guarantees in Simulation-Based Inference · ICML 2024
Computational science and engineering
astronomy
0.632023
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Variational Inference for Deblending Crowded Starfields · J. Mach. Learn. Res. 2023
Machine learning › Trustworthy machine learning › risk control
false discovery rate control
0.612022
Normalizing Flows for Knockoff-free Controlled Feature Selection · NeurIPS 2022
Machine learning › Generative modeling
variational autoencoder
0.522020
Information Constraints on Auto-Encoding Variational Bayes · NeurIPS 2018
Decision-Making with Auto-Encoding Variational Bayes · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
rao-blackwellization
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Optimization for machine learning › gradient estimation
stochastic gradient estimation
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Optimization for machine learning
variance reduction
0.412019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.312018
Information Constraints on Auto-Encoding Variational Bayes · NeurIPS 2018
Mathematical optimization
nonconvex optimization
0.312018
Stochastic Cubic Regularization for Fast Nonconvex Optimization · NeurIPS 2018
Mathematical optimization › nonconvex optimization
saddle point escape
0.312018
Stochastic Cubic Regularization for Fast Nonconvex Optimization · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › gradient-based variational inference
black-box variational inference
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Optimization for machine learning
stochastic optimization
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Reinforcement learning › policy optimization
trust region methods
0.312017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212015
A Gaussian Process Model of Quasar Spectral Energy Distributions · NIPS 2015
Natural language and speech › Language models and text generation
generative inference
0.212015
Celeste: Variational inference for a generative model of astronomical images · ICML 2015
Machine learning › Probabilistic and Bayesian machine learning
discrete distribution
0.112019
Rao-Blackwellized Stochastic Gradients for Discrete Distributions · ICML 2019
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing
0.112018
Information Constraints on Auto-Encoding Variational Bayes · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning
reparameterization trick
0.112017
Fast Black-box Variational Inference through Stochastic Trust-Region Optimization · NIPS 2017

Methods — techniques the papers use, named apart from their topics

variational inference · 2.4MCMC · 1.9forward KL divergence · 1.3reproducing kernel hilbert space · 0.8neural tangent kernel · 0.8neural posterior estimation · 0.8evidence lower bound · 0.8conformal prediction · 0.8normalizing flow · 0.6knockoffs · 0.6stochastic gradient · 0.3kernel independence measure · 0.3hilbert-schmidt independence criterion · 0.3hessian-vector product · 0.3cubic-regularized newton method · 0.3poisson model · 0.2latent variable model · 0.2gaussian process · 0.2
YearPublicationVenuePosition
2024 Sequential Monte Carlo for Inclusive KL Minimization in Amortized Variational Inference
abstract
For training an encoder network to perform amortized variational inference, the Kullback-Leibler (KL) divergence from the exact posterior to its approximation, known as the inclusive or forward KL, is an increasingly popular choice of variational objective due to the mass-covering property of its minimizer. However, minimizing this objective is challenging. A popular existing approach, Reweighted Wake-Sleep (RWS), suffers from heavily biased gradients and a circular pathology that results in highly concentrated variational distributions. As an alternative, we propose SMC-Wake, a procedure for fitting an amortized variational approximation that uses likelihood-tempered sequential Monte Carlo samplers to estimate the gradient of the inclusive KL divergence. We propose three gradient estimators, all of which are asymptotically unbiased in the number of iterations and two of which are strongly consistent. Our method interleaves stochastic gradient updates, SMC samplers, and iterative improvement to an estimate of the normalizing constant to reduce bias from self-normalization. In experiments with both simulated and real datasets, SMC-Wake fits variational distributions that approximate the posterior more accurately than existing methods.
Declan McNamara, Jackson Loper, Jeffrey Regier
AISTATS3
2024 Variational Inference with Coverage Guarantees in Simulation-Based Inference
abstract
Amortized variational inference is an often employed framework in simulation-based inference that produces a posterior approximation that can be rapidly computed given any new observation. Unfortunately, there are few guarantees about the quality of these approximate posteriors. We propose Conformalized Amortized Neural Variational Inference (CANVI), a procedure that is scalable, easily implemented, and provides guaranteed marginal coverage. Given a collection of candidate amortized posterior approximators, CANVI constructs conformalized predictors based on each candidate, compares the predictors using a metric known as predictive efficiency, and returns the most efficient predictor. CANVI ensures that the resulting predictor constructs regions that contain the truth with a user-specified level of probability. CANVI is agnostic to design decisions in formulating the candidate approximators and only requires access to samples from the forward model, permitting its use in likelihood-free settings. We prove lower bounds on the predictive efficiency of the regions produced by CANVI and explore how the quality of a posterior approximation relates to the predictive efficiency of prediction regions based on that approximation. Finally, we demonstrate the accurate calibration and high predictive efficiency of CANVI on a suite of simulation-based inference benchmark tasks and an important scientific task: analyzing galaxy emission spectra.
Yash P. Patel, Declan McNamara, Jackson Loper, Jeffrey Regier, Ambuj Tewari
ICML4
2024 Globally Convergent Variational Inference
abstract
In variational inference (VI), an approximation of the posterior distribution is selected from a family of distributions through numerical optimization. With the most common variational objective function, known as the evidence lower bound (ELBO), only convergence to a *local* optimum can be guaranteed. In this work, we instead establish the *global* convergence of a particular VI method. This VI method, which may be considered an instance of neural posterior estimation (NPE), minimizes an expectation of the inclusive (forward) KL divergence to fit a variational distribution that is parameterized by a neural network. Our convergence result relies on the neural tangent kernel (NTK) to characterize the gradient dynamics that arise from considering the variational objective in function space. In the asymptotic regime of a fixed, positive-definite neural tangent kernel, we establish conditions under which the variational objective admits a unique solution in a reproducing kernel Hilbert space (RKHS). Then, we show that the gradient descent dynamics in function space converge to this unique function. In ablation studies and practical problems, we demonstrate that our results explain the behavior of NPE in non-asymptotic finite-neuron settings, and show that NPE outperforms ELBO-based optimization, which often converges to shallow local optima.
Declan McNamara, Jackson Loper, Jeffrey Regier
NeurIPS3
2023 Variational Inference for Deblending Crowded Starfields
abstract
In images collected by astronomical surveys, stars and galaxies often overlap visually. Deblending is the task of distinguishing and characterizing individual light sources in survey images. We propose StarNet, a Bayesian method to deblend sources in astronomical images of crowded star fields. StarNet leverages recent advances in variational inference, including amortized variational distributions and an optimization objective targeting an expectation of the forward KL divergence. In our experiments with SDSS images of the M2 globular cluster, StarNet is substantially more accurate than two competing methods: Probabilistic Cataloging (PCAT), a method that uses MCMC for inference, and DAOPHOT, a software pipeline employed by SDSS for deblending. In addition, the amortized approach to inference gives StarNet the scaling characteristics necessary to perform Bayesian inference on modern astronomical surveys.
Runjing Liu, Jon D. McAuliffe, Jeffrey Regier
J. Mach. Learn. Res.3
2022 Normalizing Flows for Knockoff-free Controlled Feature Selection
abstract
Controlled feature selection aims to discover the features a response depends on while limiting the false discovery rate (FDR) to a predefined level. Recently, multiple deep-learning-based methods have been proposed to perform controlled feature selection through the Model-X knockoff framework. We demonstrate, however, that these methods often fail to control the FDR for two reasons. First, these methods often learn inaccurate models of features. Second, the "swap" property, which is required for knockoffs to be valid, is often not well enforced. We propose a new procedure called FlowSelect to perform controlled feature selection that does not suffer from either of these two problems. To more accurately model the features, FlowSelect uses normalizing flows, the state-of-the-art method for density estimation. Instead of enforcing the "swap" property, FlowSelect uses a novel MCMC-based procedure to calculate p-values for each feature directly. Asymptotically, FlowSelect computes valid p-values. Empirically, FlowSelect consistently controls the FDR on both synthetic and semi-synthetic benchmarks, whereas competing knockoff-based approaches do not. FlowSelect also demonstrates greater power on these benchmarks. Additionally, FlowSelect correctly infers the genetic variants associated with specific soybean traits from GWAS data.
Derek Hansen, Brian Manzo, Jeffrey Regier
NeurIPS3
2020 Decision-Making with Auto-Encoding Variational Bayes
abstract
To make decisions based on a model fit with auto-encoding variational Bayes (AEVB), practitioners often let the variational distribution serve as a surrogate for the posterior distribution. This approach yields biased estimates of the expected risk, and therefore leads to poor decisions for two reasons. First, the model fit with AEVB may not equal the underlying data distribution. Second, the variational distribution may not equal the posterior distribution under the fitted model. We explore how fitting the variational distribution based on several objective functions other than the ELBO, while continuing to fit the generative model based on the ELBO, affects the quality of downstream decisions. For the probabilistic principal component analysis model, we investigate how importance sampling error, as well as the bias of the model parameter estimates, varies across several approximate posteriors when used as proposal distributions. Our theoretical results suggest that a posterior approximation distinct from the variational distribution should be used for making decisions. Motivated by these theoretical results, we propose learning several approximate proposals for the best model and combining them using multiple importance sampling for decision-making. In addition to toy examples, we present a full-fledged case study of single-cell RNA sequencing. In this challenging instance of multiple hypothesis testing, our proposed approach surpasses the current state of the art.
Romain Lopez, Pierre Boyeau, Nir Yosef, Michael I. Jordan, Jeffrey Regier
NeurIPS5
2019 Rao-Blackwellized Stochastic Gradients for Discrete Distributions
abstract
We wish to compute the gradient of an expectation over a finite or countably infinite sample space having K $\leq$ $\infty$ categories. When K is indeed infinite, or finite but very large, the relevant summation is intractable. Accordingly, various stochastic gradient estimators have been proposed. In this paper, we describe a technique that can be applied to reduce the variance of any such estimator, without changing its bias{—}in particular, unbiasedness is retained. We show that our technique is an instance of Rao-Blackwellization, and we demonstrate the improvement it yields on a semi-supervised classification problem and a pixel attention task.
Runjing Liu, Jeffrey Regier, Nilesh Tripuraneni, Michael I. Jordan, Jon D. McAuliffe
ICML2
2019 Cataloging the visible universe through Bayesian inference in Julia at petascale
Jeffrey Regier, Keno Fischer, Kiran Pamnany, Andreas Noack 0001, Jarrett Revels, Maximilian Lam, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
J. Parallel Distributed Comput.1
2018 Cataloging the Visible Universe Through Bayesian Inference at Petascale
abstract
Astronomical catalogs derived from wide-field imaging surveys are an important tool for understanding the Universe. We construct an astronomical catalog from 55 TB of imaging data using Celeste, a Bayesian variational inference code written entirely in the high-productivity programming language Julia. Using over 1.3 million threads on 650,000 Intel Xeon Phi cores of the Cori Phase II supercomputer, Celeste achieves a peak rate of 1.54 DP PFLOP/s. Celeste is able to jointly optimize parameters for 188M stars and galaxies, loading and processing 178 TB across 8192 nodes in 14.6 minutes. To achieve this, Celeste exploits parallelism at multiple levels (cluster, node, and thread) and accelerates I/O through Cori's Burst Buffer. Julia's native performance enables Celeste to employ high-level constructs without resorting to hand-written or generated low-level code (C/C++/Fortran), and yet achieve petascale performance.
Jeffrey Regier, Kiran Pamnany, Keno Fischer, Andreas Noack 0001, Maximilian Lam, Jarrett Revels, Steve Howard, Ryan Giordano, David Schlegel, Jon D. McAuliffe, Rollin C. Thomas, Prabhat
IPDPS1
2018 Information Constraints on Auto-Encoding Variational Bayes
abstract
Parameterizing the approximate posterior of a generative model with neural networks has become a common theme in recent machine learning research. While providing appealing flexibility, this approach makes it difficult to impose or assess structural constraints such as conditional independence. We propose a framework for learning representations that relies on Auto-Encoding Variational Bayes and whose search space is constrained via kernel-based measures of independence. In particular, our method employs the $d$-variable Hilbert-Schmidt Independence Criterion (dHSIC) to enforce independence between the latent representations and arbitrary nuisance factors. We show how to apply this method to a range of problems, including the problems of learning invariant representations and the learning of interpretable representations. We also present a full-fledged application to single-cell RNA sequencing (scRNA-seq). In this setting the biological signal in mixed in complex ways with sequencing errors and sampling effects. We show that our method out-performs the state-of-the-art in this domain.
Romain Lopez, Jeffrey Regier, Michael I. Jordan, Nir Yosef
NeurIPS2
2018 Stochastic Cubic Regularization for Fast Nonconvex Optimization
abstract
This paper proposes a stochastic variant of a classic algorithm---the cubic-regularized Newton method [Nesterov and Polyak]. The proposed algorithm efficiently escapes saddle points and finds approximate local minima for general smooth, nonconvex functions in only $\mathcal{\tilde{O}}(\epsilon^{-3.5})$ stochastic gradient and stochastic Hessian-vector product evaluations. The latter can be computed as efficiently as stochastic gradients. This improves upon the $\mathcal{\tilde{O}}(\epsilon^{-4})$ rate of stochastic gradient descent. Our rate matches the best-known result for finding local minima without requiring any delicate acceleration or variance-reduction techniques.
Nilesh Tripuraneni, Mitchell Stern, Chi Jin 0001, Jeffrey Regier, Michael I. Jordan
NeurIPS4
2017 Fast Black-box Variational Inference through Stochastic Trust-Region Optimization
abstract
We introduce TrustVI, a fast second-order algorithm for black-box variational inference based on trust-region optimization and the reparameterization trick. At each iteration, TrustVI proposes and assesses a step based on minibatches of draws from the variational distribution. The algorithm provably converges to a stationary point. We implemented TrustVI in the Stan framework and compared it to two alternatives: Automatic Differentiation Variational Inference (ADVI) and Hessian-free Stochastic Gradient Variational Inference (HFSGVI). The former is based on stochastic first-order optimization. The latter uses second-order information, but lacks convergence guarantees. TrustVI typically converged at least one order of magnitude faster than ADVI, demonstrating the value of stochastic second-order information. TrustVI often found substantially better variational distributions than HFSGVI, demonstrating that our convergence theory can matter in practice.
Jeffrey Regier, Michael I. Jordan, Jon D. McAuliffe
NIPS1
2015 Celeste: Variational inference for a generative model of astronomical images
abstract
We present a new, fully generative model of optical telescope image sets, along with a variational procedure for inference. Each pixel intensity is treated as a Poisson random variable, with a rate parameter dependent on latent properties of stars and galaxies. Key latent properties are themselves random, with scientific prior distributions constructed from large ancillary data sets. We check our approach on synthetic images. We also run it on images from a major sky survey, where it exceeds the performance of the current state-of-the-art method for locating celestial bodies and measuring their colors.
Jeffrey Regier, Andrew C. Miller, Jon D. McAuliffe, Ryan P. Adams, Matthew Hoffman 0001, Dustin Lang, David Schlegel, Prabhat
ICML1
2015 A Gaussian Process Model of Quasar Spectral Energy Distributions
abstract
We propose a method for combining two sources of astronomical data, spectroscopy and photometry, that carry information about sources of light (e.g., stars, galaxies, and quasars) at extremely different spectral resolutions. Our model treats the spectral energy distribution (SED) of the radiation from a source as a latent variable that jointly explains both photometric and spectroscopic observations. We place a flexible, nonparametric prior over the SED of a light source that admits a physically interpretable decomposition, and allows us to tractably perform inference. We use our model to predict the distribution of the redshift of a quasar from five-band (low spectral resolution) photometric data, the so called ``photo-z'' problem. Our method shows that tools from machine learning and Bayesian statistics allow us to leverage multiple resolutions of information to make accurate predictions with well-characterized uncertainties.
Andrew C. Miller, Albert Wu, Jeffrey Regier, Jon D. McAuliffe, Dustin Lang, Prabhat, David Schlegel, Ryan P. Adams
NIPS3