David R. Burt

dblp:238/1347 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Probabilistic and Bayesian machine learning · 75% Learning theory · 14% Trustworthy machine learning · 10%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%
Theoretical computer science
2 papers
Mathematical optimization · 79% Information theory · 21%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
1.942024
Numerically Stable Sparse Gaussian Processes via Minimum Separation using Cover Trees · J. Mach. Learn. Res. 2024
Sparse Gaussian Process Hyperparameters: Optimize or Integrate? · NeurIPS 2022
Rates of Convergence for Sparse Variational Gaussian Process Regression · ICML 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
1.022025
Smooth Sailing: Lipschitz-Driven Uncertainty Quantification for Spatial Associations · NeurIPS 2025
On the Expressiveness of Approximate Inference in Bayesian Neural Networks · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
0.922021
Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients · ICML 2021
Convergence of Sparse Variational Inference in Gaussian Processes Regression · J. Mach. Learn. Res. 2020
Computational science and engineering
spatial statistics
0.912025
Smooth Sailing: Lipschitz-Driven Uncertainty Quantification for Spatial Associations · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › sparse gaussian process
sparse variational gaussian process
0.822020
Convergence of Sparse Variational Inference in Gaussian Processes Regression · J. Mach. Learn. Res. 2020
Rates of Convergence for Sparse Variational Gaussian Process Regression · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.622022
On the Expressiveness of Approximate Inference in Bayesian Neural Networks · NeurIPS 2020
Sparse Gaussian Process Hyperparameters: Optimize or Integrate? · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.612022
Sparse Gaussian Process Hyperparameters: Optimize or Integrate? · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse gaussian process
0.612022
Sparse Gaussian Process Hyperparameters: Optimize or Integrate? · NeurIPS 2022
Machine learning › Learning theory
generalization bounds
0.512021
How Tight Can PAC-Bayes be in the Small Data Regime? · NeurIPS 2021
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds
0.512021
How Tight Can PAC-Bayes be in the Small Data Regime? · NeurIPS 2021
Mathematical optimization › iterative methods
conjugate gradient method
0.512021
Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.412020
On the Expressiveness of Approximate Inference in Bayesian Neural Networks · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412020
On the Expressiveness of Approximate Inference in Bayesian Neural Networks · NeurIPS 2020
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
monte carlo dropout
0.412020
On the Expressiveness of Approximate Inference in Bayesian Neural Networks · NeurIPS 2020
Machine learning › Deep learning architectures and training
small-data regime
0.112021
How Tight Can PAC-Bayes be in the Small Data Regime? · NeurIPS 2021
Information theory › information measures › divergence measures
kullback-leibler divergence
0.112020
Convergence of Sparse Variational Inference in Gaussian Processes Regression · J. Mach. Learn. Res. 2020

Methods — techniques the papers use, named apart from their topics

lipschitz smoothness · 1.7gaussian error model · 1.7helmholtz decomposition · 1.3gaussian process · 1.3variational inference · 1.3minimum separation · 0.8inducing points · 0.8cover tree · 0.8stochastic variational inference · 0.6markov chain monte carlo · 0.6variational method · 0.5lower bound maximization · 0.5conjugate gradient · 0.5matrix operations · 0.4inducing variables · 0.4
YearPublicationVenuePosition
2025 Consistent Validation for Predictive Methods in Spatial Settings
abstract
Spatial prediction tasks are key to weather forecasting, studying air pollution impacts, and other scientific endeavors. Determining how much to trust predictions made by statistical or physical methods is essential for the credibility of scientific conclusions. Unfortunately, classical approaches for validation fail to handle mismatch between locations available for validation and (test) locations where we want to make predictions. This mismatch is often not an instance of covariate shift (as commonly formalized) because the validation and test locations are fixed (e.g., on a grid or at select points) rather than i.i.d. from two distributions. In the present work, we formalize a check on validation methods: that they become arbitrarily accurate as validation data becomes arbitrarily dense. We show that classical and covariate-shift methods can fail this check. We propose a method that builds from existing ideas in the covariate-shift literature, but adapts them to the validation data at hand. We prove that our proposal passes our check. And we demonstrate its advantages empirically on simulated and real data.
David R. Burt, Yunyi Shen, Tamara Broderick
AISTATS1
2025 Smooth Sailing: Lipschitz-Driven Uncertainty Quantification for Spatial Associations
abstract
Estimating associations between spatial covariates and responses — rather than merely predicting responses — is central to environmental science, epidemiology, and economics. For instance, public health officials might be interested in whether air pollution has a strictly positive association with a health outcome, and the magnitude of any effect. Standard machine learning methods often provide accurate predictions but offer limited insight into covariate-response relationships. And we show that existing methods for constructing confidence (or credible) intervals for associations can fail to provide nominal coverage in the face of model misspecification and nonrandom locations — despite both being essentially always present in spatial problems. We introduce a method that constructs valid frequentist confidence intervals for associations in spatial settings. Our method requires minimal assumptions beyond a form of spatial smoothness and a homoskedastic Gaussian error assumption. In particular, we do not require model correctness or covariate overlap between training and target locations. Our approach is the first to guarantee nominal coverage in this setting and outperforms existing techniques in both real and simulated experiments. Our confidence intervals are valid in finite samples when the noise of the Gaussian error is known, and we provide an asymptotically consistent estimation procedure for this noise variance when it is unknown.
David R. Burt, Renato Berlinghieri, Stephen Bates, Tamara Broderick
NeurIPS1
2024 Numerically Stable Sparse Gaussian Processes via Minimum Separation using Cover Trees
abstract
Gaussian processes are frequently deployed as part of larger machine learning and decision-making systems, for instance in geospatial modeling, Bayesian optimization, or in latent Gaussian models. Within a system, the Gaussian process model needs to perform in a stable and reliable manner to ensure it interacts correctly with other parts of the system. In this work, we study the numerical stability of scalable sparse approximations based on inducing points. To do so, we first review numerical stability, and illustrate typical situations in which Gaussian process models can be unstable. Building on stability theory originally developed in the interpolation literature, we derive sufficient and in certain cases necessary conditions on the inducing points for the computations performed to be numerically stable. For low-dimensional tasks such as geospatial modeling, we propose an automated method for computing inducing points satisfying these conditions. This is done via a modification of the cover tree data structure, which is of independent interest. We additionally propose an alternative sparse approximation for regression with a Gaussian likelihood which trades off a small amount of performance to further improve stability. We provide illustrative examples showing the relationship between stability of calculations and predictive performance of inducing point methods on spatial tasks.
Alexander Terenin, David R. Burt, Artem Artemev, Seth R. Flaxman, Mark van der Wilk, Carl E. Rasmussen
J. Mach. Learn. Res.2
2023 Gaussian processes at the Helm(holtz): A more fluid model for ocean currents
abstract
Oceanographers are interested in predicting ocean currents and identifying divergences in a current vector field based on sparse observations of buoy velocities. Since we expect current dynamics to be smooth but highly non-linear, Gaussian processes (GPs) offer an attractive model. But we show that applying a GP with a standard stationary kernel directly to buoy data can struggle at both current prediction and divergence identification – due to some physically unrealistic prior assumptions. To better reflect known physical properties of currents, we propose to instead put a standard stationary kernel on the divergence and curl-free components of a vector field obtained through a Helmholtz decomposition. We show that, because this decomposition relates to the original vector field just via mixed partial derivatives, we can still perform inference given the original data with only a small constant multiple of additional computational expense. We illustrate the benefits of our method on synthetic and real oceans data.
Renato Berlinghieri, Brian L. Trippe, David R. Burt, Ryan Giordano, Kaushik Srinivasan, Tamay M. Özgökmen, Junfei Xia, Tamara Broderick
ICML3
2022 Wide Mean-Field Bayesian Neural Networks Ignore the Data
abstract
Bayesian neural networks (BNNs) combine the expressive power of deep learning with the advantages of Bayesian formalism. In recent years, the analysis of wide, deep BNNs has provided theoretical insight into their priors and posteriors. However, we have no analogous insight into their posteriors under approximate inference. In this work, we show that mean-field variational inference entirely fails to model the data when the network width is large and the activation function is odd. Specifically, for fully-connected BNNs with odd activation functions and a homoscedastic Gaussian likelihood, we show that the optimal mean-field variational posterior predictive (i.e., function space) distribution converges to the prior predictive distribution as the width tends to infinity. We generalize aspects of this result to other likelihoods. Our theoretical results are suggestive of underfitting behavior previously observered in BNNs. While our convergence bounds are non-asymptotic and constants in our analysis can be computed, they are currently too loose to be applicable in standard training regimes. Finally, we show that the optimal approximate posterior need not tend to the prior if the activation function is not odd, showing that our statements cannot be generalized arbitrarily.
Beau Coker, Wessel P. Bruinsma, David R. Burt, Finale Doshi-Velez
AISTATS3
2022 Sparse Gaussian Process Hyperparameters: Optimize or Integrate?
abstract
The kernel function and its hyperparameters are the central model selection choice in a Gaussian process (Rasmussen and Williams, 2006).Typically, the hyperparameters of the kernel are chosen by maximising the marginal likelihood, an approach known as Type-II maximum likelihood (ML-II). However, ML-II does not account for hyperparameter uncertainty, and it is well-known that this can lead to severely biased estimates and an underestimation of predictive uncertainty. While there are several works which employ fully Bayesian characterisation of GPs, relatively few propose such approaches for the sparse GPs paradigm. In this work we propose an algorithm for sparse Gaussian process regression which leverages MCMC to sample from the hyperparameter posterior within the variational inducing point framework of (Titsias, 2009). This work is closely related to (Hensman et al, 2015b) but side-steps the need to sample the inducing points, thereby significantly improving sampling efficiency in the Gaussian likelihood case. We compare this scheme against natural baselines in literature along with stochastic variational GPs (SVGPs) along with an extensive computational analysis.
Vidhi Lalchand, Wessel P. Bruinsma, David R. Burt, Carl E. Rasmussen
NeurIPS3
2021 Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients
abstract
We propose a lower bound on the log marginal likelihood of Gaussian process regression models that can be computed without matrix factorisation of the full kernel matrix. We show that approximate maximum likelihood learning of model parameters by maximising our lower bound retains many benefits of the sparse variational approach while reducing the bias introduced into hyperparameter learning. The basis of our bound is a more careful analysis of the log-determinant term appearing in the log marginal likelihood, as well as using the method of conjugate gradients to derive tight lower bounds on the term involving a quadratic form. Our approach is a step forward in unifying methods relying on lower bound maximisation (e.g. variational methods) and iterative approaches based on conjugate gradients for training Gaussian processes. In experiments, we show improved predictive performance with our model for a comparable amount of training time compared to other conjugate gradient based approaches.
Artem Artemev, David R. Burt, Mark van der Wilk
ICML2
2021 How Tight Can PAC-Bayes be in the Small Data Regime?
abstract
In this paper, we investigate the question: _Given a small number of datapoints, for example $N = 30$, how tight can PAC-Bayes and test set bounds be made?_ For such small datasets, test set bounds adversely affect generalisation performance by withholding data from the training procedure. In this setting, PAC-Bayes bounds are especially attractive, due to their ability to use all the data to simultaneously learn a posterior and bound its generalisation risk. We focus on the case of i.i.d. data with a bounded loss and consider the generic PAC-Bayes theorem of Germain et al. While their theorem is known to recover many existing PAC-Bayes bounds, it is unclear what the tightest bound derivable from their framework is. For a fixed learning algorithm and dataset, we show that the tightest possible bound coincides with a bound considered by Catoni; and, in the more natural case of distributions over datasets, we establish a lower bound on the best bound achievable in expectation. Interestingly, this lower bound recovers the Chernoff test set bound if the posterior is equal to the prior. Moreover, to illustrate how tight these bounds can be, we study synthetic one-dimensional classification tasks in which it is feasible to meta-learn both the prior and the form of the bound to numerically optimise for the tightest bounds possible. We find that in this simple, controlled scenario, PAC-Bayes bounds are competitive with comparable, commonly used Chernoff test set bounds. However, the sharpest test set bounds still lead to better guarantees on the generalisation error than the PAC-Bayes bounds we consider.
Andrew Y. K. Foong, Wessel P. Bruinsma, David R. Burt, Richard E. Turner
NeurIPS3
2020 Bandit optimisation of functions in the Matérn kernel RKHS
abstract
We consider the problem of optimising functions in the reproducing kernel Hilbert space (RKHS) of a Matérn kernel with smoothness parameter $u$ over the domain $[0,1]^d$ under noisy bandit feedback. Our contribution, the $\pi$-GP-UCB algorithm, is the first practical approach with guaranteed sublinear regret for all $u>1$ and $d \geq 1$. Empirical validation suggests better performance and drastically improved computational scalablity compared with its predecessor, Improved GP-UCB.
David Janz, David R. Burt, Javier González 0002
AISTATS2
2020 On the Expressiveness of Approximate Inference in Bayesian Neural Networks
abstract
While Bayesian neural networks (BNNs) hold the promise of being flexible, well-calibrated statistical models, inference often requires approximations whose consequences are poorly understood. We study the quality of common variational methods in approximating the Bayesian predictive distribution. For single-hidden layer ReLU BNNs, we prove a fundamental limitation in function-space of two of the most commonly used distributions defined in weight-space: mean-field Gaussian and Monte Carlo dropout. We find there are simple cases where neither method can have substantially increased uncertainty in between well-separated regions of low uncertainty. We provide strong empirical evidence that exact inference does not have this pathology, hence it is due to the approximation and not the model. In contrast, for deep networks, we prove a universality result showing that there exist approximate posteriors in the above classes which provide flexible uncertainty estimates. However, we find empirically that pathologies of a similar form as in the single-hidden layer case can persist when performing variational inference in deeper networks. Our results motivate careful consideration of the implications of approximate inference methods in BNNs.
Andrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. Turner
NeurIPS2
2020 Convergence of Sparse Variational Inference in Gaussian Processes Regression
abstract
Gaussian processes are distributions over functions that are versatile and mathematically convenient priors in Bayesian modelling. However, their use is often impeded for data with large numbers of observations, $N$, due to the cubic (in $N$) cost of matrix operations used in exact inference. Many solutions have been proposed that rely on $M \ll N$ inducing variables to form an approximation at a cost of $\mathcal{O}\left(NM^2\right)$. While the computational cost appears linear in $N$, the true complexity depends on how $M$ must scale with $N$ to ensure a certain quality of the approximation. In this work, we investigate upper and lower bounds on how $M$ needs to grow with $N$ to ensure high quality approximations. We show that we can make the KL-divergence between the approximate model and the exact posterior arbitrarily small for a Gaussian-noise regression model with $M \ll N$. Specifically, for the popular squared exponential kernel and $D$-dimensional Gaussian distributed covariates, $M = \mathcal{O}((\log N)^D)$ suffice and a method with an overall computational cost of $\mathcal{O}\left(N(\log N)^{2D}(\log \log N)^2\right)$ can be used to perform inference.
David R. Burt, Carl E. Rasmussen, Mark van der Wilk
J. Mach. Learn. Res.1
2019 Rates of Convergence for Sparse Variational Gaussian Process Regression
abstract
Excellent variational approximations to Gaussian process posteriors have been developed which avoid the $\mathcal{O}\left(N^3\right)$ scaling with dataset size $N$. They reduce the computational cost to $\mathcal{O}\left(NM^2\right)$, with $M\ll N$ the number of inducing variables, which summarise the process. While the computational cost seems to be linear in $N$, the true complexity of the algorithm depends on how $M$ must increase to ensure a certain quality of approximation. We show that with high probability the KL divergence can be made arbitrarily small by growing $M$ more slowly than $N$. A particular case is that for regression with normally distributed inputs in D-dimensions with the Squared Exponential kernel, $M=\mathcal{O}(\log^D N)$ suffices. Our results show that as datasets grow, Gaussian process posteriors can be approximated cheaply, and provide a concrete rule for how to increase $M$ in continual learning scenarios.
David R. Burt, Carl E. Rasmussen, Mark van der Wilk
ICML1