Christopher van der Heide

dblp:259/1618 · also Chris van der Heide · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 28% Probabilistic and Bayesian machine learning · 25% Learning theory · 21%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
uncertainty estimation
1.522025
Uncertainty Quantification with the Empirical Neural Tangent Kernel · NeurIPS 2025
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty quantification
0.912025
Uncertainty Quantification with the Empirical Neural Tangent Kernel · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.912025
Uncertainty Quantification with the Empirical Neural Tangent Kernel · NeurIPS 2025
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.912025
Uncertainty Quantification with the Empirical Neural Tangent Kernel · NeurIPS 2025
Machine learning and data management
kernel methods
0.912025
Determinant Estimation under Memory Constraints and Neural Scaling Laws · ICML 2025
Algorithms and data structures
numerical linear algebra
0.912025
Determinant Estimation under Memory Constraints and Neural Scaling Laws · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.822023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks · AAAI 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
empirical bayes
0.712023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Machine learning › Optimization for machine learning
hyperparameter optimization
0.712023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning
marginal likelihood
0.712023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Machine learning › Deep learning architectures and training › overparameterized neural network
infinite-width neural networks
0.512021
Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks · AAAI 2021
Machine learning › Learning theory › neural network theory
neural network kernels
0.512021
Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks · AAAI 2021
Machine learning › Deep learning architectures and training
scaling laws
0.312025
Determinant Estimation under Memory Constraints and Neural Scaling Laws · ICML 2025
Machine learning › Learning theory › over-parameterization
double descent
0.212023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023
Machine learning › Learning theory
generalization
0.212023
Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes · ICML 2023

Methods — techniques the papers use, named apart from their topics

pseudo-determinant · 2.6LDL decomposition · 2.6block-wise computation · 1.7gradient-descent sampling · 0.9gaussian process posterior approximation · 0.9blockwise computation · 0.9cross-validation · 0.7bayesian inference · 0.7fixed-point dynamics analysis · 0.5covariance function derivation · 0.5
YearPublicationVenuePosition
2025 Determinant Estimation under Memory Constraints and Neural Scaling Laws
abstract
Calculating or accurately estimating log-determinants of large positive definite matrices is of fundamental importance in many machine learning tasks. While its cubic computational complexity can already be prohibitive, in modern applications, even storing the matrices themselves can pose a memory bottleneck. To address this, we derive a novel hierarchical algorithm based on block-wise computation of the LDL decomposition for large-scale log-determinant calculation in memory-constrained settings. In extreme cases where matrices are highly ill-conditioned, accurately computing the full matrix itself may be infeasible. This is particularly relevant when considering kernel matrices at scale, including the empirical Neural Tangent Kernel (NTK) of neural networks trained on large datasets. Under the assumption of neural scaling laws in the test error, we show that the ratio of pseudo-determinants satisfies a power-law relationship, allowing us to derive corresponding scaling laws. This enables accurate estimation of NTK log-determinants from a tiny fraction of the full dataset; in our experiments, this results in a $\sim$100,000$\times$ speedup with improved accuracy over competing approximations. Using these techniques, we successfully estimate log-determinants for dense matrices of extreme sizes, which were previously deemed intractable and inaccessible due to their enormous scale and computational demands.
Siavash Ameli, Christopher van der Heide, Liam Hodgkinson, Fred (Farbod) Roosta, Michael W. Mahoney
ICML2
2025 Spectral Estimation with Free Decompression
abstract
Computing eigenvalues of very large matrices is a critical task in many machine learning applications, including the evaluation of log-determinants, the trace of matrix functions, and other important metrics. As datasets continue to grow in scale, the corresponding covariance and kernel matrices become increasingly large, often reaching magnitudes that make their direct formation impractical or impossible. Existing techniques typically rely on matrix-vector products, which can provide efficient approximations, if the matrix spectrum behaves well. However, in settings like distributed learning, or when the matrix is defined only indirectly, access to the full data set can be restricted to only very small sub-matrices of the original matrix. In these cases, the matrix of nominal interest is not even available as an implicit operator, meaning that even matrix-vector products may not be available. In such settings, the matrix is "impalpable", in the sense that we have access to only masked snapshots of it. We draw on principles from free probability theory to introduce a novel method of "free decompression" to estimate the spectrum of such matrices. Our method can be used to extrapolate from the empirical spectral densities of small submatrices to infer the eigenspectrum of extremely large (impalpable) matrices (that we cannot form or even evaluate with full matrix-vector products). We demonstrate the effectiveness of this approach through a series of examples, comparing its performance against known limiting distributions from random matrix theory in synthetic settings, as well as applying it to submatrices of real-world datasets, matching them with their full empirical eigenspectra.
Siavash Ameli, Christopher van der Heide, Liam Hodgkinson, Michael W. Mahoney
NeurIPS2
2025 Uncertainty Quantification with the Empirical Neural Tangent Kernel
abstract
While neural networks have demonstrated impressive performance across various tasks, accurately quantifying uncertainty in their predictions is essential to ensure their trustworthiness and enable widespread adoption in critical systems. Several Bayesian uncertainty quantification (UQ) methods exist that are either cheap or reliable, but not both. We propose a post-hoc, sampling-based UQ method for overparameterized networks at the end of training. Our approach constructs efficient and meaningful deep ensembles by employing a (stochastic) gradient-descent sampling process on appropriately linearized networks. We demonstrate that our method effectively approximates the posterior of a Gaussian Process using the empirical Neural Tangent Kernel. Through a series of numerical experiments, we show that our method not only outperforms competing approaches in computational efficiency--often reducing costs by multiple factors--but also maintains state-of-the-art performance across a variety of UQ metrics for both regression and classification tasks.
Joseph Wilson, Christopher van der Heide, Liam Hodgkinson, Fred (Farbod) Roosta
NeurIPS2
2025 Temperature Optimization for Bayesian Deep Learning
abstract
The Cold Posterior Effect (CPE) is a phenomenon in Bayesian Deep Learning (BDL), where tempering the posterior to a cold temperature often improves the predictive performance of the posterior predictive distribution (PPD). Although the term ‘CPE’ suggests colder temperatures are inherently better, the BDL community increasingly recognizes that this is not always the case. Despite this, there remains no systematic method for finding the optimal temperature beyond grid search. In this work, we propose a data-driven approach to select the temperature that maximizes test log-predictive density, treating the temperature as a model parameter and estimating it directly from the data. We empirically demonstrate that our method performs comparably to grid search, at a fraction of the cost, across both regression and classification tasks. Finally, we highlight the differing perspectives on CPE between the BDL and Generalized Bayes communities: while the former primarily emphasizes the predictive performance of the PPD, the latter prioritizes the utility of the posterior under model misspecification; these distinct objectives lead to different temperature preferences.
Kenyon Ng, Christopher van der Heide, Liam Hodgkinson, Susan Wei
UAI2
2023 Monotonicity and Double Descent in Uncertainty Estimation with Gaussian Processes
abstract
Despite their importance for assessing reliability of predictions, uncertainty quantification (UQ) measures in machine learning models have only recently begun to be rigorously characterized. One prominent issue is the curse of dimensionality: it is commonly believed that the marginal likelihood should be reminiscent of cross-validation metrics and both should deteriorate with larger input dimensions. However, we prove that by tuning hyperparameters to maximize marginal likelihood (the empirical Bayes procedure), performance, as measured by the marginal likelihood, improves monotonically with the input dimension. On the other hand, cross-validation metrics exhibit qualitatively different behavior that is characteristic of double descent. Cold posteriors, which have recently attracted interest due to their improved performance in certain settings, appear to exacerbate these phenomena. We verify empirically that our results hold for real data, beyond our considered assumptions, and we explore consequences involving synthetic covariates.
Liam Hodgkinson, Christopher van der Heide, Fred (Farbod) Roosta, Michael W. Mahoney
ICML2
2021 Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite Networks
abstract
Analysing and computing with Gaussian processes arising from infinitely wide neural networks has recently seen a resurgence in popularity. Despite this, many explicit covariance functions of networks with activation functions used in modern networks remain unknown. Furthermore, while the kernels of deep networks can be computed iteratively, theoretical understanding of deep kernels is lacking, particularly with respect to fixed-point dynamics. Firstly, we derive the covariance functions of multi-layer perceptrons (MLPs) with exponential linear units (ELU) and Gaussian error linear units (GELU) and evaluate the performance of the limiting Gaussian processes on some benchmarks. Secondly, and more generally, we analyse the fixed-point dynamics of iterated kernels corresponding to a broad range of activation functions. We find that unlike some previously studied neural network kernels, these new kernels exhibit non-trivial fixed-point dynamics which are mirrored in finite-width neural networks. The fixed point behaviour present in some networks explains a mechanism for implicit regularisation in overparameterised deep models. Our results relate to both the static iid parameter conjugate kernel and the dynamic neural tangent kernel constructions
Russell Tsuchida, Tim Pearce, Christopher van der Heide, Fred (Farbod) Roosta, Marcus Gallagher
AAAI3
2021 Shadow Manifold Hamiltonian Monte Carlo
abstract
Hamiltonian Monte Carlo and its descendants have found success in machine learning and computational statistics due to their ability to draw samples in high dimensions with greater efficiency than classical MCMC. One of these derivatives, Riemannian manifold Hamiltonian Monte Carlo (RMHMC), better adapts the sampler to the geometry of the target density, allowing for improved performances in sampling problems with complex geometric features. Other approaches have boosted acceptance rates by sampling from an integrator-dependent “shadow density” and compensating for the induced bias via importance sampling. We combine the benefits of RMHMC with those attained by sampling from the shadow density, by deriving the shadow Hamiltonian corresponding to the generalized leapfrog integrator used in RMHMC. This leads to a new algorithm, shadow manifold Hamiltonian Monte Carlo, that shows improved performance over RMHMC, and leaves the target density invariant.
Christopher van der Heide, Fred (Farbod) Roosta, Liam Hodgkinson, Dirk P. Kroese
AISTATS1
2021 Stochastic continuous normalizing flows: training SDEs as ODEs
abstract
We provide a general theoretical framework for stochastic continuous normalizing flows, an extension of continuous normalizing flows for density estimation of stochastic differential equations (SDEs). Using the theory of rough paths, the underlying Brownian motion is treated as a latent variable and approximated. Doing so enables the treatment of SDEs as random ordinary differential equations, which can be trained using existing techniques. For scalar loss functions, this approach naturally recovers the stochastic adjoint method of Li et al. [2020] for training neural SDEs, while supporting a more flexible class of approximations.
Liam Hodgkinson, Christopher van der Heide, Fred (Farbod) Roosta, Michael W. Mahoney
UAI2