VLDB 2026 Research / reviewers in the wild / expert
Jon Cockayne
dblp:180/5684 · also Jonathan Cockayne
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-3287-199XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
2 papers |
Mathematical optimization · 87% Algorithms and data structures · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 50% Learning theory · 50% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering › scientific machine learning
surrogate modeling |
0.9 | 1 | 2025 | SMRS: advocating a unified reporting standard for surrogate models in the artificial intelligence era · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
calibration |
0.6 | 1 | 2022 | Testing Whether a Learning Procedure is Calibrated · J. Mach. Learn. Res. 2022 |
Machine learning › Learning theory
hypothesis testing |
0.6 | 1 | 2022 | Testing Whether a Learning Procedure is Calibrated · J. Mach. Learn. Res. 2022 |
Mathematical optimization › numerical computation
iterative linear solver |
0.5 | 1 | 2021 | Probabilistic Iterative Methods for Linear Systems · J. Mach. Learn. Res. 2021 |
Mathematical optimization
uncertainty quantification |
0.5 | 1 | 2021 | Probabilistic Iterative Methods for Linear Systems · J. Mach. Learn. Res. 2021 |
Mathematical optimization
importance sampling |
0.3 | 1 | 2017 | On the Sampling Problem for Kernel Quadrature · ICML 2017 |
Mathematical optimization › numerical analysis › numerical integration › quadrature rules
kernel quadrature |
0.3 | 1 | 2017 | On the Sampling Problem for Kernel Quadrature · ICML 2017 |
Algorithms and data structures › randomized algorithms
monte carlo methods |
0.3 | 1 | 2017 | On the Sampling Problem for Kernel Quadrature · ICML 2017 |
Mathematical optimization › numerical analysis
numerical integration |
0.3 | 1 | 2017 | On the Sampling Problem for Kernel Quadrature · ICML 2017 |
Computational science and engineering
computational reproducibility |
0.3 | 1 | 2025 | SMRS: advocating a unified reporting standard for surrogate models in the artificial intelligence era · NeurIPS 2025 |
Computational science and engineering › scientific data management
data standardization |
0.3 | 1 | 2025 | SMRS: advocating a unified reporting standard for surrogate models in the artificial intelligence era · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
bayesian inference · 1.1hypothesis testing · 0.6probabilistic inference · 0.5sequential monte carlo · 0.3bayesian monte carlo · 0.3adaptive tempering · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrated Computation-Aware Gaussian ProcessesabstractGaussian processes are notorious for scaling cubically with the size of the training set, preventing application to very large regression problems. Computation-aware Gaussian processes (CAGPs) tackle this scaling issue by exploiting probabilistic linear solvers to reduce complexity, widening the posterior with additional \emph{computational} uncertainty due to reduced computation. However, the most commonly used CAGP framework results in (sometimes dramatically) conservative uncertainty quantification, making the posterior difficult to use in practice. In this work, we prove that if the utilised probabilistic linear solver is \emph{calibrated}, in a rigorous statistical sense, then so too is the induced CAGP. We thus propose a new CAGP framework, CAGP-GS, based on using Gauss-Seidel iterations for the underlying probabilistic linear solver. CAGP-GS performs favourably compared to existing approaches when the test set is low-dimensional and few iterations are performed. We test the calibratedness on a synthetic problem, and compare the performance to existing approaches on a large-scale global temperature regression problem. Disha Hegde, Mohamed Adil, Jon Cockayne |
AISTATS | 3 |
| 2025 | Computation-Aware Kalman Filtering and SmoothingabstractKalman filtering and smoothing are the foundational mechanisms for efficient inference in Gauss-Markov models. However, their time and memory complexities scale prohibitively with the size of the state space. This is particularly problematic in spatiotemporal regression problems, where the state dimension scales with the number of spatial observations. Existing approximate frameworks leverage low-rank approximations of the covariance matrix. But since they do not model the error introduced by the computational approximation, their predictive uncertainty estimates can be overly optimistic. In this work, we propose a probabilistic numerical method for inference in high-dimensional Gauss-Markov models which mitigates these scaling issues. Our matrix-free iterative algorithm leverages GPU acceleration and crucially enables a tunable trade-off between computational cost and predictive uncertainty. Finally, we demonstrate the scalability of our method on a large-scale climate dataset. Marvin Pförtner, Jonathan Wenger, Jon Cockayne, Philipp Hennig |
AISTATS | 3 |
| 2025 | SMRS: advocating a unified reporting standard for surrogate models in the artificial intelligence eraabstractSurrogate models are widely used to approximate complex systems across science and engineering to reduce computational costs. Despite their widespread adoption, the field lacks standardisation across key stages of the modelling pipeline, including data sampling, model selection, evaluation, and downstream analysis. This fragmentation limits reproducibility and cross-domain utility – a challengefurther exacerbated by the rapid proliferation of AI-driven surrogate models. We argue for the urgent need to establish a structured reporting standard, the Surrogate Model Reporting Standard (SMRS), that systematically captures essential design and evaluation choices while remaining agnostic to implementation specifics. By promoting a standardised yet flexible framework, we aim to improve the reliabilityof surrogate modelling, foster interdisciplinary knowledge transfer, and, as a result, accelerate scientific progress in the AI era. Elizaveta Semenova, Siobhan Mackenzie Hall, Timothy James Hitge, Alisa Sheinkman, Jon Cockayne |
NeurIPS | 5 |
| 2022 | Testing Whether a Learning Procedure is CalibratedabstractA learning procedure takes as input a dataset and performs inference for the parameters $\theta$ of a model that is assumed to have given rise to the dataset. Here we consider learning procedures whose output is a probability distribution, representing uncertainty about $\theta$ after seeing the dataset. Bayesian inference is a prime example of such a procedure, but one can also construct other learning procedures that return distributional output. This paper studies conditions for a learning procedure to be considered calibrated, in the sense that the true data-generating parameters are plausible as samples from its distributional output. A learning procedure whose inferences and predictions are systematically over- or under-confident will fail to be calibrated. On the other hand, a learning procedure that is calibrated need not be statistically efficient. A hypothesis-testing framework is developed in order to assess, using simulation, whether a learning procedure is calibrated. Several vignettes are presented to illustrate different aspects of the framework. Jon Cockayne, Matthew M. Graham, Chris J. Oates, Timothy John Sullivan, Onur Teymur |
J. Mach. Learn. Res. | 1 |
| 2021 | Probabilistic Iterative Methods for Linear SystemsabstractThis paper presents a probabilistic perspective on iterative methods for approximating the solution $\mathbf{x} \in \mathbb{R}^d$ of a nonsingular linear system $\mathbf{A} \mathbf{x} = \mathbf{b}$. Classically, an iterative method produces a sequence $\mathbf{x}_m$ of approximations that converge to $\mathbf{x}$ in $\mathbb{R}^d$. Our approach, instead, lifts a standard iterative method to act on the set of probability distributions, $\mathcal{P}(\mathbb{R}^d)$, outputting a sequence of probability distributions $\mu_m \in \mathcal{P}(\mathbb{R}^d)$. The output of a probabilistic iterative method can provide both a "best guess" for $\mathbf{x}$, for example by taking the mean of $\mu_m$, and also probabilistic uncertainty quantification for the value of $\mathbf{x}$ when it has not been exactly determined. A comprehensive theoretical treatment is presented in the case of a stationary linear iterative method, where we characterise both the rate of contraction of $\mu_m$ to an atomic measure on $\mathbf{x}$ and the nature of the uncertainty quantification being provided. We conclude with an empirical illustration that highlights the potential for probabilistic iterative methods to provide insight into solution uncertainty. Jon Cockayne, Ilse C. F. Ipsen, Chris J. Oates, Tim W. Reid |
J. Mach. Learn. Res. | 1 |
| 2017 | On the Sampling Problem for Kernel QuadratureabstractThe standard Kernel Quadrature method for numerical integration with random point sets (also called Bayesian Monte Carlo) is known to converge in root mean square error at a rate determined by the ratio s/d, where s and d encode the smoothness and dimension of the integrand. However, an empirical investigation reveals that the rate constant C is highly sensitive to the distribution of the random points. In contrast to standard Monte Carlo integration, for which optimal importance sampling is well-understood, the sampling distribution that minimises C for Kernel Quadrature does not admit a closed form. This paper argues that the practical choice of sampling distribution is an important open problem. One solution is considered; a novel automatic approach based on adaptive tempering and sequential Monte Carlo. Empirical results demonstrate a dramatic reduction in integration error of up to 4 orders of magnitude can be achieved with the proposed method. François-Xavier Briol, Chris J. Oates, Jon Cockayne, Wilson Ye Chen, Mark A. Girolami |
ICML | 3 |