Matthew Ashman

dblp:277/1623 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 61% Deep learning architectures and training · 30% Knowledge representation and reasoning · 5%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
neural processes
3.142025
Gridded Transformer Neural Processes for Spatio-Temporal Data · ICML 2025
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Approximately Equivariant Neural Processes · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › neural processes
transformer neural processes
1.622025
Gridded Transformer Neural Processes for Spatio-Temporal Data · ICML 2025
Translation Equivariant Transformer Neural Processes · ICML 2024
Machine learning › Deep learning architectures and training
equivariant neural network
1.522024
Approximately Equivariant Neural Processes · NeurIPS 2024
Translation Equivariant Transformer Neural Processes · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
bayesian causal discovery
0.912025
A Meta-Learning Approach to Bayesian Causal Discovery · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.912025
A Meta-Learning Approach to Bayesian Causal Discovery · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
Gridded Transformer Neural Processes for Spatio-Temporal Data · ICML 2025
Machine learning › Deep learning architectures and training › equivariant neural network
approximate equivariance
0.812024
Approximately Equivariant Neural Processes · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.812024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Machine learning › Deep learning architectures and training › equivariant neural network
shift equivariance
0.812024
Translation Equivariant Transformer Neural Processes · ICML 2024
Privacy and data protection › differential privacy
differentially private learning
0.812024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Privacy and data protection › differential privacy › differentially private learning
differentially private regression
0.812024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Privacy and data protection
differential privacy
0.812024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.712023
Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
latent variable causal discovery
0.712023
Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023
Machine learning › Trustworthy machine learning › calibration
prediction calibration
0.212024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212024
Noise-Aware Differentially Private Regression via Meta-Learning · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

meta-learning · 3.1functional differential privacy mechanism · 1.5convolutional conditional neural process · 1.5gridded pseudo-tokens · 0.9equivariant architectures · 0.9bayesian inference · 0.9transformer · 0.8permutation invariant set functions · 0.8equivariant deep learning · 0.8ADMG learning · 0.7
YearPublicationVenuePosition
2025 A Meta-Learning Approach to Bayesian Causal Discovery
abstract
Discovering a unique causal structure is difficult due to both inherent identifiability issues, and the consequences of finite data. As such, uncertainty over causal structures, such as those obtained from a Bayesian posterior, are often necessary for downstream tasks. Finding an accurate approximation to this posterior is challenging, due to the large number of possible causal graphs, as well as the difficulty in the subproblem of finding posteriors over the functional relationships of the causal edges. Recent works have used Bayesian meta learning to view the problem of posterior estimation as a supervised learning task. Yet, these methods are limited as they cannot reliably sample from the posterior over causal structures and fail to encode key properties of the posterior, such as correlation between edges and permutation equivariance with respect to nodes. To address these limitations, we propose a Bayesian meta learning model that allows for sampling causal structures from the posterior and encodes these key properties. We compare our meta-Bayesian causal discovery against existing Bayesian causal discovery methods, demonstrating the advantages of directly learning a posterior over causal structure.
Anish Dhir, Matthew Ashman, James Requeima, Mark van der Wilk
ICLR2
2025 Gridded Transformer Neural Processes for Spatio-Temporal Data
abstract
Effective modelling of large-scale spatio-temporal datasets is essential for many domains, yet existing approaches often impose rigid constraints on the input data, such as requiring them to lie on fixed-resolution grids. With the rise of foundation models, the ability to process diverse, heterogeneous data structures is becoming increasingly important. Neural processes (NPs), particularly transformer neural processes (TNPs), offer a promising framework for such tasks, but struggle to scale to large spatio-temporal datasets due to the lack of an efficient attention mechanism. To address this, we introduce gridded pseudo-token TNPs which employ specialised encoders and decoders to handle unstructured data and utilise a processor comprising gridded pseudo-tokens with efficient attention mechanisms. Furthermore, we develop equivariant gridded TNPs for applications where exact or approximate translation equivariance is a useful inductive bias, improving accuracy and training efficiency. Our method consistently outperforms a range of strong baselines in various synthetic and real-world regression tasks involving large-scale data, while maintaining competitive computational efficiency. Experiments with weather data highlight the potential of gridded TNPs and serve as just one example of a domain where they can have a significant impact.
Matthew Ashman, Cristiana Diaconu, Eric Langezaal, Adrian Weller, Richard E. Turner
ICML1
2024 Translation Equivariant Transformer Neural Processes
abstract
The effectiveness of neural processes (NPs) in modelling posterior prediction maps—the mapping from data to posterior predictive distributions—has significantly improved since their inception. This improvement can be attributed to two principal factors: (1) advancements in the architecture of permutation invariant set functions, which are intrinsic to all NPs; and (2) leveraging symmetries present in the true posterior predictive map, which are problem dependent. Transformers are a notable development in permutation invariant set functions, and their utility within NPs has been demonstrated through the family of models we refer to as TNPs. Despite significant interest in TNPs, little attention has been given to incorporating symmetries. Notably, the posterior prediction maps for data that are stationary—a common assumption in spatio-temporal modelling—exhibit translation equivariance. In this paper, we introduce of a new family of translation equivariant TNPs that incorporate translation equivariance. Through an extensive range of experiments on synthetic and real-world spatio-temporal data, we demonstrate the effectiveness of TE-TNPs relative to their non-translation-equivariant counterparts and other NP baselines.
Matthew Ashman, Cristiana Diaconu, Junhyuck Kim, Lakee Sivaraya, Stratis Markou, James Requeima, Wessel P. Bruinsma, Richard E. Turner
ICML1
2024 Approximately Equivariant Neural Processes
abstract
Equivariant deep learning architectures exploit symmetries in learning problems to improve the sample efficiency of neural-network-based models and their ability to generalise. However, when modelling real-world data, learning problems are often not *exactly* equivariant, but only approximately. For example, when estimating the global temperature field from weather station observations, local topographical features like mountains break translation equivariance. In these scenarios, it is desirable to construct architectures that can flexibly depart from exact equivariance in a data-driven way. Current approaches to achieving this cannot usually be applied out-of-the-box to any architecture and symmetry group. In this paper, we develop a general approach to achieving this using existing equivariant architectures. Our approach is agnostic to both the choice of symmetry group and model architecture, making it widely applicable. We consider the use of approximately equivariant architectures in neural processes (NPs), a popular family of meta-learning models. We demonstrate the effectiveness of our approach on a number of synthetic and real-world regression experiments, showing that approximately equivariant NP models can outperform both their non-equivariant and strictly equivariant counterparts.
Matthew Ashman, Cristiana Diaconu, Adrian Weller, Wessel P. Bruinsma, Richard E. Turner
NeurIPS1
2024 Noise-Aware Differentially Private Regression via Meta-Learning
abstract
Many high-stakes applications require machine learning models that protect user privacy and provide well-calibrated, accurate predictions. While Differential Privacy (DP) is the gold standard for protecting user privacy, standard DP mechanisms typically significantly impair performance. One approach to mitigating this issue is pre-training models on simulated data before DP learning on the private data. In this work we go a step further, using simulated data to train a meta-learning model that combines the Convolutional Conditional Neural Process (ConvCNP) with an improved functional DP mechanism of Hall et al. (2013), yielding the DPConvCNP. DPConvCNP learns from simulated data how to map private data to a DP predictive model in one forward pass, and then provides accurate, well-calibrated predictions. We compare DPConvCNP with a DP Gaussian Process (GP) baseline with carefully tuned hyperparameters. The DPConvCNP outperforms the GP baseline, especially on non-Gaussian data, yet is much faster at test time and requires less tuning.
Ossi Räisä, Stratis Markou, Matthew Ashman, Wessel P. Bruinsma, Marlon Tobaben, Antti Honkela, Richard E. Turner
NeurIPS3
2023 Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning
Matthew Ashman, Chao Ma 0019, Agrin Hilmkil, Joel Jennings, Cheng Zhang 0005
ICLR1
2021 Scalable Gaussian Process Variational Autoencoders
abstract
Conventional variational autoencoders fail in modeling correlations between data points due to their use of factorized priors. Amortized Gaussian process inference through GP-VAEs has led to significant improvements in this regard, but is still inhibited by the intrinsic complexity of exact GP inference. We improve the scalability of these methods through principled sparse inference approaches. We propose a new scalable GP-VAE model that outperforms existing approaches in terms of runtime and memory footprint, is easy to implement, and allows for joint end-to-end optimization of all components.
Metod Jazbec, Matthew Ashman, Vincent Fortuin, Michael Pearce, Stephan Mandt, Gunnar Rätsch
AISTATS2