Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

S. T. John

dblp:218/6590 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 63% Representation and self-supervised learning · 15% Motion planning and robot control · 7%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
2.652023
Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models · ICML 2023
Causal Modeling of Policy Interventions From Treatment-Outcome Sequences · ICML 2023
Memory-Based Dual Gaussian Processes for Sequential Learning · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.332023
Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models · ICML 2023
Learning Invariances using the Marginal Likelihood · NeurIPS 2018
Large-Scale Cox Process Inference using Variational Fourier Features · ICML 2018
Machine learning › Representation and self-supervised learning › causal representation learning › identifiability
identifiable representation learning
0.912025
Identifying latent state transitions in non-linear dynamical systems · ICLR 2025
Machine learning › Representation and self-supervised learning › blind source separation › independent component analysis
nonlinear ICA
0.912025
Identifying latent state transitions in non-linear dynamical systems · ICLR 2025
Robotics › Motion planning and robot control › system identification
nonlinear system identification
0.912025
Identifying latent state transitions in non-linear dynamical systems · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.712023
Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.712023
Causal Modeling of Policy Interventions From Treatment-Outcome Sequences · ICML 2023
Machine learning › Optimization for machine learning
hyperparameter optimization
0.712023
Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.712023
Causal Modeling of Policy Interventions From Treatment-Outcome Sequences · ICML 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.712023
Memory-Based Dual Gaussian Processes for Sequential Learning · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › sparse gaussian process
sparse variational gaussian process
0.712023
Memory-Based Dual Gaussian Processes for Sequential Learning · ICML 2023
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
0.512021
Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021
Bioinformatics and computational biology
gene expression analysis
0.512021
Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021
Bioinformatics and computational biology › transcriptomics › spatial transcriptomics
spatially variable gene detection
0.512021
Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics
0.512021
Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
cox process
0.312018
Large-Scale Cox Process Inference using Variational Fourier Features · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning
marginal likelihood
0.312018
Learning Invariances using the Marginal Likelihood · NeurIPS 2018
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random fourier features
0.312018
Large-Scale Cox Process Inference using Variational Fourier Features · ICML 2018
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.212023
Memory-Based Dual Gaussian Processes for Sequential Learning · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
expectation propagation
0.212023
Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models · ICML 2023

Methods — techniques the papers use, named apart from their topics

variational inference · 1.6variational autoencoder · 0.9nonlinear ICA · 0.9point process · 0.7laplace approximation · 0.7gaussian process · 0.7expectation propagation · 0.7dual sparse variational gaussian process · 0.7causal modeling · 0.7variational bayesian inference · 0.5negative binomial likelihood · 0.5gaussian process regression · 0.5MCMC · 0.3
YearPublicationVenuePosition
2025 Identifying latent state transitions in non-linear dynamical systems
abstract
This work aims to recover the underlying states and their time evolution in a latent dynamical system from high-dimensional sensory measurements. Previous works on identifiable representation learning in dynamical systems focused on identifying the latent states, often with linear transition approximations. As such, they cannot identify nonlinear transition dynamics, and hence fail to reliably predict complex future behavior. Inspired by the advances in nonlinear ICA, we propose a state-space modeling framework in which we can identify not just the latent states but also the unknown transition function that maps the past states to the present. Our identifiability theory relies on two key assumptions: (i) sufficient variability in the latent noise, and (ii) the bijectivity of the augmented transition function. Drawing from this theory, we introduce a practical algorithm based on variational auto-encoders. We empirically demonstrate that it improves generalization and interpretability of target dynamical systems by (i) recovering latent state dynamics with high accuracy, (ii) correspondingly achieving high future prediction accuracy, and (iii) adapting fast to new environments. Additionally, for complex real-world dynamics, (iv) it produces state-of the-art future prediction results for long horizons, highlighting its usefulness for practical scenarios.
Caglar Hizli, Çagatay Yildiz, Matthias Bethge, S. T. John, Pekka Marttinen
ICLR4
2024 Learning relevant contextual variables within Bayesian optimization
abstract
Contextual Bayesian Optimization (CBO) efficiently optimizes black-box functions with respect to design variables, while simultaneously integrating _contextual_ information regarding the environment, such as experimental conditions. However, the relevance of contextual variables is not necessarily known beforehand. Moreover, contextual variables can sometimes be optimized themselves at additional cost, a setting overlooked by current CBO algorithms. Cost-sensitive CBO would simply include optimizable contextual variables as part of the design variables based on their cost. Instead, we adaptively select a subset of contextual variables to include in the optimization, based on the trade-off between their _relevance_ and the additional cost incurred by optimizing them compared to leaving them to be determined by the environment. We learn the relevance of contextual variables by sensitivity analysis of the posterior surrogate model while minimizing the cost of optimization by leveraging recent developments on early stopping for BO. We empirically evaluate our proposed Sensitivity-Analysis-Driven Contextual BO (_SADCBO_) method against alternatives on both synthetic and real-world experiments, together with extensive ablation studies, and demonstrate a consistent improvement across examples.
Julien Martinelli, Ayush Bharti, Armi Tiihonen, S. T. John, Louis Filstroff, Sabina Sloman, Patrick Rinke, Samuel Kaski
UAI4
2023 Memory-Based Dual Gaussian Processes for Sequential Learning
abstract
Sequential learning with Gaussian processes (GPs) is challenging when access to past data is limited, for example, in continual and active learning. In such cases, errors can accumulate over time due to inaccuracies in the posterior, hyperparameters, and inducing points, making accurate learning challenging. Here, we present a method to keep all such errors in check using the recently proposed dual sparse variational GP. Our method enables accurate inference for generic likelihoods and improves learning by actively building and updating a memory of past data. We demonstrate its effectiveness in several applications involving Bayesian optimization, active learning, and continual learning.
Paul E. Chang, Prakhar Verma, S. T. John, Arno Solin, Mohammad Emtiyaz Khan
ICML3
2023 Causal Modeling of Policy Interventions From Treatment-Outcome Sequences
abstract
A treatment policy defines when and what treatments are applied to affect some outcome of interest. Data-driven decision-making requires the ability to predict what happens if a policy is changed. Existing methods that predict how the outcome evolves under different scenarios assume that the tentative sequences of future treatments are fixed in advance, while in practice the treatments are determined stochastically by a policy and may depend, for example, on the efficiency of previous treatments. Therefore, the current methods are not applicable if the treatment policy is unknown or a counterfactual analysis is needed. To handle these limitations, we model the treatments and outcomes jointly in continuous time, by combining Gaussian processes and point processes. Our model enables the estimation of a treatment policy from observational sequences of treatments and outcomes, and it can predict the interventional and counterfactual progression of the outcome after an intervention on the treatment policy (in contrast with the causal effect of a single treatment). We show with real-world and semi-synthetic data on blood glucose progression that our method can answer causal queries more accurately than existing alternatives.
Caglar Hizli, S. T. John, Anne Juuti, Tuure Saarinen, Kirsi Pietiläinen, Pekka Marttinen
ICML2
2023 Improving Hyperparameter Learning under Approximate Inference in Gaussian Process Models
abstract
Approximate inference in Gaussian process (GP) models with non-conjugate likelihoods gets entangled with the learning of the model hyperparameters. We improve hyperparameter learning in GP models and focus on the interplay between variational inference (VI) and the learning target. While VI’s lower bound to the marginal likelihood is a suitable objective for inferring the approximate posterior, we show that a direct approximation of the marginal likelihood as in Expectation Propagation (EP) is a better learning objective for hyperparameter optimization. We design a hybrid training procedure to bring the best of both worlds: it leverages conjugate-computation VI for inference and uses an EP-like marginal likelihood approximation for hyperparameter learning. We compare VI, EP, Laplace approximation, and our proposed training procedure and empirically demonstrate the effectiveness of our proposal across a wide range of data sets.
S. T. John, Arno Solin
ICML2
2022 Non-separable Spatio-temporal Graph Kernels via SPDEs
abstract
Gaussian processes (GPs) provide a principled and direct approach for inference and learning on graphs. However, the lack of justified graph kernels for spatio-temporal modelling has held back their use in graph problems. We leverage an explicit link between stochastic partial differential equations (SPDEs) and GPs on graphs, introduce a framework for deriving graph kernels via SPDEs, and derive non-separable spatio-temporal graph kernels that capture interaction across space and time. We formulate the graph kernels for the stochastic heat equation and wave equation. We show that by providing novel tools for spatio-temporal GP modelling on graphs, we outperform pre-existing graph kernels in real-world applications that feature diffusion, oscillation, and other complicated interactions.
Alexander Nikitin 0002, S. T. John, Arno Solin, Samuel Kaski
AISTATS2
2021 Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments
abstract
MOTIVATION: The negative binomial distribution has been shown to be a good model for counts data from both bulk and single-cell RNA-sequencing (RNA-seq). Gaussian process (GP) regression provides a useful non-parametric approach for modelling temporal or spatial changes in gene expression. However, currently available GP regression methods that implement negative binomial likelihood models do not scale to the increasingly large datasets being produced by single-cell and spatial transcriptomics. RESULTS: The GPcounts package implements GP regression methods for modelling counts data using a negative binomial likelihood function. Computational efficiency is achieved through the use of variational Bayesian inference. The GP function models changes in the mean of the negative binomial likelihood through a logarithmic link function and the dispersion parameter is fitted by maximum likelihood. We validate the method on simulated time course data, showing better performance to identify changes in over-dispersed counts data than methods based on Gaussian or Poisson likelihoods. To demonstrate temporal inference, we apply GPcounts to single-cell RNA-seq datasets after pseudotime and branching inference. To demonstrate spatial inference, we apply GPcounts to data from the mouse olfactory bulb to identify spatially variable genes and compare to two published GP methods. We also provide the option of modelling additional dropout using a zero-inflated negative binomial. Our results show that GPcounts can be used to model temporal and spatial counts data in cases where simpler Gaussian and Poisson likelihoods are unrealistic. AVAILABILITY AND IMPLEMENTATION: GPcounts is implemented using the GPflow library in Python and is available at https://github.com/ManchesterBioinference/GPcounts along with the data, code and notebooks required to reproduce the results presented here. The version used for this paper is archived at https://doi.org/10.5281/zenodo.5027066. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nuha Bintayyash, Sokratia Georgaka, S. T. John, Sumon Ahmed, Alexis Boukouvalas, James Hensman, Magnus Rattray
Bioinform.3
2020 Amortized variance reduction for doubly stochastic objective
abstract
Approximate inference in complex probabilistic models such as deep Gaussian processes requires the optimisation of doubly stochastic objective functions. These objectives incorporate randomness both from mini-batch subsampling of the data and from Monte Carlo estimation of expectations. If the gradient variance is high, the stochastic optimisation problem becomes difficult with a slow rate of convergence. Control variates can be used to reduce the variance, but past approaches do not take into account how mini-batch stochasticity affects sampling stochasticity, resulting in sub-optimal variance reduction. We propose a new approach in which we use a recognition network to cheaply approximate the optimal control variate for each mini-batch, with no additional model gradient computations. We illustrate the properties of this proposal and test its performance on logistic regression and deep Gaussian processes.
Ayman Boustati, Sattar Vakili, James Hensman, S. T. John
UAI4
2019 Gaussian Process Modulated Cox Processes under Linear Inequality Constraints
abstract
Gaussian process (GP) modulated Cox processes are widely used to model point patterns. Existing approaches require a mapping (link function) between the unconstrained GP and the positive intensity function. This commonly yields solutions that do not have a closed form or that are restricted to specific covariance functions. We introduce a novel finite approximation of GP-modulated Cox processes where positiveness conditions can be imposed directly on the GP, with no restrictions on the covariance function. Our approach can also ensure other types of inequality constraints (e.g. monotonicity, convexity), resulting in more versatile models that can be used for other classes of point processes (e.g. renewal processes). We demonstrate on both synthetic and real-world data that our framework accurately infers the intensity functions. Where monotonicity is a feature of the process, our ability to include this in the inference improves results.
Andrés F. López-Lopera, S. T. John, Nicolas Durrande
AISTATS2
2018 Large-Scale Cox Process Inference using Variational Fourier Features
abstract
Gaussian process modulated Poisson processes provide a flexible framework for modeling spatiotemporal point patterns. So far this had been restricted to one dimension, binning to a pre-determined grid, or small data sets of up to a few thousand data points. Here we introduce Cox process inference based on Fourier features. This sparse representation induces global rather than local constraints on the function space and is computationally efficient. This allows us to formulate a grid-free approximation that scales well with the number of data points and the size of the domain. We demonstrate that this allows MCMC approximations to the non-Gaussian posterior. In practice, we find that Fourier features have more consistent optimization behavior than previous approaches. Our approximate Bayesian method can fit over 100 000 events with complex spatiotemporal patterns in three dimensions on a single GPU.
S. T. John, James Hensman
ICML1
2018 Learning Invariances using the Marginal Likelihood
abstract
In many supervised learning tasks, learning what changes do not affect the predic-tion target is as crucial to generalisation as learning what does. Data augmentationis a common way to enforce a model to exhibit an invariance: training data is modi-fied according to an invariance designed by a human and added to the training data.We argue that invariances should be incorporated the model structure, and learnedusing themarginal likelihood, which can correctly reward the reduced complexityof invariant models. We incorporate invariances in a Gaussian process, due to goodmarginal likelihood approximations being available for these models. Our maincontribution is a derivation for a variational inference scheme for invariant Gaussianprocesses where the invariance is described by a probability distribution that canbe sampled from, much like how data augmentation is implemented in practice
Mark van der Wilk, Matthias Bauer 0001, S. T. John, James Hensman
NeurIPS3