VLDB 2026 Research / reviewers in the wild / expert
James Hensman
dblp:116/2940
· DBLP profile ↗
36ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0002-4989-3589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
22 papers |
Probabilistic and Bayesian machine learning · 61% Efficient and distributed learning · 14% Language models and text generation · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 52, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
5.2 | 15 | 2022 | Additive Gaussian Processes Revisited · ICML 2022 Deep Neural Networks as Point Estimates for Deep Gaussian Processes · NeurIPS 2021 Sparse Gaussian Processes with Spherical Harmonic Features · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
3.0 | 10 | 2019 | Deep Gaussian Processes with Importance-Weighted Variational Inference · ICML 2019 Overcoming Mean-Field Approximations in Recurrent Gaussian Process Models · ICML 2019 Learning Invariances using the Marginal Likelihood · NeurIPS 2018 |
Machine learning › Efficient and distributed learning
model compression |
1.5 | 2 | 2024 | QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs · NeurIPS 2024 SliceGPT: Compress Large Language Models by Deleting Rows and Columns · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › hierarchical gaussian process
deep gaussian process |
0.9 | 2 | 2021 | Deep Neural Networks as Point Estimates for Deep Gaussian Processes · NeurIPS 2021 Deep Gaussian Processes with Importance-Weighted Variational Inference · ICML 2019 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge-augmented language models |
0.9 | 1 | 2025 | KBLaM: Knowledge Base augmented Language Model · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition
knowledge base enrichment |
0.9 | 1 | 2025 | KBLaM: Knowledge Base augmented Language Model · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
knowledge base integration |
0.9 | 1 | 2025 | KBLaM: Knowledge Base augmented Language Model · ICLR 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | KBLaM: Knowledge Base augmented Language Model · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.8 | 1 | 2024 | SliceGPT: Compress Large Language Models by Deleting Rows and Columns · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression › sparsity
post-training sparsity |
0.8 | 1 | 2024 | SliceGPT: Compress Large Language Models by Deleting Rows and Columns · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.8 | 1 | 2024 | QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.8 | 1 | 2024 | SliceGPT: Compress Large Language Models by Deleting Rows and Columns · ICLR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › sparse gaussian process
sparse variational gaussian process |
0.7 | 2 | 2020 | Sparse Gaussian Processes with Spherical Harmonic Features · ICML 2020 MCMC for Variationally Sparse Gaussian Processes · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.6 | 2 | 2019 | Pseudo-Extended Markov chain Monte Carlo · NeurIPS 2019 MCMC for Variationally Sparse Gaussian Processes · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › gaussian process regression
additive gaussian processes |
0.6 | 1 | 2022 | Additive Gaussian Processes Revisited · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
kernel design |
0.6 | 1 | 2022 | Additive Gaussian Processes Revisited · ICML 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
neural network uncertainty |
0.5 | 1 | 2021 | Deep Neural Networks as Point Estimates for Deep Gaussian Processes · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › uncertainty estimation
predictive uncertainty |
0.5 | 1 | 2021 | Deep Neural Networks as Point Estimates for Deep Gaussian Processes · NeurIPS 2021 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.5 | 1 | 2021 | Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021 |
Bioinformatics and computational biology
gene expression analysis |
0.5 | 1 | 2021 | Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021 |
Bioinformatics and computational biology › transcriptomics › spatial transcriptomics
spatially variable gene detection |
0.5 | 1 | 2021 | Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021 |
Bioinformatics and computational biology › transcriptomics
spatial transcriptomics |
0.5 | 1 | 2021 | Non-parametric modelling of temporal and spatial counts data from RNA-seq experiments · Bioinform. 2021 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
importance weighted variational inference |
0.4 | 1 | 2019 | Deep Gaussian Processes with Importance-Weighted Variational Inference · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation |
0.4 | 1 | 2019 | Overcoming Mean-Field Approximations in Recurrent Gaussian Process Models · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › sampling › posterior sampling
multimodal posterior sampling |
0.4 | 1 | 2019 | Pseudo-Extended Markov chain Monte Carlo · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
collapsed variational inference |
0.4 | 2 | 2015 | Fast Nonparametric Clustering of Structured Time-Series · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Fast Variational Inference in the Conjugate Exponential Family · NIPS 2012 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference |
0.3 | 1 | 2018 | Infinite-Horizon Gaussian Processes · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
0.3 | 1 | 2018 | Gaussian Process Conditional Density Estimation · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation |
0.3 | 1 | 2018 | Gaussian Process Conditional Density Estimation · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
cox process |
0.3 | 1 | 2018 | Large-Scale Cox Process Inference using Variational Fourier Features · ICML 2018 |
Methods — techniques the papers use, named apart from their topics
variational inference · 2.2sparse gaussian process · 1.2gaussian process · 1.1sentence encoder · 0.9attention mechanism · 0.9transformer architecture · 0.8multistage extraction · 0.8language model · 0.8hadamard rotation · 0.8computational invariance · 0.8variational bayesian inference · 0.5negative binomial likelihood · 0.5gaussian process regression · 0.5spherical harmonic features · 0.4fourier features · 0.4collapsed variational inference · 0.4variational bayes · 0.2markov chain monte carlo · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KBLaM: Knowledge Base augmented Language ModelabstractIn this paper, we propose Knowledge Base augmented Language Model (KBLAM), a new method for augmenting Large Language Models (LLMs) with external knowledge. KBLAM works with a knowledge base (KB) constructed from a corpus of documents, transforming each piece of knowledge in the KB into continuous key-value vector pairs via pre-trained sentence encoders with linear adapters and
integrating them into pre-trained LLMs via a specialized rectangular attention mechanism. Unlike Retrieval-Augmented Generation, KBLAM eliminates external retrieval modules, and unlike in-context learning, its computational overhead scales linearly with KB size rather than quadratically. Our approach enables integrating a large KB of more than 10K triples into an 8B pre-trained LLM of only 8K context window on one single A100 80GB GPU and allows for dynamic updates without model fine-tuning or retraining. Experiments demonstrate KBLAM’s effectiveness in various tasks, including question-answering and open-ended reasoning, while providing interpretable insights into its use of the augmented knowledge. Code and datasets are available at https://github.com/microsoft/KBLaM/ Taketomo Isazawa, Liana Mikaelyan, James Hensman |
ICLR | 4 |
| 2024 | Learning to Extract Structured Entities Using Language ModelsabstractRecent advances in machine learning have significantly impacted the field of information extraction, with Language Models (LMs) playing a pivotal role in extracting structured information from unstructured text.Prior works typically represent information extraction as triplet-centric and use classical metrics such as precision and recall for evaluation.We reformulate the task to be entity-centric, enabling the use of diverse metrics that can provide more insights from various perspectives.We contribute to the field by introducing Structured Entity Extraction and proposing the Approximate Entity Set OverlaP (AESOP) metric, designed to appropriately assess model performance.Later, we introduce a new Multistage Structured Entity Extraction (MuSEE) model that harnesses the power of LMs for enhanced effectiveness and efficiency by decomposing the extraction task into multiple stages.Quantitative and human side-by-side evaluations confirm that our model outperforms baselines, offering promising directions for future advancements in structured entity extraction.Our source code is available at https://github.com/microsoft/Structured- Entity-Extraction. Haolun Wu, Ye Yuan 0017, Liana Mikaelyan, Alexander Meulemans, Xue Liu 0004, James Hensman, Bhaskar Mitra 0001 |
EMNLP | 6 |
| 2024 | SliceGPT: Compress Large Language Models by Deleting Rows and ColumnsabstractLarge language models have become the cornerstone of natural language processing, but their use comes with substantial costs in terms of compute and memory resources. Sparsification provides a solution to alleviate these resource constraints, and recent works have shown that trained models can be sparsified post-hoc. Existing sparsification techniques face challenges as they need additional data structures and offer constrained speedup with current hardware. In this paper we present SliceGPT, a new post-training sparsification scheme which replaces each weight matrix with a smaller (dense) matrix, reducing the embedding dimension of the network. Through extensive experimentation we show that SliceGPT can remove up to 25% of the model parameters (including embeddings) for LLAMA-2 70B, OPT 66B and Phi-2 models while maintaining 99%, 99% and 90% zero-shot task performance of the dense model respectively. Our sliced models run on fewer GPUs and run faster without any additional code optimization: on 24GB consumer GPUs we reduce the total compute for inference on LLAMA-2 70B to 64% of that of the dense model; on 40GB A100 GPUs we reduce it to 66%. We offer a new insight, computational invariance in transformer networks, which enables SliceGPT and we hope it will inspire and enable future avenues to reduce memory and computation demands for pre-trained models. Saleh Ashkboos, Maximilian L. Croci, Marcelo Gennari, Torsten Hoefler, James Hensman |
ICLR | 5 |
| 2024 | QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsabstractWe introduce QuaRot, a new Quantization scheme based on Rotations, which is able to quantize LLMs end-to-end, including all weights, activations, and KV cache in 4 bits. QuaRot rotates LLMs in a way that removes outliers from the hidden state without changing the output, making quantization easier. This computational invariance is applied to the hidden state (residual) of the LLM, as well as to the activations of the feed-forward components, aspects of the attention mechanism, and to the KV cache. The result is a quantized model where all matrix multiplications are performed in 4 bits, without any channels identified for retention in higher precision. Our 4-bit quantized LLAMA2-70B model has losses of at most 0.47 WikiText-2 perplexity and retains 99% of the zero-shot performance. We also show that QuaRot can provide lossless 6 and 8 bit LLAMA-2 models without any calibration data using round-to-nearest quantization. Code is available at github.com/spcl/QuaRot. Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, James Hensman |
NeurIPS | 9 |
| 2022 | Additive Gaussian Processes RevisitedabstractGaussian Process (GP) models are a class of flexible non-parametric models that have rich representational power. By using a Gaussian process with additive structure, complex responses can be modelled whilst retaining interpretability. Previous work showed that additive Gaussian process models require high-dimensional interaction terms. We propose the orthogonal additive kernel (OAK), which imposes an orthogonality constraint on the additive functions, enabling an identifiable, low-dimensional representation of the functional relationship. We connect the OAK kernel to functional ANOVA decomposition, and show improved convergence rates for sparse computation methods. With only a small number of additive low-dimensional terms, we demonstrate the OAK model achieves similar or better predictive performance compared to black-box models, while retaining interpretability. Alexis Boukouvalas, James Hensman |
ICML | 3 |
| 2021 | Deep Neural Networks as Point Estimates for Deep Gaussian ProcessesabstractNeural networks and Gaussian processes are complementary in their strengths and weaknesses. Having a better understanding of their relationship comes with the promise to make each method benefit from the strengths of the other. In this work, we establish an equivalence between the forward passes of neural networks and (deep) sparse Gaussian process models. The theory we develop is based on interpreting activation functions as interdomain inducing features through a rigorous analysis of the interplay between activation functions and kernels. This results in models that can either be seen as neural networks with improved uncertainty prediction or deep Gaussian processes with increased prediction accuracy. These claims are supported by experimental results on regression and classification datasets. Vincent Dutordoir, James Hensman, Mark van der Wilk, Carl Henrik Ek, Zoubin Ghahramani, Nicolas Durrande |
NeurIPS | 2 |
| 2021 | Non-parametric modelling of temporal and spatial counts data from RNA-seq experimentsabstractMOTIVATION: The negative binomial distribution has been shown to be a good model for counts data from both bulk and single-cell RNA-sequencing (RNA-seq). Gaussian process (GP) regression provides a useful non-parametric approach for modelling temporal or spatial changes in gene expression. However, currently available GP regression methods that implement negative binomial likelihood models do not scale to the increasingly large datasets being produced by single-cell and spatial transcriptomics. RESULTS: The GPcounts package implements GP regression methods for modelling counts data using a negative binomial likelihood function. Computational efficiency is achieved through the use of variational Bayesian inference. The GP function models changes in the mean of the negative binomial likelihood through a logarithmic link function and the dispersion parameter is fitted by maximum likelihood. We validate the method on simulated time course data, showing better performance to identify changes in over-dispersed counts data than methods based on Gaussian or Poisson likelihoods. To demonstrate temporal inference, we apply GPcounts to single-cell RNA-seq datasets after pseudotime and branching inference. To demonstrate spatial inference, we apply GPcounts to data from the mouse olfactory bulb to identify spatially variable genes and compare to two published GP methods. We also provide the option of modelling additional dropout using a zero-inflated negative binomial. Our results show that GPcounts can be used to model temporal and spatial counts data in cases where simpler Gaussian and Poisson likelihoods are unrealistic. AVAILABILITY AND IMPLEMENTATION: GPcounts is implemented using the GPflow library in Python and is available at https://github.com/ManchesterBioinference/GPcounts along with the data, code and notebooks required to reproduce the results presented here. The version used for this paper is archived at https://doi.org/10.5281/zenodo.5027066. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nuha Bintayyash, Sokratia Georgaka, S. T. John, Sumon Ahmed, Alexis Boukouvalas, James Hensman, Magnus Rattray |
Bioinform. | 6 |
| 2020 | Doubly Sparse Variational Gaussian ProcessesabstractThe use of Gaussian process models is typically limited to datasets with a few tens of thousands of observations due to their complexity and memory footprint.The two most commonly used methods to overcome this limitation are 1) the variational sparse approximation which relies on inducing points and 2) the state-space equivalent formulation of Gaussian processes which can be seen as exploiting some sparsity in the precision matrix.In this work, we propose to take the best of both worlds: we show that the inducing point framework is still valid for state space models and that it can bring further computational and memory savings. Furthermore, we provide the natural gradient formulation for the proposed variational parameterisation.Finally, this work makes it possible to use the state-space formulation inside deep Gaussian process models as illustrated in one of the experiments. Vincent Adam, Stefanos Eleftheriadis, Artem Artemev, Nicolas Durrande, James Hensman |
AISTATS | 5 |
| 2020 | Bayesian Image Classification with Deep Convolutional Gaussian ProcessesabstractIn decision-making systems, it is important to have classifiers that have calibrated uncertainties, with an optimisation objective that can be used for automated model selection and training. Gaussian processes (GPs) provide uncertainty estimates and a marginal likelihood objective, but their weak inductive biases lead to inferior accuracy. This has limited their applicability in certain tasks (e.g. image classification). We propose a translation insensitive convolutional kernel, which relaxes the translation invariance constraint imposed by previous convolutional GPs. We show how we can use the marginal likelihood to learn the degree of insensitivity. We also reformulate GP image-to-image convolutional mappings as multi-output GPs, leading to deep convolutional GPs. We show experimentally that our new kernel improves performance in both single-layer and deep models. We also demonstrate that our fully Bayesian approach improves on dropout-based Bayesian deep learning methods in terms of uncertainty and marginal likelihood estimates. Vincent Dutordoir, Mark van der Wilk, Artem Artemev, James Hensman |
AISTATS | 4 |
| 2020 | Sparse Gaussian Processes with Spherical Harmonic FeaturesabstractWe introduce a new class of inter-domain variational Gaussian processes (GP) where data is mapped onto the unit hypersphere in order to use spherical harmonic representations. Our inference scheme is comparable to variational Fourier features, but it does not suffer from the curse of dimensionality, and leads to diagonal covariance matrices between inducing variables. This enables a speed-up in inference, because it bypasses the need to invert large covariance matrices. Our experiments show that our model is able to fit a regression model for a dataset with 6 million entries two orders of magnitude faster compared to standard sparse GPs, while retaining state of the art accuracy. We also demonstrate competitive performance on classification with non-conjugate likelihoods. Vincent Dutordoir, Nicolas Durrande, James Hensman |
ICML | 3 |
| 2020 | Amortized variance reduction for doubly stochastic objectiveabstractApproximate inference in complex probabilistic models such as deep Gaussian processes requires the optimisation of doubly stochastic objective functions. These objectives incorporate randomness both from mini-batch subsampling of the data and from Monte Carlo estimation of expectations. If the gradient variance is high, the stochastic optimisation problem becomes difficult with a slow rate of convergence. Control variates can be used to reduce the variance, but past approaches do not take into account how mini-batch stochasticity affects sampling stochasticity, resulting in sub-optimal variance reduction. We propose a new approach in which we use a recognition network to cheaply approximate the optimal control variate for each mini-batch, with no additional model gradient computations. We illustrate the properties of this proposal and test its performance on logistic regression and deep Gaussian processes. Ayman Boustati, Sattar Vakili, James Hensman, S. T. John |
UAI | 3 |
| 2019 | Banded Matrix Operators for Gaussian Markov Models in the Automatic Differentiation EraabstractBanded matrices can be used as precision matrices in several models including linear state-space models, some Gaussian processes, and Gaussian Markov random fields. The aim of the paper is to make modern inference methods (such as variational inference or gradient-based sampling) available for Gaussian models with banded precision. We show that this can efficiently be achieved by equipping an automatic differentiation framework, such as TensorFlow or PyTorch, with some linear algebra operators dedicated to banded matrices. This paper studies the algorithmic aspects of the required operators, details their reverse-mode derivatives, and show that their complexity is linear in the number of observations. Nicolas Durrande, Vincent Adam, Lucas Bordeaux, Stefanos Eleftheriadis, James Hensman |
AISTATS | 5 |
| 2019 | Overcoming Mean-Field Approximations in Recurrent Gaussian Process ModelsabstractWe identify a new variational inference scheme for dynamical systems whose transition function is modelled by a Gaussian process. Inference in this setting has either employed computationally intensive MCMC methods, or relied on factorisations of the variational posterior. As we demonstrate in our experiments, the factorisation between latent system states and transition function can lead to a miscalibrated posterior and to learning unnecessarily large noise terms. We eliminate this factorisation by explicitly modelling the dependence between state trajectories and the low-rank representation of our Gaussian process posterior. Samples of the latent states can then be tractably generated by conditioning on this representation. The method we obtain gives better predictive performance and more calibrated estimates of the transition function, yet maintains the same time and space complexities as mean-field methods. Alessandro Davide Ialongo, Mark van der Wilk, James Hensman, Carl E. Rasmussen |
ICML | 3 |
| 2019 | Deep Gaussian Processes with Importance-Weighted Variational InferenceabstractDeep Gaussian processes (DGPs) can model complex marginal densities as well as complex mappings. Non-Gaussian marginals are essential for modelling real-world data, and can be generated from the DGP by incorporating uncorrelated variables to the model. Previous work in the DGP model has introduced noise additively, and used variational inference with a combination of sparse Gaussian processes and mean-field Gaussians for the approximate posterior. Additive noise attenuates the signal, and the Gaussian form of variational distribution may lead to an inaccurate posterior. We instead incorporate noisy variables as latent covariates, and propose a novel importance-weighted objective, which leverages analytic results and provides a mechanism to trade off computation for improved accuracy. Our results demonstrate that the importance-weighted objective works well in practice and consistently outperforms classical variational inference, especially for deeper models. Hugh Salimbeni, Vincent Dutordoir, James Hensman, Marc Peter Deisenroth |
ICML | 3 |
| 2019 | Pseudo-Extended Markov chain Monte CarloabstractSampling from posterior distributions using Markov chain Monte Carlo (MCMC) methods can require an exhaustive number of iterations, particularly when the posterior is multi-modal as the MCMC sampler can become trapped in a local mode for a large number of iterations. In this paper, we introduce the pseudo-extended MCMC method as a simple approach for improving the mixing of the MCMC sampler for multi-modal posterior distributions. The pseudo-extended method augments the state-space of the posterior using pseudo-samples as auxiliary variables. On the extended space, the modes of the posterior are connected, which allows the MCMC sampler to easily move between well-separated posterior modes. We demonstrate that the pseudo-extended approach delivers improved MCMC sampling over the Hamiltonian Monte Carlo algorithm on multi-modal posteriors, including Boltzmann machines and models with sparsity-inducing priors. Christopher Nemeth, Fredrik Lindsten, Maurizio Filippone, James Hensman |
NeurIPS | 4 |
| 2018 | Natural Gradients in Practice: Non-Conjugate Variational Inference in Gaussian Process ModelsabstractThe natural gradient method has been used effectively in conjugate Gaussian process models, but the non-conjugate case has been largely unexplored. We examine how natural gradients can be used in non-conjugate stochastic settings, together with hyperparameter learning. We conclude that the natural gradient can significantly improve performance in terms of wall-clock time. For ill-conditioned posteriors the benefit of the natural gradient method is especially pronounced, and we demonstrate a practical setting where ordinary gradients are unusable. We show how natural gradients can be computed efficiently and automatically in any parameterization, using automatic differentiation. Hugh Salimbeni, Stefanos Eleftheriadis, James Hensman |
AISTATS | 3 |
| 2018 | Large-Scale Cox Process Inference using Variational Fourier FeaturesabstractGaussian process modulated Poisson processes provide a flexible framework for modeling spatiotemporal point patterns. So far this had been restricted to one dimension, binning to a pre-determined grid, or small data sets of up to a few thousand data points. Here we introduce Cox process inference based on Fourier features. This sparse representation induces global rather than local constraints on the function space and is computationally efficient. This allows us to formulate a grid-free approximation that scales well with the number of data points and the size of the domain. We demonstrate that this allows MCMC approximations to the non-Gaussian posterior. In practice, we find that Fourier features have more consistent optimization behavior than previous approaches. Our approximate Bayesian method can fit over 100 000 events with complex spatiotemporal patterns in three dimensions on a single GPU. S. T. John, James Hensman |
ICML | 2 |
| 2018 | Gaussian Process Conditional Density EstimationabstractConditional Density Estimation (CDE) models deal with estimating conditional distributions. The conditions imposed on the distribution are the inputs of the model. CDE is a challenging task as there is a fundamental trade-off between model complexity, representational capacity and overfitting. In this work, we propose to extend the model's input with latent variables and use Gaussian processes (GP) to map this augmented input onto samples from the conditional distribution. Our Bayesian approach allows for the modeling of small datasets, but we also provide the machinery for it to be applied to big data using stochastic variational inference. Our approach can be used to model densities even in sparse data regions, and allows for sharing learned structure between conditions. We illustrate the effectiveness and wide-reaching applicability of our model on a variety of real-world problems, such as spatio-temporal density estimation of taxi drop-offs, non-Gaussian noise modeling, and few-shot learning on omniglot images. Vincent Dutordoir, Hugh Salimbeni, James Hensman, Marc Peter Deisenroth |
NeurIPS | 3 |
| 2018 | Infinite-Horizon Gaussian ProcessesabstractGaussian processes provide a flexible framework for forecasting, removing noise, and interpreting long temporal datasets. State space modelling (Kalman filtering) enables these non-parametric models to be deployed on long datasets by reducing the complexity to linear in the number of data points. The complexity is still cubic in the state dimension m which is an impediment to practical application. In certain special cases (Gaussian likelihood, regular spacing) the GP posterior will reach a steady posterior state when the data are very long. We leverage this and formulate an inference scheme for GPs with general likelihoods, where inference is based on single-sweep EP (assumed density filtering). The infinite-horizon model tackles the cubic cost in the state dimensionality and reduces the cost in the state dimension m to O(m^2) per data point. The model is extended to online-learning of hyperparameters. We show examples for large finite-length modelling problems, and present how the method runs in real-time on a smartphone on a continuous data stream updated at 100 Hz. Arno Solin, James Hensman, Richard E. Turner |
NeurIPS | 2 |
| 2018 | Learning Invariances using the Marginal LikelihoodabstractIn many supervised learning tasks, learning what changes do not affect the predic-tion target is as crucial to generalisation as learning what does. Data augmentationis a common way to enforce a model to exhibit an invariance: training data is modi-fied according to an invariance designed by a human and added to the training data.We argue that invariances should be incorporated the model structure, and learnedusing themarginal likelihood, which can correctly reward the reduced complexityof invariant models. We incorporate invariances in a Gaussian process, due to goodmarginal likelihood approximations being available for these models. Our maincontribution is a derivation for a variational inference scheme for invariant Gaussianprocesses where the invariance is described by a probability distribution that canbe sampled from, much like how data augmentation is implemented in practice Mark van der Wilk, Matthias Bauer 0001, S. T. John, James Hensman |
NeurIPS | 4 |
| 2018 | Scalable Joint Models for Reliable Uncertainty-Aware Event PredictionabstractMissing data and noisy observations pose significant challenges for reliably predicting events from irregularly sampled multivariate time series (longitudinal) data. Imputation methods, which are typically used for completing the data prior to event prediction, lack a principled mechanism to account for the uncertainty due to missingness. Alternatively, state-of-the-art joint modeling techniques can be used for jointly modeling the longitudinal and event data and compute event probabilities conditioned on the longitudinal observations. These approaches, however, make strong parametric assumptions and do not easily scale to multivariate signals with many observations. Our proposed approach consists of several key innovations. First, we develop a flexible and scalable joint model based upon sparse multiple-output Gaussian processes. Unlike state-of-the-art joint models, the proposed model can explain highly challenging structure including non-Gaussian noise while scaling to large data. Second, we derive an optimal policy for predicting events using the distribution of the event occurrence estimated by the joint model. The derived policy trades-off the cost of a delayed detection versus incorrect assessments and abstains from making decisions when the estimated event probability does not satisfy the derived confidence criteria. Experiments on a large dataset show that the proposed framework significantly outperforms state-of-the-art techniques in event prediction. Hossein Soleimani, James Hensman, Suchi Saria |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Identification of Gaussian Process State Space ModelsabstractThe Gaussian process state space model (GPSSM) is a non-linear dynamical system, where unknown transition and/or measurement mappings are described by GPs. Most research in GPSSMs has focussed on the state estimation problem, i.e., computing a posterior of the latent state given the model. However, the key challenge in GPSSMs has not been satisfactorily addressed yet: system identification, i.e., learning the model. To address this challenge, we impose a structured Gaussian variational posterior distribution over the latent states, which is parameterised by a recognition model in the form of a bi-directional recurrent neural network. Inference with this structure allows us to recover a posterior smoothed over sequences of data. We provide a practical algorithm for efficiently computing a lower bound on the marginal likelihood using the reparameterisation trick. This further allows for the use of arbitrary kernels within the GPSSM. We demonstrate that the learnt GPSSM can efficiently generate plausible future trajectories of the identified system after only observing a small number of episodes from the true system. Stefanos Eleftheriadis, Tom Nicholson, Marc Peter Deisenroth, James Hensman |
NIPS | 4 |
| 2017 | Convolutional Gaussian ProcessesabstractWe present a practical way of introducing convolutional structure into Gaussian processes, making them more suited to high-dimensional inputs like images. The main contribution of our work is the construction of an inter-domain inducing point approximation that is well-tailored to the convolutional kernel. This allows us to gain the generalisation benefit of a convolutional kernel, together with fast but accurate posterior inference. We investigate several variations of the convolutional kernel, and apply it to MNIST and CIFAR-10, where we obtain significant improvements over existing Gaussian process models. We also show how the marginal likelihood can be used to find an optimal weighting between convolutional and RBF kernels to further improve performance. This illustration of the usefulness of the marginal likelihood may help automate discovering architectures in larger models. Mark van der Wilk, Carl E. Rasmussen, James Hensman |
NIPS | 3 |
| 2017 | Variational Fourier Features for Gaussian Processes
James Hensman, Nicolas Durrande, Arno Solin |
J. Mach. Learn. Res. | 1 |
| 2017 | GPflow: A Gaussian Process Library using TensorFlowabstractGPflow is a Gaussian process library that uses TensorFlow for its core computations and Python for its front end. The distinguishing features of GPflow are that it uses variational inference as the primary approximation method, provides concise code through the use of automatic differentiation, has been engineered with a particular emphasis on software testing and is able to exploit GPU hardware. Alexander G. de G. Matthews, Mark van der Wilk, Tom Nickson, Keisuke Fujii 0002, Alexis Boukouvalas, Pablo León-Villagrá, Zoubin Ghahramani, James Hensman |
J. Mach. Learn. Res. | 8 |
| 2016 | On Sparse Variational Methods and the Kullback-Leibler Divergence between Stochastic ProcessesabstractThe variational framework for learning inducing variables (Titsias, 2009) has had a large impact on the Gaussian process literature. The framework may be interpreted as minimizing a rigorously defined Kullback-Leibler divergence between the approximating and posterior processes. To our knowledge this connection has thus far gone unremarked in the literature. In this paper we give a substantial generalization of the literature on this topic. We give a new proof of the result for infinite index sets which allows inducing points that are not data points and likelihoods that depend on all function values. We then discuss augmented index sets and show that, contrary to previous works, marginal consistency of augmentation is not enough to guarantee consistency of variational inference with the original model. We then characterize an extra condition where such a guarantee is obtainable. Finally we show how our framework sheds light on interdomain sparse approximations and sparse approximations for Cox processes. Alexander G. de G. Matthews, James Hensman, Richard E. Turner, Zoubin Ghahramani |
AISTATS | 2 |
| 2016 | Chained Gaussian ProcessesabstractGaussian process models are flexible, Bayesian non-parametric approaches to regression. Properties of multivariate Gaussians mean that they can be combined linearly in the manner of additive models and via a link function (like in generalized linear models) to handle non-Gaussian data. However, the link function formalism is restrictive, link functions are always invertible and must convert a parameter of interest to an linear combination of the underlying processes. There are many likelihoods and models where a non-linear combination is more appropriate. We term these more general models "Chained Gaussian Processes": the transformation of the GPs to the likelihood parameters will not generally be invertible, and that implies that linearisation would only be possible with multiple (localized) links, i.e a chain. We develop an approximate inference procedure for Chained GPs that is scalable and applicable to any factorized likelihood. We demonstrate the approximation on a range of likelihood functions. Alan D. Saul, James Hensman, Aki Vehtari, Neil D. Lawrence |
AISTATS | 2 |
| 2015 | Scalable Variational Gaussian Process ClassificationabstractGaussian process classification is a popular method with a number of appealing properties. We show how to scale the model within a variational inducing point framework, out-performing the state of the art on benchmark datasets. Importantly, the variational formulation an be exploited to allow classification in problems with millions of data points, as we demonstrate in experiments. James Hensman, Alexander G. de G. Matthews, Zoubin Ghahramani |
AISTATS | 1 |
| 2015 | MCMC for Variationally Sparse Gaussian ProcessesabstractGaussian process (GP) models form a core part of probabilistic machine learning. Considerable research effort has been made into attacking three issues with GP models: how to compute efficiently when the number of data is large; how to approximate the posterior when the likelihood is not Gaussian and how to estimate covariance function parameter posteriors. This paper simultaneously addresses these, using a variational approximation to the posterior which is sparse in sup- port of the function but otherwise free-form. The result is a Hybrid Monte-Carlo sampling scheme which allows for a non-Gaussian approximation over the function values and covariance parameters simultaneously, with efficient computations based on inducing-point sparse GPs. James Hensman, Alexander G. de G. Matthews, Maurizio Filippone, Zoubin Ghahramani |
NIPS | 1 |
| 2015 | Fast and accurate approximate inference of transcript expression from RNA-seq dataabstractMOTIVATION: Assigning RNA-seq reads to their transcript of origin is a fundamental task in transcript expression estimation. Where ambiguities in assignments exist due to transcripts sharing sequence, e.g. alternative isoforms or alleles, the problem can be solved through probabilistic inference. Bayesian methods have been shown to provide accurate transcript abundance estimates compared with competing methods. However, exact Bayesian inference is intractable and approximate methods such as Markov chain Monte Carlo and Variational Bayes (VB) are typically used. While providing a high degree of accuracy and modelling flexibility, standard implementations can be prohibitively slow for large datasets and complex transcriptome annotations. RESULTS: We propose a novel approximate inference scheme based on VB and apply it to an existing model of transcript expression inference from RNA-seq data. Recent advances in VB algorithmics are used to improve the convergence of the algorithm beyond the standard Variational Bayes Expectation Maximization algorithm. We apply our algorithm to simulated and biological datasets, demonstrating a significant increase in speed with only very small loss in accuracy of expression level estimation. We carry out a comparative study against seven popular alternative methods and demonstrate that our new algorithm provides excellent accuracy and inter-replicate consistency while remaining competitive in computation time. AVAILABILITY AND IMPLEMENTATION: The methods were implemented in R and C++, and are available as part of the BitSeq project at github.com/BitSeq. The method is also available through the BitSeq Bioconductor package. The source code to reproduce all simulation results can be accessed via github.com/BitSeq/BitSeqVB_benchmarking. James Hensman, Panagiotis Papastamoulis, Peter Glaus, Antti Honkela, Magnus Rattray |
Bioinform. | 1 |
| 2015 | Fast Nonparametric Clustering of Structured Time-SeriesabstractIn this publication, we combine two Bayesian nonparametric models: the Gaussian Process (GP) and the Dirichlet Process (DP). Our innovation in the GP model is to introduce a variation on the GP prior which enables us to model structured time-series data, i.e., data containing groups where we wish to model inter- and intra-group variability. Our innovation in the DP model is an implementation of a new fast collapsed variational inference procedure which enables us to optimize our variational approximation significantly faster than standard VB approaches. In a biological time series application we show how our model better captures salient features of the data, leading to better consistency with existing biological classifications, while the associated inference algorithm provides a significant speed-up over EM-based variational inference. James Hensman, Magnus Rattray, Neil D. Lawrence |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Tilted Variational BayesabstractWe present a novel method for approximate inference. Using some of the constructs from expectation propagation (EP), we derive a lower bound of the marginal likelihood in a similar fashion to variational Bayes (VB). The method combines some of the benefits of VB and EP: it can be used with light-tailed likelihoods (where traditional VB fails), and it provides a lower bound on the marginal likelihood. We apply the method to Gaussian process classification, a situation where the Kullback-Leibler divergence minimized in traditional VB can be infinite, and to robust Gaussian process regression, where the inference process is dramatically simplified in comparison to EP. Code to reproduce all the experiments can be found at github.com/SheffieldML/TVB. James Hensman, Max Zwiessele, Neil D. Lawrence |
AISTATS | 1 |
| 2014 | Hybrid Discriminative-Generative Approach with Gaussian ProcessesabstractMachine learning practitioners are often faced with a choice between a discriminative and a generative approach to modelling. Here, we present a model based on a hybrid approach that breaks down some of the barriers between the discriminative and generative points of view, allowing continuous dimensionality reduction of hybrid discrete-continuous data, discriminative classification with missing inputs and manifold learning informed by class labels. Ricardo Andrade Pacheco, James Hensman, Max Zwiessele, Neil D. Lawrence |
AISTATS | 2 |
| 2013 | Gaussian Processes for Big Data
James Hensman, Nicolò Fusi, Neil D. Lawrence |
UAI | 1 |
| 2013 | Hierarchical Bayesian modelling of gene expression time series across irregularly sampled replicates and clustersabstractBACKGROUND: Time course data from microarrays and high-throughput sequencing experiments require simple, computationally efficient and powerful statistical models to extract meaningful biological signal, and for tasks such as data fusion and clustering. Existing methodologies fail to capture either the temporal or replicated nature of the experiments, and often impose constraints on the data collection process, such as regularly spaced samples, or similar sampling schema across replications. RESULTS: We propose hierarchical Gaussian processes as a general model of gene expression time-series, with application to a variety of problems. In particular, we illustrate the method's capacity for missing data imputation, data fusion and clustering.The method can impute data which is missing both systematically and at random: in a hold-out test on real data, performance is significantly better than commonly used imputation methods. The method's ability to model inter- and intra-cluster variance leads to more biologically meaningful clusters. The approach removes the necessity for evenly spaced samples, an advantage illustrated on a developmental Drosophila dataset with irregular replications. CONCLUSION: The hierarchical Gaussian process model provides an excellent statistical basis for several gene-expression time-series tasks. It has only a few additional parameters over a regular GP, has negligible additional complexity, is easily implemented and can be integrated into several existing algorithms. Our experiments were implemented in python, and are available from the authors' website: http://staffwww.dcs.shef.ac.uk/people/J.Hensman/. James Hensman, Neil D. Lawrence, Magnus Rattray |
BMC Bioinform. | 1 |
| 2012 | Fast Variational Inference in the Conjugate Exponential FamilyabstractWe present a general method for deriving collapsed variational inference algorithms for probabilistic models in the conjugate exponential family. Our method unifies many existing approaches to collapsed variational inference. Our collapsed variational inference leads to a new lower bound on the marginal likelihood. We exploit the information geometry of the bound to derive much faster optimization methods based on conjugate gradients for these models. Our approach is very general and is easily applied to any model where the mean field update equations have been derived. Empirically we show significant speed-ups for probabilistic models optimized using our bound. James Hensman, Magnus Rattray, Neil D. Lawrence |
NIPS | 1 |