Artem Artemev

dblp:236/5889 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 68% Reinforcement learning · 27% Efficient and distributed learning · 5%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.812024
Numerically Stable Sparse Gaussian Processes via Minimum Separation using Cover Trees · J. Mach. Learn. Res. 2024
Compilers and program optimization › compiler optimization
machine learning for compiler optimization
0.612022
Memory safe computations with XLA compiler · NeurIPS 2022
Compilers and program optimization
memory optimization
0.612022
Memory safe computations with XLA compiler · NeurIPS 2022
Machine learning › Reinforcement learning
bandit
0.512021
Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
0.512021
Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
sparse gaussian process
0.512021
Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021
Machine learning › Reinforcement learning
thompson sampling
0.512021
Scalable Thompson Sampling using Sparse Gaussian Process Models · NeurIPS 2021
Mathematical optimization › iterative methods
conjugate gradient method
0.512021
Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients · ICML 2021
Machine learning › Efficient and distributed learning
memory-efficient training
0.212022
Memory safe computations with XLA compiler · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

dataflow graph rewriting · 1.1XLA · 1.1variational method · 1.0lower bound maximization · 1.0conjugate gradient · 1.0minimum separation · 0.8inducing points · 0.8cover tree · 0.8sparse gaussian process · 0.5regret analysis · 0.5
YearPublicationVenuePosition
2024 Numerically Stable Sparse Gaussian Processes via Minimum Separation using Cover Trees
abstract
Gaussian processes are frequently deployed as part of larger machine learning and decision-making systems, for instance in geospatial modeling, Bayesian optimization, or in latent Gaussian models. Within a system, the Gaussian process model needs to perform in a stable and reliable manner to ensure it interacts correctly with other parts of the system. In this work, we study the numerical stability of scalable sparse approximations based on inducing points. To do so, we first review numerical stability, and illustrate typical situations in which Gaussian process models can be unstable. Building on stability theory originally developed in the interpolation literature, we derive sufficient and in certain cases necessary conditions on the inducing points for the computations performed to be numerically stable. For low-dimensional tasks such as geospatial modeling, we propose an automated method for computing inducing points satisfying these conditions. This is done via a modification of the cover tree data structure, which is of independent interest. We additionally propose an alternative sparse approximation for regression with a Gaussian likelihood which trades off a small amount of performance to further improve stability. We provide illustrative examples showing the relationship between stability of calculations and predictive performance of inducing point methods on spatial tasks.
Alexander Terenin, David R. Burt, Artem Artemev, Seth R. Flaxman, Mark van der Wilk, Carl E. Rasmussen
J. Mach. Learn. Res.3
2022 Memory safe computations with XLA compiler
abstract
Software packages like TensorFlow and PyTorch are designed to support linear algebra operations, and their speed and usability determine their success. However, by prioritising speed, they often neglect memory requirements. As a consequence, the implementations of memory-intensive algorithms that are convenient in terms of software design can often not be run for large problems due to memory overflows. Memory-efficient solutions require complex programming approaches with significant logic outside the computational framework. This impairs the adoption and use of such algorithms. To address this, we developed an XLA compiler extension that adjusts the computational data-flow representation of an algorithm according to a user-specified memory limit. We show that k-nearest neighbour, sparse Gaussian process regression methods and Transformers can be run on a single device at a much larger scale, where standard implementations would have failed. Our approach leads to better use of hardware resources. We believe that further focus on removing memory constraints at a compiler level will widen the range of machine learning methods that can be developed in the future.
Artem Artemev, Yuze An, Tilman Roeder, Mark van der Wilk
NeurIPS1
2021 Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients
abstract
We propose a lower bound on the log marginal likelihood of Gaussian process regression models that can be computed without matrix factorisation of the full kernel matrix. We show that approximate maximum likelihood learning of model parameters by maximising our lower bound retains many benefits of the sparse variational approach while reducing the bias introduced into hyperparameter learning. The basis of our bound is a more careful analysis of the log-determinant term appearing in the log marginal likelihood, as well as using the method of conjugate gradients to derive tight lower bounds on the term involving a quadratic form. Our approach is a step forward in unifying methods relying on lower bound maximisation (e.g. variational methods) and iterative approaches based on conjugate gradients for training Gaussian processes. In experiments, we show improved predictive performance with our model for a comparable amount of training time compared to other conjugate gradient based approaches.
Artem Artemev, David R. Burt, Mark van der Wilk
ICML1
2021 Scalable Thompson Sampling using Sparse Gaussian Process Models
abstract
Thompson Sampling (TS) from Gaussian Process (GP) models is a powerful tool for the optimization of black-box functions. Although TS enjoys strong theoretical guarantees and convincing empirical performance, it incurs a large computational overhead that scales polynomially with the optimization budget. Recently, scalable TS methods based on sparse GP models have been proposed to increase the scope of TS, enabling its application to problems that are sufficiently multi-modal, noisy or combinatorial to require more than a few hundred evaluations to be solved. However, the approximation error introduced by sparse GPs invalidates all existing regret bounds. In this work, we perform a theoretical and empirical analysis of scalable TS. We provide theoretical guarantees and show that the drastic reduction in computational complexity of scalable TS can be enjoyed without loss in the regret performance over the standard TS. These conceptual claims are validated for practical implementations of scalable TS on synthetic benchmarks and as part of a real-world high-throughput molecular design task.
Sattar Vakili, Henry B. Moss, Artem Artemev, Vincent Dutordoir, Victor Picheny
NeurIPS3
2020 Doubly Sparse Variational Gaussian Processes
abstract
The use of Gaussian process models is typically limited to datasets with a few tens of thousands of observations due to their complexity and memory footprint.The two most commonly used methods to overcome this limitation are 1) the variational sparse approximation which relies on inducing points and 2) the state-space equivalent formulation of Gaussian processes which can be seen as exploiting some sparsity in the precision matrix.In this work, we propose to take the best of both worlds: we show that the inducing point framework is still valid for state space models and that it can bring further computational and memory savings. Furthermore, we provide the natural gradient formulation for the proposed variational parameterisation.Finally, this work makes it possible to use the state-space formulation inside deep Gaussian process models as illustrated in one of the experiments.
Vincent Adam, Stefanos Eleftheriadis, Artem Artemev, Nicolas Durrande, James Hensman
AISTATS3
2020 Bayesian Image Classification with Deep Convolutional Gaussian Processes
abstract
In decision-making systems, it is important to have classifiers that have calibrated uncertainties, with an optimisation objective that can be used for automated model selection and training. Gaussian processes (GPs) provide uncertainty estimates and a marginal likelihood objective, but their weak inductive biases lead to inferior accuracy. This has limited their applicability in certain tasks (e.g. image classification). We propose a translation insensitive convolutional kernel, which relaxes the translation invariance constraint imposed by previous convolutional GPs. We show how we can use the marginal likelihood to learn the degree of insensitivity. We also reformulate GP image-to-image convolutional mappings as multi-output GPs, leading to deep convolutional GPs. We show experimentally that our new kernel improves performance in both single-layer and deep models. We also demonstrate that our fully Bayesian approach improves on dropout-based Bayesian deep learning methods in terms of uncertainty and marginal likelihood estimates.
Vincent Dutordoir, Mark van der Wilk, Artem Artemev, James Hensman
AISTATS3
2020 Automatic Tuning of Stochastic Gradient Descent with Bayesian Optimisation
Victor Picheny, Vincent Dutordoir, Artem Artemev, Nicolas Durrande
ECML/PKDD (3)3