Thomas Cass

dblp:256/0784 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Kernel, tree and ensemble methods · 47% Deep learning architectures and training · 21% Probabilistic and Bayesian machine learning · 12%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Time series and sequential data
anomaly detection
0.912025
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection · J. Mach. Learn. Res. 2025
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.912025
Random Feature Representation Boosting · ICML 2025
Machine learning › Kernel, tree and ensemble methods
gradient boosting
0.912025
Random Feature Representation Boosting · ICML 2025
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.912025
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection · J. Mach. Learn. Res. 2025
Machine learning › Trustworthy machine learning
novelty detection
0.912025
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection · J. Mach. Learn. Res. 2025
Machine learning › Deep learning architectures and training › feedforward neural network
random feature neural network
0.912025
Random Feature Representation Boosting · ICML 2025
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
0.912025
Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection · J. Mach. Learn. Res. 2025
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.912025
Random Feature Representation Boosting · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.512021
SigGPDE: Scaling Sparse Gaussian Processes on Sequential Data · ICML 2021
Machine learning › Kernel, tree and ensemble methods › kernel methods
signature kernel
0.512021
SigGPDE: Scaling Sparse Gaussian Processes on Sequential Data · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › sparse gaussian process
sparse variational gaussian process
0.512021
SigGPDE: Scaling Sparse Gaussian Processes on Sequential Data · ICML 2021

Methods — techniques the papers use, named apart from their topics

tikhonov regularization · 0.9random features · 0.9quadratically constrained least squares · 0.9empirical measure estimation · 0.9boosting theory · 0.9inducing variables · 0.5hyperbolic partial differential equations · 0.5evidence lower bound · 0.5
YearPublicationVenuePosition
2025 Random Feature Representation Boosting
abstract
We introduce Random Feature Representation Boosting (RFRBoost), a novel method for constructing deep residual random feature neural networks (RFNNs) using boosting theory. RFRBoost uses random features at each layer to learn the functional gradient of the network representation, enhancing performance while preserving the convex optimization benefits of RFNNs. In the case of MSE loss, we obtain closed-form solutions to greedy layer-wise boosting with random features. For general loss functions, we show that fitting random feature residual blocks reduces to solving a quadratically constrained least squares problem. Through extensive numerical experiments on tabular datasets for both regression and classification, we show that RFRBoost significantly outperforms RFNNs and end-to-end trained MLP ResNets in the small- to medium-scale regime where RFNNs are typically applied. Moreover, RFRBoost offers substantial computational benefits, and theoretical guarantees stemming from boosting theory.
Nikita Zozoulenko, Thomas Cass, Lukas Gonon
ICML2
2025 Infinite-dimensional Mahalanobis Distance with Applications to Kernelized Novelty Detection
abstract
The Mahalanobis distance is a classical tool used to measure the covariance-adjusted distance between points in $\mathbb{R}^d$. In this work, we extend the concept of Mahalanobis distance to separable Banach spaces by reinterpreting it as a Cameron-Martin norm associated with a probability measure. This approach leads to a basis-free, data-driven notion of anomaly distance through the so-called variance norm, which can naturally be estimated using empirical measures of a sample. Our framework generalizes the classical $\mathbb{R}^d$, functional $(L^2[0,1])^d$, and kernelized settings; importantly, it incorporates non-injective covariance operators. We prove that the variance norm is invariant under invertible bounded linear transformations of the data, extending previous results which are limited to unitary operators. In the Hilbert space setting, we connect the variance norm to the RKHS of the covariance operator, and establish consistency and convergence results for estimation using empirical measures with Tikhonov regularization. Using the variance norm, we introduce the notion of a kernelized nearest-neighbour Mahalanobis distance, and study some of its finite-sample concentration properties. In an empirical study on 12 real-world data sets, we demonstrate that the kernelized nearest-neighbour Mahalanobis distance outperforms the traditional kernelized Mahalanobis distance for multivariate time series novelty detection, using state-of-the-art time series kernels such as the signature, global alignment, and Volterra reservoir kernels.
Nikita Zozoulenko, Thomas Cass, Lukas Gonon
J. Mach. Learn. Res.2
2021 SigGPDE: Scaling Sparse Gaussian Processes on Sequential Data
abstract
Making predictions and quantifying their uncertainty when the input data is sequential is a fundamental learning challenge, recently attracting increasing attention. We develop SigGPDE, a new scalable sparse variational inference framework for Gaussian Processes (GPs) on sequential data. Our contribution is twofold. First, we construct inducing variables underpinning the sparse approximation so that the resulting evidence lower bound (ELBO) does not require any matrix inversion. Second, we show that the gradients of the GP signature kernel are solutions of a hyperbolic partial differential equation (PDE). This theoretical insight allows us to build an efficient back-propagation algorithm to optimize the ELBO. We showcase the significant computational gains of SigGPDE compared to existing methods, while achieving state-of-the-art performance for classification tasks on large datasets of up to 1 million multivariate time series.
Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V. Bonilla, Theodoros Damoulas, Terry J. Lyons
ICML3