Harald Oberhauser

dblp:175/1262 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-2644-8906ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning to Forget: Bayesian Time Series Forecasting using Recurrent Sparse Spectrum Signature Gaussian Processes
abstract
The signature kernel is a kernel between time series of arbitrary length and comes with strong theoretical guarantees from stochastic analysis. It has found applications in machine learning such as covariance functions for Gaussian processes. A strength of the underlying signature features is that they provide a structured global description of a time series. However, this property can quickly become a curse when local information is essential and forgetting is required; so far this has only been addressed with ad-hoc methods such as slicing the time series into smaller segments. To overcome this, we propose a principled and data-driven approach by introducing a novel forgetting mechanism for signature features. This allows the model to dynamically adapt its observed context length to focus on more recent information. To achieve this, we revisit the recently introduced Random Fourier Signature Features, and develop Random Fourier Decayed Signature Features (RFDSF) with Gaussian processes (GPs). This results in a Bayesian time series forecasting algorithm with variational inference, that offers a scalable probabilistic algorithm that processes and transforms a time series into a joint predictive distribution over the time steps in one pass using recurrence. For example, processing a sequence of length $10^4$ steps in less than $10^{-2}$ seconds and in $\approx$ 1GB of GPU memory. We demonstrate that the algorithm outperforms other GP-based alternatives and competes with state-of-the-art probabilistic time series forecasting algorithms.
Csaba Tóth, Masaki Adachi, Michael A. Osborne, Harald Oberhauser
AISTATS4
2025 Generalized Time Series Classification via Component Decomposition and Alignment
abstract
The objective of domain generalization is to develop a model that can handle the domain shift problem without access to the target domain. In this paper, we propose a new domain generalization approach called Decomposition Framework with Dynamic Component Alignment (DFDCA), which employs signal decomposition on input data and conducts domain alignment on each component, providing another perspective on domain generalization for time series classification. Specifically, we first utilize a neural decomposition module to decompose the original time series data into several components, and design loss functions to guide the network to effectively perform signal decomposition for class-wise domain alignment on the decomposed components. The denoising attention mechanism is then introduced to enhance informative components while suppressing task-irrelevant components. Our proposed approach is evaluated on four publicly available datasets based on the cross-domain setting where the training and test samples are drawn from different distributions. The results demonstrate that it outperforms other baseline methods, achieving state-of-the-art performance.
Yichuan Cheng, Darrick Lee, Harald Oberhauser, Haoliang Li
IEEE Trans. Big Data3
2024 Adaptive Batch Sizes for Active Learning: A Probabilistic Numerics Approach
abstract
Active learning parallelization is widely used, but typically relies on fixing the batch size throughout experimentation. This fixed approach is inefficient because of a dynamic trade-off between cost and speed—larger batches are more costly, smaller batches lead to slower wall-clock run-times—and the trade-off may change over the run (larger batches are often preferable earlier). To address this trade-off, we propose a novel Probabilistic Numerics framework that adaptively changes batch sizes. By framing batch selection as a quadrature task, our integration-error-aware algorithm facilitates the automatic tuning of batch sizes to meet predefined quadrature precision objectives, akin to how typical optimizers terminate based on convergence thresholds. This approach obviates the necessity for exhaustive searches across all potential batch sizes. We also extend this to scenarios with constrained active learning and constrained optimization, interpreting constraint violations as reductions in the precision requirement, to subsequently adapt batch construction. Through extensive experiments, we demonstrate that our approach significantly enhances learning efficiency and flexibility in diverse Bayesian batch active learning and Bayesian optimization applications.
Masaki Adachi, Satoshi Hayakawa, Martin Jørgensen, Xingchen Wan, Harald Oberhauser, Michael A. Osborne
AISTATS6
2023 Sampling-based Nyström Approximation and Kernel Quadrature
abstract
We analyze the Nyström approximation of a positive definite kernel associated with a probability measure. We first prove an improved error bound for the conventional Nyström approximation with i.i.d. sampling and singular-value decomposition in the continuous regime; the proof techniques are borrowed from statistical learning theory. We further introduce a refined selection of subspaces in Nyström approximation with theoretical guarantees that is applicable to non-i.i.d. landmark points. Finally, we discuss their application to convex kernel quadrature and give novel theoretical guarantees as well as numerical observations.
Satoshi Hayakawa, Harald Oberhauser, Terry J. Lyons
ICML2
2023 Kernelized Cumulants: Beyond Kernel Mean Embeddings
abstract
In $\mathbb{R}^d$, it is well-known that cumulants provide an alternative to moments that can achieve the same goals with numerous benefits such as lower variance estimators. In this paper we extend cumulants to reproducing kernel Hilbert spaces (RKHS) using tools from tensor algebras and show that they are computationally tractable by a kernel trick. These kernelized cumulants provide a new set of all-purpose statistics; the classical maximum mean discrepancy and Hilbert-Schmidt independence criterion arise as the degree one objects in our general construction. We argue both theoretically and empirically (on synthetic, environmental, and traffic data analysis) that going beyond degree one has several advantages and can be achieved with the same computational complexity and minimal overhead in our experiments.
Patric Bonnier, Harald Oberhauser
NeurIPS2
2022 Fast Bayesian Inference with Batch Bayesian Quadrature via Kernel Recombination
abstract
Calculation of Bayesian posteriors and model evidences typically requires numerical integration. Bayesian quadrature (BQ), a surrogate-model-based approach to numerical integration, is capable of superb sample efficiency, but its lack of parallelisation has hindered its practical applications. In this work, we propose a parallelised (batch) BQ method, employing techniques from kernel quadrature, that possesses an empirically exponential convergence rate.Additionally, just as with Nested Sampling, our method permits simultaneous inference of both posteriors and model evidence.Samples from our BQ surrogate model are re-selected to give a sparse set of samples, via a kernel recombination algorithm, requiring negligible additional time to increase the batch size.Empirically, we find that our approach significantly outperforms the sampling efficiency of both state-of-the-art BQ techniques and Nested Sampling in various real-world datasets, including lithium-ion battery analytics.
Masaki Adachi, Satoshi Hayakawa, Martin Jørgensen, Harald Oberhauser, Michael A. Osborne
NeurIPS4
2022 Positively Weighted Kernel Quadrature via Subsampling
abstract
We study kernel quadrature rules with convex weights. Our approach combines the spectral properties of the kernel with recombination results about point measures. This results in effective algorithms that construct convex quadrature rules using only access to i.i.d. samples from the underlying measure and evaluation of the kernel and that result in a small worst-case error. In addition to our theoretical results and the benefits resulting from convex weights, our experiments indicate that this construction can compete with the optimal bounds in well-known examples.
Satoshi Hayakawa, Harald Oberhauser, Terry J. Lyons
NeurIPS2
2022 Capturing Graphs with Hypo-Elliptic Diffusions
abstract
Convolutional layers within graph neural networks operate by aggregating information about local neighbourhood structures; one common way to encode such substructures is through random walks. The distribution of these random walks evolves according to a diffusion equation defined using the graph Laplacian. We extend this approach by leveraging classic mathematical results about hypo-elliptic diffusions. This results in a novel tensor-valued graph operator, which we call the hypo-elliptic graph Laplacian. We provide theoretical guarantees and efficient low-rank approximation algorithms. In particular, this gives a structured approach to capture long-range dependencies on graphs that is robust to pooling. Besides the attractive theoretical properties, our experiments show that this method competes with graph transformers on datasets requiring long-range reasoning but scales only linearly in the number of edges as opposed to quadratically in nodes.
Csaba Tóth, Darrick Lee, Celia Hacker, Harald Oberhauser
NeurIPS4
2022 Signature Moments to Characterize Laws of Stochastic Processes
abstract
The sequence of moments of a vector-valued random variable can characterize its law. We study the analogous problem for path-valued random variables, that is stochastic processes, by using so-called robust signature moments. This allows us to derive a metric of maximum mean discrepancy type for laws of stochastic processes and study the topology it induces on the space of laws of stochastic processes. This metric can be kernelized using the signature kernel which allows to efficiently compute it. As an application, we provide a non-parametric two-sample hypothesis test for laws of stochastic processes.
Ilya Chevyrev, Harald Oberhauser
J. Mach. Learn. Res.2
2021 Seq2Tens: An Efficient Representation of Sequences by Low-Rank Tensor Projections
Csaba Tóth, Patric Bonnier, Harald Oberhauser
ICLR3
2020 Bayesian Learning from Sequential Data using Gaussian Processes with Signature Covariances
abstract
We develop a Bayesian approach to learning from sequential data by using Gaussian processes (GPs) with so-called signature kernels as covariance functions. This allows to make sequences of different length comparable and to rely on strong theoretical results from stochastic analysis. Signatures capture sequential structure with tensors that can scale unfavourably in sequence length and state space dimension. To deal with this, we introduce a sparse variational approach with inducing tensors. We then combine the resulting GP with LSTMs and GRUs to build larger models that leverage the strengths of each of these approaches and benchmark the resulting GPs on multivariate time series (TS) classification datasets.
Csaba Tóth, Harald Oberhauser
ICML2
2020 A Randomized Algorithm to Reduce the Support of Discrete Measures
abstract
Given a discrete probability measure supported on $N$ atoms and a set of $n$ real-valued functions, there exists a probability measure that is supported on a subset of $n+1$ of the original $N$ atoms and has the same mean when integrated against each of the $n$ functions. If $ N \gg n$ this results in a huge reduction of complexity. We give a simple geometric characterization of barycenters via negative cones and derive a randomized algorithm that computes this new measure by ``greedy geometric sampling''. We then study its properties, and benchmark it on synthetic and real-world data to show that it can be very beneficial in the $N\gg n$ regime. A Python implementation is available at \url{https://github.com/FraCose/Recombination_Random_Algos}.
Francesco Cosentino, Harald Oberhauser, Alessandro Abate
NeurIPS2
2020 Persistence Paths and Signature Features in Topological Data Analysis
abstract
We introduce a new feature map for barcodes as they arise in persistent homology computation. The main idea is to first realize each barcode as a path in a convenient vector space, and to then compute its path signature which takes values in the tensor algebra of that vector space. The composition of these two operations-barcode to path, path to tensor series-results in a feature map that has several desirable properties for statistical learning, such as universality and characteristicness, and achieves state-of-the-art results on common classification benchmarks.
Ilya Chevyrev, Vidit Nanda, Harald Oberhauser
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Kernels for Sequentially Ordered Data
abstract
We present a novel framework for learning with sequential data of any kind, such as multivariate time series, strings, or sequences of graphs. The main result is a ”sequentialization” that transforms any kernel on a given domain into a kernel for sequences in that domain. This procedure preserves properties such as positive definiteness, the associated kernel feature map is an ordered variant of sample (cross-)moments, and this sequentialized kernel is consistent in the sense that it converges to a kernel for paths if sequences converge to paths (by discretization). Further, classical kernels for sequences arise as special cases of this method. We use dynamic programming and low-rank techniques for tensors to provide efficient algorithms to compute this sequentialized kernel.
Franz J. Király, Harald Oberhauser
J. Mach. Learn. Res.2