Mattes Mollenhauer

dblp:217/3144 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Learning theory · 62% Kernel, tree and ensemble methods · 24% Probabilistic and Bayesian machine learning · 8%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
statistical learning theory
2.742024
Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm · J. Mach. Learn. Res. 2024
Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms · NeurIPS 2024
Kernel Autocovariance Operators of Stationary Processes: Estimation and Convergence · J. Mach. Learn. Res. 2022
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression
1.322024
Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm · J. Mach. Learn. Res. 2024
Optimal Rates for Regularized Conditional Mean Embedding Learning · NeurIPS 2022
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates
1.322024
Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms · NeurIPS 2024
Optimal Rates for Regularized Conditional Mean Embedding Learning · NeurIPS 2022
Machine learning › Learning theory
excess risk bounds
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Learning theory
generalization bounds
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Learning theory › nonparametric regression
sobolev norm rates
0.812024
Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm · J. Mach. Learn. Res. 2024
Machine learning › Deep learning architectures and training › regularization
spectral regularization
0.812024
Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods › kernel embedding
conditional mean embedding
0.612022
Optimal Rates for Regularized Conditional Mean Embedding Learning · NeurIPS 2022
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.612022
Kernel Autocovariance Operators of Stationary Processes: Estimation and Convergence · J. Mach. Learn. Res. 2022
Machine learning › Learning theory
integral operator
0.312025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
0.312025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

reproducing kernel hilbert space · 1.9kernel ridge regression · 1.3fuk-nagaev inequality · 0.9eigenvalue decay · 0.9tensor product construction · 0.8lower bounds · 0.8interpolation space · 0.8concentration inequalities · 0.8spectral analysis · 0.6interpolation space theory · 0.6
YearPublicationVenuePosition
2025 Regularized least squares learning with heavy-tailed noise is minimax optimal
abstract
This paper examines the performance of ridge regression in reproducing kernel Hilbert spaces in the presence of noise that exhibits a finite number of higher moments. We establish excess risk bounds consisting of subgaussian and polynomial terms based on the well known integral operator framework. The dominant subgaussian component allows to achieve convergence rates that have previously only been derived under subexponential noise—a prevalent assumption in related work from the last two decades. These rates are optimal under standard eigenvalue decay conditions, demonstrating the asymptotic robustness of regularized least squares against heavy- tailed noise. Our derivations are based on a Fuk–Nagaev inequality for Hilbert-space valued random variables.
Mattes Mollenhauer, Nicole Mücke, Dimitri Meunier, Arthur Gretton
NeurIPS1
2024 Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms
abstract
We study theoretical properties of a broad class of regularized algorithms with vector-valued output. These spectral algorithms include kernel ridge regression, kernel principal component regression and various implementations of gradient descent. Our contributions are twofold. First, we rigorously confirm the so-called saturation effect for ridge regression with vector-valued output by deriving a novel lower bound on learning rates; this bound is shown to be suboptimal when the smoothness of the regression function exceeds a certain level. Second, we present an upper bound on the finite sample risk for general vector-valued spectral algorithms, applicable to both well-specified and misspecified scenarios (where the true regression function lies outside of the hypothesis space), and show that this bound is minimax optimal in various regimes. All of our results explicitly allow the case of infinite-dimensional output variables, proving consistency of recent practical applications.
Dimitri Meunier, Zikai Shen, Mattes Mollenhauer, Arthur Gretton
NeurIPS3
2024 Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm
abstract
We present the first optimal rates for infinite-dimensional vector-valued ridge regression on a continuous scale of norms that interpolate between L2 and the hypothesis space, which we consider as a vector-valued reproducing kernel Hilbert space. These rates allow to treat the misspecified case in which the true regression function is not contained in the hypothesis space. We combine standard assumptions on the capacity of the hypothesis space with a novel tensor product construction of vector-valued interpolation spaces in order to characterize the smoothness of the regression function. Our upper bound not only attains the same rate as real-valued kernel ridge regression, but also removes the assumption that the target regression function is bounded. For the lower bound, we reduce the problem to the scalar setting using a projection argument. We show that these rates are optimal in most cases and independent of the dimension of the output space. We illustrate our results for the special case of vector-valued Sobolev spaces.
Dimitri Meunier, Mattes Mollenhauer, Arthur Gretton
J. Mach. Learn. Res.3
2022 Optimal Rates for Regularized Conditional Mean Embedding Learning
abstract
We address the consistency of a kernel ridge regression estimate of the conditional mean embedding (CME), which is an embedding of the conditional distribution of $Y$ given $X$ into a target reproducing kernel Hilbert space $\mathcal{H}_Y$. The CME allows us to take conditional expectations of target RKHS functions, and has been employed in nonparametric causal and Bayesian inference.We address the misspecified setting, where the target CME isin the space of Hilbert-Schmidt operators acting from an input interpolation space between $\mathcal{H}_X$ and $L_2$, to $\mathcal{H}_Y$. This space of operators is shown to be isomorphic to a newly defined vector-valued interpolation space. Using this isomorphism, we derive a novel and adaptive statistical learning rate for the empirical CME estimator under the misspecified setting. Our analysis reveals that our rates match the optimal $O(\log n / n)$ rates without assuming $\mathcal{H}_Y$ to be finite dimensional. We further establish a lower bound on the learning rate, which shows that the obtained upper bound is optimal.
Dimitri Meunier, Mattes Mollenhauer, Arthur Gretton
NeurIPS3
2022 Kernel Autocovariance Operators of Stationary Processes: Estimation and Convergence
abstract
We consider autocovariance operators of a stationary stochastic process on a Polish space that is embedded into a reproducing kernel Hilbert space. We investigate how empirical estimates of these operators converge along realizations of the process under various conditions. In particular, we examine ergodic and strongly mixing processes and obtain several asymptotic results as well as finite sample error bounds. We provide applications of our theory in terms of consistency results for kernel PCA with dependent data and the conditional mean embedding of transition probabilities. Finally, we use our approach to examine the nonparametric estimation of Markov transition operators and highlight how our theory can give a consistency analysis for a large family of spectral analysis methods including kernel-based dynamic mode decomposition.
Mattes Mollenhauer, Stefan Klus, Christof Schütte, Péter Koltai
J. Mach. Learn. Res.1
2020 Kernel Conditional Density Operators
abstract
We introduce a novel conditional density estimationmodel termed the conditional densityoperator (CDO). It naturally captures multivariate,multimodal output densities andshows performance that is competitive withrecent neural conditional density models andGaussian processes. The proposed model isbased on a novel approach to the reconstructionof probability densities from their kernelmean embeddings by drawing connections toestimation of Radon-Nikodym derivatives inthe reproducing kernel Hilbert space (RKHS).We prove finite sample bounds for the estimationerror in a standard density reconstructionscenario, independent of problem dimensionality.Interestingly, when a kernel is used thatis also a probability density, the CDO allowsus to both evaluate and sample the outputdensity efficiently. We demonstrate the versatilityand performance of the proposed modelon both synthetic and real-world data.
Ingmar Schuster, Mattes Mollenhauer, Stefan Klus, Krikamol Muandet
AISTATS2