Nicole Mücke

dblp:218/6255 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0002-5708-1820ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 50% Probabilistic and Bayesian machine learning · 28% Kernel, tree and ensemble methods · 10%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
excess risk bounds
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Learning theory
generalization bounds
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression
0.912025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression
least squares regression
0.412019
Beating SGD Saturation with Tail-Averaging and Minibatching · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › non-parametric methods
nonparametric learning
0.412019
Beating SGD Saturation with Tail-Averaging and Minibatching · NeurIPS 2019
Machine learning › Optimization for machine learning
stochastic gradient descent
0.412019
Beating SGD Saturation with Tail-Averaging and Minibatching · NeurIPS 2019
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.312018
Parallelizing Spectrally Regularized Kernel Algorithms · J. Mach. Learn. Res. 2018
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates
0.312018
Parallelizing Spectrally Regularized Kernel Algorithms · J. Mach. Learn. Res. 2018
Machine learning › Deep learning architectures and training › regularization
spectral regularization
0.312018
Parallelizing Spectrally Regularized Kernel Algorithms · J. Mach. Learn. Res. 2018
Machine learning › Learning theory
statistical learning theory
0.312018
Parallelizing Spectrally Regularized Kernel Algorithms · J. Mach. Learn. Res. 2018
Machine learning › Learning theory
integral operator
0.312025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
0.312025
Regularized least squares learning with heavy-tailed noise is minimax optimal · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

fuk-nagaev inequality · 0.9eigenvalue decay · 0.9tail averaging · 0.4minibatching · 0.4kernel ridge regression · 0.3bias-variance analysis · 0.3
YearPublicationVenuePosition
2025 Regularized least squares learning with heavy-tailed noise is minimax optimal
abstract
This paper examines the performance of ridge regression in reproducing kernel Hilbert spaces in the presence of noise that exhibits a finite number of higher moments. We establish excess risk bounds consisting of subgaussian and polynomial terms based on the well known integral operator framework. The dominant subgaussian component allows to achieve convergence rates that have previously only been derived under subexponential noise—a prevalent assumption in related work from the last two decades. These rates are optimal under standard eigenvalue decay conditions, demonstrating the asymptotic robustness of regularized least squares against heavy- tailed noise. Our derivations are based on a Fuk–Nagaev inequality for Hilbert-space valued random variables.
Mattes Mollenhauer, Nicole Mücke, Dimitri Meunier, Arthur Gretton
NeurIPS2
2025 Empirical risk minimization in the interpolating regime with application to neural network learning
abstract
Abstract A common strategy to train deep neural networks (DNNs) is to use very large architectures and to train them until they (almost) achieve zero training error. Empirically observed good generalization performance on test data, even in the presence of lots of label noise, corroborate such a procedure. On the other hand, in statistical learning theory it is known that over-fitting models may lead to poor generalization properties, occurring in e.g. empirical risk minimization (ERM) over too large hypotheses classes. Inspired by this contradictory behavior, so-called interpolation methods have recently received much attention, leading to consistent and optimally learning methods for, e.g., some local averaging schemes with zero training error. We extend this analysis to ERM-like methods for least squares regression and show that for certain, large hypotheses classes called inflated histograms, some interpolating empirical risk minimizers enjoy very good statistical guarantees while others fail in the worst sense. Moreover, we show that the same phenomenon occurs for DNNs with zero training error and sufficiently large architectures.
Nicole Mücke, Ingo Steinwart
Mach. Learn.1
2022 Data-splitting improves statistical performance in overparameterized regimes
abstract
While large training datasets generally offer improvement in model performance, the training process becomes computationally expensive and time consuming. Distributed learning is a common strategy to reduce the overall training time by exploiting multiple computing devices. Recently, it has been observed in the single machine setting that overparameterization is essential for benign overfitting in ridgeless regression in Hilbert spaces. We show that in this regime, data splitting has a regularizing effect, hence improving statistical performance and computational complexity at the same time. We further provide a unified framework that allows to analyze both the finite and infinite dimensional setting. We numerically demonstrate the effect of different model parameters.
Nicole Mücke, Enrico Reiss, Jonas Rungenhagen
AISTATS1
2021 Stochastic Gradient Descent Meets Distribution Regression
abstract
Stochastic gradient descent (SGD) provides a simple and efficient way to solve a broad range of machine learning problems. Here, we focus on distribution regression (DR), involving two stages of sampling: Firstly, we regress from probability measures to real-valued responses. Secondly, we sample bags from these distributions for utilizing them to solve the overall regression problem. Recently, DR has been tackled by applying kernel ridge regression and the learning properties of this approach are well understood. However, nothing is known about the learning properties of SGD for two stage sampling problems. We fill this gap and provide theoretical guarantees for the performance of SGD for DR. Our bounds are optimal in a mini-max sense under standard assumptions.
Nicole Mücke
AISTATS1
2019 Reducing training time by efficient localized kernel regression
abstract
We study generalization properties of kernel regularized least squares regression based on a partitioning approach. We show that optimal rates of convergence are preserved if the number of local sets grows sufficiently slowly with the sample size. Moreover, the partitioning approach can be efficiently combined with local Nyström subsampling, improving computational cost twofold.
Nicole Mücke
AISTATS1
2019 Beating SGD Saturation with Tail-Averaging and Minibatching
abstract
While stochastic gradient descent (SGD) is one of the major workhorses in machine learning, the learning properties of many practically used variants are still poorly understood. In this paper, we consider least squares learning in a nonparametric setting and contribute to filling this gap by focusing on the effect and interplay of multiple passes, mini-batching and averaging, in particular tail averaging. Our results show how these different variants of SGD can be combined to achieve optimal learning rates, also providing practical insights. A novel key result is that tail averaging allows faster convergence rates than uniform averaging in the nonparametric setting. Further, we show that a combination of tail-averaging and minibatching allows more aggressive step-size choices than using any one of said components.
Nicole Mücke, Gergely Neu, Lorenzo Rosasco
NeurIPS1
2018 Parallelizing Spectrally Regularized Kernel Algorithms
abstract
We consider a distributed learning approach in supervised learning for a large class of spectral regularization methods in an reproducing kernel Hilbert space (RKHS) framework. The data set of size $n$ is partitioned into $m=O(n^\alpha)$, $\alpha < \frac{1}{2}$, disjoint subsamples. On each subsample, some spectral regularization method (belonging to a large class, including in particular Kernel Ridge Regression, $L^2$-boosting and spectral cut-off) is applied. The regression function $f$ is then estimated via simple averaging, leading to a substantial reduction in computation time. We show that minimax optimal rates of convergence are preserved if $m$ grows sufficiently slowly (corresponding to an upper bound for $\alpha$) as $n \to \infty$, depending on the smoothness assumptions on $f$ and the intrinsic dimensionality. In spirit, the analysis relies on a classical bias/stochastic error analysis.
Nicole Mücke, Gilles Blanchard
J. Mach. Learn. Res.1