Johannes Schmidt-Hieber

dblp:205/3078 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-2699-4990ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Optimization for machine learning · 27% Probabilistic and Bayesian machine learning · 25% Deep learning architectures and training · 22%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
generalization bounds
1.522025
On the VC dimension of deep group convolutional neural networks · NeurIPS 2025
On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization · IEEE Trans. Inf. Theory 2023
Machine learning › Optimization for machine learning
stochastic gradient descent
1.522025
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent · NeurIPS 2025
On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization · IEEE Trans. Inf. Theory 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process › hierarchical gaussian process
deep gaussian process
1.222023
Posterior Contraction for Deep Gaussian Process Priors · J. Mach. Learn. Res. 2023
On the inability of Gaussian process regression to optimally learn compositional functions · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian asymptotics
posterior contraction rates
1.222023
Posterior Contraction for Deep Gaussian Process Priors · J. Mach. Learn. Res. 2023
On the inability of Gaussian process regression to optimally learn compositional functions · NeurIPS 2022
Machine learning › Optimization for machine learning
convergence analysis
0.912025
Spike-timing-dependent Hebbian learning as noisy gradient descent · NeurIPS 2025
Machine learning › Deep learning architectures and training › equivariant neural network
group convolutional network
0.912025
On the VC dimension of deep group convolutional neural networks · NeurIPS 2025
Machine learning › Representation and self-supervised learning
hebbian learning
0.912025
Spike-timing-dependent Hebbian learning as noisy gradient descent · NeurIPS 2025
Machine learning › Optimization for machine learning › stochastic optimization
noisy gradient descent
0.912025
Spike-timing-dependent Hebbian learning as noisy gradient descent · NeurIPS 2025
Machine learning › Deep learning architectures and training › spiking neural network
spike-timing-dependent plasticity
0.912025
Spike-timing-dependent Hebbian learning as noisy gradient descent · NeurIPS 2025
Machine learning › Learning theory › computational learning theory › VC theory
VC dimension
0.912025
On the VC dimension of deep group convolutional neural networks · NeurIPS 2025
Mathematical optimization
statistical guarantees
0.912025
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent · NeurIPS 2025
Machine learning › Deep learning architectures and training › regularization
dropout
0.812024
Dropout Regularization Versus l2-Penalization in the Linear Model · J. Mach. Learn. Res. 2024
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.812024
Dropout Regularization Versus l2-Penalization in the Linear Model · J. Mach. Learn. Res. 2024
Machine learning › Deep learning architectures and training
regularization
0.812024
Dropout Regularization Versus l2-Penalization in the Linear Model · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.712023
Posterior Contraction for Deep Gaussian Process Priors · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process prior
0.712023
Posterior Contraction for Deep Gaussian Process Priors · J. Mach. Learn. Res. 2023
Machine learning › Optimization for machine learning
implicit regularization
0.712023
On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization · IEEE Trans. Inf. Theory 2023
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry
0.712023
On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization · IEEE Trans. Inf. Theory 2023
Machine learning › Learning theory
statistical learning theory
0.712023
Posterior Contraction for Deep Gaussian Process Priors · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.612022
On the inability of Gaussian process regression to optimally learn compositional functions · NeurIPS 2022
Machine learning › Learning theory › statistical estimation › minimax estimation
minimax rates
0.612022
On the inability of Gaussian process regression to optimally learn compositional functions · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

ruppert-polyak averaging · 1.7high-dimensional time series · 1.7coupling technique · 1.7stochastic approximation · 0.9non-convex optimization · 0.9equivariant neural network · 0.9ReLU activation · 0.9l2-penalization · 0.8gradient descent · 0.8dropout · 0.8
YearPublicationVenuePosition
2025 Understanding the Effect of GCN Convolutions in Regression Tasks
abstract
Graph Convolutional Networks (GCNs) have become a pivotal method in machine learning for modeling functions over graphs. Despite their widespread success across various applications, their statistical properties (e.g., consistency, convergence rates) remain ill-characterized. To begin addressing this knowledge gap, we consider networks for which the graph structure implies that neighboring nodes exhibit similar signals and provide statistical theory for the impact of convolution operators. Focusing on estimators based solely on neighborhood aggregation, we examine how two common convolutions—the original GCN and GraphSAGE convolutions—affect the learning error as a function of the neighborhood topology and the number of convolutional layers. We explicitly characterize the bias variance type trade-off incurred by GCNs as a function of the neighborhood size and identify specific graph topologies where convolution operators are less effective. Our theoretical findings are corroborated by synthetic experiments, and provide a start to a deeper quantitative understanding of convolutional effects in GCNs for offering rigorous guidelines for practitioners.
Juntong Chen, Johannes Schmidt-Hieber, Claire Donnat, Olga Klopp
AISTATS2
2025 Spike-timing-dependent Hebbian learning as noisy gradient descent
abstract
Hebbian learning is a key principle underlying learning in biological neural networks. We relate a Hebbian spike-timing-dependent plasticity rule to noisy gradient descent with respect to a non-convex loss function on the probability simplex. Despite the constant injection of noise and the non-convexity of the underlying optimization problem, one can rigorously prove that the considered Hebbian learning dynamic identifies the presynaptic neuron with the highest activity and that the convergence is exponentially fast in the number of iterations. This is non-standard and surprising as typically noisy gradient descent with fixed noise level only converges to a stationary regime where the noise causes the dynamic to fluctuate around a minimiser.
Niklas Dexheimer, Sascha Gaudlitz, Johannes Schmidt-Hieber
NeurIPS3
2025 Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
abstract
Stochastic Gradient Descent (SGD) and its Ruppert–Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide rigorous statistical guarantees for constant learning-rate SGD and ASGD in high-dimensional regimes. Our key innovation is to transfer powerful tools from high-dimensional time series to online learning. Specifically, by viewing SGD as a nonlinear autoregressive process and adapting existing coupling techniques, we prove the geometric-moment contraction of high-dimensional SGD for constant learning rates, thereby establishing asymptotic stationarity of the iterates. Building on this, we derive the $q$-th moment convergence of SGD and ASGD for any $q\ge2$ in general $\ell^s$-norms, and, in particular, the $\ell^{\infty}$-norm that is frequently adopted in high-dimensional sparse or structured models. Furthermore, we provide sharp high-probability concentration analysis which entails the probabilistic bound of high-dimensional ASGD. Beyond closing a critical gap in SGD theory, our proposed framework offers a novel toolkit for analyzing a broad class of high-dimensional learning algorithms.
Jiaqi Li 0032, Zhipeng Lou, Johannes Schmidt-Hieber, Wei Biao Wu
NeurIPS3
2025 On the VC dimension of deep group convolutional neural networks
abstract
Recent works have introduced new equivariant neural networks, motivated by their improved generalization compared to traditional deep neural networks. While experiments support this advantage, the theoretical understanding of their generalization properties remains limited. In this paper, we analyze the generalization capabilities of Group Convolutional Neural Networks (GCNNs) with the ReLU activation function through the lens of Vapnik-Chervonenkis (VC) dimension theory. We investigate how architectural factors—such as the number of layers, weights, and input dimensions—affect the VC dimension. A key challenge in our analysis is proving a lower bound on the VC dimension, for which we introduce new techniques, establishing a novel connection between GCNNs and standard deep neural networks. Additionally, we compare our derived bounds to those known for fully connected neural networks. Our results extend previous findings on the VC dimension of continuous GCNNs with two layers, offering new insights into their generalization behavior, particularly their dependence on input resolution.
Anna Sepliarskaia, Sophie Langer, Johannes Schmidt-Hieber
NeurIPS3
2024 Dropout Regularization Versus l2-Penalization in the Linear Model
abstract
We investigate the statistical behavior of gradient descent iterates with dropout in the linear regression model. In particular, non-asymptotic bounds for the convergence of expectations and covariance matrices of the iterates are derived. The results shed more light on the widely cited connection between dropout and $\ell_2$-regularization in the linear model. We indicate a more subtle relationship, owing to interactions between the gradient descent dynamics and the additional randomness induced by dropout. Further, we study a simplified variant of dropout which does not have a regularizing effect and converges to the least squares estimator.
Gabriel Clara, Sophie Langer, Johannes Schmidt-Hieber
J. Mach. Learn. Res.3
2023 Posterior Contraction for Deep Gaussian Process Priors
abstract
We study posterior contraction rates for a class of deep Gaussian process priors in the nonparametric regression setting under a general composition assumption on the regression function. It is shown that the contraction rates can achieve the minimax convergence rate (up to log n factors), while being adaptive to the underlying structure and smoothness of the target function. The proposed framework extends the Bayesian nonparametric theory for Gaussian process priors.
Gianluca Finocchio, Johannes Schmidt-Hieber
J. Mach. Learn. Res.2
2023 On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization
abstract
The classical statistical learning theory implies that fitting too many parameters leads to overfitting and poor performance. That modern deep neural networks generalize well despite a large number of parameters contradicts this finding and constitutes a major unsolved problem towards explaining the success of deep learning. While previous work focuses on the implicit regularization induced by stochastic gradient descent (SGD), we study here how the local geometry of the energy landscape around local minima affects the statistical properties of SGD with Gaussian gradient noise. We argue that under reasonable assumptions, the local geometry forces SGD to stay close to a low dimensional subspace and that this induces another form of implicit regularization and results in tighter bounds on the generalization error for deep neural networks. To derive generalization error bounds for neural networks, we first introduce a notion of stagnation sets around the local minima and impose a local essential convexity property of the population risk. Under these conditions, lower bounds for SGD to remain in these stagnation sets are derived. If stagnation occurs, we derive a bound on the generalization error of deep neural networks involving the spectral norms of the weight matrices but not the number of network parameters. Technically, our proofs are based on controlling the change of parameter values in the SGD iterates and local uniform convergence of the empirical loss functions based on the entropy of suitable neighborhoods around local minima.
Masaaki Imaizumi, Johannes Schmidt-Hieber
IEEE Trans. Inf. Theory2
2022 On the inability of Gaussian process regression to optimally learn compositional functions
abstract
We rigorously prove that deep Gaussian process priors can outperform Gaussian process priors if the target function has a compositional structure. To this end, we study information-theoretic lower bounds for posterior contraction rates for Gaussian process regression in a continuous regression model. We show that if the true function is a generalized additive function, then the posterior based on any mean-zero Gaussian process can only recover the truth at a rate that is strictly slower than the minimax rate by a factor that is polynomially suboptimal in the sample size $n$.
Matteo Giordano, Kolyan Ray, Johannes Schmidt-Hieber
NeurIPS3
2021 The Kolmogorov-Arnold representation theorem revisited
abstract
There is a longstanding debate whether the Kolmogorov-Arnold representation theorem can explain the use of more than one hidden layer in neural networks. The Kolmogorov-Arnold representation decomposes a multivariate function into an interior and an outer function and therefore has indeed a similar structure as a neural network with two hidden layers. But there are distinctive differences. One of the main obstacles is that the outer function depends on the represented function and can be wildly varying even if the represented function is smooth. We derive modifications of the Kolmogorov-Arnold representation that transfer smoothness properties of the represented function to the outer function and can be well approximated by ReLU networks. It appears that instead of two hidden layers, a more natural interpretation of the Kolmogorov-Arnold representation is that of a deep neural network where most of the layers are required to approximate the interior function.
Johannes Schmidt-Hieber
Neural Networks1
2019 A comparison of deep networks with ReLU activation function and linear spline-type methods
abstract
Deep neural networks (DNNs) generate much richer function spaces than shallow networks. Since the function spaces induced by shallow networks have several approximation theoretic drawbacks, this explains, however, not necessarily the success of deep networks. In this article we take another route by comparing the expressive power of DNNs with ReLU activation function to linear spline methods. We show that MARS (multivariate adaptive regression splines) is improper learnable by DNNs in the sense that for any given function that can be expressed as a function in MARS with M parameters there exists a multilayer neural network with O(Mlog(M∕ε)) parameters that approximates this function up to sup-norm error ε. We show a similar result for expansions with respect to the Faber-Schauder system. Based on this, we derive risk comparison inequalities that bound the statistical risk of fitting a neural network by the statistical risk of spline-based methods. This shows that deep networks perform better or only slightly worse than the considered spline methods. We provide a constructive proof for the function approximations.
Konstantin Eckle, Johannes Schmidt-Hieber
Neural Networks2