Zhipeng Lou

dblp:313/8152 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Optimization for machine learning · 74% Learning theory · 26%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
stochastic gradient descent
2.332025
Gaussian Approximation and Concentration of Constant Learning-Rate Stochastic Gradient Descent · NeurIPS 2025
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent · NeurIPS 2025
Beyond Sub-Gaussian Noises: Sharp Concentration Analysis for Stochastic Gradient Descent · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning
convergence analysis
0.912025
Gaussian Approximation and Concentration of Constant Learning-Rate Stochastic Gradient Descent · NeurIPS 2025
Mathematical optimization
statistical guarantees
0.912025
Statistical Guarantees for High-Dimensional Stochastic Gradient Descent · NeurIPS 2025
Machine learning › Learning theory
generalization bounds
0.612022
Beyond Sub-Gaussian Noises: Sharp Concentration Analysis for Stochastic Gradient Descent · J. Mach. Learn. Res. 2022
Machine learning › Learning theory › generalization bounds
high probability bounds
0.612022
Beyond Sub-Gaussian Noises: Sharp Concentration Analysis for Stochastic Gradient Descent · J. Mach. Learn. Res. 2022

Methods — techniques the papers use, named apart from their topics

ruppert-polyak averaging · 1.7high-dimensional time series · 1.7coupling technique · 1.7nagaev-type inequality · 0.9berry-esseen bound · 0.9autoregressive approximation · 0.9nagaev inequality · 0.6moment bounds · 0.6averaged stochastic gradient descent · 0.6
YearPublicationVenuePosition
2025 Statistical Guarantees for High-Dimensional Stochastic Gradient Descent
abstract
Stochastic Gradient Descent (SGD) and its Ruppert–Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide rigorous statistical guarantees for constant learning-rate SGD and ASGD in high-dimensional regimes. Our key innovation is to transfer powerful tools from high-dimensional time series to online learning. Specifically, by viewing SGD as a nonlinear autoregressive process and adapting existing coupling techniques, we prove the geometric-moment contraction of high-dimensional SGD for constant learning rates, thereby establishing asymptotic stationarity of the iterates. Building on this, we derive the $q$-th moment convergence of SGD and ASGD for any $q\ge2$ in general $\ell^s$-norms, and, in particular, the $\ell^{\infty}$-norm that is frequently adopted in high-dimensional sparse or structured models. Furthermore, we provide sharp high-probability concentration analysis which entails the probabilistic bound of high-dimensional ASGD. Beyond closing a critical gap in SGD theory, our proposed framework offers a novel toolkit for analyzing a broad class of high-dimensional learning algorithms.
Jiaqi Li 0032, Zhipeng Lou, Johannes Schmidt-Hieber, Wei Biao Wu
NeurIPS2
2025 Gaussian Approximation and Concentration of Constant Learning-Rate Stochastic Gradient Descent
abstract
We establish a comprehensive finite-sample and asymptotic theory for stochastic gradient descent (SGD) with constant learning rates. First, we propose a novel linear approximation technique to provide a quenched central limit theorem (CLT) for SGD iterates with refined tail properties, showing that regardless of the chosen initialization, the fluctuations of the algorithm around its target point converge to a multivariate normal distribution. Our conditions are substantially milder than those required in the classical CLTs for SGD, yet offering a stronger convergence result. Furthermore, we derive the first Berry-Esseen bound -- the Gaussian approximation error -- for the constant learning-rate SGD, which is sharp compared to the decaying learning-rate schemes in the literature. Beyond the moment convergence, we also provide the Nagaev-type inequality for the SGD tail probabilities by adopting the autoregressive approximation techniques, which entails non-asymptotic large-deviation guarantees. These results are verified via numerical simulations, paving the way for theoretically grounded uncertainty quantification, especially with non-asymptotic validity.
Ziyang Wei, Jiaqi Li 0032, Zhipeng Lou, Wei Biao Wu
NeurIPS3
2022 Beyond Sub-Gaussian Noises: Sharp Concentration Analysis for Stochastic Gradient Descent
abstract
In this paper, we study the concentration property of stochastic gradient descent (SGD) solutions. In existing concentration analyses, researchers impose restrictive requirements on the gradient noise, such as boundedness or sub-Gaussianity. We consider a much richer class of noise where only finitely-many moments are required, thus allowing heavy-tailed noises. In particular, we obtain Nagaev type high-probability upper bounds for the estimation errors of averaged stochastic gradient descent (ASGD) in a linear model. Specifically, we prove that, after $T$ steps of SGD, the ASGD estimate achieves an $O(\sqrt{\log(1/\delta)/T} + (\delta T^{q-1})^{-1/q})$ error rate with probability at least $1-\delta$, where $q>2$ controls the tail of the gradient noise. In comparison, one has the $O(\sqrt{\log(1/\delta)/T})$ error rate for sub-Gaussian noises. We also show that the Nagaev type upper bound is almost tight through an example, where the exact asymptotic form of the tail probability can be derived. Our concentration analysis indicates that, in the case of heavy-tailed noises, the polynomial dependence on the failure probability $\delta$ is generally unavoidable for the error rate of SGD.
Wanrong Zhu, Zhipeng Lou, Wei Biao Wu
J. Mach. Learn. Res.2