Yicheng Li 0004

dblp:422/5721 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Learning theory · 65% Kernel, tree and ensemble methods · 29% Deep learning architectures and training · 6%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
minimax optimality
3.752025
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025
On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression
3.652025
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025
On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024
On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023
Machine learning › Learning theory
generalization bounds
3.042024
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024
Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024
On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods
kernel methods
2.232024
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024
On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024
On the Saturation Effect of Kernel Ridge Regression · ICLR 2023
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
1.522024
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024
On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024
Machine learning › Learning theory
spectral methods
1.522024
On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024
On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space
1.422024
On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024
On the Optimality of Misspecified Kernel Ridge Regression · ICML 2023
Machine learning › Learning theory
generalization error
0.912025
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025
Machine learning › Learning theory
statistical learning theory
0.912025
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025
Machine learning › Learning theory
curse of dimensionality
0.812024
On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024
Machine learning › Learning theory › nonparametric regression
kernel regression
0.812024
Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024
Machine learning › Learning theory
over-parameterization
0.812024
Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024
Machine learning › Deep learning architectures and training › weight initialization
random initialization
0.812024
On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.812024
On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024
Machine learning › Learning theory › overfitting
benign overfitting
0.712023
On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023
Machine learning › Learning theory
generalization
0.712023
On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023
Machine learning › Learning theory
learning curves
0.712023
On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023
Mathematical optimization
regularization
0.212023
On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

gradient flow · 1.5minimax analysis · 1.4source condition · 0.9neural tangent kernel · 0.9reproducing kernel hilbert space · 0.8over-parameterization · 0.8interpolation space · 0.8gradient flow analysis · 0.8early stopping · 0.8RKHS interpolation · 0.8power-law decay · 0.7bias-variance trade-off · 0.7asymptotic analysis · 0.7
YearPublicationVenuePosition
2025 Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions
abstract
Motivated by studies of neural networks, particularly the neural tangent kernel theory, we investigate the large-dimensional behavior of kernel ridge regression, where the sample size satisfies $n $ is proportion to $ d^{\gamma}$ for some $\gamma > 0$. Given a reproducing kernel Hilbert space $H$ associated with an inner product kernel defined on the unit sphere $S^{d}$, we assume that the true function $f_{\rho}^{*}$ belongs to the interpolation space $[H]^{s}$ for some $s>0$ (source condition). We first establish the exact order (both upper and lower bounds) of the generalization error of KRR for the optimally chosen regularization parameter $\lambda$. Furthermore, we show that KRR is minimax optimal when $01$, KRR fails to achieve minimax optimality, exhibiting the saturation effect. Our results illustrate that the convergence rate with respect to dimension $d$ varying along $\gamma$ exhibits a periodic plateau behavior, and the convergence rate with respect to sample size $n$ exhibits a multiple descent behavior. Interestingly, our work unifies several recent studies on kernel regression in the large-dimensional setting, which correspond to $s=0$ and $s=1$, respectively. [abs][pdf][bib] © JMLR 2025. (edit, beta) Mastodon
Haobo Zhang 0004, Yicheng Li 0004, Weihao Lu 0002
J. Mach. Learn. Res.2
2024 On the Saturation Effects of Spectral Algorithms in Large Dimensions
abstract
The saturation effects, which originally refer to the fact that kernel ridge regression (KRR) fails to achieve the information-theoretical lower bound when the regression function is over-smooth, have been observed for almost 20 years and were rigorously proved recently for kernel ridge regression and some other spectral algorithms over a fixed dimensional domain. The main focus of this paper is to explore the saturation effects for a large class of spectral algorithms (including the KRR, gradient descent, etc.) in large dimensional settings where $n \asymp d^{\gamma}$. More precisely, we first propose an improved minimax lower bound for the kernel regression problem in large dimensional settings and show that the gradient flow with early stopping strategy will result in an estimator achieving this lower bound (up to a logarithmic factor). Similar to the results in KRR, we can further determine the exact convergence rates (both upper and lower bounds) of a large class of (optimal tuned) spectral algorithms with different qualification $\tau$'s. In particular, we find that these exact rate curves (varying along $\gamma$) exhibit the periodic plateau behavior and the polynomial approximation barrier. Consequently, we can fully depict the saturation effects of the spectral algorithms and reveal a new phenomenon in large dimensional settings (i.e., the saturation effect occurs in large dimensional setting as long as the source condition $s>\tau$ while it occurs in fixed dimensional setting as long as $s>2\tau$).
Weihao Lu 0002, Haobo Zhang 0004, Yicheng Li 0004
NeurIPS3
2024 On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory
abstract
This paper aims to discuss the impact of random initialization of neural networks in the neural tangent kernel (NTK) theory, which is ignored by most recent works in the NTK theory. It is well known that as the network's width tends to infinity, the neural network with random initialization converges to a Gaussian process \(f^{\mathrm{GP}}\), which takes values in \(L^{2}(\mathcal{X})\), where \(\mathcal{X}\) is the domain of the data. In contrast, to adopt the traditional theory of kernel regression, most recent works introduced a special mirrored architecture and a mirrored (random) initialization to ensure the network's output is identically zero at initialization. Therefore, it remains a question whether the conventional setting and mirrored initialization would make wide neural networks exhibit different generalization capabilities. In this paper, we first show that the training dynamics of the gradient flow of neural networks with random initialization converge uniformly to that of the corresponding NTK regression with random initialization \(f^{\mathrm{GP}}\). We then show that \(\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 1\) for any \(s < \frac{3}{d+1}\) and \(\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 0\) for any \(s \geq \frac{3}{d+1}\), where \([\mathcal{H}^{\mathrm{NT}}]^{s}\) is the real interpolation space of the RKHS \(\mathcal{H}^{\mathrm{NT}}\) associated with the NTK. Consequently, the generalization error of the wide neural network trained by gradient descent is \(\Omega(n^{-\frac{3}{d+3}})\), and it still suffers from the curse of dimensionality. Thus, the NTK theory may not explain the superior performance of neural networks.
Guhan Chen, Yicheng Li 0004
NeurIPS2
2024 Improving Adaptivity via Over-Parameterization in Sequence Models
abstract
It is well known that eigenfunctions of a kernel play a crucial role in kernel regression. Through several examples, we demonstrate that even with the same set of eigenfunctions, the order of these functions significantly impacts regression outcomes. Simplifying the model by diagonalizing the kernel, we introduce an over-parameterized gradient descent in the realm of sequence model to capture the effects of various orders of a fixed set of eigen-functions. This method is designed to explore the impact of varying eigenfunction orders. Our theoretical results show that the over-parameterization gradient flow can adapt to the underlying structure of the signal and significantly outperform the vanilla gradient flow method. Moreover, we also demonstrate that deeper over-parameterization can further enhance the generalization capability of the model. These results not only provide a new perspective on the benefits of over-parameterization and but also offer insights into the adaptivity and generalization potential of neural networks beyond the kernel regime.
Yicheng Li 0004
NeurIPS1
2024 On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains
abstract
In this paper, we provide a strategy to determine the eigenvalue decay rate (EDR) of a large class of kernel functions defined on a general domain rather than $\mathbb{S}^{d}$. This class of kernel functions include but are not limited to the neural tangent kernel associated with neural networks with different depths and various activation functions. After proving that the dynamics of training the wide neural networks uniformly approximated that of the neural tangent kernel regression on general domains, we can further illustrate the minimax optimality of the wide neural network provided that the underground truth function $f\in [\mathcal H_{\mathrm{NTK}}]^{s}$, an interpolation space associated with the RKHS $\mathcal{H}_{\mathrm{NTK}}$ of NTK. We also showed that the overfitted neural network can not generalize well. We believe our approach for determining the EDR of kernels might be also of independent interests.
Yicheng Li 0004, Zixiong Yu, Guhan Chen
J. Mach. Learn. Res.1
2024 On the Optimality of Misspecified Spectral Algorithms
abstract
In the misspecified spectral algorithms problem, researchers usually assume the underground true function $f_{\rho}^{*} \in [\mathcal{H}]^{s}$, a less-smooth interpolation space of a reproducing kernel Hilbert space (RKHS) $\mathcal{H}$ for some $s\in (0,1)$. The existing minimax optimal results require $\|f_{\rho}^{*}\|_{L^{\infty}}<\infty$ which implicitly requires $s > \alpha_{0}$ where $\alpha_{0}\in (0,1)$ is the embedding index, a constant depending on $\mathcal{H}$. Whether the spectral algorithms are optimal for all $s\in (0,1)$ is an outstanding problem lasting for years. In this paper, we show that spectral algorithms are minimax optimal for any $\alpha_{0}-\frac{1}{\beta} < s < 1$, where $\beta$ is the eigenvalue decay rate of $\mathcal{H}$. We also give several classes of RKHSs whose embedding index satisfies $ \alpha_0 = \frac{1}{\beta} $. Thus, the spectral algorithms are minimax optimal for all $s\in (0,1)$ on these RKHSs.
Haobo Zhang 0004, Yicheng Li 0004
J. Mach. Learn. Res.2
2023 On the Saturation Effect of Kernel Ridge Regression
Yicheng Li 0004, Haobo Zhang 0004
ICLR1
2023 On the Optimality of Misspecified Kernel Ridge Regression
abstract
In the misspecified kernel ridge regression problem, researchers usually assume the underground true function $f_{\rho}^{\star} \in [\mathcal{H}]^{s}$, a less-smooth interpolation space of a reproducing kernel Hilbert space (RKHS) $\mathcal{H}$ for some $s\in (0,1)$. The existing minimax optimal results require $\left\Vert f_{\rho}^{\star} \right \Vert_{L^{\infty}} < \infty$ which implicitly requires $s > \alpha_{0}$ where $\alpha_{0} \in (0,1) $ is the embedding index, a constant depending on $\mathcal{H}$. Whether the KRR is optimal for all $s\in (0,1)$ is an outstanding problem lasting for years. In this paper, we show that KRR is minimax optimal for any $s\in (0,1)$ when the $\mathcal{H}$ is a Sobolev RKHS.
Haobo Zhang 0004, Yicheng Li 0004, Weihao Lu 0002
ICML2
2023 On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay
abstract
The widely observed 'benign overfitting phenomenon' in the neural network literature raises the challenge to the `bias-variance trade-off' doctrine in the statistical learning theory. Since the generalization ability of the 'lazy trained' over-parametrized neural network can be well approximated by that of the neural tangent kernel regression, the curve of the excess risk (namely, the learning curve) of kernel ridge regression attracts increasing attention recently. However, most recent arguments on the learning curve are heuristic and are based on the 'Gaussian design' assumption. In this paper, under mild and more realistic assumptions, we rigorously provide a full characterization of the learning curve in the asymptotic sense under a power-law decay condition of the eigenvalues of the kernel and also the target function. The learning curve elaborates the effect and the interplay of the choice of the regularization parameter, the source condition and the noise. In particular, our results suggest that the 'benign overfitting phenomenon' exists in over-parametrized neural networks only when the noise level is small.
Yicheng Li 0004, Haobo Zhang 0004
NeurIPS1