VLDB 2026 Research / reviewers in the wild / expert
Yicheng Li 0004
dblp:422/5721
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Learning theory · 65% Kernel, tree and ensemble methods · 29% Deep learning architectures and training · 6% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
minimax optimality |
3.7 | 5 | 2025 | Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025 On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024 On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression |
3.6 | 5 | 2025 | Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025 On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024 On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023 |
Machine learning › Learning theory
generalization bounds |
3.0 | 4 | 2024 | On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024 Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024 On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
2.2 | 3 | 2024 | On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024 On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024 On the Saturation Effect of Kernel Ridge Regression · ICLR 2023 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
1.5 | 2 | 2024 | On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024 On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024 |
Machine learning › Learning theory
spectral methods |
1.5 | 2 | 2024 | On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024 On the Saturation Effects of Spectral Algorithms in Large Dimensions · NeurIPS 2024 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
1.4 | 2 | 2024 | On the Optimality of Misspecified Spectral Algorithms · J. Mach. Learn. Res. 2024 On the Optimality of Misspecified Kernel Ridge Regression · ICML 2023 |
Machine learning › Learning theory
generalization error |
0.9 | 1 | 2025 | Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025 |
Machine learning › Learning theory
statistical learning theory |
0.9 | 1 | 2025 | Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions · J. Mach. Learn. Res. 2025 |
Machine learning › Learning theory
curse of dimensionality |
0.8 | 1 | 2024 | On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024 |
Machine learning › Learning theory › nonparametric regression
kernel regression |
0.8 | 1 | 2024 | Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024 |
Machine learning › Learning theory
over-parameterization |
0.8 | 1 | 2024 | Improving Adaptivity via Over-Parameterization in Sequence Models · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › weight initialization
random initialization |
0.8 | 1 | 2024 | On the Impacts of the Random Initialization in the Neural Tangent Kernel Theory · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks |
0.8 | 1 | 2024 | On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General Domains · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory › overfitting
benign overfitting |
0.7 | 1 | 2023 | On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023 |
Machine learning › Learning theory
generalization |
0.7 | 1 | 2023 | On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023 |
Machine learning › Learning theory
learning curves |
0.7 | 1 | 2023 | On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023 |
Mathematical optimization
regularization |
0.2 | 1 | 2023 | On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law Decay · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
gradient flow · 1.5minimax analysis · 1.4source condition · 0.9neural tangent kernel · 0.9reproducing kernel hilbert space · 0.8over-parameterization · 0.8interpolation space · 0.8gradient flow analysis · 0.8early stopping · 0.8RKHS interpolation · 0.8power-law decay · 0.7bias-variance trade-off · 0.7asymptotic analysis · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimal Rates of Kernel Ridge Regression under Source Condition in Large DimensionsabstractMotivated by studies of neural networks, particularly the neural tangent kernel theory, we investigate the large-dimensional behavior of kernel ridge regression, where the sample size satisfies $n $ is proportion to $ d^{\gamma}$ for some $\gamma > 0$. Given a reproducing kernel Hilbert space $H$ associated with an inner product kernel defined on the unit sphere $S^{d}$, we assume that the true function $f_{\rho}^{*}$ belongs to the interpolation space $[H]^{s}$ for some $s>0$ (source condition). We first establish the exact order (both upper and lower bounds) of the generalization error of KRR for the optimally chosen regularization parameter $\lambda$. Furthermore, we show that KRR is minimax optimal when $01$, KRR fails to achieve minimax optimality, exhibiting the saturation effect. Our results illustrate that the convergence rate with respect to dimension $d$ varying along $\gamma$ exhibits a periodic plateau behavior, and the convergence rate with respect to sample size $n$ exhibits a multiple descent behavior. Interestingly, our work unifies several recent studies on kernel regression in the large-dimensional setting, which correspond to $s=0$ and $s=1$, respectively. [abs][pdf][bib] © JMLR 2025. (edit, beta) Mastodon Haobo Zhang 0004, Yicheng Li 0004, Weihao Lu 0002 |
J. Mach. Learn. Res. | 2 |
| 2024 | On the Saturation Effects of Spectral Algorithms in Large DimensionsabstractThe saturation effects, which originally refer to the fact that kernel ridge regression (KRR) fails to achieve the information-theoretical lower bound when the regression function is over-smooth, have been observed for almost 20 years and were rigorously proved recently for kernel ridge regression and some other spectral algorithms over a fixed dimensional domain. The main focus of this paper is to explore the saturation effects for a large class of spectral algorithms (including the KRR, gradient descent, etc.) in large dimensional settings where $n \asymp d^{\gamma}$. More precisely, we first propose an improved minimax lower bound for the kernel regression problem in large dimensional settings and show that the gradient flow with early stopping strategy will result in an estimator achieving this lower bound (up to a logarithmic factor). Similar to the results in KRR, we can further determine the exact convergence rates (both upper and lower bounds) of a large class of (optimal tuned) spectral algorithms with different qualification $\tau$'s. In particular, we find that these exact rate curves (varying along $\gamma$) exhibit the periodic plateau behavior and the polynomial approximation barrier. Consequently, we can fully depict the saturation effects of the spectral algorithms and reveal a new phenomenon in large dimensional settings (i.e., the saturation effect occurs in large dimensional setting as long as the source condition $s>\tau$ while it occurs in fixed dimensional setting as long as $s>2\tau$). Weihao Lu 0002, Haobo Zhang 0004, Yicheng Li 0004 |
NeurIPS | 3 |
| 2024 | On the Impacts of the Random Initialization in the Neural Tangent Kernel TheoryabstractThis paper aims to discuss the impact of random initialization of neural networks in the neural tangent kernel (NTK) theory, which is ignored by most recent works in the NTK theory. It is well known that as the network's width tends to infinity, the neural network with random initialization converges to a Gaussian process \(f^{\mathrm{GP}}\), which takes values in \(L^{2}(\mathcal{X})\), where \(\mathcal{X}\) is the domain of the data. In contrast, to adopt the traditional theory of kernel regression, most recent works introduced a special mirrored architecture and a mirrored (random) initialization to ensure the network's output is identically zero at initialization. Therefore, it remains a question whether the conventional setting and mirrored initialization would make wide neural networks exhibit different generalization capabilities. In this paper, we first show that the training dynamics of the gradient flow of neural networks with random initialization converge uniformly to that of the corresponding NTK regression with random initialization \(f^{\mathrm{GP}}\). We then show that \(\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 1\) for any \(s < \frac{3}{d+1}\) and \(\mathbf{P}(f^{\mathrm{GP}} \in [\mathcal{H}^{\mathrm{NT}}]^{s}) = 0\) for any \(s \geq \frac{3}{d+1}\), where \([\mathcal{H}^{\mathrm{NT}}]^{s}\) is the real interpolation space of the RKHS \(\mathcal{H}^{\mathrm{NT}}\) associated with the NTK. Consequently, the generalization error of the wide neural network trained by gradient descent is \(\Omega(n^{-\frac{3}{d+3}})\), and it still suffers from the curse of dimensionality. Thus, the NTK theory may not explain the superior performance of neural networks. Guhan Chen, Yicheng Li 0004 |
NeurIPS | 2 |
| 2024 | Improving Adaptivity via Over-Parameterization in Sequence ModelsabstractIt is well known that eigenfunctions of a kernel play a crucial role in kernel regression.
Through several examples, we demonstrate that even with the same set of eigenfunctions, the order of these functions significantly impacts regression outcomes.
Simplifying the model by diagonalizing the kernel, we introduce an over-parameterized gradient descent in the realm of sequence model to capture the effects of various orders of a fixed set of eigen-functions.
This method is designed to explore the impact of varying eigenfunction orders.
Our theoretical results show that the over-parameterization gradient flow can adapt to the underlying structure of the signal and significantly outperform the vanilla gradient flow method.
Moreover, we also demonstrate that deeper over-parameterization can further enhance the generalization capability of the model.
These results not only provide a new perspective on the benefits of over-parameterization and but also offer insights into the adaptivity and generalization potential of neural networks beyond the kernel regime. Yicheng Li 0004 |
NeurIPS | 1 |
| 2024 | On the Eigenvalue Decay Rates of a Class of Neural-Network Related Kernel Functions Defined on General DomainsabstractIn this paper, we provide a strategy to determine the eigenvalue decay rate (EDR) of a large class of kernel functions defined on a general domain rather than $\mathbb{S}^{d}$. This class of kernel functions include but are not limited to the neural tangent kernel associated with neural networks with different depths and various activation functions. After proving that the dynamics of training the wide neural networks uniformly approximated that of the neural tangent kernel regression on general domains, we can further illustrate the minimax optimality of the wide neural network provided that the underground truth function $f\in [\mathcal H_{\mathrm{NTK}}]^{s}$, an interpolation space associated with the RKHS $\mathcal{H}_{\mathrm{NTK}}$ of NTK. We also showed that the overfitted neural network can not generalize well. We believe our approach for determining the EDR of kernels might be also of independent interests. Yicheng Li 0004, Zixiong Yu, Guhan Chen |
J. Mach. Learn. Res. | 1 |
| 2024 | On the Optimality of Misspecified Spectral AlgorithmsabstractIn the misspecified spectral algorithms problem, researchers usually assume the underground true function $f_{\rho}^{*} \in [\mathcal{H}]^{s}$, a less-smooth interpolation space of a reproducing kernel Hilbert space (RKHS) $\mathcal{H}$ for some $s\in (0,1)$. The existing minimax optimal results require $\|f_{\rho}^{*}\|_{L^{\infty}}<\infty$ which implicitly requires $s > \alpha_{0}$ where $\alpha_{0}\in (0,1)$ is the embedding index, a constant depending on $\mathcal{H}$. Whether the spectral algorithms are optimal for all $s\in (0,1)$ is an outstanding problem lasting for years. In this paper, we show that spectral algorithms are minimax optimal for any $\alpha_{0}-\frac{1}{\beta} < s < 1$, where $\beta$ is the eigenvalue decay rate of $\mathcal{H}$. We also give several classes of RKHSs whose embedding index satisfies $ \alpha_0 = \frac{1}{\beta} $. Thus, the spectral algorithms are minimax optimal for all $s\in (0,1)$ on these RKHSs. Haobo Zhang 0004, Yicheng Li 0004 |
J. Mach. Learn. Res. | 2 |
| 2023 | On the Saturation Effect of Kernel Ridge Regression
Yicheng Li 0004, Haobo Zhang 0004 |
ICLR | 1 |
| 2023 | On the Optimality of Misspecified Kernel Ridge RegressionabstractIn the misspecified kernel ridge regression problem, researchers usually assume the underground true function $f_{\rho}^{\star} \in [\mathcal{H}]^{s}$, a less-smooth interpolation space of a reproducing kernel Hilbert space (RKHS) $\mathcal{H}$ for some $s\in (0,1)$. The existing minimax optimal results require $\left\Vert f_{\rho}^{\star} \right \Vert_{L^{\infty}} < \infty$ which implicitly requires $s > \alpha_{0}$ where $\alpha_{0} \in (0,1) $ is the embedding index, a constant depending on $\mathcal{H}$. Whether the KRR is optimal for all $s\in (0,1)$ is an outstanding problem lasting for years. In this paper, we show that KRR is minimax optimal for any $s\in (0,1)$ when the $\mathcal{H}$ is a Sobolev RKHS. Haobo Zhang 0004, Yicheng Li 0004, Weihao Lu 0002 |
ICML | 2 |
| 2023 | On the Asymptotic Learning Curves of Kernel Ridge Regression under Power-law DecayabstractThe widely observed 'benign overfitting phenomenon' in the neural network literature raises the challenge to the `bias-variance trade-off' doctrine in the statistical learning theory.
Since the generalization ability of the 'lazy trained' over-parametrized neural network can be well approximated by that of the neural tangent kernel regression,
the curve of the excess risk (namely, the learning curve) of kernel ridge regression attracts increasing attention recently.
However, most recent arguments on the learning curve are heuristic and are based on the 'Gaussian design' assumption.
In this paper, under mild and more realistic assumptions, we rigorously provide a full characterization of the learning curve in the asymptotic sense
under a power-law decay condition of the eigenvalues of the kernel and also the target function.
The learning curve elaborates the effect and the interplay of the choice of the regularization parameter, the source condition and the noise.
In particular, our results suggest that the 'benign overfitting phenomenon' exists in over-parametrized neural networks only when the noise level is small. Yicheng Li 0004, Haobo Zhang 0004 |
NeurIPS | 1 |