Zheng-Chu Guo

dblp:94/9816 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
2since 2021 · last 2025
0009-0001-7134-6725ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 1 since 2021Theory of computation · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
abstract
Recent advances have significantly improved our understanding of the generalization performance of gradient descent (GD) methods in deep neural networks. A natural and fundamental question is whether GD can achieve generalization rates comparable to the minimax optimal rates established in the kernel setting. Existing results either yield suboptimal rates of $O(1/\sqrt{n})$, or focus on networks with smooth activation functions, incurring exponential dependence on network depth $L$. In this work, we establish optimal generalization rates for GD with deep ReLU networks by carefully trading off optimization and generalization errors, achieving only polynomial dependence on depth. Specifically, under the assumption that the data are NTK separable from the margin $\gamma$, we prove an excess risk rate of $\widetilde{O}(L^4 (1 + \gamma L^2) / (n \gamma^2))$, which aligns with the optimal SVM-type rate $\widetilde{O}(1 / (n \gamma^2))$ up to depth-dependent factors. A key technical contribution is our novel control of activation patterns near a reference model, enabling a sharper Rademacher complexity bound for deep ReLU networks trained with gradient descent.
Yuanfan Li, Yunwen Lei, Zheng-Chu Guo, Yiming Ying
NeurIPS3
2024 Online regularized learning algorithm for functional data
Yuan Mao, Zheng-Chu Guo
J. Complex.2
2020 Optimal learning rates for distribution regression
Zhiying Fang, Zheng-Chu Guo, Ding-Xuan Zhou
J. Complex.2
2020 Realizing Data Features by Deep Nets
abstract
This article considers the power of deep neural networks (deep nets) in realizing data features. Based on refined covering number estimates, we find that, to realize data features such as the locality, rotation invariance, and manifold structure, deep nets essentially improve the performances of shallow neural networks (shallow nets) without requiring additional capacity costs. Conversely, to realize some data features, such as the smoothness, we show that deep nets perform similar as shallow nets, provided the depth is not extremely large. Both sides show the advantages and limitations of deep nets in realizing data features and demonstrate that deep nets are not always better than shallow nets.
Zheng-Chu Guo, Lei Shi 0010, Shaobo Lin
IEEE Trans. Neural Networks Learn. Syst.1
2017 Learning from Networked Examples
abstract
Many machine learning algorithms are based on the assumption that training examples are drawn independently. However, this assumption does not hold anymore when learning from a networked sample because two or more training examples may share some common objects, and hence share the features of these shared objects. We show that the classic approach of ignoring this problem potentially can have a harmful effect on the accuracy of statistics, and then consider alternatives. One of these is to only use independent examples, discarding other information. However, this is clearly suboptimal. We analyze sample error bounds in this networked setting, providing significantly improved results. An important component of our approach is formed by efficient sample weighting schemes, which leads to novel concentration inequalities.
Yuyi Wang 0001, Zheng-Chu Guo, Jan Ramon
ALT2
2017 Learning Theory of Distributed Regression with Bias Corrected Regularization Kernel Network
abstract
Distributed learning is an effective way to analyze big data. In distributed regression, a typical approach is to divide the big data into multiple blocks, apply a base regression algorithm on each of them, and then simply average the output functions learnt from these blocks. Since the average process will decrease the variance, not the bias, bias correction is expected to improve the learning performance if the base regression algorithm is a biased one. Regularization kernel network is an effective and widely used method for nonlinear regression analysis. In this paper we will investigate a bias corrected version of regularization kernel network. We derive the error bounds when it is applied to a single data set and when it is applied as a base algorithm in distributed regression. We show that, under certain appropriate conditions, the optimal learning rates can be reached in both situations.
Zheng-Chu Guo, Lei Shi 0010, Qiang Wu 0003
J. Mach. Learn. Res.1
2017 Convergence of Unregularized Online Learning Algorithms
Yunwen Lei, Lei Shi 0010, Zheng-Chu Guo
J. Mach. Learn. Res.3
2016 Generalization bounds for metric and similarity learning
Qiong Cao, Zheng-Chu Guo, Yiming Ying
Mach. Learn.2
2014 Guaranteed Classification via Regularized Similarity Learning
abstract
Learning an appropriate (dis)similarity function from the available data is a central problem in machine learning, since the success of many machine learning algorithms critically depends on the choice of a similarity function to compare examples. Despite many approaches to similarity metric learning that have been proposed, there has been little theoretical study on the links between similarity metric learning and the classification performance of the resulting classifier. In this letter, we propose a regularized similarity learning formulation associated with general matrix norms and establish their generalization bounds. We show that the generalization error of the resulting linear classifier can be bounded by the derived generalization bound of similarity learning. This shows that a good generalization of the learned similarity function guarantees a good classification of the resulting linear classifier. Our results extend and improve those obtained by Bellet, Habrard, and Sebban (2012). Due to the techniques dependent on the notion of uniform stability (Bousquet & Elisseeff, 2002), the bound obtained there holds true only for the Frobenius matrix-norm regularization. Our techniques using the Rademacher complexity (Bartlett & Mendelson, 2002) and its related Khinchin-type inequality enable us to establish bounds for regularized similarity learning formulations associated with general matrix norms, including sparse L1-norm and mixed (2,1)-norm.
Zheng-Chu Guo, Yiming Ying
Neural Comput.1