VLDB 2026 Research / reviewers in the wild / expert
Yuqing Wang 0005
dblp:60/5086-5
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0004-9922-3067ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Optimization for machine learning · 52% Deep learning architectures and training · 19% Generative modeling · 15% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Algorithms and data structures · 50% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › training dynamics
edge of stability |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Learning theory
implicit bias |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
accelerated gradient methods |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Optimization for machine learning
convergence analysis |
0.8 | 1 | 2024 | Evaluating the design space of diffusion-based generative models · NeurIPS 2024 |
Machine learning › Generative modeling › score matching
denoising score matching |
0.8 | 1 | 2024 | Evaluating the design space of diffusion-based generative models · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Evaluating the design space of diffusion-based generative models · NeurIPS 2024 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Mathematical optimization
nonconvex optimization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.7 | 1 | 2023 | Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport · ICLR 2023 |
Machine learning › Optimization for machine learning
riemannian optimization |
0.7 | 1 | 2023 | Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport · ICLR 2023 |
Machine learning › Optimization for machine learning › riemannian optimization
stiefel manifold optimization |
0.7 | 1 | 2023 | Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport · ICLR 2023 |
Machine learning › Optimization for machine learning
learning rate |
0.6 | 1 | 2022 | Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect · ICLR 2022 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.4 | 1 | 2020 | Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel Perspective · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › convolutional neural network
residual network |
0.4 | 1 | 2020 | Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel Perspective · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
0.2 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Optimization for machine learning
optimal transport |
0.2 | 1 | 2023 | Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
gradient descent · 2.3unbalanced initialization · 1.5nesterov's accelerated gradient · 1.5convergence analysis · 1.4lyapunov analysis · 0.9non-asymptotic convergence analysis · 0.8stiefel manifold optimization · 0.7momentum optimization · 0.7theoretical analysis · 0.4neural tangent kernel · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapultabstractLarge learning rates, when applied to gradient descent for nonconvex optimization, yield various implicit biases including the edge of stability, balancing, and catapult. These phenomena cannot be well explained by classical optimization theory. Though significant theoretical progress has been made in understanding these implicit biases, it remains unclear for which objective functions they are more likely to occur --- more precisely, for which functions there exists a larger set of initial conditions that lead to these phenomena? This paper provides an initial step in answering this question and also shows that these implicit biases are in fact various tips of the same iceberg. To establish these results, we develop a global convergence theory under large learning rates, for a family of nonconvex functions without globally Lipschitz continuous gradient, which was typically assumed in existing convergence analysis. Specifically, these phenomena are more likely to occur when the optimization objective function has good regularity. This regularity, together with gradient descent using a large learning rate that favors flatter regions, results in these nontrivial dynamical behaviors. Another corollary is the first non-asymptotic convergence rate bound for large-learning-rate gradient descent optimization of nonconvex functions. Although our theory only applies to specific functions so far, the possibility of extrapolating it to neural networks is also experimentally validated, for which different choices of loss, activation functions, and other techniques such as batch normalization can all affect regularity significantly and lead to very different training dynamics. Yuqing Wang 0005, Zhenghao Xu, Tuo Zhao, Molei Tao |
J. Mach. Learn. Res. | 1 |
| 2024 | Evaluating the design space of diffusion-based generative modelsabstractMost existing theoretical investigations of the accuracy of diffusion models, albeit significant, assume the score function has been approximated to a certain accuracy, and then use this a priori bound to control the error of generation. This article instead provides a first quantitative understanding of the whole generation process, i.e., both training and sampling. More precisely, it conducts a non-asymptotic convergence analysis of denoising score matching under gradient descent. In addition, a refined sampling error analysis for variance exploding models is also provided. The combination of these two results yields a full error analysis, which elucidates (again, but this time theoretically) how to design the training and sampling processes for effective generation. For instance, our theory implies a preference toward noise distribution and loss weighting in training that qualitatively agree with the ones used in [Karras et al., 2022]. It also provides perspectives on the choices of time and variance schedules in sampling: when the score is well trained, the design in [Song et al., 2021] is more preferable, but when it is less trained, the design in [Karras et al., 2022] becomes more preferable. Yuqing Wang 0005, Ye He 0003, Molei Tao |
NeurIPS | 1 |
| 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural NetworksabstractWe study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solutions $\mathbf{X}_T\in\mathbb{R}^{m\times d}$ and $\mathbf{Y}_T\in\mathbb{R}^{n\times d}$, where $d\geq r$, satisfying $\lVert\mathbf{X}_T\mathbf{Y}_T^\top-\mathbf{A}\rVert_F\leq\epsilon\lVert\mathbf{A}\rVert_F$ in $T=O(\kappa^2\log\frac{1}{\epsilon})$ iterations with high probability, where $\kappa$ denotes the condition number of $\mathbf{A}$. Furthermore, we prove that Nesterov's accelerated gradient (NAG) attains an iteration complexity of $O(\kappa\log\frac{1}{\epsilon})$, which is the best-known bound of first-order methods for rectangular matrix factorization. Different from small balanced random initialization in the existing literature, we adopt an unbalanced initialization, where $\mathbf{X}_0$ is large and $\mathbf{Y}_0$ is $0$. Moreover, our initialization and analysis can be further extended to linear neural networks, where we prove that NAG can also attain an accelerated linear convergence rate. In particular, we only require the width of the network to be greater than or equal to the rank of the output label matrix. In contrast, previous results achieving the same rate require excessive widths that additionally depend on the condition number and the rank of the input data matrix. Zhenghao Xu, Yuqing Wang 0005, Tuo Zhao, Rachel Ward, Molei Tao |
NeurIPS | 2 |
| 2023 | Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport
Yuqing Wang 0005, Molei Tao |
ICLR | 2 |
| 2022 | Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
Yuqing Wang 0005, Minshuo Chen, Tuo Zhao, Molei Tao |
ICLR | 1 |
| 2020 | Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveabstractDeep residual networks (ResNets) have demonstrated better generalization performance than deep feedforward networks (FFNets). However, the theory behind such a phenomenon is still largely unknown. This paper studies this fundamental problem in deep learning from a so-called ``neural tangent kernel'' perspective. Specifically, we first show that under proper conditions, as the width goes to infinity, training deep ResNets can be viewed as learning reproducing kernel functions with some kernel function. We then compare the kernel of deep ResNets with that of deep FFNets and discover that the class of functions induced by the kernel of FFNets is asymptotically not learnable, as the depth goes to infinity. In contrast, the class of functions induced by the kernel of ResNets does not exhibit such degeneracy. Our discovery partially justifies the advantages of deep ResNets over deep FFNets in generalization abilities. Numerical results are provided to support our claim. Kaixuan Huang, Yuqing Wang 0005, Molei Tao, Tuo Zhao |
NeurIPS | 2 |