EDBT 2026 Demo / reviewers in the wild / expert
Yanming Lai
dblp:286/8790
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0008-5510-8760ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Optimization for machine learning · 67% Deep learning architectures and training · 33% | |
| Theoretical computer science
1 paper |
Approximation and online algorithms · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.9 | 1 | 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz Method · IEEE Trans. Inf. Theory 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
projected gradient descent |
0.9 | 1 | 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz Method · IEEE Trans. Inf. Theory 2025 |
Machine learning › Deep learning architectures and training › feedforward neural network › multilayer neural network
three-layer neural network |
0.9 | 1 | 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz Method · IEEE Trans. Inf. Theory 2025 |
Approximation and online algorithms
approximation |
0.3 | 1 | 2026 | Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective · J. Mach. Learn. Res. 2026 |
Methods — techniques the papers use, named apart from their topics
self-attention · 1.0kolmogorov-arnold superposition theorem · 1.0optimization error · 0.9generalization error · 0.9error analysis · 0.9approximation error · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation PerspectiveabstractThe Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of feedforward neural networks to Transformer research. Yuling Jiao, Yanming Lai, Yang Wang 0020, Bokai Yan |
J. Mach. Learn. Res. | 2 |
| 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz MethodabstractMachine learning is a rapidly advancing field with diverse applications across various domains. One prominent area of research is the utilization of deep learning techniques for solving partial differential equations(PDEs). In this work, we specifically focus on employing a three-layer tanh neural network within the framework of the deep Ritz method(DRM) to solve second-order elliptic equations with three different types of boundary conditions. We perform projected gradient descent(PDG) to train the three-layer network and we establish its global convergence. To the best of our knowledge, we are the first to provide a comprehensive error analysis of using overparameterized networks to solve PDE problems, as our analysis simultaneously includes estimates for approximation error, generalization error, and optimization error. We present error bound in terms of the sample size n and our work provides guidance on how to set the network depth, width, step size, and number of iterations for the projected gradient descent algorithm. Importantly, our assumptions in this work are classical and we do not require any additional assumptions on the solution of the equation. This ensures the broad applicability and generality of our results. Yuling Jiao, Yanming Lai, Yang Wang 0020 |
IEEE Trans. Inf. Theory | 2 |