VLDB 2026 Research / reviewers in the wild / expert
Yongqiang Cai
dblp:228/6809
· DBLP profile ↗
6ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning theory · 39% Deep learning architectures and training · 33% Language models and text generation · 10% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 88% Automata and formal languages · 12% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
2.1 | 3 | 2024 | Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024 Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023 Achieve the Minimum Width of Neural Networks for Universal Approximation · ICLR 2023 |
Machine learning › Learning theory › neural network theory
neural network approximation theory |
1.3 | 2 | 2023 | Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023 Achieve the Minimum Width of Neural Networks for Universal Approximation · ICLR 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025 |
Mathematical optimization
approximation theory |
0.9 | 1 | 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025 |
Mathematical optimization › approximation theory
universal approximation |
0.9 | 1 | 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › model representation
compositional model |
0.8 | 1 | 2024 | Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024 |
Machine learning › Deep learning architectures and training
neural network expressivity |
0.7 | 1 | 2023 | Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.4 | 1 | 2019 | A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019 |
Machine learning › Optimization for machine learning
convergence analysis |
0.4 | 1 | 2019 | A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.4 | 1 | 2019 | A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019 |
Machine learning › Deep learning architectures and training
positional encoding |
0.3 | 1 | 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025 |
Automata and formal languages
regular languages |
0.2 | 1 | 2024 | Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
positional encoding · 1.7approximation theory · 1.7mapping composition · 1.5constructive approximation · 1.5topological theory · 0.7lift-flow-discretization · 0.7ordinary least squares analysis · 0.4condition number analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vocabulary In-Context Learning in Transformers: Benefits of Positional EncodingabstractNumerous studies have demonstrated that the Transformer architecture possesses the capability for in-context learning (ICL). In scenarios involving function approximation, context can serve as a control parameter for the model, endowing it with the universal approximation property (UAP). In practice, context is represented by tokens from a finite set, referred to as a vocabulary, which is the case considered in this paper, i.e., vocabulary in-context learning (VICL). We demonstrate that VICL in single-layer Transformers, without positional encoding, does not possess the UAP; however, it is possible to achieve the UAP when positional encoding is included. Several sufficient conditions for the positional encoding are provided. Our findings reveal the benefits of positional encoding from an approximation theory perspective in the context of in-context learning. Ruoxiang Xu, Yongqiang Cai |
NeurIPS | 3 |
| 2025 | Neural networks trained by weight permutation are universal approximatorsabstractThe universal approximation property is fundamental to the success of neural networks, and has traditionally been achieved by training networks without any constraints on their parameters. However, recent experimental research proposed a novel permutation-based training method, which exhibited a desired classification performance without modifying the exact weight values. In this paper, we provide a theoretical guarantee of this permutation training method by proving its ability to guide a ReLU network to approximate one-dimensional continuous functions. Our numerical results further validate this method’s efficiency in regression tasks with various initializations. The notable observations during weight permutation suggest that permutation training can provide an innovative tool for describing network learning behavior. • We prove the universal approximation property of permutation-trained networks to one-dimensional continuous functions. • The numerical experiments emphasize the crucial role played by the initializations in the permutation training scenario. • Observation of permutation patterns indicates that permutation training holds promise in describing intricate learning behaviors. Yongqiang Cai, Gaohang Chen, Zhonghua Qiao |
Neural Networks | 1 |
| 2024 | Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping CompositionsabstractIn recent years, deep learning-based sequence modelings, such as language models, have received much attention and success, which pushes researchers to explore the possibility of transforming non-sequential problems into a sequential form. Following this thought, deep neural networks can be represented as composite functions of a sequence of mappings, linear or nonlinear, where each composition can be viewed as a word. However, the weights of linear mappings are undetermined and hence require an infinite number of words. In this article, we investigate the finite case and constructively prove the existence of a finite vocabulary $V$=$\phi_i: \mathbb{R}^d \to \mathbb{R}^d | i=1,...,n$ with $n=O(d^2)$ for the universal approximation. That is, for any continuous mapping $f: \mathbb{R}^d \to \mathbb{R}^d$, compact domain $\Omega$ and $\varepsilon>0$, there is a sequence of mappings $\phi_{i_1}, ..., \phi_{i_m} \in V, m \in \mathbb{Z}^+$, such that the composition $\phi_{i_m} \circ ... \circ \phi_{i_1} $ approximates $f$ on $\Omega$ with an error less than $\varepsilon$. Our results demonstrate an unusual approximation power of mapping compositions and motivate a novel compositional model for regular languages. Yongqiang Cai |
ICML | 1 |
| 2023 | Achieve the Minimum Width of Neural Networks for Universal Approximation
Yongqiang Cai |
ICLR | 1 |
| 2023 | Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal ApproximationabstractThe study of universal approximation properties (UAP) for neural networks (NN) has a long history. When the network width is unlimited, only a single hidden layer is sufficient for UAP. In contrast, when the depth is unlimited, the width for UAP needs to be not less than the critical width $w^*_{\min}=\max(d_x,d_y)$, where $d_x$ and $d_y$ are the dimensions of the input and output, respectively. Recently, (Cai, 2022) shows that a leaky-ReLU NN with this critical width can achieve UAP for $L^p$ functions on a compact domain $\mathcal{K}$, i.e., the UAP for $L^p(\mathcal{K},\mathbb{R}^{d_y})$. This paper examines a uniform UAP for the function class $C(\mathcal{K},\mathbb{R}^{d_y})$ and gives the exact minimum width of the leaky-ReLU NN as $w_{\min}=\max(d_x+1,d_y)+1_{d_y=d_x+1}$, which involves the effects of the output dimensions. To obtain this result, we propose a novel lift-flow-discretization approach that shows that the uniform UAP has a deep connection with topological theory. Li'ang Li, Yifei Duan, Guanghua Ji, Yongqiang Cai |
ICML | 4 |
| 2019 | A Quantitative Analysis of the Effect of Batch Normalization on Gradient DescentabstractDespite its empirical success and recent theoretical progress, there generally lacks a quantitative analysis of the effect of batch normalization (BN) on the convergence and stability of gradient descent. In this paper, we provide such an analysis on the simple problem of ordinary least squares (OLS), where the precise dynamical properties of gradient descent (GD) is completely known, thus allowing us to isolate and compare the additional effects of BN. More precisely, we show that unlike GD, gradient descent with BN (BNGD) converges for arbitrary learning rates for the weights, and the convergence remains linear under mild conditions. Moreover, we quantify two different sources of acceleration of BNGD over GD – one due to over-parameterization which improves the effective condition number and another due having a large range of learning rates giving rise to fast descent. These phenomena set BNGD apart from GD and could account for much of its robustness properties. These findings are confirmed quantitatively by numerical experiments, which further show that many of the uncovered properties of BNGD in OLS are also observed qualitatively in more complex supervised learning problems. Yongqiang Cai, Qianxiao Li, Zuowei Shen |
ICML | 1 |