Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yongqiang Cai

dblp:228/6809 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Learning theory · 39% Deep learning architectures and training · 33% Language models and text generation · 10%
Theoretical computer science
2 papers
Mathematical optimization · 88% Automata and formal languages · 12%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
2.132024
Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024
Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023
Achieve the Minimum Width of Neural Networks for Universal Approximation · ICLR 2023
Machine learning › Learning theory › neural network theory
neural network approximation theory
1.322023
Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023
Achieve the Minimum Width of Neural Networks for Universal Approximation · ICLR 2023
Natural language and speech › Language models and text generation
in-context learning
0.912025
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025
Mathematical optimization
approximation theory
0.912025
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025
Mathematical optimization › approximation theory
universal approximation
0.912025
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › model representation
compositional model
0.812024
Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024
Machine learning › Deep learning architectures and training
neural network expressivity
0.712023
Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation · ICML 2023
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.412019
A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019
Machine learning › Optimization for machine learning
convergence analysis
0.412019
A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.412019
A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent · ICML 2019
Machine learning › Deep learning architectures and training
positional encoding
0.312025
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding · NeurIPS 2025
Automata and formal languages
regular languages
0.212024
Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions · ICML 2024

Methods — techniques the papers use, named apart from their topics

positional encoding · 1.7approximation theory · 1.7mapping composition · 1.5constructive approximation · 1.5topological theory · 0.7lift-flow-discretization · 0.7ordinary least squares analysis · 0.4condition number analysis · 0.4
YearPublicationVenuePosition
2025 Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
abstract
Numerous studies have demonstrated that the Transformer architecture possesses the capability for in-context learning (ICL). In scenarios involving function approximation, context can serve as a control parameter for the model, endowing it with the universal approximation property (UAP). In practice, context is represented by tokens from a finite set, referred to as a vocabulary, which is the case considered in this paper, i.e., vocabulary in-context learning (VICL). We demonstrate that VICL in single-layer Transformers, without positional encoding, does not possess the UAP; however, it is possible to achieve the UAP when positional encoding is included. Several sufficient conditions for the positional encoding are provided. Our findings reveal the benefits of positional encoding from an approximation theory perspective in the context of in-context learning.
Ruoxiang Xu, Yongqiang Cai
NeurIPS3
2025 Neural networks trained by weight permutation are universal approximators
abstract
The universal approximation property is fundamental to the success of neural networks, and has traditionally been achieved by training networks without any constraints on their parameters. However, recent experimental research proposed a novel permutation-based training method, which exhibited a desired classification performance without modifying the exact weight values. In this paper, we provide a theoretical guarantee of this permutation training method by proving its ability to guide a ReLU network to approximate one-dimensional continuous functions. Our numerical results further validate this method’s efficiency in regression tasks with various initializations. The notable observations during weight permutation suggest that permutation training can provide an innovative tool for describing network learning behavior. • We prove the universal approximation property of permutation-trained networks to one-dimensional continuous functions. • The numerical experiments emphasize the crucial role played by the initializations in the permutation training scenario. • Observation of permutation patterns indicates that permutation training holds promise in describing intricate learning behaviors.
Yongqiang Cai, Gaohang Chen, Zhonghua Qiao
Neural Networks1
2024 Vocabulary for Universal Approximation: A Linguistic Perspective of Mapping Compositions
abstract
In recent years, deep learning-based sequence modelings, such as language models, have received much attention and success, which pushes researchers to explore the possibility of transforming non-sequential problems into a sequential form. Following this thought, deep neural networks can be represented as composite functions of a sequence of mappings, linear or nonlinear, where each composition can be viewed as a word. However, the weights of linear mappings are undetermined and hence require an infinite number of words. In this article, we investigate the finite case and constructively prove the existence of a finite vocabulary $V$=$\phi_i: \mathbb{R}^d \to \mathbb{R}^d | i=1,...,n$ with $n=O(d^2)$ for the universal approximation. That is, for any continuous mapping $f: \mathbb{R}^d \to \mathbb{R}^d$, compact domain $\Omega$ and $\varepsilon>0$, there is a sequence of mappings $\phi_{i_1}, ..., \phi_{i_m} \in V, m \in \mathbb{Z}^+$, such that the composition $\phi_{i_m} \circ ... \circ \phi_{i_1} $ approximates $f$ on $\Omega$ with an error less than $\varepsilon$. Our results demonstrate an unusual approximation power of mapping compositions and motivate a novel compositional model for regular languages.
Yongqiang Cai
ICML1
2023 Achieve the Minimum Width of Neural Networks for Universal Approximation
Yongqiang Cai
ICLR1
2023 Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation
abstract
The study of universal approximation properties (UAP) for neural networks (NN) has a long history. When the network width is unlimited, only a single hidden layer is sufficient for UAP. In contrast, when the depth is unlimited, the width for UAP needs to be not less than the critical width $w^*_{\min}=\max(d_x,d_y)$, where $d_x$ and $d_y$ are the dimensions of the input and output, respectively. Recently, (Cai, 2022) shows that a leaky-ReLU NN with this critical width can achieve UAP for $L^p$ functions on a compact domain $\mathcal{K}$, i.e., the UAP for $L^p(\mathcal{K},\mathbb{R}^{d_y})$. This paper examines a uniform UAP for the function class $C(\mathcal{K},\mathbb{R}^{d_y})$ and gives the exact minimum width of the leaky-ReLU NN as $w_{\min}=\max(d_x+1,d_y)+1_{d_y=d_x+1}$, which involves the effects of the output dimensions. To obtain this result, we propose a novel lift-flow-discretization approach that shows that the uniform UAP has a deep connection with topological theory.
Li'ang Li, Yifei Duan, Guanghua Ji, Yongqiang Cai
ICML4
2019 A Quantitative Analysis of the Effect of Batch Normalization on Gradient Descent
abstract
Despite its empirical success and recent theoretical progress, there generally lacks a quantitative analysis of the effect of batch normalization (BN) on the convergence and stability of gradient descent. In this paper, we provide such an analysis on the simple problem of ordinary least squares (OLS), where the precise dynamical properties of gradient descent (GD) is completely known, thus allowing us to isolate and compare the additional effects of BN. More precisely, we show that unlike GD, gradient descent with BN (BNGD) converges for arbitrary learning rates for the weights, and the convergence remains linear under mild conditions. Moreover, we quantify two different sources of acceleration of BNGD over GD – one due to over-parameterization which improves the effective condition number and another due having a large range of learning rates giving rise to fast descent. These phenomena set BNGD apart from GD and could account for much of its robustness properties. These findings are confirmed quantitatively by numerical experiments, which further show that many of the uncovered properties of BNGD in OLS are also observed qualitatively in more complex supervised learning problems.
Yongqiang Cai, Qianxiao Li, Zuowei Shen
ICML1