Hude Liu

dblp:405/6660 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Learning theory · 67% Deep learning architectures and training · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
approximation theory
0.912025
Attention Mechanism, Max-Affine Partition, and Universal Approximation · NeurIPS 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Attention Mechanism, Max-Affine Partition, and Universal Approximation · NeurIPS 2025
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
0.912025
Attention Mechanism, Max-Affine Partition, and Universal Approximation · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

self-attention · 0.9cross-attention · 0.9
YearPublicationVenuePosition
2025 Attention Mechanism, Max-Affine Partition, and Universal Approximation
abstract
We establish the universal approximation capability of single-layer, single-head self- and cross-attention mechanisms with minimal attached structures. Our key insight is to interpret single-head attention as an input domain-partition mechanism that assigns distinct values to subregions. This allows us to engineer the attention weights such that this assignment imitates the target function. Building on this, we prove that a single self-attention layer, preceded by sum-of-linear transformations, is capable of approximating any continuous function on a compact domain under the $L_\infty$-norm. Furthermore, we extend this construction to approximate any Lebesgue integrable function under $L_p$-norm for $1\leq p <\infty$. Lastly, we also extend our techniques and show that, for the first time, single-head cross-attention achieves the same universal approximation guarantees.
Hude Liu, Jerry Yao-Chieh Hu, Zhao Song 0002, Han Liu 0001
NeurIPS1