Namjun Kim

dblp:302/6685 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 60% Deep learning architectures and training · 40%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
1.622025
Minimum Width for Universal Approximation using Squashable Activation Functions · ICML 2025
Minimum width for universal approximation using ReLU networks on compact domain · ICLR 2024
Machine learning › Deep learning architectures and training
activation function
0.912025
Minimum Width for Universal Approximation using Squashable Activation Functions · ICML 2025
Machine learning › Learning theory
neural network theory
0.912025
Minimum Width for Universal Approximation using Squashable Activation Functions · ICML 2025

Methods — techniques the papers use, named apart from their topics

composition · 0.9affine transformation · 0.9lp approximation · 0.8ReLU networks · 0.8
YearPublicationVenuePosition
2025 Minimum Width for Universal Approximation using Squashable Activation Functions
abstract
The exact minimum width that allows for universal approximation of unbounded-depth networks is known only for ReLU and its variants. In this work, we study the minimum width of networks using general activation functions. Specifically, we focus on squashable functions that can approximate the identity function and binary step function by alternatively composing with affine transformations. We show that for networks using a squashable activation function to universally approximate $L^p$ functions from $[0,1]^{d_x}$ to $\mathbb R^{d_y}$, the minimum width is $\max\\{d_x,d_y,2\\}$ unless $d_x=d_y=1$; the same bound holds for $d_x=d_y=1$ if the activation function is monotone. We then provide sufficient conditions for squashability and show that all non-affine analytic functions and a class of piecewise functions are squashable, i.e., our minimum width result holds for those general classes of activation functions.
Jonghyun Shin, Namjun Kim, Geonho Hwang
ICML2
2024 Minimum width for universal approximation using ReLU networks on compact domain
abstract
It has been shown that deep neural networks of a large enough width are universal approximators but they are not if the width is too small. There were several attempts to characterize the minimum width $w_{\min}$ enabling the universal approximation property; however, only a few of them found the exact values. In this work, we show that the minimum width for $L^p$ approximation of $L^p$ functions from $[0,1]^{d_x}$ to $\mathbb R^{d_y}$ is exactly $\max\\{d_x,d_y,2\\}$ if an activation function is ReLU-Like (e.g., ReLU, GELU, Softplus). Compared to the known result for ReLU networks, $w_{\min}=\max\\{d_x+1,d_y\\}$ when the domain is ${\mathbb R^{d_x}}$, our result first shows that approximation on a compact domain requires smaller width than on ${\mathbb R^{d_x}}$. We next prove a lower bound on $w_{\min}$ for uniform approximation using general activation functions including ReLU: $w_{\min}\ge d_y+1$ if $d_x<d_y\le2d_x$. Together with our first result, this shows a dichotomy between $L^p$ and uniform approximations for general activation functions and input/output dimensions.
Namjun Kim, Chanho Min
ICLR1