Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shunta Akiyama

dblp:280/3821 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Learning theory · 31% Optimization for machine learning · 25% Deep learning architectures and training · 21%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
1.522025
Survival Analysis via Density Estimation · ICML 2025
Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023
Machine learning › Learning theory
generalization bounds
1.422025
Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021
Machine learning › Optimization for machine learning
non-convex optimization
1.422025
Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021
Machine learning › Deep learning architectures and training
ReLU networks
1.222023
Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021
Machine learning › Deep learning architectures and training
teacher-student framework
1.222023
Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021
Machine learning › Optimization for machine learning › convergence analysis
global convergence guarantees
0.912025
Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025
Machine learning › Learning theory › generalization bounds
rademacher complexity
0.912025
Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025
Bioinformatics and computational biology
survival analysis
0.912025
Survival Analysis via Density Estimation · ICML 2025
Machine learning › Efficient and distributed learning › federated learning
communication-efficient federated learning
0.812024
SILVER: Single-loop variance reduction and application to federated learning · ICML 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
SILVER: Single-loop variance reduction and application to federated learning · ICML 2024
Machine learning › Optimization for machine learning
variance reduction
0.812024
SILVER: Single-loop variance reduction and application to federated learning · ICML 2024
Machine learning › Generative modeling
diffusion model
0.712023
Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023
Machine learning › Learning theory › statistical learning theory
excess risk
0.712023
Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023
Machine learning › Learning theory › statistical estimation
minimax estimation
0.712023
Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023
Machine learning › Learning theory › approximation theory
neural network approximation
0.712023
Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023
Machine learning › Learning theory
excess risk bounds
0.512021
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence
0.512021
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021
Machine learning › Optimization for machine learning › stochastic optimization
noisy gradient descent
0.512021
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021
Machine learning › Deep learning architectures and training
training dynamics
0.512021
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021
Machine learning › Deep learning architectures and training › feedforward neural network
two-layer neural network
0.512021
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021
Mathematical optimization
nonconvex optimization
0.212024
SILVER: Single-loop variance reduction and application to federated learning · ICML 2024
Algorithms and data structures
kernel methods
0.212023
Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023
Machine learning › Learning theory
over-parameterization
0.112021
On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021

Methods — techniques the papers use, named apart from their topics

kernel methods · 1.8density estimation · 1.7copula · 1.7variance-reduced gradient estimator · 1.5local updates · 1.5teacher-student · 1.3rademacher complexity · 0.9block coordinate descent · 0.9total variation bound · 0.7score matching · 0.7
YearPublicationVenuePosition
2025 Survival Analysis via Density Estimation
abstract
This paper introduces a novel framework for survival analysis by reinterpreting it as a form of density estimation. Our algorithm post-processes density estimation outputs to derive survival functions, enabling the application of any density estimation model to effectively estimate survival functions. This approach broadens the toolkit for survival analysis and enhances the flexibility and applicability of existing techniques. Our framework is versatile enough to handle various survival analysis scenarios, including competing risk models for multiple event types. It can also address dependent censoring when prior knowledge of the dependency between event time and censoring time is available in the form of a copula. In the absence of such information, our framework can estimate the upper and lower bounds of survival functions, accounting for the associated uncertainty.
Hiroki Yanagisawa, Shunta Akiyama
ICML2
2025 Block Coordinate Descent for Neural Networks Provably Finds Global Minima
abstract
In this paper, we consider a block coordinate descent (BCD) algorithm for training deep neural networks and provide a new global convergence guarantee under strictly monotonically increasing activation functions. While existing works demonstrate convergence to stationary points for BCD in neural networks, our contribution is the first to prove convergence to global minima, ensuring arbitrarily small loss. We show that the loss with respect to the output layer decreases exponentially while the loss with respect to the hidden layers remains well-controlled. Additionally, we derive generalization bounds using the Rademacher complexity framework, demonstrating that BCD not only achieves strong optimization guarantees but also provides favorable generalization performance. Moreover, we propose a modified BCD algorithm with skip connections and non-negative projection, extending our convergence guarantees to ReLU activation, which are not strictly monotonic. Empirical experiments confirm our theoretical findings, showing that the BCD algorithm achieves a small loss for strictly monotonic and ReLU activations.
Shunta Akiyama
NeurIPS1
2024 SILVER: Single-loop variance reduction and application to federated learning
abstract
Most variance reduction methods require multiple times of full gradient computation, which is time-consuming and hence a bottleneck in application to distributed optimization. We present a single-loop variance-reduced gradient estimator named SILVER (SIngle-Loop VariancE-Reduction) for the finite-sum non-convex optimization, which does not require multiple full gradients but nevertheless achieves the optimal gradient complexity. Notably, unlike existing methods, SILVER provably reaches second-order optimality, with exponential convergence in the Polyak-Łojasiewicz (PL) region, and achieves further speedup depending on the data heterogeneity. Owing to these advantages, SILVER serves as a new base method to design communication-efficient federated learning algorithms: we combine SILVER with local updates which gives the best communication rounds and number of communicated gradients across all range of Hessian heterogeneity, and, at the same time, guarantees second-order optimality and exponential convergence in the PL region.
Kazusato Oko, Shunta Akiyama, Denny Wu, Tomoya Murata, Taiji Suzuki
ICML2
2023 Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods
Shunta Akiyama, Taiji Suzuki
ICLR1
2023 Diffusion Models are Minimax Optimal Distribution Estimators
abstract
While efficient distribution learning is no doubt behind the groundbreaking success of diffusion modeling, its theoretical guarantees are quite limited. In this paper, we provide the first rigorous analysis on approximation and generalization abilities of diffusion modeling for well-known function spaces. The highlight of this paper is that when the true density function belongs to the Besov space and the empirical score matching loss is properly minimized, the generated data distribution achieves the nearly minimax optimal estimation rates in the total variation distance and in the Wasserstein distance of order one. Furthermore, we extend our theory to demonstrate how diffusion models adapt to low-dimensional data distributions. We expect these results advance theoretical understandings of diffusion modeling and its ability to generate verisimilar outputs.
Kazusato Oko, Shunta Akiyama, Taiji Suzuki
ICML2
2021 Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Taiji Suzuki, Shunta Akiyama
ICLR2
2021 On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting
abstract
Deep learning empirically achieves high performance in many applications, but its training dynamics has not been fully understood theoretically. In this paper, we explore theoretical analysis on training two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown teacher network through its outputs. We show that with a specific regularization and sufficient over-parameterization, the student network can identify the parameters of the teacher network with high probability via gradient descent with a norm dependent stepsize even though the objective function is highly non-convex. The key theoretical tool is the measure representation of the neural networks and a novel application of a dual certificate argument for sparse estimation on a measure space. We analyze the global minima and global convergence property in the measure space.
Shunta Akiyama, Taiji Suzuki
ICML1