VLDB 2026 Research / reviewers in the wild / expert
Shunta Akiyama
dblp:280/3821
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Learning theory · 31% Optimization for machine learning · 25% Deep learning architectures and training · 21% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
1.5 | 2 | 2025 | Survival Analysis via Density Estimation · ICML 2025 Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023 |
Machine learning › Learning theory
generalization bounds |
1.4 | 2 | 2025 | Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025 Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021 |
Machine learning › Optimization for machine learning
non-convex optimization |
1.4 | 2 | 2025 | Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025 Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021 |
Machine learning › Deep learning architectures and training
ReLU networks |
1.2 | 2 | 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023 On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Machine learning › Deep learning architectures and training
teacher-student framework |
1.2 | 2 | 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023 On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Machine learning › Optimization for machine learning › convergence analysis
global convergence guarantees |
0.9 | 1 | 2025 | Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025 |
Machine learning › Learning theory › generalization bounds
rademacher complexity |
0.9 | 1 | 2025 | Block Coordinate Descent for Neural Networks Provably Finds Global Minima · NeurIPS 2025 |
Bioinformatics and computational biology
survival analysis |
0.9 | 1 | 2025 | Survival Analysis via Density Estimation · ICML 2025 |
Machine learning › Efficient and distributed learning › federated learning
communication-efficient federated learning |
0.8 | 1 | 2024 | SILVER: Single-loop variance reduction and application to federated learning · ICML 2024 |
Machine learning › Efficient and distributed learning
federated learning |
0.8 | 1 | 2024 | SILVER: Single-loop variance reduction and application to federated learning · ICML 2024 |
Machine learning › Optimization for machine learning
variance reduction |
0.8 | 1 | 2024 | SILVER: Single-loop variance reduction and application to federated learning · ICML 2024 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023 |
Machine learning › Learning theory › statistical learning theory
excess risk |
0.7 | 1 | 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023 |
Machine learning › Learning theory › statistical estimation
minimax estimation |
0.7 | 1 | 2023 | Diffusion Models are Minimax Optimal Distribution Estimators · ICML 2023 |
Machine learning › Learning theory › approximation theory
neural network approximation |
0.7 | 1 | 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023 |
Machine learning › Learning theory
excess risk bounds |
0.5 | 1 | 2021 | Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021 |
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence |
0.5 | 1 | 2021 | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Machine learning › Optimization for machine learning › stochastic optimization
noisy gradient descent |
0.5 | 1 | 2021 | Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods · ICLR 2021 |
Machine learning › Deep learning architectures and training
training dynamics |
0.5 | 1 | 2021 | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Machine learning › Deep learning architectures and training › feedforward neural network
two-layer neural network |
0.5 | 1 | 2021 | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Mathematical optimization
nonconvex optimization |
0.2 | 1 | 2024 | SILVER: Single-loop variance reduction and application to federated learning · ICML 2024 |
Algorithms and data structures
kernel methods |
0.2 | 1 | 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods · ICLR 2023 |
Machine learning › Learning theory
over-parameterization |
0.1 | 1 | 2021 | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student Setting · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
kernel methods · 1.8density estimation · 1.7copula · 1.7variance-reduced gradient estimator · 1.5local updates · 1.5teacher-student · 1.3rademacher complexity · 0.9block coordinate descent · 0.9total variation bound · 0.7score matching · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Survival Analysis via Density EstimationabstractThis paper introduces a novel framework for survival analysis by reinterpreting it as a form of density estimation. Our algorithm post-processes density estimation outputs to derive survival functions, enabling the application of any density estimation model to effectively estimate survival functions. This approach broadens the toolkit for survival analysis and enhances the flexibility and applicability of existing techniques. Our framework is versatile enough to handle various survival analysis scenarios, including competing risk models for multiple event types. It can also address dependent censoring when prior knowledge of the dependency between event time and censoring time is available in the form of a copula. In the absence of such information, our framework can estimate the upper and lower bounds of survival functions, accounting for the associated uncertainty. Hiroki Yanagisawa, Shunta Akiyama |
ICML | 2 |
| 2025 | Block Coordinate Descent for Neural Networks Provably Finds Global MinimaabstractIn this paper, we consider a block coordinate descent (BCD) algorithm for training deep neural networks and provide a new global convergence guarantee under strictly monotonically increasing activation functions. While existing works demonstrate convergence to stationary points for BCD in neural networks, our contribution is the first to prove convergence to global minima, ensuring arbitrarily small loss. We show that the loss with respect to the output layer decreases exponentially while the loss with respect to the hidden layers remains well-controlled. Additionally, we derive generalization bounds using the Rademacher complexity framework, demonstrating that BCD not only achieves strong optimization guarantees but also provides favorable generalization performance. Moreover, we propose a modified BCD algorithm with skip connections and non-negative projection, extending our convergence guarantees to ReLU activation, which are not strictly monotonic. Empirical experiments confirm our theoretical findings, showing that the BCD algorithm achieves a small loss for strictly monotonic and ReLU activations. Shunta Akiyama |
NeurIPS | 1 |
| 2024 | SILVER: Single-loop variance reduction and application to federated learningabstractMost variance reduction methods require multiple times of full gradient computation, which is time-consuming and hence a bottleneck in application to distributed optimization. We present a single-loop variance-reduced gradient estimator named SILVER (SIngle-Loop VariancE-Reduction) for the finite-sum non-convex optimization, which does not require multiple full gradients but nevertheless achieves the optimal gradient complexity. Notably, unlike existing methods, SILVER provably reaches second-order optimality, with exponential convergence in the Polyak-Łojasiewicz (PL) region, and achieves further speedup depending on the data heterogeneity. Owing to these advantages, SILVER serves as a new base method to design communication-efficient federated learning algorithms: we combine SILVER with local updates which gives the best communication rounds and number of communicated gradients across all range of Hessian heterogeneity, and, at the same time, guarantees second-order optimality and exponential convergence in the PL region. Kazusato Oko, Shunta Akiyama, Denny Wu, Tomoya Murata, Taiji Suzuki |
ICML | 2 |
| 2023 | Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods
Shunta Akiyama, Taiji Suzuki |
ICLR | 1 |
| 2023 | Diffusion Models are Minimax Optimal Distribution EstimatorsabstractWhile efficient distribution learning is no doubt behind the groundbreaking success of diffusion modeling, its theoretical guarantees are quite limited. In this paper, we provide the first rigorous analysis on approximation and generalization abilities of diffusion modeling for well-known function spaces. The highlight of this paper is that when the true density function belongs to the Besov space and the empirical score matching loss is properly minimized, the generated data distribution achieves the nearly minimax optimal estimation rates in the total variation distance and in the Wasserstein distance of order one. Furthermore, we extend our theory to demonstrate how diffusion models adapt to low-dimensional data distributions. We expect these results advance theoretical understandings of diffusion modeling and its ability to generate verisimilar outputs. Kazusato Oko, Shunta Akiyama, Taiji Suzuki |
ICML | 2 |
| 2021 | Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Taiji Suzuki, Shunta Akiyama |
ICLR | 2 |
| 2021 | On Learnability via Gradient Method for Two-Layer ReLU Neural Networks in Teacher-Student SettingabstractDeep learning empirically achieves high performance in many applications, but its training dynamics has not been fully understood theoretically. In this paper, we explore theoretical analysis on training two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown teacher network through its outputs. We show that with a specific regularization and sufficient over-parameterization, the student network can identify the parameters of the teacher network with high probability via gradient descent with a norm dependent stepsize even though the objective function is highly non-convex. The key theoretical tool is the measure representation of the neural networks and a novel application of a dual certificate argument for sparse estimation on a measure space. We analyze the global minima and global convergence property in the measure space. Shunta Akiyama, Taiji Suzuki |
ICML | 1 |