Yuri Kinoshita

dblp:317/4944 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 29% Learning theory · 20% Optimization for machine learning · 20%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › neural network theory › feature learning theory
information exponent
0.912025
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning · ICML 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning · ICML 2025
Machine learning › Learning theory › statistical estimation › semiparametric inference
single-index model
0.912025
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning · ICML 2025
Machine learning › Deep learning architectures and training › training optimization
stochastic gradient descent dynamics
0.912025
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning · ICML 2025
Machine learning › Deep learning architectures and training › feedforward neural network
convex neural network
0.812024
A provable control of sensitivity of neural networks through a direct parameterization of the overall bi-Lipschitzness · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
A provable control of sensitivity of neural networks through a direct parameterization of the overall bi-Lipschitzness · NeurIPS 2024
Machine learning › Generative modeling › variational autoencoder
posterior collapse
0.712023
Controlling Posterior Collapse by an Inverse Lipschitz Constraint on the Decoder Network · ICML 2023
Machine learning › Generative modeling
variational autoencoder
0.712023
Controlling Posterior Collapse by an Inverse Lipschitz Constraint on the Decoder Network · ICML 2023
Machine learning › Optimization for machine learning
convergence analysis
0.612022
Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization · NeurIPS 2022
Machine learning › Optimization for machine learning
non-convex optimization
0.612022
Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo › langevin dynamics
stochastic gradient langevin dynamics
0.612022
Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization · NeurIPS 2022
Machine learning › Optimization for machine learning
variance reduction
0.612022
Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

stochastic gradient descent · 0.9mixture of experts · 0.9legendre-fenchel duality · 0.8inverse lipschitz neural network · 0.7log-sobolev inequality · 0.6KL divergence · 0.6
YearPublicationVenuePosition
2025 Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
abstract
Mixture of Experts (MoE), an ensemble of specialized models equipped with a router that dynamically distributes each input to appropriate experts, has achieved successful results in the field of machine learning. However, theoretical understanding of this architecture is falling behind due to its inherent complexity. In this paper, we theoretically study the sample and runtime complexity of MoE following the stochastic gradient descent when learning a regression task with an underlying cluster structure of single index models. On the one hand, we show that a vanilla neural network fails in detecting such a latent organization as it can only process the problem as a whole. This is intrinsically related to the concept of *information exponent* which is low for each cluster, but increases when we consider the entire task. On the other hand, with a MoE, we show that it succeeds in dividing the problem into easier subproblems by leveraging the ability of each expert to weakly recover the simpler function corresponding to an individual cluster. To the best of our knowledge, this work is among the first to explore the benefits of the MoE framework by examining its SGD dynamics in the context of nonlinear regression.
Ryotaro Kawata, Kohsei Matsutani, Yuri Kinoshita, Naoki Nishikawa, Taiji Suzuki
ICML3
2024 A provable control of sensitivity of neural networks through a direct parameterization of the overall bi-Lipschitzness
abstract
While neural networks can enjoy an outstanding flexibility and exhibit unprecedented performance, the mechanism behind their behavior is still not well-understood. To tackle this fundamental challenge, researchers have tried to restrict and manipulate some of their properties in order to gain new insights and better control on them. Especially, throughout the past few years, the concept of *bi-Lipschitzness* has been proved as a beneficial inductive bias in many areas. However, due to its complexity, the design and control of bi-Lipschitz architectures are falling behind, and a model that is precisely designed for bi-Lipschitzness realizing a direct and simple control of the constants along with solid theoretical analysis is lacking. In this work, we investigate and propose a novel framework for bi-Lipschitzness that can achieve such a clear and tight control based on convex neural networks and the Legendre-Fenchel duality. Its desirable properties are illustrated with concrete experiments to illustrate its broad range of applications.
Yuri Kinoshita, Taro Toyoizumi
NeurIPS1
2023 Controlling Posterior Collapse by an Inverse Lipschitz Constraint on the Decoder Network
abstract
Variational autoencoders (VAEs) are one of the deep generative models that have experienced enormous success over the past decades. However, in practice, they suffer from a problem called posterior collapse, which occurs when the posterior distribution coincides, or collapses, with the prior taking no information from the latent structure of the input data into consideration. In this work, we introduce an inverse Lipschitz neural network into the decoder and, based on this architecture, provide a new method that can control in a simple and clear manner the degree of posterior collapse for a wide range of VAE models equipped with a concrete theoretical guarantee. We also illustrate the effectiveness of our method through several numerical experiments.
Yuri Kinoshita, Kenta Oono, Kenji Fukumizu, Yuichi Yoshida, Shin-ichi Maeda
ICML1
2022 Improved Convergence Rate of Stochastic Gradient Langevin Dynamics with Variance Reduction and its Application to Optimization
abstract
The stochastic gradient Langevin Dynamics is one of the most fundamental algorithms to solve sampling problems and non-convex optimization appearing in several machine learning applications. Especially, its variance reduced versions have nowadays gained particular attention. In this paper, we study two variants of this kind, namely, the Stochastic Variance Reduced Gradient Langevin Dynamics and the Stochastic Recursive Gradient Langevin Dynamics. We prove their convergence to the objective distribution in terms of KL-divergence under the sole assumptions of smoothness and Log-Sobolev inequality which are weaker conditions than those used in prior works for these algorithms. With the batch size and the inner loop length set to $\sqrt{n}$, the gradient complexity to achieve an $\epsilon$-precision is $\tilde{O}((n+dn^{1/2}\epsilon^{-1})\gamma^2 L^2\alpha^{-2})$, which is an improvement from any previous analyses. We also show some essential applications of our result to non-convex optimization.
Yuri Kinoshita, Taiji Suzuki
NeurIPS1