EDBT 2026 Demo / reviewers in the wild / expert
Ryo Karakida
dblp:183/0467
· DBLP profile ↗
24ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0003-2294-8946ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 9 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target PropagationabstractLocal learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters due to the locality, making it challenging to identify desirable settings where the algorithm progresses in a stable manner.
To provide theoretical and quantitative insights, we introduce maximal update parameterization ($\mu$P) in the infinite-width limit for two representative designs of local targets: predictive coding (PC) and target propagation (TP). We verify that $\mu$P enables hyperparameter transfer across models of different widths.
Furthermore, our analysis reveals unique and intriguing properties of $\mu$P that are not present in conventional BP. By analyzing deep linear networks, we find that PC's gradients interpolate between first-order and Gauss-Newton-like gradients, depending on the parameterization.
We demonstrate that, in specific standard settings, PC in the infinite-width limit behaves more similarly to the first-order gradient.
For TP, even with the standard scaling of the last layer differing from classical $\mu$P, its local loss optimization favors the feature learning regime over the kernel regime. Satoki Ishikawa, Rio Yokota, Ryo Karakida |
ICLR | 3 |
| 2025 | Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor ProgramsabstractIn modern theoretical analyses of neural networks, the infinite-width limit is often invoked to justify Gaussian approximations of neuron preactivations (e.g., via neural network Gaussian processes or Tensor Programs). However, these Gaussian-based asymptotic theories have so far been unable to capture the behavior of attention layers, except under special regimes such as infinitely many heads or tailored scaling schemes. In this paper, leveraging the Tensor Programs framework, we rigorously identify the infinite-width limit distribution of variables within a single attention layer under realistic architectural dimensionality and standard $1/\sqrt{n}$-scaling with $n$ dimensionality. We derive the exact form of this limit law without resorting to infinite-head approximations or tailored scalings, demonstrating that it departs fundamentally from Gaussianity. This limiting distribution exhibits non-Gaussianity from a hierarchical structure, being Gaussian conditional on the random similarity scores. Numerical experiments validate our theoretical predictions, confirming the effectiveness of our theory at finite width and accurate description of finite-head attentions. Beyond characterizing a standalone attention layer, our findings lay the groundwork for developing a unified theory of deep Transformer architectures in the infinite-width regime. Mana Sakai, Ryo Karakida, Masaaki Imaizumi |
NeurIPS | 2 |
| 2025 | Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from JacobiansabstractThe theoretical understanding of self-attention (SA) has been steadily progressing. A prominent line of work studies a class of SA layers that admit an energy function decreased by state updates. While it provides valuable insights into inherent biases in signal propagation, it often relies on idealized assumptions or additional constraints not necessarily present in standard SA. Thus, to broaden our understanding, this work aims to relax these energy constraints and provide an energy-agnostic characterization of inference dynamics by dynamical systems analysis. In more detail, we first consider relaxing the symmetry and single-head constraints traditionally required in energy-based formulations. Next, we show that analyzing the Jacobian matrix of the state is highly valuable when investigating more general SA architectures without necessarily admitting an energy function. It reveals that the normalization layer plays an essential role in suppressing the Lipschitzness of SA and the Jacobian's complex eigenvalues, which correspond to the oscillatory components of the dynamics. In addition, the Lyapunov exponents computed from the Jacobians demonstrate that the normalized dynamics lie close to a critical state, and this criticality serves as a strong indicator of high inference performance.
Furthermore, the Jacobian perspective also enables us to develop regularization methods for training and a pseudo-energy for monitoring inference dynamics. Akiyoshi Tomihari, Ryo Karakida |
NeurIPS | 2 |
| 2025 | Continual learning in inertial measurement unit-based human activity recognition with user-centric class-incremental learning scenarioabstractThe development of neural networks and wearable sensing technologies has increased the focus on modules for human activity recognition (HAR) using inertial measurement units (IMUs). Updating a pre-trained network has the potential to enhance user-centric interfaces. However, updating the network parameters with new class data can lead to catastrophic forgetting. Continual learning (CL) methods have been proposed to address this issue. The amount of new class data from the user (target subject) is considerably small compared to the amount of existing class data from the pre-measured subjects that are used for pre-training the network. The efficiency of CL methods in IMU-based HAR with user-centric class-incremental learning (Class-IL) scenarios is unclear. Thus, we compared 12 CL methods employing the regularization , architectural, and replay approaches. The evaluation was performed using five well-known IMU-based HAR datasets. Among the three approaches, the replay-based methods effectively prevented forgetting in IMU-based HAR, even with a small amount of data, by providing several samples per class. Moreover, in some cases, hybrid regularization and replay methods performed better than replay-based methods. The findings of this study highlight the challenging nature of suppressing forgetting in the Class-IL scenario, particularly when incrementally incorporating a limited amount of new class data from a target subject. Our future work will focus on the development of a hybrid method for IMU-based HAR with online user-centric Class-IL scenarios. Suguru Kanoga, Yuya Tsukiji, Ryo Karakida |
Expert Syst. Appl. | 3 |
| 2025 | Optimal layer selection for latent data augmentation
Tomoumi Takase, Ryo Karakida |
Neural Networks | 2 |
| 2024 | On the Parameterization of Second-Order Optimization Effective towards the Infinite WidthabstractSecond-order optimization has been developed to accelerate the training of deep neural networks and it is being applied to increasingly larger-scale models. In this study, towards training on further larger scales, we identify a specific parameterization for second-order optimization that promotes feature learning in a stable manner even if the network width increases significantly. Inspired by a maximal update parametrization, we consider a one-step update of the gradient and reveal the appropriate scales of hyperparameters including random initialization, learning rates, and damping terms. Our approach covers two major second-order optimization algorithms, K-FAC and Shampoo, and we demonstrate that our parametrization achieves higher generalization performance in feature learning.
In particular, it enables us to transfer the hyperparameters across models with different widths. Satoki Ishikawa, Ryo Karakida |
ICLR | 2 |
| 2024 | Self-attention Networks Localize When QK-eigenspectrum ConcentratesabstractThe self-attention mechanism prevails in modern machine learning. It has an interesting functionality of adaptively selecting tokens from an input sequence by modulating the degree of attention localization, which many researchers speculate is the basis of the powerful model performance but complicates the underlying mechanism of the learning dynamics. In recent years, mainly two arguments have connected attention localization to the model performances. One is the rank collapse, where the embedded tokens by a self-attention block become very similar across different tokens, leading to a less expressive network. The other is the entropy collapse, where the attention probability approaches non-uniform and entails low entropy, making the learning dynamics more likely to be trapped in plateaus. These two failure modes may apparently contradict each other because the rank and entropy collapses are relevant to uniform and non-uniform attention, respectively. To this end, we characterize the notion of attention localization by the eigenspectrum of query-key parameter matrices and reveal that a small eigenspectrum variance leads attention to be localized. Interestingly, the small eigenspectrum variance prevents both rank and entropy collapse, leading to better model expressivity and trainability. Han Bao 0002, Ryuichiro Hataya, Ryo Karakida |
ICML | 3 |
| 2024 | Understanding MLP-Mixer as a wide and sparse MLPabstractMulti-layer perceptron (MLP) is a fundamental component of deep learning, and recent MLP-based architectures, especially the MLP-Mixer, have achieved significant empirical success. Nevertheless, our understanding of why and how the MLP-Mixer outperforms conventional MLPs remains largely unexplored. In this work, we reveal that sparseness is a key mechanism underlying the MLP-Mixers. First, the Mixers have an effective expression as a wider MLP with Kronecker-product weights, clarifying that the Mixers efficiently embody several sparseness properties explored in deep learning. In the case of linear layers, the effective expression elucidates an implicit sparse regularization caused by the model architecture and a hidden relation to Monarch matrices, which is also known as another form of sparse parameterization. Next, for general cases, we empirically demonstrate quantitative similarities between the Mixer and the unstructured sparse-weight MLPs. Following a guiding principle proposed by Golubeva, Neyshabur and Gur-Ari (2021), which fixes the number of connections and increases the width and sparsity, the Mixers can demonstrate improved performance. Tomohiro Hayase, Ryo Karakida |
ICML | 2 |
| 2023 | Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit BiasabstractGradient regularization (GR) is a method that penalizes the gradient norm of the training loss during training. While some studies have reported that GR can improve generalization performance, little attention has been paid to it from the algorithmic perspective, that is, the algorithms of GR that efficiently improve the performance. In this study, we first reveal that a specific finite-difference computation, composed of both gradient ascent and descent steps, reduces the computational cost of GR. Next, we show that the finite-difference computation also works better in the sense of generalization performance. We theoretically analyze a solvable model, a diagonal linear network, and clarify that GR has a desirable implicit bias to so-called rich regime and finite-difference computation strengthens this bias. Furthermore, finite-difference GR is closely related to some other algorithms based on iterative ascent and descent steps for exploring flat minima. In particular, we reveal that the flooding method can perform finite-difference GR in an implicit way. Thus, this work broadens our understanding of GR for both practice and theory. Ryo Karakida, Tomoumi Takase, Tomohiro Hayase, Kazuki Osawa |
ICML | 1 |
| 2023 | Attention in a Family of Boltzmann Machines Emerging From Modern Hopfield NetworksabstractHopfield networks and Boltzmann machines (BMs) are fundamental energy-based neural network models. Recent studies on modern Hopfield networks have broadened the class of energy functions and led to a unified perspective on general Hopfield networks, including an attention module. In this letter, we consider the BM counterparts of modern Hopfield networks using the associated energy functions and study their salient properties from a trainability perspective. In particular, the energy function corresponding to the attention module naturally introduces a novel BM, which we refer to as the attentional BM (AttnBM). We verify that AttnBM has a tractable likelihood function and gradient for certain special cases and is easy to train. Moreover, we reveal the hidden connections between AttnBM and some single-layer models, namely the gaussian-Bernoulli restricted BM and the denoising autoencoder with softmax units coming from denoising score matching. We also investigate BMs introduced by other energy functions and show that the energy function of dense associative memory models gives BMs belonging to exponential family harmoniums. Toshihiro Ota, Ryo Karakida |
Neural Comput. | 2 |
| 2023 | Deep learning in random neural fields: Numerical experiments via neural tangent kernel
Kaito Watanabe, Kotaro Sakamoto, Ryo Karakida, Sho Sonoda, Shun-ichi Amari |
Neural Networks | 3 |
| 2022 | Learning curves for continual learning in neural networks: Self-knowledge transfer and forgetting
Ryo Karakida, Shotaro Akaho |
ICLR | 1 |
| 2021 | The Spectrum of Fisher Information of Deep Networks Achieving Dynamical IsometryabstractThe Fisher information matrix (FIM) is fundamental to understanding the trainability of deep neural nets (DNN), since it describes the parameter space’s local metric. We investigate the spectral distribution of the conditional FIM, which is the FIM given a single sample, by focusing on fully-connected networks achieving dynamical isometry. Then, while dynamical isometry is known to keep specific backpropagated signals independent of the depth, we find that the parameter space’s local metric linearly depends on the depth even under the dynamical isometry. More precisely, we reveal that the conditional FIM’s spectrum concentrates around the maximum and the value grows linearly as the depth increases. To examine the spectrum, considering random initialization and the wide limit, we construct an algebraic methodology based on the free probability theory. As a byproduct, we provide an analysis of the solvable spectral distribution in two-hidden-layer cases. Lastly, experimental results verify that the appropriate learning rate for the online training of DNNs is in inverse proportional to depth, which is determined by the conditional FIM’s spectrum. Tomohiro Hayase, Ryo Karakida |
AISTATS | 2 |
| 2021 | Self-paced data augmentation for training neural networks
Tomoumi Takase, Ryo Karakida, Hideki Asoh |
Neurocomputing | 2 |
| 2021 | Pathological Spectra of the Fisher Information Metric and Its Variants in Deep Neural NetworksabstractThe Fisher information matrix (FIM) plays an essential role in statistics and machine learning as a Riemannian metric tensor or a component of the Hessian matrix of loss functions. Focusing on the FIM and its variants in deep neural networks (DNNs), we reveal their characteristic scale dependence on the network width, depth, and sample size when the network has random weights and is sufficiently wide. This study covers two widely used FIMs for regression with linear output and for classification with softmax output. Both FIMs asymptotically show pathological eigenvalue spectra in the sense that a small number of eigenvalues become large outliers depending on the width or sample size, while the others are much smaller. It implies that the local shape of the parameter space or loss landscape is very sharp in a few specific directions while almost flat in the other directions. In particular, the softmax output disperses the outliers and makes a tail of the eigenvalue density spread from the bulk. We also show that pathological spectra appear in other variants of FIMs: one is the neural tangent kernel; another is a metric for the input signal and feature space that arises from feedforward signal propagation. Thus, we provide a unified perspective on the FIM and its variants that will lead to more quantitative understanding of learning in large-scale DNNs. Ryo Karakida, Shotaro Akaho, Shun-ichi Amari |
Neural Comput. | 1 |
| 2020 | Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksabstractNatural Gradient Descent (NGD) helps to accelerate the convergence of gradient descent dynamics, but it requires approximations in large-scale deep neural networks because of its high computational cost. Empirical studies have confirmed that some NGD methods with approximate Fisher information converge sufficiently fast in practice. Nevertheless, it remains unclear from the theoretical perspective why and under what conditions such heuristic approximations work well. In this work, we reveal that, under specific conditions, NGD with approximate Fisher information achieves the same fast convergence to global minima as exact NGD. We consider deep neural networks in the infinite-width limit, and analyze the asymptotic training dynamics of NGD in function space via the neural tangent kernel. In the function space, the training dynamics with the approximate Fisher information are identical to those with the exact Fisher information, and they converge quickly. The fast convergence holds in layer-wise approximations; for instance, in block diagonal approximation where each block corresponds to a layer as well as in block tri-diagonal and K-FAC approximations. We also find that a unit-wise approximation achieves the same fast convergence under some assumptions. All of these different approximations have an isotropic gradient in the function space, and this plays a fundamental role in achieving the same convergence properties in training. Thus, the current study gives a novel and unified theoretical foundation with which to understand NGD methods in deep learning. Ryo Karakida, Kazuki Osawa |
NeurIPS | 1 |
| 2019 | Fisher Information and Natural Gradient Learning in Random Deep NetworksabstractThe parameter space of a deep neural network is a Riemannian manifold, where the metric is defined by the Fisher information matrix. The natural gradient method uses the steepest descent direction in a Riemannian manifold, but it requires inversion of the Fisher matrix, however, which is practically difficult. The present paper uses statistical neurodynamical method to reveal the properties of the Fisher information matrix in a net of random connections. We prove that the Fisher information matrix is unit-wise block diagonal supplemented by small order terms of off-block-diagonal elements. We further prove that the Fisher information matrix of a single unit has a simple reduced form, a sum of a diagonal matrix and a rank 2 matrix of weight-bias correlations. We obtain the inverse of Fisher information explicitly. We then have an explicit form of the approximate natural gradient, without relying on the matrix inversion. Shun-ichi Amari, Ryo Karakida, Masafumi Oizumi |
AISTATS | 2 |
| 2019 | Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field ApproachabstractThe Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large width limits, which enables us to utilize mean field theories. We investigate the asymptotic statistics of the FIM’s eigenvalues and reveal that most of them are close to zero while the maximum eigenvalue takes a huge value. Because the landscape of the parameter space is defined by the FIM, it is locally flat in most dimensions, but strongly distorted in others. Moreover, we demonstrate the potential usage of the derived statistics in learning strategies. First, small eigenvalues that induce flatness can be connected to a norm-based capacity measure of generalization ability. Second, the maximum eigenvalue that induces the distortion enables us to quantitatively estimate an appropriately sized learning rate for gradient methods to converge. Ryo Karakida, Shotaro Akaho, Shun-ichi Amari |
AISTATS | 1 |
| 2019 | The Normalization Method for Alleviating Pathological Sharpness in Wide Neural NetworksabstractNormalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information matrix (FIM), which also corresponds to the local shape of the loss landscape under certain conditions. We analyze deep neural networks with random initialization, which is known to suffer from a pathologically sharp shape of the landscape when the network becomes sufficiently wide. We reveal that batch normalization in the last layer contributes to drastically decreasing such pathological sharpness if the width and sample number satisfy a specific condition. In contrast, it is hard for batch normalization in the middle hidden layers to alleviate pathological sharpness in many settings. We also found that layer normalization cannot alleviate pathological sharpness either. Thus, we can conclude that batch normalization in the last layer significantly contributes to decreasing the sharpness induced by the FIM. Ryo Karakida, Shotaro Akaho, Shun-ichi Amari |
NeurIPS | 1 |
| 2019 | Information Geometry for Regularized Optimal Transport and Barycenters of PatternsabstractWe propose a new divergence on the manifold of probability distributions, building on the entropic regularization of optimal transportation problems. As Cuturi ( 2013 ) showed, regularizing the optimal transport problem with an entropic term is known to bring several computational benefits. However, because of that regularization, the resulting approximation of the optimal transport cost does not define a proper distance or divergence between probability distributions. We recently tried to introduce a family of divergences connecting the Wasserstein distance and the Kullback-Leibler divergence from an information geometry point of view (see Amari, Karakida, & Oizumi, 2018 ). However, that proposal was not able to retain key intuitive aspects of the Wasserstein geometry, such as translation invariance, which plays a key role when used in the more general problem of computing optimal transport barycenters. The divergence we propose in this work is able to retain such properties and admits an intuitive interpretation. Shun-ichi Amari, Ryo Karakida, Masafumi Oizumi, Marco Cuturi |
Neural Comput. | 2 |
| 2018 | Dynamics of Learning in MLP: Natural Gradient and Singularity RevisitedabstractThe dynamics of supervised learning play a main role in deep learning, which takes place in the parameter space of a multilayer perceptron (MLP). We review the history of supervised stochastic gradient learning, focusing on its singular structure and natural gradient. The parameter space includes singular regions in which parameters are not identifiable. One of our results is a full exploration of the dynamical behaviors of stochastic gradient learning in an elementary singular network. The bad news is its pathological nature, in which part of the singular region becomes an attractor and another part a repulser at the same time, forming a Milnor attractor. A learning trajectory is attracted by the attractor region, staying in it for a long time, before it escapes the singular region through the repulser region. This is typical of plateau phenomena in learning. We demonstrate the strange topology of a singular region by introducing blow-down coordinates, which are useful for analyzing the natural gradient dynamics. We confirm that the natural gradient dynamics are free of critical slowdown. The second main result is the good news: the interactions of elementary singular networks eliminate the attractor part and the Milnor-type attractors disappear. This explains why large-scale networks do not suffer from serious critical slowdowns due to singularities. We finally show that the unit-wise natural gradient is effective for learning in spite of its low computational cost. Shun-ichi Amari, Tomoko Ozeki, Ryo Karakida, Masato Okada |
Neural Comput. | 3 |
| 2016 | Maximum likelihood learning of RBMs with Gaussian visible units on the Stiefel manifold
Ryo Karakida, Masato Okada, Shun-ichi Amari |
ESANN | 1 |
| 2016 | Adaptive Natural Gradient Learning Algorithms for Unnormalized Statistical Models
Ryo Karakida, Masato Okada, Shun-ichi Amari |
ICANN (1) | 1 |
| 2016 | Dynamical analysis of contrastive divergence learning: Restricted Boltzmann machines with Gaussian visible units
Ryo Karakida, Masato Okada, Shun-ichi Amari |
Neural Networks | 1 |