Tomohiro Hayase

dblp:218/5945 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-6453-4317ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Deep learning architectures and training · 43% Efficient and distributed learning · 23% Learning theory · 21%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › feedforward neural network
MLP-based architecture
0.812024
Understanding MLP-Mixer as a wide and sparse MLP · ICML 2024
Machine learning › Deep learning architectures and training › feedforward neural network › MLP-based architecture
MLP-Mixer
0.812024
Understanding MLP-Mixer as a wide and sparse MLP · ICML 2024
Machine learning › Learning theory
neural network theory
0.812024
Understanding MLP-Mixer as a wide and sparse MLP · ICML 2024
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.812024
Understanding MLP-Mixer as a wide and sparse MLP · ICML 2024
Machine learning › Efficient and distributed learning › model compression › sparsity
sparse parameterization
0.812024
Understanding MLP-Mixer as a wide and sparse MLP · ICML 2024
Machine learning › Optimization for machine learning › gradient estimation
finite-difference approximation
0.712023
Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias · ICML 2023
Machine learning › Deep learning architectures and training › regularization
gradient regularization
0.712023
Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias · ICML 2023
Machine learning › Learning theory
implicit bias
0.712023
Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias · ICML 2023
Machine learning › Deep learning architectures and training
regularization
0.712023
Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias · ICML 2023
Machine learning › Optimization for machine learning › optimization landscape
flat minima
0.212023
Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias · ICML 2023

Methods — techniques the papers use, named apart from their topics

sparsity analysis · 0.8monarch matrices · 0.8gradient descent · 0.7gradient ascent · 0.7finite-difference computation · 0.7
YearPublicationVenuePosition
2024 Understanding MLP-Mixer as a wide and sparse MLP
abstract
Multi-layer perceptron (MLP) is a fundamental component of deep learning, and recent MLP-based architectures, especially the MLP-Mixer, have achieved significant empirical success. Nevertheless, our understanding of why and how the MLP-Mixer outperforms conventional MLPs remains largely unexplored. In this work, we reveal that sparseness is a key mechanism underlying the MLP-Mixers. First, the Mixers have an effective expression as a wider MLP with Kronecker-product weights, clarifying that the Mixers efficiently embody several sparseness properties explored in deep learning. In the case of linear layers, the effective expression elucidates an implicit sparse regularization caused by the model architecture and a hidden relation to Monarch matrices, which is also known as another form of sparse parameterization. Next, for general cases, we empirically demonstrate quantitative similarities between the Mixer and the unstructured sparse-weight MLPs. Following a guiding principle proposed by Golubeva, Neyshabur and Gur-Ari (2021), which fixes the number of connections and increases the width and sparsity, the Mixers can demonstrate improved performance.
Tomohiro Hayase, Ryo Karakida
ICML1
2023 Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias
abstract
Gradient regularization (GR) is a method that penalizes the gradient norm of the training loss during training. While some studies have reported that GR can improve generalization performance, little attention has been paid to it from the algorithmic perspective, that is, the algorithms of GR that efficiently improve the performance. In this study, we first reveal that a specific finite-difference computation, composed of both gradient ascent and descent steps, reduces the computational cost of GR. Next, we show that the finite-difference computation also works better in the sense of generalization performance. We theoretically analyze a solvable model, a diagonal linear network, and clarify that GR has a desirable implicit bias to so-called rich regime and finite-difference computation strengthens this bias. Furthermore, finite-difference GR is closely related to some other algorithms based on iterative ascent and descent steps for exploring flat minima. In particular, we reveal that the flooding method can perform finite-difference GR in an implicit way. Thus, this work broadens our understanding of GR for both practice and theory.
Ryo Karakida, Tomoumi Takase, Tomohiro Hayase, Kazuki Osawa
ICML3
2022 Downstream Augmentation Generation For Contrastive Learning
abstract
Contrastive learning has become one of the most promising approaches for learning image representations. However, it heavily relies on heuristic data augmentation techniques, such as Gaussian blurring and color jittering, for making image pairs to be contrastively compared. These augmentations are not always appropriate for downstream tasks that each have their own camera and illumination settings. In this paper, we aim at improving the augmentation process and propose an augmentation generator, a network that learns to augment images for contrastive learning. Under the assumption that each downstream task has an optimal implicit augmentation function, the augmentation generator enhances the contrastive learning by estimating it. We demonstrate the effectiveness of our learning framework on two combined datasets, EMNIST-Omniglot and ImageNet-DAISO.
Tomohiro Hayase, Suguru Yasutomi, Nakamasa Inoue
ICASSP1
2021 The Spectrum of Fisher Information of Deep Networks Achieving Dynamical Isometry
abstract
The Fisher information matrix (FIM) is fundamental to understanding the trainability of deep neural nets (DNN), since it describes the parameter space’s local metric. We investigate the spectral distribution of the conditional FIM, which is the FIM given a single sample, by focusing on fully-connected networks achieving dynamical isometry. Then, while dynamical isometry is known to keep specific backpropagated signals independent of the depth, we find that the parameter space’s local metric linearly depends on the depth even under the dynamical isometry. More precisely, we reveal that the conditional FIM’s spectrum concentrates around the maximum and the value grows linearly as the depth increases. To examine the spectrum, considering random initialization and the wide limit, we construct an algebraic methodology based on the free probability theory. As a byproduct, we provide an analysis of the solvable spectral distribution in two-hidden-layer cases. Lastly, experimental results verify that the appropriate learning rate for the online training of DNNs is in inverse proportional to depth, which is determined by the conditional FIM’s spectrum.
Tomohiro Hayase, Ryo Karakida
AISTATS1
2021 Layer-Wise Interpretation of Deep Neural Networks using Identity Initialization
abstract
The interpretability of neural networks (NNs) is a challenging but essential topic for transparency in the decision-making process using machine learning. One of the reasons for the lack of interpretability is random weight initialization, where the input is randomly embedded into a different feature space in each layer. In this paper, we propose an interpretation method for a deep multilayer perceptron, which is the most general architecture of NNs, based on identity initialization (namely, initialization using identity matrices). The proposed method allows us to analyze the contribution of each neuron to classification and class likelihood in each hidden layer. As a property of the identity-initialized perceptron, the weight matrices remain near the identity matrices even after learning. This property enables us to treat the change of features from the input to each hidden layer as the contribution to classification. Furthermore, we can separate the output of each hidden layer into a contribution map that depicts the contribution to classification and class likelihood, by adding extra dimensions to each layer according to the number of classes, thereby allowing the calculation of the recognition accuracy in each layer and thus revealing the roles of independent layers, such as feature extraction and classification.
Shohei Kubota, Hideaki Hayashi, Tomohiro Hayase, Seiichi Uchida
ICASSP3