Shingo Yashima

dblp:254/3096 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 61% Probabilistic and Bayesian machine learning · 12% Kernel, tree and ensemble methods · 12%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 54% Hardware accelerators and domain-specific architectures · 46%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.422024
SAS: Structured Activation Sparsification · ICLR 2024
Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity
0.812024
SAS: Structured Activation Sparsification · ICLR 2024
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core
0.812024
SAS: Structured Activation Sparsification · ICLR 2024
Machine learning › Efficient and distributed learning › model compression
pruning
0.712023
Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023
Hardware accelerators and domain-specific architectures
sparse computation
0.712023
Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning
neural network ensemble
0.612022
Feature Space Particle Inference for Neural Network Ensembles · ICML 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Feature Space Particle Inference for Neural Network Ensembles · ICML 2022
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.212024
SAS: Structured Activation Sparsification · ICLR 2024
Machine learning › Efficient and distributed learning › model compression
sparse neural network
0.212023
Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023

Methods — techniques the papers use, named apart from their topics

structured sparsity · 1.5sparsetensorcore · 1.5n:m sparsity · 1.5sparse dot-product · 1.3bit-pruning · 1.3feature space particle optimization · 0.6bayesian inference · 0.6
YearPublicationVenuePosition
2024 SAS: Structured Activation Sparsification
abstract
Wide networks usually yield better accuracy than their narrower counterpart at the expense of the massive $\texttt{mult}$ cost. To break this tradeoff, we advocate a novel concept of $\textit{Structured Activation Sparsification}$, dubbed SAS, which boosts accuracy without increasing computation by utilizing the projected sparsity in activation maps with a specific structure. Concretely, the projected sparse activation is allowed to have N nonzero value among M consecutive activations. Owing to the local structure in sparsity, the wide $\texttt{matmul}$ between a dense weight and the sparse activation is executed as an equivalent narrow $\texttt{matmul}$ between a dense weight and dense activation, which is compatible with NVIDIA's $\textit{SparseTensorCore}$ developed for the N:M structured sparse weight. In extensive experiments, we demonstrate that increasing sparsity monotonically improves accuracy (up to 7% on CIFAR10) without increasing the $\texttt{mult}$ count. Furthermore, we show that structured sparsification of $\textit{activation}$ scales better than that of $\textit{weight}$ given the same computational budget.
Yusuke Sekikawa, Shingo Yashima
ICLR2
2023 Bit-Pruning: A Sparse Multiplication-Less Dot-Product
Yusuke Sekikawa, Shingo Yashima
ICLR2
2022 Multi-task Curriculum Learning based on Gradient Similarity
Hiroaki Igarashi, Kenichi Yoneji, Kohta Ishikawa, Rei Kawakami, Teppei Suzuki, Shingo Yashima, Ikuro Sato
BMVC6
2022 Feature Space Particle Inference for Neural Network Ensembles
abstract
Ensembles of deep neural networks demonstrate improved performance over single models. For enhancing the diversity of ensemble members while keeping their performance, particle-based inference methods offer a promising approach from a Bayesian perspective. However, the best way to apply these methods to neural networks is still unclear: seeking samples from the weight-space posterior suffers from inefficiency due to the over-parameterization issues, while seeking samples directly from the function-space posterior often leads to serious underfitting. In this study, we propose to optimize particles in the feature space where activations of a specific intermediate layer lie to alleviate the abovementioned difficulties. Our method encourages each member to capture distinct features, which are expected to increase the robustness of the ensemble prediction. Extensive evaluation on real-world datasets exhibits that our model significantly outperforms the gold-standard Deep Ensembles on various metrics, including accuracy, calibration, and robustness.
Shingo Yashima, Teppei Suzuki, Kohta Ishikawa, Ikuro Sato, Rei Kawakami
ICML1
2021 Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features
abstract
Although kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to derive computationally efficient learning algorithms. We consider solving a binary classification problem using random features and stochastic gradient descent, both of which are common and widely used in practical large-scale problems. Although there are plenty of previous works investigating the efficiency of these algorithms in terms of the convergence of the objective loss function, these results suggest that the computational gain comes at expense of the learning accuracy when dealing with general Lipschitz loss functions such as logistic loss. In this study, we analyze the properties of these algorithms in terms of the convergence not of the loss function, but the classification error under the strong low-noise condition, which reflects a realistic property of real-world datasets. We extend previous studies on SGD to a random features setting, examining a novel analysis about the error induced by the approximation of random features in terms of the distance between the generated hypothesis to show that an exponential convergence of the expected classification error is achieved even if random features approximation is applied. We demonstrate that the convergence rate does not depend on the number of features and there is a significant computational benefit in using random features in classification problems under the strong low-noise condition.
Shingo Yashima, Atsushi Nitanda, Taiji Suzuki
AISTATS1