EDBT 2026 Demo / reviewers in the wild / expert
Shingo Yashima
dblp:254/3096
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 61% Probabilistic and Bayesian machine learning · 12% Kernel, tree and ensemble methods · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 54% Hardware accelerators and domain-specific architectures · 46% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.4 | 2 | 2024 | SAS: Structured Activation Sparsification · ICLR 2024 Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.8 | 1 | 2024 | SAS: Structured Activation Sparsification · ICLR 2024 |
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core |
0.8 | 1 | 2024 | SAS: Structured Activation Sparsification · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.7 | 1 | 2023 | Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023 |
Hardware accelerators and domain-specific architectures
sparse computation |
0.7 | 1 | 2023 | Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
neural network ensemble |
0.6 | 1 | 2022 | Feature Space Particle Inference for Neural Network Ensembles · ICML 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.6 | 1 | 2022 | Feature Space Particle Inference for Neural Network Ensembles · ICML 2022 |
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks |
0.2 | 1 | 2024 | SAS: Structured Activation Sparsification · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
sparse neural network |
0.2 | 1 | 2023 | Bit-Pruning: A Sparse Multiplication-Less Dot-Product · ICLR 2023 |
Methods — techniques the papers use, named apart from their topics
structured sparsity · 1.5sparsetensorcore · 1.5n:m sparsity · 1.5sparse dot-product · 1.3bit-pruning · 1.3feature space particle optimization · 0.6bayesian inference · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SAS: Structured Activation SparsificationabstractWide networks usually yield better accuracy than their narrower counterpart at the expense of the massive $\texttt{mult}$ cost.
To break this tradeoff, we advocate a novel concept of $\textit{Structured Activation Sparsification}$, dubbed SAS, which boosts accuracy without increasing computation by utilizing the projected sparsity in activation maps with a specific structure.
Concretely, the projected sparse activation is allowed to have N nonzero value among M consecutive activations.
Owing to the local structure in sparsity, the wide $\texttt{matmul}$ between a dense weight and the sparse activation is executed as an equivalent narrow $\texttt{matmul}$ between a dense weight and dense activation, which is compatible with NVIDIA's $\textit{SparseTensorCore}$ developed for the N:M structured sparse weight.
In extensive experiments, we demonstrate that increasing sparsity monotonically improves accuracy (up to 7% on CIFAR10) without increasing the $\texttt{mult}$ count.
Furthermore, we show that structured sparsification of $\textit{activation}$ scales better than that of $\textit{weight}$ given the same computational budget. Yusuke Sekikawa, Shingo Yashima |
ICLR | 2 |
| 2023 | Bit-Pruning: A Sparse Multiplication-Less Dot-Product
Yusuke Sekikawa, Shingo Yashima |
ICLR | 2 |
| 2022 | Multi-task Curriculum Learning based on Gradient Similarity
Hiroaki Igarashi, Kenichi Yoneji, Kohta Ishikawa, Rei Kawakami, Teppei Suzuki, Shingo Yashima, Ikuro Sato |
BMVC | 6 |
| 2022 | Feature Space Particle Inference for Neural Network EnsemblesabstractEnsembles of deep neural networks demonstrate improved performance over single models. For enhancing the diversity of ensemble members while keeping their performance, particle-based inference methods offer a promising approach from a Bayesian perspective. However, the best way to apply these methods to neural networks is still unclear: seeking samples from the weight-space posterior suffers from inefficiency due to the over-parameterization issues, while seeking samples directly from the function-space posterior often leads to serious underfitting. In this study, we propose to optimize particles in the feature space where activations of a specific intermediate layer lie to alleviate the abovementioned difficulties. Our method encourages each member to capture distinct features, which are expected to increase the robustness of the ensemble prediction. Extensive evaluation on real-world datasets exhibits that our model significantly outperforms the gold-standard Deep Ensembles on various metrics, including accuracy, calibration, and robustness. Shingo Yashima, Teppei Suzuki, Kohta Ishikawa, Ikuro Sato, Rei Kawakami |
ICML | 1 |
| 2021 | Exponential Convergence Rates of Classification Errors on Learning with SGD and Random FeaturesabstractAlthough kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to derive computationally efficient learning algorithms. We consider solving a binary classification problem using random features and stochastic gradient descent, both of which are common and widely used in practical large-scale problems. Although there are plenty of previous works investigating the efficiency of these algorithms in terms of the convergence of the objective loss function, these results suggest that the computational gain comes at expense of the learning accuracy when dealing with general Lipschitz loss functions such as logistic loss. In this study, we analyze the properties of these algorithms in terms of the convergence not of the loss function, but the classification error under the strong low-noise condition, which reflects a realistic property of real-world datasets. We extend previous studies on SGD to a random features setting, examining a novel analysis about the error induced by the approximation of random features in terms of the distance between the generated hypothesis to show that an exponential convergence of the expected classification error is achieved even if random features approximation is applied. We demonstrate that the convergence rate does not depend on the number of features and there is a significant computational benefit in using random features in classification problems under the strong low-noise condition. Shingo Yashima, Atsushi Nitanda, Taiji Suzuki |
AISTATS | 1 |