EDBT 2026 Demo / reviewers in the wild / expert
Tsuguchika Tabaru
dblp:174/7809
· DBLP profile ↗
7ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0001-6568-1968ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A novel structured sparse fully connected layer in convolutional neural networksabstractAbstract Convolutional Neural Networks (CNNs) are one of the factors supporting the rapid development of artificial intelligent techniques. However, as the ability of the network increases, the size of the network becomes larger. Thus far, several works related to reduction of the network size have been tackled. In many cases, these approaches produce an unstructured network which prevents efficient parallel computation. To avoid this problem, we propose a novel structured sparse fully connected layer (FCL) in the CNNs. The aim of our proposed approach is reduction of the number of network parameters in the FCLs which occupy a large part of network parameters. Unlike the general FCLs used in the popular CNNs such as VGG‐16, the proposed approach reduces the connection between the last convolutional layer and the first FCL. In addition, we show an implementation for the proposed sparse FCLs on the GPU using cuBLAS. As a result for ILSVRC‐2012 dataset, the proposed approach achieves a 21.3 times compression with 0.68% top‐1 accuracy and 0.31% top‐5 accuracy decreases for VGG‐16. The implementation of the proposed FCLs achieves speed‐up factor 14.97 and 16.67 for forward and backward propagation compared to that for the noncompressed FCLs, respectively. Naoki Matsumura, Yasuaki Ito, Koji Nakano, Akihiko Kasagi, Tsuguchika Tabaru |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | NEEBS: Nonexpert large-scale environment building system for deep neural networkabstractSummary Deep neural networks (DNNs) have greatly improved the accuracy of various tasks in areas such as natural language processing (NLP). Obtaining a highly accurate DNN model requires multiple repetitions of training on a huge dataset, which requires a large‐scale cluster the compute nodes of which are tightly connected by high‐speed interconnects to exchange a large amount of intermediate data with very short latency. However, fully using the computational power of a large‐scale cluster for training requires knowledge of its components such as a distributed file system, an interconnection, and optimized high‐performance libraries. We have developed a Non‐Expert large‐scale Environment Building System (NEEBS) that aids a user in building a fast‐running training environment on a large‐scale cluster. It automatically installs and configures the applications and necessary libraries. It also optimally prepares tools to stage both data and executable programs, and launcher scripts suitable for both the applications and job submission systems of the cluster. NEEBS achieves 93.91% throughput scalability in NLP pretraining. We also present an approach to reduce pretraining time of highly accurate DNN model for NLP using a large‐scale computation environment built using NEEBS. We trained a Bidirectional Encoder Representations from Transformers (BERT)‐3.9b and a BERT‐xlarge using a dense masked language model (MLM) on Megatron‐LM framework and evaluated the improvement in learning time and learning efficiency for a Japanese language dataset using 768 graphics processing units (GPUs) on the AI Bridging Cloud Infrastructure (ABCI). Our implementation NEEBS improved learning efficiency per iteration by a factor of 10 and completed the pretraining of BERT‐xlarge in 4.7 h. This pretraining takes 5 months on a single GPU. To determine if the BERT models are correctly pretrained, we evaluated their accuracy in two tasks, Stanford Natural Language Inference Corpus translated into Japanese (JSNLI) and Twitter reputation analysis (TwitterRA). BERT‐3.9b achieved 94.30% accuracy for JSNLI, and BERT‐xlarge achieved 90.63% accuracy for TwitterRA. We constructed pretrained models with comparable accuracy to other Japanese BERT models in a shorter time. Yoshiharu Tajima, Masahiro Asaoka, Akihiro Tabuchi, Akihiko Kasagi, Tsuguchika Tabaru |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | Automatic Pruning Rate Derivation for Structured Pruning of Deep Neural NetworksabstractTo compress the neural network model, structured pruning has been proposed. However, finding a proper pruning rate to suppress the accuracy degradation of pruned models is difficult because existing structured pruning methods assign the pruning rate manually. As described herein, we propose an automatic pruning rate derivation method for structured pruning to reduce the workload of inefficient manual pruning rate assignment. The value of the pruning error (L1-norm of the pruned weight) depends on the pruning rate. Therefore, to derive the pruning rate, our method compares the pruning error and the threshold. When the pruning error is less than the threshold, the degradation of the pruned model accuracy is suppressed. We demonstrate the superiority of our proposed method over state-of-the-art methods on CIFAR-10 and ImageNet using various ResNets. For example, the proposed method reduces 56.2% parameters of ResNet-50 with similar accuracy of 75.32% to earlier works on ImageNet. Yasufumi Sakai, Akinori Iwakawa, Tsuguchika Tabaru, Atsuki Inoue, Hiroshi Kawaguchi 0001 |
ICPR | 3 |
| 2021 | The 16, 384-node Parallelism of 3D-CNN Training on An Arm CPU based SupercomputerabstractAs the computational cost and datasets available for deep neural network training continue to increase, there is a significant demand for fast distributed training on supercomputers. However, porting and tuning applications for new advanced supercomputers requires tremendous amount of development efforts. Therefore, we present software tuning best practice for a 3D-CNN model training on a new Arm CPU based supercomputer, Fugaku. We (i) tune computation in DL by a JIT translator for aarch64, (ii) optimize collective communication such as Allreduce for 6D mesh/torus network topology, (iii) tune I/O by data staging with compression and data loader with caching, and (iv) parallelize training in data and model parallelism. We apply the proposed methods to a CosmoFlow 3D-CNN model, and achieve the training in 30 minutes using 16,384 nodes consisting of 4096 data- and 4 model-parallelism. This is the fastest result of any CPU-based systems in MLPerf HPC v0.7 in the world. Akihiro Tabuchi, Koichi Shirahata, Masafumi Yamazaki, Akihiko Kasagi, Takumi Honda, Kouji Kurihara, Kentaro Kawakami, Tsuguchika Tabaru, Naoto Fukumoto, Akiyoshi Kuroda, Takaaki Fukai, Kento Sato |
HiPC | 8 |
| 2020 | An Efficient Technique for Large Mini-batch Challenge of DNNs Training on Large Scale ClusterabstractDistributed deep learning using large mini-batches is a key strategy to perform the deep learning as fast as possible, but it represents a great challenge as it is difficult to achieve high scaling efficiency when using large clusters without compromising accuracy. The particular problem in this challenge is decreasing the number of model update iterations in whole of training. Thus, we need a technique which can converge the validation accuracy with a small number of iterations to address this challenge. In this paper, we introduce a novel technique, Final Polishing. This technique adjusts the means and variances in the batch normalization and mitigates the difference of normalization between validation datasets and augmented training datasets. By applying the technique, we achieved top-1 validation accuracy of 75.08% with mini-batch size of 81,920, with 2,048 GPUs and completed the training of ResNet-50 in 74.7 seconds.In addition, targeting top-1 validation accuracy of 75.9% or more, we tried additional parameters tuning. Then, we adjusted the number of GPUs and hyperparameters of DNNs with Final Polishing, and we also achieved top-1 validation accuracy of 75.97% with mini-batch size of 86,016, with 3,072 GPUs and completed the training of ResNet-50 in 62.1 seconds. Akihiko Kasagi, Akihiro Tabuchi, Masafumi Yamazaki, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, Kohta Nakashima |
HPDC | 7 |
| 2020 | Huffman Coding with Gap Arrays for GPU AccelerationabstractHuffman coding is a fundamental lossless data compression scheme used in many data compression file formats such as gzip, zip, png, and jpeg. Huffman encoding is easily parallelized, because all 8-bit symbols can be converted into codewords independently. On the other hand, since an encoded codeword sequence has no separator to identify each codeword, parallelizing Huffman decoding is a much harder task. This work presents a new data structure called gap array to be attached to an encoded codeword sequence of Huffman coding for accelerating parallel Huffman decoding. In addition, it also shows that GPU Huffman encoding and decoding can be accelerated by several techniques including (1) the Single Kernel Soft Synchronization (SKSS), (2) wordwise global memory access and (3) compact codebooks. The experimental results for 10 files on NVIDIA Tesla V100 GPU show that our GPU Huffman encoding and decoding run 2.87x-7.70x times and 1.26x-2.63x times faster than previously presented GPU Huffman encoding and decoding, respectively. Also, Huffman decoding can be further accelerated by a factor of 1.67x-6450x if a gap array is attached to an encoded codeword sequence. Since the size and computing overhead of gap arrays in Huffman encoding are small, we can conclude that gap arrays should be introduced for GPU Huffman encoding and decoding. Naoya Yamamoto, Koji Nakano, Yasuaki Ito, Daisuke Takafuji, Akihiko Kasagi, Tsuguchika Tabaru |
ICPP | 6 |
| 2020 | Efficient convolution pooling on the GPU
Shunsuke Suita, Takahiro Nishimura, Hiroki Tokura, Koji Nakano, Yasuaki Ito, Akihiko Kasagi, Tsuguchika Tabaru |
J. Parallel Distributed Comput. | 7 |