EDBT 2026 Demo / reviewers in the wild / expert
Akihiko Kasagi
dblp:125/2799
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-5793-335XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPU-Accelerated One-Electron Integral Computation for Quantum ChemistryabstractABSTRACT In Quantum chemical computation, numerical schemes such as the Hartree–Fock (HF) and density functional theory (DFT) are widely used to solve the Schrödinger equation numerically, to realize experiment‐free prediction and analysis of key molecular properties such as structure and energy. Computing one‐electron integrals, such as kinetic energy integrals and nuclear attraction integrals, is essential in both HF and DFT to characterize the molecular electronic states. However, as molecules of practical interest grow in size and angular momentum, computing one‐electron orbitals becomes computationally expensive in most cases. Although computing kinetic energy integrals on CPUs is straightforward, bottlenecks in CPU‐GPU data transfer have often been overlooked. In this study, we propose an efficient method to compute both the kinetic‐energy and nuclear‐attractive integrals on GPUs. First, we explicitly and symbolically expand recurrence relations based on the Obara–Saika and McMurchie–Davidson methods to eliminate redundant operations, thus improving computational efficiency. Second, we implemented a hybrid method that selects the best/fastest of both methods depending on the integration task. Third, we achieved further speedups by using CUDA streams to parallelize the execution of multiple kernels and efficiently utilize multiprocessor resources on the GPU. Computational experiments using NVIDIA A100 GPUs and Intel Xeon Gold 6338 CPU on relevant molecules of interest demonstrated the superiority of our one‐electron integral GPU implementations, achieving a speedup of 20.2 times over PySCF, and a speedup of 132.6 times over GPU4PySCF. Nobuya Yokogawa, Yasuaki Ito, Satoki Tsuji, Haruto Fujii, Kanta Suzuki, Koji Nakano, Victor Parque, Akihiko Kasagi |
Concurr. Comput. Pract. Exp. | 8 |
| 2025 | Efficient GPU Implementations of Three-Center Two-Electron Repulsion IntegralsabstractABSTRACT In computational quantum chemistry, the computation of three‐center two‐electron repulsion integrals (also termed three‐center ERIs) is essential for density fitting. Due to the large number of integral elements and the induced combinatorial computational complexity, the community has actively pursued the acceleration/speedup of ERI calculations to achieve pragmatic levels of efficiency. From the perspective of GPU acceleration, atomicAdd is known to incur significant memory overhead: The frequent collisions and retrials of value aggregation in global GPU memory lead to substantial performance degradation. To tackle this issue, we propose new thread mapping strategies for three‐center two‐electron integrals on GPUs, aiming at reducing the computational cost associated with value aggregation. Our methods are based on the idea of suitable substitutions of device‐level reduction ( atomicAdd ) with efficient warp‐ and thread‐level reduction, such as warp‐shuffle and register accumulation. As a result, our computational experiments using an Intel Xeon Gold 6338 CPU, an NVIDIA A100 GPU, and relevant molecules of interest show the superiority against the conventional thread mapping scheme, achieving up to 2.76 speedups to compute three‐center ERIs more efficiently. Moreover, compared to well‐known quantum chemistry software such as PySCF and GPU4PySCF, our method achieved up to speedups over PySCF and up to speedups over GPU4PySCF. Our method has the potential to further enhance the performance, extensibility, and versatility of GPU‐accelerated quantum chemical computations. Kanta Suzuki, Yasuaki Ito, Haruto Fujii, Nobuya Yokogawa, Satoki Tsuji, Koji Nakano, Victor Parque, Akihiko Kasagi |
Concurr. Comput. Pract. Exp. | 8 |
| 2025 | GPU Acceleration of the Boys Function Evaluation in Computational Quantum ChemistryabstractABSTRACT The Boys function, a mathematical integral function, plays a pivotal role and is frequently evaluated in ab initio molecular orbital computations. The main contribution of this paper is to accelerate the bulk evaluation of the Boys function through the effective utilization of GPUs. The proposed GPU implementation addresses GPU‐specific programming issues such as warp divergence and coalesced/stride access to global memory, and we employ the optimal numerical evaluation method from four methods based on input values to ensure efficient computation with sufficient accuracy. Moreover, to consider actual computation of molecular integrals, we have implemented and evaluated the proposed method in two scenarios: single evaluation, which computes a single value of the Boys function for a single input, and incremental evaluation, which computes multiple values of the Boys function incrementally. The execution time of the proposed GPU implementation was evaluated for both scenarios using an NVIDIA A100 Tensor Core GPU. As a result, the GPU‐accelerated bulk evaluation has achieved a throughput of computing the values of the Boys function times per second for the single evaluation and times per second for the incremental evaluation, respectively. Our parallelized CPU and GPU implementation is available at https://github.com/sstsuji/Boys‐function‐GPU‐library . Satoki Tsuji, Yasuaki Ito, Koji Nakano, Akihiko Kasagi |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Performance Analysis of Quantum Computer Simulators Across Different EnvironmentsabstractQuantum computers can achieve extremely fast computations for certain problems using quantum properties. Due to these factors, research and development in the field of quantum computers have been vigorously pursued. It is crucial to develop a quantum computer that operates accurately and, concurrently, create applications compatible with quantum computing. Quantum computer simulators are tools that represent the behavior of a quantum computer on classical computers, and they are highly conducive to the development of quantum applications that can run on actual quantum computers. On the other hand, quantum computer simulators run on classical computers, which are extremely slow compared to quantum computers. Therefore, it is desirable for the development of quantum applications to be able to fully utilize the computational resources on a classical computer and run at high speed. In this paper, we analyze the performance of the Qiskit Aer simulator, the Qulacs simulator, and mpiQulacs on various types of single servers. We then compare the performance of these quantum computer simulators across different computer environments. Nozomi Aoki, Masafumi Yamazaki, Akira Hirai, Mari Yamaoka, Naoto Fukumoto, Akihiko Kasagi, Masato Oguchi |
SERA | 6 |
| 2023 | A novel structured sparse fully connected layer in convolutional neural networksabstractAbstract Convolutional Neural Networks (CNNs) are one of the factors supporting the rapid development of artificial intelligent techniques. However, as the ability of the network increases, the size of the network becomes larger. Thus far, several works related to reduction of the network size have been tackled. In many cases, these approaches produce an unstructured network which prevents efficient parallel computation. To avoid this problem, we propose a novel structured sparse fully connected layer (FCL) in the CNNs. The aim of our proposed approach is reduction of the number of network parameters in the FCLs which occupy a large part of network parameters. Unlike the general FCLs used in the popular CNNs such as VGG‐16, the proposed approach reduces the connection between the last convolutional layer and the first FCL. In addition, we show an implementation for the proposed sparse FCLs on the GPU using cuBLAS. As a result for ILSVRC‐2012 dataset, the proposed approach achieves a 21.3 times compression with 0.68% top‐1 accuracy and 0.31% top‐5 accuracy decreases for VGG‐16. The implementation of the proposed FCLs achieves speed‐up factor 14.97 and 16.67 for forward and backward propagation compared to that for the noncompressed FCLs, respectively. Naoki Matsumura, Yasuaki Ito, Koji Nakano, Akihiko Kasagi, Tsuguchika Tabaru |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | NEEBS: Nonexpert large-scale environment building system for deep neural networkabstractSummary Deep neural networks (DNNs) have greatly improved the accuracy of various tasks in areas such as natural language processing (NLP). Obtaining a highly accurate DNN model requires multiple repetitions of training on a huge dataset, which requires a large‐scale cluster the compute nodes of which are tightly connected by high‐speed interconnects to exchange a large amount of intermediate data with very short latency. However, fully using the computational power of a large‐scale cluster for training requires knowledge of its components such as a distributed file system, an interconnection, and optimized high‐performance libraries. We have developed a Non‐Expert large‐scale Environment Building System (NEEBS) that aids a user in building a fast‐running training environment on a large‐scale cluster. It automatically installs and configures the applications and necessary libraries. It also optimally prepares tools to stage both data and executable programs, and launcher scripts suitable for both the applications and job submission systems of the cluster. NEEBS achieves 93.91% throughput scalability in NLP pretraining. We also present an approach to reduce pretraining time of highly accurate DNN model for NLP using a large‐scale computation environment built using NEEBS. We trained a Bidirectional Encoder Representations from Transformers (BERT)‐3.9b and a BERT‐xlarge using a dense masked language model (MLM) on Megatron‐LM framework and evaluated the improvement in learning time and learning efficiency for a Japanese language dataset using 768 graphics processing units (GPUs) on the AI Bridging Cloud Infrastructure (ABCI). Our implementation NEEBS improved learning efficiency per iteration by a factor of 10 and completed the pretraining of BERT‐xlarge in 4.7 h. This pretraining takes 5 months on a single GPU. To determine if the BERT models are correctly pretrained, we evaluated their accuracy in two tasks, Stanford Natural Language Inference Corpus translated into Japanese (JSNLI) and Twitter reputation analysis (TwitterRA). BERT‐3.9b achieved 94.30% accuracy for JSNLI, and BERT‐xlarge achieved 90.63% accuracy for TwitterRA. We constructed pretrained models with comparable accuracy to other Japanese BERT models in a shorter time. Yoshiharu Tajima, Masahiro Asaoka, Akihiro Tabuchi, Akihiko Kasagi, Tsuguchika Tabaru |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | GPU implementations of deflate encoding and decodingabstractSummary Deflate coding is a very popular lossless data compression method used in zlib, gzip (GNU zip), and zip, which performs the LZSS compression algorithm with Huffman coding. Deflate encoding and decoding involve sequential operations and their parallel acceleration using a GPU is quite hard. The main purpose of this paper is to present GPU implementations for encoding and decoding of Deflate coding. For efficient GPU implementations of Deflate coding, we have used multiple small hash tables for finding matching subsequences in the dictionary by multiple threads in parallel and applied the Single Kernel Soft Synchronization (SKSS) technique to fully utilize GPU computing resources. We have also adopted Huffman coding with gap arrays to accelerate parallel Huffman decoding. We have evaluated the performance of our GPU implementations using an NVIDIA A100 GPU and compared them with parallel/sequential Deflate encoding and decoding on the Intel X86 multicore CPUs using multiple threads/a single thread. Our GPU implementation of Deflate decoding is 1.66x–8.33x faster than the multiple thread implementation and 4.13x–36.56x faster than the single thread implementation. Daisuke Takafuji, Koji Nakano, Yasuaki Ito, Akihiko Kasagi |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | BERT-Based Scientific Paper Quality Prediction
Taiki Sasaki, Yasuaki Ito, Koji Nakano, Akihiko Kasagi |
ICANN (4) | 4 |
| 2021 | The 16, 384-node Parallelism of 3D-CNN Training on An Arm CPU based SupercomputerabstractAs the computational cost and datasets available for deep neural network training continue to increase, there is a significant demand for fast distributed training on supercomputers. However, porting and tuning applications for new advanced supercomputers requires tremendous amount of development efforts. Therefore, we present software tuning best practice for a 3D-CNN model training on a new Arm CPU based supercomputer, Fugaku. We (i) tune computation in DL by a JIT translator for aarch64, (ii) optimize collective communication such as Allreduce for 6D mesh/torus network topology, (iii) tune I/O by data staging with compression and data loader with caching, and (iv) parallelize training in data and model parallelism. We apply the proposed methods to a CosmoFlow 3D-CNN model, and achieve the training in 30 minutes using 16,384 nodes consisting of 4096 data- and 4 model-parallelism. This is the fastest result of any CPU-based systems in MLPerf HPC v0.7 in the world. Akihiro Tabuchi, Koichi Shirahata, Masafumi Yamazaki, Akihiko Kasagi, Takumi Honda, Kouji Kurihara, Kentaro Kawakami, Tsuguchika Tabaru, Naoto Fukumoto, Akiyoshi Kuroda, Takaaki Fukai, Kento Sato |
HiPC | 4 |
| 2020 | An Efficient Technique for Large Mini-batch Challenge of DNNs Training on Large Scale ClusterabstractDistributed deep learning using large mini-batches is a key strategy to perform the deep learning as fast as possible, but it represents a great challenge as it is difficult to achieve high scaling efficiency when using large clusters without compromising accuracy. The particular problem in this challenge is decreasing the number of model update iterations in whole of training. Thus, we need a technique which can converge the validation accuracy with a small number of iterations to address this challenge. In this paper, we introduce a novel technique, Final Polishing. This technique adjusts the means and variances in the batch normalization and mitigates the difference of normalization between validation datasets and augmented training datasets. By applying the technique, we achieved top-1 validation accuracy of 75.08% with mini-batch size of 81,920, with 2,048 GPUs and completed the training of ResNet-50 in 74.7 seconds.In addition, targeting top-1 validation accuracy of 75.9% or more, we tried additional parameters tuning. Then, we adjusted the number of GPUs and hyperparameters of DNNs with Final Polishing, and we also achieved top-1 validation accuracy of 75.97% with mini-batch size of 86,016, with 3,072 GPUs and completed the training of ResNet-50 in 62.1 seconds. Akihiko Kasagi, Akihiro Tabuchi, Masafumi Yamazaki, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, Kohta Nakashima |
HPDC | 1 |
| 2020 | Huffman Coding with Gap Arrays for GPU AccelerationabstractHuffman coding is a fundamental lossless data compression scheme used in many data compression file formats such as gzip, zip, png, and jpeg. Huffman encoding is easily parallelized, because all 8-bit symbols can be converted into codewords independently. On the other hand, since an encoded codeword sequence has no separator to identify each codeword, parallelizing Huffman decoding is a much harder task. This work presents a new data structure called gap array to be attached to an encoded codeword sequence of Huffman coding for accelerating parallel Huffman decoding. In addition, it also shows that GPU Huffman encoding and decoding can be accelerated by several techniques including (1) the Single Kernel Soft Synchronization (SKSS), (2) wordwise global memory access and (3) compact codebooks. The experimental results for 10 files on NVIDIA Tesla V100 GPU show that our GPU Huffman encoding and decoding run 2.87x-7.70x times and 1.26x-2.63x times faster than previously presented GPU Huffman encoding and decoding, respectively. Also, Huffman decoding can be further accelerated by a factor of 1.67x-6450x if a gap array is attached to an encoded codeword sequence. Since the size and computing overhead of gap arrays in Huffman encoding are small, we can conclude that gap arrays should be introduced for GPU Huffman encoding and decoding. Naoya Yamamoto, Koji Nakano, Yasuaki Ito, Daisuke Takafuji, Akihiko Kasagi, Tsuguchika Tabaru |
ICPP | 5 |
| 2020 | Efficient convolution pooling on the GPU
Shunsuke Suita, Takahiro Nishimura, Hiroki Tokura, Koji Nakano, Yasuaki Ito, Akihiko Kasagi, Tsuguchika Tabaru |
J. Parallel Distributed Comput. | 6 |
| 2014 | Parallel Algorithms for the Summed Area Table on the Asynchronous Hierarchical Memory Machine, with GPU implementationsabstractThe Hierarchical Memory Machine (HMM) is a theoretical parallel computing model that captures the essence of computing on CUDA-enabled GPUs. The summed area table (SAT) of a matrix is a data structure frequently used in the area of computer vision which can be obtained by computing the column-wise prefix-sums and then the row-wise prefix-sums. The main contribution of this paper is to introduce the asynchronous Hierarchical Memory Machine (asynchronous HMM), which supports asynchronous execution of CUDA blocks, and show a global-memory-access-optimal parallel algorithm for computing the SAT on the asynchronous HMM. A straightforward algorithm (2R2W SAT algorithm) on the asynchronous HMM, which computes the prefix-sums in every column using one thread each and then computes the prefix-sums in every row, performs 2 read operations and 2 write operations per element of a matrix. The previously published best algorithm (2R1W SAT algorithm) performs 2 read operations and 1 write operation per element. We present a more efficient algorithm (1R1W SAT algorithm) which performs 1 read operation and 1 write operation per element. Clearly, since every element in a matrix must be read at least once, and all resulting values must be written, our 1R1W SAT algorithm is optimal in terms of the global memory access. We also show a combined algorithm ((1 + r)R1W SAT algorithm) of 2R1W and 1R1W SAT algorithms that may have better performance. We have implemented several algorithms including 2R2W, 2R1W, 1R1W, (1 + r)R1W SAT algorithms on GeForce GTX 780 Ti. The experimental results show that our (1 + r)R1W SAT algorithm runs faster than any other SAT algorithms for large input matrices. Also, it runs more than 100 times faster than the best SAT algorithm using a single CPU. Akihiko Kasagi, Koji Nakano, Yasuaki Ito |
ICPP | 1 |
| 2013 | An Optimal Offline Permutation Algorithm on the Hierarchical Memory Machine, with the GPU ImplementationabstractThe Hierarchical Memory Machine (HMM) is a theoretical parallel computing model that captures the essence of computation on CUDA-enabled GPUs. The offline permutation is a task to copy numbers stored in an array a of size n to an array b of the same size along a permutation P given in advance. A conventional algorithm can complete the offline permutation by executing b[p[i]] ← a[i] for all i in parallel, where an array p stores the permutation P. This conventional algorithm simply performs three rounds of memory access for reading from a, reading from p, and writing in b. The main contribution of this paper is to present an optimal offline permutation algorithm running in O(n/w + L) time units using n threads on the HMM with width w and latency L. We also implement our optimal offline permutation algorithm on GeForce GTX-680 GPU and evaluate the performance. Quite surprisingly, our optimal offline permutation algorithm achieves better performance than the conventional algorithm in most permutations, although it performs 32 rounds of memory access. For example, the bit-reversal permutation for 4M float (32-bit) numbers can be completed in 780ms by our optimal permutation algorithm, while the conventional algorithm takes 2328ms. We can say that the experimental results of this paper provide a good example of GPU computation showing that a complicated but ingenious implementation with a larger constant factor in computing time can outperform a much simpler conventional algorithm. Akihiko Kasagi, Koji Nakano, Yasuaki Ito |
ICPP | 1 |