EDBT 2026 Demo / reviewers in the wild / expert
Naoki Matsumura
dblp:232/6587
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Active Learning for Graph Neural Networks Training in Catalyst Energy PredictionabstractThe electrification of major industrial processes constitutes an important step in reducing global carbon emissions. Thus, the identification of materials able to serve as catalysts for these processes is of great interest. To this end, materials informatics is increasingly being used to accelerate material discovery. The reactivity of a given material is determined by its electronic structure, i.e., by the availability and distribution of electrons, which nowadays is calculated using computationally demanding quantum mechanics-based density functional theory (DFT) simulations, wherein the total electronic energy is a key observable. Recently, the use of graph neural networks (GNNs) has been proposed as a viable path towards faster energy prediction. However, obtaining enough labeled data for GNN training is challenging due to the extensive DFT calculations required. In this work, we propose active learning for GNN training, aiming to improve the energy prediction accuracy of GNNs with less labeled training data. In the search for materials, promising materials are discovered from a vast array of material data, each with diverse structures and energies. Therefore, GNNs for material search should be trained on data that captures this diversity. To sample such data our method calculates the Euclidean distances among the expanded features, which include the structure and energy information of the material data that are sampled to generate the training data. Then, the data used for GNN training are sampled from the candidates in the order of the longest of the minimum distances calculated for each feature. We demonstrate the superiority of our proposed method using the energy prediction GNNs of PaiNN and EquiformerV2 on various catalyst datasets, including the published Open Catalyst 2022 dataset. When compared to existing active learning methods, GNNs trained with the proposed active learning method achieve a lower mean absolute error in predicted energy using less training data under all experimental conditions. Yasufumi Sakai, Naoki Matsumura, Atsuki Inoue, Hiroshi Kawaguchi 0001, Thang Dang, Atsushi Ishikawa, Árni Björn Höskuldsson, Egill Skúlason |
IJCNN | 2 |
| 2023 | A novel structured sparse fully connected layer in convolutional neural networksabstractAbstract Convolutional Neural Networks (CNNs) are one of the factors supporting the rapid development of artificial intelligent techniques. However, as the ability of the network increases, the size of the network becomes larger. Thus far, several works related to reduction of the network size have been tackled. In many cases, these approaches produce an unstructured network which prevents efficient parallel computation. To avoid this problem, we propose a novel structured sparse fully connected layer (FCL) in the CNNs. The aim of our proposed approach is reduction of the number of network parameters in the FCLs which occupy a large part of network parameters. Unlike the general FCLs used in the popular CNNs such as VGG‐16, the proposed approach reduces the connection between the last convolutional layer and the first FCL. In addition, we show an implementation for the proposed sparse FCLs on the GPU using cuBLAS. As a result for ILSVRC‐2012 dataset, the proposed approach achieves a 21.3 times compression with 0.68% top‐1 accuracy and 0.31% top‐5 accuracy decreases for VGG‐16. The implementation of the proposed FCLs achieves speed‐up factor 14.97 and 16.67 for forward and backward propagation compared to that for the noncompressed FCLs, respectively. Naoki Matsumura, Yasuaki Ito, Koji Nakano, Akihiko Kasagi, Tsuguchika Tabaru |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Tile art image generation using parallel greedy algorithm on the GPU and its approximation with machine learningabstractSummary Tile art image generation is one of the non‐photorealistic rendering methods. The generated digital image resembles artistic representation given digital photos and illustrations. The first contribution of this paper is to propose a tile image generation based on the greedy approach. The greedy approach is based on the characteristic of the human visual system to optimize generated images. In addition, to shorten the computation time, we show the parallel algorithm and its GPU acceleration technique. We have implemented it on NVIDIA Tesla V100 GPU. The experimental result shows that the GPU implementation attains a speed‐up factor of 318 and 16.19 over the sequential CPU implementation and the parallel multi‐core CPU implementation with 160 threads, respectively. The second contribution of this paper is to propose an approximation method using machine learning with deep neural networks. After learning the network with the tile art images generated by the greedy approach as training dataset, it can generate tile art images that well‐reproduce the original images with tile patterns. Moreover, we show an additional machine learning technique by repeating the forwarding computation for the generated tile art image as an input image. As a result, using this technique, we can generate a tile art images with clear shape of tiles. Naoki Matsumura, Hiroki Tokura, Yuki Kuroda, Yasuaki Ito, Koji Nakano |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Efficient implementations of Bloom filter using block RAMs and DSP slices on the FPGAabstractSummary This paper presents efficient FPGA implementations for the Bloom filter, in which a large set P of L‐byte patterns are registered beforehand. Our Bloom filter circuit performs the byte stream pattern test such that it receives an input byte stream t and outputs the bit stream in every clock cycle. Each bit of the output bit stream is 1 if an L‐byte sequence of t starting from the corresponding position is identical with one of the patterns in P. Our circuits use rolling hash functions to compute signatures of all patterns in P registered in Ultra RAMs of the Xilinx UltraScale+ FPGA VU9P. We present two types of implementations, DSP‐based implementation and RAM‐based implementation to compute rolling hash functions of L‐byte sequences using DSP slices and Block RAMs in the FPGA, respectively. The experimental results show that both DSP‐based and RAM‐based Bloom filter circuits for 4800K patterns of length 1024 can perform the byte stream pattern test for 1.1 Gps and 1.3 Gbps input byte streams, respectively, with false positive probability 10−12. Moreover, we can configure DSP‐based and RAM‐based Bloom filter circuits for 100K patterns to work for 54.9 Gbps and 62.2 Gbps input byte streams, respectively, with false positive probability 10−12. Takuma Wada, Naoki Matsumura, Ryota Yasudo, Koji Nakano, Yasuaki Ito |
Concurr. Comput. Pract. Exp. | 2 |