Takumi Honda

dblp:131/9150 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 A 590-Nanosecond 757-Gbps FPGA Lossy Compressed Network
abstract
Inter-FPGA communication bandwidth has become a limiting factor in scaling memory-intensive workloads on FPGA-based systems. While modern FPGAs integrate high- bandwidth memory (HBM) to increase local memory throughput, network interfaces often lag behind, creating an imbalance between computation and communication resources. Data compression is a technique to increase effective communication bandwidth by reducing the amount of data transferred, but existing solutions struggle to meet the performance and operation latency requirements of FPGA-based platforms. This paper presents a high- throughput lossy compression framework that enables sub-microsecond latency communication in FPGA clusters. The proposed design addresses the challenge of aligning variable-length compressed data with fixed-width network channels by using transpose circuits, memory-bank reordering, and word- wise operations. A run-length encoding scheme with bounded error is employed to compress floating-point and fixed-point data without relying on complex fine-grained bit-level manipulations, enabling low-latency and scalable implementation. The proposed architecture is implemented on a custom Stratix 10 MX2100 FPGA card equipped with eight 50 Gbps network ports and silicon photonics transceivers. The system achieves up to 757 Gbps of aggregate bandwidth per FPGA in collective communication operations. Compression and decompression are performed within 590 ns total latency, while maintaining the quality of results in a GradAllReduce workload for deep learning.
Michihiro Koibuchi, Takumi Honda, Naoto Fukumoto, Shoichi Hirasawa, Koji Nakano
IEEE Trans. Parallel Distributed Syst.2
2025 Lossy Compressed Collective Inter-FPGA Communications
abstract
A cutting-edge FPGA can be equipped with many memory channels of HBMs.The bottleneck for inter-FPGA memory communication would become the aggregate network bandwidth.To fill the gap between memory and network bandwidth, this study presents an approximate inter-FPGA memory network that works at the expense of output quality.Lossy data compression is a typical way of providing approximate communication.Prior works illustrated that simple lossy compression algorithms with a higher than 1.8 compression ratio were acceptable for communications generated in typical parallel applications, such as two-dimensional Lattice Boltzmann Method (2D-LBM), k-means data clustering, fast Fourier transform (FFT), and conjugate gradient method (CG) applications.Although the lossy compression for approximate communication has been widely studied, its feasibility and evaluation studies using high-bandwidth interconnection networks on real systems were rarely done.We design and evaluate lossy compressed collective communications so that all the memory and network bandwidth are fully used in a custom Stratix10 MX2100 FPGA card.Our evaluation results show that the custom card transfers up to 549.9-Gbps of collective (scatter, allgather, alltoall) data on an FPGA cluster, while the original design without the compression transfers up to 360 Gbps.The FPGA sender and receiver overhead, including (de)compression, are 312.8nsand 365.7ns.Our cycle-accurate network simulation shows that a high compression ratio significantly improves the effective network throughput of typical synthetic traffic patterns, especially for long messages.
Michihiro Koibuchi, Yoshinobu Ishida, Shoichi Hirasawa, Yao Hu 0005, Takumi Honda, Yusuke Nagasaka, Naoto Fukumoto
HPC Asia5
2023 Accelerating Hybrid DFT Simulations Using Performance Modeling on Supercomputers
abstract
Density Functional Theory (DFT) is an electronic-structure theory that computes the electronic energy of atoms and molecules from their electron density. Among several DFT methods, one called “hybrid DFT” adds the Hartree-Fock exchange energy to the original DFT exchange energy, and it improves the accuracy of the estimation of energy. However, this introduces additional computational costs, preventing its wide application for large-scale calculations. In light of those issues, a performance model to tune the computational configurations for hybrid DFT software automatically is proposed. The proposed model makes it possible to exhaustively search for parameters to minimize computation time without having to execute actual calculations with all parameter combinations. Several techniques for optimizing hybrid DFT, specially designed for the Fugaku supercomputer, are also proposed. It is concluded that combining all approaches reduces node-time cost by 2.23x and 2.68x for a 52-atom input on Fugaku and ABCI, respectively.
Yosuke Oyama, Takumi Honda, Atsushi Ishikawa, Koichi Shirahata
CCGrid2
2023 Big Data Assimilation: Real-time 30-second-refresh Heavy Rain Forecast Using Fugaku During Tokyo Olympics and Paralympics
abstract
Real-time 30-second-refresh numerical weather prediction (NWP) was performed with exclusive use of 11,580 nodes (~7%) of supercomputer Fugaku during Tokyo Olympics and Paralympics in 2021. Total 75,248 forecasts were disseminated in the 1-month period mostly stably with time-to-solution less than 3 minutes for 30-minute forecast. Japan's Big Data Assimilation (BDA) project developed the novel NWP system for precise prediction of hazardous rains toward solving the global climate crisis. Compared with typical 1-hour-refresh systems, the BDA system offered two orders of magnitude increase in problem size and revealed the effectiveness of 30-second refresh for highly nonlinear, rapidly evolving convective rains. To achieve the required time-to-solution for real-time 30-second refresh with high accuracy, the core BDA software incorporated single precision and enhanced parallel I/O with properly selected configurations of 1000 ensemble members and 500-m-mesh weather model. The massively parallel, I/O intensive real-time BDA computation demonstrated a promising future direction.
Takemasa Miyoshi, Arata Amemiya, Shigenori Otsuka, Yasumitsu Maejima, Takumi Honda, Hirofumi Tomita, Seiya Nishizawa, Kenta Sueki, Tsuyoshi Yamaura, Yutaka Ishikawa, Shinsuke Satoh, Tomoo Ushio, Kana Koike, Atsuya Uno
SC6
2021 The 16, 384-node Parallelism of 3D-CNN Training on An Arm CPU based Supercomputer
abstract
As the computational cost and datasets available for deep neural network training continue to increase, there is a significant demand for fast distributed training on supercomputers. However, porting and tuning applications for new advanced supercomputers requires tremendous amount of development efforts. Therefore, we present software tuning best practice for a 3D-CNN model training on a new Arm CPU based supercomputer, Fugaku. We (i) tune computation in DL by a JIT translator for aarch64, (ii) optimize collective communication such as Allreduce for 6D mesh/torus network topology, (iii) tune I/O by data staging with compression and data loader with caching, and (iv) parallelize training in data and model parallelism. We apply the proposed methods to a CosmoFlow 3D-CNN model, and achieve the training in 30 minutes using 16,384 nodes consisting of 4096 data- and 4 model-parallelism. This is the fastest result of any CPU-based systems in MLPerf HPC v0.7 in the world.
Akihiro Tabuchi, Koichi Shirahata, Masafumi Yamazaki, Akihiko Kasagi, Takumi Honda, Kouji Kurihara, Kentaro Kawakami, Tsuguchika Tabaru, Naoto Fukumoto, Akiyoshi Kuroda, Takaaki Fukai, Kento Sato
HiPC5
2020 An Efficient Technique for Large Mini-batch Challenge of DNNs Training on Large Scale Cluster
abstract
Distributed deep learning using large mini-batches is a key strategy to perform the deep learning as fast as possible, but it represents a great challenge as it is difficult to achieve high scaling efficiency when using large clusters without compromising accuracy. The particular problem in this challenge is decreasing the number of model update iterations in whole of training. Thus, we need a technique which can converge the validation accuracy with a small number of iterations to address this challenge. In this paper, we introduce a novel technique, Final Polishing. This technique adjusts the means and variances in the batch normalization and mitigates the difference of normalization between validation datasets and augmented training datasets. By applying the technique, we achieved top-1 validation accuracy of 75.08% with mini-batch size of 81,920, with 2,048 GPUs and completed the training of ResNet-50 in 74.7 seconds.In addition, targeting top-1 validation accuracy of 75.9% or more, we tried additional parameters tuning. Then, we adjusted the number of GPUs and hyperparameters of DNNs with Final Polishing, and we also achieved top-1 validation accuracy of 75.97% with mini-batch size of 86,016, with 3,072 GPUs and completed the training of ResNet-50 in 62.1 seconds.
Akihiko Kasagi, Akihiro Tabuchi, Masafumi Yamazaki, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, Kohta Nakashima
HPDC4
2017 Simple and Fast Parallel Algorithms for the Voronoi Map and the Euclidean Distance Map, with GPU Implementations
abstract
The complete Voronoi map of a binary image with black and white pixels is a matrix of the same size such that each element is the closest black pixel of the corresponding pixel. The complete Voronoi map visualizes the influence region of each black pixel. However, each region may not be connected due to exclave pixels. The connected Voronoi map is a modification of the complete Voronoi map so that all regions are connected. The Euclidean distance map of a binary image is a matrix, in which each element is the distance to the closest black pixel. It has many applications of image processing such as dilation, erosion, blurring effects, skeletonization and matching. The main contribution of this paper is to present simple and fast parallel algorithms for computing the complete/connected Voronoi maps and the Euclidean distance map and implement them in the GPU. Our parallel algorithm first computes the mixed Voronoi map, which is a mixture of the complete and connected Voronoi maps, and then converts it into the complete/connected Voronoi by exposing/hiding all exclave pixels. After that, the complete Voronoi map is converted into the Euclidean distance map by computing the distance to the closest black pixel for every pixel in an obvious way. The experimental results on GeForce GTX~1080 GPU show that the computing time for these conversions is relatively small. The throughput of our GPU implementation for computing the Euclidean distance maps of 2K × 2K binary images is up to 2.08 times larger than the previously published best GPU implementation, and up to 172 times larger than CPU implementation using Intel Core i7-4790.
Takumi Honda, Shinnosuke Yamamoto, Hiroaki Honda, Koji Nakano, Yasuaki Ito
ICPP1
2017 Accelerating digital halftoning using the local exhaustive search on the GPU
abstract
Summary Digital halftoning is an important process to convert a grayscale image into a binary image with black and white pixels. Local exhaustive search‐based halftoning is one of the halftoning methods that can generate high‐quality binary images. However, considering the computing time, it is not realistic for most applications. As a first contribution, this paper proposes a graphics processing unit (GPU) implementation for digital halftoning employing local exhaustive search to produce high‐quality binary images. Programming issues of the GPU architecture have been carefully assessed for implementing the proposed method. Experimental results show that the proposed GPU implementation on NVIDIA (Santa Clara, CA, USA) GeForce GTX TITAN X attains a speed‐up factor of up to 48 over a CPU implementation. Our second contribution is a GPU implementation for cluster‐dot halftoning tailored for local exhaustive search. This implementation attains a speed‐up factor of 92 over a sequential CPU implementation. Copyright © 2016 John Wiley & Sons, Ltd.
Hiroaki Koge, Takumi Honda, Toru Fujita, Yasuaki Ito, Koji Nakano, Jacir Luiz Bordim
Concurr. Comput. Pract. Exp.2
2014 GPU-Accelerated Verification of the Collatz Conjecture
Takumi Honda, Yasuaki Ito, Koji Nakano
ICA3PP (1)1