EDBT 2026 Demo / reviewers in the wild / expert
Naoto Fukumoto
dblp:154/2381
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0003-2103-881XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 590-Nanosecond 757-Gbps FPGA Lossy Compressed NetworkabstractInter-FPGA communication bandwidth has become a limiting factor in scaling memory-intensive workloads on FPGA-based systems. While modern FPGAs integrate high- bandwidth memory (HBM) to increase local memory throughput, network interfaces often lag behind, creating an imbalance between computation and communication resources. Data compression is a technique to increase effective communication bandwidth by reducing the amount of data transferred, but existing solutions struggle to meet the performance and operation latency requirements of FPGA-based platforms. This paper presents a high- throughput lossy compression framework that enables sub-microsecond latency communication in FPGA clusters. The proposed design addresses the challenge of aligning variable-length compressed data with fixed-width network channels by using transpose circuits, memory-bank reordering, and word- wise operations. A run-length encoding scheme with bounded error is employed to compress floating-point and fixed-point data without relying on complex fine-grained bit-level manipulations, enabling low-latency and scalable implementation. The proposed architecture is implemented on a custom Stratix 10 MX2100 FPGA card equipped with eight 50 Gbps network ports and silicon photonics transceivers. The system achieves up to 757 Gbps of aggregate bandwidth per FPGA in collective communication operations. Compression and decompression are performed within 590 ns total latency, while maintaining the quality of results in a GradAllReduce workload for deep learning. Michihiro Koibuchi, Takumi Honda, Naoto Fukumoto, Shoichi Hirasawa, Koji Nakano |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Lossy Compressed Collective Inter-FPGA CommunicationsabstractA cutting-edge FPGA can be equipped with many memory channels of HBMs.The bottleneck for inter-FPGA memory communication would become the aggregate network bandwidth.To fill the gap between memory and network bandwidth, this study presents an approximate inter-FPGA memory network that works at the expense of output quality.Lossy data compression is a typical way of providing approximate communication.Prior works illustrated that simple lossy compression algorithms with a higher than 1.8 compression ratio were acceptable for communications generated in typical parallel applications, such as two-dimensional Lattice Boltzmann Method (2D-LBM), k-means data clustering, fast Fourier transform (FFT), and conjugate gradient method (CG) applications.Although the lossy compression for approximate communication has been widely studied, its feasibility and evaluation studies using high-bandwidth interconnection networks on real systems were rarely done.We design and evaluate lossy compressed collective communications so that all the memory and network bandwidth are fully used in a custom Stratix10 MX2100 FPGA card.Our evaluation results show that the custom card transfers up to 549.9-Gbps of collective (scatter, allgather, alltoall) data on an FPGA cluster, while the original design without the compression transfers up to 360 Gbps.The FPGA sender and receiver overhead, including (de)compression, are 312.8nsand 365.7ns.Our cycle-accurate network simulation shows that a high compression ratio significantly improves the effective network throughput of typical synthetic traffic patterns, especially for long messages. Michihiro Koibuchi, Yoshinobu Ishida, Shoichi Hirasawa, Yao Hu 0005, Takumi Honda, Yusuke Nagasaka, Naoto Fukumoto |
HPC Asia | 7 |
| 2024 | Performance Analysis of Quantum Computer Simulators Across Different EnvironmentsabstractQuantum computers can achieve extremely fast computations for certain problems using quantum properties. Due to these factors, research and development in the field of quantum computers have been vigorously pursued. It is crucial to develop a quantum computer that operates accurately and, concurrently, create applications compatible with quantum computing. Quantum computer simulators are tools that represent the behavior of a quantum computer on classical computers, and they are highly conducive to the development of quantum applications that can run on actual quantum computers. On the other hand, quantum computer simulators run on classical computers, which are extremely slow compared to quantum computers. Therefore, it is desirable for the development of quantum applications to be able to fully utilize the computational resources on a classical computer and run at high speed. In this paper, we analyze the performance of the Qiskit Aer simulator, the Qulacs simulator, and mpiQulacs on various types of single servers. We then compare the performance of these quantum computer simulators across different computer environments. Nozomi Aoki, Masafumi Yamazaki, Akira Hirai, Mari Yamaoka, Naoto Fukumoto, Akihiko Kasagi, Masato Oguchi |
SERA | 5 |
| 2022 | Efficient Collision-Free MTTKRP Algorithm for Multi-core CPUs with Less Memory UsageabstractTensor decomposition is often used to extract underlying features in the analysis of large and multi-dimensional data. For the tensor data with sparse characteristics, Sparse Matricized Tensor Times Khatri-Rao Product (MTTKRP) is a performance bottleneck in the Alternating Least Squares (ALS) method, which is widely used for tensor decomposition. Since MTTKRP is computed for each mode of the tensor in the ALS method, it is required to compute MTTKRP efficiently not only for a specific mode but also for all modes. In addition, it is hard to convert input tensor data into the optimal format for each mode in terms of memory usage for large-scale data. We propose a fast MTTKRP calculation algorithm with a single replica for all modes based on the widely used COO format. It is a scalable, faster, and memory-saving method that does not require exclusive control, which is necessary for the conventional MTTKRP algorithm with the COO format. As a result of performance evaluation, we have achieved a significant speed-up from the existing method. For the MTTKRP calculation, we have achieved a maximum performance improvement of x4.8 and an average performance improvement of x1.3 on Intel Xeon, and a maximum performance improvement of x9.9 and an average performance improvement of x2.3 on Fujitsu A64FX. For the ALS method, our approach on the Fujitsu A64FX achieved a performance improvement of up to x1.8 over the existing method. Yusuke Nagasaka, Naoto Fukumoto |
CCGRID | 2 |
| 2021 | Towards Straggler-Tolerant and Accuracy-Aware Distributed DNN Training in CloudsabstractThis study investigated how straggler mitigation affects accuracy during distributed training. While distributed training is one promising way to shorten training time, we cannot obtain a maximum performance improvement if some workers become stragglers due to a slowdown. We can avoid the performance impact from stragglers by excluding them from training. However, we face another problem of decreasing accuracy during training because gradients computed by excluded stragglers are not reflected in the training. To achieve both straggler tolerance and accuracy awareness, we should implement distributed training so that excluded stragglers return to the training once they recover from slowdown. Although some implementations supporting auto-scaling have been created, they have not evaluated the training accuracy during scale-in to exclude stragglers from training and scale-out to return them to training. Therefore, we conducted experiments to see how excluding stragglers causes accuracy degradation. We also discussed how the exclusion time and training data assignment can have an impact on accuracy during training. Shingo Okuno, Masahiro Miwa, Naoto Fukumoto |
CCGRID | 3 |
| 2021 | The 16, 384-node Parallelism of 3D-CNN Training on An Arm CPU based SupercomputerabstractAs the computational cost and datasets available for deep neural network training continue to increase, there is a significant demand for fast distributed training on supercomputers. However, porting and tuning applications for new advanced supercomputers requires tremendous amount of development efforts. Therefore, we present software tuning best practice for a 3D-CNN model training on a new Arm CPU based supercomputer, Fugaku. We (i) tune computation in DL by a JIT translator for aarch64, (ii) optimize collective communication such as Allreduce for 6D mesh/torus network topology, (iii) tune I/O by data staging with compression and data loader with caching, and (iv) parallelize training in data and model parallelism. We apply the proposed methods to a CosmoFlow 3D-CNN model, and achieve the training in 30 minutes using 16,384 nodes consisting of 4096 data- and 4 model-parallelism. This is the fastest result of any CPU-based systems in MLPerf HPC v0.7 in the world. Akihiro Tabuchi, Koichi Shirahata, Masafumi Yamazaki, Akihiko Kasagi, Takumi Honda, Kouji Kurihara, Kentaro Kawakami, Tsuguchika Tabaru, Naoto Fukumoto, Akiyoshi Kuroda, Takaaki Fukai, Kento Sato |
HiPC | 9 |
| 2021 | Low-Latency Low-Energy Memory-Cube Networks using Dual-Voltage DatapathsabstractThree-dimensional stack memory that provides both high-bandwidth access and large capacity is a promising technology for next-generation computer systems. While a large number of memory cubes increase the aggregate memory capacity, the communication latency and power consumption would become significant due to its low-radix large-diameter packet network. In this context, we propose a memory-cube network called Diagonal Memory Network (DMN). A diagonal network topology, its floor layout, and its lightweight router are designed for low-latency and low-voltage memory-read communication. Our evaluation results show that a DMN router decreases 31% of the hardware resources than a conventional virtual-channel router. The DMN router reduces 13% and 67% energy consumption to transit a packet along with the original datapath and bypassing datapath, respectively. Yoshiya Shikama, Ryuta Kawano, Hiroki Matsutani, Hideharu Amano, Yusuke Nagasaka, Naoto Fukumoto, Michihiro Koibuchi |
PDP | 6 |
| 2020 | An Efficient Technique for Large Mini-batch Challenge of DNNs Training on Large Scale ClusterabstractDistributed deep learning using large mini-batches is a key strategy to perform the deep learning as fast as possible, but it represents a great challenge as it is difficult to achieve high scaling efficiency when using large clusters without compromising accuracy. The particular problem in this challenge is decreasing the number of model update iterations in whole of training. Thus, we need a technique which can converge the validation accuracy with a small number of iterations to address this challenge. In this paper, we introduce a novel technique, Final Polishing. This technique adjusts the means and variances in the batch normalization and mitigates the difference of normalization between validation datasets and augmented training datasets. By applying the technique, we achieved top-1 validation accuracy of 75.08% with mini-batch size of 81,920, with 2,048 GPUs and completed the training of ResNet-50 in 74.7 seconds.In addition, targeting top-1 validation accuracy of 75.9% or more, we tried additional parameters tuning. Then, we adjusted the number of GPUs and hyperparameters of DNNs with Final Polishing, and we also achieved top-1 validation accuracy of 75.97% with mini-batch size of 86,016, with 3,072 GPUs and completed the training of ResNet-50 in 62.1 seconds. Akihiko Kasagi, Akihiro Tabuchi, Masafumi Yamazaki, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, Kohta Nakashima |
HPDC | 6 |
| 2017 | Understanding storage traffic characteristics on enterprise virtual desktop infrastructureabstractDespite the growing popularity of enterprise virtual desktop infrastructure (VDI), little is known about its storage traffic characteristics. In addition, no prior work has considered the detailed characteristics of virtual machine (VM) behavior on VDI. In this paper, we analyze the enterprise storage traffic on commercial office VDI using designated VMs. For 28 consecutive days, we gathered various types of traces, including a usage questionnaire and active and passive measurements. To characterize the storage traffic, we focused on two perspectives: fibre channel (FC) traffic and VM behavior. From the FC traffic perspective, we found that read traffic is dominant, although the applications are similar to those in a previous small-scale VDI. In particular, the write response time of large transactions, e.g.,128 KiB, is strongly affected by a slight decrease in cache hits during an update storm. From the VM behavior, we found that all active user VMs generate only 25% of traffic. Although a few VMs generate massive traffic, their impact is small. These characteristics are unique in comparison with the small-scale VDI. Our results have significant implications for designing the next generation of VDI and improving its performance. Chunghan Lee, Tatsuo Kumano, Tatsuma Matsuki, Hiroshi Endo, Naoto Fukumoto, Mariko Sugawara |
SYSTOR | 5 |