Takaaki Miyajima

dblp:125/2854 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0001-9409-0913ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Real-time High-resolution X-Ray Computed Tomography
abstract
Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.
Du Wu, Peng Chen 0035, Xiao Wang 0004, Isaac Lyngaas, Takaaki Miyajima, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib
ICS5
2023 ESSPER: Elastic and Scalable FPGA-Cluster System for High-Performance Reconfigurable Computing with Supercomputer Fugaku
abstract
FPGA clusters have yet to be a mainstream of HPC, even for accelerators, and several challenges exist in their architecture and system organization. This work presents ESSPER, a flexible and scalable FPGA cluster prototype system for reconfigurable HPC to meet the concept of customizability, scalability, and interoperability with existing HPC systems. Based on our classification of FPGA cluster architectures, we propose a new category of FPGA clusters with a host-FPGA bridging network using software-bridged APIs for the use of remote FPGAs. We have designed, implemented, verified, and demonstrated a proof-of-concept system of ESSPER, as a functional extension of the supercomputer Fugaku.
Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno
HPC Asia3
2022 ESSPER: Elastic and Scalable System for High-Performance Reconfigurable Computing with Software-bridged APIs
abstract
Many-core CPUs and GPUs, present mainstream architectures for HPC, are facing difficulty in maintaining the same performance improvement rate because of the recent slow-down in the semiconductor scaling, the dark silicon problem, and wasteful mechanisms required for accelerating general-purpose computing such as a branch predictor and an out-of-order mechanism. Also, the power efficiency of HPC systems is significantly important to achieve higher performance.
Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno
FPT3
2021 A memory bandwidth improvement with memory space partitioning for single-precision floating-point FFT on Stratix 10 FPGA
abstract
The Fast Fourier Transform (FFT) is one of the fundamental computational methods used in the fields of computational science and high-performance computing. Single-precision floating-point complex FFT itself is known as a memory bandwidth bottleneck and often becomes a bottleneck of application acceleration in these fields. We are researching and developing a parallel FFT on FPGA(s) to overcome this problem. In this paper, we discuss the memory bandwidth of the single-precision floating-point complex FFT on an FPGA. Our FFT implementation is based on a state-of-the-art OpenCL implementation provided by Intel. We first show that the computational performance of the FFT on Intel PAC D5005 is proportional to the effective memory bandwidth of the main memory. Then we propose a memory sub-system to improve the effective memory bandwidth. Specifically, a memory space partitioning and the sub-modules that access each memory space individually. In our FPGA design running at 270 MHz, two memory channels of DDR4-2400 memory are used for both reading and writing, respectively. Our proposed memory sub-system achieved an effective memory bandwidth of 22.57 [GB/s] (65.3% of the theoretical peak of this implementation) was achieved when the number of data points for FFT was 16,777,216.
Takaaki Miyajima, Kentaro Sano
CLUSTER1
2020 Performance Evaluation and Power Analysis of Teraflop-scale Fluid Simulation with Stratix 10 FPGA
abstract
Stream computing is a suitable approach to improve both performance and power efficiency of numerical computations with FPGAs. To achieve further performance gain, temporal and spatial parallelism were exploited: the first one deepens and the latter duplicates pipelines of streamed computation cores. These two types of parallelism were previously evaluated with Arria 10 FPGA. However, it has not been verified if they are also effective for the latest FPGA, Stratix 10, which has a larger amount of logic elements (i.e., 2.4X of Arria 10) and is equipped with a new feature to improve the maximum clock frequency (i.e., HyperFlex architecture). To show the scalability for such state-of-the-art FPGAs, in this paper, we firstly implemented a streamed fluid simulation accelerator with both parallelism types for Stratix 10. We then thoroughly evaluated it by obtaining computational performance (FLOPS), power efficiency (FLOPS/W), resource utilization, and maximum clock frequency (Fmax). From the results, we found that this implementation excessively used DSP blocks due to inefficient mapping of floating-point operations, which reduced Fmax and the number of pipelined cores. To improve the scalability, we optimized the implementation to reduce the DSP block usage by utilizing a Multiply-Add function in a single DSP block. As a result, the optimized fluid simulation achieves 1.06 TFLOPS and 12.6 GFLOPS/W, which is 1.36X and 1.24X higher than the non-optimized version, respectively. Moreover, we estimate that the fluid simulation with Stratix 10 could outperform GPU-based implementation with Tesla V100 by optimizing it for HyperFlex architecture.
Atsushi Koshiba, Kouki Watanabe, Takaaki Miyajima, Kentaro Sano
FPGA3
2013 Task level pipelining with PEACH2: An FPGA switching fabric for high performance computing
abstract
We demonstrate task level pipelining on multiple accelerators with PEACH2. PEACH2 is implmented on FPGA, and enables ultra low latency direct communication among multiple accelerators over computational nodes. By installing PEACH2, typical high performance computation nodes are tightly coupled. In this environment, application can be accelerated by exploiting not only data level parallelism, but also task level pipelined operation. Furthermore, we can processe multiple task on multiple accelerators in a pipelined manner. In our demonstration, application achieves 44% speed up compared to a single GPU.
Takaaki Miyajima, Takuya Kuhara, Toshihiro Hanawa, Hideharu Amano, Taisuke Boku
FPT1
2013 Artificial intelligence of Blokus Duo on FPGA using Cyber Work Bench
abstract
This paper presents a design of an FPGA-based Blokus Duo solver. It searches a game tree by using the miniMax algorithm with alpha-beta pruning and move ordering. In addition, HLS tool called CyberWorkBench (CWB) is used to implement hardware. By making the use of functions in CWB, parallel fully pipelined design is generated. The implemented solver works at 100MHz with Xilinx Spartan-6 XC6SLX45 FPGA on the Digilent Atlys board. It can search states after three moves in most cases.
Naru Sugimoto, Takaaki Miyajima, Takuya Kuhara, Yuki Katuta, Takushi Mitsuichi, Hideharu Amano
FPT2
2013 A speculative gather system for Cool Mega-Array
abstract
Cool Mega Array (CMA) is a low power reconfigurable processor array for battery driven mobile devices. A prototype chip CMA-1 consists of a 8 × 8 PE (Processing Element) array and a micro-controller for controlling data alignment. Because the PE array of CMA is built with a combinatorial circuit, it does not have a signal which tells that operation in the PE array was completed. A propagate delay of the whole PE array corresponding to the operation time was estimated by using the data path and mapping information in the design stage of the application. The timing information for gathering the data was specified in the microcode of the controller. However, since this timing is fixed, it cannot treat the variation of environment temperature and voltage scaling for the PE array. Here, a speculative gather system is proposed which sets the timing of collecting operation results from the PE array dynamically. By collecting results twice and comparing them, it guarantees the correctness of the operation results and adjusts the gather timing automatically. The speculative gather system is implemented in the CMA, and evaluation results appear that the performance is improved by 25.3% on average with the overhead of 0.5% in area and 3.1% in power consumption.
Rie Uno, Nobuaki Ozaki, Mai Izawa, Akihito Tsusaka, Takaaki Miyajima, Hideharu Amano
FPT5