Takuya Ina

dblp:194/6638 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-3989-5011ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 70% GPUs and heterogeneous computing · 15% Hardware accelerators and domain-specific architectures · 12%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › large-scale simulation
fusion plasma simulation
0.412020
Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020
High-performance computing
gyrokinetic simulation
0.412020
Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020
High-performance computing
scientific computing
0.412020
Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020
Hardware accelerators and domain-specific architectures
accelerator optimization
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
GPUs and heterogeneous computing › GPU memory management
GPU memory access optimization
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
High-performance computing
stencil computation
0.312017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
High-performance computing › iterative methods
krylov subspace method
0.112020
Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020
Storage systems
data layout
0.112017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017
GPUs and heterogeneous computing
GPU computing
0.112017
Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017

Methods — techniques the papers use, named apart from their topics

mixed precision · 0.4FP16 preconditioner · 0.4temporal blocking · 0.3register reuse · 0.3data layout transformation · 0.3
YearPublicationVenuePosition
2023 A new data conversion method for mixed precision Krylov solvers with FP16/BF16 Jacobi preconditioners
abstract
Mixed precision Krylov solvers with the Jacobi preconditioner often show significant convergence degradation when the Jacobi preconditioner is computed in low precision such as FP16 and BF16. It is found that this convergence degradation is attributed to loss of diagonal dominance due to roundoff errors in data conversion. To resolve this issue, we propose a new data conversion method, which is designed to keep diagonal dominance of the original matrix data. The proposed method is tested by computing the Poisson equation using the conjugate gradient method, the general minimum residual method, and the biconjugate gradient stabilized method with the FP16/BF16 Jacobi preconditioner on NVIDIA V100 GPUs. Here, the new data conversion is implemented by switching the round-nearest, round-up, round-down, and round-towards-zero intrinsics in CUDA, and is called once before the main iteration. Therefore, the cost of the new data conversion is negligible. When the coefficients of matrix is continuously changed by scaling the linear system, the conventional data conversion based on the round-nearest intrinsic shows periodic changes of the convergence property depending on the difference of the roundoff errors between diagonal and off-diagonal coefficients. Here, the period and magnitude of the convergence degradation depend on the bit length of significand. On the other hand, the proposed data conversion method is shown to fully avoid the convergence degradation, and robust mixed precision computing is enabled for the Jacobi preconditioner without extra overheads.
Takuya Ina, Yasuhiro Idomura, Toshiyuki Imamura, Naoyuki Onodera
HPC Asia1
2020 Prompt Report on Exa-Scale HPL-AI Benchmark
abstract
Our performance benchmark of HPL-AI on the supercomputer Fugaku was awarded in the 55th top500 at ISC20. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of exascale in a floating-point arithmetic benchmark. Due to the novelty of HPL-AI, there are few guidelines for large systems and several drawbacks to the large-scale benchmark. It is not enough to replace FP64 operations solely to those on FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and introduction of optimization techniques on extensive computing such as on Fugaku. In the poster, we give some comments on the accuracy, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku.
Shuhei Kudo, Keigo Nitadori, Takuya Ina, Toshiyuki Imamura
CLUSTER3
2020 Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method
abstract
The multi-scale full-f simulation of the next generation experimental fusion reactor ITER based on a five dimensional (5D) gyrokinetic model is one of the most computationally demanding problems in fusion science. In this work, a Gyrokinetic Toroidal 5D Eulerian code (GT5D) is accelerated by a new mixed-precision communication-avoiding (CA) Krylov method. The bottleneck of global collective communication on accelerated computing platforms is resolved using a CA Krylov method. In addition, a new FP16 preconditioner, which is designed using the new support for FP16 SIMD operations on A64FX, reduces both the number of iterations (halo data communication) and the computational cost. The performance of the proposed method for ITER size simulations with ~0.1 trillion grids on 1,440 CPUs/GPUs on Fugaku and Summit shows 2.8× and 1.9× speedups respectively from the conventional non-CA Krylov method, and excellent strong scaling is obtained up to 5,760 CPUs/GPUs.
Yasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki Imamura
SC2
2019 Implementation and performance evaluation of a communication-avoiding GMRES method for stencil-based code on GPU cluster
Kazuya Matsumoto, Yasuhiro Idomura, Takuya Ina, Akie Mayumi, Susumu Yamada
J. Supercomput.3
2017 Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns
abstract
This paper describes optimization for high-dimensional stencil computations on accelerators involving complex memory access patterns, which appear in five dimensional fusion plasma turbulence codes, GYSELA and GT5D. They include different types of memory access patterns, the indirect memory access in GYSELA with a Semi-Lagrangian scheme and the strided memory access in GT5D with a Finite-Difference scheme. We focus on the affinity of the memory access patterns to accelerators such as GPGPUs and Xeon Phi coprocessors. On both devices, the Array of Structure of Array (AoSoA) data layout is preferable for contiguous memory accesses. It is shown that the effective local cache usage by improving spatial and temporal data locality is critical on Xeon Phi. On GPGPU, the texture memory usage improves the performance of the indirect memory accesses in the Semi-Lagrangian scheme. The reuse of registers by taking account of the physical symmetry of the Finite-Difference scheme reduces the amount of memory accesses. Through these optimizations, we achieve acceleration of 3.9 (8.1) on Xeon Phi (GPGPU) for the Semi-Lagrangian scheme and of 1.4 (3.9) on Xeon Phi (GPGPU) for the Finite-Different scheme with respect to the fully optimized codes on Sandy Bridge.
Yuuichi Asahi, Guillaume Latu, Takuya Ina, Yasuhiro Idomura, Virginie Grandgirard, Xavier Garbet
IEEE Trans. Parallel Distributed Syst.3