EDBT 2026 Demo / reviewers in the wild / expert
Yasuhiro Idomura
dblp:01/7390
· DBLP profile ↗
10ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0002-2829-0498ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 70% GPUs and heterogeneous computing · 15% Hardware accelerators and domain-specific architectures · 12% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › large-scale simulation
fusion plasma simulation |
0.4 | 1 | 2020 | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020 |
High-performance computing
gyrokinetic simulation |
0.4 | 1 | 2020 | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020 |
High-performance computing
scientific computing |
0.4 | 1 | 2020 | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020 |
Hardware accelerators and domain-specific architectures
accelerator optimization |
0.3 | 1 | 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
GPUs and heterogeneous computing › GPU memory management
GPU memory access optimization |
0.3 | 1 | 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
High-performance computing
stencil computation |
0.3 | 1 | 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
High-performance computing › iterative methods
krylov subspace method |
0.1 | 1 | 2020 | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov method · SC 2020 |
Storage systems
data layout |
0.1 | 1 | 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access Patterns · IEEE Trans. Parallel Distributed Syst. 2017 |
Methods — techniques the papers use, named apart from their topics
mixed precision · 0.4FP16 preconditioner · 0.4temporal blocking · 0.3register reuse · 0.3data layout transformation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A new data conversion method for mixed precision Krylov solvers with FP16/BF16 Jacobi preconditionersabstractMixed precision Krylov solvers with the Jacobi preconditioner often show significant convergence degradation when the Jacobi preconditioner is computed in low precision such as FP16 and BF16. It is found that this convergence degradation is attributed to loss of diagonal dominance due to roundoff errors in data conversion. To resolve this issue, we propose a new data conversion method, which is designed to keep diagonal dominance of the original matrix data. The proposed method is tested by computing the Poisson equation using the conjugate gradient method, the general minimum residual method, and the biconjugate gradient stabilized method with the FP16/BF16 Jacobi preconditioner on NVIDIA V100 GPUs. Here, the new data conversion is implemented by switching the round-nearest, round-up, round-down, and round-towards-zero intrinsics in CUDA, and is called once before the main iteration. Therefore, the cost of the new data conversion is negligible. When the coefficients of matrix is continuously changed by scaling the linear system, the conventional data conversion based on the round-nearest intrinsic shows periodic changes of the convergence property depending on the difference of the roundoff errors between diagonal and off-diagonal coefficients. Here, the period and magnitude of the convergence degradation depend on the bit length of significand. On the other hand, the proposed data conversion method is shown to fully avoid the convergence degradation, and robust mixed precision computing is enabled for the Jacobi preconditioner without extra overheads. Takuya Ina, Yasuhiro Idomura, Toshiyuki Imamura, Naoyuki Onodera |
HPC Asia | 2 |
| 2021 | AMR-Net: Convolutional Neural Networks for Multi-resolution Steady Flow PredictionabstractWe develop a convolutional neural network model to predict multi-resolution steady flow data. Based on the image-to-image translation model pix2pixHD, our model can predict high resolution flow fields from the set of patched signed distance functions. By patching the high resolution data, our model uses roughly the one third of memory used by pix2pixHD. The accuracy of our model is almost the same as the U-Net model using the unpatched high resolution data. Yuuichi Asahi, Sora Hatayama, Takashi Shimokawabe, Naoyuki Onodera, Yuta Hasegawa, Yasuhiro Idomura |
CLUSTER | 6 |
| 2021 | GPU Acceleration of Multigrid Preconditioned Conjugate Gradient Solver on Block-Structured Cartesian GridabstractWe develop a multigrid preconditioned conjugate gradient (MG-CG) solver for the pressure Poisson equation in a two-phase flow CFD code JUPITER. The JUPITER code is redesigned to realize efficient CFD simulations including complex boundaries and objects based on a block-structured Cartesian grid system. The code is written in CUDA, and is tuned to achieve high performance on GPU based supercomputers. The main kernels of the MG-CG solver achieve more than 90% of the roofline performance. The MG preconditioner is constructed based on the geometric MG method with a three-stage V-cycle, and a red-black SOR (RB-SOR) smoother and its variant with cache-reuse optimization (CR-SOR) are applied at each stage. The numerical experiments are conducted for two-phase flows in a fuel bundle of a nuclear reactor. Thanks to the block-structured data format, grids inside fuel pins are removed without performance degradation, and the total number of grids is reduced to 2.26 × 109, which is about 70% of the original Cartesian grid. The MG-CG solvers with the RB-SOR and CR-SOR smoothers reduce the number of iterations to less than 15% and 9% of the original preconditioned CG method, leading to 3.1- and 5.9-times speedups, respectively. In the strong scaling test, the MG-CG solver with the CR-SOR smoother is accelerated by 2.1 times between 64 and 256 GPUs. The obtained performance indicates that the MG-CG solver designed for the block-structured grid is highly efficient and enables large-scale simulations of two-phase flows on GPU based supercomputers. Naoyuki Onodera, Yasuhiro Idomura, Yuta Hasegawa, Susumu Yamashita, Takashi Shimokawabe, Takayuki Aoki |
HPC Asia | 2 |
| 2021 | Tree cutting approach for domain partitioning on forest-of-octrees-based block-structured static adaptive mesh refinement with lattice Boltzmann methodabstractThe aerodynamics simulation code based on the lattice Boltzmann method (LBM) using forest-of-octrees-based block-structured adaptive mesh refinement (AMR) with temporary-fixed refinement was implemented, and its performance was evaluated on GPU-based supercomputers. Although the Space-Filling-Curve-based (SFC) domain partitioning algorithm for the octree-based AMR has been widely used on conventional CPU-based supercomputers, accelerated computation on GPU-based supercomputers revealed a bottleneck due to costly halo data communication. Our new tree cutting approach adopts a hybrid domain partitioning with the coarse structured block decomposition and the SFC partitioning in each block. This hybrid approach improved the locality and the topology of the partitioned sub-domains and reduced the amount of the halo communication to one-third of the original SFC approach. In the strong scaling test, the code achieved maximum ×1.82 speedup at the performance of 2207 MLUPS (mega-lattice update per second) on 128 GPUs (NVIDIA® Tesla® V100). In the weak scaling test, the code achieved 9620 MLUPS at 128 GPUs with 4.473 billion grid points, while keeping the parallel efficiency of 93.4% from 8 to 128 GPUs. Yuta Hasegawa, Takayuki Aoki, Hiromichi Kobayashi, Yasuhiro Idomura, Naoyuki Onodera |
Parallel Comput. | 4 |
| 2020 | Acceleration of fusion plasma turbulence simulations using the mixed-precision communication-avoiding krylov methodabstractThe multi-scale full-f simulation of the next generation experimental fusion reactor ITER based on a five dimensional (5D) gyrokinetic model is one of the most computationally demanding problems in fusion science. In this work, a Gyrokinetic Toroidal 5D Eulerian code (GT5D) is accelerated by a new mixed-precision communication-avoiding (CA) Krylov method. The bottleneck of global collective communication on accelerated computing platforms is resolved using a CA Krylov method. In addition, a new FP16 preconditioner, which is designed using the new support for FP16 SIMD operations on A64FX, reduces both the number of iterations (halo data communication) and the computational cost. The performance of the proposed method for ITER size simulations with ~0.1 trillion grids on 1,440 CPUs/GPUs on Fugaku and Summit shows 2.8× and 1.9× speedups respectively from the conventional non-CA Krylov method, and excellent strong scaling is obtained up to 5,760 CPUs/GPUs. Yasuhiro Idomura, Takuya Ina, Yussuf Ali, Toshiyuki Imamura |
SC | 1 |
| 2020 | Overlapping communications in gyrokinetic codes on accelerator-based platformsabstractSummary Communication and computation overlapping techniques have been introduced in the five‐dimensional gyrokinetic codes GYSELA and GKV. In order to anticipate some of the exa‐scale requirements, these codes were ported to the modern accelerators, Xeon Phi KNL and Tesla P 100 GPU. On accelerators, a serial version of GYSELA on KNL and GKV on GPU are respectively 1.3× and 7.4× faster than those on a single Skylake processor (a single socket). For the scalability, we have measured GYSELA performance on Xeon Phi KNL from 16 to 512 KNLs (1024 to 32k cores) and GKV performance on Tesla P 100 GPU from 32 to 256 GPUs. In their parallel versions, transpose communication in semi‐Lagrangian solver in GYSELA or Convolution kernel in GKV turned out to be a main bottleneck. This indicates that in the exa‐scale, the network constraints would be critical. In order to mitigate the communication costs, the pipeline and task‐based overlapping techniques have been implemented in these codes. The GYSELA 2D advection solver has achieved a 33% to 92% speed up, and the GKV 2D convolution kernel has achieved a factor of 2 speed up with pipelining. The task‐based approach gives 11% to 82% performance gain in the derivative computation of the electrostatic potential in GYSELA. We have shown that the pipeline‐based approach is applicable with the presence of symmetry, while the task‐based approach can be applicable to more general situations. Yuuichi Asahi, Guillaume Latu, Julien Bigot, Shinya Maeyama, Virginie Grandgirard, Yasuhiro Idomura |
Concurr. Comput. Pract. Exp. | 6 |
| 2019 | Implementation and performance evaluation of a communication-avoiding GMRES method for stencil-based code on GPU cluster
Kazuya Matsumoto, Yasuhiro Idomura, Takuya Ina, Akie Mayumi, Susumu Yamada |
J. Supercomput. | 2 |
| 2017 | Optimization of Fusion Kernels on Accelerators with Indirect or Strided Memory Access PatternsabstractThis paper describes optimization for high-dimensional stencil computations on accelerators involving complex memory access patterns, which appear in five dimensional fusion plasma turbulence codes, GYSELA and GT5D. They include different types of memory access patterns, the indirect memory access in GYSELA with a Semi-Lagrangian scheme and the strided memory access in GT5D with a Finite-Difference scheme. We focus on the affinity of the memory access patterns to accelerators such as GPGPUs and Xeon Phi coprocessors. On both devices, the Array of Structure of Array (AoSoA) data layout is preferable for contiguous memory accesses. It is shown that the effective local cache usage by improving spatial and temporal data locality is critical on Xeon Phi. On GPGPU, the texture memory usage improves the performance of the indirect memory accesses in the Semi-Lagrangian scheme. The reuse of registers by taking account of the physical symmetry of the Finite-Difference scheme reduces the amount of memory accesses. Through these optimizations, we achieve acceleration of 3.9 (8.1) on Xeon Phi (GPGPU) for the Semi-Lagrangian scheme and of 1.4 (3.9) on Xeon Phi (GPGPU) for the Finite-Different scheme with respect to the fully optimized codes on Sandy Bridge. Yuuichi Asahi, Guillaume Latu, Takuya Ina, Yasuhiro Idomura, Virginie Grandgirard, Xavier Garbet |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Improved strong scaling of a spectral/finite difference gyrokinetic code for multi-scale plasma turbulence
Shinya Maeyama, Tomohiko Watanabe, Yasuhiro Idomura, Motoki Nakata, Masanori Nunami, Akihiro Ishizawa |
Parallel Comput. | 3 |
| 2013 | Nuclear Fusion Simulation Code Optimization on GPU ClustersabstractGT5D is a nuclear fusion simulation program which aims to analyze the turbulence phenomena in tokamak plasma. In this research, we optimize it for GPU clusters with multiple GPUs on a node. Based on the profile result of GT5D on a CPU node, we decide to offload the whole of the time development part of the program to GPUs except MPI communication. We achieved 3.37 times faster performance in maximum in function level evaluation, and 2.03 times faster performance in total than the case of CPU-only execution, both in the measurement on high density GPU cluster HA-PACS where each computation node consists of four NVIDIA M2090 GPUs and two Intel Xeon E5-2670 (Sandy Bridge) to provide 16 cores in total. These performance improvements on single GPU corresponds to four CPU cores, not compared with a single CPU core. It includes 53% performance gain with overlapping the communication between MPI processes with GPU calculation. Norihisa Fujita, Hideo Nuga, Taisuke Boku, Yasuhiro Idomura |
ICPADS | 4 |