Keigo Nitadori

dblp:84/7552 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0001-7374-4236ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 68% GPUs and heterogeneous computing · 21% Performance modeling and evaluation · 6%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
performance optimization at scale
0.322014
24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014
4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012
High-performance computing
scientific computing systems
0.322014
24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014
4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012
Programming languages and type systems
domain-specific languages
0.212016
Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016
High-performance computing
stencil computation
0.212016
Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster
0.222010
190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
GPUs and heterogeneous computing › GPU computing › GPGPU acceleration
GPU-accelerated supercomputing
0.212014
24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014
High-performance computing
n-body simulation
0.212014
24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014
High-performance computing › performance optimization at scale
parallel scalability
0.112012
4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012
Interconnection networks and networks-on-chip › cluster interconnect
infiniband
0.112010
190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010
High-performance computing › n-body simulation
treecode
0.112010
190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010
Performance modeling and evaluation › numerical algorithms
fast multipole method
0.112009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
High-performance computing › n-body simulation
hierarchical n-body methods
0.112009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
GPUs and heterogeneous computing
GPU computing
0.112014
24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014
Performance modeling and evaluation › design trade-off analysis
cost-performance analysis
0.012010
190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010
Computational science and engineering › computational fluid dynamics
turbulence simulation
0.012009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009

Methods — techniques the papers use, named apart from their topics

domain-specific language · 0.5auto-tuning · 0.5MPI · 0.5treecode · 0.3parallel efficiency analysis · 0.2gravitational tree-code · 0.2tree algorithm · 0.1particle-mesh algorithm · 0.1TreePM · 0.1hierarchical n-body · 0.1fast multipole method · 0.1
YearPublicationVenuePosition
2020 Prompt Report on Exa-Scale HPL-AI Benchmark
abstract
Our performance benchmark of HPL-AI on the supercomputer Fugaku was awarded in the 55th top500 at ISC20. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of exascale in a floating-point arithmetic benchmark. Due to the novelty of HPL-AI, there are few guidelines for large systems and several drawbacks to the large-scale benchmark. It is not enough to replace FP64 operations solely to those on FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and introduction of optimization techniques on extensive computing such as on Fugaku. In the poster, we give some comments on the accuracy, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku.
Shuhei Kudo, Keigo Nitadori, Takuya Ina, Toshiyuki Imamura
CLUSTER2
2016 Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes
abstract
Stencil computation has many applications in science and engineering, thus many optimization techniques such as temporal blocking have been developed. They are, however, rarely used in real-world applications, since a large amount of careful programming is required for even the simplest of stencils. We introduce Formura, a domain specific language that provides easy access to optimized stencil computations. Higher-order integration schemes can be defined using mathematical notations. Formura generates C code with MPI calls and performs autotuning. Hence its performance is portable to most distributed-memory computers. We show the scientific applicability of Formura by performing magnetohydrodynamics (MHD) and belowground biology simulations. Ability to reach bytes-per-flops ratio only attainable by temporal blocking is demonstrated. We also demonstrate scaling up to the full nodes of the K computer, with 1.184 Pflops, 11.62% floating-pointoperation efficiency, and 31.26% memory throughput efficiency.
Takayuki Muranushi, Hideyuki Hotta, Junichiro Makino, Seiya Nishizawa, Hirofumi Tomita, Keigo Nitadori, Masaki Iwasawa, Natsuki Hosono, Yutaka Maruyama, Hikaru Inoue, Hisashi Yashiro, Yoshifumi Nakamura
SC6
2014 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs
abstract
We have simulated, for the first time, the long term evolution of the Milky Way Galaxy using 51 billion particles on the Swiss Piz Daint supercomputer with our N-body gravitational tree-code Bonsai. Herein, we describe the scientific motivation and numerical algorithms. The Milky Way model was simulated for 6 billion years, during which the bar structure and spiral arms were fully formed. This improves upon previous simulations by using 1000 times more particles, and provides a wealth of new data that can be directly compared with observations. We also report the scalability on both the Swiss Piz Daint and the US ORNL Titan. On Piz Daint the parallel efficiency of Bonsai was above 95%. The highest performance was achieved with a 242 billion particle Milky Way model using 18600 GPUs on Titan, thereby reaching a sustained GPU and application performance of 33.49 Pflops and 24.77 Pflops respectively.
Jeroen Bédorf, Evghenii Gaburov, Michiko S. Fujii, Keigo Nitadori, Tomoaki Ishiyama, Simon Portegies Zwart
SC4
2012 4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem
abstract
As an entry for the 2012 Gordon-Bell performance prize, we report performance results of astrophysical N-body simulations of one trillion particles performed on the full system of K computer. This is the first gravitational trillion-body simulation in the world. We describe the scientific motivation, the numerical algorithm, the parallelization strategy, and the performance analysis. Unlike many previous Gordon-Bell prize winners that used the tree algorithm for astrophysical N-body simulations, we used the hybrid TreePM method, for similar level of accuracy in which the short-range force is calculated by the tree algorithm, and the long-range force is solved by the particle-mesh algorithm. We developed a highly-tuned gravity kernel for short-range forces, and a novel communication algorithm for long-range forces. The average performance on 24576 and 82944 nodes of K computer are 1.53 and 4.45 Pflops, which correspond to 49% and 42% of the peak speed.
Tomoaki Ishiyama, Keigo Nitadori, Junichiro Makino
SC2
2010 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs
abstract
We present the results of a hierarchical N-body simulation on DEGIMA, a cluster of PCs with 576 graphic processing units (GPUs) and using an InfiniBand interconnect. DEGIMA stands for DEstination for GPU Intensive MAchine, and is located at Nagasaki Advanced Computing Center (NACC), Nagasaki University. In this work, we have upgraded DEGIMA_s interconnect using InfiniBand. DEGIMA is composed by 144 nodes with 576 GT200 GPUs. An astrophysical N-body simulation with 3,278,982,596 particles using a treecode algorithm shows a sustained performance of 190.5 Tflops on DEGIMA. The overall cost of the hardware was $411,921 dollars. The maximum corrected performance is 104.8 Tflops for the simulation, resulting in a cost performance of 254.4 MFlops/$. This corrections is performed by counting the FLOPS based on the most efficient CPU algorithm. Any extra FLOPS that arise from the GPU implementation and parameter differences are not included in the 254.4 MFLOPS/$.
Tsuyoshi Hamada, Keigo Nitadori
SC2
2009 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence
abstract
As an entry for the 2009 Gordon Bell price/performance prize, we present the results of two different hierarchical N-body simulations on a cluster of 256 graphics processing units (GPUs). Unlike many previous N-body simulations on GPUs that scale as O(N2), the present method calculates the O(N log N) treecode and O(N) fast multipole method (FMM) on the GPUs with unprecedented efficiency. We demonstrate the performance of our method by choosing one standard application --a gravitational N-body simulation-- and one non-standard application --simulation of turbulence using vortex particles. The gravitational simulation using the treecode with 1,608,044,129 particles showed a sustained performance of 42.15 TFlops. The vortex particle simulation of homogeneous isotropic turbulence using the periodic FMM with 16,777,216 particles showed a sustained performance of 20.2 TFlops. The overall cost of the hardware was 228,912 dollars. The maximum corrected performance is 28.1TFlops for the gravitational simulation, which results in a cost performance of 124 MFlops/$. This correction is performed by counting the Flops based on the most efficient CPU algorithm. Any extra Flops that arise from the GPU implementation and parameter differences are not included in the 124 MFlops/$.
Tsuyoshi Hamada, Tetsu Narumi, Rio Yokota, Kenji Yasuoka, Keigo Nitadori, Makoto Taiji
SC5