VLDB 2026 Research / reviewers in the wild / expert
Keigo Nitadori
dblp:84/7552
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0001-7374-4236ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 68% GPUs and heterogeneous computing · 21% Performance modeling and evaluation · 6% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
performance optimization at scale |
0.3 | 2 | 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014 4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012 |
High-performance computing
scientific computing systems |
0.3 | 2 | 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014 4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012 |
Programming languages and type systems
domain-specific languages |
0.2 | 1 | 2016 | Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016 |
High-performance computing
stencil computation |
0.2 | 1 | 2016 | Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016 |
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.2 | 2 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
GPUs and heterogeneous computing › GPU computing › GPGPU acceleration
GPU-accelerated supercomputing |
0.2 | 1 | 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014 |
High-performance computing
n-body simulation |
0.2 | 1 | 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014 |
High-performance computing › performance optimization at scale
parallel scalability |
0.1 | 1 | 2012 | 4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problem · SC 2012 |
Interconnection networks and networks-on-chip › cluster interconnect
infiniband |
0.1 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
High-performance computing › n-body simulation
treecode |
0.1 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
Performance modeling and evaluation › numerical algorithms
fast multipole method |
0.1 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
High-performance computing › n-body simulation
hierarchical n-body methods |
0.1 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs · SC 2014 |
Performance modeling and evaluation › design trade-off analysis
cost-performance analysis |
0.0 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
Computational science and engineering › computational fluid dynamics
turbulence simulation |
0.0 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
Methods — techniques the papers use, named apart from their topics
domain-specific language · 0.5auto-tuning · 0.5MPI · 0.5treecode · 0.3parallel efficiency analysis · 0.2gravitational tree-code · 0.2tree algorithm · 0.1particle-mesh algorithm · 0.1TreePM · 0.1hierarchical n-body · 0.1fast multipole method · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Prompt Report on Exa-Scale HPL-AI BenchmarkabstractOur performance benchmark of HPL-AI on the supercomputer Fugaku was awarded in the 55th top500 at ISC20. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of exascale in a floating-point arithmetic benchmark. Due to the novelty of HPL-AI, there are few guidelines for large systems and several drawbacks to the large-scale benchmark. It is not enough to replace FP64 operations solely to those on FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and introduction of optimization techniques on extensive computing such as on Fugaku. In the poster, we give some comments on the accuracy, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku. Shuhei Kudo, Keigo Nitadori, Takuya Ina, Toshiyuki Imamura |
CLUSTER | 2 |
| 2016 | Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codesabstractStencil computation has many applications in science and engineering, thus many optimization techniques such as temporal blocking have been developed. They are, however, rarely used in real-world applications, since a large amount of careful programming is required for even the simplest of stencils. We introduce Formura, a domain specific language that provides easy access to optimized stencil computations. Higher-order integration schemes can be defined using mathematical notations. Formura generates C code with MPI calls and performs autotuning. Hence its performance is portable to most distributed-memory computers. We show the scientific applicability of Formura by performing magnetohydrodynamics (MHD) and belowground biology simulations. Ability to reach bytes-per-flops ratio only attainable by temporal blocking is demonstrated. We also demonstrate scaling up to the full nodes of the K computer, with 1.184 Pflops, 11.62% floating-pointoperation efficiency, and 31.26% memory throughput efficiency. Takayuki Muranushi, Hideyuki Hotta, Junichiro Makino, Seiya Nishizawa, Hirofumi Tomita, Keigo Nitadori, Masaki Iwasawa, Natsuki Hosono, Yutaka Maruyama, Hikaru Inoue, Hisashi Yashiro, Yoshifumi Nakamura |
SC | 6 |
| 2014 | 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUsabstractWe have simulated, for the first time, the long term evolution of the Milky Way Galaxy using 51 billion particles on the Swiss Piz Daint supercomputer with our N-body gravitational tree-code Bonsai. Herein, we describe the scientific motivation and numerical algorithms. The Milky Way model was simulated for 6 billion years, during which the bar structure and spiral arms were fully formed. This improves upon previous simulations by using 1000 times more particles, and provides a wealth of new data that can be directly compared with observations. We also report the scalability on both the Swiss Piz Daint and the US ORNL Titan. On Piz Daint the parallel efficiency of Bonsai was above 95%. The highest performance was achieved with a 242 billion particle Milky Way model using 18600 GPUs on Titan, thereby reaching a sustained GPU and application performance of 33.49 Pflops and 24.77 Pflops respectively. Jeroen Bédorf, Evghenii Gaburov, Michiko S. Fujii, Keigo Nitadori, Tomoaki Ishiyama, Simon Portegies Zwart |
SC | 4 |
| 2012 | 4.45 Pflops astrophysical N-body simulation on K computer: the gravitational trillion-body problemabstractAs an entry for the 2012 Gordon-Bell performance prize, we report performance results of astrophysical N-body simulations of one trillion particles performed on the full system of K computer. This is the first gravitational trillion-body simulation in the world. We describe the scientific motivation, the numerical algorithm, the parallelization strategy, and the performance analysis. Unlike many previous Gordon-Bell prize winners that used the tree algorithm for astrophysical N-body simulations, we used the hybrid TreePM method, for similar level of accuracy in which the short-range force is calculated by the tree algorithm, and the long-range force is solved by the particle-mesh algorithm. We developed a highly-tuned gravity kernel for short-range forces, and a novel communication algorithm for long-range forces. The average performance on 24576 and 82944 nodes of K computer are 1.53 and 4.45 Pflops, which correspond to 49% and 42% of the peak speed. Tomoaki Ishiyama, Keigo Nitadori, Junichiro Makino |
SC | 2 |
| 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUsabstractWe present the results of a hierarchical N-body simulation on DEGIMA, a cluster of PCs with 576 graphic processing units (GPUs) and using an InfiniBand interconnect. DEGIMA stands for DEstination for GPU Intensive MAchine, and is located at Nagasaki Advanced Computing Center (NACC), Nagasaki University. In this work, we have upgraded DEGIMA_s interconnect using InfiniBand. DEGIMA is composed by 144 nodes with 576 GT200 GPUs. An astrophysical N-body simulation with 3,278,982,596 particles using a treecode algorithm shows a sustained performance of 190.5 Tflops on DEGIMA. The overall cost of the hardware was $411,921 dollars. The maximum corrected performance is 104.8 Tflops for the simulation, resulting in a cost performance of 254.4 MFlops/$. This corrections is performed by counting the FLOPS based on the most efficient CPU algorithm. Any extra FLOPS that arise from the GPU implementation and parameter differences are not included in the 254.4 MFLOPS/$. Tsuyoshi Hamada, Keigo Nitadori |
SC | 2 |
| 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulenceabstractAs an entry for the 2009 Gordon Bell price/performance prize, we present the results of two different hierarchical N-body simulations on a cluster of 256 graphics processing units (GPUs). Unlike many previous N-body simulations on GPUs that scale as O(N2), the present method calculates the O(N log N) treecode and O(N) fast multipole method (FMM) on the GPUs with unprecedented efficiency. We demonstrate the performance of our method by choosing one standard application --a gravitational N-body simulation-- and one non-standard application --simulation of turbulence using vortex particles. The gravitational simulation using the treecode with 1,608,044,129 particles showed a sustained performance of 42.15 TFlops. The vortex particle simulation of homogeneous isotropic turbulence using the periodic FMM with 16,777,216 particles showed a sustained performance of 20.2 TFlops. The overall cost of the hardware was 228,912 dollars. The maximum corrected performance is 28.1TFlops for the gravitational simulation, which results in a cost performance of 124 MFlops/$. This correction is performed by counting the Flops based on the most efficient CPU algorithm. Any extra Flops that arise from the GPU implementation and parameter differences are not included in the 124 MFlops/$. Tsuyoshi Hamada, Tetsu Narumi, Rio Yokota, Kenji Yasuoka, Keigo Nitadori, Makoto Taiji |
SC | 5 |