Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hisashi Yashiro

dblp:149/6572 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-2678-526XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 70% Processor architecture and microarchitecture · 30%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing
data assimilation
0.412020
A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020
Processor architecture and microarchitecture
instruction set architecture
0.412020
Co-design for A64FX manycore processor and "Fugaku" · SC 2020
High-performance computing › large-scale simulation
numerical weather prediction
0.412020
A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020
High-performance computing
scientific computing systems
0.412020
A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020
High-performance computing
supercomputer architecture
0.412020
Co-design for A64FX manycore processor and "Fugaku" · SC 2020
Programming languages and type systems
domain-specific languages
0.212016
Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016
High-performance computing
stencil computation
0.212016
Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes · SC 2016

Methods — techniques the papers use, named apart from their topics

ensemble kalman filter · 0.9approximate computing · 0.9domain-specific language · 0.5auto-tuning · 0.5MPI · 0.5performance analysis · 0.4co-design · 0.4
YearPublicationVenuePosition
2025 Large Scale Ensemble Coupling of Non-hydrostatic Atmospheric Model NICAM
Takashi Arakawa, Hisashi Yashiro, Shinji Sumimoto, Kengo Nakajima
HPC Asia2
2024 Parallelized Remapping Algorithms for km-scale Global Weather and Climate Simulations with Icosahedral Grid System
abstract
In weather and climate research, latitude–longitude grid data are typically used for analysis and visualization, and remapping from model native grids to latitude–longitude grids typically requires a significant amount of time. Here, we developed a series of parallelized remapping algorithms for NICAM, a global weather and climate model with an icosahedral grid system, and demonstrated their performance with global 14–0.87-km mesh model data on the supercomputer Fugaku. The original remapping tool in NICAM supports parallelization only in reading and interpolating data. In our proposed algorithms, the process of data writing is parallelized by separating output files or using the MPI-IO library, both of which enable us to remap 0.87-km mesh data with 670 million horizontal grid points and 94 vertical levels. The benchmark with 14-km mesh data shows that the developed algorithms significantly outperform the original algorithm in terms of elapsed time (by 7.4–8.7 times) and memory usage (by 2.8–5.0 times). Among the proposed algorithms, the separation of output files, along with reduced MPI communication size, leads to a better performance in the elapsed time and its scalability, and the use of the MPI-IO library leads to a better performance in memory usage. The remapping year per wall-clock day, assuming a six-hourly output interval, is up to 0.56 with 3.5-km mesh data, demonstrating the feasibility of handling global cloud-resolving climate simulation data in a practical time. This study demonstrates the importance of IO performance, including MPI-IO, in accelerating weather and climate research on future supercomputers.
Chihiro Kodama, Hisashi Yashiro, Takashi Arakawa, Daisuke Takasuka, Shuhei Matsugishi, Hirofumi Tomita
HPC Asia2
2024 WaitIO-Hybrid: Communication for Coupling MPI Programs Among Heterogeneous Systems
Shinji Sumimoto, Takashi Arakawa, Yoshio Sakaguchi, Hiroya Matsuba, Satoshi Ohshima, Hisashi Yashiro, Toshihiro Hanawa, Kengo Nakajima
PDCAT6
2022 Development of a coupler h3-Open-UTIL/MP
abstract
In this work, we develop a coupler, h3-Open-UTIL/MP, as part of the h3-Open-BDEC project. The h3-Open-UTIL/MP coupler has an ensemble coupling capability to run multiple coupled programs in parallel and an interface to enable Python application coupling. In this paper, first, we outlined the features of h3-Open-UTIL/MP and its ensemble coupling capabilities. Then, we reported the details of Python application coupling and a coupling case study of the atmospheric model NICAM using a machine learning (ML) framework. To couple a Python application, we added a Python wrapper to the h3-Open-UTIL/MP interface provided as a Fortran module. Subsequently, we load Fortran routines as a library in the Python application. We used the ML framework PyTorch as a coupling example. We targeted the NICAM cloud physics subroutine and trained it to generate output variables from input variables. The results show a room for improvement in terms of reproducibility. Reproducibility can be improved in several ways: data selection, preprocess improvement, and changes in ML algorithms. In addition, the ML perspective was the bottleneck of execution time. Hence, there is a clear need to improve the computational performance. One effective method to improve computational performance is to run NICAM on a conventional high-performance computer and PyTorch on a GPU machine. We are currently developing a heterogeneous machine coupling program to implement this coupling method, which is briefly described in the final section of this paper.
Takashi Arakawa, Hisashi Yashiro, Kengo Nakajima
HPC Asia2
2022 A System-Wide Communication to Couple Multiple MPI Programs for Heterogeneous Computing
Shinji Sumimoto, Takashi Arakawa, Yoshio Sakaguchi, Hiroya Matsuba, Hisashi Yashiro, Toshihiro Hanawa, Kengo Nakajima
PDCAT5
2020 Co-design for A64FX manycore processor and "Fugaku"
abstract
We have been carrying out the FLAGSHIP 2020 Project to develop the Japanese next-generation flagship supercomputer, the Post-K, recently named “Fugaku”. We have designed an original many core processor based on Armv8 instruction sets with the Scalable Vector Extension (SVE), an A64FX processor, as well as a system including interconnect and a storage subsystem with the industry partner, Fujitsu. The “co-design” of the system and applications is a key to making it power efficient and high performance. We determined many architectural parameters by reflecting an analysis of a set of target applications provided by applications teams. In this paper, we present the pragmatic practice of our co-design effort for “Fugaku”. As a result, the system has been proven to be a very power-efficient system, and it is confirmed that the performance of some target applications using the whole system is more than 100 times the performance of the K computer.
Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama, Tetsuya Odajima, Miwako Tsuji, Hisashi Yashiro, Masaki Aoki, Naoyuki Shida, Ikuo Miyoshi, Kouichi Hirai, Atsushi Furuya, Akira Asato, Kuniki Morita, Toshiyuki Shimizu
SC7
2020 A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations
abstract
Numerical weather prediction (NWP) supports our daily lives. Weather models require higher spatiotemporal resolutions to prepare for extreme weather disasters and reduce the uncertainty of predictions. The accuracy of the initial state of the weather simulation is also critical; thus, we need more advanced data assimilation (DA) technology. By combining resolution and ensemble size, we have achieved the world’s largest weather DA experiment using a global cloud-resolving model and an ensemble Kalman filter method. The number of grid points was $\sim$4.4 trillion, and 1.3 PiB of data was passed from the model simulation part to the DA part. We adopted a data-centric application design and approximate computing to speed up the overall system of DA. Our DA system, named NICAM-LETKF, scales to 131,072 nodes (6,291,456 cores) of the supercomputer Fugaku with a sustained performance of 29 PFLOPS and 79 PFLOPS for the simulation and DA parts, respectively.
Hisashi Yashiro, Koji Terasaki, Yuta Kawai, Shuhei Kudo, Takemasa Miyoshi, Toshiyuki Imamura, Kazuo Minami, Hikaru Inoue, Tatsuo Nishiki, Takayuki Saji, Masaki Satoh, Hirofumi Tomita
SC1
2017 CONeP: A cost-effective online nesting procedure for regional atmospheric models
abstract
We propose a cost-effective online nesting procedure (CONeP) for regional atmospheric models to improve computational efficiency. The conventional procedure of online nesting is ineffective because computations are executed sequentially for each domain, and it does not enable users freely to determine the number of computational nodes. However, CONeP can completely avoid this limitation through three actions: 1) splitting the processes into multiple subgroups; 2) making each subgroup manage just one domain; and 3) executing the computations for each domain simultaneously. Since users can assign an optimal number of nodes to each domain, the model with CONeP is computationally efficient. We demonstrate the computational advantage of CONeP over the conventional procedure, comparing the elapsed times with both procedures on a supercomputer. The elapsed time with CONeP is markedly shorter than that observed with the conventional procedure using the same number of computational nodes. This advantage becomes more significant as the number of nesting domains increases.
Ryuji Yoshida, Seiya Nishizawa, Hisashi Yashiro, Sachiho A. Adachi, Yousuke Sato, Tsuyoshi Yamaura, Hirofumi Tomita
Parallel Comput.3
2016 Simulations of below-ground dynamics of fungi: 1.184 pflops attained by automated generation and autotuning of temporal blocking codes
abstract
Stencil computation has many applications in science and engineering, thus many optimization techniques such as temporal blocking have been developed. They are, however, rarely used in real-world applications, since a large amount of careful programming is required for even the simplest of stencils. We introduce Formura, a domain specific language that provides easy access to optimized stencil computations. Higher-order integration schemes can be defined using mathematical notations. Formura generates C code with MPI calls and performs autotuning. Hence its performance is portable to most distributed-memory computers. We show the scientific applicability of Formura by performing magnetohydrodynamics (MHD) and belowground biology simulations. Ability to reach bytes-per-flops ratio only attainable by temporal blocking is demonstrated. We also demonstrate scaling up to the full nodes of the K computer, with 1.184 Pflops, 11.62% floating-pointoperation efficiency, and 31.26% memory throughput efficiency.
Takayuki Muranushi, Hideyuki Hotta, Junichiro Makino, Seiya Nishizawa, Hirofumi Tomita, Keigo Nitadori, Masaki Iwasawa, Natsuki Hosono, Yutaka Maruyama, Hikaru Inoue, Hisashi Yashiro, Yoshifumi Nakamura
SC11
2014 Scalable rank-mapping algorithm for an icosahedral grid system on the massive parallel computer with a 3-D torus network
abstract
In this paper, we develop a rank-mapping algorithm for an icosahedral grid system on a massive parallel computer with the 3-D torus network topology, specifically on the K computer. Our aim is to improve the weak scaling performance of the point-to-point communications for exchanging grid-point values between adjacent grid regions on a sphere. We formulate a new rank-mapping algorithm to reduce the maximum number of hops for the point-to-point communications. We evaluate both the new algorithm and the standard ones on the K computer, using the communication kernel of the Nonhydrostatic Icosahedral Atmospheric Model (NICAM), a global atmospheric model with an icosahedral grid system. We confirm that, unlike the standard algorithms, the new one achieves almost perfect performance in the weak scaling on the K computer, even for 10,240 nodes. Results of additional experiments imply that the high scalability of the new rank-mapping algorithm on the K computer is achieved by reducing network congestion in the links between adjacent nodes.
Chihiro Kodama, Masaaki Terai, Akira T. Noda, Yohei Yamada, Masaki Satoh, Tatsuya Seiki, Shin-ichi Iga, Hisashi Yashiro, Hirofumi Tomita, Kazuo Minami
Parallel Comput.8