Tsuyoshi Ichimura

dblp:39/8426 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 2 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
High-performance computing · 83% Electronic design automation · 14% Parallel and multicore computing · 4%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › scientific computing systems
earthquake simulation
1.342022
Extreme Scale Earthquake Simulation with Uncertainty Quantification · SC 2022
A fast scalable implicit solver for nonlinear time-evolution earthquake city problem on low-ordered unstructured finite elements with artificial intelligence and transprecision computing · SC 2018
Implicit nonlinear wave simulation with 1.08T DOF and 0.270T unstructured finite elements to enhance comprehensive earthquake simulation · SC 2015
High-performance computing
scientific computing
1.342022
Extreme Scale Earthquake Simulation with Uncertainty Quantification · SC 2022
A fast scalable implicit solver for nonlinear time-evolution earthquake city problem on low-ordered unstructured finite elements with artificial intelligence and transprecision computing · SC 2018
Implicit nonlinear wave simulation with 1.08T DOF and 0.270T unstructured finite elements to enhance comprehensive earthquake simulation · SC 2015
Electronic design automation
uncertainty quantification
0.612022
Extreme Scale Earthquake Simulation with Uncertainty Quantification · SC 2022
High-performance computing
large-scale simulation
0.512021
High-Performance Computing Implementations of Agent-Based Economic Models for Realizing 1: 1 Scale Simulations of Large Economies · IEEE Trans. Parallel Distributed Syst. 2021

Methods — techniques the papers use, named apart from their topics

stochastic finite element method · 0.6mixed-precision implicit iterative solver · 0.6cache-efficient algorithms · 0.5OpenMP · 0.5MPI · 0.5unstructured finite element method · 0.4transprecision computing · 0.3artificial intelligence · 0.3implicit nonlinear solver · 0.2MPI-OpenMP hybrid parallelization · 0.2
YearPublicationVenuePosition
2022 152K-computer-node parallel scalable implicit solver for dynamic nonlinear earthquake simulation
abstract
We have used data learning and low-precision computation to develop an implicit solver that demonstrates high performance up to 152,352 computer nodes (609,408 MPI processes × 12 OpenMP threads = 7,312,896 parallel computation) and conducted an unprecedented ultra-large-scale analysis of ultra-high-fidelity fault-structure systems using nonlinear dynamic finite element analysis on three-dimensional low-order unstructured elements. The developed solver achieved 25.45-fold speedup from the state of the art solver on Fugaku and attained weak scaling efficiency of 93.7% from 9.391 billion [email protected] computer nodes to 1.201 trillion [email protected],984 computer nodes on performance measurement problems. Moreover, a realistic 324 billion DOF application example, which is difficult to obtain performance for, was computed in high performance. Since the developed solver is based on a highly generalizable algorithm, it is expected to contribute not only to earthquake simulation on Fugaku but also to the enhancement of similar applications in other fields and on other supercomputers.
Tsuyoshi Ichimura, Kohei Fujita, Kentaro Koyama, Ryota Kusakabe, Yuma Kikuchi, Takane Hori, Muneo Hori, Lalith Maddegedara, Noriyuki Ohi, Tatsuo Nishiki, Hikaru Inoue, Kazuo Minami, Seiya Nishizawa, Miwako Tsuji, Naonori Ueda
HPC Asia1
2022 Extreme Scale Earthquake Simulation with Uncertainty Quantification
abstract
We develop a stochastic finite element method with ultra-large degrees of freedom that discretize probabilistic and physical spaces using unstructured second-order tetrahedral elements with double precision using a mixed-precision implicit iterative solver that scales to the full Fugaku system and enables fast Uncertainty Quantification (UQ). The developed solver designed to attain high performance on a variety of CPU/GPU-based supercomputers enabled solving 37 trillion degrees-of-freedom problem with 19.8% peak FP64 performance on full Fugaku (89.8 PFLOPS) with 87.7% weak scaling efficiency, corresponding to 224-fold speedup over the state of the art solver running on full Summit. This method, which has shown its effectiveness via solving huge (32-trillion degrees-of-freedom) practical problems, is expected to be a breakthrough in damage mitigation, and is expected to facilitate the scientific understanding of earthquake phenomena and have a ripple effect on other fields that similarly require UQ.
Tsuyoshi Ichimura, Kohei Fujita, Ryota Kusakabe, Kentaro Koyama, Sota Murakami, Yuma Kikuchi, Takane Hori, Muneo Hori, Hikaru Inoue, Takafumi Nose, Takahiro Kawashima, Lalith Maddegedara
SC1
2021 High-Performance Computing Implementations of Agent-Based Economic Models for Realizing 1: 1 Scale Simulations of Large Economies
abstract
We present a scalable high-performance computing implementation of an agent-based economic model using distributed + shared-memory hybrid parallelization paradigms, capable of simulating 1:1 scale models of large economies like the eurozone. Agent-based economic models consist of millions of agents interacting over several graphs, which are either centralized or scale-free in nature. While most of the interactions are bi-directional, the interaction graphs are dense and random and keep evolving as the simulation progresses. These characteristics cause a very large and unknown number of random communications among MPI processes, posing challenges to developing scalable parallel extensions. Further, random access to large volume of data makes the algorithms highly memory-bound, severely degrading computational performance. Adopting various strategies inspired by the real-world functioning of economies, we reduce the large unknown number of communications to a known handful number. Memory-intensive algorithms are improved to make these cache-efficient, and advanced MPI functions are used to minimize communication overhead, thereby attaining higher performance and scalability. Further, an MPI + OpenMP hybrid model is developed to best utilize modern many-core computing nodes with low per-core memory capacity. It is demonstrated that our implementation can simulate a full fledged economic model with 331 million agents within 108 seconds using 128 CPU cores attaining 70 percent strong scalability.
Amit Gill, Lalith Maddegedara, Sebastian Poledna, Muneo Hori, Kohei Fujita, Tsuyoshi Ichimura
IEEE Trans. Parallel Distributed Syst.6
2020 The Effectiveness of Low-Precision Floating Arithmetic on Numerical Codes: A Case Study on Power Consumption
abstract
The low-precision floating point arithmetic that performs computation by reducing numerical accuracy with narrow bit-width is attracting since it can improve the performance of the numerical programs. Small memory footprint, faster computing speed, and energy saving are expected by performing calculation with low precision data. However, there have not been many studies on how low-precision arithmetics affects power and energy consumption of numerical codes. In this study, we investigate the power efficiency improvement by aggressively using low-precision arithmetics for HPC applications. In our evaluations, we analyze power characteristics of the Poisson's equation and the ground motion simulation programs with double precision and single precision floating point arithmetics. We confirm that energy efficiency improves by using low-precision arithmetics but it is heavily influenced by parameters such as data division and the number of OpenMP threads.
Ryuichi Sakamoto, Masaaki Kondo, Kohei Fujita, Tsuyoshi Ichimura, Kengo Nakajima
HPC Asia4
2019 Matched Filtering Accelerated by Tensor Cores on Volta GPUs With Improved Accuracy Using Half-Precision Variables
abstract
Matched Filtering can be applied to various fields owing to its ability to compute a correlation coefficient of two vectors and detect many template events. With an improvement in observation techniques, massive observation data and templates have been accumulated, in which a reduction of computation cost of Matched Filtering has become an important issue. This computation is mainly matrix-matrix product and Tensor Core on NVIDIA Volta GPU is expected to compute it rapidly. However, actual performance of Tensor Core is usually limited by the bandwidth of shared memory or global memory. In addition, only lower-precision data types are supported in the current API for Tensor Core. Therefore, we have to prevent a decline in accuracy in the computation. In this letter, we designed a Matched Filtering algorithm to solve these problems mentioned above and utilized high arithmetic capacity on Tensor Core. Specifically, we reduced the number of memory access to global memory and shared memory by using low-level description. In addition, we introduced local normalization to reduce the numerical error. We applied our developed kernel to template matching of seismic observation data and compared the performance and the accuracy with cuBLAS, a common library in GPU computation. When we compared the performance with the function in cuBLAS that offered almost the same accuracy as our kernel, we reduced the elapsed time by a factor of 4.74.
Takuma Yamaguchi, Tsuyoshi Ichimura, Kohei Fujita, Aitaro Kato, Shigeki Nakagawa
IEEE Signal Process. Lett.2
2018 Wave Propagation Simulation of Complex Multi-Material Problems with Fast Low-Order Unstructured Finite-Element Meshing and Analysis
abstract
Many wave-propagation analyses with varying geometries and material properties are expected to be useful for model optimization. Low-order unstructured finite-element methods are suitable for such analyses, as they are capable of modeling multi-material problems with complex geometries; however, the meshing and analysis cost is large. Therefore, in this paper, we developed a fast mesh-generator and analysis method. The robust mesh generator was 17.4-fold faster than a conventional mesh generator, and the predictor algorithm for dynamic implicit finite-element solvers showed a 1.69-fold increase in speed relative to conventional solvers and a 91.3% size-up efficiency on the full Oakforest-PACS system. We demonstrated the usability of the developed meshing and analysis methods via a wave-propagation simulation on a 1.9 billion unstructured tetrahedral-element model using half of the K computer system (41,472 compute nodes).
Kohei Fujita, Keisuke Katsushima, Tsuyoshi Ichimura, Masashi Horikoshi, Kengo Nakajima, Muneo Hori, Lalith Maddegedara
HPC Asia3
2018 A Fast Scalable Implicit Solver with Concentrated Computation for Nonlinear Time-Evolution Problems on Low-Order Unstructured Finite Elements
abstract
Many supercomputers are shifting to architectures with low B (byte/s; memory transfer capability) per F (FLOPS capability) ratios. However, utilizing increased F is difficult for applications that inherently require large B. Targeting an implicit unstructured low-order finite-element analysis solver, which typically requires large B, we have developed a concentrated computation algorithm that yields significant performance improvements on low B/F supercomputers. 35.7% peak performance was achieved for a sparse matrix-vector multiplication kernel, and 15.6% peak performance was achieved for the whole solver on the second generation Xeon Phi-based Oakforest-PACS. This is 5.02 times faster than (and 6.90 times the peak performance of) the state-of-the-art solver (the SC14 Gordon Bell finalist solver). On Oakforest-PACS, the proposed solver was approximately 2.42 times faster than the state-of-the-art solver running on the K computer. The proposed approach has implications for systems and applications and is expected to have significant impact on various fields that use finite-element methods for nonlinear time evolution problems.
Tsuyoshi Ichimura, Kohei Fujita, Masashi Horikoshi, Larry Meadows, Kengo Nakajima, Takuma Yamaguchi, Kentaro Koyama, Hikaru Inoue, Akira Naruse, Keisuke Katsushima, Muneo Hori, Lalith Maddegedara
IPDPS1
2018 A fast scalable implicit solver for nonlinear time-evolution earthquake city problem on low-ordered unstructured finite elements with artificial intelligence and transprecision computing
Tsuyoshi Ichimura, Kohei Fujita, Takuma Yamaguchi, Akira Naruse, Jack C. Wells, Thomas C. Schulthess, Tjerk P. Straatsma, Christopher Zimmer 0001, Maxime Martinasso, Kengo Nakajima, Muneo Hori, Lalith Maddegedara
SC1
2016 Automatic Evacuation Management Using a Multi Agent System and Parallel Meta-Heuristic Search
Leonel Aguilar Melgar, Lalith Maddegedara, Tsuyoshi Ichimura, Muneo Hori
PRIMA3
2015 Implicit nonlinear wave simulation with 1.08T DOF and 0.270T unstructured finite elements to enhance comprehensive earthquake simulation
abstract
This paper presents a new heroic computing method for unstructured, low-order, finite-element, implicit nonlinear wave simulation: 1.97 PFLOPS (18.6% of peak) was attained on the full K computer when solving a 1.08T degrees-of-freedom (DOF) and 0.270T-element problem. This is 40.1 times more DOF and elements, a 2.68-fold improvement in peak performance, and 3.67 times faster in time-to-solution compared to the SC14 Gordon Bell finalist's state-of-the-art simulation. The method scales up to the full K computer with 663,552 CPU cores with 96.6% sizeup efficiency, enabling solving of a 1.08T DOF problem in 29.7 s per time step. Using such heroic computing, we solved a practical problem involving an area 23.7 times larger than the state-of-the-art, and conducted a comprehensive earthquake simulation by combining earthquake wave propagation analysis and evacuation analysis. Application at such scale is a groundbreaking accomplishment and is expected to change the quality of earthquake disaster estimation and contribute to society.
Tsuyoshi Ichimura, Kohei Fujita, Pher Errol Balde Quinay, Lalith Maddegedara, Muneo Hori, Seizo Tanaka, Yoshihisa Shizawa, Hiroshi Kobayashi, Kazuo Minami
SC1
2014 A Scalable Workbench for Large Urban Area Simulations, Comprised of Resources for Behavioural Models, Interactions and Dynamic Environments
Leonel Aguilar Melgar, Lalith Maddegedara, Muneo Hori, Tsuyoshi Ichimura, Seizo Tanaka
PRIMA4
2014 Physics-Based Urban Earthquake Simulation Enhanced by 10.7 BlnDOF × 30 K Time-Step Unstructured FE Non-Linear Seismic Wave Simulation
abstract
With the aim of dramatically improving the reliability of urban earthquake response analyses, we developed an unstructured 3-D finite-element-based MPI-OpenMP hybrid seismic wave amplification simulation code, GAMERA. On the K computer, GAMERA was able to achieve a size-up efficiency of 87.1% up to the full K computer. Next, we applied GAMERA to a physics-based urban earthquake response analysis for Tokyo. Using 294,912 CPU cores of the K computer for 11 h, 32 min, we analyzed the 3-D non-linear ground motion of a 10.7 BlnDOF problem with 30 K time steps. Finally, we analyzed the stochastic response of 13,275 building structures in the domain considering uncertainty in structural parameters using 3 h, 56 min of 80,000 CPU cores of the K computer. Although a large amount of computer resources is needed presently, such analyses can change the quality of disaster estimations and are expected to become standard in the future.
Tsuyoshi Ichimura, Kohei Fujita, Seizo Tanaka, Muneo Hori, Lalith Maddegedara, Yoshihisa Shizawa, Hiroshi Kobayashi
SC1
2013 On the Development of an MAS Based Evacuation Simulation System: Autonomous Navigation and Collision Avoidance
Leonel Aguilar Melgar, Lalith Maddegedara, Muneo Hori, Tsuyoshi Ichimura, Seizo Tanaka
PRIMA4