EDBT 2026 Demo / reviewers in the wild / expert
Yoichi Shimomura
dblp:266/2443 · also Youichi Shimomura
· DBLP profile ↗
13ranked-venue papers
1as first author
12since 2021 · last 2025
0009-0008-3231-2179ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Workflow Batch Job Scheduling with Considering Task Dependencies
Kaito Yanai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
JSSPP | 3 |
| 2024 | Modernizing an Operational Real-Time Tsunami Simulator to Support Diverse Hardware PlatformsabstractTo issue early warnings and rapidly initiate disaster responses after tsunami damage, various tsunami inundation forecast systems have been deployed worldwide. Japan's Cabinet Office operates a forecast system that utilizes supercomputers to perform tsunami propagation and inundation simulation in real time. Although this real-time approach is able to produce significantly more accurate forecasts than the conventional database-driven approach, its wider adoption was hindered because it was specifically developed for vector supercomputers. In this paper, we migrate the simulation code to modern CPUs and GPUs in a minimally invasive manner to reduce the testing and maintenance costs. A directive-based approach is employed to retain the structure of the original code while achieving performance portability, and hardware-specific optimizations including load balance improvement for GPUs are applied. The migrated code runs efficiently on recent CPUs, GPUs and vector processors: a six-hour tsunami simulation using over 47 million cells completes in less than 2.5 minutes on 32 Intel Sapphire Rapids CPUs and 1.5 minutes on 32 NVIDIA H100 GPUs. These results demonstrate that the code enables broader access to accurate tsunami inundation forecasts. Keichi Takahashi, Takashi Abe, Akihiro Musa, Yoshihiko Sato, Yoichi Shimomura, Hiroyuki Takizawa, Shunichi Koshimura |
CLUSTER | 5 |
| 2024 | Clustering Based Job Runtime Prediction for Backfilling Using Classification
Hang Cui 0005, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
JSSPP | 3 |
| 2024 | Maximizing Energy Budget Utilization Using Dynamic Power Cap Control
Sho Ishii, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
JSSPP | 3 |
| 2024 | A Node Selection Method for on-Demand Job Execution with Considering Deadline Constraints
Daiki Nakai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
JSSPP | 3 |
| 2024 | Conflict-aware workload co-execution on SX-aurora TSUBASAabstractAbstract NEC SX-Aurora TSUBASA (SX-AT) is the latest vector supercomputer, consisting of host processors called Vector Hosts (VHs) and vector processors called Vector Engines (VEs). The goal of this work is to simultaneously use both VHs and VEs to increase the resource utilization and improve the system throughput by co-executing more workloads. One difficulty is that performance interferences among VH and VE workloads could occur because they share some computing resources and potentially compete to use the same resource at the same time, so-called resource conflicts. To achieve efficient workload co-execution, first, this paper experimentally investigates the performance interference between a VH and a VE, when each of the two processors executes a different workload. It is empirically shown that the frequency of system calls from the VE workload could be a good indicator to predict if the co-execution could cause severe performance interference, even though monitoring system calls requires a huge runtime overhead and it is impractical to simply use it for decision making of co-execution. Then, this paper proposes a workload co-execution strategy based on a practical approach to identifying a pair of VE and VH workloads that could cause severe performance interferences. Our evaluation results clearly demonstrate that the system call frequency can be used to predict if the workload can affect the performance of another co-executing workload, and VH’s CPU load can be a good approximation of the system call frequency. The proposed approach based on the CPU loads could accurately identify a pair of workloads causing frequent resource conflicts, and thus reduce the risk of severe performance interferences between co-executing workloads on an SX-AT system, resulting in shorter makespan without significantly increasing the turn-around time. Riku Nunokawa, Yoichi Shimomura, Mulya Agung, Ryusuke Egawa, Hiroyuki Takizawa |
CCF Trans. High Perform. Comput. | 2 |
| 2022 | A Real-time Flood Inundation Prediction on SX-Aurora TSUBASAabstractDue to extreme weather, record-breaking heavy rainfalls frequently cause severe flood damages. Thus, there is a strong demand for predicting flood scales to mitigate damages. In this paper, we propose a real-time flood inundation prediction system on a shared HPC system. Although the Rainfall-Runoff Inundation (RRI) model has been developed for predicting large-scale flood inundation, it is necessary to improve the performance for real-time prediction. Since the RRI model is highly memory-bound, we port the RRI simulation code to the latest vector computing system, SX-Aurora TSUBASA (SX-AT), which provides high sustained memory bandwidth. We discuss performance optimization of the RRI code at the node level and MPI parallelization strategies. The RRI code also needs to output intermediate results at a high frequency. Thus, the RRI code is split into file I/O operation and kernel computation, which are assigned to different kinds of processors using the heterogeneity of SX-AT. Furthermore, we discuss a resource demand estimation method to minimize the amount of shared computing resources used for prediction in order to reduce the impact on other users sharing the system. In our evaluation, we demonstrate that SX-AT with only 32 cores can meet the real-time simulation requirement of simulating 7-hour flood inundation for the Tohoku region of Japan within 20 minutes. The evaluation results also demonstrate that the proposed method can adaptively adjust the computing resource amount used for the real-time simulation, and thus reduce the computing resource by 75% in comparison with the worst-case scenario of conservative static resource allocation. Yoichi Shimomura, Akihiro Musa, Yoshihiko Sato, Atsuhiko Konja, Guoqing Cui, Rei Aoyagi, Keichi Takahashi, Hiroyuki Takizawa |
HIPC | 1 |
| 2022 | Toward Building a Digital Twin of Job Scheduling and Power Management on an HPC System
Tatsuyoshi Ohmura, Yoichi Shimomura, Ryusuke Egawa, Hiroyuki Takizawa |
JSSPP | 2 |
| 2022 | A Task-Parallel Runtime for Heterogeneous Multi-node Vector Systems
Kazuki Ide, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
PDCAT | 3 |
| 2022 | Towards Priority-Flexible Task Mapping for Heterogeneous Multi-core NUMA Systems
Mulya Agung, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa |
PDCAT | 4 |
| 2022 | Equivalence Checking of Code Transformation by Numerical and Symbolic Approaches
Shunpei Sugawara, Keichi Takahashi, Yoichi Shimomura, Ryusuke Egawa, Hiroyuki Takizawa |
PDCAT | 3 |
| 2021 | Towards Conflict-Aware Workload Co-execution on SX-Aurora TSUBASA
Riku Nunokawa, Yoichi Shimomura, Mulya Agung, Ryusuke Egawa, Hiroyuki Takizawa |
PDCAT | 2 |
| 2009 | Performance evaluation of NEC SX-9 using real science and engineering applicationsabstractThis paper describes a new-generation vector parallel supercomputer, NEC SX-9 system. The SX-9 processor has an outstanding core to achieve over 100Gflop/s, and a software-controllable on-chip cache to keep the high ratio of the memory bandwidth to the floating-point operation rate. Moreover, its large SMP nodes of 16 vector processors with 1.6Tflop/s performance and 1TB memory are connected with dedicated network switches, which can achieve inter-node communication at 128GB/s per direction. The sustained performance of the SX-9 processor is evaluated using six practical applications in comparison with conventional vector processors and the latest scalar processor such as Nehalem-EP. Based on the results, this paper discusses the performance tuning strategies for new-generation vector systems. An SX-9 system of 16 nodes is also evaluated by using the HPC challenge benchmark suite and a CFD code. Those evaluation results clarify the highest sustained performance and scalability of the SX-9 system. Takashi Soga, Akihiro Musa, Yoichi Shimomura, Ryusuke Egawa, Ken'ichi Itakura, Hiroyuki Takizawa, Koki Okabe, Hiroaki Kobayashi |
SC | 3 |