EDBT 2026 Demo / reviewers in the wild / expert
Shuhei Kudo
dblp:64/10222
· DBLP profile ↗
8ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0003-7168-5764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing
data assimilation |
0.4 | 1 | 2020 | A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020 |
High-performance computing › large-scale simulation
numerical weather prediction |
0.4 | 1 | 2020 | A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020 |
High-performance computing
scientific computing systems |
0.4 | 1 | 2020 | A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulations · SC 2020 |
Methods — techniques the papers use, named apart from their topics
ensemble kalman filter · 0.9approximate computing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximate Block Diagonalization of Symmetric Matrices Using the D-Wave Advantage Quantum AnnealerabstractABSTRACT Approximate block diagonalization is a problem of transforming a given symmetric matrix as close to block diagonal as possible by symmetric permutations of its rows and columns. This problem arises as a preprocessing stage of various scientific calculations and has been shown to be NP‐complete. In this paper, we consider solving this problem approximately using the D‐Wave Advantage quantum annealer. For this purpose, several steps are needed. First, we have to reformulate the problem as a quadratic unconstrained binary optimization (QUBO) problem. Second, the QUBO has to be embedded into the physical qubit network of the quantum annealer. Third, and optionally, reverse annealing for improving the solution can be applied. We propose two QUBO formulations and four embedding strategies for the problem and discuss their advantages and disadvantages. Through numerical experiments, it is shown that the combination of domain‐wall encoding and D‐Wave's automatic embedding is the most efficient in terms of usage of physical qubits, while the combination of one‐hot encoding and automatic embedding is superior in terms of the probability of obtaining a feasible solution. It is also shown that reverse annealing is effective in improving the solution for medium‐sized problems. Koushi Teramoto, Evgeniy Mishchenko, Keisuke Kawamura, Shuhei Kudo, Yasuhiko Takenaga, Yusaku Yamamoto |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Approximate Block Diagonalization of Symmetric Matrices Using Quantum AnnealingabstractWe consider the problem of transforming a given symmetric matrix into a nearly block diagonal form by permutation of its rows and columns. Such a transformation is useful as preconditioning to accelerate the convergence of an eigenvalue solver, but the problem of finding an optimal permutation that maximizes the Frobenius norms of the diagonal blocks is NP-complete. We formulate this problem as QUBO (Quadratic Unconstrained Binary Optimization) and solve it using D-Wave Advantage quantum annealing machine. Experimental results on small problems show that the true minimum can be obtained with high probability. We also discuss how to improve the mapping of the problem onto the physical qubit network to increase the size of the problems that can be solved. Koushi Teramoto, Masaki Kugaya, Shuhei Kudo, Yasuhiko Takenaga, Yusaku Yamamoto |
HPC Asia | 3 |
| 2024 | Automatic performance tuning using the ATMathCoreLib tool: Two experimental studies related to dense symmetric eigensolversabstractSummary We consider automatic performance tuning of dense symmetric eigenvalue problems using ATMathCoreLib, which is a library to assist automatic tuning. We deal with two problems, namely, automatic code selection for the symmetric generalized eigenvalue problem in distributed‐memory parallel environments and automatic parameter tuning in tridiagonalization of dense symmetric matrices on multicore processors. As for the first problem, numerical experiments show that ATMathCoreLib can choose the fastest solver for a given computing environment and problem size quickly even if the fluctuation in the execution time is as high as 40%. As for the second problem, ATMathCoreLib was able to select nearly optimal combinations of the algorithm and its parameter reliably and efficiently for various computing environments and matrix sizes. The performance of auto‐tuning was further enhanced by incorporating a user‐provided execution‐time model into ATMathCoreLib. Yusuke Hirota, Shuhei Kudo, Takeo Hoshi, Yusaku Yamamoto |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | A Rapid Euclidean Norm Calculation Algorithm that Reduces Overflow and Underflow
Takeyuki Harayama, Shuhei Kudo, Daichi Mukunoki, Toshiyuki Imamura, Daisuke Takahashi |
ICCSA (1) | 2 |
| 2020 | Prompt Report on Exa-Scale HPL-AI BenchmarkabstractOur performance benchmark of HPL-AI on the supercomputer Fugaku was awarded in the 55th top500 at ISC20. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of exascale in a floating-point arithmetic benchmark. Due to the novelty of HPL-AI, there are few guidelines for large systems and several drawbacks to the large-scale benchmark. It is not enough to replace FP64 operations solely to those on FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and introduction of optimization techniques on extensive computing such as on Fugaku. In the poster, we give some comments on the accuracy, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku. Shuhei Kudo, Keigo Nitadori, Takuya Ina, Toshiyuki Imamura |
CLUSTER | 1 |
| 2020 | A 1024-member ensemble data assimilation with 3.5-km mesh global weather simulationsabstractNumerical weather prediction (NWP) supports our daily lives. Weather models require higher spatiotemporal resolutions to prepare for extreme weather disasters and reduce the uncertainty of predictions. The accuracy of the initial state of the weather simulation is also critical; thus, we need more advanced data assimilation (DA) technology. By combining resolution and ensemble size, we have achieved the world’s largest weather DA experiment using a global cloud-resolving model and an ensemble Kalman filter method. The number of grid points was $\sim$4.4 trillion, and 1.3 PiB of data was passed from the model simulation part to the DA part. We adopted a data-centric application design and approximate computing to speed up the overall system of DA. Our DA system, named NICAM-LETKF, scales to 131,072 nodes (6,291,456 cores) of the supercomputer Fugaku with a sustained performance of 29 PFLOPS and 79 PFLOPS for the simulation and DA parts, respectively. Hisashi Yashiro, Koji Terasaki, Yuta Kawai, Shuhei Kudo, Takemasa Miyoshi, Toshiyuki Imamura, Kazuo Minami, Hikaru Inoue, Tatsuo Nishiki, Takayuki Saji, Masaki Satoh, Hirofumi Tomita |
SC | 4 |
| 2019 | Cache-efficient implementation and batching of tridiagonalization on manycore CPUsabstractWe herein propose an efficient implementation of tridiagonalization (TRD) for small matrices on manycore CPUs. Tridiagonalization is a matrix decomposition that is used as a preprocessor for eigenvalue computations. Further, TRD for such small matrices appears even in the HPC environment as a subproblem of large computations. Shuhei Kudo, Toshiyuki Imamura |
HPC Asia | 1 |
| 2017 | Performance analysis and optimization of the parallel one-sided block Jacobi SVD algorithm with dynamic ordering and variable blockingabstractSummary The one‐sided block Jacobi (OSBJ) method is known to be an efficient method for computing the singular value decomposition on a parallel computer. In this paper, we focus on the most recent variant of the OSBJ method, the one with parallel dynamic ordering and variable blocking, and present both theoretical and experimental analyses of the algorithm. In the first part of the paper, we provide a detailed theoretical analysis of its convergence properties. In the second part, based on preliminary performance measurement on the Fujitsu FX10 and SGI Altix ICE parallel computers, we identify two performance bottlenecks of the algorithm and propose new implementations to resolve the problem. Experimental results show that they are effective and can achieve up to 1.8 and 1.4 times speedup of the total execution time on the FX10 and the Altix ICE, respectively. Comparison with the ScaLAPACK SVD routine PDGESVD shows that our OSBJ solver is efficient when solving small to medium sized problems (n < 10000) using modest number ( < 100) of computing nodes. Copyright © 2016 John Wiley & Sons, Ltd. Shuhei Kudo, Yusaku Yamamoto, Martin Becka, Marián Vajtersic |
Concurr. Comput. Pract. Exp. | 1 |