VLDB 2026 Research / reviewers in the wild / expert
Tomonori Kouya
dblp:153/2103
· DBLP profile ↗
4ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0003-0178-5519ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Acceleration of LU decomposition supporting double-double, triple-double, and quadruple-double precision floating-point arithmetic with AVX2abstractIn this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision LU decomposition with Intel's Advanced Vector Extensions 2 (AVX2). We targeted double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic designed using certain types of error-free transformation (EFT) arithmetic. We implemented accelerated DD, TD, and QD precision addition and multiplication using SIMDized EFT functions with AVX2, which perform simultaneous computation using four binary64 numbers on the x86_64 computing environment, and with these, we were able to develop multiple-precision LU decomposition based on SIMDized matrix multiplication. Our LU decomposition is up to three times faster than the non-accelerated one. Tomonori Kouya |
ARITH | 1 |
| 2021 | Acceleration of Multiple Precision Matrix Multiplication Based on Multi-component Floating-Point Arithmetic Using AVX2abstractIn this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision matrix multiplication with AVX2. We target double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic designed by certain types of error-free transformation (EFT) arithmetic. Furthermore, we implement SIMDized EFT functions, which simultaneously compute with four binary64 numbers on x86_64 computing environment, and by using help of them, we also develop SIMDized DD, TD, and QD additions and multiplications. In addition, AVX2 load/store functions were adopted to efficiently speed up reading and storing matrix elements from/to memory. Owing to these combined techniques, our implemented multiple precision matrix multiplications have been accelerated more than three times compared with non-accelerated ones. Our accelerated matrix multiplication modifies the performance of parallelization with OpenMP. Tomonori Kouya |
ICCSA (5) | 1 |
| 2020 | Performance Evaluation of Strassen Matrix Multiplication Supporting Triple-Double Precision Floating-Point Arithmetic
Tomonori Kouya |
ICCSA (5) | 1 |
| 2019 | Performance Evaluation of an Efficient Double-Double BLAS1 Function With Error-Free Transformation and its Application to Explicit Extrapolation MethodsabstractError-free transformation (EFT) has been recently applied to solve ill-conditioned problems. This transformation can reduce the number of arithmetic operations required compared to multiple precision arithmetic. In this study, we implement double-double (DD) BLAS1 functions with EFT and propose the application of the approach to explicit extrapolation methods for solving initial value problems of ordinary differential equations (ODEs). The presented routines can be effective for a large system of linear ODEs, especially when a harmonic sequence is used. Tomonori Kouya |
ARITH | 1 |