Takahide Yoshikawa

dblp:10/3595 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Extraction and Representation of Sparsity Patterns for Efficient Data Transfer on Accelerators
abstract
Sparse computations are common in practical HPC, AI and graph-based applications. Such computations often exhibit scattered and fragmented data accesses, which negatively impact data transfer efficiency to/from accelerators. We propose, implement and evaluate an algorithm for extracting or mining sparsity patterns that exist in sparse matrices. The algorithm extracts multiple pattern types in a matrix, including blocks, bands, triangles or regular compositions of each. It does so without a priori knowledge of the presence of these patterns in the matrix. The patterns may contain, under user control, zero elements, or imperfections, to facilitate the extraction of larger patterns. Additionally, we introduce the Compressed Sparse Pattern (CSP), a novel compressed representation for sparse matrices that is based on these patterns. The use of CSP combined with extensions to Address Generation Units (AGUs) of accelerators regularize data accesses and improve data transfer efficiency. Evaluation of the pattern mining algorithm and CSP using 26 real-world sparse matrices is conducted on an Ubuntu system with an 8 core Intel CPU (3.6 GHz i7-9700K) and 32 GB of memory. The evaluation shows that patterns of different sizes and shapes are common, representing ∼82% of the non-zero elements in these matrices. The patterns can be efficiently extracted in time, with an average of 4.6 seconds. The evaluation also shows that the mining of composite patterns contributes ∼8% to the number of non-zeros in patterns and that imperfections increase pattern sizes with a minimal impact of only ∼7% zero elements in patterns. Finally, using CSP leads to up to 90% reduction in data transfer overhead, compared to CSR and CSC, both common compressed sparse matrix representations. These results validate our approach of extracting and representing patterns to improve data transfer efficiency.
Toshiyuki Ichiba, Katsuhiro Yoda, Yasuhiro Watanabe, Takahide Yoshikawa, Tarek S. Abdelrahman
SBAC-PAD5
2024 Raising Compute Density of Molecular Dynamics Simulation Through Approximate Memoization
abstract
Molecular dynamics (MD) simulation involves simulating the interactions of particles. MD has many applications in basic biological sciences, drug discovery, materials science, and other fields. Simulating$1\mu \mathrm{s}$of a 100K-atom system can take hours or days11https://www.bdr.riken.jp/en/research/labs/taiji-mlmdgrape4.html, where the compute-heavy aspect of MD is calculating the long-range forces between pairs of particles. In this paper, we explore the application of approximate computing in MD as a means to improve compute density. Specifically, we employ approximate memoization, where previously computed forces (and more) are stored in a table, and are retrieved in subsequent force calculations, provided the inputs to the force calculation are the same or similar. If the prior-computed table values can be used, significant computational work is avoided. In an experimental study, we apply software simulation to understand the degree to which approximation is feasible. We then propose a hardware implementation of memoization to be used within an ASIC MD simulator, MDGRAPE-4A [18]. We show that compute density, measured as$\text{pair-interactions}/(s\cdot\mu m^{2})$is improved substantially, between 40 % and 70 % for the studied cases. This is contingent on the particular system being simulated, the table size, and permitted level of approximation.
Salim Khemira, Xinyuan Wang 0003, Yutaka Tamiya, Makoto Taiji, Takahide Yoshikawa, Jason Helge Anderson
ASAP6
2023 Thermal Scaffolding for Ultra-Dense 3D Integrated Circuits
abstract
We address the thermal challenge of ultra-dense 3D (e.g., monolithic 3D) integrated circuits with multiple high-speed computing engines in the 3D stack. We present a new thermal scaffolding approach achieved through a combination of (1) new Back-End-of-Line (BEOL)-compatible dielectric materials for simultaneous high thermal conductivity and low dielectric constant, (2) new 3D physical co-design of BEOL dielectrics with thermal metal structures for uniform heat conduction with minimal metal insertion overhead, and (3) previous, experimentally demonstrated heatsink advances. Physical designs of thermal scaffolding enable 12-tier 7nm ultra-dense 3D IC with max temperatures ≤125 degrees Celsius: an iso-footprint, iso-delay, 4x improvement in stacked tiers.
Dennis Rich, Anna Kasperovich, Mohamadali Malakoutian, Robert M. Radway, Shiho Hagiwara, Takahide Yoshikawa, Srabanti Chowdhury, Subhasish Mitra
DAC6
2023 Out-of-Step Pipeline for Gather/Scatter Instructions
abstract
Wider SIMD units suffer from low scalability of gather/scatter instructions that appear in sparse matrix calculations. We address this problem with an out-of-step pipeline which tolerates bank conflicts of a multibank L1D by allowing element operations of SIMD instructions to proceed out of step with each other. We evaluated it with a sparse matrix-vector product kernel for matrices from HPCG and SuiteSparse Matrix Collection. The results show that, for the SIMD width of 1024 bit, it achieves 1.91 times improvement over a model of a conventional pipeline.
Yi Ge, Katsuhiro Yoda, Makiko Ito, Toshiyuki Ichiba, Takahide Yoshikawa, Ryota Shioya, Masahiro Goshima
DATE5
2022 Memory Bandwidth Conservation for SpMV Kernels Through Adaptive Lossy Data Compression
Makiko Ito, Takahide Yoshikawa, Yuan He 0002, Masaaki Kondo
PDCAT3
2018 The Tofu Interconnect D
abstract
In this paper, we introduce a new and highly scalable interconnect called Tofu interconnect D that will be used in the post-K machine. This machine will officially be operational around 2021. The letter D represents high “density” node and “dynamic” packet slicing for “dual-rail” transfer. Herein we describe the design and the evaluation results of TofuD. Due to the high-density packaging, the optical link ratio of TofuD has decreased to 25% from the 66% optical link ratio of Tofu2. TofuD applies a new technique called dynamic packet slicing to reduce latency and to improve fault resilience. The evaluation results show that the one-way 8-byte Put latency is 0.49 μs. This is 31% lower than the latency of Tofu2. The injection rate per node is 38.1 GB/s which is approximately 83% of the injection rate of Tofu2. The link efficiency is as high as approximately 93%.
Yuichiro Ajima, Takahiro Kawashima, Takayuki Okamoto, Naoyuki Shida, Kouichi Hirai, Toshiyuki Shimizu, Shinya Hiramoto, Yoshiro Ikeda, Takahide Yoshikawa, Kenji Uchida, Tomohiro Inoue
CLUSTER9