Kazushi Kawamura

dblp:125/0294 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-0795-2974ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BingoGCN: Towards Scalable and Efficient GNN Acceleration with Fine-Grained Partitioning and SLT
abstract
Graph Neural Networks (GNNs) are increasingly popular due to their wide applicability to tasks requiring the understanding of unstructured graph data, such as those in social network analysis and autonomous driving.However, real-time, large-scale GNN inference faces challenges due to the large size of node features and adjacency matrices, leading to memory communication and buffer size overheads caused by irregular memory access patterns.While graph partitioning can help with localized access patterns and reduction in on-chip buffer size, fine-grained partitioning results in increased inter-partition edges and off-chip memory accesses, negatively impacting overall performance.To overcome these limitations, we propose BingoGCN, a scalable GNN acceleration framework that introduces multidimensional dynamic feature summarization called Cross-Partition Message Quantization (CMQ) for inter-partition message passing.This eliminates irregular off-chip memory access without additional training and accuracy loss, even with fine-grained partitioning.By shifting the bottleneck from memory to computation, BingoGCN allows for further performance optimization through the Strong Lottery Ticket (SLT) theory using randomly generated weights.BingoGCN addresses the challenge of SLT's unstructured sparsity in hardware acceleration with a novel training algorithm and random weight generator designs, enabling fine-grained (FG) sparsity and improved load balancing.We integrated CMQ and FG-SLT into the messagepassing of GNNs and designed an efficient hardware architecture to support this flow.Our FPGA-based implementation achieves a significant reduction in memory accesses while preserving accuracy comparable to the original models.
Jiale Yan, Hiroaki Ito, Yuta Nagahara, Kazushi Kawamura, Masato Motomura, Thiem Van Chu, Daichi Fujiki
ISCA4
2025 Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
abstract
This paper proposes a novel matrix quantization method, Binary Quadratic Quan- tization (BQQ). In contrast to conventional first-order quantization approaches— such as uniform quantization and binary coding quantization—that approximate real-valued matrices via linear combinations of binary bases, BQQ leverages the expressive power of binary quadratic expressions while maintaining an extremely compact data format. We validate our approach with two experiments: a matrix compression benchmark and post-training quantization (PTQ) on pretrained Vision Transformer-based models. Experimental results demonstrate that BQQ consistently achieves a superior trade-off between memory efficiency and reconstruction error than conventional methods for compressing diverse matrix data. It also delivers strong PTQ performance, even though we neither target state-of-the-art PTQ accuracy under tight memory constraints nor rely on PTQ-specific binary matrix optimization. For example, our proposed method outperforms the state-of- the-art PTQ method by up to 2.2% and 59.1% on the ImageNet dataset under the calibration-based and data-free scenarios, respectively, with quantization equivalent to 2 bits. These findings highlight the surprising effectiveness of binary quadratic expressions for efficient matrix approximation and neural network compression.
Kyo Kuroki, Yasuyuki Okoshi, Thiem Van Chu, Kazushi Kawamura, Masato Motomura
NeurIPS4
2025 DMSA: An Efficient Architecture for Sparse-Sparse Matrix Multiplication Based on Distribute-Merge Product Dataflow
abstract
The sparse–sparse matrix multiplication (SpMSpM) is a fundamental operation in various applications. Existing SpMSpM accelerators based on inner product (IP) and outer product (OP) suffer from low computational efficiency and high memory traffic due to inefficient index matching and merging overheads. Gustavson’s product (GP)-based accelerators mitigate some of these challenges but struggle with workload imbalance and irregular memory access patterns, limiting computational parallelism. To overcome these limitations, we propose a distribute-merge product (DMP), a novel SpMSpM dataflow that evenly distributes workloads across multiple computation streams and merges partial results efficiently. We design and implement DMP-based SpMSpM architecture (DMSA), incorporating four key techniques to fully exploit the parallelism of DMP and efficiently handle irregular memory accesses. Implemented on a Xilinx ZCU106 FPGA, DMSA achieves speedups of up to$3.38\times $and$1.73\times $over two state-of-the-art FPGA-based SpMSpM accelerators while maintaining comparable hardware resource usage. In addition, compared to CPU and GPU implementations on an NVIDIA Jetson AGX Xavier, DMSA is$4.96\times $and$1.53\times $faster while achieving$6.67\times $and$2.33\times $better energy efficiency, respectively.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Daichi Fujiki, Masato Motomura, Thiem Van Chu
IEEE Trans. Very Large Scale Integr. Syst.3
2024 Sparse-Sparse Matrix Multiplication Accelerator on FPGA featuring Distribute-Merge Product Dataflow
abstract
Sparse-Sparse matrix multiplication (SpMSpM) is a critical computation in various fields such as computational science and graph analysis. It poses computational challenges for general-purpose CPUs and GPUs due to its requirements for random memory access and the inherently low spatial/temporal locality. Given the increasing importance of SpMSpM, numerous accelerators have been recently proposed. However, they suffer from various issues such as low input utilization, heavy computational load, and excessive memory traffic during the merging process of intermediate results. This paper introduces a novel Distribute-Merge Product (DMP) SpMSpM dataflow and a DMP-based SpMSpM Architecture (DMSA). DMP distributes the workload into balanced streams, generates partial matrices based on these streams, and merges the partial results in a parallel and pipelined fashion. We have designed DMSA as a highly scalable architecture, implemented it on a Xilinx ZCU106 Evaluation Kit, and evaluated it on a set of benchmarks from the SuiteSparse matrix collection. When compared to a latest SpMSpM accelerator with approximately the same amount of hardware resources on the same FPGA platform, DMSA achieves 2.72 × speedup, by facilitating the parallelism of partial matrix generation and merging. The speedup on the same platform reaches 4.80 × when the parallelism explored in the merging process is doubled, evidencing the DMSA’s superb scalability.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Masato Motomura, Thiem Van Chu
ASPDAC3
2024 Classical Thermodynamics-based Parallel Annealing Algorithm for High-speed and Robust Combinatorial Optimization
abstract
In recent years, quantum annealing has triggered active research on annealing methods for solving various combinatorial optimization problems (COPs) by mapping them to the Ising model based on spin glass theory. In particular, parallel annealing algorithms (PAAs) that can update all variables simultaneously attract attention due to fast optimization using parallel computers, either as an extension of Simulated Annealing rooted in classical thermodynamics or as a quantum-inspired algorithm. However, both types of PAAs face their own challenges. The classical thermodynamics-based PAAs (c-PAAs) perform inferior to the quantum-inspired PAAs (q-PAAs), whereas the q-PAAs require more parameters to be tuned than the c-PAAs. This paper proposes a new c-PAA based on Mean Field Annealing, which has the unique feature of updating analog variables deterministically. The proposed PAA achieves high speed and robustness despite fewer parameters than the q-PAAs, which means the proposed PAA breaks through the challenges of conventional PAAs. We demonstrate its performance through experiments on four types of COPs: Maximum Cut Problem, Graph Coloring Problem, Maximum Independent Set Problem, and Traveling Salesman Problem. These results imply that unless a real physical phenomenon is used, quantum-inspired algorithms cannot be considered superior to classical thermodynamics-based algorithms.
Kyo Kuroki, Satoru Jimbo, Thiem Van Chu, Masato Motomura, Kazushi Kawamura
GECCO5
2024 ETreeNet: Ensemble Model Fusing Decision Trees and Neural Networks for Small Tabular Data
abstract
In real-world machine learning applications, addressing the challenges associated with small tabular data is essential. While Decision Tree (DT)-based models are known to be effective for tabular data, their suitability diminishes when confronted with applications involving diverse data modalities beyond tabular data. Then, many studies focusing on tabular data propose Neural Networks (NN)-based models. To cope with the issue of limited data availability, most NN-based models for small tabular data utilize transformer architectures with techniques such as transfer learning, pre-training, and data augmentation. However, training or retraining a transformer-based model requires substantial data. This problem raises the question of whether it is appropriate to employ a transformer-based model for limited tabular data. We try to answer this question by proposing an ensemble model fusing DTs and NNs, called ETreeNet, which can outperform state-of-the-art transformer-based and DT-based models. ETreeNet comprises three methods: (1) Ensembling Tree-structured Neural Networks (TNNs), allowing training on small data due to reduced training parameters; (2) Sampling of features observed in Random Forest (RF) to enhance accuracy by reducing the influence of uninformative features; (3) Ensembling RF and TNNs to improve accuracy further. We conduct experiments using 500 instances of tabular training data, and the results show that ETreeNet achieves up to a 5% enhancement over state-of-the-art transformer-based and DT-based models.
Tsukasa Yamakura, Kazushi Kawamura, Masato Motomura, Thiem Van Chu
IJCNN2
2023 Decision Forest Training Accelerator Based on Binary Feature Decomposition
abstract
In recent years, while Deep Neural Networks (DNNs) have revolutionized various fields, it is widely acknowledged that they are not always the optimal solution, and complementary Machine Learning (ML) tools are necessary. For instance, developing DNN models that can effectively handle tabular data with rows and columns remains a challenging open question. Additionally, the difficulty of interpreting DNN models poses a significant obstacle that hinders their use in many practical applications where the interpretability of the inference results and the ability to offer advice on how to modify input for desired output are required. In such cases, Decision Forests (DFs) have been widely considered a promising solution.
Thiem Van Chu, Yu Mizutani, Yuta Nagahara, Shungo Kumazawa, Kazushi Kawamura, Jaehoon Yu, Masato Motomura
FCCM5
2022 Multicoated Supermasks Enhance Hidden Networks
abstract
Hidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing backpropagated information. To address this we propose a method that extends Hidden Networks by training an overlay of multiple hierarchical supermasks{—}a multicoated supermask. This method shows that using multiple supermasks for a single task achieves higher accuracy without additional training cost. Experiments on CIFAR-10 and ImageNet show that Multicoated Supermasks enhance the tradeoff between accuracy and model size. A ResNet-101 using a 7-coated supermask outperforms its Hidden Networks counterpart by 4%, matching the accuracy of a dense ResNet-50 while being an order of magnitude smaller.
Yasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando, Kazushi Kawamura, Thiem Van Chu, Masato Motomura, Jaehoon Yu
ICML5
2021 A High-Performance and Flexible FPGA Inference Accelerator for Decision Forests Based on Prior Feature Space Partitioning
abstract
Recent studies have demonstrated the potential of FPGAs for accelerating the inference computation of decision forests (DFs). However, designing a high-performance architecture that is flexible enough to be adopted in various scenarios of FPGA resource requirements remains a challenge. To address this, we propose a DF inference method that makes a transformation from traversing trees into traversing feature spaces. Specifically, as a preprocessing step, we partition each feature space into multiple regions based on thresholds. The inference task for an input data point is then conducted by (1) determining which region in each feature space the data point belongs to and (2) combining the inference information in these regions. The regularity of the computation allows us to design a DF inference architecture, called FT-DFP (Feature-space Traversing Decision Forest Processor), that can be flexibly configured for different performance and FPGA resource usage requirements. We prototype FT-DFP on a low-end FPGA (Artix-7) board and evaluate it using four real-world datasets. The evaluation results show that (1) the flexibility of FT-DFP allows us to fit a wide variety of DF models into low-end FPGA devices with limited resources; (2) FT-DFP's performance is comparable to the best of existing accelerators implemented on high-end FPGA devices and 3.04 × higher than Hummingbird, a state-of-the-art GPU-optimized implementation, running on a high-end GPU; and (3) FT-DFP is 130.96 × more energy-efficient than Hummingbird.
Thiem Van Chu, Ryuichi Kitajima, Kazushi Kawamura, Jaehoon Yu, Masato Motomura
FPT3
2021 Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation Aligner
abstract
A 4b-quantized convolutional neural network (CNN) inference engine for edge-AI is presented featuring a Cartesian-product MAC array and pipelined activation aligners targeting deep-/random-pruned models. A 40nm prototype with 32x32 MACs and 5Mb SRAM runs at 534 MHz, 1.07 TOPS, 352 mW at 1.1V, and attains 5.30 dense TOPS/W, 234 MHz at 0.8V. Sparse TOPS/W reaches 26.5 when running a randomly pruned model (after 88% pruning). Training algorithms for obtaining highly efficient sparse/quantized models are also proposed.
Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura
HCS5
2020 FPGA-based Heterogeneous Solver for Three-Dimensional Routing
abstract
A heuristic algorithm is one of the approaches to solve an NP-hard problem. In order to enhance the capability of the system, heterogeneous computing is often adapted. In this paper, we propose an FPGA-based heterogeneous solver for three-dimensional routing. The proposed system is implemented into multiple FPGA boards and a single-board computer. The experimental results demonstrate that the proposed system outperforms a single FPGA system.
Kento Hasegawa, Ryota Ishikawa, Makoto Nishizawa, Kazushi Kawamura, Masashi Tawada, Nozomu Togawa
ASP-DAC4
2013 A partial redundant fault-secure high-level synthesis algorithm for RDR architectures
abstract
In this paper, we propose a partial redundant fault-secure high-level synthesis algorithm for RDR architectures, where we duplicate a part of the original CDFG and maximize its reliability under a timing constraint. Firstly, our algorithm allocates some new additional functional units to vacant spaces on RDR islands for recomputation and increases the number of duplicated operation nodes. Secondly, it minimizes the number of inserted comparator nodes through re-scheduling/re-binding the recomputation CDFG's nodes. As a result, we will obtain a scheduled/bound recomputation CDFG and renewed functional unit allocation with high reliability. Experimental results demonstrate that our algorithm improves reliability by up to 52% compared with the conventional approach.
Kazushi Kawamura, Sho Tanaka, Masao Yanagisawa, Nozomu Togawa
ISCAS1