Masato Motomura

dblp:89/1208 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-1543-1252ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
abstract
The strong lottery ticket hypothesis (SLTH) conjectures that high-performing subnetworks, called strong lottery tickets (SLTs), are hidden in randomly initialized neural networks. Although recent theoretical studies have established the SLTH across various neural architectures, the SLTH for transformer architectures still lacks theoretical understanding. In particular, the current theory of the SLTH does not yet account for the multi-head attention (MHA) mechanism, a core component of transformers. To address this gap, we introduce a theoretical analysis of the existence of SLTs within MHAs. We prove that, if a randomly initialized MHA of H heads and input dimension d has the hidden dimension O(d log(Hd^(3/2))) for the key and value, it contains an SLT that approximates an arbitrary MHA with the same input dimension with high probability. Furthermore, by leveraging this theory for MHAs, we extend the SLTH to transformers without normalization layers. We empirically validate our theoretical findings, demonstrating that the approximation error between the SLT within a source model (MHA and transformer) and an approximate target counterpart decreases exponentially by increasing the hidden dimension of the source model.
Hikari Otsuka, Daiki Chijiwa, Yasuyuki Okoshi, Daichi Fujiki, Susumu Takeuchi, Masato Motomura
AAAI6
2026 AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation Quantization
abstract
Processing-in-Memory (PIM) architectures offer a promising solution to the memory bottlenecks in data-intensive machine learning, yet often overlook the growing challenge of activation memory footprint. Conventional PIM approaches struggle with massive KV cache sizes generated in long-context scenarios by Transformer-based models, frequently exceeding PIM's limited memory capacity, while techniques like sparse attention can conflict with PIM's need for data locality. Existing PIM approaches and quantization methods are often insufficient or poorly suited for leveraging the unique characteristics of activations. This work identifies an opportunity for PIMspecialized activation quantization to enhance bandwidth and compute efficiency. We explore clustering-based vector quantization approaches, which align well with activation characteristics and PIM's internal bandwidth capabilities. Building on this, we introduce AQPIM, a novel PIM-aware activation quantization framework based on Product Quantization (PQ), optimizing it for modern Large Language Models (LLMs). By performing quantization directly within memory, AQPIM leverages PIM's high internal bandwidth and enables direct computation on compressed data, significantly reducing both memory footprint and computational overhead for attention computation. AQPIM addresses PQ's accuracy challenges by introducing several algorithmic optimizations. Evaluations demonstrate that AQPIM achieves significant performance improvements, drastically reducing of GPU-CPU communication that can account for$90 \sim 98.5 \%$of decoding latency, together with$3.4 \times$speedup over a SOTA PIM approach.
Kosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi Fujiki
HPCA3
2026 Efficient Vision Transformers via Token Merging with Head-Wise Attention Correction
abstract
Vision Transformers (ViTs) offer strong performance by modeling global relationships across image patches, but their scalability is limited by the quadratic cost of self-attention. To mitigate this, Token Merging (ToMe) reduces computation by merging similar tokens. This approach relies on proportional attention to preserve the original balance of attention weights after merging. Yet, proportional attention does not fully resolve the attention distortion. It only compensates for a merged token’s influence on other tokens, while ignoring the fundamental distortion of the token’s own self-attention score. This change creates unaddressed distortions that vary across attention heads.In this work, we conduct a detailed analysis of these attention distortions and reveal their dependence on the query–key projection weights of each head. Based on this finding, we propose Head-wise Attention Correction (HAC), a method that adjusts attention scores after token merging by accounting for head-specific characteristics. HAC effectively mitigates the distortions overlooked by proportional attention, maintaining model accuracy while significantly reducing computation. Experiments on ImageNet demonstrate that our method effectively improves the trade-off between efficiency and performance, advancing the development of efficient Vision Transformers via token merging. Code is available at https://github.com/ychikawa/HAC-ToMe
Yuki Ichikawa, Masato Motomura, Thiem Van Chu, Daichi Fujiki
WACV2
2025 BingoGCN: Towards Scalable and Efficient GNN Acceleration with Fine-Grained Partitioning and SLT
abstract
Graph Neural Networks (GNNs) are increasingly popular due to their wide applicability to tasks requiring the understanding of unstructured graph data, such as those in social network analysis and autonomous driving.However, real-time, large-scale GNN inference faces challenges due to the large size of node features and adjacency matrices, leading to memory communication and buffer size overheads caused by irregular memory access patterns.While graph partitioning can help with localized access patterns and reduction in on-chip buffer size, fine-grained partitioning results in increased inter-partition edges and off-chip memory accesses, negatively impacting overall performance.To overcome these limitations, we propose BingoGCN, a scalable GNN acceleration framework that introduces multidimensional dynamic feature summarization called Cross-Partition Message Quantization (CMQ) for inter-partition message passing.This eliminates irregular off-chip memory access without additional training and accuracy loss, even with fine-grained partitioning.By shifting the bottleneck from memory to computation, BingoGCN allows for further performance optimization through the Strong Lottery Ticket (SLT) theory using randomly generated weights.BingoGCN addresses the challenge of SLT's unstructured sparsity in hardware acceleration with a novel training algorithm and random weight generator designs, enabling fine-grained (FG) sparsity and improved load balancing.We integrated CMQ and FG-SLT into the messagepassing of GNNs and designed an efficient hardware architecture to support this flow.Our FPGA-based implementation achieves a significant reduction in memory accesses while preserving accuracy comparable to the original models.
Jiale Yan, Hiroaki Ito, Yuta Nagahara, Kazushi Kawamura, Masato Motomura, Thiem Van Chu, Daichi Fujiki
ISCA5
2025 Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
abstract
Test-time scaling (TTS) has proven effective in enhancing the reasoning capabilities of large language models (LLMs). Verification plays a key role in TTS, simultaneously influencing (1) reasoning performance and (2) compute efficiency, due to the quality and computational cost of verification. In this work, we challenge the conventional paradigms of verification, and make the first attempt toward systematically investigating the impact of verification granularity—that is, how frequently the verifier is invoked during generation, beyond verifying only the final output or individual generation steps. To this end, we introduce Variable Granularity Search (VG-Search), a unified algorithm that generalizes beam search and Best-of-N sampling via a tunable granularity parameter $g$. Extensive experiments with VG-Search under varying compute budgets, generator-verifier configurations, and task attributes reveal that dynamically selecting $g$ can improve the compute efficiency and scaling behavior. Building on these findings, we propose adaptive VG-Search strategies that achieve accuracy gains of up to 3.1\% over Beam Search and 3.6\% over Best-of-N, while reducing FLOPs by over 52\%. We will open-source the code to support future research.
Hao Mark Chen, Guanxi Lu, Yasuyuki Okoshi, Zhiwen Mo, Masato Motomura, Hongxiang Fan
NeurIPS5
2025 Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
abstract
This paper proposes a novel matrix quantization method, Binary Quadratic Quan- tization (BQQ). In contrast to conventional first-order quantization approaches— such as uniform quantization and binary coding quantization—that approximate real-valued matrices via linear combinations of binary bases, BQQ leverages the expressive power of binary quadratic expressions while maintaining an extremely compact data format. We validate our approach with two experiments: a matrix compression benchmark and post-training quantization (PTQ) on pretrained Vision Transformer-based models. Experimental results demonstrate that BQQ consistently achieves a superior trade-off between memory efficiency and reconstruction error than conventional methods for compressing diverse matrix data. It also delivers strong PTQ performance, even though we neither target state-of-the-art PTQ accuracy under tight memory constraints nor rely on PTQ-specific binary matrix optimization. For example, our proposed method outperforms the state-of- the-art PTQ method by up to 2.2% and 59.1% on the ImageNet dataset under the calibration-based and data-free scenarios, respectively, with quantization equivalent to 2 bits. These findings highlight the surprising effectiveness of binary quadratic expressions for efficient matrix approximation and neural network compression.
Kyo Kuroki, Yasuyuki Okoshi, Thiem Van Chu, Kazushi Kawamura, Masato Motomura
NeurIPS5
2025 DMSA: An Efficient Architecture for Sparse-Sparse Matrix Multiplication Based on Distribute-Merge Product Dataflow
abstract
The sparse–sparse matrix multiplication (SpMSpM) is a fundamental operation in various applications. Existing SpMSpM accelerators based on inner product (IP) and outer product (OP) suffer from low computational efficiency and high memory traffic due to inefficient index matching and merging overheads. Gustavson’s product (GP)-based accelerators mitigate some of these challenges but struggle with workload imbalance and irregular memory access patterns, limiting computational parallelism. To overcome these limitations, we propose a distribute-merge product (DMP), a novel SpMSpM dataflow that evenly distributes workloads across multiple computation streams and merges partial results efficiently. We design and implement DMP-based SpMSpM architecture (DMSA), incorporating four key techniques to fully exploit the parallelism of DMP and efficiently handle irregular memory accesses. Implemented on a Xilinx ZCU106 FPGA, DMSA achieves speedups of up to$3.38\times $and$1.73\times $over two state-of-the-art FPGA-based SpMSpM accelerators while maintaining comparable hardware resource usage. In addition, compared to CPU and GPU implementations on an NVIDIA Jetson AGX Xavier, DMSA is$4.96\times $and$1.53\times $faster while achieving$6.67\times $and$2.33\times $better energy efficiency, respectively.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Daichi Fujiki, Masato Motomura, Thiem Van Chu
IEEE Trans. Very Large Scale Integr. Syst.5
2024 Sparse-Sparse Matrix Multiplication Accelerator on FPGA featuring Distribute-Merge Product Dataflow
abstract
Sparse-Sparse matrix multiplication (SpMSpM) is a critical computation in various fields such as computational science and graph analysis. It poses computational challenges for general-purpose CPUs and GPUs due to its requirements for random memory access and the inherently low spatial/temporal locality. Given the increasing importance of SpMSpM, numerous accelerators have been recently proposed. However, they suffer from various issues such as low input utilization, heavy computational load, and excessive memory traffic during the merging process of intermediate results. This paper introduces a novel Distribute-Merge Product (DMP) SpMSpM dataflow and a DMP-based SpMSpM Architecture (DMSA). DMP distributes the workload into balanced streams, generates partial matrices based on these streams, and merges the partial results in a parallel and pipelined fashion. We have designed DMSA as a highly scalable architecture, implemented it on a Xilinx ZCU106 Evaluation Kit, and evaluated it on a set of benchmarks from the SuiteSparse matrix collection. When compared to a latest SpMSpM accelerator with approximately the same amount of hardware resources on the same FPGA platform, DMSA achieves 2.72 × speedup, by facilitating the parallelism of partial matrix generation and merging. The speedup on the same platform reaches 4.80 × when the parallelism explored in the merging process is doubled, evidencing the DMSA’s superb scalability.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Masato Motomura, Thiem Van Chu
ASPDAC4
2024 Classical Thermodynamics-based Parallel Annealing Algorithm for High-speed and Robust Combinatorial Optimization
abstract
In recent years, quantum annealing has triggered active research on annealing methods for solving various combinatorial optimization problems (COPs) by mapping them to the Ising model based on spin glass theory. In particular, parallel annealing algorithms (PAAs) that can update all variables simultaneously attract attention due to fast optimization using parallel computers, either as an extension of Simulated Annealing rooted in classical thermodynamics or as a quantum-inspired algorithm. However, both types of PAAs face their own challenges. The classical thermodynamics-based PAAs (c-PAAs) perform inferior to the quantum-inspired PAAs (q-PAAs), whereas the q-PAAs require more parameters to be tuned than the c-PAAs. This paper proposes a new c-PAA based on Mean Field Annealing, which has the unique feature of updating analog variables deterministically. The proposed PAA achieves high speed and robustness despite fewer parameters than the q-PAAs, which means the proposed PAA breaks through the challenges of conventional PAAs. We demonstrate its performance through experiments on four types of COPs: Maximum Cut Problem, Graph Coloring Problem, Maximum Independent Set Problem, and Traveling Salesman Problem. These results imply that unless a real physical phenomenon is used, quantum-inspired algorithms cannot be considered superior to classical thermodynamics-based algorithms.
Kyo Kuroki, Satoru Jimbo, Thiem Van Chu, Masato Motomura, Kazushi Kawamura
GECCO4
2024 ETreeNet: Ensemble Model Fusing Decision Trees and Neural Networks for Small Tabular Data
abstract
In real-world machine learning applications, addressing the challenges associated with small tabular data is essential. While Decision Tree (DT)-based models are known to be effective for tabular data, their suitability diminishes when confronted with applications involving diverse data modalities beyond tabular data. Then, many studies focusing on tabular data propose Neural Networks (NN)-based models. To cope with the issue of limited data availability, most NN-based models for small tabular data utilize transformer architectures with techniques such as transfer learning, pre-training, and data augmentation. However, training or retraining a transformer-based model requires substantial data. This problem raises the question of whether it is appropriate to employ a transformer-based model for limited tabular data. We try to answer this question by proposing an ensemble model fusing DTs and NNs, called ETreeNet, which can outperform state-of-the-art transformer-based and DT-based models. ETreeNet comprises three methods: (1) Ensembling Tree-structured Neural Networks (TNNs), allowing training on small data due to reduced training parameters; (2) Sampling of features observed in Random Forest (RF) to enhance accuracy by reducing the influence of uninformative features; (3) Ensembling RF and TNNs to improve accuracy further. We conduct experiments using 500 instances of tabular training data, and the results show that ETreeNet achieves up to a 5% enhancement over state-of-the-art transformer-based and DT-based models.
Tsukasa Yamakura, Kazushi Kawamura, Masato Motomura, Thiem Van Chu
IJCNN3
2023 Decision Forest Training Accelerator Based on Binary Feature Decomposition
abstract
In recent years, while Deep Neural Networks (DNNs) have revolutionized various fields, it is widely acknowledged that they are not always the optimal solution, and complementary Machine Learning (ML) tools are necessary. For instance, developing DNN models that can effectively handle tabular data with rows and columns remains a challenging open question. Additionally, the difficulty of interpreting DNN models poses a significant obstacle that hinders their use in many practical applications where the interpretability of the inference results and the ability to offer advice on how to modify input for desired output are required. In such cases, Decision Forests (DFs) have been widely considered a promising solution.
Thiem Van Chu, Yu Mizutani, Yuta Nagahara, Shungo Kumazawa, Kazushi Kawamura, Jaehoon Yu, Masato Motomura
FCCM7
2022 Multicoated Supermasks Enhance Hidden Networks
abstract
Hidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing backpropagated information. To address this we propose a method that extends Hidden Networks by training an overlay of multiple hierarchical supermasks{—}a multicoated supermask. This method shows that using multiple supermasks for a single task achieves higher accuracy without additional training cost. Experiments on CIFAR-10 and ImageNet show that Multicoated Supermasks enhance the tradeoff between accuracy and model size. A ResNet-101 using a 7-coated supermask outperforms its Hidden Networks counterpart by 4%, matching the accuracy of a dense ResNet-50 while being an order of magnitude smaller.
Yasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando, Kazushi Kawamura, Thiem Van Chu, Masato Motomura, Jaehoon Yu
ICML7
2022 Real-Time Tone Mapping: A Survey and Cross-Implementation Hardware Benchmark
abstract
The rising demand for high quality display has ensued active research in high dynamic range (HDR) imaging, which has the potential to replace the standard dynamic range imaging. This is due to HDR’s features like accurate reproducibility of a scene with its entire spectrum of visible lighting and color depth. But this capability comes with expensive capture, display, storage and distribution resource requirements. Also, display of HDR images/video content on an ordinary display device with limited dynamic range requires some form of adaptation. Many adaptation algorithms, widely known as tone mapping (TM) operators, have been studied and proposed in the last few decades. In this article, we present a comprehensive survey of 60 TM algorithms that have been implemented on hardware for acceleration and real-time performance. In this state-of-the-art survey, we will discuss those TM algorithms which have been implemented on GPU, FPGA, and ASIC in terms of their hardware specifications and performance. Output image quality is an important metric for TM algorithms. From our literature survey we found that, various objective quality metrics have been used to demonstrate the quality of those algorithms hardware implementation. We have compiled those metrics used in this survey, and analyzed the relationship between hardware cost, image quality and computational efficiency. Currently, machine learning-based (ML) algorithms have become an important tool to solve many image processing tasks, and this article concludes with a discussion on the future research directions to realize ML-based TM operators on hardware.
Yafei Ou, Prasoon Ambalathankandy, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai, Masayuki Ikebe
IEEE Trans. Circuits Syst. Video Technol.4
2021 Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
Ángel López García-Arias, Masanori Hashimoto, Masato Motomura, Jaehoon Yu
BMVC3
2021 A High-Performance and Flexible FPGA Inference Accelerator for Decision Forests Based on Prior Feature Space Partitioning
abstract
Recent studies have demonstrated the potential of FPGAs for accelerating the inference computation of decision forests (DFs). However, designing a high-performance architecture that is flexible enough to be adopted in various scenarios of FPGA resource requirements remains a challenge. To address this, we propose a DF inference method that makes a transformation from traversing trees into traversing feature spaces. Specifically, as a preprocessing step, we partition each feature space into multiple regions based on thresholds. The inference task for an input data point is then conducted by (1) determining which region in each feature space the data point belongs to and (2) combining the inference information in these regions. The regularity of the computation allows us to design a DF inference architecture, called FT-DFP (Feature-space Traversing Decision Forest Processor), that can be flexibly configured for different performance and FPGA resource usage requirements. We prototype FT-DFP on a low-end FPGA (Artix-7) board and evaluate it using four real-world datasets. The evaluation results show that (1) the flexibility of FT-DFP allows us to fit a wide variety of DF models into low-end FPGA devices with limited resources; (2) FT-DFP's performance is comparable to the best of existing accelerators implemented on high-end FPGA devices and 3.04 × higher than Hummingbird, a state-of-the-art GPU-optimized implementation, running on a high-end GPU; and (3) FT-DFP is 130.96 × more energy-efficient than Hummingbird.
Thiem Van Chu, Ryuichi Kitajima, Kazushi Kawamura, Jaehoon Yu, Masato Motomura
FPT5
2021 Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation Aligner
abstract
A 4b-quantized convolutional neural network (CNN) inference engine for edge-AI is presented featuring a Cartesian-product MAC array and pipelined activation aligners targeting deep-/random-pruned models. A 40nm prototype with 32x32 MACs and 5Mb SRAM runs at 534 MHz, 1.07 TOPS, 352 mW at 1.1V, and attains 5.30 dense TOPS/W, 234 MHz at 0.8V. Sparse TOPS/W reaches 26.5 when running a randomly pruned model (after 88% pruning). Training algorithms for obtaining highly efficient sparse/quantized models are also proposed.
Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura
HCS7
2021 A 96-MB 3D-Stacked SRAM Using Inductive Coupling With 0.4-V Transmitter, Termination Scheme and 12: 1 SerDes in 40-nm CMOS
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 35% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates for the degradation of the first transmitted bit, a coil termination scheme that aims to eliminate the ringing of 3D inductive coupling bus, and a 12:1 SerDes that minimizes power consumption and area overhead in inductive coupling channels. Low-power, large-capacity, 3-cycle latency 3D-stacked SRAM for a DNN accelerator is achieved with the combination of these techniques to serve as a replacement of 3D-stacked DRAM. The performance of the proposed 3D-SRAM is compared with HBM DRAM and achieves more than 50% lower energy consumption. The scaling scenario of the SRAM module is discussed in light of the scaling of the inductive coupling technology and logic process.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.5
2020 A 3D-Stacked SRAM using Inductive Coupling with Low-Voltage Transmitter and 12: 1 SerDes
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 45% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates the degradation of the first signal, a coil termination scheme that aims to eliminate the noise of 3D inductive coupling bus, and a 12:1 SerDes. The data density of the SRAM should reach 12.3-MB/mm3, which extends beyond that of state-of-the-art stacked DRAMs.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Kota Ando, Kazutoshi Hirose, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
ISCAS7
2020 An Adaptive Global and Local Tone Mapping Algorithm Implemented on FPGA
abstract
We present a fast global and locally adaptive tone mapping algorithm and its field-programmable gate array (FPGA) implementation. The specially designed tone mapping function, which is based on local histogram equalization, controls global, and local characteristics individually. In contrast to other tonemap operators, our algorithm manages light/dark halos separately and by using local tonemap function alone, it can effectively suppress noise. We validated the effectiveness of our algorithms using subjective and objective assessment. Using an average of the bins, we achieve fast smoothed local histogram estimation with fewer bins while maintaining high accuracy. Our new implementation method requires minimal data access and reduced memory as it operates with a downscaled frame size of 240 × 135 pixels. Relative local area size is 248 × 248 @Full-HD resolution (1920 × 1080). For low-latency pixel output, the system performs the tone mapping using pixel information from the previous frame. When we implemented the system on FPGA (TB-7K-325TIMG and Xilinx Kintex-7), we achieved lightweight hardware as the total usage rate is about 25% of the available FPGA resource. Using an online 1080p video we demonstrate, a real-time video processing using our hardware tone mapping system.
Prasoon Ambalathankandy, Masayuki Ikebe, Takayuki Yoshida, Takeshi Shimada, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai
IEEE Trans. Circuits Syst. Video Technol.6
2019 DeltaNet: Differential Binary Neural Network
abstract
Energy-constrained neural network processing is in high demanded for various mobile applications. Binarized neural network (BNN) aggressively enhances the computational efficiency, and in contrast, it suffers from degradation of accuracy due to its extreme approximation. We propose a neural network model using a new activation function "Delta" based on binarization of differences between weighted-sums. The "Delta" retains the magnitude relation between numerical values, and conveys richer information than ordinary binarization. We can design the hardware architecture for the proposed model with almost the same elements as BNN. The evaluation shows that it achieves higher recognition accuracy than a conventional BNN with almost the same hardware configuration.
Yuka Oba, Kota Ando, Tetsuya Asai, Masato Motomura, Shinya Takamaeda-Yamazaki
ASAP4
2018 Dither NN: An Accurate Neural Network with Dithering for Low Bit-Precision Hardware
abstract
Energy-constrained neural network processing is in high demanded for various mobile applications. Binary neural network aggressively enhances the computational efficiency, and in contrast, it suffers from degradation of accuracy due to its extreme approximation. We propose a novel accurate neural network model based on binarization and "dithering" that distributes the quantization error to neighboring pixels. The quantization errors in the binarization are distributed in the plane, so that a pixel in the multi-level source expression more accurately represented in the resulting binarized plane by multiple pixels. We designed a low-overhead binary-based hardware architecture for the proposed model. The evaluation results show that this method can be realized with a few additional lightweight hardware components.
Kota Ando, Kodai Ueyoshi, Yuka Oba, Kazutoshi Hirose, Ryota Uematsu, Takumi Kudo, Masayuki Ikebe, Tetsuya Asai, Shinya Takamaeda-Yamazaki, Masato Motomura
FPT10
2018 Sparse Disparity Estimation Using Global Phase Only Correlation for Stereo Matching Acceleration
abstract
In this study, we propose an efficient stereo matching method which estimates sparse disparities using global phase only correlation (POC). Conventionally, cost functions are to be calculated for all disparity candidates and the associated computational cost has been impediment in achieving a realtime performance. Therefore, we consider to use full image 2D phase only correlation (FIPOC) for detecting the valid disparities candidates. This would require comparatively fewer calculations for the same number of disparities. Since, the FIPOC output indicates the disparity distribution of two stereo images, we can sort the disparity candidates and choose them for sparse calculation. In our proposed method, the searchable disparity range is half of the input image size, which is much wider than that of the conventional methods. When we apply the FIPOC to naive sum of absolute difference (SAD) stereo matching method, the combined algorithm would require fewer calculations while maintaining the same accuracy. In our evaluation, the proposed method achieves 194 disparity stereo matching in 70 ms on$398 \times 288$images without the need for SIMD instruction, multi-thread operation, or additional hardware while using a Intel Core i5-5257U.
Takeshi Shimada, Masayuki Ikebe, Prasoon Ambalathankandy, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai
ICASSP5
2018 Analysis of Smoothed LHE Methods for Processing Images with Optical Illusions
abstract
To replicate human visual perception, we analyze processing images with optical illusion using edge preserving filters and smoothed local histogram equalization (LHE). Images with the optical illusions are good models for gradual/rapid changes in contrast and strong edges, which are good cases for assessing the robustness of image filters. Here, we study and analyze the performance of smoothed LHE filters while processing perceptual illusion. Our studies conclude that, smoothed LHEs are useful in retaining actual edge forms in these images as they can operate using large kernel sizes. These large kernel size filters can construct sawtooth like edge and it corresponds to adequately wide halos. We also demonstrate the usefulness of smoothed LHE like tone mapping techniques in preserving naturalness, and we confirmed it by performing subjective visual test.
Prasoon Ambalathankandy, Takeshi Shimada, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai, Masayuki Ikebe
VCIP4
2017 An image sensor/processor 3D stacked module featuring ThruChip interfaces
abstract
A 1,000-fps motion vector (MV) estimation and classification engine for high-speed computational imaging in a 3D stacked imager/processor module is proposed, prototyped, assembled and tested. The module features 1) ThruChip interfaces for high fps image transfer, 2) orders of magnitude more area/power efficient MV estimation architecture compared to conventional ones, and 3) a cognitive classification scheme employed on MV patterns, enabling the classification of moving objects not possible in conventional proposals.
Masayuki Ikebe, Tetsuya Asai, Masafumi Mori, Toshiyuki Itou, Daisuke Uchida, Yasuhiro Take, Tadahiro Kuroda, Masato Motomura
ASP-DAC8
2017 A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only)
Hiroki Nakahara, Haruyoshi Yonekawa, Hisashi Iwamoto, Masato Motomura
FPGA4
2017 FPGA implementation of edge-guided pattern generation for motion-vector estimation of textureless objects
abstract
The widely accepted block-matching technique, which is required to identify motion vectors, fails in cases in which texture is not existent. In [1], we proposed a hardware-oriented cellular-automaton algorithm that generates spatial patterns on textureless objects and backgrounds, aiming at motion-vector estimation of textureless moving objects. This demonstration presents a field-programmable gate array (FPGA) system that supports real-time processing. This system provides motion-vectors in moving textureless objects and enables enhanced processing of motion vector classification.
Aoi Tanibata, Alexandre Schmid, Shinya Takamaeda-Yamazaki, Masayuki Ikebe, Masato Motomura, Tetsuya Asai
FPL5
2017 Exploring optimized accelerator design for binarized convolutional neural networks
abstract
The convolutional neural network (CNN) is a state-of-the-art model that can achieve significantly high accuracy in many machine-learning tasks. Recently, for further developing the practical applications of CNNs, efficient hardware platforms for accelerating CNN have been throughly studied. A binarized neural network has been reported to minimize the multipliers, which consume a large amount of resources, with a minimal decrease in accuracy. In this study, we analyzed the optimal performance of CNN implemented on an field programmable gate array (FPGA) considering its logic resources and a memory bandwidth, using multiple types of parallelisms such as kernels, pixels, and channels both in conventional and binarized CNNs. As a result, it became clear that all the parallelisms are required for the binarized neural network to obtain the best performance of 8.38 TOPS.
Kodai Ueyoshi, Kota Ando, Kentaro Orimo, Masayuki Ikebe, Tetsuya Asai, Masato Motomura
IJCNN6
2017 Live demonstration: Feature extraction system using restricted Boltzmann machines on FPGA
abstract
Real-time results obtained from an unsupervised feature extraction system using Restricted Boltzmann Machines (RBMs) implemented on FPGA are presented. The feature extraction application is demonstrated using the MNIST dataset, and the weights storing features are visualized in real-time. A digit classification is also performed based on the learning results. Our demonstration system performs 134 times faster than the compared conventional CPU.
Kodai Ueyoshi, Takao Marukame, Tetsuya Asai, Masato Motomura, Alexandre Schmid
ISCAS4
2016 A memory-based realization of a binarized deep convolutional neural network
abstract
A pre-trained deep convolutional neural network (CNN) is a feed-forward computation perspective, which is widely used for the embedded systems, requires high power-and-area efficiency. This paper realizes a binarized CNN which treats only binary 2-values (+1/-1) for the inputs and the weights. In this case, the multiplier is replaced with an EX-NOR circuit. To reduce both power and area, we realize the 2-valued CNN by off- and on-chip memories. Since our 2D convolution operations are realized by the on-chip memory, our implementation consumes lower power than the DSP block. We decompose the memory part, and realize them by a cascade of memories (LUT cascade). By introducing a batch normalization technique, the classification error for the binarized CNN can be improved. We implemented the CIFAR-10 benchmark on the NetFPGA-SUME board, which has the Xilinx Inc. Virtex 7 FPGA and three on-chip QDR II+ Synchronous SRAMs. Compared with the conventional FPGA realizations, the performance is 2.82 times faster, the power efficiency is 1.76 times, and the area efficiency is 29.13 times better.
Hiroki Nakahara, Haruyoshi Yonekawa, Tsutomu Sasao, Hisashi Iwamoto, Masato Motomura
FPT5
2016 Memory-error tolerance of scalable and highly parallel architecture for restricted Boltzmann machines in Deep Belief Network
abstract
A key aspect of constructing highly scalable Deep-learning microelectronic systems is to implement fault tolerance in the learning sequence. Error-injection analyses for memory is performed using a custom hardware model implementing parallelized restricted Boltzmann machines (RBMs). It is confirmed that the RBMs in Deep Belief Networks (DBNs) provides remarkable robustness against memory errors. Fine-tuning has significant effects on recovery of accuracy for static errors injected to the structural data of RBMs during and after learning, which are either at cell-level or block level. The memory-error tolerance is observable using our hardware networks with fine-graded memory distribution.
Kodai Ueyoshi, Takao Marukame, Tetsuya Asai, Masato Motomura, Alexandre Schmid
ISCAS4
2014 Caching memcached at reconfigurable network interface
abstract
Memcached is a technology that improves response speed of web servers by caching data on DRAMs in distributed servers. In order to achieve higher performance, memcached has been evaluated on various platforms. Among them, FPGA seems to be the most efficient platform to run memcached, and several research groups are trying to achieve higher throughput with it. However, it is difficult to utilize a large amount of memory (several dozen gigabytes) with an FPGA. Some groups are trying to solve this problem by using an embedded CPU for memory allocation and another group is employing an SSD. Unlike other approaches that try to replace memcached itself on FPGAs, our approach augments the software memcached running on the host CPU by caching its data and some operations at the FPGA-equipped network interface card (NIC) mounted on the server. The locality of memcached data enables the FPGA NIC to have a fairly high hit rate with a smaller memory. We first explore the cache parameters by software simulations and estimate the effectiveness of our approach, and then prototype a system to prove its effectiveness. Through our evaluation with YCSB, a standard key-value store (KVS) benchmarking tool, we estimate that the latency improved by an order of magnitude over software memcached running on a high performance CPU.
Eric Shun Fukuda, Hiroaki Inoue, Takashi Takenaka, Dahoo Kim, Tsunaki Sadahisa, Tetsuya Asai, Masato Motomura
FPL7
2014 Achieving higher performance of memcached by caching at network interface
abstract
As the volume of data that web services handle is becoming larger, many web service providers are utilizing memcached, an in-memory key-value store to improve their web server's performance. While memcached usually runs on a server with a high performance processor, various hardware platforms has been evaluated for running memcached in order to achieve higher performance. Recently, several works that use FPGAs have successfully achieved higher performance than Xeon. These works, however, struggles to utilize a large memory with FPGAs. In this paper, we propose a system that enables us to overcome this problem and enhances memcached by caching a part of software memcached's commands and data to the network interface card equipped with an FPGA and a DRAM. Our evaluation showed that the NIC cache has less than 30% of hit rate for workload with Latest key selection distribution, and 30% to 60% for Zipf distribution workloads.
Eric Shun Fukuda, Hiroaki Inoue, Takashi Takenaka, Dahoo Kim, Tsunaki Sadahisa, Tetsuya Asai, Masato Motomura
FPT7
2013 C-Based Complex Event Processing on Reconfigurable Hardware
abstract
This brief presents an efficient complex event-processing framework, designed to process a large number of sequential events on field-programmable gate arrays (FPGAs). Unlike conventional structured query language based approaches, our approach features logic automation constructed with a new C-based event language that supports regular expressions on the basis of C functions, so that a wide variety of event-processing applications can be efficiently mapped to FPGAs. Evaluations on an FPGA-based network interface card show that we can achieve 12.3 times better event-processing performance than does a CPU software in a financial trading application.
Hiroaki Inoue, Takashi Takenaka, Masato Motomura
IEEE Trans. Very Large Scale Integr. Syst.3
2011 20Gbps C-Based Complex Event Processing
abstract
This paper presents the world's fastest complex event processing system, designed to process a large number of events on FPGAs. Unlike conventional SQL-based approaches, our approach features logic automation constructed with a new C-based event language that supports regular expressions on the basis of C functions, so that a wide variety of event-processing applications can be efficiently mapped to FPGAs. Evaluations on an FPGA-based NIC show that we have achieved 20Gbps event processing performance in a financial trading application with only a small logic increase, 2.7%.
Hiroaki Inoue, Takashi Takenaka, Masato Motomura
FPL3
2011 Test compression for dynamically reconfigurable processors
abstract
We present the world's first test compression technique that features automation of compression rules for test time reduction on dynamically reconfigurable processors. Evaluations on an actual 40-nm product show that our technique achieves a 2.7 times compression ratio for original configuration information (better than does GZIP), the peak decompression bandwidth of 1.6 GB/s, and 2.7 times shorter test times.
Hiroaki Inoue, Junya Yamada, Hideyuki Yoneda, Katsumi Togawa, Masato Motomura, Koichiro Furuta
ACM Trans. Reconfigurable Technol. Syst.5
2004 Implementing and Evaluating Stream Applications on the Dynamically Reconfigurable Processor
abstract
Dynamically reconfigurable processor (DRP) developed by NEC electronics is a coarse grain reconfigurable processor that selects a data path from the on-chip repository of sixteen circuit configurations, or contexts, to implement different logic on one single DRP chip. Several stream applications have been implemented on DRP-1, the first prototype chip, and evaluation results are presented. By computing parallelly using the processing elements(PEs) and distributed memory modules, DRP-1 outperformed pentium III/4 and embedded CPU MIPS64 in some stream application examples. We also present programming techniques applicable on reconfigurable processors and discuss their feasibility in boosting system performance.
Noriaki Suzuki, Shunsuke Kurotaki, Masayasu Suzuki, Naoto Kaneko, Yutaka Yamada, Katsuaki Deguchi, Yohei Hasegawa, Hideharu Amano, Kenichiro Anjo, Masato Motomura, Kazutoshi Wakabayashi, Takeo Toi, Toru Awashima
FCCM10
2004 Stream applications on the dynamically reconfigurable processor
abstract
Dynamically reconfigurable processor (DRP) developed by NEC Electronics is a coarse grain reconfigurable processor that selects a data path from the on-chip repository of sixteen circuit configurations, or contexts, to implement different logic on one single DRP chip. Several stream applications have been implemented on the DRP-1, the first prototype chip, and evaluation results are presented. By pipelining the executions, DRP-1 outperformed Pentium III/4, embedded CPU MIPS64, and Texas Instruments DSP TMS320C67J3 in some stream application examples. We also present programming techniques applicable on dynamically reconfigurable processors and discuss their feasibility in boosting system performance.
Masayasu Suzuki, Yohei Hasegawa, Yutaka Yamada, Naoto Kaneko, Katsuaki Deguchi, Hideharu Amano, Kenichiro Anjo, Masato Motomura, Kazutoshi Wakabayashi, Takao Toi, Toru Awashima
FPT8
2004 A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable Systems
abstract
The increase in complexity of programmable hardware platforms results in the need to develop efficient high-level synthesis tools since that allows more efficient exploration of the design space while predicting the effects of technology specific tools on the design space. Much of the previous work, however, neglects the delay of interconnects (e.g. multiplexers) which can heavily influence the overall performance of the design. In addition, in the case of dynamic reconfigurable logic circuits, unless an appropriate design methodology is followed, an unnecessarily large number of configurable logic blocks may end up being used for communication between contexts, rather than for implementing function units. The aim of this paper is to present a new technique to perform interconnect-sensitive synthesis, targeting dynamic reconfigurable circuits. Further, the proposed technique exploits multiple hardware contexts to achieve efficient designs. Experimental results on several benchmarks, which have been done on our DRL LSI circuit [M. Meribout et al. [200]], [M. Meribout et al. (1997)], demonstrate that, by jointly optimizing the interconnect communication, and function unit cost, we can achieve higher quality designs than is possible with such previous techniques as Force-Directed-Scheduling.
Mahmoud Méribout, Masato Motomura
IEEE Trans. Computers2
2004 Efficient metrics and high-level synthesis for dynamically reconfigurable logic
abstract
The increase in complexity of programmable hardware platforms results in the need to develop efficient high-level synthesis (HLS) tools since it allows more efficient exploration of the design space while predicting the effects of technology specific tools on the design space. Much of the previous works however neglect the delay of interconnects (e.g. multiplexer) which can indeed contribute heavily on the overall performance of the design. In addition, in the case of dynamic reconfigurable logic (DRL) circuits, unless an appropriate design methodology is followed, large number of configurable logic blocks (CLBs) could be used for communication between contexts, rather than for implementing functional units (FUs). The aim of this paper is to present a new technique to perform interconnect-sensitive synthesis, targeting dynamic reconfigurable circuits. Further, the proposed technique exploits multiple hardware contexts to achieve efficient designs. Experimental results on several benchmarks, which have been done on our DRL LSI circuit (Meribout, 2000 and Motomura, 1997), demonstrate that by jointly optimizing the interconnect, communication, and function-unit cost, higher quality designs than other previous techniques (e.g. force-directed scheduling) can be achieved.
Mahmoud Méribout, Masato Motomura
IEEE Trans. Very Large Scale Integr. Syst.2
2003 New design methodology with efficient prediction of quality metrics for logic level design towards dynamic reconfigurable logic
Mahmoud Méribout, Masato Motomura
J. Syst. Archit.2
2000 Reconfigurable computing: its concept and a practical embodiment using newly developed dynamically reconfigurable logic (DRL) LSI: invited talk
abstract
This paper first outlines a broad range of reconfigurable computing research activities from a perspective of system LSI designs. Then, the paper focuses onto dynamically reconfigurable logic (DRL) LSI, a prototype chip that we developed to evaluate the reconfigurable computing concept. Through its ability to exchange hardware contexts quickly, this chip can accelerate media/communication applications with customized hardware configurations, yet maintaining scalability towards varying application sizes.
Masakazu Yamashina, Masato Motomura
ASP-DAC2
2000 A Virtual Hardware System on a Dynamically Reconfigurable Logic Device
abstract
WASMII is virtual hardware using a multi-context reconfigurable device with a data driven control. Since implementation of WASMII was infeasible due to the unavailability of such a device, the system has been only evaluated using an emulator so far. However, the first reconfigurable multi-context device called DRL has been developed by NEC. Making the use of its flexible reconfigurability, we have implemented a mechanism of WASMII on DRL.
Yuichiro Shibata, Masaki Uno, Hideharu Amano, Koichiro Furuta, Taro Fujii, Masato Motomura
FCCM6
2000 A Study of Channeled DRAM Memory Architectures
abstract
Channeled DRAM features small on-chip buffers called channels that are placed in front of the DRAM core. In this study various techniques to efficiently control the channels were investigated. Different techniques of caching and prefetching were adapted to the unique features of Channeled DRAM. An existing execution-driven processor simulator was extended by a memory simulation library and three benchmarks were run on four different memory system configurations of this simulator to evaluate the performance of the different control strategies. As a result, using Channeled DRAM as replacement for conventional SDRAM improves the memory system performance by reducing the average access latency up to 50%.
Lars Friebe, Yoshikazu Yabe, Masato Motomura
ICCD3
1998 An Embedded DRAM-FPGA Chip with Instantaneous Logic Reconfiguration
abstract
Reconfigurable computing is attracting wide attention as a novel general purpose computing paradigm for accelerating compute intensive and/or data-parallel applications, such as compression, encryption, searching, sorting, and image processing. A key enabling technology for a reconfigurable computer is in-system logic reconfiguration of SRAM-based FPGAs, through which its hardware architecture is dynamically customized for a specific task on demand. Quicker a reconfiguration is, more frequent the reconfigurations can become: i.e., a reconfigurable computer can adapt to applications which have more dynamic behavior. A whole-chip reconfiguration in conventional FPGAs, however, takes at least 100/spl mu/s. With this long latency, a reconfigurable computer is adaptable only to static applications, substantially losing the general-purposeness of the original concept. Integrating a DRAM with an FPGA can become an ideal solution to this problem. The on-chip DRAM can store hundreds of configuration programs, and the logic reconfiguration can get extremely faster by context-switching among the programs utilizing huge bandwidth internal to the DRAM core. Being driven by this observation, we have conducted prototype design of an embedded DRAM-FPGA chip.
Masato Motomura, Yoshiharu Aimoto, Atsufkni Shibayama, Yoshikazu Yabe, Masakazu Yamashina
FCCM1
1995 Ordered multithreading: a novel technique for exploiting thread-level parallelism
Masato Motomura, Toshiaki Inoue, Sunao Torii, Akihiko Konagaya
PACT1