Shinya Takamaeda-Yamazaki

dblp:91/10760 · also Shinya Takamaeda · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-3441-1695ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Relational Hoare Logic for High-Level Synthesis of Hardware Accelerators
Izumi Tanaka, Ken Sakayori, Shinya Takamaeda-Yamazaki, Naoki Kobayashi 0001
ESOP (2)3
2026 Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal Coding
abstract
Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in a broad range of data-intensive workloads from databases to machine learning. Due to its low computational intensity, the execution of this operation tends to be memory-bound, especially for large vectors, thereby limiting the utilization of compute resources. Processing-using-DRAM (PuD) is an emerging computing paradigm that performs massively parallel bitwise operations directly within the DRAM array, alleviating off-chip data movement. Unfortunately, no prior work proposes an efficient PuD-based solution tailored to vector-scalar comparisons. Existing PuD-based approaches require many DRAM commands because the comparison's algorithmic complexity grows with operand bit-width in the bit-serial execution model, which is inherently induced by current PuD architectures. As a result, this command overhead becomes the dominant performance bottleneck, limiting application-level speed up. We propose Clutch, a novel data representation and comparison algorithm for accelerating vector-scalar comparisons in PuD systems with high efficiency and scalability. Our key idea is twofold. First, to reduce the number of DRAM commands required for comparison, Clutch adopts temporal coding for vectors, where each value is encoded as a sequence of leading ones. This enables lookup-based comparisons, where comparing against a scalar input simply involves accessing the corresponding DRAM row. Second, Clutch leverages our key insight that a divide-and-conquer approach enables scalable lookup-based comparisons without incurring a prohibitive memory footprint at high bit-precision. Specifically, Clutch partitions the operand into multiple multi-bit chunks which can be compared independently using compact lookup tables, and merges per-chunk results through a procedure designed to execute efficiently on PuD.Clutch provides a flexible tradeoff between throughput and memory usage by adjusting chunk count. Experimental results on two applications, predicate evaluation and decision tree inference, demonstrate that Clutch improves end-to-end application throughput (and energy efficiency) by an average of 12 × (69 ×) over highly-optimized CPU and GPU execution and 2.9 × (3.0 ×) over the state-of-the-art bit-serial PuD implementation. Notably, we present, to our knowledge, the first mapping of decision tree inference to PuD execution, extending PuD to a new application domain. Our results demonstrate that DRAM can serve as a high-performance and energy-efficient computing substrate for comparison-intensive workloads.
Daichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, A. Giray Yaglikçi, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki
ICS11
2026 PuDghost: Experimental Analysis of Computation Result Corruption in Processing-Using-Dram Operations on Real Dram Chips and Implications for Future Systems
Daichi Tokuda, Ismail Emir Yuksel, Tatsuya Kubo, Ataberk Olgun, Haocong Luo, Nisa Bostanci, Jikun Wang, A. Giray Yaglikçi, Shinya Takamaeda-Yamazaki, Onur Mutlu
ISCA9
2025 Scalable Moment Propagation and Analysis of Variational Distributions for Practical Bayesian Deep Learning
abstract
Bayesian deep learning is one of the key frameworks employed in handling predictive uncertainty. Variational inference (VI), an extensively used inference method, derives the predictive distributions by Monte Carlo (MC) sampling. The drawback of MC sampling is its extremely high computational cost compared to that of ordinary deep learning. In contrast, the moment propagation (MP)-based approach propagates the output moments of each layer to derive predictive distributions instead of MC sampling. Because of this computational property, it is expected to realize faster inference than MC-based approaches. However, the applicability of the MP-based method in deep models has not been explored sufficiently, even though some studies have demonstrated the effectiveness of MP only in small toy models. One of the reasons is that it is difficult to train deep models by MP because of the large variance in activations. To realize MP in deep models, some normalization layers are required but have not yet been studied. In addition, it is still difficult to design well-calibrated MP-based models, because the effectiveness of MP-based methods under various variational distributions has also not been investigated. In this study, we propose a fast and reliable MP-based Bayesian deep-learning method. First, to train deep-learning models using MP, we introduce a batch normalization layer extended to random variables to prevent increases in the variance of activations. Second, to identify the appropriate variational distribution in MP, we investigate the treatment of moments of several variational distributions and evaluate their uncertainty quality of predictions. Experiments with regression tasks demonstrate that the MP-based method provides qualitatively and quantitatively equivalent predictive performance to MC-based methods regardless of variational distributions. In the classification tasks, we show that we can train MP-based deep models by extended batch normalization. We also show that the MP-based approach realizes 2.0-2.8 times faster inference than the MC-based approach while maintaining the predictive performance. The results of this study can help realize a fast and well-calibrated uncertainty estimation method that can be deployed in a wider range of reliability-aware applications.
Yuki Hirayama, Shinya Takamaeda-Yamazaki
IEEE Trans. Neural Networks Learn. Syst.2
2025 DF-BETA: An FPGA-based Memory Locality Aware Decision Forest Accelerator via Bit-Level Early Termination
abstract
Decision forests, particularly Gradient Boosting Decision Trees (GBDT), are popular due to their high prediction performance and computational efficiency, making them suitable for embedded systems with circuit size and available energy constraints. In this study, we propose a new lightweight GBDT inference acceleration mechanism through the hardware and algorithm co-design. First, we present LoADPack, a hardware-friendly GBDT algorithm that enhances memory access locality. LoADPack obtains trees where the features and thresholds used across the entire ensemble are regular regardless of a branching direction by unifying some nodes and aligning the memory access patterns. Furthermore, we present DF-BETA, a resource-efficient accelerator for the LoADPack algorithm. DF-BETA utilizes MSB-first bit-serial computation to enable early determination of comparison calculations of 32-bit floating-point numbers, optimizing the operation for determining a branch direction. The hardware complexity and computation termination speed vary with the granularity of bit-serial computation. Therefore, we conduct design space exploration of DF-BETA to identify the optimal configuration. Our findings reveal that using 4-bit-serial comparators minimizes circuit size while achieving the leading throughput. Compared to running unconstrained GBDT on a typical accelerator with 32-bit bit-parallel comparators, our accelerator achieves 1.6 times higher throughput on average while maintaining comparable accuracy.
Daichi Tokuda, Shinya Takamaeda-Yamazaki
ACM Trans. Reconfigurable Technol. Syst.2
2024 OSA-HCIM: On-The-Fly Saliency-Aware Hybrid SRAM CIM with Dynamic Precision Configuration
abstract
Computing-in-Memory (CIM) has shown great potential for enhancing efficiency and performance for deep neural networks (DNNs). However, the lack of flexibility in CIM leads to an unnecessary expenditure of computational resources on less critical operations, and a diminished Signal-to-Noise Ratio (SNR) when handling more complex tasks, significantly hindering the overall performance. Hence, we focus on the integration of CIM with Saliency-Aware Computing—a paradigm that dynamically tailors computing precision based on the importance of each input. We propose On-the-fly Saliency-Aware Hybrid CIM (OSA-HCIM) offering three primary contributions: (1) On-the-fly Saliency-Aware (OSA) precision configuration scheme, which dynamically sets the precision of each multiply-and-accumulate (MAC) operation based on its saliency, (2) Hybrid CIM Array (HCIMA), which enables simultaneous operation of digital-domain CIM (DCIM) and analog-domain CIM (ACIM) via split-port 6T SRAM, and (3) an integrated framework combining OSA and HCIMA to fulfill diverse accuracy and power demands.Implemented on a 65nm CMOS process, OSA-HCIM demon-strates an exceptional balance between accuracy and resource utilization. Notably, it is the first CIM design to incorporate a dynamic digital-to-analog boundary, providing unprecedented flexibility for saliency-aware computing. OSA-HCIM achieves a 1. 95x enhancement in energy efficiency, while maintaining minimal accuracy loss compared to DCIM when tested on CIFAR100 dataset.
Yung-Chin Chen, Shimpei Ando, Daichi Fujiki, Shinya Takamaeda-Yamazaki, Kentaro Yoshioka
ASPDAC4
2024 FS-Boost: Communication-Efficient Federated Subtree-Based Gradient Boosting Decision Trees
abstract
Federated learning (FL) is a secure and distributed machine learning method in which clients learn cooperatively without disclosing private data to others. some decision tree-based FL have been proposed that employ gradient boosting decision trees (GBDT). However, previously proposed GBDT-based FL methods require sharing the entire decision tree, including the tree structure and leaf weights, for each synchronization step in training. This process inevitably results in extensive communication. To solve this problem, we propose FS-Boost-Federated Subtree-based GBDT, a horizontal FL method that reduces the communication cost. Suppose the maximum depth of the decision tree is set to d. In that case, a conventional GBDT trains only a single decision tree of depth d within one training round. In contrast, FS- Boost utilizes subtrees from depth 1 to d-1 generated in the learning process in addition to the whole trees. Sharing still-growing subtrees with other clients reduces the total amount of communication cost through accelerating model convergence. Our experiment results indicate that FS-Boost significantly reduced the communication cost by at least half in most cases while maintaining the accuracy.
Kotaro Shimamura, Shinya Takamaeda-Yamazaki
CCNC2
2024 PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
abstract
Approximate computing emerges as a promising approach to enhance the efficiency of compute-in-memory (CiM) systems in deep neural network processing. However, traditional approximate techniques often significantly trade off accuracy for power efficiency, and fail to reduce data transfer between main memory and CiM banks, which dominates power consumption. This paper introduces a novel probabilistic approximate computation (PAC) method that leverages statistical techniques to approximate multiply-and-accumulation (MAC) operations, reducing approximation error by 4× compared to existing approaches. PAC enables efficient sparsity-based computation in CiM systems by simplifying complex MAC vector computations into scalar calculations. Moreover, PAC enables sparsity encoding and eliminates the LSB activations transmission, significantly reducing data reads and writes. This sets PAC apart from traditional approximate computing techniques, minimizing not only computation power but also memory accesses by 50%, thereby boosting system-level efficiency. We developed PACiM, a sparsity-centric architecture that fully exploits sparsity to reduce bit-serial cycles by 81% and achieves a peak 8b/8b efficiency of 14.63 TOPS/W in 65 nm CMOS while maintaining high accuracy of 93.85/72.36/66.02% on CIFAR-10/CIFAR-100/ImageNet benchmarks using a ResNet-18 model, demonstrating the effectiveness of our PAC methodology. Software simulation framework is available at GitHub.
Wenlun Zhang, Shimpei Ando, Yung-Chin Chen, Satomi Miyagi, Shinya Takamaeda-Yamazaki, Kentaro Yoshioka
ICCAD5
2023 SPinS-FL: Communication-Efficient Federated Subnetwork Learning
abstract
Federated learning (FL) is a distributed machine learning method in which edge devices collaboratively train a unified model without disclosing their private training data to others. Unlike data centers, edge devices often stand in low-bandwidth traffic environments, which can be a bottleneck in the training process. To tackle this problem, we first consider a naive FL based on the Edge-Popup algorithm. The Edge-Popup-based neural network constructs a subnetwork by assigning a score to each randomly initialized weight and selecting weights based on the score without updating the weights. Forward propagation, based on the subnetwork consisting of only the live weights with high scores, is performed, and then backpropagation is performed to update scores and search for a subnetwork that achieves high accuracy. This algorithm has two unique properties: first, the score itself is unnecessary for calculating the score's gradient, as long as there is a supermask, a bitmask indicating which weights are participating in the subnetwork then; and second, the intensity of the turnover in the score rankings varies according to the scores' rank range. Exploiting these properties, we propose SPinS-FL, a novel FL method based on the Edge-Popup algorithm. Instead of communicating all the scores with the server, SPinS-FL only communicates some supermasks and a partial set of scores of the weights around the boundary between selected and unselected in the subnetwork; this dramatically improves communication efficiency. Numerical experiments using the CIFAR-10 and CIFAR-100 datasets show that SPinS-FL succeeds in reducing the traffic volume by up to 79.8% compared with communicating all the scores. Moreover, SPinS-FL achieves comparable accuracy as the federated averaging method but with up to 69.9% less traffic cost.
Masayoshi Tsutsui, Shinya Takamaeda-Yamazaki
CCNC2
2022 FADEC: FPGA-based Acceleration of Video Depth Estimation by HW/SW Co-design
abstract
3D reconstruction from videos has become increasingly popular for various applications, including navigation for autonomous driving of robots and drones, augmented reality (AR), and 3D modeling. This task often combines traditional image/video processing algorithms and deep neural networks (DNNs). Although recent developments in deep learning have improved the accuracy of the task, the large number of cal-culations involved results in low computation speed and high power consumption. Although there are various domain-specific hardware accelerators for DNNs, it is not easy to accelerate the entire process of applications that alternate between traditional image/video processing algorithms and DNNs. Thus, FPGA-based end-to-end acceleration is required for such complicated applications in low-power embedded environments. This paper proposes a novel FPGA-based accelerator for DeepVideoMVS, which is a DNN-based depth estimation method for 3D reconstruction. We employ HW/SW co-design to appropriately utilize heterogeneous components in modern SoC FPGAs, such as programmable logic (PL) and CPU, according to the inherent characteristics of the method. As some operations are unsuitable for hardware implementation, we determine the operations to be implemented in software through analyzing the number of times each operation is performed and its memory access pattern, and then considering comprehensive aspects: the ease of hardware implementation and degree of expected acceleration by hardware. The hardware and software implementations are executed in parallel on the PL and CPU to hide their execution latencies. The proposed accelerator was developed on a Xilinx ZCUI04 board by using NNgen, an open-source high-level synthesis (HLS) tool. Experiments showed that the proposed accelerator operates 60.2 times faster than the software-only implementation on the same FPGA board with minimal accuracy degradation. Code available: https://github.com/casys-utokyo/fadec/
Nobuho Hashimoto, Shinya Takamaeda-Yamazaki
FPT2
2022 Model-based Federated Reinforcement Distillation
abstract
Reinforcement learning (RL) is a framework for learning highly rewarding policies through interactions with the environment. The more the agent knows about the environment, the more easily it learns. Therefore, exploration is often performed using multiple agents. However, information gathered by edge devices is not always available to all devices, including the server. Federated learning (FL) is a framework for collaborative learning while preserving the privacy of the training data. Most of the FL methods exchange the trained model parameters instead of the private data. In contrast, some FL methods incorporate knowledge distillation. These methods are known to be superior to parameter exchange methods in terms of communication efficiency. In this study, we apply the idea of distillation-based FL to RL problems. We highlight some difficulties unique to RL and give a solution to them by introducing an environment model. We call this newly proposed method model-based federated reinforcement distillation (Model-based FRD). As with existing distillation-based FL methods, our proposed method achieves high communication efficiency. Our experimental results show that our method reduces communication costs by approximately 821 times.
Sefutsu Ryu, Shinya Takamaeda-Yamazaki
GLOBECOM2
2022 Real-Time Tone Mapping: A Survey and Cross-Implementation Hardware Benchmark
abstract
The rising demand for high quality display has ensued active research in high dynamic range (HDR) imaging, which has the potential to replace the standard dynamic range imaging. This is due to HDR’s features like accurate reproducibility of a scene with its entire spectrum of visible lighting and color depth. But this capability comes with expensive capture, display, storage and distribution resource requirements. Also, display of HDR images/video content on an ordinary display device with limited dynamic range requires some form of adaptation. Many adaptation algorithms, widely known as tone mapping (TM) operators, have been studied and proposed in the last few decades. In this article, we present a comprehensive survey of 60 TM algorithms that have been implemented on hardware for acceleration and real-time performance. In this state-of-the-art survey, we will discuss those TM algorithms which have been implemented on GPU, FPGA, and ASIC in terms of their hardware specifications and performance. Output image quality is an important metric for TM algorithms. From our literature survey we found that, various objective quality metrics have been used to demonstrate the quality of those algorithms hardware implementation. We have compiled those metrics used in this survey, and analyzed the relationship between hardware cost, image quality and computational efficiency. Currently, machine learning-based (ML) algorithms have become an important tool to solve many image processing tasks, and this article concludes with a discussion on the future research directions to realize ML-based TM operators on hardware.
Yafei Ou, Prasoon Ambalathankandy, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai, Masayuki Ikebe
IEEE Trans. Circuits Syst. Video Technol.3
2021 ASBNN: Acceleration of Bayesian Convolutional Neural Networks by Algorithm-hardware Co-design
abstract
Bayesian Convolutional Neural Networks (BCNNs) have been proposed to address the problem of model uncertainty in conventional neural networks. By treating weights as distributions rather than deterministic values, BCNNs mitigate the problem of overfitting, training with a small amount of data, and uncertainty evaluations. However, computing the distributions of BCNN outputs is time- and energy-consuming because it requires computing multiple forward passes.To address this computational problem, we propose a novel algorithm-hardware co-design approach with an approximation algorithm and hardware support for the rapid computation of BCNN. Our observations of the absolute number of each layer’s input and the input difference among multiple forward passes show that most of these values are significantly small compared with other large values. Our algorithm treats these small values as zero and makes them sparser. The extracted sparsity allows us to skip most multiplications. As a result, it achieves a computation reduction of 81.1 % in classification tasks and 77.7 % in regression tasks. Additionally, to support the algorithm-level approximation on hardware, we propose a novel dataflow that is specialized for our algorithm, and develop a new accelerator architecture, accelerator for sparse Bayesian Neural Networks (ASBNN), that can handle sparsity extracted by the algorithm. Our evaluation demonstrates that the ASBNN successfully exploits the algorithmic computation reduction to improve the computation time by 3.3× and energy efficiency by 3.7× compared with the naive implementation of dense BCNN accelerators.
Yoshiki Fujiwara, Shinya Takamaeda-Yamazaki
ASAP2
2021 An FPGA-Based Fully Pipelined Bilateral Grid for Real-Time Image Denoising
abstract
The bilateral filter (BF) is widely used in image processing because it can perform denoising while preserving edges. It has disadvantages in that it is nonlinear, and its computational complexity and hardware resources are directly proportional to its window size. Thus far, several approximation methods and hardware implementations have been proposed to solve these problems. However, processing large-scale and high-resolution images in real time under severe hardware resource constraints remains a challenge. This paper proposes a real-time image denoising system that uses an FPGA based on the bilateral grid (BG). In the BG, a 2D image consisting of x- and y-axes is projected onto a 3D space called a “grid,” which consists of axes that correlate to the x-component, y-component, and intensity value of the input image. This grid is then blurred using the Gaussian filter, and the output image is generated by interpolating the grid. Although it is possible to change the window size in the BF, it is impossible to change it on the input image in the BG. This makes it difficult to associate the BG with the BF and to obtain the property of suppressing the increase in hardware resources when the window radius is enlarged. This study demonstrates that a BG with a variable-sized window can be realized by introducing the window radius parameter wherein the window radius on the grid is always 1. We then implement this BG on an FPGA in a fully pipelined manner. Further, we verify that our design suppresses the increase in hardware resources even when the window size is enlarged and outperforms the existing designs in terms of computation speed and hardware resources.
Nobuho Hashimoto, Shinya Takamaeda-Yamazaki
FPL2
2021 A 96-MB 3D-Stacked SRAM Using Inductive Coupling With 0.4-V Transmitter, Termination Scheme and 12: 1 SerDes in 40-nm CMOS
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 35% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates for the degradation of the first transmitted bit, a coil termination scheme that aims to eliminate the ringing of 3D inductive coupling bus, and a 12:1 SerDes that minimizes power consumption and area overhead in inductive coupling channels. Low-power, large-capacity, 3-cycle latency 3D-stacked SRAM for a DNN accelerator is achieved with the combination of these techniques to serve as a replacement of 3D-stacked DRAM. The performance of the proposed 3D-SRAM is compared with HBM DRAM and achieves more than 50% lower energy consumption. The scaling scenario of the SRAM module is discussed in light of the scaling of the inductive coupling technology and logic process.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 A 3D-Stacked SRAM using Inductive Coupling with Low-Voltage Transmitter and 12: 1 SerDes
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 45% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates the degradation of the first signal, a coil termination scheme that aims to eliminate the noise of 3D inductive coupling bus, and a 12:1 SerDes. The data density of the SRAM should reach 12.3-MB/mm3, which extends beyond that of state-of-the-art stacked DRAMs.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Kota Ando, Kazutoshi Hirose, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
ISCAS6
2020 An Adaptive Global and Local Tone Mapping Algorithm Implemented on FPGA
abstract
We present a fast global and locally adaptive tone mapping algorithm and its field-programmable gate array (FPGA) implementation. The specially designed tone mapping function, which is based on local histogram equalization, controls global, and local characteristics individually. In contrast to other tonemap operators, our algorithm manages light/dark halos separately and by using local tonemap function alone, it can effectively suppress noise. We validated the effectiveness of our algorithms using subjective and objective assessment. Using an average of the bins, we achieve fast smoothed local histogram estimation with fewer bins while maintaining high accuracy. Our new implementation method requires minimal data access and reduced memory as it operates with a downscaled frame size of 240 × 135 pixels. Relative local area size is 248 × 248 @Full-HD resolution (1920 × 1080). For low-latency pixel output, the system performs the tone mapping using pixel information from the previous frame. When we implemented the system on FPGA (TB-7K-325TIMG and Xilinx Kintex-7), we achieved lightweight hardware as the total usage rate is about 25% of the available FPGA resource. Using an online 1080p video we demonstrate, a real-time video processing using our hardware tone mapping system.
Prasoon Ambalathankandy, Masayuki Ikebe, Takayuki Yoshida, Takeshi Shimada, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai
IEEE Trans. Circuits Syst. Video Technol.5
2019 DeltaNet: Differential Binary Neural Network
abstract
Energy-constrained neural network processing is in high demanded for various mobile applications. Binarized neural network (BNN) aggressively enhances the computational efficiency, and in contrast, it suffers from degradation of accuracy due to its extreme approximation. We propose a neural network model using a new activation function "Delta" based on binarization of differences between weighted-sums. The "Delta" retains the magnitude relation between numerical values, and conveys richer information than ordinary binarization. We can design the hardware architecture for the proposed model with almost the same elements as BNN. The evaluation shows that it achieves higher recognition accuracy than a conventional BNN with almost the same hardware configuration.
Yuka Oba, Kota Ando, Tetsuya Asai, Masato Motomura, Shinya Takamaeda-Yamazaki
ASAP5
2018 Dither NN: An Accurate Neural Network with Dithering for Low Bit-Precision Hardware
abstract
Energy-constrained neural network processing is in high demanded for various mobile applications. Binary neural network aggressively enhances the computational efficiency, and in contrast, it suffers from degradation of accuracy due to its extreme approximation. We propose a novel accurate neural network model based on binarization and "dithering" that distributes the quantization error to neighboring pixels. The quantization errors in the binarization are distributed in the plane, so that a pixel in the multi-level source expression more accurately represented in the resulting binarized plane by multiple pixels. We designed a low-overhead binary-based hardware architecture for the proposed model. The evaluation results show that this method can be realized with a few additional lightweight hardware components.
Kota Ando, Kodai Ueyoshi, Yuka Oba, Kazutoshi Hirose, Ryota Uematsu, Takumi Kudo, Masayuki Ikebe, Tetsuya Asai, Shinya Takamaeda-Yamazaki, Masato Motomura
FPT9
2018 Sparse Disparity Estimation Using Global Phase Only Correlation for Stereo Matching Acceleration
abstract
In this study, we propose an efficient stereo matching method which estimates sparse disparities using global phase only correlation (POC). Conventionally, cost functions are to be calculated for all disparity candidates and the associated computational cost has been impediment in achieving a realtime performance. Therefore, we consider to use full image 2D phase only correlation (FIPOC) for detecting the valid disparities candidates. This would require comparatively fewer calculations for the same number of disparities. Since, the FIPOC output indicates the disparity distribution of two stereo images, we can sort the disparity candidates and choose them for sparse calculation. In our proposed method, the searchable disparity range is half of the input image size, which is much wider than that of the conventional methods. When we apply the FIPOC to naive sum of absolute difference (SAD) stereo matching method, the combined algorithm would require fewer calculations while maintaining the same accuracy. In our evaluation, the proposed method achieves 194 disparity stereo matching in 70 ms on$398 \times 288$images without the need for SIMD instruction, multi-thread operation, or additional hardware while using a Intel Core i5-5257U.
Takeshi Shimada, Masayuki Ikebe, Prasoon Ambalathankandy, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai
ICASSP4
2018 Analysis of Smoothed LHE Methods for Processing Images with Optical Illusions
abstract
To replicate human visual perception, we analyze processing images with optical illusion using edge preserving filters and smoothed local histogram equalization (LHE). Images with the optical illusions are good models for gradual/rapid changes in contrast and strong edges, which are good cases for assessing the robustness of image filters. Here, we study and analyze the performance of smoothed LHE filters while processing perceptual illusion. Our studies conclude that, smoothed LHEs are useful in retaining actual edge forms in these images as they can operate using large kernel sizes. These large kernel size filters can construct sawtooth like edge and it corresponds to adequately wide halos. We also demonstrate the usefulness of smoothed LHE like tone mapping techniques in preserving naturalness, and we confirmed it by performing subjective visual test.
Prasoon Ambalathankandy, Takeshi Shimada, Shinya Takamaeda-Yamazaki, Masato Motomura, Tetsuya Asai, Masayuki Ikebe
VCIP3
2017 CPRring: A Structure-Aware Ring-Based Checkpointing Architecture for FPGA Computing
abstract
In this paper, we present a new architecture for FPGA checkpointing along with an efficient mechanism. We then provide a static analysis of original HDL source code to reduce the cost of hardware for checkpointing functionality. Our evaluations show that with the proposals, checkpointing hardware causes small degradation in maximum clock frequency (less than10%). The LUT overhead varies from 14.4% (Dijkstra) to 103.84%(Matrix Multiplication).
Hoang Gia Vu, Shinya Takamaeda-Yamazaki, Takashi Nakada, Yasuhiko Nakashima
FCCM2
2017 FPGA implementation of edge-guided pattern generation for motion-vector estimation of textureless objects
abstract
The widely accepted block-matching technique, which is required to identify motion vectors, fails in cases in which texture is not existent. In [1], we proposed a hardware-oriented cellular-automaton algorithm that generates spatial patterns on textureless objects and backgrounds, aiming at motion-vector estimation of textureless moving objects. This demonstration presents a field-programmable gate array (FPGA) system that supports real-time processing. This system provides motion-vectors in moving textureless objects and enables enhanced processing of motion vector classification.
Aoi Tanibata, Alexandre Schmid, Shinya Takamaeda-Yamazaki, Masayuki Ikebe, Masato Motomura, Tetsuya Asai
FPL3
2014 Ultrasmall: The smallest MIPS soft processor
abstract
Soft processors have been commonly used in FPGAbased designs to perform various useful functions. Some of these functions are not performance-critical and required to be implemented using very few FPGA resources. For such cases, it is desired to reduce circuit area of the soft processor as much as possible. This paper proposes Ultrasmall, a small soft processor for FPGAs. Ultrasmall supports a subset of the MIPS-I ISA and is designed for microcontrollers in FPGA-based SoCs. Ultrasmall employs an area efficient architecture to minimize the use of FPGA resources. While supporting the 32-bit ISA, Ultrasmall adopts the 2-bit wide serial ALU architecture. This approach significantly reduces the amount of FPGA resource usage. In addition to the device-independent optimizations for any FPGAs, we apply primitives-based optimizations for the Xilinx Spartan-3E FPGA series with 4-input LUTs, thereby further reducing the total number of occupied slices. The evaluation result shows that, on the Xilinx Spartan-3E XC3S500E FPGA, Ultrasmall occupies only 137 slices which is 84% of the number of occupied slices of Supersmall, a very small soft processor with the same design concept as Ultrasmall. On the other hand, in term of performance, Ultrasmall is 2.9× faster than Supersmall.
Hiroshi Nakatsuka, Yuichiro Tanaka, Thiem Van Chu, Shinya Takamaeda-Yamazaki, Kenji Kise
FPL4
2014 flipSyrup: Cycle-accurate hardware simulation framework on abstract FPGA platforms
abstract
FPGA-based rapid prototyping is widely applied for fast simulations of hardware structure verifications. In this paper, we propose flipSyrup, a prototyping framework for cycle-accurate hardware simulations on abstract FPGA platforms. In order to mitigate the development complexity of FPGA-based simulators, the framework provides two abstractions of resources on FPGA platforms: Memory systems and inter-FPGA interconnections on multi-FPGA platforms. The framework enables designers to draw up a target hardware using abstract interfaces as ideal memory systems and interconnections on FPGA platforms. Our evaluation result shows that the slowdowns in simulation speed under the abstractions by using the framework are not critical.
Shinya Takamaeda-Yamazaki, Kenji Kise
FPL1