Hakaru Tamukoh

dblp:54/5674 · DBLP profile ↗
← Back
49ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0002-3669-1371ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 4 first-author · 11 since 2021Systems, architecture and hardware · 20 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards Optimizing Physical Reservoir Computing: A Hybrid PCPO-ESN Framework for Time Series Forecasting
abstract
In neuroscience-inspired computing, reservoir computing (RC) has emerged as a powerful model that leverages recurrent dynamics to mimic neural systems. Echo State Networks (ESNs), a prominent RC framework, are effective for modeling temporal data. However, physical RC using pulse-coupled phase oscillator (PCPO) networks, while promising for hardware systems, often struggle with flexibility, limiting their performance in time series forecasting. To address these challenges, this paper introduces a novel hybrid RC framework that combines the both states from ESNs and PCPO using two oscillators towards optimization physical RC. This hybrid-state using the oscillating events and enhance the representational capacity of the reservoir, thereby enabling the model to learn pertinent features for more accurate predictions. We evaluate our framework using the synthetic NARMA-10 dataset and the real-world Numenta Anomaly Benchmark (NAB). Experimental results show that the proposed method achieves lower mean squared error (MSE) values 7.19×10−3on NARMA-10 and 33.60×10−3on NAB, which outperforming the baseline ESN model.
Dinda Pramanta, Ninnart Fuengfusin, Arie Rachmad Syulistyo, Hakaru Tamukoh
ISCAS4
2025 An in-memory computing circuit with carry-over thermometer coding for a hippocampus-inspired model
abstract
The memory-based hippocampus-inspired model (MBHIM) that has been proposed by the authors is a model that reproduces episodic memory suitable for VLSI implementation. While an MBHIM VLSI circuit can achieve high energy efficiency using the in-memory computing architecture with thermometer coding, its computational functionality was limited to addition operation. The use of thermometer coding, though enabling simple VLSI implementation, remains an integration density challenge. In this paper, we propose a memory cell circuit for MBHIM, which supports addition and subtraction using in-memory computing, enabling broader applications, including time-dependent forgetting mechanisms. We also propose a carry-over thermometer coding scheme for efficient value representation, achieving a reduction of 60-90% transistor count compared to conventional approaches. This scheme maintains the advantages of thermometer coding while achieving high integration density, making it particularly effective for memory devices with high write error rates, such as MRAM and ReRAM. The proposed scheme is promising for realizing brain-inspired AI hardware that achieves low power consumption, high computational efficiency, and high integration density.
Yuka Shishido, Tomoro Marcus Jones, Hakaru Tamukoh, Takashi Morie
ISCAS3
2025 Reservoir Computing with VCO-Based Spiking Neurons for Regression and Classification
abstract
Reservoir computing (RC) significantly reduces the requirement on hardware and training resources, making it suitable for edge-computing applications. This work proposes using voltage-controlled oscillator (VCO)-based spiking neurons for RC to leverage the intrinsic randomness and variability of the neuron circuit for low-power operations. We describe the underlying circuit design and propose a network architecture based on the spiking neuron for RC. We demonstrate the effectiveness of the proposed RC network using VCO-based spiking neurons through benchmark tasks and electrocardiogram (ECG) classification.
Kanta Yoshioka, Parker Allred, Taylor Barton, Bibhu Datta Sahoo 0003, Yen-Cheng Kuan, Shiuh-Hua Wood Chiang, Hakaru Tamukoh
ISCAS7
2024 Robust Binary Encoding for Ternary Neural Networks Toward Deployment on Emerging Memory
abstract
Deep neural networks (DNNs) have enabled state-of-the-art performance across various applications. However, their deployment is often hindered by high energy demands. One solution is to deploy DNNs on hardware equipped with emerging non-volatile memory, which does not require energy to maintain memory state. Nonetheless, this approach may introduce bit-flips in DNN parameters, leading to a drop in model performance. To mitigate this issue, ternary neural networks (TNNs), which are known for their robustness against bit-flips, can be utilized. To accelerate TNNs, it is necessary to define binary representations (BRs) for ternary values to enable bit-wise operations. This paper proposes a framework for identifying a set of BRs that minimizes the discrepancy between the Top-1 Accuracy of TNNs before and after bit-flips, thereby enhancing their robustness. The BRs are evaluated using randomly generated data, the CIFAR-10 dataset, and the ImageNet 2012 dataset. The experimental results demonstrate that our BRs can significantly reduce or maintain the Top-1 Accuracy degradation caused by bit-flips, compared to conventional BRs.
Ninnart Fuengfusin, Hakaru Tamukoh, Osamu Nomura, Takashi Morie
IJCNN2
2024 A Hippocampus-Inspired Environment-Specific Knowledge Acquisition System Utilizing Common Knowledge with Contextual Information
abstract
Home service robots acquire environment-specific knowledge through experiences (episodes) in a home to perform tasks autonomously. We propose a hippocampus-inspired memory system comprising an environment-specific knowledge module that handles episodic memory and a common knowledge module. The robot works in a home and needs to learn the locations of objects in the home with minimal user help to deliver or store objects. We employ a large language model (LLM) that runs on an edge device as the common knowledge module to protect the privacy of users. The environment-specific knowledge module acquires episodes and generates contextual information. The LLM receives contextual information generated by the environment-specific knowledge module, with which it infers an appropriate possible location. We verified that the LLM could infer the possible object locations and that the proposed system could minimize the required user effort.
Akinobu Mizutani, Yuichiro Tanaka, Hakaru Tamukoh, Osamu Nomura, Katsumi Tateno, Takashi Morie
IJCNN3
2024 CMOS digital-analog mixed signal VLSI implementation of a hippocampus-inspired model
abstract
For artificial intelligence (AI) to be useful in the home, it is required to acquire unique knowledge of the home obtained through interaction with space and environment. This is difficult for deep learning-based AI. The human brain can learn unique knowledge from few experiences. The entorhinal cortex and hippocampus play essential roles for episodic memory formation and recall. While entorhinal-hippocampal models have been proposed that can reproduce episodic memory, hardware systems that implement such models face challenges related to high computational complexity, power consumption, and processing speed. In this paper, we propose a digital-analog mixed-signal CMOS VLSI implementation of a hippocampus-inspired model that can memorize and associate place and object information essential for the formation of episodic memory. By using both analog and digital in-memory computing architecture, the proposed circuit has achieved a computational efficiency of 22 TOPS/W, which is very high for AI hardware with a learning function. The proposed circuit was fabricated, measured, and evaluated. The results of an experiment using a fabricated chip and a control system showed that the proposed circuit can memorize and process place and object information, and can acquire environment-unique knowledge through interaction with a space.
Yuka Shishido, Osamu Nomura, Katsumi Tateno, Hakaru Tamukoh, Takashi Morie
IJCNN4
2024 Enhancing Memory Capacity of Reservoir Computing with Delayed Input and Efficient Hardware Implementation with Shift Registers
abstract
To use reservoir computing (RC) for practical tasks, both a high memory capacity and nonlinearity are required; however, some RC models have the problem of a low memory capacity. We propose a delay mechanism for increasing the memory capacity in RC as well as a simple and small-scale digital circuit for implementing the delay mechanism. The proposed delay mechanism is integrated into the input layer of the RC model and is expected to be implemented in several RC models, such as material reservoirs and chaotic Boltzmann machine (CBM)-RC. We conducted experiments using a CBM-RC with a delay mechanism (CBM-RC-DL) and evaluated the performance improvement achieved by introducing a delay mechanism. We used CBM-RC as the base model because it is an appropriate model for the hardware implementation of large networks but has a low memory capacity. The experimental results for CBM-RC-DL indicated that the delay mechanism significantly increased the memory capacity of CBM-RC with the addition of a small-scale circuit. Furthermore, the entire synthesized CBM-RC-DL was sufficiently small-scale to be implemented in a field-programmable gate array for edge computing, and it outperformed conventional methods in nonlinear autoregressive moving average 10 (NARMA10)—a benchmark task for time-series data processing. The proposed delay mechanism can facilitate the use of many RC models because of its simple structure.
Soshi Hirayae, Kanta Yoshioka, Atsuki Yokota, Ichiro Kawashima, Yuichiro Tanaka, Yuichi Katori, Osamu Nomura, Takashi Morie, Hakaru Tamukoh
ISCAS9
2024 FPGA Implementation for Large Scale Reservoir Computing based on Chaotic Boltzmann Machine
abstract
This paper reports on a field programmable gate array (FPGA) implementation of Chaotic Boltzmann Machine Reservoir Computing (CBM-RC). The reservoir will be large-scale, as it is expected to be applied to sensor information prediction for autonomous mobile robots. Therefore, we employ a design premised on storing the weight information into a large memory outside the FPGA. We propose an efficient compression method for the weight matrix and a parallel processing system, by considering both the characteristics of CBM-RC and the fact that the weight matrix of a large-scale reservoir is generally a sparse matrix. Our RC system, which has more than 8000 neurons and 1024 inputs/outputs, has been implemented on an AMD Alveo U50 FPGA board. This RC is the largest scale compared to those in related studies. We have performed the NARMA10 task and demonstrated that we can estimate 1024 predictions at once with NMSE accuracy that is even or better to conventional RC.
Shigeki Matsumoto, Yuki Ichikawa, Nobuki Kajihara, Hakaru Tamukoh
ISCAS4
2024 Unified Understanding of Environment, Task, and Human for Human-Robot Interaction in Real-World Environments
abstract
To facilitate human–robot interaction (HRI) tasks in real-world scenarios, service robots must adapt to dynamic environments and understand the required tasks while effectively communicating with humans. To accomplish HRI in practice, we propose a novel indoor dynamic map, task understanding system, and response generation system. The indoor dynamic map optimizes robot behavior by managing an occupancy grid map and dynamic information, such as furniture and humans, in separate layers. The task understanding system targets tasks that require multiple actions, such as serving ordered items. Task representations that predefine the flow of necessary actions are applied to achieve highly accurate understanding. The response generation system is executed in parallel with task understanding to facilitate smooth HRI by informing humans of the subsequent actions of the robot. In this study, we focused on waiter duties in a restaurant setting as a representative application of HRI in a dynamic environment. We developed an HRI system that could perform tasks such as serving food and cleaning up while communicating with customers. In experiments conducted in a simulated restaurant environment, the proposed HRI system successfully communicated with customers and served ordered food with 90% accuracy. In a questionnaire administered after the experiment, the HRI system of the robot received 4.2 points out of 5. These outcomes indicated the effectiveness of the proposed method and HRI system in executing waiter tasks in real-world environments.
Yuga Yano, Akinobu Mizutani, Yukiya Fukuda, Daiju Kanaoka, Tomohiro Ono, Hakaru Tamukoh
RO-MAN6
2024 Hibikino-Musashi@Home RoboCup@Home DSPL Champion 2024
Akinobu Mizutani, Kosei Isomoto, Kosei Yamao, Ryohei Kobayashi 0003, Soma Fumoto, Koshun Arimura, Naoki Yamaguchi, Tomoya Shiba, Kouki Kimizuka, Yuta Ohno, Ryo Terashima, Hiromasa Yamaguchi, Tomoaki Fujino, Ryoga Maruno, Wataru Yoshimura, Kazuhito Mine, Tang Phu Thien Nhan, Yuga Yano, Yuichiro Tanaka, Takeshi Nishida, Takashi Morie, Hakaru Tamukoh
RoboCup22
2023 ManifoldNeRF: View-dependent Image Feature Supervision for Few-shot Neural Radiance Fields
Daiju Kanaoka, Motoharu Sonogashira, Hakaru Tamukoh, Yasutomo Kawanishi
BMVC3
2023 Efficient Repetition Coding for Deep Learning Towards Implementation Using Emerging Non-Volatile Memory with Write-Errors
abstract
Emerging non-volatile memory devices, such as resistive random access memory (ReRAM) and voltage-controlled magnetoresistive random access memory (VC-MRAM), promise low energy consumption for artificial intelligence applications. However, when implementing deep neural networks (DNNs) using such memory devices, write-error may cause millions of bit-flipping to DNN. This easily degrades the DNN performance. To address this problem, we propose a novel repetition coding for deep-learning (RC-DL), which is a repetition coding designed to protect IEEE 32-bit floating-point (FP32) DNN models. Compared to conventional repetition coding, the proposed RC-DL exploits FP32 non-uniform magnitude encoding by increasing the repeat rates to protect sensitive bit positions and reduce the repeat rates to insensitive bit positions. Hence, RC-DL uses a number of bits equivalent to a 3-bit repetition code while delivering the performance close to 11-bit repetition code. We perform extensive Monte Carlo simulations to simulate the write-error property with ImageNet 2012 pretrained models. The DNN models with RC-DL are shown to be operable in the extremely imperfect environment while delivering with only minor reductions in DNN performance.
Ninnart Fuengfusin, Hakaru Tamukoh, Yuichiro Tanaka, Osamu Nomura, Takashi Morie
IJCNN2
2023 FPGA Implementation of a Chaotic Boltzmann Machine Annealer
abstract
Ising machines are attracting attention for their ability to solve large-scale combinatorial optimization problems because these problems are difficult to solve. To accelerate the computing of Ising machines, implementation of Ising machines with digital circuits such as simulated annealing (SA) machines is in progress. However, these Ising machines on digital circuits require random number generators, which are implemented with large circuit resources. This work focuses on chaotic Boltzmann machines (CBMs), which imitate the stochastic behavior of Boltzmann machines (BMs) with deterministic chaotic dynamics. CBMs are one of the models that work as chaotic simulated annealing (CSA) machines within Ising machines. Therefore, we can implement the Ising machines without random number generators by using CBMs. In conventional work, CSA machines using CBMs (CBM-CSAs) are implemented with some hardware-oriented algorithms, but the CBM-CSA circuit is not optimized for these hardware-oriented algorithms. In the conventional CBM-CSA circuit, memory circuits are implemented separately, which prevents making the CBM-CSA from larger, and neuron circuits require the reset of accumulated values, which causes the increase in the calculation time. To solve these problems, we implement only one large memory circuit to make the CBM-CSA larger and improve the neuron circuits to allow dynamic changes of inputs to arithmetic circuits to inhibit the increase in the calculation time. As a result, we implement a CBM-CSA with 4096 nodes on an FPGA (Alveo U250), and the CBM-CSA can control 16-bit width weights and run at 100MHz. We evaluate the implemented CBM-CSA by solving K4000, max-cut problem, which is one of the combinatorial optimization problems. The best solution of CBM-CSA is comparable to that of the SA on the central processing unit (CPU). Moreover, the CBM-CSA is approximately 600 times as fast as the SA on the CPU and approximately twice as fast as the conventional Ising machine on an FPGA based on the improvements in this work. Furthermore, this work implements one of the highest-performance Ising machines on a single FPGA.
Kanta Yoshioka, Yuichi Katori, Yuichiro Tanaka, Osamu Nomura, Takashi Morie, Hakaru Tamukoh
IJCNN6
2023 In-material reservoir implementation of reservoir-based convolution
abstract
This study aims to implement a reservoir-based convolutional neural network (CNN) on physical reservoir computing (RC) to develop an efficient image recognition system for edge AI. Therefore, we propose a novel reservoir-based convolution circuit system that uses in-material reservoir computing, a type of physical RC made from a sulfonated polyaniline network. The experimental results demonstrate that the proposed circuit system extracts image features in the same way as the original CNN and that a reservoir-based CNN on the in-material RC achieves an accuracy rate of 81.7% in an image classification task while an echo state network-based CNN achieves 87.7%.
Yuichiro Tanaka, Yuki Usami, Hirofumi Tanaka, Hakaru Tamukoh
ISCAS4
2023 Dense Traversability Estimation System for Extreme Environments
abstract
Traversability estimation is essential for safe path planning in robotics and autonomous driving. In this study, we propose an accurate and dense traversability estimation system for extreme environments such as disaster areas. Traversability estimations often occur undefined regions due to the shielding and sparsity of sensor data. These regions may lead to the selecting of dangerous paths. The proposed system uses point cloud accumulation and the Bayesian generalized kernel (BGK) elevation estimation to create a dense Digital elevation map (DEM) with no undefined regions from point clouds obtained from light detection and ranging (LiDAR). Subsequently, traversability is estimated from road surface roughness, slope and vehicle performance using fuzzy logic in our system. In our experiments, our system is installed in an experimental vehicle, and the experiments conducted on rocky terrain, slopes, and craters to verify its effectivity in extreme environments. Results show that our system can create a dense DEM with few undefined regions and estimate valid traversability for all obstacles.
Yukiya Fukuda, Yuya Mii, Yuga Yano, Hidenari Iwai, Shintaro Inoue, Hakaru Tamukoh
IV6
2023 Autonomous Waiter Robot System for Recognizing Customers, Taking Orders, and Serving Food
Yuga Yano, Kosei Isomoto, Tomohiro Ono, Hakaru Tamukoh
RoboCup4
2022 Desgin and Implementation of ROS2-based Autonomous Tiny Robot Car with Integration of Multiple ROS2 FPGA Nodes
abstract
This paper introduces an autonomous tiny robot car equipped with a camera-based lane detection function and a traffic signal/obstacle, pedestrian recognition function. Each function is integrated by Robot Operating System 2 (ROS2), a middleware for robot system development. Autonomous driving without the need for a driver requires not only lane-following driving but also traffic signal recognition and obstacle recognition. These functions are implemented on FPGA, and we evaluated them. According to these results, the execution time of traffic signal recognition by FPGA was 1.2 to 3.4 times faster than CPU execution. YOLOv4 is used for obstacle recognition, which improved mAP by 3.79 points compared to YOLO v3-Tiny.
Hayato Mori, Hayato Amano, Akinobu Mizutani, Eisuke Okazaki, Yuki Konno, Kohei Sada, Tomohiro Ono, Yuma Yoshimoto, Hakaru Tamukoh, Takeshi Ohkawa, Midori Sugaya
FPT9
2022 A memory-based entorhinal-hippocampal model and its FPGA implementation by on-chip RAMs
abstract
Artificial general intelligence, which imitates the human brain, is aspired. Episodic memories are considered to be a key feature in building human brain functions. This paper proposes a memory-based entorhinal-hippocampal model that encodes spatial and non-spatial information, essential to realize episodic memories. The model works as a memory that stores the location of objects and events as neural activity packets. This paper also proposes an area-efficient hardware implementation method for field-programmable gate arrays (FPGAs). Our proposal utilizes on-chip random access memories (RAMs) to achieve a large-scale implementation of our model. Circuit simulations validated the behavior of our hardware-friendly model. The results of logic synthesis revealed the area efficiency of the FPGA implementation method that utilizes on-chip RAMs.
Ichiro Kawashima, Katsumi Tateno, Takashi Morie, Hakaru Tamukoh
ISCAS4
2021 A dataset generation for object recognition and a tool for generating ROS2 FPGA node
abstract
This paper introduces our autonomous driving system equipped with recognition processing units from a camera image for hazard object / human-doll detection and drive lane detection. In particular, this paper focuses on a dataset generation method for neural networks and a generation tool “FPGA Oriented Easy Synthesizer Tool (FOrEST)” for ROS2-FPGA nodes. The results show that mAP of a neural network trained by the generated dataset is 94%, and a overhead of ROS2-FPGA communication by the FOrEST is 2–3 ms.
Hayato Amano, Hayato Mori, Akinobu Mizutani, Tomohiro Ono, Yuma Yoshimoto, Takeshi Ohkawa, Hakaru Tamukoh
FPT7
2021 An area-efficient multiply-accumulation architecture and implementations for time-domain neural processing
abstract
In our work, a new area-efficient multiply-accumulation scheme for time-domain neural processing named differential multiply-accumulation is proposed. Our new scheme reduces hardware resources utilization of multiply-accumulation with suppressing the increasing computational time resulting from the time-multiplexing. As a result, 2,048 neurons of fully connected CBM and RC-CBM were synthesized for a single field-programmable gate array (FPGA).
Ichiro Kawashima, Yuichi Katori, Takashi Morie, Hakaru Tamukoh
FPT4
2021 FPGA Implementation of Pulse-Coupled Phase Oscillators working as a Reservoir at the Edge of Chaos
abstract
In the field of neuroscience, reservoir computing (RC) has been viewed as a model of the neuron computational system. RC is a framework for constructing recurrent neural networks, which is used for modeling the parts of the brain to solve the temporal problem. We construct the network inside the reservoir using the Pulse-Coupled Phase Oscillator (PCPO) with neighbor topology connections on field-programmable-gate- array (FPGA). Winfree model is used for PCPO spiking-based. We investigate the stability of the edge phenomenon using the zero one test (Z1-Test) methodology. We evaluate the proposed model on the time series generation tasks. We have successfully implemented and verified using FPGA that the 3×3 and 10×10 PCPO working as a Reservoir on the network.
Dinda Pramanta, Hakaru Tamukoh
ISCAS2
2021 An efficient hardware-oriented dropout algorithm
Yoeng Jye Yeoh, Takashi Morie, Hakaru Tamukoh
Neurocomputing3
2020 Design and Implementation of Pulse-Coupled Phase Oscillators on a Field-Programmable Gate Array for Reservoir Computing
Dinda Pramanta, Hakaru Tamukoh
ICONIP (5)2
2020 Live Demonstration: Hardware-Oriented Dual Stream Object Recognition System using Binarized Neural Networks
abstract
This live demonstration presents a “Binarized Dual Stream VGG-16 (BDS-VGG16)” model which is a convolutional neural networks model for object recognition. The model is designed for hardware accelerators (i.e., field programmable gate arrays) for implementation in robots. The model uses RGB images and Depth images for increasing accuracy. In addition, the model achieved 99.2% in the experiment using RGB-D Object Dataset. In the demonstration, visitors will learn the operation of the BDS-VGG16 model which is an object recognition system for the service robots.
Yuma Yoshimoto, Hakaru Tamukoh
ISCAS2
2020 Hardware-Oriented Dual Stream Object Recognition System using Binarized Neural Networks
abstract
Service robots require an object recognition system to ensure their effective functioning in situations where Convolutional Neural Networks (CNN) are the mainstream machine learning technique employed by the system. Particularly, “Dual Stream VGG-16 (DS-VGG16),” which uses RGB and depth images, has been reported to have high accuracy for object recognition. However, it is difficult to implement CNN in robots, because it requires high computation power and consumes a huge amount of power. Although, implementing CNN with Field Programmable Gate Array (FPGA) solves the electric power problem, however it is difficult due to limited resources available. This paper proposes “Binarized Dual Stream VGG-16 (BDS-VGG16),” which is Hardware-Oriented DS-VGG16. With the concept of Binarized Neural Networks (BNN), BDS-VGG16 is effective when implemented on FPGA. In results, the accuracy of BDS-VGG16 is 99.2%. It is higher than that of Eitel's model by 5.1 points. Further, we developed an object recognition system based on the Robot Operating System (ROS) which is a well-known middleware for robots, it uses our proposed method for service robot application.
Yuma Yoshimoto, Hakaru Tamukoh
ISCAS2
2019 High-Speed Synchronization of Pulse-Coupled Phase Oscillators on Multi-FPGA
Dinda Pramanta, Hakaru Tamukoh
ICONIP (5)2
2019 Reservoir Computing Based on Dynamics of Pseudo-Billiard System in Hypercube
abstract
Reservoir computing (RC) is a framework for constructing recurrent neural networks with simple training rule and sparsely and randomly connected nonlinear units. The network (called reservoir) generates complex motion that can be used for many tasks including time series generation and prediction. We construct a reservoir based on the dynamics of the pseudo-billiard system that produce complex motion in a high-dimensional hypercube. In particular, we use the chaotic Boltzmann machine (CBM) whose units exhibit chaotic behavior in the hypercube. The units interact with each other in a time-domain manner through its binary state, and thus an efficient hardware implementation of the system is expected. In order to utilize the CBM as the reservoir, it is necessary to control its chaotic behavior for ensuring the echo state property of RC and establish encoding and decoding for input and output signal. For this purpose, we introduce a reference clock and analyze effects and properties of the reference input. We evaluate the proposed model on the time series generation tasks and show that the model works properly on a broad range of parameter values. Our approach presents a novel mechanism for time-domain information processing and a fundamental technology for a brain like artificial intelligence system.
Yuichi Katori, Hakaru Tamukoh, Takashi Morie
IJCNN2
2019 A Chaotic Boltzmann Machine Working as a Reservoir and Its Analog VLSI Implementation
abstract
Reservoir computing is attracting great interest because of its high computing ability especially for time-series prediction, despite its simple structure and learning scheme. This paper proposes a reservoir computing hardware model using a chaotic Boltzmann machine (CBM) as the reservoir, which can achieve complex motion in a dynamical system on a high-dimensional hypercube. The CBM uses analog nonlinear dynamics, unlike the stochastic operation of the original Boltzmann machine model. To utilize CBMs as a reservoir, chaotic operation must be suppressed, and the echo state property should be satisfied. We modify the CBM model for simpler analog complementary metal-oxide-semiconductor very-large-scale integration (CMOS VLSI) implementation, and propose its use as a reservoir by adding an external reference clock signal. We then verify its proper operation by numerical simulation. We also refine the CMOS VLSI circuit design based on the proposed modified CBM model to improve power consumption and calculation precision.
Masatoshi Yamaguchi, Yuichi Katori, Daichi Kamimura, Hakaru Tamukoh, Takashi Morie
IJCNN4
2019 Live Demonstration: Hardware Implementation of Brain-Inspired Amygdala Model
abstract
This live demonstration presents a brain-inspired amygdala model. Amygdala is an area of the brain that is associated with fear conditioning, which is a type of classical conditioning. The model can learn preferences through humanrobot interactions by application of classical conditioning to the model. Additionally, to develop a high speed and low power system, we design a hardware of the amygdala model, and implemented the hardware into field programmable gate array.
Yuichiro Tanaka, Hakaru Tamukoh
ISCAS2
2019 Hardware Implementation of Brain-Inspired Amygdala Model
abstract
Deep neural networks (DNNs) have achieved state-of-the-art results in several computing tasks. However, the performance of these DNNs is reliant on the availability of large amounts of training data, which is not always present. We approached this problem by developing a brain-inspired amygdala model to achieve computer learning based on limited training data. The amygdala is an area of the brain associated with classical fear conditioning. The proposed amygdala model is composed of a single layer of deep self-organizing map network (deep SOM network) and a fully-connected neural network (FCNN), which imitates the function and structure of an amygdala. We applied the proposed amygdala model to a robot waiter task in a restaurant. The experimental results show that the model learned a customer's preferences after only a few human robot interactions. To develop the digital hardware of the amygdala model, we designed hardware for the deep SOM network and the FCNN and implemented them in an XCZU9EG field programmable gate array (FPGA). Our FPGA implementation of a deep SOM network with 272 neurons and an FCNN with three output neurons outperformed a software implementation on an Intel Core i5-3470 CPU by over 600 times.
Yuichiro Tanaka, Hakaru Tamukoh
ISCAS2
2019 Live Demonstration: A VLSI Implementation of Time-Domain Analog Weighted-Sum Calculation Model for Intelligent Processing on Robots
abstract
This live demonstration presents a VLSI chip based on “Time-domain Analog Computing with Transient states (TACT)” approach for intelligent processing on robots. This TACT chip, fabricated using 250-nm CMOS technology, implements a time-domain analog weighted-sum calculation model with very high energy efficiency. We integrate the TACT chip into a robot via Robot Operating System (ROS) interfaces. A human tracking robot demonstration is performed by the TACT chip with energy efficiency of 300 TOPS/W.
Masatoshi Yamaguchi, Gouki Iwamoto, Yushi Abe, Yuichiro Tanaka, Yutaro Ishida, Hakaru Tamukoh, Takashi Morie
ISCAS6
2018 Mixed Precision Weight Networks: Training Neural Networks with Varied Precision Weights
Ninnart Fuengfusin, Hakaru Tamukoh
ICONIP (2)2
2018 Live Demonstration: A Hardware Accelerated Robot Middleware Package for Intelligent Processing on Robots
abstract
This live demonstration presents a "connective object for middleware to accelerator (COMTA)," an intelligent processing system that uses hardware accelerators (i.e., field programmable gate arrays (FPGAs)) and robot middleware. The key idea of COMTA is to automatically generate the system via robot middleware interfaces. To realize the proposed system, we have developed a block of programs called an "object" in a hardware/software complex system. We demonstrate an implementation of a human tracking image processing application on a vehicle robot accelerated by COMTA. The demonstration system achieved 3.3 times better power efficiency than a general PCs.
Yutaro Ishida, Takashi Morie, Hakaru Tamukoh
ISCAS3
2018 A Hardware Accelerated Robot Middleware Package for Intelligent Processing on Robots
abstract
Service robots require implementation of intelligent processing, e.g., image processing. However, the computational resources of standard PCs typically used in service robots are not sufficient for such processes. Furthermore, robot middleware is widely used in many robots because such systems facilitate integration and are suitable for rapid prototyping. We propose a "connective object for middleware to accelerator (COMTA)," which is a processing system that uses hardware accelerators, i.e., field programmable gate arrays (FPGAs), and robot middleware. Users can access the FPGAs in the proposed system via middleware interfaces; thus, complex internal circuits are not required. For human tracking using image processing, the proposed system can automatically generate from a single configuration file. The proposed system performs 3.3 times more efficiently relative to computation than standard PCs in robots.
Yutaro Ishida, Takashi Morie, Hakaru Tamukoh
ISCAS3
2017 A Hardware-Oriented Dropout Algorithm for Efficient FPGA Implementation
Yoeng Jye Yeoh, Takashi Morie, Hakaru Tamukoh
ICONIP (6)3
2017 A CMOS chaotic Boltzmann machine circuit and three-neuron network operation
abstract
This paper proposes CMOS VLSI implementation of a chaotic Boltzmann machine (CBM) model, which uses analog nonlinear dynamics instead of stochastic operation as in the original Boltzmann machine model. The CBM model is suitable for efficient VLSI implementation of Boltzmann machines because it requires no random number generator circuits, which consume a considerable footprint on a VLSI chip as well as considerable power. We describe the design results of CMOS circuits of neuron and synapse units. The neuron circuit uses subthreshold operation of MOSFETs to realize the exponential function used in the CBM model. We also provide measurement results of a fabricated CMOS chip for single-neuron unit circuit operation and demonstrate chaotic behavior in a three-neuron network.
Masatoshi Yamaguchi, Hakaru Tamukoh, Hideyuki Suzuki, Takashi Morie
IJCNN2
2016 Restricted Boltzmann Machines Without Random Number Generators for Efficient Digital Hardware Implementation
Sansei Hori, Takashi Morie, Hakaru Tamukoh
ICANN (1)3
2016 FPGA Implementation of Autoencoders Having Shared Synapse Architecture
Akihiro Suzuki, Takashi Morie, Hakaru Tamukoh
ICONIP (1)3
2016 Time-Domain Weighted-Sum Calculation for Ultimately Low Power VLSI Neural Networks
Hakaru Tamukoh, Takashi Morie
ICONIP (1)2
2016 A CMOS Unit Circuit Using Subthreshold Operation of MOSFETs for Chaotic Boltzmann Machines
Masatoshi Yamaguchi, Takashi Kato, Hideyuki Suzuki, Hakaru Tamukoh, Takashi Morie
ICONIP (1)5
2015 A Color Quantization Based on Vector Error Diffusion and Particle Swarm Optimization Considering Human Visibility
Ryosuke Kubota, Hakaru Tamukoh, Hideaki Kawano, Noriaki Suetake, Byungki Cha, Takashi Aso
PSIVT2
2015 Parameterized digital hardware design of pulse-coupled phase oscillator networks
Yasuhiro Suedomi, Hakaru Tamukoh, Kenji Matsuzaka, Michio Tanaka, Takashi Morie
Neurocomputing2
2014 Morphological Associative Memory Employing a Split Store Method
Hakaru Tamukoh, Kensuke Koga, Hideaki Harada, Takashi Morie
ICONIP (3)1
2013 Parameterized Digital Hardware Design of Pulse-Coupled Phase Oscillator Model toward Spike-Based Computing
Yasuhiro Suedomi, Hakaru Tamukoh, Michio Tanaka, Kenji Matsuzaka, Takashi Morie
ICONIP (3)2
2013 Design of networked hw/sw complex system using hardware object model and its application
abstract
In this paper, we describe a design example of a networked hw/sw complex system based on the hardware object model. In the design flow of the proposed system, a hardware unit can be handled as an object in object-oriented design. This hardware object can be dynamically constructed, executed, and destructed from any remote applications on the network. The hardware object is loaded into the FPGA's virtual hardware circuit space, and accelerates user applications. To realize these functions, we propose virtualization technologies to remotely reconfigure and control the hardware object. As the platform of the networked hw/sw complex system, we have developed a hwModule FPGA board and its System Development Kit. In order to demonstrate the proposed architecture's effectiveness, a video streaming application based on the networked hw/sw complex system has been developed.
Hakaru Tamukoh, Masatoshi Sekine
IECON1
2012 Live demonstration: "Internet Booster" a novel WEB application platform accelerated by reconfigurable virtual hardware circuits
abstract
This live demonstration presents an Internet Booster (IB) which is a novel WEB application platform accelerated by virtual hardware circuits. Akey-idea of IB is obtaining relevant WEB applications virtual hardware circuit through the Internet. The downloaded virtual hardware circuit is dynamically reconfigured into an FPGA and cooperates with software and network to accelerate WEB application. In order to realize the concept of IB, we develop a hwModule VC FPGA Cardbus board and a networked hw/sw complex system to access virtual hardware circiuts via the Internet. Furthermore, we also develop a virtual hardware video codec and a TCP/IP hardware stack to construct a demonstration system. In the live demonstration, we show a WEB streaming application accelerated by IB on the local area network environment. The demonstration system achieves full-color VGA size (24 bits RGB, 640 × 480 pixels) video streaming with around 20 frames per second.
Hakaru Tamukoh, Nadav Bergstein, Kotoko Fujita, Masatoshi Sekine
ISCAS1
2010 A Dynamically Reconfigurable Platform for Self-Organizing Neural Network Hardware
Hakaru Tamukoh, Masatoshi Sekine
ICONIP (2)1
2006 A Digital Hardware Architecture of Self-Organizing Relationship (SOR) Network
Hakaru Tamukoh, Keiichi Horio, Takeshi Yamakawa
ICONIP (3)1
2004 Self-organizing map hardware accelerator system and its application to realtime image enlargement
abstract
We propose a new fast learning algorithm for SOM and its digital hardware design based on the massively parallel architecture. When this proposed algorithm is realized by using Xilinx XC2V6000-6 FPGA, a maximum performance of 17500 MCUPS is achieved and up to 256 competing units (16 /spl times/ 16 map) can be implemented. Each competing unit have a weight vector which is represented by 128 elements of 16 bits accuracy. Furthermore, we applied the proposed hardware to a realtime digital image enlargement system. In the case of full color (24 bits) image enlargement from QQVGA (160 /spl times/ 120 pixel) to QVGA (320 /spl times/ 240 pixel), a proposed hardware requires only 0.12 second per image, while the personal computer (Intel XEON, 2.8 GHz Dual) requires more than 5 seconds per image.
Hakaru Tamukoh, Taksahi Aso, Keiichi Horio, Takeshi Yamakawa
IJCNN1