Haruyoshi Yonekawa

dblp:194/3975 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 87% Embedded and real-time systems · 7% Reconfigurable computing and FPGAs · 6%
Artificial intelligence
2 papers
Efficient and distributed learning · 64% Image recognition and object detection · 36%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.622018
A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA · FPGA 2018
A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only) · FPGA 2017
Computer vision › Image recognition and object detection
object detection
0.312018
A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA · FPGA 2018
Hardware accelerators and domain-specific architectures › vision accelerator
object detection accelerator
0.312018
A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA · FPGA 2018
Machine learning › Efficient and distributed learning › model compression › quantization › quantized neural network
binary neural network
0.312017
A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only) · FPGA 2017
Machine learning › Efficient and distributed learning
model compression
0.312017
A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only) · FPGA 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
binary neural network accelerator
0.312017
A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only) · FPGA 2017
Embedded and real-time systems › real-time embedded systems › multimedia embedded systems
embedded vision system
0.112018
A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA · FPGA 2018
Reconfigurable computing and FPGAs
FPGA implementation
0.112017
A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only) · FPGA 2017

Methods — techniques the papers use, named apart from their topics

support vector regression · 0.7binarized neural network · 0.7batch normalization removal · 0.6
YearPublicationVenuePosition
2018 A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA
abstract
A frame object detection problem consists of two problems: one is a regression problem to spatially separated bounding boxes, the second is the associated classification of the objects within realtime frame rate. It is widely used in the embedded systems, such as robotics, autonomous driving, security, and drones - all of which require high-performance and low-power consumption. This paper implements the YOLO (You only look once) object detector on an FPGA, which is faster and has a higher accuracy. It is based on the convolutional deep neural network (CNN), and it is a dominant part both the performance and the area. However, the object detector based on the CNN consists of a bounding box prediction (regression) and a class estimation (classification). Thus, the conventional all binarized CNN fails to recognize in most cases. In the paper, we propose a lightweight YOLOv2, which consists of the binarized CNN for a feature extraction and the parallel support vector regression (SVR) for both a classification and a localization. To our knowledge, this is the first time binarized CNN»s have been successfully used in object detection. We implement a pipelined based architecture for the lightweight YOLOv2 on the Xilinx Inc. zcu102 board, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC. The implemented object detector archived 40.81 frames per second (FPS). Compared with the ARM Cortex-A57, it was 177.4 times faster, it dissipated 1.1 times more power, and its performance per power efficiency was 158.9 times better. Also, compared with the nVidia Pascall embedded GPU, it was 27.5 times faster, it dissipated 1.5 times lower power, and its performance per power efficiency was 42.9 times better. Thus, our method is suitable for the frame object detector for an embedded vision system.
Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Shimpei Sato
FPGA2
2017 A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only)
Hiroki Nakahara, Haruyoshi Yonekawa, Hisashi Iwamoto, Masato Motomura
FPGA2
2017 A demonstration of the GUINNESS: A GUI based neural NEtwork SyntheSizer for an FPGA
abstract
Summary form only given. The GUINNESS is a tool flow for the deep neural network toward FPGA implementation [3,4,5] based on the GUI (Graphical User Interface) including both the binarized deep neural network training on GPUs and the inference on an FPGA. It generates the trained the Binarized deep neural network [2] on the desktop PC, then, it generates the bitstream by using standard the FPGA CAD tool flow. All the operation is done on the GUI, thus, the designer is not necessary to write any scripts to descript the neural network structure, training behaviour, only specify the values for hyper parameters. After finished the training, it automatically generates C++ codes to synthesis the bitstream using the Xilinx SDSoC system design tool flow. Thus, our tool flow is suitable for the software programmers who are not familiar with the FPGA design.
Hiroyuki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Masayuki Shimoda, Simpei Sato
FPL2
2017 An object detector based on multiscale sliding window search using a fully pipelined binarized CNN on an FPGA
abstract
An object detection problem consists of two problems: one is classification of detected object category and the other is localization. Frame object detection is used in an embedded vision systems, such as a robot, an automobile, a security camera, and a drone. These applications require high-performance computation and low-power consumption by an inexpensive device. This paper proposes multiscale sliding window based object detector using a fully pipelined binarized deep convolutional neural network (BCNN) on an FPGA. It consists of a sliding window part, a fully pipelined BCNN classifier, and an ARM processing unit for detection. Duplicate detections were filtered by using a non-maximum suppression algorithm running on the ARM processor. We propose the fully pipelined layers for the BCNN and its architecture for FPGA realization. Since the proposed BCNN circuit uses on-chip memories on the FPGA, its throughput is higher than a GPU based one with practical recognition accuracy. We trained the VGG11 based BCNN using the KITTI vision benchmark for the car detection scenario. Then, we implemented the proposed object detector on the Xilinx Inc. Zynq UltraScale+ MPSoC zcu102 evaluation board. The GPU based object detectors were too slow for the realtime application requirement (HD frame rate), with the exception of YOLOv2. As compared with the GPU implementation of YOLOv2, the proposed FPGA detector had higher recognition accuracy and lower power consumption. Compared with the YOLOv2, the proposed FPGA one is higher with respect to recognition accuracy, and its power consumption is lower than the GPU based YOLOv2. Thus, the FPGA based object detector suitable for the embedded realtime applications.
Hiroki Nakahara, Haruyoshi Yonekawa, Shimpei Sato
FPT2
2016 A memory-based realization of a binarized deep convolutional neural network
abstract
A pre-trained deep convolutional neural network (CNN) is a feed-forward computation perspective, which is widely used for the embedded systems, requires high power-and-area efficiency. This paper realizes a binarized CNN which treats only binary 2-values (+1/-1) for the inputs and the weights. In this case, the multiplier is replaced with an EX-NOR circuit. To reduce both power and area, we realize the 2-valued CNN by off- and on-chip memories. Since our 2D convolution operations are realized by the on-chip memory, our implementation consumes lower power than the DSP block. We decompose the memory part, and realize them by a cascade of memories (LUT cascade). By introducing a batch normalization technique, the classification error for the binarized CNN can be improved. We implemented the CIFAR-10 benchmark on the NetFPGA-SUME board, which has the Xilinx Inc. Virtex 7 FPGA and three on-chip QDR II+ Synchronous SRAMs. Compared with the conventional FPGA realizations, the performance is 2.82 times faster, the power efficiency is 1.76 times, and the area efficiency is 29.13 times better.
Hiroki Nakahara, Haruyoshi Yonekawa, Tsutomu Sasao, Hisashi Iwamoto, Masato Motomura
FPT2