EDBT 2026 Demo / reviewers in the wild / expert
Hiroki Nakahara
dblp:80/2794
· DBLP profile ↗
37ranked-venue papers
23as first author
5since 2021 · last 2023
0000-0002-5701-7466ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 23 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Remarn: A Reconfigurable Multi-threaded Multi-core Accelerator for Recurrent Neural NetworksabstractThis work introduces Remarn, a reconfigurable multi-threaded multi-core accelerator supporting both spatial and temporal co-execution of Recurrent Neural Network (RNN) inferences. It increases processing capabilities and quality of service of cloud-based neural processing units (NPUs) by improving their hardware utilization and by reducing design latency, with two innovations. First, a custom coarse-grained multi-threaded RNN/Long Short-Term Memory (LSTM) hardware architecture, switching tasks among threads when RNN computational engines meet data hazards. Second, the partitioning of this hardware architecture into multiple full-fledged sub-accelerator cores, enabling spatially co-execution of multiple RNN/LSTM inferences. These innovations improve the exploitation of the available parallelism to increase runtime hardware utilization and boost design throughput. Evaluation results show that a dual-threaded quad-core Remarn NPU achieves 2.91 times higher performance while only occupying 5.0% more area than a single-threaded one on a Stratix 10 FPGA. When compared with a Tesla V100 GPU implementation, our design achieves 6.5 times better performance and 15.6 times higher power efficiency, showing that our approach contributes to high performance and energy-efficient FPGA-based multi-RNN inference designs for datacenters. Zhiqiang Que, Hiroki Nakahara, Hongxiang Fan, He Li 0008, Jiuxi Meng, Kuen Hung Tsoi, Xinyu Niu, Eriko Nurvitadhi, Wayne Luk |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | PrefaceabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Yun Liang 0001, Hiroki Nakahara, Wei Zhang 0012, Fubing Mao, Ray C. C. Cheung |
FPT | 2 |
| 2022 | Message from the General Chair and Program Co-ChairsabstractOn behalf of the FPT'22 Organizing Committee, we appreciate all of you for joining FPT'22 both in-person or virtually. We wished we could say “Welcome to Hong Kong” to all the attendees, but the ongoing Covid-19 travel restrictions still make overseas travel inconvenient for some of the attendees. Hence, we have worked very hard to give good support to both the in-person and virtual attendees. We bring a hybrid mode FPT'22 and hope to connect all the attendees together to enjoy the exciting program. Wei Zhang 0012, Ray C. C. Cheung, Yun Liang 0001, Hiroki Nakahara |
FPT | 4 |
| 2022 | Recurrent Neural Networks With Column-Wise Matrix-Vector Multiplication on FPGAsabstractThis article presents a reconfigurable accelerator for REcurrent Neural networks with fine-grained cOlumn-Wise matrix–vector multiplicatioN (RENOWN). We propose a novel latency-hiding architecture for recurrent neural network (RNN) acceleration using column-wise matrix–vector multiplication (MVM) instead of the state-of-the-art row-wise operation. This hardware (HW) architecture can eliminate data dependencies to improve the throughput of RNN inference systems. Besides, we introduce a configurable checkerboard tiling strategy which allows large weight matrices, while incorporating various configurations of element-based parallelism (EP) and vector-based parallelism (VP). These optimizations improve the exploitation of parallelism to increase HW utilization and enhance system throughput. Evaluation results show that our design can achieve over 29.6 tera operations per second (TOPS) which would be among the highest for field-programmable gate array (FPGA)-based RNN designs. Compared to state-of-the-art accelerators on FPGAs, our design achieves 3.7–14.8 times better performance and has the highest HW utilization. Zhiqiang Que, Hiroki Nakahara, Eriko Nurvitadhi, Andrew Boutros, Hongxiang Fan, Chenglong Zeng, Jiuxi Meng, Kuen Hung Tsoi, Xinyu Niu, Wayne Luk |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation AlignerabstractA 4b-quantized convolutional neural network (CNN) inference engine for edge-AI is presented featuring a Cartesian-product MAC array and pipelined activation aligners targeting deep-/random-pruned models. A 40nm prototype with 32x32 MACs and 5Mb SRAM runs at 534 MHz, 1.07 TOPS, 352 mW at 1.1V, and attains 5.30 dense TOPS/W, 234 MHz at 0.8V. Sparse TOPS/W reaches 26.5 when running a randomly pruned model (after 88% pruning). Training algorithms for obtaining highly efficient sparse/quantized models are also proposed. Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura |
HCS | 4 |
| 2020 | Tiny On-Chip Memory Realization of Weight Sparseness Split-CNNs on Low-end FPGAsabstractConsidering implementation in a low-end FPGA with more restrictions on on-chip memory resources and external memory bandwidth, existing methods are Limited by communication with external memory. The on-chip memory in a low-end FPGA is small, hence cannot store an entire feature map. It is imperative to rely on external memory for buffering. Although an on-chip memory in an FPGA is fast, implementing sparse CNN in low-end FPGAs is hindered by their limited memory size. To address the limitations of external memory bandwidth and its size, we employ a split-CNN [1] that splits an input image into small spatial patches and tests each patch using a CNN model as shown in Fig. 1. Since the feature-map to be processed at one time is reduced by the splitting, the amount of memory required for buffering is reduced. Akira Jinguji, Shimpei Sato, Hiroki Nakahara |
FCCM | 3 |
| 2020 | High-Throughput Convolutional Neural Network on an FPGA by Customized JPEG CompressionabstractThe growing interest in using FPGAs to accelerate convolutional neural network (CNN) workloads is driving the deployment of FPGAs on cloud services such as Amazon AWS and Microsoft Azure. Such current cloud-based FPGAs have serious problems concerning data transfer bandwidth. In this paper, we compress a transfer image using customized JPEG coding and implement a customized image decoder architecture. We analyze the trade-off between data transfer speed-up and recognition accuracy drop. Based on this compression scheme, we design a high-throughput CNN inference engine. Almost all existing FPGA-based CNN accelerators are based with the same idea as their GPU counterparts, where operations from different network layers are mapped onto the same hardware units working in a multiplexed way. Our fully pipelined architecture maps all the network layers on-chip and transfers the computation from different layers to their unit with independent optimization. We apply two CNN optimization techniques to a residual network, one is a channel shift and point-wise approximation, and the other is a binary weight quantization. We implement the proposed CNN inference accelerator on the Xilinx Virtex UltraScale+ XCVU9P FPGA. Our system peak-performance achieves 2.41 TOPS. Our compressed JPEG image transfer only consumes 4% of the system resource, drops 0.3 points of accuracy and achieves 81,120 FPS which is 65.27 times faster than the conventional straightforward RGB data transfer. Thus, our proposed data transfer architecture is sufficient to increase system performance. As for the system throughput, our system is 3.84-34.41 times higher than existing FPGA implementations. Compared with the Xeon CPU, it achieves 138.38 times higher throughput, and it dissipates 1.2 times lower power, so its efficiency is 177.12 times better. Compared with the Tesla V100 GPU, it achieves 9.48 times higher throughput, dissipates 3.9 times lower power, and its efficiency is 37.52 times better. Thus, our parallel architecture on an FPGA provides superior throughput for the acceleration of a CNN. Hiroki Nakahara, Zhiqiang Que, Wayne Luk |
FCCM | 1 |
| 2020 | Optimizing Reconfigurable Recurrent Neural NetworksabstractThis paper proposes a novel latency-hiding hardware architecture based on column-wise matrix-vector multiplication to eliminate data dependency, improving the throughput of systems of RNN models. In addition, a flexible checkerboard tiling strategy is introduced to allow large weight matrices, while supporting element-based parallelism and vector-based parallelism. These optimizations improve the exploitation of the available parallelism to increase run-time hardware utilization and boost inference throughput. Furthermore, a quantization scheme with fine-tuning is proposed to achieve high accuracy. Evaluation results show that the proposed architecture can enhance performance and energy efficiency with little accuracy loss. It achieves 1.05 to 3.35 times better performance and 1.22 to 3.92 times better hardware utilization than a state-of-theart FPGA-based LSTM design, which shows that our approach contributes to high performance FPGA-based LSTM systems. Zhiqiang Que, Hiroki Nakahara, Eriko Nurvitadhi, Hongxiang Fan, Chenglong Zeng, Jiuxi Meng, Xinyu Niu, Wayne Luk |
FCCM | 2 |
| 2020 | R2CNN: Recurrent Residual Convolutional Neural Network on FPGAabstractOver the past years, feed-forward convolutional neural networks (CNNs) have evolved from a simple feed-forward architecture to deep and residual (skip-connection) architectures, demonstrating increasingly higher object categorization accuracy and increasingly better explanatory power of both neural and behavioral responses. However, from the neuroscientist point of view, the relationship between such deep architectures and the ventral visual pathway is incomplete. For example, current state-of-the-art CNNs appear to be too complex (e.g., now over 100 layers for ResNet) compared with the relatively shallow cortical hierarchy (4-8 layers). We introduce new CNNs with shallow recurrent architectures and skip connections requiring fewer parameters. With higher accuracy for classification, we propose an architecture for recurrent residual convolutional neural network (R2CNN) on FPGA, which efficiently utilizes on-chip memory bandwidth. We propose an Output-Kernel- Input-Parallel (OKIP) convolution circuit for a recurrent residual convolution stage. We implement the inference hardware on a Xilinx ZCU104 evaluation board with high-level synthesis. Our R2CNN accelerator achieves top-5 accuracy of 90.08% on ImageNet bench- mark, which has higher accuracy than conventional FPGA implementations. Hiroki Nakahara, Zhiqiang Que, Akira Jinguji, Wayne Luk |
FPGA | 1 |
| 2020 | An FPGA-Based Low-Latency Accelerator for Randomly Wired Neural NetworksabstractConvolutional neural networks (CNNs) are widely used for image tasks in both embedded systems and data centers. Particularly, when deploying CNNs in a data center, achieving high accuracy and low latency are important for various tasks such as the image processing of streaming videos. However, conventional CNN accelerators often use architectures with a deep pipeline, resulting in not only high throughput but also high latency. We propose an FPGA-based inference accelerator for randomly wired neural networks (RWNNs), whose layer structures are based on random graph models. Because RWNN can be processed in parallel, we can reduce the latency by concurrently using multiple computational units. We use the HBM2 to store feature maps, as multiple computing units need to simultaneously access different feature maps. In addition, the HBM channels and computational units are connected using a crossbar switch to efficiently transfer the feature maps. We allocate each layer to computational units using a simple heuristic algorithm. In addition, we allocate each layer to the HBM channels by coloring a conflict graph built based on the allocated schedule. This makes it possible for the computational units to access HBM channels in parallel. We implemented our architecture on an Alveo U50 FPGA and compared it with a FPGA-based inference accelerator that targets ResNet-50. In the ImageNet image classification task, we could process an image in 16.6 ms, which is 43% lower than that for a conventional accelerator. In addition, our accelerator could reduce the latency by 5.53 times compared with a CPU and 1.98 times compared with a GPU. Ryosuke Kuramochi, Hiroki Nakahara |
FPL | 2 |
| 2019 | An FPGA-based Fine Tuning Accelerator for a Sparse CNNabstractFine-tuning learns abundant feature expression for a wide range of natural images by using a pre-trained CNN model. It can be applied to a wide range of the neural network (NN)based computer vision problems. This paper proposes an FPGA-based fine-tuning accelerator for a sparse convolutional neural network (CNN). The proposed architecture consists of sparse convolutional units and pooling units with distributed stacks those are suitable for a sparse CNN. Additionally, this paper presents a fine-tuning scheme, which loads a pre-trained sparse CNN to reduce the memory size for the training step. Thus, our fine-tuning scheme stores all parameters on BRAMs and Ultra RAMs in the case of the Xilinx Virtex UltraScale+ FPGA to accelerate the training computation and reduce power consumption by eliminating energy-costly DRAM accesses. We implemented on a Xilinx Virtex UltraScale+ VC1525 acceleration development kit. Experimental results show that the proposed sparse finetuning accelerator on the FPGA can achieve four times faster, 2.9 times lower power consumption, and 11.6 times better performance per power, compared to the existing NVidia GTX1080Ti GPU. Hiroki Nakahara, Akira Jinguji, Masayuki Shimoda, Shimpei Sato |
FPGA | 1 |
| 2019 | Real-Time Multi-Pedestrian Detection in Surveillance Camera using FPGAabstractIn surveillance cameras, pedestrians and objects are detected using Convolutional Neural Network (CNN) based Object Detection such as YOLO and SSD. Since the size of the CNN input image is fixed, small objects cannot be detected when the high-resolution image is resized. By splitting image and applying object detection to each split image, CNN can detect small objects and take advantage of the high performance of the FPGA. We demonstrate an object detector by splitting the given image as stream data to YOLOv2 on an FPGA which has a very high performance. It achieved both practical speed and accuracy. Due to the difference in scale between training data and test data, the detection of small objects fails when the granularity of split is small. Splitting the image matches the scale of the data. When detection throughput is high enough, it detects many objects with a practical speed. We demonstrate the comparison of FPGA and mobile GPU in the proposed image split method of object detection. We implement Object Detection Systems by Xilinx ZCU104 FPGA board and NVIDIA Mobile GPU boards with a USB camera. The image split method can be adapted as it is to all implementations of object detection. We compare the performance of the FPGA and GPU. As a result of the experiment, FPGA achieves 547.0 FPS per image and is 3.9 times faster than Mobile GPU. When the given image is split into 4 × 4 grids, the system realizes 34.2 FPS and it satisfies the real-time requirement on the standard camera(30FPS). We showed that ultra-fast object detection can be used to improve accuracy. Akira Jinguji, Youki Sada, Hiroki Nakahara |
FPL | 3 |
| 2019 | FPGA-Based Training Accelerator Utilizing Sparseness of Convolutional Neural NetworkabstractTraining of convolutional neural networks (CNNs) is almost exclusively performed on large clusters of GPUs. However, it consumes vast amounts of power. Thus, high-speed training systems superior in low-power consumption are desired. This paper proposes an FPGA-based training accelerator utilizing a sparseness of a CNN, which consists of universal convolutional units and pooling units with distributed stacks. The proposed universal convolution architecture supports various convolution operations, such as the point-wise, depth-wise, large kernel and atrous convolutions used in the modern CNN. Additionally, we utilize a fine-tuning scheme, which loads a pre-trained dense CNN to reduce the memory size for the training process, while it considers important connectivity to preserve recognition accuracy. Our training scheme reduces 85% parameters to accelerate the training computation and reduce on-chip size. Thus, it eliminates energy-consuming DRAM accesses. We implemented the proposed training accelerator on a Xilinx Virtex UltraScale+ VC1525 acceleration development board. Experimental results show that the proposed sparse CNN training accelerator on the FPGA can achieve four times faster, 2.9 times lower power consumption, and 11.6 times better performance per power, compared to the existing NVIDIA RTX2080Ti GPU. Hiroki Nakahara, Youki Sada, Masayuki Shimoda, Kouki Sayama, Akira Jinguji, Shimpei Sato |
FPL | 1 |
| 2019 | An FPGA Implementation of Real-Time Object Detection with a Thermal CameraabstractWe demonstrate a sparse YOLOv2-based object detector with a thermal camera. A thermal camera outputs pixel values which represent heat (temperature), and the output is gray-scale images. Since the thermal cameras do not depend on whether there is the light or not unlike other visible range cameras, object detection using the thermal camera is reliable without ambient surrounding. This topic is of a broad interest in object surveillance and action recognition. However, since it is challenging to extract informative features from the thermal images, the implementation challenges of the object detector with high accuracy remain. In recent works, convolutional neural networks (CNNs) outperform conventional techniques, and a variety of object detectors based on the CNNs have been proposed. The representative networks are single-shot detectors that consist of a CNN and infer locations and classes simultaneously (e.g., SSD and YOLOv2). Although the primary advantage of the type is that it enables to train detection and classification simultaneously, the resulting increased computation time and area requirements can cause problems of implementation on an FPGA. Also, as for the proposed networks on RGB three channel images, one of the problems is false positive; the realization of more reliable object detector is required. To realize a real-time reliable object detector, we investigate an FPGA implementation of a sparse YOLOv2-based one whose inputs are four-channel images that consist of both the RGB and the thermal ones. In this demonstration, we show a performance comparison between an RGB-based detector and our proposed one on FPGAs. Masayuki Shimoda, Youki Sada, Ryosuke Kuramochi, Hiroki Nakahara |
FPL | 4 |
| 2018 | A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGAabstractA frame object detection problem consists of two problems: one is a regression problem to spatially separated bounding boxes, the second is the associated classification of the objects within realtime frame rate. It is widely used in the embedded systems, such as robotics, autonomous driving, security, and drones - all of which require high-performance and low-power consumption. This paper implements the YOLO (You only look once) object detector on an FPGA, which is faster and has a higher accuracy. It is based on the convolutional deep neural network (CNN), and it is a dominant part both the performance and the area. However, the object detector based on the CNN consists of a bounding box prediction (regression) and a class estimation (classification). Thus, the conventional all binarized CNN fails to recognize in most cases. In the paper, we propose a lightweight YOLOv2, which consists of the binarized CNN for a feature extraction and the parallel support vector regression (SVR) for both a classification and a localization. To our knowledge, this is the first time binarized CNN»s have been successfully used in object detection. We implement a pipelined based architecture for the lightweight YOLOv2 on the Xilinx Inc. zcu102 board, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC. The implemented object detector archived 40.81 frames per second (FPS). Compared with the ARM Cortex-A57, it was 177.4 times faster, it dissipated 1.1 times more power, and its performance per power efficiency was 158.9 times better. Also, compared with the nVidia Pascall embedded GPU, it was 27.5 times faster, it dissipated 1.5 times lower power, and its performance per power efficiency was 42.9 times better. Thus, our method is suitable for the frame object detector for an embedded vision system. Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Shimpei Sato |
FPGA | 1 |
| 2018 | A Demonstration of FPGA-Based You Only Look Once Version2 (YOLOv2)abstractWe implement the YOLO (You only look once) object detector on an FPGA, which is faster and has higher accuracy. It is based on the convolutional deep neural network (CNN), and it is a dominant part of both the performance and the area. It is widely used in the embedded systems, such as robotics, autonomous driving, security, and drones, all of which require high-performance and low-power consumption. A frame object detection problem consists of two problems: one is a regression problem to spatially separated bounding boxes, the second is the associated classification of the objects within realtime frame rate. We used the binary (1 bit) precision CNN for feature extraction and the half-precision (16 bit) precision CNN for both classification and localization. We implement a pipelined based architecture for the mixed-precision YOLOv2 on the Xilinx Inc. zcu102 board, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC. The implemented object detector archived 35.71 frames per second (FPS), which is faster than the standard video speed (29.9 FPS). Compared with a CPU and a GPU, an FPGA based accelerator was superior in power performance efficiency. Our method is suitable for the frame object detector for an embedded vision system. Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato |
FPL | 1 |
| 2018 | Demonstration of Object Detection for Event-Driven Cameras on FPGAs and GPUsabstractWe demonstrate an object detection system using a sliding window method for an event-driven camera[3] which outputs a subtracted frame (usually a binary value) when changes are detected in captured images. Fig. 1 shows the overall architecture of the object detector using a sliding window. First, our system extracts the object region candidates by applying a sliding window to the picture output by an event-driven camera. Since the event-driven camera outputs are binary, the sliding window determines whether an object exists or not based on the proportion of white pixels inside the bounding box as shown in Fig. 2. It reduces the number of proposed regions from 1,485 to average 10. The extracted pictures are resized to 40 × 40 size and applied to the ABCNN (All Binarized Convolutional Neural Network[7]), which is a BCNN (Binarized Convolutional Neural Network[4]) in which the first convolutional layer is done in binary form as well. All binarization is found to improve both the area and computation time, while the classification accuracy decreases. If the ABCNN infers the detected object as human, then it draws a box on the proposed region. Since more than one bounding box is commonly drawn against a single detected object, non-maximum suppression is used to reduce the extras to a single box. The proposed system is implemented on an FPGA and a GPU to show the comparison of them. Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara |
FPL | 3 |
| 2018 | An FPGA Realization of OpenPose Based on a Sparse Weight Convolutional Neural NetworkabstractThe OpenPose is a kind of a deep learning based pose estimator which achieved a top accuracy for multiple person pose estimations. Even if using the OpenPose, it is necessary to used high-performance GPU since it requires massive parameters access with high-bandwidth off-chip GDDR5 memories and a higher operation clock frequency. Thus, the power consumption becomes a critical issue to realization. Also, its computation time is slower than the current video standard frame speed (29.97 FPS). In the paper, we introduce a sparse weight CNN to reduce the amount of memory size for weights, which is Then, we offer the indirect memory access architecture to realize the sparse CNN convolutional operation efficiently. Also, to increase throughput further, we applied the six stages of pipeline architecture with a pipeline buffer memory realization. Our implementation satisfied the timing constraint for real-time applications. Since our architecture computed an image with 42.6 msec, the number of frames per second (FPS) was 23.43. We measured the total board power consumption: It was 55 Watt. Thus, the performance per power efficiency was 0.444 (FPS/W). Compared with the NVidia Titan X Pascal architecture GPU, it was 3.49 times faster, it dissipated 3.54 times lower power, and its performance per power efficiency was 13.05 times better. As far as we know, this work is the first FPGA implementation of the OpenPose. Akira Jinguji, Tomoya Fujii, Shimpei Sato, Hiroki Nakahara |
FPT | 4 |
| 2018 | A Tri-State Weight Convolutional Neural Network for an FPGA: Applied to YOLOv2 Object DetectorabstractA frame object detection, such as the YOLO (You only look once), is used in embedded vision systems, such as a robot, an automobile, a security camera, and a drone. However, it requires highly performance-per-power detection by an inexpensive device. In the paper, we propose a tri-state weight CNN, which is a generalization of a low-precision and sparse (pruning) for CNN weight. In the former part, we set a weight {-1,0,+1} as a ternary CNN, while in the latter part, we set a {-w,0,+w} as a sparse weight CNN. The proposed tri-state CNN is a kind of a mixed-precision one, which is suitable for an object detector consisting of a bounding box prediction (regression) and a class estimation (classification). We apply an indirect memory access architecture to skip zero part and propose the weight parallel 2D convolutional circuit. It can efficiently be applied to the AlexNet based CNN, which has different size kernels. We design the AlexNet based YOLOv2 to reduce the number of layers toward low-latency computation. In the experiment, the proposed tri-state scheme CNN reduces the memory size for weight by 92%. We implement the proposed tri-state weight YOLOv2 on the AvNet Inc. UltraZed-EG starter kit, which has the Xilinx Inc. Zynq Ultrascale+ MPSoC ZU3EG. It archived 61.70 frames per second (FPS), which exceeds the standard video frame rate (29.97 FPS). Compared with the ARM Cortex-A57, it was 268.2 times faster, and its performance per power efficiency was 313.51 times better. Also, compared with the NVidia Pascal embedded GPU, it was 4.0 times faster, and its power performance efficiency was 11.35 times better. Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato |
FPT | 1 |
| 2018 | A High-speed Low-power Deep Neural Network on an FPGA based on the Nested RNS: Applied to an Object DetectorabstractA pre-trained convolutional deep neural network (CNN) is the feed-forward computation perspective, and it is widely used for the embedded vision systems. One of the applications of the CNN is a frame object detection problem. It is widely used in the embedded systems, such as a robot, an automobile, a security camera, and a drone, that require a highly performance-power efficient device. In the CNN, the 2D convolutional operation occupies more than 90time. Since the 2D convolutional operation performs massive multiply-accumulation (MAC) operations, conventional realizations could not implement a fully parallel CNN. The RNS decomposes an integer into a tuple of integers by residues of moduli set. Since no pair of modulus has a common factor with any other, the conventional RNS decomposes the MAC unit into circuits with different sizes means that the RNS could not utilize resources of an FPGA with uniform size. In this paper, we use the nested RNS (NRNS), which recursively decompose the RNS. It can decompose the MAC unit into circuits with small sizes. In the CNN using the NRNS, a MAC unit is decomposed into 4-bit ones realized by look-up tables of the FPGA. Thus, it leads to a high clock frequency with less hardware. We designed the Tiny YOLOv2 for the practical object detection, and it using the CNN based on the NRNS is implemented on a Digilent NetFPGA-SUME FPGA board. Compared with the NVidia GTX1080Ti (Pascal architecture) for the designed Tiny YOLOv2, the FPGA using the NRNS was 3.19 times better than the GPU as for the performance-power efficiency. Hiroki Nakahara, Tsutomu Sasao |
ISCAS | 1 |
| 2017 | A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only)
Hiroki Nakahara, Haruyoshi Yonekawa, Hisashi Iwamoto, Masato Motomura |
FPGA | 1 |
| 2017 | A fully connected layer elimination for a binarizec convolutional neural network on an FPGAabstractA pre-trained convolutional deep neural network (CNN) is widely used for embedded systems, which requires highly power-and-area efficiency. In that case, the CPU is too slow, the embedded GPU dissipates much power, and the ASIC cannot keep up with the rapidly progress of the CNN variations. This paper uses a binarized CNN which treats only binary 2-values for the inputs and the weights. Since the multiplier is replaced into an XNOR circuit, we can realize a high-performance MAC circuit by using many XNOR circuits. In the paper, we eliminate internal FC layers excluding the last one, then, insert a binarized average pooling layer, which can be realized by a majority circuit for binarized (1/0) values. In that case, since the weight memory is replaced into the 1's counter, we can realize a compact and faster CNN than the conventional ones. We implemented the VGG-11 benchmark CNN for the CIFAR10 image classification task on the Xilinx Inc. Zedboard. Compared with the conventional binarized implementations on an FPGA, the classification accuracy was almost the same, the performance per power efficiency is 5.1 better, as for the performance per area efficiency, it is 8.0 times better, and as for the performance per memory, it is 8.2 times better. Hiroki Nakahara, Tomoya Fujii, Shimpei Sato |
FPL | 1 |
| 2017 | An object detector based on multiscale sliding window search using a fully pipelined binarized CNN on an FPGAabstractAn object detection problem consists of two problems: one is classification of detected object category and the other is localization. Frame object detection is used in an embedded vision systems, such as a robot, an automobile, a security camera, and a drone. These applications require high-performance computation and low-power consumption by an inexpensive device. This paper proposes multiscale sliding window based object detector using a fully pipelined binarized deep convolutional neural network (BCNN) on an FPGA. It consists of a sliding window part, a fully pipelined BCNN classifier, and an ARM processing unit for detection. Duplicate detections were filtered by using a non-maximum suppression algorithm running on the ARM processor. We propose the fully pipelined layers for the BCNN and its architecture for FPGA realization. Since the proposed BCNN circuit uses on-chip memories on the FPGA, its throughput is higher than a GPU based one with practical recognition accuracy. We trained the VGG11 based BCNN using the KITTI vision benchmark for the car detection scenario. Then, we implemented the proposed object detector on the Xilinx Inc. Zynq UltraScale+ MPSoC zcu102 evaluation board. The GPU based object detectors were too slow for the realtime application requirement (HD frame rate), with the exception of YOLOv2. As compared with the GPU implementation of YOLOv2, the proposed FPGA detector had higher recognition accuracy and lower power consumption. Compared with the YOLOv2, the proposed FPGA one is higher with respect to recognition accuracy, and its power consumption is lower than the GPU based YOLOv2. Thus, the FPGA based object detector suitable for the embedded realtime applications. Hiroki Nakahara, Haruyoshi Yonekawa, Shimpei Sato |
FPT | 1 |
| 2017 | All binarized convolutional neural network and its implementation on an FPGAabstractA pre-trained convolutional neural network (CNN) is a feed-forward computation perspective, which is widely used in the embedded systems requiring highly power-and-area efficiency. This paper realizes a binarized CNN which treats only binarized values (+1/-1) for the weights and the activation value. In this case, the multiplier is replaced by an XNOR circuit instead of a dedicated DSP block. Binarization for both weights and activation are more suitable for hardware implementation. However, the first convolutional layer still calculates in integer precision, since the input value is 8 bit RGB pixel, and not binarized. In this paper, we decompose the input value into maps of which each pixel is in 1-bit precision. The proposed method enables a binarized CNN to use bitwise operation in all layers, and shares a binarized convolutional circuit among all convolutional layers. We call this all binarized CNN. We compared our proposal with conventional ones. Since all binarized CNN do not require a dedicated DSP block, our proposal is smaller and 1.2 times faster than the typical CNNs, and almost maintains baseline classification accuracy. In addition, pipelined all binarized CNN achieved 1840 FPS, consumed 0.3 watts and its accuracy was 82.8%. Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara |
FPT | 3 |
| 2016 | An acceleration of a random forest classification using Altera SDK for OpenCLabstractA random forest (RF) is a kind of ensemble machine learning algorithm used for a classification and a regression. It consists of multiple decision trees that are built from randomly sampled data. The RF has a simple, fast learning, and identification capability compared with other machine learning algorithms. It is widely used for applicable to various recognition systems. Since it is necessary to un-balanced trace for each tree and requires communication for all the ones, the random forest is not suitable in SIMD architectures such as GPUs. Although the accelerators using the FPGA have been proposed, such implementations were based on HDL design. Thus, they required longer design time compared with the soft-ware based realizations. In this paper, we show the accelerator for the RF using the Altera SDK for OpenCL, which is a kind of high-level synthesis. To accelerate the RF classification, we propose the fully pipelined architecture to increase the memory bandwidth using on-chip memories on the FPGA. Also, we apply appropriate bit fixed point representation instead of 32 bit floating point one in order to reduce the hardware size, power consumption, and increase the memory bandwidth. We implemented the RF on the Terasic Corp. DE5-NET FPGA board, and compared with the CPU and the GPU implementations, As for the LPS (lookups per second), the FPGA realization was 10.7 times faster than the GPU one, and it was 14.0 times faster than the CPU one. As for the LPS per power consumption, the FPGA realization was 61.3 times better than the GPU one, and it was 12.1 times better than the CPU one. Hiroki Nakahara, Akira Jinguji, Tomonori Fujii, Simpei Sato |
FPT | 1 |
| 2016 | A memory-based realization of a binarized deep convolutional neural networkabstractA pre-trained deep convolutional neural network (CNN) is a feed-forward computation perspective, which is widely used for the embedded systems, requires high power-and-area efficiency. This paper realizes a binarized CNN which treats only binary 2-values (+1/-1) for the inputs and the weights. In this case, the multiplier is replaced with an EX-NOR circuit. To reduce both power and area, we realize the 2-valued CNN by off- and on-chip memories. Since our 2D convolution operations are realized by the on-chip memory, our implementation consumes lower power than the DSP block. We decompose the memory part, and realize them by a cascade of memories (LUT cascade). By introducing a batch normalization technique, the classification error for the binarized CNN can be improved. We implemented the CIFAR-10 benchmark on the NetFPGA-SUME board, which has the Xilinx Inc. Virtex 7 FPGA and three on-chip QDR II+ Synchronous SRAMs. Compared with the conventional FPGA realizations, the performance is 2.82 times faster, the power efficiency is 1.76 times, and the area efficiency is 29.13 times better. Hiroki Nakahara, Haruyoshi Yonekawa, Tsutomu Sasao, Hisashi Iwamoto, Masato Motomura |
FPT | 1 |
| 2015 | A deep convolutional neural network based on nested residue number systemabstractA pre-trained deep convolutional neural network (DCNN) is the feed-forward computation perspective which is widely used for the embedded vision systems. In the DCNN, the 2D convolutional operation occupies more than 90% of the computation time. Since the 2D convolutional operation performs massive multiply-accumulation (MAC) operations, conventional realizations could not implement a fully parallel DCNN. The RNS decomposes an integer into a tuple of L integers by residues of moduli set. Since no pair of modulus have a common factor with any other, the conventional RNS decomposes the MAC unit into circuits with different sizes. It means that the RNS could not utilize resources of an FPGA with uniform size. In this paper, we propose the nested RNS (NRNS), which recursively decompose the RNS. It can decompose the MAC unit into circuits with small sizes. In the DCNN using the NRNS, a 48-bit MAC unit is decomposed into 4-bit ones realized by look-up tables of the FPGA. In the system, we also use binary to NRNS converters and NRNS to binary converters. The binary to NRNS converter is realized by on-chip BRAMs, while the NRNS to binary one is realized by DSP blocks and BRAMs. Thus, a balanced usage of FPGA resources leads to a high clock frequency with less hardware. The ImageNet DCNN using the NRNS is implemented on a Xilinx Virtex VC707 evaluation board. As for the performance per area GOPS (Giga operations per second) per a slice, the proposed one is 5.86 times better than the existing best realization. Hiroki Nakahara, Tsutomu Sasao |
FPL | 1 |
| 2013 | A packet classifier using LUT cascades based on EVMDDS (k)abstractThis paper presents a packet classifier using multiple LUT cascades based on edge-valued multi-valued decision diagrams (EVMDDs (k)). First, a set of rules for a packet classifier is partitioned into groups. Second, they are decomposed into field functions and Cartesian product functions. Third, they are represented by EVMDDs (k), and finally, they are converted to LUT cascades using adders. We implemented the proposed circuit on a Virtex 7 VC707 evaluation board. The system throughput is 345.60 Gbps for minimum packet size (40 Bytes). As for the normalized throughput (efficiency), the proposed one is 7.14 times better than existing FPGA implementations. Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura |
FPL | 1 |
| 2013 | A high-speed FFT based on a six-step algorithm: Applied to a radio telescope for a solar radio burstabstractA radio telescope analyzes the radio frequency (RF) received from celestial objects. It consists of an antenna, a receiver, and a spectrometer. The spectrometer converts the time domain into the frequency domain with an FFT operation. A solar radio burst observation requires a high-speed FFT. This paper proposes a P parallel N point FFT for fixed point data based on a six-step algorithm. We analyze the hardware resources for the P parallel N point FFT. We implemented 32 parallel N point FFT circuits on a Xilinx Virtex 7 VC707 board. Comparison with the existing FFT implementations shows that the proposed one is 4.52–22.64 times faster. Hiroki Nakahara, Kazumasa Iwai, Hiroyuki Nakanishi |
FPT | 1 |
| 2012 | On a Wideband Fast Fourier Transform Using Piecewise Linear Approximations: Application to a Radio Telescope Spectrometer
Hiroki Nakahara, Hiroyuki Nakanishi, Tsutomu Sasao |
ICA3PP (1) | 1 |
| 2010 | A Packet Classifier Using a Parallel Branching Program MachineabstractA branching program machine (BM) is a special purpose processor that uses only two kinds of instructions: Branch and output instructions. Thus, the architecture for the BM is much simpler than that for a general purpose processor (MPU). Since the BM uses the dedicated instructions for a special purpose application, it is faster than the MPU. This paper presents a packet classifier using a parallel branching program machine (PBM). To reduce computation time and code size, first, a set of rules for the packet classifier is partitioned into groups. Then, they are evaluated by the PBM in parallel. Also, this paper shows a method to estimate the number of necessary BMs to realize the packet classifier. The PBM32 consisting of 32 BMs has been implemented on an FPGA, and compared with the Intel's [email protected]. The PBM32 is 8.1-11.1 times faster than the Core2Duo, and the PBM32 requires only 0.2-10.3 percent of the memory for the Core2Duo. Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura |
DSD | 1 |
| 2009 | The Parallel Sieve Method for a Virus Scanning EngineabstractThis paper shows a new architecture for a virus scanning system, which is different from that of an intrusion detection system. The proposed method uses two-stage matching: In the first stage, a hardware filter quickly scans the text to find partial matches, and in the second stage, the MPU scans the text to find a total match in the ClamAV 514,287 virus pattern set. To make the hardware filter simple, we use a finite-input memory machine (FIMM). To reduce the memory size of the FIMM, we introduce the parallel sieve method. The proposed method is memorybased, so it is quickly reconfigurable and dissipates lower power than a TCAM-based method. The system is implemented on the Stratix III FPGA with three off-chip SRAMs and an SDRAM, where all ClamAV 514,287 virus patterns are stored. Compared with existing methods, our method achieves 1.41-31.36 times more efficient area-throughput ratio. Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura |
DSD | 1 |
| 2009 | A virus scanning engine using a parallel finite-input memory machine and MPUsabstractThis paper presents a virus scanning engine. After showing the difference between ClamAV (an anti-virus software) and SNORT (an intrusion detection software), we show a new architecture for the virus scanning engine, which is different from that of the intrusion detection engine. The new architecture consists of a parallel finite-input memory machine (PFIMM) and general purpose MPUs. It uses two-stage matching. That is, in the first stage, the parallel hardware filter quickly scans the text to find partial matches, and in the second stage, the MPU scan the text to find the total match. To reduce the memory size, compressed match vectors are used. The system is implemented on the Stratix III FPGA, where 65,536 ClamAV virus patterns are stored. As for the area-performance ratio, our system is 1.2-26.3 times more efficient than existing ones. Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura |
FPL | 1 |
| 2007 | Implementations of Reconfigurable Logic Arrays on FPGAsabstractThis paper presents a method to implement a reconfigurable logic array on an FPGA. To design circuits with 2-valued k-input LUTs, 2k-valued logic is introduced. Standard benchmark functions as well as symmetric functions are efficiently implemented by a logic array with 2k-valued variables. Number of products and number of bits to represent functions by the expressions with 2k-valued variables for k = 1,2,3,4, and 5 are compared. Both sum-of-products expressions and EXOR sum-of-products expressions of 2k-valued logic significantly reduces needed FPGA resources, when 2 les k les 5. Experimental results for benchmark functions and symmetric functions are shown. Implementations of arrays with 16-valued variables on Xilinx and Altera FPGAs are also shown. Tsutomu Sasao, Hiroki Nakahara |
FPT | 2 |
| 2007 | A CAM Emulator Using Look-Up Table CascadesabstractAn address table relates k different registered vectors to the addresses from 1 to k. An address generation function represents the address table. This paper presents a realization of an address generation function with an LUT cascade on an FPGA. The address generation function is implemented by BRAMs of an FPGA, while the addition and the deletion of registered vectors are implemented by an embedded processor on the FPGA. Compared with CAMs produced by the Xilinx Core Generator, our implementations are smaller and faster. This paper also shows that the addition and deletion of a registered vector can be done in time that is proportional to the number of cells in the LUT cascade. Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura |
IPDPS | 1 |
| 2006 | A fast logic simulator using a look up table cascade emulatorabstractThis paper shows a new type of a cycle-based logic simulation method using a look-up table (LUT) cascade emulator. The method first transforms a given circuit into LUT cascades through BDD (binary decision diagram). Then, it stores LUT data to the memory of an LUT cascade emulator. Next, it generates the C code representing the control circuit of the LUT cascade emulator. And, finally, it converts the C code into the execution code. This method is compared with a levelized compiled code (LCC) simulator with respect to the simulation time and setup time. Although we used standard PC to simulate the circuit, experimental results show that this method is 12-64 times faster than the LCC Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura |
ASP-DAC | 1 |
| 2006 | A Soft Error Tolerant LUT Cascade EmulatorabstractAn LUT cascade emulator realizes an arbitrary sequential circuit. Given a sequential circuit, we convert the combinational part into one or more LUT cascades, and store LUT (cell) data into a memory in the LUT cascade emulator. The emulator evaluates multi-output logic functions by reading cell data sequentially. To improve the tolerance to soft errors, cell data in the memory are encoded by error correcting codes. Also, error-correcting circuits and checking circuits that periodically scan the memories are appended. When a soft error is detected, it removes the error by rewriting the correct data into the memory. To mask soft errors in flip-flops, a TMR (triple module redundancy) technique is employed. Our system detects a soft error in a single bit. Also, the mission time of the system is more than 1000times of time of an ordinary LUT cascade emulator Hiroki Nakahara, Tsutomu Sasao |
ATS | 1 |