Yuichiro Shibata

dblp:76/990 · DBLP profile ↗
← Back
50ranked-venue papers
1as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2023 A Mobile-Oriented GPU Implementation of a Convolutional Neural Network for Object Detection
Yasutoshi Araki, Takuho Kawazu, Taito Manabe, Yoichi Ishizuka, Yuichiro Shibata
CISIS5
2023 Efficient FPGA Implementation of a Convolutional Neural Network for Surgical Image Segmentation Focusing on Recursive Structure
Takehiro Miura, Shuto Abe, Taito Manabe, Yuichiro Shibata, Taiichiro Kosaka, Tomohiko Adachi
CISIS4
2022 FPGA Implementation of an Object Recognition System with Low Power Consumption Using a YOLOv3-tiny-based CNN
Yasutoshi Araki, Masatomo Matsuda, Taito Manabe, Yoichi Ishizuka, Yuichiro Shibata
CISIS5
2022 A Lane Detection Hardware Algorithm Based on Helmholtz Principle and Its Application to Unmanned Mobile Vehicles
abstract
We are developing an SoC FPGA-based unmanned mobile vehicle for the FPGA design competition. For the vehicle to follow roads successfully, it must be able to detect not only straight lines but also curved lines accurately. Therefore, we implemented a lane detection algorithm that is robust not only against straight lines but also against curves to improve driving performance. We implemented an autonomous driving system employing this algorithm on Digilent Zybo Z7-20. We evaluated the lane detection algorithm based on simulations and showed that this algorithm can reduce false detection of lane features compared to the classical Canny filter.
Katsuaki Kamimae, Shintaro Matsui, Yasutoshi Araki, Takehiro Miura, Keigo Motoyoshi, Keizo Yamashita, Haruto Ikehara, Takuho Kawazu, Huang Yuwei, Masahiro Nishimura, Shuto Abe, Kenyu Okino, Yuta Hashiguchi, Koki Fukuda, Kengo Yanagihara, Taito Manabe, Yuichiro Shibata
FPT17
2022 FPGA implementation of HDR synthesis processing with image compression techniques
abstract
This paper presents an FPGA implementation of real-time high dynamic range (HDR) synthesis, which expresses a wide dynamic range by combining multiple images with different exposures using image pyramids. We have implemented a pipeline that performs streaming processing on images without using external memory. However, implementation for high-resolution images has been difficult due to large memory usage for line buffers. Therefore, we propose an image compression algorithm based on adaptive differential pulse code modulation (ADPCM). Compression modules based on the algorithm can be easily integrated into the pipeline. When the image resolution is 4K and the pyramid depth is 7, memory usage can be halved from 168.48 % to 84.32 % by introducing the compression modules, resulting in better quality.
Masahiro Nishimura, Yuta Imamura, Taito Manabe, Yuichiro Shibata
FPT4
2021 SoC FPGA implementation of an unmanned mobile vehicle with an image transmission system over VNC
abstract
We are developing the unmanned mobile vehicle implemented on SoC FPGA for the FPGA design competition. For highly productive development of image-based self-driving mobile vehicles, a remote verification and debugging environment with real-time image transmission is important. This paper presents an image transmission system with which we can monitor onboard camera images of the vehicle and feature detection results over VNC Wi-Fi connection. We implemented the whole system on a Xilinx Zynq-7000 with a maximum operating frequency of 125 MHz. The evaluation of the system showed that the resolution of 640×720 is the most beneficial for VNC in this experiment in terms of the performance ratio of VNC to SSH X11 forwarding. We also shortly describe other components to be used to develop the autonomous driving system in this paper.
Keigo Motoyoshi, Yuta Imamura, Taichi Saikai, Koki Fujita, Daiki Furukawa, Masatomo Matsuda, Tatsuma Mori, Yasutoshi Araki, Takehiro Miura, Keizo Yamashita, Haruto Ikehara, Kaito Ohira, Katsuaki Kamimae, Takuho Kawazu, Masahiro Nishimura, Shintaro Matsui, Koki Tomonaga, Taito Manabe, Yuichiro Shibata
FPT19
2019 A Self-partial Reconfiguration Framework with Configuration Data Compression for Intel FPGAs
Shota Fukui, Yuichi Kawamata, Yuichiro Shibata
CISIS3
2019 A Simple Heterogeneous Redundant Design Method for Finite State Machines on FPGAs
Takanori Itagawa, Ryo Kamasaka, Yuichiro Shibata
CISIS3
2019 Pipelined FPGA Implementation of a Wave-Front-Fetch Graph Cut System
Naofumi Yoshinaga, Ryo Kamasaka, Yuichiro Shibata, Kiyoshi Oguri
CISIS3
2018 Storing and Compressing Video into Neural Networks by Overfitting
Hiroki Egawa, Yuichiro Shibata
CISIS2
2018 FPGA Implementation of Lightweight Communication Protocol Processing for IoT
Ryouhei Tsugami, Yuichiro Shibata
CISIS2
2018 Light-Weight Fine-Grain Dynamic Partial Reconfiguration on Xilinx FPGAs
Kunihiro Ueda, Keisuke Dohi, Yuichiro Shibata
CISIS3
2018 Discussion on High Level Synthesis FPGA Design of Camera Calibration
Kazuya Uetsuhara, Hiroki Nagayama, Yuichiro Shibata, Kiyoshi Oguri
CISIS3
2017 HLS-Based FPGA Acceleration of Building-Cube Stencil Computation
Rie Soejima, Yuichiro Shibata, Kiyoshi Oguri
CISIS2
2017 Power Performance Analysis of FPGA-Based Particle Filtering for Realtime Object Tracking
Akane Tahara, Yoshiki Hayashida, Theint Theint Thu, Yuichiro Shibata, Kiyoshi Oguri
CISIS4
2017 FPGA implementation of a real-time super-resolution system with a CNN based on a residue number system
abstract
A super-resolution technology is used for filling the gap between high-resolution displays and lower-resolution images. One of various algorithms to interpolate the lost information is to use a convolutional neural network (CNN). This paper shows an FPGA implementation and a performance evaluation of our CNN-based super-resolution system, which can process moving images in real time. We apply horizontal and/or vertical flips to input images instead of pre-enlargement. This method prevents information loss and enables the network to make the best use of its input size. In addition, we adopted the residue number system (RNS) to reduce resource utilization. The proposed system can perform super-resolution from 960×540 to 1920×1080 at 60fps with a latency of less than 1ms. In spite of resource restriction of the FPGA, the system generates clear super-resolution images with smooth edges. The evaluation results also revealed the superior quality in terms of the peak signal-to-noise ratio (PSNR), compared to other systems using pre-enlargement.
Taito Manabe, Yuichiro Shibata, Kiyoshi Oguri
FPT2
2016 FPGA implementation of a real-time super-resolution system using a convolutional neural network
abstract
Super-resolution technologies are used to fill the gap between high-resolution displays and lower-resolution contents. There are various algorithms to interpolate information, one of which is using a convolutional neural network (CNN). This paper shows FPGA implementation and performance evaluation of a CNN-based super-resolution system, which can process moving images in real time. We apply horizontal and/or vertical flips to network input images instead of commonly used pre-enlargement techniques. This method prevents information loss and enables the network to utilize the best of its input image size. Our system can perform super-resolution from 960×540 pixels to 1920×1080 pixels at not less than 48fps with a latency of less than 1 ms. Even though the network scale and the size of filters are limited due to resource restriction of the FPGA, the system generates clear super-resolution images with smooth edges. The evaluation results also reveal that the proposed system achieves superior quality in terms of the structural similarity (SSIM) index, compared to other systems using pre-enlargement.
Taito Manabe, Yuichiro Shibata, Kiyoshi Oguri
FPT2
2014 A soft-core processor for finite field arithmetic with a variable word size accelerator
abstract
This paper presents implementation and evaluation of an accelerator architecture for soft-cores to speed up reduction process for the arithmetic on GF(2m) used in Elliptic Curve Cryptography (ECC) systems. In this architecture, the word size of the accelerator can be customized when the architecture is configured on an FPGA. Focusing on the fact that the number of the reduction processing operations on GF(2m) is affected by the irreducible polynomial and the word size, we propose to employ an unconventional word size for the accelerator depending on a given irreducible polynomial and implement a MIPS-based soft-core processor coupled with a variable-word size accelerator. As a result of evaluation with several polynomials, it was shown that the performance improvement of up to 10.2 times was obtained compared to the 32-bit word size, even taking into account the maximum frequency degradation of 20.4% caused by changing the word size. The advantage of using unconventional word sizes was also shown, suggesting the promise of this approach for low-power ECC systems.
Aiko Iwasaki, Keisuke Dohi, Yuichiro Shibata, Kiyoshi Oguri, Ryuichi Harasawa
FPL3
2012 Deep-pipelined FPGA implementation of ellipse estimation for eye tracking
abstract
This paper presents a deep-pipelined FPGA implementation of real-time ellipse estimation for eye tracking. The system is constructed by the Starburst algorithm on a stream-oriented architecture and the RANSAC algorithm without any external memories. In particular, the paper presents comparative results between three different hypothesis generators for the RANSAC algorithm based on Cramer's rule, Gauss-Jordan elimination and LU decomposition. Comparison criteria include resource usage, throughput and energy consumption. The result shows that the three implementations have different characteristics and the optimal algorithm needs to be chosen depending on the amount of resources on FPGAs and required performance.
Keisuke Dohi, Yuma Hatanaka, Kazuhiro Negi, Yuichiro Shibata, Kiyoshi Oguri
FPL4
2011 Pattern Compression of FAST Corner Detection for Efficient Hardware Implementation
abstract
This paper shows stream-oriented FPGA implementation of the machine-learned Features from Accelerated Segment Test (FAST) corner detection, which is used in the parallel tracking and mapping (PTAM) for augmented reality (AR). One of the difficulties of compact hardware implementation of the FAST corner detection is a matching process with a large number of corner patterns. We propose corner pattern compression methods focusing on discriminant division and pattern symmetry for rotation and inversion. This pattern compression enables implementation of the corner pattern matching with a combinational circuit. Our prototype implementation achieves real-time execution performance with 7-9% of available slices of a Virtex-5 FPGA.
Keisuke Dohi, Yuji Yorita, Yuichiro Shibata, Kiyoshi Oguri
FPL3
2011 Deep pipelined one-chip FPGA implementation of a real-time image-based human detection algorithm
abstract
In this paper, deep pipelined FPGA implementation of a real-time image-based human detection algorithm is presented. By using binary patterned HOG features, AdaBoost classifiers generated by offline training, and some approximation arithmetic strategies, our architecture can be efficiently fitted on a low-end FPGA without any external memory modules. Empirical evaluation reveals that our system achieves 62.5 fps of the detection throughput, showing 96.6% and 20.7% of the detection rate and the false positive rate, respectively. Moreover, if a highspeed camera device is available, the maximum throughput of 112 fps is expected to be accomplished, which is 7.5 times faster than software implementation.
Kazuhiro Negi, Keisuke Dohi, Yuichiro Shibata, Kiyoshi Oguri
FPT3
2011 Steering Time-Dependent Estimation of Posteriors with Hyperparameter Indexing in Bayesian Topic Models
Tomonari Masada, Atsuhiro Takasu, Yuichiro Shibata, Kiyoshi Oguri
PAKDD (1)3
2010 Highly efficient mapping of the Smith-Waterman algorithm on CUDA-compatible GPUs
abstract
This paper describes a multi-threaded parallel design and implementation of the Smith-Waterman (SW) algorithm on graphic processing units (GPUs) with NVIDIA corporation's Compute Unified Device Architecture (CUDA). Central to this is a divide and conquer approach which divides the computation of a whole pairwise sequence alignment matrix into multiple sub-matrices (or parallelograms) each running efficiently on the available hardware resources of the GPU in hand, with temporary intermediate data stored in global memory. Moreover, we use thread warps and padding techniques in order to decrease the cost of thread synchronization, as well as loop unrolling in order to reduce the cost of conditional branches. While intermediate data is stored in global memory for large queries, the most inner loop in our implementation will only access shared memory and registers. As a result of these optimizations, our implementation of the SW algorithm achieves a throughput ranging between 9.09 GCUPS (Giga Cell Update per Second) and 12.71 GCUPS on a single-GPU version, and a throughput between 29.46 GCUPS and 43.05 GCUPS on a quad-GPU platform. Compared with the best GPU implementation of the SW algorithm reported to date, our implementation achieves up to 46 % improvement in speed. The source code of our implementation is available in the public domain for Bioinformaticians to benefit from its performance.
Keisuke Dohi, Khaled Benkrid, Cheng Ling, Tsuyoshi Hamada, Yuichiro Shibata
ASAP5
2010 A datapath classification method for FPGA-based scientific application accelerator systems
abstract
Resource reduction design techniques play an important role to implement large-scale FPGA-based accelerator systems in floating point applications since available resources on FPGAs are limited. This paper proposes a dataflow graph classification method which makes groups of graphs based on their similarity in order to bring out efficient graph combining. Aiming at finding effective parameters for the k-means algorithm, various parameter combinations are evaluated and compared in terms of resource reduction effects and performance. The experimental results using an FPGA-based biochemical simulator reveal that the graph clustering that uses information on the maximum common subgraphs achieve 73.3% of resource reduction rate while alleviating the performance degradation.
Yui Ogawa, Tomonori Ooya, Yasunori Osana, Masato Yoshimi, Yuri Nishikawa, Akira Funahashi, Noriko Hiroi, Hideharu Amano, Yuichiro Shibata, Kiyoshi Oguri
FPT9
2010 Modeling Topical Trends over Continuous Time with Priors
Tomonari Masada, Daiji Fukagawa, Atsuhiro Takasu, Yuichiro Shibata, Kiyoshi Oguri
ISNN (2)4
2009 Bayesian Multi-topic Microarray Analysis with Hyperparameter Reestimation
Tomonari Masada, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri
ADMA3
2009 Dynamic hyperparameter optimization for bayesian topical trend analysis
abstract
This paper presents a new Bayesian topical trend analysis. We regard the parameters of topic Dirichlet priors in latent Dirichlet allocation as a function of document timestamps and optimize the parameters by a gradient-based algorithm. Since our method gives similar hyperparameters to the documents having similar timestamps, topic assignment in collapsed Gibbs sampling is affected by timestamp similarities. We compute TFIDF-based document similarities by using a result of collapsed Gibbs sampling and evaluate our proposal by link detection task of Topic Detection and Tracking.
Tomonari Masada, Daiji Fukagawa, Atsuhiro Takasu, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri
CIKM5
2009 Configuring area and performance: Empirical evaluation on an FPGA-based biochemical simulator
abstract
One of the obvious advantages of FPGA-based reconfigurable computing is customizability of a tradeoff point between performance and hardware costs. However, this tradeoff has rarely been discussed in a whole application level, which is the most important view for application users. This paper presents empirical evaluation of a hardware module sharing technique which can shift a tradeoff point of area and performance on an FPGA-based biochemical simulator. The biochemical simulation results are discussed in terms of hardware costs, simulation throughput, parallelism extracted in simulation hardware, and data transfer overheads.
Tomonori Ooya, Hideki Yamada, Tomoya Ishimori, Yuichiro Shibata, Yasunori Osana, Kiyoshi Oguri, Masato Yoshimi, Yuri Nishikawa, Akira Funahashi, Noriko Hiroi, Hideharu Amano
FPL4
2009 Accelerating Collapsed Variational Bayesian Inference for Latent Dirichlet Allocation with Nvidia CUDA Compatible Devices
Tomonari Masada, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri
IEA/AIE3
2008 Retrieving 3-d information with FPGA-based stream processing
abstract
The traditional computer system has the feature where data processing proceeds as the data goes repeatedly to and from the microprocessor and the memory. However, in this paper, attention is paid to stream processing, where the input data from the sensor goes through the dedicated circuits and becomes the end result. As dedicated circuits depend on the process and its progress, the reconfigurable hardware is necessary for stream processing. Almost ten years ago, we proposed LSI architecture named PCA (Plastic Cell Architecture), in which a circuit can configure another circuit and can operate together [1]. Recently we have finished the architectural design of PCA. We are therefore evaluating the stream processing itself by using FPGA. The stream processing is thought of as the brain type processing, so we chose computer vision. Currently, computer vision is thought as the technology to retrieve 3-D information from the multi set of 2-D data which are stored in the memory in advance from cameras. Negating the use of random access memory, we can get the ideas of stream processing. Pipelined template matching is one of those ideas. In the paper, we present the new idea of stereo matching. Two streams of pixel data from the completely synchronized two cameras in epipolar constraint are directly inserted into two delay lines which are followed by a lots of correlation checkers. This structure enables both stereo matching and feature points detecting
Hidenori Matsubayashi, Shinsuke Nino, Toru Aramaki, Yuichiro Shibata, Kiyoshi Oguri
FPGA4
2008 An optimization method of DMA transfer for a general purpose reconfigurable machine
abstract
DMA transfer between a CPU and an FPGA often becomes a bottleneck of current reconfigurable machines. The DMA transfer of the machines like SRC-6 supports streaming processing with on-board memory interleaving, but as a pre-processing of the interleaving, the CPU must reorder the data for applications with severe FPGA resource constraints. This paper empirically evaluates this overhead to reveal the trade-off point. The results show that a speedup is achieved by interleaved streaming DMA when 150KB or lower data strings are transferred.
Sayaka Shida, Yuichiro Shibata, Kiyoshi Oguri, Duncan A. Buell
FPL2
2008 Practical implementation of a network-based stochastic biochemical simulation system on an FPGA
abstract
Stochastic simulation of biochemical reaction networks are widely focused by life scientists to represent stochastic behaviors in cellular processes. Stochastic algorithm has loop-and thread-level parallelism, and it is suitable for running on application specific hardware to achieve high performance with low cost. We have implemented and evaluated the FPGA-based stochastic simulator according to theoretical research of the algorithm. This paper introduces an improved architecture for accelerating a stochastic simulation algorithm called the Next Reaction Method. This new architecture has scalability to various size of FPGA. As the result with a middle-range FPGA, 5.38 times higher throughput was obtained compared to software running on a Core 2 Quad Q6600 2.40GHz.
Masato Yoshimi, Yuri Nishikawa, Yasunori Osana, Akira Funahashi, Yuichiro Shibata, Hideki Yamada, Noriko Hiroi, Hiroaki Kitano, Hideharu Amano
FPL5
2007 Implementation of a barotropic operator for ocean model simulation using a reconfigurable machine
abstract
This paper presents and discusses implementation of a barotropic operator used in ocean model simulation called Parallel Ocean Program (POP) using SRC-6 MAP. While a lot of high-end reconfigurable machines on which users can implement applications with a programming language are now available, enough implementation experience has not been accumulated for practical applications. In this paper, several implementation techniques accompanied by modification on original application source code are empirically evaluated and analyzed. The results show that appropriate use of internal memory and streaming DMA make 100 MHz FPGAs achieve comparative performance with GHz processors by using 100 MHz FPGAs.
Sayaka Shida, Yuichiro Shibata, Kiyoshi Oguri, Duncan A. Buell
FPL2
2007 A Combining technique of rate law functions for a cost-effective reconfigurable biological simulator
abstract
In order to simulate large scale biological models with a reconfigurable FPGA-based biochemical simulator system, reduction of required resources are essential. This paper proposes a method which combines common terms in rate law functions appeared in biochemical models and generates a shared hardware module used for numerical integration. In this approach, two functions are combined in a tree structure level, followed by pipeline scheduling and arithmetic module binding. The evaluation result reveals that this approach reduces hardware resources by 31.4% on average at the cost of 14.4% throughput degradation.
Hideki Yamada, Naoki Iwanaga, Yuichiro Shibata, Yasunori Osana, Masato Yoshimi, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Hideharu Amano, Akira Funahashi, Noriko Hiroi, Hiroaki Kitano, Kiyoshi Oguri
FPL3
2007 FPGA Implementation of a Data-Driven Stochastic Biochemical Simulator with the Next Reaction Method
abstract
This paper introduces a scalable FPGA implementation of a stochastic simulation algorithm (SSA) called the Next Reaction Method. There are some hardware approaches of SSAs that obtained high-throughput on reconfigurable devices such as FPGAs, but these works lacked in scalability. The design of this work can accommodate to the increasing size of target biochemical models, or to make use of increasing capacity of FPGAs. Interconnection network between arithmetic circuits and multiple simulation circuits aims to perform a data-driven multi-threading simulation. Approximately 8 times speedup was obtained compared to an excution on Xeon 2.80GHz.
Masato Yoshimi, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Yasunori Osana, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hideki Yamada, Hiroaki Kitano, Hideharu Amano
FPL8
2007 FPGA Implementation of a Statically Reconfigurable Java Environment for Embedded Systems
abstract
A demand for low power and high performance Java environments is now growing in the embedded systems field. One approach is dedicated Java processors which directly execute Java bytecode. We have proposed a novel reconfigurable Java environment which consists of a general purpose core processor with configurable bytecode processing units, a bytecode compiler, and a software Java Virtual Machine (JVM). This paper discusses design of a memory system for the reconfigurable Java architecture focusing on a hardware custom stack. Empirical evaluation using prototype systems reveals that adding a 2-word hardware stack shows the best results, achieving the performance enhancement of up to 15.9% compared to software execution.
Shinsuke Nino, Takayuki Mori, YoungHun Ko, Yuichiro Shibata, Kiyoshi Oguri
FPT4
2007 A Framework for Implementing a Network-Based Stochastic Biochemical Simulator on an FPGA
abstract
This paper studies several designs of network-based FPGA implementation of a stochastic simulation algorithm called the next reaction method, known for its large number of calculation involved. The procedure is divided into several subdivisions which will be implemented as independent modules, and they are connected with configurable interconnection networks so as to provide high throughput. By performing a multi-threading simulation, 3.6 times speedup was obtained compared with an execution on general purpose processors.
Masato Yoshimi, Yuri Nishikawa, Toshinori Kojima, Yasunori Osana, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Hideki Yamada, Hiroaki Kitano, Hideharu Amano
FPT7
2006 An Implementation Technique of Multi-Cycled Arithmetic Functions For a Dynamically Reconfigurable Processor
abstract
Dynamically reconfigurable processor (DRP) released by NEC Electronics is expected to have potential for high degree of parallel processing. Applications for DRP are described in C language, and parallelism in the source code is automatically extracted by a compiler. On the other hand, it is also important to optimize descriptions so that the potential performance of the device is effectively brought out. In this paper, arithmetic algorithms and an optimized coding technique to efficiently implement applications with multi-cycled arithmetic functions on DRP are discussed, focusing on the required number of the states. In this technique, the same kind of multi-cycled functions are aggregated into single functions, and arithmetic algorithms whose behavior is steady on operand values are utilized. The effects of the technique are evaluated with fixed-point arithmetic functions and polynomial arithmetic functions over a finite field, showing 2.68~3.09 times performance improvement without large increase in the number of states nor severe degradation of the frequency
Miwa Miyata, Hideyuki Tsuchiya, Yuichiro Shibata, Kiyoshi Oguri
FPL3
2006 Performance Evaluation of an Fpga-Based Biochemical Simulator ReCSip
abstract
ReCSiP is an FPGA-based biochemical simulator to accelerate kinetic simulations of biochemical pathways. Biochemical models are described as a set of ordinary differential equations (ODEs). Each equation in the model is called "rate law function", which represents the velocity of corresponding biochemical reaction mechanism. ReCSiP achieves high-throughput simulation with statically pipelined rate law function modules and numerical integration modules on an FPGA. This paper shows the basic structure of ReCSiP, and results of evaluation in 2 aspects: area and throughput. As the summary of evaluation, 1) about 64% of the total circuit area is occupied by floating-point arithmetic units, and 2) with an XC2VP70, ReCSiP at 90MHz can achieve 20times or more speedup compared to Intel's Pentium4 microprocessor at 3.2GHz
Yasunori Osana, Masato Yoshimi, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano
FPL5
2006 An FPGA Implementation of High Throughput Stochastic Simulator for Large-Scale Biochemical Systems
abstract
Stochastic simulation of biochemical systems has become one of major approaches to study life processes as system, yet is a computational challenge to run the simulation due to its vast calculation cost. This paper shows the implementation and evaluation of a stochastic simulation algorithm (SSA) called "first reaction method" on an FPGA-based biochemical simulator. It achieves high throughput by (1) consecutively throwing data into deeply-pipelined floating point arithmetic units, and (2) by distributing multiple simulators for parallel execution. As the result of evaluation on an FPGA-based simulation platform called ReC-SiP2, the simulator outperforms execution on Xeon 2.80 GHz by approximately 80 times, even with large-scale biochemical systems
Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano
FPL8
2005 Evaluation of Space Allocation Circuits
Shinya Kyusaka, Hayato Higuchi, Taichi Nagamoto, Yuichiro Shibata, Kiyoshi Oguri
EUC4
2005 New Area Management Method Based on "Pressure" for Plastic Cell Architecture
Taichi Nagamoto, Satoshi Yano, Mitsuru Uchida, Yuichiro Shibata, Kiyoshi Oguri
EUC4
2005 Efficient Scheduling of Rate Law Functions for ODE-Based Multimodel Biochemical Simulation on an FPGA
abstract
A reconfigurable biochemical simulator by solving ordinary differential equations has received attention as a personal high speed environment for biochemical researchers. For efficient use of the reconfigurable hardware, static scheduling of high-throughput arithmetic pipeline structures is essential. This paper shows and compares some scheduling alternatives, and analyzes the tradeoffs between performance and hardware amount. Through the evaluation, it is shown that the sharing first scheduling reduces the hardware cost by 33.8% in average, with the up to 11.5% throughput degradation. Effects of sharing of rate law functions are also analyzed.
Naoki Iwanaga, Yuichiro Shibata, Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Tomonori Fukushima, Hideharu Amano, Akira Funahashi, Noriko Hiroi, Hiroaki Kitano, Kiyoshi Oguri
FPL2
2005 A Framework for ODE-Based Multimodel Biochemical Simulations on an FPGA
abstract
Today, mathematical modeling and simulation of biochemical pathways take a major role in biological researches. However, modern microprocessors cannot provide enough throughputs to explore the large parameter space of target pathways. To address this problem, ReCSiP (a reconfigurable cell simulation platform), an FPGA-based biochemical simulator is proposed. It's an ODE-based simulator, which solves the rate-law functions. The framework proposed in this paper, enables to simulate pathways consisting many different types of chemical reactions by connecting the rate-law modules (solvers) on an FPGA. It provides the solver-to-solver communication mechanism on an FPGA and automatic configuration software to generate the circuit.
Yasunori Osana, Yow Iwaoka, Tomonori Fukushima, Masato Yoshimi, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano
FPL7
2005 The Design of Scalable Stochastic Biochemical Simulator on FPGA
Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano
FPT6
2004 Implementation of the Extended Euclidean Algorithm for the Tate Pairing on FPGA
Takehiro Ito, Yuichiro Shibata, Kiyoshi Oguri
FPL2
2001 A prototype chip of multicontext FPGA with DRAM for virtual hardware
abstract
DRAM-type multicontext FPGA is hopeful for Virtual Hardware. Since it is possible to implement a large number of contexts in a single chip. However, it has been reported only a few examples because of the difficulty of mixed process of DRAM and logic. Here we try to implement a prototype multi-context FPGA with DRAM for Virtual Hardware.
Daisuke Kawakami, Yuichiro Shibata, Hideharu Amano
ASP-DAC2
2000 A Virtual Hardware System on a Dynamically Reconfigurable Logic Device
abstract
WASMII is virtual hardware using a multi-context reconfigurable device with a data driven control. Since implementation of WASMII was infeasible due to the unavailability of such a device, the system has been only evaluated using an emulator so far. However, the first reconfigurable multi-context device called DRL has been developed by NEC. Making the use of its flexible reconfigurability, we have implemented a mechanism of WASMII on DRL.
Yuichiro Shibata, Masaki Uno, Hideharu Amano, Koichiro Furuta, Taro Fujii, Masato Motomura
FCCM1
2000 A Reconfigurable Stochastic Model Simulator for Analysis of Parallel Systems
abstract
Markov chain and queueing model are convenient tools with which to analyze parallel systems for architects. For a high speed simulation and easy modeling, a reconfigurable Markov chain/queueing model simulation system called RSMS (Reconfigurable Stochastic Model Simulator) is proposed. A user describes the target system in a dedicated description language called Taico. The description is automatically translated into the HDL description of the Markov chain/queueing model simulator. Then, the simulator is implemented on the FPGA devices of the reconfigurable system, and directly executed. Evaluation results with example parallel systems demonstrate that the performance of the proposed system is much superior to that of a common workstation.
Ou Yamamoto, Yuichiro Shibata, Hitoshi Kurosawa, Hideharu Amano
FCCM2
1998 Reconfigurable Systems: Activities in Asia and South Pacific (Embedded Tutorial)
abstract
Systems and researches on reconfigurable systems in Asia and South Pacific are picked up and introduced. Like Northern America and European countries, various platforms, application specific systems and education platforms have been proposed and developed.
Hideharu Amano, Yuichiro Shibata
ASP-DAC2