Yoshiki Yamaguchi

dblp:39/2795 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Sub-Microsecond HuffYUV-Based FPGA Accelerator for Real-Time Video Compression over High-Bandwidth Networks
abstract
This paper proposes a sub-microsecond FPGA-based video (de)compression system using the HuffYUV lossless format, designed to meet the ultra-low-latency demands of future high-speed networks such as IOWN and Beyond 5G. Unlike conventional codecs like H.264, HEVC, and JPEG XS, which introduce tens to hundreds of microseconds of latency, the proposed architecture achieves an end-to-end latency of just 1.1µs on Full HD 60fps video. By employing a bit-parallel Huffman decoding scheme, the system enables deterministic latency and efficient pipelining. Implemented on a Xilinx Versal FPGA (VPK120), the system consumes only a small portion of resources (e.g., 18.6% BRAMs) and achieves a throughput required to process 4K videos. Evaluation with a complete video pass-through system demonstrates its real-time performance and practical deployability. The architecture is highly suitable for edge applications requiring extreme responsiveness, including remote musical performance, telemedicine, and interactive XR.
Takeo Kurosawa, Keisuke Sugiura, Yoshiki Yamaguchi, Ryouhei Tsugami, Toshihito Fujiwara, Tatsuya Fukui, Satoshi Narikawa
CCNC3
2026 FusionSense: Tri-Stage Near-Sensor Learning for Runtime-Adaptive Multimodal Edge Intelligence
abstract
Autonomous systems and smart-industry deployments increasingly split computation across near-sensor, edge, and cloud resources, where tight energy, latency, and reliability budgets demand runtime adaptivity. In practice, deciding what to compute and transmit at each point is pivotal; yet as multimodal sensor suites (cameras, LiDAR/depth, etc.) proliferate at the edge, most prior approaches either (i) fuse modalities on powerful servers or (ii) apply uni-modal near-sensor filters that ignore cross-modal dependencies, leading to redundant transmissions or missed events. We present Fusion-Sense, a fusion-aware intelligent sensing framework for energy-constrained autonomous edge systems. Lightweight near-sensor classifiers are trained via a three-step procedure: (i) a server-side fusion model learns the downstream task, (ii) filter-out-safe (FoS) labels quantify each modality's necessity relative to the fused decision, and (iii) an edge-side fusion model is compacted by injecting near-sensor predictions as auxiliary signals. The result is a runtime decision layer that jointly reduces compute and communication while scaling linearly with sensor count. On a dual-modality (RGB+Depth/LiDAR) setup with SynDrone, FusionSense sustains task quality at substantially higher data-reduction rates than unimodal filters and delivers large end-to-end gains: up to 33× lower energy at 1% FoI prevalence, 11× at 10%, a 92.3% reduction in quality loss at a fixed 30% data reduction, and roughly 1.5× higher energy savings than the best prior filtering baseline.
Sanggeon Yun, Ryozo Masukawa, Minhyoung Na, Hyunwoo Oh, Yoshiki Yamaguchi, Wenjun Huang 0001, Sungheon Jeong 0001, Mohsen Imani
ISLPED5
2025 Low-Latency Immersive Display Systems with FPGA for Remote Applications
abstract
This paper presents an FPGA-based hardware system designed to enhance immersive display and reduce latency in video streaming for remote-controlled and autonomous driving applications. The system improves operator immersion and response accuracy by utilizing a spherical display with real-time distortion correction. The system corrects spherical display distortions in real time by addressing challenges such as video distortion and data buffering through dedicated hardware solutions. It achieves a buffer reduction of approximately 40 %, enabling seamless video streaming even over limited network connections. Experimental evaluations confirm the system's capability to minimize latency while delivering high-quality, immersive visuals, positively impacting remote vehicle control precision.
Den Tabata, Taiga Kobori, Yoshiki Yamaguchi, Ryouhei Tsugami, Toshihito Fujiwara, Tatsuya Fukui, Satoshi Narikawa
CCNC3
2025 Marker-Based Recognition for Autonomous Micro-Drone Flight: An FPGA-Optimized Feasibility Study
abstract
Drones are nowadays essential in construction and structural maintenance, providing high-resolution images for structural assessments and precise interventions. However, complex maneuvering in confined spaces and reduced communication with external positioning systems, such as GPS, pose challenges for conventional drones. Micro-drones, with their small size that minimizes failure impact and enhances adaptability in restricted or hazardous environments, have gained increasing research interest for expanding drone applications across various fields. This study presents an energy-efficient design that incorporates an optimized combination of basic color separation and a custom downsized version of the Sobel operator to detect markers placed along the trajectory of the drone in real time. The flight instructions stem from the marker identification to fine-tune altitude, horizontal position, and orientation of the micro-drone before moving forward to the next marker. The proposed strategy demonstrated accurate marker detection under favorable lighting conditions 97% of the time, while significantly reducing BlockRAM usage.
Diego Marcelo Ramirez Jove, Keisuke Sugiura, Yoshiki Yamaguchi
VLSI-SoC3
2025 PV-VTT: A Privacy-Centric Dataset for Mission-Specific Anomaly Detection and Natural Language Interpretation
abstract
Video crime detection is a significant application of computer vision and artificial intelligence. However, existing datasets primarily focus on detecting severe crimes by analyzing entire video clips, often neglecting the precursor ac-tivities (i.e., privacy violations) that could potentially pre-vent these crimes. To address this limitation, we present PV-VTT (f_rivacyY._iolation Y._ideo T_Iext), a unique mul-timodal dataset aimed at identifying privacy violations. PV-VTT provides detailed annotations for both video and text in scenarios. To ensure the privacy of individuals in the videos, we only provide video feature vectors, avoiding the release of any raw video data. This privacy-focused approach allows researchers to use the dataset while protecting participant confidentiality. Recognizing that privacy violations are often ambiguous and context-dependent, we propose a Graph Neural Network (GNN)-based video de-scription model. Our model generates a GNN-based prompt with image for Large Language Model (LLM), which deliver cost-effective and high-quality video descriptions. By leveraging a single video frame along with relevant text, our method reduces the number of input tokens required, maintaining descriptive quality while optimizing LLM API-usage. Extensive experiments validate the effectiveness and interpretability of our approach in video description tasks and flexibility of our PV-VTT dataset.11Dataset: https://ryozomasukawa.github.io/PV-VTT.github.io/
Ryozo Masukwa, Sanggeon Yun, Yoshiki Yamaguchi, Mohsen Imani
WACV3
2025 Qu-Trefoil: Large-Scale Quantum Circuit Simulator Working on FPGA With SATA Storages
abstract
Quantum circuits are fundamental components of quantum computing, and state-vector-based quantum circuit simulation is a widely used technique for tracking qubit behavior throughout circuit evolution. However, simulating a circuit with$n$qubits requires$2^{n+4}$bytes of memory, making simulations of more than 40 qubits feasible only on supercomputers. To address this limitation, we propose the Qu-Trefoil, a system designed for large-scale quantum circuit simulations on an FPGA-based platform called Trefoil. Trefoil is a multi-FPGA system connected to eight storage subsystems, each equipped with 32 SATA disks. Qu-Trefoil integrates a suite of HLS-based universal quantum gates, including Clifford gates (Hadamard (H), Pauli-Z (Z), Phase (S), Controlled-NOT (CNOT)), the T gate, and unitary matrix computation, along with HDL-designed modules for system-wide integration. Our extensive evaluation demonstrates the system's robustness and flexibility, covering quantum gate performance, chunk size, disk extensibility, and efficiency across different SATA generations. We successfully simulated quantum circuits with over 43 qubits, which required more than 128 TB of memory, in approximately 3.72 to 13.06 hours on a single storage subsystem equipped with one FPGA. This achievement represents a significant milestone in the advancement of quantum computing simulations. Furthermore, thanks to its unique architecture, Qu-Trefoil is more accessible, flexible, and cost-efficient than other existing simulators for large-scale quantum circuit simulations, making it a viable option for researchers with limited access to supercomputers.
Kaijie Wei, Hideharu Amano, Ryohei Niwase, Yoshiki Yamaguchi, Takefumi Miyoshi
IEEE Trans. Computers4
2023 GPU Acceleration of Multi-Object Tracking with Motion Vector Interpolation and Affine Transformation
abstract
In recent studies of object detection and tracking, neural networks have been widely used, and their accuracy has improved. However, its computational complexity is very high and requires the use of high-end GPUs. In order to achieve realtime inference on edge devices, it is necessary to reduce the computational complexity of the network by scaling it down, but this leads to a loss of accuracy. To avoid this loss of accuracy, a method has been proposed in which object detection is performed using a neural network at regular intervals, and in the frames in between, the detected object positions are interpolated using motion prediction. In this research, we propose a method to improve the accuracy of interpolation even when the camera is moving by using an affine transformation used for image stabilization. We also show its realtime computation method on Jetson TX2, one of the lowest power embedded GPUs. The proposed method enables realtime processing of object detection using Yolov5s and tracking of the detected objects at the edge.
Yoshiki Kunimoto, Qiong Chang, Yoshiki Yamaguchi, Tsutomu Maruyama
ASAP3
2023 GPU-FPGA-accelerated Radiative Transfer Simulation with Inter-FPGA Communication
abstract
The complementary use of graphics processing units (GPUs) and field programmable gate arrays (FPGAs) is a major topic of interest in the high-performance computing (HPC) field. GPU–FPGA-accelerated computing is an effective tool for multiphysics simulations, which encompass multiple physical models and simultaneous physical phenomena. Because the constituent operations in multiphysics simulations exhibit varying characteristics, accelerating these operations solely using GPUs is often challenging. Hence, FPGAs are frequently implemented for this purpose. The objective of the present study was to further improve application performance by employing both GPUs and FPGAs in a complementary manner. Recently, this approach has been applied to the radiative transfer simulation code for astrophysics known as ARGOT, with evaluation results quantitatively demonstrating the resulting improvement in performance. However, the evaluation results in question came from the use of a single node equipped with both a GPU and FPGA. In this study, we extended the GPU–FPGA-accelerated ARGOT code to operate on multiple nodes using the message passing interface (MPI) and an FPGA-to-FPGA communication technology scheme called Communication Integrated Reconfigurable CompUting System (CIRCUS). We evaluated the performance of the ARGOT code with multiple GPUs and FPGAs under weak scaling conditions, and found it to achieve up to 12.8x speedup compared to the GPU-only execution.
Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura
HPC Asia3
2023 A Scalable Many-core Overlay Architecture on an HBM2-enabled Multi-Die FPGA
abstract
The overlay architecture enables to raise the abstraction level of hardware design and enhances hardware-accelerated applications’ portability. In FPGAs, there is a growing awareness of the overlay structure as typified by many-core architecture. It works in theory; however, it is difficult in practice, because it is beset with serious design issues. For example, the size of FPGAs is bigger than before. It is exacerbating the issue of the place-and-route. Besides, a single FPGA is actually the sum of small-to-middle FPGAs by advancing packaging technology like silicon interposers. Thus, the tightly coupled many-core designs will face this covert issue that the wires among the regions are extremely restricted. This article proposes efficient essential processing elements, micro-architecture design, and the interconnect architecture toward a scalable many-core overlay design. In particular, our work proposes a novel compact buffering technique to reduce memory resource utilization in tightly connected overlays while preserving computational efficiency. This technique reduces the utilization of BlockRAM to nearly 50% while achieving a best-case computational efficiency of 91.93% in a three-dimensional Jacobi benchmark. Besides, the proposed enhancements led to around 2× and 3× improvement in performance and power efficiency, respectively. Moreover, the improved scalability allowed increasing compute resources and delivering around 4× better performance and power efficiency, as compared to the baseline Dynamically Re-programmable Architecture of Gather-scatter Overlay Nodes overlay.
Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku
ACM Trans. Reconfigurable Technol. Syst.2
2022 Accelerating Radiative Transfer Simulation on NVIDIA GPUs with OpenACC
Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura
PDCAT3
2021 HBM2 Memory System for HPC Applications on an FPGA
abstract
Field Programmable Gate Arrays (FPGAs) have been targeted as a new accelerator of the HPC field. This is because the barrier to using FPGAs has been gradually lowered due to the widespread use of high-level synthesis (HLS) technology. In addition, the bandwidth of external memory in FPGAs is much lower than that of other accelerators widely used in HPC, such as NVIDIA V100 GPUs. However, the latest FPGAs can use High Bandwidth Memory 2 (HBM2), which has a memory bandwidth of up to 512GB/s. Therefore, we believe FPGAs will be a viable option for speeding up applications. However, unlike CPUs and GPUs, FPGAs do not have caches and memory networks to exploit the full potential of HBM2, which may limit the efficiency of the application. In this paper, we propose a memory system for HBM2 and HPC applications. We show the prototype implementation of the system and evaluate its performance. We also demonstrate the use of the proposed system from an application developed in High-Level Synthesis (HLS) written in C++.
Norihisa Fujita, Ryohei Kobayashi 0001, Yoshiki Yamaguchi, Taisuke Boku
CLUSTER3
2021 An FPGA-based storage control with load balancing
abstract
In the last decade, the number of cloud computing companies adopting FPGAs for performance gain has increased considerably. Indeed, the adoption of FPGA has contributed to the improvement of network performance in HPC (High-Performance Computing) systems. Nevertheless, data storage performance has not followed the trend and remained one hard-to-solve bottleneck lowering the processing capacity of a whole system, in particular, systems with high demands for real-time computations. This paper proposes a high-speed, large-capacity, and low-latency storage system that makes a breakthrough in storage performance by using inherent FPGA parallelism to control multiple SATA devices. It also has a load balancing function that can inhibit the slowest among multiple connected storage devices from hindering the throughput of the entire storage system. To verify the proposed approach, an FPGA board was custom manufactured, which embeds a Xilinx Kintex Ultrascale FPGA. It can accept connections from up to 16 SATA devices simultaneously. The SATA controller on an FPGA was almost developed from scratch and written by Verilog HDL. It enables the system to achieve low-latency processing. In our experimental results, it shows 16 SATA devices (SAMSUNG EVO 860 SSDs) work simultaneously and adequately. Besides, the load-balancing function was evaluated on the board.
Naoya Umezu, Yoshiki Yamaguchi, Taisuke Boku
CLUSTER2
2021 An efficient RTL buffering scheme for an FPGA-accelerated simulation of diffuse radiative transfer
abstract
This paper proposes the efficient buffering approach for implementing radiative transfer equations to bridge the performance gap between processing elements and HBM memory bandwidth. The radiation transfer equation originally focuses on the fundamental physics process in astrophysics. Besides, it has become the focus of a lot of attention in recent years because of the wealth of applications such as medical bioimaging. However, the acceleration requires a complicated memory access pattern with low latency, and the earlier studies unveil conventional memory access based on software control has no aptitude for this computation. Thus, this article introduced an HBM FPGA and proposed an application-specific buffering mechanism called PRISM (PRefetchable and Instantly accessible Scratchpad Memory) to efficiently bridge the computational unit and the HBM. The proposed approach was evaluated on a XILINX Alveo U280 FPGA, and the experimental results are also discussed.
Kazuki Furukawa, Ryohei Kobayashi 0001, Tomoya Yokono, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Masayuki Umemura
FPT5
2020 Condensing an overload of parallel computing ingredients into a single architecture recipe
abstract
General-purpose processors offer the best programming flexibility to address a wide range of problems. Nonetheless, they still lack behind special-purpose processors when it comes to sustained computational performance. Here, we leverage the best from both worlds and we propose a flexible, highly scalable, high-performance computing architecture with versatility in mind. The proposed architecture code-named DRAGON, benefits from several forms of parallelism such as SIMD, VLIW, Memory Broadcasting and even vector processing.
Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku
ASAP2
2020 Accelerating Radiative Transfer Simulation with GPU-FPGA Cooperative Computation
abstract
Field-programmable gate arrays (FPGAs) have garnered significant interest in research on high-performance computing. This is ascribed to the drastic improvement in their computational and communication capabilities in recent years owing to advances in semiconductor integration technologies that rely on Moore’s Law. In addition to these performance improvements, toolchains for the development of FPGAs in OpenCL have been offered by FPGA vendors to reduce the programming effort required. These improvements suggest the possibility of implementing the concept of enabling on-the-fly offloading computation at which CPUs/GPUs perform poorly relative to FPGAs while performing low-latency data transfers. We consider this concept to be of key importance to improve the performance of heterogeneous supercomputers that employ accelerators such as a GPU. In this study, we propose GPU–FPGA-accelerated simulation based on this concept and demonstrate the implementation of the proposed method with CUDA and OpenCL mixed programming. The experimental results showed that our proposed method can increase the performance by up to $17.4 \times$ compared with GPU-based implementation. This performance is still $1.32 \times$ higher even when solving problems with the largest size, which is the fastest problem size for GPU-based implementation. We consider the realization of GPU–FPGA-accelerated simulation to be the most significant difference between our work and previous studies.
Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura
ASAP3
2020 Toward OpenACC-enabled GPU-FPGA Accelerated Computing
abstract
Field-programmable gate arrays (FPGAs) have garnered significant interest in research on high-performance computing because their computation and communication capabilities have drastically improved in recent years due to advances in semiconductor integration technologies that rely on Moore's Law. These improvements reveal the possibility of implementing a concept to enable on-the-fly offloading computation at which CPUs/GPUs perform poorly to FPGAs while performing low-latency data movement. We think that this concept is key to improving the performance of heterogeneous supercomputers using accelerators such as the GPU. In this paper, we propose a GPU-FPGA-accelerated simulation based on the concept and show preliminary results of the proposed concept.
Norihisa Fujita, Ryohei Kobayashi 0001, Yoshiki Yamaguchi, Kohji Yoshikawa, Makito Abe, Masayuki Umemura
CLUSTER3
2019 MITRACA: Manycore Interlinked Torus Reconfigurable Accelerator Architecture
abstract
Big data, Artificial Intelligence, and cloud services are emerging technologies whose power consumption due to the tremendous amount of computing resources became a significant issue in data centers. FPGA (Field Programmable Gate Array) based accelerators may offer a convenient solution for high-performance and energy-efficient computing. However, designing these accelerators using hardware description language is a burdensome task and requires specialized skill sets. To help not only FPGA engineers but also software programmers to implement their applications quickly, an overlay architecture on an FPGA will be a good candidate. Thus, this paper proposes a coarse-grained overlay architecture with SIMD (Single Instruction Multiple Data) instructions.
Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku
ASAP2
2018 OpenCL-ready High Speed FPGA Network for Reconfigurable High Performance Computing
abstract
Field programmable gate arrays (FPGAs) have gained attention in high-performance computing (HPC) research because their computation and communication capabilities have dramatically improved in recent years as a result of improvements to semiconductor integration technologies that depend on Moore's Law. In addition to FPGA performance improvements, OpenCL-based FPGA development toolchains have been developed and offered by FPGA vendors, which reduces the programming effort required as compared to the past. These improvements reveal the possibilities of realizing a concept to enable on-the-fly offloading computation at which CPUs/GPUs perform poorly to FPGAs while performing low-latency data movement. We think that this concept is one of the keys to more improve the performance of modern heterogeneous supercomputers using accelerators like GPUs. In this paper, we propose high-performance inter-FPGA Ethernet communication using OpenCL and Verilog HDL mixed programming in order to demonstrate the feasibility of realizing this concept. OpenCL is used to program application algorithms and data movement control when Verilog HDL is used to implement low-level components for Ethernet communication. Experimental results using ping-pong programs showed that our proposed approach achieves a latency of 0.99 μs and as much as 4.97 GB/s between FPGAs over different nodes, thus confirming that the proposed method is effective at realizing this concept.
Ryohei Kobayashi 0001, Yuma Oobata, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku
HPC Asia4
2013 The study of three-dimensional multiphase-flow simulator
abstract
This paper presents an FPGA-based system that aims to perform three-dimensional multiphase-flow simulations. In this implementation, the immiscible lattice-gas automata (LGA) were selected as the target simulation model. The immiscible LGA are classified as the cellular automata (CA), which are a discrete dynamic model. The simulation box consists of an array of cells. The structure of the box should be chosen carefully since it decides the limitation of simulation behaviors. On the other hand, all the lattice structures should be allocated in order so that they enable us to describe the LGA as stencil computation. Here, we expect that the FPGAs have a great possibility in achieving dramatic speedup. Experimental result shows speedups that achieves two orders of magnitude in the immiscible LGA with the face-centred hyper cubic lattice structure.
Kenta Fujinami, Yoshiki Yamaguchi, Akira Sugiura, Yuetsu Kodama
FPL2
2011 A comparison of FPGAs, GPUS and CPUS for Smith-Waterman algorithm (abstract only)
abstract
The Smith-Waterman algorithm is a key technique for comparing genetic sequences. This paper presents a comprehensive study of a systolic design for Smith-Waterman algorithm. It is parameterized in terms of the sequence length, the amount of parallelism, and the number of FPGAs. Two methods of organizing the parallelism, the line-based and the lattice-based methods, are introduced. Our analytical treatment reveals how these two methods perform relative to peak performance when the level of parallelism varies. A novel systolic design is then described, showing how the parametric description can be effectively implemented, with specific focus on enhancing parallelism and on optimizing the total size of memory and circuits; in particular, we develop efficient realizations for compressing score matrices and for reducing affine gap cost functions. Promising results have been achieved showing, for example, a single XC5VLX330 FPGA at 131MHz can be three times faster than a platform with two NVIDIA GTX295 at 1242MHz.
Yoshiki Yamaguchi, Kuen Hung Tsoi, Wayne Luk
FPGA1
2009 Performance comparison of FPGA, GPU and CPU in image processing
abstract
Many applications in image processing have high inherent parallelism. FPGAs have shown very high performance in spite of their low operational frequency by fully extracting the parallelism. In recent micro processors, it also becomes possible to utilize the parallelism using multi-cores which support improved SIMD instructions, though programmers have to use them explicitly to achieve high performance. Recent GPUs support a large number of cores, and have a potential for high performance in many applications. However, the cores are grouped, and data transfer between the groups is very limited. Programming tools for FPGA, SIMD instructions on CPU and a large number of cores on GPU have been developed, but it is still difficult to achieve high performance on these platforms. In this paper, we compare the performance of FPGA, GPU and CPU using three applications in image processing; two-dimensional filters, stereo-vision and k-means clustering, and make it clear which platform is faster under which conditions.
Shuichi Asano, Tsutomu Maruyama, Yoshiki Yamaguchi
FPL3
2009 Dynamic reconfiguration system for real-time video processing
abstract
Dynamic reconfiguration is a powerful approach for realizing a high-speed and low-power-consumption system. It enables a system to gather as much specific circuits as the system requires, and then to comprise general-purpose computation. This paper introduces a real-time video-streaming processing system with dynamic reconfigurations. In our case studies, our system achieved high performance (4times VGA, 30 fps) and low power consumption (< 1 watt).
Saya Hinaga, Yoshiki Yamaguchi, Tetsuhiko Yao, Tohru Kawabe
FPL2
2008 How fast is an FPGA in image processing?
abstract
In image processing, FPGAs have shown very high performance in spite of their low operational frequency. This high performance comes from (1) high parallelism in applications in image processing, (2) high ratio of 8 bit operations, and (3) a large number of internal memory banks on FPGAs which can be accessed in parallel. In the recent micro processors, it becomes possible to execute SIMD instructions on 128 bit data in one clock cycle. Furthermore, these processors support multi-cores and large cache memory which can hold all image data for each core. In this paper, we compare the performance of FPGAs with those processors using three applications in image processing; two-dimensional filters, stereo-vision and k-means clustering, and make it clear how fast is an FPGA in image processing, and how many hardware resources are required to achieve the performance.
Takashi Saegusa, Tsutomu Maruyama, Yoshiki Yamaguchi
FPL3
2008 An adaptive pattern recognition hardware with on-chip shift register-based partial reconfiguration
abstract
A pattern recognition system that can process a large amount of image data at high speed is required in many fields. In this paper, we propose an on-chip pattern recognition system that utilizes the reconfigurability of the FPGA. The features of the system are not only very high recognition speed but also an adaptive function. For example, when objects to be detected change appearance, recognition parameters must be changed to retain the recognition accuracy. The system can automatically adjust by executing on-chip partial reconfiguration. The system runs at 25MHz and can return a recognition result in one clock cycle, 40ns. To update the system, all processes needed for searching for the best recognition parameters, generating configuration data and reconfiguring the system are carried out within 30s.
Hiroyuki Kawai, Yoshiki Yamaguchi, Moritoshi Yasunaga, Kyrre Glette, Jim Tørresen
FPT2
2007 An FPGA Implementation of Multiple Sequence Alignment Based on Carrillo-Lipman Method
abstract
Multiple sequence alignment problems in computational biology have been focused recently because of the rapid growth of sequence databases. By computing alignment, we can understand similarity among the sequences. In this paper, we describe a compact system with an FPGA board and a host computer for multiple sequence alignment based on Carrillo-Lipman method. In our system, two dimensional dynamic programming is repeatedly applied along other dimensions to realize multidimensional search with a simple and common architecture, and unnecessary parts of the search space for finding the optimal alignment are skipped using Carrillo-Lipman method to reduce the computation time.
Shingo Masuno, Tsutomu Maruyama, Yoshiki Yamaguchi, Akihiko Konagaya
FPL3
2007 High speed tablation system using an FPGA designed for distribution tables of frequent DNA subsequences
abstract
A method is described for enumerating the frequencies of DNA subsequences on a system comprising a host computer and a field programmable gate array (FPGA) board with one FPGA. Frequencies of subsequences with lengths of up to K0+ K1+ K2(24 in the current implementation) are enumerated in three phases. In these three phases, subsequences with lengths of up to K0, K0+ K1, and K0+ K1+ K2, respectively, are enumerated; these three phases are executed simultaneously on a pipelined circuit, resulting in high performance. The enumeration of frequent subsequences in databases, which are becoming larger and larger, will enable subsequences that are unique and/or repeatedly used in many parts of the sequences to be found.
Yoshiki Yamaguchi, Tsutomu Maruyama, Fumikazu Konishi, Akihiko Konagaya
FPL1
2007 Bio-Inspired Functional Asymmetry Camera System
Yoshiki Yamaguchi, Noriyuki Aibe, Moritoshi Yasunaga, Yorihisa Yamamoto, Takaaki Awano, Ikuo Yoshihara
ICONIP (2)1
2006 Multidimensional Dynamic Programming for Homology Search on Distributed Systems
Shingo Masuno, Tsutomu Maruyama, Yoshiki Yamaguchi, Akihiko Konagaya
Euro-Par3
2005 Multidimensional Dynamic Programming for Homology Search
abstract
Alignment problems in computational biology have been focused recently because of the rapid growth of sequence databases. By computing alignment, we can understand similarity among the sequences. Many systems for alignment have been proposed to date, but most of them are designed for two-dimensional alignment (alignment between two sequences). In this paper, we describe a compact system with an off-the-shelf FPGA board and a host computer for more than three-dimensional alignment based on dynamic programming. In our approach, high performance is achieved (1) by configuring optimal circuit for each dimensional alignment, and (2) by two phase search in each dimension by reconfiguration. In order to realize multidimensional search with a common architecture, two-dimensional dynamic programming is repeated along other dimensions. With this approach, we can minimize the size of units for alignment and achieve high parallelism. Our system with one XC2V6000 enables about 300-fold speedup as compared with single Intel Pentium 4 2GHz processor for four-dimensional alignment, and 100-fold speedup for five-dimensional alignment.
Shingo Masuno, Tsutomu Maruyama, Yoshiki Yamaguchi, Akihiko Konagaya
FPL3
2005 Spatiotemporal Simulation of a Single Living Cell
Yoshiki Yamaguchi, Tsutomu Maruyama, Ryuzo Azuma, Akihiko Konagaya
FPT1
2004 Three-Dimensional Dynamic Programming for Homology Search
Yoshiki Yamaguchi, Tsutomu Maruyama, Akihiko Konagaya
FPL1
2002 High Speed Homology Search Using Run-Time Reconfiguration
Yoshiki Yamaguchi, Yosuke Miyajima, Tsutomu Maruyama, Akihiko Konagaya
FPL1
2001 An Approach to Real-Time Visualization of PIV Method with FPGA
Tsutomu Maruyama, Yoshiki Yamaguchi, Atsushi Kawase
FPL2