Masahiro Iida

dblp:85/2316 · DBLP profile ↗
← Back
38ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 53% Electronic design automation · 47%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › reconfigurable architecture
embedded FPGA
1.012026
Improving Area Efficiency in Synthesizable eFPGA with Multi-output Logic Cell and Domain-Specific Routing Architecture · FPGA 2026
Electronic design automation
logic synthesis
1.012026
Improving Area Efficiency in Synthesizable eFPGA with Multi-output Logic Cell and Domain-Specific Routing Architecture · FPGA 2026
Reconfigurable computing and FPGAs
FPGA routing architecture
0.312026
Improving Area Efficiency in Synthesizable eFPGA with Multi-output Logic Cell and Domain-Specific Routing Architecture · FPGA 2026
Electronic design automation › physical design › placement and routing
FPGA placement and routing
0.212013
A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013
Electronic design automation
physical design
0.212013
A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013
Reconfigurable computing and FPGAs › FPGA physical design
FPGA clustering
0.112006
Effective clustering technique to optimize routability of outer cluster nets · FPGA 2006
Electronic design automation › physical design › routing › routability
routability optimization
0.112006
Effective clustering technique to optimize routability of outer cluster nets · FPGA 2006

Methods — techniques the papers use, named apart from their topics

multi-output logic cell · 1.0AIG-based logic cell · 1.0object-oriented programming · 0.2HDL template generation · 0.2evaluation functions · 0.1
YearPublicationVenuePosition
2026 Improving Area Efficiency in Synthesizable eFPGA with Multi-output Logic Cell and Domain-Specific Routing Architecture
abstract
We address the area overhead of synthesizable eFPGAs by proposing the ''PAE Cell,'' an AIG-based multi-output logic cell, and a domain-specific routing architecture for control logic. Evaluations demonstrate that the PAE Cell achieves 4-LUT equivalence at approximately 0.4× the area, while our routing architecture reduces total eFPGA area by up to 80% compared to conventional islandstyle designs.
Ryo Iwasaki, Yumi Iseki, Sota Kohata, Miyu Yoshida, Kenshu Seto, Masahiro Iida
FPGA7
2022 FPL Demo: An FPGA-IP Prototype Chip for MEC devices
abstract
This demonstration shows a prototype chip of SLM (Scalable Logic Module) for a novel FPGA-IP embedded in various chips for edge computing. In this paper, the authors briefly describe the architecture of our FPGA-IP and the evaluation environment for the prototype chip.
Morihiro Kuga, Masahiro Iida, Hideharu Amano
FPL2
2020 Image Search System Based on Feature Vectors of Convolutional Neural Network
abstract
Edge computing offers real-time applications because the edge device closes with the data source such as the end device. This condition gives the challenge to implement deep learning in the edge device. Unfortunately, deep learning requires high computing resources, but often edge-side devices have limitations. In this study, we built an image search system based on CNN (Convolutional Neural Network)'s feature vectors to address the challenges by enlarging the implementation of CNN in the edge device such as Raspberry Pi 3. The image search system applied these informative features vector to get similar images in the image searching task by using cosine similarity. We used a 102-flower categories dataset and we prepared a light database to run the system as an off-line system in the edge device. The MobileNetV2 as CNN's model reached 70.02% of the top 1 accuracy and 92.84 % for the top 5 accuracy. As a result, the image search system showed five images result with the most similar image from the same image category. Image resolution, model complexity, and hardware capability give the significant time in this image search system. The framework of this system can be simply used for other deep learning models and applications by updating the model, dataset, database, and hardware.
Mery Diana, Motoki Amagasaki, Masahiro Iida
TENCON3
2017 hCODE 2.0: An open-source toolkit for building efficient FPGA-enabled clouds
abstract
Major cloud service providers have started employing field-programmable gate arrays (FPGAs) to implement high-performance and low-power-consumption cloud capability. However, building or utilizing an FPGA-enabled cloud is still challenging due to the lack of fundamental tools. In our previous work, we proposed an hCODE base system for managing portable accelerator IPs on different hardware. In this paper, we extend the previous work and introduce the hCODE 2.0, which is an open-source toolkit for building efficient FPGA-enabled clouds. First, we provide a fundamental toolkit to simplify HW project management and FPGA management at a cluster scale. Second, we implement on-chip resource virtualization and accelerator scheduling capabilities to show possibilities of improving FPGA utilization efficiency with our tools.
Qian Zhao 0001, Hendarmawan, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT4
2016 hCODE: An open-source platform for FPGA accelerators
abstract
Field-programmable gate arrays (FPGAs) have demonstrated great speed performance and power efficiency advantages over conventional computers in various domains. However, it is still difficult for general software engineers to employ FPGA-based hardware accelerators because of the gap between hardware and software development methods. In this paper, we propose a heterogeneous computing oriented development environment (hCODE) to simplify the creation, sharing, and project integration of hardware accelerators. The hCODE defines hardware specifications and interface design rules. Hardware developers can provide designs that follow these rules, allowing software engineers to easily search, download, and integrate accelerators in their applications without caring about the details of the hardware.
Qian Zhao 0001, Takuya Nakamichi, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT4
2016 A heuristic method of generating diameter 3 graphs for order/degree problem (invited paper)
abstract
We propose a heuristic method that generates a graph for order/degree problem. Target graphs of our heuristics have large order (> 4000) and diameter 3. We describe the observation of smaller graphs and basic structure of our heuristics. We also explain an evaluation function of each edge for efficient 2-opt local search. Using them, we found the best solutions for several graphs.
Teruaki Kitasuka, Masahiro Iida
NOCS2
2016 A novel soft error tolerant FPGA architecture
abstract
Due to reaching the nanoscale transistor size, effect of single event upset (SEU) to the memory has become conspicuous. In small device geometries, a single particle strike might affect multiple adjacent cells in a memory array resulting in a multiple bit upset (MBU). Traditional fault tolerance technologies such as triple modular redundancy (TMR) and error correcting code (ECC) occupy the large area and have vulnerability to MBU. In this research, we propose DMR based error correct circuit and employ a combination of proposed circuit and the interleaving technique to mitigate MBU. In addition, we explain soft error simulator developed to calculate bit interleaving distance. The results show that the area of proposed circuit is the smallest when we compare the proposed circuit, ECC based error correct circuit and TMR. Simulation results show that the interleaving distance which can conceal all MBU patterns is 4.
Motoki Amagasaki, Yuji Nakamura, Takuya Teraoka, Masahiro Iida, Toshinori Sueyoshi
VLSI-SoC4
2015 Architecture exploration of 3D FPGA to minimize internal layer connection
abstract
A three-dimensional (3D) integration based on wafer-to-wafer bonding using through-silicon vias (TSVs) has been developed for the fabrication of new 3D large-scale integrated chips. To balance between cost and performance, and to explore 3D field-programmable gate array (FPGA) with realistic 3D integration processes, we propose spatially distributed and functionally distributed types of 3D FPGA architectures. The functionally distributed architecture consists of two wafers, a logic layer and a routing layer, and is stacked by a face-down process technology. Since vertical wires pass through microbumps, no TSVs are needed. In contrast, the spatially distributed architecture is divided into multiple layers with the same structure, unlike in the functionally distributed type. This architecture can be expanded to more than two layers by stacking multiples of the same die. The goal of this paper is to elucidate the advantages and disadvantages of these two types of 3D FPGAs. According to our evaluation, when only two layers are used, the functionally distributed architecture is more effective. When higher performance is achieved by using more than two layers, the spatially distributed architecture achieves better performance.
Motoki Amagasaki, Yuto Takeuchi, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC4
2014 A logic cell architecture exploiting the shannon expansion for the reduction of configuration memory
abstract
Most modern field-programmable gate arrays (FPGAs) employ a look-up table (LUT) as their basic logic cell. Although a k-input LUT can implement any k-input logic, its functionality relies on a large amount of configuration memory. As FPGA scales improve, the increased quantity of configuration memory cells required for FPGAs will require a larger area and consume more power. Moreover, the soft-error rate per device will also increase as more configuration memory cells are embedded. We propose scalable logic modules (SLMs), logic cells requiring less configuration memory, reducing configuration memory by making use of partial functions of Shannon expansion for frequently appearing logics. Experimental results show that SLM-based FPGAs use much less configuration memory and have smaller area than conventional LUT-based FPGAs.
Qian Zhao 0001, Kyosei Yanagida, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL4
2014 A novel three-dimensional FPGA architecture with high-speed serial communication links
abstract
Three-dimensional (3D) integrated circuit technology is expected to offer continual improvement to very-large-scale integration performance as the process of miniaturization approaches physical limits. However, because the through-silicon vias (TSVs) that are used to create interlayer vertical connections are much larger area than transistors, there is an inherent tradeoff between connectivity and small size. Field-programmable gate arrays (FPGAs) are particularly noted for requiring a high level of routing resources, which means that it is unrealistic to make the same number of connections vertically as horizontally. In previous research, we proposed a method for creating a two-layer compact 3D FPGA with face-down integration (the base FPGA). In this paper, we discuss stacking multiple base FPGAs by the face-up method and propose a method for achieving highspeed interlayer communications with TSV serial connections. The proposed architecture improves FPGA performance by using smaller TSVs. The evaluation results show that the proposed 3D FPGA can achieve a total area that is as low as 67% the equivalent two-dimensional FPGA.
Takuya Kajiwara, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morituro Kuga, Toshinori Sueyoshi
FPT4
2014 Zyndroid: An Android platform for software/hardware coprocessing
abstract
High performance is required of many Android systems because embedded systems written for this operating system are used in several fields and rely on increasingly complicated processing. To accommodate this, we present a software/hardware (SW/HW) coprocessing platform implemented on a programmable system-on-a-chip (Xilinx Inc.: Zynq). This platform provides a unified architecture, extended OS kernel, application framework, and application distribution model to simplify the development and use of Android SW/HW coprocessing applications.
Susumu Mashimo, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT3
2014 Blokus Duo engine on a Zynq
abstract
In this article, we present a design of a Blokus Duo engine for the ICFPT 2014 Design Competition. Our design is implemented on a Xilinx Zynq-7000 SoC ZC706 Evaluation Kit and we employ the minimax algorithm with alpha-beta pruning. The ARM processor runs the search algorithm, and the handwritten hardware accelerator calculate within 1 second under the competition constraint. One of the keys to a stronger Blokus Duo player is to evaluate more states of a game; our Blokus Duo engine evaluates 12.3 times as many nodes of a game search tree as the Intel Core i7-3770T.
Susumu Mashimo, Kansuke Fukuda, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPT4
2013 A novel FPGA design framework with VLSI post-routing performance analysis (abstract only)
abstract
The most widely used open-source field-programmable gate array (FPGA) placement and routing tool is VPR, which can define the target FPGA, perform placement and routing, and report area and timing information. However, it cannot be used in FPGA IP design efficiently for two reasons. First, for most newly developed FPGA architectures, VPR cannot support them directly. Modifying the C-coded VPR for using it to evaluate a number of new architectures requires a long time. Second, the accuracy of the VPR performance results is not enough for the evaluation of a complete synthesizable FPGA IP in the design that targets the productions of LSI. We propose a FPGA design framework that in particular improves FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely EasyRouter. EasyRouter is developed using the C# language. When an object-oriented programming method is used, the source codes are fewer and easier manage compared to VPR, which shortens the development time. By using simple HDL templates, EasyRouter can automatically generate entire chip HDL codes and the configuration bitstream. With these files, the FPGA IP can be evaluated with commercial VLSI CADs with high accuracy and reliability.
Qian Zhao 0001, Kazuki Inoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPGA4
2013 Defect-robust FPGA architectures for intellectual property cores in system LSI
abstract
In this paper, we propose fault-tolerant field-programmable gate array (FPGA) architectures and their computer-aid design (CAD) for intellectual property (IP) cores in system large-scale integration (LSI). Unlike discrete FPGAs, in which the integration scale can be made relatively large, programmable IP cores must correspond to arrays of various sizes. The key features of our architectures are regular tile structure, spare modules and bypass wires for fault avoidance, and configuration mechanism for single-cycle reconfiguration. In addition, we develop routing tools, namely EasyRouter for proposed architecture. This tool can handle various array sizes corresponding to developed programmable IP cores. In this evaluation, we compared the performances of conventional FPGA and the proposed fault-tolerant FPGA architectures. On average, our architectures have less than 2.2 times the area and 1.3 times the delay compared with conventional FPGA architectures. At the same time, conventional FP-GAs cannot tolerate faults, whereas our architectures perform with a 90% success rate in fault avoidance for a ratio of faulty tiles of 1% or less.
Motoki Amagasaki, Kazuki Inoue, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL4
2013 An automatic FPGA design and implementation framework
abstract
Conventional FPGA design and implementation processes involve two separate flows. The FPGA architecture is determined by academic FPGA design flow. However, in the implementation phase, commercial VLSI design flow are used. In this research, we propose an FPGA design framework in order to improve synthesizable FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely the EasyRouter, which can bridge the two flows efficiently. With this design flow, accurate physical information can be reported when a new FPGA IP architecture is evaluated with reliable commercial VLSI CADs.
Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL3
2013 Three-dimensional stacking FPGA architecture using face-to-face integration
abstract
In recent years, as VLSI process scales have developed into deep sub-micrometer dimensions, routing delay problems have become critical. For reconfigurable logic devices (RLDs) like field-programmable gate arrays (FPGAs) in particular, routing resources occupy major parts of the available area and hinder performance. In order to balance cost and performance, and to explore 3D FPGA architectures with realistic 3D LSI processes, we proposed a novel two-layers 3D FPGA architecture based on 3D connections on logic block input and output pins. Evaluation shows that this novel RLD with two layers of 3D routing architecture uses 48.75% less on-board area and 30.54% less critical path delay than does a conventional 2D 4-lookup table island-style FPGA on average.
Tetsuro Hamada, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC4
2012 Designing Flexible Reconfigurable Regions to Relocate Partial Bitstreams
abstract
Current commercial SRAM-based FPGAs, such as Virtex-6 and Stratix-V, can perform dynamic partial reconfiguration (DPR). Partial reconfiguration (PR) can change a part of the device without reconfiguring the whole chip. Thus, we can switch the part of system with continuing the operation. However, the authorized design flow by Xilinx creates different PR bit stream (PRB) for each partially reconfigurable region (PRR) even if it is the same circuit. This indicates that N × M PRBs must be prepared to implement M types modules on N PRRs. This increases design time and memory usage to store PRBs. This paper presents a uniforming design technique for PRRs to relocate a PRB among them. In addition, uniformed PRRs can be used to implement large module by combining adjacent PRRs. In this work, we use Xilinx Virtex-6 XC6VLX240T and Integrated Software Environment 13.3 (ISE) to verify the proposed technique.
Yoshihiro Ichinomiya, Sadaki Usagawa, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FCCM4
2012 Fault detection and avoidance of FPGA in various granularities
abstract
Although redundancy techniques are generally used to provide fault tolerance for the system large scale integrations (LSIs), these techniques have significantly costs. However, field programmable gate arrays (FPGAs) can easily provide high reliability due to their reconfiguration ability. The present paper proposes an effective fault detection method for the global interconnects of the FPGA. In the proposed method, by combining different kinds of test patterns, the fault resources can be identified. We also developed placement and routing tools to avoid fault resources in two cases, namely, tile-level avoidance and multiplexer-level avoidance. In the evaluation of the proposed technique, the proposed detection method diagnosed a defect multiplexer with six test configurations. We found that the fault FPGA can achieve the same performance as normal FPGA in multiplexer-level avoidance.
Kazuki Inoue, Yuki Nishitani, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FPL4
2012 Accelerated evaluation of SEU failure-in-time using frame-based partial reconfiguration
abstract
SRAM-based field programmable gate arrays (FPGAs) are vulnerable to soft-error. To improve circuit dependability, various dependable design techniques have been studied. By the same token, evaluation techniques are required to ensure dependability. The most popular evaluation technique is reconfiguration-based fault-injection (FI) analysis. However, most FI analyses are inadequate for the evaluation of a dependable circuit because they don't consider fault accumulation. The critical issue is the reconfiguration time for injecting many faults. This paper presents an FI analysis system using frame-based partial reconfiguration and a bootstrap method to accelerate evaluation. As a result, our system can accelerate FI time by about a factor of 5 ~ 10 relative to the full-reconfiguration FI system. Further, the number of reconfiguration times is reduced to one out of several dozen by applying the bootstrap method.
Yoshihiro Ichinomiya, Kohei Takano, Motoki Amagasaki, Morihiro Kuga, Masahiro Iida, Toshinori Sueyoshi
FPT5
2012 Fault Recovery Technique for TMR Softcore Processor System Using Partial Reconfiguration
Makoto Fujino, Hiroki Tanaka, Yoshihiro Ichinomiya, Motoki Amagasaki, Morihiro Kuga, Masahiro Iida, Toshinori Sueyoshi
ICA3PP (1)6
2012 A Bitstream Relocation Technique to Improve Flexibility of Partial Reconfiguration
Yoshihiro Ichinomiya, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
ICA3PP (1)3
2012 Evaluation of fault tolerant technique based on homogeneous FPGA architecture
Yuki Nishitani, Kazuki Inoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC4
2011 An Easily Testable Routing Architecture and Efficient Test Technique
abstract
Generally, a programmable LSI such as an FPGA is difficult to test as compared to an ASIC. There are two major reasons for this. One is that automatic test pattern generator (ATPG) cannot be used because of the programmability of the FPGA. The other reason is that the FPGA architecture is very complex. In this paper, we propose a novel FPGA architecture that will simplify the testing of the device. The architecture is very simple and has several types of circuit blocks and orderly wire connections. This paper also presents efficient test configurations for our proposed architecture. We tested the interconnects of our architecture by using our configurations and achieved 100% test coverage for a short test time.
Kazuki Inoue, Hiroki Yosho, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
FPL4
2011 An easily testable routing architecture of FPGA
abstract
Generally, a programmable LSI such as an FPGA is difficult to test compared to an ASIC. There are two major reasons for this. One is that automatic test pattern generator (ATPG) cannot be used because of the programmability of the FPGA. The other reason is that the FPGA architecture is very complex. In this paper, we propose a new FPGA architecture that will simplify the testing of the device. The base of our architecture is general island-style FPGA architecture, but it consists of a few types of circuit blocks and orderly wire connections. This paper also presents efficient test configurations for our proposed architecture. We tested the interconnects of our architecture by using our configurations and achieved 100% test coverage for a short test time.
Masahiro Iida, Kazuki Inoue, Motoki Amagasaki, Toshinori Sueyoshi
VLSI-SoC1
2010 Improving the Robustness of a Softcore Processor against SEUs by Using TMR and Partial Reconfiguration
abstract
SRAM-based field programmable gate arrays (FPGAs) are vulnerable to a single event upset (SEU), which is induced by radiation effect. This paper presents a technique for ensuring reliable softcore processor implementation on SRAM-based FPGAs. Although an FPGA is susceptible to SEUs, these faults can be corrected as a result of its reconfigurability. We propose techniques for SEU mitigation and recovery of a softcore processor using triple modular redundancy (TMR) and partial reconfiguration (PR) with state synchronization. By carrying out an experiment, we confirm that a faulty softcore processor can be recovered and synchronized with other softcore processors. The proposed technique requires 4.315 times the resource usage and 62.491% of the operating frequency of the base processor. However, the proposed recovery process only takes 6 μs under TMR and PR. As a result of reliability estimation, the proposed system achieved about 2.713 times longer MTBF comparing with the previous system.
Yoshihiro Ichinomiya, Shiro Tanoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
FCCM4
2010 First Prototype of a Genuine Power-Gatable Reconfigurable Logic Chip with FeRAM Cells
abstract
An advantage of a RLD (Reconfigurable logic device) such as an FPGA (Field programmable gate array) is that it can be customized after being manufactured. However, there is a problem related to standby power when using it in SoC used in embedded systems. Power gating, which is one of the power reduction techniques, is difficult to use in SRAM-based RLDs because of the high overhead - data hibernation and reconfiguration time - and SRAM being volatile. In this paper, we describe a chip that we developed - are configurable logic chip based on FeRAM (Ferroelectric random access memory) technology. The chip employs island-style routing architecture and uses a variable grain logic cell as a logic block. A NV-FF (Non-Volatile FlipFlop), which contains FeRAM, aFF, and power-gating control circuits, is used as configuration memory. The NV-FF can transmit data between FeRAM and FF automatically when power to the chip is turned off/on. Thus, chip-level power gating is possible. The hibernate/restore time is less than 1 ms. The chip has 18 x 18 logic blocks and an area of 54.76 mm2.
Masahiro Koga, Masahiro Iida, Motoki Amagasaki, Yoshinobu Ichida, Mitsuro Saji, Jun Iida, Toshinori Sueyoshi
FPL2
2010 COGRE: A Configuration Memory Reduced Reconfigurable Logic Cell Architecture for Area Minimization
abstract
Because of the redundancy factors of FPGAs, there is a performance gap between FPGAs and ASICs. In this paper, we propose a small-memory logic cell, COGRE, to minimize the FPGA area. Our approach is to investigate the appearance ratio of the logic functions in a circuit implementation. Moreover, we group the logic functions on the basis of the NPN-equivalence class. The results of our investigation show that only small portions of the NPN-equivalence class can cover large portions of the logic functions used to implement circuits. Further, we found that NPN-equivalence classes with a high appearance ratio can be implemented by using a small number of AND gates, OR gates, and NOT gates. On the basis of this observation, we develop 5-input and 6-inputCOGRE architectures composed of several NAND gates and programmable inverters. The experimental results show that the logic area in 6-COGRE is 46.3% smaller than that in 6-LUT. The logic area of 5-COGRE is 32.6% smaller than that of 5-LUT and 10.0% smaller than that of 4-LUT. Further, the total number of configuration memory bits in 6-COGRE is32.1% smaller than the number of configuration memory bits in 6-LUT.
Yasuhiro Okamoto, Yoshihiro Ichinomiya, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
FPL4
2010 A robust reconfigurable logic device based on less configuration memory logic cell
abstract
As the size of integrated circuit has reached the nanoscale, embedded memories are more sensitive to single event upset (SEU), because of their low threshold voltage. In particular field-programmable gate arrays (FPGAs), which contain large amounts of configuration memories to implement customer circuits, are more likely to suffer from soft errors caused by SEU. In this research, we first develop a Hamming code based error detect and correct (EDC) circuit that can prevent the configuration memory of a reconfigurable device from SEU. We then propose a novel reconfigurable logic element, namely COGRE, which will use much less configuration memory than the conventional FPGA 4-, 5- or 6-LUTs (lookup tables). Evaluation revealed that compared to the 6-LUT FPGAs with triple modular redundancy (TMR) configuration memory blocks, the 5- and 6-input proposed architecture save about 75.44 and 74.29% memories on average, respectively. And the dependability of the proposed architectures is about 6.8 to 10 times better than the LUTs with a tile level TMR structure on average. Moreover, with the consideration of the on the fly scrubbing advantage of the EDC, SEUs cannot be accumulated, so a much higher dependability can be achieved.
Qian Zhao 0001, Yoshihiro Ichinomiya, Yasuhiro Okamoto, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
FPT5
2010 Power-aware FPGA routing fabrics and design tools
abstract
The performance of field-programmable gate arrays (FPGAs) has been significantly improved due to a new process technology. However, several problems have arisen in the new generation FPGAs. Specifically, the issue of power consumption is a serious issue, because FPGAs have many routing resources. We report on the improvement of both the FPGA routing structure and electronic design automation (EDA) tools in order to solve this issue. In order to reduce the power consumption, high activity nets are assigned to low load lines, which is the routing structure used for small-world networks in FPGAs. In addition, the clustering and routing algorithms of the EDA tools are improved to complement the routing structure. This report demonstrates that the power can be reduced. Based on evaluation results, a maximum power consumption improvement of 48.4% was obtained, and the average improvement was 22.9% when using the proposed routing structure and EDA tools.
Shoichi Nishida, Jyunya Eto, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi
VLSI-SoC4
2010 A Variable-Grain Logic Cell and Routing Architecture for a Reconfigurable IP Core
abstract
In the present study, we investigate the use of reconfigurable logic devices (RLDs) as intellectual properties (IPs) for system on a chip (SoC). Using RLDs, SoCs can achieve both high performance and high flexibility. However, conventional RLDs have problems related to performance, area, and power consumption. In order to resolve these problems, we investigated the features of RLD architecture. RLDs are classified into fine-grained and coarse-grained devices based on their architecture. Generally, the granularity of an RLD is limited to either type, which means that a device can only achieve high performance in applications that are suited to its architecture. Therefore, we propose a variable-grain logic cell (VGLC) architecture that can overcome the trade-off between fine-grained and coarse-grained architectures, which are required for the implementation of random and arithmetic logics, respectively. The VGLC is based on a 4-bit adder including configuration bits, which can perform arithmetic and random logic operations unlike the LUT. In the present paper, a local interconnection architecture for the VGLC is proposed. Several types of local interconnections composed of different crossbars are compared, and the trade-off between hardware resources and flexibility is discussed. Using local interconnection, the routing area is reduced by a maximum of 49%.
Kazuki Inoue, Qian Zhao 0001, Yasuhiro Okamoto, Hiroki Yosho, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi
ACM Trans. Reconfigurable Technol. Syst.6
2009 Improvement of Execution Efficiency on the MX Core
abstract
SIMD (Single Instruction/Multiple Data) type processors have the advantage of smaller area as compared with a general processor, DSP and many MIMD (Multiple Instruction/Multiple Data) type processors. On the other hand, the performance depends on the parallel degree of data in an application. MX Core which was developed by Renesas technology Corp., is a massively parallel SIMD type accelerator. We propose a method to improve execution efficiency of the MX Core in this paper. Our methodology includes optimization of calculation precision and change data transfer structure to MIMD type. As a result of evaluation, we improved parallel operation degree. We achieved a speedup of 2.92 times at the maximum and double parallel degree improvement in RSA than conventional implementation technique. The proposal structure reduced data transfer processing of IMDCT by 90%, and speed up processing time of IMDCT by 2.65 times compared with traditional implementation of the MX Core.
Mitsutaka Nakano, Masahiro Iida, Toshinori Sueyoshi
PDCAT2
2007 A Novel Technique to Create Energy-Efficient Contexts for Reconfigurable Logic
abstract
High power consumption is a constraining factor for the growth of programmable logic devices. We propose two techniques in order to reduce power consumption. The first is a technique for creating contexts. This technique uses data-dependent circuits and wire sharing between contexts. The second is a technique for switching the contexts. In this paper, we evaluate the capability of the two techniques to reduce power consumption using a multi-context logic device. As a result, as compared with the original circuit, our multi-context circuits can reduce the power consumption by 9.1% on an average and by a maximum of 19.0%. Furthermore, applying our resource sharing technique to these circuits, we achieved a reduction of 10.6% on an average and a maximum reduction of 18.8%.
Hiroshi Shinohara, Hideaki Monji, Masahiro Iida, Toshinori Sueyoshi
FCCM3
2007 A Novel Technique to Create Energy-Efficient Contexts for Reconfigurable Logic
abstract
High power consumption is a constraining factor for the growth of programmable logic devices. We propose two techniques in order to reduce power consumption. The first is a technique for creating contexts. This technique uses data-dependent circuits and wire sharing between contexts. The second is a technique for switching the contexts. In this paper, we evaluate the capability of the two techniques to reduce power consumption using a multi-context logic device. As a result, as compared with the original circuit, our multi-context circuits can reduce the power consumption by 9.1% on an average and by a maximum of 19.0%. Furthermore, applying our resource sharing technique to these circuits, we achieved a reduction of 10.6% on an average and a maximum reduction of 18.8%.
Hiroshi Shinohara, Hideaki Monji, Masahiro Iida, Toshinori Sueyoshi
FCCM3
2007 A Variable Grain Logic Cell Architecture for Reconfigurable Logic Cores
abstract
Reconfigurable Logic Devices are classified as the fine-grained or coarse-grained type on the basis of their basic logic cell architecture. In general, each architecture has its own merit; therefore, it is difficult to achieve a balance between the operation speed and implementation area in various applications. In this paper, we propose a Variable Grain Logic Cell (VGLC) architecture, which consists of a 4-bit ripple carry adder with configuration memory bits and also develop technology mapping tool. Its key feature is the variable granularity being a trade-off between coarse-grained and fine-grained types required for the implementation arithmetic and random logic, respectively. As a result, critical path delay, and number of configuration memory bits are reduced by 49.7%, and 48.5%, respectively, in the benchmark circuits.
Motoki Amagasaki, Ryoichi Yamaguchi, Kazunori Matsuyama, Masahiro Iida, Toshinori Sueyoshi
FPL4
2007 An Embedded Reconfigurable Logic Core based on Variable Grain Logic Cell Architecture
abstract
Reconfigurable computing is becoming increasingly attractive for many applications. It involves the use of reconfigurable logic devices (RLDs). RLDs are classified as the fine-grained or coarse-grained type on the basis of their basic logic cell architecture. In general, each architecture has its own merit; therefore, it is difficult to achieve a balance between the operation speed and implementation area in various applications. In this paper, we propose a variable grain logic cell (VGLC) architecture, which consists of a 4-bit ripple carry adder with configuration memory bits and also develop technology mapping tool. Its key feature is the variable granularity being a tradeoff between coarse-grained and fine-grained types required for the implementation arithmetic and random logic, respectively. As a result, critical path delay, and number of configuration memory bits are reduced by 49.7%, and 48.5%, respectively, in the benchmark circuits. In addition, when implementing DSP benchmarks on trials, the result is comparable with the highest performance processors today.
Yoshiaki Satou, Motoki Amagasaki, Hiroshi Miura, Kazunori Matsuyama, Ryoichi Yamaguchi, Masahiro Iida, Toshinori Sueyoshi
FPT6
2006 Effective clustering technique to optimize routability of outer cluster nets
abstract
We also study a clustering technique for a cluster-based FPGA to optimize the routability of outer cluster nets. Prior research of FPGA clustering aims to take in inter-cluster connections to utilize the various advantages of the local interconnection in the cluster. Taking in the inter-cluster connections to the cluster can improve the FPGA speed and area, and can lighten the burden of the placement and routing tool. However, many inter-cluster connections still remain in the outside cluster. The condition of these connections give the performance of the circuit big influence. Therefore, we propose the effective clustering technique to optimize routability of outer cluster nets which are not taken in. In order to reduce the used routing resources in FPGA, our technique uses two evaluation functions. One evaluation function can be reduced routing resources in the outside cluster. The second evaluation function can utilize various characteristics of the local routing resources in the inside cluster. Our clustering technique has the unusual ability in the optimization of routing resources concurrently. As a result, our method resulted in 21.4% improvement in terms of the routing area (16.6% on average), and the critical path delay is reduced by 28.0% (11.8% on average) compared to existing clustering method in the benchmark circuits.
Masaki Kobata, Masahiro Iida, Toshinori Sueyoshi
FPGA2
2006 Evaluation of Variable Grain Logic Cell Architecture for Reconfigurable Device
abstract
Reconfigurable logic devices are usually classified on the basis of their basic logic cell architecture as fine-grained or coarse-grained. In general, each architecture is suitable on its own merit; therefore, it is difficult to achieve a balance between the operation speed and area-efficiency in applications. In order to solve this problem, we propose a new logic cell architecture based on a 4-bit ripple carry adder that includes configuration memory bits. This is called the variable grain logic cell architecture, VGLC. It is possible to realize two features by using the VGLC: one is a high device speed of a coarse-grained cell and the other is the versatile logic of a fine-grained cell. This paper demonstrates the transistor-level optimization of our proposed logic cell. Moreover, based on the results of the evaluation, the authors show that the critical path delay can be reduced by a maximum of 37% when using the proposed logic cell architecture is used in a 32-bit multiplier
Motoki Amagasaki, Takurou Shimokawa, Kazunori Matsuyama, Ryoichi Yamaguchi, Hideaki Nakayama, Naoto Hamabe, Masahiro Iida, Toshinori Sueyoshi
VLSI-SoC7
2005 Applying the Small-World Network to Routing Structure of FPGAs
abstract
The degree of integration and the operating frequency of programmable logic have improved dramatically with the development of new process technologies. However, for the deep sub-micron processes, the delay, reliability, cost, and power tend to be determined by interconnections. In conventional programmable logic, reducing the number of switches on a critical path is important because the wiring delay is considerably smaller than the switch delay. However, to achieve a decrease in the critical path delay it is also necessary to consider the wiring delay for the deep sub-micron processes. This paper proposes a novel routing structure using a small-world network structure for the interconnection of programmable logic. This paper demonstrates that the critical path delay can be reduced. Based on the results of an evaluation, the authors show that the critical path delay can be reduced by a maximum of 15% and the amount of routing resources can be reduced by a maximum of 23% when using the small-world network structure.
Hisashi Tsukiashi, Masahiro Iida, Toshinori Sueyoshi
FPL2