EDBT 2026 Demo / reviewers in the wild / expert
Toshinori Sueyoshi
dblp:63/4349
· DBLP profile ↗
43ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Electronic design automation · 55% Reconfigurable computing and FPGAs · 35% Interconnection networks and networks-on-chip · 9% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation › physical design › placement and routing
FPGA placement and routing |
0.2 | 1 | 2013 | A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013 |
Electronic design automation
physical design |
0.2 | 1 | 2013 | A novel FPGA design framework with VLSI post-routing performance analysis (abstract only) · FPGA 2013 |
Reconfigurable computing and FPGAs › FPGA physical design
FPGA clustering |
0.1 | 1 | 2006 | Effective clustering technique to optimize routability of outer cluster nets · FPGA 2006 |
Electronic design automation › physical design › routing › routability
routability optimization |
0.1 | 1 | 2006 | Effective clustering technique to optimize routability of outer cluster nets · FPGA 2006 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 2001 | Recursive Diagonal Torus: An Interconnection Network for Massively Parallel Computers · IEEE Trans. Parallel Distributed Syst. 2001 |
Interconnection networks and networks-on-chip
routing algorithms |
0.0 | 1 | 2001 | Recursive Diagonal Torus: An Interconnection Network for Massively Parallel Computers · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel architecture
massively parallel processor |
0.0 | 1 | 2001 | Recursive Diagonal Torus: An Interconnection Network for Massively Parallel Computers · IEEE Trans. Parallel Distributed Syst. 2001 |
Methods — techniques the papers use, named apart from their topics
object-oriented programming · 0.2HDL template generation · 0.2evaluation functions · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | hCODE 2.0: An open-source toolkit for building efficient FPGA-enabled cloudsabstractMajor cloud service providers have started employing field-programmable gate arrays (FPGAs) to implement high-performance and low-power-consumption cloud capability. However, building or utilizing an FPGA-enabled cloud is still challenging due to the lack of fundamental tools. In our previous work, we proposed an hCODE base system for managing portable accelerator IPs on different hardware. In this paper, we extend the previous work and introduce the hCODE 2.0, which is an open-source toolkit for building efficient FPGA-enabled clouds. First, we provide a fundamental toolkit to simplify HW project management and FPGA management at a cluster scale. Second, we implement on-chip resource virtualization and accelerator scheduling capabilities to show possibilities of improving FPGA utilization efficiency with our tools. Qian Zhao 0001, Hendarmawan, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPT | 6 |
| 2016 | hCODE: An open-source platform for FPGA acceleratorsabstractField-programmable gate arrays (FPGAs) have demonstrated great speed performance and power efficiency advantages over conventional computers in various domains. However, it is still difficult for general software engineers to employ FPGA-based hardware accelerators because of the gap between hardware and software development methods. In this paper, we propose a heterogeneous computing oriented development environment (hCODE) to simplify the creation, sharing, and project integration of hardware accelerators. The hCODE defines hardware specifications and interface design rules. Hardware developers can provide designs that follow these rules, allowing software engineers to easily search, download, and integrate accelerators in their applications without caring about the details of the hardware. Qian Zhao 0001, Takuya Nakamichi, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPT | 6 |
| 2016 | A novel soft error tolerant FPGA architectureabstractDue to reaching the nanoscale transistor size, effect of single event upset (SEU) to the memory has become conspicuous. In small device geometries, a single particle strike might affect multiple adjacent cells in a memory array resulting in a multiple bit upset (MBU). Traditional fault tolerance technologies such as triple modular redundancy (TMR) and error correcting code (ECC) occupy the large area and have vulnerability to MBU. In this research, we propose DMR based error correct circuit and employ a combination of proposed circuit and the interleaving technique to mitigate MBU. In addition, we explain soft error simulator developed to calculate bit interleaving distance. The results show that the area of proposed circuit is the smallest when we compare the proposed circuit, ECC based error correct circuit and TMR. Simulation results show that the interleaving distance which can conceal all MBU patterns is 4. Motoki Amagasaki, Yuji Nakamura, Takuya Teraoka, Masahiro Iida, Toshinori Sueyoshi |
VLSI-SoC | 5 |
| 2015 | Architecture exploration of 3D FPGA to minimize internal layer connectionabstractA three-dimensional (3D) integration based on wafer-to-wafer bonding using through-silicon vias (TSVs) has been developed for the fabrication of new 3D large-scale integrated chips. To balance between cost and performance, and to explore 3D field-programmable gate array (FPGA) with realistic 3D integration processes, we propose spatially distributed and functionally distributed types of 3D FPGA architectures. The functionally distributed architecture consists of two wafers, a logic layer and a routing layer, and is stacked by a face-down process technology. Since vertical wires pass through microbumps, no TSVs are needed. In contrast, the spatially distributed architecture is divided into multiple layers with the same structure, unlike in the functionally distributed type. This architecture can be expanded to more than two layers by stacking multiples of the same die. The goal of this paper is to elucidate the advantages and disadvantages of these two types of 3D FPGAs. According to our evaluation, when only two layers are used, the functionally distributed architecture is more effective. When higher performance is achieved by using more than two layers, the spatially distributed architecture achieves better performance. Motoki Amagasaki, Yuto Takeuchi, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
VLSI-SoC | 6 |
| 2014 | A logic cell architecture exploiting the shannon expansion for the reduction of configuration memoryabstractMost modern field-programmable gate arrays (FPGAs) employ a look-up table (LUT) as their basic logic cell. Although a k-input LUT can implement any k-input logic, its functionality relies on a large amount of configuration memory. As FPGA scales improve, the increased quantity of configuration memory cells required for FPGAs will require a larger area and consume more power. Moreover, the soft-error rate per device will also increase as more configuration memory cells are embedded. We propose scalable logic modules (SLMs), logic cells requiring less configuration memory, reducing configuration memory by making use of partial functions of Shannon expansion for frequently appearing logics. Experimental results show that SLM-based FPGAs use much less configuration memory and have smaller area than conventional LUT-based FPGAs. Qian Zhao 0001, Kyosei Yanagida, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPL | 6 |
| 2014 | A novel three-dimensional FPGA architecture with high-speed serial communication linksabstractThree-dimensional (3D) integrated circuit technology is expected to offer continual improvement to very-large-scale integration performance as the process of miniaturization approaches physical limits. However, because the through-silicon vias (TSVs) that are used to create interlayer vertical connections are much larger area than transistors, there is an inherent tradeoff between connectivity and small size. Field-programmable gate arrays (FPGAs) are particularly noted for requiring a high level of routing resources, which means that it is unrealistic to make the same number of connections vertically as horizontally. In previous research, we proposed a method for creating a two-layer compact 3D FPGA with face-down integration (the base FPGA). In this paper, we discuss stacking multiple base FPGAs by the face-up method and propose a method for achieving highspeed interlayer communications with TSV serial connections. The proposed architecture improves FPGA performance by using smaller TSVs. The evaluation results show that the proposed 3D FPGA can achieve a total area that is as low as 67% the equivalent two-dimensional FPGA. Takuya Kajiwara, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morituro Kuga, Toshinori Sueyoshi |
FPT | 6 |
| 2014 | Zyndroid: An Android platform for software/hardware coprocessingabstractHigh performance is required of many Android systems because embedded systems written for this operating system are used in several fields and rely on increasingly complicated processing. To accommodate this, we present a software/hardware (SW/HW) coprocessing platform implemented on a programmable system-on-a-chip (Xilinx Inc.: Zynq). This platform provides a unified architecture, extended OS kernel, application framework, and application distribution model to simplify the development and use of Android SW/HW coprocessing applications. Susumu Mashimo, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPT | 5 |
| 2014 | Blokus Duo engine on a ZynqabstractIn this article, we present a design of a Blokus Duo engine for the ICFPT 2014 Design Competition. Our design is implemented on a Xilinx Zynq-7000 SoC ZC706 Evaluation Kit and we employ the minimax algorithm with alpha-beta pruning. The ARM processor runs the search algorithm, and the handwritten hardware accelerator calculate within 1 second under the competition constraint. One of the keys to a stronger Blokus Duo player is to evaluate more states of a game; our Blokus Duo engine evaluates 12.3 times as many nodes of a game search tree as the Intel Core i7-3770T. Susumu Mashimo, Kansuke Fukuda, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPT | 6 |
| 2013 | A novel FPGA design framework with VLSI post-routing performance analysis (abstract only)abstractThe most widely used open-source field-programmable gate array (FPGA) placement and routing tool is VPR, which can define the target FPGA, perform placement and routing, and report area and timing information. However, it cannot be used in FPGA IP design efficiently for two reasons. First, for most newly developed FPGA architectures, VPR cannot support them directly. Modifying the C-coded VPR for using it to evaluate a number of new architectures requires a long time. Second, the accuracy of the VPR performance results is not enough for the evaluation of a complete synthesizable FPGA IP in the design that targets the productions of LSI. We propose a FPGA design framework that in particular improves FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely EasyRouter. EasyRouter is developed using the C# language. When an object-oriented programming method is used, the source codes are fewer and easier manage compared to VPR, which shortens the development time. By using simple HDL templates, EasyRouter can automatically generate entire chip HDL codes and the configuration bitstream. With these files, the FPGA IP can be evaluated with commercial VLSI CADs with high accuracy and reliability. Qian Zhao 0001, Kazuki Inoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPGA | 6 |
| 2013 | Defect-robust FPGA architectures for intellectual property cores in system LSIabstractIn this paper, we propose fault-tolerant field-programmable gate array (FPGA) architectures and their computer-aid design (CAD) for intellectual property (IP) cores in system large-scale integration (LSI). Unlike discrete FPGAs, in which the integration scale can be made relatively large, programmable IP cores must correspond to arrays of various sizes. The key features of our architectures are regular tile structure, spare modules and bypass wires for fault avoidance, and configuration mechanism for single-cycle reconfiguration. In addition, we develop routing tools, namely EasyRouter for proposed architecture. This tool can handle various array sizes corresponding to developed programmable IP cores. In this evaluation, we compared the performances of conventional FPGA and the proposed fault-tolerant FPGA architectures. On average, our architectures have less than 2.2 times the area and 1.3 times the delay compared with conventional FPGA architectures. At the same time, conventional FP-GAs cannot tolerate faults, whereas our architectures perform with a 90% success rate in fault avoidance for a ratio of faulty tiles of 1% or less. Motoki Amagasaki, Kazuki Inoue, Qian Zhao 0001, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPL | 6 |
| 2013 | An automatic FPGA design and implementation frameworkabstractConventional FPGA design and implementation processes involve two separate flows. The FPGA architecture is determined by academic FPGA design flow. However, in the implementation phase, commercial VLSI design flow are used. In this research, we propose an FPGA design framework in order to improve synthesizable FPGA IP design efficiency. A novel FPGA routing tool is developed in this framework, namely the EasyRouter, which can bridge the two flows efficiently. With this design flow, accurate physical information can be reported when a new FPGA IP architecture is evaluated with reliable commercial VLSI CADs. Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPL | 5 |
| 2013 | Three-dimensional stacking FPGA architecture using face-to-face integrationabstractIn recent years, as VLSI process scales have developed into deep sub-micrometer dimensions, routing delay problems have become critical. For reconfigurable logic devices (RLDs) like field-programmable gate arrays (FPGAs) in particular, routing resources occupy major parts of the available area and hinder performance. In order to balance cost and performance, and to explore 3D FPGA architectures with realistic 3D LSI processes, we proposed a novel two-layers 3D FPGA architecture based on 3D connections on logic block input and output pins. Evaluation shows that this novel RLD with two layers of 3D routing architecture uses 48.75% less on-board area and 30.54% less critical path delay than does a conventional 2D 4-lookup table island-style FPGA on average. Tetsuro Hamada, Qian Zhao 0001, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
VLSI-SoC | 6 |
| 2012 | Designing Flexible Reconfigurable Regions to Relocate Partial BitstreamsabstractCurrent commercial SRAM-based FPGAs, such as Virtex-6 and Stratix-V, can perform dynamic partial reconfiguration (DPR). Partial reconfiguration (PR) can change a part of the device without reconfiguring the whole chip. Thus, we can switch the part of system with continuing the operation. However, the authorized design flow by Xilinx creates different PR bit stream (PRB) for each partially reconfigurable region (PRR) even if it is the same circuit. This indicates that N × M PRBs must be prepared to implement M types modules on N PRRs. This increases design time and memory usage to store PRBs. This paper presents a uniforming design technique for PRRs to relocate a PRB among them. In addition, uniformed PRRs can be used to implement large module by combining adjacent PRRs. In this work, we use Xilinx Virtex-6 XC6VLX240T and Integrated Software Environment 13.3 (ISE) to verify the proposed technique. Yoshihiro Ichinomiya, Sadaki Usagawa, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FCCM | 6 |
| 2012 | Fault detection and avoidance of FPGA in various granularitiesabstractAlthough redundancy techniques are generally used to provide fault tolerance for the system large scale integrations (LSIs), these techniques have significantly costs. However, field programmable gate arrays (FPGAs) can easily provide high reliability due to their reconfiguration ability. The present paper proposes an effective fault detection method for the global interconnects of the FPGA. In the proposed method, by combining different kinds of test patterns, the fault resources can be identified. We also developed placement and routing tools to avoid fault resources in two cases, namely, tile-level avoidance and multiplexer-level avoidance. In the evaluation of the proposed technique, the proposed detection method diagnosed a defect multiplexer with six test configurations. We found that the fault FPGA can achieve the same performance as normal FPGA in multiplexer-level avoidance. Kazuki Inoue, Yuki Nishitani, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FPL | 6 |
| 2012 | Accelerated evaluation of SEU failure-in-time using frame-based partial reconfigurationabstractSRAM-based field programmable gate arrays (FPGAs) are vulnerable to soft-error. To improve circuit dependability, various dependable design techniques have been studied. By the same token, evaluation techniques are required to ensure dependability. The most popular evaluation technique is reconfiguration-based fault-injection (FI) analysis. However, most FI analyses are inadequate for the evaluation of a dependable circuit because they don't consider fault accumulation. The critical issue is the reconfiguration time for injecting many faults. This paper presents an FI analysis system using frame-based partial reconfiguration and a bootstrap method to accelerate evaluation. As a result, our system can accelerate FI time by about a factor of 5 ~ 10 relative to the full-reconfiguration FI system. Further, the number of reconfiguration times is reduced to one out of several dozen by applying the bootstrap method. Yoshihiro Ichinomiya, Kohei Takano, Motoki Amagasaki, Morihiro Kuga, Masahiro Iida, Toshinori Sueyoshi |
FPT | 6 |
| 2012 | Fault Recovery Technique for TMR Softcore Processor System Using Partial Reconfiguration
Makoto Fujino, Hiroki Tanaka, Yoshihiro Ichinomiya, Motoki Amagasaki, Morihiro Kuga, Masahiro Iida, Toshinori Sueyoshi |
ICA3PP (1) | 7 |
| 2012 | A Bitstream Relocation Technique to Improve Flexibility of Partial Reconfiguration
Yoshihiro Ichinomiya, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
ICA3PP (1) | 5 |
| 2012 | Evaluation of fault tolerant technique based on homogeneous FPGA architecture
Yuki Nishitani, Kazuki Inoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
VLSI-SoC | 6 |
| 2011 | Anomaly Detection Using Chi-square Values Based on the Typical Features and the Time DeviationabstractIn the research of the anomaly detection system analyzing the packet header on the Internet, previous researches have proposed the anomaly detection system using chi-square values in terms of the source IP address and/or the destination port number. In these previous researches, the chi-square values were calculated from one feature causing the degradation in the False-Positive when the same symbol appears sequentially. Therefore, we propose the anomaly detection technique using chi-square values based on multi features. We also propose dynamic BIN division technique to deal with the traffic fluctuations such as day and night traffic differences. Applying our method, the chi-square values based on the time division were able to decrease the False-Positive. Our method was also able to adapt the traffic variations by applying the dynamic BIN division technique. Shunsuke Oshima, Takuo Nakashima, Toshinori Sueyoshi |
AINA | 3 |
| 2011 | An Easily Testable Routing Architecture and Efficient Test TechniqueabstractGenerally, a programmable LSI such as an FPGA is difficult to test as compared to an ASIC. There are two major reasons for this. One is that automatic test pattern generator (ATPG) cannot be used because of the programmability of the FPGA. The other reason is that the FPGA architecture is very complex. In this paper, we propose a novel FPGA architecture that will simplify the testing of the device. The architecture is very simple and has several types of circuit blocks and orderly wire connections. This paper also presents efficient test configurations for our proposed architecture. We tested the interconnects of our architecture by using our configurations and achieved 100% test coverage for a short test time. Kazuki Inoue, Hiroki Yosho, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi |
FPL | 5 |
| 2011 | An easily testable routing architecture of FPGAabstractGenerally, a programmable LSI such as an FPGA is difficult to test compared to an ASIC. There are two major reasons for this. One is that automatic test pattern generator (ATPG) cannot be used because of the programmability of the FPGA. The other reason is that the FPGA architecture is very complex. In this paper, we propose a new FPGA architecture that will simplify the testing of the device. The base of our architecture is general island-style FPGA architecture, but it consists of a few types of circuit blocks and orderly wire connections. This paper also presents efficient test configurations for our proposed architecture. We tested the interconnects of our architecture by using our configurations and achieved 100% test coverage for a short test time. Masahiro Iida, Kazuki Inoue, Motoki Amagasaki, Toshinori Sueyoshi |
VLSI-SoC | 4 |
| 2010 | Early DoS/DDoS Detection Method using Short-term StatisticsabstractEarly detection methods are required to prevent the DoS / DDoS attacks. The detection methods using the entropy have been classified into the long-term entropy based on the observation of more than 10,000 packets and the short-term entropy that of less than 10,000 packets. The long-term entropy have less fluctuation leading to easy detection of anomaly accesses using the threshold, while having the defects in detection at the early attacking stage and of difficulty to trace the short term attacks. In this paper, we propose and evaluate the DoS/DDoS detection method based on the short-term entropy focusing on the early detection. Firstly, the pre-experiment extracted the effective window width; 50 for DDoS and 500 for slow DoS attacks. Secondly, we showed that classifying the type of attacks can be made possible using the distribution of the average and standard deviation of the entropy. In addition, we generated the pseudo attacking packets under a normal condition to calculate the entropy and carry out a test of significance. When the number of attacking packets is equal to the number of arriving packets, the high detection results with False-negative = 5% was extracted, and the effectiveness of the proposed method was shown. Shunsuke Oshima, Takuo Nakashima, Toshinori Sueyoshi |
CISIS | 3 |
| 2010 | Improving the Robustness of a Softcore Processor against SEUs by Using TMR and Partial ReconfigurationabstractSRAM-based field programmable gate arrays (FPGAs) are vulnerable to a single event upset (SEU), which is induced by radiation effect. This paper presents a technique for ensuring reliable softcore processor implementation on SRAM-based FPGAs. Although an FPGA is susceptible to SEUs, these faults can be corrected as a result of its reconfigurability. We propose techniques for SEU mitigation and recovery of a softcore processor using triple modular redundancy (TMR) and partial reconfiguration (PR) with state synchronization. By carrying out an experiment, we confirm that a faulty softcore processor can be recovered and synchronized with other softcore processors. The proposed technique requires 4.315 times the resource usage and 62.491% of the operating frequency of the base processor. However, the proposed recovery process only takes 6 μs under TMR and PR. As a result of reliability estimation, the proposed system achieved about 2.713 times longer MTBF comparing with the previous system. Yoshihiro Ichinomiya, Shiro Tanoue, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
FCCM | 6 |
| 2010 | First Prototype of a Genuine Power-Gatable Reconfigurable Logic Chip with FeRAM CellsabstractAn advantage of a RLD (Reconfigurable logic device) such as an FPGA (Field programmable gate array) is that it can be customized after being manufactured. However, there is a problem related to standby power when using it in SoC used in embedded systems. Power gating, which is one of the power reduction techniques, is difficult to use in SRAM-based RLDs because of the high overhead - data hibernation and reconfiguration time - and SRAM being volatile. In this paper, we describe a chip that we developed - are configurable logic chip based on FeRAM (Ferroelectric random access memory) technology. The chip employs island-style routing architecture and uses a variable grain logic cell as a logic block. A NV-FF (Non-Volatile FlipFlop), which contains FeRAM, aFF, and power-gating control circuits, is used as configuration memory. The NV-FF can transmit data between FeRAM and FF automatically when power to the chip is turned off/on. Thus, chip-level power gating is possible. The hibernate/restore time is less than 1 ms. The chip has 18 x 18 logic blocks and an area of 54.76 mm2. Masahiro Koga, Masahiro Iida, Motoki Amagasaki, Yoshinobu Ichida, Mitsuro Saji, Jun Iida, Toshinori Sueyoshi |
FPL | 7 |
| 2010 | COGRE: A Configuration Memory Reduced Reconfigurable Logic Cell Architecture for Area MinimizationabstractBecause of the redundancy factors of FPGAs, there is a performance gap between FPGAs and ASICs. In this paper, we propose a small-memory logic cell, COGRE, to minimize the FPGA area. Our approach is to investigate the appearance ratio of the logic functions in a circuit implementation. Moreover, we group the logic functions on the basis of the NPN-equivalence class. The results of our investigation show that only small portions of the NPN-equivalence class can cover large portions of the logic functions used to implement circuits. Further, we found that NPN-equivalence classes with a high appearance ratio can be implemented by using a small number of AND gates, OR gates, and NOT gates. On the basis of this observation, we develop 5-input and 6-inputCOGRE architectures composed of several NAND gates and programmable inverters. The experimental results show that the logic area in 6-COGRE is 46.3% smaller than that in 6-LUT. The logic area of 5-COGRE is 32.6% smaller than that of 5-LUT and 10.0% smaller than that of 4-LUT. Further, the total number of configuration memory bits in 6-COGRE is32.1% smaller than the number of configuration memory bits in 6-LUT. Yasuhiro Okamoto, Yoshihiro Ichinomiya, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi |
FPL | 5 |
| 2010 | A robust reconfigurable logic device based on less configuration memory logic cellabstractAs the size of integrated circuit has reached the nanoscale, embedded memories are more sensitive to single event upset (SEU), because of their low threshold voltage. In particular field-programmable gate arrays (FPGAs), which contain large amounts of configuration memories to implement customer circuits, are more likely to suffer from soft errors caused by SEU. In this research, we first develop a Hamming code based error detect and correct (EDC) circuit that can prevent the configuration memory of a reconfigurable device from SEU. We then propose a novel reconfigurable logic element, namely COGRE, which will use much less configuration memory than the conventional FPGA 4-, 5- or 6-LUTs (lookup tables). Evaluation revealed that compared to the 6-LUT FPGAs with triple modular redundancy (TMR) configuration memory blocks, the 5- and 6-input proposed architecture save about 75.44 and 74.29% memories on average, respectively. And the dependability of the proposed architectures is about 6.8 to 10 times better than the LUTs with a tile level TMR structure on average. Moreover, with the consideration of the on the fly scrubbing advantage of the EDC, SEUs cannot be accumulated, so a much higher dependability can be achieved. Qian Zhao 0001, Yoshihiro Ichinomiya, Yasuhiro Okamoto, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi |
FPT | 6 |
| 2010 | Power-aware FPGA routing fabrics and design toolsabstractThe performance of field-programmable gate arrays (FPGAs) has been significantly improved due to a new process technology. However, several problems have arisen in the new generation FPGAs. Specifically, the issue of power consumption is a serious issue, because FPGAs have many routing resources. We report on the improvement of both the FPGA routing structure and electronic design automation (EDA) tools in order to solve this issue. In order to reduce the power consumption, high activity nets are assigned to low load lines, which is the routing structure used for small-world networks in FPGAs. In addition, the clustering and routing algorithms of the EDA tools are improved to complement the routing structure. This report demonstrates that the power can be reduced. Based on evaluation results, a maximum power consumption improvement of 48.4% was obtained, and the average improvement was 22.9% when using the proposed routing structure and EDA tools. Shoichi Nishida, Jyunya Eto, Motoki Amagasaki, Masahiro Iida, Morihiro Kuga, Toshinori Sueyoshi |
VLSI-SoC | 6 |
| 2010 | A Variable-Grain Logic Cell and Routing Architecture for a Reconfigurable IP CoreabstractIn the present study, we investigate the use of reconfigurable logic devices (RLDs) as intellectual properties (IPs) for system on a chip (SoC). Using RLDs, SoCs can achieve both high performance and high flexibility. However, conventional RLDs have problems related to performance, area, and power consumption. In order to resolve these problems, we investigated the features of RLD architecture. RLDs are classified into fine-grained and coarse-grained devices based on their architecture. Generally, the granularity of an RLD is limited to either type, which means that a device can only achieve high performance in applications that are suited to its architecture. Therefore, we propose a variable-grain logic cell (VGLC) architecture that can overcome the trade-off between fine-grained and coarse-grained architectures, which are required for the implementation of random and arithmetic logics, respectively. The VGLC is based on a 4-bit adder including configuration bits, which can perform arithmetic and random logic operations unlike the LUT. In the present paper, a local interconnection architecture for the VGLC is proposed. Several types of local interconnections composed of different crossbars are compared, and the trade-off between hardware resources and flexibility is discussed. Using local interconnection, the routing area is reduced by a maximum of 49%. Kazuki Inoue, Qian Zhao 0001, Yasuhiro Okamoto, Hiroki Yosho, Motoki Amagasaki, Masahiro Iida, Toshinori Sueyoshi |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2009 | A novel states recovery technique for the TMR softcore processorabstractThe present paper describes a technique for ensuring re- liable softcore processor implementation on SRAM-based field programmable gate arrays (FPGAs), which can handle the effects of single event upsets (SEUs). We propose the triple modular redundancy (TMR) scheme coupled with dynamic partial reconfiguration to remove SEUs from the configuration memory of the FPGA. Although the FPGA is subject to SEUs, these errors can be corrected as a result of its reconfigurability. Furthermore, we consider the synchronization after a partial reconfiguration using an interrupt process of an RTOS. Experimental results reveal that one faulty softcore processor is recovered and synchronized with the other softcore processors. The present study demonstrates that a softcore processor can recover from an SEU using the proposed dynamic partial reconfiguration and the synchronization process. Shiro Tanoue, Tomoyuki Ishida, Yoshihiro Ichinomiya, Motoki Amagasaki, Morihiro Kuga, Toshinori Sueyoshi |
FPL | 6 |
| 2009 | Improvement of Execution Efficiency on the MX CoreabstractSIMD (Single Instruction/Multiple Data) type processors have the advantage of smaller area as compared with a general processor, DSP and many MIMD (Multiple Instruction/Multiple Data) type processors. On the other hand, the performance depends on the parallel degree of data in an application. MX Core which was developed by Renesas technology Corp., is a massively parallel SIMD type accelerator. We propose a method to improve execution efficiency of the MX Core in this paper. Our methodology includes optimization of calculation precision and change data transfer structure to MIMD type. As a result of evaluation, we improved parallel operation degree. We achieved a speedup of 2.92 times at the maximum and double parallel degree improvement in RSA than conventional implementation technique. The proposal structure reduced data transfer processing of IMDCT by 90%, and speed up processing time of IMDCT by 2.65 times compared with traditional implementation of the MX Core. Mitsutaka Nakano, Masahiro Iida, Toshinori Sueyoshi |
PDCAT | 3 |
| 2008 | Analysis of Queueing Property for Self-Similar TrafficabstractThe scale-invariant burstiness or self-similarity has been found in real network. Relation between self- similarity and network parameter is mainly discussed in the context of application layer, and this self- similarity is caused by the file size of Web servers or the duration of user sessions. Traffic dynamics, however, are mainly generated by physical conditions such as the resource restrainment and the queue size on intermediate routers. The purpose of this research is to provide the basic concept of resource provisioning. In this paper, we have investigated queueing property on the bottleneck link for self-similar traffic using network simulator. We examined the distribution property of Pareto distribution on ns-2 simulator, then analyzing the simulated results, following properties were extracted. Firstly, the traffics with self-similar property are transmitted consuming the queue resources less on the bottleneck link. Secondly, queueing dynamics tend to exceed the linear manner for increase of packet flows. Thirdly, bandwidth expansion improve the consumption of the bottleneck queue while small delay does not affect to consume the queue. Finally, the effective throughput performance gain is measured under the restricted resource conditions. Takuo Nakashima, Toshinori Sueyoshi |
AINA | 2 |
| 2008 | Extraction of Characteristics of Anomaly Accessed IP Packets by the Entropy-Based AnalysisabstractTo defend DoS (denial of service) attacks, the access filtering mechanism is adopted on the end servers or the IDS (intrusion detection system). The difficulty to define the filtering rules comes from the hardness to identify normal and anomaly packets from the incoming packets. The purpose of our research is to explore the early detective method for anomaly accesses based on statistic analysis. In this paper, we firstly define the entropy-based analysis, then analyze the amount of incoming packets to our collage. As the results, we were able to extract the following features for the entropy analysis. Firstly, fluctuations for first octet aggregation lead to similar pattern compared to that of first and second octets aggregation. Secondly, sliding time of 10 minutes of entropy window was sensitive to detect anomaly accesses. Finally, differential entropy detected the small amount of 80/TCP anomaly accesses while analysis of frequency was hard to find that. Takuo Nakashima, Shunsuke Oshima, Yusuke Nishikido, Toshinori Sueyoshi |
CISIS | 4 |
| 2007 | Performance Estimation of TCP under SYN Flood AttacksabstractThe SYN flood attack is a DoS (denial of service) method affecting hosts to retain the half-open state and causing to exhaust its memory resources. This attack is hardly filtered by the router in such a case that the source IP address is spoofed. In this paper, we present the performance estimation of TCP under SYN flood attacks and propose a detective method at an early stage. We implement an attacking program and observe response packets from the server on different OS's. Our performance estimation explores the metric to detect a condition caused by SYN flood attacks. Firstly, the observation of response packets leads to find the most sensitive metric and its threshold. Secondly, the packet loss rate is adopted as the metric to identify whether the server is attacked or not. Finally, we detect the slight variations of response packet if the value exceeds the pre-determined threshold value, then the detective host sends the RST packet to release the half-open state on TCP Takuo Nakashima, Toshinori Sueyoshi |
CISIS | 2 |
| 2007 | A Novel Technique to Create Energy-Efficient Contexts for Reconfigurable LogicabstractHigh power consumption is a constraining factor for the growth of programmable logic devices. We propose two techniques in order to reduce power consumption. The first is a technique for creating contexts. This technique uses data-dependent circuits and wire sharing between contexts. The second is a technique for switching the contexts. In this paper, we evaluate the capability of the two techniques to reduce power consumption using a multi-context logic device. As a result, as compared with the original circuit, our multi-context circuits can reduce the power consumption by 9.1% on an average and by a maximum of 19.0%. Furthermore, applying our resource sharing technique to these circuits, we achieved a reduction of 10.6% on an average and a maximum reduction of 18.8%. Hiroshi Shinohara, Hideaki Monji, Masahiro Iida, Toshinori Sueyoshi |
FCCM | 4 |
| 2007 | A Novel Technique to Create Energy-Efficient Contexts for Reconfigurable LogicabstractHigh power consumption is a constraining factor for the growth of programmable logic devices. We propose two techniques in order to reduce power consumption. The first is a technique for creating contexts. This technique uses data-dependent circuits and wire sharing between contexts. The second is a technique for switching the contexts. In this paper, we evaluate the capability of the two techniques to reduce power consumption using a multi-context logic device. As a result, as compared with the original circuit, our multi-context circuits can reduce the power consumption by 9.1% on an average and by a maximum of 19.0%. Furthermore, applying our resource sharing technique to these circuits, we achieved a reduction of 10.6% on an average and a maximum reduction of 18.8%. Hiroshi Shinohara, Hideaki Monji, Masahiro Iida, Toshinori Sueyoshi |
FCCM | 4 |
| 2007 | A Variable Grain Logic Cell Architecture for Reconfigurable Logic CoresabstractReconfigurable Logic Devices are classified as the fine-grained or coarse-grained type on the basis of their basic logic cell architecture. In general, each architecture has its own merit; therefore, it is difficult to achieve a balance between the operation speed and implementation area in various applications. In this paper, we propose a Variable Grain Logic Cell (VGLC) architecture, which consists of a 4-bit ripple carry adder with configuration memory bits and also develop technology mapping tool. Its key feature is the variable granularity being a trade-off between coarse-grained and fine-grained types required for the implementation arithmetic and random logic, respectively. As a result, critical path delay, and number of configuration memory bits are reduced by 49.7%, and 48.5%, respectively, in the benchmark circuits. Motoki Amagasaki, Ryoichi Yamaguchi, Kazunori Matsuyama, Masahiro Iida, Toshinori Sueyoshi |
FPL | 5 |
| 2007 | An Embedded Reconfigurable Logic Core based on Variable Grain Logic Cell ArchitectureabstractReconfigurable computing is becoming increasingly attractive for many applications. It involves the use of reconfigurable logic devices (RLDs). RLDs are classified as the fine-grained or coarse-grained type on the basis of their basic logic cell architecture. In general, each architecture has its own merit; therefore, it is difficult to achieve a balance between the operation speed and implementation area in various applications. In this paper, we propose a variable grain logic cell (VGLC) architecture, which consists of a 4-bit ripple carry adder with configuration memory bits and also develop technology mapping tool. Its key feature is the variable granularity being a tradeoff between coarse-grained and fine-grained types required for the implementation arithmetic and random logic, respectively. As a result, critical path delay, and number of configuration memory bits are reduced by 49.7%, and 48.5%, respectively, in the benchmark circuits. In addition, when implementing DSP benchmarks on trials, the result is comparable with the highest performance processors today. Yoshiaki Satou, Motoki Amagasaki, Hiroshi Miura, Kazunori Matsuyama, Ryoichi Yamaguchi, Masahiro Iida, Toshinori Sueyoshi |
FPT | 7 |
| 2006 | Effective clustering technique to optimize routability of outer cluster netsabstractWe also study a clustering technique for a cluster-based FPGA to optimize the routability of outer cluster nets. Prior research of FPGA clustering aims to take in inter-cluster connections to utilize the various advantages of the local interconnection in the cluster. Taking in the inter-cluster connections to the cluster can improve the FPGA speed and area, and can lighten the burden of the placement and routing tool. However, many inter-cluster connections still remain in the outside cluster. The condition of these connections give the performance of the circuit big influence. Therefore, we propose the effective clustering technique to optimize routability of outer cluster nets which are not taken in. In order to reduce the used routing resources in FPGA, our technique uses two evaluation functions. One evaluation function can be reduced routing resources in the outside cluster. The second evaluation function can utilize various characteristics of the local routing resources in the inside cluster. Our clustering technique has the unusual ability in the optimization of routing resources concurrently. As a result, our method resulted in 21.4% improvement in terms of the routing area (16.6% on average), and the critical path delay is reduced by 28.0% (11.8% on average) compared to existing clustering method in the benchmark circuits. Masaki Kobata, Masahiro Iida, Toshinori Sueyoshi |
FPGA | 3 |
| 2006 | Evaluation of Variable Grain Logic Cell Architecture for Reconfigurable DeviceabstractReconfigurable logic devices are usually classified on the basis of their basic logic cell architecture as fine-grained or coarse-grained. In general, each architecture is suitable on its own merit; therefore, it is difficult to achieve a balance between the operation speed and area-efficiency in applications. In order to solve this problem, we propose a new logic cell architecture based on a 4-bit ripple carry adder that includes configuration memory bits. This is called the variable grain logic cell architecture, VGLC. It is possible to realize two features by using the VGLC: one is a high device speed of a coarse-grained cell and the other is the versatile logic of a fine-grained cell. This paper demonstrates the transistor-level optimization of our proposed logic cell. Moreover, based on the results of the evaluation, the authors show that the critical path delay can be reduced by a maximum of 37% when using the proposed logic cell architecture is used in a 32-bit multiplier Motoki Amagasaki, Takurou Shimokawa, Kazunori Matsuyama, Ryoichi Yamaguchi, Hideaki Nakayama, Naoto Hamabe, Masahiro Iida, Toshinori Sueyoshi |
VLSI-SoC | 8 |
| 2005 | Applying the Small-World Network to Routing Structure of FPGAsabstractThe degree of integration and the operating frequency of programmable logic have improved dramatically with the development of new process technologies. However, for the deep sub-micron processes, the delay, reliability, cost, and power tend to be determined by interconnections. In conventional programmable logic, reducing the number of switches on a critical path is important because the wiring delay is considerably smaller than the switch delay. However, to achieve a decrease in the critical path delay it is also necessary to consider the wiring delay for the deep sub-micron processes. This paper proposes a novel routing structure using a small-world network structure for the interconnection of programmable logic. This paper demonstrates that the critical path delay can be reduced. Based on the results of an evaluation, the authors show that the critical path delay can be reduced by a maximum of 15% and the amount of routing resources can be reduced by a maximum of 23% when using the small-world network structure. Hisashi Tsukiashi, Masahiro Iida, Toshinori Sueyoshi |
FPL | 3 |
| 2004 | EXPRESS-1: a dynamically reconfigurable platform using embedded processor FPGAabstractThis work presents a dynamically reconfigurable platform called EXPRESS-1, which uses the commercially available embedded processor FPGA. The system makes the most of a fully reconfigurable logic part to explore the area of fine-grained reconfigurable computing. A dynamic reconfiguration mechanism is implemented utilizing a real-time operating system, so device reconfiguration in response to application demand works without suspending other services. EXPRESS-1 features a transparent execution mechanism. Whether a function is executed by the hardware or software, the mechanism frees users from awareness of its execution manner. Furthermore, there is no need to explicitly specify reconfiguration commands into a program because the system determines if reconfiguration is needed based on current conditions. The development of EXPRESS-1 and the runtime reconfiguration mechanism of the fully reconfigurable logic are described. System capabilities are also reported through fundamental evaluations with some practical applications such as JavaVM, encryption processing, and image processing. Hidetomo Shibamura, Masayuki Fukuyama, Daisuke Uchida, Seiji Ikeda, Morihiro Kuga, Toshinori Sueyoshi |
FPT | 6 |
| 2001 | Recursive Diagonal Torus: An Interconnection Network for Massively Parallel ComputersabstractRecursive Diagonal Torus (RDT), a class of interconnection network is proposed for massively parallel computers with up to 2/sup 16/ nodes. By making the best use of a recursively structured diagonal mesh (torus) connection, the RDT has a smaller diameter (e.g., it is 11 for 2/sup 10/ nodes) with a smaller number of links per node (i.e., 8 links per node) than those of the hypercube. A simple routing algorithm, called vector routing, which is near-optimal and easy to implement is also proposed. Although the congestion on upper rank tori sometimes degrades the performance under the random traffic, the RDT provides much better performance than that of a 2D/3D torus in most cases and, under hot spot traffic, the RDT provides much better performance than that of a 2D/3D/4D torus. The RDT router chip which provides a message multicast for maintaining cache consistency is available. Using the 0.5 /spl mu/m BICMOS SOG technology, versatile functions, including hierarchical multicasting, combining acknowledge packets, shooting down/restart mechanism, and time-out/setup mechanisms, work at a 60 MHz clock rate. Yulu Yang, Akira Funahashi, Akiya Jouraku, Hiroaki Nishi, Hideharu Amano, Toshinori Sueyoshi |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 1989 | The Kyushu University reconfigurable parallel processor: design of memory and intercommunicaiton architecturesabstractThe reconfigurable parallel processor system under development at Kyushu University is an MIMD-type multiprocessor which consists of N processing-elements (currently N is 128) fully connected by S N × N crossbar networks (currently S is 1). Each PE (Processing Element) employs a Fujitsu SPARC MB86900/10 chip-set, a Weitek WTL1164/65 chip-set, an MMU (Memory Management Unit) with 64K bytes of cache, 4M bytes of memory, and an MCU (Message Communication Unit). The modular 128 × 128 crossbar network is implemented by arranging 256 identical 8 × 8 crossbar LSI-modules in a 16 × 16 matrix form. The full 128-PE configuration achieves supercomputer levels of performance by providing 1.28 GIPS and 205 MFLOPS of computing power, 512M bytes of memory, and 2.56G bytes/s of inter-PE communication bandwidth. At the same time, it exploits unique reconfigurability in the memory and intercommunication architectures. By utilizing these two types of reconfigurability, we believe that the system can be effectively tailored to a wide spectrum of applications such as numerical computation, image processing, computer graphics, artificial intelligence, neurocomputing, and so on. Kazuaki J. Murakami, Shin-ichiro Mori, Akira Fukuda, Toshinori Sueyoshi, Shinji Tomita |
ICS | 4 |