EDBT 2026 Demo / reviewers in the wild / expert
Hideharu Amano
dblp:59/1558
· DBLP profile ↗
215ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0002-9371-2060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 200 · 15 first-author · 13 since 2021Software engineering, systems software and programming languages · 8 · 2 first-authorArtificial intelligence and machine learning · 3Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 3-day ASIC Design Hands-on with the Minimal Fab
Hideharu Amano, Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Tohru Mogami, Yoshio Mita |
ISCAS | 1 |
| 2025 | An SoC Design and Fabrication Hands-On Educational Course within One Week Using Structured ASICabstractSince the design and manufacturing of semiconductor chips takes a considerable amount of time, it is difficult to fully understand the entire process and conduct student experiments that provide a hands-on experience of chip creation. The lack of student experiments that allow beginners to easily experience the process from chip design to manufacturing is one of the reasons why there are few students interested in semiconductor research. This, in turn, exacerbates the shortage of skilled workers in the semiconductor industry. The Agile-Chip platform is a method that enables the rapid and low-cost production of small quantities of chips by manufacturing only the top wiring layer using minimal fab technology on cut wafers, which are pre-cut from mass-produced wafers except for the top-most wiring. In this paper, we propose a student experiment method for semiconductor beginners using the Agile-Chip platform and provide an implementation example. For the base chip, a gate array is used, which allows the configuration of various gates using only the top wiring layer. The students design a 61-stage ring oscillator using the gate array, verify its operation through simulation, and then proceed with the corresponding layout design. After confirming the consistency of both, they generate a GDS file. This wiring layer is then fabricated either in a minimal fab or a cleanroom, and finally, the chip is mounted on a substrate for measurement. Hideharu Amano, Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Tohru Mogami, Yoshio Mita |
ISCAS | 1 |
| 2025 | Qu-Trefoil: Large-Scale Quantum Circuit Simulator Working on FPGA With SATA StoragesabstractQuantum circuits are fundamental components of quantum computing, and state-vector-based quantum circuit simulation is a widely used technique for tracking qubit behavior throughout circuit evolution. However, simulating a circuit with$n$qubits requires$2^{n+4}$bytes of memory, making simulations of more than 40 qubits feasible only on supercomputers. To address this limitation, we propose the Qu-Trefoil, a system designed for large-scale quantum circuit simulations on an FPGA-based platform called Trefoil. Trefoil is a multi-FPGA system connected to eight storage subsystems, each equipped with 32 SATA disks. Qu-Trefoil integrates a suite of HLS-based universal quantum gates, including Clifford gates (Hadamard (H), Pauli-Z (Z), Phase (S), Controlled-NOT (CNOT)), the T gate, and unitary matrix computation, along with HDL-designed modules for system-wide integration. Our extensive evaluation demonstrates the system's robustness and flexibility, covering quantum gate performance, chunk size, disk extensibility, and efficiency across different SATA generations. We successfully simulated quantum circuits with over 43 qubits, which required more than 128 TB of memory, in approximately 3.72 to 13.06 hours on a single storage subsystem equipped with one FPGA. This achievement represents a significant milestone in the advancement of quantum computing simulations. Furthermore, thanks to its unique architecture, Qu-Trefoil is more accessible, flexible, and cost-efficient than other existing simulators for large-scale quantum circuit simulations, making it a viable option for researchers with limited access to supercomputers. Kaijie Wei, Hideharu Amano, Ryohei Niwase, Yoshiki Yamaguchi, Takefumi Miyoshi |
IEEE Trans. Computers | 2 |
| 2025 | Agile-X: A Structured-ASIC Created With a Mask-Less Lithography System Enabling Low-Cost and Agile Chip FabricationabstractScaling to finer CMOS process nodes necessitates more masks, resulting in higher costs and extended turnaround times (TATs). High costs and long TATs have hindered researchers outside the field of integrated circuits, including those in medicine, physics, and science from prototyping their own chips. Therefore, opportunities for diverse innovations in integrated circuits and talent development have been limited. We have developed the Agile-X platform for low-cost, rapid manufacturing of system-on-chips. Users can implement their own dedicated circuits with gate-array circuits on a base chip, which has common intellectual properties (IPs) such as RISC-V CPUs, various IOs, and ADCs. The base chip is manufactured in a foundry up to the intermediate metal layers and shipped with metal deposition on its surface. By directly drawing wiring patterns on this base chip with a mask-less lithography system, custom chips can be manufactured on-site without masks. As this process only requires wiring and eliminates masks, production time is drastically reduced compared to traditional full-mask wafer processes and multiproject wafer (MPW) shuttles. Development and manufacturing costs for the base chip, including preintegrated IPs, are shared among all Agile-X users. This reduces both IP and base-chip wafer costs per user. We prototyped wafers using a 0.18-$\mu $m CMOS process and tested the proposed structured ASIC platform and manufacturing process using mask-less lithography systems. The results indicate that the process from inputting GDS data to lithography and dry etching can be completed within 30 min, and custom application-specific integrated circuits (ASICs) can be manufactured within a day. Compared with full-mask wafer design and manufacturing, the manufacturing cost per chip, including IP costs, is reduced from 271000 USD to 22 USD, a reduction of 1/12252, and the manufacturing period is reduced from 20 days to 30 min, a reduction of 1/960. Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Hideharu Amano, Tohru Mogami, Yoshio Mita, Tadahiro Kuroda |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Optimized Two-Step Store Control for MTJ-Based Nonvolatile Flip-Flops to Minimize Store Energy Under Process and Temperature VariationsabstractIntroducing a magnetic tunneling junction (MTJ) into a flip-flop enables nonvolatile power gating (PG) but large store energy to MTJ is a critical concern. We propose an optimized two-step store (TSS) control to first perform a short store with an optimal time for all nonvolatile flip-flops (NVFFs) and then perform a long store only at the failed ones for reducing the store energy. As the key technologies to realize this, we present a verify-and-retryable NVFF (VR-NVFF) circuit enabling the TSS control and an analytical expression for the optimal short-store time (${T}_{\text {short}{\_}\text{opt}}$) minimizing the store energy. To examine the effectiveness of the optimized TSS control and the validity of analytically derived${T}_{\text {short}{\_}\text{opt}}$, we implemented the TSS control on a coarse-grained reconfigurable array (CGRA)-based accelerator chip and fabricated it in a 40-nm CMOS/MTJ hybrid process technology. Results demonstrated that analytical${T}_{\text {short}{\_}\text{opt}}$showed a good agreement with the measured value (within 8% difference) under process and temperature variations. The TSS control with${T}_{\text {short}{\_}\text{opt}}$reduced the store energy to$0.32\times $of that of the conventional long-store-only technique. The break-even time (BET), which is the minimum power-gating time to get the gain in energy savings, was shortened to 0.51–$0.7\times $by the TSS control, achieving the BET of 50–$923 \mu \text{s}$in the range of 0 °C–80 °C. Kimiyoshi Usami, Daiki Yokoyama, Aika Kamei, Hideharu Amano, Kenta Suzuki, Keizo Hiraga, Kazuhiro Bessho |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Designing low-diameter interconnection networks with multi-ported host-switch graphsabstractSummary A host‐switch graph was originally proposed as a graph that represents a network topology of a computer systems with 1‐port host computers and ‐port switches. It has been studied from both theoretical and practical aspects in terms of the diameter, the average shortest path length, and the performance of real applications. In recent high‐performance computing systems, however, a host computer is connected to multiple switches by using InfiniBand, NVSwitch, or Omni‐Path, and consequently they provide high bandwidths. Since a host‐switch graph cannot represent such systems, this article extends a host‐switch graph so that it can represent such systems. As a result, a host‐switch graph can include multi‐ported hosts. Furthermore, we propose to use multi‐port hosts for reducing the diameter. We show that the diameter minimization is equivalent to solving the degree diameter problem for bipartite graphs of diameter three. Our experimental results show that we can drastically reduce the diameter as well as increasing the bandwidth and improves performance of MPI applications by up to 162% as compared with networks with single‐ported hosts. Ryota Yasudo, Koji Nakano, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | A Variation-Aware MTJ Store Energy Estimation Model for Edge Devices With Verify-and-Retryable Nonvolatile Flip-FlopsabstractWhile the spin-transfer torque (STT) magnetic tunnel junction (MTJ) is a promising technique for enabling nonvolatile flip-flops (NVFFs) to perform power gating to reduce leakage power without any data losses, the large store energy (the energy to make a store operation) of MTJs needs to be addressed. The nonvolatile cool mega array series is an edge-oriented coarse-grained reconfigurable accelerator that implements an improved MTJ-based NVFF with a verify-and-retryable store method that should ideally reduce the store energy under the presence of the switching time variation originating from the stochastic nature of the MTJs. However, the energy reduction effect of the method has not been formulated or evaluated thoroughly enough to make the best use of the method in actual applications. In this study, we propose an analytical model to estimate the store energy in typical operational conditions under the assumption of switching time variations following the normal distributions based on the measurements of a real chip fabricated with a 40-nm perpendicular MTJ/CMOS hybrid process. In contrast to the tedious measurement on each different condition, the proposed model allows for an instantaneous determination of the best storing method for minimizing the store energy, with an energy reduction of up to 69% compared with a conventional one-time attempt storing method. This model is expected to be used for system-level energy simulations and, ultimately, for design explorations in pursuit of energy-optimized memory. Aika Kamei, Hideharu Amano, Takuya Kojima, Daiki Yokoyama, Kimiyoshi Usami, Keizo Hiraga, Kenta Suzuki, Kazuhiro Bessho |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Optimizing Application Mapping for Multi-FPGA Systems with Multi-ejection STDM SwitchesabstractMulti-FPGA systems have received an attention as a computing cluster for multi-access edge computing (MEC). Also, they can process time-critical jobs with their hardwired logic. For this purpose, the static time-division multiplexing (STDM) network is adopted because it enables to predict latency and bandwidth. However, the overall performance of the STDM network depends on the number of time slots. This paper proposes a new mapping tool that optimizes the application mapping so that the number of slots is minimized. Our tool handles multicasts and multi-ejection function which are effective techniques for STDM switches implemented on an FPGA cluster. For applications with all-to-all communication, our experimental results show that the tool reduces the number of time slots by 59–68% with both multicasts and multi-ejection switches. Kohei Ito, Ryota Yasudo, Hideharu Amano |
FPL | 3 |
| 2022 | FPL Demo: An FPGA-IP Prototype Chip for MEC devicesabstractThis demonstration shows a prototype chip of SLM (Scalable Logic Module) for a novel FPGA-IP embedded in various chips for edge computing. In this paper, the authors briefly describe the architecture of our FPGA-IP and the evaluation environment for the prototype chip. Morihiro Kuga, Masahiro Iida, Hideharu Amano |
FPL | 3 |
| 2022 | A Hardware Trojan Exploiting Coherence Protocol on NoCs
Yoshiya Shikama, Michihiro Koibuchi, Hideharu Amano |
PDCAT | 3 |
| 2022 | An efficient compilation of coarse-grained reconfigurable architectures utilizing pre-optimized sub-graph mappingsabstractIn recent years, IoT devices have become widespread, and energy-efficient coarse-grained reconfigurable architectures (CGRAs) have attracted attention. CGRAs comprise several processing units called processing elements (PEs) arranged in a two-dimensional array. The operations of PEs and the interconnections between them are adaptively changed depending on a target application, and this contributes to a higher energy efficiency compared to general-purpose processors. The application kernel executed on CGRAs is represented as a data flow graph (DFG), and CGRA compilers are responsible for mapping the DFG onto the PE array. Thus, mapping algorithms significantly influence the performance and power efficiency of CGRAs as well as the compile time. This paper proposes POCOCO, a compiler framework for CGRAs that can use pre-optimized subgraph mappings. This contributes to reducing the compiler optimization task. To leverage the subgraph mappings, we extend an existing mapping method based on a genetic algorithm. Experiments on three architectures demonstrated that the proposed method reduces the optimization time by 48%, on an average, for the best case of the three architectures. Ayaka Ohwada, Takuya Kojima, Hideharu Amano |
PDP | 3 |
| 2022 | Mapping-Aware Kernel Partitioning Method for CGRAs Assisted by Deep LearningabstractCoarse-grained reconfigurable architectures (CGRAs) provide high energy efficiency with word-level programmability rather than bit-level ones such as FPGAs. The coarser reconfigurability brings about higher energy efficiency and reduces the complexity of compiler tasks compared to the FPGAs. However, application mapping process for CGRAs is still time-consuming. When the compiler tries to map a large and complicated application data-flow-graph(DFG) onto the reconfigurable fabric, it tends to result in inefficient resource use or to fail in mapping. In case of failure, the compiler must divide it into several sub-DFGs and goes back to the same flow. In this work, we propose a novel partitioning method based on a genetic algorithm to eliminate the unmappable DFGs and improve the mapping quality. In order not to generate unmappable sub-DFGs, we also propose an estimation model which predicts the mappability and resource requirements using a DGCNN (Deep Graph Convolutional Neural Network). The genetic algorithm with this model can seek the most resource-efficient mapping without the back-end mapping process. Our model can predict the mappability with more than 98% accuracy and resource usage with a negligible error for two studied CGRAs. Besides, the proposed partitioning method demonstrates 53-75% of memory saving, 1.28-1.39x higher throughput, and better mapping quality over three comparative approaches. Takuya Kojima, Ayaka Ohwada, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | TCI Tester: Tester for Through Chip InterfaceabstractNo abstract available. Hideto Kayashima, Hideharu Amano |
ASP-DAC | 2 |
| 2021 | Resource-saving FPGA Implementation of the Satisfiability Problem Solver: AmoebaSATslimabstractThe Boolean satisfiability problem (SAT) is an NP-complete combinatorial optimization problem, where fast SAT solvers are useful for various smart society applications. Since these edge-oriented applications require time-critical control, a high speed SAT solver on FPGA is a promising approach. Here the authors propose a novel FPGA implementation of a bio-inspired stochastic local search algorithm called ‘AmoebaSAT’ on a Zynq board. Previous studies on FPGA-AmoebaSATs tackled relatively smaller-sized 3-SAT instances with a few hundred variables and found the solutions in several milli seconds. These implementations, however, adopted an instance-specific approach, which requires synthesis of FPGA configuration every time when the targeted instance is altered. In this paper, a slimmed version of AmoebaSAT named ‘AmoebaSATslim,’ which omits the most resource-consuming part of interactions among variables, is proposed. The FPGA-AmoebaSATslim enables to tackle significantly larger-sized 3-SAT instances, accepting 30,000 variables with 130, 800 clauses. It achieves up to approximately 24 times faster execution speed than the software-AmoebaSATslim implemented on a CPU of the x86 server. Ying Jie Yan, Hideharu Amano, Masashi Aono, Kaori Ohkoda, Shingo Fukuda, Kenta Saito, Seiya Kasai |
FPT | 2 |
| 2021 | A Case for Low-Latency Network-on-Chip using Compression RoutersabstractThe communication latency is a primary concern for designing Network-on-Chips (NoCs) since it significantly affects the parallel application performance on a many-core computer system. To reduce the communication latency, we propose an on-chip router that (de)compresses the contents of an incoming packet before completing switch arbitration. The compression router thus has no latency penalty for the compression operation, whereas it shortens a packet length that decreases the network injection-and-ejection latency. Evaluation results show that the compression router improves 7.7% of the parallel application performance (IS, CG, FT, and TSP) and 49% of the effective network throughput by 1.8 compression ratio on NoC. The drawback is that the router area and its energy consumption per bit increase by 0.12mm2and 1.4 times compared to the conventional virtual-channel router. Naoya Niwa, Yoshiya Shikama, Hideharu Amano, Michihiro Koibuchi |
PDP | 3 |
| 2021 | Low-Latency Low-Energy Memory-Cube Networks using Dual-Voltage DatapathsabstractThree-dimensional stack memory that provides both high-bandwidth access and large capacity is a promising technology for next-generation computer systems. While a large number of memory cubes increase the aggregate memory capacity, the communication latency and power consumption would become significant due to its low-radix large-diameter packet network. In this context, we propose a memory-cube network called Diagonal Memory Network (DMN). A diagonal network topology, its floor layout, and its lightweight router are designed for low-latency and low-voltage memory-read communication. Our evaluation results show that a DMN router decreases 31% of the hardware resources than a conventional virtual-channel router. The DMN router reduces 13% and 67% energy consumption to transit a packet along with the original datapath and bypassing datapath, respectively. Yoshiya Shikama, Ryuta Kawano, Hiroki Matsutani, Hideharu Amano, Yusuke Nagasaka, Naoto Fukumoto, Michihiro Koibuchi |
PDP | 4 |
| 2021 | Analytical Performance Estimation for Large-Scale Reconfigurable Dataflow PlatformsabstractNext-generation high-performance computing platforms will handle extreme data- and compute-intensive problems that are intractable with today’s technology. A promising path in achieving the next leap in high-performance computing is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators such as FPGAs, which have been shown to speed up compute-intensive tasks with reduced power consumption. However, assessing the feasibility of large-scale heterogeneous systems requires fast and accurate performance prediction. This article proposes Performance Estimation for Reconfigurable Kernels and Systems (PERKS), a novel performance estimation framework for reconfigurable dataflow platforms. PERKS makes use of an analytical model with machine and application parameters for predicting the performance of multi-accelerator systems and detecting their bottlenecks. Model calibration is automatic, making the model flexible and usable for different machine configurations and applications, including hypothetical ones. Our experimental results show that PERKS can predict the performance of current workloads on reconfigurable dataflow platforms with an accuracy above 91%. The results also illustrate how the modelling scales to large workloads, and how performance impact of architectural features can be estimated in seconds. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker, Ce Guo 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2020 | Accelerating Deep Learning using Multiple GPUs and FPGA-Based 10GbE SwitchabstractA back-propagation algorithm following a gradient descent approach is used for training deep neural networks. Since it iteratively performs a large number of matrix operations to compute the gradients, GPUs (Graphics Processing Units) are efficient especially for the training phase. Thus, a cluster of computers each of which equips multiple GPUs can significantly accelerate the training phase. Although the gradient computation is still a major bottleneck of the training, gradient aggregation and parameter optimization impose both communication and computation overheads, which should also be reduced for further shortening the training time. To address this issue, in this paper, multiple GPUs are interconnected with a PCI Express (PCIe) over 10 Gbit Ethernet (10GbE) technology. Since these remote GPUs are interconnected via network switches, gradient aggregation and optimizers (e.g., SGD, Adagrad, Adam, and SMORMS3) are offloaded to an FPGA-based network switch between a host machine and remote GPUs; thus, the gradient aggregation and optimization are completed in the network. Evaluation results using four remote GPUs connected via the FPGA-based 10GbE switch that implements the four optimizers demonstrate that these optimization algorithms are accelerated by up to 3. 0x and 1. 25x compared to CPU and GPU implementations, respectively. Also, the gradient aggregation throughput by the FPGA-based switch achieves 98.3% of the 10GbE line rate. Tomoya Itsubo, Michihiro Koibuchi, Hideharu Amano, Hiroki Matsutani |
PDP | 3 |
| 2020 | Extracting Success from IBM's 20-Qubit Machines Using Error-Aware CompilationabstractNISQ (Noisy, Intermediate-Scale Quantum) computing requires error mitigation to achieve meaningful computation. Our compilation tool development focuses on the fact that the error rates of individual qubits are not equal, with a goal of maximizing the success probability of real-world subroutines such as an adder circuit. We begin by establishing a metric for choosing among possible paths and circuit alternatives for executing gates between variables placed far apart within the processor, and test our approach on two IBM 20-qubit systems named Tokyo and Poughkeepsie. We find that a single-number metric describing the fidelity of individual gates is a useful but imperfect guide. Our compiler uses this subsystem and maps complete circuits onto the machine using a beam search-based heuristic that will scale as processor and program sizes grow. To evaluate the whole compilation process, we compiled and executed adder circuits, then calculated the Kullback–Leibler divergence (KL-divergence, a measure of the distance between two probability distributions). For a circuit within the capabilities of the hardware, our compilation increases estimated success probability and reduces KL-divergence relative to an error-oblivious placement. Shin Nishio, Yulu Pan, Takahiko Satoh, Hideharu Amano, Rodney Van Meter |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | GenMap: A Genetic Algorithmic Approach for Optimizing Spatial Mapping of Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) are expected to be used for embedded systems, Internet of Things (IoT) devices, and edge computing thanks to their high-energy efficiency and programmability. In essence, a CGRA is an array of numerous processing elements. To exploit this abundant computation resource, a compiler for CGRAs has to fulfill more tasks compared that for general-purpose processors. Therefore, many studies have proposed optimization methods, especially for application mapping, because the performance and energy efficiency strongly depend on optimization at compile time. However, many works focus only on performance improvement or resource minimization, although such optimization objectives are not always appropriate when considering various use cases. In this work, we propose GenMap, an application mapping framework using multiobjective optimization based on a genetic algorithm so that users can set optimization criteria as needed. Besides, it provides aggressive power optimization using our dynamic power model and leakage minimization technique. The proposed method is applied to three fabricated CGRA chips for evaluation. Experimental results show that GenMap achieves 15.7% reduction of wire length while keeping processing element utilization when compared with conventional methods. In addition, according to real chip experiments, 12.1%-46.8% of energy consumption is reduced, and up to $2\times $ speedup is archived for several architectures when compared with other two approaches. Takuya Kojima, Nguyen Anh Vu Doan, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Sparse 3-D NoCs with Inductive CouplingabstractWireless interconnects based on inductive coupling technology are compelling propositions for designing 3-D integrated chips. This work addresses the heat dissipation problem on such systems. Although effective cooling technologies have been proposed for systems designed based on Through Silicon Via (TSV), their application to systems that use inductive coupling is problematic because of increased wireless-communication distance. For this reason, we propose two methods for designing sparse 3-D chips layouts and Networks on Chip (NoCs) based on inductive coupling. The first method computes an optimized 3-D chip layout and then generates a randomized network topology for this layout. The second method uses a standard stack chip layout with a standard network topology as a starting point, and then deterministically transforms it into either a "staircase" or a "checkerboard" layout. We quantitatively compare the designs produced by these two methods in terms of network and application performance. Our main finding is that the first method produces designs that ultimately lead to higher parallel application performance, as demonstrated for nine OpenMP applications in the NAS Parallel Benchmarks. Michihiro Koibuchi, Lambert T. Leong, Tomohiro Totoki, Naoya Niwa, Hiroki Matsutani, Hideharu Amano, Henri Casanova |
DAC | 6 |
| 2019 | Demonstration of Flow-in-Cloud: A Multi-FPGA SystemabstractFlow-in-Cloud(FiC) is an acceleration platform designed to make a virtual monolithic large FPGA image from a number of mid-range economical FPGAs. We will show the live demonstration of the acceleration example of FiC with 24 boards through the network. Kazuei Hironaka, Kensuke Iizuka, Akram Ben Ahmed, M. M. Imdad Ullah, Yugo Yamauchi, Yuxi Sun 0001, Miho Yamakura, Aoi Hiruma, Hideharu Amano |
FPL | 9 |
| 2019 | Demonstration of Low Power Stream Processing Using a Variable Pipelined CGRAabstractVPCMA (Variable Pipelined Cool Mega Array) is a low power CGRA (Coarse-Grained Reconfigurable Architecture) which we previously proposed in [1]. CC-SOTB2 is a real chip implementation of the VPCMA using Renesas 65-nm SOTB technology [2]. In this demonstration, we will show the power consumption of the CC-SOTB2 while performing a real image processing. Takuya Kojima, Naoki Ando, Yusuke Matsushita 0001, Hideharu Amano |
FPL | 4 |
| 2019 | Multi-FPGA Management on Flow-in-Cloud Prototype SystemabstractFiC (Flow-in-Cloud) is a multi-FPGA system to achieve monolithic large scale FPGA with across multiple FPGAs for computational-heavy applications. The FiC system employs multiple economical mid-range FPGAs and connect them directly with flexible and high bandwidth interconnection network. We developed FiCSW board as a 1st generation of FPGA board in the FiC project. The FiCSW design, since the FPGA is used as both computational resource and the network circuit switch, there are no host server for FPGA needed like other server-based multi FPGA architecture. This design strategy helps scalable FPGA fabric in low-cost. However, from the system manageability's perspective, the strategy is hard to handle multiple FPGA nodes without management system. In this paper, we mainly focused on the management system on the FiCSW multi-FPGA system, introducing the system architecture and implementation, and working application on the FiCSW prototype system. Kazuei Hironaka, Akram Ben Ahmed, Hideharu Amano |
SNPD | 3 |
| 2019 | Designing High-Performance Interconnection Networks with Host-Switch GraphsabstractThis paper aims at establishing a method for designing high-performance network topologies to bridge a gap between theoretical and practical studies. To this end, we present a novel graph called a host-switch graph, which consists of host vertices and switch vertices with maximum degree 1 and$r$, respectively. This graph represents a network topology of a practical parallel/distributed computer system with host computers connected by$r$-port switches. We discuss important metrics for designing high-performance interconnection networks: the host-to-host average shortest path length (h-ASPL) and the bisection width (BiW). In particular, we explore a method for constructing host-switch graphs with low h-ASPL and high BiW that connect the fixed number of hosts via any number of$r$-port switches. We demonstrate that the number of switches that provides the minimum h-ASPL can mathematically be approximated, and the minimum number of switches that provides a certain BiW can experimentally be approximated. On the basis of the approximations, we propose a randomized algorithm for searching host-switch graphs. We then apply the graphs to interconnection networks and compare them with typical network topologies. As compared with the torus, the dragonfly, and the fat-tree, our networks attain higher performance and smaller power and costs. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2018 | Superpixel Accelerator for Computer Vision Applications on Arria 10 SoCabstractSuperpixel segmentation is a very popular image segmentation technique used in various computer vision tasks. Recently, a number of superpixel algorithms have been proposed in literature. One such algorithm is considered as the-state-of-the-art in superpixel segmentation: Simple Linear Iterative Clustering or SLIC. However, its original implementation has a long execution time on high performance processors designed within the common mobile and enterprise applications, as well on high-end processors such as Intel Xeon. Overall, the execution time for single-threaded implementation is considered critical for real-time or near real-time applications. In this paper, we explore the possibility of accelerating parts of the SLIC image segmentation critical for performance, by designing the image segmentation accelerator for Intel's Arria 10 SoC. We propose a novel architecture to enable hardware acceleration by addressing the problem of hardware/software partitioning to minimize the overall program latency. Amila Akagic, Emir Buza, Razija Turcinhodzic, Hana Haseljic, Hiroyuki Noda, Hideharu Amano |
DDECS | 6 |
| 2018 | Performance Prediction for Large-Scale Heterogeneous PlatformsabstractThis paper presents an approach for analysing, modelling and predicting application performance of large-scale heterogeneous platforms. Our approach combines analytical and statistical modelling techniques, and aims to: (1) identify and characterise code regions that are the most promising candidates to benefit from acceleration; (2) provide statistical models that predict application behaviour for unobserved inputs; and (3) predict performance gain with different system architectures. Ryota Yasudo, Ana Lucia Varbanescu, José Gabriel F. Coutinho, Wayne Luk, Hideharu Amano |
FCCM | 5 |
| 2018 | A Configuration Data Multicasting Method for Coarse-Grained Reconfigurable ArchitecturesabstractThis paper proposes a novel configuration data compression technique for coarse-grained reconfigurable architectures (CGRAs). The proposed technique is based on a multicast configuration technique called RoMultiC, which reduces the configuration time by multicasting the same data to multiple PEs(Processing Elements) with two bit-maps. Scheduling algorithms for an optimizing the order of multicasting have been proposed. In general, configuration data for CGRAs can be divided into some fields like machine code formats. The proposed scheme confines a part of fields for multicasting so that the possibility of multicasting more PEs can be increased. This paper analyzes algorithms to find a configuration pattern which maximizes the number of multicasted PEs. We implemented the proposed scheme to CMA (Cool Mega Array), a straight forward CGRA as a case study. Experimental results show that the proposed method achieves 40.0% smaller configuration for an image processing application at maximum. Furthermore, it achieves 35.6% reduction of the power consumption for the configuration with a negligible area overhead. Takuya Kojima, Hideharu Amano |
FPL | 2 |
| 2018 | Accelerator-in-Switch: A Novel Cooperation Framework for FPGAs and GPUsabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. A large-scale FPGAs have been used as high-performance switches which connect powerful computational accelerators including GPUs. In most cases, the FPGA has a room of implementing additional accelerators which can treat on-the-fly data in the switch directly. Here, we introduce two implementation examples: one is for a low latency switching hub PEACH3 for high-performance scientific computing, and the other is FiC-SW for AI computing in a cloud. The key technique is how to use the partial reconfiguration and HLS description for the accelerator-in-switch. Hideharu Amano |
FPT | 1 |
| 2018 | FPGA Design for Autonomous Vehicle Driving Using Binarized Neural NetworksabstractWe propose an autonomous vehicle controlled by FPGAs. In our design, considering embedded systems, we apply the binarized neural networks (BNNs) which can realize a satisfying result in high speed and accuracy to recognize pedestrians and some obstacles on a given road. To detect the traffic light, a passive camera-based pipeline is applied. Furthermore, the implementation of road lane detection is based on color selection algorithm, Canny Edge Detection, and Hough Transformation. The proposed design is realized by two Xilinx boards: PYNQ-Z1 and Zynq-Xc7Z010. These two FPGA boards cooperate with each other through a shared network cable. In the proposed design, the resource used by Zynq-Xc7Z010 can be greatly reduced and the inference time on the FPGA has been thousands times faster than the software implementation. Kaijie Wei, Koki Honda, Hideharu Amano |
FPT | 3 |
| 2018 | Performance Estimation for Exascale Reconfigurable Dataflow PlatformsabstractThe next generation high-performance computing platforms will need to support exascale computing. A promising path in achieving exascale is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators. However, assessing the feasibility of heterogeneous exascale systems requires fast and accurate performance prediction. This paper proposes PERKS, a novel performance estimation frame-work for reconfigurable dataflow platforms (RDPs). PERKS uses machine and application parameters to build an analytical model for predicting the performance of multi-accelerator systems. Moreover, model calibration is automatic, making the model flexible and usable for different machine configurations and applications. Our experimental results demonstrate that PERKS can predict the performance of current workloads and RDPs with an accuracy above 95%. We also demonstrate how the modelling scales to exascale workloads and exascale platforms. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker |
FPT | 5 |
| 2018 | AxNoC: Low-power Approximate Network-on-Chips using Critical-Path IsolationabstractVarious parallel applications, such as numerical convergent computation and multimedia processing, have intrinsic tolerance to inaccuracies that allow soft errors, i.e. bit flips, on a chip. However, existing Network-on-Chips (NoCs) guarantee error-free data transfer; thus, encountering limits to reduce the power consumption. In this context, we propose an approximate dual-voltage NoC, called AxNoC. An AxNoC router uses a per-flit look-ahead power management so that headers and important-data flits are perfectly transferred at a high voltage while the remaining flits may incur bit flips by decreasing the supply voltage. An AxNoC router isolates the critical path when the supply voltage is low since such a critical path is enabled only at high voltage. The critical path isolation enables low-voltage operation to work at the same operating frequency at high voltage. An AxNoC router was implemented using a 28nm process and the evaluation results illustrate its efficiency to reduce the power consumption reaching up to 43% while incurring a small area overhead that does not exceed 6.2%. We also demonstrate that AxNoC exhibits an acceptable accuracy illustrated in a sufficiently small geomean of error. Akram Ben Ahmed, Daichi Fujiki, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
NOCS | 5 |
| 2018 | Asymmetric Body Bias Control With Low-Power FD-SOI Technologies: Modeling and Power Optimization
Hayate Okuhara, Akram Ben Ahmed, Johannes Maximilian Kühn, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | A Case for Uni-directional Network Topologies in Large-Scale ClustersabstractDesigning low-latency network topologies of switches is a key objective for next-generation large-scale clusters. Low latency is preconditioned on low hop counts, but existing network topologies have hop counts much larger than theoretical lower bounds. To alleviate this problem, we propose building network topologies based on uni-directional graphs that are known to have hop counts close to theoretical lower bounds. A practical difficulty with uni-directional topologies is switch-by-switch flow control, which we resolve by using hot-potato routing. Cycle-accurate network simulation experiments for various traffic patterns on uni-directional topologies show that hot-potato routing achieves performance comparable to that of conventional deadlock-free routing. Similar experiments are used to compare several uni-directional topologies to bi-directional topologies, showing that the former achieve significantly lower latency and higher throughput. We quantify end-to-end application performance for parallel application benchmarks via discrete-even simulation, showing that uni-directional topologies can lead to large application performance improvements over their bi-directional counterparts. Finally, we discuss practical issues for uni-directional topologies such as cabling complexity and cost, power consumption, and soft-error tolerance. Our results make a compelling case for considering uni-directional topologies for upcoming large-scale clusters. Michihiro Koibuchi, Tomohiro Totoki, Hiroki Matsutani, Hideharu Amano, Fabien Chaix, Ikki Fujiwara, Henri Casanova |
CLUSTER | 4 |
| 2017 | Body bias optimization for variable pipelined CGRAabstractVariable Pipeline Cool Mega Array (VPCMA) is an low power Coarse Grained Reconfigurable Architecture (CGRA) based on the concept of CMA (Cool Mega Array). It implements a pipeline structure that can be configured depending on performance requirements, and the silicon on thin buried oxide (SOTB) technology that allows to control its body bias voltage to balance performance and leakage power. In this paper, we propose a methodology to optimize exactly with an Integer Linear Program the VPCMA body bias while considering simultaneously its variable pipeline structure. For the studied applications, we evaluate that it is possible to achieve an average reduction of energy consumption of 19.3% and 11.8% when compared to respectively the zero bias (without body bias control) and the uniform (control of the whole PE array) cases, while respecting performance constraints. Besides, with appropriate body bias control, it is possible to extend the possible performance, hence enabling broader trade-off analyzes between consumption and performance. These promising results show that applying an adequate optimization technique for the body bias control while simultaneously considering pipeline structures can not only enable further power reduction than previous methods, but also allow more trade-off analysis possibilities. Takuya Kojima, Naoki Ando, Hayate Okuhara, Nguyen Anh Vu Doan, Hideharu Amano |
FPL | 5 |
| 2017 | In-switch approximate processing: Delayed tasks management for MapReduce applicationsabstractIn MapReduce, the parallel processing performance is often limited by only a few compute nodes that delay to complete given tasks. Although various techniques have been invented to handle such stragglers, these techniques mostly impose a burden on master node to monitor the progress of all the compute nodes, resulting in a new bottleneck as the number of compute nodes increases. As an alternative approach, in this paper, we propose to move such straggler management burden from master node to network switch that connects the master and compute nodes, because all the information goes through the switch. More specifically, the proposed network switch monitors output packets from Map tasks to detect stragglers. When detected, the proposed switch generates a response instead of the straggler based on the outputs of the other normal Map tasks, so that Reduce tasks can be started without delay. We introduce some approximate techniques for the proxy computation and response at the switch; thus our switch is called "ApproxSW." We implement ApproxSW on NetFPGA-SUME board that has four 10Gbit Ethernet (10GbE) interfaces and a Virtex-7 FPGA. An experiment shows that the ApproxSW functions do not degrade the original 10GbE switch performance. We also analyze the accuracy of the proxy computation and response for stragglers and show that the proposed approximation based on task similarity achieves the best accuracy. Koya Mitsuzuka, Ami Hayashi, Michihiro Koibuchi, Hideharu Amano, Hiroki Matsutani |
FPL | 4 |
| 2017 | Accelerator-in-switch: A framework for tightly coupled switching hub and an accelerator with FPGAabstractAccelerator-in-Switch (AiS) is a framework for building an accelerator logic tightly coupled with a switching hub in a single FPGA for high performance computation with heterogeneous environment with CPUs and GPUs. AiS is implemented on a partial reconfigurable region of an FPGA whose permanent region is used for a switching hub. A port of the switching hub is connected to the registers and local memory of AiS directly. AiS has a standard interface for a standard bus (Avalon MM bus, here) to exchange data between on-board DDR SDRAM and the local memory, and various types of accelerators can be implemented just by providing such an interface. The data input and output are performed with the DMA controller inside the switching hub with the shared memory model between host CPUs and GPUs. We implemented two example accelerators: a reduction calculator for a radiation transfer equations solver (RED) and LET generator for N-body simulation (LET) were implemented as the AiS on PEACH3, a switching hub for a PCIe direct interconnection network with Altera's Stratix V. The use of partial reconfiguration makes it possible to switch multiple accelerators without stopping the switching hub. As a result, we reduced the time for place&route of an accelerator by 47% compared to the case of the design combining the accelerator into the switching hub. Chiharu Tsuruta, Takahiro Kaneda, Naoki Nishikawa, Hideharu Amano |
FPL | 4 |
| 2017 | FPGA-based accelerator for losslessly quantized convolutional neural networksabstractConvolutional Neural Networks (CNN) have been widely used for various computer vision tasks. While GPUs are the most common platform for CNN implementation, FPGAs are promising alternatives to provide better energy efficiency. Recent work demonstrates the potential of network quantization to reduce the model size and enhance computation efficiency while maintaining comparable accuracy to the full precision counterparts. Quantized CNN is especially suitable for FPGA implementation due to the presence of values with non-trivial bitwidth. In this paper, we present the design of an FPGA-based accelerator for losslessly quantized CNNs using High Level Synthesis tool. The experiment result shows that our design achieves 12.9 GOPS/Watt for quantized Alexnet on Imagnet Dataset. Mankit Sit, Ryosuke Kazami, Hideharu Amano |
FPT | 3 |
| 2017 | High-Bandwidth Low-Latency Approximate Interconnection NetworksabstractComputational applications are subject to various kinds of numerical errors, ranging from deterministic roundoff errors to soft errors caused by non-deterministic bit flips, which do not lead to application failure but corrupt application results. Non-deterministic bit flips are typically mitigated in hardware using various error correcting codes (ECC). But in practice, due to performance and cost concerns, these techniques do not guarantee error-free execution. On large-scale computing platforms, soft errors occur with non-negligible probability in RAM and on the CPU, and it has become clear that applications must tolerate them. For some applications, this tolerance is intrinsic as result quality can remain acceptable even in the presence of soft errors (e.g., data analysis applications, multimedia applications). Tolerance can also be built into the application, resolving data corruptions in software during application execution. By contrast, today's optical networks hold on to a rigid error-free standard, which imposes limits on network performance scalability. In this work we propose high-bandwidth, low-latency approximate networks with the following three features: (1) Optical links that exploit multi-level quadrature amplitude modulation (QAM) for achieving high bandwidth; (2) Avoidance of forward error correction (FEC), which makes optical link error-prone but affords lower latency; and (3) The use of symbol mapping coding between bit sequence and QAM to ensure data integrity that is sufficient for practical soft-error-tolerant applications. Discrete-event simulation results for application benchmarks show that approx networks achieve speedups up to 2.94 when compared to conventional networks. Daichi Fujiki, Kiyo Ishii, Ikki Fujiwara, Hiroki Matsutani, Hideharu Amano, Henri Casanova, Michihiro Koibuchi |
HPCA | 5 |
| 2017 | HiRy: An Advanced Theory on Design of Deadlock-Free Adaptive Routing for Arbitrary TopologiesabstractRecently proposed irregular networks can reduce the latency for both on-chip and off-chip systems with a large number of computing nodes and thus can improve the performance of parallel application. However, these networks usually suffer from deadlocks in routing packets when using a naive minimal path routing algorithm. To solve this problem, we focus attention on a lately proposed theory that generalizes the turn model to maintain the network performance with deadlock-freedom. The theorems remain a challenge of applying themselves to arbitrary topologies including fully irregular networks. In this paper, we advance the theorems to completely general ones. To apply the idea of the turn model to arbitrary topologies, we introduce a concept of regions that define continuous directions of channels on an n-dimensional space. Moreover, we provide a feasible implementation of a deadlock-free routing method based on our advanced theorem. To reduce the latency and the number of required Virtual Channels (VCs) with this method, a heuristic approach is introduced to reduce the number of prohibited turns between channels. Experimental results show that the routing method based on our proposed theorem can improve the network throughput by up to 138 % compared to a conventional deterministic minimal routing method. Moreover, it can reduce the latency by up to 2.9 % compared to another fully adaptive routing method. Ryuta Kawano, Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICPADS | 5 |
| 2017 | Order/Radix Problem: Towards Low End-to-End Latency Interconnection NetworksabstractWe introduce a novel graph called a host-switch graph, which consists of host vertices and switch vertices. Using host-switch graphs, we formulate a graph problem called an order/radix problem (ORP) for designing low end-to-end latency interconnection networks. Our focus is on reducing the host-to-host average shortest path length (h-ASPL), since the shortest path length between hosts in a host-switch graph corresponds to the end-to-end latency of a network. We hence define ORP as follows: given order (the number of hosts) and radix (the number of ports per switch), find a host-switch graph with the minimum h-ASPL. We demonstrate that the optimal number of switches can mathematically be predicted. On the basis of the prediction, we carry out a randomized algorithm to find a host-switch graph with the minimum h-ASPL. Interestingly, our solutions include a host-switch graph such that switches have the different number of hosts. We then apply host-switch graphs to interconnection networks and evaluate them practically. As compared with the three conventional interconnection networks (the torus, the dragonfly, and the fat-tree), we demonstrate that our networks provide higher performance while the number of switches can decrease. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
ICPP | 5 |
| 2017 | XYZ-Randomization using TSVs for Low-Latency Energy Efficient 3D-NoCsabstractIn this paper, we propose a method to design low latency and low energy networks for 3D Network-on-Chip (3D-NoC). Recent many-core processors require low-latency interconnection networks since the increasing number of cores limits the network performance. To achieve high performance in such many-core chips, small-world or random networks have been applied in the NoC field. However, the actual diameters and average shortest path lengths (ASPL) of these networks are far from the theoretical lower bound. In this work, we propose an approach based on the graph theory to design ultra low-latency topologies. We introduce a method to design a network that has low values of diameter and ASPL, with configurable upper bound of wire length, called opt ASPL. We also show that irregular topology, such as the topology used in opt ASPL, has a higher average energy consumption than general regular topology like 3D torus. In NoCs, energy budget and link length are limited, and thus such parameters must be carefully considered. Therefore, we introduce a multi-objective optimization for the ASPL and energy consumption called opt A/e which can obtain the Pareto optimal set useful for NoC designers. In a router with 64 nodes per chips and 4 chips stacked with a 3D-NoC, our proposed network optimized for energy consumption has a lower ASPL by 26.8% and a lower energy consumption by 10.9% compared to a 3D torus. Hiroshi Nakahara, Nguyen Anh Vu Doan, Ryota Yasudo, Hideharu Amano |
NOCS | 4 |
| 2017 | Implementation of Bitsliced AES Encryption on CUDA-Enabled GPU
Naoki Nishikawa, Hideharu Amano, Keisuke Iwai |
NSS | 2 |
| 2017 | Level-shifter-less approach for multi-VDD design to use body bias control in FD-SOIabstractLevel shifters to convert signal swings from low-voltage (VDDL) to high-voltage (VDDH) are required at the boundary of voltage domains in SoC employing multiple supply voltages. However, they cost delay, power and area in addition to increasing the complexity of physical design. This paper proposes a level-shifter-less (LSL) approach to use a reverse body bias (RBB) in the VDDH domain and superior threshold-voltage modulation capability of FD-SOI devices. Simulation results and measurements of a fabricated chip demonstrated that the chip applying the LSL approach correctly operates at VDDL=0.6V and VDDH=1.2V under RBB of 2V for pMOS transistors while suppressing the static dc current in the VDDH domain. Kimiyoshi Usami, Shunsuke Kogure, Yusuke Yoshida, Ryo Magasaki, Hideharu Amano |
VLSI-SoC | 5 |
| 2017 | Scalable Networks-on-Chip with Elastic Links Demarcated by Decentralized RoutersabstractAs the number of cores on a chip increases, Networks-on-Chip (NoCs) that connect many cores would face long links to reduce hop counts. The long links become bottlenecks in terms of both energy and RC delays as technology advances. To alleviate the negative impact of long links, we propose decentralized routers for NoCs. A decentralized router consists of multiple submodules that are positioned on a link, and hence the long links are segmented. Furthermore, we illustrate the design of an entire network that uses decentralized routers to obtain a good tradeoff between hop counts and wire delays per hop. Decentralized routers are effective especially in high-radix topologies, such as the flattened butterfly, and energy-delay product is reduced by greater than 60 percent. As NoCs become larger and more complex, the benefit of the decentralized routers will become more significant. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
IEEE Trans. Computers | 4 |
| 2017 | The First 25 Years of the FPL Conference: Significant PapersabstractA summary of contributions made by significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented. The 27 papers chosen represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2017 | Power Optimization Methodology for Ultralow Power Microcontroller With Silicon on Thin BOX MOSFETabstractIn this brief, a practical power optimization method that calculates the optimal power supply and body bias voltages, for a given target operational frequency and a temperature, is proposed and evaluated. The proposed optimization method is based upon a simple power model in which several coefficients for leakage power, switching power, temperature, and operational frequency are obtained from accurate real chip measurements. The calculated optimal-voltage settings by the proposed model can achieve minimum accuracies of 93.8%, 91.6%, and 79.5% for room-temperature, 50 °C, and 65 °C, respectively. Since the proposed methodology is based on well-known power formulas, it can be applied to the latest FD-SOI technologies. Hayate Okuhara, Yu Fujita, Kimiyoshi Usami, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | ACRO: Assignment of channels in reverse order to make arbitrary routing deadlock-freeabstractDistributed routing methods with small routing tables are scalable design on irregular networks for large-scale High Performance Computing (HPC) systems. Recently proposed compact routing methods, however, do not guarantee deadlock-freedom. Cyclic channel dependencies on arbitrary routing are typically removed with multiple Virtual Channels (VCs). However, challenges still remain to provide good trade-offs between a number of required VCs and a time complexity of an algorithm for assignment of VCs to paths. In this work, a novel algorithm ACRO is proposed for enriching arbitrary routing functions with deadlock-freedom with a reasonable number of VCs and a time complexity. Experimental results show that ACRO can reduce the average number of required VCs by up to 63% compared with the conventional algorithm that has the same time complexity. Furthermore, ACRO reduces a time complexity by a factor of O(|N| · log|N|) compared with that of the other conventional algorithm that needs almost the same number of VCs. Ryuta Kawano, Hiroshi Nakahara, Seiichi Tade, Ikki Fujiwara, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICIS | 7 |
| 2016 | Leveraging FDSOI through body bias domain partitioning and bias searchabstractIn FDSOI, sophisticated body biasing schemes can greatly reduce leakage or improve performance as well as efficiency. This paper proposes algorithms to determine body bias domain candidates which then merge those to reach a desired number of domains. Domain candidates are determined using an activation based approach, analyzing mapped verilog netlists to identify which parts of the design are used under specified conditions. Body bias domain partitionings are then determined based on activation and the timing of the partitioned parts. The algorithms include a body bias assignment algorithm to reach given timing goals with multiple domains and cross-domain resource sharing. The approach is compatible with any synthesis optimization and is resource sharing aware. Using an implementation of the proposed algorithms, overall leakage can be significantly reduced in all scenarios while obtaining the same benefits of body biasing. The method is evaluated in STMicro's 28nm FDSOI and Renesas's 65nm SOTB. Johannes Maximilian Kühn, Hideharu Amano, Oliver Bringmann 0001, Wolfgang Rosenstiel |
DAC | 2 |
| 2016 | Body bias grain size exploration for a coarse grained reconfigurable acceleratorabstractThis paper explores the grain of domain size of an energy efficient coarse grained reconfigurable array called CMA (Cool Mega Array). By using Genetic Algorithm based body bias assignment method, the leakage reduction of various grain size was evaluated. As a result, a domain with 2×1 PEs achieved about 40% power reduction with a 6% area overhead. Yusuke Matsushita 0001, Hayate Okuhara, Koichiro Masuyama, Yu Fujita, Ryuta Kawano, Hideharu Amano |
FPL | 6 |
| 2016 | Variable pipeline structure for Coarse Grained Reconfigurable Array CMAabstractCool mega-array (CMA) is a kind of coarse grained reconfigurable architecture (CGRA) which has shown its ability of ultra low-power computation. However, as CMA completely eliminates clock trees and registers, the performance improvement has been limited. In this paper, we introduce a variable pipeline structure to CMA with the minimum essential registers to provide more wide trade-off between performance and energy. Comparing with the baseline CMA (non-pipelined structure), an average of 77% improvement for performance was achieved with a small power overhead. Moreover, the energy efficiency was 1461 MOPS/mW at most which was about 2× that of the baseline structure. The best pipeline depth for an arbitrary energy-performance trade-off became selectable with only 11% area overhead. Naoki Ando, Koichiro Masuyama, Hayate Okuhara, Hideharu Amano |
FPT | 4 |
| 2016 | Trax solver on Zynq using incremental update algorithmabstractThis paper proposes a software/hardware co-design system for a Trax solver. Since development with Hardware Description Language (HDL) is tough work, we selected an approach: 1.writing C++ code, 2.finding bottleneck of the problem, and 3. re-writing the bottleneck part in the High Level Synthesis (HLS). Generally, game AI algorithm is divided into 3 parts, data structure, searching, and board evaluation. In this design, we focused on the data structure and implemented an incremental update algorithm. As a result, we can detect a line and an attack with O(1). Also, we implemented a local pruning, which reduces search area when the search depth becomes large in alpha beta tree search. The implemented solver works with 150MHz clock on Xilinx XC7Z020-CLG484 of Digilent ZedBoard. Hiroshi Nakahara, Tetsui Ohkubo, Hideki Shimura, Ryotaro Sakai, Chiharu Tsuruta, Takahiro Kaneda, Hideharu Amano |
FPT | 7 |
| 2016 | Randomizing Packet Memory Networks for Low-Latency Processor-Memory CommunicationabstractThree-dimensional stacked memory is considered to be one of the innovative elements for the next-generation computing system, for it provides high bandwidth and energy efficiency. Particularly, packet routing ability of Hybrid Memory Cubes (HMCs) enables new interconnects for the memories, giving flexibility to its topological design space. Since memory-processor communication is latency-sensitive, our challenge is to alleviate latency of the memory interconnection network, which is subject to high overheads from hop-count increase. Interestingly, random network topologies are known to have remarkably low diameter that is even comparable to theoretical Moore graph. In this context, we first propose to exploit the random topologies for the memory networks. Second, we also propose several optimizations to leverage the random topologies to be further adaptive to the latency-sensitive memory-processor communication: communication path length based selection, deterministic minimal routing, and page-size granularity memory mapping. Finally, we present interesting results of our evaluation: the random networks with universal memory access outperformed non-random networks of which memory access was optimally localized. Daichi Fujiki, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
PDP | 4 |
| 2016 | Efficient 3-D Bus Architectures for Inductive-Coupling ThruChip InterfacesabstractWireless 3-D network-on-chips (NoCs) with inductive-coupling ThruChip interfaces provide a large degree of flexibility for customizing the number of arbitrary chips in a package after chips have been fabricated. To simplify the vertical communication interfaces, static time division multiple access (TDMA) is used for the vertical broadcast buses, while arbitrary or customized topologies can be used for the intrachip network. This paper proposes two techniques to break through the simple static TDMA-based vertical buses while maintaining a simple communication interface. The first technique is headfirst sliding (HS) routing to reduce the waiting time for acquiring the communication time-slot. HS routing selects the best vertical bus based on the current time, taking advantage of static TDMA. The second technique extends carrier sense multiple access with collision detection (CSMA/CD) for vertical broadcast buses. We introduce a packet collision detection technique for inductive-coupling buses and propose two retransmission strategies to reduce the waiting time for packet retransmissions caused by collisions. Network simulation results show that HS routing reduces the communication latency by 39.1% compared with the conventional static TDMA bus-based 3-D NoC that uses the shortest path routing. The proposed CSMA/CD bus also improves the latency by 52.5% and throughput by 34.1%. The full-system simulation results show that HS routing and the proposed CSMA/CD technique reduce the application execution time accordingly while maintaining the average flit transfer energy overhead modest. Takahiro Kagami, Hiroki Matsutani, Michihiro Koibuchi, Yasuhiro Take, Tadahiro Kuroda, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2015 | A metamorphotic Network-on-Chip for various types of parallel applicationsabstractA metamorphotic Network-on-Chip (NoC) architecture is proposed in order to customize for performance or energy consumption on a per-application basis. Adding reconfigurability on conventional topologies has been studied so far especially for application workloads that can be statically analyzed. In this context, we propose such a platform to take care of both the static and the dynamic cases where application workloads cannot be statically analyzed while performance or energy constraints are given. Our metamorphotic NoC reconfigures its topology, routing, operating frequency, and supply voltage based on the following three modes. 1) Regular mode uses a traditional mesh topology for neighboring communications. As the link length is short and uniform, it can be operated at a higher frequency and higher voltage, while a long-range communication increases the path length. 2) Random mode uses a random topology for unknown workloads to reduce the path length by exploiting the small-world effect. As the path length is reduced but the wire delay is increased, it is intended for a lower operating frequency and lower voltage. 3) Custom mode uses an optimized topology for given workloads. To support Random and Custom modes, assembled multiplexers are embedded into the metamorphotic NoC. Random and Regular/Custom modes are generated by randomly or selectively reconfiguring these multiplexers, respectively, based on the performance or energy constraints. This paper explores the design space of assembled multiplexers and provides a reasonable design recommendation through a graph analysis. It is demonstrated based on experimental results on the area overhead, operating frequency, network performance, and energy consumption. The results show that Regular mode can operate at 1.27GHz and Random mode can reduce the average network latency by 19.6% and the energy consumption by 44.2% compared with a traditional NoC that has mesh topology with little overhead. Custom mode can reduce them as well as Random mode. Seiichi Tade, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
ASAP | 3 |
| 2015 | Spatial and temporal granularity limits of body biasing in UTBB-FDSOI
Johannes Maximilian Kühn, Dustin Peterson, Hideharu Amano, Oliver Bringmann 0001, Wolfgang Rosenstiel |
DATE | 3 |
| 2015 | Reduction calculator in an FPGA based switching Hub for high performance clustersabstractUnused logic in the field-programmable gate array (FPGA) for the switching hub is one potential resource to accelerate the computation of data exchanged through the hub. However, for large scale scientific computation, it is difficult to implement such an accelerator on the FPGA used in high performance computers. Here, a reduction calculator for executing ARGOT (accelerated radiative transfer on grids using oct-tree) to solve the radiative transfer equation used for simulation of astronomical objects is implemented on the FPGA of PEACH2 (PCI Express Adaptive Communication Hub ver2), a low latency switching hub for high performance GPU (graphics processor unit) clusters. The implemented reduction calculator uses a pipelined tree of adders and works with a 150-MHz clock without affecting the switching hub functions. Use of the DMA (direct memory access) transfer with descriptors made it possible to improve the performance of CPU excution by a maximum of about 45 times in a real system. Takuya Kuhara, Chiharu Tsuruta, Toshihiro Hanawa, Hideharu Amano |
FPL | 4 |
| 2015 | Significant papers from the first 25 years of the FPL conferenceabstractThe list of significant papers from the first 25 years of the Field-Programmable Logic and Applications conference (FPL) is presented in this paper. These 27 papers represent those which have most strongly influenced theory and practice in the field. Philip H. W. Leong, Hideharu Amano, Jason Helge Anderson, Koen Bertels, João M. P. Cardoso, Oliver Diessel, Guy Gogniat, Mike Hutton, Wayne Luk, Patrick Lysaght, Marco Platzner, Viktor Prasanna 0001, Tero Rissa, Cristina Silvano, Hayden Kwok-Hay So, Yu Wang 0002 |
FPL | 2 |
| 2015 | 7MOPS/lemon-battery image processing demonstration with an ultra-low power reconfigurable accelerator CMA-SOTB-2abstractCool Mega Array (CMA)-SOTB-2 is an ultra-low energy Coarse Grained Reconfigurable Architecture[1] (CGRA) for recent advanced sensor networks, Internet of Things and wearable computing. It has a large Processing Element (PE) array without memory elements for mapping an application's data-flow graph, a small simple programmable μ-controller for data management, and data memory. Unlike traditional coarse grained reconfigurable processors, the power consumption for hardware context switching, storing intermediate data in registers, and clock distribution for them are eliminated from PE array which occupies large area of a chip. It is implemented by using Silicon on Thin BOX (SOTB) CMOS, a new process technology developed by the Low-power Electronics Association & Project (LEAP). Koichiro Masuyama, Yu Fujita, Hayate Okuhara, Hideharu Amano |
FPL | 4 |
| 2015 | Trax solver on Zynq with Deep Q-NetworkabstractA software/hardware co-design system for a Trax solver is proposed. Implementation of Trax AI is challenging due to its complicated rules, so we adopted an embedded system called Zynq (Zynq-7000 AP SoC) and introduced a High Level Synthesis (HLS) design. We also added Deep Q-Network, a machine learning algorithm, to the system for use as an evaluation function. Our solver automatically optimizes its own evaluation function through games with humans or other AIs. The implemented solver works with a 150-MHz clock on the Xilinx XC7Z020-CLG484 of a Digilent ZedBoard. A part of the Deep Q-Network job can be executed on the FPGA of the Zynq board more than 26 times faster than with ARM Coretex-A9 650-MHz software. Naru Sugimoto, Takuji Mitsuishi, Takahiro Kaneda, Chiharu Tsuruta, Ryotaro Sakai, Hideki Shimura, Hideharu Amano |
FPT | 7 |
| 2015 | An optimal power supply and body bias voltage for a ultra low power micro-controller with silicon on thin box MOSFETabstractBody bias control is an efficient means of balancing the trade-off between leakage power and performance especially for chips with silicon on thin buried oxide (SOTB), a type of FD-SOI technology. In this work, a method for finding the optimal combination of the supply voltage and body bias voltage to the core and memory is proposed and applied to a real micro-controller chip using SOTB CMOS technology. By obtaining several coefficients of equations for leakage power, switching power and operational frequency from the real chip measurements, the optimized voltage setting can be obtained for the target operational frequency. The power consumption lost by the error of optimization is 12.6% at maximum, and it can save at most 73.1% of power from the cases where only the body bias voltage is optimized. This method can be applied to the latest FD-SOI technologies. Hayate Okuhara, Kuniaki Kitamori, Yu Fujita, Kimiyoshi Usami, Hideharu Amano |
ISLPED | 5 |
| 2015 | On-Chip Decentralized Routers with Balanced Pipelines for Avoiding Interconnect BottleneckabstractTechnology scaling makes designers face difficulties dealing with wire delay of long global interconnects, especially for high-radix networks. In this context, we propose decentralization of on-chip packet routers. A decentralized router consists of submodules, each of which has particular functionality and they are scattered on a link, thereby long wires are segmented. Our starting point is from a conventional router architecture, and we illustrate four case studies to generalize our proposal. We also propose a new buffer design and how to balance pipelines of a router. A proof-of-concept is shown in 28-nm process technology. Our results demonstrate that the decentralization of an on-chip router enables Link Traversal (LT) stages to be eliminated, and the critical path delay is improved by up to 45% with the reduced area compared with a conventional router. As technology advances, the benefit of the decentralized routers become more substantial in the nano-scale era. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
NOCS | 4 |
| 2015 | Optimized Core-Links for Low-Latency NoCsabstractIn recent many-core architectures, the number of cores has been steadily increasing and thus the network latency between cores becomes an important issue for parallel application programs. Because packet-switched network structures are widely used for core-to-core communications, a topology among cores has a major impact on the network latency. It has been reported that a small-world Network-on-Chip that adds links between randomly-selected routers on a regular router topology is effective for reducing the network latency. In this study, we extend this framework by connecting multiple links between a single core and quasi-optimally selected neigh boring routers to form multiple links from each core on a 2D MESH router topology. Results obtained by a flit-level discrete event simulator show that our optimized core-link topologies can achieve the average latency up to 48% lower than that of baseline topologies. Furthermore, full-system CMP simulation results show that by using optimized core-links we can improve the application execution time on the NAS Parallel Benchmarks by up to 10.1%. Ryuta Kawano, Seiichi Tade, Ikki Fujiwara, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
PDP | 5 |
| 2014 | Design and control methodology for fine grain power gating based on energy characterization and code profiling of microprocessorsabstractThis paper presents a design and control scheme of a microprocessor whose internal function units are power gated at instruction-by-instruction basis. Enabling/disabling the power gating is adaptively controlled under the support of on-chip leakage monitors and the operating system to minimize energy overhead due to sleep-in and wakeup. Measured results of the fabricated chip in the 65nm CMOS technology demonstrated that our approach reduces energy to 21-35% in the range of 25-85°C as compared to the non power-gated case. Energy dissipation was reduced by up to 15% as compared to the conventional fine-grain power gating technique in the same temperature range. Kimiyoshi Usami, Masaru Kudo, Kensaku Matsunaga, Tsubasa Kosaka, Yoshihiro Tsurui, Hideharu Amano, Hiroaki Kobayashi, Ryuichi Sakamoto, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ASP-DAC | 7 |
| 2014 | Design and evaluation of fine-grained power-gating for embedded microprocessorsabstractPower-performance efficiency is still remaining a primary concern for microprocessor designers. One of the sources of power inefficiency for recent LSI chips is increasing leakage power consumption. Power-gating is a well known technique to reduce leakage power consumption by switching off the power supply to idle logic blocks. Recently, fine-grained power-gating is emerged as a technique to minimize leakage current during the active processor cycles by switching on and off a logic blocks in much finer temporal/spatial granularity. Though fine-grained power-gating is useful, a comprehensive evaluation and analysis has not been conducted on a real LSI chips. In this paper, we evaluate fine-grained run-time power-gating for microprocessors' functional units using a real embedded microprocessor. We also introduce an architecture and compiler co-operative power-gating scheme which mitigates negative power reduction caused by the energy overhead associated with finegrained power-gating. The experimental results with a fabricated core shows that a hardware-based scheme saves power consumption of functional units by 44% and hardware compiler co-operative scheme further improves power efficiency by 5.9% when core temperature is 25 ˚C. Masaaki Kondo, Hiroaki Kobayashi, Ryuichi Sakamoto, Motoki Wada, Jun Tsukamoto, Mitaro Namiki, Hideharu Amano, Kensaku Matsunaga, Masaru Kudo, Kimiyoshi Usami, Toshiya Komoda, Hiroshi Nakamura |
DATE | 8 |
| 2014 | Low-latency wireless 3D NoCs via randomized shortcut chipsabstractIn this paper, we demonstrate that we can reduce the communication latency significantly by inserting a fraction of randomness into a wireless 3D NoC (where CMOS wireless links are used for vertical inter-chip communication) when considering the physical constraints of the 3D design space. Towards this end, we consider two cases, namely 1) replacing existing horizontal 2D links in a wireless 3D NoC with randomized shortcut NoC links and 2) enabling full connectivity by adding a randomized NoC layer to a wireless 3D platform with partial or no horizontal connectivity. Consequently, the packet routing is optimized by exploiting both the existing and the newly added random NoC. At the same time, by adding randomly wired shortcut NoCs to a wireless 3D platform, a good balance can be established between the modularity of the design and the minimum randomness needed to achieve low latency, and experimental results show that by adding a random NoC chip to wireless 3D CMPs without built-in horizontal connectivity, the communication latency can be reduced by as much as 26.2% when compared to adding a 2D mesh NoC. Also, the application execution time and average flit transfer energy can be improved accordingly. Hiroki Matsutani, Michihiro Koibuchi, Ikki Fujiwara, Takahiro Kagami, Yasuhiro Take, Tadahiro Kuroda, Paul Bogdan, Radu Marculescu, Hideharu Amano |
DATE | 9 |
| 2014 | A high speed design and implementation of dynamically reconfigurable processor using 28NM SOI technologyabstractAlthough dynamically reconfigurable processor arrays (DRPAs) are advantageous for embedded devices because of their high energy efficiency, many of the recent mobile devices are required to execute increasingly performance-centric jobs. One fairly straingtfoward way of increasing the clock frequency is introducing a pipelined structure into each PE. However, this results in frequent pipeline stalls due to the data hazard between multiple PEs. In order to mitigate the effect of data hazard between PEs, we propose a tiny vector instruction mechanism. With a single vector instruction, a small amount of data is continuously processed in the pipeline of the PE. Pipeline stalls are removed without increasing the number of hardware contexts, and thus the amount of configuration data. Evaluation results based on the implementation using 28nm SOI process technology, a DRPA with tiny vector instructions (DRPA-TVI) improves the performance by 2.4 three times compared to a base DRPA with just a small increase of area and power consumption. Toru Katagiri, Hideharu Amano |
FPL | 2 |
| 2014 | Body bias control for a coarse grained reconfigurable accelerator implemented with Silicon on Thin BOX technologyabstractFor low power yet high performance processing in battery driven devices, a coarse grained reconfigurable accelerator called Cool Mega Array (CMA)-SOTB is implemented by using Silicon on Thin BOX (SOTB), a new process technology developed by the Low-power Electronics Association & Project (LEAP). A real chip using a 65nm experimental process achieved a sustained performance of 192MOPS with a power supply of 0.4V and power consumption of 1.7mW. A clock frequency of 89MHz was achieved with a power supply of just 0.4V when a forward bias voltage was given. When using a reverse bias, the leakage current could be suppressed to less than 20µW in the stand-by mode. The key concept of CMA-SOTB is maintaining a balance between performance and leakage current by independently controlling the bias voltages of the PE array and the microcontroller. Evaluations of the operational frequency and power consumption of filter application programs shed light on how to find the combination of bias voltages that achieves the best energy efficiency for a required performance. The range of advantageous power supply voltage for a required performance considering the body bias was also found. Honlian Su, Yu Fujita, Hideharu Amano |
FPL | 3 |
| 2014 | Image processing by A 0.3V 2MW coarse-grained reconfigurable accelerator CMA-SOTB with a solar batteryabstractCool mega array with silicon on thin box (CMA-SOTB) is an extremely low power coarse grained reconfigurable accelerator. It was implemented by using the SOTB technology developed by a Japanese national project, low-power electronics association & project (LEAP). Making the best use of such a device and low energy architectural techniques, CMA-SOTB works more than 25MHz clock with less than 0.3V supply voltage. Various kind of optimization can be done by controlling the body bias voltage for PE array and micro-controller independently. The demonstration using CMA-SOTB first shows that a simple image processing application can work with a 0.25V-0.4V solar battery. Then the leakage power control by changing the body bias is demonstrated. In the stand-by mode, less than 20μW power is consumed by using strong reverse bias. Yu Fujita, Koichiro Masuyama, Hideharu Amano |
FPT | 3 |
| 2014 | Hardware/software co-design architecture for Blokus Duo solverabstractThis paper presents a software and hardware design of an FPGA-based Blokus Duo solver. We used Embedded system called ZYNQ-7000 All Programmable SoC to implement the solver. By combining hardware with software, efficient acceleration is performed. Our system searches a game tree by using the miniMax algorithm with alpha-beta pruning. The implemented solver works at 75MHz with Xilinx Zynq-7000 AP SoC XC7Z020-CLG484 on the Digilent ZedBoard. It can search states after three moves in most cases. Naru Sugimoto, Hideharu Amano |
FPT | 2 |
| 2014 | A perpetuum mobile 32bit CPU on 65nm SOTB CMOS technology with reverse-body-bias assisted sleep modeabstractPresents a conference poster that addresses a perpetuum mobile 32bit central processing unit that resides on 65nm CMOS technology using reverse-body-bias via assisted sleep mode. Shiro Kamohara, Nobuyuki Sugii, Koichiro Ishibashi, Kimiyoshi Usami, Hideharu Amano, Kazutoshi Kobayashi, Cong-Kha Pham |
Hot Chips Symposium | 5 |
| 2014 | Design of a low power NoC router using Marching Memory Through typeabstractPower consumption of Network-on-Chip (NoC) is becoming more important in many core processors. Input buffers utilized in routers consume a significant part of the total power of NoCs. In order to reduce this power consumption, a novel power efficient memory called Marching Memory Through type (MMTH) is introduced. By connecting transparent latches in tandem, MMTH achieves high speed operation with a low power consumption. MMTH, however, requires a certain overhead at read operation, and hence we propose a latency reduction scheme based on the look-ahead routing. The proposed router was designed in Renesas's 40nm process and compared with a standard router using conventional register-based FIFOs in terms of the network performance, application performance, and power consumption. The result of evaluation shows that the proposed router reduces the power consumption by 42.4% on average at 2GHz and the expense of only 0.5-2.0% performance overhead. Ryota Yasudo, Takahiro Kagami, Hideharu Amano, Yasunobu Nakase, Masashi Watanabe, Tsukasa Oishi, Toru Shimizu, Tadao Nakamura |
NOCS | 3 |
| 2014 | 3D NoC with Inductive-Coupling Links for Building-Block SiPsabstractA wireless 3D NoC architecture is described for building-block SiPs, in which the number of hardware components (or chips) in a package can be changed after chips have been fabricated. The architecture uses inductive-coupling links that can connect more than two examined dies without wire connections. Each chip has data transceivers for the uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a vertical unidirectional ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring, we use bubble flow control, which does not rely on the conventional VC-based deadlock avoidance mechanism. In addition, we propose a bidirectional communication scheme to form a bidirectional ring network by using the inductive-coupling transceivers that can dynamically change the communication modes, such as TX, RX, and Idle modes. This paper illustrates the inductive-coupling transceiver circuits, which can carry high data transfer rates of up to 8 Gbps per channel, for the wireless 3D NoC. It also illustrates an implementation of a wireless 3D NoC that has on-chip routers and transceivers implemented with a 65 nm process in order to show the feasibility of our proposal. The vertical bubble flow control and conventional VC-based approach on the uni- and bidirectional ring networks are compared with the vertical broadcast bus in terms of throughput, hardware amount, and application performance using a full system multiprocessor simulator. The results show that the proposed bidirectional communication scheme efficiently improves application performance without adding any inductive-coupling transceivers. In addition, the proposed vertical bubble flow network outperforms the conventional VC-based approach by 7.9-12.5 percent with a 33.5 percent smaller router area for building-block SiPs connecting up to eight chips. Yasuhiro Take, Hiroki Matsutani, Daisuke Sasaki, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
IEEE Trans. Computers | 6 |
| 2013 | A case for wireless 3D NoCs for CMPsabstractInductive-coupling is yet another 3D integration technique that can be used to stack more than three known-good-dies in a SiP without wire connections. We present a topology-agnostic 3D CMP architecture using inductive-coupling that offers great flexibility in customizing the number of processor chips, SRAM chips, and DRAM chips in a SiP after chips have been fabricated. In this paper, first, we propose a routing protocol that exchanges the network information between all chips in a given SiP to establish efficient deadlock-free routing paths. Second, we propose its optimization technique that analyzes the application traffic patterns and selects different spanning tree roots so as to minimize the average hop counts and improve the application performance. Hiroki Matsutani, Paul Bogdan, Radu Marculescu, Yasuhiro Take, Daisuke Sasaki, Hao Zhang 0020, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
ASP-DAC | 9 |
| 2013 | Demonstration of a heterogeneous multi-core processor with 3-D inductive coupling linksabstractCube-1 is a heterogeneous multi-core processor which can achieve the required performance with the least energy consumption as possible. It can control the performance and energy with two levels: (1) the number of accelerators can be easily changed by increasing or decreasing the number of stacked chips after fabrication, as they are connected with inductive coupling links. (2) The supply voltage for PE array of the accelerator can be controlled by the host CPU so that the required performance can be obtained with a minimum supply voltage. Yusuke Koizumi, Noriyuki Miura, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 6 |
| 2013 | A fully pipelined FPGA architecture for stochastic simulation of chemical systemsabstractSimulation of chemical systems allows bio-chemists to understand how the interactions of individual molecules can lead to cellular and organism level behaviour. When the concentration of moleculesis very small, it is necessary to model every single chemical interaction in a Monte-Carlo simulation, presenting a huge computational burden. This paper presents a new fully pipelined architecture for chemical simulation, which avoids the traditional approach of optimising for minimum operation count, and instead optimises for throughput and parallelism. We show that even though this leads to a higher asymptotic operation count per simulation step, it allows for a much greater degree of spatial and pipeline parallelism, and the increased area is offset by much greater throughput. The new architecture is implemented in a Virtex-6 SX475T and can sustain a rate of over 1 billion reactions per second for problems with less than 64 reactions. Compared against existing chemical simulators on small to medium size chemical models, the new architecture is 30-100 times faster than a commercial software simulator running on an 8-core 3.4GHz Core i7, and 12-30 times faster than the best existing FPGA simulators. David B. Thomas, Hideharu Amano |
FPL | 2 |
| 2013 | A hardware complete detection mechanism for an energy efficient reconfigurable accelerator CMAabstractCool Mega Array (CMA) is an energy efficient Coarse Grained Reconfigurable processor Array (CGRA) consisting of a large PE (Processing Element) array. In order to reduce the power for storing intermediate results and clock tree, the PE array is consisting of combinatorial circuits. A hardware completion detection mechanism for CMA is proposed, implemented and evaluated. Each PE uses serially connected buffers with selectable taps, and the delay is decided according to the operation executed in the PE. Since the completion signal is transferred exactly on the same paths that for computation, the delay in the switch and wires are accounted. The post layout simulation revealed that the same performance without the mechanism can be obtained only with 5.1% area overhead and less than 6% extra power consumption. With the mechanism, a single micro-code can be used for various supply voltages to PE array. Akihito Tsusaka, Mai Izawa, Rie Uno, Nobuyuki Ozaki, Hideharu Amano |
FPL | 5 |
| 2013 | Task level pipelining with PEACH2: An FPGA switching fabric for high performance computingabstractWe demonstrate task level pipelining on multiple accelerators with PEACH2. PEACH2 is implmented on FPGA, and enables ultra low latency direct communication among multiple accelerators over computational nodes. By installing PEACH2, typical high performance computation nodes are tightly coupled. In this environment, application can be accelerated by exploiting not only data level parallelism, but also task level pipelined operation. Furthermore, we can processe multiple task on multiple accelerators in a pipelined manner. In our demonstration, application achieves 44% speed up compared to a single GPU. Takaaki Miyajima, Takuya Kuhara, Toshihiro Hanawa, Hideharu Amano, Taisuke Boku |
FPT | 4 |
| 2013 | A low power reconfigurable accelerator using a back-gate bias control techniqueabstractLeakage power is a serious problem especially for accerelators which use a large size Processing Element (PE) array. Here, a low power reconfigurable accelerator called Cool Mega Array (CMA) with back-gate bias control (CMA-bb) is implemented and evaluated. In CMA-bb, the back-gate bias of the microcontroller and PE array can be controlled independently. In the idle mode, reverse bias is given to the both parts to suppress the leakage current. When high performance is required, forward bias is used to increase the clock frequency. For simple applications, the operational power can be suppressed by using reverse bias only in the PE array. The real chip is implemented with a 65nm experimental process for low leakage applications. The evaluation results show that the leakage current can be suppressed to 300μA by using the reverse bias. The operational frequency is increased from 39MHz to 50MHz with up to 21% increase of operational power by using the forward bias. For simple applications, 8% to 9.4% of operational power is saved by giving reverse bias only to the PE array. Hongliang Su, Kuniaki Kitamori, Hideharu Amano |
FPT | 4 |
| 2013 | Artificial intelligence of Blokus Duo on FPGA using Cyber Work BenchabstractThis paper presents a design of an FPGA-based Blokus Duo solver. It searches a game tree by using the miniMax algorithm with alpha-beta pruning and move ordering. In addition, HLS tool called CyberWorkBench (CWB) is used to implement hardware. By making the use of functions in CWB, parallel fully pipelined design is generated. The implemented solver works at 100MHz with Xilinx Spartan-6 XC6SLX45 FPGA on the Digilent Atlys board. It can search states after three moves in most cases. Naru Sugimoto, Takaaki Miyajima, Takuya Kuhara, Yuki Katuta, Takushi Mitsuichi, Hideharu Amano |
FPT | 6 |
| 2013 | Partially reconfigurable flux calculation scheme in advection term computationabstractFast Aerodynamics Routines (FaSTAR) is one of the most recent fluid dynamics software package. The problem of FaSTAR is hard to be executed in parallel machines because of its irregular and unpredictable data structure. Exploiting reconfigurable hardware with their advantages to make up for the inadequacy of the existing high performance computers had gradually become the solutions. However, a single FPGA is not enough for the FaSTAR package because the whole module is very large. Instead of using many FPGAs, partially reconfigurable hardware available in recent FPGAs is explored for this application. Advection term computation module in FaSTAR is chosen as a target subroutine. We proposed a reconfigurable flux calculation scheme using partial reconfiguration technique to save hardware resources to fit in a single FPGA. We developed flux computational module and five flux calculation schemes are implemented as reconfigurable modules. This implementation has advantages of up to 62.75% resource saving and enhancing the configuration speed by 6.28 times. Performance evaluation also shows that 2.65 times acceleration is achieved compared to Intel Core 2 Duo at 2.4 GHz. Mohamad Sofian Abu Talip, Takayuki Akamine, Mao Hatto, Yasunori Osana, Naoyuki Fujita, Hideharu Amano |
FPT | 6 |
| 2013 | A speculative gather system for Cool Mega-ArrayabstractCool Mega Array (CMA) is a low power reconfigurable processor array for battery driven mobile devices. A prototype chip CMA-1 consists of a 8 × 8 PE (Processing Element) array and a micro-controller for controlling data alignment. Because the PE array of CMA is built with a combinatorial circuit, it does not have a signal which tells that operation in the PE array was completed. A propagate delay of the whole PE array corresponding to the operation time was estimated by using the data path and mapping information in the design stage of the application. The timing information for gathering the data was specified in the microcode of the controller. However, since this timing is fixed, it cannot treat the variation of environment temperature and voltage scaling for the PE array. Here, a speculative gather system is proposed which sets the timing of collecting operation results from the PE array dynamically. By collecting results twice and comparing them, it guarantees the correctness of the operation results and adjusts the gather timing automatically. The speculative gather system is implemented in the CMA, and evaluation results appear that the performance is improved by 25.3% on average with the overhead of 0.5% in area and 3.1% in power consumption. Rie Uno, Nobuaki Ozaki, Mai Izawa, Akihito Tsusaka, Takaaki Miyajima, Hideharu Amano |
FPT | 6 |
| 2013 | A scalable 3D heterogeneous multi-core processor with inductive-coupling thruchip interface
Noriyuki Miura, Yusuke Koizumi, Eiichi Sasaki, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
Hot Chips Symposium | 7 |
| 2013 | A Routing Strategy for Inductive-Coupling Based Wireless 3-D NoCs by Maximizing Topological Regularity
Daisuke Sasaki, Hao Zhang 0020, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICA3PP (2) | 5 |
| 2013 | Headfirst sliding routing: A time-based routing scheme for bus-NoC hybrid 3-D architectureabstractA contact-less approach that connects chips in vertical dimension has a great potential to customize components in 3-D chip multiprocessors (CMPs), assuming card-style components inserted to a single cartridge communicate each other wirelessly using inductive-coupling technology. To simplify the vertical communication interfaces, static Time Division Multiple Access (TDMA) is used for the vertical broadcast buses, while arbitrary or customized topologies can be used for intra-chip networks. In this paper, we propose the Headfirst sliding routing scheme to overcome the simple static TDMA-based vertical buses. Each vertical bus grants a communication time-slot for different chips at the same time periodically, which means these buses work with different phases. Depending on the current time, packets are routed toward the best vertical bus (elevator) just before the elevator acquires its communication time-slot. Network simulations show that Headfirst sliding routing reduces the communication latency by up to 32.7%, and full-system CMP simulations show that it reduces application execution time by 9.9%. Synthesis results show that the area and critical path delay overheads are modest. Takahiro Kagami, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
NOCS | 4 |
| 2012 | A multi-Vdd dynamic variable-pipeline on-chip router for CMPsabstractWe propose a multi-voltage (multi-Vdd) variable pipeline router to reduce the power consumption of Network-on-Chips (NoCs) designed for chip multi-processors (CMPs). Our multi-Vdd variable pipeline router adjusts its pipeline depth (i.e., communication latency) and supply voltage level in response to the applied workload. Unlike dynamic voltage and frequency scaling (DVFS) routers, the operating frequency is the same for all routers throughout the CMP; thus, there is no need to synchronize neighboring routers working at different frequencies. In this paper, we implemented the multi-Vdd variable pipeline router, which selects two supply voltage levels and pipeline modes, using a 65nm CMOS process and evaluated it using a full-system CMP simulator. Evaluation results show that although the application performance degraded by 1.0% to 2.1%, the standby power of NoCs reduced by 10.4% to 44.4%. Hiroki Matsutani, Yuto Hirata, Michihiro Koibuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
ASP-DAC | 6 |
| 2012 | Performance analysis of fully-adaptable CRC accelerators on an FPGAabstractWe present a methodology for designing high-speed fully-adaptable Cyclic Redundancy Check (CRC) accelerators capable of supporting wide range of CRC standards. We extend our previous research with a module for generating contents of look-up tables, and we design new overlapped pipelined architecture. The resulting integration requires minimal resource and it ensures fast table re-generating process. Our accelerators achieve highest throughput when compared to related work, with possibility of additionally increasing throughput by extending the number of bits processed at a time. On the Xilinx Virtex 6 LX550T board they occupy between 1-2% area to produce maximum of 289.8Gbps with BRAM, or between 1.6 - 14% of area for 418.8Gbps without BRAM. Amila Akagic, Hideharu Amano |
FPL | 2 |
| 2012 | Reconfigurable out-of-order mechanism generator for unstructured grid computation in computational fluid dynamicsabstractFaSTAR developed by JAXA is a leading edge CFD (Computational Fluid Dynamics) program package which supports various solvers based on unstructured grids. The computation based on unstructured grid causes a lot of pipeline stalls by RAW (Read After Write) hazard when reconfigurable accelerators are implemented in FPGAs. In order to cope with this problem, the OoO (Out-of-Order) mechanism generator is proposed. By setting parameters depending on the target computation, the OoO mechanism with appropriate structure of the execution unit and waiting buffer is generated. The OoO mechanisms are applied to five subroutines in FaSTAR, and it achieved 2.6 times performance as the case of in-order execution, and 2.9 times as the software executed by Intel Core2Duo processor with reasonable amount of overhead. Takayuki Akamine, Kenta Inakagata, Yasunori Osana, Naoyuki Fujita, Hideharu Amano |
FPL | 5 |
| 2012 | CMA-Cube: A scalable reconfigurable accelerator with 3-D wireless inductive coupling interconnectabstractCMA-Cube is the second prototype of building block scalable reconfigurable accelerator using inductive coupling interconnect. It uses the wireless inductive coupling interconnect as a packet switching network which connects accelerators. As an accelerator core, CMA (Cool Mega Array), which consists of a large coarse-grained PE array with combinatorial circuits and tiny micro-controller, is applied. Evaluation results of Cube-1 Quad Core which consists of a host embedded CPU and three CMA-Cubes achieved 3.15 times performance acceleration as that without accelerators when JPEG decoder is executed. Yusuke Koizumi, Eiichi Sasaki, Hideharu Amano, Hiroki Matsutani, Yasuhiro Take, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 3 |
| 2012 | A study of adaptable co-processors for Cyclic Redundancy Check on an FPGAabstractCyclic Redundancy Check (CRC) is a well known error detection scheme used to detect corruption of digital content in digital networks and storage devices. In this paper, we present a study of different approaches of designing highly adaptable co-processors for CRC on an FPGA which are used in many network and server applications. The results of our research are two new architectures: adaptable and dynamically re-configurable CRC co-processors. Both architectures are highly flexible in terms of a number of CRC standards they support. We explored their scalability by processing different amount of input messages at a time. Results show that throughput doubles when we double the amount of data processed at a time. Our experimental results on adaptable CRC co-processor demonstrate re-generation latency as low as .9 - 4.52μs and throughput between 27.8 - 418Gbps (64 - 1024 bits of an input message). The re-configuration latency of dynamic parts of other CRC co-processor was significantly higher .3 - .45s, but area utilization was the least. The throughput of this architecture was between 29.25 - 347.37 Gbps. Amila Akagic, Hideharu Amano |
FPT | 2 |
| 2012 | Dynamic power control with a heterogeneous multi-core system using a 3-D wireless inductive coupling interconnectabstractCube-2 is a prototype of building block scalable reconfigurable accelerator using an inductive coupling interconnect. It is consisting of a ultra low leakage embedded processor Geyser and coarse-grained reconfigurable accelerators CMA (Cool Mega Array). A Geyser chip and multiple CMA chips are stacked, and a powerful network is formed by using the inductive coupling interconnect. The performance can be enhanced by increasing the number of CMA chips. JPEG decoder is implemented with a cooperation of Geyser and CMAs, and low power execution by controlling the power supply voltage of CMAs is demonstrated. Yusuke Koizumi, Hideharu Amano, Hiroki Matsutani, Noriyuki Miura, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPT | 2 |
| 2012 | A case for random shortcut topologies for HPC interconnectsabstractAs the scales of parallel applications and platforms increase the negative impact of communication latencies on performance becomes large. Fortunately, modern High Performance Computing (HPC) systems can exploit low-latency topologies of high-radix switches. In this context, we propose the use of random shortcut topologies, which are generated by augmenting classical topologies with random links. Using graph analysis we find that these topologies, when compared to non-random topologies of the same degree, lead to drastically reduced diameter and average shortest path length. The best results are obtained when adding random links to a ring topology, meaning that good random shortcut topologies can easily be generated for arbitrary numbers of switches. Using flit-level discrete event simulation we find that random shortcut topologies achieve throughput comparable to and latency lower than that of existing non-random topologies such as hypercubes and tori. Finally, we discuss and quantify practical challenges for random shortcut topologies, including routing scalability and larger physical cable lengths. Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano, D. Frank Hsu, Henri Casanova |
ISCA | 3 |
| 2012 | An OpenCL Runtime Library for Embedded Multi-Core AcceleratorabstractIn recent years, improvements of energy efficiency and computational performance have become a major issue, because smartphones and tablets become popular. To implement high performance, multi-core accelerator consists of general purpose processors and accelerators are often used. But to use these multi-core accelerator efficiently, programmers have to consider synchronization and data transfer between accelerators, memory and I/O. Therefore frameworks such as OpenCL have been proposed for effective use of parallel computing resources of multi-core processors. OpenCL frameworks for GPGPU, Intel multi-core and many other kind of multi-core processors have been developed on Linux. However, in order to use those multi-core accelerators effectively, it is necessary to control accelerators with software environment aiming effective use of accelerators. Generally user mode application program calls OS to synchronize of accelerator's data transfer and termination of the its execution with overhead. In this paper, accelerators are controlled from collaborated OS and OpenCL library, instead of user mode application program. Our feature of these methods is that OS cooperates with library to reduce overhead of control accelerators. The proposed framework in this paper integrates the execution and data transfer control in an OpenCL library and an embedded OS. The OpenCL library creates task automaticity generated from the description of the user program. The OpenCL library gives scheduling information such as execution order about tasks to the scheduler of the embedded OS. The embedded OS's scheduler handles events of execution, termination and synchronization of data transfer via interrupts from accelerators. The scheduler determines to execute next task from the scheduling information passed by the OpenCL library. By the schemes to control accelerators are managed in OS, accelerators can be worked efficiently. The poster presentation shows the implemented OpenCL and OS framework for "Cube" processor embedded multi-core accelerator. Ryuichi Sakamoto, Mikiko Sato, Yusuke Koizumi, Hideharu Amano, Mitaro Namiki |
RTCSA | 4 |
| 2011 | Geyser-2: The second prototype CPU with fine-grained run-time power gatingabstractGeyser-2 is the second prototype MIPS CPU which provides a fine-grained run-time power gating (PG) controlled by instructions. Geyser-l, the first prototype only provides the fine-grained run-time PG core. Although it demonstrated the leakage power reduction on a real chip, the operational frequency is limited at 60MHz because of the limitation of the I/O speed. Geyser-2 with cache and TLB mechanism is implemented to show (1) run-time PG works at least with 200MHz which is commonly used clock for embedded systems, and (2) it is also efficient on the environment with real application programs with an operating system. Daisuke Ikebuchi, Yoshiki Saito, M. Kamata, Naomi Seki, Yu Kojima, Hideharu Amano, Satoshi Koyama, Tatsunori Hashida, Y. Umahashi, D. Masuda, Kimiyoshi Usami, Mitaro Namiki, Seidai Takeda, Hiroshi Nakamura, Masaaki Kondo |
ASP-DAC | 7 |
| 2011 | The realtime image processing demonstration with CMA-1: An ultra low-power reconfigurable acceleratorabstractCMA (Cool Mega Array) is an ultra low-power reconfigurable accelerator with a large PE (Processing Element) array consisting of combinational circuits. Although the configuration is static during execution, various types of application can be implemented by using the versatile data manupilation instructions of the attached micro-controller. By using a real CMA-1 chip, we will demonstrate that CMA-1 can process various image processing with extremely low power only required for the computation. Kazuei Hironaka, Nobuaki Ozaki, Hideharu Amano |
FPT | 3 |
| 2011 | Reducing power for dynamically reconfigurable processor array by reducing number of reconfigurationsabstractA power-consumption-centric assignment algorithm called partially fixed configuration mapping (PFCM) is proposed for multi-context dynamically reconfigurable processors. By assigning the same operations into the same PE (processing element) as many as possible, the amount of changing configuration data for dynamic reconfiguration can be reduced, resulting in the redundant power consumed for changing the configuration also being reduced. The proposed algorithm was implemented in a compiler for a dynamically reconfigurable processor for research. Evaluation results showed that the consumed power was reduced by 10% on average without increasing the execution time. Masayuki Kimura, Kazuei Hironaka, Hideharu Amano |
FPT | 3 |
| 2011 | Cool Mega-Array: A highly energy efficient reconfigurable acceleratorabstractA highly energy efficient reconfigurable accelerator called CMA (Cool Mega-Array) is proposed. It consists of a large Processing Element (PE) array without memory elements for maintain result of ALU and configuration data, a small simple programmable micro controller for data management, and the data memory. Unlike traditional coarse grained reconfigurable processors, the power consumption for hardware context switching, storing intermediate data in registers, and clock distribution for them are eliminated from PE array which occupies large area of a chip. Configuration registers are collected to small area of micro controller. The data flow graph mapped on the PE array is static during execution. Various application programs can be implemented by making the best use of flexible data management instructions with the micro controller. When the delay time in the PE array is longer than the data handling time with the micro controller, the supply voltage for the PE array is scaled to reduce the power consumption without degrading the performance. In the opposite case, wave pipelining is applied to enhance PE array performance. A prototype chip CMA-1 with 8 × 8 PE array with 24-bit data width was fabricated in 2.1 × 4.2mm265-nm CMOS technology, and achieves 2.4-GOPS/11.2-mW sustained performance. This energy efficiency is comparable to that of the most energy efficient accelerators that have been reported. Nobuaki Ozaki, Yoshihiro Yasuda, Yoshiki Saito, Daisuke Ikebuchi, Masayuki Kimura, Hideharu Amano, Hiroshi Nakamura, Kimiyoshi Usami, Mitaro Namiki, Masaaki Kondo |
FPT | 6 |
| 2011 | On-chip detection methodology for break-even time of power gated function units
Kimiyoshi Usami, Yuya Goto, Kensaku Matsunaga, Satoshi Koyama, Daisuke Ikebuchi, Hideharu Amano, Hiroshi Nakamura |
ISLPED | 6 |
| 2011 | A vertical bubble flow network using inductive-coupling for 3-D CMPsabstractA wireless 3-D NoC architecture for CMPs, in which the number of processor and cache chips stacked in a package can be changed after the chip fabrication, is proposed by using the inductive coupling technology that can connect more than two known-good-dies without wire connections. Each chip has data transceivers for uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a single vertical ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring network, we use the bubble flow control which is more flexible and efficient compared to the conventional VC-based deadlock avoidance. We implemented a real 3-D chip that has on-chip routers and inductive-coupling data transceivers using a 65nm process in order to show the feasibility of our proposal. The vertical bubble flow control is compared with the conventional VC-based approach and vertical bus in terms of the throughput, hardware amount, and application performance using a full system CMP simulator. The results show that the proposed vertical bubble flow network outperforms the VC-based approach by 7.9%-12.5% with a 33.5% smaller router area. Hiroki Matsutani, Yasuhiro Take, Daisuke Sasaki, Masayuki Kimura, Yuki Ono, Yukinori Nishiyama, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
NOCS | 9 |
| 2011 | A Dynamic Link-Width Optimization for Network-on-ChipabstractNetwork-on-Chip (NoC) is considered to be a promising approach to implement many-core systems and a large number of on-chip router optimization studies have been proposed. In this paper, we propose to dynamically adjust link-width of each port on a router optimized to spatially biased traffic. Different from the previous No Coptimization approaches, in which the optimization is almost performed in the NoC design step, the proposed method achieves a dynamical link-width optimization at run-time. Daihan Wang, Michihiro Koibuchi, Tomohiro Yoneda, Hiroki Matsutani, Hideharu Amano |
RTCSA (2) | 5 |
| 2011 | An analytical network performance model for SIMD processor CSX600 interconnects
Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Kenichi Miura, Hideharu Amano |
J. Syst. Archit. | 5 |
| 2011 | Prediction Router: A Low-Latency On-Chip Router Architecture with Multiple PredictorsabstractMulti and many-core applications are sensitive to interprocessor communication latencies, suggesting the need for low-latency on-chip networks. We propose a low-latency router architecture that predicts the output channel to be used by the next packet transfer and speculatively completes the switch arbitration to reduce communication latency. The packets coming into the prediction routers are transferred without waiting for the routing computation and switch arbitration if the prediction hits. Thus, the primary concern for reducing communication latency is the hit rates of the prediction algorithms, which vary based on network environments, such as the network topology, routing algorithm, and traffic pattern. Although typical low-latency routers that skip one or more pipeline stages use a bypass data path that is based on a static or single bypassing policy (e.g., accelerating the packets moving in the same dimension), our prediction router architecture predictively forwards packets based on the prediction algorithm selected from among several candidates in response to the network environment. We analyze the prediction hit rates of five prediction algorithms on meshes, tori, fat trees, and Spidergons. Then, we present four case studies, each of which assumes different many-core architectures. We implemented the prediction routers for each case study by using a 45 nm CMOS process, and evaluated them in terms of the prediction hit rate, zero-load latency, hardware amount, and energy consumption. A typical prediction router with two or three predictors shows that although the area and energy are increased by 4.8-12.0 percent and 5.3 percent, respectively, up to 89.8 percent of the prediction hit rate is achieved in real applications, which provides favorable trade-offs between modest hardware/energy overheads and significant latency saving. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tsutomu Yoshinaga |
IEEE Trans. Computers | 3 |
| 2011 | Performance, Area, and Power Evaluations of Ultrafine-Grained Run-Time Power-Gating Routers for CMPsabstractThis paper proposes the ultrafine-grained run-time power gating of on-chip routers, in which the power supply to each router component (e.g., virtual-channel buffer, virtual-channel multiplexer, and crossbar multiplexer and output latch) can be individually controlled based on the applied workload. Since only the router components that are transferring a packet are activated, the leakage power of the on-chip network can be reduced to a near-optimal level. However, such techniques inherently increase the communication latency and degrade the application performance, since a certain amount of wakeup latency is required to activate the sleeping components. To mitigate this wakeup latency, an early wakeup method that can preliminarily detect the next packet arrival and activate the corresponding components is essential. We designed and implemented an ultrafine-grained power-gating router using a commercial 65 nm process. We propose four early wakeup methods and combine them with the power-gating router. The proposed router with the early wakeup methods is evaluated in terms of its application performance, area overhead, and leakage power reduction taking into account the on/off energy overhead. The simulation results showed that it reduces the leakage power by 54.4-59.9% on average even when the application programs are fully running, at the expense of 4.6% of the area and 0.7-3.7% of the performance overheads when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2011 | A Switch-Tagged Routing Methodology for PC Clusters with VLAN EthernetabstractEthernet has been used for connecting hosts in PC clusters, besides its use in local area networks. Although a layer-2 Ethernet topology is limited to a tree structure because of the need to avoid broadcast storms and deadlocks of frames, various deadlock-free routing algorithms on topologies that include loops suitable for parallel processing can be employed by the application of IEEE 802.1Q VLAN technology. However, the MPI communication libraries used in current PC clusters do not always support tagged VLAN technology; therefore, at present, the design of VLAN-based Ethernet cannot be applied to such PC clusters. In this study, we propose a switch-tagged routing methodology in order to implement various deadlock-free routing algorithms on such PC clusters by using at most the same number of VLANs as the degree of a switch. Since the MPI communication libraries do not need to perform VLAN operations, the proposed methodology has advantages in both simple host configuration and high portability. In addition, when it is used with on/off and multispeed link regulation, the power consumption of Ethernet switches can be reduced. Evaluation results using NAS parallel benchmarks showed that the performance of the topologies that include loops using the proposed methodology was comparable to that of an ideal one-switch (full crossbar) network, and the torus topology in particular had up to a 27 percent performance improvement compared with a tree topology with link aggregation. Michihiro Koibuchi, Tomohiro Otsuka, Tomohiro Kudoh, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2010 | Geyser-1: a MIPS R3000 CPU core with fine-grained run-time power gatingabstractGeyser-1 is a MIPS CPU which provides a fine-grained run-time power gating (PG) controlled by instructions. Unlike traditional PGs, it uses special standard cells in which the virtual ground (VGND) is separated from the real ground, and a certain number of the sleep transistors are inserted for quick power shut-down and wake-up. In Geyser-1, the fine-grained run-time PG is applied to computational modules in the execution stage. The power shut-down and wakeup are controlled with architectural and software level. This implementation is the first available CPU with this type of run-time PG technique. Geyser-1 has both time and spatial fine-grained PG and works well with a real chip. Daisuke Ikebuchi, Naomi Seki, Yu Kojima, M. Kamata, Hideharu Amano, Toshiaki Shirai, Satoshi Koyama, Tatsunori Hashida, Y. Umahashi, Hiroki Masuda, Kimiyoshi Usami, Seidai Takeda, Hiroshi Nakamura, Mitaro Namiki, Masaaki Kondo |
ASP-DAC | 6 |
| 2010 | MuCCRA-3: a low power dynamically reconfigurable processor arrayabstractMuCCRA-3 is a low power coarse-grained Dynamically Reconfigurable Processor Array (DRPA) for a flexible off-loading engine in various SoC (System-on-a-Chip). Similar to the other DRPAs, it has an array of processing elements (PEs), a simple coarse-grained processor, consisting of an ALU and a register file, and dynamic reconfiguration of the array enables time-multiplexed execution. DRPAs including MuCCRA-3 provide multiple sets of configuration data called hardware contexts, and switch them in a clock cycle. For low power computation, the PE array structure of MuCCRA-3 is optimized according to the evaluation results of previous prototypes, MuCCRA-1 and 2, and was implemented with 65 nm low power CMOS process from Fujitsu. By using a real chip, the power consumption and performance are evaluated. The evaluation results suggest that MuCCRA-3 works with extremely low power: 10 mW-13 mW. Yoshiki Saito, Toru Sano, Masaru Kato, Vasutan Tunbunheng, Yoshihiro Yasuda, Masayuki Kimura, Hideharu Amano |
ASP-DAC | 7 |
| 2010 | Reducing power consumption for Dynamically Reconfigurable Processor Array with Partially Fixed Configuration MappingabstractThe Partially Fixed Configuration Mapping (PFCM) is a context mapping technique for Dynamically Reconfigurable Processor Array (DRPA) focusing on reducing the power consumption. It assigns operations into Processing Elements (PEs) so as to keep the configuration of the previous context as possible. It reduces the changing part of the datapath structure on the PE array as well as its switching frequency. Preliminary evaluation results show that it can reduce the computing power by 6.7% - 11.3%. The demonstration shows the power reduction directly by using the real chip MuCCRA-3, a prototype of DRPA executing signal processing applications with and without applying PFCM. The design environment for using PFCM is also exhibited. Kazuei Hironaka, Masayuki Kimura, Yoshiki Saito, Toru Sano, Masaru Kato, Vasutan Tunbunheng, Yoshihiro Yasuda, Hideharu Amano |
FPT | 8 |
| 2010 | A datapath classification method for FPGA-based scientific application accelerator systemsabstractResource reduction design techniques play an important role to implement large-scale FPGA-based accelerator systems in floating point applications since available resources on FPGAs are limited. This paper proposes a dataflow graph classification method which makes groups of graphs based on their similarity in order to bring out efficient graph combining. Aiming at finding effective parameters for the k-means algorithm, various parameter combinations are evaluated and compared in terms of resource reduction effects and performance. The experimental results using an FPGA-based biochemical simulator reveal that the graph clustering that uses information on the maximum common subgraphs achieve 73.3% of resource reduction rate while alleviating the performance degradation. Yui Ogawa, Tomonori Ooya, Yasunori Osana, Masato Yoshimi, Yuri Nishikawa, Akira Funahashi, Noriko Hiroi, Hideharu Amano, Yuichiro Shibata, Kiyoshi Oguri |
FPT | 8 |
| 2010 | Wire congestion aware synthesis for a dynamically reconfigurable processorabstractThis paper presents two iterative synthesis techniques between a high-level synthesizer (HLS) and the place and route tool to shorten the prolonged wire delay for a dynamically reconfigurable processor. At first, we use feedback wire delays for each context to a scheduler in the HLS. The experimental results showed that a critical-path delay was shorten 21% on average for applications with timing closure problems. Second, we skip the routing and estimate wire delays based on their congestions. The synthesis time was cut by 1/3 with only four points down on delay improvement rate. Takao Toi, Takumi Okamoto, Toru Awashima, Kazutoshi Wakabayashi, Hideharu Amano |
FPT | 5 |
| 2010 | Stabilizing Path Modification of Power-Aware On/Off Interconnection NetworksabstractPower saving is required for interconnects of modern PC clusters as well as the performance improvement. To reduce the power consumption of switches with maintaining the performance, on/off link regulations that activate and deactivate the links based on the traffic load have been widely developed in interconnection networks. Depending on which operation is selected, link activation or deactivation, the available network resources are changed, thus requiring paths to be reconfigured. To maintain deadlock freedom of packet transfers, connectivity, and performance during the path changes, we propose to apply dynamic reconfiguration techniques that process packet transfer uninterruptedly to power-aware on/off interconnection networks. The dynamic network reconfiguration techniques stabilize the update of paths that are quite crucial to use power-aware on/off link techniques in interconnects of PC clusters. We investigate the performance and behavior of network reconfiguration technique as soon as the link activation or deactivation occurs. Evaluation results show that the simple dynamic reconfiguration techniques slightly reduce the peak packet latency and reconfiguration time of the change compared with existing static reconfiguration in on/off interconnection networks. A reconfiguration technique called Double Scheme reduces by up to 95% the peak packet latency caused by the on/off link operation. José Miguel Montañana, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
NAS | 4 |
| 2010 | A Deadlock-Free Non-minimal Fully Adaptive Routing Using Virtual Cut-Through SwitchingabstractSystem area networks (SANs), which usually employ virtual cut-through switching, have been used to connect hosts in modern PC clusters and massively parallel computers. In this paper, we propose a non-minimal fully adaptive deadlock-free routing mechanism for virtual-cut-through networks called “Semi-deflection”. Semi-deflection routing guarantees deadlock-free packet transfer without use of virtual channels by allowing non-blocking transfer between specific pairs of routers. As the result of throughput evaluation, Semi-deflection routing improved throughput by up to 26 percent compared with that of north-last turn model, which is a typical adaptive routing, and also reduced latency. Yuri Nishikawa, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
NAS | 4 |
| 2010 | Ultra Fine-Grained Run-Time Power Gating of On-chip Routers for CMPsabstractThis paper proposes an ultra fine-grained run-time power gating of on-chip router, in which power supply to each router component (e.g., VC queue, crossbar MUX, and output latch) can be individually controlled in response to the applied workload. As only the router components which are just transferring a packet are activated, the leakage power of the on-chip network can be reduced to the near-optimal level. However, a certain amount of wakeup latency is required to activate the sleeping components, and the application performance will be degraded. In this paper, we estimate the wakeup latency for each component based on circuit simulations using a 65 nm process. Then we propose four early wakeup methods to overcome the wakeup latency. The proposed router with the early wakeup methods is evaluated in terms of the application performance, area, and leakage power. As a result, it reduces the leakage power by 78.9%, at the expense of the 4.3% area and 4.0% performance when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
NOCS | 6 |
| 2009 | Modularizing flux limiter functions for a Computational Fluid Dynamics accelerator on FPGAsabstractCFD is taken notice as a cost effective design tool for aircraft components. UPACS is a convenient CFD platform, since it supports a large degree of versatility using various kinds of solvers. However, its major drawback is a long simulation time. We have developed a UPACS accelerator named FLOPS-2D with multiple FPGA boards, and implemented some core functions. Here, by using flexibility of FPGAs, the selectable functions are proposed. All possible flux limiter functions in MUSCLs are implemented independently, and only required functions are selected and implemented with other modules to form the optimal structure. All implemented functions achieved at least 24 times higher performance than that with the Core 2 Duo, and configurable solvers using a desired flux limiter function and the number of arithmetic pipelines are developed. Kenta Inakagata, Hirokazu Morishita, Yasunori Osana, Naoyuki Fujita, Hideharu Amano |
FPL | 5 |
| 2009 | Configuring area and performance: Empirical evaluation on an FPGA-based biochemical simulatorabstractOne of the obvious advantages of FPGA-based reconfigurable computing is customizability of a tradeoff point between performance and hardware costs. However, this tradeoff has rarely been discussed in a whole application level, which is the most important view for application users. This paper presents empirical evaluation of a hardware module sharing technique which can shift a tradeoff point of area and performance on an FPGA-based biochemical simulator. The biochemical simulation results are discussed in terms of hardware costs, simulation throughput, parallelism extracted in simulation hardware, and data transfer overheads. Tomonori Ooya, Hideki Yamada, Tomoya Ishimori, Yuichiro Shibata, Yasunori Osana, Kiyoshi Oguri, Masato Yoshimi, Yuri Nishikawa, Akira Funahashi, Noriko Hiroi, Hideharu Amano |
FPL | 11 |
| 2009 | MuCCRA-Cube: A 3D dynamically reconfigurable processor with inductive-coupling linkabstractMuCCRA-Cube is a scalable three dimensional dynamically reconfigurable processor. By stacking multiple dies connected with inductive-coupling links, the number of PE array can be increased so that the required performance is achieved. A prototype chip with 90nm CMOS process consisting of four dies each of which has a 4 × 4 PE array was implemented. The vertical link achieved 7.2Gb/s/chip, and the average execution time is reduced to 31% compared to that using a single chip. Shotaro Saito, Yoshinori Kohama, Yasufumi Sugimori, Yohei Hasegawa, Hiroki Matsutani, Toru Sano, Kazutaka Kasuga, Yoichi Yoshida, Kiichi Niitsu, Noriyuki Miura, Tadahiro Kuroda, Hideharu Amano |
FPL | 12 |
| 2009 | Fine Grain Partial Reconfiguration for energy saving in Dynamically Reconfigurable ProcessorsabstractBased on the power consumption analysis of a real dynamically reconfigurable processor array (DRPA) prototype MuCCRA-3, it appears that the key of power saving is keeping the datapath on the processing element (PE) array as possible. Fine grain partial reconfiguration (FGPR) is a simple technique to minimize the change of configuration code in a hardware context switching. In FGPR, a configuration code is divided into several components and only the configuration data for the required components are changed. Evaluation results demonstrate that about 15% of the power consumption is reduced with only 0.7% hardware overhead. The total amount of configuration data and its loading time can be also reduced by 37% in average. Toru Sano, Yoshiki Saito, Masaru Kato, Hideharu Amano |
FPL | 4 |
| 2009 | Leakage power reduction for coarse-grained dynamically reconfigurable processor arrays using Dual Vt cellsabstractOne of benefit of coarse-grained dynamically reconfigurable processor arrays (DRPAs) is their low dynamic power consumption by operating a number of processing element (PE) in parallel with a low frequency clock. However, in the future advanced process, the leakage power will occupy a considerable part of the total power consumption, and it may degrade the advantage of DRPAs. In order to reduce the leakage power of DRPA without severe performance degradation, eight designs (Mult, Sw, MultSw, LowHalf, 1Row, ColHalf, Sw+Half and Sw+Mult) using Dual-Vt cells are evaluated based on a prototype DRPA called MuCCRA-3T. Evaluation results show that Sw in which Low-Vt cells are only used in switching elements of the array achieved the best power-delay product. If performance of Sw is not enough, Sw+Half in which Low-Vt cells are used for a lower half PEs and all switching elements improves 24% of the leakage power with 5%-14% of extra delay time of the design with all Low-Vt cells. Keiichiro Hirai, Masaru Kato, Yoshiki Saito, Hideharu Amano |
FPT | 4 |
| 2009 | Prediction router: Yet another low latency on-chip router architectureabstractNetwork-on-Chips (NoCs) are quite latency sensitive, since their communication latency strongly affects the application performance on recent many-core architectures. To reduce the communication latency, we propose a low-latency router architecture that predicts an output channel being used by the next packet transfer and speculatively completes the switch arbitration. In the prediction routers, incoming packets are transferred without waiting the routing computation and switch arbitration if the prediction hits. Thus, the primary concern for reducing the communication latency is the hit rates of prediction algorithms, which vary from the network environments, such as the network topology, routing algorithm, and traffic pattern. Although typical low-latency routers that speculatively skip one or more pipeline stages use a bypass datapath for specific packet transfers (e.g., packets moving on the same dimension), our prediction router predictively forwards packets based on a prediction algorithm selected from several candidates in response to the network environments. In this paper, we analyze the prediction hit rates of six prediction algorithms on meshes, tori, and fat trees. Then we provide three case studies, each of which assumes different many-core architecture. We have implemented a prediction router for each case study by using a 65 nm CMOS process, and evaluated them in terms of the prediction hit rate, zero load latency, hardware amount, and energy consumption. The results show that although the area and energy are increased by 6.4-15.9% and 8.0-9.5% respectively, up to 89.8% of the prediction hit rate is achieved in real applications, which provide favorable trade-offs between the modest hardware/energy overheads and the latency saving. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tsutomu Yoshinaga |
HPCA | 3 |
| 2009 | An on/off link activation method for low-power ethernet in PC clustersabstractThe power consumption of interconnects is increased as the link bandwidth is improved in PC clusters. In this paper, we propose an on/off link activation method that uses the static analysis of the traffic in order to reduce the power consumption of Ethernet switches while maintaining the performance of PC clusters. When a link whose utilization is low is deactivated, the proposed method renews the VLAN-based paths that avoid it without creating broadcast storms. Since each host does not need to process VLAN tags, the proposed method has advantages in both simple host configuration and high portability. Evaluation results using NAS Parallel Benchmarks show that the proposed method reduces the power consumption of switches by up to 37% without performance degradation. Michihiro Koibuchi, Tomohiro Otsuka, Hiroki Matsutani, Hideharu Amano |
IPDPS | 4 |
| 2009 | Evaluation of a multicore reconfigurable architecture with variable core sizesabstractA multicore architecture for processors has emerged as a dominant trend in the chip making industry. As reconfigurable devices gradually prove their capability in improving computation power while preserving flexibility, we are examining a multicore reconfigurable architecture consisting of multiple reconfigurable computational cores connected by an interconnection network. Using an NEC Electronics' DRP-1 as a core for the multicore architecture, a comparison with a tile-based architecture is performed by implementing several streaming applications with various versions. By using wider communication channels and assigning more resources for computations, it is possible to improve throughput over implementations for the tile-based architecture. Another evaluation with different core sizes is examined in order to see the effect of core size in a homogeneous multicore system on performance and internal fragmentation. Evaluation results show that the size of core is a trade-off between throughput and resource usage. Vu Manh Tuan, Naohiro Katsura, Hiroki Matsutani, Hideharu Amano |
IPDPS | 4 |
| 2009 | Performance Analysis of ClearSpeed's CSX600 InterconnectsabstractClearSpeed's CSX600 that consists of 96 Processing Elements (PEs) employs a one-dimensional array topology for a simple SIMD processing. To clearly show the performance factors and practical issues of NoCs in an existing modern many-core SIMD system, this paper measures and analyzes NoCs of CSX600 called Swazzle and ClearConnect. Evaluation and analysis results show that the sending and receiving overheads are the major limitation factors to the effective network bandwidth. We found that (1) the number of used PEs, (2) the size of transferred data, and (3) data alignment of a shared memory are three main points to make the best use of bandwidth. In addition, we estimated the best- and worst-case latencies of data transfers in parallel applications. Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Akihiro Shitara, Kenichi Miura, Hideharu Amano |
ISPA | 6 |
| 2009 | Fat H-Tree: A Cost-Efficient Tree-Based On-Chip NetworkabstractThe topological explorations of on-chip networks are important for efficiently using their enormous wire resources for low-latency and high-throughput communications using a modest silicon budget. In this paper, we propose a novel tree-based interconnection network called Fat H-Tree that meets these requirements. A Fat H-Tree provides a torus structure by combining two folded H-Tree networks and is an attractive alternative to tree-based networks such as the Fat Trees in a microarchitecture domain. We introduce its chip layout schemes based on a folding technique for 2D and 3D ICs. Three deadlock-free routing schemes are proposed for Fat H-Tree. We evaluate the performance of Fat H-Tree and other tree-based networks using real application traces. In addition, the network logic area, wire resource, and energy consumption of Fat H-Tree are compared with other topologies, based on a typical implementation of on-chip routers synthesized with a 90-nm standard cell library. The results show that (1) a Fat H-Tree outperforms a Fat Tree with two upward and four downward connections in terms of the throughput and average hop count, (2) a Fat H-Tree requires 19.8 percent-27.8 percent smaller network logic area than the Fat Tree, (3) a Fat H-Tree consumes slightly less energy than the Fat Tree does, and (4) a Fat H-Tree uses slightly more wire resources than the Fat Tree, but the current process technology can provide sufficient wire resources for implementing Fat-H-Tree-based on-chip networks. Hiroki Matsutani, Michihiro Koibuchi, Yutaka Yamada, D. Frank Hsu, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2009 | Guest Editors' Introduction: ICFPT 2007abstractNo abstract available. Hideharu Amano, Tadao Nakamura |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2008 | Run-time power gating of on-chip routers using look-ahead routingabstractSince on-chip routers in Network-on-Chips play a key role in on-chip communication between cores, they should be always preparing for packet injections even if a part of cores are in standby mode, resulting in a larger standby power of routers compared with cores. The run-time power gating of individual channels in a router is one of attractive solutions to reduce the standby power of chip without affecting the on-chip communication. However, a state transition between sleep and active mode incurs the performance penalty, and turning a power switch on or off dissipates the overhead energy, which means a short-term sleep adversely increases the power consumption. In this paper, we propose a sleep control method based on look-ahead routing that detects the arrival of packets two hops ahead, so as to hide the wake-up delay and reduce the short-term sleeps of channels. Simulation results using real application traces show that the proposed method conceals the wake-up delay of less than five cycles, and more leakage power can be saved compared with the original naive method. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Daihan Wang |
ASP-DAC | 3 |
| 2008 | Power reduction techniques for Dynamically Reconfigurable Processor ArraysabstractThe power consumption of Dynamically Reconfigurable Processing Array (DRPA) is quantitatively analyzed by using a real chip layout and applications taking into account the reconfiguration power. Evaluation result shows that processing power for PEs is dominant and reconfiguration power is about 20.7% of the total dynamic power consumption. Based on the above evaluation results, we proposed two dynamic power reduction techniques: functional unit-level operand isolation and selective context fetch. Evaluation results demonstrate that the functional unit-level operand isolation can reduce up to 20.8% of the dynamic power with only 2.2% area overhead. On the selective context fetch, the power reduction is limited by the increasing of the additional hardware. Takashi Nishimura, Keiichiro Hirai, Yoshiki Saito, Takuro Nakamura, Yohei Hasegawa, Satoshi Tsutsumi, Vasutan Tunbunheng, Hideharu Amano |
FPL | 8 |
| 2008 | Instruction buffer mode for multi-context Dynamically Reconfigurable ProcessorsabstractIn multi-context Dynamically Reconfigurable Processor Array (DRPA), the required number of contexts is often increased by those with low resource usage. In order to execute such contexts without wasting a context memory, we propose a new execution mode called instruction buffer mode in addition to the normal multi-context mode. In this mode, a configuration code from the central configuration memory is stored in the instruction buffer and executed directly. Furthermore, by exploiting a multicast method, a single configuration code loaded to the buffer can be executed by multiple processing elements in a SIMD fashion. We also investigate a mode selection policy based on simple formulas. From the result of implementation and evaluation by using a prototype DRPA called MuCCRA-1, it appears that the total execution time is reduced 12% by using the instruction buffer mode, while 12% of the semiconductor area is increased. Toru Sano, Masaru Kato, Satoshi Tsutsumi, Yohei Hasegawa, Hideharu Amano |
FPL | 5 |
| 2008 | A link removal methodology for Networks-on-Chip on reconfigurable systemsabstractWhile the regular 2-D mesh topology has been utilized for most of Network-on-Chips (NoCs) on FPGAs, spatially biased traffic in some applications make some customization method feasible. A link removal strategy that customizes the router in NoC is proposed for reconfigurable systems in order to minimize required hardware amount. Based on the pre-analyzed traffic information, links on which the communication amount is small are removed to reduce the hardware cost with enough performance being kept. Two policies are proposed to avoid deadlocks and better performance can be achieved compared with up*/down* routing on the irregular topology with links removed. In the image recognition application susan, the proposed method can save 30% of the hardware amount without performance degradation. Daihan Wang, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
FPL | 3 |
| 2008 | Practical implementation of a network-based stochastic biochemical simulation system on an FPGAabstractStochastic simulation of biochemical reaction networks are widely focused by life scientists to represent stochastic behaviors in cellular processes. Stochastic algorithm has loop-and thread-level parallelism, and it is suitable for running on application specific hardware to achieve high performance with low cost. We have implemented and evaluated the FPGA-based stochastic simulator according to theoretical research of the algorithm. This paper introduces an improved architecture for accelerating a stochastic simulation algorithm called the Next Reaction Method. This new architecture has scalability to various size of FPGA. As the result with a middle-range FPGA, 5.38 times higher throughput was obtained compared to software running on a Core 2 Quad Q6600 2.40GHz. Masato Yoshimi, Yuri Nishikawa, Yasunori Osana, Akira Funahashi, Yuichiro Shibata, Hideki Yamada, Noriko Hiroi, Hiroaki Kitano, Hideharu Amano |
FPL | 9 |
| 2008 | Exploiting memory hierarchy for a Computational Fluid Dynamics accelerator on FPGAsabstractComputational Fluid Dynamics (CFD) is an important tool for aeronautical engineers. Instead of expensive super-computers or clusters, using custom pipelines built on FPGAs is expected to be a cost effective solution to accelerate CFD. The problem is that to keep the pipeline busy is difficult because of the memory bandwidth. To deal with this problem, an effective memory access method using Block-RAMs is implemented based on a careful survey about memory access pattern. This work is targetting on two major subroutines in UPACS, a CFD software package. As a result, the amount of data transfer is reduced about 40%. This shows 46–170fold speed-up is expected by several Virtex-4 FPGAs compared to Itanium2 processor. Hirokazu Morishita, Yasunori Osana, Naoyuki Fujita, Hideharu Amano |
FPT | 4 |
| 2008 | Exploring the optimal size for multicasting configuration data of dynamically reconfigurable processorsabstractThe configuration data transfer time of a dynamically reconfigurable processor often bottlenecks the hardware context switching time and degrades its computation performance. In order to reduce data transferring time from a central memory to hardware context memory modules in all processing elements (PEs) and switching elements (SEs), a multicasting mechanism called RoMultiC (row-muticast configuration) was proposed. However, the original Ro-MultiC used the whole PE or SE as a unit of multicast, the reduction of transfers is limited. Here, the trade-off between the granularity of multicast and hardware increase are evaluated, and the best way to make the multicast bit-map is explored. Evaluation results show that time for transfer is reduced up to 42% compared with the original RoMultiC with only 2% hardware overhead. Takuro Nakamura, Toru Sano, Yohei Hasegawa, Satoshi Tsutsumi, Vasutan Tunbunheng, Hideharu Amano |
FPT | 6 |
| 2008 | Leakage power reduction for coarse grained dynamically reconfigurable processor arrays with fine grained Power Gating techniqueabstractOne of the benefits of coarse grained dynamically reconfigurable processor array(DRPA) is its low dynamic power consumption by operating a number of processing elements(PE) in parallel with low clock frequency. However, in the future advanced processes, leakage power will occupy a considerable part of the total power consumption, and it may degrade the advantage of DRPAs. In order to reduce the leakage power, a fine grained Power Gating(PG) is applied to a DRPA, MuCCRA-2.32b, and leakage power and area overhead are measured. We evaluated the effect of two control modes; Pair and Unit Individual based on layout design and real applications. It appears that by applying PG for ALUs and SMUs in PEs individually, 48% of leakage power can be reduced with 9.0% of area overhead. Yoshiki Saito, Tomoaki Shirai, Takuro Nakamura, Takashi Nishimura, Yohei Hasegawa, Satoshi Tsutsumi, Toshihiro Kashima, Mitsutaka Nakata, Seidai Takeda, Kimiyoshi Usami, Hideharu Amano |
FPT | 11 |
| 2008 | A fine-grain dynamic sleep control scheme in MIPS R3000abstractA fine-grain dynamic power gating is proposed for saving the leakage power in MIPS R3000 by sleep control and applied to a processor pipeline. An execution unit is divided into four small units: multiplier, divider, shifter and other (CLU). The power of each unit is cut off dynamically, based on the operation. We tape-outed the prototype chip Geyser-0, which provides an R3000 Core with the power reduction technique, 16 KB caches and translation lookaside buffer (TLB) using 90 nm CMOS technology. The evaluation results of four benchmark programs for embedded applications show that 47% of the leakage power is reduced on average with 41% area overhead. Naomi Seki, Jo Kei, Daisuke Ikebuchi, Yu Kojima, Yohei Hasegawa, Hideharu Amano, Toshihiro Kashima, Seidai Takeda, Toshiaki Shirai, Mitsutaka Nakata, Kimiyoshi Usami, Tetsuya Sunata, Jun Kanai, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ICCD | 7 |
| 2008 | A Lightweight Fault-Tolerant Mechanism for Network-on-Chip
Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano, Timothy M. Pinkston |
NOCS | 3 |
| 2008 | Adding Slow-Silent Virtual Channels for Low-Power On-Chip Networks
Hiroki Matsutani, Michihiro Koibuchi, Daihan Wang, Hideharu Amano |
NOCS | 4 |
| 2007 | Design Methodology and Trade-offs Analysis for Parameterized Dynamically Reconfigurable Processor ArraysabstractIn this paper, we propose a Dynamically Reconfigurable Processor Array (DRPA) generator which can generate various types of DRPAs. Our target DRPA architecture is fully parameterized. By specifying architectural parameters, it can automatically generate RTL model, simulation environment, and finally chip layout. In our DRPA generator, although the fundamental design of a processing element (PE) and an inter-PE connection is fixed, the array size, PE granularity, and connection flexibilities of intra/inter PE are selectable. In this paper, we have generated various types of DRPAs and evaluated semiconductor area and speed by using the ASPLA/STARC 90-nm CMOS technology. From evaluation results, fundamental trade-offs between architectural parameters and area/delay are analyzed. Yohei Hasegawa, Hideharu Amano |
FPL | 2 |
| 2007 | A High Speed License Plate Recognition System on an FPGAabstractA high speed FPGA off-loading engine for detecting the license plate itself in order to avoid the traffic accident is proposed. A complicated algorithm is written in Handel-C, and parallel processing is explicitly utilized in every level of implementation; an input image is segmented into 16 areas, and each area is processed in parallel by a multiple calculation unit executing pipeline processing and a distributed memory module. A prototype circuit implemented on a general purpose FPGA board achieved 4.16 times performance as software execution on a Pentium-III desktop PC. The highest performance in literature; 100 frames per second; can be achieved. Takamasa Kanamori, Hideharu Amano, Masatoshi Arai, Daisuke Konno, Tomomichi Nanba, Yoshiaki Ajioka |
FPL | 2 |
| 2007 | A Temporal Correlation Based Port Combination Methodology for Networks-on-chip on Reconfigurable SystemsabstractA temporal correlation based port combination algorithm that customizes the router design in Network-on-Chip (NoC) is proposed for reconfigurable systems in order to minimize required hardware amount. Given the traffic characteristics of the target application and the expected hardware amount reduction rate, the algorithm automatically makes the port combination plan for the networks. Since the port combination technique has the advantage of almost keeping the topology, it does not affect the design of the other layers, such as task mapping and scheduling. The algorithm shows much better efficiency than the algorithm without temporal correlation. For the multimedia stream processing application, the algorithm can save 55% of the hardware amount without performance degradation, while the non-temporal correlation algorithm suffers from 30% performance loss. Daihan Wang, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
FPL | 4 |
| 2007 | A Combining technique of rate law functions for a cost-effective reconfigurable biological simulatorabstractIn order to simulate large scale biological models with a reconfigurable FPGA-based biochemical simulator system, reduction of required resources are essential. This paper proposes a method which combines common terms in rate law functions appeared in biochemical models and generates a shared hardware module used for numerical integration. In this approach, two functions are combined in a tree structure level, followed by pipeline scheduling and arithmetic module binding. The evaluation result reveals that this approach reduces hardware resources by 31.4% on average at the cost of 14.4% throughput degradation. Hideki Yamada, Naoki Iwanaga, Yuichiro Shibata, Yasunori Osana, Masato Yoshimi, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Hideharu Amano, Akira Funahashi, Noriko Hiroi, Hiroaki Kitano, Kiyoshi Oguri |
FPL | 9 |
| 2007 | FPGA Implementation of a Data-Driven Stochastic Biochemical Simulator with the Next Reaction MethodabstractThis paper introduces a scalable FPGA implementation of a stochastic simulation algorithm (SSA) called the Next Reaction Method. There are some hardware approaches of SSAs that obtained high-throughput on reconfigurable devices such as FPGAs, but these works lacked in scalability. The design of this work can accommodate to the increasing size of target biochemical models, or to make use of increasing capacity of FPGAs. Interconnection network between arithmetic circuits and multiple simulation circuits aims to perform a data-driven multi-threading simulation. Approximately 8 times speedup was obtained compared to an excution on Xeon 2.80GHz. Masato Yoshimi, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Yasunori Osana, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hideki Yamada, Hiroaki Kitano, Hideharu Amano |
FPL | 12 |
| 2007 | Overwrite Configuration Technique in Multicast Configuration Scheme for Dynamically Reconfigurable Processor ArraysabstractA new configuration scheduling algorithm in multicast configuration scheme is proposed and evaluated over reduction ratio of configuration data transfer cycles and power/energy overhead on a coarse-grained dynamically reconfigurable processor array (DRPA). As a case study, the proposed methods are applied to some real applications on a DRPA architecture MuCCRA-1. As a result, we confirmed that the proposed overwrite configuration technique for DRPAs reduced an application configuration cycles 66.5% at maximum compared to one without multicast and 20.2% compared to one without the overwrite configuration. It also decreased the configuration energy consumption 18.7% at maximum beyond the overhead of memory overwrites. Satoshi Tsutsumi, Vasutan Tunbunheng, Yohei Hasegawa, Adepu Parimala, Takuro Nakamura, Takashi Nishimura, Hideharu Amano |
FPT | 7 |
| 2007 | A Mapping Method for Multi-Process Execution on Dynamically Reconfigurable ProcessorsabstractThemulti-processexecution in dynamically reconfigurable processors is a technique to enhance throughput by trying to exploit more inherent parallelism of applications. In order to improve the efficiency of themulti-processexecution, the paper proposes a systematic method for mapping an application modeled as a Kahn Process Network onto a dynamically reconfigurable processing array. Using real applications, the impact on the performance from different versions mapped onto the Dynamically Reconfigurable Processor (DRP) is evaluated. Evaluation results show that our proposed mapping algorithm achieves the best performance in terms of the throughput and the execution time. Vu Manh Tuan, Hideharu Amano |
FPT | 2 |
| 2007 | A Framework for Implementing a Network-Based Stochastic Biochemical Simulator on an FPGAabstractThis paper studies several designs of network-based FPGA implementation of a stochastic simulation algorithm called the next reaction method, known for its large number of calculation involved. The procedure is divided into several subdivisions which will be implemented as independent modules, and they are connected with configurable interconnection networks so as to provide high throughput. By performing a multi-threading simulation, 3.6 times speedup was obtained compared with an execution on general purpose processors. Masato Yoshimi, Yuri Nishikawa, Toshinori Kojima, Yasunori Osana, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Hideki Yamada, Hiroaki Kitano, Hideharu Amano |
FPT | 10 |
| 2007 | Tightly-Coupled Multi-Layer Topologies for 3-D NoCsabstractThree-dimensional network-on-chip (3-D NoC) is an emerging research topic exploring the network architecture of 3-D ICs that stack several smaller wafers for reducing wire length and wire delay. Although the network topology of 3-D NoC has been explored for a couple of years, there is still only a narrow range of choices. In this paper, we propose a class of 3-D topologies called Xbar-connected network-on-tiers (XNoTs), which consist of multiple network layers tightly connected via crossbar switches. To make the best use of the short delay and high density of inter-wafer links, XNoTs topologies have crossbar switches that connect different layers and their cores. The planar topology on every layer can be independently customized so as to meet the cost-performance requirements, as far as network connectivity is at least guaranteed with the bottom layer. We also propose their routing algorithm, which guarantees deadlock-freedom by restricting the inter-layer packet transfer from a lower-numbered layer to a higher-numbered layer. Path sets at the bottom layer close to the heat sink of the chip can be selectively employed in order to mitigate the heat-dissipation problem of 3-D ICs. Several forms of XNoTs topologies including meshes, tori, and/or trees are created, and they are evaluated in terms of performance, cost, and energy consumption. As a result, we show that even with the flexibilities mentioned above, XNoTs achieve at least as high throughput as existing 3-D topologies for equivalent chip sizes. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICPP | 3 |
| 2007 | Performance Improvement Methodology for ClearSpeed's CSX600abstractThis paper focuses on a performance of network-on-a- chip (NoC) and I/O of ClearSpeed's CSX600 coprocessor with 96 multithread processing elements. Two versions of the Himeno benchmark were implemented on the CSX600 to evaluate its performance when it encounters frequent memory transfers between shared and local memories, or between local memories. In order to efficiently use the NoC bandwidth, the dataflow was customized to the one- dimensional array structure of CSX600's NoC. The results of evaluation and profiling indicate that the performance was lower than 1/50 of the sustained performance. We show three key points to improve the performance on such a case: 1) exploiting bandwidth between mono and poly memory, 2) further program tuning, and 3) architectural reform. Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Kenichi Miura, Hideharu Amano |
ICPP | 5 |
| 2007 | Performance, Cost, and Energy Evaluation of Fat H-Tree: A Cost-Efficient Tree-Based On-Chip NetworkabstractFat H-Tree is a novel tree-based interconnection network providing a torus structure, which is formed by combining two folded H-Tree networks, and is an attractive alternative to tree-based networks such as Fat Trees in a micro architecture domain. In this paper, we introduce Fat H-Tree and its deadlock-free routing algorithms. The performance of Fat H-Tree is evaluated using real application traces, and the result is compared with those of other tree-based networks. The network logic area and wire resources for Fat H-Tree are computed based on a typical implementation of on-chip routers using a 0.18mum standard cell library. In addition, the energy consumption is estimated based on the gate-level power analysis. The results show that 1) Fat H-Tree outperforms Fat Tree with two upward and four downward connections in terms of throughput and average hop count; 2) Fat H-Tree requires 19.3%-26.4% smaller network logic area compared with the Fat Tree; 3) Fat H-Tree consumes 8.3%-8.6% less energy compared with the Fat Tree due to its short average hop count; 4) Fat H-Tree uses slightly more wire resources compared with the Fat Tree, but the current process technology can provide sufficient wire resources for implementing Fat H-Tree based on-chip networks. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
IPDPS | 3 |
| 2007 | An Effective Design of Deadlock-Free Routing Algorithms Based on 2D Turn Model for Irregular NetworksabstractSystem area networks (SANs), which usually accept arbitrary topologies, have been used to connect hosts in PC clusters. Although deadlock-free routing is often employed for low-latency communications using wormhole or virtual cut-through switching, the interconnection adaptivity introduces difficulties in establishing deadlock-free paths. An up*/down* routing algorithm, which has been widely used to avoid deadlocks in irregular networks, tends to make unbalanced paths as it employs a one-dimensional directed graph. The current study introduces a two-dimensional directed graph on which adaptive routings called left-up first turn (L-turn) routings and right-down last turn (R-turn) routings are proposed to make the paths as uniformly distributed as possible. This scheme guarantees deadlock-freedom because it uses the turn model approach, and the extra degree of freedom in the two-dimensional graph helps to ensure that the prohibited turns are well-distributed. Simulation results show that better throughput and latency results from uniformly distributing the prohibited turns by which the traffic would be more distributed toward the leaf nodes. The L-turn routings, which meet this condition, improve throughput by up to 100 percent compared with two up*/down*-based routings, and also reduce latency Akiya Jouraku, Michihiro Koibuchi, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCsabstractIn this paper, “Martini,” a network interface controller chip for our original network called RHiNET is described. Martini is designed to provide high-bandwidth and low-latency communication with small overhead. To obtain high performance communication, protected user-level zero-copy RDMA communication functions are completely implemented by a hardwired logic. Also, to reduce the communication latency efficiently, we have proposed PIO-based communication mechanisms called “On-the-fly (OTF)” and have implemented them on Martini. The evaluation results show that Martini connected to a 64bit/66MHz PCI-bus achieves 470MByte/s maximum bidirectional bandwidth and 1.74 μsec minimum latency on host-to-host memory copying. Konosuke Watanabe, Tomohiro Otsuka, Junichiro Tsuchiya, Hiroaki Nishi, Junji Yamamoto, Noboru Tanabe, Tomohiro Kudoh, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2006 | A Context Dependent Clock Control Mechanism for Dynamically Reconfigurable ProcessorsabstractDynamically reconfigurable processors improve the area-efficiency by executing a task with multiple hardware contexts. The maximum operational frequency is limited with a context which has the largest delay time, and it causes a certain overhead when each context has various delay time. A context dependent dynamic clock control method, which changes the clock so as to fit the current operational context, is proposed for NEC electronics' DRP-1. A clock generator consisting of a preset-able counter associated with the state transition table for controlling the context switching is proposed. Performance evaluation using several applications reveals that the proposed method improves the performance from 10% to 110% with a small increasing of the power consumption Hideharu Amano, Yohei Hasegawa, Shohei Abe, Kenichiro Ishikawa, Shunsuke Tsutsumi, Shunsuke Kurotaki, Takuro Nakamura, Takashi Nishimura |
FPL | 1 |
| 2006 | Performance Evaluation of an Fpga-Based Biochemical Simulator ReCSipabstractReCSiP is an FPGA-based biochemical simulator to accelerate kinetic simulations of biochemical pathways. Biochemical models are described as a set of ordinary differential equations (ODEs). Each equation in the model is called "rate law function", which represents the velocity of corresponding biochemical reaction mechanism. ReCSiP achieves high-throughput simulation with statically pipelined rate law function modules and numerical integration modules on an FPGA. This paper shows the basic structure of ReCSiP, and results of evaluation in 2 aspects: area and throughput. As the summary of evaluation, 1) about 64% of the total circuit area is occupied by floating-point arithmetic units, and 2) with an XC2VP70, ReCSiP at 90MHz can achieve 20times or more speedup compared to Intel's Pentium4 microprocessor at 3.2GHz Yasunori Osana, Masato Yoshimi, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano |
FPL | 8 |
| 2006 | An FPGA Implementation of High Throughput Stochastic Simulator for Large-Scale Biochemical SystemsabstractStochastic simulation of biochemical systems has become one of major approaches to study life processes as system, yet is a computational challenge to run the simulation due to its vast calculation cost. This paper shows the implementation and evaluation of a stochastic simulation algorithm (SSA) called "first reaction method" on an FPGA-based biochemical simulator. It achieves high throughput by (1) consecutively throwing data into deeply-pipelined floating point arithmetic units, and (2) by distributing multiple simulators for parallel execution. As the result of evaluation on an FPGA-based simulation platform called ReC-SiP2, the simulator outperforms execution on Xeon 2.80 GHz by approximately 80 times, even with large-scale biochemical systems Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Yuri Nishikawa, Toshinori Kojima, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano |
FPL | 11 |
| 2006 | An adaptive Viterbi decoder on the dynamically reconfigurable processorabstractIn order to evaluate practical adaptive computing on dynamically reconfigurable processors, several Viterbi decoders with different constraint variables are implemented on NEC Electronics' DRP-1. By switching designs, its throughput varies from 4.71 Mbps to 9.95 Mbps and its power consumption does from 423.93 mW to 1028.97 mW at the fixed throughput in response to the signal to noise ratio. The power can be saved up to 58.3% and the throughput can be improved 2.1 times by switching designs appropriately when the distance of the base station and the mobile terminal is not very long Shohei Abe, Yohei Hasegawa, Takao Toi, Takeshi Inuo, Hideharu Amano |
FPT | 5 |
| 2006 | Switch-tagged VLAN Routing Methodology for PC Clusters with EthernetabstractEthernet has been used for connecting hosts in the area of high performance-per-cost PC clusters. Although L2 Ethernet topology is limited to a tree structure, various routing algorithms on topologies suitable for parallel processing can be employed by applying IEEE 802.1Q VLAN technology. However, communication library used in PC clusters does not always support VLANs, so the design of VLAN-based routing method cannot be applied for such PC clusters. In this paper, we propose a switch-tagged VLAN methodology to flexibly set the route of frames on such PC clusters. Since each host does not need to process VLAN tags, the proposed method has advantages in both simple host configuration and high portability. Evaluation results using NAS Parallel Benchmarks showed that performance of topologies supported by the proposed method was comparable with that of an ideal 1-switch (full crossbar) network in the case of a 16-host PC cluster Tomohiro Otsuka, Michihiro Koibuchi, Tomohiro Kudoh, Hideharu Amano |
ICPP | 4 |
| 2006 | Performance and power analysis of time-multiplexed execution on dynamically reconfigurable processorabstractDynamically reconfigurable processor (DRP) developed by NEC Electronics is a coarse grain reconfigurable processor that selects a datapath called a context from the on-chip repository of sixteen circuit configurations at runtime. The time-multiplexed execution based on the multi-context functionality is expected to drastically improve area and power efficiency. To demonstrate the impact of the time-multiplexed execution, we have implemented several stream applications on DRP with various context sizes. Throughout the evaluation based on real application designs, we analyzed the impact of the time-multiplexed execution on performance and power dissipation quantitatively. Yohei Hasegawa, Shohei Abe, Shunsuke Kurotaki, Vu Manh Tuan, Naohiro Katsura, Takuro Nakamura, Takashi Nishimura, Hideharu Amano |
IPDPS | 8 |
| 2006 | A cost-effective context memory structure for dynamically reconfigurable processorsabstractMulticontext reconfigurable processors can switch its configuration in a single clock cycle by providing a context memory in each of the processing elements. Although these processors have proven to be powerful in many applications, the number of contexts is often not enough. The context translation table which translates the global instruction pointer, or the global logical context number, into a local physical context number is proposed to realize a larger application while reducing the actual context memories. Our evaluation using NEC Electronics' DRP-1 shows that the proposed method is effective when the size of the tile is small and the number of context is large. In the most efficient case, the required number of contexts is reduced to 25%, and the total amount of configuration data becomes 6.9%. The template configuration method which extends this idea harnesses the power of multicontext devices by storing basic contexts as templates and combining them to form the actual contexts. While effective in theory, our evaluation shows that the return in adopting such mechanisms in more finer processors as the DRP-1 is minimal where the size of the context memory adds up relative to the number of processing units. Masayasu Suzuki, Yohei Hasegawa, Vu Manh Tuan, Shohei Abe, Hideharu Amano |
IPDPS | 5 |
| 2006 | Enforcing Dimension-Order Routing in On-Chip Torus Networks Without Virtual Channels
Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ISPA | 3 |
| 2006 | A Simple Data Transfer Technique Using Local Address for Networks-on-ChipsabstractNetworks-on-chips (NoCs) have been studied to connect a number of modules in a chip by introducing a network structure which is similar to that in parallel computers. Since embedded streaming applications usually generate predictable small-sized data traffic, the network structure can be customized to the target traffic. Accordingly, we develop a data transfer technique for simplifying routers for predictable small-sized communication in simple tile-based architectures. A data structure is split into single-flit packets, and a label is attached to each of them in order to route them independently. A label is transferred on dedicated wires beside data lines in a channel by taking advantage of relaxed pin count limitations of a channel. To reduce the wiring area for the label, the label is locally assigned according to a preanalysis of required communication pairs of nodes. Analysis results show that only a 3-bit local label is sufficient to route all data of evaluated streaming applications in the case of a 16-node 2D torus. The required amount of hardware for a router is reduced by 37 percent compared with that for a wormhole packet router with the same number of routing table entries. Michihiro Koibuchi, Kenichiro Anjo, Yutaka Yamada, Akiya Jouraku, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2005 | Performance and Cost Analysis of Time-Multiplexed Execution on the Dynamically Reconfigurable ProcessorabstractDynamically reconfigurable processors with multi-context facility have been used for various applications. The relationship between context size and performance of such processors is analyzed based on real designs. The parallelism diagram which shows the required PEs in each step of the algorithm is introduced as the basis of the analysis, and models for performance and cost are shown. Evaluation results show that the performance is degraded about 23% when the size of a context becomes 1/2. The performance per cost is improved 7-14 times than that of the case without time-multiplexed execution. Hideharu Amano, Shohei Abe, Yohei Hasegawa, Katsuaki Deguchi, Masayasu Suzuki |
FCCM | 1 |
| 2005 | Time-multiplexed execution on the dynamically reconfigurable processor: a performance/cost evaluationabstractDynamically Reconfigurable Processor (DRP) developed by NEC Electronics is a coarse grain reconfigurable processor that selects a data path from the on-chip repository of sixteen circuit configurations, or contexts, to implement different logic on one single DRP chip.The impact of time-multiplexed execution to performance and cost is analyzed based on real designs including an IPsec router. The Parallelism Diagram which shows the required PEs in each step of the algorithm is introduced as the basis of the analysis, and models for performance and cost are shown. Evaluation results show that the time-multiplexed execution improves the performance per cost around 4.5 to 14 times than that of the case without time-multiplexed execution. Yohei Hasegawa, Shohei Abe, Katsuaki Deguchi, Masayasu Suzuki, Hideharu Amano |
FPGA | 5 |
| 2005 | An I/O mechanism on a Dynamically Reconfigurable Processor - Which should be moved: Data or Configuration?abstractIn some applications on dynamically reconfigurable processor (DRP), the input/output of data stream occupies about 20% of total execution time. In order to hide the overhead of input/output, the separation of the I/O context and double-buffering mechanism are proposed. Using the mechanism, I/O overhead of six streaming processing programs can be completely hidden. Based on the analysis of I/O performance, two alternatives for executing parallel processing with multiple DRP cores are compared and discussed. Hideharu Amano, Shohei Abe, Katsuaki Deguchi, Yohei Hasegawa |
FPL | 1 |
| 2005 | Efficient Scheduling of Rate Law Functions for ODE-Based Multimodel Biochemical Simulation on an FPGAabstractA reconfigurable biochemical simulator by solving ordinary differential equations has received attention as a personal high speed environment for biochemical researchers. For efficient use of the reconfigurable hardware, static scheduling of high-throughput arithmetic pipeline structures is essential. This paper shows and compares some scheduling alternatives, and analyzes the tradeoffs between performance and hardware amount. Through the evaluation, it is shown that the sharing first scheduling reduces the hardware cost by 33.8% in average, with the up to 11.5% throughput degradation. Effects of sharing of rate law functions are also analyzed. Naoki Iwanaga, Yuichiro Shibata, Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Tomonori Fukushima, Hideharu Amano, Akira Funahashi, Noriko Hiroi, Hiroaki Kitano, Kiyoshi Oguri |
FPL | 7 |
| 2005 | A Framework for ODE-Based Multimodel Biochemical Simulations on an FPGAabstractToday, mathematical modeling and simulation of biochemical pathways take a major role in biological researches. However, modern microprocessors cannot provide enough throughputs to explore the large parameter space of target pathways. To address this problem, ReCSiP (a reconfigurable cell simulation platform), an FPGA-based biochemical simulator is proposed. It's an ODE-based simulator, which solves the rate-law functions. The framework proposed in this paper, enables to simulate pathways consisting many different types of chemical reactions by connecting the rate-law modules (solvers) on an FPGA. It provides the solver-to-solver communication mechanism on an FPGA and automatic configuration software to generate the circuit. Yasunori Osana, Yow Iwaoka, Tomonori Fukushima, Masato Yoshimi, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano |
FPL | 10 |
| 2005 | An Adaptive Cryptographic Accelerator for IPsec on Dynamically Reconfigurable Processor
Yohei Hasegawa, Shohei Abe, Hiroki Matsutani, Hideharu Amano, Kenichiro Anjo, Toru Awashima |
FPT | 4 |
| 2005 | RoMultiC: Fast and Simple Configuration Data Multicasting Scheme for Coarse Grain Reconfigurable Devices
Vasutan Tunbunheng, Masayasu Suzuki, Hideharu Amano |
FPT | 3 |
| 2005 | The Design of Scalable Stochastic Biochemical Simulator on FPGA
Masato Yoshimi, Yasunori Osana, Yow Iwaoka, Akira Funahashi, Noriko Hiroi, Yuichiro Shibata, Naoki Iwanaga, Hiroaki Kitano, Hideharu Amano |
FPT | 9 |
| 2005 | VLAN-Based Minimal Paths in PC Cluster with Ethernet on Mesh and TorusabstractIn a PC cluster with Ethernet, well-distributed multiple paths among hosts can be obtained by applying VLAN technology. In this paper, we propose VLAN topology sets and path assignment methods in mesh and torus. The proposed VLAN-based methods on mesh require N/sup M-1/ and /spl lfloor/N/sup M-1//2/spl rfloor/+1 VLANs to provide balanced minimal paths and partially balanced ones respectively, where N is the number of switches per dimension and M is the number of dimensions. Similarly, those on torus require 2N/sup M-1/ and N/sup M-1/+2 VLANs respectively. Simulation results show that the proposed methods improve up to 902% and 706% of throughput respectively. Tomohiro Otsuka, Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
ICPP | 4 |
| 2005 | Implementation of active direction-pass filter on dynamically reconfigurable processorabstractIn this paper, we report the design and implementation of a sound source separation system using a dynamically reconfigurable device. A robot in real-world environments should have an ability to treat a mixture of multiple sound signals. Active direction-pass filter (ADPF) which extracts sound from a specific direction by using a pair of microphones has been developed as such a method of sound source separation. The ADPF was used as a front-end for an automatic speech recognition system, and recognition of three simultaneous speech signals has been reported. The ADPF, however, requires a lot of computational power, while the battery capacity and the physical size of the robot are limited. To reduce the power consumption and the size of the system, we adopted the dynamically reconfigurable device, DRP developed by NEC Electronics. We implemented the ADPF on DRP, and investigated the effectiveness of dynamically reconfigurable device for these applications. The preliminary experiment shows that ADPF on DRP separates a mixture of sound sources in real-time with practical accuracy. Shunsuke Kurotaki, Noriaki Suzuki, Kazuhiro Nakadai, Hiroshi G. Okuno, Hideharu Amano |
IROS | 5 |
| 2005 | Evaluation of Network Interface Controller on DIMMnet-2 Prototype BoardabstractBy recent performance improvement of interconnection networks for a PC cluster, standard I/O bus which connects network interface becomes the performance bottleneck. DIMMnet is a network interface which can solve the problem by using the memory bus instead of PCI bus or other I/O buses. The second generation network interface DIMMnet-2 can be connected with DDR-SDRAM slot by using the indirect accessing to memory and buffers. Although the current board is a prototype using an FPGA, the latency for 8 Bytes data transfer is only 0.441µs. Akira Kitamura, Yasuo Miyabe, Tetsu Izawa, Tomotaka Miyashiro, Konosuke Watanabe, Tomohiro Otsuka, Hideharu Amano, Yoshihiro Hamada, Noboru Tanabe, Hironori Nakajo |
PDCAT | 7 |
| 2005 | Path selection algorithm: the strategy for designing deterministic routing from alternative paths
Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
Parallel Comput. | 3 |
| 2005 | The performance of SNAIL-2 (a SSS-MIN connected multiprocessor with cache coherent mechanism)
Takashi Midorikawa, Daisuke Shiraishi, Masayoshi Shigeno, Yasuki Tanabe, Toshihiro Hanawa, Hideharu Amano |
Parallel Comput. | 6 |
| 2005 | Performance Evaluation of Deterministic Routings, Multicasts, and Topologies on RHiNET-2 ClusterabstractSystem area networks (SANs), which usually accept arbitrary topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. Although deadlock-free routings, multicasts, and topologies for SANs have been widely developed, their evaluation on real PC clusters was rarely done. Thus, the evaluation of routings, multicasts, and topologies in real systems is important to analyze their impact on the total systems and validate their simulation results. In this paper, we implement and evaluate deadlock-free routings and unicast-based multicasts under various topologies and channel buffer sizes on a PC cluster called RHiNET-2 with 64 hosts. Execution results show that descending layers (DL) routing and structured channel pools improve up to 57 percent of bandwidth and 34 percent of barrier synchronization time compared with up*/down* routing. They also show that, by visiting hosts in numerical order, execution time of unicast-based barrier synchronization is improved up to 28 percent compared with that in random order. However, channel buffer sizes don't affect the bandwidth in the RHiNET-2 cluster. In addition to fundamental evaluation, we appraise them using NAS Parallel Benchmarks, and the DL routing achieves 3.2 percent improvement on their execution time compared with up*/down* routing. Michihiro Koibuchi, Konosuke Watanabe, Tomohiro Otsuka, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2004 | Future reconfigurable computing system
Masahiko Kawamura, Hideharu Amano |
ASP-DAC | 2 |
| 2004 | ReCSiP: a reconfigurable cell simulation platform: accelerating biological applications with FPGA
Yasunori Osana, Tomonori Fukushima, Hideharu Amano |
ASP-DAC | 3 |
| 2004 | Folded Fat H-Tree: An Interconnection Topology for Dynamically Reconfigurable Processor Array
Yutaka Yamada, Hideharu Amano, Michihiro Koibuchi, Akiya Jouraku, Kenichiro Anjo, Katsunobu Nishimura |
EUC | 2 |
| 2004 | Implementing and Evaluating Stream Applications on the Dynamically Reconfigurable ProcessorabstractDynamically reconfigurable processor (DRP) developed by NEC electronics is a coarse grain reconfigurable processor that selects a data path from the on-chip repository of sixteen circuit configurations, or contexts, to implement different logic on one single DRP chip. Several stream applications have been implemented on DRP-1, the first prototype chip, and evaluation results are presented. By computing parallelly using the processing elements(PEs) and distributed memory modules, DRP-1 outperformed pentium III/4 and embedded CPU MIPS64 in some stream application examples. We also present programming techniques applicable on reconfigurable processors and discuss their feasibility in boosting system performance. Noriaki Suzuki, Shunsuke Kurotaki, Masayasu Suzuki, Naoto Kaneko, Yutaka Yamada, Katsuaki Deguchi, Yohei Hasegawa, Hideharu Amano, Kenichiro Anjo, Masato Motomura, Kazutoshi Wakabayashi, Takeo Toi, Toru Awashima |
FCCM | 8 |
| 2004 | Techniques for Virtual Hardware on a Dynamically Reconfigurable Processor - An Approach to Tough Cases
Hideharu Amano, Takeshi Inuo, Hirokazu Kami, Taro Fujii, Masayasu Suzuki |
FPL | 1 |
| 2004 | Stochastic Simulation for Biochemical Reactions on FPGA
Masato Yoshimi, Yasunori Osana, Tomonori Fukushima, Hideharu Amano |
FPL | 4 |
| 2004 | Stream applications on the dynamically reconfigurable processorabstractDynamically reconfigurable processor (DRP) developed by NEC Electronics is a coarse grain reconfigurable processor that selects a data path from the on-chip repository of sixteen circuit configurations, or contexts, to implement different logic on one single DRP chip. Several stream applications have been implemented on the DRP-1, the first prototype chip, and evaluation results are presented. By pipelining the executions, DRP-1 outperformed Pentium III/4, embedded CPU MIPS64, and Texas Instruments DSP TMS320C67J3 in some stream application examples. We also present programming techniques applicable on dynamically reconfigurable processors and discuss their feasibility in boosting system performance. Masayasu Suzuki, Yohei Hasegawa, Yutaka Yamada, Naoto Kaneko, Katsuaki Deguchi, Hideharu Amano, Kenichiro Anjo, Masato Motomura, Kazutoshi Wakabayashi, Takao Toi, Toru Awashima |
FPT | 6 |
| 2004 | BLACK-BUS: A New Data-Transfer Technique Using Local Address on Networks-on-ChipsabstractSummary form only given. Network-on-a-chip (NoC) has received attention as a high-performance interconnect, because traditional buses, which can't transfer more than one data-stream simultaneously, are more likely to become a bottleneck. Since some concepts of NoC have been proposed by simply borrowing the networking structure of parallel computers or system area networks (SANs), it tends to require complicated network interface logic in all the nodes. We propose a novel data-transfer method called Black-Bus as a NoC. In Black-Bus, a local identifier (ID) is attached to each raw data as routing information. Unlike the traditional packet transfer, the local ID is transferred on dedicated wires attached to data lines to remove complicated packet generation procedure in a node. Only a small-sized local ID is required to specify routing tags to the destination, and intermediate routers change it to solve local ID conflicts between paths on a physical channel. The required local ID and routing table sizes for the Black-Bus router are evaluated with access trace data of NAS parallel benchmarks for on-chip multiprocessors, and JPEG codec as stream processing. Evaluation results show that most of the applications require only at most 3 bits for the local ID in a 16-node system. And the Black-Bus data-transfer reduces up to 75% of routing tags compared with global addressing scheme used in the traditional packet networks. Kenichiro Anjo, Yutaka Yamada, Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
IPDPS | 5 |
| 2003 | MAPLE chip: a processing element for a static scheduling centric multiprocessorabstractA custom processor called MAPLE, which supports static scheduling by automatic parallelizing compilers, is implemented and evaluated. MAPLE has a high performance floating point arithmetic unit and low latency data transfer mechanism for other MAPLE chips. The maximum operational frequency is 80 MHz in simulation, and the operation on the prototype board with 23 MHz clock is confirmed. It requires about 0.56 W at 23 MHz operation. Kenta Yasufuku, Riku Ogawa, Keisuke Iwai, Hideharu Amano |
ASP-DAC | 4 |
| 2003 | Performance Evaluation of RHiNET-2/NI: A Network Interface for Distributed Parallel Computing SystemsabstractRHiNET-2/NI is a network interface for a parallel and distributed computing system with network connected PCs. The core of the network interface is an ASIC network controller chip Martini, which provides low-latency and large-bandwidth communication. Evaluation results show that it achieves almost full bandwidth of the 66MHz/64bit PCI bus, which is much larger than that of Myrinet-2000. The performance of a small prototype parallel system achieves almost linear speed up. Konosuke Watanabe, Tomohiro Otsuka, Junichiro Tsuchiya, Hideharu Amano, Hiroshi Harada, Junji Yamamoto, Hiroaki Nishi, Tomohiro Kudoh |
CCGRID | 4 |
| 2003 | Performance Evaluation of Routing Algorithms in RHiNET-2 ClusterabstractSystem area networks (SANs), which usually accept irregular topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. A lot of deadlock-free routings for SANs have been proposed, and their evaluation on simulations have been widely done. However, these simulation results may differ from that of real PC clusters, since hosts, network interfaces and switches used in the simulation are simplified for achieving enough simulation speed. In this paper, we implement deadlock-free routings on a high-performance PC cluster called RHiNET-2, and evaluate their performance. Execution results show that the DL routing and the structured channel pools achieve almost the same total bandwidth and execution time of the barrier synchronization. Compared with the simple Up*/Down* routing, they improve 51% of total bandwidth and 29% improvement on execution time of the barrier synchronization. Michihiro Koibuchi, Konosuke Watanabe, Kenichi Kono, Akiya Jouraku, Hideharu Amano |
CLUSTER | 5 |
| 2003 | A Dynamically Adaptive Switching Fabric on a Multicontext Reconfigurable Device
Hideharu Amano, Akiya Jouraku, Kenichiro Anjo |
FPL | 1 |
| 2003 | Reducing the Configuration Loading Time of a Coarse Grain Multicontext Reconfigurable Device
Toshiro Kitaoka, Hideharu Amano, Kenichiro Anjo |
FPL | 2 |
| 2003 | Implementation of ReCSiP: A ReConfigurable Cell SImulation Platform
Yasunori Osana, Tomonori Fukushima, Hideharu Amano |
FPL | 3 |
| 2003 | An implementation of the Rijndael on Async-WASMIIabstractWASMII is a data driven virtual hardware system based on a multi-context device. Implementation experience of WASMII revealed that the total frequency was often severely degraded by the design of the most complicated context. An approach to address this problem is using asynchronized operation. In this paper, we propose and Async-WASMII, in which each context works in asynchronous manner and implemented it on an asynchronous reconfigurable device PCA (Plastic Cell Architecture). As an application, an encryption algorithm, Rijndael is implemented and evaluated. Yoshinori Adachi, Kenichiro Ishikawa, Satoshi Tsutsumi, Hideharu Amano |
FPT | 4 |
| 2003 | Descending Layers Routing: A Deadlock-Free Deterministic Routing using Virtual Channels in System Area Networks with Irregular TopologiesabstractSystem area networks (SANs), which usually accept irregular topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. Since wormhole or virtual cut-through transfer is used for low latency communication, deadlock-free routings are essential in SANs. We propose a novel deadlock-free deterministic routing called descending layers (DL) routing for SANs. In order to reduce both nonminimal paths and traffic congestion, the network is divided into layers of subnetworks with the same topology using virtual channels, and a large number of paths across multiple subnetworks are established. The DL routing is implemented on a real PC cluster called RHiNET-2, and execution results show that its throughput is improved up to 33% compared with that of up*/down* routing. Its execution time of a barrier synchronization is also improved 29% compared with that of up*/down* routing. Simulation results of various sizes and topologies also show that the DL routing achieves up to 266% improvement on throughput compared with up*/down* routing. Michihiro Koibuchi, Akiya Jouraku, Konosuke Watanabe, Hideharu Amano |
ICPP | 4 |
| 2002 | RHiNET/NI: A Reconfigurable Network Interface for Cluster Computing
Naoyuki Izu, Tomonori Yokoyama, Junichiro Tsuchiya, Konosuke Watanabe, Hideharu Amano |
FPL | 5 |
| 2002 | A General Hardware Design Model for Multicontext FPGAs
Naoto Kaneko, Hideharu Amano |
FPL | 2 |
| 2001 | A prototype chip of multicontext FPGA with DRAM for virtual hardwareabstractDRAM-type multicontext FPGA is hopeful for Virtual Hardware. Since it is possible to implement a large number of contexts in a single chip. However, it has been reported only a few examples because of the difficulty of mixed process of DRAM and logic. Here we try to implement a prototype multi-context FPGA with DRAM for Virtual Hardware. Daisuke Kawakami, Yuichiro Shibata, Hideharu Amano |
ASP-DAC | 3 |
| 2001 | L-Turn Routing: An Adaptive Routing in Irregular NetworksabstractNetwork-based parallel processing using commodity personal computers has been widely developed. Since such systems require high degree of flexibility and scalability of wiring, a high-speed network with an irregular topology is often needed. In traditional routing algorithms for irregular networks, available paths are considerably restricted in order to avoid deadlocks. In this paper we propose a novel routing algorithm called left-up-first turn routing (L-turn routing), which makes a better traffic balancing in irregular networks by building a specific spanning tree. Result of simulations shows that L-turn routing achieves better performance than traditional ones with each topology. Michihiro Koibuchi, Akira Funahashi, Akiya Jouraku, Hideharu Amano |
ICPP | 4 |
| 2001 | Recursive Diagonal Torus: An Interconnection Network for Massively Parallel ComputersabstractRecursive Diagonal Torus (RDT), a class of interconnection network is proposed for massively parallel computers with up to 2/sup 16/ nodes. By making the best use of a recursively structured diagonal mesh (torus) connection, the RDT has a smaller diameter (e.g., it is 11 for 2/sup 10/ nodes) with a smaller number of links per node (i.e., 8 links per node) than those of the hypercube. A simple routing algorithm, called vector routing, which is near-optimal and easy to implement is also proposed. Although the congestion on upper rank tori sometimes degrades the performance under the random traffic, the RDT provides much better performance than that of a 2D/3D torus in most cases and, under hot spot traffic, the RDT provides much better performance than that of a 2D/3D/4D torus. The RDT router chip which provides a message multicast for maintaining cache consistency is available. Using the 0.5 /spl mu/m BICMOS SOG technology, versatile functions, including hierarchical multicasting, combining acknowledge packets, shooting down/restart mechanism, and time-out/setup mechanisms, work at a 60 MHz clock rate. Yulu Yang, Akira Funahashi, Akiya Jouraku, Hiroaki Nishi, Hideharu Amano, Toshinori Sueyoshi |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2000 | A floating point arithmetic unit for a static scheduling and compiler oriented multiprocessor systemabstractNo abstract available. Takahiro Kawaguchi, Takayuki Suzuki, Hideharu Amano |
ASP-DAC | 3 |
| 2000 | MEMOnet : Network interface plugged into a memory slotabstractThe communication architecture of the DIMMnet-1 network interface, based on MEMOnet, is described. MEMOnet is an architecture consisting of a network interface plugged into a memory slot. The DIMMnet-1 prototype will have two banks of PC133 based SO-DIMM slots and an 8 Gbps full duplex optical link or two 448 MB/s full duplex LVDS channel links. The software overhead incurred to generate a message is only I CPU cycle and the estimated hardware delay is less than 100 ns using the atomic on-the-fly sending with header TLB. The estimated achievable communication bandwidth with block on-the-fly sending with protection stampable window memory is 440 MB/s which was observed in our experiments writing to the DIMM area with a write combining attribute. This is 3.3 times higher than the maximum bandwidth of PCI. This high performance distributed computing environment is available using economical personal computers with DIMM slots. Noboru Tanabe, Junji Yamamoto, Hiroaki Nishi, Tomohiro Kudoh, Yoshihiro Hamada, Hironori Nakajo, Hideharu Amano |
CLUSTER | 7 |
| 2000 | A Virtual Hardware System on a Dynamically Reconfigurable Logic DeviceabstractWASMII is virtual hardware using a multi-context reconfigurable device with a data driven control. Since implementation of WASMII was infeasible due to the unavailability of such a device, the system has been only evaluated using an emulator so far. However, the first reconfigurable multi-context device called DRL has been developed by NEC. Making the use of its flexible reconfigurability, we have implemented a mechanism of WASMII on DRL. Yuichiro Shibata, Masaki Uno, Hideharu Amano, Koichiro Furuta, Taro Fujii, Masato Motomura |
FCCM | 3 |
| 2000 | A Reconfigurable Stochastic Model Simulator for Analysis of Parallel SystemsabstractMarkov chain and queueing model are convenient tools with which to analyze parallel systems for architects. For a high speed simulation and easy modeling, a reconfigurable Markov chain/queueing model simulation system called RSMS (Reconfigurable Stochastic Model Simulator) is proposed. A user describes the target system in a dedicated description language called Taico. The description is automatically translated into the HDL description of the Markov chain/queueing model simulator. Then, the simulator is implemented on the FPGA devices of the reconfigurable system, and directly executed. Evaluation results with example parallel systems demonstrate that the performance of the proposed system is much superior to that of a common workstation. Ou Yamamoto, Yuichiro Shibata, Hitoshi Kurosawa, Hideharu Amano |
FCCM | 4 |
| 2000 | A Local Area System Network RHinet-1: A Network for High Performance Parallel ComputingabstractThe Real World Computing Partnership (RWCP) has developed a local area system network (LASN) called RHiNET-1 (RWCP High-performance NETwork, version 1) using 1.33-Gbps optical interconnections for high-performance computing using personal computers distributed in an office or laboratory environment. The network interface, RHiNET-1/NI, uses a complex programmable logic device (CPLD) based protocol controller to provide an easy evaluation platform for various protocols. It fits in a 32-bit/33-MHz PCI bus. The switch, RHiNET-1/SW, consists of a single-chip CMOS switch and external SRAM. It provides low-latency, reliable communication with a flexible topology design. We are currently evaluating protocols on RHiNET-1. RHiNET-1 will enable a new form of high-performance computing environment. We are also developing the second implementation, RHiNET-2. RHiNET-2/NI will support a 64-bit/66-MHz PCI bus. RHiNET-2/SW is an 8-Gbps/port 8/spl times/8 single-chip ASIC switch. The aggregate bandwidth of RHiNET-2/SW is 64 Gbps. Hiroaki Nishi, Koji Tasho, Junji Yamamoto, Tomohiro Kudoh, Hideharu Amano |
HPDC | 5 |
| 1999 | Performance evaluation of SNAIL: A multiprocessor based on the simple serial synchronized multistage interconnection network architecture
Junji Yamamoto, Takashi Fujiwara, T. Komeda, Takayuki Kamei, Toshihiro Hanawa, Hideharu Amano |
Parallel Comput. | 6 |
| 1998 | Reconfigurable Systems: Activities in Asia and South Pacific (Embedded Tutorial)abstractSystems and researches on reconfigurable systems in Asia and South Pacific are picked up and introduced. Like Northern America and European countries, various platforms, application specific systems and education platforms have been proposed and developed. Hideharu Amano, Yuichiro Shibata |
ASP-DAC | 1 |
| 1998 | The MINC (Multistage Interconnection Network with Cache Control Mechanism) ChipabstractAlthough bus connected multiprocessors have been widely used as high-end workstations or servers, the number of connected processors is strictly limited by the maximum bandwidth of the shared bus. Instead of them, a switch connected multiprocessor which uses a crossbar or Multistage Interconnection Networks (MINs) for connecting processors and memory modules is a hopeful candidate. However, in such a system, a snoop cache technique in bus connected multiprocessors cannot be used, and consistency problems must be solved for providing the cache memory between a processor and the switch. To address this problem, hardware approaches by making the best use of advanced VLSI technology have been proposed. However, traditional methods require a large memory outside the switching element and it causes not only a large additional hardware but also the extra latency by accessing the outside memory. Moreover, the complicated MIN with cache or directory must also treat data packet which should be transferred quickly. In order to solve these problems, we proposed the MINC (MIN with Cache control mechanism). In the MINC, the MIN which only transfers a part of the address and cache coherent messages is separated from the data transfer network, and pushed into an LSI chip called the MINC chip. The coherent control is done based on the directory using the reduced Hierarchical Bit-map Directory scheme (RHBD). In order to reduce unnecessary packets, the pruning cache which is a small cache enough to implement inside the chip is introduced in the MINC chip. Takashi Midorikawa, Takayuki Kamei, Toshihiro Hanawa, Hideharu Amano |
ASP-DAC | 4 |
| 1997 | An LSI implementation of the simple serial synchronized multistage interconnection networkabstractA high speed switch is a critical component of multiprocessors. Multistage interconnection network (MIN) has been utilized as a switch for connection processors and memory modules in multiprocessors. Unlike the crossbar, it consists of small switching elements, and provides a high bandwidth with relatively small hardware. Most of traditional MINs are blocking networks and packets are transferred in the store-and-forward manner between switching elements with bit-parallel (8-64bits) lines. Since the width of communication paths and transferred manner cause pin-limitation problems and complicated structure, the high density implementation and high speed clock is not utilized. In order to solve these problems, we implemented the SSS-PBSF chip. This switch uses the PBSF connection structure which can obtain a higher bandwidth than that of crossbar with connecting banyan networks in a 3D direction. A simple serial synchronized (SSS) style control mechanism is adopted both for high speed operation and solving the pin-limitation problem. Takayuki Kamei, Masashi Sasahara, Hideharu Amano |
ASP-DAC | 3 |
| 1997 | The RDT network router chipabstractThe RDT network router chip is a versatile router for the massively parallel computer prototype JUMP-1. The major goal of this project is to establish techniques for building an efficient distributed shared memory on a massively parallel processor. For this purpose, the reduced hierarchical bit-map directory (RHBD) schemes are used for efficient cache management of the distributed shared memory. In order to implement (RHBD) schemes efficiently, we proposed a novel interconnection network RDT (recursive diagonal torus), and developed a sophisticated router chip for the RDT which equips a hierarchical multicast mechanism without deadlock and acknowledge combining mechanism. By using the 0.5/spl mu/BiCMOS SOG technology it can transfer all packets synchronized with a unique CPU clock(60MHz). Long coaxial cables are directly driven with the ECL interface of this chip. The mixed design approach with schematic and VHDL permits the development of the complicated chip with 90,522 gates in a year. Hiroaki Nishi, Hideharu Amano, Katsunobu Nishimura, Kenichiro Anjo, Tomohiro Kudoh |
ASP-DAC | 2 |
| 1997 | Shared vs. Snoop: Evaluation of Cache Structure for Single-Chip Multiprocessors
Toru Kisuki, Masaki Wakabayashi, Junji Yamamoto, Keisuke Inoue, Hideharu Amano |
Euro-Par | 5 |
| 1995 | Hierarchical Bit-Map Directory Schemes on the RDT Interconnection Network for a Massively Parallel Processor JUMP-1
Tomohiro Kudoh, Hideharu Amano, Takashi Matsumoto 0002, Kei Hiraki, Yulu Yang, Katsunobu Nishimura, Koichi Yoshimura, Yasuhito Fukushima |
ICPP (1) | 2 |
| 1995 | Neural network parallel computing for multi-layer channel routing problems
Kyotaro Suzuki, Hideharu Amano, Yoshiyasu Takefuji |
Neurocomputing | 2 |
| 1995 | A Performance Evaluation of the Multiprocessor Testbed ATTEMPT-0
Takuya Terasawa, Ou Yamamoto, Tomohiro Kudoh, Hideharu Amano |
Parallel Comput. | 4 |
| 1995 | WASMII: An MPLD with data-driven control on a virtual hardware
Xiao-ping Ling, Hideharu Amano |
J. Supercomput. | 2 |
| 1994 | Multistage Interconnection Networks with Multiple OutletsabstractMultistage Interconnection Networks(MINs) with multiple outlets are networks which can support higher bandwidth than that of nonblocking networks by passing multiple packets to the same destination. A novel MIN topology with multiple outlets called Piled Banyan Switching Fabrics (PBSF) is proposed for the Simple Serial Synchronized (SSS)-MIN used in multiprocessors, and analyzed with other two types of MIN with multiple outlets called Multi-Banyan Switching Fabrics (MBSF) and Tandem Banyan Switching Fabrics (TBSF). The throughput of these MINs is evaluated and compared with both the theoretical model and simulation. The PBSF supports the best throughput and latency used for the SSS-MIN. Although the latency of the TBSF is large, the pass-through ratio is close to 1 if the number of connected banyan networks are more than 4. Therefore, the TBSF is useful for the ATM switching networks in which the relatively large latency is tolerable. The conflict-free access of these MINs is also analyzed, and it appears that rows, column, forward and backward diagonal of the matrix can be accessed without conflict. Toshihiro Hanawa, Hideharu Amano, Yoshifumi Fujikawa |
ICPP (1) | 2 |
| 1994 | SNAIL: A Multiprocessor Based on the Simple Serial Synchronized Multistage Interconnection Network ArchitectureabstractSimple Serial Synchronized (SSS) Multistage Interconnection Network (MIN) is a novel MIN architecture for connecting processors and memory modules in multiprocessors. Synchronized bit-serial communication simplifies the structure/control, and also solves the pin-limitation problem. Here, design, implementation, and evaluation of a multiprocessor prototype called SNAIL with the SSS-MIN are presented. The heart of SNAIL is the prototype 1 /mu CMOS SSS-MIN gate array chip which exchanges packets from 16 inputs with 50MHz clock. The message combining is implemented only with 20% increases of the hardware. From the empirical evaluation with some application programs, it appears that the latency and synchronization overhead of the SSSMIN are tolerable, and the bandwidth of the SSS-MIN is sufficient. Although the performance improvement with the bit serial message combine is not so large (1%) when instructions are stored in the local memory, it becomes up to 400% when instructions are stored in the shared memory. Masashi Sasahara, Jun Terada, Luo Zhou, Kalidou Gaye, Jun-ichi Yamato, Satoshi Ogura, Hideharu Amano |
ICPP (1) | 7 |
| 1992 | A Parallel Logic Simulation Algorithm Based on Query
Tomohiro Kudoh, Tetsuro Kimura, Hideharu Amano, Takuya Terasawa |
ICPP (3) | 3 |
| 1991 | A Batcher Double Omega Network with Combining
Hideharu Amano, Kalidou Gaye |
ICPP (1) | 1 |
| 1990 | A Fault Tolerant Batcher Network
Hideharu Amano |
ICPP (1) | 1 |
| 1990 | (SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculationsabstract(SM)/sup 2/-II is a large-scale parallel machine dedicated to scientific computation which includes sparse matrix calculations. In order to connect thousands of microprocessors and utilize a high degree of parallelism, the whole (SM)/sup 2/-II system is designed based on a simple computational model called the node and connecting-line (NC) model. The concept and the architecture of (SM)/sup 2/-II are described. The NC-model and a language called node oriented concurrent C (NCC) are derived. The concurrent process controller is briefly introduced. Receiver selectable multicast (RSM) is proposed, and the structure which allows connection of a large number of processing units is described. The performance of the RSM is analyzed. Some connection structures for clusters are evaluated. An operational prototype is introduced.> Hideharu Amano, Taisuke Boku, Tomohiro Kudoh |
IEEE Trans. Computers | 1 |
| 1988 | IMPULSE: A High Performance Processing Unit for Multiprocessors for Scientific CalculationabstractA high-performance processing unit for multiprocessor systems for scientific calculations, called Impulse, is described. Impulse is equipped with a hardware process-control mechanism, and a powerful floating-point processor and its controller. The process-control method is based on the concurrent process model called the NC model. In the NC model, the processes and their communicating channels are static, and it is relatively easy to implement the interprocess communication server and process scheduler in hardware according to this model. To enhance the system performance, Impulse is composed of three parts, the task engine, IPC engine, and FFP engine. From the results of simulations, it appears that if IPC engine provides efficient process control even if the granularity of the processes is very fine.> Taisuke Boku, Shigehiro Nomura, Hideharu Amano |
ISCA | 3 |
| 1985 | (SM)²-II: A New Version of the Sparse Matrix Solving Machineabstractarticle Free Access Share on (SM)2-II: a new version of the sparse matrix solving machine Authors: Hideharu Amano Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Taisuke Boku Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Tomohiro Kudoh Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Hideo Aiso Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile Authors Info & Claims ACM SIGARCH Computer Architecture NewsVolume 13Issue 3June 1985 pp 100–107https://doi.org/10.1145/327070.327137Published:01 June 1985Publication History 16citation208DownloadsMetricsTotal Citations16Total Downloads208Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Hideharu Amano, Taisuke Boku, Tomohiro Kudoh, Hideo Aiso |
ISCA | 1 |
| 1983 | (SM)2: Sparse Matrix Solving MachineabstractIn analyzing electronic circuits, it is usually necessary to solve simultaneous linear equations which provide with a sparse coefficient matrix. In order to treat these problems effectively, we propose a dedicated parallel machine called 'Sparse Matrix Solving Machine', or (SM)2 for short. Hideharu Amano, Takaichi Yoshida, Hideo Aiso |
ISCA | 1 |