VLDB 2026 Research / reviewers in the wild / expert
Michihiro Koibuchi
dblp:66/4132
· DBLP profile ↗
101ranked-venue papers
17as first author
16since 2021 · last 2026
0000-0002-5790-6992ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 75 · 16 first-author · 9 since 2021Computer networks · 7 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 590-Nanosecond 757-Gbps FPGA Lossy Compressed NetworkabstractInter-FPGA communication bandwidth has become a limiting factor in scaling memory-intensive workloads on FPGA-based systems. While modern FPGAs integrate high- bandwidth memory (HBM) to increase local memory throughput, network interfaces often lag behind, creating an imbalance between computation and communication resources. Data compression is a technique to increase effective communication bandwidth by reducing the amount of data transferred, but existing solutions struggle to meet the performance and operation latency requirements of FPGA-based platforms. This paper presents a high- throughput lossy compression framework that enables sub-microsecond latency communication in FPGA clusters. The proposed design addresses the challenge of aligning variable-length compressed data with fixed-width network channels by using transpose circuits, memory-bank reordering, and word- wise operations. A run-length encoding scheme with bounded error is employed to compress floating-point and fixed-point data without relying on complex fine-grained bit-level manipulations, enabling low-latency and scalable implementation. The proposed architecture is implemented on a custom Stratix 10 MX2100 FPGA card equipped with eight 50 Gbps network ports and silicon photonics transceivers. The system achieves up to 757 Gbps of aggregate bandwidth per FPGA in collective communication operations. Compression and decompression are performed within 590 ns total latency, while maintaining the quality of results in a GradAllReduce workload for deep learning. Michihiro Koibuchi, Takumi Honda, Naoto Fukumoto, Shoichi Hirasawa, Koji Nakano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Lossy Compressed Collective Inter-FPGA CommunicationsabstractA cutting-edge FPGA can be equipped with many memory channels of HBMs.The bottleneck for inter-FPGA memory communication would become the aggregate network bandwidth.To fill the gap between memory and network bandwidth, this study presents an approximate inter-FPGA memory network that works at the expense of output quality.Lossy data compression is a typical way of providing approximate communication.Prior works illustrated that simple lossy compression algorithms with a higher than 1.8 compression ratio were acceptable for communications generated in typical parallel applications, such as two-dimensional Lattice Boltzmann Method (2D-LBM), k-means data clustering, fast Fourier transform (FFT), and conjugate gradient method (CG) applications.Although the lossy compression for approximate communication has been widely studied, its feasibility and evaluation studies using high-bandwidth interconnection networks on real systems were rarely done.We design and evaluate lossy compressed collective communications so that all the memory and network bandwidth are fully used in a custom Stratix10 MX2100 FPGA card.Our evaluation results show that the custom card transfers up to 549.9-Gbps of collective (scatter, allgather, alltoall) data on an FPGA cluster, while the original design without the compression transfers up to 360 Gbps.The FPGA sender and receiver overhead, including (de)compression, are 312.8nsand 365.7ns.Our cycle-accurate network simulation shows that a high compression ratio significantly improves the effective network throughput of typical synthetic traffic patterns, especially for long messages. Michihiro Koibuchi, Yoshinobu Ishida, Shoichi Hirasawa, Yao Hu 0005, Takumi Honda, Yusuke Nagasaka, Naoto Fukumoto |
HPC Asia | 1 |
| 2024 | A Bandwidth-Optimal All-to-All Communication in Two-Dimensional Fully Connected NetworkabstractHigh-radix direct interconnection networks enable many simultaneous, non-blocking point-to-point communications at every compute node of high-performance computing systems. This feature provides a novel opportunity to improve the performance of collective communications. In this study, we focus on two-dimensional fully connected network topologies (2dfc) which are basic blocks of cutting-edge high-radix network topologies such as Dragonfly and HyperX. We aim to design a bandwidth-optimal algorithm for the All-to-All collective communication in the 2dfc network. By using parallel point-to-point communications, the proposed algorithm enables parallel non-blocking intra- and inter-group communications with striping data placement. Simulation results using SimGrid illustrated that the proposed algorithm outperforms the conventional algorithms significantly. Compared with topology-aware algorithms, the proposed approach reduces 1.66× and 1.58× communication time of an existing and a topology-aware collective algorithm in a 16 × 16 2dfc network, respectively. Kien Trung Pham, Truong Thao Nguyen, Michihiro Koibuchi |
CCGrid | 3 |
| 2023 | An Auto-Tuning Method for High-Bandwidth Low-Latency Approximate Interconnection NetworksabstractAhstract-The next-generation interconnection networks, such as 400 GbE specification, impose Forwarding Error Correction (FEC) operation, such as RS-FEC (544,514), to incoming packets at every switch. The significant FEC latency increases the end-to-end communication latency that degrades the application performance in parallel computers. To resolve the FEC latency problem, a prior work presented error-prone high-bandwidth low-latency networks that do not perform the FEC operation. They enable high-bandwidth approximate data transfer and low-bandwidth perfect data transfer to support various kinds of parallel applications subject to different levels of probability of bit-flip occurrence. As the number of approximate data transfers increases, the parallel applications can obtain a significant speedup of their execution at the expense of the moderate degraded quality of results (QoRs). However, it is difficult for users to identify whether each communication should be approximate or not, so as to obtain the shortest execution time with enough QoRs for a given parallel application. In this study, we apply an auto-tuning framework for approximate interconnection networks; it automatically identifies whether each communication should be approximate data transfer or not, by attempting thousands executions of a given parallel application. An auto-tuning attempts a large number of program executions by varying the possible communication parameters to find out the best execution configuration of the program. The multiple executions would generate different positions of bit flips on communication data that may provide different qualities of results even if the same parameters are taken. Although this uncertainty introduces difficulties in the optimization of the auto-tuning, many offline trials lead to a high probability of the program's success execution. Evaluation results show that high-performance MPI applications with our auto-tuning method result in 1.30 average performance improvement on error-prone high-performance approximate networks. Shoichi Hirasawa, Michihiro Koibuchi |
PDP | 2 |
| 2023 | Designing low-diameter interconnection networks with multi-ported host-switch graphsabstractSummary A host‐switch graph was originally proposed as a graph that represents a network topology of a computer systems with 1‐port host computers and ‐port switches. It has been studied from both theoretical and practical aspects in terms of the diameter, the average shortest path length, and the performance of real applications. In recent high‐performance computing systems, however, a host computer is connected to multiple switches by using InfiniBand, NVSwitch, or Omni‐Path, and consequently they provide high bandwidths. Since a host‐switch graph cannot represent such systems, this article extends a host‐switch graph so that it can represent such systems. As a result, a host‐switch graph can include multi‐ported hosts. Furthermore, we propose to use multi‐port hosts for reducing the diameter. We show that the diameter minimization is equivalent to solving the degree diameter problem for bipartite graphs of diameter three. Our experimental results show that we can drastically reduce the diameter as well as increasing the bandwidth and improves performance of MPI applications by up to 162% as compared with networks with single‐ported hosts. Ryota Yasudo, Koji Nakano, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Effective switchless inter-FPGA memory networks
Truong Thao Nguyen, Kien Trung Pham, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
J. Parallel Distributed Comput. | 5 |
| 2022 | A Scalable Distributed Radix Sorter for FPGA Clusters using High-Bandwidth Memory NetworksabstractA modern FPGA card can be equipped with high bandwidth memory, such as HBM2. Since the amount of the memory is limited on an FPGA, highly parallel data processing becomes crucial on tightly coupled FPGAs by high-density optical integration, e.g., onboard Si-Photonics transceivers. This study presents a scalable distributed radix sorter, and implements it on an eight-FPGA cluster. Each custom Stratix10 MX2100 FPGA card has 819-Gbps memory bandwidth with two HBM2 memories and 800-Gbps network bandwidth with eight custom embedded optical modules. Existing FPGA sorter typically relies on a merge sort. However, it has a severe performance bottleneck at the final stage of data merge, which cannot make the best use of the high memory-to-memory bandwidth on the FPGA cluster. Instead, we implement a radix sort for a 32-bit key range consisting of eight 4-bit counting sorts optimized to the memory-network structure. Each counting sort needs memory read/write access only once through global and local pipelines. We demonstrated a sorting throughput of 37.2 GB/s. Yutaka Urino, Takanori Shimizu, Hiroshi Yamaguchi, Kenji Mizutani, Shigeru Nakamura, Tatsuya Usuki, Michihiro Koibuchi |
FCCM | 7 |
| 2022 | Scalable Low-Latency Inter-FPGA NetworksabstractA cutting-edge FPGA card can be equipped with many high-bandwidth I/Os by means of high-density optical integration, e.g., onboard Si-photonics transceivers, to provide high network bandwidth for memory-to-memory inter-FPGA communication. This study presents its scalable switchless net-work architecture by exploiting an indirect path, consisting of two one-hop paths, for enabling a diameter-2 network topology. It then takes a Kautz network topology with a diameter of two for connecting d(d + 1) FPGAs with a degree of$d$, which is close to the theoretical upper bound. The Kautz network topologies have bi-directional links and uni-directional links which form triangles. Uni-directional links introduce difficulty in avoiding channel buffer overflow because the existing link-level flow control assumes a bi-directional link. This study presents an indirect flow control along a uni-directional triangle embedded in the Kautz network topology. It then develops a combination of unicasts that forms multi-port collective communications to mitigate the influence of the startup latency on the execution time. Since a high-degree FPGA card introduces difficulty in storing many I/O ports at the panel of a 1- U compute server, we propose using WDM (Wavelength Division Multiplexing) as an alternative and present its efficient mapping onto arrayed waveguide grating (AWG). The required number of wavelengths becomes d on d+ 1 AWG equipments. Based on our experimental results with OPTWEB of custom Stratix10 FPGA cards, SimGrid simulation results show that our collective communication is 7 × faster than that of Dragonfly with 272 FPGAs. Kien Trung Pham, Truong Thao Nguyen, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IPDPS | 5 |
| 2022 | A High-Radix Circulant Network Topology for Efficient Collective Communication
Ke Cui, Michihiro Koibuchi |
PDCAT | 2 |
| 2022 | A Hardware Trojan Exploiting Coherence Protocol on NoCs
Yoshiya Shikama, Michihiro Koibuchi, Hideharu Amano |
PDCAT | 2 |
| 2021 | Packet Forwarding Cache of Commodity Switches for Parallel ComputersabstractSwitch delay dominates communication latencies in interconnection networks, especially for short messages because switch delays are massive relative to the link and packet injection delays. At a conventional switch, routing decision is based on CAM (Content Addressable Memory) table lookup, and it imposes a significant delay. Reducing the access latency to CAM is crucial for the upcoming low-delay switch in parallel computers. Besides the CAM latency problem, the packet forwarding rate is not proportional to the switching capacity on cutting-edge commodity switches of interconnection networks. A switch will not able to forward incoming packets at the maximum line rate. To resolve the latency and throughput problems, we explore an on-chip packet forwarding cache to a switch. An incoming packet avoids large-latency accessing a CAM forwarding table if the cache hits. It supports an almost 100% hit rate (no capacity miss nor conflict miss) for packets generated in up to 2K-node jobs. For 100% hit rate on larger jobs, we present a switchable hash function to refer to a packet forwarding table on a switch. The switchable hash function is optimized to typical network topologies, i.e., k-ary n-cubes, fat trees, and Dragonfly. The main idea is that a large number of packet destinations share a same index tag, resulting in the same required number of cache entries as the number of output ports. This design can be enabled by the path regularity of the above network topologies. Our evaluation results show that the reasonable packet forwarding cache supports a 933-Gbps line rate even for incoming shortest packets on the above network topologies. We illustrate that parallel applications obtain the performance gain of 5.07x speed up using the cache switches since the impact of the switch delay and link bandwidth is significant on the end-to-end communication performance. Shoichi Hirasawa, Hayato Yamaki, Michihiro Koibuchi |
CLUSTER | 3 |
| 2021 | The Case for Disjoint Job Mapping on High-Radix Networked Parallel Computers
Yao Hu 0005, Michihiro Koibuchi |
ICA3PP (2) | 2 |
| 2021 | Accelerating MPI Communication Using Floating-point Compression on Lossy Interconnection NetworksabstractApproximate communication has arisen as a new opportunity for improving the efficiency of communication in parallel computer systems, which can significantly reduce the communication time by transmitting partial or imprecise messages. In this study, we explore application-level approximate communication techniques on lossy high-performance interconnection networks that leave bit flips. Although existing compression techniques do not assume lossy interconnection networks, our approximate-communication challenge is to co-design a lossy floating-point compression algorithm and a critical bit-flip recovery scheme optimized to a given bit error rate (BER) on a target lossy interconnection network. Our objective is to transfer the maximum amount of approximate data over the least amount of compression overhead and bit-flip recovery time. Our scheme is implemented into several representative communication-intensive MPI (Message Passing Interface) applications, and it is shown that our approximate communication scheme effectively speeds up the total execution time without much loss in the quality of the result. Yao Hu 0005, Michihiro Koibuchi |
LCN | 2 |
| 2021 | A Case for Low-Latency Network-on-Chip using Compression RoutersabstractThe communication latency is a primary concern for designing Network-on-Chips (NoCs) since it significantly affects the parallel application performance on a many-core computer system. To reduce the communication latency, we propose an on-chip router that (de)compresses the contents of an incoming packet before completing switch arbitration. The compression router thus has no latency penalty for the compression operation, whereas it shortens a packet length that decreases the network injection-and-ejection latency. Evaluation results show that the compression router improves 7.7% of the parallel application performance (IS, CG, FT, and TSP) and 49% of the effective network throughput by 1.8 compression ratio on NoC. The drawback is that the router area and its energy consumption per bit increase by 0.12mm2and 1.4 times compared to the conventional virtual-channel router. Naoya Niwa, Yoshiya Shikama, Hideharu Amano, Michihiro Koibuchi |
PDP | 4 |
| 2021 | Low-Latency Low-Energy Memory-Cube Networks using Dual-Voltage DatapathsabstractThree-dimensional stack memory that provides both high-bandwidth access and large capacity is a promising technology for next-generation computer systems. While a large number of memory cubes increase the aggregate memory capacity, the communication latency and power consumption would become significant due to its low-radix large-diameter packet network. In this context, we propose a memory-cube network called Diagonal Memory Network (DMN). A diagonal network topology, its floor layout, and its lightweight router are designed for low-latency and low-voltage memory-read communication. Our evaluation results show that a DMN router decreases 31% of the hardware resources than a conventional virtual-channel router. The DMN router reduces 13% and 67% energy consumption to transit a packet along with the original datapath and bypassing datapath, respectively. Yoshiya Shikama, Ryuta Kawano, Hiroki Matsutani, Hideharu Amano, Yusuke Nagasaka, Naoto Fukumoto, Michihiro Koibuchi |
PDP | 7 |
| 2021 | OPTWEB: A Lightweight Fully Connected Inter-FPGA Network for Efficient CollectivesabstractModern FPGA accelerators can be equipped with many high-bandwidth network I/Os, e.g., 64 x 50 Gbps, enabled by onboard optics or co-packaged optics. Some dozens of tightly coupled FPGA accelerators form an emerging computing platform for distributed data processing. However, a conventional indirect packet network using Ethernet's Intellectual Properties imposes an unacceptably large amount of the logic for handling such high-bandwidth interconnects on an FPGA. Besides the indirect network, another approach builds a direct packet network. Existing direct inter-FPGA networks have a low-radix network topology, e.g., 2-D torus. However, the low-radix network has the disadvantage of a large diameter and large average shortest path length that increases the latency of collectives. To mitigate both problems, we propose a lightweight, fully connected inter-FPGA network called OPTWEB for efficient collectives. Since all end-to-end separate communication paths are statically established using onboard optics, raw block data can be transferred with simple link-level synchronization. Once each source FPGA assigns a communication stream to a path by its internal switch logic between memory-mapped and stream interfaces for remote direct memory access (RDMA), a one-hop transfer is provided. Since each FPGA performs input/output of the remote memory access between all FPGAs simultaneously, multiple RDMAs efficiently form collectives. The OPTWEB network provides 0.71-μsec start-up latency of collectives among multiple Intel Stratix 10 MX FPGA cards with onboard optics. The OPTWEB network consumes 31.4 and 57.7 percent of adaptive logic modules for aggregate 400-Gbps and 800-Gbps interconnects on a custom Stratix 10 MX 2100 FPGA, respectively. The OPTWEB network reduces by 40 percent the cost compared to a conventional packet network. Kenji Mizutani, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IEEE Trans. Computers | 4 |
| 2020 | Dual-Plane Isomorphic Hypercube NetworkabstractWe propose a multi-plane isomorphic network that increases network throughput and reduces network latency by effectively configuring multi-plane networks. In the proposed network, each plane adopts the same graph topology but different switch-to-switch connections. We evaluate the dual-plane isomorphic hypercube network by graph analysis and cycle level simulation. Results of the graph analysis show that the dual-plane isomorphic 8-hypercube reduces the average shortest path length by 22% and improves throughput by 28% compared with the dual-plane hypercube. Similar improvements are confirmed from the results of the cycle level simulation. We also examine the dual-plane isomorphic folded-hypercube network. Finally, we discuss the effect of longer cable length caused by the isomorphic network on the network cost and latency. Takeo Hosomi, Ryota Yasudo, Michihiro Koibuchi, Shinji Shimojo |
HPC Asia | 3 |
| 2020 | Accelerating Deep Learning using Multiple GPUs and FPGA-Based 10GbE SwitchabstractA back-propagation algorithm following a gradient descent approach is used for training deep neural networks. Since it iteratively performs a large number of matrix operations to compute the gradients, GPUs (Graphics Processing Units) are efficient especially for the training phase. Thus, a cluster of computers each of which equips multiple GPUs can significantly accelerate the training phase. Although the gradient computation is still a major bottleneck of the training, gradient aggregation and parameter optimization impose both communication and computation overheads, which should also be reduced for further shortening the training time. To address this issue, in this paper, multiple GPUs are interconnected with a PCI Express (PCIe) over 10 Gbit Ethernet (10GbE) technology. Since these remote GPUs are interconnected via network switches, gradient aggregation and optimizers (e.g., SGD, Adagrad, Adam, and SMORMS3) are offloaded to an FPGA-based network switch between a host machine and remote GPUs; thus, the gradient aggregation and optimization are completed in the network. Evaluation results using four remote GPUs connected via the FPGA-based 10GbE switch that implements the four optimizers demonstrate that these optimization algorithms are accelerated by up to 3. 0x and 1. 25x compared to CPU and GPU implementations, respectively. Also, the gradient aggregation throughput by the FPGA-based switch achieves 98.3% of the 10GbE line rate. Tomoya Itsubo, Michihiro Koibuchi, Hideharu Amano, Hiroki Matsutani |
PDP | 2 |
| 2019 | Sparse 3-D NoCs with Inductive CouplingabstractWireless interconnects based on inductive coupling technology are compelling propositions for designing 3-D integrated chips. This work addresses the heat dissipation problem on such systems. Although effective cooling technologies have been proposed for systems designed based on Through Silicon Via (TSV), their application to systems that use inductive coupling is problematic because of increased wireless-communication distance. For this reason, we propose two methods for designing sparse 3-D chips layouts and Networks on Chip (NoCs) based on inductive coupling. The first method computes an optimized 3-D chip layout and then generates a randomized network topology for this layout. The second method uses a standard stack chip layout with a standard network topology as a starting point, and then deterministically transforms it into either a "staircase" or a "checkerboard" layout. We quantitatively compare the designs produced by these two methods in terms of network and application performance. Our main finding is that the first method produces designs that ultimately lead to higher parallel application performance, as demonstrated for nine OpenMP applications in the NAS Parallel Benchmarks. Michihiro Koibuchi, Lambert T. Leong, Tomohiro Totoki, Naoya Niwa, Hiroki Matsutani, Hideharu Amano, Henri Casanova |
DAC | 1 |
| 2019 | Diameter/ASPL-Based Mapping of Applications with Uncertain Communication over Random Interconnection NetworksabstractThe steadily increasing number of compute nodes in high-performance computing (HPC) systems calls for more efficient interconnection networks and application mapping strategies. Traditionally, the user applications are mapped onto a regular network with nearby compute nodes to reduce the network distances. However, it is likely that the rigid mapping limitation on regular topologies may largely delay the dispatch time when the workload is heavy at a short period. In this work, we target the applications with unknown communication patterns over random interconnection networks, which have drawn increasing attention in the HPC world due to their low diameter and low average shortest path length (ASPL). We propose to use topology embedding metrics, i.e., diameter and ASPL, for application mapping to overcome the rigid limitation on regular topologies. We investigate the time-space tradeoff among several application mapping algorithms using a compound application workload. Evaluation results show that, when compared to the baseline random connected topology mapping, the proposed diameter/ASPL-based topology mapping algorithms reduce up to 48.0% makespan and up to 78.1% average turnaround time and improve up to 1.9X system utilization when the instantaneous workload is heavy over random interconnection networks. Yao Hu 0005, Michihiro Koibuchi |
ICPADS | 2 |
| 2019 | The Case for Water-Immersion Computer BoardsabstractA key concern for a high-power processor is heat dissipation, which limits the power, and thus the operating frequencies, of chips so as not to exceed some temperature threshold. In particular, 3-D chip integration will further increase power density, thus requiring more efficient cooling technology. While air, fluorinert and mineral oil have been traditionally used as coolants, in this study, we propose to directly use tap or natural water due to its superior thermal conductivity. We have developed the "in-water computer" prototypes that rely on a parylene film insulation coating. Our prototypes can support direct water-immersion cooling by taking and draining natural water, while existing cooling requires the secondary coolant (e.g. outside air in cold climates) for cooling the primary coolants that contact chips. Our prototypes successfully reduce by 20 degrees the chip temperature of commodity processor chips. Our analysis results show that the in-water cooling increases the acceptable amount of power density of chips, thus achieving higher operating frequencies of chips. Through a full-system simulation, our results show that the water-immersion chip multiprocessors outperform the counterpart water-pipe cooled and oil-immersion chips by up to 14% and 4.5%, respectively, in terms of execution times of NAS Parallel Benchmarks. Michihiro Koibuchi, Ikki Fujiwara, Naoya Niwa, Tomohiro Totoki, Shoichi Hirasawa |
ICPP | 1 |
| 2019 | Designing High-Performance Interconnection Networks with Host-Switch GraphsabstractThis paper aims at establishing a method for designing high-performance network topologies to bridge a gap between theoretical and practical studies. To this end, we present a novel graph called a host-switch graph, which consists of host vertices and switch vertices with maximum degree 1 and$r$, respectively. This graph represents a network topology of a practical parallel/distributed computer system with host computers connected by$r$-port switches. We discuss important metrics for designing high-performance interconnection networks: the host-to-host average shortest path length (h-ASPL) and the bisection width (BiW). In particular, we explore a method for constructing host-switch graphs with low h-ASPL and high BiW that connect the fixed number of hosts via any number of$r$-port switches. We demonstrate that the number of switches that provides the minimum h-ASPL can mathematically be approximated, and the minimum number of switches that provides a certain BiW can experimentally be approximated. On the basis of the approximations, we propose a randomized algorithm for searching host-switch graphs. We then apply the graphs to interconnection networks and compare them with typical network topologies. As compared with the torus, the dragonfly, and the fat-tree, our networks attain higher performance and smaller power and costs. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Low-Reliable Low-Latency Networks Optimized for HPC Parallel ApplicationsabstractHigh-end network standards, such as 400GbE, have been introduced Forwarding Error Correction (FEC) for maintaining the same bit error rate (BER) as that in traditional low-bandwidth interconnection networks. However, FEC operation latency overhead surprisingly becomes higher than the sum of all the other switch operation overheads, e.g., routing computation and switch allocation. FEC operation latency overhead significantly degrades the performance of parallel applications in HPC systems. Instead, in this study, we exploit the low-latency network design using a Hamming code that does not provide rigid error-free communication. Since it is consistent with existing frame format based on standard Reed-Solomon RS(544,514) with DC(64b/66b) direct linecode and TC(256b/257b) transcode, respectively, the influences upon the other network layer design are limited. Interestingly, a large number of parallel applications can accept the BER in such a Hamming code. Since lowering such a BER improves switch operation latency, the proposed network using the Hamming code improves the execution time of NAS Parallel Benchmarks by 56% on average when compared to the counterpart RS-FEC networks. Truong Thao Nguyen, Hiroki Matsutani, Michihiro Koibuchi |
NCA | 3 |
| 2018 | AxNoC: Low-power Approximate Network-on-Chips using Critical-Path IsolationabstractVarious parallel applications, such as numerical convergent computation and multimedia processing, have intrinsic tolerance to inaccuracies that allow soft errors, i.e. bit flips, on a chip. However, existing Network-on-Chips (NoCs) guarantee error-free data transfer; thus, encountering limits to reduce the power consumption. In this context, we propose an approximate dual-voltage NoC, called AxNoC. An AxNoC router uses a per-flit look-ahead power management so that headers and important-data flits are perfectly transferred at a high voltage while the remaining flits may incur bit flips by decreasing the supply voltage. An AxNoC router isolates the critical path when the supply voltage is low since such a critical path is enabled only at high voltage. The critical path isolation enables low-voltage operation to work at the same operating frequency at high voltage. An AxNoC router was implemented using a 28nm process and the evaluation results illustrate its efficiency to reduce the power consumption reaching up to 43% while incurring a small area overhead that does not exceed 6.2%. We also demonstrate that AxNoC exhibits an acceptable accuracy illustrated in a sufficiently small geomean of error. Akram Ben Ahmed, Daichi Fujiki, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
NOCS | 4 |
| 2017 | Cable-geometric error-prone approach for low-latency interconnection networksabstractInterconnection network is a main concern in the architecture design of highly parallel systems such as high density data centers and supercomputers that reach millions of endpoints, e.g., 10M cores for Sunway TaihuLight system. As the number of endpoints of such systems has gradually increased to meet the higher computing and storage demand, the interconnection network is required to provide a low latency and high communication bandwidth, i.e., less than 1-microsecond latency across systems with a link bandwidth is greater than 100GB/s. In the low-latency design context, our primary aim is to provide a novel solution based on the technology-driven approaches: (1) cable-geometric small-world network topology with custom routing, and (2) the use of an FEC(forward error correction)-free error-prone mechanism for a high-speed low-reliable link. Both can present a logarithmic diameter low-radix network with low-latency and high-bandwidth switches, although approximate computing or ABFT (algorithm-based fault tolerance) design is required for parallel applications. Thao-Nguyen Nguyen, Michihiro Koibuchi |
CCGrid | 2 |
| 2017 | A Case for Uni-directional Network Topologies in Large-Scale ClustersabstractDesigning low-latency network topologies of switches is a key objective for next-generation large-scale clusters. Low latency is preconditioned on low hop counts, but existing network topologies have hop counts much larger than theoretical lower bounds. To alleviate this problem, we propose building network topologies based on uni-directional graphs that are known to have hop counts close to theoretical lower bounds. A practical difficulty with uni-directional topologies is switch-by-switch flow control, which we resolve by using hot-potato routing. Cycle-accurate network simulation experiments for various traffic patterns on uni-directional topologies show that hot-potato routing achieves performance comparable to that of conventional deadlock-free routing. Similar experiments are used to compare several uni-directional topologies to bi-directional topologies, showing that the former achieve significantly lower latency and higher throughput. We quantify end-to-end application performance for parallel application benchmarks via discrete-even simulation, showing that uni-directional topologies can lead to large application performance improvements over their bi-directional counterparts. Finally, we discuss practical issues for uni-directional topologies such as cabling complexity and cost, power consumption, and soft-error tolerance. Our results make a compelling case for considering uni-directional topologies for upcoming large-scale clusters. Michihiro Koibuchi, Tomohiro Totoki, Hiroki Matsutani, Hideharu Amano, Fabien Chaix, Ikki Fujiwara, Henri Casanova |
CLUSTER | 1 |
| 2017 | In-switch approximate processing: Delayed tasks management for MapReduce applicationsabstractIn MapReduce, the parallel processing performance is often limited by only a few compute nodes that delay to complete given tasks. Although various techniques have been invented to handle such stragglers, these techniques mostly impose a burden on master node to monitor the progress of all the compute nodes, resulting in a new bottleneck as the number of compute nodes increases. As an alternative approach, in this paper, we propose to move such straggler management burden from master node to network switch that connects the master and compute nodes, because all the information goes through the switch. More specifically, the proposed network switch monitors output packets from Map tasks to detect stragglers. When detected, the proposed switch generates a response instead of the straggler based on the outputs of the other normal Map tasks, so that Reduce tasks can be started without delay. We introduce some approximate techniques for the proxy computation and response at the switch; thus our switch is called "ApproxSW." We implement ApproxSW on NetFPGA-SUME board that has four 10Gbit Ethernet (10GbE) interfaces and a Virtex-7 FPGA. An experiment shows that the ApproxSW functions do not degrade the original 10GbE switch performance. We also analyze the accuracy of the proxy computation and response for stragglers and show that the proposed approximation based on task similarity achieves the best accuracy. Koya Mitsuzuka, Ami Hayashi, Michihiro Koibuchi, Hideharu Amano, Hiroki Matsutani |
FPL | 3 |
| 2017 | High-Bandwidth Low-Latency Approximate Interconnection NetworksabstractComputational applications are subject to various kinds of numerical errors, ranging from deterministic roundoff errors to soft errors caused by non-deterministic bit flips, which do not lead to application failure but corrupt application results. Non-deterministic bit flips are typically mitigated in hardware using various error correcting codes (ECC). But in practice, due to performance and cost concerns, these techniques do not guarantee error-free execution. On large-scale computing platforms, soft errors occur with non-negligible probability in RAM and on the CPU, and it has become clear that applications must tolerate them. For some applications, this tolerance is intrinsic as result quality can remain acceptable even in the presence of soft errors (e.g., data analysis applications, multimedia applications). Tolerance can also be built into the application, resolving data corruptions in software during application execution. By contrast, today's optical networks hold on to a rigid error-free standard, which imposes limits on network performance scalability. In this work we propose high-bandwidth, low-latency approximate networks with the following three features: (1) Optical links that exploit multi-level quadrature amplitude modulation (QAM) for achieving high bandwidth; (2) Avoidance of forward error correction (FEC), which makes optical link error-prone but affords lower latency; and (3) The use of symbol mapping coding between bit sequence and QAM to ensure data integrity that is sufficient for practical soft-error-tolerant applications. Discrete-event simulation results for application benchmarks show that approx networks achieve speedups up to 2.94 when compared to conventional networks. Daichi Fujiki, Kiyo Ishii, Ikki Fujiwara, Hiroki Matsutani, Hideharu Amano, Henri Casanova, Michihiro Koibuchi |
HPCA | 7 |
| 2017 | SINET5: A low-latency and high-bandwidth backbone network for SDN/NFV EraabstractSINET5 is a new 100-Gbps-based academic backbone network, which started full-scale operations in April 2016. It uses multi-protocol label switching-transport profile (MPLS-TP) systems and reconfigurable optical add-drop multiplexers (ROADMs) to create a nationwide network and has more than 50 backbone IP routers to provide a wide range of services, such as several virtual private network (VPN) services. It provides end-to-end data communications up to 100 Gbps throughput, minimized-latency, and software-defined networking (SDN)-friendly functions to researchers in every Japanese prefecture. SINET5 is also a platform for dynamic inter-cloud connections and network functions visualization (NFV) services. This paper brief review of the network architecture, and describes new featured services, SDN-oriented layer-2 on-demand VPN services, and NFV functions. Field test results for SINET5 performance are also reported. Takashi Kurimoto, Shigeo Urushidani, Kenjiro Yamanaka, Motonori Nakamura, Shunji Abe, Kensuke Fukuda, Michihiro Koibuchi, Hiroki Takakura, Shigeki Yamada, Yusheng Ji |
ICC | 8 |
| 2017 | HiRy: An Advanced Theory on Design of Deadlock-Free Adaptive Routing for Arbitrary TopologiesabstractRecently proposed irregular networks can reduce the latency for both on-chip and off-chip systems with a large number of computing nodes and thus can improve the performance of parallel application. However, these networks usually suffer from deadlocks in routing packets when using a naive minimal path routing algorithm. To solve this problem, we focus attention on a lately proposed theory that generalizes the turn model to maintain the network performance with deadlock-freedom. The theorems remain a challenge of applying themselves to arbitrary topologies including fully irregular networks. In this paper, we advance the theorems to completely general ones. To apply the idea of the turn model to arbitrary topologies, we introduce a concept of regions that define continuous directions of channels on an n-dimensional space. Moreover, we provide a feasible implementation of a deadlock-free routing method based on our advanced theorem. To reduce the latency and the number of required Virtual Channels (VCs) with this method, a heuristic approach is introduced to reduce the number of prohibited turns between channels. Experimental results show that the routing method based on our proposed theorem can improve the network throughput by up to 138 % compared to a conventional deterministic minimal routing method. Moreover, it can reduce the latency by up to 2.9 % compared to another fully adaptive routing method. Ryuta Kawano, Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICPADS | 4 |
| 2017 | Order/Radix Problem: Towards Low End-to-End Latency Interconnection NetworksabstractWe introduce a novel graph called a host-switch graph, which consists of host vertices and switch vertices. Using host-switch graphs, we formulate a graph problem called an order/radix problem (ORP) for designing low end-to-end latency interconnection networks. Our focus is on reducing the host-to-host average shortest path length (h-ASPL), since the shortest path length between hosts in a host-switch graph corresponds to the end-to-end latency of a network. We hence define ORP as follows: given order (the number of hosts) and radix (the number of ports per switch), find a host-switch graph with the minimum h-ASPL. We demonstrate that the optimal number of switches can mathematically be predicted. On the basis of the prediction, we carry out a randomized algorithm to find a host-switch graph with the minimum h-ASPL. Interestingly, our solutions include a host-switch graph such that switches have the different number of hosts. We then apply host-switch graphs to interconnection networks and evaluate them practically. As compared with the three conventional interconnection networks (the torus, the dragonfly, and the fat-tree), we demonstrate that our networks provide higher performance while the number of switches can decrease. Ryota Yasudo, Michihiro Koibuchi, Koji Nakano, Hiroki Matsutani, Hideharu Amano |
ICPP | 2 |
| 2017 | A Case of Electrical Circuit Switched Interconnection Network for Parallel ComputersabstractCircuit switching is a way to minimize network latency and maximize network bandwidth when a limited number of source-and-destination pairs exchange messages which are predictable. Although there are a large number of studies of optical circuit switching (OCS) on HPC systems and datacenters, it is still not mature. In this context, we explore the use of electrical circuit switching (ECS) for the low-latency purpose on HPC systems and datacenters. ECS has the same link bandwidth as existing electrical packet switched networks, and inherits quick update of input-and-output connections from electrical switches. We develop a network topology generator for ECS to minimize the number of time slots optimized to target applications whose traffic patterns are predictable. By performing a quantitative discrete-event simulation, we present that an ideal ECS network outperforms counterpart EPS networks. Evaluation results show that the minimum necessary number of slots (MNNS) can be reduced to a small number in a generated topology while keeping resource amount less than that in a standard mesh network. Yao Hu 0005, Tomohiro Kudoh, Michihiro Koibuchi |
PDCAT | 3 |
| 2017 | Scalable Networks-on-Chip with Elastic Links Demarcated by Decentralized RoutersabstractAs the number of cores on a chip increases, Networks-on-Chip (NoCs) that connect many cores would face long links to reduce hop counts. The long links become bottlenecks in terms of both energy and RC delays as technology advances. To alleviate the negative impact of long links, we propose decentralized routers for NoCs. A decentralized router consists of multiple submodules that are positioned on a link, and hence the long links are segmented. Furthermore, we illustrate the design of an entire network that uses decentralized routers to obtain a good tradeoff between hop counts and wire delays per hop. Decentralized routers are effective especially in high-radix topologies, such as the flattened butterfly, and energy-delay product is reduced by greater than 60 percent. As NoCs become larger and more complex, the benefit of the decentralized routers will become more significant. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
IEEE Trans. Computers | 3 |
| 2017 | Distributed Shortcut Networks: Low-Latency Low-Degree Non-Random Topologies Targeting the Diameter and Cable Length Trade-OffabstractLow communication latency becomes a main concern in highly parallel computers and supercomputers that reach millions of processing cores. Random network topologies are better suited to achieve low average shortest path length and low diameter in terms of the hop counts between nodes. However, random topologies lead to two problems: (1) increased aggregate cable length on a machine room floor that would become dominant for communication latency in next-generation custom supercomputers, and (2) high routing complexity that typically requires a routing table at each node (e.g., topology-agnostic deadlock-free routing). In this context, we first propose low-degree non-random topologies that exploit the small-world effect, which has been well modeled by some random network models. Our main idea is to carefully design a set of various-length shortcuts that keep the diameter small while maintaining a short cable length for economical passive electric cables. We also propose custom routing that uses the regularity of the various-length shortcuts. Our experimental graph analyses show that our proposed topology has low diameter and low average shortest path length, which are considerably better than those of the counterpart 3-D torus and are near to those of a random topology with the same average degree. The proposed topology has average cable length drastically shorter than that of the counterpart random topology, which leads to low cost of interconnection networks. Our custom routing takes non-minimal paths to provide lower zero-load latency than the minimal custom routings on different counterpart topologies. Our discrete-event simulation results using SimGrid show that our proposed topology is suitable for applications that have irregular communication patterns or non-nearest neighbor collective communication patterns. Nguyen T. Truong, Ikki Fujiwara, Michihiro Koibuchi, Khanh-Van Nguyen |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | ACRO: Assignment of channels in reverse order to make arbitrary routing deadlock-freeabstractDistributed routing methods with small routing tables are scalable design on irregular networks for large-scale High Performance Computing (HPC) systems. Recently proposed compact routing methods, however, do not guarantee deadlock-freedom. Cyclic channel dependencies on arbitrary routing are typically removed with multiple Virtual Channels (VCs). However, challenges still remain to provide good trade-offs between a number of required VCs and a time complexity of an algorithm for assignment of VCs to paths. In this work, a novel algorithm ACRO is proposed for enriching arbitrary routing functions with deadlock-freedom with a reasonable number of VCs and a time complexity. Experimental results show that ACRO can reduce the average number of required VCs by up to 63% compared with the conventional algorithm that has the same time complexity. Furthermore, ACRO reduces a time complexity by a factor of O(|N| · log|N|) compared with that of the other conventional algorithm that needs almost the same number of VCs. Ryuta Kawano, Hiroshi Nakahara, Seiichi Tade, Ikki Fujiwara, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICIS | 6 |
| 2016 | HPC Job Mapping over Reconfigurable Wireless LinksabstractWireless supercomputers and datacenters with 60GHz radio or free-space optics (FSO) have been proposed so that a diverse application workload can be better supported by changing network topologies by swapping the endpoints of wireless links. In this study we proposed the use of such wireless links for the purpose of improving job mapping. We investigated various trade-offs of the number of wireless links, time overhead of wireless link reconfiguration, topology embedding and job sizes. Our simulation results demonstrate that the wired job mapping heavily degrades the system utilization of supercomputers and datacenters under a conventional fixed network topology. By contrast, wireless interconnection networks can have an ideal job mapping by directly reconnecting non-neighboring computing nodes. It improves the system utilization by up to 17.7% for user jobs on a supercomputer and thus can shorten the whole service time especially for dealing with dozens of intensively incoming jobs. Furthermore, we confirmed that either workload or scheduling policy does not impact the fact that the ideal job mapping on wireless supercomputers outperforms that on wired networks in terms of system utilization and whole service time. Finally, our evaluation shows that a constrained and reasonable more use of partial wireless links can achieve shorter queuing length and time. Yao Hu 0005, Ikki Fujiwara, Michihiro Koibuchi |
CCGrid | 3 |
| 2016 | Deterministic Construction of Regular Geometric Graphs with Short Average Distance and Limited Edge Length
Satoshi Fujita, Koji Nakano, Michihiro Koibuchi, Ikki Fujiwara |
ICA3PP | 3 |
| 2016 | Randomly Optimized Grid Graph for Low-Latency Interconnection NetworksabstractIn this work we present randomly optimized grid graphs that maximize the performance measure, such as diameter and average shortest path length (ASPL), with subject to limited edge length on a grid surface. We also provide theoretical lower bounds of the diameter and the ASPL, which prove optimality of our randomly optimized grid graphs. We further present a diagonal grid layout that significantly reduces the diameter compared to the conventional one under the edge-length limitation. We finally show their applications to three case studies of off-and on-chip interconnection networks. Our design efficiently improves their performance measures, such as end-to-end communication latency, network power consumption, cost, and execution time of parallel benchmarks. Koji Nakano, Daisuke Takafuji, Satoshi Fujita, Hiroki Matsutani, Ikki Fujiwara, Michihiro Koibuchi |
ICPP | 6 |
| 2016 | Suitability of the Random Topology for HPC ApplicationsabstractWith each technology improvement, parallel systems get larger, and the impact of interconnection networks becomes more prominent. Random topologies and their variants received more and more attention lately due to their low diameter, low average shortest path length and high scalability. However, existing supercomputers still prefer torus and fat-tree topologies, because a number of existing parallel algorithms are optimized for them and the interconnect implementation is more straight-forward in terms of floor layout. In this paper, we investigate the performance of traditional and emerging parallel workloads on these network topologies, using a event-discrete simulation called SimGrid. We observe that random topology is better for Fourier Transform (FT), Graph500, Himeno benchmarks, and its improvement over the counterpart torus is 18 percent in average. Through this study, our recommendation is to use random topology in current and future supercomputers for these scientific and big-data analysis parallel applications. Fabien Chaix, Ikki Fujiwara, Michihiro Koibuchi |
PDP | 3 |
| 2016 | Randomizing Packet Memory Networks for Low-Latency Processor-Memory CommunicationabstractThree-dimensional stacked memory is considered to be one of the innovative elements for the next-generation computing system, for it provides high bandwidth and energy efficiency. Particularly, packet routing ability of Hybrid Memory Cubes (HMCs) enables new interconnects for the memories, giving flexibility to its topological design space. Since memory-processor communication is latency-sensitive, our challenge is to alleviate latency of the memory interconnection network, which is subject to high overheads from hop-count increase. Interestingly, random network topologies are known to have remarkably low diameter that is even comparable to theoretical Moore graph. In this context, we first propose to exploit the random topologies for the memory networks. Second, we also propose several optimizations to leverage the random topologies to be further adaptive to the latency-sensitive memory-processor communication: communication path length based selection, deterministic minimal routing, and page-size granularity memory mapping. Finally, we present interesting results of our evaluation: the random networks with universal memory access outperformed non-random networks of which memory access was optimally localized. Daichi Fujiki, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
PDP | 3 |
| 2016 | Efficient 3-D Bus Architectures for Inductive-Coupling ThruChip InterfacesabstractWireless 3-D network-on-chips (NoCs) with inductive-coupling ThruChip interfaces provide a large degree of flexibility for customizing the number of arbitrary chips in a package after chips have been fabricated. To simplify the vertical communication interfaces, static time division multiple access (TDMA) is used for the vertical broadcast buses, while arbitrary or customized topologies can be used for the intrachip network. This paper proposes two techniques to break through the simple static TDMA-based vertical buses while maintaining a simple communication interface. The first technique is headfirst sliding (HS) routing to reduce the waiting time for acquiring the communication time-slot. HS routing selects the best vertical bus based on the current time, taking advantage of static TDMA. The second technique extends carrier sense multiple access with collision detection (CSMA/CD) for vertical broadcast buses. We introduce a packet collision detection technique for inductive-coupling buses and propose two retransmission strategies to reduce the waiting time for packet retransmissions caused by collisions. Network simulation results show that HS routing reduces the communication latency by 39.1% compared with the conventional static TDMA bus-based 3-D NoC that uses the shortest path routing. The proposed CSMA/CD bus also improves the latency by 52.5% and throughput by 34.1%. The full-system simulation results show that HS routing and the proposed CSMA/CD technique reduce the application execution time accordingly while maintaining the average flit transfer energy overhead modest. Takahiro Kagami, Hiroki Matsutani, Michihiro Koibuchi, Yasuhiro Take, Tadahiro Kuroda, Hideharu Amano |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | A metamorphotic Network-on-Chip for various types of parallel applicationsabstractA metamorphotic Network-on-Chip (NoC) architecture is proposed in order to customize for performance or energy consumption on a per-application basis. Adding reconfigurability on conventional topologies has been studied so far especially for application workloads that can be statically analyzed. In this context, we propose such a platform to take care of both the static and the dynamic cases where application workloads cannot be statically analyzed while performance or energy constraints are given. Our metamorphotic NoC reconfigures its topology, routing, operating frequency, and supply voltage based on the following three modes. 1) Regular mode uses a traditional mesh topology for neighboring communications. As the link length is short and uniform, it can be operated at a higher frequency and higher voltage, while a long-range communication increases the path length. 2) Random mode uses a random topology for unknown workloads to reduce the path length by exploiting the small-world effect. As the path length is reduced but the wire delay is increased, it is intended for a lower operating frequency and lower voltage. 3) Custom mode uses an optimized topology for given workloads. To support Random and Custom modes, assembled multiplexers are embedded into the metamorphotic NoC. Random and Regular/Custom modes are generated by randomly or selectively reconfiguring these multiplexers, respectively, based on the performance or energy constraints. This paper explores the design space of assembled multiplexers and provides a reasonable design recommendation through a graph analysis. It is demonstrated based on experimental results on the area overhead, operating frequency, network performance, and energy consumption. The results show that Regular mode can operate at 1.27GHz and Random mode can reduce the average network latency by 19.6% and the energy consumption by 44.2% compared with a traditional NoC that has mesh topology with little overhead. Custom mode can reduce them as well as Random mode. Seiichi Tade, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
ASAP | 4 |
| 2015 | Augmenting low-latency HPC network with free-space optical linksabstractVarious network topologies can be used for deploying High Performance Computing (HPC) clusters. The network topology, which connects switches In cabinets on a machine room floor, is typically defined once and for all at system deployment time. For a diverse application workload, there are downsides to having a single wired topology. In this work, we propose using free-space optics (FSO) in large-scale systems so that a diverse application workload can be better supported. A high-density layout of FSO terminals on top of the cabinets is determined that allows line-of-sight communication between arbitrary cabinet pairs. We first show that our proposal reduces both end-to-end network latency and total cable length when compared to a wired topology. We then demonstrate that the use of FSO links improves the embedding/partitioning capabilities of a wired topology. More specifically, we show that a recently proposed random low-latency topology can be augmented with a reasonable number of FSO links to support multiple k-ary n-cube and fat tree embedded topologies. Finally, we investigate power-aware on/off link regulation techniques and show how adding/reconfiguring FSO links leads to both performance and power efficiency improvements. Ikki Fujiwara, Michihiro Koibuchi, Tomoya Ozaki, Hiroki Matsutani, Henri Casanova |
HPCA | 2 |
| 2015 | On-Chip Decentralized Routers with Balanced Pipelines for Avoiding Interconnect BottleneckabstractTechnology scaling makes designers face difficulties dealing with wire delay of long global interconnects, especially for high-radix networks. In this context, we propose decentralization of on-chip packet routers. A decentralized router consists of submodules, each of which has particular functionality and they are scattered on a link, thereby long wires are segmented. Our starting point is from a conventional router architecture, and we illustrate four case studies to generalize our proposal. We also propose a new buffer design and how to balance pipelines of a router. A proof-of-concept is shown in 28-nm process technology. Our results demonstrate that the decentralization of an on-chip router enables Link Traversal (LT) stages to be eliminated, and the critical path delay is improved by up to 45% with the reduced area compared with a conventional router. As technology advances, the benefit of the decentralized routers become more substantial in the nano-scale era. Ryota Yasudo, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tadao Nakamura |
NOCS | 3 |
| 2015 | Optimized Core-Links for Low-Latency NoCsabstractIn recent many-core architectures, the number of cores has been steadily increasing and thus the network latency between cores becomes an important issue for parallel application programs. Because packet-switched network structures are widely used for core-to-core communications, a topology among cores has a major impact on the network latency. It has been reported that a small-world Network-on-Chip that adds links between randomly-selected routers on a regular router topology is effective for reducing the network latency. In this study, we extend this framework by connecting multiple links between a single core and quasi-optimally selected neigh boring routers to form multiple links from each core on a 2D MESH router topology. Results obtained by a flit-level discrete event simulator show that our optimized core-link topologies can achieve the average latency up to 48% lower than that of baseline topologies. Furthermore, full-system CMP simulation results show that by using optimized core-links we can improve the application execution time on the NAS Parallel Benchmarks by up to 10.1%. Ryuta Kawano, Seiichi Tade, Ikki Fujiwara, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
PDP | 6 |
| 2015 | Swap-And-Randomize: A Method for Building Low-Latency HPC InterconnectsabstractRandom network topologies have been proposed to create low-diameter, low-latency interconnection networks in large-scale computing systems. However, these topologies are difficult to deploy in practice, especially when re-designing existing systems, because they lead to increased total cable length and cable packaging complexity. In this work we propose a new method for creating random topologies without increasing cable length: randomly swap link endpoints in a non-random topology that is already deployed across several cabinets in a machine room. We quantitatively evaluate topologies created in this manner using both graph analysis and cycle-accurate network simulation, including comparisons with non-random topologies and previously-proposed random topologies. Ikki Fujiwara, Michihiro Koibuchi, Hiroki Matsutani, Henri Casanova |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Low-latency wireless 3D NoCs via randomized shortcut chipsabstractIn this paper, we demonstrate that we can reduce the communication latency significantly by inserting a fraction of randomness into a wireless 3D NoC (where CMOS wireless links are used for vertical inter-chip communication) when considering the physical constraints of the 3D design space. Towards this end, we consider two cases, namely 1) replacing existing horizontal 2D links in a wireless 3D NoC with randomized shortcut NoC links and 2) enabling full connectivity by adding a randomized NoC layer to a wireless 3D platform with partial or no horizontal connectivity. Consequently, the packet routing is optimized by exploiting both the existing and the newly added random NoC. At the same time, by adding randomly wired shortcut NoCs to a wireless 3D platform, a good balance can be established between the modularity of the design and the minimum randomness needed to achieve low latency, and experimental results show that by adding a random NoC chip to wireless 3D CMPs without built-in horizontal connectivity, the communication latency can be reduced by as much as 26.2% when compared to adding a 2D mesh NoC. Also, the application execution time and average flit transfer energy can be improved accordingly. Hiroki Matsutani, Michihiro Koibuchi, Ikki Fujiwara, Takahiro Kagami, Yasuhiro Take, Tadahiro Kuroda, Paul Bogdan, Radu Marculescu, Hideharu Amano |
DATE | 2 |
| 2014 | Layout-aware expandable low-degree topologyabstractSystem expandability becomes a major concern for highly-parallel computers and datacenters, because their number of nodes gradually increases year by year. In this context we propose a low-degree expandable topology and its floor layout in which a cabinet or node set can be newly inserted by connecting short cables to a single existing cabinet. Our graph analysis shows that the proposed topology has low diameter, low average shortest path length and short aggregate cable length comparable to existing topologies with the same degree. When incrementally adding nodes and cabinets to the proposed topology, its diameter and average shortest path length increase modestly. Flit-level network simulation results show that the proposed topology has lower latency for three synthetic traffic patterns as expected from graph analysis. Our event-driven network simulation results show that the proposed topology provides a comparable performance to 2-D torus even for bandwidth-sensitive parallel applications. Nguyen T. Truong, Van K. Nguyen, Nhat T. X. Le, Ikki Fujiwara, Fabien Chaix, Michihiro Koibuchi |
ICPADS | 6 |
| 2014 | Skywalk: A Topology for HPC Networks with Low-Delay SwitchesabstractWith low-delay switches on the horizon, end-to-end latency in large-scale High Performance Computing (HPC) interconnects will be dominated by cable delays. In this context we define a new network topology, Skywalk, for deploying low-latency interconnects in upcoming HPC systems. Skywalk uses randomness to achieve low latency, but does so in a way that accounts for the physical layout of the topology so as to lead to further cable length and thus latency reductions. Via graph analysis and discrete-event simulation we show that Skywalk compares favorably (in terms of latency, cable length, and throughput) to traditional low-degree torus and moderate-degree hypercube topologies, to high-degree fully-connected Dragonfly topologies, to the HyperX topology, and to recently proposed fully random topologies. Ikki Fujiwara, Michihiro Koibuchi, Hiroki Matsutani, Henri Casanova |
IPDPS | 2 |
| 2014 | Hierarchical Network Coding for Collective Communication on HPC InterconnectsabstractNetwork bandwidth is a performance concern especially for collective communication because the bisection bandwidth of recent supercomputers is far less than their full bisection bandwidth. In this context we propose to exploit the use of a network coding technique to reduce the number of unicasts and the size of transferred data generated by latency-sensitive collective communication in supercomputers. Our proposed network coding scheme has a hierarchical multicasting structure with intra-group and inter-group unicasts. Quantitative analysis show that the aggregate path hop counts by our hierarchical network coding decrease as much as 94% when compared to conventional unicast-based multicasts. We validate these results by cycle-accurate network simulations. In 1,024-switch networks, the network reduces the execution time of collective communication as much as 64%. We also show that our hierarchical network coding is beneficial for any packet size. Ahmed Shalaby 0001, Mohamed El-Sayed Ragab, Victor Goulart, Ikki Fujiwara, Michihiro Koibuchi |
PDP | 5 |
| 2014 | 3D NoC with Inductive-Coupling Links for Building-Block SiPsabstractA wireless 3D NoC architecture is described for building-block SiPs, in which the number of hardware components (or chips) in a package can be changed after chips have been fabricated. The architecture uses inductive-coupling links that can connect more than two examined dies without wire connections. Each chip has data transceivers for the uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a vertical unidirectional ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring, we use bubble flow control, which does not rely on the conventional VC-based deadlock avoidance mechanism. In addition, we propose a bidirectional communication scheme to form a bidirectional ring network by using the inductive-coupling transceivers that can dynamically change the communication modes, such as TX, RX, and Idle modes. This paper illustrates the inductive-coupling transceiver circuits, which can carry high data transfer rates of up to 8 Gbps per channel, for the wireless 3D NoC. It also illustrates an implementation of a wireless 3D NoC that has on-chip routers and transceivers implemented with a 65 nm process in order to show the feasibility of our proposal. The vertical bubble flow control and conventional VC-based approach on the uni- and bidirectional ring networks are compared with the vertical broadcast bus in terms of throughput, hardware amount, and application performance using a full system multiprocessor simulator. The results show that the proposed bidirectional communication scheme efficiently improves application performance without adding any inductive-coupling transceivers. In addition, the proposed vertical bubble flow network outperforms the conventional VC-based approach by 7.9-12.5 percent with a 33.5 percent smaller router area for building-block SiPs connecting up to eight chips. Yasuhiro Take, Hiroki Matsutani, Daisuke Sasaki, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
IEEE Trans. Computers | 4 |
| 2013 | A case for wireless 3D NoCs for CMPsabstractInductive-coupling is yet another 3D integration technique that can be used to stack more than three known-good-dies in a SiP without wire connections. We present a topology-agnostic 3D CMP architecture using inductive-coupling that offers great flexibility in customizing the number of processor chips, SRAM chips, and DRAM chips in a SiP after chips have been fabricated. In this paper, first, we propose a routing protocol that exchanges the network information between all chips in a given SiP to establish efficient deadlock-free routing paths. Second, we propose its optimization technique that analyzes the application traffic patterns and selects different spanning tree roots so as to minimize the average hop counts and improve the application performance. Hiroki Matsutani, Paul Bogdan, Radu Marculescu, Yasuhiro Take, Daisuke Sasaki, Hao Zhang 0020, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
ASP-DAC | 7 |
| 2013 | Layout-conscious random topologies for HPC off-chip interconnectsabstractAs the scales of parallel applications and platforms increase the negative impact of communication latencies on performance becomes large. Random network topologies can be used to achieve low hop counts between nodes and thus low latency. However, random topologies lead to increased aggregate cable length and cable packaging complexity on a machine room floor. In this work we propose two new methods for generating random topologies and their physical layout on a floorplan: randomize links after optimizing the physical layout, or optimize the layout after randomizing links. The first method randomly swaps link endpoints for a given non-random topology for which a good physical layout is known. The resulting topology has the same cable length and cable packaging as the original topology, but achieves lower communication latency. The second method creates a random topology with random links picked so that they will not lead to a long physical cable length, and then solves a constrained optimization problem to compute a physical layout that minimizes aggregate cable length. We quantitatively compare these two methods using both graph analysis and cycle-accurate network simulation, including comparisons with previously proposed random topologies and non-random topologies. Michihiro Koibuchi, Ikki Fujiwara, Hiroki Matsutani, Henri Casanova |
HPCA | 1 |
| 2013 | A Routing Strategy for Inductive-Coupling Based Wireless 3-D NoCs by Maximizing Topological Regularity
Daisuke Sasaki, Hao Zhang 0020, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICA3PP (2) | 4 |
| 2013 | Distributed Shortcut Networks: Layout-Aware Low-Degree Topologies Exploiting Small-World EffectabstractLow communication latency becomes a main concern in highly parallel computers and supercomputers. Random network topologies are best to achieve low average shortest path length and low diameter in hop counts between nodes and thus low communication latency. However, random topologies lead to a problem of increased aggregate cable length on a machine room floor. In this context we propose low-degree non-random topologies that exploit the small-world effect, which has been typically well modeled by some random network models. Our main idea is to carefully design a set of various-length shortcuts that keep the diameter small while maintain an economical cable length. Our experimental graph analysis showed that our proposed topology has low diameter and low average shortest path length, which is considerably better than those of a counterpart 2-D torus and is near to those of a counterpart random topology with the same average degree. Meanwhile, the proposed topology has average cable length drastically shorter than that of the counterpart random topology. Our cycle-accurate network simulation results show that the proposed topology has lower latency by 15% and almost the same throughput when compared to torus with the same degree. Van K. Nguyen, Nhat T. X. Le, Ikki Fujiwara, Michihiro Koibuchi |
ICPP | 4 |
| 2013 | PopCache: Cache more or less based on content popularity for information-centric networkingabstractDue to a mismatch between downloading and caching content, the network may not gain significant benefit from the sophisticated in-network caching of information-centric networking (ICN) architectures by using a basic caching mechanism. This paper aims to seek an effective caching decision policy to improve the content dissemination in ICN. We propose PopCache-a caching decision policy with respect to the content popularity-that allows an individual ICN router to cache content more or less in accordance with the popularity characteristic of the content. We propose an analytical model to evaluate the performance of different caching decision policies in terms of the server-hit rate and expected round-trip time. The analysis confirmed by simulation results shows that PopCache yields the lowest expected round-trip time compared with three benchmark caching decision policies, i.e., the always, fixed probability and path-capacity-based probability, and PopCache provides the server-hit rate comparable to the lowest ones. Kalika Suksomboon, Saran Tarnoi, Yusheng Ji, Michihiro Koibuchi, Kensuke Fukuda, Shunji Abe, Motonori Nakamura, Michihiro Aoki, Shigeo Urushidani, Shigeki Yamada |
LCN | 4 |
| 2013 | Headfirst sliding routing: A time-based routing scheme for bus-NoC hybrid 3-D architectureabstractA contact-less approach that connects chips in vertical dimension has a great potential to customize components in 3-D chip multiprocessors (CMPs), assuming card-style components inserted to a single cartridge communicate each other wirelessly using inductive-coupling technology. To simplify the vertical communication interfaces, static Time Division Multiple Access (TDMA) is used for the vertical broadcast buses, while arbitrary or customized topologies can be used for intra-chip networks. In this paper, we propose the Headfirst sliding routing scheme to overcome the simple static TDMA-based vertical buses. Each vertical bus grants a communication time-slot for different chips at the same time periodically, which means these buses work with different phases. Depending on the current time, packets are routed toward the best vertical bus (elevator) just before the elevator acquires its communication time-slot. Network simulations show that Headfirst sliding routing reduces the communication latency by up to 32.7%, and full-system CMP simulations show that it reduces application execution time by 9.9%. Synthesis results show that the area and critical path delay overheads are modest. Takahiro Kagami, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
NOCS | 3 |
| 2012 | A multi-Vdd dynamic variable-pipeline on-chip router for CMPsabstractWe propose a multi-voltage (multi-Vdd) variable pipeline router to reduce the power consumption of Network-on-Chips (NoCs) designed for chip multi-processors (CMPs). Our multi-Vdd variable pipeline router adjusts its pipeline depth (i.e., communication latency) and supply voltage level in response to the applied workload. Unlike dynamic voltage and frequency scaling (DVFS) routers, the operating frequency is the same for all routers throughout the CMP; thus, there is no need to synchronize neighboring routers working at different frequencies. In this paper, we implemented the multi-Vdd variable pipeline router, which selects two supply voltage levels and pipeline modes, using a 65nm CMOS process and evaluated it using a full-system CMP simulator. Evaluation results show that although the application performance degraded by 1.0% to 2.1%, the standby power of NoCs reduced by 10.4% to 44.4%. Hiroki Matsutani, Yuto Hirata, Michihiro Koibuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
ASP-DAC | 3 |
| 2012 | A case for random shortcut topologies for HPC interconnectsabstractAs the scales of parallel applications and platforms increase the negative impact of communication latencies on performance becomes large. Fortunately, modern High Performance Computing (HPC) systems can exploit low-latency topologies of high-radix switches. In this context, we propose the use of random shortcut topologies, which are generated by augmenting classical topologies with random links. Using graph analysis we find that these topologies, when compared to non-random topologies of the same degree, lead to drastically reduced diameter and average shortest path length. The best results are obtained when adding random links to a ring topology, meaning that good random shortcut topologies can easily be generated for arbitrary numbers of switches. Using flit-level discrete event simulation we find that random shortcut topologies achieve throughput comparable to and latency lower than that of existing non-random topologies such as hypercubes and tori. Finally, we discuss and quantify practical challenges for random shortcut topologies, including routing scalability and larger physical cable lengths. Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano, D. Frank Hsu, Henri Casanova |
ISCA | 1 |
| 2012 | Cabinet Layout Optimization of Supercomputer Topologies for Shorter Cable LengthabstractAs the scales of supercomputers increase total cable length becomes enormous, e.g., up to thousands of kilometers. Recent high-radix switches with dozens of ports make switch layout and system packaging more complex. In this study, we study the optimization of the physical layout of topologies of switches on a machine room floor with the goal of reducing cable length. For a given topology, using graph clustering algorithms, we group switches logically into cabinets so that the number of inter-cabinet cables is small. Then, we map the cabinets onto a physical floor space so as to minimize total cable length. This is done by modeling and optimizing the mapping problem as a facility location problem. Our evaluation results show that, when compared to standard clustering/mapping approaches and for popular network topologies, our clustering approach can reduce the number of inter-cabinet cables by up to 40.3% and our mapping approach can reduce the inter-rack cable length by up to 39.6%. Ikki Fujiwara, Michihiro Koibuchi, Henri Casanova |
PDCAT | 2 |
| 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing AlgorithmsabstractMost standard cluster interconnect technologies are flexible with respect to network topology. This has spawned a substantial amount of research on topology-agnostic routing algorithms, which make no assumption about the network structure, thus providing the flexibility needed to route on irregular networks. Actually, such an irregularity should be often interpreted as minor modifications of some regular interconnection pattern, such as those induced by faults. In fact, topology-agnostic routing algorithms are also becoming increasingly useful for networks on chip (NoCs), where faults may make the preferred 2D mesh topology irregular. Existing topology-agnostic routing algorithms were developed for varying purposes, giving them different and not always comparable properties. Details are scattered among many papers, each with distinct conditions, making comparison difficult. This paper presents a comprehensive overview of the known topology-agnostic routing algorithms. We classify these algorithms by their most important properties, and evaluate them consistently. This provides significant insight into the algorithms and their appropriateness for different on- and off-chip environments. José Flich, Tor Skeie, Andres Mejia, Olav Lysne, Pedro López 0001, Antonio Robles, José Duato, Michihiro Koibuchi, Tomas Rokicki, José Carlos Sancho |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2011 | A vertical bubble flow network using inductive-coupling for 3-D CMPsabstractA wireless 3-D NoC architecture for CMPs, in which the number of processor and cache chips stacked in a package can be changed after the chip fabrication, is proposed by using the inductive coupling technology that can connect more than two known-good-dies without wire connections. Each chip has data transceivers for uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a single vertical ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring network, we use the bubble flow control which is more flexible and efficient compared to the conventional VC-based deadlock avoidance. We implemented a real 3-D chip that has on-chip routers and inductive-coupling data transceivers using a 65nm process in order to show the feasibility of our proposal. The vertical bubble flow control is compared with the conventional VC-based approach and vertical bus in terms of the throughput, hardware amount, and application performance using a full system CMP simulator. The results show that the proposed vertical bubble flow network outperforms the VC-based approach by 7.9%-12.5% with a 33.5% smaller router area. Hiroki Matsutani, Yasuhiro Take, Daisuke Sasaki, Masayuki Kimura, Yuki Ono, Yukinori Nishiyama, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano |
NOCS | 7 |
| 2011 | A Dynamic Link-Width Optimization for Network-on-ChipabstractNetwork-on-Chip (NoC) is considered to be a promising approach to implement many-core systems and a large number of on-chip router optimization studies have been proposed. In this paper, we propose to dynamically adjust link-width of each port on a router optimized to spatially biased traffic. Different from the previous No Coptimization approaches, in which the optimization is almost performed in the NoC design step, the proposed method achieves a dynamical link-width optimization at run-time. Daihan Wang, Michihiro Koibuchi, Tomohiro Yoneda, Hiroki Matsutani, Hideharu Amano |
RTCSA (2) | 2 |
| 2011 | An analytical network performance model for SIMD processor CSX600 interconnects
Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Kenichi Miura, Hideharu Amano |
J. Syst. Archit. | 2 |
| 2011 | Prediction Router: A Low-Latency On-Chip Router Architecture with Multiple PredictorsabstractMulti and many-core applications are sensitive to interprocessor communication latencies, suggesting the need for low-latency on-chip networks. We propose a low-latency router architecture that predicts the output channel to be used by the next packet transfer and speculatively completes the switch arbitration to reduce communication latency. The packets coming into the prediction routers are transferred without waiting for the routing computation and switch arbitration if the prediction hits. Thus, the primary concern for reducing communication latency is the hit rates of the prediction algorithms, which vary based on network environments, such as the network topology, routing algorithm, and traffic pattern. Although typical low-latency routers that skip one or more pipeline stages use a bypass data path that is based on a static or single bypassing policy (e.g., accelerating the packets moving in the same dimension), our prediction router architecture predictively forwards packets based on the prediction algorithm selected from among several candidates in response to the network environment. We analyze the prediction hit rates of five prediction algorithms on meshes, tori, fat trees, and Spidergons. Then, we present four case studies, each of which assumes different many-core architectures. We implemented the prediction routers for each case study by using a 45 nm CMOS process, and evaluated them in terms of the prediction hit rate, zero-load latency, hardware amount, and energy consumption. A typical prediction router with two or three predictors shows that although the area and energy are increased by 4.8-12.0 percent and 5.3 percent, respectively, up to 89.8 percent of the prediction hit rate is achieved in real applications, which provides favorable trade-offs between modest hardware/energy overheads and significant latency saving. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tsutomu Yoshinaga |
IEEE Trans. Computers | 2 |
| 2011 | Performance, Area, and Power Evaluations of Ultrafine-Grained Run-Time Power-Gating Routers for CMPsabstractThis paper proposes the ultrafine-grained run-time power gating of on-chip routers, in which the power supply to each router component (e.g., virtual-channel buffer, virtual-channel multiplexer, and crossbar multiplexer and output latch) can be individually controlled based on the applied workload. Since only the router components that are transferring a packet are activated, the leakage power of the on-chip network can be reduced to a near-optimal level. However, such techniques inherently increase the communication latency and degrade the application performance, since a certain amount of wakeup latency is required to activate the sleeping components. To mitigate this wakeup latency, an early wakeup method that can preliminarily detect the next packet arrival and activate the corresponding components is essential. We designed and implemented an ultrafine-grained power-gating router using a commercial 65 nm process. We propose four early wakeup methods and combine them with the power-gating router. The proposed router with the early wakeup methods is evaluated in terms of its application performance, area overhead, and leakage power reduction taking into account the on/off energy overhead. The simulation results showed that it reduces the leakage power by 54.4-59.9% on average even when the application programs are fully running, at the expense of 4.6% of the area and 0.7-3.7% of the performance overheads when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | A Switch-Tagged Routing Methodology for PC Clusters with VLAN EthernetabstractEthernet has been used for connecting hosts in PC clusters, besides its use in local area networks. Although a layer-2 Ethernet topology is limited to a tree structure because of the need to avoid broadcast storms and deadlocks of frames, various deadlock-free routing algorithms on topologies that include loops suitable for parallel processing can be employed by the application of IEEE 802.1Q VLAN technology. However, the MPI communication libraries used in current PC clusters do not always support tagged VLAN technology; therefore, at present, the design of VLAN-based Ethernet cannot be applied to such PC clusters. In this study, we propose a switch-tagged routing methodology in order to implement various deadlock-free routing algorithms on such PC clusters by using at most the same number of VLANs as the degree of a switch. Since the MPI communication libraries do not need to perform VLAN operations, the proposed methodology has advantages in both simple host configuration and high portability. In addition, when it is used with on/off and multispeed link regulation, the power consumption of Ethernet switches can be reduced. Evaluation results using NAS parallel benchmarks showed that the performance of the topologies that include loops using the proposed methodology was comparable to that of an ideal one-switch (full crossbar) network, and the torus topology in particular had up to a 27 percent performance improvement compared with a tree topology with link aggregation. Michihiro Koibuchi, Tomohiro Otsuka, Tomohiro Kudoh, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2010 | Stabilizing Path Modification of Power-Aware On/Off Interconnection NetworksabstractPower saving is required for interconnects of modern PC clusters as well as the performance improvement. To reduce the power consumption of switches with maintaining the performance, on/off link regulations that activate and deactivate the links based on the traffic load have been widely developed in interconnection networks. Depending on which operation is selected, link activation or deactivation, the available network resources are changed, thus requiring paths to be reconfigured. To maintain deadlock freedom of packet transfers, connectivity, and performance during the path changes, we propose to apply dynamic reconfiguration techniques that process packet transfer uninterruptedly to power-aware on/off interconnection networks. The dynamic network reconfiguration techniques stabilize the update of paths that are quite crucial to use power-aware on/off link techniques in interconnects of PC clusters. We investigate the performance and behavior of network reconfiguration technique as soon as the link activation or deactivation occurs. Evaluation results show that the simple dynamic reconfiguration techniques slightly reduce the peak packet latency and reconfiguration time of the change compared with existing static reconfiguration in on/off interconnection networks. A reconfiguration technique called Double Scheme reduces by up to 95% the peak packet latency caused by the on/off link operation. José Miguel Montañana, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
NAS | 2 |
| 2010 | A Deadlock-Free Non-minimal Fully Adaptive Routing Using Virtual Cut-Through SwitchingabstractSystem area networks (SANs), which usually employ virtual cut-through switching, have been used to connect hosts in modern PC clusters and massively parallel computers. In this paper, we propose a non-minimal fully adaptive deadlock-free routing mechanism for virtual-cut-through networks called “Semi-deflection”. Semi-deflection routing guarantees deadlock-free packet transfer without use of virtual channels by allowing non-blocking transfer between specific pairs of routers. As the result of throughput evaluation, Semi-deflection routing improved throughput by up to 26 percent compared with that of north-last turn model, which is a typical adaptive routing, and also reduced latency. Yuri Nishikawa, Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano |
NAS | 2 |
| 2010 | Ultra Fine-Grained Run-Time Power Gating of On-chip Routers for CMPsabstractThis paper proposes an ultra fine-grained run-time power gating of on-chip router, in which power supply to each router component (e.g., VC queue, crossbar MUX, and output latch) can be individually controlled in response to the applied workload. As only the router components which are just transferring a packet are activated, the leakage power of the on-chip network can be reduced to the near-optimal level. However, a certain amount of wakeup latency is required to activate the sleeping components, and the application performance will be degraded. In this paper, we estimate the wakeup latency for each component based on circuit simulations using a 65 nm process. Then we propose four early wakeup methods to overcome the wakeup latency. The proposed router with the early wakeup methods is evaluated in terms of the application performance, area, and leakage power. As a result, it reduces the leakage power by 78.9%, at the expense of the 4.3% area and 4.0% performance when we assume a 1 GHz operation. Hiroki Matsutani, Michihiro Koibuchi, Daisuke Ikebuchi, Kimiyoshi Usami, Hiroshi Nakamura, Hideharu Amano |
NOCS | 2 |
| 2009 | Prediction router: Yet another low latency on-chip router architectureabstractNetwork-on-Chips (NoCs) are quite latency sensitive, since their communication latency strongly affects the application performance on recent many-core architectures. To reduce the communication latency, we propose a low-latency router architecture that predicts an output channel being used by the next packet transfer and speculatively completes the switch arbitration. In the prediction routers, incoming packets are transferred without waiting the routing computation and switch arbitration if the prediction hits. Thus, the primary concern for reducing the communication latency is the hit rates of prediction algorithms, which vary from the network environments, such as the network topology, routing algorithm, and traffic pattern. Although typical low-latency routers that speculatively skip one or more pipeline stages use a bypass datapath for specific packet transfers (e.g., packets moving on the same dimension), our prediction router predictively forwards packets based on a prediction algorithm selected from several candidates in response to the network environments. In this paper, we analyze the prediction hit rates of six prediction algorithms on meshes, tori, and fat trees. Then we provide three case studies, each of which assumes different many-core architecture. We have implemented a prediction router for each case study by using a 65 nm CMOS process, and evaluated them in terms of the prediction hit rate, zero load latency, hardware amount, and energy consumption. The results show that although the area and energy are increased by 6.4-15.9% and 8.0-9.5% respectively, up to 89.8% of the prediction hit rate is achieved in real applications, which provide favorable trade-offs between the modest hardware/energy overheads and the latency saving. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Tsutomu Yoshinaga |
HPCA | 2 |
| 2009 | Implementation and Evaluation of Layer-1 Bandwidth-on-Demand Capabilities in SINET3abstractThis paper describes the implementation and evaluation of layer-1 bandwidth-on-demand (BoD) capabilities in the Japanese academic backbone network, called SINET3. The network has a nationwide GMPLS-based layer-1 platform and provides reservation-based and signaling-based BoD services. The overall architecture for providing BoD services including its capabilities, user interface, path calculation, and interface to drive the layer-1 platform are described. Actual examples of BoD services and evaluations of the path setup/release time in the network are also presented. Shigeo Urushidani, Kensuke Fukuda, Yusheng Ji, Michihiro Koibuchi, Shunji Abe, Motonori Nakamura, Shigeki Yamada, Kaori Shimizu, Rie Hayashi, Ichiro Inoue, Kohei Shiomoto |
ICC | 4 |
| 2009 | An on/off link activation method for low-power ethernet in PC clustersabstractThe power consumption of interconnects is increased as the link bandwidth is improved in PC clusters. In this paper, we propose an on/off link activation method that uses the static analysis of the traffic in order to reduce the power consumption of Ethernet switches while maintaining the performance of PC clusters. When a link whose utilization is low is deactivated, the proposed method renews the VLAN-based paths that avoid it without creating broadcast storms. Since each host does not need to process VLAN tags, the proposed method has advantages in both simple host configuration and high portability. Evaluation results using NAS Parallel Benchmarks show that the proposed method reduces the power consumption of switches by up to 37% without performance degradation. Michihiro Koibuchi, Tomohiro Otsuka, Hiroki Matsutani, Hideharu Amano |
IPDPS | 1 |
| 2009 | Efficient Scheduling Algorithms on Bandwidth Reservation Service of Internet Using MetaheuristicsabstractNetwork services that dynamically allocate bandwidth resources, such as QoS and layer-1 bandwidth-on-demand(BoD), are increasingly required to advanced Internet backbones, such as science information networks (SINET) in Japan. In this paper, we propose scheduling algorithms for BoD service which allocate parts of full bandwidth dedicated to specific users according to their requests in advanced Internet backbones. The scheduling algorithms maximize the number of accepted requests, fairness, or the total bandwidth in BoD. Simulation results show that the proposed algorithms achieve high utilization of network resources and user's fairness, compared with a simple random-based algorithm. Tomoyuki Hiroyasu, Kozo Kawasaki, Michihiro Koibuchi, Shigeo Urushidani, Mitsunori Miki, Masato Yoshimi |
ISDA | 3 |
| 2009 | Performance Analysis of ClearSpeed's CSX600 InterconnectsabstractClearSpeed's CSX600 that consists of 96 Processing Elements (PEs) employs a one-dimensional array topology for a simple SIMD processing. To clearly show the performance factors and practical issues of NoCs in an existing modern many-core SIMD system, this paper measures and analyzes NoCs of CSX600 called Swazzle and ClearConnect. Evaluation and analysis results show that the sending and receiving overheads are the major limitation factors to the effective network bandwidth. We found that (1) the number of used PEs, (2) the size of transferred data, and (3) data alignment of a shared memory are three main points to make the best use of bandwidth. In addition, we estimated the best- and worst-case latencies of data transfers in parallel applications. Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Akihiro Shitara, Kenichi Miura, Hideharu Amano |
ISPA | 2 |
| 2009 | Design of versatile academic infrastructure for multilayer network servicesabstractThis paper describes the network design and configurations of the new Japanese academic infrastructure, called SINET3, which provides a rich variety of network services to more than 700 universities and research institutions. Since the start of full-scale operations in June 2007, the network has expanded its services to include multi-layer transfer services (IP, Ethernet, and layer-1), enriched virtual private network services (L3VPN, L2VPN, VPLS, and L1VPN), enhanced QoS services (packet-based and circuit-based), and brand-new layer-1 bandwidth-on-demand (BoD) services. This paper explains how the network provides these various network services on a single network platform by effectively configuring leading-edge networking components, such as high-performance IP routers, layer- 1 switches, and a BoD server. Evaluations of the network design and configurations confirmed that the networking functions were effectively coordinated. The procedures and techniques related to the configuration validation that covered all phases of the network design and construction are also presented. Shigeo Urushidani, Shunji Abe, Yusheng Ji, Kensuke Fukuda, Michihiro Koibuchi, Motonori Nakamura, Shigeki Yamada, Kaori Shimizu, Rie Hayashi, Ichiro Inoue, Kohei Shiomoto |
IEEE J. Sel. Areas Commun. | 5 |
| 2009 | Fat H-Tree: A Cost-Efficient Tree-Based On-Chip NetworkabstractThe topological explorations of on-chip networks are important for efficiently using their enormous wire resources for low-latency and high-throughput communications using a modest silicon budget. In this paper, we propose a novel tree-based interconnection network called Fat H-Tree that meets these requirements. A Fat H-Tree provides a torus structure by combining two folded H-Tree networks and is an attractive alternative to tree-based networks such as the Fat Trees in a microarchitecture domain. We introduce its chip layout schemes based on a folding technique for 2D and 3D ICs. Three deadlock-free routing schemes are proposed for Fat H-Tree. We evaluate the performance of Fat H-Tree and other tree-based networks using real application traces. In addition, the network logic area, wire resource, and energy consumption of Fat H-Tree are compared with other topologies, based on a typical implementation of on-chip routers synthesized with a 90-nm standard cell library. The results show that (1) a Fat H-Tree outperforms a Fat Tree with two upward and four downward connections in terms of the throughput and average hop count, (2) a Fat H-Tree requires 19.8 percent-27.8 percent smaller network logic area than the Fat Tree, (3) a Fat H-Tree consumes slightly less energy than the Fat Tree does, and (4) a Fat H-Tree uses slightly more wire resources than the Fat Tree, but the current process technology can provide sufficient wire resources for implementing Fat-H-Tree-based on-chip networks. Hiroki Matsutani, Michihiro Koibuchi, Yutaka Yamada, D. Frank Hsu, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | Run-time power gating of on-chip routers using look-ahead routingabstractSince on-chip routers in Network-on-Chips play a key role in on-chip communication between cores, they should be always preparing for packet injections even if a part of cores are in standby mode, resulting in a larger standby power of routers compared with cores. The run-time power gating of individual channels in a router is one of attractive solutions to reduce the standby power of chip without affecting the on-chip communication. However, a state transition between sleep and active mode incurs the performance penalty, and turning a power switch on or off dissipates the overhead energy, which means a short-term sleep adversely increases the power consumption. In this paper, we propose a sleep control method based on look-ahead routing that detects the arrival of packets two hops ahead, so as to hide the wake-up delay and reduce the short-term sleeps of channels. Simulation results using real application traces show that the proposed method conceals the wake-up delay of less than five cycles, and more leakage power can be saved compared with the original naive method. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano, Daihan Wang |
ASP-DAC | 2 |
| 2008 | Impact of topology and link aggregation on a PC cluster with EthernetabstractIn addition to its use in local area networks, Ethernet has been used for connecting hosts in the area of high-performance computing. Here, we investigated the impact of topology and link aggregation on a large-scale PC cluster with Ethernet. Ethernet topology that allows loops and its routing can be implemented by the VLAN routing method without creating broadcast storms. To simplify the system configuration without modifying system software, the VLAN tag is added to a frame at switches in our implementation of topologies. Each host creates VLAN interfaces that have different local network addresses on a physical interface, so that a switch learns the MAC addresses of hosts in a PC cluster by broadcast. Evaluation results showed that the performance characteristics of an eight-switch network are comparable to those of an ideal 1-switch (full crossbar) network in the execution of High-Performance LINPACK Benchmark (HPL) on a 225-host PC cluster. On the other hand, evaluation results using NAS Parallel Benchmarks indicated that topologies achieved by the proposed methodology showed performance improvements of up to about 650% as compared to the simple tree topology. These results indicate that topology and link aggregation have marked impacts and commodity switches can be used instead of expensive and high functional switches. Takafumi Watanabe, Masahiro Nakao, Tomoyuki Hiroyasu, Tomohiro Otsuka, Michihiro Koibuchi |
CLUSTER | 5 |
| 2008 | A link removal methodology for Networks-on-Chip on reconfigurable systemsabstractWhile the regular 2-D mesh topology has been utilized for most of Network-on-Chips (NoCs) on FPGAs, spatially biased traffic in some applications make some customization method feasible. A link removal strategy that customizes the router in NoC is proposed for reconfigurable systems in order to minimize required hardware amount. Based on the pre-analyzed traffic information, links on which the communication amount is small are removed to reduce the hardware cost with enough performance being kept. Two policies are proposed to avoid deadlocks and better performance can be achieved compared with up*/down* routing on the irregular topology with links removed. In the image recognition application susan, the proposed method can save 30% of the hardware amount without performance degradation. Daihan Wang, Hiroki Matsutani, Hideharu Amano, Michihiro Koibuchi |
FPL | 4 |
| 2008 | A Lightweight Fault-Tolerant Mechanism for Network-on-Chip
Michihiro Koibuchi, Hiroki Matsutani, Hideharu Amano, Timothy M. Pinkston |
NOCS | 1 |
| 2008 | Adding Slow-Silent Virtual Channels for Low-Power On-Chip Networks
Hiroki Matsutani, Michihiro Koibuchi, Daihan Wang, Hideharu Amano |
NOCS | 2 |
| 2007 | A Temporal Correlation Based Port Combination Methodology for Networks-on-chip on Reconfigurable SystemsabstractA temporal correlation based port combination algorithm that customizes the router design in Network-on-Chip (NoC) is proposed for reconfigurable systems in order to minimize required hardware amount. Given the traffic characteristics of the target application and the expected hardware amount reduction rate, the algorithm automatically makes the port combination plan for the networks. Since the port combination technique has the advantage of almost keeping the topology, it does not affect the design of the other layers, such as task mapping and scheduling. The algorithm shows much better efficiency than the algorithm without temporal correlation. For the multimedia stream processing application, the algorithm can save 55% of the hardware amount without performance degradation, while the non-temporal correlation algorithm suffers from 30% performance loss. Daihan Wang, Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
FPL | 3 |
| 2007 | Layer-1 Bandwidth on Demand Services in SINET3abstractThis paper describes brand-new layer-1 bandwidth on demand (BoD) services implemented in the new Japanese academic backbone network, called SINET3. SINET3 is an advanced converged network that provides multi-layer transfer, enriched VPN, enhanced QoS, and layer-1 BoD services. The layer-1 BoD services are dynamic layer-1 resource allocation services directly triggered by users and artfully achieved on the multi-service platform by using a layer-1 BoD server. This paper first explains how the network accommodates a wide variety of network services by effectively combining leading-edge technologies. The paper next describes the overall mechanism for the dynamic layer-1 path setup on the multi-service platform and details the functions of the BoD server in many aspects. The designs focus on the tangible achievement of these services over a nationwide network composed of 75 layer-1 switches and 12 IP/MPLS routers. Shigeo Urushidani, Jun Matsukata, Kensuke Fukuda, Shunji Abe, Yusheng Ji, Michihiro Koibuchi, Shigeki Yamada, Kaori Shimizu, Tomonori Takeda, Ichiro Inoue, Kohei Shiomoto |
GLOBECOM | 6 |
| 2007 | Investigating QoS Performance on a Testbed NetworkabstractQuality of Service (QoS) in Layer 3 is essential for satisfying various types of Internet-application requirements by making the best use of limited bandwidth. A large number of such applications still use TCP/IP in order to connect various types of computer nodes. This paper investigates QoS performance in a network equipment testbed. We examine the major Class of Service (CoS) functions provided by the Juniper T320 router, and measure their performance. In addition to fundamental analysis of the QoS behavior, we show the impact of QoS operations on a parallel system distributed in multi-domain networks as a practical case study of grid environments. Jumpot Phuritatkul, Kien Nguyen 0002, Michihiro Koibuchi, Yusheng Ji |
ICCCN | 3 |
| 2007 | Tightly-Coupled Multi-Layer Topologies for 3-D NoCsabstractThree-dimensional network-on-chip (3-D NoC) is an emerging research topic exploring the network architecture of 3-D ICs that stack several smaller wafers for reducing wire length and wire delay. Although the network topology of 3-D NoC has been explored for a couple of years, there is still only a narrow range of choices. In this paper, we propose a class of 3-D topologies called Xbar-connected network-on-tiers (XNoTs), which consist of multiple network layers tightly connected via crossbar switches. To make the best use of the short delay and high density of inter-wafer links, XNoTs topologies have crossbar switches that connect different layers and their cores. The planar topology on every layer can be independently customized so as to meet the cost-performance requirements, as far as network connectivity is at least guaranteed with the bottom layer. We also propose their routing algorithm, which guarantees deadlock-freedom by restricting the inter-layer packet transfer from a lower-numbered layer to a higher-numbered layer. Path sets at the bottom layer close to the heat sink of the chip can be selectively employed in order to mitigate the heat-dissipation problem of 3-D ICs. Several forms of XNoTs topologies including meshes, tori, and/or trees are created, and they are evaluated in terms of performance, cost, and energy consumption. As a result, we show that even with the flexibilities mentioned above, XNoTs achieve at least as high throughput as existing 3-D topologies for equivalent chip sizes. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ICPP | 2 |
| 2007 | Performance Improvement Methodology for ClearSpeed's CSX600abstractThis paper focuses on a performance of network-on-a- chip (NoC) and I/O of ClearSpeed's CSX600 coprocessor with 96 multithread processing elements. Two versions of the Himeno benchmark were implemented on the CSX600 to evaluate its performance when it encounters frequent memory transfers between shared and local memories, or between local memories. In order to efficiently use the NoC bandwidth, the dataflow was customized to the one- dimensional array structure of CSX600's NoC. The results of evaluation and profiling indicate that the performance was lower than 1/50 of the sustained performance. We show three key points to improve the performance on such a case: 1) exploiting bandwidth between mono and poly memory, 2) further program tuning, and 3) architectural reform. Yuri Nishikawa, Michihiro Koibuchi, Masato Yoshimi, Kenichi Miura, Hideharu Amano |
ICPP | 2 |
| 2007 | Performance, Cost, and Energy Evaluation of Fat H-Tree: A Cost-Efficient Tree-Based On-Chip NetworkabstractFat H-Tree is a novel tree-based interconnection network providing a torus structure, which is formed by combining two folded H-Tree networks, and is an attractive alternative to tree-based networks such as Fat Trees in a micro architecture domain. In this paper, we introduce Fat H-Tree and its deadlock-free routing algorithms. The performance of Fat H-Tree is evaluated using real application traces, and the result is compared with those of other tree-based networks. The network logic area and wire resources for Fat H-Tree are computed based on a typical implementation of on-chip routers using a 0.18mum standard cell library. In addition, the energy consumption is estimated based on the gate-level power analysis. The results show that 1) Fat H-Tree outperforms Fat Tree with two upward and four downward connections in terms of throughput and average hop count; 2) Fat H-Tree requires 19.3%-26.4% smaller network logic area compared with the Fat Tree; 3) Fat H-Tree consumes 8.3%-8.6% less energy compared with the Fat Tree due to its short average hop count; 4) Fat H-Tree uses slightly more wire resources compared with the Fat Tree, but the current process technology can provide sufficient wire resources for implementing Fat H-Tree based on-chip networks. Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
IPDPS | 2 |
| 2007 | An Effective Design of Deadlock-Free Routing Algorithms Based on 2D Turn Model for Irregular NetworksabstractSystem area networks (SANs), which usually accept arbitrary topologies, have been used to connect hosts in PC clusters. Although deadlock-free routing is often employed for low-latency communications using wormhole or virtual cut-through switching, the interconnection adaptivity introduces difficulties in establishing deadlock-free paths. An up*/down* routing algorithm, which has been widely used to avoid deadlocks in irregular networks, tends to make unbalanced paths as it employs a one-dimensional directed graph. The current study introduces a two-dimensional directed graph on which adaptive routings called left-up first turn (L-turn) routings and right-down last turn (R-turn) routings are proposed to make the paths as uniformly distributed as possible. This scheme guarantees deadlock-freedom because it uses the turn model approach, and the extra degree of freedom in the two-dimensional graph helps to ensure that the prohibited turns are well-distributed. Simulation results show that better throughput and latency results from uniformly distributing the prohibited turns by which the traffic would be more distributed toward the leaf nodes. The L-turn routings, which meet this condition, improve throughput by up to 100 percent compared with two up*/down*-based routings, and also reduce latency Akiya Jouraku, Michihiro Koibuchi, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | Switch-tagged VLAN Routing Methodology for PC Clusters with EthernetabstractEthernet has been used for connecting hosts in the area of high performance-per-cost PC clusters. Although L2 Ethernet topology is limited to a tree structure, various routing algorithms on topologies suitable for parallel processing can be employed by applying IEEE 802.1Q VLAN technology. However, communication library used in PC clusters does not always support VLANs, so the design of VLAN-based routing method cannot be applied for such PC clusters. In this paper, we propose a switch-tagged VLAN methodology to flexibly set the route of frames on such PC clusters. Since each host does not need to process VLAN tags, the proposed method has advantages in both simple host configuration and high portability. Evaluation results using NAS Parallel Benchmarks showed that performance of topologies supported by the proposed method was comparable with that of an ideal 1-switch (full crossbar) network in the case of a 16-host PC cluster Tomohiro Otsuka, Michihiro Koibuchi, Tomohiro Kudoh, Hideharu Amano |
ICPP | 2 |
| 2006 | Enforcing Dimension-Order Routing in On-Chip Torus Networks Without Virtual Channels
Hiroki Matsutani, Michihiro Koibuchi, Hideharu Amano |
ISPA | 2 |
| 2006 | A Simple Data Transfer Technique Using Local Address for Networks-on-ChipsabstractNetworks-on-chips (NoCs) have been studied to connect a number of modules in a chip by introducing a network structure which is similar to that in parallel computers. Since embedded streaming applications usually generate predictable small-sized data traffic, the network structure can be customized to the target traffic. Accordingly, we develop a data transfer technique for simplifying routers for predictable small-sized communication in simple tile-based architectures. A data structure is split into single-flit packets, and a label is attached to each of them in order to route them independently. A label is transferred on dedicated wires beside data lines in a channel by taking advantage of relaxed pin count limitations of a channel. To reduce the wiring area for the label, the label is locally assigned according to a preanalysis of required communication pairs of nodes. Analysis results show that only a 3-bit local label is sufficient to route all data of evaluated streaming applications in the case of a 16-node 2D torus. The required amount of hardware for a router is reduced by 37 percent compared with that for a wormhole packet router with the same number of routing table entries. Michihiro Koibuchi, Kenichiro Anjo, Yutaka Yamada, Akiya Jouraku, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | VLAN-Based Minimal Paths in PC Cluster with Ethernet on Mesh and TorusabstractIn a PC cluster with Ethernet, well-distributed multiple paths among hosts can be obtained by applying VLAN technology. In this paper, we propose VLAN topology sets and path assignment methods in mesh and torus. The proposed VLAN-based methods on mesh require N/sup M-1/ and /spl lfloor/N/sup M-1//2/spl rfloor/+1 VLANs to provide balanced minimal paths and partially balanced ones respectively, where N is the number of switches per dimension and M is the number of dimensions. Similarly, those on torus require 2N/sup M-1/ and N/sup M-1/+2 VLANs respectively. Simulation results show that the proposed methods improve up to 902% and 706% of throughput respectively. Tomohiro Otsuka, Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
ICPP | 2 |
| 2005 | Enforcing in-order packet delivery in system area networks with adaptive routing
Michihiro Koibuchi, José Flich, Antonio Robles, Pedro López 0001, José Duato |
J. Parallel Distributed Comput. | 1 |
| 2005 | Path selection algorithm: the strategy for designing deterministic routing from alternative paths
Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
Parallel Comput. | 1 |
| 2005 | Performance Evaluation of Deterministic Routings, Multicasts, and Topologies on RHiNET-2 ClusterabstractSystem area networks (SANs), which usually accept arbitrary topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. Although deadlock-free routings, multicasts, and topologies for SANs have been widely developed, their evaluation on real PC clusters was rarely done. Thus, the evaluation of routings, multicasts, and topologies in real systems is important to analyze their impact on the total systems and validate their simulation results. In this paper, we implement and evaluate deadlock-free routings and unicast-based multicasts under various topologies and channel buffer sizes on a PC cluster called RHiNET-2 with 64 hosts. Execution results show that descending layers (DL) routing and structured channel pools improve up to 57 percent of bandwidth and 34 percent of barrier synchronization time compared with up*/down* routing. They also show that, by visiting hosts in numerical order, execution time of unicast-based barrier synchronization is improved up to 28 percent compared with that in random order. However, channel buffer sizes don't affect the bandwidth in the RHiNET-2 cluster. In addition to fundamental evaluation, we appraise them using NAS Parallel Benchmarks, and the DL routing achieves 3.2 percent improvement on their execution time compared with up*/down* routing. Michihiro Koibuchi, Konosuke Watanabe, Tomohiro Otsuka, Hideharu Amano |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2004 | Folded Fat H-Tree: An Interconnection Topology for Dynamically Reconfigurable Processor Array
Yutaka Yamada, Hideharu Amano, Michihiro Koibuchi, Akiya Jouraku, Kenichiro Anjo, Katsunobu Nishimura |
EUC | 3 |
| 2004 | BLACK-BUS: A New Data-Transfer Technique Using Local Address on Networks-on-ChipsabstractSummary form only given. Network-on-a-chip (NoC) has received attention as a high-performance interconnect, because traditional buses, which can't transfer more than one data-stream simultaneously, are more likely to become a bottleneck. Since some concepts of NoC have been proposed by simply borrowing the networking structure of parallel computers or system area networks (SANs), it tends to require complicated network interface logic in all the nodes. We propose a novel data-transfer method called Black-Bus as a NoC. In Black-Bus, a local identifier (ID) is attached to each raw data as routing information. Unlike the traditional packet transfer, the local ID is transferred on dedicated wires attached to data lines to remove complicated packet generation procedure in a node. Only a small-sized local ID is required to specify routing tags to the destination, and intermediate routers change it to solve local ID conflicts between paths on a physical channel. The required local ID and routing table sizes for the Black-Bus router are evaluated with access trace data of NAS parallel benchmarks for on-chip multiprocessors, and JPEG codec as stream processing. Evaluation results show that most of the applications require only at most 3 bits for the local ID in a 16-node system. And the Black-Bus data-transfer reduces up to 75% of routing tags compared with global addressing scheme used in the traditional packet networks. Kenichiro Anjo, Yutaka Yamada, Michihiro Koibuchi, Akiya Jouraku, Hideharu Amano |
IPDPS | 3 |
| 2003 | Performance Evaluation of Routing Algorithms in RHiNET-2 ClusterabstractSystem area networks (SANs), which usually accept irregular topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. A lot of deadlock-free routings for SANs have been proposed, and their evaluation on simulations have been widely done. However, these simulation results may differ from that of real PC clusters, since hosts, network interfaces and switches used in the simulation are simplified for achieving enough simulation speed. In this paper, we implement deadlock-free routings on a high-performance PC cluster called RHiNET-2, and evaluate their performance. Execution results show that the DL routing and the structured channel pools achieve almost the same total bandwidth and execution time of the barrier synchronization. Compared with the simple Up*/Down* routing, they improve 51% of total bandwidth and 29% improvement on execution time of the barrier synchronization. Michihiro Koibuchi, Konosuke Watanabe, Kenichi Kono, Akiya Jouraku, Hideharu Amano |
CLUSTER | 1 |
| 2003 | Descending Layers Routing: A Deadlock-Free Deterministic Routing using Virtual Channels in System Area Networks with Irregular TopologiesabstractSystem area networks (SANs), which usually accept irregular topologies, have been used to connect nodes in PC/WS clusters or high-performance storage systems. Since wormhole or virtual cut-through transfer is used for low latency communication, deadlock-free routings are essential in SANs. We propose a novel deadlock-free deterministic routing called descending layers (DL) routing for SANs. In order to reduce both nonminimal paths and traffic congestion, the network is divided into layers of subnetworks with the same topology using virtual channels, and a large number of paths across multiple subnetworks are established. The DL routing is implemented on a real PC cluster called RHiNET-2, and execution results show that its throughput is improved up to 33% compared with that of up*/down* routing. Its execution time of a barrier synchronization is also improved 29% compared with that of up*/down* routing. Simulation results of various sizes and topologies also show that the DL routing achieves up to 266% improvement on throughput compared with up*/down* routing. Michihiro Koibuchi, Akiya Jouraku, Konosuke Watanabe, Hideharu Amano |
ICPP | 1 |
| 2001 | L-Turn Routing: An Adaptive Routing in Irregular NetworksabstractNetwork-based parallel processing using commodity personal computers has been widely developed. Since such systems require high degree of flexibility and scalability of wiring, a high-speed network with an irregular topology is often needed. In traditional routing algorithms for irregular networks, available paths are considerably restricted in order to avoid deadlocks. In this paper we propose a novel routing algorithm called left-up-first turn routing (L-turn routing), which makes a better traffic balancing in irregular networks by building a specific spanning tree. Result of simulations shows that L-turn routing achieves better performance than traditional ones with each topology. Michihiro Koibuchi, Akira Funahashi, Akiya Jouraku, Hideharu Amano |
ICPP | 1 |