Toshihiro Hanawa

dblp:96/5345 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-2970-6037ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 3 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 Automated Configuration of Power-Management Knobs for Optimal HPC Job Executions
Francesco Antici, Andrea Proia, Ryoma Ohara, Toshihiro Hanawa, Zeynep Kiziltan, Andrea Bartolini, Jens Domke
CCGrid4
2024 WaitIO-Hybrid: Communication for Coupling MPI Programs Among Heterogeneous Systems
Shinji Sumimoto, Takashi Arakawa, Yoshio Sakaguchi, Hiroya Matsuba, Satoshi Ohshima, Hisashi Yashiro, Toshihiro Hanawa, Kengo Nakajima
PDCAT7
2022 Optimizations of H-matrix-vector Multiplication for Modern Multi-core Processors
abstract
Hierarchical matrices (H-matrices) can robustly approximate the dense matrices that appear in the boundary element method (BEM). To accelerate the solving of linear systems in the BEM, we must speed up the matrix-vector multiplication in the iterative linear solver. However, speed-up approaches are usually developed for dense or sparse matrices, and are rarely reported for hierarchical matrix-vector multiplication (HiMV). The HiMV algorithm generates a large number of matrix-vector multiplications, which have not been sufficiently discussed. Therefore, the efficiency of HiMV has not reached its potential. This paper discusses optimization methodologies of HiMV for modern multi-core CPUs: an H-matrix storage method for efficient memory access, a method that avoids write contentions during reduction operations on the solution vector, an inter-thread load-balancing method, and blocking and sub-matrix sorting methods for cache efficiency. We demonstrate that these optimizations significantly improve the performance of modern CPU-based supercomputers. Relative to the target performance of dense matrix-vector multiplication (DGEMV), the HiMV flops reached 84.8%, 100.7%, and 98.7% during single-socket execution on the A64FX, AMD EPYC, and Intel Xeon Cascade Lake processors, respectively. Optimization of memory performance and cache efficiency is especially important for the A64FX with high-speed high-bandwidth memory.
Tetsuya Hoshino, Akihiro Ida, Toshihiro Hanawa
CLUSTER3
2022 A System-Wide Communication to Couple Multiple MPI Programs for Heterogeneous Computing
Shinji Sumimoto, Takashi Arakawa, Yoshio Sakaguchi, Hiroya Matsuba, Hisashi Yashiro, Toshihiro Hanawa, Kengo Nakajima
PDCAT6
2021 Automatic Graph Partitioning for Very Large-scale Deep Learning
abstract
This work proposes RaNNC (Rapid Neural Network Connector) as middleware for automatic hybrid parallelism. In recent deep learning research, as exemplified by T5 and GPT-3, the size of neural network models continues to grow. Since such models do not fit into the memory of accelerator devices, they need to be partitioned by model parallelism techniques. Moreover, to accelerate training for huge training data, we need a combination of model and data parallelisms, i.e., hybrid parallelism. Given a model description for PyTorch without any specification for model parallelism, RaNNC automatically partitions the model into a set of subcomponents so that (1) each subcomponent fits a device memory and (2) a high training throughput for pipeline parallelism is achieved by balancing the computation times of the subcomponents. Since the search space for partitioning models can be extremely large, RaNNC partitions a model through the following three phases. First, it identifies atomic subcomponents using simple heuristic rules. Next it groups them into coarser-grained blocks while balancing their computation times. Finally, it uses a novel dynamic programming-based algorithm to efficiently search for combinations of blocks to determine the final partitions. In our experiments, we compared RaNNC with two popular frameworks, Megatron-LM (hybrid parallelism) and GPipe (originally proposed for model parallelism, but a version allowing hybrid parallelism also exists), for training models with increasingly greater numbers of parameters. In the pre-training of enlarged BERT models, RaNNC successfully trained models five times larger than those Megatron-LM could, and RaNNC's training throughputs were comparable to Megatron-LM's when pre-training the same models. RaNNC also achieved better training throughputs than GPipe on both the enlarged BERT model pre-training (GPipe with hybrid parallelism) and the enlarged ResNet models (GPipe with model parallelism) in all of the settings we tried. These results are remarkable, since RaNNC automatically partitions models without any modification to their descriptions; Megatron-LM and GPipe require users to manually rewrite the models' descriptions.
Masahiro Tanaka, Kenjiro Taura, Toshihiro Hanawa, Kentaro Torisawa
IPDPS3
2020 Analysis of Cooling Water Temperature Impact on Computing Performance and Energy Consumption
abstract
The hot water cooling technique has been widely accepted as one of the standard techniques to improve the energy efficiency for the HPC and Data Centers. However, the higher operating temperature may impact the CPU power consumption due to the leakage current. Moreover, it may degrade the computational performance due to the DVFS mechanism activated to maintain the power and temperature within the TDP limit. In that sense, to fairly evaluate the efficiency of the hot water cooling technique, it becomes important to take into consideration not only the energy reduction on the HPC facility side (cooling system) but also the impact on the power consumption and the performance degradation on the HPC system side. In this paper, we utilized the Oakforest-PACS system and its facility, jointly administrated by University of Tsukuba and The University of Tokyo, in order to execute a quantitative and systematic analysis on the impact of the cooling water temperature onto the HPC system and its facility. For this purpose, we utilized lower (9 °C) and higher (18 °C) cooling water temperature other than the regular operational temperature (12 °C). Contrary to the gain in the energy consumption, on the HPC facility side, when using higher cooling water temperature, we observed an increase in the number of nodes suffering from performance degradation on the HPC system side. As a result, it can directly increase the probability of including low-performance nodes on multiple node jobs, and thus affecting their overall performance, especially during barrier synchronizations.
Jorji Nonaka, Toshihiro Hanawa, Fumiyoshi Shoji
CLUSTER2
2020 Development of training environment for deep learning with medical images on supercomputer system based on asynchronous parallel Bayesian optimization
Yukihiro Nomura, Issei Sato, Toshihiro Hanawa, Shohei Hanaoka, Takahiro Nakao, Tomomi Takenaga, Tetsuya Hoshino, Yuji Sekiya, Soichiro Miki, Takeharu Yoshikawa, Naoto Hayashi, Osamu Abe
J. Supercomput.3
2018 Load-Balancing-Aware Parallel Algorithms of H-Matrices with Adaptive Cross Approximation for GPUs
abstract
Hierarchical matrices (H-matrices) are an approximation technique for dense matrices, such as the coefficient matrix of the boundary element method (BEM). An H-matrix is expressed by a set of low-rank approximated and small dense sub-matrices, each of which has various ranks. The use of H-matrices reduces the required memory footprint of dense matrices from O(N^2) to O(NlogN) and is suitable for many-core processors that have relatively small memory capacities compared to traditional CPUs. However, existing parallel adaptive cross approximation (ACA) algorithms, which are low-rank approximation algorithms used to construct H-matrices, are not designed to exploit many-core processors in terms of load balancing. In existing parallel algorithms, the ACA process is independently applied to each sub-matrix. The computational load of the ACA process for each sub-matrix depends on the sub-matrix's rank; however, the rank is defined after the ACA process is applied. This makes it difficult to balance the load. We propose load-balancing-aware parallel ACA algorithms for H-matrices that focus on many-core processors. We implemented the proposed algorithms into HACApK, which is an open-source H-matrix library originally developed for CPU-based clusters. The proposed algorithms were evaluated using BEM problems on an NVIDIA Tesla P100 GPU (P100) and an Intel Xeon Broadwell processor. The evaluation results demonstrate the improved performance of the proposed algorithms in all GPU cases. For example, in a case where it is difficult for existing parallel algorithms to balance the load, the proposed algorithms achieved a 12.9 times performance improvement for the P100.
Tetsuya Hoshino, Akihiro Ida, Toshihiro Hanawa, Kengo Nakajima
CLUSTER3
2015 Evaluation of FFT for GPU Cluster Using Tightly Coupled Accelerators Architecture
abstract
Inter-node communications between accelerators in heterogeneous clusters require extra latency because of the time required to transfer data copies between the host and accelerator. Such communication latencies inhibit the optimal performance of affected applications. To address this problem, we proposed the Tightly Coupled Accelerators (TCA) architecture and designed an interconnection router chip named PEACH2. Accelerators in the TCA architecture communicate directly via the PCIe protocol, which is the current fundamental interface for all the accelerators and the host CPU, to eliminate protocol and data copy overheads. In this paper, we apply the TCA architecture to the Fast Fourier Transform (FFT) program, which is commonly used in scientific computations. First, we implemented all-to-all communication to TCA. The all-to-all communication was then applied to FFTE, which is one of the implementations of FFT. Based on the evaluation results using the HA-PACS/TCA system, we achieved the speedup of 2.7 with TCA in comparison with that with MPI using 16 nodes on the medium size.
Toshihiro Hanawa, Hisafumi Fujii, Norihisa Fujita, Tetsuya Odajima, Kazuya Matsumoto, Taisuke Boku
CLUSTER1
2015 Improving Strong-Scaling on GPU Cluster Based on Tightly Coupled Accelerators Architecture
abstract
The Tightly Coupled Accelerators (TCA) architecture that we proposed in previous work enables direct ommunication between accelerators over nodes. In this paper, we present a proof-of-concept GPU cluster called the HA-PACS/TCA using the PEACH2 chip that we designed as an interconnection router chip based on the TCA architecture. Our system demonstrated 2.0 ?sec of latency on inter-node GPU-to-GPU communication with a PCIe Gen2 x8 by RDMA, reducing minimum latency to just 44% of the InfiniBand-QDR and MPI using GPUDirect for RDMA. Through results of Himeno benchmark tests, we demonstrated that our TCA architecture improved performance scalability with the small-sized problem by up to 61%.
Toshihiro Hanawa, Hisafumi Fujii, Norihisa Fujita, Tetsuya Odajima, Kazuya Matsumoto, Yuetsu Kodama, Taisuke Boku
CLUSTER1
2015 Hybrid Communication with TCA and InfiniBand on a Parallel Programming Language XcalableACC for GPU Clusters
abstract
For the execution of parallel HPC applications on GPU-ready clusters, high communication latency between GPUs over nodes will be a serious problem on strong scalability. To reduce the communication latency between GPUs, we proposed the Tightly Coupled Accelerator (TCA) architecture and developed the PEACH2 board as a proof-of-concept interconnection system for TCA. Although PEACH2 provides very low communication latency, there are some hardware limitations due to its implementation depending on PCIe technology, such as the practical number of nodes in a system which is 16 currently named sub-cluster. More number of nodes should be connected by conventional interconnections such as InfiniBand, and the entire network system is configured as a hybrid one with global conventional network and local high-speed network by PEACH2. For ease of user programmability, it is desirable to operate such a complicated communication system at the library or language level (which hides the system). In this paper, we develop a hybrid interconnection network system combining PEACH2 and InfiniBand, and implement it based on a high-level PGAS language for accelerated clusters named XcalableACC (XACC). A preliminary performance evaluation confirms that the hybrid network improves the performance based on the Himeno benchmark for stencil computation by up to 40%, relative to MVAPICH2 with GDR on InfiniBand. Additionally, Allgather collective communication with a hybrid network improves the performance by up to 50% for networks of 8 to 16 nodes. The combination of local communication, supported by the low latency of PEACH2 and global communication supported by the high bandwidth and scalability of InfiniBand, results in an improvement of overall performance.
Tetsuya Odajima, Taisuke Boku, Toshihiro Hanawa, Hitoshi Murai, Masahiro Nakao, Akihiro Tabuchi, Mitsuhisa Sato
CLUSTER3
2015 Reduction calculator in an FPGA based switching Hub for high performance clusters
abstract
Unused logic in the field-programmable gate array (FPGA) for the switching hub is one potential resource to accelerate the computation of data exchanged through the hub. However, for large scale scientific computation, it is difficult to implement such an accelerator on the FPGA used in high performance computers. Here, a reduction calculator for executing ARGOT (accelerated radiative transfer on grids using oct-tree) to solve the radiative transfer equation used for simulation of astronomical objects is implemented on the FPGA of PEACH2 (PCI Express Adaptive Communication Hub ver2), a low latency switching hub for high performance GPU (graphics processor unit) clusters. The implemented reduction calculator uses a pipelined tree of adders and works with a 150-MHz clock without affecting the switching hub functions. Use of the DMA (direct memory access) transfer with descriptors made it possible to improve the performance of CPU excution by a maximum of about 45 times in a real system.
Takuya Kuhara, Chiharu Tsuruta, Toshihiro Hanawa, Hideharu Amano
FPL3
2013 Task level pipelining with PEACH2: An FPGA switching fabric for high performance computing
abstract
We demonstrate task level pipelining on multiple accelerators with PEACH2. PEACH2 is implmented on FPGA, and enables ultra low latency direct communication among multiple accelerators over computational nodes. By installing PEACH2, typical high performance computation nodes are tightly coupled. In this environment, application can be accelerated by exploiting not only data level parallelism, but also task level pipelined operation. Furthermore, we can processe multiple task on multiple accelerators in a pipelined manner. In our demonstration, application achieves 44% speed up compared to a single GPU.
Takaaki Miyajima, Takuya Kuhara, Toshihiro Hanawa, Hideharu Amano, Taisuke Boku
FPT3
2013 Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage
ICA3PP (2)4
2012 DS-Bench Toolset: Tools for dependability benchmarking with simulation and assurance
abstract
Today's information systems have become large and complex because they must interact with each other via networks. This makes testing and assuring the dependability of systems much more difficult than ever before. DS-Bench Toolset has been developed to address this issue, and it includes D-Case Editor, DS-Bench, and D-Cloud. D-Case Editor is an assurance case editor. It makes a tool chain with DS-Bench and D-Cloud, and exploits the test results as evidences of the dependability of the system. DS-Bench manages dependability benchmarking tools and anomaly loads according to benchmarking scenarios. D-Cloud is a test environment for performing rapid system tests controlled by DS-Bench. It combines both a cluster of real machines for performance-accurate benchmarks and a cloud computing environment as a group of virtual machines for exhaustive function testing with a fault-injection facility. DS-Bench Toolset enables us to test systems satisfactorily and to explain the dependability of the systems to the stakeholders.
Hajime Fujita 0002, Yutaka Matsuno, Toshihiro Hanawa, Mitsuhisa Sato, Shinpei Kato, Yutaka Ishikawa
DSN3
2011 XMCAPI: Inter-core Communication Interface on Multi-chip Embedded Systems
abstract
Multi-core processor technology has been applied to the processors in embedded systems as well as in ordinary PC systems. In multi-core embedded processors, however, a processor may consist of heterogeneous CPU cores that are not configured with a shared memory and do not have a communication mechanism for inter-core communication. MCAPI is a highly portable API standard for providing inter-core communication independent of the architecture heterogeneity. In this paper, we extend the current MCAPI to a multi-chip in a distributed memory configuration and propose its portable implementation, named XMCAPI, on a commodity network stack. With XMCAPI, the inter-core communication method for intra-chip cores is extended to inter-chip cores. We evaluate the XMCAPI implementation, xmcapi/ip, on a standard socket in a portable software development environment.
Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato
EUC2
2010 D-Cloud: Design of a Software Testing Environment for Reliable Distributed Systems Using Cloud Computing Technology
abstract
In this paper, we propose a software testing environment, called D-Cloud, using cloud computing technology and virtual machines with fault injection facility. Nevertheless, the importance of high dependability in a software system has recently increased, and exhaustive testing of software systems is becoming expensive and time-consuming, and, in many cases, sufficient software testing is not possible. In particular, it is often difficult to test parallel and distributed systems in the real world after deployment, although reliable systems, such as high-availability servers, are parallel and distributed systems. D-Cloud is a cloud system which manages virtual machines with fault injection facility. D-Cloud sets up a test environment on the cloud resources using a given system configuration file and executes several tests automatically according to a given scenario. In this scenario, D-Cloud enables fault tolerance testing by causing device faults by virtual machine. We have designed the D-Cloud system using Eucalyptus software and a description language for system configuration and the scenario of fault injection written in XML. We found that the D-Cloud system, which allows a user to easily set up and test a distributed system on the cloud and effectively reduces the cost and time of testing.
Takayuki Banzai, Hitoshi Koizumi, Ryo Kanbayashi, Takayuki Imada, Toshihiro Hanawa, Mitsuhisa Sato
CCGRID5
2010 Customizing Virtual Machine with Fault Injector by Integrating with SpecC Device Model for a Software Testing Environment D-Cloud
abstract
D-Cloud is a software testing environment for dependable parallel and distributed systems using cloud computing technology. We use Eucalyptus as cloud management software to manage virtual machines designed based on QEMU, called FaultVM, which have a fault injection mechanism. D-Cloud enables the test procedures to be automated using a large amount of computing resources in the cloud by interpreting the system configuration and the test scenario written in XML in D-Cloud front end and enables tests including hardware faults by emulating hardware faults by FaultVM flexibly. In the present paper, we describe the customization facility of FaultVM used to add new device models. We use SpecC, which is a system description language, to describe the behavior of devices, and a simulator generated from the description by SpecC is linked and integrated into FaultVM. This also makes the definition and injection of faults flexible without the modification of the original QEMU source codes. This facility allows D-Cloud to be used to test distributed systems with customized devices.
Toshihiro Hanawa, Hitoshi Koizumi, Takayuki Banzai, Mitsuhisa Sato, Shin'ichi Miura, Tadatoshi Ishii, Hidehisa Takamizawa
PRDC1
2009 Flexible Multi-link Ethernet Binding System for PC Clusters with Asymmetric Topology
abstract
In current high-performance PC clusters, the performance and cost of interconnection network are essential issues. Very cost-effective Ethernets, such as Gigabit Ethernet, as well as high performance SANs, such as Infiniband and Myrinet, are still widely used. The authors have been developing a multi-link binding network system for Ethernet, called RI2N, for high-throughput and fault-tolerant interconnection with Gigabit Ethernet. It can be used both for internode communication in MPI programs and traditional UNIX network services such as NFS. In this paper, the authors propose an optimized version of RI2N, called RI2N+, that allows asymmetrical multi-link connection for fitting to various cost-effective system configurations. Such a configuration cannot be supported by Linux Channel Bonding, which is widely used in standard Linux distributions. RI2N+ automatically detects the asymmetric network configuration and controls the traffic distribution to multiple links. In the basic performance evaluation under a high traffic rate, it was confirmed that the throughput of the network with the proposed scheme is improved by approximately 30\% compared with the original RI2N. RI2N+ also maintains high performance even in asymmetric configurations, that is up to 86\% of the relative performance compared with the symmetric case.
Taiga Yonemoto, Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato
ICPADS3
2009 RI2N/DRV: Multi-link ethernet for high-bandwidth and fault-tolerant network on PC clusters
abstract
Although recent high-end interconnection network devices and switches provide a high performance to cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, Linux kernels are equipped with software named Linux Channel Bonding (LCB), which is based IEEE802.3ad Link Aggregation technology. However, standard LCB has the disadvantage of mismatch with the TCP protocol; consequently, both large latency and bandwidth instability can occur. Fault-tolerance feature is supported by LCB, but the usability is not sufficient. We developed a new implementation similar to LCB named Redundant Interconnection with Inexpensive Network with Driver (RI2N/DRV) for use on Gigabit Ethernet. RI2N/DRV has a complete software stack that is very suitable for TCP, an upper layer protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets, even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as network file system (NFS), without any modification of other modules.
Shin'ichi Miura, Toshihiro Hanawa, Taiga Yonemoto, Taisuke Boku, Mitsuhisa Sato
IPDPS2
2009 Towards an Open Dependable Operating System
abstract
This paper introduces a new dependable operating system project, called DEOS, started in 2006, and scheduled to continue for six years. In this project, a safety extension mechanism called P-Bus is to be designed, and implemented in the Linux kernel so that a future dependability attribute is implemented with P-Bus. A hardware abstraction layer, called SPUMONE, is introduced so that a light-weight operating system, called ArcOS, and a monitoring service on top of ArcOS monitors the Linux kernel to provide a safety-net for the Linux kernel. New dependability metrics are being designed to enable developers and users to decide which hardware or software solution meets their dependability requirements, and thus can be used.
Yutaka Ishikawa, Hajime Fujita 0002, Toshiyuki Maeda, Motohiko Matsuda, Midori Sugaya, Mitsuhisa Sato, Toshihiro Hanawa, Shin'ichi Miura, Taisuke Boku, Yuki Kinebuchi, Tatsuo Nakajima, Jin Nakazawa, Hideyuki Tokuda
ISORC7
2008 RI2N: High-bandwidth and fault-tolerant network with multi-link Ethernet for PC clusters
abstract
Although recent high-end interconnection network devices and switches provide a high performance/cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, a Linux kernel is equipped with a software solution named Linux Channel Bonding (LCB), which is based on IEEE802.3ad Link Aggregation technology. However, standard LCB has the problem of mismatching with the commonly used TCP protocol, which consequently implies several problems of both large latency and instability on bandwidth improvement. The fault-tolerant feature is also supported, but the usability is not sufficient. We have developed a new implementation similar to LCB named RI2N/DRV (Redundant Interconnection with Inexpensive Network with Driver) for use on a Gigabit Ethernet with a complete software stack that is very compatible with the TCP protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as NFS, without any modification of other modules.
Shin'ichi Miura, Takayuki Okamoto, Taisuke Boku, Toshihiro Hanawa, Mitsuhisa Sato
CLUSTER4
2008 A dynamic routing control system for high-performance PC cluster with multi-path Ethernet connection
abstract
VLAN-based Flexible, Reliable and Expandable Commodity Network (VFREC-Net) is a network construction technology for PC clusters that allows multi-path network routing to be configured using inexpensive Layer-2 Ethernet switches based on tagged-VLAN technology. Current VFREC-Net system encounters problems with traffic balancing when the communication pattern of the application does not fit the network topology, due to its static routing scheme.
Shin'ichi Miura, Taisuke Boku, Takayuki Okamoto, Toshihiro Hanawa
IPDPS4
2005 The performance of SNAIL-2 (a SSS-MIN connected multiprocessor with cache coherent mechanism)
Takashi Midorikawa, Daisuke Shiraishi, Masayoshi Shigeno, Yasuki Tanabe, Toshihiro Hanawa, Hideharu Amano
Parallel Comput.5
1999 Performance evaluation of SNAIL: A multiprocessor based on the simple serial synchronized multistage interconnection network architecture
Junji Yamamoto, Takashi Fujiwara, T. Komeda, Takayuki Kamei, Toshihiro Hanawa, Hideharu Amano
Parallel Comput.5
1998 The MINC (Multistage Interconnection Network with Cache Control Mechanism) Chip
abstract
Although bus connected multiprocessors have been widely used as high-end workstations or servers, the number of connected processors is strictly limited by the maximum bandwidth of the shared bus. Instead of them, a switch connected multiprocessor which uses a crossbar or Multistage Interconnection Networks (MINs) for connecting processors and memory modules is a hopeful candidate. However, in such a system, a snoop cache technique in bus connected multiprocessors cannot be used, and consistency problems must be solved for providing the cache memory between a processor and the switch. To address this problem, hardware approaches by making the best use of advanced VLSI technology have been proposed. However, traditional methods require a large memory outside the switching element and it causes not only a large additional hardware but also the extra latency by accessing the outside memory. Moreover, the complicated MIN with cache or directory must also treat data packet which should be transferred quickly. In order to solve these problems, we proposed the MINC (MIN with Cache control mechanism). In the MINC, the MIN which only transfers a part of the address and cache coherent messages is separated from the data transfer network, and pushed into an LSI chip called the MINC chip. The coherent control is done based on the directory using the reduced Hierarchical Bit-map Directory scheme (RHBD). In order to reduce unnecessary packets, the pruning cache which is a small cache enough to implement inside the chip is introduced in the MINC chip.
Takashi Midorikawa, Takayuki Kamei, Toshihiro Hanawa, Hideharu Amano
ASP-DAC3
1994 Multistage Interconnection Networks with Multiple Outlets
abstract
Multistage Interconnection Networks(MINs) with multiple outlets are networks which can support higher bandwidth than that of nonblocking networks by passing multiple packets to the same destination. A novel MIN topology with multiple outlets called Piled Banyan Switching Fabrics (PBSF) is proposed for the Simple Serial Synchronized (SSS)-MIN used in multiprocessors, and analyzed with other two types of MIN with multiple outlets called Multi-Banyan Switching Fabrics (MBSF) and Tandem Banyan Switching Fabrics (TBSF). The throughput of these MINs is evaluated and compared with both the theoretical model and simulation. The PBSF supports the best throughput and latency used for the SSS-MIN. Although the latency of the TBSF is large, the pass-through ratio is close to 1 if the number of connected banyan networks are more than 4. Therefore, the TBSF is useful for the ATM switching networks in which the relatively large latency is tolerable. The conflict-free access of these MINs is also analyzed, and it appears that rows, column, forward and backward diagonal of the matrix can be accessed without conflict.
Toshihiro Hanawa, Hideharu Amano, Yoshifumi Fujikawa
ICPP (1)1