Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tomohiro Kudoh

dblp:02/6204 · DBLP profile ↗
← Back
36ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 2 first-authorSoftware engineering, systems software and programming languages · 3Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Interconnection networks and networks-on-chip · 72% High-performance computing · 23% Parallel and multicore computing · 3%
Computer networks
1 paper
Internet architecture and protocols · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › routing algorithms
deadlock-free routing
0.112011
A Switch-Tagged Routing Methodology for PC Clusters with VLAN Ethernet · IEEE Trans. Parallel Distributed Syst. 2011
High-performance computing › cluster computing
PC cluster
0.112011
A Switch-Tagged Routing Methodology for PC Clusters with VLAN Ethernet · IEEE Trans. Parallel Distributed Syst. 2011
Interconnection networks and networks-on-chip
low-latency communication
0.112007
Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCs · IEEE Trans. Parallel Distributed Syst. 2007
Interconnection networks and networks-on-chip › network interface
network interface card
0.112007
Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCs · IEEE Trans. Parallel Distributed Syst. 2007
Interconnection networks and networks-on-chip
remote direct memory access
0.112007
Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCs · IEEE Trans. Parallel Distributed Syst. 2007
Interconnection networks and networks-on-chip
cluster interconnect
0.022007
A Local Area System Network RHinet-1: A Network for High Performance Parallel Computing · HPDC 2000
Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCs · IEEE Trans. Parallel Distributed Syst. 2007
Internet architecture and protocols › network topology
ethernet layer-2 topology
0.012011
A Switch-Tagged Routing Methodology for PC Clusters with VLAN Ethernet · IEEE Trans. Parallel Distributed Syst. 2011
Interconnection networks and networks-on-chip
optical interconnection networks
0.012000
A Local Area System Network RHinet-1: A Network for High Performance Parallel Computing · HPDC 2000
Parallel and multicore computing › multiprocessor system
large-scale multiprocessor
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Interconnection networks and networks-on-chip
multicast
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Processor architecture and microarchitecture
multiprocessor architecture
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Parallel and multicore computing › parallel computing
parallel scientific computing
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
High-performance computing › sparse linear algebra
sparse matrix computation
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Parallel and multicore computing › parallel programming models
concurrent programming languages
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Parallel and multicore computing
parallel programming models
0.011990
(SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations · IEEE Trans. Computers 1990
Electronic design automation
circuit analysis
0.011985
(SM)²-II: A New Version of the Sparse Matrix Solving Machine · ISCA 1985
High-performance computing › sparse linear algebra
sparse linear systems
0.011985
(SM)²-II: A New Version of the Sparse Matrix Solving Machine · ISCA 1985

Methods — techniques the papers use, named apart from their topics

deadlock-free routing algorithms · 0.2VLAN tagging · 0.2hardwired logic · 0.1PIO-based communication · 0.1CPLD protocol controller · 0.0CMOS switches · 0.0simulation · 0.0performance analysis · 0.0
YearPublicationVenuePosition
2019 Optimizing Weight Value Quantization for CNN Inference
abstract
The size and complexity of CNN models are increasing and as a result they are requiring more computational and memory resources to be used effectively. Use of a lower bit width numerical representation such as binary, ternary or several bit width has been studied extensively so as to reduce the required resources. However, the representation capability of such extremely low bit width is not always sufficient and the accuracy obtained for some CNN models and data is low. There are some prior studies that use moderate lower bit width with well-known numerical representations such as fixed point or logarithmic representation. It is not apparent, however, whether those representations are optimal for maintaining high accuracy. In this paper, we investigated the numerical quantization from the ground up, and introduced a novel "Variable Bin-size Quantization (VBQ)" representation in which quantization bin boundaries are optimized to obtain maximum accuracy for each CNN model. A genetic algorithm was employed to optimize the bin boundaries of VBQ. Additionally, since the appropriate bit width to obtain sufficient accuracy cannot be determined in advance, we attempted to use the parameters obtained by a training process using higher precision representation (FP32), and used quantization in inference only. This reduced the required large computational resource cost for training. During the process of tuning VBQ boundaries using a genetic algorithm, we discovered that the optimal distribution of bins can be approximated by an equation with two parameters. We then used simulated annealing for finding the optimal parameters of the equation for AlexNet and VGG16. As a result, AlexNet and VGG16 with our 4-bit quantization achieved top-5 accuracy at 74.8% and 86.3% respectively, which were comparable to 76.3% and 88.1% obtained by FP32. Thus, VBQ combined with the approximate equation and the simulated annealing scheme can achieve similar levels of accuracy with less resources and reduced computational cost compared to other current approaches.
Wakana Nogami, Tsutomu Ikegami, Shin-ichi O'Uchi, Ryousei Takano, Tomohiro Kudoh
IJCNN5
2018 Image-Classifier Deep Convolutional Neural Network Training by 9-bit Dedicated Hardware to Realize Validation Accuracy and Energy Efficiency Superior to the Half Precision Floating Point Format
abstract
We propose a 9-bit floating point format for training image-classifier deep convolutional neural networks. The proposed floating point format has a 5-bit exponent, a 3-bit mantissa with the hidden most significant bit (MSB), and a sign bit. The 9-bit floating point format reduces not only the transistor count of the multiplier in the multipy-accumulate (MAC) unit, but also the data traffic for the forward and backward propagations and the weight update. Both of the reductions realize a power efficient training. To maintain the validation accuracy, the accumulator is implemented with an internal longer-bit-length floating point format while the multiplier accepts the 9-bit format. We examined this format in the training of the AlexNet and the ResNet-50 with the ILSVRC 2012 data set. The trained 9-bit AlexNet and ResNet-50 exhibited the validation accuracy superior to the 16-bit floating point format training by 1.2 % and 0.5 %, respectively. The transistor count in the 9-bit MAC unit is estimated to be reduced by 84% as compared to the 32-bit counterpart.
Shin-ichi O'Uchi, Hiroshi Fuketa, Tsutomu Ikegami, Wakana Nogami, Takashi Matsukawa, Tomohiro Kudoh, Ryousei Takano
ISCAS6
2017 A Case of Electrical Circuit Switched Interconnection Network for Parallel Computers
abstract
Circuit switching is a way to minimize network latency and maximize network bandwidth when a limited number of source-and-destination pairs exchange messages which are predictable. Although there are a large number of studies of optical circuit switching (OCS) on HPC systems and datacenters, it is still not mature. In this context, we explore the use of electrical circuit switching (ECS) for the low-latency purpose on HPC systems and datacenters. ECS has the same link bandwidth as existing electrical packet switched networks, and inherits quick update of input-and-output connections from electrical switches. We develop a network topology generator for ECS to minimize the number of time slots optimized to target applications whose traffic patterns are predictable. By performing a quantitative discrete-event simulation, we present that an ideal ECS network outperforms counterpart EPS networks. Evaluation results show that the minimum necessary number of slots (MNNS) can be reduced to a small number in a generated topology while keeping resource amount less than that in a standard mesh network.
Yao Hu 0005, Tomohiro Kudoh, Michihiro Koibuchi
PDCAT2
2016 Flow-centric computing leveraged by photonic circuit switching for the post-moore era
abstract
Artificial Intelligence (AI)-based big data analysis is emerging for broad application fields, including manufacturing, autonomous car, health care, agriculture, and so on. Data centers are under severe pressure to meet ever-increasing demands on clouds. On the other hand, the amount of data processing is bounded by the total power budget, where the reasonable number is up to 20 MW. Therefore, the energy efficiency becomes a critical issue both now and the future. A combination of special-purpose processing units communicating through huge bandwidth optical circuit switched network allows to dramatically improve the energy efficiency of big data processing. Based on this idea, we propose flow-centric computing, a software-defined data center architecture focusing on data flow processing. Compute, storage, and network resources are disaggregated and dynamically composes slices, i.e., data processing environments, based on workload-specific demands. We are conducting a feasibility study of the concept and developing system software technologies.
Ryousei Takano, Tomohiro Kudoh
NOCS2
2014 Iris: An Inter-cloud Resource Integration System for Elastic Cloud Data Centers
abstract
This paper proposes a new cloud computing service model, Hardware as a Service (HaaS), that is based on the idea of implementing ``elastic data centers'' that provide a data center administrator with resources located at different data centers as demand requires. To demonstrate the feasibility of the proposed model, we have developed what we call an Inter-cloud Resource Integration System (Iris) by using nested virtualization and OpenFlow technologies. Iris dynamically configures and provides a virtual infrastructure over inter-cloud resources, on which an IaaS cloud can run. Using Iris, we have confirmed an IaaS cloud can seamlessly extend and manage resources over multiple data centers. The experimental results on an emulated inter-cloud environment show that the overheads of the HaaS layer are acceptable when the network latency is less than 10 msec. We believe these results provide new insight to help establish inter-cloud computing.
Ryousei Takano, Atsuko Takefusa, Hidemoto Nakada, Seiya Yanagita, Tomohiro Kudoh
CLOSER5
2012 Stream processing with BigData: SSS-MapReduce
abstract
We propose a Map Reduce based stream processing system, called SSS, which is capable of processing stream along with large scale static data. Unlike the existing stream processing systems that can work only on the relatively small on-memory data-set, SSS can process incoming streamed data consulting the stored data. SSS processes streamed data with continuous Mappers and Reducers, which are periodically invoked by the system. It also supports merge operation on two sets of data, which enables stream data processing with large static data. This poster shows overview of SSS stream processing and preliminary evaluation results.
Hidemoto Nakada, Hirotaka Ogawa, Tomohiro Kudoh
CloudCom3
2012 Virtual Machine packing algorithms for lower power consumption
abstract
Virtual Machine(VM)-based flexible capacity management is an effective scheme to reduce total power consumption in the data centers. However, there remain the following issues, trade-off between power-saving and user experience, decision on VM packing plans within a feasible calculation time, and collision avoidance for multiple VM live migration processes. In order to resolve these issues, we propose two VM packing algorithms, a matching-based (MBA) and a greedy-type heuristic (GREEDY). MBA enables to decide an optimal plan in polynomial time, while GREEDY is an aggressive packing approach faster than MBA. We investigate the basic performance and the feasibility of proposed algorithms under both artificial and realistic simulation scenarios, respectively. The basic performance experiments show that the algorithms reduce total power consumption by between 18% and 50%, and MBA makes suitable VM packing plans within a feasible calculation time. The feasibility experiments show that the proposed algorithms are feasible to make packing plans for an actual supercomputer, and GREEDY has the advantage in power consumption, but MBA shows the better performance in user experience.
Satoshi Takahashi, Hidemoto Nakada, Atsuko Takefusa, Tomohiro Kudoh, Maiko Shigeno, Akiko Yoshise
CloudCom4
2012 A distributed application execution system for an infrastructure with dynamically configured networks
abstract
We have been developing a middleware suite called GridARS that enables co-allocation of computing and network resources from multiple administration sites. In such middleware, it is important to provide each user application with a slice which is a set of dynamically allocated resources distributed across sites. However, there are the following issues in constructing such a slice automatically: 1) multi-site administration heterogeneity, 2) dynamic determination of application configuration information, 3) distributed resource monitoring, and 4) asymmetric network reachability. We design and implement an application execution system that provides each application with a slice, that mimics a conventional computing cluster system over the dynamically allocated resources. From the demonstration of the proposed system on an emulated wide area network environment, we confirmed that: first, the proposed system can fully automate resource allocation, slice construction, application invocation, and resource monitoring, in coordination with GridARS. Second, the proposed system can setup a slice quickly, even if the allocated resources are widely distributed and their communication latencies are high. This is because the overhead for gathering and distributing contextualization information is small, and OS-level virtualization and stackable file system technologies accelerate the contextualization process at each node.
Ryousei Takano, Hidemoto Nakada, Atsuko Takefusa, Tomohiro Kudoh
CloudCom4
2012 Cooperative VM migration for a virtualized HPC cluster with VMM-bypass I/O devices
abstract
An HPC cloud, a flexible and robust cloud computing service specially dedicated to high performance computing, is a promising future e-Science platform. In cloud computing, virtualization is widely used to achieve flexibility and security. Virtualization makes migration or checkpoint/restart of computing elements (virtual machines) easy, and such features are useful for realizing fault tolerance and server consolidations. However, in widely used virtualization schemes, I/O devices are also virtualized, and thus I/O performance is severely degraded. To cope with this problem, VMM-bypass I/O technologies, including PCI passthrough and SR-IOV, in which the I/O overhead can be significantly reduced, have been introduced. However, such VMM-bypass I/O technologies make it impossible to migrate or checkpoint/restart virtual machines, since virtual machines are directly attached to hardware devices. This paper proposes a novel and practical mechanism, called Symbiotic Virtualization (SymVirt), for enabling migration and checkpoint/restart on a virtualized cluster with VMM-bypass I/O devices, without the virtualization overhead during normal operations. SymVirt allows a VMM to cooperate with a message passing layer on the guest OS, then it realizes VM-level migration and checkpoint/restart by using a combination of a PCI hotplug and coordination of distributed VMMs. We have implemented the proposed mechanism on top of QEMU/KVM and the Open MPI system. All PCI devices, including Infiniband and Myrinet, are supported without implementing specific para-virtualized drivers; and it is not necessary to modify either of the MPI runtime and applications. Using the proposed mechanism, we demonstrate reactive and proactive FT mechanisms on a virtualized Infiniband cluster. We have confirmed the effectiveness using both a memory intensive micro benchmark and the NAS parallel benchmark. Moreover, we also show that postcopy live migration enables us to reduce the down time of an application as the memory footprint increases.
Ryousei Takano, Hidemoto Nakada, Takahiro Hirofuchi, Yoshio Tanaka, Tomohiro Kudoh
eScience5
2011 GridARS: A Grid Advanced Resource Management System Framework for Intercloud
abstract
Intercloud is a promising technology for data intensive applications. However, an important issue for Intercloud applications is orchestration of various virtualized and performance-assured resources, not only computers, but also network and storage, provided from multiple domains. We have been developing an advance reservation-based resource management framework, called Grid ARS, which can integrate heterogeneous resources and construct a performance-assured virtual infrastructure over Intercloud environment. Grid ARS provides four services that address resource management, resource allocation planning, provisioning and monitoring of the constructed virtual infrastructure. Grid ARS has been developed using common Web services technologies and standards. In this paper, we present overview of Grid ARS and its service components and describe Grid ARS demonstration challenges, demonstration at GLIF2010 and SC10 and OGF NSI interoperation in 2011.
Atsuko Takefusa, Hidemoto Nakada, Ryousei Takano, Tomohiro Kudoh, Yoshio Tanaka
CloudCom4
2011 A Switch-Tagged Routing Methodology for PC Clusters with VLAN Ethernet
abstract
Ethernet has been used for connecting hosts in PC clusters, besides its use in local area networks. Although a layer-2 Ethernet topology is limited to a tree structure because of the need to avoid broadcast storms and deadlocks of frames, various deadlock-free routing algorithms on topologies that include loops suitable for parallel processing can be employed by the application of IEEE 802.1Q VLAN technology. However, the MPI communication libraries used in current PC clusters do not always support tagged VLAN technology; therefore, at present, the design of VLAN-based Ethernet cannot be applied to such PC clusters. In this study, we propose a switch-tagged routing methodology in order to implement various deadlock-free routing algorithms on such PC clusters by using at most the same number of VLANs as the degree of a switch. Since the MPI communication libraries do not need to perform VLAN operations, the proposed methodology has advantages in both simple host configuration and high portability. In addition, when it is used with on/off and multispeed link regulation, the power consumption of Ethernet switches can be reduced. Evaluation results using NAS parallel benchmarks showed that the performance of the topologies that include loops using the proposed methodology was comparable to that of an ideal one-switch (full crossbar) network, and the torus topology in particular had up to a 27 percent performance improvement compared with a tree topology with link aggregation.
Michihiro Koibuchi, Tomohiro Otsuka, Tomohiro Kudoh, Hideharu Amano
IEEE Trans. Parallel Distributed Syst.3
2010 SSS: An Implementation of Key-Value Store Based MapReduce Framework
abstract
MapReduce has been very successful in implementing large-scale data-intensive applications. Because of its simple programming model, MapReduce has also begun being utilized as a programming tool for more general distributed and parallel applications, e.g., HPC applications. However, its applicability is limited due to relatively inefficient runtime performance and hence insufficient support for flexible workflow. In particular, the performance problem is not negligible in iterative MapReduce applications. On the other hand, today, HPC community is going to be able to utilize very fast and energy-efficient Solid State Drives (SSDs) with 10 Gbit/sec-class read/write performance. This fact leads us to the possibility to develop "High-Performance MapReduce'', so called. From this perspective, we have been developing a new MapReduce framework called "SSS'' based on distributed key-value store (KVS). In this paper, we first discuss the limitations of existing MapReduce implementations and present the design and implementation of SSS. Although our implementation of SSS is still in a prototype stage, we conduct two benchmarks for comparing the performance of SSS and Hadoop. The results indicate that SSS performs 1-10 times faster than Hadoop.
Hirotaka Ogawa, Hidemoto Nakada, Ryousei Takano, Tomohiro Kudoh
CloudCom4
2010 An Advance Reservation-Based Co-allocation Algorithm for Distributed Computers and Network Bandwidth on QoS-Guaranteed Grids
Atsuko Takefusa, Hidemoto Nakada, Tomohiro Kudoh, Yoshio Tanaka
JSSPP3
2008 High Performance Relay Mechanism for MPI Communication Libraries Run on Multiple Private IP Address Clusters
abstract
We have been developing a Grid-enabled MPI communication library called GridMPI, which is designed to run on multiple clusters connected to a wide-area network. Some of these clusters may use private IP addresses. Therefore, some mechanism to enable communication between private IP address clusters is required. Such a mechanism should be widely adoptable, and should provide high communication performance. In this paper, we propose a message relay mechanism to support private IP address clusters in the manner of the Interoperable MPI (IMPI) standard. Therefore, any MPI implementations which follow the IMPI standard can communicate with the relay. Furthermore, we also propose a trunking method in which multiple pairs of relay nodes simultaneously communicate between clusters to improve the available communication bandwidth. While the relay mechanism introduces an one-way latency of about 25 musec, the extra overhead is negligible, since the communication latency through a wide area network is a few hundred times as large as this. By using trunking, the inter-cluster communication bandwidth can improve as the number of trunks increases. We confirmed the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. When relay nodes with 1 Gbps NICs are used, the performance of most of the NAS Parallel Benchmarks improved proportional to the number of trunks. Especially, using 8 trunks, FT and IS are 4.4 and 3.4 times faster, respectively, compared with the single trunk case. The results showed that the proposed method is effective for running MPI programs over high bandwidth-delay product networks.
Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa, Yasufumi Yoshizawa
CCGRID3
2007 Effects of packet pacing for MPI programs in a Grid environment
abstract
Improving the performance of TCP communication is the key to the successful deployment of MPI programs in a Grid environment in which multiple clusters are connected through high performance dedicated networks. To efficiently utilize the inter-cluster bandwidth, a traffic control mechanism is required so as not to allow the aggregate transmission bandwidth to exceed the inter-cluster bandwidth when multiple nodes communicate at one time. In this paper, we propose a traffic control method for MPI programs, in which an application or the MPI runtime controls the transmission rate based on the communication pattern by using certain MPI attributes. Packet pacing is used at each node preventing microscopic burst transmission to thus avoid congestion. We confirm the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. We show most of the NAS Parallel benchmarks improve the performance, since the proposed method reduces packet losses due to traffic congestion on the inter-cluster network. The results have indicated that it is feasible to connect multiple clusters and run large-scale scientific applications over distances up to 1000 kilometers, if an appropriate network is available.
Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa
CLUSTER3
2007 Topic 13 High-Performance Networks
Thilo Kielmann, Pascale Vicat-Blanc Primet, Tomohiro Kudoh, Bruce Lowekamp
Euro-Par3
2007 GridARS: An Advance Reservation-Based Grid Co-allocation Framework for Distributed Computing and Network Resources
Atsuko Takefusa, Hidemoto Nakada, Tomohiro Kudoh, Yoshio Tanaka, Satoshi Sekiguchi
JSSPP3
2007 Martini: A Network Interface Controller Chip for High Performance Computing with Distributed PCs
abstract
In this paper, “Martini,” a network interface controller chip for our original network called RHiNET is described. Martini is designed to provide high-bandwidth and low-latency communication with small overhead. To obtain high performance communication, protected user-level zero-copy RDMA communication functions are completely implemented by a hardwired logic. Also, to reduce the communication latency efficiently, we have proposed PIO-based communication mechanisms called “On-the-fly (OTF)” and have implemented them on Martini. The evaluation results show that Martini connected to a 64bit/66MHz PCI-bus achieves 470MByte/s maximum bidirectional bandwidth and 1.74 μsec minimum latency on host-to-host memory copying.
Konosuke Watanabe, Tomohiro Otsuka, Junichiro Tsuchiya, Hiroaki Nishi, Junji Yamamoto, Noboru Tanabe, Tomohiro Kudoh, Hideharu Amano
IEEE Trans. Parallel Distributed Syst.7
2006 Efficient MPI Collective Operations for Clusters in Long-and-Fast Networks
abstract
Several MPI systems for grid environment, in which clusters are connected by wide-area networks, have been proposed. However, the algorithms of collective communication in such MPI systems assume relatively low bandwidth wide-area networks, and they are not designed for the fast wide-area networks that are becoming available. On the other hand, for cluster MPI systems, a beast algorithm by van de Geijn et al. and an allreduce algorithm by Rabenseifner have been proposed, which are efficient in a high bisection bandwidth environment. We modify those algorithms so as to effectively utilize fast wide-area inter-cluster networks and to control the number of nodes which can transfer data simultaneously through wide-area networks to avoid congestion. We confirmed the effectiveness of the modified algorithms by experiments using a 10 Gbps emulated WAN environment. The environment consists of two clusters, where each cluster consists of nodes with 1 Gbps Ethernet links and a switch with a 10 Gbps upper link. The two clusters are connected through a 10 Gbps WAN emulator which can insert latency. In a 10 millisecond latency environment, when the message size is 32 MB, the proposed beast and allreduce are 1.6 and 3.2 times faster, respectively, than the algorithms used in existing MPI systems for grid environment
Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa
CLUSTER2
2006 Switch-tagged VLAN Routing Methodology for PC Clusters with Ethernet
abstract
Ethernet has been used for connecting hosts in the area of high performance-per-cost PC clusters. Although L2 Ethernet topology is limited to a tree structure, various routing algorithms on topologies suitable for parallel processing can be employed by applying IEEE 802.1Q VLAN technology. However, communication library used in PC clusters does not always support VLANs, so the design of VLAN-based routing method cannot be applied for such PC clusters. In this paper, we propose a switch-tagged VLAN methodology to flexibly set the route of frames on such PC clusters. Since each host does not need to process VLAN tags, the proposed method has advantages in both simple host configuration and high portability. Evaluation results using NAS Parallel Benchmarks showed that performance of topologies supported by the proposed method was comparable with that of an ideal 1-switch (full crossbar) network in the case of a 16-host PC cluster
Tomohiro Otsuka, Michihiro Koibuchi, Tomohiro Kudoh, Hideharu Amano
ICPP3
2006 G-lambda: Coordination of a Grid scheduler and lambda path service over GMPLS
Atsuko Takefusa, Michiaki Hayashi, Naohide Nagatsu, Hidemoto Nakada, Tomohiro Kudoh, Takahiro Miyamoto, Tomohiro Otani, Hideaki Tanaka, Masatoshi Suzuki, Yasunori Sameshima
Future Gener. Comput. Syst.5
2005 TCP Adaptation for MPI on Long-and-Fat Networks
abstract
Typical MPI applications work in phases of computation and communication, and messages are exchanged in relatively small chunks. This behavior is not optimal for TCP because TCP is designed only to handle a contiguous flow of messages efficiently. This behavior anomaly is well-known, but fixes are not integrated into today's TCP implementations, even though performance is seriously degraded, especially for MPI applications. This paper proposes three improvements in the Linux TCP stack: i.e., pacing at start-up, reducing Retransmit-Timeout time, and TCP parameter switching at the transition of computation phases in an MPI application. Evaluation of these improvements using the NAS parallel benchmarks shows that the BT, CG, IS, and SP benchmarks achieved 10 to 30 percent improvements. On the other hand, the FT and MG benchmarks showed no improvement because they have the steady communication that TCP assumes, and the LU benchmark became slightly worse because it has very little communication
Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa
CLUSTER2
2004 GNET-1: gigabit Ethernet network testbed
abstract
GNET-1 is a fully programmable network testbed. It provides functions such as wide area network emulation, network instrumentation, traffic shaping, and traffic generation at gigabit Ethernet wire speeds by programming the core FPGA. GNET-1 is a powerful tool for developing network-aware grid software. It is also a network monitoring and traffic-shaping tool that provides high-performance communication over wide area networks. This work describes several sample uses of GNET-1 and presents its architecture.
Yuetsu Kodama, Tomohiro Kudoh, Ryousei Takano, Hitoshi Sato, Osamu Tatebe, Satoshi Sekiguchi
CLUSTER2
2004 The design and implementation of an asynchronous communication mechanism for the MPI communication model
abstract
Many implementations of an MPI communication library are realized on top of the socket interface which is based on connection-oriented stream communication. This work addresses a mismatch between the MPI communication model and the socket interface. In order to overcome a mismatch and implement an efficient MPI library for large-scale commodity-based clusters, a new communication mechanism, called 02G, is designed and implemented. O2G integrates receive queue management of MPI into a TCP/IP protocol handler, without modifying the protocol stacks. Received data is extracted from the TCP receive buffer and copied into the user space within the TCP/IP protocol handler invoked by interrupts. It totally avoids polling of sockets and reduces system call overhead, which becomes dominant in large-scale clusters. In addition, its immediate and asynchronous receive operation avoids message flow disruption due to a shortage of capacity in the receive buffer, and keeps the bandwidth high. An evaluation using the NAS Parallel Benchmarks shows that 02G made an MPI implementation up to 30 percent faster than the original one. An evaluation on bandwidth also shows that 02G made an MPI implementation independent of the number of connections, while an implementation with sockets was greatly affected by the number of connections.
Motohiko Matsuda, Tomohiro Kudoh, H. Tazuka, Yutaka Ishikawa
CLUSTER2
2003 Evaluation of MPI Implementations on Grid-connected Clusters using an Emulated WAN Environmen
abstract
The MPICH-SCore high performance communication library for cluster computing is integrated into the MPICHG-2 library in order to adapt PC clusters to a Grid environment. The integrated library is called MPICH-G2/SCore. In addition, for the purpose of comparison with other approaches, MPICH-SCore itself is extended to encapsulate its network packet into a UDP packet so that packets are delivered via L3 switches. This extension is called UDP-encapsulated MPICH-SCore. In this paper, three implementations of the MPI library, UDP-encapsulated MPICH-SCore, MPICH-G2/SCore, and MPICH-P4, are evaluated using an emulated WAN environment where two clusters, each consisting of sixteen hosts, are connected by a router PC. The router PC controls the latency of message delivery between clusters, and the added latency is varied from I millisecond to 4 milliseconds in round-trip time. Experiments are performed using the NAS Parallel Benchmarks, which show UDP-encapsulated MPICH-SCore most often performs better than other implementations. However, the differences are not critical for the benchmarks. The preliminary results show that the performance of the LU benchmark scales up linearly with under 4 millisecond round-trip latency. The CG and MG benchmarks show the scalability of 1.13 and 1.24 times with 4 millisecond round-trip latency, respectively.
Motohiko Matsuda, Tomohiro Kudoh, Yutaka Ishikawa
CCGRID2
2003 Design and implementation of PVFS-PM: a cluster file system on SCore
abstract
This paper discusses the design and implementation of a cluster file system, called PVFS-PM, on the SCore cluster system software. This is the first attempt to implement a cluster file system on the SCore system. It is based on the PVFS cluster file system but replaces TCP with the PMv2 communication library supported by SCore to provide a scalable, high-performance cluster file system. PVFS-PM improves the performance by factors of 1.07 and 1.93 for writing and reading, respectively, with 8 I/O nodes, compared with the original PVFS on TCP on a Gigabit Ethernet-connected SCore cluster.
Koji Segawa, Osamu Tatebe, Yuetsu Kodama, Tomohiro Kudoh, Toshiyuki Shimizu
CCGRID4
2003 Performance Evaluation of RHiNET-2/NI: A Network Interface for Distributed Parallel Computing Systems
abstract
RHiNET-2/NI is a network interface for a parallel and distributed computing system with network connected PCs. The core of the network interface is an ASIC network controller chip Martini, which provides low-latency and large-bandwidth communication. Evaluation results show that it achieves almost full bandwidth of the 66MHz/64bit PCI bus, which is much larger than that of Myrinet-2000. The performance of a small prototype parallel system achieves almost linear speed up.
Konosuke Watanabe, Tomohiro Otsuka, Junichiro Tsuchiya, Hideharu Amano, Hiroshi Harada, Junji Yamamoto, Hiroaki Nishi, Tomohiro Kudoh
CCGRID8
2000 MEMOnet : Network interface plugged into a memory slot
abstract
The communication architecture of the DIMMnet-1 network interface, based on MEMOnet, is described. MEMOnet is an architecture consisting of a network interface plugged into a memory slot. The DIMMnet-1 prototype will have two banks of PC133 based SO-DIMM slots and an 8 Gbps full duplex optical link or two 448 MB/s full duplex LVDS channel links. The software overhead incurred to generate a message is only I CPU cycle and the estimated hardware delay is less than 100 ns using the atomic on-the-fly sending with header TLB. The estimated achievable communication bandwidth with block on-the-fly sending with protection stampable window memory is 440 MB/s which was observed in our experiments writing to the DIMM area with a write combining attribute. This is 3.3 times higher than the maximum bandwidth of PCI. This high performance distributed computing environment is available using economical personal computers with DIMM slots.
Noboru Tanabe, Junji Yamamoto, Hiroaki Nishi, Tomohiro Kudoh, Yoshihiro Hamada, Hironori Nakajo, Hideharu Amano
CLUSTER4
2000 A Local Area System Network RHinet-1: A Network for High Performance Parallel Computing
abstract
The Real World Computing Partnership (RWCP) has developed a local area system network (LASN) called RHiNET-1 (RWCP High-performance NETwork, version 1) using 1.33-Gbps optical interconnections for high-performance computing using personal computers distributed in an office or laboratory environment. The network interface, RHiNET-1/NI, uses a complex programmable logic device (CPLD) based protocol controller to provide an easy evaluation platform for various protocols. It fits in a 32-bit/33-MHz PCI bus. The switch, RHiNET-1/SW, consists of a single-chip CMOS switch and external SRAM. It provides low-latency, reliable communication with a flexible topology design. We are currently evaluating protocols on RHiNET-1. RHiNET-1 will enable a new form of high-performance computing environment. We are also developing the second implementation, RHiNET-2. RHiNET-2/NI will support a 64-bit/66-MHz PCI bus. RHiNET-2/SW is an 8-Gbps/port 8/spl times/8 single-chip ASIC switch. The aggregate bandwidth of RHiNET-2/SW is 64 Gbps.
Hiroaki Nishi, Koji Tasho, Junji Yamamoto, Tomohiro Kudoh, Hideharu Amano
HPDC4
1997 The RDT network router chip
abstract
The RDT network router chip is a versatile router for the massively parallel computer prototype JUMP-1. The major goal of this project is to establish techniques for building an efficient distributed shared memory on a massively parallel processor. For this purpose, the reduced hierarchical bit-map directory (RHBD) schemes are used for efficient cache management of the distributed shared memory. In order to implement (RHBD) schemes efficiently, we proposed a novel interconnection network RDT (recursive diagonal torus), and developed a sophisticated router chip for the RDT which equips a hierarchical multicast mechanism without deadlock and acknowledge combining mechanism. By using the 0.5/spl mu/BiCMOS SOG technology it can transfer all packets synchronized with a unique CPU clock(60MHz). Long coaxial cables are directly driven with the ECL interface of this chip. The mixed design approach with schematic and VHDL permits the development of the complicated chip with 90,522 gates in a year.
Hiroaki Nishi, Hideharu Amano, Katsunobu Nishimura, Kenichiro Anjo, Tomohiro Kudoh
ASP-DAC5
1995 Test Suite Generation Methods for Concurrent Systems Based on Colored Petri Nets
abstract
Automatic generation of test suites for concurrent systems is a newly exploited area of conformance tests. A few methods based on a finite state machine (FSM) have been proposed. However, these methods require a large amount of computation cost. By using coloured Petri nets (CPN), the required amount of computation costs can be reduced, and the length of the test suites can be reduced by using the equivalent marking technique on CPN. In addition, these methods allow us to test interaction parameters of concurrent systems. We propose two CPN based test suite generation methods for conformance tests: the coloured Petri net tree (CPT) method and the coloured Petri net graph (CPG) method. An experimental test suite generator (TSG) based on CPT method is presented. To show the advantages of the CPT and CPG methods, the effects of test suite length reduction by equivalent markings are evaluated.
Harumi Watanabe, Tomohiro Kudoh
APSEC2
1995 Hierarchical Bit-Map Directory Schemes on the RDT Interconnection Network for a Massively Parallel Processor JUMP-1
Tomohiro Kudoh, Hideharu Amano, Takashi Matsumoto 0002, Kei Hiraki, Yulu Yang, Katsunobu Nishimura, Koichi Yoshimura, Yasuhito Fukushima
ICPP (1)1
1995 A Performance Evaluation of the Multiprocessor Testbed ATTEMPT-0
Takuya Terasawa, Ou Yamamoto, Tomohiro Kudoh, Hideharu Amano
Parallel Comput.3
1992 A Parallel Logic Simulation Algorithm Based on Query
Tomohiro Kudoh, Tetsuro Kimura, Hideharu Amano, Takuya Terasawa
ICPP (3)1
1990 (SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculations
abstract
(SM)/sup 2/-II is a large-scale parallel machine dedicated to scientific computation which includes sparse matrix calculations. In order to connect thousands of microprocessors and utilize a high degree of parallelism, the whole (SM)/sup 2/-II system is designed based on a simple computational model called the node and connecting-line (NC) model. The concept and the architecture of (SM)/sup 2/-II are described. The NC-model and a language called node oriented concurrent C (NCC) are derived. The concurrent process controller is briefly introduced. Receiver selectable multicast (RSM) is proposed, and the structure which allows connection of a large number of processing units is described. The performance of the RSM is analyzed. Some connection structures for clusters are evaluated. An operational prototype is introduced.>
Hideharu Amano, Taisuke Boku, Tomohiro Kudoh
IEEE Trans. Computers3
1985 (SM)²-II: A New Version of the Sparse Matrix Solving Machine
abstract
article Free Access Share on (SM)2-II: a new version of the sparse matrix solving machine Authors: Hideharu Amano Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Taisuke Boku Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Tomohiro Kudoh Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Hideo Aiso Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile Authors Info & Claims ACM SIGARCH Computer Architecture NewsVolume 13Issue 3June 1985 pp 100–107https://doi.org/10.1145/327070.327137Published:01 June 1985Publication History 16citation208DownloadsMetricsTotal Citations16Total Downloads208Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Hideharu Amano, Taisuke Boku, Tomohiro Kudoh, Hideo Aiso
ISCA3