EDBT 2026 Demo / reviewers in the wild / expert
Atsushi Hori
dblp:81/7017
· DBLP profile ↗
42ranked-venue papers
15as first author
5since 2021 · last 2023
0000-0002-7010-8098ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 9 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PiP-MColl: Process-in-Process-based Multi-object MPI CollectivesabstractIn the era of exascale computing, the adoption of a large number of CPU cores and nodes by high-performance computing (HPC) applications has made MPI collective performance increasingly crucial. As the number of cores and nodes increases, the importance of optimizing MPI collective performance becomes more evident. Current collective algorithms, including kernel-assisted inter-process data exchange techniques and data sharing based shared-memory approaches, are prone to significant performance degradation due to the overhead of system calls and page faults or the cost of extra data-copy latency. These issues can negatively impact the efficiency and scalability of HPC applications. To address these issues, we propose PiP-MColl, a Process-in-Process-based Multi-object Interprocess MPI Collective design that maximizes small message MPI collective performance at scale. We also present specific designs to boost the performance for larger messages, such that we observe a comprehensive improvement for a series of message sizes beyond small messages. PiP-MColl features efficient multiple sender and receiver collective algorithms and leverages Process-in-Process shared memory techniques to eliminate unnecessary system call, page fault overhead and extra data copy, which results in improved intra- and inter-node message rate and throughput. Experimental results demonstrate that PiP-MColl significantly outperforms popular MPI libraries, including OpenMPI, MVAPICH2, and Intel MPI, by up to 4.6X for the MPI collectives MPI_Scatter, MPI_Allgather, and MPI_Allreduce. Jiajun Huang 0001, Kaiming Ouyang, Jinyang Liu 0003, Min Si, Kenneth Raffenetti, Hui Zhou 0012, Atsushi Hori, Zizhong Chen, Yanfei Guo, Rajeev Thakur |
CLUSTER | 8 |
| 2023 | Accelerating MPI Collectives with Process-in-Process-based Multi-object TechniquesabstractIn the exascale computing era, optimizing MPI collective performance in high-performance computing (HPC) applications is critical. Current algorithms face performance degradation due to system call overhead, page faults, or data-copy latency, affecting HPC applications' efficiency and scalability. To address these issues, we propose PiP-MColl, a Process-in-Process-based Multi-object Inter-process MPI Collective design that maximizes small message MPI collective performance at scale. PiP-MColl features efficient multiple sender and receiver collective algorithms and leverages Process-in-Process shared memory techniques to eliminate unnecessary system call, page fault overhead, and extra data copy, improving intra- and inter-node message rate and throughput. Our design also boosts performance for larger messages, resulting in comprehensive improvement for various message sizes. Experimental results show that PiP-MColl outperforms popular MPI libraries, including OpenMPI, MVAPICH2, and Intel MPI, by up to 4.6X for MPI collectives like MPI_Scatter and MPI_Allgather. Jiajun Huang 0001, Kaiming Ouyang, Jinyang Liu 0003, Min Si, Kenneth Raffenetti, Hui Zhou 0012, Atsushi Hori, Zizhong Chen, Yanfei Guo, Rajeev Thakur |
HPDC | 8 |
| 2021 | Daps: A Dynamic Asynchronous Progress Stealing Model for MPI CommunicationabstractMPI provides nonblocking point-to-point and one-sided communication models to help applications achieve communication and computation overlap. These models provide the opportunity for MPI to offload data transfer to low level network hardware while the user process is computing. In practice, however, MPI implementations have to often handle complex data transfer in software due to limited capability of network hardware. Therefore, additional asynchronous progress is necessary to ensure prompt progress of these software-handled communication. Traditional mechanisms either spawn an additional background thread on each MPI process or launch a fixed number of helper processes on each node. Both mechanisms may degrade performance in user computation due to statically occupied CPU resources. The user has to fine-tune the progress resource deployment to gain overall performance. For complex multiphase applications, unfortunately, severe performance degradation may occur due to dynamically changing communication characteristics and thus changed progress requirement. This paper proposes a novel Dynamic Asynchronous Progress Stealing model, called Daps, to completely address the asynchronous progress complication. Daps is implemented inside the MPI runtime. It dynamically leverages idle MPI processes to steal communication progress tasks from other busy computing processes located on the same node. The basic concept of Daps is straightforward; however, various implementation challenges have to be resolved due to the unique requirements of interprocess data and code sharing. We present our design that ensures high performance while maintaining strict program correctness. We compare Daps with state-of-the-art asynchronous progress approaches by utilizing both microbenchmarks and HPC proxy applications. Kaiming Ouyang, Min Si, Atsushi Hori, Zizhong Chen, Pavan Balaji |
CLUSTER | 3 |
| 2021 | A Scalability Study of Data Exchange in HPC Multi-component WorkflowsabstractMulti-component workflows play a significant role in High-Performance Computing and Big Data applications. They usually contain multiple, independently developed components that execute side-by-side to perform sophisticated computation and data exchange through file I/O over parallel file system. However, file I/O can become an impediment in such systems and cause undesirable performance degradation due to its relatively low speed (compared to the interconnect fabric), which is unacceptable especially for applications with strict time constraints. The Data Transfer Framework (DTF) is an I/O arbitration layer working with the PnetCDF I/O library aiming at eliminating the bottleneck by transparently redirecting file I/O operations through the parallel file system to message passing via the high-speed interconnect between coupled components. Scalable and high-speed data transfer between components can be thus easily achieved with minimal development effort by using DTF. However, previous work provides insufficient scalability evaluation of the framework. In order to comprehensively evaluate the scalability of an I/O middleware like DTF and highlight its major advantages, we develop an I/O benchmark for multicomponent workflows. Using the benchmark we conduct large-scale scalability evaluation using up to 32,768 compute nodes on supercomputer Fugaku and 2,048 compute nodes on Oakforest-PACS by comparing direct data transfer to file I/O performed on Lustre file system and Fugaku’s Lightweight Layered IO-Accelerator (LLIO). We provide insights into DTF’s scalability and performance enhancements with the intention to impact future I/O middleware and inter-component data exchange design in multi-component workflows. Atsushi Hori, Balazs Gerofi, Yutaka Ishikawa |
CLUSTER | 2 |
| 2021 | An international survey on MPI users
Atsushi Hori, Emmanuel Jeannot, George Bosilca, Takahiro Ogura, Balazs Gerofi, Yutaka Ishikawa |
Parallel Comput. | 1 |
| 2020 | CAB-MPI: exploring interprocess work-stealing towards balanced MPI communicationabstractLoad balance is essential for high-performance applications. Unbalanced communication can cause severe performance degradation, even in computation-balanced BSP applications. Designing communication-balanced applications is challenging, however, because of the diverse communication implementations at the underlying runtime system. In this paper, we address this challenge through an interprocess workstealing scheme based on process-memory-sharing techniques. We present CAB-MPI, an MPI implementation that can identify idle processes inside MPI and use these idle resources to dynamically balance communication workload on the node. We design throughput-optimized strategies to ensure efficient stealing of the data movement tasks. We demonstrate the benefit of work stealing through several internal processes in MPI, including intranode data transfer, pack/unpack for noncontiguous communication, and computation in one-sided accumulates. The implementation is evaluated through a set of microbenchmarks and proxy applications on Intel Xeon and Xeon Phi platforms. Kaiming Ouyang, Min Si, Atsushi Hori, Zizhong Chen, Pavan Balaji |
SC | 3 |
| 2019 | Comparing the performance of rigid, moldable and grid-shaped applications on failure-prone HPC platforms
Valentin Le Fèvre, Thomas Hérault, Yves Robert, Aurelien Bouteiller, Atsushi Hori, George Bosilca, Jack J. Dongarra |
Parallel Comput. | 5 |
| 2018 | Improving Collective MPI-IO Using Topology-Aware Stepwise Data Aggregation with I/O ThrottlingabstractMPI-IO has been used in an internal I/O interface layer of HDF5 or PnetCDF, where collective MPI-IO plays a big role in parallel I/O to manage a huge scale of scientific data. However, existing collective MPI-IO optimization named two-phase I/O has not been tuned enough for recent supercomputers consisting of mesh/torus interconnects and a huge scale of parallel file systems due to lack of topology-awareness in data transfers and optimization for parallel file systems. In this paper, we propose I/O throttling and topology-aware stepwise data aggregation in two-phase I/O of ROMIO, which is a representative MPI-IO library, in order to improve collective MPI-IO performance even if we have multiple processes per compute node. Throttling I/O requests going to a target file system mitigates I/O request contention, and consequently I/O performance improvements are achieved in file access phase of two-phase I/O. Topology-aware aggregator layout with paying attention to multiple aggregators per compute node alleviates contention in data aggregation phase of two-phase I/O. In addition, stepwise data aggregation improves data aggregation performance. HPIO benchmark results on the K computer indicate that the proposed optimization has achieved up to about 73% and 39% improvements in write performance compared with the original implementation using 12,288 and 24,576 processes on 3,072 and 6,144 compute nodes, respectively. Yuichi Tsujita, Atsushi Hori, Toyohisa Kameyama, Atsuya Uno, Fumiyoshi Shoji, Yutaka Ishikawa |
HPC Asia | 2 |
| 2018 | Process-in-process: techniques for practical address-space sharingabstractThe two most common parallel execution models for many-core CPUs today are multiprocess (e.g., MPI) and multithread (e.g., OpenMP). The multiprocess model allows each process to own a private address space, although processes can explicitly allocate shared-memory regions. The multithreaded model shares all address space by default, although threads can explicitly move data to thread-private storage. In this paper, we present a third model called process-in-process (PiP), where multiple processes are mapped into a single virtual address space. Thus, each process still owns its process-private storage (like the multiprocess model) but can directly access the private storage of other processes in the same virtual address space (like the multithread model). Atsushi Hori, Min Si, Balazs Gerofi, Masamichi Takagi, Jai Dayal, Pavan Balaji, Yutaka Ishikawa |
HPDC | 1 |
| 2018 | Deep Learning on Large-Scale Muticore ClustersabstractConvolutional neural networks (CNNs) have achieved outstanding accuracy among conventional machine learning algorithms. Recent works have shown that large and complicated models, which take significant cost for training are needed to get higher accuracy. To train these models efficiently in high performance computers (HPCs), many parallelization techniques for CNNs have been developed. However, most techniques are mainly targeting GPUs and parallelizations for CPUs are not fully investigated. This paper explores CNN training performance on large-scale multicore clusters by optimizing intra-node processing and applying techniques of inter-node parallelization for multiple GPUs. Detailed experiments conducted on state-of-the-art multi-core processors using the openMP API and MPI framework demonstrated that Caffe-based CNNs can be accelerated by using well-designed multithreaded programs. We achieved at most 1.64 times speedup in convolution operations with devised lowering strategy compared to conventional lowering and acquired 772 times speedup with 864 nodes compared to one node. Kazumasa Sakivama, Shinpei Kato, Yutaka Ishikawa, Atsushi Hori, Abraham Monrroy Cano |
SBAC-PAD | 4 |
| 2016 | On the Scalability, Performance Isolation and Device Driver Transparency of the IHK/McKernel Hybrid Lightweight KernelabstractExtreme degree of parallelism in high-end computing requires low operating system noise so that large scale, bulk-synchronous parallel applications can be run efficiently. Noiseless execution has been historically achieved by deploying lightweight kernels (LWK), which, on the other hand, can provide only a restricted set of the POSIX API in exchange for scalability. However, the increasing prevalence of more complex application constructs, such as in-situ analysis and workflow composition, dictates the need for the rich programming APIs of POSIX/Linux. In order to comply with these seemingly contradictory requirements, hybrid kernels, where Linux and a lightweight kernel (LWK) are run side-by-side on compute nodes, have been recently recognized as a promising approach. Although multiple research projects are now pursuing this direction, the questions of how node resources are shared between the two types of kernels, how exactly the two kernels interact with each other and to what extent they are integrated, remain subjects of ongoing debate. In this paper, we describe IHK/McKernel, a hybrid software stack that seamlessly blends an LWK with Linux by selectively offloading system services from the lightweight kernel to Linux. Specifically, we are focusing on transparent reuse of Linux device drivers and detail the design of our framework that enables the LWK to naturally leverage the Linux driver codebase without sacrificing scalability or the POSIX API. Through rigorous evaluation on a medium size cluster we demonstrate how McKernel provides consistent, isolated performance for simulations even in face of competing, in-situ workloads. Balazs Gerofi, Masamichi Takagi, Atsushi Hori, Gou Nakamura, Tomoki Shirasawa, Yutaka Ishikawa |
IPDPS | 3 |
| 2015 | Sliding Substitution of Failed NodesabstractThis paper considers the questions of how spare nodes should be allocated, how to substitute them for faulty nodes, and how much the communication performance is affected by such a substitution. The third question stems from the modification of the rank mapping by node substitutions, which can incur additional message collisions. In a stencil computation, rank mapping is done in a straightforward way on a Cartesian network without incurring any message collisions. However, once a substitution has occurred, the node- rank mapping may be destroyed. Therefore, these questions must be answered in a way that minimizes the degradation of communication performance. Atsushi Hori, Kazumi Yoshinaga, Thomas Hérault, Aurelien Bouteiller, George Bosilca, Yutaka Ishikawa |
EuroMPI | 1 |
| 2014 | Interface for heterogeneous kernels: A framework to enable hybrid OS designs targeting high performance computing on manycore architecturesabstractTurning towards exascale systems and beyond, it has been widely argued that the currently available systems software is not going to be feasible due to various requirements such as the ability to deal with heterogeneous architectures, the need for systems level optimization targeting specific applications, elimination of OS noise, and at the same time, compatibility with legacy applications. To cope with these issues, a hybrid design of operating systems where light-weight specialized kernels can cooperate with a traditional OS kernel seems adequate, and a number of recent research projects are now heading into this direction. This paper presents Interface for Heterogeneous Kernels (IHK), a general framework enabling hybrid kernel designs in systems equipped with manycore processors and/or accelerators. IHK provides a range of capabilities, such as resource partitioning, management of heterogeneous OS kernels, as well as a low-level communication layer among the kernels. We describe IHK's interface and demonstrate its feasibility for hybrid kernel designs through executing various different lightweight OS kernels on top of it, which are specialized for certain types of applications. We use the Intel Xeon Phi, Intel's latest manycore coprocessor, as our experimental platform. Taku Shimosawa, Balazs Gerofi, Masamichi Takagi, Gou Nakamura, Tomoki Shirasawa, Yuji Saeki, Masaaki Shimizu, Atsushi Hori, Yutaka Ishikawa |
HiPC | 8 |
| 2014 | CMCP: a novel page replacement policy for system level hierarchical memory management on many-coresabstractThe increasing prevalence of co-processors such as the Intel Xeon Phi, has been reshaping the high performance computing (HPC) landscape. The Xeon Phi comes with a large number of power efficient CPU cores, but at the same time, it's a highly memory constraint environment leaving the task of memory management entirely up to application developers. To reduce programming complexity, we are focusing on application transparent, operating system (OS) level hierarchical memory management. Balazs Gerofi, Akio Shimada, Atsushi Hori, Masamichi Takagi, Yutaka Ishikawa |
HPDC | 3 |
| 2014 | Multithreaded Two-Phase I/O: Improving Collective MPI-IO Performance on a Lustre File SystemabstractROMIO, a representative MPI-IO implementation, has been widely used in recent large-scale parallel computations. The two-phase I/O optimization scheme of ROMIO improves I/O performance for non-contiguous access patterns, however, this scheme still has room to improve performance to make it suitable for recent data-intensive computing. We propose overlapping data exchange operations with file I/O operations by using a multithreaded scheme to achieve further I/O throughput improvement. We show up to 60% improvement by the multithreaded two-phase I/O relative to the original two-phase I/O in performance evaluation of collective write operations on a Lustre file system of a Linux PC cluster. Yuichi Tsujita, Kazumi Yoshinaga, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
PDP | 3 |
| 2013 | Partially Separated Page Tables for Efficient Operating System Assisted Hierarchical Memory Management on Heterogeneous ArchitecturesabstractHeterogeneous architectures, where a multicore processor is accompanied with a large number of simpler, but more power-efficient CPU cores optimized for parallel workloads, are receiving a lot of attention recently. At present, these co-processors, such as the Intel Xeon Phi product family, come with limited on-board memory, which requires partitioning computational problems manually into pieces that can fit into the device's RAM, as well as efficiently overlapping computation and communication. In this paper we propose an application transparent, operating system (OS)assisted hierarchical memory management system, where the OS orchestrates data movement between the host and the device and updates the process virtual memory address space accordingly. We identify the main scalability issues of frequent address space changes, such as the increasing price of TLB invalidations with the growing number of CPU cores, and propose partially separated page tables with address-range CPU masks to overcome the problem. With partially separated page tables each core maintains its own set of mappings of the computation area, enabling the OS to perform address space updates in a scalable manner, and involve a particular CPU core in TLB invalidation only if it is absolutely necessary. Furthermore, we propose dedicated data movement cores in order to efficiently overlap computation and communication. We provide experimental results on stencil computation, a common HPCkernel, and show that OS assisted memory management has the potential for scalable transparent data movement. Balazs Gerofi, Akio Shimada, Atsushi Hori, Yutaka Ishikawa |
CCGRID | 3 |
| 2013 | A Delegation Mechanism on Many-Core Oriented Hybrid Parallel Computers for Scalability of Communicators and Communications in MPIabstractThis paper describes a delegation based high throughput MPIcommunication mechanism under tough memory utilization constrains on a many-core oriented hybrid parallel computer. Towards the Exascale era, hybrid parallel computers consisting of many-core and multi-core architectures both on the same node are focused. Although many-core architectures such as GPU or Intel MIC has high potential in computing power by the large number of computing cores, per-core computing power is lower than that of multi-core CPUs. Furthermore, available memory resources for the many-core CPUs are quite smaller than those for multi-core CPUs. Thus we may have a sort of penalty in memory utilization in MPI communications when we utilize a normal MPI library. Here we deploy a delegatee process on each node to merge MPI communications and minimize memory utilization for an MPI communicator. Another advantage of the delegatee process scheme is minimization of memory utilization on many-core CPUs by delegating MPI requests to associated delegatee process on multi-core CPUs. In this paper, we show performance advantages and effective resource utilization by our proposed scheme compared with the original MPI implementation. Kazumi Yoshinaga, Yuichi Tsujita, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
PDP | 3 |
| 2013 | Optimization of MPI persistent communicationabstractThis paper proposes a novel optimization technique for MPI persistent communication, that utilizes multiple RDMA engines to carry out low latency communication. Because the interconnects used in modern supercomputers have multiple RDMA engines, the multiple communication requests specified by a persistent communication invocation can be scheduled onto the RDMA engines in an optimal way and thus result in better communication performance. Such a scheduling algorithm is not only a packing problem, but also avoids interconnect resource contentions as much as possible. The scheduling algorithm proposed in this paper balances load of RDMA engines and mitigates network link contentions in case of neighbor communication patterns such like in stencil computation. The proposed scheduling algorithm is implemented in Open MPI of K computer using RDMA functions provided as Fujitsu MPI extensions. A typical 2D stencil computation of a climate simulation code is used as a benchmark program. The experimental result shows that a factor of two speedup of communication time is achieved. Masayuki Hatanaka, Atsushi Hori, Yutaka Ishikawa |
EuroMPI | 2 |
| 2013 | Revisiting rendezvous protocols in the context of RDMA-capable host channel adapters and many-core processorsabstractWe revisit RDMA-based rendezvous protocols in MPI in the context of cluster computer with RDMA-capable HCA and many-core processors, and propose two improved protocols. The conventional sender-initiate rendezvous protocols cause costly processor-device communications via PCI bus on detecting completion of RDMA transfer. The conventional receiver-initiate rendezvous protocols need to send extra control messages when a value of the memory-slot to poll in the receive buffer has the same value as the send buffer. The first proposed protocol implements polling on a memory-slot in the receive buffer to eliminate the processor-device communications. The second proposed protocol randomizes the value of the memory-slot to poll to reduce extra control messages. We have evaluated the proposed protocols using micro-benchmarks and NAS Parallel Benchmarks. One of the proposed protocols has a benefit compared to the conventional protocols. And the second proposed protocol reduces the execution time by up to 11.14% compared to the first protocol. Masamichi Takagi, Yuichi Nakamura 0002, Atsushi Hori, Balazs Gerofi, Yutaka Ishikawa |
EuroMPI | 3 |
| 2012 | clone_n(): Parallel Thread Creation for Upcoming Many-Core ArchitecturesabstractHeterogeneous architectures, where a multicore processor, which is optimized for fast single-thread performance, is accompanied with a large number of simpler, but more power-efficient cores optimized for parallel workloads, such as NVIDIA's GPUs or Intel's Many Integrated Core (MIC), have been receiving a lot attention recently. Although NVIDIA's GPUs include built-in support for parallelism control, the MIC uses classical software thread creation and scheduling done by the operating system (OS). While efficient thread creation is desired in such many-core environments, current OS APIs provide the facility of creating only one thread at a time. In this paper, we propose a new system call for parallel thread creation on many-core coprocessors and show that it can perform up to 6.9 times better than the sequential version when executed on Intel's MIC software development platform. Balazs Gerofi, Atsushi Hori, Yutaka Ishikawa |
CLUSTER | 2 |
| 2012 | An Efficient Kernel-Level Blocking MPI Implementation
Atsushi Hori, Toyohisa Kameyama, Yuichi Tsujita, Mitaro Namiki, Yutaka Ishikawa |
EuroMPI | 1 |
| 2012 | Revisiting Persistent Communication in MPI
Yutaka Ishikawa, Kengo Nakajima, Atsushi Hori |
EuroMPI | 3 |
| 2012 | Delegation-Based MPI Communications for a Hybrid Parallel Computer with Many-Core Architecture
Kazumi Yoshinaga, Yuichi Tsujita, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
EuroMPI | 3 |
| 2012 | Audit: A new synchronization API for the GET/PUT protocol
Atsushi Hori, Jinpil Lee, Mitsuhisa Sato |
J. Parallel Distributed Comput. | 1 |
| 2011 | Xruntime: A Seamless Runtime Environment for High Performance ComputingabstractMost HPC clusters are now based on an x86 architecture and Linux. In a Grid consisting of such clusters, users might think that a single executable file, a data set, and a single job script will work in all cluster environments. However, due to the lack of interoperability of MPI library implementations, file systems, and batch job systems, users need to be conscious of the runtime environments over the Grid. In order to overcome such differences, we propose a runtime environment called Xruntime consisting of the following three components. MPI-Adapter is middleware between the user program and an MPI implementation to make it possible to execute a single executable code on different MPI implementations. Catwalk is transparent file staging middleware that makes remote files accessed by an application visible as if they were located in the local system. Xruntime allows a user to use various clusters with single batch job script language the user is familiar with, without learning other batch job script languages used on other clusters. Keiji Yamamoto, Atsushi Hori, Shinji Sumimoto, Yutaka Ishikawa |
HPCC | 2 |
| 2011 | Catwalk-ROMIO: A Cost-Effective MPI-IOabstractThe nature of highly parallelized parallel file access which often consists of lots of fine grain, non-contiguous I/O requests, can degrade the I/O performance severely. To tackle this problem, a novel technique to maximize the bandwidth of the MPI-IO is proposed. This proposed technique is utilize a ring communication topology. This technique is implemented as an ADIO device of ROMIO, named Catwalk-ROMIO, and evaluated. The evaluation shows that Catwalk-ROMIO utilizing only one disk can exhibit comparable performance with parallel files systems, PVFS2 and Lustre, utilizing several file servers and disks. The evaluation also shows that Catwalk-ROMIO performance is almost independent from file access patterns, in contrast to the performance of parallel file systems performing only well with collective I/O. Catwalk-ROMIO only requires TCP/IP network for the ring communication topology and one file server which are common in HPC clusters without any additional cost. Thus, Catwalk-ROMIO is considered to be a very cost-effective MPI-IO implementation. Atsushi Hori, Keiji Yamamoto, Yutaka Ishikawa |
ICPADS | 1 |
| 2009 | On-demand file staging system for Linux clustersabstractAn on-demand file staging system, Catwalk, is proposed. Catwalk is designed so that it can run on any Linux clusters without any special or additional hardware. By having hook functions on the system calls of file operations, a file staging system can be transparent from the view of users, and users can be free from having wrong file staging scripts. In Catwalk, the file copying is done via normal TCP protocol so that Catwalk can run over ordinary, widely-used Ethernet. The stage-in file copy is pipelined to maximize the bandwidth from single file server. The performance of Catwalk is evaluated and compared with NFS using synthetic but realistic workloads. The evaluations show the stage-in performance with the pipeline technique is much better than the performance of NFS. The stage-out performance is comparable with the NFS performance despite the extra copying of files, and the file server is lightly loaded with the Catwalk stage-out while NFS entails much heavier server loads. The biggest problems of NFS are its centralized design and lack of scheduling for the parallel workloads. The performance of Catwalk shows that remote file access performance can be improved much better if file accesses are scheduled in a proper way. Thus the proposed file staging system can be a strong complement to NFS, especially for small clusters often having no dedicated parallel file system. Atsushi Hori, Yoshikazu Kamoshida, Hiroya Matsuba, Kazuki Ohta, Takashi Yasui, Shinji Sumimoto, Yutaka Ishikawa |
CLUSTER | 1 |
| 2001 | SCore: An Integrated Cluster System Software Package for High Performance Cluster ComputingabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. SCore Cluster System is an integrated software package for high performance clusters, and is used by worldwide users. SCore and the PMv2 high performance communication library, which supports various networks available today, was designed simultaneously. As a result, SCore can provide not only high performance cluster computing environment, but also a number of cluster programming facilities including Ethernet trunking, network interleaving, gang-scheduling, checkpoint and restart, real-time monitoring of user programs, deadlock detection, debugging support, etc. SCore Cluster System Software is developed by the Real World Computing (RWC) Project funded by Japanese government. The RWC project will be terminated by the end of March, 2002, however, PC Cluster Consortium is formed and the development of SCore Cluster System Software is continued by the consortium. In this talk, after the brief introduction of SCore Cluster System Software, target and the current status of the PC Cluster Consortium will be presented. Atsushi Hori |
CLUSTER | 1 |
| 2000 | SCore: An Integrated Cluster System Software Package for High Performance Cluster Computing
Atsushi Hori |
CLUSTER | 1 |
| 2000 | Consistent Checkpointing for High Performance Clusters
Toshihiro Nishioka, Atsushi Hori, Yutaka Ishikawa |
CLUSTER | 2 |
| 2000 | High Performance Communication using a Commodity Network for Cluster SystemsabstractProposes a scheme to realize a high-performance communication facility using a commodity network. This scheme does not require any special hardware or hardware-specific device drivers in order to adapt to many kinds of network interface cards (NICs). In this scheme, a reliable lightweight network protocol is handled directly on a data link layer called by a network device driver. An interrupt reaping technique is proposed to eliminate the hardware interrupt overhead when an application waits for a message. PM/Ethernet, an instance of the scheme, is implemented on Linux with minimal modification to the Linux kernel, and existing network device drivers are used without any modification. Using Pentium III 500-MHz PCs on Packet Engine's G-NIC II Gigabit Ethernet NIC, it achieves 77.5 MB/s bandwidth and 37.6 /spl mu/s round-trip time latency compared to that of TCP/IP, which achieves 46.7 MB/s bandwidth and 89.6 /spl mu/s round-trip time latency. The NAS parallel benchmark IS results show that MPI on PM/Ethernet achieves 75% better performance than MPI on TCP/IP and is 7.8% slower than that of MPI on Myrinet PM. Shinji Sumimoto, Hiroshi Tezuka, Atsushi Hori, Hiroshi Harada, Toshiyuki Takahashi, Yutaka Ishikawa |
HPDC | 3 |
| 2000 | PM2: A High Performance Communication Middleware for Heterogeneous Network EnvironmentsabstractThis paper introduces a high performance communication middle layer, called PM2, for hetero-geneous network environments. PM2 currently supports Myrinet, Ethernet, and SMP. Binary code written in PM2 or written in a communication library, such as MPICH-SCore on top of PM2, may run on any combination of those networks without re-compilation. According to a set of NAS parallel benchmark results, MPICH-SCore performance is better than dedicated communication libraries such as MPICH-BIP/SMP and MPICH-GM when running some benchmark programs. Toshiyuki Takahashi, Shinji Sumimoto, Atsushi Hori, Hiroshi Harada, Yutaka Ishikawa |
SC | 3 |
| 1999 | The design and evaluation of high performance communication using a Gigabit EthernetabstractA high performance communication facility, called the GigaE PM, has been designed and implemented for parallel applications on clusters of computers using a Gigabit Ethernet. The GigaE PM provides not only a reliable high bandwidth and low latency communication function, but also supports existing network protocols such as TCP/IP. In the design of the GigaE PM, it is assumed that the Gigabit Ethernet card used has a dedicated processor and its program can be modied. A reliable communication mechanism for a parallel application is implemented on the rmware while existing network protocols are handled by an operating system kernel. A prototype system has been implemented using an Essential Communications Gigabit Ethernet card. The performance results show that a 48.3 s round trip time for a four byte user message, and 56.7 MBytes/sec bandwidth for a 1,468 byte message have been achieved on Intel Pentium II 400 MHz PCs. We have implemented MPICH-PM on top of the GigaE PM, and evaluat... Shinji Sumimoto, Hiroshi Tezuka, Atsushi Hori, Hiroshi Harada, Toshiyuki Takahashi, Yutaka Ishikawa |
International Conference on Supercomputing | 3 |
| 1998 | The Design and Implementation of Zero Copy MPI Using Commodity Hardware with a High Performance NetworkabstractThis paper designs an implementation of the MPI message passing interface using a zero copy message transfer primitive supported by a lower communication layer to realize a high performance communication library.The zero copy message transfer primitive requires a memory area pinned down to physical memory, which is a restricted quantity resource under a paging memory system.Allocation of pinned down memory by multiple simultaneous requests for sending and receiving without any control can cause deadlock.To avoid this deadlock, we have introduced: i) separate of control of send/receive pin-down memory areas to ensure that at least one send and receive may be processed concurrently, and ii) delayed queues to handle the postponed message passing operations which could not be pinned-down. Francis O'Carroll, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa |
International Conference on Supercomputing | 3 |
| 1998 | Overhead Analysis of Preemptive Gang Scheduling
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
JSSPP | 1 |
| 1998 | Highly Efficient Gang Scheduling ImplementationabstractA new and more highly efficient gang scheduling implementation technique is the basis for this paper. Network preemption, in which network interface contexts are saved and restored, has already been proposed to enable parallel applications to perform efficent user-level communication. This network preemption technique can be used to for detecting global state, such as deadlock, of a parallel program execution. A gang scheduler, SCore-D, using the network preemption technique is implemented with PM, a user-level communication library. This paper evaluates network preemption gang scheduling overhead using eight NAS parallel benchmark programs. The results of this evaluation illustrate that the saving and restoring network contexts occupies almost half of the total gang scheduling overhead. A new mechanism, having multiple network contexts and merely switching the context pointers without saving and restoring the network contexts, is proposed. The NAS parallel benchmark evaluation shows that gang scheduling overhead is almost halved. The maximum gang scheduling overhead among benchmark programs is less than 10%, with a 40msec time slice on 64 single-way PentiumPros, connected by Myrinet to form a PC cluster. The numbers of secondary cache misses are counted, and it is found that network preemption with multiple network contexts is more cache-effective than a single network context. The observed scheduling overhead for applications running on 64 nodes can only be a small percent of the execution time. The gang scheduling overheads of switching two NAS parallel benchmark programs are also evaluated. The additional overheads are less than 2% in most cases, with a 100msec time slice on 64 nodes. This slightly higher scheduling overheads than for switching a single parallel process comes from more frequent cache misses. This paper contributes the following findings; i) gang scheduling overhead with network preemption can be sufficiently low, ii) proposed network preemption with multiple network contexts is more cache-effective than a single network context, and, iii) network preemption can be applied to detect global states of user parallel processes. SCore-D gang scheduler realized by network preemption can utilize processor resources by the detecting the global state of user parallel processes. Network preemption with multiple contexts exhibits highly efficient gang scheduling. The combination of low scheduling overhead and the global state detection mechanism achieves an interactive parallel programming where parallel program development and the production run of parallel programs can be mixed freely. Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
SC | 1 |
| 1998 | Ninf and PM: Communication libraries for global computing and high-performance cluster computing
Mitsuhisa Sato, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa, Satoshi Sekiguchi, Hidemoto Nakada, Satoshi Matsuoka, Umpei Nagashima |
Future Gener. Comput. Syst. | 3 |
| 1997 | Global State Detection Using Network Preemption
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
JSSPP | 1 |
| 1996 | Implementation of Gang-Scheduling on Workstation Cluster
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa, Noriyuki Soda, Hiroki Konaka, Munenori Maeda |
JSSPP | 1 |
| 1995 | Time Space Sharing Scheduling: A Simulation Analysis
Atsushi Hori, Yutaka Ishikawa, Jörg Nolte, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo |
Euro-Par | 1 |
| 1995 | A prototype router for the massively parallel computer RWC-1abstractThe RWC-1 is a massively parallel computer based on a multi-threaded architecture. This architecture requires extremely high communication performance with reasonable hardware cost. ln this paper, we first introduce a new class of direct interconnection networks called MDCE (Multidimensional Directed Cycles Ensemble extension). MDCE has many desirable features for RWC-1 including small degree, low latency, and high throughput. MDCE is thus adopted for a RWC-1 network. We have designed an MDCE router and fabricated an experimental VLSI chip. We explain the design details in this paper. The chip employs operating system support features as well as communication functions, and enables advanced resource management, A prototype chip with about 125,000 gates has been fabricated using 0.6-/spl mu/m CMOS gate array technology. Its clock runs at 50 MHz and a transmission rate of 300 M bytes per second per communication port is achieved. Takashi Yokota, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono, Atsushi Hori, Shuichi Sakai |
ICCD | 5 |
| 1995 | Time Space Sharing Scheduling and Architectural Support
Atsushi Hori, Takashi Yokota, Yutaka Ishikawa, Shuichi Sakai, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo, Jörg Nolte, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono |
JSSPP | 1 |