VLDB 2026 Research / reviewers in the wild / expert
Mitsuhisa Sato
dblp:04/2564
· DBLP profile ↗
119ranked-venue papers
8as first author
17since 2021 · last 2027
0000-0003-0543-7116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 101 · 7 first-author · 15 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorArtificial intelligence and machine learning · 3 · 2 since 2021Security and privacy · 3Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Closed-loop calculations of electronic structure on a quantum processor and a classical supercomputer at full scaleabstractQuantum computers must operate in concert with classical computers to deliver on the promise of quantum advantage for practical problems. To achieve that, it is important to understand how quantum and classical computing can interact together, and how one can characterize the scalability and efficiency of hybrid quantum–classical workflows. So far, early experiments with quantum-centric supercomputing workflows have been limited in scale and complexity. Here, we use a Heron quantum processor deployed on premises with the entire supercomputer Fugaku to perform the largest computation of electronic structure involving quantum and classical high-performance computing. We design a closed-loop workflow between the quantum processors and 152,064 classical nodes of Fugaku, to approximate the electronic structure of chemistry models beyond the reach of exact diagonalization, with accuracy comparable to some all-classical approximation methods. Our work pushes the limits of the integration of quantum and classical high-performance computing, showcasing computational resource orchestration at the largest scale possible for current classical supercomputers. Tomonori Shirakawa, Javier Robledo Moreno, Toshinari Itoko, Vinay Tripathi, Kento Ueda, Yukio Kawashima, Lukas Broers, William M. Kirby, Himadri Pathak, Hanhee Paik, Miwako Tsuji, Yuetsu Kodama, Mitsuhisa Sato, Constantinos Evangelinos, Seetharami Seelam, Robert Walkup, Seiji Yunoki, Mario Motta, Petar Jurcevic, Hiroshi Horii, Antonio Mezzacapo |
Future Gener. Comput. Syst. | 13 |
| 2025 | Massively Parallel CMA-ES With Increasing PopulationabstractABSTRACT The Increasing Population Covariance Matrix Adaptation Evolution Strategy (IPOP‐CMA‐ES) algorithm is a reference stochastic optimizer dedicated to blackbox optimization, where no prior knowledge about the underlying problem structure is available. This paper aims to accelerate IPOP‐CMA‐ES thanks to high‐performance computing and parallelism when solving large optimization problems. We first show how BLAS and LAPACK routines can be introduced in linear algebra operations, and we then propose two strategies for deploying IPOP‐CMA‐ES efficiently on large‐scale parallel architectures with up to thousands of CPU cores. The first parallel strategy processes the multiple searches in the same ordering as the sequential IPOP‐CMA‐ES, while the second one processes concurrently these multiple searches. These strategies are implemented in MPI+OpenMP and compared on 6144 cores of the supercomputer Fugaku. We manage to obtain substantial speedups (up to several thousand) and even super‐linear ones, and we provide an in‐depth analysis of our results to understand precisely the superior performance of our second strategy. These results are finally confirmed on a local compute cluster with 512 cores. David Redon, Pierre Fortin 0001, Bilel Derbel, Miwako Tsuji, Mitsuhisa Sato |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | A Parallel and Asynchronous Approach for Anomaly DetectionabstractThis article addresses the pressing need for accurate anomaly detection techniques, particularly for cybersecurity applications. We emphasize the effectiveness of ensemble and machine learning techniques, as well as the parallelizability of the Unite and Conquer approach, to improve the efficiency, speed, and accuracy of calculations. More precisely, we introduce a variant of an existing framework for its optimization by taking into account the asynchronicity of the communications and evaluate its large-scale performance on the Fugaku supercomputer. Our evaluation focuses on the detection rate and response time of expertise using extensive datasets, including the UNSW-NB15 dataset, in the cybersecurity domain. Additionally, we discuss the framework’s expanded functionality and its potential integration into existing Security Orchestration, Automation, and Response (SOAR) systems, thereby strengthening cyber threat detection and response capabilities. Zineb Ziani, Nahid Emad, Miwako Tsuji, Mitsuhisa Sato, Ahmed Bouaziz |
IEEE Big Data | 4 |
| 2024 | Advancements in Traffic Simulations with multiMATSim's Distributed FrameworkabstractInternational audience Sara Moukir, Miwako Tsuji, Nahid Emad, Mitsuhisa Sato, Stéphane Baudelocq |
ICAART (1) | 4 |
| 2024 | Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer FugakuabstractBreadth-first search (BFS) is a fundamental building block of various high-performance computing applications beyond graph analysis and also known as a benchmark problem in the Graph 500 list. The increasing volume of global data demands efficient distributed BFS, which, however, is hindered by the high communication costs of exchanging vertex data between compute nodes. To address this challenge, this paper introduces four techniques: (i) forest pruning, which reduces the number of vertices by eliminating those unnecessary for the search; (ii) group reordering and (iii) multilevel bitmap compression, which decrease the memory footprint of graph data, thereby enabling fewer nodes to manage larger graphs; and (iv) adaptive parameter tuning, which quickly optimizes the hyperparameters of the BFS algorithm. In the evaluation using 152,064 nodes of the supercomputer Fugaku, our implementation achieved 198 tera-traversed edges per second, doubling the performance reported in the latest Graph500-related study on Fugaku. Junya Arai, Masahiro Nakao, Yuto Inoue, Kanto Teranishi, Koji Ueno, Keiichiro Yamamura, Mitsuhisa Sato, Katsuki Fujisawa |
SC | 7 |
| 2024 | Quantum-centric supercomputing for materials science: A perspective on challenges and future directions
Yuri Alexeev, Maximilian Amsler, Marco Antonio Barroca, Sanzio Bassini, Torey Battelle, Daan Camps, David Casanova, Young Jay Choi, Fred Chong, Charles Chung, Christopher Codella, Antonio D. Córcoles, James Cruise, Alberto Di Meglio, Ivan Duran, Thomas Eckl, Sophia E. Economou, Stephan J. Eidenbenz, Bruce Elmegreen, Clyde Fare, Ismael Faro, Cristina Sanz Fernández, Rodrigo Neumann Barros Ferreira, Keisuke Fuji, Bryce Fuller, Laura Gagliardi, Giulia Galli, Jennifer R. Glick, Isacco Gobbi, Pranav Gokhale, Salvador de la Puente Gonzalez, Johannes Greiner, William Gropp, Michele Grossi, Emanuel Gull, Burns Healy, Matthew R. Hermes, Benchen Huang, Travis S. Humble, Nobuyasu Ito, Artur F. Izmaylov, Ali Javadi-Abhari, Douglas M. Jennewein, Shantenu Jha, Bert de Jong, Petar Jurcevic, William M. Kirby, Stefan Kister, Masahiro Kitagawa, Joel Klassen, Katherine Klymko, Kwangwon Koh, Masaaki Kondo, Doga Murat Kürkçüoglu, Krzysztof Kurowski, Teodoro Laino, Ryan Landfield, Matthew L. Leininger, Vicente Leyton-Ortega, Ang Li 0006, Meifeng Lin, Junyu Liu, Nicolás Lorente, André Luckow, Simon Martiel, Francisco Martín-Fernández, Margaret Martonosi, Claire Marvinney, Arcesio Castañeda Medina, Dirk Merten, Antonio Mezzacapo, Kristel Michielsen, Abhishek Mitra, Tushar Mittal, Kyungsun Moon, Joel Moore, Sarah Mostame, Mario Motta, Young-Hye Na, Yunseong Nam, Prineha Narang, Yu-ya Ohnishi, Daniele Ottaviani, Matthew Otten, Scott Pakin, Vincent R. Pascuzzi, Edwin Pednault, Tomasz Piontek, Jed W. Pitera, Patrick Rall, Gokul Subramanian Ravi, Niall Robertson, Matteo A. C. Rossi, Piotr Rydlichowski, Hoon Ryu, Georgy Samsonidze, Mitsuhisa Sato, Nishant Saurabh, Kunal Sharma, Soyoung Shin, George Slessman, Mathias Steiner, Iskandar Sitdikov, In-Saeng Suh, Eric D. Switzer, Joel Thompson, Synge Todo, Minh C. Tran, Dimitar Trenev, Christian Trott, Huan-Hsin Tseng, Norm M. Tubman, Esin Tureci, David García Valiñas, Sofia Vallecorsa, Christopher Wever, Konrad W. Wojciechowski, Xiaodi Wu 0001, Shinjae Yoo, Nobuyuki Yoshioka, Victor Wen-zhe Yu, Seiji Yunoki, Sergiy Zhuk, Dmitry Zubarev |
Future Gener. Comput. Syst. | 99 |
| 2024 | Large-scale and cooperative graybox parallel optimization on the supercomputer Fugaku
Lorenzo Canonne, Bilel Derbel, Miwako Tsuji, Mitsuhisa Sato |
J. Parallel Distributed Comput. | 4 |
| 2024 | Design and performance evaluation of UCX for the Tofu Interconnect D on Fugaku towards efficient multithreaded communicationabstractAbstract The increasing trend of manycore processors makes multithreaded communication more important to avoid costly global synchronization among cores. One of the representative approaches that require multithreaded communication is the global task-based programming model. In the model, a program is divided into tasks, and tasks are asynchronously executed by each node, and independent thread-to-thread communications are expected. However, the Message passing interface (MPI) based approach is not efficient because of design issues. In this research, we design and implement the utofu transport layer in an abstracted communication library called Unified communication-X (UCX) for efficient remote direct memory access (RDMA) based multithreaded communication on Tofu Interconnect D. The evaluation results on Fugaku show that UCX can significantly improve the multithreaded performance over MPI, while maintaining portability between systems thanks to UCX. UCX shows about 32.8 times lower latency than Fujitsu MPI with 24 threads in the multithreaded pingpong benchmark and about 37.8 times higher update rate than Fujitsu MPI with 24 threads on 256 nodes in multithreaded GUPs benchmark. Yutaka Watanabe, Miwako Tsuji, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
J. Supercomput. | 5 |
| 2024 | Correction: Design and performance evaluation of UCX for the Tofu Interconnect D on Fugaku towards efficient multithreaded communication
Yutaka Watanabe, Miwako Tsuji, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
J. Supercomput. | 5 |
| 2022 | Performance analysis of a state vector quantum circuit simulation on A64FX processorabstractAlong with the recent development of quantum computers, quantum computer simulators are also exploited to verify and evaluate quantum computers. Since the amount of computation in the quantum circuit simulations increases exponentially with the number of qubits, which indicates the size of the quantum computer, it is important to study the performance of the quantum circuit simulations. In this paper, we analyze the performance of a state vector quantum circuit simulation on A64FX processor. We consider the implementation with the optimization control lines (OCLs), and that written with the Arm C Language Extensions (ACLE) to enhance vectorization, and compare them with the original implementation. Our experiments show that the vectorization improves the performance if the quantum gate simulations involve floating-point arithmetic operations. Miwako Tsuji, Mitsuhisa Sato |
CLUSTER | 2 |
| 2022 | Pushing the Frontier in the Design of Laser-Based Electron Accelerators with Groundbreaking Mesh-Refined Particle-In-Cell Simulations on Exascale-Class Supercomputersabstract(150 word max) We present a first-of-kind mesh-refined (MR) massively parallel Particle-In-Cell (PIC) code for kinetic plasma simulations optimized on the Frontier, Fugaku, Summit, and Perlmutter supercomputers. Major innovations, implemented in the WarpX PIC code, include: (i) a three level parallelization strategy that demonstrated performance portability and scaling on millions of A64FX cores and tens of thousands of AMD and Nvidia GPUs (ii) a groundbreaking mesh refinement capability that provides between 1.5 x to 4 x savings in computing requirements on the science case reported in this paper, (iii) an efficient load balancing strategy between multiple MR levels. The MR PIC code enabled 3D simulations of laser-matter interactions on Frontier, Fugaku, and Summit, which have so far been out of the reach of standard codes. These simulations helped remove a major limitation of compact laser-based electron accelerators, which are promising candidates for next generation high-energy physics experiments and ultra-high dose rate FLASH radiotherapy. Luca Fedeli, Axel Huebl, France Boillod-Cerneux, Thomas Clark, Kevin Gott, Conrad Hillairet, Stephan Jaure, Adrien Leblanc, Rémi Lehe, Andrew Myers 0001, Christelle Piechurski, Mitsuhisa Sato, Neïl Zaïm, Weiqun Zhang, Jean-Luc Vay, Henri Vincenti |
SC | 12 |
| 2021 | Sequences of Sparse Matrix-Vector Multiplication on Fugaku's A64FX processorsabstractWe implement parallel and distributed versions of the sparse matrix-vector product and the sequence of matrix-vector product operations, using OpenMP, MPI, and the ARM SVE intrinsic functions, for different matrix storage formats. We investigate the efficiency of these implementations on one and two A64FX processors, using a variety of sparse matrices as input. The matrices have different properties in size, sparsity and regularity. We observe that a parallel and distributed implementation shows good scaling on two nodes for cases where the matrix is close to a diagonal matrix, but the performances degrade quickly with variations to the sparsity or regularity of the input. Jérôme Gurhem, Maxence Vandromme, Miwako Tsuji, Serge G. Petiton, Mitsuhisa Sato |
CLUSTER | 5 |
| 2021 | Evaluation of SPEC CPU and SPEC OMP on the A64FXabstractWe evaluated the A64FX processor used in the supercomputer Fugaku using the SPEC CPU and SPEC OMP benchmark suites. As a result, we found the performance of the A64FX processor, which had 48 cores, was lower than that of the Xeon with dual sockets of 24 cores each in SPEC CPU int and fp. In SPEC OMP, due to the effect of the Xeon’s Hyperthread, the A64FX performance was lower than the performance of the Xeon with single socket of 28 cores. But in several benchmarks in SPEC CPU fp and SPEC OMP, the A64FX performance was higher due to its high memory bandwidth. In addition, by comparing the performance and power using the power control mechanism of the A64FX, it was confirmed that power can be reduced without affecting the performance when not using all cores. Yuetsu Kodama, Masaaki Kondo, Mitsuhisa Sato |
CLUSTER | 3 |
| 2021 | Performance Evaluation and Analysis of A64FX many-core Processor for the Fiber Miniapp SuiteabstractIn recent years, there has been growing interest in Arm-based processors for high performance computing systems such as supercomputer Fugaku using A64FX Arm-based processor. We have evaluated the performance of A64FX processor using Fiber Miniapp suite and have investigated various numbers of MPI processes, OpenMP threads as well as different methods to assign MPI processes and OpenMP threads. In addition to the performance evaluation, the performance comparison with other processors and some performance analysis are shown. Our experiments suggest that while shorter OpenMP thread strides perform better in most mini applications, MPI process allocation methods have not had a large impact on the performance. For some applications of “as-is” with small data set, A64FX shows poor performance, but it can be improved by enhancing the SIMD vectorization and changing instruction scheduling during the compilation. The performance of the A64FX is better or comparable with other processors for other applications and data sets. Miwako Tsuji, Mitsuhisa Sato |
CLUSTER | 2 |
| 2021 | Graph optimization algorithm for low-latency interconnection networksabstractIn various industrial products including parallel computing systems, it is expected that the performance of the whole system will be improved by applying a network topology with smaller diameter and average shortest path length (ASPL). Such network topologies can be defined as the order/degree problem in graph theory by modeling a network topology as an undirected graph. Although previous research indicates that the random graph has both small diameter and ASPL due to the small-world effect, there is still room for improvement. In this paper, we propose an algorithm for the order/degree problem that optimizes random graphs. The feature of the proposed algorithm is that by giving symmetry to the graph, the solution search performance is improved and the calculation time for obtaining the diameter and APSL is greatly reduced. The proposed algorithm is evaluated using various graphs, including a huge graph with one million vertices presented by Graph Golf, an international competition for the order/degree problem. The results show that the proposed algorithm can generate a network topology with sufficiently small diameter and APSL. In addition, we also simulate the latency, performance of the parallel benchmarks, and bisection on the generated network topologies. The results show that the generated network topology performs better than the random topology and a conventional k-ary n-cube topology. Masahiro Nakao, Maaki Sakai, Yoshiko Hanada, Hitoshi Murai, Mitsuhisa Sato |
Parallel Comput. | 5 |
| 2021 | Performance and power consumption analysis of Arm Scalable Vector Extension
Tetsuya Odajima, Yuetsu Kodama, Mitsuhisa Sato |
J. Supercomput. | 3 |
| 2021 | A new sustained system performance metric for scientific performance evaluationabstractAbstract Because of the increasing complexities of systems and applications, the performance of many traditional HPC benchmarks, such as HPL or HPCG, no longer correlates strongly with the actual performance of real applications. To address the discrepancy between simple benchmarks and real applications, and to better understand the application performance of systems, some metrics use a set of either real applications or mini applications. In particular, the Sustained System Performance (SSP) metric Kramer et al. (The NERSC sustained system performance (SSP) metric. Tech Rep LBNL-58868, 2005), which indicates the expected throughput of different applications executing with different datasets, is widely used. Whereas such a metric should lead to direct insights on the actual performance of real applications, sometimes more effort is necessary to port and evaluate complex applications. In this study, to obtain the approximate performance of SSP representing real applications, without running real applications, we propose a metric called the Simplified Sustained System Performance (SSSP) metric, which is computed based on several benchmark scores and their respective weighting factors, and we construct a method evaluating the SSSP metric of a system. The weighting factors are obtained by minimizing the gap between the SSP and SSSP scores based on a small set of reference systems. We evaluated the applicability of the SSSP method using eight systems and demonstrated that our proposed SSSP metrics produce appropriate performance projections of the SSP metrics of these systems, even when we adopted a simple method for computing the weighting factors. Additionally, the robustness of our SSSP metric was confirmed via computation of the weighting factors based on a smaller set of reference systems and computation of the SSSP metrics of other systems. Miwako Tsuji, William T. Kramer, Jean-Christophe Weill, Jean-Philippe Nomine, Mitsuhisa Sato |
J. Supercomput. | 5 |
| 2020 | Evaluation of Power Management Control on the Supercomputer FugakuabstractThe supercomputer “Fugaku”, which recently ranked number one on multiple supercomputing lists, including the Top500 in June 2020, has various power control features, such as (1) an eco mode that utilizes only one of two floating-point pipelines while decreasing the power supply to the chip; (2) a boost mode that increases clock frequency; and (3) a core retention function that turns unused cores into a low-power state. By orchestrating these power-performance features while considering the characteristics of currently running applications, we can potentially gain even better system-level energy efficiency. In this article, we report on the effectiveness of these features using the pre-evaluation environment for Fugaku. As a result, we confirmed several prominent results useful for the operation of the Fugaku system, including: remarkable power reduction and energy-efficiency improvement by coordinating the eco mode and the core retention feature in the memory intensive case; a 10% speed-up with a 17% power consumption increase using the boost mode in the CPU intensive case; and considerable power variations across over 20K nodes. Yuetsu Kodama, Tetsuya Odajima, Eishi Arima, Mitsuhisa Sato |
CLUSTER | 4 |
| 2020 | Performance Evaluation of Supercomputer Fugaku using Breadth-First Search Benchmark in Graph500abstractThere is increasing demand for the high-speed processing of large-scale graphs in various fields. However, such graph processing requires irregular calculations, making it difficult to scale performance on large-scale distributed memory systems. Against this background, Graph500, a competition for evaluating the performance of large-scale graph processing, has been held. We developed breadth-first search (BFS), which is one of the benchmark kernels used in Graph500, and took the top spot a total of 10 times using the K computer. In this paper, we tune BFS performance and evaluate it using the supercomputer Fugaku, which is the successor to the K computer. The results of evaluating BFS for a large-scale graph composed of about 1.1 trillion vertices and 17.6 trillion edges using 92,160 nodes of Fugaku indicate that Fugaku has 2.27 times the performance of the K computer. Fugaku took the top spot on Graph500 in June 2020. Masahiro Nakao, Koji Ueno, Katsuki Fujisawa, Yuetsu Kodama, Mitsuhisa Sato |
CLUSTER | 5 |
| 2020 | Preliminary Performance Evaluation of the Fujitsu A64FX Using HPC ApplicationsabstractRIKEN Center for Computational Science has been installing the supercomputer Fugaku. The Fujitsu A64FX, based on the Armv8.2-A+SVE architecture, is used in the system. In this paper, we evaluated the seven HPC applications and benchmarks on the A64FX. In a performance comparison with Marvell (Cavium) ThunderX2 processor and Intel Xeon Skylake processor, the A64FX achieved higher performance in a memory bandwidth-intensive application thanks to its high memory bandwidth. However, we confirmed that the performance of the A64FX decreased from a lack of out-of-order resources. To mitigate this problem, the “loop fission” function of the Fujitsu compiler was used to improve the performance. Tetsuya Odajima, Yuetsu Kodama, Miwako Tsuji, Motohiko Matsuda, Yutaka Maruyama, Mitsuhisa Sato |
CLUSTER | 6 |
| 2020 | Accuracy Improvement of Memory System Simulation for Modern Shared Memory ProcessorabstractFor the purpose of developing applications for supercomputer Fugaku at an early stage, RIKEN has developed a processor simulator. This simulator is based on the general-purpose processor simulator gem5. It does not simulate the actual hardware of a Fugaku processor. However, we believe that sufficient simulation accuracy can be obtained since it simulates the instruction pipeline of out-of-order execution with cycle-level accuracy along with performing detailed parameter tuning of out-of-order resources. In order to estimate the accurate execution time of a program, it is necessary to simulate with accuracy not only the instruction execution time, but also the access time of the cache memory hierarchy. Therefore, in the RIKEN simulator, we expanded gem5 to match the performance of the cache memory hierarchy to that of a Fugaku processor. In this simulator, we aim to estimate the execution cycles of one node application on a Fugaku processor with accuracy that enables relative evaluation and application tuning. In this paper, we show the details of the implementation of this simulator and verify its accuracy compared with that of a Fugaku processor test chip. In the evaluation of the total 46 kernel benchmarks, it was confirmed that the difference is 13% or less for 85% of the kernels. In the multithreaded execution of Stream Triad benchmark, scalable performance according to the number of threads was confirmed, and achieved over 80% of memory throughput with enough accuracy. Yuetsu Kodama, Tetsuya Odajima, Akira Asato, Mitsuhisa Sato |
HPC Asia | 4 |
| 2020 | Parallelization of All-Pairs-Shortest-Path Algorithms in Unweighted GraphabstractThe design of the network topology of a large-scale parallel computer system can be represented as an order/degree problem in graph theory. To solve the order/degree problem, it is necessary to obtain an all-pairs-shortest-path (APSP) of the graph. Thus, this paper evaluates two parallel algorithms that quickly find the APSP in unweighted graphs and compares their performance. The first APSP algorithm is based on the breadth-first search (BFS-APSP) and the second is based on the adjacency matrix (ADJ-APSP). First, we develop serial algorithms and threaded algorithms using OpenMP, and show that ADJ-APSP is up to 32.34 times faster than BFS-APSP. Next, we develop hybrid-parallel algorithms using OpenMP and MPI, and show that BFS-APSP is faster than ADJ-APSP under certain conditions because the maximum number of processes in BFS-APSP is greater than in ADJ-APSP. In addition, we parallelize ADJ-APSP using a single GPU (NVIDIA Tesla V100) and achieve a speed increase of up to 16.53-fold compared to that of a single CPU. Finally, we evaluate the performance of the algorithms using 128 GPUs and achieve a computation time 101.10 times faster than that using a single GPU. Moreover, it is shown that the calculation time of both algorithms can be greatly reduced when the input graphs are symmetric. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 3 |
| 2020 | The Supercomputer "Fugaku" and Arm-SVE enabled A64FX processor for energy-efficiency and sustained application performanceabstractWe have been carrying out the FLAGSHIP 2020 to develop the Japanese next-generation flagship supercomputer, Post-K, named as “Fugaku” recently. In the project, we have designed a new Arm-SVE enabled processor, called A64FX, as well as the system including interconnect with the industry partner, Fujitsu. The processor is designed for energy-efficiency and sustained application performance. In the design of the system, the “co-design” with the system and applications is a key to make it efficient and high-performance. We analyzed a set of the target applications provided from applications teams for the design of the processor architecture and the decision of many architectural parameters. The “Fugaku” is being installed and scheduled to be put into operation for public service around 2021. In this talk, several features and some preliminary performance results of the “Fugaku” system and A64FX manycore processor will be presented as well as the overview of the system. Mitsuhisa Sato |
ISPDC | 1 |
| 2020 | Co-design for A64FX manycore processor and "Fugaku"abstractWe have been carrying out the FLAGSHIP 2020 Project to develop the Japanese next-generation flagship supercomputer, the Post-K, recently named “Fugaku”. We have designed an original many core processor based on Armv8 instruction sets with the Scalable Vector Extension (SVE), an A64FX processor, as well as a system including interconnect and a storage subsystem with the industry partner, Fujitsu. The “co-design” of the system and applications is a key to making it power efficient and high performance. We determined many architectural parameters by reflecting an analysis of a set of target applications provided by applications teams. In this paper, we present the pragmatic practice of our co-design effort for “Fugaku”. As a result, the system has been proven to be a very power-efficient system, and it is confirmed that the performance of some target applications using the whole system is more than 100 times the performance of the K computer. Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama, Tetsuya Odajima, Miwako Tsuji, Hisashi Yashiro, Masaki Aoki, Naoyuki Shida, Ikuo Miyoshi, Kouichi Hirai, Atsushi Furuya, Akira Asato, Kuniki Morita, Toshiyuki Shimizu |
SC | 1 |
| 2020 | Design and evaluation of efficient global data movement in partitioned global address space
Hitoshi Murai, Mitsuhisa Sato |
Parallel Comput. | 2 |
| 2020 | InKS: a programming model to decouple algorithm from optimization in HPC codes
Ksander Ejjaaouani, Olivier Aumage, Julien Bigot, Michel Mehrenberger, Hitoshi Murai, Masahiro Nakao, Mitsuhisa Sato |
J. Supercomput. | 7 |
| 2019 | Distributed and Parallel Programming Paradigms on the K computer and a ClusterabstractIn this paper, we focus on a distributed and parallel programming paradigm for massively multicore supercomputers. We introduce YML, a development and execution environment for parallel and distributed applications based on a graph of task components scheduled at runtime and optimized for several middlewares. Then we show why YML may be well adapted to applications running on a lot of cores. The tasks are developed with the PGAS language XMP based on directives. We use YML/XMP to implement the block-wise Gaussian elimination to solve linear systems. We also implemented it with XMP and MPI without blocks. ScaLAPACK was also used to created an non-block implementation of the resolution of a dense linear system through LU factorization. Furthermore, we run it with different amount of blocks and number of processes per task. We find out that a good compromise between the number of blocks and the number of processes per task gives interesting results. YML/XMP obtains results faster than XMP on the K computer and close to XMP, MPI and ScaLAPACK on clusters of CPUs. We conclude that parallel and distributed multilevel programming paradigms like YML/XMP may be interesting solutions for extreme scale computing. Jérôme Gurhem, Miwako Tsuji, Serge G. Petiton, Mitsuhisa Sato |
HPC Asia | 4 |
| 2019 | Multi-accelerator extension in OpenMP based on PGAS modelabstractMany systems used in HPC field have multiple accelerators on a single compute node. However, programming for multiple accelerators is more difficult than that for a single accelerator. Therefore, in this paper, we propose an OpenMP extension that allows easy programming for multiple accelerators. We extend existing OpenMP syntax to create Partitioned Global Address Space (PGAS) on separated memories of several accelerators. The feature enables users to perform programming to use multiple accelerators in ease. In performance evaluation, we implement the STREAM Triad and the HIMENO benchmarks using the proposed OpenMP extension. As a result of evaluating the performance on a compute node equipped with up to four GPUs, we confirm that the proposed OpenMP extension demonstrates sufficient performance. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 3 |
| 2019 | A Method for Order/Degree Problem Based on Graph Symmetry and Simulated Annealing with MPI/OpenMP ParallelizationabstractThe network topology in various systems, such as large-scale data centers, high-performance computing systems, and Network on Chip, is strongly related to network latency. Designing a network topology with low latency can be defined as an order/degree problem (ODP) in graph theory by modeling the network topology as an undirected graph. This study proposes a method for efficiently solving ODPs based on graph symmetry and simulated annealing (SA). This method makes the network topology symmetrical, thereby improving the solution search performance of SA and drastically reducing the calculation time. The proposed method is applied to several problems from an international competition for ODPs called Graph Golf to find network topologies with sufficiently low latency. The symmetry-based calculation achieves a speed up of 31.76 times for one of the problems. Furthermore, to reduce calculation time, the proposed method is extended to use hybrid parallelization with MPI and OpenMP. As a result, a maximum speed up of 209.80 times was achieved on 20 compute nodes consisting of 400 CPU cores. Even faster performance was achieved by combining the symmetry-based calculation and hybrid parallelization. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 3 |
| 2019 | Scalable communication performance prediction using auto-generated pseudo MPI event traceabstractFor the co-design of HPC systems and applications, it is important to study how application performance is affected by the characteristics of the future systems, not just on a computation node but also for the parallel processing including inter-node communications. Trace-driven network simulators have been widely used because of its simplicity. However, they require the trace files corresponding to the simulated system size. Therefore, if a future system is larger than a current system, we can not adopt the trace files directly; that is, it is difficult to simulate a system larger than the current system. In order to address the scaling problem in the trace-driven network simulation, we have proposed a method called SCAlable Mpi Profiler (SCAMP). The SCAMP method runs an application on a current system, obtains MPI-event trace files, copies and edits the real trace files to create a large amount of pseudo MPI-event trace files for a future system, and finally drives a network simulator by inputting the pseudo MPI-event trace files. We also implemented a pseudo MPI-event trace file generator based on the analysis of LLVM's intermediate representations. We aim to easily obtain a first-order approximation of the communication performances for various network configurations and applications. In this paper, we describe the SCAMP system design and implementation as well as several performance evaluation results. Miwako Tsuji, Taisuke Boku, Mitsuhisa Sato |
HPC Asia | 3 |
| 2018 | A Source-to-Source Translation of Coarray Fortran with MPI for High PerformanceabstractCoarray Fortran (CAF) is a partitioned global address space (PGAS) language that is a part of standard Fortran 2008. We have implemented it as a source-to-source translator as a part of the Omni XcalebleMP compiler. Since the output is written in Fortran standard, the translator must utilize Fortran conventions such as the assumed-shape array and generic function in order to reduce both development costs and runtime overhead. The runtime library uses either GASNet, MPI-3, or Fujitsu's low-level Remote Direct Memory Access (RDMA) interface (FJ-RDMA) for one-sided communication. Hidetoshi Iwashita, Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 4 |
| 2018 | Multi-tasking Execution in PGAS Language XcalableMP and Communication Optimization on Many-core ClustersabstractLarge-scale clusters based on many-core processors such as Intel Xeon Phi have recently been deployed. Multi-tasking execution using task dependencies in OpenMP 4.0 is a promising candidate for facilitating the parallelization of such many-core processors, because this enables users to avoid global synchronization through fine-grained task-to-task synchronization using user-specified data dependencies. Recently, the partitioned global address space (PGAS) model has emerged as a usable distributed-memory programming model. In this paper, we propose a multi-tasking execution model in the PGAS language XcalableMP (XMP) for many-core clusters. The model provides a method to describe interactions between tasks based on point-to-point communications on the global address space. A communication is executed non-collectively among nodes. We implemented the proposed execution model in XMP, and designed a simple code transformation algorithm to MPI and OpenMP. We implemented two benchmarks using our model for preliminary evaluation, namely blocked Cholesky factorization and the Laplace equation solver. Most of the implementations using our model outperform the conventional barrier-based data-parallel model. To improve the performance in many-core clusters, we propose a communication optimization method by dedicating a single thread for communications, to avoid performance problems related to the current multi-threaded MPI execution. As a result, the performances of blocked Cholesky factorization and the Laplace equation solver using this communication optimization are improved to 138% and 119% compared with the barrier-based implementation in Intel Xeon Phi KNL clusters, respectively. From the viewpoint of productivity, the program implemented by our model in XMP is almost the same as the implementation based on the OpenMP task depend clause, because XMP enables the parallelization of the serial source code with additional directives and small changes as well as OpenMP. Keisuke Tsugane, Jinpil Lee, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 4 |
| 2017 | Implementation and Evaluation of One-sided PGAS Communication in XcalableACC for Accelerated ClustersabstractClusters equipped with accelerators such as graphics processing unit (GPU) and Many Integrated Core (MIC) are widely used. For such clusters, programmers write programs for their applications by combining MPI with one of the available accelerator programming models. In particular, OpenACC enables programmers to develop their applications easily, but with lower productivity owing to complex MPI programming. XcalableACC (XACC) is a new programming model, which is an "orthogonal" integration of a partitioned global address space (PGAS) language XcalableMP (XMP) and OpenACC. While XMP enables distributed-memory programming on both global-view and local-view models, OpenACC allows operations to be offloaded to a set of accelerators. In the local-view model, programmers can describe communication with the coarray features adopted from Fortran 2008, and we extend them to communication between accelerators. We have designed and implemented an XACC compiler for NVIDIA GPU and evaluated its performance and productivity by using two benchmarks, Himeno benchmark and NAS Parallel Benchmarks CG (NPB-CG). The performance of the XACC version with the Himeno benchmark and NPB-CG are over 85% and 97% in the local-view model against the MPI+OpenACC version, respectively. Moreover, using non-blocking communication makes the performance of local-view version over 89% with the Himeno benchmark. From the viewpoint of productivity, the local-view model provides an intuitive form of array assignment statement for communication. Akihiro Tabuchi, Masahiro Nakao, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
CCGrid | 5 |
| 2017 | Preliminary Performance Evaluation of Application Kernels Using ARM SVE with Multiple Vector LengthsabstractModern high performance processors are equipped with very wide SIMD instruction set. SVE (Scalable Vector Extension) is an ARM® SIMD technology that supports vector lengths from 128 bits to 2048 bits. One of its promising features is to offer "vector-length agnostic" programming to allow the same SVE code to run on hardware of any vector length without any modification of the code. This feature would be useful to explore the best vector length with appropriate hardware resources in the space of various combinations of hardware parameters in order to make more efficient use of hardware resources, since we can use the same vectorized SIMDcode. In this paper, we report the performance of application kernelsusing ARM SVE with multiple vector lengths while keeping the hardware resource the same. We have confirmed that when the performance of the program is limited by a bottleneck of a long chain of arithmetic operations or instruction issues, the performance can be improved by increasing the vector length. However, it was necessary to prepare a sufficient number of physical registers for performance improvement, and when the number of physical registers was too small, it was found that with such a program, the performance might be reduced. When the performance is limited by memory access bandwidth to cache and memory, the vector length does not affect the performance significantly. Yuetsu Kodama, Tetsuya Odajima, Motohiko Matsuda, Miwako Tsuji, Jinpil Lee, Mitsuhisa Sato |
CLUSTER | 6 |
| 2017 | Implementing Lattice QCD Application with XcalableACC Language on Accelerated ClusterabstractAccelerated clusters, which are distributed memory systems equipped with accelerators, have been used in various fields. For accelerated clusters, programmers often implement their applications by a combination of MPI and CUDA (MPI+CUDA). However, the approach faces programming complexity issues. This paper introduces the XcalableACC (XACC) language, which is a hybrid model of XcalableMP (XMP) and OpenACC. While XMP is a directive-based language for distributed memory systems, OpenACC is also a directive-based language for accelerators. XACC enables programmers to develop applications on accelerated clusters with ease. To evaluate XACC performance and productivity levels, we implemented a lattice quantum chromodynamics (Lattice QCD) application using XACC on 64 compute nodes and 256 GPUs and found its performance was almost the same as that of MPI+CUDA. Moreover, we found that XACC requires much less change from the serial Lattice QCD code than MPI+CUDA to implement the parallel Lattice QCD code. Masahiro Nakao, Hitoshi Murai, Hidetoshi Iwashita, Akihiro Tabuchi, Taisuke Boku, Mitsuhisa Sato |
CLUSTER | 6 |
| 2017 | A Performance Projection of Mini-Applications onto Benchmarks Toward the Performance Projection of Real-ApplicationsabstractWidely used benchmarks, such as High Performance Linpack (HPL), do not always provide direct insights are notoriously poor indicators of into the actual application performance of systems. When real applications are used, and there have been are criticisms indicating that the performance of simplified benchmarks such as HPL no longer strongly correlate to real application performance. In contrast, performance evaluations based on real or mini applications may give a direct estimation into application performance. The Sustained System Performance (SSP) metric, which is used to evaluate systems based on the performance at scale of various applications, has been successfully adopted to procure systems at the National Energy Research Scientific Computing Center (NERSC), the National Center for Supercomputing Applications (NCSA) and other facilities. However, significant effort is required to tune and optimize several mini applications for each of systems. In this paper, we propose a new performance metric - the Simplified Sustained System Performance (SSSP) metric - based on a suite of simple benchmarks, which enables performance projection that correlates with full applications, but use onto a suite of mini applications. While the SSP metric is calculated over a set of applications, the SSSP metric applies its methodology to a set of benchmarks. Preliminary weighting factors for benchmarks are introduced to approximate the original SSP metric more accurately by the SSSP metric. To define the weighting factors, we perform a simple learning algorithm. Our preliminary experiments show that even though our metric is still easy to measure because it is based on a combination of simple benchmarks, it can provide projections of the performance of applications. Miwako Tsuji, William T. Kramer, Mitsuhisa Sato |
CLUSTER | 3 |
| 2016 | Hybrid-view programming of nuclear fusion simulation code in the PGAS parallel programming language XcalableMP
Keisuke Tsugane, Taisuke Boku, Hitoshi Murai, Mitsuhisa Sato, William Tang 0002, Bei Wang 0002 |
Parallel Comput. | 4 |
| 2015 | Hybrid Communication with TCA and InfiniBand on a Parallel Programming Language XcalableACC for GPU ClustersabstractFor the execution of parallel HPC applications on GPU-ready clusters, high communication latency between GPUs over nodes will be a serious problem on strong scalability. To reduce the communication latency between GPUs, we proposed the Tightly Coupled Accelerator (TCA) architecture and developed the PEACH2 board as a proof-of-concept interconnection system for TCA. Although PEACH2 provides very low communication latency, there are some hardware limitations due to its implementation depending on PCIe technology, such as the practical number of nodes in a system which is 16 currently named sub-cluster. More number of nodes should be connected by conventional interconnections such as InfiniBand, and the entire network system is configured as a hybrid one with global conventional network and local high-speed network by PEACH2. For ease of user programmability, it is desirable to operate such a complicated communication system at the library or language level (which hides the system). In this paper, we develop a hybrid interconnection network system combining PEACH2 and InfiniBand, and implement it based on a high-level PGAS language for accelerated clusters named XcalableACC (XACC). A preliminary performance evaluation confirms that the hybrid network improves the performance based on the Himeno benchmark for stencil computation by up to 40%, relative to MVAPICH2 with GDR on InfiniBand. Additionally, Allgather collective communication with a hybrid network improves the performance by up to 50% for networks of 8 to 16 nodes. The combination of local communication, supported by the low latency of PEACH2 and global communication supported by the high bandwidth and scalability of InfiniBand, results in an improvement of overall performance. Tetsuya Odajima, Taisuke Boku, Toshihiro Hanawa, Hitoshi Murai, Masahiro Nakao, Akihiro Tabuchi, Mitsuhisa Sato |
CLUSTER | 7 |
| 2014 | A PGAS Execution Model for Efficient Stencil Computation on Many-Core ProcessorsabstractA efficient PGAS execution model on many-core processor for stencil computation is proposed and implemented. We use XcalableMP as a base language and we modify its runtime well fit in many-core processors. The runtime uses processes for parallel execution and global arrays of the stencil codes are broken into blocked sub-arrays placed on shared memory. Using two stencil codes, Laplace and Himeno, we evaluated its performance. In the evaluation, we show (1) Blocking improves locality of memory access during computation therefore improves total CPU execution time. (2) Direct data access using shared memory can relieve communication burden of sub-array halo exchanges. Mitsuru Ikei, Mitsuhisa Sato |
CCGRID | 2 |
| 2014 | Grid-Oriented Process Clustering System for Partial Message LoggingabstractIn a computer cluster composed of many nodes, the mean time between failures becomes shorter as the number of nodes increases. This may mean that lengthy tasks cannot be performed, because they will be interrupted by failure. Therefore, fault tolerance has become an essential part of high-performance computing. Partial message logging forms clusters of processes, and coordinates a series of checkpoints to log messages between groups. Our study proposes a system of two features to improve the efficiency of partial message logging: 1) the communication log used in the clustering is recorded at runtime, and 2) a graph partitioning algorithm reduces the complexity of the system by geometrically partitioning a grid graph. The proposed system is evaluated by executing a scientific application. The results of process clustering are compared to existing methods in terms of the clustering performance and quality. Hideyuki Jitsumoto, Yuki Todoroki, Yutaka Ishikawa, Mitsuhisa Sato |
DSN | 4 |
| 2014 | Hybrid-view programming of nuclear fusion simulation code in the PGAS parallel programming language XcalableMPabstractRecently, the Partitioned Global Address Space (PGAS) parallel programming model has emerged as a usable distributed memory programming model. XcalableMP (XMP) is a PGAS parallel programming language that extends base languages such as C and Fortran with directives in OpenMP-like style. XMP supports a global-view model that allows programmers to define global data and to map them to a set of processors, which execute the distributed global data as a single thread. In XMP, the concept of a coarray is also employed for local-view programming. In this study, we port Gyrokinetic Toroidal Code - Princeton (GTC-P), which is a three-dimensional PIC code developed at Princeton University to study the microturbulence phenomenon in magnetically confined fusion plasmas, to XMP as an example of hybrid memory model coding with the global-view and local-view programming models. In local-view programming, the coarray notation is simple and intuitive compared with Message Passing Interface (MPI) programming while the performance is comparable to that of the MPI version. Thus, because the global-view programming model is suitable for expressing the data parallelism for a field of grid space data, we implement a hybrid-view version using a global-view programming model to compute the field and a local-view programming model to compute the movement of particles. The performance is degraded by 5-25% compared with the original MPI version, but the hybrid-view version facilitates more natural data expression for static grid space data (in global-view model) and dynamic particle data (in local-view model), and it also increases the readability of the code for higher productivity. Keisuke Tsugane, Hideo Nuga, Taisuke Boku, Hitoshi Murai, Mitsuhisa Sato, William Tang 0002, Bei Wang 0002 |
ICPADS | 5 |
| 2014 | Victim Selection and Distributed Work Stealing Performance: A Case StudyabstractWork stealing is a popular solution to perform dynamic load balancing of irregular computations, both for shared memory and distributed memory systems. While shared memory performance of work stealing is well understood, distributing this algorithm to several thousands of nodes can introduce new performance issues. In particular, most studies of work stealing assume that all participating processes are equidistant from each other, in terms of communication latency. This paper presents a new performance evaluation of the popular UTS benchmark, in its work stealing implementation, on the scale of ten thousands of compute nodes. Taking advantage of the physical scale of the K Computer, we investigate in details the performance impact of communication latencies on work stealing. In particular, we introduce a new performance metric to assess the time needed by the work stealing scheduler to distribute work among all processes. Using this metric, we identify a previously overlooked issue: the victim selection function used by the work stealing application can severely impact its performance at large scale. To solve this issue, we introduce a new strategy taking into account the physical distance between nodes and achieve significant performance improvements. Swann Perarnau, Mitsuhisa Sato |
IPDPS | 2 |
| 2013 | Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage |
ICA3PP (2) | 3 |
| 2013 | Multiple-SPMD Programming Environment Based on PGAS and Workflow toward Post-petascale ComputingabstractIn this paper, we propose a new development and execution environment based on workflow and PGAS methodologies for parallel programmings in post-petascale systems. It is expected that post-petascale systems will have a huge and highly hierarchical architecture with nodes of many-core processors and accelerators. For current parallel programs, MPI, MPI/OpenMP hybrid, and so on, it would be sometimes difficult to exploit the post-petascale systems efficiently. The proposed environment, called FP2C (Framework for Post-Petascale Computing), supports multi-program methodologies across multi-architectural levels. It introduces a PGAS parallel programming language called XcalableMP (XMP) to describe tasks into a workflow environment called YML. FP2C is composed of three layers: (1) workflow programming, (2)distributed programming, and (3) shared-memory parallel programming/accelerator. Computational experiments suggest that effective use of cores and memories can be achieved by controlling the level of hierarchization using FP2C. Miwako Tsuji, Mitsuhisa Sato, Maxime R. Hugues, Serge G. Petiton |
ICPP | 2 |
| 2013 | A communication library between multiple sets of MPI processes for a MPMD modelabstractBecause current high-end parallel systems have more than several thousands of nodes, a new programming model is required to exploit different levels of parallelism. For example, a conventional master-worker program assumes that a worker is running as a single MPI process. In some cases, a worker may run as a set of MPI processes to exploit a different level of parallelism or solve large problems within each worker. As the number of nodes increases, communication may be needed between the multiple sets of MPI processes. On the other hand, in order to make use of a large-scale system more effectively, it has been recently common for researchers to couple parallel components to conduct simulations of multiple physics models. This method may also requires a way to integrate multiple MPI programs. Takenori Shimosaka, Hitoshi Murai, Mitsuhisa Sato |
EuroMPI | 3 |
| 2012 | Productivity and Performance of Global-View Programming with XcalableMP PGAS LanguageabstractXcalableMP (XMP) is a PGAS parallel language with a directive-based extension of C and Fortran. While it sup- ports “coarray” as a local-view programming model, an XMP global-view programming model is useful when parallelizing data-parallel programs by adding directives with minimum code modification. This paper considers the productivity and performance of the XMP global-view programming model. In the global-view programming model, a programmer describes data distributions and work-mapping to map the computations to nodes, where the computed data are located. Global-view communication directives are used to move a part of the distributed data globally and to maintain consistency in the shadow area. Rich sets of XMP global-view programming model can reduce the cost for parallelization significantly, and optimization of “privatization” is not necessary. For productivity and performance study, the Omni XMP compiler and the Berkeley Unified Parallel C compiler are used. Experimental results show that XMP can implement the benchmarks with a smaller programming cost than UPC. Furthermore, XMP has higher access performance for global data, which has an affinity with own process than UPC. In addition, the XMP coarray function can effectively tune the application's performance. Masahiro Nakao, Jinpil Lee, Taisuke Boku, Mitsuhisa Sato |
CCGRID | 4 |
| 2012 | An asynchronous parallel genetic algorithm for the maximum likelihood phylogenetic tree searchabstractA phylogenetic tree represents the evolutionary relationships among biological species. Although parallel computation is essential for the phylogenetic tree searches, it is not easy to maintain the diversity of population in a parallel genetic algorithm. In this paper, we design a new asynchronous parallel genetic algorithm for tree optimization which maintain the diversity of population without any communication or synchronization. Miwako Tsuji, Mitsuhisa Sato, Akifumi S. Tanabe, Yuji Inagaki, Tetsuo Hashimoto |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | DS-Bench Toolset: Tools for dependability benchmarking with simulation and assuranceabstractToday's information systems have become large and complex because they must interact with each other via networks. This makes testing and assuring the dependability of systems much more difficult than ever before. DS-Bench Toolset has been developed to address this issue, and it includes D-Case Editor, DS-Bench, and D-Cloud. D-Case Editor is an assurance case editor. It makes a tool chain with DS-Bench and D-Cloud, and exploits the test results as evidences of the dependability of the system. DS-Bench manages dependability benchmarking tools and anomaly loads according to benchmarking scenarios. D-Cloud is a test environment for performing rapid system tests controlled by DS-Bench. It combines both a cluster of real machines for performance-accurate benchmarks and a cloud computing environment as a group of virtual machines for exhaustive function testing with a fault-injection facility. DS-Bench Toolset enables us to test systems satisfactorily and to explain the dependability of the systems to the stakeholders. Hajime Fujita 0002, Yutaka Matsuno, Toshihiro Hanawa, Mitsuhisa Sato, Shinpei Kato, Yutaka Ishikawa |
DSN | 4 |
| 2012 | Audit: A new synchronization API for the GET/PUT protocol
Atsushi Hori, Jinpil Lee, Mitsuhisa Sato |
J. Parallel Distributed Comput. | 3 |
| 2012 | Preface
Chao-Tung Yang, Kuan-Chou Lai, Mitsuhisa Sato, Tzung-Shi Chen |
J. Supercomput. | 3 |
| 2011 | XMCAPI: Inter-core Communication Interface on Multi-chip Embedded SystemsabstractMulti-core processor technology has been applied to the processors in embedded systems as well as in ordinary PC systems. In multi-core embedded processors, however, a processor may consist of heterogeneous CPU cores that are not configured with a shared memory and do not have a communication mechanism for inter-core communication. MCAPI is a highly portable API standard for providing inter-core communication independent of the architecture heterogeneity. In this paper, we extend the current MCAPI to a multi-chip in a distributed memory configuration and propose its portable implementation, named XMCAPI, on a commodity network stack. With XMCAPI, the inter-core communication method for intra-chip cores is extended to inter-chip cores. We evaluate the XMCAPI implementation, xmcapi/ip, on a standard socket in a portable software development environment. Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato |
EUC | 4 |
| 2011 | Introduction
Mitsuhisa Sato, Denis Barthou, Pedro C. Diniz, P. Saddayapan |
Euro-Par (1) | 1 |
| 2011 | A distributed architecture of Sensing Web for sharing open sensor nodes
Ryo Kanbayashi, Mitsuhisa Sato |
Future Gener. Comput. Syst. | 2 |
| 2010 | D-Cloud: Design of a Software Testing Environment for Reliable Distributed Systems Using Cloud Computing TechnologyabstractIn this paper, we propose a software testing environment, called D-Cloud, using cloud computing technology and virtual machines with fault injection facility. Nevertheless, the importance of high dependability in a software system has recently increased, and exhaustive testing of software systems is becoming expensive and time-consuming, and, in many cases, sufficient software testing is not possible. In particular, it is often difficult to test parallel and distributed systems in the real world after deployment, although reliable systems, such as high-availability servers, are parallel and distributed systems. D-Cloud is a cloud system which manages virtual machines with fault injection facility. D-Cloud sets up a test environment on the cloud resources using a given system configuration file and executes several tests automatically according to a given scenario. In this scenario, D-Cloud enables fault tolerance testing by causing device faults by virtual machine. We have designed the D-Cloud system using Eucalyptus software and a description language for system configuration and the scenario of fault injection written in XML. We found that the D-Cloud system, which allows a user to easily set up and test a distributed system on the cloud and effectively reduces the cost and time of testing. Takayuki Banzai, Hitoshi Koizumi, Ryo Kanbayashi, Takayuki Imada, Toshihiro Hanawa, Mitsuhisa Sato |
CCGRID | 6 |
| 2010 | Runtime Energy Adaptation with Low-Impact Instrumented Code in a Power-Scalable Cluster SystemabstractRecently, improving the energy efficiency of high performance PC clusters has become important. In order to reduce the energy consumption of the microprocessor, many high performance microprocessors have a Dynamic Voltage and Frequency Scaling (DVFS) mechanism. This paper proposes a new DVFS method called the Code-Instrumented Runtime (CI-Runtime) DVFS method, in which a combination of voltage and frequency, which is called a P-State, is managed in the instrumented code at runtime. The proposed CI-Runtime DVFS method achieves better energy saving than the Interrupt based Runtime DVFS method, since it selects the appropriate P-State in each defined region based on the characteristics of program execution. Moreover, the proposed CI-Runtime DVFS method is more useful than the Static DVFS method, since it does not acquire exhaustive profiles for each P-State. The method consists of two parts. In the first part of the proposed CI-Runtime DVFS method, the instrumented codes are inserted by defining regions that have almost the same characteristics. The instrumented code must be inserted at the appropriate point, because the performance of the application decreases greatly if the instrumented code is called too many times in a short period. A method for automatically defining regions is proposed in this paper. The second part of the proposed method is the energy adaptation algorithm which is used at runtime. Two types of DVFS control algorithms energy adaptation with estimated energy consumption and energy adaptation with only performance information, are compared. The proposed CI-Runtime DVFS method was implemented on a power-scalable PC cluster. The results show that the proposed CI-Runtime with energy adaptation using estimated energy consumption could achieve an energy saving of 14.2% which is close to the optimal value, without obtaining exhaustive profiles for every available P-State setting. Hideaki Kimura 0003, Takayuki Imada, Mitsuhisa Sato |
CCGRID | 3 |
| 2010 | Customizing Virtual Machine with Fault Injector by Integrating with SpecC Device Model for a Software Testing Environment D-CloudabstractD-Cloud is a software testing environment for dependable parallel and distributed systems using cloud computing technology. We use Eucalyptus as cloud management software to manage virtual machines designed based on QEMU, called FaultVM, which have a fault injection mechanism. D-Cloud enables the test procedures to be automated using a large amount of computing resources in the cloud by interpreting the system configuration and the test scenario written in XML in D-Cloud front end and enables tests including hardware faults by emulating hardware faults by FaultVM flexibly. In the present paper, we describe the customization facility of FaultVM used to add new device models. We use SpecC, which is a system description language, to describe the behavior of devices, and a simulator generated from the description by SpecC is linked and integrated into FaultVM. This also makes the definition and injection of faults flexible without the modification of the original QEMU source codes. This facility allows D-Cloud to be used to test distributed systems with customized devices. Toshihiro Hanawa, Hitoshi Koizumi, Takayuki Banzai, Mitsuhisa Sato, Shin'ichi Miura, Tadatoshi Ishii, Hidehisa Takamizawa |
PRDC | 4 |
| 2009 | Using a cluster as a memory resource: A fast and large virtual memory on MPIabstractThe 64-bit OS provides ample memory address space that is beneficial for applications using a large amount of data. This paper proposes using a cluster as a memory resource for sequential applications requiring a large amount of memory. This system is an extension of our previously proposed socket-based distributed large memory system (DLM), which offers large virtual memory by using remote memory distributed over nodes in a cluster. The newly designed DLM is based on MPI (message passing interface) to exploit higher portability. MPI-based DLM provides fast and large virtual memory on widely available open clusters managed with an MPI batch queuing system. To access this remote memory, we rely on swap protocols adequate for MPI thread support levels. In experiments, we confirmed that it achieves 493 MB/s and 613 MB/s of remote memory bandwidth with the STREAM benchmark on 2.5 GB/s and 5 GB/s links (Myri-10G x2, x4) and high performance of applications with NPB and Himeno benchmarks. Additionally, this system enables users unfamiliar with parallel programming to use a cluster. Hiroko Midorikawa, Kazuhiro Saito, Mitsuhisa Sato, Taisuke Boku |
CLUSTER | 3 |
| 2009 | A Distributed Architecture of Sensing Web for Sharing Open Sensor Nodes
Ryo Kanbayashi, Mitsuhisa Sato |
GPC | 2 |
| 2009 | Reliable Software Distributed Shared Memory Using Page MigrationabstractReliability has recently become an important issue in PC cluster technology. This research proposes a software distributed shared memory system, named SCASH-FT, as an execution platform for high performance and highly reliable parallel system for commodity PC clusters. To achieve fault tolerance, each node has redundant page data that allows recovery from node failure using SCASH-FT. All page data is checkpointed and duplicated to another node when a user explicitly calls the checkpoint function. When failure occurs, SCASH-FT invokes the rollback function by restarting an execution from the last checkpoint data. SCASH-FT takes charge of processes such as detecting failure and restarting execution. So, all you have to do is just adding checkpoint function calls in the source code to determine the timing of each checkpoint. Evaluation results show that the checkpoint cost and the rollback penalty depend on the data access pattern and the checkpoint frequency. Thus, users can control their application performance by adjusting checkpoint frequency. Jinpil Lee, Mitsuhisa Sato |
ICPADS | 2 |
| 2009 | Flexible Multi-link Ethernet Binding System for PC Clusters with Asymmetric TopologyabstractIn current high-performance PC clusters, the performance and cost of interconnection network are essential issues. Very cost-effective Ethernets, such as Gigabit Ethernet, as well as high performance SANs, such as Infiniband and Myrinet, are still widely used. The authors have been developing a multi-link binding network system for Ethernet, called RI2N, for high-throughput and fault-tolerant interconnection with Gigabit Ethernet. It can be used both for internode communication in MPI programs and traditional UNIX network services such as NFS. In this paper, the authors propose an optimized version of RI2N, called RI2N+, that allows asymmetrical multi-link connection for fitting to various cost-effective system configurations. Such a configuration cannot be supported by Linux Channel Bonding, which is widely used in standard Linux distributions. RI2N+ automatically detects the asymmetric network configuration and controls the traffic distribution to multiple links. In the basic performance evaluation under a high traffic rate, it was confirmed that the throughput of the network with the proposed scheme is improved by approximately 30\% compared with the original RI2N. RI2N+ also maintains high performance even in asymmetric configurations, that is up to 86\% of the relative performance compared with the symmetric case. Taiga Yonemoto, Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato |
ICPADS | 5 |
| 2009 | RI2N/DRV: Multi-link ethernet for high-bandwidth and fault-tolerant network on PC clustersabstractAlthough recent high-end interconnection network devices and switches provide a high performance to cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, Linux kernels are equipped with software named Linux Channel Bonding (LCB), which is based IEEE802.3ad Link Aggregation technology. However, standard LCB has the disadvantage of mismatch with the TCP protocol; consequently, both large latency and bandwidth instability can occur. Fault-tolerance feature is supported by LCB, but the usability is not sufficient. We developed a new implementation similar to LCB named Redundant Interconnection with Inexpensive Network with Driver (RI2N/DRV) for use on Gigabit Ethernet. RI2N/DRV has a complete software stack that is very suitable for TCP, an upper layer protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets, even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as network file system (NFS), without any modification of other modules. Shin'ichi Miura, Toshihiro Hanawa, Taiga Yonemoto, Taisuke Boku, Mitsuhisa Sato |
IPDPS | 5 |
| 2009 | Towards an Open Dependable Operating SystemabstractThis paper introduces a new dependable operating system project, called DEOS, started in 2006, and scheduled to continue for six years. In this project, a safety extension mechanism called P-Bus is to be designed, and implemented in the Linux kernel so that a future dependability attribute is implemented with P-Bus. A hardware abstraction layer, called SPUMONE, is introduced so that a light-weight operating system, called ArcOS, and a monitoring service on top of ArcOS monitors the Linux kernel to provide a safety-net for the Linux kernel. New dependability metrics are being designed to enable developers and users to decide which hardware or software solution meets their dependability requirements, and thus can be used. Yutaka Ishikawa, Hajime Fujita 0002, Toshiyuki Maeda, Motohiko Matsuda, Midori Sugaya, Mitsuhisa Sato, Toshihiro Hanawa, Shin'ichi Miura, Taisuke Boku, Yuki Kinebuchi, Tatsuo Nakajima, Jin Nakazawa, Hideyuki Tokuda |
ISORC | 6 |
| 2008 | Runtime DVFS control with instrumented Code in power-scalable cluster systemabstractRecently, several energy reduction techniques using DVFS have been presented for PC clusters. This work proposes a Code-instrumented Runtime DVFS control, in which the combination of frequency and voltage (called a gear) is managed at the instrumented code at runtime. The codes are inserted by defining the program regions that have the same characteristics. The Code-instrumented Runtime DVFS control method is better than the Interrupt-based Runtime DVFS control method, in which the gear is managed by periodic interrupt, because it can reflect the program information to control DVFS. Though Static DVFS control, which makes use of the power profile before execution, gives better energy reduction, the proposed Code-instrumented Runtime DVFS control is easier to use, because it requires no information such as profile. The proposed DVFS control method was designed and implemented. The beta-adaptation was used as the runtime algorithm to choose the appropriate gear. The results show that the proposed method can improve the performance and energy consumption compared with Interrupt-based Runtime DVFS control. Although our Code-instrumented Runtime DVFS control can select lower voltages and frequencies than the present Runtime DVFS control given a certain deadline, unfortunately, it was also found to increase power consumption of the PC cluster due to an increase in the execution time. Hideaki Kimura 0003, Mitsuhisa Sato, Takayuki Imada, Yoshihiko Hotta |
CLUSTER | 2 |
| 2008 | DLM: A distributed Large Memory System using remote memory swapping over cluster nodesabstractEmerging 64bitOS’s supply a huge amount of memory address space that is essential for new applications using very large data. It is expected that the memory in connected nodes can be used to store swapped pages efficiently, especially in a dedicated cluster which has a high-speed network such as 10GbE and Infiniband. In this paper, we propose the Distributed Large Memory System (DLM), which provides very large virtual memory by using remote memory distributed over the nodes in a cluster. The performance of DLM programs using remote memory is compared to ordinary programs using local memory. The results of STREAM, NPB and Himeno benchmarks show that the DLM achieves better performance than other remote paging schemes using a block swap device to access remote memory. In addition to performance, DLM offers the advantages of easy availability and high portability, because it is a user-level software without the need for special hardware. To obtain high performance, the DLM can tune its parameters independently from kernel swap parameters. We also found that DLM’s independence of kernel swapping provides more stable behavior. Hiroko Midorikawa, Motoyoshi Kurokawa, Ryutaro Himeno, Mitsuhisa Sato |
CLUSTER | 4 |
| 2008 | RI2N: High-bandwidth and fault-tolerant network with multi-link Ethernet for PC clustersabstractAlthough recent high-end interconnection network devices and switches provide a high performance/cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, a Linux kernel is equipped with a software solution named Linux Channel Bonding (LCB), which is based on IEEE802.3ad Link Aggregation technology. However, standard LCB has the problem of mismatching with the commonly used TCP protocol, which consequently implies several problems of both large latency and instability on bandwidth improvement. The fault-tolerant feature is also supported, but the usability is not sufficient. We have developed a new implementation similar to LCB named RI2N/DRV (Redundant Interconnection with Inexpensive Network with Driver) for use on a Gigabit Ethernet with a complete software stack that is very compatible with the TCP protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as NFS, without any modification of other modules. Shin'ichi Miura, Takayuki Okamoto, Taisuke Boku, Toshihiro Hanawa, Mitsuhisa Sato |
CLUSTER | 5 |
| 2008 | Performance Evaluation of Data Management Layer by Data Sharing Patterns for Grid RPC Applications
Yoshihiro Nakajima, Yoshiaki Aida, Mitsuhisa Sato, Osamu Tatebe |
Euro-Par | 3 |
| 2008 | Power management of distributed web savers by controlling server power state and traffic prediction for QoSabstractIn this paper, we propose a scheme based on server node state control, including stand-by/wake-up and processor power control, to achieve aggressive power education while satisfying Quality of Service (QoS). Decreasing power consumption on Web servers is currently a challenging new problem to be solved in a data center or warehouse. Although Web servers are configured to have maximum performance, the actual access rate to the servers can be small in a specific period, such as midnight, so it may be possible to reduce the power consumption of the servers while satisfying QoS with lower server performance. In order to reduce power consumption on the server nodes, we now have to consider the power consumption of the entire node rather than only processor power by Dynamic Voltage and Frequency Scaling (DVFS). We implemented the proposed scheme to the distributed Web server system using the power-profile of server nodes and considering load increment based on traffic prediction method and evaluated the proposed scheme with a Web server benchmark workload based on SPECWeb99. The result reveals that the proposed scheme achieved an energy saving of approximately 17% with sufficient QoS performance on the distributed Web server system. Takayuki Imada, Mitsuhisa Sato, Yoshihiko Hotta, Hideaki Kimura 0003 |
IPDPS | 2 |
| 2008 | A parallel method for large sparse generalized eigenvalue problems using a GridRPC system
Tetsuya Sakurai, Yoshihisa Kodaki, Hiroto Tadano, Daisuke Takahashi, Mitsuhisa Sato, Umpei Nagashima |
Future Gener. Comput. Syst. | 5 |
| 2008 | Integrating Computing Resources on Multiple Grid-Enabled Job Scheduling Systems Through a Grid RPC System
Yoshihiro Nakajima, Mitsuhisa Sato, Yoshiaki Aida, Taisuke Boku, Franck Cappello |
J. Grid Comput. | 2 |
| 2007 | Resolution of large symmetric eigenproblems on a world wide gridabstractWe propose an explicit restarted Lanczos algorithm on a world-wide heterogeneous grid platform. This method computes one or few eigenpairs of a large sparse real symmetric matrix. We take the specificities of computational resources into account and deal with communications over the Internet by means of techniques such as out-of-core and data persistence. We also show that a restarted algorithm and the combination of several paradigms of parallelism are interesting in this context. We perform many experimentations using several parameters related to the Lanczos method and the configuration of the platform. Depending on the number of computed Ritz eigenpairs, the results underline how critical the choice of the dimension of the working subspace is. Moreover, the size of platform has to be scaled to the order of the eigenproblem because of communications over the Internet. Laurent Choy, Serge G. Petiton, Mitsuhisa Sato |
CCGRID | 3 |
| 2007 | Toward power-aware computing with dynamic voltage scaling for heterogeneous platformsabstractEnergy conservation is a dynamic topic of research in high performance computing and cluster computing. Power-aware computing for heterogeneous world-wide Grid is a new track of research. In this work, we study and evaluate the impact of the heterogeneity of the nodes of a computing platform on the energy consumption. We propose to take advantage of this heterogeneity in order to save energy with no significant loss of performance by using dynamic voltage scaling (DVS) in a distributed eigensolver. We show that using DVS only during the slack-time does not penalize the performances but it does not provide significant energy savings. If DVS is applied to all the execution, we get important global and local energy savings (respectively up to 9% and 20%) without a significant rise of the wall-clock times. Laurent Choy, Serge G. Petiton, Mitsuhisa Sato |
CLUSTER | 3 |
| 2007 | Bandwidth-Aware Design of Large-Scale Clusters for Scientific Computations
Mitsuhisa Sato |
HPCC | 1 |
| 2007 | RI2N/UDP: High bandwidth and fault-tolerant network for a PC-cluster based on multi-link EthernetabstractPC-clusters with high performance/cost ratio have been one of the typical platforms for high performance computing. To lower costs, Gigabit Ethernet is often used for intercommunication networks. However, the reliability of Ethernet is limited due to hardware failures and tentative errors in the network switches. To solve this problem, we propose an interconnection network system based on multi-link Ethernet named RI2N. In this paper, we developed a user level implementation of RI2N using UDP/IP that is called RI2N/UDP. When this new system was evaluated for performance and fault tolerance, the bandwidth on a 2-link Gigabit Ethernet was 246 MB/s, and the system could remain active during network link failure to provide high system reliability. Takayuki Okamoto, Shin'ichi Miura, Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi |
IPDPS | 4 |
| 2007 | Direct Execution of Linux Binary on Windows for Grid RPC WorkersabstractLocal area or campus-type networks consist of PCs using different operating systems such as Windows and Linux. These PCs are expected to have enormous potential computing power for grid computing. The majority of PCs in this type of environment run on Windows, while grid applications and middleware are often developed on Linux. The challenge is to absorb the heterogeneity of operating systems. Grid RPC is a promising programming model for the development of grid applications. We have designed and implemented an agent called BEE, which enables direct execution of Linux binary programs on Windows for a grid RPC worker. We have integrated the BEE agent into an OmniRPC system in order to make use of Windows PCs as computing resources in a hybrid grid environment combining Windows PCs into grid computing resources. The BEE agent allows Linux binaries of the program of the OmniRPC worker to be exported and run under Windows without any modification of its Linux binaries. The results of our experiments show that the performance of a worker program using BEE is almost the same as that of Windows native binary and Cygwin and is better that that of using VMware. We have demonstrated a hybrid grid environment combining Window PCs in a conventional grid of Linux nodes. Yoshifumi Uemura, Yoshihiro Nakajima, Mitsuhisa Sato |
IPDPS | 3 |
| 2006 | PACS-CS: A Large-Scale Bandwidth-Aware PC Cluster for Scientific ComputationsabstractWe have been developing a large scale PC cluster named PACS-CS (Parallel Array Computer System for Computational Sciences) at Center for Computational Sciences, University of Tsukuba, for wide variety of computational science applications such as computational physics, computational material science, computational biology, etc. We consider the most important issue on the computation node is the memory access bandwidth, then a node is equipped with a single CPU which is different from ordinary high-end PC clusters. The interconnection network for parallel processing is configured as a multi-dimensional hyper-crossbar network based on trunking of Gigabit Ethernet to support large scale scientific computation with physical space modeling. Based on the above concept, we are developing an original mother board to configure a single CPU node with 8 ports of Gigabit Ethernet, which can be implemented in the half size of 19 inch rack-mountable 1U size platform. Under the preliminary performance evaluation, we confirmed that the computation part in practical Lattice QCD code will be able to achieve 30% of peak performance, and up to 600 Mbyte/sec of bandwidth at single directed neighboring communication will be achieved. PACS-CS will start its operation on July 2006 with 2560 CPUs and 14.3 Tflops of peak performance. Taisuke Boku, Mitsuhisa Sato, Akira Ukawa, Daisuke Takahashi, Shinji Sumimoto, Kouichi Kumon, Takashi Moriyama, Masaaki Shimizu |
CCGRID | 2 |
| 2006 | Integrating Computing Resources on Multiple Grid-enabled Job Scheduling Systems Through a Grid RPC SystemabstractWe present a framework for a parallel programming model by remote procedure calls bridging between largescale computing resource pools managed by multiple gridenabled job scheduling systems. With this system, the user can exploit not only each remote servers and clusters, but also computing resources provided with grid-enabled job scheduling systems located on different sites. This framework requires a Grid RPC system to decouple the computation in a remote node from the Grid RPC mechanism and uses document-based communication rather than connection-based communication. We implemented the proposed framework as an extension of the OmniRPC system, which is a Grid RPC system for parallel programming in a grid environment. We designed a general interface to adapt the OmniRPC system to various grid-enabled job scheduling systems easily and applied the proposed system to several grid-enabled job scheduling systems, including XtremWeb, CyberGRIP, Condor and Grid Engine. we show the preliminary performance of these implementations using a phylogenetic application. We found that the proposed system can achieve approximately the same performance as using OmniRPC and can handle interruptions in worker programs on remote nodes. Yoshihiro Nakajima, Mitsuhisa Sato, Yoshiaki Aida, Taisuke Boku, Franck Cappello |
CCGRID | 2 |
| 2006 | Emprical study on Reducing Energy of Parallel Programs using Slack Reclamation by DVFS in a Power-scalable High Performance ClusterabstractIt has become important to improve the energy efficiency of high performance PC clusters. In PC clusters, high-performance microprocessors have a dynamic voltage and frequency scaling (DVFS) mechanism, which allows the voltage and frequency to be set for reduction in energy consumption. In this paper, we proposed a new algorithm that reduces energy consumption in a parallel program executed on a power-scalable cluster using DVFS. Whenever the computational load is not balanced, parallel programs encounter slack time, that is, they must wait for synchronization of the tasks. Our algorithm reclaims slack time by changing the voltage and frequency, which allows a reduction in energy consumption without impacting on the performance of the program. Our algorithm can be applied to parallel programs represented by a directed acyclic task graph (DAG). It selects an appropriate set of voltages and frequencies (called the gear) that allow the tasks to execute at the lowest frequency that does not increase the overall execution time, but at the same time allows the tasks to be executed as uniformly as possible in frequency. We built two different types of power-scalable clusters using AMD Turion and Transmeta Crusoe. For the empirical study on energy reduction in PC clusters, we designed a toolkit called PowerWatch that includes power monitoring tools and the DVFS control library. This toolkit precisely measures the power consumption of the entire cluster in real time. The experimental results using benchmark problems show that our algorithm reduces energy consumption by 25% with only a 1 % loss in performance Hideaki Kimura 0003, Mitsuhisa Sato, Yoshihiko Hotta, Taisuke Boku, Daisuke Takahashi |
CLUSTER | 2 |
| 2006 | Performance Improvement by Data Management Layer in a Grid RPC System
Yoshiaki Aida, Yoshihiro Nakajima, Mitsuhisa Sato, Tetsuya Sakurai, Daisuke Takahashi, Taisuke Boku |
GPC | 3 |
| 2006 | A scalable communication layer for multi-dimensional hyper crossbar network using multiple gigabit ethernetabstractThis paper proposes a scalable communication layer for a multi-dimensional hyper crossbar network using multiple Gigabit Ethernet for the PACS-CS system which consists of 2560 single-processor nodes and a 16 x 16 x 10 three dimensional hyper-crossbar network (3D-HXB). To realize a high performance communication layer using multiple existing Ethernet networks, the host processor usage for the communication processing must be reduced to less than the appropriate packet processing time which is calculated from a message size and a target communication bandwidth. To overcome this problem, we have developed the PM/Ethernet-HXB communication facility. PM/Ethernet-HXB realizes communication protocol processing without exclusion even for Zero-copy communication between the communication buffers of nodes. We have implemented the PM/Ethernet-HXB on SCore cluster system software, and evaluated its communication and application performance. PM/Ethernet-HXB achieves a unidirectional communication bandwidth of 1065 MB/s using nine Gigabit Ethernet links on a single dimension network. It also realizes a unidirectional communication bandwidth of 741 MB/s (98.8% of the theoretical performance) and a bidirectional bandwidth of 1401 MB/s (93.4% of the theoretical performance) on the three dimensional connections (3D-HXB: a total of six Ethernet links). The results of MPI communication bandwidth are a unidirectional communication bandwidth of 960 MB/s and a bidirectional bandwidth of 1008 MB/s using eight links on a single dimension network. These results show that PM/Ethernet-HXB realizes a comparative performance using multiple Gigabit Ethernet networks to dedicated cluster networks such as InfiniBand 4x (1000 MB/s). The speedups of IS and CG Class C NAS parallel benchmarks are scalable up to using four links on eight node cluster, and performance degradation between 3D-HXB (2 x 2 x 2) and 1-dimensional network is small. Shinji Sumimoto, Kazuichi Oe, Kouichi Kumon, Taisuke Boku, Mitsuhisa Sato, Akira Ukawa |
ICS | 5 |
| 2006 | MegaProto/E: power-aware high-performance cluster with commodity technologyabstractIn our research project named "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we have been developing a prototype cluster not based on ASIC or FPGA but instead only using commodity technology. Its packaging is extremely compact and dense, and its performance/power ratio is very high. Our previous prototype system named "MegaProto" demonstrated that one cluster unit, which consists of 16 commodity low-power processors, can be successfully implemented on just 1U height chassis and it is capable of up to 2.8 times higher performance/power ratio than ordinary high-performance dual-Xeon 1U server units. We have improved MegaProto by replacing the CPU and enhancing the I/O performance. The new cluster unit named "MegaProto/E" with 16 Transmeta Efficeon processors achieves 32 GFlops of peak performance, which is 2.2-fold greater than that of the original one. The cluster unit is equipped with an independent dual network of Gigabit Ethernet, including dual 24-port switches. The maximum power consumption of the cluster unit is 320 W, which is comparable with that of today's high-end PC servers for high performance clusters. Performance evaluation using NPB kernels and HPL shows that the performance of MegaProto/E exceeds that of a dual-Xeon server in all the benchmarks, and its performance ratio ranges from 1.3 to 3.7. These results reveal that our solution of implementing a number of ultra low-power processors in compact packaging is an excellent way to achieve extremely high performance in applications with a certain degree of parallelism. We are now building a multi-unit cluster with 128 CPUs (8 units) to prove that this advantage still holds with higher scalability Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi, Hiroshi Nakashima, Hiroshi Nakamura, Satoshi Matsuoka, Yoshihiko Hotta |
IPDPS | 2 |
| 2006 | Profile-based optimization of power performance by using dynamic voltage scaling on a PC clusterabstractCurrently, several of the high performance processors used in a PC cluster have a DVS (dynamic voltage scaling) architecture that can dynamically scale processor voltage and frequency. Adaptive scheduling of the voltage and frequency enables us to reduce power dissipation without a performance slowdown during communication and memory access. In this paper, we propose a method of profiled-based power-performance optimization by DVS scheduling in a high-performance PC cluster. We divide the program execution into several regions and select the best gear for power efficiency. Selecting the best gear is not straightforward since the overhead of DVS transition is not free. We propose an optimization algorithm to select a gear using the execution and power profile by taking the transition overhead into account. We have built and designed a power-profiling system, PowerWatch. With this system we examined the effectiveness of our optimization algorithm on two types of power-scalable clusters (Crusoe and Turion). According to the results of benchmark tests, we achieved almost 40% reduction in terms of EDP (energy-delay product) without performance impact (less than 5%) compared to results using the standard clock frequency. Yoshihiko Hotta, Mitsuhisa Sato, Hideaki Kimura 0003, Satoshi Matsuoka, Taisuke Boku, Daisuke Takahashi |
IPDPS | 2 |
| 2006 | Storage challenge - High performance data analysis for particle physics using the Gfarm file systemabstractThe Belle experiment operates at the KEKB accelerator, a high luminosity asymmetric energy e+ e- collider. The Belle collaboration studies CP violations in decays of B mesons to answer one of the fundamental questions of Nature, the matter-anti-matter asymmetry. Currently, Belle accumulates more than one million B Bbar meson pairs, corresponding to about 1.2 TB of raw data, per day.The challenge is how to realize the required high performance data access and scalable data computing. The Gfarm file system is a Grid-wide network shared file system that federates local storage of cluster nodes; moreover it provides scalable I/O performance with distributed data access. In the challenge, we will construct a Gfarm file system with 40 TB capacity and 30 GB/sec I/O bandwidth, integrating the local disks of 800 compute nodes in the KEKB computing facility, and demonstrate high-performance Belle data analysis. Nobuhiko Katayama, Mitsuhisa Sato, Taisuke Boku, Akira Ukawa, Shohei Nishida, Ichiro Adachi, Osamu Tatebe |
SC | 2 |
| 2006 | Editorial: Special Issue on Global and Peer-to-Peer Computing
Franck Cappello, Adriana Iamnitchi, Mitsuhisa Sato |
J. Grid Comput. | 3 |
| 2005 | Grid and Cluster Matrix Computation with Persistent Storage and Out-of-core ProgrammingabstractIn this paper we present a performance evaluation of a large-scale numerical application on a cluster and a global grid/cluster platform. The computational resources are a cluster of clusters (34 nodes, 84 processors) and a local area network grid (128 nodes), distributed on two geographic sites: Tsukuba University (Japan) and University of Lille I (France). We compare a classical MPI (message passing interface) version with global grid/cluster versions. We also present and test some techniques for numerical applications on a grid/cluster infrastructure based on out-of-core programming and an efficient data placement. We discuss the performances of a block-based Gauss-Jordan method for large matrix inversion. As experimental grid middleware we use the XtremWeb system to manage non-dedicated distributed resources Lamine M. Aouad, Serge G. Petiton, Mitsuhisa Sato |
CLUSTER | 3 |
| 2005 | MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity TechnologyabstractIn our research project "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we claim that a million-scale parallel system could be built with densely mounted low-power commodity processors. "MegaProto" is a proof-of-concept low-power and highperformance cluster build only with commodity components to implement this claim. A one-rack system is composed of 32 motherboard "cluster units" of 1 U-height and commodity switches to interconnect them mutually as well as with other racks. Each cluster unit houses 16 low-power dollarbill- sized commodity PC-architecture daughterboards, together with a high bandwidth, 2 Gbps per processor embedded switched network based on Gigabit Ethernet. The peak performance of a one-rack system is 0.48 TFlops for the first version and will improve to 1.02 TFlops in the second version through a processor/daughterboard upgrade. The system consumes about 10 kW or less per rack, resulting in 100 MFlops/W power efficiency with a power-aware intrarack network of 32 Gbps bisection bandwidth, while additional 2.4 kW will boost this to sufficiently large 256 Gbps. Performance studies show that even the first version significantly outperforms a conventional high-end 1U server comprised of dual power-hungry processors in a majority of NPB programs. It is also investigated how the current automated DVS control could save power for the HPC parallel programs along with its limitation. Hiroshi Nakashima, Hiroshi Nakamura, Mitsuhisa Sato, Taisuke Boku, Satoshi Matsuoka, Daisuke Takahashi, Yoshihiko Hotta |
SC | 3 |
| 2005 | OpenGR: A directive-based grid programming environment
Motonori Hirano, Mitsuhisa Sato, Yoshio Tanaka |
Parallel Comput. | 2 |
| 2005 | Editorial
Daniel A. Reed, Mitsuhisa Sato, Denis Trystram |
Parallel Comput. | 2 |
| 2004 | Implementation and performance evaluation of CONFLEX-G: grid-enabled molecular conformational space search program with OmniRPCabstractCONFLEX-G is the grid-enabled version of a molecular conformational space search program called CONFLEX. We have implemented CONFLEX-G using a grid RPC system called OmniRPC. In this paper, we report the performance of CONFLEX-G in a grid testbed of several geographically distributed PC clusters. In order to explore many conformation of large bio-molecules, CONFLEX-G generates trial structures of the molecules and allocates jobs to optimize a trial structure with a reliable molecular mechanics method in the grid. OmniRPC provides a restricted persistence model to support the parametric search applications. In this model, when the initialization procedure is defined in the RPC module, the module is automatically initialized at the time of invocation by calling the initialization procedure. This can eliminate unnecessary communication and initialization at each call in CONFLEX-G. CONFLEX-G can achieve performance comparable to CONFLEX MPI and can exploit more computing resources by allowing the use of a cluster of multiple clusters in the grid. The experimental result shows that CONFLEX-G achieved a speedup of 56.5 times in the case of the 1BL1 molecule, where the molecule consists of a large number of atoms, and each trial structure optimization requires significant time. The load imbalance of the optimization time of the trial structure may also cause performance degradation. Yoshihiro Nakajima, Mitsuhisa Sato, Hitoshi Gotoh, Taisuke Boku, Daisuke Takahashi |
ICS | 2 |
| 2004 | Parallel Implementation of Strassen's Matrix Multiplication Algorithm for Heterogeneous ClustersabstractSummary form only given. We propose a new distribution scheme for a parallel Strassen's matrix multiplication algorithm on heterogeneous clusters. In the heterogeneous clustering environment, appropriate data distribution is the most important factor for achieving maximum overall performance. However, Strassen's algorithm reduces the total operation count to about 7/8 times per one recursion and, hence, the recursion level has an effect on the total operation count. Thus, we need to consider not only load balancing but also the recursion level in Strassen's algorithm. Our scheme achieves both load balancing and reduction of the total operation count. As a result, we achieve a speedup of nearly 21.7% compared to the conventional parallel Strassen's algorithm in a heterogeneous clustering environment. Yuhsuke Ohtaki, Daisuke Takahashi, Taisuke Boku, Mitsuhisa Sato |
IPDPS | 4 |
| 2003 | HMCS-G: Grid-enabled Hybrid Computing System for Computational AstrophysicsabstractThe authors have developed a hybrid computing system called HMCS-G, a Grid-enabled Heterogeneous Multi-Computer System, that provides a multiple cluster environment centered around a dedicated machine for gravity calculation. The purpose of HMCS-G is to provide an ideal computational environment for astrophysical study involving multiple physical phenomena. The worker cluster may comprise general-purpose PCs to perform tasks such as hydrodynamics computations, while the special-purpose machine, in this case a GRAPE-6 cluster, performs gravity calculations for all pairs of particles in the system. These systems are connected by OmniRPC, a grid-enabled RPC system that supports Globus and ssh for authentication. HMCS-G effectively provides worldwide access to a GRAPE-6 cluster, thereby securing several TFLOPS performance for intensive computations such as gravity calculation. All participating PC-clusters share this resource in a time-based manner using grid technology. The actual turn-around response time was measured for a system implemented over a number of institutions, and it was confirmed that HMCS-G provides acceptable real-world application performance. Precise simulations of galaxy formation are currently being performed on clusters in several institutes, involving smoothed particle hydrodynamics and radiative transfer in the context of complete gravity calculation as the first real application of HMCS-G. Taisuke Boku, Mitsuhisa Sato, Kenji Onuma, Junichiro Makino, Hajime Susa, Daisuke Takahashi, Masayuki Umemura, Akira Ukawa |
CCGRID | 2 |
| 2003 | Performance of Cluster-enabled OpenMP for the SCASH Software Distributed Shared Memory SystemabstractOpenMP has attracted widespread interest because it is an easy-to-use parallel programming model for shared memory multiprocessor systems. Implementation of a "cluster-enabled" OpenMP compiler is presented. Compiled programs are linked to the page-based software distributed-shared-memory system, SCASH, which runs on PC clusters. This allows OpenMP programs to be run transparently in a distributed memory environment. The compiler converts programs written for OpenMP into parallel programs using the SCASH static library, moving all shared global variables into SCASH shared address space at runtime. As data mapping has a great impact on the performance of OpenMP programs compiled for software distributed-shared-memory, extensions to OpenMP directives are defined for specifying data mapping and loop scheduling behavior, allowing data to be allocated to the node where it is to be processed. Experimental results of benchmark programs on PC clusters using both Myrinet and fast Ethernet are reported. Yoshinori Ojima, Mitsuhisa Sato, Hiroshi Harada, Yutaka Ishikawa |
CCGRID | 2 |
| 2003 | Preliminary Evaluation of Dynamic Load Balancing Using Loop Re-partitioning on Omni/SCASHabstractIncreasingly large-scale clusters of PC/WS continue to become majority platform in HPC field. Such a commodity cluster environment, there may be incremental upgrade due to several reasons, such as rapid progress in processor technologies, or user needs and it may cause the performance heterogeneity between nodes from which the application programmer will suffer as load imbalances. To overcome these problems, some dynamic load balancing mechanisms are needed. In this paper, we report our ongoing work on dynamic load balancing extension to Omni/SCASH which is an implementation of OpenMP on Software Distributed Shared Memory, SLASH. Using our dynamic load balancing mechanisms, we expect that programmers can have load imbalances adjusted automatically by the runtime system without explicit definition of data and task placements in a commodity cluster environment with possibly heterogeneous performance nodes. Yoshiaki Sakae, Mitsuhisa Sato, Satoshi Matsuoka, Hiroshi Harada |
CCGRID | 2 |
| 2003 | OmniRPC: a Grid RPC ystem for Parallel Programming in Cluster and Grid EnvironmentabstractWe have designed and implemented a Grid RPC system called OmniRPC, for parallel programming in cluster and grid environments. While OmniRPC inherits its API from Ninf, the programmer can use OpenMP for easy-to-use parallel programming because the API is designed to be thread-safe. To support typical master-worker grid applications such as a parametric execution, OmniRPC provides an automatic-initializable remote module to send and store data to a remote executable invoked in the remote host. Since it may accept several requests for subsequent calls by keeping the connection alive, the data set by the initialization is re-used, resulting in efficient execution by reducing the amount of communication. The OmniRPC system also supports a local environment with "rsh", a grid environment with Globus, and remote hosts with "ssh". Furthermore, the user can use the same program over OmniRPC for both clusters and grids because a typical grid resource is regarded simply as a cluster of clusters distributed geographically. For a cluster over a private network, an agent process running the server host functions as a proxy to relay communications between the client and the remote executables by multiplexing the communications into one connection to the client. This feature allows a single client to use a thousand of remote computing hosts. Mitsuhisa Sato, Taisuke Boku, Daisuke Takahashi |
CCGRID | 1 |
| 2002 | A Blocking Algorithm for Parallel 1-D FFT on Clusters of PCs
Daisuke Takahashi, Taisuke Boku, Mitsuhisa Sato |
Euro-Par | 3 |
| 2002 | Exploiting cluster networks for distributed object groups and collective operations
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Future Gener. Comput. Syst. | 2 |
| 2001 | TACO-Exploiting Cluster Networks for High-Level Collective OperationsabstractTACO (Topologies and Collections) is a template library that introduces the flavour of distributed data parallel processing by means of reusable topology classes and C++ templates. The paper introduces TACO's basic abstractions and provides a performance analysis for basic collective operations on various cluster architectures with several different networks. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
CCGRID | 2 |
| 2000 | TACO -- Dynamic Distributed Collections with Templates and Topologies
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Euro-Par | 2 |
| 2000 | Performance Evaluation of a Firewall-Compliant Globus-based Wide-Area Cluster SystemabstractPresents a performance evaluation of a wide-area cluster system based on a firewall-enabled Globus metacomputing toolkit. In order to establish communication links beyond the firewall, we have designed and implemented a resource manager called RMF (Resource Manager beyond the Firewall) and the Nexus Proxy, which relays TCP communication links beyond the firewall. In order to extend the Globus metacomputing toolkit to become firewall-enabled, we have built the Nexus Proxy into the Globus toolkit. We have built a firewall-enabled Globus-based wide-area cluster system in Japan and run some benchmarks on it. In this paper, we report various performance results, such as the communication bandwidth and latencies obtained, as well as application performance involving a tree search problem. In a wide-area environment, the communication latency through the Nexus Proxy is approximately six times larger when compared to that of direct communications. As the message size increases, however, the communication overhead caused by the Nexus Proxy can be negligible. We have developed a tree search problem using MPICH-G. We used a self-scheduling algorithm, which is considered to be suitable for a distributed heterogeneous metacomputing environment since it performs dynamic load balancing with low overhead. The performance results indicate that the communication overhead caused by the Nexus Proxy is not a severe problem in metacomputing environments. Yoshio Tanaka, Motonori Hirano, Mitsuhisa Sato, Hidemoto Nakada, Satoshi Sekiguchi |
HPDC | 3 |
| 2000 | Template Based Structured CollectionsabstractCollective operations on distributed data sets foster a high-level data-parallel programming style that eases many aspects of parallel programming significantly. In this paper we describe how higher-order collective operations on distributed object sets can be introduced in a structured way by means of reusable topology classes and C++ templates. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
IPDPS | 2 |
| 2000 | Network interface active messages for low overhead communication on SMP PC clusters
Motohiko Matsuda, Yoshio Tanaka, Kazuto Kubota, Mitsuhisa Sato |
Future Gener. Comput. Syst. | 4 |
| 1999 | Design and implementations of Ninf: towards a global computing infrastructure
Hidemoto Nakada, Mitsuhisa Sato, Satoshi Sekiguchi |
Future Gener. Comput. Syst. | 2 |
| 1998 | Practical Simulation of Large-Scale Parallel Programs and Its Performance Analysis of the NAS Parallel Benchmarks
Kazuto Kubota, Ken'ichi Itakura, Mitsuhisa Sato, Taisuke Boku |
Euro-Par | 3 |
| 1998 | Ninflet: a migratable parallel objects framework using JavaabstractNinflet is a Java-based global computing system that builds on our experiences with the Ninf system which facilitated RPC-based computing of numerical tasks in a wide-area network. The goal of Ninflet is to become a new generation of concurrent object-oriented systems which harness abundant idle computing powers, and also seamlessly integrate global as well as local network parallel computing. Ninflet is designed to make use of Java features to implement important features in global computing, such as resource allocation, inter-Ninflet communication, security, checkpointing, object migration, and easy server management via HTTP. © 1998 John Wiley & Sons, Ltd. Hiromitsu Takagi, Satoshi Matsuoka, Hidemoto Nakada, Satoshi Sekiguchi, Mitsuhisa Sato, Umpei Nagashima |
Concurr. Pract. Exp. | 5 |
| 1998 | Ninf and PM: Communication libraries for global computing and high-performance cluster computing
Mitsuhisa Sato, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa, Satoshi Sekiguchi, Hidemoto Nakada, Satoshi Matsuoka, Umpei Nagashima |
Future Gener. Comput. Syst. | 1 |
| 1997 | Multi-client LAN/WAN Performance Analysis of Ninf: a High-Performance Global Computing SystemabstractRapid increase in speed and availability of network of supercomputers is making high-performance global computing possible, including our Ninf system. However, critical issues regarding system performance characteristics in global computing have been little investigated, especially under multi-client, multi-site WAN settings. In order to investigate the feasibility of Ninf and similar systems, we conducted benchmarks under various LAN and WAN environments, and observed the following results: 1) Given sufficient communication bandwidth, Ninf performance quickly overtakes client local performance, 2) current supercomputers are sufficient platforms for supporting Ninf and similar systems in terms of performance and OS fault resiliency, 3) for a vector-parallel machine (Cray J90), employing optimized data-parallel library is a better choice compared to conventional task-parallel execution employed for non-numerical data servers, 4) computationally intensive tasks such as EP can readily be supported under the current Ninf infrastructure, and 5) for communication-intensive applications such as Linpack, server CPU utilization dominates LAN performance, while communication bandwidth dominates WAN performance, and furthermore, aggregate bandwidth could be sustained for multiple clients located at different Internet sites; as a result, distribution of multiple tasks to computing servers on different networks would be essential for achieving higher client-observed performance. Our results are not necessarily restricted to the Ninf system, but rather, would be applicable to other similar global computing systems. Atsuko Takefusa, Satoshi Matsuoka, Hirotaka Ogawa, Hidemoto Nakada, Hiromitsu Takagi, Mitsuhisa Sato, Satoshi Sekiguchi, Umpei Nagashima |
SC | 6 |
| 1997 | Fine-Grain Multithreading with the EM-X MultiprocessorabstractArticle Fine-grain multithreading with the EM-X multiprocessor Share on Authors: Andrew Sohn Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Yuetsu Kodama Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Jui Ku Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Mitsuhisa Sato Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, Japan Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, JapanView Profile , Hirofumi Sakane Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Hayato Yamana Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Shuichi Sakai Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Yoshinori Yamaguchi Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile Authors Info & Claims SPAA '97: Proceedings of the ninth annual ACM symposium on Parallel algorithms and architecturesJune 1997 Pages 189–198https://doi.org/10.1145/258492.258511Online:01 June 1997Publication History 7citation121DownloadsMetricsTotal Citations7Total Downloads121Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Andrew Sohn, Yuetsu Kodama, Jui Ku, Mitsuhisa Sato, Hirofumi Sakane, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi |
SPAA | 4 |
| 1997 | Data and Workload Distribution in a Multithreaded Architecture
Andrew Sohn, Mitsuhisa Sato, Namhoon Yoo, Jean-Luc Gaudiot |
J. Parallel Distributed Comput. | 2 |
| 1997 | Performance Evaluation of a Workstation Cluster, TMC CM-5, and Intel Paragon/XP Using a Parallel Homology Analysis Program
Satoko Sakata, Umpei Nagashima, Mitsuhisa Sato, Satoshi Sekiguchi, Haruo Hosoya |
Parallel Comput. | 3 |
| 1995 | Multithreading with the EM-4 distributed-memory multiprocessor
Andrew Sohn, Chinhyun Kim, Mitsuhisa Sato |
PACT | 3 |
| 1995 | A Macrotask-level Unlimited Speculative Execution on MultiprocessorsabstractThe purpose of this paper is to propose a new fast EM-4 multiprocessor to decrease the control overhead of macrotasks.Preliminary evaluations show that the control overhead of the proposed scheme is smaller than that of the other control schemes.Moreover, it is confirmed that the distributed control can be implemented by using software when the average macrotssk execution time is larger than 14.4ps on the EM-4 multiprocessor. Hayato Yamana, Mitsuhisa Sato, Yuetsu Kodama, Hirofumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi |
International Conference on Supercomputing | 2 |
| 1995 | The EM-X Parallel Computer: Architecture and Basic PerformanceabstractLatency tolerance is essential in achieving high performance on parallel computers for remote function calls and fine-grained remote memory accesses. EM-X supports interprocessor communication on an execution pipeline with small and simple packets. It can create a packet in one cycle, and receive a packet from the network in the on-chip buffer without interruption. EM-X invokes threads on packet arrival, minimizing the overhead of thread switching. It can tolerate communication latency by using efficient multi-threading and optimizing packet flow of fine grain communication. EM-X also supports the synchronization of two operands, direct remote memory read/write operations and flexible packet scheduling with priority. This paper describes distinctive features of the EM-X architecture and reports the performance of small synthetic programs and larger more realistic programs. Yuetsu Kodama, Hirohumi Sakane, Mitsuhisa Sato, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi |
ISCA | 3 |
| 1995 | An Experience with Super-Linear Speedup Achieved by Parallel Computing on a Workstation Cluster: Parallel Calculation of Density of States of Large Scale Cyclic Polyacenes
Umpei Nagashima, Sachiko Hyugaji, Satoshi Sekiguchi, Mitsuhisa Sato, Haruo Hosoya |
Parallel Comput. | 4 |
| 1995 | Reduced Interprocessor-Communication Architecture and its Implementation on EM-4
Shuichi Sakai, Yuetsu Kodama, Mitsuhisa Sato, Andrew Shaw, Hiroshi Matsuoka, Hideo Hirono, Kazuaki Okamoto, Takashi Yokota |
Parallel Comput. | 3 |
| 1994 | Nonnumeric search results on the EM-4 distributed-memory multiprocessorabstractNumeric scientific problems have been the main focus of supercomputing as their numerous implementations on various multiprocessors indicate. Nonnumeric problems on the other hand have received very little attention for parallel implementation due to their irregular behaviors and a large amount of resource usage. This report presents our experiences implementing difficult nonnumeric problems on the EM-4 multiprocessor, believing that supercomputers should also be able to effectively execute nonnumeric problems if they are to be considered 'supercomputers'. We selected two typical search problems, the Eight-Puzzle and the Tower-of-Hanoi. Two parallel search techniques we used to implement the search problems, unidirectional and bidirectional heuristic search. A total of eight different programs have been implemented on the EM-4 multiprocessor with realistic problem sizes. Execution results demonstrate that the parallel bidirectional heuristic search can solve the tree depth 20 to 40 of the Eight-Puzzle in an optimal or near optimal number of iterations in less than two seconds, and is highly scalable as it gives over 40-fold speedup for both problems on 80 processors.> Andrew Sohn, Mitsuhisa Sato, Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi |
SC | 2 |
| 1993 | EMC-Y: Parallel Processing Element Optimizing Communication and ComputationabstractEMC-Y is a new processing element for highly parallel computers designed to achieve high performance parallel computation by fusing a dataflow mechanism and a von Neumann execution pipeline. We have already developed EMC-R, which is the processing element used in the EM-4 prototype. EMC-Y improves on EMC-R's packet communication performance, allowing it to tolerate a more network traffic. This paper presents the architecture of EMC-Y, concentrating on the principles of packet communication. EMC-Y uses an output packet buffer and optimal packet routing to improve the performance of packet sending and transferring. EMC-Y changes the memory access priority for input packet buffer operation to improve the performance of receiving packets. Since the EMC-Y processor not only improves the performance of packet input and output but also balances them, it can tolerate a large amount of traffic and can improve the execution performance. We evaluate the improvements of EMC-Y architecture using a clock level simulator. The results show that EMC-Y improves performance by 50% to 70% in several programs over EMC-R at the same clock speed. Yuetsu Kodama, Yasuhito Koumura, Mitsuhisa Sato, Hirohumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi |
International Conference on Supercomputing | 3 |
| 1993 | Data Stream Control Optimization in Dataflow ArchitecturesabstractIn order to achieve implementation control in a dataflow architecture, it is necessary to bring data streams under control and branch them in accord with prevailing conditions. This study suggests methodologies for modeling and formulating the control overhead of data streams. From these methodologies, techniques are derived to minimize overhead. In order to verify the practicality of these techniques, they were first applied to three representative Japanese dataflow architectures to show their basic ability to reduce overhead, and for further verification, they were next applied to several real applications to test their more specific abilities, as well as their value as techniques for optimizing compilation. Shorin Kyo, Satoshi Sekiguchi, Mitsuhisa Sato |
International Conference on Supercomputing | 3 |
| 1993 | Super-Threading: Architectural and Software Mechanisms for Optimizing Parallel ComputationabstractThis paper presents super-threading, which generically means the architectural and software mechanisms for optimizing parallel computation. Super-threading includes architectural optimization of a processing element (PE), mechanism for supporting fast communication and computation, techniques of a compiler and a run time system for optimizing thread creation, thread allocation, tuning of granularity and data allocation to physically distributed storage.This paper states what super-threading is and examines some of the technologies belonging to it. The processor architecture based on super-threading is proposed and its implementation on a highly parallel computer EM-4 is shown with performance data. Software issues about super-threading are also examined mainly from the viewpoint of granularity optimization. Dynamic granularity optimization methods are proposed here, and evaluated on EM-4. The performance data indicate that super-threading is a key technology for realizing an efficient massively parallel computer. Shuichi Sakai, Kazuaki Okamoto, Hiroshi Matsuoka, Hideo Hirono, Yuetsu Kodama, Mitsuhisa Sato |
International Conference on Supercomputing | 6 |
| 1992 | Thread-based Programming for the EM-4 Hybrid Dataflow MachineabstractIn this paper, we present a thread-based programming model for the EM-4 hybrid dataflow machine, where parallelism and synchronization among threads of sequential execution are described explicitly by the programmer. Although EM-4 was originally designed as a dataflow machine, we demonstrate that it provides effective architectural support for a variety of programming styles, including message passing and distributed data sharing in imperative languages. Our approach allows the programmer to control the parallelism and maintain data locality explicitly to achieve high performance. EM-4 can be thought of as a multi-threaded architecture that can exploit both von Neumann and dataflow compiling technology. Thread-based programming provides the first step to explore better programming/compiling technology for a hybrid dataflow machine as well as EM-4. Mitsuhisa Sato, Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi, Yasuhito Koumura |
ISCA | 1 |
| 1989 | Run-Time Checking in Lisp by Integrating Memory Addressing and Range CheckingabstractThis paper describes the BL addressing mode and the address tag in FLATS2 machine, which is a general-purpose MIMD computer now under construction. The BL addressing mode integrates memory accessing and range checking by hardware. Address tag is a bit in word, which indicates the capability for memory access. Combining them together, efficient memory protection is provided at run-time. It reduces the cost of run-time type checking in Lisp by checking the address tag and the address of a pointer against the range of the region associated to a type, in parallel with the memory access. The arithmetic instructions check the address tags of operands to support the generic arithmetic in Lisp. We can also make use of this scheme to check the number of arguments and multiple return values and to check array-bounds to support faster execution of Common Lisp program. These facilities are not specific to Lisp, so that they can be used more generally than other tagged architectures. Mitsuhisa Sato, Shuichi Ichikawa, Eiichi Goto |
ISCA | 1 |