EDBT 2026 Demo / reviewers in the wild / expert
Masahiro Nakao
dblp:16/4082
· DBLP profile ↗
15ranked-venue papers
8as first author
4since 2021 · last 2024
0000-0001-7848-1172ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Analysis Towards Energy-Aware Image-based In Situ Visualization on the FugakuabstractEnergy efficiency has become a serious concern when running applications on HPC systems. Although these systems were designed to mainly run simulation codes as fast as possible, due to the ever-increasing size of the simulation outputs, the in situ visualization has gained increasing attention. In situ visualization uses the same HPC system to execute a part or even the entire visualization processing, and there are currently a variety of tools and libraries, that facilitate domain scientists to integrate them with their simulation codes. Among different approaches, image- and video-based in situ visualization has been widely adopted as an effective approach for the subsequent offline visual analysis. In this approach, a large number of renderings are required at every visualization time step and can consume a considerable computational resource. Fugaku adopted PowerAPI which enables the users to set the power mode for their jobs. However, simulation and visualization codes may have different processing behaviors requiring different power settings for obtaining the most energy-efficient runnings. In this work, we tried to shed light on the energy efficiency of the visualization portion that was not considered before. We investigated the computational cost and energy consumption of some rendering techniques by using the PowerAPI and KVS (Kyoto Visualization System) on the Fugaku, and hope that the obtained findings will be useful for potential users looking to run in situ visualization on the Fugaku and other PowerAPI-enabled HPC systems. Razil Tahir, Jorji Nonaka, Ken Iwata, Taisei Matsushima, Naohisa Sakamoto, Chongke Bi, Masahiro Nakao, Hitoshi Murai |
HPC Asia | 7 |
| 2024 | Doubling Graph Traversal Efficiency to 198 TeraTEPS on the Supercomputer FugakuabstractBreadth-first search (BFS) is a fundamental building block of various high-performance computing applications beyond graph analysis and also known as a benchmark problem in the Graph 500 list. The increasing volume of global data demands efficient distributed BFS, which, however, is hindered by the high communication costs of exchanging vertex data between compute nodes. To address this challenge, this paper introduces four techniques: (i) forest pruning, which reduces the number of vertices by eliminating those unnecessary for the search; (ii) group reordering and (iii) multilevel bitmap compression, which decrease the memory footprint of graph data, thereby enabling fewer nodes to manage larger graphs; and (iv) adaptive parameter tuning, which quickly optimizes the hyperparameters of the BFS algorithm. In the evaluation using 152,064 nodes of the supercomputer Fugaku, our implementation achieved 198 tera-traversed edges per second, doubling the performance reported in the latest Graph500-related study on Fugaku. Junya Arai, Masahiro Nakao, Yuto Inoue, Kanto Teranishi, Koji Ueno, Keiichiro Yamamura, Mitsuhisa Sato, Katsuki Fujisawa |
SC | 2 |
| 2022 | Graph optimization algorithm using symmetry and host bias for low-latency indirect networkabstractIt is known that an indirect network with a small host-to-host Average Shortest Path Length (h-ASPL) improves overall system performance in a parallel computer system . As a means to discuss such indirect networks in graph theory, the Order/Radix Problem (ORP) has been proposed. ORP involves finding a graph with a minimum h-ASPL that satisfies a given number of hosts and radix. A graph in ORP represents an indirect network and has two types of vertices: host and switch. We propose an optimization algorithm to generate graphs with a sufficiently small h-ASPL. The primary features of the proposed algorithm are the symmetry of the graph and the bias of the hosts adjacent to each switch. These features reduce the computational time to calculate the h-ASPL and improve the search performance of the algorithm. The performance of the proposed algorithm is evaluated using problems presented by Graph Golf, an international ORP competition. Our results show that the proposed algorithm can generate graphs with a smaller h-ASPL than the existing algorithm. To evaluate the performance of the graphs generated by the proposed algorithm, we use the parallel simulation framework SimGrid and the parallel benchmark collection NAS Parallel Benchmarks. Our results also show that the graphs generated by the proposed algorithm have higher performance than those generated by the existing algorithm. Masahiro Nakao, Masaki Tsukamoto, Yoshiko Hanada, Keiji Yamamoto |
Parallel Comput. | 1 |
| 2021 | Graph optimization algorithm for low-latency interconnection networksabstractIn various industrial products including parallel computing systems, it is expected that the performance of the whole system will be improved by applying a network topology with smaller diameter and average shortest path length (ASPL). Such network topologies can be defined as the order/degree problem in graph theory by modeling a network topology as an undirected graph. Although previous research indicates that the random graph has both small diameter and ASPL due to the small-world effect, there is still room for improvement. In this paper, we propose an algorithm for the order/degree problem that optimizes random graphs. The feature of the proposed algorithm is that by giving symmetry to the graph, the solution search performance is improved and the calculation time for obtaining the diameter and APSL is greatly reduced. The proposed algorithm is evaluated using various graphs, including a huge graph with one million vertices presented by Graph Golf, an international competition for the order/degree problem. The results show that the proposed algorithm can generate a network topology with sufficiently small diameter and APSL. In addition, we also simulate the latency, performance of the parallel benchmarks, and bisection on the generated network topologies. The results show that the generated network topology performs better than the random topology and a conventional k-ary n-cube topology. Masahiro Nakao, Maaki Sakai, Yoshiko Hanada, Hitoshi Murai, Mitsuhisa Sato |
Parallel Comput. | 1 |
| 2020 | Performance Evaluation of Supercomputer Fugaku using Breadth-First Search Benchmark in Graph500abstractThere is increasing demand for the high-speed processing of large-scale graphs in various fields. However, such graph processing requires irregular calculations, making it difficult to scale performance on large-scale distributed memory systems. Against this background, Graph500, a competition for evaluating the performance of large-scale graph processing, has been held. We developed breadth-first search (BFS), which is one of the benchmark kernels used in Graph500, and took the top spot a total of 10 times using the K computer. In this paper, we tune BFS performance and evaluate it using the supercomputer Fugaku, which is the successor to the K computer. The results of evaluating BFS for a large-scale graph composed of about 1.1 trillion vertices and 17.6 trillion edges using 92,160 nodes of Fugaku indicate that Fugaku has 2.27 times the performance of the K computer. Fugaku took the top spot on Graph500 in June 2020. Masahiro Nakao, Koji Ueno, Katsuki Fujisawa, Yuetsu Kodama, Mitsuhisa Sato |
CLUSTER | 1 |
| 2020 | Parallelization of All-Pairs-Shortest-Path Algorithms in Unweighted GraphabstractThe design of the network topology of a large-scale parallel computer system can be represented as an order/degree problem in graph theory. To solve the order/degree problem, it is necessary to obtain an all-pairs-shortest-path (APSP) of the graph. Thus, this paper evaluates two parallel algorithms that quickly find the APSP in unweighted graphs and compares their performance. The first APSP algorithm is based on the breadth-first search (BFS-APSP) and the second is based on the adjacency matrix (ADJ-APSP). First, we develop serial algorithms and threaded algorithms using OpenMP, and show that ADJ-APSP is up to 32.34 times faster than BFS-APSP. Next, we develop hybrid-parallel algorithms using OpenMP and MPI, and show that BFS-APSP is faster than ADJ-APSP under certain conditions because the maximum number of processes in BFS-APSP is greater than in ADJ-APSP. In addition, we parallelize ADJ-APSP using a single GPU (NVIDIA Tesla V100) and achieve a speed increase of up to 16.53-fold compared to that of a single CPU. Finally, we evaluate the performance of the algorithms using 128 GPUs and achieve a computation time 101.10 times faster than that using a single GPU. Moreover, it is shown that the calculation time of both algorithms can be greatly reduced when the input graphs are symmetric. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 1 |
| 2020 | InKS: a programming model to decouple algorithm from optimization in HPC codes
Ksander Ejjaaouani, Olivier Aumage, Julien Bigot, Michel Mehrenberger, Hitoshi Murai, Masahiro Nakao, Mitsuhisa Sato |
J. Supercomput. | 6 |
| 2019 | Multi-accelerator extension in OpenMP based on PGAS modelabstractMany systems used in HPC field have multiple accelerators on a single compute node. However, programming for multiple accelerators is more difficult than that for a single accelerator. Therefore, in this paper, we propose an OpenMP extension that allows easy programming for multiple accelerators. We extend existing OpenMP syntax to create Partitioned Global Address Space (PGAS) on separated memories of several accelerators. The feature enables users to perform programming to use multiple accelerators in ease. In performance evaluation, we implement the STREAM Triad and the HIMENO benchmarks using the proposed OpenMP extension. As a result of evaluating the performance on a compute node equipped with up to four GPUs, we confirm that the proposed OpenMP extension demonstrates sufficient performance. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 1 |
| 2019 | A Method for Order/Degree Problem Based on Graph Symmetry and Simulated Annealing with MPI/OpenMP ParallelizationabstractThe network topology in various systems, such as large-scale data centers, high-performance computing systems, and Network on Chip, is strongly related to network latency. Designing a network topology with low latency can be defined as an order/degree problem (ODP) in graph theory by modeling the network topology as an undirected graph. This study proposes a method for efficiently solving ODPs based on graph symmetry and simulated annealing (SA). This method makes the network topology symmetrical, thereby improving the solution search performance of SA and drastically reducing the calculation time. The proposed method is applied to several problems from an international competition for ODPs called Graph Golf to find network topologies with sufficiently low latency. The symmetry-based calculation achieves a speed up of 31.76 times for one of the problems. Furthermore, to reduce calculation time, the proposed method is extended to use hybrid parallelization with MPI and OpenMP. As a result, a maximum speed up of 209.80 times was achieved on 20 compute nodes consisting of 400 CPU cores. Even faster performance was achieved by combining the symmetry-based calculation and hybrid parallelization. Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 1 |
| 2018 | A Source-to-Source Translation of Coarray Fortran with MPI for High PerformanceabstractCoarray Fortran (CAF) is a partitioned global address space (PGAS) language that is a part of standard Fortran 2008. We have implemented it as a source-to-source translator as a part of the Omni XcalebleMP compiler. Since the output is written in Fortran standard, the translator must utilize Fortran conventions such as the assumed-shape array and generic function in order to reduce both development costs and runtime overhead. The runtime library uses either GASNet, MPI-3, or Fujitsu's low-level Remote Direct Memory Access (RDMA) interface (FJ-RDMA) for one-sided communication. Hidetoshi Iwashita, Masahiro Nakao, Hitoshi Murai, Mitsuhisa Sato |
HPC Asia | 2 |
| 2017 | Implementation and Evaluation of One-sided PGAS Communication in XcalableACC for Accelerated ClustersabstractClusters equipped with accelerators such as graphics processing unit (GPU) and Many Integrated Core (MIC) are widely used. For such clusters, programmers write programs for their applications by combining MPI with one of the available accelerator programming models. In particular, OpenACC enables programmers to develop their applications easily, but with lower productivity owing to complex MPI programming. XcalableACC (XACC) is a new programming model, which is an "orthogonal" integration of a partitioned global address space (PGAS) language XcalableMP (XMP) and OpenACC. While XMP enables distributed-memory programming on both global-view and local-view models, OpenACC allows operations to be offloaded to a set of accelerators. In the local-view model, programmers can describe communication with the coarray features adopted from Fortran 2008, and we extend them to communication between accelerators. We have designed and implemented an XACC compiler for NVIDIA GPU and evaluated its performance and productivity by using two benchmarks, Himeno benchmark and NAS Parallel Benchmarks CG (NPB-CG). The performance of the XACC version with the Himeno benchmark and NPB-CG are over 85% and 97% in the local-view model against the MPI+OpenACC version, respectively. Moreover, using non-blocking communication makes the performance of local-view version over 89% with the Himeno benchmark. From the viewpoint of productivity, the local-view model provides an intuitive form of array assignment statement for communication. Akihiro Tabuchi, Masahiro Nakao, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
CCGrid | 2 |
| 2017 | Implementing Lattice QCD Application with XcalableACC Language on Accelerated ClusterabstractAccelerated clusters, which are distributed memory systems equipped with accelerators, have been used in various fields. For accelerated clusters, programmers often implement their applications by a combination of MPI and CUDA (MPI+CUDA). However, the approach faces programming complexity issues. This paper introduces the XcalableACC (XACC) language, which is a hybrid model of XcalableMP (XMP) and OpenACC. While XMP is a directive-based language for distributed memory systems, OpenACC is also a directive-based language for accelerators. XACC enables programmers to develop applications on accelerated clusters with ease. To evaluate XACC performance and productivity levels, we implemented a lattice quantum chromodynamics (Lattice QCD) application using XACC on 64 compute nodes and 256 GPUs and found its performance was almost the same as that of MPI+CUDA. Moreover, we found that XACC requires much less change from the serial Lattice QCD code than MPI+CUDA to implement the parallel Lattice QCD code. Masahiro Nakao, Hitoshi Murai, Hidetoshi Iwashita, Akihiro Tabuchi, Taisuke Boku, Mitsuhisa Sato |
CLUSTER | 1 |
| 2015 | Hybrid Communication with TCA and InfiniBand on a Parallel Programming Language XcalableACC for GPU ClustersabstractFor the execution of parallel HPC applications on GPU-ready clusters, high communication latency between GPUs over nodes will be a serious problem on strong scalability. To reduce the communication latency between GPUs, we proposed the Tightly Coupled Accelerator (TCA) architecture and developed the PEACH2 board as a proof-of-concept interconnection system for TCA. Although PEACH2 provides very low communication latency, there are some hardware limitations due to its implementation depending on PCIe technology, such as the practical number of nodes in a system which is 16 currently named sub-cluster. More number of nodes should be connected by conventional interconnections such as InfiniBand, and the entire network system is configured as a hybrid one with global conventional network and local high-speed network by PEACH2. For ease of user programmability, it is desirable to operate such a complicated communication system at the library or language level (which hides the system). In this paper, we develop a hybrid interconnection network system combining PEACH2 and InfiniBand, and implement it based on a high-level PGAS language for accelerated clusters named XcalableACC (XACC). A preliminary performance evaluation confirms that the hybrid network improves the performance based on the Himeno benchmark for stencil computation by up to 40%, relative to MVAPICH2 with GDR on InfiniBand. Additionally, Allgather collective communication with a hybrid network improves the performance by up to 50% for networks of 8 to 16 nodes. The combination of local communication, supported by the low latency of PEACH2 and global communication supported by the high bandwidth and scalability of InfiniBand, results in an improvement of overall performance. Tetsuya Odajima, Taisuke Boku, Toshihiro Hanawa, Hitoshi Murai, Masahiro Nakao, Akihiro Tabuchi, Mitsuhisa Sato |
CLUSTER | 5 |
| 2012 | Productivity and Performance of Global-View Programming with XcalableMP PGAS LanguageabstractXcalableMP (XMP) is a PGAS parallel language with a directive-based extension of C and Fortran. While it sup- ports “coarray” as a local-view programming model, an XMP global-view programming model is useful when parallelizing data-parallel programs by adding directives with minimum code modification. This paper considers the productivity and performance of the XMP global-view programming model. In the global-view programming model, a programmer describes data distributions and work-mapping to map the computations to nodes, where the computed data are located. Global-view communication directives are used to move a part of the distributed data globally and to maintain consistency in the shadow area. Rich sets of XMP global-view programming model can reduce the cost for parallelization significantly, and optimization of “privatization” is not necessary. For productivity and performance study, the Omni XMP compiler and the Berkeley Unified Parallel C compiler are used. Experimental results show that XMP can implement the benchmarks with a smaller programming cost than UPC. Furthermore, XMP has higher access performance for global data, which has an affinity with own process than UPC. In addition, the XMP coarray function can effectively tune the application's performance. Masahiro Nakao, Jinpil Lee, Taisuke Boku, Mitsuhisa Sato |
CCGRID | 1 |
| 2008 | Impact of topology and link aggregation on a PC cluster with EthernetabstractIn addition to its use in local area networks, Ethernet has been used for connecting hosts in the area of high-performance computing. Here, we investigated the impact of topology and link aggregation on a large-scale PC cluster with Ethernet. Ethernet topology that allows loops and its routing can be implemented by the VLAN routing method without creating broadcast storms. To simplify the system configuration without modifying system software, the VLAN tag is added to a frame at switches in our implementation of topologies. Each host creates VLAN interfaces that have different local network addresses on a physical interface, so that a switch learns the MAC addresses of hosts in a PC cluster by broadcast. Evaluation results showed that the performance characteristics of an eight-switch network are comparable to those of an ideal 1-switch (full crossbar) network in the execution of High-Performance LINPACK Benchmark (HPL) on a 225-host PC cluster. On the other hand, evaluation results using NAS Parallel Benchmarks indicated that topologies achieved by the proposed methodology showed performance improvements of up to about 650% as compared to the simple tree topology. These results indicate that topology and link aggregation have marked impacts and commodity switches can be used instead of expensive and high functional switches. Takafumi Watanabe, Masahiro Nakao, Tomoyuki Hiroyasu, Tomohiro Otsuka, Michihiro Koibuchi |
CLUSTER | 2 |