EDBT 2026 Demo / reviewers in the wild / expert
Taisuke Boku
dblp:26/1323
· DBLP profile ↗
70ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0001-8730-2228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 7 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating General Relativistic Radiation Magnetohydrodynamic Simulations with GPUs
Ryohei Kobayashi 0001, Hiroyuki R. Takahashi, Akira Nukada, Yuta Asahina, Taisuke Boku, Ken Ohsuga |
HPC Asia | 5 |
| 2024 | CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous SystemsabstractPerformance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems. Norihisa Fujita, Beau Johnston, Narasinga Rao Miniskar, Ryohei Kobayashi 0001, Mohammad Alaul Haque Monil, Keita Teranishi, Seyong Lee, Jeffrey S. Vetter, Taisuke Boku |
e-Science | 9 |
| 2024 | Improving Performance on Replica-Exchange Molecular Dynamics Simulations by Optimizing GPU Core UtilizationabstractWhile GPUs are the main players of the accelerating devices on high performance computing systems, their performance depends on how to utilize a numerous number of cores in parallel on each device. Typically, a loop structure with a number of iterations is assigned to a device to utilize their cores to map calculations in iterations so that there must be enough count of iterations to fill the thousands of GPU cores in the high-end GPUs. Taisuke Boku, Masatake Sugita, Ryohei Kobayashi 0001, Shinnosuke Furuya, Takuya Fujie, Masahito Ohue, Yutaka Akiyama |
ICPP | 1 |
| 2024 | Design and performance evaluation of UCX for the Tofu Interconnect D on Fugaku towards efficient multithreaded communicationabstractAbstract The increasing trend of manycore processors makes multithreaded communication more important to avoid costly global synchronization among cores. One of the representative approaches that require multithreaded communication is the global task-based programming model. In the model, a program is divided into tasks, and tasks are asynchronously executed by each node, and independent thread-to-thread communications are expected. However, the Message passing interface (MPI) based approach is not efficient because of design issues. In this research, we design and implement the utofu transport layer in an abstracted communication library called Unified communication-X (UCX) for efficient remote direct memory access (RDMA) based multithreaded communication on Tofu Interconnect D. The evaluation results on Fugaku show that UCX can significantly improve the multithreaded performance over MPI, while maintaining portability between systems thanks to UCX. UCX shows about 32.8 times lower latency than Fujitsu MPI with 24 threads in the multithreaded pingpong benchmark and about 37.8 times higher update rate than Fujitsu MPI with 24 threads on 256 nodes in multithreaded GUPs benchmark. Yutaka Watanabe, Miwako Tsuji, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
J. Supercomput. | 4 |
| 2024 | Correction: Design and performance evaluation of UCX for the Tofu Interconnect D on Fugaku towards efficient multithreaded communication
Yutaka Watanabe, Miwako Tsuji, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
J. Supercomput. | 4 |
| 2023 | GPU-FPGA-accelerated Radiative Transfer Simulation with Inter-FPGA CommunicationabstractThe complementary use of graphics processing units (GPUs) and field programmable gate arrays (FPGAs) is a major topic of interest in the high-performance computing (HPC) field. GPU–FPGA-accelerated computing is an effective tool for multiphysics simulations, which encompass multiple physical models and simultaneous physical phenomena. Because the constituent operations in multiphysics simulations exhibit varying characteristics, accelerating these operations solely using GPUs is often challenging. Hence, FPGAs are frequently implemented for this purpose. The objective of the present study was to further improve application performance by employing both GPUs and FPGAs in a complementary manner. Recently, this approach has been applied to the radiative transfer simulation code for astrophysics known as ARGOT, with evaluation results quantitatively demonstrating the resulting improvement in performance. However, the evaluation results in question came from the use of a single node equipped with both a GPU and FPGA. In this study, we extended the GPU–FPGA-accelerated ARGOT code to operate on multiple nodes using the message passing interface (MPI) and an FPGA-to-FPGA communication technology scheme called Communication Integrated Reconfigurable CompUting System (CIRCUS). We evaluated the performance of the ARGOT code with multiple GPUs and FPGAs under weak scaling conditions, and found it to achieve up to 12.8x speedup compared to the GPU-only execution. Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura |
HPC Asia | 4 |
| 2023 | A Scalable Many-core Overlay Architecture on an HBM2-enabled Multi-Die FPGAabstractThe overlay architecture enables to raise the abstraction level of hardware design and enhances hardware-accelerated applications’ portability. In FPGAs, there is a growing awareness of the overlay structure as typified by many-core architecture. It works in theory; however, it is difficult in practice, because it is beset with serious design issues. For example, the size of FPGAs is bigger than before. It is exacerbating the issue of the place-and-route. Besides, a single FPGA is actually the sum of small-to-middle FPGAs by advancing packaging technology like silicon interposers. Thus, the tightly coupled many-core designs will face this covert issue that the wires among the regions are extremely restricted. This article proposes efficient essential processing elements, micro-architecture design, and the interconnect architecture toward a scalable many-core overlay design. In particular, our work proposes a novel compact buffering technique to reduce memory resource utilization in tightly connected overlays while preserving computational efficiency. This technique reduces the utilization of BlockRAM to nearly 50% while achieving a best-case computational efficiency of 91.93% in a three-dimensional Jacobi benchmark. Besides, the proposed enhancements led to around 2× and 3× improvement in performance and power efficiency, respectively. Moreover, the improved scalability allowed increasing compute resources and delivering around 4× better performance and power efficiency, as compared to the baseline Dynamically Re-programmable Architecture of Gather-scatter Overlay Nodes overlay. Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | An FPGA-based Accelerator for Regular Path Queries over Edge-labeled GraphsabstractEdge-labeled directed graphs are commonly used to represent various information in different applications, such as social networks, knowledge graphs, etc., and regular path queries (RPQs) allow us to extract pairs of nodes that are reachable from one to another through a labeled path matching with the query pattern represented as a regular expression. It is useful for us to extract complicated or semantically meaningful information from a graph, but it gives rise to a challenge when dealing with large graphs. This is due to the long execution time caused by the explosive growth of intermediate results, but, on the other hand, some applications require fast query executions. To address this problem, we propose an FPGA-based RPQ accelerator. The idea is to exploit FPGA’s parallelism in traversing the target graph and matching the regular path expression in parallel with the pipeline manner. To validate the performance of the proposed method, we conducted a set of experiments. From the results, we observed that the proposed method achieves shorter elapsed times for RPQs against social graphs extracted from the real world, up to three orders of magnitude compared with baseline methods. Kento Miura, Ryohei Kobayashi 0001, Toshiyuki Amagasa, Hiroyuki Kitagawa, Norihisa Fujita, Taisuke Boku |
IEEE Big Data | 6 |
| 2022 | Multi-hetero Acceleration by GPU and FPGA for Astrophysics Simulation on oneAPI EnvironmentabstractGPU (Graphics Processing Unit) computing is one of the most popular accelerating methods for various high-performance computing applications. For scientific computations based on multi-physical phenomena, however, a single device solution on a GPU is insufficient, where the single timescale or degree of parallelism is not simply supported by a simple GPU-only solution. We have been researching a combination of a GPU and FPGA (Field Programmable Gate Array) for such complex physical simulations. The most challenging issue is how to program these multiple devices using a single code. Ryuta Kashino, Ryohei Kobayashi 0001, Norihisa Fujita, Taisuke Boku |
HPC Asia | 4 |
| 2022 | Accelerating Radiative Transfer Simulation on NVIDIA GPUs with OpenACC
Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura |
PDCAT | 4 |
| 2021 | HBM2 Memory System for HPC Applications on an FPGAabstractField Programmable Gate Arrays (FPGAs) have been targeted as a new accelerator of the HPC field. This is because the barrier to using FPGAs has been gradually lowered due to the widespread use of high-level synthesis (HLS) technology. In addition, the bandwidth of external memory in FPGAs is much lower than that of other accelerators widely used in HPC, such as NVIDIA V100 GPUs. However, the latest FPGAs can use High Bandwidth Memory 2 (HBM2), which has a memory bandwidth of up to 512GB/s. Therefore, we believe FPGAs will be a viable option for speeding up applications. However, unlike CPUs and GPUs, FPGAs do not have caches and memory networks to exploit the full potential of HBM2, which may limit the efficiency of the application. In this paper, we propose a memory system for HBM2 and HPC applications. We show the prototype implementation of the system and evaluate its performance. We also demonstrate the use of the proposed system from an application developed in High-Level Synthesis (HLS) written in C++. Norihisa Fujita, Ryohei Kobayashi 0001, Yoshiki Yamaguchi, Taisuke Boku |
CLUSTER | 4 |
| 2021 | An FPGA-based storage control with load balancingabstractIn the last decade, the number of cloud computing companies adopting FPGAs for performance gain has increased considerably. Indeed, the adoption of FPGA has contributed to the improvement of network performance in HPC (High-Performance Computing) systems. Nevertheless, data storage performance has not followed the trend and remained one hard-to-solve bottleneck lowering the processing capacity of a whole system, in particular, systems with high demands for real-time computations. This paper proposes a high-speed, large-capacity, and low-latency storage system that makes a breakthrough in storage performance by using inherent FPGA parallelism to control multiple SATA devices. It also has a load balancing function that can inhibit the slowest among multiple connected storage devices from hindering the throughput of the entire storage system. To verify the proposed approach, an FPGA board was custom manufactured, which embeds a Xilinx Kintex Ultrascale FPGA. It can accept connections from up to 16 SATA devices simultaneously. The SATA controller on an FPGA was almost developed from scratch and written by Verilog HDL. It enables the system to achieve low-latency processing. In our experimental results, it shows 16 SATA devices (SAMSUNG EVO 860 SSDs) work simultaneously and adequately. Besides, the load-balancing function was evaluated on the board. Naoya Umezu, Yoshiki Yamaguchi, Taisuke Boku |
CLUSTER | 3 |
| 2021 | An efficient RTL buffering scheme for an FPGA-accelerated simulation of diffuse radiative transferabstractThis paper proposes the efficient buffering approach for implementing radiative transfer equations to bridge the performance gap between processing elements and HBM memory bandwidth. The radiation transfer equation originally focuses on the fundamental physics process in astrophysics. Besides, it has become the focus of a lot of attention in recent years because of the wealth of applications such as medical bioimaging. However, the acceleration requires a complicated memory access pattern with low latency, and the earlier studies unveil conventional memory access based on software control has no aptitude for this computation. Thus, this article introduced an HBM FPGA and proposed an application-specific buffering mechanism called PRISM (PRefetchable and Instantly accessible Scratchpad Memory) to efficiently bridge the computational unit and the HBM. The proposed approach was evaluated on a XILINX Alveo U280 FPGA, and the experimental results are also discussed. Kazuki Furukawa, Ryohei Kobayashi 0001, Tomoya Yokono, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Masayuki Umemura |
FPT | 6 |
| 2021 | Performance Evaluation of OpenCL-Enabled Inter-FPGA Optical Link Communication Framework CIRCUS and SMIabstractIn recent years, Field Programmable Gate Array (FPGAs) have attracted much attention as accelerators in the research area of HighPerformance Computing (HPC). One of the strong features of current FPGA devices is their ability to achieve high-bandwidth communication performance with direct optical links to construct multi-FPGA platforms as well as their adjustability. However, FPGA programming is not easily performed on user applications. By more user-friendly programming environments, FPGAs can be applied to various HPC applications on multi-FPGA platforms. Ryuta Kashino, Ryohei Kobayashi 0001, Norihisa Fujita, Taisuke Boku |
HPC Asia | 4 |
| 2021 | High Resolution of City-Level Climate Simulation by GPU with Multi-physical Phenomena
Koei Watanabe, Kohei Kikuchi, Taisuke Boku, Takuto Sato, Hiroyuki Kusaka |
NPC | 3 |
| 2020 | Condensing an overload of parallel computing ingredients into a single architecture recipeabstractGeneral-purpose processors offer the best programming flexibility to address a wide range of problems. Nonetheless, they still lack behind special-purpose processors when it comes to sustained computational performance. Here, we leverage the best from both worlds and we propose a flexible, highly scalable, high-performance computing architecture with versatility in mind. The proposed architecture code-named DRAGON, benefits from several forms of parallelism such as SIMD, VLIW, Memory Broadcasting and even vector processing. Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku |
ASAP | 3 |
| 2020 | Accelerating Radiative Transfer Simulation with GPU-FPGA Cooperative ComputationabstractField-programmable gate arrays (FPGAs) have garnered significant interest in research on high-performance computing. This is ascribed to the drastic improvement in their computational and communication capabilities in recent years owing to advances in semiconductor integration technologies that rely on Moore’s Law. In addition to these performance improvements, toolchains for the development of FPGAs in OpenCL have been offered by FPGA vendors to reduce the programming effort required. These improvements suggest the possibility of implementing the concept of enabling on-the-fly offloading computation at which CPUs/GPUs perform poorly relative to FPGAs while performing low-latency data transfers. We consider this concept to be of key importance to improve the performance of heterogeneous supercomputers that employ accelerators such as a GPU. In this study, we propose GPU–FPGA-accelerated simulation based on this concept and demonstrate the implementation of the proposed method with CUDA and OpenCL mixed programming. The experimental results showed that our proposed method can increase the performance by up to $17.4 \times$ compared with GPU-based implementation. This performance is still $1.32 \times$ higher even when solving problems with the largest size, which is the fastest problem size for GPU-based implementation. We consider the realization of GPU–FPGA-accelerated simulation to be the most significant difference between our work and previous studies. Ryohei Kobayashi 0001, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku, Kohji Yoshikawa, Makito Abe, Masayuki Umemura |
ASAP | 4 |
| 2020 | Parallelized GPU Code of City-Level Large Eddy SimulationabstractIn this paper, we describe the GPU implementation of our City-LES code, which is developed at the Center for Computational Sciences (CCS), University of Tsukuba for detailed large eddy simulations, including surface conditions such as buildings, surface materials, and sunlight effect Wefocus on the 1) performance comparison between CUDA and OpenACC, and 2) how to reduce the data exchange between CPU and GPU memories. Using a number of GPU devices of NVIDIA Tesla V 100, we found that the current OpenACC compiler by PGI can achieve a comparable performance with CUDA in the main part of the LES calculation. We also apply OpenACC aggressively even for performances that are lower than that of a CPU to avoid data copying between the GPU and CPU, encapsulating all the data only on the GPU memory. In our optimized OpenACC (partially in CUDA) code, the results show that the performance of the full GPU version of code is doubled, and most of the GPU-CPU data copying is removed from the original GPU code. For the scaling performance test, a full GPU version achieves a 4. 7 x to 10x performance of the CPU version; this is done on a GPU cluster Cygnus at CCS, where each node is equipped with two Intel Xeon CPUs and four NVIDIA Tesla V100 GPUs, with strong scaling up to 32 nodes with 128 GPUs. For weak scaling, the full GPU version achieves a performance of more than 9x that of the CPU version for up to 32 nodes with 128 GPUs of parallel execution. Daisuke Tsuji, Taisuke Boku, Ryosaku Ikeda, Takuto Sato, Hiroto Tadano, Hiroyuki Kusaka |
ISPDC | 2 |
| 2019 | MITRACA: Manycore Interlinked Torus Reconfigurable Accelerator ArchitectureabstractBig data, Artificial Intelligence, and cloud services are emerging technologies whose power consumption due to the tremendous amount of computing resources became a significant issue in data centers. FPGA (Field Programmable Gate Array) based accelerators may offer a convenient solution for high-performance and energy-efficient computing. However, designing these accelerators using hardware description language is a burdensome task and requires specialized skill sets. To help not only FPGA engineers but also software programmers to implement their applications quickly, an overlay architecture on an FPGA will be a good candidate. Thus, this paper proposes a coarse-grained overlay architecture with SIMD (Single Instruction Multiple Data) instructions. Riadh Ben Abdelhamid, Yoshiki Yamaguchi, Taisuke Boku |
ASAP | 3 |
| 2019 | Scalable communication performance prediction using auto-generated pseudo MPI event traceabstractFor the co-design of HPC systems and applications, it is important to study how application performance is affected by the characteristics of the future systems, not just on a computation node but also for the parallel processing including inter-node communications. Trace-driven network simulators have been widely used because of its simplicity. However, they require the trace files corresponding to the simulated system size. Therefore, if a future system is larger than a current system, we can not adopt the trace files directly; that is, it is difficult to simulate a system larger than the current system. In order to address the scaling problem in the trace-driven network simulation, we have proposed a method called SCAlable Mpi Profiler (SCAMP). The SCAMP method runs an application on a current system, obtains MPI-event trace files, copies and edits the real trace files to create a large amount of pseudo MPI-event trace files for a future system, and finally drives a network simulator by inputting the pseudo MPI-event trace files. We also implemented a pseudo MPI-event trace file generator based on the analysis of LLVM's intermediate representations. We aim to easily obtain a first-order approximation of the communication performances for various network configurations and applications. In this paper, we describe the SCAMP system design and implementation as well as several performance evaluation results. Miwako Tsuji, Taisuke Boku, Mitsuhisa Sato |
HPC Asia | 2 |
| 2018 | Performance Evaluation of Large Scale Electron Dynamics Simulation under Many-core Cluster based on Knights LandingabstractWe have been developing an advanced scientific code called "ARTED" for an electron dynamics simulation using the first-order computation of materials to be ported to various large-scale parallel systems including the "K" Computer, which was previously Japan's fastest supercomputer. In this paper, the implementation and performance evaluation of the ARTED code used in Intel's latest many-core processor, the Knights Landing (KNL) stand-alone cluster, are described based on past research on porting the code to the Knights Corner (KNC) accelerator. Our target system is Oakforest-PACS, which is currently the fastest supercomputer in Japan. For performance tuning on KNL, the largest issue is how to utilize multiple levels of parallelism, such as the instruction level (512-bit SIMD instruction), hardware thread (4 threads/core), and large number of cores. We focus on the dominant computation part of the code, where 25 points of a 3D stencil computation are required. Yuta Hirokawa, Taisuke Boku, Shunsuke A. Sato, Kazuhiro Yabana |
HPC Asia | 2 |
| 2018 | OpenCL-ready High Speed FPGA Network for Reconfigurable High Performance ComputingabstractField programmable gate arrays (FPGAs) have gained attention in high-performance computing (HPC) research because their computation and communication capabilities have dramatically improved in recent years as a result of improvements to semiconductor integration technologies that depend on Moore's Law. In addition to FPGA performance improvements, OpenCL-based FPGA development toolchains have been developed and offered by FPGA vendors, which reduces the programming effort required as compared to the past. These improvements reveal the possibilities of realizing a concept to enable on-the-fly offloading computation at which CPUs/GPUs perform poorly to FPGAs while performing low-latency data movement. We think that this concept is one of the keys to more improve the performance of modern heterogeneous supercomputers using accelerators like GPUs. In this paper, we propose high-performance inter-FPGA Ethernet communication using OpenCL and Verilog HDL mixed programming in order to demonstrate the feasibility of realizing this concept. OpenCL is used to program application algorithms and data movement control when Verilog HDL is used to implement low-level components for Ethernet communication. Experimental results using ping-pong programs showed that our proposed approach achieves a latency of 0.99 μs and as much as 4.97 GB/s between FPGAs over different nodes, thus confirming that the proposed method is effective at realizing this concept. Ryohei Kobayashi 0001, Yuma Oobata, Norihisa Fujita, Yoshiki Yamaguchi, Taisuke Boku |
HPC Asia | 5 |
| 2018 | Performance and Scalability of Lightweight Multi-kernel Based Operating SystemsabstractMulti-kernels leverage today's multi-core chips to run multiple operating system (OS) kernels, typically a Light Weight Kernel (LWK) and a Linux kernel, simultaneously. The LWK provides high performance and scalability, while the Linux kernel provides compatibility. Multi-kernels show the promise of being able to meet tomorrow's extreme-scale computing needs while providing strong isolation, yielding high performance and scalability needed by classical HPC applications. McKernel and mOS started as independent research initiatives to explore the above potential. Previous work described their design and architecture advantages. This paper deploys the two LWKs and presents results from running them on a 2,048-node system with Intel Xeon Phi processors (KNL) connected by Intel Omni-Path Fabric. We compare the performance of McKernel, mOS, and Linux. Although the two multi-kernel efforts approached the problem from different angles, the results show a median performance improvement of 9% with some applications as high as 280% validating the efficacy of the multi-kernel approach. We provide insight into the performance improvements and discuss the strengths of the two different multi-kernel approaches. Balazs Gerofi, Rolf Riesen, Masamichi Takagi, Taisuke Boku, Kengo Nakajima, Yutaka Ishikawa, Robert W. Wisniewski |
IPDPS | 4 |
| 2017 | Implementation and Evaluation of One-sided PGAS Communication in XcalableACC for Accelerated ClustersabstractClusters equipped with accelerators such as graphics processing unit (GPU) and Many Integrated Core (MIC) are widely used. For such clusters, programmers write programs for their applications by combining MPI with one of the available accelerator programming models. In particular, OpenACC enables programmers to develop their applications easily, but with lower productivity owing to complex MPI programming. XcalableACC (XACC) is a new programming model, which is an "orthogonal" integration of a partitioned global address space (PGAS) language XcalableMP (XMP) and OpenACC. While XMP enables distributed-memory programming on both global-view and local-view models, OpenACC allows operations to be offloaded to a set of accelerators. In the local-view model, programmers can describe communication with the coarray features adopted from Fortran 2008, and we extend them to communication between accelerators. We have designed and implemented an XACC compiler for NVIDIA GPU and evaluated its performance and productivity by using two benchmarks, Himeno benchmark and NAS Parallel Benchmarks CG (NPB-CG). The performance of the XACC version with the Himeno benchmark and NPB-CG are over 85% and 97% in the local-view model against the MPI+OpenACC version, respectively. Moreover, using non-blocking communication makes the performance of local-view version over 89% with the Himeno benchmark. From the viewpoint of productivity, the local-view model provides an intuitive form of array assignment statement for communication. Akihiro Tabuchi, Masahiro Nakao, Hitoshi Murai, Taisuke Boku, Mitsuhisa Sato |
CCGrid | 4 |
| 2017 | Implementing Lattice QCD Application with XcalableACC Language on Accelerated ClusterabstractAccelerated clusters, which are distributed memory systems equipped with accelerators, have been used in various fields. For accelerated clusters, programmers often implement their applications by a combination of MPI and CUDA (MPI+CUDA). However, the approach faces programming complexity issues. This paper introduces the XcalableACC (XACC) language, which is a hybrid model of XcalableMP (XMP) and OpenACC. While XMP is a directive-based language for distributed memory systems, OpenACC is also a directive-based language for accelerators. XACC enables programmers to develop applications on accelerated clusters with ease. To evaluate XACC performance and productivity levels, we implemented a lattice quantum chromodynamics (Lattice QCD) application using XACC on 64 compute nodes and 256 GPUs and found its performance was almost the same as that of MPI+CUDA. Moreover, we found that XACC requires much less change from the serial Lattice QCD code than MPI+CUDA to implement the parallel Lattice QCD code. Masahiro Nakao, Hitoshi Murai, Hidetoshi Iwashita, Akihiro Tabuchi, Taisuke Boku, Mitsuhisa Sato |
CLUSTER | 5 |
| 2016 | Hybrid-view programming of nuclear fusion simulation code in the PGAS parallel programming language XcalableMP
Keisuke Tsugane, Taisuke Boku, Hitoshi Murai, Mitsuhisa Sato, William Tang 0002, Bei Wang 0002 |
Parallel Comput. | 2 |
| 2015 | Evaluation of FFT for GPU Cluster Using Tightly Coupled Accelerators ArchitectureabstractInter-node communications between accelerators in heterogeneous clusters require extra latency because of the time required to transfer data copies between the host and accelerator. Such communication latencies inhibit the optimal performance of affected applications. To address this problem, we proposed the Tightly Coupled Accelerators (TCA) architecture and designed an interconnection router chip named PEACH2. Accelerators in the TCA architecture communicate directly via the PCIe protocol, which is the current fundamental interface for all the accelerators and the host CPU, to eliminate protocol and data copy overheads. In this paper, we apply the TCA architecture to the Fast Fourier Transform (FFT) program, which is commonly used in scientific computations. First, we implemented all-to-all communication to TCA. The all-to-all communication was then applied to FFTE, which is one of the implementations of FFT. Based on the evaluation results using the HA-PACS/TCA system, we achieved the speedup of 2.7 with TCA in comparison with that with MPI using 16 nodes on the medium size. Toshihiro Hanawa, Hisafumi Fujii, Norihisa Fujita, Tetsuya Odajima, Kazuya Matsumoto, Taisuke Boku |
CLUSTER | 6 |
| 2015 | Improving Strong-Scaling on GPU Cluster Based on Tightly Coupled Accelerators ArchitectureabstractThe Tightly Coupled Accelerators (TCA) architecture that we proposed in previous work enables direct ommunication between accelerators over nodes. In this paper, we present a proof-of-concept GPU cluster called the HA-PACS/TCA using the PEACH2 chip that we designed as an interconnection router chip based on the TCA architecture. Our system demonstrated 2.0 ?sec of latency on inter-node GPU-to-GPU communication with a PCIe Gen2 x8 by RDMA, reducing minimum latency to just 44% of the InfiniBand-QDR and MPI using GPUDirect for RDMA. Through results of Himeno benchmark tests, we demonstrated that our TCA architecture improved performance scalability with the small-sized problem by up to 61%. Toshihiro Hanawa, Hisafumi Fujii, Norihisa Fujita, Tetsuya Odajima, Kazuya Matsumoto, Yuetsu Kodama, Taisuke Boku |
CLUSTER | 7 |
| 2015 | Hybrid Communication with TCA and InfiniBand on a Parallel Programming Language XcalableACC for GPU ClustersabstractFor the execution of parallel HPC applications on GPU-ready clusters, high communication latency between GPUs over nodes will be a serious problem on strong scalability. To reduce the communication latency between GPUs, we proposed the Tightly Coupled Accelerator (TCA) architecture and developed the PEACH2 board as a proof-of-concept interconnection system for TCA. Although PEACH2 provides very low communication latency, there are some hardware limitations due to its implementation depending on PCIe technology, such as the practical number of nodes in a system which is 16 currently named sub-cluster. More number of nodes should be connected by conventional interconnections such as InfiniBand, and the entire network system is configured as a hybrid one with global conventional network and local high-speed network by PEACH2. For ease of user programmability, it is desirable to operate such a complicated communication system at the library or language level (which hides the system). In this paper, we develop a hybrid interconnection network system combining PEACH2 and InfiniBand, and implement it based on a high-level PGAS language for accelerated clusters named XcalableACC (XACC). A preliminary performance evaluation confirms that the hybrid network improves the performance based on the Himeno benchmark for stencil computation by up to 40%, relative to MVAPICH2 with GDR on InfiniBand. Additionally, Allgather collective communication with a hybrid network improves the performance by up to 50% for networks of 8 to 16 nodes. The combination of local communication, supported by the low latency of PEACH2 and global communication supported by the high bandwidth and scalability of InfiniBand, results in an improvement of overall performance. Tetsuya Odajima, Taisuke Boku, Toshihiro Hanawa, Hitoshi Murai, Masahiro Nakao, Akihiro Tabuchi, Mitsuhisa Sato |
CLUSTER | 2 |
| 2014 | Hybrid-view programming of nuclear fusion simulation code in the PGAS parallel programming language XcalableMPabstractRecently, the Partitioned Global Address Space (PGAS) parallel programming model has emerged as a usable distributed memory programming model. XcalableMP (XMP) is a PGAS parallel programming language that extends base languages such as C and Fortran with directives in OpenMP-like style. XMP supports a global-view model that allows programmers to define global data and to map them to a set of processors, which execute the distributed global data as a single thread. In XMP, the concept of a coarray is also employed for local-view programming. In this study, we port Gyrokinetic Toroidal Code - Princeton (GTC-P), which is a three-dimensional PIC code developed at Princeton University to study the microturbulence phenomenon in magnetically confined fusion plasmas, to XMP as an example of hybrid memory model coding with the global-view and local-view programming models. In local-view programming, the coarray notation is simple and intuitive compared with Message Passing Interface (MPI) programming while the performance is comparable to that of the MPI version. Thus, because the global-view programming model is suitable for expressing the data parallelism for a field of grid space data, we implement a hybrid-view version using a global-view programming model to compute the field and a local-view programming model to compute the movement of particles. The performance is degraded by 5-25% compared with the original MPI version, but the hybrid-view version facilitates more natural data expression for static grid space data (in global-view model) and dynamic particle data (in local-view model), and it also increases the readability of the code for higher productivity. Keisuke Tsugane, Hideo Nuga, Taisuke Boku, Hitoshi Murai, Mitsuhisa Sato, William Tang 0002, Bei Wang 0002 |
ICPADS | 3 |
| 2013 | Task level pipelining with PEACH2: An FPGA switching fabric for high performance computingabstractWe demonstrate task level pipelining on multiple accelerators with PEACH2. PEACH2 is implmented on FPGA, and enables ultra low latency direct communication among multiple accelerators over computational nodes. By installing PEACH2, typical high performance computation nodes are tightly coupled. In this environment, application can be accelerated by exploiting not only data level parallelism, but also task level pipelined operation. Furthermore, we can processe multiple task on multiple accelerators in a pipelined manner. In our demonstration, application achieves 44% speed up compared to a single GPU. Takaaki Miyajima, Takuya Kuhara, Toshihiro Hanawa, Hideharu Amano, Taisuke Boku |
FPT | 5 |
| 2013 | Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage |
ICA3PP (2) | 2 |
| 2013 | Nuclear Fusion Simulation Code Optimization on GPU ClustersabstractGT5D is a nuclear fusion simulation program which aims to analyze the turbulence phenomena in tokamak plasma. In this research, we optimize it for GPU clusters with multiple GPUs on a node. Based on the profile result of GT5D on a CPU node, we decide to offload the whole of the time development part of the program to GPUs except MPI communication. We achieved 3.37 times faster performance in maximum in function level evaluation, and 2.03 times faster performance in total than the case of CPU-only execution, both in the measurement on high density GPU cluster HA-PACS where each computation node consists of four NVIDIA M2090 GPUs and two Intel Xeon E5-2670 (Sandy Bridge) to provide 16 cores in total. These performance improvements on single GPU corresponds to four CPU cores, not compared with a single CPU core. It includes 53% performance gain with overlapping the communication between MPI processes with GPU calculation. Norihisa Fujita, Hideo Nuga, Taisuke Boku, Yasuhiro Idomura |
ICPADS | 3 |
| 2012 | Productivity and Performance of Global-View Programming with XcalableMP PGAS LanguageabstractXcalableMP (XMP) is a PGAS parallel language with a directive-based extension of C and Fortran. While it sup- ports “coarray” as a local-view programming model, an XMP global-view programming model is useful when parallelizing data-parallel programs by adding directives with minimum code modification. This paper considers the productivity and performance of the XMP global-view programming model. In the global-view programming model, a programmer describes data distributions and work-mapping to map the computations to nodes, where the computed data are located. Global-view communication directives are used to move a part of the distributed data globally and to maintain consistency in the shadow area. Rich sets of XMP global-view programming model can reduce the cost for parallelization significantly, and optimization of “privatization” is not necessary. For productivity and performance study, the Omni XMP compiler and the Berkeley Unified Parallel C compiler are used. Experimental results show that XMP can implement the benchmarks with a smaller programming cost than UPC. Furthermore, XMP has higher access performance for global data, which has an affinity with own process than UPC. In addition, the XMP coarray function can effectively tune the application's performance. Masahiro Nakao, Jinpil Lee, Taisuke Boku, Mitsuhisa Sato |
CCGRID | 3 |
| 2011 | XMCAPI: Inter-core Communication Interface on Multi-chip Embedded SystemsabstractMulti-core processor technology has been applied to the processors in embedded systems as well as in ordinary PC systems. In multi-core embedded processors, however, a processor may consist of heterogeneous CPU cores that are not configured with a shared memory and do not have a communication mechanism for inter-core communication. MCAPI is a highly portable API standard for providing inter-core communication independent of the architecture heterogeneity. In this paper, we extend the current MCAPI to a multi-chip in a distributed memory configuration and propose its portable implementation, named XMCAPI, on a commodity network stack. With XMCAPI, the inter-core communication method for intra-chip cores is extended to inter-chip cores. We evaluate the XMCAPI implementation, xmcapi/ip, on a standard socket in a portable software development environment. Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato |
EUC | 3 |
| 2011 | Introduction
Wolfgang Karl, Samuel Thibault, Stanimire Tomov, Taisuke Boku |
Euro-Par (2) | 4 |
| 2011 | First-principles calculations of electron states of a silicon nanowire with 100, 000 atoms on the K computerabstractReal space DFT (RSDFT) is a simulation technique most suitable for massively-parallel architectures to perform first-principles electronic-structure calculations based on density functional theory. We here report unprecedented simulations on the electron states of silicon nanowires with up to 107,292 atoms carried out during the initial performance evaluation phase of the K computer being developed at RIKEN. Yukihiro Hasegawa, Jun-ichi Iwata, Miwako Tsuji, Daisuke Takahashi, Atsushi Oshiyama, Kazuo Minami, Taisuke Boku, Fumiyoshi Shoji, Atsuya Uno, Motoyoshi Kurokawa, Hikaru Inoue, Ikuo Miyoshi, Mitsuo Yokokawa |
SC | 7 |
| 2009 | Using a cluster as a memory resource: A fast and large virtual memory on MPIabstractThe 64-bit OS provides ample memory address space that is beneficial for applications using a large amount of data. This paper proposes using a cluster as a memory resource for sequential applications requiring a large amount of memory. This system is an extension of our previously proposed socket-based distributed large memory system (DLM), which offers large virtual memory by using remote memory distributed over nodes in a cluster. The newly designed DLM is based on MPI (message passing interface) to exploit higher portability. MPI-based DLM provides fast and large virtual memory on widely available open clusters managed with an MPI batch queuing system. To access this remote memory, we rely on swap protocols adequate for MPI thread support levels. In experiments, we confirmed that it achieves 493 MB/s and 613 MB/s of remote memory bandwidth with the STREAM benchmark on 2.5 GB/s and 5 GB/s links (Myri-10G x2, x4) and high performance of applications with NPB and Himeno benchmarks. Additionally, this system enables users unfamiliar with parallel programming to use a cluster. Hiroko Midorikawa, Kazuhiro Saito, Mitsuhisa Sato, Taisuke Boku |
CLUSTER | 4 |
| 2009 | Flexible Multi-link Ethernet Binding System for PC Clusters with Asymmetric TopologyabstractIn current high-performance PC clusters, the performance and cost of interconnection network are essential issues. Very cost-effective Ethernets, such as Gigabit Ethernet, as well as high performance SANs, such as Infiniband and Myrinet, are still widely used. The authors have been developing a multi-link binding network system for Ethernet, called RI2N, for high-throughput and fault-tolerant interconnection with Gigabit Ethernet. It can be used both for internode communication in MPI programs and traditional UNIX network services such as NFS. In this paper, the authors propose an optimized version of RI2N, called RI2N+, that allows asymmetrical multi-link connection for fitting to various cost-effective system configurations. Such a configuration cannot be supported by Linux Channel Bonding, which is widely used in standard Linux distributions. RI2N+ automatically detects the asymmetric network configuration and controls the traffic distribution to multiple links. In the basic performance evaluation under a high traffic rate, it was confirmed that the throughput of the network with the proposed scheme is improved by approximately 30\% compared with the original RI2N. RI2N+ also maintains high performance even in asymmetric configurations, that is up to 86\% of the relative performance compared with the symmetric case. Taiga Yonemoto, Shin'ichi Miura, Toshihiro Hanawa, Taisuke Boku, Mitsuhisa Sato |
ICPADS | 4 |
| 2009 | RI2N/DRV: Multi-link ethernet for high-bandwidth and fault-tolerant network on PC clustersabstractAlthough recent high-end interconnection network devices and switches provide a high performance to cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, Linux kernels are equipped with software named Linux Channel Bonding (LCB), which is based IEEE802.3ad Link Aggregation technology. However, standard LCB has the disadvantage of mismatch with the TCP protocol; consequently, both large latency and bandwidth instability can occur. Fault-tolerance feature is supported by LCB, but the usability is not sufficient. We developed a new implementation similar to LCB named Redundant Interconnection with Inexpensive Network with Driver (RI2N/DRV) for use on Gigabit Ethernet. RI2N/DRV has a complete software stack that is very suitable for TCP, an upper layer protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets, even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as network file system (NFS), without any modification of other modules. Shin'ichi Miura, Toshihiro Hanawa, Taiga Yonemoto, Taisuke Boku, Mitsuhisa Sato |
IPDPS | 4 |
| 2009 | Towards an Open Dependable Operating SystemabstractThis paper introduces a new dependable operating system project, called DEOS, started in 2006, and scheduled to continue for six years. In this project, a safety extension mechanism called P-Bus is to be designed, and implemented in the Linux kernel so that a future dependability attribute is implemented with P-Bus. A hardware abstraction layer, called SPUMONE, is introduced so that a light-weight operating system, called ArcOS, and a monitoring service on top of ArcOS monitors the Linux kernel to provide a safety-net for the Linux kernel. New dependability metrics are being designed to enable developers and users to decide which hardware or software solution meets their dependability requirements, and thus can be used. Yutaka Ishikawa, Hajime Fujita 0002, Toshiyuki Maeda, Motohiko Matsuda, Midori Sugaya, Mitsuhisa Sato, Toshihiro Hanawa, Shin'ichi Miura, Taisuke Boku, Yuki Kinebuchi, Tatsuo Nakajima, Jin Nakazawa, Hideyuki Tokuda |
ISORC | 9 |
| 2008 | RI2N: High-bandwidth and fault-tolerant network with multi-link Ethernet for PC clustersabstractAlthough recent high-end interconnection network devices and switches provide a high performance/cost ratio, most of the small to medium sized PC clusters are still built on the commodity network, Ethernet. To enhance performance on commonly used Gigabit Ethernet networks, link aggregation or binding technology is used. Currently, a Linux kernel is equipped with a software solution named Linux Channel Bonding (LCB), which is based on IEEE802.3ad Link Aggregation technology. However, standard LCB has the problem of mismatching with the commonly used TCP protocol, which consequently implies several problems of both large latency and instability on bandwidth improvement. The fault-tolerant feature is also supported, but the usability is not sufficient. We have developed a new implementation similar to LCB named RI2N/DRV (Redundant Interconnection with Inexpensive Network with Driver) for use on a Gigabit Ethernet with a complete software stack that is very compatible with the TCP protocol. Our algorithm suppresses unnecessary ACK packets and retransmission of packets even in imbalanced network traffic and link failures on multiple links. It provides both high-bandwidth and fault-tolerant communication on multi-link Gigabit Ethernet. We confirmed that this system improves the performance and reliability of the network, and our system can be applied to ordinary UNIX services such as NFS, without any modification of other modules. Shin'ichi Miura, Takayuki Okamoto, Taisuke Boku, Toshihiro Hanawa, Mitsuhisa Sato |
CLUSTER | 3 |
| 2008 | A dynamic routing control system for high-performance PC cluster with multi-path Ethernet connectionabstractVLAN-based Flexible, Reliable and Expandable Commodity Network (VFREC-Net) is a network construction technology for PC clusters that allows multi-path network routing to be configured using inexpensive Layer-2 Ethernet switches based on tagged-VLAN technology. Current VFREC-Net system encounters problems with traffic balancing when the communication pattern of the application does not fit the network topology, due to its static routing scheme. Shin'ichi Miura, Taisuke Boku, Takayuki Okamoto, Toshihiro Hanawa |
IPDPS | 2 |
| 2008 | Integrating Computing Resources on Multiple Grid-Enabled Job Scheduling Systems Through a Grid RPC System
Yoshihiro Nakajima, Mitsuhisa Sato, Yoshiaki Aida, Taisuke Boku, Franck Cappello |
J. Grid Comput. | 4 |
| 2007 | RI2N/UDP: High bandwidth and fault-tolerant network for a PC-cluster based on multi-link EthernetabstractPC-clusters with high performance/cost ratio have been one of the typical platforms for high performance computing. To lower costs, Gigabit Ethernet is often used for intercommunication networks. However, the reliability of Ethernet is limited due to hardware failures and tentative errors in the network switches. To solve this problem, we propose an interconnection network system based on multi-link Ethernet named RI2N. In this paper, we developed a user level implementation of RI2N using UDP/IP that is called RI2N/UDP. When this new system was evaluated for performance and fault tolerance, the bandwidth on a 2-link Gigabit Ethernet was 246 MB/s, and the system could remain active during network link failure to provide high system reliability. Takayuki Okamoto, Shin'ichi Miura, Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi |
IPDPS | 3 |
| 2006 | PACS-CS: A Large-Scale Bandwidth-Aware PC Cluster for Scientific ComputationsabstractWe have been developing a large scale PC cluster named PACS-CS (Parallel Array Computer System for Computational Sciences) at Center for Computational Sciences, University of Tsukuba, for wide variety of computational science applications such as computational physics, computational material science, computational biology, etc. We consider the most important issue on the computation node is the memory access bandwidth, then a node is equipped with a single CPU which is different from ordinary high-end PC clusters. The interconnection network for parallel processing is configured as a multi-dimensional hyper-crossbar network based on trunking of Gigabit Ethernet to support large scale scientific computation with physical space modeling. Based on the above concept, we are developing an original mother board to configure a single CPU node with 8 ports of Gigabit Ethernet, which can be implemented in the half size of 19 inch rack-mountable 1U size platform. Under the preliminary performance evaluation, we confirmed that the computation part in practical Lattice QCD code will be able to achieve 30% of peak performance, and up to 600 Mbyte/sec of bandwidth at single directed neighboring communication will be achieved. PACS-CS will start its operation on July 2006 with 2560 CPUs and 14.3 Tflops of peak performance. Taisuke Boku, Mitsuhisa Sato, Akira Ukawa, Daisuke Takahashi, Shinji Sumimoto, Kouichi Kumon, Takashi Moriyama, Masaaki Shimizu |
CCGRID | 1 |
| 2006 | Integrating Computing Resources on Multiple Grid-enabled Job Scheduling Systems Through a Grid RPC SystemabstractWe present a framework for a parallel programming model by remote procedure calls bridging between largescale computing resource pools managed by multiple gridenabled job scheduling systems. With this system, the user can exploit not only each remote servers and clusters, but also computing resources provided with grid-enabled job scheduling systems located on different sites. This framework requires a Grid RPC system to decouple the computation in a remote node from the Grid RPC mechanism and uses document-based communication rather than connection-based communication. We implemented the proposed framework as an extension of the OmniRPC system, which is a Grid RPC system for parallel programming in a grid environment. We designed a general interface to adapt the OmniRPC system to various grid-enabled job scheduling systems easily and applied the proposed system to several grid-enabled job scheduling systems, including XtremWeb, CyberGRIP, Condor and Grid Engine. we show the preliminary performance of these implementations using a phylogenetic application. We found that the proposed system can achieve approximately the same performance as using OmniRPC and can handle interruptions in worker programs on remote nodes. Yoshihiro Nakajima, Mitsuhisa Sato, Yoshiaki Aida, Taisuke Boku, Franck Cappello |
CCGRID | 4 |
| 2006 | Emprical study on Reducing Energy of Parallel Programs using Slack Reclamation by DVFS in a Power-scalable High Performance ClusterabstractIt has become important to improve the energy efficiency of high performance PC clusters. In PC clusters, high-performance microprocessors have a dynamic voltage and frequency scaling (DVFS) mechanism, which allows the voltage and frequency to be set for reduction in energy consumption. In this paper, we proposed a new algorithm that reduces energy consumption in a parallel program executed on a power-scalable cluster using DVFS. Whenever the computational load is not balanced, parallel programs encounter slack time, that is, they must wait for synchronization of the tasks. Our algorithm reclaims slack time by changing the voltage and frequency, which allows a reduction in energy consumption without impacting on the performance of the program. Our algorithm can be applied to parallel programs represented by a directed acyclic task graph (DAG). It selects an appropriate set of voltages and frequencies (called the gear) that allow the tasks to execute at the lowest frequency that does not increase the overall execution time, but at the same time allows the tasks to be executed as uniformly as possible in frequency. We built two different types of power-scalable clusters using AMD Turion and Transmeta Crusoe. For the empirical study on energy reduction in PC clusters, we designed a toolkit called PowerWatch that includes power monitoring tools and the DVFS control library. This toolkit precisely measures the power consumption of the entire cluster in real time. The experimental results using benchmark problems show that our algorithm reduces energy consumption by 25% with only a 1 % loss in performance Hideaki Kimura 0003, Mitsuhisa Sato, Yoshihiko Hotta, Taisuke Boku, Daisuke Takahashi |
CLUSTER | 4 |
| 2006 | Performance Improvement by Data Management Layer in a Grid RPC System
Yoshiaki Aida, Yoshihiro Nakajima, Mitsuhisa Sato, Tetsuya Sakurai, Daisuke Takahashi, Taisuke Boku |
GPC | 6 |
| 2006 | A scalable communication layer for multi-dimensional hyper crossbar network using multiple gigabit ethernetabstractThis paper proposes a scalable communication layer for a multi-dimensional hyper crossbar network using multiple Gigabit Ethernet for the PACS-CS system which consists of 2560 single-processor nodes and a 16 x 16 x 10 three dimensional hyper-crossbar network (3D-HXB). To realize a high performance communication layer using multiple existing Ethernet networks, the host processor usage for the communication processing must be reduced to less than the appropriate packet processing time which is calculated from a message size and a target communication bandwidth. To overcome this problem, we have developed the PM/Ethernet-HXB communication facility. PM/Ethernet-HXB realizes communication protocol processing without exclusion even for Zero-copy communication between the communication buffers of nodes. We have implemented the PM/Ethernet-HXB on SCore cluster system software, and evaluated its communication and application performance. PM/Ethernet-HXB achieves a unidirectional communication bandwidth of 1065 MB/s using nine Gigabit Ethernet links on a single dimension network. It also realizes a unidirectional communication bandwidth of 741 MB/s (98.8% of the theoretical performance) and a bidirectional bandwidth of 1401 MB/s (93.4% of the theoretical performance) on the three dimensional connections (3D-HXB: a total of six Ethernet links). The results of MPI communication bandwidth are a unidirectional communication bandwidth of 960 MB/s and a bidirectional bandwidth of 1008 MB/s using eight links on a single dimension network. These results show that PM/Ethernet-HXB realizes a comparative performance using multiple Gigabit Ethernet networks to dedicated cluster networks such as InfiniBand 4x (1000 MB/s). The speedups of IS and CG Class C NAS parallel benchmarks are scalable up to using four links on eight node cluster, and performance degradation between 3D-HXB (2 x 2 x 2) and 1-dimensional network is small. Shinji Sumimoto, Kazuichi Oe, Kouichi Kumon, Taisuke Boku, Mitsuhisa Sato, Akira Ukawa |
ICS | 4 |
| 2006 | MegaProto/E: power-aware high-performance cluster with commodity technologyabstractIn our research project named "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we have been developing a prototype cluster not based on ASIC or FPGA but instead only using commodity technology. Its packaging is extremely compact and dense, and its performance/power ratio is very high. Our previous prototype system named "MegaProto" demonstrated that one cluster unit, which consists of 16 commodity low-power processors, can be successfully implemented on just 1U height chassis and it is capable of up to 2.8 times higher performance/power ratio than ordinary high-performance dual-Xeon 1U server units. We have improved MegaProto by replacing the CPU and enhancing the I/O performance. The new cluster unit named "MegaProto/E" with 16 Transmeta Efficeon processors achieves 32 GFlops of peak performance, which is 2.2-fold greater than that of the original one. The cluster unit is equipped with an independent dual network of Gigabit Ethernet, including dual 24-port switches. The maximum power consumption of the cluster unit is 320 W, which is comparable with that of today's high-end PC servers for high performance clusters. Performance evaluation using NPB kernels and HPL shows that the performance of MegaProto/E exceeds that of a dual-Xeon server in all the benchmarks, and its performance ratio ranges from 1.3 to 3.7. These results reveal that our solution of implementing a number of ultra low-power processors in compact packaging is an excellent way to achieve extremely high performance in applications with a certain degree of parallelism. We are now building a multi-unit cluster with 128 CPUs (8 units) to prove that this advantage still holds with higher scalability Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi, Hiroshi Nakashima, Hiroshi Nakamura, Satoshi Matsuoka, Yoshihiko Hotta |
IPDPS | 1 |
| 2006 | Profile-based optimization of power performance by using dynamic voltage scaling on a PC clusterabstractCurrently, several of the high performance processors used in a PC cluster have a DVS (dynamic voltage scaling) architecture that can dynamically scale processor voltage and frequency. Adaptive scheduling of the voltage and frequency enables us to reduce power dissipation without a performance slowdown during communication and memory access. In this paper, we propose a method of profiled-based power-performance optimization by DVS scheduling in a high-performance PC cluster. We divide the program execution into several regions and select the best gear for power efficiency. Selecting the best gear is not straightforward since the overhead of DVS transition is not free. We propose an optimization algorithm to select a gear using the execution and power profile by taking the transition overhead into account. We have built and designed a power-profiling system, PowerWatch. With this system we examined the effectiveness of our optimization algorithm on two types of power-scalable clusters (Crusoe and Turion). According to the results of benchmark tests, we achieved almost 40% reduction in terms of EDP (energy-delay product) without performance impact (less than 5%) compared to results using the standard clock frequency. Yoshihiko Hotta, Mitsuhisa Sato, Hideaki Kimura 0003, Satoshi Matsuoka, Taisuke Boku, Daisuke Takahashi |
IPDPS | 5 |
| 2006 | Storage challenge - High performance data analysis for particle physics using the Gfarm file systemabstractThe Belle experiment operates at the KEKB accelerator, a high luminosity asymmetric energy e+ e- collider. The Belle collaboration studies CP violations in decays of B mesons to answer one of the fundamental questions of Nature, the matter-anti-matter asymmetry. Currently, Belle accumulates more than one million B Bbar meson pairs, corresponding to about 1.2 TB of raw data, per day.The challenge is how to realize the required high performance data access and scalable data computing. The Gfarm file system is a Grid-wide network shared file system that federates local storage of cluster nodes; moreover it provides scalable I/O performance with distributed data access. In the challenge, we will construct a Gfarm file system with 40 TB capacity and 30 GB/sec I/O bandwidth, integrating the local disks of 800 compute nodes in the KEKB computing facility, and demonstrate high-performance Belle data analysis. Nobuhiko Katayama, Mitsuhisa Sato, Taisuke Boku, Akira Ukawa, Shohei Nishida, Ichiro Adachi, Osamu Tatebe |
SC | 3 |
| 2005 | MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity TechnologyabstractIn our research project "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we claim that a million-scale parallel system could be built with densely mounted low-power commodity processors. "MegaProto" is a proof-of-concept low-power and highperformance cluster build only with commodity components to implement this claim. A one-rack system is composed of 32 motherboard "cluster units" of 1 U-height and commodity switches to interconnect them mutually as well as with other racks. Each cluster unit houses 16 low-power dollarbill- sized commodity PC-architecture daughterboards, together with a high bandwidth, 2 Gbps per processor embedded switched network based on Gigabit Ethernet. The peak performance of a one-rack system is 0.48 TFlops for the first version and will improve to 1.02 TFlops in the second version through a processor/daughterboard upgrade. The system consumes about 10 kW or less per rack, resulting in 100 MFlops/W power efficiency with a power-aware intrarack network of 32 Gbps bisection bandwidth, while additional 2.4 kW will boost this to sufficiently large 256 Gbps. Performance studies show that even the first version significantly outperforms a conventional high-end 1U server comprised of dual power-hungry processors in a majority of NPB programs. It is also investigated how the current automated DVS control could save power for the HPC parallel programs along with its limitation. Hiroshi Nakashima, Hiroshi Nakamura, Mitsuhisa Sato, Taisuke Boku, Satoshi Matsuoka, Daisuke Takahashi, Yoshihiko Hotta |
SC | 4 |
| 2004 | Implementation and performance evaluation of CONFLEX-G: grid-enabled molecular conformational space search program with OmniRPCabstractCONFLEX-G is the grid-enabled version of a molecular conformational space search program called CONFLEX. We have implemented CONFLEX-G using a grid RPC system called OmniRPC. In this paper, we report the performance of CONFLEX-G in a grid testbed of several geographically distributed PC clusters. In order to explore many conformation of large bio-molecules, CONFLEX-G generates trial structures of the molecules and allocates jobs to optimize a trial structure with a reliable molecular mechanics method in the grid. OmniRPC provides a restricted persistence model to support the parametric search applications. In this model, when the initialization procedure is defined in the RPC module, the module is automatically initialized at the time of invocation by calling the initialization procedure. This can eliminate unnecessary communication and initialization at each call in CONFLEX-G. CONFLEX-G can achieve performance comparable to CONFLEX MPI and can exploit more computing resources by allowing the use of a cluster of multiple clusters in the grid. The experimental result shows that CONFLEX-G achieved a speedup of 56.5 times in the case of the 1BL1 molecule, where the molecule consists of a large number of atoms, and each trial structure optimization requires significant time. The load imbalance of the optimization time of the trial structure may also cause performance degradation. Yoshihiro Nakajima, Mitsuhisa Sato, Hitoshi Gotoh, Taisuke Boku, Daisuke Takahashi |
ICS | 4 |
| 2004 | Parallel Implementation of Strassen's Matrix Multiplication Algorithm for Heterogeneous ClustersabstractSummary form only given. We propose a new distribution scheme for a parallel Strassen's matrix multiplication algorithm on heterogeneous clusters. In the heterogeneous clustering environment, appropriate data distribution is the most important factor for achieving maximum overall performance. However, Strassen's algorithm reduces the total operation count to about 7/8 times per one recursion and, hence, the recursion level has an effect on the total operation count. Thus, we need to consider not only load balancing but also the recursion level in Strassen's algorithm. Our scheme achieves both load balancing and reduction of the total operation count. As a result, we achieve a speedup of nearly 21.7% compared to the conventional parallel Strassen's algorithm in a heterogeneous clustering environment. Yuhsuke Ohtaki, Daisuke Takahashi, Taisuke Boku, Mitsuhisa Sato |
IPDPS | 3 |
| 2003 | HMCS-G: Grid-enabled Hybrid Computing System for Computational AstrophysicsabstractThe authors have developed a hybrid computing system called HMCS-G, a Grid-enabled Heterogeneous Multi-Computer System, that provides a multiple cluster environment centered around a dedicated machine for gravity calculation. The purpose of HMCS-G is to provide an ideal computational environment for astrophysical study involving multiple physical phenomena. The worker cluster may comprise general-purpose PCs to perform tasks such as hydrodynamics computations, while the special-purpose machine, in this case a GRAPE-6 cluster, performs gravity calculations for all pairs of particles in the system. These systems are connected by OmniRPC, a grid-enabled RPC system that supports Globus and ssh for authentication. HMCS-G effectively provides worldwide access to a GRAPE-6 cluster, thereby securing several TFLOPS performance for intensive computations such as gravity calculation. All participating PC-clusters share this resource in a time-based manner using grid technology. The actual turn-around response time was measured for a system implemented over a number of institutions, and it was confirmed that HMCS-G provides acceptable real-world application performance. Precise simulations of galaxy formation are currently being performed on clusters in several institutes, involving smoothed particle hydrodynamics and radiative transfer in the context of complete gravity calculation as the first real application of HMCS-G. Taisuke Boku, Mitsuhisa Sato, Kenji Onuma, Junichiro Makino, Hajime Susa, Daisuke Takahashi, Masayuki Umemura, Akira Ukawa |
CCGRID | 1 |
| 2003 | OmniRPC: a Grid RPC ystem for Parallel Programming in Cluster and Grid EnvironmentabstractWe have designed and implemented a Grid RPC system called OmniRPC, for parallel programming in cluster and grid environments. While OmniRPC inherits its API from Ninf, the programmer can use OpenMP for easy-to-use parallel programming because the API is designed to be thread-safe. To support typical master-worker grid applications such as a parametric execution, OmniRPC provides an automatic-initializable remote module to send and store data to a remote executable invoked in the remote host. Since it may accept several requests for subsequent calls by keeping the connection alive, the data set by the initialization is re-used, resulting in efficient execution by reducing the amount of communication. The OmniRPC system also supports a local environment with "rsh", a grid environment with Globus, and remote hosts with "ssh". Furthermore, the user can use the same program over OmniRPC for both clusters and grids because a typical grid resource is regarded simply as a cluster of clusters distributed geographically. For a cluster over a private network, an agent process running the server host functions as a proxy to relay communications between the client and the remote executables by multiplexing the communications into one connection to the client. This feature allows a single client to use a thousand of remote computing hosts. Mitsuhisa Sato, Taisuke Boku, Daisuke Takahashi |
CCGRID | 2 |
| 2002 | A Blocking Algorithm for Parallel 1-D FFT on Clusters of PCs
Daisuke Takahashi, Taisuke Boku, Mitsuhisa Sato |
Euro-Par | 2 |
| 2002 | Heterogeneous multi-computer system: a new platform for multi-paradigm scientific simulationabstractHMCS (Heterogeneous Multi-Computer System) is a new parallel processing platform combining massively parallel processors for continuum simulation and particle simulation to realize multi-scale computational physics simulations. We are constructing a prototype system of HMCS with a general purpose scientific parallel processor CP-PACS and a gravity calculation parallel processor GRAPE-6 connecting them via commodity-base parallel network.On the prototype of HMCS, a microscopic gravity calculation on GRAPE-6 and a macroscopic radiation hydrodynamic calculation on CP-PACS are performed simultaneously to realize very detailed simulation on computational astrophysics. Both systems are connected via parallel network controlling system named PIO (Parallel I/O System).In this paper, we report the overall concept, design and implementation of HMCS, and our application result for radiative transfer smoothed particle hydrodynamics with self gravity for analysis of galaxy generation. Taisuke Boku, Masayuki Umemura, Junichiro Makino, Toshiyuki Fukushige, Hajime Susa, Akira Ukawa |
ICS | 1 |
| 2000 | SCIMA: Software Controlled Integrated Memory Architecture for High Performance ComputingabstractProcessor performance has been improved due to clock acceleration and ILP extraction techniques. Performance of main memory, however, has not been improved so much. The performance gap between processor and memory will be growing further in the future. This is very serious problem in high performance computing because effective performance is limited by memory ability in most cases. In order to overcome this problem, we propose a new VLSI architecture called SCIMA which integrates software controllable memory into a processor chip. Most of data access is regular in high performance computing. The software controllable memory is more suitable for making good use of the regularity than conventional cache. This paper presents its architecture and performance evaluation. The evaluation results reveal the superiority of SCIMA compared with conventional cache-based architecture. Masaaki Kondo, Hideki Okawara, Hiroshi Nakamura, Taisuke Boku |
ICCD | 4 |
| 1999 | Performance of lattice QCD programs on CP-PACS
Sinya Aoki, R. Burkhalter, Kazuyuki Kanaya, T. Yoshié, Taisuke Boku, Hiroshi Nakamura, Yoshiyuki Yamashita |
Parallel Comput. | 5 |
| 1999 | CP-PACS: A massively parallel processor at the University of Tsukuba
Kisaburo Nakazawa, Hiroshi Nakamura, Taisuke Boku, Ikuo Nakata, Yoshiyuki Yamashita |
Parallel Comput. | 3 |
| 1998 | Practical Simulation of Large-Scale Parallel Programs and Its Performance Analysis of the NAS Parallel Benchmarks
Kazuto Kubota, Ken'ichi Itakura, Mitsuhisa Sato, Taisuke Boku |
Euro-Par | 4 |
| 1997 | Advanced processor design using hardware description language AIDLabstractIn order to design advanced processors in a short time, designers must simulate their designs and reflect the results to the designs at the very early stages. However, conventional hardware description languages (HDLs) do not have enough ability to describe designs easily and accurately at these stages. Thus, we have proposed a new HDL called AIDL (Architecture- and Implementation-level Description Language). In this paper, in order to evaluate the effectiveness of AIDL, we describe and compare three processors in both AIDL and VHDL descriptions. Takayuki Morimoto, Kazushi Saito, Hiroshi Nakamura, Taisuke Boku, Kisaburo Nakazawa |
ASP-DAC | 4 |
| 1997 | CP-PACS: A Massively Parallel Processor for Large Scale Scientific Calculations
Taisuke Boku, Ken'ichi Itakura, Hiroshi Nakamura, Kisaburo Nakazawa |
International Conference on Supercomputing | 1 |
| 1993 | A Scalar Architecture for Pseudo Vector Processing Based on Slide-Windowed RegistersabstractIn this paper, we present a new scalar architecture for high-speed vector processing. Without using cache memory, the proposed architecture tolerates main memory access latency by introducing slide-windowed floating-point registers with data preloading feature and pipelined memory. The architecture can hold upward compatibility with existing scalar architectures. In the new architecture, software can control the window structure. This is the advantage compared with our previous work of register-windows. Because of this advantage, registers are utilized more flexibly and computational efficiency is largely enhanced. Furthermore, this flexibility helps the compiler to generate efficient object codes easily. Hiroshi Nakamura, Taisuke Boku, Hideo Wada, Hiromitsu Imori, Ikuo Nakata, Yasuhiro Inagami, Kisaburo Nakazawa, Yoshiyuki Yamashita |
International Conference on Supercomputing | 2 |
| 1990 | (SM)²-II: A Large-Scale Multiprocessor for Sparse Matrix Calculationsabstract(SM)/sup 2/-II is a large-scale parallel machine dedicated to scientific computation which includes sparse matrix calculations. In order to connect thousands of microprocessors and utilize a high degree of parallelism, the whole (SM)/sup 2/-II system is designed based on a simple computational model called the node and connecting-line (NC) model. The concept and the architecture of (SM)/sup 2/-II are described. The NC-model and a language called node oriented concurrent C (NCC) are derived. The concurrent process controller is briefly introduced. Receiver selectable multicast (RSM) is proposed, and the structure which allows connection of a large number of processing units is described. The performance of the RSM is analyzed. Some connection structures for clusters are evaluated. An operational prototype is introduced.> Hideharu Amano, Taisuke Boku, Tomohiro Kudoh |
IEEE Trans. Computers | 2 |
| 1988 | IMPULSE: A High Performance Processing Unit for Multiprocessors for Scientific CalculationabstractA high-performance processing unit for multiprocessor systems for scientific calculations, called Impulse, is described. Impulse is equipped with a hardware process-control mechanism, and a powerful floating-point processor and its controller. The process-control method is based on the concurrent process model called the NC model. In the NC model, the processes and their communicating channels are static, and it is relatively easy to implement the interprocess communication server and process scheduler in hardware according to this model. To enhance the system performance, Impulse is composed of three parts, the task engine, IPC engine, and FFP engine. From the results of simulations, it appears that if IPC engine provides efficient process control even if the granularity of the processes is very fine.> Taisuke Boku, Shigehiro Nomura, Hideharu Amano |
ISCA | 1 |
| 1985 | (SM)²-II: A New Version of the Sparse Matrix Solving Machineabstractarticle Free Access Share on (SM)2-II: a new version of the sparse matrix solving machine Authors: Hideharu Amano Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Taisuke Boku Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Tomohiro Kudoh Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile , Hideo Aiso Department of Electrical Engineering, Keio University, Yokohoma 223 Japan Department of Electrical Engineering, Keio University, Yokohoma 223 JapanView Profile Authors Info & Claims ACM SIGARCH Computer Architecture NewsVolume 13Issue 3June 1985 pp 100–107https://doi.org/10.1145/327070.327137Published:01 June 1985Publication History 16citation208DownloadsMetricsTotal Citations16Total Downloads208Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Hideharu Amano, Taisuke Boku, Tomohiro Kudoh, Hideo Aiso |
ISCA | 2 |