Yuetsu Kodama

dblp:11/3575 · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
5since 2021 · last 2027
0000-0001-5787-0363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2027 Closed-loop calculations of electronic structure on a quantum processor and a classical supercomputer at full scale
abstract
Quantum computers must operate in concert with classical computers to deliver on the promise of quantum advantage for practical problems. To achieve that, it is important to understand how quantum and classical computing can interact together, and how one can characterize the scalability and efficiency of hybrid quantum–classical workflows. So far, early experiments with quantum-centric supercomputing workflows have been limited in scale and complexity. Here, we use a Heron quantum processor deployed on premises with the entire supercomputer Fugaku to perform the largest computation of electronic structure involving quantum and classical high-performance computing. We design a closed-loop workflow between the quantum processors and 152,064 classical nodes of Fugaku, to approximate the electronic structure of chemistry models beyond the reach of exact diagonalization, with accuracy comparable to some all-classical approximation methods. Our work pushes the limits of the integration of quantum and classical high-performance computing, showcasing computational resource orchestration at the largest scale possible for current classical supercomputers.
Tomonori Shirakawa, Javier Robledo Moreno, Toshinari Itoko, Vinay Tripathi, Kento Ueda, Yukio Kawashima, Lukas Broers, William M. Kirby, Himadri Pathak, Hanhee Paik, Miwako Tsuji, Yuetsu Kodama, Mitsuhisa Sato, Constantinos Evangelinos, Seetharami Seelam, Robert Walkup, Seiji Yunoki, Mario Motta, Petar Jurcevic, Hiroshi Horii, Antonio Mezzacapo
Future Gener. Comput. Syst.12
2024 MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs
abstract
Modern High-Performance Computing (HPC) systems play a fundamental role in driving scientific research, as they execute computationally intensive jobs originating from diverse domains. However, HPC jobs are characterized by conflicting computational requirements, which may cause inefficiencies in resource usage, system throughput and energy consumption. One approach to tackling this problem is to distinguish between memory-bound and compute-bound jobs at their submission time, with the goal of making informed decisions about their execution. In this paper, we present MCBound, the first online data-driven framework to classify HPC jobs as memory/compute-bound before job execution, without user intervention. We propose a systematic characterization technique to generate a reference dataset from historical data for initial classification model training. Using the proposed characterization technique, we analyze the data of 2.2 million job runs on the Supercomputer Fugaku1, a production HPC system installed at the RIKEN Center for Computational Science, in Japan. We implement MCBound for Fugaku and classify the jobs executed during February 2024. Our approach is proven effective, as it obtains an F1-macro average score of at least 0.89 as prediction quality, while incurring a negligible overhead on the system’s operations. Our Python-based implementation of MCBound can be seamlessly configured and deployed in other HPC systems.1https://www.fujitsu.com/global/about/innovation/fugaku/
Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, Özalp Babaoglu, Yuetsu Kodama
SC5
2023 At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC Workloads
abstract
Over the last three decades, innovations in the memory subsystem were primarily targeted at overcoming the data movement bottleneck. In this paper, we focus on a specific market trend in memory technology: 3D-stacked memory and caches. We investigate the impact of extending the on-chip memory capabilities in future HPC-focused processors, particularly by 3D-stacked SRAM. First, we propose a method oblivious to the memory subsystem to gauge the upper-bound in performance improvements when data movement costs are eliminated. Then, using the gem5 simulator, we model two variants of a hypothetical LARge Cache processor (LARC), fabricated in 1.5 nm and enriched with high-capacity 3D-stacked cache. With a volume of experiments involving a broad set of proxy-applications and benchmarks, we aim to reveal how HPC CPU performance will evolve, and conclude an average boost of 9.56× for cache-sensitive HPC applications, on a per-chip basis. Additionally, we exhaustively document our methodological exploration to motivate HPC centers to drive their own technological agenda through enhanced co-design.
Jens Domke, Emil Vatai, Balazs Gerofi, Yuetsu Kodama, Mohamed Wahib, Artur Podobas, Sparsh Mittal, Miquel Pericàs, Lingqi Zhang 0001, Peng Chen 0035, Aleksandr Drozd, Satoshi Matsuoka
ACM Trans. Archit. Code Optim.4
2021 Evaluation of SPEC CPU and SPEC OMP on the A64FX
abstract
We evaluated the A64FX processor used in the supercomputer Fugaku using the SPEC CPU and SPEC OMP benchmark suites. As a result, we found the performance of the A64FX processor, which had 48 cores, was lower than that of the Xeon with dual sockets of 24 cores each in SPEC CPU int and fp. In SPEC OMP, due to the effect of the Xeon’s Hyperthread, the A64FX performance was lower than the performance of the Xeon with single socket of 28 cores. But in several benchmarks in SPEC CPU fp and SPEC OMP, the A64FX performance was higher due to its high memory bandwidth. In addition, by comparing the performance and power using the power control mechanism of the A64FX, it was confirmed that power can be reduced without affecting the performance when not using all cores.
Yuetsu Kodama, Masaaki Kondo, Mitsuhisa Sato
CLUSTER1
2021 Performance and power consumption analysis of Arm Scalable Vector Extension
Tetsuya Odajima, Yuetsu Kodama, Mitsuhisa Sato
J. Supercomput.2
2020 Evaluation of Power Management Control on the Supercomputer Fugaku
abstract
The supercomputer “Fugaku”, which recently ranked number one on multiple supercomputing lists, including the Top500 in June 2020, has various power control features, such as (1) an eco mode that utilizes only one of two floating-point pipelines while decreasing the power supply to the chip; (2) a boost mode that increases clock frequency; and (3) a core retention function that turns unused cores into a low-power state. By orchestrating these power-performance features while considering the characteristics of currently running applications, we can potentially gain even better system-level energy efficiency. In this article, we report on the effectiveness of these features using the pre-evaluation environment for Fugaku. As a result, we confirmed several prominent results useful for the operation of the Fugaku system, including: remarkable power reduction and energy-efficiency improvement by coordinating the eco mode and the core retention feature in the memory intensive case; a 10% speed-up with a 17% power consumption increase using the boost mode in the CPU intensive case; and considerable power variations across over 20K nodes.
Yuetsu Kodama, Tetsuya Odajima, Eishi Arima, Mitsuhisa Sato
CLUSTER1
2020 Performance Evaluation of Supercomputer Fugaku using Breadth-First Search Benchmark in Graph500
abstract
There is increasing demand for the high-speed processing of large-scale graphs in various fields. However, such graph processing requires irregular calculations, making it difficult to scale performance on large-scale distributed memory systems. Against this background, Graph500, a competition for evaluating the performance of large-scale graph processing, has been held. We developed breadth-first search (BFS), which is one of the benchmark kernels used in Graph500, and took the top spot a total of 10 times using the K computer. In this paper, we tune BFS performance and evaluate it using the supercomputer Fugaku, which is the successor to the K computer. The results of evaluating BFS for a large-scale graph composed of about 1.1 trillion vertices and 17.6 trillion edges using 92,160 nodes of Fugaku indicate that Fugaku has 2.27 times the performance of the K computer. Fugaku took the top spot on Graph500 in June 2020.
Masahiro Nakao, Koji Ueno, Katsuki Fujisawa, Yuetsu Kodama, Mitsuhisa Sato
CLUSTER4
2020 Preliminary Performance Evaluation of the Fujitsu A64FX Using HPC Applications
abstract
RIKEN Center for Computational Science has been installing the supercomputer Fugaku. The Fujitsu A64FX, based on the Armv8.2-A+SVE architecture, is used in the system. In this paper, we evaluated the seven HPC applications and benchmarks on the A64FX. In a performance comparison with Marvell (Cavium) ThunderX2 processor and Intel Xeon Skylake processor, the A64FX achieved higher performance in a memory bandwidth-intensive application thanks to its high memory bandwidth. However, we confirmed that the performance of the A64FX decreased from a lack of out-of-order resources. To mitigate this problem, the “loop fission” function of the Fujitsu compiler was used to improve the performance.
Tetsuya Odajima, Yuetsu Kodama, Miwako Tsuji, Motohiko Matsuda, Yutaka Maruyama, Mitsuhisa Sato
CLUSTER2
2020 Accuracy Improvement of Memory System Simulation for Modern Shared Memory Processor
abstract
For the purpose of developing applications for supercomputer Fugaku at an early stage, RIKEN has developed a processor simulator. This simulator is based on the general-purpose processor simulator gem5. It does not simulate the actual hardware of a Fugaku processor. However, we believe that sufficient simulation accuracy can be obtained since it simulates the instruction pipeline of out-of-order execution with cycle-level accuracy along with performing detailed parameter tuning of out-of-order resources. In order to estimate the accurate execution time of a program, it is necessary to simulate with accuracy not only the instruction execution time, but also the access time of the cache memory hierarchy. Therefore, in the RIKEN simulator, we expanded gem5 to match the performance of the cache memory hierarchy to that of a Fugaku processor. In this simulator, we aim to estimate the execution cycles of one node application on a Fugaku processor with accuracy that enables relative evaluation and application tuning. In this paper, we show the details of the implementation of this simulator and verify its accuracy compared with that of a Fugaku processor test chip. In the evaluation of the total 46 kernel benchmarks, it was confirmed that the difference is 13% or less for 85% of the kernels. In the multithreaded execution of Stream Triad benchmark, scalable performance according to the number of threads was confirmed, and achieved over 80% of memory throughput with enough accuracy.
Yuetsu Kodama, Tetsuya Odajima, Akira Asato, Mitsuhisa Sato
HPC Asia1
2020 Co-design for A64FX manycore processor and "Fugaku"
abstract
We have been carrying out the FLAGSHIP 2020 Project to develop the Japanese next-generation flagship supercomputer, the Post-K, recently named “Fugaku”. We have designed an original many core processor based on Armv8 instruction sets with the Scalable Vector Extension (SVE), an A64FX processor, as well as a system including interconnect and a storage subsystem with the industry partner, Fujitsu. The “co-design” of the system and applications is a key to making it power efficient and high performance. We determined many architectural parameters by reflecting an analysis of a set of target applications provided by applications teams. In this paper, we present the pragmatic practice of our co-design effort for “Fugaku”. As a result, the system has been proven to be a very power-efficient system, and it is confirmed that the performance of some target applications using the whole system is more than 100 times the performance of the K computer.
Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama, Tetsuya Odajima, Miwako Tsuji, Hisashi Yashiro, Masaki Aoki, Naoyuki Shida, Ikuo Miyoshi, Kouichi Hirai, Atsushi Furuya, Akira Asato, Kuniki Morita, Toshiyuki Shimizu
SC4
2017 Preliminary Performance Evaluation of Application Kernels Using ARM SVE with Multiple Vector Lengths
abstract
Modern high performance processors are equipped with very wide SIMD instruction set. SVE (Scalable Vector Extension) is an ARM® SIMD technology that supports vector lengths from 128 bits to 2048 bits. One of its promising features is to offer "vector-length agnostic" programming to allow the same SVE code to run on hardware of any vector length without any modification of the code. This feature would be useful to explore the best vector length with appropriate hardware resources in the space of various combinations of hardware parameters in order to make more efficient use of hardware resources, since we can use the same vectorized SIMDcode. In this paper, we report the performance of application kernelsusing ARM SVE with multiple vector lengths while keeping the hardware resource the same. We have confirmed that when the performance of the program is limited by a bottleneck of a long chain of arithmetic operations or instruction issues, the performance can be improved by increasing the vector length. However, it was necessary to prepare a sufficient number of physical registers for performance improvement, and when the number of physical registers was too small, it was found that with such a program, the performance might be reduced. When the performance is limited by memory access bandwidth to cache and memory, the vector length does not affect the performance significantly.
Yuetsu Kodama, Tetsuya Odajima, Motohiko Matsuda, Miwako Tsuji, Jinpil Lee, Mitsuhisa Sato
CLUSTER1
2015 Improving Strong-Scaling on GPU Cluster Based on Tightly Coupled Accelerators Architecture
abstract
The Tightly Coupled Accelerators (TCA) architecture that we proposed in previous work enables direct ommunication between accelerators over nodes. In this paper, we present a proof-of-concept GPU cluster called the HA-PACS/TCA using the PEACH2 chip that we designed as an interconnection router chip based on the TCA architecture. Our system demonstrated 2.0 ?sec of latency on inter-node GPU-to-GPU communication with a PCIe Gen2 x8 by RDMA, reducing minimum latency to just 44% of the InfiniBand-QDR and MPI using GPUDirect for RDMA. Through results of Himeno benchmark tests, we demonstrated that our TCA architecture improved performance scalability with the small-sized problem by up to 61%.
Toshihiro Hanawa, Hisafumi Fujii, Norihisa Fujita, Tetsuya Odajima, Kazuya Matsumoto, Yuetsu Kodama, Taisuke Boku
CLUSTER6
2013 The study of three-dimensional multiphase-flow simulator
abstract
This paper presents an FPGA-based system that aims to perform three-dimensional multiphase-flow simulations. In this implementation, the immiscible lattice-gas automata (LGA) were selected as the target simulation model. The immiscible LGA are classified as the cellular automata (CA), which are a discrete dynamic model. The simulation box consists of an array of cells. The structure of the box should be chosen carefully since it decides the limitation of simulation behaviors. On the other hand, all the lattice structures should be allocated in order so that they enable us to describe the LGA as stencil computation. Here, we expect that the FPGAs have a great possibility in achieving dramatic speedup. Experimental result shows speedups that achieves two orders of magnitude in the immiscible LGA with the face-centred hyper cubic lattice structure.
Kenta Fujinami, Yoshiki Yamaguchi, Akira Sugiura, Yuetsu Kodama
FPL4
2013 Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage
ICA3PP (2)5
2008 High Performance Relay Mechanism for MPI Communication Libraries Run on Multiple Private IP Address Clusters
abstract
We have been developing a Grid-enabled MPI communication library called GridMPI, which is designed to run on multiple clusters connected to a wide-area network. Some of these clusters may use private IP addresses. Therefore, some mechanism to enable communication between private IP address clusters is required. Such a mechanism should be widely adoptable, and should provide high communication performance. In this paper, we propose a message relay mechanism to support private IP address clusters in the manner of the Interoperable MPI (IMPI) standard. Therefore, any MPI implementations which follow the IMPI standard can communicate with the relay. Furthermore, we also propose a trunking method in which multiple pairs of relay nodes simultaneously communicate between clusters to improve the available communication bandwidth. While the relay mechanism introduces an one-way latency of about 25 musec, the extra overhead is negligible, since the communication latency through a wide area network is a few hundred times as large as this. By using trunking, the inter-cluster communication bandwidth can improve as the number of trunks increases. We confirmed the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. When relay nodes with 1 Gbps NICs are used, the performance of most of the NAS Parallel Benchmarks improved proportional to the number of trunks. Especially, using 8 trunks, FT and IS are 4.4 and 3.4 times faster, respectively, compared with the single trunk case. The results showed that the proposed method is effective for running MPI programs over high bandwidth-delay product networks.
Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa, Yasufumi Yoshizawa
CCGRID4
2007 Effects of packet pacing for MPI programs in a Grid environment
abstract
Improving the performance of TCP communication is the key to the successful deployment of MPI programs in a Grid environment in which multiple clusters are connected through high performance dedicated networks. To efficiently utilize the inter-cluster bandwidth, a traffic control mechanism is required so as not to allow the aggregate transmission bandwidth to exceed the inter-cluster bandwidth when multiple nodes communicate at one time. In this paper, we propose a traffic control method for MPI programs, in which an application or the MPI runtime controls the transmission rate based on the communication pattern by using certain MPI attributes. Packet pacing is used at each node preventing microscopic burst transmission to thus avoid congestion. We confirm the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. We show most of the NAS Parallel benchmarks improve the performance, since the proposed method reduces packet losses due to traffic congestion on the inter-cluster network. The results have indicated that it is feasible to connect multiple clusters and run large-scale scientific applications over distances up to 1000 kilometers, if an appropriate network is available.
Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa
CLUSTER4
2006 Efficient MPI Collective Operations for Clusters in Long-and-Fast Networks
abstract
Several MPI systems for grid environment, in which clusters are connected by wide-area networks, have been proposed. However, the algorithms of collective communication in such MPI systems assume relatively low bandwidth wide-area networks, and they are not designed for the fast wide-area networks that are becoming available. On the other hand, for cluster MPI systems, a beast algorithm by van de Geijn et al. and an allreduce algorithm by Rabenseifner have been proposed, which are efficient in a high bisection bandwidth environment. We modify those algorithms so as to effectively utilize fast wide-area inter-cluster networks and to control the number of nodes which can transfer data simultaneously through wide-area networks to avoid congestion. We confirmed the effectiveness of the modified algorithms by experiments using a 10 Gbps emulated WAN environment. The environment consists of two clusters, where each cluster consists of nodes with 1 Gbps Ethernet links and a switch with a 10 Gbps upper link. The two clusters are connected through a 10 Gbps WAN emulator which can insert latency. In a 10 millisecond latency environment, when the message size is 32 MB, the proposed beast and allreduce are 1.6 and 3.2 times faster, respectively, than the algorithms used in existing MPI systems for grid environment
Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa
CLUSTER3
2005 TCP Adaptation for MPI on Long-and-Fat Networks
abstract
Typical MPI applications work in phases of computation and communication, and messages are exchanged in relatively small chunks. This behavior is not optimal for TCP because TCP is designed only to handle a contiguous flow of messages efficiently. This behavior anomaly is well-known, but fixes are not integrated into today's TCP implementations, even though performance is seriously degraded, especially for MPI applications. This paper proposes three improvements in the Linux TCP stack: i.e., pacing at start-up, reducing Retransmit-Timeout time, and TCP parameter switching at the transition of computation phases in an MPI application. Evaluation of these improvements using the NAS parallel benchmarks shows that the BT, CG, IS, and SP benchmarks achieved 10 to 30 percent improvements. On the other hand, the FT and MG benchmarks showed no improvement because they have the steady communication that TCP assumes, and the LU benchmark became slightly worse because it has very little communication
Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa
CLUSTER3
2004 GNET-1: gigabit Ethernet network testbed
abstract
GNET-1 is a fully programmable network testbed. It provides functions such as wide area network emulation, network instrumentation, traffic shaping, and traffic generation at gigabit Ethernet wire speeds by programming the core FPGA. GNET-1 is a powerful tool for developing network-aware grid software. It is also a network monitoring and traffic-shaping tool that provides high-performance communication over wide area networks. This work describes several sample uses of GNET-1 and presents its architecture.
Yuetsu Kodama, Tomohiro Kudoh, Ryousei Takano, Hitoshi Sato, Osamu Tatebe, Satoshi Sekiguchi
CLUSTER1
2003 Design and implementation of PVFS-PM: a cluster file system on SCore
abstract
This paper discusses the design and implementation of a cluster file system, called PVFS-PM, on the SCore cluster system software. This is the first attempt to implement a cluster file system on the SCore system. It is based on the PVFS cluster file system but replaces TCP with the PMv2 communication library supported by SCore to provide a scalable, high-performance cluster file system. PVFS-PM improves the performance by factors of 1.07 and 1.93 for writing and reading, respectively, with 8 I/O nodes, compared with the original PVFS on TCP on a Gigabit Ethernet-connected SCore cluster.
Koji Segawa, Osamu Tatebe, Yuetsu Kodama, Tomohiro Kudoh, Toshiyuki Shimizu
CCGRID3
1999 Communication Studies of Single-Threaded and Multithreaded Distributed-Memory Multiprocessors
abstract
This report explicates the communication overlapping capabilities of three distributed-memory machines, SGI/Cray T3E, IBM SP-2 with wide nodes, and the ETL EM-X. Bitonic sorting and Fast Fourier Transform are selected for experiments. Various message sizes are used to determine when, where, how much and why the overlapping takes place. Experimental results with up to 64 processors indicated that the communication performance of EM-X is insensitive to various message sizes while SP-2 is the most sensitive. T3E stayed in between. The EM-X gave the highest communication overlapping capability while T3E did the lowest. The experimental results are compared with the analytical results based on LogP and LogGP communication models.
Andrew Sohn, Yunheung Paek, Jui-Yuan Ku, Yuetsu Kodama, Yoshinori Yamaguchi
HPCA4
1998 Load Balanced Parallel Radix Sort
abstract
Radix sort suffers from the unequal number of input keys due to the unknown characteristics of input keys.We present in this report a new radix sorting algorithm, called balanced radix sort which guarantees that each processor has exactly the same number of keys regardless of the data characteristics.The main idea of balanced radix sort is to store any processor which has over n/P keys to its neighbor processor, where n is the total number of keys and P is the number of processors.We have implemented balanced radix sort on two distributed-memory machines IBM SPZWN and Cray T3E.Multiple versions of 32-bit and 64-bit integers and 64-bit doubles are implemented in Message Passing Interface for portability.The sequential and parallel versions consist of approximately 50 and 150 lines of C code respectively including parallel constructs.Experimental results indicate that balanced radix sort can sort OSG integers in 20 seconds and 128M doubles in 15 seconds on a 64-processor SPZWN while yielding over 40-fold speedup.When compared with other radix sorting algorithms, balanced radix sort outperformed, showing two to six times faster.When compared with sample sorting algorithms, which are known to outperform all similar methods, balanced radix sort is 30% to 100% faster based on the same machine and key initialization.
Andrew Sohn, Yuetsu Kodama
International Conference on Supercomputing2
1998 Highly Efficient Implementation of MPI Point-to-Point Communication Using Remote Memory Operations
abstract
MPl point-to-point communication is a basic operation, however it requires runtime-matching of send and receive that causes to reduce performance.This paper proposes a new approach to send messages by remote memory write without inquiring of the receiver under a communication pattern such that nonblocking receive is issued in advance.Basically, this approach makes it possible to gain low latency and high bandwidth as the hardware specification.MPI-EMX, our implementation of the MPI on the EM-X multiprocessor, achieves a zero-byte latency of 13.4 psec.and a maximum bandwidth of 31.4MB/s, which can compete with commercial MPPs.This approach to reduce communication latency is widely applicable to other systems and is quite a promising technique for achieving low latency and high bandwidth.
Osamu Tatebe, Yuetsu Kodama, Satoshi Sekiguchi, Yoshinori Yamaguchi
International Conference on Supercomputing2
1998 Fast Speculative Search Engine on the Highly Parallel Computer EM-X
abstract
No abstract available.
Hayato Yamana, Hanpei Koike, Yuetsu Kodama, Hirofumi Sakane, Yoshinori Yamaguchi
SIGIR3
1997 Fine-Grain Multithreading with the EM-X Multiprocessor
abstract
Article Fine-grain multithreading with the EM-X multiprocessor Share on Authors: Andrew Sohn Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Yuetsu Kodama Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Jui Ku Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJ Computer and Information Science Dept., New Jersey Institute of Technology, Newark, NJView Profile , Mitsuhisa Sato Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, Japan Real World Computing Tsukuba Research Center, Tsukuba, Ibaraki, 305, JapanView Profile , Hirofumi Sakane Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Hayato Yamana Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Shuichi Sakai Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile , Yoshinori Yamaguchi Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, Japan Computer Architecture Section, Electrotechnical Laboratory, Tsukuba-shi, Ibaraki 305, JapanView Profile Authors Info & Claims SPAA '97: Proceedings of the ninth annual ACM symposium on Parallel algorithms and architecturesJune 1997 Pages 189–198https://doi.org/10.1145/258492.258511Online:01 June 1997Publication History 7citation121DownloadsMetricsTotal Citations7Total Downloads121Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Andrew Sohn, Yuetsu Kodama, Jui Ku, Mitsuhisa Sato, Hirofumi Sakane, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi
SPAA2
1995 A Macrotask-level Unlimited Speculative Execution on Multiprocessors
abstract
The purpose of this paper is to propose a new fast EM-4 multiprocessor to decrease the control overhead of macrotasks.Preliminary evaluations show that the control overhead of the proposed scheme is smaller than that of the other control schemes.Moreover, it is confirmed that the distributed control can be implemented by using software when the average macrotssk execution time is larger than 14.4ps on the EM-4 multiprocessor.
Hayato Yamana, Mitsuhisa Sato, Yuetsu Kodama, Hirofumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi
International Conference on Supercomputing3
1995 The EM-X Parallel Computer: Architecture and Basic Performance
abstract
Latency tolerance is essential in achieving high performance on parallel computers for remote function calls and fine-grained remote memory accesses. EM-X supports interprocessor communication on an execution pipeline with small and simple packets. It can create a packet in one cycle, and receive a packet from the network in the on-chip buffer without interruption. EM-X invokes threads on packet arrival, minimizing the overhead of thread switching. It can tolerate communication latency by using efficient multi-threading and optimizing packet flow of fine grain communication. EM-X also supports the synchronization of two operands, direct remote memory read/write operations and flexible packet scheduling with priority. This paper describes distinctive features of the EM-X architecture and reports the performance of small synthetic programs and larger more realistic programs.
Yuetsu Kodama, Hirohumi Sakane, Mitsuhisa Sato, Hayato Yamana, Shuichi Sakai, Yoshinori Yamaguchi
ISCA1
1995 Reduced Interprocessor-Communication Architecture and its Implementation on EM-4
Shuichi Sakai, Yuetsu Kodama, Mitsuhisa Sato, Andrew Shaw, Hiroshi Matsuoka, Hideo Hirono, Kazuaki Okamoto, Takashi Yokota
Parallel Comput.2
1994 Nonnumeric search results on the EM-4 distributed-memory multiprocessor
abstract
Numeric scientific problems have been the main focus of supercomputing as their numerous implementations on various multiprocessors indicate. Nonnumeric problems on the other hand have received very little attention for parallel implementation due to their irregular behaviors and a large amount of resource usage. This report presents our experiences implementing difficult nonnumeric problems on the EM-4 multiprocessor, believing that supercomputers should also be able to effectively execute nonnumeric problems if they are to be considered 'supercomputers'. We selected two typical search problems, the Eight-Puzzle and the Tower-of-Hanoi. Two parallel search techniques we used to implement the search problems, unidirectional and bidirectional heuristic search. A total of eight different programs have been implemented on the EM-4 multiprocessor with realistic problem sizes. Execution results demonstrate that the parallel bidirectional heuristic search can solve the tree depth 20 to 40 of the Eight-Puzzle in an optimal or near optimal number of iterations in less than two seconds, and is highly scalable as it gives over 40-fold speedup for both problems on 80 processors.>
Andrew Sohn, Mitsuhisa Sato, Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi
SC4
1993 EMC-Y: Parallel Processing Element Optimizing Communication and Computation
abstract
EMC-Y is a new processing element for highly parallel computers designed to achieve high performance parallel computation by fusing a dataflow mechanism and a von Neumann execution pipeline. We have already developed EMC-R, which is the processing element used in the EM-4 prototype. EMC-Y improves on EMC-R's packet communication performance, allowing it to tolerate a more network traffic. This paper presents the architecture of EMC-Y, concentrating on the principles of packet communication. EMC-Y uses an output packet buffer and optimal packet routing to improve the performance of packet sending and transferring. EMC-Y changes the memory access priority for input packet buffer operation to improve the performance of receiving packets. Since the EMC-Y processor not only improves the performance of packet input and output but also balances them, it can tolerate a large amount of traffic and can improve the execution performance. We evaluate the improvements of EMC-Y architecture using a clock level simulator. The results show that EMC-Y improves performance by 50% to 70% in several programs over EMC-R at the same clock speed.
Yuetsu Kodama, Yasuhito Koumura, Mitsuhisa Sato, Hirohumi Sakane, Shuichi Sakai, Yoshinori Yamaguchi
International Conference on Supercomputing1
1993 Super-Threading: Architectural and Software Mechanisms for Optimizing Parallel Computation
abstract
This paper presents super-threading, which generically means the architectural and software mechanisms for optimizing parallel computation. Super-threading includes architectural optimization of a processing element (PE), mechanism for supporting fast communication and computation, techniques of a compiler and a run time system for optimizing thread creation, thread allocation, tuning of granularity and data allocation to physically distributed storage.This paper states what super-threading is and examines some of the technologies belonging to it. The processor architecture based on super-threading is proposed and its implementation on a highly parallel computer EM-4 is shown with performance data. Software issues about super-threading are also examined mainly from the viewpoint of granularity optimization. Dynamic granularity optimization methods are proposed here, and evaluated on EM-4. The performance data indicate that super-threading is a key technology for realizing an efficient massively parallel computer.
Shuichi Sakai, Kazuaki Okamoto, Hiroshi Matsuoka, Hideo Hirono, Yuetsu Kodama, Mitsuhisa Sato
International Conference on Supercomputing5
1993 Design and Implementation of a Circular Omega Network in the EM-4
Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi
Parallel Comput.2
1992 Thread-based Programming for the EM-4 Hybrid Dataflow Machine
abstract
In this paper, we present a thread-based programming model for the EM-4 hybrid dataflow machine, where parallelism and synchronization among threads of sequential execution are described explicitly by the programmer. Although EM-4 was originally designed as a dataflow machine, we demonstrate that it provides effective architectural support for a variety of programming styles, including message passing and distributed data sharing in imperative languages. Our approach allows the programmer to control the parallelism and maintain data locality explicitly to achieve high performance. EM-4 can be thought of as a multi-threaded architecture that can exploit both von Neumann and dataflow compiling technology. Thread-based programming provides the first step to explore better programming/compiling technology for a hybrid dataflow machine as well as EM-4.
Mitsuhisa Sato, Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi, Yasuhito Koumura
ISCA2
1992 A prototype of a highly parallel dataflow machine EM-4 and its preliminary evaluation
Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi
Future Gener. Comput. Syst.1
1992 Methodologies in development and testing of the dataflow machine EM-4
Kazuaki Okamoto, Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi
Parallel Comput.2
1991 Design and Implementation of a Versatile Interconnection Network in the EM-4
Shuichi Sakai, Yuetsu Kodama, Yoshinori Yamaguchi
ICPP (1)2
1991 Load balancing by function distribution on the EM-4 prototype
abstract
The EM-4 is a highly parallel daiaflow machine that will eventually have more than 1,000 processing elements (PEs). This paper presents load balancing methods by function distribution in the EM-4 and their evaluations on the EM-4 prototype, which con-sists of 80 PEs. The EM-4 can distribute function in-stances statically by using several different allocation functions, including a method exploiting the locality of the network. Furthermore, the EM-4 can dynami-cally distribute function instances using MLPE packets which circulate through the PEs and detect the local-minimum load PE. These function distn’bution meth-ods are evaluated by executing a divide-and-conquer program and a game tree searching program, and ex-amining its dynamic characteristics on every PE. 1
Yuetsu Kodama, Shuichi Sakai, Yoshinori Yamaguchi
SC1
1989 An Architecture of a Dataflow Single Chip Processor
abstract
A highly parallel (more than a thousand) dataflow machine EM-4 is now under development. The EM-4 design principle is to construct a high performance computer using a compact architecture by overcoming several defects of dataflow machines. Constructing the EM-4, it is essential to fabricate a processing element (PE) on a single chip for reducing operation speed, system size, design complexity and cost. In the EM-4, the PE , called EMC-R, has been specially designed using a 50,000-gate gate array chip. This paper focuses on an architecture of the EMC-R. The distinctive features of it are: a strongly connected arc dataflow model; a direct matching scheme; a RISC-based design; a deadlock-free on-chip packet switch; and an integration of a packet-based circular pipeline and a register-based advanced control pipeline. These features are intensively examined, and the instruction set architecture and the configuration architecture which exploit them are described.
Shuichi Sakai, Yoshinori Yamaguchi, Kei Hiraki, Yuetsu Kodama, Toshitsugu Yuba
ISCA4