EDBT 2026 Demo / reviewers in the wild / expert
Satoshi Matsuoka
dblp:57/4464
· DBLP profile ↗
178ranked-venue papers
15as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 138 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 1 first-authorDatabases, data management, data science and information retrieval · 10 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 10 · 4 first-authorSoftware engineering, systems software and programming languages · 9 · 4 first-authorHuman-computer interaction and ubiquitous computing · 9Theory of computation · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix MultiplicationabstractDistributed Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental operation in high-performance computing and deep learning applications. The major performance bottleneck in distributed SpMM lies in substantial communication overhead, which limits both performance and scalability. In this paper, we identify two key sources of communication inefficiency in distributed SpMM: redundant data transfer due to sparsity unawareness, and suboptimal utilization of hierarchical network topology. To address these, we propose (1) a fine-grained, sparsity-aware communication strategy that reduces communication overhead by exploiting the sparsity pattern of the sparse matrix, and (2) a hierarchical communication strategy that maps the sparsity-aware strategy onto two-tier GPU network architectures, minimizing redundant data movement across slower inter-node links. We implement these optimizations in a comprehensive distributed SpMM framework, SHIRO. Extensive evaluations on real-world datasets show that SHIRO demonstrates strong scalability up to 128 GPUs, achieving geometric mean speedups of 221.5 ×, 56.0 ×, 23.4 ×, and 8.8 × in SpMM over four state-of-the-art baselines (CAGNET, SPA, BCL, and CoLa, respectively) at this scale. Chen Zhuang, Lingqi Zhang 0001, Benjamin Brock, Du Wu, Peng Chen 0035, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 7 |
| 2026 | RoWD: Automated rogue workload detector for HPC securityabstractThe increasing reliance on High-Performance Computing (HPC) systems to execute complex scientific and industrial workloads raises significant security concerns related to the misuse of HPC resources for unauthorized or malicious activities. Rogue job executions can threaten the integrity, confidentiality, and availability of HPC infrastructures. Given the scale and heterogeneity of HPC job submissions, manual or ad hoc monitoring is inadequate to effectively detect such misuse. Therefore, automated solutions capable of systematically analyzing job submissions are essential to detect rogue workloads. To address this challenge, we present RoWD (Rogue Workload Detector), the first framework for automated and systematic security screening of the HPC job-submission pipeline. RoWD is composed of modular plug-ins that classify different types of workloads and enable the detection of rogue jobs through the analysis of job scripts and associated metadata. We deploy RoWD on the Supercomputer Fugaku to classify AI workloads and release SCRIPT-AI, the first dataset of annotated job scripts labeled with workload characteristics. We evaluate RoWD on approximately 50K previously unseen jobs executed on Fugaku between 2021 and 2025. Our results show that RoWD accurately classifies AI jobs (achieving an F1 score of 95%), is robust against adversarial behavior, and incurs low runtime overhead, making it suitable for strengthening the security of HPC environments and for real-time deployment in production systems. Francesco Antici, Jens Domke, Andrea Bartolini, Zeynep Kiziltan, Satoshi Matsuoka |
Future Gener. Comput. Syst. | 5 |
| 2025 | Scaling Large-scale GNN Training to Thousands of Processors on CPU-based SupercomputersabstractGraph Convolutional Networks (GCNs), particularly for largescale graphs, are crucial across numerous domains.However, training distributed full-batch GCNs on large-scale graphs suffers from inefficient memory access patterns and high communication overhead.To address these challenges, we introduce SuperGCN, an efficient and scalable distributed GCN Chen Zhuang, Lingqi Zhang 0001, Du Wu, Peng Chen 0035, Jiajun Huang 0001, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 10 |
| 2025 | A General and Scalable GCN Training Framework on CPU SupercomputersabstractGraph Convolutional Networks (GCNs) are widely used in various domains. However, training distributed full-batch GCNs on large-scale graphs poses challenges due to inefficient memory access patterns and high communication overhead. This paper presents a general and efficient GCN training framework on CPU supercomputers. It comprises a general aggregation kernel designed to optimize irregular memory access and a quantization method with label propagation to reduce communication overhead. Experimental results show that our method achieves a speedup of up to 4.1× compared with the SoTA implementations. Chen Zhuang, Peng Chen 0035, Xin Liu 0020, Rio Yokota, Nikoli Dryden, Lingqi Zhang 0001, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
PPoPP | 8 |
| 2024 | Real-time High-resolution X-Ray Computed TomographyabstractComputed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes. Du Wu, Peng Chen 0035, Xiao Wang 0004, Isaac Lyngaas, Takaaki Miyajima, Toshio Endo, Satoshi Matsuoka, Mohamed Wahib |
ICS | 7 |
| 2023 | PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU ApplicationsabstractIterative memory-bound solvers commonly occur in HPC codes. Typical GPU implementations have a loop on the host side that invokes the GPU kernel as much as time/algorithm steps there are. The termination of each kernel implicitly acts the barrier required after advancing the solution every time step. We propose an execution model for running memory-bound iterative GPU kernels: PERsistent KernelS (PERKS). In this model, the time loop is moved inside persistent kernel, and device-wide barriers are used for synchronization. We then reduce the traffic to device memory by caching subset of the output in each time step in the unused registers and shared memory. PERKS can be generalized to any iterative solver: they largely independent of the solver's implementation. We explain the design principle of PERKS and demonstrate effectiveness of PERKS for a wide range of iterative 2D/3D stencil benchmarks (geomean speedup of 2.12x for 2D stencils and 1.24x for 3D stencils over state-of-art libraries), and a Krylov subspace conjugate gradient solver (geomean speedup of 4.86x in smaller SpMV datasets from SuiteSparse and 1.43x in larger SpMV datasets over a state-of-art library). All PERKS-based implementations available at: https://github.com/neozhang307/PERKS. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 7 |
| 2023 | Revisiting Temporal Blocking Stencil OptimizationsabstractIterative stencils are used widely across the spectrum of High Performance Computing (HPC) applications. Many efforts have been put into optimizing stencil GPU kernels, given the prevalence of GPU-accelerated supercomputers. To improve the data locality, temporal blocking is an optimization that combines a batch of time steps to process them together. Under the observation that GPUs are evolving to resemble CPUs in some aspects, we revisit temporal blocking optimizations for GPUs. We explore how temporal blocking schemes can be adapted to the new features in the recent Nvidia GPUs, including large scratchpad memory, hardware prefetching, and device-wide synchronization. We propose a novel temporal blocking method, EBISU, which champions low device occupancy to drive aggressive deep temporal blocking on large tiles that are executed tile-by-tile. We compare EBISU with state-of-the-art temporal blocking libraries: STENCILGEN and AN5D. We also compare with state-of-the-art stencil auto-tuning tools that are equipped with temporal blocking optimizations: ARTEMIS and DRSTENCIL. Over a wide range of stencil benchmarks, EBISU achieves speedups up to 2.53x and a geometric mean speedup of 1.49x over the best state-of-the-art performance in each stencil benchmark. Lingqi Zhang 0001, Mohamed Wahib, Peng Chen 0035, Jintao Meng 0001, Xiao Wang 0004, Toshio Endo, Satoshi Matsuoka |
ICS | 7 |
| 2023 | Efficient checkpoint/Restart of CUDA applications
Akira Nukada, Taichiro Suzuki, Satoshi Matsuoka |
Parallel Comput. | 3 |
| 2023 | At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC WorkloadsabstractOver the last three decades, innovations in the memory subsystem were primarily targeted at overcoming the data movement bottleneck. In this paper, we focus on a specific market trend in memory technology: 3D-stacked memory and caches. We investigate the impact of extending the on-chip memory capabilities in future HPC-focused processors, particularly by 3D-stacked SRAM. First, we propose a method oblivious to the memory subsystem to gauge the upper-bound in performance improvements when data movement costs are eliminated. Then, using the gem5 simulator, we model two variants of a hypothetical LARge Cache processor (LARC), fabricated in 1.5 nm and enriched with high-capacity 3D-stacked cache. With a volume of experiments involving a broad set of proxy-applications and benchmarks, we aim to reveal how HPC CPU performance will evolve, and conclude an average boost of 9.56× for cache-sensitive HPC applications, on a per-chip basis. Additionally, we exhaustively document our methodological exploration to motivate HPC centers to drive their own technological agenda through enhanced co-design. Jens Domke, Emil Vatai, Balazs Gerofi, Yuetsu Kodama, Mohamed Wahib, Artur Podobas, Sparsh Mittal, Miquel Pericàs, Lingqi Zhang 0001, Peng Chen 0035, Aleksandr Drozd, Satoshi Matsuoka |
ACM Trans. Archit. Code Optim. | 12 |
| 2023 | Simeuro: A Hybrid CPU-GPU Parallel Simulator for Neuromorphic Computing ChipsabstractWith the success of deep learning, there have been numerous efforts to build hardware for it. One approach that is gaining momentum is neuromorphic computing with spiking neural networks (SNNs), which are multiplication-free and open the possibility of using analog computing via novel technologies. However, to design effective and efficient hardware for such architectures, a fast and accurate software simulator is key. This article presents Simeuro, a fast and scalable system-level simulator for SNN models used in neuromorphic accelerators. The simulator uses spike-level details and configurable architectural constraints that are independent of the underlying hardware implementation. Simeuro supports a wide range of features including analog computing, novel memory (currently, RRAM is supported), and a full network-on-chip. The simulator can provide detailed simulation results such as routing statistics, energy consumption, delay, and accuracy of arbitrarily defined SNN architectures. Our simulator leverages a CPU-GPU hybrid environment to expedite the simulation by scaling out to multi-nodes equipped with multi-GPUs. We are able to conduct core simulations for a system-scale SNN chip of 20,000 neuromorphic cores on up to 512 A100 GPUs in a few minutes. Huaipeng Zhang, Nhut-Minh Ho, Dogukan Yigit Polat, Peng Chen 0035, Mohamed Wahib, Truong Thao Nguyen, Jintao Meng 0001, Rick Siow Mong Goh, Satoshi Matsuoka, Tao Luo 0014, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2021 | Performance portable back-projection algorithms on CPUs: agnostic data locality and vectorization optimizationsabstractComputed Tomography (CT) is a key 3D imaging technology that fundamentally relies on the compute-intense back-projection operation to generate 3D volumes. GPUs are typically used for back-projection in production CT devices. However, with the rise of power-constrained micro-CT devices, and also the emergence of CPUs comparable in performance to GPUs, back-projection for CPUs could become favorable. Unlike GPUs, extracting parallelism for back-projection algorithms on CPUs is complex given that parallelism and locality are not explicitly defined and controlled by the programmer, as is the case when using CUDA for instance. We propose a collection of novel back-projection algorithms that reduce the arithmetic computation, robustly enable vectorization, enforce a regular memory access pattern, and maximize the data locality. We also implement the novel algorithms as efficient back-projection kernels that are performance portable over a wide range of CPUs. Performance evaluation using a variety of CPUs from different vendors and generations demonstrates that our back-projection implementation achieves on average 5.2 times speedup over the multi-threaded implementation of the most widely used, and optimized, open library. With a state‐of‐the‐art CPU, we reach performance that rivals top-performing GPUs. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Shin'ichiro Takizawa, Takahiro Hirofuchi, Hirotaka Ogawa, Satoshi Matsuoka |
ICS | 7 |
| 2021 | Matrix Engines for High Performance Computing: A Paragon of Performance or Grasping at Straws?abstractMatrix engines or units, in different forms and affinities, are becoming a reality in modern processors; CPUs and otherwise. The current and dominant algorithmic approach to Deep Learning merits the commercial investments in these units, and deduced from the No. 1 benchmark in supercomputing, namely High Performance Linpack, one would expect an awakened enthusiasm by the HPC community, too. Hence, our goal is to identify the practical added benefits for HPC and machine learning applications by having access to matrix engines. For this purpose, we perform an in-depth survey of software stacks, proxy applications and benchmarks, and historical batch job records. We provide a cost-benefit analysis of matrix engines, both asymptotically and in conjunction with state-of-the-art processors. While our empirical data will temper the enthusiasm, we also outline opportunities to “misuse” these dense matrix-multiplication engines if they come for free. Jens Domke, Emil Vatai, Aleksandr Drozd, Peng Chen 0035, Yosuke Oyama, Lingqi Zhang 0001, Shweta Salaria, Daichi Mukunoki, Artur Podobas, Mohamed Wahib, Satoshi Matsuoka |
IPDPS | 11 |
| 2021 | Scalable FBP decomposition for cone-beam CT reconstructionabstractFiltered Back-Projection (FBP) is a fundamental compute intense algorithm used in tomographic image reconstruction. Cone-Beam Computed Tomography (CBCT) devices use a cone-shaped X-ray beam, in comparison to the parallel beam used in older CT generations. Distributed image reconstruction of cone-beam datasets typically relies on dividing batches of images into different nodes. This simple input decomposition, however, introduces limits on input/output sizes and scalability. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Takahiro Hirofuchi, Hirotaka Ogawa, Ander Biguri, Richard P. Boardman, Thomas Blumensath, Satoshi Matsuoka |
SC | 9 |
| 2021 | The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs With Hybrid ParallelismabstractWe present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy. Yosuke Oyama, Naoya Maruyama, Nikoli Dryden, Erin McCarthy, Peter Harrington, Jan Balewski, Satoshi Matsuoka, Peter Nugent, Brian Van Essen |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | A Template-based Framework for Exploring Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are being considered as a complementary addition to modern High-Performance Computing (HPC) systems. These reconfigurable devices overcome many of the limitations of the (more popular) FPGA, by providing higher operating frequency, denser compute capacity, and lower power consumption. Today, CGRAs have been used in several embedded applications, including automobile, telecommunication, and mobile systems, but the literature on CGRAs in HPC is sparse and the field full of research opportunities. In this work, we introduce our CGRA simulator infrastructure for use in evaluating future HPC CGRA systems. Our CGRA simulator is built on synthesizable VHDL and is highly parametrizable, including support for connectivity, SIMD, data-type width, and heterogeneity. Unlike other related work, our framework supports co-integration with third-party memory simulators or evaluation of future memory architecture, which is crucial to reason around memory-bound applications. We demonstrate how our framework can be used to explore the performance of multiple different kernels, showing the impact of different configuration and design-space options. Artur Podobas, Kentaro Sano, Satoshi Matsuoka |
ASAP | 3 |
| 2020 | AN5D: automated stencil framework for high-degree temporal blocking on GPUsabstractStencil computation is one of the most widely-used compute patterns in high performance computing applications. Spatial and temporal blocking have been proposed to overcome the memory-bound nature of this type of computation by moving memory pressure from external memory to on-chip memory on GPUs. However, correctly implementing those optimizations while considering the complexity of the architecture and memory hierarchy of GPUs to achieve high performance is difficult. We propose AN5D, an automated stencil framework which is capable of automatically transforming and optimizing stencil patterns in a given C source code, and generating corresponding CUDA code. Parameter tuning in our framework is guided by our performance model. Our novel optimization strategy reduces shared memory and register pressure in comparison to existing implementations, allowing performance scaling up to a temporal blocking degree of 10. We achieve the highest performance reported so far for all evaluated stencil benchmarks on the state-of-the-art Tesla V100 GPU. Kazuaki Matsumura, Hamid Reza Zohouri, Mohamed Wahib, Toshio Endo, Satoshi Matsuoka |
CGO | 5 |
| 2020 | A Study of Single and Multi-device Synchronization Methods in Nvidia GPUsabstractGPUs are playing an increasingly important role in general-purpose computing. Many algorithms require synchronizations at different levels of granularity in a single GPU. Additionally, the emergence of dense GPU nodes also calls for multi-GPU synchronization. Nvidia's latest CUDA provides a variety of synchronization methods. Until now, there is no full understanding of the characteristics of those synchronization methods. This work explores important undocumented features and provides an in-depth analysis of the performance considerations and pitfalls of the state-of-art synchronization methods for Nvidia GPUs. The provided analysis would be useful when making design choices for applications, libraries, and frameworks running on single and/or multi-GPU environments. We provide a case study of the commonly used reduction operator to illustrate how the knowledge gained in our analysis can be useful. We also describe our micro-benchmarks and measurement methods. Lingqi Zhang 0001, Mohamed Wahib, Satoshi Matsuoka |
IPDPS | 4 |
| 2020 | A Formal Model for a Linear Time Correctness Condition of Proof Nets of Multiplicative Linear Logic
Satoshi Matsuoka |
LOPSTR | 1 |
| 2020 | Scaling distributed deep learning workloads beyond the memory capacity with KARMAabstractThe dedicated memory of hardware accelerators can be insufficient to store all weights and/or intermediate states of large deep learning models. Although model parallelism is a viable approach to reduce the memory pressure issue, significant modification of the source code and considerations for algorithms are required. An alternative solution is to use out-of-core methods instead of, or in addition to, data parallelism. We propose a performance model based on the concurrency analysis of out-of-core training behavior, and derive a strategy that combines layer swapping and redundant recomputing. We achieve an average of 1. 52x speedup in six different models over the state-of-the-art out-of-core methods. We also introduce the first method to solve the challenging problem of out-of-core multi-node training by carefully pipelining gradient exchanges and performing the parameter updates on the host. Our data parallel out-of-core solution can outperform complex hybrid model parallelism in training large models, e.g. Megatron-LM and Turning-NLG. Mohamed Wahib, Truong Thao Nguyen, Aleksandr Drozd, Jens Domke, Lingqi Zhang 0001, Ryousei Takano, Satoshi Matsuoka |
SC | 8 |
| 2019 | Batched Sparse Matrix Multiplication for Accelerating Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) are recently getting much attention in bioinformatics and chemoinformatics as a state-of-the-art machine learning approach with high accuracy. GCNs process convolutional operations along with graph structures, and GPUs are used to process enormous operations including sparse-dense matrix multiplication (SpMM) when the graph structure is expressed as an adjacency matrix with sparse matrix format. However, the SpMM operation on small graph, where the number of nodes is tens or hundreds, hardly exploits high parallelism or compute power of GPU. Therefore, SpMM becomes a bottleneck of training and inference in GCNs applications. In order to improve the performance of GCNs applications, we propose new SpMM algorithm especially for small sparse matrix and Batched SpMM, which exploits high parallelism of GPU by processing multiple SpMM operations with single CUDA kernel. To the best of our knowledge, this is the first work of batched approach for SpMM. We evaluated the performance of the GCNs application on TSUBAME3.0 implementing NVIDIA Tesla P100 GPU, and our batched approach shows significant speedups of up to 1.59x and 1.37x in training and inference, respectively. Yusuke Nagasaka, Akira Nukada, Ryosuke Kojima, Satoshi Matsuoka |
CCGRID | 4 |
| 2019 | Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural NetworksabstractLarge-scale distributed training of deep neural networks suffers from the generalization gap caused by the increase in the effective mini-batch size. Previous approaches try to solve this problem by varying the learning rate and batch size over epochs and layers, or some ad hoc modification of the batch normalization. We propose an alternative approach using a second order optimization method that shows similar generalization capability to first order methods, but converges faster and can handle larger mini-batches. To test our method on a benchmark where highly optimized first order methods are available as references, we train ResNet-50 on ImageNet-1K. We converged to 75% Top-1 validation accuracy in 35 epochs for mini-batch sizes under 16,384, and achieved 75% even with a mini-batch size of 131,072, which took only 978 iterations. Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, Satoshi Matsuoka |
CVPR | 6 |
| 2019 | Double-Precision FPUs in High-Performance Computing: An Embarrassment of Riches?abstractAmong the (uncontended) common wisdom in High-Performance Computing (HPC) is the applications' need for large amount of double-precision support in hardware. Hardware manufacturers, the TOP500 list, and (rarely revisited) legacy software have without doubt followed and contributed to this view. In this paper, we challenge that wisdom, and we do so by exhaustively comparing a large number of HPC proxy applications on two processors: Intel's Knights Landing (KNL) and Knights Mill (KNM). Although similar, the KNL and KNM architecturally deviate at one important point: the silicon area devoted to double-precision arithmetics. This fortunate discrepancy allows us to empirically quantify the performance impact in reducing the amount of hardware double-precision arithmetic. Our analysis shows that this common wisdom might not always be right. We find that the investigated HPC proxy applications do allow for a (significant) reduction in double-precision with little-to-no performance implications. With the advent of a failing of Moore's law, our results partially reinforce the view taken by modern industry (e.g., upcoming Fujitsu ARM64FX) to integrate hybrid-precision hardware units. Jens Domke, Kazuaki Matsumura, Mohamed Wahib, Keita Yashima, Toshiki Tsuchikawa, Yohei Tsuji, Artur Podobas, Satoshi Matsuoka |
IPDPS | 9 |
| 2019 | A versatile software systolic execution model for GPU memory-bound kernelsabstractThis paper proposes a versatile high-performance execution model, inspired by systolic arrays, for memory-bound regular kernels running on CUDA-enabled GPUs. We formulate a systolic model that shifts partial sums by CUDA warp primitives for the computation. We also employ register files as a cache resource in order to operate the entire model efficiently. We demonstrate the effectiveness and versatility of the proposed model for a wide variety of stencil kernels that appear commonly in HPC, and also convolution kernels (increasingly important in deep learning workloads). Our algorithm outperforms the top reported state-of-the-art stencil implementations, including implementations with sophisticated temporal and spatial blocking techniques, on the two latest Nvidia architectures: Tesla V100 and P100. For 2D convolution of general filter sizes and shapes, our algorithm is on average 2.5× faster than Nvidia's NPP on V100 and P100 GPUs. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 5 |
| 2019 | iFDK: a scalable framework for instant high-resolution image reconstructionabstractComputed Tomography (CT) is a widely used technology that requires compute-intense algorithms for image reconstruction. We propose a novel back-projection algorithm that reduces the projection computation cost to 1/6 of the standard algorithm. We also propose an efficient implementation that takes advantage of the heterogeneity of GPU-accelerated systems by overlapping the filtering and back-projection stages on CPUs and GPUs, respectively. Finally, we propose a distributed framework for high-resolution image reconstruction on state-of-the-art GPU-accelerated supercomputers. The framework relies on an elaborate interleave of MPI collective communication steps to achieve scalable communication. Evaluation on a single Tesla V100 GPU demonstrates that our back-projection kernel performs up to 1.6× faster than the standard FDK implementation. We also demonstrate the scalability and instantaneous CT capability of the distributed framework by using up to 2,048 V100 GPUs to solve 4K and 8K problems within 30 seconds and 2 minutes, respectively (including I/O). Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
SC | 5 |
| 2019 | HyperX topology: first at-scale implementation and comparison to the fat-treeabstractThe de-facto standard topology for modern HPC systems and data-centers are Folded Clos networks, commonly known as Fat-Trees. The number of network endpoints in these systems is steadily increasing. The switch radix increase is not keeping up, forcing an increased path length in these multi-level trees that will limit gains for latency-sensitive applications. Additionally, today's Fat-Trees force the extensive use of active optical cables which carries a prohibitive cost-structure at scale. To tackle these issues, researchers proposed various low-diameter topologies, such as Dragonfly. Another novel, but only theoretically studied, option is the HyperX. We built the world's first 3 Pflop/s supercomputer with two separate networks, a 3--level Fat-Tree and a 12×8 HyperX. This dual-plane system allows us to perform a side-by-side comparison using a broad set of benchmarks. We show that the HyperX, together with our novel communication pattern-aware routing, can challenge the performance of, or even outperform, traditional Fat-Trees. Jens Domke, Satoshi Matsuoka, Ivan R. Ivanov, Yuki Tsushima, Tomoya Yuki, Akihiro Nomura 0002, Shin'ichi Miura, Nic McDonald, Dennis Lee Floyd, Nicolas Dubé |
SC | 2 |
| 2019 | Scaling Word2Vec on Big CorpusabstractWord embedding has been well accepted as an important feature in the area of natural language processing (NLP). Specifically, the Word2Vec model learns high-quality word embeddings and is widely used in various NLP tasks. The training of Word2Vec is sequential on a CPU due to strong dependencies between word–context pairs. In this paper, we target to scale Word2Vec on a GPU cluster. To do this, one main challenge is reducing dependencies inside a large training batch. We heuristically design a variation of Word2Vec, which ensures that each word–context pair contains a non-dependent word and a uniformly sampled contextual word. During batch training, we “freeze” the context part and update only on the non-dependent part to reduce conflicts. This variation also directly controls the training iterations by fixing the number of samples and treats high-frequency and low-frequency words equally. We conduct extensive experiments over a range of NLP tasks. The results show that our proposed model achieves a 7.5 times acceleration on 16 GPUs without accuracy drop. Moreover, by using high-level Chainer deep learning framework, we can easily implement Word2Vec variations such as CNN-based subword-level models and achieves similar scaling results. Bofang Li, Aleksandr Drozd, Yuhe Guo, Tao Liu 0001, Satoshi Matsuoka, Xiaoyong Du 0001 |
Data Sci. Eng. | 5 |
| 2019 | Performance optimization, modeling and analysis of sparse matrix-matrix products on multi-core and many-core processors
Yusuke Nagasaka, Satoshi Matsuoka, Ariful Azad, Aydin Buluç |
Parallel Comput. | 2 |
| 2018 | Optimizing Preconditioned Conjugate Gradient on TaihuLight for OpenFOAMabstractPorting the domain-specific software OpenFOAM onto the TaihuLight supercomputer is a challenging task, due to the highly memory-bound nature of both the supercomputer's processor (SW26010) and the software's liner solvers. Our study tackles this technical challenge, in three steps, by optimizing the linear solvers, such as Preconditioned Conjugate Gradient (PCG), on the SW26010. First, in order to minimize the all_reduce communication cost of PCG, we developed a new algorithm RNPCG, a non-blocking PCG leveraging the on-chip register communication. Second, we optimized three key kernels of the PCG, including proposing a localized version of the Diagonal-based Incomplete Cholesky (LDIC) preconditioner. Third, to scale the RNPCG on TaihuLight, we designed the three-level non-blocking all_reduce operations. With these three steps, we implemented the RNPCG in OpenFOAM. The experimental results on TaihuLight show that 1) compared with the default implementations of OpenFOAM, the RNPCG and the LDIC on a single-core group of SW26010 can achieve a maximum speedup of 8.9X and 3.1X, respectively; 2) the scalable RNPCG can outperform the standard PCG both in the strong and the weak scaling up to 66,560 cores. James Lin 0001, Minhua Wen, Delong Meng, Xin Liu 0020, Akira Nukada, Satoshi Matsuoka |
CCGrid | 6 |
| 2018 | Efficient Algorithms for the Summed Area Tables Primitive on GPUsabstractTwo-dimensional Summed Area Tables (SAT) is a fundamental primitive used in image processing and machine learning applications. We present a collection of optimization methods for computing SAT on CUDA-enabled GPUs. Conventional approaches rely on computing the prefix sum in one dimension in parallel, transposing the matrix, then computing the prefix sum for the other dimension in parallel. Additionally, conventional methods use the scratchpad memory as cache. We propose a collection of algorithms that are scalable with respect to problem size. We use the register cache technique instead of the scratchpad memory and also employ a naive serial scan on the thread level for computing the prefix sum for one of the dimensions. Using a novel transpose-in-registers method we increase the inter-thread parallelism and outperform conventional SAT implementations. In addition, we significantly reduce both the communication between threads and the number of arithmetic instructions. On an Nvidia Pascal P100 GPU and Volta V100, our evaluations demonstrate that our implementations outperform state of the art libraries and yield up to 2.3x and 3.2x speedup over OpenCV and Nvidia NPP libraries, respectively. Peng Chen 0035, Mohamed Wahib, Shin'ichiro Takizawa, Ryousei Takano, Satoshi Matsuoka |
CLUSTER | 5 |
| 2018 | Accelerating Deep Learning Frameworks with Micro-BatchesabstractcuDNN is a low-level library that provides GPU kernels frequently used in deep learning. Specifically, cuDNN implements several equivalent convolution algorithms, whose performance and memory footprint may vary considerably, depending on the layer dimensions. When an algorithm is automatically selected by cuDNN, the decision is performed on a per-layer basis, and thus it often resorts to slower algorithms that fit the workspace size constraints. We present u-cuDNN, a thin wrapper library for cuDNN that transparently divides layers' mini-batch computation into multiple micro-batches, both on a single GPU and a heterogeneous set of GPUs. Based on Dynamic Programming and Integer Linear Programming (ILP), u-cuDNN enables faster algorithms by decreasing the workspace requirements. At the same time, u-cuDNN does not decrease the accuracy of the results, effectively decoupling statistical efficiency from the hardware efficiency. We demonstrate the effectiveness of u-cuDNN for the Caffe and TensorFlow frameworks, achieving speedups of 1.63x for AlexNet and 1.21x for ResNet-18 on the P100-SXM2 GPU. We also show that u-cuDNN achieves speedups of up to 4.54x, and 1.60x on average for DeepBench's convolutional layers on the V100-SXM2 GPU. In a distributed setting, u-cuDNN attains a speedup of 2.20x when training ResNet-18 on a heterogeneous GPU cluster over a single GPU. These results indicate that using micro-batches can seamlessly increase the performance of deep learning, while maintaining the same overall memory footprint. Yosuke Oyama, Tal Ben-Nun, Torsten Hoefler, Satoshi Matsuoka |
CLUSTER | 4 |
| 2018 | Predicting Performance Using Collaborative FilteringabstractPerformance prediction of parallel applications across systems becomes increasingly important in today's diverse computing environments. A wide range of choices in execution platforms pose new challenges to researchers in choosing a system which best fits their workloads and administrators in scheduling applications to the best performing systems. While previous studies have employed simulation-or profile-based prediction approaches, such solutions are time-consuming to be deployed on multiple platforms. To address this problem, we use two collaborative filtering techniques to build analytical models which can quickly and accurately predict the performance of workloads across different multicore systems. The first technique leverages information gained from performance observed for certain applications on a subset of systems and use it to discover similarities among applications as well as systems. The second collaborative filtering based model learns latent features of systems and workloads automatically and use these features to characterize the performance of applications on different platforms. We evaluated both the methods using 30 workloads chosen from NAS Parallel Benchmarks, BOTS and Rodinia benchmarking suites on ten different systems. Our results show that such collaborative filtering methods can make predictions with RMSE as low as 0.6 and with an average RMSE of 1.6. Shweta Salaria, Aleksandr Drozd, Artur Podobas, Satoshi Matsuoka |
CLUSTER | 4 |
| 2018 | Combined Spatial and Temporal Blocking for High-Performance Stencil Computation on FPGAs Using OpenCLabstractRecent developments in High Level Synthesis tools have attracted software programmers to accelerate their high-performance computing applications on FPGAs. Even though it has been shown that FPGAs can compete with GPUs in terms of performance for stencil computation, most previous work achieve this by avoiding spatial blocking and restricting input dimensions relative to FPGA on-chip memory. In this work we create a stencil accelerator using Intel FPGA SDK for OpenCL that achieves high performance without having such restrictions. We combine spatial and temporal blocking to avoid input size restrictions, and employ multiple FPGA-specific optimizations to tackle issues arisen from the added design complexity. Accelerator parameter tuning is guided by our performance model, which we also use to project performance for the upcoming Intel Stratix 10 devices. On an Arria 10 GX 1150 device, our accelerator can reach up to 760 and 375 GFLOP/s of compute performance, for 2D and 3D stencils, respectively, which rivals the performance of a highly-optimized GPU implementation. Furthermore, we estimate that the upcoming Stratix 10 devices can achieve a performance of up to 3.5 TFLOP/s and 1.6 TFLOP/s for 2D and 3D stencil computation, respectively. Hamid Reza Zohouri, Artur Podobas, Satoshi Matsuoka |
FPGA | 3 |
| 2018 | Adaptive Pattern Matching with Reinforcement Learning for Dynamic GraphsabstractGraph pattern matching algorithms to handle million-scale dynamic graphs are widely used in many applications such as social network analytics and suspicious transaction detections from financial networks. On the other hand, the computation complexity of many graph pattern matching algorithms is expensive, and it is not affordable to extract patterns from million-scale graphs. Moreover, most real-world networks are time-evolving, updating their structures continuously, which makes it harder to update and output newly matched patterns in real time. Many incremental graph pattern matching algorithms which reduce the number of updates have been proposed to handle such dynamic graphs. However, it is still challenging to recompute vertices in the incremental graph pattern matching algorithms in a single process, and that prevents the real-time analysis. We propose an incremental graph pattern matching algorithm to deal with time-evolving graph data and also propose an adaptive optimization system based on reinforcement learning to recompute vertices in the incremental process more efficiently. Then we discuss the qualitative efficiency of our system with several types of data graphs and pattern graphs. We evaluate the performance using million-scale attributed and time-evolving social graphs. Our incremental algorithm is up to 10.1 times faster than an existing graph pattern matching and 1.95 times faster with the adaptive systems in a computation node than naive incremental processing. Hiroki Kanezashi, Toyotaro Suzumura, Dario Garcia-Gasulla, Min-hwan Oh, Satoshi Matsuoka |
HiPC | 5 |
| 2018 | Cambrian explosion of computing and big data in the post-moore eraabstractThe so-called "Moore's Law", by which the performance of the processors will increase exponentially by factor of 4 every 3 years or so, is slated to be ending in 10--15 year timeframe due to the lithography of VLSIs reaching its limits around that time, and combined with other physical factors. We are now embarking on a project to revolutionize the total system architectural stack in a holistic fashion in the Post-Moore era, from devices and hardware, abstracted by system software and programming models and languages, and optimized according to the device characteristics with new algorithms and applications that exploit them. Such systems will have multitudes of varieties according to the matching characteristics of applications to the underlying architecture, leading to what can be metaphorically described as Cambrian Explosion of computing systems. The diverse elements of such systems will be interconnected with next-generation terabit optics and networks, allowing metropolitan-scale computing infrastructure that would truly realize high performance parallel and distributed computing. However, which algorithms and applications would benefit the most from such future computing, given that some physical constants, e.g., communication latency, cannot be improved? We speculate on some of the scenarios that would change the nature of current Cloud-centric infrastructures towards the Post-Moore era. Satoshi Matsuoka |
HPDC | 1 |
| 2018 | Explorations of Data Swapping on Burst BufferabstractBurst buffers have been widely deployed in many supercomputers to absorb bursty I/O and accelerate I/O performance. Previous work has shown that with burst buffer systems, I/O operations from computer nodes can be greatly accelerated. While the lack of data swapping supports on burst buffer leads to under-utilization and application failure issues. In addition, the effects of data replacement algorithms on application performance and the suitability of each algorithm for the target application are unclear. In this paper, we address these challenges by simulating data swapping on burst buffers with different data replacement strategies. Trace logs from a set of real-world HPC applications are used with different data replacement algorithms to show the behavior of representative HPC applications. From the results, we found that most HPC applications can still achieve full performance when using a buffer size that is far less than the total access space of the application, which can lead to a huge reduction on the required capacity for burst buffer. Moreover, we found that data replacement algorithms can have significant impact on the application performance. Our finding show the importance of having data swapping in reducing the required capacity and guide the future design of the usage of burst buffer. Kento Sato, Satoshi Matsuoka |
ICPADS | 3 |
| 2018 | Interference between I/O and MPI Traffic on Fat-tree NetworksabstractNetwork congestion arising from simultaneous data transfers can be a significant performance bottleneck for many applications, especially when network resources are shared by multiple concurrently running jobs. Many studies have focused on the impact of network congestion on either MPI performance or I/O performance but the interaction between MPI and I/O traffic is rarely studied and not well understood. In this paper, we analyze and characterize the interference between MPI and I/O traffic on fat-tree networks, highlighting the role of important factors such as message sizes, communication intervals, and job sizes. We also investigate several strategies for reducing MPI-I/O interference, and the benefits and trade-offs of each approach for different scenarios. Kevin A. Brown, Satoshi Matsuoka, Martin Schulz 0001, Abhinav Bhatele |
ICPP | 3 |
| 2018 | Efficient Solving of Scan Primitive on Multi-GPU SystemsabstractGPUs fulfill high computation demands, but it is necessary to develop code carefully, selecting algorithms well suited to the GPU architecture and applying different optimizations. This article presents a GPU-suitable algorithm and a tuning strategy for performing the scan primitive over large problem sizes in CUDA. This tuning strategy defines different performance premises to find the GPU execution parameters that maximize performance. Taking these premises into consideration, we easily develop the kernels using CUDA skeletons to ensure efficiency and portability. Based on this, we describe an optimal proposal analyzed over different multiple GPU environments, the first multiple-GPU batch scan proposal to the best of our knowledge. The resulting implementations outperform other well-known libraries in most cases, such as CUDPP, ModernGPU, Thrust, CUB and LightScan. Adrián Pérez Diéguez, Margarita Amor, Ramón Doallo, Akira Nukada, Satoshi Matsuoka |
IPDPS | 5 |
| 2018 | DRAGON: breaking GPU memory capacity limits with direct NVM access
Pak Markthub, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Satoshi Matsuoka |
SC | 5 |
| 2018 | Evaluating the SW26010 many-core processor with a micro-benchmark suite for performance optimizations
James Lin 0001, Zhigeng Xu, Linjin Cai, Akira Nukada, Satoshi Matsuoka |
Parallel Comput. | 5 |
| 2017 | Being "BYTES-oriented" in HPC leads to an open big data/AI ecosystem and further advances into the post-moore eraabstractWith rapid rise and increase of Big Data and AI as a new breed of high-performance workloads on supercomputers, we need to accommodate them at scale, traditional simulation-based HPC and BD/AI will converge. Our TSUBAME3 supercomputer at Tokyo Institute of Technology became online in Aug. 2017, and became the greenest supercomputer in the world on the Green 500 ranking at 14.11 GFlops/W; the other aspect of TSUBAME3, is to embody various Data or “BYTES-oriented” features to allow for HPC to BD/AI convergence at scale, including significant scalable horizontal bandwidth as well as support for deep memory hierarchy and capacity, along with high flops in low precision arithmetic for deep learning. Furthermore, TSUBAM3's technologies will be commoditized to construct one of the world's largest BD/AI focused and “open-source” cloud infrastructure called ABCI (AI-Based Bridging Cloud Infrastructure), hosted by AIST-AIRC (AI Research Center), the largest public funded AI research center in Japan. The performance of the machine is slated to be several hundred AI-Petaflops for machine learning; the true nature of the machine however, is its BYTES-oriented, optimization acceleration in the memory hiearchy, I/O, the interconnect etc, for high-performance BD/AI. ABCI will be online Spring 2018 and its archiecture, software, as well as the datacenter infrastructure design itself will be made open to drive rapid adoptions and improvements by the community, unlike the concealed cloud infrastructures of today. Finally, transcending from FLOPS-centric mindset to being BYTES-oriented will be one of the key solutions to the upcoming “end-of-Moore's law” in the mind 2020s, upon which FLOPS increase will cease and BYTES-oriented advances will be the new source of performance increases over time in general for any computing. Satoshi Matsuoka |
IEEE BigData | 1 |
| 2017 | Evaluation of HPC-Big Data Applications Using Cloud PlatformsabstractThe path to HPC-Big Data convergence has resulted in numerous researches that demonstrate the performance trade-off between running applications on supercomputers and cloud platforms. Previous studies typically focus either on scientific HPC benchmarks or previous cloud configurations, failing to consider all the new opportunities offered by current cloud offerings. We present a comparative study of the performance of representative big data benchmarks, or "Big Data Ogres", and HPC benchmarks running on supercomputer and cloud. Our work distinguishes itself from previous studies in a way that we explore the latest generation of compute-optimized Amazon Elastic Compute Cloud instances, C4 for our experimentation on cloud. Our results reveal that Amazon C4 instances with increased compute performance and low variability in results make EC2-based cluster feasible for scientific computing and its applications in simulations, modeling and analysis. Shweta Salaria, Kevin A. Brown, Hideyuki Jitsumoto, Satoshi Matsuoka |
CCGrid | 4 |
| 2017 | Co-locating Graph Analytics and HPC ApplicationsabstractWe evaluate the on-node interference caused when co-locating traditional high-performance computing applications with a big-data application. Using kernel benchmarks from the NPB suite and a state-of-art graph analytics code, we explore different process placements and effects they have on application performance. Our results show that the most memory intensive HPC application (MG) experienced the highest performance variation during co-location. Kevin A. Brown, Satoshi Matsuoka |
CLUSTER | 2 |
| 2017 | Evaluating high-level design strategies on FPGAs for high-performance computingabstractField-Programmable Gate Arrays (FPGAs) are gaining considerable momentum in mainstream high-performance systems in recent years due to their flexibility and low power consumption. Still, FPGAs remain largely unavailable to software programmers due to programming and debugging difficulties that are inherent to standard Hardware Description Languages. The performance that hardware-oblivious software engineers can expect from migrating legacy code to FPGAs remains shrouded in mystery. To gain insight on how to use FPGAs in high-performance computing, we created four different systems and evaluated them using benchmarks from the Rodinia benchmark suite. The systems we evaluated were diverse in both programming model and generality, and range from a custom-built 30-core manycore system to FSM-based accelerators using LegUP and deep data-flow pipelines using Intel FPGA SDK for OpenCL. We found that the original version of LegUp does not achieve very good performance out of the box; still, with some non-trivial modification in the architecture, we improved its performance by up to 10 times. Despite this, we found Intel FPGA SDK for OpenCL to perform up to two orders of magnitude faster than LegUp. We also found our general-purpose manycore system to have comparable performance with LegUp. Artur Podobas, Hamid Reza Zohouri, Naoya Maruyama, Satoshi Matsuoka |
FPL | 4 |
| 2017 | Evaluating high-level design strategies on FPGAs for high-performance computingabstractField-Programmable Gate Arrays (FPGAs) are gaining considerable momentum in mainstream high-performance systems in recent years due to their flexibility and low power consumption. Still, FPGAs remain largely unavailable to software programmers due to programming and debugging difficulties that are inherent to standard Hardware Description Languages. The performance that hardware-oblivious software engineers can expect from migrating legacy code to FPGAs remains shrouded in mystery. To gain insight on how to use FPGAs in high-performance computing, we created four different systems and evaluated them using benchmarks from the Rodinia benchmark suite. The systems we evaluated were diverse with respect to both programming model and generality, and range from a custom-built 30-core manycore system to FSM-based accelerators using LegUP and deep data-flow pipelines using Intel FPGA SDK for OpenCL. We found that the original version of LegUp does not achieve very good performance out of the box; still, with some non-trivial modification in the architecture, we improved its performance by up to 10 times. Despite this, we found Intel FPGA SDK for OpenCL to perform up to two orders of magnitude faster than LegUp. We also found our general-purpose manycore system to have comparable performance with LegUp. Artur Podobas, Hamid Reza Zohouri, Naoya Maruyama, Satoshi Matsuoka |
FPL | 4 |
| 2017 | Designing and accelerating spiking neural networks using OpenCL for FPGAsabstractSpiking Neural Networks (SNNs) are artificial networks inspired by the biological brain and are used to study various aspects of brainlike-computing. SNNs are typically computed using general-purpose processors. However, with the recent maturity of High-Level Synthesis (HLS) tools, system developers can now accelerate their design on Field-Programmable Gate-Arrays (FPGA) with small to minimal impact in productivity. The present work focuses on accelerating SNNs on FPGAs. We show how OpenCL for FPGAs can be leveraged to accelerate two different but equally important neuron models, including axons and synapses, and empirically quantify the performance of our designs, reaching performance up to 2.25 GSpikes/second. Artur Podobas, Satoshi Matsuoka |
FPT | 2 |
| 2017 | Asynchronous, Data-Parallel Deep Convolutional Neural Network Training with Linear Prediction Model for Parameter Transition
Ikuro Sato, Ryo Fujisaki, Yosuke Oyama, Akihiro Nomura 0002, Satoshi Matsuoka |
ICONIP (2) | 5 |
| 2017 | Optimizations of Two Compute-Bound Scientific Kernels on the SW26010 Many-Core ProcessorabstractThe home-grown SW26010 many-core processor enabled the production of China's first independently developed number-one ranked supercomputer - the Sunway TaihuLight. The design of the limited off-chip memory bandwidth, however, renders the SW26010 a highly memory-bound processor. To compensate for this limitation, the processor was designed with a unique hardware feature, "Register Level Communication" (RLC), to share register data among its 8 × 8 computing processing elements (CPEs) via a 2D onchip network. Such a radical architecture has sparked global researchers' concerns regarding the programming challenges this may cause. To address these concerns, we adopted two compute-bound scientific kernels as benchmarks to identify the potential programming challenges. The first kernel is doubleprecision general matrix-multiplication (DGEMM). An RLCfriendly algorithm was designed for this kernel to reuse the data that already reside in the registers of 64 CPEs. This novel optimization enables the kernel to achieve up to 88.7% efficiency in one core group of the SW26010. This paper reveals, for the first time, the details of how the highly efficient DGEMM is implemented on the home-grown processor. The second kernel that we used is N-body. Due to the inefficient hardware support for transcendental operations on the SW26010, we replaced the reciprocal square root (rsqrt) instruction of N-body with a software routine to tackle the problem. Based on the programming challenges identified through these two optimized kernels, we proposed a three-level programming guideline for the SW26010. The paper concludes with our crucial finding that the critical step towards bridging the ninja performance gap on the SW26010 is to design an RLC-friendly algorithm to increase arithmetic intensity. James Lin 0001, Zhigeng Xu, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
ICPP | 5 |
| 2017 | High-Performance and Memory-Saving Sparse General Matrix-Matrix Multiplication for NVIDIA Pascal GPUabstractSparse general matrix-matrix multiplication (SpGEMM) is one of the key kernels of preconditioners such as algebraic multigrid method or graph algorithms. However, the performance of SpGEMM is quite low on modern processors due to random memory access to both input and output matrices. As well as the number and the pattern of non-zero elements in the output matrix, important for achieving locality, are unknown before the execution. Moreover, the state-of-the-art GPU implementations of SpGEMM requires large amounts of memory for temporary results, limiting the matrix size computable on fast GPU device memory. We propose a new fast SpGEMM algorithm requiring small amount of memory and achieving high performance. Calculation of the pattern and value in output matrix is optimized by using GPU's on-chip shared memory and a hash table. Additionally, our algorithm launches multiple kernels running concurrently to improve the utilization of GPU resources. The kernels for the calculation of each row of output matrix are chosen based on the number of non-zero elements. Performance evaluation using matrices from the Sparse Matrix Collection of University Florida on NVIDIA's Pascal generation GPU shows that our approach achieves speedups of up to x4.3 in single precision and x4.4 in double precision compared to existing SpGEMM libraries. Furthermore, the memory usage is reduced by 14.7% in single precision and 10.9% in double precision on average, allowing larger matrices to be computed. Yusuke Nagasaka, Akira Nukada, Satoshi Matsuoka |
ICPP | 3 |
| 2017 | Efficient Breadth-First Search on Massively Parallel and Distributed-Memory MachinesabstractThere are many large-scale graphs in real world such as Web graphs and social graphs. The interest in large-scale graph analysis is growing in recent years. Breadth-First Search (BFS) is one of the most fundamental graph algorithms used as a component of many graph algorithms. Our new method for distributed parallel BFS can compute BFS for one trillion vertices graph within half a second, using large supercomputers such as the K-Computer. By the use of our proposed algorithm, the K-Computer was ranked 1st in Graph500 using all the 82,944 nodes available on June and November 2015 and June 2016 38,621.4 GTEPS. Based on the hybrid BFS algorithm by Beamer (Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing Workshops and PhD Forum, IPDPSW ’13, IEEE Computer Society, Washington, 2013 ), we devise sets of optimizations for scaling to extreme number of nodes, including a new efficient graph data structure and several optimization techniques such as vertex reordering and load balancing. Our performance evaluation on K-Computer shows that our new BFS is 3.19 times faster on 30,720 nodes than the base version using the previously known best techniques. Koji Ueno, Toyotaro Suzumura, Naoya Maruyama, Katsuki Fujisawa, Satoshi Matsuoka |
Data Sci. Eng. | 5 |
| 2016 | Predicting statistics of asynchronous SGD parameters for a large-scale distributed deep learning system on GPU supercomputersabstractMany studies have shown that Deep Convolutional Neural Networks (DCNNs) exhibit great accuracies given large training datasets in image recognition tasks. Optimization technique known as asynchronous mini-batch Stochastic Gradient Descent (SGD) is widely used for deep learning because it gives fast training speed and good recognition accuracies, while it may increases generalization error if training parameters are in inappropriate ranges. We propose a performance model of a distributed DCNN training system called SPRINT that uses asynchronous GPU processing based on mini-batch SGD. The model considers the probability distribution of mini-batch size and gradient staleness that are the core parameters of asynchronous SGD training. Our performance model takes DCNN architecture and machine specifications as input parameters, and predicts time to sweep entire dataset, mini-batch size and staleness with 5%, 9% and 19% error in average respectively on several supercomputers with up to thousands of GPUs. Experimental results on two different supercomputers show that our model can steadily choose the fastest machine configuration that nearly meets a target mini-batch size. Yosuke Oyama, Akihiro Nomura 0002, Ikuro Sato, Hiroki Nishimura, Yukimasa Tamatsu, Satoshi Matsuoka |
IEEE BigData | 6 |
| 2016 | I/O chunking and latency hiding approach for out-of-core sorting acceleration using GPU and flash NVMabstractWe propose an out-of-core sorting acceleration technique, called xtr2sort, that deals with multi-level memory hierarchies of device memory (GPU), host memory (CPU), and semi-external non-volatile memory (Flash NVM) for leveraging the high computational performance and memory bandwidth of GPUs, while offloading bandwidth-oblivious operations onto semi-external memory in order to significantly increasing the memory capacity available for the sort data, well beyond the that of the GPU as well as of the CPU. xtr2sort splits the input records into several chunks to fit in GPU device memory and overlaps (1) I/O operations between semi-external and host memory, (2) data transfers between host and device memory, and (3) sorting on the GPU device in an asynchronous manner for hiding latency. Experimental results show that xtr2sort can sort records up to 256 times larger than is possible with in-core GPU sorting and 16 times larger than is possible with in-core CPU sorting. xtr2sort also achieves 4.39 times faster than out-of-core CPU sorting using 72 threads on 204.8 giga records with int32_t, even though the input records could not fit in the host memory, let alone the GPU device memory. These results indicate that I/O chunking and latency hiding/overlapping maintains sorting performance, despite slow Flash NVM performance, by utilizing GPUs along with good algorithms. Such an approach is viable for accelerating future computing systems with deep memory hierarchies. Hitoshi Sato, Ryo Mizote, Satoshi Matsuoka, Hirotaka Ogawa |
IEEE BigData | 3 |
| 2016 | Extreme scale breadth-first search on supercomputersabstractBreadth-First Search(BFS) is one of the most fundamental graph algorithms used as a component of many graph algorithms. Our new method for distributed parallel BFS can compute BFS for one trillion vertices graph within half a second, using large supercomputers such as the K-Computer. By the use of our proposed algorithm, the K-Computer was ranked 1st in Graph500 using all the 82,944 nodes available on June and November 2015 and June 2016 38,621.4 GTEPS. Based on the hybrid-BFS algorithm by Beamer[3], we devise sets of optimizations for scaling to extreme number of nodes, including a new efficient graph data structure and optimization techniques such as vertex reordering and load balancing. Performance evaluation on the K shows our new BFS is 3.19 times faster on 30,720 nodes than the base version using the previously-known best techniques. Koji Ueno, Toyotaro Suzumura, Naoya Maruyama, Katsuki Fujisawa, Satoshi Matsuoka |
IEEE BigData | 5 |
| 2016 | Serving More GPU Jobs, with Low Penalty, Using Remote GPU Execution and MigrationabstractRemote GPU execution has been proven to increase GPU occupancy and reduce job waiting time in multi-GPU batch-queue systems, by allowing jobs to utilize remote GPUs when there are not enough unoccupied local GPUs available. However, for GPU communication intensive applications, remote GPU communication overhead may account for more than 70% of the applications' execution times. The need for using a remote GPU exists when there are not enough local GPUs available on a node assigned to the job, but a local GPU could become available afterward. We propose mrCUDA, a middleware for migrating execution on a remote GPU to a local GPU on-demand. Our evaluation shows that for long-running jobs mrCUDA overhead accounts for less than 1% of their total execution times. In addition, by applying mrCUDA to the first-come-first-serve (FCFS) job scheduling algorithm, we could reduce job lifetimes (waiting + execution times) as much as 30% on average without changing the scheduling policy. Pak Markthub, Akihiro Nomura 0002, Satoshi Matsuoka |
CLUSTER | 3 |
| 2016 | Word Embeddings, Analogies, and Machine Learning: Beyond king - man + woman = queenabstractSolving word analogies became one of the most popular benchmarks for word embeddings on the assumption that linear relations between word pairs (such as king:man :: woman:queen) are indicative of the quality of the embedding. We question this assumption by showing that the information not detected by linear offset may still be recoverable by a more sophisticated search method, and thus is actually encoded in the embedding. The general problem with linear offset is its sensitivity to the idiosyncrasies of individual words. We show that simple averaging over multiple word pairs improves over the state-of-the-art. A further improvement in accuracy (up to 30% for some embeddings and relations) is achieved by combining cosine similarity with an estimation of the extent to which a candidate answer belongs to the correct word class. In addition to this practical contribution, this work highlights the problem of the interaction between word embeddings and analogy retrieval algorithms, and its implications for the evaluation of word embeddings and the use of analogies in extrinsic tasks. Aleksandr Drozd, Anna Rogers, Satoshi Matsuoka |
COLING | 3 |
| 2016 | Routing on the Dependency Graph: A New Approach to Deadlock-Free High-Performance RoutingabstractLossless interconnection networks are omnipresent in high performance computing systems, data centers and network-on-chip architectures. Such networks require efficient and deadlock-free routing functions to utilize the available hardware. Topology-aware routing functions become increasingly inapplicable, due to irregular topologies, which either are irregular by design or as a result of hardware failures. Existing topology-agnostic routing methods either suffer from poor load balancing or are not bounded in the number of virtual channels needed to resolve deadlocks in the routing tables. We propose a novel topology-agnostic routing approach which implicitly avoids deadlocks during the path calculation instead of solving both problems separately. We present a model implementation, called Nue, of a destination-based and oblivious routing function. Nue routing heuristically optimizes the load balancing while enforcing deadlock-freedom without exceeding a given number of virtual channels, which we demonstrate based on the InfiniBand architecture. Jens Domke, Torsten Hoefler, Satoshi Matsuoka |
HPDC | 3 |
| 2016 | Tapas: An Implicitly Parallel Programming Framework for Hierarchical N-Body AlgorithmsabstractTapas is our new C++ programming framework for hierarchical algorithms such as N-body, on large scale heterogeneous supercomputers. Although N-body and their variants are widely used in scientific applications, their correct implementations are often difficult on such modern machines, as the algorithms are irregular, complex, and involve explicit task parallel programming over distributed nodes. Encapsulating the complexities in a library or a framework has been challenging due to irregular data access over massively distributed memory. Tapas solves this by converting the users clean implicit-style parallel program into an inspector-executor style code on heterogeneous multi-core, multi-node environment solely by the use of C++ template metaprogramming. A prototype implementation of the Fast Multipole Method on Tapas demonstrates a comparable performance and scaling as ExaFMM, the fastest hand-tuned implementation of FMM, as well as efficient usage of hundreds of GPUs. Specifically, the serial performance is 95% of ExaFMM, whereas the distributed-memory strong-scaling evaluation using up to 1500 CPU cores demonstrates 64% to 81% of the ExaFMM performance. The multi-GPU version of the Tapas-based FMM achieves a 5.15x speedup when executed on 100 nodes of TSUBAME2.5 with 300 GPUs. Keisuke Fukuda, Motohiko Matsuda, Naoya Maruyama, Rio Yokota, Kenjiro Taura, Satoshi Matsuoka |
ICPADS | 6 |
| 2016 | CloudBB: Scalable I/O Accelerator for Shared Cloud StorageabstractCurrent shared cloud storage cannot provide sufficient I/O throughput for data-intensive HPC applications. Moreover, the consistency policy used in most shared cloud storage can cause parallel I/O applications to fail due to unexpected file inconsistencies. In order to resolve these problems, we propose a novel fast, scalable and fault tolerant filesystem called CloudBB (Cloud-based Burst Buffer). Unlike conventional filesystems, CloudBB creates an on-demand two-level hierarchical storage system and caches popular files to accelerate I/O performance. Since CloudBB supports multiple metadata servers, CloudBB is also highly scalable. In addition, by using file replication, failure detection and recovery techniques, CloudBB is resilient to failures. Furthermore, we implement CloudBB by using FUSE so that existing applications can run seamlessly and benefit from all of the CloudBB's capabilities without code modification. To validate the effectiveness of CloudBB, we evaluate performance of real data-intensive HPC applications in Amazon EC2/S3. The results show CloudBB improves performance by up to 28.7 times while reducing cost by up to 94.7% compared to the ones without CloudBB. Kento Sato, Satoshi Matsuoka |
ICPADS | 3 |
| 2016 | Evaluating and optimizing OpenCL kernels for high performance computing with FPGAsabstractWe evaluate the power and performance of the Rodinia benchmark suite using the Altera SDK for OpenCL targeting a Stratix V FPGA against a modern CPU and GPU. We study multiple OpenCL kernels per benchmark, ranging from direct ports of the original GPU implementations to loop-pipelined kernels specifically optimized for FPGAs. Based on our results, we find that even though OpenCL is functionally portable across devices, direct ports of GPU-optimized code do not perform well compared to kernels optimized with FPGA-specific techniques such as sliding windows. However, by exploiting FPGA-specific optimizations, it is possible to achieve up to 3.4x better power efficiency using an Altera Stratix V FPGA in comparison to an NVIDIA K20c GPU, and better run time and power efficiency in comparison to CPU. We also present preliminary results for Arria 10, which, due to hardened FPUs, exhibits noticeably better performance compared to Stratix V in floating-point-intensive benchmarks. Hamid Reza Zohouri, Naoya Maruyama, Aaron Smith, Motohiko Matsuda, Satoshi Matsuoka |
SC | 5 |
| 2016 | Special Issue on Cluster Computing
Michela Taufer, Pavan Balaji, Satoshi Matsuoka |
Parallel Comput. | 3 |
| 2016 | GPU-Accelerated Large-Scale Distributed Sorting Coping with Device Memory CapacityabstractSplitter-based parallel sorting algorithms are known to be highly efficient for distributed sorting due to their low communication complexity. Although using GPU accelerators could help to reduce the computation cost in general, their effectiveness in distributed sorting algorithms remains unclear. We investigate applicability of using GPU devices to the splitter-based algorithms and extend HykSort, an existing splitter-based algorithm by offloading costly computation phases to GPUs. To cope with the volumes of data exceeding the GPU memory capacity, out-of-core local sort is used with small overhead about 7.5 percent when the data size is tripled. We evaluate the performance of our implementation on the TSUBAME2.5 supercomputer that comprises over 4,000 NVIDIA K20x GPUs. Weak scaling analysis shows 389 times speedup with 0.25 TB/s throughput when sorting 4 TB of 64 bit integer values on 1,024 nodes compared to running on one node; this is 1.40 times faster than the reference CPU implementation. Detailed analysis however reveals that the performance is mostly bottlenecked by the CPU-GPU host-to-device bandwidth. With orders of magnitude improvements announced for next generation GPUs, the performance boost will be tremendous in accordance with other successful GPU accelerations. Hideyuki Shamoto, Koichi Shirahata, Aleksandr Drozd, Hitoshi Sato, Satoshi Matsuoka |
IEEE Trans. Big Data | 5 |
| 2015 | Characterizing MPI and Hybrid MPI+Threads Applications at Scale: Case Study with BFSabstractWith the increasing prominence of many-core architectures and decreasing per-core resources on large supercomputers, a number of applications developers are investigating the use of hybrid MPI+threads programming to utilize computational units while sharing memory. An MPI-only model that uses one MPI process per system core is capable of effectively utilizing the processing units, but it fails to fully utilize the memory hierarchy and relies on fine-grained internodes communication. Hybrid MPI+threads models, on the other hand, can handle internodes parallelism more effectively and alleviate some of the overheads associated with internodes communication by allowing more coarse-grained data movement between address spaces. The hybrid model, however, can suffer from locking and memory consistency overheads associated with data sharing. In this paper, we use a distributed implementation of the breadth-first search algorithm in order to understand the performance characteristics of MPI-only and MPI+threads models at scale. We start with a baseline MPI-only implementation and propose MPI+threads extensions where threads independently communicate with remote processes while cooperating for local computation. We demonstrate how the coarse-grained communication of MPI+threads considerably reduces time and space overheads that grow with the number of processes. At large scale, however, these overheads constitute performance barriers for both models and require fixing the root causes, such as the excessive polling for communication progress and inefficient global synchronizations. To this end, we demonstrate various techniques to reduce such overheads and show performance improvements on up to 512K cores of a Blue Gene/Q system. Abdelhalim Amer, Huiwei Lu, Pavan Balaji, Satoshi Matsuoka |
CCGRID | 4 |
| 2015 | Modeling Gather and Scatter with Hardware Performance Counters for Xeon PhiabstractIntel Initial Many-Core Instructions (IMCI) for Xeon Phi introduces hardware-implemented Gather and Scatter (G/S) load/store contents of SIMD registers from/to non-contiguous memory locations. However, they can be one of key performance bottlenecks for Xeon Phi. Modelling G/S can provide insights to the performance on Xeon Phi, however, the existing solution needs a hand-written assembly implementation. Therefore, we modeled G/S with hardware performance counters which can be profiled by the tools like PAPI. We profiled Address Generation Interlock (AGI) events as the number of G/S, estimated the average latency of G/S with VPU_DATA_READ, and combined them to model the total latencies of G/S. We applied our model to the 3D 7-point stencil and the result showed G/S spent nearly 40% of total kernel time. We also validated the model by implementing a G/S- free version with intrinsics. The contribution of the work is a performance model for G/S built with hardware counters. We believe the model can be generally applicable to CPU as well. James Lin 0001, Akira Nukada, Satoshi Matsuoka |
CCGRID | 3 |
| 2015 | Efficient Execution of Multiple CUDA Applications Using Transparent Suspend, Resume and Migration
Taichiro Suzuki, Akira Nukada, Satoshi Matsuoka |
Euro-Par | 3 |
| 2015 | Hardware-Centric Analysis of Network Performance for MPI ApplicationsabstractAs the scale of high-performance computing systems increases, optimizing inter-process communication becomes more challenging while being critical for ensuring good performance. However, the hardware layer abstraction provided by MPI makes it difficult to study application communication performance over the network hardware, especially for collective operations. We present a new approach to network performance analysis based on exposing low-level communication metrics in a flexible manner and conducting hardware-centric analysis of these metrics. We show how low-level network metrics can be revealed using Open MPI's Peruse utility, without interfacing with the hardware layer. A lightweight profiler, ibprof, was developed to aggregate these metrics from message passing events at a cost of <;1% runtime overhead for communication in NPB kernel and application benchmarks. We also developed a flexible visualization module for the Boxfish analysis tool to analyze our communication profile over the physical topology of the network. Using case studies, we demonstrate how our approach can identify communication anomalies in network applications and guide performance optimization strategies. Kevin A. Brown, Jens Domke, Satoshi Matsuoka |
ICPADS | 3 |
| 2015 | Realizing Extremely Large-Scale Stencil Applications on GPU SupercomputersabstractThe problem of deepening memory hierarchy towards exascale is becoming serious for applications such as those based on stencil kernels, as it is difficult to satisfy both high memory bandwidth ad capacity requirements simultaneously. This is evident even today, where problem sizes of stencil-based applications on GPU supercomputers are limited by aggregated capacity of GPU device memory. Locality improvement techniques such as temporal blocking is known to preserve performance, but integrating the technique into existing stencil applications results in substantially higher programming cost, especially for complex applications and as a result are not typically utilized. We alleviate this problem with a run-time GPU-MPI process virtualization library we call HHRT that automates data movement across the memory hierarchy, and a systematic methodology to convert and optimize the code to accommodate temporal blocking. The proposed methodology has shown to significantly eases the adaptation of real applications, such as the whole-city airflow simulator embodying more than 12,000 lines of code; with careful tuning, we successfully maintain up to 85% performance even with problems whose footprint is four time larger than GPU device memory capacity, and scale to hundreds of GPUs on the TSUBAME2.5 supercomputer. Toshio Endo, Yuki Takasaki, Satoshi Matsuoka |
ICPADS | 3 |
| 2015 | Exploration of Lossy Compression for Application-Level Checkpoint/RestartabstractThe scale of high performance computing (HPC) systems is exponentially growing, potentially causing prohibitive shrinkage of mean time between failures (MTBF) while the overall increase in the I/O performance of parallel file systems will be far behind the increase in scale. As such, there have been various attempts to decrease the checkpoint overhead, one of which is to employ compression techniques to the checkpoint files. While most of the existing techniques focus on lossless compression, their compression rates and thus effectiveness remain rather limited. Instead, we propose a loss compression technique based on wavelet transformation for checkpoints, and explore its impact to application results. Experimental application of our loss compression technique to a production climate application, NICAM, shows that the overall checkpoint time including compression is reduced by 81%, while relative error remains fairly constant at approximately 1.2% on overall average of all variables of compressed physical quantities compared to original checkpoint without compression. Naoto Sasaki, Kento Sato, Toshio Endo, Satoshi Matsuoka |
IPDPS | 4 |
| 2015 | MPI+Threads: runtime contention and remediesabstractHybrid MPI+Threads programming has emerged as an alternative model to the “MPI everywhere” model to better handle the increasing core density in cluster nodes. While the MPI standard allows multithreaded concurrent communication, such flexibility comes with the cost of maintaining thread safety within the MPI implementation, typically implemented using critical sections. In contrast to previous works that studied the importance of critical-section granularity in MPI implementations, in this paper we investigate the implication of critical-section arbitration on communication performance. We first analyze the MPI runtime when multithreaded concurrent communication takes place on hierarchical memory systems. Our results indicate that the mutex-based approach that most MPI implementations use today can incur performance penalties due to unfair arbitration. We then present methods to mitigate these penalties with a first-come, first-served arbitration and a priority locking scheme that favors threads doing useful work. Through evaluations using several benchmarks and applications, we demonstrate up to 5-fold improvement in performance. Abdelhalim Amer, Huiwei Lu, Yanjie Wei, Pavan Balaji, Satoshi Matsuoka |
PPoPP | 5 |
| 2014 | NVM-based Hybrid BFS with memory efficient data structureabstractWe introduce a memory efficient implementation for the NVM-based Hybrid BFS algorithm that merges redundant data structures to a single graph data structure, while offloading infrequent accessed graph data on NVMs based on the detailed analysis of access patterns, and demonstrate extremely fast BFS execution for large-scale unstructured graphs whose size exceed the capacity of DRAM on the machine. Experimental results of Kronecker graphs compliant to the Graph500 benchmark on a 2-way INTEL Xeon E5-2690 machine with 256 GB of DRAM show that our proposed implementation can achieve 4.14 GTEPS for a SCALE31 graph problem with 231vertices and 235edges, whose size is 4 times larger than the size of graphs that the machine can accommodate only using DRAM with only 14.99 % performance degradation. We also show that the power efficiency of our proposed implementation achieves 11.8 MTEPS/W. Based on the implementation, we have achieved the 3rd and 4th position of the Green Graph500 list (2014 June) in the Big Data category. Keita Iwabuchi, Hitoshi Sato, Yuichiro Yasui, Katsuki Fujisawa, Satoshi Matsuoka |
IEEE BigData | 5 |
| 2014 | Large-scale distributed sorting for GPU-based heterogeneous supercomputersabstractSplitter-based parallel sorting algorithms are known to be highly efficient for distributed sorting due to their low communication complexity. Although using GPU accelerators could help to reduce the computation cost in general, their effectiveness in distributed sorting algorithms on large-scale heterogeneous GPU-based systems remains unclear. We investigate applicability of using GPU devices to the splitter-based algorithms and extend HykSort, an existing splitter-based algorithm by offloading costly computation phases to GPUs. We also handle GPU memory overflows by introducing an iterative approach which sorts multiple chunks and merges them into one array. We evaluate the performance of our implementation with local sort acceleration on the TSUBAME2.5 supercomputer that comprises over 4000 NVIDIA K20x GPUs. Performance evaluation of weak scaling shows that we achieve 389 times speedup with 0.25TB/s throughput when sorting 4TB 64bit integer on 1024 nodes compared to running on 1 node; on the other hand, for CPU vs. GPU comparison, our implementation achieves only 1.40 times speedup using 1024 nodes. Detailed analysis however reveals that the limitation is almost entirely due to the bottleneck in CPU-GPU host-to-device bandwidth. With orders of magnitude improvements planned for next generation GPUs, the performance boost will be tremendous in accordance with other successful GPU accelerations. Hideyuki Shamoto, Koichi Shirahata, Aleksandr Drozd, Hitoshi Sato, Satoshi Matsuoka |
IEEE BigData | 5 |
| 2014 | A User-Level InfiniBand-Based File System and Checkpoint Strategy for Burst BuffersabstractCheckpoint/Restart is an indispensable fault tolerance technique commonly used by high-performance computing applications that run continuously for hours or days at a time. However, even with state-of-the-art checkpoint/restart techniques, high failure rates at large scale will limit application efficiency. To alleviate the problem, we consider using burst buffers. Burst buffers are dedicated storage resources positioned between the compute nodes and the parallel file system, and this new tier within the storage hierarchy fills the performance gap between node-local storage and parallel file systems. With burst buffers, an application can quickly store checkpoints with increased reliability. In this work, we explore how burst buffers can improve efficiency compared to using only node-local storage. To fully exploit the bandwidth of burst buffers, we develop a user-level Infini Band-based file system (IBIO). We also develop performance models for coordinated and uncoordinated checkpoint/restart strategies, and we apply those models to investigate the best checkpoint strategy using burst buffers on future large-scale systems. Kento Sato, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 7 |
| 2014 | How file access patterns influence interference among cluster applicationsabstractOn large-scale clusters, tens to hundreds of applications can simultaneously access a parallel file system, leading to contention and in its wake to degraded application performance. However, the degree of interference depends on the specific file access pattern. On the basis of synchronized time-slice profiles, we compare the interference potential of different file access patterns. We consider both micro-benchmarks, to study the effects of certain patterns in isolation, and realistic applications to gauge the severity of such interference under production conditions. In particular, we found that writing large files simultaneously with small files can slow down the latter at small chunk sizes but the former at larger chunk sizes. We further show that such effects can seriously affect the runtime of real applications-up to a factor of five in one instance. In the future, both our insights and profiling techniques can be used to automatically classify the interference potential between applications and to adjust scheduling decisions accordingly. Chih-Song Kuo, Aamer Shah, Akihiro Nomura 0002, Satoshi Matsuoka, Felix Wolf 0001 |
CLUSTER | 4 |
| 2014 | Out-of-core GPU memory management for MapReduce-based large-scale graph processingabstractGPUs can accelerate edge scan performance of graph processing applications; however, the capacity of device memory on GPUs limits the size of graph to process, whereas efficient techniques to handle GPU memory overflows, including overflow detection and performance analysis in large-scale systems, are not well investigated. To address the problem, we propose a MapReduce-based out-of-core GPU memory management technique for processing large-scale graph applications on heterogeneous GPU-based supercomputers. Our proposed technique automatically handles memory overflows from GPUs by dynamically dividing graph data into multiple chunks and overlaps CPU-GPU data transfer and computation on GPUs as much as possible. Our experimental results on TSUBAME2.5 using 1024 nodes (12288 CPU cores, 3072 GPUs) exhibit that our GPU-based implementation performs 2.10x faster than running on CPU when graph data size does not fit on GPUs. We also study the performance characteristics of our proposed out-of-core GPU memory management technique, including application's performance and power efficiency of scale-up and scale-out approaches. Koichi Shirahata, Hitoshi Sato, Satoshi Matsuoka |
CLUSTER | 3 |
| 2014 | TSUBAME-KFC: A modern liquid submersion cooling prototype towards exascale becoming the greenest supercomputer in the worldabstractModern supercomputer performance is principally limited by power. TSUBAME-KFC is a state-of-the-art prototype for our next-generation TSUBAME3.0 supercomputer and towards future exascale. In collaboration with Green Revolution Cooling and others, TSUBAME-KFC submerges compute nodes configured with extremely high processor/component density, into non-toxic, low viscosity oil with high 260 Celsius flash point, and cooled using ambient / evaporative cooling tower. This minimizes cooling power while all semiconductor components kept at low temperature to lower leakage current. Numerous off-line in addition to on-line power and temperature sensors are facilitated throughout and constantly monitored to immediately observe the effect of voltage/frequency control. As a result, TSUBAME-KFC achieved world No.1 on the Green500 in Nov. 2013 and Jun. 2014, by over 20% c.f. the nearest competitors. Toshio Endo, Akira Nukada, Satoshi Matsuoka |
ICPADS | 3 |
| 2014 | Cache-aware sparse matrix formats for Kepler GPUabstractScientific simulations often require solving extremely large sparse linear equations, whose dominant kernel is sparse matrix vector multiplication. On modern many-core processors such as GPU or MIC, the operation has been known to pose significant bottleneck and thus would result in extremely poor efficiency, because of limited processor-to-memory bandwidth and low cache hit ratio due to random access to the input vector. Our family of new sparse matrix formats for many-core processors significantly increases the cache hit ratio and thus performance by segmenting the matrix along the columns, dividing the work among the many core up to the internal cache capacity, and aggregating the result later on. Performance studies show that we achieve up to x3.0 speedup in SpMV and x1.68 in multi-node CG, compared to the best vendor libraries and competing new formats that have been recently proposed such as SELL-C-σ. Yusuke Nagasaka, Akira Nukada, Satoshi Matsuoka |
ICPADS | 3 |
| 2014 | Scalable analysis of multicore data reuse and sharingabstractThe performance and energy efficiency of multicore systems are increasingly dominated by the costs of communication. As hardware parallelism grows, developers require more powerful tools to assess the data sharing and reuse properties of their algorithms. The reuse distance is an effective metric to study the temporal locality of programs and model private and shared caches. But the application of this method is challenging. First, generating memory traces is very expensive in storage and very intrusive on execution, possibly distorting the parallel schedule. And second, the algorithm is computationally very expensive, limiting the length, memory size and parallelism of analyzable programs. Miquel Pericàs, Kenjiro Taura, Satoshi Matsuoka |
ICS | 3 |
| 2014 | Petascale General Solver for Semidefinite Programming Problems with Over Two Million ConstraintsabstractThe semi definite programming (SDP) problem is one of the central problems in mathematical optimization. The primal-dual interior-point method (PDIPM) is one of the most powerful algorithms for solving SDP problems, and many research groups have employed it for developing software packages. However, two well-known major bottlenecks, i.e., the generation of the Schur complement matrix (SCM) and its Cholesky factorization, exist in the algorithmic framework of the PDIPM. We have developed a new version of the semi definite programming algorithm parallel version (SDPARA), which is a parallel implementation on multiple CPUs and GPUs for solving extremely large-scale SDP problems with over a million constraints. SDPARA can automatically extract the unique characteristics from an SDP problem and identify the bottleneck. When the generation of the SCM becomes a bottleneck, SDPARA can attain high scalability using a large quantity of CPU cores and some processor affinity and memory interleaving techniques. SDPARA can also perform parallel Cholesky factorization using thousands of GPUs and techniques for overlapping computation and communication if an SDP problem has over two million constraints and Cholesky factorization constitutes a bottleneck. We demonstrate that SDPARA is a high-performance general solver for SDPs in various application fields through numerical experiments conducted on the TSUBAME 2.5 supercomputer, and we solved the largest SDP problem (which has over 2.33 million constraints), thereby creating a new world record. Our implementation also achieved 1.713 PFlops in double precision for large-scale Cholesky factorization using 2,720 CPUs and 4,080 GPUs. Katsuki Fujisawa, Toshio Endo, Yuichiro Yasui, Hitoshi Sato, Naoki Matsuzawa, Satoshi Matsuoka, Hayato Waki |
IPDPS | 6 |
| 2014 | FMI: Fault Tolerant Messaging Interface for Fast and Transparent RecoveryabstractFuture supercomputers built with more components will enable larger, higher-fidelity simulations, but at the cost of higher failure rates. Traditional approaches to mitigating failures, such as checkpoint/restart (C/R) to a parallel file system incur large overheads. On future, extreme-scale systems, it is unlikely that traditional C/R will recover a failed application before the next failure occurs. To address this problem, we present the Fault Tolerant Messaging Interface (FMI), which enables extremely low-latency recovery. FMI accomplishes this using a survivable communication runtime coupled with fast, in-memory C/R, and dynamic node allocation. FMI provides message-passing semantics similar to MPI, but applications written using FMI can run through failures. The FMI runtime software handles fault tolerance, including check pointing application state, restarting failed processes, and allocating additional nodes when needed. Our tests show that FMI runs with similar failure-free performance as MPI, but FMI incurs only a 28% overhead with a very high mean time between failures of 1 minute. Kento Sato, Adam Moody, Kathryn Mohror, Todd Gamblin, Bronis R. de Supinski, Naoya Maruyama, Satoshi Matsuoka |
IPDPS | 7 |
| 2014 | Using rCUDA to Reduce GPU Resource-Assignment Fragmentation Caused by Job SchedulerabstractIn heterogeneous supercomputers such as TSUBAME2.5, GPUs on some nodes in GPU batch queues are left idle even though there are jobs waiting in the queues, this is caused by GPU resource-assignment fragmentation problem. For example, in the case that each node has three GPUs like TSUBAME2.5's, if a node has already been assigned to a job requesting two GPUs per node, that node cannot be assigned to another job requesting more than one GPU per node until the ongoing job finishes, hence, one GPU is left idle on that node. We examine this problem on TSUBAME2.5's GPU batch-queue system and present a scheduling algorithm that assigns rCUDA (a remote CUDA execution technology) to some processes of some jobs. Because rCUDA allows jobs to utilize the idle GPUs, the proposed scheduling algorithm can alleviate the problem. Using a job pattern obtained from a scheduler log of a TSUBAME2.5's GPU queue, our simulation shows that the proposed algorithm can decrease jobs' lifetime (from the time when a job arrives until finishes) by about 5% on average. Moreover, it can reduce the average number of idle GPUs by about 15%. Also, even reducing the number of nodes serving jobs by around 4%, the proposed algorithm can maintain the average jobs' lifetime around the same as the scheduling algorithm currently used in the TSUBAME2.5's GPU queue. Pak Markthub, Akihiro Nomura 0002, Satoshi Matsuoka |
PDCAT | 3 |
| 2014 | Fail-in-Place Network Design: Interaction Between Topology, Routing Algorithm and FailuresabstractThe growing system size of high performance computers results in a steady decrease of the mean time between failures. Exchanging network components often requires whole system downtime which increases the cost of failures. In this work, we study a fail-in-place strategy where broken network elements remain untouched. We show, that a fail-in-place strategy is feasible for todays networks and the degradation is manageable, and provide guidelines for the design. Our network failure simulation tool chain allows system designers to extrapolate the performance degradation based on expected failure rates, and it can be used to evaluate the current state of a system. In a case study of real-world HPC systems, we will analyze the performance degradation throughout the systems lifetime under the assumption that faulty network components are not repaired, which results in a recommendation to change the used routing algorithm to improve the network performance as well as the fail-in-place characteristic. Jens Domke, Torsten Hoefler, Satoshi Matsuoka |
SC | 3 |
| 2013 | CUDA vs OpenACC: Performance Case Studies with Kernel Benchmarks and a Memory-Bound CFD ApplicationabstractOpenACC is a new accelerator programming interface that provides a set of OpenMP-like loop directives for the programming of accelerators in an implicit and portable way. It allows the programmer to express the offloading of data and computations to accelerators, such that the porting process for legacy CPU-based applications can be significantly simplified. This paper focuses on the performance aspects of OpenACC using two micro benchmarks and one real-world computational fluid dynamics application. Both evaluations show that in general OpenACC performance is approximately 50\% lower than CUDA. However, for some applications it can reach up to 98\% with careful manual optimizations. The results also indicate several limitations of the OpenACC specification that hamper full use of the GPU hardware resources, resulting in a significant performance gap when compared to a fully tuned CUDA code. The lack of a programming interface for the shared memory in particular results in as much as three times lower performance. Tetsuya Hoshino, Naoya Maruyama, Satoshi Matsuoka, Ryoji Takaki |
CCGRID | 3 |
| 2013 | A Scalable Implementation of a MapReduce-based Graph Processing Algorithm for Large-Scale Heterogeneous SupercomputersabstractFast processing for extremely large-scale graph is becoming increasingly important in various domains such as health care, social networks, intelligence, system biology, and electric power grids. The GIM-V algorithm based on MapReduce programing model is designed as a general graph processing method for supporting petabyte-scale graph data. On the other hand, recent large-scale data-intensive computing systems tend to employ GPU accelerators to gain good peak performance and high memory bandwidth, however, the validity of acceleration, including optimization techniques, of the GIM-V algorithm using GPUs is an open problem. To address the problem, we implemented a multi-GPU-based GIM-V application with load balance optimization between GPU devices. Our implementation extends the existing MapReduce library for supporting multi-GPU-environments using the MPI library and optimizes load balance between GPU devices by employing task scheduling-based graph partitioning. We conducted our implementation on the TSUBAME2.0 supercomputer using 256 nodes (6144 hyper-threaded CPU cores, 768 GPUs). The results exhibit that our GPU-based implementation performed 87.04 ME/s on 230(1.07 billion) vertices and 234(17.2 billion) edges, and 1.52 times faster than the CPU-based naive implementation with 2^29 vertices and 233edges. We also studied the performance characteristics of our implementation and load balance optimization technique. Koichi Shirahata, Hitoshi Sato, Toyotaro Suzumura, Satoshi Matsuoka |
CCGRID | 4 |
| 2013 | A parallel optimization method for stencil computation on the domain that is bigger than memory capacity of GPUsabstractThe problem size of the stencil computation on GPU cluster is limited by the memory capacity GPUs, which is typically smaller than that of host memories. This paper proposes and evaluates parallel optimization method for stencil computation to achieve scalability, larger problem size than the memory capacity of GPUs and high performance. It uses 2D decomposition to achieve scalability over GPUs. Then it enables bigger sub-domain on each GPU to achieve bigger problem size. It applies temporal blocking method to improve memory access locality of stencil computation and reuses former result to solve redundant problem to get higher performance. Evaluation of stencil simulation on 3D domain shows that our new method for 7-point and 19-point on GPUs achieves good scalability which is 1.45 times and 1.72 times better than other methods on average. Guanghao Jin, Toshio Endo, Satoshi Matsuoka |
CLUSTER | 3 |
| 2013 | Improving the Computing Efficiency of HPC Systems Using a Combination of Proactive and Preventive CheckpointingabstractAs the failure frequency is increasing with the components count in modern and future supercomputers, resilience is becoming critical for extreme scale systems. The association of failure prediction with proactive checkpointing seeks to reduce the effect of failures in the execution time of parallel applications. Unfortunately, proactive checkpointing does not systematically avoid restarting from scratch. To mitigate this issue, failure prediction and proactive checkpointing can be coupled with periodic checkpointing. However, blind use of these techniques does not always improves system efficiency, because everyone of them comes with a mix of overheads and benefits. In order to study and understand the combination of these techniques and their improvement in the system's efficiency, we developed: (i) a prototype combining state of the art failure prediction, fast proactive checkpointing and preventive checkpointing; (ii) a mathematical model that reflects the expected computing efficiency of the combination and computes the optimal checkpointing interval in this context; (iii) a discrete event simulator to evaluate the computing efficiency of the combination for system parameters corresponding to the current and projected large scale HPC systems. We evaluate our proposed technique on a large supercomputer (i.e. TSUBAME2) with production-level HPC applications and we show that failure prediction, proactive and preventive checkpointing can be coupled successfully, imposing only about 2% to 6% of overhead in comparison with preventive checkpointing only. Moreover, our model-based simulations show that the optimal solution improves the computing efficiency up to 30% in comparison with classic periodic checkpointing. We show that the prediction recall has a much higher impact on execution efficiency than the prediction precision. This result suggests that researchers on failure prediction algorithms should focus on improving the recall. We also show that the combination of these techniques can significantly improve (by a factor 2, for a particular configuration) the mean time between failures (MTBF) perceived by the application. Mohamed-Slim Bouguerra, Ana Gainaru, Leonardo Arturo Bautista-Gomez, Franck Cappello, Satoshi Matsuoka, Naoya Maruyama |
IPDPS | 5 |
| 2012 | Design and Implementation of Portable and Efficient Non-blocking Collective CommunicationabstractNon-blocking communications are widely used in parallel applications for hiding communication overheads through overlapped computation and communication. While most of the existing implementations provide a non-blocking version of point-to-point communications, there is no portable and efficient implementation of non-blocking collectives, partly because application execution contexts need to be interrupted by dependent communications. This paper presents a portable and efficient user-level implementation technique of non-blocking communications. It allows users to design non-blocking collectives by declaring their operations and dependencies using provided APIs without being concerned with complicated management of their progression. While user-level implementations can be less efficient than kernel-level ones due to the cost of OS context switches, we solve this problem by employing the Marcel user level light-weight thread library when invoking communication operations. More specifically, each communication operation is mapped to one Marcel thread and scheduled to be executed when each operation's dependencies are satisfied by certain events. All executable operations and main user thread are executed simultaneously without any explicit invocations. Performance evaluations with micro benchmarks demonstrate the effectiveness of our proposed technique. Compared to existing OS-thread based method, it reduces CPU load to less than 10% while achieving similar level of communication latencies. We also discuss and compare the descriptive power of internal expressions for non-blocking communications. Akihiro Nomura 0002, Yutaka Ishikawa, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 4 |
| 2012 | Hierarchical Clustering Strategies for Fault Tolerance in Large Scale HPC SystemsabstractFuture high performance computing systems will need to use novel techniques to allow scientific applications to progress despite frequent failures. Checkpoint-Restart is currently the most popular way to mitigate the impact of failures during long-running executions. Different techniques try to reduce the cost of Checkpoint-Restart, some of them such as local check pointing and erasure codes aim to reduce the time to checkpoint while others such as uncoordinated checkpoint and message-logging aim to decrease the cost of recovery. In this paper, we study how to combine all these techniques together in order to optimize both: check pointing and recovery. We present several clustering and topology challenges that lead us to an optimization problem in a four-dimensional space: reliability level, recovery cost, encoding time and message logging overhead. We propose a novel clustering method inspired from brain topology studies in neuroscience and evaluate it with a Tsunami simulation application in TSUBAME2. Our evaluation with 1024 processes shows that our novel clustering method can guarantee good performance for all of the four mentioned dimensions of our optimization problem. Leonardo Arturo Bautista-Gomez, Thomas Ropars, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
CLUSTER | 5 |
| 2012 | Scalable Reed-Solomon-Based Reliable Local Storage for HPC Applications on IaaS Clouds
Leonardo Arturo Bautista-Gomez, Bogdan Nicolae, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
Euro-Par | 5 |
| 2012 | Topic 16: GPU and Accelerators Computing
Alex Ramírez, Dimitrios S. Nikolopoulos, David R. Kaeli, Satoshi Matsuoka |
Euro-Par | 4 |
| 2012 | Using Bittorrent and SVC for efficient video sharing and streamingabstractMassive and large scale content distribution over Internet is attracting a lot of research efforts as many challenges remain to be solved. Recent studies show that Internet video including video-to-TV and video calling is dominating the Internet traffic. As Internet becomes widely accessible to wired, mobile and wireless users, it is important to design a system that can ensure video streaming across variable network conditions while simultaneously handling devices and end-user heterogeneities. Most of the proposed solutions, such as CDN and peer-to-peer (P2P), solve the scalability problem but fail to handle receiver's heterogeneity. In this paper, we combine P2P network and SVC (Scalable Video Coding) to provide an efficient video sharing and streaming system. Our solution consists of an SVC layered extension of the widely used Bittorrent protocol to support real-time content delivery with different video qualities given the receivers capabilities. Thus, we propose different optimization techniques to organize peers in an overlay. The results, obtained by means of simulation, show that our system outperforms solutions that relay on single layer streams such as AVC (Advanced Video Coding) and this in terms of receivers perceived QoS. Abdelhalim Amer, Ahmed Touflk, Walid-Khaled Hidouci, Satoshi Matsuoka |
ISCC | 4 |
| 2012 | High-performance general solver for extremely large-scale semidefinite programming problemsabstractSemidefinite programming (SDP) is one of the most important problems among optimization problems at present. It is relevant to a wide range of fields such as combinatorial optimization, structural optimization, control theory, economics, quantum chemistry, sensor network location and data mining. The capability to solve extremely large-scale SDP problems will have a significant effect on the current and future applications of SDP. In 1995, Fujisawa et al. started the SDPA(Semidefinite programming algorithm) Project aimed at solving large-scale SDP problems with high numerical stability and accuracy. SDPA is one of the main codes to solve general SDPs. SDPARA is a parallel version of SDPA on multiple processors with distributed memory, and it replaces two major bottleneck parts (the generation of the Schur complement matrix and its Cholesky factorization) of SDPA by their parallel implementation. In particular, it has been successfully applied to combinatorial optimization and truss topology optimization. The new version of SDPARA (7.5.0-G) on a large-scale supercomputer called TSUBAME 2.0 at the Tokyo Institute of Technology has successfully been used to solve the largest SDP problem (which has over 1.48 million constraints), and created a new world record. Our implementation has also achieved 533 TFlops in double precision for large-scale Cholesky factorization using 2,720 CPUs and 4,080 GPUs. Katsuki Fujisawa, Hitoshi Sato, Satoshi Matsuoka, Toshio Endo, Makoto Yamashita, Maho Nakata |
SC | 3 |
| 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputerabstractFor scalable 3-D FFT computation using multiple GPUs, efficient all-to-all communication between GPUs is the most important factor in good performance. Implementations with point-to-point MPI library functions and CUDA memory copy APIs typically exhibit very large overheads especially for small message sizes in all-to-all communications between many nodes. We propose several schemes to minimize the overheads, including employment of lower-level API of InfiniBand to effectively overlap intra- and inter-node communication, as well as auto-tuning strategies to control scheduling and determine rail assignments. As a result we achieve very good strong scalability as well as good performance, up to 4.8TFLOPS using 256 nodes of TSUBAME 2.0 Supercomputer (768 GPUs) in double precision. Akira Nukada, Kento Sato, Satoshi Matsuoka |
SC | 3 |
| 2012 | Design and modeling of a non-blocking checkpointing systemabstractAs the capability and component count of systems increase, the MTBF decreases. Typically, applications tolerate failures with checkpoint/restart to a parallel file system (PFS). While simple, this approach can suffer from contention for PFS resources. Multi-level checkpointing is a promising solution. However, while multi-level checkpointing is successful on today's machines, it is not expected to be sufficient for exascale class machines, which are predicted to have orders of magnitude larger memory sizes and failure rates. Our solution combines the benefits of non-blocking and multi-level checkpointing. In this paper, we present the design of our system and model its performance. Our experiments show that our system can improve efficiency by 1.1 to 2.0x on future machines. Additionally, applications using our checkpointing system can achieve high efficiency even when using a PFS with lower bandwidth. Kento Sato, Naoya Maruyama, Kathryn Mohror, Adam Moody, Todd Gamblin, Bronis R. de Supinski, Satoshi Matsuoka |
SC | 7 |
| 2012 | A coding theoretic study of MLL proof netsabstractIn this paper we propose a novel approach for analysing proof nets of Multiplicative Linear Logic (MLL) using coding theory. We define families of proof structures called PS-families and introduce a metric space for each family. In each family: (1) an MLL proof net is a true code element; and (2) a proof structure that is not an MLL proof net is a false (or corrupted) code element. The definition of our metrics elegantly reflects the duality of the multiplicative connectives. We show that in our framework one-error-detection is always possible but one-error-correction is always impossible. We also demonstrate the importance of our main result by presenting two proof-net enumeration algorithms for a given PS-family: the first searches proof nets naively and exhaustively without help from our main result, while the second uses our main result to carry out an intelligent search. In some cases, the first algorithm visits proof structures exponentially, while the second does so only polynomially. Satoshi Matsuoka |
Math. Struct. Comput. Sci. | 1 |
| 2011 | Dealing with Grid-Computing Authorization Using Identity-Based Certificateless Proxy SignatureabstractIn this paper, we propose a new Identity-Based Certificateless Proxy Signature scheme, for the grid environment, in order to enable attribute-based authorization, fine-grained delegation and enhanced delegation chain establishment and validation, all without relying on any kind of PKI Certificates or proxy certificates. We show that our scheme is correct and secure. We also give an evaluation of the computational and communication overhead of the proposed scheme. Simulations shows satisfying results. Mohamed Amin Jabri, Satoshi Matsuoka |
CCGRID | 2 |
| 2011 | Multi-ring Structured Overlay Network for the Inter-cloud Computing Environment
Sumeth Lerthirunwong, Hitoshi Sato, Satoshi Matsuoka |
CLOSER | 3 |
| 2011 | Panel StatementabstractSummary form only given, as follows. The 25th year of IPDPS gives us the opportunity to look back and (to attempt) to assess what has gone wrong, what has gone well, and what came as a surprise, in the field of parallel and distributed processing. The panel members will give a few examples of striking events that took place in their area (covering Algorithms/ Applications/ Architectures/ Software). They will also give a short statement on how they would summarize the evolution of the field as a whole over the last 25 years. Yves Robert, William J. Dally, Jack J. Dongarra, Satoshi Matsuoka, Robert Schreiber, Horst D. Simon, Uzi Vishkin |
IPDPS | 4 |
| 2011 | Making TSUBAME2.0, the world's greenest production supercomputer, even greener: challenges to the architects
Satoshi Matsuoka |
ISLPED | 1 |
| 2011 | FTI: high performance fault tolerance interface for hybrid systemsabstractLarge scientific applications deployed on current petascale systems expend a significant amount of their execution time dumping checkpoint files to remote storage. New fault tolerant techniques will be critical to efficiently exploit post-petascale systems. In this work, we propose a low-overhead high-frequency multi-level checkpoint technique in which we integrate a highly-reliable topology-aware Reed-Solomon encoding in a three-level checkpoint scheme. We efficiently hide the encoding time using one Fault-Tolerance dedicated thread per node. We implement our technique in the Fault Tolerance Interface FTI. We evaluate the correctness of our performance model and conduct a study of the reliability of our library. To demonstrate the performance of FTI, we present a case study of the Mw9.0 Tohoku Japan earthquake simulation with SPECFEM3D on TSUBAME2.0. We demonstrate a checkpoint overhead as low as 8% on sustained 0.1 petaflops runs (1152 GPUs) while checkpointing at high frequency. Leonardo Arturo Bautista-Gomez, Seiji Tsuboi, Dimitri Komatitsch, Franck Cappello, Naoya Maruyama, Satoshi Matsuoka |
SC | 6 |
| 2011 | Petaflop biofluidics simulations on a two million-core systemabstractWe present a computational framework for multi-scale simulations of real-life biofluidic problems. The framework allows to simulate suspensions composed by hundreds of millions of bodies interacting with each other and with a surrounding fluid in complex geometries. We apply the methodology to the simulation of blood flow through the human coronary arteries with a spatial resolution comparable with the size of red blood cells, and physiological levels of hematocrit (the red blood cell volume fraction). The simulation exhibits excellent scalability on a cluster of 4000 M2050 Nvidia GPUs and achieves close to 1 Petaflop aggregate performance, which demonstrates the capability to predicting the evolution of biofluidic phenomena of clinical significance. The combination of novel mathematical models, computational algorithms, hardware technology, code tuning and optimization required to achieve these results are presented. Massimo Bernaschi, Mauro Bisson, Toshio Endo, Satoshi Matsuoka, Massimiliano Fatica, Simone Melchionna |
SC | 4 |
| 2011 | Physis: an implicitly parallel programming model for stencil computations on large-scale GPU-accelerated supercomputersabstractThis paper proposes a compiler-based programming framework that automatically translates user-written structured grid code into scalable parallel implementation code for GPU-equipped clusters. To enable such automatic translations, we design a small set of declarative constructs that allow the user to express stencil computations in a portable and implicitly parallel manner. Our framework translates the user-written code into actual implementation code in CUDA for GPU acceleration and MPI for node-level parallelization with automatic optimizations such as computation and communication overlapping. We demonstrate the feasibility of such automatic translations by implementing several structured grid applications in our framework. Experimental results on the TSUBAME2.0 GPU-based supercomputer show that the performance is comparable as hand-written code and good strong and weak scalability up to 256 GPUs. Naoya Maruyama, Tatsuo Nomura, Kento Sato, Satoshi Matsuoka |
SC | 4 |
| 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputerabstractThe mechanical properties of metal materials largely depend on their intrinsic internal microstructures. To develop engineering materials with the expected properties, predicting patterns in solidified metals would be indispensable. The phase-field simulation is the most powerful method known to simulate the micro-scale dendritic growth during solidification in a binary alloy. To evaluate the realistic description of solidification, however, phase-field simulation requires computing a large number of complex nonlinear terms over a fine-grained grid. Due to such heavy computational demand, previous work on simulating three-dimensional solidification with phase-field methods was successful only in describing simple shapes. Our new simulation techniques achieved scales unprecedentedly large, sufficient for handling complex dendritic structures required in material science. Our simulations on the GPU-rich TSUBAME 2.0 supercomputer at the Tokyo Institute of Technology have demonstrated good weak scaling and achieved 1.017 PFlops in single precision for our largest configuration, using 4,000 GPUs along with 16,000 CPU cores. Takashi Shimokawabe, Takayuki Aoki, Tomohiro Takaki, Toshio Endo, Akinori Yamanaka, Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
SC | 8 |
| 2010 | Dynamic Load-Balanced Multicast for Data-Intensive Applications on CloudsabstractData-intensive parallel applications on clouds need to deploy large data sets from the cloud's storage facility to all compute nodes as fast as possible. Many multicast algorithms have been proposed for clusters and grid environments. The most common approach is to construct one or more spanning trees based on the network topology and network monitoring data in order to maximize available bandwidth and avoid bottleneck links. However, delivering optimal performance becomes difficult once the available bandwidth changes dynamically. In this paper, we focus on Amazon EC2/S3 (the most commonly used cloud platform today) and propose two high performance multicast algorithms. These algorithms make it possible to efficiently transfer large amounts of data stored in Amazon S3 to multiple Amazon EC2 nodes. The three salient features of our algorithms are (1) to construct an overlay network on clouds without network topology information, (2) to optimize the total throughput dynamically, and (3) to increase the download throughput by letting nodes cooperate with each other. The two algorithms differ in the way nodes cooperate: the first `non-steal' algorithm lets each node download an equal share of all data, while the second `steal' algorithm uses work stealing to counter the effect of heterogeneous download bandwidth. As a result, all nodes can download files from S3 quickly, even when the network performance changes while the algorithm is running. We evaluate our algorithms on EC2/S3, and show that they are scalable and consistently achieve high throughput. Both algorithms perform much better than having each node downloading all data directly from S3. Tatsuhiro Chiba, Mathijs den Burger, Thilo Kielmann, Satoshi Matsuoka |
CCGRID | 4 |
| 2010 | Distributed Diskless Checkpoint for Large Scale SystemsabstractIn high performance computing (HPC), the applications are periodically check pointed to stable storage to increase the success rate of long executions. Nowadays, the overhead imposed by disk-based checkpoint is about 20% of execution time and in the next years it will be more than 50% if the checkpoint frequency increases as the fault frequency increases. Diskless checkpoint has been introduced as a solution to avoid the IO bottleneck of disk-based checkpoint. However, the encoding time, the dedicated resources (the spares) and the memory overhead imposed by diskless checkpoint are significant obstacles against its adoption. In this work, we address these three limitations: 1) we propose a fault tolerant model able to tolerate up to 50% of process failures with a low check pointing overhead 2) our fault tolerance model works without spare node, while still guarantying high reliability, 3) we use solid state drives to significantly increase the checkpoint performance and avoid the memory overhead of classic diskless checkpoint. Leonardo Arturo Bautista-Gomez, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
CCGRID | 4 |
| 2010 | Hybrid Map Task Scheduling for GPU-Based Heterogeneous ClustersabstractMapReduce is a programming model that enables efficient massive data processing in large-scale computing environments such as supercomputers and clouds. Such large-scale computers employ GPUs to enjoy its good peak performance and high memory bandwidth. Since the performance of each job is depending on running application characteristics and underlying computing environments, scheduling MapReduce tasks onto CPU cores and GPU devices for efficient execution is difficult. To address this problem, we have proposed a hybrid scheduling technique for GPU-based computer clusters, which minimizes the execution time of a submitted job using dynamic profiles of Map tasks running on CPU cores and GPU devices. We have implemented a prototype of our proposed scheduling technique by extending MapReduce framework, Hadoop. We have conducted some experiments for this prototype by using a K-means application as a benchmark on a supercomputer. The results show that the proposed technique achieves 1.93 times faster than the Hadoop original scheduling algorithm at 64 nodes (1024 CPU cores and 128 GPU devices). The results also indicate that the performance of map tasks, including both CPU and GPU tasks, is significantly affected by the overhead of map task invocation in the Hadoop framework. Koichi Shirahata, Hitoshi Sato, Satoshi Matsuoka |
CloudCom | 3 |
| 2010 | Low-overhead diskless checkpoint for hybrid computing systemsabstractAs the size of new supercomputers scales to tens of thousands of sockets, the mean time between failures (MTBF) is decreasing to just several hours and long executions need some kind of fault tolerance method to survive failures. Checkpoint\Restart is a popular technique used for this purpose; but writing the state of a big scientific application to remote storage will become prohibitively expensive in the near future. Diskless checkpoint was proposed as a solution to avoid the I/O bottleneck of disk-based checkpoint. However, the complex time-consuming encoding techniques hinder its scalability. At the same time, heterogeneous computing is becoming more and more popular in high performance computing (HPC), with new clusters combining CPUs and graphic processing units (GPUs). However, hybrid applications cannot always use all the resources available on the nodes, leaving some idle resources suc h us GPUs or CPU cores. In this work, we propose a hybrid diskless checkpoint (HDC) technique for GPU-accelerated clusters, that can checkpoint CPU/GPU applications, does not require spare nodes and can tolerate up to 50% of process failures with a low, sometimes negligible, checkpoint overhead. Leonardo Arturo Bautista-Gomez, Akira Nukada, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
HiPC | 5 |
| 2010 | Authorization within grid-computing using certificateless identity-based proxy signatureabstractEnsuring security and privacy within the Grid computing environment is a fundamental and key requirement for users in order to adopt and use secure and trusted Grids. In light of this fact, the majority of grids, nowadays, provide security services by relying on the de facto standard GSI framework which makes use of Public key End Entity Certificates (EEC) and Proxy Certificates as its foundation. But, due to the huge burden in managing, distributing and revoking compromised EEC, Public Key Certificates are becoming an impediment to the wide adoption and use of Grids at scale. Thus, many research efforts stressed how compelling is the adoption of Identity-Based cryptography and Certificateless Public Key Cryptography within the Grid-Computing environment, which could alleviate the burden of using PKI certificates while, still, providing the intended secure, flexible and easy use of the Grid to both non expert users and Grid's administrators. Mohamed Amin Jabri, Satoshi Matsuoka |
HPDC | 2 |
| 2010 | Linpack evaluation on a supercomputer with heterogeneous acceleratorsabstractWe report Linpack benchmark results on the TSUBAME supercomputer, a large scale heterogeneous system equipped with NVIDIA Tesla GPUs and ClearSpeed SIMD accelerators. With all of 10,480 Opteron cores, 640 Xeon cores, 648 ClearSpeed accelerators and 624 NVIDIA Tesla GPUs, we have achieved 87.01TFlops, which is the third record as a heterogeneous system in the world. This paper describes careful tuning and load balancing method required to achieve this performance. On the other hand, since the peak speed is 163 TFlops, the efficiency is 53%, which is lower than other systems. This paper also analyses this gap from the aspect of system architecture. Toshio Endo, Akira Nukada, Satoshi Matsuoka, Naoya Maruyama |
IPDPS | 3 |
| 2010 | A high-performance fault-tolerant software framework for memory on commodity GPUsabstractAs GPUs are increasingly used to accelerate HPC applications by allowing more flexibility and programmability, their fault tolerance is becoming much more important than before when they were used only for graphics. The current generation of GPUs, however, does not have standard error detection and correction capabilities, such as SEC-DED ECC for DRAM, which is almost always exercised in HPC servers. We present a high-performance software framework to enhance commodity off-the-shelf GPUs with DRAM fault tolerance. It combines data coding for detecting bit-flip errors and checkpointing for recovering computations when such errors are detected. We analyze performance of data coding in GPUs and present optimizations geared toward memory-intensive GPU applications. We present performance studies of the prototype implementation of the framework and show that the proposed framework can be realized with negligible overheads in compute intensive applications such as N-body problem and matrix multiplication, and as low as 35% in a highly-efficient memory intensive 3-D FFT kernel. Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
IPDPS | 3 |
| 2010 | An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production CodeabstractRegional weather forecasting demands fast simulation over fine-grained grids, resulting in extremely memory- bottlenecked computation, a difficult problem on conventional supercomputers. Early work on accelerating mainstream weather code WRF using GPUs with their high memory performance, however, resulted in only minor speedup due to partial GPU porting of the huge code. Our full CUDA porting of the high- resolution weather prediction model ASUCA is the first such one we know to date; ASUCA is a next-generation, production weather code developed by the Japan Meteorological Agency, similar to WRF in the underlying physics (non-hydrostatic model). Benchmark on the 528 (NVIDIA GT200 Tesla) GPU TSUBAME Supercomputer at the Tokyo Institute of Technology demonstrated over 80-fold speedup and good weak scaling achieving 15.0 TFlops in single precision for 6956 x 6052 x 48 mesh. Further benchmarks on TSUBAME 2.0, which will embody over 4000 NVIDIA Fermi GPUs and deployed in October 2010, will be presented. Takashi Shimokawabe, Takayuki Aoki, Chiashi Muroi, Junichi Ishida, Kohei Kawano, Toshio Endo, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
SC | 9 |
| 2010 | Global-scale distributed I/O with ParaMEDICabstractAbstract Achieving high performance for distributed I/O on a wide‐area network continues to be an elusive holy grail. Despite enhancements in network hardware as well as software stacks, achieving high‐performance remains a challenge. In this paper, our worldwide team took a completely new and non‐traditional approach to distributed I/O, calledParaMEDIC: Parallel Metadata Environment for Distributed I/O and Computing, by utilizing application‐specifictransformationof data to orders of magnitude smaller metadata before performing the actual I/O. Specifically, this paper details our experiences in deploying a large‐scale system to facilitate the discovery of missing genes and constructing a genome similarity tree by encapsulating the mpiBLAST sequence‐search algorithm into ParaMEDIC. The overall project involved nine computational sites spread across the U.S. and generated more than a petabyte of data that was ‘teleported’ to a large‐scale facility in Tokyo for storage. Copyright © 2010 John Wiley & Sons, Ltd. Pavan Balaji, Wu-chun Feng, Heshan Lin, Jeremy S. Archuleta, Satoshi Matsuoka, Andrew S. Warren, João Carlos Setubal, Ewing L. Lusk, Rajeev Thakur, Ian T. Foster, Daniel S. Katz, Shantenu Jha, K. Shinpaugh, Susan Coghlan, Daniel A. Reed |
Concurr. Comput. Pract. Exp. | 5 |
| 2009 | Aspects of GPU for general purpose high performance computingabstractWe discuss hardware and software aspects of GPGPU, specifically focusing on NVIDIA cards and CUDA, from the viewpoints of parallel computing. The major weak points of GPU against newest supercomputers are identified to be and summarized as only four points: large SIMD vector length, small memory, absence of fast L2 cache, and high register spill penalty. As software concerns, we derive optimal scheduling algorithm for latency hiding of host-device data transfer, and discuss SPMD parallelism on GPUs. Reiji Suda, Takayuki Aoki, Shoichi Hirasawa, Akira Nukada, Hiroki Honda, Satoshi Matsuoka |
ASP-DAC | 6 |
| 2009 | Adaptive Resource Indexing Technique for Unstructured Peer-to-Peer NetworksabstractSearching for particular resources in a large-scale decentralized unstructured network can be very difficult since there is no centralized management to provide the specific location of resources. Moreover, the dynamic behavior of networks and the diversity of user behavior cause the search more complex and may not guarantee success. To address the problems, we propose a new adaptive resource indexing technique that aims to increase both efficiency and quality of the search by reducing both messages and time required for each query. Our approach consists of two complementary techniques. One is an index selection technique that selectively keeps the indices at each peer to increase the chance of successful queries with minimum space requirement. Another is an index distribution technique that automatically adjusts index distribution rate based on the search performance to optimize both the search performance and overhead. We simulate the technique in various network conditions and the results show that our technique is effective in decreasing hop counts and messages needed for resolving queries with only small overhead. It decreases the average hop count by up to 44% with 75%-less messages when used with flooding based queries even facing high churn. Furthermore, the query success rate with a limited timeout condition also increases, approaching nearly to 100%. Sumeth Lerthirunwong, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 3 |
| 2009 | File Clustering Based Replication Algorithm in a Grid EnvironmentabstractReplication in grid file systems can significantly improve I/O performance of data-intensive applications. However, most of existing replication techniques apply to individual files, which may introduce inefficient replication overheads for a large number of files. We propose a file clustering based replication algorithm for grid file systems. Our algorithm groups files according to a relationship of simultaneous accesses between files and stores replicas of the clustered files into storage nodes, to satisfy expected most of future read access times to the clustered files and replication times for individual files being minimized under the given storage capacity limitation. Our experiments on a given grid environment, 20 nodes of 5 sites, suggest that the proposed algorithm achieves accurate file clustering and efficient replica management; our clustering policy with the file cluster size limit of 5120 MB and the storage capacity limit for replicas of 10240 MB exhibits 1.58 times efficiency than the policy that never groups related files. The results also indicate that the overheads required for introducing our algorithm significantly affect I/O performance of running applications. Hitoshi Sato, Satoshi Matsuoka, Toshio Endo |
CCGRID | 2 |
| 2009 | A Model-Based Algorithm for Optimizing I/O Intensive Applications in Clouds Using VM-Based MigrationabstractFederated storage resources in geographically distributed environments are becoming viable platforms for data-intensive cloud and grid applications. To improve I/O performance in such environments, we propose a novel model-based I/O performance optimization algorithm for data-intensive applications running on a virtual cluster, which determines virtual machine (VM) migration strategies,i.e., when and where a VM should be migrated, while minimizing the expected value of file access time. We solve this problem as a shortest path problem of a weighted direct acyclic graph (DAG), where the weighted vertex represents a location of a VM and expected file access time from the location, and the weighted edge represents a migration of a VM and time. We construct the DAG from our Markov model which represents the dependency of files. Our simulation-based studies suggest that our proposed algorithm can achieve higher performance than simple techniques, such as ones that never migrate VMs: 38% or always migrate VMs onto the locations that hold target files: 47%. Kento Sato, Hitoshi Sato, Satoshi Matsuoka |
CCGRID | 3 |
| 2009 | Power-aware dynamic task scheduling for heterogeneous accelerated clustersabstractRecent accelerators such as GPUs achieve better cost-performance and watt-performance ratio, while the range of their application is more limited than general CPUs. Thus heterogeneous clusters and supercomputers equipped both with accelerators and general CPUs are becoming popular, such as LANL's Roadrunner and our own TSUBAME supercomputer. Under the assumption that many applications will run both on CPUs and accelerators but with varying speed and power consumption characteristics, we propose a task scheduling scheme that optimize overall energy consumption of the system. We model task scheduling in terms of the scheduling makespan and energy to be consumed for each scheduling decision. We define acceleration factor to normalize the effect of acceleration per each task. The proposed scheme attempts to improve energy efficiency by effectively adjusting the schedule based on the acceleration factor. Although in the paper we adopted the popular EDP (Energy-Delay Product) as the optimization metric, our scheme is agnostic on the optimization function. Simulation studies on various sets of tasks with mixed acceleration factors, the overall makespan closely matched the theoretical optimal, while the energy consumption was reduced up to 13.8%. Tomoaki Hamano, Toshio Endo, Satoshi Matsuoka |
IPDPS | 3 |
| 2009 | Auto-tuning 3-D FFT library for CUDA GPUsabstractExisting implementations of FFTs on GPUs are optimized for specific transform sizes like powers of two, and exhibit unstable and peaky performance i.e., do not perform as well in other sizes that appear in practice. Our new auto-tuning 3-D FFT on CUDA generates high performance CUDA kernels for FFTs of varying transform sizes, alleviating this problem. Although auto-tuning has been implemented on GPUs for dense kernels such as DGEMM and stencils, this is the first instance that has been applied comprehensively to bandwidth intensive and complex kernels such as 3-D FFTs. Bandwidth intensive optimizations such as selecting the number of threads and inserting padding to avoid bank conflicts on shared memory are systematically applied. Our resulting autotuner is fast and results in performance that essentially beats all 3-D FFT implementations on a single processor to date, and moreover exhibits stable performance irrespective of problem sizes or the underlying GPU hardware. Akira Nukada, Satoshi Matsuoka |
SC | 2 |
| 2009 | Interoperation of world-wide production e-Science infrastructuresabstractAbstract Many production Grid and e‐Science infrastructures have begun to offer services to end‐users during the past several years with an increasing number of scientific applications that require access to a wide variety of resources and services in multiple Grids. Therefore, the Grid Interoperation Now—Community Group of the Open Grid Forum—organizes and manages interoperation efforts among those production Grid infrastructures to reach the goal of a world‐wide Grid vision on a technical level in the near future. This contribution highlights fundamental approaches of the group and discusses open standards in the context of production e‐Science infrastructures. Copyright © 2009 John Wiley & Sons, Ltd. Morris Riedel, Erwin Laure, Thomas Soddemann, Laurence Field, John-Paul Navarro, James Casey, Maarten Litmaath, Jean-Philippe Baud, Birger Koblitz, Charles E. Catlett, Dane Skow, Cindy Zheng, Philip M. Papadopoulos, Mason J. Katz, Neha Sharma 0001, Oxana Smirnova, Balázs Kónya, Peter W. Arzberger, Frank Würthwein, Abhishek Singh Rana, Terrence Martin, M. Wan, Von Welch, Tony Rimovsky, Steven J. Newhouse, Andrea Vanni, Yoshio Tanaka, Yusuke Tanimura, Tsutomu Ikegami, David Abramson 0001, Colin Enticott, Graham Jenkins, Ruth Pordes, Steven Timm, Gidon Moont, Mona Aggarwal, Dave Colling, Olivier van der Aa, Alex Sim, Vijaya Natarajan, Arie Shoshani, Junmin Gu, Gerson Galang, Riccardo Zappi, Luca Magnoni, Vincenzo Ciaschini, Michele Pace, Valerio Venturi, Moreno Marzolla, Paolo Andreetto, Robert Cowles, Shaowen Wang 0001, Yuji Saeki, Hitoshi Sato, Satoshi Matsuoka, Putchong Uthayopas, Somsak Sriprayoonsakul, Oscar Koeroo, Matthew Viljoen, Laura Pearlman, Stephen Pickles, David Wallom, Glenn Moloney, Jerome Lauret, Jim Marsteller, Paul Sheldon, Surya Pathak, Shaun De Witt, Jirí Mencák, Jens Jensen, Matt Hodges, Derek Ross, Sugree Phatanapherom, Gilbert Netzer, Anders Rhod Gregersen, Mike Jones 0002, Péter Kacsuk, Achim Streit, Daniel Mallmann, Felix Wolf 0001, Thomas Lippert, Thierry Delaitre, Eduardo Huedo, Neil Geddes |
Concurr. Comput. Pract. Exp. | 56 |
| 2008 | Time-Stamping Authority GridabstractDistributed time-stamping enables tolerance to distributed denial of service attacks. However, they involve high cost due to the requirement that all time- stamping units (TSUs), whose number could be numerous, must be audited by trusted third parties. We propose a distributed multi-generation time-stamping scheme with predictable expectation time of notary and minimizing the cost of administration. The scheme is called the "Time-Stamping Authority Grid" with "K = L + M among N in G generations scheme ", where K is the number of issued requests, L is the number of requests to reliable TSUs, M is the number of requests to randomly chosen TSUs, N is the total number of TSUs, and G is the number of the generation that the requests are propagated. Our scheme solves the problems involving both centralized and previous distributed time-stamping schemes by constructing a peer-to-peer network of TSUs, thereby attaining the scalability and dramatically reduced cost. Takeshi Nishikawa, Satoshi Matsuoka |
CCGRID | 2 |
| 2008 | Environmental-aware optimization of MPI checkpointing intervalsabstractFault-tolerance for HPC systems with long-running applications of massive and growing scale is now essential. Although checkpointing with rollback recovery is a popular technique, automated checkpointing is becoming troublesome in a real system, due to the extremely large size of collective application memory. Therefore, automated optimization of the checkpoint interval is essential, but the optimal point depends on hardware failure rates and I/O bandwidth. Our new model and an algorithm, which is an extension of Vaidyapsilas model, solve the problem by taking such parameters into account. Prototype implementation on our fault-tolerant MPI framework ABARIS showed approximately 5.5% improvement over statically user-determined cases. Hideyuki Jitsumoto, Toshio Endo, Satoshi Matsuoka |
CLUSTER | 3 |
| 2008 | Massive supercomputing coping with heterogeneity of modern acceleratorsabstractHeterogeneous supercomputers with combined general-purpose and accelerated CPUs promise to be the future major architecture due to their wide-ranging generality and superior performance / power ratio. However, developing applications that achieve effective scalability is still very difficult, and in fact unproven on large-scale machines in such combined setting. We show that an effective method for such heterogeneous systems so that the porting from applications written with homogeneous assumptions could be achieved. For this goal, we divide porting of applications into several steps, analyze performance of the kernel computation, create processes that virtualize the underlying processors, tune parameters with preferences to accelerators, and balance the load between heterogeneous nodes. We apply our method to the parallel Linpack benchmark on the TSUBAME heterogeneous supercomputer. We efficiently utilize both 10,000 general purpose CPU cores and 648 SIMD accelerators in a combined fashion—the resulting 56.43 TFlops utilized the entire machine, and not only ranked significantly on the Top500 supercomputer list, but also it is the highest Linpack performance on heterogeneous systems in the world. Toshio Endo, Satoshi Matsuoka |
IPDPS | 2 |
| 2008 | Performance evaluation of parallel applications on next generation memory architecture with power-aware paging methodabstractWith increasing demand for low power high performance computing, reducing power of not only CPUs but also memory is becoming important. In typical general-purpose HPC environments, DRAM is installed in an over-provisioned fashion to avoid swapping, although in most cases not all such memory is used, leading to unnecessary and excessive power consumption, even in a standby state. We propose a next generation low power memory system that reduces required DRAM capacity while minimizing application performance degradation. In this system, both DRAM and MRAM, fast non-volatile memory, are used as main memory, while flash memory is used as a swap device. Our profile-based paging algorithm optimizes memory accesses by using faster memory as much as possible, reducing accesses to slower memory. Simulated results of our architecture show that the overall energy consumption of the memory system can be reduced to 25% by in the best case by reducing DRAM capacity, with only 17% performance loss in application benchmarks. Y. Hosogaya, Toshio Endo, Satoshi Matsuoka |
IPDPS | 3 |
| 2008 | Model-based fault localization in large-scale computing systemsabstractWe propose a new fault localization technique for software bugs in large-scale computing systems. Our technique always collects per-process function call traces of a target system, and derives a concise execution model that reflects its normal function calling behaviors using the traces. To find the cause of a failure, we compare the derived model with the traces collected when the system failed, and compute a suspect score that quantifies how likely a particular part of call traces explains the failure. The execution model consists of a call probability of each function in the system that we estimate using the normal traces. Functions with low probabilities in the model give high anomaly scores when called upon a failure. Frequently-called functions in the model also give high scores when not called. Finally, we report the function call sequences ranked with the suspect scores to the human analyst, narrowing further manual localization down to a small part of the overall system. We have applied our proposed method to fault localization of a known non-deterministic bug in a distributed parallel job manager. Experimental results on a three-site, 78-node distributed environment demonstrate that our method quickly locates an anomalous event that is highly correlated with the bug, indicating the effectiveness of our approach. Naoya Maruyama, Satoshi Matsuoka |
IPDPS | 2 |
| 2008 | An efficient, model-based CPU-GPU heterogeneous FFT libraryabstractGeneral-Purpose computing on Graphics Processing Units (GPGPU) is becoming popular in HPC because of its high peak performance. However, in spite of the potential performance improvements as well as recent promising results in scientific computing applications, its real performance is not necessarily higher than that of the current high-performance CPUs, especially with recent trends towards increasing the number of cores on a single die. This is because the GPU performance can be severely limited by such restrictions as memory size and bandwidth and programming using graphics-specific APIs. To overcome this problem, we propose a model-based, adaptive library for 2D FFT that automatically achieves optimal performance using available heterogeneous CPU-GPU computing resources. To find optimal load distribution ratios between CPUs and GPUs, we construct a performance model that captures the respective contributions of CPU vs. GPU, and predicts the total execution time of 2D-FFT for arbitrary problem sizes and load distribution. The performance model divides the FFT computation into several small sub steps, and predicts the execution time of each step using profiling results. Preliminary evaluation with our prototype shows that the performance model can predict the execution time of problem sizes that are 16 times as large as the profile runs with less than 20% error, and that the predicted optimal load distribution ratios have less than 1% error. We show that the resulting performance improvement using both CPUs and GPUs can be as high as 50% compared to using either a CPU core or a GPU. Yasuhiko Ogata, Toshio Endo, Naoya Maruyama, Satoshi Matsuoka |
IPDPS | 4 |
| 2008 | Locality aware MPI communication on a commodity opto-electronic hybrid networkabstractFuture supercomputers with millions of processors would pose significant challenges in their interconnection networks due to difficulty in design constraints such as space, cable length, cost, power consumption, etc. Instead of huge switches or bisection bandwidth restricted topologies such as a torus, we propose a network which utilizes both fully-connected lower-bandwidth electronic packet switching (EPS) network and low-power optical circuit switching (OCS) network. Optical circuits, connected sparingly to only a limited set of nodes to conserve power and cost, are used in a supplemental fashion as “shortcut” routes only when a node communicates substantially across EPS switches, while short latency communication is handled by EPS only. Our MPI inter-node communication algorithm accommodates for such a network by appropriate scheduling of nodes according to application communication patterns, in particular utilizing relatively high EPS local switch bandwidth to forward messages to nodes with optical connections for shortcutting in order to maximize overall throughput. Simulation studies confirm that our proposal effectively avoids contentions in the network in high-bandwidth applications with nominal additions of optical circuitry to existing machines. Shin'ichiro Takizawa, Toshio Endo, Satoshi Matsuoka |
IPDPS | 3 |
| 2008 | Bandwidth intensive 3-D FFT kernel for GPUs using CUDAabstractMost GPU performance “hypes” have focused around tightly-coupled applications with small memory bandwidth requirements e.g., N-body, but GPUs are also commodity vector machines sporting substantial memory bandwidth; however, effective programming methodologies thereof have been poorly studied. Our new 3-D FFT kernel, written in NVIDIA CUDA, achieves nearly 80 GFLOPS on a top-end GPU, being more than three times faster than any existing FFT implementations on GPUs including CUFFT. Careful programming techniques are employed to fully exploit modern GPU hardware characteristics while overcoming their limitations, including on-chip shared memory utilization, optimizing the number of threads and registers through appropriate localization, and avoiding low-speed stride memory accesses. Our kernel applied to real applications achieves orders of magnitude boost in power&cost vs. performance metrics. The off-card bandwidth limitation is still an issue, which could be alleviated somewhat with application kernels confinement within the card, while ideal solution being facilitation of faster GPU interfaces. Akira Nukada, Yasuhiko Ogata, Toshio Endo, Satoshi Matsuoka |
SC | 4 |
| 2008 | Intelligent data staging with overlapped execution of grid applications
Yuya Machida, Shin'ichiro Takizawa, Hidemoto Nakada, Satoshi Matsuoka |
Future Gener. Comput. Syst. | 4 |
| 2007 | High-Performance MPI Broadcast Algorithm for Grid Environments Utilizing Multi-lane NICsabstractThe performance of MPI collective operations, such as broadcast and reduction, is heavily affected by network topologies, especially in grid environments. Many techniques to construct efficient broadcast trees have been proposed for grids. On the other hand, recent high performance computing nodes are often equipped with multi-lane network interface cards (NICs), most previous collective communication methods fail to harness effectively. Our new broadcast algorithm for grid environments harnesses almost all downward and upward bandwidths of multi-lane NICs; A message to be broadcast is split into two pieces, which are broadcast along two independent binary trees in a pipelined fashion, and swapped between both trees. The salient feature of our algorithm is generality; it works effectively on both large clusters and grid environments. It can be also applied to nodes with a single NIC, by making multiple sockets share the NIC. Experimentations on a emulated network environment show that we achieve higher performance than traditional methods, regardless of network topologies or the message sizes. Tatsuhiro Chiba, Toshio Endo, Satoshi Matsuoka |
CCGRID | 3 |
| 2007 | Virtual Clusters on the Fly - Fast, Scalable, and Flexible InstallationabstractOne of the advantages in virtualized computing clusters compared to traditional shared HPC environments is their ability to accommodate user-specific system customization. However, past attempts to providing virtual clusters are not scalable with increasing number of VMs, nor do they allow fine-grained customization of VMs, assuming that preconfigured VM images are always available on the grid. We propose a new virtual cluster installation technique that achieves efficiency and scalability, and yet simultaneously fine-grained customizability. It allows the user to create VMs on the fly for fine-grained customization of VMs, and pipelined data transfer for scalable installation with increasing number of VMs. To achieve efficiency in the presence of such full customization, it automatically caches frequently-constructed virtual disk images to save software installation time in common cases. Our experimental studies using a prototype implementation show that installation of a 190-node virtual cluster can be done in 40 seconds. From this result along with a scalability study, we estimate that installation of a 1000-node virtual cluster could be done in less than two minutes. Hideo Nishimura, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 3 |
| 2007 | Grid'BnB : A Parallel Branch and Bound Framework for Grids
Denis Caromel, Alexandre di Costanzo, Laurent Baduel, Satoshi Matsuoka |
HiPC | 4 |
| 2007 | A Peer-to-Peer Infrastructure for Autonomous Grid MonitoringabstractModern grids have become very complex by their size and their heterogeneity. It makes the deployment and maintenance of systems a difficult task requiring lots of efforts from administrators and programmers. Our goal is to investigate the concepts that underlie autonomic computing systems, especially for grid environment. We believe that peer-to-peer overlay networks are a valuable basis to support some of the main issues of autonomic computing in the particular case of grids. This article presents the construction of an autonomous, decentralized, scalable, and efficient grid monitoring system. The components of this application negotiate through a peer-to-peer network in order to provide autonomic behaviors and exchange data. We present a solution based on a gossip broadcast protocol upon a hierarchical, directed, and acyclic graph to rapidly diffuse information in the system while limiting the number of messages. The software architecture is detailed, and then the first results of its performance are presented and analyzed. Laurent Baduel, Satoshi Matsuoka |
IPDPS | 2 |
| 2007 | ABARIS: An Adaptable Fault Detection/Recovery Component Framework for MPIsabstractLong-running MPI applications on clusters and grids that are prone to node and network failures, motivates the use of fault tolerant MPI implementations. However, previous fault tolerant MPIs lack the ability to allow the user to easily choose appropriate fault recovery strategies according to the execution environment, independent of the application codes-rather, the user often had to hard-code restoration strategies in accordance to diverse sets of fault patterns, which could be numerous: for instance, if the fault is transient to a particular process, we merely have to restart the process on the same computing node; on the other hand, if the fault is due to repetitive hardware unreliability, we must migrate the process to a new node in its recovery. ABARIS is our new fault/recovery model aware component framework for MPI, where users can customize MPI fault detection and recovery algorithms according to their application and execution environmental requirements by merely selecting appropriate fault/recovery components, independent of the application code. Currently, the ARABIS framework prototype is implemented on top of MPICH-P4MPD. Preliminary evaluation of the prototype using NPB on our MPI fault simulator demonstrates that overhead compared to the original MPICH-P4MPD is almost negligible (less than 1%) under normal execution, and when faults occur, appropriate selections and pairings of fault model and recovery method components for corresponding to the execution environment is significant to the overall execution time. Hideyuki Jitsumoto, Toshio Endo, Satoshi Matsuoka |
IPDPS | 3 |
| 2007 | A Decentralized, Scalable, and Autonomous Grid Monitoring System
Laurent Baduel, Satoshi Matsuoka |
OPODIS | 2 |
| 2007 | Weak typed Böhm theorem on IMLL
Satoshi Matsuoka |
Ann. Pure Appl. Log. | 1 |
| 2006 | MegaProto/E: power-aware high-performance cluster with commodity technologyabstractIn our research project named "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we have been developing a prototype cluster not based on ASIC or FPGA but instead only using commodity technology. Its packaging is extremely compact and dense, and its performance/power ratio is very high. Our previous prototype system named "MegaProto" demonstrated that one cluster unit, which consists of 16 commodity low-power processors, can be successfully implemented on just 1U height chassis and it is capable of up to 2.8 times higher performance/power ratio than ordinary high-performance dual-Xeon 1U server units. We have improved MegaProto by replacing the CPU and enhancing the I/O performance. The new cluster unit named "MegaProto/E" with 16 Transmeta Efficeon processors achieves 32 GFlops of peak performance, which is 2.2-fold greater than that of the original one. The cluster unit is equipped with an independent dual network of Gigabit Ethernet, including dual 24-port switches. The maximum power consumption of the cluster unit is 320 W, which is comparable with that of today's high-end PC servers for high performance clusters. Performance evaluation using NPB kernels and HPL shows that the performance of MegaProto/E exceeds that of a dual-Xeon server in all the benchmarks, and its performance ratio ranges from 1.3 to 3.7. These results reveal that our solution of implementing a number of ultra low-power processors in compact packaging is an excellent way to achieve extremely high performance in applications with a certain degree of parallelism. We are now building a multi-unit cluster with 128 CPUs (8 units) to prove that this advantage still holds with higher scalability Taisuke Boku, Mitsuhisa Sato, Daisuke Takahashi, Hiroshi Nakashima, Hiroshi Nakamura, Satoshi Matsuoka, Yoshihiko Hotta |
IPDPS | 6 |
| 2006 | Profile-based optimization of power performance by using dynamic voltage scaling on a PC clusterabstractCurrently, several of the high performance processors used in a PC cluster have a DVS (dynamic voltage scaling) architecture that can dynamically scale processor voltage and frequency. Adaptive scheduling of the voltage and frequency enables us to reduce power dissipation without a performance slowdown during communication and memory access. In this paper, we propose a method of profiled-based power-performance optimization by DVS scheduling in a high-performance PC cluster. We divide the program execution into several regions and select the best gear for power efficiency. Selecting the best gear is not straightforward since the overhead of DVS transition is not free. We propose an optimization algorithm to select a gear using the execution and power profile by taking the transition overhead into account. We have built and designed a power-profiling system, PowerWatch. With this system we examined the effectiveness of our optimization algorithm on two types of power-scalable clusters (Crusoe and Turion). According to the results of benchmark tests, we achieved almost 40% reduction in terms of EDP (energy-delay product) without performance impact (less than 5%) compared to results using the standard clock frequency. Yoshihiko Hotta, Mitsuhisa Sato, Hideaki Kimura 0003, Satoshi Matsuoka, Taisuke Boku, Daisuke Takahashi |
IPDPS | 4 |
| 2006 | Design and Implementation of NAREGI SuperScheduler Based on the OGSA Architecture
Satoshi Matsuoka, Masayuki Hatanaka, Yasumasa Nakano, Yuji Iguchi, Toshio Ohno, Kazushige Saga, Hidemoto Nakada |
J. Comput. Sci. Technol. | 1 |
| 2005 | MegaProto: 1 TFlops/10kW Rack Is Feasible Even with Only Commodity TechnologyabstractIn our research project "Mega-Scale Computing Based on Low-Power Technology and Workload Modeling", we claim that a million-scale parallel system could be built with densely mounted low-power commodity processors. "MegaProto" is a proof-of-concept low-power and highperformance cluster build only with commodity components to implement this claim. A one-rack system is composed of 32 motherboard "cluster units" of 1 U-height and commodity switches to interconnect them mutually as well as with other racks. Each cluster unit houses 16 low-power dollarbill- sized commodity PC-architecture daughterboards, together with a high bandwidth, 2 Gbps per processor embedded switched network based on Gigabit Ethernet. The peak performance of a one-rack system is 0.48 TFlops for the first version and will improve to 1.02 TFlops in the second version through a processor/daughterboard upgrade. The system consumes about 10 kW or less per rack, resulting in 100 MFlops/W power efficiency with a power-aware intrarack network of 32 Gbps bisection bandwidth, while additional 2.4 kW will boost this to sufficiently large 256 Gbps. Performance studies show that even the first version significantly outperforms a conventional high-end 1U server comprised of dual power-hungry processors in a majority of NPB programs. It is also investigated how the current automated DVS control could save power for the HPC parallel programs along with its limitation. Hiroshi Nakashima, Hiroshi Nakamura, Mitsuhisa Sato, Taisuke Boku, Satoshi Matsuoka, Daisuke Takahashi, Yoshihiko Hotta |
SC | 5 |
| 2005 | Japanese Computational Grid Research Project: NAREGIabstractThe National Research Grid Initiative (NAREGI) is one of the major Japanese national IT projects currently being conducted. NAREGI will cover the period 2003-2007, and collaboration among industry, academia, and the government will play a key role in its success. The Center for Grid Research and Development has been established as a center for R&D of high-performance, scalable grid middleware technologies, which are aimed at enabling major computing centers to host grids over high-speed networks to provide a future computational infrastructure for scientific and engineering research in the 21st century. As an example of utilizing such grid computing technologies, the Center for Application Research and Development is conducting research on leading-edge, grid-enabled nanoscience and nanotechnology simulation applications, which will lead to the discovery and development of new materials and next-generation nanodevices. These two centers are collaborating to establish daily research use of a multiteraflop grid testbed infrastructure, which will be built to demonstrate the advantages of grid technologies for future applications in all areas of science and engineering. Satoshi Matsuoka, S. Shinjo, Mutsumi Aoyagi, Satoshi Sekiguchi, Hitohide Usami, Kenichi Miura |
Proc. IEEE | 1 |
| 2004 | A Java-based programming environment for hierarchical Grid: JojoabstractDespite recent developments in higher-level middleware for the Grid supporting high level of ease-of-programming, hurdles for widespread adoption of Grids remain high, due to (1) assumption of peer-to-peer connectivity of all Grid nodes, as well as (2) lack of scalable programming and deployment support. We propose a Java-based programming environment for a hierarchically organized Grid named Jojo, that allow seamless utilization of privately addressed clusters. Jojo provides several features, including secure private remote invocation using Globus GRAM and ssh/rsh to privately addressed nodes in clusters, intuitive message passing API suitable for overlapped execution using multiple threads, and automatic user/system program staging. Using Jojo, users can easily construct and execute parallel distributed applications on the Grid. We show the design and implementation of its programming API, a working example, as well as preliminary performance evaluation results that prove the effectiveness of hierarchal execution. Hidemoto Nakada, Satoshi Matsuoka |
CCGRID | 2 |
| 2003 | Evaluation of the inter-cluster data transfer on Grid environmentabstractHigh-performance peer-to-peer transfer between clusters will be fundamental technology base for various Grid middleware, such as large-scale data transfer in DataGrid settings, or collective communication in Grid-wide MPIs. There, two major factors are involved: on one hand network pipes with large RTT /spl times/ bandwidth typically become data-starved, resulting in bandwidth loss; on the other hand when multiple nodes on the clusters attempt simultaneous transfer, the network pipe could become saturated, resulting in packet loss which again may result in bandwidth degradation in large RTT /spl times/ bandwidth networks. By dynamically and automatically adjusting transfer parameters between the two clusters, such as the number of network nodes, number of socket stripes, we could achieve optimal bandwidth even when the network is under heavy contention. In order to arrive at a proper performance model for automated adjustment, we have conducted several simulations by which we have discovered that such automatic tuning would beneficial, but the ideal number of network pipes does not exactly match the simple transfer model of traditional peer-to-peer settings between single nodes. Shoji Ogura, Satoshi Matsuoka, Hidemoto Nakada |
CCGRID | 2 |
| 2003 | Preliminary Evaluation of Dynamic Load Balancing Using Loop Re-partitioning on Omni/SCASHabstractIncreasingly large-scale clusters of PC/WS continue to become majority platform in HPC field. Such a commodity cluster environment, there may be incremental upgrade due to several reasons, such as rapid progress in processor technologies, or user needs and it may cause the performance heterogeneity between nodes from which the application programmer will suffer as load imbalances. To overcome these problems, some dynamic load balancing mechanisms are needed. In this paper, we report our ongoing work on dynamic load balancing extension to Omni/SCASH which is an implementation of OpenMP on Software Distributed Shared Memory, SLASH. Using our dynamic load balancing mechanisms, we expect that programmers can have load imbalances adjusted automatically by the runtime system without explicit definition of data and task placements in a commodity cluster environment with possibly heterogeneous performance nodes. Yoshiaki Sakae, Mitsuhisa Sato, Satoshi Matsuoka, Hiroshi Harada |
CCGRID | 3 |
| 2003 | Performance Analysis of Scheduling and Replication Algorithms on Grid Datafarm Architecture for High-Energy Physics ApplicationsabstractData Grid is a Grid for ubiquitous access and analysis of large-scale data. Because Data Grid is in the early stages of development, the performance of its petabyte-scale models in a realistic data processing setting has not been well investigated. By enhancing our Bricks Grid simulator to accommodated Data Grid scenarios, we investigate and compare the performance of different Data Grid models. These are categorized mainly as either central or tier models; they employ various scheduling and replication strategies under realistic assumptions of job processing for CERN LHC experiments on the Grid Datafarm system. Our results show that the central model is efficient but that the tier model, with its greater resources and its speculative class of background replication policies, are quite effective and achieve higher performance, while each tier is smaller than the central model. Atsuko Takefusa, Osamu Tatebe, Satoshi Matsuoka, Youhei Morita |
HPDC | 3 |
| 2003 | Ninf-G: A Reference Implementation of RPC-based Programming Middleware for Grid Computing
Yoshio Tanaka, Hidemoto Nakada, Satoshi Sekiguchi, Toyotaro Suzumura, Satoshi Matsuoka |
J. Grid Comput. | 5 |
| 2002 | Grid Datafarm Architecture for Petascale Data Intensive ComputingabstractThe Grid Datafarm (Gfarm) architecture is designed for global petascale data-intensive computing. It provides a global parallel filesystem with online petascale storage, scalable I/O bandwidth, and scalable parallel processing, and it can exploit local I/O in a grid of clusters with tens of thousands of nodes. Gfarm parallel I/O APIs and commands provide a single filesystem image and manipulate filesystem metadata consistently. Fault tolerance and load balancing are automatically managed by file duplication or recomputation using a command history log. Preliminary performance evaluation has shown scalable disk I/O and network bandwidth on 64 nodes of the Presto III Athlon cluster. The Gfarm parallel I/O write and read operations has achieved data transfer rates of 1.74 GB/s and 1.97 GB/s, respectively, using 64 cluster nodes. The Gfarm parallel file copy reached 443 MB/s with 23 parallel streams on the Myrinet 2000. The Gfarm architecture is expected to enable petascale data-intensive Grid computing with an I/O bandwidth scales to the TB/s range and scalable computational power. Osamu Tatebe, Youhei Morita, Satoshi Matsuoka, Noriyuki Soda, Satoshi Sekiguchi |
CCGRID | 3 |
| 2002 | First Light of the Earth Simulator and Its PC Cluster ApplicationsabstractThe Earth Simulator (ES) is the largest parallel vector processor in the world that is mainly dedicated to large-scale simulation studies of global change. Development of the ES system started in 1997 and was completed at the end of February, 2002. The system consists of 640 processor nodes that are connected via a very fast single-stage crossbar network (12.3 GB/s). The total peak performance and main memory of the system are 40 TFLOPS and 10 TB, respectively. Studies to evaluate the performance of the ES were made using an atmospheric circulation model Afes (Atmospheric General Circulation Model for ES) and LINPACK benchmark test. The sustained performance of Afes for T1279L96 (the equivalent horizontal resolution given by T1279 is about 10 km and the total number of layers is 96) was as high as 14.5 TFLOPS on a half system of the ES with 2,560 PEs (320 nodes). The sustained-to-peak performance ratio was 70.8%. The ES also achieved a LINPACK world record of 35.86 TFLOPS. This rating exceeded the previous record, set by the ASCI White, by about 5 times. The Earth Simulator is now running. Huge amounts of output data will arise from the huge computer system. For example, the data volume of simulation results from the Afes is of the order of 10-100 TB. In the phase of operation, management of huge output datafiles and interactive visual monitoring of many terabytes of simulation results are extremely important for the ES. The ES has introduced a prototype PC cluster to seek the best solution to these problems. The PC cluster comprises 64 PCs that are interconnected with a Myrinet2000 switch. Each PC has a Pentium III (1 GHz), 1 GB of main memory and 120 GB of disk space. An outline of the Earth Simulator system, recent results on performance evaluation using real applications and the LINPACK benchmark test, and an outline of the PC cluster system are presented. Keiji Tani, Takayuki Aoki, Satoshi Matsuoka, Satoru Ohkura, Hitoshi Uehara, Tetsuo Aoyagi |
CLUSTER | 3 |
| 2002 | Evaluating Web Services Based Implementations of GridRPCabstractGridRPC is a class of Grid middleware for scientific computing. Interoperability has been an important issue, because current GridRPC systems each employ its own protocol. Web services, where XML-based standards such as SOAP and WSDL are expected to see widespread use, could be the medium of interoperability; however it is not clear if 1) XML-based schemas have sufficient expressive power for GridRPC, and 2) whether performance could be made sufficient. Our experiments indicate that the use of such technologies are more promising. than previously reported. Although a naive implementation of SOAP-based GridRPC has severe performance overhead, application of a series of optimizations improves performance. However encoding of various features of GridRPC proved to be somewhat difficult due to WSDL limitations. The results show that GridRPC systems can be based on Web technologies, but there needs to be work to extend WSDL specifications, possibly impacting OGSA-based Grid services directions. Satoshi Shirasuna, Hidemoto Nakada, Satoshi Matsuoka, Satoshi Sekiguchi |
HPDC | 3 |
| 2001 | Grid RPC meets Data Grid: Network Enabled Services for Data Farming on the GridabstractThe Computational Grid[1] is a promising platform for running large-scale scientific applications. It provides a base software infrastructure that allows for the development of middleware aimed at deploying applications on Grid resources. The question is, how do you program it---in this regard, Network-Enabled Server (NES) paradigm, which enables Grid-based RPC, or GridRPC for short is a good candidate as a viable Grid middleware that offers a simple yet powerful programming paradigm for programming on the Grid. Several systems that facilitate whole or parts of the paradigm are already in existence, such as Neos[7], Netsolve[3], Nimrod/G[4], Ninf[2], and RCS[6], and we feel that pursuit of a common design in GridRPC, as had been done for MPI for message passing, will bring benefits of standardized programming model to the Grid world. This talk will introduce the NES/Grid RPC features, discuss early user experiences, and touch upon the Grid Data Farm project, based on Grid RPC, which involving processing Petabytes of collider accelerator data streaming over the Euro-Japanese link with thousands-node scale cluster possibly spread over several Japanese institutions. Compared to traditional RPC systems, such as CORBA, designed for applications that facilitate nonscientific applications, GridRPC systems offer features and capabilities that make it easy to program mediumto coarse-grained, task parallel applications that involve hundreds to thousands or more high-performance nodes, either concentrated as a tightly coupled cluster, or a set of them spread over a wide-area network. Such applications will often require handling of shipping megabytes of multidimensional array data in a user-transparent and efficient way, as well as requiring the support of RPC calls that range anywhere from 100s of milliseconds up to several days or even weeks. There are other necessary features of Grid RPC systems such as dynamic resource discovery, dynamic load balancing, fault tolerance, security (multisite authentication, delegation of authentication, adapting to multiple security policies, etc.), easy-to-use client/server management, firewall and private address considerations, remote large file and I/O support etc. These features are essentially what is needed for the Grid RPC systems to execute well on the Grid---features either missing or incomplete in traditional `closed world’ RPC systems ---and in fact are what are provided by lower level Grid substrates such as Condor[10], Globus[8], and Legion[9]. As such GridRPC systems either provide these features themselves, or builds upon the features provided by such substrates. Satoshi Matsuoka |
CCGRID | 1 |
| 2001 | A Study of Deadline Scheduling for Client-Server Systems on the Computational GridabstractThe Computational Grid is a promising platform for the deployment of various high-performance computing applications. A number of projects have addressed the idea of software as a service on the network. These systems usually implement client-server architectures with many servers running on distributed Grid resources and have commonly been referred to as network-enabled servers (NES). An important question is that of scheduling in this multi-client multi-server scenario. Note that in this context most requests are computationally intensive as they are generated by high-performance computing applications. The Bricks simulation framework has been developed and extensively used to evaluate scheduling strategies for NES systems. The authors first present recent developments and extensions to the Bricks simulation models. They discuss a deadline scheduling strategy that is appropriate for the multi-client multi-server case, and augment it with "Load Correction" and "Fallback" mechanisms which could improve the performance of the algorithm. We then give Bricks simulation results. The results show that future NES systems should use deadline scheduling with multiple fallbacks and it is possible to allow users to make a trade-off between failure-rate and cost by adjusting the level of conservatism of deadline scheduling algorithms. Atsuko Takefusa, Satoshi Matsuoka, Henri Casanova, Francine Berman |
HPDC | 2 |
| 2001 | An Evaluation of Multiple Pointing Input Systems
Kentaro Fukuchi, Satoshi Matsuoka |
INTERACT | 2 |
| 2001 | A Jini-based computing portal systemabstractJiPANG(A Jini-based Portal Augmenting Grids) is a portal system and a toolkit which provides uniform access interface layer to a variety of Grid systems, and is built on top of Jini distributed object technology. JiPANG performs uniform higher-level management of the computing services and resources being managed by individual Grid systems such as Ninf, NetSolve, Globus, etc. In order to give the user a uniform interface to the Grids JiPANG provides a set of simple Java APIs called the JiPANG Toolkits, and furthermore, allows the user to interact with Grid systems, again in a uniform way, using the JiPANG Browser application. With JiPANG, users need not install any client packages before-hand to interact with Grid systems, nor be concerned about updating to the latest version. Such uniform, transparent services available in a ubiquitous manner we believe is essential for the success of Grid as a viable computing platform for the next generation. Toyotaro Suzumura, Satoshi Matsuoka, Hidemoto Nakada |
SC | 2 |
| 2001 | Towards performance evaluation of high-performance computing on multiple Java platforms
Satoshi Matsuoka, Shigeo Itou |
Future Gener. Comput. Syst. | 1 |
| 2000 | OpenJIT: An Open-Ended, Reflective JIT Compiler Framework for Java
Hirotaka Ogawa, Kouya Shimura, Satoshi Matsuoka, Fuyuhiko Maruyama, Yukihiko Sohda, Yasunori Kimura |
ECOOP | 3 |
| 2000 | Are Global Computing Systems Useful? Comparison of Client-server Global Computing Systems Ninf, NetSolve Versus CORBabstractRecent developments of global computing systems such as Ninf, NetSolve and Globus have opened up the opportunities for providing high-performance computing services over wide-area networks. However, most research focused on the individual architectural aspects of the system, or application deployment examples, instead of the necessary characteristics such systems should intrinsically satisfy, nor how such systems relate with each other. Our comparative study performs deployment of example applications of network-based libraries using Ninf, NetSolve, and CORBA systems. There, we discover that dedicated systems for global computing such as Ninf and NetSolve have management, programmability, and it does not suffer performance disadvantages over more generic distributed computing capabilities provided by CORBA. Such results indicate the advantage of dedicated global computing systems over general systems, stemming further basic research is necessary across multiple systems to identify the ideal software architectures for global computing. Toyotaro Suzumura, Takayuki Nakagawa, Satoshi Matsuoka, Hidemoto Nakada, Satoshi Sekiguchi |
IPDPS | 3 |
| 1999 | Overview of a Performance Evaluation System for Global Computing Scheduling AlgorithmsabstractWhile there have been several proposals of high-performance global computing systems, scheduling schemes for the systems have not been well investigated. The reason is difficulties of evaluation by large-scale benchmarks with reproducible results. Our Bricks performance evaluation system allows the analysis and comparison of various scheduling schemes in a typical high-performance global computing setting. Bricks can simulate various behaviors of global computing systems, especially the behavior of networks and resource scheduling algorithms. Moreover, Bricks is partitioned into components such that not only can its constituents be replaced to simulate various different system algorithms, but it also allows the incorporation of existing global computing components via its foreign interface. To test the validity of the latter characteristics, we incorporated the NWS (Network Weather Service) system, which monitors and forecasts global computing systems behavior. Experiments were conducted by running NWS under a real environment versus a Bricks-simulated environment, given the observed parameters of the real environment. We observed that Bricks behaved in the same manner as the real environment, and NWS also behaved similarly, making quite comparative forecasts under both environments. Atsuko Takefusa, Satoshi Matsuoka, Hidemoto Nakada, Kento Aida, Umpei Nagashima |
HPDC | 2 |
| 1999 | Teddy: A Sketching Interface for 3D Freeform DesignabstractArticle Teddy: a sketching interface for 3D freeform design Share on Authors: Takeo Igarashi University of Tokyo University of TokyoView Profile , Satoshi Matsuoka Tokyo Institute of Technology Tokyo Institute of TechnologyView Profile , Hidehiko Tanaka University of Tokyo University of TokyoView Profile Authors Info & Claims SIGGRAPH '99: Proceedings of the 26th annual conference on Computer graphics and interactive techniquesJuly 1999 Pages 409–416https://doi.org/10.1145/311535.311602Published:01 July 1999 830citation3,369DownloadsMetricsTotal Citations830Total Downloads3,369Last 12 Months47Last 6 weeks13 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Takeo Igarashi, Satoshi Matsuoka, Hidehiko Tanaka |
SIGGRAPH | 2 |
| 1998 | A Performance Evaluation Model for Effective Job Scheduling in Global Computing SystemsabstractThe paper proposes a performance evaluation model for effective job scheduling in global computing systems. The proposed model represents a global computing system by a queueing network, in which servers and networks are represented by queueing systems. Evaluation of the proposed model showed that the model could simulate behavior of an actual global computing system and job scheduling on the system effectively. Kento Aida, Atsuko Takefusa, Hidemoto Nakada, Satoshi Matsuoka, Umpei Nagashima |
HPDC | 4 |
| 1998 | Popup Vernier: A Tool for Sub-Pixel-Pitch Dragging with Smooth Mode TransitionabstractDragging is one of the most useful and popular techniques in direct manipulation graphical user interfaces.However, dragging has inherent restrictions caused by pixel resolution of a display.Although in some situations the restriction could be negligible, certain kinds of applications, e.g., real world applications where the range of adjustable parameters vastly exceed the screen resolution, require sub-pixel-pitch dragging.We propose a sub-pixel-pitch dragging tool, popup vernier, plus a methodology to transfer smoothly into 'vernier mode' during dragging.A popup vernier consists of locally zoomed grids and vernier scales displayed around them.Verniers provide intuitive manipulation and feedback of fine grain dragging, in that pixel-pitch movements of the grids represent sub-pixel-pitch movements of a dragged object, and the vernier scales show the object's position at a sub-pixel accuracy.The effectiveness of our technique is verified with a proposed evaluation measure that captures the smoothness of transition from standard mode to vernier mode, based on the Fitts' law. Yuji Ayatsuka, Jun Rekimoto, Satoshi Matsuoka |
ACM Symposium on User Interface Software and Technology | 3 |
| 1998 | Ninflet: a migratable parallel objects framework using JavaabstractNinflet is a Java-based global computing system that builds on our experiences with the Ninf system which facilitated RPC-based computing of numerical tasks in a wide-area network. The goal of Ninflet is to become a new generation of concurrent object-oriented systems which harness abundant idle computing powers, and also seamlessly integrate global as well as local network parallel computing. Ninflet is designed to make use of Java features to implement important features in global computing, such as resource allocation, inter-Ninflet communication, security, checkpointing, object migration, and easy server management via HTTP. © 1998 John Wiley & Sons, Ltd. Hiromitsu Takagi, Satoshi Matsuoka, Hidemoto Nakada, Satoshi Sekiguchi, Mitsuhisa Sato, Umpei Nagashima |
Concurr. Pract. Exp. | 2 |
| 1998 | Ninf and PM: Communication libraries for global computing and high-performance cluster computing
Mitsuhisa Sato, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa, Satoshi Sekiguchi, Hidemoto Nakada, Satoshi Matsuoka, Umpei Nagashima |
Future Gener. Comput. Syst. | 7 |
| 1997 | A Methodology for Specifying Data Distribution Using Only Standard Object-Oriented FeaturesabstractObject-oriented class frameworks are promising alternatives to traditional languages and their parallelizing compilers due to their higher at&action facilities for encapsulating parallelism and distribution.We claim that language feature6 already available in cxisiting object-oriented languages such aa C* for constructing class t&reworks can be maximally exploited for achieving 6cp6ration of parallelism and data distribution, without language extaxsions or ad-hoc methodologies.We first model parallel computation a6 crtiered application of a function onto 6tructurcd elements.and, ba6cd on the model, we construct a class fmmcwcrk which formulates 1) data distribution 66 hierarchical decomposition of the 6tructured elements to layered objects and 2) parallelism as nested rruversuls across these objects.We evaluate the feasibiity of our proposal on the Pujitsu APlOtXl parallel computer by extending the EPEE class framework. Naohito Sato, Satoshi Matsuoka, Jean-Marc Jézéquel, Akinori Yonezawa |
International Conference on Supercomputing | 2 |
| 1997 | In Search for an Ideal Computer-Assisted Drawing System
Takeo Igarashi, Sachiko Kawachiya, Satoshi Matsuoka, Hidehiko Tanaka |
INTERACT | 3 |
| 1997 | Multi-client LAN/WAN Performance Analysis of Ninf: a High-Performance Global Computing SystemabstractRapid increase in speed and availability of network of supercomputers is making high-performance global computing possible, including our Ninf system. However, critical issues regarding system performance characteristics in global computing have been little investigated, especially under multi-client, multi-site WAN settings. In order to investigate the feasibility of Ninf and similar systems, we conducted benchmarks under various LAN and WAN environments, and observed the following results: 1) Given sufficient communication bandwidth, Ninf performance quickly overtakes client local performance, 2) current supercomputers are sufficient platforms for supporting Ninf and similar systems in terms of performance and OS fault resiliency, 3) for a vector-parallel machine (Cray J90), employing optimized data-parallel library is a better choice compared to conventional task-parallel execution employed for non-numerical data servers, 4) computationally intensive tasks such as EP can readily be supported under the current Ninf infrastructure, and 5) for communication-intensive applications such as Linpack, server CPU utilization dominates LAN performance, while communication bandwidth dominates WAN performance, and furthermore, aggregate bandwidth could be sustained for multiple clients located at different Internet sites; as a result, distribution of multiple tasks to computing servers on different networks would be essential for achieving higher client-observed performance. Our results are not necessarily restricted to the Ninf system, but rather, would be applicable to other similar global computing systems. Atsuko Takefusa, Satoshi Matsuoka, Hirotaka Ogawa, Hidemoto Nakada, Hiromitsu Takagi, Mitsuhisa Sato, Satoshi Sekiguchi, Umpei Nagashima |
SC | 2 |
| 1997 | Interactive Beautification: A Technique for Rapid Geometric DesignabstractWe propose interactive beautification, a technique for rapid geometric design, and introduce the technique and its algorithm with a prototype system Pegasus.The motivation is to solve a problem with current drawing systems: too many complex commands and unintuitiveprocedures to satisfy geometric constraints.The Interactive beautification system receives the user'sfree stroke and beautifies it by considering geometric constraints among segments.A single stroke is beautified one after another, preventing accumulation of recognition errors or catastrophic deformation.Supported geometric constraints include perpendicularity, congruence, symmetry, etc., which were not seen in existing free stroke recognition systems.In addition, the system generates multiple candidates as a result of beautification to solve the problem of ambiguity.Using this technique, the user can draw precise diagrams rapidly satisfying geometric relations without using any editing commands.Interactive beautification is achieved by three sequential processes: 1) inferring underlining geometric constraints based on the spatial relationships among the input stroke and the existing segments, 2) generating multiple candidates by combining inferred constraints appropriately, and 3) evaluating the candidates to find the most plausible candidate and to remove the inappropriate candidates.A user study was performed using the prototypesystem, a commercial CAD tppl, and an OObased drawing system.The result showed that users can draw required diagrams more rapidly and moreprecisely using the prototype system. Takeo Igarashi, Satoshi Matsuoka, Sachiko Kawachiya, Hidehiko Tanaka |
ACM Symposium on User Interface Software and Technology | 2 |
| 1996 | Generalized Local Propagation: A Framework for Solving Constraint Hierarchies
Hiroshi Hosobe, Satoshi Matsuoka, Akinori Yonezawa |
CP | 2 |
| 1996 | OMPI: Optimizing MPI Programs using Partial EvaluationabstractMPI is gaining acceptance as a standard for message-passing in high-performance computing, due to its powerful and flexible support of various communication styles. However, the complexity of its API poses significant software overhead, and as a result, applicability of MPI has been restricted to rather regular, coarse-grained computations. Our OMPI (Optimizing MPI) system removes much of the excess overhead by employing partial evaluation techniques, which exploit static information of MPI calls. Because partial evaluation alone is insufficient, we also utilize template functions for further optimization. To validate the effectiveness for our OMPI system, we performed baseline as well as more extensive benchmarks on a set of application cores with different communication characteristics, on the 64-node Fujitsu AP1000 MPP. Benchmarks show that OMPI improves execution efficiency by as much as factor of two for communication-intensive application core with minimal code increase. It also performs significantly better than previous dynamic optimization technique. Hirotaka Ogawa, Satoshi Matsuoka |
SC | 2 |
| 1996 | Penumbrae for 3D InteractionsabstractNo abstract available. Yuji Ayatsuka, Satoshi Matsuoka, Jun Rekimoto |
ACM Symposium on User Interface Software and Technology | 2 |
| 1995 | Compiling Away the Meta-Level in Object-Oriented Concurrent Reflective Languages Using Partial EvaluationabstractMeta-level programmability is beneficial for parallel/distributed object-oriented computing to improve performance, etc. The major problem, however, is interpretation overhead due to mta-circular interpretation. To solve this problem, we propose a compilation framework for object-oriented concurrent reflective languages using partial evaluation. Since traditional partial evaluators do not allow us to directly deal with meta-circular interpreters written with concurrent objects, we devised techniques such as pre-/post-processing, a new proposed preaction extension to partial evaluation in order to handle side-effects, etc. Benchmarks of a prototype compiler for our language ABCL/R3 indicate that (1) the meta-level interpretation is essentially 'compiled away,' and (2) mta-level optimizations in a parallel application, running on a Fujitsu MPP AP1000, exhibits only 10--30% overhead compared to the hand-crafted source-level optimization in a non-reflective language. Hidehiko Masuhara, Satoshi Matsuoka, Kenichi Asai, Akinori Yonezawa |
OOPSLA | 2 |
| 1994 | Efficient parallel global garbage collection on massively parallel computersabstractOn distributed-memory high-performance massively parallel computers (MPPs) where processors are interconnected by an asynchronous network, efficient garbage collection (GC) becomes difficult, due to inter-node references and references within pending, unprocessed messages. Our parallel global GC algorithm (1) takes advantage of reference locality, (2) efficiently traverses references over nodes, (3) admits a minimum pause time for the ongoing computations, and (4) has been shown to scale up to 1024-node MPPs. The algorithm employs a global weight counting scheme to substantially reduce message traffic. Two methods for confirming the arrival of pending messages are used: one counts the number of messages and the other uses network 'bulldozing'. Performance evaluations in actual implementations on a multicomputer with from 32 to 1024 nodes, the Fujitsu AP1000, reveals various favorable properties of the algorithm.> Tomio Kamada, Satoshi Matsuoka, Akinori Yonezawa |
SC | 2 |
| 1994 | Interactive Generation of Graphical User Interfaces by Multiple Visual ExamplesabstractThe construction of application-specific Graphical User Interfaces (GUI) still needs considerable programming partly because the mapping between application data and its visual representation is complicated. This study proposes a system which generates GUIs by generalizing multiple sets of application data and its visualization examples. The most notable characteristic of the system is that programmers can interactively modify the mapping by “correcting” the system-generated visualization examples that represent the system's current notion of programmer's intentions. Conflicting mappings are automatically resolved via the use of constraint hierarchies. Ken Miyashita, Satoshi Matsuoka, Shin Takahashi, Akinori Yonezawa |
ACM Symposium on User Interface Software and Technology | 2 |
| 1993 | Highly Efficient and Encapsulated Re-use of Synchronization Code in Concurrent Object-Oriented LanguagesabstractRe-use of synchronization code in concurrent OOlanguages has been considered difficult due to inheritance anomaly, which we minimize with our new proposal. Designed with high practicality in mind, we propose language primitives (plus their implementation) with the following characteristics: (1) it allows multiple synchronization schemes---the language schemes for programming synchronization---to coexist and be integrated, (2) re-use of synchronization code is done similarly to sequential OO-languages for user familiarity, (3) it offers high degree of encapsulation---even synchronization schemes could be encapsulated in superclasses in many cases, and (4) it can be efficiently implemented on conventional MPPs. We demonstrate the effectiveness of our proposal with solutions to the example inheritance anomaly cases from [16]. We also give an overview of the implementation architecture, along with preliminary benchmarks. The proposed language primitives are being incorporated into our AB... Satoshi Matsuoka, Kenjiro Taura, Akinori Yonezawa |
OOPSLA | 1 |
| 1993 | An Efficient Implementation Scheme of Concurrent Object-Oriented Languages on Stock MulticomputersabstractSeveral novel techniques for efficient implementtion of concurrent object-oriented languages on general purpose, stock multicomputers are presented. These techniques have been developed in implementing our concurrent object-oriented language ABCL on a Fujitsu Laboratory's experimental multicomputer AP1000 consisting of 512 SPARC chips. The propsed intra-node scheduling mechanism reduces the cost of local message passing. The cost of intra-node asynchronous message passing is about 20 SPARC instructions in the bst case, including locality checking, dynamic method lookup, and scheduling. The minimum latency of asynchronous internode message passing is about 9μs, or about 120 instructions, employing the self-dispatching mechanism independently proposed by Eicken et al. A large scale benchmark which involves 9,000,000 message passings shows 440 times speedup on the 512 nodes system compared to the sequential version of the same algorithm. We rely on simple hardware support for message passing and use no specialized architectural supports for object-oriented computing. Thus, we are able to enjoy the benefits of future progress in standard processor technology. Our result shows that concurrent object-oriented languages can be implemented efficiently on conventional multicomputers. Kenjiro Taura, Satoshi Matsuoka, Akinori Yonezawa |
PPoPP | 2 |
| 1992 | ABCL/onEM-4: a new software/hardware architecture for object-oriented concurrent computing on an extended dataflow supercomputerabstractThe trend towards object-oriented software construction is becoming more and more prevalent, and parallel programming cannot be an exception. In the context of parallel computation, it is often natural to model the computation as message passing between autonomous, concurrently active objects. The problem was, as some previous studies had indicated, that the overhead from message reception to dynamic method dispatching consumes a significant amount of execution time (e.g., as much as 4000 machine cycles or 500 μseconds at 8 MHz block for some language/hardware combination). Our ABCL/onEM-4, a software/hardware implementation architecture for a concurrent object-oriented language, overcomes this problem with technologies such as address-specifiable reactive packet-driven architecture, zero-overhead context switching, and packet-driven allocation of message boxes. Preliminary performance measurements on a real hardware EM-4 confirm our claim, achieving the performance of up to nearly 10 μseconds (130 clocks) total for a remote object-creation followed by a request message send to the created object and a reply reception from the object, for a 12.5 MHz clock speed. Our results indicate that the concurrent object-oriented computational model and languages are highly viable with proper implementational software/hardware architectures. Masahiro Yasugi, Satoshi Matsuoka, Akinori Yonezawa |
ICS | 2 |
| 1992 | Object-Oriented Concurrent Reflective Languages can be Implemented EfficientlyabstractComputational reflection is beneficial in concurrent computing in offering a linguistic mechanism for incorporating user-specific policies. New challenges are (1) how to implement them, and (2) how to do so efficiently. We present efficient implementation schemes for object-oriented concurrent reflective languages using our language ABCL/R2 as an example. The schemes include: efficient lazy creation of metaobjects/meta-groups, partial compilation of scripts (methods), dynamic progression, self-reification, and light-weight objects, all appropriately integrated so that the user-level semantics remain consistent with the meta-circular definition so that the full power of reflection is retained, while achieving practical efficiency. ABCL/R2 exhibits two orders of magnitude speed improvement over its predecessor, ABCL/R, and in fact compares favorably to the ABCL/1 compiler and also C + Sun LWP, neither supporting reflection. 3 To be presented at ACM OOPSLA'92, Vancouver, Canada, Oct. 199... Hidehiko Masuhara, Satoshi Matsuoka, Takuo Watanabe, Akinori Yonezawa |
OOPSLA | 2 |
| 1992 | Declarative Programming of Graphical Interfaces by Visual ExamplesabstractGraphical user interfaces (GUI) provide intuitive and easy means for users to communicate with computers. However, construction of GUI software requires complex programming that is far from being intuitive. Because of the "semantic gap" between the textual application program and its graphical interface, the programmer himself must conceptually maintain the correspondence between the textual programming and the graphical image of the resulting interface. Instead, we propose a programming environment based on the programming by visual example (PBVE) scheme, which allows the GUI designers to "program" visual interfaces for their applications by "drawing" the example visualization of application data with a direct manipulation interface. Our system, TRIP3, realizes this with (1) the bi-directional translation model between the (abstract) application data and the pictorial data of the GUI, and (2) the ability to generate mapping rules for the translation from example application data and ... Ken Miyashita, Satoshi Matsuoka, Shin Takahashi, Akinori Yonezawa, Tomihisa Kamada |
ACM Symposium on User Interface Software and Technology | 2 |
| 1992 | A General Framework for Bidirectional Translation between Abstract and Pictorial DataabstractThe merits of direct manipulation are now widely recognized. However, direct manipulation interfaces incur high cost in their creation. To cope with this problem, we present a model of bidirectional translation between pictures and abstract application data, and a prototype system, TRIP2, based on this model. Using this model, general mapping from abstract data to pictures and from pictures to abstract data is realized merely by giving declarative mapping rules, allowing fast and easy creation of direct manipulation interfaces. We apply the prototype system to the generation of the interfaces for kinship diagrams, Graph Editors, E-R diagrams, and an Othello game. Satoshi Matsuoka, Shin Takahashi, Tomihisa Kamada, Akinori Yonezawa |
ACM Trans. Inf. Syst. | 1 |
| 1991 | Hybrid Group Reflective Architecture for Object-Oriented Concurrent Reflective Programming
Satoshi Matsuoka, Takuo Watanabe, Akinori Yonezawa |
ECOOP | 1 |
| 1991 | A general framework for Bi-directional translation between abstract and pictorial dataabstractand Pictorial Data Shin Takahashi Satoshi Matsuoka Akinori Yonezawa 3 Department of Information Science, University of Tokyo 7-3-1 Hongo, Bunkyo-ku, Tokyo, 113 Japan Tomihisa Kamada y Research and Development, ACCESS CO., LTD. 1-7-1 Sarugaku-cho, Chiyoda-ku, Tokyo, 101 Japan Abstract The merits of direct manipulation(DM) are now widely recognized. However, DM interfaces incur high cost in their creation. To cope with this problem, we present a model of bi-directional translation between internal abstract data of applications and pictures, and create a prototype system TRIP2 based on this model. Using this model, general mapping from abstract data to pictures, and from pictures to abstract data, is realized merely by giving declarative mapping rules, allowing fast and effortless creation of DM interfaces. We also apply the prototype system to the generation of the interfaces for kinship diagrams, graph diagrams, and an Othello game. keywords : Bi-Directional Translation, Direct m... Shin Takahashi, Satoshi Matsuoka, Akinori Yonezawa, Tomihisa Kamada |
UIST | 2 |
| 1989 | Asymptotic Evaluation of Window Visibility
Satoshi Matsuoka, Tomihisa Kamada, Satoru Kawai |
Inf. Process. Lett. | 1 |
| 1988 | Using Tuple Space Communication in Distributed Object-Oriented LanguagesabstractWhen Object-Oriented languages are applied to distributed problem solving, the form of communication restricted to direct message sending is not flexible enough to naturally express complex interactions among the objects. We transformed the Tuple Space Communication Model[29] for better affinity with Object-Oriented computation, and integrated it as an alternative method of communication among the distributed objects. To avoid the danger of potential bottleneck, we formulated an algorithm that makes concurrent pattern matching activities within the Tuple Space possible. Satoshi Matsuoka, Satoru Kawai |
OOPSLA | 1 |