VLDB 2026 Research / reviewers in the wild / expert
Yutaka Ishikawa
dblp:97/6217
· DBLP profile ↗
119ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-2286-9770ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 19 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 since 2021Security and privacy · 9Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Design and Implementation of an Authentication and Authorization Mechanism Based on Unix Domain Sockets
Haruka Kita, Yutaka Ishikawa, Atsuko Takefusa, Masato Oguchi |
COMPSAC | 2 |
| 2026 | The Zero Trust IoT (ZT-IoT) Project
Atsuko Takefusa, Atsushi Igarashi, Taro Sekiyama, Kuniyasu Suzaki, Toshihiro Matsui, Atsuya Osaki, Naoki Yamashita, Nobuo Aoki, Sewon Park 0001, Terunobu Inaba, Lélio Brun, Yutaka Ishikawa, Kento Aida, Yasushi Ono, Kensuke Fukuda, Eisaku Sakane, Ichiro Hasuo |
COMPSAC | 13 |
| 2025 | A Lightweight Monitoring and Anomaly Detection Framework for IoT DevicesabstractThe always-online nature and the lack of sufficient built-in protection make Internet of Things (IoT) devices highly susceptible to various cyberthreats. An efficient and effective anomaly detection system is an essential need for IoT device security. However, it’s challenging to apply advanced host-based anomaly detection techniques from conventional systems to IoT devices due to the device resource constraints. This paper introduces a monitoring and anomaly detection framework based on thread-level system call streams for IoT devices. It leverages the execution pattern of IoT applications and detects anomalies by analyzing system call arguments and associated I/O attributes, in addition to the invocation sequence in real-time. The evaluation results highlight the feasibility of the proposed approach in terms of both performance and security. Yutaka Ishikawa, Atsuko Takefusa |
COMPSAC | 2 |
| 2024 | ZT-OTA Update Framework for IoT Devices Toward Zero Trust IoTabstractInternet of Things (IoT) is widely used as a fundamental technology for realizing various services. An IoT-based service system comprising cloud servers and many IoT devices connected via networks may be risky owing to the possibility of the entire system being be affected by cyberattacks on an IoT device. Moreover, new software vulnerabilities are frequently reported. From the perspective of security, IoT device software must be reliable and resilient. Consequently, a secure software update mechanism for IoT, assurance of software reliability, and mitigation mechanisms are required. Overall, this study proposed a zero-trust over-the-air (ZT-OTA) update framework for the reliable and resilient software update management of IoT devices via OTA software updates from remote locations. The ZT-OTA update framework provides a strict software version-management mechanism that enhances the security of software updates. Further, the proposed framework collaborates with a Software Assessment Service (SAS) as an authorized third-party assessment organization and deploys reliable software approved by the service to IoT devices. Upon discovering a software vulnerability following the deployment of the software, the SAS proactively revokes the reliability of the software by revoking its certificates for signing software codes and the vulnerability assessment result. Subsequently, the ZT-OTA server notifies developers and users of IoT devices to arrange for new software that must be fixed and automatically acts on behalf of IoT devices that must be retired. This study introduces the detailed design of the ZT-OTA update framework and describes to demonstrate its feasibility. Nobuo Aoki, Atsuko Takefusa, Yutaka Ishikawa, Yasushi Ono, Eisaku Sakane, Kento Aida |
COMPSAC | 3 |
| 2024 | Formal Support for Threat Modeling with Attack Decision DiagramsabstractSystem threat analysis requires a wide range of knowledge and is time-consuming. In this study, we propose FSTM system, a method to visualize threats in the system to support threat modeling using a formal verification tool. Concretely, given a security requirement and a system model, we represent what kind of attacks are enabled to prevent the system from satisfying the security requirement using an AND-OR tree based on exhaustive formal verification. To generate an exhaustive AND-OR tree, a large number of system models modified to enable various attack patterns and verification costs are required because verification results must be obtained for all threat patterns. We used the property that we call monotonicity of security to reduce the number of verifications and automatically generate a verification model for each threat pattern from a single verification model. We implemented FSTM system using Tamarin Prover, a formal verification tool, and evaluated it with case studies. Misato Nakabayashi, Taro Sekiyama, Ichiro Hasuo, Yutaka Ishikawa |
COMPSAC | 4 |
| 2024 | Rabbit: A Language to Model and Verify Data Flow in Networked SystemsabstractThreat modeling is an effective approach to fortifying security, where system designers model a target system and investigate potential security weaknesses during the design phase. To investigate security flaws in the system rigorously, the model has to be detailed to the extent that the flow of each asset is minutely tracked. For this purpose, we propose Rabbit, a language to model a networked system with the ability to describe data flows within the system. Rabbit also aims to formally verify security properties under the specified system model and attacker model. In this paper, we model a simple client-server system and demonstrate the possibility of automatic verification using Rabbit. In an experiment we conducted, we have found a non-trivial execution path that violates a desired security property, demonstrating the potential of automatic security verification using Rabbit. Terunobu Inaba, Yutaka Ishikawa, Atsushi Igarashi, Taro Sekiyama |
ISNCC | 2 |
| 2024 | Distributed Dataset Framework for Large Language Models Pre-training
Nao Souma, Yui Obara, Yasuhiko Yokote, Yutaka Ishikawa, Kimio Kuramitsu |
PKAW | 4 |
| 2024 | Page Type-Aware Full-Sequence Program Scheduling via Reinforcement Learning in High Density SSDsabstractFull-sequence program (FSP) can program multiple bits simultaneously, and thus complete a multiple-page write at one time for naturally enhancing write performance of high density 3-D solid-state drives (SSDs). This article proposes an FSP scheduling approach for the 3-D quad-level cell (QLC) SSDs, to further boost their read responsiveness. Considering each FSP operation in QLC SSDs spansfourdifferent types of QLC pages having dissimilar read latency, we introduce matching four pages of application data to the suited QLC pages and flush them together with the one-shot program of FSP. To this end, we employ reinforcement learning to classify the (cached) application data intofourcategories on the basis of their historical access frequency and the associating request size. Thus, the frequently read data can be mapped to the QLC pages having less access latency, meanwhile the other data can be flushed onto the slow QLC pages. Then, we can group four different categories of data pages and flush them together into a four-page unit of 3-D QLC SSDs with an FSP operation. In addition, a proactive rewrite method is also triggered for grouping the hot read data with the cached data to form an FSP unit. Through a series of emulation tests on several realistic disk traces, we show that the proposed mechanisms yields notable performance improvement on the read responsiveness. Jun Li 0062, Zhigang Cai, Balazs Gerofi, Yutaka Ishikawa, Jianwei Liao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A Linux Audit and MQTT-based Security Monitoring FrameworkabstractAlong with the significant growth in the number of connected Internet of Things (IoT) devices and increasingly aggressive cyberattacks, IoT cybersecurity has been facing more and more challenges. Security monitoring systems as one of the predominant security hardening approaches are often introduced to computer systems for detecting anomaly activities and ongoing intrusion. System auditing is one of the prevalent approaches for realizing such systems. However, most of the existing monitoring techniques for IoT systems heavily rely on network traffic analysis. In this work, we emphasize the device endpoint itself and propose a flexible and extensible monitoring framework for Linux-based IoT systems. We present the feasibility of the framework by implementing a monitoring prototype and an application simulating real-world IoT surveillance scenario, and conducting comprehensive evaluations on an ARM device. The evaluation results showcase the minimal overhead cost of the proposed monitoring framework and demonstrate the practicability of security monitoring on constrained IoT devices. Yutaka Ishikawa, Atsuko Takefusa |
COMPSAC | 2 |
| 2023 | Big Data Assimilation: Real-time 30-second-refresh Heavy Rain Forecast Using Fugaku During Tokyo Olympics and ParalympicsabstractReal-time 30-second-refresh numerical weather prediction (NWP) was performed with exclusive use of 11,580 nodes (~7%) of supercomputer Fugaku during Tokyo Olympics and Paralympics in 2021. Total 75,248 forecasts were disseminated in the 1-month period mostly stably with time-to-solution less than 3 minutes for 30-minute forecast. Japan's Big Data Assimilation (BDA) project developed the novel NWP system for precise prediction of hazardous rains toward solving the global climate crisis. Compared with typical 1-hour-refresh systems, the BDA system offered two orders of magnitude increase in problem size and revealed the effectiveness of 30-second refresh for highly nonlinear, rapidly evolving convective rains. To achieve the required time-to-solution for real-time 30-second refresh with high accuracy, the core BDA software incorporated single precision and enhanced parallel I/O with properly selected configurations of 1000 ensemble members and 500-m-mesh weather model. The massively parallel, I/O intensive real-time BDA computation demonstrated a promising future direction. Takemasa Miyoshi, Arata Amemiya, Shigenori Otsuka, Yasumitsu Maejima, Takumi Honda, Hirofumi Tomita, Seiya Nishizawa, Kenta Sueki, Tsuyoshi Yamaura, Yutaka Ishikawa, Shinsuke Satoh, Tomoo Ushio, Kana Koike, Atsuya Uno |
SC | 11 |
| 2022 | Pattern-Based Prefetching with Adaptive Cache Management Inside of Solid-State DrivesabstractThis article proposes a pattern-based prefetching scheme with the support of adaptive cache management, at the flash translation layer of solid-state drives ( SSDs ). It works inside of SSDs and has features of OS dependence and uses transparency. Specifically, it first mines frequent block access patterns that reflect the correlation among the occurred I/O requests. Then, it compares the requests in the current time window with the identified patterns to direct prefetching data into the cache of SSDs. More importantly, to maximize the cache use efficiency, we build a mathematical model to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the written data. Experimental results show that our proposal can yield improvements on average read latency by 1.8 %– 36.5 % without noticeably increasing the write latency, in contrast to conventional SSD-inside prefetching schemes. Jun Li 0062, Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Kenli Li 0001, Balazs Gerofi, Yutaka Ishikawa |
ACM Trans. Storage | 7 |
| 2021 | A Scalability Study of Data Exchange in HPC Multi-component WorkflowsabstractMulti-component workflows play a significant role in High-Performance Computing and Big Data applications. They usually contain multiple, independently developed components that execute side-by-side to perform sophisticated computation and data exchange through file I/O over parallel file system. However, file I/O can become an impediment in such systems and cause undesirable performance degradation due to its relatively low speed (compared to the interconnect fabric), which is unacceptable especially for applications with strict time constraints. The Data Transfer Framework (DTF) is an I/O arbitration layer working with the PnetCDF I/O library aiming at eliminating the bottleneck by transparently redirecting file I/O operations through the parallel file system to message passing via the high-speed interconnect between coupled components. Scalable and high-speed data transfer between components can be thus easily achieved with minimal development effort by using DTF. However, previous work provides insufficient scalability evaluation of the framework. In order to comprehensively evaluate the scalability of an I/O middleware like DTF and highlight its major advantages, we develop an I/O benchmark for multicomponent workflows. Using the benchmark we conduct large-scale scalability evaluation using up to 32,768 compute nodes on supercomputer Fugaku and 2,048 compute nodes on Oakforest-PACS by comparing direct data transfer to file I/O performed on Lustre file system and Fugaku’s Lightweight Layered IO-Accelerator (LLIO). We provide insights into DTF’s scalability and performance enhancements with the intention to impact future I/O middleware and inter-component data exchange design in multi-component workflows. Atsushi Hori, Balazs Gerofi, Yutaka Ishikawa |
CLUSTER | 4 |
| 2021 | Linux vs. lightweight multi-kernels for high performance computing: experiences at pre-exascaleabstractThe long standing consensus in the High-Performance Computing (HPC) Operating Systems (OS) community is that lightweight kernel (LWK) based OSes have the potential to outperform Linux at extreme scale. To explore if LWKs live up to their expectation we developed IHK/McKernel, a lightweight multi-kernel OS designed for HPC, and deployed it on two high-end supercomputers to compare its performance against Linux. Oakforest-PACS, an Intel Xeon Phi (x86) based supercomputer, runs a moderately tuned Linux distribution. Fugaku, the world's fastest supercomputer at the time of writing this paper, is based on Fujitsu's A64FX (aarch64) CPU that runs a highly tuned Linux environment. Balazs Gerofi, Kohei Tarumizu, Takayuki Okamoto, Masamichi Takagi, Shinji Sumimoto, Yutaka Ishikawa |
SC | 7 |
| 2021 | An international survey on MPI users
Atsushi Hori, Emmanuel Jeannot, George Bosilca, Takahiro Ogura, Balazs Gerofi, Yutaka Ishikawa |
Parallel Comput. | 7 |
| 2021 | Mitigating Negative Impacts of Read Disturb in SSDsabstractRead disturb is a circuit-level noise in solid-state drives (SSDs), which may corrupt existing data in SSD blocks and then cause high read error rate and longer read latency. The approach of read refresh is commonly used to avoid read disturb errors by periodically migrating the hot read data to other free blocks, but it places considerable negative impacts on I/O (Input/Output) responsiveness. This article proposes scheduling approaches on write data and read refresh operations, to mitigate the negative effects caused by read disturb. To be specific, we first construct a model to classify SSD blocks into two categories according to the estimated read error rate by referring to the factors of block’s P/E (Program/Erase) cycle and the accumulated read count to the block. Then, the data being intensively read will be redirected to the block having a small read error rate, as it is not sensitive to read disturb even though the data will be heavily requested. Moreover, we take advantage of reinforcement learning to predict the idle interval between two I/O requests for purposely conducting (partial) read refresh operations. As a result, it is able to minimize negative impacts toward subsequent incoming I/O requests and to ensure I/O responsiveness. Through a series of emulation tests on several realistic disk traces, we demonstrate that the proposed mechanisms can noticeably yield performance improvements on the metrics of read error rate and I/O latency. Jun Li 0062, Zhibing Sha, Zhigang Cai, Jianwei Liao 0001, Balazs Gerofi, Yutaka Ishikawa |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2020 | Frequent Access Pattern-based Prefetching Inside of Solid-State DrivesabstractThis paper proposes an SSD-inside data prefetching scheme, which has features of OS-dependence and use transparency. To be specific, it first mines frequent block access patterns that reflect the correlation among the occurred requests. Then it compares the requests in the current time window with the identified patterns, to direct fetching data in advance. Furthermore, to maximize the cache use efficiency, we construct a method to adaptively determine the cache partition on the basis of I/O workload characteristics, for separately buffering the prefetched data and the write data. Experimental results demonstrate that our proposal can yield improvements on average read latency by 6.3% to 9.3% without noticeably increasing write latency, in contrast to conventional SSD-inside prefetching schemes. Xiaofei Xu 0002, Zhigang Cai, Jianwei Liao 0001, Yutaka Ishikawa |
DATE | 4 |
| 2020 | Co-design for A64FX manycore processor and "Fugaku"abstractWe have been carrying out the FLAGSHIP 2020 Project to develop the Japanese next-generation flagship supercomputer, the Post-K, recently named “Fugaku”. We have designed an original many core processor based on Armv8 instruction sets with the Scalable Vector Extension (SVE), an A64FX processor, as well as a system including interconnect and a storage subsystem with the industry partner, Fujitsu. The “co-design” of the system and applications is a key to making it power efficient and high performance. We determined many architectural parameters by reflecting an analysis of a set of target applications provided by applications teams. In this paper, we present the pragmatic practice of our co-design effort for “Fugaku”. As a result, the system has been proven to be a very power-efficient system, and it is confirmed that the performance of some target applications using the whole system is more than 100 times the performance of the K computer. Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama, Tetsuya Odajima, Miwako Tsuji, Hisashi Yashiro, Masaki Aoki, Naoyuki Shida, Ikuo Miyoshi, Kouichi Hirai, Atsushi Furuya, Akira Asato, Kuniki Morita, Toshiyuki Shimizu |
SC | 2 |
| 2018 | Improving Collective MPI-IO Using Topology-Aware Stepwise Data Aggregation with I/O ThrottlingabstractMPI-IO has been used in an internal I/O interface layer of HDF5 or PnetCDF, where collective MPI-IO plays a big role in parallel I/O to manage a huge scale of scientific data. However, existing collective MPI-IO optimization named two-phase I/O has not been tuned enough for recent supercomputers consisting of mesh/torus interconnects and a huge scale of parallel file systems due to lack of topology-awareness in data transfers and optimization for parallel file systems. In this paper, we propose I/O throttling and topology-aware stepwise data aggregation in two-phase I/O of ROMIO, which is a representative MPI-IO library, in order to improve collective MPI-IO performance even if we have multiple processes per compute node. Throttling I/O requests going to a target file system mitigates I/O request contention, and consequently I/O performance improvements are achieved in file access phase of two-phase I/O. Topology-aware aggregator layout with paying attention to multiple aggregators per compute node alleviates contention in data aggregation phase of two-phase I/O. In addition, stepwise data aggregation improves data aggregation performance. HPIO benchmark results on the K computer indicate that the proposed optimization has achieved up to about 73% and 39% improvements in write performance compared with the original implementation using 12,288 and 24,576 processes on 3,072 and 6,144 compute nodes, respectively. Yuichi Tsujita, Atsushi Hori, Toyohisa Kameyama, Atsuya Uno, Fumiyoshi Shoji, Yutaka Ishikawa |
HPC Asia | 6 |
| 2018 | PicoDriver: fast-path device drivers for multi-kernel operating systemsabstractLightweight kernel (LWK) operating systems (OS) in high-end supercomputing have a proven track record of excellent scalability. However, the lack of full Linux compatibility and limited availability of device drivers in LWKs have prohibited their wide-spread deployment. Multi-kernels, where an LWK is run side-by-side with Linux on many-core CPUs, have been proposed to address these shortcomings. In a multi-kernel system the LWK implements only performance critical kernel services and the rest of the OS functionality is offloaded to Linux. Access to device drivers is usually attained via offloading. Although high-performance interconnects are commonly driven from user-space, there are networks (e.g., Intel's OmniPath or Cray's Gemini) that require device driver interaction for a number of performance sensitive operations, which in turn can be adversely impacted by system call offloading. Balazs Gerofi, Aram Santogidis, Dominique Martinet, Yutaka Ishikawa |
HPDC | 4 |
| 2018 | Process-in-process: techniques for practical address-space sharingabstractThe two most common parallel execution models for many-core CPUs today are multiprocess (e.g., MPI) and multithread (e.g., OpenMP). The multiprocess model allows each process to own a private address space, although processes can explicitly allocate shared-memory regions. The multithreaded model shares all address space by default, although threads can explicitly move data to thread-private storage. In this paper, we present a third model called process-in-process (PiP), where multiple processes are mapped into a single virtual address space. Thus, each process still owns its process-private storage (like the multiprocess model) but can directly access the private storage of other processes in the same virtual address space (like the multithread model). Atsushi Hori, Min Si, Balazs Gerofi, Masamichi Takagi, Jai Dayal, Pavan Balaji, Yutaka Ishikawa |
HPDC | 7 |
| 2018 | Performance and Scalability of Lightweight Multi-kernel Based Operating SystemsabstractMulti-kernels leverage today's multi-core chips to run multiple operating system (OS) kernels, typically a Light Weight Kernel (LWK) and a Linux kernel, simultaneously. The LWK provides high performance and scalability, while the Linux kernel provides compatibility. Multi-kernels show the promise of being able to meet tomorrow's extreme-scale computing needs while providing strong isolation, yielding high performance and scalability needed by classical HPC applications. McKernel and mOS started as independent research initiatives to explore the above potential. Previous work described their design and architecture advantages. This paper deploys the two LWKs and presents results from running them on a 2,048-node system with Intel Xeon Phi processors (KNL) connected by Intel Omni-Path Fabric. We compare the performance of McKernel, mOS, and Linux. Although the two multi-kernel efforts approached the problem from different angles, the results show a median performance improvement of 9% with some applications as high as 280% validating the efficacy of the multi-kernel approach. We provide insight into the performance improvements and discuss the strengths of the two different multi-kernel approaches. Balazs Gerofi, Rolf Riesen, Masamichi Takagi, Taisuke Boku, Kengo Nakajima, Yutaka Ishikawa, Robert W. Wisniewski |
IPDPS | 6 |
| 2018 | Deep Learning on Large-Scale Muticore ClustersabstractConvolutional neural networks (CNNs) have achieved outstanding accuracy among conventional machine learning algorithms. Recent works have shown that large and complicated models, which take significant cost for training are needed to get higher accuracy. To train these models efficiently in high performance computers (HPCs), many parallelization techniques for CNNs have been developed. However, most techniques are mainly targeting GPUs and parallelizations for CPUs are not fully investigated. This paper explores CNN training performance on large-scale multicore clusters by optimizing intra-node processing and applying techniques of inter-node parallelization for multiple GPUs. Detailed experiments conducted on state-of-the-art multi-core processors using the openMP API and MPI framework demonstrated that Caffe-based CNNs can be accelerated by using well-designed multithreaded programs. We achieved at most 1.64 times speedup in convolution operations with devised lowering strategy compared to conventional lowering and acquired 772 times speedup with 864 nodes compared to one node. Kazumasa Sakivama, Shinpei Kato, Yutaka Ishikawa, Atsushi Hori, Abraham Monrroy Cano |
SBAC-PAD | 3 |
| 2018 | Dynamic Adaptable Asynchronous Progress Model for MPI RMA Multiphase ApplicationsabstractCasper is a process-based asynchronous progress model for MPI one-sided communication on multi- and many-core architectures. The one-sided communication is not truly one-sided in most MPI implementations: the target process still relies on software progress to complete incoming operations. Casper allows the user to specify an arbitrary number of cores dedicated to background ghost processes and transparently redirects the RMA operations to ghost processes by utilizing the PMPI redirection and MPI-3 shared-memory technologies. Although Casper benefits applications that suffer from lack of asynchronous progress, the operation redirection design might not support complex multiphase applications effectively, which often involve dynamically changing communication density and computing workloads. In this paper, we present an adaptive mechanism in Casper to address the limitation of static asynchronous progress in multiphase applications. We exploit two adaptive strategies, a user-guided strategy and a fully transparent and automatic strategy based on self-profiling and prediction, to dynamically reconfigure the asynchronous progress in Casper according to real-time performance characteristics during multiphase execution. We evaluate the adaptive approaches in both microbenchmarks and a real quantum chemistry application suite, NWChem, on the Cray XC30 supercomputer and an Intel Omni-Path cluster. Min Si, Antonio J. Peña, Jeff R. Hammond, Pavan Balaji, Masamichi Takagi, Yutaka Ishikawa |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2017 | A flexible I/O arbitration framework for netCDF-based big data processing workflows on high-end supercomputersabstractSummary On the verge of the convergence between high‐performance computing and Big Data processing, it has become increasingly prevalent to deploy large‐scale data analytics workloads on high‐end supercomputers. Such applications often come in the form of complex workflows with various different components, assimilating data from scientific simulations as well as from measurements streamed from sensor networks, such as radars and satellites. For example, as part of the Flagship 2020 (post‐K) supercomputer project of Japan, RIKEN is investigating the feasibility of a highly accurate weather forecasting system that would provide a real‐time outlook for severe guerrilla rainstorms. One of the main performance bottlenecks of this application is the lack of efficient communication among workflow components, which currently takes place over the parallel file system.In this paper, we present an initial study of a direct communication framework designed for complex workflows that eliminates unnecessary file I/O among components. Specifically, we propose an I/O arbitration layer that provides direct parallel data transfer (both synchronous and asynchronous) among job components that rely on the netCDF interface for performing I/O operations. Our solution requires only minimal modifications to application code. Moreover, we propose a configuration file–based approach that allows users to specify the desired data transfer pattern among workflow components, offering a general solution for different application contexts. We present a preliminary evaluation of the proposed framework on the K Computer (running on up to 4800 compute nodes) using RIKEN's experimental weather forecasting workflow as a case study. Jianwei Liao 0001, Balazs Gerofi, Guo-Yuan Lien, Takemasa Miyoshi, Seiya Nishizawa, Hirofumi Tomita, Wei-keng Liao, Alok N. Choudhary, Yutaka Ishikawa |
Concurr. Comput. Pract. Exp. | 9 |
| 2017 | Performing Initiative Data Prefetching in Distributed File Systems for Cloud ComputingabstractThis paper presents an initiative data prefetching scheme on the storage servers in distributed file systems for cloud computing. In this prefetching technique, the client machines are not substantially involved in the process of data prefetching, but the storage servers can directly prefetch the data after analyzing the history of disk I/O access events, and then send the prefetched data to the relevant client machines proactively. To put this technique to work, the information about client nodes is piggybacked onto the real client I/O requests, and then forwarded to the relevant storage server. Next, two prediction algorithms have been proposed to forecast future block access operations for directing what data should be fetched on storage servers in advance. Finally, the prefetched data can be pushed to the relevant client machine from the storage server. Through a series of evaluation experiments with a collection of application benchmarks, we have demonstrated that our presented initiative prefetching technique can benefit distributed file systems for cloud environments to achieve better I/O performance. In particular, configuration-limited client machines in the cloud are not responsible for predicting I/O access operations, which can definitely contribute to preferable system performance on them. Jianwei Liao 0001, François Trahay, Guoqiang Xiao 0001, Li Li 0006, Yutaka Ishikawa |
IEEE Trans. Cloud Comput. | 5 |
| 2016 | Toward a General I/O Arbitration Framework for netCDF Based Big Data Processing
Jianwei Liao 0001, Balazs Gerofi, Guo-Yuan Lien, Seiya Nishizawa, Takemasa Miyoshi, Hirofumi Tomita, Yutaka Ishikawa |
Euro-Par | 7 |
| 2016 | On the Scalability, Performance Isolation and Device Driver Transparency of the IHK/McKernel Hybrid Lightweight KernelabstractExtreme degree of parallelism in high-end computing requires low operating system noise so that large scale, bulk-synchronous parallel applications can be run efficiently. Noiseless execution has been historically achieved by deploying lightweight kernels (LWK), which, on the other hand, can provide only a restricted set of the POSIX API in exchange for scalability. However, the increasing prevalence of more complex application constructs, such as in-situ analysis and workflow composition, dictates the need for the rich programming APIs of POSIX/Linux. In order to comply with these seemingly contradictory requirements, hybrid kernels, where Linux and a lightweight kernel (LWK) are run side-by-side on compute nodes, have been recently recognized as a promising approach. Although multiple research projects are now pursuing this direction, the questions of how node resources are shared between the two types of kernels, how exactly the two kernels interact with each other and to what extent they are integrated, remain subjects of ongoing debate. In this paper, we describe IHK/McKernel, a hybrid software stack that seamlessly blends an LWK with Linux by selectively offloading system services from the lightweight kernel to Linux. Specifically, we are focusing on transparent reuse of Linux device drivers and detail the design of our framework that enables the LWK to naturally leverage the Linux driver codebase without sacrificing scalability or the POSIX API. Through rigorous evaluation on a medium size cluster we demonstrate how McKernel provides consistent, isolated performance for simulations even in face of competing, in-situ workloads. Balazs Gerofi, Masamichi Takagi, Atsushi Hori, Gou Nakamura, Tomoki Shirasawa, Yutaka Ishikawa |
IPDPS | 6 |
| 2016 | Revisiting RDMA Buffer Registration in the Context of Lightweight Multi-kernelsabstractLightweight multi-kernel architectures, where HPC specialized lightweight kernels (LWKs) run side-by-side with Linux on compute nodes, have received a great deal of attention recently due to their potential for addressing many of the challenges system software faces as we move towards exascale and beyond. LWKs in multi-kernels implement only a limited set of kernel functionality and the rest is supported by Linux, for example, device drivers for high-performance interconnects. While most of the operations of modern high-performance interconnects are driven entirely by user-space, memory registration for remote direct memory access (RDMA) usually involves interaction with the Linux device driver and thus comes at the price of service offloading. Balazs Gerofi, Masamichi Takagi, Yutaka Ishikawa |
EuroMPI | 3 |
| 2016 | "Big Data Assimilation" Toward Post-Petascale Severe Weather Prediction: An Overview and ProgressabstractFollowing the invention of the telegraph, electronic computer, and remote sensing, “big data” is bringing another revolution to weather prediction. As sensor and computer technologies advance, orders of magnitude bigger data are produced by new sensors and high-precision computer simulation or “big simulation.” Data assimilation (DA) is a key to numerical weather prediction (NWP) by integrating the real-world sensor data into simulation. However, the current DA and NWP systems are not designed to handle the “big data” from next-generation sensors and big simulation. Therefore, we propose “big data assimilation” (BDA) innovation to fully utilize the big data. Since October 2013, the Japan's BDA project has been exploring revolutionary NWP at 100-m mesh refreshed every 30 s, orders of magnitude finer and faster than the current typical NWP systems, by taking advantage of the fortunate combination of next-generation technologies: the 10-petaflops K computer, phased array weather radar, and geostationary satellite Himawari-8. So far, a BDA prototype system was developed and tested with real-world retrospective local rainstorm cases. This paper summarizes the activities and progress of the BDA project, and concludes with perspectives toward the post-petascale supercomputing era. Takemasa Miyoshi, Guo-Yuan Lien, Shinsuke Satoh, Tomoo Ushio, Kotaro Bessho, Hirofumi Tomita, Seiya Nishizawa, Ryuji Yoshida, Sachiho A. Adachi, Jianwei Liao 0001, Balazs Gerofi, Yutaka Ishikawa, Masaru Kunii, Yasumitsu Maejima, Shigenori Otsuka, Michiko Otsuka, Kozo Okamoto, Hiromu Seko |
Proc. IEEE | 12 |
| 2016 | Prefetching on Storage Servers through Mining Access Patterns on BlocksabstractDistributed file systems have been widely deployed as back-end storage systems to offer I/O services for parallel/distributed applications that process large amounts of data. Data prefetching in distributed file systems is a well-known optimization technique which can mask both network and disk latency and consequently boost I/O performance. Traditionally, data prefetching is initiated by the client file systems, however, conventional prefetching schemes are not well suited for client machines that have limited memory and computing capacity. To offer an efficient prefetching approach for resource-limited client machines, this paper proposes a novel server-side prefetching mechanism. Specifically, we propose to piggyback client identification to I/O requests so that server side block access history can be put into context. On the server side, we utilize the horizontal visibility graph technique to transform per-client time series of block access sequences into a connected graph for which we employ Tarjan's algorithm to disclose cut points in the connected graph. We express these patterns with feature tuples and we propose the X-step pattern matching algorithm to find a matching access pattern (i.e., a feature tuple) for a given block access history. Experimental results indicate that our newly proposed prefetching mechanism can ease client machines and their applications from the process of data prefetching, boosting client performance accordingly, and that it yields an attractive increase in data throughput as well. Jianwei Liao 0001, François Trahay, Balazs Gerofi, Yutaka Ishikawa |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Techniques for Enabling Highly Efficient Message Passing on Many-Core ArchitecturesabstractMany-core architecture provides a massively parallel environment with dozens of cores and hundreds of hardware threads. Scientific application programmers are increasingly looking at ways to utilize such large numbers of lightweight cores for various programming models. Efficiently executing these models on massively parallel many-core environments is not easy, however and performance may be degraded in various ways. The first author's doctoral research focuses on exploiting the capabilities of many-core architectures on widely used MPI implementations. While application programmers have studied several approaches to achieve better parallelism and resource sharing, many of those approaches still face communication problems that degrade performance. In the thesis, we investigate the characteristics of MPI on such massively threaded architectures and propose two efficient strategies -- a multi-threaded MPI approach and a process-based asynchronous model -- to optimize MPI communication for modern scientific applications. Min Si, Pavan Balaji, Yutaka Ishikawa |
CCGRID | 3 |
| 2015 | Scaling NWChem with Efficient and Portable Asynchronous Communication in MPI RMAabstractNWChem is one of the most widely used computational chemistry application suites for chemical and biological systems. Despite its vast success, the computational efficiency of NWChem is still low. This is especially true in higher accuracy methods such as the CCSD(T) coupled cluster method, where it currently achieves a mere 50% computational efficiency when run at large scales. In this paper, we demonstrate the most computationally efficient scaling of NWChem CCSD(T) to date, and use it to solve large water clusters. We use our recently proposed process-based asynchronous progress framework for MPI RMA, called Casper, to scale the computation on water clusters at near-100% computational efficiency on up to 12288 cores. Min Si, Antonio J. Peña, Jeff R. Hammond, Pavan Balaji, Yutaka Ishikawa |
CCGRID | 5 |
| 2015 | Casper: An Asynchronous Progress Model for MPI RMA on Many-Core ArchitecturesabstractIn this paper we present "Casper," a process-based asynchronous progress solution for MPI one-sided communication on multi- and many-core architectures. Casper uses transparent MPI call redirection through PMPI and MPI-3 shared-memory windows to map memory from multiple user processes into the address space of one or more ghost processes, thus allowing for asynchronous progress where needed while allowing native hardware-based communication where available. Unlike traditional thread- and interrupt-based asynchronous progress models, Casper provides the capability to dedicate an arbitrary number of ghost processes for asynchronous progress, thus balancing application requirements with the capabilities of the underlying MPI implementation. We present a detailed design of the proposed architecture including several techniques for maintaining correctness per the MPI-3 standard as well as performance optimizations where possible. We also compare Casper with traditional thread- and interrupt-based asynchronous progress models and demonstrate its performance improvements with a variety of micro benchmarks and a production chemistry application. Min Si, Antonio J. Peña, Jeff R. Hammond, Pavan Balaji, Masamichi Takagi, Yutaka Ishikawa |
IPDPS | 6 |
| 2015 | Toward Operating System Support for Scalable Multithreaded Message PassingabstractModern CPU architectures provide a large number of processing cores and application programmers are increasingly looking at hybrid programming models, where multiple threads of a single process interact with the MPI library simultaneously. Moreover, recent high-speed interconnection networks are being designed with capabilities targeting communication explicitly from multiple processor cores. As a result, scalability of the MPI library so that multithreaded applications can efficiently drive independent network communication has become a major concern. Balazs Gerofi, Masamichi Takagi, Yutaka Ishikawa |
EuroMPI | 3 |
| 2015 | Sliding Substitution of Failed NodesabstractThis paper considers the questions of how spare nodes should be allocated, how to substitute them for faulty nodes, and how much the communication performance is affected by such a substitution. The third question stems from the modification of the rank mapping by node substitutions, which can incur additional message collisions. In a stencil computation, rank mapping is done in a straightforward way on a Cartesian network without incurring any message collisions. However, once a substitution has occurred, the node- rank mapping may be destroyed. Therefore, these questions must be answered in a way that minimizes the degradation of communication performance. Atsushi Hori, Kazumi Yoshinaga, Thomas Hérault, Aurelien Bouteiller, George Bosilca, Yutaka Ishikawa |
EuroMPI | 6 |
| 2014 | Grid-Oriented Process Clustering System for Partial Message LoggingabstractIn a computer cluster composed of many nodes, the mean time between failures becomes shorter as the number of nodes increases. This may mean that lengthy tasks cannot be performed, because they will be interrupted by failure. Therefore, fault tolerance has become an essential part of high-performance computing. Partial message logging forms clusters of processes, and coordinates a series of checkpoints to log messages between groups. Our study proposes a system of two features to improve the efficiency of partial message logging: 1) the communication log used in the clustering is recorded at runtime, and 2) a graph partitioning algorithm reduces the complexity of the system by geometrically partitioning a grid graph. The proposed system is evaluated by executing a scientific application. The results of process clustering are compared to existing methods in terms of the clustering performance and quality. Hideyuki Jitsumoto, Yuki Todoroki, Yutaka Ishikawa, Mitsuhisa Sato |
DSN | 3 |
| 2014 | Interface for heterogeneous kernels: A framework to enable hybrid OS designs targeting high performance computing on manycore architecturesabstractTurning towards exascale systems and beyond, it has been widely argued that the currently available systems software is not going to be feasible due to various requirements such as the ability to deal with heterogeneous architectures, the need for systems level optimization targeting specific applications, elimination of OS noise, and at the same time, compatibility with legacy applications. To cope with these issues, a hybrid design of operating systems where light-weight specialized kernels can cooperate with a traditional OS kernel seems adequate, and a number of recent research projects are now heading into this direction. This paper presents Interface for Heterogeneous Kernels (IHK), a general framework enabling hybrid kernel designs in systems equipped with manycore processors and/or accelerators. IHK provides a range of capabilities, such as resource partitioning, management of heterogeneous OS kernels, as well as a low-level communication layer among the kernels. We describe IHK's interface and demonstrate its feasibility for hybrid kernel designs through executing various different lightweight OS kernels on top of it, which are specialized for certain types of applications. We use the Intel Xeon Phi, Intel's latest manycore coprocessor, as our experimental platform. Taku Shimosawa, Balazs Gerofi, Masamichi Takagi, Gou Nakamura, Tomoki Shirasawa, Yuji Saeki, Masaaki Shimizu, Atsushi Hori, Yutaka Ishikawa |
HiPC | 9 |
| 2014 | CMCP: a novel page replacement policy for system level hierarchical memory management on many-coresabstractThe increasing prevalence of co-processors such as the Intel Xeon Phi, has been reshaping the high performance computing (HPC) landscape. The Xeon Phi comes with a large number of power efficient CPU cores, but at the same time, it's a highly memory constraint environment leaving the task of memory management entirely up to application developers. To reduce programming complexity, we are focusing on application transparent, operating system (OS) level hierarchical memory management. Balazs Gerofi, Akio Shimada, Atsushi Hori, Masamichi Takagi, Yutaka Ishikawa |
HPDC | 5 |
| 2014 | MT-MPI: multithreaded MPI for many-core environmentsabstractMany-core architectures, such as the Intel Xeon Phi, provide dozens of cores and hundreds of hardware threads. To utilize such architectures, application programmers are increasingly looking at hybrid programming models, where multiple threads interact with the MPI library (frequently called "MPI+X" models). A common mode of operation for such applications uses multiple threads to parallelize the computation, while one of the threads also issues MPI operations (i.e., MPI FUNNELED or SERIALIZED thread-safety mode). In MPI+OpenMP applications, this is achieved, for example, by placing MPI calls in OpenMP critical sections or outside the OpenMP parallel regions. However, such a model often means that the OpenMP threads are active only during the parallel computation phase and idle during the MPI calls, resulting in wasted computational resources. In this paper, we present MT-MPI, an internally multithreaded MPI implementation that transparently coordinates with the threading runtime system to share idle threads with the application. It is designed in the context of OpenMP and requires modifications to both the MPI implementation and the OpenMP runtime in order to share appropriate information between them. We demonstrate the benefit of such internal parallelism for various aspects of MPI processing, including derived datatype communication, shared-memory communication, and network I/O operations. Min Si, Antonio J. Peña, Pavan Balaji, Masamichi Takagi, Yutaka Ishikawa |
ICS | 5 |
| 2014 | Multithreaded Two-Phase I/O: Improving Collective MPI-IO Performance on a Lustre File SystemabstractROMIO, a representative MPI-IO implementation, has been widely used in recent large-scale parallel computations. The two-phase I/O optimization scheme of ROMIO improves I/O performance for non-contiguous access patterns, however, this scheme still has room to improve performance to make it suitable for recent data-intensive computing. We propose overlapping data exchange operations with file I/O operations by using a multithreaded scheme to achieve further I/O throughput improvement. We show up to 60% improvement by the multithreaded two-phase I/O relative to the original two-phase I/O in performance evaluation of collective write operations on a Lustre file system of a Linux PC cluster. Yuichi Tsujita, Kazumi Yoshinaga, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
PDP | 6 |
| 2013 | Partially Separated Page Tables for Efficient Operating System Assisted Hierarchical Memory Management on Heterogeneous ArchitecturesabstractHeterogeneous architectures, where a multicore processor is accompanied with a large number of simpler, but more power-efficient CPU cores optimized for parallel workloads, are receiving a lot of attention recently. At present, these co-processors, such as the Intel Xeon Phi product family, come with limited on-board memory, which requires partitioning computational problems manually into pieces that can fit into the device's RAM, as well as efficiently overlapping computation and communication. In this paper we propose an application transparent, operating system (OS)assisted hierarchical memory management system, where the OS orchestrates data movement between the host and the device and updates the process virtual memory address space accordingly. We identify the main scalability issues of frequent address space changes, such as the increasing price of TLB invalidations with the growing number of CPU cores, and propose partially separated page tables with address-range CPU masks to overcome the problem. With partially separated page tables each core maintains its own set of mappings of the computation area, enabling the OS to perform address space updates in a scalable manner, and involve a particular CPU core in TLB invalidation only if it is absolutely necessary. Furthermore, we propose dedicated data movement cores in order to efficiently overlap computation and communication. We provide experimental results on stencil computation, a common HPCkernel, and show that OS assisted memory management has the potential for scalable transparent data movement. Balazs Gerofi, Akio Shimada, Atsushi Hori, Yutaka Ishikawa |
CCGRID | 4 |
| 2013 | A Delegation Mechanism on Many-Core Oriented Hybrid Parallel Computers for Scalability of Communicators and Communications in MPIabstractThis paper describes a delegation based high throughput MPIcommunication mechanism under tough memory utilization constrains on a many-core oriented hybrid parallel computer. Towards the Exascale era, hybrid parallel computers consisting of many-core and multi-core architectures both on the same node are focused. Although many-core architectures such as GPU or Intel MIC has high potential in computing power by the large number of computing cores, per-core computing power is lower than that of multi-core CPUs. Furthermore, available memory resources for the many-core CPUs are quite smaller than those for multi-core CPUs. Thus we may have a sort of penalty in memory utilization in MPI communications when we utilize a normal MPI library. Here we deploy a delegatee process on each node to merge MPI communications and minimize memory utilization for an MPI communicator. Another advantage of the delegatee process scheme is minimization of memory utilization on many-core CPUs by delegating MPI requests to associated delegatee process on multi-core CPUs. In this paper, we show performance advantages and effective resource utilization by our proposed scheme compared with the original MPI implementation. Kazumi Yoshinaga, Yuichi Tsujita, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
PDP | 6 |
| 2013 | Optimization of MPI persistent communicationabstractThis paper proposes a novel optimization technique for MPI persistent communication, that utilizes multiple RDMA engines to carry out low latency communication. Because the interconnects used in modern supercomputers have multiple RDMA engines, the multiple communication requests specified by a persistent communication invocation can be scheduled onto the RDMA engines in an optimal way and thus result in better communication performance. Such a scheduling algorithm is not only a packing problem, but also avoids interconnect resource contentions as much as possible. The scheduling algorithm proposed in this paper balances load of RDMA engines and mitigates network link contentions in case of neighbor communication patterns such like in stencil computation. The proposed scheduling algorithm is implemented in Open MPI of K computer using RDMA functions provided as Fujitsu MPI extensions. A typical 2D stencil computation of a climate simulation code is used as a benchmark program. The experimental result shows that a factor of two speedup of communication time is achieved. Masayuki Hatanaka, Atsushi Hori, Yutaka Ishikawa |
EuroMPI | 3 |
| 2013 | Revisiting rendezvous protocols in the context of RDMA-capable host channel adapters and many-core processorsabstractWe revisit RDMA-based rendezvous protocols in MPI in the context of cluster computer with RDMA-capable HCA and many-core processors, and propose two improved protocols. The conventional sender-initiate rendezvous protocols cause costly processor-device communications via PCI bus on detecting completion of RDMA transfer. The conventional receiver-initiate rendezvous protocols need to send extra control messages when a value of the memory-slot to poll in the receive buffer has the same value as the send buffer. The first proposed protocol implements polling on a memory-slot in the receive buffer to eliminate the processor-device communications. The second proposed protocol randomizes the value of the memory-slot to poll to reduce extra control messages. We have evaluated the proposed protocols using micro-benchmarks and NAS Parallel Benchmarks. One of the proposed protocols has a benefit compared to the conventional protocols. And the second proposed protocol reduces the execution time by up to 11.14% compared to the first protocol. Masamichi Takagi, Yuichi Nakamura 0002, Atsushi Hori, Balazs Gerofi, Yutaka Ishikawa |
EuroMPI | 5 |
| 2013 | Utilizing memory content similarity for improving the performance of highly available virtual machines
Balazs Gerofi, Zoltan Vass, Yutaka Ishikawa |
Future Gener. Comput. Syst. | 3 |
| 2012 | Design and Implementation of Portable and Efficient Non-blocking Collective CommunicationabstractNon-blocking communications are widely used in parallel applications for hiding communication overheads through overlapped computation and communication. While most of the existing implementations provide a non-blocking version of point-to-point communications, there is no portable and efficient implementation of non-blocking collectives, partly because application execution contexts need to be interrupted by dependent communications. This paper presents a portable and efficient user-level implementation technique of non-blocking communications. It allows users to design non-blocking collectives by declaring their operations and dependencies using provided APIs without being concerned with complicated management of their progression. While user-level implementations can be less efficient than kernel-level ones due to the cost of OS context switches, we solve this problem by employing the Marcel user level light-weight thread library when invoking communication operations. More specifically, each communication operation is mapped to one Marcel thread and scheduled to be executed when each operation's dependencies are satisfied by certain events. All executable operations and main user thread are executed simultaneously without any explicit invocations. Performance evaluations with micro benchmarks demonstrate the effectiveness of our proposed technique. Compared to existing OS-thread based method, it reduces CPU load to less than 10% while achieving similar level of communication latencies. We also discuss and compare the descriptive power of internal expressions for non-blocking communications. Akihiro Nomura 0002, Yutaka Ishikawa, Naoya Maruyama, Satoshi Matsuoka |
CCGRID | 2 |
| 2012 | clone_n(): Parallel Thread Creation for Upcoming Many-Core ArchitecturesabstractHeterogeneous architectures, where a multicore processor, which is optimized for fast single-thread performance, is accompanied with a large number of simpler, but more power-efficient cores optimized for parallel workloads, such as NVIDIA's GPUs or Intel's Many Integrated Core (MIC), have been receiving a lot attention recently. Although NVIDIA's GPUs include built-in support for parallelism control, the MIC uses classical software thread creation and scheduling done by the operating system (OS). While efficient thread creation is desired in such many-core environments, current OS APIs provide the facility of creating only one thread at a time. In this paper, we propose a new system call for parallel thread creation on many-core coprocessors and show that it can perform up to 6.9 times better than the sequential version when executed on Intel's MIC software development platform. Balazs Gerofi, Atsushi Hori, Yutaka Ishikawa |
CLUSTER | 3 |
| 2012 | DS-Bench Toolset: Tools for dependability benchmarking with simulation and assuranceabstractToday's information systems have become large and complex because they must interact with each other via networks. This makes testing and assuring the dependability of systems much more difficult than ever before. DS-Bench Toolset has been developed to address this issue, and it includes D-Case Editor, DS-Bench, and D-Cloud. D-Case Editor is an assurance case editor. It makes a tool chain with DS-Bench and D-Cloud, and exploits the test results as evidences of the dependability of the system. DS-Bench manages dependability benchmarking tools and anomaly loads according to benchmarking scenarios. D-Cloud is a test environment for performing rapid system tests controlled by DS-Bench. It combines both a cluster of real machines for performance-accurate benchmarks and a cloud computing environment as a group of virtual machines for exhaustive function testing with a fault-injection facility. DS-Bench Toolset enables us to test systems satisfactorily and to explain the dependability of the systems to the stakeholders. Hajime Fujita 0002, Yutaka Matsuno, Toshihiro Hanawa, Mitsuhisa Sato, Shinpei Kato, Yutaka Ishikawa |
DSN | 6 |
| 2012 | Partial Replication of Metadata to Achieve High Metadata Availability in Parallel File SystemsabstractThis paper presents PARTE, a prototype parallel file system with active/standby configured metadata servers (MDSs). PARTE replicates and distributes a part of files' metadata to the corresponding metadata stripes on the storage servers (OSTs) with a per-file granularity, meanwhile the client file system (client) keeps certain sent metadata requests. If the active MDS has crashed for some reason, these client backup requests will be replayed by the standby MDS to restore the lost metadata. In case one or more backup requests are lost due to network problems or dead clients, the latest metadata saved in the associated metadata stripes will be used to construct consistent and up-to-date metadata on the standby MDS. Moreover, the clients and OSTs can work in both normal mode and recovery mode in the PARTE file system. This differs from conventional active/standby configured MDSs parallel file systems, which hang all I/O requests and metadata requests during restoration of the lost metadata. In the PARTE file system, previously connected clients can continue to perform I/O operations and relevant metadata operations, because OSTs work as temporary MDSs during that period by using the replicated metadata in the relevant metadata stripes. Through examination of experimental results, we show the feasibility of the main ideas presented in this paper for providing high availability metadata service with only a slight overhead effect on I/O performance. Furthermore, since previously connected clients are never hanged during metadata recovery, in contrast to conventional systems, a better overall I/O data throughput can be achieved with PARTE. Jianwei Liao 0001, Yutaka Ishikawa |
ICPP | 2 |
| 2012 | High Performance Checksum Computation for Fault-Tolerant MPI over Infiniband
Alexandre Denis 0001, François Trahay, Yutaka Ishikawa |
EuroMPI | 3 |
| 2012 | An Efficient Kernel-Level Blocking MPI Implementation
Atsushi Hori, Toyohisa Kameyama, Yuichi Tsujita, Mitaro Namiki, Yutaka Ishikawa |
EuroMPI | 5 |
| 2012 | Revisiting Persistent Communication in MPI
Yutaka Ishikawa, Kengo Nakajima, Atsushi Hori |
EuroMPI | 1 |
| 2012 | Delegation-Based MPI Communications for a Hybrid Parallel Computer with Many-Core Architecture
Kazumi Yoshinaga, Yuichi Tsujita, Atsushi Hori, Mikiko Sato, Mitaro Namiki, Yutaka Ishikawa |
EuroMPI | 6 |
| 2012 | Enhancing TCP throughput of highly available virtual machines via speculative communicationabstractCheckpoint-recovery based virtual machine (VM) replication is an attractive technique for accommodating VM installations with high-availability. It provides seamless failover for the entire software stack executed in the VM regardless the application or the underlying operating system (OS), it runs on commodity hardware, and it is inherently capable of dealing with shared memory non-determinism of symmetric multiprocessing (SMP) configurations. There have been several studies aiming at alleviating the overhead of replication, however, due to consistency requirements, network performance of the basic replication mechanism remains extremely poor., Balazs Gerofi, Yutaka Ishikawa |
VEE | 2 |
| 2011 | EZTrace: A Generic Framework for Performance AnalysisabstractModern supercomputers with multi-core nodes enhanced by accelerators, as well as hybrid programming models introduce more complexity in modern applications. Exploiting efficiently all the resources requires a complex analysis of the performance of applications in order to detect time-consuming sections. We present eztrace, a generic trace generation framework that aims at providing a simple way to analyze applications. eztrace is based on plugins that allow it to trace different programming models such as MPI, pthread or OpenMP as well as user-defined libraries or applications. eztrace uses two steps: one to collect the basic information during execution and one post-mortem analysis. This permits tracing the execution of applications with low overhead while allowing to refine the analysis after the execution. We also present a script language for eztrace that gives the user the opportunity to easily define the functions to instrument without modifying the source code of the application. François Trahay, François Rué, Mathieu Faverge, Yutaka Ishikawa, Raymond Namyst, Jack J. Dongarra |
CCGRID | 4 |
| 2011 | RDMA Based Replication of Multiprocessor Virtual Machines over High-Performance InterconnectsabstractWith the growing prevalence of cloud computing and the increasing number of CPU cores in modern processors, symmetric multiprocessing (SMP) Virtual Machines (VM), i.e. virtual machines with multiple virtual CPUs, are gaining significance. However, accommodating SMP virtual machines with high availability at low overhead is still an open problem. Checkpoint-recovery based VM replication is an emerging approach, but it comes with the price of significant performance degradation of the application executed in the VM due to the large amount of state that needs to be synchronized between the primary and the backup machines. Advanced features of high performance interconnects, such as Remote Direct Memory Access (RDMA), on the other hand, offer extreme network throughput. As such feature may provide an opportunity for acceptable performance degradation even for multi-core replicated virtual machines, the impact of such technologies in the domain of VM replication is important to assess. In this paper, we take a first look at the performance advantages of RDMA for SMP virtual machine replication. Moreover, in order to alleviate VM downtime during replication, we propose fine-grained copy-on-write (COW), which protects only memory pages that need to be transferred to the backup host allowing simultaneous execution of the VM with the replication. We find that the performance of replicated virtual machines over high performance interconnects scales well with the number of vCPUs in multiprocessor virtual machines, and that RDMA based replication in conjunction with fine-grained COW imposes acceptable overhead compared to the native VM execution when applied to virtual machines with up to 16 vCPUs. Balazs Gerofi, Yutaka Ishikawa |
CLUSTER | 2 |
| 2011 | Xruntime: A Seamless Runtime Environment for High Performance ComputingabstractMost HPC clusters are now based on an x86 architecture and Linux. In a Grid consisting of such clusters, users might think that a single executable file, a data set, and a single job script will work in all cluster environments. However, due to the lack of interoperability of MPI library implementations, file systems, and batch job systems, users need to be conscious of the runtime environments over the Grid. In order to overcome such differences, we propose a runtime environment called Xruntime consisting of the following three components. MPI-Adapter is middleware between the user program and an MPI implementation to make it possible to execute a single executable code on different MPI implementations. Catwalk is transparent file staging middleware that makes remote files accessed by an application visible as if they were located in the local system. Xruntime allows a user to use various clusters with single batch job script language the user is familiar with, without learning other batch job script languages used on other clusters. Keiji Yamamoto, Atsushi Hori, Shinji Sumimoto, Yutaka Ishikawa |
HPCC | 4 |
| 2011 | Catwalk-ROMIO: A Cost-Effective MPI-IOabstractThe nature of highly parallelized parallel file access which often consists of lots of fine grain, non-contiguous I/O requests, can degrade the I/O performance severely. To tackle this problem, a novel technique to maximize the bandwidth of the MPI-IO is proposed. This proposed technique is utilize a ring communication topology. This technique is implemented as an ADIO device of ROMIO, named Catwalk-ROMIO, and evaluated. The evaluation shows that Catwalk-ROMIO utilizing only one disk can exhibit comparable performance with parallel files systems, PVFS2 and Lustre, utilizing several file servers and disks. The evaluation also shows that Catwalk-ROMIO performance is almost independent from file access patterns, in contrast to the performance of parallel file systems performing only well with collective I/O. Catwalk-ROMIO only requires TCP/IP network for the ring communication topology and one file server which are common in HPC clusters without any additional cost. Thus, Catwalk-ROMIO is considered to be a very cost-effective MPI-IO implementation. Atsushi Hori, Keiji Yamamoto, Yutaka Ishikawa |
ICPADS | 3 |
| 2011 | Workload Adaptive Checkpoint Scheduling of Virtual Machine ReplicationabstractCheckpoint-recovery based Virtual Machine (VM) replication is an emerging approach towards accommodating VM installations with high availability, especially, due to its inherent capability of tackling with symmetric multiprocessing (SMP) virtual machines, i.e. VMs with multiple virtual CPUs (vCPUs). However, it comes with the price of significant performance degradation of the application executed in the VM because of the large amount of state that needs to be synchronized between the primary and the backup machines. Previous research improving VM replication performance focused primarily on decreasing the amount of data transferred over the network, while relying on constant checkpoint frequency. Our goal is to investigate how and to what extent performance degradation can be mitigated by adjusting the checkpoint period dynamically. We provide a comprehensive analysis of various workloads from the aspect of VM replication, paying special attention to their behavior over the increasing number of vCPUs in the system. We propose several heuristics for scheduling replication checkpoints in order to improve quality of service. Our algorithm adapts dynamically to the properties of the workload being executed in the VM, such as changes in the number of dirtied memory pages, network and disk I/O operations, as well as to the network bandwidth available for replication. We evaluate our scheduling algorithm over two network architectures, Gigabit Ethernet and Infiniband, a high-performance interconnect fabric. We find that checkpoint scheduling has a great impact on the performance of replicated virtual machines, and show that replicated virtual machines with up to 16 vCPUs can attain performance close to the native VM execution, not only over high-performance, but also over commercial network architectures. Balazs Gerofi, Yutaka Ishikawa |
PRDC | 2 |
| 2011 | Resource Sharing in GPU-Accelerated Windowing SystemsabstractRecent windowing systems allow graphics applications to directly access the graphics processing unit (GPU) for fast rendering. However, application tasks that render frames on the GPU contend heavily with the windowing server that also accesses the GPU to blit the rendered frames to the screen. This resource-sharing nature of direct rendering introduces core challenges of priority inversion and temporal isolation in multi-tasking environments. In this paper, we identify and address resource-sharing problems raised in GPU-accelerated windowing systems. Specifically, we propose two protocols that enable application tasks to efficiently share the GPU resource in the X Window System. The Priority Inheritance with X server (PIX) protocol eliminates priority inversion caused in accessing the GPU, and the Reserve Inheritance with X server (RIX) protocol addresses the same problem for resource-reservation systems. Our design and implementation of these protocols highlight the fact that neither the X server nor user applications need modifications to use our solutions. Our evaluation demonstrates that multiple GPU-accelerated graphics applications running concurrently in the X Window System can be correctly prioritized and isolated by the PIX and the RIX protocols. Shinpei Kato, Karthik Lakshmanan, Yutaka Ishikawa, Ragunathan Rajkumar |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2011 | RGEM: A Responsive GPGPU Execution Model for Runtime EnginesabstractGeneral-purpose computing on graphics processing units, also known as GPGPU, is a burgeoning technique to enhance the computation of parallel programs. Applying this technique to real-time applications, however, requires additional support for timeliness of execution. In particular, the non-preemptive nature of GPGPU, associated with copying data to/from the device memory and launching code onto the device, needs to be managed in a timely manner. In this paper, we present a responsive GPGPU execution model (RGEM), which is a user-space runtime solution to protect the response times of high-priority GPGPU tasks from competing workload. RGEM splits a memory-copy transaction into multiple chunks so that preemption points appear at chunk boundaries. It also ensures that only the highest-priority GPGPU task launches code onto the device at any given time, to avoid performance interference caused by concurrent launches. A prototype implementation of an RGEM-based CUDA runtime engine is provided to evaluate the real-world impact of RGEM. Our experiments demonstrate that the response times of high-priority GPGPU tasks can be protected under RGEM, whereas their response times increase in an unbounded fashion without RGEM support, as the data sizes of competing workload increase. Shinpei Kato, Karthik Lakshmanan, Mihir Kelkar, Yutaka Ishikawa, Ragunathan Rajkumar |
RTSS | 5 |
| 2011 | Scalable Distributed Monte-Carlo Tree SearchabstractMonte-Carlo Tree Search (MCTS) is remarkably successful in two-player games, but parallelizing MCTS has been notoriously difficult to scale well, especially in distributed environments. For a distributed parallel search, transposition-table driven scheduling (TDS) is known to be efficient in several domains. We present a massively parallel MCTS algorithm, that applies the TDS parallelism to the Upper Confidence bound Applied to Trees (UCT) algorithm, which is the most representative MCTS algorithm. To drastically decrease communication overhead, we introduce a reformulation of UCT called Depth-First UCT. The parallel performance of the algorithm is evaluated on clusters using up to 1,200 cores in artificial game-trees. We show that this approach scales well, achieving 740-fold speedups in the best case. Kazuki Yoshizoe, Akihiro Kishimoto, Tomoyuki Kaneko, Haruhiro Yoshimoto, Yutaka Ishikawa |
SOCS | 5 |
| 2011 | TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments
Shinpei Kato, Karthik Lakshmanan, Ragunathan Rajkumar, Yutaka Ishikawa |
USENIX ATC | 4 |
| 2011 | CPU scheduling and memory management for interactive real-time applications
Shinpei Kato, Yutaka Ishikawa, Ragunathan Rajkumar |
Real Time Syst. | 2 |
| 2010 | An Efficient Process Live Migration Mechanism for Load Balanced Distributed Virtual EnvironmentsabstractDistributed virtual environments (DVE), such as multi-player online games and distributed simulations may involve a massive amount of concurrent clients. Deploying distributed server architectures is currently the most prevalent way of providing such large-scale services, where typically the virtual space is divided into several distinct regions requiring each server to handle only part of the virtual world. Inequalities in client distribution may, however, cause certain servers to become overloaded, which potentially degrades the interactivity of the environment and thus renders the load balancing problem a crucial issue. Prior research has shown several approaches for avoiding uneven workload, nevertheless, addressing the problem mainly at the application layer. In this paper we focus on solving the DVE load balancing problem at the operating system level. We propose an efficient process live migration mechanism, which is optimized for processes maintaining a massive amount of network connections. Building on top of it, we have implemented a decentralized middleware that instruments process migration among the cluster nodes, attempting to equalize loads on all machines. We demonstrate the performance of the live migration mechanism on a real-world multiplayer game server and show the behavior of the load balancing engine through a realistic DVE simulation. Balazs Gerofi, Hajime Fujita 0002, Yutaka Ishikawa |
CLUSTER | 3 |
| 2010 | Optimization Techniques at the I/O Forwarding LayerabstractI/O is the critical bottleneck for data-intensive scientific applications on HPC systems and leadership-class machines. Applications running on these systems may encounter bottlenecks because the I/O systems cannot handle the overwhelming intensity and volume of I/O requests. Applications and systems use I/O forwarding to aggregate and delegate I/O requests to storage systems. In this paper, we present two optimization techniques at the I/O forwarding layer to further reduce I/O bottlenecks on leadership-class computing systems. The first optimization pipelines data transfers so that I/O requests overlap at the network and file system layer. The second optimization merges I/O requests and schedules I/O request delegation to the back-end parallel file systems. We implemented these optimizations in the I/O Forwarding Scalability Layer and them on the T2K Open Supercomputer at the University of Tokyo and the Surveyor Blue Gene/P system at the Argonne Leadership Computing Facility. On both systems, the optimizations improved application I/O throughput, but highlighted additional areas of I/O contention at the I/O forwarding layer that we plan to address. Kazuki Ohta, Dries Kimpe, Jason Cope, Kamil Iskra, Robert B. Ross, Yutaka Ishikawa |
CLUSTER | 6 |
| 2010 | A New Concurrent Checkpoint Mechanism for Real-Time and Interactive ProcessesabstractThis paper presents a new concurrent checkpoint mechanism that allows the checkpointed process to run without stopping while checkpoints are set. The checkpointed process can keep running until a memory access request is captured by tracing TLB misses while dumping memory pages (the most time-consuming step when setting a checkpoint). At that time, the checkpointer in the kernel will copy the memory access target page to the designated memory buffer for constructing a consistent state of the checkpointed process, and then resume the memory access. From the experimental results, in contrast to non-concurrent checkpoint techniques, this mechanism can reduce the downtime time of the checkpointed process by 47.4% - 89.8% to ensure concurrency between setting a checkpoint and execution of the checkpointed process. In addition, compared with a traditional concurrent checkpoint system, this mechanism saves more than 2.2% of the checkpoint time and decreases the downtime of the checkpointed process by more than 10%. Jianwei Liao 0001, Yutaka Ishikawa |
COMPSAC | 2 |
| 2010 | AIRS: Supporting Interactive Real-Time Applications on Multicore PlatformsabstractModern real-time systems increasingly operate with multiple interactive applications. While these systems often require reliable quality of service (QoS) for the applications, even under heavy workloads, many existing CPU schedulers are not very capable of satisfying such requirements. In this paper, we design and implement an Advanced Interactive and Real-time Scheduler, called AIRS. AIRS is aimed at supporting systems that run multiple interactive real-time applications, particularly on multicore platforms. It provides a new CPU reservation mechanism to enhance the QoS of the overall system. The reservation algorithm is based on the prior Constant Bandwidth Server (CBS) algorithm, but is more flexible and efficient, when multiple applications reserve CPU bandwidth. It also provides a new multicore scheduler to improve the absolute CPU bandwidth available for the applications to perform well. The scheduling algorithm is subject to the prior Earliest Deadline First with Window-constraint Migration (EDF-WM) algorithm, but is extended to work with the new CPU reservation mechanism. Experimental evaluation shows that AIRS delivers higher quality to simultaneous playback of multiple movies than the existing real-time scheduler. It also demonstrates that AIRS offers hard timing guarantees for randomly-generated task sets with heavy workloads. Shinpei Kato, Ragunathan Rajkumar, Yutaka Ishikawa |
ECRTS | 3 |
| 2010 | A Multi-core Approach to Providing Fault Tolerance for Non-deterministic ServicesabstractWith the advent of multi- and many-core architectures, new opportunities in fault-tolerant computing have become available. In this paper we propose a novel process replication method that provides transparent failover of non-deterministic TCP services by utilizing spare CPU cores. Our method does not require any changes to the TCP protocol, does not require any changes to the client software, and unlike existing solutions, it does not require any changes to the server applications either. We measure performance overhead on two real-world applications, a multimedia streaming service and an Internet Relay Chat daemon and show that the imposed overhead is minimal as the price of seamless failover. Our prototype implementation consists of a kernel module for Linux 2.6 without any changes to the existing kernel code. Balazs Gerofi, Yutaka Ishikawa |
NCA | 2 |
| 2010 | P-Bus: Programming Interface Layer for Safe OS Kernel ExtensionsabstractP-Bus, a new programming interface layer for safe kernel extensions is proposed. P-Bus introduces a new programming interface on top of the Linux kernel in order to give formal specifications to the interface, and to improve portability of extensions. New extensions, called P-Components, are verified with a model checker MKencha to see whether a component is compliant with rules which should be obeyed to implement extensions properly. A network driver has been implemented as a P-Component and verified with MKencha. MKencha has found two bugs in the component. Hajime Fujita 0002, Motohiko Matsuda, Toshiyuki Maeda, Shin'ichi Miura, Yutaka Ishikawa |
PRDC | 5 |
| 2010 | Design and Implementation of a Fault Tolerant Single IP Address ClusterabstractAn F-FTCS mechanism that develops a fault tolerant single IP address cluster for TCP applications is proposed. The FTCS mechanism performs fine grain load balancing by handling all incoming TCP connection requests with a master node. Three fail-over algorithms are designed and implemented to carry out the fault tolerant FTCS mechanism. Discarding and Gathering Algorithms discard and gather TCP connections whose state is SYN-RECEIVED, respectively, at failure. A Scattering Algorithm synchronizes the information between nodes in the failure-free phase. These three algorithms are evaluated on Core 2 Duo machines. The Discarding Algorithm recovers from a failure from 440 to 950 msec earlier than the Gathering Algorithm, but it requires reprocessing the discarded TCP connection requests. The Scattering Algorithm requires from 120 to 160 usec more overhead during processing of a TCP connection request than that of the original FTCS mechanism. Jun Kato 0002, Hajime Fujita 0002, Yutaka Ishikawa |
PRDC | 3 |
| 2010 | Towards a Language for Communication among StakeholdersabstractComputers are now present almost everywhere and connected into ever more complex networks. This means not only that embedded systems are more complicated, but also that communication among the diverse stakeholders of systems is much harder than before. This paper introduces the D-Case approach to a systematic explanation of embedded-systems dependability. A D-Case is a structured document that argues for the dependability of a system, supported by evidence. This extends the notion of safety cases commonly used in (European) safety-critical sectors. The goal is to develop the D-Case language for communication systems dependability among the stakeholders. The paper reports the experience in constructing a D-Case for the remote test surveillance system developed to demonstrate certain dependability system components. D-Case construction is shown to be an effective method in explaining how each system component contributes to the overall dependability of the system. Another experiment shows how the D-Case approach can promote dependability through the life cycle of a larger system. Finally, the paper presents some comments on the difficulties and insights for future work. Yutaka Matsuno, Jin Nakazawa, Makoto Takeyama, Midori Sugaya, Yutaka Ishikawa |
PRDC | 5 |
| 2010 | Design of Kernel-Level Asynchronous Collective Communication
Akihiro Nomura 0002, Yutaka Ishikawa |
EuroMPI | 2 |
| 2009 | Improving Parallel Write by Node-Level Request SchedulingabstractIn a cluster of multiple processors or cpu-cores, many processes may run on each compute node. Each process tends to issue contiguous I/O requests for snapshot, checkpointing or so, however, if large number of processes enter the I/O phase at the same time, the requests from the same process may be interrupted by the requests of other processes. Then, the I/O nodes receive these requests as non-contiguous way. This interleaved access pattern causes performance degradation in parallel file systems. In order to overcome the problem, we have designed the gather-arrange-scatter (GAS) I/O architecture, for optimizing the parallel write performance. The GAS is an architecture for capturing write operations, buffering them in the memory, and scheduling them to reduce I/O cost at I/O nodes. The scheduling is done per compute node, and the requests are sent to the remote disks in parallel. In this paper, after introducing the GAS architecture in detail, its efficiency and scalability are evaluated using the NAS Parallel Benchmark BTIO. GAS is 5.2%faster than ROMIO collective I/O on PVFS2 in BTIO with 16 nodes/64 processes, and 34.9% faster than MPI noncollective I/O in the same configuration. Kazuki Ohta, Hiroya Matsuba, Yutaka Ishikawa |
CCGRID | 3 |
| 2009 | On-demand file staging system for Linux clustersabstractAn on-demand file staging system, Catwalk, is proposed. Catwalk is designed so that it can run on any Linux clusters without any special or additional hardware. By having hook functions on the system calls of file operations, a file staging system can be transparent from the view of users, and users can be free from having wrong file staging scripts. In Catwalk, the file copying is done via normal TCP protocol so that Catwalk can run over ordinary, widely-used Ethernet. The stage-in file copy is pipelined to maximize the bandwidth from single file server. The performance of Catwalk is evaluated and compared with NFS using synthetic but realistic workloads. The evaluations show the stage-in performance with the pipeline technique is much better than the performance of NFS. The stage-out performance is comparable with the NFS performance despite the extra copying of files, and the file server is lightly loaded with the Catwalk stage-out while NFS entails much heavier server loads. The biggest problems of NFS are its centralized design and lack of scheduling for the parallel workloads. The performance of Catwalk shows that remote file access performance can be improved much better if file accesses are scheduled in a proper way. Thus the proposed file staging system can be a strong complement to NFS, especially for small clusters often having no dedicated parallel file system. Atsushi Hori, Yoshikazu Kamoshida, Hiroya Matsuba, Kazuki Ohta, Takashi Yasui, Shinji Sumimoto, Yutaka Ishikawa |
CLUSTER | 7 |
| 2009 | Semi-partitioned Scheduling of Sporadic Task Systems on MultiprocessorsabstractThis paper presents a new algorithm for scheduling of sporadic task systems with arbitrary deadlines on identical multiprocessor platforms. The algorithm is based on the concept of semi-partitioned scheduling, in which most tasks are fixed to specific processors, while a few tasks migrate across processors. Particularly, we design the algorithm so that tasks are qualified to migrate only if a task set cannot be partitioned any more, and such migratory tasks migrate from one processor to another processor only once in each period. The scheduling policy is then subject to earliest deadline first. simulation results show that the algorithm delivers competitive scheduling performance to the state-of-the-art, with a smaller number of context switches. Shinpei Kato, Nobuyuki Yamasaki, Yutaka Ishikawa |
ECRTS | 3 |
| 2009 | Towards an Open Dependable Operating SystemabstractThis paper introduces a new dependable operating system project, called DEOS, started in 2006, and scheduled to continue for six years. In this project, a safety extension mechanism called P-Bus is to be designed, and implemented in the Linux kernel so that a future dependability attribute is implemented with P-Bus. A hardware abstraction layer, called SPUMONE, is introduced so that a light-weight operating system, called ArcOS, and a monitoring service on top of ArcOS monitors the Linux kernel to provide a safety-net for the Linux kernel. New dependability metrics are being designed to enable developers and users to decide which hardware or software solution meets their dependability requirements, and thus can be used. Yutaka Ishikawa, Hajime Fujita 0002, Toshiyuki Maeda, Motohiko Matsuda, Midori Sugaya, Mitsuhisa Sato, Toshihiro Hanawa, Shin'ichi Miura, Taisuke Boku, Yuki Kinebuchi, Tatsuo Nakajima, Jin Nakazawa, Hideyuki Tokuda |
ISORC | 1 |
| 2009 | Delayed Processing Technique in Critical Sections for Real-Time LinuxabstractIn a real-time Linux system, the critical sections are thought to be one of the main factors causing problems with the start of real-time tasks. Traditional approaches for overcoming this issue either provide less of a guarantee on the worst-case latency time of real-time tasks, or have heavy overhead on normal Linux tasks. In this paper, to guarantee the start time of a real-time task, the execution of a normal Linux task will be delayed, made to wait at the beginning of a critical section, on the assumption that the future execution of this section would lead to an unacceptable delay time in the start of the coming real-time task. In addition to this, to reduce the latency time of the real-time task, a technique is proposed in which hardware interrupts will not be prohibited in the most kernel's critical sections, so the timer interrupt can enter the kernel with no or less delay time. Experimental results showed that the worst-case start latency of a real-time task is reduced to 16.7% of that in Linux 2.6.20, and the penalty to the normal tasks is light, in contrast to traditional approaches. The proposed technique is useful not only for constructing a real-time Linux, but also for developing other real-time systems in which the critical sections are significantly long. Maobing Dai, Yutaka Ishikawa |
PRDC | 2 |
| 2009 | Gang EDF Scheduling of Parallel Task SystemsabstractThe preemptive real-time scheduling of sporadic parallel task systems is studied. We present an algorithm, called gang EDF, which applies the earliest deadline first (EDF) policy to the traditional gang scheduling scheme. We also provide schedulability analysis of gang EDF. Specifically, the total amount of interference that is necessary to cause a deadline miss is first identified. The contribution of each task to the interference is then bounded. Finally, verifying that the total amount of contribution does not exceed the necessary interference for every task, the schedulability test is derived. Although the techniques proposed herein are based on the prior results for the sequential task model, we introduce new ideas for the parallel task model. Shinpei Kato, Yutaka Ishikawa |
RTSS | 2 |
| 2008 | TCP Connection Scheduler in Single IP Address ClusterabstractA broadcast-based single IP cluster aims at being both scalable and available. However, existing systems can only employ static traffic assignment based on incoming packets. In this paper we propose FTCS, a new TCP connection dispatching mechanism that enables a single IP cluster to use more flexible load balancing algorithms. In this mechanism, one of the cluster nodes acts as a master node. A centralized connection scheduler runs on the master node in order to dispatch TCP connections to nodes of the clusters. Since connections are scheduled by a single scheduler, the master node is able to employ arbitrary scheduling algorithms. Once a TCP connection is established on a node, succeeding communication is handled without involving the master node. When the master node fails, one of the nodes takes over the role of the master node. Therefore the master node does not become a single point of failure. Benchmark results using SPECweb2005 Support benchmark show that a four-node Linux cluster using FTCS balances workloads well and successfully handles 13% more requests than the existing method, on average. Hajime Fujita 0002, Hiroya Matsuba, Yutaka Ishikawa |
CCGRID | 3 |
| 2008 | High Performance Relay Mechanism for MPI Communication Libraries Run on Multiple Private IP Address ClustersabstractWe have been developing a Grid-enabled MPI communication library called GridMPI, which is designed to run on multiple clusters connected to a wide-area network. Some of these clusters may use private IP addresses. Therefore, some mechanism to enable communication between private IP address clusters is required. Such a mechanism should be widely adoptable, and should provide high communication performance. In this paper, we propose a message relay mechanism to support private IP address clusters in the manner of the Interoperable MPI (IMPI) standard. Therefore, any MPI implementations which follow the IMPI standard can communicate with the relay. Furthermore, we also propose a trunking method in which multiple pairs of relay nodes simultaneously communicate between clusters to improve the available communication bandwidth. While the relay mechanism introduces an one-way latency of about 25 musec, the extra overhead is negligible, since the communication latency through a wide area network is a few hundred times as large as this. By using trunking, the inter-cluster communication bandwidth can improve as the number of trunks increases. We confirmed the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. When relay nodes with 1 Gbps NICs are used, the performance of most of the NAS Parallel Benchmarks improved proportional to the number of trunks. Especially, using 8 trunks, FT and IS are 4.4 and 3.4 times faster, respectively, compared with the single trunk case. The results showed that the proposed method is effective for running MPI programs over high bandwidth-delay product networks. Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa, Yasufumi Yoshizawa |
CCGRID | 6 |
| 2008 | Gather-arrange-scatter: Node-level request reordering for parallel file systems on multi-core clustersabstractMultiple processors or multi-core CPUs are now in common, and the number of processes running concurrently is increasing in a cluster. Each process issues contiguous I/O requests individually, but they can be interrupted by the requests of other processes if all the processes enter the I/O phase together. Then, I/O nodes handle these requests as non-contiguous. This increases the disk seek time, and causes performance degradation. Kazuki Ohta, Hiroya Matsuba, Yutaka Ishikawa |
CLUSTER | 3 |
| 2008 | Logical Partitioning without Architectural SupportsabstractAs a method for running multiple operating systems on one machine, we propose a new resource partitioning method we have named "single hardware with independent multiple operating systems" (SHIMOS). In SHIMOS, CPU and memory resources are partitioned by multiple native kernels without any architectural virtualization supports. There is nearly no slowdown, unlike VMs, because the kernel and user programs are executed directly by the real CPUs. To evaluate our method, we compare an implementation on x86 with VMs by running benchmarks simultaneously. From this comparison, we show that the method proposed is as fast as a real machine, and that it can be twice as fast as the existing VM method in executing I/O- oriented processes. Taku Shimosawa, Hiroya Matsuba, Yutaka Ishikawa |
COMPSAC | 3 |
| 2008 | A Light Lock Management Mechanism for Optimizing Real-Time and Non-Real-Time Performance in Embedded LinuxabstractIn a real-time Linux system, the critical sections are known as the main factor delaying the execution of real-time tasks. Traditional approaches to overcoming this issue have given less consideration to both real-time and non-real-time tasks. In this paper, we propose a new lock management mechanism to improve the real-time performance with a small penalty for non-real-time tasks. Using this mechanism, we guarantee the deadlines of real-time tasks while keeping the penalties accruing for non-real-time tasks small. We implemented a prototype system in Linux 2.6.20. Experimental results showed that the worst-case OS latency of real-time task is reduced to 19% of the original one, while the penalty for a non-real-time task is 10.1% of the original. The results also showed that the lock management mechanism proposed in this paper is efficient and useful for a future real-time Linux system. Maobing Dai, Toshihiro Matsui, Yutaka Ishikawa |
EUC (1) | 3 |
| 2007 | Network performance model for TCP/IP based cluster computingabstractA new communication model, called the PlogPT model, is proposed to predict communication performance in a commodity cluster where computing nodes communicate using TCP/IP. This model extends the PlogP model in order to consider the variation of bandwidth brought about by bottleneck links in network switches and the delay of packet retransmission by TCP/IP handling. Network switches are modeled as binary tree connections. To demonstrate the PlogPT model’s modeling capability, the execution time of two all-to-all communication algorithms are estimated and compared with the actual execution time and that of the PlogP model. The PlogPT model predicts the execution time of those two algorithms more precisely than the existing model. Akihiro Nomura 0002, Hiroya Matsuba, Yutaka Ishikawa |
CLUSTER | 3 |
| 2007 | Effects of packet pacing for MPI programs in a Grid environmentabstractImproving the performance of TCP communication is the key to the successful deployment of MPI programs in a Grid environment in which multiple clusters are connected through high performance dedicated networks. To efficiently utilize the inter-cluster bandwidth, a traffic control mechanism is required so as not to allow the aggregate transmission bandwidth to exceed the inter-cluster bandwidth when multiple nodes communicate at one time. In this paper, we propose a traffic control method for MPI programs, in which an application or the MPI runtime controls the transmission rate based on the communication pattern by using certain MPI attributes. Packet pacing is used at each node preventing microscopic burst transmission to thus avoid congestion. We confirm the effectiveness of the proposed method by experiments using a 10 Gbps emulated WAN environment. We show most of the NAS Parallel benchmarks improve the performance, since the proposed method reduces packet losses due to traffic congestion on the inter-cluster network. The results have indicated that it is feasible to connect multiple clusters and run large-scale scientific applications over distances up to 1000 kilometers, if an appropriate network is available. Ryousei Takano, Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Fumihiro Okazaki, Yutaka Ishikawa |
CLUSTER | 6 |
| 2007 | Single IP Address Cluster for Internet ServersabstractOperating a cluster on a single IP address is required when the cluster is used to provide certain Internet services. This paper proposes SAPS, a new method to assign a single IP address to a cluster. The TCP/IP protocol is handled at a single node called the I/O server. The other nodes, called application nodes, provide the socket interface to applications. The I/O server and applications nodes are connected using a cluster-dedicated network, such as the Myrinet network. The key benefit of the proposed method is that the TCP/IP protocol does not care about congestion and packet loss in the cluster, which often happens if multiple nodes send packets to the bottleneck router. Instead, the cluster-dedicated network manages the packet congestion more efficiently than the TCP/IP protocol. The result of the bandwidth benchmark shows SAPS fully utilizes the bandwidth of the Gigabit Ethernet. The result of the SPEC Web benchmark shows SAPS handles 7.9% more requests than the existing method. Hiroya Matsuba, Yutaka Ishikawa |
IPDPS | 2 |
| 2007 | Fault Detection System Activated by Failure InformationabstractWe propose a fault detection system activated by an application when the application recognizes the occurrence of a failure, in order to realize self managing systems that automatically find the source of a failure. In existing detection systems, there are three issues for constructing self managing applications: i) the detection results are not sent to the applications, ii) they can not identify the source failure from all of the detected failures, and iii) configuring the detection system for networked system is hard work. For overcoming these issues, the proposed system takes three approaches: i) the system receives failure information from an application and returns a result set to the application, ii) the system identifies the source failure using relationships among errors, and Hi) the system obtains information of the monitored system from a database. The relationship is expressed by a tree. This tree is called error relationship tree. The database provides information which are system entities such as hardware devices, software object, and network topology. When the proposed system starts looking for the source of a failure, causal relations from an error relation tree are referred to, and the correspondence of error definitions and actual objects is derived using the database. We show the design of the detection operation activated by the failure information and the architecture of the proposed system. Masato Sakai, Hiroya Matsuba, Yutaka Ishikawa |
PRDC | 3 |
| 2006 | Efficient MPI Collective Operations for Clusters in Long-and-Fast NetworksabstractSeveral MPI systems for grid environment, in which clusters are connected by wide-area networks, have been proposed. However, the algorithms of collective communication in such MPI systems assume relatively low bandwidth wide-area networks, and they are not designed for the fast wide-area networks that are becoming available. On the other hand, for cluster MPI systems, a beast algorithm by van de Geijn et al. and an allreduce algorithm by Rabenseifner have been proposed, which are efficient in a high bisection bandwidth environment. We modify those algorithms so as to effectively utilize fast wide-area inter-cluster networks and to control the number of nodes which can transfer data simultaneously through wide-area networks to avoid congestion. We confirmed the effectiveness of the modified algorithms by experiments using a 10 Gbps emulated WAN environment. The environment consists of two clusters, where each cluster consists of nodes with 1 Gbps Ethernet links and a switch with a 10 Gbps upper link. The two clusters are connected through a 10 Gbps WAN emulator which can insert latency. In a 10 millisecond latency environment, when the message size is 32 MB, the proposed beast and allreduce are 1.6 and 3.2 times faster, respectively, than the algorithms used in existing MPI systems for grid environment Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa |
CLUSTER | 5 |
| 2006 | Portable Execution Time Analysis MethodabstractWe propose a new execution time prediction method that combines measurement-based execution time analysis and simulation-based memory access analysis. In measurement-based execution time analysis, the target program is divided into basic blocks, to each of which a memory area accessed by the block is allocated, so that all the basic block execution times are measured on a real machine. Since the execution behavior of such a basic block is not a real case, simulation-based memory access analysis is introduced to calculate the memory access cost. The method has been implemented using the intermediate expressions (both TREE and RTL expressions) used in GCC (Gnu compiler collection). This paper demonstrates that the proposed method predicts the execution time safely in different architecture environments, i.e. Pentium-M and XScale Keiji Yamamoto, Yutaka Ishikawa, Toshihiro Matsui |
RTCSA | 2 |
| 2005 | TCP Adaptation for MPI on Long-and-Fat NetworksabstractTypical MPI applications work in phases of computation and communication, and messages are exchanged in relatively small chunks. This behavior is not optimal for TCP because TCP is designed only to handle a contiguous flow of messages efficiently. This behavior anomaly is well-known, but fixes are not integrated into today's TCP implementations, even though performance is seriously degraded, especially for MPI applications. This paper proposes three improvements in the Linux TCP stack: i.e., pacing at start-up, reducing Retransmit-Timeout time, and TCP parameter switching at the transition of computation phases in an MPI application. Evaluation of these improvements using the NAS parallel benchmarks shows that the BT, CG, IS, and SP benchmarks achieved 10 to 30 percent improvements. On the other hand, the FT and MG benchmarks showed no improvement because they have the steady communication that TCP assumes, and the LU benchmark became slightly worse because it has very little communication Motohiko Matsuda, Tomohiro Kudoh, Yuetsu Kodama, Ryousei Takano, Yutaka Ishikawa |
CLUSTER | 5 |
| 2005 | Workshop on Dependable Software - Tools and Methods - Workshop Abstract
Takuya Katayama, Yutaka Ishikawa, Yoshiki Kinoshita |
DSN | 2 |
| 2005 | Distributed Real-Time Processing for Humanoid RobotsabstractA humanoid robot is a real-time system controlled by a complex computer system that requires huge computing power for perception and planning, high energy efficiency for self-contained control, reduction of physical dimensions, and high reliability. This paper proposes a distributed architecture for the humanoid robot control substituting conventional centralized control architectures. In addition to the parallelism that provides scalable computing power at low clock namely at low energy, the distributed architecture contributes to reliable operations by replacing many fragile analog signal wires with a digital network with redundant routes. In order to accomplish a real-time control over the network, RMTP (responsive multi-threaded processor) for parallel and real-time computation has been newly designed. RMTP can synchronize more than thirty nodes distributed over a robot body in less than 5 micro second with a real time network called the responsive link (RL). Architectures of RMTP, RL and Linux-based real-time system software are presented. Toshihiro Matsui, Hirohisa Hirukawa, Yutaka Ishikawa, Nobuyuki Yamasaki, Satoshi Kagami, Fumio Kanehiro, Hajime Saito, Tetsuya Inamura |
RTCSA | 3 |
| 2004 | The design and implementation of an asynchronous communication mechanism for the MPI communication modelabstractMany implementations of an MPI communication library are realized on top of the socket interface which is based on connection-oriented stream communication. This work addresses a mismatch between the MPI communication model and the socket interface. In order to overcome a mismatch and implement an efficient MPI library for large-scale commodity-based clusters, a new communication mechanism, called 02G, is designed and implemented. O2G integrates receive queue management of MPI into a TCP/IP protocol handler, without modifying the protocol stacks. Received data is extracted from the TCP receive buffer and copied into the user space within the TCP/IP protocol handler invoked by interrupts. It totally avoids polling of sockets and reduces system call overhead, which becomes dominant in large-scale clusters. In addition, its immediate and asynchronous receive operation avoids message flow disruption due to a shortage of capacity in the receive buffer, and keeps the bandwidth high. An evaluation using the NAS Parallel Benchmarks shows that 02G made an MPI implementation up to 30 percent faster than the original one. An evaluation on bandwidth also shows that 02G made an MPI implementation independent of the number of connections, while an implementation with sockets was greatly affected by the number of connections. Motohiko Matsuda, Tomohiro Kudoh, H. Tazuka, Yutaka Ishikawa |
CLUSTER | 4 |
| 2003 | Evaluation of MPI Implementations on Grid-connected Clusters using an Emulated WAN EnvironmenabstractThe MPICH-SCore high performance communication library for cluster computing is integrated into the MPICHG-2 library in order to adapt PC clusters to a Grid environment. The integrated library is called MPICH-G2/SCore. In addition, for the purpose of comparison with other approaches, MPICH-SCore itself is extended to encapsulate its network packet into a UDP packet so that packets are delivered via L3 switches. This extension is called UDP-encapsulated MPICH-SCore. In this paper, three implementations of the MPI library, UDP-encapsulated MPICH-SCore, MPICH-G2/SCore, and MPICH-P4, are evaluated using an emulated WAN environment where two clusters, each consisting of sixteen hosts, are connected by a router PC. The router PC controls the latency of message delivery between clusters, and the added latency is varied from I millisecond to 4 milliseconds in round-trip time. Experiments are performed using the NAS Parallel Benchmarks, which show UDP-encapsulated MPICH-SCore most often performs better than other implementations. However, the differences are not critical for the benchmarks. The preliminary results show that the performance of the LU benchmark scales up linearly with under 4 millisecond round-trip latency. The CG and MG benchmarks show the scalability of 1.13 and 1.24 times with 4 millisecond round-trip latency, respectively. Motohiko Matsuda, Tomohiro Kudoh, Yutaka Ishikawa |
CCGRID | 3 |
| 2003 | Performance of Cluster-enabled OpenMP for the SCASH Software Distributed Shared Memory SystemabstractOpenMP has attracted widespread interest because it is an easy-to-use parallel programming model for shared memory multiprocessor systems. Implementation of a "cluster-enabled" OpenMP compiler is presented. Compiled programs are linked to the page-based software distributed-shared-memory system, SCASH, which runs on PC clusters. This allows OpenMP programs to be run transparently in a distributed memory environment. The compiler converts programs written for OpenMP into parallel programs using the SCASH static library, moving all shared global variables into SCASH shared address space at runtime. As data mapping has a great impact on the performance of OpenMP programs compiled for software distributed-shared-memory, extensions to OpenMP directives are defined for specifying data mapping and loop scheduling behavior, allowing data to be allocated to the node where it is to be processed. Experimental results of benchmark programs on PC clusters using both Myrinet and fast Ethernet are reported. Yoshinori Ojima, Mitsuhisa Sato, Hiroshi Harada, Yutaka Ishikawa |
CCGRID | 4 |
| 2002 | Exploiting cluster networks for distributed object groups and collective operations
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Future Gener. Comput. Syst. | 3 |
| 2001 | TACO-Exploiting Cluster Networks for High-Level Collective OperationsabstractTACO (Topologies and Collections) is a template library that introduces the flavour of distributed data parallel processing by means of reusable topology classes and C++ templates. The paper introduces TACO's basic abstractions and provides a performance analysis for basic collective operations on various cluster architectures with several different networks. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
CCGRID | 3 |
| 2000 | Consistent Checkpointing for High Performance Clusters
Toshihiro Nishioka, Atsushi Hori, Yutaka Ishikawa |
CLUSTER | 3 |
| 2000 | TACO -- Dynamic Distributed Collections with Templates and Topologies
Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
Euro-Par | 3 |
| 2000 | High Performance Communication using a Commodity Network for Cluster SystemsabstractProposes a scheme to realize a high-performance communication facility using a commodity network. This scheme does not require any special hardware or hardware-specific device drivers in order to adapt to many kinds of network interface cards (NICs). In this scheme, a reliable lightweight network protocol is handled directly on a data link layer called by a network device driver. An interrupt reaping technique is proposed to eliminate the hardware interrupt overhead when an application waits for a message. PM/Ethernet, an instance of the scheme, is implemented on Linux with minimal modification to the Linux kernel, and existing network device drivers are used without any modification. Using Pentium III 500-MHz PCs on Packet Engine's G-NIC II Gigabit Ethernet NIC, it achieves 77.5 MB/s bandwidth and 37.6 /spl mu/s round-trip time latency compared to that of TCP/IP, which achieves 46.7 MB/s bandwidth and 89.6 /spl mu/s round-trip time latency. The NAS parallel benchmark IS results show that MPI on PM/Ethernet achieves 75% better performance than MPI on TCP/IP and is 7.8% slower than that of MPI on Myrinet PM. Shinji Sumimoto, Hiroshi Tezuka, Atsushi Hori, Hiroshi Harada, Toshiyuki Takahashi, Yutaka Ishikawa |
HPDC | 6 |
| 2000 | Template Based Structured CollectionsabstractCollective operations on distributed data sets foster a high-level data-parallel programming style that eases many aspects of parallel programming significantly. In this paper we describe how higher-order collective operations on distributed object sets can be introduced in a structured way by means of reusable topology classes and C++ templates. Jörg Nolte, Mitsuhisa Sato, Yutaka Ishikawa |
IPDPS | 3 |
| 2000 | PM2: A High Performance Communication Middleware for Heterogeneous Network EnvironmentsabstractThis paper introduces a high performance communication middle layer, called PM2, for hetero-geneous network environments. PM2 currently supports Myrinet, Ethernet, and SMP. Binary code written in PM2 or written in a communication library, such as MPICH-SCore on top of PM2, may run on any combination of those networks without re-compilation. According to a set of NAS parallel benchmark results, MPICH-SCore performance is better than dedicated communication libraries such as MPICH-BIP/SMP and MPICH-GM when running some benchmark programs. Toshiyuki Takahashi, Shinji Sumimoto, Atsushi Hori, Hiroshi Harada, Yutaka Ishikawa |
SC | 5 |
| 1999 | The design and evaluation of high performance communication using a Gigabit EthernetabstractA high performance communication facility, called the GigaE PM, has been designed and implemented for parallel applications on clusters of computers using a Gigabit Ethernet. The GigaE PM provides not only a reliable high bandwidth and low latency communication function, but also supports existing network protocols such as TCP/IP. In the design of the GigaE PM, it is assumed that the Gigabit Ethernet card used has a dedicated processor and its program can be modied. A reliable communication mechanism for a parallel application is implemented on the rmware while existing network protocols are handled by an operating system kernel. A prototype system has been implemented using an Essential Communications Gigabit Ethernet card. The performance results show that a 48.3 s round trip time for a four byte user message, and 56.7 MBytes/sec bandwidth for a 1,468 byte message have been achieved on Intel Pentium II 400 MHz PCs. We have implemented MPICH-PM on top of the GigaE PM, and evaluat... Shinji Sumimoto, Hiroshi Tezuka, Atsushi Hori, Hiroshi Harada, Toshiyuki Takahashi, Yutaka Ishikawa |
International Conference on Supercomputing | 6 |
| 1998 | The Design and Implementation of Zero Copy MPI Using Commodity Hardware with a High Performance NetworkabstractThis paper designs an implementation of the MPI message passing interface using a zero copy message transfer primitive supported by a lower communication layer to realize a high performance communication library.The zero copy message transfer primitive requires a memory area pinned down to physical memory, which is a restricted quantity resource under a paging memory system.Allocation of pinned down memory by multiple simultaneous requests for sending and receiving without any control can cause deadlock.To avoid this deadlock, we have introduced: i) separate of control of send/receive pin-down memory areas to ensure that at least one send and receive may be processed concurrently, and ii) delayed queues to handle the postponed message passing operations which could not be pinned-down. Francis O'Carroll, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa |
International Conference on Supercomputing | 4 |
| 1998 | Overhead Analysis of Preemptive Gang Scheduling
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
JSSPP | 3 |
| 1998 | Highly Efficient Gang Scheduling ImplementationabstractA new and more highly efficient gang scheduling implementation technique is the basis for this paper. Network preemption, in which network interface contexts are saved and restored, has already been proposed to enable parallel applications to perform efficent user-level communication. This network preemption technique can be used to for detecting global state, such as deadlock, of a parallel program execution. A gang scheduler, SCore-D, using the network preemption technique is implemented with PM, a user-level communication library. This paper evaluates network preemption gang scheduling overhead using eight NAS parallel benchmark programs. The results of this evaluation illustrate that the saving and restoring network contexts occupies almost half of the total gang scheduling overhead. A new mechanism, having multiple network contexts and merely switching the context pointers without saving and restoring the network contexts, is proposed. The NAS parallel benchmark evaluation shows that gang scheduling overhead is almost halved. The maximum gang scheduling overhead among benchmark programs is less than 10%, with a 40msec time slice on 64 single-way PentiumPros, connected by Myrinet to form a PC cluster. The numbers of secondary cache misses are counted, and it is found that network preemption with multiple network contexts is more cache-effective than a single network context. The observed scheduling overhead for applications running on 64 nodes can only be a small percent of the execution time. The gang scheduling overheads of switching two NAS parallel benchmark programs are also evaluated. The additional overheads are less than 2% in most cases, with a 100msec time slice on 64 nodes. This slightly higher scheduling overheads than for switching a single parallel process comes from more frequent cache misses. This paper contributes the following findings; i) gang scheduling overhead with network preemption can be sufficiently low, ii) proposed network preemption with multiple network contexts is more cache-effective than a single network context, and, iii) network preemption can be applied to detect global states of user parallel processes. SCore-D gang scheduler realized by network preemption can utilize processor resources by the detecting the global state of user parallel processes. Network preemption with multiple contexts exhibits highly efficient gang scheduling. The combination of low scheduling overhead and the global state detection mechanism achieves an interactive parallel programming where parallel program development and the production run of parallel programs can be mixed freely. Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
SC | 3 |
| 1998 | Ninf and PM: Communication libraries for global computing and high-performance cluster computing
Mitsuhisa Sato, Hiroshi Tezuka, Atsushi Hori, Yutaka Ishikawa, Satoshi Sekiguchi, Hidemoto Nakada, Satoshi Matsuoka, Umpei Nagashima |
Future Gener. Comput. Syst. | 4 |
| 1997 | Global State Detection Using Network Preemption
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa |
JSSPP | 3 |
| 1996 | Implementation of Gang-Scheduling on Workstation Cluster
Atsushi Hori, Hiroshi Tezuka, Yutaka Ishikawa, Noriyuki Soda, Hiroki Konaka, Munenori Maeda |
JSSPP | 3 |
| 1995 | Time Space Sharing Scheduling: A Simulation Analysis
Atsushi Hori, Yutaka Ishikawa, Jörg Nolte, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo |
Euro-Par | 2 |
| 1995 | Time Space Sharing Scheduling and Architectural Support
Atsushi Hori, Takashi Yokota, Yutaka Ishikawa, Shuichi Sakai, Hiroki Konaka, Munenori Maeda, Takashi Tomokiyo, Jörg Nolte, Hiroshi Matsuoka, Kazuaki Okamoto, Hideo Hirono |
JSSPP | 3 |
| 1994 | Object Location Control Using Meta-level Programming
Hideaki Okamura, Yutaka Ishikawa |
ECOOP | 2 |
| 1992 | Communication Mechanism on Autonomous ObjectsabstractIn the concurrent object-oriented programming methodology, a system is described by concurrent objects which communicate with each others by various communication facilities, i.e., synchronous/asynchronous(future) message passing.Those facilities help up to implement application programs based on the client/server model.It is, however, difficult to describe application programs such that concurrent objects may simultaneously initiate communication with each other.Such objects are called autonomous objects.In this paper, we propose the notion of the visible and intensive sets, and a communication mechanism using those sets which enables us to handle communication among autonomous objects safely and easily. Yutaka Ishikawa |
OOPSLA | 1 |
| 1990 | Distributed Hartstone: A Distributed Real-Time Benchmark SuiteabstractAn extension of the uniprocessor Hartstone benchmark for the distributed real-time environment, called the Distributed Hartstone benchmark, is described. The Distributed Hartstone measures system performance in the critical areas of communication latency and bandwidth, protocol preemptability, and priority queueing at the protocol and media access levels. Areas of the system which are particularly important for distributed, real-time computing are described. On the basis of the requirements that specify various areas of the system that a distributed real-time benchmark must stress, a series of task sets in the style of the Hartstone benchmarks are given. The benchmark results from a distributed real-time operating system (ARTS testbed) are given.> Clifford W. Mercer, Yutaka Ishikawa, Hideyuki Tokuda |
ICDCS | 2 |
| 1989 | Priority Inversions in Real-Time CommunicationabstractThe priority-inversion problems in real-time communication are addressed, and solutions developed for the ARTS distributed real-time operating system are presented. The performance results of the multi-thread-based protocol implementation are compared with those of other implementation schemes, and the schedulability is analyzed. Experimental results indicate that the multi-thread-based protocol implementation could eliminate potential priority-inversion problems and also demonstrate the same schedulability as the softint implementation scheme in spite of about 10% additional implementation overhead.> Hideyuki Tokuda, Clifford W. Mercer, Yutaka Ishikawa, Thomas E. Marchok |
RTSS | 3 |
| 1986 | A Concurrent Object-Oriented Knowledge Representation Language Orient84/K: Its Features and ImplementationabstractOrient84/K is an object oriented concurrent programming language for describing knowledge systems. In Orient84/K, an object is composed of the behavior part, the knowledge-base part, and the monitor part, in order to provide object-oriented, logic-based, demon-oriented, and concurrent-programming paradigms in the object framework. Every object is capable of concurrent execution in Orient84/K. Yutaka Ishikawa, Mario Tokoro |
OOPSLA | 1 |
| 1984 | The Design of an Object Oriented ArchitectureabstractThis paper proposes a new object model, called the distributed object model, wherein the model is unified as a protection unit, as a method of data abstraction, and as a computational unit, so as to realize reliable, maintainable, and secure systems. An object oriented architecture called ZOOM is designed based on this object model. A software simulator and cross assembler for this architecture have been implemented. The feasibility and performance of the architecture are discussed according to program sizes and estimated hardware size and execution speed. Yutaka Ishikawa, Mario Tokoro |
ISCA | 1 |
| 1983 | Design of LSI speech spectrum analyzer using switched capacitor filter techniquesabstractThis paper presents a design approach to an LSI speech spectrum analyzer, which constructs one board speech recognition systems with other LSIs[4]. Particularly, the purpose of this LSI speech spectrum analyzer is to achieve high performances, such as a high resolution filter bank, powerful interface to CPU and variable spectrum integration intervals. In order to realize the above spectrum analyzer on one chip, switched capacitor filter (SCF) techniques are applied to filter bank synthesis. Design techniques to considerably reduce a SCF's area are introduced. A breadboard model constructed with discrete components shows good filter performances and high speech recognition rates. It is recognized through LSI design that this LSI can be realized on a 42 mm2chip with power dissipation of about 300 mW using the latest CMOS technology. Kenji Nakayama, Yutaka Ishikawa, Yoshiaki Kuraishi |
ICASSP | 2 |