EDBT 2026 Demo / reviewers in the wild / expert
Osamu Tatebe
dblp:51/3802
· DBLP profile ↗
41ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0003-4714-2164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-authorDatabases, data management, data science and information retrieval · 8Artificial intelligence and machine learning · 7Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SRAP: Sender-side Receiver-Aware Port Selection for Deterministic Flow-to-Core Mapping in Hash-Based Flow Steering
Shingo Hattori, Osamu Tatebe |
HPDC | 2 |
| 2026 | Lustre Query: Periodic Offline Metadata Monitoring from MDT BackupsabstractParallel file systems in HPC manage metadata for billions of files across distributed storage servers. The structure of this metadata, how files distribute by size, how users concentrate across servers, which storage policies are actually in use, determines operational decisions about capacity planning, load balancing, and data migration. Despite decades of HPC storage research, these structural properties remain underreported in the published literature for HPC parallel file systems. Runtime I/O behavior has been profiled extensively at the application level. Aggregate monitoring captures quotas and throughput. But the metadata that accumulates on disk, the artifact of all user activity over the life of a system, has received little systematic study. Existing tools can extract inode-level detail in principle, but each imposes barriers that discourage routine analysis: online queries load the metadata server, database replicas require ETL pipelines, and low-level utilities demand scripting that few administrators undertake. Sohei Koyama, Osamu Tatebe |
HPDC | 2 |
| 2024 | FINCHFS: Design of Ad-Hoc File System for I/O Heavy HPC WorkloadsabstractAlthough the performance improvements in parallel file systems have been significant, the rise of data science and deep learning using Python has introduced new I/O requirements. We redefine ad-hoc file systems as a complement to parallel file systems by specializing in requirements that cannot be met by improvements to parallel file systems. Ad-hoc file systems require extremely high metadata performance and the scalability to create and read large numbers of files in parallel in a single directory. Therefore, it is necessary to process a large number of small RPCs with low latency and high throughput, which cannot be achieved with existing ad-hoc file systems. In this research, we propose a new ad-hoc file system, FINCHFS, and realize the truly desired ad-hoc file system using intra-node shared-nothing architecture and number-aware hashing. The contribution of the intra-node shared-nothing architecture is the ability to run multiple iterative server processes on each node to achieve high IOPS using many-core. The contribution of number-aware hashing is to alleviate tail latency degradation due to the uneven distribution of requests to servers. In the evaluation using 64 nodes on the Pegasus supercomputer, we achieved 32.5 MIOPS in mdtest-hard-write. Sohei Koyama, Kohei Hiraga, Osamu Tatebe |
CLUSTER | 3 |
| 2024 | PEANUTS: A Persistent Memory-Based Network Unilateral Transfer System for Enhanced MPI-IO Data Transfer
Kohei Hiraga, Osamu Tatebe |
Euro-Par (2) | 2 |
| 2022 | CHFS: Parallel Consistent Hashing File System for Node-local Persistent MemoryabstractThis paper proposes a design for CHFS, an ad hoc parallel file system that utilizes the persistent memory of compute nodes. The design is based entirely on a highly scalable distributed key-value store with consistent hashing. CHFS improves the scalability of parallel data-access performance and metadata performance in terms of the number of compute nodes by eliminating dedicated metadata servers, sequential execution, and centralized data management. The implementation efficiently utilizes multicore and manycore CPUs, high-performance networks, and remote direct memory access by the Mochi-Margo library. With a 4-node persistent memory cluster, CHFS performs 9.9 times better than the state-of-the-art DAOS distributed object storage and 6.0 times better than GekkoFS on the IOR hard write benchmark. Regarding scalability, CHFS displays better scalability and performance for both bandwidth and metadata compared with BeeOND and GekkoFS. CHFS is a promising building block for HPC storage layers. Osamu Tatebe, Kazuki Obata, Kohei Hiraga, Hiroki Ohtsuji |
HPC Asia | 1 |
| 2020 | Gfarm/BB - Gfarm File System for Node-Local Burst Buffer
Osamu Tatebe, Shukuko Moriwake, Yoshihiro Oyama |
J. Comput. Sci. Technol. | 1 |
| 2020 | An Analysis of Concurrency Control Protocols for In-Memory Database with CCBenchabstractThis paper presents yet another concurrency control analysis platform, CCBench. CCBench supports seven protocols (Silo, TicToc, MOCC, Cicada, SI, SI with latch-free SSN, 2PL) and seven versatile optimization methods and enables the configuration of seven workload parameters. We analyzed the protocols and optimization methods using various workload parameters and a thread count of 224. Previous studies focused on thread scalability and did not explore the space analyzed here. We classified the optimization methods on the basis of three performance factors: CPU cache, delay on conflict, and version lifetime. Analyses using CCBench and 224 threads, produced six insights. The performance of optimistic concurrency control protocol for a read-only workload rapidly degrades as cardinality increases even without L3 cache misses. (I2) Silo can outperform TicToc for some write-intensive workloads by using invisible reads optimization. (I3) The effectiveness of two approaches to coping with conflict (wait and no-wait) depends on the situation. (I4) OCC reads the same record two or more times if a concurrent transaction interruption occurs, which can improve performance. (I5) Mixing different implementations is inappropriate for deep analysis. (I6) Even a state-of-the-art garbage collection method cannot improve the performance of multi-version protocols if there is a single long transaction mixed into the workload. On the basis of I4, we defined the read phase extension optimization in which an artificial delay is added to the read phase. On the basis of I6, we defined the aggressive garbage collection optimization in which even visible versions are collected. The code for CCBench and all the data in this paper are available online at GitHub. Takayuki Tanabe, Takashi Hoshino 0002, Hideyuki Kawashima, Osamu Tatebe |
Proc. VLDB Endow. | 4 |
| 2019 | In-situ Data Analysis System for High Resolution Meteorological Large Eddy Simulation ModelabstractIn meteorology, high-resolution simulations using large eddy simulation (LES) models have attracted strong attention. In these LES models, the output data become large, because the spatio-temporal resolution is higher than that of conventional meteorological simulations. Therefore, I/O requirements become much more demanding. In addition, it becomes challenging to analyze the results of these simulations, owing to a large data set size. To overcome these problems, it is necessary to consider hardware and software architectures for in-situ data analysis, for accommodating these I/O requirements. In this study, we propose a large-scale fluid analysis system combining a meteorological LES model and dynamic mode decomposition (DMD). We propose to use a burst buffer as a temporary storage space for storing large-scale output of LES models, and to read these data from DMD. The burst buffer is an all-flash storage system between compute nodes and a parallel file system. The I/O performance of the LES output and DMD input of this workflow is analyzed, and it is shown that the proposed burst buffer is an effective medium for in-situ data analysis. Takuto Sato, Osamu Tatebe, Hiroyuki Kusaka |
BDCAT | 2 |
| 2019 | Accelerating Machine Learning I/O by Overlapping Data Staging and Mini-batch GenerationsabstractThe training dataset used in deep neural networks (DNNs) keeps on increasing. When a training dataset grows larger, the reading performance of such a large training dataset becomes a problem. A high-performance computing (HPC) cluster has high performance I/O storage devices, for example, NVMe SSD, as local storage on each compute node. This high-performance I/O storage can mitigate the I/O bottleneck. However, such storage devices provide only temporary storage, therefore the users have to copy the training dataset from shared storage (such as Lustre) into local storage. Large datasets (over a few hundred GiB) takes a long time to copy the datasets between local storage and shared storage. To solve this problem, we propose a method to conceal the time spent on copying dataset to local storage by overlapping the copying and reading of training data. We implemented the proposed method at the machine learning framework Chainer. The results of our experiments showed that the read I/O bandwidth of our method improved from 1.38 times to 6.19 times compared with reading the dataset from Lustre directly using Chainer standard class. Moreover, evaluation of data parallel training showed that our method improved the performance from 1.26 times to 1.74 times for the same comparison. Kazuhiro Serizawa, Osamu Tatebe |
BDCAT | 2 |
| 2019 | GHOSTZ PW/GF: Distributed Parallel Homology Search System for Large-scale Metagenomic AnalysisabstractGenomic data, a typical example of big data, is increasing annually owing to the extensive usage and improved performance of next-generation sequencers. Homology search, which is used for genome analysis, is a technique in which homologous DNA from a known DNA sequence database is collected using unknown DNA sequences as a query. However, the conventional method suffers from performance issues and memory shortage. In this study, we proposed and implemented a distributed parallel homology search system GHOSTZ PW/GF using Gfarm, a distributed file system, and Pwrake, a dynamic workflow engine and evaluated them in TSUBAME3.0. The results indicate the high scalability of the proposed system; additionally, using a prebuilt non-redundant database comprising approximately 100 million records and sequence data comprising approximately 500 million records, the proposed system completed the execution of all processes on 180 nodes in approximately 2 h. To the best of the authors' knowledge, this is the first large-scale metagenomic analysis attempted using such a large number of queries and DBs. Kenta Machida, Osamu Tatebe |
IEEE BigData | 2 |
| 2019 | Accelerating Sequence Operator with Reduced ExpressionabstractSequence operators are effective for efficiently combining multiple events when state recognition is performed by combining time series events. Since sensor data are inherently noisy, one can take a strict attitude to deal with them: it is conceivable that all of time series events are regarded as false positives. Then, all complex events should be constructed carefully. Such an attitude is called the skip-till-any-match model in the sequence operator. When using this model, huge amounts of potential complex events are generated. A sequence operator usually supports both Kleene closure and non-Kleene closure. While efficient methods have been studied for Kleene closure so far, that for non-Kleene closure have been still explored. In this paper, we propose the reduced expression method to improve the efficiency of sequence operator processing for the skip-till-any-match model. Experimental results showed that the processing time and memory size were more efficient compared with SASE, which is the conventional method, and that degree is up to several thousand times. Hideyuki Kawashima, Osamu Tatebe |
EJC | 2 |
| 2018 | Skew-Aware Collective Communication for MapReduce ShufflingabstractThis paper proposes and examines the three in-memory shuffling methods designed to address problems in MapReduce shuffling caused by skewed data. Coupled Shuffle Architecture (CSA) employs a single pairwise all-to-all exchange to shuffle both blocks, units of shuffle transfer, and meta-blocks, which contain the metadata of corresponding blocks. Decoupled Shuffle Architecture (DSA) separates the shuffling of meta-blocks and blocks, and applies different all-to-all exchange algorithms to each shuffling process, attempting to mitigate the impact of stragglers in strongly skewed distributions. Decoupled Shuffle Architecture with Skew-Aware Meta-Shuffle (DSA w/ SMS) autonomously determines the proper placement of blocks based on the memory consumption of each worker process. This approach targets extremely skewed situations where some worker processes could exceed their node memory limitation. This study evaluates implementations of the three shuffling methods in our prototype in-memory MapReduce engine, which employs high performance interconnects such as InfiniBand and Intel Omni-Path. Our results suggest that DSA w/ SMS is the only viable solution for extremely skewed data distributions, but this solution is only valid on systems equipped with high performance interconnects. We also present a detailed investigation of the performance of CSA and DSA in various skew situations. Harunobu Daikoku, Hideyuki Kawashima, Osamu Tatebe |
IEEE BigData | 3 |
| 2018 | Applying Pwrake Workflow System and Gfarm File System to Telescope Data ProcessingabstractIn this paper, we describe a use case applying a scientific workflow system and a distributed file system to improve the performance of telescope data processing. The application is pipeline processing of data generated by Hyper Suprime-Cam (HSC) which is a focal plane camera mounted on the Subaru telescope. In this paper, we focus on the scalability of parallel I/O and core utilization. The IBM Spectrum Scale (GPFS) used for actual operation has a limit on scalability due to the configuration using storage servers. Therefore, we introduce the Gfarm file system which uses the storage of the worker node for parallel I/O performance. To improve core utilization, we introduce the Pwrake workflow system instead of the parallel processing framework developed for the HSC pipeline. Descriptions of task dependencies are necessary to further improve core utilization by overlapping different types of tasks. We discuss the usefulness of the workflow description language with the function of scripting language for defining complex task dependency. In the experiment, the performance of the pipeline is evaluated using a quarter of the observation data per night (input files: 80 GB, output files: 1.2 TB). Measurements on strong scaling from 48 to 576 cores show that the processing with Gfarm file system is more scalable than that with GPFS. Measurement using 576 cores shows that our method improves the processing speed of the pipeline by 2.2 times compared with the method used in actual operation. Masahiro Tanaka, Osamu Tatebe, Hideyuki Kawashima |
CLUSTER | 2 |
| 2018 | Performing External Join Operator on PostgreSQL with Data Transfer ApproachabstractWith the development of sensing devices, the size of data managed by human being has been rapidly increasing. To manage such huge data, relational database management system (RDBMS) plays a key role. RDBMS models the real world data as n-ary relational tables. Join operator is one of the most important relational operators, and its acceleration has been studied widely and deeply. How can an RDBMS provide such an efficient join operator? The performance improvement of join operator has been deeply studied for a decade, and many techniques are proposed already. The problem that we face is how to actually use such excellent techniques in real RDBMSs. We propose to implement an efficient join technique by the data transfer approach. The approach makes a hook point inside an RDBMS internal, and pulls data streams from the operator pipeline in the RDBMS, and applies our original join operator to the data, and finally returns the result to the operator pipeline in the RDBMS. The result of the experiment showed that our proposed method achieved 1.42x speedup compared with PostgreSQL. Our code is available on GitHub. Ryota Takizawa, Hideyuki Kawashima, Ryuya Mitsuhashi, Osamu Tatebe |
HPC Asia | 4 |
| 2018 | Integration of Parallel Write Ahead Logging and Cicada Concurrency Control MethodabstractWe proposed the idea of combining the Cicada concurrency control mechanism with P-WAL (parallel write-ahead logging) and experimentally evaluated the proposed control mechanism. Our results show that P-WAL can be combined suitably with Cicada without decreasing Cicada's performance. We then experimentally evaluated the most optimal parallel write-ahead logging parameter among (1) conservative lock release; (2) early lock release, the early release of locks on database objects; and (3) group commit, which writes the transaction log to storage all at once. The results show that early lock release and group commit were the most optimal parameters for the combination of Cicada and P-WAL. Takayuki Tanabe, Hideyuki Kawashima, Osamu Tatebe |
SMARTCOMP | 3 |
| 2016 | Three-dimensional spatial join count exploiting CPU optimized STR R-treeabstractIn this study, we attempt to address the issue regarding the spatial join count, where in the number of particles around a halo is counted only once for a given simulation result. An efficient spatial index is necessary for accelerated counting; therefore, we propose a CPU optimized sort-tile-recursive R-tree that employs a parallel radix sort and node packing with thread pool and single instruction multiple data instructions. In an experiment conducted with astronomical data, the proposed method demonstrates an improvement in performance by 26.8 times compared with that using a conventional CPU optimized R-tree. We also propose a partial materialization approach to handle large amount of data that exceeds the capacity of main memory. To accelerate the approach, we propose a construct-search-destruct pipeline that exploits a thread pool to conceal the latency of the construction and destruction of the index. The pipelining method achieves an improvement in performance by 27.5 times compared with that of a conventional CPU optimized R-tree. All our codes are available on GitHub. Ryuya Mitsuhashi, Hideyuki Kawashima, Takahiro Nishimichi, Osamu Tatebe |
IEEE BigData | 4 |
| 2016 | Fast window aggregate on array database by recursive incremental computationabstractAn array database is effective for managing and analyzing multidimensional scientific big data, and the window aggregate is an important operator in array databases. This paper proposes a method that exploits the scheme of incremental computation to accelerate the execution of window aggregates considerably. Six types of aggregate are improved using different designs of buffer tools to eliminate redundant computation. Our proposed recursive incremental computation method completely eliminates all redundant computation and achieves an improvement factor of the total window size compared with the naive method. This proposed method is fully implemented in SciDB. It improved performance by a factor of 10 on an earth science benchmark and by a factor of 64 on synthetic workloads with a certain data setting when compared with SciDB's built-in window operator. Hideyuki Kawashima, Osamu Tatebe |
eScience | 3 |
| 2016 | Experimental analysis of operating system jitter caused by page reclaim
Yoshihiro Oyama, Shun Ishiguro, Jun Murakami, Shin Sasaki, Ryo Matsumiya, Osamu Tatebe |
J. Supercomput. | 6 |
| 2015 | Server-Side Efficient Parity Generation for Cluster-Wide RAID SystemabstractExascale computing requires storage systems to provide both reliability and performance, which are in the relationship of trade-off. This paper presents a method to build a zero-overhead network storage system (Cluster-wide RAID), which consists of active storage nodes. Active storage nodes themselves have an ability to process data blocks and our proposed method utilizes them to generate parity blocks. In other words, our proposed method moves the parity calculation process to server-side from client-side. Our implementation of the proposed method utilizes InfiniBand remote direct memory access (RDMA), which shows the significant performance improement of RAID-5 write. Especially, the performance of RAID-5 is almost same as RAID-0 (no redundancy). It means that the proposed method can add redundancy to network storage systems without performance degradation. Hiroki Ohtsuji, Osamu Tatebe |
CloudCom | 2 |
| 2015 | RDMA-Based Direct Transfer of File Data to Remote Page CacheabstractThe performance of a distributed file system significantly affects data-intensive applications that frequently execute I/O operations on large amounts of data. Although many modern distributed file systems are geared to provide highly efficient I/O performance, their operations are nonetheless affected by runtime overhead in data transfer between client nodes and I/O servers. A large part of the overhead is caused by memory copies executed by the client interface using the FUSE framework or a special kernel module. In this paper, we propose a method based on InfiniBand RDMA that improves data transfer performance between client and server in a distributed file system. The major characteristic of the method is that it transfers file data directly from a server's memory to the page cache of a client node. The method minimizes memory copies that are otherwise executed in the client interface or the operating system kernel. We implemented the proposed method in the Gfarm distributed file system and tested it using I/O benchmark software and real applications. The experimental results showed that our method effected a performance improvement of up to 78.4% and 256.0% in sequential and random file reads, respectively, and an improvement of up to 6.3% in data-intensive applications. Shin Sasaki, Kazushi Takahashi, Yoshihiro Oyama, Osamu Tatebe |
CLUSTER | 4 |
| 2015 | RDMA-Based Cooperative Caching for a Distributed File SystemabstractEfficient caching of file data is critical in order to achieve high performance in data-intensive applications. However, only a limited amount of memory is usually available to cache files in client nodes even on high-performance computing platforms. Cooperative caching is an approach that enables client nodes to share memory for file caching and thereby provide a large amount of memory for the file cache in the aggregate. Many studies have confirmed the efficacy of applying cooperative caching to distributed file systems. However, to the best of our knowledge, no study has evaluated an implementation of cooperative caching integrated into a modern distributed file system running on a high-speed network. In this paper, we propose a method that improves the performance of a distributed file system oriented to high-performance computing by integrating cooperative caching into it. In the proposed method, the metadata server of the distributed file system maintains information about the cache in all client nodes, and provides clients with the predicted cache location of any requested file. Further, InfiniBand RDMA is utilized to achieve fast cache transfer between the page caches of client nodes. Implementation of the proposed method in the Gfarm distributed file system and measurement of the performance of three real-world data-intensive applications indicate that the proposed method achieves a maximum speedup of 5.8%. Shin Sasaki, Ryo Matsumiya, Kazushi Takahashi, Yoshihiro Oyama, Osamu Tatebe |
ICPADS | 5 |
| 2014 | Incremental window aggregates over array databaseabstractWe propose an efficient window aggregation method over multi-dimensional array data based on incremental computation. We improve several aggregations with different data structures exploited to achieve efficient computation: list for sum and avg, heap for max and min, and balanced binary search tree for percentile. We present time complexity analysis for the methods, and then evaluate performance with experiments in SciDB array database system with both synthetic and JRA55 meteorological dataset. Our analysis shows that performance improvement is proportional to the window size in the last dimension in theory, and the result of experiment is consistent with the analysis. In certain cases, it shows an acceleration factor more than 13 by the proposed method with percentile, while a factor over 28 with maximum. Hideyuki Kawashima, Osamu Tatebe |
IEEE BigData | 3 |
| 2014 | Preliminary evaluation of optimized transfer method for cluster-wide RAID-4abstractExa-scale storage systems will require a reliable and high bandwidth I/O mechanism. Replication and erasure coding among storage nodes are often used to improve reliability; however, the replication has a problem in storage capacity, and the erasure coding among storage nodes requires additional calculation cost and network traffic that may reduce the I/O performance. This paper proposes method to reduce client-side calculation cost and network traffic, and describes the zero-copy implementation of the cluster-wide RAID system using the InfiniBand RDMA mechanism. The evaluation result shows 34% improvement of I/O bandwidth. The performance is almost same as no parity mode. It means that the proposed mechanism is a breakthrough technology to build almost zero overhead cluster-wide RAID systems. Hiroki Ohtsuji, Osamu Tatebe |
CLUSTER | 2 |
| 2014 | Disk cache-aware task scheduling for data-intensive and many-task workflowabstractWorkflow scheduling to maximize I/O performance is one of the key issues in data-intensive, many-task computing. In our previous work, we proposed locality-aware workflow scheduling method using the Multi-Constraint Graph Partitioning. In this work, we focus on read performance of input files from the disk cache (buffer cache or page cache on main memory). In order to maximize the disk cache hit rate of input files, a LIFO-order scheduling is effective since created intermediate files may be read soon. However, LIFO policy has a disadvantage of so-called “trailing task problem.” We propose a hybrid scheduling strategy of LIFO and HRF (Highest Rank First). In our strategy, one of two policies is applied depending on the number of highest-rank tasks in the queue to avoid the problem. In addition, scheduling for the overlap of computation and I/O is proposed. We implement our scheduling strategy for the Pwrake workflow system and the Gfarm distributed file system and evaluate it by executing data-intensive workflows using a computer cluster. Our scheduling strategy improves the performance of copyfile workflow by 30% due to increase in disk cache hit rate, and the performance of Montage workflow by 12% due to increase in core utilization. Masahiro Tanaka, Osamu Tatebe |
CLUSTER | 2 |
| 2012 | Workflow Scheduling to Minimize Data Movement Using Multi-constraint Graph PartitioningabstractAmong scheduling algorithms of scientific workflows, the graph partitioning is a technique to minimize data transfer between nodes or clusters. However, when the graph partitioning is simply applied to a complex workflow DAG, tasks in each parallel phase are not always evenly assigned to computation nodes since the graph partitioning algorithm is not aware of edge directions that represent task dependencies. Thus, we propose a new method of task assignment based on Multi-Constraint Graph Partitioning. This method relates the dimension of weight vectors to the rank of a task phase defined by traversing the task graph. Our algorithm is implemented in the Pwrake workflow system and evaluated the performance of the Montage workflow using a computer cluster. The result shows that the file size accessed from remote nodes is reduced from 88% to 14% of the total file size accessed during the workflow and that the elapsed time is reduced by 31%. Masahiro Tanaka, Osamu Tatebe |
CCGRID | 2 |
| 2011 | MPI-IO/Gfarm: An Optimized Implementation of MPI-IO for the Gfarm File SystemabstractThis paper proposes a design and implementation of an MPI-IO implementation of the Gfarm file system, called MPI-IO/Gfarm. The Gfarm file system is a global file system that federates the local storage of compute nodes among several clusters. It has a scale-out architecture designed to support distributed data-intensive computing. However Gfarm file system does not achieve scalable performance in the case of parallel writes to a single file, a typical file operation in MPI-IO. This paper proposes an optimization technique to improve the parallel write performance to a single file. In the evaluation, MPI-IO/Gfarm achieves scalable parallel I/O performance. Hiroki Kimura, Osamu Tatebe |
CCGRID | 2 |
| 2010 | Pwrake: a parallel and distributed flexible workflow management tool for wide-area data intensive computingabstractThis paper proposes Pwrake, a parallel and distributed flexible workflow management tool based on Rake, a domain specific language for building applications in the Ruby programming language. Rake is a similar tool to make and ant. It uses a Rakefile that is equivalent to a Makefile in make, but written in Ruby. Due to a flexible and extensible language feature, Rake would be a powerful workflow management language. The Pwrake extends Rake to manage distributed and parallel workflow executions that include remote job submission and management of parallel executions. This paper discusses the design and implementation of the Pwrake, and demonstrates its power of language and extensibility of the system using a practical e-Science data-intensive workflow in astronomical data analysis on the Gfarm file system as a case study. Extending a scheduling algorithm to be aware of file locations, 20% of speed up is observed using 8 nodes (32 cores) in a PC cluster. Using two PC clusters located in different institutions, the file location aware scheduling shows scalable speedup. The extensible Pwrake is a promising workflow management tool even for wide-area data analysis. Masahiro Tanaka, Osamu Tatebe |
HPDC | 2 |
| 2008 | Implement the Grid Workflow Scheduling for Data Intensive Applications with CSF4abstractGrid computing technology is able to integrate and share large-scale distributed computation and data resource to facilitate the scientific researches. Recently, the grid workflow support and large-scale distributed data management are becoming two main requirements of scientists and researchers in many fields, such as bioinformatics, high-energy physics etc. In this paper, we proposed to support grid workflow for data intensive applications using CSF4 scheduling plug-ins. The grid workflow scheduling and data aware scheduling policies are implemented in two scheduling plug-ins, grid workflow plug-in and grid data aware plug-in, respectively. The two scheduling plug-ins can work together smoothly. The data aware plug-in will automatically dispatch the workflow tasks to the grid sites which are close to data replicas. At last, the experiment results are given to show the improvement of system performance and optimization of scheduling. Zhaohui Ding, Xiaohui Wei 0002, Yaoguang Yuan, Wilfred W. Li, Osamu Tatebe |
eScience | 6 |
| 2008 | Building Hierarchical Grid Storage Using the GfarmGlobal File System and the JuxMemGrid Data-Sharing Service
Gabriel Antoniu, Loïc Cudennec, Majd Ghareeb, Osamu Tatebe |
Euro-Par | 4 |
| 2008 | Performance Evaluation of Data Management Layer by Data Sharing Patterns for Grid RPC Applications
Yoshihiro Nakajima, Yoshiaki Aida, Mitsuhisa Sato, Osamu Tatebe |
Euro-Par | 4 |
| 2006 | Building Cyberinfrastructure for Bioinformatics Using Service Oriented Architecture
Wilfred W. Li, Sriram Krishnan, Kurt Mueller, Kohei Ichikawa, Susumu Date, Sargis Dallakyan, Michel F. Sanner, Chris Misleh, Zhaohui Ding, Xiaohui Wei 0002, Osamu Tatebe, Peter W. Arzberger |
CCGRID | 11 |
| 2006 | The PRAGMA Testbed - Building a Multi-Application International Grid
Cindy Zheng, David Abramson 0001, Peter W. Arzberger, Shahaan Ayyub, Colin Enticott, Slavisa Garic, Mason J. Katz, Jae-Hyuck Kwak, Bu-Sung Lee, Philip M. Papadopoulos, Sugree Phatanapherom, Somsak Sriprayoonsakul, Yoshio Tanaka, Yusuke Tanimura, Osamu Tatebe, Putchong Uthayopas |
CCGRID | 15 |
| 2006 | Storage challenge - High performance data analysis for particle physics using the Gfarm file systemabstractThe Belle experiment operates at the KEKB accelerator, a high luminosity asymmetric energy e+ e- collider. The Belle collaboration studies CP violations in decays of B mesons to answer one of the fundamental questions of Nature, the matter-anti-matter asymmetry. Currently, Belle accumulates more than one million B Bbar meson pairs, corresponding to about 1.2 TB of raw data, per day.The challenge is how to realize the required high performance data access and scalable data computing. The Gfarm file system is a Grid-wide network shared file system that federates local storage of cluster nodes; moreover it provides scalable I/O performance with distributed data access. In the challenge, we will construct a Gfarm file system with 40 TB capacity and 30 GB/sec I/O bandwidth, integrating the local disks of 800 compute nodes in the KEKB computing facility, and demonstrate high-performance Belle data analysis. Nobuhiko Katayama, Mitsuhisa Sato, Taisuke Boku, Akira Ukawa, Shohei Nishida, Ichiro Adachi, Osamu Tatebe |
SC | 7 |
| 2005 | Integrating Local Job Scheduler - LSFTM with GfarmTM
Xiaohui Wei 0002, Wilfred W. Li, Osamu Tatebe, Gaochao Xu, Liang Hu 0001, Jiubin Ju |
ISPA | 3 |
| 2004 | GNET-1: gigabit Ethernet network testbedabstractGNET-1 is a fully programmable network testbed. It provides functions such as wide area network emulation, network instrumentation, traffic shaping, and traffic generation at gigabit Ethernet wire speeds by programming the core FPGA. GNET-1 is a powerful tool for developing network-aware grid software. It is also a network monitoring and traffic-shaping tool that provides high-performance communication over wide area networks. This work describes several sample uses of GNET-1 and presents its architecture. Yuetsu Kodama, Tomohiro Kudoh, Ryousei Takano, Hitoshi Sato, Osamu Tatebe, Satoshi Sekiguchi |
CLUSTER | 5 |
| 2003 | Design and implementation of PVFS-PM: a cluster file system on SCoreabstractThis paper discusses the design and implementation of a cluster file system, called PVFS-PM, on the SCore cluster system software. This is the first attempt to implement a cluster file system on the SCore system. It is based on the PVFS cluster file system but replaces TCP with the PMv2 communication library supported by SCore to provide a scalable, high-performance cluster file system. PVFS-PM improves the performance by factors of 1.07 and 1.93 for writing and reading, respectively, with 8 I/O nodes, compared with the original PVFS on TCP on a Gigabit Ethernet-connected SCore cluster. Koji Segawa, Osamu Tatebe, Yuetsu Kodama, Tomohiro Kudoh, Toshiyuki Shimizu |
CCGRID | 2 |
| 2003 | Performance Analysis of Scheduling and Replication Algorithms on Grid Datafarm Architecture for High-Energy Physics ApplicationsabstractData Grid is a Grid for ubiquitous access and analysis of large-scale data. Because Data Grid is in the early stages of development, the performance of its petabyte-scale models in a realistic data processing setting has not been well investigated. By enhancing our Bricks Grid simulator to accommodated Data Grid scenarios, we investigate and compare the performance of different Data Grid models. These are categorized mainly as either central or tier models; they employ various scheduling and replication strategies under realistic assumptions of job processing for CERN LHC experiments on the Grid Datafarm system. Our results show that the central model is efficient but that the tier model, with its greater resources and its speculative class of background replication policies, are quite effective and achieve higher performance, while each tier is smaller than the central model. Atsuko Takefusa, Osamu Tatebe, Satoshi Matsuoka, Youhei Morita |
HPDC | 2 |
| 2002 | Grid Datafarm Architecture for Petascale Data Intensive ComputingabstractThe Grid Datafarm (Gfarm) architecture is designed for global petascale data-intensive computing. It provides a global parallel filesystem with online petascale storage, scalable I/O bandwidth, and scalable parallel processing, and it can exploit local I/O in a grid of clusters with tens of thousands of nodes. Gfarm parallel I/O APIs and commands provide a single filesystem image and manipulate filesystem metadata consistently. Fault tolerance and load balancing are automatically managed by file duplication or recomputation using a command history log. Preliminary performance evaluation has shown scalable disk I/O and network bandwidth on 64 nodes of the Presto III Athlon cluster. The Gfarm parallel I/O write and read operations has achieved data transfer rates of 1.74 GB/s and 1.97 GB/s, respectively, using 64 cluster nodes. The Gfarm parallel file copy reached 443 MB/s with 23 parallel streams on the Myrinet 2000. The Gfarm architecture is expected to enable petascale data-intensive Grid computing with an I/O bandwidth scales to the TB/s range and scalable computational power. Osamu Tatebe, Youhei Morita, Satoshi Matsuoka, Noriyuki Soda, Satoshi Sekiguchi |
CCGRID | 1 |
| 2001 | Design and implementation of FMPL, a fast message-passing library for remote memory operationsabstractA fast message-passing library FMPL has been designed and developed to maximize communication performance by utilizing general architectural communication support such as remote memory operations, as well as to maximize total performance by eliminating dynamic communication overhead and overlapping communication and computation. FMPL provides a low-cost general-purpose point-to-point communication and collective communication such as broadcast, barrier synchronization and reduction. On a Hitachi SR8000, FMPL achieves an 8-byte latency of 12.8μsec., while MPI achieves 20μsec. FMPL is designed for building more highly functional message-passing libraries like BLACS as well as applications that need maximum performance. Osamu Tatebe, Umpei Nagashima, Satoshi Sekiguchi, Hisayoshi Kitabayashi, Yoshiyuki Hayashida |
SC | 1 |
| 1998 | Highly Efficient Implementation of MPI Point-to-Point Communication Using Remote Memory OperationsabstractMPl point-to-point communication is a basic operation, however it requires runtime-matching of send and receive that causes to reduce performance.This paper proposes a new approach to send messages by remote memory write without inquiring of the receiver under a communication pattern such that nonblocking receive is issued in advance.Basically, this approach makes it possible to gain low latency and high bandwidth as the hardware specification.MPI-EMX, our implementation of the MPI on the EM-X multiprocessor, achieves a zero-byte latency of 13.4 psec.and a maximum bandwidth of 31.4MB/s, which can compete with commercial MPPs.This approach to reduce communication latency is widely applicable to other systems and is quite a promising technique for achieving low latency and high bandwidth. Osamu Tatebe, Yuetsu Kodama, Satoshi Sekiguchi, Yoshinori Yamaguchi |
International Conference on Supercomputing | 1 |
| 1994 | Efficient implementation of the multigrid preconditioned conjugate gradient method on distributed memory machinesabstractA multigrid preconditioned conjugate gradient (MGCG) method, which uses the multigrid method as a preconditioner for the conjugate gradient method, has a good convergence rate even for problems on which the standard multigrid method does not converge efficiently. This paper considers a parallelization of the MGCG method and proposes an efficient parallel MGCG method on distributed memory machines. For a good convergence rate of the MGCG method, several difficulties in parallelizing the multigrid method are successfully settled. It is also shown that the parallel MGCG method has high performance on the Fujitsu AP1000 multicomputer, and it is more than 10 times faster than the scaled conjugate gradient (SCG) method.> Osamu Tatebe, Yoshio Oyanagi |
SC | 1 |