EDBT 2026 Demo / reviewers in the wild / expert
Yusuke Tanimura
dblp:12/5779
· DBLP profile ↗
20ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLaTEC: An efficient federated learning scheme across the Thing-Edge-Cloud environment
Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 3 |
| 2024 | SFETEC: Split-FEderated Learning Scheme Optimized for Thing-Edge-Cloud EnvironmentabstractThis paper introduces SFETEC, an innovative federated learning framework addressing the limitations of traditional methods like FedAvg. SFETEC splits training into base models and core models, reducing communication overhead, mitigating non-IID data issues, and enhancing training speed. Base models are trained at client devices, while core models are trained at edge servers and aggregated at the cloud. Preliminary results show that SFETEC significantly reduces communication overhead and training duration compared to the state-of-the-art baselines while enhancing privacy. Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
e-Science | 3 |
| 2024 | Poster: Implementing Data Reduction at the Middle Point on the Computing ContinuumabstractIn the computing continuum paradigm, data is expected to be properly managed from edge to cloud for advanced digital services, in terms of response time, data movement, security including privacy, etc. Then we proposed edge-side common data processing services on the computing continuum, for reducing an amount of data transfer between edge and cloud, in wider such applications. This poster presents our work-in-progress implementation of our proposed depot service by using the Valkey high-performance key-value data store, with an example data compression module. Our initial experiment with an example data compression module shows a negligible overhead of the current prototype and potential use of efficient data collection. Yusuke Tanimura |
SEC | 1 |
| 2023 | Taming Metadata-intensive HPC Jobs Through Dynamic, Application-agnostic QoS ControlabstractModern I/O applications that run on HPC infrastructures are increasingly becoming read and metadata intensive. However, having multiple applications submitting large amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to overall performance degradation and I/O unfairness. We present PADLL, an application and file system agnostic storage middleware that enables QoS control of data and metadata workflows in HPC storage systems. It adopts ideas from Software-Defined Storage, building data plane stages that mediate and rate limit POSIX requests submitted to the shared file system, and a control plane that holistically coordinates how all I/O workflows are handled. We demonstrate its performance and feasibility under multiple QoS policies using synthetic benchmarks, real-world applications, and traces collected from a production file system. Results show that PADLL can enforce complex storage QoS policies over concurrent metadata-aggressive jobs, ensuring fairness and prioritization. Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, José Pereira 0001, João Paulo 0001 |
CCGrid | 3 |
| 2022 | Protecting Metadata Servers From Harm Through Application-level I/O ControlabstractModern large-scale I/O applications that run on HPC infrastructures are increasingly becoming metadata-intensive. Unfortunately, having multiple concurrent applications submitting massive amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to unresponsiveness of the storage backend and overall performance degradation. To address these challenges, we present Padll, a storage middleware that enables system administrators to proactively control and ensure QoS over metadata workflows in HPC storage systems. We demonstrate its performance and feasibility by controlling the rate of both synthetic and realistic I/O workloads. Results show that Padll can dynamically control metadata-aggressive workloads, prevent I/O burstiness, and ensure I/O fairness and prioritization. Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, João Paulo 0001 |
CLUSTER | 3 |
| 2022 | PAIO: General, Portable I/O Optimizations With Minor Application Modifications
Ricardo Macedo, Yusuke Tanimura, Jason H. Haga, Vijay Chidambaram, José Pereira 0001, João Paulo 0001 |
FAST | 2 |
| 2021 | The Case for Storage Optimization Decoupling in Deep Learning FrameworksabstractDeep Learning (DL) training requires efficient access to large collections of data, leading DL frameworks to implement individual I/O optimizations to take full advantage of storage performance. However, these optimizations are intrinsic to each framework, limiting their applicability and portability across DL solutions, while making them inefficient for scenarios where multiple applications compete for shared storage resources.We argue that storage optimizations should be decoupled from DL frameworks and moved to a dedicated storage layer. To achieve this, we propose a new Software-Defined Storage architecture for accelerating DL training performance. The data plane implements self-contained, generally applicable I/O optimizations, while the control plane dynamically adapts them to cope with workload variations and multi-tenant environments.We validate the applicability and portability of our approach by developing and integrating an early prototype with the TensorFlow and PyTorch frameworks. Results show that our I/O optimizations significantly reduce DL training time by up to 54% and 63% for TensorFlow and PyTorch baseline configurations, while providing similar performance benefits to framework-intrinsic I/O mechanisms provided by TensorFlow. Ricardo Macedo, Cláudia Correia, Marco Dantas, Cláudia Brito, Weijia Xu, Yusuke Tanimura, Jason H. Haga, João Paulo 0001 |
CLUSTER | 6 |
| 2020 | Building and Evaluation of Cloud Storage and Datasets Services on AI and HPC Converged InfrastructureabstractAI Bridging Cloud Infrastructure (ABCI) is a world-leading open AI computing infrastructure, for accelerating R&D activities of artificial intelligence. In order to share and reuse AI software assets with ease, ABCI supports container-based application deployment and fine-grained resource allocation on top of the conventional HPC architecture, and provides tens of peta-bytes of high performance storage. One of the on-going major challenges in ABCI, however, is to more efficiently and flexibly exchange and share machine learning models and data related to AI, with other services deployed outside of ABCI in the real world. Our new services called as ABCI Cloud Storage and ABCI Public Datasets are designed for tackling the challenge and taking a role of "Data Harbor" of ABCI. The services allow users to store input and output data of jobs to be run on the ABCI compute nodes, and to share them with not only ABCI users but also non-ABCI users. This paper presents our design and integration of the services to conventional HPC architecture, as a case of ABCI, and reports performance evaluation of them. Based on our attempt and experience, the paper finally summarizes discussion about future direction of the S3 based front data/storage service of the AI and HPC converged system. Yusuke Tanimura, Shin'ichiro Takizawa, Hirotaka Ogawa, Takahiro Hamanishi |
IEEE BigData | 1 |
| 2017 | Understanding and improving disk-based intermediate data caching in SparkabstractApache Spark is a parallel data processing framework that executes fast for iterative calculations and interactive processing, by caching intermediate data in memory with a lineage-based data recovery from faults. The Spark system can also manage data sets larger than memory capacity by placing some cache or all of them on disks on processing nodes. However, the disadvantage is potential performance degradation due to disk I/O and/or serialization. This study aims to clarify efficient/inefficient use of disks in intermediate data caching in Spark and also to improve the usability of disks for end users. In order to achieve the purpose, influence of disk use in data caching was firstly investigated in various aspects, such as caching options, data abstractions and storage devices. The results indicate that serialization cost is dominant rather than disk I/O in most cases. Secondly, a method of combined use of memory and disk was further evaluated under a high memory pressure. Then the method was improved to avoid an excessive re-caching problem, which achieved at most 20-30% reduction of total execution time under a high memory pressure and did not degrade the performance under a low memory pressure, in our experiment with 4 machine learning benchmarks. Finally, this paper summarizes important factors and potential improvements for efficiently using disks in data caching in Spark. Kaihui Zhang, Yusuke Tanimura, Hidemoto Nakada, Hirotaka Ogawa |
IEEE BigData | 2 |
| 2014 | A High Performance, QoS-Enabled, S3-Based Object StoreabstractA scale-out and reliable object store is an important building block of a cloud service, for storing virtual machine images, backups and large application data. As such an object store, Amazon S3 is available for Amazon EC2 users and other S3-compatible storage systems are also used in private clouds. However, there is concern about performance instability when many applications concurrently access the storage service, due to the characteristics of shared use. This paper presents an approach to introducing a QoS-enabled function into the S3-based object store. The object store accepts an explicit performance request as an advanced reservation, and enables QoS in the access with the extended S3 Restful interface. Implicit and static performance setting is also possible for the unmodified S3 interface. Papio S3, an object store which supports both of these S3 interfaces, is developed for implementing the approach, along with achievement of high performance upload/download using the multipart data transfer. The evaluation confirms the performance of Papio S3, and its QoS capability at the S3 data transfer, in several situations where multiple S3 clients concurrently access the same Papio S3 system. In addition, the QoS effect is compared with a load balancing approach in an existing object store, as part of the experiment. Yusuke Tanimura, Seiya Yanagita, Takahiro Hamanishi |
CCGRID | 1 |
| 2014 | Applying Selectively Parallel I/O Compression to Parallel Storage Systems
Rosa Filgueira, Malcolm P. Atkinson 0001, Yusuke Tanimura, Isao Kojima |
Euro-Par | 3 |
| 2013 | A knowledge-based support method for autonomous service operations after disastersabstractAfter the 2011 earthquake off the Pacific coast of Tohoku, the importance of network services, like IP phone and e-mail, as a mean of communication in an emergency was hugely increased, but are likely to be discontinued in these situations. If that happens, network administrators have to repair the network and restart the services promptly. It is desirable that novice administrators also take part in network recovery operations, because expert administrators are not always stationed all day long. In this paper, we propose a knowledge-based support method for autonomous service operations in emergency situations. We use the Active Information Resource based Network Management System (AIR-NMS) to reduce the burden on administrators and to enable even novice administrators to operate network services. Finally, we show the effectiveness of the proposed method through experiments using a prototype system. Yusuke Tanimura, Johan Sveholm, Kazuto Sasai, Gen Kitagata, Tetsuo Kinoshita |
ICIS | 1 |
| 2013 | MPI collective I/O based on advanced reservations to obtain performance guarantees from shared storage systemsabstractAs more data-intensive computing applications are executed on high performance computing clusters, resource contention on the shared storage system attached to the clusters becomes significant. The contention might cause I/O performance degradation and spoil performance improvement of coordinated parallel I/O by the MPI-IO implementation. In order to solve this problem, an advanced reservation approach where storage resources are managed based on the reservations to satisfy the I/O performance requirements, has been proposed. In this paper, we apply the concept of reserved data access to MPI-IO, in particular to Two-Phase collective I/O which is primarily used for I/O aggregation in non-contiguous access by MPI applications. We developed a prototype by using Dynamic-CoMPI which supports further improvement of Two-Phase I/O by using a locality aware strategy, and Papio which is a parallel storage system providing performance reservation functionality. After describing our prototype design and implementation, we show leverage of the concept by comparing our implementation with other existing MPI-IO implementations backed by OrangeFS and Lustre. The evaluation experiment confirms that the optimization benefit of Two-Phase I/O can be preserved by our approach, under the resource contention situation. Yusuke Tanimura, Rosa Filgueira, Isao Kojima, Malcolm P. Atkinson 0001 |
CLUSTER | 1 |
| 2011 | Dynamic Data Redistribution for MapReduce JoinsabstractMapReduce has become a popular method for data processing, in particular for large scale datasets, due to its accessibility as a scalable yet convenient programming paradigm. Data processing tasks often involve joins, and the repartition and fragment-replicate joins are two widely-used join algorithms utilised within the MapReduce framework. This paper presents a multi-join supporting tuple redistribution, building on both the repartition and fragment-replicate joins. Hadoop is used to demonstrate how reduce tasks may improve performance by passing intermediate results to other reduce tasks that are better able to process them using Apache ZooKeeper as a means of communication and data transfer. A performance analysis is presented showing the technique has the potential to reduce response times when processing multiple joins in single MapReduce jobs. Steven J. Lynden, Yusuke Tanimura, Isao Kojima, Akiyoshi Matono |
CloudCom | 2 |
| 2010 | ADERIS: Adaptively Integrating RDF Data from SPARQL Endpoints
Steven J. Lynden, Isao Kojima, Akiyoshi Matono, Yusuke Tanimura |
DASFAA (2) | 4 |
| 2009 | Interoperation of world-wide production e-Science infrastructuresabstractAbstract Many production Grid and e‐Science infrastructures have begun to offer services to end‐users during the past several years with an increasing number of scientific applications that require access to a wide variety of resources and services in multiple Grids. Therefore, the Grid Interoperation Now—Community Group of the Open Grid Forum—organizes and manages interoperation efforts among those production Grid infrastructures to reach the goal of a world‐wide Grid vision on a technical level in the near future. This contribution highlights fundamental approaches of the group and discusses open standards in the context of production e‐Science infrastructures. Copyright © 2009 John Wiley & Sons, Ltd. Morris Riedel, Erwin Laure, Thomas Soddemann, Laurence Field, John-Paul Navarro, James Casey, Maarten Litmaath, Jean-Philippe Baud, Birger Koblitz, Charles E. Catlett, Dane Skow, Cindy Zheng, Philip M. Papadopoulos, Mason J. Katz, Neha Sharma 0001, Oxana Smirnova, Balázs Kónya, Peter W. Arzberger, Frank Würthwein, Abhishek Singh Rana, Terrence Martin, M. Wan, Von Welch, Tony Rimovsky, Steven J. Newhouse, Andrea Vanni, Yoshio Tanaka, Yusuke Tanimura, Tsutomu Ikegami, David Abramson 0001, Colin Enticott, Graham Jenkins, Ruth Pordes, Steven Timm, Gidon Moont, Mona Aggarwal, Dave Colling, Olivier van der Aa, Alex Sim, Vijaya Natarajan, Arie Shoshani, Junmin Gu, Gerson Galang, Riccardo Zappi, Luca Magnoni, Vincenzo Ciaschini, Michele Pace, Valerio Venturi, Moreno Marzolla, Paolo Andreetto, Robert Cowles, Shaowen Wang 0001, Yuji Saeki, Hitoshi Sato, Satoshi Matsuoka, Putchong Uthayopas, Somsak Sriprayoonsakul, Oscar Koeroo, Matthew Viljoen, Laura Pearlman, Stephen Pickles, David Wallom, Glenn Moloney, Jerome Lauret, Jim Marsteller, Paul Sheldon, Surya Pathak, Shaun De Witt, Jirí Mencák, Jens Jensen, Matt Hodges, Derek Ross, Sugree Phatanapherom, Gilbert Netzer, Anders Rhod Gregersen, Mike Jones 0002, Péter Kacsuk, Achim Streit, Daniel Mallmann, Felix Wolf 0001, Thomas Lippert, Thierry Delaitre, Eduardo Huedo, Neil Geddes |
Concurr. Comput. Pract. Exp. | 28 |
| 2006 | Deploying Scientific Applications to the PRAGMA Grid Testbed: Strategies and LessonsabstractRecent advances in grid infrastructure and middleware development have enabled various types of applications in science and engineering to be deployed on the grid. The characteristics of these applications and the diverse infrastructure and middleware solutions developed, utilized or adapted by PRAGMA member institutes are summarized. The applications include those for climate modeling, computational chemistry, bioinformatics and computational genomics, remote control of instruments, and distributed databases. Many of the applications are deployed to the PRAGMA grid testbed in routine basis experiments. Strategies for deploying applications without modifications, and those taking advantage of new programming models on the grid are explored and valuable lessons learned are reported. Comprehensive end to end solutions from PRAGMA member institutes that provide important grid middleware components and generalized models of integrating applications and instruments on the grid are also described. David Abramson 0001, Amanda Lynch, Hiroshi Takemiya, Yusuke Tanimura, Susumu Date, Haruki Nakamura, Karpjoo Jeong, Suntae Hwang, Zhonghua Lu, Céline Amoreira, Kim K. Baldridge, Hurng-Chun Lee, Chi-Wei Wang, Horng-Liang Shih, Tomas E. Molina, Wilfred W. Li, Peter W. Arzberger |
CCGRID | 4 |
| 2006 | The PRAGMA Testbed - Building a Multi-Application International Grid
Cindy Zheng, David Abramson 0001, Peter W. Arzberger, Shahaan Ayyub, Colin Enticott, Slavisa Garic, Mason J. Katz, Jae-Hyuck Kwak, Bu-Sung Lee, Philip M. Papadopoulos, Sugree Phatanapherom, Somsak Sriprayoonsakul, Yoshio Tanaka, Yusuke Tanimura, Osamu Tatebe, Putchong Uthayopas |
CCGRID | 14 |
| 2006 | Implementation of Fault-Tolerant GridRPC Applications
Yusuke Tanimura, Tsutomu Ikegami, Hidemoto Nakada, Yoshio Tanaka, Satoshi Sekiguchi |
J. Grid Comput. | 1 |
| 2003 | Discussion on searching capability of distributed genetic algorithm on the gridabstractThe computational grid has become popular recently. Since the grid has the tremendous power, it is expected that the numerical optimization method like genetic algorithms (GA) performs well on the grid. In the former works, only the simple model of GA is applied on the grid. In this paper, when the distributed GA (DGA) is executed on the grid, the considerable issues and problems are discussed for the scalability, dynamic changes and heterogeneity. Through the numerical experiments, it is found that the DGA model has the following features on the grid; DGA has the scalability for searching the solutions with respect to the number of the resources and the results of the DGA are not influenced very much by the dynamic reduction of the number of resources. It is also addressed the affect of the asynchronous migration. As a result, the guideline how to implement the DGA on the grid is described. Yusuke Tanimura, Tomoyuki Hiroyasu, Mitsunori Miki |
IEEE Congress on Evolutionary Computation | 1 |