Keichi Takahashi

dblp:151/8272 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-1607-5694ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ElasticHub: A Cost-Efficient JupyterHub Platform via Automated Scaling with Kubernetes on Hybrid Cloud
Ryutaro Matsumoto, Kohei Taniguchi, Tomonori Hayami, Keichi Takahashi, Susumu Date
CLOSER4
2025 Performance Analysis of mdx II: A Next-Generation Cloud Platform for Cross-Disciplinary Data Science Research
Keichi Takahashi, Tomonori Hayami, Yu Mukaizono, Yuki Teramae, Susumu Date
CLOSER1
2025 Workflow Batch Job Scheduling with Considering Task Dependencies
Kaito Yanai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP2
2025 Load-Aware Multi-Objective Optimization of Controller and Datastore Placement in Distributed Sdns
abstract
ABSTRACT In distributed Software Defined Networking (SDN), multiple controllers need to maintain a consistent view of the network state among themselves using consensus algorithms, introducing additional communication overhead and network delay, especially in large‐scale networks. Therefore, optimizing controller placement presents significant challenges, as it must account not only for the delay between switches and controllers but also for the delay introduced by consensus algorithms. Additionally, SDN controllers have limited capacity in terms of the number of switches they can manage and the network events they can process. Improper placement of controllers can lead to longer message processing times, increased queuing delays, or even controller failures. Thus, achieving balanced workloads among controllers is essential. This study introduces and validates a practical Flow Setup Time (FST) model to measure controller response times. We proposed an advanced multi‐objective optimization approach that incorporates the Variance of Load Balancing (VOLB), to determine the optimal placements of controllers and datastore nodes involved in processing consensus algorithms. Furthermore, we applied this optimization method to different types of real networks from the Internet Topology Zoo dataset. Based on experimental findings, we identified key factors to consider when selecting optimal placement strategies, including the trade‐offs between the number of controllers, the number of datastore nodes, FST, and VOLB.
Xingyuan Kang, Keichi Takahashi, Chawanat Nakasan, Kohei Ichikawa, Hajimu Iida
Concurr. Comput. Pract. Exp.2
2025 Developing an End-to-End 3D X-Ray Ptychography Workflow Using Surrogate Models
abstract
ABSTRACT Recently, X‐ray ptychography has attracted significant attention as a non‐destructive imaging technique with high spatial resolution. However, its application to real‐time imaging is limited by the long execution time required for iterative phase retrieval, which reconstructs sample images from diffraction patterns. To address this issue, deep learning‐based surrogate models have been proposed to accelerate iterative phase retrieval by directly predicting sample images. While these surrogate models achieve significant speed‐ups, they typically ignore the time needed for model training and dataset preparation, which can diminish their benefits. Consequently, conventional iterative phase retrieval may outperform surrogate‐based approaches in end‐to‐end performance. This study aims to implement real‐time X‐ray ptychography using surrogate models that explicitly incorporate model training and dataset preparation into the workflow. Specifically, we propose a method that constructs a sample‐specific surrogate model on‐the‐fly using a small subset of observed diffraction patterns and uses its predictions as initial estimates for iterative phase retrieval. The proposed method is up to 2.72 times faster than conventional iterative phase retrieval, even when including training and dataset preparation times. Moreover, the proposed method ensures that the reconstructed images satisfy physical constraints. Comprehensive performance evaluations further demonstrate that the trade‐off between model accuracy and preparation time is critical for optimizing the total execution time in the X‐ray ptychography workflow.
Ryota Koda, Keichi Takahashi, Hiroyuki Takizawa, Nozomu Ishiguro, Yukio Takahashi
Concurr. Comput. Pract. Exp.2
2024 Modernizing an Operational Real-Time Tsunami Simulator to Support Diverse Hardware Platforms
abstract
To issue early warnings and rapidly initiate disaster responses after tsunami damage, various tsunami inundation forecast systems have been deployed worldwide. Japan's Cabinet Office operates a forecast system that utilizes supercomputers to perform tsunami propagation and inundation simulation in real time. Although this real-time approach is able to produce significantly more accurate forecasts than the conventional database-driven approach, its wider adoption was hindered because it was specifically developed for vector supercomputers. In this paper, we migrate the simulation code to modern CPUs and GPUs in a minimally invasive manner to reduce the testing and maintenance costs. A directive-based approach is employed to retain the structure of the original code while achieving performance portability, and hardware-specific optimizations including load balance improvement for GPUs are applied. The migrated code runs efficiently on recent CPUs, GPUs and vector processors: a six-hour tsunami simulation using over 47 million cells completes in less than 2.5 minutes on 32 Intel Sapphire Rapids CPUs and 1.5 minutes on 32 NVIDIA H100 GPUs. These results demonstrate that the code enables broader access to accurate tsunami inundation forecasts.
Keichi Takahashi, Takashi Abe, Akihiro Musa, Yoshihiko Sato, Yoichi Shimomura, Hiroyuki Takizawa, Shunichi Koshimura
CLUSTER1
2024 Clustering Based Job Runtime Prediction for Backfilling Using Classification
Hang Cui 0005, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP2
2024 Maximizing Energy Budget Utilization Using Dynamic Power Cap Control
Sho Ishii, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP2
2024 A Node Selection Method for on-Demand Job Execution with Considering Deadline Constraints
Daiki Nakai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP2
2024 Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
Keichi Takahashi, Hiroyuki Takizawa
PDCAT2
2022 Acar: An application-aware network routing system using SRv6
abstract
The optimal path varies depending on the communication characteristics of each application. However, existing routing protocols such as BGP and OSPF do not take this fact into account. Although Software Defined Networking (SDN) has been considered as a possible solution to this problem, SDN technologies relying on a centralized controller has scalability issues. SRv6, which is a source routing protocol that enables SDN, handles routing decisions in a decentralized manner and is expected to scale better than previous SDN technologies.This paper proposes Acar, an adaptive routing system using SRv6 that adaptively controls routing by considering the bandwidth requirements of applications and the link utilization of the network. We conducted experiments on a virtual network and demonstrated that Acar achieves better load balancing between links and higher throughput compared to ECMP.
Tomoki Sugiura, Keichi Takahashi, Kohei Ichikawa, Hajimu Iida
CCNC2
2022 A Real-time Flood Inundation Prediction on SX-Aurora TSUBASA
abstract
Due to extreme weather, record-breaking heavy rainfalls frequently cause severe flood damages. Thus, there is a strong demand for predicting flood scales to mitigate damages. In this paper, we propose a real-time flood inundation prediction system on a shared HPC system. Although the Rainfall-Runoff Inundation (RRI) model has been developed for predicting large-scale flood inundation, it is necessary to improve the performance for real-time prediction. Since the RRI model is highly memory-bound, we port the RRI simulation code to the latest vector computing system, SX-Aurora TSUBASA (SX-AT), which provides high sustained memory bandwidth. We discuss performance optimization of the RRI code at the node level and MPI parallelization strategies. The RRI code also needs to output intermediate results at a high frequency. Thus, the RRI code is split into file I/O operation and kernel computation, which are assigned to different kinds of processors using the heterogeneity of SX-AT. Furthermore, we discuss a resource demand estimation method to minimize the amount of shared computing resources used for prediction in order to reduce the impact on other users sharing the system. In our evaluation, we demonstrate that SX-AT with only 32 cores can meet the real-time simulation requirement of simulating 7-hour flood inundation for the Tohoku region of Japan within 20 minutes. The evaluation results also demonstrate that the proposed method can adaptively adjust the computing resource amount used for the real-time simulation, and thus reduce the computing resource by 75% in comparison with the worst-case scenario of conservative static resource allocation.
Yoichi Shimomura, Akihiro Musa, Yoshihiko Sato, Atsuhiko Konja, Guoqing Cui, Rei Aoyagi, Keichi Takahashi, Hiroyuki Takizawa
HIPC7
2022 Sparse Communication for Federated Learning
abstract
Federated learning trains a model on a centralized server using datasets distributed over a massive amount of edge devices. Since federated learning does not send local data from edge devices to the server, it preserves data privacy. It transfers the local models from edge devices instead of the local data. However, communication costs are frequently a problem in federated learning. This paper proposes a novel method to reduce the required communication cost for federated learning by transferring only top updated parameters in neural network models. The proposed method allows adjusting the criteria of updated parameters to trade-off the reduction of communication costs and the loss of model accuracy. We evaluated the proposed method using diverse models and datasets and found that it can achieve comparable performance to transfer original models for federated learning. As a result, the proposed method has achieved a reduction of the required communication costs around 90% when compared to the conventional method for VGG16. Furthermore, we found out that the proposed method is able to reduce the communication cost of a large model more than of a small model due to the different threshold of updated parameters in each model architecture.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Chawanat Nakasan, Pattara Leelaprute, Hajimu Iida
ICFEC2
2022 A Task-Parallel Runtime for Heterogeneous Multi-node Vector Systems
Kazuki Ide, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
PDCAT2
2022 Towards Priority-Flexible Task Mapping for Heterogeneous Multi-core NUMA Systems
Mulya Agung, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
PDCAT3
2022 Equivalence Checking of Code Transformation by Numerical and Symbolic Approaches
Shunpei Sugawara, Keichi Takahashi, Yoichi Shimomura, Ryusuke Egawa, Hiroyuki Takizawa
PDCAT2
2022 A codesign framework for online data analysis and reduction
abstract
Abstract Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation‐analysis‐reduction pipelines coupled in‐memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application‐ and system‐specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. This toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next‐generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers.
Kshitij Mehta, Bryce Allen, Matthew Wolf, Jeremy Logan, Eric Suchyta, Swati Singhal, Jong Choi 0001, Keichi Takahashi, Kevin A. Huck, Igor Yakushin, Alan Sussman, Todd S. Munson, Ian T. Foster, Scott Klasky
Concurr. Comput. Pract. Exp.8
2021 Comparative Performance Study of Lightweight Hypervisors Used in Container Environment
Keichi Takahashi, Kohei Ichikawa, Hajimu Iida, Pree Thiengburanathum, Passakorn Phannachitta
CLOSER2
2020 Federated Learning of Neural Network Models with Heterogeneous Structures
abstract
Federated learning trains a model on a centralized server using datasets distributed over a large number of edge devices. Applying federated learning ensures data privacy because it does not transfer local data from edge devices to the server. Existing federated learning algorithms assume that all deployed models share the same structure. However, it is often infeasible to distribute the same model to every edge device because of hardware limitations such as computing performance and storage space. This paper proposes a novel federated learning algorithm to aggregate information from multiple heterogeneous models. The proposed method uses weighted average ensemble to combine the outputs from each model. The weight for the ensemble is optimized using black box optimization methods. We evaluated the proposed method using diverse models and datasets and found that it can achieve comparable performance to conventional training using centralized datasets. Furthermore, we compared six different optimization methods to tune the weights for the weighted average ensemble and found that tree parzen estimator achieves the highest accuracy among the alternatives.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Hajimu Iida, Chawanat Nakasan
ICMLA2
2020 Massively Parallel Causal Inference of Whole Brain Dynamics at Single Neuron Resolution
abstract
Empirical Dynamic Modeling (EDM) is a nonlinear time series causal inference framework. The latest implementation of EDM, cppEDM, has only been used for small datasets due to computational cost. With the growth of data collection capabilities, there is a great need to identify causal relationships in large datasets. We present mpEDM, a parallel distributed implementation of EDM optimized for modern GPU-centric supercomputers. We improve the original algorithm to reduce redundant computation and optimize the implementation to fully utilize hardware resources such as GPUs and SIMD units. As a use case, we run mpEDM on AI Bridging Cloud Infrastructure (ABCI) using datasets of an entire animal brain sampled at single neuron resolution to identify dynamical causation patterns across the brain. mpEDM is 1,530× faster than cppEDM and a dataset containing 101,729 neuron was analyzed in 199 seconds on 512 nodes. This is the largest EDM causal inference achieved to date.
Wassapon Watanakeesuntorn, Keichi Takahashi, Kohei Ichikawa, Joseph Park, George Sugihara, Ryousei Takano, Jason H. Haga, Gerald M. Pao
ICPADS2
2020 Retraining Quantized Neural Network Models with Unlabeled Data
abstract
Running neural network models on edge devices is attracting much attention by neural network researchers since edge computing technology is becoming more powerful than ever. However, deploying large neural network models on edge devices is challenging due to the limitation in available computing resources and storage space. Therefore, model compression techniques have been recently studied to reduce the model size and fit models on resource-limited edge devices. Compressing neural network models reduces the size of a model, but also degrades the accuracy of the model since it reduces the precision of weights in the model. Consequently, a retraining method is required to recover the accuracy of compressed models. Most existing retraining methods require the original labeled training datasets to retrain the models, but labeling is a time-consuming process. In particular, we cannot always access the original labeled datasets because of privacy policies and license limitations. In this paper, we propose a method to retrain a compressed neural network model with an unlabeled dataset that is different from the original labeled dataset. We compress the neural network model using quantization to decrease the size of the model. Subsequently, the compressed model is retrained by our proposed retraining method without using a labeled dataset to recover the accuracy of the model. We compared the proposed retraining method against the conventional retraining. The proposed method reduced the size of VGG-16 and ResNet-50 by 81.10% and 52.45%, respectively without significant accuracy loss. In addition, our proposed retraining method is clearly faster than the conventional retraining method.
Kundjanasith Thonglek, Keichi Takahashi, Kohei Ichikawa, Chawanat Nakasan, Hidemoto Nakada, Ryousei Takano, Hajimu Iida
IJCNN2
2019 Improving Resource Utilization in Data Centers using an LSTM-based Prediction Model
abstract
Data centers are centralized facilities where computing and networking hardware are aggregated to handle large amounts of data and computation. In a data center, computing resources such as CPU and memory are usually managed by a resource manager. The resource manager accepts resource requests from users and allocates resources to their applications. A commonly known problem in resource management is that users often request more resources than their applications actually use. This leads to the degradation of overall resource utilization in a data center. This paper aims to improve resource utilization in data centers by predicting the required resource for each application. We designed and implemented a neural network model based on Long Short-Term Memory (LSTM) to predict more efficient resource allocation for a job based on historical data. Our model has two LSTM layers each of which learns the relationship between: (1) allocation and usage, and (2) CPU and memory. We used Googles cluster-usage trace, which contains a trace of resource allocation and usage for each job executed on a Google data center, to train our neural network. Googles cluster scheduler simulator was used to evaluate our proposed method. Our simulation indicated that the proposed method improved the CPU utilization and memory utilization by 10.71% and 47.36%, respectively, compared to a conventional resource manager. Moreover, we discovered that increasing the memory cell size of our LSTM model improves the accuracy of the prediction in return for longer training time.
Kundjanasith Thonglek, Kohei Ichikawa, Keichi Takahashi, Hajimu Iida, Chawanat Nakasan
CLUSTER3
2017 Highly Reconfigurable Computing Platform for High Performance Computing Infrastructure as a Service: Hi-IaaS
Akihiro Misawa, Susumu Date, Keichi Takahashi, Takashi Yoshikawa, Masahiko Takahashi, Masaki Kan, Yasuhiro Watashiba, Yoshiyuki Kido, Chonho Lee, Shinji Shimojo
CLOSER3
2017 PFAnalyzer: A Toolset for Analyzing Application-Aware Dynamic Interconnects
abstract
Recent rapid scale out of high performance computing systems has rapidly and continuously increased the scale and complexity of the interconnects. As a result, current static and over-provisioned interconnects are becoming cost-ineffective. Against this background, we have been working on the integration of network programmability into the interconnect control, based on the idea that dynamically controlling the packet flow in the interconnect according to the communication pattern of applications can increase the utilization of interconnects and improve application performance. Interconnect simulators come in handy especially when investigating the performance characteristics of interconnects with different topologies and parameters. However, little effort has been put towards the simulation of packet flow in dynamically controlled interconnects, while simulators for static interconnects have been extensively researched and developed. To facilitate analysis on the performance characteristics of dynamic interconnects, we have developed PFAnalyzer. PFAnalyzer is a toolset composed of PFSim, an interconnect simulator specialized for dynamic interconnects, and PFProf, a profiler. PFSim allows interconnect researchers and designers to investigate congestion in the interconnect for an arbitrary cluster configuration and a set of communication patterns collected by PFProf. PFAnalyzer is used to demonstrate how dynamically controlling the interconnects can reduce congestion and potentially improve the performance of applications.
Keichi Takahashi, Susumu Date, Khureltulga Dashdavaa, Yoshiyuki Kido, Shinji Shimojo
CLUSTER1
2016 Network Access Control Towards Fully-Controlled Cloud Infrastructure
abstract
Recently, researchers' and scientists' interest and concern to Internet of Things (IoT) have been remarkably increasing. A diversity of IoT devices such as mobile phones, sensors and even scientific measurement facilities have been connected to the Internet and then generating an enormous amount of data. From the demands on computational resources enough to analyze such data, the utilization of the cloud has been a major trend in these days. Taking aggregation and distribution of data from and to IoT devices on the cloud into consideration, however, access control to such data gives rise to an important problem. Each of IoT devices may have a security policy and each user may have a different attribute. For achieving safe access control to data, a fully-controlled infrastructure where access to network resources is controlled as well as computational resources is required. From such a consideration, this paper proposes an access-controlled networking mechanism that dynamically organizes a flexible and secure network linking IoT devices, computational resources and users on the cloud, based on user's attribute and IoT device security policies. The architecture of FlowSieve, which we have designed and implemented in this preliminary stage of the research, is presented as well as our envisaged fully access-controlled cloud for secure data access.
Takuya Yamada, Keichi Takahashi, Masaya Muraki, Susumu Date, Shinji Shimojo
CloudCom2