Hiroyuki Takizawa

dblp:80/5444 · DBLP profile ↗
← Back
53ranked-venue papers
10as first author
20since 2021 · last 2026
0000-0003-2858-3140ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorSecurity and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Fully Homomorphic Encryption Inference of Neural Networks Using CKKS-TFHE Scheme Switching and Accelerated Linear Layers
abstract
Fully homomorphic encryption (FHE) allows neural network inference to be performed directly on encrypted data and models, preserving end-to-end privacy. One of the most widely used FHE schemes, CKKS, supports approximate arithmetic over real numbers and is well-suited for computing polynomial operations. However, CKKS does not natively support conditional branching, making it unsuitable for non-linear functions such as ReLU, which behaves differently depending on whether the input is negative or nonnegative. To address this, ReLU is typically approximated using polynomials. Unfortunately, this approach introduces two key drawbacks: (1) the approximation is only valid within a certain input range, and (2) identifying this range typically requires prior exposure of cleartext data to perform profiling, thus partially compromising privacy. This requirement is problematic in practice, as small profiling errors can lead to substantial accuracy degradation. On the other hand, the TFHE scheme supports binary gate-level computation, enabling precise implementation of conditional operations such as ReLU. Yet TFHE lacks the SIMD capabilities of CKKS, leading to significant computational latency. This paper explores the use of scheme switching for encrypted neural network inference: leveraging CKKS for polynomial-compatible layers and switching to TFHE for layers that require conditional logic, such as ReLU. We demonstrate that this hybrid approach substantially improves the numerical fidelity of encrypted neural network inference: while polynomial approximations suffer from numerical divergence, scheme switching more closely matches non-encrypted inference.
Anas Banta Seutia, Muhammad Zaky Firdaus, Muhammad Alfi Ramadhan, Kabul Kurniawan, Muhammad Husni Santriaji, Alfian Amrizal, Reza Pulungan, Hiroyuki Takizawa
AsiaCCS8
2025 Workflow Batch Job Scheduling with Considering Task Dependencies
Kaito Yanai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP4
2025 Developing an End-to-End 3D X-Ray Ptychography Workflow Using Surrogate Models
abstract
ABSTRACT Recently, X‐ray ptychography has attracted significant attention as a non‐destructive imaging technique with high spatial resolution. However, its application to real‐time imaging is limited by the long execution time required for iterative phase retrieval, which reconstructs sample images from diffraction patterns. To address this issue, deep learning‐based surrogate models have been proposed to accelerate iterative phase retrieval by directly predicting sample images. While these surrogate models achieve significant speed‐ups, they typically ignore the time needed for model training and dataset preparation, which can diminish their benefits. Consequently, conventional iterative phase retrieval may outperform surrogate‐based approaches in end‐to‐end performance. This study aims to implement real‐time X‐ray ptychography using surrogate models that explicitly incorporate model training and dataset preparation into the workflow. Specifically, we propose a method that constructs a sample‐specific surrogate model on‐the‐fly using a small subset of observed diffraction patterns and uses its predictions as initial estimates for iterative phase retrieval. The proposed method is up to 2.72 times faster than conventional iterative phase retrieval, even when including training and dataset preparation times. Moreover, the proposed method ensures that the reconstructed images satisfy physical constraints. Comprehensive performance evaluations further demonstrate that the trade‐off between model accuracy and preparation time is critical for optimizing the total execution time in the X‐ray ptychography workflow.
Ryota Koda, Keichi Takahashi, Hiroyuki Takizawa, Nozomu Ishiguro, Yukio Takahashi
Concurr. Comput. Pract. Exp.3
2024 Modernizing an Operational Real-Time Tsunami Simulator to Support Diverse Hardware Platforms
abstract
To issue early warnings and rapidly initiate disaster responses after tsunami damage, various tsunami inundation forecast systems have been deployed worldwide. Japan's Cabinet Office operates a forecast system that utilizes supercomputers to perform tsunami propagation and inundation simulation in real time. Although this real-time approach is able to produce significantly more accurate forecasts than the conventional database-driven approach, its wider adoption was hindered because it was specifically developed for vector supercomputers. In this paper, we migrate the simulation code to modern CPUs and GPUs in a minimally invasive manner to reduce the testing and maintenance costs. A directive-based approach is employed to retain the structure of the original code while achieving performance portability, and hardware-specific optimizations including load balance improvement for GPUs are applied. The migrated code runs efficiently on recent CPUs, GPUs and vector processors: a six-hour tsunami simulation using over 47 million cells completes in less than 2.5 minutes on 32 Intel Sapphire Rapids CPUs and 1.5 minutes on 32 NVIDIA H100 GPUs. These results demonstrate that the code enables broader access to accurate tsunami inundation forecasts.
Keichi Takahashi, Takashi Abe, Akihiro Musa, Yoshihiko Sato, Yoichi Shimomura, Hiroyuki Takizawa, Shunichi Koshimura
CLUSTER6
2024 Clustering Based Job Runtime Prediction for Backfilling Using Classification
Hang Cui 0005, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP4
2024 Maximizing Energy Budget Utilization Using Dynamic Power Cap Control
Sho Ishii, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP4
2024 A Node Selection Method for on-Demand Job Execution with Considering Deadline Constraints
Daiki Nakai, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
JSSPP4
2024 Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
Keichi Takahashi, Hiroyuki Takizawa
PDCAT3
2024 Conflict-aware workload co-execution on SX-aurora TSUBASA
abstract
Abstract NEC SX-Aurora TSUBASA (SX-AT) is the latest vector supercomputer, consisting of host processors called Vector Hosts (VHs) and vector processors called Vector Engines (VEs). The goal of this work is to simultaneously use both VHs and VEs to increase the resource utilization and improve the system throughput by co-executing more workloads. One difficulty is that performance interferences among VH and VE workloads could occur because they share some computing resources and potentially compete to use the same resource at the same time, so-called resource conflicts. To achieve efficient workload co-execution, first, this paper experimentally investigates the performance interference between a VH and a VE, when each of the two processors executes a different workload. It is empirically shown that the frequency of system calls from the VE workload could be a good indicator to predict if the co-execution could cause severe performance interference, even though monitoring system calls requires a huge runtime overhead and it is impractical to simply use it for decision making of co-execution. Then, this paper proposes a workload co-execution strategy based on a practical approach to identifying a pair of VE and VH workloads that could cause severe performance interferences. Our evaluation results clearly demonstrate that the system call frequency can be used to predict if the workload can affect the performance of another co-executing workload, and VH’s CPU load can be a good approximation of the system call frequency. The proposed approach based on the CPU loads could accurately identify a pair of workloads causing frequent resource conflicts, and thus reduce the risk of severe performance interferences between co-executing workloads on an SX-AT system, resulting in shorter makespan without significantly increasing the turn-around time.
Riku Nunokawa, Yoichi Shimomura, Mulya Agung, Ryusuke Egawa, Hiroyuki Takizawa
CCF Trans. High Perform. Comput.5
2022 A Real-time Flood Inundation Prediction on SX-Aurora TSUBASA
abstract
Due to extreme weather, record-breaking heavy rainfalls frequently cause severe flood damages. Thus, there is a strong demand for predicting flood scales to mitigate damages. In this paper, we propose a real-time flood inundation prediction system on a shared HPC system. Although the Rainfall-Runoff Inundation (RRI) model has been developed for predicting large-scale flood inundation, it is necessary to improve the performance for real-time prediction. Since the RRI model is highly memory-bound, we port the RRI simulation code to the latest vector computing system, SX-Aurora TSUBASA (SX-AT), which provides high sustained memory bandwidth. We discuss performance optimization of the RRI code at the node level and MPI parallelization strategies. The RRI code also needs to output intermediate results at a high frequency. Thus, the RRI code is split into file I/O operation and kernel computation, which are assigned to different kinds of processors using the heterogeneity of SX-AT. Furthermore, we discuss a resource demand estimation method to minimize the amount of shared computing resources used for prediction in order to reduce the impact on other users sharing the system. In our evaluation, we demonstrate that SX-AT with only 32 cores can meet the real-time simulation requirement of simulating 7-hour flood inundation for the Tohoku region of Japan within 20 minutes. The evaluation results also demonstrate that the proposed method can adaptively adjust the computing resource amount used for the real-time simulation, and thus reduce the computing resource by 75% in comparison with the worst-case scenario of conservative static resource allocation.
Yoichi Shimomura, Akihiro Musa, Yoshihiko Sato, Atsuhiko Konja, Guoqing Cui, Rei Aoyagi, Keichi Takahashi, Hiroyuki Takizawa
HIPC8
2022 A Method for Reducing Time-to-Solution in Quantum Annealing Through Pausing
abstract
Recent research has shown that alternative annealing schedules provide the means for improving performance in modern quantum annealing devices. One such type of schedule is forward annealing with a pause, in which there is a period of time when system evolution is paused. While the results from using this type of schedule have been promising, effectively using a pause is not a trivial task. One challenge associated with introducing a pause into the schedule is determining the point in the anneal at which the pause will start. Additionally, tuning the schedule in real-time requires a significant amount of time. A second challenge is that while a pause may increase the number of correct solutions returned from the annealer, the time-to-solution, a standard metric for measuring performance in quantum annealing, will not necessarily be improved. We propose a method for constructing annealing schedules containing a pause that avoids the costly process of determining the optimal pause location in an online manner. We also evaluate our method on the subset sum problem, a problem of practical significance, and show that our method is able to achieve a 70% reduction in time-to-solution from a standard schedule containing no pause.
Michael R. Zielewski, Hiroyuki Takizawa
HPC Asia2
2022 Toward Building a Digital Twin of Job Scheduling and Power Management on an HPC System
Tatsuyoshi Ohmura, Yoichi Shimomura, Ryusuke Egawa, Hiroyuki Takizawa
JSSPP4
2022 A Task-Parallel Runtime for Heterogeneous Multi-node Vector Systems
Kazuki Ide, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
PDCAT4
2022 Towards Priority-Flexible Task Mapping for Heterogeneous Multi-core NUMA Systems
Mulya Agung, Keichi Takahashi, Yoichi Shimomura, Hiroyuki Takizawa
PDCAT5
2022 An Advantage Actor-Critic Deep Reinforcement Learning Method for Power Management in HPC Systems
Fitra Rahmani Khasyah, Gemilang Santiyuda, Gabriel Kaunang, Faizal Makhrus, Alfian Amrizal, Hiroyuki Takizawa
PDCAT6
2022 Equivalence Checking of Code Transformation by Numerical and Symbolic Approaches
Shunpei Sugawara, Keichi Takahashi, Yoichi Shimomura, Ryusuke Egawa, Hiroyuki Takizawa
PDCAT5
2021 neoSYCL: a SYCL implementation for SX-Aurora TSUBASA
abstract
Recently, the high-performance computing world has moved to more heterogeneous architectures. Thus, it has become a standard practice to offload a part of application execution to dedicated accelerators. However, the disadvantage in productivity is still a problem in programming for accelerators. This paper proposes neoSYCL: a SYCL implementation for SX-Aurora TSUBASA, aiming to improve productivity and achieve comparable performance with native implementations. Unlike other implementations, neoSYCL can identify and separate the kernel part of the SYCL code at the source code level. Thus, this approach can easily be moved to any heterogeneous architectures using the offload programming model. In this paper, we show the evaluation results on SX-Aurora TSUBASA. To quantitatively discuss not only performance but also the productivity, we use two different benchmarks and code-complexity metrics for the evaluation. The results show that neoSYCL can improve productivity while reaching the same performance as native implementations.
Yinan Ke, Mulya Agung, Hiroyuki Takizawa
HPC Asia3
2021 Evaluating the Performance and Conformance of a SYCL Implementation for SX-Aurora TSUBASA
Mulya Agung, Hiroyuki Takizawa
PDCAT3
2021 Towards Conflict-Aware Workload Co-execution on SX-Aurora TSUBASA
Riku Nunokawa, Yoichi Shimomura, Mulya Agung, Ryusuke Egawa, Hiroyuki Takizawa
PDCAT5
2021 OpenCL-like offloading with metaprogramming for SX-Aurora TSUBASA
abstract
This paper presents an OpenCL-like offload programming framework for NEC SX-Aurora TSUBASA (SX-Aurora) and also discusses the benefit of employing metaprogramming to describe architecture-specific parts of the programs. Unlike traditional vector systems, one node of an SX-Aurora system consists of a host processor and some vector processors on PCI-Express cards, which are called a vector host and vector engines, respectively. Since the standard OpenCL execution model does not naturally fit in the vector engine, this paper discusses how to adapt the OpenCL specification to SX-Aurora while considering the trade off between performance and code portability. This paper employs OpenCL to minimize non-portable parts of an application code for offload programming, and then metaprogramming to describe the non-portable parts. Performance evaluation results clearly demonstrate that, with a moderate programming effort, the proposed framework can express the collaboration between a vector host and a vector engine so as to make a good use of both of the two different processors. By delegating the right task to the right processor, an OpenCL-like program can fully exploit the performance of SX-Aurora. Moreover, metaprogramming can express vectorization-aware performance optimization to enhance the performance portability across different architectures including SX-Aurora.
Hiroyuki Takizawa, Shinji Shiotsuki, Naoki Ebata, Ryusuke Egawa
Parallel Comput.1
2020 Failure Prediction in Datacenters Using Unsupervised Multimodal Anomaly Detection
abstract
Predicting hard drive failures in datacenters can help avoid wasting resources and waiting time for recovery. Anomaly detection from sensing data is commonly used for predicting failures. Usually, conventional threshold-based anomaly detection methods consider each sensor independently. However, deciding an optimal threshold for each type of sensors is not trivial, especially for large-scale systems in datacenters. To detect failures that cannot conventionally be detected, multimodal anomaly detection becomes crucial integrating sensing data from different types of sensors. This work proposes a correlation-based multimodal anomaly detection approach. This approach is applied to a Network-Attached Storage (NAS) system with multiple hard disk drives (HDDs) and three sensors, which are a thermal camera, a microphone, and system performance logs. The unimodal results show that the auditory and system performance model can detect temporal anomalies, and the thermal model can detect spatial anomalies. The multimodal results show that even with a simple filter and detection algorithms, the multimodal approach was able to detect failure signs before the real failure and also earlier than the auditory unimodal approach.
Minglu Zhao, Reo Furuhata, Mulya Agung, Hiroyuki Takizawa, Tomoya Soma
IEEE BigData4
2020 Xevolver: A code transformation framework for separation of system-awareness from application codes
abstract
Summary This paper introduces the Xevolver code transformation framework to separate system‐aware code optimizations from HPC application codes. System‐aware code optimizations often make it difficult for programmers to maintain HPC application codes. On the other side, system‐aware code optimizations are mandatory to exploit the performance of target HPC systems. To achieve both high maintainability and high performance, the Xevolver framework provides an easy way to express system‐aware code optimizations as user‐defined code transformation rules. Those rules can be defined separately from HPC application codes. As a result, an HPC application code is converted into its optimized version for a particular target system just before the compilation, and standard HPC programmers do not usually need to maintain the optimized version that could be complicated and difficult‐to‐maintain. In this paper, three important components of the Xevolver framework are described, and then their practicality and benefits are demonstrated through six case studies. Accordingly, the user‐defined code transformation approach behind the Xevolver framework is promising to express system‐awareness for extracting the performance of an HPC system, and also for sharing expert knowledge and experiences about code optimizations. As the complexity and diversity of HPC system architectures are increasing in an extreme‐scale computing era, system‐aware code optimization without overcomplicating the code as discussed in this paper will become more and more important in the future.
Kazuhiko Komatsu, Ayumu Gomi, Ryusuke Egawa, Daisuke Takahashi, Reiji Suda, Hiroyuki Takizawa
Concurr. Comput. Pract. Exp.6
2019 An OpenCL-Like Offload Programming Framework for SX-Aurora TSUBASA
abstract
This paper presents an OpenCL-like offload programming framework for NEC SX-Aurora TSUBASA (SXAurora). Unlike traditional vector systems, one node of an SXAurora system consists of a host processor and some vector processors on PCI-Express cards, which are called a vector host and vector engines, respectively. Since the standard OpenCL execution model does not naturally fit in the vector engine, this paper discusses how to adapt the OpenCL specification to SXAurora while considering the trade off between performance and code portability. Performance evaluation results clearly demonstrate that, with a moderate programming effort, the proposed framework can express the collaboration between a vector host and a vector engine so as to make a good use of both of the two different processors. By delegating the right task to the right processor, an OpenCL-like program can fully exploit the performance of SX-Aurora.
Hiroyuki Takizawa, Shinji Shiotsuki, Naoki Ebata, Ryusuke Egawa
PDCAT1
2018 Performance Estimation of Deeply Pipelined Fluid Simulation on Multiple FPGAs with High-speed Communication Subsystem
abstract
To precisely evaluate the sustained performance and scalability of pipelined multiple FPGAs, this paper presents the implementation of a high-speed communication subsystem for a deeply pipelined stream computing platform. Internally, a pipeline of hardware modules for a domain-specific application is implemented in the FPGAs, where they are directly connected through their serial transceiver links. The necessary inter-FPGA communication subsystem with a flow control mechanism is implemented with Intel Arria 10 FPGAs, where the resource consumption and sustained network throughput are obtained, which averages at 7.92 GB/s. Performance estimation of a fluid simulation using the measured inter-FPGA network parameters is shown for a varied pipeline depth with explored temporal and spatial parallel options. Results show that the proposed platform with 16 FPGAs is estimated to achieve a sustained performance of 4.8 TFlops and may be scaled further by deepening the pipeline with even up to 128 FPGAs.
Antoniette Mondigo, Kentaro Sano, Hiroyuki Takizawa
ASAP3
2018 Automatic Hyperparameter Tuning of Machine Learning Models under Time Constraints
abstract
Most machine learning models use hyperparameters empirically defined in advance of their training processes in a time-consuming and try-and-error fashion. Hence, there is a strong demand for systematically finding an appropriate hyperparameter configuration in a practical time. Recent works have been interested in Bayesian Optimization to tune the hyperparameters with a less number of trials, using a Gaussian Process to determine the next hyperparameter configuration being sampled for evaluation. Most of the works use some criteria including the probability of improving (GP-PI), the expected improvement (GP-EI), and the upper confidence bounds (GP-UCB), without consideration of the execution time of each trial. In this paper, we focus on minimizing the total execution time to find an appropriate configuration. Specifically, we propose to take the execution time of each trial into account. We demonstrate the feasibility of the proposed approach and show that our proposal can find an optimal or suboptimal hyperparameter configuration faster than other Bayesian optimization-based approaches in terms of execution time.
Mulya Agung, Ryusuke Egawa, Reiji Suda, Hiroyuki Takizawa
IEEE BigData5
2018 A Failure Prediction-Based Adaptive Checkpointing Method with Less Reliance on Temperature Monitoring for HPC Applications
abstract
Checkpointing with a constant checkpoint interval, a so-called constant checkpointing method, is commonly used in HPC field and has been proved to be the optimal solution for failures whose inter-arrival times are distributed exponentially. On the other hand, previous works have shown that there is a high correlation between processor temperature and its failure rate. By analyzing the results of the temperature monitoring on a parallel application, we noticed that the failure rate is dynamically changing and the failure inter-arrival times do not follow an exponential distribution. Under such a scenario, the constant checkpointing method is not the optimal solution and thus a checkpointing method with an adaptive checkpoint interval, called an adaptive checkpointing method, is required to achieve high performance. However, to use the adaptive method, the processor temperature must be constantly monitored in order to decide the timing for checkpointing. In this paper, we propose an adaptive checkpointing method with less reliance on the temperature monitoring. Our proposed method uses the timings of already occurred failures, called the prior failures, to estimate the mean time to failure (MTTF) of the next failure, called the posterior failure. The timing of the posterior failure is predicted based on the characteristic of a truncated Weibull distribution. The simulation results show that the proposed method can reduce the total wasted time compared to the constant checkpointing method with a considerably small temperature monitoring period.
Alfian Amrizal, Mulya Agung, Ryusuke Egawa, Hiroyuki Takizawa
CLUSTER5
2018 Enhancing Memory Bandwidth in a Single Stream Computation with Multiple FPGAs
abstract
Stream computing is an area where FPGAs can be suitably utilized to meet high performance and high scalability demands. To achieve these, a deep computing pipeline is implemented on multiple FPGAs where stream computing is performed. This paper presents an approach to utilize two masters in a 1D ring network of multiple FPGAs for a single stream computation. Each master FPGA will be reading and writing to their respective DDR3 memories alternately, while streaming through the slave FPGAs. This is done in order to synchronize the computational results on physically separate memory units. Due to this, the aggregate memory bandwidth is doubled, which suggests enhanced performance. The introduction of this streaming concept lays the groundwork towards full utilization of memories in all the FPGAs, as an identified future work.
Antoniette Mondigo, Kentaro Sano, Hiroyuki Takizawa
FPT3
2017 Performance and Power Analysis of SX-ACE Using HP-X Benchmark Programs
abstract
As the SIMD width of modern microprocessors has been widening for keeping up with the computational demand for HPC systems, recently the vector architecture comes back to spotlight. Besides, a modern vector architecture that has been keeping a large SIMD width and a high B/F ratio has survived and evolved in the HPC community. In this paper, to clarify the potential of the modern vector architecture, we present the performance and power analysis of a modern vector supercomputer SX-ACE using HP-X benchmark programs (HPL, HPCG, and HPGMG). Furthermore, the implementation and optimization of these benchmarks on SX-ACE are discussed. The evaluation results show that SX-ACE achieves the highest efficiencies in the HPGMG and HPCG ranking lists. These facts clearly indicate that the powerful vector processing mechanism with a high B/F ratio is mandatory to achieve a high sustained performance in the future HPC systems.
Ryusuke Egawa, Kazuhiko Komatsu, Yoko Isobe, Toshihiro Kato, Soya Fujimoto, Hiroyuki Takizawa, Akihiro Musa, Hiroaki Kobayashi
CLUSTER6
2017 Vectorization-Aware Loop Optimization with User-Defined Code Transformations
abstract
The cost of maintaining an application code would significantly increase if the application code is branched into multiple versions, each of which is optimized for a different architecture. In this work, default and vector versions of a realworld application code are refactored to be a single version, and the differences between the versions are expressed as user-defined code transformations. As a result, application developers can maintain only the single version, and transform it to its vector version just before the compilation. Although code optimizations for a vector processor are sometimes different from those for other processors, application developers can enjoy the performance of the vector processor without increasing the code complexity. Evaluation results demonstrate that vectorization-aware loop optimization for a vector processor can be expressed as user-defined code transformation rules, and thereby significantly improve the performance of a vector processor without major code modifications.
Hiroyuki Takizawa, Thorsten Reimann, Kazuhiko Komatsu, Takashi Soga, Ryusuke Egawa, Akihiro Musa, Hiroaki Kobayashi
CLUSTER1
2017 A Memory Congestion-Aware MPI Process Placement for Modern NUMA Systems
abstract
MPI process placement is an important step to achieve scalable performance on modern non-uniform memory access (NUMA) systems. A recent study on NUMA architectures has shown that, on modern NUMA systems, the memory congestion problem could cause more severe performance degradation than the data locality problem because heavy congestion on memory controllers could cause long latencies. However, conventional work on MPI process placement has focused on locality to minimize the remote-access communication. Moreover, maximizing the locality may actually degrade performance because the load imbalance among nodes in a modern NUMA system may increase. Thus, a process placement algorithm must be designed to consider memory congestion. In this paper, a method to reconcile both the locality and the memory congestion on modern NUMA systems is proposed. This method statically analyzes the application communication pattern to optimize the process placement. A data clustering method is applied to the time-series data of the MPI communications in order to identify data traffics that potentially cause memory congestion. The proposed method has been evaluated with the NPB kernels on a real NUMA system and a simulation environment. Experimental results show that the proposed method can achieve 1.6x performance improvement compared with the current state-of-the-art strategy.
Mulya Agung, Alfian Amrizal, Kazuhiko Komatsu, Ryusuke Egawa, Hiroyuki Takizawa
HiPC5
2017 Optimizing Energy Consumption on HPC Systems with a Multi-Level Checkpointing Mechanism
abstract
Coordinated checkpointing is a widely-used checkpoint/restart (CPR) technique for fault-tolerance in large-scale HPC systems. However, this CPR technique will involve massive amounts of I/O concentration, resulting in considerably high checkpoint overhead and high energy consumption. This paper focuses on multi-level checkpointing that allows the use of different kinds of fast but less reliable storages to reduce the checkpointing frequency to parallel file system (PFS). This paper presents an energy model of multi-level checkpointing and proposes an iterative algorithm that minimizes energy consumption by optimizing the checkpoint interval of each level and selecting the best combination of checkpoint levels. It is confirmed that the algorithm is very fast and effective since it can reach convergence in a relatively small number of iteration steps. This paper also clarifies the fact that it is actually unnecessary to use all the available checkpoint levels in a multi-level CPR mechanism. By selectively using only appropriate checkpoint levels, a significant increase in energy efficiency (9 to 21%) is observed.
Alfian Amrizal, Hiroyuki Takizawa
NAS2
2017 Potential of a modern vector supercomputer for practical applications: performance evaluation of SX-ACE
abstract
Achieving a high sustained simulation performance is the most important concern in the HPC community. To this end, many kinds of HPC system architectures have been proposed, and the diversity of the HPC systems grows rapidly. Under this circumstance, a vector-parallel supercomputer SX-ACE has been designed to achieve a high sustained performance of memory-intensive applications by providing a high memory bandwidth commensurate with its high computational capability. This paper examines the potential of the modern vector-parallel supercomputer through the performance evaluation of SX-ACE using practical engineering and scientific applications. To improve the sustained simulation performances of practical applications, SX-ACE adopts an advanced memory subsystem with several new architectural features. This paper discusses how these features, such as MSHR, a large on-chip memory, and novel vector processing mechanisms, are beneficial to achieve a high sustained performance for large-scale engineering and scientific simulations. Evaluation results clearly indicate that the high sustained memory performance per core enables the modern vector supercomputer to achieve outstanding performances that are unreachable by simply increasing the number of fine-grain scalar processor cores. This paper also discusses the performance of the HPCG benchmark to evaluate the potentials of supercomputers with balanced memory and computational performance against heterogeneous and cutting-edge scalar parallel systems.
Ryusuke Egawa, Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Hiroyuki Takizawa, Hiroaki Kobayashi
J. Supercomput.6
2014 Xevolver: An XML-based code translation framework for supporting HPC application migration
abstract
This paper proposes an extensible programming framework to separate platform-specific optimizations from application codes. The framework allows programmers to define their own code translation rules for special demands of individual systems, compilers, libraries, and applications. Code translation rules associated with user-defined compiler directives are defined in an external file, and the application code is just annotated by the directives. For code transformations based on the rules, the framework exposes the abstract syntax tree (AST) of an application code as an XML document to expert programmers. Hence, the XML document of an AST can be transformed using any XML-based technologies. Our case studies using real applications demonstrate that the framework is effective to separate platform-specific optimizations from application codes, and to incrementally improve the performance of an existing application without messing up the code.
Hiroyuki Takizawa, Shoichi Hirasawa, Yasuharu Hayashi, Ryusuke Egawa, Hiroaki Kobayashi
HiPC1
2012 GPU implementation of phase-based stereo correspondence and its application
abstract
This paper proposes a Graphics Processing Unit (GPU) implementation of the stereo correspondence matching using Phase-Only Correlation (POC). The use of high-accuracy stereo correspondence matching based on POC makes it possible to measure accurate 3D shape of the object using stereo vision, while the drawback of POC-based approach is its high computational cost. Addressing this problem, we propose a GPU implementation of POC-based correspondence matching. Through a set of experiments using a variety of GPUs, we demonstrate that the proposed implementation is high-speed and high-efficiency compared with the CPU implementation. We also apply the proposed approach to a real-time 3D measurement system.
Mamoru Miura, Kinya Fudano, Koichi Ito 0001, Takafumi Aoki, Hiroyuki Takizawa, Hiroaki Kobayashi
ICIP5
2011 CheCL: Transparent Checkpointing and Process Migration of OpenCL Applications
abstract
In this paper, we propose a new transparent checkpoint/restart (CPR) tool, named CheCL, for high-performance and dependable GPU computing. CheCL can perform CPR on an OpenCL application program without any modification and recompilation of its code. A conventional check pointing system fails to checkpoint a process if the process uses OpenCL. Therefore, in CheCL, every API call is forwarded to another process called an API proxy, and the API proxy invokes the API function, two processes, an application process and an API proxy, are launched for an OpenCL application. In this case, as the application process is not an OpenCL process but a standard process, it can be safely check pointed. While CheCL intercepts all API calls, it records the information necessary for restoring OpenCL objects. The application process does not hold any OpenCL handles, but CheCL handles to keep such information. Those handles are automatically converted to OpenCL handles and then passed to API functions. Upon restart, OpenCL objects are automatically restored based on the recorded information. This paper demonstrates the feasibility of transparent check pointing of OpenCL programs including MPI applications, and quantitatively evaluates the runtime overheads. It is also discussed that CheCL can enable process migration of OpenCL applications among distinct nodes, and among different kinds of compute devices such as a CPU and a GPU.
Hiroyuki Takizawa, Kentaro Koyama, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi
IPDPS1
2011 A History-Based Performance Prediction Model with Profile Data Classification for Automatic Task Allocation in Heterogeneous Computing Systems
abstract
In this paper, we propose a runtime performance prediction model for automatic selection of accelerators to execute kernels in OpenCL. The proposed method is a history-based approach that uses profile data for performance prediction. The profile data are classified into some groups, from each of which its own performance model is derived. As the execution time of a kernel depends on some runtime parameters such as kernel arguments, the proposed method first identifies parameters affecting the execution time by calculating the correlation between each parameter and the execution time. A parameter with weak correlation is used for the classification of the profile data and the selection of the performance prediction model. A parameter with strong correlation is used for building a linear model for the prediction of the kernel execution time by using only the classified profile data. Experimental results clearly indicate that the proposed method can achieve more accurate performance prediction than conventional history-based approaches because of the profile data classification.
Katsuto Sato, Kazuhiko Komatsu, Hiroyuki Takizawa, Hiroaki Kobayashi
ISPA3
2010 A Load-Forwarding Mechanism for the Vector Architecture in Multimedia Applications
abstract
Nowadays, multimedia applications (MMAs) form an important workload for general purpose processors. Although the vector architecture is considered the most potential candidate for media processing, the traditional vector architecture has inefficiencies to execute MMAs. This paper proposes a media-oriented vector architecture, which improves the traditional one with a load-forwarding mechanism. The load-forwarding mechanism overcomes the inefficiency on utilization of the memory bandwidth. As a result, the proposed architecture achieves a higher performance with lower hardware cost than the traditional one. This paper evaluates the proposed architecture with architectural design parameters and finds out the most efficient size for the vector architecture when performing MMAs.
Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi
DSD3
2010 A voting-based working set assessment scheme for dynamic cache resizing mechanisms
abstract
Considering the trade-off between performance and power consumption has become significantly important in multi-core processor design. Under this situation, one promising approach is to employ a power-aware dynamic cache partitioning mechanism. This mechanism individually manages activation of each cache way, and exclusively allocates the minimum number of required ways to each thread. In the mechanism, an appropriate number of ways for a thread is decided based on locality assessment. However, sampling results of cache accesses that are used for locality assessment are disturbed by exceptional behaviors of cache accesses, which happen in a very short period. Such sampling results may change locality assessment results to ones that are not along with the overall trend in a long access-sampling period. These assessment results will excessively adapt the cache to exceptional behaviors, and deteriorate energy efficiency. To avoid such excessive adaptation by the exceptional behaviors, this paper proposes a voting-based working set assessment scheme, in which the number of activated ways is adjusted based on majority voting of locality assessment of several short sampling periods. By using the majority voting, the proposed scheme can identify the periods including exceptional behaviors, and ignore the assessment results of these periods. As a result, the proposed scheme makes the cache resizing mechanism more stable and robust. The experimental results indicate that the proposed scheme can reduce energy consumption by up to 24%, and 10% on an average without significant performance degradation in multi-thread execution on a 2-core CMP.
Masayuki Sato 0001, Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi
ICCD3
2009 CheCUDA: A Checkpoint/Restart Tool for CUDA Applications
abstract
In this paper, a tool named CheCUDA is designed to checkpoint CUDA applications that use GPUs as accelerators. As existing checkpoint/restart implementations do not support checkpointing the GPU status, CheCUDA hooks a part of basic CUDA driver API calls in order to record the status changes on the main memory. At checkpointing, CheCUDA stores the status changes in a file after copying all necessary data in the video memory to the main memory and then disabling the CUDA runtime. At restarting, CheCUDA reads the file, re-initializes the CUDA runtime, and recovers the resources on GPUs so as to restart from the stored status. This paper demonstrates that a prototype implementation of CheCUDA can correctly checkpoint and restart a CUDA application written with basic APIs. This also indicates that CheCUDA can migrate a process from one PC to another even if the process uses a GPU. Accordingly, CheCUDA is useful not only to enhance the dependability of CUDA applications but also to enable dynamic task scheduling of CUDA applications required especially on heterogeneous GPU cluster systems. This paper also shows the timing overhead for checkpointing.
Hiroyuki Takizawa, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi
PDCAT1
2009 Performance evaluation of NEC SX-9 using real science and engineering applications
abstract
This paper describes a new-generation vector parallel supercomputer, NEC SX-9 system. The SX-9 processor has an outstanding core to achieve over 100Gflop/s, and a software-controllable on-chip cache to keep the high ratio of the memory bandwidth to the floating-point operation rate. Moreover, its large SMP nodes of 16 vector processors with 1.6Tflop/s performance and 1TB memory are connected with dedicated network switches, which can achieve inter-node communication at 128GB/s per direction. The sustained performance of the SX-9 processor is evaluated using six practical applications in comparison with conventional vector processors and the latest scalar processor such as Nehalem-EP. Based on the results, this paper discusses the performance tuning strategies for new-generation vector systems. An SX-9 system of 16 nodes is also evaluated by using the HPC challenge benchmark suite and a CFD code. Those evaluation results clarify the highest sustained performance and scalability of the SX-9 system.
Takashi Soga, Akihiro Musa, Yoichi Shimomura, Ryusuke Egawa, Ken'ichi Itakura, Hiroyuki Takizawa, Koki Okabe, Hiroaki Kobayashi
SC6
2008 A Performance Study of Secure Data Mining on the Cell Processor
abstract
This paper examines the potential of the Cell processor as a platform for secure data mining on the future volunteer computing systems. Volunteer computing platforms have the potential to provide massive computing power. However, privacy and security concerns prevent using volunteer computing for data mining of sensitive data. The Cell processor comes with a hardware security feature. The secure volunteer data mining can be achieved by using this hardware security feature. In this paper, we present a general security scheme for the volunteer computing, and a secure parallelized K-Means clustering algorithm for the Cell processor. We also evaluate the performance of the algorithm on the Cell secure system simulator. Evaluation results indicate that the proposed secure data clustering outperforms a non-secure clustering algorithm on the general purpose CPU, but incurs a huge performance overhead introduced by the decryption process of the Cell security features.
Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi
CCGRID2
2008 SPRAT: Runtime processor selection for energy-aware computing
abstract
A commodity personal computer (PC) can be seen as a hybrid computing system equipped with two different kinds of processors, i.e. CPU and a graphics processing unit (GPU). Since the superiorities of GPUs in the performance and the power efficiency strongly depend on the system configuration and the data size determined at the runtime, a programmer cannot always know which processor should be used to execute a certain kernel. Therefore, this paper presents a runtime environment that dynamically selects an appropriate processor so as to improve the energy efficiency. The evaluation results clearly indicate that the runtime processor selection at executing each kernel with given data streams is promising for energy-aware computing on a hybrid computing system.
Hiroyuki Takizawa, Katsuto Sato, Hiroaki Kobayashi
CLUSTER1
2008 Implementation and evaluation of a distributed and cooperative load-balancing mechanism for dependable volunteer computing
abstract
This paper proposes a P2P-based dynamic load balancing mechanism to increase the dependability of volunteer computing. The proposed mechanism is incorporated into a volunteer computing middleware, called the Berkeley Open Infrastructure for Network Computing(BOINC). The proposed mechanism provides two additional features: decentralized load balancing and proxy download. The former feature reduces the variation of the execution times for individual tasks, which are usually aggravated by dynamic and unpredictable load changes on volunteer computing resources. The latter offers another way to assign tasks to idle computing resources when the BOINC project server fails in the task assignment. Using a prototype implementation, this paper examines the effect of the proposed mechanism on the performance of a real volunteer computing system. The experimental results show that the proposed mechanism can reduce the maximum turnaround time by 42% and further improve the total throughput of the volunteer computing system by 27%.
Yoshitomo Murata, Tsutomu Inaba, Hiroyuki Takizawa, Hiroaki Kobayashi
DSN3
2008 Effects of MSHR and Prefetch Mechanisms on an On-Chip Cache of the Vector Architecture
abstract
Vector supercomputers have been encountering the memory wall problem and their memory bandwidth per flop/s rate has decreased. To cover the insufficient memory bandwidth per flop/s rate, an on-chip vector cache has been proposed for the vector processors. Although vector caching is effective to increase the sustained performance to a certain degree, it still needs software and hardware supporting mechanisms to extract its potential. To this end, we propose miss status handling registers (MSHR) and a prefetch mechanism. This paper evaluates the performance of the vector cache with the MSHR and the prefetch mechanism on the vector supercomputer across three leading scientific applications. The MSHR is an effective mechanism for handling subsequent vector loads of the same data, which frequently appear in different schemes. The experimental results indicate that the MSHR can improve the computational performance of scientific applications by 1.45×. Moreover, we examine the performance of the prefetch mechanism on the vector cache. The prefetch mechanism increases the computational performance by 1.6×. Accordingly, the MSHR and the prefetching mechanism are very effective optimization options for vector caching of future vector supercomputers even if the vector supercomputers cannot maintain the current memory bandwidth per flop/s rate.
Akihiro Musa, Yoshiei Sato, Takashi Soga, Ryusuke Egawa, Hiroyuki Takizawa, Koki Okabe, Hiroaki Kobayashi
ISPA5
2008 A Utility-Based Double Auction Mechanism for Efficient Grid Resource Allocation
abstract
In Grid Computing, harnessing the power of idle resources in a distributed environment is one of the important features. However, to fully benefit from this computing model, an appropriated resource allocation method needs to be carefully chosen and deployed. A number of studies have been done on this area and one of the promising approaches is to adopt a marketing scheme called the auction model, which has been drawing much attention during past several years. In this paper, we propose a new utility-aware resource allocation protocol to make external scheduling decision in Grid. Users and service providers specify one or more weight values, and then, an auctioneer uses these values for calculating both userspsila and service providerspsila utility values which reflect preference upon the matched members in the different group. Then, we map the scheduling problem with these utility values into the problem in a weighted bipartite graph, and propose a new matching algorithm based on the existing SMP (Stable Marriage Problem) matching algorithm. Finally, the performance of this auctionpsilas awarding technique is evaluated.
Chainan Satayapiwat, Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi
ISPA3
2007 A dependable Peer-to-Peer computing platform
Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi
Future Gener. Comput. Syst.2
2007 Partial distortion entropy maximization for online data clustering
Hiroyuki Takizawa, Hiroaki Kobayashi
Neural Networks1
2006 Implications of Memory Performance for Highly Efficient Supercomputing of Scientific Applications
Akihiro Musa, Hiroyuki Takizawa, Koki Okabe, Takashi Soga, Hiroaki Kobayashi
ISPA2
2006 Hierarchical parallel processing of large scale data clustering on a PC cluster with GPU co-processing
Hiroyuki Takizawa, Hiroaki Kobayashi
J. Supercomput.1
2005 A Workflow Management Mechanism for Peer-to-Peer Computing Platforms
Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi
ISPA2
2004 Multi-grain Parallel Processing of Data-Clustering on Programmable Graphics Hardware
Hiroyuki Takizawa, Hiroaki Kobayashi
ISPA1
2004 Efficient parallel processing of competitive learning algorithms
Kentaro Sano, Shintaro Momose, Hiroyuki Takizawa, Hiroaki Kobayashi, Tadao Nakamura
Parallel Comput.3
1999 A self-organizing network system forming memory from nonstationary probability distributions
abstract
We propose an artificial neural system that forms memory by receiving input vectors obeying an unknown nonstationary probability density function (PDF). The system consists of a set of neural vector quantizer (NVQs), each of which can approximate nonstationary PDFs. Each NVQ exclusively learns a stationary piece of the nonstationary PDF and stores its approximated representation, where the nonstationary PDF consists of some stationary pieces. Experimental results show that the system has functions "memorization", "retention", and "recall" of information which is required in memory systems. The results also illustrate that the system receives inputs from a nonstationary PDF and stores statistical information by distributing it equally over the system. The system can also be used to model nonstationary phenomena. This ability is desirable for various applications, for example, process control, economical modeling, etc.
Taira Nakajima, Hiroyuki Takizawa, Hiroaki Kobayashi, Tadao Nakamura
IJCNN2