Kazuhiko Komatsu

dblp:88/2482 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-4463-8359ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adaptive Parallelization Based on Frame-Level and Tile-Level Parallelisms for VVC Encoding
abstract
ABSTRACT To meet the growing demand for high‐efficiency video compression standards, Versatile Video Coding (VVC) has been developed as a successor to High Efficiency Video Coding (HEVC). VVC offers approximately a 50% reduction in bitrate compared to HEVC while maintaining comparable visual quality. However, the improved performance of VVC comes at the cost of a significantly higher computational complexity, resulting in longer encoding times. Consequently, accelerating the VVC encoding process remains a critical challenge for its practical deployment. Leveraging the evolution of multi‐core processors and encoding tools designed for parallel processing, this study introduces an adaptive parallelization method that integrates frame‐level and tile‐level parallelisms. This method dynamically selects the number of concurrently processed frames and tiles by considering reference dependencies and their effects on coding efficiency. Moreover, the proposed method employs a content‐aware, cyclic adjustment of tile configurations to further reduce the encoding time. The evaluation results demonstrate that the proposed method can achieve a 10.58× speedup over the baseline single‐threaded implementation on average, while limiting the increase in BD‐BR to 3.23% and the degradation in BD‐PSNR to only −0.058 dB. These findings confirm that the method substantially decreases the encoding time without degrading coding efficiency. Furthermore, the results also demonstrate that the scalability of the proposed method is better than that of the conventional parallel method, Wavefront parallel processing.
Karin Onouchi, Masayuki Sato 0001, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi
Concurr. Comput. Pract. Exp.4
2023 Performance Evaluation of Tsunami Evacuation Route Planning on Multiple Annealing Machines
abstract
This paper focuses on the performance evaluation of annealing machines based on quantum annealing and simulated annealing, i.e., D-wave advantage, Fixstars Amplify Annealing Engine, D-wave Neal, and Vector Annealing, by solving a combinatorial optimization problem, tsunami evacuation route planning. First, the problem is modeled into mathematical formulas and converted into the Quadratic Unconstrained Binary Optimization (QUBO) form as an input, and will be executed on the four annealing machines. Then, the relationship between constraint weight in the objective function and the constraint compliance conditions is also discussed. Finally, the evaluation and discussion are conducted from three aspects, the maximum number of evacuees, Time to Solution (TTS), and the solutions quality.
Yihui Liu, Kazuhiko Komatsu, Masahito Kumagai, Masayuki Sato 0001, Hiroaki Kobayashi
CF2
2023 Multi-scale Loss based Electron Microscopic Image Pair Matching Method
abstract
Nanodiffraction Imaging (NDI), a novel imaging technique based on the scanning transmission electron mi-croscopy (STEM), helps the understanding of the relationships between micro-structure and macro-properties. However, the analysis requires image pair matching tasks through meticulous and time-consuming observation and selection by domain experts. Therefore, this paper proposes an image pair matching method for NDI images. The proposed method adopts a two-step training approach. The first step is to perform pre-training by contrastive learning on specialized NDI images. The second step is fine-tuning by a customized model on image pair matching tasks. In the second step, the training is performed with the incorporation of a special loss function, OriDist loss. This loss function is designed to focus on orientation and distribution of multi-scale features. The evaluation results demonstrate the ability of the proposed method to achieve high-accuracy NDI image pair matching, and efficiently reduce the search space of candidate matching images, resulting in a significant reduction in the human workload. Through an extensive ablation study, each component of the proposed method shows positive contributions to the overall performance.
Chunting Duan, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
ICMLA2
2023 A dynamic parameter tuning method for SpMM parallel execution
abstract
Summary Sparse matrix‐matrix multiplication (SpMM) is a basic kernel that is used by many algorithms. Several researches focus on various optimizations for SpMM parallel execution. However, a division of a task for parallelization is not well considered yet. Generally, a matrix is equally divided into blocks for processes even though the sparsities of input matrices are different. The parameter that divides a task into multiple processes for parallelization is fixed. As a result, load imbalance among the processes occurs. To balance the loads among the processes, this article proposes a dynamic parameter tuning method by analyzing the sparsities of input matrices. The experimental results show that the proposed method improves the performance of SpMM for examined matrices by up to 39.5% on a single vector engine and 3.49 on a single CPU.
Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
Concurr. Comput. Pract. Exp.2
2022 Analysis of Precision Vectors for Ising-Based Linear Regression
Kaho Aoyama, Kazuhiko Komatsu, Masahito Kumagai, Hiroaki Kobayashi
PDCAT2
2022 A Partitioned Memory Architecture with Prefetching for Efficient Video Encoders
Masayuki Sato 0001, Yuya Omori, Ryusuke Egawa, Ken Nakamura, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi
PDCAT7
2021 Register Flush-free Runahead Execution for Modern Vector Processors
abstract
Modern vector processors have been designed to achieve high sustained performance, especially in HPC applications, because of their powerful instruction set oriented to data-level parallelism. Additionally, the latest vector processor adopts the out-of-order execution of the vector instructions to exploit instruction-level parallelism due to a significant gap in latency between vector arithmetic instructions and vector load/store instructions. In spite of the effort, this gap still brings a deterioration of sustained performance of the modern vector processors. This paper proposes a runahead execution mechanism for the modern vector processors to fill the latency gap by further exploiting instruction-level parallelism. If the processor stalls due to a long latency instruction, the conventional runahead execution mechanism changes the processor state from a normal mode to a runahead mode, and the processor speculatively executes the subsequent instructions that can cause stalls and their dependencies. However, the conventional runahead execution mechanisms flush the registers' values calculated in the runahead mode after finishing this mode and cannot reuse them in the subsequent normal mode. Since the vector processors have many values even in one vector register, these flushes and re-executions waste the bandwidth between cores and caches. Thus, to solve this problem of the conventional runahead mechanism, our proposed mechanism leaves the registers containing the results in the runahead mode in order for the processor to use the registers even after returning to the normal mode. For correctly using these registers after exiting the runahead mode, the proposed mechanism newly realizes functions to inherit the commit order information and the register aliasing information of the runahead-executed instructions into the normal mode. The evaluation results show that the proposed mechanism improves the performance by up to 20% and 3% on average by the conventional mechanism.
Hikaru Takayashiki, Masayuki Sato 0001, Kazuhiko Komatsu, Hiroaki Kobayashi
SBAC-PAD3
2021 VGL: a high-performance graph processing framework for the NEC SX-Aurora TSUBASA vector architecture
Ilya V. Afanasyev, Vladimir V. Voevodin, Kazuhiko Komatsu, Hiroaki Kobayashi
J. Supercomput.3
2020 A Dynamic Parameter Tuning Method for High Performance SpMM
Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
PDCAT2
2020 Xevolver: A code transformation framework for separation of system-awareness from application codes
abstract
Summary This paper introduces the Xevolver code transformation framework to separate system‐aware code optimizations from HPC application codes. System‐aware code optimizations often make it difficult for programmers to maintain HPC application codes. On the other side, system‐aware code optimizations are mandatory to exploit the performance of target HPC systems. To achieve both high maintainability and high performance, the Xevolver framework provides an easy way to express system‐aware code optimizations as user‐defined code transformation rules. Those rules can be defined separately from HPC application codes. As a result, an HPC application code is converted into its optimized version for a particular target system just before the compilation, and standard HPC programmers do not usually need to maintain the optimized version that could be complicated and difficult‐to‐maintain. In this paper, three important components of the Xevolver framework are described, and then their practicality and benefits are demonstrated through six case studies. Accordingly, the user‐defined code transformation approach behind the Xevolver framework is promising to express system‐awareness for extracting the performance of an HPC system, and also for sharing expert knowledge and experiences about code optimizations. As the complexity and diversity of HPC system architectures are increasing in an extreme‐scale computing era, system‐aware code optimization without overcomplicating the code as discussed in this paper will become more and more important in the future.
Kazuhiko Komatsu, Ayumu Gomi, Ryusuke Egawa, Daisuke Takahashi, Reiji Suda, Hiroyuki Takizawa
Concurr. Comput. Pract. Exp.1
2018 Performance evaluation of a vector supercomputer SX-aurora TSUBASA
Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Mitsuo Yokokawa, Toshikazu Aoyama, Masayuki Sato 0001, Hiroaki Kobayashi
SC1
2017 Performance and Power Analysis of SX-ACE Using HP-X Benchmark Programs
abstract
As the SIMD width of modern microprocessors has been widening for keeping up with the computational demand for HPC systems, recently the vector architecture comes back to spotlight. Besides, a modern vector architecture that has been keeping a large SIMD width and a high B/F ratio has survived and evolved in the HPC community. In this paper, to clarify the potential of the modern vector architecture, we present the performance and power analysis of a modern vector supercomputer SX-ACE using HP-X benchmark programs (HPL, HPCG, and HPGMG). Furthermore, the implementation and optimization of these benchmarks on SX-ACE are discussed. The evaluation results show that SX-ACE achieves the highest efficiencies in the HPGMG and HPCG ranking lists. These facts clearly indicate that the powerful vector processing mechanism with a high B/F ratio is mandatory to achieve a high sustained performance in the future HPC systems.
Ryusuke Egawa, Kazuhiko Komatsu, Yoko Isobe, Toshihiro Kato, Soya Fujimoto, Hiroyuki Takizawa, Akihiro Musa, Hiroaki Kobayashi
CLUSTER2
2017 Vectorization-Aware Loop Optimization with User-Defined Code Transformations
abstract
The cost of maintaining an application code would significantly increase if the application code is branched into multiple versions, each of which is optimized for a different architecture. In this work, default and vector versions of a realworld application code are refactored to be a single version, and the differences between the versions are expressed as user-defined code transformations. As a result, application developers can maintain only the single version, and transform it to its vector version just before the compilation. Although code optimizations for a vector processor are sometimes different from those for other processors, application developers can enjoy the performance of the vector processor without increasing the code complexity. Evaluation results demonstrate that vectorization-aware loop optimization for a vector processor can be expressed as user-defined code transformation rules, and thereby significantly improve the performance of a vector processor without major code modifications.
Hiroyuki Takizawa, Thorsten Reimann, Kazuhiko Komatsu, Takashi Soga, Ryusuke Egawa, Akihiro Musa, Hiroaki Kobayashi
CLUSTER3
2017 A Memory Congestion-Aware MPI Process Placement for Modern NUMA Systems
abstract
MPI process placement is an important step to achieve scalable performance on modern non-uniform memory access (NUMA) systems. A recent study on NUMA architectures has shown that, on modern NUMA systems, the memory congestion problem could cause more severe performance degradation than the data locality problem because heavy congestion on memory controllers could cause long latencies. However, conventional work on MPI process placement has focused on locality to minimize the remote-access communication. Moreover, maximizing the locality may actually degrade performance because the load imbalance among nodes in a modern NUMA system may increase. Thus, a process placement algorithm must be designed to consider memory congestion. In this paper, a method to reconcile both the locality and the memory congestion on modern NUMA systems is proposed. This method statically analyzes the application communication pattern to optimize the process placement. A data clustering method is applied to the time-series data of the MPI communications in order to identify data traffics that potentially cause memory congestion. The proposed method has been evaluated with the NPB kernels on a real NUMA system and a simulation environment. Experimental results show that the proposed method can achieve 1.6x performance improvement compared with the current state-of-the-art strategy.
Mulya Agung, Alfian Amrizal, Kazuhiko Komatsu, Ryusuke Egawa, Hiroyuki Takizawa
HiPC3
2017 Potential of a modern vector supercomputer for practical applications: performance evaluation of SX-ACE
abstract
Achieving a high sustained simulation performance is the most important concern in the HPC community. To this end, many kinds of HPC system architectures have been proposed, and the diversity of the HPC systems grows rapidly. Under this circumstance, a vector-parallel supercomputer SX-ACE has been designed to achieve a high sustained performance of memory-intensive applications by providing a high memory bandwidth commensurate with its high computational capability. This paper examines the potential of the modern vector-parallel supercomputer through the performance evaluation of SX-ACE using practical engineering and scientific applications. To improve the sustained simulation performances of practical applications, SX-ACE adopts an advanced memory subsystem with several new architectural features. This paper discusses how these features, such as MSHR, a large on-chip memory, and novel vector processing mechanisms, are beneficial to achieve a high sustained performance for large-scale engineering and scientific simulations. Evaluation results clearly indicate that the high sustained memory performance per core enables the modern vector supercomputer to achieve outstanding performances that are unreachable by simply increasing the number of fine-grain scalar processor cores. This paper also discusses the performance of the HPCG benchmark to evaluate the potentials of supercomputers with balanced memory and computational performance against heterogeneous and cutting-edge scalar parallel systems.
Ryusuke Egawa, Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Hiroyuki Takizawa, Hiroaki Kobayashi
J. Supercomput.2
2011 CheCL: Transparent Checkpointing and Process Migration of OpenCL Applications
abstract
In this paper, we propose a new transparent checkpoint/restart (CPR) tool, named CheCL, for high-performance and dependable GPU computing. CheCL can perform CPR on an OpenCL application program without any modification and recompilation of its code. A conventional check pointing system fails to checkpoint a process if the process uses OpenCL. Therefore, in CheCL, every API call is forwarded to another process called an API proxy, and the API proxy invokes the API function, two processes, an application process and an API proxy, are launched for an OpenCL application. In this case, as the application process is not an OpenCL process but a standard process, it can be safely check pointed. While CheCL intercepts all API calls, it records the information necessary for restoring OpenCL objects. The application process does not hold any OpenCL handles, but CheCL handles to keep such information. Those handles are automatically converted to OpenCL handles and then passed to API functions. Upon restart, OpenCL objects are automatically restored based on the recorded information. This paper demonstrates the feasibility of transparent check pointing of OpenCL programs including MPI applications, and quantitatively evaluates the runtime overheads. It is also discussed that CheCL can enable process migration of OpenCL applications among distinct nodes, and among different kinds of compute devices such as a CPU and a GPU.
Hiroyuki Takizawa, Kentaro Koyama, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi
IPDPS4
2011 A History-Based Performance Prediction Model with Profile Data Classification for Automatic Task Allocation in Heterogeneous Computing Systems
abstract
In this paper, we propose a runtime performance prediction model for automatic selection of accelerators to execute kernels in OpenCL. The proposed method is a history-based approach that uses profile data for performance prediction. The profile data are classified into some groups, from each of which its own performance model is derived. As the execution time of a kernel depends on some runtime parameters such as kernel arguments, the proposed method first identifies parameters affecting the execution time by calculating the correlation between each parameter and the execution time. A parameter with weak correlation is used for the classification of the profile data and the selection of the performance prediction model. A parameter with strong correlation is used for building a linear model for the prediction of the kernel execution time by using only the classified profile data. Experimental results clearly indicate that the proposed method can achieve more accurate performance prediction than conventional history-based approaches because of the profile data classification.
Katsuto Sato, Kazuhiko Komatsu, Hiroyuki Takizawa, Hiroaki Kobayashi
ISPA2
2009 CheCUDA: A Checkpoint/Restart Tool for CUDA Applications
abstract
In this paper, a tool named CheCUDA is designed to checkpoint CUDA applications that use GPUs as accelerators. As existing checkpoint/restart implementations do not support checkpointing the GPU status, CheCUDA hooks a part of basic CUDA driver API calls in order to record the status changes on the main memory. At checkpointing, CheCUDA stores the status changes in a file after copying all necessary data in the video memory to the main memory and then disabling the CUDA runtime. At restarting, CheCUDA reads the file, re-initializes the CUDA runtime, and recovers the resources on GPUs so as to restart from the stored status. This paper demonstrates that a prototype implementation of CheCUDA can correctly checkpoint and restart a CUDA application written with basic APIs. This also indicates that CheCUDA can migrate a process from one PC to another even if the process uses a GPU. Accordingly, CheCUDA is useful not only to enhance the dependability of CUDA applications but also to enable dynamic task scheduling of CUDA applications required especially on heterogeneous GPU cluster systems. This paper also shows the timing overhead for checkpointing.
Hiroyuki Takizawa, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi
PDCAT3
2006 Ray Tracing Hardware System Using Plane-Sphere Intersections
abstract
Ray tracing is a global illumination based rendering method widely used in computer graphics. Although it generates photo-realistic images, it requires a large number of computations. In ray tracing, the ray-object intersection test is one of the dominant factors for the processing speed. To accelerate the intersection test, we propose a new method based on a plane-sphere intersection algorithm, and show a hardware system using an FPGA. The computations used in the method are highly pipelined and parallelized by optimizing the balance between the computation speed and the memory data bandwidth. As a result, the prototype makes full use of 512 DSP cores built in Xilinx Vertex-4 SX FPGA, and the average utilization of the DSP cores is close to 90%. The simulation results show that the proposed system running at 160MHz performs the intersection test a few hundred times faster than a commodity PC with a 3.4GHz Pentium 4
Yoshiyuki Kaeriyama, Daichi Zaitsu, Kazuhiko Komatsu, Ken-Ichi Suzuki, Tadao Nakamura, Nobuyuki Ohba
FPL3
1987 The Outline Procedure in Pattern Data Preparation for Vector-Scan Electron-Beam Lithography
abstract
This paper describes a new algorithm and some applications of an outline procedure for LSI mask patterns. The outline procedure, which extracts the outline of designed primitive shapes, is required in pattern data preparation for electron-beam (e-beam) writing. A novel algorithm is developed to extract the outlines of patterns from a large number of designed shapes within a whole chip area. The algorithm is able to extract the outline with no limit to the number of vertices in a reasonable time. A program in which the algorithm is implemented is able to compensate for some process biases and leads to a successful 1-Mbit DRAM e-beam fabrication. Some other related applications, such as proximity effect correction, are also presented.
Kazuhiko Komatsu, Masanori Suzuki
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1