Masayuki Sato 0001

dblp:27/617-1 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-4186-5014ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adaptive Parallelization Based on Frame-Level and Tile-Level Parallelisms for VVC Encoding
abstract
ABSTRACT To meet the growing demand for high‐efficiency video compression standards, Versatile Video Coding (VVC) has been developed as a successor to High Efficiency Video Coding (HEVC). VVC offers approximately a 50% reduction in bitrate compared to HEVC while maintaining comparable visual quality. However, the improved performance of VVC comes at the cost of a significantly higher computational complexity, resulting in longer encoding times. Consequently, accelerating the VVC encoding process remains a critical challenge for its practical deployment. Leveraging the evolution of multi‐core processors and encoding tools designed for parallel processing, this study introduces an adaptive parallelization method that integrates frame‐level and tile‐level parallelisms. This method dynamically selects the number of concurrently processed frames and tiles by considering reference dependencies and their effects on coding efficiency. Moreover, the proposed method employs a content‐aware, cyclic adjustment of tile configurations to further reduce the encoding time. The evaluation results demonstrate that the proposed method can achieve a 10.58× speedup over the baseline single‐threaded implementation on average, while limiting the increase in BD‐BR to 3.23% and the degradation in BD‐PSNR to only −0.058 dB. These findings confirm that the method substantially decreases the encoding time without degrading coding efficiency. Furthermore, the results also demonstrate that the scalability of the proposed method is better than that of the conventional parallel method, Wavefront parallel processing.
Karin Onouchi, Masayuki Sato 0001, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi
Concurr. Comput. Pract. Exp.2
2023 Performance Evaluation of Tsunami Evacuation Route Planning on Multiple Annealing Machines
abstract
This paper focuses on the performance evaluation of annealing machines based on quantum annealing and simulated annealing, i.e., D-wave advantage, Fixstars Amplify Annealing Engine, D-wave Neal, and Vector Annealing, by solving a combinatorial optimization problem, tsunami evacuation route planning. First, the problem is modeled into mathematical formulas and converted into the Quadratic Unconstrained Binary Optimization (QUBO) form as an input, and will be executed on the four annealing machines. Then, the relationship between constraint weight in the objective function and the constraint compliance conditions is also discussed. Finally, the evaluation and discussion are conducted from three aspects, the maximum number of evacuees, Time to Solution (TTS), and the solutions quality.
Yihui Liu, Kazuhiko Komatsu, Masahito Kumagai, Masayuki Sato 0001, Hiroaki Kobayashi
CF4
2023 Multi-scale Loss based Electron Microscopic Image Pair Matching Method
abstract
Nanodiffraction Imaging (NDI), a novel imaging technique based on the scanning transmission electron mi-croscopy (STEM), helps the understanding of the relationships between micro-structure and macro-properties. However, the analysis requires image pair matching tasks through meticulous and time-consuming observation and selection by domain experts. Therefore, this paper proposes an image pair matching method for NDI images. The proposed method adopts a two-step training approach. The first step is to perform pre-training by contrastive learning on specialized NDI images. The second step is fine-tuning by a customized model on image pair matching tasks. In the second step, the training is performed with the incorporation of a special loss function, OriDist loss. This loss function is designed to focus on orientation and distribution of multi-scale features. The evaluation results demonstrate the ability of the proposed method to achieve high-accuracy NDI image pair matching, and efficiently reduce the search space of candidate matching images, resulting in a significant reduction in the human workload. Through an extensive ablation study, each component of the proposed method shows positive contributions to the overall performance.
Chunting Duan, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
ICMLA3
2023 A dynamic parameter tuning method for SpMM parallel execution
abstract
Summary Sparse matrix‐matrix multiplication (SpMM) is a basic kernel that is used by many algorithms. Several researches focus on various optimizations for SpMM parallel execution. However, a division of a task for parallelization is not well considered yet. Generally, a matrix is equally divided into blocks for processes even though the sparsities of input matrices are different. The parameter that divides a task into multiple processes for parallelization is fixed. As a result, load imbalance among the processes occurs. To balance the loads among the processes, this article proposes a dynamic parameter tuning method by analyzing the sparsities of input matrices. The experimental results show that the proposed method improves the performance of SpMM for examined matrices by up to 39.5% on a single vector engine and 3.49 on a single CPU.
Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
Concurr. Comput. Pract. Exp.3
2022 A Partitioned Memory Architecture with Prefetching for Efficient Video Encoders
Masayuki Sato 0001, Yuya Omori, Ryusuke Egawa, Ken Nakamura, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi
PDCAT1
2021 Register Flush-free Runahead Execution for Modern Vector Processors
abstract
Modern vector processors have been designed to achieve high sustained performance, especially in HPC applications, because of their powerful instruction set oriented to data-level parallelism. Additionally, the latest vector processor adopts the out-of-order execution of the vector instructions to exploit instruction-level parallelism due to a significant gap in latency between vector arithmetic instructions and vector load/store instructions. In spite of the effort, this gap still brings a deterioration of sustained performance of the modern vector processors. This paper proposes a runahead execution mechanism for the modern vector processors to fill the latency gap by further exploiting instruction-level parallelism. If the processor stalls due to a long latency instruction, the conventional runahead execution mechanism changes the processor state from a normal mode to a runahead mode, and the processor speculatively executes the subsequent instructions that can cause stalls and their dependencies. However, the conventional runahead execution mechanisms flush the registers' values calculated in the runahead mode after finishing this mode and cannot reuse them in the subsequent normal mode. Since the vector processors have many values even in one vector register, these flushes and re-executions waste the bandwidth between cores and caches. Thus, to solve this problem of the conventional runahead mechanism, our proposed mechanism leaves the registers containing the results in the runahead mode in order for the processor to use the registers even after returning to the normal mode. For correctly using these registers after exiting the runahead mode, the proposed mechanism newly realizes functions to inherit the commit order information and the register aliasing information of the runahead-executed instructions into the normal mode. The evaluation results show that the proposed mechanism improves the performance by up to 20% and 3% on average by the conventional mechanism.
Hikaru Takayashiki, Masayuki Sato 0001, Kazuhiko Komatsu, Hiroaki Kobayashi
SBAC-PAD2
2020 A Dynamic Parameter Tuning Method for High Performance SpMM
Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi
PDCAT3
2018 Performance evaluation of a vector supercomputer SX-aurora TSUBASA
Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Mitsuo Yokokawa, Toshikazu Aoyama, Masayuki Sato 0001, Hiroaki Kobayashi
SC8
2010 A voting-based working set assessment scheme for dynamic cache resizing mechanisms
abstract
Considering the trade-off between performance and power consumption has become significantly important in multi-core processor design. Under this situation, one promising approach is to employ a power-aware dynamic cache partitioning mechanism. This mechanism individually manages activation of each cache way, and exclusively allocates the minimum number of required ways to each thread. In the mechanism, an appropriate number of ways for a thread is decided based on locality assessment. However, sampling results of cache accesses that are used for locality assessment are disturbed by exceptional behaviors of cache accesses, which happen in a very short period. Such sampling results may change locality assessment results to ones that are not along with the overall trend in a long access-sampling period. These assessment results will excessively adapt the cache to exceptional behaviors, and deteriorate energy efficiency. To avoid such excessive adaptation by the exceptional behaviors, this paper proposes a voting-based working set assessment scheme, in which the number of activated ways is adjusted based on majority voting of locality assessment of several short sampling periods. By using the majority voting, the proposed scheme can identify the periods including exceptional behaviors, and ignore the assessment results of these periods. As a result, the proposed scheme makes the cache resizing mechanism more stable and robust. The experimental results indicate that the proposed scheme can reduce energy consumption by up to 24%, and 10% on an average without significant performance degradation in multi-thread execution on a 2-core CMP.
Masayuki Sato 0001, Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi
ICCD1