Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yukinori Sato

dblp:78/6334 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
autotuning
0.412019
An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019
Compilers and program optimization
loop optimization
0.412019
An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019
Compilers and program optimization › loop optimization
loop tiling
0.412019
An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019
Compilers and program optimization › loop transformation
polyhedral compilation
0.412019
An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019
Processor architecture and microarchitecture
many-core architecture
0.112019
An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019

Methods — techniques the papers use, named apart from their topics

polyhedral compilation · 0.8auto-tuning · 0.8LLVM/Polly · 0.8
YearPublicationVenuePosition
2025 Comprehensive Performance Evaluation of Microservices on Confidential Containers in Mutli-access Edge Computing Environments
abstract
Multi-access Edge Computing (MEC) in Edge-cloud computing continuum environments is an emerging technology that enables locality-aware low-latency processing for requests from edge devices. Confidential Containers (CoCo) are an emerging confidential computing technology designed for protecting in-use data in cloud datacenters. Although containers in a MEC layer can help reduce latency for requests from edge devices, the current increasing demands for security features such as confidential computing and memory integrity protection will seriously affect expected end-to-end latencies of microservice applications. In this paper, we build a CoCo environment using Kata containers on the MEC layer and attempt to evaluate microservice-level application behavior for the advanced security features. Especially, we focus on breaking down the latency into memory access level, microservice component level, and the end-to-end latency level. From the results, we observe that 9% overhead is incurred for each memory access when applying confidential computing with memory integrity protection. At the individual microservices level, we observed the latency for gRPC communication and memcached access are increased up to 40 %. We also reveal that the latency overheads in microservice-level and application-level are from twice to 4 times larger than those observed in memory access level.
Itsuki Nakai, Takaaki Fukai, Takahiro Hirofuchi, Yukinori Sato
IC2E4
2025 A data-augmented model routing framework for efficient LLM deployment in edge-cloud environments
abstract
Abstract Large language model (LLM)-based program generation tasks are hindered by high computational demands. These challenges, along with high deployment costs, often pose a barrier to practical applications. To address these, we propose a novel data-augmented multi-LLM model routing approach that classifies prompts based on whether they should be processed on a weak LLM engine or a strong LLM. Experimental results show up to 16 times better efficiency compared to the existing cascaded approaches, while preserving the inference accuracy. Thus, the proposed method optimally allocates prompts across multiple LLMs, reducing computational costs while maintaining inference accuracy.
Muhammad Syafiq Mohd Pozi, Yukinori Sato
J. Supercomput.2
2022 Apple Brand Texture Classification Using Neural Network Model
Shigeru Kato, Renon Toyosaki, Fuga Kitano, Shunsaku Kume, Naoki Wada, Tomomichi Kagawa, Takanori Hino, Kazuki Shiogai, Yukinori Sato, Muneyuki Unehara, Hajime Nobuhara
AINA (3)9
2021 Hodgkin-Huxley-Based Neural Simulation with Networks Connecting to Near-Neighbor Neurons
abstract
Neural simulation is a very useful methodology for improving our understanding of the functions of brains, and expected to be applied to a number of practical applications related to autonomous systems and machines. However, computation time required for biophysically-meaningful simulation is often the limiting factor to mimic the large-scale neural network even if we implement custom hardware simulators on FPGAs. To overcome this issue, we focus on the spatial network connectivity of biological neurons and attempt to design a truly dataflow pipeline on FPGAs aiming at real-time Hodgkin-Huxley based simulation against more than 10,000 neurons. We implement our design on Xilinx Alveo U200 with HLS toolchains provided by Maxeler. From the results of evaluation, we find that mapping locality of neurons to the network connecting only to neighboring ones dramatically contributes to reducing the overheads for accumulation at the gap junction calculation. We demonstrate that our accelerator design can perform real-time simulation of network consisting 20736 neurons, which is 66.8 times larger than the existing implementation.
Masashi Ogaki, Yukinori Sato
ASAP2
2019 An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation
abstract
On modern many-core CPUs, performance tuning against complex memory subsystems and scalability for parallelism is mandatory to achieve their potential. In this article, we focus on loop tiling, which plays an important role in performance tuning, and develop a novel framework that analytically models the load balance and empirically autotunes unpredictable cache behaviors through iterative polyhedral compilation using LLVM/Polly. From an evaluation on many-core CPUs, we demonstrate that our autotuner achieves a performance superior to those that use conventional static approaches and well-known autotuning heuristics. Moreover, our autotuner achieves almost the same performance as a brute-force search-based approach.
Yukinori Sato, Tomoya Yuki, Toshio Endo
ACM Trans. Archit. Code Optim.1
2017 An Accurate Simulator of Cache-Line Conflicts to Exploit the Underlying Cache Performance
Yukinori Sato, Toshio Endo
Euro-Par1
2010 A FPGA implementation of the two-dimensional Digital Huygens' Model
abstract
Modeling acoustical behavior in a room is complicated and computationally intense. Many methods have been proposed to analyze the distribution of sound field by using computer simulation. However, the procedure is time-consuming with sound space increasing. In this paper, a hardware solution based on Digital Huygens' Model (DHM) and Field Programmable Gate Array (FPGA) technology is proposed to simulate and rebuild the distribution of sound field in a room. Two schemes of DHM are derived to analyze the sound propagation in a 2D space and implemented by FPGA. In a 2D space with 35 × 35 nodes and surrounded by rigid walls, the results got by hardware meet well with the analytical results in case of different incidences except having three-cycle delays. The designed hardware system consumes about 0.016s to process the computations during 10000 time steps while the software solution developed by C++ programming language costs about 0.14s. The hardware implementation occupies 76% of LUTs and 39% of FDCs in a Xilinx FPGA chip XC5VLX330T-FF1738.
Tan Yiyu, Yukinori Sato, Eiko Sugawara, Yasushi Inoguchi, Makoto Otani, Yukio Iwaya, Hiroshi Matsuoka, Takao Tsuchiya
FPT2
2009 Improving accuracy of host load predictions on computational grids by artificial neural networks
abstract
The capability to predict the host load of a system is significant for computational grids to make efficient use of shared resources. This paper attempts to improve the accuracy of host load predictions by applying a neural network predictor to reach the goal of best performance and load balance. We describe feasibility of the proposed predictor in a dynamic environment, and perform experimental evaluation using collected load traces. The results show that the neural network achieves a consistent performance improvement with surprisingly low overhead. Compared with the best previously proposed method, the typical 20:10:1 network reduces the mean and standard deviation of the prediction errors by approximately 60% and 70%, respectively. The training and testing time is extremely low, as this network needs only a couple of seconds to be trained with more than 100,000 samples in order to make tens of thousands of accurate predictions within just a second.
Truong Vinh Truong Duy, Yukinori Sato, Yasushi Inoguchi
IPDPS2
2005 Cooperation of Neighboring PEs in Clustered Architectures
abstract
Clustered architectures which intend to process data within a localized PE are one of the approaches to increase the performance under the difficulties of the wire delay problems. The performance of clustered architectures depends on the amount of parallel execution of instructions and the amount of inter-PE communication to synchronize dependent instructions. In this paper, we propose an arrangement of PEs cooperating with the adjacent PEs by means of adding communication structures between the adjacent PEs in order to relax the inter-PE communication and workload imbalance in an effective manner. We evaluate the proposed configurations and compare them with the existing one so far considered. The results show that the proposed adjacent forwarding network configuration with the instruction steering scheme that concerns both the register fanout and available free register can achieve higher instructions per clock (IPC) with the small number of registers per PE than the other configurations.
Yukinori Sato, Ken-Ichi Suzuki, Tadao Nakamura
SBAC-PAD1