EDBT 2026 Demo / reviewers in the wild / expert
Yukinori Sato
dblp:78/6334
· DBLP profile ↗
9ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
autotuning |
0.4 | 1 | 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019 |
Compilers and program optimization
loop optimization |
0.4 | 1 | 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019 |
Compilers and program optimization › loop optimization
loop tiling |
0.4 | 1 | 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019 |
Compilers and program optimization › loop transformation
polyhedral compilation |
0.4 | 1 | 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral Compilation · ACM Trans. Archit. Code Optim. 2019 |
Methods — techniques the papers use, named apart from their topics
polyhedral compilation · 0.8auto-tuning · 0.8LLVM/Polly · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comprehensive Performance Evaluation of Microservices on Confidential Containers in Mutli-access Edge Computing EnvironmentsabstractMulti-access Edge Computing (MEC) in Edge-cloud computing continuum environments is an emerging technology that enables locality-aware low-latency processing for requests from edge devices. Confidential Containers (CoCo) are an emerging confidential computing technology designed for protecting in-use data in cloud datacenters. Although containers in a MEC layer can help reduce latency for requests from edge devices, the current increasing demands for security features such as confidential computing and memory integrity protection will seriously affect expected end-to-end latencies of microservice applications. In this paper, we build a CoCo environment using Kata containers on the MEC layer and attempt to evaluate microservice-level application behavior for the advanced security features. Especially, we focus on breaking down the latency into memory access level, microservice component level, and the end-to-end latency level. From the results, we observe that 9% overhead is incurred for each memory access when applying confidential computing with memory integrity protection. At the individual microservices level, we observed the latency for gRPC communication and memcached access are increased up to 40 %. We also reveal that the latency overheads in microservice-level and application-level are from twice to 4 times larger than those observed in memory access level. Itsuki Nakai, Takaaki Fukai, Takahiro Hirofuchi, Yukinori Sato |
IC2E | 4 |
| 2025 | A data-augmented model routing framework for efficient LLM deployment in edge-cloud environmentsabstractAbstract Large language model (LLM)-based program generation tasks are hindered by high computational demands. These challenges, along with high deployment costs, often pose a barrier to practical applications. To address these, we propose a novel data-augmented multi-LLM model routing approach that classifies prompts based on whether they should be processed on a weak LLM engine or a strong LLM. Experimental results show up to 16 times better efficiency compared to the existing cascaded approaches, while preserving the inference accuracy. Thus, the proposed method optimally allocates prompts across multiple LLMs, reducing computational costs while maintaining inference accuracy. Muhammad Syafiq Mohd Pozi, Yukinori Sato |
J. Supercomput. | 2 |
| 2022 | Apple Brand Texture Classification Using Neural Network Model
Shigeru Kato, Renon Toyosaki, Fuga Kitano, Shunsaku Kume, Naoki Wada, Tomomichi Kagawa, Takanori Hino, Kazuki Shiogai, Yukinori Sato, Muneyuki Unehara, Hajime Nobuhara |
AINA (3) | 9 |
| 2021 | Hodgkin-Huxley-Based Neural Simulation with Networks Connecting to Near-Neighbor NeuronsabstractNeural simulation is a very useful methodology for improving our understanding of the functions of brains, and expected to be applied to a number of practical applications related to autonomous systems and machines. However, computation time required for biophysically-meaningful simulation is often the limiting factor to mimic the large-scale neural network even if we implement custom hardware simulators on FPGAs. To overcome this issue, we focus on the spatial network connectivity of biological neurons and attempt to design a truly dataflow pipeline on FPGAs aiming at real-time Hodgkin-Huxley based simulation against more than 10,000 neurons. We implement our design on Xilinx Alveo U200 with HLS toolchains provided by Maxeler. From the results of evaluation, we find that mapping locality of neurons to the network connecting only to neighboring ones dramatically contributes to reducing the overheads for accumulation at the gap junction calculation. We demonstrate that our accelerator design can perform real-time simulation of network consisting 20736 neurons, which is 66.8 times larger than the existing implementation. Masashi Ogaki, Yukinori Sato |
ASAP | 2 |
| 2019 | An Autotuning Framework for Scalable Execution of Tiled Code via Iterative Polyhedral CompilationabstractOn modern many-core CPUs, performance tuning against complex memory subsystems and scalability for parallelism is mandatory to achieve their potential. In this article, we focus on loop tiling, which plays an important role in performance tuning, and develop a novel framework that analytically models the load balance and empirically autotunes unpredictable cache behaviors through iterative polyhedral compilation using LLVM/Polly. From an evaluation on many-core CPUs, we demonstrate that our autotuner achieves a performance superior to those that use conventional static approaches and well-known autotuning heuristics. Moreover, our autotuner achieves almost the same performance as a brute-force search-based approach. Yukinori Sato, Tomoya Yuki, Toshio Endo |
ACM Trans. Archit. Code Optim. | 1 |
| 2017 | An Accurate Simulator of Cache-Line Conflicts to Exploit the Underlying Cache Performance
Yukinori Sato, Toshio Endo |
Euro-Par | 1 |
| 2010 | A FPGA implementation of the two-dimensional Digital Huygens' ModelabstractModeling acoustical behavior in a room is complicated and computationally intense. Many methods have been proposed to analyze the distribution of sound field by using computer simulation. However, the procedure is time-consuming with sound space increasing. In this paper, a hardware solution based on Digital Huygens' Model (DHM) and Field Programmable Gate Array (FPGA) technology is proposed to simulate and rebuild the distribution of sound field in a room. Two schemes of DHM are derived to analyze the sound propagation in a 2D space and implemented by FPGA. In a 2D space with 35 × 35 nodes and surrounded by rigid walls, the results got by hardware meet well with the analytical results in case of different incidences except having three-cycle delays. The designed hardware system consumes about 0.016s to process the computations during 10000 time steps while the software solution developed by C++ programming language costs about 0.14s. The hardware implementation occupies 76% of LUTs and 39% of FDCs in a Xilinx FPGA chip XC5VLX330T-FF1738. Tan Yiyu, Yukinori Sato, Eiko Sugawara, Yasushi Inoguchi, Makoto Otani, Yukio Iwaya, Hiroshi Matsuoka, Takao Tsuchiya |
FPT | 2 |
| 2009 | Improving accuracy of host load predictions on computational grids by artificial neural networksabstractThe capability to predict the host load of a system is significant for computational grids to make efficient use of shared resources. This paper attempts to improve the accuracy of host load predictions by applying a neural network predictor to reach the goal of best performance and load balance. We describe feasibility of the proposed predictor in a dynamic environment, and perform experimental evaluation using collected load traces. The results show that the neural network achieves a consistent performance improvement with surprisingly low overhead. Compared with the best previously proposed method, the typical 20:10:1 network reduces the mean and standard deviation of the prediction errors by approximately 60% and 70%, respectively. The training and testing time is extremely low, as this network needs only a couple of seconds to be trained with more than 100,000 samples in order to make tens of thousands of accurate predictions within just a second. Truong Vinh Truong Duy, Yukinori Sato, Yasushi Inoguchi |
IPDPS | 2 |
| 2005 | Cooperation of Neighboring PEs in Clustered ArchitecturesabstractClustered architectures which intend to process data within a localized PE are one of the approaches to increase the performance under the difficulties of the wire delay problems. The performance of clustered architectures depends on the amount of parallel execution of instructions and the amount of inter-PE communication to synchronize dependent instructions. In this paper, we propose an arrangement of PEs cooperating with the adjacent PEs by means of adding communication structures between the adjacent PEs in order to relax the inter-PE communication and workload imbalance in an effective manner. We evaluate the proposed configurations and compare them with the existing one so far considered. The results show that the proposed adjacent forwarding network configuration with the instruction steering scheme that concerns both the register fanout and available free register can achieve higher instructions per clock (IPC) with the small number of registers per PE than the other configurations. Yukinori Sato, Ken-Ichi Suzuki, Tadao Nakamura |
SBAC-PAD | 1 |