VLDB 2026 Research / reviewers in the wild / expert
Hayato Yamaki
dblp:192/7527
· DBLP profile ↗
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-2521-430XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 7 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comprehensive analysis of the impact of sub 10-nm CNFET technology on 64-bit parallel prefix adders and 32-bit matrix multiply units
Chenlin Shi, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
Integr. | 4 |
| 2025 | CACTI-CNFET: an Analytical Tool for Timing, Power, and Area of SRAMs with Carbon Nanotube Field Effect TransistorsabstractCarbon nanotube field effect transistors (CNFETs) are expected to replace silicon-based metal oxide semiconductor field effect transistors (MOSFETs) to improve the power efficiency and performance of microprocessors. However, the design of CNFET processors is not as mature as that of silicon-based processors because there is no architecture-level analytical tool for CNFET processors. Since circuit-level analysis such as RTL analysis is a very time-consuming and troublesome task, architecture-level analysis is needed for the rapid design of processors optimized for CNFETs. Shinobu Miwa, Eiichiro Sekikawa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 5 |
| 2025 | VAHRM: Variation-Aware Resource Management in Heterogeneous Supercomputing SystemsabstractIn this paper, we propose a novel resource management technique for heterogeneous supercomputing systems affected by manufacturing variability. Our proposed technique called VAHRM (Variation-Aware Heterogeneous Resource Management) takes a holistic approach to job scheduling on highly heterogeneous computing resources. VAHRM preferentially allocates energy-efficient computing resources to an energy-consuming job in a job queue, considering the impact on both the job turnaround time and the power consumption of individual resources. Furthermore, we have developed a novel approach to modeling the power consumption of computing resources that have manufacturing variability. Our approach called TSMVA (Two-Stage Modeling with Variation Awareness) enables us to generate the first variation-aware GPU power models, which can correctly estimate the power consumption of each GPU for a given job. Our experimental results show that, compared to conventional first-come-first-serve (FCFS) and state-of-the-art variation-aware scheduling algorithms, VAHRM can achieve respective improvements in system energy efficiency of up to 5.8% and 5.4% (4.5% and 4.2% on average) while reducing the average turnaround time of 21.2% and 11.9%, respectively, for various workloads obtained from a production system. Kohei Yoshida, Ryuichi Sakamoto, Kento Sato, Abhinav Bhatele, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Analyzing the impact of CUDA versions on GPU applicationsabstractCUDA toolkits are widely used to develop applications running on NVIDIA GPUs. They include compilers and are frequently updated to integrate state-of-the-art compilation techniques. Hence, many HPC users believe that the latest CUDA toolkit will improve application performance; however, considering results from CPU compilers, there are cases where this is not true. In this paper, we thoroughly evaluate the impact of CUDA toolkit version on the performance, power consumption, and energy consumption of GPU applications with three GPU architectures. Our results show that though the latest CUDA toolkit obtains the best performance, power consumption, and energy consumption for many applications in most cases, but we found a few exceptions. For such applications, we conducted an in-depth analysis using the SASS to identify why older CUDA toolkit achieve performance improvement. Our analysis showed that the factors that caused them are by three phenomena: aggressive loop unrolling, inefficient instruction scheduling, and the impact of host compilers. Kohei Yoshida, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
Parallel Comput. | 3 |
| 2023 | CNFET7: An Open Source Cell Library for 7-nm CNFET TechnologyabstractIn this paper, we propose CNFET7, the first open-source cell library for 7-nm carbon nanotube field-effect transistor (CNFET) technology. CNFET7 is based on an open-source CNFET SPICE model called VS-CNFET, and various model parameters such as the channel width and carbon nanotube diameter are carefully tuned to mimic the predictive 7-nm CNFET technology presented in a published paper. Some nondisclosure parameters, such as the cell size and pin layout, are derived from those of the NanGate 15-nm open-source cell library in the same way as for an open-source framework for CNFET circuit design. CNFET7 includes two types of delay model (i.e., the composite current source and nonlinear delay model), each having 56 cells, such as INV_X1 and BUF_X1. CNFET7 supports both logic synthesis and timing-driven place and route in the Cadence design flow. Our experimental results for several synthesized circuits show that CNFET7 has reductions of up to 96%, 62% and 82% in dynamic and static power consumption and critical-path delay, respectively, when compared with ASAP7. Chenlin Shi, Shinobu Miwa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 5 |
| 2022 | Analyzing Performance and Power-Efficiency Variations among NVIDIA GPUsabstractUnderstanding the variations in performance and power-efficiency of compute nodes is important for enhancing these factors in modern supercomputing systems. Previous studies have focused on variations in CPUs and DRAMs, but there has been little attention on GPUs. This is despite many current supercomputing systems employing GPUs (which consume a significant fraction of the power of such systems) as power-efficient accelerators for HPC applications. This paper describes the first thorough evaluation of performance and power-efficiency variations in GPUs. Specifically, we execute 48 CUDA kernels on 856 devices selected from three generations of NVIDIA GPUs (P100, V100, and A100), and analyze the impact of differences in both the CUDA kernels and GPU generation on performance and power-efficiency. Our analysis shows that there are non-negligible variations in both performance and power-efficiency, and that these variations are strongly affected by both the kernels that are running and the GPU generation. Kohei Yoshida, Rio Sageyama, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
ICPP | 4 |
| 2021 | Packet Forwarding Cache of Commodity Switches for Parallel ComputersabstractSwitch delay dominates communication latencies in interconnection networks, especially for short messages because switch delays are massive relative to the link and packet injection delays. At a conventional switch, routing decision is based on CAM (Content Addressable Memory) table lookup, and it imposes a significant delay. Reducing the access latency to CAM is crucial for the upcoming low-delay switch in parallel computers. Besides the CAM latency problem, the packet forwarding rate is not proportional to the switching capacity on cutting-edge commodity switches of interconnection networks. A switch will not able to forward incoming packets at the maximum line rate. To resolve the latency and throughput problems, we explore an on-chip packet forwarding cache to a switch. An incoming packet avoids large-latency accessing a CAM forwarding table if the cache hits. It supports an almost 100% hit rate (no capacity miss nor conflict miss) for packets generated in up to 2K-node jobs. For 100% hit rate on larger jobs, we present a switchable hash function to refer to a packet forwarding table on a switch. The switchable hash function is optimized to typical network topologies, i.e., k-ary n-cubes, fat trees, and Dragonfly. The main idea is that a large number of packet destinations share a same index tag, resulting in the same required number of cache entries as the number of output ports. This design can be enabled by the path regularity of the above network topologies. Our evaluation results show that the reasonable packet forwarding cache supports a 933-Gbps line rate even for incoming shortest packets on the above network topologies. We illustrate that parallel applications obtain the performance gain of 5.07x speed up using the cache switches since the impact of the switch delay and link bandwidth is significant on the end-to-end communication performance. Shoichi Hirasawa, Hayato Yamaki, Michihiro Koibuchi |
CLUSTER | 2 |
| 2020 | Evaluating architecture-level optimization in packet processing cachesabstractThe next-generation internet routers need ultra high speed such as 1 Tbps to process increased internet traffic. Efficient table lookup is key to realize high-speed packet processing and Packet Processing Cache (PPC) was proposed for this purpose. However, PPC has not been well optimized in the view of architecture and its ability to improve table lookup has therefore been underestimated so far. In this paper, we revisit PPC to reveal its real potential in table lookup. For this, we apply three architecture-level optimization techniques (hierarchization, pipelining and port addition), which are widely used to improve throughput of caches in microprocessors, and their combinations to PPC. Through the analysis of hundreds of billions of candidates of PPC configurations, we clarify the impact of these optimization techniques on the design space of internet routers with PPC. Our experimental results show that our best configuration can achieve 1045.7 Gbps packet processing with the power of 4273.3 mW and the area of 50.20 mm2, which is 3.35x improvement in Gbps per watt when compared to conventional PPC. Kyosuke Tanaka, Hayato Yamaki, Shinobu Miwa, Hiroki Honda |
Comput. Networks | 2 |
| 2020 | Footprint-Based DIMM HotplugabstractPower-efficiency has become one of the most critical concerns for HPC as we continue to scale computational capabilities. A significant fraction of system power is spent on large main memories, mainly caused by the substantial amount of DIMM standby power needed. However, while necessary for some workloads, for many workloads large memory configurations are too rich, i.e., these workloads only make use of a fraction of the available memory, causing unnecessary power usage. This observation opens new opportunities for power reduction by powering DIMMs on and off depending on the current workload. In this article, we propose footprint-based DIMM hotplug that enables a compute node to adjust the number of DIMMs that are powered on depending on the memory footprint of a running job. Our technique relies on two main subcomponents-memory footprint monitoring and DIMM management-which we both implement as part of an optimized page management system with small control overhead. Using Linux's memory hotplug capabilities, we implement our approach on a real system, and our results show that our proposed technique can save 50.6-52.1 percent of the DIMM standby energy and the CPU+DRAM energy of up to 1.50 Wh for various small-memory-footprint applications without loss of performance. Shinobu Miwa, Masaya Ishihara, Hayato Yamaki, Hiroki Honda, Martin Schulz 0001 |
IEEE Trans. Computers | 3 |
| 2018 | Data prediction for response flows in packet processing cacheabstractWe propose a technique to reduce compulsory misses of packet processing cache (PPC), which largely affects both throughput and energy of core routers. Rather than prefetching data, our technique called response prediction cache (RPC) speculatively stores predicted data into PPC without additional access to the low-throughput and power-consuming memory (i.e., TCAM). RPC predicts the data related to a response flow at the arrival of the corresponding request flow, based on the request-response model of internet communications. RPC can improve the cache miss rate, throughput, and energy-efficiency of PPC systems by 15.3%, 17.9%, and 17.8%, respectively. Hayato Yamaki, Hiroaki Nishi, Shinobu Miwa, Hiroki Honda |
DAC | 1 |