EDBT 2026 Demo / reviewers in the wild / expert
Hiroki Honda
dblp:36/539
· DBLP profile ↗
20ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-6257-7380ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 since 2021Software engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comprehensive analysis of the impact of sub 10-nm CNFET technology on 64-bit parallel prefix adders and 32-bit matrix multiply units
Chenlin Shi, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
Integr. | 5 |
| 2025 | CACTI-CNFET: an Analytical Tool for Timing, Power, and Area of SRAMs with Carbon Nanotube Field Effect TransistorsabstractCarbon nanotube field effect transistors (CNFETs) are expected to replace silicon-based metal oxide semiconductor field effect transistors (MOSFETs) to improve the power efficiency and performance of microprocessors. However, the design of CNFET processors is not as mature as that of silicon-based processors because there is no architecture-level analytical tool for CNFET processors. Since circuit-level analysis such as RTL analysis is a very time-consuming and troublesome task, architecture-level analysis is needed for the rapid design of processors optimized for CNFETs. Shinobu Miwa, Eiichiro Sekikawa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 6 |
| 2025 | VAHRM: Variation-Aware Resource Management in Heterogeneous Supercomputing SystemsabstractIn this paper, we propose a novel resource management technique for heterogeneous supercomputing systems affected by manufacturing variability. Our proposed technique called VAHRM (Variation-Aware Heterogeneous Resource Management) takes a holistic approach to job scheduling on highly heterogeneous computing resources. VAHRM preferentially allocates energy-efficient computing resources to an energy-consuming job in a job queue, considering the impact on both the job turnaround time and the power consumption of individual resources. Furthermore, we have developed a novel approach to modeling the power consumption of computing resources that have manufacturing variability. Our approach called TSMVA (Two-Stage Modeling with Variation Awareness) enables us to generate the first variation-aware GPU power models, which can correctly estimate the power consumption of each GPU for a given job. Our experimental results show that, compared to conventional first-come-first-serve (FCFS) and state-of-the-art variation-aware scheduling algorithms, VAHRM can achieve respective improvements in system energy efficiency of up to 5.8% and 5.4% (4.5% and 4.2% on average) while reducing the average turnaround time of 21.2% and 11.9%, respectively, for various workloads obtained from a production system. Kohei Yoshida, Ryuichi Sakamoto, Kento Sato, Abhinav Bhatele, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | Analyzing the impact of CUDA versions on GPU applicationsabstractCUDA toolkits are widely used to develop applications running on NVIDIA GPUs. They include compilers and are frequently updated to integrate state-of-the-art compilation techniques. Hence, many HPC users believe that the latest CUDA toolkit will improve application performance; however, considering results from CPU compilers, there are cases where this is not true. In this paper, we thoroughly evaluate the impact of CUDA toolkit version on the performance, power consumption, and energy consumption of GPU applications with three GPU architectures. Our results show that though the latest CUDA toolkit obtains the best performance, power consumption, and energy consumption for many applications in most cases, but we found a few exceptions. For such applications, we conducted an in-depth analysis using the SASS to identify why older CUDA toolkit achieve performance improvement. Our analysis showed that the factors that caused them are by three phenomena: aggressive loop unrolling, inefficient instruction scheduling, and the impact of host compilers. Kohei Yoshida, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
Parallel Comput. | 4 |
| 2023 | CNFET7: An Open Source Cell Library for 7-nm CNFET TechnologyabstractIn this paper, we propose CNFET7, the first open-source cell library for 7-nm carbon nanotube field-effect transistor (CNFET) technology. CNFET7 is based on an open-source CNFET SPICE model called VS-CNFET, and various model parameters such as the channel width and carbon nanotube diameter are carefully tuned to mimic the predictive 7-nm CNFET technology presented in a published paper. Some nondisclosure parameters, such as the cell size and pin layout, are derived from those of the NanGate 15-nm open-source cell library in the same way as for an open-source framework for CNFET circuit design. CNFET7 includes two types of delay model (i.e., the composite current source and nonlinear delay model), each having 56 cells, such as INV_X1 and BUF_X1. CNFET7 supports both logic synthesis and timing-driven place and route in the Cadence design flow. Our experimental results for several synthesized circuits show that CNFET7 has reductions of up to 96%, 62% and 82% in dynamic and static power consumption and critical-path delay, respectively, when compared with ASAP7. Chenlin Shi, Shinobu Miwa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 6 |
| 2022 | Analyzing Performance and Power-Efficiency Variations among NVIDIA GPUsabstractUnderstanding the variations in performance and power-efficiency of compute nodes is important for enhancing these factors in modern supercomputing systems. Previous studies have focused on variations in CPUs and DRAMs, but there has been little attention on GPUs. This is despite many current supercomputing systems employing GPUs (which consume a significant fraction of the power of such systems) as power-efficient accelerators for HPC applications. This paper describes the first thorough evaluation of performance and power-efficiency variations in GPUs. Specifically, we execute 48 CUDA kernels on 856 devices selected from three generations of NVIDIA GPUs (P100, V100, and A100), and analyze the impact of differences in both the CUDA kernels and GPU generation on performance and power-efficiency. Our analysis shows that there are non-negligible variations in both performance and power-efficiency, and that these variations are strongly affected by both the kernels that are running and the GPU generation. Kohei Yoshida, Rio Sageyama, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
ICPP | 5 |
| 2020 | Evaluating architecture-level optimization in packet processing cachesabstractThe next-generation internet routers need ultra high speed such as 1 Tbps to process increased internet traffic. Efficient table lookup is key to realize high-speed packet processing and Packet Processing Cache (PPC) was proposed for this purpose. However, PPC has not been well optimized in the view of architecture and its ability to improve table lookup has therefore been underestimated so far. In this paper, we revisit PPC to reveal its real potential in table lookup. For this, we apply three architecture-level optimization techniques (hierarchization, pipelining and port addition), which are widely used to improve throughput of caches in microprocessors, and their combinations to PPC. Through the analysis of hundreds of billions of candidates of PPC configurations, we clarify the impact of these optimization techniques on the design space of internet routers with PPC. Our experimental results show that our best configuration can achieve 1045.7 Gbps packet processing with the power of 4273.3 mW and the area of 50.20 mm2, which is 3.35x improvement in Gbps per watt when compared to conventional PPC. Kyosuke Tanaka, Hayato Yamaki, Shinobu Miwa, Hiroki Honda |
Comput. Networks | 4 |
| 2020 | Footprint-Based DIMM HotplugabstractPower-efficiency has become one of the most critical concerns for HPC as we continue to scale computational capabilities. A significant fraction of system power is spent on large main memories, mainly caused by the substantial amount of DIMM standby power needed. However, while necessary for some workloads, for many workloads large memory configurations are too rich, i.e., these workloads only make use of a fraction of the available memory, causing unnecessary power usage. This observation opens new opportunities for power reduction by powering DIMMs on and off depending on the current workload. In this article, we propose footprint-based DIMM hotplug that enables a compute node to adjust the number of DIMMs that are powered on depending on the memory footprint of a running job. Our technique relies on two main subcomponents-memory footprint monitoring and DIMM management-which we both implement as part of an optimized page management system with small control overhead. Using Linux's memory hotplug capabilities, we implement our approach on a real system, and our results show that our proposed technique can save 50.6-52.1 percent of the DIMM standby energy and the CPU+DRAM energy of up to 1.50 Wh for various small-memory-footprint applications without loss of performance. Shinobu Miwa, Masaya Ishihara, Hayato Yamaki, Hiroki Honda, Martin Schulz 0001 |
IEEE Trans. Computers | 4 |
| 2018 | Data prediction for response flows in packet processing cacheabstractWe propose a technique to reduce compulsory misses of packet processing cache (PPC), which largely affects both throughput and energy of core routers. Rather than prefetching data, our technique called response prediction cache (RPC) speculatively stores predicted data into PPC without additional access to the low-throughput and power-consuming memory (i.e., TCAM). RPC predicts the data related to a response flow at the arrival of the corresponding request flow, based on the request-response model of internet communications. RPC can improve the cache miss rate, throughput, and energy-efficiency of PPC systems by 15.3%, 17.9%, and 17.8%, respectively. Hayato Yamaki, Hiroaki Nishi, Shinobu Miwa, Hiroki Honda |
DAC | 4 |
| 2009 | Aspects of GPU for general purpose high performance computingabstractWe discuss hardware and software aspects of GPGPU, specifically focusing on NVIDIA cards and CUDA, from the viewpoints of parallel computing. The major weak points of GPU against newest supercomputers are identified to be and summarized as only four points: large SIMD vector length, small memory, absence of fast L2 cache, and high register spill penalty. As software concerns, we derive optimal scheduling algorithm for latency hiding of host-device data transfer, and discuss SPMD parallelism on GPUs. Reiji Suda, Takayuki Aoki, Shoichi Hirasawa, Akira Nukada, Hiroki Honda, Satoshi Matsuoka |
ASP-DAC | 5 |
| 2007 | F-Omega: A Framework for Steering GridRPC ApplicationsabstractSteering grid applications is needed so that they can run several days or several weeks without restarting their computation. Existing grid middleware, such as GridRPC middleware, have room for improvement in point of steering grid applications. For instance, to manage GridRPCrelated resources remains as a complicated task for programmers. And, to monitor constraint information about future server usage is a laborious task for application users. GridRPC middleware has to assist in these tasks. In this paper, we propose a framework F-Omega for these tasks. F-Omega provides programmers common application modules for automatic management of GridRPC-related resources. For users, F-Omega provides an automatic visualization system for constraint information about future server usage. Experimental results show that the proposed features of F-Omega mitigate programmers' and users' burdens of the above-mentioned tasks. Hiromasa Watanabe, Shoichi Hirasawa, Hiroki Honda |
eScience | 3 |
| 2006 | ABCLibScript: a directive to support specification of an auto-tuning facility for numerical software
Takahiro Katagiri, Kenji Kise, Hiroki Honda, Toshitsugu Yuba |
Parallel Comput. | 3 |
| 2006 | ABCLib_DRSSED: A parallel eigensolver with an auto-tuning facility
Takahiro Katagiri, Kenji Kise, Hiroki Honda, Toshitsugu Yuba |
Parallel Comput. | 3 |
| 2005 | Macro-Dataflow using Software Distributed Shared MemoryabstractMacro-dataflow processing, which exploits the parallelism among coarse-grain tasks (macrotasks) such as loops and subroutines, is considered promising to break the performance limits of loop parallelism. To realize macro-dataflow processing on distributed memory systems, "data reaching conditions", a method to make the sender-receiver pair of a data transfer determined at runtime, has previously been proposed. However, irregular data accesses induce extra data transfers, which lead to performance deterioration. This paper proposes an implementation method using software distributed shared memory, which enables on-demand data fetching. This paper describes the implementation using two well-accepted, page-based software distributed shared memory systems, TreadMarks and JI-AJIA. Evaluation results on a PC cluster show the software distributed memory approach is as much as 25% faster than the data reaching conditions Hiroshi Tanabe, Hiroki Honda, Toshitsugu Yuba |
CLUSTER | 2 |
| 2005 | Towards scalable and simple software-DSM systemsabstractLarge-scale PC clusters, such as ACSI Lightning with 2,816 processors and the AIST supercluster with more than 3,000 processors, have been constructed. Although software distributed shared memory (S-DSM) provides an attractive parallel programming model, almost all S-DSM systems proposed are only useful on a cluster of less than or equal to 16 nodes. They lack scalability. Kenji Kise, Takahiro Katagiri, Hiroki Honda, Toshitsugu Yuba |
SOSP | 3 |
| 2005 | A time-to-live based reservation algorithm on fully decentralized resource discovery in Grid computing
Sanya Tangpongprasit, Takahiro Katagiri, Kenji Kise, Hiroki Honda, Toshitsugu Yuba |
Parallel Comput. | 4 |
| 2000 | Scalable Data Mining with Log Based Consistency DSM for High Performance Distributed ComputingabstractMining the large Web based online distributed databases to discover new knowledge and financial gain is an important research problem. These computations require high performance distributed and parallel computing environments. Traditional data mining techniques such as classification, association, clustering can be extended to find new efficient solutions. The paper presents the scalable data mining problem, proposes the use of software DSM (distributed shared memory) with a new mechanism as an effective solution and discusses both the implementation and performance evaluation results. It is observed that the overhead of a software DSM is very large for scalable data mining programs. A new Log Based Consistency (LBC) mechanism, especially designed for scalable data mining on the software DSM is proposed to overcome this overhead. Traditional association rule based data mining programs frequently modify the same fields by count-up operations. In contrast, the LBC mechanism keeps up the consistency by broadcasting the count-up operation logs among the multiple nodes. Hideaki Hirayama, Hiroki Honda, Toshitsugu Yuba |
ICECCS | 2 |
| 1999 | Distributed Shared Memory with Log Based Consistency for Scalable Data MiningabstractThe paper presents the scalable data mining problem, proposes the use of software DSM (Distributed Shared Memory) with a new mechanism as an effective solution and discusses both the implementation and performance evaluation results. It is observed that the overhead of a software DSM is very large for scalable data mining programs. A new Log Based Consistency (LBC) mechanism, especially designed for scalable data mining on the software DSM is proposed to overcome this overhead. Traditional association rule based data mining programs frequently modify the same fields by count-up operations. In contrast, the LBC mechanism keeps up the consistency by broadcasting the count-up operation logs among the multiple nodes. Hideaki Hirayama, Hiroki Honda, Toshitsugu Yuba |
COMPSAC | 2 |
| 1990 | A Compilation Scheme for Macro-Dataflow Computation on Hierarchical Multiprocessor Systems
Hironori Kasahara, Hiroki Honda, Masahiko Iwata, M. Hirota |
ICPP (2) | 2 |
| 1990 | Parallel processing of near fine grain tasks using static scheduling OSCAR (optimally scheduled advanced multiprocessor)abstractThe authors propose a compilation scheme for parallel processing near fine-grain tasks, each of which consists of several instructions or a statement, on a multiprocessor system called OSCAR. The scheme allows one to minimize synchronization and data transfer overheads and to optimally use registers of each processor by employing a static scheduling algorithm considering data transfer. This scheme can effectively be combined with macro-dataflow computation and with making the loop concurrent. A compiler using the proposed scheme has been implemented on OSCAR, which has been designed to take full advantage of the static scheduling. A performance evaluation of the scheme on OSCAR is also described.> Hironori Kasahara, Hiroki Honda, Seinosuke Narita |
SC | 2 |