EDBT 2026 Demo / reviewers in the wild / expert
Shinobu Miwa
dblp:74/6159
· DBLP profile ↗
23ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-0315-3216ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comprehensive analysis of the impact of sub 10-nm CNFET technology on 64-bit parallel prefix adders and 32-bit matrix multiply units
Chenlin Shi, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
Integr. | 6 |
| 2025 | CACTI-CNFET: an Analytical Tool for Timing, Power, and Area of SRAMs with Carbon Nanotube Field Effect TransistorsabstractCarbon nanotube field effect transistors (CNFETs) are expected to replace silicon-based metal oxide semiconductor field effect transistors (MOSFETs) to improve the power efficiency and performance of microprocessors. However, the design of CNFET processors is not as mature as that of silicon-based processors because there is no architecture-level analytical tool for CNFET processors. Since circuit-level analysis such as RTL analysis is a very time-consuming and troublesome task, architecture-level analysis is needed for the rapid design of processors optimized for CNFETs. Shinobu Miwa, Eiichiro Sekikawa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 1 |
| 2025 | VAHRM: Variation-Aware Resource Management in Heterogeneous Supercomputing SystemsabstractIn this paper, we propose a novel resource management technique for heterogeneous supercomputing systems affected by manufacturing variability. Our proposed technique called VAHRM (Variation-Aware Heterogeneous Resource Management) takes a holistic approach to job scheduling on highly heterogeneous computing resources. VAHRM preferentially allocates energy-efficient computing resources to an energy-consuming job in a job queue, considering the impact on both the job turnaround time and the power consumption of individual resources. Furthermore, we have developed a novel approach to modeling the power consumption of computing resources that have manufacturing variability. Our approach called TSMVA (Two-Stage Modeling with Variation Awareness) enables us to generate the first variation-aware GPU power models, which can correctly estimate the power consumption of each GPU for a given job. Our experimental results show that, compared to conventional first-come-first-serve (FCFS) and state-of-the-art variation-aware scheduling algorithms, VAHRM can achieve respective improvements in system energy efficiency of up to 5.8% and 5.4% (4.5% and 4.2% on average) while reducing the average turnaround time of 21.2% and 11.9%, respectively, for various workloads obtained from a production system. Kohei Yoshida, Ryuichi Sakamoto, Kento Sato, Abhinav Bhatele, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | Analyzing the impact of CUDA versions on GPU applicationsabstractCUDA toolkits are widely used to develop applications running on NVIDIA GPUs. They include compilers and are frequently updated to integrate state-of-the-art compilation techniques. Hence, many HPC users believe that the latest CUDA toolkit will improve application performance; however, considering results from CPU compilers, there are cases where this is not true. In this paper, we thoroughly evaluate the impact of CUDA toolkit version on the performance, power consumption, and energy consumption of GPU applications with three GPU architectures. Our results show that though the latest CUDA toolkit obtains the best performance, power consumption, and energy consumption for many applications in most cases, but we found a few exceptions. For such applications, we conducted an in-depth analysis using the SASS to identify why older CUDA toolkit achieve performance improvement. Our analysis showed that the factors that caused them are by three phenomena: aggressive loop unrolling, inefficient instruction scheduling, and the impact of host compilers. Kohei Yoshida, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
Parallel Comput. | 2 |
| 2023 | CNFET7: An Open Source Cell Library for 7-nm CNFET TechnologyabstractIn this paper, we propose CNFET7, the first open-source cell library for 7-nm carbon nanotube field-effect transistor (CNFET) technology. CNFET7 is based on an open-source CNFET SPICE model called VS-CNFET, and various model parameters such as the channel width and carbon nanotube diameter are carefully tuned to mimic the predictive 7-nm CNFET technology presented in a published paper. Some nondisclosure parameters, such as the cell size and pin layout, are derived from those of the NanGate 15-nm open-source cell library in the same way as for an open-source framework for CNFET circuit design. CNFET7 includes two types of delay model (i.e., the composite current source and nonlinear delay model), each having 56 cells, such as INV_X1 and BUF_X1. CNFET7 supports both logic synthesis and timing-driven place and route in the Cadence design flow. Our experimental results for several synthesized circuits show that CNFET7 has reductions of up to 96%, 62% and 82% in dynamic and static power consumption and critical-path delay, respectively, when compared with ASAP7. Chenlin Shi, Shinobu Miwa, Tongxin Yang, Ryota Shioya, Hayato Yamaki, Hiroki Honda |
ASP-DAC | 2 |
| 2022 | Analyzing Performance and Power-Efficiency Variations among NVIDIA GPUsabstractUnderstanding the variations in performance and power-efficiency of compute nodes is important for enhancing these factors in modern supercomputing systems. Previous studies have focused on variations in CPUs and DRAMs, but there has been little attention on GPUs. This is despite many current supercomputing systems employing GPUs (which consume a significant fraction of the power of such systems) as power-efficient accelerators for HPC applications. This paper describes the first thorough evaluation of performance and power-efficiency variations in GPUs. Specifically, we execute 48 CUDA kernels on 856 devices selected from three generations of NVIDIA GPUs (P100, V100, and A100), and analyze the impact of differences in both the CUDA kernels and GPU generation on performance and power-efficiency. Our analysis shows that there are non-negligible variations in both performance and power-efficiency, and that these variations are strongly affected by both the kernels that are running and the GPU generation. Kohei Yoshida, Rio Sageyama, Shinobu Miwa, Hayato Yamaki, Hiroki Honda |
ICPP | 3 |
| 2021 | PredCom: A Predictive Approach to Collecting Approximated Communication TracesabstractCommunication traces collected from MPI applications are an important source of information for performance optimization as they can help analysts determine communication patterns and identify inefficiencies. However, their collection, especially at scale, is time consuming, since it usually requires running the complete target application on a large number of nodes. In this work, we present PredCom, a tool-chain to generate a predictive communication proxy based on information gathered from a few small scale runs, which allows us to extract approximate communication traces with an accuracy high enough for most analysis goals. For this, we combine LLVM passes on the original source code (to capture static program structure) with parameter prediction (to capture dynamic and scaling behavior). This approach drastically reduces the time needed for collecting the communication traces, even for traces on large numbers of MPI processes. We demonstrate that PredCom generates communication traces of various applications up to 1612x faster with an accuracy loss of 0.11 on average compared to the original large-scale traces, and we show that the generated traces can be used to optimize process placement. Shinobu Miwa, Ignacio Laguna, Martin Schulz 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Evaluating architecture-level optimization in packet processing cachesabstractThe next-generation internet routers need ultra high speed such as 1 Tbps to process increased internet traffic. Efficient table lookup is key to realize high-speed packet processing and Packet Processing Cache (PPC) was proposed for this purpose. However, PPC has not been well optimized in the view of architecture and its ability to improve table lookup has therefore been underestimated so far. In this paper, we revisit PPC to reveal its real potential in table lookup. For this, we apply three architecture-level optimization techniques (hierarchization, pipelining and port addition), which are widely used to improve throughput of caches in microprocessors, and their combinations to PPC. Through the analysis of hundreds of billions of candidates of PPC configurations, we clarify the impact of these optimization techniques on the design space of internet routers with PPC. Our experimental results show that our best configuration can achieve 1045.7 Gbps packet processing with the power of 4273.3 mW and the area of 50.20 mm2, which is 3.35x improvement in Gbps per watt when compared to conventional PPC. Kyosuke Tanaka, Hayato Yamaki, Shinobu Miwa, Hiroki Honda |
Comput. Networks | 3 |
| 2020 | Footprint-Based DIMM HotplugabstractPower-efficiency has become one of the most critical concerns for HPC as we continue to scale computational capabilities. A significant fraction of system power is spent on large main memories, mainly caused by the substantial amount of DIMM standby power needed. However, while necessary for some workloads, for many workloads large memory configurations are too rich, i.e., these workloads only make use of a fraction of the available memory, causing unnecessary power usage. This observation opens new opportunities for power reduction by powering DIMMs on and off depending on the current workload. In this article, we propose footprint-based DIMM hotplug that enables a compute node to adjust the number of DIMMs that are powered on depending on the memory footprint of a running job. Our technique relies on two main subcomponents-memory footprint monitoring and DIMM management-which we both implement as part of an optimized page management system with small control overhead. Using Linux's memory hotplug capabilities, we implement our approach on a real system, and our results show that our proposed technique can save 50.6-52.1 percent of the DIMM standby energy and the CPU+DRAM energy of up to 1.50 Wh for various small-memory-footprint applications without loss of performance. Shinobu Miwa, Masaya Ishihara, Hayato Yamaki, Hiroki Honda, Martin Schulz 0001 |
IEEE Trans. Computers | 1 |
| 2019 | Evaluating the Impact of Energy Efficient Networks on HPC WorkloadsabstractInterconnection networks grow larger as supercomputers include more nodes and require higher bandwidth for performance. This scaling significantly increases the fraction of power consumed by the network, by increasing the number of network components (links and switches). Typically, network links consume power continuously once they are turned on. However, recent proposals for energy efficient interconnects have introduced low-power operation modes for periods when network links are idle. Low-power operation can increase messaging time when switching a link from low-power to active operation. We extend the TraceR-CODES network simulator for power modeling to evaluate the impact of energy efficient networking on power and performance. Our evaluation presents the first study on both single-job and multi-job execution to realistically simulate power consumption and performance under congestion for a large-scale HPC network. Results on several workloads consisting of HPC proxy applications show that single-job and multi-job execution favor different modes of low power operation to have significant power savings at the cost of minimal performance degradation. Giorgis Georgakoudis, Takatsugu Ono, Koji Inoue, Shinobu Miwa, Abhinav Bhatele |
HiPC | 5 |
| 2018 | Data prediction for response flows in packet processing cacheabstractWe propose a technique to reduce compulsory misses of packet processing cache (PPC), which largely affects both throughput and energy of core routers. Rather than prefetching data, our technique called response prediction cache (RPC) speculatively stores predicted data into PPC without additional access to the low-throughput and power-consuming memory (i.e., TCAM). RPC predicts the data related to a response flow at the arrival of the corresponding request flow, based on the request-response model of internet communications. RPC can improve the cache miss rate, throughput, and energy-efficiency of PPC systems by 15.3%, 17.9%, and 17.8%, respectively. Hayato Yamaki, Hiroaki Nishi, Shinobu Miwa, Hiroki Honda |
DAC | 3 |
| 2015 | Immediate sleep: Reducing energy impact of peripheral circuits in STT-MRAM cachesabstractImplementing last level caches (LLCs) with STT-MRAM is a promising approach for designing energy efficient microprocessors due to high density and low leakage power of its memory cells. However, peripheral circuits of an STT-MRAM cache still suffer from leakage power because large and leaky transistors are required to drive large write current to STT-MRAM element. To overcome this problem, we propose a new power management scheme called Immediate Sleep (IS). IS immediately turns off a subarray of an STT-MRAM cache if the next access is predicted to be not critical in performance. Thus, IS can effectively reduce leakage energy with little impact on performance. Our experimental results show that our technique can save the leakage energy of an STT-MRAM LLC by 32% compared to an STT-MRAM LLC with the conventional scheme at the same performance. Eishi Arima, Hiroki Noguchi, Takashi Nakada, Shinobu Miwa, Susumu Takeda, Shinobu Fujita, Hiroshi Nakamura |
ICCD | 4 |
| 2015 | Runtime multi-optimizations for energy efficient on-chip interconnections1abstractOn-chip interconnection (or NoC) is a major performance and power contributor to modern and future multicore processors. So far, many optimization techniques have been developed to improve its bandwidth, latency and power consumption. But it is not clear how energy efficiency is affected since an optimization technique normally comes with overheads. This paper thus attempts to address when and how such optimization techniques should be applied and tuned to help achieve better energy efficiency. We firstly model the performance and energy impacts of representative NoC optimization techniques. These models help us more easily understand the consequences when applying these optimization techniques and their combinations under different circumstances. Moreover, based on such modeling, we propose and implement an adaptive control over these NoC optimization techniques to improve both performance and energy efficiency of the network. Our results show that, this proposal can achieve an average improvement of 26% and 57% on network performance and energy delay product, respectively. Yuan He 0002, Masaaki Kondo, Takashi Nakada, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
ICCD | 5 |
| 2015 | Profile-based power shifting in interconnection networks with on/off linksabstractOverprovisioning hardware devices and coordinating their power budgets are proposed to improve the application performance of future power-constrained HPC systems. This coordination process is called power shifting. Meanwhile, recent studies have revealed that on/off links can save network power in HPC systems. Future HPC systems will thus adopt on/off links in addition to power shifting. This paper explores power shifting in interconnection networks with on/off links. Given that on/off links keep network power low at application runtime, we can transfer appreciable quantities of power budgets on networks to other devices before an application runs. We thus propose a profile-based power shifting technique that allows HPC users to transfer the power budget remaining on networks to other devices at the time of job dispatch. Experimental results show that the proposed technique appreciably improves application performance under various power constraints. Shinobu Miwa, Hiroshi Nakamura |
SC | 1 |
| 2014 | Normally-off computing project: Challenges and opportunitiesabstractNormally-Off is a way of computing which aggressively powers off components of computer systems when they need not to operate. Simple power gating cannot fully take the chances of power reduction because volatile memories lose data when power is turned off. Recently, new non-volatile memories (NVMs) have appeared. High attention has been paid to normally-off computing using these NVMs. In this paper, its expectation and challenges are addressed with a brief introduction of our project started in 2011. Hiroshi Nakamura, Takashi Nakada, Shinobu Miwa |
ASP-DAC | 3 |
| 2013 | McRouter: Multicast within a router for high performance network-on-chipsabstractThe inevitable advent of the multi-core era has driven an increasing demand for low latency on-chip inter-connection networks (or NoCs). Being a critical part of the memory hierarchy for modern chip multi-processors (CMPs), these networks face stringent design constraints to provide fast communication with tight power budget. Modern NoC's first-order concern is clearly its latency, while we also find that internal bandwidth of its routers is relatively plentiful; thus, we present a low latency router design utilizing a technique we call “multicast within a router” or McRouter, which allows productive utilization of remaining bandwidth inside a NoC router. McRouter allows a single cycle transfer of flits which shortens the communication latency when there is enough remaining bandwidth within the router. The key idea is to transmit a header flit to all possible output ports (multicast) so that it is always transmitted to the correct output port without relying on route computation. In addition, we find it is affordable with marginal power overhead while still being a stand-alone design by maintaining portability and modularity (unlike look-ahead routing based designs). Our evaluation with application traffic shows that McRouter helps achieving system speed-ups of 1.28, 1.17 and 1.05 over the conventional router (CR), the VSA router (VSAR) and the prediction router (PR), respectively. Yuan He 0002, Hiroshi Sasaki 0001, Shinobu Miwa, Hiroshi Nakamura |
PACT | 3 |
| 2013 | D-MRAM cache: enhancing energy efficiency with 3T-1MTJ DRAM/MRAM hybrid memoryabstractThis paper describes a proposal of non-volatile cache architecture utilizing novel DRAM / MRAM cell-level hybrid structured memory (D-MRAM) that enables effective power reduction for high performance mobile SoCs without area overhead. Here, the key point to reduce active power is intermittent refresh process for the DRAM-mode. D-MRAM has advantage to reduce static power consumptions compared to the conventional SRAM, because there are no static leakage paths in the D-MRAM cell and it is not needed to supply voltage to its cells when used as the MRAM-mode. Besides, with advanced perpendicular magnetic tunnel junctions (p-MTJ), which decreases the write energy and latency without shortening its retention time, D-MRAM is capable of power reduction by replacing the traditional SRAM caches. Considering the 65-nm CMOS technology, the access latencies of 1MB memory macro are 2.2 ns / 1.5 ns for read / write in DRAM mode, and 2.2 ns / 4.5 ns in MRAM mode, while those of SRAM are 1.17 ns. The SPEC CPU2006 benchmarks have revealed that the energy per instruction (EPI) of the total cache memory can be dramatically reduced by 71 % on average, and the instruction per cycle (IPC) performance of the D-MRAM cache architecture degraded only by approximately 4 % on average in spite of its latency overhead. Hiroki Noguchi, Kumiko Nomura, Keiko Abe, Shinobu Fujita, Eishi Arima, Kyundong Kim, Takashi Nakada, Shinobu Miwa, Hiroshi Nakamura |
DATE | 8 |
| 2013 | Power capping of CPU-GPU heterogeneous systems through coordinating DVFS and task mappingabstractFuture computer systems are built under much stringent power budget due to the limitation of power delivery and cooling systems. To this end, sophisticated power management techniques are required. Power capping is a technique to limit the power consumption of a system to the predetermined level, and has been extensively studied in homogeneous systems. However, few studies about the power capping of CPU-GPU heterogeneous systems have been done yet. In this paper, we propose an efficient power capping technique through coordinating DVFS and task mapping in a single computing node equipped with GPUs. In CPU-GPU heterogeneous systems, settings of the device frequencies have to be considered with task mapping between the CPUs and the GPUs because the frequency scaling can incurs load imbalance between them. To guide the settings of DVFS and task mapping for avoiding power violation and the load imbalance, we develop new empirical models of the performance and the maximum power consumption of a CPU-GPU heterogeneous system. The models enable us to set near-optimal settings of the device frequencies and the task mapping in advance of the application execution. We evaluate the proposed technique with five data-parallel applications on a machine equipped with a single CPU and a single GPU. The experimental result shows that the performance achieved by the proposed power capping technique is comparable to the ideal one. Toshiya Komoda, Shingo Hayashi, Takashi Nakada, Shinobu Miwa, Hiroshi Nakamura |
ICCD | 4 |
| 2013 | Integrating Multi-GPU Execution in an OpenACC CompilerabstractGPUs have become promising computing devices in current and future computer systems due to its high performance, high energy efficiency, and low price. However, lack of high level GPU programming models hinders the wide spread of GPU applications. To resolve this issue, OpenACC is developed as the first industry standard of a directive-based GPU programming model and several implementations are now available. Although early evaluations of the OpenACC systems showed significant performance improvement with modest programming efforts, they also revealed the limitations of the systems. One of the biggest limitations is that the current OpenACC compilers do not automate the utilization of multiple GPUs. In this paper, we present an OpenACC compiler with the capability to execute single GPU OpenACC programs on multiple GPUs. By orchestrating the compiler and the runtime system, the proposed system can efficiently manage the necessary data movements among multiple GPUs memories. To enable advanced communication optimizations in the proposed system, we propose a small set of directives as extensions of OpenACC API. The directives allow programmers to express the patterns of memory accesses in the parallel loops to be offloaded. Inserting a few directives into an OpenACC program can reduce a large amount of unnecessary data movements and thus helps the proposed system drawing great performance from multi-GPU systems. We implemented and evaluated the prototype system on top of CUDA with three data parallel applications. The proposed system achieves up to 6.75x of the performance compared to OpenMP in the 1CPU with 2GPU machine, and up to 2.95x of the performance compared to OpenMP in the 2CPU with 3GPU machine. In addition, in two of the three applications, the multi-GPU OpenACC compiler outperforms the single GPU system where hand-written CUDA programs run. Toshiya Komoda, Shinobu Miwa, Hiroshi Nakamura, Naoya Maruyama |
ICPP | 2 |
| 2013 | Performance modeling for designing NoC-based multiprocessorsabstractNetwork-on-Chip (NoC) based multiprocessors have become popular as a scalable alternative to classical bus architectures. The performance evaluation of NoC-based multiprocessors is largely based on simulation. However, precise simulation is extremely slow. Additionally, there are many design parameters that affect the total performance. Therefore, it is practically impossible to use the precise simulation for the design space exploration purposes. To alleviate this problem, prototyping NoC systems and estimating their performances are critically important. In this paper, we present a generalized novel performance model that combined with the simulations for designing NoC-based multiprocessors. We revealed that the performance impact of cache and network latencies are dominant. Moreover, network congestion rarely happens under near appropriate configuration. Thus, the performance model is mainly constructed using the hardware parameters and the statistics that obtained from a simple cache simulation that is separated from the network behavior. The proposed performance model is used not only to obtain fast and accurate performance, but also to guide the NoC-based multiprocessor design space exploration. The accuracy of our approach and its practical use are illustrated through simulation. The results showed that proposed model can estimate performance with only 3.4% error on average and 21% at worst. We also confirmed that our evaluation framework can estimate 360 times faster than the brute force full system simulation. Takashi Nakada, Shinobu Miwa, Keisuke Y. Yano, Hiroshi Nakamura |
RSP | 2 |
| 2012 | A novel power-gating scheme utilizing data retentiveness on cachesabstractCaches are one of the most leakage consuming components in modern processor because of massive amount of transistors. To reduce leakage power of caches, several techniques using power-gating(PG) were proposed. Despite of its high leakage saving, a side effect of PG for caches is the loss of data during a sleep. If useful data is lost in sleep mode, it should be fetched again from a lower level memory. This consumes a considerable amount of energy, which very unfortunately mitigates the leakage saving. This paper proposes a new PG scheme considering data retentiveness of SRAM. After entering the sleep mode, data of an SRAM cell is not lost immediately and is usable by checking the validity of the data. Therefore, we utilize data retentiveness of SRAM to avoid energy overhead for data recovery, which results in further chance of leakage saving. To check availability, we introduce a simple hardware whose overhead is ignorable. We also examined leakage saving potential of our approach. For both L1 data and instruction caches, our scheme results in more than 2 times of smaller leakage energy compared to conventional PG scheme. Kyundong Kim, Seidai Takeda, Shinobu Miwa, Hiroshi Nakamura |
ACM Great Lakes Symposium on VLSI | 3 |
| 2012 | Stepwise sleep depth control for run-time leakage power savingabstractRecently, run-time sleep control scheme using multiple sleep modes have been studied. In those studies, each sleep mode has its own sleep depth. Deeper sleep mode provides higher leakage saving but incurs larger overhead energy.Use of multiple modes is helpful for further leakage saving if an appropriate mode is selected, but the best mode depends on the idle period whose length cannot be told in advance. Although the implementations how to realize different sleep depths have been well studied, few attention has been paid to the method of how to select the best sleep depth dynamically during execution. This paper proposes a simple but novel sleep control scheme, called stepwise sleep depth control, which aims to select the best depth among provided multiple sleep depths.Our scheme automatically applies deeper depth in a step-by-step manner after an idle state starts. It successfully reduces leakage energy while only a small modification is required for circuit implementation. This paper also proposes a methodology for optimizing control parameters of our sleep control scheme according to program behavior and temperature. Experimental result shows that stepwise sleep depth control applied to body biasing circuit improves net leakage saving of up to 43% for FPAlu at 1.0GHz, 75°C compared to conventional reverse body biasing. Seidai Takeda, Shinobu Miwa, Kimiyoshi Usami, Hiroshi Nakamura |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | A Fine-Grained Runtime Power/Performance Optimization Method for Processors with Adaptive Pipeline Depth
Jun Yao 0001, Shinobu Miwa, Hajime Shimada, Shinji Tomita |
J. Comput. Sci. Technol. | 2 |