VLDB 2026 Research / reviewers in the wild / expert
Ryuichi Sakamoto
dblp:119/1671
· DBLP profile ↗
15ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0002-2999-100XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VAHRM: Variation-Aware Resource Management in Heterogeneous Supercomputing SystemsabstractIn this paper, we propose a novel resource management technique for heterogeneous supercomputing systems affected by manufacturing variability. Our proposed technique called VAHRM (Variation-Aware Heterogeneous Resource Management) takes a holistic approach to job scheduling on highly heterogeneous computing resources. VAHRM preferentially allocates energy-efficient computing resources to an energy-consuming job in a job queue, considering the impact on both the job turnaround time and the power consumption of individual resources. Furthermore, we have developed a novel approach to modeling the power consumption of computing resources that have manufacturing variability. Our approach called TSMVA (Two-Stage Modeling with Variation Awareness) enables us to generate the first variation-aware GPU power models, which can correctly estimate the power consumption of each GPU for a given job. Our experimental results show that, compared to conventional first-come-first-serve (FCFS) and state-of-the-art variation-aware scheduling algorithms, VAHRM can achieve respective improvements in system energy efficiency of up to 5.8% and 5.4% (4.5% and 4.2% on average) while reducing the average turnaround time of 21.2% and 11.9%, respectively, for various workloads obtained from a production system. Kohei Yoshida, Ryuichi Sakamoto, Kento Sato, Abhinav Bhatele, Hayato Yamaki, Hiroki Honda, Shinobu Miwa |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Extendable MQTT Broker for Feedback-based Resource Management in Large-scale Computing EnvironmentsabstractHigh-performance computing (HPC) systems demand continuous monitoring to ensure efficient resource allocation and application performance. Recent studies indicate that real-time resource utilization monitoring can significantly improve the performance of dynamic scheduling algorithms. However, latency induced by protocol stack heavily impacts the effectiveness of dynamic scheduling. In this paper, we propose a novel monitoring system that implements the protocol stack on a Field-Programmable Gate Array (FPGA) and adopts a publish/subscribe (pub/sub) communication protocol. Specifically, by introducing an FPGA-based protocol stack, we substantially reduce the latency of protocol stack processing and enable the implementation of custom plugins at the L7 layer. Our experiments demonstrate that the proposed system effectively reduces protocol stack latency and, with the extensibility provided by user-defined plugins, offers great potential for a wide range of HPC monitoring and feedback applications. Ryo Ouchi, Ryuichi Sakamoto |
APNet | 2 |
| 2022 | GraphDEAR: An Accelerator Architecture for Exploiting Cache Locality in Graph Analytics ApplicationsabstractData structure is the key in Edge Computing where various types of data are continuously generated by ubiquitous devices. Within all common data structures, graphs are used to express relationships and dependencies among human identities, objects, and locations; and they are expected to become one of the most important data infrastructure in the near future. Furthermore, as graph processing often requires random accesses to vast memory spaces, conventional memory hierarchies with caches cannot perform efficiently. To alleviate such memory access bottlenecks in graph processing, we present a solution through vertex accesses scheduling and edge array re-ordering, in parallel with the execution of graph processing application to improve both temporal and spatial locality of memory accesses, especially for edge-centric graphs which are popular means in handling dynamic graphs. Our proposed architecture is evaluated and tested through both trace-based cache simulations and cycle-accurate FPGA-based prototyping. Evaluation results show that our proposal has a potential of significantly reducing the quantity of Miss-Per-Kilo-Instructions (MPKI) for Last Level Cache (LLC) by 56.27% on average. Masaaki Kondo, Yuan He 0002, Ryuichi Sakamoto, Hiroshi Nakamura |
PDP | 4 |
| 2021 | Lexicographic and reverse lexicographic quadratic Gröbner bases of cut ideals
Ryuichi Sakamoto |
J. Symb. Comput. | 1 |
| 2020 | The Effectiveness of Low-Precision Floating Arithmetic on Numerical Codes: A Case Study on Power ConsumptionabstractThe low-precision floating point arithmetic that performs computation by reducing numerical accuracy with narrow bit-width is attracting since it can improve the performance of the numerical programs. Small memory footprint, faster computing speed, and energy saving are expected by performing calculation with low precision data. However, there have not been many studies on how low-precision arithmetics affects power and energy consumption of numerical codes. In this study, we investigate the power efficiency improvement by aggressively using low-precision arithmetics for HPC applications. In our evaluations, we analyze power characteristics of the Poisson's equation and the ground motion simulation programs with double precision and single precision floating point arithmetics. We confirm that energy efficiency improves by using low-precision arithmetics but it is heavily influenced by parameters such as data division and the number of OpenMP threads. Ryuichi Sakamoto, Masaaki Kondo, Kohei Fujita, Tsuyoshi Ichimura, Kengo Nakajima |
HPC Asia | 1 |
| 2018 | Analyzing Resource Trade-offs in Hardware Overprovisioned SupercomputersabstractHardware overprovisioned systems have recently been proposed as a viable alternative for a power-efficient design of next-generation supercomputers. A key challenge for such systems is to determine the degree of overprovisioning, which refers to the number of extra nodes that need to be installed under a given power constraint. In this paper, we first show that the degree of overprovisioning depends on dynamic parameters, such as the job mix as well as the global power constraint, and that static decisions can result in limited system throughput. We then study an exhaustive combination of adaptive resource management strategies that span three job scheduling algorithms, four power capping techniques, and three node boot-up mechanisms to understand the trade-off space involved. We then draw conclusions about how these strategies can adaptively control the degree of overprovisioning and analyze their impact on job throughput and power utilization. Ryuichi Sakamoto, Tapasya Patki, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 1 |
| 2018 | OpenCL Runtime for OS-Driven Task Pipelining on Heterogeneous AcceleratorsabstractTask pipelining on accelerators is suitable for streaming applications, while its performance can decrease due to frequent user/OS interactions caused by hardware control via device drivers. Our previous research has proposed PPM, an OS support that efficiently manages multiple accelerators by eliminating the user/OS interactions during the pipelined execution. To allow users to develop and execute pipeline applications using PPM, this paper introduces a customized OpenCL runtime library. When executing an OpenCL application, the runtime library dynamically analyzes OpenCL API calls and creates a data flow graph required for the PPM execution. With the runtime library, users easily execute applications written in common OpenCL pipeline model on PPM. Atsushi Koshiba, Ryuichi Sakamoto, Mitaro Namiki |
RTCSA | 2 |
| 2017 | Production Hardware Overprovisioning: Real-World Performance Optimization Using an Extensible Power-Aware Resource Management FrameworkabstractLimited power budgets will be one of the biggest challenges for deploying future exascale supercomputers. One of the promising ways to deal with this challenge is hardware over provisioning, that is, installing more hardware resources than can be fully powered under a given power limit coupled with software mechanisms to steer the limited power to where it is needed most. Prior research has demonstrated the viability of this approach, but could only rely on small-scale simulations of the software stack. While such research is useful to understand the boundaries of performance benefits that can be achieved, it does not cover any deployment or operational concerns of using overprovisioning on production systems. This paper is the first to present an extensible power-aware resource management framework for production-sized overprovisioned systems based on the widely established SLURM resource manager. Our framework provides flexible plugin interfaces and APIs for power management that can be easily extended to implement site-specific strategies and for comparison of different power management techniques. We demonstrate our framework on a 965-node HA8000 production system at Kyushu University. Our results indicate that it is indeed possible to safely overprovision hardware in production. We also find that the power consumption of idle nodes, which depends on the degree of overprovisioning, can become a bottleneck. Using real-world data, we then draw conclusions about the impact of the total number of nodes provided in an overprovisioned environment. Ryuichi Sakamoto, Masaaki Kondo, Koji Inoue, Masatsugu Ueda, Tapasya Patki, Daniel A. Ellsworth, Barry Rountree, Martin Schulz 0001 |
IPDPS | 1 |
| 2014 | Design and control methodology for fine grain power gating based on energy characterization and code profiling of microprocessorsabstractThis paper presents a design and control scheme of a microprocessor whose internal function units are power gated at instruction-by-instruction basis. Enabling/disabling the power gating is adaptively controlled under the support of on-chip leakage monitors and the operating system to minimize energy overhead due to sleep-in and wakeup. Measured results of the fabricated chip in the 65nm CMOS technology demonstrated that our approach reduces energy to 21-35% in the range of 25-85°C as compared to the non power-gated case. Energy dissipation was reduced by up to 15% as compared to the conventional fine-grain power gating technique in the same temperature range. Kimiyoshi Usami, Masaru Kudo, Kensaku Matsunaga, Tsubasa Kosaka, Yoshihiro Tsurui, Hideharu Amano, Hiroaki Kobayashi, Ryuichi Sakamoto, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ASP-DAC | 9 |
| 2014 | Design and evaluation of fine-grained power-gating for embedded microprocessorsabstractPower-performance efficiency is still remaining a primary concern for microprocessor designers. One of the sources of power inefficiency for recent LSI chips is increasing leakage power consumption. Power-gating is a well known technique to reduce leakage power consumption by switching off the power supply to idle logic blocks. Recently, fine-grained power-gating is emerged as a technique to minimize leakage current during the active processor cycles by switching on and off a logic blocks in much finer temporal/spatial granularity. Though fine-grained power-gating is useful, a comprehensive evaluation and analysis has not been conducted on a real LSI chips. In this paper, we evaluate fine-grained run-time power-gating for microprocessors' functional units using a real embedded microprocessor. We also introduce an architecture and compiler co-operative power-gating scheme which mitigates negative power reduction caused by the energy overhead associated with finegrained power-gating. The experimental results with a fabricated core shows that a hardware-based scheme saves power consumption of functional units by 44% and hardware compiler co-operative scheme further improves power efficiency by 5.9% when core temperature is 25 ˚C. Masaaki Kondo, Hiroaki Kobayashi, Ryuichi Sakamoto, Motoki Wada, Jun Tsukamoto, Mitaro Namiki, Hideharu Amano, Kensaku Matsunaga, Masaru Kudo, Kimiyoshi Usami, Toshiya Komoda, Hiroshi Nakamura |
DATE | 3 |
| 2013 | Demonstration of a heterogeneous multi-core processor with 3-D inductive coupling linksabstractCube-1 is a heterogeneous multi-core processor which can achieve the required performance with the least energy consumption as possible. It can control the performance and energy with two levels: (1) the number of accelerators can be easily changed by increasing or decreasing the number of stacked chips after fabrication, as they are connected with inductive coupling links. (2) The supply voltage for PE array of the accelerator can be controlled by the host CPU so that the required performance can be obtained with a minimum supply voltage. Yusuke Koizumi, Noriyuki Miura, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 7 |
| 2013 | A scalable 3D heterogeneous multi-core processor with inductive-coupling thruchip interface
Noriyuki Miura, Yusuke Koizumi, Eiichi Sasaki, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
Hot Chips Symposium | 8 |
| 2012 | CMA-Cube: A scalable reconfigurable accelerator with 3-D wireless inductive coupling interconnectabstractCMA-Cube is the second prototype of building block scalable reconfigurable accelerator using inductive coupling interconnect. It uses the wireless inductive coupling interconnect as a packet switching network which connects accelerators. As an accelerator core, CMA (Cool Mega Array), which consists of a large coarse-grained PE array with combinatorial circuits and tiny micro-controller, is applied. Evaluation results of Cube-1 Quad Core which consists of a host embedded CPU and three CMA-Cubes achieved 3.15 times performance acceleration as that without accelerators when JPEG decoder is executed. Yusuke Koizumi, Eiichi Sasaki, Hideharu Amano, Hiroki Matsutani, Yasuhiro Take, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPL | 7 |
| 2012 | Dynamic power control with a heterogeneous multi-core system using a 3-D wireless inductive coupling interconnectabstractCube-2 is a prototype of building block scalable reconfigurable accelerator using an inductive coupling interconnect. It is consisting of a ultra low leakage embedded processor Geyser and coarse-grained reconfigurable accelerators CMA (Cool Mega Array). A Geyser chip and multiple CMA chips are stacked, and a powerful network is formed by using the inductive coupling interconnect. The performance can be enhanced by increasing the number of CMA chips. JPEG decoder is implemented with a cooperation of Geyser and CMAs, and low power execution by controlling the power supply voltage of CMAs is demonstrated. Yusuke Koizumi, Hideharu Amano, Hiroki Matsutani, Noriyuki Miura, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura |
FPT | 6 |
| 2012 | An OpenCL Runtime Library for Embedded Multi-Core AcceleratorabstractIn recent years, improvements of energy efficiency and computational performance have become a major issue, because smartphones and tablets become popular. To implement high performance, multi-core accelerator consists of general purpose processors and accelerators are often used. But to use these multi-core accelerator efficiently, programmers have to consider synchronization and data transfer between accelerators, memory and I/O. Therefore frameworks such as OpenCL have been proposed for effective use of parallel computing resources of multi-core processors. OpenCL frameworks for GPGPU, Intel multi-core and many other kind of multi-core processors have been developed on Linux. However, in order to use those multi-core accelerators effectively, it is necessary to control accelerators with software environment aiming effective use of accelerators. Generally user mode application program calls OS to synchronize of accelerator's data transfer and termination of the its execution with overhead. In this paper, accelerators are controlled from collaborated OS and OpenCL library, instead of user mode application program. Our feature of these methods is that OS cooperates with library to reduce overhead of control accelerators. The proposed framework in this paper integrates the execution and data transfer control in an OpenCL library and an embedded OS. The OpenCL library creates task automaticity generated from the description of the user program. The OpenCL library gives scheduling information such as execution order about tasks to the scheduler of the embedded OS. The embedded OS's scheduler handles events of execution, termination and synchronization of data transfer via interrupts from accelerators. The scheduler determines to execute next task from the scheduling information passed by the OpenCL library. By the schemes to control accelerators are managed in OS, accelerators can be worked efficiently. The poster presentation shows the implemented OpenCL and OS framework for "Cube" processor embedded multi-core accelerator. Ryuichi Sakamoto, Mikiko Sato, Yusuke Koizumi, Hideharu Amano, Mitaro Namiki |
RTCSA | 1 |