Satoshi Kawakami

dblp:132/6521 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-5044-744XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 PrometheusFree: Concurrent Detection of Laser Fault Injection Attacks in Optical Neural Networks
Kota Nishida, Yoshihiro Midoh, Noriyuki Miura, Satoshi Kawakami, Alex Orailoglu, Jun Shiomi
ASP-DAC4
2026 SFQ-Based CJoin Gate Implementation for Ultra-Low-Power Brownian Logic Circuits
abstract
The increasing energy consumption of data centers has highlighted the urgent need for ultra-low-power computing solutions. Single Flux Quantum (SFQ) circuits, which utilize superconducting Josephson junctions, offer promising advantages in terms of speed and energy efficiency; however, their power consumption has been limited by the need for noise suppression mechanisms. To address this, we propose the practical SFQ Brownian Logic Circuits (SBLCs), which exploit thermal noise for stochastic signal propagation, significantly reducing static power consumption. Especially, we address the three challenges: susceptibility to manufacturing variations, the absence of an actual CJoin gate implementation, which is essential for processing, and a lack of demonstrated practical advantages. This paper evaluates the robustness of SBLCs to manufacturing variations, proposes the first SFQ-based CJoin gate implementation, and demonstrates a Ripple Carry Adder (RCA) with a 3167x improvement in energy efficiency compared to traditional SFQ and CMOS circuits, confirming the superiority of SBLCs for future computing systems.
Soshi Takagi, Masamitsu Tanaka, Koji Inoue, Satoshi Kawakami
DATE4
2024 Late Breaking Results: Single Flux Quantum Based Brownian Circuits for Ultra-Law-Power Computing
abstract
This paper proposes a random walk circuit imple-mentation with single flux quantum devices, essential for Brownian circuits, to reduce processing energy consumption dramatically. SPICE-based simulation demonstrating its functional operation and random walks can be achieved via the Shapiro- Wilk test. Furthermore, we developed a Monte Carlo simulator for Brownian circuits, enabling functionality verification and computation step distribution analysis. Latency/energy evaluation using a half-adder as a case study revealed that proposed circuits could reduce energy consumption by 1/1260 and offer an opportunity for low-power computing systems.
Satoshi Kawakami, Yusuke Ohtusbo, Koji Inoue, Masamitsu Tanaka
DATE1
2023 Evaluating floating-point multipliers with opto-electrical hybrid circuits
abstract
Nanophotonic technology has ultra-low latency and low energy consumption characteristics, and several previous research have shown the potential of photonic-based computation. However, existing optical circuits are based on low-bit-width integer operations with analog manners, so there is less research on optical floating-point operations. Providing the benefits of analog optical circuits (i.e., low latency and low energy) to floating-point multipliers (FMs) requires careful consideration. In particular, introducing optical circuits unnecessarily increases the overhead of optical-to-electrical (O/E) and analog-to-digital (A/D) conversion, thus reducing system performance efficiency. Appropriate processing balancing between electrical and optical circuits is an important issue in OE hybrid circuit design. This study evaluates FM latency and energy consumption with different combinations of optical and electrical circuits. We design two versions of an Opto-Electrical FM (OEFM): M-OEFM, an optical implementation of an integer multiplier in an FM, and MA-OEFM, an optical implementation of an integer multiplier and adder functions. These designs are based on a policy of priority optical circuit implementation of the dominant components of conventional electrical circuits. Experimental results indicate that the M-OEFM achieved a 56 % reduction in latency and a 41 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM achieved an 88 % reduction in latency and a 19 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM is superior to M-OEFM in terms of energy-delay product (EDP). The results also indicate that appropriate co-design of analog optical and digital electrical circuits enables highly efficient FMs.
Takumi Inaba, Takatsugu Ono, Koji Inoue, Satoshi Kawakami
CF4
2022 Design of Variable Bit-Width Arithmetic Unit Using Single Flux Quantum Device
abstract
This paper presents the design of an ultra-high-speed, low-power arithmetic unit that supports variable bit-width operations with single flux quantum (SFQ) technology. Because of the high-speed nature of superconductor devices, we can achieve extremely high power-performance efficiency that cannot be achieved by state-of-the-art CMOS devices. To implement the complex function to support the variable bit-width feature, we introduce a novel circuit architecture to maintain the high-speed operation over 50GHz. Our prototype chip design successfully demonstrated 53.5GHz 1.59mW operations.
Iori Ishikawa, Ikki Nagaoka, Ryota Kashima, Koki Ishida, Kosuke Fukumitsu, Keitarou Oka, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Akira Fujimaki, Koji Inoue
ISCAS8
2020 Enhancing a manycore-oriented compressed cache for GPGPU
abstract
GPUs can achieve high performance by exploiting massive-thread parallelism. However, some factors limit performance on GPUs, one of which is the negative effects of L1 cache misses. In some applications, GPUs are likely to suffer from L1 cache conflicts because a large number of cores share a small L1 cache capacity. A cache architecture that is based on data compression is a strong candidate for solving this problem as it can reduce the number of cache misses. Unlike previous studies, our data compression scheme attempts to exploit the value locality existing within not only intra cache lines but also inter cache lines. We enhance the structure of a last-level compression cache proposed for general purpose manycore processors to optimize against shared L1 caches on GPUs. The experimental results reveal that our proposal outperforms the other compression cache for GPUs by 11 points on average.
Keitarou Oka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Koji Inoue
HPC Asia2
2020 SuperNPU: An Extremely Fast Neural Processing Unit Using Superconducting Logic Devices
abstract
Superconductor single-flux-quantum (SFQ) logic family has been recognized as a highly promising solution for the post-Moore's era, thanks to its ultra-fast and low-power switching characteristics. Therefore, researchers have made a tremendous amount of effort in various aspects to promote the technology and automate its circuit design process (e.g., low-cost fabrication, design tool development). However, there has been no progress in designing a convincing SFQ-based architectural unit due to the architects' lack of understanding of the technology's potentials and limitations at the architecture level. In this paper, we present how to architect an SFQ-based architectural unit by providing design principles with an extreme-performance neural processing unit (NPU). To achieve the goal, we first implement an architecture-level simulator to model an SFQ-based NPU accurately. We validate this model using our die-level prototypes, design tools, and logic cell library. This simulator accurately measures the NPU's performance, power consumption, area, and cooling overheads. Next, driven by the modeling, we identify key architectural challenges for designing a performance-effective SFQ-based NPU (e.g., expensive on-chip data movements and buffering). Lastly, we present SuperNPU, our example SFQ-based NPU architecture, which effectively resolves the challenges. Our evaluation shows that the proposed design outperforms a conventional state-of-the-art NPU by 23 times. With free cooling provided as done in quantum computing, the performance per chip power increases up to 490 times. Our methodology can also be applied to other architecture designs with SFQ-friendly characteristics.
Koki Ishida, Ilkwon Byun, Ikki Nagaoka, Kosuke Fukumitsu, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Jangwoo Kim, Koji Inoue
MICRO6