EDBT 2026 Demo / reviewers in the wild / expert
Takatsugu Ono
dblp:15/7201
· DBLP profile ↗
9ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-2348-0249ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Design of a Correlation-Insensitive HFQ Stochastic Adder by Local Two-Phase ClockingabstractComputing technologies based on superconducting Josephson junctions, such as single-flux-quantum (SFQ) and half-flux-quantum (HFQ) circuits, are promising as next-generation computing platforms for their exceptional speed and low power consumption. Stochastic computing (SC), a non-deterministic computation approach known for its superior area efficiency, has attracted interest due to its compatibility with SFQ logic and its potential applications in hardware for machine learning. Existing multiplexer (MUX)-based SC adders using SFQ technology encounter challenges with accuracy, area efficiency, and power consumption. In this paper, we propose a low-power, compact, and correlation-insensitive SC adder that leverages the unique characteristics of the confluence buffer, a fundamental circuit element in SFQ and HFQ technologies. By adopting local two-phase clocking, our SC adders avoid computation inaccuracy due to correlation of input bit streams. Our evaluation results show that, compared with existing SFQ MUX-based SC adders, the proposed SFQ-based design achieves approximately 50% reductions in the number of Josephson junctions, implementation area, and power consumption. Furthermore, the proposed HFQ-based design reduces power consumption to approximately 1/3rd to 1/5th of SFQ-based designs for the same manufacturing technology. Yuki Matsumoto, Masamitsu Tanaka, Takatsugu Ono |
ISLPED | 3 |
| 2023 | Evaluating floating-point multipliers with opto-electrical hybrid circuitsabstractNanophotonic technology has ultra-low latency and low energy consumption characteristics, and several previous research have shown the potential of photonic-based computation. However, existing optical circuits are based on low-bit-width integer operations with analog manners, so there is less research on optical floating-point operations. Providing the benefits of analog optical circuits (i.e., low latency and low energy) to floating-point multipliers (FMs) requires careful consideration. In particular, introducing optical circuits unnecessarily increases the overhead of optical-to-electrical (O/E) and analog-to-digital (A/D) conversion, thus reducing system performance efficiency. Appropriate processing balancing between electrical and optical circuits is an important issue in OE hybrid circuit design. This study evaluates FM latency and energy consumption with different combinations of optical and electrical circuits. We design two versions of an Opto-Electrical FM (OEFM): M-OEFM, an optical implementation of an integer multiplier in an FM, and MA-OEFM, an optical implementation of an integer multiplier and adder functions. These designs are based on a policy of priority optical circuit implementation of the dominant components of conventional electrical circuits. Experimental results indicate that the M-OEFM achieved a 56 % reduction in latency and a 41 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM achieved an 88 % reduction in latency and a 19 % reduction in energy consumption compared with conventional electric circuits. The MA-OEFM is superior to M-OEFM in terms of energy-delay product (EDP). The results also indicate that appropriate co-design of analog optical and digital electrical circuits enables highly efficient FMs. Takumi Inaba, Takatsugu Ono, Koji Inoue, Satoshi Kawakami |
CF | 2 |
| 2022 | Design of Variable Bit-Width Arithmetic Unit Using Single Flux Quantum DeviceabstractThis paper presents the design of an ultra-high-speed, low-power arithmetic unit that supports variable bit-width operations with single flux quantum (SFQ) technology. Because of the high-speed nature of superconductor devices, we can achieve extremely high power-performance efficiency that cannot be achieved by state-of-the-art CMOS devices. To implement the complex function to support the variable bit-width feature, we introduce a novel circuit architecture to maintain the high-speed operation over 50GHz. Our prototype chip design successfully demonstrated 53.5GHz 1.59mW operations. Iori Ishikawa, Ikki Nagaoka, Ryota Kashima, Koki Ishida, Kosuke Fukumitsu, Keitarou Oka, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Akira Fujimaki, Koji Inoue |
ISCAS | 10 |
| 2020 | Enhancing a manycore-oriented compressed cache for GPGPUabstractGPUs can achieve high performance by exploiting massive-thread parallelism. However, some factors limit performance on GPUs, one of which is the negative effects of L1 cache misses. In some applications, GPUs are likely to suffer from L1 cache conflicts because a large number of cores share a small L1 cache capacity. A cache architecture that is based on data compression is a strong candidate for solving this problem as it can reduce the number of cache misses. Unlike previous studies, our data compression scheme attempts to exploit the value locality existing within not only intra cache lines but also inter cache lines. We enhance the structure of a last-level compression cache proposed for general purpose manycore processors to optimize against shared L1 caches on GPUs. The experimental results reveal that our proposal outperforms the other compression cache for GPUs by 11 points on average. Keitarou Oka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Koji Inoue |
HPC Asia | 4 |
| 2020 | SuperNPU: An Extremely Fast Neural Processing Unit Using Superconducting Logic DevicesabstractSuperconductor single-flux-quantum (SFQ) logic family has been recognized as a highly promising solution for the post-Moore's era, thanks to its ultra-fast and low-power switching characteristics. Therefore, researchers have made a tremendous amount of effort in various aspects to promote the technology and automate its circuit design process (e.g., low-cost fabrication, design tool development). However, there has been no progress in designing a convincing SFQ-based architectural unit due to the architects' lack of understanding of the technology's potentials and limitations at the architecture level. In this paper, we present how to architect an SFQ-based architectural unit by providing design principles with an extreme-performance neural processing unit (NPU). To achieve the goal, we first implement an architecture-level simulator to model an SFQ-based NPU accurately. We validate this model using our die-level prototypes, design tools, and logic cell library. This simulator accurately measures the NPU's performance, power consumption, area, and cooling overheads. Next, driven by the modeling, we identify key architectural challenges for designing a performance-effective SFQ-based NPU (e.g., expensive on-chip data movements and buffering). Lastly, we present SuperNPU, our example SFQ-based NPU architecture, which effectively resolves the challenges. Our evaluation shows that the proposed design outperforms a conventional state-of-the-art NPU by 23 times. With free cooling provided as done in quantum computing, the performance per chip power increases up to 490 times. Our methodology can also be applied to other architecture designs with SFQ-friendly characteristics. Koki Ishida, Ilkwon Byun, Ikki Nagaoka, Kosuke Fukumitsu, Masamitsu Tanaka, Satoshi Kawakami, Teruo Tanimoto, Takatsugu Ono, Jangwoo Kim, Koji Inoue |
MICRO | 8 |
| 2019 | Evaluating the Impact of Energy Efficient Networks on HPC WorkloadsabstractInterconnection networks grow larger as supercomputers include more nodes and require higher bandwidth for performance. This scaling significantly increases the fraction of power consumed by the network, by increasing the number of network components (links and switches). Typically, network links consume power continuously once they are turned on. However, recent proposals for energy efficient interconnects have introduced low-power operation modes for periods when network links are idle. Low-power operation can increase messaging time when switching a link from low-power to active operation. We extend the TraceR-CODES network simulator for power modeling to evaluate the impact of energy efficient networking on power and performance. Our evaluation presents the first study on both single-job and multi-job execution to realistically simulate power consumption and performance under congestion for a large-scale HPC network. Results on several workloads consisting of HPC proxy applications show that single-job and multi-job execution favor different modes of low power operation to have significant power savings at the cost of minimal performance degradation. Giorgis Georgakoudis, Takatsugu Ono, Koji Inoue, Shinobu Miwa, Abhinav Bhatele |
HiPC | 3 |
| 2016 | Evaluating the impacts of code-level performance tunings on power efficiencyabstractAs the power consumption of HPC systems will be a primary constraint for exascale computing, a main objective in HPC communities is recently becoming to maximize power efficiency (i.e., performance per watt) rather than performance. Although programmers have spent a considerable effort to improve performance by tuning HPC programs at a code level, tunings for improving power efficiency is now required. In this work, we select two representative HPC programs (Graph500 and SDPARA) and evaluate how traditional code-level performance tunings applied to these programs affect power efficiency. We also investigate the impacts of the tunings on power efficiency at various operating frequencies of CPUs and/or GPUs. The results show that the tunings significantly improve power efficiency, and different types of tunings exhibit different trends in power efficiency by varying CPU frequency. Finally, the scalability and power efficiency of state-of-the-art Graph500 implementations are explored on both a single-node platform and a 960-node supercomputer. With their high scalability, they achieve 27.43 MTEPS/Watt with 129.76 GTEPS on the single-node system and 4.39 MTEPS/Watt with 1,085.24 GTEPS on the supercomputer. Satoshi Imamura, Keitarou Oka, Yuichiro Yasui, Yuichi Inadomi, Katsuki Fujisawa, Toshio Endo, Koji Ueno, Keiichiro Fukazawa, Nozomi Hata, Yuta Kakibuka, Koji Inoue, Takatsugu Ono |
IEEE BigData | 12 |
| 2014 | FlexDAS: A flexible direct attached storage for I/O intensive applicationsabstractBig data analysis and a data storing applications require a huge volume of storage and a high I/O performance. Applications can achieve high levels of performance and cost efficiency by exploiting the high I/O performances of direct attached storages (DAS) such as internal HDDs. With the size of stored data ever increasing, it will be difficult to replace servers since internal HDDs contain huge amounts of data. In response to this issue, we propose FlexDAS, which improves the flexibility of direct attached storage by using a disk area network (DAN) without degrading the I/O performance. We developed a prototype FlexDAS switch and quantitatively evaluated the architecture. Results show that the FlexDAS switch can disconnect and connect the HDD to the server in just 1.16 seconds. The I/O performances of the disks connected via the FlexDAS switch were almost the same as the conventional DAS architecture. Takatsugu Ono, Yotaro Konishi, Teruo Tanimoto, Noboru Iwamatsu, Takashi Miyoshi, Jun Tanaka |
IEEE BigData | 1 |
| 2014 | Hardware-assisted scalable flow control of shared receive queueabstractThe total number of processor cores in supercomputers is increasing while memory size per core is decreasing due to the adoption of processors with multiple cores. Shared Receive Queue is a technique that effectively reduces the memory usage of buffers, but the absence of flow control results in excess buffer pools. We propose a hardware-assisted flow control that reduces flow control latency by 95.1%, thus enabling scalable supercomputers with multi-core processors. Teruo Tanimoto, Takatsugu Ono, Kohta Nakashima, Takashi Miyoshi |
ICS | 2 |