Mototsugu Hamada

dblp:94/5977 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-0461-4208ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 1 first-author · 16 since 2021
YearPublicationVenuePosition
2026 Analysis and Design of Oblong Coils and Standard-Cell-Based Receiver for Area-Efficient Edge-Coupled Inductive Coupling Transceiver
abstract
Proximity inductive coupling interfaces provide a low-cost, high-yield solution for 3D assembly, thanks to their compatibility with standard CMOS processes. However, they suffer from challenges related to the design complexity of both the coil and the receiver. To address these issues, this work proposes a comprehensive approach that includes an analytical coil design methodology applicable to edge-coupled configurations, an oblong coil structure to improve layout efficiency, and a standard-cell-based receiver architecture that enables simplified and scalable implementation. The proposed oblong coil achieves a 4.5 times improvement in area efficiency compared to traditional square coils, while maintaining adequate coupling strength and crosstalk tolerance, as validated through a test chip fabricated in a 40 nm CMOS process. The proposed receiver leverages bias sharing and a digitally tunable, standard-cell-based hysteresis comparator, resulting in 0.23 times the area and 0.37 times the energy consumption relative to a conventional analog comparator, as confirmed through simulations in a 16 nm FinFET process.
Yuki Mitarai, Mototsugu Hamada, Atsutake Kosuge
ASP-DAC2
2026 A 28-nm 0.8M-Weights/mm2 9.1-TOPS/mm2 All-Analog SRAM-Based Compute-in-Memory Macro Using Fine-Grained Structured Pruning With Adaptive-Ranging ADC
Kota Shiba, Zhijie Zhan, Koji Nii, Yih Wang, Tsung-Yung Jonathan Chang, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 A Coarse- and Fine-Grained LUT Segmentation Method Enabling Single FPGA Implementation of Wired-Logic DNN Processor
abstract
A coarse- and fine-grained LUT segmentation technique is developed for wired-logic AI processors to improve FPGA resource utilization efficiency. By applying the proposed technique to FPGA-based wired-logic processors used for CIFAR-10 classification and keyword spotting, the hardware resource requirements for nonlinear functions were reduced by 92% and 92.8%, respectively, with negligible accuracy degradation.
Dongzhu Li, Mototsugu Hamada, Atsutake Kosuge
ASP-DAC3
2025 A 83.7% Resource Reduced FPGA-based Wired-Logic DNN Processor by Using Mixed-Precision Module Embedding Into Non-Linear Function LUT
abstract
Wired-logic processor architecture is a promising technology for energy-efficient FPGA-based DNN processors by eliminating power-intensive DRAM/BRAM accesses. A key challenge of wired-logic architectures is the substantial hardware resource requirement to implement all neurons and synapses on a single FPGA. While our proposed non-linear neural network (NNN) mitigates this issue by leveraging its high sparsity and binarized weights, the long bit-width of activation values remains a bottleneck, leading to considerable resource consumption and limiting the scalability of DNN models. In this paper, two techniques are proposed to address this challenge: (1) a mixed-precision activation quantization and dequantization module embedded within non-linear function look-up table (NLF-LUT), and (2) input bit-width compression for the NLF-LUT using a clip module and non-uniform step approximation (NSA). These optimizations achieve an 83.7% reduction in hardware resource usage without incurring additional computational overhead or accuracy degradation.
Mototsugu Hamada, Atsutake Kosuge
ISCAS2
2024 Efficient FPGA Resource Utilization in Wired-Logic Processors Using Coarse and Fine Segmentation of LUTs for Non-Linear Functions
abstract
A coarse- and fine-grained lookup table (LUT) segmentation technique is developed for wired-logic artificial intelligence (AI) processors to improve field-programmable gate array (FPGA) resource utilization efficiency. While wired-logic processors have achieved several orders of magnitude higher energy efficiency than conventional FPGA-based deep neural network (DNN) processors on the CIFAR-10 dataset by eliminating DRAM/BRAM access during inference processing, huge hardware resources are required for the large-scale DNNs with long-bit-width data. Implementing even small DNNs proves challenging as they surpass the hardware resources available in commercial FPGAs. To address these issues and enable the implementation of larger-scale neural networks alongside the processing of long-bit-width data, two techniques are proposed: (1) an LUT segmentation technique based on coarse and fine granularity, and (2) accuracy optimization through the incorporation of redundant bits. The application of these proposed techniques to state-of-the-art wired-logic processors markedly enhances the scalability of a single FPGA, thereby facilitating the implementation of larger-scale neural networks across various tasks, including CIFAR-10 classification and keyword spotting. The hardware resource requirements for non-linear functions in processing elements decreased by 92%, and 92.8%, respectively. Remarkably, the recognition accuracy for CIFAR-10 remains consistent, while there is a negligibly small degradation in accuracy for the keyword spotting task by 1.2%.
Dongzhu Li, Kenji Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ISCAS5
2023 A Fully Synthesized 13.7μJ/Prediction 88% Accuracy CIFAR-10 Single-Chip Data-Reusing Wired-Logic Processor Using Non-Linear Neural Network
abstract
An FPGA-based wired-logic CNN processor is presented that can process CIFAR-10 at 13.7μJ/prediction with an 88% accuracy, which is 2,036 times more energy-efficient than the prior state-of-the-art FPGA-based processor. Energy efficiency is greatly improved by implementing all processing elements and wirings in parallel on a single FPGA chip to eliminate the memory access. By utilizing both (1) a non-linear neural network which saves on neurons and synapses and (2) a shift register-based wired-logic architecture, hardware resource usage is reduced by three orders of magnitude.
Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC5
2023 A 1.2nJ/Classification Fully Synthesized All-Digital Asynchronous Wired-Logic Processor Using Quantized Non-Linear Function Blocks in 0.18μm CMOS
abstract
A 5.3 times smaller and 2.6 times more energy-efficient all-digital wired-logic processor which infers MNIST with 90.6% accuracy and 1.2nJ of energy consumption has been developed. To improve area efficiency of wired-logic architecture, nonlinear neural network (NNN), which is a neuron and synapse efficient network, and logical compression technology to implement it with area-saving and low-power digital circuits by logic synthesis are proposed, and asynchronous digital combinational circuit DNN hardware has been developed.
Rei Sumikawa, Kota Shiba, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC4
2023 An Occlusion-Resilient mmWave Imaging Radar-Based Object Recognition System Using Synthetic Training Data Generation Technique
abstract
An occlusion-resilient mmWave imaging radar-based object recognition system for advanced driver-assistance systems (ADAS) of construction machinery application is developed. As ADAS for construction sites, millimeter wave application is required in poor visibility environments such as nighttime, bad weather, and muddy conditions where object recognition by RGB cameras and LiDAR is difficult. A remaining technical challenge for ADAS is occlusion. Two techniques are proposed to improve the accuracy in occlusion scenes. First is a technique which generates simulated training data for occlusion environment to improve accuracy while reducing the cost for the training data preparation. The second is a parallel inference DNN architecture which enables object recognition with high accuracy in both normal and occlusion scenes by running two DNNs optimized respectively for normal and occlusion scenes in parallel. The object recognition accuracy of mAP50in occlusion scenes improves by 15 points compared to the conventional technique. The decrease in recognition accuracy in non-occlusion scenes is only 4 points.
Eitaro Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IECON3
2023 A 0.13mJ/Prediction CIFAR-100 Raster-Scan- Based Wired-Logic Processor Using Non-Linear Neural Network
abstract
A 0.13mJ/prediction with 68.6% accuracy single- chip wired-logic artificial intelligence (AI) processor is developed in a 16nm field-programmable gate array (FPGA). Compared with conventional von-Neumann architecture-based AI processors, the energy efficiency is greatly improved by eliminating the DRAM/BRAM access. A technical challenge of the conventional wired-logic processor is the large amount of hardware resources required. To implement a large convolutional neural network (CNN) into a single FPGA chip, two techniques are used: (1) a sparse neural network which is called non-linear neural network (NNN), and (2) a newly developed raster-scan-based wired-logic architecture. The amount of hardware resources required is reduced by a factor of 5.4. Compared with the state-of-the-art FPGA-based processor, 238 times better energy efficiency is achieved with the same accuracy on the CIFAR-I00 task. In addition, 7 times better energy efficiency is achieved compared with the state-of- the-art application-specific integrated circuit (ASIC) processor.
Dongzhu Li, Yao-Chung Hsu, Rei Sumikawa, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ISCAS5
2023 Polyomino: A 3D-SRAM-Centric Accelerator for Randomly Pruned Matrix Multiplication With Simple Reordering Algorithm and Efficient Compression Format in 180-nm CMOS
abstract
We have developed a sparse matrix reordering algorithm with a novel 3D-SRAM-centric Polyomino accelerator that enables efficient processing of the reordered matrix for parameter compression. By reordering randomly pruned, irregularly structured sparse matrices into regularly structured matrices, both the compression ratio of the data and the efficiency of the hardware processing increase. The reordering algorithm can be implemented simply by attributing it to the widely known k-sum problem. We also developed a compression format for storing the reordered matrices and show that the reordered regular structure can reduce the amount of required memory by 63% compared with the conventional method. The proposed Polyomino accelerator can efficiently process reordered matrices by using a 3D stacked SRAM, which is an external memory with random accessibility and low latency. The measurement results using a test chip fabricated in a 180-nm CMOS process demonstrate that the proposed accelerator can achieve high area-efficiency and high energy-efficiency and scales well with the pruning rate.
Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 A 5.2GHz RFID Chip Contactlessly Mountable on FPC at any 90-Degree Rotation and Face Orientation
abstract
This paper presents an RFID Chip contactlessly mountable on an FPC having an antenna pattern. Inductive coupling between the FPC and the chip realizes low-cost bonding-less implementation. It is also possible to place the chip on the FPC at any angle of 0/90/180/270 degrees and face-up or face-down. Simulation shows the antenna gain is almost the same irrespective of the chip placement angle and face orientation. The experimental results confirmed that the proposed RFID chip works at upto 20cm away from a reader whose output power is 15dBm, achieving the same figure-of-merit as a conventionally bonded module.
Reiji Miura, Saito Shibata, Masahiro Usui, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC5
2022 A 13.7μJ/prediction 88% Accuracy CIFAR-10 Single-Chip Wired-logic Processor in 16-nm FPGA using Non-Linear Neural Network
abstract
• In this study, we propose a 13.7mJ/prediction 88% accuracy CIFAR-10 single-chip wired-logic processor in 16-nm FPGA by utilizing a newly developed 98%-pruned ultra-sparse, binary-weight nonlinear neural network (NNN) and a shift-register based pipelined wired-logic architecture. Compared with the state-of-the-art FPGA-based processor, 2,036 times better energy efficiency is achieved.
Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
HCS5
2022 A 7-nm FinFET 1.2-TB/s/mm2 3D-Stacked SRAM with an Inductive Coupling Interface Using Over-SRAM Coils and Manchester-Encoded Synchronous Transceivers
abstract
A 0.7-pJ/bit, 8.5-Gbps/link inductive coupling inter-chip wireless communication interface for a 3D-stacked SRAM has been developed in a 7-nm FinFET process. A new physical placement method that allows coils to be placed over off-the-shelf SRAM macros with small magnetic field attenuation, together with the use of synchronous communication using Manchester encoding and a clocked comparator to enable the detection of small-swing signals, achieve a 26% reduction in SRAM die area compared to TSV-based stacking. Inter-chip communication at 0.7-pJ/bit, 8.5-Gbps/link was confirmed using test chips. A 4-hi 3D-stacked SRAM module using the proposed interface is estimated to achieve a 1.2-TB/s/mm2area efficiency, representing a two-orders-of-magnitude improvement over state-of-the-art 3D-stacked SRAM.
Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
HCS4
2021 Sub-10-μm Coil Design for Multi-Hop Inductive Coupling Interface
abstract
Sub-10-μm on-chip coils are designed and prototyped for the multi-hop inductive coupling interface in a 40-nm CMOS. Multi-layer coils and a new receiver circuit are employed to compensate the decrease of the coupling coefficient due to the small coil size. The prototype emulates a 3D stacked module with 8 dies in a 7-nm CMOS and shows that a 0.1-pJ/bit and 41-Tb/s/mm2 inductive coupling interface is achievable.
Tatsuo Omori, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC3
2021 A 3D-Stacked SRAM Using Inductive Coupling Technology for AI Inference Accelerator in 40-nm CMOS
abstract
A 3D-stacked SRAM using an inductive coupling wireless inter-chip communication technology (TCI) is presented for an AI inference accelerator. The energy and area efficiency are improved thanks to the introduction of a proposed low-voltage NMOS push-pull transmitter and a 12:1 SerDes. A termination scheme to short unused open coils is proposed to eliminate the ringing in an inductive coupling bus. Test chips were fabricated in a 40-nm CMOS technology confirming 0.40-V operation of the proposed transmitter with successful stacked SRAM operation.
Kota Shiba, Tatsuo Omori, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC3
2021 A 96-MB 3D-Stacked SRAM Using Inductive Coupling With 0.4-V Transmitter, Termination Scheme and 12: 1 SerDes in 40-nm CMOS
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 35% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates for the degradation of the first transmitted bit, a coil termination scheme that aims to eliminate the ringing of 3D inductive coupling bus, and a 12:1 SerDes that minimizes power consumption and area overhead in inductive coupling channels. Low-power, large-capacity, 3-cycle latency 3D-stacked SRAM for a DNN accelerator is achieved with the combination of these techniques to serve as a replacement of 3D-stacked DRAM. The performance of the proposed 3D-SRAM is compared with HBM DRAM and achieves more than 50% lower energy consumption. The scaling scenario of the SRAM module is discussed in light of the scaling of the inductive coupling technology and logic process.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.6
2020 A 3D-Stacked SRAM using Inductive Coupling with Low-Voltage Transmitter and 12: 1 SerDes
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 45% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates the degradation of the first signal, a coil termination scheme that aims to eliminate the noise of 3D inductive coupling bus, and a 12:1 SerDes. The data density of the SRAM should reach 12.3-MB/mm3, which extends beyond that of state-of-the-art stacked DRAMs.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Kota Ando, Kazutoshi Hirose, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
ISCAS8
2019 Live Demonstration: A Non-Contact Transmission Line Connector for USB3.1 HD-Video Streaming
abstract
This demonstration shows a high data rate, non-contact connector using a transmission line coupler (TLC) composed of two differential transmission lines. The TLC is an impedance-matched and wide-bandwidth coupler, therefore it enables the high-speed baseband communication at 5Gbps/lane for USB3.1. Through a non-contact connector using the TLC, a 4K monitor is connected to a smartphone via the super speed wired signal lanes(SSTX/RX). The TLC connector can be used under harsh environment such as water, dust and misalignment. This demonstration provides visitors with opportunities to find utility of the TLC connector.
Tomoya Arakawa, Joshin Sone, Mitsuji Okada, Mototsugu Hamada, Tadahiro Kuroda
ISCAS4
2018 Design Methodology in Wireless Power Transfer System for 3-D Stacked Multiple Receivers
abstract
This paper proposes a design methodology in a wireless power transfer system for 3-D stacked multiple receivers. A 1:m selective power transfer system is realized by introducing a frequency/time division multiplexing system. The power transfer function is analytically formulated and an optimization methodology of multiple tuning capacitors values is proposed and compared with simulation results. By using the optimized values, power transfer efficiencies at 6.78MHz and 13.56MHz are simulated to be 80% and 84%, respectively. Crosstalk and load transient performance is also compared with 1:1 system and confirmed that its degradation is not significant.
Shusuke Yanagawa, Ryota Shimizu, Mototsugu Hamada, Toru Shimizu, Tadahiro Kuroda
ISCAS3
2009 RF-analog circuit design in scaled SoC
abstract
Downscaling of process technology increases the development cost of RFCMOS SoC. Therefore, designers have to minimize the number of respins, and have to try to obtain higher yield. RFCMOS SoC consists of RF-analog, mixed-signal, logic and memory circuits. In order to realize a small number of respins number and higher yield, key issues are robust design methodology of RF-analog circuits, and full-chip verification. This paper describes practical techniques corresponding to those issues.
Nobuyuki Itoh, Mototsugu Hamada
ASP-DAC2
2007 An automated runtime power-gating scheme
abstract
An automated runtime power-gating scheme to reduce the leakage power in the active mode is presented in this paper. We propose a circuit that generates a sleep control signal from a clock-gating control signal automatically. By the combination of selective MT-CMOS scheme, the generated sleep control signal, and a novel flip-flop circuit with an additional latch function, a zero-wait transition from a sleep mode to an active mode is enabled. The additional latch function required for the zero-wait transition is achieved by only 6 transistors in addition to a conventional flip-flop. By the scheme, any design with the clock-gating scheme can be transformed automatically to a power-gated design while keeping the system operation the same in terms of the cycle accuracy. The scheme is applied to an MPEG4/H.264 audio/video codec and 21% power saving is achieved in the active mode while keeping the area overhead only 16% in a 90nm CMOS design.
Mototsugu Hamada, Takeshi Kitahara, Naoyuki Kawabe, Hironori Sato, Tsuyoshi Nishikawa, Takayoshi Shimazawa, Takahiro Yamashita, Hiroyuki Hara, Yukihito Oowaki
ICCD1
2006 Conditional Data Mapping Flip-Flops for Low-Power and High-Performance Systems
abstract
This paper introduces a new family of low-power and high-performance flip-flops, namely conditional data mapping flip-flops (CDMFFs), which reduce their dynamic power by mapping their inputs to a configuration that eliminates redundant internal transitions. We present two CDMFFs, having differential and single-ended structures, respectively, and compare them to the state-of-the-art flip-flops. The results indicate that both CDMFFs have the best power-delay product in their groups, respectively. In the aspect of power dissipation, the single-ended and differential CDMFFs consume the least power at data activity less than 50%, and are 31% and 26% less power than the conditional capture flip-flops at 25% data activity, respectively. In the aspect of performance, CDMFFs achieve small data-to-output delays, comparable to those of the transmission-gate pulsed latch and the modified-sense-amplifier flip-flop. In the aspect of timing reliability, CDMFFs have the best internal race immunity among pulse-triggered flip-flops. A post-layout case study is demonstrated with comparison to a transmission-gate flip-flop. The results indicate the single-ended CDMFF has 34% less in data-to-output delay and 28% less in power at 25% data activity, in spite of the 34% increase in size
Chen Kong Teh, Mototsugu Hamada, Tetsuya Fujita, Hiroyuki Hara, N. Ikumi, Yukihito Oowaki
IEEE Trans. Very Large Scale Integr. Syst.2
1998 Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques
abstract
This paper describes a fully automated low-power design methodology in which three different voltage-scaling techniques are combined together. Supply voltage is scaled globally, selectively, and adaptively while keeping the performance. This methodology enabled us to design an MPEG4 codec core with 58% less power than the original in three week turn-around-time.
Kimiyoshi Usami, Mutsunori Igarashi, Takashi Ishikawa, Masahiro Kanazawa, Masafumi Takahashi, Mototsugu Hamada, Hideho Arakida, Toshihiro Terazawa, Tadahiro Kuroda
DAC6