EDBT 2026 Demo / reviewers in the wild / expert
Atsutake Kosuge
dblp:68/11175
· DBLP profile ↗
22ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-3394-2227ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 4 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analysis and Design of Oblong Coils and Standard-Cell-Based Receiver for Area-Efficient Edge-Coupled Inductive Coupling TransceiverabstractProximity inductive coupling interfaces provide a low-cost, high-yield solution for 3D assembly, thanks to their compatibility with standard CMOS processes. However, they suffer from challenges related to the design complexity of both the coil and the receiver. To address these issues, this work proposes a comprehensive approach that includes an analytical coil design methodology applicable to edge-coupled configurations, an oblong coil structure to improve layout efficiency, and a standard-cell-based receiver architecture that enables simplified and scalable implementation. The proposed oblong coil achieves a 4.5 times improvement in area efficiency compared to traditional square coils, while maintaining adequate coupling strength and crosstalk tolerance, as validated through a test chip fabricated in a 40 nm CMOS process. The proposed receiver leverages bias sharing and a digitally tunable, standard-cell-based hysteresis comparator, resulting in 0.23 times the area and 0.37 times the energy consumption relative to a conventional analog comparator, as confirmed through simulations in a 16 nm FinFET process. Yuki Mitarai, Mototsugu Hamada, Atsutake Kosuge |
ASP-DAC | 3 |
| 2026 | A 3-day ASIC Design Hands-on with the Minimal Fab
Hideharu Amano, Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Tohru Mogami, Yoshio Mita |
ISCAS | 2 |
| 2026 | A 28-nm 0.8M-Weights/mm2 9.1-TOPS/mm2 All-Analog SRAM-Based Compute-in-Memory Macro Using Fine-Grained Structured Pruning With Adaptive-Ranging ADC
Kota Shiba, Zhijie Zhan, Koji Nii, Yih Wang, Tsung-Yung Jonathan Chang, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A Coarse- and Fine-Grained LUT Segmentation Method Enabling Single FPGA Implementation of Wired-Logic DNN ProcessorabstractA coarse- and fine-grained LUT segmentation technique is developed for wired-logic AI processors to improve FPGA resource utilization efficiency. By applying the proposed technique to FPGA-based wired-logic processors used for CIFAR-10 classification and keyword spotting, the hardware resource requirements for nonlinear functions were reduced by 92% and 92.8%, respectively, with negligible accuracy degradation. Dongzhu Li, Mototsugu Hamada, Atsutake Kosuge |
ASP-DAC | 4 |
| 2025 | An SoC Design and Fabrication Hands-On Educational Course within One Week Using Structured ASICabstractSince the design and manufacturing of semiconductor chips takes a considerable amount of time, it is difficult to fully understand the entire process and conduct student experiments that provide a hands-on experience of chip creation. The lack of student experiments that allow beginners to easily experience the process from chip design to manufacturing is one of the reasons why there are few students interested in semiconductor research. This, in turn, exacerbates the shortage of skilled workers in the semiconductor industry. The Agile-Chip platform is a method that enables the rapid and low-cost production of small quantities of chips by manufacturing only the top wiring layer using minimal fab technology on cut wafers, which are pre-cut from mass-produced wafers except for the top-most wiring. In this paper, we propose a student experiment method for semiconductor beginners using the Agile-Chip platform and provide an implementation example. For the base chip, a gate array is used, which allows the configuration of various gates using only the top wiring layer. The students design a 61-stage ring oscillator using the gate array, verify its operation through simulation, and then proceed with the corresponding layout design. After confirming the consistency of both, they generate a GDS file. This wiring layer is then fabricated either in a minimal fab or a cleanroom, and finally, the chip is mounted on a substrate for measurement. Hideharu Amano, Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Tohru Mogami, Yoshio Mita |
ISCAS | 2 |
| 2025 | A 83.7% Resource Reduced FPGA-based Wired-Logic DNN Processor by Using Mixed-Precision Module Embedding Into Non-Linear Function LUTabstractWired-logic processor architecture is a promising technology for energy-efficient FPGA-based DNN processors by eliminating power-intensive DRAM/BRAM accesses. A key challenge of wired-logic architectures is the substantial hardware resource requirement to implement all neurons and synapses on a single FPGA. While our proposed non-linear neural network (NNN) mitigates this issue by leveraging its high sparsity and binarized weights, the long bit-width of activation values remains a bottleneck, leading to considerable resource consumption and limiting the scalability of DNN models. In this paper, two techniques are proposed to address this challenge: (1) a mixed-precision activation quantization and dequantization module embedded within non-linear function look-up table (NLF-LUT), and (2) input bit-width compression for the NLF-LUT using a clip module and non-uniform step approximation (NSA). These optimizations achieve an 83.7% reduction in hardware resource usage without incurring additional computational overhead or accuracy degradation. Mototsugu Hamada, Atsutake Kosuge |
ISCAS | 3 |
| 2025 | Agile-X: A Structured-ASIC Created With a Mask-Less Lithography System Enabling Low-Cost and Agile Chip FabricationabstractScaling to finer CMOS process nodes necessitates more masks, resulting in higher costs and extended turnaround times (TATs). High costs and long TATs have hindered researchers outside the field of integrated circuits, including those in medicine, physics, and science from prototyping their own chips. Therefore, opportunities for diverse innovations in integrated circuits and talent development have been limited. We have developed the Agile-X platform for low-cost, rapid manufacturing of system-on-chips. Users can implement their own dedicated circuits with gate-array circuits on a base chip, which has common intellectual properties (IPs) such as RISC-V CPUs, various IOs, and ADCs. The base chip is manufactured in a foundry up to the intermediate metal layers and shipped with metal deposition on its surface. By directly drawing wiring patterns on this base chip with a mask-less lithography system, custom chips can be manufactured on-site without masks. As this process only requires wiring and eliminates masks, production time is drastically reduced compared to traditional full-mask wafer processes and multiproject wafer (MPW) shuttles. Development and manufacturing costs for the base chip, including preintegrated IPs, are shared among all Agile-X users. This reduces both IP and base-chip wafer costs per user. We prototyped wafers using a 0.18-$\mu $m CMOS process and tested the proposed structured ASIC platform and manufacturing process using mask-less lithography systems. The results indicate that the process from inputting GDS data to lithography and dry etching can be completed within 30 min, and custom application-specific integrated circuits (ASICs) can be manufactured within a day. Compared with full-mask wafer design and manufacturing, the manufacturing cost per chip, including IP costs, is reduced from 271000 USD to 22 USD, a reduction of 1/12252, and the manufacturing period is reduced from 20 days to 30 min, a reduction of 1/960. Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Hideharu Amano, Tohru Mogami, Yoshio Mita, Tadahiro Kuroda |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Efficient FPGA Resource Utilization in Wired-Logic Processors Using Coarse and Fine Segmentation of LUTs for Non-Linear FunctionsabstractA coarse- and fine-grained lookup table (LUT) segmentation technique is developed for wired-logic artificial intelligence (AI) processors to improve field-programmable gate array (FPGA) resource utilization efficiency. While wired-logic processors have achieved several orders of magnitude higher energy efficiency than conventional FPGA-based deep neural network (DNN) processors on the CIFAR-10 dataset by eliminating DRAM/BRAM access during inference processing, huge hardware resources are required for the large-scale DNNs with long-bit-width data. Implementing even small DNNs proves challenging as they surpass the hardware resources available in commercial FPGAs. To address these issues and enable the implementation of larger-scale neural networks alongside the processing of long-bit-width data, two techniques are proposed: (1) an LUT segmentation technique based on coarse and fine granularity, and (2) accuracy optimization through the incorporation of redundant bits. The application of these proposed techniques to state-of-the-art wired-logic processors markedly enhances the scalability of a single FPGA, thereby facilitating the implementation of larger-scale neural networks across various tasks, including CIFAR-10 classification and keyword spotting. The hardware resource requirements for non-linear functions in processing elements decreased by 92%, and 92.8%, respectively. Remarkably, the recognition accuracy for CIFAR-10 remains consistent, while there is a negligibly small degradation in accuracy for the keyword spotting task by 1.2%. Dongzhu Li, Kenji Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
ISCAS | 4 |
| 2023 | A Fully Synthesized 13.7μJ/Prediction 88% Accuracy CIFAR-10 Single-Chip Data-Reusing Wired-Logic Processor Using Non-Linear Neural NetworkabstractAn FPGA-based wired-logic CNN processor is presented that can process CIFAR-10 at 13.7μJ/prediction with an 88% accuracy, which is 2,036 times more energy-efficient than the prior state-of-the-art FPGA-based processor. Energy efficiency is greatly improved by implementing all processing elements and wirings in parallel on a single FPGA chip to eliminate the memory access. By utilizing both (1) a non-linear neural network which saves on neurons and synapses and (2) a shift register-based wired-logic architecture, hardware resource usage is reduced by three orders of magnitude. Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda |
ASP-DAC | 2 |
| 2023 | A 1.2nJ/Classification Fully Synthesized All-Digital Asynchronous Wired-Logic Processor Using Quantized Non-Linear Function Blocks in 0.18μm CMOSabstractA 5.3 times smaller and 2.6 times more energy-efficient all-digital wired-logic processor which infers MNIST with 90.6% accuracy and 1.2nJ of energy consumption has been developed. To improve area efficiency of wired-logic architecture, nonlinear neural network (NNN), which is a neuron and synapse efficient network, and logical compression technology to implement it with area-saving and low-power digital circuits by logic synthesis are proposed, and asynchronous digital combinational circuit DNN hardware has been developed. Rei Sumikawa, Kota Shiba, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
ASP-DAC | 3 |
| 2023 | An Occlusion-Resilient mmWave Imaging Radar-Based Object Recognition System Using Synthetic Training Data Generation TechniqueabstractAn occlusion-resilient mmWave imaging radar-based object recognition system for advanced driver-assistance systems (ADAS) of construction machinery application is developed. As ADAS for construction sites, millimeter wave application is required in poor visibility environments such as nighttime, bad weather, and muddy conditions where object recognition by RGB cameras and LiDAR is difficult. A remaining technical challenge for ADAS is occlusion. Two techniques are proposed to improve the accuracy in occlusion scenes. First is a technique which generates simulated training data for occlusion environment to improve accuracy while reducing the cost for the training data preparation. The second is a parallel inference DNN architecture which enables object recognition with high accuracy in both normal and occlusion scenes by running two DNNs optimized respectively for normal and occlusion scenes in parallel. The object recognition accuracy of mAP50in occlusion scenes improves by 15 points compared to the conventional technique. The decrease in recognition accuracy in non-occlusion scenes is only 4 points. Eitaro Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
IECON | 2 |
| 2023 | A 0.13mJ/Prediction CIFAR-100 Raster-Scan- Based Wired-Logic Processor Using Non-Linear Neural NetworkabstractA 0.13mJ/prediction with 68.6% accuracy single- chip wired-logic artificial intelligence (AI) processor is developed in a 16nm field-programmable gate array (FPGA). Compared with conventional von-Neumann architecture-based AI processors, the energy efficiency is greatly improved by eliminating the DRAM/BRAM access. A technical challenge of the conventional wired-logic processor is the large amount of hardware resources required. To implement a large convolutional neural network (CNN) into a single FPGA chip, two techniques are used: (1) a sparse neural network which is called non-linear neural network (NNN), and (2) a newly developed raster-scan-based wired-logic architecture. The amount of hardware resources required is reduced by a factor of 5.4. Compared with the state-of-the-art FPGA-based processor, 238 times better energy efficiency is achieved with the same accuracy on the CIFAR-I00 task. In addition, 7 times better energy efficiency is achieved compared with the state-of- the-art application-specific integrated circuit (ASIC) processor. Dongzhu Li, Yao-Chung Hsu, Rei Sumikawa, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
ISCAS | 4 |
| 2023 | Polyomino: A 3D-SRAM-Centric Accelerator for Randomly Pruned Matrix Multiplication With Simple Reordering Algorithm and Efficient Compression Format in 180-nm CMOSabstractWe have developed a sparse matrix reordering algorithm with a novel 3D-SRAM-centric Polyomino accelerator that enables efficient processing of the reordered matrix for parameter compression. By reordering randomly pruned, irregularly structured sparse matrices into regularly structured matrices, both the compression ratio of the data and the efficiency of the hardware processing increase. The reordering algorithm can be implemented simply by attributing it to the widely known k-sum problem. We also developed a compression format for storing the reordered matrices and show that the reordered regular structure can reduce the amount of required memory by 63% compared with the conventional method. The proposed Polyomino accelerator can efficiently process reordered matrices by using a 3D stacked SRAM, which is an external memory with random accessibility and low latency. The measurement results using a test chip fabricated in a 180-nm CMOS process demonstrate that the proposed accelerator can achieve high area-efficiency and high energy-efficiency and scales well with the pruning rate. Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A 5.2GHz RFID Chip Contactlessly Mountable on FPC at any 90-Degree Rotation and Face OrientationabstractThis paper presents an RFID Chip contactlessly mountable on an FPC having an antenna pattern. Inductive coupling between the FPC and the chip realizes low-cost bonding-less implementation. It is also possible to place the chip on the FPC at any angle of 0/90/180/270 degrees and face-up or face-down. Simulation shows the antenna gain is almost the same irrespective of the chip placement angle and face orientation. The experimental results confirmed that the proposed RFID chip works at upto 20cm away from a reader whose output power is 15dBm, achieving the same figure-of-merit as a conventionally bonded module. Reiji Miura, Saito Shibata, Masahiro Usui, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
ASP-DAC | 4 |
| 2022 | A 13.7μJ/prediction 88% Accuracy CIFAR-10 Single-Chip Wired-logic Processor in 16-nm FPGA using Non-Linear Neural Networkabstract• In this study, we propose a 13.7mJ/prediction 88% accuracy CIFAR-10 single-chip wired-logic processor in 16-nm FPGA by utilizing a newly developed 98%-pruned ultra-sparse, binary-weight nonlinear neural network (NNN) and a shift-register based pipelined wired-logic architecture. Compared with the state-of-the-art FPGA-based processor, 2,036 times better energy efficiency is achieved. Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda |
HCS | 2 |
| 2022 | A 7-nm FinFET 1.2-TB/s/mm2 3D-Stacked SRAM with an Inductive Coupling Interface Using Over-SRAM Coils and Manchester-Encoded Synchronous TransceiversabstractA 0.7-pJ/bit, 8.5-Gbps/link inductive coupling inter-chip wireless communication interface for a 3D-stacked SRAM has been developed in a 7-nm FinFET process. A new physical placement method that allows coils to be placed over off-the-shelf SRAM macros with small magnetic field attenuation, together with the use of synchronous communication using Manchester encoding and a clocked comparator to enable the detection of small-swing signals, achieve a 26% reduction in SRAM die area compared to TSV-based stacking. Inter-chip communication at 0.7-pJ/bit, 8.5-Gbps/link was confirmed using test chips. A 4-hi 3D-stacked SRAM module using the proposed interface is estimated to achieve a 1.2-TB/s/mm2area efficiency, representing a two-orders-of-magnitude improvement over state-of-the-art 3D-stacked SRAM. Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda |
HCS | 3 |
| 2022 | Proximity Wireless Communication Technologies: An Overview and Design GuidelinesabstractThis paper presents an overview of proximity wireless communication (PWC) technologies, their principles, design guidelines and practical applications. In particular, two different applications of PWC are reviewed. One is PWC between stacked chips. Both communication distance and coupler size are several tens of microns. Area and energy efficient design techniques are introduced. Another is PWC between module boards. Both communication distance and coupler size are several millimeters. Energy and area efficient practical designs are introduced for mobile and industrial machinery applications. Atsutake Kosuge, Tadahiro Kuroda |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2019 | A 4.8x Faster FPGA-Based Iterative Closest Point Accelerator for Object Pose Estimation of Picking Robot ApplicationsabstractAn FPGA-based accelerator for the iterative-closest-point (ICP) algorithm has been proposed, which achieves 4.8-times-faster object-pose estimation by a picking robot compared with the state-of-the-art technique. Experiments of the proposed FPGA-based ICP accelerator using Amazon Picking Contest data sets have confirmed that the object-pose estimation by the ICP takes only 0.6 seconds, and the entire picking process takes 2.0 seconds with power consumption of 6.0 W. Atsutake Kosuge, Keisuke Yamamoto, Yukinori Akamine, Taizo Yamawaki, Takashi Oshima |
FCCM | 1 |
| 2016 | Analytical thruchip inductive coupling channel design optimizationabstractThruChip interface (TCI) is an emerging 3-D integrated circuit stacking technology. TCI utilizes on-chip inductor to build vertical communication channel in near field distance and has been proved to stand comparison with through-siliconvia (TSV) in data rate, power, and reliability. Moreover, it is also cost-effective in manufacturing due to its wireless nature. In this paper, an analytical method is proposed to find near-optimal TCI inductive coupling channel solution. The experiment results show an average 16.8% transmitting current reduction and shrink design time from days to a few minutes. Li-Chung Hsu, Junichiro Kadomoto, So Hasegawa, Atsutake Kosuge, Yasuhiro Take, Tadahiro Kuroda |
ASP-DAC | 4 |
| 2015 | Design and analysis for ThruChip design for manufacturing (DFM)abstractA 1GB/s ThruChip interface (TCI) test chip for wafer thinning, power mesh, and dummy metal fill impacts are analyzed and evaluated with test chip measurement and field solver simulation. The measurement results show that TCI coil dimension can be sized down as wafer thinning by following D/Z=3 rule. However, the experiment shows 20% power reduction by enlarging TCI coil (D/Z=6). The power mesh lies between TCI coils can dramatically decrease the TCI magnetic pulse strength and hence cause TCI to fail. Dummy metal within TCI coils has no impact on TCI transmission Li-Chung Hsu, Yasuhiro Take, Atsutake Kosuge, So Hasegawa, Junichiro Kadomoto, Tadahiro Kuroda |
ASP-DAC | 3 |
| 2015 | Circuit and package design for 44GB/s inductive-coupling DRAM/SoC interfaceabstractA 44GB/s inductive-coupling DRAM/SoC interface is developed by PoP integration. It utilizes the advantages of both TSV and LPDDR by using a ThruChip Interface (TCI) and an ultra-thin fan-out wafer level package (UT-FOWLP). The TCI allows data communication between the stacked chips while the UT-FOWLP thins the chips stacking distance and provides the chips with power. This proposed DRAM/SoC interface outperforms WIO2 with TSV in terms of area efficiency (4× better), immunity from simultaneous switching output (SSO) noise (32× better) and manufacturing cost (40% cheaper). In addition, it outperforms LPDDR4 in PoP in terms of power dissipation (5× lower) and timing control easiness. The inductive-coupling interface is newly designed to allow 12× improvement on its area efficiency. By using overlapping coils with quadrature phase division multiplexing (PDM), the coil density is increased by 4 times. The coil density is further increased by 3 times by shortening communication distance with the UT-FOWLP. Akira Okada, Abdul Raziz Junaidi, Yasuhiro Take, Atsutake Kosuge, Tadahiro Kuroda |
ASP-DAC | 4 |
| 2013 | A 12.5Gb/s/link non-contact multi drop bus system with impedance-matched Transmission Line Couplers and Dicode partial-response channel transceiversabstractA reduced-reflection multi-drop bus system using Dicode (1-D) partial response signaling transceiver is presented for the first time in the world. Directional couplers on transmission lines arranged with equi-energy distributing and exact impedance matched conditions allow the bus to reach to 12.5Gbps/link speed, which is the world's fastest data link speed with multi-drop bus architecture. Dicode partial-response signaling method with a half-rate architecture was used where a precoder is placed in the transmitter to make the signal best fit for the channel to eliminate inter symbol interference (ISI). Atsutake Kosuge, Wataru Mizuhara, Noriyuki Miura, Masao Taguchi, Hiroki Ishikuro, Tadahiro Kuroda |
ASP-DAC | 1 |