Tadahiro Kuroda

dblp:97/1102 · DBLP profile ↗
← Back
58ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0003-0617-1057ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A 28-nm 0.8M-Weights/mm2 9.1-TOPS/mm2 All-Analog SRAM-Based Compute-in-Memory Macro Using Fine-Grained Structured Pruning With Adaptive-Ranging ADC
Kota Shiba, Zhijie Zhan, Koji Nii, Yih Wang, Tsung-Yung Jonathan Chang, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 Agile-X: A Structured-ASIC Created With a Mask-Less Lithography System Enabling Low-Cost and Agile Chip Fabrication
abstract
Scaling to finer CMOS process nodes necessitates more masks, resulting in higher costs and extended turnaround times (TATs). High costs and long TATs have hindered researchers outside the field of integrated circuits, including those in medicine, physics, and science from prototyping their own chips. Therefore, opportunities for diverse innovations in integrated circuits and talent development have been limited. We have developed the Agile-X platform for low-cost, rapid manufacturing of system-on-chips. Users can implement their own dedicated circuits with gate-array circuits on a base chip, which has common intellectual properties (IPs) such as RISC-V CPUs, various IOs, and ADCs. The base chip is manufactured in a foundry up to the intermediate metal layers and shipped with metal deposition on its surface. By directly drawing wiring patterns on this base chip with a mask-less lithography system, custom chips can be manufactured on-site without masks. As this process only requires wiring and eliminates masks, production time is drastically reduced compared to traditional full-mask wafer processes and multiproject wafer (MPW) shuttles. Development and manufacturing costs for the base chip, including preintegrated IPs, are shared among all Agile-X users. This reduces both IP and base-chip wafer costs per user. We prototyped wafers using a 0.18-$\mu $m CMOS process and tested the proposed structured ASIC platform and manufacturing process using mask-less lithography systems. The results indicate that the process from inputting GDS data to lithography and dry etching can be completed within 30 min, and custom application-specific integrated circuits (ASICs) can be manufactured within a day. Compared with full-mask wafer design and manufacturing, the manufacturing cost per chip, including IP costs, is reduced from 271000 USD to 22 USD, a reduction of 1/12252, and the manufacturing period is reduced from 20 days to 30 min, a reduction of 1/960.
Atsutake Kosuge, Hirofumi Sumi, Naonobu Shimamoto, Yukinori Ochiai, Yurie Inoue, Hideharu Amano, Tohru Mogami, Yoshio Mita, Tadahiro Kuroda
IEEE Trans. Very Large Scale Integr. Syst.10
2024 Efficient FPGA Resource Utilization in Wired-Logic Processors Using Coarse and Fine Segmentation of LUTs for Non-Linear Functions
abstract
A coarse- and fine-grained lookup table (LUT) segmentation technique is developed for wired-logic artificial intelligence (AI) processors to improve field-programmable gate array (FPGA) resource utilization efficiency. While wired-logic processors have achieved several orders of magnitude higher energy efficiency than conventional FPGA-based deep neural network (DNN) processors on the CIFAR-10 dataset by eliminating DRAM/BRAM access during inference processing, huge hardware resources are required for the large-scale DNNs with long-bit-width data. Implementing even small DNNs proves challenging as they surpass the hardware resources available in commercial FPGAs. To address these issues and enable the implementation of larger-scale neural networks alongside the processing of long-bit-width data, two techniques are proposed: (1) an LUT segmentation technique based on coarse and fine granularity, and (2) accuracy optimization through the incorporation of redundant bits. The application of these proposed techniques to state-of-the-art wired-logic processors markedly enhances the scalability of a single FPGA, thereby facilitating the implementation of larger-scale neural networks across various tasks, including CIFAR-10 classification and keyword spotting. The hardware resource requirements for non-linear functions in processing elements decreased by 92%, and 92.8%, respectively. Remarkably, the recognition accuracy for CIFAR-10 remains consistent, while there is a negligibly small degradation in accuracy for the keyword spotting task by 1.2%.
Dongzhu Li, Kenji Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ISCAS6
2023 A Fully Synthesized 13.7μJ/Prediction 88% Accuracy CIFAR-10 Single-Chip Data-Reusing Wired-Logic Processor Using Non-Linear Neural Network
abstract
An FPGA-based wired-logic CNN processor is presented that can process CIFAR-10 at 13.7μJ/prediction with an 88% accuracy, which is 2,036 times more energy-efficient than the prior state-of-the-art FPGA-based processor. Energy efficiency is greatly improved by implementing all processing elements and wirings in parallel on a single FPGA chip to eliminate the memory access. By utilizing both (1) a non-linear neural network which saves on neurons and synapses and (2) a shift register-based wired-logic architecture, hardware resource usage is reduced by three orders of magnitude.
Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC6
2023 A 1.2nJ/Classification Fully Synthesized All-Digital Asynchronous Wired-Logic Processor Using Quantized Non-Linear Function Blocks in 0.18μm CMOS
abstract
A 5.3 times smaller and 2.6 times more energy-efficient all-digital wired-logic processor which infers MNIST with 90.6% accuracy and 1.2nJ of energy consumption has been developed. To improve area efficiency of wired-logic architecture, nonlinear neural network (NNN), which is a neuron and synapse efficient network, and logical compression technology to implement it with area-saving and low-power digital circuits by logic synthesis are proposed, and asynchronous digital combinational circuit DNN hardware has been developed.
Rei Sumikawa, Kota Shiba, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC5
2023 An Occlusion-Resilient mmWave Imaging Radar-Based Object Recognition System Using Synthetic Training Data Generation Technique
abstract
An occlusion-resilient mmWave imaging radar-based object recognition system for advanced driver-assistance systems (ADAS) of construction machinery application is developed. As ADAS for construction sites, millimeter wave application is required in poor visibility environments such as nighttime, bad weather, and muddy conditions where object recognition by RGB cameras and LiDAR is difficult. A remaining technical challenge for ADAS is occlusion. Two techniques are proposed to improve the accuracy in occlusion scenes. First is a technique which generates simulated training data for occlusion environment to improve accuracy while reducing the cost for the training data preparation. The second is a parallel inference DNN architecture which enables object recognition with high accuracy in both normal and occlusion scenes by running two DNNs optimized respectively for normal and occlusion scenes in parallel. The object recognition accuracy of mAP50in occlusion scenes improves by 15 points compared to the conventional technique. The decrease in recognition accuracy in non-occlusion scenes is only 4 points.
Eitaro Kobayashi, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IECON4
2023 A 0.13mJ/Prediction CIFAR-100 Raster-Scan- Based Wired-Logic Processor Using Non-Linear Neural Network
abstract
A 0.13mJ/prediction with 68.6% accuracy single- chip wired-logic artificial intelligence (AI) processor is developed in a 16nm field-programmable gate array (FPGA). Compared with conventional von-Neumann architecture-based AI processors, the energy efficiency is greatly improved by eliminating the DRAM/BRAM access. A technical challenge of the conventional wired-logic processor is the large amount of hardware resources required. To implement a large convolutional neural network (CNN) into a single FPGA chip, two techniques are used: (1) a sparse neural network which is called non-linear neural network (NNN), and (2) a newly developed raster-scan-based wired-logic architecture. The amount of hardware resources required is reduced by a factor of 5.4. Compared with the state-of-the-art FPGA-based processor, 238 times better energy efficiency is achieved with the same accuracy on the CIFAR-I00 task. In addition, 7 times better energy efficiency is achieved compared with the state-of- the-art application-specific integrated circuit (ASIC) processor.
Dongzhu Li, Yao-Chung Hsu, Rei Sumikawa, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ISCAS6
2023 Polyomino: A 3D-SRAM-Centric Accelerator for Randomly Pruned Matrix Multiplication With Simple Reordering Algorithm and Efficient Compression Format in 180-nm CMOS
abstract
We have developed a sparse matrix reordering algorithm with a novel 3D-SRAM-centric Polyomino accelerator that enables efficient processing of the reordered matrix for parameter compression. By reordering randomly pruned, irregularly structured sparse matrices into regularly structured matrices, both the compression ratio of the data and the efficiency of the hardware processing increase. The reordering algorithm can be implemented simply by attributing it to the widely known k-sum problem. We also developed a compression format for storing the reordered matrices and show that the reordered regular structure can reduce the amount of required memory by 63% compared with the conventional method. The proposed Polyomino accelerator can efficiently process reordered matrices by using a 3D stacked SRAM, which is an external memory with random accessibility and low latency. The measurement results using a test chip fabricated in a 180-nm CMOS process demonstrate that the proposed accelerator can achieve high area-efficiency and high energy-efficiency and scales well with the pruning rate.
Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 A 5.2GHz RFID Chip Contactlessly Mountable on FPC at any 90-Degree Rotation and Face Orientation
abstract
This paper presents an RFID Chip contactlessly mountable on an FPC having an antenna pattern. Inductive coupling between the FPC and the chip realizes low-cost bonding-less implementation. It is also possible to place the chip on the FPC at any angle of 0/90/180/270 degrees and face-up or face-down. Simulation shows the antenna gain is almost the same irrespective of the chip placement angle and face orientation. The experimental results confirmed that the proposed RFID chip works at upto 20cm away from a reader whose output power is 15dBm, achieving the same figure-of-merit as a conventionally bonded module.
Reiji Miura, Saito Shibata, Masahiro Usui, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC6
2022 A 13.7μJ/prediction 88% Accuracy CIFAR-10 Single-Chip Wired-logic Processor in 16-nm FPGA using Non-Linear Neural Network
abstract
• In this study, we propose a 13.7mJ/prediction 88% accuracy CIFAR-10 single-chip wired-logic processor in 16-nm FPGA by utilizing a newly developed 98%-pruned ultra-sparse, binary-weight nonlinear neural network (NNN) and a shift-register based pipelined wired-logic architecture. Compared with the state-of-the-art FPGA-based processor, 2,036 times better energy efficiency is achieved.
Yao-Chung Hsu, Atsutake Kosuge, Rei Sumikawa, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
HCS6
2022 A 7-nm FinFET 1.2-TB/s/mm2 3D-Stacked SRAM with an Inductive Coupling Interface Using Over-SRAM Coils and Manchester-Encoded Synchronous Transceivers
abstract
A 0.7-pJ/bit, 8.5-Gbps/link inductive coupling inter-chip wireless communication interface for a 3D-stacked SRAM has been developed in a 7-nm FinFET process. A new physical placement method that allows coils to be placed over off-the-shelf SRAM macros with small magnetic field attenuation, together with the use of synchronous communication using Manchester encoding and a clocked comparator to enable the detection of small-swing signals, achieve a 26% reduction in SRAM die area compared to TSV-based stacking. Inter-chip communication at 0.7-pJ/bit, 8.5-Gbps/link was confirmed using test chips. A 4-hi 3D-stacked SRAM module using the proposed interface is estimated to achieve a 1.2-TB/s/mm2area efficiency, representing a two-orders-of-magnitude improvement over state-of-the-art 3D-stacked SRAM.
Kota Shiba, Mitsuji Okada, Atsutake Kosuge, Mototsugu Hamada, Tadahiro Kuroda
HCS5
2022 Proximity Wireless Communication Technologies: An Overview and Design Guidelines
abstract
This paper presents an overview of proximity wireless communication (PWC) technologies, their principles, design guidelines and practical applications. In particular, two different applications of PWC are reviewed. One is PWC between stacked chips. Both communication distance and coupler size are several tens of microns. Area and energy efficient design techniques are introduced. Another is PWC between module boards. Both communication distance and coupler size are several millimeters. Energy and area efficient practical designs are introduced for mobile and industrial machinery applications.
Atsutake Kosuge, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Sub-10-μm Coil Design for Multi-Hop Inductive Coupling Interface
abstract
Sub-10-μm on-chip coils are designed and prototyped for the multi-hop inductive coupling interface in a 40-nm CMOS. Multi-layer coils and a new receiver circuit are employed to compensate the decrease of the coupling coefficient due to the small coil size. The prototype emulates a 3D stacked module with 8 dies in a 7-nm CMOS and shows that a 0.1-pJ/bit and 41-Tb/s/mm2 inductive coupling interface is achievable.
Tatsuo Omori, Kota Shiba, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC4
2021 A 3D-Stacked SRAM Using Inductive Coupling Technology for AI Inference Accelerator in 40-nm CMOS
abstract
A 3D-stacked SRAM using an inductive coupling wireless inter-chip communication technology (TCI) is presented for an AI inference accelerator. The energy and area efficiency are improved thanks to the introduction of a proposed low-voltage NMOS push-pull transmitter and a 12:1 SerDes. A termination scheme to short unused open coils is proposed to eliminate the ringing in an inductive coupling bus. Test chips were fabricated in a 40-nm CMOS technology confirming 0.40-V operation of the proposed transmitter with successful stacked SRAM operation.
Kota Shiba, Tatsuo Omori, Mototsugu Hamada, Tadahiro Kuroda
ASP-DAC4
2021 A 96-MB 3D-Stacked SRAM Using Inductive Coupling With 0.4-V Transmitter, Termination Scheme and 12: 1 SerDes in 40-nm CMOS
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 35% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates for the degradation of the first transmitted bit, a coil termination scheme that aims to eliminate the ringing of 3D inductive coupling bus, and a 12:1 SerDes that minimizes power consumption and area overhead in inductive coupling channels. Low-power, large-capacity, 3-cycle latency 3D-stacked SRAM for a DNN accelerator is achieved with the combination of these techniques to serve as a replacement of 3D-stacked DRAM. The performance of the proposed 3D-SRAM is compared with HBM DRAM and achieves more than 50% lower energy consumption. The scaling scenario of the SRAM module is discussed in light of the scaling of the inductive coupling technology and logic process.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
IEEE Trans. Circuits Syst. I Regul. Pap.7
2020 A 3D-Stacked SRAM using Inductive Coupling with Low-Voltage Transmitter and 12: 1 SerDes
abstract
A 28.8-GB/s 96-MB 3D-stacked SRAM is presented. A total of eight SRAM dies, designed in a 40-nm CMOS process, are vertically stacked and connected using an inductive coupling wireless link with a low-voltage NMOS push-pull transmitter that reduces the power of the link by 45% with a 0.4-V power supply. The SRAM utilizes an inverted bit insertion scheme that compensates the degradation of the first signal, a coil termination scheme that aims to eliminate the noise of 3D inductive coupling bus, and a 12:1 SerDes. The data density of the SRAM should reach 12.3-MB/mm3, which extends beyond that of state-of-the-art stacked DRAMs.
Kota Shiba, Tatsuo Omori, Kodai Ueyoshi, Kota Ando, Kazutoshi Hirose, Shinya Takamaeda-Yamazaki, Masato Motomura, Mototsugu Hamada, Tadahiro Kuroda
ISCAS9
2019 Live Demonstration: A Non-Contact Transmission Line Connector for USB3.1 HD-Video Streaming
abstract
This demonstration shows a high data rate, non-contact connector using a transmission line coupler (TLC) composed of two differential transmission lines. The TLC is an impedance-matched and wide-bandwidth coupler, therefore it enables the high-speed baseband communication at 5Gbps/lane for USB3.1. Through a non-contact connector using the TLC, a 4K monitor is connected to a smartphone via the super speed wired signal lanes(SSTX/RX). The TLC connector can be used under harsh environment such as water, dust and misalignment. This demonstration provides visitors with opportunities to find utility of the TLC connector.
Tomoya Arakawa, Joshin Sone, Mitsuji Okada, Mototsugu Hamada, Tadahiro Kuroda
ISCAS5
2018 Design Methodology in Wireless Power Transfer System for 3-D Stacked Multiple Receivers
abstract
This paper proposes a design methodology in a wireless power transfer system for 3-D stacked multiple receivers. A 1:m selective power transfer system is realized by introducing a frequency/time division multiplexing system. The power transfer function is analytically formulated and an optimization methodology of multiple tuning capacitors values is proposed and compared with simulation results. By using the optimized values, power transfer efficiencies at 6.78MHz and 13.56MHz are simulated to be 80% and 84%, respectively. Crosstalk and load transient performance is also compared with 1:1 system and confirmed that its degradation is not significant.
Shusuke Yanagawa, Ryota Shimizu, Mototsugu Hamada, Toru Shimizu, Tadahiro Kuroda
ISCAS5
2017 An image sensor/processor 3D stacked module featuring ThruChip interfaces
abstract
A 1,000-fps motion vector (MV) estimation and classification engine for high-speed computational imaging in a 3D stacked imager/processor module is proposed, prototyped, assembled and tested. The module features 1) ThruChip interfaces for high fps image transfer, 2) orders of magnitude more area/power efficient MV estimation architecture compared to conventional ones, and 3) a cognitive classification scheme employed on MV patterns, enabling the classification of moving objects not possible in conventional proposals.
Masayuki Ikebe, Tetsuya Asai, Masafumi Mori, Toshiyuki Itou, Daisuke Uchida, Yasuhiro Take, Tadahiro Kuroda, Masato Motomura
ASP-DAC7
2016 Analytical thruchip inductive coupling channel design optimization
abstract
ThruChip interface (TCI) is an emerging 3-D integrated circuit stacking technology. TCI utilizes on-chip inductor to build vertical communication channel in near field distance and has been proved to stand comparison with through-siliconvia (TSV) in data rate, power, and reliability. Moreover, it is also cost-effective in manufacturing due to its wireless nature. In this paper, an analytical method is proposed to find near-optimal TCI inductive coupling channel solution. The experiment results show an average 16.8% transmitting current reduction and shrink design time from days to a few minutes.
Li-Chung Hsu, Junichiro Kadomoto, So Hasegawa, Atsutake Kosuge, Yasuhiro Take, Tadahiro Kuroda
ASP-DAC6
2016 Low-energy algorithm for self-controlled Wireless Sensor Nodes
abstract
In Internet of Things (IoT), the lifespan of Wireless Sensor Networks (WSN) has often become an issue. Sensor nodes are typically battery powered. However, high energy consumption by Radio Frequency (RF) module limits the lifespan of sensor nodes. In conventional WSN, the frequency of data transmission is normally fixed or adjusted according to requests from the gateway. In this paper, we present a WSN system for intelligent sensing. We propose a low-energy algorithm for sensor data transmission from sensor nodes for such system. In this algorithm, the sensor nodes are able to self-control their data transmission according to the trends of data. We adopt Adaptive Duty Cycle for adjustment of data transmission frequency and Compressive Sensing (CS) for sensor data compression. The simulation results show that Collective Transmission with CS-based data compression achieves 83.34% of RF energy reduction for the best-case transmission and 83.31% of RF energy reduction in the worst-case transmission, compared to the Continuous Transmission.
Ahmad Muzaffar bin Baharudin, Mika Saari, Pekka Sillberg, Petri Rantanen, Jari Soini, Tadahiro Kuroda
WINCOM6
2016 Efficient 3-D Bus Architectures for Inductive-Coupling ThruChip Interfaces
abstract
Wireless 3-D network-on-chips (NoCs) with inductive-coupling ThruChip interfaces provide a large degree of flexibility for customizing the number of arbitrary chips in a package after chips have been fabricated. To simplify the vertical communication interfaces, static time division multiple access (TDMA) is used for the vertical broadcast buses, while arbitrary or customized topologies can be used for the intrachip network. This paper proposes two techniques to break through the simple static TDMA-based vertical buses while maintaining a simple communication interface. The first technique is headfirst sliding (HS) routing to reduce the waiting time for acquiring the communication time-slot. HS routing selects the best vertical bus based on the current time, taking advantage of static TDMA. The second technique extends carrier sense multiple access with collision detection (CSMA/CD) for vertical broadcast buses. We introduce a packet collision detection technique for inductive-coupling buses and propose two retransmission strategies to reduce the waiting time for packet retransmissions caused by collisions. Network simulation results show that HS routing reduces the communication latency by 39.1% compared with the conventional static TDMA bus-based 3-D NoC that uses the shortest path routing. The proposed CSMA/CD bus also improves the latency by 52.5% and throughput by 34.1%. The full-system simulation results show that HS routing and the proposed CSMA/CD technique reduce the application execution time accordingly while maintaining the average flit transfer energy overhead modest.
Takahiro Kagami, Hiroki Matsutani, Michihiro Koibuchi, Yasuhiro Take, Tadahiro Kuroda, Hideharu Amano
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Design and analysis for ThruChip design for manufacturing (DFM)
abstract
A 1GB/s ThruChip interface (TCI) test chip for wafer thinning, power mesh, and dummy metal fill impacts are analyzed and evaluated with test chip measurement and field solver simulation. The measurement results show that TCI coil dimension can be sized down as wafer thinning by following D/Z=3 rule. However, the experiment shows 20% power reduction by enlarging TCI coil (D/Z=6). The power mesh lies between TCI coils can dramatically decrease the TCI magnetic pulse strength and hence cause TCI to fail. Dummy metal within TCI coils has no impact on TCI transmission
Li-Chung Hsu, Yasuhiro Take, Atsutake Kosuge, So Hasegawa, Junichiro Kadomoto, Tadahiro Kuroda
ASP-DAC6
2015 Circuit and package design for 44GB/s inductive-coupling DRAM/SoC interface
abstract
A 44GB/s inductive-coupling DRAM/SoC interface is developed by PoP integration. It utilizes the advantages of both TSV and LPDDR by using a ThruChip Interface (TCI) and an ultra-thin fan-out wafer level package (UT-FOWLP). The TCI allows data communication between the stacked chips while the UT-FOWLP thins the chips stacking distance and provides the chips with power. This proposed DRAM/SoC interface outperforms WIO2 with TSV in terms of area efficiency (4× better), immunity from simultaneous switching output (SSO) noise (32× better) and manufacturing cost (40% cheaper). In addition, it outperforms LPDDR4 in PoP in terms of power dissipation (5× lower) and timing control easiness. The inductive-coupling interface is newly designed to allow 12× improvement on its area efficiency. By using overlapping coils with quadrature phase division multiplexing (PDM), the coil density is increased by 4 times. The coil density is further increased by 3 times by shortening communication distance with the UT-FOWLP.
Akira Okada, Abdul Raziz Junaidi, Yasuhiro Take, Atsutake Kosuge, Tadahiro Kuroda
ASP-DAC5
2015 An 8 bit 0.3-0.8 V 0.2-40 MS/s 2-bit/Step SAR ADC With Successively Activated Threshold Configuring Comparators in 40 nm CMOS
abstract
A 0.3-0.8 V low-power 2-bit/step asynchronous successive approximation register analog-to-digital converter (ADC) is presented. A low-power 2-bit/step operation technique is proposed which uses dynamic threshold configuring comparator instead of multiple digital-to-analog converters (DACs). Power and area overhead is minimized by successively activated comparators. The comparator threshold is configured by simple Vcm biased current source, which keep the ADC free from power supply variations over 10%. Simple digital calibration is enabled by generating the reference internally. The prototype ADC fabricated in a 40 nm CMOS achieved a 44.3 dB signal-to-noise-plus-distortion ratio (SNDR) with 6.14 MS/s at a single supply voltage of 0.5 V. The ADC achieved a peak FoM of 4.8 fJ/conv-step at 0.4 V and operates down to 0.3 V.
Kentaro Yoshioka, Akira Shikata, Ryota Sekimoto, Tadahiro Kuroda, Hiroki Ishikuro
IEEE Trans. Very Large Scale Integr. Syst.4
2014 An 8b extremely area efficient threshold configuring SAR ADC with source voltage shifting technique
abstract
An extremely low power and area efficient threshold configuring ADC (TC-ADC) for time interleaved ADC is proposed. The threshold configuring comparator (TCC) performs a binary search. 5b conversion is carried out by TCC with source voltage shifting technique. Additional 2b resolution is achieved by the proposed threshold interpolation (TI) technique with only 15% power overhead. Prototype ADC in 40nm CMOS occupies a core area of only 0.0038mm2and when calibration circuit included, 0.0058 mm2. With a supply voltage of 0.7V, the ADC achieves 7.0 ENOB with 24MS/s. Peak FoM of 9.8fJ/conv. is obtained at 0.5V supply, which is over 15x improvement compared with conventional TC-ADC.
Kentaro Yoshioka, Akira Shikata, Ryota Sekimoto, Tadahiro Kuroda, Hiroki Ishikuro
ASP-DAC4
2014 Low-latency wireless 3D NoCs via randomized shortcut chips
abstract
In this paper, we demonstrate that we can reduce the communication latency significantly by inserting a fraction of randomness into a wireless 3D NoC (where CMOS wireless links are used for vertical inter-chip communication) when considering the physical constraints of the 3D design space. Towards this end, we consider two cases, namely 1) replacing existing horizontal 2D links in a wireless 3D NoC with randomized shortcut NoC links and 2) enabling full connectivity by adding a randomized NoC layer to a wireless 3D platform with partial or no horizontal connectivity. Consequently, the packet routing is optimized by exploiting both the existing and the newly added random NoC. At the same time, by adding randomly wired shortcut NoCs to a wireless 3D platform, a good balance can be established between the modularity of the design and the minimum randomness needed to achieve low latency, and experimental results show that by adding a random NoC chip to wireless 3D CMPs without built-in horizontal connectivity, the communication latency can be reduced by as much as 26.2% when compared to adding a 2D mesh NoC. Also, the application execution time and average flit transfer energy can be improved accordingly.
Hiroki Matsutani, Michihiro Koibuchi, Ikki Fujiwara, Takahiro Kagami, Yasuhiro Take, Tadahiro Kuroda, Paul Bogdan, Radu Marculescu, Hideharu Amano
DATE6
2014 3D NoC with Inductive-Coupling Links for Building-Block SiPs
abstract
A wireless 3D NoC architecture is described for building-block SiPs, in which the number of hardware components (or chips) in a package can be changed after chips have been fabricated. The architecture uses inductive-coupling links that can connect more than two examined dies without wire connections. Each chip has data transceivers for the uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a vertical unidirectional ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring, we use bubble flow control, which does not rely on the conventional VC-based deadlock avoidance mechanism. In addition, we propose a bidirectional communication scheme to form a bidirectional ring network by using the inductive-coupling transceivers that can dynamically change the communication modes, such as TX, RX, and Idle modes. This paper illustrates the inductive-coupling transceiver circuits, which can carry high data transfer rates of up to 8 Gbps per channel, for the wireless 3D NoC. It also illustrates an implementation of a wireless 3D NoC that has on-chip routers and transceivers implemented with a 65 nm process in order to show the feasibility of our proposal. The vertical bubble flow control and conventional VC-based approach on the uni- and bidirectional ring networks are compared with the vertical broadcast bus in terms of throughput, hardware amount, and application performance using a full system multiprocessor simulator. The results show that the proposed bidirectional communication scheme efficiently improves application performance without adding any inductive-coupling transceivers. In addition, the proposed vertical bubble flow network outperforms the conventional VC-based approach by 7.9-12.5 percent with a 33.5 percent smaller router area for building-block SiPs connecting up to eight chips.
Yasuhiro Take, Hiroki Matsutani, Daisuke Sasaki, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano
IEEE Trans. Computers5
2013 A 12.5Gb/s/link non-contact multi drop bus system with impedance-matched Transmission Line Couplers and Dicode partial-response channel transceivers
abstract
A reduced-reflection multi-drop bus system using Dicode (1-D) partial response signaling transceiver is presented for the first time in the world. Directional couplers on transmission lines arranged with equi-energy distributing and exact impedance matched conditions allow the bus to reach to 12.5Gbps/link speed, which is the world's fastest data link speed with multi-drop bus architecture. Dicode partial-response signaling method with a half-rate architecture was used where a precoder is placed in the transmitter to make the signal best fit for the channel to eliminate inter symbol interference (ISI).
Atsutake Kosuge, Wataru Mizuhara, Noriyuki Miura, Masao Taguchi, Hiroki Ishikuro, Tadahiro Kuroda
ASP-DAC6
2013 A case for wireless 3D NoCs for CMPs
abstract
Inductive-coupling is yet another 3D integration technique that can be used to stack more than three known-good-dies in a SiP without wire connections. We present a topology-agnostic 3D CMP architecture using inductive-coupling that offers great flexibility in customizing the number of processor chips, SRAM chips, and DRAM chips in a SiP after chips have been fabricated. In this paper, first, we propose a routing protocol that exchanges the network information between all chips in a given SiP to establish efficient deadlock-free routing paths. Second, we propose its optimization technique that analyzes the application traffic patterns and selects different spanning tree roots so as to minimize the average hop counts and improve the application performance.
Hiroki Matsutani, Paul Bogdan, Radu Marculescu, Yasuhiro Take, Daisuke Sasaki, Hao Zhang 0020, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano
ASP-DAC8
2013 A 0.35-0.8V 8b 0.5-35MS/s 2bit/step extremely-low power SAR ADC
abstract
An extremely low-voltage operating high speed and low power 2bit/step asynchronous SAR ADC is presented. Wide range dynamic threshold configuring comparator is proposed to enable power and area efficient 2bit/step operation. By configuring the comparator threshold by simple Vcmbiased current sources, the ADC holds immunity against 10% power supply variation. The prototype ADC fabricated in 40nm CMOS achieved 44.3 dB SNDR with 6.14 MS/s at a single supply voltage of 0.5 V. The ADC achieved a peak FoM of 5.9fJ/conv-step at 0.4V and operates down to 0.35V.
Kentaro Yoshioka, Akira Shikata, Ryota Sekimoto, Tadahiro Kuroda, Hiroki Ishikuro
ASP-DAC4
2013 Demonstration of a heterogeneous multi-core processor with 3-D inductive coupling links
abstract
Cube-1 is a heterogeneous multi-core processor which can achieve the required performance with the least energy consumption as possible. It can control the performance and energy with two levels: (1) the number of accelerators can be easily changed by increasing or decreasing the number of stacked chips after fabrication, as they are connected with inductive coupling links. (2) The supply voltage for PE array of the accelerator can be controlled by the host CPU so that the required performance can be obtained with a minimum supply voltage.
Yusuke Koizumi, Noriyuki Miura, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura
FPL5
2013 A scalable 3D heterogeneous multi-core processor with inductive-coupling thruchip interface
Noriyuki Miura, Yusuke Koizumi, Eiichi Sasaki, Yasuhiro Take, Hiroki Matsutani, Tadahiro Kuroda, Hideharu Amano, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura
Hot Chips Symposium6
2013 Adaptive window search using semantic texton forests for real-time object detection
abstract
We propose a new window search method to realize real-time object detection. Our method generates windows adaptively for objects' shapes and scales to detect various size objects. It also achieves real-time window search by using fast estimation of object's location based on existence probability of an object. Experiment results demonstrate that the proposed method reduces the number of windows drastically compared with exhaustive search. Furthermore, our method reduces the processing time while maintaining recall compared with the state-of-the art method when the numbers of searched windows are same in Pascal VOC 2007 dataset.
Yuki Ono, Abdul Raziz Junaidi, Tadahiro Kuroda
ICIP3
2013 A 1.26mW/Gbps 8 locking cycles versatile all-digital CDR with TDC combined DLL
abstract
This paper presents an all-digital CDR with TDC combined DLL which can be used for not only NRZ signaling but also pulse-based communication. The TDC combined DLL can realize a small area, low power and fast locking by sharing the delay line of the TDC with the DLL. The proposed CDR can recover the clock by evaluating a waveform of one cycle, and detect edge from a pulse-based signal. Locking time is within 8 clock cycles, and power efficiency is 1.26mW/GHz at 1Gbps and 0.7V power supply. The rms and peak-to-peak jitter at 2.3Gbps and 1.0V are 5.44ps and 37.4ps, respectively. Die area is 0.0297mm2.
Yuki Urano, Won-Joo Yun, Tadahiro Kuroda, Hiroki Ishikuro
ISCAS3
2012 Simultaneous data and power transmission using nested clover coils
abstract
This paper presents a simultaneous data and power transmission utilizing inductive-coupling interfaces for a non-contact memory card application. Nested clover coils are proposed to reduce interference from a power link. In order to maximize power transfer efficiency, the power transmitter tracks and predicts power consumption patterns of the memory card, and adjusts power transfer level. A test-chip prototype fabricated in a 65 nm CMOS process demonstrates 6 Gb/s data rate and 10% power transfer efficiency across a 0.1–2 kΩ load range.
Yasuhiro Take, Hayun Chung, Noriyuki Miura, Tadahiro Kuroda
ASP-DAC4
2012 CMA-Cube: A scalable reconfigurable accelerator with 3-D wireless inductive coupling interconnect
abstract
CMA-Cube is the second prototype of building block scalable reconfigurable accelerator using inductive coupling interconnect. It uses the wireless inductive coupling interconnect as a packet switching network which connects accelerators. As an accelerator core, CMA (Cool Mega Array), which consists of a large coarse-grained PE array with combinatorial circuits and tiny micro-controller, is applied. Evaluation results of Cube-1 Quad Core which consists of a host embedded CPU and three CMA-Cubes achieved 3.15 times performance acceleration as that without accelerators when JPEG decoder is executed.
Yusuke Koizumi, Eiichi Sasaki, Hideharu Amano, Hiroki Matsutani, Yasuhiro Take, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura
FPL6
2012 Dynamic power control with a heterogeneous multi-core system using a 3-D wireless inductive coupling interconnect
abstract
Cube-2 is a prototype of building block scalable reconfigurable accelerator using an inductive coupling interconnect. It is consisting of a ultra low leakage embedded processor Geyser and coarse-grained reconfigurable accelerators CMA (Cool Mega Array). A Geyser chip and multiple CMA chips are stacked, and a powerful network is formed by using the inductive coupling interconnect. The performance can be enhanced by increasing the number of CMA chips. JPEG decoder is implemented with a cooperation of Geyser and CMAs, and low power execution by controlling the power supply voltage of CMAs is demonstrated.
Yusuke Koizumi, Hideharu Amano, Hiroki Matsutani, Noriyuki Miura, Tadahiro Kuroda, Ryuichi Sakamoto, Mitaro Namiki, Kimiyoshi Usami, Masaaki Kondo, Hiroshi Nakamura
FPT5
2012 A 65fJ/b Inter-Chip Inductive-Coupling Data Transceivers Using Charge-Recycling Technique for Low-Power Inter-Chip Communication in 3-D System Integration
abstract
This paper presents a low-power inductive-coupling link in 90-nm CMOS. Our newly proposed transmitter circuit uses a charge-recycling technique for power-aware 3-D system integration. The cross-type daisy chain enables charge recycling and achieves power reduction without sacrificing communication performance such as a high timing margin, low bit error rate and high bandwidth. There are two design issues in the cross-type daisy chain: pulse amplitude reduction and another is inter-channel skew. To compensate for these issues, an inductor design and a replica circuit are proposed and investigated. Test chips were designed and fabricated in 90-nm CMOS to verify the validity of the proposed transmitter. Measurements revealed that the proposed cross-type daisy chain transmitter achieved an energy efficiency of 65 fJ/bit without degrading the timing margin, data rate, or bit error rate. In order to investigate the compatibility of the transmitter with technology scaling, a simulation of each technology node was performed. The simulation results indicate that the energy dissipation can be potentially reduced to less than 10 fJ/bit in 22 nm CMOS with proposed cross-type daisy chain.
Kiichi Niitsu, Shusuke Kawai, Noriyuki Miura, Hiroki Ishikuro, Tadahiro Kuroda
IEEE Trans. Very Large Scale Integr. Syst.5
2011 A vertical bubble flow network using inductive-coupling for 3-D CMPs
abstract
A wireless 3-D NoC architecture for CMPs, in which the number of processor and cache chips stacked in a package can be changed after the chip fabrication, is proposed by using the inductive coupling technology that can connect more than two known-good-dies without wire connections. Each chip has data transceivers for uplink and downlink in order to communicate with its neighboring chips in the package. These chips form a single vertical ring network so as to fully exploit the flexibility of the wireless approach that enables us to add, remove, and swap the chips in the ring. To avoid protocol and structural deadlocks in the ring network, we use the bubble flow control which is more flexible and efficient compared to the conventional VC-based deadlock avoidance. We implemented a real 3-D chip that has on-chip routers and inductive-coupling data transceivers using a 65nm process in order to show the feasibility of our proposal. The vertical bubble flow control is compared with the conventional VC-based approach and vertical bus in terms of the throughput, hardware amount, and application performance using a full system CMP simulator. The results show that the proposed vertical bubble flow network outperforms the VC-based approach by 7.9%-12.5% with a 33.5% smaller router area.
Hiroki Matsutani, Yasuhiro Take, Daisuke Sasaki, Masayuki Kimura, Yuki Ono, Yukinori Nishiyama, Michihiro Koibuchi, Tadahiro Kuroda, Hideharu Amano
NOCS8
2011 ThruChip interface (TCI) for 3D networks on chip
abstract
This paper presents a wireless interconnection for 3D Networks on Chip, namely ThruChip Interface (TCI). TCI employs near field inductive coupling that is suitable for high-density parallel channel arrangement with small cross talk. It is less expensive than TSV by over 20c/chip, since it is implemented by digital circuits in a standard CMOS. It bears comparison with TSV in terms of data rate (10Gb/s/ch), reliability (BER<;10-14), and energy dissipation (0.1pJ/b). ESD protection devices can be eliminated to lower delay, power, and area. It provides with an AC coupling link to make interface design easy under multiple/variable VDD's. The cost/performance will further be improved exponentially by thinning chip thickness. This talk will cover basics, applications, and future perspectives of the TCI.
Tadahiro Kuroda
VLSI-SoC1
2011 A 14-GHz AC-Coupled Clock Distribution Scheme With Phase Averaging Technique Using Single LC-VCO and Distributed Phase Interpolators
abstract
In this paper, we report the world's first ac-coupled clock distribution circuit for low-power and high-frequency clock distribution. By employing the proposed ac-coupled LC-based voltage-controlled oscillator (LC-VCO) and phase interpolators, the use of conventional current-mode-logic (CML) buffers with large power requirements can be prevented, and power consumption for clock distribution can be reduced. With the aim of verifying the effectiveness of the proposed circuit, test chips were designed and fabricated in 0.18-$\mu$m mixed-signal CMOS technology. The measured results indicated a 14.007 GHz clock distribution to four points whose pitches are 450$\mu$m, with 6.9 mW of power. The phase noise was measured to be$-$79.06 dBc/Hz at a 100 kHz offset,$-$101.66 dBc/Hz at a 1 MHz offset, and$-$107.25 dBc/Hz at a 10 MHz offset, with a clock frequency of 12.96 GHz. Furthermore, a phase averaging technique for reducing phase deviation was proposed and theoretically investigated.
Kiichi Niitsu, Shinmo Kang, Hiroki Ishikuro, Tadahiro Kuroda
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Analysis and Techniques for Mitigating Interference From Power/Signal Lines and to SRAM Circuits in CMOS Inductive-Coupling Link for Low-Power 3-D System Integration
abstract
This paper discusses analysis and techniques for mitigating interference of an inductive-coupling inter-chip link. Electromagnetic interference from power/signal lines and to SRAM circuits was simulated and measured. In order to verify the interference, test chips were designed and fabricated using 65-nm CMOS technology. The measurement results revealed that: 1) interference from power lines depends on the shape of the power lines; 2) interference from signal lines can be canceled by increasing transmitter power by only 9%; and 3) interference with SRAM circuits is less important than other issues under ordinary conditions. Based on the measurement results, interference mitigation techniques are proposed and investigated.
Kiichi Niitsu, Yasufumi Sugimori, Yoshinori Kohama, Kenichi Osada, Naohiko Irie, Hiroki Ishikuro, Tadahiro Kuroda
IEEE Trans. Very Large Scale Integr. Syst.7
2010 A versatile recognition processor for sensor network applications
abstract
A versatile recognition processor is presented that comprises 2.1M transistors using a 90 nm CMOS technology. It performs detection and recognition from image/video, sound and acceleration signals with energy consumption of sub-mJ/frame. The versatility and the power efficiency are attributed to optimal architecture design employing Haar-like Feature and Cascaded Classifier.
Risako Takashima, Yuya Hanai, Yuichi Hori, Tadahiro Kuroda
ASP-DAC4
2010 Modeling and Experimental Verification of Misalignment Tolerance in Inductive-Coupling Inter-Chip Link for Low-Power 3-D System Integration
abstract
Modeling and experimental verification of misalignment tolerance in inductive-coupling inter-chip links for 3-D system integration is introduced for the first time. Misalignment between stacked chips reduces coupling coefficiency of on-chip inductors and increases transmitter power. We proposed a modeling which estimates the increase in transmitter power by considering misalignment as an additional communication distance. Proposed model was verified by electromagnetic simulations and by measurements using testchips fabricated in 65-nm CMOS technology. The results calculated by the proposed modeling match well with measurement results. Measurement results show that misalignment tolerance of inductive-coupling link is well high and can be ignored in common conditions.
Kiichi Niitsu, Yoshinori Kohama, Yasufumi Sugimori, Kazutaka Kasuga, Kenichi Osada, Naohiko Irie, Hiroki Ishikuro, Tadahiro Kuroda
IEEE Trans. Very Large Scale Integr. Syst.8
2009 A wireless real-time on-chip bus trace system
abstract
A 480Mb/s wireless real-time bus trace system with a pulse-based inductive coupling channel array was developed using a 0.25μm CMOS digital process. The size and pitch of the inductor array are determined by numerical calculation to optimize the tradeoff between the channel coupling, crosstalk, and alignment tolerance. A low-power quasi-synchronous system is proposed to obtain an enough timing margin for RX pulse detection under the presence of the clock skew.
Shusuke Kawai, Takayuki Ikari, Yutaka Takikawa, Hiroki Ishikuro, Tadahiro Kuroda
ASP-DAC5
2009 A 1 GHz CMOS comparator with dynamic offset control technique
abstract
A dynamic offset control technique that employs charge compensation by timing control is proposed for comparator design in scaled CMOS technology. The analysis has been verified by fabricating a 65 nm CMOS 1.2 V 1 GHz comparator that occupies 25 × 65 μm2 and consumes 380 μW. Circuits for offset control occupies 21% of the areas and 12% of the power consumption of the whole comparator chip.
Xiaolei Zhu 0002, Sanroku Tsukamoto, Tadahiro Kuroda
ASP-DAC3
2009 MuCCRA-Cube: A 3D dynamically reconfigurable processor with inductive-coupling link
abstract
MuCCRA-Cube is a scalable three dimensional dynamically reconfigurable processor. By stacking multiple dies connected with inductive-coupling links, the number of PE array can be increased so that the required performance is achieved. A prototype chip with 90nm CMOS process consisting of four dies each of which has a 4 × 4 PE array was implemented. The vertical link achieved 7.2Gb/s/chip, and the average execution time is reduced to 31% compared to that using a single chip.
Shotaro Saito, Yoshinori Kohama, Yasufumi Sugimori, Yohei Hasegawa, Hiroki Matsutani, Toru Sano, Kazutaka Kasuga, Yoichi Yoshida, Kiichi Niitsu, Noriyuki Miura, Tadahiro Kuroda, Hideharu Amano
FPL11
2009 Face detection through compact classifier using Adaptive Look-Up-Table
abstract
Face detection has been well studied in terms of accuracy and speed. However, required memory size reduction is still poorly studied, which is becoming a critical issue as platforms for face detection go tiny. In this paper, we propose a novel compact weak classifier using Adaptive Look-Up-Table (ALUT) for face detection on resource-constrained devices such as wearable sensor nodes. ALUT gives good approximation of log-likelihood with fewer data, thus enabling the drastic reduction of classifier data size, keeping high accuracy and low computation cost. To generate an optimal ALUT, a new cost function called Weighted Sum of Absolute Difference (WSAD) is also proposed for further improvement. In our experiment, the classifier data size is reduced by 43% and the computation cost is reduced by 15% with same accuracy, compared to a conventional fixed LUT classifier.
Yuya Hanai, Tadahiro Kuroda
ICIP2
2008 Speaker Siglet Detection for Business Microscope
abstract
"Business Microscope" is our sensornet application in the age of knowledge, which visualizes knowledge workers' interactions by sensing their face-to-face communications. Due to the limitation of energy consumption of sensor nodes and privacy concerns, very short (0.1s) intermittently sensed (10s interval) noise-like signals called siglet is used to for detection task. To detect the speaker from the limited input, "self" vs "others" classification problem is introduced. For this new classification problem, new classifier called AdaBoost LVQ is studied to explore the application of AdaBoost to reduce the error rate of the conventional classifier with strictly limited inputs. As a result, AdaBoost LVQ achieved highest recognition accuracy of 96.45% with 19.86% error rate improvement relative to best conventional classifier.
Jun Nishimura, Nobuo Sato, Tadahiro Kuroda
ICMLA3
2008 Speech "Siglet" Detection for Business Microscope (concise contribution)
abstract
"Business Microscope" is a tool which provides knowledge workers with a bird-eye view of their daily communication. To meet the problem of the energy consumption of sensor nodes and privacy concerns for wearers and non-wearers, "siglet" sensing is proposed. Siglet sensing is a way to capture very short and noise-like signals by sensors operating on a low duty ratio. To extract the useful information on workers' communication, speech siglet detection is studied. The LBG trained speech and workplace nonspeech models with Mel frequency cepstrum coefficients (MFCCs) as feature vectors are utilized. A hierarchical pruning technique is studied to reduce the calculation cost of the matching process to nearly 25% and refine the classification accuracy. Our approach achieved average speech and nonspeech classification accuracy of 99.96% on 0. Is long test siglets.
Jun Nishimura, Nobuo Sato, Tadahiro Kuroda
PerCom3
2007 A 1Tb/s 3W Inductive-Coupling Transceiver Chip
abstract
A 1Tb/s 3W inter-chip transceiver transmits clock and data by inductive coupling at a clock rate of 1GHz and data rate of 1Gb/s per channel. 1024 data transceivers are arranged with a pitch of 30 mum in a layout area of 1mm2. The total layout area including 16 clock transceivers is 2mm2in 0.18 mum CMOS and the chip thickness is reduced to 10 mum. Simple yet accurate model of inductive coupling is utilized for transceiver design. Bi-phase modulation (BPM) is employed for the data link to improve noise immunity, reducing power in the transceiver. 4-phase time division multiplexing (TDM) reduces crosstalk and channel pitch. The BER is lower than 10-13with 150ps timing margin.
Noriyuki Miura, Tadahiro Kuroda
ASP-DAC2
2004 Practical methodology of post-layout gate sizing for 15% more power saving
Noriyuki Miura, Naoki Kato, Tadahiro Kuroda
ASP-DAC3
2002 Optimization and control of VDD and VTH for low-power, high-speed CMOS design
abstract
It is essential to control VDD and VTH for low-power, high-speed CMOS design. In this paper, it is shown that these two parameters can be controlled by designers as objectives of design optimization to find better trade-offs between power and speed. Quantitative analysis of trade-offs between power and speed is presented. Some of the popular circuit techniques and design examples to control VDD and VTH are introduced. A simple theory to compute optimum multiple VDD's and VTH's is described. Scaling scenarios of variable and/or multiple VDD's and VTH's is discussed to show future technology directions.
Tadahiro Kuroda
ICCAD1
2002 Low-Power, High-Speed CMOS VLSI Design
abstract
Ubiquitous computing is a next generation information technology where computers and communications will be scaled further, merged together, and materialized in consumer applications. Computers will be invisible behind broadband networks as servers, while terminals will come closer to us as wearable/implantable devices, more friendly devices with sophisticated human-computer interactions. IC chips will be implanted everywhere so that things can think and talk for distributed information processing. Key technologies here are low power, low cost, and good interfaces, especially for wireless data communications. Low-power, high-speed CMOS circuit techniques are presented in this paper, including low-voltage design with variable/multiple V/sub DD//V/sub TH/ control, embedded memory technology for reducing capacitance, and low-switching activity design.
Tadahiro Kuroda
ICCD1
1999 Variable supply-voltage scheme with 95%-efficiency DC-DC converter for MPEG-4 codec
abstract
A variable supply-voltage (VS) scheme with a high powerconversion-efficiency DC-DC converter is presented.A new pulse width modulation (PWM) circuit for the DC-DC converter is proposed to reduce both of power consumption and chip area.The power conversion efficiency reaches up to 95%, and the area is less than half of the conventional design.The VS scheme contains critical path replica circuits of an MPEG4 codec LSI, and its output voltage is controlled by monitoring delay time of the replica circuits.Consequently the VS scheme can automatically generate minimal internal supply voltage that meets the demand from the operation frequency of an MPEG4 codec LSI.The advantages of this circuit are successfully demonstrated through fabrication of a test chip using a 0.3pm CMOS technology.
Fuyuki Ichiba, Kojiro Suzuki, Shinji Mita, Tadahiro Kuroda, Tohru Furuyama
ISLPED4
1998 Design Methodology of Ultra Low-Power MPEG4 Codec Core Exploiting Voltage Scaling Techniques
abstract
This paper describes a fully automated low-power design methodology in which three different voltage-scaling techniques are combined together. Supply voltage is scaled globally, selectively, and adaptively while keeping the performance. This methodology enabled us to design an MPEG4 codec core with 58% less power than the original in three week turn-around-time.
Kimiyoshi Usami, Mutsunori Igarashi, Takashi Ishikawa, Masahiro Kanazawa, Masafumi Takahashi, Mototsugu Hamada, Hideho Arakida, Toshihiro Terazawa, Tadahiro Kuroda
DAC9
1996 Substrate noise influence on circuit performance in variable threshold-voltage scheme
abstract
This paper investigates substrate noise influence on circuit performance in a variable threshold-voltage scheme (VT scheme) where threshold voltage is dynamically varied by substrate-bias control to reduce active power dissipation. It is experimentally examined that substrate-bias can be controlled stably with very few substrate-contacts. Measured tracking jitter of a delay-locked loop implemented by interconnections in an 8 mm-square gate array does not degrade even when substrate-contacts are removed except for one at every strip of p-sub and n-well: A 2 mm-square discrete cosine transform core processor with no substrate-contact except in its periphery operates at supply voltages from 1.3 V to above 3 V even though it employs small-swing differential dynamic pass-transistor logic. No performance degradation nor latchup is observed in these chips even when 100 k/spl Omega/ resistance is added to the substrate. These experimental results demonstrate noise immunity of the VT scheme, and indicate the possibility that the VT scheme can be applied to existing macro design easily.
Tadahiro Kuroda, Tetsuya Fujita, Shinji Mita, Toshiaki Mori, Kenji Matsuo, Masakazu Kakumu, Takayasu Sakurai
ISLPED1