EDBT 2026 Demo / reviewers in the wild / expert
Takahiro Hanyu
dblp:67/5488
· DBLP profile ↗
43ranked-venue papers
8as first author
8since 2021 · last 2025
0000-0002-4397-8290ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI HardwareabstractIn this study, we propose an implementation methodology of real-time few-shot learning on tiny FPGA SoCs such as the PYNQ-Z1 board with arbitrary fixed-point bit-widths. Tensil-based conventional design environments limited hardware implementations to fixed-point bit-widths of 16 or 32 bits. To address this, we adopt the FINN framework, enabling implementations with arbitrary bit-widths. Several customizations and minor adjustments are made, including: 1.Optimization of Transpose nodes to resolve data format mismatches, 2.Addition of handling for converting the final "reduce mean" operation to Global Average Pooling (GAP). These adjustments allow us to reduce the bit-width while maintaining the same accuracy as the conventional realization, and achieve approximately twice the throughput in evaluations using CIFAR-10 dataset. R. Kanda, Hugo Le Blevec, Naoya Onizawa, Mathieu Léonardon, Vincent Gripon, Takahiro Hanyu |
ISCAS | 6 |
| 2025 | Design of a Low-Energy MTJ-Based Nonvolatile Register Based on a Differential Information Storing SchemeabstractIn intermittent computing enabling continuous information processing under unstable energy supply, it is important to minimize the energy consumption required to maintain intermediate information during frequent power supply interruptions. In this paper, we propose a new configuration of a magnetic-tunnel-junction (MTJ)-based nonvolatile register that can save intermediate information during intermittent operation in a short time with low energy consumption. The proposed nonvolatile register adopts a differential information storing scheme that expresses one-bit data by assigning ‘1’ when the states of two adjacent MTJ devices are the same and ‘0’ when they are different. This reduces the circuit area and energy consumption by reducing the number of MTJ devices in a register, as well as achieving quick backup and restoration by accessing MTJ devices in parallel. The evaluation using a 55 nm CMOS/MTJ-hybrid process technology shows that the proposed nonvolatile register can reduce the circuit area and energy consumption by 34% and 49%, respectively, compared to those of a conventional one. Tomoo Yoshida, Masanori Natsui, Takahiro Hanyu |
ISCAS | 3 |
| 2024 | Design of a High-Speed and Low-Power Threshold Adjustment Unit for Battery-Free Edge DevicesabstractIt is an attractive feature to make the capacity of power storage as small as possible in battery-free devices, while it may cause task failure frequently because of its unsatisfactory amount of power storage. This paper describes the design of a power control circuit using a threshold adjustment unit (TAU), that can adjust the energy harvested by a battery-less device for different tasks. The scheme sets a properly sized capacitor as power storage, which can adjust its power capacity prepared for tasks with different power consumptions by changing the charging and discharge threshold, thus avoiding task failure caused by insufficient power supply. Meanwhile, the high-speed and low-power design of this scheme can reduce the overhead caused by judging the charging/discharging voltage and the threshold voltage. By applying this unit to a battery-free configuration that assumes RF power feeding, charge and discharge path control at various thresholds in intermittent computing can be completed automatically and with low overhead. With the simulation of HSPICE based on 45nm CMOS, the average power consumption and response delay of the proposed trigger circuit with double-rail discrete-time comparator are only 45.9μW and 39.0μs, respectively. Fangcen Zhong, Masanori Natsui, Takahiro Hanyu |
IJCNN | 3 |
| 2023 | High-Performance/Low-Area Power-Gating Switch Linear Array for Energy-Efficient LSIs with an Optimum Switch-Timing ControlabstractPower gating (PG) is an important technique of energy-efficient LSIs to realize additional applications for Internet-of-Things devices in the future, but it also has issues including circuit performance degradation and increased area overhead. In this paper, we designed a new PG switch array with its control unit based on an optimum switch-timing control scheme to obtain a high-performance/low-area implementation. According to the simulation results based on 45nm CMOS under HSPICE, the maximum inrush current is reduced by up to 95.6% and the minimum wake-up-time is only 0.22ns. Fangcen Zhong, Masanori Natsui, Takahiro Hanyu |
ISCAS | 3 |
| 2023 | Fast-Converging Simulated Annealing for Ising Models Based on Integral Stochastic ComputingabstractProbabilistic bits (p-bits) have recently been presented as a spin (basic computing element) for the simulated annealing (SA) of Ising models. In this brief, we introduce fast-converging SA based on p-bits designed using integral stochastic computing. The stochastic implementation approximates a p-bit function, which can search for a solution to a combinatorial optimization problem at lower energy than conventional p-bits. Searching around the global minimum energy can increase the probability of finding a solution. The proposed stochastic computing-based SA method is compared with conventional SA and quantum annealing (QA) with a D-Wave Two quantum annealer on the traveling salesman, maximum cut (MAX-CUT), and graph isomorphism (GI) problems. The proposed method achieves a convergence speed a few orders of magnitude faster while dealing with an order of magnitude larger number of spins than the other methods. Naoya Onizawa, Kota Katsuki, Duckgyu Shin, Warren J. Gross, Takahiro Hanyu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | High Convergence Rates of CMOS Invertible Logic Circuits Based on Many-Body HamiltoniansabstractThis paper introduces CMOS invertible-logic (CIL) circuits based on many-body Hamiltonians. CIL can realize probabilistic forward and backward operations of a function by annealing a corresponding Hamiltonian using stochastic computing. We have created a Hamiltonian that includes three-body interaction of spins (probabilistic nodes). It provides some degrees of freedom to design a simpler landscape of Hamiltonian (energy) than that of the conventional two-body Hamiltonian. The simpler landscape makes it easier to reach the global minimum energy. The proposed three-body CIL circuits are designed and evaluated with the conventional two- body CIL circuits, resulting in few-times higher convergence rates with negligible area overhead on FPGA. Naoya Onizawa, Takahiro Hanyu |
ISCAS | 2 |
| 2021 | A Design Framework for Invertible LogicabstractInvertible logic using a probabilistic magnetoresistive device model has been recently presented that can compute functions in bidirectional ways and solve several problems quickly, such as factorization and combinational optimization. In this article, we present a design framework for invertible logic circuits. Our approach makes use of linear programming to create a Hamiltonian library with the minimum number of nodes for small invertible-logic functions. In addition, as the device model is approximated based on stochastic computing in synthesizable SystemVerilog, a faster simulation using the compiled SystemC binary is realized than a conventional SPICE-level simulation and is verified using field-programmable gate array (FPGA) as prototyping. Using our design framework, several invertible-logic circuits are designed and emulated (verified) in SystemC, exhibiting five order-of-magnitude faster simulation than conventional work. Naoya Onizawa, Kaito Nishino, Sean C. Smithson, Brett H. Meyer, Warren J. Gross, Hitoshi Yamagata, Hiroyuki Fujita, Takahiro Hanyu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2021 | Multi-Context TCAM-Based Selective Computing: Design Space Exploration for a Low-Power NNabstractIn this paper, we propose a low-power memory-based computing architecture, called selective computing architecture (SCA). It consists of multipliers and an LUT (Look-Up Table)-based component, that is multi-context ternary content-addressable memory (MC-TCAM). Either of them is selected by input-data conditions in neural-networks (NNs). Compared with quantized NNs, a higher accurate multiplication can be performed with low-power consumption in the proposed architecture. If input data stored in the MC-TCAM appears, the corresponding multiplication results for multiple weights are obtained. The MC-TCAM stores only shorter length of input data, resulting in achieving a low-power computing. The performance of the SCA is determined by three physical parameters concerning the configuration of MC-TCAM. The power dissipation of the target NN can be minimized by exploring these parameters in the design space. The hardware based on the proposed architecture is evaluated using TSMC 65 nm CMOS technology and MTJ model. In the case of speech command recognition, the power consumption at the multiplication of the first convolutional layer in a convolutional NN is reduced by 67% compared to the solution relying only on multipliers. Ren Arakawa, Naoya Onizawa, Jean-Philippe Diguet, Takahiro Hanyu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Memristive Computational Memory Using Memristor Overwrite Logic (MOL)abstractIn this article, we present a novel logic design style, namely, memristor overwrite logic (MOL), associated with an original MOL-based computational memory. MOL relies on a fully digital representation of memristor and can operate with different memristive device technologies. Its integration in memristive crossbar arrays and computational memories allows the execution of bit and vector-level primitive logic operations in two computational steps at most. Promising features and performances are demonstrated through the implementation of N -bit full addition using the proposed MOL-based computational memory. Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah, Naoya Onizawa, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2020 | High-Throughput/Low-Energy MTJ-Based True Random Number Generator Using a Multi-Voltage/Current ConverterabstractThis article introduces high-throughput/low-energy true random number generators (TRNGs) based on CMOS and three-terminal magnetic tunnel junction (MTJ) devices. MTJs are fast and probabilistic switching devices, which can be used as random number sources for TRNGs. However, as the switching probability is quite sensitive to the write current given to MTJs, precise closed-loop control is necessary. Thus, a high-complexity current control circuit is required, such as high precision digital-to-analog converters (DACs), occupying large area and causing large energy dissipation. In order to address the issue, we propose a multi-voltage/current (V/I) converter capable of multilevel coarse current switching and fine adjusting within each level. The fine adjusting can be done by DACs with fewer bits, resulting in much smaller size and energy dissipation than a conventional single-V/I converter. In addition, a multiple-writing scheme for three-terminal MTJs is proposed for increasing the throughput while maintaining the write power. The proposed TRNGs are designed using TSMC 65-nm CMOS and a threeterminal MTJ model that achieves a throughput of 333 Mb/s, an energy dissipation of 0.66 pJ/bit and an area of 2040 μm2. This result exhibits a 5× throughput, a 93% energy reduction and an 86% area reduction in comparison with a conventional CMOS/MTJ-based TRNG. Naoya Onizawa, Shogo Mukaida, Akira Tamakoshi, Hitoshi Yamagata, Hiroyuki Fujita, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Design of an MTJ-Based Nonvolatile LUT Circuit with a Data-Update Minimized Shift Operation for an Ultra-Low-Power FPGA: (Abstract Only)abstractNonvolatile FPGAs (NV-FPGAs) have a potential advantage to eliminate wasted standby power which is increasingly serious in recent standard SRAM-based FPGAs. However, functionality of the conventional NV-FPGAs are not sufficient compared to that of standard SRAM-based FPGAs. For example, an effective circuit structure to perform shift-register (SR) function has not been proposed yet. In this paper, a magnetic tunnel junction (MTJ) based nonvolatile lookup table (NV-LUT) circuit that can perform SR function with low power consumption is proposed. The MTJ device is the best candidate in terms of virtually unlimited endurance, CMOS compatibility, and 3D stacking capability. On the other hand, large power consumption to perform SR function a serious design issue for the MTJ-based NV-LUT circuit. Since the write current for the MTJ device is large and all the data must be updated after the SR operation using CMOS-oriented method, large power consumption is indispensable. To overcome this issue, the address for read/write access is incremented at each cycle instead of direct data shifting in the proposed LUT circuit. In this way, the number of data update per 1-bit shift is minimized to one, which results in great power saving. Moreover, since the selector is shared both read (logic) and write operation, its hardware cost is small. In fact, 99% of power reduction and 52% of transistor counts reduction compared to those of SRAM-based LUT circuit are performed. The authors would like to acknowledge ImPACT of CSTI, CIES consortium program, JST-OPERA, and JSPS KAKENHI Grant No. 17H06093. Daisuke Suzuki, Takahiro Hanyu |
FPGA | 2 |
| 2018 | High-Precision Stochastic State-Space Digital Filters Based on Minimum Roundoff Noise StructureabstractDigital filters based on stochastic computation have recently gained considerable attention because the stochastic computation attains significant reduction of hardware complexity of digital filters compared with the classical deterministic binary computation-based filters. For stochastic IIR filters, Liu and Parhi proposed an elegant method that achieves high precision by means of the normalized state-space lattice structure. This paper extends the result and presents stochastic state-space IIR filters with higher precision than the conventional method. Instead of using the normalized lattice structure, our method makes use of the minimum roundoff noise structure in realizing stochastic state-space filters, leading to further improvement of arithmetic precision of IIR filtering compared with the conventional method. Experimental results show that our method gives higher signal-to-error ratio at the filter outputs than the conventional method. Shunsuke Koshita, Naoya Onizawa, Masahide Abe, Takahiro Hanyu, Masayuki Kawamata |
ISCAS | 4 |
| 2018 | Networked Power-Gated MRAMs for Memory-Based ComputingabstractEmerging nonvolatile memory technologies open new perspectives for original computing architectures. In this paper, we propose a new type of flexible and energy-efficient architecture that relies on power-gated distributed magnetoresistive random access memory (MRAM). The proposed architecture uses a network-on-chip (NoC) to interconnect MRAM-based clusters, processing elements, and managers. The NoC distributes application-specific commands to MRAM devices by means of packets. Configurable network interfaces allow to transform MRAM devices into smart units able to respond to incoming commands. In this context, three types of MRAM designs are proposed with different power-gating policies and granularities. A relevant database search engine case study is considered to illustrate the benefits of this proposed architecture. It is implemented with a sparse-neural-network approach and simulated in SystemC with different scenarios including hundreds of database queries. Hardware designs and accurate power estimations have been conducted. The obtained results demonstrate important power reduction with database hit rates of about 94%. Targeting 65-nm technology, energy savings reach 87% when compared with an static random access memory-based implementation. Moreover, a new asymmetric read/write MRAM type provides from 39% to 50% energy reduction with respect to the other fixed-granularity models. This results in a low-power, highly scalable, and configurable implementation of memory-based computing. Jean-Philippe Diguet, Naoya Onizawa, Mostafa Rizk, Martha Johanna Sepúlveda, Amer Baghdadi, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2017 | Three-terminal MTJ-based nonvolatile logic circuits with self-terminated writing mechanism for ultra-low-power VLSI processorabstractMagnetic-Tunnel Junction (MTJ)-based non-volatile logic circuits have some possibility to solve the power-dissipation problem seriously focusing on the present CMOS-only-based VLSI processors. Three terminal MTJ devices are the promising candidate as nonvolatile storage device to realize such a nonvolatile logic circuit. However, its writing energy is still serious in comparison with conventional CMOS-only-based logic circuits. In this paper, a new MTJ-based nonvolatile logic circuit with self-terminated mechanism is proposed and its energy efficiency is evaluated in comparison with the corresponding previous work. In addition, some recent research topics related to MTJ-based nonvolatile logic-circuit design and its application, such as a computer-aided-design (CAD) tool considering a stochastic MTJ-switching behavior and the application to a resilient “die-hard” VLSI processor against sudden power-supply outage, are also demonstrated. Takahiro Hanyu, Daisuke Suzuki, Naoya Onizawa, Masanori Natsui |
DATE | 1 |
| 2017 | VLSI Implementation of Deep Neural Network Using Integral Stochastic ComputingabstractThe hardware implementation of deep neural networks (DNNs) has recently received tremendous attention: many applications in fact require high-speed operations that suit a hardware implementation. However, numerous elements and complex interconnections are usually required, leading to a large area occupation and copious power consumption. Stochastic computing (SC) has shown promising results for low-power area-efficient hardware implementations, even though existing stochastic algorithms require long streams that cause long latencies. In this paper, we propose an integer form of stochastic computation and introduce some elementary circuits. We then propose an efficient implementation of a DNN based on integral SC. The proposed architecture has been implemented on a Virtex7 field-programmable gate array, resulting in 45% and 62% average reductions in area and latency compared with the best reported architecture in the literature. We also synthesize the circuits in a 65-nm CMOS technology, and we show that the proposed integral stochastic architecture results in up to 21% reduction in energy consumption compared with the binary radix implementation at the same misclassification rate. Due to fault-tolerant nature of stochastic architectures, we also consider a quasi-synchronous implementation that yields 33% reduction in energy consumption with respect to the binary radix implementation without any compromise on performance. Arash Ardakani, François Leduc-Primeau, Naoya Onizawa, Takahiro Hanyu, Warren J. Gross |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Area/Energy-Efficient Gammatone Filters Based on Stochastic ComputationabstractThis paper introduces area/energy-efficient gammatone filters based on stochastic computation. The gammatone filter well expresses the performance of human auditory peripheral mechanism and has a potential of improving advanced speech communications systems, especially hearing assisting devices and noise robust speech-recognition systems. Using stochastic computation, a power-and-area hungry multiplier used in a digital filter is replaced by a simple logic gate, leading to area-efficient hardware. However, a straightforward implementation of the stochastic gammatone filter suffers from significantly low accuracy in computation, which results in a low dynamic range (a ratio of the maximum to minimum magnitude) due to a small value of a filter gain. To improve the computation accuracy, gain-balancing techniques are presented that represent the original gain as the product of multiple larger gains introduced at the second-order sections. In addition, dynamic scaling techniques are proposed that scales up small values only on stochastic domain in order to reduce the number of stochastic bits required while maintaining the computation accuracy. For performance comparisons, the proposed stochastic gammatone filters are designed and evaluated on taiwan semiconductor manufacturing company (TSMC) 65-nm CMOS technology. As a result, the proposed filter achieves an area reduction of 90.7% and an energy reduction of 91.8% in comparison with a fixed-point gammatone filter at the same sampling frequency and a comparable dynamic range. Naoya Onizawa, Shunsuke Koshita, Shuichi Sakamoto, Masahide Abe, Masayuki Kawamata, Takahiro Hanyu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | A low-power MTJ-based nonvolatile FPGA using self-terminated logic-in-memory structureabstractA nonvolatile field-programmable gate array (NVFPGA), where both magnetic tunnel junction (MTJ) devices and greedy power-saving techniques are utilized, is proposed. Because the circuit components are shared among several MTJ devices by the use of logic-in-memory (LIM) structure, the number of leakage current paths is reduced, which results in leakage power reduction during power-on. Moreover, the use of the self-termination scheme, which automatically turns off the write current immediately after the desired data is written, makes it possible to minimize power consumption during the backup operation. In fact, the proposed NVFPGA exhibits a 90 % power reduction in comparison with that of a conventional SRAM-based FPGA under typical benchmark-circuit implementations. Daisuke Suzuki, Takahiro Hanyu |
FPL | 2 |
| 2016 | Gammatone filter based on stochastic computationabstractThis paper introduces a design of a gammatone filter based on stochastic computation for area-efficient hardware. The gammatone filter well expresses the performance of human auditory peripheral mechanism and has a potential of improving advanced speech communications systems, especially hearing assisting devices and noise robust speech recognition systems. Using stochastic computation, a power-and-area hungry multiplier used in a digital filter is replaced by a simple logic gate, leading to area-efficient hardware. However, a straightforward implementation of the stochastic gammatone filter suffers from significantly low accuracy in computation, which results in a low dynamic range (a ratio of the maximum to minimum magnitude) due to a small value of a filter gain. To improve the computational accuracy, gain-balancing techniques are presented that represent the original gain as the product of multiple larger gains introduced at the second-order sections. As a result, the proposed techniques maintain the original gain of the filter while improving the computational accuracy. The proposed stochastic gammatone filters are designed and evaluated using MATLAB that achieves a high dynamic range of 71.71 dB compared with a low dynamic range of 5.47 dB in the straightforward implementation. Naoya Onizawa, Shunsuke Koshita, Shuichi Sakamoto, Masahide Abe, Masayuki Kawamata, Takahiro Hanyu |
ICASSP | 6 |
| 2016 | Stochastic behavior-considered VLSI CAD environment for MTJ/MOS-hybrid microprocessor designabstractA new VLSI CAD environment considering stochastic behavior of MTJ devices is proposed for the evaluation of not only the performance but also the reliability of MTJ/MOS-hybrid logic LSI. The proposed simulator allows users to support the design of MTJ/MOS-hybrid LSI by RTL/gate-level hardware description, whose simulation considering stochastic switching behavior of MTJ device can be done by analog-mixed-signal simulation with de-facto standard EDA tools. Through the design of a nonvolatile logic LSI based on a general purpose 32-bit microprocessor, the impact of the proposed design flow is demonstrated. Masanori Natsui, Akira Tamakoshi, Akira Mochizuki, Hiroki Koike, Hideo Ohno, Tetsuo Endoh, Takahiro Hanyu |
ISCAS | 7 |
| 2016 | Standby-Power-Free Integrated Circuits Using MTJ-Based VLSI ComputingabstractNonvolatile spintronic devices have potential advantages, such as fast read/write and high endurance together with back-end-of-the-line compatibility, which offers the possibility of constructing not only stand-alone RAMs and embedded RAMs that can be used in conventional VLSI circuits and systems but also standby-power-free high-performance nonvolatile CMOS logic employing logic-in-memory architecture. The advantages of employing spintronic devices, especially magnetic tunnel junction (MTJ) devices with CMOS circuits, are discussed, and the current status of the MTJ-based VLSI computing paradigm is presented along with its prospects and remaining challenges. Takahiro Hanyu, Tetsuo Endoh, Daisuke Suzuki, Hiroki Koike, Yitao Ma, Naoya Onizawa, Masanori Natsui, Shoji Ikeda, Hideo Ohno |
Proc. IEEE | 1 |
| 2015 | Spintronics-based nonvolatile logic-in-memory architecture towards an ultra-low-power and highly reliable VLSI computing paradigm
Takahiro Hanyu, Daisuke Suzuki, Naoya Onizawa, Shoun Matsunaga, Masanori Natsui, Akira Mochizuki |
DATE | 1 |
| 2015 | Design of a computational nonvolatile RAM for a greedy energy-efficient VLSI processorabstractA computational nonvolatile RAM (C-NVRAM), where magneto-resistive random access memory with spin-transfer torque magnetic tunnel junctions (STT-MTJs) is used as an on-chip storage element, combined with bit-parallel arithmetic modules is proposed for a greedy energy-efficient VLSI processors in the wide range of consumer electronics and mobile applications such as internet-of-things. A judicious combination of bit-serial/word-parallel (at a C-NVRAM) and bit-parallel processing manners makes the calculation cycles reduced in parallel computing such as an image processing. Moreover, since data-access rate of the MTJ-based nonvolatile memory is negligible (is only 1/104percent of memory cell array), its power dissipation is dominated by its static power dissipation. Therefore, the use of nonvolatile memory makes the total power reduced greatly. As a typical application, it is demonstrated in parallel image processing with 8-bit-intensity 256×256 pixels that the energy (computing-time-power-dissipation product) of the proposed hardware is less than 1/15 in comparison with that of the corresponding CMOS-only-based one under a 90nm-CMOS/100nm-MTJ process technologies. Akira Mochizuki, Naoto Yube, Takahiro Hanyu |
IECON | 3 |
| 2015 | Gabor Filter Based on Stochastic ComputationabstractThis letter introduces a design and proof-of-concept implementation of Gabor filters based on stochastic computation for area-efficient hardware. The Gabor filter exhibits a powerful image feature extraction capability, but it requires significant computational power. Using stochastic computation, a sine function used in the Gabor filter is approximated by exploiting several stochastic tanh functions designed based on a state machine. A stochastic Gabor filter realized using the stochastic sine shaper and a stochastic exponential function is simulated and compared with the original Gabor filter that shows almost equivalent behaviour at various frequencies and variance. A root-mean-square error of 0.043 at most is observed. In order to reduce long latency due to stochastic computation, 68 parallel stochastic Gabor filters are implemented in Silterra 0.13 μm CMOS technology. As a result, the proposed Gabor filters achieve a 78% area reduction compared with a conventional Gabor filter while maintaining the comparable speed. Naoya Onizawa, Daisaku Katagiri, Kazumichi Matsumiya, Warren J. Gross, Takahiro Hanyu |
IEEE Signal Process. Lett. | 5 |
| 2014 | A delay circuit with 4-terminal magnetic-random-access-memory device for power-efficient time- domain signal processingabstractA delay circuit using four-terminal magnetic-random-access-memory (MRAM) devices was designed for power-efficient time-domain signal processing. A cell area of 6.4 μm2was obtained using 90-nm CMOS/MRAM technologies. The basic operations to both store the data and control the delay time were confirmed on the fabricated test chips. In addition, we proposed a power-efficient neuromorphic core using the delay circuit. Ryusuke Nebashi, Noboru Sakimura, Hiroaki Honjo, Ayuka Morioka, Yukihide Tsuji, Kunihiko Ishihara, Keiichi Tokutome, Sadahiko Miura, Shunsuke Fukami, Keizo Kinoshita, Takahiro Hanyu, Tetsuo Endoh, Naoki Kasai, Hideo Ohno, Tadahiko Sugibayashi |
ISCAS | 11 |
| 2014 | Energy-aware current-mode inter-chip link for a dependable GALS NoC platformabstractAn inter-chip communication link with high-speed and low-energy capabilities is proposed for a dependable globally-asynchronous-locally-synchronous (GALS) network-on-chip (NoC) platform. The use of a dynamic current feedback mechanism in the link makes a current driving capability high and a signal voltage swing small, which accelerates the switching speed. The power-gating technique is also applied to greatly reduce the power dissipation since the inter-chip communication links are supposed to have long idle time in the dependable GALS NoC platform. It is demonstrated that the proposed circuit with a 1cm transmission line achieves the transmission rate of 2.8Gbps while consuming 1.1uW in a 130nm CMOS technology at the power supply of 1.2V. Hirokatsu Shirahama, Akira Mochizuki, Yuma Watanabe, Takahiro Hanyu |
ISCAS | 4 |
| 2014 | High-Throughput Compact Delay-Insensitive Asynchronous NoC RouterabstractA new asynchronous delay-insensitive data-transmission method based on level-encoded dual-rail (LEDR) encoding with novel packet-structure restriction is proposed to realize a high-throughput network-on-chip (NoC) router together with a compact hardware. The use of LEDR encoding makes communication steps and the registers being used half in comparison with four-phase dual-rail encoding because the spacer information of the four-phase one is eliminated, which significantly improves the network throughput. By using the proposed packet structure, the phase information of header and tail flits is uniquely determined. Since the router can be asynchronously controlled by ignoring the phase information, the circuit is compactly implemented. As a result, the proposed asynchronous NoC router on a 0.13-μm CMOS technology, has a 90 percent increase in throughput and a 34 percent decrease in energy dissipation with 25 percent area overhead in comparison with a conventional four-phase asynchronous NoC router under a postlayout simulation. Under a random traffic pattern in a 4 x 4 2D mesh topology, the proposed asynchronous NoC has a 140 percent increase in throughput and half packet latency compared with the conventional one. We also fabricate the asynchronous NoC based on the proposed router on a 0.13-μm CMOS technology and demonstrate the chip correctly operates under a supply voltage of 0.6 to 1.8 V. Naoya Onizawa, Atsushi Matsumoto, Tomoyoshi Funazaki, Takahiro Hanyu |
IEEE Trans. Computers | 4 |
| 2013 | Challenge of MTJ/MOS-hybrid logic-in-memory architecture for nonvolatile VLSI processorabstractA new logic-circuit style based on nonvolatile logic-in-memory architecture is proposed for realizing compact, low-power logic and highly reliable VLSI processors with parallel data accessibility. Since nonvolatile storage elements such as magnetic tunnel junction (MTJ) devices are distributed over a logic-circuit plane in the proposed style, wide memory bandwidth as well as instant power gating without escaping/reloading data can be realized. As typical examples, and an MTJ-based nonvolatile Ternary Content-Addressable Memory, an MTJ-based nonvolatile look-up table circuit for an instant power-ON/OFF field programmable gate array and a post-process variation-resilient logic-circuit design using MTJ devices are implemented and their superior performances are demonstrated in comparison with a corresponding CMOS-only-based realization. Takahiro Hanyu |
ISCAS | 1 |
| 2013 | MTJ/MOS-hybrid logic-circuit design flow for nonvolatile logic-in-memory LSIabstractA cell-based design flow for MTJ/MOS-hybrid logic circuits is presented towards the realization of practical-scale logic LSI based on nonvolatile logic-in-memory architecture. Newly-developed supplementary design tools including a precise MTJ device model enable to design MTJ/MOS-hybrid logic's intellectual properties (IPs) accurately. By the use of the IPs, various pattern layouts of the MOS and MTJ/MOS-hybrid logic-circuit cells can be automatically synthesized. The effectiveness of the proposed design flow is demonstrated through typical arithmetic-circuit design examples with a nonvolatile storage capability. Masanori Natsui, Takahiro Hanyu, Noboru Sakimura, Tadahiko Sugibayashi |
ISCAS | 2 |
| 2012 | Implementation of a perpendicular MTJ-based read-disturb-tolerant 2T-2R nonvolatile TCAM based on a reversed current reading schemeabstractA perpendicular magnetic-tunnel-junction (MTJ)-based 2T-2R ternary content-addressable memory (TCAM) cell is proposed for a high-density nonvolatile word-parallel/bit-serial TCAM. The use of MOS/MTJ-hybrid logic makes it possible to implement a compact nonvolatile TCAM cell with 2.5 μm2of a cell size in a 0.14-μm CMOS and a 100-nm perpendicular-MTJ technologies. By reversed-current reading through the perpendicular MTJ device, tolerability of read disturb is greatly enhanced. Moreover, fine-grained power gating based on bit-level equality-search scheme achieves ultra-low activity rate of 4.1% in a fabricated 72-bit × 128-word nonvolatile TCAM, which results in ultra-low active power and standby power. Shoun Matsunaga, Masanori Natsui, Shoji Ikeda, Katsuya Miura, Tetsuo Endoh, Hideo Ohno, Takahiro Hanyu |
ASP-DAC | 7 |
| 2012 | Variation-resilient current-mode logic circuit design using MTJ devicesabstractA current-mode logic-circuit style using MTJ devices as threshold voltage (Vth) variation compensating elements is proposed for realizing a process-variation-aware VLSI processor with maintaining a higher performance capability. The faulty logic-operation behavior due to Vthvariation of each MOS transistor can be neglected by adjusting resistance values of MTJ devices that are connected to the source electrode of MOS transistors in series. By using HSPICE simulation under a 90nm CMOS technology, it is demonstrated that basic current-mode logic gates using the proposed method are robust against the Vthvariation. Youngkeun Kim, Masanori Natsui, Takahiro Hanyu |
ISCAS | 3 |
| 2012 | High-speed simulator including accurate MTJ models for spintronics integrated circuit designabstractAn extremely practical simulation program with integrated circuits emphasis (SPICE) incorporating model parameters of magnetic tunnel junction (MTJ) was developed. The simulator provides reliable simulation results in spintronics circuit design because it can accurately calculate various MTJ characteristics that actual devices have, that considerably influence the operation margin and power dissipation. It can also accelerate the simulation speed, which makes it possible to simulate three times or more large-scale circuits than when a conventional macro-model is used. Noboru Sakimura, Ryusuke Nebashi, Yukihide Tsuji, Hiroaki Honjo, Tadahiko Sugibayashi, Hiroki Koike, Takashi Ohsawa, Shunsuke Fukami, Takahiro Hanyu, Hideo Ohno, Tetsuo Endoh |
ISCAS | 9 |
| 2012 | Multi-chip NoCs for Automotive ApplicationsabstractThis paper proposes a multi-chip NoC approach for implementing centralized ECUs. Unlike the conventional approach where ECUs and sensors/actuators are connected tightly, it has potential to implement efficient and reliable systems for automotive applications. Then, this paper reports our experience of implementing our first chip designed for the multi-chip NoC platform, and shows some experimental results. Tomohiro Yoneda, Masashi Imai, Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu |
PRDC | 5 |
| 2011 | Interconnect-fault-resilient delay-insensitive asynchronous communication link based on current-flow monitoringabstractDelay-insensitive asynchronous on-chip communication links are a key element to realize a highly reliable asynchronous Network-on-Chip system. However, even a single permanent fault, such as an interconnect fault, causes a deadlock state in the system. This paper presents an interconnect-fault-resilient delay-insensitive asynchronous communication link based on current-flow monitoring. Since current flow upon an interconnect is cut off by an open fault in the interconnect, the current is fed back to a transmitter, which increases a feedback current monotonically. Monitoring the feedback current makes it possible to detect the interconnect fault with delay insensitivity. The proposed link is evaluated by a 0.13μm CMOS technology with a Triple Modular Redundancy (TMR)-based asynchronous communication link which is resilient to the interconnect fault without the delay insensitivity. As a result, the energy consumption and the number of wires of the proposed link are reduced to 57% and 33%, respectively, in comparison with those of the conventional one. Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu |
DATE | 3 |
| 2011 | Instant power-on nonvolatile FPGA based on MTJ/MOS-hybrid circuitryabstractNo abstract available. Takahiro Hanyu |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Adjacent-State monitoring based fine-grained power-gating scheme for a low-power asynchronous pipelined systemabstractA new gate-level power-gating scheme with a small power-gating controller is proposed for greedily power-aware asynchronous pipelined system. The power supply of each standby stage consisting of a combinational block and a pipeline latch can be cut off by a sleep transistor, because a condition of an asynchronous operation is always monitored by using signal conditions in adjacent stages, which completely eliminates wasted power dissipation in standby stages. Since sleep-transistor control signals in each stage are simply generated by just modifying asynchronous control signals in its adjacent stage, the power- gating controller can be realized by inserting a few basic logic gates. The efficiency of the proposed scheme is demonstrated by using HSPICE simulation. The leakage power dissipation of the asynchronous circuit using the proposed method is reduced to 11.8% in comparison with that of the asynchronous one using a conventional power-gating method. Takao Kawano, Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu |
ISCAS | 4 |
| 2010 | High-throughput protocol converter based on an independent encoding/decoding scheme for asynchronous Network-on-ChipabstractThis paper presents a high-throughput asynchronous protocol converter between two-phase communication links and four-phase pipelined routers for asynchronous Network-on-Chip. In the proposed protocol converter, two-phase input and output signals are encoded to and decoded from the four-phase signals, respectively, by using two controllers which are attached to the router. Since the two controls are realized respectively by using only the input signal and the output signal, the “two-to-four-phase” and “four-to-two-phase” conversions are independently performed. Therefore, each conversion is completely performed without input-output dependency. As a result, the proposed protocol converter achieves an up to 77 % increase in throughput under a comparable energy consumption with respect to that of a conventional protocol converter on a Silterra 0.13-μm CMOS technology. Naoya Onizawa, Takahiro Hanyu |
ISCAS | 2 |
| 2010 | Special session 8B: New topic MOS/MTJ-hybrid circuit with nonvolatile logic-in-memory architecture and its impactabstractSummary form only given. Nonvolatile logic-in-memory architecture, where nonvolatile memory elements are distributed over a logic-circuit plane, is expected to realize both ultra-low-power and reduced interconnection delay. This paper presents novel nonvolatile logic circuits based on logic-in-memory architecture using magnetic tunnel junctions (MTJs) in combination with MOS transistors. Since the MTJ with a spin-injection write capability is only one device that has all the following superior features as large resistance ratio, virtually unlimited endurance, fast read/write accessibility, scalability, complementary MOS (CMOS)-process compatibility, and nonvolatility, it is very suited to implement the MOS/MTJ-hybrid logic circuit with logic-in-memory architecture. As concrete examples of the proposed circuitry, a nonvolatile full adder for motion-vector extraction, a nonvolatile ternary content-addressable memory, and a nonvolatile FPGA, have been proposed. The proposed circuitry also makes it possible to re-program the desired logical threshold of the switching gate even if the transistor characteristics are changed due to the Vthvariation. Figure 1 shows the basic structure of the proposed circuitry. The influence of the Vthvariation can be neglected by adjusting the source voltage of the MOS transistor. Therefore, the Vth-variation compensation is realized by programming the resistance value of the MTJ device. The usefulness of the proposed compensation circuitry is demonstrated at the comparators as shown in Fig.2. [1] S. Matsunaga, et al., APEX, 1, 9, 091301, Aug. 2008. [2] S. Matsunaga, et al., APEX, 2, 2, 023004, Feb. 2009. [3] D. Suzuki, et al., IEEE 2009 Symposia on VLSI Circuits, 80/81, June 2009. Takahiro Hanyu |
VTS | 1 |
| 2010 | Design of High-Throughput Fully Parallel LDPC Decoders Based on Wire PartitioningabstractWe present a method to design high-throughput fully parallel low-density parity-check (LDPC) decoders. With our method, a decoder's longest wires are divided into several short wires with pipeline registers. Log-likelihood ratio messages transmitted along with these pipelined paths are thus sent over multiple clock cycles, and the decoder's critical path delay can be reduced while maintaining comparable bit error rate performance. The number of registers inserted into paths is estimated by using wiring information extracted from initial placement and routing information with a conventional LDPC decoder, and thus only necessary registers are inserted. Also, by inserting an even number of registers into the longer wires, two different codewords can be simultaneously decoded, which improves the throughput at a small penalty in area. We present our design flow as well as post-layout simulation results for several versions of a length-1024, (3,6)-regular LDPC code. Using our technique, we achieve a maximum uncoded throughput of 13.21 Gb/s with an energy consumption of 0.098 nJ per uncoded bit atEb/N0= 5 dB. This represents a 28% increase in throughput, a 30% decrease in energy per bit, and a 1.6% increase in core area with respect to a conventional parallel LDPC decoder, using a 90-nm CMOS technology. Naoya Onizawa, Takahiro Hanyu, Vincent C. Gaudet |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | MTJ-based nonvolatile logic-in-memory circuit, future prospects and issuesabstractNonvolatile logic-in-memory architecture, where nonvolatile memory elements are distributed over a logic-circuit plane, is expected to realize both ultra-low-power and reduced interconnection delay. This paper presents novel nonvolatile logic circuits based on logic-in-memory architecture using magnetic tunnel junctions (MTJs) in combination with MOS transistors. Since the MTJ with a spin-injection write capability is only one device that has all the following superior features as large resistance ratio, virtually unlimited endurance, fast read/write accessibility, scalability, complementary MOS (CMOS)-process compatibility, and nonvolatility, it is very suited to implement the MOS/MTJ-hybrid logic circuit with logic-in-memory architecture. A concrete nonvolatile logic-in-memory circuit is designed and fabricated using a 0.18 mum CMOS/MTJ process, and its future prospects and issues are discussed. Shoun Matsunaga, Jun Hayakawa, Shoji Ikeda, Katsuya Miura, Tetsuo Endoh, Hideo Ohno, Takahiro Hanyu |
DATE | 7 |
| 2009 | High-performance Asynchronous Intra-chip Communication Link based on a Multiple-valued Current-mode Single-track SchemeabstractThis paper presents a high-performance asynchronous data-transfer circuit based on a multiple-valued current-mode single-track scheme for on-chip communication. Since one-bit data and control information are represented by using a multi-level signal in the proposed single-track scheme, one-bit data can be transmitted asynchronously using a single wire between modules. The use of current-mode signaling makes the voltage swing on wires reduced, which achieves high-speed data transfer. Moreover, as the number of current sources is reduced by the reduction of wires, it is possible to achieve low power dissipation. Using the proposed circuit, we achieve a throughput of 0.65 Gbps/wires with power consumption of 0.29 mW at 5 mm wire length. This presents a 400% increase in throughput, a 57% decrease in power consumption with respect to a conventional asynchronous circuit, using a 90 nm CMOS process. Yo Ohtake, Naoya Onizawa, Takahiro Hanyu |
ISCAS | 3 |
| 2007 | Implementation of a Standby-Power-Free CAM Based on Complementary Ferroelectric-Capacitor LogicabstractA complementary ferroelectric-capacitor (CFC) logic-circuit style is proposed for a compact and standby-power-free content-addressable memory (CAM). Since the use of the CFC logic circuit in designing a CAM cell makes it possible to merge both logic and non-volatile storage elements into serially connected ferroelectric capacitors, the CAM becomes compact. The standby power of the CAM is completely eliminated because the supply voltage can be cut off with maintaining stored data in the CAM. The test chip is fabricated by using 0.35-mum ferroelectric CMOS, and the basic behavior can be also measured. Shoun Matsunaga, Takahiro Hanyu, Hiromitsu Kimura, Hidemi Takasu |
ASP-DAC | 2 |
| 2000 | Integration of asynchronous and self-checking multiple-valued current-mode circuits based on dual-rail differential logicabstractA new multiple-valued current-mode (MVCM) integrated circuit based on dual-rail differential logic, whose current-driving capability is high at a low supply voltage, is proposed to realize a totally self-checking circuit and an asynchronous-control circuit. Two nMOS transistors with different threshold voltages are used as complementary pass switches in the proposed differential-pair circuit (DPC), so that the outputs of the DPC always become stable even when non-code-word input (1, 1) are applied, which makes it possible to design a self-checking circuit by using the MVCM circuit. In addition, the dual-rail MVCM circuit technique can be naturally utilized for efficient realization of a two-color dual-rail data-transfer scheme in asynchronous communication. In fact, it is demonstrated that the performance of both a self-checking multiplier and a simple asynchronous control circuit is superior to that of the corresponding ordinary implementation. Takahiro Hanyu, Tsukasa Ike, Michitaka Kameyama |
PRDC | 1 |
| 1997 | Low-power multiple-valued current-mode integrated circuit with current-source control and its applicationabstractA new current-source control technique is proposed to design a low-power high-speed multiple-valued current-mode (MVCM) integrated circuit in a low supply voltage. The use of a differential logic circuit (DLC) with a pair of dual-rail inputs makes the input voltage swing small, which results in a high driving capability at a lower supply voltage, while having large static power dissipation. In the proposed DLC using switched current control, the static power dissipation is greatly reduced because current sources in non-active circuit blocks are switched off. In the current control, no additional transistors are required to control the current sources because a current-control circuit is already used in the threshold detector. As a typical example of arithmetic circuits, a new 1.5 V-supply 54/spl times/54-bit multiplier based on a 0.8 /spl mu/m standard CMOS technology is also designed. Its performance is about 1.3 times faster than that of a binary fastest multiplier under the normalized power dissipation. Takahiro Hanyu, Satoshi Kazama, Michitaka Kameyama |
ASP-DAC | 1 |