Naoya Onizawa

dblp:63/2907 · DBLP profile ↗
← Back
28ranked-venue papers
12as first author
5since 2021 · last 2025
0000-0002-4855-7081ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 9 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
abstract
In this study, we propose an implementation methodology of real-time few-shot learning on tiny FPGA SoCs such as the PYNQ-Z1 board with arbitrary fixed-point bit-widths. Tensil-based conventional design environments limited hardware implementations to fixed-point bit-widths of 16 or 32 bits. To address this, we adopt the FINN framework, enabling implementations with arbitrary bit-widths. Several customizations and minor adjustments are made, including: 1.Optimization of Transpose nodes to resolve data format mismatches, 2.Addition of handling for converting the final "reduce mean" operation to Global Average Pooling (GAP). These adjustments allow us to reduce the bit-width while maintaining the same accuracy as the conventional realization, and achieve approximately twice the throughput in evaluations using CIFAR-10 dataset.
R. Kanda, Hugo Le Blevec, Naoya Onizawa, Mathieu Léonardon, Vincent Gripon, Takahiro Hanyu
ISCAS3
2023 Fast-Converging Simulated Annealing for Ising Models Based on Integral Stochastic Computing
abstract
Probabilistic bits (p-bits) have recently been presented as a spin (basic computing element) for the simulated annealing (SA) of Ising models. In this brief, we introduce fast-converging SA based on p-bits designed using integral stochastic computing. The stochastic implementation approximates a p-bit function, which can search for a solution to a combinatorial optimization problem at lower energy than conventional p-bits. Searching around the global minimum energy can increase the probability of finding a solution. The proposed stochastic computing-based SA method is compared with conventional SA and quantum annealing (QA) with a D-Wave Two quantum annealer on the traveling salesman, maximum cut (MAX-CUT), and graph isomorphism (GI) problems. The proposed method achieves a convergence speed a few orders of magnitude faster while dealing with an order of magnitude larger number of spins than the other methods.
Naoya Onizawa, Kota Katsuki, Duckgyu Shin, Warren J. Gross, Takahiro Hanyu
IEEE Trans. Neural Networks Learn. Syst.1
2021 High Convergence Rates of CMOS Invertible Logic Circuits Based on Many-Body Hamiltonians
abstract
This paper introduces CMOS invertible-logic (CIL) circuits based on many-body Hamiltonians. CIL can realize probabilistic forward and backward operations of a function by annealing a corresponding Hamiltonian using stochastic computing. We have created a Hamiltonian that includes three-body interaction of spins (probabilistic nodes). It provides some degrees of freedom to design a simpler landscape of Hamiltonian (energy) than that of the conventional two-body Hamiltonian. The simpler landscape makes it easier to reach the global minimum energy. The proposed three-body CIL circuits are designed and evaluated with the conventional two- body CIL circuits, resulting in few-times higher convergence rates with negligible area overhead on FPGA.
Naoya Onizawa, Takahiro Hanyu
ISCAS1
2021 A Design Framework for Invertible Logic
abstract
Invertible logic using a probabilistic magnetoresistive device model has been recently presented that can compute functions in bidirectional ways and solve several problems quickly, such as factorization and combinational optimization. In this article, we present a design framework for invertible logic circuits. Our approach makes use of linear programming to create a Hamiltonian library with the minimum number of nodes for small invertible-logic functions. In addition, as the device model is approximated based on stochastic computing in synthesizable SystemVerilog, a faster simulation using the compiled SystemC binary is realized than a conventional SPICE-level simulation and is verified using field-programmable gate array (FPGA) as prototyping. Using our design framework, several invertible-logic circuits are designed and emulated (verified) in SystemC, exhibiting five order-of-magnitude faster simulation than conventional work.
Naoya Onizawa, Kaito Nishino, Sean C. Smithson, Brett H. Meyer, Warren J. Gross, Hitoshi Yamagata, Hiroyuki Fujita, Takahiro Hanyu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Multi-Context TCAM-Based Selective Computing: Design Space Exploration for a Low-Power NN
abstract
In this paper, we propose a low-power memory-based computing architecture, called selective computing architecture (SCA). It consists of multipliers and an LUT (Look-Up Table)-based component, that is multi-context ternary content-addressable memory (MC-TCAM). Either of them is selected by input-data conditions in neural-networks (NNs). Compared with quantized NNs, a higher accurate multiplication can be performed with low-power consumption in the proposed architecture. If input data stored in the MC-TCAM appears, the corresponding multiplication results for multiple weights are obtained. The MC-TCAM stores only shorter length of input data, resulting in achieving a low-power computing. The performance of the SCA is determined by three physical parameters concerning the configuration of MC-TCAM. The power dissipation of the target NN can be minimized by exploring these parameters in the design space. The hardware based on the proposed architecture is evaluated using TSMC 65 nm CMOS technology and MTJ model. In the case of speech command recognition, the power consumption at the multiplication of the first convolutional layer in a convolutional NN is reduced by 67% compared to the solution relying only on multipliers.
Ren Arakawa, Naoya Onizawa, Jean-Philippe Diguet, Takahiro Hanyu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Memristive Computational Memory Using Memristor Overwrite Logic (MOL)
abstract
In this article, we present a novel logic design style, namely, memristor overwrite logic (MOL), associated with an original MOL-based computational memory. MOL relies on a fully digital representation of memristor and can operate with different memristive device technologies. Its integration in memristive crossbar arrays and computational memories allows the execution of bit and vector-level primitive logic operations in two computational steps at most. Promising features and performances are demonstrated through the implementation of N -bit full addition using the proposed MOL-based computational memory.
Khaled Alhaj Ali, Mostafa Rizk, Amer Baghdadi, Jean-Philippe Diguet, Jalal Jomaah, Naoya Onizawa, Takahiro Hanyu
IEEE Trans. Very Large Scale Integr. Syst.6
2020 High-Throughput/Low-Energy MTJ-Based True Random Number Generator Using a Multi-Voltage/Current Converter
abstract
This article introduces high-throughput/low-energy true random number generators (TRNGs) based on CMOS and three-terminal magnetic tunnel junction (MTJ) devices. MTJs are fast and probabilistic switching devices, which can be used as random number sources for TRNGs. However, as the switching probability is quite sensitive to the write current given to MTJs, precise closed-loop control is necessary. Thus, a high-complexity current control circuit is required, such as high precision digital-to-analog converters (DACs), occupying large area and causing large energy dissipation. In order to address the issue, we propose a multi-voltage/current (V/I) converter capable of multilevel coarse current switching and fine adjusting within each level. The fine adjusting can be done by DACs with fewer bits, resulting in much smaller size and energy dissipation than a conventional single-V/I converter. In addition, a multiple-writing scheme for three-terminal MTJs is proposed for increasing the throughput while maintaining the write power. The proposed TRNGs are designed using TSMC 65-nm CMOS and a threeterminal MTJ model that achieves a throughput of 333 Mb/s, an energy dissipation of 0.66 pJ/bit and an area of 2040 μm2. This result exhibits a 5× throughput, a 93% energy reduction and an 86% area reduction in comparison with a conventional CMOS/MTJ-based TRNG.
Naoya Onizawa, Shogo Mukaida, Akira Tamakoshi, Hitoshi Yamagata, Hiroyuki Fujita, Takahiro Hanyu
IEEE Trans. Very Large Scale Integr. Syst.1
2018 High-Precision Stochastic State-Space Digital Filters Based on Minimum Roundoff Noise Structure
abstract
Digital filters based on stochastic computation have recently gained considerable attention because the stochastic computation attains significant reduction of hardware complexity of digital filters compared with the classical deterministic binary computation-based filters. For stochastic IIR filters, Liu and Parhi proposed an elegant method that achieves high precision by means of the normalized state-space lattice structure. This paper extends the result and presents stochastic state-space IIR filters with higher precision than the conventional method. Instead of using the normalized lattice structure, our method makes use of the minimum roundoff noise structure in realizing stochastic state-space filters, leading to further improvement of arithmetic precision of IIR filtering compared with the conventional method. Experimental results show that our method gives higher signal-to-error ratio at the filter outputs than the conventional method.
Shunsuke Koshita, Naoya Onizawa, Masahide Abe, Takahiro Hanyu, Masayuki Kawamata
ISCAS2
2018 Networked Power-Gated MRAMs for Memory-Based Computing
abstract
Emerging nonvolatile memory technologies open new perspectives for original computing architectures. In this paper, we propose a new type of flexible and energy-efficient architecture that relies on power-gated distributed magnetoresistive random access memory (MRAM). The proposed architecture uses a network-on-chip (NoC) to interconnect MRAM-based clusters, processing elements, and managers. The NoC distributes application-specific commands to MRAM devices by means of packets. Configurable network interfaces allow to transform MRAM devices into smart units able to respond to incoming commands. In this context, three types of MRAM designs are proposed with different power-gating policies and granularities. A relevant database search engine case study is considered to illustrate the benefits of this proposed architecture. It is implemented with a sparse-neural-network approach and simulated in SystemC with different scenarios including hundreds of database queries. Hardware designs and accurate power estimations have been conducted. The obtained results demonstrate important power reduction with database hit rates of about 94%. Targeting 65-nm technology, energy savings reach 87% when compared with an static random access memory-based implementation. Moreover, a new asymmetric read/write MRAM type provides from 39% to 50% energy reduction with respect to the other fixed-granularity models. This results in a low-power, highly scalable, and configurable implementation of memory-based computing.
Jean-Philippe Diguet, Naoya Onizawa, Mostafa Rizk, Martha Johanna Sepúlveda, Amer Baghdadi, Takahiro Hanyu
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Three-terminal MTJ-based nonvolatile logic circuits with self-terminated writing mechanism for ultra-low-power VLSI processor
abstract
Magnetic-Tunnel Junction (MTJ)-based non-volatile logic circuits have some possibility to solve the power-dissipation problem seriously focusing on the present CMOS-only-based VLSI processors. Three terminal MTJ devices are the promising candidate as nonvolatile storage device to realize such a nonvolatile logic circuit. However, its writing energy is still serious in comparison with conventional CMOS-only-based logic circuits. In this paper, a new MTJ-based nonvolatile logic circuit with self-terminated mechanism is proposed and its energy efficiency is evaluated in comparison with the corresponding previous work. In addition, some recent research topics related to MTJ-based nonvolatile logic-circuit design and its application, such as a computer-aided-design (CAD) tool considering a stochastic MTJ-switching behavior and the application to a resilient “die-hard” VLSI processor against sudden power-supply outage, are also demonstrated.
Takahiro Hanyu, Daisuke Suzuki, Naoya Onizawa, Masanori Natsui
DATE3
2017 VLSI Implementation of Deep Neural Network Using Integral Stochastic Computing
abstract
The hardware implementation of deep neural networks (DNNs) has recently received tremendous attention: many applications in fact require high-speed operations that suit a hardware implementation. However, numerous elements and complex interconnections are usually required, leading to a large area occupation and copious power consumption. Stochastic computing (SC) has shown promising results for low-power area-efficient hardware implementations, even though existing stochastic algorithms require long streams that cause long latencies. In this paper, we propose an integer form of stochastic computation and introduce some elementary circuits. We then propose an efficient implementation of a DNN based on integral SC. The proposed architecture has been implemented on a Virtex7 field-programmable gate array, resulting in 45% and 62% average reductions in area and latency compared with the best reported architecture in the literature. We also synthesize the circuits in a 65-nm CMOS technology, and we show that the proposed integral stochastic architecture results in up to 21% reduction in energy consumption compared with the binary radix implementation at the same misclassification rate. Due to fault-tolerant nature of stochastic architectures, we also consider a quasi-synchronous implementation that yields 33% reduction in energy consumption with respect to the binary radix implementation without any compromise on performance.
Arash Ardakani, François Leduc-Primeau, Naoya Onizawa, Takahiro Hanyu, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Area/Energy-Efficient Gammatone Filters Based on Stochastic Computation
abstract
This paper introduces area/energy-efficient gammatone filters based on stochastic computation. The gammatone filter well expresses the performance of human auditory peripheral mechanism and has a potential of improving advanced speech communications systems, especially hearing assisting devices and noise robust speech-recognition systems. Using stochastic computation, a power-and-area hungry multiplier used in a digital filter is replaced by a simple logic gate, leading to area-efficient hardware. However, a straightforward implementation of the stochastic gammatone filter suffers from significantly low accuracy in computation, which results in a low dynamic range (a ratio of the maximum to minimum magnitude) due to a small value of a filter gain. To improve the computation accuracy, gain-balancing techniques are presented that represent the original gain as the product of multiple larger gains introduced at the second-order sections. In addition, dynamic scaling techniques are proposed that scales up small values only on stochastic domain in order to reduce the number of stochastic bits required while maintaining the computation accuracy. For performance comparisons, the proposed stochastic gammatone filters are designed and evaluated on taiwan semiconductor manufacturing company (TSMC) 65-nm CMOS technology. As a result, the proposed filter achieves an area reduction of 90.7% and an energy reduction of 91.8% in comparison with a fixed-point gammatone filter at the same sampling frequency and a comparable dynamic range.
Naoya Onizawa, Shunsuke Koshita, Shuichi Sakamoto, Masahide Abe, Masayuki Kawamata, Takahiro Hanyu
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Gammatone filter based on stochastic computation
abstract
This paper introduces a design of a gammatone filter based on stochastic computation for area-efficient hardware. The gammatone filter well expresses the performance of human auditory peripheral mechanism and has a potential of improving advanced speech communications systems, especially hearing assisting devices and noise robust speech recognition systems. Using stochastic computation, a power-and-area hungry multiplier used in a digital filter is replaced by a simple logic gate, leading to area-efficient hardware. However, a straightforward implementation of the stochastic gammatone filter suffers from significantly low accuracy in computation, which results in a low dynamic range (a ratio of the maximum to minimum magnitude) due to a small value of a filter gain. To improve the computational accuracy, gain-balancing techniques are presented that represent the original gain as the product of multiple larger gains introduced at the second-order sections. As a result, the proposed techniques maintain the original gain of the filter while improving the computational accuracy. The proposed stochastic gammatone filters are designed and evaluated using MATLAB that achieves a high dynamic range of 71.71 dB compared with a low dynamic range of 5.47 dB in the straightforward implementation.
Naoya Onizawa, Shunsuke Koshita, Shuichi Sakamoto, Masahide Abe, Masayuki Kawamata, Takahiro Hanyu
ICASSP1
2016 Standby-Power-Free Integrated Circuits Using MTJ-Based VLSI Computing
abstract
Nonvolatile spintronic devices have potential advantages, such as fast read/write and high endurance together with back-end-of-the-line compatibility, which offers the possibility of constructing not only stand-alone RAMs and embedded RAMs that can be used in conventional VLSI circuits and systems but also standby-power-free high-performance nonvolatile CMOS logic employing logic-in-memory architecture. The advantages of employing spintronic devices, especially magnetic tunnel junction (MTJ) devices with CMOS circuits, are discussed, and the current status of the MTJ-based VLSI computing paradigm is presented along with its prospects and remaining challenges.
Takahiro Hanyu, Tetsuo Endoh, Daisuke Suzuki, Hiroki Koike, Yitao Ma, Naoya Onizawa, Masanori Natsui, Shoji Ikeda, Hideo Ohno
Proc. IEEE6
2015 Spintronics-based nonvolatile logic-in-memory architecture towards an ultra-low-power and highly reliable VLSI computing paradigm
Takahiro Hanyu, Daisuke Suzuki, Naoya Onizawa, Shoun Matsunaga, Masanori Natsui, Akira Mochizuki
DATE3
2015 Gabor Filter Based on Stochastic Computation
abstract
This letter introduces a design and proof-of-concept implementation of Gabor filters based on stochastic computation for area-efficient hardware. The Gabor filter exhibits a powerful image feature extraction capability, but it requires significant computational power. Using stochastic computation, a sine function used in the Gabor filter is approximated by exploiting several stochastic tanh functions designed based on a state machine. A stochastic Gabor filter realized using the stochastic sine shaper and a stochastic exponential function is simulated and compared with the original Gabor filter that shows almost equivalent behaviour at various frequencies and variance. A root-mean-square error of 0.043 at most is observed. In order to reduce long latency due to stochastic computation, 68 parallel stochastic Gabor filters are implemented in Silterra 0.13 μm CMOS technology. As a result, the proposed Gabor filters achieve a 78% area reduction compared with a conventional Gabor filter while maintaining the comparable speed.
Naoya Onizawa, Daisaku Katagiri, Kazumichi Matsumiya, Warren J. Gross, Takahiro Hanyu
IEEE Signal Process. Lett.1
2015 Algorithm and Architecture for a Low-Power Content-Addressable Memory Based on Sparse Clustered Networks
abstract
We propose a low-power content-addressable memory (CAM) employing a new algorithm for associativity between the input tag and the corresponding address of the output data. The proposed architecture is based on a recently developed sparse clustered network using binary connections that on-average eliminates most of the parallel comparisons performed during a search. Therefore, the dynamic energy consumption of the proposed design is significantly lower compared with that of a conventional low-power CAM design. Given an input tag, the proposed architecture computes a few possibilities for the location of the matched tag and performs the comparisons on them to locate a single valid match. TSMC 65-nm CMOS technology was used for simulation purposes. Following a selection of design parameters, such as the number of CAM entries, the energy consumption and the search delay of the proposed design are 8%, and 26% of that of the conventional NAND architecture, respectively, with a 10% area overhead. A design methodology based on the silicon area and power budgets, and performance requirements is discussed.
Hooman Jarollahi, Vincent Gripon, Naoya Onizawa, Warren J. Gross
IEEE Trans. Very Large Scale Integr. Syst.3
2014 High-Throughput Compact Delay-Insensitive Asynchronous NoC Router
abstract
A new asynchronous delay-insensitive data-transmission method based on level-encoded dual-rail (LEDR) encoding with novel packet-structure restriction is proposed to realize a high-throughput network-on-chip (NoC) router together with a compact hardware. The use of LEDR encoding makes communication steps and the registers being used half in comparison with four-phase dual-rail encoding because the spacer information of the four-phase one is eliminated, which significantly improves the network throughput. By using the proposed packet structure, the phase information of header and tail flits is uniquely determined. Since the router can be asynchronously controlled by ignoring the phase information, the circuit is compactly implemented. As a result, the proposed asynchronous NoC router on a 0.13-μm CMOS technology, has a 90 percent increase in throughput and a 34 percent decrease in energy dissipation with 25 percent area overhead in comparison with a conventional four-phase asynchronous NoC router under a postlayout simulation. Under a random traffic pattern in a 4 x 4 2D mesh topology, the proposed asynchronous NoC has a 140 percent increase in throughput and half packet latency compared with the conventional one. We also fabricate the asynchronous NoC based on the proposed router on a 0.13-μm CMOS technology and demonstrate the chip correctly operates under a supply voltage of 0.6 to 1.8 V.
Naoya Onizawa, Atsushi Matsumoto, Tomoyoshi Funazaki, Takahiro Hanyu
IEEE Trans. Computers1
2013 A low-power Content-Addressable Memory based on clustered-sparse networks
abstract
A low-power Content-Addressable Memory (CAM) is introduced employing a new mechanism for associativity between the input tags and the corresponding address of the output data. The proposed architecture is based on a recently developed clustered-sparse network using binary-weighted connections that on-average will eliminate most of the parallel comparisons performed during a search. Therefore, the dynamic energy consumption of the proposed design is significantly lower compared to that of a conventional low-power CAM design. Given an input tag, the proposed architecture computes a few possibilities for the location of the matched tag and performs the comparisons on them to locate a single valid match. A 0.13μm CMOS technology was used for simulation purposes. The energy consumption and the search delay of the proposed design are 9.5%, and 30.4% of that of the conventional NAND architecture respectively with a 3.4% higher number of transistors.
Hooman Jarollahi, Vincent Gripon, Naoya Onizawa, Warren J. Gross
ASAP3
2013 Low-power area-efficient large-scale IP lookup engine based on binary-weighted clustered networks
abstract
We propose a novel architecture for low-power area-efficient large-scale IP lookup engines. The proposed architecture greatly increases memory efficiency by storing associations between IP addresses and their output rules instead of storing these data themselves. The rules can be determined by simple hardware using a few associations read from SRAMs, eliminating a power-hungry search of input addresses in TCAMs. The proposed hardware that stores 100,000 144-bit entries is evaluated under TSMC 65nm CMOS technology. The dynamic power dissipation and the area of the proposed hardware are 4.6% and 30.6% of a traditional TCAM, respectively while maintaining comparable throughput.
Naoya Onizawa, Warren J. Gross
DAC1
2013 Reduced-complexity binary-weight-coded associative memories
abstract
Associative memories retrieve stored information given partial or erroneous input patterns. Recently, a new family of associative memories based on Clustered-Neural-Networks (CNNs) was introduced that can store many more messages than classical Hopfield-Neural Networks (HNNs). In this paper, we propose hardware architectures of such memories for partial or erroneous inputs. The proposed architectures eliminate winner-take-all modules and thus reduce the hardware complexity by consuming 65% fewer FPGA lookup tables and increase the operating frequency by approximately 1.9 times compared to that of previous work.
Hooman Jarollahi, Naoya Onizawa, Vincent Gripon, Warren J. Gross
ICASSP2
2012 Architecture and implementation of an associative memory using sparse clustered networks
abstract
Associative memories are alternatives to indexed memories that when implemented in hardware can benefit many applications such as data mining. The classical neural network based methodology is impractical to implement since in order to increase the size of the memory, the number of information bits stored per memory bit (efficiency) approaches zero. In addition, the length of a message to be stored and retrieved needs to be the same size as the number of nodes in the network causing the total number of messages the network is capable of storing (diversity) to be limited. Recently, a novel algorithm based on sparse clustered neural networks has been proposed that achieves nearly optimal efficiency and large diversity. In this paper, a proof-of-concept hardware implementation of these networks is presented. The limitations and possible future research areas are discussed.
Hooman Jarollahi, Naoya Onizawa, Vincent Gripon, Warren J. Gross
ISCAS2
2012 Multi-chip NoCs for Automotive Applications
abstract
This paper proposes a multi-chip NoC approach for implementing centralized ECUs. Unlike the conventional approach where ECUs and sensors/actuators are connected tightly, it has potential to implement efficient and reliable systems for automotive applications. Then, this paper reports our experience of implementing our first chip designed for the multi-chip NoC platform, and shows some experimental results.
Tomohiro Yoneda, Masashi Imai, Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu
PRDC3
2011 Interconnect-fault-resilient delay-insensitive asynchronous communication link based on current-flow monitoring
abstract
Delay-insensitive asynchronous on-chip communication links are a key element to realize a highly reliable asynchronous Network-on-Chip system. However, even a single permanent fault, such as an interconnect fault, causes a deadlock state in the system. This paper presents an interconnect-fault-resilient delay-insensitive asynchronous communication link based on current-flow monitoring. Since current flow upon an interconnect is cut off by an open fault in the interconnect, the current is fed back to a transmitter, which increases a feedback current monotonically. Monitoring the feedback current makes it possible to detect the interconnect fault with delay insensitivity. The proposed link is evaluated by a 0.13μm CMOS technology with a Triple Modular Redundancy (TMR)-based asynchronous communication link which is resilient to the interconnect fault without the delay insensitivity. As a result, the energy consumption and the number of wires of the proposed link are reduced to 57% and 33%, respectively, in comparison with those of the conventional one.
Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu
DATE1
2011 Adjacent-State monitoring based fine-grained power-gating scheme for a low-power asynchronous pipelined system
abstract
A new gate-level power-gating scheme with a small power-gating controller is proposed for greedily power-aware asynchronous pipelined system. The power supply of each standby stage consisting of a combinational block and a pipeline latch can be cut off by a sleep transistor, because a condition of an asynchronous operation is always monitored by using signal conditions in adjacent stages, which completely eliminates wasted power dissipation in standby stages. Since sleep-transistor control signals in each stage are simply generated by just modifying asynchronous control signals in its adjacent stage, the power- gating controller can be realized by inserting a few basic logic gates. The efficiency of the proposed scheme is demonstrated by using HSPICE simulation. The leakage power dissipation of the asynchronous circuit using the proposed method is reduced to 11.8% in comparison with that of the asynchronous one using a conventional power-gating method.
Takao Kawano, Naoya Onizawa, Atsushi Matsumoto, Takahiro Hanyu
ISCAS2
2010 High-throughput protocol converter based on an independent encoding/decoding scheme for asynchronous Network-on-Chip
abstract
This paper presents a high-throughput asynchronous protocol converter between two-phase communication links and four-phase pipelined routers for asynchronous Network-on-Chip. In the proposed protocol converter, two-phase input and output signals are encoded to and decoded from the four-phase signals, respectively, by using two controllers which are attached to the router. Since the two controls are realized respectively by using only the input signal and the output signal, the “two-to-four-phase” and “four-to-two-phase” conversions are independently performed. Therefore, each conversion is completely performed without input-output dependency. As a result, the proposed protocol converter achieves an up to 77 % increase in throughput under a comparable energy consumption with respect to that of a conventional protocol converter on a Silterra 0.13-μm CMOS technology.
Naoya Onizawa, Takahiro Hanyu
ISCAS1
2010 Design of High-Throughput Fully Parallel LDPC Decoders Based on Wire Partitioning
abstract
We present a method to design high-throughput fully parallel low-density parity-check (LDPC) decoders. With our method, a decoder's longest wires are divided into several short wires with pipeline registers. Log-likelihood ratio messages transmitted along with these pipelined paths are thus sent over multiple clock cycles, and the decoder's critical path delay can be reduced while maintaining comparable bit error rate performance. The number of registers inserted into paths is estimated by using wiring information extracted from initial placement and routing information with a conventional LDPC decoder, and thus only necessary registers are inserted. Also, by inserting an even number of registers into the longer wires, two different codewords can be simultaneously decoded, which improves the throughput at a small penalty in area. We present our design flow as well as post-layout simulation results for several versions of a length-1024, (3,6)-regular LDPC code. Using our technique, we achieve a maximum uncoded throughput of 13.21 Gb/s with an energy consumption of 0.098 nJ per uncoded bit atEb/N0= 5 dB. This represents a 28% increase in throughput, a 30% decrease in energy per bit, and a 1.6% increase in core area with respect to a conventional parallel LDPC decoder, using a 90-nm CMOS technology.
Naoya Onizawa, Takahiro Hanyu, Vincent C. Gaudet
IEEE Trans. Very Large Scale Integr. Syst.1
2009 High-performance Asynchronous Intra-chip Communication Link based on a Multiple-valued Current-mode Single-track Scheme
abstract
This paper presents a high-performance asynchronous data-transfer circuit based on a multiple-valued current-mode single-track scheme for on-chip communication. Since one-bit data and control information are represented by using a multi-level signal in the proposed single-track scheme, one-bit data can be transmitted asynchronously using a single wire between modules. The use of current-mode signaling makes the voltage swing on wires reduced, which achieves high-speed data transfer. Moreover, as the number of current sources is reduced by the reduction of wires, it is possible to achieve low power dissipation. Using the proposed circuit, we achieve a throughput of 0.65 Gbps/wires with power consumption of 0.29 mW at 5 mm wire length. This presents a 400% increase in throughput, a 57% decrease in power consumption with respect to a conventional asynchronous circuit, using a 90 nm CMOS process.
Yo Ohtake, Naoya Onizawa, Takahiro Hanyu
ISCAS2