Edouard Giacomin

dblp:196/3714 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0002-5415-1870ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2023 3D SRAM Macro Design in 3D Nanofabric Process Technology
abstract
In this paper, we introduce a novel design of a 3D static random-access memory (SRAM) macro in a 3D Nanofabric process technology. The 3D Nanofabric technology is based on enabling the processing of N stack of identical layers simultaneously regardless of the number of stacked layers which consequently reduces the fabrication cost as well as the footprint of SRAM macros. To enable simultaneous patterning of stacked layers, 3D Nanofabric requires circuit topology and layout that rely on a single layer where the device channel, poly, and metal wires are all embedded without any other crossing than the gate on top of the device channel. Accordingly, we modify the layouts of the conventional SRAM bit-cell and periphery circuits which are complex and contain several metal crossings. Furthermore, we propose a new overall organization of the 3D SRAM macro that incorporates a stack of multiple identical layers each consisting of an equal size 2D array of bit-cells and the periphery circuits. We show that the proposed 3D Nanofabric SRAM macro offers 71.2% footprint gain and 36.3% read access speed improvement compared to equal size 2D SRAM macro in 3 nm FinFET.
Dawit Burusie Abdi, Shairfe Muhammad Salahuddin, Jürgen Bömmels, Edouard Giacomin, Pieter Weckx, Julien Ryckaert, Geert Hellings, Francky Catthoor
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Low Latency SEU Detection in FPGA CRAM With In-Memory ECC Checking
abstract
In harsh environments such as space, radiation and charged particles cause Single-Event Effects, faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy, are very effective against Single-Event Transients (SETs), but incur a minimum of$3\times $cost in area. Single-Event Upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process, going through every memory word looking for errors to repair. It involves a non-negligible Time To Detect (TTD) before repair, during which other events can occur and compromise the system. Field Programmable Gate Arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design integrity in harsh conditions. In this paper, we propose In-Memory Error Code Correction Checking (IMECCC), a method to replace memory scrubbing and improve FPGA configuration memory protection in high radiation environments. Our method allows asynchronous SEU detection, and replaces the scrubbing’s variable time to detect with a fixed TTD. We show that IMECCC reduces FPGA’s TTD by at least 116,$000\times $on average, with an area increase of$1.56\times $, using a test architecture resembling a Xilinx Virtex 5 QV at a 60MHz scrubbing frequency.
Aurélien Alacchi, Edouard Giacomin, Scott Temple, Roman Gauchi, Michael J. Wirthlin, Pierre-Emmanuel Gaillardon
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Smart-Redundancy With In Memory ECC Checking: Low-Power SEE-Resistant FPGA Architectures
abstract
In harsh environments, such as space, radiation, and charged particles cause single-event effects (SEEs), faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy (TMR), are very effective against single-event transients (SETs) but incur a minimum of$3\times $cost in the area. Single-event upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process going through every memory word, looking for errors to repair. Scrubbing involves a nonnegligible amount of time before an error is detected, during which other events can occur and compromise the system. Field-programmable gate arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design sustainability in harsh conditions. In this article, we propose an alternative mitigation method based on sensor integration and FPGA architecture modification, called smart-redundancy with in-memory error correction code checking (SRIMECCC). The sensors allow asynchronous SEU detection, reducing the time to detect by 57$250\times $on average and enabling local reconfiguration. Our method includes built-in dual redundancy that reduces the power consumption by 87% on average, benefiting embedded systems. SRIMECCC is also an area-efficient technique that saves 28% of total effective area compared to TMRed designs implemented in FPGAs.
Aurélien Alacchi, Edouard Giacomin, Roman Gauchi, Szymon Kulis, Pierre-Emmanuel Gaillardon
IEEE Trans. Very Large Scale Integr. Syst.2
2021 Smart-Redundancy: An Alternative SEU/SET Mitigation Method for FPGAs
abstract
Field Programmable Gate Arrays (FPGAs) reconfigurability is a key asset for many critical applications. State- of-the-art Radiation-Hardening (Rad-Hard) methods for FPGAs consist of triplicating the logic, reinforcing the memories, and bitstream scrubbing with partial reconfiguration. These methods involve a 3× reduction of Maximal Design Capacity (MDC) and an average Time-In-Error (TIE) proportional to design sizes. In this paper, we propose an alternative: Smart-Redundancy (SR), a new method based on the detection of possible events via process and hardware modifications. Thanks to integrated particle sensors, only dual redundancy is required. Results show up to 33.33% improvement in MDC over actual Rad-Hard methods, and an average TIE decrease of at least 10,000× compared to bitstream's scrubbing, at a cost of 41.08% in area using a commercial 40nm technology node.
Aurélien Alacchi, Edouard Giacomin, Xifan Tang, Pierre-Emmanuel Gaillardon
ISCAS2
2021 Area-Efficient Multiplier Designs Using a 3D Nanofabric Process Flow
abstract
In the past few years, the demand for computationally intensive applications, such as digital signal processing or convolutional neural networks, has grown exponentially. As they often rely on a significant number of multiply-and-accumulate cells, it is crucial to optimize their area and cost. Recently, a 3D Nanofabric flow has been proposed, where logic circuits are designed by stacking N identical vertical tiers on top of each other. Exploiting identical layers allows a fabrication process similar to the Vertical-NAND flash, where all the layers can be patterned at once. While the 3D Nanofabric flow presents several layout constraints (single metal routing and identical vertical layers), it can decrease the area by around one order of magnitude, leading to area-efficient and cost-effective circuits. In this paper, we propose to use the 3D Nanofabric process flow to design low-area multipliers. As multipliers can be designed using a regular array organization, we show how they can be spread across multiple vertical layers using the 3D Nanofabric flow, while respecting the different layout constraints. We then provide thorough circuit-level evaluations, including parasitics, to showcase the benefits of our proposed 3D multipliers at the circuit-level. We show that by stacking up to 64 layers to build a 64-input bit multiplier, the area and area-delay- product can be decreased by 28.6x and 25.5x, respectively, compared to a traditional 2D implementation using a 28nm FDSOI technology, with only a 10% and 35% delay and power consumption overheads, respectively.
Edouard Giacomin, Francky Catthoor, Pierre-Emmanuel Gaillardon
ISCAS1
2021 A Novel High-Gain Amplifier Circuit Using Super-Steep-Subthreshold-Slope Field-Effect Transistors
abstract
The benefits of steep-Subthreshold Swing (SS) devices, though plentiful at the device-level, have yet to be fully exploited at the circuit-level. This is evident from a look at the Three-Independent-Gate Field-Effect Transistor (TIGFET), a device renown for its ability for polarity reconfiguration. At the same time, its demonstrated dynamic control of the subthreshold slope beyond the thermal limit has only been studied at the device-level. This latter benefit is referred to as Super-Steep Subthreshold Slope (S4) operation and can lead to unprecedented gain, which is ideal for use in an amplifier circuit. In this paper, we investigate the impact of S4 operations when designing differential-amplifier circuits when using TIGFET technology. We demonstrate the benefits of our implementation both from a theoretical standpoint and through circuit-level analyses. More specifically, we show that the TIGFET -based amplifier gain is 95.5 $\times$ better, and that the gain-bandwidth product is improved by 13.8, compared to an equivalent MOSFET-based design at the 90 nm node. Besides, we show that at equivalent gains, the TIGFET-based amplifier decreases the area and power by 22.8 $\times$ and 7.2, respectively, against its MOSFET counterpart.
Matthieu Couriol, Patsy Cadareanu, Edouard Giacomin, Pierre-Emmanuel Gaillardon
VLSI-SoC3
2021 A 12-pA Resolution Sigma Delta ADC Topology for Chemiresistive Sensor-Based Applications
abstract
Chemiresistive sensor technologies can detect chemical trace in the air down to the low Part Per Billion (ppb)/Part Per Trillion (ppt) level. Achieving such low detection levels requires low-noise and high-accuracy analog front-end interfaces. One such interface uses a Sigma Delta $(\Sigma\triangle)$ Analog to Digital Converter (ADC) topology that leads to lower noise at low frequency and better effective resolution. The slow charging dynamics of chemicals in the environment is well-suited for $\Sigma\triangle$ converters. While $\Sigma\triangle$ converters are a well-studied topic, the potential resolution of chemical sensing using chemiresistive sensors has not yet been demonstrated in a fully integrated mixed-signal interface. In this paper, we develop a $\Sigma\triangle$ architecture combined with a 10-pA resolution Digitally Controlled Current Sources (DCCS) and an integrated Proportional Integrator Derivative (PID) controller feedback loop. The PID drives back the current sources at 250 kHz to maintain a constant voltage across the sensor, acting as a variable resistor Implemented on a 180 nm CMOS process with an 8-bits current source resolution at 250 kHz and coupled with nanofiber chemiresistive sensors, it translates into lppb level detection. Compared to the state-of-the-art chemical sensing interface, our circuit shows a 5$\times$improvement in detection level and a $ 10\times$ improvement in power efficiency, given the same detection level.
Matthieu Couriol, Edouard Giacomin, Pierre-Emmanuel Gaillardon
VLSI-SoC2
2020 A RRAM-based FPGA for Energy-efficient Edge Computing
abstract
The shift from centralized cloud to edge computing demands hardware systems with data processing capability at ultra-low power. Reconfigurable solutions such as Field-Programmable Gate Arrays (FPGAs) offer a high flexibility in terms of hardware implementation and are thus popular for use in many edge computing systems. However, breaking through the energy wall of FPGAs is a challenge, as low-power operation often requires compromising performances. In this paper, we study a low-power high-performance FPGA architecture exploiting Resistive Random Access Memory (RRAM) technology. To perform a comprehensive analysis, we introduce a novel design flow which can rapidly prototype FPGA fabrics from which accurate area, delay, and power results can be obtained. Based on full-chip layouts and SPICE simulations, we show that RRAM-based FPGAs can improve up to 8%/22%/16% in area/delay/power compared to SRAM-based counterparts at nominal voltage. Even when operated at a near-Vtsupply, the proposed RRAM-based FPGA can improve the Energy-Delay Product by about 2× without any delay overhead, when compared to an SRAM-based FPGA. In addition, Monte Carlo simulations showed that the proposed RRAM-based FPGA architecture stays robust under different CMOS process corners as well as under a 30% RRAM resistance standard deviation.
Xifan Tang, Edouard Giacomin, Patsy Cadareanu, Ganesh Gore, Pierre-Emmanuel Gaillardon
DATE2
2020 SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are expected to be part of the future AI portfolio, with heavy investment from industry and government, e.g., IBM TrueNorth, Intel Loihi. While Artificial Neural Network (ANN) architectures have taken large strides, few works have targeted SNN hardware efficiency. Our analysis of SNN baselines shows that at modest spike rates, SNN implementations exhibit significantly lower efficiency than accelerators for ANNs. This is primarily because SNN dataflows must consider neuron potentials for several ticks, introducing a new data structure and a new dimension to the reuse pattern. We introduce a novel SNN architecture, SpinalFlow, that processes a compressed, time-stamped, sorted sequence of input spikes. It adopts an ordering of computations such that the outputs of a network layer are also compressed, time-stamped, and sorted. All relevant computations for a neuron are performed in consecutive steps to eliminate neuron potential storage overheads. Thus, with better data reuse, we advance the energy efficiency of SNN accelerators by an order of magnitude. Even though the temporal aspect in SNNs prevents the exploitation of some reuse patterns that are more easily exploited in ANNs, at 4-bit input resolution and 90% input sparsity, SpinalFlow reduces average energy by 1.8×, compared to a 4-bit Eyeriss baseline. These improvements are seen for a range of networks and sparsity/resolution levels; SpinalFlow consumes 5× less energy and 5.4× less time than an 8-bit version of Eyeriss. We thus show that, depending on the level of observed sparsity, SNN architectures can be competitive with ANN architectures in terms of latency and energy for inference, thus lowering the barrier for practical deployment in scenarios demanding real-time learning.
Surya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin, Pierre-Emmanuel Gaillardon
ISCA4
2020 Layout Considerations of Logic Designs Using an N-layer 3D Nanofabric Process Flow
abstract
In the past few years, novel fabrication schemes such as parallel and monolithic 3D integration have been proposed to keep sustaining the need for more powerful integrated circuits. By stacking several devices, wafers, or dies, the footprint, delay, and power can be decreased when compared to traditional 2D implementations. While parallel 3D does not enable very fine-grained vertical connections, monolithic 3D currently only offers a limited number of transistor tiers due to the high cost of the additional masks and processing steps, limiting the benefits of using the third dimension. In this paper, we introduce an innovative planar circuit netlist and layout approach, which enables a new 3D integration flow called 3D Nanofabric. The flow, consisting of$N$identical vertical tiers, is aimed at single instruction multiple data processor Arithmetic Logic Units (ALUs). By using a single metal routing layer for each vertical tier, the process flow is significantly simplified since multiple vertical layers can potentially be patterned at once, similar to the 3D NAND flash process. In our study, we thoroughly investigate the layout constraints arising from the Nanofabric flow and the unique metal layer rule and propose several ways to overcome them. We then show that by stacking 32 layers to build a 32-bit ALU, the footprint is reduced by$8.7\times$when compared to a conventional 7nm FinFET implementation.
Edouard Giacomin, Jürgen Bömmels, Julien Ryckaert, Francky Catthoor, Pierre-Emmanuel Gaillardon
VLSI-SOC1
2019 OpenFPGA: An Opensource Framework Enabling Rapid Prototyping of Customizable FPGAs
abstract
Driven by the strong need in data processing applications, Field Programmable Gate Arrays (FPGAs) are playing an ever-increasing role as programmable accelerators in modern computing systems. To fully unlock processing capabilities for domain-specific applications, FPGA architectures have to be tailored for seamless cooperation with other computing resources. However, prototyping and bringing to production a customized FPGA is a costly and complex endeavor even for industrial vendors. In this paper, we introduce OpenFPGA, an opensource framework that enables rapid prototyping of customizable FPGA architectures through a semi-custom design approach. We propose an XML-to-Prototype design flow, where the Verilog netlists of a full FPGA fabric can be autogenerated using an extension of the XML language from the VTR framework and then fed into a back-end flow to generate production-ready layouts. OpenFPGA also includes a general-purpose Verilog-to-Bitstream generator for any FPGA described by the XML language. We demonstrate the capability of this automatic design flow with a Stratix IV-like FPGA architecture using a commercial 40nm technology node, and perform a detailed comparison to its academic and commercial counterparts. Compared to the current state-of-art academic results, our FPGA fabric reduces the area by 1:75 and the delay by 3 on average. In addition, OpenFPGA significantly reduces the gap between semi-custom designed FPGAs and fully-optimized commercial products with a penalty of only 60% in area and 30% in delay, respectively.
Xifan Tang, Edouard Giacomin, Aurélien Alacchi, Baudouin Chauviere, Pierre-Emmanuel Gaillardon
FPL2
2019 Wire-Aware Architecture and Dataflow for CNN Accelerators
abstract
In spite of several recent advancements, data movement in modern CNN accelerators remains a significant bottleneck. Architectures like Eyeriss implement large scratchpads within individual processing elements, while architectures like TPU v1 implement large systolic arrays and large monolithic caches. Several data movements in these prior works are therefore across long wires, and account for much of the energy consumption. In this work, we design a new wire-aware CNN accelerator, WAX, that employs a deep and distributed memory hierarchy, thus enabling data movement over short wires in the common case. An array of computational units, each with a small set of registers, is placed adjacent to a subarray of a large cache to form a single tile. Shift operations among these registers allow for high reuse with little wire traversal overhead. This approach optimizes the common case, where register fetches and access to a few-kilobyte buffer can be performed at very low cost. Operations beyond the tile require traversal over the cache's H-tree interconnect, but represent the uncommon case. For high reuse of operands, we introduce a family of new data mappings and dataflows. The best dataflow, WAXFlow-3, achieves a 2× improvement in performance and a 2.6-4.4× reduction in energy, relative to Eyeriss. As more WAX tiles are added, performance scales well until 128 tiles.
Sumanth Gudaparthi, Surya Narayanan, Rajeev Balasubramonian, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon
MICRO4
2019 GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment
abstract
Precision Medicine will rely on frequent genomic analysis, especially for patients undergoing cancer treatments or suffering from rare diseases. Sequence alignment is invoked in multiple stages of the genomic analysis pipeline. Recent projects have introduced accelerators, GenAx and Darwin, for 2nd and 3rd generation sequencers respectively. In this work, we improve upon the GenAx design by increasing its parallelism and reducing its memory bandwidth demands. This is achieved with a combination of hardware and software innovations. We first integrate in-cache operators from prior work into the GenAx memory hierarchy; we then augment the in-cache peripheral circuit to support additional new operators. We then re-structure the sequence alignment algorithm to (i) leverage the many in-cache operators, (ii) exploit the common case in genomic datasets, (iii) use Bloom Filters to reduce futile accesses, and (iv) maximize data reuse within a re-organized memory hierarchy. While the baseline GenAx accelerator processes a batch of reads in 194 seconds while nearly saturating the 153.6 GB/s memory bandwidth, the proposed GenCache architecture processes the same batch of reads in 37 seconds at an improved energy efficiency of 8.6×, while demanding 20 GB/s average memory bandwidth. Our hardware and software techniques thus interact synergistically to target both memory and compute bottlenecks, while not affecting the outputs of the application. We show that the basic principles in GenCache can also be exploited by 3rd generation sequence aligners.
Anirban Nag, C. N. Ramachandra, Rajeev Balasubramonian, Ryan Stutsman, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon
MICRO5
2019 A Predictive Process Design Kit for Three-Independent-Gate Field-Effect Transistors
abstract
The Three-Independent-Gate Field-Effect Transistor (TIGFET) is a promising beyond-CMOS technology which offers many unique properties, such as (i) dynamic control of the device polarity, (ii) dual threshold operation and (iii) more expressive logic capabilities. The efficient exploitation of these properties provides opportunity to design area and power optimized logic circuits. However, the evaluation of TIGFET-based design currently relies on a close approximation for the Power, Performance, and Area (PPA) rather than traditional layout-based methods. There is a need for a publicly available Process Design Kit (PDK) enabling systematic evaluation of the design area. In this paper, we propose Predictive PDK for the 10 nm-diameter silicon-nanowire TIGFET device. This work consists of a SPICE model and full custom physical design files including a Design Rule Manual, a Design Rule Check, and a Layout Versus Schematic decks for Calibre®. We then validate the design rules through the implementation of basic logic gates and a full-adder and compare extracted metrics with FreePDK15nm™ PDK. We show 26% and 41% area reduction in the case of an XOR gate and a 1-bit full-adder design respectively.
Ganesh Gore, Patsy Cadareanu, Edouard Giacomin, Pierre-Emmanuel Gaillardon
VLSI-SoC3
2019 A Product Engine for Energy-Efficient Execution of Binary Neural Networks Using Resistive Memories
abstract
The need for running complex Machine Learning (ML) algorithms, such as Convolutional Neural Networks (CNNs), in edge devices, which are highly constrained in terms of computing power and energy, makes it important to execute such applications efficiently. The situation has led to the popularization of Binary Neural Networks (BNNs), which significantly reduce execution time and memory requirements by representing the weights (and possibly the data being operated) using only one bit. Because approximately 90% of the operations executed by CNNs and BNNs are convolutions, a significant part of the memory transfers consists of fetching the convolutional kernels. Such kernels are usually small (e.g., 3×3 operands), and particularly in BNNs redundancy is expected. Therefore, equal kernels can be mapped to the same memory addresses, requiring significantly less memory to store them. In this context, this paper presents a custom Binary Dot Product Engine (BDPE) for BNNs that exploits the features of Resistive Random-Access Memories (RRAMs). This new engine allows accelerating the execution of the inference phase of BNNs. The novel BDPE locally stores the most used binary weights and performs binary convolution using computing capabilities enabled by the RRAMs. The system-level gem5 architectural simulator was used together with a C-based ML framework to evaluate the system's performance and obtain power results. Results show that this novel BDPE improves performance by 11.3%, energy efficiency by 7.4% and reduces the number of memory accesses by 10.7% at a cost of less than 0.3% additional die area, when integrated with a 28 nm Fully Depleted Silicon On Insulator ARMv8 in-order core, in comparison to a fully-optimized baseline of YoloV3 XNOR-Net running in a unmodified Central Processing Unit.
João Vieira, Edouard Giacomin, Yasir Mahmood Qureshi, Marina Zapater, Xifan Tang, Shahar Kvatinsky, David Atienza 0001, Pierre-Emmanuel Gaillardon
VLSI-SoC2
2019 FPGA-SPICE: A Simulation-Based Architecture Evaluation Framework for FPGAs
abstract
In this paper, we developed a simulation-based architecture evaluation framework for field-programmable gate arrays (FPGAs), called FPGA-SPICE, which enables automatic layout-level estimation and electrical simulations of FPGA architectures. FPGA-SPICE can automatically generate Verilog and SPICE netlists based on realistic FPGA configurations and a high-level eTtensible Markup Language-based FPGA architectural description language. The outputted Verilog netlists can be used to generate layouts of full FPGA fabrics through a semicustom design flow. SPICE simulation decks can be generated at three levels of complexity, namely, full-chip-level, grid-level, and component-level, providing different tradeoff between accuracy and simulation time. In order to enable such level of analysis, we presented two SPICE netlist partitioning techniques: loads extraction and parasitic net activity estimation. Electrical simulations showed that averaged over the selected benchmarks, the grid-/component-level approach can achieve 6.1×/7.5× execution speed-up with 9.9%/8.3% accuracy loss, respectively, compared to the full-chip level simulation. FPGA-SPICE was showcased through three different case studies: (1) an area breakdown analysis for static random access memory-based FPGAs, showing that configuration memories are a dominant factor; (2) a power breakdown comparison to analytical models, analyzing the source of accuracy loss; and (3) a robustness evaluation against process corners, studying their impact on energy consumption of full FPGA fabrics.
Xifan Tang, Edouard Giacomin, Giovanni De Micheli, Pierre-Emmanuel Gaillardon
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Differential Power Analysis Mitigation Technique Using Three-Independent-Gate Field Effect Transistors
abstract
Hardware security vulnerabilities are a major concern for embedded computing devices which are now used in many application such as credit cards, SIM cards, or financial systems, putting sensible data at risk. Such systems are often targeted by differential power attacks, where the power trace can be monitored in order to get access to the sensible data. To alleviate this issue, a possible technique proposed in literature is to use a complementary gate (e.g., computing both XOR and XNOR operations in parallel) in order to have a symmetrical power trace for all possible input combinations. However, this technique results in a large area and power overhead since it approximatively requires twice the number of transistors. Recently, novel technologies such as Three-Independent-Gate Field Effect Transistors (TIGFETs) have been shown to be able to realize compact logic gates using less transistors when compared to Complementary Metal Oxide Semiconductor (CMOS) technology. In this paper, we investigate the benefits of using TIGFETs in terms of hardware security. First, we show that using the complementary gate technique with TIGFETs can reduce the transistor count, the power trace variation, the switching power and leakage by 2×, 57%, 36% and 8× respectively, when compared to CMOS. In addition, we show that for the same transistor count and similar switching power, using TIGFETs can reduce the power trace variation and the leakage by 81% and 6.7× respectively when compared to CMOS.
Edouard Giacomin, Pierre-Emmanuel Gaillardon
VLSI-SoC1
2017 Physical Design Considerations of One-level RRAM-based Routing Multiplexers
abstract
Resistive Random Access Memory(RRAM) technology opens the opportunity for granting both high-performance and low-power features to routing multiplexers. In this paper, we study the physical design considerations related to RRAM-based routing multiplexers and particularly the integration of 4T(ransistor)1R(RAM) programming structures within their routing tree. We first analyze the limitations in the physical design of a naive one-level 4T1R-based multiplexer, such as co-integration of low-voltage nominal power supply and high voltage programming supply, as well as the use of long metal wires across different isolating wells. To address the limitations, we improve the one-level 4T1R-based multiplexer by re-arranging the nominal and programming voltage domains, and also study the optimal location of RRAMs in terms of performance. The improved design can effectively reduce the length of long metal wires by 50%. Electrical simulations show that using a 7nm FinFET transistor technology, the improved 4T1R-based multiplexers improve delay by 69% as compared to the basic design. At nominal working voltage, considering an input size ranging from 2 to 32, the improved 4T1R-based multiplexers outperform the best CMOS multiplexers in area by 1.4x, delay by 2x and power by 2x respectively. The improved 4T1R-based multiplexers operating at near-Vt regime can improve Power-Delay Product by up to 5.8x when compare to the best CMOS multiplexers working at nominal voltage.
Xifan Tang, Edouard Giacomin, Giovanni De Micheli, Pierre-Emmanuel Gaillardon
ISPD2