VLDB 2026 Research / reviewers in the wild / expert
Pierre-Emmanuel Gaillardon
dblp:48/9513
· DBLP profile ↗
119ranked-venue papers
15as first author
36since 2021 · last 2026
0000-0003-3634-3999ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 116 · 15 first-author · 35 since 2021Software engineering, systems software and programming languages · 25 · 4 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Special Day - A Design Blueprint for Scalable Multi-Agent Architectures in Complex EDA WorkflowsabstractElectronic Design Automation (EDA) workflows involve complex, tightly coupled tools and artifacts that require high reliability and traceability. Recent advances in Large Language Models (LLMs) have opened new avenues for AI-driven automation through Multi-Agent Systems (MASs), which can decompose complex tasks into manageable subtasks. However, deploying MASs in EDA remains challenging due to weak coordination, unstructured communication, and limited reproducibility. In this paper, we propose a design blueprint for scalable, LLM-based, MASs tailored to EDA workflows, emphasizing hierarchical orchestration, explicit task interfaces, tool-grounded execution, structured communication, modular memory management, and observability with recovery paths. We also introduce Nexus, an open-source Software Development Kit (SDK) that implements these principles, thus enabling low-code workflow specification and robust agent interactions. We validate our approach on open-source benchmarks, achieving 100% functional accuracy on RTL generation (VerilogEval-Human), up to 98.78% functional pass rate on HumanEval, and timing closure with 26.64% average LUT reduction and almost 30% lower total power on VTR designs. Valerio Tenace, Pierre-Emmanuel Gaillardon |
DATE | 2 |
| 2026 | OpenFPGA-NoC: Automated Fabric and Bitstream Generation for NoC-based FPGAsabstractAs the demand for high-performance and flexible hardware accelerators increases, Network-on-Chip (NoC)-based Field Programmable Gate Arrays (FPGAs) offer a scalable solution for complex, data-intensive applications. While commercial FPGA vendors like Xilinx, Altera, and Achronix offer hardened NoCs in their flagship architectures, there are no academic or open source FPGAs with embedded NoCs. Although many open source soft NoC implementations exist, they pose challenges like high resource utilization and low frequencies, making them unsuitable for high-performance applications. Vendor-supplied FPGAs with fixed hard NoC topologies may not satisfy the requirements of new application domains, motivating the need to enable automatic design of customized NoC-based FPGAs. To address this need, we introduce OpenFPGA-NoC, an automated flow that generates fabric netlists and bitstreams for NoC-based FPGAs. Our work extends the OpenFPGA framework by adding an NoC-specific tag in the OpenFPGA architecture description, supporting custom configuration ports to handle address mapping of NoC routers, automating the generation of architecture files, and enabling a custom RTL-to-bitstream flow. OpenFPGA-NoC provides an easy-to-use interface that allows the FPGA architect to exploit the flexibility provided by the framework- providing NoC parameters like topology, number of routers, and key router parameters like data widths and buffer depths. By providing push-button flows, OpenFPGA-NoC significantly lowers the barrier to designing high-performance FPGA fabrics. Ruthwik Reddy Sunketa, Ganesh Gore, Allen Boston, Pierre-Emmanuel Gaillardon, Aman Arora 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2025 | EDA-Aware RTL Generation with Large Language ModelsabstractLarge Language Models (LLMs) have become increasingly popular for generating RTL code. However, producing error-free RTL code in a zero-shot setting remains highly challenging even for state-of-the-art LLMs, often leading to issues that require manual, iterative refinement. This additional debugging process can dramatically increase the verification workload, underscoring the need for robust, automated correction mechanisms to ensure code correctness from the start. In this work, we introduce AIVRIL2, a self-verifying, LLM-agnostic agentic framework aimed at enhancing RTL code generation through iterative corrections of both syntax and functional errors. Our approach leverages a collaborative multi-agent system that incorporates feedback from error logs generated by EDA tools to automatically identify and resolve design flaws. Experimental results, conducted on the VerilogEval-Human benchmark suite, demonstrate that our framework significantly improves code quality, achieving nearly a 3.4× enhancement over prior methods. In the best-case scenario, functional pass rates of 77% for Verilog and 66% for VHDL were obtained, thus substantially improving the reliability of LLM-driven RTL code generation. Mubashir ul Islam, Humza Sami, Pierre-Emmanuel Gaillardon, Valerio Tenace |
DATE | 3 |
| 2025 | OpenMFDA: Microfluidic Design Automation in Three DimensionsabstractCurrent microfluidic design automation (MFDA) solutions are limited by the planarity requirements of current manufacturing techniques. Recent advances in stereolithography 3D printing create an opportunity for new MFDA design methodologies. We propose a methodology for the placement of microfluidic components and the routing of flow and control channels in three dimensions. Additionally, we propose a methodology for generating a printable 3D structure from the layout. We then present OpenMFDA, an open-source MFDA design flow implementing the proposed methodologies. This design flow takes a structural netlist and produces a sliced design for manufacturing using an SLA 3D printer. Our methodology demonstrates short run times and generates devices with 2–20 x smaller area compared to state-of-the-art MFDA tools. Ashton Snelgrove, Daniel Wakeham, Skylar Stockham, Scott Temple, Pierre-Emmanuel Gaillardon |
DATE | 5 |
| 2025 | Lightweight Congruence Profiling for Early Design Exploration of Heterogeneous FPGAsabstractField-Programmable Gate Arrays (FPGAs) have evolved from uniform logic arrays into heterogeneous fabrics integrating digital signal processors (DSPs), memories, and specialized accelerators to support emerging workloads such as machine learning. While these enhancements improve power, performance, and area (PPA), they complicate design space exploration and application optimization due to complex resource interactions. To address these challenges, we propose a lightweight profiling methodology inspired by the Roofline model. It introduces three congruence scores that quickly identify bottlenecks related to heterogeneous resources, fabric, and application logic. Evaluated on the Koios and VPR benchmark suites using a Stratix 10-like FPGA, this approach enables efficient FPGA architecture codesign to improve heterogeneous FPGA performance. Allen Boston, Biruk B. Seyoum, Luca P. Carloni, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 4 |
| 2025 | Automated Generation of Microfluidic Netlists using Large Language ModelsabstractMicrofluidic devices have emerged as powerful tools in various laboratory applications, but the complexity of their design limits accessibility for many practitioners. While progress has been made in microfluidic design automation (MFDA), a practical and intuitive solution is still needed to connect microfluidic practitioners with MFDA techniques. This work introduces the first practical application of large language models (LLMs) in this context, providing a preliminary demonstration. Building on prior research in hardware description language (HDL) code generation with LLMs, we propose an initial methodology to convert natural language microfluidic device specifications into system-level structural Verilog netlists. We demonstrate the feasibility of our approach by generating structural netlists for practical benchmarks representative of typical microfluidic designs with correct functional flow and an average syntactical accuracy of $\mathbf{8 8} \boldsymbol{\%}$. Jasper Davidson, Skylar Stockham, Allen Boston, Ashton Snelgrove, Valerio Tenace, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 6 |
| 2025 | Changing Idling Behavior Through Dynamic Idle Detection and Air Quality MessagingabstractAir quality impacts on human health are an increasing concern globally. Vehicle pollution is a particular concern because of its multiple adverse health effects, and discretionary vehicle idling contributes significantly to local-scale poor air quality. This study introduces a novel approach to traditional static (non-changing) anti-idling signage. Here, we demonstrate a system, called SmartAir, that provides dynamic social-norm messages to drivers coupled with information about idling status or vehicle emissions in the area. A machine learning algorithm with audio and video inputs determines vehicle idling status. Vehicle emissions are measured using a suite of low-cost air quality nodes. In this study, we show that the SmartAir system reduces idling time by 28.0% and local CO2concentrations by 29.5% compared to background. Tristalee Mangin, Xiwen Li, Saba Mahmoudi, Rehman Mohammed, Nathan Page, Sara Peck, Ashton Snelgrove, Evan Blanchard, Dillon Tang, Lizzie Pinegar, Owen Leishman, J. Nicholas Rice, Gregory Madden, Pierre-Emmanuel Gaillardon, Ross T. Whitaker, Kerry E. Kelly |
IEEE Internet Things J. | 14 |
| 2025 | ARIANNA: An Automatic Design Flow for Fabric Customization and eFPGA RedactionabstractIn the modern global Integrated Circuit (IC) supply chain, protecting intellectual property (IP) is a complex challenge, and balancing IP loss risk and added cost for theft countermeasures is hard to achieve. Using embedded configurable logic allows designers to completely hide the functionality of selected design portions from parties that do not have access to the configuration string (bitstream). However, the design space of redacted solutions is huge, with tradeoffs between the portions selected for redaction and the configuration of the configurable embedded logic. We propose ARIANNA, a complete flow that aids the designer in all the stages, from selecting the logic to be hidden to tailoring the bespoke fabrics for the configurable logic used to hide it. We present a security evaluation of the considered fabrics and introduce two heuristics for the novel bespoke fabric flow. We evaluate the heuristics against an exhaustive approach. We also evaluate the complete flow using a selection of benchmarks. Results show that using ARIANNA to customize the redaction fabrics yields up to 3.3× lower overheads and 4× higher eFPGA fabric utilization than a one-fits-all fabric as proposed in prior works. Luca Collini, Jitendra Bhandari, Chiara Muscari Tomajoli, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Xifan Tang, Pierre-Emmanuel Gaillardon, Ramesh Karri, Christian Pilato |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2024 | Benchmarking Microfluidic Design Automation FlowsabstractIn this paper, we propose a methodology for measuring figures of merit relevant to microfluidics practitioners. We also present a benchmark suite for microfluidics design automation (MFDA). The suite is composed of generated circuits designed to challenge tools for specific figures of merit identified in the measurement methodology. We survey the MFDA literature to evaluate the state-of-the-art for benchmark and measurement methodologies. We include in the benchmark distribution a complete set of transcriptions of the major benchmarks used in the MFDA literature, including designs previously unavailable. Benchmarks are distributed in the major hardware description languages along with reported measurements Ashton Snelgrove, Skylar Stockham, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 3 |
| 2023 | Open-source and FPGAs: Hardware, Software, Both or None?abstractFollowing the footsteps of the open-source software movement that is at the foundation of many fundamental infrastructures today, e.g., Linux, the internet, etc., a growing amount of open-source hardware initiatives have been impacting our field, e.g., the RISC-V ISA, Open chiplet standards, etc. Dana How, Tim Ansell, Vaughn Betz, Chris Lavin, Ted Speers, Pierre-Emmanuel Gaillardon |
FPGA | 6 |
| 2023 | FlowTune: End-to-End Automatic Logic Optimization Exploration via Domain-Specific Multiarmed BanditabstractDesign flows are the explicit combinations of design transformations, primarily involved in synthesis, placement, and routing processes, to accomplish the design of integrated circuits (ICs) and system-on-chip (SoC). Mostly, the flows are developed based on the knowledge of the experts. However, due to the large search space of design flows and the increasing design complexity, developing intellectual property (IP)-specific synthesis flows providing high quality of result (QoR) is extremely challenging. In recent years, machine learning (ML) has been increasingly used in electronic design automation (EDA), with the goal of reducing manual labor and speeding up the design closure process in current toolflows. Existing techniques, on the other hand, either necessitate a huge amount of labeled data and time-consuming training, or are constrained in terms of practical EDA toolflow integration due to computational overhead. This article presents a generic end-to-end sequential decision making framework FlowTune for synthesis tooflow optimization, with a novel high-performance domain-specific, multistage multiarmed bandit (MAB) approach. This framework addresses a wide range of optimization problems on Boolean optimization problems, such as And-Inv-Graphs (AIGs), conjunction normal form (CNF) minimization (# clauses) for Boolean satisfiability; logic synthesis and technology mapping, and, more importantly, end-to-end post place-and-route (PnR) optimizations. Moreover, we demonstrate the high extensibility and generalizability of the proposed domain-specific MAB approach with end-to-end FPGA design flow, evaluated at post-routing stage, with two different FPGA backend tools (OpenFPGA and VPR) and two different logic synthesis representations [AIGs and Majority-Inv-Graph (MIG)]. FlowTune is fully integrated with ABC (Mishchenko et al., 2010), Yosys (Wolf, 2016), VTR (Luu et al., 2014), LSOracle (Neto et al., 2019), OpenFPGA (Tang et al., 2019), and industrial tools, and is released publicly. The experimental results conducted on various design stages in the flow all demonstrate that our framework outperforms both handcrafted flows (Mishchenko et al., 2010) and ML explored flows (Yu et al., 2018), (Hosny et al., 2019) in QoRs, and is orders of magnitude faster compared to ML-based approaches. Walter Lau Neto, Pierre-Emmanuel Gaillardon, Cunxi Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Low Latency SEU Detection in FPGA CRAM With In-Memory ECC CheckingabstractIn harsh environments such as space, radiation and charged particles cause Single-Event Effects, faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy, are very effective against Single-Event Transients (SETs), but incur a minimum of$3\times $cost in area. Single-Event Upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process, going through every memory word looking for errors to repair. It involves a non-negligible Time To Detect (TTD) before repair, during which other events can occur and compromise the system. Field Programmable Gate Arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design integrity in harsh conditions. In this paper, we propose In-Memory Error Code Correction Checking (IMECCC), a method to replace memory scrubbing and improve FPGA configuration memory protection in high radiation environments. Our method allows asynchronous SEU detection, and replaces the scrubbing’s variable time to detect with a fixed TTD. We show that IMECCC reduces FPGA’s TTD by at least 116,$000\times $on average, with an area increase of$1.56\times $, using a test architecture resembling a Xilinx Virtex 5 QV at a 60MHz scrubbing frequency. Aurélien Alacchi, Edouard Giacomin, Scott Temple, Roman Gauchi, Michael J. Wirthlin, Pierre-Emmanuel Gaillardon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning AccelerationabstractBlock random access memories (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using logic blocks and digital signal processing slices. We propose modifying BRAMs to convert them to CoMeFa ( Co mpute-in- Me mory Blocks for F PG A s) random access memories (RAMs). These RAMs provide highly parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual-port nature of FPGA BRAMs and contain multiple configurable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute with any precision, which is extremely important for applications like deep learning (DL). Adding CoMeFa RAMs to FPGAs significantly increases their compute density while also reducing data movement. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying static RAM technology like simultaneously activating multiple wordlines on the same port, and are practical to implement. CoMeFa RAMs are especially suitable for parallel and compute-intensive applications like DL, but these versatile blocks find applications in diverse applications like signal processing and databases, among others. By augmenting an Intel Arria 10–like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55× (1.85×) across microbenchmarks from various applications and a geomean speedup of up to 2.5× across multiple deep neural networks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of DL workloads. Aman Arora 0001, Atharva Bhamburkar, Aatman Borda, Tanmay Anand, Rishabh Sehgal, Bagus Hanindhito, Pierre-Emmanuel Gaillardon, Jaydeep P. Kulkarni, Lizy Kurian John |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2023 | Smart-Redundancy With In Memory ECC Checking: Low-Power SEE-Resistant FPGA ArchitecturesabstractIn harsh environments, such as space, radiation, and charged particles cause single-event effects (SEEs), faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy (TMR), are very effective against single-event transients (SETs) but incur a minimum of$3\times $cost in the area. Single-event upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process going through every memory word, looking for errors to repair. Scrubbing involves a nonnegligible amount of time before an error is detected, during which other events can occur and compromise the system. Field-programmable gate arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design sustainability in harsh conditions. In this article, we propose an alternative mitigation method based on sensor integration and FPGA architecture modification, called smart-redundancy with in-memory error correction code checking (SRIMECCC). The sensors allow asynchronous SEU detection, reducing the time to detect by 57$250\times $on average and enabling local reconfiguration. Our method includes built-in dual redundancy that reduces the power consumption by 87% on average, benefiting embedded systems. SRIMECCC is also an area-efficient technique that saves 28% of total effective area compared to TMRed designs implemented in FPGAs. Aurélien Alacchi, Edouard Giacomin, Roman Gauchi, Szymon Kulis, Pierre-Emmanuel Gaillardon |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Not All Fabrics Are Created Equal: Exploring eFPGA Parameters for IP RedactionabstractSemiconductor design houses rely on third-party foundries to manufacture their integrated circuits (ICs). While this trend allows them to tackle fabrication costs, it introduces security concerns as external (and potentially malicious) parties can access critical parts of the designs and steal or modify the intellectual property (IP). Embedded field-programmable gate array (eFPGA) redaction is a promising technique to protect critical IPs of an ASIC by redacting (i.e., removing) critical parts and mapping them onto a custom reconfigurable fabric. Only trusted parties will receive the correct bitstream to restore the redacted functionality. While previous studies imply that using an eFPGA is a sufficient condition to provide security against IP threats like reverse-engineering, whether this truly holds for all eFPGA architectures is unclear, thus motivating the study in this article. We examine the security of eFPGA fabrics generated by varying different FPGA design parameters. We characterize the power, performance, and area (PPA) characteristics and evaluate each fabric’s resistance to Boolean satisfiability (SAT)-based bitstream recovery. Our results encourage designers to work with custom eFPGA fabrics rather than off-the-shelf commercial FPGAs and reveals that only considering a redaction fabric’s bitstream size is inadequate for gauging security. Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Christian Pilato, Ganesh Gore, Xifan Tang, Scott Temple, Pierre-Emmanuel Gaillardon, Ramesh Karri |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2023 | A Scalable and Area-Efficient Configuration Circuitry for Semi-Custom FPGA DesignabstractConfiguration circuitry is an essential component of a field-programmable gate array (FPGA) fabric, which enables the configuration of each programmable logic and takes over 40% of the FPGA chip area. Following the recent trends of automated custom FPGA design, it is essential to study the impact on the area and power of the configuration circuitry. This study compares the performance of different configuration circuitries using a strictly automated and complete standard cell-based semi-custom design methodology. We leverage an open-source framework, OpenFPGA, and extended it to support two configuration protocols: shift-register-based configuration (SRC) and memory-bank-based configuration (MBC) circuitries and their variants. We proposed area optimization strategies to improve the physical implementation of each configuration circuitry. Our results show that compared with naive SRC implementation, the proposed optimization strategies minimize the area overhead by more than 30% and power dissipation during programming by 20%. Whereas compared with MBC implementation, the optimized SRC implementation requires approximately a similar area. However, considering the more practical implementation of MBC with the write-verify functionality with optimized SRC implementation, the MBC implementation requires an 11% higher area and 30% higher routing wirelength, but results in a 62% reduction in the power dissipation during programming. Ganesh Gore, Xifan Tang, Pierre-Emmanuel Gaillardon |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Improving LUT-based optimization for ASICsabstractLUT-based optimization techniques are finding new applications in synthesis of ASIC designs. Intuitively, packing logic into LUTs provides a better balance between functionality and structure in logic optimization. On this basis, the LUT-engine framework [1] was introduced to enhance the ASIC synthesis. In this paper, we present key improvements, at both algorithmic and flow levels, making a much stronger LUT-engine. We restructure the flow of LUT-engine, to benefit from a heterogeneous mixture of LUT sizes, and revisit its requirements for maximum scalability. We propose a dedicated LUT mapper for the new flow, based on FlowMap, natively balancing LUT-count and NAND2-count for a wide range LUT sizes. We describe a specialized Boolean factoring technique, exploiting the fanin bounds in LUT networks, resulting in a very fast LUT-based AIG minimization. By using the proposed methodology, we improve 9 of the best area results in the ongoing EPFL synthesis competition. Integrated in a complete EDA flow for ASICs, the new LUT-engine performs well on a set of 87 benchmarks: -4.60% area and -3.41% switching power at +5% runtime, compared to the baseline flow without LUT-based optimizations, and -3.02% area and -2.54% switching power with -1% runtime, compared to the original LUT-engine. Walter Lau Neto, Luca G. Amarù, Vinicius Possani, Patrick Vuillod, Jiong Luo, Alan Mishchenko, Pierre-Emmanuel Gaillardon |
DAC | 7 |
| 2022 | ALICE: an automatic design flow for eFPGA redactionabstractFabricating an integrated circuit is becoming unaffordable for many semiconductor design houses. Outsourcing the fabrication to a third-party foundry requires methods to protect the intellectual property of the hardware designs. Designers can rely on embedded reconfigurable devices to completely hide the real functionality of selected design portions unless the configuration string (bitstream) is provided. However, selecting such portions and creating the corresponding reconfigurable fabrics are still open problems. We propose ALICE, a design flow that addresses the EDA challenges of this problem. ALICE partitions the RTL modules between one or more reconfigurable fabrics and the rest of the circuit, automating the generation of the corresponding redacted design. Chiara Muscari Tomajoli, Luca Collini, Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Xifan Tang, Pierre-Emmanuel Gaillardon, Ramesh Karri, Christian Pilato |
DAC | 7 |
| 2022 | Programmable logic elements using multigate ambipolar transistorsabstractWe propose a general purpose logic element with eight variations, built using multigate ambipolar transistors, sufficiently capable to replace LUTs in FPGAs. We simulate the new logic element using a 10nm silicon-nanowire three-input-gate transistor model, and compare the proposed element to lookup tables and reconfigurable logic elements from the literature implemented using the same technology model. We compare the different elements for delay, power, and number of transistors, specifically accounting for the cost of configuration storage. Compared to an equivalent LUT, the logic element variation with the most available boolean functions uses 90% of the transistors, with a penalty in delay of 102%, and improved dynamic and static power of 97% and 91%, respectively. The smallest variation uses 42% of the transistors, with improved delay of 76%, and improved dynamic and static power of 43% and 43%, respectively. Ashton Snelgrove, Pierre-Emmanuel Gaillardon |
DDECS | 2 |
| 2022 | An Open-source Three-Independent-Gate FET Standard Cell Library for Mixed Logic SynthesisabstractThree-Independent-Gate FET (TIGFET) technology is one of the most promising candidates to succeed CMOS and FinFET technologies due to its low off-current, compact surface area, reconfigurable logic, and CMOS compatibility. In this paper, we present an open-source standard cell library based on silicon nano-wire TIGFETs, which enables efficient implementation of novel XOR-and-majority-based circuit designs. We also discuss logic synthesis methods tailored to take advantage of TIGFET capabilities to allow their potential to be realized at the system level. By combining the 10nm TIGFET technology with a mixed logic synthesis tool, the PicoRV core design shows a 2.3× lower area and a 5.7× lower energy consumption, compared to an equivalent low-power 12nm FinFET implementation. Roman Gauchi, Ashton Snelgrove, Pierre-Emmanuel Gaillardon |
ISCAS | 3 |
| 2022 | An Energy-Efficient Three-Independent-Gate FET Cell Library for Low-Power Edge ComputingabstractWith the increasing demand for compute-intensive applications for IoT devices, new technologies that enable a power reduction at the device-level are needed to improve energy savings at the system-level. Unfortunately, the scaling of standard CMOS technologies is not as fast as the scaling of computing performance, which leads to the so-called "power wall". The Three-Independent-Gate Field-Effect Transistor (TIGFET) is a promising technology that enhances the device functionality to create more compact logic gates and provide silicon-nanowire structures that meet the requirements of low leakage power systems. However, the evaluation of complex designs is currently limited to the intrinsic model of the transistor and does not consider the parasitic effects of cell layouts. In this paper, we propose a standard cell library for 10-nm silicon-nanowire TIGFET devices, including combinational and sequential gates, to evaluate a production RISC-V core targeting low energy consumption budget. After synthesis, the core achieves 4 × lower energy consumption up to a frequency of 340 MHz compared to an equivalent low-power 12-nm FinFET technology node. Michael Keyser, Roman Gauchi, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 3 |
| 2022 | A Two-Level Approximate Logic Synthesis Combining Cube Insertion and RemovalabstractApproximate computing is an attractive paradigm for reducing the design complexity of error-resilient systems, therefore, improving performance and saving power consumption. In this work, we propose a new two-level approximate logic synthesis method based on cube insertion and removal procedures. The experimental results have shown significant literal count and runtime reduction compared to the state-of-the-art approach. The method scalability is illustrated for a high error threshold over large benchmark circuits. The obtained solutions have presented a literal number reduction up to 38%, 56%, and 93% with respect to an error rate of 1%, 3%, and 5%, respectively. Gabriel Ammes, Walter Lau Neto, Paulo F. Butzen, Pierre-Emmanuel Gaillardon, Renato P. Ribas |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | NEMO-CNN: An Efficient Near-Memory Accelerator for Convolutional Neural NetworksabstractThe relevance of Deep Learning applications has skyrocketed in the last few years, exposing key weaknesses of traditional Von Neumann hardware architectures. With high amounts of data to be fetched from memory, the efficiency of these systems gets adversely impacted by an order of magnitude for each memory hierarchy that is traversed (e.g., data cache, on- chip SRAM, and off-chip DRAM). Such an issue is even more relevant when we consider that Convolutional Neural Networks (CNNs) are composed of tens of millions of parameters that imply billions of operations per second to achieve an acceptable performance. In order to remove this so-called memory wall problem, we introduce NEMO-CNN: a high-performance hardware accelerator built around the Near-Memory Computing paradigm, i.e., a design methodology based on distributed memory blocks enhanced with nearby processing elements. Coupled with a smart mapping strategy that slices the CNN structure along its depth, our solution drastically reduces the amount of data exchanged between off- and on-chip memories by executing each slice concurrently on dedicated processing elements that only leverage local data. Experimental results using VGG-16, DarkNet-19, and TinyYOLOv2 networks demonstrate that our solution achieves a top efficiency of 60.7 FPS/W, outperforming existing CNN accelerators. Grant Brown, Valerio Tenace, Pierre-Emmanuel Gaillardon |
ASAP | 3 |
| 2021 | Read your Circuit: Leveraging Word Embedding to Guide Logic OptimizationabstractTo tackle the involved complexity, Electronic Design Automation (EDA) tools are broken in well-defined steps, each operating at different abstraction levels. Higher levels of abstraction shorten the flow run-time while sacrificing correlation with the physical circuit implementation. Bridging this gap between Logic Synthesis tool and Physical Design (PnR) tools is key to improve Quality of Results (QoR), while possibly shorting the time-to-market. To address this problem, in this work, we formalize logic paths as sentences, with the gates being a bag of words. Thus, we show how word embedding can be leveraged to represent generic paths and predict if a given path is likely to be critical post-PnR. We present the effectiveness of our approach, with accuracy over than 90% for our test-cases. Finally, we give a step further and introduce an intelligent and non-intrusive flow that uses this information to guide optimization. Our flow presents up to 15.53% area delay product (ADP) and 18.56% power delay product (PDP), compared to a standard flow. Walter Lau Neto, Matheus T. Moreira, Luca G. Amarù, Cunxi Yu, Pierre-Emmanuel Gaillardon |
ASP-DAC | 5 |
| 2021 | SLAP: A Supervised Learning Approach for Priority Cuts Technology MappingabstractRecently we have seen many works that leverage Machine Learning (ML) techniques in optimizing Electronic Design Automation (EDA) process. However, the uses of ML techniques are limited to learning forecasting models of existing EDA algorithms, instead of developing novel algorithms. In this work, we focus on designing an novel cut-based technology mapping algorithms assisted by ML techniques, which matches results of exhaustive cut exploration but preserving a small footprint of utilized cuts. The proposed approach has been demonstrated with a wide range of benchmarks with 24% reductions in number of cuts utilized compared to the state-of-the-art, while improving the circuit delay, and Area-Delay-Product (ADP), by average about 10%, 7%, respectively, with a 2% area penalty. Compared to the exhaustive approach, i.e., considering all the cuts, we achieve similar or better results while saving over than $2 \times $ the number of considered cuts (runtime) on average. Finally, we provide a comprehensive explanation of heuristics learned by the ML model by feature ranking. Walter Lau Neto, Matheus T. Moreira, Luca G. Amarù, Cunxi Yu, Pierre-Emmanuel Gaillardon |
DAC | 6 |
| 2021 | Invited: Getting the Most out of your Circuits with Heterogeneous Logic SynthesisabstractHigh Level Synthesis (HLS) speeds hardware development and opens the door to non-expert designers, focusing on functionality rather than implementation. The expense and rigidity of commercial electronic design automation (EDA) tool-chains can be an obstacle for these users. LSOracle is an opensource logic synthesis tool which leverages multiple underlying data structures, including and-inverter graphs (AIGs), majority-inverter graphs (MIGs), and xor-and graphs (XAGs) to automatically optimize circuits using the best representation for each region of a design, without manual intervention. The use of MIGs and XAGs gives particularly strong performance in arithmetic logic, cryptography cores, and machine-learning accelerators; applications which may be of particular interest for HLS users. Here we present an overview of the approach and demonstrate an open-source HLS-to-GDS II workflow using LSOracle, Bambu, and OpenROAD. We test the integration on a small benchmark suite and show a reduction in delay of up to 31%. Scott Temple, Walter Lau Neto, Ashton Snelgrove, Xifan Tang, Pierre-Emmanuel Gaillardon |
DAC | 5 |
| 2021 | A Deep Learning Approach to Sensor Fusion Inference at the EdgeabstractThe advent of large scale urban sensor networks has enabled a paradigm shift of how we collect and interpret data. By equipping these sensor nodes with emerging low-power hardware accelerators, they become powerful edge devices, capable of locally inferring latent features and trends from their fused multivariate data. Unfortunately, traditional inference techniques are not well suited for operation in edge devices, or simply fail to capture many statistical aspects of these low-cost sensors. As a result, these methods struggle to accurately model nonlinear events. In this work, we propose a deep learning methodology that is able to infer unseen data by learning complex trends and the distribution of the fused time-series inputs. This novel hybrid architecture combines a multivariate Long Short-Term Memory (LSTM) branch and two convolutional branches to extract time-series trends as well as short-term features. By normalizing each input vector, we are able to magnify features and better distinguish trends between series. As a demonstration of the broad applicability of this technique, we use data from a currently deployed pollution monitoring network of low-cost sensors to infer hourly ozone concentrations at the device level. Results indicate that our technique greatly outperforms traditional linear regression techniques by 6 × as well as state-of-the-art multivariate time-series techniques by 1.4 × in mean squared error. Remarkably, we also show that inferred quantities can achieve lower variability than the primary sensors which produce the input data. Thomas Becnel, Pierre-Emmanuel Gaillardon |
DATE | 2 |
| 2021 | Logic Synthesis Meets Machine Learning: Trading Exactness for GeneralizationabstractLogic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning. Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee |
DATE | 39 |
| 2021 | Taping out an FPGA in 24 hours with OpenFPGA: The SOFA ProjectabstractThis paper highlights the Skywater Open-source embedded FpgAs (SOFA) project, which is a series of open-source embedded FPGA IPs built with the Skywater 130nm technology. The SOFA project showcases an agile prototyping methodology for FPGAs, enabled by the OpenFPGA framework, whose fabrication-ready layouts are generated in 24 hours. We also present the associated Verilog-to-Bitstream toolchain for end users. Xifan Tang, Ganesh Gore, Grant Brown, Pierre-Emmanuel Gaillardon |
FPL | 4 |
| 2021 | Exploring eFPGA-based Redaction for IP ProtectionabstractRecently, eFPGA-based redaction has been proposed as a promising solution for hiding parts of a digital design from untrusted entities, where legitimate end-users can restore functionality by loading the withheld bitstream after fabrication. However, when deciding which parts of a design to redact, there are a number of practical issues that designers need to consider, including area and timing overheads, as well as security factors. Adapting an open-source FPGA fabric generation flow, we perform a case study to explore the trade-offs when redacting different modules of open-source intellectual property blocks (IPs) and explore how different parts of an eFPGA contribute to the security. We provide new insights into the feasibility and challenges of using eFPGA-based redaction as a security solution. Jitendra Bhandari, Abdul Khader Thalakkattu Moosa, Benjamin Tan 0001, Christian Pilato, Ganesh Gore, Xifan Tang, Scott Temple, Pierre-Emmanuel Gaillardon, Ramesh Karri |
ICCAD | 8 |
| 2021 | Smart-Redundancy: An Alternative SEU/SET Mitigation Method for FPGAsabstractField Programmable Gate Arrays (FPGAs) reconfigurability is a key asset for many critical applications. State- of-the-art Radiation-Hardening (Rad-Hard) methods for FPGAs consist of triplicating the logic, reinforcing the memories, and bitstream scrubbing with partial reconfiguration. These methods involve a 3× reduction of Maximal Design Capacity (MDC) and an average Time-In-Error (TIE) proportional to design sizes. In this paper, we propose an alternative: Smart-Redundancy (SR), a new method based on the detection of possible events via process and hardware modifications. Thanks to integrated particle sensors, only dual redundancy is required. Results show up to 33.33% improvement in MDC over actual Rad-Hard methods, and an average TIE decrease of at least 10,000× compared to bitstream's scrubbing, at a cost of 41.08% in area using a commercial 40nm technology node. Aurélien Alacchi, Edouard Giacomin, Xifan Tang, Pierre-Emmanuel Gaillardon |
ISCAS | 4 |
| 2021 | Area-Efficient Multiplier Designs Using a 3D Nanofabric Process FlowabstractIn the past few years, the demand for computationally intensive applications, such as digital signal processing or convolutional neural networks, has grown exponentially. As they often rely on a significant number of multiply-and-accumulate cells, it is crucial to optimize their area and cost. Recently, a 3D Nanofabric flow has been proposed, where logic circuits are designed by stacking N identical vertical tiers on top of each other. Exploiting identical layers allows a fabrication process similar to the Vertical-NAND flash, where all the layers can be patterned at once. While the 3D Nanofabric flow presents several layout constraints (single metal routing and identical vertical layers), it can decrease the area by around one order of magnitude, leading to area-efficient and cost-effective circuits. In this paper, we propose to use the 3D Nanofabric process flow to design low-area multipliers. As multipliers can be designed using a regular array organization, we show how they can be spread across multiple vertical layers using the 3D Nanofabric flow, while respecting the different layout constraints. We then provide thorough circuit-level evaluations, including parasitics, to showcase the benefits of our proposed 3D multipliers at the circuit-level. We show that by stacking up to 64 layers to build a 64-input bit multiplier, the area and area-delay- product can be decreased by 28.6x and 25.5x, respectively, compared to a traditional 2D implementation using a 28nm FDSOI technology, with only a 10% and 35% delay and power consumption overheads, respectively. Edouard Giacomin, Francky Catthoor, Pierre-Emmanuel Gaillardon |
ISCAS | 3 |
| 2021 | A Scalable and Robust Hierarchical Floorplanning to Enable 24-hour Prototyping for 100k-LUT FPGAsabstractPhysical design for Field Programmable Gate Array (FPGA) is challenging and time-consuming, primarily due to the use of a full-custom approach for aggressively optimize Performance, Power and Area (P.P.A.) of the FPGA design. The growing number of FPGA applications demands novel architectures and shorter development cycles. The use of an automated toolchain is essential to reduce end-to-end development time. This paper presents scalable and adaptive hierarchical floorplanning strategies to significantly reduce the physical design runtime and enable millions-of-LUT FPGA layout implementations using standard ASIC toolchains. This approach mainly exploits the regularity of the design and performs necessary feedthrough creations for global and clock nets to eliminate any requirement of global optimizations. To validate this approach, we implemented full-chip layouts for modern FPGA fabric with logic capacity ranging from 40 to 100k LUTs using a commercial 12nm technology. Our results show that the physical implementation of a 128k-LUT FPGA fabric can be achieved within 24-hours, which has not been demonstrated by any previous work. Compared to previous work, the runtime reduction of 8x is obtained for implementing 2.5k LUTs FPGA device. Ganesh Gore, Xifan Tang, Pierre-Emmanuel Gaillardon |
ISPD | 3 |
| 2021 | A Novel High-Gain Amplifier Circuit Using Super-Steep-Subthreshold-Slope Field-Effect TransistorsabstractThe benefits of steep-Subthreshold Swing (SS) devices, though plentiful at the device-level, have yet to be fully exploited at the circuit-level. This is evident from a look at the Three-Independent-Gate Field-Effect Transistor (TIGFET), a device renown for its ability for polarity reconfiguration. At the same time, its demonstrated dynamic control of the subthreshold slope beyond the thermal limit has only been studied at the device-level. This latter benefit is referred to as Super-Steep Subthreshold Slope (S4) operation and can lead to unprecedented gain, which is ideal for use in an amplifier circuit. In this paper, we investigate the impact of S4 operations when designing differential-amplifier circuits when using TIGFET technology. We demonstrate the benefits of our implementation both from a theoretical standpoint and through circuit-level analyses. More specifically, we show that the TIGFET -based amplifier gain is 95.5 $\times$ better, and that the gain-bandwidth product is improved by 13.8, compared to an equivalent MOSFET-based design at the 90 nm node. Besides, we show that at equivalent gains, the TIGFET-based amplifier decreases the area and power by 22.8 $\times$ and 7.2, respectively, against its MOSFET counterpart. Matthieu Couriol, Patsy Cadareanu, Edouard Giacomin, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 4 |
| 2021 | A 12-pA Resolution Sigma Delta ADC Topology for Chemiresistive Sensor-Based ApplicationsabstractChemiresistive sensor technologies can detect chemical trace in the air down to the low Part Per Billion (ppb)/Part Per Trillion (ppt) level. Achieving such low detection levels requires low-noise and high-accuracy analog front-end interfaces. One such interface uses a Sigma Delta $(\Sigma\triangle)$ Analog to Digital Converter (ADC) topology that leads to lower noise at low frequency and better effective resolution. The slow charging dynamics of chemicals in the environment is well-suited for $\Sigma\triangle$ converters. While $\Sigma\triangle$ converters are a well-studied topic, the potential resolution of chemical sensing using chemiresistive sensors has not yet been demonstrated in a fully integrated mixed-signal interface. In this paper, we develop a $\Sigma\triangle$ architecture combined with a 10-pA resolution Digitally Controlled Current Sources (DCCS) and an integrated Proportional Integrator Derivative (PID) controller feedback loop. The PID drives back the current sources at 250 kHz to maintain a constant voltage across the sensor, acting as a variable resistor Implemented on a 180 nm CMOS process with an 8-bits current source resolution at 250 kHz and coupled with nanofiber chemiresistive sensors, it translates into lppb level detection. Compared to the state-of-the-art chemical sensing interface, our circuit shows a 5$\times$improvement in detection level and a $ 10\times$ improvement in power efficiency, given the same detection level. Matthieu Couriol, Edouard Giacomin, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 3 |
| 2021 | multiPULPly: A Multiplication Engine for Accelerating Neural Networks on Ultra-low-power ArchitecturesabstractComputationally intensive neural network applications often need to run on resource-limited low-power devices. Numerous hardware accelerators have been developed to speed up the performance of neural network applications and reduce power consumption; however, most focus on data centers and full-fledged systems. Acceleration in ultra-low-power systems has been only partially addressed. In this article, we present multiPULPly, an accelerator that integrates memristive technologies within standard low-power CMOS technology, to accelerate multiplication in neural network inference on ultra-low-power systems. This accelerator was designated for PULP, an open-source microcontroller system that uses low-power RISC-V processors. Memristors were integrated into the accelerator to enable power consumption only when the memory is active, to continue the task with no context-restoring overhead, and to enable highly parallel analog multiplication. To reduce the energy consumption, we propose novel dataflows that handle common multiplication scenarios and are tailored for our architecture. The accelerator was tested on FPGA and achieved a peak energy efficiency of 19.5 TOPS/W, outperforming state-of-the-art accelerators by 1.5× to 4.5×. Adi Eliahu, Ronny Ronen, Pierre-Emmanuel Gaillardon, Shahar Kvatinsky |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | A Scalable Mixed Synthesis Framework for Heterogeneous NetworksabstractWe present a new logic synthesis framework which produces efficient post-technology mapped results on heterogeneous networks containing a mix of different types of logic. This framework accomplishes this by breaking down the circuit into sections using a hypergraph k-way partitioner and then determines the best-fit logic representation for each partition between two Boolean networks, And-Inverter Graphs (AIG) and Majority-Inverter Graphs (MIG), which have been shown to perform better over each other on different types of logic. Experimental results show that over a set of Open Piton Design Benchmarks (OPDB) and OpenCores benchmarks, our proposed methodology outperforms state-of-the-art academic tools in Area-Delay Product (ADP), Power-Delay Product (PDP), and Energy-Delay Product (EDP) by 5%, 2%, and 15% respectively after performing Application Specific Integrated Circuits (ASIC) technology mapping as well as showing a 54% improvement in runtime over conventional MIG optimization. Max Austin, Scott Temple, Walter Lau Neto, Luca G. Amarù, Xifan Tang, Pierre-Emmanuel Gaillardon |
DATE | 6 |
| 2020 | A Novel TIGFET-based DFF Design for Improved Resilience to Power Side-Channel AttacksabstractSide-channel attacks (SCAs) represent a significant security threat, and aim to reveal otherwise secret data by analyzing a relevant circuit's behavior, e.g., its power consumption. While all circuit components are potential power side channels, D-flip-flops (DFFs) are often the primary source of information leakage to an SCA. This paper proposes a DFF design based on the three-independent-gate field-effect transistor (TTGFET) that reduces side-channel vulnerabilities of sequential circuits. Notably, we find that the I-V characteristics of the TIGFET itself leads to inherent side-channel resilience, which in turn enables simpler and more efficient cryptographic hardware. Our proposed design is based on a prior TIGFET-based true single-phase clock (TSPC) DFF design, which offers high performance and reduced area. More specifically, our modified TSPC (mTSPC) design exploits the symmetric I-V characteristics of TIGFETs, which results in pull-up and pull-down currents that are nearly identical. When combined with additional circuit modifications (made possible by the unique characteristics of the TIGFET), the mTSPC circuit draws almost the same amount of supply currents under all possible input transitions (less than 1% variation for different transitions), which can in turn mask information leakage. Using a 10nm TIGFET technology model, simulation results show that the proposed TIGFET-based DFF circuit leads to decreased power consumption (up to 96.9% when compared to the prior secured designs), has a low delay (15.2 ps), and employs only 12 TIGFET devices. Furthermore, an 8-bit S-box whose output is sampled by a group of eight mTSPC DFFs was simulated. A correlation power analysis attack on the simulated S-box with 256 power traces shows that the key is not revealed, which confirms the SCA resiliency of the proposed DFF design. Mohammad Mehdi Sharifi, Ramin Rajaei, Patsy Cadareanu, Pierre-Emmanuel Gaillardon, Yier Jin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 4 |
| 2020 | A RRAM-based FPGA for Energy-efficient Edge ComputingabstractThe shift from centralized cloud to edge computing demands hardware systems with data processing capability at ultra-low power. Reconfigurable solutions such as Field-Programmable Gate Arrays (FPGAs) offer a high flexibility in terms of hardware implementation and are thus popular for use in many edge computing systems. However, breaking through the energy wall of FPGAs is a challenge, as low-power operation often requires compromising performances. In this paper, we study a low-power high-performance FPGA architecture exploiting Resistive Random Access Memory (RRAM) technology. To perform a comprehensive analysis, we introduce a novel design flow which can rapidly prototype FPGA fabrics from which accurate area, delay, and power results can be obtained. Based on full-chip layouts and SPICE simulations, we show that RRAM-based FPGAs can improve up to 8%/22%/16% in area/delay/power compared to SRAM-based counterparts at nominal voltage. Even when operated at a near-Vtsupply, the proposed RRAM-based FPGA can improve the Energy-Delay Product by about 2× without any delay overhead, when compared to an SRAM-based FPGA. In addition, Monte Carlo simulations showed that the proposed RRAM-based FPGA architecture stays robust under different CMOS process corners as well as under a 30% RRAM resistance standard deviation. Xifan Tang, Edouard Giacomin, Patsy Cadareanu, Ganesh Gore, Pierre-Emmanuel Gaillardon |
DATE | 5 |
| 2020 | SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are expected to be part of the future AI portfolio, with heavy investment from industry and government, e.g., IBM TrueNorth, Intel Loihi. While Artificial Neural Network (ANN) architectures have taken large strides, few works have targeted SNN hardware efficiency. Our analysis of SNN baselines shows that at modest spike rates, SNN implementations exhibit significantly lower efficiency than accelerators for ANNs. This is primarily because SNN dataflows must consider neuron potentials for several ticks, introducing a new data structure and a new dimension to the reuse pattern. We introduce a novel SNN architecture, SpinalFlow, that processes a compressed, time-stamped, sorted sequence of input spikes. It adopts an ordering of computations such that the outputs of a network layer are also compressed, time-stamped, and sorted. All relevant computations for a neuron are performed in consecutive steps to eliminate neuron potential storage overheads. Thus, with better data reuse, we advance the energy efficiency of SNN accelerators by an order of magnitude. Even though the temporal aspect in SNNs prevents the exploitation of some reuse patterns that are more easily exploited in ANNs, at 4-bit input resolution and 90% input sparsity, SpinalFlow reduces average energy by 1.8×, compared to a 4-bit Eyeriss baseline. These improvements are seen for a range of networks and sparsity/resolution levels; SpinalFlow consumes 5× less energy and 5.4× less time than an 8-bit version of Eyeriss. We thus show that, depending on the level of observed sparsity, SNN architectures can be competitive with ANN architectures in terms of latency and energy for inference, thus lowering the barrier for practical deployment in scenarios demanding real-time learning. Surya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin, Pierre-Emmanuel Gaillardon |
ISCA | 5 |
| 2020 | Layout Considerations of Logic Designs Using an N-layer 3D Nanofabric Process FlowabstractIn the past few years, novel fabrication schemes such as parallel and monolithic 3D integration have been proposed to keep sustaining the need for more powerful integrated circuits. By stacking several devices, wafers, or dies, the footprint, delay, and power can be decreased when compared to traditional 2D implementations. While parallel 3D does not enable very fine-grained vertical connections, monolithic 3D currently only offers a limited number of transistor tiers due to the high cost of the additional masks and processing steps, limiting the benefits of using the third dimension. In this paper, we introduce an innovative planar circuit netlist and layout approach, which enables a new 3D integration flow called 3D Nanofabric. The flow, consisting of$N$identical vertical tiers, is aimed at single instruction multiple data processor Arithmetic Logic Units (ALUs). By using a single metal routing layer for each vertical tier, the process flow is significantly simplified since multiple vertical layers can potentially be patterned at once, similar to the 3D NAND flash process. In our study, we thoroughly investigate the layout constraints arising from the Nanofabric flow and the unique metal layer rule and propose several ways to overcome them. We then show that by stacking 32 layers to build a 32-bit ALU, the footprint is reduced by$8.7\times$when compared to a conventional 7nm FinFET implementation. Edouard Giacomin, Jürgen Bömmels, Julien Ryckaert, Francky Catthoor, Pierre-Emmanuel Gaillardon |
VLSI-SOC | 5 |
| 2019 | Rebooting Our Computing ModelsabstractInnovative and new computing paradigms must be considered as we reach the limits of von Neumann computing caused by the growth in necessary data processing. This paper provides an introduction to three emerging computing models that have established themselves as likely post-CMOS and post-von Neumann solutions. The first of these ideas is quantum computing, for which we discuss the challenges and potential of quantum computer architectures. Next, a computational system using intrinsic oscillators is introduced and an example is provided which shows its superiority in comparison to a typical von Neumann computational system. Finally, digital memcomputing using self-organizing logic gates is explained and then discussed as a method for optimization problems and machine learning. Patsy Cadareanu, N. Reddy C, Carmen G. Almudéver, A. Khanna, Arijit Raychowdhury, Suman Datta, Koen Bertels, Vijayakrishan Narayanan, Massimiliano Di Ventra, Pierre-Emmanuel Gaillardon |
DATE | 10 |
| 2019 | Scalable Boolean Methods in a Modern Synthesis FlowabstractWith the continuous push to improve Quality of Results (QoR) in EDA, Boolean methods in logic synthesis have been recently drawing the attention of researchers. Boolean methods achieve better QoR than algebraic methods but require higher computational cost. In this paper, we introduce the Scalable Boolean Method (SBM) framework. The SBM consists of 4 optimization engines designed to be scalable in a modern synthesis flow. The first presented engine is a generalized resubstitution framework based on computing, and implementing, the Boolean difference between two nodes. The second consists of a gradient-based AIG optimization, while the third one is based on heterogeneous elimination for kerneling. The last proposed engine is a revisiting of maximum set of permissible functions computation with BDDs. Altogether, the SBM framework enables significant synthesis results. We improve 12 of the best known area results in the EPFL synthesis competition. Embedded in a commercial EDA flow, the new Boolean methods enable -2.20% combinational area savings and -5.99% total negative slack reduction, after physical implementation, at contained runtime cost. Eleonora Testa, Luca G. Amarù, Mathias Soeken, Alan Mishchenko, Patrick Vuillod, Jiong Luo, Christopher Casares, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DATE | 8 |
| 2019 | OpenFPGA: An Opensource Framework Enabling Rapid Prototyping of Customizable FPGAsabstractDriven by the strong need in data processing applications, Field Programmable Gate Arrays (FPGAs) are playing an ever-increasing role as programmable accelerators in modern computing systems. To fully unlock processing capabilities for domain-specific applications, FPGA architectures have to be tailored for seamless cooperation with other computing resources. However, prototyping and bringing to production a customized FPGA is a costly and complex endeavor even for industrial vendors. In this paper, we introduce OpenFPGA, an opensource framework that enables rapid prototyping of customizable FPGA architectures through a semi-custom design approach. We propose an XML-to-Prototype design flow, where the Verilog netlists of a full FPGA fabric can be autogenerated using an extension of the XML language from the VTR framework and then fed into a back-end flow to generate production-ready layouts. OpenFPGA also includes a general-purpose Verilog-to-Bitstream generator for any FPGA described by the XML language. We demonstrate the capability of this automatic design flow with a Stratix IV-like FPGA architecture using a commercial 40nm technology node, and perform a detailed comparison to its academic and commercial counterparts. Compared to the current state-of-art academic results, our FPGA fabric reduces the area by 1:75 and the delay by 3 on average. In addition, OpenFPGA significantly reduces the gap between semi-custom designed FPGAs and fully-optimized commercial products with a penalty of only 60% in area and 30% in delay, respectively. Xifan Tang, Edouard Giacomin, Aurélien Alacchi, Baudouin Chauviere, Pierre-Emmanuel Gaillardon |
FPL | 5 |
| 2019 | LSOracle: a Logic Synthesis Framework Driven by Artificial Intelligence: Invited PaperabstractThe increasing complexity of modern Integrated Circuits (ICs) leads to systems composed of various different Intellectual Property (IPs) blocks, known as System-on-Chip (SoC). Such complexity requires strong expertise from engineers, that rely on expansive commercial EDA tools. To overcome such a limitation, an automated open-source logic synthesis flow is required. In this context, this work proposes LSOracle: a novel automated mixed logic synthesis framework. LSOracle is the first to exploit state-of-the-art And-Inverter Graph (AIG) and Majority-Inverter Graph (MIG) logic optimizers and relies on a Deep Neural Network (DNN) to automatically decide which optimizer should handle different portions of the circuit. To do so, LSOracle applies k-way partitioning to split a DAG into multiple partitions and uses a to chose the best-fit optimizer. Post-tech mapping ASIC results, targeting the 7nm ASAP standard cell library, for a set of mixed-logic circuits, show an average improvement in area-delay product of 6.87% (up to 10.26%) and 2.70% (up to 6.27%) when compared to AIG and MIG, respectively. In addition, we show that for the considered circuits, LSOracle achieves an area close to AIGs (which delivered smaller circuits) with a similar performance of MIGs, which delivered faster circuits. Walter Lau Neto, Max Austin, Scott Temple, Luca G. Amarù, Xifan Tang, Pierre-Emmanuel Gaillardon |
ICCAD | 6 |
| 2019 | Wire-Aware Architecture and Dataflow for CNN AcceleratorsabstractIn spite of several recent advancements, data movement in modern CNN accelerators remains a significant bottleneck. Architectures like Eyeriss implement large scratchpads within individual processing elements, while architectures like TPU v1 implement large systolic arrays and large monolithic caches. Several data movements in these prior works are therefore across long wires, and account for much of the energy consumption. In this work, we design a new wire-aware CNN accelerator, WAX, that employs a deep and distributed memory hierarchy, thus enabling data movement over short wires in the common case. An array of computational units, each with a small set of registers, is placed adjacent to a subarray of a large cache to form a single tile. Shift operations among these registers allow for high reuse with little wire traversal overhead. This approach optimizes the common case, where register fetches and access to a few-kilobyte buffer can be performed at very low cost. Operations beyond the tile require traversal over the cache's H-tree interconnect, but represent the uncommon case. For high reuse of operands, we introduce a family of new data mappings and dataflows. The best dataflow, WAXFlow-3, achieves a 2× improvement in performance and a 2.6-4.4× reduction in energy, relative to Eyeriss. As more WAX tiles are added, performance scales well until 128 tiles. Sumanth Gudaparthi, Surya Narayanan, Rajeev Balasubramonian, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon |
MICRO | 6 |
| 2019 | GenCache: Leveraging In-Cache Operators for Efficient Sequence AlignmentabstractPrecision Medicine will rely on frequent genomic analysis, especially for patients undergoing cancer treatments or suffering from rare diseases. Sequence alignment is invoked in multiple stages of the genomic analysis pipeline. Recent projects have introduced accelerators, GenAx and Darwin, for 2nd and 3rd generation sequencers respectively. In this work, we improve upon the GenAx design by increasing its parallelism and reducing its memory bandwidth demands. This is achieved with a combination of hardware and software innovations. We first integrate in-cache operators from prior work into the GenAx memory hierarchy; we then augment the in-cache peripheral circuit to support additional new operators. We then re-structure the sequence alignment algorithm to (i) leverage the many in-cache operators, (ii) exploit the common case in genomic datasets, (iii) use Bloom Filters to reduce futile accesses, and (iv) maximize data reuse within a re-organized memory hierarchy. While the baseline GenAx accelerator processes a batch of reads in 194 seconds while nearly saturating the 153.6 GB/s memory bandwidth, the proposed GenCache architecture processes the same batch of reads in 37 seconds at an improved energy efficiency of 8.6×, while demanding 20 GB/s average memory bandwidth. Our hardware and software techniques thus interact synergistically to target both memory and compute bottlenecks, while not affecting the outputs of the application. We show that the basic principles in GenCache can also be exploited by 3rd generation sequence aligners. Anirban Nag, C. N. Ramachandra, Rajeev Balasubramonian, Ryan Stutsman, Edouard Giacomin, Hari Kambalasubramanyam, Pierre-Emmanuel Gaillardon |
MICRO | 7 |
| 2019 | A Predictive Process Design Kit for Three-Independent-Gate Field-Effect TransistorsabstractThe Three-Independent-Gate Field-Effect Transistor (TIGFET) is a promising beyond-CMOS technology which offers many unique properties, such as (i) dynamic control of the device polarity, (ii) dual threshold operation and (iii) more expressive logic capabilities. The efficient exploitation of these properties provides opportunity to design area and power optimized logic circuits. However, the evaluation of TIGFET-based design currently relies on a close approximation for the Power, Performance, and Area (PPA) rather than traditional layout-based methods. There is a need for a publicly available Process Design Kit (PDK) enabling systematic evaluation of the design area. In this paper, we propose Predictive PDK for the 10 nm-diameter silicon-nanowire TIGFET device. This work consists of a SPICE model and full custom physical design files including a Design Rule Manual, a Design Rule Check, and a Layout Versus Schematic decks for Calibre®. We then validate the design rules through the implementation of basic logic gates and a full-adder and compare extracted metrics with FreePDK15nm™ PDK. We show 26% and 41% area reduction in the case of an XOR gate and a 1-bit full-adder design respectively. Ganesh Gore, Patsy Cadareanu, Edouard Giacomin, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 4 |
| 2019 | A Product Engine for Energy-Efficient Execution of Binary Neural Networks Using Resistive MemoriesabstractThe need for running complex Machine Learning (ML) algorithms, such as Convolutional Neural Networks (CNNs), in edge devices, which are highly constrained in terms of computing power and energy, makes it important to execute such applications efficiently. The situation has led to the popularization of Binary Neural Networks (BNNs), which significantly reduce execution time and memory requirements by representing the weights (and possibly the data being operated) using only one bit. Because approximately 90% of the operations executed by CNNs and BNNs are convolutions, a significant part of the memory transfers consists of fetching the convolutional kernels. Such kernels are usually small (e.g., 3×3 operands), and particularly in BNNs redundancy is expected. Therefore, equal kernels can be mapped to the same memory addresses, requiring significantly less memory to store them. In this context, this paper presents a custom Binary Dot Product Engine (BDPE) for BNNs that exploits the features of Resistive Random-Access Memories (RRAMs). This new engine allows accelerating the execution of the inference phase of BNNs. The novel BDPE locally stores the most used binary weights and performs binary convolution using computing capabilities enabled by the RRAMs. The system-level gem5 architectural simulator was used together with a C-based ML framework to evaluate the system's performance and obtain power results. Results show that this novel BDPE improves performance by 11.3%, energy efficiency by 7.4% and reduces the number of memory accesses by 10.7% at a cost of less than 0.3% additional die area, when integrated with a 28 nm Fully Depleted Silicon On Insulator ARMv8 in-order core, in comparison to a fully-optimized baseline of YoloV3 XNOR-Net running in a unmodified Central Processing Unit. João Vieira, Edouard Giacomin, Yasir Mahmood Qureshi, Marina Zapater, Xifan Tang, Shahar Kvatinsky, David Atienza 0001, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 8 |
| 2019 | A Distributed Low-Cost Pollution Monitoring PlatformabstractPersonal exposure to heightened levels of fine airborne particulate matter (PM) has been linked to numerous adverse health effects in sensitive groups. However, researchers investigating these correlations are struggling to find the spatiotemporal datasets that are sufficient for study. Current airborne PM monitoring solutions are highly accurate, but expensive. Therefore, they are not feasible candidates for spatially dense deployments, and cannot be used to analyze the effects of exposure to pollution microclimates. In this article, we present a low-cost pollution monitoring station that operates as a single node in a wireless network. Each node periodically collects airborne pollution and supporting meteorological data and uploads measurements to a central, open-source database. A total of 50 nodes were deployed across a large metropolitan area (roughly 100 km2) over a six-month campaign. The experimental results show good correlation (R2= 0.88) between devices co-located with the federal equivalent methods, which have an accuracy traceable to the National Institute of Standards and Technology. By applying linear corrections derived from in situ field measurements to each PM sensor, we were able to demonstrate a 1.8× decrease in root-mean-squared error over the raw measurements. Thomas Becnel, Kyle Tingey, Jonathan Whitaker, Tofigh Sayahi, Katrina Lê, Pascal Goffin, Anthony Butterfield, Kerry E. Kelly, Pierre-Emmanuel Gaillardon |
IEEE Internet Things J. | 9 |
| 2019 | Devices and Circuits Using Novel 2-D Materials: A Perspective for Future VLSI SystemsabstractHere, we review the most recent developments in the field of 2-D electronics. We focus first on the synthesis of 2-D materials, discussing the different growth techniques currently available and assessing their strengths and weaknesses. Moreover, we describe a possible roadmap to enable CMOS compatible integration of 2-D materials. We then shift our attention to 2-D devices and circuits and review the state of the art. Among the plethora of device concepts, we look closely at 2-D tunnel FETs (TFETs) and negative-capacitance FETs (NC-FETs) for low-power applications. We also put a particular emphasis on doping-free polarity-controllable systems that use electrostatic doping to eliminate the need for physical or chemical doping. We conclude with an analysis of simulations of scaled devices and discuss the possibilities enabled at circuit level by 2-D electronics. Giovanni V. Resta, Alessandra Leonhardt, Yashwanth Balaji, Stefan De Gendt, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | FPGA-SPICE: A Simulation-Based Architecture Evaluation Framework for FPGAsabstractIn this paper, we developed a simulation-based architecture evaluation framework for field-programmable gate arrays (FPGAs), called FPGA-SPICE, which enables automatic layout-level estimation and electrical simulations of FPGA architectures. FPGA-SPICE can automatically generate Verilog and SPICE netlists based on realistic FPGA configurations and a high-level eTtensible Markup Language-based FPGA architectural description language. The outputted Verilog netlists can be used to generate layouts of full FPGA fabrics through a semicustom design flow. SPICE simulation decks can be generated at three levels of complexity, namely, full-chip-level, grid-level, and component-level, providing different tradeoff between accuracy and simulation time. In order to enable such level of analysis, we presented two SPICE netlist partitioning techniques: loads extraction and parasitic net activity estimation. Electrical simulations showed that averaged over the selected benchmarks, the grid-/component-level approach can achieve 6.1×/7.5× execution speed-up with 9.9%/8.3% accuracy loss, respectively, compared to the full-chip level simulation. FPGA-SPICE was showcased through three different case studies: (1) an area breakdown analysis for static random access memory-based FPGAs, showing that configuration memories are a dominant factor; (2) a power breakdown comparison to analytical models, analyzing the source of accuracy loss; and (3) a robustness evaluation against process corners, studying their impact on energy consumption of full FPGA fabrics. Xifan Tang, Edouard Giacomin, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Towards high-performance polarity-controllable FETs with 2D materialsabstractAs scaling of conventional silicon-based electronics is reaching its ultimate limit, two-dimensional semiconducting materials of the transition-metal-dichalcogenides family, such as MoS2 and WSe2, are considered as viable candidates for next-generation electronic devices. Fully relying on electrostatic doping, polarity-controllable devices, that use additional gate terminals to modulate the Schottky barriers at source and drain, can strongly take advantages of 2D materials to achieve high on/off ratio and low leakage floor. Here, we provide an overview of the latest advances in 2D material processes and growth. Then, we report on the experimental demonstration of polarity-controllable devices fabricated on 2D-WSe2 and study the scaling trends of such devices using ballistic self-consistent quantum simulations. Finally, we discuss the circuit-level opportunities of such technology. Giovanni V. Resta, Jorge Romero Gonzalez, Yashwanth Balaji, Tarun Agarwal, Francky Catthoor, Iuliana P. Radu, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
DATE | 9 |
| 2018 | Practical challenges in delivering the promises of real processing-in-memory machinesabstractProcessing-in-Memory (PiM) machines promise to overcome the von Neumann bottleneck in order to further scale performance and energy efficiency of computing systems by reducing the extent of data transfer and offering ample parallelism. In this paper, we take the memristive Memory Processing Unit (mMPU) as a case study of a PiM machine and scrutinize it in practical scenarios. Specifically, we explore the limitations of parallelism and data transfer elimination. We argue that lack of operand locality and arrangement might make data transfer inevitable in the mMPU. We then devise techniques to move data within the mMPU, without transferring it off-chip, and quantify their costs. Additionally, we present electrical parameters that might limit the parallelism offered by the mMPU and evaluate their impact. Using benchmarks from the LGsynth91 suite, their vector extensions, and a few synthetic data-parallel workloads, we show that the internal data transfer results in an increase of up to 1.5× in the execution time, while the parallelism can be limited in some cases to 256 gates, resulting in an increase in execution time by 1.1× to 2×. Nishil Talati, Ameer Haj-Ali, Rotem Ben Hur, Nimrod Wald, Ronny Ronen, Pierre-Emmanuel Gaillardon, Shahar Kvatinsky |
DATE | 6 |
| 2018 | Emerging reconfigurable nanotechnologies: can they support future electronics?abstractSeveral emerging reconfigurable technologies have been explored in recent years offering device level runtime reconfigurability. These technologies offer the freedom to choose between p- and n-type functionality from a single transistor. In order to optimally utilize the feature-sets of these technologies, circuit designs and storage elements require novel design to complement the existing and future electronic requirements. An important aspect to sustain such endeavors is to supplement the existing design flow from the device level to the circuit level. This should be backed by a thorough evaluation so as to ascertain the feasibility of such explorations. Additionally, since these technologies offer runtime reconfigurability and often encapsulate more than one functions, hardware security features like polymorphic logic gates and on-chip key storage come naturally cheap with circuits based on these reconfigurable technologies. This paper presents innovative approaches devised for circuit designs harnessing the reconfigurable features of these nanotechnologies. New circuit design paradigms based on these nano devices will be discussed to brainstorm on exciting avenues for novel computing elements. Shubham Rai, Srivatsa Rangachar Srinivasa, Patsy Cadareanu, Xunzhao Yin, Xiaobo Sharon Hu, Pierre-Emmanuel Gaillardon, Narayanan Vijaykrishnan, Akash Kumar 0001 |
ICCAD | 6 |
| 2018 | Differential Power Analysis Mitigation Technique Using Three-Independent-Gate Field Effect TransistorsabstractHardware security vulnerabilities are a major concern for embedded computing devices which are now used in many application such as credit cards, SIM cards, or financial systems, putting sensible data at risk. Such systems are often targeted by differential power attacks, where the power trace can be monitored in order to get access to the sensible data. To alleviate this issue, a possible technique proposed in literature is to use a complementary gate (e.g., computing both XOR and XNOR operations in parallel) in order to have a symmetrical power trace for all possible input combinations. However, this technique results in a large area and power overhead since it approximatively requires twice the number of transistors. Recently, novel technologies such as Three-Independent-Gate Field Effect Transistors (TIGFETs) have been shown to be able to realize compact logic gates using less transistors when compared to Complementary Metal Oxide Semiconductor (CMOS) technology. In this paper, we investigate the benefits of using TIGFETs in terms of hardware security. First, we show that using the complementary gate technique with TIGFETs can reduce the transistor count, the power trace variation, the switching power and leakage by 2×, 57%, 36% and 8× respectively, when compared to CMOS. In addition, we show that for the same transistor count and similar switching power, using TIGFETs can reduce the power trace variation and the leakage by 81% and 6.7× respectively when compared to CMOS. Edouard Giacomin, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 2 |
| 2018 | Logic Synthesis for RRAM-Based In-Memory ComputingabstractDesign of nonvolatile in-memory computing devices has attracted high attention to resistive random access memories (RRAMs). We present a comprehensive approach for the synthesis of resistive in-memory computing circuits using binary decision diagrams, and-inverter graphs, and the recently proposed majority-inverter graphs for logic representation and manipulation. The proposed approach allows to perform parallel computing on a multirow crossbar architecture for the logic representations of the given Boolean functions throughout a level-by-level implementation methodology. It also provides alternative implementations utilizing two different logic operations for each representation, and optimizes them with respect to the number of RRAM devices and operations, addressing area, and delay, respectively. Experiments show that upper bounds of the aforementioned cost metrics for the implementations obtained by our synthesis approach are considerably improved in comparison with the corresponding existing methods in both area and especially latency. Saeideh Shirinzadeh, Mathias Soeken, Pierre-Emmanuel Gaillardon, Rolf Drechsler |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Guest Editorial Memristive-Device-Based ComputingabstractToday’s and emerging computing tasks are extremely demanding in terms of storage, energy efficiency, and computing efficiency; data-intensive/big-data applications and Internet-of-Things are couple of examples. In addition, today’s computer architectures and device technologies are facing major challenges making them incapable to deliver the required functionalities and features. Computers are facing the three well-known walls[1)]: the memory wall, the instruction level parallelism wall, and the power wall. Similarly, nanoscale CMOS technology is facing three walls[2)]: the reliability wall, the leakage wall, and the cost wall. In order for computing systems to continue to deliver sustainable benefits for the foreseeable future society, alternative computing architectures have to be explored in the light of emerging new device technologies. Using memristive device technology[3)]to enable new computing paradigms such as computation-in-memory architecture[4)]–[7)]is one of the emerging alternatives that could provide a huge potential in terms of energy and computing efficiency. Said Hamdioui, Pierre-Emmanuel Gaillardon, Dietmar Fey, Tajana Rosing |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Multi-level logic benchmarks: An exactness studyabstractIn this paper, we study exact multi-level logic benchmarks. We refer to an exact logic benchmark, or exact benchmark in short, as the optimal implementation of a given Boolean function, in terms of minimum number of logic levels and/or nodes. Exact benchmarks are of paramount importance to design automation because they allow engineers to test the efficiency of heuristic techniques used in practice. When dealing with two-level logic circuits, tools to generate exact benchmarks are available, e.g., espresso-exact, and scale up to relatively large size. However, when moving to modern multi-level logic circuits, the problem of deriving exact benchmarks is inherently more complex. Indeed, few solutions are known. In this paper, we present a scalable method to generate exact multi-level benchmarks with the optimum, or provably close to the optimum, number of logic levels. Our technique involves concepts from graph theory and joint support decomposition. Experimental results show an asymptotic exponential gap between state-of-the-art synthesis techniques and our exact results. Our findings underline the need for strong new research in logic synthesis. Luca G. Amarù, Mathias Soeken, Winston Haaswijk, Eleonora Testa, Patrick Vuillod, Jiong Luo, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ASP-DAC | 7 |
| 2017 | A novel basis for logic rewritingabstractGiven a set of logic primitives and a Boolean function, exact synthesis finds the optimum representation (e.g., depth or size) of the function in terms of the primitives. Due to its high computational complexity, the use of exact synthesis is limited to small networks. Some logic rewriting algorithms use exact synthesis to replace small subnetworks by their optimum representations. However, conventional approaches have two major drawbacks. First, their scalability is limited, as Boolean functions are enumerated to precompute their optimum representations. Second, the strategies used to replace subnetworks are not satisfactory. We show how the use of exact synthesis for logic rewriting can be improved. To this end, we propose a novel method that includes various improvements over conventional approaches: (i) we improve the subnetwork selection strategy, (ii) we show how enumeration can be avoided, allowing our method to scale to larger subnetworks, and (iii) we introduce XOR Majority Graphs (XMGs) as compact logic representations that make exact synthesis more efficient. We show a 45.8% geometric mean reduction (taken over size, depth, and switching activity), a 6.5% size reduction, and depth · size reductions of 8.6%, compared to the academic state-of-the-art. Finally, we outperform 3 over 9 of the best known size results for the EPFL benchmark suite, reducing size by up to 11.5% and depth up to 46.7%. Winston Haaswijk, Mathias Soeken, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ASP-DAC | 4 |
| 2017 | Endurance management for resistive Logic-In-Memory computing architecturesabstractResistive Random Access Memory (RRAM) is a promising non-volatile memory technology which enables modern in-memory computing architectures. Although RRAMs are known to be superior to conventional memories in many aspects, they suffer from a low write endurance. In this paper, we focus on balancing memory write traffic as a solution to extend the lifetime of resistive crossbar architectures. As a case study, we monitor the write traffic in a Programmable Logic-in-Memory (PLiM) architecture, and propose an endurance management scheme for it. The proposed endurance-aware compilation is capable of handling different trade-offs between write balance, latency, and area of the resulting PLiM implementations. Experimental evaluations on a set of benchmarks including large arithmetic and control functions show that the standard deviation of writes can be reduced by 86.65% on average compared to a naive compiler, while the average number of instructions and RRAM devices also decreases by 36.45% and 13.67%, respectively. Saeideh Shirinzadeh, Mathias Soeken, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Rolf Drechsler |
DATE | 3 |
| 2017 | Wave pipelining for majority-based beyond-CMOS technologiesabstractThe performance of some emerging nanotechnologies benefits from wave pipelining. The design of such circuits requires new models and algorithms. Thus we show how Majority-Inverter Graphs (MIG) can be used for this purpose and we extend the related optimization algorithms. The resulting designs have increased throughput, something that has traditionally been a weak point for the majority of non-charge-based technologies. We benchmark the algorithm on MIG netlists with three different technologies, Spin Wave Devices (SWD), Quantum-dot Cellular Automata (QCA), and NanoMagnetic Logic (NML). We find that the wave pipelined version of the netlists have an improvement in throughput over power of 23×, 13×, and 5× for SWD, QCA, and NML, respectively. In terms of throughput over area ratio, the improvement is 5×, 8×, and 3×, respectively. Odysseas Zografos, A. De Meester, Eleonora Testa, Mathias Soeken, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Luca G. Amarù, Praveen Raghavan, Francky Catthoor, Rudy Lauwereins |
DATE | 5 |
| 2017 | Improving Circuit Mapping Performance Through MIG-based Synthesis for Carry ChainsabstractHard-wired carry chains in FPGAs are designed to improve efficiency of important arithmetic primitives. Although they are proven to be effective for arithmetic-rich functions, there are very few studies on the optimization opportunities of carry chains for general logic that is poor in arithmetic operations. Recently, Majority-Inverter Graphs (MIGs) were proposed for efficient Boolean logic optimization. MIGs open an opportunity for efficient mapping of critical paths onto hard carry chains, as the carry logic of a full adder is naturally a majority (MAJ) gate. In this paper, we propose an MIG-based synthesis method to exploit hard adders in FPGAs for the mapping of general logic. The proposed heuristic algorithm selects MAJ nodes to be mapped on the carry chains and the associated LUTs; then, the efficiency of carry chain mapping is examined theoretically for efficient LUT utilization. The experimental results show that, compared to traditional design flow Verilog-to-Routing (VTR 7.0), the proposed approach can improve delay by up to 25% with an average of 8%, while the channel width is reduced by up to 20% with an average of 6%. Zhufei Chu, Xifan Tang, Mathias Soeken, Ana Petkovska, Grace Zgheib, Luca G. Amarù, Yinshui Xia, Paolo Ienne, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ACM Great Lakes Symposium on VLSI | 10 |
| 2017 | Enabling exact delay synthesisabstractGiven (i) a Boolean function, (ii) a set of arrival times at the inputs, and (iii) a gate library with associated delay values, the exact delay synthesis problem asks for a circuit implementation which minimizes the arrival time at the output(s). The exact delay synthesis problem, with given input arrival times, relates to computing the communication complexity of a Boolean function, which is an intractable problem. Input arrival times are variable and can take any value, thereby making the exact delay synthesis search space infinite. This paper presents theory and algorithms for exact delay synthesis. We introduce the theory of equioptimizable arrival times, which allows us to partition all arrival time patterns into a finite set of equivalence classes. Thanks to this new theory, we create for the first time exact delay circuit databases covering all Boolean functions up to 5 variables and all possible arrival time patterns. We describe further arrival time compression techniques which enable the creation of larger databases. We propose an enhanced delay synthesis flow capable of dealing with large circuits, combining exact delay logic rewriting and Boolean optimization techniques, attaining unprecedented results. We improve 9/10 of the best known results in the EPFL arithmetic delay synthesis competition, outperforming previous best results up to 3x. Embedded in a commercial EDA flow for ASICs, our exact delay synthesis techniques reduce the total negative slack by 12.17%, after physical implementation, at negligible area and runtime costs. Luca G. Amarù, Mathias Soeken, Patrick Vuillod, Jiong Luo, Alan Mishchenko, Pierre-Emmanuel Gaillardon, Janet Olson, Robert K. Brayton, Giovanni De Micheli |
ICCAD | 6 |
| 2017 | Design methodology for area and energy efficient OxRAM-based non-volatile flip-flopabstractWith the introduction of the Internet of Things (IoT), power consumption became a major design issue in modern system-on-chips. In advanced technologies, leakage power has become a dominant component, especially during sleep periods. Leakage mainly comes from volatile memory elements, e.g., flip-flops that cannot be power-gated in order to retain their states. Non-Volatile Flip-Flop (NVFF) using emerging memory technologies, such as Resistive Random Access Memories (RRAM), are popular solutions to address this issue. In NVFF design, the resistance values of the memory element have a direct impact on the area and energy overhead of the structure. In this paper, we present a design methodology for area and energy efficient RRAM-based NVFF. By characterizing the optimal lower bound of the RRAM resistance ratio required for properly restoring the FF, the store and restore operations can be performed using optimal programming circuit area and energy. Four Transmission-Gate (TG) NVFF topologies implemented in 180nm CMOS technology were analyzed using the proposed methodology. The presented methodology shows that differential NVFF provides minimum restore resistance ratio down to 1.02 considering CMOS and RRAM variability. This enables improvements in terms of store energy (34%) and area overhead (40%) compared to reported state-of-the-art NV-TGFFs design approaches. Mahesh Nataraj, Alexandre Levisse, Bastien Giraud, Jean-Philippe Noël, Pascal Andreas Meinerzhagen, Jean-Michel Portal, Pierre-Emmanuel Gaillardon |
ISCAS | 7 |
| 2017 | An efficient electronic measurement interface for memristive biosensorsabstractReducing sensing time is one major concern in clinical diagnostics. In the present work, a robust measurement system is developed aiming at the faster and easier signal acquisition of memristive biosensors. Sensing chips consisting of nanofabricated silicon wires exhibiting memristive electrical response and metallic extension electrodes allowing an integrated measurement procedure are designed and fabricated. Furthermore, the electrical response of these particular nanofabricated structures is for the first time acquired using an embedded-system-based measurement front-end. The suggested prototype significantly simplifies the measurement procedure and provides conveniently the response signal of the devices. Such optimized co-design of memristive biosensors with electronic platforms hold great promise for PoC (point-of-care) applications. Sebastien Naus, Ioulia Tzouvadaki, Pierre-Emmanuel Gaillardon, Armando Biscontini, Giovanni De Micheli, Sandro Carrara |
ISCAS | 3 |
| 2017 | RM3 based logic synthesis (Special session paper)abstractIn-memory computing devices, such as resistive RAMs, natively implement material implication or a variant of the majority-of-three operation called RM3. This operation generalizes material implication and has been used as target operation in several logic synthesis algorithms for in-memory computing applications. In this work, we investigate a homogeneous logic network data structure that uses RM3as only logic operation. Such a data structure makes an ideal fit for the use in design automation algorithms for in-memory computing. We show how to derive RM3networks from well-known logic synthesis data structures and a technique how to obtain such networks using technology mapping. Mathias Soeken, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 2 |
| 2017 | Physical Design Considerations of One-level RRAM-based Routing MultiplexersabstractResistive Random Access Memory(RRAM) technology opens the opportunity for granting both high-performance and low-power features to routing multiplexers. In this paper, we study the physical design considerations related to RRAM-based routing multiplexers and particularly the integration of 4T(ransistor)1R(RAM) programming structures within their routing tree. We first analyze the limitations in the physical design of a naive one-level 4T1R-based multiplexer, such as co-integration of low-voltage nominal power supply and high voltage programming supply, as well as the use of long metal wires across different isolating wells. To address the limitations, we improve the one-level 4T1R-based multiplexer by re-arranging the nominal and programming voltage domains, and also study the optimal location of RRAMs in terms of performance. The improved design can effectively reduce the length of long metal wires by 50%. Electrical simulations show that using a 7nm FinFET transistor technology, the improved 4T1R-based multiplexers improve delay by 69% as compared to the basic design. At nominal working voltage, considering an input size ranging from 2 to 32, the improved 4T1R-based multiplexers outperform the best CMOS multiplexers in area by 1.4x, delay by 2x and power by 2x respectively. The improved 4T1R-based multiplexers operating at near-Vt regime can improve Power-Delay Product by up to 5.8x when compare to the best CMOS multiplexers working at nominal voltage. Xifan Tang, Edouard Giacomin, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ISPD | 4 |
| 2017 | Exact Synthesis of Majority-Inverter Graphs and Its ApplicationsabstractWe propose effective algorithms for exact synthesis of Boolean logic networks using satisfiability modulo theories (SMTs) solvers. Since exact synthesis is a difficult problem, it can only be applied efficiently to very small functions, having up to six variables. Key in our approach is to use majority-inverter graphs (MIGs) as underlying logic representation as they are simple (homogeneous logic representation) and expressive (contain AND/OR-inverter graphs) at the same time. This has a positive impact on the problem formulation: it simplifies the encoding as SMT constraints and also allows for various techniques to break symmetries in the search space due to the regular data structure. Our algorithm optimizes with respect to the MIG's size or depth and uses different ways to encode the problem and several methods to improve solving time, with symmetry breaking techniques being the most effective ones. We discuss several applications of exact synthesis and motivate them by experiments on a set of large arithmetic benchmarks. Using the proposed techniques, we are able to improve both area and delay after lookup table (LUT)-based technology mapping beyond the current results achieved by state-of-the-art logic synthesis algorithms. Mathias Soeken, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Majority-based synthesis for nanotechnologiesabstractWe study the logic synthesis of emerging nanotechnologies whose elementary devices abstraction is a majority voter. We argue that synthesis tools, natively supporting the majority logic abstraction, are the technology enablers. This is because they allow designers to validate majority-based nanotechnologies on large-scale benchmarks. We describe models and data-structures for logic design with majority-based nanotechnologies and we show results of applying new synthesis algorithms and tools. We conclude that new logic synthesis methods are required to achieve a fair assessment on emerging nanotechnologies. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ASP-DAC | 2 |
| 2016 | An MIG-based compiler for programmable logic-in-memory architecturesabstractResistive memories have gained high research attention for enabling design of in-memory computing circuits and systems. We propose for the first time an automatic compilation methodology suited to a recently proposed computer architecture solely based on resistive memory arrays. Our approach uses Majority-Inverter Graphs (MIGs) to manage the computational operations. In order to obtain a performance and resource efficient program, we employ optimization techniques both to the underlying MIG as well as to the compilation procedure itself. In addition, our proposed approach optimizes the program with respect to memory endurance constraints which is of particular importance for in-memory computing architectures. Mathias Soeken, Saeideh Shirinzadeh, Pierre-Emmanuel Gaillardon, Luca G. Amarù, Rolf Drechsler, Giovanni De Micheli |
DAC | 3 |
| 2016 | Exploiting inherent characteristics of reversible circuits for faster combinational equivalence checking
Luca G. Amarù, Pierre-Emmanuel Gaillardon, Robert Wille, Giovanni De Micheli |
DATE | 2 |
| 2016 | The Programmable Logic-in-Memory (PLiM) computer
Pierre-Emmanuel Gaillardon, Luca G. Amarù, Anne Siemon, Eike Linn, Rainer Waser, Anupam Chattopadhyay, Giovanni De Micheli |
DATE | 1 |
| 2016 | Fast logic synthesis for RRAM-based in-memory computing using Majority-Inverter Graphs
Saeideh Shirinzadeh, Mathias Soeken, Pierre-Emmanuel Gaillardon, Rolf Drechsler |
DATE | 3 |
| 2016 | Optimizing Majority-Inverter Graphs with functional hashing
Mathias Soeken, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DATE | 3 |
| 2016 | A Full-Capacity Local RoutingArchitecture for FPGAs (Abstract Only)abstractReconfigurable systems employ highly-routable local routing architecture to interconnect generic fine-grain logic blocks. Commercial FPGAs employ 50% sparse crossbars rather than fully-connected crossbars in their local routing architecture to trade off between the area and routability of the Logic Blocks (LBs). While the input crossbar provides good routability and logic equivalence for the inputs of the LB, the outputs of the LBs are typically assigned to a physical location. This lack of flexibility brings strong constraints to the global net router. Here, we propose a novel local routing architecture that guarantees full logic equivalence on all input and output pins of the LBs. First, we introduce full-capacity crossbars to interconnect the outputs of the fine-grain Logic Elements (LEs) to the output pins of the LBs. Second, in the local routing, we use a combination of fully-connected and full-capacity crossbars. The full-capacity crossbars are used for the feedback connections in place of the standard fully-connected crossbars to ensure a full routability while reducing the area footprint. Fully-connected crossbars are still employed for the input connections to maintain the logic equivalence of the inputs. As a result, the novel local routing architecture enhances the routability of the LB clusters without any area overhead. By granting the outputs with logic equivalence, the proposed local routing architecture unlocks the full optimization potential of FPGA routers. Architectural simulations show that without any modification on Verilog-to-Routing (VTR) tool suites, when a commercial FPGA architecture is considered and over a wide set of benchmarks, the novel local routing architecture can reduce 10% channel width and 11% routing area with 10% less area×delay×power on average. Therefore, the novel local routing architecture enhances the routability of FPGA, and brings opportunities in realizing larger implementations on a single FPGA chip. Xifan Tang, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
FPGA | 2 |
| 2016 | Digital, analog and RF design opportunities of three-independent-gate transistorsabstractField-Effect Transistors (FETs) with Three Independent Gates (TIG) can achieve different modes of operation according to the bias of the gate terminals. In particular, TIG FETs were recently shown capable of (i) device-level polarity control, (ii) dynamic threshold modulation and (iii) subthreshold slope tuning down to ultra-steep-slope operation. Experimentally demonstrated using both contemporary FinFETs and emerging silicon nanowires channel technologies, TIGFETs unlock several design opportunities. In this paper, we comment on the digital, analog and RF design capabilities offered by this new class of transistors. Pierre-Emmanuel Gaillardon, Mehdi Hasan, Luca G. Amarù, Ross M. Walker, Berardi Sensale Rodriguez |
ISCAS | 1 |
| 2016 | Emerging Technology-Based Design of Primitives for Hardware SecurityabstractHardware security concerns such as intellectual property (IP) piracy and hardware Trojans have triggered research into circuit protection and malicious logic detection from various design perspectives. In this article, emerging technologies are investigated by leveraging their unique properties for applications in the hardware security domain. Security, for the first time, will be treated as one design metric for emerging nano-architecture. Five example circuit structures including camouflaging gates, polymorphic gates, current/voltage-based circuit protectors, and current-based XOR logic are designed to show the high efficiency of silicon nanowire FETs and graphene SymFET in applications such as circuit protection and IP piracy prevention. Simulation results indicate that highly efficient and secure circuit structures can be achieved via the use of non-CMOS devices. Yu Bi, Kaveh Shamsi, Jiann-Shiun Yuan, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Xunzhao Yin, Xiaobo Sharon Hu, Michael T. Niemier, Yier Jin |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2016 | A Fault-Tolerant Ripple-Carry Adder with Controllable-Polarity TransistorsabstractThis article first explores the effects of faults on circuits implemented with controllable-polarity transistors. We propose a new fault model that suits the characteristics of these devices, and we report the results of a SPICE-based analysis of the effects of faults on the behavior of some basic gates implemented with them. Hence, we show that the considered devices are able to intrinsically tolerate a rather high number of faults. We finally exploit this property to build a robust and scalable adder whose area, performance, and leakage power characteristics are improved by 15%, 18%, and 12%;, respectively, when compared to an equivalent FinFET solution at 22nm technology node. Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Jian Zhang 0067, Giovanni De Micheli, Ernesto Sánchez 0001, Matteo Sonza Reorda |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | A Sound and Complete Axiomatization of Majority-n LogicabstractManipulating logic functions via majority operators recently drew the attention of researchers in computer science. For example, circuit optimization based on majority operators enables superior results as compared to traditional synthesis tools. Also, the Boolean satisfiability problem finds new solution approaches when described in terms of majority decisions. To support computer logic applications based on majority, a sound and complete set of axioms is required. Most of the recent advances in majority logic deal only with ternary majority (MAJ-3) operators because the axiomatization with solely MAJ-3 and complementation operators is well understood. However, it is of interest extending such axiomatization to$n$-ary majority operators (MAJ-$n$) from both the theoretical and practical perspective. In this work, we address this issue by introducing a sound and complete axiomatization of MAJ-$n$logic. Our axiomatization naturally includes existing MAJ-3 and MAJ-5 axiomatic systems. Based on this general set of axioms, computer applications can now fully exploit the expressive power of majority logic. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Anupam Chattopadhyay, Giovanni De Micheli |
IEEE Trans. Computers | 2 |
| 2016 | Majority-Inverter Graph: A New Paradigm for Logic OptimizationabstractIn this paper, we propose a paradigm shift in representing and optimizing logic by using only majority (MAJ) and inversion (INV) functions as basic operations. We represent logic functions by majority-inverter graph (MIG): a directed acyclic graph consisting of three-input majority nodes and regular/complemented edges. We optimize MIGs via a new Boolean algebra, based exclusively on majority and inversion operations, that we formally axiomatize in this paper. As a complement to MIG algebraic optimization, we develop powerful Boolean methods exploiting global properties of MIGs, such as bit-error masking. MIG algebraic and Boolean methods together attain very high optimization quality. Considering the set of IWLS'05 benchmarks, our MIG optimizer (MIGhty) enables a 7% depth reduction in LUT-6 circuits mapped by ABC while also reducing size and power activity, with respect to similar and-inverter graph (AIG) optimization. Focusing on arithmetic intensive benchmarks instead, MIGhty enables a 16% depth reduction in LUT-6 circuits mapped by ABC, again with respect to similar AIG optimization. Employed as front-end to a delay-critical 22-nm application-specified integrated circuit flow (logic synthesis + physical design) MIGhty reduces the average delay/area/power by 13%/4%/3%, respectively, over 31 academic and industrial benchmarks. We also demonstrate delay/area/power improvements by 10%/10%/5% for a commercial FPGA flow. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Efficient Statistical Parameter Selection for Nonlinear Modeling of Process/Performance VariationabstractWith the growing number of process variation (PV) sources in deeply nano-scaled technologies, parameterized device and circuit modeling is becoming very important for chip design and verification. However, the high dimensionality of parameter space, for PV analysis, is a serious modeling challenge for emerging VLSI technologies. These parameters correspond to various interdie and intradie variations, and considerably increase the difficulties of design validation. Today's response surface models and most commonly used parameter reduction methods, such as principal component analysis and independent component analysis, limit parameter reduction to linear or quadratic form and they do not address the higher order of nonlinearity among process and performance parameters. In this paper, we propose and validate a feature selection method to reduce the circuit modeling complexity associated with high parameter dimensionality. This method relies on a learning-based nonlinear sparse regression, and performs a parameter selection in the input space rather than creating a new space. This method is capable of dealing with mixed Gaussian and non-Gaussian parameters and results in a more precise parameter selection considering statistical nonlinear dependencies among input and output parameters. The application of this method is demonstrated in digital circuit timing analysis in both FinFET and Silicon Nanowire technologies. The results confirm the efficiency of this method to significantly reduce the number of required simulations while keeping estimation error small. Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Multiple Independent Gate FETs: How many gates do we need?abstractMultiple Independent Gate Field Effect Transistors (MIGFETs) are expected to push FET technology further into the semiconductor roadmap. In a MIGFET, supplementary gates either provide (i) enhanced conduction properties or (ii) more intelligent switching functions. In general, each additional gate also introduces a side implementation cost. To enable more efficient digital systems, MIGFETs must leverage their expressive power to realize complex logic circuits with few physical resources. Researchers face then the question: How many gates do we need? In this paper, we address the logic side of this question. We determine whether or not an increasing number of gates leads to more compact logic implementations. For this purpose, we develop a logic synthesis flow that intrinsically exploits a MIGFET switching function. Using simplified design assumptions and device/interconnect models, we synthesize MCNC benchmarks on 5 promising MIGFET devices, with number of gates ranging from 1 to 7. Experimental results evidence nontrivial area/delay/energy minima, located between 1 and 4 gates, depending on a MIGFET switching function and device/interconnect technology. Luca G. Amarù, Gage Hills, Pierre-Emmanuel Gaillardon, Subhasish Mitra, Giovanni De Micheli |
ASP-DAC | 3 |
| 2015 | A ultra-low-power FPGA based on monolithically integrated RRAMs
Pierre-Emmanuel Gaillardon, Xifan Tang, Jury Sandrini, Maxime Thammasack, Somayyeh Rahimian Omam, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DATE | 1 |
| 2015 | Fault modeling in controllable polarity silicon nanowire circuits
Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DATE | 2 |
| 2015 | Towards More Efficient Logic Blocks By Exploiting Biconditional Expansion (Abstract Only)abstractNowadays, Field Programmable Gate Arrays (FPGA) exploit Look-Up Tables (LUTs) to generate logic functions. A K-input LUT can implement any Boolean functions with K inputs. Thanks to this flexibility, LUTs remained conceptually unchanged in FPGAs, only the number of inputs increased in time. Unfortunately, the flexibility does not come for free and LUTs have non-negligible costs in both circuit-level performances (large number of memories, area or delay penalties) and logic-level capabilities (limited fan-out). Here, we propose an FPGA fabric based on two novel logic blocks. First, we introduce a new LUT design showing reduced power consumption with no sacrifice in the logic flexibility. Then, we present a block suited to arithmetic functions but preserving enough versatility to implement general logic functions. The two blocks are supported by a recently introduced logic representation called Biconditional Binary Decision Diagrams (BBDDs). Using architectural-level benchmarking, we showed that an FPGA architecture exploiting the novel blocks performs significantly better than current state-of-the-art FPGA architectures at 40nm technological node over a large set of test circuits. While reducing the power consumption of MCNC big20 benchmarks by 29%, the proposed architecture is able to efficiently implement arithmetic circuits as compared to its traditional LUT-based FPGA counterpart. For instance, a 256-bit adder can be realized with a 43% gain in area×delay product. While considering large general and arithmetic logic benchmarks, we observe, on average, 4%, 3% and 10% improvements in area, delay and power respectively. Pierre-Emmanuel Gaillardon, Gain Kim, Xifan Tang, Luca G. Amarù, Giovanni De Micheli |
FPGA | 1 |
| 2015 | Accurate power analysis for near-Vt RRAM-based FPGAabstractResistive Random Access Memory (RRAM)-based FPGA architectures employ RRAMs not only as memories to store the configuration but embed them in the datapaths of programmable routing resources to propagate signals with improved performances. Sources of power consumption have been intensively studied for conventional Static Random Access Memories (SRAM)-based FPGAs. However, very limited works focused so far on studying the power characteristics of RRAM-based FPGAs. In this paper, we first analyze the power characteristics of RRAM-based multiplexer at circuit level and then use electrical simulations to study power consumption of RRAM-based FPGA architectures. Experimental results show that RRAM-based FPGAs achieve a Power-Delay Product reduced by 50% compared to SRAM-based FPGA at nominal voltage and 20% compared to near-VtSRAM-based FPGA, respectively. Xifan Tang, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
FPL | 2 |
| 2015 | Exploiting the Expressive Power of Graphene Reconfigurable Gates via Post-Synthesis OptimizationabstractAs an answer to the new electronics market demands, semiconductor industry is looking for different materials, new process technologies and alternative design solutions that can support Silicon replacement in the VLSI domain. The recent introduction of graphene, together with the option of electrostatically controlling its doping profile, has shown a possible way to implement fast and power efficient Reconfigurable Gates (RGs). Also, and this is the most important feature considered in this work, those graphene RGs show higher expressive power, i.e., they implement more complex functions, like Majority, MUX, XOR, with less area w.r.t. CMOS counterparts. Unfortunately, state-of-the-art synthesis tools, which have been customized for standard NAND/NOR CMOS gates, do not exploit the aforementioned feature of graphene RGs. Sandeep Miryala, Valerio Tenace, Andrea Calimera, Enrico Macii, Massimo Poncino, Luca G. Amarù, Giovanni De Micheli, Pierre-Emmanuel Gaillardon |
ACM Great Lakes Symposium on VLSI | 8 |
| 2015 | Reliable and high performance STT-MRAM architectures based on controllable-polarity devicesabstractSource degeneration of access devices in the parallel (P)_ anti-parallel (AP) switching in Spin Transfer Torque Magnetic Random Access Memories (STT-MRAM) has ultimately been a limiting factor in the operational speed of these types of memories. In this work, new architectures for memory single-cells and arrays of cells are presented that utilize Schottky-Barrier Silicon Nanowire Field Effect Transistors with polarity control capabilities (e.g., SiNW-FETs), to substantially increase the performance of STT-MRAM, specifically Multi-Level Cell (MLC) STT-MRAM. The proposed design offers built-in reliability improvement as it omits one of the available four states in the MLC STT-MRAM memory facilitating the resistance level detection for peripheral circuitry. Our simulation results of the developed memory cell show 49.7% reductions in P-AP switching time, as well as 51.3% increases in available drive current under 1.4V supply voltage when compared to FinFET 22imi technology. With respect to memory arrays, the proposed architecture demonstrates an average write latency reduction of 37% in comparison with FinFET 22nm technology node. Kaveh Shamsi, Yu Bi, Yier Jin, Pierre-Emmanuel Gaillardon, Michael T. Niemier, Xiaobo Sharon Hu |
ICCD | 4 |
| 2015 | FPGA-SPICE: A simulation-based power estimation framework for FPGAsabstractMainstream Field Programmable Gate Array (FPGA) power estimation tools are based on probabilistic activity estimation and analytical power models. The power consumption of the programmable resources of FPGAs is highly sensitive to their configurations. Due to their highly flexible nature, the configurations of FPGAs routing multiplexers or Look Up Tables (LUTs) are really different from a design to another but current analytical power models cannot accurately capture the associated power differences. In this paper, we introduce a simulation-based power estimation framework for FPGAs, called FPGA-SPICE, which supports any FPGA architecture that can be described with an architectural description language. Our power estimation engine automatically generates accurate SPICE netlists according to the FPGA configurations and enables precise power analysis of FPGA architectures. SPICE testbenches can be generated at different level of complexity, denoted as full-chip-level, grid-level and component-level testbenches. Full-chip-level testbenches dump the netlists associated with the complete FPGA fabric. To reduce simulation time, FPGA-SPICE can split the full-chip-level testbenches into grid-level testbenches, each of which consisting of a complete logic block netlist, or component-level testbenches, which consider individual circuit elements, i.e., multiplexers, LUTs, flip-flops, etc., separately. We show that the grid/component-level approach can achieve 14 × speed-up with a moderate 14% accuracy loss, compared to the full-chip level. We also use FPGA-SPICE to study the power characteristics of a commercial FPGA architecture at different technology nodes. Experimental results show that the global routing architecture consumes 50% of the total power, the local routing architecture claims for 40% of the total power, and the remaining 10% comes from the LUTs and flip-flops. Xifan Tang, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ICCD | 2 |
| 2015 | A Survey on Low-Power Techniques with Emerging Technologies: From Devices to SystemsabstractNowadays, power consumption is one of the main limitations of electronic systems. In this context, novel and emerging devices provide new opportunities to extend the trend toward low-power design. In this survey article, we present a transversal survey on energy-efficient techniques ranging from devices to architectures. The actual trends of device research, with fully depleted planar devices, tri-gate geometries, and gate-all-around structures, allows us to reach an increasingly higher level of performance while reducing the associated power. In addition, beyond the simple device property enhancements, emerging devices also lead to innovations at the circuit and architectural levels. In particular, devices whose properties can be tuned through additional terminals enable a fine and dynamic control of device threshold. They also enable designers to realize logic gates and to implement power-related techniques in a compact way unreachable to standard technologies. These innovations reduce power consumption at the gate level and unlock new means of actuation in architectural solutions like adaptive voltage and frequency scaling. Pierre-Emmanuel Gaillardon, Edith Beigné, Suzanne Lesecq, Giovanni De Micheli |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2015 | New Logic Synthesis as Nanotechnology EnablerabstractNanoelectronics comprises a variety of devices whose electrical properties are more complex as compared to CMOS, thus enabling new computational paradigms. The potentially large space for innovation has to be explored in the search for technologies that can support large-scale and high-performance circuit design. Within this space, we analyze a set of emerging technologies characterized by a similar computational abstraction at the design level, i.e., a binary comparator or a majority voter. We demonstrate that new logic synthesis techniques, natively supporting this abstraction, are the technology enablers. We describe models and data-structures for logic design using emerging technologies and we show results of applying new synthesis algorithms and tools. We conclude that new logic synthesis methods are required to both evaluate emerging technologies and to achieve the best results in terms of area, power and performance. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Subhasish Mitra, Giovanni De Micheli |
Proc. IEEE | 2 |
| 2015 | A Novel FPGA Architecture Based on Ultrafine Grain Reconfigurable Logic CellsabstractIn this paper, we investigate the opportunity brought by controllable-polarity transistors to design efficient reconfigurable circuits. Controllable-polarity transistors are devices whose polarity can be electrostatically programmed to be either n- or p-type. Such devices are used to build ultrafine grain computation cells. These cells are arranged into regular matrices, called MClusters, with a fixed and incomplete interconnection pattern, employed to minimize the reconfigurable interconnection overhead. We subsequently use them into field-programmable gate arrays (FPGAs). To assess this architectural scheme in an efficient and objective manner, we present a complete benchmarking tool flow and focus on the packing algorithm developed to handle the architecture. We finally perform the evaluation with widely used benchmark circuits. Leveraging the ultrafine grain cells compactness from a system-level perspective, we show that FPGAs exploiting MClusters demonstrate average savings of 43% and 23% in area and delay, respectively, as compared with the CMOS lookup table FPGA counterpart at 22-nm technological node. Pierre-Emmanuel Gaillardon, Xifan Tang, Gain Kim, Giovanni De Micheli |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Data compression via logic synthesisabstractNowadays, most software and hardware applications are committed to reduce the footprint and resource usage of data. In this general context, lossless data compression is a beneficial technique that encodes information using fewer (or at most equal number of) bits as compared to the original representation. A traditional compression flow consists of two phases: data decorrelation and entropy encoding. Data decorrelation, also called entropy reduction, aims at reducing the autocorrelation of the input data stream to be compressed in order to enhance the efficiency of entropy encoding. Entropy encoding reduces the size of the previously decorrelated data by using techniques such as Huffman coding, arithmetic coding, and others. When the data decorrelation is optimal, entropy encoding produces the strongest lossless compression possible. While efficient solutions for entropy encoding exist, data decorrelation is still a challenging problem limiting ultimate lossless compression opportunities. In this paper, we use logic synthesis to remove redundancy in binary data aiming to unlock the full potential of lossless compression. Embedded in a complete lossless compression flow, our logic synthesis based methodology is capable to identify the underlying function correlating a data set. Experimental results on data sets deriving from different causal processes show that the proposed approach achieves the highest compression ratio compared to state-of-art compression tools such as ZIP, bzip2 and 7zip. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Andreas Peter Burg, Giovanni De Micheli |
ASP-DAC | 2 |
| 2014 | Leveraging Emerging Technology for Hardware Security - Case Study on Silicon Nanowire FETs and Graphene SymFETsabstractHardware security concerns such as IP piracy and hardware Trojans have triggered research into circuit protection and malicious logic detection from various design perspectives. In this paper, emerging technologies are investigated by leveraging their unique properties for applications in the hardware security domain. Three example circuit structures including camouflaging gates, polymorphic gates and power regulators are designed to prove the high efficiency of silicon nanowire FETs and graphene Sym FET in applications such as circuit protection and IP piracy prevention. Simulation results indicate that highly efficient and secure circuit structures can be achieved via the use of emerging technologies. Yu Bi, Pierre-Emmanuel Gaillardon, Xiaobo Sharon Hu, Michael T. Niemier, Jiann-Shiun Yuan, Yier Jin |
ATS | 2 |
| 2014 | Majority-Inverter Graph: A Novel Data-Structure and Algorithms for Efficient Logic OptimizationabstractIn this paper, we present Majority-Inverter Graph (MIG), a novel logic representation structure for efficient optimization of Boolean functions. An MIG is a directed acyclic graph consisting of three-input majority nodes and regular/complemented edges. We show that MIGs include any AND/OR/Inverter Graphs (AOIGs), containing also the well-known AIGs. In order to support the natural manipulation of MIGs, we introduce a new Boolean algebra, based exclusively on majority and inverter operations, with a complete axiomatic system. Theoretical results show that it is possible to explore the entire MIG representation space by using only five primitive transformation rules. Such feature opens up a great opportunity for logic optimization and synthesis. We showcase the MIG potential by proposing a delay-oriented optimization technique. Experimental results over MCNC benchmarks show that MIG optimization reduces the number of logic levels by 18%, on average, with respect to AIG optimization performed by ABC academic tool. Employed in a traditional optimization-mapping circuit synthesis flow, MIG optimization enables an average reduction of {22%, 14%, 11%} in the estimated {delay, area, power} metrics, before physical design, as compared to academic/commercial synthesis flows. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DAC | 2 |
| 2014 | An efficient manipulation package for Biconditional Binary Decision DiagramsabstractBiconditional Binary Decision Diagrams (BBDDs) are a novel class of binary decision diagrams where the branching condition, and its associated logic expansion, is biconditional on two variables. Reduced and ordered BBDDs are remarkably compact and unique for a given Boolean function. In order to exploit BBDDs in Electronic Design Automation (EDA) applications, efficient manipulation algorithms must be developed and integrated in a software package. In this paper, we present the theory for efficient BBDD manipulation and its practical software implementation. The key features of the proposed approach are strong canonical form pre-conditioning of stored BBDD nodes, recursive formulation of Boolean operations in terms of biconditional expansions, performance-oriented memory management and dedicated BBDD re-ordering techniques. Experimental results show that the developed BBDD package achieves an average node count reduction of 19.48% and a speed-up factor of 1.63x with respect to a state-of-art decision diagram manipulation package. Employed in the synthesis of datapath circuits, the BBDD manipulation package is capable to advantageously restructure arithmetic operations producing 11.02% smaller and 32.29% faster circuits as compared to a commercial synthesis flow. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DATE | 2 |
| 2014 | Advanced system on a chip design based on controllable-polarity FETsabstractField-Effect Transistors (FETs) with on-line controllable-polarity are promising candidates to support next generation System-on-Chip (SoC). Thanks to their enhanced functionality, controllable-polarity FETs enable a superior design of critical components in a SoC, such as processing units and memories, while also providing native solutions to control power consumption. In this paper, we present the efficient design of a SoC core with controllable-polarity FET. Processing units are speeded-up at the datapath level, as arithmetic operations require fewer physical resources than in standard CMOS. Power consumption is decreased via embedded power-gating techniques and tunable high-performance/low-power devices operation. Memory cells are made smaller by merging the access interface with the storage circuitry. We foresee the advantages deriving from these techniques, by evaluating their impact on the design of SoC for a contemporary telecommunication application. Using a 22-nm vertically-stacked silicon nanowire technology, a coarse-grain evaluation at the block level estimates a delay and power reduction of 20% and 19% respectively, at a cost of a moderate area overhead of 15%, with respect to a state-of-art FinFET technology. Pierre-Emmanuel Gaillardon, Luca G. Amarù, Jian Zhang 0067, Giovanni De Micheli |
DATE | 1 |
| 2014 | Majority Logic Synthesis for Spin Wave TechnologyabstractSpin Wave Devices (SWDs) are promising beyond-CMOS candidates. Unlike traditional charge-based technologies, SWDs use spin as information carrier that propagates in waves. In this scenario, the logic primitive for computation is the majority gate. The majority gate has a greater expressive power than standard NAND/NOR gates, allowing SWD circuits to be more compact than CMOS, already at the logic level. Also, because there is not charge carrier transport, SWDs are estimated to have ultra-low power consumption. However, in order to exploit this opportunity, a native majority synthesis methodology is needed to fit the SWD technology needs. In this paper, we employ Majority-Inverter Graphs (MIGs) to naturally represent and synthesize SWD circuits. Thanks to the correspondence between the functionality of SWD primitive gates and MIG elements, MIG optimization intrinsically aims at minimum cost SWD implementations. Experimental results over MCNC benchmarks validate the efficiency of MIGs in SWD synthesis. As compared to traditional AND-Inverter Graph (AIG) synthesis, MIGs generate, on average, SWD circuits with 1.30X smaller area-delay-power product (ADP), improving their delay performance by 18%. Odysseas Zografos, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Praveen Raghavan, Giovanni De Micheli |
DSD | 3 |
| 2014 | A new basic logic structure for data-path computation (abstract only)abstractNowadays, Field Programmable Gate Arrays (FPGA) implement arithmetic functions using specific circuits at the logic block level, such as the carry paths, or at the structure level adopting Digital Signal Processing (DSP) blocks. Nevertheless, all these approaches, introduced to ease the realization of specific functions, are lacking of generality. In this paper, we introduce a new logic block that natively realizes arithmetic functions while preserving the versatility to implement general logic functions. It consists of a partially interconnected matrix of signal routers driven by comparators. We demonstrate that this structure can realize (i) any 2-output 2-input logic function or (ii) any single-output 3-input logic function or (iii) specific logic, such as arithmetic functions, with up to 4-output and 8-inputs. As compared to a standard 6-input Look Up Table (LUT), the proposed block requires roughly the same area but is 35.3% faster. Even though the proposed block has not the same exhaustive configurability of a 6-input LUT, there are arithmetic functions realizable in a single block that do not fit in one, or even more, 6-input LUT. For example, a single block inherently implements an entire 3-bit adder that requires 3× more resources with LUTs plus also custom circuitry. From a system level perspective, we show that a 256-bit adder is implemented with a gain on area×delay product of 31% as compared to its traditional LUT-based counterpart. Pierre-Emmanuel Gaillardon, Luca G. Amarù, Giovanni De Micheli |
FPGA | 1 |
| 2014 | Pattern-based FPGA logic block and clustering algorithmabstractIn classical FPGA, LUTs and DFFs are pre-packed into BLEs and then BLEs are grouped into logic blocks. We propose a novel logic block architecture with fast combinational paths between LUTs, called pattern-based logic blocks. A new clustering algorithm is developed to release the potential of pattern-based logic blocks. Experimental results show that the novel architecture and the associated clustering algorithm lead to a 14% performance gain and a 8% wirelength reduction with a 3% area overhead compared to conventional architecture in large control-instensive benchmarks. Xifan Tang, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
FPL | 2 |
| 2014 | A high-performance low-power near-Vt RRAM-based FPGAabstractThe routing architecture, heavily using programmable switches, dominates the area, delay and power of Field Programmable Gate Arrays (FPGAs). Resistive Random Access Memories (RRAMs) enable high-performance routing architectures through the replacement of Static Random Access Memory (SRAM)-based programming switches. Exploiting the very low on-resistance state achievable by RRAMs, RRAM-based routing multiplexers can be used to significantly reduce the FPGA routing delays. In addition, RRAM-based routing architectures are less sensitive to supply voltage reductions and show promises in low-power FPGA designs. In this paper, we propose a near-Vt low-power RRAM-based FPGA where both delay and power reductions are achieved. Experimental results demonstrate that a near-Vi RRAM-based FPGA design leads to a 15% area shrink, a 10% delay reduction, and a 65% power improvement, compared to a conventional FPGA design for a given technology node. To achieve low on-resistance values, RRAMs typically require high programming currents. In other word, they need relatively large programming transistors, potentially resulting in area, delay and power inefficiencies. We also present a design methodology to properly size the programming transistors of RRAMs in order to further improve the area-efficiency. Experimental results show that a correct programming transistor sizing strategy contributes to further 18% area and 2% delay shrink, compared to the initial near-Vi RRAM-based FPGA. Xifan Tang, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
FPT | 2 |
| 2014 | TSPC Flip-Flop circuit design with three-independent-gate silicon nanowire FETsabstractTrue Single-Phase Clock (TSPC) Flip-Flops, based on dynamic logic implementation, are area-saving and high-speed compared to standard static flip-flops. Furthermore, logic gates can be embedded into TSPC flip-flops which significantly improves performance. As a promising approach to keep the pace of Moore's Law, functionality-enhanced devices with multiple independent gates have drown many recent interests. In particular, Three-Independent-Gate Silicon Nanowire FETs (TIG SiNWFETs) can realize the functionality of two serial transistors in a single device. Therefore, they open new opportunities to compact designs in both arithmetic and control circuits. In this paper, we propose TSPC flip-flop implementation with asynchronous set and reset using the compactness of TIG SiNWFET. Electrical simulations show that TIG SiNWFET-based TSPC flip-flop improves nearly 20%, 30% and 7% in area, delay and leakage power respectively as compared to its LSTP FinFET counterpart at 22nm. Xifan Tang, Jian Zhang 0067, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 3 |
| 2014 | Novel grid-based power routing scheme for regular controllable-polarity FET arrangementsabstractPolarity-controllable transistors have emerged in the last few years as an adequate successor of current CMOS FinFETs. Due to the additional polarity terminal, novel physical design techniques are required. We present a novel grid-based power routing scheme able to mitigate the polarity terminal impact. The logic cells are organized in regular arrangements and easily configured using the novel power routing scheme. The impact of the placement and routing techniques used is gauged in terms of routing metal distribution, speed and area performance. Benchmark circuits are synthesized, placed and routed using commercial tools and performances are extracted. Post place and route results show 28% faster circuits compared to 22nm FinFET regular layout-based designs. Odysseas Zografos, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 2 |
| 2014 | System Level Benchmarking with Yield-Enhanced Standard Cell Library for Carbon Nanotube VLSI CircuitsabstractThe quest for technologies with superior device characteristics has showcased Carbon-Nanotube Field-Effect Transistors (CNFET) into limelight. In this work we present physical design techniques to improve the yield of CNFET circuits in the presence of Carbon Nanotube (CNT) imperfections. Various layout schemes are studied for enhancing the yield of CNFET standard cell library. With the help of existing ASIC design flow, we perform system-level benchmarking of CNFET circuits and compare them to CMOS circuits at various technology nodes. With CNFET technology, we observe maximum performance gains for circuits with gate-dominated delays. Averaged across various benchmarks at 16 nm, we report 8× improvement in Energy-Delay-Product (EDP) with CNFET circuits when compared to CMOS counterpart. We also study the performance of a complete OpenRISC processor, where we see 1.5× improvement in EDP over CMOS at 16 nm technology node. Voltage scaling enabled by CNFETs can be explored in the future for further performance benefits. Shashikanth Bobba, Jie Zhang 0007, Pierre-Emmanuel Gaillardon, H.-S. Philip Wong, Subhasish Mitra, Giovanni De Micheli |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2013 | MIXSyn: An efficient logic synthesis methodology for mixed XOR-AND/OR dominated circuitsabstractWe present a new logic synthesis methodology, called MIXSyn, that produces area-efficient results for mixed XOR-AND/OR dominated logic functions. MIXSyn is a two step synthesis process. The first step is a hybrid logic optimization that enables selective and distinct optimization of AND/OR and XOR-intensive portions of the logic circuit. The second step is a library-free technology mapping that enhances design flexibility with a tractable computational cost. MIXSyn has been tested on a set of large MCNC benchmarks. Experimental results indicate that MIXSyn produces CMOS circuits with 18.0% and 9.2% fewer devices, on the average, with respect to state-of-art academic and commercial synthesis tools, respectively. MIXSyn is also capable to exploit the opportunity of novel XOR implementations offered by the use of double-gate ambipolar devices. Experimental results show that MIXSyn can reduce the number of ambipolar transistors by 20.9% and 15.3%, on the average, with respect to state-of-art academic and commercial synthesis tools, respectively. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ASP-DAC | 2 |
| 2013 | BDS-MAJ: a BDD-based logic synthesis tool exploiting majority logic decompositionabstractDespite the impressive advance of logic synthesis during the past decades, a general methodology capable of efficiently synthesizing both control and datapath logic is still missing. Indeed, while synthesis techniques for random control logic (AND/OR-intensive) are well established, no dominant method for automated synthesis of datapath logic (XOR/MAJ-intensive) has yet emerged. Recently, Binary Decision Diagrams (BDDs) have been adopted to create an optimization system, named BDS, that supports integrated synthesis of both AND/OR- and XOR-intensive functions through functional logic decomposition on the BDD structure. However, it does not support direct decomposition and manipulation of majority logic which, instead, is widely used in datapath circuits. In this paper, we present the first BDD-based majority logic decomposition method and a logic decomposition system, BDS-MAJ, that enables efficient logic synthesis for both random control and datapath circuits. Experimental results show that logic synthesis based on BDS-MAJ produces CMOS circuits having on average 28.8% and 26.4% less area and, at the same time, 12.8% and 20.9% smaller delay with respect to academic ABC and BDS synthesis tools. Compared to commercial Synopsys Design Compiler synthesis tool, BDS-MAJ reduces on average the circuit area by 6.0% and decreases the delay by 7.8%. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DAC | 2 |
| 2013 | Towards structured ASICs using polarity-tunable Si nanowire transistorsabstractIn addition to scaling semiconductor devices down to their physical limit, novel devices show enhanced functionality compared to conventional CMOS. At advanced technology nodes, many devices exhibit ambipolar behavior, i.e., they show n- and p-type characteristics simultaneously. This phenomenon can be tamed using double-gate structures. In this paper, we present a complete framework relying on Double-Gate-all-around Vertically stacked NanoWire FETs (DG-NWFETs). Such device enables a compact realization of arithmetic logic functions and presents unprecedented interest for structured ASIC applications. Pierre-Emmanuel Gaillardon, Michele De Marchi, Luca G. Amarù, Shashikanth Bobba, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DAC | 1 |
| 2013 | Biconditional BDD: a novel canonical BDD for logic synthesis targeting XOR-rich circuitsabstractWe present a novel class of decision diagrams, called Biconditional Binary Decision Diagrams (BBDDs), that enable efficient logic synthesis for XOR-rich circuits. BBDDs are binary decision diagrams where the Shannon's expansion is replaced by the biconditional expansion. Since the biconditional expansion is based on the XOR/XNOR operations, XOR-rich logic circuits are efficiently represented and manipulated with canonical Reduced and Ordered BBDDs (ROBBDDs). Experimental results show that ROBBDDs have 37% fewer nodes on average compared to traditional ROBDDs. To exploit this opportunity in logic synthesis for XOR-rich circuits, we developed a BBDD-based One-Pass Synthesis (OPS) methodology. The BBDD-based OPS is capable to harness the potential of novel XOR-efficient devices, such as ambipolar transistors. Experimental results show that our logic synthesis methodology reduces the number of ambipolar transistors by 49.7% on average with respect to state-of-art commercial logic synthesis tool. Considering CMOS technology, the BBBD-based OPS reduces the device count by 31.5% on average compared to commercial synthesis tool. Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
DATE | 2 |
| 2013 | Vertically-stacked double-gate nanowire FETs with controllable polarity: from devices to regular ASICsabstractVertically stacked nanowire FETs (NWFETs) with gate-all-around structure are the natural and most advanced extension of FinFETs. At advanced technology nodes, many devices exhibit ambipolar behavior, i.e., the device shows n- and p-type characteristics simultaneously. In this paper, we show that, by engineering of the contacts and by constructing independent double-gate structures, the device polarity can be electrostatically programmed to be either n- or p-type. Such a device enables a compact realization of XOR-based logic functions at the cost of a denser interconnect. To mitigate the added area/routing overhead caused by the additional gate, an approach for designing an efficient regular layout, called Sea-of-Tiles is presented. Then, specific logic synthesis techniques, supporting the higher expressive power provided by this technology, are introduced and used to showcase the performance of the controllable-polarity NWFETs circuits in comparison with traditional CMOS circuits. Pierre-Emmanuel Gaillardon, Luca G. Amarù, Shashikanth Bobba, Michele De Marchi, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DATE | 1 |
| 2013 | 3.5-D integration: A case studyabstractTwo diverse manufacturing techniques for building 3-D integrated systems are vertical integration with Through-Silicon-Vias (TSVs), also referred as 3-D TSV integration, and 3D monolithic integration. In this paper, we present a hybrid integration scheme that combines these two approaches, taking into account their existing technology limits, into a disruptive paradigm called 3.5-D integration. Our novel integration supports circuit-partitioning both at the gate and block level with unprecedented benefits in cost. To demonstrate the effectiveness of 3.5-D integration, we chose as case study a 288-core MPSoC and we made hypothesis on the manufacturing and test cost. We argue a potential 20% decrease in the manufacturing cost and 30% decrease in the test cost when compared to 3-D TSV integration. In order to study the performance improvement of the MPSoC, we benchmarked various blocks of the core and the on-chip interconnection network, connecting all the cores. Our study shows large improvement in performance of the core (average of 11.5%) and latency (average of 24%) of the Network-on-Chip (NoC) for the 3.5-D integration when compared to the corresponding 3-D TSV implementation. Shashikanth Bobba, Pierre-Emmanuel Gaillardon, Ciprian Seiculescu, Vasilis F. Pavlidis, Giovanni De Micheli |
ISCAS | 2 |
| 2013 | Self-checking ripple-carry adder with Ambipolar Silicon NanoWire FETabstractFor the rapid adoption of new and aggressive technologies such as ambipolar Silicon NanoWire (SiNW), addressing fault-tolerance is necessary. Traditionally, transient fault detection implies large hardware overhead or performance decrease compared to permanent fault detection. In this paper, we focus on on-line testing and its application to ambipolar SiNW. We demonstrate on self-checking ripple-carry adder how ambipolar design style can help reduce the hardware overhead. When compared with equivalent CMOS process, ambipolar SiNW design shows a reduction in area of at least 56% (28%) with a decreased delay of 62% (6%) for Static (Transmission Gate) design style. Ogun Turkyilmaz, Fabien Clermidy, Luca G. Amarù, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 4 |
| 2013 | Dual-threshold-voltage configurable circuits with three-independent-gate silicon nanowire FETsabstractWe extend ambipolar silicon nanowire transistors by using three independent gates and show an efficient approach to implement dual-threshold-voltage configurable circuits. Polarity and threshold voltage of uncommitted devices are determined by applying different bias patterns to the three gates. Uncommitted logic gates can thus be configured to implement different logic functions for dual-threshold-voltage design using a wiring scheme, to target either high-performance or low-leakage applications. Synthesis of benchmark circuits with these devices shows comparable performance and 54% reduction of leakage power consumption compared to low-standby-power FinFET technology. Jian Zhang 0067, Pierre-Emmanuel Gaillardon, Giovanni De Micheli |
ISCAS | 2 |
| 2012 | GMS: Generic memristive structure for non-volatile FPGAsabstractThe invention of the memristor enables new possibilities for computation and non-volatile memory storage. In this paper we propose a Generic Memristive Structure (GMS) for 3-D FPGA applications. The GMS cell is demonstrated to be utilized for steering logic useful for multiplexing signals, thus replacing the traditional pass-gates in FPGAs. Moreover, the same GMS cell can be utilized for programmable memories as a replacement for the SRAMs employed in the look-up tables of FPGAs. A fabricated GMS cell is presented and its use in FPGA architecture is demonstrated by the area and delay improvement for several architectural benchmarks. Pierre-Emmanuel Gaillardon, Davide Sacchetto, Shashikanth Bobba, Yusuf Leblebici, Giovanni De Micheli |
VLSI-SoC | 1 |
| 2011 | Can we go towards true 3-D architectures?abstractThanks to recent technology advances, the exploration of the vertical dimension has been shown to be more than a dream for designers. Among those technologies, the vertical transistor has not been exploited yet. This paper describes a novel implementation of logic gates fully benefiting of nanowire-based vertical transistors embedded within the metal lines. The logic design in this technology is explored and its performance is evaluated. A comparison made on an equivalent technology node shows that our cells reduce area and delay by a factor of 31x and 2x respectively. Large reconfigurable logic circuits have been benchmarked showing an improvement of area and delay by 46% and 48% on average. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Paul-Henry Morel, Jean-Philippe Noël, Fabien Clermidy, Ian O'Connor |
DAC | 1 |
| 2011 | Evaluation of a crossbar multiplexer in a lithography-based nanowire technologyabstractSilicon Nanowire technology has been demonstrated to be a promising candidate to fabricate nanowire crossbars. The use of such devices in a real architectural as well as in a design environment is an ongoing research topic. In this paper, we investigate the use of a lithography-based industrial process for designing a 4-to-1 multiplexer in a crossbar circuit. We show that by considering the line parasitic, the crossbar demonstrates poor performance in a 65-nm technology, while the area and power savings are about 6× and 1.5× respectively vs. the CMOS implementation. However, extrapolation to the 9-nm node shows a 2× better performance and 67× area saving. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Fabien Clermidy, Ian O'Connor |
ISCAS | 1 |
| 2011 | Matrix Nanodevice-Based Logic Architectures and Associated Functional Mapping MethodabstractThis article describes a novel computing architecture organization based on nanoscale logic cells. We propose the use of a cluster of matrix arrangements of cells. In order to interconnect such fine-grained logic cells within a matrix, conventional techniques are not suitable due to a large interconnect overhead. Therefore, we propose the use of static and incomplete interconnect topologies to create matrices of cells. We also propose a method to map functions onto such architectures. We then explore the main parameters of the structure (size of matrices and interconnect topologies) and their impact on the main performance metrics (packing efficiency, speed, and fault tolerance). A cluster packing method also allows the evaluation of the number of matrices used by complex functions and the fill factor for various matrix sizes. The analyses show that this approach is particularly suited for matrices of 16 cells interconnected by modified omega networks. We can conclude that this architecture could improve the scalability of traditional FPGAs by a factor of 8.5. Pierre-Emmanuel Gaillardon, Fabien Clermidy, Ian O'Connor, Maimouna Amadou, Gabriela Nicolescu |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2010 | Phase-change-memory-based storage elements for configurable logicabstractBack-end-of-line non-volatile resistive memories like Phase Change Memories (PCMs) are promising to solve memory issues in different architectures. In this paper, we investigate the usage of PCM to build an elementary configuration memory node for reconfigurable logic, such as Field-Programmable Gate Arrays (FPGAs). We propose an elementary circuit realized by 2 resistive memories and 1 programming transistor able to store a configuration voltage. We investigate the proposed node in terms of area and write time and we assess its impact on complex circuits. We show that the elementary memory node yields an improvement in area and write time of 1.5x and 16x respectively vs. a regular Flash implementation. Implemented in FPGAs, the memory node yields a delay reduction up to 51%, thanks to the reduction of dimensions and low on-resistance of PCMs. Pierre-Emmanuel Gaillardon, M. Haykel Ben Jamaa, Marina Reyboz, Giovanni Beneventi, Fabien Clermidy, Luca Perniola, Ian O'Connor |
FPT | 1 |
| 2009 | Emerging Technologies and Nanoscale Computing Fabricsabstract6-8 July 2014 Ian O'Connor, Kotb Jabeur, Nataliya Yakymets, Renaud Daviot, David Navarro, Pierre-Emmanuel Gaillardon, Fabien Clermidy, Maimouna Amadou, Gabriela Nicolescu |
VLSI-SoC | 7 |