Roman Gauchi

dblp:254/9489 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-2948-9466ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Low Latency SEU Detection in FPGA CRAM With In-Memory ECC Checking
abstract
In harsh environments such as space, radiation and charged particles cause Single-Event Effects, faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy, are very effective against Single-Event Transients (SETs), but incur a minimum of$3\times $cost in area. Single-Event Upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process, going through every memory word looking for errors to repair. It involves a non-negligible Time To Detect (TTD) before repair, during which other events can occur and compromise the system. Field Programmable Gate Arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design integrity in harsh conditions. In this paper, we propose In-Memory Error Code Correction Checking (IMECCC), a method to replace memory scrubbing and improve FPGA configuration memory protection in high radiation environments. Our method allows asynchronous SEU detection, and replaces the scrubbing’s variable time to detect with a fixed TTD. We show that IMECCC reduces FPGA’s TTD by at least 116,$000\times $on average, with an area increase of$1.56\times $, using a test architecture resembling a Xilinx Virtex 5 QV at a 60MHz scrubbing frequency.
Aurélien Alacchi, Edouard Giacomin, Scott Temple, Roman Gauchi, Michael J. Wirthlin, Pierre-Emmanuel Gaillardon
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Smart-Redundancy With In Memory ECC Checking: Low-Power SEE-Resistant FPGA Architectures
abstract
In harsh environments, such as space, radiation, and charged particles cause single-event effects (SEEs), faults occurring randomly on any electronic component. These must be mitigated to ensure device functionality. Modern mitigation methods, such as triple modular redundancy (TMR), are very effective against single-event transients (SETs) but incur a minimum of$3\times $cost in the area. Single-event upsets (SEUs) affect sequential elements and are regularly repaired using memory scrubbing. Scrubbing is a slow serial process going through every memory word, looking for errors to repair. Scrubbing involves a nonnegligible amount of time before an error is detected, during which other events can occur and compromise the system. Field-programmable gate arrays (FPGAs) rely heavily on sequential elements to store their configuration; thus, FPGA’s SEU detection time is critical to ensuring design sustainability in harsh conditions. In this article, we propose an alternative mitigation method based on sensor integration and FPGA architecture modification, called smart-redundancy with in-memory error correction code checking (SRIMECCC). The sensors allow asynchronous SEU detection, reducing the time to detect by 57$250\times $on average and enabling local reconfiguration. Our method includes built-in dual redundancy that reduces the power consumption by 87% on average, benefiting embedded systems. SRIMECCC is also an area-efficient technique that saves 28% of total effective area compared to TMRed designs implemented in FPGAs.
Aurélien Alacchi, Edouard Giacomin, Roman Gauchi, Szymon Kulis, Pierre-Emmanuel Gaillardon
IEEE Trans. Very Large Scale Integr. Syst.3
2022 An Open-source Three-Independent-Gate FET Standard Cell Library for Mixed Logic Synthesis
abstract
Three-Independent-Gate FET (TIGFET) technology is one of the most promising candidates to succeed CMOS and FinFET technologies due to its low off-current, compact surface area, reconfigurable logic, and CMOS compatibility. In this paper, we present an open-source standard cell library based on silicon nano-wire TIGFETs, which enables efficient implementation of novel XOR-and-majority-based circuit designs. We also discuss logic synthesis methods tailored to take advantage of TIGFET capabilities to allow their potential to be realized at the system level. By combining the 10nm TIGFET technology with a mixed logic synthesis tool, the PicoRV core design shows a 2.3× lower area and a 5.7× lower energy consumption, compared to an equivalent low-power 12nm FinFET implementation.
Roman Gauchi, Ashton Snelgrove, Pierre-Emmanuel Gaillardon
ISCAS1
2022 An Energy-Efficient Three-Independent-Gate FET Cell Library for Low-Power Edge Computing
abstract
With the increasing demand for compute-intensive applications for IoT devices, new technologies that enable a power reduction at the device-level are needed to improve energy savings at the system-level. Unfortunately, the scaling of standard CMOS technologies is not as fast as the scaling of computing performance, which leads to the so-called "power wall". The Three-Independent-Gate Field-Effect Transistor (TIGFET) is a promising technology that enhances the device functionality to create more compact logic gates and provide silicon-nanowire structures that meet the requirements of low leakage power systems. However, the evaluation of complex designs is currently limited to the intrinsic model of the transistor and does not consider the parasitic effects of cell layouts. In this paper, we propose a standard cell library for 10-nm silicon-nanowire TIGFET devices, including combinational and sequential gates, to evaluate a production RISC-V core targeting low energy consumption budget. After synthesis, the core achieves 4 × lower energy consumption up to a frequency of 340 MHz compared to an equivalent low-power 12-nm FinFET technology node.
Michael Keyser, Roman Gauchi, Pierre-Emmanuel Gaillardon
VLSI-SoC2
2022 Towards a Truly Integrated Vector Processing Unit for Memory-bound Applications Based on a Cost-competitive Computational SRAM Design Solution
abstract
This article presents Computational SRAM (C-SRAM) solution combining In- and Near-Memory Computing approaches. It allows performing arithmetic, logic, and complex memory operations inside or next to the memory without transferring data over the system bus, leading to significant energy reduction. Operations are performed on large vectors of data occupying the entire physical row of C-SRAM array, leading to high performance gains. We introduce the C-SRAM solution in this article as an integrated vector processing unit to be used by a scalar processor as an energy-efficient and high performing co-processor. We detail the C-SRAM system design on different levels: (i) circuit design and silicon proof of concept, (ii) system interface and instruction set architecture, and (iii) high-level software programming and simulation. Experimental results on two complete memory-bound applications, AES and MobileNetV2, show that the C-SRAM implementation achieves up to 70× timing speedup and 37× energy reduction compared to scalar architecture, and up to 17× timing speedup and 5× energy reduction compared to SIMD architecture.
Maha Kooli, Antoine Heraud, Henri-Pierre Charles, Bastien Giraud, Roman Gauchi, Mona Ezzadeen, Kevin Mambu, Valentin Egloff, Jean-Philippe Noël
ACM J. Emerg. Technol. Comput. Syst.5
2021 Storage Class Memory with Computing Row Buffer: A Design Space Exploration
abstract
Today computing centric von Neumann architectures face strong limitations in the data-intensive context of numerous applications, such as deep learning. One of these limitations corresponds to the well known von Neumann bottleneck. To overcome this bottleneck, the concepts of In-Memory Computing (IMC) and Near-Memory Computing (NMC) have been proposed. IMC solutions based on volatile memories, such as SRAM and DRAM, with nearly infinite endurance, solve only partially the data transfer problem from the Storage Class Memory (SCM). Computing in SCM is extremely limited by the intrinsic poor endurance of the Non-Volatile Memory (NVM) technologies. In this paper, we propose to take the best of both solutions, by introducing a Computing Row Buffer (C-RB), using a Computing SRAM (C-SRAM) model, in place of the standard Row Buffer (RB) in the SCM. The principle is to keep operations on large vectors in the C-RB of the SCM, minimizing data movement to and from the CPU, thus drastically reducing energy consumption of the overall system. To evaluate the proposed architecture, we use an instruction accurate platform based on Intel Pin software. Pin instruments run time binaries in order to get applications' full memory traces of our solution. We achieve energy reduction up to 7.9x on average and up to 45x for the best case and speedup up to 3.8x on average and up to 13x for the best case, and a reduction of write accesses in the SCM up to 18 %, compared to SIMD 512-bit architecture.
Valentin Egloff, Jean-Philippe Noël, Maha Kooli, Bastien Giraud, Lorenzo Ciampolini, Roman Gauchi, César Fuguet Tortolero, Eric Guthmuller, Mathieu Moreau, Jean-Michel Portal
DATE6
2020 Computational SRAM Design Automation using Pushed-Rule Bitcells for Energy-Efficient Vector Processing
abstract
This paper presents a new methodology for automating the Computational SRAM (C-SRAM) design based on off-the-shelf memory compilers and a configurable RTL IP. The main goal is to drastically reduce the development effort compared to a full-custom design, while offering a flexibility of use and a high-yield production. The proposed C-SRAM architecture has been developed to process energy-efficient vector data coupled with a scalar processor, while limiting the data transfer on the system bus. The results obtained by post P&R simulations show that 2RW and 4RW C-SRAM configurations using the double pumping technique achieved the highest performance to process vectorized MAC operations compared to the others configurations. Moreover, it has been shown that the impact of the digital wrapper decoding and executing the instructions can be mitigated by increasing the memory cut size to represent less than 10% in area and 20% in power consumption.
Jean-Philippe Noël, Valentin Egloff, Maha Kooli, Roman Gauchi, Jean-Michel Portal, Henri-Pierre Charles, Pascal Vivet, Bastien Giraud
DATE4
2020 Reconfigurable tiles of computing-in-memory SRAM architecture for scalable vectorization
abstract
For big data applications, bringing computation to the memory is expected to reduce drastically data transfers, which can be done using recent concepts of Computing-In-Memory (CIM). To address kernels with larger memory data sets, we propose a reconfigurable tile-based architecture composed of Computational-SRAM (C-SRAM) tiles, each enabling arithmetic and logic operations within the memory. The proposed horizontal scalability and vertical data communication are combined to select the optimal vector width for maximum performance. These schemes allow to use vector-based kernels available on existing SIMD engines onto the targeted CIM architecture. For architecture exploration, we propose an instruction-accurate simulation platform using SystemC/TLM to quantify performance and energy of various kernels. For detailed performance evaluation, the platform is calibrated with data extracted from the Place&Route C-SRAM circuit, designed in 22nm FDSOI technology. Compared to 512-bit SIMD architecture, the proposed CIM architecture achieves an EDP reduction up to 60× and 34× for memory bound kernels and for compute bound kernels, respectively.
Roman Gauchi, Valentin Egloff, Maha Kooli, Jean-Philippe Noël, Bastien Giraud, Pascal Vivet, Subhasish Mitra, Henri-Pierre Charles
ISLPED1
2019 Memory Sizing of a Scalable SRAM In-Memory Computing Tile Based Architecture
abstract
Modern computing applications require more and more data to be processed. Unfortunately, the trend in memory technologies does not scale as fast as the computing performances, leading to the so called memory wall. New architectures are currently explored to solve this issue, for both embedded and off-chip memories. Recent techniques that bringing computing as close as possible to the memory array such as, In-Memory Computing (IMC), Near-Memory Computing (NMC), Processing-In-Memory (PIM), allow to reduce the cost of data movement between computing cores and memories. For embedded computing, In-Memory Computing scheme presents advantageous computing and energy gains for certain class of applications. However, current solutions are not scaling to large size memories and high amount of data to compute. In this paper, we propose a new methodology to tile a SRAM/IMC based architecture and scale the memory requirements according to an application set. By using a high level LLVM-based simulation platform, we extract IMC memory requirements for a certain class of applications. Then, we detail the physical and performance costs of tiling SRAM instances. By exploring multi-tile SRAM Place&Route in 28nm FD-SOI, we explore the respective performance, energy and cost of memory interconnect. As a result, we obtain a detailed wire cost model in order to explore memory sizing trade-offs. To achieve a large capacity IMC memory, by splitting the memory in multiple sub-tiles, we can achieve lower energy (up to 78% gain) and faster (up to 49% gain) IMC tile compared to a single large IMC memory instance.
Roman Gauchi, Maha Kooli, Pascal Vivet, Jean-Philippe Noël, Edith Beigné, Subhasish Mitra, Henri-Pierre Charles
VLSI-SoC1