Luca Sterpone

dblp:66/3508 · DBLP profile ↗
← Back
107ranked-venue papers
21as first author
33since 2021 · last 2026
0000-0002-3080-2560ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 99 · 19 first-author · 30 since 2021Software engineering, systems software and programming languages · 31 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Robust Quantum Communication for Space Systems Through an Heterogeneous RISC-V based FPGA/SoC Platform
abstract
Quantum Key Distribution (QKD) is increasingly adopted in security-critical links and future scenarios will include also space applications where radiation-induced faults can compromise the correctness and availability of the protocol. This risk is amplified on SRAM-based FPGAs, where configuration upsets can alter circuit behavior. Through this paper, we present a robust, heterogeneous FPGA/SoC platform to validate the reliability of QKD sifting on commercial off-the-shelf hardware. A lightweight RISC-V processor supervise the entire procedure, while the sifting accelerator is coupled with a lightweight monitoring unit to detects faults and triggers error correction through partial reconfiguration. The platform has been implemented on AMD ZCU102 UltraScale+ development boards and evaluated through fault injections. The adopted mitigation techniques allows downtime reduction by about 12000 × and lowers the sifting-module error rate by around 2%.
Giorgio Cora, Arash Amini Bardpareh, Gonzalo Miguel Joaquin Fernandez Lobo, Corrado De Sio, Sarah Azimi, Andrea Stanco, Luca Sterpone
CF7
2026 POSTER: A Reliable Multi-FPGA RISC-V Based Cluster for Space AI Inference
Giorgio Cora, Morgana Duni, Corrado De Sio, Sarah Azimi, Luca Sterpone
CF5
2026 Late Breaking Results: Never-Stopping Inference: Self-Healing AI Accelerators on SRAM-FPGAs
abstract
This work presents a self-healing runtime for AI accelerators on SRAM-based FPGAs that combines online fault detection with fine-grained partial reconfiguration to ensure continuous inference execution. The framework dynamically isolates and repairs faulty regions while remapping workloads to healthy resources, eliminating the need for redundant hardware and system downtime. The proposed approach reduces recovery latency by 3 orders of magnitude compared to the state-of-the-art.
Eleonora Vacca, Giorgio Cora, Luca Sterpone
DATE3
2026 Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks
Federico Buccellato, Corrado De Sio, Sarah Azimi, Luca Sterpone
IOLTS4
2026 In-Hardware Fault-Tolerance Controller for Multi-FPGA Clustered Architectures
Giorgio Cora, Daniele Rizzieri, Corrado De Sio, Sarah Azimi, Luca Sterpone
IEEE Trans. Computers5
2025 POSTER: AI-Powered Anomaly Detection for Satellite Telemetry
abstract
Reliable anomaly detection in satellite telemetry is critical for mission success, yet traditional threshold-based methods struggle with complex and evolving patterns.This work presents machine learning (ML) techniques to analyze high-dimensional telemetry data.Evaluations of real-world satellite telemetry datasets demonstrate the potential of ML to enhance spacecraft health monitoring and reduce manual intervention.
Federico Buccellato, Davide Nicolini, Eleonora Vacca, Corrado De Sio, Luca Sterpone
CF5
2025 On-Hardware Resilience Analysis of DPUAccelerated CNNs on FPGA-Based Systems
abstract
In recent years, Reconfigurable SoCs have emerged as a high-performance solution for embedded systems, addressing the increasing complexity of neural networks, balancing performance, cost, and adaptability. Flexible hardware accelerators, such as AMD’s Deep Learning Processing Units (DPUs), enable efficient computation across various domains, including safety-critical applications. However, soft errors remain a significant reliability concern, especially in harsh environments like space, where radiationinduced corruption of configuration memory poses a significant threat to FPGA-based systems. Most research on the reliability and robustness of deep learning models against soft errors has focused on application-level analyses, with comparatively little attention paid to architectural hardware faults. This paper introduces a resilience evaluation framework targeting AMD’s state-of-the-art DPU, comparing traditional application-level fault injection with hardware-aware fault injection performed on an actual hardware platform, a Kria KV260. Applying this methodology, we evaluated fourteen different deep neural network architectures and demonstrated that hardware-aware fault injections reveal critical vulnerabilities that applicationonly approaches fail to detect. Moreover, we investigated the source of different faults at the hardware level, enabling the identification of architectural resources that are more susceptible to errors. These insights are valuable to support the development of more robust deployment strategies and mitigation techniques tailored to FPGA-based deep learning accelerators.
Federico Buccellato, Corrado De Sio, Sarah Azimi, Luca Sterpone
DSD4
2025 Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
abstract
Network Interface Cards (NICs) greatly evolved from simple basic devices moving traffic in and out of the network to complex heterogeneous systems offloading host CPUs from performing complex tasks on in-transit packets. These latter comprise different types of devices, ranging from NICs accelerating fixed specific functions (e.g., on-the-fly data compression/decompression, checksum computation, data encryption, etc.) to complex Systems-on-Chip (SoC) equipped with both general purpose processors and specialized engines (Smart-NICs). Similarly, Field Programmable Gate Arrays (FPGAs) moved from pure reprogrammable devices to modern heterogeneous systems comprising general-purpose processors, real-time cores and even AI-oriented engines. Furthermore, the availability of high-speed network interfaces (e.g., SFPs) makes modern FPGAs a good choice for implementing Smart-NICs. In this work, we extended the functionalities offered by an open-source NIC implementation (Corundum) by enabling time-aware traffic management in hardware, and using this feature to control the bandwidth associated with different traffic classes. By exposing dedicated control registers on the AXI bus, the driver of the NIC can easily configure the transmission bandwidth of different prioritized queues. Basically, each control register is associated with a specific transmission queue (Corundum can expose up to thousands of transmission and receiving queues), and sets up the fraction of time in a transmission window which the queue is supposed to get access the output port and transmit the packets. Queues are then prioritized and associated to different traffic classes through the Linux QDISC mechanism. Experimental evaluation demonstrates that the approach allows to properly manage the bandwidth reserved to the different transmission flows.
Alberto Scionti, Paolo Savio, Francesco Lubrano, Federico Stirano, Antonino Nespola, Olivier Terzo, Corrado De Sio, Luca Sterpone
DSD8
2025 Routino: Accelerating FPGA Routing Through Efficient Memory Representation
abstract
The rapid increase in the complexity of Field-Programmable Gate Arrays (FPGAs) is significantly impacting the efficiency of the design implementation flow. In particular, the routing process presents challenges in achieving computational efficiency and reducing time-to-solution due to the increasing on-chip resources and device complexity. This work introduces a novel router that leverages optimized data structures and memory access patterns to minimize memory consumption. Experimental results prove how the proposed approach can significantly reduce time-to-solution, identifying memory consumption as a barrier to achieving scalability and proposing solutions based on FPGA modular architecture to face it, achieving an average memory usage reduction of about 90 % and an average decrease of routing time of 40 %.
Davide Nicolini, Corrado De Sio, Eleonora Vacca, Luca Sterpone
FPL4
2024 A Novel Robust Core for Detecting Node Failures in FPGA Clusters
abstract
Field Programmable Gate Arrays (FPGAs) are gaining popularity in different fields, including space applications, where high computational capabilities are required; for this reason, FPGAs are often used as nodes in clusters. When considering mission-critical systems, reliability must be ensured, even in radiation environments such as space. Thus, it is necessary to define a way of monitoring the entire system, ensuring the correct behavior of each node. This work introduces the Beacon Controller, a module to be implemented on the FPGA elements of a cluster for real-time monitoring of the computational elements of the node.
Giorgio Cora, Corrado De Sio, Sarah Azimi, Luca Sterpone
CF4
2024 Scalable K-Nearest Neighbors Implementation using Distributed Embedded Systems
abstract
The distributed embedded systems paradigm is a promising platform for high-performance embedded applications. We present a distributed algorithm and system based on cost-effective devices. The proof of concept shows how a parallelized approach leveraging a distributed embedded platform can address the computational of the Machine Learning K-Nearest Neighbors (K-NN) algorithm with large and heterogeneous datasets.
Corrado De Sio, Andrea Avignone, Luca Sterpone, Silvia Chiusano
CF3
2024 A New Reliability Analysis of RISC-V Soft Processor for Safety-Critical Systems
abstract
RISC-V soft processors are attractive for various applications, including mission-critical ones, thanks to their reduced costs and high flexibility. Despite their growing popularity, reliability analysis of such platforms is still in an early stage, mainly relying on system-level analysis only, leaving module-level assessment unexplored. Such limitations hinder the development of mitigation strategies that could effectively focus on vulnerabilities within a RISC-V soft processor system. We propose a methodology for evaluating the module-wise reliability of a RISC-V soft processor based on fine-grained fault injection, custom layout placement, and fault analysis. Through this approach, we can provide insights into the critical elements of the processor, identifying the most susceptible to faults, both at the module and system levels. The presented results enhance comprehension of weak points within the processor, paving the way for creating robust and dependable RISC-V systems.
Giorgio Cora, Corrado De Sio, Daniele Rizzieri, Sarah Azimi, Luca Sterpone
DDECS5
2024 On the Fault Tolerance of Self-Supervised Training in Convolutional Neural Networks
abstract
Deep neural networks (DNNs) are increasingly used in critical applications from healthcare to autonomous driving. However, their predictions were shown to degrade in the presence of transient hardware faults, leading to potentially catastrophic and unpredictable errors. Consequently, several techniques have been proposed to increase the fault tolerance of DNNs by modifying network structures and/or training procedures, thereby reducing the need for costly hardware redundancy. There are, however, design or training choices whose impact on fault propagation has been overlooked in the literature. In particular, self-supervised learning (SSL), as a pretraining technique, was shown to improve the robustness of the learned features, resulting in better performance in downstream tasks. This study investigates the fault tolerance of several SSL techniques on image classification benchmarks, including several related to Earth Observation. Experimental results suggests that SSL pretraining, alone or in combination with fault mitigation techniques, generally improves DNNs' fault tolerance, although the performance gap vary among datasets and SSL techniques.
Rosario Milazzo, Vincenzo De Marco, Corrado De Sio, Sophie M. Fosson, Lia Morra, Luca Sterpone
DDECS6
2024 Toward Fault-Tolerant Applications on Reconfigurable Systems-on-Chip
abstract
FPGAs have become a well-established solution for systems aiming for high performance and flexibility. FPGA-based SoCs have facilitated the integration of software programmability and custom hardware acceleration. The sensitivity to disturbance, the scarcity of CAD tools dedicated to evaluating robustness, and the drastic increase in on-chip components pose challenges to analyzing and assessing the reliability of reconfigurable systems in safety-critical domains. The current work proposes methodologies for accurate and efficient robustness analysis of Reconfigurable SoCs. It focuses on the heterogeneous components and modules embedded in Reconfigurable SoCs, such as soft and hard processors, host-device interfacing systems, and custom hardware accelerators. The research explores and provides the methodology and practical tools for developing and evaluating reliable applications on Reconfigurable SoCs, enabling detailed analyses of systems and their components.
Corrado De Sio, Luca Sterpone
ITC2
2024 CNN-Oriented Placement Algorithm for High-Performance Accelerators on Rad-Hard FPGAs
abstract
Convolutional Neural Networks (CNNs) are quickly becoming one of the most common applications running on hardware accelerators. Considering Field Programmable Gate Arrays (FPGAs), due to their high flexibility and computational performance, they are suitable for fast classification tasks and therefore, pave the way for new machine learning inference approaches. In this work, we first designed a fully interconnected CNN architecture implementable on a single FPGA. Secondly, we developed a new Neural Node-oriented placement algorithm to enable resilient CNN accelerators on space-grade FPGAs. The proposed solution reduces the single event transient error sensitivity of CNN single neuron cores while achieving high performance and effective overall convolutional architecture fault tolerance. The developed approach has been applied and integrated into a state-of-the-art Radiation Tolerant FPGAs (RTG4) implementation flow. The experimental evaluation has been performed on a Microchip test board through benchmark application performance evaluation and transient error analysis. Experimental results demonstrate an improvement of 27.2% of the maximal working frequency and a reduction of the transient error sensitivity of about three times with respect to the previous mitigation approaches.
Luca Sterpone, Sarah Azimi, Corrado De Sio
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Assessing the Robustness of Real-Time Operating System on Soft Processor against Multiple Bit Upset
abstract
Field-Programmable Gate Arrays (FPGAs) are becoming increasingly important for space applications due to their high flexibility, performance, and complexity. In particular, soft-core processors like Xilinx Microblaze are commonly implemented using the programmable logic of FPGAs, making them suitable for embedded applications. However, the netlist of soft microprocessors can be corrupted by the effects of radiation-induced soft errors, in particular Single-Event Effects (SEUs) [1--5]. This work presents a detailed evaluation of the impact of radiation-induced architectural faults affecting the application benchmarks running on the FreeRTOS of the Microblaze embedded soft processor. The effect of Multiple Bits Upsets (MBU) faults during the execution of different software applications on FreeRTOS supported by Microblaze implemented on Zynq-7020 FPGA was evaluated, and the outcome of the software application was investigated. Please, notice that the developed platform is targeting only hardware faults and their impact on the execution of the software running in the operating system.
Andrea Portaluri, Corrado De Sio, Luca Sterpone
CF3
2023 Reliability Analysis of Microarchitectural Faults in GPGPU-based HPC Systems
abstract
As GPGPUs gain popularity in HPC applications, there is a growing need to investigate their reliability for performance improvement and reduced computation overhead. In this paper, the authors propose a novel fault injection environment for NVIDIA GPGPU devices that can automatically inject faults into instructions at the SASS level by instrumenting the CUDA binary executable file. It can categorize faults into Silent Data Corruption, Detected Unrecoverable Error, and Hang, making it an effective tool for targeting the reliability evaluation of specific threads.
Corrado De Sio, Luca Sterpone, Sarah Azimi
CF2
2023 A Framework for Uniformly Analyze and Mitigate Radiation-effects on FPGAs for Aerospace
abstract
This paper describes an architecture-modulable FPGA framework comprising synthesis, mapping, place and route, and bitstream analysis and mitigation for circuits mapped on FPGAs suitable for aerospace applications. The framework has several benefits, including analysis of soft-error effects and comparison between different vendors and parts, compatibility with commercial and radiation-hardened FPGA, and the ability to individuate single point of failure and embeds different mitigation strategies such as Triple Modular Redundancy (TMR) and Single Event Transient (SET) filtering at the synthesis or the place and route level. In this work, we provide the description of the framework and the radiation-effects analysis and mitigation on a set of benchmark circuits implemented using the most recent FPGA devices families for aerospace such as AMD Xilinx Ultrascale, NanoXplore NG-Medium, Microchip Radiation-Tolerant ProASIC3, and Radiation-Tolerant G4. Finally, we present a comparative analysis of a benchmark circuit's performance and radiation sensitivity mapped on the different FPGA device manufacturers.
Luca Sterpone, Sarah Azimi, Corrado De Sio
CF1
2023 EuFRATE: European FPGA Radiation-hardened Architecture for Telecommunications
abstract
The EuFRATE project aims to research, develop and test radiation-hardening methods for telecommunication payloads deployed for Geostationary-Earth Orbit (GEO) using Commercial-Off- The-Shelf Field Programmable Gate Arrays (FPGAs). This project is conducted by Argotec Group (Italy) with the collaboration of two partners: Politecnico di Torino (Italy) and Technische Universität Dresden (Germany). The idea of the project focuses on high-performance telecommunication algorithms and the design and implementation strategies for connecting an FPGA device into a robust and efficient cluster of multi-FPGA systems. The radiation-hardening techniques currently under development are addressing both device and cluster levels, with redundant datapaths on multiple devices, comparing the results and isolating fatal errors. This paper introduces the current state of the project's hardware design description, the composition of the FPGA cluster node, the proposed cluster topology, and the radiation hardening techniques. Intermediate stage experimental results of the FPGA communication layer performance and fault detection techniques are presented. Finally, a wide summary of the project's impact on the scientific community is provided.1
Ludovica Bozzoli, Antonino Catanese, Emilio Fazzoletto, Eugenio Scarpa, Diana Göhringer, Sergio A. Pertuz 0001, Lester Kalms, Cornelia Wulf, Najdet Charaf, Luca Sterpone, Sarah Azimi, Daniele Rizzieri, Salvatore Gabriele La Greca, David Merodio Codinachs
DATE10
2023 Assessing Convolutional Neural Networks Reliability through Statistical Fault Injections
abstract
Assessing the reliability of modern devices running CNN algorithms is a very difficult task. Actually, the complexity of the state-of-the-art devices makes exhaustive Fault Injection (FI) campaigns impractical and typically out of the computational capabilities. A possible solution consists of resorting to statistical FI campaigns that allow a reduction in the number of needed experiments by injecting only a carefully selected small part of it. Under specific hypothesis, statistical FIs guarantee an accurate picture of the problem, albeit selecting a reduced sample size. The main problems today are related to the choice of the sample size, the location of the faults, and the correct understanding of the statistical assumptions. The intent of this paper is twofold: first, we describe how to correctly specify statistical FIs for Convolutional Neural Networks; second, we propose a data analysis on the CNN parameters that drastically reduces the number of FIs needed to achieve statistically significant results without compromising the validity of the proposed method. The methodology is experimentally validated on two CNNs, ResNet-20 and MobileNetV2, and the results show that a statistical FI campaign on about 1.21% and 0.55% of the possible faults, provides very precise information of the CNN reliability. The statistical results have been confirmed by the exhaustive FI campaigns on the same cases of study.
Annachiara Ruospo, Gabriele Gavarini, Corrado De Sio, Juan-David Guerrero-Balaguera, Luca Sterpone, Matteo Sonza Reorda, Ernesto Sánchez 0001, Riccardo Mariani, Joseph Aribido, Jyotika Athavale
DATE5
2023 A Comprehensive Analysis of Transient Errors on Systolic Arrays
abstract
In recent years, the growth of interest in adopting deep neural network techniques across various domains led to new architectures for supporting the required computational effort. Tensor Processing Units (TPUs), which are based on a systolic array matrix multiplication unit (MMU), became widely popular thanks to their specific structure suitable for Artificial Intelligence. This work investigates Single Event Transient (SET) effects on TPU’s MMU. The analysis demonstrates the impact of SETs on the functionality of MMU when executing digital image processing filtering. The experimental results identify the static and dynamic SET sensitivity of TPU and depict meaningful information on the data dependency of the filters’ kernel values.
Eleonora Vacca, Sarah Azimi, Luca Sterpone
DDECS3
2023 RunSAFER: A Novel Runtime Fault Detection Approach for Systolic Array Accelerators
abstract
In this study, we introduce a new runtime fault detection technique for systolic array accelerators oriented to neural network applications. The method exploits the functional path of the systolic array to compute and process checksum values during the execution of the current application instructions flow and integrates self-testing capabilities within the systolic array Instruction Set Architecture. The proposed technique does not require additional hardware self-testing modules, and the test pattern penalty is limited to 3 clock cycles independent of the size of the systolic array. Experimental analysis performed with fault injection campaigns demonstrates full fault detection capabilities of stuck-at faults with an average computing overhead 4 times lower than state-of-the-art solutions. Additionally, our approach exhibits diminished hardware overhead in contrast to conventional techniques.
Eleonora Vacca, Giorgio Ajmone, Luca Sterpone
ICCD3
2023 Design Techniques for Multi-Core Neural Network Accelerators on Radiation-Hardened FPGAs
abstract
Radiation-Hardened-By-Design (RHBD) FPGAs have gained a lot of attention thanks to their excellent compromise between costs and performance. Being of very limited use due to a lack of performance a few years ago, these devices are now capable of implementing a wide range of applications requiring high computational capabilities.This work describes an implementation of a Very Long Instruction Word (VLIW) soft-core convolutional accelerator in the NanoXplore RHBD NG-Medium chip. Feasibility and timing performances have been analyzed in order to discover whether and how multi-core solutions can affect parallel acceleration. Placement also showed to heavily affect the delays, up to 70% more, based on the proximity to the output buffers.
Andrea Portaluri, Sarah Azimi, Luca Sterpone
ISPDC3
2023 Enhanced Video Surveillance Systems for "Signal for Help" Detection on Edge Devices
abstract
The COVID-19 pandemic triggered a concerning rise in violence against women and children, known as The Shadow Pandemic. To address this, a Canadian foundation introduced the “Signal for Help” gesture to discreetly alert others in danger. However, the effectiveness of this approach depends on individuals recognizing and responding to the signal. In this paper, we propose an innovative solution that adopts the technology available in smart cities to detect the “Signal for Help” in real-time through surveillance footage. We developed and implemented a recognition algorithm on an affordable device that achieves accurate detection of the signal in 94 % of cases. This approach has the potential to improve the response to instances of violence, providing a reliable means of alerting authorities and support networks.
Sarah Azimi, Corrado De Sio, Luca Sterpone
ISTAS3
2022 Test, Reliability and Functional Safety Trends for Automotive System-on-Chip
abstract
This paper encompasses three contributions by industry professionals and university researchers. The contributions describe different trends in automotive products, including both manufacturing test and run-time reliability strategies. The subjects considered in this session deal with critical factors, from optimizing the final test before shipment to market to in-field reliability during operative life.
Francesco Angione, Davide Appello, Joseph Aribido, Jyotika Athavale, Nicolò Bellarmino, Paolo Bernardi 0002, Riccardo Cantoro, Corrado De Sio, Tommaso Foscale, Gabriele Gavarini, Juan-David Guerrero-Balaguera, Martin Huch, Giusy Iaria, Tobias Kilian, Riccardo Mariani, Raffaele Martone, Annachiara Ruospo, Ernesto Sánchez 0001, Ulf Schlichtmann, Giovanni Squillero, Matteo Sonza Reorda, Luca Sterpone, Vincenzo Tancorre, Roberto Ugioli
ETS22
2022 Radiation-induced Effects on DMA Data Transfer in Reconfigurable Devices
abstract
As the adoption of SRAM-based FPGAs and Reconfigurable SoCs for High-Performance Computing increased in the last years, the use of Direct Memory Access for data transfer becomes a key feature of many reconfigurable applications even in the space industry. For such kinds of applications, radiation-induced effects are a serious issue that mines the correctness and success of mission-critical tasks. In this paper, we evaluate the effects of proton-induced errors on a DMA-based application implemented on a Xilinx Zynq-7020 FPGA in order to quantify the robustness of this module in a typical hardware-accelerated configuration. The obtained results confirm the high criticality of the DMA module on programmable logic. Moreover, the Multiple Bits Upsets effect has been evaluated. The most recurring patterns have been reported in order to provide further tools to better characterize the behavior of these systems under future fault injection campaigns, as demonstrated in the experimental results.
Andrea Portaluri, Sarah Azimi, Corrado De Sio, Luca Sterpone, David Merodio Codinachs
IOLTS4
2022 Analysis and Mitigation of Soft-Errors on High Performance Embedded GPUs
abstract
Multiprocessor system-on-chip such as embedded GPUs are becoming very popular in safety-critical applications, such as autonomous and semi-autonomous vehicles. However, these devices can suffer from the effects of soft-errors, such as those produced by radiation effects. These effects are able to generate unpredictable misbehaviors. Fault tolerance oriented to multi-threaded software introduces severe performance degradations due to the redundancy, voting and correction threads operations. In this paper, we propose a new fault injection environment for NVIDIA GPGPU devices and a fault tolerance approach based on error detection and correction threads executed during data transfer operations on embedded GPUs. The fault injection environment is capable of automatically injecting faults into the instructions at SASS level by instrumenting the CUDA binary executable file. The mitigation approach is based on concurrent error detection threads running simultaneously with the memory stream device to host data transfer operations. With several benchmark applications, we evaluate the impact of soft- errors classifying Silent Data Corruption, Detection, Unrecoverable Error and Hang. Finally, the proposed mitigation approach has been validated by soft-error fault injection campaigns on an NVIDIA Pascal Architecture GPU controlled by Quad-Core A57 ARM processor (JETSON TX2) demonstrating an advantage of more than 37% with respect to state of the art solution.
Luca Sterpone, Sarah Azimi, Corrado De Sio, Filippo Parisi
ISPDC1
2022 Evaluating low-level software-based hardening techniques for configurable GPU architectures
Marcio Gonçalves, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Luca Sterpone, José Rodrigo Azambuja
J. Supercomput.4
2021 A 3-D LUT Design for Transient Error Detection Via Inter-Tier In-Silicon Radiation Sensor
abstract
Three-dimensional Integrated Circuits (3-D ICs) have gained much attention as a promising approach to increase IC performance due to their several advantages in terms of integration density, power dissipation, and achievable clock frequencies. However, achieving a 3-D ICs resilient to soft errors resulting from radiation effects is a challenging problem. Traditional Radiation-Hardened-by-Design (RHBD) techniques are costly in terms of area, power, and performance overheads. In this work, we propose a new 3-D LUT design integrating error detection capabilities. The LUT has been designed on a two tiers IC model improving radiation resiliency by selective upsizing of sensitive transistors. Besides, an in-silicon radiation sensor adopting inverters chain has been implemented within the free volume of the 3-D structure. The proposed design shows a 37% reduction in sensitivity to SETs and an effective error detection rate of 83% without introducing any area overhead.
Sarah Azimi, Corrado De Sio, Luca Sterpone
DATE3
2021 A New Domains-based Isolation Design Flow for Reconfigurable SoCs
abstract
Reconfigurable SoCs are widely adopted in mission-critical tasks in aerospace and automotive. Though, one of their main drawbacks is the susceptibility to high-energy particles both in space and at sea level. Isolation Design Flow is a promising implementation approach to improve the reliability of circuits. However, considering the high number of modules in a complex circuit, especially when redundant techniques are applied, IDF requires a complex floorplanning stage. In this paper, the benefits of using IDF are evaluated, both for plain and hardened-by-redundancy designs. We propose an implementation methodology to tackle the complexity of applying IDF to TMR-based circuits that usually make the implementation approach unfeasible. The impact of different design policies on the reliability of the system is evaluated through fault injection campaigns. The proposed method is applied to the TMR-hardened CORDIC core implemented on Zynq AP-SoC and compared with other possible solutions. The results report a significant improvement in the TMR effectiveness when the proposed domains-based IDF is applied.
Andrea Portaluri, Corrado De Sio, Sarah Azimi, Luca Sterpone
IOLTS4
2021 On the Evaluation of SEEs on Open-Source Embedded Static RAMs
abstract
Static RAM modules are widely adopted in high performance systems. Single Event Effects (SEEs) resilient memories are required in many embedded systems applied in automotive and aerospace applications to increase their overall resiliency against SEEs. The current SEE resilient SRAM modules are obtained by applying radiation-hardened by design solutions which leads to elevated area overhead and difficulty to tune the resiliency capability with respect to the particle's radiation profile. To overcome these limitations, we propose a methodology for the analysis and mitigation of embedded SRAMs generated by the OpenRAM memory compiler. A technology-oriented radiation analysis tool is presented to support the interaction of the charged radiation particles with the SRAM layout and depict the sensitive transistors of the SRAM memory. A selective duplication of the sensitive transistors has been applied to the 6T-SRAM cell designed at the layout level. The designed cell is included in the OpenRAM compiler and used to generate a mitigated 8Kb SRAM-bank. We evaluated the SEEs sensitivity by comparative simulation-based radiation analysis observing a reduction more than 6 times with respect to the original 6T-SRAM cell for the SEE sensitivity at high energy heavy ions particles, with negligible degradation of operations margins and power consumption and area overhead of less than $\sim$ 4%.
Sarah Azimi, Corrado De Sio, Luca Sterpone
VLSI-SoC3
2021 DYRE: a DYnamic REconfigurable solution to increase GPGPU's reliability
abstract
Abstract General-purpose graphics processing units (GPGPUs) are extensively used in high-performance computing. However, it is well known that these devices’ reliability may be limited by the rising of faults at the hardware level. This work introduces a flexible solution to detect and mitigate permanent faults affecting the execution units in these parallel devices. The proposed solution is based on adding some spare modules to perform two in-field operations: detecting and mitigating faults. The solution takes advantage of the regularity of the execution units in the device to avoid significant design changes and reduce the overhead. The proposed solution was evaluated in terms of reliability improvement and area, performance, and power overhead costs. For this purpose, we resorted to a micro-architectural open-source GPGPU model (FlexGripPlus). Experimental results show that the proposed solution can extend the reliability by up to 57%, with overhead costs lower than 2% and 8% in area and power, respectively.
Josie E. Rodriguez Condia, Pierpaolo Narducci, Matteo Sonza Reorda, Luca Sterpone
J. Supercomput.4
2021 A Radiation-Hardened CMOS Full-Adder Based on Layout Selective Transistor Duplication
abstract
Single event transients (SETs) have become increasingly problematic for modern CMOS circuits due to the continuous scaling of feature sizes and higher operating frequencies. Especially when involving safety-critical or radiation-exposed applications, the circuits must be designed using hardening techniques. In this brief, we present a new radiation-hardened-by-design full-adder cell on 45-nm technology. The proposed design is hardened against transient errors by selective duplication of sensitive transistors based on a comprehensive radiation-sensitivity analysis. Experimental results show a 62% reduction in the SET sensitivity of the proposed design with respect to the unhardened one. Moreover, the proposed hardening technique leads to improvement in performance and power overhead and zero area overhead with respect to the state-of-the-art techniques applied to the unhardened full-adder cell.
Sarah Azimi, Corrado De Sio, Luca Sterpone
IEEE Trans. Very Large Scale Integr. Syst.3
2020 RESCUE: Interdependent Challenges of Reliability, Security and Quality in Nanoelectronic Systems
abstract
The recent trends for nanoelectronic computing systems include machine-to-machine communication in the era of Internet-of-Things (IoT) and autonomous systems, complex safety-critical applications, extreme miniaturization of implementation technologies and intensive interaction with the physical world. These set tough requirements on mutually dependent extra-functional design aspects. The H2020 MSCAITN project RESCUE is focused on key challenges for reliability, security and quality, as well as related electronic design automation tools and methodologies. The objectives include both research advancements and cross-sectoral training of a new generation of interdisciplinary researchers. Notable interdisciplinary collaborative research results for the first halfperiod include novel approaches for test generation, soft-error and transient faults vulnerability analysis, cross-layer fault-tolerance and error-resilience, functional safety validation, reliability assessment and run-time management, HW security enhancement and initial implementation of these into holistic EDA tools.
Maksim Jenihhin, Said Hamdioui, Matteo Sonza Reorda, Milos Krstic, Peter Langendörfer, Christian Sauer 0001, Anton Klotz, Michael Hübner 0001, Jörg Nolte, Heinrich Theodor Vierhaus, Georgios N. Selimis, Dan Alexandrescu, Mottaqiallah Taouil, Geert Jan Schrijen, Jaan Raik, Luca Sterpone, Giovanni Squillero, Zoya Dyka
DATE16
2020 A dynamic hardware redundancy mechanism for the in-field fault detection in cores of GPGPUs
abstract
In the past, in most General-Purpose Graphic Processing Units (GPGPUs) application fields (e.g., multimedia and gaming), the reliability features were not so relevant. Nowadays, GPGPUs are used in new domains, such as the automotive one, where reliability plays a significant role. In this work, we describe a dynamic duplication with a comparison (DDWC) mechanism intended to harden the Scalar Processor (SP) units located in the Streaming multiprocessors (SM) of a GPGPU. The proposed mechanism targets the permanent faults that may arise inside the SPs. One additional SP unit is included in the system to compute redundantly the same operations of a selected SP. Results are compared, and possible failures detected. A custom reconfiguration instruction allows the dynamic selection of the target SP to be monitored. Experimental results show that the proposed mechanism introduces a limited area overhead while it provides a significant increase in the in-field fault detection capabilities of the GPGPU. Its flexibility allows selecting the best trade-off between fault detection latency and performance overhead.
Josie E. Rodriguez Condia, Pierpaolo Narducci, Matteo Sonza Reorda, Luca Sterpone
DDECS4
2020 In-Circuit Mitigation Approach of Single Event Transients for 45nm Flip-Flops
abstract
Nowadays, radiation-induced Single Event Transients are a leading cause of critical errors in CMOS nanometric integrated circuits. In this work, we propose a workflow for analyzing and mitigating nanometric CMOS integrated circuits to radiation-induced transient errors. The analysis phase starts with the developed Rad-Ray tool for mimicking the passage of the radiation particles through the silicon matter of the cells to identify the features of the generated transient pulses. The tool is integrated with an electrical simulator to evaluate the dynamic behavior of the transient pulses inserted and propagated in the circuit. A tunable mitigation solution is proposed by inserting the filtering block before the storage element, tuned based on the duration and amplitude of the expected transient pulse, identified in the analysis phase. Experimental results are achieved by applying the proposed approach on the 45 nm Flip-Flop component available in the FreePDK design kit, comparing the Dynamic Error Rate for the original Flip-Flop and the mitigated one which shows a reduction of sensitivity up to 56% with respect of the original version, with negligible degradation of performances.
Sarah Azimi, Corrado De Sio, Luca Sterpone
IOLTS3
2020 Machine Learning Clustering Techniques for Selective Mitigation of Critical Design Features
abstract
Selective mitigation or selective hardening is an effective technique to obtain a good trade-off between the improvements in the overall reliability of a circuit and the hardware overhead induced by the hardening techniques. Selective mitigation relies on preferentially protecting circuit instances according to their susceptibility and criticality. However, ranking circuit parts in terms of vulnerability usually requires computationally intensive fault-injection simulation campaigns. This paper presents a new methodology which uses machine learning clustering techniques to group flip-flops with similar expected contributions to the overall functional failure rate, based on the analysis of a compact set of features combining attributes from static elements and dynamic elements. Fault simulation campaigns can then be executed on a per-group basis, significantly reducing the time and cost of the evaluation. The effectiveness of grouping similar sensitive flip-flops by machine learning clustering algorithms is evaluated on a practical example.Different clustering algorithms are applied and the results are compared to an ideal selective mitigation obtained by exhaustive fault-injection simulation.
Thomas Lange, Aneesh Balakrishnan, Maximilien Glorieux, Dan Alexandrescu, Luca Sterpone
IOLTS5
2020 Digital Design Techniques for Dependable High Performance Computing
abstract
As today's process technologies continuously scale down, circuits become increasingly more vulnerable to radiation-induced soft errors in nanoscale VLSI technologies. The reduction of node capacitance and supply voltages coupled with increasingly denser chips are raising soft error rates and making them an important design issue. This research work is focused on the development of design techniques for high-reliability modern VLSI technologies, focusing mainly on Radiation-induced Single Event Transient. In this work, we evaluate the complete life-cycle of the SET pulse from the generation to the mitigation. A new simulation tool, Rad-Ray, has been developed to simulate and model the passage of heavy ion into the silicon matter of modern Integrated Circuit and predict the transient voltage pulse taking into account the physical description of the design. An analysis and mitigation tool has been developed to evaluate the propagation of the predicted SET pulses within the circuit and apply a selective mitigation technique to the sensitive nodes of the circuit. The analysis and mitigation tools have been applied to many industrial projects as well as the EUCLID space mission project, including more than ten modules. The obtained results demonstrated the effectiveness of the proposed tools.
Sarah Azimi, Luca Sterpone
ITC2
2020 A dynamic reconfiguration mechanism to increase the reliability of GPGPUs
abstract
General Purpose Graphic Processing Units (GPGPUs) are effective solutions for high-demanding data processing applications. Recently, they started to be used even in safety-critical applications, such as autonomous car driving systems. GPGPUs are implemented using the latest semiconductor technologies, which are more prone to faults arising during the lifetime operation. However, until now fault mitigation solutions were not extensively included in GPGPUs, due to the limited reliability requirements of the applications they were originally intended for (e.g., gaming or multimedia). This work proposes a dynamically configurable self- repairing mechanism aimed at mitigating the impact of permanent faults in the Scalar Processor (SP) cores in GPGPUs. The mechanism is based on spare modules that can be used to replace faulty SPs when a fault is detected. A configuration instruction allows dynamically controlling in software the selection of the set of active SPs in the SM. The method is extremely flexible since it does not require any change in the application software. Experimental results show that the solution introduces a moderate area overhead while allowing continue working even in the case of any permanent faults affecting the SPs.
Josie E. Rodriguez Condia, Pierpaolo Narducci, Matteo Sonza Reorda, Luca Sterpone
VTS4
2019 A new FPGA-based Detection Method for Spurious Variations in PCBA Power Distribution Network
abstract
Nowadays, increasing demand for High-Performance Systems produces significant growth in usage of Field Programmable Gate Arrays (FPGAs) for different applications thanks to their flexibility and high level of parallelism. Such systems rely on complex multi-layer Printed Circuit Board Assemblies (PCBA)with a few dozens of hidden layers, stacked microvias and high-density interconnects. Along with creating new test challenges, the increasing PCBA complexity elevates the criticality of defects in various subsystems. One of such sub-systems is a Power-Delivery-Network (PDN) with operating margin progressively reduced due to increasingly strict requirements of High-Performance applications. As a consequence, Marginal Defects and process variations in a PDN may create latent problems that will manifest in a particular condition thus compromising the overall system performance and causing malfunctions. In this paper we propose a new FPGA-based non-intrusive method to detect Marginal Defects in a PCBA PDN. The method is based on a monitoring circuit that measures signal delays caused by PDN variations and thus detects relevant anomalies. Additional ad-hoc PDN stress circuits have been developed to validate the measurement technique. Experimental results demonstrating the consistency of the proposed approach are obtained by comparing stress and non-stress scenarios.
Sergei Odintsov, Ludovica Bozzoli, Corrado De Sio, Luca Sterpone, Artur Jutman
DDECS4
2019 On the Evaluation of the PIPB Effect within SRAM-based FPGAs
abstract
SRAM-based FPGAs are widely used in mission critical applications. Due to the increasing working frequency and technology scaling of ultra-nanometer technology, Single Event Transients (SETs) are becoming a major source of errors for these devices. In this paper, we propose an approach for evaluating the Propagation-induced Pulse Broadening (PIPB) effect introduced by the logic resources traversed by transient pulses. The proposed methodology is applicable to any recent technology to provide SET analysis, necessary for an efficient mitigation technology. Experimental results achieved from a set of benchmarks are compared with fault injection experiments executed on a 28 nm SRAM-based FPGA to demonstrate the effectiveness of our technique.
Corrado De Sio, Sarah Azimi, Luca Sterpone
ETS3
2019 Machine Learning to Tackle the Challenges of Transient and Soft Errors in Complex Circuits
abstract
The Functional Failure Rate analysis of today's complex circuits is a difficult task and requires a significant investment in terms of human efforts, processing resources and tool licenses. Thereby, de-rating or vulnerability factors are a major instrument of failure analysis efforts. Usually computationally intensive fault-injection simulation campaigns are required to obtain a fine-grained reliability metrics for the functional level. Therefore, the use of machine learning algorithms to assist this procedure and thus, optimising and enhancing fault injection efforts, is investigated in this paper. Specifically, machine learning models are used to predict accurate per-instance Functional De-Rating data for the full list of circuit instances, an objective that is difficult to reach using classical methods. The described methodology uses a set of per-instance features, extracted through an analysis approach, combining static elements (cell properties, circuit structure, synthesis attributes) and dynamic elements (signal activity). Reference data is obtained through first-principles fault simulation approaches. One part of this reference dataset is used to train the machine learning model and the remaining is used to validate and benchmark the accuracy of the trained tool. The presented methodology is applied on a practical example and various machine learning models are evaluated and compared.
Thomas Lange, Aneesh Balakrishnan, Maximilien Glorieux, Dan Alexandrescu, Luca Sterpone
IOLTS5
2019 A new CAD tool for Single Event Transient Analysis and mitigation on Flash-based FPGAs
Sarah Azimi, Boyang Du, Luca Sterpone, David Merodio Codinachs, Raoul Grimoldi, L. Cattaneo
Integr.3
2018 On the mitigation of single event transients on flash-based FPGAs
abstract
Thanks to the immunity against Single Event Upsets in configuration memory, Flash-based FPGA is becoming widely adopted in mission- and safety-critical applications, such as in aerospace field. However, the decreasing of device feature size leads to an increasing of the device sensitivity regarding Single Event Transients (SETs). In this paper, we developed a new workflow to evaluate SET phenomena in a specific convergence case and introduce a new mitigation of SET pulse without introducing any performance penalization to the original netlist.
Sarah Azimi, Boyang Du, Luca Sterpone
ETS3
2018 About the functional test of the GPGPU scheduler
abstract
General Purpose Graphical Processing Units (GPGPUs) are increasingly used in safety critical applications such as the automotive ones. Hence, techniques are required to test them during the operational phase with respect to possible permanent faults arising when the device is already deployed in the field. Functional tests adopting Software-based Self-test (SBST) are an effective solution since they provide benefits in terms of intrusiveness, flexibility and test duration. While the development of the functional test code addressing the several computational cores composing a GPGPU can be done resorting to known methods developed for CPUs, for other modules which are typical of a GPGPU we still miss effective solutions. This paper focuses on one of the most relevant module consists on the scheduler core which is in charge of managing different scalar computational cores and the different executed threads. At first, we propose a method for evaluating the fault coverage that can be achieved using an application program. Then, we provide some guidelines for improving the achieved fault coverage. Experimental results are provided on an open-source VHDL model of a GPGPU.
Boyang Du, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Luca Sterpone
IOLTS4
2018 IbIS: Interface-based Interconnection Structure for Dynamically Reconfigurable FPGAs
abstract
Nowadays SRAM-based FPGAs, already widely used for their advantages in terms of size, flexibility and performances, are becoming even more attractive since they may dynamically change their functionalities based on the elaboration demand, thanks to Dynamic Partial Reconfiguration Feature. In these systems the communication infrastructure represents the major performance bottleneck due to routing congestions and to the needs to guarantee signal integrity at the module boundaries. In this paper we propose an Interface-based communication architecture, which simplify the interaction mechanism and the DRPM architecture, reducing both delay and resources overhead with respect to the state-of-the-art solutions.
Ludovica Bozzoli, Luca Sterpone
ISCAS2
2018 PyXEL: An Integrated Environment for the Analysis of Fault Effects in SRAM-Based FPGA Routing
abstract
In the last decades, FPGAs have been increasingly used in many different mission critical applications, such as the avionics and aerospace ones. Thus, research interest in studying faults in FPGAs has seen a sharp increase, especially for those applications that require high dependability and must operate in harsh environments. The increase of resources available in FPGA devices has caused a huge growth in routing complexity. Nowadays, more than 80% of transistors in modern FPGAs are related to the routing infrastructure. The analysis of faults related to routing structure of FPGA devices is a hard task due to the lack of tools working at low-level, limited information availability about interconnection structure from vendors and, above all, no automated testing workflow for such kind of resources. In this paper, we introduce PyXEL, an integrated environment realized to automatize the analysis of fault effects in FPGAs routing structure. PyXEL is a Python-based framework that allows to easily manipulate FPGAs bitstreams in order to inject specific faults and to analyze their behavior. Moreover, PyXEL provides an easy way to build and run experimental workflow interacting directly with Xilinx Vivado and ISE allowing to select routing resources to test and logically analyze results. We demonstrated the feasibility and the advantages of our approach exploiting PyXEL to gain insight into the electrical effects of faults in the routing interconnections of the Xilinx Artix-7.
Ludovica Bozzoli, Corrado De Sio, Luca Sterpone, Cinzia Bernardeschi
RSP3
2017 Effective Mitigation of Radiation-induced Single Event Transient on Flash-based FPGAs
abstract
Due to the decreasing feature sizes of VLSI circuits, radiation induced Single Event Transients (SETs) are increasingly dominating the event ratio on modern VLSI devices. In particular, Flash-based FPGAs are characterized by the main concern of radiation-induced voltage glitches or SETs in the combinational logic. Transient pulses can be sampled by a storage element and can propagate through the circuit up to the outputs and leading to an error. In this paper, we propose a complete implementation flow including sensitivity analysis, fault tolerant mapping and fault tolerance-oriented place and route for the effective design of SET tolerant circuits on Flash-based FPGAs. In details, the proposed method allows accurate measurement of the transient pulse source induced by radiation particles and estimation of the SET error rate on the overall circuit. Besides the developed method provides a netlist mapping and place and route tool for the selective mitigation of SET effects. The proposed method has been applied to an industrial design oriented to the Euclid European Space Agency mission including more than ten different modules. The obtained results show an improvement of the total filtering capability of around 43 times with respect to the original netlist without affecting the timing constraints of the circuit.
Luca Sterpone, Sarah Azimi, Boyang Du, David Merodio Codinachs, Raoul Grimoldi
ACM Great Lakes Symposium on VLSI1
2017 Analysis of radiation-induced cross domain errors in TMR architectures on SRAM-based FPGAs
abstract
SRAM-Based FPGAs represent a low-cost alternative to ASIC device thanks to their high performance and design flexibility. In particular, for aerospace and avionics application fields, SRAM-based FPGAs are increasingly adopted for their configurability features making them a viable solution for long-time applications. However, these fields are characterized by a radiation environment that makes the technology extremely sensitive to radiation-induced Single Event Upsets (SEUs) in the SRAM-based FPGA's configuration memory. Configuration scrubbing and Triple Modular Redundancy (TMR) have been widely adopted in order to cope with SEU effects. However, modern FPGA devices are characterized by a heterogeneous routing resource distribution and a complex configuration memory mapping causing an increasing sensitivity to Cross Domain Errors affecting the TMR structure. In this paper we developed a new methodology to calculate the reliability of TMR architecture considering the intrinsic characteristics of the new generation of SRAM-based FPGAs. The method includes the analysis of the configuration bit sharing phenomena and of the routing long lines. We experimentally evaluate the method of various benchmark circuits evaluating the Mean Upset To Failure (MUTF). Finally, we used the results of the developed method to implement an improved design achieving 29x improvement of the MUTF.
Luca Sterpone, Luca Boragno
IOLTS1
2017 Fault tolerant electronic system design
abstract
Due to technology scaling, which means smaller transistor, lower voltage and more aggressive clock frequency, VLSI devices are becoming more susceptible against soft errors. Especially for those devices deployed in safety- and mission-critical applications, dependability and reliability are becoming increasingly important constraints during the development of system on/around them. Other phenomena (e.g. aging and wear-out effects) also have negative impacts on reliability of modern circuits. Furthermore, as recent researches show that even at sea level, radiation particles can still induce soft errors in electronic systems, for avionic and space applications, certain fault tolerant strategy must be applied to guarantee system reliability throughout application lifetime. In this paper, we focus on two aspects: testing for System-on-Chip/System-on-Programmable-Chip by exploiting debug infrastructures and analysis and mitigation of Single Event Effects on FPGA devices.
Boyang Du, Luca Sterpone
ITC2
2017 Evaluation of transient errors in GPGPUs for safety critical applications: An effective simulation-based fault injection environment
Sarah Azimi, Boyang Du, Luca Sterpone
J. Syst. Archit.3
2017 A novel tool-flow for zero-overhead cross-domain error resilient partially reconfigurable X-TMR for SRAM-based FPGAs
Luis Andrés Cardona, Anees Ullah, Luca Sterpone, Carles Ferrer 0001
J. Syst. Archit.3
2017 An Error-Detection and Self-Repairing Method for Dynamically and Partially Reconfigurable Systems
abstract
Reconfigurable systems are gaining an increasing interest in the domain of safety-critical applications, for example in the space and avionic domains. In fact, the capability of reconfiguring the system during run-time execution and the high computational power of modern Field Programmable Gate Arrays (FPGAs) make these devices suitable for intensive data processing tasks. Moreover, such systems must also guarantee the abilities of self-awareness, self-diagnosis and self-repair in order to cope with errors due to the harsh conditions typically existing in some environments. In this paperwe propose a self-repairing method for partially and dynamically reconfigurable systems applied at a fine-grain granularity level. Our method is able to detect correct and recover errors using the run-time capabilities offered by modern SRAM-based FPGAs. Fault injection campaigns have been executed on a dynamically reconfigurable system embedding a number of benchmark circuits. Experimental results demonstrate that our method achieves full detection of single and multiple errors, while significantly improving the system availability with respect to traditional error detection and correction methods.
Matteo Sonza Reorda, Luca Sterpone, Anees Ullah
IEEE Trans. Computers2
2016 A new EDA flow for the mitigation of SEUs in dynamic reconfigurable FPGAs
abstract
This work presents a new EDA flow that aims to increase the design robustness versus transient errors when the dynamic reconfigurable computing paradigm is adopted. In brief, we propose a modification of the existing commercial tool-chain flow to make transient error aware designs. Aiming at that scope, a new algorithm for the design mapping has been developed reducing Single Event Upsets on the routing interactions between reconfigurable placed modules. The performance evaluation of the EDA flow has been evaluated with neutron-based radiation test experiments and fault injection using a proper dynamic reconfiguration context. Results prove a reduction of the transient error sensitivity about 3 orders of magnitude without any area overhead and with a performance degradation of less than 10% on the average.
Boyang Du, Luca Sterpone, David Merodio Codinachs
ETS2
2016 Scalable FPGA graph model to detect routing faults
abstract
The SRAM cells that form the configuration memory of an SRAM-based FPGA make such FPGAs particularly vulnerable to soft errors. A soft error occurs when ionizing radiation corrupts the data stored in a circuit. The error persists until new data is written. Soft errors have long been recognized as a potential problem as radiation can come from a variety of sources. This paper presents an FPGA fault model focusing on routing aspects. A graph model of SRAM nodes behavior in case of fault, starting from netlist description of well known FPGA models, is presented. It is also performed a classification of possible logical effects of a soft error in the configuration bit controlling, providing statistics on the possible numbers of faults. Finally it is reported the definition of fault metrics computed on a set of complex benchmarks proving the effectiveness of our approach.
Luca Sterpone, Gianpiero Cabodi, Sebastiano F. Finocchiaro, Carmelo Loiacono, Francesco Savarese, Boyang Du
IOLTS1
2016 An FPGA-based testing platform for the validation of automotive powertrain ECU
abstract
Over the past decade, the complexity of electronic devices in the automotive systems increased significantly. The modern high level vehicles include more than 70 Electronic Control Units (ECUs) aimed at managing the powertrain of the vehicle, and improving passengers' comfort and safety. ECU microcontrollers aimed at the control of the fuel injection system have a key role. In this paper we present a new FPGA-based platform able to supervise and validate Commercial-Off-The-Shelf timer modules used in today state-of-the-art software applications for automotive fuel injection system with an accuracy improvement of more than 20% with respect to traditional approach. The proposed approach allows an effective and accurate validation of timing signals and it has two main advantages: can be customized with the exact timing module configurations to meet the exigency of new tests and allows effective modularity design test. As case study two industrial Time Modules manufactured by Freescale and Bosch have been used. The experimental analysis demonstrates the capability of the proposed approach providing a timing and angular precision of 10 ns and 10−5degrees respectively.
Boyang Du, Luca Sterpone
VLSI-SoC2
2016 UA2TPG: An untestability analyzer and test pattern generator for SEUs in the configuration memory of SRAM-based FPGAs
Cinzia Bernardeschi, Luca Cassano, Andrea Domenici, Luca Sterpone
Integr.4
2016 Online Test of Control Flow Errors: A New Debug Interface-Based Approach
abstract
Detecting the effects of transient faults is a key point in many processor-based safety-critical applications. This paper proposes to adopt the debug interface module existing today in several processors/controllers available on the market. In this way, we can achieve a good detection capability and small latency with respect to control flow errors, while the cost for adopting the proposed technique is rather limited and does not involve any change either in the processor hardware or in the application software. The method works even if the processor uses caches and we experimentally evaluated its characteristics demonstrating the advantages and showing the limitations on two pipelined processors. Experimental results performed by fault injection using different software applications demonstrate that the method is able to archieve high fault coverage (more than 95 percent in nearly all the considered cases) with a limited cost in terms of area and performance degradation.
Boyang Du, Matteo Sonza Reorda, Luca Sterpone, Luis Parra, Marta Portela-García, Almudena Lindoso, Luis Entrena
IEEE Trans. Computers3
2014 Fault injection and fault tolerance methodologies for assessing device robustness and mitigating against ionizing radiation
abstract
Traditionally, heavy ion radiation effects affecting digital systems working in safety critical application systems has been of huge interest. Nowadays, due to the shrinking technology process, Integrated Circuits became sensitive also to other kinds of radiation particles such as neutron that can exist at the earth surface and affects ground-level safety critical applications such as automotive or medical systems. The process of analyzing and hardening digital devices against soft errors implies rising the final cost due to time expensive fault injection campaigns and radiation tests, as well as reducing system performance due to the insertion of redundancy-based mitigation solutions. The main industrial problem arising is the localization of the critical elements in the circuit in order to apply optimal mitigation techniques. The proposal of this tutorial is to present and discuss different solutions currently available for assessing and implementing the fault tolerance of digital circuits, not only when the unique design description is provided but also at the component level, especially when Commercial-of-the-shelf (COTS) devices are selected.
Dan Alexandrescu, Luca Sterpone, Celia López-Ongil
ETS2
2014 Reconfigurable high performance architectures: How much are they ready for safety-critical applications?
abstract
Reconfigurable architectures are increasingly employed in a large range of embedded applications, mainly due to their ability to provide high performance and high flexibility, combined with the possibility to be tuned according to the specific task they address. Reconfigurable systems are today used in several application areas, and are also suitable for systems employed in safety-critical environments. The actual development trend in this area is focused on the usage of the reconfigurable features to improve the fault tolerance and the self-test and the self-repair capabilities of the considered systems. The state-of-the-art of the reconfigurable systems is today represented by Very Long Instruction Word (VLIW) processors and reconfigurable systems based on partially reconfigurable SRAM-based FPGAs. In this paper, we present an overview and accurate analysis of these two type of reconfigurable systems. The content of the paper is focused on analyzing design features, fail-safe and reconfigurable features oriented to self-adaptive mitigation and redundancy approaches applied during the design phase. Experimental results reporting a clear status of the test data and fault tolerance robustness are detailed and commented.
Davide Sabena, Luca Sterpone, Mario Schölzel, Tobias Koal, Heinrich Theodor Vierhaus, S. Wong, Robért Glein, Florian Rittner, C. Stender, Mario Porrmann, Jens Hagemeyer
ETS2
2014 Analysis and mitigation of single event effects on flash-based FPGAS
abstract
In the present paper, we propose a new design flow for the analysis and the implementation of circuits on Flash-based FPGAs hardened against Single Event Effects (SEEs). The solution we developed is based on two phases: 1) an analyzer algorithm able to evaluate the propagations of SETs through logic gates; 2) a hardening algorithm able to place and route a circuit by means of optimal electrical filtering and selective guard gates insertions. The effectiveness of the proposed design flow has been evaluated by performing hardening on seven benchmark circuits and comparing the results using different implementation approaches on 130nm Flash-based technology. The obtained results have been validated against radiation-beam testing using heavy-ions and demonstrated that our solution is able to decrease the circuits sensitivity versus SEE by two orders of magnitude with a reduction of resource overhead of 83 % with respect to traditional mitigation approaches.
Luca Sterpone, Boyang Du
ETS1
2014 Effective emulation of permanent faults in ASICs through dynamically reconfigurable FPGAs
abstract
Hardware fault emulation for Application Specific Integrated Circuits (ASICs) on FPGAs can considerably reduce the time required for the fault simulation. This paper presents a methodology to emulate ASIC faults on state-of-the-art FPGAs. The fault emulation is achieved by following a fully automated process consisting of: constrained technology mapping of ASIC net-list; creation of fault dictionary, generation of faulty partial bit-streams and fault emulation. The proposed approach exploits run-time partial reconfiguration techniques for fault injection and avoids full net-list re-compilations. The method's feasibility is assessed through carefully selected circuits and overhead in terms of area and timing is reported.
Ernesto Sánchez 0001, Luca Sterpone, Anees Ullah
FPL2
2014 Fault injection in GPGPU cores to validate and debug robust parallel applications
abstract
General Purpose Graphic Processing Units (GPGPUs) are more efficient than CPUs for processing parallel data. Unfortunately, GPGPUs are sensible to radiation. Hence, several software mitigation techniques, as well as robust algorithms, are being developed to overcome reliability problems. In this paper we propose a software debugger-based fault injection mechanism to evaluate the resiliency of applications running on a GPGPU and to validate the software hardening techniques it possibly embeds. We report some experimental results gathered on selected case studies to show the proposed approach advantages and limitations.
M. De Carvalho, Davide Sabena, Matteo Sonza Reorda, Luca Sterpone, Paolo Rech, Luigi Carro
IOLTS4
2014 Validation of a tool for estimating the effects of soft-errors on modern SRAM-based FPGAs
abstract
Predicting soft errors on SRAM-based FPGAs without a wasteful time-consuming or a high-cost has always been a very difficult goal. Among the available methods, we proposed an updated version of analytical approach to predict Single Event Effects (SEEs) based on the analysis of the circuit the FPGA implements. In this paper, we provide an experimental validation of this approach, by comparing the results it provides with a fault injection campaign. We adopted our analytical method for computing the error-rate of a design implemented on SRAM-based FPGA. Furthermore, we compared the obtained soft-error figure with the one measured by fault injection. Experimental analysis demonstrated the analytical method closely match the effective soft-error rates becoming a viable solution for the soft-error estimation at early design phases.
Marco Desogus, Luca Sterpone, David Merodio Codinachs
IOLTS2
2014 A new solution to on-line detection of Control Flow Errors
abstract
Transient faults can affect the behavior of electronic systems, and represent a major issue in many safety-critical applications. This paper focuses on Control Flow Errors (CFEs) and extends a previously proposed method, based on the usage of the debug interface existing in several processors/controllers. The new method achieves a good detection capability with very limited impact on the system development flow and reduced hardware cost: moreover, the proposed technique does not involve any change either in the processor hardware or in the application software, and works even if the processor uses caches. Experimental results are reported, showing both the advantages and the costs of the method.
Boyang Du, Matteo Sonza Reorda, Luca Sterpone, Luis Parra, Marta Portela-García, Almudena Lindoso, Luis Entrena
IOLTS3
2014 Soft error effects analysis and mitigation in VLIW safety-critical applications
abstract
VLIW architectures are widely employed in several embedded signal applications since they offer the opportunity to obtain high computational performances while maintaining reduced clock rate and power consumption. Recently, VLIW processors are being considered for employment in various embedded processing systems, including safety-critical ones (e.g., in the aerospace, automotive and rail transport domains). Terrestrial safety-critical applications based on newer nano-scale technologies raise increasing concerns about transient errors induced by neutrons. Therefore, techniques to effectively estimate and improve the reliability of VLIW processors are of great interest. In this paper, we present a novel technique aimed to further improve the efficiency of the Triple Modular Redundancy (TMR) hardening-technique applied at the software level on VLIW processors. In particular, we first experimentally demonstrate that the TMR-based software technique, when applied at the C code level, is not able to cope with most of the failures affecting user logic resources. Then, we propose a method able to analyze and modify the TMR-based code for a generic VLIW processor in order to improve the fault tolerance of the executed application without modifying the VLIW processor. In details, the proposed technique is able to reduce the number of cross-domain errors affecting the TMR-hardened code of a VLIW processor data path. We provide figures about performance and fault coverage for both the unprotected and protected versions of a set of benchmark applications, thus demonstrating the benefits and limitations of our approach.
Davide Sabena, Matteo Sonza Reorda, Luca Sterpone
VLSI-SoC3
2014 Recovery Time and Fault Tolerance Improvement for Circuits mapped on SRAM-based FPGAs
Anees Ullah, Luca Sterpone
J. Electron. Test.2
2014 ASSESS: A Simulator of Soft Errors in the Configuration Memory of SRAM-Based FPGAs
abstract
In this paper a simulator of soft errors (SEUs) in the configuration memory of SRAM-based FPGAs is presented. The simulator, named ASSESS, adopts fault models for SEUs affecting the configuration bits controlling both logic and routing resources that have been demonstrated to be much more accurate than classical fault models adopted by academic and industrial fault simulators currently available. The simulator permits the propagation of faulty values to be traced in the circuit, thus allowing the analysis of the faulty circuit not only by observing its output, but also by studying fault activation and error propagation. ASSESS has been applied to several designs, including the miniMIPS microprocessor, chosen as a realistic test case to evaluate the capabilities of the simulator. The ASSESS simulations have been validated comparing their results with a fault injection campaign on circuits from the ITC'99 benchmark, resulting in an average error of only 0.1%.
Cinzia Bernardeschi, Luca Cassano, Andrea Domenici, Luca Sterpone
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2014 On the Automatic Generation of Optimized Software-Based Self-Test Programs for VLIW Processors
abstract
Very long instruction word (VLIW) processors are increasingly employed in a large range of embedded signal processing applications, mainly due to their ability to provide high performances with reduced clock rate and power consumption. At the same time, there is an increasing request for efficient and optimal test techniques able to detect permanent faults in VLIW processors. Software-based self-test (SBST) methods are a consolidated and effective solution to detect faults in a processor both at the end of the production phase or during the operational life; however, when traditional SBST techniques are applied to VLIW processors, they may prove to be ineffective (especially in terms of size and duration), due to their inability to exploit the parallelism intrinsic in these architectures. In this paper, we present a new method for the automatic generation of efficient test programs specifically oriented to VLIW processors. The method starts from existing test programs based on generic SBST algorithms and automatically generates effective test programs able to reach the same fault coverage, while minimizing the test duration and the test code size. The method consists of four parametric phases and can deal with different VLIW processor models. The main goal of the paper is to show that in the case of VLIW processors, it is possible to automatically generate an effective test program able to achieve high fault coverage with minimal test time and required resources. Experimental data gathered on a case study demonstrate the effectiveness of the proposed approach; results show that this method is able to exploit the intrinsic parallelism of the VLIW processor, taming the growth in size, and duration of the test program when the processor size grows.
Davide Sabena, Matteo Sonza Reorda, Luca Sterpone
IEEE Trans. Very Large Scale Integr. Syst.3
2013 On-line testing of permanent radiation effects in reconfigurable systems
abstract
Partially reconfigurable systems are more and more employed in many application fields, including aerospace. SRAM-based FPGAs represent an extremely interesting hardware platform for this kind of systems, because they offer flexibility as well as processing power. In this paper we report about the ongoing development of a software flow for the generation of hard macros for on-line testing and diagnosing of permanent faults due to radiation in SRAM-FPGAs used in space missions. Once faults have been detected and diagnosed the flow allows to generate fine-grained patch hard macros that can be used to mask out the discovered faulty resources, allowing partially faulty regions of the FPGA to be available for further use.
Luca Cassano, Dario Cozzi, Sebastian Korf, Jens Hagemeyer, Mario Porrmann, Luca Sterpone
DATE6
2013 An error-detection and self-repairing method for dynamically and partially reconfigurable systems
abstract
Reconfigurable systems are gaining an increasing interest in the domain of safety-critical applications, for example in space and avionic applications. In fact, the capability of reconfiguring the system during run-time execution and the high computational power of modern Field Programmable Gate Arrays (FPGAs) makes these devices suitable for data processing. Moreover, such systems must also guarantee the abilities of self-awareness, self-diagnosis and self-repair in order to cope with errors due to the harsh conditions typically existing in some environments. In this paper we propose a self-repairing method for partially and dynamically reconfigurable systems applied at a fine-grain granularity level. Our method is able to recover and correct errors using the run-time partial reconfiguration capabilities offered by modern SRAM-based FPGAs. Fault injection experiments have been executed on a dynamically reconfigurable system embedding a number of benchmark circuits. Results demonstrate that the method can achieve full detection of single and multiple errors, while significantly improving the system availability with respect to traditional error detection and correction methods.
Matteo Sonza Reorda, Luca Sterpone, Anees Ullah
ETS2
2013 Unexcitability analysis of SEus affecting the routing structure of SRAM-based FPGAs
abstract
Testing SEUs in the configuration memory of SRAM-based FPGAs is very costly due to their large configuration memory, therefore it is necessary to optimize the generation of test patterns. In particular, in order to reduce the effort required of automatic test pattern generators, it is useful to identify early the unexcitable faults, i.e., those faults that cannot be excited by any combination of input signals. In this paper, the unexcitability of SEUs affecting the configuration bits controlling the routing resources of SRAM-based FPGAs is considered. Since this part of the configuration memory contains the largest number of configuration bits, its testing is particularly onerous. Faults in the routing resources are modeled considering the actual electrical behavior of the affected interconnections, thus the resulting fault model is more accurate than the classical open/short model usually considered. This paper introduces a methodology to prove the unexcitability of these faults. The methodology has been implemented in a tool based on a formal specification language (SAL) and a model checker (SAL-SMC). Results from the application of the tool to some circuits from the ITC'99 benchmark are reported.
Cinzia Bernardeschi, Luca Cassano, Andrea Domenici, Luca Sterpone
ACM Great Lakes Symposium on VLSI4
2013 Exploiting the debug interface to support on-line test of control flow errors
abstract
Detecting the effects of transient faults is a key point in many safety-critical applications. This paper explores the possibility of using for this purpose the debug interface existing today in several processors/controllers on the market. In this way one can achieve a good detection capability with respect to control flow errors with very small latency, while the cost for adopting the proposed technique is rather limited and does not involve any change either in the processor hardware or in the application software. The method works even if the processor uses caches. Experimental results are reported, showing both the advantages and the costs of the method.
Boyang Du, Matteo Sonza Reorda, Luca Sterpone, Luis Parra, Marta Portela-García, Almudena Lindoso, Luis Entrena
IOLTS3
2013 On the development of diagnostic test programs for VLIW processors
abstract
Software-Based Self-Test (SBST) approaches have shown to be an effective solution to detect permanent faults, both at the end of the production process, and during the operational phase. When partial reconfiguration is adopted to deal with permanent faults, we also need to identify the faulty module, which is then substituted with a spare one. Software-based Diagnosis techniques can be exploited for this purpose, too. When Very Long Instruction Word (VLIW) processors are addressed, these techniques can effectively exploit the parallelism intrinsic in these architectures. In this paper we propose a new approach that starting from existing detection-oriented programs generates a diagnosis-oriented test program which in most cases is able to identify the faulty module. Experimental results gathered on a case study show the effectiveness of the proposed approach.
Davide Sabena, Matteo Sonza Reorda, Luca Sterpone
VLSI-SoC3
2013 A Novel Fault Tolerant and Runtime Reconfigurable Platform for Satellite Payload Processing
abstract
Reconfigurable hardware is gaining a steadily growing interest in the domain of space applications. The ability to reconfigure the information processing infrastructure at runtime together with the high computational power of today's FPGA architectures at relatively low power makes these devices interesting candidates for data processing in space applications. Partial dynamic reconfiguration of FPGAs enables maximum flexibility and can be utilized for performance optimization, for improving energy efficiency, and for enhanced fault tolerance. To be able to prove the effectiveness of these novel approaches for satellite payload processing, a highly scalable prototyping environment has been developed, combining dynamically reconfigurable FPGAs with the required interfaces such as SpaceWire, MIL-STD-1553B, and SpaceFibre. The developed systems have been enabled to space harsh environments thanks to an analytical analysis of the radiation effects on its most critical reconfigurable components. Aiming at that scope, a new algorithm for the analysis of critical radiation effects, in particular, related to Single Event Upsets (SEUs) and Multiple Event Upsets (MEUs) has been developed to obtain an effective estimation of the radiation impact and enabling the tuning of the component mapping reducing the routing interaction between the reconfigurable placed modules in their different feasible positions. The experimental performance of the system has been evaluated by a proper dynamic reconfiguration scenario, demonstrating a partial reconfiguration at 400 MByte/s, blind and readback scrubbing is supported and the scrub rate can be adapted individually for different parts of the design. The fault tolerance capability has been proven by means of a new analysis algorithm and by fault injection campaigns of SEUs and MCUs into the FPGA configuration memory.
Luca Sterpone, Mario Porrmann, Jens Hagemeyer
IEEE Trans. Computers1
2012 A new SBST algorithm for testing the register file of VLIW processors
abstract
Feature size reduction drastically influences permanent faults occurrence in nanometer technology devices. Among the various test techniques, Software-Based Self-Test (SBST) approaches have been demonstrated to be an effective solution for detecting logic defects, although achieving complete fault coverage is a challenging issue due to the functional-based nature of this methodology. When VLIW processors are considered, standard processor-oriented SBST approaches result deficient since not able to cope with most of the failures affecting VLIW multiple parallel domains. In this paper we present a novel SBST algorithm specifically oriented to test the register files of VLIW processors. In particular, our algorithm addresses the cross-bar switch architecture of the VLIW register file by completely covering the intrinsic faults generated between the multiple computational domains. Fault simulation campaigns comparing previously developed methods with our solution demonstrate its effectiveness. The results show that the developed algorithm achieves a 97.12% fault coverage which is about twice better than previously developed SBST algorithms. Further advantages of our solution are the limited overhead in terms of execution cycles and memory occupation.
Davide Sabena, Matteo Sonza Reorda, Luca Sterpone
DATE3
2012 A New Fault Injection Approach for Testing Network-on-Chips
abstract
Packet-based on-chip interconnection networks, or Network-on-Chips (NoCs) are progressively replacing global on-chip interconnections in Multi-processor System-on-Chips (MP-SoCs) thanks to better performances and lower power consumption. However, modern generations of MP-SoCs have an increasing sensitivity to faults due to the progressive shrinking technology. Consequently, in order to evaluate the fault sensitivity in NoC architectures, there is the need of accurate test solution which allows to evaluate the fault tolerance capability of NoCs. This paper presents an innovative test architecture based on a dual-processor system which is able to extensively test mesh based NoCs. The proposed solution improves previously developed methods since it is based on a NoC physical implementation which allows to investigate the effects induced by several kind of faults thanks to the execution of on-line fault injection within all the network interface and router resources during NoC run-time operations. The solution has been physically implemented on an FPGA platform using a NoC emulation model adopting standard communication protocols. The obtained results demonstrated the effectiveness of the developed solution in term of testability and diagnostic capabilities and make our solutions suitable for testing large scale NoC design.
Luca Sterpone, Davide Sabena, Matteo Sonza Reorda
PDP1
2012 On the optimized generation of Software-Based Self-Test programs for VLIW processors
abstract
Software-Based Self-Test (SBST) approaches have shown to be an effective solution to detect permanent faults, both at the end of the production process, and during the operational phase. However, when Very Long Instruction Word (VLIW) processors are addressed these techniques require some optimization steps in order to properly exploit the parallelism intrinsic in these architectures. In this paper we present a new method that, starting from previously known algorithms, automatically generates an effective test program able to still reach high fault coverage on the VLIW processor under test, while reducing the test duration and the test code size. The method consists of three parametric phases and can deal with different VLIW processor models. The main goal of the proposed method is to automatically obtain a test program able to effectively reduce the test time and the required resources. Experimental results gathered on a case study show the effectiveness of the proposed approach.
Davide Sabena, Matteo Sonza Reorda, Luca Sterpone
VLSI-SoC3
2011 A new reconfigurable clock-gating technique for low power SRAM-based FPGAs
abstract
Power consumption is dramatically increasing for Static Random Access Memory Field Programmable Gate Arrays (SRAM-FPGAs), therefore lower power FPGA circuitry and new CAD tools are needed. Clock-gating methodologies have been applied in low power FPGA designs with only minor success in reducing the total average power consumption. In this paper, we developed a new structural clock-gating technique based on internal partial reconfiguration and topological modifications. The solution is based on the dynamic partial reconfiguration of the configuration memory frames related to the clock routing resources. For a set of design cases, figures of static and dynamic power consumption were obtained. The analyses have been performed on a synchronous FIFO and on a r-VEX VLIW processor. The experimental results shown that the efficiency in the total average power consumptions ranges from about 28% to 39% with respect to standard clock-gating approaches. Besides, the proposed method is not intrusive, and presents a very limited cost in term of area overhead.
Luca Sterpone, Luigi Carro, Debora Matos, Stephan Wong, F. Fakhar
DATE1
2011 Fault injection analysis of transient faults in clustered VLIW processors
abstract
VLIW architectures are widely employed in several embedded signal applications mainly because they offer the opportunity to gain high computational performances while maintaining reduced clock rate and power consumption. Recently, VLIW processors became more and more suitable to be employed in various embedded processing systems including safety critical applications such as aerospace, automotive and rail transport. Therefore, techniques to effectively estimate and improve the reliability of VLIW processor are of great interest. Terrestrial safety-critical applications based on newer nano-scale technologies raise increasing concerns about transient errors induced by neutrons. In this paper, we analyze the cross-domain failures affecting redundant mitigation techniques implemented on a statistically scheduled data path VLIW processor and we describe a fault injection analysis of transient faults affecting the r-VEX VLIW processor implemented on an FPGA platform. For a large set of benchmark applications, figures of application performances and errors analysis are provided and commented.
Luca Sterpone, Davide Sabena, Salvatore Campagna, Matteo Sonza Reorda
DDECS1
2011 A Low-Cost Emulation System for Fast Co-verification and Debug
abstract
A flexible system for SoC co-verification is proposed, built around an Infrastructure Microprocessor (IM), providing improved controllability and observability in a fast self-contained FPGA-based emulation environment. In addition, software debug is supported by enabling observation of critical signals, breakpoint setting and step-by-step execution with total memory accessibility. Experimental results in an industrial case study confirm the effectiveness of the approach for validating and debugging hardware and software.
Jorge Luis Lagos-Benites, Michelangelo Grosso, Luca Sterpone, Matteo Sonza Reorda, G. Audisio, Mauro Pipponzi, Marco Sabatini
ETS3
2010 A new placement algorithm for the mitigation of multiple cell upsets in SRAM-based FPGAs
abstract
Modern FPGAs have been designed with advanced integrated circuit techniques that allow high speed and low power performance, joined to reconfiguration capabilities. This makes new FPGA devices very advantageous for space and avionics computing. However, larger levels of integration makes FPGA's configuration memory more prone to suffer Multi-Cell Upset errors (MCUs), caused by a single radiation particle that can flip the content of multiple nearby cells. In particular, MCUs are on the rise for the new generation of SRAM-based FPGAs, since their configuration memory is based on volatile programming cells designed with smaller geometries that result more sensitive to proton- and heavy ion-induced effects. MCUs drastically limits the capabilities of specific hardening techniques adopted in space-based electronic systems, mainly based on Triple Modular Redundancy (TMR). In this paper we describe a new placement algorithm for hardening TMR circuits mapped on SRAM-based FPGAs against the effects of MCUs. The algorithm is based on layout information of the FPGA's configuration memory and on metrics related to the logic and interconnection resources locations. Experimental results obtained from MCU static analysis on a set of benchmark circuits hardened by the proposed algorithm prove the efficiency of our approach.
Luca Sterpone, Niccolò Battezzati
DATE1
2010 On the mitigation of SET broadening effects in integrated circuits
abstract
Nowadays, the integrated circuits design and manufacturing process are decreasing the minimum transistor size and this advancement, accompanied by increasing operating frequencies and lower power supplies voltages, leads, on the one side, to the availability of fast and low power circuits with very small noise margins but, on the other side, makes integrated circuits more sensitive to Single Event Transient (SET) pulses that may be generated and propagated through the combinational logic, leading to misbehaviors. SETs are mainly generated by high-energy particles that strikes the circuit near a junction, resulting in a significant charge injection/depletion process, that may produce spurious pulses. These can propagate and change their shape traversing the combinational logic paths, sometimes being broadened and amplified sometimes being filtered. In this paper, we present a place and route algorithm for integrated circuit design, which is able to mitigate and filter the erroneous effects of SETs. The proposed solution has been experimentally evaluated by means of electrical pulse injection within logic resources of several benchmark Integrated Circuits (ICs) implemented in a Flash-based FPGA and by accurate timing analyses. Preliminary results confirm the mitigation of SET broadening effects by acting on physical place and route constraints. On the selected benchmark circuit the algorithm decreases the SET sensitiveness more than 70% with respect to not hardened circuits. Besides, the solution does not introduce any area overhead or delay penalties.
Luca Sterpone, Niccolò Battezzati
DDECS1
2010 An integrated flow for the design of hardened circuits on SRAM-based FPGAs
abstract
This paper presents an enhanced design flow for the implementation of hardened systems on SRAM-based FPGAs, able to cope with the occurrence of Single Event Upsets (SEUs). The framework integrates three strategies independently designed to tackle the problem of SEUs; first a systematic methodology is used to harden the circuit exploiting an enhanced TMR-based technique, coupled with partial dynamic reconfiguration. Then, a robustness analysis is performed to identify possible TMR failures, eventually solved by a specific local re-design of the critical portions of the implementation. We present the overall flow and the benefits of the solution, experimentally evaluated on a realistic circuit.
Cristiana Bolchini, Antonio Miele, Chiara Sandionigi, Niccolò Battezzati, Luca Sterpone, Massimo Violante
ETS5
2010 A novel scalable and reconfigurable emulation platform for embedded systems verification
abstract
Modern embedded systems are characterized by a heterogeneous architecture including several modules (e.g., DSPs, memories and mixed-signal IPs) often integrated with one or more microprocessor cores controlling the system functionalities by means of embedded software programs. The verification of such a kind of systems has become a challenge due to their increasing complexity that makes traditional simulation and emulation techniques unaffordable methods for current quality and time-to-market constraints. This paper presents a new platform for the hardware and software verification of modern embedded systems based on a reconfigurable device. The main novelty consists in an infrastructure architecture containing a signal processing IP and a microprocessor core flexibly interfaced with the device under validation, aimed at the overall reduction of the design verification time. It also provides a dynamic interface supporting the software verification of the embedded system microprocessors. The proposed environment is fully scalable and adaptable to the requirements of a general purpose embedded system, enabling advanced verification flows at different phases of design and integration without time expensive interface modification. Experimental and performance analysis on a real industrial case study are reported proving the effectiveness of the proposed solution.
M. Di Marzio, Michelangelo Grosso, Matteo Sonza Reorda, Luca Sterpone, G. Audisio, Marco Sabatini
ISCAS4
2010 A new software tool for static analysis of SET sensitiveness in Flash-based FPGAs
abstract
The higher resiliency of Flash-based FPGAs to Single Event Upsets (SEUs) with respect to other non radiation-hardened devices, such as SRAM-based FPGAs, are increasing more and more their demand for avionic and space applications, where a harsh environment rich in ionizing radiation has to be faced. In this type of devices other transient faults tend to dominate over SEUs, especially when the device operates at high frequency. In this scenario, it is expected that Single Event Transient (SET) faults will predominate. As a result, designers will still need prediction techniques to forecast the effects of ionizing radiation in their designs. Although radiation testing is a feasible method for evaluating circuit sensitiveness against SETs, it is hard to implement, very expensive, and it can be used only in later phases of the design process, when a prototype of the system is available. On the other hand, simulation techniques need a first technology characterization step and also require a very detailed model for being effective; moreover they are application dependent. In this paper we propose a new software tool for analyzing designs implemented in Flash-based FPGAs and estimating SET sensitiveness. The evaluation process is static, as it does not entail any simulation. In particular, it provides worst-case results, thus being intrinsically more conservative than other dynamic methods. Experimental results are presented comparing the ones coming from radiation testing and the results provided by the presented tool. They validate the proposed approach.
Niccolò Battezzati, Luca Sterpone, Massimo Violante, Filomena Decuzzi
VLSI-SoC2
2010 A New Timing Driven Placement Algorithm for Dependable Circuits on SRAM-based FPGAs
abstract
Electronic systems for safety critical applications such as space and avionics need the maximum level of dependability for guarantee the success of their missions. Simultaneously the computation capabilities required in these fields are constantly increasing for afford the implementation of different kind of applications ranging from signal processing to networking. SRAM-based FPGAs are the candidate devices to achieve this goal thanks to their high versatility of implementing complex circuits with a very short development time. However, in critical environments, the presence of Single Event Upsets (SEUs) affecting the FPGA’s functionalities, requires the adoption of specific fault tolerant techniques, like Triple Modular Redundancy (TMR), able to increase the protection capability against radiation effects, but on the other side, introducing a dramatic penalty in terms of performances. In this paper, it is proposed a new timing-driven placement algorithm for implementing soft-errors resilient circuits on SRAM-based FPGAs with a negligible degradation of performance. The algorithm is based on a placement heuristic able to remove the crossing error domains while decreasing the routing congestions and delay inserted by the TMR routing and voting scheme. Experimental analysis performed by timing analysis and SEU static analysis point out a performance improvement of 29% on the average with respect to standard TMR approach and an increased robustness against SEU affecting the FPGA’s configuration memory. Accurate analyses of SEUs sensitivity and performance optimization have been performed on a real microprocessor core demonstrating an improvement of performances of more than 62%.
Luca Sterpone
ACM Trans. Reconfigurable Technol. Syst.1
2009 A study of the Single Event Effects impact on functional mapping within Flash-based FPGAs
abstract
Flash-based FPGAs are increasingly demanded in safety critical fields, in particular space and avionic ones, due to their non-volatile configuration memory. Although they are almost immune to permanent loss of the configuration data, they are composed of floating gate based switches that can suffer transient effects if hit by high energetic particles with critical consequences on the implemented logic. This paper presents a new way for the analysis of the impact of single event effects in flash-based FPGAs. We proposed a new methodology to identify the most critical switches inside the configuration logic block and the most redundant and robust configuration selection for each logic function. The experimental results achieved by fault injection demonstrated the feasibility of the proposed method and show that by using the most robust functional mapping it is possible to enhance the reliability of the entire design with respect to a not robust ones.
Francesco Abate, Luca Sterpone, Massimo Violante, Fernanda Lima Kastensmidt
DATE2
2009 Soft errors in Flash-based FPGAs: Analysis methodologies and first results
abstract
The paper presents the development of three different analysis methodologies in order to evaluate soft errors effects in flash-based FPGAs. They are complementary and can be used in different design stages, from the device characterization up to the design sensitiveness estimation. First results are very promising, proving that such methodologies are valid and open new ways of investigation. In particular, we are going to upgrade the experimental setup in order to support higher frequencies (up to 250 MHz) for further characterizing SEE effects. Moreover, a benchmark circuit should be defined in order to correctly predict the expected number of SETs for real circuits, taking into account other side effects, like broadening and logical masking. We expect that from the analysis results we will able to delight suitable hardening techniques that will undergo to both radiation test and prediction analysis.
Niccolò Battezzati, Filomena Decuzzi, Luca Sterpone, Massimo Violante
FPL3
2008 Differential gene expression graphs: A data structure for classification in DNA microarrays
abstract
This paper proposes an innovative data structure to be used as a backbone in designing microarray phenotype sample classifiers. The data structure is based on graphs and it is built from a differential analysis of the expression levels of healthy and diseased tissue samples in a microarray dataset. The proposed data structure is built in such a way that, by construction, it shows a number of properties that are perfectly suited to address several problems like feature extraction, clustering, and classification.
Alfredo Benso, Stefano Di Carlo, Gianfranco Politano, Luca Sterpone
BIBE4
2008 A graph-based representation of Gene Expression profiles in DNA microarrays
abstract
This paper proposes a new and very flexible data model, called Gene Expression Graph (GEG), for genes expression analysis and classification. Three features differentiate GEGs from other available microarray data representation structures: (i) the memory occupation of a GEG is independent of the number of samples used to built it; (ii) a GEG more clearly expresses relationships among expressed and non expressed genes in both healthy and diseased tissues experiments; (iii) GEGs allow to easily implement very efficient classifiers. The paper also presents a simple classifier for sample-based classification to show the flexibility and user-friendliness of the proposed data structure.
Alfredo Benso, Stefano Di Carlo, Gianfranco Politano, Luca Sterpone
CIBCB4
2008 On the design of tunable fault tolerant circuits on SRAM-based FPGAs for safety critical applications
abstract
Mission-critical applications such as space or avionics increasingly demand high fault tolerance capabilities of their electronic systems. Among the fault tolerance characteristics, the performance and costs of an electronic system remain the leader factors in the space and avionics market. In particular, when considering SRAM-based FPGAs, specific hardening techniques generally based on Triple Modular Redundancy need to be adopted in order to guarantee the desired fault tolerance degree. While effectively increasing the fault tolerance capability, these techniques introduce an important performance degradation and a dramatic area overhead, that results in higher design costs. In this paper, we propose an innovative design flow that allow the implementation of fault tolerance circuits in SRAM-based FPGA devices with different fault tolerance capability degrees. We introduce a new metric that allows a designer to precisely estimate and set the desired fault tolerance capabilities. Experimental analysis performed on a realistic industrial-type case study demonstrates the efficiency of our methodology.
Luca Sterpone, Miguel A. Aguirre, Jonathan Noel Tombs, Hipólito Guzmán-Miranda
DATE1
2008 On the Evaluation of Radiation-Induced Transient Faults in Flash-Based FPGAs
abstract
Field programmable gate arrays (FPGAs) are getting more and more attractive for military and aerospace applications, among others devices. The usage of non volatile FPGAs, like Flash-based ones, reduces permanent radiation effects but transient faults are still a concern. In this paper we propose a new methodology for effectively measuring the width of radiation-induced transient faults thus allowing tuning known mitigation techniques accordingly. Radiation experiments results are presented and commented demonstrating that the proposed methodology is a viable solution to measure the transient pulses width.
Niccolò Battezzati, Simone Gerardin, Andrea Manuzzato, Alessandro Paccagnella, Sana Rezgui, Luca Sterpone, Massimo Violante
IOLTS6
2008 Software and Hardware Techniques for SEU Detection in IP Processors
Cristiana Bolchini, Antonio Miele, Fabio Rebaudengo, Fabio Salice, Donatella Sciuto, Luca Sterpone, Massimo Violante
J. Electron. Test.6
2007 Static and Dynamic Analysis of SEU Effects in SRAM-Based FPGAs
abstract
SRAM-based Field Programmable Gate Arrays (FPGAs) are very sensitive to Single Event Upsets (SEUs) affecting their configuration memory. SEUs may have critical effects on the circuit FPGA devices implement. In order to deploy safety- or mission-critical applications on SRAM-based FPGAs, designers need to adopt hardening techniques, as well as methodologies for estimating and validating the SEU's sensitivity of the obtained applications in the early design phase. In this paper we describe a new methodology for predict the effects of SEUs by combining static and dynamic analysis of the circuit's FPGA implements. The proposed methodology is able to identify the critical single event upset locations within the configuration memory and to provide a detailed classification of the provoked effects. Experimental results on several realistic applications demonstrate the feasibility of the proposed methodology.
Luca Sterpone, Massimo Violante
ETS1
2007 A new decompression system for the configuration process of SRAM-based FPGAS
abstract
Nowadays Field Programmable Gate Arrays (FPGAs) are an improved technology for developing high-performance embedded systems. SRAM-based FPGAs offers the possibility of in-the-field reconfiguration that results in the ability to adapt the product to modified user's requirements, to enrich the product's features, or simply to correct bugs. With the advent of multi-million gate FPGAs, the size of the configuration information that defines what circuit the FPGA implements has increased drastically, and thus the amount of external memory needed to keep the configuration data is increasing dramatically. In this work we developed a novel configuration compression system that exploits internal configuration mechanism of modern SRAM-based FPGAs and results in high compression efficiency. The proposed system is applicable to any modern SRAM-based FPGA devices having an embedded microprocessor core since the configuration data are processed as raw data. Moreover, the proposed approach does not require any external hardware support and allows high speed dynamic reconfiguration. Experimental results on Xilinx SRAM-based FPGAs platform implementing several real-world circuits demonstrated 82% savings in memory on the average.
Luca Sterpone, Massimo Violante
ACM Great Lakes Symposium on VLSI1
2007 A new hardware architecture for performing the gridding of DNA microarray images
abstract
DNA microarray technologies are an essential part of modern biomedical research. The analysis of DNA microarray images allows the identification of gene expressions in order to drawn biologically meaningful conclusions for applications that ranges from the genetic profiling to the diagnosis of oncology diseases. Unfortunately, DNA microarray technology has a high variation of data quality. Therefore, in order to obtain reliable results, complex and extensive image analysis algorithms should be applied before actual DNA microarray information can be used for biomedical purpose. In this paper, we present a novel hardware acceleration architecture specifically designed to process DNA microarray images. The proposed architecture uses several units working in a single instruction-multiple data fashion managed by a microprocessor core. An FPGA-based prototypal implementation of the developed architecture is presented. Experimental results on several realistic DNA microarray images show a reduction of the computation time of one order of magnitude if compared with previously developed software-based approach.
Luca Sterpone, Massimo Violante
ACM Great Lakes Symposium on VLSI1
2007 Self Checking Circuit Optimization by means of Fault Injection Analysis: A Case Study on Reed Solomon Decoders
abstract
This paper shows how the use of exhaustive fault injection campaigns in conjunction with the analysis of the property of a circuit, allows to improve the efficiency of the checker of self checking circuits. Experimental results coming from fault injection campaigns on a Reed-Solomon decoder demonstrated that by observing the occurred errors and the correspondent detection module has been possible to reduce the number of detection module, while paying a small reduction of the percentage of SEUs that can be detected.
Salvatore Pontarelli, Luca Sterpone, Gian Carlo Cardarilli, Marco Re, Matteo Sonza Reorda, Adelio Salsano, Massimo Violante
IOLTS2
2007 Evaluating Different Solutions to Design Fault Tolerant Systems with SRAM-based FPGAs
Luca Sterpone, Matteo Sonza Reorda, Massimo Violante, Fernanda Lima Kastensmidt, Luigi Carro
J. Electron. Test.1
2006 Fault Injection-based Reliability Evaluation of SoPCs
abstract
Systems-on-programmable-chip (SoPCs) include processors, memories and programmable logic that allow to catch multiple application requirements such as high performance, reconfigurability and low-costs. Due to these characteristics, they are also becoming very attractive for safety-critical applications. However, the issue of assessing the reliability they can provide and debugging the possible safety-related mechanisms they embed is still open. In this paper, we present a new fault-injection approach for evaluating the impact of transient faults in SoPCs. Fault-injection experiments are reported on a case study consisting of a Web server implemented on a Xilinx Virtex-II FPGA embedding a PowerPC 405 and running the whole TCP/IP stack
Matteo Sonza Reorda, Luca Sterpone, Massimo Violante, Marta Portela-García, Celia López-Ongil, Luis Entrena
ETS2
2006 Dependability Evaluation of Transient Fault Effects in Reconfigurable Compute Fabric Devices
abstract
Reconfigurable compute fabrics (RCFs) are cellular architectures in which an array of computing elements and a configurable interconnection fabric are combined with a general-purpose processor. RCFs can play an important role in safety- or mission-critical applications, provided that a clear understanding of their dependability is available. In this paper, we report an evaluation of the effects induced by transient faults within the resources of an RCF Motorola MRC6011 and we resorted to extensive fault injection to investigate the effects of transient faults
Luca Sterpone, Massimo Violante
IOLTS1
2006 A New Reliability-Oriented Place and Route Algorithm for SRAM-Based FPGAs
abstract
The very high integration levels reached by VLSI technologies for SRAM-based field programmable gate arrays (FPGAs) lead to high occurrence-rate of transient faults induced by single event upsets (SEUs) in FPGAs' configuration memory. Since the configuration memory defines which circuit an SRAM-based FPGA implements, any modification induced by SEUs may dramatically change the implemented circuit. When such devices are used in safety-critical applications, fault-tolerant techniques are needed to mitigate the effects of SEUs in FPGAs' configuration memory. In this paper, we analyze the effects induced by the SEUs in the configuration memory of SRAM-based FPGAs. The reported analysis outlines that SEUs in the FPGA's configuration memory are particularly critical since they are able to escape well-known fault masking techniques such as triple modular redundancy (TMR). We then present a reliability-oriented place and route algorithm that, coupled with TMR, is able to effectively mitigate the effects of the considered faults. The effectiveness of the new reliability-oriented place and route algorithm is demonstrated by extensive fault injection experiments showing that the capability of tolerating SEU effects in the FPGA's configuration memory increases up to 85 times with respect to a standard TMR design technique
Luca Sterpone, Massimo Violante
IEEE Trans. Computers1
2005 On the Optimal Design of Triple Modular Redundancy Logic for SRAM-based FPGAs
abstract
Triple modular redundancy (TMR) is a suitable fault tolerant technique for SRAM-based FPGA. However, one of the main challenges in achieving 100% robustness in designs protected by TMR running on programmable platforms is to prevent upsets in the routing from provoking undesirable connections between signals from distinct redundant logic parts, which can generate an error in the output. This paper investigates the optimal design of the TMR logic (e.g., by cleverly inserting voters) to ensure robustness. Four different versions of a TMR digital filter were analyzed by fault injection. Faults were randomly inserted straight into the bitstream of the FPGA. The experimental results presented in this paper demonstrate that the number and placement of voters in the TMR design can directly affect the fault tolerance, ranging from 4.03% to 0.98% the number of upsets in the routing able to cause an error in the TMR circuit.
Fernanda Lima Kastensmidt, Luca Sterpone, Luigi Carro, Matteo Sonza Reorda
DATE2
2005 Multiple errors produced by single upsets in FPGA configuration memory: a possible solution
abstract
The very high integration levels reached by SRAM-based field programmable gate arrays (FPGAs) lead to high occurrence rate of single event upsets (SEUs) in their configuration memory, which can produce multiple errors affecting routing resources. Based on detailed analysis of this phenomenon, we devised a reliability-oriented place and route algorithm able to significantly improve the reliability of SRAM-based FPGAs with limited costs in terms of performance degradation and resource occupation. To evaluate the effectiveness of the algorithm we performed extensive fault injection experiments.
Matteo Sonza Reorda, Luca Sterpone, Massimo Violante
ETS2
2005 New evolutionary techniques for test-program generation for complex microprocessor cores
abstract
Checking if microprocessor cores are fully functional at the end of the productive process has become a major issue. Traditional functional approaches are not sufficient when considering modern designs. This paper describes new improvements for an existing evolutionary algorithm, called µGP, able to generate Turing-complete programs; these are exploited, along with hardware acceleration techniques, to add content to a qualifying test campaign by automatically generating assembly programs. The approach is suitable for medium-sized processor cores. The experimental evaluation performed on a SPARCv8 clearly shows the potentiality of the approach, and the effectiveness of the enhancements to the evolutionary core.
Ernesto Sánchez 0001, Massimiliano Schillaci, Matteo Sonza Reorda, Giovanni Squillero, Luca Sterpone, Massimo Violante
GECCO5
2005 Efficient Estimation of SEU Effects in SRAM-Based FPGAs
abstract
SRAM-based FPGAs are becoming very appealing for several applications where high dependability is a mandatory requirement. Unfortunately, the technology of SRAM-based FPGAs is very sensitive to single event upsets (SEUs) and particular concerns arise from SEUs affecting the FPGAs' configuration memory. In this paper we propose a new method for assessing the impact of faults in the configuration memory on the FPGA dependability. The method uses static analysis, thus reducing greatly the time for performing dependability evaluation.
Matteo Sonza Reorda, Luca Sterpone, Massimo Violante
IOLTS2
2004 On the Evaluation of SEU Sensitiveness in SRAM-Based FPGAs
Paolo Bernardi 0002, Matteo Sonza Reorda, Luca Sterpone, Massimo Violante
IOLTS3