Paolo Rech

dblp:64/5340 · DBLP profile ↗
← Back
70ranked-venue papers
8as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 63 · 8 first-author · 29 since 2021Software engineering, systems software and programming languages · 17 · 4 first-author · 7 since 2021Security and privacy · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Thinking Inside the Box: Injecting Realistic Radiation Faults in ML Accelerators
Bruno Loureiro Coelho, Mani Sadati, Abraham Chan, Alex Hands, Karthik Pattabiraman, Paolo Rech
DSN6
2026 Project Highlights - Reliability Evaluation for ARCHYTAS AI hardware accelerators
Angeliki Kritikakou, Fernando Santos 0001, Marcello Traiola, Rafael Billig Tonetto, Olivier Sentieys, Paolo Rech, Haralampos-G. D. Stratigopoulos, Georgios Keramidas
IOLTS6
2026 State of practice: Evaluating GPU performance of state vector and tensor network methods
abstract
The frontier of quantum computing (QC) simulation on classical hardware is quickly reaching the hard scalability limits for computational feasibility. Nonetheless, there is still a need to simulate large quantum systems classically, as the Noisy Intermediate Scale Quantum (NISQ) devices are yet to be considered fault tolerant and performant enough in terms of operations per second. Each of the two main exact simulation techniques, state vector and tensor network simulators, boasts specific limitations. This article investigates the limits of current state-of-the-art simulation techniques on a test bench made of eight widely used quantum subroutines, each in different configurations, with a special emphasis on performance. We perform both single process and distributed scaleability experiments on a supercomputer. We correlate the performance measures from such experiments with the metrics that characterise the benchmark circuits, identifying the main reasons behind the observed performance trends. Specifically, we perform distributed sliced tensor contractions, and we analyse the impact of pathfinding quality on contraction time, correlating both results with topological circuit characteristics. From our observations, given the structure of a quantum circuit and the number of qubits, we highlight how to select the best simulation strategy, demonstrating how preventive circuit analysis can guide and improve simulation performance by more than an order of magnitude.
Marzio Vallero, Paolo Rech, Flavio Vella
Future Gener. Comput. Syst.2
2026 Optimized Hyperdimensional Edge AI Evaluation for Efficiency and Reliability under Real Radiation
abstract
Hyperdimensional Computing (HDC) is an emerging AI algorithm, touted to be an efficient, neuro-inspired and reliable alternative to neural networks for Edge AI. HDC utilizes hypervectors with several thousand elements; the number of elements in these hypervectors denotes the HDC dimension. This dimension can be optimized for improving the efficiency and reliability of HDC inference against errors such as bit-flips, which can be caused by environmental radiation-induced soft errors. We hypothesize that, by reducing the runtime chip area and execution time utilized by HDC inference through lowering dimensionality, both efficiency and reliability against soft error-induced bit-flips can be simultaneously improved while trading off a negligible amount of accuracy and error threshold. We tested our hypothesis by executing an HDC inference algorithm with two different dimension values, 10000 (10k) and 1024, on a commercially available, low-power, bare-metal ARM platform with a Cortex-M4 processor. We conducted the efficiency analysis by measuring the CPU cycles and energy required for executing the algorithm, and the reliability analysis using real-world atmospheric-like neutron radiation from the ChipIr facility in Oxfordshire, UK. Analyses revealed that, by lowering the HDC dimension from 10k to 1024, the reliability of HDC inference against soft error-induced bit-flips was 3.5 times better and efficiency improved by more than 16 times. This innovative observation contrasts the prevailing understanding in the community that increasing the HDC dimension always improves robustness or reliability. To the best of our knowledge, our work is the first to study the reliability of HDC inference using real-world radiation.
Justus Rajappa Anuj, Laura Smets, Philippe Reiter, Paolo Rech, Ynte Vanderhoydonc, Ritesh Kumar Singh, Siegfried Mercelis, Jeroen Famaey
ACM Trans. Embed. Comput. Syst.4
2025 European Test Symposium Teams: an Anniversary Snapshot
abstract
The IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing.
Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand
ETS50
2024 Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers Models
abstract
Vision Transformers (ViTs) are highly accurate Machine Learning (ML) models. However, their large size and complexity increase the expected error rate due to hardware faults. Measuring the error rate of large ViT models is challenging, as conventional microarchitectural fault simulations can take years to produce statistically significant data. This paper proposes a two-level evaluation based on data collected through more than 70 hours of neutron beam experiments and more than 600 hours of software fault simulation. We consider 12 ViT models executed in 2 NVIDIA GPU architectures. We first characterize the fault model in ViT's kernels to identify the faults more likely to propagate to the output. We then design dedicated procedures efficiently integrated into the ViT to locate and correct these faults. We propose Maximum corrupted Malicious values (MaxiMals), an experimentally tuned low-cost mitigation solution to reduce the impact of transient faults on ViTs. We demonstrate that MaxiMals can correct 90.7% of critical failures, with execution time overheads as low as 5.61%.
Lucas Roquet, Fernando Santos 0001, Paolo Rech, Marcello Traiola, Olivier Sentieys, Angeliki Kritikakou
DAC3
2024 Reliability and Security of AI Hardware
abstract
In recent years, Artificial Intelligence (AI) systems have achieved revolutionary capabilities, providing intelligent solutions that surpass human skills in many cases. However, such capabilities come with power-hungry computation workloads. Therefore, the implementation of hardware acceleration becomes as fundamental as the software design to improve energy efficiency, silicon area, and latency of AI systems. Thus, innovative hardware platforms, architectures, and compiler-level approaches have been used to accelerate AI workloads. Crucially, innovative AI acceleration platforms are being adopted in application domains for which dependability must be paramount, such as autonomous driving, healthcare, banking, space exploration, and industry 4.0. Unfortunately, the complexity of both AI software and hardware makes the dependability evaluation and improvement extremely challenging. Studies have been conducted on both the security and reliability of AI systems, such as vulnerability assessments and countermeasures to random faults and analysis for side-channel attacks. This paper describes and discusses various reliability and security threats in AI systems, and presents representative case studies along with corresponding efficient countermeasures.
Dennis Gnad, Martin Gotthard, Jonas Krautter, Angeliki Kritikakou, Vincent Meyers, Paolo Rech, Josie E. Rodriguez Condia, Annachiara Ruospo, Ernesto Sánchez 0001, Fernando Santos 0001, Olivier Sentieys, Mehdi Baradaran Tahoori, Russell Tessier, Marcello Traiola
ETS6
2024 Neutron Beam Evaluation of Probabilistic Data Structure-based Online Checkers
abstract
High-criticality applications are vulnerable to Single Event Effects (SEEs) and require highly reliable and customizable microprocessors. Online checkers have been used to detect security and reliability issues in such systems. Popular hardware redundancy techniques such as Triple Modular Redundancy (TMR) and Dual Modular Redundancy (DMR) provide a high error coverage at the cost of substantial redundancy; therefore, there is an interest in introducing lightweight checkers that could offer the same detection ability as DMR with a much lower overhead. A possible implementation of these online checkers can be based on Probabilistic Data Structure (PDS) such as the Bloom Filter (BF). They are a form of information redundancy and an excellent complement to Single Error Correction Double Error Detection (SECDED) codes because they allow for detecting higher-order upsets. In this work, we integrate an online checker into the open-source RISC-V core NEORV32 and deploy it on a flash-based FPGA. This paper presents the evaluation of the online checker’s performance conducted under a neutron beam. The neutron beam experiments demonstrate that the real-life error rates of such structures are comparably worse than the initial simulation would indicate and that other factors can impact their performance.
Bruno Endres Forlin, Edian B. Annink, Elijah Cishugi, Carlo Cazzaniga, Paolo Rech, Gerard K. Rauwerda, Gianluca Furano, Marco Ottavi
IOLTS5
2024 Interleaved Execution of Approximated CUDA Kernels in Iterative Applications
abstract
Fine-tuning the floating-point precision of arithmetic operations in applications can be extremely challenging and time-consuming, especially in iterative applications where the output of one iteration serves as the input for the subsequent iteration. Consequently, the accuracy loss can be magnified throughout the execution. Therefore, we propose an alternative approach based on the interleaved execution of multiple approximated kernel versions designed with different precision levels. We demonstrate that creating an interleaved execution configuration of multiple CUDA kernel versions based on their accuracy loss profiles enables us to enhance performance, improve energy efficiency, and manage the accuracy loss of scientific simulation applications in various Target Output Quality (TOQ) scenarios. For a TOQ loss of approximately 3 %, we achieve a speedup of up to 1. 7x and reduce energy consumption by nearly 40 %.
Gabriel Freytag, Cristiano A. Künas, Paolo Rech, Philippe Olivier Alexandre Navaux
PDP3
2024 On the Efficacy of Surface Codes in Compensating for Radiation Events in Superconducting Devices
abstract
Reliability is fundamental for developing large-scale quantum computers. Since the benefit of technological advancements to the qubit’s stability is saturating, algorithmic solutions, such as quantum error correction (QEC) codes, are needed to bridge the gap to reliable computation. Unfortunately, the deployment of the first quantum computers has identified faults induced by natural radiation as an additional threat to qubits reliability. The high sensitivity of qubits to radiation hinders the large-scale adoption of quantum computers, since the persistence and area-of-effect of the fault can potentially undermine the efficacy of the most advanced QEC. In this paper, we investigate the resilience of various implementations of state-of-the-art QEC codes to radiation-induced faults. We report data from over 400 million fault injections and correlate hardware faults with the logical error observed after decoding the code output, extrapolating physical-to-logical error rates. We compare the code’s radiation-induced logical error rate over the code distance, the number and role in the QEC of physical qubits, the underlying quantum computer topology, and particle energy spread in the chip. We show that, by simply selecting and tuning properly the surface code, thus without introducing any overhead, the probability of correcting a radiation-induced fault is increased by up to 10%. Finally, we provide indications and guidelines for the design of future QEC codes to further increase their effectiveness against radiation-induced events.
Marzio Vallero, Gioele Casagranda, Flavio Vella, Paolo Rech
SC4
2024 Assessing the Impact of Compiler Optimizations on GPUs Reliability
abstract
Graphics Processing Units (GPUs) compilers have evolved in order to support general-purpose programming languages for multiple architectures. NVIDIA CUDA Compiler (NVCC) has many compilation levels before generating the machine code and applies complex optimizations to improve performance. These optimizations modify how the software is mapped in the underlying hardware; thus, as we show in this article, they can also affect GPU reliability. We evaluate the effects on the GPU error rate of the optimization flags applied at the NVCC Parallel Thread Execution (PTX) compiling phase by analyzing two NVIDIA GPU architectures (Kepler and Volta) and two compiler versions (NVCC 10.2 and 11.3). We compare and combine fault propagation analysis based on software fault injection, hardware utilization distribution obtained with application-level profiling, and machine instructions radiation-induced error rate measured with beam experiments. We consider eight different workloads and 144 combinations of compilation flags, and we show that optimizations can impact the GPUs’ error rate of up to an order of magnitude. Additionally, through accelerated neutron beam experiments on a NVIDIA Kepler GPU, we show that the error rate of the unoptimized GEMM (-O0 flag) is lower than the optimized GEMM’s (-O3 flag) error rate. When the performance is evaluated together with the error rate, we show that the most optimized versions (-O1 and -O3) always produce a higher amount of correct data than the unoptimized code (-O0).
Fernando Santos 0001, Luigi Carro, Flavio Vella, Paolo Rech
ACM Trans. Archit. Code Optim.4
2024 A Systematic Methodology to Compute the Quantum Vulnerability Factors for Quantum Circuits
abstract
Quantum computing is one of the most promising technology advances of the latest years. Qubits are highly sensitive to noise, which can make the output useless. Lately, it has been shown that superconducting qubits are extremely susceptible to external sources of faults, such as ionizing radiation. When adopted in large scale, radiation-induced errors are expected to become a serious challenge for qubits reliability. We propose an evaluation of the impact of transient faults in the execution of quantum circuits on superconducting chips. Inspired by the Architectural and Program Vulnerability Factors, widely used for classical computation, we propose the Quantum Vulnerability Factor (QVF) to measure the impact of qubit corruption on the circuit output. We model faults, and design a fault injector, based on the latest studies on real machines and radiation experiments. We report the finding of more than 388,000,000 fault injections, considering single and double faults, on three algorithms, identifying the faults and qubits that are more likely to impact the output. We give guidelines on how to map the qubits in real devices to reduce the output error and to reduce the probability of having a radiation-induced corruption modifying the output. Finally, we compare simulations with experiments on physical quantum computers.
Daniel Oliveira 0002, Edoardo Giusto, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech
IEEE Trans. Dependable Secur. Comput.6
2024 Can GPU performance increase faster than the code error rate?
abstract
Abstract Graphics processing units (GPUs) are the reference architecture to accelerate high-performance computing applications and the training/interference of convolutional neural networks. For both these domains, performance and reliability are two of the main constraints. It is believed that the only way to increase reliability is to sacrifice performance, e.g., using redundancies. We show in this paper that this is not always the case. As a very promising result, we found that most GPUs performance improvements also bring the benefit of increasing the number of executions correctly completed before experiencing a silent data corruption (SDC). We consider four different common GPUs’ performance optimizations: architectural solutions, software implementations, compiler optimizations, and threads degree of parallelism. We compare different implementations of a variety of parallel codes and, through beam experiments and applications profiling, we show that the performance improvement typically (but not necessarily) increases the GPU SDC rate. Nevertheless, for the vast majority of the configurations the performance gain is much higher than the SDC rate increase, allowing to process a higher amount of correct data. As we show, the programmer choices can increase up to $$25\times {}$$ 25 × the number of correctly completed executions without redesigning the algorithm nor including specific hardening solutions.
Fernando Santos 0001, Paolo Rech
J. Supercomput.2
2023 An unprotected RISC-V Soft-core processor on an SRAM FPGA: Is it as bad as it sounds?
abstract
Fast development, low cost, and reconfigurability are becoming critical factors for aerospace applications, making SRAM FPGAs attractive. However, SRAM FPGAs are prone to errors in the user and on the configuration bits. For their correct functioning, they must be capable of withstanding failures without sacrificing much performance. When adjusting a soft core for these applications, it is essential to know where redundancies are necessary, to avoid unnecessary overhead. We characterize the reliability of an unprotected RISC-V microcontroller using an accelerated neutron beam. Our investigation shows that, for our chosen benchmark and processor, the user data in the memory banks is the leading cause of the total number of errors in the application. By reversing the benchmark operations, we could root cause the origin of the observed errors and found that most of the data corruption detected during the runs stem from previously corrupt input data or from output data that were corrupted while transmitting.
Bruno Endres Forlin, Wouter van Huffelen, Carlo Cazzaniga, Paolo Rech, Nikolaos Alachiotis 0001, Marco Ottavi
ETS4
2023 Understanding and Improving GPUs' Reliability Combining Beam Experiments with Fault Simulation
abstract
Graphics Processing Units (GPUs) are being employed in High Performance Computing (HPC) and safety-critical applications, such as autonomous vehicles. This market shift led to significant improvements in the programming frameworks and performance evaluation tools and concerns about their reliability. GPU reliability evaluation is extremely challenging due to the parallel nature and high complexity of GPU architectures. We conducted the first cross-layer GPU reliability evaluation to unveil (and mitigate) GPU vulnerabilities. The proposed evaluation is achieved by comparing and combining extensive high-energy neutron beam experiments, massive fault simulation campaigns at both Register-Transfer Level (RTL) and software levels, and application profiling. Based on this extensive and detailed analysis, a novel accurate methodology to accurately estimate GPUs application FIT rate is proposed. Moreover, by employing the knowledge obtained from the cross-layer reliability evaluation, two novel hardening solutions for HPC and safety-critical applications are proposed: (1) Reduced Precision Duplication With Comparison (RP-DWC), which executes a redundant copy in a reduced precision. RP-DWC delivers excellent fault coverage, up to 86%, with minimal execution time and energy consumption overheads (13% and 24%, respectively). (2) Dedicated software solutions for hardening Convolutional Neural Networks (CNNs) that can correct up to 98% of the CNN errors.
Fernando Santos 0001, Luigi Carro, Paolo Rech
ETS3
2023 Understanding and Improving GPUs' Reliability Combining Beam Experiments with Fault Simulation
abstract
Graphics Processing Units (GPUs) are essential in High Performance Computing (HPC) and safety-critical applications like autonomous vehicles. This market shift led to significant improvements in the programming frameworks and evaluation tools and concerns about their reliability. However, GPUs' high complexity poses challenges in evaluating their reliability. We conducted the first cross-layer GPU reliability evaluation to unveil and mitigate GPU vulnerabilities. The proposed evaluation is achieved by comparing and combining extensive neutron beam experiments, fault simulation campaigns, and application profiling. Based on this detailed analysis, a novel methodology to accurately estimate GPUs application FIT rate is proposed. The cross-layer evaluation enables two novel hardening solutions: (1) Reduced Precision Duplication With Comparison (RP-DWC) executes a redundant copy in reduced precision. RP-DWC delivers excellent fault coverage, up to 86%, with minimal execution time and energy consumption overheads (13% and 24%, respectively). (2) Dedicated software solutions for hardening Convolutional Neural Networks (CNNs) can detect up to 98% of errors.
Fernando Santos 0001, Luigi Carro, Paolo Rech
ITC3
2023 Understanding the Effects of Permanent Faults in GPU's Parallelism Management and Control Units
abstract
Modern Graphics Processing Units (GPUs) demand life expectancy extended to many years, exposing the hardware to aging (i.e., permanent faults arising after the end-of-manufacturing test). Hence, techniques to assess permanent fault impacts in GPUs are strongly required, especially in safety-critical domains.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Fernando Santos 0001, Matteo Sonza Reorda, Paolo Rech
SC5
2023 Efficient Error Detection for Matrix Multiplication With Systolic Arrays on FPGAs
abstract
Matrix multiplication has always been a cornerstone in computer science. In fact, linear algebra tools permeate a wide variety of applications: from weather forecasting, to financial market prediction, radio signal processing, computer vision, and more. Since many of the aforementioned applications typically impose strict performance and/or fault tolerance constraints, the demand for fast and reliable matrix multiplication (MxM) is at an all-time high. Typically, increased reliability is achieved through redundancy. However, coarse-grain duplication incurs an often prohibitive overhead, higher than 100%. Thanks to the peculiar characteristics of the MxM algorithm, more efficient algorithm-based hardening solutions have been designed to detect (and even correct) some types of errors with lower overhead. We show that, despite being more efficient, current solutions are still sub-optimal in certain scenarios, particularly when considering persistent faults in Field-Programmable Gate-Arrays (FPGAs). Based on a thorough analysis of the fault model, we propose an error detection technique for MxM that decreases both algorithmic and architectural costs by over a polynomial degree, when compared to existing algorithm-based strategies. Furthermore, we report arithmetic overheads at the application level to be under 1% for three state-of-the-art Convolutional Neural Networks (CNNs).
Fabiano Libano, Paolo Rech, John S. Brunhaver
IEEE Trans. Computers2
2022 Reliability of Google's Tensor Processing Units for Embedded Applications
abstract
Convolutional Neural Networks (CNNs) have become the most used and efficient way to identify and classify objects in a scene. CNNs are today fundamental not only for autonomous vehicles, but also for Internet of Things (IoT) and smart cities or smart homes. Vendors are developing low-power, efficient, and low-cost dedicated accelerators to allow the execution of the computational-demanding CNNs even in embedded applications with strict power and cost budgets. Google's Coral Tensor Processing Unit (TPU) is one of the latest low power accelerators for CNNs. In this paper we investigate the reliability of TPUs to atmospheric neutrons, reporting experimental data equivalent to more than 30 million years of natural irradiation. We analyze the behavior of TPUs executing atomic operations (standard or depthwise convolutions) with increasing input sizes as well as eight CNN designs typical of embedded applications, including transfer learning and reduced data-set configurations. We found that, despite the high error rate, most neutrons-induced errors only slightly modify the convolution output and do not change the CNNs detection or classification. By reporting details about the fault model and error rate, we provide valuable information on how to evaluate and improve the reliability of CNNs executed on a TPU.
Rubens Luiz Rech Junior, Paolo Rech
DATE2
2022 QuFI: a Quantum Fault Injector to Measure the Reliability of Qubits and Quantum Circuits
abstract
Quantum computing is an up-and-coming technology that is expected to revolutionize the computation paradigm in the next few years. Qubits, the primary computing elements of quantum circuits, exploit the quantum physics proprieties to increase the parallelism and speed of computation drastically. Unfortunately, besides being intrinsically noisy, qubits have also been shown to be highly susceptible to external sources of faults, such as ionizing radiation. The latest discoveries highlight a much higher radiation sensitivity of qubits than traditional transistors and identify a much more complex fault model than bit-flip.We propose a framework to identify the quantum circuits sensitivity to radiation-induced faults and the probability for a fault in a qubit to propagate to the output. Based on the latest studies and radiation experiments performed on real quantum machines, we model the transient faults in a qubit as a phase shift with a parametrized magnitude. Additionally, our framework can inject multiple qubit faults, tuning the phase shift magnitude based on the proximity of the qubit to the particle strike location. As we show in the paper, the proposed fault injector is highly flexible, and it can be used on both quantum circuit simulators and real quantum machines. We report the finding of more than 285, 249, 536 injections on the Qiskit simulator and 53, 248 injections on real IBM machines. We consider three quantum algorithms and identify the faults and qubits that are more likely to impact the output. We also consider the fault propagation dependence on the circuit scale, showing that the reliability profile for some quantum algorithms is scale-dependent, with increased impact from radiation-induced faults as we increase the number of qubits. Finally, we also consider multi qubits faults, showing that they are much more critical than single faults. The fault injector and the data presented in this paper are available in a public repository to allow further analysis.
Daniel Oliveira 0002, Edoardo Giusto, Emanuele Dri, Nadir Casciola, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech
DSN8
2022 Understanding the Impact of Cutting in Quantum Circuits Reliability to Transient Faults
abstract
Quantum Computing is a highly promising new computation paradigm. Unfortunately, quantum bits (qubits) are extremely fragile and their state can be gradually or suddenly modified by intrinsic noise or external perturbation. In this paper, we target the sensitivity of quantum circuits to radiation-induced transient faults. We consider quantum circuit cuts that split the circuit into smaller independent portions, and understand how faults propagate in each portion. As we show, the cuts have different vulnerabilities, and our methodology successfully identifies the circuit portion that is more likely to contribute to the overall circuit error rate. Our evaluation shows that a circuit cut can have a 4.6 x higher probability than the other cuts, when corrupted, to modify the circuit output. Our study, identifying the most critical cuts, moves towards the possibility of implementing a selective hardening for quantum circuits.
Nadir Casciola, Edoardo Giusto, Emanuele Dri, Daniel Oliveira 0002, Paolo Rech, Bartolomeo Montrucchio
IOLTS5
2022 Transient-Fault-Aware Design and Training to Enhance DNNs Reliability with Zero-Overhead
abstract
Deep Neural Networks (DNNs) enable a wide series of technological advancements, ranging from clinical imaging, to predictive industrial maintenance and autonomous driving. However, recent findings indicate that transient hardware faults may corrupt the models prediction dramatically. For instance, the radiation-induced misprediction probability can be so high to impede a safe deployment of DNNs models at scale, urging the need for efficient and effective hardening solutions. In this work, we propose to tackle the reliability issue both at training and model design time. First, we show that vanilla models are highly affected by transient faults, that can induce a performances drop up to 37%. Hence, we provide three zero-overhead solutions, based on DNN re-design and re-train, that can improve DNNs reliability to transient faults up to one order of magnitude. We complement our work with extensive ablation studies to quantify the gain in performances of each hardening component.
Niccolò Cavagnero, Fernando Santos 0001, Marco Ciccone, Giuseppe Averta, Tatiana Tommasi, Paolo Rech
IOLTS6
2022 A Multi-level Approach to Evaluate the Impact of GPU Permanent Faults on CNN's Reliability
abstract
Graphics processing units (GPUs) are widely used to accelerate Artificial Intelligence applications, such as those based on Convolutional Neural Networks (CNNs). Since in some domains in which CNNs are heavily employed (e.g., automotive and robotics) the expected lifetime of GPUs is over ten years, it is of paramount importance to study the impact of permanent faults (e.g. due to aging). Crucially, while the impact of transient faults on GPUs running CNNs has been widely studied, an accurate evaluation of the impact of permanent faults is still lacking. Performing this evaluation is challenging due to the complexity of GPU devices and the software implementing a CNN. In this work, we propose a methodology that combines the accuracy of gate-level fault simulation with the speed and flexibility of software fault injection to evaluate the effects of permanent hardware faults affecting a GPU. First, we profile the executed low-level GPU instructions during the CNN inference. Then, using extensive gate-level fault injection campaigns, we provide an accurate analysis of the effects of permanent faults on the internal modules executing the targeted instructions. Finally, we propagate these effects using fast software-based fault injection. The method allows, for the first time, to estimate the percentage of permanent faults leading the CNN to produce wrong results (i.e., changing the result of its work). The method's feasibility, which allows for flexibly trade-off accuracy with the required computational effort, is shown using LeNet running on an Ampere Nvidia GPU as a case study. The method reduces the computational effort for the evaluation by several orders of magnitude with respect to plain gate- and RTL-level faults simulation.
Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Fernando Santos 0001, Matteo Sonza Reorda, Paolo Rech
ITC5
2022 Soft Error Effects on Arm Microprocessors: Early Estimations versus Chip Measurements
abstract
Extensive research efforts are being carried out to evaluate and improve the reliability of computing devices either through beam experiments or simulation-based fault injection. Unfortunately, it is still largely unclear to which extend fault injection can provide an accurate error rate estimation at early stages and if beam experiments can be used to identify the weakest resources in a device. The importance and challenges associated with a timely, but yet realistic reliability evaluation grow with the increase of complexity in both the hardware domain, with the integration of different types of cores in an SoC (System-on-Chip), and the software domain, with the OS (operating system) required to take full advantage of the available resources. In this paper, we combine and analyze data gathered with extensive beam experiments (on thefinalphysical CPU hardware) and microarchitectural fault injections (onearlymicroarchitectural CPU models). We target a standalone Arm Cortex-A5 CPU and an Arm Cortex-A9 CPU integrated into an SoC and evaluate their reliability in bare-metal and Linux-based configurations. Combining experimental data that covers more than 18 million years of device time with the result of more than 176,000 injections we find that both the SoC integration and the presence of the OS increase the system DUEs (Detected Unrecoverable Errors) rate (for different reasons) but do not significantly impact the SDCs (Silent Data Corruptions) rate which is solely attributed to the CPU core. Our reliability analysis demonstrates that even considering SoC integration and OS inclusion, early, pre-silicon microarchitecture-level fault injection delivers accurate SDC rates estimations and lower bounds for the DUE rates.
Pablo Bodmann, George Papadimitriou 0001, Rubens Luiz Rech Junior, Dimitris Gizopoulos, Paolo Rech
IEEE Trans. Computers5
2022 Reduced Precision DWC: An Efficient Hardening Strategy for Mixed-Precision Architectures
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing devices. However, it introduces performance and energy consumption overheads that could be unsuitable for high-performance computing or real-time safety-critical applications. In this article, we present Reduced-Precision Duplication with Comparison (RP-DWC) as a means to lower the overhead of DWC by executing the redundant copy in reduced precision. RP-DWC is particularly suitable for modern mixed-precision architectures, such as NVIDIA GPUs, that feature dedicated functional units for computing with programmable accuracy. We discuss the benefits and challenges associated with RP-DWC and show that the intrinsic difference between the mixed-precision copies allows for detecting most, but not all, errors. However, as the undetected faults are the ones that fall into the difference between precisions, they are the ones that produce a much smaller impact on the application output and, thus, might be tolerated. We investigate RP-DWC impact into fault detection, performance, and energy consumption on Volta GPUs. Through fault injection and beam experiment, using three microbenchmarks and four real applications, we show that RP-DWC achieves an excellent coverage (up to 86 percent) with minimal overheads (as low as 0.1 percent time and 24 percent energy consumption overhead).
Fernando Santos 0001, Marcelo Brandalero, Michael B. Sullivan 0001, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IEEE Trans. Computers7
2021 Revealing GPUs Vulnerabilities by Combining Register-Transfer and Software-Level Fault Injection
abstract
The complexity of both hardware and software makes GPUs reliability evaluation extremely challenging. A low level fault injection on a GPU model, despite being accurate, would take a prohibitively long time (months to years), while software fault injection, despite being quick, cannot access critical resources for GPUs and typically uses synthetic fault models (e.g., single bit-flips) that could result in unrealistic evaluations. This paper proposes to combine the accuracy of Register- Transfer Level (RTL) fault injection with the efficiency of software fault injection. First, on an RTL GPU model (FlexGripPlus), we inject over 1.5 million faults in low-level resources that are unprotected and hidden to the programmer, and characterize their effects on the output of common instructions. We create a pool of possible fault effects on the operation output based on the instruction opcode and input characteristics. We then inject these fault effects, at the application level, using an updated version of a software framework (NVBitFI). Our strategy reduces the fault injection time from the tens of years an RTL evaluation would need to tens of hours, thus allowing, for the first time on GPUs, to track the fault propagation from the hardware to the output of complex applications. Additionally, we provide a more realistic fault model and show that single bit-flip injection would underestimate the error rate of six HPC applications and two convolutional neural networks by up to 48parcent (18parcent on average). The RTL fault models and the injection framework we developed are made available in a public repository to enable third-party evaluations and ease results reproducibility.
Fernando Santos 0001, Josie E. Rodriguez Condia, Luigi Carro, Matteo Sonza Reorda, Paolo Rech
DSN5
2021 Protecting GPU's Microarchitectural Vulnerabilities via Effective Selective Hardening
abstract
Graphics Processing Units (GPUs) are today adopted in several domains for which reliability is fundamental, such as self-driving cars and autonomous machines. Unfortunately, on one side GPUs have been shown to have a high error rate and, on the other side, the constraints imposed by real-time safety-critical applications make traditional, costly, replication-based hardening solutions inadequate. This paper proposes an effective microarchitectural selective hardening of GPU modules to mitigate those faults that affect instructions correct execution. We first characterize, through Register-Transfer Level (RTL) fault injections, the architectural vulnerabilities of a GPU model (FlexGripPlus). We specifically target transient faults in the functional units and pipeline registers of a GPU core. Then, we apply selective hardening by triplicating the locations in each module that we found to be more critical. The results show that selective hardening using Triple Modular Redundancy (TMR) can correct 85% to 99% of faults in the pipeline registers and from 50% to 100% of faults in the functional units. The proposed selective TMR strategy reduces the hardware overhead by up to 65% when compared with traditional TMR.
Josie E. Rodriguez Condia, Paolo Rech, Fernando Santos 0001, Luigi Carro, Matteo Sonza Reorda
IOLTS2
2021 Demystifying GPU Reliability: Comparing and Combining Beam Experiments, Fault Simulation, and Profiling
abstract
Graphics Processing Units (GPUs) have moved from being dedicated devices for multimedia and gaming applications to general-purpose accelerators employed in High-Performance Computing (HPC) and safety-critical applications such as autonomous vehicles. This market shift led to a burst in the GPU's computing capabilities and efficiency, significant improvements in the programming frameworks and performance evaluation tools, and a concern about their hardware reliability. In this paper, we compare and combine high-energy neutron beam experiments that account for more than 13 million years of natural terrestrial exposure, extensive architectural-level fault simulations that required more than 350 GPU hours (using SASSIFI and NVBitFI), and detailed application-level profiling. Our main goal is to answer one of the fundamental open questions in GPU reliability evaluation: whether fault simulation provides representative results that can be used to predict the failure rates of workloads running on GPUs. We show that, in most cases, fault simulation-based prediction for silent data corruptions is sufficiently close (differences lower than 5×) to the experimentally measured rates. We also analyze the reliability of some of the main GPU functional units (including mixed-precision and tensor cores). We find that the way GPU resources are instantiated plays a critical role in the overall system reliability and that faults outside the functional units generate most detectable errors.
Fernando Santos 0001, Siva Kumar Sastry Hari, Pedro Martins Basso, Luigi Carro, Paolo Rech
IPDPS5
2021 The Impact of SoC Integration and OS Deployment on the Reliability of Arm Processors
abstract
Arm CPU architectures, thanks to their efficiency and flexibility, have been widely adopted in portable user devices such as smartphones, tablets, and laptops. Recently, the high computing efficiency, together with the unique possibility that Arm offers to adapt the architecture for a specific application, pushed the adoption of Arm-based systems both in HPC (High Performance Computing) applications and autonomous vehicles. The possibility of modifying Arm architecture can potentially be extremely beneficial, as selective fault tolerance solutions can be added at the microarchitectural level. The current trend in the design of computing devices is to integrate several functionalities on the same SoC (System-On-Chip). Modern SoCs usually integrate (one or more) CPUs and (one or more) accelerators, such as GPUs (Graphics Processing Units) or FPGAs (Field Programmable Gate Arrays). These SoCs typically allow the computing cores to share common memories, which significantly improves the performance and reduces the total power consumption but may impact the system's reliability. In this work, we have evaluated the impact of SoC integration and OS deployment using beam experiments and microarchitectural fault injection.
Pablo Bodmann, George Papadimitriou 0001, Dimitris Gizopoulos, Paolo Rech
ISPASS4
2021 Combining Architectural Simulation and Software Fault Injection for a Fast and Accurate CNNs Reliability Evaluation on GPUs
abstract
Graphic Processing Units (GPUs) are commonly used to accelerate Convolutional Neural Networks (CNNs) for object detection and classification. As CNNs are employed in safety-critical applications, such as autonomous vehicles, their reliability must be carefully evaluated. In this work, we combine the accuracy of microarchitectural simulation with the speed of software fault injection to investigate the reliability of CNNs executed in GPUs. First, with a detailed microarchitectural fault injection on a GPU model (FlexGripPlus), we characterize the effects of faults in critical and user-hidden modules (such as the Warp Scheduler and the Pipeline Registers) in the computation of convolution over a suitably selected subset of tiles. Then, with software fault injection, we propagate the fault effects in the CNN. Thanks to our approach we are able, for the first time, to analyze the impact of faults affecting GPUs' hidden modules on a whole CNN execution (LeNET) without undermining the reliability evaluation correctness.
Josie E. Rodriguez Condia, Fernando Santos 0001, Matteo Sonza Reorda, Paolo Rech
VTS4
2021 Collaborative execution of fluid flow simulation using non-uniform decomposition on heterogeneous architectures
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.4
2021 Thermal neutrons: a possible threat for supercomputer reliability
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Andrea Favalli, Opale Schappert, Stephen Wender, Carlo Cazzaniga, Christopher Frost 0002, Paolo Rech
J. Supercomput.12
2020 Thermal Neutrons: a Possible Threat for Supercomputers and Safety Critical Applications
abstract
The high performance, high efficiency, and low cost of Commercial Off-The-Shelf (COTS) devices make them attractive for applications with strict reliability constraints. Today, COTS devices are adopted in HPC and safety-critical applications such as autonomous driving. Unfortunately, the cheap natural Boron widely used in COTS chip manufacturing process makes them highly susceptible to thermal (low energy) neutrons. In this paper, we demonstrate that thermal neutrons are a significant threat to COTS device reliability. For our study, we consider an AMD APU, three NVIDIA GPUs, an Intel accelerator, and an FPGA executing a relevant set of algorithms. We consider different scenarios that impact the thermal neutron flux such as weather, concrete walls and floors, and HPC liquid cooling systems. We show that thermal neutrons FIT rate could be comparable to the high energy neutron FIT rate.
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Carlo Cazzaniga, Christopher Frost 0002, Robert C. Baumann, Paolo Rech
ETS10
2020 Reduced-Precision DWC for Mixed-Precision GPUs
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing systems, including Graphics Processing Units (GPUs). DWC, however, introduces performance and energy consumption overheads that could be unacceptable for High-Performance Computing (HPC) or real-time safety-critical applications. In this work, we propose Reduced-Precision DWC (RP-DWC): an improvement over the traditional DWC approach that uses mixed-precision GPUs hardware resources to implement fault detection. We investigate, through both fault injection campaigns and accelerated neutron beam experiments, the impact of RPDWC onto performance, energy consumption, and its fault detection capabilites. We show that RP-DWC achieves on average 74% fault coverage (up to 86%) with very small overheads (0.1% time and 24% energy consumption overhead, in the best case).
Fernando Santos 0001, Marcelo Brandalero, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IOLTS6
2019 Identifying the Most Reliable Collaborative Workload Distribution in Heterogeneous Devices
abstract
The constant need of higher performances and reduced power consumption has lead vendors to design heterogeneous devices that embed traditional CPU and an accelerator, like a GPU or FPGA. When the CPU and the accelerator are used collaboratively the device computational performances reach their peak. However, the higher amount of resources employed for computation has, potentially, the side effect of increasing soft error rate. In this paper we evaluate the reliability behavior of AMD Kaveri Accelerated Processing Units executing a set of heterogeneous applications. We distribute the workload between the CPU and GPU and evaluate which configuration provides the lowest error rate or allows the computation of the highest amount of data before experiencing a failure. We show that, in most cases, the most reliable workload distribution is the one that delivers the highest performances. As experimentally proven, by choosing the correct workload distribution the device reliability can increase of up to 9x.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
DATE4
2019 Demystifying Soft Error Assessment Strategies on ARM CPUs: Microarchitectural Fault Injection vs. Neutron Beam Experiments
abstract
Fault injection in early microarchitecture-level simulation CPU models and beam experiments on the final physical CPU chip are two established methodologies to access the soft error reliability of a microprocessor at different stages of its design flow. Beam experiments, on one hand, estimate the devices expected soft error rate in realistic physical conditions by exposing it to accelerated particles fluxes. Fault injection in microarchitectural models of the processor, on the other hand, provides deep insights on faults propagation through the entire system stack, including the operating system. Combining beam experiments and fault injection data can deliver deep insights about the devices expected reliability when deployed in the field. However, it is yet largely unclear if the fault injection error rates can be compared to those reported by beam experiments and how this comparison can lead to informed soft error protection decisions in early stages of the system design. In this paper, we present and analyze data gathered with extensive beam experiments (on physical CPU hardware) and microarchitectural fault injections (on an equivalent CPU model on Gem5) performed with 13 different benchmarks executed on top of Linux on an ARM Cortex-A9 microprocessor. We combine experimental data that cover more than 2.9 million years of natural exposure with the result of more than 80,000 injections. We then compare the soft error rate estimations that are based on neutron beam and fault injection experiments. We show that, for most benchmarks, fault injection can be very accurately used to predict the Silent Data Corruptions (SDCs) rate and the Application Crash rate. The System Crash rate measured with beam experiments, however is much larger than the one estimated by fault injection due to unknown proprietary parts of the physical hardware platform that can't be modeled in the simulator. Overall, our analysis shows that the relative difference between the total error rates of the beam experiments and the fault injection experiments is limited within a narrow range of values and is always smaller than one order of magnitude. This narrow range of the expected failure rate of the CPU provides invaluable assistance to the designers in making effective soft error protection decisions in early design stages.
Athanasios Chatzidimitriou, Pablo Bodmann, George Papadimitriou 0001, Dimitris Gizopoulos, Paolo Rech
DSN5
2019 Impact of Reduced Precision in the Reliability of Deep Neural Networks for Object Detection
abstract
Modern Graphics Processing Units (GPUs) have dedicated hardware to execute floating-point operations with different precisions (64-bit double, 32-bit single, and 16-bit half). Using reduced precision for specific applications like Deep Neural Networks (DNNs) has been shown to reduce both the execution time and power consumption with negligible effects on the DNNs' accuracy. As GPUs are playing a critical role in DNN for object detection and get into safety-critical environments, their reliability is becoming a growing concern. In this paper, we evaluate the reliability of a DNN implemented in three different precisions (half, single, and double) on NVIDIA mixed-precision GPUs. We evaluate not only the error rate of the applications but also the effects of the errors on the final detection. We perform extensive fault-injection campaign on the register file of NVIDIA mixed-precision GPUs. We found that reducing data and operation precision increases the probability for the fault to impact the DNN detection and classification. Then, we complement the fault injection study with beam experiments. We exposed YOLOv3 running on Tesla V100s to neutron beams and found that the use of half precision reduces the error rate of up to 2x. The smaller exposed area and improved performances brought by reduced precision is then likely to increase the DNN reliability.
Fernando Santos 0001, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
ETS4
2019 Reliability Evaluation of Mixed-Precision Architectures
abstract
Novel computing architectures offer the possibility to execute float point operations with different precisions. The execution of reduced precision operations, when acceptable for certain applications, is likely to reduce both the execution time and the power consumption. However, the application's error rate and the device's reliability can also be impacted by these precision changes. In this paper, we study the impact of data and operation precision changes on the reliability of modern architectures. We consider Xilinx Field-Programmable Gate-Arrays (FPGA), Intel Xeon Phis, and NVIDIA Graphics Processing Units (GPUs) executing a set of codes implemented in double, single, and half-precision IEEE754-compliant float point data. On FPGAs, the reduced area and performance improvements brought by reduced precision operations increase reliability. On Xeon Phis the compiler biases significantly double and single-precision instructions execution. This raises the drawback of increasing single-precision error rates when compared to double-precision operations. NVIDIA GPUs make use of dedicated mixed-precision cores, which draw nontrivial effects on the device reliability. Generically speaking, on GPUs half-precision allows a higher number of executions to be correctly completed before experimenting a failure. Finally, we also evaluate how transient faults impact the output correctness. Our study shows that for most applications faults in a single or half-precision data or operation are more likely to significantly modify the output value than errors in double-precision data.
Fernando Santos 0001, Caio B. Lunardi, Daniel Oliveira 0002, Fabiano Libano, Paolo Rech
HPCA5
2019 Detecting Errors in Convolutional Neural Networks Using Inter Frame Spatio-Temporal Correlation
abstract
Object detection, a critical feature for autonomous vehicles, is performed today using Convolutional Neural Networks (CNNs). Errors in a CNN execution can modify the way the vehicle sense the surrounding environment, potentially causing accidents or unexpected behaviors. The high computational requirements of CNNs combined with the need to perform detection in real-time allow little margin for implementing error detection. In this paper, we present an extremely efficient error detection solution for CNN based on the observation that, in the absence of errors, the differences between the input frames and the detection provided by the CNN should be strictly correlated. In other words, if the image between two subsequent frames does not change significantly, the detection should also be very similar. Similarly, if the detection varies considerably from a frame to the next, then the input image should also have been different. Whenever input images and output detection don't correlate we can detect a error. After formalizing and evaluating the inter-frame and output correlation thresholds, we implement and validate the detection strategy, utilizing data from previous radiation experiments. Exploiting the intrinsic efficiency in processing images of devices used to execute CNNs, we can detect up to 80% of errors while adding low overhead.
Lucas Draghetti, Fernando Santos 0001, Luigi Carro, Paolo Rech
IOLTS4
2019 Impact of Workload Distribution on Energy Consumption, Performance, and Reliability of Heterogeneous Devices
abstract
Devices integrating cores of different nature in the same chip achieve very high computation efficiency by reducing the power consumption and latencies of moving data from a chip to an external device. Heterogeneous devices are commonly used in high performance computing applications and, lately, their market has expanded from portable and gaming to safety-critical applications. In this paper, we evaluate how the collaborative workload distribution impact the energy consumption, performance, and reliability of heterogeneous devices. Then, we use the Energy-Delay-Fit Product (EDFP) to evaluate the trade off between the measured metrics and find how they correlate. To perform the proposed study we consider AMD Accelerated Processing Units (APUs) that embed a CPU and a GPU. We run four applications, each one representing an algorithm class, gradually distributing the workload from the CPU to the GPU and measuring both the energy consumption and execution time. Then, we take advantage of accelerated neutron beams to measure the realistic error rates of the different workload distributions. As we show in the paper, energy consumption and execution time are mold by the same trend while FIT rates highly depend on algorithm class and workload distribution. Additionally, we found that execution time is the most influencing factor for the device EDFP. The application EDFP varies of up tp 6 orders of magnitude depending on the workload distribution. An unwise configuration can then jeopardize the device efficiency and reliability.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
PDP4
2019 Non-uniform Partitioning for Collaborative Execution on Heterogeneous Architectures
abstract
Since the demand for computing power increases, new architectures arise to obtain better performance. An important class of integrated devices is heterogeneous architectures, which join different specialized hardware into a single chip, composing a System on Chip - SoC. Within this context, effectively splitting tasks between the different architectures is primal to obtain efficiency and performance. In this work, we evaluate two heterogeneous architectures: one composed of a general-purpose CPU and a graphics processing unit (GPU) integrated into a single chip (AMD Kaveri SoC), and another composed by a general-purpose CPU and a Field Programmable Gate Array (FPGA) integrated into a single chip (Intel Arria 10 SoC). We investigate how data partitioning affects the performance of each device in a collaborative execution through the decomposition of the data domain. As a case study, we apply the technique in the well-known Lattice Boltzmann Method (LBM), analyzing the performance of five kernels in both architectures. Our experimental results show that non-uniform partitioning improves LBM kernels performance by up to 11.40% and 15.15% on AMD Kaveri and Intel Arria 10, respectively.
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2019 Analyzing and Increasing the Reliability of Convolutional Neural Networks on GPUs
abstract
Graphics processing units (GPUs) are playing a critical role in convolutional neural networks (CNNs) for image detection. As GPU-enabled CNNs move into safety-critical environments, reliability is becoming a growing concern. In this paper, we evaluate and propose strategies to improve the reliability of object detection algorithms, as run on three NVIDIA GPU architectures. We consider three algorithms: 1) you only look once; 2) a faster region-based CNN (Faster R-CNN); and 3) a residual network, exposing live hardware to neutron beams. We complement our beam experiments with fault injection to better characterize fault propagation in CNNs. We show that a single fault occurring in a GPU tends to propagate to multiple active threads, significantly reducing the reliability of a CNN. Moreover, relying on error correcting codes dramatically reduces the number of silent data corruptions (SDCs), but does not reduce the number of critical errors (i.e., errors that could potentially impact safety-critical applications). Based on observations on how faults propagate on GPU architectures, we propose effective strategies to improve CNN reliability. We also consider the benefits of using an algorithm-based fault-tolerance technique for matrix multiplication, which can correct more than 87% of the critical SDCs in a CNN, while redesigning maxpool layers of the CNN to detect up to 98% of critical SDCs.
Fernando Santos 0001, Pedro Foletto Pimenta, Caio B. Lunardi, Lucas Draghetti, Luigi Carro, David R. Kaeli, Paolo Rech
IEEE Trans. Reliab.7
2018 Evaluating the impact of execution parameters on program vulnerability in GPU applications
abstract
While transient faults continue to be a major concern for the High Performance Computing (HPC) community, we still lack a clear understanding of how these faults propagate in applications. This paper addresses two particular aspects of the vulnerabilities of HPC applications as run on Graphics Processing Units (GPUs): their dependence on input data and on thread-block size. To characterize fault propagation as a function of input parameters, we leverage an ISA-level fault injection framework and carry out an extensive fault injection campaign to characterize the vulnerability of a suite of GPU applications. Our results show that the vulnerability of most of the programs studied are insensitive to changes in input values, except in less common cases when input values were highly biased, i.e., values that exhibit a special vulnerability behavior. For example, the multiplication property of any value with a zero value (zero times any number is equal to zero) makes it a biased input for multiplication operations. Our study also examines the effects of changing the GPU thread-block size and its impact on vulnerability. We found that, similar to performance, the vulnerability of an application can depend on the block size of the kernels in the application. In some applications, we found that the silent data corruption rate can vary by as much as 8% when changing the block size of a kernel.
Fritz Previlon, Charu Kalra, David R. Kaeli, Paolo Rech
DATE4
2018 Code-Dependent and Architecture-Dependent Reliability Behaviors
abstract
The increased need for computing capabilities and higher efficiency have stimulated industries to make available in the market novel architectures with increased complexity. The variety of codes that need to be executed combined with the complexity of novel architectures introduces challenges in the reliability evaluation of computing systems and applications. This paper compares the reliability behaviors of six different architectures (an Intel co-processor, three NVIDIA GPUs, an AMD APU, an embedded ARM) executing eight different codes. To support our evaluation, we present and discuss experimental beam data that covers a total of more than 352,000 years of natural exposure and fault-injection analysis based on a total of more than 120,000 injections. We first quantify both the Silent Data Corruptions and the Detected Unrecoverable Errors rates. Then, we qualify observed errors considering the difference between the corrupted and expected values as well as the portion of the output that has been corrupted. From these analyses, we identify the reliability characteristics which are related to the underlying hardware and the intrinsic behaviors of the executed code. Finally, we discuss the implications of the device- and code-dependent reliability behaviors for approximate computing. We analyze the benefits, in term of reduced error rate, of a relaxed output correctness.
Vinicius Fratin, Daniel Oliveira 0002, Caio B. Lunardi, Fernando Santos 0001, Gennaro Severino Rodrigues, Paolo Rech
DSN6
2018 Predicting the Reliability Behavior of HPC Applications
abstract
The error rate of current High Performance Computing (HPC) systems is already in the order of one per dozens of hours. Understanding the reliability behavior of HPC applications will be required for the next generation of supercomputers. Using the reliability behavior one can select efficient mitigation techniques for the application and fine-tune parameters such as checkpoint frequency. In this paper, we investigate the application of a machine learning model to predict the reliability behavior of HPC applications. We inject faults in more than 30 HPC applications executing in the Intel Xeon Phi Knights Landing (KNL) and use profiling information to build a predictive model with Support Vector Machines (SVM). We show that the model can predict the Program Vulnerability Factor (PVF) with an average relative error of 7% for certain classes of algorithm, such as linear algebra and sorting. The average relative error for all algorithm classes is 22%. Such a fast and straightforward prediction model can be effective as a filter to select the most unreliable applications to perform an in-depth analysis.
Daniel Oliveira 0002, Francis B. Moreira 0001, Paolo Rech, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2018 Special session: How approximate computing impacts verification, test and reliability
abstract
Two AxC techniques have been successfully applied to hardware components. The first one is the functional approximation [1]that modifies the circuit structure replacing the original function F with the function G. G implementation leads to area/energy reduction at the cost of reduced accuracy, meaning that some errors can be observed at the outputs of G. The observed errors are a variation between the output values of F (precise) and G (approximate). The variation is the accuracy loss measured by means of quality metric(s) [1]. The second AxC technique is the over-scaling based approximation. Basically, the HW component is forced to work outside its specified operating conditions [1]. The classical example is the reduction of the supply voltage under the minimum value.
Lukás Sekanina, Zdenek Vasícek, Alberto Bosio, Marcello Traiola, Paolo Rech, Daniel Oliveira 0002, Fernando Santos 0001, Stefano Di Carlo
VTS5
2017 Radiation-Induced Error Criticality in Modern HPC Parallel Accelerators
abstract
In this paper, we evaluate the error criticality of radiation-induced errors on modern High-Performance Computing (HPC) accelerators (Intel Xeon Phi and NVIDIA K40) through a dedicated set of metrics. We show that, as long as imprecise computing is concerned, the simple mismatch detection is not sufficient to evaluate and compare the radiation sensitivity of HPC devices and algorithms. Our analysis quantifies and qualifies radiation effects on applications' output correlating the number of corrupted elements with their spatial locality. Also, we provide the mean relative error (dataset-wise) to evaluate radiation-induced error magnitude. We apply the selected metrics to experimental results obtained in various radiation test campaigns for a total of more than 400 hours of beam time per device. The amount of data we gathered allows us to evaluate the error criticality of a representative set of algorithms from HPC suites. Additionally, based on the characteristics of the tested algorithms, we draw generic reliability conclusions for broader classes of codes. We show that arithmetic operations are less critical for the K40, while Xeon Phi is more reliable when executing particles interactions solved through Finite Difference Methods. Finally, iterative stencil operations seem the most reliable on both architectures.
Daniel Oliveira 0002, Laércio Lima Pilla, Mauricio Hanzich, Vinicius Fratin, Fernando Santos 0001, Caio B. Lunardi, José María Cela, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
HPCA10
2017 Experimental and analytical study of Xeon Phi reliability
abstract
We present an in-depth analysis of transient faults effects on HPC applications in Intel Xeon Phi processors based on radiation experiments and high-level fault injection. Besides measuring the realistic error rates of Xeon Phi, we quantify Silent Data Corruption (SDCs) by correlating the distribution of corrupted elements in the output to the application's characteristics. We evaluate the benefits of imprecise computing for reducing the programs' error rate. For example, for HotSpot a 0.5% tolerance in the output value reduces the error rate by 85%.
Daniel Oliveira 0002, Laércio Lima Pilla, Nathan DeBardeleben, Sean Blanchard, Heather M. Quinn, Israel Koren, Philippe Olivier Alexandre Navaux, Paolo Rech
SC8
2016 Evaluation of Histogram of Oriented Gradients Soft Errors Criticality for Automotive Applications
abstract
Pedestrian detection reliability is a key problem for autonomous or aided driving, and methods that use Histogram of Oriented Gradients (HOG) are very popular. Embedded Graphics Processing Units (GPUs) are exploited to run HOG in a very efficient manner. Unfortunately, GPUs architecture has been shown to be particularly vulnerable to radiation-induced failures. This article presents an experimental evaluation and analytical study of HOG reliability. We aim at quantifying and qualifying the radiation-induced errors on pedestrian detection applications executed in embedded GPUs. We analyze experimental results obtained executing HOG on embedded GPUs from two different vendors, exposed for about 100 hours to a controlled neutron beam at Los Alamos National Laboratory. We consider the number and position of detected objects as well as precision and recall to discriminate critical erroneous computations. The reported analysis shows that, while being intrinsically resilient (65% to 85% of output errors only slightly impact detection), HOG experienced some particularly critical errors that could result in undetected pedestrians or unnecessary vehicle stops. Additionally, we perform a fault-injection campaign to identify HOG critical procedures. We observe that Resize and Normalize are the most sensitive and critical phases, as about 20% of injections generate an output error that significantly impacts HOG detection. With our insights, we are able to find those limited portions of HOG that, if hardened, are more likely to increase reliability without introducing unnecessary overhead.
Fernando Santos 0001, Lucas Weigel, Cláudio R. Jung, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
ACM Trans. Archit. Code Optim.6
2016 Evaluation and Mitigation of Radiation-Induced Soft Errors in Graphics Processing Units
abstract
Graphics processing units (GPUs) are increasingly attractive for both safety-critical and High-Performance Computing applications. GPU reliability is a primary concern for both the automotive and aerospace markets and is becoming an issue also for supercomputers. In fact, the high number of devices in large data centers makes the probability of having at least a device corrupted to be very high. In this paper, we aim at giving novel insights on GPU reliability by evaluating the neutron sensitivity of modern GPUs memory structures, highlighting pattern dependence and multiple errors occurrences. Additionally, a wide set of parallel codes are exposed to controlled neutron beams to measure GPUs operative error rates. From experimental data and algorithm analysis we derive general insights on parallel algorithms and programming approaches reliability. Finally, error-correcting code, algorithm-based fault tolerance, and duplication with comparison hardening strategies are presented and evaluated on GPUs through radiation experiments. We present and compare both the reliability improvement and imposed overhead of the selected hardening solutions.
Daniel Oliveira 0002, Laércio Lima Pilla, Thiago Santini, Paolo Rech
IEEE Trans. Computers4
2016 Beyond Cross-Section: Spatio-Temporal Reliability Analysis
abstract
A computational system employed in safety-critical applications typically has reliability as a primary concern. Thus, the designer focuses on minimizing the device radiation-sensitive area, often leading to performance degradation. In this article, we present a mathematical model to evaluate system reliability in spatial (i.e., radiation-sensitive area) and temporal (i.e., performance) terms and prove that minimizing radiation-sensitive area does not necessarily maximize application reliability. To support our claim, we present an empirical counterexample where application reliability is improved even if the radiation-sensitive area of the device is increased. An extensive radiation test campaign using a 28 nm commercial-off-the-shelf ARM-based SoC was conducted, and experimental results demonstrate that, while executing the considered application at military aircraft altitude, the probability of executing a two-year mission workload without failures is increased by 5.85% if L1 caches are enabled (thus increasing the radiation-sensitive area) when compared to no cache level being enabled. However, if both L1 and L2 caches are enabled, the probability is decreased by 31.59%.
Thiago Santini, Paolo Rech, Gabriel L. Nazar, Flávio Rech Wagner
ACM Trans. Embed. Comput. Syst.2
2015 Exploiting cache conflicts to reduce radiation sensitivity of operating systems on embedded systems
abstract
In this paper, we investigate how the presence of a general purpose operating system influences the reliability of modern embedded Systems-on-Chips (SoCs). We analytically study the difference in the reliability of SoCs when executing the application bare to the metal and on top of the Linux kernel. Our analysis demonstrates that Linux presence barely affects the Silent Data Corruption rate while it greatly increases the system Functional Interruption (FI) rate (up to 7.48 times) if no preventive measures are taken. Furthermore, we analytically show that cache conflicts between the operating system and application can significantly reduce the Linux-induced FI rate increase. To support our analysis, a total of four representative embedded applications were individually executed bare to the metal and on top of Linux on a 28nm ARM-based SoC exposed to an accelerated neutron beam. Our experimental results demonstrate that, by carefully tuning cache conflicts, it is possible to successfully limit the Linux-induced FI rate increase to 3.85 times. The proposed solution is general and readily applied to a broad set of applications and embedded systems.
Thiago Santini, Paolo Rech, Luigi Carro, Flávio Rech Wagner
CASES2
2015 Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
abstract
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.
Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland
HPCA5
2015 Field, experimental, and analytical data on large-scale HPC systems and evaluation of the implications for exascale system design
abstract
Reliability is an issue for today's large scale computing systems designers, producers, and users. As we approach exascale, the resilience challenge will become critical due to increase in system-scale. It is then fundamental to understand the nature of errors, evaluate their probability of occurrence, and improve the design to reduce their impact on the overall system. In the paper we will present experimental, field, and analytical data to characterize and quantify errors on accelerators, providing a thorough understanding of errors impact on today and future large-scale systems.
Nathan DeBardeleben, Sean Blanchard, David R. Kaeli, Paolo Rech
VTS4
2014 GPGPUs: How to combine high computational power with high reliability
abstract
GPGPUs are used increasingly in several domains, from gaming to different kinds of computationally intensive applications. In many applications GPGPU reliability is becoming a serious issue, and several research activities are focusing on its evaluation. This paper offers an overview of some major results in the area. First, it shows and analyzes the results of some experiments assessing GPGPU reliability in HPC datacenters. Second, it provides some recent results derived from radiation experiments about the reliability of GPGPUs. Third, it describes the characteristics of an advanced fault-injection environment, allowing effective evaluation of the resiliency of applications running on GPGPUs.
Leonardo Arturo Bautista-Gomez, Franck Cappello, Luigi Carro, Nathan DeBardeleben, Bo Fang 0002, Sudhanva Gurumurthi, Karthik Pattabiraman, Paolo Rech, Matteo Sonza Reorda
DATE8
2014 Radiation Sensitivity of High Performance Computing Applications on Kepler-Based GPGPUs
abstract
In this paper we assess and discuss the radiation sensitivity of a set of HPC applications executed on NVIDIA K20 GPGPUs. The occurrence of both radiation-induced silent data corruption and functional interruption will be experimentally addressed for Hotspot, LavaMD, and Matrix Transponse. Each of the tested codes requires a proper computational power and elaborates a different amount of data. Both these characteristics play a significant role in the application radiations sensitivity. Additionally, an evaluation of the error rate at sea level will be provided for all the tested codes.
Daniel Oliveira 0002, Caio B. Lunardi, Laércio Lima Pilla, Paolo Rech, Philippe Olivier Alexandre Navaux, Luigi Carro
DSN4
2014 Impact of GPUs Parallelism Management on Safety-Critical and HPC Applications Reliability
abstract
Graphics Processing Units (GPUs) offer high computational power but require high scheduling strain to manage parallel processes, which increases the GPU cross section. The results of extensive neutron radiation experiments performed on NVIDIA GPUs confirm this hypothesis. Reducing the application Degree Of Parallelism (DOP) reduces the scheduling strain but also modifies the GPU parallelism management, including memory latency, thread registers number, and the processors occupancy, which influence the sensitivity of the parallel application. An analysis on the overall GPU radiation sensitivity dependence on the code DOP is provided and the most reliable configuration is experimentally detected. Finally, modifying the parallel management affects the GPU cross section but also the code execution time and, thus, the exposure to radiation required to complete computation. The Mean Workload and Executions Between Failures metrics are introduced to evaluate the workload or the number of executions computed correctly by the GPU on a realistic application.
Paolo Rech, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Luigi Carro
DSN1
2014 Aging and voltage scaling impacts under neutron-induced soft error rate in SRAM-based FPGAs
abstract
This work investigates the effects of aging and voltage scaling in neutron-induced bit-flip in SRAM-based FPGAs. Experimental results show that aging and voltage scaling can increase in at least two times the susceptibility of SRAM-based FPGAs to Soft Error Rate (SER). These results are innovative, because they combine three real effects that occur in programmable circuits operating at ground-level applications. In addition, a model at electrical simulation for aging, soft error and different voltages was described to investigate the effects observed at the practical neutron irradiation experiment. Results can guide designers to predict soft error effects during the lifetime of devices operating in different power supply mode.
Fernanda Lima Kastensmidt, Jorge L. Tonfat, Thiago Hanna Both, Paolo Rech, Gilson I. Wirth, Ricardo Augusto da Luz Reis, Florent Bruguier, Pascal Benoit, Lionel Torres, Christopher Frost 0002
ETS4
2014 Reducing embedded software radiation-induced failures through cache memories
abstract
Cache memories are traditionally disabled in space-level and safety-critical applications, since it was believed that the sensitive area they introduce would compromise the system reliability. As technology has evolved, the speed gap between logic and main memory has increased in such a way that disabling caches slows the code much more than in the past. As a result, the processor is exposed for a much longer time in order to compute the same workload. In this paper we demonstrate that, on modern embedded processors, enabling caches may bring benefits to critical systems: the larger exposed area may be compensated by the shorter exposure time, leading to an overall improved reliability. We describe the Mean Workload Between Failures, an intuitive metric to evaluate the impact of enabling caches for a given generic application error rate. The proposed metric is experimentally validated through an extensive radiation test campaign using a 28 nm off-the-shelf ARM-based SoC as a case study. The failure probability of the bare-metal application is decreased when the L1 cache is enabled but increased when L2 is also enabled. We also discuss when L2 caches could make the device more reliable.
Thiago Santini, Paolo Rech, Gabriel L. Nazar, Luigi Carro, Flávio Rech Wagner
ETS2
2014 Fault injection in GPGPU cores to validate and debug robust parallel applications
abstract
General Purpose Graphic Processing Units (GPGPUs) are more efficient than CPUs for processing parallel data. Unfortunately, GPGPUs are sensible to radiation. Hence, several software mitigation techniques, as well as robust algorithms, are being developed to overcome reliability problems. In this paper we propose a software debugger-based fault injection mechanism to evaluate the resiliency of applications running on a GPGPU and to validate the software hardening techniques it possibly embeds. We report some experimental results gathered on selected case studies to show the proposed approach advantages and limitations.
M. De Carvalho, Davide Sabena, Matteo Sonza Reorda, Luca Sterpone, Paolo Rech, Luigi Carro
IOLTS5
2014 Power dissipation effects on 28nm FPGA-based System on Chips neutron sensitivity
abstract
Modern System on Chips (SoCs) and embedded electronic devices work at very high frequencies, which have the countermeasure of increasing the power dissipation and, consequently, the silicon die temperature. The presented radiation experiments on a 28nm FPGA-based SoC demonstrate that the temperature variation caused by a higher operating frequency affects the FPGA configuration memory cross section. An evaluation and discussion of the observed reliability dependence on power dissipation effects on practical application is also presented.
Giovanni Bruni, Paolo Rech, Lucas A. Tambara, Gabriel L. Nazar, Fernanda Lima Kastensmidt, Ricardo Augusto da Luz Reis, Alessandro Paccagnella
VLSI-SoC2
2014 GPUs Neutron Sensitivity Dependence on Data Type
Paolo Rech, Christopher Frost 0002, Luigi Carro
J. Electron. Test.1
2013 Experimental evaluation of thread distribution effects on multiple output errors in GPUs
abstract
Graphic Processing Units are very prone to be corrupted by neutrons. Experimental results show that in the majority of the cases a typical application like matrix multiplication is affected by multiple output errors. In this paper we evaluate how different thread distributions impact the multiple output errors occurrence. The reported results and the performed architecture analysis give practical programming advices that may increase the reliability of a generic parallel algorithm without introducing any hardware or computation overhead.
Paolo Rech, Caroline Aguiar, Christopher Frost 0002, Luigi Carro
ETS1
2013 Experimental evaluation of GPUs radiation sensitivity and algorithm-based fault tolerance efficiency
abstract
Experimental results demonstrate that Graphic Processing Units are very prone to be corrupted by neutrons. We have performed several experimental campaigns at ISIS, UK and at LANSCE, Los Alamos, NM, USA accessing the sensitivity of the GPU internal resources as well as the error rate of common parallel algorithms. Experiments highlight output error patterns and radiation responses that can be fruitfully used to design optimized Algorithm-Based Fault Tolerance strategies and provide pragmatic programming guidelines to increase the code reliability with low computational overhead.
Paolo Rech, Luigi Carro
IOLTS1
2012 Neutron radiation test of graphic processing units
abstract
This paper reports and analyzes the results of neutrons radiation testing campaigns on a modern commercial-off-the-shelf Graphic Processing Unit (GPU). A set of guidelines for accelerated radiation experiments on CPUs is presented, emphasizing the shrewdness necessary to ease the test and gain meaningful data. Radiation test results are presented and discussed, highlighting the neutrons sensitivities of the different GPU memory and logic resources in terms of Failure In Time (FIT) due to neutrons at sea level.
Paolo Rech, Caroline Aguiar, Ronaldo Rodrigues Ferreira, Christopher Frost 0002, Luigi Carro
IOLTS1
2010 A Memory Fault Simulator for Radiation-Induced Effects in SRAMs
abstract
This paper introduces a simulator that allows analyzing the radiation induced errors on memory devices. The simulator takes all the radiation effects on SRAM into account and can be easily tuned on the base of data gained during radiation experiments and/or presented in literature. We also present a case study application in which the proposed simulator is used to validate a low-cost hardware platform for soft error detection in avionic environment by the mean of atmospheric balloons in the context of HAMLET project.
Paolo Rech, Alberto Bosio, Patrick Girard 0001, Serge Pravossoudovitch, Arnaud Virazel, Luigi Dilillo
Asian Test Symposium1
2010 Analysis of root causes of alpha sensitivity variations on microprocessors manufactured using different cell layouts
abstract
This paper reports and analyzes the results of alpha radiation testing campaigns on an embedded microprocessor manufactured with different standard cell libraries, each one enforcing Design for Manufacturing rules at a specific level. A set of analog simulations has been performed on flip-flops built with different physical layouts to reproduce and evaluate the effects of ionizing particles. The results of simulation experiments are presented and discussed, highlighting the configurations which are more likely to improve the system reliability, and then compared with radiation experiments data. Finally, we give a physical interpretation of the observed variations on radiation sensitivity.
Paolo Rech, Michelangelo Grosso, Fabio Melchiori, Domenico Loparco, Davide Appello, Luigi Dilillo, Alessandro Paccagnella, Matteo Sonza Reorda
IOLTS1
2010 A roaming memory test bench for detecting particle induced SEUs
abstract
In this paper, we propose a memory based test bench able to record soft errors that may occur to modern circuits in a certain environment. This system allows a good flexibility from different points of view. It is conceived to be modular, programmable, low power consuming and portable. Consequently, it can operate in various experimental conditions such as under artificial sources of particles as well as in natural ambience, from the earth surface to spatial environment.
Jean-Marc Gallière, Paolo Rech, Patrick Girard 0001, Luigi Dilillo
ITC2
2009 Evaluating Alpha-induced soft errors in embedded microprocessors
abstract
This paper presents the results of Alpha Single Event Upsets tests of an embedded 8051 microprocessor. Cross sections for the different memory resources (i.e., internal registers, code RAM, and user memory) are reported as well as the error rate for different codes implemented as test benchmarks. Test results are then discussed to find the contribution of each available resource to the overall device error rate.
Paolo Rech, Simone Gerardin, Alessandro Paccagnella, Paolo Bernardi 0002, Michelangelo Grosso, Matteo Sonza Reorda, Davide Appello
IOLTS1
2009 DfT Reuse for Low-Cost Radiation Testing of SoCs: A Case Study
abstract
This paper proposes an efficient low-cost strategy for collecting data during radiation experiments on systems-on-chips (SoCs), exploiting the available on-chip design for testability (DfT) structures devised for manufacturing test.The approach combines hardware test and diagnostic features with suitable software tools, which enable accurate measurements and quick transient effects data collection. Specific flows for radiation testing of different kinds of embedded cores are described. Results are shown for a radiation experiment conducted on an embedded SRAM core included in a 90 nm test-vehicle.
Davide Appello, Paolo Bernardi 0002, Simone Gerardin, Michelangelo Grosso, Alessandro Paccagnella, Paolo Rech, Matteo Sonza Reorda
VTS6