Juan-David Guerrero-Balaguera

dblp:291/3936 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0001-6852-2372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 9 first-author · 28 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 From Prompts to Pressure: Evaluating LLM-driven Agents for GPU Stress-code Generation
Aurora Gensale, Giuseppe Esposito, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Luca Cagliero, Matteo Sonza Reorda
IOLTS3
2026 A Software-Based Fault Tolerance Mechanism for Matrix Multiplication Operations in Tensor Cores
abstract
1Modern Graphics Processing Units (GPUs) are increasingly employed to enhance the performance of algorithms across scientific and machine learning domains. Given the importance of General Matrix Multiplication (GEMM) operations, GPUs feature specialized in-chip accelerators, such as Tensor Cores (TCUs), to speed them up. High-Performance Computing (HPC) and safety-critical sectors (e.g., automotive, space, and autonomous robotics) impose severe constraints concerning not only energy consumption, performance, and area but also reliability. Faults arising from advanced semiconductor technologies or sustained HPC workloads can silently propagate, potentially leading to catastrophic failures.This work introduces a hardware-aware, software-based fault tolerance method to enhance the resilience of GEMM operations on TCUs. By leveraging TCU architecture and parallel operation distribution, the method enables efficient fault detection and mitigation. It utilizes redundant executions on TCU arithmetic cores (Dot-Product Units) to detect and, if required, correct fault effects. The method’s flexibility supports online detection and correction of transient and permanent hardware faults in TCU’s arithmetic units. Experimental results on real GPUs show the proposed mechanism introduces minimal and constant memory overhead and a negligible performance overhead (up to 1.13 times) across operand sizes. Thus, this solution offers an effective and complementary hardening strategy for TCU operations.
Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
IEEE Trans. Computers2
2025 Analysis and Mitigation of Soft-errors in GPU-accelerated Hyperspectral Image Classifiers
abstract
This work assesses the reliability of a hyperspectral image classifier for edge devices under transient faults by using a fine-grain strategy based on the Hardware Injection Through Program Transformation (HITPT) technique. The results identified the most vulnerable software parts and the corruption effects due to hardware faults (from 5.1% to 100.0% of accuracy drop). Then, the results supported the adoption of a selective-hardening software mechanism (based on the Duplication with Comparison strategy) to effectively mitigate the most critical effects under limited costs.
Sergiu-Mohamed Abed, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Gianluca De Lucia, Marco Lapegna, Matteo Sonza Reorda
DDECS2
2025 AI-Based Classification of Adversarial Attacks vs. Hardware Fault Corruptions in the Split Computing Context
abstract
Split Computing has emerged as a promising paradigm for deploying Deep Neural Networks in Edge and Inter-net of Things systems, enabling inference tasks to be distributed between resource-constrained edge devices and cloud servers. This approach is particularly attractive for autonomous systems, where security and reliability may be critical. However, interme-diate feature maps transmitted between devices are vulnerable to corruption, which may result from intentional adversarial attacks or unintentional hardware faults. Distinguishing whether corruption originates from an external adversary or an inherent system fault is crucial for implementing appropriate counter-measures-reinforcing security mechanisms against attacks or improving system reliability to mitigate the effects of hardware-related faults. To the best of our knowledge, this work is the first to propose a machine learning-based classification mechanism capable of differentiating adversarial attacks from hardware defects in Split Computing systems. The proposed approach analyzes the intermediate feature maps transmitted from the edge device to the server, classifying the source of corruption to guide appropriate responses. Experimental results demonstrate that one of the proposed classifiers can distinguish between intentional and unintentional feature map corruptions with an accuracy of 93.91 %.
Giuseppe Esposito, Enrico Magliano, Nicola Scarano, Tamer Eltaras, Juan-David Guerrero-Balaguera, Luca Mannella, Josie E. Rodriguez Condia, Annachiara Ruospo, Stefano Di Carlo, Marco Levorato, Alessandro Savino 0001, Matteo Sonza Reorda
IOLTS5
2025 Early Reliability Estimation in Hardware Accelerators using Improved Colored Petri Nets
abstract
This work exploits Colored-Petri-Nets (CPN) for the early reliability estimation of hardware accelerators, significantly reducing the complexity during early design stages aimed at safety-critical systems. Our method builds high-level models of complex hardware accelerators to estimate reliability, integrating circuit characterization and fine-grain fault simulations on fundamental structures. We evaluate our methodology using six architecture variants of an on-chip hardware accelerator for deep learning (GPUs’ Tensor Cores). The results demonstrate that our approach reduces evaluation costs by 118x, achieving accuracy levels of up to 93.5% compared to exhaustive RT-level fault injection campaigns, while enhancing engineering productivity for early-stage designs.
Ernesto Villegas Castillo, Felipe Augusto da Silva, Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Michael Glaß
ITC4
2025 Effective Fault Effects Evaluation for Permanent Faults in GPUs executing DNNs
abstract
Deep Neural Networks (DNNs) have permeated multiple applications, including cutting-edge safety-critical domains, which require relevant computational power, often provided by Graphic Processing Units (GPUs). GPUs are manufactured with advanced semiconductor technologies that can be affected by faults during the operational phase (e.g., due to wear-out, aging, or environmental harshness), whose effects possibly reach the DNN outputs, in some cases leading to catastrophic consequences. Hence, hardware-aware reliability assessments of DNNs are crucial to be considered in the context of safety-critical systems (following regulations/standards of specific application domains). Application-level fault injection (FI) techniques (i.e., DNN parameter corruption) are often adopted for the reliability evaluation of DNNs; unfortunately, these approaches hardly represent fault effects from GPU hardware. This work proposes an FI strategy based on Hardware-Injection-Through-Program-Transformation (HITPT) to mimic the effect of permanent faults (PFs) at the GPU instruction level, enabling effective assessment of PFs on DNN’s reliability. Our approach provides a good trade-off between the fault effect evaluation’s accuracy and the required computational time. Using the proposed approach, for the first time, we systematically assessed the effects of PF in GPUs executing some DNN sample cases. The results indicate that the faults injected closer to the hardware, using our evaluation strategy, can produce a higher accuracy degradation than the evaluations performed by the typical application-level FI that modify only the DNN parameters. Furthermore, the proposed FI methodology provides insightful results to identify the most suitable fault-tolerance solutions (e.g., selective hardening or design diversity) for their application at thread levels inside GPU’s kernels.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
ACM Trans. Design Autom. Electr. Syst.1
2024 Evaluating Different Fault Injection Abstractions on the Assessment of DNN SW Hardening Strategies
abstract
1The reliability of Neural Networks has gained significant attention, prompting efforts to develop SW-based hardening techniques for safety-critical scenarios. However, evaluating hardening techniques using application-level fault injection (FI) strategies, which are commonly hardware-agnostic, may yield misleading results. This study for the first time compares two FI approaches (at the application level (APP) and instruction level (ISA)) to evaluate deep neural network SW hardening strategies. Results show that injecting permanent faults at ISA (a more detailed abstraction level than APP) changes completely the ranking of SW hardening techniques, in terms of both reliability and accuracy. These results highlight the relevance of using an adequate analysis abstraction for evaluating such techniques.
Giuseppe Esposito, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
ATS2
2024 Analyzing the Structural and Operational Impact of Faults in Floating-Point and Posit Arithmetic Cores for CNN Operations
abstract
1This work reports a first attempt to evaluate the fine-grain impact of permanent faults in the structures of arithmetic hardware cores implementing two number formats (Posit and FP). We assess and analyze errors in the cores for two operations (Add, and Multiply), which are the most used in several modern applications, including machine learning. The results show that Posit cores are structurally more vulnerable to fault propagation and induce more output corruptions than FP cores (from 3.3% up to 6.2%). Moreover, we found that the average absolute error in faulty FP cores is higher by up to 2 orders of magnitude than in Posit ones.
Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Robert Limas Sierra, Matteo Sonza Reorda
ETS2
2024 Enhancing the Reliability of Split Computing Deep Neural Networks
abstract
Artificial intelligence is becoming increasingly popular for IoT applications in safety-critical fields (e.g., autonomous systems and biomedical, robots). Unfortunately, the inference’s workload process alone increases as the model size grows. To meet the computational power limitations of mobile devices running IoT applications, modern services sometimes resort to the Split Computing paradigm. Split Computing divides the inference process of a Neural Network into Head and Tail for their execution in a mobile device and a server, respectively, which also allows the reduction of the overall IoT device’s computational cost. Nonetheless, Split Computing can be used in safety-critical fields where reliability is crucial, especially when mobile devices have computational and cost restrictions. This paper introduces hardening techniques acting on the software to mitigate the effects of hardware faults on Split Computing models. The proposed hardening techniques consist of i) a bounded activation function whose thresholds are refined by training, and ii) a per-channel bounding of the bottleneck quantization of the split points. To quantitatively assess their effectiveness, we resorted to two different split configurations of a model for image classification. In addition, we considered a Split Computing model for object detection. Our findings indicate that the proposed approaches effectively reduces fault effects by $\mathbf{3. 5 \%}$ for image classifiers and $5.73 \%$ for object detectors when compared with other hardening approaches for general DNNs.
Giuseppe Esposito, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Marco Levorato, Matteo Sonza Reorda
IOLTS2
2024 Effective Application-level Error Modeling of Permanent Faults on AI Accelerators
abstract
The deployment of Machine Learning (ML) applications extensively leverages Matrix Multiplication (MM) operations on modern and advanced accelerators, like Graphic Processing Units (GPUs), which employ Tensor Core Units (TCUs) to optimize MM’s execution efficiently. However, reliability concerns arise in devices with cutting-edge semiconductor technologies (7 nm or less), as faults can compromise some structures (e.g., TCUs) during their operation. In safety-critical applications, this can lead to wrong DNN outcomes and cause unpredictable and unacceptable actions. Thus, the impact evaluation of such faults is crucial to ensure that TCUs and GPUs meet the safety standard requirements (e.g., ISO26262). Currently, the reliability assessment of complex applications concerning hardware faults involves fault injection (FI) campaigns. Unfortunately, low-level FI campaigns might be computationally prohibitive for GPUs when these execute massive applications like DNNs. In this work, we propose an error modeling approach to accurately describe corruptions from permanent faults on TCUs, during the operation of MMs. This approach enables realistic reliability evaluations of computationally expensive MM-based workloads, resulting in a huge acceleration (up to 225X) compared with hardware-level FIs. Our experimental results show a very good accuracy (up to $93 \%$ correlation between our error modeling approach and FI campaigns conducted on TCUs).
Francesco Pessia, Juan-David Guerrero-Balaguera, Robert Limas Sierra, Josie E. Rodriguez Condia, Marco Levorato, Matteo Sonza Reorda
IOLTS2
2024 Reliability Assessment of Large DNN Models: Trading Off Performance and Accuracy
abstract
The adoption of Deep Neural Networks (DNNs) in several domains allows for increased effectiveness in applications that deal with massive data-intensive and complex data inputs. When employed in safety-critical scenarios, such as automotive, aerospace, healthcare, and autonomous robotics, assessing the DNNs' reliability and functional safety is crucial to ensure their correct in-field operation, even in the presence of hardware faults. However, the system complexity and the massive amounts of data to be processed by DNNs prevent the effective adoption of traditional strategies for reliability characterization and for identifying the most fault-sensitive structures. Accurate fault assessment strategies usually require unacceptable computational power and large evaluation times. On the other hand, faster strategies commonly lack accuracy in correctly representing system faults. Consequently, it is necessary to develop effective strategies that trade-off between performance and accuracy. This work analyses three reliability assessment strategies for deep neural networks and their underlying hardware, highlighting the main solutions and challenges in terms of evaluation performance and fault characterization accuracy. We overview different solutions to evaluate the hardware accelerators implementing DNNs at three abstraction levels:$i$) by physically injecting faults on a GPU running DNNs, ii) by performing microarchitectural characterization of GPUs to develop application-accurate error models, and iii) by using structure-aware cross-layer error modeling on DNN hardware accelerators. Our experimental results indicate that accurate error representation requires structural features from the targeted hardware.
Junchao Chen 0001, Giuseppe Esposito, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Angeliki Kritikakou, Milos Krstic, Robert Limas Sierra, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Marcello Traiola, Alessandro Veronesi
VLSI-SoC4
2024 Special Session: Reliability Assessment Recipes for DNN Accelerators
abstract
Reliability assessment is mandatory to guarantee the correct behavior of Deep Neural Network (DNN) hardware accelerators in safety-critical applications. While fault injection stands out as a well-established, practical and robust method for reliability assessment, it is still a very time-consuming process. This paper contributes with three recipes for optimizing the efficiency of the reliability assessment: a) hybrid analytical and hierarchical FI-based reliability assessment for systolic-array-based DNN accelerators; b) mixing techniques for the reliability assessment of in-chip AI accelerators in GPUs; c) reliability assessment of DNN hardware accelerators through physical fault injection. The experimental results demonstrate the efficiency of the proposed methods applied to their target DNN HW accelerator platforms.
Mohammad Hasan Ahmadilivani, Alberto Bosio, Bastien Deveautour, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Maksim Jenihhin, Angeliki Kritikakou, Robert Limas Sierra, Salvatore Pappalardo, Jaan Raik, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Mahdi Taheri, Marcello Traiola
VTS5
2024 Evaluating the Reliability of Supervised Compression for Split Computing
abstract
Recent advances in Internet-of-things (IoT) and 5G infrastructures promote new computational paradigms such as Split Computing (SC) for deploying Deep Neural Networks (DNNs) on mobile applications. In SC, DNNs are partitioned into head and tail sub-models that are executed on the mobile device and cloud/edge servers, respectively. Modern SC models resort to head compression techniques to balance energy consumption, transmission data, and model size while preserving the outstanding accuracy of large state-of-the-art DNNs. These features make SC DNNs suitable for mobile applications, including safety-critical systems (e.g., self-driving vehicles, autonomous robots, and healthcare equipment), where reliability is a paramount factor mandated by strict safety standards. Despite there are many studies available about the reliability of DNNs, the SC models are still unexplored, especially when hardware faults threaten the operation of a mobile device. In this work, we present for the first time $i)$ an application-level fault injection strategy for modeling hardware faults on mobile GPUs executing SC DNNs and ii) an evaluation of the resilience of supervised compression methods utilized by SC systems. The preliminary results gathered on some representative benchmark networks and configurations show the feasibility and effectiveness of the approach. They also demonstrate that aggressive compression strategies lead to high accuracy degradation $(\approx$ 40%), increasing the overall vulnerability of the DNN and the system.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Marco Levorato, Matteo Sonza Reorda
VTS1
2024 Analyzing the Impact of Scheduling Policies on the Reliability of GPUs Running CNN Operations
abstract
The programming flexibility and parallelism of Graphics Processing Units (GPUs) contribute to their effective adoption in complex and data-intensive fields like Machine Learning, especially in the deployment of Convolutional Neural Networks (CNNs). CNNs are also used in some safety-critical applications with severe reliability constraints, such as autonomous driving and robotics. Modern GPUs efficiently combine hardware schedulers controllers and in-chip accelerators (e.g., Tensor Core Units, or TCUs) to enhance CNN’s performance. Interestingly, fine-grain reliability analyses combining the operation of task scheduling policies in GPUs and TCUs have remained unexplored. This work analyses the reliability impact of scheduling policies on GPUs when permanent faults affect TCUs, during the execution of CNN operations. We developed a configurable architectural GPU model (in terms of clusters and parallel cores) that implements five selectable scheduling policies and supports the instruction-accurate execution of TCUs. Our results indicate that the GPU’s architecture and the scheduling policy play a crucial role in the application’s corruption from faulty TCUs. From the experiments, we found that some policies can reduce the corruption effects by up to 22% for large GPUs. In addition, we evaluated the dynamic variability of the scheduling policies and their complexity on identifying deterministic effects on the application’s outputs.
Robert Limas Sierra, Juan-David Guerrero-Balaguera, Francesco Pessia, Josie E. Rodriguez Condia, Matteo Sonza Reorda
VTS2
2024 Investigating and Reducing the Architectural Impact of Transient Faults in Special Function Units for GPUs
abstract
Abstract Ensuring the reliability of GPUs and their internal components is paramount, especially in safety-critical domains like autonomous machines and self-driving cars. These cutting-edge applications heavily rely on GPUs to implement complex algorithms due to their implicit programming flexibility and parallelism, which is crucial for efficient operation. However, as integration technologies advance, there is a growing concern regarding the potential increase in fault sensitivity of the internal components of current GPU generations. In particular, Special Function Unit (SFU) cores inside GPUs are used in multimedia, High-Performance Computing, and neural network training. Despite their frequent usage and critical role in several domains, reliability evaluations on SFUs and the development of effective mitigation solutions have yet to be studied and remain unexplored. This work evaluates the impact of transient faults in the main hardware structures of SFUs in GPUs. In addition, we analyze the main overhead costs and benefits of developing selective-hardening mechanisms for SFUs. We focus on evaluating and analyzing two SFU architectures for GPUs (’fused’ and ’modular’) and their relations to energy, area, and reliability impact on parallel applications. The experiments resort to fine-grain fault injection campaigns on an RTL GPU model (FlexGripPlus) instrumented with both SFUs. The results on both SFU architectures indicate that fused SFUs (in commercial-grade devices) require lower area overhead (about 27%) for their integration in GPUs but are more vulnerable to transient faults (in up to 47% for the analyzed cases) and less power efficient (in up to 36.6%) than modular SFUs. Moreover, the reliability estimation shows that Modular SFUs are structurally more resilient than Fused ones in up to one order of magnitude. Similarly, selective-hardening mechanism based on Triple-Modular Redundancy (TMR) shows that coarse-grain strategies might increase the reliability of the overall SFUs under feasible overhead costs.
Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Edward Javier Patiño Nuñez, Robert Limas Sierra, Matteo Sonza Reorda
J. Electron. Test.2
2023 Assessing Convolutional Neural Networks Reliability through Statistical Fault Injections
abstract
Assessing the reliability of modern devices running CNN algorithms is a very difficult task. Actually, the complexity of the state-of-the-art devices makes exhaustive Fault Injection (FI) campaigns impractical and typically out of the computational capabilities. A possible solution consists of resorting to statistical FI campaigns that allow a reduction in the number of needed experiments by injecting only a carefully selected small part of it. Under specific hypothesis, statistical FIs guarantee an accurate picture of the problem, albeit selecting a reduced sample size. The main problems today are related to the choice of the sample size, the location of the faults, and the correct understanding of the statistical assumptions. The intent of this paper is twofold: first, we describe how to correctly specify statistical FIs for Convolutional Neural Networks; second, we propose a data analysis on the CNN parameters that drastically reduces the number of FIs needed to achieve statistically significant results without compromising the validity of the proposed method. The methodology is experimentally validated on two CNNs, ResNet-20 and MobileNetV2, and the results show that a statistical FI campaign on about 1.21% and 0.55% of the possible faults, provides very precise information of the CNN reliability. The statistical results have been confirmed by the exhaustive FI campaigns on the same cases of study.
Annachiara Ruospo, Gabriele Gavarini, Corrado De Sio, Juan-David Guerrero-Balaguera, Luca Sterpone, Matteo Sonza Reorda, Ernesto Sánchez 0001, Riccardo Mariani, Joseph Aribido, Jyotika Athavale
DATE4
2023 A Reliability-aware Environment for Design Exploration for GPU Devices
abstract
1Nowadays, GPU platforms have gained wide importance in applications that require high processing power. Unfortunately, the advanced semiconductor technologies used for their manufacturing are prone to different types of faults. Hence, solutions are required to support the exploration of the resilience to faults of different architectures. Based on this motivation, this work presents an environment dedicated to the analysis of the impact of permanent faults on GPU platforms. This environment is based on GPGPU-Sim, with the objective of exploiting the configuration features of this tool and, thus, analyzing the effects of faults when changing the target architecture. To validate the environment and show its usability, a fault campaign has been carried out where three different GPU architectures (Kepler, Volta, and Turing) were used. In addition, each GPU has been modified with an arbitrary number of parallel processing cores (or SMs). Three representative applications (Vector Add, Scalar Product, and Matrix Multiply) were executed on each GPU, and the behavior of each architecture in the presence of permanent faults in the functional (i.e., integer unit and floating-point) units was analyzed. This fault campaign shows the usability of the environment and demonstrates its potential use to support decisions on the best architectural parameters for a given application.
Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
DDECS2
2023 Evaluating the Prevalence of SFUs in the Reliability of GPUs
abstract
1Currently, Graphics Processing Units (GPUs) are extensively used in several safety-critical domains to support the implementation of complex operations where reliability is a major concern. Some internal cores, such as Special Function Units (SFUs), are increasingly adopted, being crucial to achieving the necessary performance in multimedia, scientific computing, and neural network training. Unfortunately, these cores are highly unexplored in terms of their impact on reliability.In this work, we evaluate the incidence of SFUs on the reliability of GPUs when affected by soft errors. First, we analyze the impact of SFU cores on the GPU’s reliability and the running workloads. We resort to applications configured to use or not the SFU cores and evaluate the effect of soft errors by using a software-based fault injection environment (NVBITFI) in an NVIDIA Ampere GPU. Then, we focus on evaluating the impact of soft errors arising in the SFUs. A fine-grain RTL evaluation determines the soft error effects on two SFUs architectures for GPUs (’fused’ and ’modular’). The experiments use an open-source GPU (FlexGripPlus) instrumented with both SFU architectures. The results suggest that workloads using SFUs are more vulnerable to faults (from 1 up to 5 orders of magnitude for the analyzed applications). Moreover, the RTL results show that modular SFUs are less vulnerable to faults (in up to 47% for the analyzed workloads) in comparison with fused SFUs (base of commercial devices), so allowing us to identify the more robust SFU architecture.
Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Edward Javier Patiño Nuñez, Robert Limas Sierra, Matteo Sonza Reorda
ETS2
2023 Understanding the Effects of Permanent Faults in GPU's Parallelism Management and Control Units
abstract
Modern Graphics Processing Units (GPUs) demand life expectancy extended to many years, exposing the hardware to aging (i.e., permanent faults arising after the end-of-manufacturing test). Hence, techniques to assess permanent fault impacts in GPUs are strongly required, especially in safety-critical domains.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Fernando Santos 0001, Matteo Sonza Reorda, Paolo Rech
SC1
2023 Analyzing the Impact of Different Real Number Formats on the Structural Reliability of TCUs in GPUs
abstract
1Modern Graphics Processing Units (GPUs) boost the execution of tiled matrix multiplications by extensively using in-chip accelerators (Tensor Core Units or TCUs). Unfortunately, cutting-edge semiconductor technologies are increasingly prone to fault defects. Indeed, faults may affect TCUs when processing massive amounts of data under classical floating-point formats, raising reliability concerns when used in the safety-critical and High-Performance Computing (HPC) domains. In this scenario, the characterization of faulty TCUs supporting different arithmetic formats is still missed. This work for the first time quantitatively evaluates the effects of hardware faults arising in TCU structures when using two different formats for real number representation (i.e., Floating-Point and Posit). For the experimental evaluation, we resort to an architectural description of a TCU core (PyOpenTCU) and perform 60 fault simulation campaigns, injecting 57,344 faults per campaign and requiring around 24 days of computation. The experimental results indicate a relation between the corrupted spatial areas in the output matrices and the TCU’s scheduling policies. Moreover, the numeric analysis shows that hardware faults in TCUs in most cases affect up to 2 bits in the output results for both considered formats. The results also demonstrate that the Posit formats are less affected by faults than Floating-Point formats by up to one order of magnitude.
Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
VLSI-SoC2
2022 A Compaction Method for STLs for GPU in-field test
abstract
Nowadays, Graphics Processing Units (GPUs) are effective platforms for implementing complex algorithms (e.g., for Artificial Intelligence) in different domains (e.g., automotive and robotics), where massive parallelism and high computational effort are required. In some domains, strict safety-critical requirements exist, mandating the adoption of mechanisms to detect faults during the operational phases of a device. An effective test solution is based on Self-Test Libraries (STLs) aiming at testing devices functionally. This solution is frequently adopted for CPUs, but can also be used with GPUs. Nevertheless, the in-field constraints restrict the size and duration of acceptable STLs. This work proposes a method to automatically compact the test programs of a given STL targeting GPUs. The proposed method combines a multi-level abstraction analysis resorting to logic simulation to extract the microarchitectural operations triggered by the test program and the information about the thread-level activity of each instruction and to fault simulation to know its ability to propagate faults to an observable point. The main advantage of the proposed method is that it requires a single fault simulation to perform the compaction. The effectiveness of the proposed approach was evaluated, resorting to several test programs developed for an open-source GPU model (FlexGripPlus) compatible with NVIDIA GPUs. The results show that the method can compact test programs by up to 98.64% in code size and by up to 98.42% in terms of duration, with minimum effects on the achieved fault coverage.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
DATE1
2022 Test, Reliability and Functional Safety Trends for Automotive System-on-Chip
abstract
This paper encompasses three contributions by industry professionals and university researchers. The contributions describe different trends in automotive products, including both manufacturing test and run-time reliability strategies. The subjects considered in this session deal with critical factors, from optimizing the final test before shipment to market to in-field reliability during operative life.
Francesco Angione, Davide Appello, Joseph Aribido, Jyotika Athavale, Nicolò Bellarmino, Paolo Bernardi 0002, Riccardo Cantoro, Corrado De Sio, Tommaso Foscale, Gabriele Gavarini, Juan-David Guerrero-Balaguera, Martin Huch, Giusy Iaria, Tobias Kilian, Riccardo Mariani, Raffaele Martone, Annachiara Ruospo, Ernesto Sánchez 0001, Ulf Schlichtmann, Giovanni Squillero, Matteo Sonza Reorda, Luca Sterpone, Vincenzo Tancorre, Roberto Ugioli
ETS11
2022 Effective fault simulation of GPU's permanent faults for reliability estimation of CNNs
abstract
Convolutional Neural Networks (CNNs) and Graphic Processing Units (GPUs) are now increasingly adopted in many cutting edge safety-critical applications. Consequently, it is crucial to evaluate the reliability of these systems, since the hardware can be affected by several phenomena (e.g., wear out of the device), producing permanent defects in the GPU. These defects may induce wrong outcomes in the CNN that may endanger the application. Traditionally, the study of the effects of permanent faults on CNNs has been approached by resorting to application-level fault injection (e.g., acting on the weights). However, this approach has restricted scope, and it may not reveal the actual vulnerabilities in the GPU device. Hence, a more accurate evaluation of the fault effects is required, considering more in-depth details of the device’s hardware. This work introduces a more elaborated experimental evaluation of the impact of GPU’s permanent faults on the reliability of a CNN by resorting to a Software-Implemented Fault Injection(SWIFI) strategy, considering faults at the hardware level. The results of the fault simulation campaigns we performed on the GPU data-path cores are compared with those at the application level, proving that the latter ones are generally optimistic.
Juan-David Guerrero-Balaguera, Robert Limas Sierra, Matteo Sonza Reorda
IOLTS1
2022 Evaluating the impact of Permanent Faults in a GPU running a Deep Neural Network
abstract
Currently, Deep Neural Networks (DNNs) are fun-damental computational structures deployed in a wide range of modern application domains (e.g., data analysis, healthcare, automotive, robotics). The computational complexity is inherent in these cognitive models, which demand high-performance devices like Graphics Processing Units (GPUs). Therefore, the implementation of DNNs on GPU devices is becoming increasingly frequent, even for cutting-edge safety-critical applications (e.g., autonomous and semi-autonomous cars). Thus, the reliability evaluation of these applications is mandatory because several phenomena (including aging) may produce permanent defects in the GPU, thus inducing the DNN to produce wrong results. Until now, the effects of permanent faults on DNNs have been mainly investigated at the application level, only, e.g., acting on the parameters of the network. This paper presents an environment allowing for the first time a more detailed experimental evaluation of the impact of permanent faults in a GPU on the reliability of a DNN running on it, based on considering faults at the architectural level. The results of the fault injection campaigns we performed on the GPU register files are compared with those at the application level, proving that the latter ones are generally optimistic.
Juan-David Guerrero-Balaguera, Luigi Galasso, Robert Limas Sierra, Ernesto Sánchez 0001, Matteo Sonza Reorda
ITC-Asia1
2022 A Multi-level Approach to Evaluate the Impact of GPU Permanent Faults on CNN's Reliability
abstract
Graphics processing units (GPUs) are widely used to accelerate Artificial Intelligence applications, such as those based on Convolutional Neural Networks (CNNs). Since in some domains in which CNNs are heavily employed (e.g., automotive and robotics) the expected lifetime of GPUs is over ten years, it is of paramount importance to study the impact of permanent faults (e.g. due to aging). Crucially, while the impact of transient faults on GPUs running CNNs has been widely studied, an accurate evaluation of the impact of permanent faults is still lacking. Performing this evaluation is challenging due to the complexity of GPU devices and the software implementing a CNN. In this work, we propose a methodology that combines the accuracy of gate-level fault simulation with the speed and flexibility of software fault injection to evaluate the effects of permanent hardware faults affecting a GPU. First, we profile the executed low-level GPU instructions during the CNN inference. Then, using extensive gate-level fault injection campaigns, we provide an accurate analysis of the effects of permanent faults on the internal modules executing the targeted instructions. Finally, we propagate these effects using fast software-based fault injection. The method allows, for the first time, to estimate the percentage of permanent faults leading the CNN to produce wrong results (i.e., changing the result of its work). The method's feasibility, which allows for flexibly trade-off accuracy with the required computational effort, is shown using LeNet running on an Ampere Nvidia GPU as a case study. The method reduces the computational effort for the evaluation by several orders of magnitude with respect to plain gate- and RTL-level faults simulation.
Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Fernando Santos 0001, Matteo Sonza Reorda, Paolo Rech
ITC2
2022 A New Method to Generate Software Test Libraries for In-Field GPU Testing Resorting to High-Level Languages
abstract
Self-Test Libraries (STLs) are widely used by companies for in-field fault detection in CPU devices. Their usage is now extending to GPUs, due to their increasing adoption in safety-critical applications. Using STLs provided by GPU manufacturers, system companies can effectively test these devices during their operative life, as required by functional safety standards. In the automotive domain, GPUs are often used to process a high amount of sensitive information in real-time (e.g., object recognition and path tracking). Thus, GPU devices in this field must guarantee functional safety features (e.g., ISO26262) by using one or more functional safety mechanisms. This paper presents a methodology to develop STLs resorting to High-Level Languages (HLLs) (e.g., CUDA), reducing the complexity of encoding at the assembly level. Moreover, we describe the main advantages and discuss the challenges and constraints when developing STLs with HLLs for GPUs. In particular, we describe those cases that demand the usage of a Low-Level Language (LLL). Additionally, we highlight a method to develop STLs resorting to HLLs, at least for some modules. The FlexGripPlus GPU model was employed to evaluate and validate the proposed strategies experimentally. The results show that STLs based on HLLs can be effectively developed for regular modules in the GPU.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
VTS1
2021 A Novel Compaction Approach for SBST Test Programs
abstract
In-field test of processor-based devices is a must when considering safety-critical systems (e.g., in robotics, aerospace, and automotive applications). During in-field testing, different solutions can be adopted, depending on the specific constraints of each scenario. In the last years, Self-Test Libraries (STLs) developed by IP or semiconductor companies became widely adopted. Given the strict constraints of in-field test, the size and time duration of a STL is a crucial parameter. This work introduces a novel approach to compress functional test programs belonging to an STL. The proposed approach is based on analyzing (via logic simulation) the interaction between the micro-architectural operation performed by each instruction and its capacity to propagate fault effects on any observable output, reducing the required fault simulations to only one. The proposed compaction strategy was validated by resorting to a RISC-V processor and several test programs stemming from diverse generation strategies. Results showed that the proposed compaction approach can reduce the length of test programs by up to 93.9% and their duration by up to 95%, with minimal effect on fault coverage.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
ATS1
2021 On the Functional Test of Special Function Units in GPUs
abstract
The Graphics Processing Units (GPUs) usage has extended from graphic applications to others where their high computational power is exploited (e.g., to implement Artificial Intelligence algorithms). These complex applications usually need highly intensive computations based on floating-point transcendental functions. GPUs may efficiently compute these functions in hardware using ad hoc Special Function Units (SFUs). However, a permanent fault in such units could be very critical (e.g., in safety-critical automotive applications). Thus, test methodologies for SFUs are strictly required to achieve the target reliability and safety levels. In this work, we present a functional test method based on a Software-Based Self-Test (SBST) approach targeting the SFUs in GPUs. This method exploits different approaches to build a test program and applies several optimization strategies to exploit the GPU parallelism to speed up the test procedure and reduce the required memory. The effectiveness of this methodology was proven by resorting to an open-source GPU model (FlexGripPlus) compatible with NVIDIA GPUs. The experimental results show that the proposed technique achieves 90.75% of fault coverage and up to 94.26% of Testable Fault Coverage, reducing the required memory and test duration with respect to pseudorandom strategies proposed by other authors.
Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda
DDECS1