EDBT 2026 Demo / reviewers in the wild / expert
Robert Limas Sierra
dblp:321/0658
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-5206-3757ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Early-Stage Reliability Assessment of Tensor-Based Deep Learning Accelerators
Robert Limas Sierra, Alessandro Veronesi, Josie E. Rodriguez Condia, Letícia Maria Veiras Bolzani, Matteo Sonza Reorda |
IOLTS | 1 |
| 2026 | A Software-Based Fault Tolerance Mechanism for Matrix Multiplication Operations in Tensor Coresabstract1Modern Graphics Processing Units (GPUs) are increasingly employed to enhance the performance of algorithms across scientific and machine learning domains. Given the importance of General Matrix Multiplication (GEMM) operations, GPUs feature specialized in-chip accelerators, such as Tensor Cores (TCUs), to speed them up. High-Performance Computing (HPC) and safety-critical sectors (e.g., automotive, space, and autonomous robotics) impose severe constraints concerning not only energy consumption, performance, and area but also reliability. Faults arising from advanced semiconductor technologies or sustained HPC workloads can silently propagate, potentially leading to catastrophic failures.This work introduces a hardware-aware, software-based fault tolerance method to enhance the resilience of GEMM operations on TCUs. By leveraging TCU architecture and parallel operation distribution, the method enables efficient fault detection and mitigation. It utilizes redundant executions on TCU arithmetic cores (Dot-Product Units) to detect and, if required, correct fault effects. The method’s flexibility supports online detection and correction of transient and permanent hardware faults in TCU’s arithmetic units. Experimental results on real GPUs show the proposed mechanism introduces minimal and constant memory overhead and a negligible performance overhead (up to 1.13 times) across operand sizes. Thus, this solution offers an effective and complementary hardening strategy for TCU operations. Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda |
IEEE Trans. Computers | 1 |
| 2024 | Analyzing the Structural and Operational Impact of Faults in Floating-Point and Posit Arithmetic Cores for CNN Operationsabstract1This work reports a first attempt to evaluate the fine-grain impact of permanent faults in the structures of arithmetic hardware cores implementing two number formats (Posit and FP). We assess and analyze errors in the cores for two operations (Add, and Multiply), which are the most used in several modern applications, including machine learning. The results show that Posit cores are structurally more vulnerable to fault propagation and induce more output corruptions than FP cores (from 3.3% up to 6.2%). Moreover, we found that the average absolute error in faulty FP cores is higher by up to 2 orders of magnitude than in Posit ones. Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Robert Limas Sierra, Matteo Sonza Reorda |
ETS | 3 |
| 2024 | Effective Application-level Error Modeling of Permanent Faults on AI AcceleratorsabstractThe deployment of Machine Learning (ML) applications extensively leverages Matrix Multiplication (MM) operations on modern and advanced accelerators, like Graphic Processing Units (GPUs), which employ Tensor Core Units (TCUs) to optimize MM’s execution efficiently. However, reliability concerns arise in devices with cutting-edge semiconductor technologies (7 nm or less), as faults can compromise some structures (e.g., TCUs) during their operation. In safety-critical applications, this can lead to wrong DNN outcomes and cause unpredictable and unacceptable actions. Thus, the impact evaluation of such faults is crucial to ensure that TCUs and GPUs meet the safety standard requirements (e.g., ISO26262). Currently, the reliability assessment of complex applications concerning hardware faults involves fault injection (FI) campaigns. Unfortunately, low-level FI campaigns might be computationally prohibitive for GPUs when these execute massive applications like DNNs. In this work, we propose an error modeling approach to accurately describe corruptions from permanent faults on TCUs, during the operation of MMs. This approach enables realistic reliability evaluations of computationally expensive MM-based workloads, resulting in a huge acceleration (up to 225X) compared with hardware-level FIs. Our experimental results show a very good accuracy (up to $93 \%$ correlation between our error modeling approach and FI campaigns conducted on TCUs). Francesco Pessia, Juan-David Guerrero-Balaguera, Robert Limas Sierra, Josie E. Rodriguez Condia, Marco Levorato, Matteo Sonza Reorda |
IOLTS | 3 |
| 2024 | Reliability Assessment of Large DNN Models: Trading Off Performance and AccuracyabstractThe adoption of Deep Neural Networks (DNNs) in several domains allows for increased effectiveness in applications that deal with massive data-intensive and complex data inputs. When employed in safety-critical scenarios, such as automotive, aerospace, healthcare, and autonomous robotics, assessing the DNNs' reliability and functional safety is crucial to ensure their correct in-field operation, even in the presence of hardware faults. However, the system complexity and the massive amounts of data to be processed by DNNs prevent the effective adoption of traditional strategies for reliability characterization and for identifying the most fault-sensitive structures. Accurate fault assessment strategies usually require unacceptable computational power and large evaluation times. On the other hand, faster strategies commonly lack accuracy in correctly representing system faults. Consequently, it is necessary to develop effective strategies that trade-off between performance and accuracy. This work analyses three reliability assessment strategies for deep neural networks and their underlying hardware, highlighting the main solutions and challenges in terms of evaluation performance and fault characterization accuracy. We overview different solutions to evaluate the hardware accelerators implementing DNNs at three abstraction levels:$i$) by physically injecting faults on a GPU running DNNs, ii) by performing microarchitectural characterization of GPUs to develop application-accurate error models, and iii) by using structure-aware cross-layer error modeling on DNN hardware accelerators. Our experimental results indicate that accurate error representation requires structural features from the targeted hardware. Junchao Chen 0001, Giuseppe Esposito, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Angeliki Kritikakou, Milos Krstic, Robert Limas Sierra, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Marcello Traiola, Alessandro Veronesi |
VLSI-SoC | 7 |
| 2024 | Special Session: Reliability Assessment Recipes for DNN AcceleratorsabstractReliability assessment is mandatory to guarantee the correct behavior of Deep Neural Network (DNN) hardware accelerators in safety-critical applications. While fault injection stands out as a well-established, practical and robust method for reliability assessment, it is still a very time-consuming process. This paper contributes with three recipes for optimizing the efficiency of the reliability assessment: a) hybrid analytical and hierarchical FI-based reliability assessment for systolic-array-based DNN accelerators; b) mixing techniques for the reliability assessment of in-chip AI accelerators in GPUs; c) reliability assessment of DNN hardware accelerators through physical fault injection. The experimental results demonstrate the efficiency of the proposed methods applied to their target DNN HW accelerator platforms. Mohammad Hasan Ahmadilivani, Alberto Bosio, Bastien Deveautour, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Maksim Jenihhin, Angeliki Kritikakou, Robert Limas Sierra, Salvatore Pappalardo, Jaan Raik, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Mahdi Taheri, Marcello Traiola |
VTS | 8 |
| 2024 | Analyzing the Impact of Scheduling Policies on the Reliability of GPUs Running CNN OperationsabstractThe programming flexibility and parallelism of Graphics Processing Units (GPUs) contribute to their effective adoption in complex and data-intensive fields like Machine Learning, especially in the deployment of Convolutional Neural Networks (CNNs). CNNs are also used in some safety-critical applications with severe reliability constraints, such as autonomous driving and robotics. Modern GPUs efficiently combine hardware schedulers controllers and in-chip accelerators (e.g., Tensor Core Units, or TCUs) to enhance CNN’s performance. Interestingly, fine-grain reliability analyses combining the operation of task scheduling policies in GPUs and TCUs have remained unexplored. This work analyses the reliability impact of scheduling policies on GPUs when permanent faults affect TCUs, during the execution of CNN operations. We developed a configurable architectural GPU model (in terms of clusters and parallel cores) that implements five selectable scheduling policies and supports the instruction-accurate execution of TCUs. Our results indicate that the GPU’s architecture and the scheduling policy play a crucial role in the application’s corruption from faulty TCUs. From the experiments, we found that some policies can reduce the corruption effects by up to 22% for large GPUs. In addition, we evaluated the dynamic variability of the scheduling policies and their complexity on identifying deterministic effects on the application’s outputs. Robert Limas Sierra, Juan-David Guerrero-Balaguera, Francesco Pessia, Josie E. Rodriguez Condia, Matteo Sonza Reorda |
VTS | 1 |
| 2024 | Investigating and Reducing the Architectural Impact of Transient Faults in Special Function Units for GPUsabstractAbstract Ensuring the reliability of GPUs and their internal components is paramount, especially in safety-critical domains like autonomous machines and self-driving cars. These cutting-edge applications heavily rely on GPUs to implement complex algorithms due to their implicit programming flexibility and parallelism, which is crucial for efficient operation. However, as integration technologies advance, there is a growing concern regarding the potential increase in fault sensitivity of the internal components of current GPU generations. In particular, Special Function Unit (SFU) cores inside GPUs are used in multimedia, High-Performance Computing, and neural network training. Despite their frequent usage and critical role in several domains, reliability evaluations on SFUs and the development of effective mitigation solutions have yet to be studied and remain unexplored. This work evaluates the impact of transient faults in the main hardware structures of SFUs in GPUs. In addition, we analyze the main overhead costs and benefits of developing selective-hardening mechanisms for SFUs. We focus on evaluating and analyzing two SFU architectures for GPUs (’fused’ and ’modular’) and their relations to energy, area, and reliability impact on parallel applications. The experiments resort to fine-grain fault injection campaigns on an RTL GPU model (FlexGripPlus) instrumented with both SFUs. The results on both SFU architectures indicate that fused SFUs (in commercial-grade devices) require lower area overhead (about 27%) for their integration in GPUs but are more vulnerable to transient faults (in up to 47% for the analyzed cases) and less power efficient (in up to 36.6%) than modular SFUs. Moreover, the reliability estimation shows that Modular SFUs are structurally more resilient than Fused ones in up to one order of magnitude. Similarly, selective-hardening mechanism based on Triple-Modular Redundancy (TMR) shows that coarse-grain strategies might increase the reliability of the overall SFUs under feasible overhead costs. Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Edward Javier Patiño Nuñez, Robert Limas Sierra, Matteo Sonza Reorda |
J. Electron. Test. | 4 |
| 2023 | A Reliability-aware Environment for Design Exploration for GPU Devicesabstract1Nowadays, GPU platforms have gained wide importance in applications that require high processing power. Unfortunately, the advanced semiconductor technologies used for their manufacturing are prone to different types of faults. Hence, solutions are required to support the exploration of the resilience to faults of different architectures. Based on this motivation, this work presents an environment dedicated to the analysis of the impact of permanent faults on GPU platforms. This environment is based on GPGPU-Sim, with the objective of exploiting the configuration features of this tool and, thus, analyzing the effects of faults when changing the target architecture. To validate the environment and show its usability, a fault campaign has been carried out where three different GPU architectures (Kepler, Volta, and Turing) were used. In addition, each GPU has been modified with an arbitrary number of parallel processing cores (or SMs). Three representative applications (Vector Add, Scalar Product, and Matrix Multiply) were executed on each GPU, and the behavior of each architecture in the presence of permanent faults in the functional (i.e., integer unit and floating-point) units was analyzed. This fault campaign shows the usability of the environment and demonstrates its potential use to support decisions on the best architectural parameters for a given application. Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda |
DDECS | 1 |
| 2023 | Evaluating the Prevalence of SFUs in the Reliability of GPUsabstract1Currently, Graphics Processing Units (GPUs) are extensively used in several safety-critical domains to support the implementation of complex operations where reliability is a major concern. Some internal cores, such as Special Function Units (SFUs), are increasingly adopted, being crucial to achieving the necessary performance in multimedia, scientific computing, and neural network training. Unfortunately, these cores are highly unexplored in terms of their impact on reliability.In this work, we evaluate the incidence of SFUs on the reliability of GPUs when affected by soft errors. First, we analyze the impact of SFU cores on the GPU’s reliability and the running workloads. We resort to applications configured to use or not the SFU cores and evaluate the effect of soft errors by using a software-based fault injection environment (NVBITFI) in an NVIDIA Ampere GPU. Then, we focus on evaluating the impact of soft errors arising in the SFUs. A fine-grain RTL evaluation determines the soft error effects on two SFUs architectures for GPUs (’fused’ and ’modular’). The experiments use an open-source GPU (FlexGripPlus) instrumented with both SFU architectures. The results suggest that workloads using SFUs are more vulnerable to faults (from 1 up to 5 orders of magnitude for the analyzed applications). Moreover, the RTL results show that modular SFUs are less vulnerable to faults (in up to 47% for the analyzed workloads) in comparison with fused SFUs (base of commercial devices), so allowing us to identify the more robust SFU architecture. Josie E. Rodriguez Condia, Juan-David Guerrero-Balaguera, Edward Javier Patiño Nuñez, Robert Limas Sierra, Matteo Sonza Reorda |
ETS | 4 |
| 2023 | Analyzing the Impact of Different Real Number Formats on the Structural Reliability of TCUs in GPUsabstract1Modern Graphics Processing Units (GPUs) boost the execution of tiled matrix multiplications by extensively using in-chip accelerators (Tensor Core Units or TCUs). Unfortunately, cutting-edge semiconductor technologies are increasingly prone to fault defects. Indeed, faults may affect TCUs when processing massive amounts of data under classical floating-point formats, raising reliability concerns when used in the safety-critical and High-Performance Computing (HPC) domains. In this scenario, the characterization of faulty TCUs supporting different arithmetic formats is still missed. This work for the first time quantitatively evaluates the effects of hardware faults arising in TCU structures when using two different formats for real number representation (i.e., Floating-Point and Posit). For the experimental evaluation, we resort to an architectural description of a TCU core (PyOpenTCU) and perform 60 fault simulation campaigns, injecting 57,344 faults per campaign and requiring around 24 days of computation. The experimental results indicate a relation between the corrupted spatial areas in the output matrices and the TCU’s scheduling policies. Moreover, the numeric analysis shows that hardware faults in TCUs in most cases affect up to 2 bits in the output results for both considered formats. The results also demonstrate that the Posit formats are less affected by faults than Floating-Point formats by up to one order of magnitude. Robert Limas Sierra, Juan-David Guerrero-Balaguera, Josie E. Rodriguez Condia, Matteo Sonza Reorda |
VLSI-SoC | 1 |
| 2022 | Effective fault simulation of GPU's permanent faults for reliability estimation of CNNsabstractConvolutional Neural Networks (CNNs) and Graphic Processing Units (GPUs) are now increasingly adopted in many cutting edge safety-critical applications. Consequently, it is crucial to evaluate the reliability of these systems, since the hardware can be affected by several phenomena (e.g., wear out of the device), producing permanent defects in the GPU. These defects may induce wrong outcomes in the CNN that may endanger the application. Traditionally, the study of the effects of permanent faults on CNNs has been approached by resorting to application-level fault injection (e.g., acting on the weights). However, this approach has restricted scope, and it may not reveal the actual vulnerabilities in the GPU device. Hence, a more accurate evaluation of the fault effects is required, considering more in-depth details of the device’s hardware. This work introduces a more elaborated experimental evaluation of the impact of GPU’s permanent faults on the reliability of a CNN by resorting to a Software-Implemented Fault Injection(SWIFI) strategy, considering faults at the hardware level. The results of the fault simulation campaigns we performed on the GPU data-path cores are compared with those at the application level, proving that the latter ones are generally optimistic. Juan-David Guerrero-Balaguera, Robert Limas Sierra, Matteo Sonza Reorda |
IOLTS | 2 |
| 2022 | Evaluating the impact of Permanent Faults in a GPU running a Deep Neural NetworkabstractCurrently, Deep Neural Networks (DNNs) are fun-damental computational structures deployed in a wide range of modern application domains (e.g., data analysis, healthcare, automotive, robotics). The computational complexity is inherent in these cognitive models, which demand high-performance devices like Graphics Processing Units (GPUs). Therefore, the implementation of DNNs on GPU devices is becoming increasingly frequent, even for cutting-edge safety-critical applications (e.g., autonomous and semi-autonomous cars). Thus, the reliability evaluation of these applications is mandatory because several phenomena (including aging) may produce permanent defects in the GPU, thus inducing the DNN to produce wrong results. Until now, the effects of permanent faults on DNNs have been mainly investigated at the application level, only, e.g., acting on the parameters of the network. This paper presents an environment allowing for the first time a more detailed experimental evaluation of the impact of permanent faults in a GPU on the reliability of a DNN running on it, based on considering faults at the architectural level. The results of the fault injection campaigns we performed on the GPU register files are compared with those at the application level, proving that the latter ones are generally optimistic. Juan-David Guerrero-Balaguera, Luigi Galasso, Robert Limas Sierra, Ernesto Sánchez 0001, Matteo Sonza Reorda |
ITC-Asia | 3 |