Daniel Oliveira 0002

dblp:121/1932 · also Daniel A. G. de Oliveira, Daniel Alfonso Gonçalves de Oliveira · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
4since 2021 · last 2024
0000-0002-8752-3992ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 7 first-author · 3 since 2021Security and privacy · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2024 A Systematic Methodology to Compute the Quantum Vulnerability Factors for Quantum Circuits
abstract
Quantum computing is one of the most promising technology advances of the latest years. Qubits are highly sensitive to noise, which can make the output useless. Lately, it has been shown that superconducting qubits are extremely susceptible to external sources of faults, such as ionizing radiation. When adopted in large scale, radiation-induced errors are expected to become a serious challenge for qubits reliability. We propose an evaluation of the impact of transient faults in the execution of quantum circuits on superconducting chips. Inspired by the Architectural and Program Vulnerability Factors, widely used for classical computation, we propose the Quantum Vulnerability Factor (QVF) to measure the impact of qubit corruption on the circuit output. We model faults, and design a fault injector, based on the latest studies on real machines and radiation experiments. We report the finding of more than 388,000,000 fault injections, considering single and double faults, on three algorithms, identifying the faults and qubits that are more likely to impact the output. We give guidelines on how to map the qubits in real devices to reduce the output error and to reduce the probability of having a radiation-induced corruption modifying the output. Finally, we compare simulations with experiments on physical quantum computers.
Daniel Oliveira 0002, Edoardo Giusto, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech
IEEE Trans. Dependable Secur. Comput.1
2022 QuFI: a Quantum Fault Injector to Measure the Reliability of Qubits and Quantum Circuits
abstract
Quantum computing is an up-and-coming technology that is expected to revolutionize the computation paradigm in the next few years. Qubits, the primary computing elements of quantum circuits, exploit the quantum physics proprieties to increase the parallelism and speed of computation drastically. Unfortunately, besides being intrinsically noisy, qubits have also been shown to be highly susceptible to external sources of faults, such as ionizing radiation. The latest discoveries highlight a much higher radiation sensitivity of qubits than traditional transistors and identify a much more complex fault model than bit-flip.We propose a framework to identify the quantum circuits sensitivity to radiation-induced faults and the probability for a fault in a qubit to propagate to the output. Based on the latest studies and radiation experiments performed on real quantum machines, we model the transient faults in a qubit as a phase shift with a parametrized magnitude. Additionally, our framework can inject multiple qubit faults, tuning the phase shift magnitude based on the proximity of the qubit to the particle strike location. As we show in the paper, the proposed fault injector is highly flexible, and it can be used on both quantum circuit simulators and real quantum machines. We report the finding of more than 285, 249, 536 injections on the Qiskit simulator and 53, 248 injections on real IBM machines. We consider three quantum algorithms and identify the faults and qubits that are more likely to impact the output. We also consider the fault propagation dependence on the circuit scale, showing that the reliability profile for some quantum algorithms is scale-dependent, with increased impact from radiation-induced faults as we increase the number of qubits. Finally, we also consider multi qubits faults, showing that they are much more critical than single faults. The fault injector and the data presented in this paper are available in a public repository to allow further analysis.
Daniel Oliveira 0002, Edoardo Giusto, Emanuele Dri, Nadir Casciola, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech
DSN1
2022 Understanding the Impact of Cutting in Quantum Circuits Reliability to Transient Faults
abstract
Quantum Computing is a highly promising new computation paradigm. Unfortunately, quantum bits (qubits) are extremely fragile and their state can be gradually or suddenly modified by intrinsic noise or external perturbation. In this paper, we target the sensitivity of quantum circuits to radiation-induced transient faults. We consider quantum circuit cuts that split the circuit into smaller independent portions, and understand how faults propagate in each portion. As we show, the cuts have different vulnerabilities, and our methodology successfully identifies the circuit portion that is more likely to contribute to the overall circuit error rate. Our evaluation shows that a circuit cut can have a 4.6 x higher probability than the other cuts, when corrupted, to modify the circuit output. Our study, identifying the most critical cuts, moves towards the possibility of implementing a selective hardening for quantum circuits.
Nadir Casciola, Edoardo Giusto, Emanuele Dri, Daniel Oliveira 0002, Paolo Rech, Bartolomeo Montrucchio
IOLTS4
2021 Thermal neutrons: a possible threat for supercomputer reliability
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Andrea Favalli, Opale Schappert, Stephen Wender, Carlo Cazzaniga, Christopher Frost 0002, Paolo Rech
J. Supercomput.1
2020 Thermal Neutrons: a Possible Threat for Supercomputers and Safety Critical Applications
abstract
The high performance, high efficiency, and low cost of Commercial Off-The-Shelf (COTS) devices make them attractive for applications with strict reliability constraints. Today, COTS devices are adopted in HPC and safety-critical applications such as autonomous driving. Unfortunately, the cheap natural Boron widely used in COTS chip manufacturing process makes them highly susceptible to thermal (low energy) neutrons. In this paper, we demonstrate that thermal neutrons are a significant threat to COTS device reliability. For our study, we consider an AMD APU, three NVIDIA GPUs, an Intel accelerator, and an FPGA executing a relevant set of algorithms. We consider different scenarios that impact the thermal neutron flux such as weather, concrete walls and floors, and HPC liquid cooling systems. We show that thermal neutrons FIT rate could be comparable to the high energy neutron FIT rate.
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Carlo Cazzaniga, Christopher Frost 0002, Robert C. Baumann, Paolo Rech
ETS1
2019 SPADA: a statistical program attack detection analysis
abstract
One of the main challenges in system security is the detection of vulnerability exploitation, especially valid control flow exploitation. The specificity of state-of-the-art methods, such as signature-based detection, becomes a limiting factor when detecting the latest exploits and attacks uncovered.
Francis B. Moreira 0001, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux
CF2
2019 Identifying the Most Reliable Collaborative Workload Distribution in Heterogeneous Devices
abstract
The constant need of higher performances and reduced power consumption has lead vendors to design heterogeneous devices that embed traditional CPU and an accelerator, like a GPU or FPGA. When the CPU and the accelerator are used collaboratively the device computational performances reach their peak. However, the higher amount of resources employed for computation has, potentially, the side effect of increasing soft error rate. In this paper we evaluate the reliability behavior of AMD Kaveri Accelerated Processing Units executing a set of heterogeneous applications. We distribute the workload between the CPU and GPU and evaluate which configuration provides the lowest error rate or allows the computation of the highest amount of data before experiencing a failure. We show that, in most cases, the most reliable workload distribution is the one that delivers the highest performances. As experimentally proven, by choosing the correct workload distribution the device reliability can increase of up to 9x.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
DATE2
2019 Reliability Evaluation of Mixed-Precision Architectures
abstract
Novel computing architectures offer the possibility to execute float point operations with different precisions. The execution of reduced precision operations, when acceptable for certain applications, is likely to reduce both the execution time and the power consumption. However, the application's error rate and the device's reliability can also be impacted by these precision changes. In this paper, we study the impact of data and operation precision changes on the reliability of modern architectures. We consider Xilinx Field-Programmable Gate-Arrays (FPGA), Intel Xeon Phis, and NVIDIA Graphics Processing Units (GPUs) executing a set of codes implemented in double, single, and half-precision IEEE754-compliant float point data. On FPGAs, the reduced area and performance improvements brought by reduced precision operations increase reliability. On Xeon Phis the compiler biases significantly double and single-precision instructions execution. This raises the drawback of increasing single-precision error rates when compared to double-precision operations. NVIDIA GPUs make use of dedicated mixed-precision cores, which draw nontrivial effects on the device reliability. Generically speaking, on GPUs half-precision allows a higher number of executions to be correctly completed before experimenting a failure. Finally, we also evaluate how transient faults impact the output correctness. Our study shows that for most applications faults in a single or half-precision data or operation are more likely to significantly modify the output value than errors in double-precision data.
Fernando Santos 0001, Caio B. Lunardi, Daniel Oliveira 0002, Fabiano Libano, Paolo Rech
HPCA3
2019 Impact of Workload Distribution on Energy Consumption, Performance, and Reliability of Heterogeneous Devices
abstract
Devices integrating cores of different nature in the same chip achieve very high computation efficiency by reducing the power consumption and latencies of moving data from a chip to an external device. Heterogeneous devices are commonly used in high performance computing applications and, lately, their market has expanded from portable and gaming to safety-critical applications. In this paper, we evaluate how the collaborative workload distribution impact the energy consumption, performance, and reliability of heterogeneous devices. Then, we use the Energy-Delay-Fit Product (EDFP) to evaluate the trade off between the measured metrics and find how they correlate. To perform the proposed study we consider AMD Accelerated Processing Units (APUs) that embed a CPU and a GPU. We run four applications, each one representing an algorithm class, gradually distributing the workload from the CPU to the GPU and measuring both the energy consumption and execution time. Then, we take advantage of accelerated neutron beams to measure the realistic error rates of the different workload distributions. As we show in the paper, energy consumption and execution time are mold by the same trend while FIT rates highly depend on algorithm class and workload distribution. Additionally, we found that execution time is the most influencing factor for the device EDFP. The application EDFP varies of up tp 6 orders of magnitude depending on the workload distribution. An unwise configuration can then jeopardize the device efficiency and reliability.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
PDP2
2018 Code-Dependent and Architecture-Dependent Reliability Behaviors
abstract
The increased need for computing capabilities and higher efficiency have stimulated industries to make available in the market novel architectures with increased complexity. The variety of codes that need to be executed combined with the complexity of novel architectures introduces challenges in the reliability evaluation of computing systems and applications. This paper compares the reliability behaviors of six different architectures (an Intel co-processor, three NVIDIA GPUs, an AMD APU, an embedded ARM) executing eight different codes. To support our evaluation, we present and discuss experimental beam data that covers a total of more than 352,000 years of natural exposure and fault-injection analysis based on a total of more than 120,000 injections. We first quantify both the Silent Data Corruptions and the Detected Unrecoverable Errors rates. Then, we qualify observed errors considering the difference between the corrupted and expected values as well as the portion of the output that has been corrupted. From these analyses, we identify the reliability characteristics which are related to the underlying hardware and the intrinsic behaviors of the executed code. Finally, we discuss the implications of the device- and code-dependent reliability behaviors for approximate computing. We analyze the benefits, in term of reduced error rate, of a relaxed output correctness.
Vinicius Fratin, Daniel Oliveira 0002, Caio B. Lunardi, Fernando Santos 0001, Gennaro Severino Rodrigues, Paolo Rech
DSN2
2018 Predicting the Reliability Behavior of HPC Applications
abstract
The error rate of current High Performance Computing (HPC) systems is already in the order of one per dozens of hours. Understanding the reliability behavior of HPC applications will be required for the next generation of supercomputers. Using the reliability behavior one can select efficient mitigation techniques for the application and fine-tune parameters such as checkpoint frequency. In this paper, we investigate the application of a machine learning model to predict the reliability behavior of HPC applications. We inject faults in more than 30 HPC applications executing in the Intel Xeon Phi Knights Landing (KNL) and use profiling information to build a predictive model with Support Vector Machines (SVM). We show that the model can predict the Program Vulnerability Factor (PVF) with an average relative error of 7% for certain classes of algorithm, such as linear algebra and sorting. The average relative error for all algorithm classes is 22%. Such a fast and straightforward prediction model can be effective as a filter to select the most unreliable applications to perform an in-depth analysis.
Daniel Oliveira 0002, Francis B. Moreira 0001, Paolo Rech, Philippe Olivier Alexandre Navaux
SBAC-PAD1
2018 Special session: How approximate computing impacts verification, test and reliability
abstract
Two AxC techniques have been successfully applied to hardware components. The first one is the functional approximation [1]that modifies the circuit structure replacing the original function F with the function G. G implementation leads to area/energy reduction at the cost of reduced accuracy, meaning that some errors can be observed at the outputs of G. The observed errors are a variation between the output values of F (precise) and G (approximate). The variation is the accuracy loss measured by means of quality metric(s) [1]. The second AxC technique is the over-scaling based approximation. Basically, the HW component is forced to work outside its specified operating conditions [1]. The classical example is the reduction of the supply voltage under the minimum value.
Lukás Sekanina, Zdenek Vasícek, Alberto Bosio, Marcello Traiola, Paolo Rech, Daniel Oliveira 0002, Fernando Santos 0001, Stefano Di Carlo
VTS6
2017 Radiation-Induced Error Criticality in Modern HPC Parallel Accelerators
abstract
In this paper, we evaluate the error criticality of radiation-induced errors on modern High-Performance Computing (HPC) accelerators (Intel Xeon Phi and NVIDIA K40) through a dedicated set of metrics. We show that, as long as imprecise computing is concerned, the simple mismatch detection is not sufficient to evaluate and compare the radiation sensitivity of HPC devices and algorithms. Our analysis quantifies and qualifies radiation effects on applications' output correlating the number of corrupted elements with their spatial locality. Also, we provide the mean relative error (dataset-wise) to evaluate radiation-induced error magnitude. We apply the selected metrics to experimental results obtained in various radiation test campaigns for a total of more than 400 hours of beam time per device. The amount of data we gathered allows us to evaluate the error criticality of a representative set of algorithms from HPC suites. Additionally, based on the characteristics of the tested algorithms, we draw generic reliability conclusions for broader classes of codes. We show that arithmetic operations are less critical for the K40, while Xeon Phi is more reliable when executing particles interactions solved through Finite Difference Methods. Finally, iterative stencil operations seem the most reliable on both architectures.
Daniel Oliveira 0002, Laércio Lima Pilla, Mauricio Hanzich, Vinicius Fratin, Fernando Santos 0001, Caio B. Lunardi, José María Cela, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
HPCA1
2017 Experimental and analytical study of Xeon Phi reliability
abstract
We present an in-depth analysis of transient faults effects on HPC applications in Intel Xeon Phi processors based on radiation experiments and high-level fault injection. Besides measuring the realistic error rates of Xeon Phi, we quantify Silent Data Corruption (SDCs) by correlating the distribution of corrupted elements in the output to the application's characteristics. We evaluate the benefits of imprecise computing for reducing the programs' error rate. For example, for HotSpot a 0.5% tolerance in the output value reduces the error rate by 85%.
Daniel Oliveira 0002, Laércio Lima Pilla, Nathan DeBardeleben, Sean Blanchard, Heather M. Quinn, Israel Koren, Philippe Olivier Alexandre Navaux, Paolo Rech
SC1
2016 Evaluation and Mitigation of Radiation-Induced Soft Errors in Graphics Processing Units
abstract
Graphics processing units (GPUs) are increasingly attractive for both safety-critical and High-Performance Computing applications. GPU reliability is a primary concern for both the automotive and aerospace markets and is becoming an issue also for supercomputers. In fact, the high number of devices in large data centers makes the probability of having at least a device corrupted to be very high. In this paper, we aim at giving novel insights on GPU reliability by evaluating the neutron sensitivity of modern GPUs memory structures, highlighting pattern dependence and multiple errors occurrences. Additionally, a wide set of parallel codes are exposed to controlled neutron beams to measure GPUs operative error rates. From experimental data and algorithm analysis we derive general insights on parallel algorithms and programming approaches reliability. Finally, error-correcting code, algorithm-based fault tolerance, and duplication with comparison hardening strategies are presented and evaluated on GPUs through radiation experiments. We present and compare both the reliability improvement and imposed overhead of the selected hardening solutions.
Daniel Oliveira 0002, Laércio Lima Pilla, Thiago Santini, Paolo Rech
IEEE Trans. Computers1
2015 Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
abstract
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.
Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland
HPCA7
2014 Radiation Sensitivity of High Performance Computing Applications on Kepler-Based GPGPUs
abstract
In this paper we assess and discuss the radiation sensitivity of a set of HPC applications executed on NVIDIA K20 GPGPUs. The occurrence of both radiation-induced silent data corruption and functional interruption will be experimentally addressed for Hotspot, LavaMD, and Matrix Transponse. Each of the tested codes requires a proper computational power and elaborates a different amount of data. Both these characteristics play a significant role in the application radiations sensitivity. Additionally, an evaluation of the error rate at sea level will be provided for all the tested codes.
Daniel Oliveira 0002, Caio B. Lunardi, Laércio Lima Pilla, Paolo Rech, Philippe Olivier Alexandre Navaux, Luigi Carro
DSN1