EDBT 2026 Demo / reviewers in the wild / expert
Gabriel L. Nazar
dblp:18/9853 · also Gabriel Luca Nazar
· DBLP profile ↗
34ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0001-7202-7139ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anchor-and-Adapt: HLS QoR Prediction using Ground-Truth Seeding and Few-Shot Fine-TuningabstractGraph Neural Networks (GNNs) have emerged as powerful tools for guiding Design Space Exploration (DSE) in High-Level Synthesis (HLS), but they face a critical trade-off: models that base inference on pre-HLS inputs are often inaccurate, while accurate models based on post-HLS inputs are too slow for iterative exploration. This paper introduces a novel framework that targets this dilemma. Our approach centers on an anchor-based graph representation, where a single, ground-truth hardware implementation is used to seed a design graph with rich, post-implementation data. A heterogeneous GNN is then trained to predict the QoR delta caused by applying new optimization directives to this anchor, enabling rapid and high-fidelity estimation without re-running the HLS toolchain. Furthermore, we propose a dual-strategy framework that offers both a quick-setup Ensemble Model for general use and a specialized, high-accuracy Few-Shot Fine-Tuned Model for maximum per-kernel precision. Our results demonstrate that this methodology achieves significantly higher accuracy than pre-HLS methods, reducing the prediction error by 35-50% while maintaining a comparable exploration speed. More remarkably, our few-shot specialized model surpasses the prediction accuracy of post-HLS methods by more than 10% without incurring their restrictive per-point runtime cost.1 Gabriel C. Tavares, Heitor C. De Andrade, Fábio P. Itturriet, Gabriel L. Nazar |
DATE | 4 |
| 2023 | Modular VNF Components Acceleration With FPGA OverlaysabstractNetwork Functions Virtualization (NFV) is a novel paradigm that aims to minimize operational and capital expenditures, by decoupling network functions from dedicated hardware and implementing them as Virtualized Network Functions (VNFs) instead. However, to fulfill such expectations, VNFs must be implemented efficiently, offering high performance and energy efficiency, which is not always feasible on General-Purpose Processors (GPPs). Thus, the use of reconfigurable accelerators, typically based on Field-Programmable Gate Arrays (FPGAs), has been proposed to offer higher efficiency whilst not forsaking the flexibility that is the core of the NFV paradigm. Not all VNFs or even VNF Components (VNFCs), however, are suitable for FPGA acceleration. This leads to new challenges related to identifying those VNFCs that should be deployed in FPGAs, maximizing the reuse of developed FPGA accelerators, and managing this heterogeneous infrastructure. To address these challenges, in this paper we present an enhanced design of VNFAccel, a platform to manage VNFCs in heterogeneous NFV infrastructures. We evaluate the performance and energy efficiency of the implemented functions in comparison to GPP-based solutions, showing that, when properly used, FPGAs can provide relevant benefits while maintaining the flexibility and reuse potential envisioned for NFV. Filipe Bachini Lopes, Alberto E. Schaeffer Filho, Gabriel L. Nazar |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | Constraint-Aware Multi-Technique Approximate High-Level Synthesis for FPGAsabstractNumerous approximate computing (AC) techniques have been developed to reduce the design costs in error-resilient application domains, such as signal and multimedia processing, data mining, machine learning, and computer vision, to trade-off computation accuracy with area and power savings or performance improvements. Selecting adequate techniques for each application and optimization target is complex but crucial for high-quality results. In this context, Approximate High-Level Synthesis (AHLS) tools have been proposed to alleviate the burden of hand-crafting approximate circuits by automating the exploitation of AC techniques. However, such tools are typically tied to a specific approximation technique or a difficult-to-extend set of techniques whose exploitation is not fully automated or steered by optimization targets. Therefore, available AHLS tools overlook the benefits of expanding the design space by mixing diverse approximation techniques toward meeting specific design objectives with minimum error. In this work, we propose an AHLS design methodology for FPGAs that automatically identifies efficient combinations of multiple approximation techniques for different applications and design constraints. Compared to single-technique approaches, decreases of up to 30% in mean squared error and absolute increases of up to 6.5% in percentage accuracy were obtained for a set of image, video, signal processing and machine learning benchmarks. Marcos T. Leipnitz, Gabriel L. Nazar |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | SNAP: Selective NTV Heterogeneous Architectures for Power-Efficient Edge ComputingabstractWhile there is a growing need to process ML inference on the edge for improved latency and extra security, general-purpose solutions alone cannot cope with the increasing performance demand under power restrictions. Considering that systolic arrays are a prominent, but also power-hungry solution, we propose a methodology to enable their use in edge devices. For that, we propose SNAP, a selective Near-Threshold Voltage (NTV) strategy to explore heterogeneous MPSoCs with two voltage islands, one at NTV, and another at nominal voltage. By adopting a dynamic programming approach, SNAP may selectively apply NTV to the systolic array and to an optimal subset of cores in RISC- V-based MPSoCs, enabling ML acceleration on the edge. Combined with a smart application mapping, the strategy increases performance by up to 18.9 % over a nominal design within the same power limits. Rafael Billig Tonetto, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
DSD | 3 |
| 2021 | VNFAccel: An FPGA-based Platform for Modular VNF Components Acceleration
Filipe Bachini Lopes, Gabriel L. Nazar, Alberto E. Schaeffer Filho |
IM | 2 |
| 2020 | High-level synthesis of throughput-optimized and energy-efficient approximate designsabstractApproximate accelerators for throughput-demanding error-resilient kernels can be a solution to meet design requirements with acceptable deviation from the exact implementation. However, handcrafting approximate accelerators may impose prohibitive development time and cost overheads. In this scenario, approximate High-Level Synthesis (HLS) has been proposed to deal with the increased design complexity. Nevertheless, current tools are not suitable for exploring throughput optimizations, being instead constrained to perform specific improvements on area, power, and average performance. In this work, we propose the use of HLS to generate Pareto-optimal accelerators for applications facing throughput constraints. We present an approximate HLS tool able to improve the throughput of such accelerators by up to 80% with no additional area costs, while introducing manageable error for most applications. Marcos T. Leipnitz, Gabriel L. Nazar |
CF | 2 |
| 2020 | A Machine Learning Approach for Reliability-Aware Application Mapping for Heterogeneous MulticoresabstractWe propose a transparent and runtime methodology to increase the system's Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network that makes online predictions of the core's Architectural Vulnerability Factor (AVF), which allows for reliability-aware application-to-core mappings. We experiment with different configurations of RISC-V cores and compare the MWTF of prediction-based mappings against the optimal oracle, showing that our proposed model provides MWTF as close as 5.6% to the oracle. We also compare homogeneous and heterogeneous multicores, showing that heterogeneity provides room for increasing the MWTF in up to 19.4%. Rafael Billig Tonetto, Hiago Rocha, Gabriel L. Nazar, Antonio Carlos Schneider Beck |
DAC | 3 |
| 2020 | Throughput-Oriented Spatio-Temporal Optimization in Approximate High-Level SynthesisabstractCurrent and emerging systems for high-throughput applications, such as machine learning, cloud computing, and real-time video encoding demand real-time processing computations, heavily constrained by latency and power requirements. To deal with the increasing computational complexity, designers may resort to approximate accelerators for error-resilient compute-intensive kernels to meet such requirements with acceptable deviation from the exact implementation. However, since time-to-market is crucial when dealing with evolving applications, technologies, and standards, hand-crafting approximate accelerators may impose prohibitive development time and cost overheads. In this scenario, approximate High-Level Synthesis (HLS) methodologies have been proposed to deal with the complexity of exploring approximation techniques. Nevertheless, current tools are not suitable for exploring throughput optimizations, being instead constrained to perform specific improvements on area, power, and performance. In this work, we propose the use of HLS to generate Pareto-optimal accelerators for throughput-constrained applications. Particularly, we present a throughput-oriented approximate HLS methodology that explores both delay and area optimizations to increase the reuse over time and parallelism of such accelerators. Results show that our method is able to improve throughput by up to 80 % with no additional area costs or to sustain the same throughput of the exact design with about 45 % less area while introducing manageable error for most applications. Moreover, our method can attain throughput improvements of up to 18% when compared with recent works focusing only on performance or area optimizations, with no additional costs. Marcos T. Leipnitz, Gabriel L. Nazar |
ICCD | 2 |
| 2020 | Low-Power and Memory-Aware Approximate Hardware Architecture for Fractional Motion Estimation Interpolation on HEVCabstractNowadays, current video coding standards like the High Efficiency Video Coding (HEVC) implement several complex coding tools, like the Fractional Motion Estimation (FME). An alternative to improve performance and save power is allying the hardware acceleration with approximate computing solutions, focusing on such complex tools. In this work, we present a low-power and memory-aware hardware architecture for the HEVC FME interpolator, proposing the development of two novel hardware designs for the interpolation filters, called Approximate Unified FME Filters (AUFF). These solutions exploit the usage of approximate computing at both algorithmic and data levels, leading to a reduction in dissipated power and memory bandwidth. The proposed design is capable of real-time interpolation of UHD (Ultra High Definition) 4K and 8K videos when synthesized using a 40 nm standard-cell library, with a power dissipation ranging from 22.04 to 62.06 mW. Wagner Penny, Guilherme Corrêa 0001, Luciano Volcan Agostini, Daniel Palomino 0001, Marcelo Schiavon Porto, Gabriel L. Nazar, Bruno Zatt |
ISCAS | 6 |
| 2020 | A Reliability-Oriented Machine Learning Strategy for Heterogeneous Multicore Application MappingabstractWe propose a methodology to transparently estimate near-optimal application mappings aiming at increasing the Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network (ANN) capable of estimating the vulnerability factor of RISC-V cores at runtime, which allows for efficient and dynamic application-to-core mappings targeting better MWTF and MWTF/energy tradeoffs. Results show that our ANN-based mapping yields very close-to-optimal solutions, with a difference in MWTF of only 3% when compared to the optimal mapping. When compared to a homogeneous architecture composed of only big cores, heterogeneous architectures may provide improvement in MWTF of up to 20.5% while impacting 12.2% on performance. Rafael Billig Tonetto, Hiago Rocha, Bruno Zatt, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
ISCAS | 5 |
| 2019 | High-Level Synthesis of Resource-oriented Approximate Designs for FPGAsabstractWhen attempting to make a design fit a set of the heterogeneous resources found in Field-Programmable Gate Arrays (FPGAs), designers using High-Level Synthesis (HLS) may resort to approximate approaches. However, current FPGA-oriented approximate HLS tools do not allow specifying constraints on heterogeneous resources such as lookup tables, flip-flops, and multipliers, being instead error-oriented. In this work, we propose a resource-oriented HLS methodology with which designers can specify heterogeneous resource constraints and satisfy them while minimizing the output error, attaining average improvements, over error-oriented approaches, of about 34% and 2.2 dB for mean-squared error and peak signal-to-noise ratio error metrics, respectively. Marcos T. Leipnitz, Gabriel L. Nazar |
DAC | 2 |
| 2019 | Cost-effective Resilient FPGA-based LDPC Decoder ArchitectureabstractLow-Density Parity-Check (LDPC) codes have been used in many communication standards due to their capacity-approaching performance with feasible decoding architectures. Field-Programmable Gate Arrays (FPGAs) have been shown to be appropriate for the implementation of LDPC decoders, due to their ability to exploit the fine-grained parallelism found in such codes, as well as due to their reconfigurability, which allows to easily adapt the decoder to different codes. The susceptibility of FPGAs to faults affecting their configuration memories, however, demands specific fault tolerance strategies when these devices are used in harsh environments, such as aerospace applications, or even in ground-level critical systems. Thus, in this work we present a characterization of the behavior of LDPC decoders when subject to configuration errors and show that a single error can substantially degrade decoding performance, differently from what is observed in application-specific circuits. Based on this characterization, we propose a cost-effective fault tolerance scheme able to cope with faults in the FPGA fabric. Identifying the most critical components allowed reducing performance degradation by 89 % while only covering 55 % of their area. Eduardo Nunes de Souza, Gabriel L. Nazar |
IOLTS | 2 |
| 2019 | A Knapsack Methodology for Hardware-based DMR Protection against Soft Errors in Superscalar Out-of-Order ProcessorsabstractHigh-performance superscalar processors have been adopted to satisfy the rising demand for processing applications of ever-growing complexity. This extra complexity, added to the increasing vulnerability of transistors due to technology scaling, poses a great challenge since these effects have also been proven to affect ground-level safety-critical applications. To increase microarchitectural resilience, designers may adopt Dual Modular Redundancy (DMR), which offers full fault detection. However, given that DMR incurs in high area and energy overheads, we propose a design-time methodology aiming to achieve the best tradeoff between resilience and area overhead, decreasing DMR costs and maintaining acceptable detection levels for such a complex design. This is done by adopting the Knapsack Problem (KSP) as a heuristic to identify the optimal micro-architectural structures that should be duplicated to achieve target resilience with the smallest possible area overhead. By injecting over 800k faults in 12 significant micro-architectural structures of different versions of the complex Berkeley Out-of-Order Machine (BOOM) superscalar processor modeled with RTL accuracy, we compare this optimal strategy against a greedy one, showing that 90% of vulnerability reduction may be achieved with 50.6% and 107.8% area overheads for the optimal and greedy strategies, respectively. Rafael Billig Tonetto, Douglas Maciel Cardoso, Marcelo Brandalero, Luciano Volcan Agostini, Gabriel L. Nazar, José Rodrigo Azambuja, Antonio Carlos Schneider Beck |
VLSI-SoC | 5 |
| 2019 | High-Level Synthesis of Approximate Designs under Real-Time ConstraintsabstractThe adoption of High-Level Synthesis (HLS) has increased as the latest HLS tools have evolved to provide high-quality results while improving productivity and time-to-market. Concurrently, many works have been proposing the incorporation of approximate computing techniques within HLS toolchains, allowing automated generation of inexact circuits for error-tolerant application domains with the aim of trading-off computation accuracy with area/power savings or performance improvements. Thus, when attempting to make a design meet timing requirements, designers of real-time systems using HLS may resort to approximation approaches. However, current approximate HLS tools do not allow specifying real-time constraints, being instead error-constrained to explore area, power, or performance optimizations. In this work, we propose an approximate HLS framework for real-time systems that can be integrated with state-of-the-art HLS tools. With this framework designers can specify real-time constraints and satisfy them while minimizing the output error. It uses scheduling information and Worst-Case Execution Time (WCET) analysis for iteratively exploring time-error trade-offs of approximations in the time-critical execution path. Experimental results on signal and image processing benchmarks show that we can reduce the WCET of exact designs by up to 35% with acceptable quality degradation. Marcos T. Leipnitz, Gabriel L. Nazar |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Precise evaluation of the fault sensitivity of OoO superscalar processorsabstractSince superscalar processors lead the market, their resiliency evaluation by means of fault injection grows in importance. Fault injection strategies usually trade-off their levels of accuracy: low-level HW-based methods are accurate, but very expensive, need special equipment and the actual hardware, and lack controllability; while high-level simulation-based strategies are flexible, fast, easily accessible and have high controllability, but are not accurate since they are based on models that do not always reflect the low-level implementation, mainly when it comes to complex designs like out-of-order multiple-issue processors. In this work, we propose a cycle-accurate fault injection platform for superscalar processors, which has a smart checkpointing mechanism to accelerate injection time, attenuating the short-comings imposed by the aforementioned fault injection methods while providing the same level of abstraction as detailed RTL models. Leveraging from this new platform, we evaluate a complex and parameterizable Out-of-Order processor (BOOM) by experimenting with different issue widths and analyzing the sensitivity of several hardware structures of the processor. Rafael Billig Tonetto, Gabriel L. Nazar, Antonio Carlos Schneider Beck |
DATE | 2 |
| 2018 | Fault Tolerance Mechanisms for FPGA-Based Regular Expression Matching
Marcos T. Leipnitz, Gabriel L. Nazar |
J. Electron. Test. | 2 |
| 2018 | Repair of FPGA-Based Real-Time Systems With Variable SlacksabstractField-programmable gate arrays (FPGAs) based on SRAM cells are an attractive alternative for real-time system designers, as they offer high density, low cost, and high performance. The use of SRAM cells in the FPGA’s configuration memory, while enabling these desirable characteristics, also creates a reliability hazard as RAM cells are susceptible to single-event upsets (SEUs). The usual approach is the use of double or triple redundancy allied with a correction mechanism, such as periodic scrubbing. Although scrubbing is an effective technique to remove SEU-induced errors, the repair of real-time systems presents specific challenges, such as avoiding failures by missing real-time deadlines. In this article, a novel approach is proposed to use a deadline-aware scrubbing scheme with negligible area costs that dynamically chooses the scrubbing starting position. Such a scheme allows us to avoid missing real-time deadlines while maximizing the repair probability given a bounded repair time. Our approach reduces the failure rate, considering the probability of missing deadlines due to faults, by 33.39% on average, with an average area cost of 1.23%. Leonardo P. Santos, Gabriel L. Nazar, Luigi Carro |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | An energy-efficient memory hierarchy for multi-issue processorsabstractEmbedded processors must rely on the efficient use of instruction-level parallelism to answer the performance and energy needs of modern applications. However, a limiting factor to better use available resources inside the processor concerns memory bandwidth. Adding extra ports to allow for more data accesses drastically increases costs and energy. In this paper, we present a novel memory architecture system for embedded multi-issue processors that can overcome the limited memory bandwidth without adding extra ports to the system. We combine the use of software-managed memories (SMM) with the data cache to provide a system with a higher throughput without increasing the number of ports. Compiler-automated code transformations minimize the effort of programmers to benefit from the proposed architecture. Our experimental results show an average speedup of 1.17x, while consuming 69% less dynamic energy and on average 74.7% lower energy-delay product regarding data memory in comparison to a baseline processor. Tiago T. Jost, Gabriel L. Nazar, Luigi Carro |
DATE | 2 |
| 2016 | Scalable memory architecture for soft-core processorsabstractRestrictions over memory performance have always had a great impact on soft-core processors. The reduced number of ports on FPGAs' block RAMs may limit the exploitation of parallelism on soft-core processors that are implemented on top of these devices. Multiple memory ports on FPGAs are cumbersome and do not scale well, having a high cost in area and power consumption when implemented. In order to mitigate the impact of the memory bottleneck on such devices, we propose a scalable memory architecture for soft-cores. We make use of software-managed memories to build a memory system capable of improving performance and instruction-level parallelism (ILP) on soft-core processors. Results show that our architecture overcomes the limited parallelism realized on a dual-ported processor, reducing execution time by 16.5%. These improvements come with no area costs, as the processor is kept with the same total memory. Automated code transformations implemented within the LLVM compiler keep changes in application code to a minimum. We also show that our architecture scales better when boosting the number of functional units in the system. Tiago T. Jost, Gabriel L. Nazar, Luigi Carro |
ICCD | 2 |
| 2016 | Searching with a Corrupted HeuristicabstractMemory-based heuristics are a popular and effective class of admissible heuristic functions. However, corruptions to memory they use may cause these heuristics to become inadmissible. Corruption can be caused by the physical environment due to radiation and network errors, or it can be introduced voluntarily in order to decrease energy consumption. We introduce memory error correction schemes that do not require additional memory and exploit knowledge about the behavior of consistent heuristics. This is in contrast with error correcting code approaches which can limit the amount of corruption but at the cost of additional energy and memory consumption. Search algorithms using our methods are guaranteed to find a solution if one exists and its suboptimality is bounded. Moreover, our methods are resilient to any number of memory errors that may occur. An experimental evaluation is also provided to demonstrate the applicability of our approach. Levi Lelis, Richard Anthony Valenzano, Gabriel L. Nazar, Roni Stern |
SOCS | 3 |
| 2016 | Live-Out Register Fencing: Interrupt-Triggered Soft Error Correction Based on the Elimination of Register-to-Register CommunicationabstractThis article introduces Live-Out Register Fencing (LoRF), a soft error correction mechanism that uses the novel Spill Register File as a container of checkpointing data. LoRF’s Spill Register File holds the values shared among basic blocks in the program, and, coupled with a new compilation strategy, LoRF allows for error correction in the same basic block where the error was detected. In LoRF, error correction is triggered by a hardware interrupt that restores the registers of a basic block from the Spill Register File. After these registers are restored, the basic block where the error was detected can just be re-executed, thus reducing the costs of error recovery. LoRF’s error correction policy eliminates the need for expensive architectural support for checkpointing and rollback, reducing the performance overhead of online soft error correction. LoRF relies on both a modified processor architecture and a corresponding compiler. The architecture was implemented in synthesizable VHDL, whereas the compiler was developed as an extension of the LLVM framework. Fault injection experiments support an error correction coverage of 99.35% and a mean performance overhead of 1.33 for the entire life cycle of an error from its occurrence to its elimination from the system. Ronaldo Rodrigues Ferreira, Gabriel L. Nazar, Jean da Rolt, Álvaro F. Moreira, Luigi Carro |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2016 | Beyond Cross-Section: Spatio-Temporal Reliability AnalysisabstractA computational system employed in safety-critical applications typically has reliability as a primary concern. Thus, the designer focuses on minimizing the device radiation-sensitive area, often leading to performance degradation. In this article, we present a mathematical model to evaluate system reliability in spatial (i.e., radiation-sensitive area) and temporal (i.e., performance) terms and prove that minimizing radiation-sensitive area does not necessarily maximize application reliability. To support our claim, we present an empirical counterexample where application reliability is improved even if the radiation-sensitive area of the device is increased. An extensive radiation test campaign using a 28 nm commercial-off-the-shelf ARM-based SoC was conducted, and experimental results demonstrate that, while executing the considered application at military aircraft altitude, the probability of executing a two-year mission workload without failures is increased by 5.85% if L1 caches are enabled (thus increasing the radiation-sensitive area) when compared to no cache level being enabled. However, if both L1 and L2 caches are enabled, the probability is decreased by 31.59%. Thiago Santini, Paolo Rech, Gabriel L. Nazar, Flávio Rech Wagner |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | Fine-Grained Fast Field-Programmable Gate Array ScrubbingabstractField-programmable gate arrays provide several relevant advantages for critical systems, such as flexibility and high performance. However, their use in critical systems requires efficient means to mitigate transient faults in the configuration bits. This paper focuses on an alternative mechanism to reduce the repair time of traditional scrubbing approaches. It relies on fine-grained error detection and partial reconfiguration. The fine-grained information is used to dynamically choose an optimized starting position for the scrubbing procedure, reducing the mean repair time. We explore the design space provided by the technique and propose an approach to make resilient diagnosis of configuration faults. The efficiency, scalability, and robustness of the proposed mechanisms are evaluated. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Reliable execution of statechart-generated correct embedded software under soft errorsabstractThis paper proposes a design methodology for fault-tolerant embedded systems development that starts from software specification and goes down to hardware execution. The proposed design methodology uses formally verified and correct-by-construction software created from high-level UML statechart models for software specification and implementation. On the hardware reliability side, this paper uses the MoMa architecture for reliable embedded computing which we deploy as a soft-core onto an off-the-shelf FPGA. MoMa introduces architectural innovations that support the semantics of the UML statechart execution in a reliable fashion. The proposed design methodology is evaluated with a real automotive case study based on an exhaustive FPGA-implemented fault injection campaign. Ronaldo Rodrigues Ferreira, Thomas Klotz, Thilo Vörtler, Jean da Rolt, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro, Karsten Einwich |
DDECS | 5 |
| 2014 | Adaptive Low-Power Architecture for High-Performance and Reliable Embedded ComputingabstractThis paper presents the Matrix Operation Microprocessor Architecture (MoMa) for reliable embedded computing. MoMa introduces a software execution mechanism based on transactions, which provides a localized error correction scheme that leads to reduced error correction latency and hardware redundancy without incurring on expensive execution check pointing. Coupled to the transactional software execution is a dedicated adaptive core for matrix multiplication which is protected with a hardware implementation of the Algorithm-Based Fault Tolerance technique. MoMa drives the matrix core in an adaptive fashion based on dynamically turning it on only when high-performance computation is necessary, leading to ultimate power savings and error coverage. We performed an exhaustive FPGA-implemented fault injection campaign, in which we observed an error detection coverage of almost 100% and an error correction coverage of almost 98% on average. MoMa is also evaluated in terms of power, area, and performance, showing its competitiveness against a classical TMR solution. Ronaldo Rodrigues Ferreira, Jean da Rolt, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro |
DSN | 3 |
| 2014 | Reducing embedded software radiation-induced failures through cache memoriesabstractCache memories are traditionally disabled in space-level and safety-critical applications, since it was believed that the sensitive area they introduce would compromise the system reliability. As technology has evolved, the speed gap between logic and main memory has increased in such a way that disabling caches slows the code much more than in the past. As a result, the processor is exposed for a much longer time in order to compute the same workload. In this paper we demonstrate that, on modern embedded processors, enabling caches may bring benefits to critical systems: the larger exposed area may be compensated by the shorter exposure time, leading to an overall improved reliability. We describe the Mean Workload Between Failures, an intuitive metric to evaluate the impact of enabling caches for a given generic application error rate. The proposed metric is experimentally validated through an extensive radiation test campaign using a 28 nm off-the-shelf ARM-based SoC as a case study. The failure probability of the bare-metal application is decreased when the L1 cache is enabled but increased when L2 is also enabled. We also discuss when L2 caches could make the device more reliable. Thiago Santini, Paolo Rech, Gabriel L. Nazar, Luigi Carro, Flávio Rech Wagner |
ETS | 3 |
| 2014 | Power dissipation effects on 28nm FPGA-based System on Chips neutron sensitivityabstractModern System on Chips (SoCs) and embedded electronic devices work at very high frequencies, which have the countermeasure of increasing the power dissipation and, consequently, the silicon die temperature. The presented radiation experiments on a 28nm FPGA-based SoC demonstrate that the temperature variation caused by a higher operating frequency affects the FPGA configuration memory cross section. An evaluation and discussion of the observed reliability dependence on power dissipation effects on practical application is also presented. Giovanni Bruni, Paolo Rech, Lucas A. Tambara, Gabriel L. Nazar, Fernanda Lima Kastensmidt, Ricardo Augusto da Luz Reis, Alessandro Paccagnella |
VLSI-SoC | 4 |
| 2014 | Adaptive Parallelism Exploitation under Physical and Real-Time Constraints for Resilient SystemsabstractThis article introduces the resilient adaptive algebraic architecture that aims at adapting parallelism exploitation of a matrix multiplication algorithm in a time-deterministic fashion to reduce power consumption while meeting real-time deadlines present in most DSP-like applications. The proposed architecture provides low-overhead error correction capabilities relying on the hardware implementation of the algorithm-based fault-tolerance method that is executed concurrently with matrix multiplication, providing efficient occupation of memory and power resources. The Resilient Adaptive Algebraic Architecture (RA 3 ) is evaluated using three real-time industrial case studies from the telecom and multimedia application domains to present the design space exploration and the adaptation possibilities the architecture offers to hardware designers. RA 3 is compared in its performance and energy efficiency with standard high-performance architectures, namely a GPU and an out-of-order general-purpose processor. Finally, we present the results of fault injection campaigns in order to measure the architecture resilience to soft errors. Fábio P. Itturriet, Gabriel L. Nazar, Ronaldo Rodrigues Ferreira, Álvaro F. Moreira, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2013 | Scrubbing unit repositioning for fast error repair in FPGAsabstractField Programmable Gate Arrays (FPGAs) are very successful platforms that rely on large configuration memories to store the circuit functions required by users. Faults affecting such memories are a major dependability threat for these devices, and the applicability of FPGAs on critical systems depends on efficient means to mitigate their effects. The main means to effectively remove such faults, namely configuration scrubbing, consists in rewriting the desired contents of this memory and suffers from high power consumption and a long mean time to repair (MTTR). In this work we propose Scrubbing Unit Repositioning for Fast Error Repair (SURFER), a novel approach to exploit partial dynamic reconfiguration coupled with fine-grained redundancy to greatly reduce the MTTR for FPGAs subject to upsets in their configuration memories. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
CASES | 1 |
| 2013 | Accelerated FPGA repair through shifted scrubbingabstractAs critical systems make more and more use of high performance FPGAs, several reliability aspects of these devices come into play. Whenever SRAM-based FPGAs are used, upsets in the configuration memory become a major dependability threat, and must be removed as soon as possible. This is usually accomplished through a process called scrubbing. The traditional scrubbing technique, however, suffers from high energy costs and a long mean time to repair (MTTR). In this work we propose a novel approach to minimize these drawbacks through a triggered shifted scrubbing procedure. The proposed technique exploits the non-uniform distribution of critical bits in the configuration memory of the device to reduce the repair time. It provides an average MTTR reduction of 30% without any changes in the circuit implemented in the FPGA when compared to previous works. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
FPL | 1 |
| 2012 | Resilient Adaptive Algebraic Architecture for Parallel Detection and Correction of Soft-ErrorsabstractA novel fault-tolerant microprocessor capable of detecting and correcting radiation-induced soft errors is proposed and evaluated. The Resilient Adaptive Algebraic Architecture performs time redundancy in parallel with matrix multiplication computation, guaranteeing on-the-fly detection and correction of errors disrupting data and logic with minimum overhead. We evaluate the RA3microprocessor in terms of performance, area, energy consumption, and fault coverage by performing an extensive design space exploration of the architecture. Finally, we also discuss how the proposed architecture can be used to support a novel hardened-by-construction HW/SW stack based on what we call single-program execution. Fábio P. Itturriet, Ronaldo Rodrigues Ferreira, Gustavo Girão, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro |
DSD | 4 |
| 2012 | Fast error detection through efficient use of hardwired resources in FPGAsabstractProviding high reliability for FPGAs is a demanding task, as such devices may be subject to faults in the configuration bitstream, altering the specified function. Traditional modular redundancy remains the most used technique, due to its high fault coverage and low performance overhead. When high availability and strict real-time deadlines must be considered, however, a short mean time to repair also becomes crucial. The use of fine-grained modules can accelerate error detection, fault diagnosis and bitstream correction, but with increased area costs. In this work, we propose the use of hardwired resources found in state-of-the-art FPGAs to provide fast and area efficient fine-grained error detection. Experimental results show an average speed up in error detection of 7.68 times with only 3.2% more area overhead, when compared to coarse-grained modular redundancy. Gabriel L. Nazar, Luigi Carro |
ETS | 1 |
| 2012 | Exploiting Modified Placement and Hardwired Resources to Provide High Reliability in FPGAsabstractPossible scenarios for future manufacturing technologies increase the desirable features of fault tolerance techniques, such as coping with multiple faults and reducing error latency. On the other hand, current high-end FPGAs present, besides lookup tables and flip-flops, several dedicated components that perform the most commonly required functions. In this paper, we propose an approach to use such resources to efficiently provide fault detection capabilities. We further extend the technique with placement constraints to enhance the detection of faults affecting the routing resources, which is a critical demand for such devices. Gabriel L. Nazar, Luigi Carro |
FCCM | 1 |
| 2011 | Energy efficient pseudo-cache architecture through fine-grained reconfigurabilityabstractFueled by the exponential growth in transistors available to processor designers, cache memories became a very significant percentage of the overall area, power dissipation and energy consumption of modern systems. Instruction cache memories, however, typically hold highly redundant information in each of their columns, due to the repeated use of instructions and registers by compilers. Current memory architectures do not exploit this fact to reduce energy, consuming constant amounts of power regardless of switching activity. This work proposes the use of a fine-grained reconfigurable architecture to exploit this redundancy, providing an energy efficient on-chip storage element for embedded processors. The proposed architecture reached consumes up to 86% less energy, with an average reduction of 39%. Gabriel L. Nazar, Luigi Carro |
ISCAS | 1 |