Marcelo Brandalero

dblp:160/4598 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
6since 2021 · last 2022
0000-0002-0012-7023ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2022 G-GPU: A Fully-Automated Generator of GPU-like ASIC Accelerators
abstract
Modern Systems on Chip (SoC), almost as a rule, require accelerators for achieving energy efficiency and high performance for specific tasks that are not necessarily well suited for execution in standard processing units. Considering the broad range of applications and necessity for specialization, the design of SoCs has thus become expressively more challenging. In this paper, we put forward the concept of G-GPU, a general-purpose GPU-like accelerator that is not application-specific but still gives benefits in energy efficiency and throughput. Furthermore, we have identified an existing gap for these accelerators in ASIC, for which no known automated generation platform/tool exists. Our solution, called GPUPlanner, is an open-source generator of accelerators, from RTL to GDSII, that addresses this gap. Our analysis results show that our automatically generated G-GPU designs are remarkably efficient when compared against the popular CPU architecture RISC- V, presenting speed-ups of up to 223 times in raw performance and up to 11 times when the metric is performance derated by area. These results are achieved by executing a design space exploration of the GPU-like accelerators, where the memory hierarchy is broken in a smart fashion and the logic is pipelined on demand. Finally, tapeout-ready layouts of the G-GPU in 65nm CMOS are presented.
Tiago D. Perez, Marcio Gonçalves, Leonardo Gobatto, Marcelo Brandalero, José Rodrigo Azambuja, Samuel Nascimento Pagliarini
DATE4
2022 STAP: An Architecture and Design Tool for Automata Processing on Memristor TCAMs
abstract
Accelerating finite-state automata benefits several emerging application domains that are built on pattern matching. In-memory architectures, such as the Automata Processor (AP), are efficient to speed them up, at least for outperforming traditional von-Neumann architectures. In spite of the AP’s massive parallelism, current APs suffer from poor memory density, inefficient routing architectures, and limited capabilities. Although these limitations can be lessened by emerging memory technologies, its architecture is still the major source of huge communication demands and lack of scalability. To address these issues, we present STAP , a Scalable TCAM-based architecture for Automata Processing . STAP adopts a reconfigurable array of processing elements, which are based on memristive Ternary CAMs (TCAMs), to efficiently implement Non-deterministic finite automata (NFAs) through proper encoding and mapping methods. The CAD tool for STAP integrates the design flow of automata applications, a specific mapping algorithm, and place and route tools for connecting processing elements by RRAM-based programmable interconnects. Results showed 1.47× higher throughput when processing 16-bit input symbols, and improvements of 3.9× and 25× on state and routing densities over the state-of-the-art AP, while preserving 10 4 programming cycles.
João Paulo C. de Lima, Marcelo Brandalero, Michael Hübner 0001, Luigi Carro
ACM J. Emerg. Technol. Comput. Syst.2
2022 Reduced Precision DWC: An Efficient Hardening Strategy for Mixed-Precision Architectures
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing devices. However, it introduces performance and energy consumption overheads that could be unsuitable for high-performance computing or real-time safety-critical applications. In this article, we present Reduced-Precision Duplication with Comparison (RP-DWC) as a means to lower the overhead of DWC by executing the redundant copy in reduced precision. RP-DWC is particularly suitable for modern mixed-precision architectures, such as NVIDIA GPUs, that feature dedicated functional units for computing with programmable accuracy. We discuss the benefits and challenges associated with RP-DWC and show that the intrinsic difference between the mixed-precision copies allows for detecting most, but not all, errors. However, as the undetected faults are the ones that fall into the difference between precisions, they are the ones that produce a much smaller impact on the application output and, thus, might be tolerated. We investigate RP-DWC impact into fault detection, performance, and energy consumption on Volta GPUs. Through fault injection and beam experiment, using three microbenchmarks and four real applications, we show that RP-DWC achieves an excellent coverage (up to 86 percent) with minimal overheads (as low as 0.1 percent time and 24 percent energy consumption overhead).
Fernando Santos 0001, Marcelo Brandalero, Michael B. Sullivan 0001, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IEEE Trans. Computers2
2021 Artificial Intelligence for Mass Spectrometry and Nuclear Magnetic Resonance Spectroscopy
abstract
Mass Spectrometry (MS) and Nuclear Magnetic Resonance Spectroscopy (NMR) are critical components of every industrial chemical process as they provide information on the concentrations of individual compounds and by-products. These processes are carried out manually and by a specialist, which takes a substantial amount of time and prevents their utilization for real-time closed-loop process control. This paper presents recent advances from two projects that use Artificial Neural Networks (ANNs) to address the challenges of automation and performance-efficient realizations of MS and NMR. In the first part, a complete toolchain has been developed to develop simulated spectra and train ANNs to identify compounds in MS. In the second part, a limited number of experimental NMR spectra have been augmented by simulated spectra to train an ANN with better prediction performance and speed than state-of-the-art analysis. These results suggest that, in the context of the digital transformation of the process industry, we are now on the threshold of a possible strongly simplified use of MS and MRS and the accompanying data evaluation by machine-supported procedures, and can utilize both methods much wider for reaction and process monitoring or quality control.
Florian Fricke, Safdar Mahmood, Javier Hoffmann, Marcelo Brandalero, Sascha Liehr, Simon Kern, Klas Meyer, Stefan Kowarik, Stephan Westerdick, Michael Maiwald, Michael Hübner 0001
DATE4
2021 AITIA: Embedded AI Techniques for Industrial Applications
abstract
Motivated by an increasing interest from startups in embedded Artificial Intelligence (AI) and by their limited expertise, the AITIA Project targets the development of embedded AI techniques for industrial applications. This extended abstract presents the motivation and the solutions being developed towards four use cases: smart sensors, network intrusion detection, driver-assistance systems, and Industry 4.0.
Marcelo Brandalero, Mitko Veleski, Hector Gerardo Muñoz Hernandez, Muhammad Ali 0010, Laurens Le Jeune, Toon Goedemé, Nele Mentens, Jurgen Vandendriessche, Lancelot Lhoest, Bruno da Silva 0001, Abdellah Touhafi, Diana Göhringer, Michael Hübner 0001
FPL1
2021 Multi-Target Adaptive Reconfigurable Acceleration for Low-Power IoT Processing
abstract
Low-power processors for the Internet-of-Things (IoT) demand a high degree of adaptability to efficiently execute applications with different resource requirements under varying scenarios. Current single-ISA heterogeneous Chip Multiprocessors (CMPs), such as ARM's big.LITTLE, provide multiple cores and voltage/frequency levels to address this challenge. However, finding the best possible type of core and the corresponding voltage/frequency level for all the execution scenarios, which involve different applications and phases, remains far from being reached. In this article, we propose extending such a single-ISA heterogeneous CMP with a Coarse-Grained Reconfigurable Array (CGRA) and a hardware-based dynamic binary translation (DBT) module that transparently maps application code onto the CGRA for acceleration. To achieve low-energy levels and efficiently manage the power consumption of the CGRA, we introduce an additional voltage rail that enables operation in the Near-Threshold Voltage (NTV) regime when needed, leveraging key features of the CGRA's structure to address the implementation challenges of NTV computing. For less than 35 percent area overhead to the baseline CMP, performance and energy consumption are improved as follows. Compared to: (a) power-efficient execution in the LITTLE core, MuTARe achieves 29 percent reduction in energy consumption, and$2\times$speedup; (b) performance-efficient execution in the big core, a speedup of$1.6\times$with an energy reduction of 41 percent is achieved.
Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001
IEEE Trans. Computers1
2020 Proactive Aging Mitigation in CGRAs through Utilization-Aware Allocation
abstract
Resource balancing has been effectively used to mitigate the long-term aging effects of Negative Bias Temperature Instability (NBTI) in multi-core and Graphics Processing Unit (GPU) architectures. In this work, we investigate this strategy in Coarse-Grained Reconfigurable Arrays (CGRAs) with a novel application-to-CGRA allocation approach. By introducing important extensions to the reconfiguration logic and the datapath, we enable the dynamic movement of configurations throughout the fabric and allow overutilized Functional Units (FUs) to recover from stress-induced NBTI aging. Implementing the approach in a resource-constrained state-of-the-art CGRA reveals 2.2× lifetime improvement with negligible performance overheads and less than 10% increase in area.
Marcelo Brandalero, Bernardo Neuhaus Lignati, Antonio Carlos Schneider Beck, Muhammad Shafique 0001, Michael Hübner 0001
DAC1
2020 MCEA: A Resource-Aware Multicore CGRA Architecture for the Edge
abstract
Modern IoT edge devices must address the unpredictability of applications with strict power and temperature constraints. In this scenario, heterogeneous multicore architectures have been driving many solutions due to their high energy efficiency and ability to exploit Task-Level Parallelism. However, while their performance is highly dependent on the quality of the scheduling, their adaptability and generality get restricted when they use fixed-size hardware accelerators. Considering that, this work proposes MCEA, a transparent and power-adaptive multicore reconfigurable architecture. MCEA dynamically adapts the hardware to the workload rather than migrating applications; and predicatively sizes its reconfigurable accelerators without prior knowledge of the applications' behaviors. For that, MCEA uses a synergistic and online profiling system with power gating, achieving performance levels near of homogeneous architectures with fixed and oversized reconfigurable fabric (within 99% on average) while presenting energy efficiency levels similar to heterogeneous architectures statically tuned to a specific workload (within 99% on average). Therefore, MCEA improves Energy-Delay Product in 1.55x and 1.21x when compared to their heterogeneous and homogeneous counterparts, and in 4.72x when compared to a multicore with OoO processors only. We also show that MCEA outperforms a state-of-the-art reconfigurable architecture for the edge under the same power envelope.
Guilherme Korol, Michael G. Jordan, Marcelo Brandalero, Michael Hübner 0001, Mateus B. Rutzig, Antonio Carlos Schneider Beck
FPL3
2020 Endurance-Aware RRAM-Based Reconfigurable Architecture using TCAM Arrays
abstract
Field-Programmable Gate Arrays (FPGAs) have enabled the acceleration of important applications in the networking, cloud, and artificial intelligence domains, while providing a flexible fabric that can be reprogrammed on demand. Still, the high static power dissipation of FPGAs driven by Static Random Access Memories (SRAMs) leads them to energy consumption levels that may be unacceptable for several application domains. Reconfigurable fabrics with emerging Resistive RAM (RRAM) technologies have been considered as one of the most promising solutions to address these energy issues of current FPGAs. However, the low endurance and the high variability of these emerging devices present a threat to the demands for reconfiguration cycles of current applications, pushing for novel architectures and design strategies techniques for improving the device's lifetime. To address these challenges, we propose a novel reconfigurable architecture targeting classes of applications that require high flexibility in the field. More specifically, we introduce a reconfigurable architecture based on Ternary Content-Addressable Memories (TCAMs) that meets a double mission: to accelerate and tolerate endurance and variation issues supported by a CAD tool that foresees the reuse of data configuration, allowing for an increased endurance in the field. We present the potential of the proposed architecture and its synthesis flow for processing Regular Expression Matching (REM), widely used in network intrusion detection systems. The results show that the performance can achieve up to 32Gbps throughput at 0.89W, while improving the device's lifetime by two orders of magnitude.
João Paulo C. de Lima, Marcelo Brandalero, Luigi Carro
FPL2
2020 Reduced-Precision DWC for Mixed-Precision GPUs
abstract
Duplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing systems, including Graphics Processing Units (GPUs). DWC, however, introduces performance and energy consumption overheads that could be unacceptable for High-Performance Computing (HPC) or real-time safety-critical applications. In this work, we propose Reduced-Precision DWC (RP-DWC): an improvement over the traditional DWC approach that uses mixed-precision GPUs hardware resources to implement fault detection. We investigate, through both fault injection campaigns and accelerated neutron beam experiments, the impact of RPDWC onto performance, energy consumption, and its fault detection capabilites. We show that RP-DWC achieves on average 74% fault coverage (up to 86%) with very small overheads (0.1% time and 24% energy consumption overhead, in the best case).
Fernando Santos 0001, Marcelo Brandalero, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech
IOLTS2
2019 TransRec: Improving Adaptability in Single-ISA Heterogeneous Systems with Transparent and Reconfigurable Acceleration
abstract
Single-ISA heterogeneous systems, such as ARM's big.LITTLE, use microarchitecturally-different General-Purpose Processor cores to efficiently match the capabilities of the processing resources with applications' performance and energy requirements that change at run time. However, since only a fixed and non-configurable set of cores is available, reaching the best-possible match between the available resources and applications' requirements remains a challenge, especially considering the varying and unpredictable workloads. In this work, we propose TransRec, a hardware architecture which improves over these traditional heterogeneous designs. TransRec integrates a shared, transparent (i.e., no need to change application binary) and adaptive accelerator in the form of a Coarse-Grained Reconfigurable Array that can be used by any of the General-Purpose Processor cores for on-demand acceleration. Through evaluations with cycle-accurate gem5 simulations, synthesis of real RISC-V processor designs for a 15nm technology, and considering the effects of Dynamic Voltage and Frequency Scaling, we demonstrate that TransRec provides better performance-energy tradeoffs that are otherwise unachievable with traditional big.LITTLE-like designs. In particular, for less than 40% area overhead, TransRec can improve performance in the low-energy mode (LITTLE) by 2.28×, and can improve both performance and energy efficiency by 1.32× and 1.59×, respectively, in high-performance mode (big).
Marcelo Brandalero, Muhammad Shafique 0001, Luigi Carro, Antonio Carlos Schneider Beck
DATE1
2019 A Knapsack Methodology for Hardware-based DMR Protection against Soft Errors in Superscalar Out-of-Order Processors
abstract
High-performance superscalar processors have been adopted to satisfy the rising demand for processing applications of ever-growing complexity. This extra complexity, added to the increasing vulnerability of transistors due to technology scaling, poses a great challenge since these effects have also been proven to affect ground-level safety-critical applications. To increase microarchitectural resilience, designers may adopt Dual Modular Redundancy (DMR), which offers full fault detection. However, given that DMR incurs in high area and energy overheads, we propose a design-time methodology aiming to achieve the best tradeoff between resilience and area overhead, decreasing DMR costs and maintaining acceptable detection levels for such a complex design. This is done by adopting the Knapsack Problem (KSP) as a heuristic to identify the optimal micro-architectural structures that should be duplicated to achieve target resilience with the smallest possible area overhead. By injecting over 800k faults in 12 significant micro-architectural structures of different versions of the complex Berkeley Out-of-Order Machine (BOOM) superscalar processor modeled with RTL accuracy, we compare this optimal strategy against a greedy one, showing that 90% of vulnerability reduction may be achieved with 50.6% and 107.8% area overheads for the optimal and greedy strategies, respectively.
Rafael Billig Tonetto, Douglas Maciel Cardoso, Marcelo Brandalero, Luciano Volcan Agostini, Gabriel L. Nazar, José Rodrigo Azambuja, Antonio Carlos Schneider Beck
VLSI-SoC3
2019 Predicting performance in multi-core systems with shared reconfigurable accelerators
Marcelo Brandalero, Thiago Dadalt Souto, Luigi Carro, Antonio Carlos Schneider Beck
J. Syst. Archit.1
2018 Approximate on-the-fly coarse-grained reconfigurable acceleration for general-purpose applications
abstract
Approximate functional unit designs have the potential to reduce power consumption significantly compared to their precise counterparts; however, few works have investigated composing them to build generic accelerators. In this work, we do a design-space exploration of state-of-the-art approximate designs, propose a flow for designing approximate coarse-grained reconfigurable arrays (CGRAs), and discuss compilation and runtime reconfiguration issues. We compare the energy savings of precise and approximate reconfigurable acceleration and show that the latter can provide up to 50% additional power savings under a 10% quality loss constraint for the applications in the AxBench suite.
Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001
DAC1
2018 Employing classification-based algorithms for general-purpose approximate computing
abstract
Approximate computing has recently reemerged as a design solution for additional performance and energy improvements at the cost of output quality. In this paper, we propose using a tree-based classification algorithm as an approximation tool for general-purpose applications. We show that, without any hardware support, completely implemented in software, our approach can improve performance by up to 4x (1.95x on average) and reduce EDP by up to 19x (4.04 on average) when compared to precise executions. Besides that, in some cases, our software-based mechanism can even outperform traditional hardware-based Neural Network's state-of-the-art designs.
Geraldo F. Oliveira, Larissa Rozales Gonçalves, Marcelo Brandalero, Antonio Carlos Schneider Beck, Luigi Carro
DAC3
2018 Accelerating error-tolerant applications with approximate function reuse
Marcelo Brandalero, Leonardo Almeida da Silveira, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck
Sci. Comput. Program.1
2017 A Mechanism for energy-efficient reuse of decoding and scheduling of x86 instruction streams
abstract
Current superscalar x86 processors decompose each CISC instruction (variable-length and with multiple addressing modes) into multiple RISC-like pops at runtime so they can be pipelined and scheduled for concurrent execution. This challenging and power-hungry process, however, is usually repeated several times on the same instruction sequence, inefficiently producing the very same decoded and scheduled pops. Therefore, we propose a transparent mechanism to save the decoding and scheduling transformation for later reuse, so that next time the same instruction sequence is found it can automatically bypass the costly pipeline stages involved. We use a coarse-grained reconfigurable array as a means to save this transformation, since its structure enables the recovery of pops already allocated in time and space, and also larger ILP exploitation than superscalar processors. The technique can reduce the energy consumption of a powerful 8-issue superscalar by 31.4% at low area costs, while also improving performance by 32.6%.
Marcelo Brandalero, Antonio Carlos Schneider Beck
DATE1