VLDB 2026 Research / reviewers in the wild / expert
Luigi Carro
dblp:25/5214
· DBLP profile ↗
181ranked-venue papers
9as first author
19since 2021 · last 2024
0000-0002-7402-4780ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 170 · 7 first-author · 18 since 2021Software engineering, systems software and programming languages · 48 · 4 first-author · 4 since 2021Security and privacy · 4 · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Full-Stack Optimization for CAM-Only DNN InferenceabstractThe accuracy of neural networks has greatly improved across various domains over the past years. Their ever-increasing complexity, however, leads to prohibitively high energy demands and latency in von-Neumann systems. Several computing-in-memory (CIM) systems have recently been proposed to overcome this, but trade-offs involving accuracy, hardware reliability, and scalability for large models remain a challenge. Additionally, for some CIM designs, the activation movement still requires considerable time and energy. This paper explores the combination of algorithmic optimizations for ternary weight neural networks and associative processors (APs) implemented using racetrack memory (RTM). We propose a novel compilation flow to optimize convolutions on APs by reducing their arithmetic intensity. By leveraging the benefits of RTM-based APs, this approach substantially reduces data transfers within the memory while addressing accuracy, energy efficiency, and reliability concerns. Concretely, our solution improves the energy efficiency of ResNet-18 inference on ImageNet by 7.5× compared to crossbar in-memory accelerators while retaining software accuracy. João Paulo C. de Lima, Asif Ali Khan, Luigi Carro, Jerónimo Castrillón |
DATE | 3 |
| 2024 | Assessing the Impact of Compiler Optimizations on GPUs ReliabilityabstractGraphics Processing Units (GPUs) compilers have evolved in order to support general-purpose programming languages for multiple architectures. NVIDIA CUDA Compiler (NVCC) has many compilation levels before generating the machine code and applies complex optimizations to improve performance. These optimizations modify how the software is mapped in the underlying hardware; thus, as we show in this article, they can also affect GPU reliability. We evaluate the effects on the GPU error rate of the optimization flags applied at the NVCC Parallel Thread Execution (PTX) compiling phase by analyzing two NVIDIA GPU architectures (Kepler and Volta) and two compiler versions (NVCC 10.2 and 11.3). We compare and combine fault propagation analysis based on software fault injection, hardware utilization distribution obtained with application-level profiling, and machine instructions radiation-induced error rate measured with beam experiments. We consider eight different workloads and 144 combinations of compilation flags, and we show that optimizations can impact the GPUs’ error rate of up to an order of magnitude. Additionally, through accelerated neutron beam experiments on a NVIDIA Kepler GPU, we show that the error rate of the unoptimized GEMM (-O0 flag) is lower than the optimized GEMM’s (-O3 flag) error rate. When the performance is evaluated together with the error rate, we show that the most optimized versions (-O1 and -O3) always produce a higher amount of correct data than the unoptimized code (-O0). Fernando Santos 0001, Luigi Carro, Flavio Vella, Paolo Rech |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Reprogrammable Non-Linear Circuits Using ReRAM for NN AcceleratorsabstractAs the massive usage of artificial intelligence techniques spreads in the economy, researchers are exploring new techniques to reduce the energy consumption of Neural Network (NN) applications, especially as the complexity of NNs continues to increase. Using analog Resistive RAM devices to compute matrix-vector multiplication in O (1) time complexity is a promising approach, but it is true that these implementations often fail to cover the diversity of non-linearities required for modern NN applications. In this work, we propose a novel approach where Resistive RAMs themselves can be reprogrammed to compute not only the required matrix multiplications but also the activation functions, Softmax, and pooling layers, reducing energy in complex NNs. This approach offers more versatility for researching novel NN layouts compared to custom logic. Results show that our device outperforms analog and digital field-programmable approaches by up to 8.5× in experiments on real-world human activity recognition and language modeling datasets with convolutional neural network, generative pre-trained Transformer, and long short-term memory models. Rafael Fao de Moura, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Understanding and Improving GPUs' Reliability Combining Beam Experiments with Fault SimulationabstractGraphics Processing Units (GPUs) are being employed in High Performance Computing (HPC) and safety-critical applications, such as autonomous vehicles. This market shift led to significant improvements in the programming frameworks and performance evaluation tools and concerns about their reliability. GPU reliability evaluation is extremely challenging due to the parallel nature and high complexity of GPU architectures. We conducted the first cross-layer GPU reliability evaluation to unveil (and mitigate) GPU vulnerabilities. The proposed evaluation is achieved by comparing and combining extensive high-energy neutron beam experiments, massive fault simulation campaigns at both Register-Transfer Level (RTL) and software levels, and application profiling. Based on this extensive and detailed analysis, a novel accurate methodology to accurately estimate GPUs application FIT rate is proposed. Moreover, by employing the knowledge obtained from the cross-layer reliability evaluation, two novel hardening solutions for HPC and safety-critical applications are proposed: (1) Reduced Precision Duplication With Comparison (RP-DWC), which executes a redundant copy in a reduced precision. RP-DWC delivers excellent fault coverage, up to 86%, with minimal execution time and energy consumption overheads (13% and 24%, respectively). (2) Dedicated software solutions for hardening Convolutional Neural Networks (CNNs) that can correct up to 98% of the CNN errors. Fernando Santos 0001, Luigi Carro, Paolo Rech |
ETS | 2 |
| 2023 | Understanding and Improving GPUs' Reliability Combining Beam Experiments with Fault SimulationabstractGraphics Processing Units (GPUs) are essential in High Performance Computing (HPC) and safety-critical applications like autonomous vehicles. This market shift led to significant improvements in the programming frameworks and evaluation tools and concerns about their reliability. However, GPUs' high complexity poses challenges in evaluating their reliability. We conducted the first cross-layer GPU reliability evaluation to unveil and mitigate GPU vulnerabilities. The proposed evaluation is achieved by comparing and combining extensive neutron beam experiments, fault simulation campaigns, and application profiling. Based on this detailed analysis, a novel methodology to accurately estimate GPUs application FIT rate is proposed. The cross-layer evaluation enables two novel hardening solutions: (1) Reduced Precision Duplication With Comparison (RP-DWC) executes a redundant copy in reduced precision. RP-DWC delivers excellent fault coverage, up to 86%, with minimal execution time and energy consumption overheads (13% and 24%, respectively). (2) Dedicated software solutions for hardening Convolutional Neural Networks (CNNs) can detect up to 98% of errors. Fernando Santos 0001, Luigi Carro, Paolo Rech |
ITC | 2 |
| 2023 | Plug N' PIM: An integration strategy for Processing-in-Memory accelerators
Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
Integr. | 4 |
| 2023 | Data and Computation Reuse in CNNs Using Memristor TCAMsabstractExploiting computational and data reuse in CNNs is crucial for the successful design of resource-constrained platforms. In image recognition applications, high levels of input locality and redundancy present in CNNs have become the golden goose for skipping costly arithmetic operations. One promising technique for this consists in storing function responses of some input patterns into offline lookup tables and replacing online computation with search operations, which are highly efficient when implemented by emerging non-volatile memory technologies. In this work, we rethink both algorithm and architecture for exploiting locality and reuse opportunities by replacing entire convolutions with searches on Content-addressable Memories. By previously calculating convolution results and building compact lookup tables with our novel clustering algorithm, one can evaluate activations at constant time complexity, also requiring a single read operation of the current input tensor. Then, we devise a reconfigurable array of processing elements based on memristive Ternary Content-addressable Memories to efficiently implement the algorithmic solution and meet the flexibility requirements of several CNN architectures. Results show that our design reduces the number of multiplications and memory accesses proportionally to the number of convolutional layer channels. The average performance is 1,172 and 82 FPS for AlexNet and VGG-16 models, thus outperforming state-of-the-art works by 13×. Rafael Fao de Moura, João Paulo C. de Lima, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | Aggressive Performance Improvement on Processing-in-Memory Devices by Adopting HugepagesabstractProcessing-in-Memory (PIM) devices integrated into general-purpose systems demand virtual memory support. In this way, these devices can be seamlessly coupled to the software stack, while maintaining compatibility and security provided by address management via the Operating System (OS) without requiring disruptive programming efforts. Typically, PIM intends to access large volumes of data via vector operations, and thus can suffer severe penalties due to the high cost of page misses in the Translation Look-aside Buffer (TLB). Our study demonstrates the criticality of such penalties on the system's performance and that PIM must resort to large page sizes. The presented results exploit the native large pages available on the host, and they show substantial performance improvements$(84\times)$for wide-vector PIM operations with large pages. Paulo C. Santos 0001, Bruno Endres Forlin, Marco A. Z. Alves, Luigi Carro |
ASAP | 4 |
| 2022 | Quantization-Aware In-situ Training for Reliable and Accurate Edge AIabstractIn-memory analog computation based on memristor crossbars has become the most promising approach for DNN inference. Because compute and memory requirements are larger during training, memristive crossbars are also an alternative to train DNN models within a feasible energy budget for edge devices, especially in the light of trends towards security, privacy, latency, and energy reduction, by avoiding data transfer over the Internet. To enable online training and inference on the same device, however, there are still challenges related to different minimum bitwidth needed in each phase, and memristor non-idealities to be addressed. We provide an in-situ training framework that allows the network to adapt to hardware imperfections, while practically eliminating errors from weight quantization. We validate our methodology with image classifiers, namely MNIST and CIFAR10, by training NN models with 8-bit weights and quantizing to 2 bits. The training algorithm recovers up to 12 % of the accuracy lost to quantization errors even under high variability, reduces training energy by up to 6 ×, and allows for energy-efficient inferences using a single cell per synapse, hence enhancing robustness and accuracy for a smooth training-to-inference transition. João Paulo C. de Lima, Luigi Carro |
DATE | 2 |
| 2022 | STAP: An Architecture and Design Tool for Automata Processing on Memristor TCAMsabstractAccelerating finite-state automata benefits several emerging application domains that are built on pattern matching. In-memory architectures, such as the Automata Processor (AP), are efficient to speed them up, at least for outperforming traditional von-Neumann architectures. In spite of the AP’s massive parallelism, current APs suffer from poor memory density, inefficient routing architectures, and limited capabilities. Although these limitations can be lessened by emerging memory technologies, its architecture is still the major source of huge communication demands and lack of scalability. To address these issues, we present STAP , a Scalable TCAM-based architecture for Automata Processing . STAP adopts a reconfigurable array of processing elements, which are based on memristive Ternary CAMs (TCAMs), to efficiently implement Non-deterministic finite automata (NFAs) through proper encoding and mapping methods. The CAD tool for STAP integrates the design flow of automata applications, a specific mapping algorithm, and place and route tools for connecting processing elements by RRAM-based programmable interconnects. Results showed 1.47× higher throughput when processing 16-bit input symbols, and improvements of 3.9× and 25× on state and routing densities over the state-of-the-art AP, while preserving 10 4 programming cycles. João Paulo C. de Lima, Marcelo Brandalero, Michael Hübner 0001, Luigi Carro |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | Sim2PIM: A complete simulation framework for Processing-in-Memory
Bruno Endres Forlin, Paulo C. Santos 0001, Augusto E. Becker, Marco A. Z. Alves, Luigi Carro |
J. Syst. Archit. | 5 |
| 2022 | Reduced Precision DWC: An Efficient Hardening Strategy for Mixed-Precision ArchitecturesabstractDuplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing devices. However, it introduces performance and energy consumption overheads that could be unsuitable for high-performance computing or real-time safety-critical applications. In this article, we present Reduced-Precision Duplication with Comparison (RP-DWC) as a means to lower the overhead of DWC by executing the redundant copy in reduced precision. RP-DWC is particularly suitable for modern mixed-precision architectures, such as NVIDIA GPUs, that feature dedicated functional units for computing with programmable accuracy. We discuss the benefits and challenges associated with RP-DWC and show that the intrinsic difference between the mixed-precision copies allows for detecting most, but not all, errors. However, as the undetected faults are the ones that fall into the difference between precisions, they are the ones that produce a much smaller impact on the application output and, thus, might be tolerated. We investigate RP-DWC impact into fault detection, performance, and energy consumption on Volta GPUs. Through fault injection and beam experiment, using three microbenchmarks and four real applications, we show that RP-DWC achieves an excellent coverage (up to 86 percent) with minimal overheads (as low as 0.1 percent time and 24 percent energy consumption overhead). Fernando Santos 0001, Marcelo Brandalero, Michael B. Sullivan 0001, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech |
IEEE Trans. Computers | 6 |
| 2021 | Providing Plug N' Play for Processing-in-Memory AcceleratorsabstractAlthough Processing-in-Memory (PIM) emerged as a solution to avoid unnecessary and expensive data movements to/from host and accelerators, their widespread usage is still difficult, given that to effectively use a PIM device, huge and costly modifications must be done at the host processor side to allow instructions offloading, cache coherence, virtual memory management, and communication between different PIM instances. The present work addresses these challenges by presenting non-invasive solutions for those requirements. We demonstrate that, at compile-time, and without any host modifications or programmer intervention, it is possible to exploit already available resources to allow efficient host and PIM communication and task partitioning, without disturbing neither host nor memory hierarchy. We present Plug&PIM, a plug n' play strategy for PIM adoption with minimal performance penalties. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
ASP-DAC | 3 |
| 2021 | Sim2PIM: A Fast Method for Simulating Host Independent & PIM Agnostic DesignsabstractProcessing-in-Memory (PIM), with the help of modern memory integration technologies, has emerged as a practical approach to mitigate the memory wall and improve performance and energy efficiency in contemporary applications. However, there is a need for tools capable of quickly simulating different PIMs designs and their suitable integration with different hosts. This work presents Sim2PIM, a Simple Simulator for PIM devices that seamlessly integrates any PIM architecture with the host processor and memory hierarchy. Sim2PIM's simulation environment allows the user to describe a PIM architecture in different user-defined abstraction levels. The application code runs natively on the Host, with minimal overhead from the simulator integration, allowing Sim2PIM to collect precise metrics from the Hardware Performance Counters (HPCs). Our simulator is available to download at https://pim.computer/. Paulo C. Santos 0001, Bruno Endres Forlin, Luigi Carro |
DATE | 3 |
| 2021 | Revealing GPUs Vulnerabilities by Combining Register-Transfer and Software-Level Fault InjectionabstractThe complexity of both hardware and software makes GPUs reliability evaluation extremely challenging. A low level fault injection on a GPU model, despite being accurate, would take a prohibitively long time (months to years), while software fault injection, despite being quick, cannot access critical resources for GPUs and typically uses synthetic fault models (e.g., single bit-flips) that could result in unrealistic evaluations. This paper proposes to combine the accuracy of Register- Transfer Level (RTL) fault injection with the efficiency of software fault injection. First, on an RTL GPU model (FlexGripPlus), we inject over 1.5 million faults in low-level resources that are unprotected and hidden to the programmer, and characterize their effects on the output of common instructions. We create a pool of possible fault effects on the operation output based on the instruction opcode and input characteristics. We then inject these fault effects, at the application level, using an updated version of a software framework (NVBitFI). Our strategy reduces the fault injection time from the tens of years an RTL evaluation would need to tens of hours, thus allowing, for the first time on GPUs, to track the fault propagation from the hardware to the output of complex applications. Additionally, we provide a more realistic fault model and show that single bit-flip injection would underestimate the error rate of six HPC applications and two convolutional neural networks by up to 48parcent (18parcent on average). The RTL fault models and the injection framework we developed are made available in a public repository to enable third-party evaluations and ease results reproducibility. Fernando Santos 0001, Josie E. Rodriguez Condia, Luigi Carro, Matteo Sonza Reorda, Paolo Rech |
DSN | 3 |
| 2021 | Protecting GPU's Microarchitectural Vulnerabilities via Effective Selective HardeningabstractGraphics Processing Units (GPUs) are today adopted in several domains for which reliability is fundamental, such as self-driving cars and autonomous machines. Unfortunately, on one side GPUs have been shown to have a high error rate and, on the other side, the constraints imposed by real-time safety-critical applications make traditional, costly, replication-based hardening solutions inadequate. This paper proposes an effective microarchitectural selective hardening of GPU modules to mitigate those faults that affect instructions correct execution. We first characterize, through Register-Transfer Level (RTL) fault injections, the architectural vulnerabilities of a GPU model (FlexGripPlus). We specifically target transient faults in the functional units and pipeline registers of a GPU core. Then, we apply selective hardening by triplicating the locations in each module that we found to be more critical. The results show that selective hardening using Triple Modular Redundancy (TMR) can correct 85% to 99% of faults in the pipeline registers and from 50% to 100% of faults in the functional units. The proposed selective TMR strategy reduces the hardware overhead by up to 65% when compared with traditional TMR. Josie E. Rodriguez Condia, Paolo Rech, Fernando Santos 0001, Luigi Carro, Matteo Sonza Reorda |
IOLTS | 4 |
| 2021 | Demystifying GPU Reliability: Comparing and Combining Beam Experiments, Fault Simulation, and ProfilingabstractGraphics Processing Units (GPUs) have moved from being dedicated devices for multimedia and gaming applications to general-purpose accelerators employed in High-Performance Computing (HPC) and safety-critical applications such as autonomous vehicles. This market shift led to a burst in the GPU's computing capabilities and efficiency, significant improvements in the programming frameworks and performance evaluation tools, and a concern about their hardware reliability. In this paper, we compare and combine high-energy neutron beam experiments that account for more than 13 million years of natural terrestrial exposure, extensive architectural-level fault simulations that required more than 350 GPU hours (using SASSIFI and NVBitFI), and detailed application-level profiling. Our main goal is to answer one of the fundamental open questions in GPU reliability evaluation: whether fault simulation provides representative results that can be used to predict the failure rates of workloads running on GPUs. We show that, in most cases, fault simulation-based prediction for silent data corruptions is sufficiently close (differences lower than 5×) to the experimentally measured rates. We also analyze the reliability of some of the main GPU functional units (including mixed-precision and tensor cores). We find that the way GPU resources are instantiated plays a critical role in the overall system reliability and that faults outside the functional units generate most detectable errors. Fernando Santos 0001, Siva Kumar Sastry Hari, Pedro Martins Basso, Luigi Carro, Paolo Rech |
IPDPS | 4 |
| 2021 | Machine Learning Migration for Efficient Near-Data ProcessingabstractMachine Learning (ML) rises as a highly useful tool to analyze the vast amount of data generated in every field of science nowadays. Simultaneously, data movement inside computer systems gains more focus due to its high impact an time and energy consumption. In this context, the Near-Data Processing (NDP) architectures emerged as a prominent solution to increasing data by drastically reducing the required amount of data movement For NDP, we see three main approaches, Application-Specific Integrated Circuits (ASICs), full Central Processing Units (CPUs) and Graphics Processing Units (GPIIs), or vector units integration. However, previous work considered only ASICs, CPUs and GPUs when executing ML algorithms inside the memory. In this paper, we present an approach to execute ML algorithms near-data, using a general-purpose vector architecture and applying near-data parallelism to kernels from KNN, MEP, and CNN algorithms. To facilitate this process, we also present an NDP intrinsics library to ease the evaluation and debugging tasks. Our results show speedups up to fox for KNN, 11× for MLP, and 3× for convolution when processing near-data compared to a high-performance ×86 baseline. Aline S. Cordeiro, Sairo R. dos Santos, Francis B. Moreira 0001, Paulo C. Santos 0001, Luigi Carro, Marco A. Z. Alves |
PDP | 5 |
| 2021 | Multi-Target Adaptive Reconfigurable Acceleration for Low-Power IoT ProcessingabstractLow-power processors for the Internet-of-Things (IoT) demand a high degree of adaptability to efficiently execute applications with different resource requirements under varying scenarios. Current single-ISA heterogeneous Chip Multiprocessors (CMPs), such as ARM's big.LITTLE, provide multiple cores and voltage/frequency levels to address this challenge. However, finding the best possible type of core and the corresponding voltage/frequency level for all the execution scenarios, which involve different applications and phases, remains far from being reached. In this article, we propose extending such a single-ISA heterogeneous CMP with a Coarse-Grained Reconfigurable Array (CGRA) and a hardware-based dynamic binary translation (DBT) module that transparently maps application code onto the CGRA for acceleration. To achieve low-energy levels and efficiently manage the power consumption of the CGRA, we introduce an additional voltage rail that enables operation in the Near-Threshold Voltage (NTV) regime when needed, leveraging key features of the CGRA's structure to address the implementation challenges of NTV computing. For less than 35 percent area overhead to the baseline CMP, performance and energy consumption are improved as follows. Compared to: (a) power-efficient execution in the LITTLE core, MuTARe achieves 29 percent reduction in energy consumption, and$2\times$speedup; (b) performance-efficient execution in the big core, a speedup of$1.6\times$with an energy reduction of 41 percent is achieved. Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001 |
IEEE Trans. Computers | 2 |
| 2020 | G-PUF: An Intrinsic PUF Based on GPU Error SignaturesabstractPhysically Unclonable Functions (PUFs) are security primitives that provide trustworthy hardware for key-generation and device authentication. Among them, in contrast to dedicated PUFs, intrinsic PUFs are created from existing hardware components that exploit their variability through software. In this work we focus on GPUs and present G-PUF, a PUF implemented entirely in software on CUDA and hence does not require hardware modifications. Our results show that G-PUF has comparable characteristics to SRAM and DRAM PUFs in terms of uniformity 55.61% and reliability 90.09%. Bruno Endres Forlin, Ronaldo Husemann, Luigi Carro, Cezar Reinbrecht, Said Hamdioui, Mottaqiallah Taouil |
ETS | 3 |
| 2020 | Endurance-Aware RRAM-Based Reconfigurable Architecture using TCAM ArraysabstractField-Programmable Gate Arrays (FPGAs) have enabled the acceleration of important applications in the networking, cloud, and artificial intelligence domains, while providing a flexible fabric that can be reprogrammed on demand. Still, the high static power dissipation of FPGAs driven by Static Random Access Memories (SRAMs) leads them to energy consumption levels that may be unacceptable for several application domains. Reconfigurable fabrics with emerging Resistive RAM (RRAM) technologies have been considered as one of the most promising solutions to address these energy issues of current FPGAs. However, the low endurance and the high variability of these emerging devices present a threat to the demands for reconfiguration cycles of current applications, pushing for novel architectures and design strategies techniques for improving the device's lifetime. To address these challenges, we propose a novel reconfigurable architecture targeting classes of applications that require high flexibility in the field. More specifically, we introduce a reconfigurable architecture based on Ternary Content-Addressable Memories (TCAMs) that meets a double mission: to accelerate and tolerate endurance and variation issues supported by a CAD tool that foresees the reuse of data configuration, allowing for an increased endurance in the field. We present the potential of the proposed architecture and its synthesis flow for processing Regular Expression Matching (REM), widely used in network intrusion detection systems. The results show that the performance can achieve up to 32Gbps throughput at 0.89W, while improving the device's lifetime by two orders of magnitude. João Paulo C. de Lima, Marcelo Brandalero, Luigi Carro |
FPL | 3 |
| 2020 | Leveraging reuse and endurance by efficient mapping and placement for NVM-based FPGAsabstractAlthough NVM technologies can bring higher density, near-zero leakage power, and CMOS compatibility, the reduced endurance of such devices is still a major limitation for a broad range of applications, including NVM-based FPGAs. In this work, we propose to drastically reduce LUT writing and word flips occurrence by taking reconfiguration reuse as a target of optimization in the CAD flow. Specifically, we investigate circuit reuse in LUT- and CAM-based FPGAs, and propose a new endurance-aware mapping and placement algorithm. Simulation results show that much greater reuse can be achieved with PLA-style, and hence higher endurance is expected. João Paulo C. de Lima, Rafael Fao de Moura, Luigi Carro |
IOLTS | 3 |
| 2020 | Reduced-Precision DWC for Mixed-Precision GPUsabstractDuplication with Comparison (DWC) is an effective software-level solution to improve the reliability of computing systems, including Graphics Processing Units (GPUs). DWC, however, introduces performance and energy consumption overheads that could be unacceptable for High-Performance Computing (HPC) or real-time safety-critical applications. In this work, we propose Reduced-Precision DWC (RP-DWC): an improvement over the traditional DWC approach that uses mixed-precision GPUs hardware resources to implement fault detection. We investigate, through both fault injection campaigns and accelerated neutron beam experiments, the impact of RPDWC onto performance, energy consumption, and its fault detection capabilites. We show that RP-DWC achieves on average 74% fault coverage (up to 86%) with very small overheads (0.1% time and 24% energy consumption overhead, in the best case). Fernando Santos 0001, Marcelo Brandalero, Pedro Martins Basso, Michael Hübner 0001, Luigi Carro, Paolo Rech |
IOLTS | 5 |
| 2019 | A Compiler for Automatic Selection of Suitable Processing-in-Memory InstructionsabstractAlthough not a new technique, due to the advent of 3D-stacked technologies, the integration of large memories and logic circuitry able to compute large amount of data has revived the Processing-in-Memory (PIM) techniques. PIM is a technique to increase performance while reducing energy consumption when dealing with large amounts of data. Despite several designs of PIM are available in the literature, their effective implementation still burdens the programmer. Also, various PIM instances are required to take advantage of the internal 3D-stacked memories, which further increases the challenges faced by the programmers. In this way, this work presents the Processing-In-Memory cOmpiler (PRIMO). Our compiler is able to efficiently exploit large vector units on a PIM architecture, directly from the original code. PRIMO is able to automatically select suitable PIM operations, allowing its automatic offloading. Moreover, PRIMO concerns about several PIM instances, selecting the most suitable instance while reduces internal communication between different PIM units. The compilation results of different benchmarks depict how PRIMO is able to exploit large vectors, while achieving a near-optimal performance when compared to the ideal execution for the case study PIM. PRIMO allows a speedup of 38× for specific kernels, while on average achieves 11.8 × for a set of benchmarks from PolyBench Suite. Hameeza Ahmed, Paulo C. Santos 0001, João Paulo C. de Lima, Rafael Fao de Moura, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 7 |
| 2019 | TransRec: Improving Adaptability in Single-ISA Heterogeneous Systems with Transparent and Reconfigurable AccelerationabstractSingle-ISA heterogeneous systems, such as ARM's big.LITTLE, use microarchitecturally-different General-Purpose Processor cores to efficiently match the capabilities of the processing resources with applications' performance and energy requirements that change at run time. However, since only a fixed and non-configurable set of cores is available, reaching the best-possible match between the available resources and applications' requirements remains a challenge, especially considering the varying and unpredictable workloads. In this work, we propose TransRec, a hardware architecture which improves over these traditional heterogeneous designs. TransRec integrates a shared, transparent (i.e., no need to change application binary) and adaptive accelerator in the form of a Coarse-Grained Reconfigurable Array that can be used by any of the General-Purpose Processor cores for on-demand acceleration. Through evaluations with cycle-accurate gem5 simulations, synthesis of real RISC-V processor designs for a 15nm technology, and considering the effects of Dynamic Voltage and Frequency Scaling, we demonstrate that TransRec provides better performance-energy tradeoffs that are otherwise unachievable with traditional big.LITTLE-like designs. In particular, for less than 40% area overhead, TransRec can improve performance in the low-energy mode (LITTLE) by 2.28×, and can improve both performance and energy efficiency by 1.32× and 1.59×, respectively, in high-performance mode (big). Marcelo Brandalero, Muhammad Shafique 0001, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2019 | Impact of Reduced Precision in the Reliability of Deep Neural Networks for Object DetectionabstractModern Graphics Processing Units (GPUs) have dedicated hardware to execute floating-point operations with different precisions (64-bit double, 32-bit single, and 16-bit half). Using reduced precision for specific applications like Deep Neural Networks (DNNs) has been shown to reduce both the execution time and power consumption with negligible effects on the DNNs' accuracy. As GPUs are playing a critical role in DNN for object detection and get into safety-critical environments, their reliability is becoming a growing concern. In this paper, we evaluate the reliability of a DNN implemented in three different precisions (half, single, and double) on NVIDIA mixed-precision GPUs. We evaluate not only the error rate of the applications but also the effects of the errors on the final detection. We perform extensive fault-injection campaign on the register file of NVIDIA mixed-precision GPUs. We found that reducing data and operation precision increases the probability for the fault to impact the DNN detection and classification. Then, we complement the fault injection study with beam experiments. We exposed YOLOv3 running on Tesla V100s to neutron beams and found that the use of half precision reduces the error rate of up to 2x. The smaller exposed area and improved performances brought by reduced precision is then likely to increase the DNN reliability. Fernando Santos 0001, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech |
ETS | 3 |
| 2019 | Detecting Errors in Convolutional Neural Networks Using Inter Frame Spatio-Temporal CorrelationabstractObject detection, a critical feature for autonomous vehicles, is performed today using Convolutional Neural Networks (CNNs). Errors in a CNN execution can modify the way the vehicle sense the surrounding environment, potentially causing accidents or unexpected behaviors. The high computational requirements of CNNs combined with the need to perform detection in real-time allow little margin for implementing error detection. In this paper, we present an extremely efficient error detection solution for CNN based on the observation that, in the absence of errors, the differences between the input frames and the detection provided by the CNN should be strictly correlated. In other words, if the image between two subsequent frames does not change significantly, the detection should also be very similar. Similarly, if the detection varies considerably from a frame to the next, then the input image should also have been different. Whenever input images and output detection don't correlate we can detect a error. After formalizing and evaluating the inter-frame and output correlation thresholds, we implement and validate the detection strategy, utilizing data from previous radiation experiments. Exploiting the intrinsic efficiency in processing images of devices used to execute CNNs, we can detect up to 80% of errors while adding low overhead. Lucas Draghetti, Fernando Santos 0001, Luigi Carro, Paolo Rech |
IOLTS | 3 |
| 2019 | Predicting performance in multi-core systems with shared reconfigurable accelerators
Marcelo Brandalero, Thiago Dadalt Souto, Luigi Carro, Antonio Carlos Schneider Beck |
J. Syst. Archit. | 3 |
| 2019 | Aggressive Energy Reduction for Video Inference with Software-only StrategiesabstractIn the past years, several works have proposed custom hardware and software-based techniques for the acceleration of Convolutional Neural Networks (CNNs). Most of these works focus on saving computations by changing the used precision or modifying frame processing. To reach a more aggressive energy reduction, in this paper we propose software-only modifications to the CNNs inference process. Our approach exploits the inherent locality in videos by replacing entire frame computations with a movement prediction algorithm. Furthermore, when a frame must be processed, we avoid energy-demanding floating-point operations, and at the same time reduce memory accesses by employing look-up tables in place of the original convolutions. Using the proposed approach, one can reach significant energy gains of more than 25× for security cameras, and 12× for moving vehicles applications, with only small software modifications. Larissa Rozales Gonçalves, Rafael Fao de Moura, Luigi Carro |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | Analyzing and Increasing the Reliability of Convolutional Neural Networks on GPUsabstractGraphics processing units (GPUs) are playing a critical role in convolutional neural networks (CNNs) for image detection. As GPU-enabled CNNs move into safety-critical environments, reliability is becoming a growing concern. In this paper, we evaluate and propose strategies to improve the reliability of object detection algorithms, as run on three NVIDIA GPU architectures. We consider three algorithms: 1) you only look once; 2) a faster region-based CNN (Faster R-CNN); and 3) a residual network, exposing live hardware to neutron beams. We complement our beam experiments with fault injection to better characterize fault propagation in CNNs. We show that a single fault occurring in a GPU tends to propagate to multiple active threads, significantly reducing the reliability of a CNN. Moreover, relying on error correcting codes dramatically reduces the number of silent data corruptions (SDCs), but does not reduce the number of critical errors (i.e., errors that could potentially impact safety-critical applications). Based on observations on how faults propagate on GPU architectures, we propose effective strategies to improve CNN reliability. We also consider the benefits of using an algorithm-based fault-tolerance technique for matrix multiplication, which can correct more than 87% of the critical SDCs in a CNN, while redesigning maxpool layers of the CNN to detect up to 98% of critical SDCs. Fernando Santos 0001, Pedro Foletto Pimenta, Caio B. Lunardi, Lucas Draghetti, Luigi Carro, David R. Kaeli, Paolo Rech |
IEEE Trans. Reliab. | 5 |
| 2018 | Design space exploration for PIM architectures in 3D-stacked memoriesabstractScaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3D-stacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-of-the-art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High-Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth. João Paulo C. de Lima, Paulo C. Santos 0001, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
CF | 5 |
| 2018 | Approximate on-the-fly coarse-grained reconfigurable acceleration for general-purpose applicationsabstractApproximate functional unit designs have the potential to reduce power consumption significantly compared to their precise counterparts; however, few works have investigated composing them to build generic accelerators. In this work, we do a design-space exploration of state-of-the-art approximate designs, propose a flow for designing approximate coarse-grained reconfigurable arrays (CGRAs), and discuss compilation and runtime reconfiguration issues. We compare the energy savings of precise and approximate reconfigurable acceleration and show that the latter can provide up to 50% additional power savings under a 10% quality loss constraint for the applications in the AxBench suite. Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001 |
DAC | 2 |
| 2018 | Employing classification-based algorithms for general-purpose approximate computingabstractApproximate computing has recently reemerged as a design solution for additional performance and energy improvements at the cost of output quality. In this paper, we propose using a tree-based classification algorithm as an approximation tool for general-purpose applications. We show that, without any hardware support, completely implemented in software, our approach can improve performance by up to 4x (1.95x on average) and reduce EDP by up to 19x (4.04 on average) when compared to precise executions. Besides that, in some cases, our software-based mechanism can even outperform traditional hardware-based Neural Network's state-of-the-art designs. Geraldo F. Oliveira, Larissa Rozales Gonçalves, Marcelo Brandalero, Antonio Carlos Schneider Beck, Luigi Carro |
DAC | 5 |
| 2018 | Processing in 3D memories to speed up operations on complex data structuresabstractPointer chasing has been, for years, the kernel operation employed by diverse data structures, from graphs to hash tables and dictionaries. However, due to the bewildering growth in the volume of data that current applications have to deal with, performing pointer chasing operations have become a major source of performance and energy bottleneck, due to its sparse memory access behavior. In this work, we aim to tackle this problem by taking advantage of the already available parallelism present in today's 3D-stacked memories. We present a simple mechanism that can accelerate pointer chasing operations by making use of a state-of-the-art PIM design that executes in-memory vector operations. The key idea behind our design is to run speculative loads, in parallel, based on a given memory address in a reconfigurable window of addresses. Our design can perform pointer-chasing operations on b+tree 4.9 χ faster when compared to modern baseline systems. Besides that, since our device avoids data movement, we can also reduce energy consumption by 85% when compared to the baseline. Paulo C. Santos 0001, Geraldo F. Oliveira, João Paulo C. de Lima, Marco A. Z. Alves, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 5 |
| 2018 | HIPE: HMC instruction predication extension applied on database processingabstractThe recent Hybrid Memory Cube (HMC) is a smart memory which includes functional units inside one logic layer of the 3D stacked memory design. In order to execute instructions inside the Hybrid Memory Cube (HMC), the processor needs to send instructions to be executed near data, keeping most of the pipeline complexity inside the processor. Thus, control-flow and data-flow dependencies are all managed inside the processor, in such way that only update instructions are supported by the HMC. In order to solve data-flow dependencies inside the memory, previous work proposed HMC Instruction Vector Extensions (HIVE), which embeds a high number of functional units with a interlock register bank. In this work we propose HMC Instruction Prediction Extensions (HIPE), that supports predicated execution inside the memory, in order to transform control-flow dependencies into data-flow dependencies. Our mechanism focus on removing the high latency iteration between the processor and the smart memory during the execution of branches that depends on data processed inside the memory. In this paper we evaluate a balanced design of HIVE comparing to x86 and HMC executions. After we show the HIPE mechanism results when executing a database workload, which is a strong candidate to use smart memories. We show interesting trade-offs of performance when comparing our mechanism to previous work. Diego G. Tomé, Paulo C. Santos 0001, Luigi Carro, Eduardo C. de Almeida, Marco A. Z. Alves |
DATE | 3 |
| 2018 | Repair of FPGA-Based Real-Time Systems With Variable SlacksabstractField-programmable gate arrays (FPGAs) based on SRAM cells are an attractive alternative for real-time system designers, as they offer high density, low cost, and high performance. The use of SRAM cells in the FPGA’s configuration memory, while enabling these desirable characteristics, also creates a reliability hazard as RAM cells are susceptible to single-event upsets (SEUs). The usual approach is the use of double or triple redundancy allied with a correction mechanism, such as periodic scrubbing. Although scrubbing is an effective technique to remove SEU-induced errors, the repair of real-time systems presents specific challenges, such as avoiding failures by missing real-time deadlines. In this article, a novel approach is proposed to use a deadline-aware scrubbing scheme with negligible area costs that dynamically chooses the scrubbing starting position. Such a scheme allows us to avoid missing real-time deadlines while maximizing the repair probability given a bounded repair time. Our approach reduces the failure rate, considering the probability of missing deadlines due to faults, by 33.39% on average, with an average area cost of 1.23%. Leonardo P. Santos, Gabriel L. Nazar, Luigi Carro |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | An energy-efficient memory hierarchy for multi-issue processorsabstractEmbedded processors must rely on the efficient use of instruction-level parallelism to answer the performance and energy needs of modern applications. However, a limiting factor to better use available resources inside the processor concerns memory bandwidth. Adding extra ports to allow for more data accesses drastically increases costs and energy. In this paper, we present a novel memory architecture system for embedded multi-issue processors that can overcome the limited memory bandwidth without adding extra ports to the system. We combine the use of software-managed memories (SMM) with the data cache to provide a system with a higher throughput without increasing the number of ports. Compiler-automated code transformations minimize the effort of programmers to benefit from the proposed architecture. Our experimental results show an average speedup of 1.17x, while consuming 69% less dynamic energy and on average 74.7% lower energy-delay product regarding data memory in comparison to a baseline processor. Tiago T. Jost, Gabriel L. Nazar, Luigi Carro |
DATE | 3 |
| 2017 | Operand size reconfiguration for big data processing in memoryabstractNowadays, applications that predominantly perform lookups over large databases are becoming more popular with column-stores as the database system architecture of choice. For these applications, Hybrid Memory Cubes (HMCs) can provide bandwidth of up to 320 GB/s and represents the best choice to keep the throughput for these ever increasing databases. However, even with the high available memory bandwidth and processing power, in order to achieve the peak performance, data movements through the memory hierarchy consumes an unnecessary amount of time and energy. In order to accelerate database operations, and reduce the energy consumption of the system, this paper presents the Reconfigurable Vector Unit (RVU) that enables massive and adaptive in-memory processing, extending the native HMC instructions and also increasing its effectiveness. RVU enables the programmer to reconfigure it to perform as a large vector unit or multiple small vectors units to better adjust for the application needs during different computation phases. Due to its adaptability, RVU is capable of achieving performance increase of 27 χ on average and reduce the DRAM energy consumption in 29% when compared to an x86 processor with 16 cores. Compared with the state-of-the-art mechanism capable of performing large vector operations with fixed size, inside the HMC, RVU performed up to 12% better in terms of performance and improve in 53% the energy consumption. Paulo C. Santos 0001, Geraldo F. Oliveira, Diego G. Tomé, Marco A. Z. Alves, Eduardo C. de Almeida, Luigi Carro |
DATE | 6 |
| 2017 | Radiation-Induced Error Criticality in Modern HPC Parallel AcceleratorsabstractIn this paper, we evaluate the error criticality of radiation-induced errors on modern High-Performance Computing (HPC) accelerators (Intel Xeon Phi and NVIDIA K40) through a dedicated set of metrics. We show that, as long as imprecise computing is concerned, the simple mismatch detection is not sufficient to evaluate and compare the radiation sensitivity of HPC devices and algorithms. Our analysis quantifies and qualifies radiation effects on applications' output correlating the number of corrupted elements with their spatial locality. Also, we provide the mean relative error (dataset-wise) to evaluate radiation-induced error magnitude. We apply the selected metrics to experimental results obtained in various radiation test campaigns for a total of more than 400 hours of beam time per device. The amount of data we gathered allows us to evaluate the error criticality of a representative set of algorithms from HPC suites. Additionally, based on the characteristics of the tested algorithms, we draw generic reliability conclusions for broader classes of codes. We show that arithmetic operations are less critical for the K40, while Xeon Phi is more reliable when executing particles interactions solved through Finite Difference Methods. Finally, iterative stencil operations seem the most reliable on both architectures. Daniel Oliveira 0002, Laércio Lima Pilla, Mauricio Hanzich, Vinicius Fratin, Fernando Santos 0001, Caio B. Lunardi, José María Cela, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech |
HPCA | 9 |
| 2016 | Large vector extensions inside the HMC
Marco A. Z. Alves, Matthias Diener, Paulo C. Santos 0001, Luigi Carro |
DATE | 4 |
| 2016 | A reconfigurable heterogeneous multicore with a homogeneous ISA
Jeckson Dellagostin Souza, Luigi Carro, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DATE | 2 |
| 2016 | Scalable memory architecture for soft-core processorsabstractRestrictions over memory performance have always had a great impact on soft-core processors. The reduced number of ports on FPGAs' block RAMs may limit the exploitation of parallelism on soft-core processors that are implemented on top of these devices. Multiple memory ports on FPGAs are cumbersome and do not scale well, having a high cost in area and power consumption when implemented. In order to mitigate the impact of the memory bottleneck on such devices, we propose a scalable memory architecture for soft-cores. We make use of software-managed memories to build a memory system capable of improving performance and instruction-level parallelism (ILP) on soft-core processors. Results show that our architecture overcomes the limited parallelism realized on a dual-ported processor, reducing execution time by 16.5%. These improvements come with no area costs, as the processor is kept with the same total memory. Automated code transformations implemented within the LLVM compiler keep changes in application code to a minimum. We also show that our architecture scales better when boosting the number of functional units in the system. Tiago T. Jost, Gabriel L. Nazar, Luigi Carro |
ICCD | 3 |
| 2016 | Exploring Cache Size and Core Count Tradeoffs in Systems with Reduced Memory Access LatencyabstractOne of the main challenges for computer architects is how to hide the high average memory access latency from the processor. In this context, Hybrid Memory Cubes (HMCs) can provide substantial energy and bandwidth improvements compared to traditional memory organizations. However, it is not clear how this reduced average memory access latency will impact the LLC. For applications with high cache miss ratios, the latency to search for the data inside the cache memory will impact negatively on the performance. The importance of this overhead depends on the memory access latency. In this paper, we present an evaluation of the L3 cache importance on a high performance processor using HMC also exploring chip area tradeoffs between the cache size and number of processor cores. We show that the high bandwidth provided by HMC memories can eliminate the need for L3 caches, removing hardware and making room for more processing power. Our evaluations show that performance increased 37% and the EDP improved 12% while maintaining the same original chip area in a wide range of parallel applications, when compared to DDR3 memories. Paulo C. Santos 0001, Marco A. Z. Alves, Matthias Diener, Luigi Carro, Philippe Olivier Alexandre Navaux |
PDP | 4 |
| 2016 | Exploiting Idle Hardware to Provide Low Overhead Fault Tolerance for VLIW ProcessorsabstractBecause of technology scaling, the soft error rate has been increasing in digital circuits, which affects system reliability. Therefore, modern processors, including VLIW architectures, must have means to mitigate such effects to guarantee reliable computing. In this scenario, our work proposes three low overhead fault tolerance approaches based on instruction duplication with zero latency detection, which uses a rollback mechanism to correct soft errors in the pipelanes of a configurable VLIW processor. The first uses idle issue slots within a period of time to execute extra instructions considering distinct application phases. The second works at a finer grain, adaptively exploiting idle functional units at run-time. However, some applications present high instruction-level parallelism (ILP), so the ability to provide fault tolerance is reduced: less functional units will be idle, decreasing the number of potential duplicated instructions. The third approach attacks this issue by dynamically reducing ILP according to a configurable threshold, increasing fault tolerance at the cost of performance. While the first two approaches achieve significant fault coverage with minimal area and power overhead for applications with low ILP, the latter improves fault tolerance with low performance degradation. All approaches are evaluated considering area, performance, power dissipation, and error coverage. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Luigi Carro, Fernanda Lima Kastensmidt, Stephan Wong, Antonio Carlos Schneider Beck |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2016 | Evaluation of Histogram of Oriented Gradients Soft Errors Criticality for Automotive ApplicationsabstractPedestrian detection reliability is a key problem for autonomous or aided driving, and methods that use Histogram of Oriented Gradients (HOG) are very popular. Embedded Graphics Processing Units (GPUs) are exploited to run HOG in a very efficient manner. Unfortunately, GPUs architecture has been shown to be particularly vulnerable to radiation-induced failures. This article presents an experimental evaluation and analytical study of HOG reliability. We aim at quantifying and qualifying the radiation-induced errors on pedestrian detection applications executed in embedded GPUs. We analyze experimental results obtained executing HOG on embedded GPUs from two different vendors, exposed for about 100 hours to a controlled neutron beam at Los Alamos National Laboratory. We consider the number and position of detected objects as well as precision and recall to discriminate critical erroneous computations. The reported analysis shows that, while being intrinsically resilient (65% to 85% of output errors only slightly impact detection), HOG experienced some particularly critical errors that could result in undetected pedestrians or unnecessary vehicle stops. Additionally, we perform a fault-injection campaign to identify HOG critical procedures. We observe that Resize and Normalize are the most sensitive and critical phases, as about 20% of injections generate an output error that significantly impacts HOG detection. With our insights, we are able to find those limited portions of HOG that, if hardened, are more likely to increase reliability without introducing unnecessary overhead. Fernando Santos 0001, Lucas Weigel, Cláudio R. Jung, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech |
ACM Trans. Archit. Code Optim. | 5 |
| 2016 | Live-Out Register Fencing: Interrupt-Triggered Soft Error Correction Based on the Elimination of Register-to-Register CommunicationabstractThis article introduces Live-Out Register Fencing (LoRF), a soft error correction mechanism that uses the novel Spill Register File as a container of checkpointing data. LoRF’s Spill Register File holds the values shared among basic blocks in the program, and, coupled with a new compilation strategy, LoRF allows for error correction in the same basic block where the error was detected. In LoRF, error correction is triggered by a hardware interrupt that restores the registers of a basic block from the Spill Register File. After these registers are restored, the basic block where the error was detected can just be re-executed, thus reducing the costs of error recovery. LoRF’s error correction policy eliminates the need for expensive architectural support for checkpointing and rollback, reducing the performance overhead of online soft error correction. LoRF relies on both a modified processor architecture and a corresponding compiler. The architecture was implemented in synthesizable VHDL, whereas the compiler was developed as an extension of the LLVM framework. Fault injection experiments support an error correction coverage of 99.35% and a mean performance overhead of 1.33 for the entire life cycle of an error from its occurrence to its elimination from the system. Ronaldo Rodrigues Ferreira, Gabriel L. Nazar, Jean da Rolt, Álvaro F. Moreira, Luigi Carro |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2015 | Saving memory movements through vector processing in the DRAMabstractDespite the ability of modern processors to execute a variety of algorithms efficiently through instructions based on registers with ever-increasing widths, some applications present poor performance due to the limited interconnection bandwidth between main memory and processing units. Near-data processing has started to gain acceptance as an accelerator device due to the technology constraints and high costs associated with data transfer. However, previous approaches to near-data computing do not provide general-purpose processing, or require large amounts of logic and do not fully use the potential of the DRAM devices. These issues limited its wide adoption. In this paper, we present the Memory Vector Extensions (MVX), which implement vector instructions directly inside the DRAM devices, therefore avoiding data movement between memory and processing units, while requiring a lower amount of logic than previous approaches. MVX is able to obtain up to 211× increase in performance for application kernels with a high spatial locality and a low temporal locality. Comparing to an embedded processor with 8 cores and 2 memory channels that supports AVX-512 instructions, MVX performs 24× faster on average for three well known algorithms. Marco A. Z. Alves, Paulo C. Santos 0001, Francis B. Moreira 0001, Matthias Diener, Luigi Carro |
CASES | 5 |
| 2015 | Exploiting cache conflicts to reduce radiation sensitivity of operating systems on embedded systemsabstractIn this paper, we investigate how the presence of a general purpose operating system influences the reliability of modern embedded Systems-on-Chips (SoCs). We analytically study the difference in the reliability of SoCs when executing the application bare to the metal and on top of the Linux kernel. Our analysis demonstrates that Linux presence barely affects the Silent Data Corruption rate while it greatly increases the system Functional Interruption (FI) rate (up to 7.48 times) if no preventive measures are taken. Furthermore, we analytically show that cache conflicts between the operating system and application can significantly reduce the Linux-induced FI rate increase. To support our analysis, a total of four representative embedded applications were individually executed bare to the metal and on top of Linux on a 28nm ARM-based SoC exposed to an accelerated neutron beam. Our experimental results demonstrate that, by carefully tuning cache conflicts, it is possible to successfully limit the Linux-induced FI rate increase to 3.85 times. The proposed solution is general and readily applied to a broad set of applications and embedded systems. Thiago Santini, Paolo Rech, Luigi Carro, Flávio Rech Wagner |
CASES | 3 |
| 2015 | NFRs early estimation through software metrics
Andrws Vieira, Pedro Faustini, Luigi Carro, Érika F. Cota |
DATE | 3 |
| 2015 | Understanding GPU errors on large-scale HPC systems and the implications for system design and operationabstractIncrease in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience. Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland |
HPCA | 11 |
| 2015 | Performance evaluation of hierarchical NoC topologies for stacked 3D ICsabstractThree-Dimensional (3D) integrated circuits (ICs) have emerged as a solution to attend the demand of high performance, low power and high density of the MultiProcessors Systems-on-Chip (MPSoCs). However, some important issues need to be observed in the interconnection device for 3D designs. In this paper we have presented the advantages of the 3D-HiCIT network-on-chip (NoC) when compared to other hierarchical topologies in terms of flexibility, scalability and performance. Considering all constraints of this new scenario of circuit integration, the proposed hierarchical 3D NoC verified in this work meets well with the reality of these designs, presenting gains in several aspects. Debora Matos, Max Prass, Márcio Eduardo Kreutz, Luigi Carro, Altamiro Amadeu Susin |
ISCAS | 4 |
| 2015 | Bit-Flip Aware Control-Flow Error DetectionabstractRecent increase of transient fault rates has made processor reliability a major concern. Moreover performance improvements are required for many of today's embedded systems. At the same time software implemented fault detection remains the only option for off-the-shelf processors. Software methods, however, introduce significant performance overheads due to the additional instructions required for the detection. A good observation is that often code segments not susceptible to faults are protected. In this paper we propose a technique for systematic analysis of the bit-flip effects on the program control-flow in order to identify only those locations susceptible to control-flow errors and hence minimize the number of fault detection assertions. We instrument the code with minimal overhead, while maintaining high fault coverage level. Our experiments show that using the result of our bit-flip analysis and limiting the code instrumentation to only the susceptible locations releases 28.9% (on average) of the memory while the level of fault coverage remains the same as with full instrumentation. Ghazaleh Nazarian, Diego G. Rodrigues, Álvaro F. Moreira, Luigi Carro, Georgi Gaydadjiev |
PDP | 4 |
| 2015 | Evaluation of energy savings on a VLIW processor through dynamic issue-width adaptationabstractThe development of energy efficient hardware has been a trend in microprocessor design for the last two decades. VLIW processors are a representative example, since they have a simpler design and competitive performance, because their ILP exploitation is done statically by the compiler. In this paper, we study the energy savings that could be obtained by adapting such microarchitecture according to the current program phase. Our contribution is twofold. First, by executing a set of benchmarks on the ρ-vex configurable softcore VLIW processor, and by modifying the number of issues, we show the potentials of energy reduction. Then, with this information in hand, we developed an oracle experiment to dynamically vary the issue width of the processor according to the phase behavior, considering two different phase granularites. The potential energy savings using this policy could be as high as 81.5% when compared with the static version, executing the MiBench set. Juan Sebastian Piedrahita Giraldo, Anderson Luiz Sartor, Luigi Carro, Stephan Wong, Antonio Carlos Schneider Beck |
RSP | 3 |
| 2015 | A Runtime FPGA Placement and Routing Using Low-Complexity Graph TraversalabstractDynamic Partial Reconfiguration (DPaR) enables efficient allocation of logic resources by adding new functionalities or by sharing and/or multiplexing resources over time. Placement and routing (P&R) is one of the most time-consuming steps in the DPaR flow. P&R are two independent NP-complete problems, and, even for medium size circuits, traditional P&R algorithms are not capable of placing and routing hardware modules at runtime. We propose a novel runtime P&R algorithm for Field-Programmable Gate Array (FPGA)-based designs. Our algorithm models the FPGA as an implicit graph with a direct correspondence to the target FPGA. The P&R is performed as a graph mapping problem by exploring the node locality during a depth-first traversal. We perform the P&R using a greedy heuristic that executes in polynomial time. Unlike state-of-the-art algorithms, our approach does not try similar solutions, thus allowing the P&R to execute in milliseconds. Our algorithm is also suitable for P&R in fragmented regions. We generate results for a manufacturer-independent virtual FPGA. Compared with the most popular P&R tool running the same benchmark suite, our algorithm is up to three orders of magnitude faster. Ricardo S. Ferreira 0001, Luciana Rocha, José A. M. Nacif, Stephan Wong, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2015 | Fine-Grained Fast Field-Programmable Gate Array ScrubbingabstractField-programmable gate arrays provide several relevant advantages for critical systems, such as flexibility and high performance. However, their use in critical systems requires efficient means to mitigate transient faults in the configuration bits. This paper focuses on an alternative mechanism to reduce the repair time of traditional scrubbing approaches. It relies on fine-grained error detection and partial reconfiguration. The fine-grained information is used to dynamically choose an optimized starting position for the scrubbing procedure, reducing the mean repair time. We explore the design space provided by the technique and propose an approach to make resilient diagnosis of configuration faults. The efficiency, scalability, and robustness of the proposed mechanisms are evaluated. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Optimum design of a banked memory with power management for wireless sensor networks
Leonardo Steinfeld, Marcus Ritt, Fernando Silveira, Luigi Carro |
Wirel. Networks | 4 |
| 2014 | GPGPUs: How to combine high computational power with high reliabilityabstractGPGPUs are used increasingly in several domains, from gaming to different kinds of computationally intensive applications. In many applications GPGPU reliability is becoming a serious issue, and several research activities are focusing on its evaluation. This paper offers an overview of some major results in the area. First, it shows and analyzes the results of some experiments assessing GPGPU reliability in HPC datacenters. Second, it provides some recent results derived from radiation experiments about the reliability of GPGPUs. Third, it describes the characteristics of an advanced fault-injection environment, allowing effective evaluation of the resiliency of applications running on GPGPUs. Leonardo Arturo Bautista-Gomez, Franck Cappello, Luigi Carro, Nathan DeBardeleben, Bo Fang 0002, Sudhanva Gurumurthi, Karthik Pattabiraman, Paolo Rech, Matteo Sonza Reorda |
DATE | 3 |
| 2014 | Reliable execution of statechart-generated correct embedded software under soft errorsabstractThis paper proposes a design methodology for fault-tolerant embedded systems development that starts from software specification and goes down to hardware execution. The proposed design methodology uses formally verified and correct-by-construction software created from high-level UML statechart models for software specification and implementation. On the hardware reliability side, this paper uses the MoMa architecture for reliable embedded computing which we deploy as a soft-core onto an off-the-shelf FPGA. MoMa introduces architectural innovations that support the semantics of the UML statechart execution in a reliable fashion. The proposed design methodology is evaluated with a real automotive case study based on an exhaustive FPGA-implemented fault injection campaign. Ronaldo Rodrigues Ferreira, Thomas Klotz, Thilo Vörtler, Jean da Rolt, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro, Karsten Einwich |
DDECS | 7 |
| 2014 | Adaptive Low-Power Architecture for High-Performance and Reliable Embedded ComputingabstractThis paper presents the Matrix Operation Microprocessor Architecture (MoMa) for reliable embedded computing. MoMa introduces a software execution mechanism based on transactions, which provides a localized error correction scheme that leads to reduced error correction latency and hardware redundancy without incurring on expensive execution check pointing. Coupled to the transactional software execution is a dedicated adaptive core for matrix multiplication which is protected with a hardware implementation of the Algorithm-Based Fault Tolerance technique. MoMa drives the matrix core in an adaptive fashion based on dynamically turning it on only when high-performance computation is necessary, leading to ultimate power savings and error coverage. We performed an exhaustive FPGA-implemented fault injection campaign, in which we observed an error detection coverage of almost 100% and an error correction coverage of almost 98% on average. MoMa is also evaluated in terms of power, area, and performance, showing its competitiveness against a classical TMR solution. Ronaldo Rodrigues Ferreira, Jean da Rolt, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro |
DSN | 5 |
| 2014 | Radiation Sensitivity of High Performance Computing Applications on Kepler-Based GPGPUsabstractIn this paper we assess and discuss the radiation sensitivity of a set of HPC applications executed on NVIDIA K20 GPGPUs. The occurrence of both radiation-induced silent data corruption and functional interruption will be experimentally addressed for Hotspot, LavaMD, and Matrix Transponse. Each of the tested codes requires a proper computational power and elaborates a different amount of data. Both these characteristics play a significant role in the application radiations sensitivity. Additionally, an evaluation of the error rate at sea level will be provided for all the tested codes. Daniel Oliveira 0002, Caio B. Lunardi, Laércio Lima Pilla, Paolo Rech, Philippe Olivier Alexandre Navaux, Luigi Carro |
DSN | 6 |
| 2014 | Impact of GPUs Parallelism Management on Safety-Critical and HPC Applications ReliabilityabstractGraphics Processing Units (GPUs) offer high computational power but require high scheduling strain to manage parallel processes, which increases the GPU cross section. The results of extensive neutron radiation experiments performed on NVIDIA GPUs confirm this hypothesis. Reducing the application Degree Of Parallelism (DOP) reduces the scheduling strain but also modifies the GPU parallelism management, including memory latency, thread registers number, and the processors occupancy, which influence the sensitivity of the parallel application. An analysis on the overall GPU radiation sensitivity dependence on the code DOP is provided and the most reliable configuration is experimentally detected. Finally, modifying the parallel management affects the GPU cross section but also the code execution time and, thus, the exposure to radiation required to complete computation. The Mean Workload and Executions Between Failures metrics are introduced to evaluate the workload or the number of executions computed correctly by the GPU on a realistic application. Paolo Rech, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Luigi Carro |
DSN | 4 |
| 2014 | Reducing embedded software radiation-induced failures through cache memoriesabstractCache memories are traditionally disabled in space-level and safety-critical applications, since it was believed that the sensitive area they introduce would compromise the system reliability. As technology has evolved, the speed gap between logic and main memory has increased in such a way that disabling caches slows the code much more than in the past. As a result, the processor is exposed for a much longer time in order to compute the same workload. In this paper we demonstrate that, on modern embedded processors, enabling caches may bring benefits to critical systems: the larger exposed area may be compensated by the shorter exposure time, leading to an overall improved reliability. We describe the Mean Workload Between Failures, an intuitive metric to evaluate the impact of enabling caches for a given generic application error rate. The proposed metric is experimentally validated through an extensive radiation test campaign using a 28 nm off-the-shelf ARM-based SoC as a case study. The failure probability of the bare-metal application is decreased when the L1 cache is enabled but increased when L2 is also enabled. We also discuss when L2 caches could make the device more reliable. Thiago Santini, Paolo Rech, Gabriel L. Nazar, Luigi Carro, Flávio Rech Wagner |
ETS | 4 |
| 2014 | Fault injection in GPGPU cores to validate and debug robust parallel applicationsabstractGeneral Purpose Graphic Processing Units (GPGPUs) are more efficient than CPUs for processing parallel data. Unfortunately, GPGPUs are sensible to radiation. Hence, several software mitigation techniques, as well as robust algorithms, are being developed to overcome reliability problems. In this paper we propose a software debugger-based fault injection mechanism to evaluate the resiliency of applications running on a GPGPU and to validate the software hardening techniques it possibly embeds. We report some experimental results gathered on selected case studies to show the proposed approach advantages and limitations. M. De Carvalho, Davide Sabena, Matteo Sonza Reorda, Luca Sterpone, Paolo Rech, Luigi Carro |
IOLTS | 6 |
| 2014 | Adaptive multiple switching strategy toward an ideal NoCabstractThe exigency for heterogeneous many-core systems has brought an exponential growth in the complexity of their interconnections. In this manner, other Network-on-Chip (NoC) alternatives are being sought to attend the requirements in terms of power consumption and performance. Nevertheless, several of these proposals present very complex architectures, with virtual channels, tables and extra controls. In this paper we propose the junction of two advantageous strategies: hierarchical topology with adaptability. The use of these two techniques is novel in the literature and it allows ensuring high performance even when the application has their communication rates altered. The gains in power and in performance are possible due to the use of low cost components in a hierarchical structure. Debora Matos, Márcio Eduardo Kreutz, Cezar Reinbrecht, Luigi Carro, Altamiro Amadeu Susin |
ISCAS | 4 |
| 2014 | GPUs Neutron Sensitivity Dependence on Data Type
Paolo Rech, Christopher Frost 0002, Luigi Carro |
J. Electron. Test. | 3 |
| 2014 | Adaptive Parallelism Exploitation under Physical and Real-Time Constraints for Resilient SystemsabstractThis article introduces the resilient adaptive algebraic architecture that aims at adapting parallelism exploitation of a matrix multiplication algorithm in a time-deterministic fashion to reduce power consumption while meeting real-time deadlines present in most DSP-like applications. The proposed architecture provides low-overhead error correction capabilities relying on the hardware implementation of the algorithm-based fault-tolerance method that is executed concurrently with matrix multiplication, providing efficient occupation of memory and power resources. The Resilient Adaptive Algebraic Architecture (RA 3 ) is evaluated using three real-time industrial case studies from the telecom and multimedia application domains to present the design space exploration and the adaptation possibilities the architecture offers to hardware designers. RA 3 is compared in its performance and energy efficiency with standard high-performance architectures, namely a GPU and an out-of-order general-purpose processor. Finally, we present the results of fault injection campaigns in order to measure the architecture resilience to soft errors. Fábio P. Itturriet, Gabriel L. Nazar, Ronaldo Rodrigues Ferreira, Álvaro F. Moreira, Luigi Carro |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2013 | Scrubbing unit repositioning for fast error repair in FPGAsabstractField Programmable Gate Arrays (FPGAs) are very successful platforms that rely on large configuration memories to store the circuit functions required by users. Faults affecting such memories are a major dependability threat for these devices, and the applicability of FPGAs on critical systems depends on efficient means to mitigate their effects. The main means to effectively remove such faults, namely configuration scrubbing, consists in rewriting the desired contents of this memory and suffers from high power consumption and a long mean time to repair (MTTR). In this work we propose Scrubbing Unit Repositioning for Fast Error Repair (SURFER), a novel approach to exploit partial dynamic reconfiguration coupled with fine-grained redundancy to greatly reduce the MTTR for FPGAs subject to upsets in their configuration memories. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
CASES | 3 |
| 2013 | A transparent and energy aware reconfigurable multiprocessor platform for simultaneous ILP and TLP exploitationabstractAs the number of embedded applications increases, companies are launching new platforms within short periods of time to efficiently execute software with the lowest possible energy consumption. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of Custom Reconfigurable Arrays for Multiprocessor Systems (CReAMS). CReAMS is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. We also show that CReAMS delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered for both designs. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 3 |
| 2013 | A New Memory Banking System for Energy-Efficient Wireless Sensor NetworksabstractThe ever-increasing complexity of applications covered by wireless sensor networks (WSNs) demands for increasing memory size, which in turn increases the power drain. It is well known that SRAM power consumption can be reduced by employing a banked structure, where unused banks are switched into the low leakage retention mode. In this work, we propose a new strategy for memory banking, taking advantage of the software properties of WSN, and achieving aggressive power savings. We present a detailed model for the energy saving for equally sized banks with two power management schemes: a best-oracle policy and a simple greedy policy. Thanks to our modeling, at design time the optimum number of banks can be estimated, and the design can reach huge energy savings. The memory content allocation and the power management problem were solved by an integer linear program formulation for two real wireless sensor network application (based on TinyOS and ContikiOS). Experimental results show an energy reduction of up to 77.4% for a partition overhead of 1%. Leonardo Steinfeld, Fernando Silveira, Marcus Ritt, Luigi Carro |
DCOSS | 4 |
| 2013 | Experimental evaluation of thread distribution effects on multiple output errors in GPUsabstractGraphic Processing Units are very prone to be corrupted by neutrons. Experimental results show that in the majority of the cases a typical application like matrix multiplication is affected by multiple output errors. In this paper we evaluate how different thread distributions impact the multiple output errors occurrence. The reported results and the performed architecture analysis give practical programming advices that may increase the reliability of a generic parallel algorithm without introducing any hardware or computation overhead. Paolo Rech, Caroline Aguiar, Christopher Frost 0002, Luigi Carro |
ETS | 4 |
| 2013 | A run-time graph-based Polynomial Placement and routing algorithm for virtual FPGASabstractDynamic partial reconfiguration enables efficient use of hardware resources by multiplexing system functionality in time. However, many challenges arise from partial reconfiguration implementation. The placement and routing (P&R) of the hardware modules is a computationally intensive task, and the state-of-art algorithms are not suitable to place and route modules at run-time. This paper makes several contributions: (1) Single Placement at run-time: we propose a novel P&R algorithm based on greedy heuristic where a single placement is performed at run-time in few milliseconds. (2) Implicit Graph Model: the FPGA is modelled as an implicit graph with a direct correspondence to the physical FPGA, and the P&R is performed as a graph mapping problem by exploring the node locality during the depth-first traversal. (3) Polynomial Placement: we show that even a single placement can be routed without critical path degradation. (4) Fragmented Regions: the graph approach is flexible, and it allows efficient placement even onto fragmented FPGA areas. Compared with the most popular P&R tool running the same benchmark suite our algorithm is on average 864x faster. Moreover, the bitstream for partial reconfiguration is also reduced by a factor of 4. Ricardo S. Ferreira 0001, Luciana Rocha, José A. M. Nacif, Stephan Wong, Luigi Carro |
FPL | 6 |
| 2013 | Accelerated FPGA repair through shifted scrubbingabstractAs critical systems make more and more use of high performance FPGAs, several reliability aspects of these devices come into play. Whenever SRAM-based FPGAs are used, upsets in the configuration memory become a major dependability threat, and must be removed as soon as possible. This is usually accomplished through a process called scrubbing. The traditional scrubbing technique, however, suffers from high energy costs and a long mean time to repair (MTTR). In this work we propose a novel approach to minimize these drawbacks through a triggered shifted scrubbing procedure. The proposed technique exploits the non-uniform distribution of critical bits in the configuration memory of the device to reduce the repair time. It provides an average MTTR reduction of 30% without any changes in the circuit implemented in the FPGA when compared to previous works. Gabriel L. Nazar, Leonardo P. Santos, Luigi Carro |
FPL | 3 |
| 2013 | Algorithm transformation methods to reduce software-only fault tolerance techniques' overheadabstractThis paper introduces a framework that tackles the costs in area and energy consumed by methodologies like spatial or temporal redundancy with a different approach: given an algorithm, we find a transformation in which part of the computation involved is transformed into memory accesses. The precomputed data stored in memory can be protected then by applying traditional and well established ECC algorithms to provide fault tolerant hardware designs. At the same time, the transformation increases the performance of the system by reducing its execution time, which is then used by customized software-only fault tolerant techniques to protect the system without any degradation when compared to its original form. Application of this technique to key algorithms in a MP3 player, combined with a fault injection campaign, show that this approach increases fault tolerance up to 92%, without any performance degradation. José Rodrigo Azambuja, Gustavo Brown, Fernanda Lima Kastensmidt, Luigi Carro |
IOLTS | 4 |
| 2013 | Experimental evaluation of GPUs radiation sensitivity and algorithm-based fault tolerance efficiencyabstractExperimental results demonstrate that Graphic Processing Units are very prone to be corrupted by neutrons. We have performed several experimental campaigns at ISIS, UK and at LANSCE, Los Alamos, NM, USA accessing the sensitivity of the GPU internal resources as well as the error rate of common parallel algorithms. Experiments highlight output error patterns and radiation responses that can be fruitfully used to design optimized Algorithm-Based Fault Tolerance strategies and provide pragmatic programming guidelines to increase the code reliability with low computational overhead. Paolo Rech, Luigi Carro |
IOLTS | 2 |
| 2013 | A run-time adaptive multiprocessor systemabstractBecause of the continuous increase in the number and complexity of embedded applications, new platforms have been launched within shorter periods of time to fulfill their performance requirements with the lowest energy consumption possible. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of custom reconfigurable arrays for multiprocessor systems. The proposed approach is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. Results show that our proposal delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
ISCAS | 3 |
| 2013 | Towards a multiple-ISA embedded system
Jair Fajardo Junior, Mateus B. Rutzig, Luigi Carro, Antonio Carlos Schneider Beck |
J. Syst. Archit. | 3 |
| 2012 | Embedded reconfigurable architecturesabstractIn current-day embedded systems design, one is faced with cut-throat competition to deliver new functionalities in increasingly shorter time frames. This is now achieved by incorporating processor cores into embedded systems through (re-)programmability. However, this is not always beneficial for the performance or energy consumption. Therefore, adaptable embedded systems have been proposed to deal with these negative effects by reconfiguring the critical sections of an embedded system. In these proposals, we are clearly witnessing a trend that is moving from static configurations to dynamic (re)configurations. Stephan Wong, Luigi Carro, Stamatios Kavvadias, Georgios Keramidas, Francesco Papariello, Claudio Scordino, Roberto Giorgi, Stefanos Kaxiras |
CASES | 2 |
| 2012 | Resilient Adaptive Algebraic Architecture for Parallel Detection and Correction of Soft-ErrorsabstractA novel fault-tolerant microprocessor capable of detecting and correcting radiation-induced soft errors is proposed and evaluated. The Resilient Adaptive Algebraic Architecture performs time redundancy in parallel with matrix multiplication computation, guaranteeing on-the-fly detection and correction of errors disrupting data and logic with minimum overhead. We evaluate the RA3microprocessor in terms of performance, area, energy consumption, and fault coverage by performing an extensive design space exploration of the architecture. Finally, we also discuss how the proposed architecture can be used to support a novel hardened-by-construction HW/SW stack based on what we call single-program execution. Fábio P. Itturriet, Ronaldo Rodrigues Ferreira, Gustavo Girão, Gabriel L. Nazar, Álvaro F. Moreira, Luigi Carro |
DSD | 6 |
| 2012 | Fault-Tolerant Algebraic Architecture for radiation induced soft-errorsabstractSummary form only given. A novel fault-tolerant microprocessor capable of detecting and correcting radiation-induced soft errors is proposed and evaluated. The Fault-Tolerant Algebraic Architecture (FTAA) performs time redundancy intrinsically with computation, guaranteeing on-the-fly detection and correction of errors disrupting data and logic with minimum overhead. We evaluate the FTAA microprocessor in terms of performance, area, energy consumption, and fault coverage by performing an extensive design space exploration of the architecture. Fábio P. Itturriet, Ronaldo Rodrigues Ferreira, Luigi Carro |
ETS | 3 |
| 2012 | Fast error detection through efficient use of hardwired resources in FPGAsabstractProviding high reliability for FPGAs is a demanding task, as such devices may be subject to faults in the configuration bitstream, altering the specified function. Traditional modular redundancy remains the most used technique, due to its high fault coverage and low performance overhead. When high availability and strict real-time deadlines must be considered, however, a short mean time to repair also becomes crucial. The use of fine-grained modules can accelerate error detection, fault diagnosis and bitstream correction, but with increased area costs. In this work, we propose the use of hardwired resources found in state-of-the-art FPGAs to provide fast and area efficient fine-grained error detection. Experimental results show an average speed up in error detection of 7.68 times with only 3.2% more area overhead, when compared to coarse-grained modular redundancy. Gabriel L. Nazar, Luigi Carro |
ETS | 2 |
| 2012 | Exploiting Modified Placement and Hardwired Resources to Provide High Reliability in FPGAsabstractPossible scenarios for future manufacturing technologies increase the desirable features of fault tolerance techniques, such as coping with multiple faults and reducing error latency. On the other hand, current high-end FPGAs present, besides lookup tables and flip-flops, several dedicated components that perform the most commonly required functions. In this paper, we propose an approach to use such resources to efficiently provide fault detection capabilities. We further extend the technique with placement constraints to enhance the detection of faults affecting the routing resources, which is a critical demand for such devices. Gabriel L. Nazar, Luigi Carro |
FCCM | 2 |
| 2012 | Neutron radiation test of graphic processing unitsabstractThis paper reports and analyzes the results of neutrons radiation testing campaigns on a modern commercial-off-the-shelf Graphic Processing Unit (GPU). A set of guidelines for accelerated radiation experiments on CPUs is presented, emphasizing the shrewdness necessary to ease the test and gain meaningful data. Radiation test results are presented and discussed, highlighting the neutrons sensitivities of the different GPU memory and logic resources in terms of Failure In Time (FIT) due to neutrons at sea level. Paolo Rech, Caroline Aguiar, Ronaldo Rodrigues Ferreira, Christopher Frost 0002, Luigi Carro |
IOLTS | 5 |
| 2012 | Floorplan-aware hierarchical NoC topology with GALS interfacesabstractNetworks-on-chip has been seen as an interconnect solution for complex systems. However, performance and energy issues still represent limiting factors for Multi-Processors System-on-Chip (MPSoC). Complex router architectures can be prohibitive for the embedded domain, once they dissipate too much power and energy. In this paper we propose a low power hierarchical network topology with GALS interfaces, allowing each cluster operates in a specific frequency. The clusters are composed by crossbar devices and the number of cores allocated for each cluster is defined considering floorplan information. Experimental results show that our strategy can reduce the power dissipation in up to 58% and the latency in up to 56% for the benchmarks analyzed when compared with a packet-switched mesh network-on-chip. Debora Matos, Cezar Reinbrecht, Gianluca Palermo, Jonathan Martinelli, Altamiro Amadeu Susin, Cristina Silvano, Luigi Carro |
ISCAS | 7 |
| 2012 | ATARDS: An adaptive fault-tolerant strategy to cope with massive defects in Network-on-Chip interconnections
Anelise Kologeski, Caroline Concatto, Fernanda Lima Kastensmidt, Luigi Carro |
VLSI-SoC | 4 |
| 2011 | An FPGA-based heterogeneous coarse-grained dynamically reconfigurable architectureabstractCoarse-grained reconfigurable architecture has emerged as a promising model for embedded systems as a solution to reduce the complexity of FPGA synthesis and mapping steps, consequently reducing reconfiguration time. Despite these advantages, CGRA usage has been limited due to the lack of commercial CGRA circuits. This work proposes a virtual and dynamic CGRA implemented on top of an FPGA. This approach allows the usage of commercial-off-the-shelf FPGA devices combined with the advantages of CGRAs. The proposed architecture consists of a set of heterogeneous functional units (FU) and a global interconnection network. The global network allows any FU to be used at each cycle, which reduces significantly the placement complexity. In addition, we introduce a polynomial mapping algorithm which includes scheduling, placement and routing steps (SPR). Moreover, the proposed approach performs a very fast placement and routing in comparison to similar CGRA approaches. The three SPR steps are computed in few milliseconds. The feasibility of this approach is demonstrated for a suite of digital signal processing benchmarks. Ricardo S. Ferreira 0001, Julio C. Goldner Vendramini, Lucas Mucida, Monica Magalhães Pereira, Luigi Carro |
CASES | 5 |
| 2011 | A new reconfigurable clock-gating technique for low power SRAM-based FPGAsabstractPower consumption is dramatically increasing for Static Random Access Memory Field Programmable Gate Arrays (SRAM-FPGAs), therefore lower power FPGA circuitry and new CAD tools are needed. Clock-gating methodologies have been applied in low power FPGA designs with only minor success in reducing the total average power consumption. In this paper, we developed a new structural clock-gating technique based on internal partial reconfiguration and topological modifications. The solution is based on the dynamic partial reconfiguration of the configuration memory frames related to the clock routing resources. For a set of design cases, figures of static and dynamic power consumption were obtained. The analyses have been performed on a synchronous FIFO and on a r-VEX VLIW processor. The experimental results shown that the efficiency in the total average power consumptions ranges from about 28% to 39% with respect to standard clock-gating approaches. Besides, the proposed method is not intrusive, and presents a very limited cost in term of area overhead. Luca Sterpone, Luigi Carro, Debora Matos, Stephan Wong, F. Fakhar |
DATE | 2 |
| 2011 | Improving Reliability in NoCs by Application-Specific Mapping Combined with Adaptive Fault-Tolerant Method in the LinksabstractA strategy to handle multiple defects in the No Clinks with almost no impact on the communication delay is presented. The fault-tolerant method can guarantee the functionally of the NoC with multiple defects in any link, and with multiple faulty links. The proposed technique uses information from test phase to map the application and to configure fault-tolerant features along the NoC links. Results from an application remapped in the NoC show that the communication delay is almost unaffected, with minimal impact and overhead when compared to a fault-free system. We also show that our proposal has a variable impact in performance while traditional fault-tolerant solution like Hamming Code has a constant impact. Besides our proposal can save among 15% to 100% the energy when compared Hamming Code. Anelise Kologeski, Caroline Concatto, Luigi Carro, Fernanda Lima Kastensmidt |
ETS | 3 |
| 2011 | Matrix control-flow algorithm-based fault toleranceabstractA novel software-implemented hardware fault tolerance method based on encoding both the control and the data-flow segments of programs with matrices is proposed and evaluated. Results show an average speed-up of 3 times compared to standard duplication and comparison, with coverage higher than 95% for the case studies considered, which outperforms previous works in the field. Ronaldo Rodrigues Ferreira, Álvaro F. Moreira, Luigi Carro |
IOLTS | 3 |
| 2011 | Energy efficient pseudo-cache architecture through fine-grained reconfigurabilityabstractFueled by the exponential growth in transistors available to processor designers, cache memories became a very significant percentage of the overall area, power dissipation and energy consumption of modern systems. Instruction cache memories, however, typically hold highly redundant information in each of their columns, due to the repeated use of instructions and registers by compilers. Current memory architectures do not exploit this fact to reduce energy, consuming constant amounts of power regardless of switching activity. This work proposes the use of a fine-grained reconfigurable architecture to exploit this redundancy, providing an energy efficient on-chip storage element for embedded processors. The proposed architecture reached consumes up to 86% less energy, with an average reduction of 39%. Gabriel L. Nazar, Luigi Carro |
ISCAS | 2 |
| 2011 | Two-levels of adaptive buffer for virtual channel router in NoCsabstractNoC designs are based on a compromise of latency, power dissipation or energy, usually defined at design time. However, setting all parameters at design time can cause either excessive power dissipation (originated by router underutilization), or a higher latency. Moreover, routers with virtual channels have larger buffer sizes and more complex control, increasing the total costs. The situation worsens whenever the application changes its communication pattern, i.e., when a portable phone downloads a new service. In this paper we propose the use of a two-level adaptive buffer for a virtual channel router, where the buffers units and the virtual channels are dynamically allocated to increase router efficiency in a NoC, even under rather different communication loads. With the proposed architecture the buffer and virtual channels in the input channels of the routers can be adapted at run time. The adaptive virtual channel router decreases the latency in the worst case by 10%, and a reduction of 80% in the best case is achieved when compared to previous works. Caroline Concatto, Anelise Kologeski, Luigi Carro, Fernanda Lima Kastensmidt, Gianluca Palermo, Cristina Silvano |
VLSI-SoC | 3 |
| 2011 | Reconfigurable Routers for Low Power and High PerformanceabstractNetwork-on-chip (NoC) designs are based on a compromise among latency, power dissipation, or energy, and the balance is usually defined at design time. However, setting all parameters, such as buffer size, at design time can cause either excessive power dissipation (originated by router under utilization), or a higher latency. The situation worsens whenever the application changes its communication pattern, e.g., a portable phone downloads a new service. Large buffer sizes can ensure performance during the execution of different applications, but unfortunately, these same buffers are mainly responsible for the router total power dissipation. Another aspect is that by sizing buffers for the worst case latency incurs extra dissipation for the mean case, which is much more frequent. In this paper we propose the use of a reconfigurable router, where the buffer slots are dynamically allocated to increase router efficiency in an NoC, even under rather different communication loads. In the proposed architecture, the depth of each buffer word used in the input channels of the routers can be reconfigured at run time. The reconfigurable router allows up to 52% power savings, while maintaining the same performance as that of a homogeneous router, but using a 64% smaller buffer size. Debora Matos, Caroline Concatto, Márcio Eduardo Kreutz, Fernanda Lima Kastensmidt, Luigi Carro, Altamiro Amadeu Susin |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2010 | Challenges for embedded multicore architectureabstractIn this tutorial we discuss the impact of multicore architectures for embedded devices at different levels, ranging from heterogeneous/homogeneous ISAs to the organization and software development. Luigi Carro, Georgi Gaydadjiev |
CASES | 1 |
| 2010 | A new quaternary FPGA based on a voltage-mode multi-valued circuitabstractFPGA structures are widely used due to early time-to-market and reduced non-recurring engineering costs in comparison to ASIC designs. Interconnections play a crucial role in modern FPGAs, because they dominate delay, power and area. Multiple-valued logic allows the reduction of the number of signals in the circuit, hence can serve as a mean to effectively curtail the impact of interconnections. In this work we propose a new FPGA structure based on a low-power quaternary voltage-mode device. The most important characteristics of the proposed architecture are the reduced fanout, low number of wires and switches, and the small wire length. We use a set of FIR filters as a demonstrator of the benefits of the quaternary representation in FPGAs. Results show a significant reduction on power consumption with small timing penalties. Cristiano Lazzari, Paulo F. Flores, José Monteiro 0001, Luigi Carro |
DATE | 4 |
| 2010 | System Level Hardening by Computing with MatricesabstractContinuous advances in transistor manufacturing have enabled technology scaling along the years, sustaining Moore's law. As transistors sizes rapidly shrink, and voltage scales, the amount of charge in a node also rapidly decreases. A particle hitting the core will probably cause a transient fault to spam over several clock cycles. In this scenario, embedded systems using state-of-the-art technologies will face the challenge of operating in an environment susceptible to multiple errors, but with restricted resources available to deploy fault-tolerance, as these techniques severely increase power consumption. One possible solution to this problem is the adoption of software based fault-tolerance at the system level, aiming at reduced energy levels to ensure reliability and low energy dissipation. In this paper, we claim the detection and correction of errors on generic data structures at system level by using matrices to encode any program and algorithm. With such encoding, it is possible to employ established techniques of detection and correction of errors occurring in matrices, running with inexpressive overhead of power and energy. We evaluated this proposal using two case studies significant for the embedded system domain. Using the proposed approach, we observed in some cases an overhead of only 5% in performance and 8% in program size. Ronaldo Rodrigues Ferreira, Álvaro F. Moreira, Luigi Carro |
DSD | 3 |
| 2010 | Multiple Bit Error Detection and Correction in MemoryabstractTechnology evolution provides ever increasing density of transistors in chips, lower power consumption and higher performance. In this environment the occurrence of multiple-bit upsets (MBUs) becomes a significant concern. Critical applications need high reliability, but traditional error mitigation techniques assume only the single error model, and only a few techniques to correct MBUs at algorithm level have been proposed. In this paper, a novel circuit level technique to detect and correct multiple errors in memory is proposed. Since it is implemented at circuit level, it is transparent to programmers. This technique is based in the Decimal Hamming coding and here it is compared to Reed Solomon coding at circuit level. Experimental results show that for memory words wider than 16 bits, the proposed technique is faster and imposes lower area overhead than optimized RS, while mitigating errors affecting up to 25% of the memory word. J. F. Tarillo, Nikolaos Mavrogiannakis, Carlos Arthur Lang Lisbôa, Costas Argyrides, Luigi Carro |
DSD | 5 |
| 2010 | A Cost-Effective Technique for Mapping BLUTs to QLUTs in FPGAsabstractQuaternary logic has shown to be a promising alternative for implementing FPGAs, since voltage mode quaternary circuits can reduce the circuits' cost and at the same time reduce its power consumption. In this paper, we study the implementation of circuits in quaternary logic. To obtain cost-effective implementations of quaternary circuits, we propose a mapping from binary to quaternary circuits based on integer linear programming. Our results show that the expected improvements can be achieved, reducing, in average, the number of transistors by 27% and the number of nets by 19%, compared to a binary implementation. Marcus Ritt, Carlos Arthur Lang Lisbôa, Luigi Carro, Cristiano Lazzari |
FPL | 3 |
| 2010 | Voltage-mode quaternary FPGAs: An evaluation of interconnectionsabstractThis work presents a study about FPGA interconnections and evaluates their effects on voltage-mode binary and quaternary FPGA structures. FPGAs are widely used due to the fast time-to-market and reduced non-recurring engineering costs in comparison to ASIC designs. Interconnections play a crucial role in modern FPGAs, because they dominate delay, power and area. The use of multiple-valued logic allows the reduction of the number of signals in the circuit, hence providing a mean to effectively curtail the impact of interconnections. The most important characteristic of the results are the reduced fanout, fewer number of wires and the smaller wire length presented by the quaternary devices. We use a set of arithmetic circuits to compare binary and quaternary implementations. This work presents the first step on developing quaternary circuits by mapping any binary random logic onto quaternary devices. Cristiano Lazzari, Paulo F. Flores, José Monteiro 0001, Luigi Carro |
ISCAS | 4 |
| 2010 | Associating packets of heterogeneous cores using a synchronizer wrapper for NoCsabstractMPSoCs systems are composed of heterogeneous cores, and for this reason, the cores can present different bandwidth, different clock domains or still they can require an irregular traffic behavior. When networks-on-chip (NoCs) are used to connect these cores, one very often needs some synchronization solution, and due to the mentioned problems, this might be required for synchronous or asynchronous NOCs. In this paper we show a network interface (NI) with a synchronizer wrapper solution. We verified its applicability for different channel widths and buffer depths of a NoC. These network interfaces were used to connect a H.264 decoder and the simulation results demonstrate that the wrapper provides a reliable synchronization solution, and does not compromise the latency of the network. These interfaces have been successfully implemented in a 0.18um CMOS technology. Debora Matos, Luigi Carro, Altamiro Amadeu Susin |
ISCAS | 2 |
| 2010 | A broad strategy to detect crosstalk faults in network-on-chip interconnectsabstractIn this paper, we propose a method to detect crosstalk faults within and among channels of mesh NoCs, using a global test strategy that is based on a particular set of test paths and test packet. All test paths must be activated simultaneously without resource conflict. The test packet is built using Maximal Aggressor Fault (MAF) vectors. The test strategy is capable of detecting 100% of the considered faults. The test application time grows quadratically with the NoC increase, but can be drastically improved by means of two alternative approaches also proposed in the paper. Those local approaches are based on simultaneously testing multiple victims or NoC regions unlikely to aggress each other. The test time then grows linearly and shows a very modest derivative. Mariza Botelho, Fernanda Lima Kastensmidt, Marcelo Lubaszewski, Érika F. Cota, Luigi Carro |
VLSI-SoC | 5 |
| 2010 | Network interface to synchronize multiple packets on NoC-based Systems-on-ChipabstractThe cores of a System-on-Chip (SoC) connected by Networks-on-Chip (NoCs) need interfaces to properly send and receive packets. However, in this interfacing, different situations can occur when heterogeneous cores are applied. Applications may require, for example, an irregular traffic behavior or present a large bandwidth variation. These situations may lead to problems in data synchronization. In this paper we show a simple and efficient synchronization solution, which although known in the literature, has not yet been applied to NoC-based systems scenario. Using a network interface as the synchronization mechanism, the proposed circuit handles data dependencies instead of letting each core solve the synchronization problems at higher levels. As case study, an H.264 video decoder was used to show the need and advantage of our approach. The proposed design is FIFO-based and can be applied when multiple packets from different sources need to be synchronized in a single destination. Simulations were performed to verify the functionality and efficiency of the synchronization solution. These interfaces were implemented in VHDL and synthesized using an 180 nm CMOS technology. Debora Matos, Miklecio Costa, Luigi Carro, Altamiro Amadeu Susin |
VLSI-SoC | 3 |
| 2009 | A fast error correction technique for matrix multiplication algorithmsabstractTemporal redundancy techniques will no longer be able to cope with radiation induced soft errors in technologies beyond the 45 nm node, because transients will last longer than the cycle time of circuits. The use of spatial redundancy techniques will also be precluded, due to their intrinsic high power and area overheads. The use of algorithm level techniques to detect and correct errors with low cost has been proposed in previous works, using a matrix multiplication algorithm as the case study. In this paper, a new approach to deal with this problem is proposed, in which the time required to recompute the erroneous element when an error is detected is minimized. Costas Argyrides, Carlos Arthur Lang Lisbôa, Dhiraj K. Pradhan, Luigi Carro |
IOLTS | 4 |
| 2009 | Invariant checkers: An efficient low cost technique for run-time transient errors detectionabstractSemiconductor technology evolution brings along higher soft error rates and long duration transients, which require new low cost system level approaches for error detection and mitigation. Known software based error detection techniques imply a high overhead in terms of memory usage and execution times. In this work, the use of software invariants as a means to detect transient errors affecting a system at run-time is proposed. The technique is based on the use of a publicly available tool to automate the invariant detection process, and the decomposition of complex algorithms into simpler ones, which are checked through the verification of their invariants during the execution of the program. A sample program is used as a case study, and fault injection campaigns are performed to verify the error detection capability of the proposed technique. The experimental results show that the proposed technique provides high error detection capability, with low execution time overhead. Carmela Noro Grando, Carlos Arthur Lang Lisbôa, Álvaro F. Moreira, Luigi Carro |
IOLTS | 4 |
| 2009 | A low cost and adaptable routing network for reconfigurable systemsabstractNowadays, scalability, parallelism and fault-tolerance are key features to take advantage of last silicon technology advances, and that is why reconfigurable architectures are in the spotlight. However, one of the major problems in designing reconfigurable and parallel processing elements concerns the design of a cost-effective interconnection network. This way, considering that Multistage Interconnection Network (MIN) has been successfully used in several computer system levels and applications in the past, in this work we propose the use of a MIN, at the word level, on a coarse-grained reconfigurable architecture. More precisely, this work presents a novel parallel self-placement and routing mechanism for MIN on the circuit-switching mode. We take into account one-to-one as well as multicast (one-to-many) permutations. Our approach is scalable and it is targeted to be used in run-time environments where dynamic routing among functional units is required. In addition, our algorithm is embedded in the switch structure, and it is independent of the interstage interconnection pattern. Our approach can handle blocking and non-blocking networks, symmetrical or asymmetrical topologies. As case study, we use the proposed technique in a dynamic reconfigurable system, showing a major area reduction of 30% without performance overhead. Ricardo S. Ferreira 0001, Marcone Laure, Antonio Carlos Schneider Beck, Thiago Lo, Mateus B. Rutzig, Luigi Carro |
IPDPS | 6 |
| 2009 | Simulink®-based heterogeneous multiprocessor SoC design flow for mixed hardware/software refinement and simulation
Sangil Han, Soo-Ik Chae, Lisane B. de Brisolara, Luigi Carro, Katalin Popovici, Xavier Guerin, Ahmed Amine Jerraya, Kai Huang 0002, Xiaolang Yan |
Integr. | 4 |
| 2008 | Transparent Reconfigurable Acceleration for Heterogeneous Embedded ApplicationsabstractEmbedded systems are becoming increasingly complex. Besides the additional processing capabilities, they are characterized by high diversity of computational models coexisting in a single device. Although reconfigurable architectures have already shown to be a potential solution for such systems, they just present significant speedups of very specific dataflow oriented kernels. Furthermore, reconfigurable fabric is still withheld by the need of special tools and compilers, clearly not sustaining backward software compatibility. In this paper, we propose a new technique to optimize both dataflow and control-flow oriented code in a totally transparent process, without the need of any modification in the source or binary codes. For that, we have developed a Binary Translation algorithm implemented in hardware, which works in parallel to a MIPS processor. The proposed mechanism is responsible for transforming sequences of instructions at runtime to be executed on a dynamic coarse-grain reconfigurable array, supporting speculative execution. Executing the MIBench suite, we show performance improvements of up to 2.5 times, while reducing 1.7 times the required energy, using trivial hardware resources. Antonio Carlos Schneider Beck, Mateus B. Rutzig, Georgi Gaydadjiev, Luigi Carro |
DATE | 4 |
| 2008 | Using UML as Front-end for Heterogeneous Software Code Generation StrategiesabstractIn this paper we propose an embedded software design flow, which starts from an UML model and provides automatic mapping to other models like Simulink or finite-state machines (FSM). An automatic synthesis of an executable and synthesizable Simulink model is also proposed, enabling the use of UML as front-end for a multi-model design strategy that includes a Simulink-based MPSoC target design flow. In addition, the proposed synthesis tool automatically handles processor allocation, mapping of threads to processors, and insertion of required Simulink temporal barriers, ports, and dataflow connections. Following this approach, the UML model is mapped to the more appropriated model and specialized code generators are used. Therefore, this approach allows designers to employ UML to model the whole system and reuse this model to generate code using different strategies and targeting different platforms. Lisane B. de Brisolara, Marcio Ferreira da Silva Oliveira, Ricardo Miotto Redin, Luís C. Lamb, Luigi Carro, Flávio Rech Wagner |
DATE | 5 |
| 2008 | Reducing interconnection cost in coarse-grained dynamic computing through multistage networkabstractCoarse-grained reconfigurable architectures appear as a scalable solution to embedded system design, with a reduced reconfiguration time, memory footprint, as well as placement and routing complexity. To ensure high performance, data must be efficiently delivered to the reconfigurable matrix. For that, several architectures propose the use of fully interconnected local networks, as crossbar or large multiplexers. However, these interconnections are very area consuming. Therefore, in order to reduce the interconnection complexity without losing performance, this work proposes to use Multistage Interconnection Networks. As a case study, we have implemented the proposed approach in a tightly coupled reconfigurable array, which works together with a MIPS processor. Simulation results over the Mibench Benchmark set show savings of up to 26% of the total area, with a decrease of only 1% on the average performance. Ricardo S. Ferreira 0001, Marcone Laure, Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
FPL | 5 |
| 2008 | Balancing reconfigurable data path resources according to application requirementsabstractProcessor architectures are changing mainly due to the excessive power dissipation and the future break of Moore's law. Thus, new alternatives are necessary to sustain the performance increase of the processors, while still allowing low energy computations. Reconfigurable systems are strongly emerging as one of these solutions. However, because they are very area consuming and deal with a large number of applications with diverse behaviors, new tools must be developed to automatically handle this new problem. This way, in this work we present a tool aimed to balance the reconfigurable area occupied with the performance required by a given application, calculating the exact size and shape of a reconfigurable data path. Using as case study a tightly coupled reconfigurable array and the Mibench Benchmark set, we show that the solution found by the proposed tool saves four times area in comparison with the non-optimized version of the reconfigurable logic, with a decrease of only 5.8% on average of its original performance. This way, we open new applications for reconfigurable devices as low cost accelerators. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
IPDPS | 3 |
| 2008 | Binary translation process to optimize nanowire arrays usageabstractThe last years of semiconductor research have allowed the construction of nanowires. These atomic structures promise to have more than one order of magnitude higher density when compared to 22 nm CMOS, with less power dissipation, but unfortunately with much lower switching speed. Although there are several works that deal with the problems related to specific fabrication issues, the best use of these structures from a design perspective is still an open field of research. The design solution should take into account available parallelism (to cope with these slower than CMOS devices) and reliability. Moreover, one has to tackle the software compatibility problem, in the sense that nanowire circuits are being considered as accelerators, and not a CMOS replacement. In this paper we propose the use of a nanowire array together with a binary translation mechanism that allows the coupling of the nanowire array to a regular microprocessor, and we show how can one expect high performance and low dissipation, while still considering the intrinsic reliability issue of nanowires. Eduardo Luis Rhod, Mateus B. Rutzig, Luigi Carro |
ISCAS | 3 |
| 2008 | Algorithm Level Fault Tolerance: A Technique to Cope with Long Duration Transient Faults in Matrix Multiplication AlgorithmsabstractFor technologies beyond the 45 nm node, radiation induced transients will last longer than one clock cycle. In this scenario, temporal redundancy techniques will no longer be able to cope with radiation induced soft errors, while spatial redundancy techniques still impose high power and area overheads. The solution to this impasse is the use of algorithm level techniques, able to detect and correct errors with low cost. In this paper, a new approach to deal with this problem is proposed, and applied to matrix multiplication algorithm. The proposed technique is compared to previously published fault tolerance techniques, and the costs of detection and recomputation for both approaches are compared and discussed. Carlos Arthur Lang Lisbôa, Costas Argyrides, Dhiraj K. Pradhan, Luigi Carro |
VTS | 4 |
| 2008 | Majority Logic Mapping for Soft Error Dependability
Lorenzo Petroli, Carlos Arthur Lang Lisbôa, Fernanda Lima Kastensmidt, Luigi Carro |
J. Electron. Test. | 4 |
| 2008 | Hardware and Software Transparency in the Protection of Programs Against SEUs and SETs
Eduardo Luis Rhod, Carlos Arthur Lang Lisbôa, Luigi Carro, Matteo Sonza Reorda, Massimo Violante |
J. Electron. Test. | 3 |
| 2007 | Simulink-Based MPSoC Design Flow: Case Study of Motion-JPEG and H.264abstractSystem-level design methodologies have been introduced as a solution to handle the design complexity of embedded multiprocessor SoC (MPSoC) systems. In this paper we describe a system-level design flow starting from Simulink specification, focusing on concurrent hardware and software design and verification at four different abstraction levels: Simulink Combined Algorithm and Architecture Model (CAAM), Virtual Architecture, Transaction-accurate Model and Virtual Prototype. We used two multimedia applications, Motion-JPEG and H.264, to evaluate this design flow. Experimental results show that our design flow can generate various MPSoC architectures from Simulink CAAM correctly and efficiently, allowing processor and task design space exploration at different abstraction levels. Kai Huang 0002, Sangil Han, Katalin Popovici, Lisane B. de Brisolara, Xavier Guerin, Xiaolang Yan, Soo-Ik Chae, Luigi Carro, Ahmed Amine Jerraya |
DAC | 9 |
| 2007 | A low-SER efficient core processor architecture for future technologies
Eduardo Luis Rhod, Carlos Arthur Lang Lisbôa, Luigi Carro |
DATE | 3 |
| 2007 | System Level Approaches for Mitigation of Long Duration Transient Faults in Future TechnologiesabstractThe evolution of the technology in search of smaller and faster devices brings along the need for a new paradigm in the design of circuits tolerant to soft errors. The current assumption of transient pulses shorter than the cycle time of the circuit will no longer be true, thereby precluding the use of most of the mitigation techniques proposed so far. With transient faults duration spanning more than one clock cycle of operation, new fault tolerance solutions, working at the system level, with low area and performance overheads, must be devised. In this paper we propose the first steps in the direction of using low cost verification schemes at the algorithmic level, applied to general purpose matrix multiplication applications. Experimental results obtained with two different implementations of checker circuits using the proposed technique are presented and discussed. Carlos Arthur Lang Lisbôa, Marcelo Ienczczak Erigson, Luigi Carro |
ETS | 3 |
| 2007 | Digital Generation of Signals for Low Cost RF BISTabstractRF test signals are a requirement for the implementation of effective BIST techniques in transceivers. In this work a method to encode a binary signal with the desired RF frequency is presented. The approach employs high-pass sigma delta modulators, in contrast to conventional low- pass or band-pass approaches, allowing signal generation close to the Nyquist limit of FS/2 (FS=sampling frequency). As a digital signal is used, only a 1-bit DAC is needed, reducing test costs. Practical results using a 3Gbps transceiver illustrate the performance achievable by the method. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
ETS | 2 |
| 2007 | A Digitally Testable Capacitance-Insensitive Mixed-Signal FilterabstractOne of the main problems when developing analog filters in VLSI is to achieve high accuracy regarding the cutoff frequency. This is mainly due to the difficulty in obtaining accurate time constants. Testing of such filters is also challenging, in the sense that special equipment is required. Small deviations in the resistor or capacitor values may lead to a very high mismatch between the expected and the achieved cutoff frequency. Although switched-capacitor or active-transistor techniques may produce good results, the cost to use such approaches becomes another limiting factor, and only increases the tester needs. In this work, we present the development of an analog FIR filter, which does not use passive components to tune the cutoff frequency or the quality factor. Instead, the filter coefficients and the input signal are represented in a bit stream fashion, and are digitally processed, thus avoiding the use of expensive analog-to- digital converters. The impact of this filter architecture on test cost and possible design-for-test techniques are discussed in this paper. Erik Schüler, Marcelo Negreiros, Pascal Nouet, Luigi Carro |
ETS | 4 |
| 2007 | Using built-in sensors to cope with long duration transient faults in future technologiesabstractTransients spanning more than one clock cycle will challenge soft error tolerant designs for future technologies. To face this problem, a low overhead technique that uses bulk built-in current sensors and recomputation is proposed here. Carlos Arthur Lang Lisbôa, Fernanda Lima Kastensmidt, Egas Henes Neto, Gilson I. Wirth, Luigi Carro |
ITC | 5 |
| 2007 | Reducing fine-grain communication overhead in multithread code generation for heterogeneous MPSoCabstractHeterogeneous MPSoCs present unique opportunities for emerging embedded applications, which require both high-performance and programmability. Although, software programming for these MPSoC architectures requires tedious and error-prone tasks, thereby automatic code generation tools are required. A code generation method based on fine-grain specification can provide more design space and optimization opportunities, such as exploiting fine-level parallelism and more efficient partitions. However, when partitioned, fine-grain models may require a large number of inter-processor communications, decreasing the overall system performance. This paper presents a Simulink-based multithread code generation method, which applies Message Aggregation optimization technique to reduce the number of inter-processor communications. This technique reduces the communication overheads in terms of execution time by reduction on the number of messages exchanged and in terms of memory size by the reduction on the number of channels. The paper also presents experiment results for one multimedia application, showing performance improvements and memory reduction obtained with Message Aggregation technique. Lisane B. de Brisolara, Sangil Han, Xavier Guerin, Luigi Carro, Ricardo Augusto da Luz Reis, Soo-Ik Chae, Ahmed Amine Jerraya |
SCOPES | 4 |
| 2007 | Transparent acceleration of data dependent instructions for general purpose processorsabstractAlthough transistor scaling keeps following Moore’s law, and more area is available for designers, the clock frequency and ILP rate do not present the same level of growth anymore. This way, new architectural alternatives are necessary. Reconfigurable fabric appears to be one emerging possibility: besides exploiting the parallelism among instructions, it can also accelerate sequences of data dependent ones. However, coarse grain reconfiguration wide spread usage is still withhold by the need of special tools and compilers, which clearly do not sustain the reuse of legacy code without any modification. Based on all these facts, this work proposes a new Binary Translation algorithm, implemented in hardware and working in parallel to the processor, responsible for transforming sequences of instructions at run-time to be executed on a dynamic coarse-grain reconfigurable array, tightly coupled to a traditional RISC machine. Therefore, we can take advantage of using pure combinational logic to optimize even control-flow oriented code in a totally transparent process, without any modification in the source or binary codes. Using the Simplescalar Toolset together with the embedded benchmark suite MIBench, we show performance improvements and area evaluation when comparing against traditional superscalar architectures. Antonio Carlos Schneider Beck, Luigi Carro |
VLSI-SoC | 2 |
| 2007 | RF Digital Signal Generation Beyond NyquistabstractThis paper discusses low cost RF signal generation for BIST, using only digital circuits. One major problem is the range of frequencies that can be achieved by any digital signal generator, since the Nyquist limit is the clock frequency divided by 2 (FS/2). The proposed method uses the images of the signal in order to reach frequencies beyond FS/2. Different digital signal generation techniques for RF and their limitations are addressed as well as their impact in the framework of analog RF BIST. Experimental results and comparisons are provided using a gigabit transceiver Marcelo Negreiros, Adão Antônio de Souza Jr., Luigi Carro, Altamiro Amadeu Susin |
VTS | 3 |
| 2007 | Reducing Test Time Using an Enhanced RF Loopback
Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
J. Electron. Test. | 2 |
| 2007 | Functionally Fault-tolerant DSP Microprocessor using Sigma-delta Modulated Signals
Erik Schüler, Marcelo Ienczczak Erigson, Luigi Carro |
J. Electron. Test. | 3 |
| 2007 | Evaluating Different Solutions to Design Fault Tolerant Systems with SRAM-based FPGAs
Luca Sterpone, Matteo Sonza Reorda, Massimo Violante, Fernanda Lima Kastensmidt, Luigi Carro |
J. Electron. Test. | 5 |
| 2006 | The Molen FemtoJava EngineabstractThis paper presents the Molen FemtoJava engine that is extended with concepts taken from the Molen polymorphic processor. This allows for the existing FemtoJava to be augmented with reconfigurable hardware with only a single extension of the bytecodes and thereby but still allowing the implementation of arbitrary hardware implementations. Therefore, computationally intensive functions can be moved to the reconfigurable hardware to improve their performance. Our experimental results on MP3 decoding, which is a common embedded application, can be improved by at least of 27% reduction of execution cycles with minimal additional hardware area (about 7%). Finally, our synthesis results also show that the Molen extension of the FemtoJava engine only required an additional 10% of area (in terms of FPGA slices). Júlio C. B. de Mattos, Stephan Wong, Luigi Carro |
ASAP | 3 |
| 2006 | An improved RF loopback for test time reductionabstractIn this work a method to improve the loopback test used in RF analog circuits is described. The approach is targeted to the SoC environment, being able to reuse system resources in order to minimize the test overhead. An RF sampler is used to observe spectral characteristics of the RF signal path during loopback operation. While able to improve the observability of the signal path, the method also allows faster diagnosis than conventional loopback tests, as the number of transmitted symbols can be greatly reduced. Practical results for a prototyped RF link at 860MHz are presented in order to demonstrate the relevance of the method Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
DATE | 2 |
| 2006 | Evaluating Sigma-Delta Modulated Signals to Develop Fault-Tolerant CircuitsabstractAs microelectronics evolves smaller into the nanometric scale, external interferences starts to be harmful to the system expected behavior. As classical systems do not handle adequately faults caused by such sources, new topologies are proposed. Our present work proposes a solution for this problem consisting on the use of sigma-delta modulation in order to obtain a fault-tolerance even in presence of multiple faults. This paper provides the mathematical analysis and demonstration to support the proposed approach Erik Schüler, Daniel S. Farenzena, Luigi Carro |
ETS | 3 |
| 2006 | Evaluating SEU and Crosstalk Effects in Network-on-Chip RoutersabstractThis work intends to evaluate the effect of a single event upsets (SEUs) and crosstalk faults in a NoC router architecture by developing a fault injection mechanism, allowing an accurate analysis of the impact of SEU and crosstalk over the router service. Results show that such faults may affect the router behavior, causing loss of packets, errors in packet information or even compromising the router service, provoking permanent routing problems Arthur Pereira Frantz, Luigi Carro, Érika F. Cota, Fernanda Lima Kastensmidt |
IOLTS | 2 |
| 2006 | Reconfiguration of embedded Java applicationsabstractThis work presents the development of a coarse grain reconfigurable unit to be coupled to a native Java microcontroller, which is designed for an optimized execution of the embedded application. Code fragments to be accelerated through this unit are identified by profiling the application. The unit is able to explore ILP in a simple way and allows for Java compatibility, while also reducing the number of executed instructions, thus improving the performance with simultaneous energy savings. In many cases, as demonstrated by experiments, it also allows for smaller power consumption. João Cláudio Soares Otero, Flávio Rech Wagner, Luigi Carro |
IPDPS | 3 |
| 2006 | Increasing analog programmability in SoCsabstractThe use of programmability in systems-on-chip (SoC) brings as the main advantage the possibility of reducing the time-to-market and the cost of design, specially when different systems and functions must cover different markets, going from low-power and low-frequency instrumentation to high frequency communication. This paper presents a technique that can be used to increase the analog programmability in a SoC, also allowing one to integrate more analog functions, while guaranteeing the use of the analog part in a larger range of applications. Practical results are presented showing that the proposed technique can be used from DC to RF applications Erik Schüler, Luigi Carro |
IPDPS | 2 |
| 2006 | Reconfigurable communications for image processing applicationsabstractThis work tries to reuse programmable communication resources like a network-on-chip (NoC) in the acceleration of image applications. We show a mathematical model for the computation and communication pattern of two distributed motion estimation algorithms, full search block matching algorithm and multi-resolution block matching algorithm. Experimental results show that the use of the multi-resolution method reduces not only the computation time but also the traffic of messages on the NoC. This leads to a lower power consumption in the NoC during the processing time of each image. The studied examples show the importance of the link between algorithms and their mapping onto a programmable fabric, not only regarding computation, but facing communication as well André Borin Soares, Luigi Carro, Altamiro Amadeu Susin |
IPDPS | 2 |
| 2006 | Reconfigurable analog interface for mixed signal SOCabstractThis work discusses the development, modeling and implementation results of a programmable architecture to be employed for analog signals interfacing in mixed-signal SOC. The proposed architecture is able to achieve wide frequency range, covering a large range of applications with constant performance, allied to digital configuration compatibility. The proposed approach utilizes the concept of frequency translation (mixing) and SigmaDelta modulation leading to a fairly constant analog block for an input signal uniform treatment from DC to high frequencies. The interface performance theoretical model is addressed for supporting the design space exploration and also the physical design. An interface prototype using a fourth order continuous time band-pass SigmaDelta modulator is built and characterized validating the proposed performance model. The usage of this interface as a multi-band parametric ADC is presented Eric E. Fabris, Luigi Carro, Sergio Bampi |
ISCAS | 2 |
| 2006 | Dependable Network-on-Chip Router Able to Simultaneously Tolerate Soft Errors and CrosstalkabstractAs the technology scales down into deep sub-micron domain, more IP cores are integrated in the same die and new communication architectures are used to meet performance and power constraints. However, the same technologic advance makes devices and interconnects more sensitive to new types of malfunctions and failures, such as crosstalk and transient faults. This paper proposes fault tolerant techniques to protect NoC routers against the occurrence of soft errors and crosstalk at the same time, with minimum area and performance overhead. Experimental results show that a cost-effective protection alternative can be achieved by the combination of error correction codes and time redundancy techniques Arthur Pereira Frantz, Fernanda Lima Kastensmidt, Luigi Carro, Érika F. Cota |
ITC | 3 |
| 2006 | Automatic Dataflow Execution with Reconfiguration and Dynamic Instruction MergingabstractAs Moore's law is loosing steam, one already sees the phenomenon of clock frequency reduction caused by the excessive power dissipation. New technologies that will completely or partially replace silicon are arising, and new architectural alternatives are necessary. Reconfigurable fabric appears to be one of these solutions, and has shown speed ups of critical parts of several data stream programs. However, the wide spread use of reconfigurable computing is still withhold by the need of special tools and compilers, which clearly preclude software portability and reuse of legacy code. Based on all these facts, this work proposes a coarse-grain dynamic reconfigurable array, tightly coupled to a traditional RISC machine. Besides taking advantage of using combinational logic to speed up the execution, dynamic analysis of the code at run time was implemented to reconfigure the array, maintaining full software compatibility. Using the Simplescalar Toolset together with the embedded benchmark suite MIBench, meaningful performance improvements (up to 3 times of speed up) were shown, thanks to the implementation of the proposed approach Antonio Carlos Schneider Beck, Victor F. Gomes, Luigi Carro |
VLSI-SoC | 3 |
| 2006 | A low power high performance CMOS voltage-mode quaternary full adderabstractMultiple-valued logic, despite of all its theoretical potentialities, has not provided real advantages for arithmetic circuits when compared to the binary equivalent ones until now. This paper shows a new efficient method to implement quaternary logic arithmetic circuits using multi-threshold transistors, where 3 power supply lines are used to perform quaternary circuits with low power consumption and high performance. As a demonstration, a quaternary full adder is described in TSMC 0.18μm technology and compared to regular binary circuits, presenting a 76% reduction in power consumption, and an improvement of 15% regarding speed with a 20% area overhead. Ricardo C. Goncalves da Silva, Henri Boudinov, Luigi Carro |
VLSI-SoC | 3 |
| 2005 | Comparing high-level modeling approaches for embedded system designabstractThis paper present a comparison between three different high-level modeling approaches for embedded systems design, focusing on systems that require dataflow models. The proposed evaluation investigates the facilities provided by these approaches for expressing systems requirements, functional specification, and timing constraints. Properties like model readability, testability, and implementability are also considered. Moreover, the support to different Models of Computation is also evaluated. A Crane Control System is used as case study to apply the proposed comparison criteria. Lisane B. de Brisolara, Leandro Buss Becker, Luigi Carro, Flávio Rech Wagner, Carlos Eduardo Pereira, Ricardo Augusto da Luz Reis |
ASP-DAC | 3 |
| 2005 | Time and energy efficient mapping of embedded applications onto NoCsabstractThis work analyzes, the mapping of applications onto generic regular Networks-on-Chip (NoCs). Cores must be placed considering communication requirements so as to minimize the overall application execution time and energy consumption. We expand previous mapping strategies by taking into consideration the dynamic behavior of the target application and thus potential contentions in the intercommunication of the cores. Experimental results for a suite of 22 benchmarks and various NoC sizes show that a 42% average reduction in the execution time of the mapped application can be obtained, together with a 21% average reduction in the total energy consumption for state-of-the-art technologies. César A. M. Marcon, André Borin Soares, Altamiro Amadeu Susin, Luigi Carro, Flávio Rech Wagner |
ASP-DAC | 4 |
| 2005 | Dynamic reconfiguration with binary translation: breaking the ILP barrier with software compatibilityabstractIn this paper we present the impact of dynamically translating any sequence of instructions into combinational logic. The proposed approach combines a reconfigurable architecture with a binary translation mechanism, being totally transparent for the software designer. Besides ensuring software compatibility, the technique allows porting the same code for different machines tracking technological evolutions. The target processor is a Java machine able to execute Java bytecodes. Experimental results show that even code without any available parallelism can benefit from the proposed approach. Algorithms used in the embedded systems domain were accelerated 4.6 times in the mean, while spending 10.89 times less energy in the average. We present results regarding the impact of area and power, and compare the proposed approach with other Java machines, including a VLIW one. Antonio Carlos Schneider Beck, Luigi Carro |
DAC | 2 |
| 2005 | On the Optimal Design of Triple Modular Redundancy Logic for SRAM-based FPGAsabstractTriple modular redundancy (TMR) is a suitable fault tolerant technique for SRAM-based FPGA. However, one of the main challenges in achieving 100% robustness in designs protected by TMR running on programmable platforms is to prevent upsets in the routing from provoking undesirable connections between signals from distinct redundant logic parts, which can generate an error in the output. This paper investigates the optimal design of the TMR logic (e.g., by cleverly inserting voters) to ensure robustness. Four different versions of a TMR digital filter were analyzed by fault injection. Faults were randomly inserted straight into the bitstream of the FPGA. The experimental results presented in this paper demonstrate that the number and placement of voters in the TMR design can directly affect the fault tolerance, ranging from 4.03% to 0.98% the number of upsets in the routing able to cause an error in the TMR circuit. Fernanda Lima Kastensmidt, Luca Sterpone, Luigi Carro, Matteo Sonza Reorda |
DATE | 3 |
| 2005 | Noise Figure Evaluation Using Low Cost BISTabstractA technique for noise figure evaluation, suitable for BIST implementation, is described. It is based on a low cost single-bit digitizer, which allows the simultaneous evaluation of the noise figure in several test points of the analog circuit. The method is also able to benefit from SoC resources, like memory and processing power. The theoretical background and experimental results are presented in order to demonstrate the feasibility of the approach. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
DATE | 2 |
| 2005 | Increasing Fault Tolerance to Multiple Upsets Using Digital Sigma-Delta ModulatorsabstractAs the transistor gate length goes straightforward to the sub-micron dimension, there is an increased possibility of occurrence of external interferences in these devices. The direct effect of such external and/or intrinsic interferences is, in many cases, the total mismatch between the desired answer of the system and the obtained response. So, new techniques must be studied in order to guarantee the correct operation of these systems, under multiple simultaneous faults. This work presents the use of a totally digital sigma-delta modulator that is used to develop arithmetic operations with much better results than if a common digital operator was used. Simulation results show that, even under multiple simultaneous faults, the system presents very good results, as in the addition case, where a maximum standard deviation of 0.7 is achieved for sigma-delta-modulated signals, while for the digital adder alone, this value is 57.5. Such behavior is good enough to be used in operators that tolerate small errors, like in the digital filters where these errors are embedded in the system noise. Erik Schüler, Luigi Carro |
IOLTS | 2 |
| 2005 | Low Cost BIST for Static and Dynamic Testing of ADCs
Maria Da Gloria Flores, Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin, Felipe R. Clayton, Cristiano Benevento |
J. Electron. Test. | 3 |
| 2005 | Low Cost On-Line Testing Strategy for RF Circuits
Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
J. Electron. Test. | 2 |
| 2004 | Highly Digital, Low-Cost Design of Statistic Signal Acquisition in SoCsabstractPresently, the gap between analog and digital processes is ever increasing. Although digital circuits are still obeying Moore's law, their analog counterparts follow far behind. Since signal acquisition, through ADC circuits is an often required feature, for many embedded applications the benefits of Moore's law have not been achieved. This paper presents our approach to take advantage of the increasing integration of technology for analog interfacing in SoC's, by converting the statistics of the signal. Digital self-tuning of the threshold levels, the use of less expensive and highly variable analog blocks, and stochastic convergence of resolution allow a robust acquisition process. We present the mathematics behind the approach, as well as a set of target applications and experimental results validating the concept. Adão Antônio de Souza Jr., Luigi Carro |
DATE | 2 |
| 2004 | Low Cost Analog Testing of RF Signal PathsabstractA low cost method for testing analogue RF signal paths suitable for BIST implementation in a SoC environment is described. The method is based on the use of a simple and low-cost one-bit digitizer that enables the reuse of processor and memory resources available in the SoC, while incurring little analogue area overhead. The proposed method also allows a constant load to be observed by the circuit, since no switches or muxes are needed for digitizing specific test points. Mathematical background and experimental results are presented in order to validate the test approach. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
DATE | 2 |
| 2004 | Towards a BIST technique for noise figure evaluationabstractThis work presents some results regarding the development of a BIST technique capable of noise figure evaluation. Noise figure is an important parameter in the specification and design of low noise systems, such as communications systems and biomedical instrumentation. A review of published techniques for noise figure evaluation is provided. A new technique aimed to estimate noise figure in a SoC environment is then proposed. The technique is based on the use of a simple and low cost noise generator. Simulation results are provided in order to make an initial evaluation of the feasibility of the proposed approach. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
ETS | 2 |
| 2004 | Analog Signal Processing Reconfiguration for Systems-on-Chip Using a Fixed Analog Cell Approach
Eric E. Fabris, Luigi Carro, Sergio Bampi |
FPL | 2 |
| 2004 | A Low Power FPAA for Wide Band Applications
Erik Schüler, Luigi Carro |
FPL | 2 |
| 2004 | Achieving wide frequency range in an analog FPGAabstractReconfigurability is supposed to be one of the main components in most future systems-on-chip (SoCs) devices. As digital reconfigurability can be achieved through the use of field programmable gate arrays (FPGAs), analog programmability can come from their analog counterparts named FPAAs (field programmable analog arrays). The analog front-end of SoCs should be able to execute functions from low frequency (instrumentation, for example) to high frequency (wireless communication), what is not achievable yet with the actual existing FPAAs. This work presents an analog interface able to deal with this problem, without degrading the flexibility of the FPAA and with a low power cost for the system. Practical results show how the proposed interface can deal with frequencies that commercial FPAAs do not reach, allowing analog signal processing in a wide range of frequencies in both continuous time and sampled systems. Erik Schüler, Luigi Carro |
FPT | 2 |
| 2004 | An Intrinsically Robust Technique for Fault Tolerance under Multiple Upsets
Carlos Arthur Lang Lisbôa, Luigi Carro |
IOLTS | 2 |
| 2004 | Low Cost On-Line Testing of RF Circuits
Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
IOLTS | 2 |
| 2004 | Searching for Global Test Costs Optimization in Core-Based Systems
Érika F. Cota, Luigi Carro, Marcelo Lubaszewski, Alex Orailoglu |
J. Electron. Test. | 2 |
| 2004 | A New FPGA for DSP Applications Integrating BIST Capabilities
Alex Gonsales, Marcelo Lubaszewski, Luigi Carro, Michel Renovell |
J. Electron. Test. | 3 |
| 2004 | Strategies for the integration of hardware and software IP components in embedded systems-on-chip
Flávio Rech Wagner, Wander O. Cesário, Luigi Carro, Ahmed Amine Jerraya |
Integr. | 3 |
| 2004 | Reusing an on-chip network for the test of core-based systemsabstractNetworks-on-chip are likely to become the main communication platform of systems-on-chip. To cope with the growing complexity of the test of such systems, the authors propose the reuse of the on-chip network as a test access mechanism to the cores embedded into systems that use this communication platform. An algorithm exploiting the network characteristics to minimize test time is presented. Then, the reuse strategy is evaluated considering a number of system configurations, such as different positions of the cores in the network, power consumption constraints and number of interfaces with the tester. Experimental results for the ITC'02 SOC Test Benchmarks show that the parallelization capability of the network can be exploited to reduce the system test time, whereas area and pin overhead are strongly minimized. Érika F. Cota, Luigi Carro, Marcelo Lubaszewski |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2003 | Designing fault tolerant systems into SRAM-based FPGAsabstractThis paper discusses high level techniques for designing fault tolerant systems in SRAM-based FPGAs, without modification in the FPGA architecture. Triple Modular Redundancy (TMR) has been successfully applied in FPGAs to mitigate transient faults, which are likely to occur in space applications. However, TMR comes with high area and power dissipation penalties. The new technique proposed in this paper was specifically developed for FPGAs to cope with transient faults in the user combinational and sequential logic, while also reducing pin count, area and power dissipation. The methodology was validated by fault injection experiments in an emulation board. We present some fault coverage results and a comparison with the TMR approach. Fernanda Lima Kastensmidt, Luigi Carro, Ricardo Augusto da Luz Reis |
DAC | 2 |
| 2003 | Ultimate low cost analog BISTabstractIn this work a BIST method for linear analog circuits with very low cost and the smallest possible analog overhead area is presented. The method is suitable to be implemented in the SoC environment, as it allows the reuse of resources already available in the system, and it is essentialy digital. Theoretical background is provided, and experimental results demonstrate the advantages and limits of the proposed approach. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
DAC | 2 |
| 2003 | Reducing pin and area overhead in fault-tolerant FPGA-based designsabstractThis paper proposes a new high-level technique for designing fault tolerant systems in SRAM-based FPGAs, without modifications in the FPGA architecture. Traditionally, TMR has been successfully applied in FPGAs to mitigate transient faults, which are likely to occur in space applications. However, TMR comes with high area and power dissipation penalties. The proposed technique was specifically developed for FPGAs to cope with transient faults in the user combinational and sequential logic, while also reducing pin count, area and power dissipation. The methodology was validated by fault injection experiments in an emulation board. We present some fault coverage results and a comparison with the TMR approach. Fernanda Lima Kastensmidt, Luigi Carro, Ricardo Augusto da Luz Reis |
FPGA | 2 |
| 2003 | Power-aware NoC Reuse on the Testing of Core-based SystemsabstractThis work discusses the impact of power consumption on the test time of core-based systems, when an available on-chip network is reused as test access mechanism. A previously proposed technique for the reuse of an on-chip network is extended to consider power consumption during test, while minimizing the system testing time. Experimental results with the ITC'02 SoC benchmarks show that although power constraints can preclude the full exploration of the network parallelism, this platform is still a powerful mechanism for the system test time reduction at a very low cost. Érika F. Cota, Luigi Carro, Flávio Rech Wagner, Marcelo Lubaszewski |
ITC | 2 |
| 2003 | Low Power Java Processor for Embedded Applications
Antonio Carlos Schneider Beck, Luigi Carro |
VLSI-SOC | 2 |
| 2003 | An All-Digital ADC for Instrumentation within SOCs
Adão Antônio de Souza Jr., Luigi Carro |
VLSI-SOC | 2 |
| 2003 | The Impact of NoC Reuse on the Testing of Core-based SystemsabstractThe authors propose the reuse of on-chip networks for the test of core-based systems that use this platform. Two possibilities of reuse are proposed and discussed with respect to test time minimization. An algorithm exploiting network characteristics to reduce test time is presented. Experimental results show that the parallelization capability of the network can be exploited to reduce the system test time, whereas area and pin overhead are strongly minimized. Érika F. Cota, Márcio Eduardo Kreutz, Cesar A. Zeferino, Luigi Carro, Marcelo Lubaszewski, Altamiro Amadeu Susin |
VTS | 4 |
| 2003 | Ultra Low Cost Analog BIST Using Spectral AnalysisabstractIn this work, a low cost method to implement an analog BIST scheme for the system on chip environment is presented. The method is based on spectral analysis and it is entirely digital. A simple and low cost 1-bit digitizer is used to capture analog information without the need for an AD converter or oversampling techniques. It also allows partitioning of the analog circuit for test thanks to the low analog area overhead of the digitizer. The mathematical framework and a test example are presented, with practical results illustrating limitations and advantages of the proposed technique. Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
VTS | 2 |
| 2003 | The SigmaDelta-BIST Method Applied to Analog Filters
L. Cassol, O. Betat, Luigi Carro, Marcelo Lubaszewski |
J. Electron. Test. | 3 |
| 2003 | A Statistical Sampler for a New On-Line Analog Test Method
Marcelo Negreiros, Luigi Carro, Altamiro Amadeu Susin |
J. Electron. Test. | 2 |
| 2003 | A multiple bit upset tolerant SRAM memoryabstractSRAMs are used nowadays in almost every electronic product. However, as technology shrinks transistor sizes, single and multiple bit upsets only observable in space applications previously are now reported at ground level. This article presents a high level technique to protect SRAM memories against multiple upsets based on correcting codes. The proposed technique combines Reed Solomon code and Hamming code to assure reliability in presence of multiple bit flips with reduced area and performance penalties. Multiple upsets were randomly injected in various combinations of memory cells to evaluate the robustness of the method. The experiment was emulated in a Virtex FPGA platform. Results show that 100% of the injected double faults and a large amount of multiple faults were corrected by the method. Gustavo Neuberger, Fernanda Lima Kastensmidt, Luigi Carro, Ricardo Augusto da Luz Reis |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2002 | Test Planning and Design Space Exploration in a Core-Based EnvironmentabstractThis paper proposes a comprehensive model for test planning in a core-based environment. The main contribution of this work is the use of several types of TAMs and the consideration of different optimization factors (area, ping and test time) during the global TAM and test schedule definition. This expansion of concerns makes possible an efficient yet fine-grained search in the huge design space of a reuse-based environment. Experimental results clearly show the variety of trade-offs that can be explored using the proposed model, and its effectiveness on optimizing the system test design. Érika F. Cota, Luigi Carro, Marcelo Lubaszewski, Alex Orailoglu |
DATE | 2 |
| 2001 | Synthesis of an 8051-Like Micro-Controller Tolerant to Transient Faults
Érika F. Cota, Fernanda Lima Kastensmidt, Sana Rezgui, Luigi Carro, Raoul Velazco, Marcelo Lubaszewski, Ricardo Augusto da Luz Reis |
J. Electron. Test. | 4 |
| 2000 | TI-BIST: a temperature independent analog BIST for switched-capacitor filtersabstractThis paper describes a method to obtain a temperature independent analog BIST. The test procedure is based on the reuse of existing analog circuits, configured either as stimuli generators or as signature analyzers. The paper explains the general problem of temperature deviation present in analog BIST, and shows an approach to overcome this limitation, validated by simulation results. Luigi Carro, Érika F. Cota, Marcelo Lubaszewski, Yves Bertrand, Florence Azaïs, Michel Renovell |
Asian Test Symposium | 1 |
| 2000 | System Synthesis for Multiprocessor Embedded ApplicationsabstractThis paper presents the system synthesis techniques available in S/sup 3/E/sup 2/S, a CAD environment for the specification, simulation, and synthesis of embedded electronic systems that can be modeled as a combination of analog parts, digital hardware, and software. S/sup 3/E/sup 2/S is based on a distributed, object-oriented system model, where objects are initially modeled by their abstract behavior and may be later refined into digital or analog hardware and software. System synthesis is targeted to a multiprocessor platform. Each processor, either a custom-designed one or an off-the-shelf component, can have a specialized behavior like signal processing or control processing. The environment selects processors that best match the desired application by analyzing and comparing processor and application characteristics. The paper illustrates the architecture selection process with concrete examples. Luigi Carro, Márcio Eduardo Kreutz, Flávio Rech Wagner, Márcio Oyamada |
DATE | 1 |
| 2000 | Non-Linear Components for Mixed Circuits Analog Front-EndabstractThis paper presents the development of some front-end analog circuits for mixed signals systems. The paper proposes the use of externally linear internally nonlinear analog circuits. Using this approach, analog area is greatly reduced and circuits can be built on top of completely digital technologies. Experimental results in the analog and digital domain support the proposed approach to mixed circuits design. Luigi Carro, Adão Antônio de Souza Jr., Marcelo Negreiros, Gabriel Parmegiani Jahn, Denis Teixeira Franco |
DATE | 1 |
| 2000 | Reuse of Existing Resources for Analog BIST of a Switch Capacitor FilteabstractThe objective of this paper is to discuss the possibility of reusing the existing hardware originally present in an analog application to implement test functions for a completely autonomous self-testable solution. In this first approach, a 8/sup th/ order analog linear filter is used as an application example. The required modifications in the circuit are presented with the results in terms of area overhead and fault coverage. Érika F. Cota, Michel Renovell, Florence Azaïs, Yves Bertrand, Luigi Carro, Marcelo Lubaszewski |
DATE | 5 |
| 2000 | System Design Based on Single Language and Single-Chip Java ASIP MicrocontrollerabstractMicrocontrollers have been playing an important role in the embedded market. However, the designer of microcontroller based systems must deal with different languages and tools in the hardware and software development, despite of their distinct design process. This paper presents a new design strategy to implement embedded applications described uniquely in Java, while maintaining software compatibility throughout the design process. Moreover, the target hardware is a single chip FPGA, taking benefit from their low cost and easy reconfiguration to customize the microcontroller. This papers presents the environment and some results of system synthesis. Sérgio Akira Ito, Luigi Carro, Ricardo P. Jacobi |
DATE | 2 |
| 2000 | FPGA Architecture Comparison for Non-Conventional Signal ProcessingabstractIn the design of a specific application requiring scalar processing and a neural network, it was noticed that the same underlying hardware used to process analog signals could be used as another way to implement neural networks. This paper presents the main ideas behind this design approach, and a comparison in terms of area and processing speeds of both solutions, when using an FPGA as a substrate. Denis Teixeira Franco, Luigi Carro |
IJCNN (2) | 2 |
| 1999 | A Method to Diagnose Faults in Linear Analog Circuits using an Adaptive TesterabstractThis work presents a new diagnosis method for use in an adaptive analog tester. The tester is able to detect faults in any linear circuit by learning a reference behaviour in a first step, and comparing this behaviour against the output of the circuit under test in a second step. Considering the same basic structure, the diagnosis method consists on injecting probable faults in a mathematical model of the circuit and later comparing its output with the output of the real faulty circuit. This method has been successfully applied to a case study, a biquad filter. Component soft, large, and hard deviations, and faults in operational amplifiers were considered. The results obtained from practical experiments with this analog circuit are discussed in the paper. Érika F. Cota, Luigi Carro, Marcelo Lubaszewski |
DATE | 2 |
| 1999 | Architecture Considerations for Mixed Signals FPGAsabstractNo abstract available. Luigi Carro |
FPGA | 1 |
| 1998 | Efficient Analog Test Methodology Based on Adaptive AlgorithmsabstractThis papers describes a new, fast and economical methodology to test linear analog circuits based on adaptive algorithms. To the authors knowledge, this is the first time such technique is used to test analog circuits, allowing complete fault coverage. The paper presents experimental results showing easy detection of soft, large-deviation and hard faults, with low cost instrumentation. Components variations from 5 % to 1 % have been detected, as the comparison parameter (output error power) varied from 300 % to 20%. 1 Introduction and Luigi Carro, Marcelo Negreiros |
DAC | 1 |
| 1996 | Embedded Systems Design with Frontend CompilersabstractWe describe our research in the field of application specific instruction-set processors. We first show our approach to design integrated processors, and compare it to other approaches. We focus on the importance of the front-end compiler, which can lead to different optimization strategies. This approach is different from other published works, in the sense that the processor and all its architecture changes to optimize a system are defined only after the code it must execute is known, and not the other way around. We present the results of our system in some case studies, and we conclude by showing some results of a RISC used as a microcontroller. C. Alba, Luigi Carro, A. Lima, Altamiro Amadeu Susin |
ICCD | 2 |
| 1996 | Prototyping and reengineering of microcontroller-based systemsabstractThis paper describes our current research in the field of systems design, trying to reach an Application Specific Integrated System (ASIS). Our target system is based on industry applications. We show the design approach to change presently developed boards using classical microcontrollers, migrating the Cisc architecture to an ASIP architecture. The studied examples show meaningful gains regarding the total area of the processor. Luigi Carro, C. Pereira, Altamiro Amadeu Susin |
RSP | 1 |
| 1994 | Algorithms and architectures to computational systems implementationabstractThis paper describes some techniques currently under research to explore hardware-software tradeoffs during a system development. We show that the moving of SW operations to HW can be further improved if the source code is modified in order to increase the overall parallelism of the system. We then show the limits of this approach and a new RISC architecture under research to overcome this limitations.> Luigi Carro, Altamiro Amadeu Susin |
RSP | 1 |
| 1993 | SHC-SLX: A levelized compiled, event driven interpreted VLSI simulator
Luigi Carro, César A. M. Marcon, Altamiro Amadeu Susin |
Microprocess. Microprogramming | 1 |