VLDB 2026 Research / reviewers in the wild / expert
Antonio Carlos Schneider Beck
dblp:02/7021 · also Antonio C. S. Beck, Antonio Carlos S. Beck, Antonio Carlos Schneider Beck Filho
· DBLP profile ↗
81ranked-venue papers
5as first author
33since 2021 · last 2026
0000-0002-4492-1747ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 67 · 5 first-author · 26 since 2021Software engineering, systems software and programming languages · 18 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-aware DVFS-driven workload provisioning in heterogeneous cloud FaaS architecturesabstractAbstract Cloud Warehouses are evolving with diverse computational resources, including CPUs, GPUs, and accelerators, catering to a multitude of tenant applications. While this heterogeneity promises improved performance and energy efficiency, harnessing its full potential poses challenges due to dynamic workload characteristics and variable application demands. To address this, scheduling approaches combined with optimization techniques like Dynamic Voltage and Frequency Scaling (DVFS) are crucial. However, integrating these approaches effectively can be complex, potentially leading to conflicts and diminished benefits. This research proposes two frameworks, EAPECloud and EAPECloud-DVFS, designed for energy-aware collaborative provisioning in heterogeneous CPU-GPU cloud nodes. The first approach reduces energy consumption by selecting and maintaining a static combination of the best scheduler and V-F pair for most workloads. The second approach goes further by dynamically adjusting the V-F pair of each device using DVFS techniques while selecting the optimal scheduler. While the static approach delivers strong results in most cases, the dynamic strategy achieves even greater energy savings, albeit with an additional convergence time to determine the optimal V-F pair. Although each framework has distinct advantages and use cases, our findings demonstrate that both approaches effectively reduce energy consumption in heterogeneous environments, with EAPECloud-DVFS achieving up to a 126.33% performance improvement compared to the Linux CPU Governor, highlighting its efficiency and applicability in real-time systems. Lucas Rister Machado, Gregory de Moraes Rossato, Antonio Carlos Schneider Beck, Michael G. Jordan, Mateus B. Rutzig |
J. Supercomput. | 3 |
| 2025 | Hyle: An HLS Framework for Hyperdimensional Computing Accelerators on FPGAsabstractHyperdimensional computing (HDC) is an emerging brain-inspired machine learning paradigm that exploits unique properties of high-dimensional vectors. HDC establishes a standard set of operations implemented by all of its classes, with Fourier and binary being the most common classes. The first achieves high accuracy but is built upon complex numbers, hindering its adoption in accelerators, whereas the latter is widely adopted in hardware despite providing lower accuracy. To overcome previous problems, the new CGR class was proposed to fit between Fourier and binary classes, learning better than the binary class at affordable hardware implementation w.r.t. Fourier. This paper introduces Hyle, an HLS-based framework for building HDC accelerators in CGR and binary for FPGAs. Hyle is the first proposal to accelerate CGR. We show that the learning advantage of CGR over BSC can result in faster and smaller accelerators and reduce model size up to 8× at iso-accuracy compared to binary. Caio Vieira, Antonio Carlos Schneider Beck |
ICCAD | 2 |
| 2025 | Integration framework for online thread throttling with thread and page mapping on NUMA systems
Janaina Schwarzrock, Hiago Rocha, Arthur Francisco Lorenzon, Samuel Xavier de Souza, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 5 |
| 2025 | IoT-Edge Splitting With Pruned Early-Exit CNNs for Adaptive InferenceabstractDeep neural networks (DNNs) are driving the Internet of Things (IoT) revolution. To manage latency and privacy concerns in this domain, IoT devices may offload partial or full DNN processing to nearby edge servers rather than relying only on the cloud. In this scenario, while field-programmable gate arrays (FPGAs) are efficient and provide flexibility for these resource-constrained IoT and edge devices, runtime optimizations like pruning and early exit can deliver further improvements. However, they need to be applied with careful design, since offloading and split computing may require constant synchronization of such dynamic DNN models. With that in mind, this article introduces a framework that automatically constructs inference platforms combining pruning and early exit with FPGA-based offloading. It addresses the latency-power-accuracy tradeoff, adapting inference to the unpredictable conditions of IoT-edge environments. Using convolutional neural networks (CNNs) as a case study, it achieves a reduction of up to$1.6\times $in latency and a$3.9\times $improvement in power efficiency, with minimal accuracy loss. Guilherme Korol, Antonio Carlos Schneider Beck |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Balancing Performance and Aging in Cloud Environments
Thiago Gonçalves, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
CLOSER | 2 |
| 2024 | Exploiting Virtual Layers and Pruning for FPGA-Based Adaptive Traffic ClassificationabstractTraffic classification is crucial to many network administration tasks, from resource management to QoS monitoring. While DNNs are the state-of-the-art method for classifying traffic, they are computationally intensive, posing challenges to their adoption within the networking infrastructure. One popular alternative is exploiting the FPGAs in the Smart Network Interface Cards (SmartNICs) to speed up the DNN inference processing. In this context, there are two main obstacles. First, modern networks experience high-volume and very volatile traffic flow, making the design of in-network accelerators difficult. Second, the use of FPGA-enabled SmartNICs involves reconfiguration when changing the classification task, which leads to significant time and energy overheads. In this work, we propose two different but complementary solutions to the aforementioned challenges: the use of pruning, which dynamically removes parts of the DNN to speed up its processing at the cost of controlled accuracy drops, therefore adapting the inference processing to the constantly changing traffic; and Hardware Virtual Layers (HWVL), which eliminate the need for FPGA reconfigurations for seamless and almost instantaneous task switching. Both approaches are combined in the Spyke Framework, improving throughput in up to 1.46× and reducing the energy per inference in up to 1.37× compared to a state-of-the-art FPGA accelerator. Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DSD | 5 |
| 2024 | Allok: a machine learning approach for efficient graph execution on CPU-GPU clusters
Marcelo K. Moori, Hiago Rocha, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
J. Supercomput. | 4 |
| 2024 | Synergistically Rebalancing the EDP of Container-Based Parallel ApplicationsabstractThe use of containers has become standard in cloud environments. However, many parallel applications in containers will not present gains proportional to the extra available hardware. This inefficient use of hardware naturally leads to energy consumption waste. With that in mind, we proposeTT-Autoscaling. It works at two different levels: a) in the container, by automatically and transparently tuning the number of threads at runtime of the application, in a way to optimize the trade-off between energy and performance; b) in the cloud infrastructure, by smartly transferring the released resources to other containers that may run in parallel, making better use of the available resources. We compareTT-Autoscalingto the default execution of containers (serial execution with the maximum number of threads), showing 55.8% of performance improvements, 53.6% of energy reductions, and 79.5% of EDP improvements. We also show thatTT-Autoscalingoutperforms strategies that apply vertical autoscalers proposed by orchestrator tools. Vinicius S. da Silva, Everton Camargo de Lima, Janaina Schwarzrock, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | Adaptive Inference on Reconfigurable SmartNICs for Traffic Classification
Julio Costella Vicenzi, Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
AINA (2) | 5 |
| 2023 | Pruning and Early-Exit Co-Optimization for CNN Acceleration on FPGAsabstractThe challenge of processing heavy-load ML tasks, particularly CNN-based ones at resource-constrained IoT devices, has encouraged the use of edge servers. The edge offers performance levels higher than the end devices and better latency and security levels than the Cloud. On top of that, the rising complexity of ML applications, the ever-increasing number of connected devices, and the current demands for energy efficiency require optimizing such CNN models. Pruning and early-exit are notable optimizations that have been successfully used to alleviate the computational cost of inference. However, these optimizations have not yet been exploited simultaneously: while pruning is usually applied at design time, which involves retraining the CNN before deployment, early-exit is inherently dynamic. In this work, we propose AdaPEx, a framework that exploits the intrinsic reconfigurable FPGA capabilities so both can be cooperatively employed. AdaPEx first explores the trade-off between pruning and early-exit at design-time, creating a design space never exploited in the state-of-the-art. Then, AdaPEx applies FPGA reconfiguration as a means to enable the combined use of pruning and early-exit dynamically. At runtime, this allows matching the inference processing to the current edge conditions and a user-configurable accuracy threshold. In a smart IoT application, AdaPEx processes up to 1.32× more inferences and improves EDP by up to 2.55× over the state-of-the-art FPGA-based FINN accelerator. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Jerónimo Castrillón, Antonio Carlos Schneider Beck |
DATE | 5 |
| 2023 | Automatic CPU-GPU Allocation for Graph ExecutionabstractAlthough advances in modern GPUs have accelerated the execution of heavy data processing applications, speeding up graph processing on these systems is not a trivial task: graph applications are characterized by their high volume of irregular memory access that varies with the graph structure so that they do not reach their peak performance when executing on GPUs in many times. In these cases, the CPU execution is more suitable. Given that graph structures can be identified through high-level metrics (e.g., diameter and average clustering coefficient), they may assist the designer in deciding where to execute a given input graph (GPU or CPU). Based on that, in this work, we propose GraCo: a graph processing framework to help the decision-making on where to process a batch of graph applications. Whenever a new batch is submitted to the target HPC system, GraCo decides the best machine to execute each application based only on the available high-level features, precluding any additional applications' execution. Our experimental results comparing GraCo with three other strategies executed on an HPC system comprised of 4 CPUs and 3 GPUs showed that GraCo outperforms the other strategies by at least 34.94×, 13.59×, and 492.31× in total execution time, energy, and energy-delay product. Marcelo K. Moori, Hiago Rocha, Matheus A. Silva, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 6 |
| 2023 | Improving the efficiency of graph algorithm executions on high-performance computingabstractSummary The growing need for extracting information from large graphs has been pushing the development of parallel graph algorithms. However, the highly irregular structure of the real‐world graphs limits the performance and energy improvements of graph applications. In this paper, we show that, in most cases, using all the available cores of the multiprocessor is not the best option in terms of the aforementioned non‐functional requirements. Based on that, we proposeGraphKat, a framework that enables the simultaneous processing of several algorithms/graphs instead of executing them serially (i.e., one after another), increasing efficiency in terms of performance and energy.GraphKatworks in two steps: (i) it characterizes the graph applications with a specific number of threads based on their efficiency levels; and (ii) it defines the execution order of all graph applications in the target system. Experimental results on three multicore processors (Intel and AMD) show thatGraphKatimproves the overall system's efficiency related to performance (up to ) and energy‐saving (up to 245.21), and reduces the graph applications' execution time (up to ) and energy consumption (up to 6.64) compared to the default execution of parallel applications on HPC systems. Marcelo K. Moori, Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Mitigating execution unit contention in parallel applications using instruction-aware mappingabstractSummary Parallel applications running on simultaneous multithreading (SMT) processors naturally compete for execution units when their threads are mapped to the same core. This issue is further aggravated when such threads execute similar instructions that stress the same execution unit type, making their execution to behave very similarly as if the threads were running sequentially. This, in turn, will lead to performance degradation and underutilization of hardware resources. This work proposes a completely transparent framework (no modifications to the source code are necessary) that automatically maps threads of multiple parallel applications on SMT processors. The framework focuses on improving performance by mitigating the contention on execution units, considering each thread's instruction types, which are detected at runtime by our framework. Results show performance gains of 21% (geometric mean), compared to the native scheduler of the operating system. Matheus S. Serpa, Eduardo Henrique Molina da Cruz, Matthias Diener, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 9 |
| 2023 | MVSym: Efficient symbiotic exploitation of HLS-kernel multi-versioning for collaborative CPU-FPGA cloud systems
Michael G. Jordan, Bernardo Neuhaus Lignati, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
Integr. | 5 |
| 2023 | Energy-aware fully-adaptive resource provisioning in collaborative CPU-FPGA cloud environments
Michael G. Jordan, Guilherme Korol, Tiago Knorst, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 5 |
| 2022 | Using machine learning to optimize graph execution on NUMA machinesabstractThis paper proposes PredG, a Machine Learning framework to enhance the graph processing performance by finding the ideal thread and data mapping on NUMA systems. PredG is agnostic to the input graph: it uses the available graphs' features to train an ANN to perform predictions as new graphs arrive - without any application execution after being trained. When evaluating PredG over representative graphs and algorithms on three NUMA systems, its solutions are up to 41% faster than the Linux OS Default and the Best Static - on average 2% far from the Oracle -, and it presents lower energy consumption. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
DAC | 4 |
| 2022 | AdaFlow: A Framework for Adaptive Dataflow CNN Acceleration on FPGAsabstractTo meet latency and privacy requirements, resource-hungry deep learning applications have been migrating to the Edge, where IoT devices can offload the inference processing to local Edge servers. Since FPGAs have successfully accelerated an increasing number of deep learning applications (especially CNN-based ones), they emerge as an effective alternative for Edge platforms. However, Edge applications may present highly unpredictable workloads, requiring runtime adaptability in the inference processing. Although some works apply model switching on CPU and GPU platforms by exploiting different pruning rates at runtime, so the inference can adapt according to some quality-performance trade-off, FPGA-based accelerators refrain from this approach since they are synthesized to specific CNN models. In this context, this work enables model switching on FPGAs by adding to the well-known FINN accelerator an extra level of adaptability (i.e., flexibility) and support to the dynamic use of pruning via fast model switch on flexible accelerators, at the cost of some extra logic, or via FPGA reconfigurations of fixed accelerators. From that, we developed AdaFlow: a framework that automatically builds, at design time, a library from these new available versions (flexible and fixed, pruned or not) that will be used, at runtime, to dynamically select a given version according to a user-configurable accuracy threshold and current workload conditions. We have evaluated AdaFlow under a smart Edge surveillance application with two CNN models and two datasets, showing that AdaFlow processes, on average, 1.3× more inferences and increases, on average, 1.4× the power efficiency over state-of-the-art statically deployed dataflow accelerators. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DATE | 4 |
| 2022 | SNAP: Selective NTV Heterogeneous Architectures for Power-Efficient Edge ComputingabstractWhile there is a growing need to process ML inference on the edge for improved latency and extra security, general-purpose solutions alone cannot cope with the increasing performance demand under power restrictions. Considering that systolic arrays are a prominent, but also power-hungry solution, we propose a methodology to enable their use in edge devices. For that, we propose SNAP, a selective Near-Threshold Voltage (NTV) strategy to explore heterogeneous MPSoCs with two voltage islands, one at NTV, and another at nominal voltage. By adopting a dynamic programming approach, SNAP may selectively apply NTV to the systolic array and to an optimal subset of cores in RISC- V-based MPSoCs, enabling ML acceleration on the edge. Combined with a smart application mapping, the strategy increases performance by up to 18.9 % over a nominal design within the same power limits. Rafael Billig Tonetto, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
DSD | 2 |
| 2022 | ConfAx: Exploiting Approximate Computing for Configurable FPGA CNN Acceleration at the EdgeabstractThe number of CNN-based applications executing at the Edge has been considerably increasing. Considering that CNNs are recognized error-resilient and the varied Edge conditions, we exploit hardware-level Approximate Computing to optimize FPGA-based CNN accelerators without any model retraining or other modifications. Given that, we propose ConfAx, a fully configurable multi-target Framework that navigates the accuracy-performance-resource trade-off to deploy different versions of CNN FPGA approximate accelerators. With an Edge case study (video surveillance), we show that ConfAx reduces power (up to $1.65\times)$ and energy (up to $ 1.44\times$) over a state-of-the-art accelerator at minor accuracy penalties (0.88% on average). Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ISCAS | 4 |
| 2022 | Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 7 |
| 2021 | Exploiting HLS-Generated Multi-Version Kernels to Improve CPU-FPGA Cloud SystemsabstractCloud Warehouses have been exploiting CPU-FPGA collaborative execution environments, where multiple clients share the same infrastructure to achieve to maximize resource utilization with the highest possible energy efficiency and scalability. However, the resource provisioning is challenging in these environments, since kernels may be dispatched to both CPU and FPGA concurrently in a highly variant scenario, in terms of available resources and workload characteristics. In this work, we propose MultiVers, a framework that leverages automatic HLS generation to enable further gains in such CPU-FPGA collaborative systems. MultiVers exploits the automatic generation from HLS to build libraries containing multiple versions of each incoming kernel request, greatly enlarging the available design space exploration passive of optimization by the allocation strategies in the cloud provider. Multivers makes both kernel multiversioning and allocation strategy to work symbiotically, allowing fine-tuning in terms of resource usage, performance, energy, or any combination of these parameters. We show the efficiency of MultiVers by using real-world cloud request scenarios with a diversity of benchmarks, achieving average improvements on makespan and energy of up to 4.62x and 19.04x, respectively, over traditional allocation strategies executing non-optimized kernels. Bernardo Neuhaus Lignati, Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ASP-DAC | 5 |
| 2021 | Synergically Rebalancing Parallel Execution via DCT and Turbo BoostingabstractThe increasing use of cloud and HPC systems put more pressure on the efficient utilization of hardware resources to keep costs low. Many dynamic concurrency throttling (DCT) techniques have successfully used to tune the number of executing threads to better balance a parallel application according to its available scalability. Similarly, boosting frequency strategies have been used to speed up the sequential parts’ execution. Given that, we propose Poseidon, the first transparent and automatic approach that cooperatively exploits both techniques to rebalance OpenMP applications without any preprocessing, with no code transformation, recompilation, or OS modification. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
DAC | 5 |
| 2021 | Combining Thread Throttling and Mapping to Optimize the EDP of Parallel ApplicationsabstractThread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1. Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 7 |
| 2021 | Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo BoostingabstractWith the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 7 |
| 2021 | Boosting Graph Analytics by Tuning Threads and Data Affinity on NUMA SystemsabstractThe execution of large real-world graphs, such as web searches and social networks, has been boosting by modern HPC systems. However, their irregular communication patterns and poor data locality impose many challenges, mainly when executed on NUMA systems. As we show in this paper, there is no one-fits-all configuration for threads/data mapping, and the best combination will vary according to the NUMA system, graph algorithm, and input graph at hand. Based on that, we propose Graphith: a framework that automatically enhances graph processing performance by adapting its execution considering the variables mentioned above. Graphith also goes one step further and improves the existing policies: it uses a Genetic Algorithm to fine-tune the thread-to-core allocation combined with data mapping policies. With that, Graphith improves in 21%, on average, the default execution, and is, on average, 7% better than the best possible combination of standard policies. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 4 |
| 2021 | FAIR: Fully-Adaptive Framework for Improving Resource Provisioning in Collaborative CPU-FPGA Cloud EnvironmentsabstractCloud Warehouses have been exploiting CPU-FPGA collaborative environments to accelerate multi-tenant applications to achieve scalability and maximize resource utilization. However, resource provisioning is challenging in these environments since kernels may be dispatched to CPU and FPGA concurrently in a scenario with highly variant workloads and demands. The provisioning complexity is further aggravated due to diverse CPU and FPGA architectures being used at Cloud Warehouses (e.g., different FPGA/CPU devices between nodes). That means that the resource manager needs to consider the workload to be allocated and the characteristics of the Cloud infrastructure, which can be non-uniform. This paper shows that efficient resource provisioning in CPU-FPGA cloud environments requires different strategies depending on the demand, architecture, and workload. To provide the best use of resources in this complex environment, we propose FAIR, a Fully-Adaptive approach for Improving Resource provisioning in Collaborative CPU-FPGA Cloud. FAIR is end user-transparent and, in contrast to existing approaches, exploits the benefits of multiple provisioning strategies by dynamically selecting the most appropriate depending on the warehouse needs, workload properties, and target architecture. Over a varied set of scenarios, FAIR significantly improves the performance and energy efficiency of the environment compared to the use of fixed single strategies. On average, FAIR provides 32% performance improvements over the use of the best fixed single strategy. Compared to an Oracle that always selects the best energy strategies, FAIR achieves only 3% energy degradation. Michael G. Jordan, Guilherme Korol, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
SBAC-PAD | 4 |
| 2021 | Mitigating the processor aging through dynamic concurrency throttling
Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Parallel Distributed Comput. | 5 |
| 2021 | Improving multitask performance and energy consumption with partial-ISA multicores
Jeckson Dellagostin Souza, Pedro Henrique Exenberger Becker, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 3 |
| 2021 | Low learning-cost offline strategies for EDP optimization of parallel applications
Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Samuel Xavier de Souza, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 5 |
| 2021 | Multi-Target Adaptive Reconfigurable Acceleration for Low-Power IoT ProcessingabstractLow-power processors for the Internet-of-Things (IoT) demand a high degree of adaptability to efficiently execute applications with different resource requirements under varying scenarios. Current single-ISA heterogeneous Chip Multiprocessors (CMPs), such as ARM's big.LITTLE, provide multiple cores and voltage/frequency levels to address this challenge. However, finding the best possible type of core and the corresponding voltage/frequency level for all the execution scenarios, which involve different applications and phases, remains far from being reached. In this article, we propose extending such a single-ISA heterogeneous CMP with a Coarse-Grained Reconfigurable Array (CGRA) and a hardware-based dynamic binary translation (DBT) module that transparently maps application code onto the CGRA for acceleration. To achieve low-energy levels and efficiently manage the power consumption of the CGRA, we introduce an additional voltage rail that enables operation in the Near-Threshold Voltage (NTV) regime when needed, leveraging key features of the CGRA's structure to address the implementation challenges of NTV computing. For less than 35 percent area overhead to the baseline CMP, performance and energy consumption are improved as follows. Compared to: (a) power-efficient execution in the LITTLE core, MuTARe achieves 29 percent reduction in energy consumption, and$2\times$speedup; (b) performance-efficient execution in the big core, a speedup of$1.6\times$with an energy reduction of 41 percent is achieved. Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001 |
IEEE Trans. Computers | 3 |
| 2021 | Synergistically Exploiting CNN Pruning and HLS Versioning for Adaptive Inference on Multi-FPGAs at the EdgeabstractFPGAs, because of their energy efficiency, reconfigurability, and easily tunable HLS designs, have been used to accelerate an increasing number of machine learning, especially CNN-based, applications. As a representative example, IoT Edge applications, which require low latency processing of resource-hungry CNNs, offload the inferences from resource-limited IoT end nodes to Edge servers featuring FPGAs. However, the ever-increasing number of end nodes pressures these FPGA-based servers with new performance and adaptability challenges. While some works have exploited CNN optimizations to alleviate inferences’ computation and memory burdens, others have exploited HLS to tune accelerators for statically defined optimization goals. However, these works have not tackled both CNN and HLS optimizations altogether; neither have they provided any adaptability at runtime, where the workload’s characteristics are unpredictable. In this context, we propose a hybrid two-step approach that, first, creates new optimization opportunities at design-time through the automatic training of CNN model variants (obtained via pruning) and the automatic generation of versions of convolutional accelerators (obtained during HLS synthesis); and, second, synergistically exploits these created CNN and HLS optimization opportunities to deliver a fully dynamic Multi-FPGA system that adapts its resources in a fully automatic or user-configurable manner. We implement this two-step approach as the AdaServ Framework and show, through a smart video surveillance Edge application as a case study, that it adapts to the always-changing Edge conditions: AdaServ processes at least 3.37× more inferences (using the automatic approach) and is at least 6.68× more energy-efficient (user-configurable approach) than original convolutional accelerators and CNN Models (VGG-16 and AlexNet). We also show that AdaServ achieves better results than solutions dynamically changing only the CNN model or HLS version, highlighting the importance of exploring both; and that it is always better than the best statically chosen CNN model and HLS version, showing the need for dynamic adaptability. Guilherme Korol, Michael G. Jordan, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | A Runtime and Non-Intrusive Approach to Optimize EDP by Tuning Threads and CPU Frequency for OpenMP ApplicationsabstractEfficiently exploiting thread-level parallelism has been challenging. Many parallel applications are not sufficiently balanced or CPU-bound to take advantage of the increasing number of cores and the highest possible operating frequency. Moreover, many variables may change according to the system (input set, microarchitecture, and number of cores) or during execution, influencing each parallel region in different ways. Therefore, the task of rightly choosing the ideal configuration (number of threads and DVFS) for each parallel region to deliver the best Energy-Delay Product (EDP) is not straightforward. While the significant number of variables prevents the use of exhaustive search methods, the changing nature of the problem precludes offline strategies. Few solutions are online and synergistically consider thread throttling and DVFS. However, they lack transparency (demand changes in the original code) and/or adaptability (do not automatically adjust to applications at run-time). Our proposed Hoder covers all the characteristics above, optimizing at run-time any dynamically linked OpenMP application, without requiring any code transformation or recompilation. We show Hoder's efficiency by comparing it to two exhaustive offline and two online search approaches, three state-of-the-art techniques, and regular OpenMP execution, considering different setups (Intel 44-, 16- and 12-core; AMD 8- and 12-core). Janaina Schwarzrock, Charles Cardoso De Oliveira, Marcus Ritt, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Proactive Aging Mitigation in CGRAs through Utilization-Aware AllocationabstractResource balancing has been effectively used to mitigate the long-term aging effects of Negative Bias Temperature Instability (NBTI) in multi-core and Graphics Processing Unit (GPU) architectures. In this work, we investigate this strategy in Coarse-Grained Reconfigurable Arrays (CGRAs) with a novel application-to-CGRA allocation approach. By introducing important extensions to the reconfiguration logic and the datapath, we enable the dynamic movement of configurations throughout the fabric and allow overutilized Functional Units (FUs) to recover from stress-induced NBTI aging. Implementing the approach in a resource-constrained state-of-the-art CGRA reveals 2.2× lifetime improvement with negligible performance overheads and less than 10% increase in area. Marcelo Brandalero, Bernardo Neuhaus Lignati, Antonio Carlos Schneider Beck, Muhammad Shafique 0001, Michael Hübner 0001 |
DAC | 3 |
| 2020 | Enhancing Thread-Level Parallelism in Asymmetric Multicores using Transparent Instruction OffloadingabstractAsymmetric multicore architectures (AMC) with single-ISA can accelerate multi-threaded applications by running the serial region on the big core and the parallel region on multiple small cores. In such architectures, all cores implement resource-expensive and application-specific instruction extensions (e.g., SIMD and FP). We argue that instead of implementing such extensions in the big core, the resources must be traded-off to increase the number of small cores. Furthermore, when the big core requires such instruction extensions, we offload execution to the small cores. This design mainly leverages the observation that SIMD/FP operations are more frequently executed inside parallel regions. The proposed AMC provides an additional 1.76x speedup and 12.4% energy savings compared to a traditional AMC of the same area due to enhanced thread-level parallelism (TLP) exploitation. Jeckson Dellagostin Souza, Madhavan Manivannan, Miquel Pericàs, Antonio Carlos Schneider Beck |
DAC | 4 |
| 2020 | A Machine Learning Approach for Reliability-Aware Application Mapping for Heterogeneous MulticoresabstractWe propose a transparent and runtime methodology to increase the system's Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network that makes online predictions of the core's Architectural Vulnerability Factor (AVF), which allows for reliability-aware application-to-core mappings. We experiment with different configurations of RISC-V cores and compare the MWTF of prediction-based mappings against the optimal oracle, showing that our proposed model provides MWTF as close as 5.6% to the oracle. We also compare homogeneous and heterogeneous multicores, showing that heterogeneity provides room for increasing the MWTF in up to 19.4%. Rafael Billig Tonetto, Hiago Rocha, Gabriel L. Nazar, Antonio Carlos Schneider Beck |
DAC | 4 |
| 2020 | Tuning the ISA for increased heterogeneous computation in MPSoCsabstractHeterogeneous MPSoCs are crucial to meeting energy efficiency and performance, given their combination of cores and accelerators. In this work, we propose a novel technique for MPSoCs design, increasing their specialization and task-parallelism within a given area and power budget. By removing the microarchitectural support of costly ISA extensions (e.g., FP, SIMD, crypto) from a few cores (transforming them into PartialISA Cores), we make room to add extra (full and simpler) inorder cores and hardware accelerators. While applications must migrate from Partial-ISA cores when they need the removed ISA support, they also execute at lower power consumption during their ISA-extension-free phases, since partial cores have much simpler datapaths compared to their full-ISA counterparts. On top of it, the additional cores and accelerators increase task-level parallelism and make the MPSoC more suitable for application-specific scenarios. We show the effectiveness of our approach by composing different MPSoCs in distinct execution scenarios, using the FP instructions and RISC-V ISA as a case study. To support our system, we also propose two scheduling policies, performance- and energy-oriented, to coordinate the execution of this novel design. For the former policy, we achieve 2.8× speedup for a neural network road sign detection, 1.53x speedup for a video-streaming app, and 1.2x speedup for a taskparallel scenario, consuming 68%, 75%, and 33% less energy, respectively. For the energy-oriented policy, partial-ISA reduces energy consumption by 29% over a highly efficient baseline, with increased performance. Pedro Henrique Exenberger Becker, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2020 | Enhancing Multithreaded Performance of Asymmetric Multicores with SIMD OffloadingabstractAsymmetric multicore architectures with single-ISA can accelerate multithreaded applications by running code that does not execute concurrently (i.e., the serial region) on a big core and the parallel region on a larger number of smaller cores. Nevertheless, in such architectures the big core still implements resource-expensive application-specific instruction extensions that are rarely used while running the serial region, such as Single Instruction Multiple Data (SIMD) and Floating-Point (FP) operations. In this work, we propose a design in which these extensions are not implemented in the big core, thereby freeing up area and resources to increase the number of small cores in the system, and potentially enhance thread-level parallelism (TLP). To address the case when missing instruction extensions are required while running on the big core we devise an approach to automatically offload these operations to the execution units of the small cores, where the extensions are implemented and can be executed. Our evaluation shows that, on average, the proposed architecture provides 1.76x speedup when compared to a traditional single-ISA asymmetric multicore processor with the same area, for a variety of parallel applications. Jeckson Dellagostin Souza, Madhavan Manivannan, Miquel Pericàs, Antonio Carlos Schneider Beck |
DATE | 4 |
| 2020 | MCEA: A Resource-Aware Multicore CGRA Architecture for the EdgeabstractModern IoT edge devices must address the unpredictability of applications with strict power and temperature constraints. In this scenario, heterogeneous multicore architectures have been driving many solutions due to their high energy efficiency and ability to exploit Task-Level Parallelism. However, while their performance is highly dependent on the quality of the scheduling, their adaptability and generality get restricted when they use fixed-size hardware accelerators. Considering that, this work proposes MCEA, a transparent and power-adaptive multicore reconfigurable architecture. MCEA dynamically adapts the hardware to the workload rather than migrating applications; and predicatively sizes its reconfigurable accelerators without prior knowledge of the applications' behaviors. For that, MCEA uses a synergistic and online profiling system with power gating, achieving performance levels near of homogeneous architectures with fixed and oversized reconfigurable fabric (within 99% on average) while presenting energy efficiency levels similar to heterogeneous architectures statically tuned to a specific workload (within 99% on average). Therefore, MCEA improves Energy-Delay Product in 1.55x and 1.21x when compared to their heterogeneous and homogeneous counterparts, and in 4.72x when compared to a multicore with OoO processors only. We also show that MCEA outperforms a state-of-the-art reconfigurable architecture for the edge under the same power envelope. Guilherme Korol, Michael G. Jordan, Marcelo Brandalero, Michael Hübner 0001, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
FPL | 6 |
| 2020 | A Reliability-Oriented Machine Learning Strategy for Heterogeneous Multicore Application MappingabstractWe propose a methodology to transparently estimate near-optimal application mappings aiming at increasing the Mean Workload to Failure (MWTF) in heterogeneous multicore processors. For that, we leverage an Artificial Neural Network (ANN) capable of estimating the vulnerability factor of RISC-V cores at runtime, which allows for efficient and dynamic application-to-core mappings targeting better MWTF and MWTF/energy tradeoffs. Results show that our ANN-based mapping yields very close-to-optimal solutions, with a difference in MWTF of only 3% when compared to the optimal mapping. When compared to a homogeneous architecture composed of only big cores, heterogeneous architectures may provide improvement in MWTF of up to 20.5% while impacting 12.2% on performance. Rafael Billig Tonetto, Hiago Rocha, Bruno Zatt, Antonio Carlos Schneider Beck, Gabriel L. Nazar |
ISCAS | 4 |
| 2020 | Decreasing the Learning Cost of Offline Parallel Application Optimization StrategiesabstractMany parallel applications do not scale as the number of threads increases, which means that executing them with the maximum possible number of threads will not always deliver the best outcome in performance, energy consumption, or the tradeoff between both (represented by the energy-delay product- EDP). Given that, several strategies, online and offline, have already been proposed to rightly tune the number of threads according to the application. While the former can capture some behaviors that can only be known at runtime, the latter do not impose any execution overhead and can use more efficient and costly algorithms. However, these learning algorithms in static strategics may take several hours, precluding their use or a smooth migration across different systems. In this scenario, we propose a generic methodology for such offline strategies to significantly decrease the learning time by inferring the execution behavior of parallel applications using smaller input sets than the ones used by the target applications. Through the execution of eighteen well-known benchmarks on two multicore processors, we show that our methodology is capable of converging to results that are very close to those that use the regular input set, but converging 84.7% faster, on average. We also show that such a strategy delivers better results than a dynamic one, presenting an EDP 7.7% lower, on average, when executing the applications with the number of threads found during learning. Finally, we also compare our learning methodology with an exhaustive search. It has an average learning cost (i.e., the time spent by our search algorithm to find the best configuration) of only 3.1% to optimize the EDP of the entire benchmark set1. Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 4 |
| 2019 | A Compiler for Automatic Selection of Suitable Processing-in-Memory InstructionsabstractAlthough not a new technique, due to the advent of 3D-stacked technologies, the integration of large memories and logic circuitry able to compute large amount of data has revived the Processing-in-Memory (PIM) techniques. PIM is a technique to increase performance while reducing energy consumption when dealing with large amounts of data. Despite several designs of PIM are available in the literature, their effective implementation still burdens the programmer. Also, various PIM instances are required to take advantage of the internal 3D-stacked memories, which further increases the challenges faced by the programmers. In this way, this work presents the Processing-In-Memory cOmpiler (PRIMO). Our compiler is able to efficiently exploit large vector units on a PIM architecture, directly from the original code. PRIMO is able to automatically select suitable PIM operations, allowing its automatic offloading. Moreover, PRIMO concerns about several PIM instances, selecting the most suitable instance while reduces internal communication between different PIM units. The compilation results of different benchmarks depict how PRIMO is able to exploit large vectors, while achieving a near-optimal performance when compared to the ideal execution for the case study PIM. PRIMO allows a speedup of 38× for specific kernels, while on average achieves 11.8 × for a set of benchmarks from PolyBench Suite. Hameeza Ahmed, Paulo C. Santos 0001, João Paulo C. de Lima, Rafael Fao de Moura, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 6 |
| 2019 | TransRec: Improving Adaptability in Single-ISA Heterogeneous Systems with Transparent and Reconfigurable AccelerationabstractSingle-ISA heterogeneous systems, such as ARM's big.LITTLE, use microarchitecturally-different General-Purpose Processor cores to efficiently match the capabilities of the processing resources with applications' performance and energy requirements that change at run time. However, since only a fixed and non-configurable set of cores is available, reaching the best-possible match between the available resources and applications' requirements remains a challenge, especially considering the varying and unpredictable workloads. In this work, we propose TransRec, a hardware architecture which improves over these traditional heterogeneous designs. TransRec integrates a shared, transparent (i.e., no need to change application binary) and adaptive accelerator in the form of a Coarse-Grained Reconfigurable Array that can be used by any of the General-Purpose Processor cores for on-demand acceleration. Through evaluations with cycle-accurate gem5 simulations, synthesis of real RISC-V processor designs for a 15nm technology, and considering the effects of Dynamic Voltage and Frequency Scaling, we demonstrate that TransRec provides better performance-energy tradeoffs that are otherwise unachievable with traditional big.LITTLE-like designs. In particular, for less than 40% area overhead, TransRec can improve performance in the low-energy mode (LITTLE) by 2.28×, and can improve both performance and energy efficiency by 1.32× and 1.59×, respectively, in high-performance mode (big). Marcelo Brandalero, Muhammad Shafique 0001, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 4 |
| 2019 | The Impact of Parallel Programming Interfaces on the Aging of a Multicore Embedded ProcessorabstractIn order to meet the increasing performance demand of applications, the amount of cores in a single chip package has been increasing. However, the heat has been rising at a higher scale, which accelerates the aging process in modern processors. Therefore, wisely balancing the use of resources is important to extend its longevity. Frequency performance stagnates after a certain amount of concurrent threads starts executing. In such cases, the only result is a temperature rise that directly influences the aging process, reducing the processor lifetime. This unbalance between threads can be originated from many factors, which includes the way threads communicate and synchronize. Considering that those characteristics are related to the Parallel Programming Interface (PPI) used to parallelize the application, this work proposes to evaluate three widely used PPIs executing on an embedded multicore. We show that, depending on the characteristic of the application, by only switching from one PPI to another, it is possible to reduce the effects of aging. For that, we have developed a model based on the Arrhenius equation. We show that OpenMP has a lower impact on the processor aging for memory-bound applications: up to 38% and 68% lower than PThreads and MPI, respectively. On the other hand, PThreads presents the lowest impact on the processor aging for CPU-bound applications. Ângelo Vieira, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Fábio D. Rossi, Jorji Nonaka |
ISCAS | 8 |
| 2019 | Transparent Aging-Aware Thread ThrottlingabstractTo satisfy the rising performance demands of modern applications, the number of cores in a single chip package has been increasing. However, the power dissipated and temperature have been growing at a higher rate, accelerating the aging process of new processors. Considering that a significant number of parallel applications are unbalanced, in many cases performance stagnates after a certain number of concurrent threads starts executing. In such cases, the only outcome is a temperature rise on the processor, which drastically accelerates aging. Given that, we propose an automatic and transparent approach to reduce the processor aging by automatically tuning the number of threads for OpenMP applications at run-time. Our tool, Geras, is entirely transparent to the end-user, so even already compiled binaries can be optimized. Through the execution of twelve well-known benchmarks on two multicore platforms, we show that Geras can improve the processor lifetime by up to 83% and 89% over the standard OpenMP execution and its built-in feature that dynamically adjusts the number of threads, respectively. We also show that Geras outperforms techniques that target performance or energy, which reinforces the need for a specific tool that optimizes aging1. Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
SBAC-PAD | 5 |
| 2019 | Trimming the ISA to Optimize Area and EDP in Heterogeneous CMPsabstractAsymmetric multicores of single-ISA are great solutions for efficient processor resource usage. These processors can maintain software productivity with high energy efficiency by smartly migrating tasks to low-energy cores when high performance is not necessary. This work proposes going one step further by exploiting that 1) the pipelines of specialized instructions (e.g., SIMD - Single Instruction Multiple Data and FP - Floating Point) have extremely high cost (in some cases responsible for more than half of the core area); and 2) these instructions are not used as often as the instructions from the base ISA. By trimming the ISA of some cores in a processor, it is possible to create a Partially Heterogeneous ISA (PHISA) system. PHISA is composed of heterogeneous cores that share the base ISA, but are asymmetric in functionality (only some of these cores implement the full ISA with specialized instructions). PHISA uses transparent migrations of tasks to maintain support for the extended instructions while freeing valuable area and power from the processor design. In this work, we show how PHISA can be used to improve the performance by up to 32% and reduce energy consumption by up to 82% when compared to reference processors in embedded scenarios. Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
SBAC-PAD | 2 |
| 2019 | The Impact of Turbo Frequency on the Energy, Performance, and Aging of Parallel ApplicationsabstractTechnologies that improve the performance of parallel applications by increasing the nominal operating frequency of processors respecting a given TDP (Thermal Design Power) have been widely used. However, they may impact on other non-functional requirements in different ways (e.g. increasing energy consumption or aging). Therefore, considering the huge number of configurations available, represented by the range of all possible combinations among different parallel applications, amount of threads, dynamic voltage and frequency scaling (DVFS) governors, boosting technologies and simultaneous multithreading (SMT), selecting the one that offers the best tradeoff for a non-functional requirement is extremely challenging for software designers. Given that, in this work we assess the impact of changing these configurations on the energy consumption, performance, and aging of parallel applications on a turbo-compliant processor. Results show that there is no single configuration that would provide the best solution for all nonfunctional requirements at once. For instance, we demonstrate that the configuration that offers the best performance is the same one that has the worst impact on aging, accelerating it by up to 1.75 times. With our experiments, we provide guidelines for the developer when it comes to tuning performance using turbo boosting to save as much energy as possible and increase the lifespan of the hardware components. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Alessandro Girardi, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
VLSI-SoC | 6 |
| 2019 | A Knapsack Methodology for Hardware-based DMR Protection against Soft Errors in Superscalar Out-of-Order ProcessorsabstractHigh-performance superscalar processors have been adopted to satisfy the rising demand for processing applications of ever-growing complexity. This extra complexity, added to the increasing vulnerability of transistors due to technology scaling, poses a great challenge since these effects have also been proven to affect ground-level safety-critical applications. To increase microarchitectural resilience, designers may adopt Dual Modular Redundancy (DMR), which offers full fault detection. However, given that DMR incurs in high area and energy overheads, we propose a design-time methodology aiming to achieve the best tradeoff between resilience and area overhead, decreasing DMR costs and maintaining acceptable detection levels for such a complex design. This is done by adopting the Knapsack Problem (KSP) as a heuristic to identify the optimal micro-architectural structures that should be duplicated to achieve target resilience with the smallest possible area overhead. By injecting over 800k faults in 12 significant micro-architectural structures of different versions of the complex Berkeley Out-of-Order Machine (BOOM) superscalar processor modeled with RTL accuracy, we compare this optimal strategy against a greedy one, showing that 90% of vulnerability reduction may be achieved with 50.6% and 107.8% area overheads for the optimal and greedy strategies, respectively. Rafael Billig Tonetto, Douglas Maciel Cardoso, Marcelo Brandalero, Luciano Volcan Agostini, Gabriel L. Nazar, José Rodrigo Azambuja, Antonio Carlos Schneider Beck |
VLSI-SoC | 7 |
| 2019 | Predicting performance in multi-core systems with shared reconfigurable accelerators
Marcelo Brandalero, Thiago Dadalt Souto, Luigi Carro, Antonio Carlos Schneider Beck |
J. Syst. Archit. | 4 |
| 2019 | Aurora: Seamless Optimization of OpenMP ApplicationsabstractEfficiently exploiting thread-level parallelism has been challenging for software developers. As many parallel applications do not scale with the number of cores, the task of rightly choosing the ideal amount of threads to produce the best results in performance or energy is not straightforward. Moreover, many variables may change according to the system at hand (e.g., application, input set, microarchitecture, number of cores) and even during execution. Existing solutions lack transparency (demand changes in the original code) or adaptability (do not automatically adjust to applications at run-time). In this scenario, we propose Aurora, an OpenMP framework that is completely transparent to both the designer and end-user. Without any code transformation or recompilation, it is capable of automatically finding, at run-time and with minimum overhead, the optimal number of threads for each parallel loop region and re-adapt in cases the behavior of a region changes during execution. When executing fifteen well-known benchmarks on four multi-core processors, Aurora improves the Energy-Delay Product by up to 98, 86 and 91 percent over the standard OpenMP execution, the OpenMP feature that dynamically adjusts the number of threads, and the Feedback-Driven Threading, respectively. Arthur Francisco Lorenzon, Charles Cardoso De Oliveira, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2018 | Design space exploration for PIM architectures in 3D-stacked memoriesabstractScaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3D-stacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-of-the-art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High-Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth. João Paulo C. de Lima, Paulo C. Santos 0001, Marco A. Z. Alves, Antonio Carlos Schneider Beck, Luigi Carro |
CF | 4 |
| 2018 | Adaptive and polymorphic VLIW processor to optimize fault tolerance, energy consumption, and performanceabstractBecause most traditional homogeneous and heterogeneous processors have a fixed design that limits its runtime adaptability, they are not able to cope with the varying application behavior when one considers the axes of fault tolerance, performance, and energy consumption altogether. In this context, we propose a new dynamically adaptive processor design that is capable of delivering the best trade-off among these three axes according to the application at hand, or be tuned to optimize a specific metric. This is achieved by extending a polymorphic processor that can change its issue-width during runtime with specific mechanisms for fault tolerance, energy optimization, and performance enhancement. They are controlled by an optimization algorithm that evaluates and chooses which is the best configuration according to given requirements. Considering a metric that weighs all three axes, the proposed adaptive processor delivers a result that is 94.88% of the oracle processor on average, while a static configuration (defined at design time without runtime adaptation) only achieves 28.24% at most, which means that dynamic adaptation is required to cope with different application behaviors as there is not one specific configuration that fits all applications. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Sandip Kundu, Israel Koren, Antonio Carlos Schneider Beck |
CF | 5 |
| 2018 | Approximate on-the-fly coarse-grained reconfigurable acceleration for general-purpose applicationsabstractApproximate functional unit designs have the potential to reduce power consumption significantly compared to their precise counterparts; however, few works have investigated composing them to build generic accelerators. In this work, we do a design-space exploration of state-of-the-art approximate designs, propose a flow for designing approximate coarse-grained reconfigurable arrays (CGRAs), and discuss compilation and runtime reconfiguration issues. We compare the energy savings of precise and approximate reconfigurable acceleration and show that the latter can provide up to 50% additional power savings under a 10% quality loss constraint for the applications in the AxBench suite. Marcelo Brandalero, Luigi Carro, Antonio Carlos Schneider Beck, Muhammad Shafique 0001 |
DAC | 3 |
| 2018 | Employing classification-based algorithms for general-purpose approximate computingabstractApproximate computing has recently reemerged as a design solution for additional performance and energy improvements at the cost of output quality. In this paper, we propose using a tree-based classification algorithm as an approximation tool for general-purpose applications. We show that, without any hardware support, completely implemented in software, our approach can improve performance by up to 4x (1.95x on average) and reduce EDP by up to 19x (4.04 on average) when compared to precise executions. Besides that, in some cases, our software-based mechanism can even outperform traditional hardware-based Neural Network's state-of-the-art designs. Geraldo F. Oliveira, Larissa Rozales Gonçalves, Marcelo Brandalero, Antonio Carlos Schneider Beck, Luigi Carro |
DAC | 4 |
| 2018 | Processing in 3D memories to speed up operations on complex data structuresabstractPointer chasing has been, for years, the kernel operation employed by diverse data structures, from graphs to hash tables and dictionaries. However, due to the bewildering growth in the volume of data that current applications have to deal with, performing pointer chasing operations have become a major source of performance and energy bottleneck, due to its sparse memory access behavior. In this work, we aim to tackle this problem by taking advantage of the already available parallelism present in today's 3D-stacked memories. We present a simple mechanism that can accelerate pointer chasing operations by making use of a state-of-the-art PIM design that executes in-memory vector operations. The key idea behind our design is to run speculative loads, in parallel, based on a given memory address in a reconfigurable window of addresses. Our design can perform pointer-chasing operations on b+tree 4.9 χ faster when compared to modern baseline systems. Besides that, since our device avoids data movement, we can also reduce energy consumption by 85% when compared to the baseline. Paulo C. Santos 0001, Geraldo F. Oliveira, João Paulo C. de Lima, Marco A. Z. Alves, Luigi Carro, Antonio Carlos Schneider Beck |
DATE | 6 |
| 2018 | Precise evaluation of the fault sensitivity of OoO superscalar processorsabstractSince superscalar processors lead the market, their resiliency evaluation by means of fault injection grows in importance. Fault injection strategies usually trade-off their levels of accuracy: low-level HW-based methods are accurate, but very expensive, need special equipment and the actual hardware, and lack controllability; while high-level simulation-based strategies are flexible, fast, easily accessible and have high controllability, but are not accurate since they are based on models that do not always reflect the low-level implementation, mainly when it comes to complex designs like out-of-order multiple-issue processors. In this work, we propose a cycle-accurate fault injection platform for superscalar processors, which has a smart checkpointing mechanism to accelerate injection time, attenuating the short-comings imposed by the aforementioned fault injection methods while providing the same level of abstraction as detailed RTL models. Leveraging from this new platform, we evaluate a complex and parameterizable Out-of-Order processor (BOOM) by experimenting with different issue widths and analyzing the sensitivity of several hardware structures of the processor. Rafael Billig Tonetto, Gabriel L. Nazar, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2018 | Accelerating error-tolerant applications with approximate function reuse
Marcelo Brandalero, Leonardo Almeida da Silveira, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
Sci. Comput. Program. | 4 |
| 2017 | A Mechanism for energy-efficient reuse of decoding and scheduling of x86 instruction streamsabstractCurrent superscalar x86 processors decompose each CISC instruction (variable-length and with multiple addressing modes) into multiple RISC-like pops at runtime so they can be pipelined and scheduled for concurrent execution. This challenging and power-hungry process, however, is usually repeated several times on the same instruction sequence, inefficiently producing the very same decoded and scheduled pops. Therefore, we propose a transparent mechanism to save the decoding and scheduling transformation for later reuse, so that next time the same instruction sequence is found it can automatically bypass the costly pipeline stages involved. We use a coarse-grained reconfigurable array as a means to save this transformation, since its structure enables the recovery of pops already allocated in time and space, and also larger ILP exploitation than superscalar processors. The technique can reduce the energy consumption of a powerful 8-issue superscalar by 31.4% at low area costs, while also improving performance by 32.6%. Marcelo Brandalero, Antonio Carlos Schneider Beck |
DATE | 2 |
| 2017 | LAANT: A library to automatically optimize EDP for OpenMP applicationsabstractEfficiently exploiting thread level parallelism from new multicore systems has been challenging for software developers. While blindly increasing the number of threads may lead to performance gains, it can also result in disproportionate increase in energy consumption. For this reason, rightly choosing the number of threads is essential to reach the best compromise between both. However, such task is extremely difficult: besides the huge number of variables involved, many of them will change according to different aspects of the system at hand and are only possible to be defined at run-time. To address this complex scenario, we propose LAANT, a novel library to automatically find the optimal number of threads for OpenMP applications, by dynamically considering their characteristics, input set, and the processor architecture. By executing nine well-known benchmarks on three real multicore processors, LAANT improves the EDP (Energy-Delay Product) by up to 61%, compared to the standard OpenMP execution; and by 44%, when the dynamic adjustment of the number of threads of OpenMP is activated. Arthur Francisco Lorenzon, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
DATE | 3 |
| 2017 | Improving EDP in multi-core embedded systems through multidimensional frequency scalingabstractEnergy saving management in multi-core embedded environments has been a challenge for designers. To achieve energy efficiency, most studies consider dynamic frequency scaling on one hardware component only, such as processor or memory - which will most likely also affect performance. This work proposes the use of frequency scaling considering the three most important hardware components altogether: processors, L2 cache, and RAM; seeking for the best set of frequencies for each one of them to improve the Energy-Delay Product (EDP), depending on the application's behavior. Therefore, this work addresses multidimensional frequency scaling for multi-core embedded systems. By evaluating different frequency levels, we show that the EDP can be improved in up to 46.4% when compared to the standard way that the frequencies are configured. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Mateus B. Rutzig, Fábio D. Rossi |
ISCAS | 4 |
| 2017 | A framework to automatically generate heterogeneous organization reconfigurable multiprocessingabstractHeterogeneous MPSoCs are vastly used in current embedded systems but they are highly dependent on special compilers. Dynamic reconfigurable systems are an alternative to overcome such drawback due to their adaptability. However, when such architectures are considered, one must concern about which hardware blocks should be heterogeneous and their degree of heterogeneity. In this work, we propose a framework that automatically generates heterogeneous reconfigurable multiprocessors that exploit the ideal ILP of parallel applications to improve performance/watt. Our generated system achieves, on average, 32% of performance improvements with 33% of energy savings over its manual generated counterpart, with equivalent chip area. Josimar Sfreddo, Rafael Fao de Moura, Michael G. Jordan, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck, Mateus B. Rutzig |
ISCAS | 5 |
| 2016 | Adaptive ILP control to increase fault tolerance for VLIW processorsabstractBecause of technology scaling, soft error rate has been increasing in modern processors, affecting system reliability. To mitigate such effect, we propose an adaptive fault tolerance approach that exploits, at run-time, idle functional units to execute duplicated instructions in a configurable VLIW processor. In applications with high Instruction Level Parallelism (ILP) and few functional units available for duplication, it adaptively reschedules instructions according to a configurable threshold, providing a tradeoff between performance and fault tolerance. On average, failure rate is reduced by 89.53%, performance by 5.86%; while energy consumption increases by 72% and area by 22.2%, using a fault tolerance oriented threshold. Anderson Luiz Sartor, Stephan Wong, Antonio Carlos Schneider Beck |
ASAP | 3 |
| 2016 | Run-time phase prediction for a reconfigurable VLIW processor
Anderson Luiz Sartor, Anthony Brandon, Antonio Carlos Schneider Beck, Xuehai Zhou, Stephan Wong |
DATE | 4 |
| 2016 | A reconfigurable heterogeneous multicore with a homogeneous ISA
Jeckson Dellagostin Souza, Luigi Carro, Mateus B. Rutzig, Antonio Carlos Schneider Beck |
DATE | 4 |
| 2016 | Exploiting Idle Hardware to Provide Low Overhead Fault Tolerance for VLIW ProcessorsabstractBecause of technology scaling, the soft error rate has been increasing in digital circuits, which affects system reliability. Therefore, modern processors, including VLIW architectures, must have means to mitigate such effects to guarantee reliable computing. In this scenario, our work proposes three low overhead fault tolerance approaches based on instruction duplication with zero latency detection, which uses a rollback mechanism to correct soft errors in the pipelanes of a configurable VLIW processor. The first uses idle issue slots within a period of time to execute extra instructions considering distinct application phases. The second works at a finer grain, adaptively exploiting idle functional units at run-time. However, some applications present high instruction-level parallelism (ILP), so the ability to provide fault tolerance is reduced: less functional units will be idle, decreasing the number of potential duplicated instructions. The third approach attacks this issue by dynamically reducing ILP according to a configurable threshold, increasing fault tolerance at the cost of performance. While the first two approaches achieve significant fault coverage with minimal area and power overhead for applications with low ILP, the latter improves fault tolerance with low performance degradation. All approaches are evaluated considering area, performance, power dissipation, and error coverage. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Luigi Carro, Fernanda Lima Kastensmidt, Stephan Wong, Antonio Carlos Schneider Beck |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2016 | Investigating different general-purpose and embedded multicores to achieve optimal trade-offs between performance and energy
Arthur Francisco Lorenzon, Márcia C. Cera, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 3 |
| 2015 | The Influence of Parallel Programming Interfaces on Multicore Embedded SystemsabstractThread-Level Parallelism (TLP) exploitation for embedded systems has been a challenge for software developers: while it is necessary to take advantage of the availability of multiple cores, it is also mandatory to consume less energy. To speed up the development process and make it as transparent as possible, software designers use Parallel Programming Interfaces (PPIs). However, as will be shown in this paper, each PPI implements different ways to exchange data using shared memory regions, influencing performance, energy consumption and Energy-Delay Product (EDP), which varies across different embedded processors. By evaluating four PPIs and three multicore processors (ARM A8, A9 and Intel Atom), we demonstrate that by simply switching PPI it is possible to save up to 59% in energy consumption and achieve up to 85% of EDP improvements, in the most significant case. We also show that the efficiency (i.e., The best possible use of the available resources) decreases as the number of threads increases in almost all cases, but at distinct rates. Arthur Francisco Lorenzon, Anderson Luiz Sartor, Márcia C. Cera, Antonio Carlos Schneider Beck |
COMPSAC | 4 |
| 2015 | The Impact of Virtual Machines on Embedded SystemsabstractEmbedded systems are becoming increasingly complex and, due to their tight energy requirements, all the available resources must be used in the best possible way. However, Android, the most used software platform for embedded systems, features a virtual machine to run applications. Even though it ensures flexibility so the application can execute on different underlying architectures without the need for recompilation, it burdens the system because of the introduction of an extra software layer. Considering this scenario, through the development of an extension of the Android QEMU emulator and a specific benchmark set, this work evaluates the significance of the virtual machine by comparing applications written in Java and in native language. We show that, given a fixed energy budget, a different amount of applications can be executed depending the way they were implemented. We also demonstrate that this difference varies according to the processor, by executing the applications on all officially supported Android architectures (Intel x86, ARM, and MIPS). Therefore, even though the Virtual Machine provides total transparency to the software developer, he/she must be aware of it and the underlying target micro architecture at early designs stages so as to build a low-energy application. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
COMPSAC | 3 |
| 2015 | On the influence of static power consumption in multicore embedded systemsabstractEnergy consumption in multicore embedded systems has become a constant concern. Thread-Level Parallelism exploitation may reduce energy consumption because it saves static power consumption of the processor, since performance is obtained. However, as will be shown in this paper, the influence of the static power on the energy consumption and Energy-Delay Product will depend on how significant it is in the processor. By evaluating different levels of static power in the respect to the total power consumption in two embedded processors (ARM and Atom), we demonstrate that if the right value of static power consumption is tuned during the designing and manufacturing, it is possible to save up 35% in energy consumption and achieve up to 20% of improvements in the EDP efficiency (i.e., the best possible use of the available resources). We also show that the more communication the parallel application has, the lower is the impact of static power of the processor in the total energy consumption. Arthur Francisco Lorenzon, Márcia C. Cera, Antonio Carlos Schneider Beck |
ISCAS | 3 |
| 2015 | Evaluation of energy savings on a VLIW processor through dynamic issue-width adaptationabstractThe development of energy efficient hardware has been a trend in microprocessor design for the last two decades. VLIW processors are a representative example, since they have a simpler design and competitive performance, because their ILP exploitation is done statically by the compiler. In this paper, we study the energy savings that could be obtained by adapting such microarchitecture according to the current program phase. Our contribution is twofold. First, by executing a set of benchmarks on the ρ-vex configurable softcore VLIW processor, and by modifying the number of issues, we show the potentials of energy reduction. Then, with this information in hand, we developed an oracle experiment to dynamically vary the issue width of the processor according to the phase behavior, considering two different phase granularites. The potential energy savings using this policy could be as high as 81.5% when compared with the static version, executing the MiBench set. Juan Sebastian Piedrahita Giraldo, Anderson Luiz Sartor, Luigi Carro, Stephan Wong, Antonio Carlos Schneider Beck |
RSP | 5 |
| 2013 | A transparent and energy aware reconfigurable multiprocessor platform for simultaneous ILP and TLP exploitationabstractAs the number of embedded applications increases, companies are launching new platforms within short periods of time to efficiently execute software with the lowest possible energy consumption. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of Custom Reconfigurable Arrays for Multiprocessor Systems (CReAMS). CReAMS is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. We also show that CReAMS delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered for both designs. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
DATE | 2 |
| 2013 | A run-time adaptive multiprocessor systemabstractBecause of the continuous increase in the number and complexity of embedded applications, new platforms have been launched within shorter periods of time to fulfill their performance requirements with the lowest energy consumption possible. However, for each new platform deployment, new tool chains, with additional libraries, debuggers and compilers must come along, breaking binary compatibility. This strategy implies in high hardware and software redesign costs. In this scenario, we propose the exploitation of custom reconfigurable arrays for multiprocessor systems. The proposed approach is composed of multiple adaptive reconfigurable processors that simultaneously exploit Instruction and Thread Level Parallelism. It works in a transparent fashion, so binary compatibility is maintained, with no need to change the software development process or environment. Results show that our proposal delivers higher performance per watt in comparison to a 4-issue Superscalar processor, when the same power budget is considered. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
ISCAS | 2 |
| 2013 | Towards a multiple-ISA embedded system
Jair Fajardo Junior, Mateus B. Rutzig, Luigi Carro, Antonio Carlos Schneider Beck |
J. Syst. Archit. | 4 |
| 2009 | A low cost and adaptable routing network for reconfigurable systemsabstractNowadays, scalability, parallelism and fault-tolerance are key features to take advantage of last silicon technology advances, and that is why reconfigurable architectures are in the spotlight. However, one of the major problems in designing reconfigurable and parallel processing elements concerns the design of a cost-effective interconnection network. This way, considering that Multistage Interconnection Network (MIN) has been successfully used in several computer system levels and applications in the past, in this work we propose the use of a MIN, at the word level, on a coarse-grained reconfigurable architecture. More precisely, this work presents a novel parallel self-placement and routing mechanism for MIN on the circuit-switching mode. We take into account one-to-one as well as multicast (one-to-many) permutations. Our approach is scalable and it is targeted to be used in run-time environments where dynamic routing among functional units is required. In addition, our algorithm is embedded in the switch structure, and it is independent of the interstage interconnection pattern. Our approach can handle blocking and non-blocking networks, symmetrical or asymmetrical topologies. As case study, we use the proposed technique in a dynamic reconfigurable system, showing a major area reduction of 30% without performance overhead. Ricardo S. Ferreira 0001, Marcone Laure, Antonio Carlos Schneider Beck, Thiago Lo, Mateus B. Rutzig, Luigi Carro |
IPDPS | 3 |
| 2008 | Transparent Reconfigurable Acceleration for Heterogeneous Embedded ApplicationsabstractEmbedded systems are becoming increasingly complex. Besides the additional processing capabilities, they are characterized by high diversity of computational models coexisting in a single device. Although reconfigurable architectures have already shown to be a potential solution for such systems, they just present significant speedups of very specific dataflow oriented kernels. Furthermore, reconfigurable fabric is still withheld by the need of special tools and compilers, clearly not sustaining backward software compatibility. In this paper, we propose a new technique to optimize both dataflow and control-flow oriented code in a totally transparent process, without the need of any modification in the source or binary codes. For that, we have developed a Binary Translation algorithm implemented in hardware, which works in parallel to a MIPS processor. The proposed mechanism is responsible for transforming sequences of instructions at runtime to be executed on a dynamic coarse-grain reconfigurable array, supporting speculative execution. Executing the MIBench suite, we show performance improvements of up to 2.5 times, while reducing 1.7 times the required energy, using trivial hardware resources. Antonio Carlos Schneider Beck, Mateus B. Rutzig, Georgi Gaydadjiev, Luigi Carro |
DATE | 1 |
| 2008 | Reducing interconnection cost in coarse-grained dynamic computing through multistage networkabstractCoarse-grained reconfigurable architectures appear as a scalable solution to embedded system design, with a reduced reconfiguration time, memory footprint, as well as placement and routing complexity. To ensure high performance, data must be efficiently delivered to the reconfigurable matrix. For that, several architectures propose the use of fully interconnected local networks, as crossbar or large multiplexers. However, these interconnections are very area consuming. Therefore, in order to reduce the interconnection complexity without losing performance, this work proposes to use Multistage Interconnection Networks. As a case study, we have implemented the proposed approach in a tightly coupled reconfigurable array, which works together with a MIPS processor. Simulation results over the Mibench Benchmark set show savings of up to 26% of the total area, with a decrease of only 1% on the average performance. Ricardo S. Ferreira 0001, Marcone Laure, Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
FPL | 4 |
| 2008 | Balancing reconfigurable data path resources according to application requirementsabstractProcessor architectures are changing mainly due to the excessive power dissipation and the future break of Moore's law. Thus, new alternatives are necessary to sustain the performance increase of the processors, while still allowing low energy computations. Reconfigurable systems are strongly emerging as one of these solutions. However, because they are very area consuming and deal with a large number of applications with diverse behaviors, new tools must be developed to automatically handle this new problem. This way, in this work we present a tool aimed to balance the reconfigurable area occupied with the performance required by a given application, calculating the exact size and shape of a reconfigurable data path. Using as case study a tightly coupled reconfigurable array and the Mibench Benchmark set, we show that the solution found by the proposed tool saves four times area in comparison with the non-optimized version of the reconfigurable logic, with a decrease of only 5.8% on average of its original performance. This way, we open new applications for reconfigurable devices as low cost accelerators. Mateus B. Rutzig, Antonio Carlos Schneider Beck, Luigi Carro |
IPDPS | 2 |
| 2007 | Transparent acceleration of data dependent instructions for general purpose processorsabstractAlthough transistor scaling keeps following Moore’s law, and more area is available for designers, the clock frequency and ILP rate do not present the same level of growth anymore. This way, new architectural alternatives are necessary. Reconfigurable fabric appears to be one emerging possibility: besides exploiting the parallelism among instructions, it can also accelerate sequences of data dependent ones. However, coarse grain reconfiguration wide spread usage is still withhold by the need of special tools and compilers, which clearly do not sustain the reuse of legacy code without any modification. Based on all these facts, this work proposes a new Binary Translation algorithm, implemented in hardware and working in parallel to the processor, responsible for transforming sequences of instructions at run-time to be executed on a dynamic coarse-grain reconfigurable array, tightly coupled to a traditional RISC machine. Therefore, we can take advantage of using pure combinational logic to optimize even control-flow oriented code in a totally transparent process, without any modification in the source or binary codes. Using the Simplescalar Toolset together with the embedded benchmark suite MIBench, we show performance improvements and area evaluation when comparing against traditional superscalar architectures. Antonio Carlos Schneider Beck, Luigi Carro |
VLSI-SoC | 1 |
| 2006 | Automatic Dataflow Execution with Reconfiguration and Dynamic Instruction MergingabstractAs Moore's law is loosing steam, one already sees the phenomenon of clock frequency reduction caused by the excessive power dissipation. New technologies that will completely or partially replace silicon are arising, and new architectural alternatives are necessary. Reconfigurable fabric appears to be one of these solutions, and has shown speed ups of critical parts of several data stream programs. However, the wide spread use of reconfigurable computing is still withhold by the need of special tools and compilers, which clearly preclude software portability and reuse of legacy code. Based on all these facts, this work proposes a coarse-grain dynamic reconfigurable array, tightly coupled to a traditional RISC machine. Besides taking advantage of using combinational logic to speed up the execution, dynamic analysis of the code at run time was implemented to reconfigure the array, maintaining full software compatibility. Using the Simplescalar Toolset together with the embedded benchmark suite MIBench, meaningful performance improvements (up to 3 times of speed up) were shown, thanks to the implementation of the proposed approach Antonio Carlos Schneider Beck, Victor F. Gomes, Luigi Carro |
VLSI-SoC | 1 |
| 2005 | Dynamic reconfiguration with binary translation: breaking the ILP barrier with software compatibilityabstractIn this paper we present the impact of dynamically translating any sequence of instructions into combinational logic. The proposed approach combines a reconfigurable architecture with a binary translation mechanism, being totally transparent for the software designer. Besides ensuring software compatibility, the technique allows porting the same code for different machines tracking technological evolutions. The target processor is a Java machine able to execute Java bytecodes. Experimental results show that even code without any available parallelism can benefit from the proposed approach. Algorithms used in the embedded systems domain were accelerated 4.6 times in the mean, while spending 10.89 times less energy in the average. We present results regarding the impact of area and power, and compare the proposed approach with other Java machines, including a VLIW one. Antonio Carlos Schneider Beck, Luigi Carro |
DAC | 1 |
| 2003 | Low Power Java Processor for Embedded Applications
Antonio Carlos Schneider Beck, Luigi Carro |
VLSI-SOC | 1 |