VLDB 2026 Research / reviewers in the wild / expert
Arthur Francisco Lorenzon
dblp:160/4663
· DBLP profile ↗
65ranked-venue papers
8as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 5 first-author · 21 since 2021Computer networks · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Cost-Performance Efficiency of Scientific Cloud Workflows through GPU Sharing
Matheus M. Costa, Tiago Ferreto, César A. F. De Rose, Odej Kao, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 6 |
| 2026 | Why Large Language Models Struggle with Cloud Instance Selection
Matheus Machado, Matheus M. Costa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Arthur Francisco Lorenzon |
CLOSER | 5 |
| 2026 | When Faster Kernels Do Not Mean Faster Applications in Exascale GPU Systems
Mariana Toledo Costa, Antigoni Georgiadou, James B. White, Woong Shin, Bruno Villasenor Alvarez, Jorda Polo, Karl W. Schulz, Philippe Olivier Alexandre Navaux, O. E. Bronson Messer, Arthur Francisco Lorenzon |
Euro-Par (1) | 10 |
| 2025 | Energy-Aware Node Selection for Cloud-Based Parallel Workloads with Machine Learning and Infrastructure as Code
Denis B. Citadin, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 5 |
| 2025 | WFQ-Based SLA-Aware Edge Applications Provisioning
Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 2 |
| 2025 | LLM-Based Adaptive Digital Twin Allocation for Microservice Workloads
Pedro Henrique Sachete Garcia, Ester S. Oribes, Ivan Mangini Lopes Júnior, Braulio Marques de Souza, Ângelo Vieira, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 6 |
| 2025 | Towards Optimizing Cost and Performance for Parallel Workloads in Cloud Computing
William Maas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 5 |
| 2025 | One GPU, Many Ranks: Enabling Performance and Energy-Efficient In-Transit Visualization via Resource SharingabstractIn-transit visualization has become essential in high-performance computing (HPC) to reduce I/O overheads and enable real-time data analysis. However, as simulations grow in scale and complexity, these visualization tasks increasingly demand substantial computational resources, exacerbating energy consumption and limiting system scalability. As we show in this paper, a key bottleneck is the conventional one-rank-per-GPU allocation model, which leads to irregular GPU utilization and waste of hardware resources. To tackle this challenge, we propose using GPU-sharing strategies to improve energy efficiency in in-transit visualization without compromising performance. We evaluate six distinct configurations built upon three NVIDIA GPU-sharing mechanisms: the default CUDA model with context switching between processes, Multi-Process Service (MPS), which enables dynamic context sharing, and Multi-Instance GPU (MIG), which provides hardware-level partitioning. Using the WarpX simulation code and Ascent visualization framework, our experiments on the Polaris supercomputer span multiple rendering techniques, node counts, and data sizes. Results show that GPU-sharing strategies can improve the trade-off between performance and energy, represented by the energy-delay product (EDP) metric, by up to 81.7%. We also show that workload-aware strategy selection is essential to improve performance-energy efficiency: MIG-based configurations are more effective for lightweight and regular workloads, offering up to 64.5% energy savings, while MPS better handles GPU-intensive workloads, achieving up to 71.1% EDP improvement. Finally, we demonstrate that optimized sharing strategies can reduce the required compute nodes by up to 75%, freeing system resources for concurrent workloads. Matheus M. Costa, Philippe Olivier Alexandre Navaux, Silvio Rizzi 0001, Arthur Francisco Lorenzon |
ICPP | 4 |
| 2025 | OLEO: Optimizing LEO Satellites Offloading of Cloud-Edge ApplicationsabstractLow Earth Orbit (LEO) satellite constellations enable cloud-edge computing for latency-sensitive applications. However, frequent satellite mobility challenges resource allocation and service continuity, leading to disruptions and inefficient provisioning. Existing strategies often overlook temporal constraints, resulting in frequent migrations and degraded performance. We propose OLEO, a heuristic strategy that optimizes application offloading by prioritizing satellites with higher exposure time, reducing unnecessary migrations and improving resource utilization. Experimental results show that OLEO provisions up to 1.5 X more application requests while reducing migrations by up to 20% in comparison baselines. Gabriel P. Costa, Diogo Matos, Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
ISCC | 4 |
| 2025 | Efficient Multi-Workload Execution for Sustainable GPU PerformanceabstractModern scientific research often relies on powerful computing systems that use graphics processing units (GPUs) to run complex applications. However, running these systems requires a large amount of energy, which contributes to carbon emissions and raises concerns about environmental impact. Given this scenario, we explore how sharing a single GPU between multiple applications can improve both performance and sustainability when running scientific workflows. We consider three execution strategies: running applications one after another, running two at the same time, and replacing finished tasks with new ones right away, using eighteen widely used scientific applications on three different GPUs (AMD MI250X, AMD RX 7900XT, and NVIDIA RTX 4090). To demonstrate that finding the best co-execution combination of applications improves resource efficiency, we use a mathematical approach based on linear programming to schedule which applications run together. Our results show that this approach can reduce total execution time by up to 47% and lower carbon emissions by as much as 34%, with minimal impact on the performance of individual applications. Additionally, when optimal combinations of parallel applications are used, the overall performance of a complete scientific workflow can improve by 36%, while carbon emissions are reduced by 25%. Matheus M. Costa, Philippe Olivier Alexandre Navaux, Silvio Rizzi 0001, O. E. Bronson Messer, Arthur Francisco Lorenzon |
SBAC-PAD | 5 |
| 2025 | Evaluating Code Portability for Carbon-Efficient RTM ComputingabstractAs GPU architectures continue to diversify across high-performance computing (HPC) systems, ensuring code portability and minimizing environmental impact have become critical challenges. This paper investigates how different programming models affect the carbon-performance efficiency of a Reverse Time Migration (RTM) application, a key workload in geophysics. We provide twelve implementations of the RTM code using CUDA, HIP, Kokkos, RAJA, and OpenMP Target, and evaluate their behavior on eleven GPUs from NVIDIA and AMD. Our analysis covers execution time, energy consumption, and carbon footprint. Results show that HIP, when tuned per architectures and RAJA, achieve the highest code portability in terms of carbon-efficiency, reaching up to 93.7% and 93.4% efficiency across all platforms, respectively. Overall, while HIP and CUDA deliver peak performance and the lowest emissions when properly optimized, the gap between these native models and high-level abstractions such as RAJA and Kokkos is steadily narrowing, indicating growing potential for portable and sustainable HPC development. Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux, Alexandre Sardinha, O. E. Bronson Messer |
SBAC-PAD | 1 |
| 2025 | Integration framework for online thread throttling with thread and page mapping on NUMA systems
Janaina Schwarzrock, Hiago Rocha, Arthur Francisco Lorenzon, Samuel Xavier de Souza, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 3 |
| 2025 | MAPER: mobility-aware energy-efficient container registry migrations for edge computing infrastructures
Daniel Chaves Temp, Alexandre A. F. da Costa, Ângelo Vieira, Ester S. Oribes, Ivan M. Lopes, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Fábio D. Rossi |
J. Supercomput. | 8 |
| 2024 | An ANN-Guided Multi-Objective Framework for Power-Performance Balancing in HPC SystemsabstractPower-performance efficiency has become one of the most critical issues in evolving High-Performance Computing systems (HPC) towards Exaflops. Thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are methods widely applied to better balance the power consumption and performance improvements of parallel applications. However, selecting ideal combinations of these knobs for every application is challenging due to the massive number of possible solutions, as there is no unique combination that delivers at the same time the best performance and the lowest power consumption. Given that, we propose HPC-PPO (power-performance optimizer), a multi-objective optimization strategy driven by an artificial neural network that leverages hardware and software features of parallel applications to predict Pareto-efficient configurations of TLP degree, DVFS, and UFS that optimize the balance between power and performance. When validating HPC-PPO on three multicore processors with twenty-five applications, we show that HPC-PPO can predict combinations very close to the best ones found by an exhaustive search. We also show that the Pareto-efficient configurations predicted by HPC-PPO improve parallel applications' performance by 30.7% while spending 23.9% less power when compared to state-of-the-art strategies. William Maas, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CF | 6 |
| 2024 | Balancing Performance and Aging in Cloud Environments
Thiago Gonçalves, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
CLOSER | 3 |
| 2024 | Harnessing Data Movement Strategies to Optimize Performance-Energy Efficiency of Oil & Gas Simulations in HPC
Pedro H. C. Rigon, Brenda S. Schussler, Alexandre Sardinha, Pedro M. Silva, Fábio Oliveira, Alexandre Carissimi, Jairo Panetta, Filippo Spiga, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux |
Euro-Par (2) | 9 |
| 2024 | DigiNet: Scaling up Provisioning of Network Digital TwinabstractThe pursuit of self-driving networks is increasing pressure on adopting intelligent, edge-based networking services. However, deploying autonomous network models within operational and large-scale infrastructures entails substantial risks that require rigorous verification and validation procedures. In this context, the application of a Network Digital Twin (NDT) is emerging as a viable approach towards intelligent network decision-making based on high-fidelity models built upon digital representations of physical network devices (i.e., Digital Twins). In this paper, we take the first steps towards efficiently provisioning NDT models. To that end, we introduce the Digital Twin Network Provisioning Problem (DigiNet), which encompasses the optimal placement of NDT models and the efficient collection of telemetry data for synchronizing NDT models with their physical counterparts. We theoretically formalize DigiNet as a Mixed-Integer Linear Programming (MILP) model and present a polynomial-time heuristic. Our results show that DigiNet outperforms baseline approaches by up to 10x regarding the number of NDT models provisioned. Marcelo Caggiani Luizelli, Francisco Germano Vogt, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Rodrigo N. Calheiros, Christian Esteve Rothenberg |
NetSoft | 4 |
| 2024 | Spinner: Enabling In-network Flow Clustering Entirely in a Programmable Data PlaneabstractData plane programmability is redesigning the way we manage and operate forwarding devices. However, most of the algorithmic decisions performed by data planes are still deterministic and control-plane dependent. We argue that it is possible to break this dependency and make the data plane intelligent, so that it can learn the infrastructure state autonomously. Despite existing efforts to make data planes intelligent, little has been done to design unsupervised ML algorithms that fit the architectural constraints of programmable devices. Executing such approaches in the data plane has the potential to reduce the overall decision-making time, thus meeting packet processing deadlines (which are in the order of nanoseconds). In this paper, we propose Spinner, the first effort to deliver an unsupervised Machine Learning (ML) approach entirely in programmable devices. Spinner is a flow clustering algorithm designed to fit existing architectural constraints of SmartNICs, and that can reach line rate for most packet sizes with complexity O(k). To demonstrate the potential behind in-network clustering, we prototype and deploy Spinner in a programmable testbed and use it to enhance Explicit Congestion Notifications (ECN) at the server side. Spinner-enhanced TCP provides up to 2x higher throughput when comparing to de-facto TCP implementations. Luigi Cannarozzo, Thiago Bortoluzzi Morais, Paulo Silas Severo de Souza, Leonardo Gobatto, Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, José Rodrigo Azambuja, Arthur Francisco Lorenzon, Fábio D. Rossi, Weverton Luis da Costa Cordeiro, Marcelo Caggiani Luizelli |
NOMS | 8 |
| 2024 | BTO, Block and Thread Optimization of GPU Kernels on Geophysical ExplorationabstractThe pursuit of performance and energy efficiency of geophysical exploration applications on high-performance computing (HPC) servers has been driving the optimization of hardware resource usage in graphic processing units (GPUs). On such architectures, the execution configuration of each kernel (e.g., the number of blocks and threads per block) plays an essential role in the performance and energy consumption of these applications. However, as we show in this paper, due to the massive number of possible configurations of the number of blocks and threads per block, leveraging solely on the software developer to define the configuration execution for every GPU kernel does not lead to an ideal usage of GPU hardware resources, leading to performance loss and an increase on the energy consumption. To tackle this challenge, we propose BTO, a block and thread optimization strategy driven by a genetic algorithm. It cooperatively optimizes the number of blocks and thread per block for every GPU kernel at runtime with minimum convergence overhead regardless of the number of GPUs available on the system. When employing BTO to optimize the Fletcher modeling, a representative geophysical exploration application, on different AMD and NVIDIA GPUs, we show that BTO improves the energy-delay product (EDP - tradeoff between performance and energy) by up to 83.8% and 81.9% over the standard execution of Fletcher and the default execution of GPU applications on the target architectures. Moreover, by comparing it to an exhaustive search, we show that BTO converges to optimal execution configurations in 96.4% of all evaluated scenarios. Brenda S. Schussler, Pedro H. C. Rigon, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux |
PDP | 3 |
| 2024 | Towards Performance Portability of an Oil and Gas Application on Heterogeneous ArchitecturesabstractWith the widespread use of graphic processing units (GPUs) from different vendors in oil and gas companies, it has become increasingly important to achieve code portability. This allows for evaluating performance across diverse GPU vendors, enabling informed decisions. This capability enables companies to assess hardware solutions based on cost, performance, or energy efficiency, promoting competition and potentially driving innovation and cost reduction in GPU technologies. With that in mind, we address the challenge of achieving performance portability for Fletcher, an anisotropic wave propagation modeling application across various GPU architectures from NVIDIA and AMD. Hence, our first contribution is the parallel implementation of Fletcher on GPUs using eleven variations of portable programming models, including HIP, CUDA, Kokkos, RAJA, and OpenMP Target. Our second contribution is a comprehensive evaluation of these implementations across five generations of GPU from NVIDIA and AMD, assessing performance, performance-portability, and power-performance efficiency. Through an extensive set of experiments, we demonstrate that HIP outperforms the assessed programming models in terms of performance portability across all evaluated GPUs. Specifically, it performs 7.9% better than Kokkos, 8.8% better than RAJA, and 67.8% better than OpenMP Target. We also show that while the automatic translation of CUDA to HIP code allows for execution on AMD GPUs, optimizing for high performance is necessary. Additionally, we show that the programming models’ impact on Fletcher’s performance is heavily influenced by the size of the GPU’s L2 cache, especially given that this is a memory-intensive application. Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux, Alexandre Sardinha, O. E. Bronson Messer |
SBAC-PAD | 1 |
| 2024 | HBPB, applying reuse distance to improve cache efficiency proactively
Arthur M. Krause, Paulo C. Santos 0001, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux |
J. Parallel Distributed Comput. | 3 |
| 2024 | A neural network framework for optimizing parallel computing in cloud servers
Everton Camargo de Lima, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Arthur Francisco Lorenzon |
J. Syst. Archit. | 5 |
| 2024 | Allok: a machine learning approach for efficient graph execution on CPU-GPU clusters
Marcelo K. Moori, Hiago Rocha, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
J. Supercomput. | 3 |
| 2024 | Synergistically Rebalancing the EDP of Container-Based Parallel ApplicationsabstractThe use of containers has become standard in cloud environments. However, many parallel applications in containers will not present gains proportional to the extra available hardware. This inefficient use of hardware naturally leads to energy consumption waste. With that in mind, we proposeTT-Autoscaling. It works at two different levels: a) in the container, by automatically and transparently tuning the number of threads at runtime of the application, in a way to optimize the trade-off between energy and performance; b) in the cloud infrastructure, by smartly transferring the released resources to other containers that may run in parallel, making better use of the available resources. We compareTT-Autoscalingto the default execution of containers (serial execution with the maximum number of threads), showing 55.8% of performance improvements, 53.6% of energy reductions, and 79.5% of EDP improvements. We also show thatTT-Autoscalingoutperforms strategies that apply vertical autoscalers proposed by orchestrator tools. Vinicius S. da Silva, Everton Camargo de Lima, Janaina Schwarzrock, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | Towards Optimizing the Edge-to-Cloud Continuum Resource Allocation
Igor Ferrazza Capeletti, Ariel Góes de Castro, Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
CLOSER | 5 |
| 2023 | Latency-Aware Cost-Efficient Provisioning of Composite Applications in Multi-Provider Clouds
Daniel Chaves Temp, Igor Ferrazza Capeletti, Ariel Góes de Castro, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
CLOSER | 5 |
| 2023 | Taking Detours: An In-Network Fault-Tolerant Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel network monitoring approach mainly fostered by programmable network devices. Despite existing efforts toward the orchestration of INT, little has yet been done to provide fault-tolerant mechanisms in the data plane (e.g., to address hardware failure). In this paper, we introduce InPatching - an in-network approach to fast recover INT-based monitoring from network link failures. InPatching is implemented in the data plane and allows the application of detours in an autonomous and coordinated manner without the control plane intervention. To provide efficient detours to INT solutions, we formalize the fault-tolerant probing planning for INT by means of a MILP (Mixed-Integer Linear Programming) model. We prototype InPatching in P4 and we show that it can recover from fault conditions much faster than control plane solutions (up to 18X), while not imposing substantial overhead. Ariel Góes de Castro, Igor Capelletti, Fábio D. Rossi, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli |
ICC | 4 |
| 2023 | Automatic CPU-GPU Allocation for Graph ExecutionabstractAlthough advances in modern GPUs have accelerated the execution of heavy data processing applications, speeding up graph processing on these systems is not a trivial task: graph applications are characterized by their high volume of irregular memory access that varies with the graph structure so that they do not reach their peak performance when executing on GPUs in many times. In these cases, the CPU execution is more suitable. Given that graph structures can be identified through high-level metrics (e.g., diameter and average clustering coefficient), they may assist the designer in deciding where to execute a given input graph (GPU or CPU). Based on that, in this work, we propose GraCo: a graph processing framework to help the decision-making on where to process a batch of graph applications. Whenever a new batch is submitted to the target HPC system, GraCo decides the best machine to execute each application based only on the available high-level features, precluding any additional applications' execution. Our experimental results comparing GraCo with three other strategies executed on an HPC system comprised of 4 CPUs and 3 GPUs showed that GraCo outperforms the other strategies by at least 34.94×, 13.59×, and 492.31× in total execution time, energy, and energy-delay product. Marcelo K. Moori, Hiago Rocha, Matheus A. Silva, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 5 |
| 2023 | NeurOPar, A Neural Network-Driven EDP Optimization Strategy for Parallel WorkloadsabstractThe pursuit of energy efficiency has been driving the development of techniques to optimize hardware resource usage in high-performance computing (HPC) servers. On multicore architectures, thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are three popular methods applied to improve the trade-off between performance and energy consumption, represented by the energy-delay product (EDP). However, the complexity of selecting the optimal configuration (TLP degree, DVFS, and UFS) for each application poses a challenge to software developers and end-users due to the massive number of possible configurations. To tackle this challenge, we propose NeurOpar, an optimization strategy for parallel workloads driven by an artificial neural network (ANN). It uses representative hardware and software metrics to build and train an ANN model that predicts combinations of thread count and core/uncore frequency levels that provide optimal EDP results. Through experiments on four multicore processors using twenty-five applications, we demonstrate that NeurOPar predicts combinations that yield EDP values close to the best ones achieved by an exhaustive search and improve the overall EDP by 42% compared to the default execution of HPC applications. We also show that NeurOPar can enhance the execution of parallel applications without incurring the performance and energy penalties associated with online methods by comparing it with two state-of-the-art strategies. Cristiano A. Künas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
SBAC-PAD | 6 |
| 2023 | Improving the efficiency of graph algorithm executions on high-performance computingabstractSummary The growing need for extracting information from large graphs has been pushing the development of parallel graph algorithms. However, the highly irregular structure of the real‐world graphs limits the performance and energy improvements of graph applications. In this paper, we show that, in most cases, using all the available cores of the multiprocessor is not the best option in terms of the aforementioned non‐functional requirements. Based on that, we proposeGraphKat, a framework that enables the simultaneous processing of several algorithms/graphs instead of executing them serially (i.e., one after another), increasing efficiency in terms of performance and energy.GraphKatworks in two steps: (i) it characterizes the graph applications with a specific number of threads based on their efficiency levels; and (ii) it defines the execution order of all graph applications in the target system. Experimental results on three multicore processors (Intel and AMD) show thatGraphKatimproves the overall system's efficiency related to performance (up to ) and energy‐saving (up to 245.21), and reduces the graph applications' execution time (up to ) and energy consumption (up to 6.64) compared to the default execution of parallel applications on HPC systems. Marcelo K. Moori, Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Mitigating execution unit contention in parallel applications using instruction-aware mappingabstractSummary Parallel applications running on simultaneous multithreading (SMT) processors naturally compete for execution units when their threads are mapped to the same core. This issue is further aggravated when such threads execute similar instructions that stress the same execution unit type, making their execution to behave very similarly as if the threads were running sequentially. This, in turn, will lead to performance degradation and underutilization of hardware resources. This work proposes a completely transparent framework (no modifications to the source code are necessary) that automatically maps threads of multiple parallel applications on SMT processors. The framework focuses on improving performance by mitigating the contention on execution units, considering each thread's instruction types, which are detected at runtime by our framework. Results show performance gains of 21% (geometric mean), compared to the native scheduler of the operating system. Matheus S. Serpa, Eduardo Henrique Molina da Cruz, Matthias Diener, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 10 |
| 2023 | Mobility-Aware Registry Migration for Containerized Applications on Edge Computing Infrastructures
Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
J. Netw. Comput. Appl. | 3 |
| 2022 | Towards Efficient Selective In-Band Network Telemetry Report Using SmartNICs
Ronaldo Canofre, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 3 |
| 2022 | Multivariate Interpolation at the Edge to Infer Faulty IoT Sensor Metrics
Marcos Paulo Konzen, Patric Lincoln Ramires Izolan, Fábio Júnior Griesang, Paulo Silas Severo de Souza, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Júlio C. B. de Mattos, Cinara Ewerling da Rosa, Fábio D. Rossi |
CLOSER | 6 |
| 2022 | Using machine learning to optimize graph execution on NUMA machinesabstractThis paper proposes PredG, a Machine Learning framework to enhance the graph processing performance by finding the ideal thread and data mapping on NUMA systems. PredG is agnostic to the input graph: it uses the available graphs' features to train an ANN to perform predictions as new graphs arrive - without any application execution after being trained. When evaluating PredG over representative graphs and algorithms on three NUMA systems, its solutions are up to 41% faster than the Linux OS Default and the Best Static - on average 2% far from the Oracle -, and it presents lower energy consumption. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
DAC | 3 |
| 2022 | Seamless optimization of the GEMM kernel for task-based programming modelsabstractThe general matrix-matrix multiplication (GEMM) kernel is a fundamental building block of many scientific applications. Many libraries such as Intel MKL and BLIS provide highly optimized sequential and parallel versions of this kernel. The parallel implementations of the GEMM kernel rely on the well-known fork-join execution model to exploit multi-core systems efficiently. However, these implementations are not well suited for task-based applications as they break the data-flow execution model. In this paper, we present a task-based implementation of the GEMM kernel that can be seamlessly leveraged by task-based applications while providing better performance than the fork-join version. Our implementation leverages several advanced features of the OmpSs-2 programming model and a new heuristic to select the best parallelization strategy and blocking parameters based on the matrix and hardware characteristics. When evaluating the performance and energy consumption on two modern multi-core systems, we show that our implementations provide significant performance improvements over an optimized OpenMP fork-join implementation, and can beat vendor implementations of the GEMM (e.g., Intel MKL and AMD AOCL). We also demonstrate that a real application can leverage our optimized task-based implementation to enhance performance. Arthur Francisco Lorenzon, Sandro Matheus V. N. Marques, Antoni C. Navarro, Vicenç Beltran 0001 |
ICS | 1 |
| 2022 | DyPro: Dynamic Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel net-work monitoring mechanism that improves fine-grained net-work visibility. Despite the increasing research efforts towards the orchestration of INT data acquisition, little has yet been done to efficiently collect telemetry data from the network considering monitoring applications requirements. In this paper, we introduce DyPro - a dynamic probing planning for INT. In particular, DyP ro ensures that telemetry dependencies are always satisfied by monitoring application requirements. We theoretically formalize it as a Mixed-Integer Linear Programming (MILP) optimization model and propose a heuristic procedure to efficiently solve it. Results show that DyP ro can outperform state-of-the-art solutions by up to 5x regarding the percentage of monitoring applications satisfied. Leandro M. Dallanora, Ariel Góes de Castro, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 5 |
| 2022 | Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 8 |
| 2021 | The Actual Cost of Programmable SmartNICs: Diving into the Existing Limits
Pablo B. Viegas, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 3 |
| 2021 | Synergically Rebalancing Parallel Execution via DCT and Turbo BoostingabstractThe increasing use of cloud and HPC systems put more pressure on the efficient utilization of hardware resources to keep costs low. Many dynamic concurrency throttling (DCT) techniques have successfully used to tune the number of executing threads to better balance a parallel application according to its available scalability. Similarly, boosting frequency strategies have been used to speed up the sequential parts’ execution. Given that, we propose Poseidon, the first transparent and automatic approach that cooperatively exploits both techniques to rebalance OpenMP applications without any preprocessing, with no code transformation, recompilation, or OS modification. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
DAC | 6 |
| 2021 | Combining Dynamic Concurrency Throttling with Voltage and Frequency Scaling on Task-based Programming ModelsabstractBeing on the verge of exascale performance has shifted the prioritization of performance in applications to the inclusion of power-performance efficiency as a primary objective in the High Performance Computing (HPC) community. Simultaneously, this has surfaced hardware and software efforts that employ techniques such as dynamic voltage and frequency scaling (DVFS) for core and uncore units or dynamic concurrency throttling (DCT) to exploit hardware resources efficiently, by saving energy while maintaining performance. These techniques are complementary, so they can be used together. However, employing them is not a straightforward task, as they have to be adjusted based on the workload, and it is even more complex to combine them properly. Thus, these techniques should be applied transparently by a runtime system, without relying on application developers. In this paper, we extend a task-based runtime system with an infrastructure that categorizes workloads based on their computational profile – memory-bounded, compute-bounded, or balanced. This categorization is done in an on-line manner and with a negligible overhead. With this additional information, we enhance the CPU-manager and scheduler of OmpSs-2, a task-based parallel programming model, to automatically combine DVFS and DCT techniques based on workloads. Moreover, we show that our heuristics transparently improve energy efficiency on average by 15% with no significant performance loss and either equal or surpass the energy efficiency of the best static configuration available. Antoni Navarro Muñoz, Arthur Francisco Lorenzon, Eduard Ayguadé, Vicenç Beltran 0001 |
ICPP | 2 |
| 2021 | Combining Thread Throttling and Mapping to Optimize the EDP of Parallel ApplicationsabstractThread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1. Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 8 |
| 2021 | Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo BoostingabstractWith the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 8 |
| 2021 | Boosting Graph Analytics by Tuning Threads and Data Affinity on NUMA SystemsabstractThe execution of large real-world graphs, such as web searches and social networks, has been boosting by modern HPC systems. However, their irregular communication patterns and poor data locality impose many challenges, mainly when executed on NUMA systems. As we show in this paper, there is no one-fits-all configuration for threads/data mapping, and the best combination will vary according to the NUMA system, graph algorithm, and input graph at hand. Based on that, we propose Graphith: a framework that automatically enhances graph processing performance by adapting its execution considering the variables mentioned above. Graphith also goes one step further and improves the existing policies: it uses a Genetic Algorithm to fine-tune the thread-to-core allocation combined with data mapping policies. With that, Graphith improves in 21%, on average, the default execution, and is, on average, 7% better than the best possible combination of standard policies. Hiago Rocha, Janaina Schwarzrock, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
PDP | 3 |
| 2021 | Mitigating the processor aging through dynamic concurrency throttling
Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Parallel Distributed Comput. | 6 |
| 2021 | Low learning-cost offline strategies for EDP optimization of parallel applications
Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Samuel Xavier de Souza, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 6 |
| 2021 | A Runtime and Non-Intrusive Approach to Optimize EDP by Tuning Threads and CPU Frequency for OpenMP ApplicationsabstractEfficiently exploiting thread-level parallelism has been challenging. Many parallel applications are not sufficiently balanced or CPU-bound to take advantage of the increasing number of cores and the highest possible operating frequency. Moreover, many variables may change according to the system (input set, microarchitecture, and number of cores) or during execution, influencing each parallel region in different ways. Therefore, the task of rightly choosing the ideal configuration (number of threads and DVFS) for each parallel region to deliver the best Energy-Delay Product (EDP) is not straightforward. While the significant number of variables prevents the use of exhaustive search methods, the changing nature of the problem precludes offline strategies. Few solutions are online and synergistically consider thread throttling and DVFS. However, they lack transparency (demand changes in the original code) and/or adaptability (do not automatically adjust to applications at run-time). Our proposed Hoder covers all the characteristics above, optimizing at run-time any dynamically linked OpenMP application, without requiring any code transformation or recompilation. We show Hoder's efficiency by comparing it to two exhaustive offline and two online search approaches, three state-of-the-art techniques, and regular OpenMP execution, considering different setups (Intel 44-, 16- and 12-core; AMD 8- and 12-core). Janaina Schwarzrock, Charles Cardoso De Oliveira, Marcus Ritt, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | A Heuristic Approach for Large-Scale Orchestration of the In-band Data Plane Telemetry Problem
Rumenigue Hohemberger, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA | 2 |
| 2020 | Enhancing Resource Management Through Prediction-Based Policies
Antoni C. Navarro, Arthur Francisco Lorenzon, Eduard Ayguadé, Vicenç Beltran 0001 |
Euro-Par | 2 |
| 2020 | Decreasing the Learning Cost of Offline Parallel Application Optimization StrategiesabstractMany parallel applications do not scale as the number of threads increases, which means that executing them with the maximum possible number of threads will not always deliver the best outcome in performance, energy consumption, or the tradeoff between both (represented by the energy-delay product- EDP). Given that, several strategies, online and offline, have already been proposed to rightly tune the number of threads according to the application. While the former can capture some behaviors that can only be known at runtime, the latter do not impose any execution overhead and can use more efficient and costly algorithms. However, these learning algorithms in static strategics may take several hours, precluding their use or a smooth migration across different systems. In this scenario, we propose a generic methodology for such offline strategies to significantly decrease the learning time by inferring the execution behavior of parallel applications using smaller input sets than the ones used by the target applications. Through the execution of eighteen well-known benchmarks on two multicore processors, we show that our methodology is capable of converging to results that are very close to those that use the regular input set, but converging 84.7% faster, on average. We also show that such a strategy delivers better results than a dynamic one, presenting an EDP 7.7% lower, on average, when executing the applications with the number of threads found during learning. Finally, we also compare our learning methodology with an exhaustive search. It has an average learning cost (i.e., the time spent by our search algorithm to find the best configuration) of only 3.1% to optimize the EDP of the entire benchmark set1. Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 5 |
| 2020 | Modeling and Simulating Daily Power Budgets for Sustainable Data CentersabstractA novel energy-efficient scenario that makes possible to maintain sustainable data centers to feed part of resources through renewable energy sources has emerged. As renewable energies are accumulated in the form of power budgets, data centers must adapt a slice of the processing resources required to meet applications at those limits. This work is modeling and simulating the computing capacity of a data center according to daily power budgets from different sources of renewable energy. The results showed that based on the daily energy harvest of today's renewable energy sources, intelligent resource orchestration could use such energy so that up to 40% of what is needed to maintain a quality-of-service data center comes from non-polluting sources. Rumenigue Hohemberger, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
PDP | 2 |
| 2019 | The Impact of Parallel Programming Interfaces on the Aging of a Multicore Embedded ProcessorabstractIn order to meet the increasing performance demand of applications, the amount of cores in a single chip package has been increasing. However, the heat has been rising at a higher scale, which accelerates the aging process in modern processors. Therefore, wisely balancing the use of resources is important to extend its longevity. Frequency performance stagnates after a certain amount of concurrent threads starts executing. In such cases, the only result is a temperature rise that directly influences the aging process, reducing the processor lifetime. This unbalance between threads can be originated from many factors, which includes the way threads communicate and synchronize. Considering that those characteristics are related to the Parallel Programming Interface (PPI) used to parallelize the application, this work proposes to evaluate three widely used PPIs executing on an embedded multicore. We show that, depending on the characteristic of the application, by only switching from one PPI to another, it is possible to reduce the effects of aging. For that, we have developed a model based on the Arrhenius equation. We show that OpenMP has a lower impact on the processor aging for memory-bound applications: up to 38% and 68% lower than PThreads and MPI, respectively. On the other hand, PThreads presents the lowest impact on the processor aging for CPU-bound applications. Ângelo Vieira, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Fábio D. Rossi, Jorji Nonaka |
ISCAS | 7 |
| 2019 | Multilevel resource allocation for performance-aware energy-efficient cloud data centersabstractThe massive power consumption of data centers has been a recurring concern in current research. In cloud environments, lots of methods are being adopted that aim for energy efficiency. However, although such methods enable the decrease in power consumption, they regularly affect application performance. In this paper, we present a multilevel resource allocation approach towards dynamic network bandwidth at the physical substrate, managing different power-saving states and workload allocation at the cloud infrastructure at the same time employ virtual machine allocation and selection policies at the cloud platform. In order to evaluate our approach, tests were carried out in a simulated environment using scale-out application on a dynamic cloud infrastructure. Results showed that our proposal presents a better balance regarding a more energy-efficient data center with a smaller impact on application performance when compared with other works discussed in the literature. Fábio D. Rossi, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 6 |
| 2019 | Transparent Aging-Aware Thread ThrottlingabstractTo satisfy the rising performance demands of modern applications, the number of cores in a single chip package has been increasing. However, the power dissipated and temperature have been growing at a higher rate, accelerating the aging process of new processors. Considering that a significant number of parallel applications are unbalanced, in many cases performance stagnates after a certain number of concurrent threads starts executing. In such cases, the only outcome is a temperature rise on the processor, which drastically accelerates aging. Given that, we propose an automatic and transparent approach to reduce the processor aging by automatically tuning the number of threads for OpenMP applications at run-time. Our tool, Geras, is entirely transparent to the end-user, so even already compiled binaries can be optimized. Through the execution of twelve well-known benchmarks on two multicore platforms, we show that Geras can improve the processor lifetime by up to 83% and 89% over the standard OpenMP execution and its built-in feature that dynamically adjusts the number of threads, respectively. We also show that Geras outperforms techniques that target performance or energy, which reinforces the need for a specific tool that optimizes aging1. Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
SBAC-PAD | 6 |
| 2019 | The Impact of Turbo Frequency on the Energy, Performance, and Aging of Parallel ApplicationsabstractTechnologies that improve the performance of parallel applications by increasing the nominal operating frequency of processors respecting a given TDP (Thermal Design Power) have been widely used. However, they may impact on other non-functional requirements in different ways (e.g. increasing energy consumption or aging). Therefore, considering the huge number of configurations available, represented by the range of all possible combinations among different parallel applications, amount of threads, dynamic voltage and frequency scaling (DVFS) governors, boosting technologies and simultaneous multithreading (SMT), selecting the one that offers the best tradeoff for a non-functional requirement is extremely challenging for software designers. Given that, in this work we assess the impact of changing these configurations on the energy consumption, performance, and aging of parallel applications on a turbo-compliant processor. Results show that there is no single configuration that would provide the best solution for all nonfunctional requirements at once. For instance, we demonstrate that the configuration that offers the best performance is the same one that has the worst impact on aging, accelerating it by up to 1.75 times. With our experiments, we provide guidelines for the developer when it comes to tuning performance using turbo boosting to save as much energy as possible and increase the lifespan of the hardware components. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Alessandro Girardi, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
VLSI-SoC | 7 |
| 2019 | Aurora: Seamless Optimization of OpenMP ApplicationsabstractEfficiently exploiting thread-level parallelism has been challenging for software developers. As many parallel applications do not scale with the number of cores, the task of rightly choosing the ideal amount of threads to produce the best results in performance or energy is not straightforward. Moreover, many variables may change according to the system at hand (e.g., application, input set, microarchitecture, number of cores) and even during execution. Existing solutions lack transparency (demand changes in the original code) or adaptability (do not automatically adjust to applications at run-time). In this scenario, we propose Aurora, an OpenMP framework that is completely transparent to both the designer and end-user. Without any code transformation or recompilation, it is capable of automatically finding, at run-time and with minimum overhead, the optimal number of threads for each parallel loop region and re-adapt in cases the behavior of a region changes during execution. When executing fifteen well-known benchmarks on four multi-core processors, Aurora improves the Energy-Delay Product by up to 98, 86 and 91 percent over the standard OpenMP execution, the OpenMP feature that dynamically adjusts the number of threads, and the Feedback-Driven Threading, respectively. Arthur Francisco Lorenzon, Charles Cardoso De Oliveira, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Adaptive and polymorphic VLIW processor to optimize fault tolerance, energy consumption, and performanceabstractBecause most traditional homogeneous and heterogeneous processors have a fixed design that limits its runtime adaptability, they are not able to cope with the varying application behavior when one considers the axes of fault tolerance, performance, and energy consumption altogether. In this context, we propose a new dynamically adaptive processor design that is capable of delivering the best trade-off among these three axes according to the application at hand, or be tuned to optimize a specific metric. This is achieved by extending a polymorphic processor that can change its issue-width during runtime with specific mechanisms for fault tolerance, energy optimization, and performance enhancement. They are controlled by an optimization algorithm that evaluates and chooses which is the best configuration according to given requirements. Considering a metric that weighs all three axes, the proposed adaptive processor delivers a result that is 94.88% of the oracle processor on average, while a static configuration (defined at design time without runtime adaptation) only achieves 28.24% at most, which means that dynamic adaptation is required to cope with different application behaviors as there is not one specific configuration that fits all applications. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Sandip Kundu, Israel Koren, Antonio Carlos Schneider Beck |
CF | 2 |
| 2017 | LAANT: A library to automatically optimize EDP for OpenMP applicationsabstractEfficiently exploiting thread level parallelism from new multicore systems has been challenging for software developers. While blindly increasing the number of threads may lead to performance gains, it can also result in disproportionate increase in energy consumption. For this reason, rightly choosing the number of threads is essential to reach the best compromise between both. However, such task is extremely difficult: besides the huge number of variables involved, many of them will change according to different aspects of the system at hand and are only possible to be defined at run-time. To address this complex scenario, we propose LAANT, a novel library to automatically find the optimal number of threads for OpenMP applications, by dynamically considering their characteristics, input set, and the processor architecture. By executing nine well-known benchmarks on three real multicore processors, LAANT improves the EDP (Energy-Delay Product) by up to 61%, compared to the standard OpenMP execution; and by 44%, when the dynamic adjustment of the number of threads of OpenMP is activated. Arthur Francisco Lorenzon, Jeckson Dellagostin Souza, Antonio Carlos Schneider Beck |
DATE | 1 |
| 2017 | Improving EDP in multi-core embedded systems through multidimensional frequency scalingabstractEnergy saving management in multi-core embedded environments has been a challenge for designers. To achieve energy efficiency, most studies consider dynamic frequency scaling on one hardware component only, such as processor or memory - which will most likely also affect performance. This work proposes the use of frequency scaling considering the three most important hardware components altogether: processors, L2 cache, and RAM; seeking for the best set of frequencies for each one of them to improve the Energy-Delay Product (EDP), depending on the application's behavior. Therefore, this work addresses multidimensional frequency scaling for multi-core embedded systems. By evaluating different frequency levels, we show that the EDP can be improved in up to 46.4% when compared to the standard way that the frequencies are configured. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Mateus B. Rutzig, Fábio D. Rossi |
ISCAS | 3 |
| 2016 | Exploiting Idle Hardware to Provide Low Overhead Fault Tolerance for VLIW ProcessorsabstractBecause of technology scaling, the soft error rate has been increasing in digital circuits, which affects system reliability. Therefore, modern processors, including VLIW architectures, must have means to mitigate such effects to guarantee reliable computing. In this scenario, our work proposes three low overhead fault tolerance approaches based on instruction duplication with zero latency detection, which uses a rollback mechanism to correct soft errors in the pipelanes of a configurable VLIW processor. The first uses idle issue slots within a period of time to execute extra instructions considering distinct application phases. The second works at a finer grain, adaptively exploiting idle functional units at run-time. However, some applications present high instruction-level parallelism (ILP), so the ability to provide fault tolerance is reduced: less functional units will be idle, decreasing the number of potential duplicated instructions. The third approach attacks this issue by dynamically reducing ILP according to a configurable threshold, increasing fault tolerance at the cost of performance. While the first two approaches achieve significant fault coverage with minimal area and power overhead for applications with low ILP, the latter improves fault tolerance with low performance degradation. All approaches are evaluated considering area, performance, power dissipation, and error coverage. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Luigi Carro, Fernanda Lima Kastensmidt, Stephan Wong, Antonio Carlos Schneider Beck |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | Investigating different general-purpose and embedded multicores to achieve optimal trade-offs between performance and energy
Arthur Francisco Lorenzon, Márcia C. Cera, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 1 |
| 2015 | The Influence of Parallel Programming Interfaces on Multicore Embedded SystemsabstractThread-Level Parallelism (TLP) exploitation for embedded systems has been a challenge for software developers: while it is necessary to take advantage of the availability of multiple cores, it is also mandatory to consume less energy. To speed up the development process and make it as transparent as possible, software designers use Parallel Programming Interfaces (PPIs). However, as will be shown in this paper, each PPI implements different ways to exchange data using shared memory regions, influencing performance, energy consumption and Energy-Delay Product (EDP), which varies across different embedded processors. By evaluating four PPIs and three multicore processors (ARM A8, A9 and Intel Atom), we demonstrate that by simply switching PPI it is possible to save up to 59% in energy consumption and achieve up to 85% of EDP improvements, in the most significant case. We also show that the efficiency (i.e., The best possible use of the available resources) decreases as the number of threads increases in almost all cases, but at distinct rates. Arthur Francisco Lorenzon, Anderson Luiz Sartor, Márcia C. Cera, Antonio Carlos Schneider Beck |
COMPSAC | 1 |
| 2015 | The Impact of Virtual Machines on Embedded SystemsabstractEmbedded systems are becoming increasingly complex and, due to their tight energy requirements, all the available resources must be used in the best possible way. However, Android, the most used software platform for embedded systems, features a virtual machine to run applications. Even though it ensures flexibility so the application can execute on different underlying architectures without the need for recompilation, it burdens the system because of the introduction of an extra software layer. Considering this scenario, through the development of an extension of the Android QEMU emulator and a specific benchmark set, this work evaluates the significance of the virtual machine by comparing applications written in Java and in native language. We show that, given a fixed energy budget, a different amount of applications can be executed depending the way they were implemented. We also demonstrate that this difference varies according to the processor, by executing the applications on all officially supported Android architectures (Intel x86, ARM, and MIPS). Therefore, even though the Virtual Machine provides total transparency to the software developer, he/she must be aware of it and the underlying target micro architecture at early designs stages so as to build a low-energy application. Anderson Luiz Sartor, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck |
COMPSAC | 2 |
| 2015 | On the influence of static power consumption in multicore embedded systemsabstractEnergy consumption in multicore embedded systems has become a constant concern. Thread-Level Parallelism exploitation may reduce energy consumption because it saves static power consumption of the processor, since performance is obtained. However, as will be shown in this paper, the influence of the static power on the energy consumption and Energy-Delay Product will depend on how significant it is in the processor. By evaluating different levels of static power in the respect to the total power consumption in two embedded processors (ARM and Atom), we demonstrate that if the right value of static power consumption is tuned during the designing and manufacturing, it is possible to save up 35% in energy consumption and achieve up to 20% of improvements in the EDP efficiency (i.e., the best possible use of the available resources). We also show that the more communication the parallel application has, the lower is the impact of static power of the processor in the total energy consumption. Arthur Francisco Lorenzon, Márcia C. Cera, Antonio Carlos Schneider Beck |
ISCAS | 1 |