VLDB 2026 Research / reviewers in the wild / expert
Fábio D. Rossi
dblp:128/5117 · also Fábio Diniz Rossi
· DBLP profile ↗
59ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0002-2450-1024ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 11 since 2021Computer networks · 9 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrustEdge: Failure-Aware Orchestration for Edge Application Provisioning
Marcos Paulo Konzen, Paulo Silas Severo de Souza, Fábio D. Rossi, Júlio C. B. de Mattos |
CLOSER | 3 |
| 2026 | Why Large Language Models Struggle with Cloud Instance Selection
Matheus Machado, Matheus M. Costa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Arthur Francisco Lorenzon |
CLOSER | 4 |
| 2026 | Mininet-AI: Emulating Network Agentic AI
Pedro da Silva Santiago, Diogo Mainart Monteiro, Victor Hugo Schneider Lopes, Francisco G. Vogt, Fábio D. Rossi, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli |
NetSoft | 5 |
| 2026 | Event-Driven Neuromorphic-Inspired Task Offloading for Energy-Efficient Edge-IoT Systems
Pedro Henrique Sachete Garcia, Fábio D. Rossi |
J. Grid Comput. | 2 |
| 2025 | Energy-Aware Node Selection for Cloud-Based Parallel Workloads with Machine Learning and Infrastructure as Code
Denis B. Citadin, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 2 |
| 2025 | WFQ-Based SLA-Aware Edge Applications Provisioning
Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 5 |
| 2025 | LLM-Based Adaptive Digital Twin Allocation for Microservice Workloads
Pedro Henrique Sachete Garcia, Ester S. Oribes, Ivan Mangini Lopes Júnior, Braulio Marques de Souza, Ângelo Vieira, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Paulo Silas Severo de Souza, Fábio D. Rossi |
CLOSER | 9 |
| 2025 | Towards Optimizing Cost and Performance for Parallel Workloads in Cloud Computing
William Maas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CLOSER | 2 |
| 2025 | OLEO: Optimizing LEO Satellites Offloading of Cloud-Edge ApplicationsabstractLow Earth Orbit (LEO) satellite constellations enable cloud-edge computing for latency-sensitive applications. However, frequent satellite mobility challenges resource allocation and service continuity, leading to disruptions and inefficient provisioning. Existing strategies often overlook temporal constraints, resulting in frequent migrations and degraded performance. We propose OLEO, a heuristic strategy that optimizes application offloading by prioritizing satellites with higher exposure time, reducing unnecessary migrations and improving resource utilization. Experimental results show that OLEO provisions up to 1.5 X more application requests while reducing migrations by up to 20% in comparison baselines. Gabriel P. Costa, Diogo Matos, Pedro Henrique Sachete Garcia, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
ISCC | 5 |
| 2025 | Toward real-time IoT multi-sensor data orchestration on wireless sensor networks
Pedro Henrique Sachete Garcia, Marcelo Caggiani Luizelli, Fábio D. Rossi |
J. Supercomput. | 3 |
| 2025 | MAPER: mobility-aware energy-efficient container registry migrations for edge computing infrastructures
Daniel Chaves Temp, Alexandre A. F. da Costa, Ângelo Vieira, Ester S. Oribes, Ivan M. Lopes, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Fábio D. Rossi |
J. Supercomput. | 9 |
| 2024 | An ANN-Guided Multi-Objective Framework for Power-Performance Balancing in HPC SystemsabstractPower-performance efficiency has become one of the most critical issues in evolving High-Performance Computing systems (HPC) towards Exaflops. Thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are methods widely applied to better balance the power consumption and performance improvements of parallel applications. However, selecting ideal combinations of these knobs for every application is challenging due to the massive number of possible solutions, as there is no unique combination that delivers at the same time the best performance and the lowest power consumption. Given that, we propose HPC-PPO (power-performance optimizer), a multi-objective optimization strategy driven by an artificial neural network that leverages hardware and software features of parallel applications to predict Pareto-efficient configurations of TLP degree, DVFS, and UFS that optimize the balance between power and performance. When validating HPC-PPO on three multicore processors with twenty-five applications, we show that HPC-PPO can predict combinations very close to the best ones found by an exhaustive search. We also show that the Pareto-efficient configurations predicted by HPC-PPO improve parallel applications' performance by 30.7% while spending 23.9% less power when compared to state-of-the-art strategies. William Maas, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
CF | 4 |
| 2024 | DigiNet: Scaling up Provisioning of Network Digital TwinabstractThe pursuit of self-driving networks is increasing pressure on adopting intelligent, edge-based networking services. However, deploying autonomous network models within operational and large-scale infrastructures entails substantial risks that require rigorous verification and validation procedures. In this context, the application of a Network Digital Twin (NDT) is emerging as a viable approach towards intelligent network decision-making based on high-fidelity models built upon digital representations of physical network devices (i.e., Digital Twins). In this paper, we take the first steps towards efficiently provisioning NDT models. To that end, we introduce the Digital Twin Network Provisioning Problem (DigiNet), which encompasses the optimal placement of NDT models and the efficient collection of telemetry data for synchronizing NDT models with their physical counterparts. We theoretically formalize DigiNet as a Mixed-Integer Linear Programming (MILP) model and present a polynomial-time heuristic. Our results show that DigiNet outperforms baseline approaches by up to 10x regarding the number of NDT models provisioned. Marcelo Caggiani Luizelli, Francisco Germano Vogt, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Rodrigo N. Calheiros, Christian Esteve Rothenberg |
NetSoft | 6 |
| 2024 | Spinner: Enabling In-network Flow Clustering Entirely in a Programmable Data PlaneabstractData plane programmability is redesigning the way we manage and operate forwarding devices. However, most of the algorithmic decisions performed by data planes are still deterministic and control-plane dependent. We argue that it is possible to break this dependency and make the data plane intelligent, so that it can learn the infrastructure state autonomously. Despite existing efforts to make data planes intelligent, little has been done to design unsupervised ML algorithms that fit the architectural constraints of programmable devices. Executing such approaches in the data plane has the potential to reduce the overall decision-making time, thus meeting packet processing deadlines (which are in the order of nanoseconds). In this paper, we propose Spinner, the first effort to deliver an unsupervised Machine Learning (ML) approach entirely in programmable devices. Spinner is a flow clustering algorithm designed to fit existing architectural constraints of SmartNICs, and that can reach line rate for most packet sizes with complexity O(k). To demonstrate the potential behind in-network clustering, we prototype and deploy Spinner in a programmable testbed and use it to enhance Explicit Congestion Notifications (ECN) at the server side. Spinner-enhanced TCP provides up to 2x higher throughput when comparing to de-facto TCP implementations. Luigi Cannarozzo, Thiago Bortoluzzi Morais, Paulo Silas Severo de Souza, Leonardo Gobatto, Ivan Peter Lamb, Pedro Arthur Pinheiro Rosa Duarte, José Rodrigo Azambuja, Arthur Francisco Lorenzon, Fábio D. Rossi, Weverton Luis da Costa Cordeiro, Marcelo Caggiani Luizelli |
NOMS | 9 |
| 2024 | A neural network framework for optimizing parallel computing in cloud servers
Everton Camargo de Lima, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Arthur Francisco Lorenzon |
J. Syst. Archit. | 2 |
| 2024 | Synergistically Rebalancing the EDP of Container-Based Parallel ApplicationsabstractThe use of containers has become standard in cloud environments. However, many parallel applications in containers will not present gains proportional to the extra available hardware. This inefficient use of hardware naturally leads to energy consumption waste. With that in mind, we proposeTT-Autoscaling. It works at two different levels: a) in the container, by automatically and transparently tuning the number of threads at runtime of the application, in a way to optimize the trade-off between energy and performance; b) in the cloud infrastructure, by smartly transferring the released resources to other containers that may run in parallel, making better use of the available resources. We compareTT-Autoscalingto the default execution of containers (serial execution with the maximum number of threads), showing 55.8% of performance improvements, 53.6% of energy reductions, and 79.5% of EDP improvements. We also show thatTT-Autoscalingoutperforms strategies that apply vertical autoscalers proposed by orchestrator tools. Vinicius S. da Silva, Everton Camargo de Lima, Janaina Schwarzrock, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Towards Optimizing the Edge-to-Cloud Continuum Resource Allocation
Igor Ferrazza Capeletti, Ariel Góes de Castro, Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
CLOSER | 6 |
| 2023 | Latency-Aware Cost-Efficient Provisioning of Composite Applications in Multi-Provider Clouds
Daniel Chaves Temp, Igor Ferrazza Capeletti, Ariel Góes de Castro, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
CLOSER | 7 |
| 2023 | Taking Detours: An In-Network Fault-Tolerant Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel network monitoring approach mainly fostered by programmable network devices. Despite existing efforts toward the orchestration of INT, little has yet been done to provide fault-tolerant mechanisms in the data plane (e.g., to address hardware failure). In this paper, we introduce InPatching - an in-network approach to fast recover INT-based monitoring from network link failures. InPatching is implemented in the data plane and allows the application of detours in an autonomous and coordinated manner without the control plane intervention. To provide efficient detours to INT solutions, we formalize the fault-tolerant probing planning for INT by means of a MILP (Mixed-Integer Linear Programming) model. We prototype InPatching in P4 and we show that it can recover from fault conditions much faster than control plane solutions (up to 18X), while not imposing substantial overhead. Ariel Góes de Castro, Igor Capelletti, Fábio D. Rossi, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Christian Esteve Rothenberg, Marcelo Caggiani Luizelli |
ICC | 3 |
| 2023 | NeurOPar, A Neural Network-Driven EDP Optimization Strategy for Parallel WorkloadsabstractThe pursuit of energy efficiency has been driving the development of techniques to optimize hardware resource usage in high-performance computing (HPC) servers. On multicore architectures, thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are three popular methods applied to improve the trade-off between performance and energy consumption, represented by the energy-delay product (EDP). However, the complexity of selecting the optimal configuration (TLP degree, DVFS, and UFS) for each application poses a challenge to software developers and end-users due to the massive number of possible configurations. To tackle this challenge, we propose NeurOpar, an optimization strategy for parallel workloads driven by an artificial neural network (ANN). It uses representative hardware and software metrics to build and train an ANN model that predicts combinations of thread count and core/uncore frequency levels that provide optimal EDP results. Through experiments on four multicore processors using twenty-five applications, we demonstrate that NeurOPar predicts combinations that yield EDP values close to the best ones achieved by an exhaustive search and improve the overall EDP by 42% compared to the default execution of HPC applications. We also show that NeurOPar can enhance the execution of parallel applications without incurring the performance and energy penalties associated with online methods by comparing it with two state-of-the-art strategies. Cristiano A. Künas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
SBAC-PAD | 2 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 7 |
| 2023 | Mobility-Aware Registry Migration for Containerized Applications on Edge Computing Infrastructures
Daniel Chaves Temp, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
J. Netw. Comput. Appl. | 5 |
| 2022 | Children's Impressions of Early Spelling Assessment through Handwriting on Tablets vs. Paper-based MethodabstractTraditional paper-based children’s spelling assessments were hampered due to Covid-19 because existing technologies did not provide strategic signals to teachers, such as the child’s handwriting direction and how they read what they write. Our project emerged as a novel method to assess children’s spelling by touchscreens in this context. Hence, this paper aims to extend community knowledge concerning children’s experience and perception of handwriting spelling on tablet devices. The experiment consisted in presenting three handwriting methods (paper and pencil, finger and pen writing) and was conducted with eight Brazilian children between 4.5 and 7 years old. In addition to observation, in our experimental protocol we adopted the Fun Sorter, Again-Again Table, and the Smileyometer as evaluation tools. Our results show children were excited about handwriting using a touch pen on the tablet. Most of them even revealed they prefer the pen tablet mode to the traditional paper and pencil mode. However, the majority of children did not feel comfortable writing by finger, and it required more time than other methods. Furthermore, we observed child’s handwriting using finger looks different when compared to paper and pencil, while the tracing using a touch pen is similar to the registration produced on paper. Jaline Mombach, Fábio D. Rossi, Deborah S. A. Fernandes, Fabrízzio Alphonsus A. M. N. Soares |
IDC | 2 |
| 2022 | Towards Efficient Selective In-Band Network Telemetry Report Using SmartNICs
Ronaldo Canofre, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 4 |
| 2022 | Latency-aware Privacy-preserving Service Migration in Federated Edges
Paulo Silas Severo de Souza, Ângelo Vieira, Felipe Rubin, Tiago Ferreto, Fábio D. Rossi |
CLOSER | 5 |
| 2022 | Multivariate Interpolation at the Edge to Infer Faulty IoT Sensor Metrics
Marcos Paulo Konzen, Patric Lincoln Ramires Izolan, Fábio Júnior Griesang, Paulo Silas Severo de Souza, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Júlio C. B. de Mattos, Cinara Ewerling da Rosa, Fábio D. Rossi |
CLOSER | 10 |
| 2022 | DyPro: Dynamic Probing Planning for In-Band Network TelemetryabstractIn-band Network Telemetry (INT) is a novel net-work monitoring mechanism that improves fine-grained net-work visibility. Despite the increasing research efforts towards the orchestration of INT data acquisition, little has yet been done to efficiently collect telemetry data from the network considering monitoring applications requirements. In this paper, we introduce DyPro - a dynamic probing planning for INT. In particular, DyP ro ensures that telemetry dependencies are always satisfied by monitoring application requirements. We theoretically formalize it as a Mixed-Integer Linear Programming (MILP) optimization model and propose a heuristic procedure to efficiently solve it. Results show that DyP ro can outperform state-of-the-art solutions by up to 5x regarding the percentage of monitoring applications satisfied. Leandro M. Dallanora, Ariel Góes de Castro, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 4 |
| 2022 | Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 4 |
| 2021 | The Actual Cost of Programmable SmartNICs: Diving into the Existing Limits
Pablo B. Viegas, Ariel Góes de Castro, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA (1) | 4 |
| 2021 | Synergically Rebalancing Parallel Execution via DCT and Turbo BoostingabstractThe increasing use of cloud and HPC systems put more pressure on the efficient utilization of hardware resources to keep costs low. Many dynamic concurrency throttling (DCT) techniques have successfully used to tune the number of executing threads to better balance a parallel application according to its available scalability. Similarly, boosting frequency strategies have been used to speed up the sequential parts’ execution. Given that, we propose Poseidon, the first transparent and automatic approach that cooperatively exploits both techniques to rebalance OpenMP applications without any preprocessing, with no code transformation, recompilation, or OS modification. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
DAC | 3 |
| 2021 | Combining Thread Throttling and Mapping to Optimize the EDP of Parallel ApplicationsabstractThread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1. Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 4 |
| 2021 | Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo BoostingabstractWith the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 4 |
| 2021 | Mitigating the processor aging through dynamic concurrency throttling
Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Parallel Distributed Comput. | 3 |
| 2021 | Low learning-cost offline strategies for EDP optimization of parallel applications
Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Samuel Xavier de Souza, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 2 |
| 2020 | A Heuristic Approach for Large-Scale Orchestration of the In-band Data Plane Telemetry Problem
Rumenigue Hohemberger, Arthur Francisco Lorenzon, Fábio D. Rossi, Marcelo Caggiani Luizelli |
AINA | 3 |
| 2020 | IRENE: Interference and High Availability Aware Microservice-based Applications Placement for Edge Computing
Paulo Silas Severo de Souza, João Nascimento, Conrado Boeira, Ângelo Vieira, Felipe Rubin, Rômulo Reis de Oliveira, Fábio D. Rossi, Tiago Ferreto |
CLOSER | 7 |
| 2020 | Remote Assessing Children's Handwriting Spelling on Mobile DevicesabstractAssessment of children's spelling development stages is an activity frequent in literacy classrooms. Usually, teachers adopt dictation sessions, using a paper-and-pencil based method, which has to be conducted individually with each child. As there are many students in a class, the activity becomes laborious to be offered frequently, pushing teachers to opt to a sub-optimal amount of tests. Moreover, the number of tests to be performed in person is now limited. In this context, we aim to develop a method for conducting automated word dictation sessions for an in-person and remote assessment. Besides, our proposal aims to support teachers, parents, and other literacy professionals in identifying children's spelling development stages. In this paper, we report the conception of our computational artifact through the Design Science Research Methodology. After reviewing the existing studies, we developed a high fidelity prototype and evaluated the concept during a focus group discussion with literacy teachers. The collected qualitative result indicates the feasibility and utility of our approach, and evident limitations in existing apps to assess children's spelling, contributing to future research in this area. Jaline Mombach, Fábio D. Rossi, Juliana Paula Felix, Fabrízzio Alphonsus A. M. N. Soares |
COMPSAC | 2 |
| 2020 | Decreasing the Learning Cost of Offline Parallel Application Optimization StrategiesabstractMany parallel applications do not scale as the number of threads increases, which means that executing them with the maximum possible number of threads will not always deliver the best outcome in performance, energy consumption, or the tradeoff between both (represented by the energy-delay product- EDP). Given that, several strategies, online and offline, have already been proposed to rightly tune the number of threads according to the application. While the former can capture some behaviors that can only be known at runtime, the latter do not impose any execution overhead and can use more efficient and costly algorithms. However, these learning algorithms in static strategics may take several hours, precluding their use or a smooth migration across different systems. In this scenario, we propose a generic methodology for such offline strategies to significantly decrease the learning time by inferring the execution behavior of parallel applications using smaller input sets than the ones used by the target applications. Through the execution of eighteen well-known benchmarks on two multicore processors, we show that our methodology is capable of converging to results that are very close to those that use the regular input set, but converging 84.7% faster, on average. We also show that such a strategy delivers better results than a dynamic one, presenting an EDP 7.7% lower, on average, when executing the applications with the number of threads found during learning. Finally, we also compare our learning methodology with an exhaustive search. It has an average learning cost (i.e., the time spent by our search algorithm to find the best configuration) of only 3.1% to optimize the EDP of the entire benchmark set1. Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 2 |
| 2020 | Modeling and Simulating Daily Power Budgets for Sustainable Data CentersabstractA novel energy-efficient scenario that makes possible to maintain sustainable data centers to feed part of resources through renewable energy sources has emerged. As renewable energies are accumulated in the form of power budgets, data centers must adapt a slice of the processing resources required to meet applications at those limits. This work is modeling and simulating the computing capacity of a data center according to daily power budgets from different sources of renewable energy. The results showed that based on the daily energy harvest of today's renewable energy sources, intelligent resource orchestration could use such energy so that up to 40% of what is needed to maintain a quality-of-service data center comes from non-polluting sources. Rumenigue Hohemberger, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli, Fábio D. Rossi |
PDP | 4 |
| 2020 | A new approach to performing paper-based children's spelling tests on mobile devicesabstractIdentifying the phase or stage of children's spelling development is a regular activity in literacy classes. Usually, teachers assume some developmental theory and perform paper-and-pencil based tests. Existing digital tools often do not consider any of these known theories, nor do they capture the child's handwriting. Therefore, some teachers prefer to continue applying the tests manually. Accordingly, this study's research question is how to promote the performance of spelling tests on mobile devices, simulating the interaction between manual tests and the theoretical models already used by teachers. Through the Design Science Research Methodology (DSRM), we propose a method to apply child spelling tests in an automated way. In this work, we present the results of the child's interface usability evaluation in the developed computer artifact, using guidelines of Touchscreen Interaction Design Recommendations for Children (TIDRC). The results indicate adequacy to the recommendations of 88% of items in visual and audio features (cognitive dimension), 75% in the physical dimension, and 47% in socio-emotional dimension. These results are promising and relevant compared to previous studies that evaluated apps using the TIDRC framework. Jaline Mombach, Afonso Ueslei Da Fonseca, Thamer H. Nascimento, Wellington Galvão Rodrigues, Henrique Gressler, Fábio D. Rossi, Fabrízzio Alphonsus A. M. N. Soares |
SMC | 6 |
| 2019 | Improving the Trade-Off between Performance and Energy Saving in Mobile Devices through a Transparent Code Offloading TechniqueabstractThe popularity of mobile devices has increased significantly, and nowadays they are used for the most diverse purposes like accessing the Internet or helping on business matters. Such popularity emerged as a consequence of the compatibility of these devices with a large variety of applications. However, the complexity of these applications boosted the demand for computational resources on mobile devices. Code Offloading is a solution that aims to mitigate this problem by reducing the use of resources and battery on mobile devices by sending parts of applications to be processed in the cloud. In this sense, this paper presents an evaluation of a transparent code offloading technique, where no modification in the application source code is required to allow the smartphone to send parts of the application to be processed in the cloud. We used a face detection application for the evaluation. Results showed the technique can improve applications performance in some scenarios, achieving speed-up of 12x in the best case. Rômulo Reis de Oliveira, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Tiago Ferreto, Fábio D. Rossi |
CLOSER | 5 |
| 2019 | Towards Balancing Energy Savings and Performance for Volunteer Computing through Virtualized Approach
Fábio D. Rossi, Tiago Ferreto, Marcelo Da Silva Conterato, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Rodrigo N. Calheiros, Guilherme da Cunha Rodrigues |
CLOSER | 1 |
| 2019 | On the Integrated Professional Practice in a Computing Course Towards InnovationabstractThe integration of disciplines in higher technology courses is a significant challenge. Sometimes the student can not visualize the correlations between different knowledge in the direction of a complete formation. This paper shows an integrated professional practice that relies on an innovation discipline to discover and prospect new technological products. From this, technical disciplines can focus on developing such products, having as supporting the knowledge and capacities determined in their curricula. The results show that students were able to develop products that aligned with local and regional problems, as well as open up a possibility for the creation of startups based on the developed products. Fábio D. Rossi, Paulo Silas Severo de Souza, Jaline Mombach, Tiago Ferreto |
ICALT | 1 |
| 2019 | GAMED: Gamification-Based Assessment Methodology for Final Project DevelopmentabstractTeachers and psychologists report that students may suffer from multiple psychological issues such as lack of interest and high levels of stress during the development of final course projects. In this context, approaches such as gamification arise with the proposal of improving students motivation by bringing games elements to school. However, employing gamification into classroom is not a trivial task since, if not managed properly, students may lose their focus. In this paper we present GAMED, an assessment methodology that introduces sistematic steps to improve students engagement through gamification. We used GAMED in a class with high school students during two semesters and the results showed that it can improve aspects such as motivation, engagement, and teamwork. Paulo Silas Severo de Souza, Jaline Mombach, Fábio D. Rossi, Tiago Ferreto |
ICALT | 3 |
| 2019 | The Impact of Parallel Programming Interfaces on the Aging of a Multicore Embedded ProcessorabstractIn order to meet the increasing performance demand of applications, the amount of cores in a single chip package has been increasing. However, the heat has been rising at a higher scale, which accelerates the aging process in modern processors. Therefore, wisely balancing the use of resources is important to extend its longevity. Frequency performance stagnates after a certain amount of concurrent threads starts executing. In such cases, the only result is a temperature rise that directly influences the aging process, reducing the processor lifetime. This unbalance between threads can be originated from many factors, which includes the way threads communicate and synchronize. Considering that those characteristics are related to the Parallel Programming Interface (PPI) used to parallelize the application, this work proposes to evaluate three widely used PPIs executing on an embedded multicore. We show that, depending on the characteristic of the application, by only switching from one PPI to another, it is possible to reduce the effects of aging. For that, we have developed a model based on the Arrhenius equation. We show that OpenMP has a lower impact on the processor aging for memory-bound applications: up to 38% and 68% lower than PThreads and MPI, respectively. On the other hand, PThreads presents the lowest impact on the processor aging for CPU-bound applications. Ângelo Vieira, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Marcelo Caggiani Luizelli, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Fábio D. Rossi, Jorji Nonaka |
ISCAS | 9 |
| 2019 | Multilevel resource allocation for performance-aware energy-efficient cloud data centersabstractThe massive power consumption of data centers has been a recurring concern in current research. In cloud environments, lots of methods are being adopted that aim for energy efficiency. However, although such methods enable the decrease in power consumption, they regularly affect application performance. In this paper, we present a multilevel resource allocation approach towards dynamic network bandwidth at the physical substrate, managing different power-saving states and workload allocation at the cloud infrastructure at the same time employ virtual machine allocation and selection policies at the cloud platform. In order to evaluate our approach, tests were carried out in a simulated environment using scale-out application on a dynamic cloud infrastructure. Results showed that our proposal presents a better balance regarding a more energy-efficient data center with a smaller impact on application performance when compared with other works discussed in the literature. Fábio D. Rossi, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Arthur Francisco Lorenzon, Marcelo Caggiani Luizelli |
ISCC | 1 |
| 2019 | Transparent Aging-Aware Thread ThrottlingabstractTo satisfy the rising performance demands of modern applications, the number of cores in a single chip package has been increasing. However, the power dissipated and temperature have been growing at a higher rate, accelerating the aging process of new processors. Considering that a significant number of parallel applications are unbalanced, in many cases performance stagnates after a certain number of concurrent threads starts executing. In such cases, the only outcome is a temperature rise on the processor, which drastically accelerates aging. Given that, we propose an automatic and transparent approach to reduce the processor aging by automatically tuning the number of threads for OpenMP applications at run-time. Our tool, Geras, is entirely transparent to the end-user, so even already compiled binaries can be optimized. Through the execution of twelve well-known benchmarks on two multicore platforms, we show that Geras can improve the processor lifetime by up to 83% and 89% over the standard OpenMP execution and its built-in feature that dynamically adjusts the number of threads, respectively. We also show that Geras outperforms techniques that target performance or energy, which reinforces the need for a specific tool that optimizes aging1. Thiarles S. Medeiros, Luan Pereira, Fábio D. Rossi, Marcelo Caggiani Luizelli, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
SBAC-PAD | 3 |
| 2019 | The Impact of Turbo Frequency on the Energy, Performance, and Aging of Parallel ApplicationsabstractTechnologies that improve the performance of parallel applications by increasing the nominal operating frequency of processors respecting a given TDP (Thermal Design Power) have been widely used. However, they may impact on other non-functional requirements in different ways (e.g. increasing energy consumption or aging). Therefore, considering the huge number of configurations available, represented by the range of all possible combinations among different parallel applications, amount of threads, dynamic voltage and frequency scaling (DVFS) governors, boosting technologies and simultaneous multithreading (SMT), selecting the one that offers the best tradeoff for a non-functional requirement is extremely challenging for software designers. Given that, in this work we assess the impact of changing these configurations on the energy consumption, performance, and aging of parallel applications on a turbo-compliant processor. Results show that there is no single configuration that would provide the best solution for all nonfunctional requirements at once. For instance, we demonstrate that the configuration that offers the best performance is the same one that has the worst impact on aging, accelerating it by up to 1.75 times. With our experiments, we provide guidelines for the developer when it comes to tuning performance using turbo boosting to save as much energy as possible and increase the lifespan of the hardware components. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Fábio D. Rossi, Marcelo Caggiani Luizelli, Alessandro Girardi, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
VLSI-SoC | 3 |
| 2018 | Evaluating container-based virtualization overhead on the general-purpose IoT platformabstractVirtualization has become a key technology that provides several advantages (e.g., flexibility, migration, isolation) for a plethora of computing infrastructures. However, traditional virtualization models are not suitable for embedded IoT platforms due to the virtualization layer verhead. New virtualization proposals such as container-based approaches arise as an option where performance is not impacted. However, when working on general-purpose embedded platforms, some studies have demonstrated that applications on container-based virtualization on embedded devices present considerable performance overhead. Since most performance evaluations on platforms using containers were run on servers, this study expands the testbed scenario by analyzing several metrics that measure the overhead of container-based virtualization layer on embedded IoT devices. Results demonstrated improvements up to 23% in terms of performance and up to 32% in terms of EDP. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Fábio D. Rossi, Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Marcelo Da Silva Conterato, Tiago Ferreto |
ISCC | 3 |
| 2018 | Performance-Aware Energy-Efficient Processes Grouping for Embedded PlatformsabstractEmbedded systems are becoming more popular in several sectors of society by performing a broad range of tasks. In this context, there is a concern about improving the trade-off between performance and energy savings since they are usually battery-dependent. However, it is not a trivial task since some embedded devices are developed with strict hardware constraints. In this sense, we present an operating system level tool that groups processes dynamically on resources. Our tool manages different types of processes, and through isolation characteristics, provides a better utilization of resources. The results show that our tool can improve the trade-off between performance and power saving of embedded systems in up to 15%. Paulo Silas Severo de Souza, Wagner dos Santos Marques, Marcelo Da Silva Conterato, Tiago Ferreto, Fábio D. Rossi |
ISCC | 5 |
| 2017 | Dynamic Network Bandwidth Resizing for Big Data ApplicationsabstractBig Data concerns processing of large volumes of digital data with high velocity and variety. Big Data technologies allow the analysis of data in real time, which is critical for various eScience applications. In order to meet the growing demand of Big Data applications, the infrastructures must be flexible enough to adapt to the characteristics of the applications. Most of the solutions presented in the literature to support Big Data applications focus on scaling processors and memory to handle a variable demand from applications. In a complementary way, this article targets the problem of adapting the network bandwidth to the amount of data to be transferred to and from the applications in order to improve the performance of the applications. For this purpose, we propose the use link aggregation protocol along with Software-Defined Network capabilities for management of the network flow. Results showed that the proposed approach improves the application's performance by up to 33%. Fábio D. Rossi, Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Marcelo Da Silva Conterato |
eScience | 1 |
| 2017 | Improving EDP in multi-core embedded systems through multidimensional frequency scalingabstractEnergy saving management in multi-core embedded environments has been a challenge for designers. To achieve energy efficiency, most studies consider dynamic frequency scaling on one hardware component only, such as processor or memory - which will most likely also affect performance. This work proposes the use of frequency scaling considering the three most important hardware components altogether: processors, L2 cache, and RAM; seeking for the best set of frequencies for each one of them to improve the Energy-Delay Product (EDP), depending on the application's behavior. Therefore, this work addresses multidimensional frequency scaling for multi-core embedded systems. By evaluating different frequency levels, we show that the EDP can be improved in up to 46.4% when compared to the standard way that the frequencies are configured. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Mateus B. Rutzig, Fábio D. Rossi |
ISCAS | 6 |
| 2017 | Modeling and simulation of global and sleep states in ACPI-compliant energy-efficient cloud environmentsabstractSummary The more large‐scale data centers infrastructure costs increase, the more simulation‐based evaluations are needed to understand better the trade‐off between energy and performance and support the development of new energy‐aware resource allocation policies. Specifically, in the cloud computing field, various simulators are able to predict and measure the behavior of applications on different architectures using different resource allocation policies. Yet, only a few of them have the ability to simulate energy‐saving strategies, and none of them support the complete advanced configuration and power interface (ACPI) specification. ACPI defines a terminology for all possible power states of a machine and their associated power rate. The hardware industry has relied on ACPI to provide up‐to‐date standard interfaces for hardware discovery, configuration, power management, and monitoring, enabling a better understanding of the energy consumption level of different hardware states, referred to as ACPI G‐states, S‐states, and P‐states. In this paper, we improve the modeling and simulation of the ACPI G/S‐states and show not only that these states offer different energy‐saving levels but also that state transitions consume energy. In addition, we model the latency to transit between two states and the effects on the turnaround time when the transitions are not performed conservatively. Furthermore, the equations provide essential information to quantify the trade‐off between energy consumption and performance and assist in the analysis/decision on which strategy fits better in the environment and how it could be refined. Our expanded energy model was implemented in CloudSim and validated with simulation‐based experiments with a very high level of accuracy, with a standard deviation of at most 6%. Copyright © 2016 John Wiley & Sons, Ltd. Miguel G. Xavier, Fábio D. Rossi, César A. F. De Rose, Rodrigo N. Calheiros, Danielo Goncalves Gomes |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | E-eco: Performance-aware energy-efficient cloud data center orchestration
Fábio D. Rossi, Miguel G. Xavier, César A. F. De Rose, Rodrigo N. Calheiros, Rajkumar Buyya |
J. Netw. Comput. Appl. | 1 |
| 2015 | Modeling power consumption for DVFS policiesabstractPower-aware management strategies are a trend towards achieving energy-efficient computing environments. One of the approaches behind those strategies is dynamic frequency and voltage scaling (DVFS). Since frequency adjustments may have a negative impact on system performance, users often have to experiment with these policies to find the optimal configuration for their application and energy reduction goals. While the performance impact can be easily measured by the total execution time of an application, power consumption measurements require additional logging and frequently external equipment. The following paper presents a mathematical model to help users estimate the power consumption of their application when using different DVFS policies. A preliminary evaluation shows that the model has 94% accuracy when compared against real-time measurements. Fábio D. Rossi, Mauro Storch, Israel C. De Oliveira, César A. F. De Rose |
ISCAS | 1 |
| 2015 | On the Impact of Energy-Efficient Strategies in HPC ClustersabstractEnergy-aware management strategies are a recent trend towards achieving energy-efficient computing in HPC clusters. One of the approaches behind those strategies is to apply energy-saving states on idle nodes, alternating them among different sleep states that reflect on many power consumption levels. This paper investigated the way such energy-efficient strategies affected the job turnaround time - the elapsed time between when the job is submitted and when the job is completed, including the wait time as well as the job's actual execution time - in these clusters. Based on the results we proposed a Best-Fit Energy-Aware Strategy that switches the nodes to a sleep state, depending on the throughput of the resource manager's job queue. We simulated the proposed strategy using the SimGrid simulator. Our preliminary results showed a reduction of up to 19% in the overall energy consumption and give us a better understanding of the trade-offs involved in using energy-efficient strategies. Fábio D. Rossi, Miguel G. Xavier, Yuri J. Monti, César A. F. De Rose |
PDP | 1 |
| 2015 | A Performance Isolation Analysis of Disk-Intensive Workloads on Container-Based CloudsabstractThe popularity of Cloud computing due to the increasing number of customers has led Cloud providers to adopt resource-sharing solutions to meet growing demand for infrastructure resources. As the adoption of resource-sharing/consolidation in Cloud computing became arguably a well-established solution, the ability the underlying virtualization systems of preventing performance interferences from customers must also be understood. Virtualization systems based on containers, such as LXC, are the basis of the next-generation of Cloud computing and have become the most popular solution under PaaS/IaaS Cloud platforms with the rise of Docker -- an open platform for developers and sysadmins to build, ship, and run distributed applications. Such platforms have enticed many attentions globally, since they leverage container-based virtualization systems to offer high scalability while low performance overheads, the performance might be solely aggravated if the customers' workloads are consolidated onto the same hardware and the isolation layer does not properly isolate the shared resources. Performance isolation is an inherent concern of such systems due to the nature as they are conceived and is still an unexplored and open research topic, the consequences might influence in the adoption under shared Cloud computing platforms where Quality-of-Service is a crucial factor that cannot be disregarded. In this paper we analyze the performance interference suffered by disk-intensive workloads within very noisy-perturbed containers (different hardware components stressed). Our results show workload combinations whose performance degradation goes up to 38%, but in contrast we expose a workload-balanced scenario wherein the performance does not suffer any interference. Miguel G. Xavier, Israel C. De Oliveira, Fábio D. Rossi, Robson D. Dos Passos, Kassiano J. Matteussi, César A. F. De Rose |
PDP | 3 |
| 2014 | Green software development for multi-core architecturesabstractAdvances in computer architecture to provide higher parallelism (e.g. hyper threading and multi-core) usually incur in higher complexity in software development. Applications should be designed to use efficiently the additional resources in order to improve its performance. However, the popularity of mobile devices and recent studies in IT-related energy consumption have driven software developers to focus also on energy efficiency. Besides improving applications' performance, software developers should aim at minimizing the amount of energy consumed by the applications. Energy saving becomes an important non-functional requirement for new applications. This paper evaluates the behavior of applications on multi-core architectures and proposes energy-saving alternatives for software development. Fábio D. Rossi, Miguel G. Xavier, Endrigo D'Agostini Conte, Tiago Ferreto, César A. F. De Rose |
ISCC | 1 |
| 2013 | Performance Evaluation of Container-Based Virtualization for High Performance Computing EnvironmentsabstractThe use of virtualization technologies in high performance computing (HPC) environments has traditionally been avoided due to their inherent performance overhead. However, with the rise of container-based virtualization implementations, such as Linux VServer, OpenVZ and Linux Containers (LXC), it is possible to obtain a very low overhead leading to near-native performance. In this work, we conducted a number of experiments in order to perform an in-depth performance evaluation of container-based virtualization for HPC. We also evaluated the trade-off between performance and isolation in container-based virtualization systems and compared them with Xen, which is a representative of the traditional hypervisor-based virtualization systems used today. Miguel G. Xavier, Marcelo Veiga Neves, Fábio D. Rossi, Tiago Ferreto, Timoteo Lange, César A. F. De Rose |
PDP | 3 |