Luiz Gustavo Fernandes

dblp:f/LuizGustavoFernandes · also Luiz G. L. Fernandes, Luiz Gustavo Leão Fernandes · DBLP profile ↗
← Back
48ranked-venue papers
0as first author
19since 2021 · last 2025
0000-0002-7506-3685ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 9 since 2021Software engineering, systems software and programming languages · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Computer networks · 3Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 Performance, Portability, and Productivity of HIP on GPUs with NAS Parallel Benchmarks
abstract
Graphics Processing Units (GPUs) are powerful, massively parallel processors that have become ubiquitous in modern computing. In recent years, the GPU market has diversified, with vendors like AMD and Intel offering high-performance alternatives to NVIDIA. However, most applications are written using NVIDIA’s CUDA API, which is incompatible with non-NVIDIA GPUs, creating significant challenges for developers who must port their code to different architectures. To address this issue, AMD developed the Heterogeneous-Compute Interface for Portability (HIP), an open-source API for cross-vendor GPU programming. However, HIP is relatively new, leaving gaps in the literature regarding its performance, portability, and productivity. In this paper, we evaluate HIP using the NAS Parallel Benchmarks (NPB), a CFD-based suite maintained by NASA. We present the first HIP-based implementation of NPB and conduct experiments on integrated and discrete GPUs from NVIDIA, AMD, and Intel. Our results provide novel insights into HIP’s performance and portability, particularly for integrated GPUs and Intel discrete GPUs, which have been underrepresented in prior studies. We also assess productivity using different metrics to quantify the programming effort of HIP-based implementations. This work addresses key gaps in the literature, offering valuable data and insights for developers targeting emerging GPU architectures.
Gabriell Alves de Araujo, Dalvan Griebler, Luiz Gustavo Fernandes
SBAC-PAD3
2024 MPR: An MPI Framework for Distributed Self-adaptive Stream Processing
Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes, Walter Binder
Euro-Par (3)3
2024 Benchmarking parallel programming for single-board computers
Renato B. Hoffmann, Dalvan Griebler, Rodrigo da Rosa Righi, Luiz Gustavo Fernandes
Future Gener. Comput. Syst.4
2024 Performance and programmability of GrPPI for parallel stream processing on multi-cores
abstract
Abstract GrPPI library aims to simplify the burdening task of parallel programming. It provides a unified, abstract, and generic layer while promising minimal overhead on performance. Although it supports stream parallelism, GrPPI lacks an evaluation regarding representative performance metrics for this domain, such as throughput and latency. This work evaluates GrPPI focused on parallel stream processing. We compare the throughput and latency performance, memory usage, and programmability of GrPPI against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks and benchmarks with handwritten parallel code using the same backends supported by GrPPI. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is often competitive with handwritten parallel code, the infeasibility of fine-tuning GrPPI is a crucial drawback for emerging applications. Despite this, programmability experiments estimate that GrPPI can potentially reduce the development time of parallel applications by about three times.
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes
J. Supercomput.6
2024 Enhancing self-adaptation for efficient decision-making at run-time in streaming applications on multicores
abstract
Abstract Parallel computing is very important to accelerate the performance of computing applications. Moreover, parallel applications are expected to continue executing in more dynamic environments and react to changing conditions. In this context, applying self-adaptation is a potential solution to achieve a higher level of autonomic abstractions and runtime responsiveness. In our research, we aim to explore and assess the possible abstractions attainable through the transparent management of parallel executions by self-adaptation. Our primary objectives are to expand the adaptation space to better reflect real-world applications and assess the potential for self-adaptation to enhance efficiency. We provide the following scientific contributions: (I) A conceptual framework to improve the designing of self-adaptation; (II) A new decision-making strategy for applications with multiple parallel stages; (III) A comprehensive evaluation of the proposed decision-making strategy compared to the state-of-the-art. The results demonstrate that the proposed conceptual framework can help design and implement self-adaptive strategies that are more modular and reusable. The proposed decision-making strategy provides significant gains in accuracy compared to the state-of-the-art, increasing the parallel applications’ performance and efficiency.
Adriano Vogel, Marco Danelutto, Massimo Torquati, Dalvan Griebler, Luiz Gustavo Fernandes
J. Supercomput.5
2023 A Latency, Throughput, and Programmability Perspective of GrPPI for Streaming on Multi-cores
abstract
Several solutions aim to simplify the burdening task of parallel programming. The GrPPI library is one of them. It allows users to implement parallel code for multiple backends through a unified, abstract, and generic layer while promising minimal overhead on performance. An outspread evaluation of GrPPI regarding stream parallelism with representative metrics for this domain, such as throughput and latency, was not yet done. In this work, we evaluate GrPPI focused on stream processing. We evaluate performance, memory usage, and programming effort and compare them against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is competitive with handwritten code in some cases, in other cases, the infeasibility of fine-tuning GrPPI is a crucial drawback. Despite this, programmability experiments estimate that GrPPI has the potential to reduce by about three times the development time of parallel applications.
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, André Sacilotto Santos, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes
PDP7
2023 Revisiting self-adaptation for efficient decision-making at run-time in parallel executions
abstract
Self-adaptation is a potential alternative to provide a higher level of autonomic abstractions and run-time responsiveness in parallel executions. However, the recurrent problem is that self-adaptation is still limited in flexibility and efficiency. For instance, there is a lack of mechanisms to apply adaptation actions and efficient decision-making strategies to decide which configurations should be conveniently enforced at run-time. In this work, we are interested in providing and evaluating potential abstractions achievable with self-adaptation transparently managing parallel executions. Therefore, we provide a new mechanism to support self-adaptation in applications with multiple parallel stages executed in multi-cores. Moreover, we reproduce, reimplement, and evaluate an existing decision-making strategy in our scenario. The observations from the results show that the proposed mechanism for self-adaptation can provide new parallelism abstractions and autonomous responsiveness at run-time. On the other hand, there is a need for more accurate decision-making strategies to enable efficient executions of applications in resource-constrained scenarios like multi-cores.
Adriano Vogel, Marco Danelutto, Dalvan Griebler, Luiz Gustavo Fernandes
PDP4
2023 NAS Parallel Benchmarks with CUDA and beyond
abstract
Abstract NAS Parallel Benchmarks (NPB) is a standard benchmark suite used in the evaluation of parallel hardware and software. Several research efforts from academia have made these benchmarks available with different parallel programming models beyond the original versions with OpenMP and MPI. This work joins these research efforts by providing a new CUDA implementation for NPB. Our contribution covers different aspects beyond the implementation. First, we define design principles based on the best programming practices for GPUs and apply them to each benchmark using CUDA. Second, we provide ease of use parametrization support for configuring the number of threads per block in our version. Third, we conduct a broad study on the impact of the number of threads per block in the benchmarks. Fourth, we propose and evaluate five strategies for helping to find a better number of threads per block configuration. The results have revealed relevant performance improvement solely by changing the number of threads per block, showing performance improvements from 8% up to 717% among the benchmarks. Fifth, we conduct a comparative analysis with the literature, evaluating performance, memory consumption, code refactoring required, and parallelism implementations. The performance results have shown up to 267% improvements over the best benchmarks versions available. We also observe the best and worst design choices, concerning code size and the performance trade‐off. Lastly, we highlight the challenges of implementing parallel CFD applications for GPUs and how the computations impact the GPU's behavior.
Gabriell Alves de Araujo, Dalvan Griebler, Dinei A. Rockenbach, Marco Danelutto, Luiz Gustavo Fernandes
Softw. Pract. Exp.5
2023 Micro-batch and data frequency for stream processing on multi-cores
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes
J. Supercomput.4
2022 Analyzing Programming Effort Model Accuracy of High-Level Parallel Programs for Stream Processing
abstract
Over the years, several Parallel Programming Models (PPMs) have supported the abstraction of programming complexity for parallel computer systems. However, few studies aim to evaluate the productivity reached by such abstractions since this is a complex task that involves human beings. There are several studies to develop predictive methods to estimate the effort required to develop software applications. In order to evaluate the reliability of such metrics, it is necessary to assess the accuracy in different programming paradigms. In this work, we used the data of an experiment conducted with beginners in parallel programming to determine the effort required for implementing stream parallelism using FastFlow, SPar, and TBB. Our results show that some traditional software effort estimation models, such as COCOMO II, fall short. In contrast, Planning Poker could contribute toward a parallel-aware effort model.
Gabriella Andrade, Dalvan Griebler, Rodrigo Pereira dos Santos, Christoph W. Kessler, August Ernstsson, Luiz Gustavo Fernandes
SEAA6
2022 Evaluating Micro-batch and Data Frequency for Stream Processing Applications on Multi-cores
abstract
In stream processing, data arrives constantly and is often unpredictable. It can show large fluctuations in arrival frequency, size, complexity, and other factors. These fluctuations can strongly impact application latency and throughput, which are critical factors in this domain. Therefore, there is a significant amount of research on self-adaptive techniques involving elasticity or micro-batching as a way to mitigate this impact. However, there is a lack of benchmarks and tools for helping researchers to investigate micro-batching and data stream frequency implications. In this paper, we extend a benchmarking framework to support dynamic micro-batching and data stream frequency management. We used it to create custom benchmarks and compare latency and throughput aspects from two different parallel libraries. We validate our solution through an extensive analysis of the impact of micro-batching and data stream frequency on stream processing applications using Intel TBB and FastFlow, which are two libraries that leverage stream parallelism on multi-core architectures. Our results demonstrated up to 33% throughput gain over latency using micro-batches. Additionally, while TBB ensures lower latency, FastFlow ensures higher throughput in the parallel applications for different data stream frequency configurations.
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes
PDP4
2022 Self-adaptation on parallel stream processing: A systematic review
abstract
Summary A recurrent challenge in real‐world applications is autonomous management of the executions at run‐time. In this vein, stream processing is a class of applications that compute data flowing in the form of streams (e.g., video feeds, images, and data analytics), where parallel computing can help accelerate the executions. On the one hand, stream processing applications are becoming more complex, dynamic, and long‐running. On the other hand, it is unfeasible for humans to monitor and manually change the executions continuously. Hence, self‐adaptation can reduce costs and human efforts by providing a higher‐level abstraction with an autonomic/seamless management of executions. In this work, we aim at providing a literature review regarding self‐adaptation applied to the parallel stream processing domain. We present a comprehensive revision using a systematic literature review method. Moreover, we propose a taxonomy to categorize and classify the existing self‐adaptive approaches. Finally, applying the taxonomy made it possible to characterize the state‐of‐the‐art, identify trends, and discuss open research challenges and future opportunities.
Adriano Vogel, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes
Concurr. Comput. Pract. Exp.4
2022 OpenMP as runtime for providing high-level stream parallelism on multi-cores
Renato B. Hoffmann, Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes
J. Supercomput.4
2021 Assessing Coding Metrics for Parallel Programming of Stream Processing Programs on Multi-cores
abstract
From the popularization of multi-core architectures, several parallel APIs have emerged, helping to abstract the programming complexity and increasing productivity in application development. Unfortunately, only a few research efforts in this direction managed to show the usability pay-back of the programming abstraction created, because it is not easy and poses many challenges for conducting empirical software engineering. We believe that coding metrics commonly used in software engineering code measurements can give useful indicators on the programming effort of parallel applications and APIs. These metrics were designed for general purposes without considering the evaluation of applications from a specific domain. In this study, we aim to evaluate the feasibility of seven coding metrics to be used in the parallel programming domain. To do so, five stream processing applications implemented with different parallel APIs for multi-cores were considered. Our experiments have shown COCOMO II is a suitable model for evaluating the productivity of different parallel APIs targeting multi-cores on stream processing applications while other metrics are restricted to the code size.
Gabriella Andrade, Dalvan Griebler, Rodrigo Pereira dos Santos, Marco Danelutto, Luiz Gustavo Fernandes
SEAA5
2021 Introducing a Stream Processing Framework for Assessing Parallel Programming Interfaces
abstract
Stream Processing applications are spread across different sectors of industry and people's daily lives. The increasing data we produce, such as audio, video, image, and text are demanding quickly and efficiently computation. It can be done through Stream Parallelism, which is still a challenging task and most reserved for experts. We introduce a Stream Processing framework for assessing Parallel Programming Interfaces (PPIs). Our framework targets multi-core architectures and C++ stream processing applications, providing an API that abstracts the details of the stream operators of these applications. Therefore, users can easily identify all the basic operators and implement parallelism through different PPIs. In this paper, we present the proposed framework, implement three applications using its API, and show how it works, by using it to parallelize and evaluate the applications with the PPIs Intel TBB, FastFlow, and SPar. The performance results were consistent with the literature.
Adriano Marques Garcia, Dalvan Griebler, Luiz Gustavo Fernandes, Claudio Schepke
PDP3
2021 Towards On-the-fly Self-Adaptation of Stream Parallel Patterns
abstract
Stream processing applications compute streams of data and provide insightful results in a timely manner, where parallel computing is necessary for accelerating the application executions. Considering that these applications are becoming increasingly dynamic and long-running, a potential solution is to apply dynamic runtime changes. However, it is challenging for humans to continuously monitor and manually self-optimize the executions. In this paper, we propose self-adaptiveness of the parallel patterns used, enabling flexible on-the-fly adaptations. The proposed solution is evaluated with an existing programming framework and running experiments with a synthetic and a real-world application. The results show that the proposed solution is able to dynamically self-adapt to the most suitable parallel pattern configuration and achieve performance competitive with the best static cases. The feasibility of the proposed solution encourages future optimizations and other applicabilities.
Adriano Vogel, Gabriele Mencagli, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes
PDP5
2021 Latency-aware adaptive micro-batching techniques for streamed data compression on graphics processing units
abstract
Summary Stream processing is a parallel paradigm used in many application domains. With the advance of graphics processing units (GPUs), their usage in stream processing applications has increased as well. The efficient utilization of GPU accelerators in streaming scenarios requires to batch input elements in microbatches, whose computation is offloaded on the GPU leveraging data parallelism within the same batch of data. Since data elements are continuously received based on the input speed, the bigger the microbatch size the higher the latency to completely buffer it and to start the processing on the device. Unfortunately, stream processing applications often have strict latency requirements that need to find the best size of the microbatches and to adapt it dynamically based on the workload conditions as well as according to the characteristics of the underlying device and network. In this work, we aim at implementing latency‐aware adaptive microbatching techniques and algorithms for streaming compression applications targeting GPUs. The evaluation is conducted using the Lempel‐Ziv‐Storer‐Szymanski compression application considering different input workloads. As a general result of our work, we noticed that algorithms with elastic adaptation factors respond better for stable workloads, while algorithms with narrower targets respond better for highly unbalanced workloads.
Charles Michael Stein, Dinei A. Rockenbach, Dalvan Griebler, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes
Concurr. Comput. Pract. Exp.7
2021 The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes
Future Gener. Comput. Syst.7
2021 Providing high-level self-adaptive abstractions for stream parallelism on multicores
abstract
Abstract Stream processing applications are common computing workloads that demand parallelism to increase their performance. As in the past, parallel programming remains a difficult task for application programmers. The complexity increases when application programmers must set nonintuitive parallelism parameters, that is, the degree of parallelism. The main problem is that state‐of‐the‐art libraries use a static degree of parallelism and are not sufficiently abstracted for developing stream processing applications. In this article, we propose a self‐adaptive regulation of the degree of parallelism to provide higher‐level abstractions. Flexibility is provided to programmers with two new self‐adaptive strategies, one is for performance experts, and the other abstracts the need to set a performance goal. We evaluated our solution using compiler transformation rules to generate parallel code with the SPar domain‐specific language. The experimental results with real‐world applications highlighted higher abstraction levels without significant performance degradation in comparison to static executions. The strategy for performance experts achieved slightly higher performance than the one that works without user‐defined performance goals.
Adriano Vogel, Dalvan Griebler, Luiz Gustavo Fernandes
Softw. Pract. Exp.3
2020 The Impact of CPU Frequency Scaling on Power Consumption of Computing Infrastructures
Adriano Marques Garcia, Matheus S. Serpa, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux
ICCSA (6)5
2020 Efficient NAS Parallel Benchmark Kernels with CUDA
abstract
NAS Parallel Benchmarks (NPB) are one of the standard benchmark suites used to evaluate parallel hardware and software. There are many research efforts trying to provide different parallel versions apart from the original OpenMP and MPI. Concerning GPU accelerators, there are only the OpenCL and OpenACC available as consolidated versions. Our goal is to provide an efficient parallel implementation of the five NPB kernels with CUDA. Our contribution covers different aspects. First, best parallel programming practices were followed to implement NPB kernels using CUDA. Second, the support of larger workloads (class B and C) allow to stress and investigate the memory of robust GPUs. Third, we show that it is possible to make NPB efficient and suitable for GPUs although the benchmarks were designed for CPUs in the past. We succeed in achieving double performance with respect to the state-of-the-art in some cases as well as implementing efficient memory usage. Fourth, we discuss new experiments comparing performance and memory usage against OpenACC and OpenCL state-of-the-art versions using a relative new GPU architecture. The experimental results also revealed that our version is the best one for all the NPB kernels compared to OpenACC and OpenCL. The greatest differences were observed for the FT and EP kernels.
Gabriell Alves de Araujo, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes
PDP4
2020 Simplifying and implementing service level objectives for stream parallelism
Dalvan Griebler, Adriano Vogel, Daniele De Sensi, Marco Danelutto, Luiz Gustavo Fernandes
J. Supercomput.5
2019 Minimizing Communication Overheads in Container-based Clouds for HPC Applications
abstract
Although the industry has embraced the cloud computing model, there are still significant challenges to be addressed concerning the quality of cloud services. Network-intensive applications may not scale in the cloud due to the sharing of the network infrastructure. In the literature, performance evaluation studies are showing that the network tends to limit the scalability and performance of HPC applications. Therefore, we proposed the aggregation of Network Interface Cards (NICs) in a ready-to-use integration with the OpenNebula cloud manager using Linux containers. We perform a set of experiments using a network microbenchmark to get specific network performance metrics and NAS parallel benchmarks to analyze the performance impact on HPC applications. Our results highlight that the implementation of NIC aggregation improves network performance in terms of throughput and latency. Moreover, HPC applications have different patterns of behavior when using our approach, which depends on communication and the amount of data transferring. While network-intensive applications increased the performance up to 38%, other applications with aggregated NICs maintained the same performance or presented slightly worse performance.
Anderson M. Maliszewski, Adriano Vogel, Dalvan Griebler, Eduardo Roloff, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux
ISCC5
2019 Should PARSEC Benchmarks be More Parametric? A Case Study with Dedup
abstract
Parallel applications of the same domain can present similar patterns of behavior and characteristics. Characterizing common application behaviors can help for understanding performance aspects in the real-world scenario. One way to better understand and evaluate applications' characteristics is by using customizable/parametric benchmarks that enable users to represent important characteristics at run-time. We observed that parameterization techniques should be better exploited in the available benchmarks, especially on stream processing domain. For instance, although widely used, the stream processing benchmarks available in PARSEC do not support the simulation and evaluation of relevant and modern characteristics. Therefore, our goal is to identify the stream parallelism characteristics present in PARSEC. We also implemented a ready to use parameterization support and evaluated the application behaviors considering relevant performance metrics for stream parallelism (service time, throughput, latency). We choose Dedup to be our case study. The experimental results have shown performance improvements in our parameterization support for Dedup. Moreover, this support increased the customization space for benchmark users, which is simple to use. In the future, our solution can be potentially explored on different parallel architectures and parallel programming frameworks.
Carlos A. F. Maron, Adriano Vogel, Dalvan Griebler, Luiz Gustavo Fernandes
PDP4
2019 Memory Performance and Bottlenecks in Multicore and GPU Architectures
abstract
Nowadays, there are several different architectures available not only for the industry, but also for normal consumers. Traditional multicore processors, GPUs, accelerators such as the Sunway SW26010, or even energy efficiency-driven processors such as the ARM family, present very different architectural characteristics. This wide range of characteristics presents a challenge for the developers of applications. Developers must deal with different instruction sets, memory hierarchies, or even different programming paradigms when programming for these architectures. Therefore, the same application can perform well when executing on one architecture, but poorly on another architecture. To optimize an application, it is important to have a deep understanding of how it behaves on different architectures. The related work in this area mostly focuses on a limited analysis encompassing execution time and energy. In this paper, we perform a detailed investigation on the impact of the memory subsystem of different architectures, which is one of the most important aspects to be considered. For this study, we performed experiments in the Broadwell CPU and Pascal GPU, using applications from the Rodinia benchmark suite. In this way, we were able to understand why an application performs well on one architecture and poorly on others.
Matheus S. Serpa, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, Eduardo Henrique Molina da Cruz, Matthias Diener, Dalvan Griebler, Luiz Gustavo Fernandes
PDP7
2019 Stream Parallelism on the LZSS Data Compression Application for Multi-Cores with GPUs
abstract
GPUs have been used to accelerate different data parallel applications. The challenge consists in using GPUs to accelerate stream processing applications. Our goal is to investigate and evaluate whether stream parallel applications may benefit from parallel execution on both CPU and GPU cores. In this paper, we introduce new parallel algorithms for the Lempel-Ziv-Storer-Szymanski (LZSS) data compression application. We implemented the algorithms targeting both CPUs and GPUs. GPUs have been used with CUDA and OpenCL to exploit inner algorithm data parallelism. Outer stream parallelism has been exploited using CPU cores through SPar. The parallel implementation of LZSS achieved 135 fold speedup using a multi-core CPU and two GPUs. We also observed speedups in applications where we were not expecting to get it using the same combine data-stream parallel exploitation techniques.
Charles Michael Stein, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes
PDP4
2019 Stream parallelism with ordered data constraints on multi-core systems
Dalvan Griebler, Renato B. Hoffmann, Marco Danelutto, Luiz Gustavo Fernandes
J. Supercomput.4
2018 Performance of Data Mining, Media, and Financial Applications under Private Cloud Conditions
abstract
This paper contributes to a performance analysis of real-world workloads under private cloud conditions. We selected six benchmarks from PARSEC related to three mainstream application domains (financial, data mining, and media processing). Our goal was to evaluate these application domains in different cloud instances and deployment environments, concerning container or kernel-based instances and using dedicated or shared machine resources. Experiments have shown that performance varies according to the application characteristics, virtualization technology, and cloud environment. Results highlighted that financial, data mining, and media processing applications running in the LXC instances tend to outperform KVM when there is a dedicated machine resource environment. However, when two instances are sharing the same machine resources, these applications tend to achieve better performance in the KVM instances. Finally, financial applications achieved better performance in the cloud than media and data mining.
Dalvan Griebler, Adriano Vogel, Carlos A. F. Maron, Anderson M. Maliszewski, Claudio Schepke, Luiz Gustavo Fernandes
ISCC6
2018 Evaluating, Estimating, and Improving Network Performance in Container-based Clouds
abstract
Cloud computing has recently attracted a great deal of interest from both industry and academia, emerging as an important paradigm to improve resource utilization, efficiency, flexibility, and pay-per-use. However, cloud platforms inherently include a virtualization layer that imposes performance degradation on network-intensive applications. Thus, it is crucial to anticipate possible performance degradation to resolve system bottlenecks. This paper uses the Petri Nets approach to create different models for evaluating, estimating, and improving network performance in container-based cloud environments. Based on model estimations, we assessed the network bandwidth utilization of the system under different setups. Then, by identifying possible bottlenecks, we show how the system could be modified to improve performance. We then tested how the model would behave through real-world experiments. When the model indicates probable bandwidth saturation, we propose a link aggregation approach to increase bandwidth, using lightweight virtualization to reduce virtualization overhead. Results reveal that our model anticipates the structural and behavioral characteristics of the network in the cloud environment. Therefore, it systematically improves network efficiency, which saves effort, time, and money.
Cassiano Rista, Marcelo Teixeira, Dalvan Griebler, Luiz Gustavo Fernandes
ISCC4
2018 Efficient NAS Benchmark Kernels with C++ Parallel Programming
abstract
Benchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As a consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and an efficient parallelization on average.
Dalvan Griebler, Junior Loff, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes
PDP5
2017 A High-Level DSL for Geospatial Visualizations with Multi-core Parallelism Support
abstract
The amount of data generated worldwide associated with geolocalization has exponentially increased over the last decade due to social networks, population demographics, and the popularization of Global Positioning Systems. Several methods for geovisualization have already been developed, but many of them are focused on a specific application or require learning a variety of tools and programming languages. It becomes even more difficult when users have to manage a large amount of data because state-of-the-art alternatives require the use of third-party pre-processing tools. We present a novel Domain-Specific Language (DSL), which focuses on large data geovisualizations. Through a compiler, we support automatic visualization generations and data pre-processing. The system takes advantage of multi-core parallelism to speed-up data pre-processing abstractly. Our experiments were designated to highlight the programming effort and performance of our DSL. The results have shown a considerable programming effort reduction and efficient parallelism support with respect to the sequential version.
Cleverson Ledur, Dalvan Griebler, Isabel H. Manssour, Luiz Gustavo Fernandes
COMPSAC (1)4
2017 An Intra-Cloud Networking Performance Evaluation on CloudStack Environment
abstract
Infrastructure-as-a-Service (IaaS) is a cloud on-demand commodity built on top of virtualization technologies and managed by IaaS tools. In this scenario, performance is a relevant matter because a set of aspects may impact and increase the system overhead. Specific on the network, the use of virtualized capabilities may cause performance degradation (eg.,latency, throughput). The goal of this paper is to contribute to networking performance evaluation, providing new insights for private IaaS clouds. To achieve our goal, we deploy CloudStack environments and conduct experiments with different configurations and techniques. The research findings demonstrate that KVM-based cloud instances have small network performance degradation regarding throughput (about 0.2% for coarse-grained and 6.8% for fine-grained messages) while container-based instances have even better results. On the other hand, the KVM instances present worst latency (about 12.4% on coarse-grained and two times more on fine-grained messages w. r. t. native environment) and better in container-based instances, where the performance results are close to the native environment. Furthermore, we demonstrate a performance optimization of applications running on KVM.
Adriano Vogel, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes
PDP4
2016 Private IaaS Clouds: A Comparative Analysis of OpenNebula, CloudStack and OpenStack
abstract
Despite the evolution of cloud computing in recent years, the performance and comprehensive understanding of the available private cloud tools are still under research. This paper contributes to an analysis of the Infrastructure as a Service (IaaS) domain by mapping new insights and discussing the challenges for improving cloud services. The goal is to make a comparative analysis of OpenNebula, OpenStack and CloudStack tools, evaluating their differences on support for flexibility and resiliency. Also, we aim at evaluating these three cloud tools when they are deployed using a mutual hypervisor (KVM) for discovering new empirical insights. Our research results demonstrated that OpenStack is the most resilient and CloudStack is the most flexible for deploying an IaaS private cloud. Moreover, the performance experiments indicated some contrasts among the private IaaS cloud instances when running intensive workloads and scientific applications.
Adriano Vogel, Dalvan Griebler, Carlos A. F. Maron, Claudio Schepke, Luiz Gustavo Fernandes
PDP5
2016 A comparative study of energy-aware scheduling algorithms for computational grids
Silvana Teodoro, Andriele Busatto do Carmo, Daniel Couto Adornes, Luiz Gustavo Fernandes
J. Syst. Softw.4
2015 Towards a Domain-Specific Language for geospatial data visualization maps with Big Data sets
abstract
Data visualization is an alternative for representing information and helping people gain faster insights. However, the programming/creating of a visualization for large data sets is still a challenging task for users with low-level of software development knowledge. Our goal is to increase the productivity of experts who are familiar with the application domain. Therefore, we proposed an external Domain-Specific Language (DSL) that allows massive input of raw data and provides a small dictionary with suitable data visualization keywords. Also, we implemented it to support efficient data filtering operations and generate HTML or Javascript output code files (using Google Maps API). To measure the potential of our DSL, we evaluated four types of geospatial data visualization maps with four different technologies. The experiment results demonstrated a productivity gain when compared to the traditional way of implementing (e.g., Google Maps API, OpenLayers, and Leaflet), and efficient algorithm implementation.
Cleverson Ledur, Dalvan Griebler, Isabel H. Manssour, Luiz Gustavo Fernandes
AICCSA4
2015 A Unified MapReduce Domain-Specific Language for Distributed and Shared Memory Architectures
abstract
MapReduce is a suitable and efficient parallel programming pattern for processing big data analysis.In recent years, many frameworks/languages have implemented this pattern to achieve high performance in data mining applications, particularly for distributed memory architectures (e.g., clusters).Nevertheless, the industry of processors is now able to offer powerful processing on single machines (e.g., multi-core).Thus, these applications may address the parallelism in another architectural level.The target problems of this paper are code reuse and programming effort reduction since current solutions do not provide a single interface to deal with these two architectural levels.Therefore, we propose a unified domain-specific language in conjunction with transformation rules for code generation for Hadoop and Phoenix++.We selected these frameworks as state-of-the-art MapReduce implementations for distributed and shared memory architectures, respectively.Our solution achieves a programming effort reduction from 41.84% and up to 95.43% without significant performance losses (below the threshold of 3%) compared to Hadoop and Phoenix++.
Daniel Couto Adornes, Dalvan Griebler, Cleverson Ledur, Luiz Gustavo Fernandes
SEKE4
2015 Coding Productivity in MapReduce Applications for Distributed and Shared Memory Architectures
abstract
MapReduce was originally proposed as a suitable and efficient approach for analyzing and processing large amounts of data. Since then, many researches contributed with MapReduce implementations for distributed and shared memory architectures. Nevertheless, different architectural levels require different optimization strategies in order to achieve high-performance computing. Such strategies in turn have caused very different MapReduce programming interfaces among these researches. This paper presents some research notes on coding productivity when developing MapReduce applications for distributed and shared memory architectures. As a case study, we introduce our current research on a unified MapReduce domain-specific language with code generation for Hadoop and Phoenix++, which has achieved a coding productivity increase from 41.84% and up to 94.71% without significant performance losses (below 3%) compared to those frameworks.
Daniel Couto Adornes, Dalvan Griebler, Cleverson Ledur, Luiz Gustavo Fernandes
Int. J. Softw. Eng. Knowl. Eng.4
2014 JAR tool: using document analysis for improving the throughput of high performance printing environments
abstract
Digital printers have consistently improved their speed in the past years. Meanwhile, the need for document personalization and customization has increased. As a consequence of these two facts, the traditional rasterization process has become a highly demanding computational step in the printing workflow. Moreover, Print Service Providers are now using multiple RIP engines to speed up the whole document rasterization process, and depending on the input document characteristics the rasterization process may not achieve the print-engine speed creating a unwanted bottleneck. In this scenario, we developed a tool called Job Adaptive Router (JAR) aiming at improving the throughput of the rasterization process through a clever load balance among RIP engines which is based on information obtained by the analysis of input documents content. Furthermore, along with this tool we propose some strategies that consider relevant characteristics of documents, such as transparency and reusability of images, to split the job in a more intelligent way. The obtained results confirm that the use of the proposed tool improved the rasterization process performance.
Mariana Luderitz Kolberg, Luiz Gustavo Fernandes, Mateus Raeder, Carolina Fonseca
ACM Symposium on Document Engineering2
2014 Evaluating the Impact of Transactional Characteristics on the Performance of Transactional Memory Applications
abstract
Transactional Memory (TM) is reputed by many researchers to be a promising solution to ease parallel programming on multicore processors. This model provides the scalability of fine-grained locking while avoiding common issues of traditional mechanisms, such as deadlocks. During these almost twenty years of research, several TM systems and benchmarks have been proposed. However, TM is not yet widely adopted by the scientific community to develop parallel applications due to unanswered questions in the literature, such as "how to identify if a parallel application can exploit TM to achieve better performance?" or "what are the reasons of poor performances of some TM applications?". In this work, we contribute to answer those questions through a comparative evaluation of a set of TM applications on four different state- of-the-art TM systems. Moreover, we identify some of the most important TM characteristics that impact directly the performance of TM applications. Our results can be useful to identify opportunities for optimizations.
Fernando Rui, Márcio Castro 0001, Dalvan Griebler, Luiz Gustavo Fernandes
PDP4
2014 Performance and Usability Evaluation of a Pattern-Oriented Parallel Programming Interface for Multi-Core Architectures
Dalvan Griebler, Daniel Couto Adornes, Luiz Gustavo Fernandes
SEKE3
2012 Dynamic Thread Mapping Based on Machine Learning for Transactional Memory Applications
Márcio Castro 0001, Fabrício Góes, Luiz Gustavo Fernandes, Jean-François Méhaut
Euro-Par3
2011 Analysis and Tracing of Applications Based on Software Transactional Memory on Multicore Architectures
abstract
Transactional Memory (TM) is a new programming paradigm that offers an alternative to traditional lock-based concurrency mechanisms. It offers a higher-level programming interface and promises to greatly simplify the development of correct concurrent applications on multicore architectures. However, simplicity often comes with an important performance deterioration and given the variety of TM implementations it is still a challenge to know what kind of applications can really take advantage of TM. In order to gain some insight on these issues, helping developers to understand and improve the performance of TM applications, we propose a generic approach for collecting and tracing relevant information about transactions. Our solution can be applied to different Software Transactional Memory (STM) libraries and applications as it does not modify neither the target application nor the STM library source codes. We show that the collected information can be helpful in order to comprehend the performance of TM applications.
Márcio Castro 0001, Kiril Georgiev, Vania Marangozova-Martin, Jean-François Méhaut, Luiz Gustavo Fernandes, Miguel Santana
PDP5
2009 Job profiling in high performance printing
abstract
Digital presses have consistently improved their speed in the past ten years. Meanwhile, the need for document personalization and customization has increased. As a consequence of these two facts, the traditional RIP (Raster Image Processing) process has became a highly demanding computational step in the print workflow. Print Service Providers (PSP) are now using multiple RIP engines and parallelization strategies to speed up the whole ripping process which is currently based on a per-page base. Nevertheless, these strategies are not optimized in terms of assuring the best Return On Investment (ROI) for the RIP engines. Depending on the input document jobs characteristics, the ripping step may not achieve the print-engine speed creating a unwanted bottleneck. The aim of this paper is to present a way to improve the ROI of PSPs proposing a profiling strategy which enables the optimal usage of RIPs for specific jobs features ensuring that jobs are always consumed at least at engine speed. The profiling strategy is based on a per-page analysis of input PDF jobs identifying their key components. This work introduces a profiler tool to extract information from jobs and some metrics to predict a job ripping cost based on its profile. This information is extremely useful during the job splitting step, since jobs can be split in a clever way. This improves the load balance of the allocated RIPs engines and makes the overall process faster. Finally, experimental results are presented in order to evaluate both, the profiler and the proposed metrics.
Thiago Nunes, Fabio Giannetti, Mariana Luderitz Kolberg, Rafael Nemetz, Alexis Cabeda Faria, Luiz Gustavo Fernandes
ACM Symposium on Document Engineering6
2009 NUMA-ICTM: A parallel version of ICTM exploiting memory placement strategies for NUMA machines
abstract
In geophysics, the appropriate subdivision of a region into segments is extremely important. ICTM (interval categorizer tesselation model) is an application that categorizes geographic regions using information extracted from satellite images. The categorization of large regions is a computational intensive problem, what justifies the proposal and development of parallel solutions in order to improve its applicability. Recent advances in multiprocessor architectures lead to the emergence of NUMA (non-uniform memory access) machines. In this work, we present NUMA-ICTM: a parallel solution of ICTM for NUMA machines. First, we parallelize ICTM using OpenMP. After, we improve the OpenMP solution using the MAI (memory affinity interface) library, which allows a control of memory allocation in NUMA machines. The results show that the optimization of memory allocation leads to significant performance gains over the pure OpenMP parallel solution.
Márcio Castro 0001, Luiz Gustavo Fernandes, Christiane Pousa Ribeiro, Jean-François Méhaut, Marilton S. de Aguiar
IPDPS2
2009 Memory Affinity for Hierarchical Shared Memory Multiprocessors
abstract
Currently, parallel platforms based on large scale hierarchical shared memory multiprocessors with Non-Uniform Memory Access (NUMA) are becoming a trend in scientific High Performance Computing (HPC). Due to their memory access constraints, these platforms require a very careful data distribution. Many solutions were proposed to resolve this issue. However, most of these solutions did not include optimizations for numerical scientific data (array data structures) and portability issues. Besides, these solutions provide a restrict set of memory policies to deal with data placement. In this paper, we describe an user-level interface named Memory Affinity interface (MAi), which allows memory affinity control on Linux based cache-coherent NUMA (ccNUMA) platforms. Its main goals are, fine data control, flexibility and portability. The performance of MAi is evaluated on three ccNUMA platforms using numerical scientific HPC applications, the NAS Parallel Benchmarks and a Geophysics application. The results show important gains (up to 31\%) when compared to Linux default solution.
Christiane Pousa Ribeiro, Jean-François Méhaut, Alexandre Carissimi, Márcio Castro 0001, Luiz Gustavo Fernandes
SBAC-PAD5
2008 Parallel Verified Linear System Solver for Uncertain Input Data
abstract
This paper presents a new parallel implementation for solving dense interval linear systems with verified computing. The use of intervals appears as one possible way to handle the uncertainty of input data in real problems. A verified method using midpoint-radius arithmetic and directed roundings was combined with optimized libraries such as SCALAPACK and PBLAS to provide a free, fast, reliable and accurate solver. Accuracy and performance results for executing this implementation in a cluster are shown. It is the authors opinion that the combination of verified and parallel computing is a powerful tool that could be used for several other mathematical problems.
Mariana Luderitz Kolberg, Márcio Dorn, Luiz Gustavo Fernandes, Gerd Bohlender
SBAC-PAD3
2004 Parallel PEPS Tool Performance Analysis Using Stochastic Automata Networks
Lucas Baldo, Luiz Gustavo Fernandes, Paulo Roisenberg, Pedro Velho, Thais Webber
Euro-Par2
2003 Performance Analysis Issues for Parallel Implementations of Propagation Algorithm
abstract
We present a theoretical study to evaluate the performance of a family of parallel implementations of the propagation algorithm. The propagation algorithm is used to an image interpolation application. The theoretical performance analysis is based on the construction of generic models using stochastic automata networks (SAN) formalism to describe each implementation scheme. The prediction results can be compared to the achieved performance in some real test cases to verify the accuracy of our modeling technique. The main contribution is to point out the advantages and problems of our approach to the development of generic models of parallel implementations.
Leonardo Brenner, Luiz Gustavo Fernandes, Paulo Fernandes 0001, Afonso Sales
SBAC-PAD2