EDBT 2026 Demo / reviewers in the wild / expert
Dalvan Griebler
dblp:135/0603
· DBLP profile ↗
53ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0002-4690-3964ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 13 since 2021Software engineering, systems software and programming languages · 12 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Computer networks · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Compression of Clinical Signals: A Hybrid TinyML-Federated Learning Approach in Edge-Fog Environments
João Medeiros Dinis, Jean Schmith, Diego Kreutz, Guilherme Galante, Dalvan Griebler, Rodrigo da Rosa Righi |
CLOSER | 5 |
| 2026 | Orthodontic Extraction Decision Support Using Deep Learning on Lateral Cephalometric Radiographs
João Pedro De Moura Medeiros, Adriel Silva de Araújo, Vinicius Chrisosthemos Teixeira, Piedro Rockembach Nunes, Fernando Jung Lau, Sunna Imtiaz Ahmad, Quinn Roederer, Vinicius Dutra, Dalvan Griebler, Hakan Turkkahraman, Márcio Sarroglia Pinho |
COMPSAC | 9 |
| 2026 | Sentence representations for semantic textual similarity: A systematic reviewabstractIn natural language processing (NLP), generating semantically-rich representations of sentences can improve performance on multiple tasks, such as question answering, duplicate detection, sentiment analysis, and machine translation. Recent approaches to NLP using machine learning can produce text representations that carry syntactic and semantic information. This article surveys recent works on generating sentence representations for semantic textual similarity tasks. We conduct our survey using a systematic literature review approach. We retrieve papers from several digital libraries and summarize their key techniques and findings. We propose a taxonomy to facilitate the understanding of the semantic textual similarity task on the sentence level. In our analysis, we describe the current state-of-the-art in sentence representation for semantic textual similarity and propose a guideline for working on this task. • Identification of the main research on semantic similarity between sentences. • Taxonomy to define the field of semantic similarity. • Identification of state-of-the-art approaches for existing datasets. • Guidelines for working on semantic similarity between sentences. Larissa Guder, João Paulo Aires, Hígor Uélinton Silva, Felipe Meneguzzi, Dalvan Griebler |
Comput. Speech Lang. | 5 |
| 2026 | Enabling large-state stream processing on memory-constrained multi-core systems via key-value stores
Andrea Filippi, Gabriele Mencagli, Dalvan Griebler |
Future Gener. Comput. Syst. | 3 |
| 2026 | Scalable join operators over data streams with shared-nothing parallelismabstractStream joins are among the most computationally demanding stateful operators in stream processing. Tuples arriving from different streams must be analyzed on-the-fly to identify pairs that satisfy specific user-defined conditions. Since buffering all tuples from the input streams is infeasible due to memory constraints, stream joins are typically computed over a subset of the received tuples. This subset is often organized either by a specific time interval ( online interval joins ) or by fixed-length temporal windows with a defined slide ( window joins ). In this paper, we present various parallel patterns for stream join computation, aimed at effectively increasing overall query throughput. Our focus is on leveraging shared-nothing parallelism to provide portable parallelization strategies that can be efficiently executed on modern scale-in and scale-out Stream Processing Engines. Among the proposed patterns, the one exhibiting hybrid parallelism emerges as the most promising in terms of performance and load balancing. The experimental evaluation highlights the performance characteristics of the proposed patterns using real-world datasets and diverse key distributions, and compares them with state-of-the-art solutions, confirming the effectiveness of the parallel pattern with hybrid parallelism against the main competitors. Gabriele Mencagli, Yuriy Rymarchuk, Dalvan Griebler |
Inf. Syst. | 3 |
| 2025 | Automatic Synthesis of Specialized Hash FunctionsabstractThis paper introduces a technique for synthesizing hash functions specialized to particular byte formats. This code generation method leverages three prevalent patterns: (i) fixed-length keys, (ii) keys with common subsequences, and (iii) keys ranging on predetermined sequences of bytes. Code generation involves two algorithms: one identifies relevant regular expressions within key examples, and the other generates specialized hash functions based on these expressions. Comparative analysis demonstrates that the synthetic functions outperform the general-purpose hashes in the C++ Standard Template Library and the Google Abseil Library when keys are given in ascending, normal or uniform distribution. In applications where low-mixing hashes are acceptable, the synthetic functions achieve speedups ranging from 2% to 11% on full benchmarks, and speedups of almost 50x once only hashing speed is considered. Renato B. Hoffmann, Leonardo G. Faé, Dalvan Griebler, Xinliang David Li, Fernando Magno Quintão Pereira |
CGO | 3 |
| 2025 | A Novel AI-driven Automated Orthodontic Model Analysis to Improve Classification of Orthodontic Extraction CasesabstractMalocclusion, a prevalent dental condition worldwide, necessitates orthodontic intervention to correct tooth misalignment and improve oral health. Treatment can involve extraction of permanent teeth, depending on dental crowding, jaw relationships, and facial aesthetics. Today, clinical decision support systems have introduced machine learning (ML) to assist orthodontists in determining optimal treatment plans. This study explores the development of a novel, fully automated method for extracting dentoalveolar features from 3D intraoral scans (IOS), aiming to enhance orthodontic decision-making. Using deep learning-based IOS segmentation as basis, dental measurements were developed and utilized to train supervised ML classifiers, including support vector machines (SVM), logistic regression, decision trees, and random forests. An ensemble of SVM models demonstrated the highest accuracy (73%) in predicting extraction decisions, with these novel domain-specific features proving more informative than traditional dental arch measurements. While we can make further improvements not only in the automated segmentation but also by applying feature selection, the results highlight the potential of AI-driven analysis to streamline orthodontic workflows, reduce manual intervention and improve clinical efficiency. Sunna Imtiaz Ahmad, Adriel Silva de Araújo, Vinicius Crisosthemos Teixeira, Carlos Falcão de Azevedo Gomes, Vinicius Dutra, Quinn Roederer, R. Scott Conley, Dalvan Griebler, Márcio Sarroglia Pinho, Hakan Turkkahraman |
COMPSAC | 8 |
| 2025 | NPB-PSTL: C++ STL Algorithms with Parallel Execution Policies in NAS Parallel BenchmarksabstractThe C++ language continually evolves through formal specifications established by its standards committee, proposing new features to maintain $\mathrm{C}++$ as a relevant programming language while improving usability, performance, and portability across platforms. With the addition of parallel Standard Template Library (STL) algorithms in C++17, programmers can now leverage parallel processing capabilities via vendor-neutral parallel execution policies. This study presents an adaptation of the NAS Parallel Benchmarks (NPB)—a well-established suite of applications for evaluating parallel architectures-by porting its sequential C-style code to use C++ STL abstractions and performance-portable parallelism features. Our goals are to (1) assess the suitability of C++ STL for scientific applications like the ones in the NPB and (2) provide a comparative performance and portability of STL algorithms’ parallel execution policies across different multicore architectures (x86 and AArch64). Results indicate that the performance of parallel STL algorithms is often close to that of optimized handwritten versions (OpenMP, Intel TBB, and FastFlow) on different architectures, with notable shortfalls. Across all NPB benchmarks, the STL algorithms’ geometric mean shows sequential execution times that are between 3.76% and $\mathrm{6. 9 \%}$ higher, while parallel executions may reach a geometric mean of up to $\mathrm{2 1. 2 1 \%}$ higher execution time. Junior Loff, Renato B. Hoffmann, Nicola Bianchessi, Leonardo Mallmann, Dalvan Griebler, Walter Binder |
PDP | 5 |
| 2025 | Performance, Portability, and Productivity of HIP on GPUs with NAS Parallel BenchmarksabstractGraphics Processing Units (GPUs) are powerful, massively parallel processors that have become ubiquitous in modern computing. In recent years, the GPU market has diversified, with vendors like AMD and Intel offering high-performance alternatives to NVIDIA. However, most applications are written using NVIDIA’s CUDA API, which is incompatible with non-NVIDIA GPUs, creating significant challenges for developers who must port their code to different architectures. To address this issue, AMD developed the Heterogeneous-Compute Interface for Portability (HIP), an open-source API for cross-vendor GPU programming. However, HIP is relatively new, leaving gaps in the literature regarding its performance, portability, and productivity. In this paper, we evaluate HIP using the NAS Parallel Benchmarks (NPB), a CFD-based suite maintained by NASA. We present the first HIP-based implementation of NPB and conduct experiments on integrated and discrete GPUs from NVIDIA, AMD, and Intel. Our results provide novel insights into HIP’s performance and portability, particularly for integrated GPUs and Intel discrete GPUs, which have been underrepresented in prior studies. We also assess productivity using different metrics to quantify the programming effort of HIP-based implementations. This work addresses key gaps in the literature, offering valuable data and insights for developers targeting emerging GPU architectures. Gabriell Alves de Araujo, Dalvan Griebler, Luiz Gustavo Fernandes |
SBAC-PAD | 2 |
| 2025 | Optimization of resource-aware parallel and distributed computing: a reviewabstractThis paper presents a review of state-of-the-art solutions concerning the optimization of computing in the field of parallel and distributed systems. Firstly, we contribute by identifying resources and quality metrics in this context including servers, network interconnects, storage systems, computational devices as well as execution time/performance, energy, security, and error vulnerability, respectively. We subsequently identify commonly used problem formulations and algorithms for integer linear programming, greedy algorithms, dynamic programming, genetic algorithms, particle swarm optimization, ant colony optimization, game theory, and reinforcement learning. Afterward, we characterize frequently considered optimization problems by stating these terms in domains such as data centers, cloud, fog, blockchain, high performance, and volunteer computing. Based on the extensive analysis, we identify how particular resources and corresponding quality metrics are considered in these domains and which problem formulations are used for which system types, either parallel or distributed environments. This allows us to formulate open research problems and challenges in this field and analyze research interest in problem formulations/domains in recent years. Pawel Czarnul, Marcel Antal, Hamza Baniata, Dalvan Griebler, Attila Kertész, Christoph W. Kessler, Andreas Kouloumpris, Salko Kovacic, András Márkus, Maria K. Michael, Panagiota Nikolaou, Isil Öz, Radu Prodan, Gordana Rakic |
J. Supercomput. | 4 |
| 2024 | Multiview Machine Learning Classification of Tooth Extraction in Orthodontics Using Intraoral ScansabstractOrthodontic treatment planning often involves de-ciding whether to extract teeth, a critical and irreversible decision. Integrating machine learning (ML) can enhance decision-making. This study proposes using Intraoral Scans (IOS) 3D models to predict extraction/non-extraction binary decisions with ML models. We leverage a multiview approach, using images taken from multiple points of view of the 3D model. The methodology involved a dataset composed of preprocessed IOS from 181 subjects and an experimental procedure that evaluated multiple ML models in their ability to classify subjects using either grayscale pixel intensities or radiomic features. The results indicated that a logistic model applied to the radiomic features from the back and frontal views of the 3D models was one of the best model candidates, achieving a test accuracy of 70 % and F1 score of. 73 and. 65 for non-extraction and extraction cases, respectively. Overall, these findings indicate that a multiview approach to IOS 3D models can be used to predict extraction/non-extraction decisions. In addition, the results suggest that radiomic features provide useful information in the analysis of IOS data. Carlos Falcão de Azevedo Gomes, Adriel Silva de Araújo, Sunna Imtiaz Ahmad, Maurício Cecílio Magnaguagno, Vinicius Crisosthemos Teixeira, Anushri Singh Rajapuri, Quinn Roederer, Dalvan Griebler, Vinicius Dutra, Hakan Turkkahraman, Márcio Sarroglia Pinho |
COMPSAC | 8 |
| 2024 | MPR: An MPI Framework for Distributed Self-adaptive Stream Processing
Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes, Walter Binder |
Euro-Par (3) | 2 |
| 2024 | Benchmarking parallel programming for single-board computers
Renato B. Hoffmann, Dalvan Griebler, Rodrigo da Rosa Righi, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 2 |
| 2024 | General-purpose data stream processing on heterogeneous architectures with WindFlowabstractMany emerging applications analyze data streams by running graphs of communicating tasks called operators. To develop and deploy such applications, Stream Processing Systems (SPSs) like Apache Storm and Flink have been made available to researchers and practitioners. They exhibit imperative or declarative programming interfaces to develop operators running arbitrary algorithms working on structured or unstructured data streams. In this context, the interest in leveraging hardware acceleration with GPUs has become more pronounced in high-throughput use cases. Unfortunately, GPU acceleration has been studied for relational operators working on structured streams only, while non-relational operators have often been overlooked. This paper presents WindFlow, a library supporting the seamless GPU offloading of general partitioned-stateful operators, extending the range of operators that benefit from hardware acceleration. Its design provides high throughput still exposing a high-level API to users compared with the raw utilization of GPUs in Apache Flink. Gabriele Mencagli, Massimo Torquati, Dalvan Griebler, Alessandra Fais, Marco Danelutto |
J. Parallel Distributed Comput. | 3 |
| 2024 | Performance and programmability of GrPPI for parallel stream processing on multi-coresabstractAbstract GrPPI library aims to simplify the burdening task of parallel programming. It provides a unified, abstract, and generic layer while promising minimal overhead on performance. Although it supports stream parallelism, GrPPI lacks an evaluation regarding representative performance metrics for this domain, such as throughput and latency. This work evaluates GrPPI focused on parallel stream processing. We compare the throughput and latency performance, memory usage, and programmability of GrPPI against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks and benchmarks with handwritten parallel code using the same backends supported by GrPPI. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is often competitive with handwritten parallel code, the infeasibility of fine-tuning GrPPI is a crucial drawback for emerging applications. Despite this, programmability experiments estimate that GrPPI can potentially reduce the development time of parallel applications by about three times. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
J. Supercomput. | 2 |
| 2024 | Enhancing self-adaptation for efficient decision-making at run-time in streaming applications on multicoresabstractAbstract Parallel computing is very important to accelerate the performance of computing applications. Moreover, parallel applications are expected to continue executing in more dynamic environments and react to changing conditions. In this context, applying self-adaptation is a potential solution to achieve a higher level of autonomic abstractions and runtime responsiveness. In our research, we aim to explore and assess the possible abstractions attainable through the transparent management of parallel executions by self-adaptation. Our primary objectives are to expand the adaptation space to better reflect real-world applications and assess the potential for self-adaptation to enhance efficiency. We provide the following scientific contributions: (I) A conceptual framework to improve the designing of self-adaptation; (II) A new decision-making strategy for applications with multiple parallel stages; (III) A comprehensive evaluation of the proposed decision-making strategy compared to the state-of-the-art. The results demonstrate that the proposed conceptual framework can help design and implement self-adaptive strategies that are more modular and reusable. The proposed decision-making strategy provides significant gains in accuracy compared to the state-of-the-art, increasing the parallel applications’ performance and efficiency. Adriano Vogel, Marco Danelutto, Massimo Torquati, Dalvan Griebler, Luiz Gustavo Fernandes |
J. Supercomput. | 4 |
| 2023 | A Latency, Throughput, and Programmability Perspective of GrPPI for Streaming on Multi-coresabstractSeveral solutions aim to simplify the burdening task of parallel programming. The GrPPI library is one of them. It allows users to implement parallel code for multiple backends through a unified, abstract, and generic layer while promising minimal overhead on performance. An outspread evaluation of GrPPI regarding stream parallelism with representative metrics for this domain, such as throughput and latency, was not yet done. In this work, we evaluate GrPPI focused on stream processing. We evaluate performance, memory usage, and programming effort and compare them against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is competitive with handwritten code in some cases, in other cases, the infeasibility of fine-tuning GrPPI is a crucial drawback. Despite this, programmability experiments estimate that GrPPI has the potential to reduce by about three times the development time of parallel applications. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, André Sacilotto Santos, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
PDP | 2 |
| 2023 | Revisiting self-adaptation for efficient decision-making at run-time in parallel executionsabstractSelf-adaptation is a potential alternative to provide a higher level of autonomic abstractions and run-time responsiveness in parallel executions. However, the recurrent problem is that self-adaptation is still limited in flexibility and efficiency. For instance, there is a lack of mechanisms to apply adaptation actions and efficient decision-making strategies to decide which configurations should be conveniently enforced at run-time. In this work, we are interested in providing and evaluating potential abstractions achievable with self-adaptation transparently managing parallel executions. Therefore, we provide a new mechanism to support self-adaptation in applications with multiple parallel stages executed in multi-cores. Moreover, we reproduce, reimplement, and evaluate an existing decision-making strategy in our scenario. The observations from the results show that the proposed mechanism for self-adaptation can provide new parallelism abstractions and autonomous responsiveness at run-time. On the other hand, there is a need for more accurate decision-making strategies to enable efficient executions of applications in resource-constrained scenarios like multi-cores. Adriano Vogel, Marco Danelutto, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 3 |
| 2023 | Message from the General ChairsabstractOn behalf of the organizing committee, we welcome you to the 35th International Symposium on Computer Architecture and High-Performance Computing (SBAC-PAD 2023), in Porto Alegre, Rio Grande do Sul, Brazil. Since 1987, SBAC-PAD has continuously presented an overview of new developments, applications, and trends in parallel and distributed computing technologies. SBAC-PAD is open to faculty members, researchers, specialists, and graduate students. We have striven to continue this tradition and have considered papers across a whole range of computer system design related topics. Tiago Ferreto, Dalvan Griebler, César A. F. De Rose |
SBAC-PAD | 2 |
| 2023 | NAS Parallel Benchmarks with CUDA and beyondabstractAbstract NAS Parallel Benchmarks (NPB) is a standard benchmark suite used in the evaluation of parallel hardware and software. Several research efforts from academia have made these benchmarks available with different parallel programming models beyond the original versions with OpenMP and MPI. This work joins these research efforts by providing a new CUDA implementation for NPB. Our contribution covers different aspects beyond the implementation. First, we define design principles based on the best programming practices for GPUs and apply them to each benchmark using CUDA. Second, we provide ease of use parametrization support for configuring the number of threads per block in our version. Third, we conduct a broad study on the impact of the number of threads per block in the benchmarks. Fourth, we propose and evaluate five strategies for helping to find a better number of threads per block configuration. The results have revealed relevant performance improvement solely by changing the number of threads per block, showing performance improvements from 8% up to 717% among the benchmarks. Fifth, we conduct a comparative analysis with the literature, evaluating performance, memory consumption, code refactoring required, and parallelism implementations. The performance results have shown up to 267% improvements over the best benchmarks versions available. We also observe the best and worst design choices, concerning code size and the performance trade‐off. Lastly, we highlight the challenges of implementing parallel CFD applications for GPUs and how the computations impact the GPU's behavior. Gabriell Alves de Araujo, Dalvan Griebler, Dinei A. Rockenbach, Marco Danelutto, Luiz Gustavo Fernandes |
Softw. Pract. Exp. | 2 |
| 2023 | Micro-batch and data frequency for stream processing on multi-cores
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
J. Supercomput. | 2 |
| 2022 | Analyzing Programming Effort Model Accuracy of High-Level Parallel Programs for Stream ProcessingabstractOver the years, several Parallel Programming Models (PPMs) have supported the abstraction of programming complexity for parallel computer systems. However, few studies aim to evaluate the productivity reached by such abstractions since this is a complex task that involves human beings. There are several studies to develop predictive methods to estimate the effort required to develop software applications. In order to evaluate the reliability of such metrics, it is necessary to assess the accuracy in different programming paradigms. In this work, we used the data of an experiment conducted with beginners in parallel programming to determine the effort required for implementing stream parallelism using FastFlow, SPar, and TBB. Our results show that some traditional software effort estimation models, such as COCOMO II, fall short. In contrast, Planning Poker could contribute toward a parallel-aware effort model. Gabriella Andrade, Dalvan Griebler, Rodrigo Pereira dos Santos, Christoph W. Kessler, August Ernstsson, Luiz Gustavo Fernandes |
SEAA | 2 |
| 2022 | Evaluating Micro-batch and Data Frequency for Stream Processing Applications on Multi-coresabstractIn stream processing, data arrives constantly and is often unpredictable. It can show large fluctuations in arrival frequency, size, complexity, and other factors. These fluctuations can strongly impact application latency and throughput, which are critical factors in this domain. Therefore, there is a significant amount of research on self-adaptive techniques involving elasticity or micro-batching as a way to mitigate this impact. However, there is a lack of benchmarks and tools for helping researchers to investigate micro-batching and data stream frequency implications. In this paper, we extend a benchmarking framework to support dynamic micro-batching and data stream frequency management. We used it to create custom benchmarks and compare latency and throughput aspects from two different parallel libraries. We validate our solution through an extensive analysis of the impact of micro-batching and data stream frequency on stream processing applications using Intel TBB and FastFlow, which are two libraries that leverage stream parallelism on multi-core architectures. Our results demonstrated up to 33% throughput gain over latency using micro-batches. Additionally, while TBB ensures lower latency, FastFlow ensures higher throughput in the parallel applications for different data stream frequency configurations. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 2 |
| 2022 | Towards Parallel Data Stream Processing on System-on-Chip CPU+GPU DevicesabstractData Stream Processing is a pervasive computing paradigm with a wide spectrum of applications. Traditional streaming systems exploit the processing capabilities provided by homogeneous Clusters and Clouds. Due to the transition to streaming systems suitable for IoT/Edge environments, there has been the urgent need of new streaming frameworks and tools tailored for embedded platforms, often available as System-onChips composed of a small multicore CPU and an integrated onchip GPU. Exploiting this hybrid hardware requires special care in the runtime system design. In this paper, we discuss the support provided by the WindFlow library, showing its design principles and its effectiveness on the NVIDIA Jetson Nano board. Gabriele Mencagli, Dalvan Griebler, Marco Danelutto |
PDP | 2 |
| 2022 | Self-adaptation on parallel stream processing: A systematic reviewabstractSummary A recurrent challenge in real‐world applications is autonomous management of the executions at run‐time. In this vein, stream processing is a class of applications that compute data flowing in the form of streams (e.g., video feeds, images, and data analytics), where parallel computing can help accelerate the executions. On the one hand, stream processing applications are becoming more complex, dynamic, and long‐running. On the other hand, it is unfeasible for humans to monitor and manually change the executions continuously. Hence, self‐adaptation can reduce costs and human efforts by providing a higher‐level abstraction with an autonomic/seamless management of executions. In this work, we aim at providing a literature review regarding self‐adaptation applied to the parallel stream processing domain. We present a comprehensive revision using a systematic literature review method. Moreover, we propose a taxonomy to categorize and classify the existing self‐adaptive approaches. Finally, applying the taxonomy made it possible to characterize the state‐of‐the‐art, identify trends, and discuss open research challenges and future opportunities. Adriano Vogel, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | OpenMP as runtime for providing high-level stream parallelism on multi-cores
Renato B. Hoffmann, Junior Loff, Dalvan Griebler, Luiz Gustavo Fernandes |
J. Supercomput. | 3 |
| 2021 | Assessing Coding Metrics for Parallel Programming of Stream Processing Programs on Multi-coresabstractFrom the popularization of multi-core architectures, several parallel APIs have emerged, helping to abstract the programming complexity and increasing productivity in application development. Unfortunately, only a few research efforts in this direction managed to show the usability pay-back of the programming abstraction created, because it is not easy and poses many challenges for conducting empirical software engineering. We believe that coding metrics commonly used in software engineering code measurements can give useful indicators on the programming effort of parallel applications and APIs. These metrics were designed for general purposes without considering the evaluation of applications from a specific domain. In this study, we aim to evaluate the feasibility of seven coding metrics to be used in the parallel programming domain. To do so, five stream processing applications implemented with different parallel APIs for multi-cores were considered. Our experiments have shown COCOMO II is a suitable model for evaluating the productivity of different parallel APIs targeting multi-cores on stream processing applications while other metrics are restricted to the code size. Gabriella Andrade, Dalvan Griebler, Rodrigo Pereira dos Santos, Marco Danelutto, Luiz Gustavo Fernandes |
SEAA | 2 |
| 2021 | Introducing a Stream Processing Framework for Assessing Parallel Programming InterfacesabstractStream Processing applications are spread across different sectors of industry and people's daily lives. The increasing data we produce, such as audio, video, image, and text are demanding quickly and efficiently computation. It can be done through Stream Parallelism, which is still a challenging task and most reserved for experts. We introduce a Stream Processing framework for assessing Parallel Programming Interfaces (PPIs). Our framework targets multi-core architectures and C++ stream processing applications, providing an API that abstracts the details of the stream operators of these applications. Therefore, users can easily identify all the basic operators and implement parallelism through different PPIs. In this paper, we present the proposed framework, implement three applications using its API, and show how it works, by using it to parallelize and evaluate the applications with the PPIs Intel TBB, FastFlow, and SPar. The performance results were consistent with the literature. Adriano Marques Garcia, Dalvan Griebler, Luiz Gustavo Fernandes, Claudio Schepke |
PDP | 2 |
| 2021 | Towards On-the-fly Self-Adaptation of Stream Parallel PatternsabstractStream processing applications compute streams of data and provide insightful results in a timely manner, where parallel computing is necessary for accelerating the application executions. Considering that these applications are becoming increasingly dynamic and long-running, a potential solution is to apply dynamic runtime changes. However, it is challenging for humans to continuously monitor and manually self-optimize the executions. In this paper, we propose self-adaptiveness of the parallel patterns used, enabling flexible on-the-fly adaptations. The proposed solution is evaluated with an existing programming framework and running experiments with a synthetic and a real-world application. The results show that the proposed solution is able to dynamically self-adapt to the most suitable parallel pattern configuration and achieve performance competitive with the best static cases. The feasibility of the proposed solution encourages future optimizations and other applicabilities. Adriano Vogel, Gabriele Mencagli, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 3 |
| 2021 | Latency-aware adaptive micro-batching techniques for streamed data compression on graphics processing unitsabstractSummary Stream processing is a parallel paradigm used in many application domains. With the advance of graphics processing units (GPUs), their usage in stream processing applications has increased as well. The efficient utilization of GPU accelerators in streaming scenarios requires to batch input elements in microbatches, whose computation is offloaded on the GPU leveraging data parallelism within the same batch of data. Since data elements are continuously received based on the input speed, the bigger the microbatch size the higher the latency to completely buffer it and to start the processing on the device. Unfortunately, stream processing applications often have strict latency requirements that need to find the best size of the microbatches and to adapt it dynamically based on the workload conditions as well as according to the characteristics of the underlying device and network. In this work, we aim at implementing latency‐aware adaptive microbatching techniques and algorithms for streaming compression applications targeting GPUs. The evaluation is conducted using the Lempel‐Ziv‐Storer‐Szymanski compression application considering different input workloads. As a general result of our work, we noticed that algorithms with elastic adaptation factors respond better for stable workloads, while algorithms with narrower targets respond better for highly unbalanced workloads. Charles Michael Stein, Dinei A. Rockenbach, Dalvan Griebler, Massimo Torquati, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures
Junior Loff, Dalvan Griebler, Gabriele Mencagli, Gabriell Alves de Araujo, Massimo Torquati, Marco Danelutto, Luiz Gustavo Fernandes |
Future Gener. Comput. Syst. | 2 |
| 2021 | Providing high-level self-adaptive abstractions for stream parallelism on multicoresabstractAbstract Stream processing applications are common computing workloads that demand parallelism to increase their performance. As in the past, parallel programming remains a difficult task for application programmers. The complexity increases when application programmers must set nonintuitive parallelism parameters, that is, the degree of parallelism. The main problem is that state‐of‐the‐art libraries use a static degree of parallelism and are not sufficiently abstracted for developing stream processing applications. In this article, we propose a self‐adaptive regulation of the degree of parallelism to provide higher‐level abstractions. Flexibility is provided to programmers with two new self‐adaptive strategies, one is for performance experts, and the other abstracts the need to set a performance goal. We evaluated our solution using compiler transformation rules to generate parallel code with the SPar domain‐specific language. The experimental results with real‐world applications highlighted higher abstraction levels without significant performance degradation in comparison to static executions. The strategy for performance experts achieved slightly higher performance than the one that works without user‐defined performance goals. Adriano Vogel, Dalvan Griebler, Luiz Gustavo Fernandes |
Softw. Pract. Exp. | 2 |
| 2020 | The Impact of CPU Frequency Scaling on Power Consumption of Computing Infrastructures
Adriano Marques Garcia, Matheus S. Serpa, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux |
ICCSA (6) | 3 |
| 2020 | Performance Impact of IEEE 802.3ad in Container-Based Clouds for HPC Applications
Anderson M. Maliszewski, Eduardo Roloff, Dalvan Griebler, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux |
ICCSA (6) | 3 |
| 2020 | Performance and Cost-aware HPC in Clouds: A Network Interconnection AssessmentabstractThe availability of computing resources has significantly changed due to the growing adoption of the cloud computing paradigm. Aiming at potential advantages such as cost savings through the pay-per-use method and resource allocation in a scalable/elastic way, we witnessed consistent efforts to execute high-performance computing (HPC) applications in the cloud. Performance in this environment depends heavily upon two main system components: processing power and network interconnection. If, on the one hand, allocating more powerful hardware theoretically boosts performance, on the other hand, it increases the allocation cost. In this paper, we evaluated how the network interconnection impacts on performance and cost efficiency. Our experiments were carried out using NAS Parallel Benchmarks and Alya HPC application on Microsoft Azure public cloud provider, with three different cloud instances/network interconnections. The results revealed that through the use of the accelerated networking approach, which allows the instance to have a high-performance interconnect without additional charges, the performance of HPC applications can be significantly improved with a better cost efficiency. Anderson M. Maliszewski, Eduardo Roloff, Emmanuell D. Carreño, Dalvan Griebler, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux |
ISCC | 4 |
| 2020 | Efficient NAS Parallel Benchmark Kernels with CUDAabstractNAS Parallel Benchmarks (NPB) are one of the standard benchmark suites used to evaluate parallel hardware and software. There are many research efforts trying to provide different parallel versions apart from the original OpenMP and MPI. Concerning GPU accelerators, there are only the OpenCL and OpenACC available as consolidated versions. Our goal is to provide an efficient parallel implementation of the five NPB kernels with CUDA. Our contribution covers different aspects. First, best parallel programming practices were followed to implement NPB kernels using CUDA. Second, the support of larger workloads (class B and C) allow to stress and investigate the memory of robust GPUs. Third, we show that it is possible to make NPB efficient and suitable for GPUs although the benchmarks were designed for CPUs in the past. We succeed in achieving double performance with respect to the state-of-the-art in some cases as well as implementing efficient memory usage. Fourth, we discuss new experiments comparing performance and memory usage against OpenACC and OpenCL state-of-the-art versions using a relative new GPU architecture. The experimental results also revealed that our version is the best one for all the NPB kernels compared to OpenACC and OpenCL. The greatest differences were observed for the FT and EP kernels. Gabriell Alves de Araujo, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 2 |
| 2020 | Simplifying and implementing service level objectives for stream parallelism
Dalvan Griebler, Adriano Vogel, Daniele De Sensi, Marco Danelutto, Luiz Gustavo Fernandes |
J. Supercomput. | 1 |
| 2019 | Minimizing Communication Overheads in Container-based Clouds for HPC ApplicationsabstractAlthough the industry has embraced the cloud computing model, there are still significant challenges to be addressed concerning the quality of cloud services. Network-intensive applications may not scale in the cloud due to the sharing of the network infrastructure. In the literature, performance evaluation studies are showing that the network tends to limit the scalability and performance of HPC applications. Therefore, we proposed the aggregation of Network Interface Cards (NICs) in a ready-to-use integration with the OpenNebula cloud manager using Linux containers. We perform a set of experiments using a network microbenchmark to get specific network performance metrics and NAS parallel benchmarks to analyze the performance impact on HPC applications. Our results highlight that the implementation of NIC aggregation improves network performance in terms of throughput and latency. Moreover, HPC applications have different patterns of behavior when using our approach, which depends on communication and the amount of data transferring. While network-intensive applications increased the performance up to 38%, other applications with aggregated NICs maintained the same performance or presented slightly worse performance. Anderson M. Maliszewski, Adriano Vogel, Dalvan Griebler, Eduardo Roloff, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux |
ISCC | 3 |
| 2019 | Should PARSEC Benchmarks be More Parametric? A Case Study with DedupabstractParallel applications of the same domain can present similar patterns of behavior and characteristics. Characterizing common application behaviors can help for understanding performance aspects in the real-world scenario. One way to better understand and evaluate applications' characteristics is by using customizable/parametric benchmarks that enable users to represent important characteristics at run-time. We observed that parameterization techniques should be better exploited in the available benchmarks, especially on stream processing domain. For instance, although widely used, the stream processing benchmarks available in PARSEC do not support the simulation and evaluation of relevant and modern characteristics. Therefore, our goal is to identify the stream parallelism characteristics present in PARSEC. We also implemented a ready to use parameterization support and evaluated the application behaviors considering relevant performance metrics for stream parallelism (service time, throughput, latency). We choose Dedup to be our case study. The experimental results have shown performance improvements in our parameterization support for Dedup. Moreover, this support increased the customization space for benchmark users, which is simple to use. In the future, our solution can be potentially explored on different parallel architectures and parallel programming frameworks. Carlos A. F. Maron, Adriano Vogel, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 3 |
| 2019 | Memory Performance and Bottlenecks in Multicore and GPU ArchitecturesabstractNowadays, there are several different architectures available not only for the industry, but also for normal consumers. Traditional multicore processors, GPUs, accelerators such as the Sunway SW26010, or even energy efficiency-driven processors such as the ARM family, present very different architectural characteristics. This wide range of characteristics presents a challenge for the developers of applications. Developers must deal with different instruction sets, memory hierarchies, or even different programming paradigms when programming for these architectures. Therefore, the same application can perform well when executing on one architecture, but poorly on another architecture. To optimize an application, it is important to have a deep understanding of how it behaves on different architectures. The related work in this area mostly focuses on a limited analysis encompassing execution time and energy. In this paper, we perform a detailed investigation on the impact of the memory subsystem of different architectures, which is one of the most important aspects to be considered. For this study, we performed experiments in the Broadwell CPU and Pascal GPU, using applications from the Rodinia benchmark suite. In this way, we were able to understand why an application performs well on one architecture and poorly on others. Matheus S. Serpa, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, Eduardo Henrique Molina da Cruz, Matthias Diener, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 6 |
| 2019 | Stream Parallelism on the LZSS Data Compression Application for Multi-Cores with GPUsabstractGPUs have been used to accelerate different data parallel applications. The challenge consists in using GPUs to accelerate stream processing applications. Our goal is to investigate and evaluate whether stream parallel applications may benefit from parallel execution on both CPU and GPU cores. In this paper, we introduce new parallel algorithms for the Lempel-Ziv-Storer-Szymanski (LZSS) data compression application. We implemented the algorithms targeting both CPUs and GPUs. GPUs have been used with CUDA and OpenCL to exploit inner algorithm data parallelism. Outer stream parallelism has been exploited using CPU cores through SPar. The parallel implementation of LZSS achieved 135 fold speedup using a multi-core CPU and two GPUs. We also observed speedups in applications where we were not expecting to get it using the same combine data-stream parallel exploitation techniques. Charles Michael Stein, Dalvan Griebler, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 2 |
| 2019 | Stream parallelism with ordered data constraints on multi-core systems
Dalvan Griebler, Renato B. Hoffmann, Marco Danelutto, Luiz Gustavo Fernandes |
J. Supercomput. | 1 |
| 2018 | Performance of Data Mining, Media, and Financial Applications under Private Cloud ConditionsabstractThis paper contributes to a performance analysis of real-world workloads under private cloud conditions. We selected six benchmarks from PARSEC related to three mainstream application domains (financial, data mining, and media processing). Our goal was to evaluate these application domains in different cloud instances and deployment environments, concerning container or kernel-based instances and using dedicated or shared machine resources. Experiments have shown that performance varies according to the application characteristics, virtualization technology, and cloud environment. Results highlighted that financial, data mining, and media processing applications running in the LXC instances tend to outperform KVM when there is a dedicated machine resource environment. However, when two instances are sharing the same machine resources, these applications tend to achieve better performance in the KVM instances. Finally, financial applications achieved better performance in the cloud than media and data mining. Dalvan Griebler, Adriano Vogel, Carlos A. F. Maron, Anderson M. Maliszewski, Claudio Schepke, Luiz Gustavo Fernandes |
ISCC | 1 |
| 2018 | Evaluating, Estimating, and Improving Network Performance in Container-based CloudsabstractCloud computing has recently attracted a great deal of interest from both industry and academia, emerging as an important paradigm to improve resource utilization, efficiency, flexibility, and pay-per-use. However, cloud platforms inherently include a virtualization layer that imposes performance degradation on network-intensive applications. Thus, it is crucial to anticipate possible performance degradation to resolve system bottlenecks. This paper uses the Petri Nets approach to create different models for evaluating, estimating, and improving network performance in container-based cloud environments. Based on model estimations, we assessed the network bandwidth utilization of the system under different setups. Then, by identifying possible bottlenecks, we show how the system could be modified to improve performance. We then tested how the model would behave through real-world experiments. When the model indicates probable bandwidth saturation, we propose a link aggregation approach to increase bandwidth, using lightweight virtualization to reduce virtualization overhead. Results reveal that our model anticipates the structural and behavioral characteristics of the network in the cloud environment. Therefore, it systematically improves network efficiency, which saves effort, time, and money. Cassiano Rista, Marcelo Teixeira, Dalvan Griebler, Luiz Gustavo Fernandes |
ISCC | 3 |
| 2018 | Efficient NAS Benchmark Kernels with C++ Parallel ProgrammingabstractBenchmarking is a way to study the performance of new architectures and parallel programming frameworks. Well-established benchmark suites such as the NAS Parallel Benchmarks (NPB) comprise legacy codes that still lack portability to C++ language. As a consequence, a set of high-level and easy-to-use C++ parallel programming frameworks cannot be tested in NPB. Our goal is to describe a C++ porting of the NPB kernels and to analyze the performance achieved by different parallel implementations written using the Intel TBB, OpenMP and FastFlow frameworks for Multi-Cores. The experiments show an efficient code porting from Fortran to C++ and an efficient parallelization on average. Dalvan Griebler, Junior Loff, Gabriele Mencagli, Marco Danelutto, Luiz Gustavo Fernandes |
PDP | 1 |
| 2017 | A High-Level DSL for Geospatial Visualizations with Multi-core Parallelism SupportabstractThe amount of data generated worldwide associated with geolocalization has exponentially increased over the last decade due to social networks, population demographics, and the popularization of Global Positioning Systems. Several methods for geovisualization have already been developed, but many of them are focused on a specific application or require learning a variety of tools and programming languages. It becomes even more difficult when users have to manage a large amount of data because state-of-the-art alternatives require the use of third-party pre-processing tools. We present a novel Domain-Specific Language (DSL), which focuses on large data geovisualizations. Through a compiler, we support automatic visualization generations and data pre-processing. The system takes advantage of multi-core parallelism to speed-up data pre-processing abstractly. Our experiments were designated to highlight the programming effort and performance of our DSL. The results have shown a considerable programming effort reduction and efficient parallelism support with respect to the sequential version. Cleverson Ledur, Dalvan Griebler, Isabel H. Manssour, Luiz Gustavo Fernandes |
COMPSAC (1) | 2 |
| 2017 | An Intra-Cloud Networking Performance Evaluation on CloudStack EnvironmentabstractInfrastructure-as-a-Service (IaaS) is a cloud on-demand commodity built on top of virtualization technologies and managed by IaaS tools. In this scenario, performance is a relevant matter because a set of aspects may impact and increase the system overhead. Specific on the network, the use of virtualized capabilities may cause performance degradation (eg.,latency, throughput). The goal of this paper is to contribute to networking performance evaluation, providing new insights for private IaaS clouds. To achieve our goal, we deploy CloudStack environments and conduct experiments with different configurations and techniques. The research findings demonstrate that KVM-based cloud instances have small network performance degradation regarding throughput (about 0.2% for coarse-grained and 6.8% for fine-grained messages) while container-based instances have even better results. On the other hand, the KVM instances present worst latency (about 12.4% on coarse-grained and two times more on fine-grained messages w. r. t. native environment) and better in container-based instances, where the performance results are close to the native environment. Furthermore, we demonstrate a performance optimization of applications running on KVM. Adriano Vogel, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 2 |
| 2016 | Private IaaS Clouds: A Comparative Analysis of OpenNebula, CloudStack and OpenStackabstractDespite the evolution of cloud computing in recent years, the performance and comprehensive understanding of the available private cloud tools are still under research. This paper contributes to an analysis of the Infrastructure as a Service (IaaS) domain by mapping new insights and discussing the challenges for improving cloud services. The goal is to make a comparative analysis of OpenNebula, OpenStack and CloudStack tools, evaluating their differences on support for flexibility and resiliency. Also, we aim at evaluating these three cloud tools when they are deployed using a mutual hypervisor (KVM) for discovering new empirical insights. Our research results demonstrated that OpenStack is the most resilient and CloudStack is the most flexible for deploying an IaaS private cloud. Moreover, the performance experiments indicated some contrasts among the private IaaS cloud instances when running intensive workloads and scientific applications. Adriano Vogel, Dalvan Griebler, Carlos A. F. Maron, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 2 |
| 2015 | Towards a Domain-Specific Language for geospatial data visualization maps with Big Data setsabstractData visualization is an alternative for representing information and helping people gain faster insights. However, the programming/creating of a visualization for large data sets is still a challenging task for users with low-level of software development knowledge. Our goal is to increase the productivity of experts who are familiar with the application domain. Therefore, we proposed an external Domain-Specific Language (DSL) that allows massive input of raw data and provides a small dictionary with suitable data visualization keywords. Also, we implemented it to support efficient data filtering operations and generate HTML or Javascript output code files (using Google Maps API). To measure the potential of our DSL, we evaluated four types of geospatial data visualization maps with four different technologies. The experiment results demonstrated a productivity gain when compared to the traditional way of implementing (e.g., Google Maps API, OpenLayers, and Leaflet), and efficient algorithm implementation. Cleverson Ledur, Dalvan Griebler, Isabel H. Manssour, Luiz Gustavo Fernandes |
AICCSA | 2 |
| 2015 | A Unified MapReduce Domain-Specific Language for Distributed and Shared Memory ArchitecturesabstractMapReduce is a suitable and efficient parallel programming pattern for processing big data analysis.In recent years, many frameworks/languages have implemented this pattern to achieve high performance in data mining applications, particularly for distributed memory architectures (e.g., clusters).Nevertheless, the industry of processors is now able to offer powerful processing on single machines (e.g., multi-core).Thus, these applications may address the parallelism in another architectural level.The target problems of this paper are code reuse and programming effort reduction since current solutions do not provide a single interface to deal with these two architectural levels.Therefore, we propose a unified domain-specific language in conjunction with transformation rules for code generation for Hadoop and Phoenix++.We selected these frameworks as state-of-the-art MapReduce implementations for distributed and shared memory architectures, respectively.Our solution achieves a programming effort reduction from 41.84% and up to 95.43% without significant performance losses (below the threshold of 3%) compared to Hadoop and Phoenix++. Daniel Couto Adornes, Dalvan Griebler, Cleverson Ledur, Luiz Gustavo Fernandes |
SEKE | 2 |
| 2015 | Coding Productivity in MapReduce Applications for Distributed and Shared Memory ArchitecturesabstractMapReduce was originally proposed as a suitable and efficient approach for analyzing and processing large amounts of data. Since then, many researches contributed with MapReduce implementations for distributed and shared memory architectures. Nevertheless, different architectural levels require different optimization strategies in order to achieve high-performance computing. Such strategies in turn have caused very different MapReduce programming interfaces among these researches. This paper presents some research notes on coding productivity when developing MapReduce applications for distributed and shared memory architectures. As a case study, we introduce our current research on a unified MapReduce domain-specific language with code generation for Hadoop and Phoenix++, which has achieved a coding productivity increase from 41.84% and up to 94.71% without significant performance losses (below 3%) compared to those frameworks. Daniel Couto Adornes, Dalvan Griebler, Cleverson Ledur, Luiz Gustavo Fernandes |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2014 | Evaluating the Impact of Transactional Characteristics on the Performance of Transactional Memory ApplicationsabstractTransactional Memory (TM) is reputed by many researchers to be a promising solution to ease parallel programming on multicore processors. This model provides the scalability of fine-grained locking while avoiding common issues of traditional mechanisms, such as deadlocks. During these almost twenty years of research, several TM systems and benchmarks have been proposed. However, TM is not yet widely adopted by the scientific community to develop parallel applications due to unanswered questions in the literature, such as "how to identify if a parallel application can exploit TM to achieve better performance?" or "what are the reasons of poor performances of some TM applications?". In this work, we contribute to answer those questions through a comparative evaluation of a set of TM applications on four different state- of-the-art TM systems. Moreover, we identify some of the most important TM characteristics that impact directly the performance of TM applications. Our results can be useful to identify opportunities for optimizations. Fernando Rui, Márcio Castro 0001, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 3 |
| 2014 | Performance and Usability Evaluation of a Pattern-Oriented Parallel Programming Interface for Multi-Core Architectures
Dalvan Griebler, Daniel Couto Adornes, Luiz Gustavo Fernandes |
SEKE | 1 |