EDBT 2026 Demo / reviewers in the wild / expert
Claudio Schepke
dblp:79/5559
· DBLP profile ↗
28ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-4118-8831ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Systems, architecture and hardware · 5 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Systematic Performance Study of Parallel Programming Models for Stencil-Based CFD Applications on Heterogeneous Architectures
Mariana Padilha, Claudio Schepke, Tailí Petry, José Rufino |
ICCSA (1) | 2 |
| 2026 | Performance and Resilience Evaluation of a Hybrid Edge-Cloud IoT System Across Connected, Offline, and Reconnection Scenarios
Luís Fernando Alves Da Silva, Pedro Ramires Da Silva Amalfi Costa, Claudio Schepke |
ICCSA (2) | 3 |
| 2025 | Contributions to Accelerating a Numerical Simulation of Free Flow Parallel to a Porous PlaneabstractFlow models over flat porous surfaces have applications in natural processes, such as material, food, chemical processing, or mountain mudflow simulations. The development of simplified analytical or numerical models can predict characteristics such as velocity, pressure, deviation length, and even temperature of such flows for geophysical and engineering purposes. In this context, there is considerable interest in theoretical and experimental models. Mathematical models to represent such phenomena for fluid mechanics have continuously been developed and implemented. Given this, we propose a mathematical and simulation model to describe a free-flowing flow parallel to a porous material and its transition zone. The objective of the application is to analyze the influence of the porous matrix on the flow under different matrix properties. We implement a Computational Fluid Dynamics scheme using the Finite Volume Method to simulate and calculate the numerical solutions for case studies. However, computational applications of this type demand high performance, requiring parallel execution techniques. Due to this, it is necessary to modify the sequential version of the code. So, we propose a methodology describing the steps required to adapt and improve the code. This approach decreases 5.3% the execution time of the sequential version of the code. Next, we adopt OpenMP for parallel versions and instantiate parallel code flows and executions on multi-core. We get a speedup of 10.4 by using 12 threads. The paper provides simulations that offer the correct understanding, modeling, and construction of abrupt transitions between free flow and porous media. The process presented here could expand to the simulations of other porous media problems. Furthermore, customized simulations require little processing time, thanks to parallel processing. Claudio Schepke, Roberta A. Spigolon, José Rufino, César Flaubiano da Cruz Cristaldo, Glener L. Pizzolato |
PDP | 1 |
| 2024 | A Real-Time Visualization Tool of Hardware Resources for Flutter Applications
Felipe Bedinotto Fava, Claudio Schepke |
ICCSA (2) | 2 |
| 2024 | Assessing the Performance of Docker in Docker Containers for Microservice-Based ArchitecturesabstractWe provide a comprehensive and updated assessment of Docker versus Docker in Docker (DinD), evaluating its impact on CPU, memory, disk, and network. Using different workloads, we evaluate DinD's performance across distinct hardware platforms and GNU/Linux distributions on cloud Infrastructure as a Service (laaS) platforms like Google Compute Engine (GCE) and traditional server-based environments. We developed an automated tools suite to achieve our goal. We execute four well-known benchmarks on Docker and its nested-container variant. Our findings indicate that nested-containers require up to 7 seconds for startup, while the Docker standard containers require less than 0.5 seconds for Debian and Alpine operating systems. Our results suggest that Docker containers based on Debian consistently outperform their Alpine counter-parts, showing lower CPU latency. A key distinction among these Docker images lies in the varying number of installed libraries (e.g., stretching from 13 to 119) across different Linux distributions for the same system (e.g., MySQL). Furthermore, the number of events and CPU latency indicates that the influence of DinD over Docker proves that it is insignificant for both operating systems. In terms of memory, running containers of Debian-based images consume 20% more size of memory than those based on Alpine. No significant differences are between nested-containers and Dockers for disk and network IO. It is worth emphasizing that some of the disparities, such as a bigger memory footprint, appear to be a direct result of the software stack in use, including different kernel versions. libraries. and other essential packages. Felipe Bedinotto Fava, Luiz Felipe Laviola Leite, Luís Fernando Alves Da Silva, Pedro Ramires Da Silva Amalfi Costa, Angelo Gaspar Diniz Nogueira, Amanda Fagundes Gobus Lopes, Claudio Schepke, Diego Kreutz, Rodrigo B. Mansilha |
PDP | 7 |
| 2024 | Performance and programmability of GrPPI for parallel stream processing on multi-coresabstractAbstract GrPPI library aims to simplify the burdening task of parallel programming. It provides a unified, abstract, and generic layer while promising minimal overhead on performance. Although it supports stream parallelism, GrPPI lacks an evaluation regarding representative performance metrics for this domain, such as throughput and latency. This work evaluates GrPPI focused on parallel stream processing. We compare the throughput and latency performance, memory usage, and programmability of GrPPI against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks and benchmarks with handwritten parallel code using the same backends supported by GrPPI. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is often competitive with handwritten parallel code, the infeasibility of fine-tuning GrPPI is a crucial drawback for emerging applications. Despite this, programmability experiments estimate that GrPPI can potentially reduce the development time of parallel applications by about three times. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
J. Supercomput. | 3 |
| 2023 | Modified Differential Evolution Algorithm Applied to Economic Load Dispatch Problems
Gabriella Lopes Andrade, Claudio Schepke, Natiele Lucca, João Plínio Juchem Neto |
ICCSA (1) | 2 |
| 2023 | A Latency, Throughput, and Programmability Perspective of GrPPI for Streaming on Multi-coresabstractSeveral solutions aim to simplify the burdening task of parallel programming. The GrPPI library is one of them. It allows users to implement parallel code for multiple backends through a unified, abstract, and generic layer while promising minimal overhead on performance. An outspread evaluation of GrPPI regarding stream parallelism with representative metrics for this domain, such as throughput and latency, was not yet done. In this work, we evaluate GrPPI focused on stream processing. We evaluate performance, memory usage, and programming effort and compare them against handwritten parallel code. For this, we use the benchmarking framework SPBench to build custom GrPPI benchmarks. The basis of the benchmarks is real applications, such as Lane Detection, Bzip2, Face Recognizer, and Ferret. Experiments show that while performance is competitive with handwritten code in some cases, in other cases, the infeasibility of fine-tuning GrPPI is a crucial drawback. Despite this, programmability experiments estimate that GrPPI has the potential to reduce by about three times the development time of parallel applications. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, André Sacilotto Santos, José Daniel García, Javier Fernández 0001, Luiz Gustavo Fernandes |
PDP | 3 |
| 2023 | Parallel Directives Evaluation in Porous Media Application: A Case StudyabstractHigh-performance computing provides the acceleration of scientific applications through the use of parallelism. Applications of this type usually demand a lot of computation time for a version with a single code execution stream. The adoption of different models of parallel programming enables the development of concurrent code. In this sense, this paper evaluates parallel interfaces and their programming models. Therefore, as a case study, we evaluate a porous media application that simulates grain drying using OpenMP (loop, sections, tasks, target, and teams approach) and OpenACC programming interfaces. The results show a reduction in processing time in all test cases. The total parallel simulation time for a multicore architecture using 16 physical cores was 5.61 times less using loops, 5.96 using targets, and 7.50 using teams. Task and section directives produce around 1.20 speedup due to the limitations of concurrent task executions of the application. The reduction using a single GPU was 7.54. We also contribute with some collected traces, identifying the parallel steps and synchronization time. Natiele Lucca, Claudio Schepke, Gabriel Tremarin |
PDP | 2 |
| 2023 | Micro-batch and data frequency for stream processing on multi-cores
Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
J. Supercomput. | 3 |
| 2023 | Parallel OpenMP and OpenACC porous media simulation
Hígor Uélinton Silva, Natiele Lucca, Claudio Schepke, Dalmo Paim de Oliveira, César Flaubiano da Cruz Cristaldo |
J. Supercomput. | 3 |
| 2022 | Evaluating Micro-batch and Data Frequency for Stream Processing Applications on Multi-coresabstractIn stream processing, data arrives constantly and is often unpredictable. It can show large fluctuations in arrival frequency, size, complexity, and other factors. These fluctuations can strongly impact application latency and throughput, which are critical factors in this domain. Therefore, there is a significant amount of research on self-adaptive techniques involving elasticity or micro-batching as a way to mitigate this impact. However, there is a lack of benchmarks and tools for helping researchers to investigate micro-batching and data stream frequency implications. In this paper, we extend a benchmarking framework to support dynamic micro-batching and data stream frequency management. We used it to create custom benchmarks and compare latency and throughput aspects from two different parallel libraries. We validate our solution through an extensive analysis of the impact of micro-batching and data stream frequency on stream processing applications using Intel TBB and FastFlow, which are two libraries that leverage stream parallelism on multi-core architectures. Our results demonstrated up to 33% throughput gain over latency using micro-batches. Additionally, while TBB ensures lower latency, FastFlow ensures higher throughput in the parallel applications for different data stream frequency configurations. Adriano Marques Garcia, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 3 |
| 2022 | Parallel OpenMP and OpenACC Mixing Layer SimulationabstractIt is estimated that up to 25% of the grain crop ends up being lost in the post-harvest. The correct drying of the beans is one of the measures to contain this loss. As the grain mass is a set of solid and empty spaces, its drying could be considered a problem of the coupled open-porous medium. In this paper, a mathematical and computer simulation model was proposed, which describes the convection in a free flow with a porous obstacle applied to the drying of the grain. A computational fluid dynamics scheme was implemented in FORTRAN using Finite Volume to simulate and compute the numerical solutions. The code is parallel implemented using OpenMP and OpenACC programming interfaces. As a result, there was a significant reduction in processing time in both cases. The total simulation time was eight times less for a multicore architecture (16 physical cores) and 17.3 times using a single GPU (Quadro M5000). Hígor Uélinton Silva, Claudio Schepke, Natiele Lucca, César Flaubiano da Cruz Cristaldo, Dalmo Paim de Oliveira |
PDP | 2 |
| 2021 | Introducing a Stream Processing Framework for Assessing Parallel Programming InterfacesabstractStream Processing applications are spread across different sectors of industry and people's daily lives. The increasing data we produce, such as audio, video, image, and text are demanding quickly and efficiently computation. It can be done through Stream Parallelism, which is still a challenging task and most reserved for experts. We introduce a Stream Processing framework for assessing Parallel Programming Interfaces (PPIs). Our framework targets multi-core architectures and C++ stream processing applications, providing an API that abstracts the details of the stream operators of these applications. Therefore, users can easily identify all the basic operators and implement parallelism through different PPIs. In this paper, we present the proposed framework, implement three applications using its API, and show how it works, by using it to parallelize and evaluate the applications with the PPIs Intel TBB, FastFlow, and SPar. The performance results were consistent with the literature. Adriano Marques Garcia, Dalvan Griebler, Luiz Gustavo Fernandes, Claudio Schepke |
PDP | 4 |
| 2020 | The Impact of CPU Frequency Scaling on Power Consumption of Computing Infrastructures
Adriano Marques Garcia, Matheus S. Serpa, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux |
ICCSA (6) | 4 |
| 2020 | A New Library of Bio-Inspired Algorithms
Natiele Lucca, Claudio Schepke |
ICCSA (1) | 2 |
| 2020 | Evaluation of SIMD Instructions on Bio-Inspired AlgorithmsabstractBio-inspired algorithms are based on the collective behavior of interacting organisms and are used to solve or to approach efficient solutions for large optimization problems. This article evaluates a new parallel library of a family of bio-inspired algorithms. The library offers the implementation of some basic algorithms, being easily extensible through interfaces, and explores the parallelism using SIMD-type instructions. Evaluations of performances are presented using seven test functions applied to each of the implemented algorithms. The tests also allowed to show that parallel implementations offer higher performances in all cases, reaching up to 20 times for some functions. Natiele Lucca, Claudio Schepke |
ISPDC | 2 |
| 2020 | Acceleration of Radiofrequency Ablation Process for Liver Cancer Using GPUabstractComputational simulation is a technique used in several research areas. In medicine, the Radiofrequency Ablation Finite Element Method (RAFEM) application was developed to simulate the RadioFrequency Ablation (RFA) process, which is a medical procedure to treat hepatic cancer. This application presents a high computational time to perform a simulation, taking up to 20 hours for simulation while the RFA procedure itself lasts from four to six minutes. Some efforts have already been carried out to obtain better performance for the application. However, none of them considered the use of GPUs, which is an architecture that can be used to accelerate finite element method applications. This work aims to propose a parallelization approach to explore ways to obtain better performance for the application through the use of GPUs to reduce the execution time under a few minutes. We adopted an iterative development process to coordinate and to generate parallel versions of the code. Through this process allied to the use of GPU-accelerated code libraries, it was possible to create different versions of the application, adding improvements to each one. The results obtained with the best version showed a reduction of the computation time by up to 18 times. Therefore, this improvement makes the use of RAFEM application more feasible in health treatment by reducing the waiting time for the results. Claudio Schepke, Marcelo Miletto |
PDP | 1 |
| 2020 | PAMPAR: A new parallel benchmark for performance and energy consumption evaluationabstractSummary This paper presents PAMPAR, a new benchmark to evaluate the performance and energy consumption of different Parallel Programming Interfaces (PPIs). The benchmark is composed of 11 algorithms implemented in PThreads, OpenMP, MPI‐1, and MPI‐2 (spawn) PPIs. Previous studies have used some of these pseudo‐applications to perform this type of evaluation in different architectures since there is no benchmark that offers this variety of PPIs and communication models. In this work, we measure the energy and performance of each pseudo‐application in a single architecture, varying the number of threads/processes. We also organize the pseudo‐applications according to their memory accesses, floating‐point operations, and branches. The goal is to show that this set of pseudo‐applications has enough features to build a parallel benchmark. The results show that there is no single best case that provides both better performance and low energy consumption in the presented scenarios. Moreover, the pseudo‐applications usage of the system resources are different enough to represent different scenarios and be efficient as a benchmark. Adriano Marques Garcia, Claudio Schepke, Alessandro Girardi |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | A Dynamic Task-Based D3Q19 Lattice-Boltzmann Method for Heterogeneous ArchitecturesabstractNowadays computing platforms expose a significant number of heterogeneous processing units such as multicore processors and accelerators. The task-based programming model has been a de facto standard model for such architectures since its model simplifies programming by unfolding parallelism at runtime based on data-flow dependencies between tasks. Many studies have proposed parallel strategies over heterogeneous platforms with accelerators. However, to the best of our knowledge, no dynamic task-based strategy of the Lattice-Boltzmann Method (LBM) has been proposed to exploit CPU+GPU computing nodes. In this paper, we present a dynamic task-based D3Q19 LBM implementation using three runtime systems for heterogeneous architectures: OmpSs, StarPU, and XKaapi. We detail our implementations and compare performance over two heterogeneous platforms. Experimental results demonstrate that our task-based approach attained up to 8.8 of speedup over an OpenMP parallel loop version. João V. F. Lima, Gabriel Freytag, Vinícius Garcia Pinto, Claudio Schepke, Philippe Olivier Alexandre Navaux |
PDP | 4 |
| 2018 | Performance of Data Mining, Media, and Financial Applications under Private Cloud ConditionsabstractThis paper contributes to a performance analysis of real-world workloads under private cloud conditions. We selected six benchmarks from PARSEC related to three mainstream application domains (financial, data mining, and media processing). Our goal was to evaluate these application domains in different cloud instances and deployment environments, concerning container or kernel-based instances and using dedicated or shared machine resources. Experiments have shown that performance varies according to the application characteristics, virtualization technology, and cloud environment. Results highlighted that financial, data mining, and media processing applications running in the LXC instances tend to outperform KVM when there is a dedicated machine resource environment. However, when two instances are sharing the same machine resources, these applications tend to achieve better performance in the KVM instances. Finally, financial applications achieved better performance in the cloud than media and data mining. Dalvan Griebler, Adriano Vogel, Carlos A. F. Maron, Anderson M. Maliszewski, Claudio Schepke, Luiz Gustavo Fernandes |
ISCC | 5 |
| 2018 | Performance and Energy Consumption Analysis of Coprocessors Using Different Programming ModelsabstractThe optimization of the relation between performance and energy consumption is a strong requirement mainly in high performance environments. The top 10 Green500 supercomputers use accelerators/coprocessors as primary approach to increase the performance while reducing energy consumption. This paper presents a study on the main factors that impact this relationship, evaluating and comparing Intel programming models on an Intel Xeon Phi coprocessor architecture. The methodology applied in this work consists of evaluating performance and energy consumption on execution scenarios using Linpack and HPL 2.1 benchmarks. These scenarios consider various environment parameters and execution on the Intel host, offload and native programming models. Experimental results indicate that the host and offload models are more efficient in the performance per energy consumption relationship with shared memory and distributed memory, whereas the native model demonstrated better efficiency in energy consumption. Robson Goncalves, Alessandro Girardi, Claudio Schepke |
PDP | 3 |
| 2017 | An Intra-Cloud Networking Performance Evaluation on CloudStack EnvironmentabstractInfrastructure-as-a-Service (IaaS) is a cloud on-demand commodity built on top of virtualization technologies and managed by IaaS tools. In this scenario, performance is a relevant matter because a set of aspects may impact and increase the system overhead. Specific on the network, the use of virtualized capabilities may cause performance degradation (eg.,latency, throughput). The goal of this paper is to contribute to networking performance evaluation, providing new insights for private IaaS clouds. To achieve our goal, we deploy CloudStack environments and conduct experiments with different configurations and techniques. The research findings demonstrate that KVM-based cloud instances have small network performance degradation regarding throughput (about 0.2% for coarse-grained and 6.8% for fine-grained messages) while container-based instances have even better results. On the other hand, the KVM instances present worst latency (about 12.4% on coarse-grained and two times more on fine-grained messages w. r. t. native environment) and better in container-based instances, where the performance results are close to the native environment. Furthermore, we demonstrate a performance optimization of applications running on KVM. Adriano Vogel, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 3 |
| 2016 | Private IaaS Clouds: A Comparative Analysis of OpenNebula, CloudStack and OpenStackabstractDespite the evolution of cloud computing in recent years, the performance and comprehensive understanding of the available private cloud tools are still under research. This paper contributes to an analysis of the Infrastructure as a Service (IaaS) domain by mapping new insights and discussing the challenges for improving cloud services. The goal is to make a comparative analysis of OpenNebula, OpenStack and CloudStack tools, evaluating their differences on support for flexibility and resiliency. Also, we aim at evaluating these three cloud tools when they are deployed using a mutual hypervisor (KVM) for discovering new empirical insights. Our research results demonstrated that OpenStack is the most resilient and CloudStack is the most flexible for deploying an IaaS private cloud. Moreover, the performance experiments indicated some contrasts among the private IaaS cloud instances when running intensive workloads and scientific applications. Adriano Vogel, Dalvan Griebler, Carlos A. F. Maron, Claudio Schepke, Luiz Gustavo Fernandes |
PDP | 4 |
| 2014 | Improving OLAM with Cloud Elasticity
Guilherme Galante, Luis C. E. Bona, Claudio Schepke |
ICCSA (6) | 3 |
| 2011 | Improving Performance on Atmospheric Models through a Hybrid OpenMP/MPI ImplementationabstractThis work shows how a Hybrid MPI/OpenMP implementation can improve the performance of the Ocean-Land-Atmosphere Model (OLAM) on a multi-core cluster environment, which is a typical HPC many small files workload application. Previous experiments have shown that the scalability of this application on clusters is limited by the performance of the output operations. We show that the Hybrid MPI/OpenMP version of OLAM decreases the number of output files, resulting in better performance for I/O operations. We also observe that the MPI version of OLAM performs better for unbalanced workloads and that further parallel optimizations should be included on the hybrid version in order to improve the parallel execution time of OLAM. Carla Osthoff, Pablo Javier Grunmann, Francieli Zanon Boito, Rodrigo Kassick, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Claudio Schepke, Jairo Panetta, Nicolas Maillard, Pedro Leite da Silva Dias, Robert L. Walko |
ISPA | 7 |
| 2011 | Why Online Dynamic Mesh Refinement is Better for Parallel Climatological ModelsabstractForecast precisions of climatological models are limited by computing power and time available for the executions. As more and faster processors are used in the computation, the resolution of the mesh adopted to represent the Earth's atmosphere can be increased, and consequently the numerical forecast is more accurate and shows local phenomena. However, a finer mesh resolution, able to include local phenomena in a global atmosphere integration, is still not possible. To overcome this situation, different mesh refinement levels can be used at the same time for different areas. In this context, this paper evaluates how mesh refinement at run time can improve performance for climatological models. In order to contribute with this analysis, an online dynamic mesh refinement was developed. It increases mesh resolution in parts of a parallel distributed model, when special atmosphere conditions are registered during the execution. The results show that the parallel execution of this improvement provides better resolution for the meshes, without a significant increase of execution time. Claudio Schepke, Nicolas Maillard, Jörg Schneider 0001, Hans-Ulrich Heiß |
SBAC-PAD | 1 |
| 2007 | Performance Improvement of the Parallel Lattice Boltzmann Method Through Blocked Data DistributionsabstractThis paper presents a blocked parallel implementation of a three-diagonal version of the Lattice Boltzmann Method. This method is a numerical model used to represent and to simulate fluid flows through mesoscopic approaches. Parallel implementations are often adopted to attend the demand of an expressive memory amount and processing power of the method. However, most implementations use simple data distribution strategies to parallelize the operations on the regular fluid data set. Fluid flows simulations crossing a cavity have been used as case study to evaluate our implementation. The presented results with blocked implementations achieve a performance 31% higher than non-blocked versions for some data distributions. Thus, this work shows that blocked implementations can be efficiently used to reduce the parallel execution time of the method. Claudio Schepke, Nicolas Maillard |
SBAC-PAD | 1 |