EDBT 2026 Demo / reviewers in the wild / expert
Matheus S. Serpa
dblp:199/1097 · also Matheus da Silva Serpa
· DBLP profile ↗
16ranked-venue papers
4as first author
9since 2021 · last 2023
0000-0001-5178-1036ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Mitigating execution unit contention in parallel applications using instruction-aware mappingabstractSummary Parallel applications running on simultaneous multithreading (SMT) processors naturally compete for execution units when their threads are mapped to the same core. This issue is further aggravated when such threads execute similar instructions that stress the same execution unit type, making their execution to behave very similarly as if the threads were running sequentially. This, in turn, will lead to performance degradation and underutilization of hardware resources. This work proposes a completely transparent framework (no modifications to the source code are necessary) that automatically maps threads of multiple parallel applications on SMT processors. The framework focuses on improving performance by mitigating the contention on execution units, considering each thread's instruction types, which are detected at runtime by our framework. Results show performance gains of 21% (geometric mean), compared to the native scheduler of the operating system. Matheus S. Serpa, Eduardo Henrique Molina da Cruz, Matthias Diener, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Smart resource allocation of concurrent execution of parallel applicationsabstractAbstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy. Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 2 |
| 2021 | Lightweight Deep Learning Applications on AVX-512abstractMachine Learning and Deep Learning applications are of paramount importance these days. Different areas of academia and industry use daily workloads based on these applications. Several aspects are relevant regarding their applicability, such as the complexity and accuracy of the models and their performance and energy efficiency. Currently, there is a trend to usually favor the use of GPUs to train and execute Deep Learning models, intensified by specialized hardware. However, this article demonstrates that using a CPU with AVX-512 instructions can achieve comparable performance to current GPUs and, depending on the workload, suppress it by ≈ 1.8x. Andre Ramos Carneiro, Matheus S. Serpa, Philippe Olivier Alexandre Navaux |
ISCC | 2 |
| 2021 | Combining Thread Throttling and Mapping to Optimize the EDP of Parallel ApplicationsabstractThread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1. Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 3 |
| 2021 | Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo BoostingabstractWith the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability. Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
PDP | 3 |
| 2021 | Investigating memory prefetcher performance over parallel applications: From real to simulatedabstractAbstract Memory prefetcher algorithms are widely used in processors to mitigate the performance gap between the processors and the memory subsystem. The complexities behind the architectures and prefetcher algorithms, however, not only hinder the development of accurate architecture simulators, but also hinder understanding the prefetcher's contribution to performance, on both a real hardware and in a simulated environment. In this paper, we contribute to shed light on the memory prefetcher's role in the performance of parallel High‐Performance Computing applications, considering the prefetcher algorithms offered by both the real hardware and the simulators. We performed a careful experimental investigation, executing the NAS parallel benchmark (NPB) on a real Skylake machine, and as well in a simulated environment with the ZSim and Sniper simulators, taking into account the prefetcher algorithms offered by both Skylake and the simulators. Our experimental results show that: (i) prefetching from the L3 to L2 cache presents better performance gains, (ii) the memory contention in the parallel execution constrains the prefetcher's effect, (iii) Skylake's parallel memory contention is poorly simulated by ZSim and Sniper, and (iv) Skylake's noninclusive L3 cache hinders the accurate simulation of NPB with the Sniper's prefetchers. Valéria Soldera Girelli, Francis B. Moreira 0001, Matheus S. Serpa, Danilo Carastan-Santos, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Energy efficiency and portability of oil and gas simulations on multicore and graphics processing unit architecturesabstractSummary Reverse time migration (RTM) simulation is the basis of the seismic imaging tools used by the oil and gas industry. Developers have been porting their simulations to the new high‐performance computing architectures, providing faster and more accurate results at each new generation. However, several challenges arrive when trying to achieve high performance on these new architectures. The first one is to choose the architecture that best fits the kind of simulation. After that, researchers should choose the API used to implement the simulation code. These two decisions are strongly related to the effort, performance, and energy efficiency of the simulations. In this article, we propose three optimizations for an oil and gas application, which reduce the floating‐point operations by changing the equation derivatives. We evaluate these optimizations in different multicore and GPU architectures, investigating the impact of different APIs on the performance, energy efficiency, and portability of the code. Our experimental results show that the dedicated CUDA implementation running on the NVIDIA Volta architecture has the best performance and energy efficiency for RTM on GPUs, while the OpenMP version is the best for Intel Broadwell in the multicore. Also, the OpenACC version, which has a lower programming effort and executes on both architectures, has an up to 20% better performance and energy efficiency than the nonportable ones. Matheus S. Serpa, Pablo J. Pavan, Eduardo Henrique Molina da Cruz, Rodrigo L. Machado, Jairo Panetta, Antônio Azambuja, Alexandre Carissimi, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Collaborative execution of fluid flow simulation using non-uniform decomposition on heterogeneous architectures
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux |
J. Parallel Distributed Comput. | 2 |
| 2020 | The Impact of CPU Frequency Scaling on Power Consumption of Computing Infrastructures
Adriano Marques Garcia, Matheus S. Serpa, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux |
ICCSA (6) | 2 |
| 2020 | Task-based parallel strategies for computational fluid dynamic application in heterogeneous CPU/GPU resourcesabstractSummary Parallel applications executing in contemporary heterogeneous clusters are complex to code and optimize. The task‐based programming model is an alternative to handle the coding complexity. This model consists of splitting the problem domain into tasks with dependencies through a directed acyclic graph, and submit the set of tasks to a runtime scheduler that maps each task dynamically to resources. We consider that computational fluid dynamics applications are typical in scientific computing but not enough exploited by designs that employ the task‐based programming model. This article presents task‐based parallel strategies for a simple CFD application that targets heterogeneous multi‐CPU/multi‐GPU computing resources. We design, develop, evaluate, and compare the performance of three parallel strategies (naive, ghost‐cells, and arrow) of a task‐based heterogeneous (CPU and GPU) application that simulates the flow of an incompressible Newtonian fluid with constant viscosity. All implementations rely on the StarPU runtime, and we use the StarVZ toolkit to conduct comprehensive performance analysis. Results indicate that the ghost cell strategy provides the best speedup (77×) considering the simulation time when the GPU resources still have available memory. However, the arrow strategy achieves better results when the simulation data increases. Lucas Leandro Nesi, Matheus S. Serpa, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Memory Performance and Bottlenecks in Multicore and GPU ArchitecturesabstractNowadays, there are several different architectures available not only for the industry, but also for normal consumers. Traditional multicore processors, GPUs, accelerators such as the Sunway SW26010, or even energy efficiency-driven processors such as the ARM family, present very different architectural characteristics. This wide range of characteristics presents a challenge for the developers of applications. Developers must deal with different instruction sets, memory hierarchies, or even different programming paradigms when programming for these architectures. Therefore, the same application can perform well when executing on one architecture, but poorly on another architecture. To optimize an application, it is important to have a deep understanding of how it behaves on different architectures. The related work in this area mostly focuses on a limited analysis encompassing execution time and energy. In this paper, we perform a detailed investigation on the impact of the memory subsystem of different architectures, which is one of the most important aspects to be considered. For this study, we performed experiments in the Broadwell CPU and Pascal GPU, using applications from the Rodinia benchmark suite. In this way, we were able to understand why an application performs well on one architecture and poorly on others. Matheus S. Serpa, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, Eduardo Henrique Molina da Cruz, Matthias Diener, Dalvan Griebler, Luiz Gustavo Fernandes |
PDP | 1 |
| 2019 | Non-uniform Partitioning for Collaborative Execution on Heterogeneous ArchitecturesabstractSince the demand for computing power increases, new architectures arise to obtain better performance. An important class of integrated devices is heterogeneous architectures, which join different specialized hardware into a single chip, composing a System on Chip - SoC. Within this context, effectively splitting tasks between the different architectures is primal to obtain efficiency and performance. In this work, we evaluate two heterogeneous architectures: one composed of a general-purpose CPU and a graphics processing unit (GPU) integrated into a single chip (AMD Kaveri SoC), and another composed by a general-purpose CPU and a Field Programmable Gate Array (FPGA) integrated into a single chip (Intel Arria 10 SoC). We investigate how data partitioning affects the performance of each device in a collaborative execution through the decomposition of the data domain. As a case study, we apply the technique in the well-known Lattice Boltzmann Method (LBM), analyzing the performance of five kernels in both architectures. Our experimental results show that non-uniform partitioning improves LBM kernels performance by up to 11.40% and 15.15% on AMD Kaveri and Intel Arria 10, respectively. Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |
| 2019 | An Unsupervised Learning Approach for I/O Behavior CharacterizationabstractI/O operations are the bottleneck of several applications due to the difference between processing and data access speeds. Hence, understanding the I/O behavior is vital to find problems and propose solutions. Thus, identifying and characterizing the I/O access pattern is important, since it reflects directly on applications' performance. With this premise, we propose an I/O characterization approach that uses unsupervised learning to cluster jobs with similar I/O behavior, using information from high-level aggregated traces. As a case study, we apply our approach on four months of activity - a total of 28, 938 jobs - from the Intrepid supercomputer located at Argonne Laboratory. Our experimental results show that nine access patterns represent the I/O behavior in 73% of the clusters. From these nine patterns, we learn some aspects about the I/O such as the most accesses patterns are made using POSIX and small requests, also, the most patterns are accessing unique files. Lastly, analyzing the I/O workload over four months, we can notice that it is composed by several applications that spend a short time on I/O activity, but when compared to the others, the total I/O time represents a greater portion of the overall system. Pablo J. Pavan, Jean Luca Bez, Matheus S. Serpa, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 3 |
| 2018 | Improving Communication and Load Balancing with Thread Mapping in Manycore SystemsabstractCommunication and load balancing have a significant impact on the performance of parallel applications and have been the subject of extensive research in multicore architectures. Thread mapping has been one of the solutions adopted in multicore architectures to address both communication and load balancing. However, the impact of such issues on more recently introduced manycore architectures is still unknown. Most related work on manycore architectures focus on execution time and idleness information for scheduling decisions. In this paper, we improve the state of the art by performing a very detailed analysis of the impact of thread mapping on communication and load balancing in two manycore systems from Intel, namely Knights Corner and Knights Landing. We observed that the widely used metric of CPU time provides very inaccurate information for load balancing. We also evaluated the usage of thread mapping based on the communication and load information of the applications to improve the performance of manycore systems. Eduardo Henrique Molina da Cruz, Matthias Diener, Matheus S. Serpa, Philippe Olivier Alexandre Navaux, Laércio Lima Pilla, Israel Koren |
PDP | 3 |
| 2018 | Optimizing Machine Learning Algorithms on Multi-Core and Many-Core Architectures Using Thread and Data MappingabstractDriven by the development of new technologies such as personal assistants or autonomous cars, machine learning has rapidly become one of the most active fields in computer science. The algorithms at the core of machine learning are notoriously demanding in terms of resources. It is therefore of paramount importance to optimize their operation on modern processors. Several approaches have been proposed to accelerate machine learning on GPUs and massively parallel computers, as well as dedicated ASICs. In this paper, we focus on Intel's multi-core Xeon and many-core accelerator Xeon Phi Knights Landing, which can host several hundreds of threads on the same CPU. In such architectures, thread and data mapping are keys for performance. We study the impact of mapping strategies, revealing that, with smart mapping policies, one can indeed significantly speed up machine learning applications on many-core architectures. Execution time was reduced by up to 25.2% and 18.5% on Intel Xeon and Xeon Phi KNL, respectively. Matheus S. Serpa, Arthur M. Krause, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux, Marcelo Pasin, Pascal Felber |
PDP | 1 |