Carla Osthoff

dblp:76/6443 · also Carla Barros Osthoff Ferreira de Barros, Carla Osthoff Ferreira de Barros · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-4694-7182ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Assuming the best: Towards a reliable protocol for resource usage prediction for high-performance computing based on machine learning
abstract
In High-Performance Computing (HPC) systems, multiple processes simultaneously consume resources such as CPU time, memory, and electrical power, among others. Accurately predicting the resource consumption of a process based on its execution parameters enables more efficient resource allocation, ultimately improving the overall performance of the HPC system. While many studies have explored this topic, fewer explicitly examine the underlying assumptions of their approaches. This work contributes to filling that gap by proposing, experimenting with, and discussing a protocol to approach this problem, covering from the collection of processes footprint data to the experimental evaluation of Machine Learning models based on such data. The reported results of the assessment of this protocol in a case study of the RAxML bioinformatics application on a real supercomputer highlight not only its effectiveness ( R 2 values greater than 0.9 were achieved in most tests) but also the reasonableness of the assumptions considered.
Alexandre H. L. Porto, Micaella Coelho, Hiago Rocha, Carla Osthoff, Kary A. C. S. Ocaña, Douglas de O. Cardoso
Future Gener. Comput. Syst.4
2025 A Deep Look into the Temporal I/O Behavior of HPC Applications
abstract
The increasing gap between compute and I/O speeds in high-performance computing (HPC) systems imposes the need for techniques to improve applications' I/O performance. Such techniques must rely on assumptions about I/O behavior in order to efficiently allocate I/O resources such as burst buffers, to schedule accesses to the shared parallel file system or to delay certain applications at the batch scheduler level to prevent contention, for instance. In this paper, we verify these common assumptions about I/O behavior, specifically about temporal behavior, using over 440,000 traces from real HPC systems. By combining traces from diverse systems, we characterize the behaviors observed in real HPC workloads. Among other findings, we show that I/O activity tends to last for a few seconds, and that periodic jobs are the minority, but responsible for a large portion of the I/O time. Furthermore, we make projections for the expected improvement yielded by popular approaches for I/O performance improvement. Our work provides valuable insights to everyone working to alleviate the I/O bottleneck in HPC.
Francieli Zanon Boito, Luan Teylo, Mihail Popov, Théo Jolivel, Francois Tessier, Jakob Lüttgau, Julien Monniot, Ahmad Tarraf, Andre Ramos Carneiro, Carla Osthoff
IPDPS10
2023 An Evaluation of Direct and Indirect Memory Accesses in Fluid Flow Simulator
Stiw Harrison Herrera Taipe, Thiago Teixeira, Weber Ribeiro, Andre Ramos Carneiro, Marcio Rentes Borges, Carla Osthoff, Frederico Luís Cabral, Sanderson L. Gonzaga de Oliveira
ICCSA (1)6
2023 Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos Dumont
abstract
High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.3
2022 Reducing Cache Miss Rate Using Thread Oversubscription to Accelerate an MPI-OpenMP-Based 2-D Hopmoc Method
Frederico Luís Cabral, Carla Osthoff, Sanderson L. Gonzaga de Oliveira
ICCSA (1)2
2021 HPC Data Storage at a Glance: The Santos Dumont Experience
abstract
High-Performance Computing (HPC) platforms are used to solve the most diverse scientific problems in research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific software, which have different requirements. These requirements include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle a mixed workload scenario when storing data from the applications. Knowledge of the application set and its performance running in a supercomputer is needed to understand the storage system's usage, pinpoint possible bottlenecks, and guide optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, where we were able to identify inefficient usage and problematic factors of performance.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2020 The Influence of Reordering Algorithms on the Convergence of a Preconditioned Restarted GMRES Method
Sanderson L. Gonzaga de Oliveira, C. Carvalho, Carla Osthoff
ICCSA (1)3
2020 A Convergence Analysis of a Multistep Method Applied to an Advection-Diffusion Equation in 1-D
Diogo T. Robaina, Sanderson L. Gonzaga de Oliveira, Mauricio Kischinhevsky, Carla Osthoff, Alexandre da Costa Sena
ICCSA (1)4
2020 An evaluation of MPI and OpenMP paradigms in finite-difference explicit methods for PDEs on shared-memory multi- and manycore systems
abstract
Summary This paper focuses on parallel implementations of three two‐dimensional explicit numerical methods on Intel® Xeon® Scalable Processor and the coprocessor Knights Landing. In this study, the performance of a hybrid parallel programming with message passing interface (MPI) and Open Multi‐Processing (OpenMP) and a pure MPI implementation used with two thread binding policies is compared with an improved OpenMP‐based implementation in three explicit finite‐difference methods for solving partial differential equations on shared‐memory multicore and manycore systems. Specifically, the improved OpenMP‐based version is a strategy that synchronizes adjacent threads and eliminates the implicit barriers of a naïve OpenMP‐based implementation. The experiments show that the most suitable approach depends on several characteristics related to the nonuniform memory access (NUMA) effect and load balancing, such as the size of the MPI domain and the number of synchronization points used in the parallel implementation. In algorithms that use four and five synchronization points, hybrid MPI/OpenMP approaches yielded better speedups than the other versions did in runs performed on both systems. The pure MPI‐based strategy, however, achieved better results than the other proposed approaches did in the method that employs only one synchronization point.
Frederico Luís Cabral, Sanderson L. Gonzaga de Oliveira, Carla Osthoff, Gabriel P. Costa, Diego N. Brandão, Mauricio Kischinhevsky
Concurr. Comput. Pract. Exp.3
2020 BioinfoPortal: A scientific gateway for integrating bioinformatics applications on the Brazilian national high-performance computing network
Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Fábio Porto 0001, Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Ana Tereza Ribeiro de Vasconcelos
Future Gener. Comput. Syst.3
2019 Towards a Science Gateway for Bioinformatics: Experiences in the Brazilian System of High Performance Computing
abstract
Science gateways bring out the possibility of reproducible science as they are integrated into reusable techniques, data and workflow management systems, security mechanisms, and high performance computing (HPC). We introduce BioinfoPortal, a science gateway that integrates a suite of different bioinformatics applications using HPC and data management resources provided by the Brazilian National HPC System (SINAPAD). BioinfoPortal follows the Software as a Service (SaaS) model and the web server is freely available for academic use. The goal of this paper is to describe the science gateway and its usage, addressing challenges of designing a multiuser computational platform for parallel/distributed executions of large-scale bioinformatics applications using the Brazilian HPC resources. We also present a study of performance and scalability of some bioinformatics applications executed in the HPC environments and perform machine learning analyses for predicting features for the HPC allocation/usage that could better perform the bioinformatics applications via BioinfoPortal.
Kary A. C. S. Ocaña, Marcelo Galheigo, Carla Osthoff, Luiz M. R. Gadelha Jr., Antônio Tadeu A. Gomes, Daniel de Oliveira 0001, Fábio Porto 0001, Ana Tereza Ribeiro de Vasconcelos
CCGRID3
2019 A Variant of the George-Liu Algorithm
Sanderson L. Gonzaga de Oliveira, Alexandre Augusto Alberto Moreira de Abreu, Carla Osthoff, L. N. Henderson Guedes de Oliveira
ICCSA (1)3
2019 An Experimental Analysis of Heuristics for Profile Reduction
Sanderson L. Gonzaga de Oliveira, Carla Osthoff, L. N. Henderson Guedes de Oliveira
ICCSA (1)2
2018 A Total Variation Diminishing Hopmoc Scheme for Numerical Time Integration of Evolutionary Differential Equations
Diego N. Brandão, Sanderson L. Gonzaga de Oliveira, Mauricio Kischinhevsky, Carla Osthoff, Frederico Luís Cabral
ICCSA (1)4
2018 Collective I/O Performance on the Santos Dumont Supercomputer
abstract
The historical gap between processing and data access speeds causes many applications to spend a large portion of their execution on I/O operations. From the point of view of a large-scale, expensive, supercomputer, it is important to ensure applications achieve the best I/O performance to promote an efficient usage of the machine. In this paper, we evaluate the I/O infrastructure of the Santos Dumont supercomputer, the largest one from Latin America. More specifically, we investigate the performance of collective I/O operations. By conducting an analysis of a scientific application that uses the machine, we identify large performance differences between the available MPI implementations. We then further study the observed phenomenon using the BT-IO and IOR benchmarks, in addition to a custom microbenchmark. We conclude that the customized MPI implementation by Bull (used by more than 20% of the jobs) presents the worst performance for small collective write operations. Our results are being used to help the Santos Dumont users to achieve the best performance for their applications. Additionally, by investigating the observed phenomenon, we provide information to help improve future MPI-IO collective write implementations.
Andre Ramos Carneiro, Jean Luca Bez, Francieli Zanon Boito, Bruno Alves Fagundes, Carla Osthoff, Philippe Olivier Alexandre Navaux
PDP5
2011 Improving Performance on Atmospheric Models through a Hybrid OpenMP/MPI Implementation
abstract
This work shows how a Hybrid MPI/OpenMP implementation can improve the performance of the Ocean-Land-Atmosphere Model (OLAM) on a multi-core cluster environment, which is a typical HPC many small files workload application. Previous experiments have shown that the scalability of this application on clusters is limited by the performance of the output operations. We show that the Hybrid MPI/OpenMP version of OLAM decreases the number of output files, resulting in better performance for I/O operations. We also observe that the MPI version of OLAM performs better for unbalanced workloads and that further parallel optimizations should be included on the hybrid version in order to improve the parallel execution time of OLAM.
Carla Osthoff, Pablo Javier Grunmann, Francieli Zanon Boito, Rodrigo Kassick, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Claudio Schepke, Jairo Panetta, Nicolas Maillard, Pedro Leite da Silva Dias, Robert L. Walko
ISPA1
2005 Scientific Models Management in Computational Grids
Halisson Brito, Julia Celia M. Strauch, Jano Moreira de Souza, Carla Osthoff
SSDBM4