EDBT 2026 Demo / reviewers in the wild / expert
Francieli Zanon Boito
dblp:99/9377 · also Francieli Boito
· DBLP profile ↗
21ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-1139-0724ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOTO: Transparent I/O Tuning for HPC ApplicationsabstractHigh-performance computing applications rely on parallel file systems, where I/O performance is strongly affected by configuration parameters such as stripe count. However, the ideal stripe count is highly application- and system-dependent, making it difficult to predict and rarely tuned in practice. As a result, substantial I/O performance potential remains unexplored. We present TOTO, a transparent tool that improves I/O performance without requiring application modifications. TOTO intercepts POSIX calls, characterizes application behavior, and uses a machine learning model to select an appropriate stripe count, even for already opened files. We also introduce an allocation algorithm that balances performance and resource occupation, and describe a methodology for training the model once per system using limited data. Our results show that TOTO can improve I/O performance by up to 4.6 × compared to using a default stripe count, while imposing an overhead of at most \(8\%\). Moreover, compared to the state of the art, TOTO can optimize more applications with a lower resource occupation, which is expected to decrease contention in the I/O infrastructure. Francieli Zanon Boito, Luan Teylo, Mihail Popov, Laora Aimi, Alexis Bandet, Laércio Lima Pilla, Guillaume Pallez |
ICS | 1 |
| 2026 | On the Impact of Interference from Concurrent Jobs on Checkpointing PerformanceabstractI/O has been identified as one of the main bottlenecks in HPC. Among the most I/O-intensive operations is checkpointing, which is necessary to save the state of an application and allow it to be restarted at an advanced stage of computation. However, near the parallel file system, concurrency prevents checkpoint phases from reaching the best I/O performance. In this paper, we study I/O interference in this specific context: we look at performance of a checkpoint phase when faced with different interference patterns, exploring aspects such as scale, number of processes, operation, number of files, etc. Through an extensive experimentation, in two systems, we show the impact of these aspects on checkpoint. Moreover, we show that some configurations — e.g., an application that does random accesses — lead to degraded system I/O performance. This paper provides an important background for any effort into mitigating I/O interference and into improving checkpointing performance. Méline Trochon, Jean-Thomas Acquaviva, Francieli Zanon Boito, Brice Goglin, Francois Tessier, Luan Teylo |
SSDBM | 3 |
| 2025 | A Deep Look into the Temporal I/O Behavior of HPC ApplicationsabstractThe increasing gap between compute and I/O speeds in high-performance computing (HPC) systems imposes the need for techniques to improve applications' I/O performance. Such techniques must rely on assumptions about I/O behavior in order to efficiently allocate I/O resources such as burst buffers, to schedule accesses to the shared parallel file system or to delay certain applications at the batch scheduler level to prevent contention, for instance. In this paper, we verify these common assumptions about I/O behavior, specifically about temporal behavior, using over 440,000 traces from real HPC systems. By combining traces from diverse systems, we characterize the behaviors observed in real HPC workloads. Among other findings, we show that I/O activity tends to last for a few seconds, and that periodic jobs are the minority, but responsible for a large portion of the I/O time. Furthermore, we make projections for the expected improvement yielded by popular approaches for I/O performance improvement. Our work provides valuable insights to everyone working to alleviate the I/O bottleneck in HPC. Francieli Zanon Boito, Luan Teylo, Mihail Popov, Théo Jolivel, Francois Tessier, Jakob Lüttgau, Julien Monniot, Ahmad Tarraf, Andre Ramos Carneiro, Carla Osthoff |
IPDPS | 1 |
| 2024 | Scheduling Distributed I/O Resources in HPC Systems
Alexis Bandet, Francieli Zanon Boito, Guillaume Pallez |
Euro-Par (1) | 2 |
| 2024 | Capturing Periodic I/O Using Frequency TechniquesabstractMany HPC applications perform their I/O in bursts that follow a periodic pattern. This allows for making predictions as to when a burst occurs. System providers can take advantage of such knowledge to reduce file-system contention by actively scheduling I/O bandwidth. The effectiveness of this approach, however, depends on the ability to detect and quantify the periodicity of I/O patterns online. In this paper, we introduce FTIO, an online method to detect periodic I/O phases, which is based on discrete Fourier transform (DFT), combined with outlier detection. We provide metrics that gauge the confidence in the output and tell how far from being periodic the signal is. We validate our approach with large-scale experiments on a production system and examine its limitations extensively. Our experiments show that FTIO has a mean error below 11%. Finally, we demonstrate that FTIO allowed the I/O scheduler Set10 to boost system utilization by 26% and reduce I/O slowdown by 56%. Ahmad Tarraf, Alexis Bandet, Francieli Zanon Boito, Guillaume Pallez, Felix Wolf 0001 |
IPDPS | 3 |
| 2023 | IO-Sets: Simple and Efficient Approaches for I/O Bandwidth ManagementabstractOne of the main performance issues faced by high-performance computing platforms is the congestion caused by concurrent I/O from applications. When this happens, the platform's overall performance and utilization are harmed. From the extensive work in this field, I/O scheduling is the essential solution to this problem. The main drawback of current techniques is the amount of information needed about applications, which compromises their applicability. In this paper, we propose a novel method for I/O management,IO-Sets. We present its potential through a scheduling heuristic calledSet-10, which is simple and requires only minimal information. Our extensive experimental campaign shows the importance ofIO-Setsand the robustness ofSet-10under various workloads. In particular in most of the simulated scenarios we improve the I/O slowdown over fairshare by 50%, which corresponds in our scenarios to a platform utilization gain of 2.5%. In the practical scenarios that we did, the utilization gain varies between 10 and 30%. We also provide insights on using our proposal in practice. Francieli Zanon Boito, Guillaume Pallez, Luan Teylo, Nicolas Vidal 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | The role of storage target allocation in applications' I/O performance with BeeGFSabstractParallel file systems are at the core of HPC I/O infrastructures. Those systems minimize the I/O time of applications by separating files into fixed-size chunks and distributing them across multiple storage targets. Therefore, the I/O performance experienced with a PFS is directly linked to the capacity to retrieve these chunks in parallel. In this work, we conduct an in-depth evaluation of the impact of the stripe count (the number of targets used for striping) on the write performance of BeeGFS, one of the most popular parallel file systems today. We consider different network configurations and show the fundamental role played by this parameter, in addition to the number of compute nodes, processes and storage targets. Through a rigorous experimental evaluation, we directly contradict conclusions from related work. Notably, we show that sharing I/O targets does not lead to performance degradation and that applications should use as many storage targets as possible. Our recommendations have the potential to significantly improve the overall write performance of BeeGFS deployments and also provide valuable information for future work on storage target allocation and stripe count tuning. Francieli Zanon Boito, Guillaume Pallez, Luan Teylo |
CLUSTER | 1 |
| 2021 | Arbitration Policies for On-Demand User-Level I/O Forwarding on HPC PlatformsabstractI/O forwarding is a well-established and widely-adopted technique in HPC to reduce contention in the access to storage servers and transparently improve I/O performance. Rather than having applications directly accessing the shared parallel file system, the forwarding technique defines a set of I/O nodes responsible for receiving application requests and forwarding them to the file system, thus reshaping the flow of requests. The typical approach is to statically assign I/O nodes to applications depending on the number of compute nodes they use, which is not always necessarily related to their I/O requirements. Thus, this approach leads to inefficient usage of these resources. This paper investigates arbitration policies based on the applications I/O demands, represented by their access patterns. We propose a policy based on the Multiple-Choice Knapsack problem that seeks to maximize global bandwidth by giving more I/O nodes to applications that will benefit the most. Furthermore, we propose a user-level I/O forwarding solution as an on-demand service capable of applying different allocation policies at runtime for machines where this layer is not present. We demonstrate our approach's applicability through extensive experimentation and show it can transparently improve global I/O bandwidth by up to 85% in a live setup compared to the default static policy. Jean Luca Bez, Alberto Miranda, Ramon Nou, Francieli Zanon Boito, Toni Cortes, Philippe Olivier Alexandre Navaux |
IPDPS | 4 |
| 2020 | On low-power SoCs as storage bricks for BioinformaticsabstractSummary In the perspective of energy‐efficient storage solutions supporting scientific computing, we investigate the possibility of using low‐power Systems‐On‐Chip (SoCs) as storage bricks of a BeeGFS file system in support of a realistic genome sequencing pipeline from Bioinformatics. Metadata and data performances of such a file system made of low‐power devices are assessed via various benchmarks for an increasing number of clients accessing the data, and compared to those achieved with traditional power‐hungry architectures with a similar BeeGFS setup. Then, we show how the sequencing information streamed by portable sequencing devices, such as the Oxford Nanopore MinION, can be managed by the underlying BeeGFS file system built with low‐power SoCs. Lucia Morganti, Elena Corni, Luca Lama, Carmelo Pellegrino, Francieli Zanon Boito, Ivan Merelli, Daniele D'Agostino, Daniele Cesini |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Adaptive request scheduling for the I/O forwarding layer using reinforcement learning
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux |
Future Gener. Comput. Syst. | 2 |
| 2019 | Detecting I/O Access Patterns of HPC Workloads at RuntimeabstractIn this paper, we seek to guide optimization and tuning strategies by identifying the application's I/O access pattern. We evaluate three machine learning techniques to automatically detect the I/O access pattern of HPC applications at runtime: decision trees, random forests, and neural networks. We focus on the detection using metrics from file-level accesses as seen by the clients, I/O nodes, and parallel file system servers. We evaluated these detection strategies in a case study in which the accurate detection of the current access pattern is fundamental to adjust a parameter of an I/O scheduling algorithm. We demonstrate that such approaches correctly classify the access pattern, regarding file layout and spatiality of accesses - into the most common ones used by the community and by I/O benchmarking tools to test new I/O optimization - with up to 99% precision. Furthermore, when applied to our study case, it guides a tuning mechanism to achieve 99% of the performance of an Oracle solution. Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |
| 2019 | An Unsupervised Learning Approach for I/O Behavior CharacterizationabstractI/O operations are the bottleneck of several applications due to the difference between processing and data access speeds. Hence, understanding the I/O behavior is vital to find problems and propose solutions. Thus, identifying and characterizing the I/O access pattern is important, since it reflects directly on applications' performance. With this premise, we propose an I/O characterization approach that uses unsupervised learning to cluster jobs with similar I/O behavior, using information from high-level aggregated traces. As a case study, we apply our approach on four months of activity - a total of 28, 938 jobs - from the Intrepid supercomputer located at Argonne Laboratory. Our experimental results show that nine access patterns represent the I/O behavior in 73% of the clusters. From these nine patterns, we learn some aspects about the I/O such as the most accesses patterns are made using POSIX and small requests, also, the most patterns are accessing unique files. Lastly, analyzing the I/O workload over four months, we can notice that it is composed by several applications that spend a short time on I/O activity, but when compared to the others, the total I/O time represents a greater portion of the overall system. Pablo J. Pavan, Jean Luca Bez, Matheus S. Serpa, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 4 |
| 2019 | Energy efficiency and I/O performance of low-power architecturesabstractSummary This paper presents an energy efficiency and I/O performance analysis of low‐power architectures when compared to conventional architectures, with the goal of studying the viability of using them as storage servers. Our results show that despite the fact the power demand of the storage device amounts for a small fraction of the power demand of the whole system, significant increases in power demand are observed when accessing the storage device. We investigate the access pattern impact on power demand, looking at the whole system and at the storage device by itself, and compare all tested configurations regarding energy efficiency. Then we extrapolate the conclusions from this research to provide guidelines for when considering the replacement of traditional storage servers by low‐power alternatives. We show the choice depends on the expected workload, estimates of power demand of the systems, and factors limiting performance. These guidelines can be applied for other architectures than the ones used in this work. Pablo J. Pavan, Ricardo K. Lorenzoni, Vinícius Machado 0002, Jean Luca Bez, Edson L. Padoin, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 6 |
| 2018 | Collective I/O Performance on the Santos Dumont SupercomputerabstractThe historical gap between processing and data access speeds causes many applications to spend a large portion of their execution on I/O operations. From the point of view of a large-scale, expensive, supercomputer, it is important to ensure applications achieve the best I/O performance to promote an efficient usage of the machine. In this paper, we evaluate the I/O infrastructure of the Santos Dumont supercomputer, the largest one from Latin America. More specifically, we investigate the performance of collective I/O operations. By conducting an analysis of a scientific application that uses the machine, we identify large performance differences between the available MPI implementations. We then further study the observed phenomenon using the BT-IO and IOR benchmarks, in addition to a custom microbenchmark. We conclude that the customized MPI implementation by Bull (used by more than 20% of the jobs) presents the worst performance for small collective write operations. Our results are being used to help the Santos Dumont users to achieve the best performance for their applications. Additionally, by investigating the observed phenomenon, we provide information to help improve future MPI-IO collective write implementations. Andre Ramos Carneiro, Jean Luca Bez, Francieli Zanon Boito, Bruno Alves Fagundes, Carla Osthoff, Philippe Olivier Alexandre Navaux |
PDP | 3 |
| 2017 | TWINS: Server Access Coordination in the I/O Forwarding LayerabstractThis paper presents a study of I/O scheduling techniques applied to the I/O forwarding layer. In high-performance computing environments, applications rely on parallel file systems (PFS) to obtain good I/O performance even when handling large amounts of data. To alleviate the concurrency caused by thousands of nodes accessing a significantly smaller number of PFS servers, intermediate I/O nodes are typically applied between processing nodes and the file system. Each intermediate node forwards requests from multiple clients to the system, a setup which gives this component the opportunity to perform optimizations like I/O scheduling. We evaluate scheduling techniques that improve spatiality and request size of the access patterns. We show they are only partially effective because the access pattern is not the main factor for read performance in the I/O forwarding layer. A new scheduling algorithm, TWINS, is presented to coordinate the access of intermediate I/O nodes to the data servers. Our proposal decreases concurrency at the data servers, a factor previously proven to negatively affect performance. The proposed algorithm is able to improve read performance from shared files by up to 28% over other scheduling algorithms and by up to 50% over not forwarding I/O. Jean Luca Bez, Francieli Zanon Boito, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
PDP | 2 |
| 2017 | High Performance I/O for Seismic Wave Propagation SimulationsabstractThis paper describes our research to provide high performance I/O for seismic wave propagation simulations. Earthquake early warning systems are designed to provide near real-time prediction of strong ground motion. Such systems are crucial tools for risk mitigation and disaster prevention. The ability to accurately and quickly simulate the propagation of seismic waves in complex media lies at the heart of such systems. Besides the processing requirements, it is important for seismic simulations to leverage a high-performance storage infrastructure to output results as frequently as possible, so they can be used for the decision-making process. We propose and evaluate a series of I/O optimizations to the Ondes3D seismic wave propagation simulation, considering its different types of output files separately. These optimizations are designed while keeping the previous output formats, in order not to compromise the application interaction with the other parts of the earthquake early warning system. The optimization techniques presented in this paper have provided I/O performance improvements of up to 85% and decreased the application execution time up to 70%. Francieli Zanon Boito, Jean Luca Bez, Fabrice Dupros, Mario A. R. Dantas, Philippe Olivier Alexandre Navaux, Hideo Aochi |
PDP | 1 |
| 2016 | Automatic I/O scheduling algorithm selection for parallel file systemsabstractSummary This article presents our approach to provide input/output (I/O) scheduling with double adaptivity: to applications and devices. In high‐performance computing environments, parallel file systems provide a shared storage infrastructure to applications. In the situation where multiple applications access this shared infrastructure concurrently, their performance can be impaired because of interference. Our work focuses on I/O scheduling as a tool to improve performance by alleviating interference effects. The role of the I/O scheduler is to decide the order in which applications' requests must be processed by the parallel file system's servers, applying optimizations to adjust the resulting access pattern for improved performance. Our approach to improve I/O scheduling results is based on using information from applications' access patterns and storage devices' sensitivity to access sequentiality. We have applied machine learning to provide the ability to automatically select the best scheduling algorithm for each situation. Our approach improves performance by up to 75% over an approach that uses the same scheduling algorithm to all situations, without adaptability. Our results evidence that both aspects – applications and storage devices – are essential to make good scheduling decisions. Copyright © 2015 John Wiley & Sons, Ltd. Francieli Zanon Boito, Rodrigo Kassick, Philippe Olivier Alexandre Navaux, Yves Denneulin |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | AGIOS: Application-Guided I/O Scheduling for Parallel File SystemsabstractIn this paper, we improve the performance of server-side I/O scheduling on parallel file systems by transparently including information about the applications' access patterns. Server-side I/O scheduling is a valuable tool on multiapplication scenarios, where the applications' spatial locality suffers from interference caused by concurrent accesses to the file system. We present AGIOS, an I/O scheduling library for parallel file systems. We guide scheduler's decisions by including information about the applications' future requests. This information is obtained from traces generated by the scheduler itself, without changes in application or file system. Our approach shows performance improvements under different workloads of 46.3% on average when compared to a scenario without an I/O scheduler, and of 25.1% when compared to a scheduler which does not use information about future accesses. Francieli Zanon Boito, Rodrigo Kassick, Philippe Olivier Alexandre Navaux, Yves Denneulin |
ICPADS | 1 |
| 2011 | Improving Performance on Atmospheric Models through a Hybrid OpenMP/MPI ImplementationabstractThis work shows how a Hybrid MPI/OpenMP implementation can improve the performance of the Ocean-Land-Atmosphere Model (OLAM) on a multi-core cluster environment, which is a typical HPC many small files workload application. Previous experiments have shown that the scalability of this application on clusters is limited by the performance of the output operations. We show that the Hybrid MPI/OpenMP version of OLAM decreases the number of output files, resulting in better performance for I/O operations. We also observe that the MPI version of OLAM performs better for unbalanced workloads and that further parallel optimizations should be included on the hybrid version in order to improve the parallel execution time of OLAM. Carla Osthoff, Pablo Javier Grunmann, Francieli Zanon Boito, Rodrigo Kassick, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Claudio Schepke, Jairo Panetta, Nicolas Maillard, Pedro Leite da Silva Dias, Robert L. Walko |
ISPA | 3 |
| 2011 | Dynamic I/O Reconfiguration for a NFS-Based Parallel File SystemabstractThe large gap between the speed in which data can be processed and the performance of I/O devices makes the shared storage infrastructure of a cluster a great bottle-neck. Parallel File Systems try to smooth such difference by distributing data onto several servers, increasing the system's available bandwidth. However, most implementations use a fixed number of I/O servers, defined during the initialization of the system, and can not add new resources without a complete redistribution of the existing data. With the execution of different applications at the same time, the concurrent access to these resources can aggravate the existing bottleneck, making very hard to define an initial number of servers that satisfies the performance requirements of different applications. This paper presents a reconfiguration mechanism for the dNFSp file system that uses on-line monitoring of application's I/O behavior to detect performance contention and dedicate more I/O resources to applications with higher demands. These extra resources are taken from the available nodes of the cluster, using their I/O devices as a temporary storage. We show that this strategy is capable of increasing the I/O performance in up to 200% for access patterns with short I/O phases and 47% for longer I/O phases. Rodrigo Kassick, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux |
PDP | 2 |
| 2010 | Impact of I/O Coordination on a NFS-Based Parallel File System with Dynamic ReconfigurationabstractThe large gap between processing and I/O speed makes the storage infrastructure of a cluster a great bottleneck for UPC applications. Parallel File Systems propose a solution to this issue by distributing data onto several servers, dividing the load of I/O operations and increasing the available bandwidth. However, most parallel file systems use a fixed number of I/O servers defined during initialization and do not support addition of new resources as applications' demands grow. With the execution of different applications at the same time, the concurrent access to these resources can impact the performance and aggravate the existing bottleneck. The dNFSp File System proposes a reconfiguration mechanism that aims to include new I/O resources as application's demands grow. These resources are standard cluster nodes and are dedicated to a single application. This paper presents a study of the I/O performance of this reconfiguration mechanism under two circumstances: the use of several independent processes on a multi-core system or of a single centralized I/O process that coordinates the requests from all instances on a node. We show that the use of coordination can improve performance of applications with regular intervals between I/O phases. For applications with no such intervals, on the other hand, uncoordinated I/O presents better performance. Rodrigo Kassick, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |