Otávio O. Napoli

dblp:244/1350 · also Otávio Oliveira Napoli · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-9606-4751ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 On the Components That Enable Robust Generalization in HAR Models
abstract
Generalizing human activity recognition (HAR) models across datasets remains challenging due to variations in sensors, environments, and user behavior.Domain Generalization (DG) methods attempt to address these shifts through objective-level modifications, architecture-level augmentations, and model pretraining strategies, but prior HAR studies often evaluate these components in isolation using suboptimal baselines.We systematically assess the contribution of each DG component across multiple HAR architectures, from CNNs to Transformers, using the DAGHAR benchmark.Our results show that pretraining the model with a self-supervised learning technique provides the most substantial and consistent gains in cross-dataset generalization, while architecture-level augmentations offer complementary improvements, and objective-level methods alone yield limited benefits across architectures.This suggests that DG studies should treat model pretraining as a standard baseline rather than an optional enhancement.* This project was supported by the
Otávio O. Napoli, Edson Borin
ESANN1
2025 On Domain Generalization for Human Activity Recognition with Mix-Based Methods
abstract
Domain generalization (DG) is a challenging problem that involves adapting a model trained on source domains to an unseen target domain.In human activity recognition (HAR), domain shifts often arise from differences in sensor placement, device specifications, or environmental factors, making generalization difficult.In this work, we investigate the effectiveness of mix-based methods like MixStyle and Exact Feature Distribution Mixing (EFDM) when integrated into state-of-the-art models like ResNet and TS2Vec for DG in HAR tasks, leveraging the DAGHAR benchmark.Our results demonstrate that MixStyle significantly outperforms both EFDM and Empirical Risk Minimization approaches, highlighting its effectiveness in addressing domain shifts.* This project was
Otávio O. Napoli, Edson Borin
ESANN1
2025 SPINN: a Tool for Distributed Patch Inference on Massive Data Samples
abstract
Patched inference is a widely used technique in machine learning (ML) that enables fixed-shape models to process arbitrarily large or variably sized inputs by dividing them into smaller, compatible patches. This approach is particularly useful in domains such as seismic processing, medical imaging, and electron microscopy, where data samples often exceed the memory capacity of individual computing nodes. While patched inference is effective for leveraging pre-trained models and operating on resource-constrained hardware, there remains a lack of tools supporting its efficient, distributed execution at scale.To address this gap, we introduce SPINN (Scalable Parallel INference Network), a Python library designed to streamline and accelerate patched inference on high-performance computing (HPC) systems. SPINN supports data partitioning, patch-wise processing using user-defined ML models, and result aggregation, all while leveraging distributed computing frameworks such as Dask and Ray.We validate SPINN on two seismic interpretation tasks, fault detection and facies segmentation, using both public and large-scale private data (up to 272 GB). Experiments demonstrate that SPINN enables smoother prediction outputs via overlapping patches and achieves superlinear scalability with Dask in HPC environments, significantly outperforming conventional solutions such as the NVIDIA Triton Inference Server in large-scale scenarios. SPINN thus emerges as a robust and scalable solution for applying deep learning inference to massive data samples in memory-constrained or compute-intensive settings.
João Seródio, Júlio César Faracco, Fernando Gubitoso, Otávio O. Napoli, Alan Souza, Daniel Miranda, Carlos A. Astudillo, Edson Borin
SBAC-PAD4
2024 Memory-efficient DRASiW Models
Otávio O. Napoli, Ana de Almeida 0002, Edson Borin, Maurício Breternitz
Neurocomputing1
2023 Efficient Knowledge Aggregation Methods for Weightless Neural Networks
abstract
Weightless Neural Networks (WNN) are good candidates for Federated Learning scenarios due to their robustness and computational lightness.In this work, we show that it is possible to aggregate the knowledge of multiple WNNs using more compact data structures, such as Bloom Filters, to reduce the amount of data transferred between devices.Finally, we explore variations of Bloom Filters and found that a particular data-structure, the Count-Min Sketch (CMS), is a good candidate for aggregation.Costing at most 3% of accuracy, CMS can be up to 3x smaller when compared to previous approaches, specially for large datasets.
Otávio O. Napoli, Ana de Almeida 0002, José Miguel Salles Dias, Luís Brás Rosário, Edson Borin, Maurício Breternitz
ESANN1
2023 PB3Opt: Profile-based biased Bayesian optimization to select computing clusters on the cloud
abstract
Summary Given the wide variety of cloud computing resources for creating high‐performance computer clusters and their complex performance relationship with applications, finding the optimal, or near‐optimal, cluster is a complex problem. As a result, several approaches have been proposed to find the optimal, or near‐optimal, cluster for a given high‐performance computing workload, while reducing the search cost. Among the approaches found in the literature, Bayesian optimization is one of the most known and applied. However, it is still possible to increase its performance by integrating it with historical data related to workload behavior. In this context, we suggest the approach, which introduces a bias in the Bayesian optimization expected improvement acquisition function. The new acquisition function uses the ranking of computer clusters of previously explored workloads that have the same behavior as the workload being optimized. Our experimental results show that classifies the behavior of workloads in groups so that the average‐ranking has 88.7% similarity with the ranking of the workload. With this, finds, for almost 95% of workloads, a solution that is less than or equal to 1.2 worse than the optimal computer cluster. In addition, the works well when combined with the paramount iterations technique and is capable of reducing the search cost significantly.
Thais Aparecida Silva Camacho, Vanderson Martins do Rosário, Otávio O. Napoli, Edson Borin
Concurr. Comput. Pract. Exp.3
2023 Fast selection of compiler optimizations using performance prediction with graph neural networks
abstract
Abstract Tuning application performance on modern computing infrastructures involves choices in a vast design space as modern computing architectures can have several complex structures impacting performance. Moreover, different applications use these structures in different ways, leading to a challenging performance function. Consequently, it is hard for compilers or experts to find optimal compilation parameters for an application that maximizes such performance function. One approach to tackle this problem is to evaluate many possible optimization plans and select the best among them. However, executing an application to measure its performance for every plan can be very expensive. To tackle this problem, previous work has investigated the use of Machine Learning techniques to predict the performance of the applications without executing them quickly. In this work, we evaluate the use of graph neural networks (GNN) to make fast predictions without executing the application to guide the selection of good optimization sequences. We propose a GNN architecture to make such predictions. We train and test it using 30 thousand different compilation plans applied to 300 different applications, using ARM64 and LLVM IR code representations as input. Our results indicate that the control and data flow graph can then learn features from the control and data flow graph to outperform nongraph‐aware Machine Learning models. Our GNN architecture achieved 91% accuracy in our dataset compared to 79% when using a nongraph‐aware architecture–taking only 16ms to predict a given input. If the application been optimized took an average of 10 s to execute, and we evaluated 1000 optimization sequences, it would take almost 9 h to assess all pairs, but only 16 s with our GNN .
Vanderson Martins do Rosário, Anderson Faustino da Silva, André Felipe Zanella, Otávio O. Napoli, Edson Borin
Concurr. Comput. Pract. Exp.4
2021 Smart selection of optimizations in dynamic compilers
abstract
Summary Dynamic compilers perform compilation and generation of target code during runtime, implying that the compilation time is added into the program runtime. Thus, to build a high‐performing dynamic compilation system, it is crucial to be able to generate high‐quality code and, at the same time, have a small compilation cost. In this article, we present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Our approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high‐quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. We evaluate the proposed strategy in two LLVM‐based dynamic binary translators, namely OI‐DBT and HQEMU, and show that these two translators can achieve average speedups of 1.26x and 1.15x in MiBench and Spec Cpu benchmarks, respectively.
Vanderson Martins do Rosário, Anderson Faustino da Silva, Thais Aparecida Silva Camacho, Otávio O. Napoli, Maurício Breternitz, Edson Borin
Concurr. Comput. Pract. Exp.4