Samuel Baraldi Mafra

dblp:121/2498 · also Samuel B. Mafra, Samuel Mafra · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2026
0000-0002-5238-1795ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 2
YearPublicationVenuePosition
2026 Large-Scale Benchmarking of Intrusion Detection Datasets With GPU-Accelerated Data Pipelines, Complexity Analysis, and Model Evaluation
abstract
Intrusion detection systems (IDSs) are critical for identifying malicious activity in computer networks; however, the evaluation of machine learning (ML)–based IDS remains inconsistent and fragmented. Many existing studies rely on outdated datasets, neglect computational complexity, or use limited performance metrics. Additionally, few works leverage the full potential of modern graphics processing unit (GPU) acceleration. The objective of this study is to establish a scalable, reproducible, and standardized benchmarking framework for intrusion detection. We present an end‐to‐end, GPU‐accelerated pipeline that integrates automated data preprocessing, intrinsic dataset complexity analysis, and multiobjective hyperparameter optimization (HPO) across more than 70 publicly available datasets. Our numerical findings demonstrate that stratified sampling rates of 10% are sufficient to maintain statistical signal integrity, with class probability deviations remaining below 0.01 relative to the full population. Furthermore, feature‐reduced configurations decrease the model size by a median of 60% while maintaining weighted F 1 scores within 0.01 of the baseline. Finally, experimental complexity analysis reveals that the GPU‐accelerated modeling stages achieve empirical time‐invariance ( O (1)), reducing training latency by up to two orders of magnitude compared with traditional central processing unit (CPU) workflows. These contributions offer a rigorous quantitative view of the performance‐efficiency trade‐offs essential for next‐generation IDS evaluation.
Marcelo V. C. Aragão, Felipe A. P. de Figueiredo, Samuel Baraldi Mafra
Int. J. Intell. Syst.3
2025 Dynamic-Balancing AutoML for Imbalanced Tabular Data With Adaptive Resampling and Complexity-Aware Analysis
abstract
Handling class imbalance is a fundamental challenge in supervised learning, particularly in real‐world scenarios where minority classes are critical yet underrepresented. This paper presents a novel dynamic‐balancing pipeline that enhances automated machine learning (AutoML) performance on imbalanced tabular datasets. The proposed approach integrates both traditional and generative resampling techniques with adaptive, class‐specific thresholds, enabling automated and dataset‐sensitive balancing strategies. To assess its generalizability, the pipeline is applied uniformly across binary, multiclass, and multilabel classification tasks. Each configuration is evaluated within an AutoML framework using performance and efficiency metrics, with outcomes validated through statistical testing and effect size analysis. The study also incorporates dataset complexity measures—including feature‐label dependency and class overlap—to investigate how structural characteristics affect balancing efficacy. By combining principled resampling, exhaustive grid search, and rigorous evaluation, the pipeline enables more robust and efficient AutoML workflows. This work contributes a flexible and reproducible framework for addressing class imbalance, particularly in multilabel contexts, and establishes a foundation for scalable, complexity‐aware resampling in automated model development.
Marcelo V. C. Aragão, Tiago de M. Pereira, Mateus de Freitas Carvalho, Felipe A. P. de Figueiredo, Samuel Baraldi Mafra
Int. J. Intell. Syst.5