Luís Paulo F. Garcia

dblp:25/7075 · also Luís Paulo Faina Garcia · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0003-0679-9143ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A novel open set Energy-based Flow Classifier for Network Intrusion Detection
Manuela M. C. de Souza, Camila F. T. Pontes, João J. C. Gondim, Luís Paulo F. Garcia, Luiz A. DaSilva, Eduardo F. M. Cavalcante, Marcelo Antonio Marotta
Comput. Secur.4
2024 DODFMiner: An automated tool for Named Entity Recognition from Official Gazettes
Gabriel M. C. Guimarães, Felipe X. B. da Silva, Andrei L. Queiroz, Ricardo M. Marcacini, Thiago de Paulo Faleiros, Vinicius Ruela Pereira Borges, Luís Paulo F. Garcia
Neurocomputing7
2024 A review on preprocessing algorithm selection with meta-learning
Pedro B. Pio, Adriano Rivolli, André C. P. L. F. de Carvalho, Luís Paulo F. Garcia
Knowl. Inf. Syst.4
2023 Neural architecture search with interpretable meta-features and fast predictors
Gean Trindade Pereira, Iury Batista de Andrade Santos, Luís Paulo F. Garcia, Thierry Urruty, Muriel Visani, André C. P. L. F. de Carvalho
Inf. Sci.3
2022 Meta-features for meta-learning
abstract
a b s t r a c tMeta-learning is increasingly used to support the recommendation of machine learning algorithms and their configurations.These recommendations are made based on meta-data, consisting of performance evaluations of algorithms and characterizations on prior datasets.These characterizations, also called meta-features, describe properties of the data which are predictive for the performance of machine learning algorithms trained on them.Unfortunately, despite being used in many studies, meta-features are not uniformly described, organized and computed, making many empirical studies irreproducible and hard to compare.This paper aims to deal with this by systematizing and standardizing data characterization measures for classification datasets used in meta-learning.Moreover, it presents an extensive list of meta-features and characterization tools, which can be used as a guide for new practitioners.By identifying particularities and subtle issues related to the characterization measures, this survey points out possible future directions that the development of meta-features for meta-learning can assume.
Adriano Rivolli, Luís Paulo F. Garcia, Carlos Soares, Joaquin Vanschoren, André C. P. L. F. de Carvalho
Knowl. Based Syst.2
2021 Towards holistic Entity Linking: Survey and directions
Italo Lopes Oliveira, Renato Fileto, René Speck, Luís Paulo F. Garcia, Diego Moussallem, Jens Lehmann 0001
Inf. Syst.4
2021 Assessing the data complexity of imbalanced datasets
Victor H. Barella, Luís Paulo F. Garcia, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, André C. P. L. F. de Carvalho
Inf. Sci.2
2020 Algorithm Recommendation for Data Streams
abstract
In the last decades, many companies have taken advantage of knowledge discovery to identify valuable information in massive volumes of data generated at high frequency. Machine learning techniques can be employed for knowledge discovery since they can extract patterns from data and induce models to predict future events. However, dynamic and evolving environments usually generate non-stationary data streams. Hence, models trained in these scenarios may perish over time due to seasonality or concept drift. Periodic retraining can help, but a fixed hypothesis space may no longer be appropriate. An alternative solution is to use meta-learning for regular algorithm selection in time-changing environments, choosing the bias that best suits the current data. In this paper, we present an enhanced framework for data stream algorithm selection based on MetaStream. Our approach uses meta-learning and incremental learning to actively select the best algorithm for the current concept in a time-changing environment. Different from previous work, we use a rich set of state-of-the-art meta-features, and an incremental learning approach in the meta-level based on LightGBM. The results show that this new strategy can improve the recommendation accuracy of the best algorithm in time-changing data.
Jáder Martins Camboim de Sá, André Luis Debiaso Rossi, Gustavo Batista, Luís Paulo F. Garcia
ICPR4
2020 Boosting meta-learning with simulated data complexity measures
abstract
Meta-Learning has been largely used over the last years to support the recommendation of the most suitable machine learning algorithm(s) and hyperparameters for new datasets. Traditionally, a meta-base is created containing meta-features extracted from several datasets along with the performance of a pool of machine learning algorithms when applied to these datasets. The meta-features must describe essential aspects of the dataset and distinguish different problems and solutions. However, if one wants the use of Meta-Learning to be computationally efficient, the extraction of the meta-feature values should also show a low computational cost, considering a trade-off between the time spent to run all the algorithms and the time required to extract the meta-features. One class of measures with successful results in the characterization of classification datasets is concerned with estimating the underlying complexity of the classification problem. These data complexity measures take into account the overlap between classes imposed by the feature values, the separability of the classes and distribution of the instances within the classes. However, the extraction of these measures from datasets usually presents a high computational cost. In this paper, we propose an empirical approach designed to decrease the computational cost of computing the data complexity measures, while still keeping their descriptive ability. The proposal consists of a novel Meta-Learning system able to predict the values of the data complexity measures for a dataset by using simpler meta-features as input. In an extensive set of experiments, we show that the predictive performance achieved by Meta-Learning systems which use the predicted data complexity measures is similar to the performance obtained using the original data complexity measures, but the computational cost involved in their computation is significantly reduced.
Luís Paulo F. Garcia, Adriano Rivolli, Edesio Alcobaça, Ana Carolina Lorena, André C. P. L. F. de Carvalho
Intell. Data Anal.1
2020 MFE: Towards reproducible meta-feature extraction
abstract
Automated recommendation of machine learning algorithms is receiving a large deal of attention, not only because they can recommend the most suitable algorithms for a new task, but also because they can support efficient hyper-parameter tuning, leading to better machine learning solutions. The automated recommendation can be implemented using meta-learning, learning from previous learning experiences, to create a meta-model able to associate a data set to the predictive performance of machine learning algorithms. Although a large number of publications report the use of meta-learning, reproduction and comparison of meta-learning experiments is a difficult task. The literature lacks extensive and comprehensive public tools that enable the reproducible investigation of the different meta-learning approaches. An alternative to deal with this difficulty is to develop a meta-feature extractor package with the main characterization measures, following uniform guidelines that facilitate the use and inclusion of new meta-features. In this paper, we propose two Meta-Feature Extractor (MFE) packages, written in both Python and R, to fill this lack. The packages follow recent frameworks for meta-feature extraction, aiming to facilitate the reproducibility of meta-learning experiments.
Edesio Alcobaça, Felipe Siqueira, Adriano Rivolli, Luís Paulo F. Garcia, Jefferson Tales Oliva, André C. P. L. F. de Carvalho
J. Mach. Learn. Res.4
2019 New label noise injection methods for the evaluation of noise filters
abstract
Noise is often present in real datasets used for training Machine Learning classifiers. Their disruptive effects in the learning process may include: increasing the complexity of the induced models, a higher processing time and a reduced predictive power in the classification of new examples. Therefore, treating noisy data in a preprocessing step is crucial for improving data quality and to reduce their harmful effects in the learning process. There are various filters using different concepts for identifying noisy examples in a dataset. Their ability in noise preprocessing is usually assessed in the identification of artificial noise injected into one or more datasets. This is performed to overcome the limitation that only a domain expert can guarantee whether a real example is indeed noisy. The most frequently used label noise injection method is the noise at random method, in which a percentage of the training examples have their labels randomly exchanged. This is carried out regardless of the characteristics and example space positions of the selected examples. This paper proposes two novel methods to inject label noise in classification datasets. These methods, based on complexity measures, can produce more challenging and realistic noisy datasets by the disturbance of the labels of critical examples situated close to the decision borders and can improve the noise filtering evaluation. An extensive experimental evaluation of different noise filters is performed using public datasets with imputed label noise and the influence of the noise injection methods are compared in both data preprocessing and classification steps.
Luís Paulo F. Garcia, Jens Lehmann 0001, André C. P. L. F. de Carvalho, Ana Carolina Lorena
Knowl. Based Syst.1
2018 Classifier Recommendation Using Data Complexity Measures
abstract
Application of machine learning to new and unfamiliar domains calls for increasing automation in choosing a learning algorithm suitable for the data arising from each domain. Meta-learning could address this need since it has been largely used in the last years to support the recommendation of the most suitable algorithms for a new dataset. The use of complexity measures could increase the systematic comprehension over the meta-models and also allow to differentiate the performance of a set of techniques taking into account the overlap between classes imposed by feature values, the separability and distribution of the data points. In this paper we compare the effectiveness of several standard regression models in predicting the accuracies of classifiers for classification problems from the OpenML repository. We show that the models can predict the classifiers' accuracies with low mean-squared-error and identify the best classifier for a problem that results in statistically significant improvements over a randomly chosen classifier or a fixed classifier believed to be good on average.
Luís Paulo F. Garcia, Ana Carolina Lorena, Marcílio Carlos Pereira de Souto, Tin Kam Ho
ICPR1
2018 Data Complexity Measures for Imbalanced Classification Tasks
abstract
In imbalanced classification tasks, the training datasets may show class overlapping and classes of low density. In these scenarios, the predictions for the minority class are impaired. Although assessing the imbalance level of a training set is straightforward, it is hard to measure other aspects that may affect the predictive performance of classification algorithms in imbalanced tasks. This paper presents a set of measures designed to understand the difficulty of imbalanced classification tasks by regarding on each class individually. They are adapted from popular data complexity measures for classification problems, which are shown to perform poorly in imbalanced scenarios. Experiments on synthetic datasets with different levels of imbalance, class overlapping and density of the classes show that the proposed adaptations can better explain the difficulty of imbalanced classification tasks.
Victor H. Barella, Luís Paulo F. Garcia, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, André C. P. L. F. de Carvalho
IJCNN2
2016 Ensembles of label noise filters: a ranking approach
Luís Paulo F. Garcia, Ana Carolina Lorena, Stan Matwin, André C. P. L. F. de Carvalho
Data Min. Knowl. Discov.1
2016 Noise detection in the meta-learning level
Luís Paulo F. Garcia, André C. P. L. F. de Carvalho, Ana Carolina Lorena
Neurocomputing1
2015 Effect of label noise in the complexity of classification problems
Luís Paulo F. Garcia, André C. P. L. F. de Carvalho, Ana Carolina Lorena
Neurocomputing1
2015 Using the One-vs-One decomposition to improve the performance of class noise filters via an aggregation strategy in multi-class classification problems
abstract
Noise filters are preprocessing techniques designed to improve data quality in classification tasks by detecting and eliminating examples that contain errors or noise. However, filtering can also remove correct examples and examples containing valuable information, which could be useful for learning. This fact usually implies a margin of improvement on the noise detection accuracy for almost any noise filter. This paper proposes a scheme to improve the performance of noise filters in multi-class classification problems, based on decomposing the dataset into multiple binary subproblems. Decomposition strategies have proven to be successful in improving classification performance in multi-class problems by generating simpler binary subproblems. Similarly, we adapt the principles of the One-vs-One decomposition strategy to noise filtering, making the noise identification process simpler. In order to integrate the filtering results achieved in the binary subproblems, our proposal uses a soft voting approach considering a reliability level based on the aggregation of the noise degree prediction calculated for each binary classifier. The experimental results show that the One-vs-One decomposition strategy usually increases the performance of the noise filters studied, which can detect more accurately the noisy examples.
Luís Paulo F. Garcia, José A. Sáez, Julián Luengo, Ana Carolina Lorena, André C. P. L. F. de Carvalho, Francisco Herrera
Knowl. Based Syst.1