Antonio Rafael Sabino Parmezan

dblp:177/1796 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-1725-132XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Content-Based Macroscopic Microbial Image Retrieval
Antonio Rafael Sabino Parmezan, Angela Patricia Mestas Muñante, Diego Minatel, Solange Oliveira Rezende
IEEE Big Data1
2025 A Spatio-Temporal Approach for Identifying Microorganisms in Short Image Sequences
Antonio Rafael Sabino Parmezan, João Pedro Ribeiro da Silva, Diego Minatel, Solange Oliveira Rezende
IEEE Big Data1
2024 Fine-tuning pre-trained neural networks for medical image classification in small clinical datasets
abstract
Convolutional neural networks have been effective in several applications, arising as a promising supporting tool in a relevant Dermatology problem: skin cancer diagnosis. However, generalizing well can be difficult when little training data is available. The fine-tuning transfer learning strategy has been employed to differentiate properly malignant from non-malignant lesions in dermoscopic images. Fine-tuning a pre-trained network allows one to classify data in the target domain, occasionally with few images, using knowledge acquired in another domain. This work proposes eight fine-tuning settings based on convolutional networks previously trained on ImageNet that can be employed mainly in limited data samples to reduce overfitting risk. They differ on the architecture, the learning rate and the number of unfrozen layer blocks. We evaluated the settings in two public datasets with 104 and 200 dermoscopic images. By finding competitive configurations in small datasets, this paper illustrates that deep learning can be effective if one has only a few dozen malignant and non-malignant lesion images to study and differentiate in Dermatology. The proposal is also flexible and potentially useful for other domains. In fact, it performed satisfactorily in an assessment conducted in a larger dataset with 746 computerized tomographic images associated with the coronavirus disease.
Newton Spolaôr, Huei Diana Lee, Ana Isabel Mendes, Conceição Veloso Nogueira, Antonio Rafael Sabino Parmezan, Weber Shoity Resende Takaki, Cláudio Saddy Rodrigues Coy, Feng Chung Wu, Rui Fonseca-Pinto
Multim. Tools Appl.5
2023 DIF-SR: A Differential Item Functioning-Based Sample Reweighting Method
Diego Minatel, Antonio Rafael Sabino Parmezan, Mariana Curi, Alneu de Andrade Lopes
CIARP2
2023 Fairness-Aware Model Selection Using Differential Item Functioning
abstract
Differential Item Functioning (DIF) is a powerful tool for developing fairer tests and mitigating bias in applicant selection tests. DIF aims to detect items in a test that favor or harm groups of people based on aspects such as gender, age, and race, which should be irrelevant to the assessment. Likewise, in machine learning, selecting a model from a pool of candidates is essential to identify the one that minimizes or eliminates discriminatory effects in its decision-making process. As far as we know, research into knowledge discovery through supervised machine learning has predominantly focused on including fair-ness notions at the lowest level of the pre-processing, pattern extraction, and post-processing phases. This fact evidences a need for studies on the impact of model selection on the development of impartial models. Herein, we present a novel approach to fairness-aware model selection to fill the mentioned gap. Our proposal introduces ABC, the first group fairness metric based on DIF concepts. We experimentally evaluated our approach against two model selection strategies by employing ten datasets, six classification algorithms, one performance measure, four group fairness measures, and one statistical significance test. According to the results, our proposal stands out for achieving a trade-off between improving the sense of justice and good classifier performance. Consequently, ABC is a promising metric for selecting fairer models with high predictive power.
Diego Minatel, Antonio Rafael Sabino Parmezan, Mariana Curi, Alneu de Andrade Lopes
ICMLA2
2021 A Graph-Based Spatial Cross-Validation Approach for Assessing Models Learned with Selected Features to Understand Election Results
abstract
Elections are complex activities fundamental to any democracy. The contextualized analysis of election data allows us to understand electoral behavior and the factors that influence it. Multidisciplinary studies have been prioritized the predictive modeling of electoral features from thousands of explanatory features, considering geographic and spatial aspects inherent to the data. When building a model for such a purpose, it must be rigorously evaluated to understand its prediction error in future test cases. Although cross-validation is a widely used procedure for this task, it leads to optimistic results because the spatial independence between test and training data is not ensured in the resampling. On the other hand, alternatives to deal with spatial dependence may fall into a pessimistic scenario by assuming total spatial independence between the test and training sets regardless of the size of the first one, increasing the probability of overfitting. This paper addresses these issues by proposing a graph-based spatial cross-validation approach to assess models learned with selected features from spatially contextualized electoral datasets. Our approach takes advantage of the spatial graph structure provided by the lattice-type spatial objects to define a local training set to each test fold. We generate the local training sets by removing spatially close data that are highly correlated and irrelevant distant data that may interfere with error estimates. Experiments involving the second round of the 2018 Brazilian presidential election demonstrate that our approach contributes to the fair evaluation of models by enabling more realistic and local modeling.
Tiago Pinho da Silva, Antonio Rafael Sabino Parmezan, Gustavo Batista
ICMLA2
2021 Automatic recommendation of feature selection algorithms based on dataset characteristics
abstract
Feature selection in real-world data mining problems is essential to make the learning task efficient and more accurate. Identifying the best feature selection algorithm, among the many available, is a complex activity that still relies heavily on human experts or some random trial-and-error procedure. Thus, the automated machine learning community has taken some steps towards the automation of this process. In this paper, we address the metalearning challenge of recommending feature selection algorithms by proposing a novel meta-feature engineering model. Our model considers a broad collection of meta-features that enable the study of the relationship between the dataset properties and the feature selection algorithm performance in terms of several criteria. We arrange the input meta-features into eight categories: (i) simple, (ii) statistical, (iii) information-theoretical, (iv) complexity, (v) landmarking, (vi) based on symbolic models, (vii) based on images, and (viii) based on complex networks (graphs). The target meta-features emerge from a multi-criteria performance measure, based on five individual performance indexes, that assesses feature selection methods grounded in information, distance, dependence, consistency, and precision measures. We evaluate our proposal using a recently developed framework that extracts the input meta-features from 213 benchmark datasets, and ranks the assessed feature selection algorithms, to fill in the target meta-features in meta-bases. This evaluation uses five state-of-the-art classification methods to induce recommendation models from meta-bases: C4.5, Random Forest, XGBoost, ANN, and SVM. The results showed that it is possible to reach an average accuracy of up to 90% applying our meta-feature engineering model. This work is the first to use an extensive empirical evaluation to provide a careful discussion of the strengths and limitations of more than 160 meta-features. These meta-features, while designed to aid the task of feature selection algorithm recommendation, can readily be employed in other metalearning scenarios. Therefore, we believe our findings are a valuable contribution to the fields of automated machine learning and data mining, as well as to the feature extraction and pattern recognition communities.
Antonio Rafael Sabino Parmezan, Huei Diana Lee, Newton Spolaôr, Feng Chung Wu
Expert Syst. Appl.1
2021 Efficient unsupervised drift detector for fast and high-dimensional data streams
Vinícius M. A. de Souza, Antonio Rafael Sabino Parmezan, Farhan Asif Chowdhury, Abdullah Mueen
Knowl. Inf. Syst.2
2021 A video indexing and retrieval computational prototype based on transcribed speech
Newton Spolaôr, Huei Diana Lee, Weber Shoity Resende Takaki, Leandro Augusto Ensina, Antonio Rafael Sabino Parmezan, Jefferson Tales Oliva, Cláudio Saddy Rodrigues Coy, Feng Chung Wu
Multim. Tools Appl.5
2019 Evaluation of statistical and machine learning models for time series prediction: Identifying the state-of-the-art and the best conditions for the use of each model
abstract
The choice of the most promising algorithm to model and predict a particular phenomenon is one of the most prominent activities of the temporal data forecasting. Forecasting (or prediction), similarly to other data mining tasks, uses empirical evidence to select the most suitable model for a problem at hand since no modeling method can be considered as the best. However, according to our systematic literature review of the last decade, few scientific publications rigorously expose the benefits and limitations of the most popular algorithms for time series prediction. At the same time, there is a limited performance record of these models when applied to complex and highly nonlinear data. In this paper, we present one of the most extensive, impartial and comprehensible experimental evaluations ever done in the time series prediction field. From 95 datasets, we evaluate eleven predictors, seven parametric and four non-parametric, employing two multi-step-ahead projection strategies and four performance evaluation measures. We report many lessons learned and recommendations concerning the advantages, drawbacks, and the best conditions for the use of each model. The results show that SARIMA is the only statistical method able to outperform, but without a statistical difference, the following machine learning algorithms: ANN, SVM, and kNN-TSPI. However, such forecasting accuracy comes at the expense of a larger number of parameters. The evaluated datasets, as well detailed results achieved by different indexes as MSE, Theil’s U coefficient, POCID, and a recently-proposed multi-criteria performance measure are available online in our repository. Such repository is another contribution of this paper since other researchers can replicate our results and evaluate their methods more rigorously. The findings of this study will impact further research on this topic since they provide a broad insight into models selection, parameters setting, evaluation measures, and experimental setup.
Antonio Rafael Sabino Parmezan, Vinícius M. A. de Souza, Gustavo Batista
Inf. Sci.1
2018 Towards Hierarchical Classification of Data Streams
Antonio Rafael Sabino Parmezan, Vinícius M. A. de Souza, Gustavo Batista
CIARP1
2018 Dermoscopic assisted diagnosis in melanoma: Reviewing results, optimizing methodologies and quantifying empirical guidelines
Huei Diana Lee, Ana Isabel Mendes, Newton Spolaôr, Jefferson Tales Oliva, Antonio Rafael Sabino Parmezan, Feng Chung Wu, Rui Fonseca-Pinto
Knowl. Based Syst.5
2017 Metalearning for choosing feature selection algorithms in data mining: Proposal of a new framework
Antonio Rafael Sabino Parmezan, Huei Diana Lee, Feng Chung Wu
Expert Syst. Appl.1
2015 A Study of the Use of Complexity Measures in the Similarity Search Process Adopted by kNN Algorithm for Time Series Prediction
abstract
In the last two decades, with the rise of the Data Mining process, there is an increasing interest in the adaptation of Machine Learning methods to support Time Series non-parametric modeling and prediction. The non-parametric temporal data modeling can be performed according to local and global approaches. The most of the local prediction data strategies are based on the k-Nearest Neighbor (kNN) learning method. In this paper we propose a modification of the kNN algorithm for Time Series prediction. Our proposal differs from the literature by incorporating three techniques for obtaining amplitude and offset invariance, complexity invariance, and treatment of trivial matches. We evaluate the proposed method with six complexity measures, in order to verify the impact of these measures in the projection of the future values. Besides, we face our method with two Machine Learning regression algorithms. The experimental comparisons were performed using 55 data sets, which are available at the ICMC-USP Time Series Prediction Repository. Our results indicate that the developed method is competitive and the use of a complexity-invariant distance measure generally improves the predictive performance.
Antonio Rafael Sabino Parmezan, Gustavo Batista
ICMLA1