Ricardo Cerri

dblp:94/7227 · DBLP profile ↗
← Back
48ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0002-2582-1695ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Oxytrees: Model Trees for Bipartite Learning
abstract
Bipartite learning is a machine learning task that aims to predict interactions between pairs of instances. It has been applied to various domains, including drug-target interactions, RNA-disease associations, and regulatory network inference. Despite being widely investigated, current methods still present drawbacks, as they are often designed for a specific application and thus do not generalize to other problems or present scalability issues. To address these challenges, we propose Oxytrees: proxy-based biclustering model trees. Oxytrees compress the interaction matrix into row- and column-wise proxy matrices, significantly reducing training time without compromising predictive performance. We also propose a new leaf-assignment algorithm that significantly reduces the time taken for prediction. Finally, Oxytrees employ linear models using the Kronecker product kernel in their leaves, resulting in shallower trees and thus even faster training. Using 15 datasets, we compared the predictive performance of ensembles of Oxytrees with that of the current state-of-the-art. We achieved up to 30-fold improvement in training times compared to state-of-the-art biclustering forests, while demonstrating competitive or superior performance in most evaluation settings, particularly in the inductive setting. Finally, we provide an intuitive Python API to access all datasets, methods and evaluation measures used in this work, thus enabling reproducible research in this field.
Pedro Ilídio, Felipe Kenji Nakano, Alireza Gharahighehi, Robbe D'hondt, Ricardo Cerri, Celine Vens
AAAI5
2026 Automated Machine Learning in medical research: A systematic literature mapping study
Giovanna A. Castro, Luiza G. Barioto, Yu H. Cao, Renato Moraes Silva, Helena de Medeiros Caseli, João A. Machado-Neto, Ricardo Cerri, Aline Villavicencio, Tiago A. Almeida 0001
Artif. Intell. Medicine7
2025 Grammar-based Evolutionary Approaches for Software Effort Estimation
abstract
Software effort estimation predicts resources needed for a project, including person-hours and costs, and is vital for effective planning and budgeting. This paper compares two grammar-based evolutionary algorithms: grammar-based genetic programming (GGP) and grammatical evolution (GE). Both algorithms are tested on public project datasets and compared with machine learning models such as support vector machines, artificial neural networks, and least-squares linear regression. Results demonstrate that GGP and GE outperform alternative methods across two evaluation metrics, highlighting their effectiveness in estimating software effort.
Márcio P. Basgalupp, Rodrigo C. Barros, Ricardo Cerri, Ferrante Neri, Péricles B. C. Miranda, Teresa Bernarda Ludermir
CEC3
2025 Inductive models for structured output prediction of lncRNA-disease associations
abstract
Long non-coding RNAs have gained significant attention due to their crucial roles in the pathogenesis of complex human diseases, such as neurological diseases, cardiovascular diseases, AIDS, diabetes, and various types of cancer. In the machine learning literature, lncRNA-disease association (LDA) has been widely investigated as a binary classification problem, where each lncRNA-disease pair is seen as an independent instance. This approach presents drawbacks as it does not exploit the correlation among the diseases, aggravates the already imbalanced dataset, and substantially increases the execution time. Furthermore, the literature focuses on the transductive setting where new disease associations are predicted in lncRNAs already seen by the model, which naturally restricts its application to already seen lncRNAs. As a solution, we propose to address LDA prediction as a structured output prediction problem, namely (hierarchical) multi-label classification, where all LDAs are predicted at once for a given lncRNA. We compared several LDA methods and their structured output variants with recent (hierarchical) multi-label classification methods in an inductive setting, e.g., disease associations are predicted in unseen lncRNAs. Our experiments reveal that approaching LDA prediction with structured output prediction leads to superior or competitive results while drastically reducing the running time.
Felipe Kenji Nakano, Livia Bertoni, Ricardo Cerri, Celine Vens
CIBCB3
2025 ARM-stream: active recovery of miscategorizations in clustering-based data stream classifiers
Douglas Monteiro Cavalcanti, Ricardo Cerri, Elaine Ribeiro de Faria
Data Min. Knowl. Discov.2
2025 Multi-label classification with label clusters
Elaine Cecília Gatto, Mauri Ferrandin, Ricardo Cerri
Knowl. Inf. Syst.3
2025 Pdarts: projected differentiable architecture search for seismic inversion
Lucas C. Souza, Carlos G. C. Junior, Ricardo Cerri, Edson S. Gomi, Bruno S. Carmo, Hermes Senger, Murilo Naldi
Mach. Learn.3
2024 Classification of LTR Retrotransposons via Interaction Prediction
abstract
Transposable Elements (TEs) are genetic sequences that can relocate within the genome, promoting genetic diversity. In eukaryotes, TEs are classified into classes, subclasses, orders, superfamilies, families, and subfamilies. LTR retrotransposons (LTR-RT) constitute an order in this taxonomy. The main objective of this study is to investigate the classification of LTR retrotransposons at the superfamily level. Predictive Bi-Clustering Trees (PBCTs) were used to predict interactions between LTR-RT sequences and conserved protein domains to achieve this. Two datasets were used to investigate the relationships among different superfamilies. The first dataset contained LTR retrotransposon sequences assigned to Copia, Gypsy, and Bel-Pao superfamilies, while the second dataset included consensus sequences of the conserved domains for each superfamily. Thus, the PBCT decision tree tests could relate to the LTR-RT sequence and conserved domain attributes. In the classification process, interaction is interpreted as either the presence or absence of a domain in a given LTR-RT sequence. The sequence is then classified into the superfamily with the most predicted domains. Precision-recall curves were adopted as evaluation metrics for the method, and its performance was compared to some of the most commonly used models in the task of transposable element classification. The experiments conducted on D. melanogaster and A. thaliana showed that PBCTs are promising and comparable to other methods, especially in classifying the Gypsy superfamily.
Silvana C. S. Cardoso, Douglas Silva Domingues, Alexandre Rossi Paschoal, Carlos Fischer, Ricardo Cerri
CIBCB5
2024 Ensemble Methods for Selecting Single Nucleotide Polymorphisms Associated to Rice Phenotypes
abstract
Rice (Oryza sativa) is one of the largest collections of genetic resources among plant species of economic interest. Several genetic variability studies have been developed to increase this cultivar’s productivity. In this context, Single Nucleotide Polymorphisms (SNPs), single base variations in DNA sequences, have been widely studied, as they act as molecular markers linked to productivity and resistance in rice cultivation. Due to the ineffectiveness of conventional methods in selecting SNPs, methods based on Machine Learning have been used. For this purpose, the selection of SNPs is modeled as a Feature Selection problem. Although feature selection is widespread in the literature, there are still gaps regarding its use in the context of rice genetic improvement. To advance interesting points regarding this discussion, we propose two ensemble methods for selecting important SNPs related to different phenotypes in rice, combining feature selection algorithms to generate a robust result. Our first proposal directly selects the most important SNPs using the phenotype numeric values. The second proposal discretizes the phenotype values, selecting the most important SNPs through classification algorithms. Experiments using real-world rice datasets showed that the proposed ensembles were better or very competitive compared to other methods from the literature regarding SNPs selected and prediction performance.
Breno O. Funicheli, Claudio Brondani, Rosana P. Vianello, Ricardo Cerri
IJCNN4
2024 Deep forests with tree-embeddings and label imputation for weak-label learning
abstract
Due to recent technological advances, a massive amount of data is generated on a daily basis. Unfortunately, this is not always beneficial as such data may present weak-supervision, meaning that the output space can be incomplete, inexact, and inaccurate. This kind of problems is investigated in weakly-supervised learning. In this work, we explore weak-label learning, a structured output prediction task for weakly-supervised problems where positive annotations are reliable, whereas negatives are missing. For the first time in this class of problems, we investigate deep forest algorithms based on tree-embeddings, a recently proposed feature representation strategy leveraging the structure of decision trees. Furthermore, we propose two new procedures for label-imputation in each layer, named Strict Label Complement (SLC), which provides fixed conservative estimates for the number of missing labels and employs them to restrict imputations, and Fluid Label Addition (FLA), which performs such estimations on every layer and uses them to adjust the imputer’s predicted probabilities without any restrictions. We combine the new approaches with deep forest architectures to produce four new algorithms: SLCForest and FLAForest, using output space feature augmentation, and also the cascade forest embedders CaFE-SLC and CaFE-FLA, employing both tree-embeddings and the output space. Our results reveal that our methods provide superior or competitive performance to the state-of-the-art. Furthermore, we also noticed that our methods are associated with better results even in cases without weak-supervision.
Pedro Ilídio, Ricardo Cerri, Celine Vens, Felipe Kenji Nakano
IJCNN2
2024 Better trees: an empirical study on hyperparameter tuning of classification decision tree induction algorithms
Rafael Gomes Mantovani, Tomás Horváth, André Luis Debiaso Rossi, Ricardo Cerri, Sylvio Barbon Junior, Joaquin Vanschoren, André C. P. L. F. de Carvalho
Data Min. Knowl. Discov.4
2023 A New Time Series Framework for Forest Fire Risk Forecasting and Classification
abstract
There's an increasing concern about the occurrence and spread of forest fires across the globe, as they contribute to greenhouse gas emissions and play a major influential role in economics and public health. Thus, there's a need for accurate methods to predict and classify forest fire risk. The main known forest fire risk indexes have limitations, such as not taking into account the unique characteristics of the biome in study, and not being able to predict forest fire risk for a given number of days in the future. This last aspect, in particular, is of utmost relevance. Addressing it allows for coordinated planning and action by proper authorities with adequate anticipation. Aiming to solve this problem, we present a new framework that applies Machine Learning methods for: (1) climatic variables forecasting; and (2) forest fire risk classification. For the first objective, different time series forecasting algorithms were tested. The forecasted variables are then used as input for the second objective, for which different classification algorithms were also tested. We evaluated our proposal using Brazilian Pantanal regional biome data from 1999 to 2019, where climatic variables were collected from ground meteorological stations, and fire occurrences (hotspots) were obtained from satellite images. The experiments considered 4 climatic variables and 5 forest fire risk classes. The results were evaluated based on the average correlation between (i) the prediction of forest fire risk classes and (ii) the observation of hotspots. Our proposal proved to be better or competitive with the main forest fire risk indexes, with the advantage of predicting fire risk for a given number of days in the future.
Bruna Zamith Santos, Balbina Maria Araujo Soriano, Marcelo Gonçalves Narciso, Diego Furtado Silva, Ricardo Cerri
IJCNN5
2023 Multi-label classification via closed frequent labelsets and label taxonomies
Mauri Ferrandin, Ricardo Cerri
Soft Comput.2
2022 An Algorithm Adaptation Method for Multi-Label Stream Classification using Self-Organizing Maps
abstract
Multi-label stream classification is the task of classifying instances in two or more classes simultaneously, with instances flowing continuously in high speed. This task imposes difficult challenges, such as the detection of concept drifts, where the distributions of the instances in the stream change with time, and infinitely delayed labels, when the ground truth labels of the instances are never available to help updating the classifiers. To solve such task, the methods from the literature use the problem transformation approach, which divides the multi-label problem into different sub-problems, associating one classification model for each class. In this paper, we propose a method based on self-organizing maps that, different from the literature, uses only one model to deal with all classes simultaneously. By using the algorithm adaptation approach, our proposal better considers label dependencies, improving the results over its counterparts. Experiments using different synthetic and real-world datasets showed that our proposal obtained the overall best performance when compared to different methods from the literature.
Ricardo Cerri, Elaine Ribeiro de Faria, João Gama 0001
ICMLA1
2022 A Two-step Model for Drug-Target Interaction Prediction with Predictive Bi-Clustering Trees and XGBoost
abstract
Interaction data are obtained by observing and recording interactions between objects. The use of interaction data makes it possible to solve several complex problems. Currently, there are several ways to use this data to produce solutions, one of which is the prediction of new drug-target interactions based on already known interactions. To perform this task, supervised machine learning methods can be used. Among these methods, we highlight Predictive Bi-Clustering Trees (PBCT), a global-based multi-label method which can simultaneously predict all interactions of an object. To use it, an interaction matrix is constructed based on the true bi-partite graph containing the interactions between objects. PBCT then induces a decision tree traversing the interaction matrix where leaf nodes correspond to partitions of the original matrix. The performance of PBCT, however, is harmed when the datasets are too imbalanced, generating leaf nodes with a much higher number of negative interactions. In this work, we propose a two-step approach for improving PBCT, where Predictive Bi-Clustering Trees are used to generate partitions in the interaction matrix, and the XGboost classifier is used to predict interactions based on these partitions. Our approach was applied to drug-target interaction prediction, showing improvements when compared with the original state-of-the-art PBCT. The prediction of drug-target interactions in silico brings economy and agility in the discovery of new interactions since it can be used to induce the experimental procedure in vitro.
André Hallwas Ribeiro Alves, Ricardo Cerri
IJCNN2
2021 Feature Selection for Hierarchical Multi-label Classification
Luan V. M. da Silva, Ricardo Cerri
IDA2
2021 Exploring Autoencoders for Feature Extraction in Multi-Target Classification
abstract
Multi-target learning is a prediction task where each example is associated with multiple target variables (outputs) simultaneously. One of the challenges in this research field is related to the high dimensionality of the data and the high number of target variables with dependencies. In such scenarios, it is crucial to extract lower dimensional representations from the original input space, such that these can be provided as input to other multi-target predictors. In this paper, we proposed using Autoencoders as feature extractors in several multi-target classification datasets publicly available. Results were evaluated considering state-of-the-art multi-target classification methods and evaluation measures in the literature. The experiments showed that the neural networks were able to keep the predictive performance even when the extracted features corresponded to a dimension size equivalent to 10% of the original number of features and, in some cases, getting better results than when using the original datasets.
Brendon Gouveia Cambuí, Rafael Gomes Mantovani, Ricardo Cerri
IJCNN3
2021 Exploring Label Correlations for Partitioning the Label Space in Multi-label Classification
abstract
Recent works on Multi-Label Classification (MLC) present multiple strategies to explore label correlations in a way to improve classifiers performances. However, these works focus only in the traditional local and global approaches, i.e., transforming the original problem into a set of binary local problems, or dealing globally with all classes simultaneously. Very few works have investigated strategies to use label correlations in order to partition the label space in a different ways. While in local partitions several binary classifiers are used (one per label), global partitions use only one classifier to deal with all labels. On the contrary, here we propose a strategy that explores the correlations between labels to partition the label space aiming to find partitions in-between (hybrid) the local and global ones. We believe in-between local and global partitions better cluster similar labels, improving the multi-label classifiers ability to explore label correlations. We compared the hybrid partitions with global, local and random generated partitions. Our experimental results showed that the hybrid partitions lead to competitive results and, in general, were slightly better than global and local partitions. The random partitions were also competitive with the global and local partitions, showing that the current local and global approaches still need improvements in order to consider label correlations.
Elaine Cecília Gatto, Mauri Ferrandin, Ricardo Cerri
IJCNN3
2021 Preventing the generation of inconsistent sets of crisp classification rules
Thiago Zafalon Miranda, Diorge Brognara Sardinha, Ricardo Cerri
Expert Syst. Appl.3
2021 Beyond global and local multi-target learning
Márcio P. Basgalupp, Ricardo Cerri, Leander Schietgat, Isaac Triguero, Celine Vens
Inf. Sci.2
2021 The experience of teaching introductory programming skills to bioscientists in Brazil
abstract
Computational biology has gained traction as an independent scientific discipline over the last years in South America. However, there is still a growing need for bioscientists, from different backgrounds, with different levels, to acquire programming skills, which could reduce the time from data to insights and bridge communication between life scientists and computer scientists. Python is a programming language extensively used in bioinformatics and data science, which is particularly suitable for beginners. Here, we describe the conception, organization, and implementation of the Brazilian Python Workshop for Biological Data. This workshop has been organized by graduate and undergraduate students and supported, mostly in administrative matters, by experienced faculty members since 2017. The workshop was conceived for teaching bioscientists, mainly students in Brazil, on how to program in a biological context. The goal of this article was to share our experience with the 2020 edition of the workshop in its virtual format due to the Coronavirus Disease 2019 (COVID-19) pandemic and to compare and contrast this year's experience with the previous in-person editions. We described a hands-on and live coding workshop model for teaching introductory Python programming. We also highlighted the adaptations made from in-person to online format in 2020, the participants' assessment of learning progression, and general workshop management. Lastly, we provided a summary and reflections from our personal experiences from the workshops of the last 4 years. Our takeaways included the benefits of the learning from learners' feedback (LLF) that allowed us to improve the workshop in real time, in the short, and likely in the long term. We concluded that the Brazilian Python Workshop for Biological Data is a highly effective workshop model for teaching a programming language that allows bioscientists to go beyond an initial exploration of programming skills for data analysis in the medium to long term.
Luíza Zuvanov, Ana Letycia Basso Garcia, Fernando Henrique Correr, Rodolfo Bizarria Jr., Ailton Pereira da Costa Filho, Alisson Hayasi da Costa, Andréa T. Thomaz, Ana Lucia Mendes Pinheiro, Diego Mauricio Riaño-Pachón, Flavia Vischi Winck, Franciele Grego Esteves, Gabriel Rodrigues Alves Margarido, Giovanna Maria Stanfoca Casagrande, Henrique Cordeiro Frajacomo, Leonardo Martins, Mariana Feitosa Cavalheiro, Nathalia Graf Grachet, Raniere Gaia Costa da Silva, Ricardo Cerri, Rommel Ramos, Simone Daniela Sartorio de Medeiros, Thayana Vieira Tavares, Renato Augusto Corrêa dos Santos
PLoS Comput. Biol.19
2020 Predictive Bi-clustering Trees for Hierarchical Multi-label Classification
Bruna Zamith Santos, Felipe Kenji Nakano, Ricardo Cerri, Celine Vens
ECML/PKDD (3)3
2020 Active learning for hierarchical multi-label classification
Felipe Kenji Nakano, Ricardo Cerri, Celine Vens
Data Min. Knowl. Discov.2
2019 A lexicographic genetic algorithm for hierarchical classification rule induction
abstract
Hierarchical Classification (HC) consists of assigning an instance to multiple classes simultaneously in a hierarchical structure containing dozens or even hundreds of classes. A field that greatly benefits from HC is Bioinformatics, in which interpretable methods that make predictions automatically are still scarce. In this context, a topic that has gained attention is the classification of Transposable Elements (TEs), which are DNA fragments capable of moving inside the genome of their hosts, affecting the genes' functionalities in many species. Thus, in this paper, we propose a method called Hierarchical Classification with a Lexicographic Genetic Algorithm (HC-LGA), which evolves rules towards the HC of TEs. Our proposed method follows a Multi-Objective Lexicographic approach in order to better deal with the still relevant problem of accuracy-interpretability trade-off. Besides, to the best of our knowledge, this is the first work to combine HC with such an approach. Experiments with two popular TEs datasets showed that HC-LGA achieved competitive results compared with most of the state-of-the-art HC methods in the literature, having the advantage of generating an interpretable model. Furthermore, HC-LGA obtained a comparable performance against its simpler optimization version, however generating a more interpretable list of rules.
Gean Trindade Pereira, Paulo Henrique Ribeiro Gabriel, Ricardo Cerri
GECCO3
2019 Pruned Sets for Multi-Label Stream Classification without True Labels
abstract
In multi-label classification problems an example can be simultaneously classified into more than one class. This is also a challenging task in Data Streams (DS) classification, where unbounded and non-stationary distributed multi-label data contain multiple concepts that drift at different rates and patterns. In addition, the true labels of the examples may never become available and updating classification models in a supervised fashion is unfeasible. In this paper, we propose a Multi-Label Stream Classification (MLSC) method applying a Novelty Detection (ND) procedure task to update the classification model detecting any new patterns in the examples, which differ in some aspects from observed patterns, in an unsupervised fashion without any external feedback. Although ND is suitable for multi-class stream classification, it is still a not well-investigated task for multi-label problems. We improve a initial work proposed in [1] and extended it with a new Pruned Sets (PS) transformation strategy. The experiments showed that our method presents competitive performances over data sets with different concept drifts, and outperform, in some aspects, the baseline methods.
Joel D. Costa Júnior, Elaine Ribeiro de Faria, Jonathan de Andrade Silva, João Gama 0001, Ricardo Cerri
IJCNN5
2018 A Genetic Algorithm for Transposable Elements Hierarchical Classification Rule Induction
abstract
Genomes of animals and plants are crowded with Transposable Elements (TEs), which are DNA sequences capable to move within the genome of a cell. They can modify the functionality of host genes, which makes them extremely important for the genetic variability of species. Therefore, their correct classification is crucial to understand their role in evolution. In this paper, the classification of TEs is treated as a Hierarchical Classification problem making use of Evolutionary Computation. Thus, hierarchical datasets suitable to be used by ML methods are presented, along with a novel hierarchical global rule induction classification strategy using a Genetic Algorithm. To the best of our knowledge, this is the first attempt in the literature to apply a hierarchical global-based rule learner to induce a model which classifies TEs according to a hierarchical taxonomy. We compared our global method with local and homology-based methods, and evaluated them using measures specific for hierarchical problems. The experimental results showed that our proposal achieved better or competitive results if compared to state-of-the-art methods from the literature, having the advantage of presenting an interpretable set of classification rules.
Gean Trindade Pereira, Bruna Zamith Santos, Ricardo Cerri
CEC3
2018 Hierarchical Multi-Label Classification Networks
abstract
One of the most challenging machine learning problems is a particular case of data classification in which classes are hierarchically structured and objects can be assigned to multiple paths of the class hierarchy at the same time. This task is known as hierarchical multi-label classification (HMC), with applications in text classification, image annotation, and in bioinformatics problems such as protein function prediction. In this paper, we propose novel neural network architectures for HMC called HMCN, capable of simultaneously optimizing local and global loss functions for discovering local hierarchical class-relationships and global information from the entire class hierarchy while penalizing hierarchical violations. We evaluate its performance in 21 datasets from four distinct domains, and we compare it against the current HMC state-of-the-art approaches. Results show that HMCN substantially outperforms all baselines with statistical significance, arising as the novel state-of-the-art for HMC.
Jonatas Wehrmann, Ricardo Cerri, Rodrigo C. Barros
ICML2
2018 Multi-label Feature Selection Techniques for Hierarchical Multi-label Protein Function Prediction
abstract
Protein Function Prediction is a complex Hierarchical Multi-label Classification task where the functional classes involved are organized in a hierarchy. While many Machine Learning methods have been proposed for this task, very few studies were performed for feature selection in such hierarchical scenarios. In this paper, we investigate feature selection techniques for hierarchical multi-label classification of protein functions. As decision trees are natural feature selectors, we rely on a hierarchical multi-label decision tree induction algorithm to extract features represented by the internal nodes of the tree. We also investigated the performance of a ReliefF-based non-hierarchical multi-label feature selection technique on the hierarchical scenario. We tested the different techniques on two classifiers, based on neural networks and genetic algorithms. The experimental results show that, in very few cases, the existing feature selection techniques were able to improve the classifiers performances, showing the need for developing feature selectors specifically to consider hierarchical class relationships.
Ricardo Cerri, Rafael Gomes Mantovani, Márcio P. Basgalupp, André C. P. L. F. de Carvalho
IJCNN1
2018 Improving Hierarchical Classification of Transposable Elements using Deep Neural Networks
abstract
Transposable Elements (TEs) are DNA sequences capable of moving within a cell's genome. Their transposition has many effects in genomes, such as creating genetic variability and promoting changes in genes' functionality. Recently, TEs classification has been addressed using Machine Learning (ML), more specifically by Hierarchical Classification (HC) methods. Such works proved to be superior than previous ones in the literature. However, there is still room for improvement performance wise. In this direction, Deep Neural Networks (DNNs) have attracted a lot of attention in ML. In particular, Stacked Denoising Auto-Encoders (DAEs) and Deep Multi Layer-Perceptrons (MLPs) are known to provide outstanding results. By performing an extensive evaluation, our results point out that DNNs can enhance the performance of HC methods, being able to push further the state-of-art in TEs' classification.
Felipe Kenji Nakano, Saulo Martiello Mastelini, Sylvio Barbon Junior, Ricardo Cerri
IJCNN4
2018 Applying multi-label techniques in emotion identification of short texts
Alex Marino Gonçalves de Almeida, Ricardo Cerri, Emerson Cabrera Paraiso, Rafael Gomes Mantovani, Sylvio Barbon Junior
Neurocomputing2
2018 A machine learning based framework to identify and classify long terminal repeat retrotransposons
abstract
Transposable elements (TEs) are repetitive nucleotide sequences that make up a large portion of eukaryotic genomes. They can move and duplicate within a genome, increasing genome size and contributing to genetic diversity within and across species. Accurate identification and classification of TEs present in a genome is an important step towards understanding their effects on genes and their role in genome evolution. We introduce TE-Learner, a framework based on machine learning that automatically identifies TEs in a given genome and assigns a classification to them. We present an implementation of our framework towards LTR retrotransposons, a particular type of TEs characterized by having long terminal repeats (LTRs) at their boundaries. We evaluate the predictive performance of our framework on the well-annotated genomes of Drosophila melanogaster and Arabidopsis thaliana and we compare our results for three LTR retrotransposon superfamilies with the results of three widely used methods for TE identification or classification: RepeatMasker, Censor and LtrDigest. In contrast to these methods, TE-Learner is the first to incorporate machine learning techniques, outperforming these methods in terms of predictive performance, while able to learn models and make predictions efficiently. Moreover, we show that our method was able to identify TEs that none of the above method could find, and we investigated TE-Learner's predictions which did not correspond to an official annotation. It turns out that many of these predictions are in fact strongly homologous to a known TE.
Leander Schietgat, Celine Vens, Ricardo Cerri, Carlos Fischer, Eduardo P. Costa, Jan Ramon, Claudia M. A. Carareto, Hendrik Blockeel
PLoS Comput. Biol.3
2017 Stacking Methods for Hierarchical Classification
abstract
Hierarchical Classification (HC) consists of classification problems whose classes are structured in a hierarchical fashion. Many problems are addressed by HC, in special a decent amount of works dealt with bioinformatics related problems such as Protein Function Prediction (PFP) and Transposable Elements (TEs) classification. Both of them are still a challenging task for HC due to the noisy and imbalanced nature of the datasets. As a countermeasure, Stacking is an ensemble method capable of generalizing knowledge from many classifiers. In this work, we propose three Stacking methods for HC and evaluate its performance on PFP and TEs datasets. Our results show that, when compared to regular Stacking and state-of-art methods from the literature, our methods are able to obtain superior or competitive performances.
Felipe Kenji Nakano, Saulo Martiello Mastelini, Sylvio Barbon Junior, Ricardo Cerri
ICMLA4
2017 Incorporating instance correlations in multi-label classification via label-space
abstract
Multi-label classification is a machine learning task where instances can be classified into two or more labels simultaneously. In this task, there exist correlations between the instances belonging to same or similar sets of labels. This paper proposes the incorporation of instance correlations by modifying the multi-label datasets. We used the label-space to create new features, which represent these correlations. The original and modified datasets were used with different multi-label classification methods. Experiments have shown that better results can be obtained when instance correlations were incorporated in the classification tasks. All methods were evaluated with measures specifically designed for multi-label problems.
Iuri Bonna M. de Abreu, Rafael Gomes Mantovani, Ricardo Cerri
IJCNN3
2017 A self-organizing map-based method for multi-label classification
abstract
In Machine Learning, multi-label classification is the task of assigning an instance to two or more categories simultaneously. This is a very challenging task, since datasets can have many instances and become very unbalanced. While most of the methods in the literature use supervised learning to solve multilabel problems, in this paper we propose the use of unsupervised learning through neural networks. More specifically, we explore the power of Self-Organizing Maps (Kohonen Maps), since they have a self-organization ability and maps input instances to a map of neurons. Because instances that are assigned to similar groups of labels tend to be more similar, there is a network tendency that, after organization, training instances which are similar to each other are mapped to closer neurons in the map. Testing instances can then be mapped to specific neurons in the network, being classified in the labels assigned to training instances mapped to these neurons. Our proposal was experimentally compared to other literature methods, showing competitive performances. The evaluation was performed using freely available datasets and measures specifically designed for multi-label problems.
Gustavo G. Colombini, Iuri Bonna M. de Abreu, Ricardo Cerri
IJCNN3
2017 Top-down strategies for hierarchical classification of transposable elements with neural networks
abstract
Transposable Elements are DNA sequences that can move from one place to another inside the genome of a cell. They are important for genetic variability, and can modify the functionality of genes. The correct classification of these elements is crucial to understand their role in the evolution of species. In this paper, we investigate Transposable Elements classification as a Hierarchical Classification problem using Machine Learning. We present new hierarchical datasets suitable to be used by Machine Learning methods, and also new hierarchical top-down classification strategies using neural networks. We compared our strategies with existing ones in the literature, and evaluated them using measures specific for hierarchical problems. Experiments showed that our proposal achieved better or competitive results than those found by other methods in the literature.
Felipe Kenji Nakano, Walter José G. S. Pinto, Gisele L. Pappa, Ricardo Cerri
IJCNN4
2016 Reduction strategies for hierarchical multi-label classification in protein function prediction
abstract
BACKGROUND: Hierarchical Multi-Label Classification is a classification task where the classes to be predicted are hierarchically organized. Each instance can be assigned to classes belonging to more than one path in the hierarchy. This scenario is typically found in protein function prediction, considering that each protein may perform many functions, which can be further specialized into sub-functions. We present a new hierarchical multi-label classification method based on multiple neural networks for the task of protein function prediction. A set of neural networks are incrementally training, each being responsible for the prediction of the classes belonging to a given level. RESULTS: The method proposed here is an extension of our previous work. Here we use the neural network output of a level to complement the feature vectors used as input to train the neural network in the next level. We experimentally compare this novel method with several other reduction strategies, showing that it obtains the best predictive performance. Empirical results also show that the proposed method achieves better or comparable predictive performance when compared with state-of-the-art methods for hierarchical multi-label classification in the context of protein function prediction. CONCLUSIONS: The experiments showed that using the output in one level as input to the next level contributed to better classification results. We believe the method was able to learn the relationships between the protein functions during training, and this information was useful for classification. We also identified in which functional classes our method performed better.
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho, Yaochu Jin
BMC Bioinform.1
2015 Hierarchical classification of Gene Ontology-based protein functions with neural networks
abstract
Hierarchical Multi-label Classification (HMC) is a classification task where classes are organized in a hierarchical taxonomy, and instances can be simultaneously classified in more than one class. This paper investigates the HMC problem of classifying proteins in functions organized according to the Gene Ontology hierarchical taxonomy. This is a complex task, since the Gene Ontology hierarchy is organized as a Directed Acyclic Graph with thousands of classes hierarchically represented. We propose a neural network-based method to incorporate label-dependency during learning. The experimental results show that the proposed method achieves competitive results when compared to the state-of-the-art methods from the literature.
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho
IJCNN1
2015 Learning HMMs for nucleotide sequences from amino acid alignments
abstract
Profile hidden Markov models (profile HMMs) are known to efficiently predict whether an amino acid (AA) sequence belongs to a specific protein family. Profile HMMs can also be used to search for protein domains in genome sequences. In this case, HMMs are typically learned from AA sequences and then used to search on the six-frame translation of nucleotide (NT) sequences. However, this approach demands additional processing of the original data and search results. Here, we propose an alternative and more direct method which converts an AA alignment into an NT one, after which an NT-based HMM is trained to be applied directly on a genome.
Carlos Fischer, Claudia M. A. Carareto, Renato Augusto Corrêa dos Santos, Ricardo Cerri, Eduardo P. Costa, Leander Schietgat, Celine Vens
Bioinform.4
2015 An Extensive Evaluation of Decision Tree-Based Hierarchical Multilabel Classification Methods and Performance Measures
abstract
Hierarchical multilabel classification is a complex classification problem where an instance can be assigned to more than one class simultaneously, and these classes are hierarchically organized with superclasses and subclasses, that is, an instance can be classified as belonging to more than one path in the hierarchical structure. This article experimentally analyses the behavior of different decision tree–based hierarchical multilabel classification methods based on the local and global classification approaches. The approaches are compared using distinct hierarchy‐based and distance‐based evaluation measures, when they are applied to a variation of real multilabel and hierarchical datasets' characteristics. Also, the different evaluation measures investigated are compared according to their degrees of consistency, discriminancy, and indifferency. As a result of the experimental analysis, we recommend the use of the global classification approach and suggest the use of the Hierarchical Precision and Hierarchical Recall evaluation measures.
Ricardo Cerri, Gisele L. Pappa, André C. P. L. F. de Carvalho, Alex Alves Freitas
Comput. Intell.1
2014 A framework for bottom-up induction of oblique decision trees
Rodrigo C. Barros, Pablo A. Jaskowiak, Ricardo Cerri, André C. P. L. F. de Carvalho
Neurocomputing3
2014 Hierarchical multi-label classification using local neural networks
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho
J. Comput. Syst. Sci.1
2013 A grammatical evolution algorithm for generation of Hierarchical Multi-Label Classification rules
abstract
Hierarchical Multi-Label Classification (HMC) is a challenging task in data mining and machine learning. Each instance in HMC can be classified into two or more classes simultaneously. These classes are structured in a hierarchy, in the form of either a tree or a directed acyclic graph. Therefore, an instance can be assigned to two or more paths from the hierarchical structure, resulting in a complex classification problem with hundreds or thousands of classes. Several methods have been proposed to deal with such problems, including several algorithms based on well-known bio-inspired techniques, such as neural networks, ant colony optimization, and genetic algorithms. In this work, we propose a novel global method called GEHM, which makes use of grammatical evolution for generating HMC rules. In this approach, the grammatical evolution algorithm evolves the antecedents of classification rules, in order to assign instances from a HMC dataset to a probabilistic class vector. Our method is compared to bio-inspired HMC algorithms in protein function prediction datasets. The empirical analysis conducted in this work shows that GEHM outperforms the bio-inspired algorithms with statistical significance, which suggests that grammatical evolution is a promising alternative to deal with hierarchical multi-label classification of biological data.
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho, Alex Alves Freitas
IEEE Congress on Evolutionary Computation1
2013 A grammatical evolution approach for software effort estimation
abstract
Software effort estimation is an important task within software engineering. It is widely used for planning and monitoring software project development as a means to deliver the product on time and within budget. Several approaches for generating predictive models from collected metrics have been proposed throughout the years. Machine learning algorithms, in particular, have been widely-employed to this task, bearing in mind their capability of providing accurate predictive models for the analysis of project stakeholders. In this paper, we propose a grammatical evolution approach for software metrics estimation. Our novel algorithm, namely SEEGE, is empirically evaluated on public project data sets, and we compare its performance with state-of-the-art machine learning algorithms such as support vector machines for regression and artificial neural networks, and also to popular linear regression. Results show that SEEGE outperforms the other algorithms considering three different evaluation measures, clearly indicating its effectiveness for the effort estimation task.
Rodrigo C. Barros, Márcio P. Basgalupp, Ricardo Cerri, Tiago Silva da Silva, André C. P. L. F. de Carvalho
GECCO3
2013 Probabilistic Clustering for Hierarchical Multi-Label Classification of Protein Functions
Rodrigo C. Barros, Ricardo Cerri, Alex Alves Freitas, André C. P. L. F. de Carvalho
ECML/PKDD (2)2
2011 A bottom-up oblique decision tree induction algorithm
abstract
Decision tree induction algorithms are widely used in knowledge discovery and data mining, specially in scenarios where model comprehensibility is desired. A variation of the traditional univariate approach is the so-called oblique decision tree, which allows multivariate tests in its non-terminal nodes. Oblique decision trees can model decision boundaries that are oblique to the attribute axes, whereas univariate trees can only perform axis-parallel splits. The majority of the oblique and univariate decision tree induction algorithms perform a top-down strategy for growing the tree, relying on an impurity-based measure for splitting nodes. In this paper, we propose a novel bottom-up algorithm for inducing oblique trees named BUTIA. It does not require an impurity-measure for dividing nodes, since we know a priori the data resulting from each split. For generating the splitting hyperplanes, our algorithm implements a support vector machine solution, and a clustering algorithm is used for generating the initial leaves. We compare BUTIA to traditional univariate and oblique decision tree algorithms, C4.5, CART, OC1 and FT, as well as to a standard SVM implementation, using real gene expression benchmark data. Experimental results show the effectiveness of the proposed approach in several cases.
Rodrigo C. Barros, Ricardo Cerri, Pablo A. Jaskowiak, André C. P. L. F. de Carvalho
ISDA2
2011 Hierarchical multi-label classification for protein function prediction: A local approach based on neural networks
abstract
In Hierarchical Multi-Label Classification problems, each instance can be classified into two or more classes simultaneously, differently from conventional classification. Additionally, the classes are structured in a hierarchy, in the form of either a tree or a directed acyclic graph. Hence, an instance can be assigned to two or more paths from the hierarchical structure, resulting in a complex classification problem with possibly hundreds of classes. Many methods have been proposed to deal with such problems, some of them employing a single classifier to deal with all classes simultaneously (global methods), and others employing many classifiers to decompose the original problem into a set of subproblems (local methods). In this work, we propose a novel local method named HMC-LMLP, which uses one Multi-Layer Perceptron per hierarchical level. The predictions in one level are used as inputs to the network responsible for the predictions in the next level. We make use of two distinct Multi-Layer Perceptron algorithms: Back-propagation and Resilient Back-propagation. In addition, we make use of an error measure specially tailored to multi-label problems for training the networks. Our method is compared to state-of-the-art hierarchical multi-label classification algorithms, in protein function prediction datasets. The experimental results show that our approach presents competitive predictive accuracy, suggesting that artificial neural networks constitute a promising alternative to deal with hierarchical multi-label classification of biological data.
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho
ISDA1
2011 Adapting non-hierarchical multilabel classification methods for hierarchical multilabel classification
abstract
In most classification problems, a classifier assigns a single class to each instance and the classes form a flat (non-hierarchical) structure, without superclasses or subclasses. In hierarchical multilabel classification problems, the classes are hi
Ricardo Cerri, André C. P. L. F. de Carvalho, Alex Alves Freitas
Intell. Data Anal.1
2010 New top-down methods using SVMs for Hierarchical Multilabel Classification problems
abstract
Hierarchical Multilabel Classification is a problem where the classes involved are hierarchically structured, and examples can be assigned to more than one class simultaneously at a same hierarchical level. This paper describes and evaluates five different methods for this classification task, based on two approaches, named top-down and one-shot. In the top-down approach, the classification task is carried out by discriminating the classes, level by level, in the hierarchy. In the one-shot approach, the methods consider the whole set of classes at once in the classification. Based on the top-down approach, two new hierarchical methods (with label combination and with label decomposition) and the well-known binary hierarchical method are investigated using SVM classifiers. Other two methods from the literature, named HC4.5 and Clus-HMC, based on the one-shot approach, are also used. The methods are applied to ten biological datasets and evaluated using specific metrics for this kind of classification. The experimental results showed that the proposed methods can improve the classification accuracy.
Ricardo Cerri, André C. P. L. F. de Carvalho
IJCNN1