Alex Alves Freitas

dblp:f/AlexAlvesFreitas · DBLP profile ↗
← Back
115ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-9825-4700ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 22 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Automated machine learning for positive-unlabelled learning
abstract
Abstract Positive-Unlabelled (PU) learning is a field of machine learning that aims to learn classifiers from data consisting of labelled positive and unlabelled instances, which can be in reality positive or negative, but whose label is unknown. Many PU learning methods have been proposed over the last two decades, so many so that selecting an optimal method for a given PU learning task presents a challenge. Our previous work has addressed this by proposing GA-Auto-PU, the first Automated Machine Learning (Auto-ML) system for PU learning. In this work, we propose two new PU learning Auto-ML systems: BO-Auto-PU, based on a Bayesian Optimisation (BO) approach, and EBO-Auto-PU, based on a novel evolutionary/BO approach. We present an extensive evaluation of the three Auto-ML systems, comparing them to each other and to well-established PU learning methods across 60 datasets (20 datasets, each with 3 versions). The results of the comparison show statistically significant improvements in predictive accuracy over the baseline methods, as well as large improvements in computational time for the newly proposed Auto-PU systems over the original Auto-PU system.
Jack D. Saunders, Alex Alves Freitas
Appl. Intell.2
2024 Auto-Sklong: A New AutoML System for Longitudinal Classification
abstract
Automated Machine Learning (AutoML) addresses the challenge of selecting the best machine learning algorithm and its hyperparameter settings for a given dataset. However, existing AutoML systems typically focus on standard classification tasks and cannot directly exploit temporal information e.g. in longitudinal datasets, which contain multiple measurements of the same features over time — a common scenario in biomedical applications. We introduce Auto-Sklong, the first AutoML system that includes longitudinal classification algorithms in its search space. Experiments with 20 age-related disease datasets from the English Longitudinal Study of Ageing demonstrate that Auto-Sklong significantly outperforms a state-of-the-art AutoML system (Auto-Sklearn) and two baseline random forest methods in terms of predictive accuracy.
Simon Provost, Alex Alves Freitas
BIBM2
2023 Interpretable Ensembles of Classifiers for Uncertain Data With Bioinformatics Applications
abstract
Data uncertainty remains a challenging issue in many applications, but few classification algorithms can effectively cope with it. An ensemble approach for uncertain categorical features has recently been proposed, achieving promising results. It consists in biasing the sampling of features for each model in an ensemble so that less uncertain features are more likely to be sampled. Here we extend this idea of biased sampling and propose two new approaches: one for selecting training instances for each model in an ensemble and another for sampling features to be considered when splitting a node in a Random Forest training. We applied these approaches to classify ageing-related genes and predict drugs' side effects based on uncertain features representing protein-protein and protein-chemical interactions. We show that ensembles based on our proposed approaches achieve better predictive performance. In particular, our proposed approaches improved the performance of a Random Forest based on the most sophisticated approach for handling uncertain data in ensembles of this kind. Furthermore, we propose two new approaches for interpreting an ensemble of Naive Bayes classifiers and analyse their results on our datasets of ageing-related genes and drug's side effects.
Marcelo Rodrigues de Holanda Maia, Alexandre Plastino 0001, Alex Alves Freitas, João Pedro de Magalhães
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Fair Feature Selection with a Lexicographic Multi-objective Genetic Algorithm
James Brookhouse, Alex Alves Freitas
PPSN (2)2
2022 Machine learning-based predictions of dietary restriction associations across ageing-related genes
abstract
BACKGROUND: Dietary restriction (DR) is the most studied pro-longevity intervention; however, a complete understanding of its underlying mechanisms remains elusive, and new research directions may emerge from the identification of novel DR-related genes and DR-related genetic features. RESULTS: This work used a Machine Learning (ML) approach to classify ageing-related genes as DR-related or NotDR-related using 9 different types of predictive features: PathDIP pathways, two types of features based on KEGG pathways, two types of Protein-Protein Interactions (PPI) features, Gene Ontology (GO) terms, Genotype Tissue Expression (GTEx) expression features, GeneFriends co-expression features and protein sequence descriptors. Our findings suggested that features biased towards curated knowledge (i.e. GO terms and biological pathways), had the greatest predictive power, while unbiased features (mainly gene expression and co-expression data) have the least predictive power. Moreover, a combination of all the feature types diminished the predictive power compared to predictions based on curated knowledge. Feature importance analysis on the two most predictive classifiers mostly corroborated existing knowledge and supported recent findings linking DR to the Nuclear Factor Erythroid 2-Related Factor 2 (NRF2) signalling pathway and G protein-coupled receptors (GPCR). We then used the two strongest combinations of feature type and ML algorithm to predict DR-relatedness among ageing-related genes currently lacking DR-related annotations in the data, resulting in a set of promising candidate DR-related genes (GOT2, GOT1, TSC1, CTH, GCLM, IRS2 and SESN2) whose predicted DR-relatedness remain to be validated in future wet-lab experiments. CONCLUSIONS: This work demonstrated the strong potential of ML-based techniques to identify DR-associated features as our findings are consistent with literature and recent discoveries. Although the inference of new DR-related mechanistic findings based solely on GO terms and biological pathways was limited due to their knowledge-driven nature, the predictive power of these two features types remained useful as it allowed inferring new promising candidate DR-related genes.
Gustavo Daniel Vega Magdaleno, Vladislav Bespalov, Yalin Zheng, Alex Alves Freitas, João Pedro de Magalhães
BMC Bioinform.4
2021 An Ensemble of Naive Bayes Classifiers for Uncertain Categorical Data
abstract
Coping with uncertainty is a very challenging issue in many real-world applications. However, conventional classification models usually assume there is no uncertainty in data at all. In order to fill this gap, there has been a growing number of studies addressing the problem of classification based on uncertain data. Although some methods resort to ignoring uncertainty or artificially removing it from data, it has been shown that predictive performance can be improved by actually incorporating information on uncertainty into classification models. This paper proposes an approach for building an ensemble of classifiers for uncertain categorical data based on biased random subspaces. Using Naive Bayes classifiers as base models, we have applied this approach to classify ageing-related genes based on real data, with uncertain features representing protein-protein interactions. Our experimental results show that models based on the proposed approach achieve better predictive performance than single Naive Bayes classifiers and conventional ensembles.
Marcelo Rodrigues de Holanda Maia, Alexandre Plastino 0001, Alex Alves Freitas
ICDM3
2021 A Novel Feature Selection Method for Uncertain Features: An Application to the Prediction of Pro-/Anti-Longevity Genes
abstract
Understanding the ageing process is a very challenging problem for biologists. To help in this task, there has been a growing use of classification methods (from machine learning) to learn models that predict whether a gene influences the process of ageing or promotes longevity. One type of predictive feature often used for learning such classification models is Protein-Protein Interaction (PPI) features. One important property of PPI features is their uncertainty, i.e., a given feature (PPI annotation) is often associated with a confidence score, which is usually ignored by conventional classification methods. Hence, we propose the Lazy Feature Selection for Uncertain Features (LFSUF) method, which is tailored for coping with the uncertainty in PPI confidence scores. In addition, following the lazy learning paradigm, LFSUF selects features for each instance to be classified, making the feature selection process more flexible. We show that our LFSUF method achieves better predictive accuracy when compared to other feature selection methods that either do not explicitly take PPI confidence scores into account or deal with uncertainty globally rather than using a per-instance approach. Also, we interpret the results of the classification process using the features selected by LFSUF, showing that the number of selected features is significantly reduced, assisting the interpretability of the results. The datasets used in the experiments and the program code of the LFSUF method are freely available on the web at http://github.com/pablonsilva/FSforUncertainFeatureSpaces.
Pablo Nascimento da Silva, Alexandre Plastino 0001, Fabio Fabris, Alex Alves Freitas
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 Adapting Random Forests to Cope with Heavily Censored Datasets in Survival Analysis
Tossapol Pomsuwan, Alex Alves Freitas
ESANN2
2020 A robust experimental evaluation of automated multi-label classification methods
abstract
Automated Machine Learning (AutoML) has emerged to deal with the selection and configuration of algorithms for a given learning task. With the progression of AutoML, several effective methods were introduced, especially for traditional classification and regression problems. Apart from the AutoML success, several issues remain open. One issue, in particular, is the lack of ability of AutoML methods to deal with different types of data. Based on this scenario, this paper approaches AutoML for multi-label classification (MLC) problems. In MLC, each example can be simultaneously associated to several class labels, unlike the standard classification task, where an example is associated to just one class label. In this work, we provide a general comparison of five automated multi-label classification methods - two evolutionary methods, one Bayesian optimization method, one random search and one greedy search - on 14 datasets and three designed search spaces. Overall, we observe that the most prominent method is the one based on a canonical grammar-based genetic programming (GGP) search method, namely Auto-MEKAGGP. Auto-MEKAGGP presented the best average results in our comparison and was statistically better than all the other methods in different search spaces and evaluated measures, except when compared to the greedy search method.
Alex Guimarães Cardoso de Sá, Cristiano Guimarães Pimenta, Gisele L. Pappa, Alex Alves Freitas
GECCO4
2020 Prioritizing positive feature values: a new hierarchical feature selection method
Pablo Nascimento da Silva, Alexandre Plastino 0001, Alex Alves Freitas
Appl. Intell.3
2020 Comparing enrichment analysis and machine learning for identifying gene properties that discriminate between gene classes
abstract
Biologists very often use enrichment methods based on statistical hypothesis tests to identify gene properties that are significantly over-represented in a given set of genes of interest, by comparison with a 'background' set of genes. These enrichment methods, although based on rigorous statistical foundations, are not always the best single option to identify patterns in biological data. In many cases, one can also use classification algorithms from the machine-learning field. Unlike enrichment methods, classification algorithms are designed to maximize measures of predictive performance and are capable of analysing combinations of gene properties, instead of one property at a time. In practice, however, the majority of studies use either enrichment or classification methods (rather than both), and there is a lack of literature discussing the pros and cons of both types of method. The goal of this paper is to compare and contrast enrichment and classification methods, offering two contributions. First, we discuss the (to some extent complementary) advantages and disadvantages of both types of methods for identifying gene properties that discriminate between gene classes. Second, we provide a set of high-level recommendations for using enrichment and classification methods. Overall, by highlighting the strengths and the weaknesses of both types of methods we argue that both should be used in bioinformatics analyses.
Fabio Fabris, Daniel Palmer, João Pedro de Magalhães, Alex Alves Freitas
Briefings Bioinform.4
2020 Investigating the role of Simpson's paradox in the analysis of top-ranked features in high-dimensional bioinformatics datasets
abstract
An important problem in bioinformatics consists of identifying the most important features (or predictors), among a large number of features in a given classification dataset. This problem is often addressed by using a machine learning-based feature ranking method to identify a small set of top-ranked predictors (i.e. the most relevant features for classification). The large number of studies in this area has, however, an important limitation: they ignore the possibility that the top-ranked predictors occur in an instance of Simpson's paradox, where the positive or negative association between a predictor and a class variable reverses sign upon conditional on each of the values of a third (confounder) variable. In this work, we review and investigate the role of Simpson's paradox in the analysis of top-ranked predictors in high-dimensional bioinformatics datasets, in order to avoid the potential danger of misinterpreting an association between a predictor and the class variable. We perform computational experiments using four well-known feature ranking methods from the machine learning field and five high-dimensional datasets of ageing-related genes, where the predictors are Gene Ontology terms. The results show that occurrences of Simpson's paradox involving top-ranked predictors are much more common for one of the feature ranking methods.
Alex Alves Freitas
Briefings Bioinform.1
2020 Using deep learning to associate human genes with age-related diseases
abstract
MOTIVATION: One way to identify genes possibly associated with ageing is to build a classification model (from the machine learning field) capable of classifying genes as associated with multiple age-related diseases. To build this model, we use a pre-compiled list of human genes associated with age-related diseases and apply a novel Deep Neural Network (DNN) method to find associations between gene descriptors (e.g. Gene Ontology terms, protein-protein interaction data and biological pathway information) and age-related diseases. RESULTS: The novelty of our new DNN method is its modular architecture, which has the capability of combining several sources of biological data to predict which ageing-related diseases a gene is associated with (if any). Our DNN method achieves better predictive performance than standard DNN approaches, a Gradient Boosted Tree classifier (a strong baseline method) and a Logistic Regression classifier. Given the DNN model produced by our method, we use two approaches to identify human genes that are not known to be associated with age-related diseases according to our dataset. First, we investigate genes that are close to other disease-associated genes in a complex multi-dimensional feature space learned by the DNN algorithm. Second, using the class label probabilities output by our DNN approach, we identify genes with a high probability of being associated with age-related diseases according to the model. We provide evidence of these putative associations retrieved from the DNN model with literature support. AVAILABILITY AND IMPLEMENTATION: The source code and datasets can be found at: https://github.com/fabiofabris/Bioinfo2019. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Daniel Palmer, Khalid M. Salama, João Pedro de Magalhães, Alex Alves Freitas
Bioinform.5
2020 An evolutionary algorithm for automated machine learning focusing on classifier ensembles: An improved algorithm and extended results
João Carlos Xavier Jr., Alex Alves Freitas, Teresa Bernarda Ludermir, Antonino Feitosa Neto, Cephas A. S. Barreto
Theor. Comput. Sci.2
2019 Automated Machine Learning for Studying the Trade-Off Between Predictive Accuracy and Interpretability
Alex Alves Freitas
CD-MAKE1
2018 A Survey of Genetic Algorithms for Multi-Label Classification
abstract
In recent years, multi-label classification (MLC) has become an emerging research topic in big data analytics and machine learning. In this problem, each object of a dataset may belong to multiple class labels and the goal is to learn a classification model that can infer the correct labels of new, previously unseen, objects. This paper presents a survey of genetic algorithms (GAs) designed for MLC tasks. The study is organized in three parts. First, we propose a new taxonomy focused on GAs for MLC. In the second part, we provide an up-to-date overview of the work in this area, categorizing the approaches identified in the literature with respect to the taxonomy. In the third and last part, we discuss some new ideas for combining GAs with MLC.
Eduardo Corrêa Gonçalves, Alex Alves Freitas, Alexandre Plastino 0001
CEC2
2018 Automated Selection and Configuration of Multi-Label Classification Algorithms with Grammar-Based Genetic Programming
Alex Guimarães Cardoso de Sá, Alex Alves Freitas, Gisele L. Pappa
PPSN (2)2
2018 A Novel Genetic Algorithm for Feature Selection in Hierarchical Feature Spaces
abstract
Feature selection methods have been widely adopted to prepare high-dimensional feature spaces for the classification task of data mining. However, in many real-world datasets, the feature space is formed by binary features related via generalization-specialization relationships, also known as hierarchical feature spaces. Although there are many methods for the traditional feature selection problem, methods which properly consider hierarchical features are still very underexplored. In this work, we propose a novel genetic algorithm (GA) for hierarchical feature selection. The proposed GA has two novel hierarchical mutation operators tailored to deal with redundant features in hierarchical feature spaces. The computational experiments show that our proposed approach exhibited better predictive performance than two state-of-the-art hierarchical feature selection methods (SHSEL and HIP) and also than two traditional feature selection methods (ReliefF and CFS).
Pablo Nascimento da Silva, Alexandre Plastino 0001, Alex Alves Freitas
SDM3
2018 A new approach for interpreting Random Forest models and its application to the biology of ageing
abstract
Motivation: This work uses the Random Forest (RF) classification algorithm to predict if a gene is over-expressed, under-expressed or has no change in expression with age in the brain. RFs have high predictive power, and RF models can be interpreted using a feature (variable) importance measure. However, current feature importance measures evaluate a feature as a whole (all feature values). We show that, for a popular type of biological data (Gene Ontology-based), usually only one value of a feature is particularly important for classification and the interpretation of the RF model. Hence, we propose a new algorithm for identifying the most important and most informative feature values in an RF model. Results: The new feature importance measure identified highly relevant Gene Ontology terms for the aforementioned gene classification task, producing a feature ranking that is much more informative to biologists than an alternative, state-of-the-art feature importance measure. Availability and implementation: The dataset and source codes used in this paper are available as 'Supplementary Material' and the description of the data can be found at: https://fabiofabris.github.io/bioinfo2018/web/. Supplementary information: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Aoife Doherty, Daniel Palmer, João Pedro de Magalhães, Alex Alves Freitas
Bioinform.5
2018 Multi-objective genetic algorithms in the study of the genetic code's adaptability
abstract
Using a robustness measure based on values of the polar requirement of amino acids, Freeland and Hurst (1998) showed that less than one in one million random hypothetical codes are better than the standard genetic code. In this paper, instead of comparing the standard code with randomly generated codes, we use an optimisation algorithm to find the best hypothetical codes. This approach has been used before, but considering only one objective to be optimised. The robustness measure based on the polar requirement is considered the most effective objective to be optimised by the algorithm. We propose here that the polar requirement is not the only property to be considered when computing the robustness of the genetic code. We include the hydropathy index and molecular volume in the evaluation of the amino acids using three multi-objective approaches: the weighted formula, lexicographic and Pareto approaches. To our knowledge, this is the first work proposing multi-objective optimisation approaches with a non-restrictive encoding for studying the evolution of the genetic code. Our results indicate that multi-objective approaches considering the three amino acid properties obtain better results than those obtained by single objective approaches reported in the literature. The codes obtained by the multi-objective approach are more robust and structurally more similar to the standard code.
Lariza Laura de Oliveira, Alex Alves Freitas, Renato Tinós
Inf. Sci.2
2017 Pricing Rainfall Based Futures Using Genetic Programming
Sam Cramer, Michael Kampouridis, Alex Alves Freitas, Antonis Alexandridis 0002
EvoApplications (1)3
2017 An extensive evaluation of seven machine learning methods for rainfall prediction in weather derivatives
Sam Cramer, Michael Kampouridis, Alex Alves Freitas, Antonis Alexandridis 0002
Expert Syst. Appl.3
2017 Instance-based classification with Ant Colony Optimization
abstract
Instance-based learning (IBL) methods predict the class label of a new instance based directly on the distance between the new unlabeled instance and each labeled instance in the training set, without constructing a classification model in the training phase. In this paper, we introduce a novel cla ss-based feature weighting technique, in the context of instance-based distance methods, using the Ant Colony Optimization meta-heuristic. We address three different approaches of instance-based classification: k-Nearest Neighbours, distance-based Nearest Neighbours, and Gaussian Kernel Estimator. We present a multi-archive adaptation of the ACOℝ algorithm and apply it to the optimization of the key parameter in each IBL algorithm and of the class-based feature weights. We also propose an ensemble of classifiers approach that makes use of the archived populations of the ACOℝ algorithm. We empirically evaluate the performance of our proposed algorithms on 36 benchmark datasets, and compare them with conventional instance-based classification algorithms, using various parameter settings, as well as with a state-of-the-art coevolutionary algorithm for instance selection and feature weighting for Nearest Neighbours classifiers.
Khalid M. Salama, Ashraf M. Abdelbar, Ayah Helal, Alex Alves Freitas
Intell. Data Anal.4
2016 Feature engineering for improving financial derivatives-based rainfall prediction
abstract
Rainfall is one of the most challenging variables to predict, as it exhibits very unique characteristics that do not exist in other time series data. Moreover, rainfall is a major component and is essential for applications that surround water resource planning. In particular, this paper is interested in extending previous work carried out on the prediction of rainfall using Genetic Programming (GP) for rainfall derivatives. Currently in the rainfall derivatives literature, the process of predicting rainfall is dominated by statistical models, namely using a Markov-chain extended with rainfall prediction (MCRP). In this paper we further extend our new methodology by looking at the effect of feature engineering on the rainfall prediction process. Feature engineering will allow us to extract additional information from the data variables created. By incorporating feature engineering techniques we look to further tailor our GP to the problem domain and we compare the performance of the previous GP, which previously statistically outperformed MCRP, against our new GP using feature engineering on 21 different data sets of cities across Europe and report the results. The goal is to see whether GP can outperform its predecessor without extra features, which acts as a benchmark. Results indicate that in general GP using extra features significantly outperforms a GP without the use of extra features.
Sam Cramer, Michael Kampouridis, Alex Alves Freitas
CEC3
2016 A Genetic Decomposition Algorithm for Predicting Rainfall within Financial Weather Derivatives
abstract
Regression problems provide some of the most challenging research opportunities, where the predictions of such domains are critical to a specific application. Problem domains that exhibit large variability and are of chaotic nature are the most challenging to predict. Rainfall being a prime example, as it exhibits very unique characteristics that do not exist in other time series data. Moreover, rainfall is essential for applications that surround financial securities such as rainfall derivatives. This paper is interested in creating a new methodology for increasing the predictive accuracy of rainfall within the problem domain of rainfall derivatives. Currently, the process of predicting rainfall within rainfall derivatives is dominated by statistical models, namely Markov-chain extended with rainfall prediction (MCRP). In this paper, we propose a novel algorithm for decomposing rainfall, which is a hybrid Genetic Programming/Genetic Algorithm (GP/GA) algorithm. Hence, the overall problem becomes easier to solve. We compare the performance of our hybrid GP/GA, against MCRP, Radial Basis Function and GP without decomposition. We aim to show the effectiveness that a decomposition algorithm can have on the problem domain. Results show that in general decomposition has a very positive effect by statistically outperforming GP without decomposition and MCRP.
Sam Cramer, Michael Kampouridis, Alex Alves Freitas
GECCO3
2016 New KEGG pathway-based interpretable features for classifying ageing-related mouse proteins
abstract
MOTIVATION: The incidence of ageing-related diseases has been constantly increasing in the last decades, raising the need for creating effective methods to analyze ageing-related protein data. These methods should have high predictive accuracy and be easily interpretable by ageing experts. To enable this, one needs interpretable classification models (supervised machine learning) and features with rich biological meaning. In this paper we propose two interpretable feature types based on Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways and compare them with traditional feature types in hierarchical classification (a more challenging classification task regarding predictive performance) and binary classification (a classification task producing easier to interpret classification models). As far as we know, this work is the first to: (i) explore the potential of the KEGG pathway data in the hierarchical classification setting, (i) use the graph structure of KEGG pathways to create a feature type that quantifies the influence of a current protein on another specific protein within a KEGG pathway graph and (iii) propose a method for interpreting the classification models induced using KEGG features. RESULTS: We performed tests measuring predictive accuracy considering hierarchical and binary class labels extracted from the Mouse Phenotype Ontology. One of the KEGG feature types leads to the highest predictive accuracy among five individual feature types across three hierarchical classification algorithms. Additionally, the combination of the two KEGG feature types proposed in this work results in one of the best predictive accuracies when using the binary class version of our datasets, at the same time enabling the extraction of knowledge from ageing-related data using quantitative influence information. AVAILABILITY AND IMPLEMENTATION: The datasets created in this paper will be freely available after publication. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabio Fabris, Alex Alves Freitas
Bioinform.2
2016 Improving the Interpretability of Classification Rules Discovered by an Ant Colony Algorithm: Extended Results
abstract
Most ant colony optimization (ACO) algorithms for inducing classification rules use a ACO-based procedure to create a rule in a one-at-a-time fashion. An improved search strategy has been proposed in the cAnt-Miner[Formula: see text] algorithm, where an ACO-based procedure is used to create a complete list of rules (ordered rules), i.e., the ACO search is guided by the quality of a list of rules instead of an individual rule. In this paper we propose an extension of the cAnt-Miner[Formula: see text] algorithm to discover a set of rules (unordered rules). The main motivations for this work are to improve the interpretation of individual rules by discovering a set of rules and to evaluate the impact on the predictive accuracy of the algorithm. We also propose a new measure to evaluate the interpretability of the discovered rules to mitigate the fact that the commonly used model size measure ignores how the rules are used to make a class prediction. Comparisons with state-of-the-art rule induction algorithms, support vector machines, and the cAnt-Miner[Formula: see text] producing ordered rules are also presented.
Fernando E. B. Otero, Alex Alves Freitas
Evol. Comput.2
2016 An Extensive Empirical Comparison of Probabilistic Hierarchical Classifiers in Datasets of Ageing-Related Genes
abstract
This study comprehensively evaluates the performance of five types of probabilistic hierarchical classification methods used for predicting Gene Ontology (GO) terms related to ageing. Of those tested, a new hybrid of a Local Hierarchical Classifier (LHC) and the Predictive Clustering Tree algorithm (LHC-PCT) had the best predictive accuracy results. We also tested the impact of two types of variations in most hierarchical classification algorithms, namely: (a) changing the base algorithm (we tested Naive Bayes and Support Vector Machines), and the impact of (b) using or not the Correlation based Feature Selection (CFS) algorithm in a pre-processing step. In total, we evaluated the predictive performance of 17 variations of hierarchical classifiers across 15 datasets of ageing and longevity-related genes. We conclude that the LHC-PCT algorithm ranks better across several tests (seven out of 12). In addition, we interpreted the models generated by the PCT algorithm to show how hierarchical classification algorithms can be used to extract biological insights out of the ageing-related datasets that we compiled.
Fabio Fabris, Alex Alves Freitas, Jennifer M. A. Tullet
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 A hierarchical neural network for predicting protein functions
abstract
This paper introduces the use of a modified feedforward neural network to cope with the problem of predicting protein functions. Since this kind of classification task is inherently hierarchical, this work proposes the use of two different architectures for the modified feedforward neural network, both mimicking the hierarchical nature of the classes (protein functions) to be predicted. The first approach consists of four feed-forward neural networks in cascade, each one taking as input the classification obtained by the previous network, which means, the input to a network is the classes that could be assigned to the protein at the immediately higher (parent) level in the class hierarchy. The second approach is an extension of the first one, which also adds as input to each sub-network the attributes of the protein being classified. In both situations, it was used two kinds of feed-forward architectures: an Adaline network, which is composed of a single layer of adjustable weights, and a MLP ("Multi-Layer Perceptron"), composed by two layers of adjustable weights. Both approaches were compared with a baseline consisting of a single MLP that maps the input attributes to the classes of the lowest level in the hierarchy. The MLP was built with the input layer, plus one hidden layer and one output layer. The three approaches were compared on eight datasets, the first four involving the prediction of GPCR (G-Protein Coupled Receptor) functions and the second four datasets involving the prediction of enzymes functions. The results show that a big-bang hierarchical neural network, based on the MLP paradigm, using a top-down evaluation for new instances has better behavior in hierarchical problems, when compared to its flat version.
Júlio C. Nievola, Emerson Cabrera Paraiso, Alex Alves Freitas
BIBE3
2015 A new genetic algorithm for multi-label correlation-based feature selection
Suwimol Jungjit, Alex Alves Freitas
ESANN2
2015 Simpler is Better: a Novel Genetic Algorithm to Induce Compact Multi-label Chain Classifiers
abstract
Multi-label classification (MLC) is the task of assigning multiple class labels to an object based on the features that describe the object. One of the most effective MLC methods is known as Classifier Chains (CC). This approach consists in training q binary classifiers linked in a chain, y1 → y2 → ... → yq, with each responsible for classifying a specific label in {l1, l2, ..., lq}. The chaining mechanism allows each individual classifier to incorporate the predictions of the previous ones as additional information at classification time. Thus, possible correlations among labels can be automatically exploited. Nevertheless, CC suffers from two important drawbacks: (i) the label ordering is decided at random, although it usually has a strong effect on predictive accuracy; (ii) all labels are inserted into the chain, although some of them might carry irrelevant information to discriminate the others. In this paper we tackle both problems at once, by proposing a novel genetic algorithm capable of searching for a single optimized label ordering, while at the same time taking into consideration the utilization of partial chains. Experiments on benchmark datasets demonstrate that our approach is able to produce models that are both simpler and more accurate.
Eduardo Corrêa Gonçalves, Alexandre Plastino 0001, Alex Alves Freitas
GECCO3
2015 A Novel Extended Hierarchical Dependence Network Method Based on Non-hierarchical Predictive Classes and Applications to Ageing-Related Data
abstract
We propose a novel algorithm for hierarchical classification, the Hierarchical Dependence Network based on non-Hierarchical Predictive Classes (HDN-nHPC) algorithm. HDN-nHPC uses relationships among predictive classes that are not descendants or ancestors of each other to improve classification performance and, at the same time, provide insights to non-obvious predictive class relationships. To test our algorithm and baselines, we have used hierarchical ageing-related datasets where the classes are terms in the Gene Ontology. We have concluded, based on our experiments, that using non-hierarchical predictive class relationships improves the performance of the classification algorithm and that, considering one out of three accuracy measures, the HDN-nHPC is statistically significantly better than the other three algorithms that we have tested, while no statistical significant differences were found on the other two measures.
Fabio Fabris, Alex Alves Freitas
ICTAI2
2015 An Extensive Evaluation of Decision Tree-Based Hierarchical Multilabel Classification Methods and Performance Measures
abstract
Hierarchical multilabel classification is a complex classification problem where an instance can be assigned to more than one class simultaneously, and these classes are hierarchically organized with superclasses and subclasses, that is, an instance can be classified as belonging to more than one path in the hierarchical structure. This article experimentally analyses the behavior of different decision tree–based hierarchical multilabel classification methods based on the local and global classification approaches. The approaches are compared using distinct hierarchy‐based and distance‐based evaluation measures, when they are applied to a variation of real multilabel and hierarchical datasets' characteristics. Also, the different evaluation measures investigated are compared according to their degrees of consistency, discriminancy, and indifferency. As a result of the experimental analysis, we recommend the use of the global classification approach and suggest the use of the Hierarchical Precision and Hierarchical Recall evaluation measures.
Ricardo Cerri, Gisele L. Pappa, André C. P. L. F. de Carvalho, Alex Alves Freitas
Comput. Intell.4
2015 Ant colony algorithms for constructing Bayesian multi-net classifiers
abstract
Bayesian Multi-nets (BMNs) are a special kind of Bayesian network (BN) classifiers that consist of several local Bayesian networks, one for each predictable class, to model an asymmetric set of variable dependencies given each class value. Deterministic methods using greedy local search are the mos t frequently used methods for learning the structure of BMNs based on optimizing a scoring function. Ant Colony Optimization (ACO) is a meta-heuristic global search method for solving combinatorial optimization problems, inspired by the behavior of real ant colonies. In this paper, we propose two novel ACO-based algorithms with two different approaches to build BMN classifiers: ABC-Minerlmn and ABC-Minergmn. The former uses a local learning approach, in which the ACO algorithm completes the construction of one local BN at a time. The latter uses a global approach, which involves building a complete BMN classifier by each single ant in the colony. We experimentally evaluate the performance of our ant-based algorithms on 33 benchmark classification datasets, where our proposed algorithms are shown to be significantly better than other commonly used deterministic algorithms for learning various Bayesian classifiers in the literature, as well as competitive to other well-known classification algorithms.
Khalid M. Salama, Alex Alves Freitas
Intell. Data Anal.2
2015 Predicting the Pro-Longevity or Anti-Longevity Effect of Model Organism Genes with New Hierarchical Feature Selection Methods
abstract
Ageing is a highly complex biological process that is still poorly understood. With the growing amount of ageing-related data available on the web, in particular concerning the genetics of ageing, it is timely to apply data mining methods to that data, in order to try to discover novel patterns that may assist ageing research. In this work, we introduce new hierarchical feature selection methods for the classification task of data mining and apply them to ageing-related data from four model organisms: Caenorhabditis elegans (worm), Saccharomyces cerevisiae (yeast), Drosophila melanogaster (fly), and Mus musculus (mouse). The main novel aspect of the proposed feature selection methods is that they exploit hierarchical relationships in the set of features (Gene Ontology terms) in order to improve the predictive accuracy of the Naïve Bayes and 1-Nearest Neighbour (1-NN) classifiers, which are used to classify model organisms' genes into pro-longevity or anti-longevity genes. The results show that our hierarchical feature selection methods, when used together with Naïve Bayes and 1-NN classifiers, obtain higher predictive accuracy than the standard (without feature selection) Naïve Bayes and 1-NN classifiers, respectively. We also discuss the biological relevance of a number of Gene Ontology terms very frequently selected by our algorithms in our datasets.
Cen Wan, Alex Alves Freitas, João Pedro de Magalhães
IEEE ACM Trans. Comput. Biol. Bioinform.2
2014 Extending multi-label feature selection with KEGG pathway information for microarray data analysis
abstract
We propose three approaches to extend our previous Multi-Label Correlation-based Feature Selection (ML-CFS) method with cancer-related KEGG pathway information, in order to select a better set of genes (features) for cancer microarray data classification. In the approach which produced the best results, ML-CFS was extended with a weighted formula that combines genes' predictive power and occurrence in cancer-related KEGG pathways as criteria for gene selection. We also investigated the effect of different weights for those two criteria. That approach obtained, in general, a statistically significantly smaller hamming loss (i.e. higher predictive accuracy) when compared to the hamming loss obtained by ML-CFS without using KEGG pathway information, in two cancer-related microarray datasets, using two different multi-label classification algorithms - one based on neural networks, the other based on nearest neighbors. In addition to significantly improving predictive performance, the genes selected by that approach were found to be more biologically relevant to the analysis of our datasets than genes selected without using KEGG pathway information. To the best of our knowledge, this is the first paper to propose a KEGG pathway-based feature selection method for multi-label classification.
Suwimol Jungjit, Martin Michaelis, Alex Alves Freitas, Jindrich Cinatl
CIBCB3
2014 Dependency network methods for Hierarchical Multi-label Classification of gene functions
abstract
Hierarchical Multi-label Classification (HMC) is a challenging real-world problem that naturally emerges in several areas. This work proposes two new algorithms using a Probabilistic Graphical Model based on Dependency Networks (DN) to solve the HMC problem of classifying gene functions into pre-established class hierarchies. DNs are especially attractive for their capability of using traditional, “out-of-the-shelf”, classification algorithms to model the relationship among classes and for their ability to cope with cyclic dependencies, resulting in greater flexibility with respect to Bayesian Networks. We tested our two algorithms: the first is a stand-alone Hierarchical Dependency Network (HDN) algorithm, and the second is a hybrid between the HDN and the Predictive Clustering Tree (PCT) algorithm, a well-known classifier for HMC. Based on our experiments, the hybrid classifier, using SVMs as base classifiers, obtained higher predictive accuracy than both the standard PCT algorithm and the HDN algorithm, considering 22 bioinformatics datasets and two out of three predictive accuracy measures specific for hierarchical classification (AU(PRC) and AUPRCw).
Fabio Fabris, Alex Alves Freitas
CIDM2
2014 Distinct Chains for Different Instances: An Effective Strategy for Multi-label Classifier Chains
Pablo Nascimento da Silva, Eduardo Corrêa Gonçalves, Alexandre Plastino 0001, Alex Alves Freitas
ECML/PKDD (2)4
2014 Evolving decision trees with beam search-based initialization and lexicographic multi-objective evaluation
Márcio P. Basgalupp, Rodrigo C. Barros, André C. P. L. F. de Carvalho, Alex Alves Freitas
Inf. Sci.4
2014 Evolutionary Design of Decision-Tree Algorithms Tailored to Microarray Gene Expression Data Sets
abstract
Decision-tree induction algorithms are widely used in machine learning applications in which the goal is to extract knowledge from data and present it in a graphically intuitive way. The most successful strategy for inducing decision trees is the greedy top-down recursive approach, which has been continuously improved by researchers over the past 40 years. In this paper, we propose a paradigm shift in the research of decision trees: instead of proposing a new manually designed method for inducing decision trees, we propose automatically designing decision-tree induction algorithms tailored to a specific type of classification data set (or application domain). Following recent breakthroughs in the automatic design of machine learning algorithms, we propose a hyper-heuristic evolutionary algorithm called hyper-heuristic evolutionary algorithm for designing decision-tree algorithms (HEAD-DT) that evolves design components of top-down decision-tree induction algorithms. By the end of the evolution, we expect HEAD-DT to generate a new and possibly better decision-tree algorithm for a given application domain. We perform extensive experiments in 35 real-world microarray gene expression data sets to assess the performance of HEAD-DT, and compare it with very well known decision-tree algorithms such as C4.5, CART, and REPTree. Results show that HEAD-DT is capable of generating algorithms that significantly outperform the baseline manually designed decision-tree algorithms regarding predictive accuracy and F-measure.
Rodrigo C. Barros, Márcio P. Basgalupp, Alex Alves Freitas, André C. P. L. F. de Carvalho
IEEE Trans. Evol. Comput.3
2013 Prediction of the pro-longevity or anti-longevity effect of Caenorhabditis Elegans genes based on Bayesian classification methods
abstract
The genetic mechanisms of ageing are mysterious and sophisticated issues that attract biologists' attention. With the help of data mining techniques, some findings relevant to the ageing problem can be revealed. This paper studies the performance of Bayesian network augmented naive Bayes classifier, naive Bayes classifier and proposed feature selection methods for naive Bayes on predicting a C. elegans gene's effect on the organism's longevity. The results show that due to the hierarchical structure of predictor attribute values (Gene Ontology terms), the Bayesian network augmented naive Bayes classifier performs better than the naive Bayes classifier, and the proposed feature selection methods for naive Bayes can effectively optimize the predictive performance of naive Bayes.
Cen Wan, Alex Alves Freitas
BIBM2
2013 A grammatical evolution algorithm for generation of Hierarchical Multi-Label Classification rules
abstract
Hierarchical Multi-Label Classification (HMC) is a challenging task in data mining and machine learning. Each instance in HMC can be classified into two or more classes simultaneously. These classes are structured in a hierarchy, in the form of either a tree or a directed acyclic graph. Therefore, an instance can be assigned to two or more paths from the hierarchical structure, resulting in a complex classification problem with hundreds or thousands of classes. Several methods have been proposed to deal with such problems, including several algorithms based on well-known bio-inspired techniques, such as neural networks, ant colony optimization, and genetic algorithms. In this work, we propose a novel global method called GEHM, which makes use of grammatical evolution for generating HMC rules. In this approach, the grammatical evolution algorithm evolves the antecedents of classification rules, in order to assign instances from a HMC dataset to a probabilistic class vector. Our method is compared to bio-inspired HMC algorithms in protein function prediction datasets. The empirical analysis conducted in this work shows that GEHM outperforms the bio-inspired algorithms with statistical significance, which suggests that grammatical evolution is a promising alternative to deal with hierarchical multi-label classification of biological data.
Ricardo Cerri, Rodrigo C. Barros, André C. P. L. F. de Carvalho, Alex Alves Freitas
IEEE Congress on Evolutionary Computation4
2013 Investigating the impact of various classification quality measures in the predictive accuracy of ABC-Miner
abstract
Learning classifiers from datasets is a central problem in data mining and machine learning research. ABC-Miner is an Ant-based Bayesian Classification algorithm that employs the Ant Colony Optimization (ACO) meta-heuristics to learn the structure of Bayesian Augmented Naive-Bayes (BAN) Classifiers. One of the most important aspects of the ACO algorithm is the choice of the quality measure used to evaluate a candidate solution to update pheromone. In this paper, we explore the use of various classification quality measures for evaluating the BAN classifiers constructed by the ants. The aim of this investigation is to discover how the use of different evaluation measures affects the quality of the output classifier in terms of predictive accuracy. In our experiments, we use 6 different classification measures on 25 benchmark datasets. We found that the hypothesis that different measures produce different results is acceptable according to the Friedman's statistical test.
Khalid M. Salama, Alex Alves Freitas
IEEE Congress on Evolutionary Computation2
2013 Clustering-based Bayesian Multi-net Classifier construction with Ant Colony Optimization
abstract
Bayesian Multi-nets (BMNs) are a special kind of Bayesian network (BN) classifiers that consist of several local networks, typically, one for each predictable class, to model an asymmetric set of variable dependencies given each class value. Alternatively, multi-nets can be learnt upon arbitrary partitions of a dataset, in which each partition holds more consistent variable dependencies given the data subset in the partition. This paper proposes two contributions to the approach that clusters the dataset into separate data subsets to build asymmetric local BN classifiers, one for each subset. First, we extend the K-modes algorithm, previously used by the Case-Based Bayesian Network Classifiers (CBBN) approach to create clusters before learning the BN classifiers. Second, we introduce the Ant-Clust-B algorithm that employs Ant Colony Optimization (ACO) to learn clustering-based BMNs. Ant-Clust-B uses ACO in the clustering step before learning the local BN classifiers. Empirical results are obtained from experiments on 18 UCI datasets.
Khalid M. Salama, Alex Alves Freitas
IEEE Congress on Evolutionary Computation2
2013 An Extended Local Hierarchical Classifier for Prediction of Protein and Gene Functions
Luiz H. C. Merschmann, Alex Alves Freitas
DaWaK2
2013 Improving the interpretability of classification rules discovered by an ant colony algorithm
abstract
The vast majority of Ant Colony Optimization (ACO) algorithms for inducing classification rules use an ACO-based procedure to create a rule in an one-at-a-time fashion. An improved search strategy has been proposed in the cAnt-MinerPB algorithm, where an ACO-based procedure is used to create a complete list of rules (ordered rules) - i.e., the ACO search is guided by the quality of a list of rules, instead of an individual rule. In this paper we propose an extension of the cAnt-MinerPB algorithm to discover a set of rules (unordered rules). The main motivation for discovering a set of rules is to improve the interpretation of individual rules and evaluate the impact on the predictive accuracy of the algorithm. We also propose a new measure to evaluate the interpretability of the discovered rules to mitigate the fact that the commonly-used model size measure ignores how the rules are used to make a class prediction. Comparisons with state-of-the-art rule induction algorithms and the cAnt-MinerPB producing ordered rules are also presented.
Fernando E. B. Otero, Alex Alves Freitas
GECCO2
2013 A Genetic Algorithm for Optimizing the Label Ordering in Multi-label Classifier Chains
abstract
First proposed in 2009, the classifier chains model (CC) has become one of the most influential algorithms for multi-label classification. It is distinguished by its simple and effective approach to exploit label dependencies. The CC method involves the training of q single-label binary classifiers, where each one is solely responsible for classifying a specific label in ll, ..., lq. These q classifiers are linked in a chain, such that each binary classifier is able to consider the labels predicted by the previous ones as additional information at classification time. The label ordering has a strong effect on predictive accuracy, however it is decided at random and/or combining random orders via an ensemble. A disadvantage of the ensemble approach consists of the fact that it is not suitable when the goal is to generate interpretable classifiers. To tackle this problem, in this work we propose a genetic algorithm for optimizing the label ordering in classifier chains. Experiments on diverse benchmark datasets, followed by the Wilcoxon test for assessing statistical significance, indicate that the proposed strategy produces more accurate classifiers.
Eduardo Corrêa Gonçalves, Alexandre Plastino 0001, Alex Alves Freitas
ICTAI3
2013 Probabilistic Clustering for Hierarchical Multi-Label Classification of Protein Functions
Rodrigo C. Barros, Ricardo Cerri, Alex Alves Freitas, André C. P. L. F. de Carvalho
ECML/PKDD (2)3
2013 Two Extensions to Multi-label Correlation-Based Feature Selection: A Case Study in Bioinformatics
abstract
This paper proposes two extensions to a Multi-Label Correlation Based Feature Selection Method (ML-CFS): (1) ML-CFS using the absolute value of the correlation coefficient in the equation for evaluating a candidate feature subset, and (2) ML-CFS using Mutual Information for class label weighting. These extensions are evaluated in a bioinformatics case study addressing the multi-label classification of a cancer-related DNA micro array dataset with over 20,000 features. The results show that ML-CFS with absolute value of correlation obtained a significantly better predictive accuracy (smaller hamming loss) than the original ML-CFS. On the other hand, using Mutual Information to assign weights to labels showed some positive effect when using the ML-RBF classifier, but it showed a negative effect when using the ML-kNN classifier.
Suwimol Jungjit, Martin Michaelis, Alex Alves Freitas, Jindrich Cinatl
SMC3
2013 Automatic Design of Decision-Tree Algorithms with Evolutionary Algorithms
abstract
This study reports the empirical analysis of a hyper-heuristic evolutionary algorithm that is capable of automatically designing top-down decision-tree induction algorithms. Top-down decision-tree algorithms are of great importance, considering their ability to provide an intuitive and accurate knowledge representation for classification problems. The automatic design of these algorithms seems timely, given the large literature accumulated over more than 40 years of research in the manual design of decision-tree induction algorithms. The proposed hyper-heuristic evolutionary algorithm, HEAD-DT, is extensively tested using 20 public UCI datasets and 10 microarray gene expression datasets. The algorithms automatically designed by HEAD-DT are compared with traditional decision-tree induction algorithms, such as C4.5 and CART. Experimental results show that HEAD-DT is capable of generating algorithms which are significantly more accurate than C4.5 and CART.
Rodrigo C. Barros, Márcio P. Basgalupp, André C. P. L. F. de Carvalho, Alex Alves Freitas
Evol. Comput.4
2013 A New Sequential Covering Strategy for Inducing Classification Rules With Ant Colony Algorithms
abstract
Ant colony optimization (ACO) algorithms have been successfully applied to discover a list of classification rules. In general, these algorithms follow a sequential covering strategy, where a single rule is discovered at each iteration of the algorithm in order to build a list of rules. The sequential covering strategy has the drawback of not coping with the problem of rule interaction, i.e., the outcome of a rule affects the rules that can be discovered subsequently since the search space is modified due to the removal of examples covered by previous rules. This paper proposes a new sequential covering strategy for ACO classification algorithms to mitigate the problem of rule interaction, where the order of the rules is implicitly encoded as pheromone values and the search is guided by the quality of a candidate list of rules. Our experiments using 18 publicly available data sets show that the predictive accuracy obtained by a new ACO classification algorithm implementing the proposed sequential covering strategy is statistically significantly higher than the predictive accuracy of state-of-the-art rule induction classification algorithms.
Fernando E. B. Otero, Alex Alves Freitas, Colin G. Johnson
IEEE Trans. Evol. Comput.2
2012 AutoClustering: An estimation of distribution algorithm for the automatic generation of clustering algorithms
abstract
Most of the existing Data Mining algorithms have been manually produced, that is, have been developed by a human programmer. A prominent Artificial Intelligence research area is automatic programming - the generation of a computer program by another computer program. Clustering is an important data mining task with many useful real-world applications. Particularly, the class of clustering algorithms based on the idea of data density to identify clusters has many advantages, such as the ability to identify arbitrary-shape clusters. We propose the use of Estimation of Distribution Algorithms for the artificial generation of density-based clustering algorithms. In order to guarantee the generation of valid algorithms, a directed acyclic graph (DAG) was defined where each node represents a procedure (building block) and each edge represents a possible execution sequence between two nodes. The Building Blocks DAG specifies the alphabet of the EDA, that is, any possibly generated algorithm. Preliminary experimental results compare the clustering algorithms artificially generated by AutoClustering to DBSCAN, a well-known manually-designed algorithm.
Aruanda Simões Goncalves Meiguins, Roberto Célio Limão de Oliveira, Bianchi Serique Meiguins, Samuel F. S. Junior, Alex Alves Freitas
IEEE Congress on Evolutionary Computation5
2012 Evolving recursive programs using non-recursive scaffolding
abstract
Genetic programming has proven capable of evolving solutions to a wide variety of problems. However, the successes have largely been with programs without iteration or recursion; evolving recursive programs has turned out to be particularly challenging. The main obstacle to evolving recursive programs seems to be that they are particularly fragile to the application of search operators: a small change in a correct recursive program generally produces a completely wrong program. In this paper, we present a simple and general method that allows us to pass back and forth from a recursive program to an associated non-recursive program. Finding a recursive program can be reduced to evolving non-recursive programs followed by converting the optimum non-recursive program found to the associated optimum recursive program. This avoids the fragility problem above, as evolution does not search the space of recursive programs. We present promising experimental results on a test-bed of recursive problems.
Alberto Moraglio, Fernando E. B. Otero, Colin G. Johnson, Simon J. Thompson, Alex Alves Freitas
IEEE Congress on Evolutionary Computation5
2012 A hyper-heuristic evolutionary algorithm for automatically designing decision-tree algorithms
abstract
Decision tree induction is one of the most employed methods to extract knowledge from data, since the representation of knowledge is very intuitive and easily understandable by humans. The most successful strategy for inducing decision trees, the greedy top-down approach, has been continuously improved by researchers over the years. This work, following recent breakthroughs in the automatic design of machine learning algorithms, proposes a hyper-heuristic evolutionary algorithm for automatically generating decision-tree induction algorithms, named HEAD-DT. We perform extensive experiments in 20 public data sets to assess the performance of HEAD-DT, and we compare it to traditional decision-tree algorithms such as C4.5 and CART. Results show that HEAD-DT can generate algorithms that significantly outperform C4.5 and CART regarding predictive accuracy and F-Measure.
Rodrigo C. Barros, Márcio P. Basgalupp, André C. P. L. F. de Carvalho, Alex Alves Freitas
GECCO4
2012 A Survey of Evolutionary Algorithms for Decision-Tree Induction
abstract
This paper presents a survey of evolutionary algorithms that are designed for decision-tree induction. In this context, most of the paper focuses on approaches that evolve decision trees as an alternate heuristics to the traditional top-down divide-and-conquer approach. Additionally, we present some alternative methods that make use of evolutionary algorithms to improve particular components of decision-tree classifiers. The paper's original contributions are the following. First, it provides an up-to-date overview that is fully focused on evolutionary algorithms and decision trees and does not concentrate on any specific evolutionary approach. Second, it provides a taxonomy, which addresses works that evolve decision trees and works that design decision-tree components by the use of evolutionary algorithms. Finally, a number of references are provided that describe applications of evolutionary algorithms for decision-tree induction in different domains. At the end of this paper, we address some important issues and open questions that can be the subject of future research.
Rodrigo C. Barros, Márcio P. Basgalupp, André C. P. L. F. de Carvalho, Alex Alves Freitas
IEEE Trans. Syst. Man Cybern. Part C4
2011 A hierarchical approach to represent relational data applied to clustering tasks
abstract
Nowadays, the representation of many real word problems needs to use some type of relational model. As a consequence, information used by a wide range of systems has been stored in multi relational tables. However, from a data mining point of view, it has been a problem, since most of the traditional data mining algorithms have not been originally proposed to handle this type of data without discarding relationship information. Aiming to ameliorate this problem, we propose a hierarchical approach for handling relational data. In this approach the relational data is converted into a hierarchical structure (the main table as the root and the relations as the nodes). This hierarchical way to represent relational data can be used either for classification or clustering purposes. In this paper, we will use it in clustering algorithms. In order to do so, we propose a hierarchical distance metric to compute the similarity between the tables. In the empirical analysis, we will apply the proposed approach in two well-known clustering algorithms (k-means and agglomerative hierarchical). Finally, this paper also compares the effectiveness of our approach with one existing relational approach.
João Carlos Xavier Jr., Anne M. P. Canuto, Alex Alves Freitas, Luiz Marcos Garcia Gonçalves, Carlos Nascimento Silla Jr.
IJCNN3
2011 A survey of hierarchical classification across different application domains
Carlos Nascimento Silla Jr., Alex Alves Freitas
Data Min. Knowl. Discov.2
2011 Adapting non-hierarchical multilabel classification methods for hierarchical multilabel classification
abstract
In most classification problems, a classifier assigns a single class to each instance and the classes form a flat (non-hierarchical) structure, without superclasses or subclasses. In hierarchical multilabel classification problems, the classes are hi
Ricardo Cerri, André C. P. L. F. de Carvalho, Alex Alves Freitas
Intell. Data Anal.3
2011 Lazy attribute selection: Choosing attributes at classification time
abstract
Attribute selection is a data preprocessing step which aims at identifying relevant attributes for the target machine learning task – namely classification in this paper. In this paper, we propose a new attribute selection strategy – based on a lazy
Rafael B. Pereira, Alexandre Plastino 0001, Bianca Zadrozny, Luiz H. C. Merschmann, Alex Alves Freitas
Intell. Data Anal.5
2011 Selecting different protein representations and classification algorithms in hierarchical protein function prediction
abstract
Automatically inferring the function of unknown proteins is a challenging task in proteomics. There are two major problems in the task of computational protein function prediction, which are the choice of the protein representation and the choice of the classification algorithm. There are several w ays of extracting features from a protein, and the choice of the feature representation might be as important as the choice of the classification algorithm. These problems are aggravated in the case of hierarchical protein function prediction, where a hierarchy of classifiers is built and each of those classifiers' construction has to consider the aforementioned selection problems. In this paper we address these problem by employing three alternative selective hierarchical classification approaches: (a) selecting the best classifier given a fixed representation; (b) selecting the best representation given a fixed classifier; and (c) selecting the best classifier and representation simultaneously, in a synergistic fashion. The analysis of the results have shown that the selective representation approach is almost always ranked number 1 when compared against the different fixed representations and that the use of the selective classifier approach is not able to surpass using only the best classifier for the target problem.
Carlos Nascimento Silla Jr., Alex Alves Freitas
Intell. Data Anal.2
2011 A genetic programming method for protein motif discovery and protein classification
Denise Fukumi Tsunoda, Alex Alves Freitas, Heitor Silvério Lopes
Soft Comput.2
2010 Knowledge discovery with Artificial Immune Systems for hierarchical multi-label classification of protein functions
abstract
This work presents a system for knowledge discovery from protein databases, based on an Artificial Immune System. The discovered rules have the advantage of representing comprehensible knowledge to biologist users. This task leads to a very challenging problem since a protein can be assigned multiple classes (functions or Gene Ontology (GO) terms) across several levels of the GO's term hierarchy. To solve this problem we present two versions of an algorithm called MHC-AIS (Multi-label Hierarchical Classification with an Artificial Immune System), which is a sophisticated classification algorithm tailored to both multi-label and hierarchical classification. The first version of MHC-AIS builds a global classifier to predict all classes in the dataset, whilst the second version builds a local classifier to predict each class. The proposed versions and an algorithm chosen for comparison are evaluated on a protein dataset, and the results show that MHC-AIS outperformed the compared algorithm in general.
Roberto Teixeira Alves, Miguel R. Delgado, Alex Alves Freitas
FUZZ-IEEE3
2010 An estimation of distribution algorithm for the automatic generation of clustering algorithms
abstract
A prominent Artificial Intelligence research area is automatic programming -- the generation of a computer program by another computer program. We propose the use of Estimation of Distribution Algorithms (EDA) for the artificial generation of density-based clustering algorithms -- a class of data mining algorithms. Preliminary experimental results compare the clustering algorithms artificially generated by AutoClustering to DBSCAN, a well-known manually-designed algorithm.
Aruanda Simões Goncalves Meiguins, Alex Alves Freitas, Roberto Célio Limão de Oliveira, Samuel F. S. Junior, Bianchi Serique Meiguins
GECCO2
2010 Learning hybridization strategies in evolutionary algorithms
abstract
Evolutionary Algorithms are powerful optimization techniques which have been applied to many different problems, from complex mathematical functions to real-world applications. Some studies report performance improvements through the combination of d
Antonio LaTorre, José M. Peña 0002, Santiago Muelas, Alex Alves Freitas
Intell. Data Anal.4
2010 On the Importance of Comprehensible Classification Models for Protein Function Prediction
abstract
The literature on protein function prediction is currently dominated by works aimed at maximizing predictive accuracy, ignoring the important issues of validation and interpretation of discovered knowledge, which can lead to new insights and hypotheses that are biologically meaningful and advance the understanding of protein functions by biologists. The overall goal of this paper is to critically evaluate this approach, offering a refreshing new perspective on this issue, focusing not only on predictive accuracy but also on the comprehensibility of the induced protein function prediction models. More specifically, this paper aims to offer two main contributions to the area of protein function prediction. First, it presents the case for discovering comprehensible protein function prediction models from data, discussing in detail the advantages of such models, namely, increasing the confidence of the biologist in the system's predictions, leading to new insights about the data and the formulation of new biological hypotheses, and detecting errors in the data. Second, it presents a critical review of the pros and cons of several different knowledge representations that can be used in order to support the discovery of comprehensible protein function prediction models.
Alex Alves Freitas, Daniela Wieser, Rolf Apweiler
IEEE ACM Trans. Comput. Biol. Bioinform.1
2009 Handling continuous attributes in Ant Colony Classification algorithms
abstract
Most real-world classification problems involve continuous (real-valued) attributes, as well as, nominal (discrete) attributes. The majority of ant colony optimisation (ACO) classification algorithms have the limitation of only being able to cope with nominal attributes directly. Extending the approach for coping with continuous attributes presented by cAnt-Miner (Ant-Miner coping with continuous attributes), in this paper we propose two new methods for handling continuous attributes in ACO classification algorithms. The first method allows a more flexible representation of continuous attributes' intervals. The second method explores the problem of attribute interaction, which originates from the way that continuous attributes are handled in cAnt-Miner, in order to implement an improved pheromone updating method. Empirical evaluation on eight publicly available data sets shows that the proposed methods facilitate the discovery of more accurate classification models.
Fernando E. B. Otero, Alex Alves Freitas, Colin G. Johnson
CIDM2
2009 A Hybrid Evolutionary Approach for the Protein Classification Problem
Denise Fukumi Tsunoda, Heitor Silvério Lopes, Alex Alves Freitas
ICCCI3
2009 A Global-Model Naive Bayes Approach to the Hierarchical Prediction of Protein Functions
abstract
In this paper we propose a new global-model approach for hierarchical classification, where a single global classification model is built by considering all the classes in the hierarchy - rather than building a number of local classification models as it is more usual in hierarchical classification. The method is an extension of the flat classification algorithm naive Bayes. We present the extension made to the original algorithm as well as its evaluation on eight protein function hierarchical classification datasets. The achieved results are positive and show that the proposed global model is better than using a local model approach.
Carlos Nascimento Silla Jr., Alex Alves Freitas
ICDM2
2009 MAHATMA: A Genetic Programming-Based Tool for Protein Classification
abstract
Proteins can be grouped into families according to some features such as hydrophobicity, composition or structure, aiming to establish common biological functions. This paper presents a system that was conceived to discover features (particular sequences of amino acids, or motifs) that occur very often in proteins of a given family but rarely occur in proteins of other families. These features can be used for the classification of unknown proteins, that is, to predict their function by analyzing their primary structure. Experiments were done with a set of enzymes extracted from the protein data bank. The heuristic method used was based on genetic programming using operators specially tailored for the target problem. The final performance was measured using sensitivity (Se) and specificity (Sp). The best results obtained for the enzyme dataset suggest that the proposed evolutionary computation method is very effective to find predictive features (motifs) for protein classification.
Denise Fukumi Tsunoda, Alex Alves Freitas, Heitor Silvério Lopes
ISDA2
2009 A Hybrid Data Mining Metaheuristic for the p-Median Problem
abstract
Metaheuristics represent an important class of techniques to solve, approximately, hard combinatorial optimization problems for which the use of exact methods is impractical. In this work, we propose a hybrid version of the GRASP metaheuristic, which incorporates a data mining process, to solve the p-median problem. We believe that patterns obtained by a data mining technique, from a set of sub-optimal solutions of a combinatorial optimization problem, can be used to guide metaheuristic procedures in the search for better solutions. Traditional GRASP is an iterative metaheuristic which returns the best solution reached over all iterations. In the hybrid GRASP proposal, after executing a significant number of iterations, the data mining process extracts patterns from an elite set of sub-optimal solutions for the p-median problem. These patterns present characteristics of near optimal solutions and can be used to guide the following GRASP iterations in the search through the combinatorial solution space. Computational experiments, comparing traditional GRASP and different data mining hybrid proposals for the p-median problem, showed that employing patterns mined from an elite set of sub-optimal solutions made the hybrid GRASP find better results. Besides, the conducted experiments also evidenced that incorporating a data mining technique into a metaheuristic accelerated the process of finding near optimal and optimal solutions.
Alexandre Plastino 0001, Erick Rocha Fonseca, Richard Fuchshuber, Simone L. Martins, Alex Alves Freitas, Martino Luis, Saïd Salhi
SDM5
2009 Novel Top-Down Approaches for Hierarchical Classification and Their Application to Automatic Music Genre Classification
abstract
This paper presents two novel hierarchical classification methods which are extensions of a previously proposed selective classifier top-down approach, which consists of selecting - during the training phase - the best classifier at each node of a classifier tree. More precisely, we propose two novel selective top-down hierarchical methods. First, a method that selects the best feature set instead of the best classifier. Secondly, a method that selects both the best classifier and the best representation simultaneously. These methods are evaluated on the task of hierarchical music genre classification using four different types of feature sets extracted from each song and four classifiers.
Carlos Nascimento Silla Jr., Alex Alves Freitas
SMC2
2009 Automatically evolving rule induction algorithms tailored to the prediction of postsynaptic activity in proteins
abstract
It is well-known that no classification algorithm is the best in all application domains. The conventional approach for coping with this problem consists of trying to select the best classification algorithm for the target application domain. We prop
Gisele L. Pappa, Alex Alves Freitas
Intell. Data Anal.2
2009 Evolving rule induction algorithms with multi-objective grammar-based genetic programming
Gisele L. Pappa, Alex Alves Freitas
Knowl. Inf. Syst.2
2009 Hierarchical classification of protein function with ensembles of rules and particle swarm optimisation
Nicholas Holden, Alex Alves Freitas
Soft Comput.2
2009 A Survey of Evolutionary Algorithms for Clustering
abstract
This paper presents a survey of evolutionary algorithms designed for clustering tasks. It tries to reflect the profile of this area by focusing more on those subjects that have been given more importance in the literature. In this context, most of the paper is devoted to partitional algorithms that look for hard clusterings of data, though overlapping (i.e., soft and fuzzy) approaches are also covered in the paper. The paper is original in what concerns two main aspects. First, it provides an up-to-date overview that is fully devoted to evolutionary algorithms for clustering, is not limited to any particular kind of evolutionary approach, and comprises advanced topics like multiobjective and ensemble-based evolutionary clustering. Second, it provides a taxonomy that highlights some very important aspects in the context of evolutionary data clustering, namely, fixed or variable number of clusters, cluster-oriented or nonoriented operators, context-sensitive or context-insensitive operators, guided or unguided operators, binary, integer, or real encodings, centroid-based, medoid-based, label-based, tree-based, or graph-based representations, among others. A number of references are provided that describe applications of evolutionary algorithms for clustering in different domains, such as image processing, computer security, and bioinformatics. The paper ends by addressing some important issues and open questions that can be subject of future research.
Eduardo R. Hruschka, Ricardo J. G. B. Campello, Alex Alves Freitas, André C. P. L. F. de Carvalho
IEEE Trans. Syst. Man Cybern. Part C3
2008 Optimizing amino acid groupings for GPCR classification
abstract
MOTIVATION: There is much interest in reducing the complexity inherent in the representation of the 20 standard amino acids within bioinformatics algorithms by developing a so-called reduced alphabet. Although there is no universally applicable residue grouping, there are numerous physiochemical criteria upon which one can base groupings. Local descriptors are a form of alignment-free analysis, the efficiency of which is dependent upon the correct selection of amino acid groupings. RESULTS: Within the context of G-protein coupled receptor (GPCR) classification, an optimization algorithm was developed, which was able to identify the most efficient grouping when used to generate local descriptors. The algorithm was inspired by the relatively new computational intelligence paradigm of artificial immune systems. A number of amino acid groupings produced by this algorithm were evaluated with respect to their ability to generate local descriptors capable of providing an accurate classification algorithm for GPCRs.
Matthew N. Davies, Andrew Secker, Alex Alves Freitas, Edward Clark, Jonathan Timmis, Darren R. Flower
Bioinform.3
2008 Message-passing algorithms for the prediction of protein domain interactions from protein-protein interaction data
abstract
MOTIVATION: Cellular processes often hinge upon specific interactions among proteins, and knowledge of these processes at a system level constitutes a major goal of proteomics. In particular, a greater understanding of protein-protein interactions can be gained via a more detailed investigation of the protein domain interactions that mediate the interactions of proteins. Existing high-throughput experimental techniques assay protein-protein interactions, yet they do not provide any direct information on the interactions among domains. Inferences concerning the latter can be made by analysis of the domain composition of a set of proteins and their interaction map. This inference problem is non-trivial, however, due to the high level of noise generally present in experimental data concerning protein-protein interactions. This noise leads to contradictions, i.e. the impossibility of having a pattern of domain interactions compatible with the protein-protein interaction map. RESULTS: We formulate the problem of prediction of protein domain interactions in a form that lends itself to the application of belief propagation, a powerful algorithm for such inference problems, which is based on message passing. The input to our algorithm is an interaction map among a set of proteins, and a set of domain assignments to the relevant proteins. The output is a list of probabilities of interaction between each pair of domains. Our method is able to effectively cope with errors in the protein-protein interaction dataset and systematically resolve contradictions. We applied the method to a dataset concerning the budding yeast Saccharomyces cerevisiae and tested the quality of our predictions by cross-validation on this dataset, by comparison with existing computational predictions, and finally with experimentally available domain interactions. Results compare favourably to those by existing algorithms. AVAILABILITY: A C language implementation of the algorithm is available upon request.
Mudassar Iqbal, Alex Alves Freitas, Colin G. Johnson, Massimo Vergassola
Bioinform.2
2007 WAIRS: improving classification accuracy by weighting attributes in the AIRS classifier
abstract
AIRS (Artificial Immune Recognition System) has shown itself to be a competitive classifier. It has also proved to be the most popular immune inspired classifier. However, rather than AIRS being a classifier in its own right as previously described, we see AIRS more as a pre-processor to a KNN classifier. It is our view that by not explicitly classing it as such development of this algorithm has been rather held back. Seeing it as a pre-processor allows inspiration to be taken from the machine learning literature where such pre-processors are not uncommon. With this in mind, this paper takes a core feature of many such pre-processors, that of attribute weighting, and applies it to AIRS. The resultant algorithm called WAIRS (Weighted AIRS) uses a weighted distance function during all affinity evaluations. WAIRS is tested on 9 benchmark datasets and is found to outperform AIRS in the majority of cases.
Andrew Secker, Alex Alves Freitas
IEEE Congress on Evolutionary Computation2
2007 Induction of fuzzy rules with artificial immune systems in acgh based er status breast cancer characterization
abstract
Genomic DNA copy number aberrations are frequent in solid tumours although their underlying causes remain obscure. In this paper we show how Artificial Immune System (AIS) paradigm can be successfully employed in the elucidation of biological dynamics of cancerous processes using a novel fuzzy rule induction system for data mining (IFRAIS). Competitive results have been obtained using IFRAIS. A biological interpretation of the results, carried out using Gene Ontology, followed the statistical assessment and put in evidence interesting patterns that are currently under investigation.
Filippo Menolascina, Roberto Teixeira Alves, Stefania Tommasi, Patrizia Chiarappa, Myriam Delgado, Giuseppe Mastronardi, Angelo Paradiso, Alex Alves Freitas, Vitoantonio Bevilacqua
GECCO8
2007 EDACluster: an Evolutionary Density and Grid-Based Clustering Algorithm
abstract
This paper presents EDACluster, an estimation of distribution algorithm (EDA) applied to the clustering task. EDA is an evolutionary algorithm used here to optimize the search for adequate clusters when very little is known about the target dataset. The proposed algorithm uses a mixed approach - density and grid- based - to identify sets of dense cells in the dataset. The output is a list of items and their associated clusters. Items in low-density areas are considered noise and are not assigned to any cluster. This work uses four public domain datasets to perform the tests that compare EDACluster with DBSCAN, a conventional density-based clustering algorithm.
César S. de Oliveira, Paulo Igor Alves Godinho, Aruanda Simões Goncalves Meiguins, Bianchi Serique Meiguins, Alex Alves Freitas
ISDA5
2007 On the hierarchical classification of G protein-coupled receptors
abstract
MOTIVATION: G protein-coupled receptors (GPCRs) play an important role in many physiological systems by transducing an extracellular signal into an intracellular response. Over 50% of all marketed drugs are targeted towards a GPCR. There is considerable interest in developing an algorithm that could effectively predict the function of a GPCR from its primary sequence. Such an algorithm is useful not only in identifying novel GPCR sequences but in characterizing the interrelationships between known GPCRs. RESULTS: An alignment-free approach to GPCR classification has been developed using techniques drawn from data mining and proteochemometrics. A dataset of over 8000 sequences was constructed to train the algorithm. This represents one of the largest GPCR datasets currently available. A predictive algorithm was developed based upon the simplest reasonable numerical representation of the protein's physicochemical properties. A selective top-down approach was developed, which used a hierarchical classifier to assign sequences to subdivisions within the GPCR hierarchy. The predictive performance of the algorithm was assessed against several standard data mining classifiers and further validated against Support Vector Machine-based GPCR prediction servers. The selective top-down approach achieves significantly higher accuracy than standard data mining methods in almost all cases.
Matthew N. Davies, Andrew Secker, Alex Alves Freitas, Miguel Mendao, Jonathan Timmis, Darren R. Flower
Bioinform.3
2007 Revisiting the Foundations of Artificial Immune Systems for Data Mining
abstract
This paper advocates a problem-oriented approach for the design of artificial immune systems (AIS) for data mining. By problem-oriented approach we mean that, in real-world data mining applications the design of an AIS should take into account the characteristics of the data to be mined together with the application domain: the components of the AIS - such as its representation, affinity function, and immune process - should be tailored for the data and the application. This is in contrast with the majority of the literature, where a very generic AIS algorithm for data mining is developed and there is little or no concern in tailoring the components of the AIS for the data to be mined or the application domain. To support this problem-oriented approach, we provide an extensive critical review of the current literature on AIS for data mining, focusing on the data mining tasks of classification and anomaly detection. We discuss several important lessons to be taken from the natural immune system to design new AIS that are considerably more adaptive than current AIS. Finally, we conclude this paper with a summary of seven limitations of current AIS for data mining and ten suggested research directions.
Alex Alves Freitas, Jonathan Timmis
IEEE Trans. Evol. Comput.1
2006 Automatically Evolving Rule Induction Algorithms
Gisele L. Pappa, Alex Alves Freitas
ECML2
2006 A new ant colony algorithm for multi-label classification with applications in bioinfomatics
abstract
The conventional classification task of data mining can be called single-label classification, since there is a single class attribute to be predicted. This paper addresses a more challenging version of the classification task, where there are two or more class attributes to be predicted. We propose a new ant colony algorithm for the multi-label classification task. The new algorithm, called MuLAM (Multi-Label Ant-Miner) is a major extension of Ant-Miner, the first ant colony algorithm for discovering classification rules. We report results comparing the performance of MuLAM with the performance of three other classification techniques, namely the very simple majority classifier, the original Ant-Miner algorithm and C5.0, a very popular rule induction algorithm. The experiments were performed using five bioinformatics datasets, involving the prediction of several kinds of protein function.
Allen Chan, Alex Alves Freitas
GECCO2
2006 A new discrete particle swarm algorithm applied to attribute selection in a bioinformatics data set
abstract
Many data mining applications involve the task of building a model for predictive classification. The goal of such a model is to classify examples (records or data instances) into classes or categories of the same type. The use of variables (attributes) not related to the classes can reduce the accuracy and reliability of a classification or prediction model. Superuous variables can also increase the costs of building a model - particularly on large data sets. We propose a discrete Particle Swarm Optimization (PSO) algorithm designed for attribute selection. The proposed algorithm deals with discrete variables, and its population of candidate solutions contains particles of different sizes. The performance of this algorithm is compared with the performance of a standard binary PSO algorithm on the task of selecting attributes in a bioinformatics data set. The criteria used for comparison are: (1) maximizing predictive accuracy; and (2) finding the smallest subset of attributes.
Elon Santos Correa, Alex Alves Freitas, Colin G. Johnson
GECCO2
2006 Estimating photometric redshifts with genetic algorithms
abstract
Photometry is used as a cheap and easy way to estimate redshifts of galaxies, which would otherwise require considerable amounts of expensive telescope time. However, the analysis of photometric redshift datasets is a task where it is sometimes difficult to achieve a high classification accuracy. This work presents a custom Genetic Algorithm (GA) for mining the Hubble Deep Field North (HDF-N) datasets to achieve accurate IF-THEN classification rules. This kind of knowledge representation has the advantage of being intuitively comprehensible to the user, facilitating astronomers' interpretation of discovered knowledge. The GA is tested against the state of the art decision tree algorithm C5.0 [6] achieving significantly better results.
Nick Miles, Alex Alves Freitas, Stephen Serjeant
GECCO2
2006 A new version of the ant-miner algorithm discovering unordered rule sets
abstract
The Ant-Miner algorithm, first proposed by Parpinelli and colleagues, applies an ant colony optimization heuristic to the classification task of data mining to discover an ordered list of classification rules. In this paper we present a new version of the Ant-Miner algorithm, which we call Unordered Rule Set Ant-Miner, that produces an unordered set of classification rules. The proposed version was evaluated against the original Ant-Miner algorithm in six public-domain datasets and was found to produce comparable results in terms of predictive accuracy. However, the proposed version has the advantage of discovering more modular rules, i.e., rules that can be interpreted independently from other rules – unlike the rules in an ordered list, where the interpretation of a rule requires knowledge of the previous rules in the list. Hence, the proposed version facilitates the interpretation of discovered knowledge, an important point in data mining. Categories and Subject Descriptors I.2.6 [Artificial Intelligence]: Learning – concept learning, induction.
James Smaldon, Alex Alves Freitas
GECCO2
2005 Evaluating the Correlation Between Objective Rule Interestingness Measures and Real Human Interest
Deborah R. Carvalho, Alex Alves Freitas, Nelson F. F. Ebecken
PKDD2
2005 A hybrid particle swarm/ant colony algorithm for the classification of hierarchical biological data
abstract
This paper proposes a hybrid PSO/ACO algorithm for hierarchical classification, where the classes to be predicted are arranged in a tree-like hierarchy. The performance of the algorithm is evaluated on a challenging biological data set, involving the hierarchical functional classification of enzymes. The proposed algorithm is compared with an existing PSO for classification, which was also adapted for hierarchical classification.
Nicholas Holden, Alex Alves Freitas
SIS2
2004 An Artificial Immune System for Fuzzy-Rule Induction in Data Mining
Roberto Teixeira Alves, Myriam Delgado, Heitor Silvério Lopes, Alex Alves Freitas
PPSN4
2004 Web Page Classification with an Ant Colony Algorithm
Nicholas Holden, Alex Alves Freitas
PPSN2
2004 A constrained-syntax genetic programming system for discovering classification rules: application to medical data sets
Celia C. Bojarczuk, Heitor Silvério Lopes, Alex Alves Freitas, Edson L. Michalkiewicz
Artif. Intell. Medicine3
2004 A hybrid decision tree/genetic algorithm method for data mining
Deborah R. Carvalho, Alex Alves Freitas
Inf. Sci.2
2003 AISEC: an artificial immune system for e-mail classification
abstract
With the increase in information on the Internet, the strive to find more effective tools for distinguishing between interesting and non-interesting material is increasing. Drawing analogies from the biological immune system, this paper presents an immune-inspired algorithm called AISEC that is capable of continuously classifying electronic mail as interesting and non-interesting without the need for re-training. Comparisons are drawn with a naive Bayesian classifier and it is shown that the proposed system performs as well as the naive Bayesian system and has a great potential for augmentation.
Andrew Secker, Alex Alves Freitas, Jonathan Timmis
IEEE Congress on Evolutionary Computation2
2003 An Innovative Application of a Constrained-Syntax Genetic Programming System to the Problem of Predicting Survival of Patients
Celia C. Bojarczuk, Heitor Silvério Lopes, Alex Alves Freitas
EuroGP3
2003 Genetic Programming for Attribute Construction in Data Mining
Fernando E. B. Otero, Monique M. S. Silva, Alex Alves Freitas, Júlio C. Nievola
EuroGP3
2003 Guest editorial data mining and knowledge discovery with evolutionary algorithms
Ashish Ghosh, Alex Alves Freitas
IEEE Trans. Evol. Comput.2
2002 A Genetic Algorithm With Sequential Niching For Discovering Small-disjunct Rules
Deborah R. Carvalho, Alex Alves Freitas
GECCO2
2002 Constructing X-of-n Attributes With A Genetic Algorithm
Otavio Larsen, Alex Alves Freitas, Júlio C. Nievola
GECCO2
2002 Genetic Programming For Attribute Construction In Data Mining
Fernando E. B. Otero, Monique M. S. Silva, Alex Alves Freitas
GECCO3
2002 A Genetic Algorithm For Discovering Interesting Fuzzy Prediction Rules: Applications To Science And Technology Data
Wesley Romão, Alex Alves Freitas, Roberto Carlos dos Santos Pacheco
GECCO2
2002 Devising adaptive migration policies for cooperative distributed genetic algorithms
abstract
Distributed genetic algorithms (DGAs) constitute an interesting approach to undertake the premature convergence problem in evolutionary optimization. This is done by spatial partitioning a huge panmitic population into several semi-isolated groups, called demes, each evolving in parallel by its own pace, and possibly exploring different regions of the search space. At the center of such approach lies the migratory process that simulates the swapping of individuals belonging to different demes, in such a way to ensure the sharing of good genetic material. In this paper, we model the migration step in DGAs as an explicit means to promote cooperation among genetic agents, autonomous entities encapsulating GA instances for possibly tackling different sub-problems of a complicated task. The focus is on the characterization of adaptive migration policies in which the choice of what individuals to migrate and/or replace is not defined a priori but according to a more knowledge-oriented rule. Comparative results obtained for a data-mining task were conducted, in order to assess the performance of adaptive migration according to efficiency/effectiveness criteria.
Edgar Noda, André L. V. Coelho, Ivan Luiz Marques Ricarte, Akebo Yamakami, Alex Alves Freitas
SMC5
2002 Data mining with an ant colony optimization algorithm
abstract
The paper proposes an algorithm for data mining called Ant-Miner (ant-colony-based data miner). The goal of Ant-Miner is to extract classification rules from data. The algorithm is inspired by both research on the behavior of real ant colonies and some data mining concepts as well as principles. We compare the performance of Ant-Miner with CN2, a well-known data mining algorithm for classification, in six public domain data sets. The results provide evidence that: 1) Ant-Miner is competitive with CN2 with respect to predictive accuracy, and 2) the rule lists discovered by Ant-Miner are considerably simpler (smaller) than those discovered by CN2.
Rafael S. Parpinelli, Heitor Silvério Lopes, Alex Alves Freitas
IEEE Trans. Evol. Comput.3
2001 Discovering Fuzzy Classification Rules with Genetic Programming and Co-evolution
Roberto R. F. Mendes, Fabricio de B. Voznika, Alex Alves Freitas, Júlio C. Nievola
PKDD3
2000 Discovering comprehensible classification rules with a genetic algorithm
abstract
Presents a classification algorithm based on genetic algorithms (GAs) that discovers comprehensible IF-THEN rules, in the spirit of data mining. The proposed GA has a flexible chromosome encoding, where each chromosome corresponds to a classification rule. Although the number of genes (the genotype) is fixed, the number of rule conditions (the phenotype) is variable. The GA also has specific mutation operators for this chromosome encoding. The algorithm was evaluated on two public-domain real-world data sets (in the medical domains of dermatology and breast cancer).
M. V. Fidelis, Heitor Silvério Lopes, Alex Alves Freitas
CEC3
2000 A hybrid decision tree/genetic algorithm for coping with the problem of small disjuncts in data mining
Deborah R. Carvalho, Alex Alves Freitas
GECCO2
2000 Comparing a Genetic Algorithm with a Rule Induction Algorithm in the Data Mining Task of Dependence Modeling
Edgar Noda, Alex Alves Freitas, Heitor Silvério Lopes
GECCO2
2000 A Genetic Algorithm-Based Solution for the Problem of Small Disjuncts
abstract
In essence, small disjuncts are rules covering a small number of examples. Hence, these rules are usually error-prone, which contributes to a decrease in predictive accuracy. The problem is particularly serious because, although each small disjuncts covers few examples, the set of small disjuncts can cover a large number of examples. This paper proposes a solution to the problem of discovering accurate small-disjunct rules based on genetic algorithms. The basic idea of our method is to use a hybrid decision tree / genetic algorithm approach for classification. More precisely, examples belonging to large disjuncts are classified by rules produced by a decision-tree algorithm, while examples belonging to small disjuncts are classified by a new genetic algorithm, particularly designed for discovering small-disjunct rules. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Deborah R. Carvalho, Alex Alves Freitas
PKDD2
1999 Discovering comprehensible classification rules by using Genetic Programming: a case study in a medical domain
Celia C. Bojarczuk, Heitor Silvério Lopes, Alex Alves Freitas
GECCO3
1999 A Fuzzy Beam-Search Rule Induction Algorithm
abstract
This paper proposes a fuzzy beam search rule induction algorithm for the classification task. The use of fuzzy logic and fuzzy sets not only provides us with a powerful, flexible approach to cope with uncertainty, but also allows us to express the discovered rules in a representation more intuitive and comprehensible for the user, by using linguistic terms (such as low, medium, high) rather than continuous, numeric values in rule conditions. The proposed algorithm is evaluated in two public domain data sets. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Christina S. Fertig, Alex Alves Freitas, Lúcia V. R. Arruda, Celso A. A. Kaestner
PKDD2
1999 Interfacing knowledge discovery algorithms to large database management systems
Simon H. Lavington, Neil Dewhurst, Elwood Wilkins, Alex Alves Freitas
Inf. Softw. Technol.4
1999 On rule interestingness measures
Alex Alves Freitas
Knowl. Based Syst.1
1998 On Objective Measures of Rule Surprisingness
Alex Alves Freitas
PKDD1
1998 Scalable, High-Performance Data Mining with Parallel Processing
Alex Alves Freitas
PKDD1
1997 The Pronciple of Transformation between Efficiency and Effectiveness: Towards a Fair Evaluation of the Cost-Effectiveness of KDD Techniques
Alex Alves Freitas
PKDD1