VLDB 2026 Research / reviewers in the wild / expert
Alberto Fernández 0001
dblp:30/2816-1
· DBLP profile ↗
71ranked-venue papers
19as first author
9since 2021 · last 2025
0000-0002-6480-8434ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 16 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Overlap Number of Balls Model-Agnostic CounterFactuals (ONB-MACF): A data-morphology-based counterfactual generation method for trustworthy artificial intelligence
José Daniel Pascual-Triana, Alberto Fernández 0001, Javier Del Ser, Francisco Herrera |
Inf. Sci. | 2 |
| 2024 | A Wearable Eye-Tracking Approach for Early Autism Detection with Machine Learning: Unravelling Challenges and OpportunitiesabstractEarly detection of Autism Spectrum Disorder (ASD) is crucial to facilitate timely interventions, improve outcomes, and enhance the quality of life for individuals on the spectrum. Artificial intelligence and machine learning have significantly advanced the study of ASD by enabling sophisticated analyses of complex behavioural data, providing more accurate and timely detection methods. Nevertheless, existing research is mostly focused on either screening manual methods via psychological surveys, and/or the application of functional Magnetic Resonance Imaging. In the former case, there is a limitation due the subjectivity of the procedure; whereas in the latter the high cost represents a barrier to widespread use in diagnosis. In consideration of the aforementioned factors, this research contributes significantly to the field by integrating commodity wearable technology with psycho-healthcare, employing eye-tracking capabilities as a novel, non-invasive approach for early ASD diagnosis. Our methodology stands out by offering a comprehensive, four-step process that includes sophisticated image preprocessing, innovative feature extraction, precise classification, and a novel multi-instance aggregation strategy. We show its potential via a compelling case study that attests to the feasibility and efficacy of this innovative paradigm. Additionally, the work underscores prevailing challenges in the field, stressing some factors such as the acquisition of extensive and diverse datasets to allow for a multi-modal approach and the application of Trustworthy Artificial Intelligence. J. Lopez-Martinez, Purificación Checa, José M. Soto-Hidalgo, Isaac Triguero, Alberto Fernández 0001 |
IJCNN | 5 |
| 2023 | Fuzzy Rule-Based Explainer Systems for Deep Neural Networks: From Local Explainability to Global UnderstandingabstractExplainability of deep neural networks has been receiving increasing attention with regard to auditability and trustworthiness purposes. Of the various post-hoc explainability approaches, rule extraction methods assist to understand the logic that underpins their functioning. Whereas the rule-based solutions are directly managed and understood by practitioners, the use of intervals or crisp values in the antecedents that rely on numerical values might not be intuitive enough. In this case, the benefits of a linguistic representation based on fuzzy sets/rules are straightforward, as these semantically meaningful components ease the model understanding. This article proposes fuzzy rule-based explainer systems for deep neural networks. The algorithm learns a compact yet accurate set of fuzzy rules based on features' importance (i.e., attribution values) distilled from the trained networks. These systems can be used for both local and global explainability purposes. The evaluation results of different applications revealed that the fuzzy explainers maintained the fidelity and accuracy of the original deep neural networks while implying lower complexity and better comprehensibility. Fatemeh Aghaeipoor, Mohammad Sabokrou, Alberto Fernández 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2022 | The impact of heterogeneous distance functions on missing data imputation and classification performance
Miriam Seoane Santos, Pedro H. Abreu, Alberto Fernández 0001, Julián Luengo, João A. M. Santos |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | An efficiency curve for evaluating imbalanced classifiers considering intrinsic data characteristics: Experimental analysisabstractBalancing the accuracy rates of the majority and minority classes is challenging in imbalanced classification. Furthermore, data characteristics have a significant impact on the performance of imbalanced classifiers, which are generally neglected by existing evaluation methods. The objective of this study is to introduce a new criterion to comprehensively evaluate imbalanced classifiers. Specifically, we introduce an efficiency curve that is established using data envelopment analysis without explicit inputs (DEA-WEI), to determine the trade-off between the benefits of improved minority class accuracy and the cost of reduced majority class accuracy. In sequence, we analyze the impact of the imbalanced ratio and typical imbalanced data characteristics on the efficiency of the classifiers. Empirical analyses using 68 imbalanced data reveal that traditional classifiers such as C4.5 and the k-nearest neighbor are more effective on disjunct data, whereas ensemble and undersampling techniques are more effective for overlapping and noisy data. The efficiency of cost-sensitive classifiers decreases dramatically when the imbalanced ratio increases. Finally, we investigate the reasons for the different efficiencies of classifiers on imbalanced data and recommend steps to select appropriate classifiers for imbalanced data based on data characteristics. Xiangrui Chao, Gang Kou, Yi Peng 0001, Alberto Fernández 0001 |
Inf. Sci. | 4 |
| 2022 | FW-SMOTE: A feature-weighted oversampling approach for imbalanced classification
Sebastián Maldonado 0001, Carla Vairetti, Alberto Fernández 0001, Francisco Herrera |
Pattern Recognit. | 3 |
| 2022 | IFC-BD: An Interpretable Fuzzy Classifier for Boosting Explainable Artificial Intelligence in Big DataabstractIn current Data Science applications, the course of action has derived to adapt the system behavior for the human cognition, resulting in the emerging area of explainable artificial intelligence. Among different classification paradigms, those based on fuzzy rules are suitable solutions to stress the interpretability of the global systems. However, in case of addressing Big Data analytics, they may comprise an excessive number of rules and/or linguistic labels that not only may cause losing the system performance but also may affect the system semantic as well as the system interpretability. In this article, we propose IFC-BD, an interpretable fuzzy classifier for Big Data, aiming at boosting the horizons of explainability by learning a compact yet accurate fuzzy model. IFC-BD is developed in a cell-based distributed framework through the three working stages of initial rule learning, rule generalization, and heuristic rule selection. This whole procedure allows reaching from a high number of specific rules to less number of more general and confident rules. Additionally, in order to resolve possible rules conflict, a new estimated rule weight is proposed specifically for big data problems. IFC-BD was evaluated in comparison to the state-of-the-art approaches of the fuzzy classification paradigm, considering interpretability, accuracy, and running time. The findings of the experiments revealed that the proposed algorithm was able to improve the explainability of fuzzy rule-based classifiers as well as their predictive performance. Fatemeh Aghaeipoor, Mohammad Masoud Javidi, Alberto Fernández 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Learning interpretable multi-class models by means of hierarchical decomposition: Threshold Control for Nested Dichotomies
J. A. Fdez-Sánchez, José Daniel Pascual-Triana, Alberto Fernández 0001, Francisco Herrera |
Neurocomputing | 3 |
| 2021 | Revisiting data complexity metrics based on morphology for overlap and imbalance: snapshot, new overlap number of balls metrics and singular problems prospectabstractData Science and Machine Learning have become fundamental assets for companies and research institutions alike. As one of its fields, supervised classification allows for class prediction of new samples, learning from given training data. However, some properties can cause datasets to be problematic to classify. In order to evaluate a dataset a priori, data complexity metrics have been used extensively. They provide information regarding different intrinsic characteristics of the data, which serve to evaluate classifier compatibility and a course of action that improves performance. However, most complexity metrics focus on just one characteristic of the data, which can be insufficient to properly evaluate the dataset towards the classifiers' performance. In fact, class overlap, a very detrimental feature for the classification process (especially when imbalance among class labels is also present) is hard to assess. This research work focuses on revisiting complexity metrics based on data morphology. In accordance to their nature, the premise is that they provide both good estimates for class overlap, and great correlations with the classification performance. For that purpose, a novel family of metrics have been developed. Being based on ball coverage by classes, they are named after Overlap Number of Balls. Finally, some prospects for the adaptation of the former family of metrics to singular (more complex) problems are discussed. José Daniel Pascual-Triana, David Charte, Marta Andrés Arroyo, Alberto Fernández 0001, Francisco Herrera |
Knowl. Inf. Syst. | 4 |
| 2020 | Chi-BD-DRF: Design of Scalable Fuzzy Classifiers for Big Data via A Dynamic Rule Filtering ApproachabstractBig data classification problems are known to be no longer addressable by sequential algorithms. Therefore, it is necessary to design and develop novel solutions to provide accurate yet interpretable models in a tolerable elapsed time. In this area, Fuzzy Rule-Based Classification Systems are very advantageous due to their intrinsic interpretable and accurate capabilities. However, when these systems are applied in Big Data scenarios, the size of the rule set can become too large to be useful, whereas many of the generated rules could be associated with the non-dense areas or outliers. The presence of such rules in the rule base not only increases the running time and computation overheads but also affects on the interpretability of the fuzzy system. In this contribution, we propose a novel approach to obtain compact and accurate fuzzy models for Big data problems in a linearly scalable complex time. To do so, a dynamic filtering approach is applied to remove low supporting rules. Moreover, an efficient computation of the rules' weights is presented to improve the accuracy of the predictions. This model is developed for Big Data analytics by using Apache Spark framework. This allows taking advantage of the built-in resources and directives for a transparent distributed computing, as well as the machine learning pipeline to ease the complete processing. Experimental results, using different Big Data problems, confirmed the goodness of the proposed algorithm with respect to the baseline fuzzy classifier. Fatemeh Aghaeipoor, Mohammad Masoud Javidi, Isaac Triguero, Alberto Fernández 0001 |
FUZZ-IEEE | 4 |
| 2020 | HFER: Promoting Explainability in Fuzzy Systems via Hierarchical Fuzzy Exception RulesabstractWhen developing a Machine Learning model, the consideration of explainability as an additional design driver can improve its deployment into any application context. Given an audience, an explainable Artificial Intelligence system is one that produces details or reasons to make it's functioning clear or easy to understand. Among different paradigms that inherently support these capabilities, Fuzzy Rule Based Systems are a very accountable solution. The main issue when dealing with fuzzy systems is to select an appropriate granularity to represent (fuzzify) the input data. A low value may cause the generation of too generalist rules, causing a hinder on predictive performance, whereas a high value may lead to both overfitting and/or very complex solutions. To overcome this situation, we propose a novel hierarchical fuzzy classification system based on fuzzy exception rules. To do so, low granularity rules are first generated and their confidence is examined. For those cases in which the fuzzy confidence is below a quality threshold, new higher granularity rules are created to cover the "instances in conflict" for the general rule, which is still kept in the rule base. Experimental results show the achievement of a compact and interpretable final rule base while maintaining or improving the predictive performance in comparison with the baseline fuzzy rule based classification and hierarchical systems. José Ramón Trillo, Alberto Fernández 0001, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2020 | Discussion on Vuttipittayamongkol, P. and Elyan, E., Improved Overlap-Based Undersampling for Imbalanced Dataset Classification with Application to Epilepsy and Parkinson's Disease
Alberto Fernández 0001 |
Int. J. Neural Syst. | 1 |
| 2019 | On the Need of Interpretability for Biomedical Applications: Using Fuzzy Models for Lung Cancer Prediction with Liquid BiopsyabstractIn the latter years, we are witnessing a movement from the standard Data Mining towards a more profitable and challenging scenario known as Data Science. It can be defined as a set of quantitative and qualitative approaches that are applied to current relevant problems. In order to be able to "dig" to the deepest level considering the whole information available, the knowledge domain and the analysis of the data must have a strong synergy.There are many fields of application where it is necessary, if not essential, to give an explanation of the phenomenon under study. It is no longer enough to simply apply a Machine Learning model, but it must be comprehensible in order to provide a real decision support system. For this reason, a strong movement has emerged in favour of the eXplainable Artificial Intelligence that aims to respond to the "how" and "why" of the operation of automatic models.In this work, our objective is to show the benefits of one of the learning paradigms of Computational Intelligence: Fuzzy Rule Based Systems and Evolutionary Fuzzy Systems. To this end, we focus on biomedical applications by presenting a case study based on lung cancer prediction from samples taken by liquid biopsy. Liquid biopsy enable us to study genomic alterations for each individual independently, a step towards personalised medicine. The results show the goodness of the solution based on Evolutionary Fuzzy Systems in terms of interpretability and comprehensibility, obtaining a low number of rules with less than 3 fuzzy linguistic labels per antecedent. Nicolas Potie, Stavros Giannoukakos, Michael Hackenberg, Alberto Fernández 0001 |
FUZZ-IEEE | 4 |
| 2019 | A multi-objective evolutionary fuzzy system to obtain a broad and accurate set of solutions in intrusion detection systems
Salma Elhag, Alberto Fernández 0001, Abdulrahman H. Altalhi, Saleh Alshomrani, Francisco Herrera |
Soft Comput. | 2 |
| 2019 | A Metahierarchical Rule Decision System to Design Robust Fuzzy Classifiers Based on Data ComplexityabstractThere is a wide variety of studies that propose different classifiers to solve a large amount of problems in distinct classification scenarios. The no free lunch theorem states that if we use a big enough set of varied problems, all classifiers would be equivalent in performance. From another point of view, the performance of the classifiers is dependant of the scope and properties of the datasets. In this sense, new proposals on the topic often focus on a given context, aiming at improving the related state-of-the-art approaches. Data complexity metrics have been traditionally used to determine the inner characteristics of datasets. This way, researchers are able to categorize the problems in different scenarios. Then, this taxonomy can be applied to determine inner characteristics of the datasets in order to determine intervals of good and bad behavior for a given classifier. In this paper, we will take advantage of the data complexity metrics in order to design a fuzzy metaclassifier. The final goal is to create decision rules based on the inner characteristics of the data to apply a different version of the fuzzy classifier for a given problem. To do so, we will make use of the FARC-HD classifier, an evolutionary fuzzy system that has led to different extensions in the specialized literature. Experimental results show the goodness of this novel approach as it is able to outperform all versions of FARC-HD on a wide set of problems, and obtain competitive results (in terms of performance and interpretability) versus two selected state-of-the-art rule-based classification system, C4.5 and FURIA. Javier Cózar, Alberto Fernández 0001, Francisco Herrera, José A. Gámez 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2018 | Improving Fuzzy Rule Based Classification Systems in Big Data via Support-based FilteringabstractFuzzy Rule Based Classification Systems have the benefit of making possible to understand the decision of the classifier. Additionally, they have shown to be robust to solve complex problems. When these capabilities are applied to the context of Big Data, the benefits get multiplied.Therefore, to achieve the highest advantages of Fuzzy Rule Based Classification Systems, the output model must be both interpretable and accurate. The former is achieved by using fuzzy linguistic labels, that are related to human understanding. The latter is achieved by means of robust fuzzy rules, which are identified by means of a component known as "fuzzy rule weight". However, obtaining these rule weights is computationally expensive, resulting on a bottle-neck when applied in Big Data problems.In this work, we propose Chi-BD-SF, which stands for Chi Big Data Support Filtering. It comprises a scalable yet accurate fuzzy rule learning algorithm. It is based on the well-known Chi et al., exchanging the rule weight computation by a support metric in order to solve the conflicts between different consequent rules. In order to show the goodness of this proposal, we analyze several performance metrics, such as the quality of classification, the robustness of the rule base generated and the runtimes of the usage of traditional weights and the support of the rule. The results of our novel Chi-BD-SF approach, in contrast to related Big Data fuzzy classifiers, show that this proposal is able to out-speed the usage of rule weights also obtaining more accurate results. Luis Íñiguez, Mikel Galar, Alberto Fernández 0001 |
FUZZ-IEEE | 3 |
| 2018 | Surveying alignment-free features for Ortholog detection in related yeast proteomes by using supervised big data classifiersabstractBACKGROUND: The development of new ortholog detection algorithms and the improvement of existing ones are of major importance in functional genomics. We have previously introduced a successful supervised pairwise ortholog classification approach implemented in a big data platform that considered several pairwise protein features and the low ortholog pair ratios found between two annotated proteomes (Galpert, D et al., BioMed Research International, 2015). The supervised models were built and tested using a Saccharomycete yeast benchmark dataset proposed by Salichos and Rokas (2011). Despite several pairwise protein features being combined in a supervised big data approach; they all, to some extent were alignment-based features and the proposed algorithms were evaluated on a unique test set. Here, we aim to evaluate the impact of alignment-free features on the performance of supervised models implemented in the Spark big data platform for pairwise ortholog detection in several related yeast proteomes. RESULTS: The Spark Random Forest and Decision Trees with oversampling and undersampling techniques, and built with only alignment-based similarity measures or combined with several alignment-free pairwise protein features showed the highest classification performance for ortholog detection in three yeast proteome pairs. Although such supervised approaches outperformed traditional methods, there were no significant differences between the exclusive use of alignment-based similarity measures and their combination with alignment-free features, even within the twilight zone of the studied proteomes. Just when alignment-based and alignment-free features were combined in Spark Decision Trees with imbalance management, a higher success rate (98.71%) within the twilight zone could be achieved for a yeast proteome pair that underwent a whole genome duplication. The feature selection study showed that alignment-based features were top-ranked for the best classifiers while the runners-up were alignment-free features related to amino acid composition. CONCLUSIONS: The incorporation of alignment-free features in supervised big data models did not significantly improve ortholog detection in yeast proteomes regarding the classification qualities achieved with just alignment-based similarity measures. However, the similarity of their classification performance to that of traditional ortholog detection methods encourages the evaluation of other alignment-free protein pair descriptors in future research. Deborah Galpert, Alberto Fernández 0001, Francisco Herrera, Agostinho Antunes, Reinaldo Molina Ruiz, Guillermín Agüero-Chapín |
BMC Bioinform. | 2 |
| 2018 | SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year AnniversaryabstractThe Synthetic Minority Oversampling Technique (SMOTE) preprocessing algorithm is considered "de facto" standard in the framework of learning from imbalanced data. This is due to its simplicity in the design of the procedure, as well as its robustness when applied to different type of problems. Since its publication in 2002, SMOTE has proven successful in a variety of applications from several different domains. SMOTE has also inspired several approaches to counter the issue of class imbalance, and has also significantly contributed to new supervised learning paradigms, including multilabel classification, incremental learning, semi-supervised learning, multi-instance learning, among others. It is standard benchmark for learning from imbalanced data. It is also featured in a number of different software packages - from open source to commercial. In this paper, marking the fifteen year anniversary of SMOTE, we reflect on the SMOTE journey, discuss the current state of affairs with SMOTE, its applications, and also identify the next set of challenges to extend SMOTE for Big Data problems. Alberto Fernández 0001, Salvador García 0001, Francisco Herrera, Nitesh V. Chawla |
J. Artif. Intell. Res. | 1 |
| 2018 | Dynamic affinity-based classification of multi-class imbalanced data with one-versus-one decomposition: a fuzzy rough set approach
Sarah Vluymans, Alberto Fernández 0001, Yvan Saeys, Chris Cornelis, Francisco Herrera |
Knowl. Inf. Syst. | 2 |
| 2018 | Imbalance: Oversampling algorithms for imbalanced classification in R
Ignacio Cordón, Salvador García 0001, Alberto Fernández 0001, Francisco Herrera |
Knowl. Based Syst. | 3 |
| 2017 | Chi-Spark-RS: An Spark-built evolutionary fuzzy rule selection algorithm in imbalanced classification for big data problemsabstractThe significance and benefits of addressing classification tasks in Big Data applications is beyond any doubt. To do so, learning algorithms must be scalable to cope with such a high volume of data. The most suitable option to reach this objective is by using a MapReduce programming scheme, in which algorithms are automatically executed in a distributed and fault tolerant way. Among different available tools that support this framework, Spark has emerged as a “de facto” solution when using iterative approaches. In this work, our goal is to design and implement an Evolutionary Fuzzy Rule Selection algorithm within a Spark environment. To do so, we build different local rule bases within each Map Task that are later optimized by means of a genetic process. With this procedure, we seek to minimize the total number of rules that are gathered by each Reduce task to obtain a compact and accurate Fuzzy Rule Based Classification System. In particular, we set the experimental framework in the scenario of imbalanced classification. Therefore, the final objective will be analyzing the best synergy between the novel Evolutionary Fuzzy Rule Selection algorithm and the solutions applied to cope with skewed class distributions, namely cost-sensitive learning, random under-sampling and random-oversampling. Alberto Fernández 0001, Eva Almansa, Francisco Herrera |
FUZZ-IEEE | 1 |
| 2017 | A Pareto-based Ensemble with Feature and Instance Selection for Learning from Multi-Class Imbalanced DatasetsabstractImbalanced classification is related to those problems that have an uneven distribution among classes. In addition to the former, when instances are located into the overlapped areas, the correct modeling of the problem becomes harder. Current solutions for both issues are often focused on the binary case study, as multi-class datasets require an additional effort to be addressed. In this research, we overcome these problems by carrying out a combination between feature and instance selections. Feature selection will allow simplifying the overlapping areas easing the generation of rules to distinguish among the classes. Selection of instances from all classes will address the imbalance itself by finding the most appropriate class distribution for the learning task, as well as possibly removing noise and difficult borderline examples. For the sake of obtaining an optimal joint set of features and instances, we embedded the searching for both parameters in a Multi-Objective Evolutionary Algorithm, using the C4.5 decision tree as baseline classifier in this wrapper approach. The multi-objective scheme allows taking a double advantage: the search space becomes broader, and we may provide a set of different solutions in order to build an ensemble of classifiers. This proposal has been contrasted versus several state-of-the-art solutions on imbalanced classification showing excellent results in both binary and multi-class problems. Alberto Fernández 0001, Cristóbal J. Carmona, María José del Jesus, Francisco Herrera |
Int. J. Neural Syst. | 1 |
| 2016 | A First Approach in Evolutionary Fuzzy Systems based on the lateral tuning of the linguistic labels for Big Data classificationabstractThe treatment and processing of Big Data problems imply an essential advantage for researchers and corporations. This is due to the huge quantity of knowledge that is hidden within the vast amount of information that is available nowadays. In order to be able to address with such volume of information in an efficient way, the scalability for Big Data applications is achieved by means of the MapReduce programming model. It is designed to divide the data into several chunks or groups that are processed in parallel, and whose result is “assembled” to provide a single solution. Alberto Fernández 0001, Sara del Río, Francisco Herrera |
FUZZ-IEEE | 1 |
| 2016 | Enhancing evolutionary fuzzy systems for multi-class problems: Distance-based relative competence weighting with truncated confidences (DRCW-TC)
Alberto Fernández 0001, Mikel Elkano, Mikel Galar, José Antonio Sanz 0001, Saleh Alshomrani, Humberto Bustince, Francisco Herrera |
Int. J. Approx. Reason. | 1 |
| 2016 | Ordering-based pruning for improving the performance of ensembles of classifiers in the framework of imbalanced datasets
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera |
Inf. Sci. | 2 |
| 2015 | On the impact of Distance-based Relative Competence Weighting approach in One-vs-One classification for Evolutionary Fuzzy Systems: DRCW-FH-GBML algorithmabstractThe advantages of multi-classification schemes based on decomposition strategies, and especially the One-vs-One framework, have been stressed even for those algorithms that can address multiple classes. However, there is an inherent hitch for the One-vs-One learning scheme related to the decision process: the non-competent classifier problem. This issue refers to the case where a binary classifier outputs a score degree for a couple of classes that are not related with the input example, thus including “noise” in the score-matrix and degrading the final accuracy. For this reason, several approaches have been developed in order to address the influence of the non-competence. Among them, the distance-based combination strategy has excelled as a very robust solution. In this contribution, we aim at investigating the behaviour of this approach using Evolutionary Fuzzy Systems as baseline classifiers. We will show that the synergy between both methodologies allows a significant improvement of the results to be obtained in contrast to the standard classifier and the classical One-vs-One scheme. Alberto Fernández 0001, Mikel Galar, José Antonio Sanz 0001, Humberto Bustince, Oscar Cordón, Francisco Herrera |
FUZZ-IEEE | 1 |
| 2015 | Improving the OVO performance in Fuzzy Rule-Based Classification Systems by the genetic learning of the granularity levelabstractThis contribution proposes a genetic learning process for designing the knowledge base of Fuzzy Rule-Based classification Systems, that will be used as binary classifiers in a One-vs-One decomposition for multi-class problems. A Genetic Algorithm is designed to adapt the number of fuzzy labels per variable (granularity level) for each classifier in order to improve the accuracy rate of a multi-class classifier. The genetic learning process evolves granularity levels and needs a fuzzy rules generation method for generating the whole knowledge base of the Fuzzy System. Several data-sets from KEEL data-set repository are used in the experimental study and we compare our proposal with three related methods: the standard way to design Fuzzy Rule-Based Classification Systems using the fuzzy rules generation method chosen with and without One-vs-One decomposition, and our proposal of genetic granularity level learning without One-vs-One decomposition. Pedro Villar, Alberto Fernández 0001, Rosana Montes-Soldado, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2015 | Addressing Overlapping in Classification with Imbalanced Datasets: A First Multi-objective Approach for Feature and Instance Selection
Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
IDEAL | 1 |
| 2015 | On the combination of genetic fuzzy systems and pairwise learning for improving detection rates on Intrusion Detection Systems
Salma Elhag, Alberto Fernández 0001, Abdullah Bawakid, Saleh Alshomrani, Francisco Herrera |
Expert Syst. Appl. | 2 |
| 2015 | A proposal for evolutionary fuzzy systems using feature weighting: Dealing with overlapping in imbalanced datasets
Saleh Alshomrani, Abdullah Bawakid, Seong-O Shim, Alberto Fernández 0001, Francisco Herrera |
Knowl. Based Syst. | 4 |
| 2015 | Revisiting Evolutionary Fuzzy Systems: Taxonomy, applications, new trends and challenges
Alberto Fernández 0001, Victoria López, María José del Jesus, Francisco Herrera |
Knowl. Based Syst. | 1 |
| 2015 | DRCW-OVO: Distance-based relative competence weighting combination for One-vs-One strategy in multi-class problems
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Francisco Herrera |
Pattern Recognit. | 2 |
| 2015 | Enhancing Multiclass Classification in FARC-HD Fuzzy Classifier: On the Synergy Between $n$-Dimensional Overlap Functions and Decomposition StrategiesabstractThere are many real-world classification problems involving multiple classes, e.g., in bioinformatics, computer vision, or medicine. These problems are generally more difficult than their binary counterparts. In this scenario, decomposition strategies usually improve the performance of classifiers. Hence, in this paper, we aim to improve the behavior of fuzzy association rule-based classification model for high-dimensional problems (FARC-HD) fuzzy classifier in multiclass classification problems using decomposition strategies, and more specifically One-versus-One (OVO) and One-versus-All (OVA) strategies. However, when these strategies are applied on FARC-HD, a problem emerges due to the low-confidence values provided by the fuzzy reasoning method. This undesirable condition comes from the application of the product t-norm when computing the matching and association degrees, obtaining low values, which are also dependent on the number of antecedents of the fuzzy rules. As a result, robust aggregation strategies in OVO, such as the weighted voting obtain poor results with this fuzzy classifier. In order to solve these problems, we propose to adapt the inference system of FARC-HD replacing the product t-norm with overlap functions. To do so, we define n-dimensional overlap functions. The usage of these new functions allows one to obtain more adequate outputs from the base classifiers for the subsequent aggregation in OVO and OVA schemes. Furthermore, we propose a new aggregation strategy for OVO to deal with the problem of the weighted voting derived from the inappropriate confidences provided by FARC-HD for this aggregation method. The quality of our new approach is analyzed using 20 datasets and the conclusions are supported by a proper statistical analysis. In order to check the usefulness of our proposal, we carry out a comparison against some of the state-of-the-art fuzzy classifiers. Experimental results show the competitiveness of our method. Mikel Elkano, Mikel Galar, José Antonio Sanz 0001, Alberto Fernández 0001, Edurne Barrenechea Tartas, Francisco Herrera, Humberto Bustince |
IEEE Trans. Fuzzy Syst. | 4 |
| 2014 | Enhancing difficult classes in one-vs-one classifier fusion strategy using restricted equivalence functions
Mikel Galar, Edurne Barrenechea Tartas, Alberto Fernández 0001, Francisco Herrera |
FUSION | 3 |
| 2014 | Empowering difficult classes with a similarity-based aggregation in multi-class classification problems
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Francisco Herrera |
Inf. Sci. | 2 |
| 2014 | On the importance of the validation technique for classification with imbalanced datasets: Addressing covariate shift when data is skewed
Victoria López, Alberto Fernández 0001, Francisco Herrera |
Inf. Sci. | 2 |
| 2013 | Addressing covariate shift for Genetic Fuzzy Systems classifiers: A case of study with FARC-HD for imbalanced datasetsabstractThe estimation of the quality of the learned models in Data Mining has been traditionally carried out by means of a k-fold partition technique. However, the “random” division of the instances over the folds may results in a problem known as covariate shift, i.e. there is a different data distribution between the training and test folds. In classification with imbalanced datasets this problem is more severe. The misclassification of minority class instances due to an incorrect learning of the real boundaries caused by a not well defined data distribution, truly affects the measures of performance in this scenario. To avoid this harmful situation, we propose the use of a specific validation technique for the partitioning of the data, known as “Distribution optimally balanced stratified cross-validation”. This methodology makes the decision of placing close-by samples on different folds, so that each partition will end up with enough representatives of every region. In this contribution, we show the goodness of this methodology using Genetic Fuzzy Systems, as they are known to be robust approaches for all types of classification problems. Specifically, we have chosen the FARC-HD algorithm, a novel technique which has shown to obtain very accurate results. From the experimental analysis, which is carried out on a wide number of imbalanced datasets, we emphasize the necessity of using a proper validation methodology for extracting well founded conclusions. Victoria López, Alberto Fernández 0001, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2013 | An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics
Victoria López, Alberto Fernández 0001, Salvador García 0001, Vasile Palade, Francisco Herrera |
Inf. Sci. | 2 |
| 2013 | Analysing the classification of imbalanced data-sets with multiple classes: Binarization techniques and ad-hoc approaches
Alberto Fernández 0001, Victoria López, Mikel Galar, María José del Jesus, Francisco Herrera |
Knowl. Based Syst. | 1 |
| 2013 | A hierarchical genetic fuzzy system based on genetic programming for addressing classification with highly imbalanced and borderline data-sets
Victoria López, Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
Knowl. Based Syst. | 2 |
| 2013 | EUSBoost: Enhancing ensembles for highly imbalanced data-sets by evolutionary undersampling
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Francisco Herrera |
Pattern Recognit. | 2 |
| 2013 | Dynamic classifier selection for One-vs-One strategy: Avoiding non-competent classifiers
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera |
Pattern Recognit. | 2 |
| 2013 | IVTURS: A Linguistic Fuzzy Rule-Based Classification System Based On a New Interval-Valued Fuzzy Reasoning Method With Tuning and Rule SelectionabstractInterval-valued fuzzy sets have been shown to be a useful tool to deal with the ignorance related to the definition of the linguistic labels. Specifically, they have been successfully applied to solve classification problems, performing simple modifications on the fuzzy reasoning method to work with this representation and making the classification based on a single number. In this paper, we present IVTURS, which is a new linguistic fuzzy rule-based classification method based on a new completely interval-valued fuzzy reasoning method. This inference process uses interval-valued restricted equivalence functions to increase the relevance of the rules in which the equivalence of the interval membership degrees of the patterns and the ideal membership degrees is greater, which is a desirable behavior. Furthermore, their parametrized construction allows the computation of the optimal function for each variable to be performed, which could involve a potential improvement in the system's behavior. Additionally, we combine this tuning of the equivalence with rule selection in order to decrease the complexity of the system. In this paper, we name our method IVTURS-FARC, since we use the FARC-HD method to accomplish the fuzzy rule learning process. The experimental study is developed in three steps in order to ascertain the quality of our new proposal. First, we determine both the essential role that interval-valued fuzzy sets play in the method and the need for the rule selection process. Next, we show the improvements achieved by IVTURS-FARC with respect to the tuning of the degree of ignorance when it is applied in both an isolated way and when combined with the tuning of the equivalence. Finally, the significance of IVTURS-FARC is further depicted by means of a comparison by which it is proved to outperform the results of FARC-HD and FURIA, which are two high performing fuzzy classification algorithms. José Antonio Sanz 0001, Alberto Fernández 0001, Humberto Bustince, Francisco Herrera |
IEEE Trans. Fuzzy Syst. | 2 |
| 2012 | Cost Sensitive and Preprocessing for Classification with Imbalanced Data-sets: Similar Behaviour and Potential Hybridizations
Victoria López, Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
ICPRAM (2) | 2 |
| 2012 | Analysis of preprocessing vs. cost-sensitive learning for imbalanced classification. Open problems on intrinsic data characteristics
Victoria López, Alberto Fernández 0001, Jose G. Moreno-Torres, Francisco Herrera |
Expert Syst. Appl. | 2 |
| 2012 | Feature Selection and Granularity Learning in Genetic Fuzzy Rule-Based Classification Systems for Highly Imbalanced Data-SetsabstractThis paper proposes a Genetic Algorithm for jointly performing a feature selection and granularity learning for Fuzzy Rule-Based Classification Systems in the scenario of highly imbalanced data-sets. We refer to imbalanced data-sets when the class distribution is not uniform, a situation that it is present in many real application areas. The aim of this work is to get more compact models by selecting the adequate variables and adapting the number of fuzzy labels for each problem, improving the interpretability of the model. The experimental analysis is carried out over a wide range of highly imbalanced data-sets and uses the statistical tests suggested in the specialized literature. Pedro Villar, Alberto Fernández 0001, Ramón Alberto Carrasco, Francisco Herrera |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2012 | A Review on Ensembles for the Class Imbalance Problem: Bagging-, Boosting-, and Hybrid-Based ApproachesabstractClassifier learning with data-sets that suffer from imbalanced class distributions is a challenging problem in data mining community. This issue occurs when the number of examples that represent one class is much lower than the ones of the other classes. Its presence in many real-world applications has brought along a growth of attention from researchers. In machine learning, the ensemble of classifiers are known to increase the accuracy of single classifiers by combining several of them, but neither of these learning techniques alone solve the class imbalance problem, to deal with this issue the ensemble learning algorithms have to be designed specifically. In this paper, our aim is to review the state of the art on ensemble techniques in the framework of imbalanced data-sets, with focus on two-class problems. We propose a taxonomy for ensemble-based methods to address the class imbalance where each proposal can be categorized depending on the inner ensemble methodology in which it is based. In addition, we develop a thorough empirical comparison by the consideration of the most significant published approaches, within the families of the taxonomy proposed, to show whether any of them makes a difference. This comparison has shown the good behavior of the simplest approaches which combine random undersampling techniques with bagging or boosting ensembles. In addition, the positive synergy between sampling techniques and bagging has stood out. Furthermore, our results show empirically that ensemble-based algorithms are worthwhile since they outperform the mere use of preprocessing techniques before learning the classifier, therefore justifying the increase of complexity by means of a significant enhancement of the results. Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2011 | On the cooperation of interval-valued fuzzy sets and genetic tuning to improve the performance of fuzzy decision treesabstractFuzzy decision trees are widely employed to face classification problems since they combine the high interpretability given by the decision tree and the capability of management of the uncertainty inherent to fuzzy logic. However, the success of fuzzy systems in general depends, to a large degree, on the choice of the membership functions. For this reason, we propose to model the linguistic labels by means of Interval-Valued Fuzzy Sets to take into account the ignorance related to their definition. On the other hand, we define an evolutionary method to tune the shape of the Interval-Valued Fuzzy Sets looking for the best ignorance degree that each Interval-Valued Fuzzy Set represents. In this contribution, we will make use of the fuzzy ID3 algorithm as a base technique from which to apply our methodology. The experimental study shows how our methodology enhances the performance of the base fuzzy decision tree. Furthermore, we compare our approach with respect to four state-of-the-art fuzzy decision trees and C4.5 as a representative algorithm for crisp decision trees. The goodness of our proposal is tested on a large collection of data-sets and it is supported by an exhaustive statistical analysis. José Antonio Sanz 0001, Humberto Bustince, Alberto Fernández 0001, Francisco Herrera |
FUZZ-IEEE | 3 |
| 2011 | Studying the behavior of a multiobjective genetic algorithm to design fuzzy rule-based classification systems for imbalanced data-setsabstractThis paper studies the behavior of a multiobjective Genetic Algorithm for jointly performing a feature selection and granularity learning for Fuzzy Rule-Based Classification Systems in the scenario of imbalanced data-sets. We refer to imbalanced data-sets when the class distribution is not uniform, a situation that it is present in many real application areas. We consider two different measures, one for the precision of the model and other for its complexity as the two objectives to optimize. In one previous approach, we aggregate these two measures in a single-objective Genetic Algorithm, and thus, a multiobjective approach of that Genetic Algorithm would yield a set of models with different trade-off between high accuracy and low complexity rather than a unique model, provided by the single-objective Genetic Algorithm. The experimental analysis, carried out over a wide range of imbalanced data-sets, shows that our approach is able to obtain a set of models with good trade-off between the two objectives considered but it is an open problem how to select the solution with best prediction ability from the whole set of solutions obtained. Pedro Villar, Alberto Fernández 0001, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2011 | A genetic tuning to improve the performance of Fuzzy Rule-Based Classification Systems with Interval-Valued Fuzzy Sets: Degree of ignorance and lateral position
José Antonio Sanz 0001, Alberto Fernández 0001, Humberto Bustince, Francisco Herrera |
Int. J. Approx. Reason. | 2 |
| 2011 | An overview of ensemble methods for binary classifiers in multi-class problems: Experimental study on one-vs-one and one-vs-all schemes
Mikel Galar, Alberto Fernández 0001, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera |
Pattern Recognit. | 2 |
| 2011 | Addressing data complexity for imbalanced data sets: analysis of SMOTE-based oversampling and evolutionary undersampling
Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera |
Soft Comput. | 2 |
| 2010 | A genetic algorithm for tuning fuzzy rule-based classification systems with Interval-Valued Fuzzy SetsabstractFuzzy Rule-Based Classification Systems are a widely used tool in Data Mining because of the interpretability given by the concept of linguistic label. However, the use of this type of models implies a degree of uncertainty in the definition of the fuzzy partitions. In this work we will use the concept of Interval-Valued Fuzzy Set to deal with this problem. The aim of this contribution is to show the improvement in the performance of linguistic Fuzzy Rule-Based Classification Systems afterward the application of a cooperative tuning methodology between the tuning of the amplitude of the support and the lateral tuning (based on the 2-tuples fuzzy linguistic model) applied to the linguistic labels modeled with Interval-Valued Fuzzy Sets. José Antonio Sanz 0001, Alberto Fernández 0001, Humberto Bustince, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2010 | Multi-class Imbalanced Data-Sets with Linguistic Fuzzy Rule Based Classification Systems Based on Pairwise Learning
Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
IPMU | 1 |
| 2010 | A Genetic Algorithm for Feature Selection and Granularity Learning in Fuzzy Rule-Based Classification Systems for Highly Imbalanced Data-Sets
Pedro Villar, Alberto Fernández 0001, Francisco Herrera |
IPMU (1) | 2 |
| 2010 | A first approach for cost-sensitive classification with linguistic Genetic Fuzzy Systems in imbalanced data-setsabstractClassification in imbalanced domains has become one of the most relevant problems within the area of Machine Learning at the present. This problem has raised in significance due to its presence in many real applications and it occurs when the distribution of the available examples to carry out the learning process is very different between the classes (often for binary class data-sets). Usually, the underrepresented class is the concept of the most interest for the problem, being the cost derived from a misclassification of these examples much higher than that of the remaining examples. In this work we analyze the behaviour of a cost-sensitive learning method for Fuzzy Rule Based Classification Systems in the scenario of high imbalanced data-sets. Specifically, we focus on one representative rule learning approach for Genetic Fuzzy Systems, the Fuzzy Hybrid Genetics-Based Machine Learning algorithm. The experimental results show how our cost-sensitive approach in this type of domains will help us to obtain very accurate solutions in shorter training times and also with a lower complexity with respect to other possibilities proposed for classification with imbalanced problems such as the use of preprocessing to rebalance the class distribution. Victoria López, Alberto Fernández 0001, Francisco Herrera |
ISDA | 2 |
| 2010 | Solving multi-class problems with linguistic fuzzy rule based classification systems based on pairwise learning and preference relations
Alberto Fernández 0001, María Calderón, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera |
Fuzzy Sets Syst. | 1 |
| 2010 | On the 2-tuples based genetic tuning performance for fuzzy rule based classification systems in imbalanced data-sets
Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
Inf. Sci. | 1 |
| 2010 | Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera |
Inf. Sci. | 2 |
| 2010 | Improving the performance of fuzzy rule-based classification systems with interval-valued fuzzy sets and genetic amplitude tuning
José Antonio Sanz 0001, Alberto Fernández 0001, Humberto Bustince, Francisco Herrera |
Inf. Sci. | 2 |
| 2010 | Analysis of an evolutionary RBFN design algorithm, CO2RBFN, for imbalanced data sets
M. Dolores Pérez-Godoy, Alberto Fernández 0001, Antonio J. Rivera, María José del Jesus |
Pattern Recognit. Lett. | 2 |
| 2010 | Genetics-Based Machine Learning for Rule Induction: State of the Art, Taxonomy, and Comparative StudyabstractThe classification problem can be addressed by numerous techniques and algorithms which belong to different paradigms of machine learning. In this paper, we are interested in evolutionary algorithms, the so-called genetics-based machine learning algorithms. In particular, we will focus on evolutionary approaches that evolve a set of rules, i.e., evolutionary rule-based systems, applied to classification tasks, in order to provide a state of the art in this field. This paper has a double aim: to present a taxonomy of the genetics-based machine learning approaches for rule induction, and to develop an empirical analysis both for standard classification and for classification with imbalanced data sets. We also include a comparative study of the genetics-based machine learning (GBML) methods with some classical non-evolutionary algorithms, in order to observe the suitability and high potential of the search performed by evolutionary algorithms and the behavior of the GBML algorithms in contrast to the classical approaches, in terms of classification accuracy. Alberto Fernández 0001, Salvador García 0001, Julián Luengo, Ester Bernadó-Mansilla, Francisco Herrera |
IEEE Trans. Evol. Comput. | 1 |
| 2009 | A genetic learning of the fuzzy rule-based classification system granularity for highly imbalanced data-setsabstractIn this contribution we analyse the significance of the granularity level (number of labels) in Fuzzy Rule-Based Classification Systems in the scenario of data-sets with a high imbalance degree. We refer to imbalanced data-sets when the class distribution is not uniform, a situation that it is present in many real application areas. The aim of this work is to adapt the number of fuzzy labels for each problem, applying a fine granularity in those variables which have a higher dispersion of values and a thick granularity in the variables where an excessive number of labels may result irrelevant. We compare this methodology with the use of a fixed number of labels and with the C4.5 decision tree. Pedro Villar, Alberto Fernández 0001, Francisco Herrera |
FUZZ-IEEE | 2 |
| 2009 | Implementation and Integration of Algorithms into the KEEL Data-Mining Software Tool
Alberto Fernández 0001, Julián Luengo, Joaquín Derrac, Jesús Alcalá-Fdez, Francisco Herrera |
IDEAL | 1 |
| 2009 | Addressing Data-Complexity for Imbalanced Data-Sets: A Preliminary Study on the Use of Preprocessing for C4.5abstractIn this work we analyse the behaviour of the C4.5 classification method with respect to a bunch of imbalanced data-sets. We consider the use of two metrics of data complexity known as “maximum Fishers discriminant ratio” and “nonlinearity of 1NN classifier”, to analyse the effect of preprocessing (oversampling in this case) in order to deal with the imbalance problem. In order to do that, we analyse C4.5 over a wide range of imbalanced data-sets built from real data, and try to extract behaviour patterns from the results. We obtain rules that describe both good or bad behaviours of C4.5 in the case of using the original data-sets (absence of preprocessing) and when applying preprocessing. These rules allow us to determine the effect of the use of preprocessing and to predict the response of C4.5 to preprocessing from the data-set’s complexity metrics prior to its application, and then establish when the preprocessing would be useful to. Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera |
ISDA | 2 |
| 2009 | On the influence of an adaptive inference system in fuzzy rule based classification systems for imbalanced data-sets
Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
Expert Syst. Appl. | 1 |
| 2009 | Hierarchical fuzzy rule based classification systems with genetic rule selection for imbalanced data-sets
Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
Int. J. Approx. Reason. | 1 |
| 2009 | A study of statistical techniques and performance measures for genetics-based machine learning: accuracy and interpretability
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera |
Soft Comput. | 2 |
| 2008 | A Short Study on the Use of Genetic 2-Tuples Tuning for Fuzzy Rule Based Classification Systems in Imbalanced Data-SetsabstractIn this work our aim is to increase the performance of Fuzzy Rule Based Classifications Systems in the framework of imbalanced data-sets by means of the application of a genetic tuning step. We focus on the imbalanced data-set problem since it appears in many real application areas and, for this reason, it has become a relevant topic in the area of machine learning. This problem occurs when the number of examples that represents one of the concepts of interest (usually the most important) is much lower than that of the remaining ones. We want to adapt the 2-tuples based genetic tuning approach to classification problems and to study the positive synergy between this method and the Chi et al.'s fuzzy learning method, which is a basic approach in order to build the initial Knowledge Base. The experimental results show the improvement achieved by the 2-tuples based genetic tuning over the Fuzzy Rule Based Classification System in all types of imbalanced data, obtaining a better behaviour than the basic approach. Alberto Fernández 0001, María José del Jesus, Francisco Herrera |
HIS | 1 |
| 2008 | A study of the behaviour of linguistic fuzzy rule based classification systems in the framework of imbalanced data-sets
Alberto Fernández 0001, Salvador García 0001, María José del Jesus, Francisco Herrera |
Fuzzy Sets Syst. | 1 |
| 2006 | A Proposal of Evolutionary Prototype Selection for Class Imbalance Problems
Salvador García 0001, José Ramón Cano, Alberto Fernández 0001, Francisco Herrera |
IDEAL | 3 |