Salvador García 0001

dblp:27/5514 · DBLP profile ↗
← Back
115ranked-venue papers
16as first author
26since 2021 · last 2026
0000-0003-4494-7565ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 10 first-author · 21 since 2021Databases, data management, data science and information retrieval · 24 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 2Systems, architecture and hardware · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A consensus model considering trust-based interactive weight allocation for the procurement of new energy vehicles
Junpeng Sun, Zaiwu Gong, Yejun Xu, Yanxin Xu, Salvador García 0001
Eng. Appl. Artif. Intell.5
2026 Balancing forecast accuracy and switching costs in online optimization of energy management systems
Evgenii Genov, Julian Ruddick, Christoph Bergmeir, Majid Vafaeipour, Thierry Coosemans, Salvador García 0001, Maarten Messagie
Expert Syst. Appl.6
2026 Evolutionary Computation for the Design and Enrichment of General-Purpose Artificial Intelligence Systems: Survey and Prospects
abstract
In Artificial Intelligence, there is an increasing demand for adaptive models capable of dealing with a diverse spectrum of learning tasks, surpassing the limitations of systems devised to cope with a single task. The recent emergence of General-Purpose Artificial Intelligence Systems (GPAIS) poses model configuration and adaptability challenges at far greater complexity scales than the optimal design of traditional Machine Learning models. Evolutionary Computation (EC) has been a useful tool for both the design and optimization of Machine Learning models, endowing them with the capability to configure and/or adapt themselves to the task under consideration. Therefore, their application to GPAIS is a natural choice. This paper aims to analyze the role of EC in the field of GPAIS, exploring the use of EC for their design or enrichment. We also match GPAIS properties to Machine Learning areas in which EC has had a notable contribution, highlighting recent milestones of EC for GPAIS. Furthermore, we discuss the challenges of harnessing the benefits of EC for GPAIS, presenting different strategies to both design and improve GPAIS with EC, covering tangential areas, identifying research niches, and outlining potential research directions for EC and GPAIS.
Daniel Molina, Javier Poyatos, Javier Del Ser, Salvador García 0001, Hisao Ishibuchi, Isaac Triguero, Bing Xue 0001, Xin Yao 0001, Francisco Herrera
IEEE Trans. Evol. Comput.4
2025 RSPCA: Random Sample Partition and Clustering Approximation for ensemble learning of big data
Mohammad Sultan Mahmud, Diego García-Gil, Salvador García 0001, Joshua Zhexue Huang
Pattern Recognit.4
2025 Determination of the Number of Clusters in High-Dimensional Data With Subspace Clusters
abstract
Current big data clustering methods pose a challenge to efficiently estimate the number of clusters when dealing with a large number of instances and dimensions. However, determining the large number of clusters in a high-dimensional dataset is difficult due to the conflicting nature of its subspaces. In this paper, we propose a new distributed clustering ensemble method calledsubspaceclusterensemble (SSCE) to find the number of clusters in big high-dimensional datasets with informative subspace clusters. We represent the high-dimensional dataset as a set of random subspaces, and then in each subspace view, we investigate the number of clusters and initial cluster centers with the I-niceDP clustering scheme. To efficiently handle big data (i.e., a large number of instances with high dimensions), we use multiple random samples of a big dataset to estimate the number of clusters. With this aim, we first adopt the random sample partitioning (RSP) data model to facilitate random sample generation and then apply the subspace cluster estimation scheme. To do so, each random sample (i.e., RSP data block) is processed independently in parallel, and subspace clustering components are generated. A novel two-stage clustering ensemble, ball fusion (BF), is proposed to estimate the final result from the subspace clustering components and multiple RSP sample outcomes. Subspace and full-space clusters induce key connections within subspaces to generate the final ensemble outcome. Results from experiments on synthetic and real-world datasets demonstrated the effectiveness of the proposed method, which outperformed the state-of-the-art baselines.
Mohammad Sultan Mahmud, Joshua Zhexue Huang, Germán González-Almagro, Salvador García 0001
IEEE Trans. Big Data4
2025 Hybrid Gromov-Wasserstein Embedding for Capsule Learning
abstract
Capsule networks (CapsNets) aim to parse images into a hierarchy of objects, parts, and their relationships using a two-step process involving part-whole transformation and hierarchical component routing. However, this hierarchical relationship modeling is computationally expensive, which has limited the wider use of CapsNet despite its potential advantages. The current state of CapsNet models primarily focuses on comparing their performance with capsule baselines, falling short of achieving the same level of proficiency as deep convolutional neural network (CNN) variants in intricate tasks. To address this limitation, we present an efficient approach for learning capsules that surpasses canonical baseline models and even demonstrates superior performance compared with high-performing convolution models. Our contribution can be outlined in two aspects: first, we introduce a group of subcapsules onto which an input vector is projected. Subsequently, we present the hybrid Gromov-Wasserstein (HGW) framework, which initially quantifies the dissimilarity between the input and the components modeled by the subcapsules, followed by determining their alignment degree through optimal transport (OT). This innovative mechanism capitalizes on new insights into defining alignment between the input and subcapsules, based on the similarity of their respective component distributions. This approach enhances CapsNets' capacity to learn from intricate, high-dimensional data while retaining their interpretability and hierarchical structure. Our proposed model offers two distinct advantages: 1) its lightweight nature facilitates the application of capsules to more intricate vision tasks, including object detection; and 2) it outperforms baseline approaches in these demanding tasks. Our empirical findings illustrate that HGW capsules (HGWCapsules) exhibit enhanced robustness against affine transformations, scale effectively to larger datasets, and surpass CNN and CapsNet models across various vision tasks.
Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Eric Granger, Salvador García 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Fractional Correspondence Framework in Detection Transformer
abstract
The Detection Transformer (DETR), by incorporating the Hungarian algorithm, has significantly simplified the matching process in object detection tasks. This algorithm facilitates optimal one-to-one matching of predicted bounding boxes to ground-truth annotations during training. While effective, this strict matching process does not inherently account for the varying densities and distributions of objects, leading to suboptimal correspondences such as failing to handle multiple detections of the same object or missing small objects. To address this, we propose the Regularized Transport Plan (RTP). RTP introduces a flexible matching strategy that captures the cost of aligning predictions with ground truths to find the most accurate correspondences between these sets. By utilizing the differentiable Sinkhorn algorithm, RTP allows for soft, fractional matching rather than strict one-to-one assignments. This approach enhances the model's capability to manage varying object densities and distributions effectively. Our extensive evaluations on the MS-COCO and VOC benchmarks demonstrate the effectiveness of our approach. RTP-DETR, surpassing the performance of the Deform-DETR and the recently introduced DINO-DETR, achieving absolute gains in mAP of +3.8% and +1.7%, respectively.
Masoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou 0001, Yue Lu 0001, Salvador García 0001
ACM Multimedia5
2024 Enhancing microblog sentiment analysis through multi-level feature interaction fusion with social relationship guidance
Chenquan Gan, Xiaopeng Cao, Qingyi Zhu, Deepak Kumar Jain 0001, Salvador García 0001
Appl. Intell.5
2024 Metric learning for monotonic classification: turning the space up to the limits of monotonicity
abstract
Abstract This paper presents, for the first time, a distance metric learning algorithm for monotonic classification. Monotonic datasets arise in many real-world applications, where there exist order relations in the input and output variables, and the outputs corresponding to ordered pairs of inputs are also expected to be ordered. Monotonic classification can be addressed through several distance-based classifiers that are able to respect the monotonicity constraints of the data. The performance of distance-based classifiers can be improved with the use of distance metric learning algorithms, which are able to find the distances that best represent the similarities among each pair of data samples. However, learning a distance for monotonic data has an additional drawback: the learned distance may negatively impact the monotonic constraints of the data. In our work, we propose a new model for learning distances that does not corrupt these constraints. This methodology will also be useful in identifying and discarding non-monotonic pairs of samples that may be present in the data due to noise. The experimental analysis conducted, supported by a Bayesian statistical testing, demonstrates that the distances obtained by the proposed method can enhance the performance of several distance-based classifiers in monotonic problems.
Juan-Luis Suárez, Germán González-Almagro, Salvador García 0001, Francisco Herrera
Appl. Intell.3
2024 Robust multi-modal pedestrian detection using deep convolutional neural network with ensemble learning model
Deepak Kumar Jain 0001, Salvador García 0001, S. Neelakandan
Expert Syst. Appl.3
2024 Application of spatial uncertainty predictor in CNN-BiLSTM model using coronary artery disease ECG signals
abstract
This study aims to address the need for reliable diagnosis of coronary artery disease (CAD) using artificial intelligence (AI) models. Despite the progress made in mitigating opacity with explainable AI (XAI) and uncertainty quantification (UQ), understanding the real-world predictive reliability of AI methods remains a challenge. In this study, we propose a novel indicator called the Spatial Uncertainty Estimator (SUE) to assess the prediction reliability of classification networks in practical Electrocardiography (ECG) scenarios. SUE quantifies the spatial overlap of critical Grad-CAM (Gradient-weighted Class Activation Mapping) features, offering a confidence score for predictions. To validate SUE, we designed a deep learning network that integrates Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) mechanisms for precise ECG signal classification of CAD. This network achieved high accuracy, sensitivity, and specificity rates of 99.6%, 99.8%, and 98.2%, respectively. During test time, SUE accurately distinguishes between correctly classified and misclassified ECG segments, demonstrating the superiority of the proposed network over existing methods. The study highlights the potential of combining XAI and UQ techniques to enhance ECG analysis. The evaluation of spatial overlap among discriminative features provides quantitative insights into the network's robustness, encompassing both current prediction accuracy and the repeatability of predictions.
Silvia Seoni, Filippo Molinari, U. Rajendra Acharya, Shu Lih Oh, Prabal Datta Barua, Salvador García 0001, Massimo Salvi
Inf. Sci.6
2024 Video multimodal sentiment analysis using cross-modal feature translation and dynamical propagation
Chenquan Gan, Qingyi Zhu, Deepak Kumar Jain 0001, Salvador García 0001
Knowl. Based Syst.6
2023 Speech emotion recognition via multiple fusion under spatial-temporal parallel network
abstract
Speech, as a necessary way to express emotions, plays a vital role in human communication. With the continuous deepening of research on emotion recognition in human–computer interaction, speech emotion recognition (SER) has become an essential task to improve the human–computer interaction experience. When performing emotion feature extraction of speech, the method of cutting the speech spectrum will destroy the continuity of speech. Besides, the method of using the cascaded structure without cutting the speech spectrum cannot simultaneously extract speech spectrum information from both temporal and spatial domains. To this end, we propose a spatial–temporal parallel network for speech emotion recognition without cutting the speech spectrum. To further mix the temporal and spatial features, we design a novel fusion method (called multiple fusion) that combines the concatenate fusion and ensemble strategy. Finally, the experimental results on five datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Chenquan Gan, Qingyi Zhu, Yong Xiang 0001, Deepak Kumar Jain 0001, Salvador García 0001
Neurocomputing6
2023 Handling Imbalanced Classification Problems With Support Vector Machines via Evolutionary Bilevel Optimization
abstract
Support vector machines (SVMs) are popular learning algorithms to deal with binary classification problems. They traditionally assume equal misclassification costs for each class; however, real-world problems may have an uneven class distribution. This article introduces EBCS-SVM: evolutionary bilevel cost-sensitive SVMs. EBCS-SVM handles imbalanced classification problems by simultaneously learning the support vectors and optimizing the SVM hyperparameters, which comprise the kernel parameter and misclassification costs. The resulting optimization problem is a bilevel problem, where the lower level determines the support vectors and the upper level the hyperparameters. This optimization problem is solved using an evolutionary algorithm (EA) at the upper level and sequential minimal optimization (SMO) at the lower level. These two methods work in a nested fashion, that is, the optimal support vectors help guide the search of the hyperparameters, and the lower level is initialized based on previous successful solutions. The proposed method is assessed using 70 datasets of imbalanced classification and compared with several state-of-the-art methods. The experimental results, supported by a Bayesian test, provided evidence of the effectiveness of EBCS-SVM when working with highly imbalanced datasets.
Alejandro Rosales-Pérez, Salvador García 0001, Francisco Herrera
IEEE Trans. Cybern.2
2023 GEN: Generative Equivariant Networks for Diverse Image-to-Image Translation
abstract
Image-to-image (I2I) translation has become a key asset for generative adversarial networks. Convolutional neural networks (CNNs), despite having a significant performance, are not able to capture the spatial relationships among different parts of an object and, thus, do not qualify as the ideal representative model for image translation tasks. As a remedy to this problem, capsule networks have been proposed to represent patterns for a visual object in such a way that preserves hierarchical spatial relationships. The training of capsules is constrained by learning all pairwise relationships between capsules of consecutive layers. This design would be prohibitively expensive both in time and memory. In this article, we present a new framework for capsule networks to provide a full description of the input components at various levels of semantics, which can successfully be applied to the generator-discriminator architectures without incurring computational overhead compared to the CNNs. To successfully apply the proposed capsules in the generative adversarial network, we put forth a novel Gromov-Wasserstein (GW) distance as a differentiable loss function that compares the dissimilarity between two distributions and then guides the learned distribution toward target properties, using optimal transport (OT) discrepancy. The proposed method-which is called generative equivariant network (GEN)-is an alternative architecture for GANs with equivariance capsule layers. The proposed model is evaluated through a comprehensive set of experiments on I2I translation and image generation tasks and compared with several state-of-the-art models. Results indicate that there is a principled connection between generative and capsule models that allows extracting discriminant and invariant information from image data.
Pourya Shamsolmoali, Masoumeh Zareapoor, Swagatam Das, Salvador García 0001, Eric Granger, Jie Yang 0002
IEEE Trans. Cybern.4
2022 Monotonic Constrained Clustering: A First Approach
Germán González-Almagro, Pablo Sánchez Bermejo, Juan-Luis Suárez, José Ramón Cano, Salvador García 0001
IEA/AIE5
2022 A Preliminary Approach for using Metric Learning in Monotonic Classification
Juan-Luis Suárez, Germán González-Almagro, Salvador García 0001, Francisco Herrera
IEA/AIE3
2022 3SHACC: Three stages hybrid agglomerative constrained clustering
Germán González-Almagro, Juan-Luis Suárez, Julián Luengo, José Ramón Cano, Salvador García 0001
Neurocomputing5
2021 Ordinal Regression with Explainable Distance Metric Learning Based on Ordered Sequences: Extended Abstract
abstract
Ordinal regression addresses the problem of predicting non-numerical ordered classes. It walks a fine line between standard regression and classification, and the problem is often addressed from one of these perspectives. This can lead to suboptimal results as the ordinal information in the data may not be properly exploited. In this work we propose a distance metric learning algorithm to handle ordinal regression. Our model aims at optimizing the number of ordered sequences in local neighborhoods of the data, so that the learned distance can then be used by a distance-based predictor and improve its performance in ordinal regression problems. We evaluate our algorithm on several ordinal regression datasets and show that it outperforms the current distance metric learning for ordinal regression proposals, as well as being competitive with respect to the state-of-the-art of ordinal regression. The current paper is an extended abstract for the work [1].
Juan-Luis Suárez, Salvador García 0001, Francisco Herrera
DSAA2
2021 Neurocomputing guest editorial for the special issue: Advances in deep and shallow machine learning approaches for handling data irregularities
Swagatam Das, Salvador García 0001, Isaac Triguero
Neurocomputing2
2021 Fuzzy k-nearest neighbors with monotonicity constraints: Moving towards the robustness of monotonic noise
Sergio González, Salvador García 0001, Sheng-Tun Li, Robert Ivor John, Francisco Herrera
Neurocomputing2
2021 A tutorial on distance metric learning: Mathematical foundations, algorithms, experimental analysis, prospects and challenges
Juan-Luis Suárez, Salvador García 0001, Francisco Herrera
Neurocomputing2
2021 Synthetic Sample Generation for Label Distribution Learning
Julián Luengo, José Ramón Cano, Salvador García 0001
Inf. Sci.4
2021 BELIEF: A distance-based redundancy-proof feature selection method for Big Data
Sergio Ramírez-Gallego, Salvador García 0001, Ning Xiong 0001, Francisco Herrera
Inf. Sci.3
2021 Advances in domain adaptation for computer vision
Pourya Shamsolmoali, Salvador García 0001, Huiyu Zhou 0001, M. Emre Celebi 0001
Image Vis. Comput.2
2021 Ordinal regression with explainable distance metric learning based on ordered sequences
Juan-Luis Suárez, Salvador García 0001, Francisco Herrera
Mach. Learn.2
2020 A Hybrid Surrogate Model for Evolutionary Undersampling in Imbalanced Classification
abstract
Data preprocessing is a key stage in data mining that allows machine learning algorithms to obtain meaningful insights. Many preprocessing problems such as feature selection or instance selection can be modelled as optimisation/search problems. Evolutionary algorithms have traditionally excelled in this task when dealing with data of a moderate size. However, their application to large datasets typically involves very high computational costs. In this work, we propose a hybrid surrogate model for evolutionary undersampling in imbalanced classification problems. These are characterised by having a highly skewed distribution of classes in which evolutionary algorithms aim to balance the training data by selecting only the most relevant data. The proposed technique combines a two-stage clustering-based surrogate method with a windowing approach to quickly approximate fitness values of the chromosomes and accelerate the search. The experiments carried out in 44 standard imbalanced datasets show that the proposed hybrid surrogate model highly reduces the computational cost of the evolutionary algorithm without a considerable loss of performance.
Hoang Lam Le, Dario Landa Silva, Mikel Galar, Salvador García 0001, Isaac Triguero
CEC4
2020 Improving constrained clustering via decomposition-based multiobjective optimization with memetic elitism
abstract
Clustering has always been a topic of interest in knowledge discovery, it is able to provide us with valuable information within the unsupervised machine learning framework. It received renewed attention when it was shown to produce better results in environments where partial information about how to solve the problem is available, thus leading to a new machine learning paradigm: semi-supervised machine learning. This new type of information can be given in the form of constraints, which guide the clustering process towards quality solutions. In particular, this study considers the pairwise instance-level must-link and cannot-link constraints. Given the ill-posed nature of the constrained clustering problem, we approach it from the multiobjective optimization point of view. Our proposal consists in a memetic elitist evolutionary strategy that favors exploitation by applying a local search procedure to the elite of the population and transferring its results only to the external population, which will also be used to generate new individuals. We show the capability of this method to produce quality results for the constrained clustering problem when considering incremental levels of constraint-based information. For the comparison with state-of-the-art methods, we include previous multiobjective approaches, single-objective genetic algorithms and classic constrained clustering methods.
Germán González-Almagro, Alejandro Rosales-Pérez, Julián Luengo, José Ramón Cano, Salvador García 0001
GECCO5
2020 Preprocessing methodology for time series: An industrial world application case study
Juan Antonio Cortés-Ibáñez, Sergio González, José Javier Valle-Alonso, Julián Luengo, Salvador García 0001, Francisco Herrera
Inf. Sci.5
2020 pyDML: A Python Library for Distance Metric Learning
abstract
pyDML is an open-source python library that provides a wide range of distance metric learning algorithms. Distance metric learning can be useful to improve similarity learning algorithms, such as the nearest neighbors classifier, and also has other applications, like dimensionality reduction. The pyDML package currently provides more than 20 algorithms, which can be categorized, according to their purpose, in: dimensionality reduction algorithms, algorithms to improve nearest neighbors or nearest centroids classifiers, information theory based algorithms or kernel based algorithms, among others. In addition, the library also provides some utilities for the visualization of classifier regions, parameter tuning and a stats website with the performance of the implemented algorithms. The package relies on the scipy ecosystem, it is fully compatible with scikit-learn, and is distributed under GPLv3 license. Source code and documentation can be found at https://github.com/jlsuarezdiaz/pyDML.
Juan-Luis Suárez, Salvador García 0001, Francisco Herrera
J. Mach. Learn. Res.2
2020 Fast and Scalable Approaches to Accelerate the Fuzzy k-Nearest Neighbors Classifier for Big Data
abstract
One of the best-known and most effective methods in supervised classification is the k-nearest neighbors algorithm (kNN). Several approaches have been proposed to improve its accuracy, where fuzzy approaches prove to be among the most successful, highlighting the classical fuzzy k-nearest neighbors (FkNN). However, these traditional algorithms fail to tackle the large amounts of data that are available today. There are multiple alternatives to enable kNN classification in big datasets, spotlighting the approximate version of kNN called hybrid spill tree. Nevertheless, the existing proposals of FkNN for big data problems are not fully scalable, because a high computational load is required to obtain the same behavior as the original FkNN algorithm. This article proposes global approximate hybrid spill tree FkNN and local hybrid spill tree FkNN, two approximate approaches that speed up runtime without losing quality in the classification process. The experimentation compares various FkNN approaches for big data with datasets of up to 11 million instances. The results show an improvement in runtime and accuracy over literature algorithms.
Jesús Maillo, Salvador García 0001, Julián Luengo, Francisco Herrera, Isaac Triguero
IEEE Trans. Fuzzy Syst.2
2019 Big Data Preprocessing as the Bridge between Big Data and Smart Data: BigDaPSpark and BigDaPFlink Libraries
abstract
With the advent of Big Data, terabytes of data are generated and stored every second. This raw data is far from \nbeing perfect, it contains many imperfections (noise, missing values, etc.) and is not suitable for analysis, \nas it will led to wrong conclusions. Data preprocessing is the set of techniques devoted to polish, clean, \nfix, and improve that raw data. With this preprocessed data, we would be able to find more patterns in it, \nand to better explain the underlaying distribution of the data. This is what is called Smart Data, raw data \nthat has been preprocessed and is ready for being analyzed, data that contains valuable information that will \nled to knowledge. In this work, we present two Big Data libraries for achieving Smart Data from Big Data, \nBigDaPSpark and BigDaPFlink. They are built on top of two Big Data frameworks, Apache Spark and Apache \nFlink. Both libraries contain a series of algorithms for Big Data preprocessing, ranging from noise cleaning, \nto discretization, or data reduction, among many others. Additionally, we ilustrate the usage of the libraries \nwith two cases of use.
Diego García-Gil, Alejandro Alcalde-Barros, Julián Luengo, Salvador García 0001, Francisco Herrera
IoTBDS4
2019 A First Approach on Big Data Missing Values Imputation
abstract
Albeit most techniques and algorithms assume that the data is accurate, measurements in our analogic world are far from being perfect. Since our capabilities of storing and processing data are growing everyday, these imperfections will accumulate, generating poorer decisions and hindering any knowledge extraction process carried out over the raw data. One of the most disturbing imperfections is the presence of missing values. Many inductive algorithms assume that the data is complete, thus if they face missing data they will not work properly or the quality of the knowledge extracted will be poorer. At this point there is no sophisticated missing values treatment implemented in any major Big Data framework. In this contribution, we present two novel imputation methods based on clustering that achieve better results than simply removing the faulty examples or filling-in the missing values with the mean that can be easily ported to Spark’s MLlib.
Besay Montesdeoca, Julián Luengo, Jesús Maillo, Diego García-Gil, Salvador García 0001, Francisco Herrera
IoTBDS5
2019 From Big to Smart Data: Iterative ensemble filter for noise filtering in Big Data classification
abstract
The quality of the data is directly related to the quality of the models drawn from that data. For that reason, many research is devoted to improve the quality of the data and to amend errors that it may contain. One of the most common problems is the presence of noise in classification tasks, where noise refers to the incorrect labeling of training instances. This problem is very disruptive, as it changes the decision boundaries of the problem. Big Data problems pose a new challenge in terms of quality data due to the massive and unsupervised accumulation of data. This Big Data scenario also brings new problems to classic data preprocessing algorithms, as they are not prepared for working with such amounts of data, and these algorithms are key to move from Big to Smart Data. In this paper, an iterative ensemble filter for removing noisy instances in Big Data scenarios is proposed. Experiments carried out in six Big Data datasets have shown that our noise filter outperforms the current state-of-the-art noise filter in Big Data domains. It has also proved to be an effective solution for transforming raw Big Data into Smart Data.
Diego García-Gil, Francisco Luque Sánchez, Julián Luengo, Salvador García 0001, Francisco Herrera
Int. J. Intell. Syst.4
2019 Monotonic classification: An overview on algorithms, performance measures and data sets
José Ramón Cano, Pedro Antonio Gutiérrez, Bartosz Krawczyk, Michal Wozniak 0001, Salvador García 0001
Neurocomputing5
2019 Label noise filtering techniques to improve monotonic classification
José Ramón Cano, Julián Luengo, Salvador García 0001
Neurocomputing3
2019 Smartdata: Data preprocessing to achieve smart data in R
Ignacio Cordón, Julián Luengo, Salvador García 0001, Francisco Herrera, Francisco Charte
Neurocomputing3
2019 Enabling Smart Data: Noise filtering in Big Data classification
Diego García-Gil, Julián Luengo, Salvador García 0001, Francisco Herrera
Inf. Sci.3
2019 Chain based sampling for monotonic imbalanced classification
Sergio González, Salvador García 0001, Sheng-Tun Li, Francisco Herrera
Inf. Sci.2
2019 Instance reduction for one-class classification
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
Knowl. Inf. Syst.3
2018 A preliminary study on Hybrid Spill-Tree Fuzzy k-Nearest Neighbors for big data classification
abstract
The Fuzzy k Nearest Neighbor (Fuzzy kNN) classifier is well known for its effectiveness in supervised learning problems. kNN classifies by comparing new incoming examples with a similarity function using the samples of the training set. The fuzzy version of the kNN accounts for the underlying uncertainty in the class labels, and it is composed of two different stages. The first one is responsible for calculating the fuzzy membership degree for each sample of the problem in order to obtain smoother boundaries between classes. The second stage classifies similarly to the standard kNN algorithm but uses the previously calculated class membership degree. To deal with very large datasets, distributed versions of the Fuzzy kNN algorithm have been proposed. However, existing approaches remain not fully scalable as they aim to replicate the exact behavior of the Fuzzy kNN. In this work, we present an approximate and distributed Fuzzy kNN approach based on Hybrid Spill-Tree implemented under Apache Spark. The aim of this model is to alleviate the scalability problems and to deal with big datasets maintaining high accuracy. In our experiments, we compare in precision and runtime with the Fuzzy kNN for big data problems existing in the literature, running with datasets of up to 11 million instances. The results show an improvement in the runtime and accuracy with respect to the previous exact model.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE3
2018 Cooperative multi-objective evolutionary support vector machines for multiclass problems
abstract
In recent years, evolutionary algorithms have been found to be effective and efficient techniques to train support vector machines (SVMs) for binary classification problems while multiclass problems have been neglected. This paper proposes CMOE-SVM: Cooperative Multi-Objective Evolutionary SVMs for multiclass problems. CMOE-SVM enables SVMs to handle multiclass problems via co-evolutionary optimization, by breaking down the original M-class problem into M simpler ones, which are optimized simultaneously in a cooperative manner. Furthermore, CMOE-SVM can explicitly maximize the margin and reduce the training error (the two components of the SVM optimization), by means of multi-objective optimization. Through a comprehensive experimental evaluation using a suite of benchmark datasets, we validate the performance of CMOE-SVM. The experimental results, supported by statistical tests, give evidence of the effectiveness of the proposed approach for solving multiclass classification problems.
Alejandro Rosales-Pérez, Andrés Eduardo Gutiérrez-Rodríguez, Salvador García 0001, Hugo Terashima-Marín, Carlos A. Coello Coello, Francisco Herrera
GECCO3
2018 Online entropy-based discretization for data streaming classification
Sergio Ramírez-Gallego, Salvador García 0001, Francisco Herrera
Future Gener. Comput. Syst.2
2018 On the use of convolutional neural networks for robust classification of multiple fingerprint captures
abstract
Fingerprint classification is one of the most common approaches to accelerate the identification in large databases of fingerprints. Fingerprints are grouped into disjoint classes, so that an input fingerprint is compared only with those belonging to the predicted class, reducing the penetration rate of the search. The classification procedure usually starts by the extraction of features from the fingerprint image, frequently based on visual characteristics. In this work, we propose an approach to fingerprint classification using convolutional neural networks, which avoid the necessity of an explicit feature extraction process by incorporating the image processing within the training of the classifier. Furthermore, such an approach is able to predict a class even for low-quality fingerprints that are rejected by commonly used algorithms, such as FingerCode. The study gives special importance to the robustness of the classification for different impressions of the same fingerprint, aiming to minimize the penetration in the database. In our experiments, convolutional neural networks yielded better accuracy and penetration rate than state-of-the-art classifiers based on explicit feature extraction. The tested networks also improved on the runtime, as a result of the joint optimization of both feature extraction and classification.
Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera
Int. J. Intell. Syst.3
2018 DRCW-ASEG: One-versus-One distance-based relative competence weighting with adaptive synthetic example generation for multi-class imbalanced datasets
Zhongliang Zhang 0001, Sergio González, Salvador García 0001, Francisco Herrera
Neurocomputing4
2018 Dynamic ensemble selection for multi-class imbalanced datasets
Salvador García 0001, Zhongliang Zhang 0001, Abdulrahman H. Altalhi, Saleh Alshomrani, Francisco Herrera
Inf. Sci.1
2018 SMOTE for Learning from Imbalanced Data: Progress and Challenges, Marking the 15-year Anniversary
abstract
The Synthetic Minority Oversampling Technique (SMOTE) preprocessing algorithm is considered "de facto" standard in the framework of learning from imbalanced data. This is due to its simplicity in the design of the procedure, as well as its robustness when applied to different type of problems. Since its publication in 2002, SMOTE has proven successful in a variety of applications from several different domains. SMOTE has also inspired several approaches to counter the issue of class imbalance, and has also significantly contributed to new supervised learning paradigms, including multilabel classification, incremental learning, semi-supervised learning, multi-instance learning, among others. It is standard benchmark for learning from imbalanced data. It is also featured in a number of different software packages - from open source to commercial. In this paper, marking the fifteen year anniversary of SMOTE, we reflect on the SMOTE journey, discuss the current state of affairs with SMOTE, its applications, and also identify the next set of challenges to extend SMOTE for Big Data problems.
Alberto Fernández 0001, Salvador García 0001, Francisco Herrera, Nitesh V. Chawla
J. Artif. Intell. Res.2
2018 Imbalance: Oversampling algorithms for imbalanced classification in R
Ignacio Cordón, Salvador García 0001, Alberto Fernández 0001, Francisco Herrera
Knowl. Based Syst.2
2018 Principal Components Analysis Random Discretization Ensemble for Big Data
Diego García-Gil, Sergio Ramírez-Gallego, Salvador García 0001, Francisco Herrera
Knowl. Based Syst.3
2017 Exact fuzzy k-nearest neighbor classification for big datasets
abstract
The k-Nearest Neighbors (kNN) classifier is one of the most effective methods in supervised learning problems. It classifies unseen cases comparing their similarity with the training data. Nevertheless, it gives to each labeled sample the same importance to classify. There are several approaches to enhance its precision, with the Fuzzy k-Nearest Neighbors (Fuzzy-kNN) classifier being among the most successful ones. Fuzzy-kNN computes a fuzzy degree of membership of each instance to the classes of the problem. As a result, it generates smoother borders between classes. Apart from the existing kNN approach to handle big datasets, there is not a fuzzy variant to manage that volume of data. Nevertheless, calculating this class membership adds an extra computational cost becoming even less scalable to tackle large datasets because of memory needs and high runtime. In this work, we present an exact and distributed approach to run the Fuzzy-kNN classifier on big datasets based on Spark, which provides the same precision than the original algorithm. It presents two separately stages. The first stage transforms the training set adding the class membership degrees. The second stage classifies with the kNN algorithm the test set using the class membership computed previously. In our experiments, we study the scaling-up capabilities of the proposed approach with datasets up to 11 million instances, showing promising results.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE3
2017 Training set selection for monotonic ordinal classification
José Ramón Cano, Salvador García 0001
Data Knowl. Eng.2
2017 Prototype selection to improve monotonic nearest neighbor
José Ramón Cano, Naif R. Aljohani, Rabeeh Ayaz Abbasi, Jalal S. Alowibdi, Salvador García 0001
Eng. Appl. Artif. Intell.5
2017 A survey on data preprocessing for data stream mining: Current status and future directions
Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
Neurocomputing3
2017 CommuniMents: A Framework for Detecting Community Based Sentiments for Events
abstract
Social media has revolutionized human communication and styles of interaction. Due to its effectiveness and ease, people have started using it increasingly to share and exchange information, carry out discussions on various events, and express their opinions. Various communities may have diverse sentiments about events and it is an interesting research problem to understand the sentiments of a particular community for a specific event. In this article, the authors propose a framework CommuniMents which enables us to identify the members of a community and measure the sentiments of the community for a particular event. CommuniMents uses automated snowball sampling to identify the members of a community, then fetches their published contents (specifically tweets), pre-processes the contents and measures the sentiments of the community. The authors perform qualitative and quantitative evaluation for a variety of real world events to validate the effectiveness of the proposed framework.
Muhammad Aslam Jarwar, Rabeeh Ayaz Abbasi, Mubashar Mushtaq, Onaiza Maqbool, Naif R. Aljohani, Ali Daud, Jalal S. Alowibdi, José Ramón Cano, Salvador García 0001, Ilyoung Chong
Int. J. Semantic Web Inf. Syst.9
2017 Minutiae-based fingerprint matching decomposition: Methodology for big data frameworks
Daniel Peralta, Salvador García 0001, José Manuel Benítez 0001, Francisco Herrera
Inf. Sci.2
2017 Distributed incremental fingerprint identification with reduced database penetration rate using a hierarchical classification based on feature fusion and selection
Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera
Knowl. Based Syst.3
2017 Exploring the effectiveness of dynamic ensemble selection in the one-versus-one scheme
Zhongliang Zhang 0001, Salvador García 0001, Jiafu Tang, Francisco Herrera
Knowl. Based Syst.3
2017 MoNGEL: monotonic nested generalized exemplar learning
Javier García 0004, Habib Fardoun, Daniyal M. Alghazzawi, José Ramón Cano, Salvador García 0001
Pattern Anal. Appl.5
2017 Class Switching according to Nearest Enemy Distance for learning from highly imbalanced data-sets
Sergio González, Salvador García 0001, Marcelino Lázaro, Aníbal R. Figueiras-Vidal, Francisco Herrera
Pattern Recognit.2
2017 An Evolutionary Multiobjective Model and Instance Selection for Support Vector Machines With Pareto-Based Ensembles
abstract
Support vector machines (SVMs) are among the most powerful learning algorithms for classification tasks. However, these algorithms require a high computational cost during the training phase, which can limit their application on large-scale datasets. Moreover, it is known that their effectiveness highly depends on the hyper-parameters used to train the model. With the intention of dealing with these, this paper introduces an evolutionary multiobjective model and instance selection (IS) approach for SVMs with Pareto-based ensemble, whose goals are, precisely, to optimize the size of the training set and the classification performance attained by the selection of the instances, which can be done using either a wrapper or a filter approach. Due to the nature of multiobjective evolutionary algorithms, several Pareto optimal solutions can be found. We study several ways of using such information to perform a classification task. To accomplish this, our proposal performs a processing over the Pareto solutions in order to combine them into a single ensemble. This is done in five different ways, which are based on: 1) a global Pareto ensemble; 2) error reduction; 3) a complementary error reduction; 4) maximized margin distance; and 5) boosting. Through a comprehensive experimental study we evaluate the suitability of the proposed approach and the Pareto processing, and we show its advantages over a single-objective formulation, traditional IS techniques, and learning algorithms.
Alejandro Rosales-Pérez, Salvador García 0001, Jesus A. Gonzalez, Carlos A. Coello Coello, Francisco Herrera
IEEE Trans. Evol. Comput.2
2017 Evolutionary Fuzzy Rule-Based Methods for Monotonic Classification
abstract
In data science applications, it is very often to require predictive models satisfying monotonicity with respect to the explanatory variables involved in the dataset. In ordinal classification or regression, this occurs when the output variable or class label do not decrease when input variables increase, or vice versa. This problem is commonly known as monotonic classification, and most existing classification techniques are not able to manage this kind of constraints or they require first to monotonize the data. In the literature, the monotonicity has been considered in linguistic fuzzy models, fuzzy-inference methods, and fuzzy rule-based control systems. However, to the best of our knowledge, there is no fuzzy rule-based system designed to produce monotonic fuzzy rule-based models for classification problems. In this paper, we propose to incorporate some mechanisms based on monotonicity indexes for addressing such problems in two popular and competitive evolutionary fuzzy systems algorithms for classification and regression tasks: FARC-HD and FSmogfse+Tune. In addition, the proposals are able to handle any kind of classification dataset without the necessity of preprocessing. The quality of our approaches is analyzed using statistical analysis and comparing with well-known monotonic classifiers.
Jesús Alcalá-Fdez, Rafael Alcalá, Sergio González, Yusuke Nojima, Salvador García 0001
IEEE Trans. Fuzzy Syst.5
2017 Nearest Neighbor Classification for High-Speed Big Data Streams Using Spark
abstract
Mining massive and high-speed data streams among the main contemporary challenges in machine learning. This calls for methods displaying a high computational efficacy, with ability to continuously update their structure and handle ever-arriving big number of instances. In this paper, we present a new incremental and distributed classifier based on the popular nearest neighbor algorithm, adapted to such a demanding scenario. This method, implemented in Apache Spark, includes a distributed metric-space ordering to perform faster searches. Additionally, we propose an efficient incremental instance selection method for massive data streams that continuously update and remove outdated examples from the case-base. This alleviates the high computational requirements of the original classifier, thus making it suitable for the considered problem. Experimental study conducted on a set of real-life massive data streams proves the usefulness of the proposed solution and shows that we are able to provide the first efficient nearest neighbor solution for high-speed big and streaming data.
Sergio Ramírez-Gallego, Bartosz Krawczyk, Salvador García 0001, Michal Wozniak 0001, José Manuel Benítez 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Syst.3
2016 Evolutionary fuzzy k-nearest neighbors algorithm using interval-valued fuzzy sets
Joaquín Derrac, Francisco Chiclana, Salvador García 0001, Francisco Herrera
Inf. Sci.3
2016 Tutorial on practical tips of the most influential data preprocessing algorithms in data mining
Salvador García 0001, Julián Luengo, Francisco Herrera
Knowl. Based Syst.1
2016 Empowering one-vs-one decomposition with ensemble learning for multi-class imbalanced data
Zhongliang Zhang 0001, Bartosz Krawczyk, Salvador García 0001, Alejandro Rosales-Pérez, Francisco Herrera
Knowl. Based Syst.3
2016 Landmark-based music recognition system optimisation using genetic algorithms
Salvador Gutiérrez, Salvador García 0001
Multim. Tools Appl.2
2016 Multivariate Discretization Based on Evolutionary Cut Points Selection for Classification
abstract
Discretization is one of the most relevant techniques for data preprocessing. The main goal of discretization is to transform numerical attributes into discrete ones to help the experts to understand the data more easily, and it also provides the possibility to use some learning algorithms which require discrete data as input, such as Bayesian or rule learning. We focus our attention on handling multivariate classification problems, where high interactions among multiple attributes exist. In this paper, we propose the use of evolutionary algorithms to select a subset of cut points that defines the best possible discretization scheme of a data set using a wrapper fitness function. We also incorporate a reduction mechanism to successfully manage the multivariate approach on large data sets. Our method has been compared with the best state-of-the-art discretizers on 45 real datasets. The experiments show that our proposed algorithm overcomes the rest of the methods producing competitive discretization schemes in terms of accuracy, for C4.5, Naive Bayes, PART, and PrUning and BuiLding Integrated in Classification classifiers; and obtained far simpler solutions.
Sergio Ramírez-Gallego, Salvador García 0001, José Manuel Benítez 0001, Francisco Herrera
IEEE Trans. Cybern.2
2015 Managing Monotonicity in Classification by a Pruned Random Forest
Sergio González, Francisco Herrera, Salvador García 0001
IDEAL3
2015 MRPR: A MapReduce solution for prototype reduction in big data classification
Isaac Triguero, Daniel Peralta, Jaume Bacardit, Salvador García 0001, Francisco Herrera
Neurocomputing4
2015 A survey on fingerprint minutiae-based local matching for verification and identification: Taxonomy and experimental evaluation
Daniel Peralta, Mikel Galar, Isaac Triguero, Daniel Paternain, Salvador García 0001, Edurne Barrenechea Tartas, José Manuel Benítez 0001, Humberto Bustince, Francisco Herrera
Inf. Sci.5
2015 Self-labeled techniques for semi-supervised learning: taxonomy, software and empirical study
Isaac Triguero, Salvador García 0001, Francisco Herrera
Knowl. Inf. Syst.2
2015 A survey of fingerprint classification Part I: Taxonomies on feature extraction methods and learning models
Mikel Galar, Joaquín Derrac, Daniel Peralta, Isaac Triguero, Daniel Paternain, Carlos Lopez-Molina, Salvador García 0001, José Manuel Benítez 0001, Miguel Pagola, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera
Knowl. Based Syst.7
2015 A survey of fingerprint classification Part II: Experimental analysis and ensemble proposal
Mikel Galar, Joaquín Derrac, Daniel Peralta, Isaac Triguero, Daniel Paternain, Carlos Lopez-Molina, Salvador García 0001, José Manuel Benítez 0001, Miguel Pagola, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera
Knowl. Based Syst.7
2015 SEG-SSC: A Framework Based on Synthetic Examples Generation for Self-Labeled Semi-Supervised Classification
abstract
Self-labeled techniques are semi-supervised classification methods that address the shortage of labeled examples via a self-learning process based on supervised models. They progressively classify unlabeled data and use them to modify the hypothesis learned from labeled samples. Most relevant proposals are currently inspired by boosting schemes to iteratively enlarge the labeled set. Despite their effectiveness, these methods are constrained by the number of labeled examples and their distribution, which in many cases is sparse and scattered. The aim of this paper is to design a framework, named synthetic examples generation for self-labeled semi-supervised classification, to improve the classification performance of any given self-labeled method by using synthetic labeled data. These are generated via an oversampling technique and a positioning adjustment model that use both labeled and unlabeled examples as reference. Next, these examples are incorporated in the main stages of the self-labeling process. The principal aspects of the proposed framework are: 1) introducing diversity to the multiple classifiers used by using more (new) labeled data; 2) fulfilling labeled data distribution with the aid of unlabeled data; and 3) being applicable to any kind of self-labeled method. In our empirical studies, we have applied this scheme to four recent self-labeled methods, testing their capabilities with a large number of data sets. We show that this framework significantly improves the classification capabilities of self-labeled techniques.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Cybern.2
2014 A first attempt on evolutionary prototype reduction for nearest neighbor one-class classification
abstract
Evolutionary prototype reduction techniques are data preprocessing methods originally developed to enhance the nearest neighbor rule. They reduce the training data by selecting or generating representative examples of a given problem. These algorithms have been designed and widely analyzed in standard classification providing very competitive results. However, its application scope can be extended to many other specific domains, such as one-class classification, in which its way of working is very interesting in order to reduce computational complexity and sensitivity to noisy data. In this contribution, we perform a first study on the usefulness of evolutionary prototype reduction methods for one-class classification. To do so, we will focus on two recent evolutionary approaches that follow very different strategies: selection and generation of examples from the training data. Both alternatives provide a resulting preprocessed data set that will be used later by a nearest neighbor one-class classifier as its training data. The results achieved support that these data reduction techniques are suitable tools to improve the performance of the nearest neighbor one-class classification.
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation3
2014 A combined MapReduce-windowing two-level parallel scheme for evolutionary prototype generation
abstract
Evolutionary prototype generation techniques have demonstrated their usefulness to improve the capabilities of the nearest neighbor classifier. They act as data reduction algorithms by generating representative points of a given problem. Their main purposes are to speed up the classification process and to reduce the storage requirements and sensitivity to noise of the nearest neighbor rule. Nowadays, with the increment of available data, the use of this kind of reduction techniques becomes more important. However, their applicability can be limited to problems with no more than tens of thousands of instances. In order to address this limitation, in this work we develop a two-level parallelization scheme for evolutionary prototype generation methods. Firstly, it distributes the functioning of these algorithms in several tasks based on a MapReduce framework. Then, for each one of these tasks (mappers), we accelerate the prototype generation process by using a windowing approach. This model enables evolutionary prototype generation algorithms to be applied over large-scale classification problems without accuracy loss. Our preliminary experiments using a dataset of 1 million instances show that this proposal is an appropriate tool to improve the performance of the nearest neighbor classifier with big data.
Isaac Triguero, Daniel Peralta, Jaume Bacardit, Salvador García 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation4
2014 Addressing imbalanced classification with instance generation techniques: IPADE-ID
Victoria López, Isaac Triguero, Cristóbal J. Carmona, Salvador García 0001, Francisco Herrera
Neurocomputing4
2014 On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification
Isaac Triguero, José A. Sáez, Julián Luengo, Salvador García 0001, Francisco Herrera
Neurocomputing4
2014 Fuzzy nearest neighbor algorithms: Taxonomy, experimental analysis and prospects
Joaquín Derrac, Salvador García 0001, Francisco Herrera
Inf. Sci.2
2014 Analyzing convergence performance of evolutionary algorithms: A statistical approach
Joaquín Derrac, Salvador García 0001, Sheldon Hui, Ponnuthurai N. Suganthan, Francisco Herrera
Inf. Sci.2
2013 An insight into classification with imbalanced data: Empirical results and current trends on using data intrinsic characteristics
Victoria López, Alberto Fernández 0001, Salvador García 0001, Vasile Palade, Francisco Herrera
Inf. Sci.3
2013 On the use of evolutionary feature selection for improving fuzzy rough set based prototype selection
Joaquín Derrac, Nele Verbiest, Salvador García 0001, Chris Cornelis, Francisco Herrera
Soft Comput.3
2013 A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning
abstract
Discretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones.
Salvador García 0001, Julián Luengo, José A. Sáez, Victoria López, Francisco Herrera
IEEE Trans. Knowl. Data Eng.1
2012 A Preliminary Study on Selecting the Optimal Cut Points in Discretization by Evolutionary Algorithms
Salvador García 0001, Victoria López, Julián Luengo, Cristóbal J. Carmona, Francisco Herrera
ICPRAM (1)1
2012 Web usage mining to improve the design of an e-commerce website: OrOliveSur.com
Cristóbal J. Carmona, Sergio Ramírez-Gallego, F. J. Torres, Enrique Bernal, María José del Jesus, Salvador García 0001
Expert Syst. Appl.6
2012 Integrating a differential evolution feature weighting scheme into prototype generation
Isaac Triguero, Joaquín Derrac, Salvador García 0001, Francisco Herrera
Neurocomputing3
2012 Enhancing evolutionary instance selection algorithms by means of fuzzy rough set based feature selection
Joaquín Derrac, Chris Cornelis, Salvador García 0001, Francisco Herrera
Inf. Sci.3
2012 On the choice of the best imputation methods for missing values considering three groups of classification methods
Julián Luengo, Salvador García 0001, Francisco Herrera
Knowl. Inf. Syst.2
2012 Evolutionary-based selection of generalized instances for imbalanced classification
Salvador García 0001, Joaquín Derrac, Isaac Triguero, Cristóbal J. Carmona, Francisco Herrera
Knowl. Based Syst.1
2012 Prototype Selection for Nearest Neighbor Classification: Taxonomy and Empirical Study
abstract
The nearest neighbor classifier is one of the most used and well-known techniques for performing recognition tasks. It has also demonstrated itself to be one of the most useful algorithms in data mining in spite of its simplicity. However, the nearest neighbor classifier suffers from several drawbacks such as high storage requirements, low efficiency in classification response, and low noise tolerance. These weaknesses have been the subject of study for many researchers and many solutions have been proposed. Among them, one of the most promising solutions consists of reducing the data used for establishing a classification rule (training data) by means of selecting relevant prototypes. Many prototype selection methods exist in the literature and the research in this area is still advancing. Different properties could be observed in the definition of them, but no formal categorization has been established yet. This paper provides a survey of the prototype selection methods proposed in the literature from a theoretical and empirical point of view. Considering a theoretical point of view, we propose a taxonomy based on the main characteristics presented in prototype selection and we analyze their advantages and drawbacks. Empirically, we conduct an experimental study involving different sizes of data sets for measuring their performance in terms of accuracy, reduction capabilities, and runtime. The results obtained by all the methods studied have been verified by nonparametric statistical tests. Several remarks, guidelines, and recommendations are made for the use of prototype selection for nearest neighbor classification.
Salvador García 0001, Joaquín Derrac, José Ramón Cano, Francisco Herrera
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Integrating Instance Selection, Instance Weighting, and Feature Weighting for Nearest Neighbor Classifiers by Coevolutionary Algorithms
abstract
Cooperative coevolution is a successful trend of evolutionary computation which allows us to define partitions of the domain of a given problem, or to integrate several related techniques into one, by the use of evolutionary algorithms. It is possible to apply it to the development of advanced classification methods, which integrate several machine learning techniques into a single proposal. A novel approach integrating instance selection, instance weighting, and feature weighting into the framework of a coevolutionary model is presented in this paper. We compare it with a wide range of evolutionary and nonevolutionary related methods, in order to show the benefits of the employment of coevolution to apply the techniques considered simultaneously. The results obtained, contrasted through nonparametric statistical tests, show that our proposal outperforms other methods in the comparison, thus becoming a suitable tool in the task of enhancing the nearest neighbor classifier.
Joaquín Derrac, Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Part B3
2012 A Taxonomy and Experimental Study on Prototype Generation for Nearest Neighbor Classification
abstract
The nearest neighbor (NN) rule is one of the most successfully used techniques to resolve classification and pattern recognition tasks. Despite its high classification accuracy, this rule suffers from several shortcomings in time response, noise sensitivity, and high storage requirements. These weaknesses have been tackled by many different approaches, including a good and well-known solution that we can find in the literature, which consists of the reduction of the data used for the classification rule (training data). Prototype reduction techniques can be divided into two different approaches, which are known as prototype selection and prototype generation (PG) or abstraction. The former process consists of choosing a subset of the original training data, whereas PG builds new artificial prototypes to increase the accuracy of the NN classification. In this paper, we provide a survey of PG methods specifically designed for the NN rule. From a theoretical point of view, we propose a taxonomy based on the main characteristics presented in them. Furthermore, from an empirical point of view, we conduct a wide experimental study that involves small and large datasets to measure their performance in terms of accuracy and reduction capabilities. The results are contrasted through nonparametrical statistical tests. Several remarks are made to understand which PG models are appropriate for application to different datasets.
Isaac Triguero, Joaquín Derrac, Salvador García 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Part C3
2011 Differential evolution for optimizing the positioning of prototypes in nearest neighbor classification
Isaac Triguero, Salvador García 0001, Francisco Herrera
Pattern Recognit.2
2011 Addressing data complexity for imbalanced data sets: analysis of SMOTE-based oversampling and evolutionary undersampling
Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera
Soft Comput.3
2010 A preliminary study on the use of differential evolution for adjusting the position of examples in nearest neighbor classification
abstract
Nearest neighbor is one of the most successfully used techniques for performing classification and pattern recognition tasks. Its simplicity and effectiveness justify the use of this technique in certain domains but it however presents several drawbacks referring to time response, noise sensitivity and storage requirements. Several solutions have been proposed in order to alleviate these problems, such as improving the technique for speeding up or carrying out a data reduction process. Prototype generation is a suitable process for data reduction that allows to fit a data set for nearest neighbor classification. Position adjustment of prototypes is a successful technique within the prototype generation methodology. Evolutionary algorithms are adaptive methods based on natural evolution that may be used for search and optimization. Position adjustment of prototypes can be viewed as a search problem, thus it could be solved using evolutionary algorithms. In this paper, we perform a preliminary study on the use of differential evolution algorithms to the prototype generation problem. Differential evolution models are compared with other algorithms for adjusting the position of prototypes and the results are contrasted through non-parametrical statistical tests. The results show that some differential evolution models consistently outperform previously proposed methods.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation2
2010 A Preliminary Study on the Selection of Generalized Instances for Imbalanced Classification
Salvador García 0001, Joaquín Derrac, Isaac Triguero, Cristóbal J. Carmona, Francisco Herrera
IEA/AIE (1)1
2010 Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera
Inf. Sci.1
2010 A study on the use of imputation methods for experimentation with Radial Basis Function Network classifiers handling missing attribute values: The good synergy between RBFNs and EventCovering method
Julián Luengo, Salvador García 0001, Francisco Herrera
Neural Networks2
2010 IFS-CoCo: Instance and feature selection based on cooperative coevolution with nearest neighbor rule
Joaquín Derrac, Salvador García 0001, Francisco Herrera
Pattern Recognit.2
2010 Genetics-Based Machine Learning for Rule Induction: State of the Art, Taxonomy, and Comparative Study
abstract
The classification problem can be addressed by numerous techniques and algorithms which belong to different paradigms of machine learning. In this paper, we are interested in evolutionary algorithms, the so-called genetics-based machine learning algorithms. In particular, we will focus on evolutionary approaches that evolve a set of rules, i.e., evolutionary rule-based systems, applied to classification tasks, in order to provide a state of the art in this field. This paper has a double aim: to present a taxonomy of the genetics-based machine learning approaches for rule induction, and to develop an empirical analysis both for standard classification and for classification with imbalanced data sets. We also include a comparative study of the genetics-based machine learning (GBML) methods with some classical non-evolutionary algorithms, in order to observe the suitability and high potential of the search performed by evolutionary algorithms and the behavior of the GBML algorithms in contrast to the classical approaches, in terms of classification accuracy.
Alberto Fernández 0001, Salvador García 0001, Julián Luengo, Ester Bernadó-Mansilla, Francisco Herrera
IEEE Trans. Evol. Comput.2
2010 IPADE: Iterative Prototype Adjustment for Nearest Neighbor Classification
abstract
Nearest prototype methods are a successful trend of many pattern classification tasks. However, they present several shortcomings such as time response, noise sensitivity, and storage requirements. Data reduction techniques are suitable to alleviate these drawbacks. Prototype generation is an appropriate process for data reduction, which allows the fitting of a dataset for nearest neighbor (NN) classification. This brief presents a methodology to learn iteratively the positioning of prototypes using real parameter optimization procedures. Concretely, we propose an iterative prototype adjustment technique based on differential evolution. The results obtained are contrasted with nonparametric statistical tests and show that our proposal consistently outperforms previously proposed methods, thus becoming a suitable tool in the task of enhancing the performance of the NN classifier.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Neural Networks2
2009 A First Approach to Nearest Hyperrectangle Selection by Evolutionary Algorithms
abstract
The nested generalized exemplar theory accomplishes learning by storing objects in Euclidean n-space, as hyperrectangles. Classification of new data is performed by computing their distance to the nearest “generalized exemplar” or hyperrectangle. This learning method permits to combine the distance-based classification with the axis-parallel rectangle representation employed in most of the rule-learning systems. This contribution proposes the use of evolutionary algorithms to select the most influential hyperrectangles to obtain accurate and simple models in classification tasks. The proposal is compared with the most representative nearest hyperrectangle learning approaches and the results obtained show that the evolutionary proposal outperforms them in accuracy and requires storing a lower number of hyperrectangles.
Salvador García 0001, Joaquín Derrac, Julián Luengo, Francisco Herrera
ISDA1
2009 Addressing Data-Complexity for Imbalanced Data-Sets: A Preliminary Study on the Use of Preprocessing for C4.5
abstract
In this work we analyse the behaviour of the C4.5 classification method with respect to a bunch of imbalanced data-sets. We consider the use of two metrics of data complexity known as “maximum Fishers discriminant ratio” and “nonlinearity of 1NN classifier”, to analyse the effect of preprocessing (oversampling in this case) in order to deal with the imbalance problem. In order to do that, we analyse C4.5 over a wide range of imbalanced data-sets built from real data, and try to extract behaviour patterns from the results. We obtain rules that describe both good or bad behaviours of C4.5 in the case of using the original data-sets (absence of preprocessing) and when applying preprocessing. These rules allow us to determine the effect of the use of preprocessing and to predict the response of C4.5 to preprocessing from the data-set’s complexity metrics prior to its application, and then establish when the preprocessing would be useful to.
Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera
ISDA3
2009 Evolutionary Undersampling for Classification with Imbalanced Datasets: Proposals and Taxonomy
abstract
Learning with imbalanced data is one of the recent challenges in machine learning. Various solutions have been proposed in order to find a treatment for this problem, such as modifying methods or the application of a preprocessing stage. Within the preprocessing focused on balancing data, two tendencies exist: reduce the set of examples (undersampling) or replicate minority class examples (oversampling). Undersampling with imbalanced datasets could be considered as a prototype selection procedure with the purpose of balancing datasets to achieve a high classification rate, avoiding the bias toward majority class examples. Evolutionary algorithms have been used for classical prototype selection showing good results, where the fitness function is associated to the classification and reduction rates. In this paper, we propose a set of methods called evolutionary undersampling that take into consideration the nature of the problem and use different fitness functions for getting a good trade-off between balance of distribution of classes and performance. The study includes a taxonomy of the approaches and an overall comparison among our models and state of the art undersampling methods. The results have been contrasted by using nonparametric statistical procedures and show that evolutionary undersampling outperforms the nonevolutionary models when the degree of imbalance is increased.
Salvador García 0001, Francisco Herrera
Evol. Comput.1
2009 A study on the use of statistical tests for experimentation with neural networks: Analysis of parametric test conditions and non-parametric tests
Julián Luengo, Salvador García 0001, Francisco Herrera
Expert Syst. Appl.2
2009 Diagnose Effective Evolutionary Prototype Selection Using an Overlapping Measure
abstract
Evolutionary prototype selection has shown its effectiveness in the past in the prototype selection domain. It improves in most of the cases the results offered by classical prototype selection algorithms but its computational cost is expensive. In this paper, we analyze the behavior of the evolutionary prototype selection strategy, considering a complexity measure for classification problems based on overlapping. In addition, we have analyzed different k values for the nearest neighbour classifier in this domain of study to see its influence on the results of PS methods. The objective consists of predicting when the evolutionary prototype selection is effective for a particular problem, based on this overlapping measure.
Salvador García 0001, José Ramón Cano, Ester Bernadó-Mansilla, Francisco Herrera
Int. J. Pattern Recognit. Artif. Intell.1
2009 KEEL: a software tool to assess evolutionary algorithms for data mining problems
Jesús Alcalá-Fdez, Luciano Sánchez, Salvador García 0001, María José del Jesus, Sebastián Ventura, Josep Maria Garrell i Guiu, José Otero, Cristóbal Romero 0001, Jaume Bacardit, Víctor Manuel Rivas Santos, Juan Carlos Fernández 0001, Francisco Herrera
Soft Comput.3
2009 A study of statistical techniques and performance measures for genetics-based machine learning: accuracy and interpretability
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera
Soft Comput.1
2008 Evolutionary Training Set Selection to Optimize C4.5 in Imbalanced Problems
abstract
Classification in imbalanced domains is a recent challenge in machine learning. We refer to imbalanced classification when data presents many examples from one class and few from the other class, and the less representative class is the one which has more interest. One of the most used techniques to tackle this problem consists in preprocessing the data previously to the learning process. This preprocessing could be done through under-sampling; removing examples, mainly belonging to the majority class; and over-sampling, by means of replicating or generating new minority examples. This contribution proposes an under-sampling procedure based on evolutionary algorithms to perform a training set selection for optimizing the models obtained by the C4.5 decision tree. The proposal has been compared with other under-sampling and over-sampling techniques and the results are very competitive in terms of accuracy, and the obtained models are more interpretable.
Salvador García 0001, Francisco Herrera
HIS1
2008 Making CN2-SD subgroup discovery algorithm scalable to large size data sets using instance selection
José Ramón Cano, Francisco Herrera, Manuel Lozano 0001, Salvador García 0001
Expert Syst. Appl.4
2008 A study of the behaviour of linguistic fuzzy rule based classification systems in the framework of imbalanced data-sets
Alberto Fernández 0001, Salvador García 0001, María José del Jesus, Francisco Herrera
Fuzzy Sets Syst.2
2008 A memetic algorithm for evolutionary prototype selection: A scaling up approach
Salvador García 0001, José Ramón Cano, Francisco Herrera
Pattern Recognit.1
2008 Subgroup discover in large size data sets preprocessed using stratified instance selection for increasing the presence of minority classes
José Ramón Cano, Salvador García 0001, Francisco Herrera
Pattern Recognit. Lett.2
2006 A Proposal of Evolutionary Prototype Selection for Class Imbalance Problems
Salvador García 0001, José Ramón Cano, Alberto Fernández 0001, Francisco Herrera
IDEAL1
2006 Incorporating Knowledge in Evolutionary Prototype Selection
Salvador García 0001, José Ramón Cano, Francisco Herrera
IDEAL1