Isaac Triguero

dblp:87/8568 · DBLP profile ↗
← Back
72ranked-venue papers
15as first author
21since 2021 · last 2026
0000-0002-0150-0651ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 13 first-author · 15 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Decision-Focused Learning Enhanced by Automated Feature Engineering for Energy Storage Optimisation
abstract
Decision-making under uncertainty in energy management is complicated by unknown parameters hindering optimal strategies, particularly in Battery Energy Storage System (BESS) operations. Predict-Then-Optimise (PTO) approaches treat forecasting and optimisation as separate processes, allowing prediction errors to cascade into suboptimal decisions as models minimise forecasting errors rather than optimising downstream tasks. The emerging Decision-Focused Learning (DFL) methods overcome this limitation by integrating prediction and optimisation; however, they are relatively new and have been tested primarily on synthetic datasets with limited evidence of their practical viability. Real-world BESS applications present additional challenges, including greater variability and data scarcity due to collection constraints. Because of these challenges, this work leverages Automated Feature Engineering (AFE) to improve the nascent approach of DFL. This AFE–DFL integration automatically extracts decision-relevant features from limited energy data without requiring domain expertise, while ensuring features directly enhance BESS operational decisions rather than merely improving prediction accuracy metrics. We propose an AFE–DFL framework suitable for small datasets that forecasts electricity prices and demand while optimising BESS operations to minimise costs. We validate the framework’s effectiveness on a novel real-world UK property dataset. The evaluation compares DFL methods against PTO, with and without AFE. Results show that DFL yields lower operating costs than PTO, and adding AFE further improves DFL performance by 22.9–56.5 % compared to models without AFE. These findings provide empirical evidence for DFL’s practical viability, demonstrating that AFE-DFL integration reduces reliance on domain expertise while achieving superior economic outcomes for BESS optimisation.
Nasser Alkhulaifi, Ismail Gokay Dogan, Timothy Cargan, Alexander L. Bowler, Direnc Pekaslan, Nicholas James Watson, Isaac Triguero
Expert Syst. Appl.7
2026 Evolutionary Computation for the Design and Enrichment of General-Purpose Artificial Intelligence Systems: Survey and Prospects
abstract
In Artificial Intelligence, there is an increasing demand for adaptive models capable of dealing with a diverse spectrum of learning tasks, surpassing the limitations of systems devised to cope with a single task. The recent emergence of General-Purpose Artificial Intelligence Systems (GPAIS) poses model configuration and adaptability challenges at far greater complexity scales than the optimal design of traditional Machine Learning models. Evolutionary Computation (EC) has been a useful tool for both the design and optimization of Machine Learning models, endowing them with the capability to configure and/or adapt themselves to the task under consideration. Therefore, their application to GPAIS is a natural choice. This paper aims to analyze the role of EC in the field of GPAIS, exploring the use of EC for their design or enrichment. We also match GPAIS properties to Machine Learning areas in which EC has had a notable contribution, highlighting recent milestones of EC for GPAIS. Furthermore, we discuss the challenges of harnessing the benefits of EC for GPAIS, presenting different strategies to both design and improve GPAIS with EC, covering tangential areas, identifying research niches, and outlining potential research directions for EC and GPAIS.
Daniel Molina, Javier Poyatos, Javier Del Ser, Salvador García 0001, Hisao Ishibuchi, Isaac Triguero, Bing Xue 0001, Xin Yao 0001, Francisco Herrera
IEEE Trans. Evol. Comput.6
2025 A First Approach to Refine Semantic Spaces in Zero-Shot Learning with a Genetic Algorithm
abstract
Evolutionary computation has been successfully applied to tackle a wide variety of machine learning problems due to its generalisation and adaptability capabilities. Recently, it has shown great potential to enhance General Purpose Artificial Intelligence Systems particularly, those that work in open-world scenarios which require dynamic adaptation abilities. Zero-Shot Learning (ZSL) is an emerging paradigm within the open-world context that allows us to perform predictive tasks, such as the classification of unknown elements (i.e. unknown classes) for which a model has not been specifically trained. To do this, ZSL uses auxiliary information known as semantic space, typically in the form of attributes that define each class. This plays a crucial role in associating prior knowledge with unknown situations. However, the treatment of the semantic space has remained an underexplored area, as selecting the most relevant semantic attributes that generalise to unknown classes is a challenging problem. In this preliminary work, we propose a tailored genetic algorithm to perform feature selection of the semantic space, removing irrelevant features that negatively affect the generalisation capabilities of a well-known ZSL approach. The results on four commonly used ZSL image classification problems show that refining the semantic space may consistently boost the accuracy across all datasets.
J. J. Herrera, Francisco Herrera, Isaac Triguero
CEC3
2025 Monte Carlo-Based Interval TOPSIS for Navigating Decision Support Under Uncertainty
Jingda Ying, Christian Wagner 0002, Isaac Triguero, Shaily Kabir
IDEAL (2)3
2025 Directed Perturbations for Efficient Learning of Surrogate Losses
abstract
Decision-Focused Learning (DFL) is a paradigm to learn neural network-based predictive models tailored to a specific optimisation problem. A key challenge for DFL methods lies in the non-differentiable nature of most optimisation problems. Recent solutions use a learned, differentiable, model to act as a surrogate loss. To learn the model, the optimisation problem is solved repeatedly using random perturbations of the predictions to calculate a regret value which can be used as a target to learn the surrogate loss. However, this necessitates numerous runs of the, potentially computationally expensive, optimiser. As such, maximising the useful information from each run of the optimizer is paramount. A sample of purely random perturbations may not yield an effective distribution to learning the surrogate from. We propose using a directed perturbation strategy to generate a set of meaningful perturbations for learning a surrogate loss model. We evaluate our approach on a resource allocation problem and real-world case study focused on a critical energy challenge: optimising solar-plus-battery systems. The results show that directed perturbations learn a stronger surrogate loss with fewer runs of the optimiser, enabling a more efficient DFL.
Timothy Cargan, Dario Landa Silva, Isaac Triguero
IJCNN3
2025 Frequency-Aware Contrastive Loss for Self-Supervised EEG Seizure Detection
abstract
Contrastive Representation Learning (CRL), a sub-field of self-supervised learning, models relationships between data points by comparing pairs of samples and utilizing specifically formulated contrastive pretext tasks in which multiple views of each input example are generated via data augmentation. While contrastive learning has shown promise in EEG analysis, current methods typically rely on random negative sampling, which may not capture the subtle distinctions crucial for seizure detection. Our proposed method, Frequency-aware NT-Xent (FA-NT-Xent) loss, introduces a principled approach to hard negative selection by leveraging domain knowledge from EEG analysis. Specifically, we construct a distance matrix between batch samples based on their spectral band power profiles across five canonical frequency bands (delta, theta, alpha, beta, gamma), enabling the identification of samples that share similar frequency characteristics but belong to different classes. These hard negative examples are weighted by a learnable parameter and integrated into the contrastive loss computation, enhancing the model’s ability to distinguish between physiologically similar but clinically distinct patterns. Evaluated on the CHB-MIT dataset using a cross-patient validation scheme, our method demonstrates improvements over the standard NT-Xent loss and achieves competitive performance against state-of-the-art approaches. Visualization of the learned representations reveals improved class separation and the emergence of patient-independent seizure characteristics, indicating that our frequency-aware approach successfully captures clinically relevant features while maintaining robust cross-patient generalization.
Kristina Stemikovskaya, Diana Manjarres, Isaac Triguero
IJCNN3
2025 From Traditional Methods to GPT-based Models for 2D Video Game Level Procedural Content Generation: An Empirical Study
abstract
Procedural level generation in video games has made significant strides, yet achieving high-quality automated level design remains a major challenge. Over the years, techniques have evolved from simple constructive algorithms to advanced Artificial Intelligence (AI) models like Generative Adversarial Networks and Large Language Models. However, the lack of a standardised evaluation framework has hindered direct numerical comparisons and the ability to gauge true progress in the field. To address this gap, we propose an evaluation methodology to benchmark key generation techniques and explore the potential of general-purpose AI models. As a case study, we present a Super Mario Bros level generator powered by ChatGPT, leveraging general-purpose natural language for design tasks. The results show the different strengths and weaknesses of existing models, indicating that traditional algorithms still outperform the most advanced AI methods in this domain, highlighting the need for further innovation to bridge the gap.
Daniel Cerezo, Isaac Triguero
SMC2
2025 AutoEnergy: An automated feature engineering algorithm for energy consumption forecasting with AutoML
Nasser Alkhulaifi, Alexander L. Bowler, Direnc Pekaslan, Nicholas James Watson, Isaac Triguero
Knowl. Based Syst.5
2024 A Wearable Eye-Tracking Approach for Early Autism Detection with Machine Learning: Unravelling Challenges and Opportunities
abstract
Early detection of Autism Spectrum Disorder (ASD) is crucial to facilitate timely interventions, improve outcomes, and enhance the quality of life for individuals on the spectrum. Artificial intelligence and machine learning have significantly advanced the study of ASD by enabling sophisticated analyses of complex behavioural data, providing more accurate and timely detection methods. Nevertheless, existing research is mostly focused on either screening manual methods via psychological surveys, and/or the application of functional Magnetic Resonance Imaging. In the former case, there is a limitation due the subjectivity of the procedure; whereas in the latter the high cost represents a barrier to widespread use in diagnosis. In consideration of the aforementioned factors, this research contributes significantly to the field by integrating commodity wearable technology with psycho-healthcare, employing eye-tracking capabilities as a novel, non-invasive approach for early ASD diagnosis. Our methodology stands out by offering a comprehensive, four-step process that includes sophisticated image preprocessing, innovative feature extraction, precise classification, and a novel multi-instance aggregation strategy. We show its potential via a compelling case study that attests to the feasibility and efficacy of this innovative paradigm. Additionally, the work underscores prevailing challenges in the field, stressing some factors such as the acquisition of extensive and diverse datasets to allow for a multi-modal approach and the application of Trustworthy Artificial Intelligence.
J. Lopez-Martinez, Purificación Checa, José M. Soto-Hidalgo, Isaac Triguero, Alberto Fernández 0001
IJCNN4
2024 Exploring Automated Feature Engineering for Energy Consumption Forecasting with AutoML
abstract
Machine learning methods are widely used to predict energy consumption, aiming to enhance efficiency and support environmental goals. However, developing these models is traditionally time-consuming and expert-dependent. While Automated Machine Learning (AutoML) has emerged as a valuable approach to streamlining machine learning pipelines, including appropriate preprocessing and learning algorithms, it may still require human experts. These experts might be needed to generate new, interpretable features that could significantly improve model performance. This is particularly relevant in complex settings such as the energy domain, where deep learning's automatic feature extraction lacks interpretability. To address this challenge, this exploratory work introduces an automated feature engineering method tailored for energy forecasting problems. It involves generating a comprehensive set of features that can be fed into AutoML, thereby reducing the need for domain knowledge in feature engineering. The proposed method has been validated using eleven datasets from various energy domains, including residential buildings, renewable energy, and regional energy consumption, with state-of-the-art AutoML methods, namely H20, TPOT, AutoGluon, and FLAML. The results demonstrate a noticeable reduction in prediction errors across all the examined datasets.
Nasser Alkhulaifi, Alexander L. Bowler, Direnc Pekaslan, Isaac Triguero, Nicholas James Watson
SMC4
2024 Local-global methods for generalised solar irradiance forecasting
abstract
Abstract For efficient operation, solar power operators often require generation forecasts for multiple sites with varying data availability. Many proposed methods for forecasting solar irradiance / solar power production formulate the problem as a time-series, using current observations to generate forecasts. This necessitates a real-time data stream and enough historical observations at every location for these methods to be deployed. In this paper, we propose the use of Global methods to train generalised models. Using data from 20 locations distributed throughout the UK, we show that it is possible to learn models without access to data for all locations, enabling them to generate forecasts for unseen locations. We show a single Global model trained on multiple locations can produce more consistent and accurate results across locations. Furthermore, by leveraging weather observations and measurements from other locations we show it is possible to create models capable of accurately forecasting irradiance at locations without any real-time data. We apply our approaches to both classical and state-of-the-art Machine Learning methods, including a Transformer architecture. We compare models using satellite imagery or point observations (temperature, pressure, etc.) as weather data. These methods could facilitate planning and optimisation for both newly deployed solar farms and domestic installations from the moment they come online.
Timothy Cargan, Dario Landa Silva, Isaac Triguero
Appl. Intell.3
2024 SEGAL time series classification - Stable explanations using a generative model and an adaptive weighting method for LIME
abstract
Local Interpretability Model-agnostic Explanations (LIME) is a well-known post-hoc technique for explaining black-box models. While very useful, recent research highlights challenges around the explanations generated. In particular, there is a potential lack of stability, where the explanations provided vary over repeated runs of the algorithm, casting doubt on their reliability. This paper investigates the stability of LIME when applied to multivariate time series classification. We demonstrate that the traditional methods for generating neighbours used in LIME carry a high risk of creating 'fake' neighbours, which are out-of-distribution in respect to the trained model and far away from the input to be explained. This risk is particularly pronounced for time series data because of their substantial temporal dependencies. We discuss how these out-of-distribution neighbours contribute to unstable explanations. Furthermore, LIME weights neighbours based on user-defined hyperparameters which are problem-dependent and hard to tune. We show how unsuitable hyperparameters can impact the stability of explanations. We propose a two-fold approach to address these issues. First, a generative model is employed to approximate the distribution of the training data set, from which within-distribution samples and thus meaningful neighbours can be created for LIME. Second, an adaptive weighting method is designed in which the hyperparameters are easier to tune than those of the traditional method. Experiments on real-world data sets demonstrate the effectiveness of the proposed method in providing more stable explanations using the LIME framework. In addition, in-depth discussions are provided on the reasons behind these results.
Han Meng, Christian Wagner 0002, Isaac Triguero
Neural Networks3
2023 Hyper-Stacked: Scalable and Distributed Approach to AutoML for Big Data
Ryan Dave, Juan S. Angarita-Zapata, Isaac Triguero
CD-MAKE3
2023 Identifying bird species by their calls in Soundscapes
abstract
Abstract In many real data science problems, it is common to encounter a domain mismatch between the training and testing datasets, which means that solutions designed for one may not transfer well to the other due to their differences. An example of such was in the BirdCLEF2021 Kaggle competition, where participants had to identify all bird species that could be heard in audio recordings. Thus, multi-label classifiers, capable of coping with domain mismatch, were required. In addition, classifiers needed to be resilient to a long-tailed (imbalanced) class distribution and weak labels. Throughout the competition, a diverse range of solutions based on convolutional neural networks were proposed. However, it is unclear how different solution components contribute to overall performance. In this work, we contextualise the problem with respect to the previously existing literature, analysing and discussing the choices made by the different participants. We also propose a modular solution architecture to empirically quantify the effects of different architectures. The results of this study provide insights into which components worked well for this challenge.
Kyle D. S. Maclean, Isaac Triguero
Appl. Intell.2
2023 Explaining time series classifiers through meaningful perturbation and optimisation
abstract
Machine learning approaches have enabled increasingly powerful time series classifiers. While performance has improved drastically, the resulting classifiers generally suffer from poor explainability, limiting their applicability in critical areas. Saliency-based methods designed to highlight the critical features are one of the most promising approaches to improving this explainability. Here, current techniques commonly rely on artificially perturbing the features, using, for example, random noise or ‘zeroing’ these features. We first demonstrate that an important drawback of these methods is that the perturbations used can result in unrealistic assessments of the classifier, since the perturbations force the data outside their original distribution. We articulate how this can result in poor identification of critical features, and hence misleading explanations. In order to address this issue and identify the most important features for the output of a black-box model, we propose a dual approach through meaningful perturbation and optimisation. First, leveraging a mechanism originally proposed in image analysis, a generative model is trained to create within-distribution perturbations of the input. These are then used to reliably evaluate whether a set of features is critical. Second, a greedy-based segmentation and identification strategy is proposed to search for the smallest set of critical features. Experiments show that the proposed approach addresses the out-of-distribution problem and identifies fewer critical features than existing methods. In combination, both aspects of the proposed approach offer a qualitative advance towards generating meaningful and robust explanations in the context of time series classification.
Han Meng, Christian Wagner 0002, Isaac Triguero
Inf. Sci.3
2022 Feature Importance Identification for Time Series Classifiers
abstract
Time series classification is a challenging research area where machine learning techniques such as deep learning perform well, yet lack interpretability. Identifying the most important features for such classifiers provides a pathway to improving their interpretability. Several Feature Importance (FI) identification methods remove the contributions of features, i.e. observations at certain time steps of, from the input and evaluate the change in the classification result to measure the importance of features. As time series features cannot simply be deleted, current techniques generally rely on replacing features with constant or random values. While effective, this approach risks unexpected results in the classification and thus feature importance estimation-as the replacements used may be different to what the classifier encountered in the training phase. This is referred to as the Out-Of-Distribution problem. The OOD problem has been recognised in image and language models but have not received much attention in the context of time series classification. This work addresses the OOD problem in FI identification for time series classifiers. Specifically, we propose a method based on Conditional Variational Autoencoder to generate possible sets of within-distribution inputs, which are used to evaluate feature importance through marginalisation. Experiments on publicly accessible datasets are carried out showing that the method identifies the most important features with higher accuracy than existing methods, providing the basis for improved explainability of time series classifiers.
Han Meng, Christian Wagner 0002, Isaac Triguero
SMC3
2021 Few-Shot Learning for Postnatal Gestational Age Estimation
abstract
A baby's gestational age determines whether or not they are premature, which helps clinicians decide treatment. The most accurate dating methods use Ultrasound Scans, but these are expensive, require trained personnel and cannot always be deployed to remote areas. In the absence of such methods, the Ballard Score, a postnatal clinical examination, can be used. However, this method is highly subjective and results vary widely depending on the experience of the examiner. In the last decade, there have been efforts to exploit machine learning methods to create reliable postnatal methods deployable anywhere in the world. However, current state of the art methods remain influenced by the quality of their input data, which is a major issue in areas where data collection is difficult or impossible. Gathering images of newborns, especially those who are premature, is a very challenging task, due to the intrusiveness of taking photographs inside an ICU. This paper explores few-shot learning for gestational age estimation and investigates whether it can be used as an alternative way to explicitly deal with a lack of data. We compare two popular methods for few-shot learning, Model Agnostic Meta-Learning and Prototypical Networks, with a small dataset collected as part of the GesATional Project. Experimental results show that few-shot learning can be used to estimate gestational age postnatally. As a novel contribution to deal with the lack of data, models are trained with CelebFaces Attributes, to aid the procedure. We demonstrate that the novel integration of additional meta-sets improves performance by an average of 2.3% for both models comparing to just using a single dataset.
Stepan Romanov, Heda Song, Michel F. Valstar, Don Sharkey, Caz Henry, Isaac Triguero, Mercedes Torres Torres
IJCNN6
2021 Neurocomputing guest editorial for the special issue: Advances in deep and shallow machine learning approaches for handling data irregularities
Swagatam Das, Salvador García 0001, Isaac Triguero
Neurocomputing3
2021 L2AE-D: Learning to Aggregate Embeddings for Few-shot Learning with Meta-level Dropout
Heda Song, Mercedes Torres Torres, Ender Özcan, Isaac Triguero
Neurocomputing4
2021 Beyond global and local multi-target learning
Márcio P. Basgalupp, Ricardo Cerri, Leander Schietgat, Isaac Triguero, Celine Vens
Inf. Sci.4
2021 Multigranulation Supertrust Model for Attribute Reduction
abstract
As big data often contains a significant amount of uncertain, unstructured, and imprecise data that are structurally complex and incomplete, traditional attribute reduction methods are less effective when applied to large-scale incomplete information systems to extract knowledge. Multigranular computing provides a powerful tool for use in big data analysis conducted at different levels of information granularity. In this article, we present a novel multigranulation supertrust fuzzy-rough set-based attribute reduction (MSFAR) algorithm to support the formation of hierarchies of information granules of higher types and higher orders, which addresses newly emerging data mining problems in big data analysis. First, a multigranulation supertrust model based on the valued tolerance relation is constructed to identify the fuzzy similarity of the changing knowledge granularity with multimodality attributes. Second, an ensemble consensus compensatory scheme was adopted to calculate the multigranular trust degree based on the reputation at different granularities to create reasonable subproblems with different granulation levels. Third, an equilibrium method of multigranular coevolution is employed to ensure a wide range of balancing of exploration and exploitation, and this strategy can classify super elitists' preferences and detect noncooperative behaviors with a global convergence ability and high search accuracy. The experimental results demonstrate that the MSFAR algorithm achieves a high performance in addressing uncertain and fuzzy attribute reduction problems with a large number of multigranularity variables.
Weiping Ding 0001, Witold Pedrycz, Isaac Triguero, Zehong Cao, Chin-Teng Lin
IEEE Trans. Fuzzy Syst.3
2020 A Hybrid Surrogate Model for Evolutionary Undersampling in Imbalanced Classification
abstract
Data preprocessing is a key stage in data mining that allows machine learning algorithms to obtain meaningful insights. Many preprocessing problems such as feature selection or instance selection can be modelled as optimisation/search problems. Evolutionary algorithms have traditionally excelled in this task when dealing with data of a moderate size. However, their application to large datasets typically involves very high computational costs. In this work, we propose a hybrid surrogate model for evolutionary undersampling in imbalanced classification problems. These are characterised by having a highly skewed distribution of classes in which evolutionary algorithms aim to balance the training data by selecting only the most relevant data. The proposed technique combines a two-stage clustering-based surrogate method with a windowing approach to quickly approximate fitness values of the chromosomes and accelerate the search. The experiments carried out in 44 standard imbalanced datasets show that the proposed hybrid surrogate model highly reduces the computational cost of the evolutionary algorithm without a considerable loss of performance.
Hoang Lam Le, Dario Landa Silva, Mikel Galar, Salvador García 0001, Isaac Triguero
CEC5
2020 A Local Search with a Surrogate Assisted Option for Instance Reduction
Ferrante Neri, Isaac Triguero
EvoApplications2
2020 Chi-BD-DRF: Design of Scalable Fuzzy Classifiers for Big Data via A Dynamic Rule Filtering Approach
abstract
Big data classification problems are known to be no longer addressable by sequential algorithms. Therefore, it is necessary to design and develop novel solutions to provide accurate yet interpretable models in a tolerable elapsed time. In this area, Fuzzy Rule-Based Classification Systems are very advantageous due to their intrinsic interpretable and accurate capabilities. However, when these systems are applied in Big Data scenarios, the size of the rule set can become too large to be useful, whereas many of the generated rules could be associated with the non-dense areas or outliers. The presence of such rules in the rule base not only increases the running time and computation overheads but also affects on the interpretability of the fuzzy system. In this contribution, we propose a novel approach to obtain compact and accurate fuzzy models for Big data problems in a linearly scalable complex time. To do so, a dynamic filtering approach is applied to remove low supporting rules. Moreover, an efficient computation of the rules' weights is presented to improve the accuracy of the predictions. This model is developed for Big Data analytics by using Apache Spark framework. This allows taking advantage of the built-in resources and directives for a transparent distributed computing, as well as the machine learning pipeline to ease the complete processing. Experimental results, using different Big Data problems, confirmed the goodness of the proposed algorithm with respect to the baseline fuzzy classifier.
Fatemeh Aghaeipoor, Mohammad Masoud Javidi, Isaac Triguero, Alberto Fernández 0001
FUZZ-IEEE3
2020 General-Purpose Automated Machine Learning for Transportation: A Case Study of Auto-sklearn for Traffic Forecasting
Juan S. Angarita-Zapata, Antonio D. Masegosa, Isaac Triguero
IPMU (2)3
2020 Current trends of granular data mining for biomedical data analysis
Weiping Ding 0001, Chin-Teng Lin, Alan Wee-Chung Liew, Isaac Triguero, Wenjian Luo
Inf. Sci.4
2020 Fast and Scalable Approaches to Accelerate the Fuzzy k-Nearest Neighbors Classifier for Big Data
abstract
One of the best-known and most effective methods in supervised classification is the k-nearest neighbors algorithm (kNN). Several approaches have been proposed to improve its accuracy, where fuzzy approaches prove to be among the most successful, highlighting the classical fuzzy k-nearest neighbors (FkNN). However, these traditional algorithms fail to tackle the large amounts of data that are available today. There are multiple alternatives to enable kNN classification in big datasets, spotlighting the approximate version of kNN called hybrid spill tree. Nevertheless, the existing proposals of FkNN for big data problems are not fully scalable, because a high computational load is required to obtain the same behavior as the original FkNN algorithm. This article proposes global approximate hybrid spill tree FkNN and local hybrid spill tree FkNN, two approximate approaches that speed up runtime without losing quality in the classification process. The experimentation compares various FkNN approaches for big data with datasets of up to 11 million instances. The results show an improvement in runtime and accuracy over literature algorithms.
Jesús Maillo, Salvador García 0001, Julián Luengo, Francisco Herrera, Isaac Triguero
IEEE Trans. Fuzzy Syst.5
2019 Evolving Deep CNN-LSTMs for Inventory Time Series Prediction
abstract
Inventory forecasting is a key component of effective inventory management. In this work, we utilise hybrid deep learning models for inventory forecasting. According to the highly nonlinear and non-stationary characteristics of inventory data, the models employ Long Short-Term Memory (LSTM) to capture long temporal dependencies and Convolutional Neural Network (CNN) to learn the local trend features. However, designing optimal CNN-LSTM network architecture and tuning parameters can be challenging and would require consistent human supervision. To automate optimal architecture searching of CNN-LSTM, we implement three meta-heuristics: a Particle Swarm Optimisation (PSO) and two Differential Evolution (DE) variants. Computational experiments on real-world inventory forecasting problems are conducted to evaluate the performance of the applied meta-heuristics in terms of evolved network architectures for obtaining prediction accuracy. Moreover, the evolved CNN-LSTM models are also compared to Seasonal Auto-regressive Integrated Moving Average (SARIMA) models for inventory forecasting problems. The experimental results indicate that the evolved CNN-LSTM models are capable of dealing with complex nonlinear inventory forecasting problem.
Ning Xue, Isaac Triguero, Grazziela Patrocinio Figueredo, Dario Landa Silva
CEC2
2019 A Preliminary Approach for the Exploitation of Citizen Science Data for Fast and Robust Fuzzy k-Nearest Neighbour Classification
abstract
Citizen science is becoming mainstream in a wide variety of real-world applications in astronomy or bioinformatics, in which, for example, classification tasks by experts are very time consuming. These projects engage amateur volunteers that are tasked to manually classify unannotated examples. As a result, we obtain a larger volume of labelled data that, however, contains a great level of uncertainty due to the wide range of expertise of the volunteers. Handling that inherent uncertainty is key to building robust and fast machine learning models that maximise the outcome of citizen science projects. In this work, we introduce a preliminary approach that first transforms the original results from a citizen science project to handle the uncertainty, and then uses this as input to a fuzzy k-nearest neighbour classifier. We leverage citizen science results in such a way that it naturally speeds up the learning and classification phases of the fuzzy classifier, and improves the classification performance. As a case study, we will focus on the Galaxy Zoo project that consisted of galaxy image classification. Our experimental results show that an appropriate use of citizen science data enables a faster and more robust classification using the fuzzy k-nearest neighbour classifier.
Manuel Jiménez, Mercedes Torres Torres, Robert Ivor John, Isaac Triguero
FUZZ-IEEE4
2019 Fuzzy Hot Spot Identification for Big Data: An Initial Approach
abstract
Hot spot identification problems are present across a wide range of areas, such as transportation, health care and energy. Hot spots are locations where a certain type of event occurs with high frequency. A recent big data approach is capable of identifying hot spots in a dynamic manner, through the processing of large volumes of sensor data arriving as a stream. However, the method may produce imprecise results due to its crisp interpretation of hot spot locations and reliance on a fixed hot spot radius value. This paper presents an initial approach to addressing this shortcoming through incorporating the concept of fuzzy hot spots into the process. Experimental results on large real-world transportation datasets demonstrate the improved way in which this approach handles uncertainty in the definition of hot spots, and highlight promising future research areas for further application of fuzzy systems to the hot spot identification problem.
Rebecca Tickle, Isaac Triguero, Grazziela Patrocinio Figueredo, Ender Özcan, Mohammad Mesgarpour, Robert Ivor John
FUZZ-IEEE2
2019 A Simulation-based Optimisation Approach for Inventory Management of Highly Perishable Food
abstract
The taste and freshness of perishable foods decrease dramatically with time. Effective inventory management requires understanding of market demand as well as balancing customers needs and references with products’ shelf life. The objective is to avoid food overproduction as this leads to waste and value loss. In addition, product depletion has to be minimised, as it can result in customers reneging. This study tackles the production planning of highly perishable foods (such as freshly prepared dishes, sandwiches and desserts with shelf life varying from 6 to 12 hours), in an environment with highly variable customers demand. In the scenario considered here, the planning horizon is longer than the products’ shelf life. Therefore, food needs to be replenished several times at different intervals. Furthermore, customers demand varies significantly during the planning period. We tackle the problem by combining discrete-event simulation and particle swarm optimisation (PSO). The simulation model focuses on the behaviour of the system as parameters (i.e. replenishment time and quantity) change. PSO is employed to determine the best combination of parameter values for the simulations. The effectiveness of the proposed approach is applied to some real-world scenario corresponding to a local food shop. Experimental results show that the proposed methodology combining discrete event simulation and particle swarm optimisation is effective for inventory management of highly perishable foods with variable customers demand.
Ning Xue, Dario Landa Silva, Grazziela Patrocinio Figueredo, Isaac Triguero
ICORES4
2019 Multi-head CNN-RNN for multi-time series anomaly detection: An industrial case study
Mikel Canizo, Isaac Triguero, Angel Conde, Enrique Onieva
Neurocomputing2
2019 Handling uncertainty in citizen science data: Towards an improved amateur-based large-scale classification
Manuel Jiménez, Isaac Triguero, Robert Ivor John
Inf. Sci.2
2019 Instance reduction for one-class classification
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
Knowl. Inf. Syst.2
2019 Introduction to the Special Issue on Human-interaction-aware Data Analytics for Cyber-physical Systems
abstract
No abstract available.
Tongquan Wei, Junlong Zhou, Rajiv Ranjan 0001, Isaac Triguero, Huafeng Yu, Chun Jason Xue, Schahram Dustdar
ACM Trans. Cyber Phys. Syst.4
2018 A Preliminary Study of the Feasibility of Global Evolutionary Feature Selection for Big Datasets under Apache Spark
abstract
Designing efficient learning models capable of dealing with tons of data has become a reality in the era of big data. However, the amount of available data is too much for traditional data mining techniques to be applicable. This issue is even more serious when evolutionary algorithms are a key part of the learning algorithm. In this scenario, one typical approach is to follow a divide-and-conquer strategy, where data is divided into different chunks that are individually and independently addressed. Afterwards, the partial knowledge obtained from each chunk of data is combined in order to give a solution to the problem. Nevertheless, these kinds of local approaches do not look at data as a whole, missing a global view of the problem, which may result in less accurate models that also depend on how data is split. In this work, we focus on evolutionary feature selection algorithms. A divide-and-conquer approach to handle evolutionary feature selection in big data was already developed. We aim at designing its global counterpart, which looks at the feature selection problem from a global perspective, making use of the data as a whole to select the most appropriate features. In order to do so, we consider Apache Spark as a big data technology where our algorithm is implemented. We design a genetic algorithm capable of dealing with big datasets by selecting the proper parameters for our base algorithm (the well-known CHC) and adapting the evaluation procedure to take all the distributed data into account. Several preliminary results are discussed to study the feasibility of global evolutionary feature selection methods for big datasets.
Mikel Galar, Isaac Triguero, Humberto Bustince, Francisco Herrera
CEC2
2018 A Genetic Algorithm With Composite Chromosome for Shift Assignment of Part-time Employees
abstract
Personnel scheduling problems involve multiple tasks, including assigning shifts to workers. The purpose is usually to satisfy objectives and constraints arising from management, labour unions and employee preferences. The shift assignment problem is usually highly constrained and difficult to solve. The problem can be further complicated (i) if workers have mixed skills; (ii) if the start/end times of shifts are flexible; and (iii) if multiple criteria are considered when evaluating the quality of the assignment. This paper proposes a genetic algorithm using composite chromosome encoding to tackle the shift assignment problem that typically arises in retail stores, where most employees work part-time, have mixed-skills and require flexible shifts. Experiments on a number of problem instances extracted from a real-world retail store, show the effectiveness of the proposed approach in finding good-quality solutions. The computational results presented here also include a comparison with results obtained by formulating the problem as a mixed-integer linear programming model and then solving it with a commercial solver. Results show that the proposed genetic algorithm exhibits an effective and efficient performance in solving this difficult optimisation problem.
Ning Xue, Dario Landa Silva, Isaac Triguero, Grazziela Patrocinio Figueredo
CEC3
2018 A preliminary study on Hybrid Spill-Tree Fuzzy k-Nearest Neighbors for big data classification
abstract
The Fuzzy k Nearest Neighbor (Fuzzy kNN) classifier is well known for its effectiveness in supervised learning problems. kNN classifies by comparing new incoming examples with a similarity function using the samples of the training set. The fuzzy version of the kNN accounts for the underlying uncertainty in the class labels, and it is composed of two different stages. The first one is responsible for calculating the fuzzy membership degree for each sample of the problem in order to obtain smoother boundaries between classes. The second stage classifies similarly to the standard kNN algorithm but uses the previously calculated class membership degree. To deal with very large datasets, distributed versions of the Fuzzy kNN algorithm have been proposed. However, existing approaches remain not fully scalable as they aim to replicate the exact behavior of the Fuzzy kNN. In this work, we present an approximate and distributed Fuzzy kNN approach based on Hybrid Spill-Tree implemented under Apache Spark. The aim of this model is to alleviate the scalability problems and to deal with big datasets maintaining high accuracy. In our experiments, we compare in precision and runtime with the Fuzzy kNN for big data problems existing in the literature, running with datasets of up to 11 million instances. The results show an improvement in the runtime and accuracy with respect to the previous exact model.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE5
2018 On the use of convolutional neural networks for robust classification of multiple fingerprint captures
abstract
Fingerprint classification is one of the most common approaches to accelerate the identification in large databases of fingerprints. Fingerprints are grouped into disjoint classes, so that an input fingerprint is compared only with those belonging to the predicted class, reducing the penetration rate of the search. The classification procedure usually starts by the extraction of features from the fingerprint image, frequently based on visual characteristics. In this work, we propose an approach to fingerprint classification using convolutional neural networks, which avoid the necessity of an explicit feature extraction process by incorporating the image processing within the training of the classifier. Furthermore, such an approach is able to predict a class even for low-quality fingerprints that are rejected by commonly used algorithms, such as FingerCode. The study gives special importance to the robustness of the classification for different impressions of the same fingerprint, aiming to minimize the penetration in the database. In our experiments, convolutional neural networks yielded better accuracy and penetration rate than state-of-the-art classifiers based on explicit feature extraction. The tested networks also improved on the runtime, as a result of the joint optimization of both feature extraction and classification.
Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera
Int. J. Intell. Syst.2
2018 Self-labeling techniques for semi-supervised time series classification: an empirical study
Mabel González Castellanos, Christoph Bergmeir, Isaac Triguero, Yanet Rodríguez, José Manuel Benítez 0001
Knowl. Inf. Syst.3
2017 A first attempt on global evolutionary undersampling for imbalanced big data
abstract
The design of efficient big data learning models has become a common need in a great number of applications. The massive amounts of available data may hinder the use of traditional data mining techniques, especially when evolutionary algorithms are involved as a key step. Existing solutions typically follow a divide-and-conquer approach in which the data is split into several chunks that are addressed individually. Next, the partial knowledge acquired from every slice of data is aggregated in multiple ways to solve the entire problem. However, these approaches are missing a global view of the data as a whole, which may result in less accurate models. In this work we carry out a first attempt on the design of a global evolutionary undersampling model for imbalanced classification problems. These are characterised by having a highly skewed distribution of classes in which evolutionary models are being used to balance the dataset by selecting only the most relevant data. Using Apache Spark as big data technology, we have introduced a number of variations to the well-known CHC algorithm to work with very large chromosomes and reduce the costs associated to the fitness evaluation. We discuss some preliminary results, showing the great potential of this new kind of evolutionary big data model.
Isaac Triguero, Mikel Galar, Humberto Bustince, Francisco Herrera
CEC1
2017 Exact fuzzy k-nearest neighbor classification for big datasets
abstract
The k-Nearest Neighbors (kNN) classifier is one of the most effective methods in supervised learning problems. It classifies unseen cases comparing their similarity with the training data. Nevertheless, it gives to each labeled sample the same importance to classify. There are several approaches to enhance its precision, with the Fuzzy k-Nearest Neighbors (Fuzzy-kNN) classifier being among the most successful ones. Fuzzy-kNN computes a fuzzy degree of membership of each instance to the classes of the problem. As a result, it generates smoother borders between classes. Apart from the existing kNN approach to handle big datasets, there is not a fuzzy variant to manage that volume of data. Nevertheless, calculating this class membership adds an extra computational cost becoming even less scalable to tackle large datasets because of memory needs and high runtime. In this work, we present an exact and distributed approach to run the Fuzzy-kNN classifier on big datasets based on Spark, which provides the same precision than the original algorithm. It presents two separately stages. The first stage transforms the training set adding the class membership degrees. The second stage classifies with the kNN algorithm the test set using the class membership computed previously. In our experiments, we study the scaling-up capabilities of the proposed approach with datasets up to 11 million instances, showing promising results.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE5
2017 kNN-IS: An Iterative Spark-based design of the k-Nearest Neighbors classifier for big data
Jesús Maillo, Sergio Ramírez-Gallego, Isaac Triguero, Francisco Herrera
Knowl. Based Syst.3
2017 Distributed incremental fingerprint identification with reduced database penetration rate using a hierarchical classification based on feature fusion and selection
Daniel Peralta, Isaac Triguero, Salvador García 0001, Yvan Saeys, José Manuel Benítez 0001, Francisco Herrera
Knowl. Based Syst.2
2016 Evolutionary undersampling for extremely imbalanced big data classification under apache spark
abstract
The classification of datasets with a skewed class distribution is an important problem in data mining. Evolutionary undersampling of the majority class has proved to be a successful approach to tackle this issue. Such a challenging task may become even more difficult when the number of the majority class examples is very big. In this scenario, the use of the evolutionary model becomes unpractical due to the memory and time constrictions. Divide-and-conquer approaches based on the MapReduce paradigm have already been proposed to handle this type of problems by dividing data into multiple subsets. However, in extremely imbalanced cases, these models may suffer from a lack of density from the minority class in the subsets considered. Aiming at addressing this problem, in this contribution we provide a new big data scheme based on the new emerging technology Apache Spark to tackle highly imbalanced datasets. We take advantage of its in-memory operations to diminish the effect of the small sample size. The key point of this proposal lies in the independent management of majority and minority class examples, allowing us to keep a higher number of minority class examples in each subset. In our experiments, we analyze the proposed model with several data sets with up to 17 million instances. The results show the goodness of this evolutionary undersampling model for extremely imbalanced big data classification.
Isaac Triguero, Mikel Galar, D. Merino, Jesús Maillo, Humberto Bustince, Francisco Herrera
CEC1
2016 EPRENNID: An evolutionary prototype reduction based ensemble for nearest neighbor classification of imbalanced data
Sarah Vluymans, Isaac Triguero, Chris Cornelis, Yvan Saeys
Neurocomputing2
2016 On the stopping criteria for k-Nearest Neighbor in positive unlabeled time series classification problems
Mabel González Castellanos, Christoph Bergmeir, Isaac Triguero, Yanet Rodríguez, José Manuel Benítez 0001
Inf. Sci.3
2016 Labelling strategies for hierarchical multi-label classification techniques
Isaac Triguero, Celine Vens
Pattern Recognit.1
2015 Evolutionary undersampling for imbalanced big data classification
abstract
Classification techniques in the big data scenario are in high demand in a wide variety of applications. The huge increment of available data may limit the applicability of most of the standard techniques. This problem becomes even more difficult when the class distribution is skewed, the topic known as imbalanced big data classification. Evolutionary undersampling techniques have shown to be a very promising solution to deal with the class imbalance problem. However, their practical application is limited to problems with no more than tens of thousands of instances. In this contribution we design a parallel model to enable evolutionary undersampling methods to deal with large-scale problems. To do this, we rely on a MapReduce scheme that distributes the functioning of these kinds of algorithms in a cluster of computing elements. Moreover, we develop a windowing approach for class imbalance data in order to speed up the undersampling process without losing accuracy. In our experiments we test the capabilities of the proposed scheme with several data sets with up to 4 million instances. The results show promising scalability abilities for evolutionary undersampling within the proposed framework.
Isaac Triguero, Mikel Galar, Sarah Vluymans, Chris Cornelis, Humberto Bustince, Francisco Herrera, Yvan Saeys
CEC1
2015 MRPR: A MapReduce solution for prototype reduction in big data classification
Isaac Triguero, Daniel Peralta, Jaume Bacardit, Salvador García 0001, Francisco Herrera
Neurocomputing1
2015 A survey on fingerprint minutiae-based local matching for verification and identification: Taxonomy and experimental evaluation
Daniel Peralta, Mikel Galar, Isaac Triguero, Daniel Paternain, Salvador García 0001, Edurne Barrenechea Tartas, José Manuel Benítez 0001, Humberto Bustince, Francisco Herrera
Inf. Sci.3
2015 Self-labeled techniques for semi-supervised learning: taxonomy, software and empirical study
Isaac Triguero, Salvador García 0001, Francisco Herrera
Knowl. Inf. Syst.1
2015 A survey of fingerprint classification Part I: Taxonomies on feature extraction methods and learning models
Mikel Galar, Joaquín Derrac, Daniel Peralta, Isaac Triguero, Daniel Paternain, Carlos Lopez-Molina, Salvador García 0001, José Manuel Benítez 0001, Miguel Pagola, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera
Knowl. Based Syst.4
2015 A survey of fingerprint classification Part II: Experimental analysis and ensemble proposal
Mikel Galar, Joaquín Derrac, Daniel Peralta, Isaac Triguero, Daniel Paternain, Carlos Lopez-Molina, Salvador García 0001, José Manuel Benítez 0001, Miguel Pagola, Edurne Barrenechea Tartas, Humberto Bustince, Francisco Herrera
Knowl. Based Syst.4
2015 ROSEFW-RF: The winner algorithm for the ECBDL'14 big data competition: An extremely imbalanced big data bioinformatics problem
Isaac Triguero, Sara del Río, Victoria López, Jaume Bacardit, José Manuel Benítez 0001, Francisco Herrera
Knowl. Based Syst.1
2015 SEG-SSC: A Framework Based on Synthetic Examples Generation for Self-Labeled Semi-Supervised Classification
abstract
Self-labeled techniques are semi-supervised classification methods that address the shortage of labeled examples via a self-learning process based on supervised models. They progressively classify unlabeled data and use them to modify the hypothesis learned from labeled samples. Most relevant proposals are currently inspired by boosting schemes to iteratively enlarge the labeled set. Despite their effectiveness, these methods are constrained by the number of labeled examples and their distribution, which in many cases is sparse and scattered. The aim of this paper is to design a framework, named synthetic examples generation for self-labeled semi-supervised classification, to improve the classification performance of any given self-labeled method by using synthetic labeled data. These are generated via an oversampling technique and a positioning adjustment model that use both labeled and unlabeled examples as reference. Next, these examples are incorporated in the main stages of the self-labeling process. The principal aspects of the proposed framework are: 1) introducing diversity to the multiple classifiers used by using more (new) labeled data; 2) fulfilling labeled data distribution with the aid of unlabeled data; and 3) being applicable to any kind of self-labeled method. In our empirical studies, we have applied this scheme to four recent self-labeled methods, testing their capabilities with a large number of data sets. We show that this framework significantly improves the classification capabilities of self-labeled techniques.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Cybern.1
2014 A first attempt on evolutionary prototype reduction for nearest neighbor one-class classification
abstract
Evolutionary prototype reduction techniques are data preprocessing methods originally developed to enhance the nearest neighbor rule. They reduce the training data by selecting or generating representative examples of a given problem. These algorithms have been designed and widely analyzed in standard classification providing very competitive results. However, its application scope can be extended to many other specific domains, such as one-class classification, in which its way of working is very interesting in order to reduce computational complexity and sensitivity to noisy data. In this contribution, we perform a first study on the usefulness of evolutionary prototype reduction methods for one-class classification. To do so, we will focus on two recent evolutionary approaches that follow very different strategies: selection and generation of examples from the training data. Both alternatives provide a resulting preprocessed data set that will be used later by a nearest neighbor one-class classifier as its training data. The results achieved support that these data reduction techniques are suitable tools to improve the performance of the nearest neighbor one-class classification.
Bartosz Krawczyk, Isaac Triguero, Salvador García 0001, Michal Wozniak 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation2
2014 A combined MapReduce-windowing two-level parallel scheme for evolutionary prototype generation
abstract
Evolutionary prototype generation techniques have demonstrated their usefulness to improve the capabilities of the nearest neighbor classifier. They act as data reduction algorithms by generating representative points of a given problem. Their main purposes are to speed up the classification process and to reduce the storage requirements and sensitivity to noise of the nearest neighbor rule. Nowadays, with the increment of available data, the use of this kind of reduction techniques becomes more important. However, their applicability can be limited to problems with no more than tens of thousands of instances. In order to address this limitation, in this work we develop a two-level parallelization scheme for evolutionary prototype generation methods. Firstly, it distributes the functioning of these algorithms in several tasks based on a MapReduce framework. Then, for each one of these tasks (mappers), we accelerate the prototype generation process by using a windowing approach. This model enables evolutionary prototype generation algorithms to be applied over large-scale classification problems without accuracy loss. Our preliminary experiments using a dataset of 1 million instances show that this proposal is an appropriate tool to improve the performance of the nearest neighbor classifier with big data.
Isaac Triguero, Daniel Peralta, Jaume Bacardit, Salvador García 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation1
2014 Minutiae filtering to improve both efficacy and efficiency of fingerprint matching algorithms
Daniel Peralta, Mikel Galar, Isaac Triguero, Oscar Miguel-Hurtado, José Manuel Benítez 0001, Francisco Herrera
Eng. Appl. Artif. Intell.3
2014 Addressing imbalanced classification with instance generation techniques: IPADE-ID
Victoria López, Isaac Triguero, Cristóbal J. Carmona, Salvador García 0001, Francisco Herrera
Neurocomputing2
2014 On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification
Isaac Triguero, José A. Sáez, Julián Luengo, Salvador García 0001, Francisco Herrera
Neurocomputing1
2014 Fast fingerprint identification for large databases
Daniel Peralta, Isaac Triguero, Raul Sánchez-Reillo, Francisco Herrera, José Manuel Benítez 0001
Pattern Recognit.2
2012 Evolutionary algorithms for the design of grid-connected PV-systems
Daniel Gómez-Lorente, Isaac Triguero, Consolación Gil, A. Espín Estrella
Expert Syst. Appl.2
2012 Integrating a differential evolution feature weighting scheme into prototype generation
Isaac Triguero, Joaquín Derrac, Salvador García 0001, Francisco Herrera
Neurocomputing1
2012 Evolutionary-based selection of generalized instances for imbalanced classification
Salvador García 0001, Joaquín Derrac, Isaac Triguero, Cristóbal J. Carmona, Francisco Herrera
Knowl. Based Syst.3
2012 Time Series Modeling and Forecasting Using Memetic Algorithms for Regime-Switching Models
abstract
In this brief, we present a novel model fitting procedure for the neuro-coefficient smooth transition autoregressive model (NCSTAR), as presented by Medeiros and Veiga. The model is endowed with a statistically founded iterative building procedure and can be interpreted in terms of fuzzy rule-based systems. The interpretability of the generated models and a mathematically sound building procedure are two very important properties of forecasting models. The model fitting procedure employed by the original NCSTAR is a combination of initial parameter estimation by a grid search procedure with a traditional local search algorithm. We propose a different fitting procedure, using a memetic algorithm, in order to obtain more accurate models. An empirical evaluation of the method is performed, applying it to various real-world time series originating from three forecasting competitions. The results indicate that we can significantly enhance the accuracy of the models, making them competitive to models commonly used in the field.
Christoph Bergmeir, Isaac Triguero, Daniel Molina, José Luis Aznarte, José Manuel Benítez 0001
IEEE Trans. Neural Networks Learn. Syst.2
2012 Integrating Instance Selection, Instance Weighting, and Feature Weighting for Nearest Neighbor Classifiers by Coevolutionary Algorithms
abstract
Cooperative coevolution is a successful trend of evolutionary computation which allows us to define partitions of the domain of a given problem, or to integrate several related techniques into one, by the use of evolutionary algorithms. It is possible to apply it to the development of advanced classification methods, which integrate several machine learning techniques into a single proposal. A novel approach integrating instance selection, instance weighting, and feature weighting into the framework of a coevolutionary model is presented in this paper. We compare it with a wide range of evolutionary and nonevolutionary related methods, in order to show the benefits of the employment of coevolution to apply the techniques considered simultaneously. The results obtained, contrasted through nonparametric statistical tests, show that our proposal outperforms other methods in the comparison, thus becoming a suitable tool in the task of enhancing the nearest neighbor classifier.
Joaquín Derrac, Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Part B2
2012 A Taxonomy and Experimental Study on Prototype Generation for Nearest Neighbor Classification
abstract
The nearest neighbor (NN) rule is one of the most successfully used techniques to resolve classification and pattern recognition tasks. Despite its high classification accuracy, this rule suffers from several shortcomings in time response, noise sensitivity, and high storage requirements. These weaknesses have been tackled by many different approaches, including a good and well-known solution that we can find in the literature, which consists of the reduction of the data used for the classification rule (training data). Prototype reduction techniques can be divided into two different approaches, which are known as prototype selection and prototype generation (PG) or abstraction. The former process consists of choosing a subset of the original training data, whereas PG builds new artificial prototypes to increase the accuracy of the NN classification. In this paper, we provide a survey of PG methods specifically designed for the NN rule. From a theoretical point of view, we propose a taxonomy based on the main characteristics presented in them. Furthermore, from an empirical point of view, we conduct a wide experimental study that involves small and large datasets to measure their performance in terms of accuracy and reduction capabilities. The results are contrasted through nonparametrical statistical tests. Several remarks are made to understand which PG models are appropriate for application to different datasets.
Isaac Triguero, Joaquín Derrac, Salvador García 0001, Francisco Herrera
IEEE Trans. Syst. Man Cybern. Part C1
2011 Differential evolution for optimizing the positioning of prototypes in nearest neighbor classification
Isaac Triguero, Salvador García 0001, Francisco Herrera
Pattern Recognit.1
2010 A preliminary study on the use of differential evolution for adjusting the position of examples in nearest neighbor classification
abstract
Nearest neighbor is one of the most successfully used techniques for performing classification and pattern recognition tasks. Its simplicity and effectiveness justify the use of this technique in certain domains but it however presents several drawbacks referring to time response, noise sensitivity and storage requirements. Several solutions have been proposed in order to alleviate these problems, such as improving the technique for speeding up or carrying out a data reduction process. Prototype generation is a suitable process for data reduction that allows to fit a data set for nearest neighbor classification. Position adjustment of prototypes is a successful technique within the prototype generation methodology. Evolutionary algorithms are adaptive methods based on natural evolution that may be used for search and optimization. Position adjustment of prototypes can be viewed as a search problem, thus it could be solved using evolutionary algorithms. In this paper, we perform a preliminary study on the use of differential evolution algorithms to the prototype generation problem. Differential evolution models are compared with other algorithms for adjusting the position of prototypes and the results are contrasted through non-parametrical statistical tests. The results show that some differential evolution models consistently outperform previously proposed methods.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Congress on Evolutionary Computation1
2010 A Preliminary Study on the Selection of Generalized Instances for Imbalanced Classification
Salvador García 0001, Joaquín Derrac, Isaac Triguero, Cristóbal J. Carmona, Francisco Herrera
IEA/AIE (1)3
2010 IPADE: Iterative Prototype Adjustment for Nearest Neighbor Classification
abstract
Nearest prototype methods are a successful trend of many pattern classification tasks. However, they present several shortcomings such as time response, noise sensitivity, and storage requirements. Data reduction techniques are suitable to alleviate these drawbacks. Prototype generation is an appropriate process for data reduction, which allows the fitting of a dataset for nearest neighbor (NN) classification. This brief presents a methodology to learn iteratively the positioning of prototypes using real parameter optimization procedures. Concretely, we propose an iterative prototype adjustment technique based on differential evolution. The results obtained are contrasted with nonparametric statistical tests and show that our proposal consistently outperforms previously proposed methods, thus becoming a suitable tool in the task of enhancing the performance of the NN classifier.
Isaac Triguero, Salvador García 0001, Francisco Herrera
IEEE Trans. Neural Networks1