Julián Luengo

dblp:63/6487 · also Julián Luengo-Martín · DBLP profile ↗
← Back
59ranked-venue papers
14as first author
11since 2021 · last 2026
0000-0003-3952-3629ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSecurity and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 STOOD-X: Explainable out-of-distribution detection via nonparametric statistical testing on large-scale datasets
abstract
Out-of-distribution (OOD) detection is a critical task in machine learning, particularly in safety-sensitive applications where model failures can have serious consequences. However, current OOD detection methods often suffer from restrictive distributional assumptions, limited scalability, and a lack of interpretability. To address these challenges, we propose STOOD-X , a two-stage methodology that combines a Statistical nonparametric Test for OOD Detection with eXplainability enhancements. In the first stage, STOOD-X uses feature-space distances and a Wilcoxon-Mann-Whitney test to identify OOD samples without assuming a specific feature distribution. In the second stage, it generates user-friendly, concept-based visual explanations that reveal the features driving each decision, aligning with the BLUE XAI paradigm. Through extensive experiments on benchmark datasets and multiple architectures, STOOD-X achieves competitive performance compared to state-of-the-art post hoc OOD detectors, particularly in high-dimensional and complex settings. In addition, its explainability framework enables human oversight, bias detection, and model debugging, fostering trust and collaboration between humans and AI systems. Therefore, STOOD-X offers a robust, explainable, and scalable solution for real-world OOD detection tasks.
Iván Sevillano-García, Julián Luengo, Francisco Herrera
Pattern Recognit.2
2025 Developing Big Data anomaly dynamic and static detection algorithms: AnomalyDSD spark package
abstract
Anomaly detection is the process of identifying observations that differ greatly from the majority of data. Unsupervised anomaly detection aims to find outliers in data that is not labeled, therefore, the anomalous instances are unknown. The exponential data generation has led to the era of Big Data. This scenario brings new challenges to classic anomaly detection problems due to the massive and unsupervised accumulation of data. Traditional methods are not able to cop up with computing and time requirements of Big Data problems. In this paper, we propose four distributed algorithm designs for Big Data anomaly detection problems: HBOS_BD, LODA_BD, LSCP_BD, and XGBOD_BD. They have been designed following the MapReduce distributed methodology in order to be capable of handling Big Data problems. These algorithms have been integrated into an Spark Package, focused on static and dynamic Big Data anomaly detection tasks, namely AnomalyDSD. Experiments using a real-world case of study have shown the performance and validity of the proposals for Big Data problems. With this proposal, we have enabled the practitioner to efficiently and effectively detect anomalies in Big Data datasets, where the early detection of an anomaly can lead to a proper and timely decision.
Diego García-Gil, Daniel Argüelles-Martino, Jacinto Carrasco, Ignacio Aguilera-Martos, Julián Luengo, Francisco Herrera
Inf. Sci.6
2024 Local Attention: Enhancing the Transformer Architecture for Efficient Time Series Forecasting
abstract
Transformers have emerged as a highly effective architecture for natural language processing and computer vision. Of late, there has been a surge in initiatives aimed at refining this architecture to enhance its applicability to long sequence time-series forecasting, yielding promising outcomes.This paper introduces Local Attention, an efficient attention mechanism tailored for time series data. This mechanism exploits the continuity properties of time series and the principle of locality in order to compute less attention scores. We provide an Θ(n log n) algorithm to implement Local Attention based on tensor algebra results, which contrasts to the Θ(n2) time and memory complexity of the original attention mechanism.Our experimental analysis shows that the vanilla transformer with Local Attention outperforms state of the art models based on probabilistic attention mechanisms. These findings affirm the effectiveness of our approach and outline a spectrum of future challenges in long sequence time series forecasting.
Ignacio Aguilera-Martos, Andrés Herrera-Poyatos, Julián Luengo, Francisco Herrera
IJCNN3
2023 REVEL Framework to Measure Local Linear Explanations for Black-Box Models: Deep Learning Image Classification Case Study
abstract
Explainable artificial intelligence is proposed to provide explanations for reasoning performed by artificial intelligence. There is no consensus on how to evaluate the quality of these explanations, since even the definition of explanation itself is not clear in the literature. In particular, for the widely known local linear explanations, there are qualitative proposals for the evaluation of explanations, although they suffer from theoretical inconsistencies. The case of image is even more problematic, where a visual explanation seems to explain a decision while detecting edges is what it really does. There are a large number of metrics in the literature specialized in quantitatively measuring different qualitative aspects, so we should be able to develop metrics capable of measuring in a robust and correct way the desirable aspects of the explanations. Some previous papers have attempted to develop new measures for this purpose. However, these measures suffer from lack of objectivity or lack of mathematical consistency, such as saturation or lack of smoothness. In this paper, we propose a procedure called REVEL to evaluate different aspects concerning the quality of explanations with a theoretically coherent development which do not have the problems of the previous measures. This procedure has several advances in the state of the art: it standardizes the concepts of explanation and develops a series of metrics not only to be able to compare between them but also to obtain absolute information regarding the explanation itself. The experiments have been carried out on four image datasets as benchmark where we show REVEL’s descriptive and analytical power.
Iván Sevillano-García, Julián Luengo, Francisco Herrera
Int. J. Intell. Syst.2
2023 TSFEDL: A python library for time series spatio-temporal feature extraction and prediction using deep learning
abstract
The combination of convolutional and recurrent neural networks is a promising framework. This arrangement allows the extraction of high-quality spatio-temporal features together with their temporal dependencies. This fact is key for time series prediction problems such as forecasting, classification or anomaly detection, amongst others. In this paper, the TSFEDL library is introduced. It compiles 22 state-of-the-art methods for both time series feature extraction and prediction, employing convolutional and recurrent deep neural networks for its use in several data mining tasks. The library is built upon a set of Tensorflow + Keras and PyTorch modules under the AGPLv3 license. The performance validation of the architectures included in this proposal confirms the usefulness of this Python package.
Ignacio Aguilera-Martos, Ángel Miguel García-Vico, Julián Luengo, Sergio Damas, Francisco J. Melero 0001, José Javier Valle-Alonso, Francisco Herrera
Neurocomputing3
2023 Multi-step histogram based outlier scores for unsupervised anomaly detection: ArcelorMittal engineering dataset case of study
abstract
Anomaly detection is the task of detecting samples that behave differently from the rest of the data or that include abnormal values. Unsupervised anomaly detection is the most common scenario, which implies that the algorithms cannot train with a labeled input and do not know the anomaly behavior beforehand. Histogram-based methods are one of the most approaches in unsupervised anomaly detection, remarking a good performance and a low runtime. Despite the good performance, histogram-based anomaly detectors are not capable of processing data flows while updating their knowledge and cannot deal with a high amount of samples. In this paper, we propose a new histogram-based approach for addressing the aforementioned problems by introducing the ability to update the information inside a histogram. We have applied these strategies to design a new algorithm called Multi-step Histogram Based Outlier Scores (MHBOS), including five new histogram update mechanisms. The results have shown the performance and validity of MHBOS as well as the proposed strategies in terms of performance and computing times.
Ignacio Aguilera-Martos, Marta García-Bárzana, Diego García-Gil, Jacinto Carrasco, Julián Luengo, Francisco Herrera
Neurocomputing6
2022 The impact of heterogeneous distance functions on missing data imputation and classification performance
Miriam Seoane Santos, Pedro H. Abreu, Alberto Fernández 0001, Julián Luengo, João A. M. Santos
Eng. Appl. Artif. Intell.4
2022 3SHACC: Three stages hybrid agglomerative constrained clustering
Germán González-Almagro, Juan-Luis Suárez, Julián Luengo, José Ramón Cano, Salvador García 0001
Neurocomputing3
2021 Anomaly detection in predictive maintenance: A new evaluation framework for temporal unsupervised anomaly detection algorithms
Jacinto Carrasco, Ignacio Aguilera-Martos, Diego García-Gil, Irina Markova, Marta García-Bárzana, Manuel Arias-Rodil, Julián Luengo, Francisco Herrera
Neurocomputing8
2021 Synthetic Sample Generation for Label Distribution Learning
Julián Luengo, José Ramón Cano, Salvador García 0001
Inf. Sci.2
2021 Multiple instance classification: Bag noise filtering for negative instance noise cleaning
abstract
Data in the real world is far from being perfect. The appearance of noise is a common issue that arises from the limitations of data acquisition mechanisms and human knowledge. In classification, label noise will hinder the performance of almost all classifiers, inducing a bias in the built model. While label noise has recently attracted researchers’ attention in standard classification, it has only recently begun to be studied in multiple instance classification. In this work, we propose the usage of filtering algorithms for multiple instance classification that are able to reduce the impact of negative instances within the bags. In order to do so, we decompose the bags to form a standard classification problem that can be efficiently treated by a specialized noise filter. Such a decomposition is tackled in different ways, with the aim of exploiting the knowledge offered by the examples from opposite bags. The bags are then rebuilt, without the identified noise instances. In our experiments, we show that by applying our approach we can diminish the impact of noise and even obtain better results at 0% noise level for several classifiers. Our approach sets out a promising approach to dealing with noise in the bags of multiple instance datasets and further improve the classification rate of the built models.
Julián Luengo, Dánel Sánchez Tarragó, Ronaldo C. Prati, Francisco Herrera
Inf. Sci.1
2020 Improving constrained clustering via decomposition-based multiobjective optimization with memetic elitism
abstract
Clustering has always been a topic of interest in knowledge discovery, it is able to provide us with valuable information within the unsupervised machine learning framework. It received renewed attention when it was shown to produce better results in environments where partial information about how to solve the problem is available, thus leading to a new machine learning paradigm: semi-supervised machine learning. This new type of information can be given in the form of constraints, which guide the clustering process towards quality solutions. In particular, this study considers the pairwise instance-level must-link and cannot-link constraints. Given the ill-posed nature of the constrained clustering problem, we approach it from the multiobjective optimization point of view. Our proposal consists in a memetic elitist evolutionary strategy that favors exploitation by applying a local search procedure to the elite of the population and transferring its results only to the external population, which will also be used to generate new individuals. We show the capability of this method to produce quality results for the constrained clustering problem when considering incremental levels of constraint-based information. For the comparison with state-of-the-art methods, we include previous multiobjective approaches, single-objective genetic algorithms and classic constrained clustering methods.
Germán González-Almagro, Alejandro Rosales-Pérez, Julián Luengo, José Ramón Cano, Salvador García 0001
GECCO3
2020 Preprocessing methodology for time series: An industrial world application case study
Juan Antonio Cortés-Ibáñez, Sergio González, José Javier Valle-Alonso, Julián Luengo, Salvador García 0001, Francisco Herrera
Inf. Sci.4
2020 Fast and Scalable Approaches to Accelerate the Fuzzy k-Nearest Neighbors Classifier for Big Data
abstract
One of the best-known and most effective methods in supervised classification is the k-nearest neighbors algorithm (kNN). Several approaches have been proposed to improve its accuracy, where fuzzy approaches prove to be among the most successful, highlighting the classical fuzzy k-nearest neighbors (FkNN). However, these traditional algorithms fail to tackle the large amounts of data that are available today. There are multiple alternatives to enable kNN classification in big datasets, spotlighting the approximate version of kNN called hybrid spill tree. Nevertheless, the existing proposals of FkNN for big data problems are not fully scalable, because a high computational load is required to obtain the same behavior as the original FkNN algorithm. This article proposes global approximate hybrid spill tree FkNN and local hybrid spill tree FkNN, two approximate approaches that speed up runtime without losing quality in the classification process. The experimentation compares various FkNN approaches for big data with datasets of up to 11 million instances. The results show an improvement in runtime and accuracy over literature algorithms.
Jesús Maillo, Salvador García 0001, Julián Luengo, Francisco Herrera, Isaac Triguero
IEEE Trans. Fuzzy Syst.3
2020 COVIDGR Dataset and COVID-SDNet Methodology for Predicting COVID-19 Based on Chest X-Ray Images
abstract
Currently, Coronavirus disease (COVID-19), one of the most infectious diseases in the 21st century, is diagnosed using RT-PCR testing, CT scans and/or Chest X-Ray (CXR) images. CT (Computed Tomography) scanners and RT-PCR testing are not available in most medical centers and hence in many cases CXR images become the most time/cost effective tool for assisting clinicians in making decisions. Deep learning neural networks have a great potential for building COVID-19 triage systems and detecting COVID-19 patients, especially patients with low severity. Unfortunately, current databases do not allow building such systems as they are highly heterogeneous and biased towards severe cases. This article is three-fold: (i) we demystify the high sensitivities achieved by most recent COVID-19 classification models, (ii) under a close collaboration with Hospital Universitario Clínico San Cecilio, Granada, Spain, we built COVIDGR-1.0, a homogeneous and balanced database that includes all levels of severity, from normal with Positive RT-PCR, Mild, Moderate to Severe. COVIDGR-1.0 contains 426 positive and 426 negative PA (PosteroAnterior) CXR views and (iii) we propose COVID Smart Data based Network (COVID-SDNet) methodology for improving the generalization capacity of COVID-classification models. Our approach reaches good and stable results with an accuracy of [Formula: see text], [Formula: see text], [Formula: see text] in severe, moderate and mild COVID-19 severity levels. Our approach could help in the early detection of COVID-19. COVIDGR-1.0 along with the severity level labels are available to the scientific community through this link https://dasci.es/es/transferencia/open-data/covidgr/.
Siham Tabik, Anabel Gómez-Ríos, José Luis Martín-Rodríguez, Iván Sevillano-García, Manuel Rey-Area, David Charte, Emilio Guirado, Juan-Luis Suárez, Julián Luengo, M. A. Valero-González, P. García-Villanova, Eulalia Olmedo-Sánchez, Francisco Herrera
IEEE J. Biomed. Health Informatics9
2019 Big Data Preprocessing as the Bridge between Big Data and Smart Data: BigDaPSpark and BigDaPFlink Libraries
abstract
With the advent of Big Data, terabytes of data are generated and stored every second. This raw data is far from \nbeing perfect, it contains many imperfections (noise, missing values, etc.) and is not suitable for analysis, \nas it will led to wrong conclusions. Data preprocessing is the set of techniques devoted to polish, clean, \nfix, and improve that raw data. With this preprocessed data, we would be able to find more patterns in it, \nand to better explain the underlaying distribution of the data. This is what is called Smart Data, raw data \nthat has been preprocessed and is ready for being analyzed, data that contains valuable information that will \nled to knowledge. In this work, we present two Big Data libraries for achieving Smart Data from Big Data, \nBigDaPSpark and BigDaPFlink. They are built on top of two Big Data frameworks, Apache Spark and Apache \nFlink. Both libraries contain a series of algorithms for Big Data preprocessing, ranging from noise cleaning, \nto discretization, or data reduction, among many others. Additionally, we ilustrate the usage of the libraries \nwith two cases of use.
Diego García-Gil, Alejandro Alcalde-Barros, Julián Luengo, Salvador García 0001, Francisco Herrera
IoTBDS3
2019 A First Approach on Big Data Missing Values Imputation
abstract
Albeit most techniques and algorithms assume that the data is accurate, measurements in our analogic world are far from being perfect. Since our capabilities of storing and processing data are growing everyday, these imperfections will accumulate, generating poorer decisions and hindering any knowledge extraction process carried out over the raw data. One of the most disturbing imperfections is the presence of missing values. Many inductive algorithms assume that the data is complete, thus if they face missing data they will not work properly or the quality of the knowledge extracted will be poorer. At this point there is no sophisticated missing values treatment implemented in any major Big Data framework. In this contribution, we present two novel imputation methods based on clustering that achieve better results than simply removing the faulty examples or filling-in the missing values with the mean that can be easily ported to Spark’s MLlib.
Besay Montesdeoca, Julián Luengo, Jesús Maillo, Diego García-Gil, Salvador García 0001, Francisco Herrera
IoTBDS2
2019 Towards highly accurate coral texture images classification using deep convolutional neural networks and data augmentation
Anabel Gómez-Ríos, Siham Tabik, Julián Luengo, A. S. M. Shihavuddin, Bartosz Krawczyk, Francisco Herrera
Expert Syst. Appl.3
2019 From Big to Smart Data: Iterative ensemble filter for noise filtering in Big Data classification
abstract
The quality of the data is directly related to the quality of the models drawn from that data. For that reason, many research is devoted to improve the quality of the data and to amend errors that it may contain. One of the most common problems is the presence of noise in classification tasks, where noise refers to the incorrect labeling of training instances. This problem is very disruptive, as it changes the decision boundaries of the problem. Big Data problems pose a new challenge in terms of quality data due to the massive and unsupervised accumulation of data. This Big Data scenario also brings new problems to classic data preprocessing algorithms, as they are not prepared for working with such amounts of data, and these algorithms are key to move from Big to Smart Data. In this paper, an iterative ensemble filter for removing noisy instances in Big Data scenarios is proposed. Experiments carried out in six Big Data datasets have shown that our noise filter outperforms the current state-of-the-art noise filter in Big Data domains. It has also proved to be an effective solution for transforming raw Big Data into Smart Data.
Diego García-Gil, Francisco Luque Sánchez, Julián Luengo, Salvador García 0001, Francisco Herrera
Int. J. Intell. Syst.3
2019 Label noise filtering techniques to improve monotonic classification
José Ramón Cano, Julián Luengo, Salvador García 0001
Neurocomputing2
2019 Smartdata: Data preprocessing to achieve smart data in R
Ignacio Cordón, Julián Luengo, Salvador García 0001, Francisco Herrera, Francisco Charte
Neurocomputing2
2019 Enabling Smart Data: Noise filtering in Big Data classification
Diego García-Gil, Julián Luengo, Salvador García 0001, Francisco Herrera
Inf. Sci.2
2019 Emerging topics and challenges of learning from noisy data in nonstandard classification: a survey beyond binary class noise
Ronaldo C. Prati, Julián Luengo, Francisco Herrera
Knowl. Inf. Syst.2
2019 Coral species identification with texture or structure images using a two-level classifier based on Convolutional Neural Networks
abstract
Corals are crucial animals as they support a large part of marine life. The automatic classification of corals species based on underwater images is important as it can help experts to track and detect threatened and vulnerable coral species. However, this classification is complicated due to the nature of coral underwater images and the fact that current underwater coral datasets are unrealistic as they contain only texture images, while the images taken by autonomous underwater vehicles show the complete coral structure. The objective of this paper is two-fold. The first is to build a dataset that is representative of the problem of classifying underwater coral images, the StructureRSMAS dataset. The second is to build a classifier capable of resolving the real problem of classifying corals, based either on texture or structure images. We have achieved this by using a two-level classifier composed of three ResNet models. The first level recognizes whether the input image is a texture or a structure image. Then, the second level identifies the coral species. To do this, we have used a known texture dataset, RSMAS , and StructureRSMAS.
Anabel Gómez-Ríos, Siham Tabik, Julián Luengo, A. S. M. Shihavuddin, Francisco Herrera
Knowl. Based Syst.3
2018 A preliminary study on Hybrid Spill-Tree Fuzzy k-Nearest Neighbors for big data classification
abstract
The Fuzzy k Nearest Neighbor (Fuzzy kNN) classifier is well known for its effectiveness in supervised learning problems. kNN classifies by comparing new incoming examples with a similarity function using the samples of the training set. The fuzzy version of the kNN accounts for the underlying uncertainty in the class labels, and it is composed of two different stages. The first one is responsible for calculating the fuzzy membership degree for each sample of the problem in order to obtain smoother boundaries between classes. The second stage classifies similarly to the standard kNN algorithm but uses the previously calculated class membership degree. To deal with very large datasets, distributed versions of the Fuzzy kNN algorithm have been proposed. However, existing approaches remain not fully scalable as they aim to replicate the exact behavior of the Fuzzy kNN. In this work, we present an approximate and distributed Fuzzy kNN approach based on Hybrid Spill-Tree implemented under Apache Spark. The aim of this model is to alleviate the scalability problems and to deal with big datasets maintaining high accuracy. In our experiments, we compare in precision and runtime with the Fuzzy kNN for big data problems existing in the literature, running with datasets of up to 11 million instances. The results show an improvement in the runtime and accuracy with respect to the previous exact model.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE2
2018 CNC-NOS: Class noise cleaning by ensemble filtering and noise scoring
Julián Luengo, Seong-O Shim, Saleh Alshomrani, Abdulrahman H. Altalhi, Francisco Herrera
Knowl. Based Syst.1
2017 Exact fuzzy k-nearest neighbor classification for big datasets
abstract
The k-Nearest Neighbors (kNN) classifier is one of the most effective methods in supervised learning problems. It classifies unseen cases comparing their similarity with the training data. Nevertheless, it gives to each labeled sample the same importance to classify. There are several approaches to enhance its precision, with the Fuzzy k-Nearest Neighbors (Fuzzy-kNN) classifier being among the most successful ones. Fuzzy-kNN computes a fuzzy degree of membership of each instance to the classes of the problem. As a result, it generates smoother borders between classes. Apart from the existing kNN approach to handle big datasets, there is not a fuzzy variant to manage that volume of data. Nevertheless, calculating this class membership adds an extra computational cost becoming even less scalable to tackle large datasets because of memory needs and high runtime. In this work, we present an exact and distributed approach to run the Fuzzy-kNN classifier on big datasets based on Spark, which provides the same precision than the original algorithm. It presents two separately stages. The first stage transforms the training set adding the class membership degrees. The second stage classifies with the kNN algorithm the test set using the class membership computed previously. In our experiments, we study the scaling-up capabilities of the proposed approach with datasets up to 11 million instances, showing promising results.
Jesús Maillo, Julián Luengo, Salvador García 0001, Francisco Herrera, Isaac Triguero
FUZZ-IEEE2
2016 Evaluating the classifier behavior with noisy data considering performance and robustness: The Equalized Loss of Accuracy measure
José A. Sáez, Julián Luengo, Francisco Herrera
Neurocomputing2
2016 Tutorial on practical tips of the most influential data preprocessing algorithms in data mining
Salvador García 0001, Julián Luengo, Francisco Herrera
Knowl. Based Syst.2
2016 The influence of noise on the evolutionary fuzzy systems for subgroup discovery
Julián Luengo, Ángel Miguel García-Vico, M. Dolores Pérez-Godoy, Cristóbal J. Carmona
Soft Comput.1
2015 Naive Bayes Classifier with Mixtures of Polynomials
Julián Luengo, Rafael Rumí
ICPRAM (1)1
2015 SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera
Inf. Sci.2
2015 An automatic extraction method of the domains of competence for learning classifiers using data complexity measures
Julián Luengo, Francisco Herrera
Knowl. Inf. Syst.1
2015 Using the One-vs-One decomposition to improve the performance of class noise filters via an aggregation strategy in multi-class classification problems
abstract
Noise filters are preprocessing techniques designed to improve data quality in classification tasks by detecting and eliminating examples that contain errors or noise. However, filtering can also remove correct examples and examples containing valuable information, which could be useful for learning. This fact usually implies a margin of improvement on the noise detection accuracy for almost any noise filter. This paper proposes a scheme to improve the performance of noise filters in multi-class classification problems, based on decomposing the dataset into multiple binary subproblems. Decomposition strategies have proven to be successful in improving classification performance in multi-class problems by generating simpler binary subproblems. Similarly, we adapt the principles of the One-vs-One decomposition strategy to noise filtering, making the noise identification process simpler. In order to integrate the filtering results achieved in the binary subproblems, our proposal uses a soft voting approach considering a reliability level based on the aggregation of the noise degree prediction calculated for each binary classifier. The experimental results show that the One-vs-One decomposition strategy usually increases the performance of the noise filters studied, which can detect more accurately the noisy examples.
Luís Paulo F. Garcia, José A. Sáez, Julián Luengo, Ana Carolina Lorena, André C. P. L. F. de Carvalho, Francisco Herrera
Knowl. Based Syst.3
2014 Managing Borderline and Noisy Examples in Imbalanced Classification by Combining SMOTE with Ensemble Filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera
IDEAL2
2014 On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification
Isaac Triguero, José A. Sáez, Julián Luengo, Salvador García 0001, Francisco Herrera
Neurocomputing3
2014 Analyzing the presence of noise in multi-class problems: alleviating its influence with the One-vs-One decomposition
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera
Knowl. Inf. Syst.3
2014 Statistical computation of feature weighting schemes through data estimation for nearest neighbor classifiers
José A. Sáez, Joaquín Derrac, Julián Luengo, Francisco Herrera
Pattern Recognit.3
2013 Tackling the problem of classification with noisy data using Multiple Classifier Systems: Analysis of the performance and robustness
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera
Inf. Sci.3
2013 Predicting noise filtering efficacy with data complexity measures for nearest neighbor classification
José A. Sáez, Julián Luengo, Francisco Herrera
Pattern Recognit.2
2013 A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning
abstract
Discretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones.
Salvador García 0001, Julián Luengo, José A. Sáez, Victoria López, Francisco Herrera
IEEE Trans. Knowl. Data Eng.2
2012 A preliminary study on missing data imputation in evolutionary fuzzy systems of subgroup discovery
abstract
In real-life data, a loss of information is frequent in data mining due to the presence of missing values in the attributes. Missing values can occur due to problems in the manual data entry procedures, equipment errors or incorrect measurements. The presence of missing values in attributes conditions the results obtained by any knowledge extraction approach. Specifically, this problem could lead in subgroup discovery to a loss of quality of results obtained by subgroups on measures such as sensitivity, confidence, significance or unusualness. This paper presents an experimental study to analyse the effect of different missing data imputation mechanisms within subgroup discovery algorithms based on evolutionary fuzzy systems presented throughout the literature. The analysis is carried out with a large number of data sets obtained from KEEL repository. Among all the imputation techniques, the imputation method K-Nearest Neighbour outstands as the best option. In summary, if experts need to analyse a problem with a high percentage of missing values they must use this imputation method in order to treat data in a correct way and also to obtain a meaningful descriptive knowledge. In addition, results also show that the evolutionary fuzzy system with the best results is the algorithm NMEEF-SD in the missing values scenario.
Cristóbal J. Carmona, Julián Luengo, Pedro González 0001, María José del Jesus
FUZZ-IEEE2
2012 A Preliminary Study on Selecting the Optimal Cut Points in Discretization by Evolutionary Algorithms
Salvador García 0001, Victoria López, Julián Luengo, Cristóbal J. Carmona, Francisco Herrera
ICPRAM (1)3
2012 An analysis on the use of pre-processing methods in evolutionary fuzzy systems for subgroup discovery
Cristóbal J. Carmona, Julián Luengo, Pedro González 0001, María José del Jesus
Expert Syst. Appl.2
2012 Shared domains of competence of approximate learning models using measures of separability of classes
Julián Luengo, Francisco Herrera
Inf. Sci.1
2012 On the choice of the best imputation methods for missing values considering three groups of classification methods
Julián Luengo, Salvador García 0001, Francisco Herrera
Knowl. Inf. Syst.1
2012 Missing data imputation for fuzzy rule-based classification systems
Julián Luengo, José A. Sáez, Francisco Herrera
Soft Comput.1
2011 Fuzzy Rule Based Classification Systems versus crisp robust learners trained in presence of class noise's effects: A case of study
abstract
The presence of noise is common in any real-world dataset and may adversely affect the accuracy, construction time and complexity of the classifiers in this context. Traditionally, many algorithms have incorporated mechanisms to deal with noisy problems and reduce noise's effects on performance; they are called robust learners. The C4.5 crisp algorithm is a well-known example of this group of methods. On the other hand, models built by Fuzzy Rule Based Classification Systems are widely recognized for their robustness to imperfect data, but also for their interpretability. The aim of this contribution is to analyze the good behavior and robustness of Fuzzy Rule Based Classification Systems when noise is present in the examples' class labels, especially versus robust learners. In order to accomplish this study, a large number of datasets are created by introducing different levels of noise into the class labels in the training sets. We compare a Fuzzy Rule Based Classification System, the Fuzzy Unordered Rule Induction Algorithm, with respect to the C4.5 classic robust learner which is considered tolerant to noise. From the results obtained it is possible to observe that Fuzzy Rule Based Classification Systems have a good tolerance, in comparison to the C4.5 algorithm, to class noise.
José A. Sáez, Julián Luengo, Francisco Herrera
ISDA2
2011 Addressing data complexity for imbalanced data sets: analysis of SMOTE-based oversampling and evolutionary undersampling
Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera
Soft Comput.1
2010 An extraction method for the characterization of the Fuzzy Rule Based Classification Systems' behavior using data complexity measures: A case of study with FH-GBML
abstract
When dealing with problems using Fuzzy Rule Based Classification Systems it is difficult to know in advance whether the model will perform well or badly. In this work we present an automatic extraction method to determine the domains of competence of Fuzzy Rule Based Classification Systems As a case of study we use the Fuzzy Hybrid Genetic Based Machine Learning method. We consider twelve metrics of data complexity in order to analyze the behavior patterns of this method, obtaining intervals of such data complexity measures with good or bad performance of it. Combining these intervals we obtain rules that describe both good or bad behaviors of the Fuzzy Rule Based Classification System mentioned. These rules allow describe both good or bad behaviors of the Fuzzy Rule Based Classification Systems mentioned, allowing us to characterize the response quality of the methods from the data set complexity metrics of a given data set. Thus, we can establish the domains of competence of the Fuzzy Rule Based Classification Systems considered, making it possible to establish when the method will perform well or badly prior to its application.
Julián Luengo, Francisco Herrera
FUZZ-IEEE1
2010 Domains of competence of fuzzy rule based classification systems with data complexity measures: A case of study using a fuzzy hybrid genetic based machine learning method
Julián Luengo, Francisco Herrera
Fuzzy Sets Syst.1
2010 Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera
Inf. Sci.3
2010 A study on the use of imputation methods for experimentation with Radial Basis Function Network classifiers handling missing attribute values: The good synergy between RBFNs and EventCovering method
Julián Luengo, Salvador García 0001, Francisco Herrera
Neural Networks1
2010 Genetics-Based Machine Learning for Rule Induction: State of the Art, Taxonomy, and Comparative Study
abstract
The classification problem can be addressed by numerous techniques and algorithms which belong to different paradigms of machine learning. In this paper, we are interested in evolutionary algorithms, the so-called genetics-based machine learning algorithms. In particular, we will focus on evolutionary approaches that evolve a set of rules, i.e., evolutionary rule-based systems, applied to classification tasks, in order to provide a state of the art in this field. This paper has a double aim: to present a taxonomy of the genetics-based machine learning approaches for rule induction, and to develop an empirical analysis both for standard classification and for classification with imbalanced data sets. We also include a comparative study of the genetics-based machine learning (GBML) methods with some classical non-evolutionary algorithms, in order to observe the suitability and high potential of the search performed by evolutionary algorithms and the behavior of the GBML algorithms in contrast to the classical approaches, in terms of classification accuracy.
Alberto Fernández 0001, Salvador García 0001, Julián Luengo, Ester Bernadó-Mansilla, Francisco Herrera
IEEE Trans. Evol. Comput.3
2009 Implementation and Integration of Algorithms into the KEEL Data-Mining Software Tool
Alberto Fernández 0001, Julián Luengo, Joaquín Derrac, Jesús Alcalá-Fdez, Francisco Herrera
IDEAL2
2009 A First Approach to Nearest Hyperrectangle Selection by Evolutionary Algorithms
abstract
The nested generalized exemplar theory accomplishes learning by storing objects in Euclidean n-space, as hyperrectangles. Classification of new data is performed by computing their distance to the nearest “generalized exemplar” or hyperrectangle. This learning method permits to combine the distance-based classification with the axis-parallel rectangle representation employed in most of the rule-learning systems. This contribution proposes the use of evolutionary algorithms to select the most influential hyperrectangles to obtain accurate and simple models in classification tasks. The proposal is compared with the most representative nearest hyperrectangle learning approaches and the results obtained show that the evolutionary proposal outperforms them in accuracy and requires storing a lower number of hyperrectangles.
Salvador García 0001, Joaquín Derrac, Julián Luengo, Francisco Herrera
ISDA3
2009 Addressing Data-Complexity for Imbalanced Data-Sets: A Preliminary Study on the Use of Preprocessing for C4.5
abstract
In this work we analyse the behaviour of the C4.5 classification method with respect to a bunch of imbalanced data-sets. We consider the use of two metrics of data complexity known as “maximum Fishers discriminant ratio” and “nonlinearity of 1NN classifier”, to analyse the effect of preprocessing (oversampling in this case) in order to deal with the imbalance problem. In order to do that, we analyse C4.5 over a wide range of imbalanced data-sets built from real data, and try to extract behaviour patterns from the results. We obtain rules that describe both good or bad behaviours of C4.5 in the case of using the original data-sets (absence of preprocessing) and when applying preprocessing. These rules allow us to determine the effect of the use of preprocessing and to predict the response of C4.5 to preprocessing from the data-set’s complexity metrics prior to its application, and then establish when the preprocessing would be useful to.
Julián Luengo, Alberto Fernández 0001, Salvador García 0001, Francisco Herrera
ISDA1
2009 A study on the use of statistical tests for experimentation with neural networks: Analysis of parametric test conditions and non-parametric tests
Julián Luengo, Salvador García 0001, Francisco Herrera
Expert Syst. Appl.1
2009 A study of statistical techniques and performance measures for genetics-based machine learning: accuracy and interpretability
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera
Soft Comput.3