Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

José A. Sáez

dblp:30/11077 · also José Antonio Sáez · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0002-4592-1538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 77% Data mining · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
data preprocessing
0.212013
A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013
Data integration and cleaning › data preprocessing
discretization
0.212013
A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013
Data mining › predictive modeling
classification
0.012013
A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013
Data mining › predictive modeling
supervised learning
0.012013
A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013

Methods — techniques the papers use, named apart from their topics

nonparametric statistical tests · 0.2
YearPublicationVenuePosition
2025 When Colorimetry Meets the Cloud: IoT-Enabled Health Monitoring at Home via Highly-Specific Gas Sensing
abstract
The integration of Internet of Things (IoT) technologies into healthcare has shown great potential, particularly in remote health monitoring. This paper proposes a novel system combining gas colorimetry with cloud-based IoT schemes to enable precise, non-intrusive home health monitoring. Using highly specific gas sensing through advanced colorimetric gas sensors, this approach detects health-related pollutants and noxious gases in the ambient air, providing continuous real-time data to healthcare providers. The system includes a compact IoT-enabled device capable of simultaneously reading up to four distinct colorimetric gas sensors, complemented by environmental sensing capabilities. This new gasometric device, roughly the size of a credit card, offers a robust, scalable, and privacy-preserving solution for integrating real-time data into cloud-based health analytics. Furthermore, this device is designed to be deployed and integrated within commercial setups along with other advanced sensors, such as cardiorespiratory tracking radars, allowing for complementary analysis of physiology metrics. These features make it a cost-effective and adaptable tool for home diagnostics, bridging the gap between advanced sensing technologies and practical healthcare delivery.
Ismael Benito-Altamirano, Omar Romera-Aller, Miquel Alfaras, Zouhair Haddi, Anaïs Espinoso, Xavier Llauradó, José A. Sáez, Cristian Fàbrega
AINA (3)7
2025 A Survey of Preprocessing Techniques for Flow Cytometry Data in Classification Tasks
abstract
Flow cytometry is an advanced technique for analyzing cellular heterogeneity in biomedical research and clinical diagnostics. Its ability to generate multiparametric data has facilitated advancements in disease classification problems, particularly addressing challenges in distinguishing cell populations and predicting disease outcomes. In these classification tasks, most works consist on either labeling individual cells based on their phenotypic markers or categorizing patient samples as healthy or diseased. However, the complexity of flow cytometry data, characterized by hurdles such as spectral overlap, wide dynamic ranges or batch effects require the usage of preprocessing strategies prior to modeling analysis. This research provides a comprehensive survey of current preprocessing techniques for flow cytometry data used in classification tasks and discusses their specific applications, focusing on four key aspects of data treatment: signal compensation and transformation, batch effect mitigation, imperfect data treatment, and feature selection and class balance. Emphasis is placed on standardizing preprocessing workflows and addressing computational and analytical difficulties posed by the size of modern flow cytometry datasets. The paper also includes a discussion on future opportunities for improving preprocessing pipelines to improve the reproducibility of flow cytometry-based classification models. In short, this work serves as a reference for experts in the field, consolidating best practices in preprocessing and providing guidelines for the development of methodologies that optimize the classification of flow cytometry data.
David Núñez-Nepomuceno, José A. Sáez, Pedro Carmona-Saez
IEEE Trans. Comput. Biol. Bioinform.2
2024 Tackling the problem of noisy IoT sensor data in smart agriculture: Regression noise filters for enhanced evapotranspiration prediction
Juan Martín, José A. Sáez, Emilio Corchado
Expert Syst. Appl.2
2024 Compact Class-Conditional Attribute Category Clustering: Amino Acid Grouping for Enhanced HIV-1 Protease Cleavage Classification
abstract
Categorical attributes are common in many classification tasks, presenting certain challenges as the number of categories grows. This situation can affect data handling, negatively impacting the building time of models, their complexity and, ultimately, their classification performance. In order to mitigate these issues, this research proposes a novel preprocessing technique for grouping attribute categories in classification datasets. This approach combines the exact representation of the association between categorical values in a Euclidean space, clustering methods and attribute quality metrics to group similar attribute categories based on their contribution to the classification task. To estimate its effectiveness, the proposal is evaluated within the context of HIV-1 protease cleavage site prediction, where each attribute represents an amino acid that can take multiple possible values. The results obtained on HIV-1 real-world datasets show a significant reduction in the number of categories per attribute, with an average reduction percentage ranging from 74% to 81%. This reduction leads to simplified data representations and improved classification performances compared to not preprocessing. Specifically, improvements of up to 0.07 in accuracy and 0.19 in geometric mean are observed across different datasets and classification algorithms. Additionally, extensive simulations on synthetic datasets with varied characteristics are carried out, providing consistent and reliable results that validate the robustness of the proposal. These findings highlight the capability of the developed method to enhance cleavage prediction, which could potentially contribute to understanding viral processes and developing targeted therapeutic strategies.
José A. Sáez, José Fernando Vera
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 ANCES: A novel method to repair attribute noise in classification problems
José A. Sáez, Emilio Corchado
Pattern Recognit.1
2016 Tackling label noise with multi-class decomposition using fuzzy one-class support vector machines
abstract
Class label noise is a data-level difficulty associated with training objects with incorrectly assigned labels. This problem may originate from poorly documented historic data, errors during data generation process or mistakes made by human experts. Inclusion of such examples during the training process will mislead the classifier by presenting a falsified class distribution and consequently lead to degradation of models' generalization abilities. This phenomenon becomes even more troublesome in multi-class scenarios that may be affected by highly complex intra-class noise. Decomposition strategies with binary classifiers were proven to alleviate this difficulty by using simplified binary subtasks that are less affected by the noise. In this paper we propose to extend this approach by using the one-class classification decomposition. In this scenario each class has assigned individual one-class classifier that aims at capturing its distinguishing characteristics. This allows to create a robust data description and then apply a dedicated classifier combination in order to reconstruct the original multi-class task. We further extend this concept by using fuzzy one-class classifiers that allow to associate membership values with each training objects. This allows us to reduce the influence of uncertain and potentially noisy samples on the shape of learned decision boundary. Experimental study backed-up with statistical analysis shows that fuzzy one-class classifier decomposition offers an excellent robustness to noise in multi-class classification.
Bartosz Krawczyk, José A. Sáez, Michal Wozniak 0001
FUZZ-IEEE2
2016 Evaluating the classifier behavior with noisy data considering performance and robustness: The Equalized Loss of Accuracy measure
José A. Sáez, Julián Luengo, Francisco Herrera
Neurocomputing1
2016 Analyzing the oversampling of different classes and types of examples in multi-class imbalanced datasets
José A. Sáez, Bartosz Krawczyk, Michal Wozniak 0001
Pattern Recognit.1
2015 SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera
Inf. Sci.1
2015 Using the One-vs-One decomposition to improve the performance of class noise filters via an aggregation strategy in multi-class classification problems
abstract
Noise filters are preprocessing techniques designed to improve data quality in classification tasks by detecting and eliminating examples that contain errors or noise. However, filtering can also remove correct examples and examples containing valuable information, which could be useful for learning. This fact usually implies a margin of improvement on the noise detection accuracy for almost any noise filter. This paper proposes a scheme to improve the performance of noise filters in multi-class classification problems, based on decomposing the dataset into multiple binary subproblems. Decomposition strategies have proven to be successful in improving classification performance in multi-class problems by generating simpler binary subproblems. Similarly, we adapt the principles of the One-vs-One decomposition strategy to noise filtering, making the noise identification process simpler. In order to integrate the filtering results achieved in the binary subproblems, our proposal uses a soft voting approach considering a reliability level based on the aggregation of the noise degree prediction calculated for each binary classifier. The experimental results show that the One-vs-One decomposition strategy usually increases the performance of the noise filters studied, which can detect more accurately the noisy examples.
Luís Paulo F. Garcia, José A. Sáez, Julián Luengo, Ana Carolina Lorena, André C. P. L. F. de Carvalho, Francisco Herrera
Knowl. Based Syst.2
2014 Managing Borderline and Noisy Examples in Imbalanced Classification by Combining SMOTE with Ensemble Filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera
IDEAL1
2014 On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification
Isaac Triguero, José A. Sáez, Julián Luengo, Salvador García 0001, Francisco Herrera
Neurocomputing2
2014 Analyzing the presence of noise in multi-class problems: alleviating its influence with the One-vs-One decomposition
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera
Knowl. Inf. Syst.1
2014 Statistical computation of feature weighting schemes through data estimation for nearest neighbor classifiers
José A. Sáez, Joaquín Derrac, Julián Luengo, Francisco Herrera
Pattern Recognit.1
2013 Tackling the problem of classification with noisy data using Multiple Classifier Systems: Analysis of the performance and robustness
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera
Inf. Sci.1
2013 Predicting noise filtering efficacy with data complexity measures for nearest neighbor classification
José A. Sáez, Julián Luengo, Francisco Herrera
Pattern Recognit.1
2013 A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning
abstract
Discretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones.
Salvador García 0001, Julián Luengo, José A. Sáez, Victoria López, Francisco Herrera
IEEE Trans. Knowl. Data Eng.3
2012 Missing data imputation for fuzzy rule-based classification systems
Julián Luengo, José A. Sáez, Francisco Herrera
Soft Comput.2
2012 Study on the Impact of Partition-Induced Dataset Shift on k -Fold Cross-Validation
abstract
Cross-validation is a very commonly employed technique used to evaluate classifier performance. However, it can potentially introduce dataset shift, a harmful factor that is often not taken into account and can result in inaccurate performance estimation. This paper analyzes the prevalence and impact of partition-induced covariate shift on different k-fold cross-validation schemes. From the experimental results obtained, we conclude that the degree of partition-induced covariate shift depends on the cross-validation scheme considered. In this way, worse schemes may harm the correctness of a single-classifier performance estimation and also increase the needed number of repetitions of cross-validation to reach a stable performance estimation.
Jose G. Moreno-Torres, José A. Sáez, Francisco Herrera
IEEE Trans. Neural Networks Learn. Syst.2
2011 Fuzzy Rule Based Classification Systems versus crisp robust learners trained in presence of class noise's effects: A case of study
abstract
The presence of noise is common in any real-world dataset and may adversely affect the accuracy, construction time and complexity of the classifiers in this context. Traditionally, many algorithms have incorporated mechanisms to deal with noisy problems and reduce noise's effects on performance; they are called robust learners. The C4.5 crisp algorithm is a well-known example of this group of methods. On the other hand, models built by Fuzzy Rule Based Classification Systems are widely recognized for their robustness to imperfect data, but also for their interpretability. The aim of this contribution is to analyze the good behavior and robustness of Fuzzy Rule Based Classification Systems when noise is present in the examples' class labels, especially versus robust learners. In order to accomplish this study, a large number of datasets are created by introducing different levels of noise into the class labels in the training sets. We compare a Fuzzy Rule Based Classification System, the Fuzzy Unordered Rule Induction Algorithm, with respect to the C4.5 classic robust learner which is considered tolerant to noise. From the results obtained it is possible to observe that Fuzzy Rule Based Classification Systems have a good tolerance, in comparison to the C4.5 algorithm, to class noise.
José A. Sáez, Julián Luengo, Francisco Herrera
ISDA1