EDBT 2026 Demo / reviewers in the wild / expert
José A. Sáez
dblp:30/11077 · also José Antonio Sáez
· DBLP profile ↗
20ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0002-4592-1538ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 77% Data mining · 23% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
data preprocessing |
0.2 | 1 | 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013 |
Data integration and cleaning › data preprocessing
discretization |
0.2 | 1 | 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining › predictive modeling
classification |
0.0 | 1 | 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013 |
Data mining › predictive modeling
supervised learning |
0.0 | 1 | 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised Learning · IEEE Trans. Knowl. Data Eng. 2013 |
Methods — techniques the papers use, named apart from their topics
nonparametric statistical tests · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | When Colorimetry Meets the Cloud: IoT-Enabled Health Monitoring at Home via Highly-Specific Gas SensingabstractThe integration of Internet of Things (IoT) technologies into healthcare has shown great potential, particularly in remote health monitoring. This paper proposes a novel system combining gas colorimetry with cloud-based IoT schemes to enable precise, non-intrusive home health monitoring. Using highly specific gas sensing through advanced colorimetric gas sensors, this approach detects health-related pollutants and noxious gases in the ambient air, providing continuous real-time data to healthcare providers. The system includes a compact IoT-enabled device capable of simultaneously reading up to four distinct colorimetric gas sensors, complemented by environmental sensing capabilities. This new gasometric device, roughly the size of a credit card, offers a robust, scalable, and privacy-preserving solution for integrating real-time data into cloud-based health analytics. Furthermore, this device is designed to be deployed and integrated within commercial setups along with other advanced sensors, such as cardiorespiratory tracking radars, allowing for complementary analysis of physiology metrics. These features make it a cost-effective and adaptable tool for home diagnostics, bridging the gap between advanced sensing technologies and practical healthcare delivery. Ismael Benito-Altamirano, Omar Romera-Aller, Miquel Alfaras, Zouhair Haddi, Anaïs Espinoso, Xavier Llauradó, José A. Sáez, Cristian Fàbrega |
AINA (3) | 7 |
| 2025 | A Survey of Preprocessing Techniques for Flow Cytometry Data in Classification TasksabstractFlow cytometry is an advanced technique for analyzing cellular heterogeneity in biomedical research and clinical diagnostics. Its ability to generate multiparametric data has facilitated advancements in disease classification problems, particularly addressing challenges in distinguishing cell populations and predicting disease outcomes. In these classification tasks, most works consist on either labeling individual cells based on their phenotypic markers or categorizing patient samples as healthy or diseased. However, the complexity of flow cytometry data, characterized by hurdles such as spectral overlap, wide dynamic ranges or batch effects require the usage of preprocessing strategies prior to modeling analysis. This research provides a comprehensive survey of current preprocessing techniques for flow cytometry data used in classification tasks and discusses their specific applications, focusing on four key aspects of data treatment: signal compensation and transformation, batch effect mitigation, imperfect data treatment, and feature selection and class balance. Emphasis is placed on standardizing preprocessing workflows and addressing computational and analytical difficulties posed by the size of modern flow cytometry datasets. The paper also includes a discussion on future opportunities for improving preprocessing pipelines to improve the reproducibility of flow cytometry-based classification models. In short, this work serves as a reference for experts in the field, consolidating best practices in preprocessing and providing guidelines for the development of methodologies that optimize the classification of flow cytometry data. David Núñez-Nepomuceno, José A. Sáez, Pedro Carmona-Saez |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2024 | Tackling the problem of noisy IoT sensor data in smart agriculture: Regression noise filters for enhanced evapotranspiration prediction
Juan Martín, José A. Sáez, Emilio Corchado |
Expert Syst. Appl. | 2 |
| 2024 | Compact Class-Conditional Attribute Category Clustering: Amino Acid Grouping for Enhanced HIV-1 Protease Cleavage ClassificationabstractCategorical attributes are common in many classification tasks, presenting certain challenges as the number of categories grows. This situation can affect data handling, negatively impacting the building time of models, their complexity and, ultimately, their classification performance. In order to mitigate these issues, this research proposes a novel preprocessing technique for grouping attribute categories in classification datasets. This approach combines the exact representation of the association between categorical values in a Euclidean space, clustering methods and attribute quality metrics to group similar attribute categories based on their contribution to the classification task. To estimate its effectiveness, the proposal is evaluated within the context of HIV-1 protease cleavage site prediction, where each attribute represents an amino acid that can take multiple possible values. The results obtained on HIV-1 real-world datasets show a significant reduction in the number of categories per attribute, with an average reduction percentage ranging from 74% to 81%. This reduction leads to simplified data representations and improved classification performances compared to not preprocessing. Specifically, improvements of up to 0.07 in accuracy and 0.19 in geometric mean are observed across different datasets and classification algorithms. Additionally, extensive simulations on synthetic datasets with varied characteristics are carried out, providing consistent and reliable results that validate the robustness of the proposal. These findings highlight the capability of the developed method to enhance cleavage prediction, which could potentially contribute to understanding viral processes and developing targeted therapeutic strategies. José A. Sáez, José Fernando Vera |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | ANCES: A novel method to repair attribute noise in classification problems
José A. Sáez, Emilio Corchado |
Pattern Recognit. | 1 |
| 2016 | Tackling label noise with multi-class decomposition using fuzzy one-class support vector machinesabstractClass label noise is a data-level difficulty associated with training objects with incorrectly assigned labels. This problem may originate from poorly documented historic data, errors during data generation process or mistakes made by human experts. Inclusion of such examples during the training process will mislead the classifier by presenting a falsified class distribution and consequently lead to degradation of models' generalization abilities. This phenomenon becomes even more troublesome in multi-class scenarios that may be affected by highly complex intra-class noise. Decomposition strategies with binary classifiers were proven to alleviate this difficulty by using simplified binary subtasks that are less affected by the noise. In this paper we propose to extend this approach by using the one-class classification decomposition. In this scenario each class has assigned individual one-class classifier that aims at capturing its distinguishing characteristics. This allows to create a robust data description and then apply a dedicated classifier combination in order to reconstruct the original multi-class task. We further extend this concept by using fuzzy one-class classifiers that allow to associate membership values with each training objects. This allows us to reduce the influence of uncertain and potentially noisy samples on the shape of learned decision boundary. Experimental study backed-up with statistical analysis shows that fuzzy one-class classifier decomposition offers an excellent robustness to noise in multi-class classification. Bartosz Krawczyk, José A. Sáez, Michal Wozniak 0001 |
FUZZ-IEEE | 2 |
| 2016 | Evaluating the classifier behavior with noisy data considering performance and robustness: The Equalized Loss of Accuracy measure
José A. Sáez, Julián Luengo, Francisco Herrera |
Neurocomputing | 1 |
| 2016 | Analyzing the oversampling of different classes and types of examples in multi-class imbalanced datasets
José A. Sáez, Bartosz Krawczyk, Michal Wozniak 0001 |
Pattern Recognit. | 1 |
| 2015 | SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera |
Inf. Sci. | 1 |
| 2015 | Using the One-vs-One decomposition to improve the performance of class noise filters via an aggregation strategy in multi-class classification problemsabstractNoise filters are preprocessing techniques designed to improve data quality in classification tasks by detecting and eliminating examples that contain errors or noise. However, filtering can also remove correct examples and examples containing valuable information, which could be useful for learning. This fact usually implies a margin of improvement on the noise detection accuracy for almost any noise filter. This paper proposes a scheme to improve the performance of noise filters in multi-class classification problems, based on decomposing the dataset into multiple binary subproblems. Decomposition strategies have proven to be successful in improving classification performance in multi-class problems by generating simpler binary subproblems. Similarly, we adapt the principles of the One-vs-One decomposition strategy to noise filtering, making the noise identification process simpler. In order to integrate the filtering results achieved in the binary subproblems, our proposal uses a soft voting approach considering a reliability level based on the aggregation of the noise degree prediction calculated for each binary classifier. The experimental results show that the One-vs-One decomposition strategy usually increases the performance of the noise filters studied, which can detect more accurately the noisy examples. Luís Paulo F. Garcia, José A. Sáez, Julián Luengo, Ana Carolina Lorena, André C. P. L. F. de Carvalho, Francisco Herrera |
Knowl. Based Syst. | 2 |
| 2014 | Managing Borderline and Noisy Examples in Imbalanced Classification by Combining SMOTE with Ensemble Filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera |
IDEAL | 1 |
| 2014 | On the characterization of noise filters for self-training semi-supervised in nearest neighbor classification
Isaac Triguero, José A. Sáez, Julián Luengo, Salvador García 0001, Francisco Herrera |
Neurocomputing | 2 |
| 2014 | Analyzing the presence of noise in multi-class problems: alleviating its influence with the One-vs-One decomposition
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera |
Knowl. Inf. Syst. | 1 |
| 2014 | Statistical computation of feature weighting schemes through data estimation for nearest neighbor classifiers
José A. Sáez, Joaquín Derrac, Julián Luengo, Francisco Herrera |
Pattern Recognit. | 1 |
| 2013 | Tackling the problem of classification with noisy data using Multiple Classifier Systems: Analysis of the performance and robustness
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera |
Inf. Sci. | 1 |
| 2013 | Predicting noise filtering efficacy with data complexity measures for nearest neighbor classification
José A. Sáez, Julián Luengo, Francisco Herrera |
Pattern Recognit. | 1 |
| 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised LearningabstractDiscretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones. Salvador García 0001, Julián Luengo, José A. Sáez, Victoria López, Francisco Herrera |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Missing data imputation for fuzzy rule-based classification systems
Julián Luengo, José A. Sáez, Francisco Herrera |
Soft Comput. | 2 |
| 2012 | Study on the Impact of Partition-Induced Dataset Shift on k -Fold Cross-ValidationabstractCross-validation is a very commonly employed technique used to evaluate classifier performance. However, it can potentially introduce dataset shift, a harmful factor that is often not taken into account and can result in inaccurate performance estimation. This paper analyzes the prevalence and impact of partition-induced covariate shift on different k-fold cross-validation schemes. From the experimental results obtained, we conclude that the degree of partition-induced covariate shift depends on the cross-validation scheme considered. In this way, worse schemes may harm the correctness of a single-classifier performance estimation and also increase the needed number of repetitions of cross-validation to reach a stable performance estimation. Jose G. Moreno-Torres, José A. Sáez, Francisco Herrera |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | Fuzzy Rule Based Classification Systems versus crisp robust learners trained in presence of class noise's effects: A case of studyabstractThe presence of noise is common in any real-world dataset and may adversely affect the accuracy, construction time and complexity of the classifiers in this context. Traditionally, many algorithms have incorporated mechanisms to deal with noisy problems and reduce noise's effects on performance; they are called robust learners. The C4.5 crisp algorithm is a well-known example of this group of methods. On the other hand, models built by Fuzzy Rule Based Classification Systems are widely recognized for their robustness to imperfect data, but also for their interpretability. The aim of this contribution is to analyze the good behavior and robustness of Fuzzy Rule Based Classification Systems when noise is present in the examples' class labels, especially versus robust learners. In order to accomplish this study, a large number of datasets are created by introducing different levels of noise into the class labels in the training sets. We compare a Fuzzy Rule Based Classification System, the Fuzzy Unordered Rule Induction Algorithm, with respect to the C4.5 classic robust learner which is considered tolerant to noise. From the results obtained it is possible to observe that Fuzzy Rule Based Classification Systems have a good tolerance, in comparison to the C4.5 algorithm, to class noise. José A. Sáez, Julián Luengo, Francisco Herrera |
ISDA | 1 |