EDBT 2026 Demo / reviewers in the wild / expert
David Guijo-Rubio
dblp:185/3448
· DBLP profile ↗
20ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-8035-4057ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOC-UCO: a comprehensive repository of tabular ordinal classification datasetsabstractAn Ordinal Classification (OC) problem corresponds to a special type of classification characterised by the presence of a natural order relationship among the classes. This type of problem, that can be found in a number of real-world applications, has motivated the design and development of many ordinal methodologies over the last years. However, it is important to highlight that the development of the OC field suffers from one main disadvantage: the lack of a comprehensive set of datasets on which novel approaches to the literature are benchmarked. In order to approach this objective, this manuscript from the University of Córdoba (UCO), which has previous experience on the OC field, provides the literature with a publicly available repository of tabular data for a robust validation of novel OC approaches, namely TOC-UCO (Tabular Ordinal Classification repository of the UCO). Specifically, this repository includes a set of tabular ordinal datasets that have been preprocessed under a common framework and that have a reasonable number of patterns and an appropriate class distribution. We also provide the sources and preprocessing steps of each dataset, along with details on how to benchmark a novel approach using the TOC-UCO repository. For this, indices for different randomised train-test partitions are provided to facilitate the reproducibility of the experiments. • Introduction of a novel ordinal classification repository: TOC-UCO . • Extension of the number of datasets with a high number of classes. • Analysis of the previous ordinal classification benchmarking repository. • In-depth comparison between TOC-UCO and the previous repository. • Presentation of a baseline experimentation on the new TOC-UCO archive. Rafael Ayllón-Gavilán, David Guijo-Rubio, Antonio M. Gómez-Orellana, Francisco Bérchez-Moreno, Víctor Manuel Vargas Yun, Pedro Antonio Gutiérrez |
Neurocomputing | 2 |
| 2026 | Splitting criteria for ordinal decision trees: An experimental studyabstractOrdinal Classification (OC) addresses those classification tasks where the labels exhibit a natural order. Unlike nominal classification, which treats all classes as mutually exclusive and unordered, OC takes the ordinal relationship into account, producing more accurate and relevant results. This is particularly critical in applications where the magnitude of classification errors has significant consequences. Despite this, OC problems are often tackled using nominal methods, leading to suboptimal solutions. Although decision trees are among the most popular classification approaches, ordinal tree-based approaches have received less attention when compared to other classifiers. This work provides a comprehensive survey of ordinal splitting criteria, standardising the notations used in the literature to enhance clarity and consistency. Three ordinal splitting criteria, Ordinal Gini (OGini), Weighted Information Gain, and Ranking Impurity, are compared to the nominal counterparts of the first two (Gini and information gain), by incorporating them into a decision tree classifier. An extensive repository considering 45 publicly available OC datasets is presented, supporting the first experimental comparison of ordinal and nominal splitting criteria using well-known OC evaluation metrics. The results have been statistically analysed, highlighting that OGini stands out as the best ordinal splitting criterion to date, reducing the mean absolute error achieved by Gini by more than 3.02 % . To promote reproducibility, all source code developed, a detailed guide for reproducing the results, the 45 OC datasets, and the individual results for all the evaluated methodologies are provided. Rafael Ayllón-Gavilán, Francisco J. Martínez-Estudillo, David Guijo-Rubio, César Hervás-Martínez, Pedro Antonio Gutiérrez |
Pattern Recognit. | 3 |
| 2026 | Soft Labelling for Deep Ordinal Classification: An Experimental ReviewabstractOrdinal classification, where labels follow a natural order, has gained increasing attention, particularly in the deep learning community due to its relevance in tasks such as age estimation, medical grading, and quality assessment. Despite the growing number of deep ordinal classification methods, a comprehensive experimental analysis of their core ordinal components remains lacking. This work presents a systematic evaluation of deep ordinal classifiers by analysing the impact of three key modelling choices: the loss function, output layer, and labelling strategy. To analyse their effects, we adopt a unified architecture and evaluate one nominal and 19 ordinal configurations, resulting from combination of two loss functions, two output layers, and five labelling strategies. These configurations are assessed on 12 diverse ordinal image datasets using six performance metrics, including both ordinal and nominal measures. Results show that ordinal output layers consistently outperform softmax, and that soft labelling generally improves generalisation. While categorical cross-entropy achieves better average performance, especially on nominal metrics, no configuration performs best across all datasets. Statistical analyses indicate significant interactions between losses, outputs, labelling strategies, and datasets, highlighting the need to adapt methodological choices to specific tasks. These findings provide valuable guidance for designing robust deep ordinal classification models. Víctor Manuel Vargas Yun, David Guijo-Rubio, Rafael Ayllón-Gavilán, Antonio M. Gómez-Orellana, Pedro Antonio Gutiérrez, César Hervás-Martínez |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | dlordinal: A Python package for deep ordinal classificationabstractdlordinal is a new Python library that unifies many recent deep ordinal classification methodologies available in the literature. Developed using PyTorch as underlying framework, it implements the top performing state-of-the-art deep learning techniques for ordinal classification problems. Ordinal approaches are designed to leverage the ordering information present in the target variable. Specifically, it includes loss functions, various output layers, dropout techniques, soft labelling methodologies, and other classification strategies, all of which are appropriately designed to incorporate the ordinal information. Furthermore, as the performance metrics to assess novel proposals in ordinal classification depend on the distance between target and predicted classes in the ordinal scale, suitable ordinal evaluation metrics are also included. dlordinal is distributed under the BSD-3-Clause license and is available at https://github.com/ayrna/dlordinal. Francisco Bérchez-Moreno, Rafael Ayllón-Gavilán, Víctor Manuel Vargas Yun, David Guijo-Rubio, César Hervás-Martínez, Juan Carlos Fernández 0001, Pedro Antonio Gutiérrez |
Neurocomputing | 4 |
| 2025 | Convolutional- and Deep Learning-Based Techniques for Time Series Ordinal ClassificationabstractTime-series classification (TSC) covers the supervised learning problem where input data is provided in the form of series of values observed through repeated measurements over time, and whose objective is to predict the category to which they belong. When the class values are ordinal, classifiers that take this into account can perform better than nominal classifiers. Time-series ordinal classification (TSOC) is the field bridging this gap, yet unexplored in the literature. There are a wide range of time-series problems showing an ordered label structure, and TSC techniques that ignore the order relationship discard useful information. Hence, this article presents the first benchmarking of TSOC methodologies, exploiting the ordering of the target labels to boost the performance of current TSC state of the art. Both convolutional- and deep-learning-based methodologies (among the best performing alternatives for nominal TSC) are adapted for TSOC. For the experiments, a selection of 29 ordinal problems has been made. In this way, this article contributes to the establishment of the state of the art in TSOC. The results obtained by ordinal versions are found to be significantly better than current nominal TSC techniques in terms of ordinal performance metrics, outlining the importance of considering the ordering of the labels when dealing with this kind of problems. Rafael Ayllón-Gavilán, David Guijo-Rubio, Pedro Antonio Gutiérrez, Anthony J. Bagnall, César Hervás-Martínez |
IEEE Trans. Cybern. | 2 |
| 2024 | A Hands-on Introduction to Time Series Classification and RegressionabstractTime series classification and regression are rapidly evolving fields that find areas of application in all domains of machine learning and data science. This hands on tutorial will provide an accessible overview of the recent research in these fields, using code examples to introduce the process of implementing and evaluating an estimator. We will show how to easily reproduce published results and how to compare a new algorithm to state-of-the-art. Finally, we will work through real world examples from the field of Electroencephalogram (EEG) classification and regression. EEG machine learning tasks arise in medicine, brain-computer interface research and psychology. We use these problems to how to compare algorithms on problems from a single domain and how to deal with data with different characteristics, such as missing values, unequal length and high dimensionality. The latest advances in the fields of time series classification and regression are all available through the aeon toolkit, an open source, scikit-learn compatible framework for time series machine learning which we use to provide our code examples. Anthony J. Bagnall, Matthew Middlehurst, Germain Forestier, Ali Ismail-Fawaz, Antoine Guillaume, David Guijo-Rubio, Chang Wei Tan, Angus Dempster, Geoffrey I. Webb |
KDD | 6 |
| 2024 | Unsupervised feature based algorithms for time series extrinsic regressionabstractAbstract Time Series Extrinsic Regression (TSER) involves using a set of training time series to form a predictive model of a continuous response variable that is not directly related to the regressor series. The TSER archive for comparing algorithms was released in 2022 with 19 problems. We increase the size of this archive to 63 problems and reproduce the previous comparison of baseline algorithms. We then extend the comparison to include a wider range of standard regressors and the latest versions of TSER models used in the previous study. We show that none of the previously evaluated regressors can outperform a regression adaptation of a standard classifier, rotation forest. We introduce two new TSER algorithms developed from related work in time series classification. FreshPRINCE is a pipeline estimator consisting of a transform into a wide range of summary features followed by a rotation forest regressor. DrCIF is a tree ensemble that creates features from summary statistics over random intervals. Our study demonstrates that both algorithms, along with InceptionTime, exhibit significantly better performance compared to the other 18 regressors tested. More importantly, DrCIF is the only one that significantly outperforms a standard rotation forest regressor. David Guijo-Rubio, Matthew Middlehurst, Guilherme Arcencio, Diego Furtado Silva, Anthony J. Bagnall |
Data Min. Knowl. Discov. | 1 |
| 2024 | ORFEO: Ordinal classifier and Regressor Fusion for Estimating an Ordinal categorical targetabstractIn this paper we present a novel methodology, referenced as ORFEO (Ordinal classifier and Regressor Fusion for Estimating an Ordinal categorical target), to enhance the performance in ordinal classification problems for which the latent variable is observable. ORFEO is an artificial neural network model incorporating two outputs, one for ordinal classification, using the cumulative link model, and one for regression, using a linear model. Both outputs are simultaneously optimised considering a loss function that linearly combines both classification and regression losses. The main motivation behind developing the proposed approach is to enhance the performance of a standard ordinal classifier. This improvement is facilitated by considering the regression output, which allows the model to differentiate between patterns within the same category. The ORFEO model is applied to two problems in the field of marine and ocean engineering: short-term prediction of both significant wave height and flux of energy. Both problems are addressed considering four different coastal zones of the United States of America, using 13 datasets formed by buoys measurements and reanalysis data. A comprehensive comparison against 20 methodologies, including regression and nominal/ordinal classification approaches is performed, by using diverse nominal and ordinal performance metrics. Ranks achieved indicate that ORFEO outperforms all the compared methodologies in terms of all the performance measures, demonstrating the efficacy and robustness of the proposal. Finally, a statistical analysis is conducted, concluding that there are statistically significant differences across ordinal and nominal performance metrics in favour of the proposed ORFEO model. Antonio M. Gómez-Orellana, David Guijo-Rubio, Pedro Antonio Gutiérrez, César Hervás-Martínez, Víctor Manuel Vargas Yun |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | aeon: a Python Toolkit for Learning from Time Seriesabstractaeon is a unified Python 3 library for all machine learning tasks involving time series. The package contains modules for time series forecasting, classification, extrinsic regression and clustering, as well as a variety of utilities, transformations and distance measures designed for time series data. aeon also has a number of experimental modules for tasks such as anomaly detection, similarity search and segmentation. aeon follows the scikit-learn API as much as possible to help new users and enable easy integration of aeon estimators with useful tools such as model selection and pipelines. It provides a broad library of time series algorithms, including efficient implementations of the very latest advances in research. Using a system of optional dependencies, aeon integrates a wide variety of packages into a single interface while keeping the core framework with minimal dependencies. The package is distributed under the 3-Clause BSD license and is available at https://github.com/aeon-toolkit/aeon. Matthew Middlehurst, Ali Ismail-Fawaz, Antoine Guillaume, Christopher Holder, David Guijo-Rubio, Guzal Bulatova, Leonidas Tsaprounis, Lukasz Mentel, Martin Walter, Patrick Schäfer 0001, Anthony J. Bagnall |
J. Mach. Learn. Res. | 5 |
| 2024 | EBANO: A novel Ensemble BAsed on uNimodal Ordinal classifiers for the prediction of significant wave heightabstractIn this study, we present EBANO (Ensemble BAsed on uNimodal Ordinal classifiers), which is a novel ensemble approach of ordinal classifiers that includes four soft labelling approaches along with an ordinal logistic regression model. These models are integrated within the ensemble using a new aggregation methodology that automatically weights each individual classifier using a randomised search algorithm. In addition, the proposed EBANO methodology is applied to tackle short-term prediction of Significant Wave Height (SWH). Thus, we employ EBANO using a diverse set of eight datasets derived from reanalysis data and buoy-recorded SWH measurements. To approach the problem from an ordinal classification perspective, the SWH values are discretised into five ordered classes by applying hierarchical clustering. EBANO is compared with each of the individual classifiers integrated in the proposed ensemble along with a different ensemble technique termed HESCA. Both the average results and the ranks obtained show the superiority of EBANO over the compared methodologies, being more pronounced in the metrics that account for the imbalance present in the datasets considered. Finally, a statistical analysis is performed, confirming the statistical significance of the observed differences in all comparisons. This analysis underscores the effectiveness of EBANO in addressing the problem of SWH prediction, showcasing its excellence. Víctor Manuel Vargas Yun, Antonio M. Gómez-Orellana, Pedro Antonio Gutiérrez, César Hervás-Martínez, David Guijo-Rubio |
Knowl. Based Syst. | 5 |
| 2023 | Cluster analysis and forecasting of viruses incidence growth curves: Application to SARS-CoV-2abstractThe sanitary emergency caused by COVID-19 has compromised countries and generated a worldwide health and economic crisis. To provide support to the countries' responses, numerous lines of research have been developed. The spotlight was put on effectively and rapidly diagnosing and predicting the evolution of the pandemic, one of the most challenging problems of the past months. This work contributes to the existing literature by developing a two-step methodology to analyze the transmission rate, designing models applied to territories with similar pandemic behavior characteristics. Virus transmission is considered as bacterial growth curves to understand the spread of the virus and to make predictions about its future evolution. Hence, an analytical clustering procedure is first applied to create groups of locations where the virus transmission rate behaved similarly in the different outbreaks. A curve decomposition process based on an iterative polynomial process is then applied, obtaining meaningful forecasting features. Information of the territories belonging to the same cluster is merged to build models capable of simultaneously predicting the 14-day incidence in several locations using Evolutionary Artificial Neural Networks. The methodology is applied to Andalusia (Spain), although it is applicable to any region across the world. Individual models trained for a specific territory are carried out for comparison purposes. The results demonstrate that this methodology achieves statistically similar, or even better, performance for most of the locations. In addition to being extremely competitive, the main advantage of the proposal lies in its complexity cost reduction. The total number of parameters to be estimated is reduced up to 93.51% for the short term and 93.31% for the mid-term forecasting, respectively. Moreover, the number of required models is reduced by 73.53% and 58.82% for the short- and mid-term forecasting horizons. Miguel Díaz-Lozano, David Guijo-Rubio, Pedro Antonio Gutiérrez, César Hervás-Martínez |
Expert Syst. Appl. | 2 |
| 2023 | Generalised triangular distributions for ordinal deep learning: Novel proposal and optimisation
Víctor Manuel Vargas Yun, Antonio Manuel Durán-Rosal, David Guijo-Rubio, Pedro Antonio Gutiérrez, César Hervás-Martínez |
Inf. Sci. | 3 |
| 2022 | COVID-19 contagion forecasting framework based on curve decomposition and evolutionary artificial neural networks: A case study in Andalusia, Spain
Miguel Díaz-Lozano, David Guijo-Rubio, Pedro Antonio Gutiérrez, Antonio M. Gómez-Orellana, Isaac Túñez, Luis Ortigosa-Moreno, Armando Romanos-Rodríguez, Javier Padillo-Ruiz, César Hervás-Martínez |
Expert Syst. Appl. | 2 |
| 2021 | Enhancing the ORCA framework with a new Fuzzy Rule Base System implementation compatible with the JFML libraryabstractClassification and regression techniques are two of the main tasks considered by the Machine Learning area. They mainly depend on the target variable to predict. In this context, ordinal classification represents an intermediate task, which is focused on the prediction of nominal variables where the categories follow a specific intrinsic order given by the problem. Nevertheless, the integration of different algorithms able to solve ordinal classification problems is often unavailable in most of existing Machine Learning software, which hinders the use of new approaches. Therefore, this paper focuses on the incorporation of an ordinal classification algorithm (NSLVOrd) in one of the most complete ordinal regression frameworks, “Ordinal Regression and Classification Algorithms framework (ORCA)” by using both fuzzy rules and the JFML library. The use of NSLVOrd in the ORCA tool as well as a case study with a real database are shown where the obtained results are promising. Francisco J. Rodríguez-Lozano, David Guijo-Rubio, Pedro Antonio Gutiérrez, José M. Soto-Hidalgo, Juan Carlos Gámez |
FUZZ-IEEE | 2 |
| 2021 | Time-Series Clustering Based on the Characterization of Segment TypologiesabstractTime-series clustering is the process of grouping time series with respect to their similarity or characteristics. Previous approaches usually combine a specific distance measure for time series and a standard clustering method. However, these approaches do not take the similarity of the different subsequences of each time series into account, which can be used to better compare the time-series objects of the dataset. In this article, we propose a novel technique of time-series clustering consisting of two clustering stages. In a first step, a least-squares polynomial segmentation procedure is applied to each time series, which is based on a growing window technique that returns different-length segments. Then, all of the segments are projected into the same dimensional space, based on the coefficients of the model that approximates the segment and a set of statistical features. After mapping, a first hierarchical clustering phase is applied to all mapped segments, returning groups of segments for each time series. These clusters are used to represent all time series in the same dimensional space, after defining another specific mapping process. In a second and final clustering stage, all the time-series objects are grouped. We consider internal clustering quality to automatically adjust the main parameter of the algorithm, which is an error threshold for the segmentation. The results obtained on 84 datasets from the UCR Time Series Classification Archive have been compared against three state-of-the-art methods, showing that the performance of this methodology is very promising, especially on larger datasets. David Guijo-Rubio, Antonio Manuel Durán-Rosal, Pedro Antonio Gutiérrez, Alicia Troncoso Lora, César Hervás-Martínez |
IEEE Trans. Cybern. | 1 |
| 2020 | Time series ordinal classification via shapeletsabstractNominal time series classification has been widely developed over the last years. However, to the best of our knowledge, ordinal classification of time series is an unexplored field, and this paper proposes a first approach in the context of the shapelet transform (ST). For those time series dataset where there is a natural order between the labels and the number of classes is higher than 2, nominal classifiers are not capable of achieving the best results, because the models impose the same cost of misclassification to all the errors, regardless the difference between the predicted and the ground-truth. In this sense, we consider four different evaluation metrics to do so, three of them of an ordinal nature. The first one is the widely known Information Gain (IG), proved to be very competitive for ST methods, whereas the remaining three measures try to boost the order information by refining the quality measure. These three measures are a reformulation of the Fisher score, the Spearman's correlation coefficient (ρ), and finally, the Pearson's correlation coefficient (R2). An empirical evaluation is carried out, considering 7 ordinal datasets from the UEA & UCR time series classification repository, 4 classifiers (2 of them of nominal nature, whereas the other 2 are of ordinal nature) and 2 performance measures (correct classification rate, CCR, and average mean absolute error, AMAE). The results show that, for both performance metrics, the ST quality metric based on R2is able to obtain the best results, specially for AMAE, for which the differences are statistically significant in favour of R2. David Guijo-Rubio, Pedro Antonio Gutiérrez, Anthony J. Bagnall, César Hervás-Martínez |
IJCNN | 1 |
| 2020 | Prediction of convective clouds formation using evolutionary neural computation techniques
David Guijo-Rubio, Pedro Antonio Gutiérrez, Carlos Casanova-Mateo, Juan Carlos Fernández 0001, Antonio M. Gómez-Orellana, Pablo Salvador-González, Sancho Salcedo-Sanz, César Hervás-Martínez |
Neural Comput. Appl. | 1 |
| 2019 | A Hybrid Approach to Time Series Classification with Shapelets
David Guijo-Rubio, Pedro Antonio Gutiérrez, Romain Tavenard, Anthony J. Bagnall |
IDEAL (1) | 1 |
| 2019 | Modelling Survival by Machine Learning Methods in Liver Transplantation: Application to the UNOS Dataset
David Guijo-Rubio, Pedro J. Villalón-Vaquero, Pedro Antonio Gutiérrez, María Dolores Ayllón-Terán, Javier Briceño, César Hervás-Martínez |
IDEAL (2) | 1 |
| 2018 | Distribution-Based Discretisation and Ordinal Classification Applied to Wave Height Prediction
David Guijo-Rubio, Antonio Manuel Durán-Rosal, Antonio M. Gómez-Orellana, Pedro Antonio Gutiérrez, César Hervás-Martínez |
IDEAL (2) | 1 |