EDBT 2026 Demo / reviewers in the wild / expert
Vítor Cerqueira
dblp:166/2054
· DBLP profile ↗
28ranked-venue papers
18as first author
16since 2021 · last 2026
0000-0002-9694-8423ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 16 first-author · 15 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 2 since 2021Theory of computation · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grasynda: Graph-Based Synthetic Time Series Generation
Luis Amorim, Moisés Santos, Paulo J. Azevedo, Carlos Soares, Vítor Cerqueira |
IDA | 5 |
| 2025 | Stress-Testing of Multimodal Models in Medical Image-Based Report GenerationabstractMultimodal models, namely vision-language models, present unique possibilities through the seamless integration of different information mediums for data generation. These models mostly act as a black-box, making them lack transparency and explicability. Reliable results require accountable and trustworthy Artificial Intelligence (AI), namely when in use for critical tasks, such as the automatic generation of medical imaging reports for healthcare diagnosis. By exploring stress-testing techniques, multimodal generative models can become more transparent by disclosing their shortcomings, further supporting their responsible usage in the medical field. Flávia Carvalhido, Henrique Lopes Cardoso, Vítor Cerqueira |
AAAI | 3 |
| 2025 | Cherry-Picking in Time Series Forecasting: How to Select Datasets to Make Your Model ShineabstractThe importance of time series forecasting drives continuous research and the development of new approaches to tackle this problem. Typically, these methods are introduced through empirical studies that frequently claim superior accuracy for the proposed approaches. Nevertheless, concerns are rising about the reliability and generalizability of these results due to limitations in experimental setups. This paper addresses a critical limitation: the number and representativeness of the datasets used. We investigate the impact of dataset selection bias, particularly the practice of cherry-picking datasets, on the performance evaluation of forecasting methods. Through empirical analysis with a diverse set of benchmark datasets, our findings reveal that cherry-picking datasets can significantly distort the perceived performance of methods, often exaggerating their effectiveness. Furthermore, our results demonstrate that by selectively choosing just four datasets — what most studies report — 46% of methods could be deemed best in class, and 77% could rank within the top three. Additionally, recent deep learning-based approaches show high sensitivity to dataset selection, whereas classical methods exhibit greater robustness. Finally, our results indicate that, when empirically validating forecasting algorithms on a subset of the benchmarks, increasing the number of datasets tested from 3 to 6 reduces the risk of incorrectly identifying an algorithm as the best one by approximately 40%. Our study highlights the critical need for comprehensive evaluation frameworks that more accurately reflect real-world scenarios. Adopting such frameworks will ensure the development of robust and reliable forecasting methods. Luis Roque, Vítor Cerqueira, Carlos Soares, Luís Torgo |
AAAI | 2 |
| 2025 | Meta Subspace Analysis: Understanding Model (Mis)behavior in the Metafeature Space
Carlos Soares, Paulo J. Azevedo, Vítor Cerqueira, Luís Torgo |
DS | 3 |
| 2025 | Meta-learning and Data Augmentation for Stress Testing Forecasting Models
Ricardo Inácio, Vítor Cerqueira, Marília Barandas, Carlos Soares |
IDA | 2 |
| 2025 | Modelradar: aspect-based forecast evaluationabstractAbstract Accurate evaluation of forecasting models is essential for ensuring reliable predictions. Current practices for evaluating and comparing forecasting models focus on summarising performance into a single score, using metrics such as SMAPE. While convenient, averaging performance over all samples dilutes relevant information about model behaviour under varying conditions. This limitation is especially problematic for time series forecasting, where multiple layers of averaging–across time steps, horizons, and multiple time series in a dataset–can mask relevant performance variations. We address this limitation by proposing ModelRadar, a framework for evaluating univariate time series forecasting models across multiple aspects, such as stationarity, presence of anomalies, or forecasting horizons. We demonstrate the advantages of this framework by comparing 24 forecasting methods, including classical approaches and different machine learning algorithms. PatchTST, a state-of-the-art transformer-based neural network architecture, performs best overall but its superiority varies with forecasting conditions. For instance, concerning the forecasting horizon, we found that PatchTST (and also other neural networks) only outperforms classical approaches for multi-step ahead forecasting. Another relevant insight is that classical approaches such as ETS or Theta are notably more robust in the presence of anomalies. These and other findings highlight the importance of aspect-based model evaluation for both practitioners and researchers. ModelRadar is available as a Python package. Vítor Cerqueira, Luis Roque, Carlos Soares |
Mach. Learn. | 1 |
| 2025 | Mast: interpretable stress testing via meta-learning for forecasting model robustness evaluation
Ricardo Inácio, Vítor Cerqueira, Marília Barandas, Carlos Soares |
Mach. Learn. | 2 |
| 2024 | Forecasting with Deep Learning: Beyond Average of Average of Average Performance
Vítor Cerqueira, Luis Roque, Carlos Soares |
DS (1) | 1 |
| 2024 | VEST: automatic feature engineering for forecasting
Vítor Cerqueira, Nuno Moniz, Carlos Soares |
Mach. Learn. | 1 |
| 2023 | STUDD: a student-teacher method for unsupervised concept drift detection
Vítor Cerqueira, Heitor Murilo Gomes, Albert Bifet, Luís Torgo |
Mach. Learn. | 1 |
| 2023 | Automated imbalanced classification via layered learning
Vítor Cerqueira, Luís Torgo, Paula Branco, Colin Bellinger |
Mach. Learn. | 1 |
| 2023 | Early anomaly detection in time series: a hierarchical approach for predicting critical health episodes
Vítor Cerqueira, Luís Torgo, Carlos Soares |
Mach. Learn. | 1 |
| 2023 | Model Selection for Time Series Forecasting An Empirical Analysis of Multiple Estimators
Vítor Cerqueira, Luís Torgo, Carlos Soares |
Neural Process. Lett. | 1 |
| 2022 | A case study comparing machine learning with statistical methods for time series forecasting: size matters
Vítor Cerqueira, Luís Torgo, Carlos Soares |
J. Intell. Inf. Syst. | 1 |
| 2021 | Empirical Study on the Impact of Different Sets of Parameters of Gradient Boosting Algorithms for Time-Series Forecasting with LightGBM
Filipa Barros, Vítor Cerqueira, Carlos Soares |
PRICAI (1) | 2 |
| 2021 | Automated imbalanced classification via meta-learning
Nuno Moniz, Vítor Cerqueira |
Expert Syst. Appl. | 2 |
| 2020 | Unsupervised Concept Drift Detection Using a Student-Teacher Approach
Vítor Cerqueira, Heitor Murilo Gomes, Albert Bifet |
DS | 1 |
| 2020 | Evaluating time series forecasting models: an empirical study on performance estimation methods
Vítor Cerqueira, Luís Torgo, Igor Mozetic |
Mach. Learn. | 1 |
| 2019 | Layered Learning for Early Anomaly Detection: Predicting Critical Health Episodes
Vítor Cerqueira, Luís Torgo, Carlos Soares |
DS | 1 |
| 2019 | Arbitrage of forecasting experts
Vítor Cerqueira, Luís Torgo, Fábio Pinto, Carlos Soares |
Mach. Learn. | 1 |
| 2018 | SMOTEBoost for Regression: Improving the Prediction of Extreme ValuesabstractSupervised learning with imbalanced domains is one of the biggest challenges in machine learning. Such tasks differ from standard learning tasks by assuming a skewed distribution of target variables, and user domain preference towards under-represented cases. Most research has focused on imbalanced classification tasks, where a wide range of solutions has been tested. Still, little work has been done concerning imbalanced regression tasks. In this paper, we propose an adaptation of the SMOTEBoost approach for the problem of imbalanced regression. Originally designed for classification tasks, it combines boosting methods and the SMOTE resampling strategy. We present four variants of SMOTEBoost and provide an experimental evaluation using 30 datasets with an extensive analysis of results in order to assess the ability of SMOTEBoost methods in predicting extreme target values, and their predictive trade-off concerning baseline boosting methods. SMOTEBoost is publicly available in a software package. Nuno Moniz, Rita P. Ribeiro, Vítor Cerqueira, Nitesh V. Chawla |
DSAA | 3 |
| 2018 | Constructive Aggregation and Its Application to Forecasting with Dynamic Ensembles
Vítor Cerqueira, Fábio Pinto, Luís Torgo, Carlos Soares, Nuno Moniz |
ECML/PKDD (1) | 1 |
| 2018 | On Evaluating Floating Car Data Quality for Knowledge DiscoveryabstractFloating car data (FCD) denotes the type of data (location, speed, and destination) produced and broadcasted periodically by running vehicles. Increasingly, intelligent transportation systems take advantage of such data for prediction purposes as input to road and transit control and to discover useful mobility patterns with applications to transport service design and planning, to name just a few applications. However, there are considerable quality issues that affect the usefulness and efficacy of FCD in these many applications. In this paper, we propose a methodology to compute such quality indicators automatically for large FCD sets. It leverages on a set of statistical indicators (named Yuki-san) covering multiple dimensions of FCD such as spatio-temporal coverage, accuracy, and reliability. As such, the Yuki-san indicators provide a quick and intuitive means to assess the potential “value” and “veracity” characteristics of the data. Experimental results with two mobility-related data mining and supervised learning tasks on the basis of two real-world FCD sources show that the Yuki-san indicators are indeed consistent with how well the applications perform using the data. With a wider variety of FCD (e.g., from navigation systems and CAN buses) becoming available, further research and validation into the dimensions covered and the efficacy of the Yuki-San indicators is needed. Vítor Cerqueira, Luís Moreira-Matias, Jihed Khiari, J. W. C. van Lint |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Dynamic and Heterogeneous Ensembles for Time Series ForecastingabstractThis paper addresses the issue of learning time series forecasting models in changing environments by leveraging the predictive power of ensemble methods. Concept drift adaptation is performed in an active manner, by dynamically combining base learners according to their recent performance using a non-linear function. Diversity in the ensembles is encouraged with several strategies that include heterogeneity among learners, sampling techniques and computation of summary statistics as extra predictors. Heterogeneity is used with the goal of better coping with different dynamic regimes of the time series. The driving hypotheses of this work are that (i) heterogeneous ensembles should better fit different dynamic regimes and (ii) dynamic aggregation should allow for fast detection and adaptation to regime changes. We extend some strategies typically used in classification tasks to time series forecasting. The proposed methods are validated using Monte Carlo simulations on 16 real-world univariate time series with numerical outcome as well as an artificial series with clear regime shifts. The results provide strong empirical evidence for our hypotheses. To encourage reproducibility the proposed method is publicly available as a software package. Vítor Cerqueira, Luís Torgo, Mariana Oliveira 0001, Bernhard Pfahringer |
DSAA | 1 |
| 2017 | A Comparative Study of Performance Estimation Methods for Time Series ForecastingabstractPerformance estimation denotes a task of estimating the loss that a predictive model will incur on unseen data. These procedures are part of the pipeline in every machine learning task and are used for assessing the overall generalisation ability of models. In this paper we address the application of these methods to time series forecasting tasks. For independent and identically distributed data the most common approach is cross-validation. However, the dependency among observations in time series raises some caveats about the most appropriate way to estimate performance in these datasets and currently there is no settled way to do so. We compare different variants of cross-validation and different variants of out-of-sample approaches using two case studies: One with 53 real-world time series and another with three synthetic time series. Results show noticeable differences in the performance estimation methods in the two scenarios. In particular, empirical experiments suggest that cross-validation approaches can be applied to stationary synthetic time series. However, in real-world scenarios the most accurate estimates are produced by the out-of-sample methods, which preserve the temporal order of observations. Vítor Cerqueira, Luís Torgo, Jasmina Smailovic, Igor Mozetic |
DSAA | 1 |
| 2017 | Arbitrated Ensemble for Time Series Forecasting
Vítor Cerqueira, Luís Torgo, Fábio Pinto, Carlos Soares |
ECML/PKDD (2) | 1 |
| 2016 | Combining Boosted Trees with Metafeature Engineering for Predictive Maintenance
Vítor Cerqueira, Fábio Pinto, Cláudio Rebelo de Sá, Carlos Soares |
IDA | 1 |
| 2016 | Automated Setting of Bus Schedule Coverage Using Unsupervised Machine Learning
Jihed Khiari, Luís Moreira-Matias, Vítor Cerqueira, Oded Cats |
PAKDD (1) | 3 |