VLDB 2026 Research / reviewers in the wild / expert
Carlos Soares
dblp:50/4846
· DBLP profile ↗
43ranked-venue papers in the field
2as first author
12since 2021 · last 2026
0000-0003-4549-8917ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 30 (2 first)Information Retrieval & Web Search · 5Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grasynda: Graph-Based Synthetic Time Series Generation
Luis Amorim, Moisés Santos, Paulo J. Azevedo, Carlos Soares, Vítor Cerqueira |
IDA | 4 |
| 2025 | Read-write LSTM: A Novel Approach Integrating Backpropagation to Data in LSTMabstractTraditional recurrent neural networks operate as passive observers of data, unable to modify the information they learn from despite errors that may arise from suboptimal input representations. We introduce Read & Write LSTM (read-write LSTM), a new variant within the family of read & write machine learning (RW-ML) architectures that address this fundamental limitation by integrating input modification directly into the backpropagation process. Read-write LSTM establishes a dynamic feedback loop where input representations evolve alongside model weights through gradient transformation mechanisms. Our approach introduces a principled gradient scaling framework with an adaptive correction rate that carefully controls the extent of data modification, preserving data integrity while enhancing representational power. We comprehensively evaluate read-write LSTM against traditional LSTMs and state-of-the-art transformer models on the M4 competition and Numenta Anomaly Benchmark datasets, demonstrating significant improvements in forecasting accuracy. Notably, read-write LSTM consistently out-performs standard LSTM models in over 70% of time series with complex patterns and achieves superior performance on 55% of anomaly-rich datasets. Through extensive experimentation and analysis, we establish both the theoretical foundations and practical benefits of integrating data modification with neural computation, paving the way for a new generation of adaptive learning systems that actively reshape their inputs rather than merely adapting to them. Yassine Baghoussi, Carlos Soares, João Mendes-Moreira 0001 |
ICDM | 2 |
| 2025 | Meta-learning and Data Augmentation for Stress Testing Forecasting Models
Ricardo Inácio, Vítor Cerqueira, Marília Barandas, Carlos Soares |
IDA | 4 |
| 2025 | Evaluating Transfer Learning Methods on Real-World Data Streams: A Case Study in Financial Fraud Detection
Ricardo Ribeiro Pereira, Jacopo Bono, Hugo M. Ferreira, Pedro Ribeiro 0004, Carlos Soares, Pedro Bizarro |
ECML/PKDD (9) | 5 |
| 2024 | Kernel Corrector LSTM
Rodrigo Tuna, Yassine Baghoussi, Carlos Soares, João Mendes-Moreira 0001 |
IDA (2) | 3 |
| 2024 | RHiOTS: A Framework for Evaluating Hierarchical Time Series Forecasting AlgorithmsabstractWe introduce the Robustness of Hierarchically Organized Time Series (RHiOTS) framework, designed to assess the robustness of hierarchical time series forecasting models and algorithms on real-world datasets. Hierarchical time series, where lower-level forecasts must sum to upper-level ones, are prevalent in various contexts, such as retail sales across countries. Current empirical evaluations of forecasting methods are often limited to a small set of benchmark datasets, offering a narrow view of algorithm behavior. RHiOTS addresses this gap by systematically altering existing datasets and modifying the characteristics of individual series and their interrelations. It uses a set of parameterizable transformations to simulate those changes in the data distribution. Additionally, RHiOTS incorporates an innovative visualization component, turning complex, multidimensional robustness evaluation results into intuitive, easily interpretable visuals. This approach allows an in-depth analysis of algorithm and model behavior under diverse conditions. We illustrate the use of RHiOTS by analyzing the predictive performance of several algorithms. Our findings show that traditional statistical methods are more robust than state-of-the-art deep learning algorithms, except when the transformation effect is highly disruptive. Furthermore, we found no significant differences in the robustness of the algorithms when applying specific reconciliation methods, such as MinT. RHiOTS provides researchers with a comprehensive tool for understanding the nuanced behavior of forecasting algorithms, offering a more reliable basis for selecting the most appropriate method for a given problem. Luis Roque, Carlos Soares, Luís Torgo |
KDD | 2 |
| 2023 | GASTeN: Generative Adversarial Stress Test Networks
Carlos Soares, André Restivo, Luís F. Teixeira 0001 |
IDA | 2 |
| 2022 | Density Estimation in High-Dimensional Spaces: A Multivariate Histogram Approach
Pedro Strecht, João Mendes-Moreira 0001, Carlos Soares |
ADMA (2) | 3 |
| 2022 | On Usefulness of Outlier Elimination in Classification Tasks
Dusan Hetlerovic, Lubos Popelínský, Pavel Brazdil, Carlos Soares, Fernando Freitas |
IDA | 4 |
| 2022 | A case study comparing machine learning with statistical methods for time series forecasting: size matters
Vítor Cerqueira, Luís Torgo, Carlos Soares |
J. Intell. Inf. Syst. | 3 |
| 2021 | Promoting Fairness through Hyperparameter OptimizationabstractConsiderable research effort has been guided towards algorithmic fairness but real-world adoption of bias reduction techniques is still scarce. Existing methods are either metric-or model-specific, require access to sensitive attributes at inference time, or carry high development or deployment costs. This work explores the unfairness that emerges when optimizing ML models solely for predictive performance, and how to mitigate it with a simple and easily deployed intervention: fairness-aware hyperparameter optimization (HO). We propose and evaluate fairness-aware variants of three popular HO algorithms: Fair Random Search, Fair TPE, and Fairband. We validate our approach on a real-world bank account opening fraud case-study, as well as on three datasets from the fairness literature. Results show that, without extra training cost, it is feasible to find models with 111% mean fairness increase and just 6% decrease in performance when compared with fairness-blind HO.1 André F. Cruz, Pedro Saleiro, Catarina G. Belém, Carlos Soares, Pedro Bizarro |
ICDM | 4 |
| 2021 | Micro-MetaStream: Algorithm selection for time-changing data
André Luis Debiaso Rossi, Carlos Soares, Bruno Feres de Souza, André C. P. L. F. de Carvalho |
Inf. Sci. | 2 |
| 2018 | Analysing the Footprint of Classifiers in Overlapped and Imbalanced Contexts
Marta Mercier, Miriam Seoane Santos, Pedro H. Abreu, Carlos Soares, Jastin Pompeu Soares, João A. M. Santos |
IDA | 4 |
| 2018 | Constructive Aggregation and Its Application to Forecasting with Dynamic Ensembles
Vítor Cerqueira, Fábio Pinto, Luís Torgo, Carlos Soares, Nuno Moniz |
ECML/PKDD (1) | 4 |
| 2018 | CF4CF: recommending collaborative filtering algorithms using collaborative filteringabstractAs Collaborative Filtering becomes increasingly important in both academia and industry recommendation solutions, it also becomes imperative to study the algorithm selection task in this domain. This problem aims at finding automatic solutions which enable the selection of the best algorithms for a new problem, without performing full-fledged training and validation procedures. Existing work in this area includes several approaches using Metalearning, which relate the characteristics of the problem domain with the performance of the algorithms. This study explores an alternative approach to deal with this problem. Since, in essence, the algorithm selection problem is a recommendation problem, we investigate the use of Collaborative Filtering algorithms to select Collaborative Filtering algorithms. The proposed approach integrates subsampling landmarkers, a data characterization approach commonly used in Metalearning, with a Collaborative Filtering methodology, named CF4CF. The predictive performance obtained by CF4CF using benchmark recommendation datasets was similar or superior to that obtained with Metalearning. Tiago Cunha 0001, Carlos Soares, André C. P. L. F. de Carvalho |
RecSys | 2 |
| 2018 | Metalearning and Recommender Systems: A literature review and empirical study on the algorithm selection problem for Collaborative FilteringabstractThe problem of information overload motivated the appearance of Recommender Systems.From the several open problems in this area, the decision of which is the best recommendation algorithm for a specific problem is one of the most important and less studied.The current trend to solve this problem is the experimental evaluation of several recommendation algorithms in a handful of datasets.However, these studies require an extensive amount of computational resources, particularly processing time.To avoid these drawbacks, researchers have investigated the use of Metalearning to select the best recommendation algorithms in different scopes.Such studies allow to understand the relationships between data characteristics and the relative performance of recommendation algorithms, which can be used to select the best algorithm(s) for a new problem.The contributions of this study are two-fold: 1) to identify and discuss the key concepts of algorithm selection for recommendation algorithms via a systematic literature review and 2) to perform an experimental study on the Metalearning approaches reviewed in order to identify the most promising concepts for automatic selection of recommendation algorithms. Tiago Cunha 0001, Carlos Soares, André C. P. L. F. de Carvalho |
Inf. Sci. | 2 |
| 2017 | Arbitrated Ensemble for Time Series Forecasting
Vítor Cerqueira, Luís Torgo, Fábio Pinto, Carlos Soares |
ECML/PKDD (2) | 4 |
| 2017 | Metalearning for Context-aware Filtering: Selection of Tensor Factorization AlgorithmsabstractThis work addresses the problem of selecting Tensor Factorization algorithms for the Context-aware Filtering recommendation task using a metalearning approach. The most important challenge of applying metalearning on new problems is the development of useful measures able to characterize the data, i.e. metafeatures. We propose an extensive and exhaustive set of metafeatures to characterize Context-aware Filtering recommendation task. These metafeatures take advantage of the tensor's hierarchical structure via slice operations. The algorithm selection task is addressed as a Label Ranking problem, which ranks the Tensor Factorization algorithms according to their expected performance, rather than simply selecting the algorithm that is expected to obtain the best performance. A comprehensive experimental work is conducted on both levels, baselevel and metalevel (Tensor Factorization and Label Ranking, respectively). The results show that the proposed metafeatures lead to metamodels that tend to rank Tensor Factorization algorithms accurately and that the selected algorithms present high recommendation performance. Tiago Cunha 0001, Carlos Soares, André C. P. L. F. de Carvalho |
RecSys | 2 |
| 2017 | RELink: A Research Framework and Test Collection for Entity-Relationship RetrievalabstractImprovements of entity-relationship (E-R) search techniques have been hampered by a lack of test collections, particularly for complex queries involving multiple entities and relationships. In this paper we describe a method for generating E-R test queries to support comprehensive E-R search experiments. Queries and relevance judgments are created from content that exists in a tabular form where columns represent entity types and the table structure implies one or more relationships among the entities. Editorial work involves creating natural language queries based on relationships represented by the entries in the table. We have publicly released the RELink test collection comprising 600 queries and relevance judgments obtained from a sample of Wikipedia List-of-lists-of-lists tables. The latter comprise tuples of entities that are extracted from columns and labelled by corresponding entity types and relationships they represent. In order to facilitate research in complex E-R retrieval, we have created and released as open source the RELink Framework that includes Apache Lucene indexing and search specifically tailored to E-R retrieval. RELink includes entity and relationship indexing based on the ClueWeb-09-B Web collection with FACC1 text span annotations linked to Wikipedia entities. With ready to use search resources and a comprehensive test collection, we support community in pursuing E-R research at scale. Pedro Saleiro, Natasa Milic-Frayling, Eduarda Mendes Rodrigues, Carlos Soares |
SIGIR | 4 |
| 2016 | TimeMachine: Entity-Centric Search and Visualization of News Archives
Pedro Saleiro, Jorge Teixeira, Carlos Soares, Eugénio Oliveira |
ECIR | 3 |
| 2016 | Combining Boosted Trees with Metafeature Engineering for Predictive Maintenance
Vítor Cerqueira, Fábio Pinto, Cláudio Rebelo de Sá, Carlos Soares |
IDA | 4 |
| 2016 | Learning from the News: Predicting Entity Popularity on Twitter
Pedro Saleiro, Carlos Soares |
IDA | 2 |
| 2016 | Towards Automatic Generation of Metafeatures
Fábio Pinto, Carlos Soares, João Mendes-Moreira 0001 |
PAKDD (1) | 2 |
| 2016 | Selecting Collaborative Filtering Algorithms Using Metalearning
Tiago Cunha 0001, Carlos Soares, André C. P. L. F. de Carvalho |
ECML/PKDD (2) | 2 |
| 2016 | CHADE: Metalearning with Classifier Chains for Dynamic Combination of Classifiers
Fábio Pinto, Carlos Soares, João Mendes-Moreira 0001 |
ECML/PKDD (1) | 2 |
| 2016 | Entropy-based discretization methods for ranking data
Cláudio Rebelo de Sá, Carlos Soares, Arno J. Knobbe |
Inf. Sci. | 2 |
| 2015 | Using Metalearning for Prediction of Taxi Trip Duration Using Different Granularity Levels
Mohammad Nozari Zarmehri, Carlos Soares |
IDA | 2 |
| 2014 | TweeProfiles: Detection of Spatio-temporal Patterns on Twitter
Tiago Cunha 0001, Carlos Soares, Eduarda Mendes Rodrigues |
ADMA | 2 |
| 2014 | An Empirical Methodology to Analyze the Behavior of Bagging
Fábio Pinto, Carlos Soares, João Mendes-Moreira 0001 |
ADMA | 2 |
| 2014 | Merging Decision Trees: A Case Study in Predicting Student Performance
Pedro Strecht, João Mendes-Moreira 0001, Carlos Soares |
ADMA | 3 |
| 2013 | Space Allocation in the Retail Industry: A Decision Support System Integrating Evolutionary Algorithms and Regression Models
Fábio Pinto, Carlos Soares |
ECML/PKDD (3) | 2 |
| 2013 | Dimensions as Virtual Items: Improving the predictive ability of top-N recommender systems
Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Inf. Process. Manag. | 3 |
| 2012 | Integrating Data Mining and Optimization Techniques on Surgery Scheduling
Carlos Gomes, Bernardo Almada-Lobo, José Luís Cabral de Moura Borges, Carlos Soares |
ADMA | 4 |
| 2012 | Finding Interesting Contexts for Explaining Deviations in Bus Trip Duration Using Distribution Rules
Alípio Mário Jorge, João Mendes-Moreira 0001, Jorge Freire de Sousa, Carlos Soares, Paulo J. Azevedo |
IDA | 4 |
| 2011 | Mining Association Rules for Label Ranking
Cláudio Rebelo de Sá, Carlos Soares, Alípio Mário Jorge, Paulo J. Azevedo, Joaquim Pinto da Costa |
PAKDD (2) | 2 |
| 2011 | Exploiting Additional Dimensions as Virtual Items on Top-N Recommender SystemsabstractTraditionally, recommender systems for the web deal with applications that have two dimensions, users and items. Based on access data that relate these dimensions, a recommendation model can be built and used to identify a set of N items that will be of interest to a certain user. In this paper we propose a multidimensional approach, called DaVI (Dimensions as Virtual Items), that enables the use of common two-dimensional top-N recommender algorithms for the generation of recommendations using additional dimensions (e.g., contextual or background information). We empirically evaluate our approach with two different top-N recommender algorithms, Item-based Collaborative Filtering and Association Rules based, on two real world data sets. The empirical results demonstrate that DaVI enables the application of existing two-dimensional recommendation algorithms to exploit the useful information in multidimensional data. Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 3 |
| 2009 | The Effect of Varying Parameters and Focusing on Bus Travel Time Prediction
João Mendes-Moreira 0001, Carlos Soares, Alípio Mário Jorge, Jorge Freire de Sousa |
PAKDD | 2 |
| 2009 | UCI++: Improved Support for Algorithm Selection Using Datasetoids
Carlos Soares |
PAKDD | 1 |
| 2008 | The Impact of Contextual Information on the Accuracy of Existing Recommender Systems for Web PersonalizationabstractTraditionally, recommender systems for the Web deal with applications that have two types of entities/dimensions, users and items. With these dimensions, a recommendation model can be built and used to identify a set of N items that will be of interest to a certain user. In this paper we propose a direct method that enriches the information in the access logs with new dimensions. We empirically test this method with two recommender systems, an item-based collaborative filtering technique and association rules, on three data sets. Our results show that while collaborative filtering is not able to take advantage of the new dimensions added, association rules are capable of profiting from our direct method. Marcos Aurélio Domingues, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 3 |
| 2006 | Personalization of E-newsletters Based on Web Log Analysis and ClusteringabstractWe present a methodology for the personalization of e-newsletters based on the analysis of user access logs. To approach the problem we have used clustering on the set of users, described by their Web access patterns. Our work is evaluated using a case study with real data from e-newsletters sent by mail to users of a Web portal, and can be adapted to similar situations. Positive results were obtained, indicating that the methodology is able to automatically select contents for a personalized e-newsletter Carla Carvalho, Alípio Mário Jorge, Carlos Soares |
Web Intelligence | 3 |
| 2006 | Factor Analysis to Support the Visualization and Interpretation of Clusters of Portal UsersabstractClusterings based on many variables are difficult to visualize and interpret. We present a methodology based on factor analysis (FA) which can be used for that purpose. FA generates a small set of variables which encode most of the information in the original variables. We apply the methodology to segment the users of a Web portal, using access log data. It not only makes it simpler to visualize and understand the clusters which are obtained on the original variables but it also helps the analyst in selecting some of the original variables for further analysis of those clusters Carmen Rebelo, Pedro Quelhas Brito, Carlos Soares, Alípio Mário Jorge |
Web Intelligence | 3 |
| 2000 | A Comparison of Ranking Methods for Classification Algorithm Selection
Pavel Brazdil, Carlos Soares |
ECML | 2 |
| 2000 | Zoomed Ranking: Selection of Classification Algorithms Based on Relevant Performance Information
Carlos Soares, Pavel Brazdil |
PKDD | 1 |