VLDB 2026 Research / reviewers in the wild / expert
Karol Capala
dblp:348/1245
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-8864-0760ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 87% Machine learning and data management · 13% | |
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Stability of Machine Learning Predictive Features Under Limited Data · IEEE Trans. Knowl. Data Eng. 2025 |
Data mining › predictive modeling
classification |
0.9 | 1 | 2025 | Stability of Machine Learning Predictive Features Under Limited Data · IEEE Trans. Knowl. Data Eng. 2025 |
Data mining › dimensionality reduction
feature selection |
0.9 | 1 | 2025 | Stability of Machine Learning Predictive Features Under Limited Data · IEEE Trans. Knowl. Data Eng. 2025 |
Methods — techniques the papers use, named apart from their topics
descriptive machine learning · 1.7data abstraction · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Stability of Machine Learning Predictive Features Under Limited DataabstractIn a world where Machine Learning (ML) is increasingly used to make predictions about critical events, such as health outcomes, it is crucial to ensure decision-makers have access to explainable, consistent, and relevant predictive features. ML prediction relies on perfect data with similar distributions for testing and validation. These results are compared with those of humans, who use more noisy and limited data. Human predictions overcome those limitations by learning from abstractions. This paper addresses these issues by conducting experiments comparing traditional machine learning methods and a previously proposed method that uses data abstractions to learn predictive feature significance. The results indicate that the previously proposed descriptive ML approach maintains higher classification accuracy and ensures the stability of feature selection as data incompleteness increases, becoming valuable under limited data scenarios. It demonstrates the possibility of developing ML capable of automatic decision-making. Karol Capala, Paulina Tworek, José L. R. Sousa |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | CACTUS: A Comprehensive Abstraction and Classification Tool for Uncovering StructuresabstractThe availability of large datasets is providing the impetus for driving many current artificial intelligent developments. However, specific challenges arise in developing solutions that exploit small datasets, mainly due to practical and cost-effective deployment issues, as well as the opacity of deep learning models. To address this, the Comprehensive Abstraction and Classification Tool for Uncovering Structures (CACTUS) is presented as a means of improving secure analytics by effectively employing explainable artificial intelligence. CACTUS achieves this by providing additional support for categorical attributes, preserving their original meaning, optimising memory usage, and speeding up the computation through parallelisation. It exposes to the user the frequency of the attributes in each class and ranks them by their discriminative power. Performance is assessed by applying it to various domains, including Wisconsin Diagnostic Breast Cancer, Thyroid0387, Mushroom, Cleveland Heart Disease, and Adult Income datasets. Luca Gherardini, Varun Ravi Varma, Karol Capala, Roger F. Woods, José L. R. Sousa |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2023 | SaNDA: A small and iNcomplete dataset analyserabstractIn personalised health, small datasets with missing data are quite common. Current Machine Learning methods are unable to process such datasets in a meaningful way due to the huge data volume requirement. To address this problem, we propose a new Small and iNcomplete Dataset Analyser (SaNDA) to process such datasets in a meaningful way. Due to the characteristics of these datasets and the criticality of the domain, an explainable method is mandatory to facilitate decision-making interpretation. Thus, SaNDA prioritises explainability over efficiency by design. We evaluated our proposal against Random Forest as a baseline for explainable methods, and against gcForest as state-of-the-art for small datasets. We observed that our proposal outperforms Random Forest when there is more missing data and/or lower number of entries in the dataset, obtaining less favourable results over larger, well-curated datasets. It is also preferable than gcForest due to its explainability and privacy protection capabilities. Given the difficulties in obtaining complete, reliable data in the healthcare field, we consider that our proposal could be useful for practitioners. Alfredo Ibias, Varun Ravi Varma, Karol Capala, Luca Gherardini, José L. R. Sousa |
Inf. Sci. | 3 |