VLDB 2026 Research / reviewers in the wild / expert
José L. R. Sousa
dblp:201/6093
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-9570-6054ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explainable AI for Allergy Diagnosis: A CACTUS Framework to Address Doctor Variability
Luca Gherardini, Paulina Tworek, Maja Szczypka, Yousef Khan, Marek Mikolajczyk, Roman Lewandowski, José L. R. Sousa |
AIME (2) | 7 |
| 2025 | Machine Learning-Based Decision Support for Allergy Diagnosis: Real-World Implementation in a Hospital Setting
Paulina Tworek, Maja Szczypka, Julia Kahan, Marek Mikolajczyk, Roman Lewandowski, José L. R. Sousa |
AIME (1) | 6 |
| 2025 | Stability of Machine Learning Predictive Features Under Limited DataabstractIn a world where Machine Learning (ML) is increasingly used to make predictions about critical events, such as health outcomes, it is crucial to ensure decision-makers have access to explainable, consistent, and relevant predictive features. ML prediction relies on perfect data with similar distributions for testing and validation. These results are compared with those of humans, who use more noisy and limited data. Human predictions overcome those limitations by learning from abstractions. This paper addresses these issues by conducting experiments comparing traditional machine learning methods and a previously proposed method that uses data abstractions to learn predictive feature significance. The results indicate that the previously proposed descriptive ML approach maintains higher classification accuracy and ensures the stability of feature selection as data incompleteness increases, becoming valuable under limited data scenarios. It demonstrates the possibility of developing ML capable of automatic decision-making. Karol Capala, Paulina Tworek, José L. R. Sousa |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | CACTUS: A Comprehensive Abstraction and Classification Tool for Uncovering StructuresabstractThe availability of large datasets is providing the impetus for driving many current artificial intelligent developments. However, specific challenges arise in developing solutions that exploit small datasets, mainly due to practical and cost-effective deployment issues, as well as the opacity of deep learning models. To address this, the Comprehensive Abstraction and Classification Tool for Uncovering Structures (CACTUS) is presented as a means of improving secure analytics by effectively employing explainable artificial intelligence. CACTUS achieves this by providing additional support for categorical attributes, preserving their original meaning, optimising memory usage, and speeding up the computation through parallelisation. It exposes to the user the frequency of the attributes in each class and ranks them by their discriminative power. Performance is assessed by applying it to various domains, including Wisconsin Diagnostic Breast Cancer, Thyroid0387, Mushroom, Cleveland Heart Disease, and Adult Income datasets. Luca Gherardini, Varun Ravi Varma, Karol Capala, Roger F. Woods, José L. R. Sousa |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | SaNDA: A small and iNcomplete dataset analyserabstractIn personalised health, small datasets with missing data are quite common. Current Machine Learning methods are unable to process such datasets in a meaningful way due to the huge data volume requirement. To address this problem, we propose a new Small and iNcomplete Dataset Analyser (SaNDA) to process such datasets in a meaningful way. Due to the characteristics of these datasets and the criticality of the domain, an explainable method is mandatory to facilitate decision-making interpretation. Thus, SaNDA prioritises explainability over efficiency by design. We evaluated our proposal against Random Forest as a baseline for explainable methods, and against gcForest as state-of-the-art for small datasets. We observed that our proposal outperforms Random Forest when there is more missing data and/or lower number of entries in the dataset, obtaining less favourable results over larger, well-curated datasets. It is also preferable than gcForest due to its explainability and privacy protection capabilities. Given the difficulties in obtaining complete, reliable data in the healthcare field, we consider that our proposal could be useful for practitioners. Alfredo Ibias, Varun Ravi Varma, Karol Capala, Luca Gherardini, José L. R. Sousa |
Inf. Sci. | 5 |