Raissa Souza

dblp:325/2935 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-7455-3383ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Accounting for population structure in deep learning models for genomic analysis
abstract
BACKGROUND: Deep learning methods are becoming increasingly popular for genotype analyses in recent years. In conventional genomic analyses, it is important to account for confounders to avoid biasing the results. Genetic relatedness is one of the most common confounders in conventional genomic analyses and there is a general consensus that it should be considered in the analysis to prevent distant levels of common ancestry from affecting the identification of causal variants. In contrast, genetic relatedness is not considered or ignored in many of the recently published deep learning models. OBJECTIVE: This study investigates whether the omission of genetic relatedness in deep learning models, common in recent literature, introduces confounding effects similar to those observed in conventional genomic analyses, particularly due to ancestry-related variants. METHODS: We developed and used a deep learning model to perform classifications based on single nucleotide polymorphism data from simulated and real-world datasets to examine whether population structure is confounding the model and potentially causing shortcut learning. RESULTS: The results of this study suggest that population structure may not significantly affect the performance of the deep learning model. However, explainable AI revealed notable differences in the focus between the confounded and unconfounded models when examining SNP feature importance. CONCLUSION: While population structure may not heavily affect model performance, it is important to reduce the models' capabilities of shortcut learning when designing deep learning models for analyzing genomic datasets, by using ancestry-related variants over potentially relevant biomarkers of the disease or disorder in question. The code used to perform these analyses can be found at: https://github.com/notTrivial/populationStructure.
Gabrielle Dagasso, Matthias Wilms, Raissa Souza, Nils Daniel Forkert
J. Biomed. Informatics3
2024 Foundation model-driven distributed learning for enhanced retinal age prediction
abstract
OBJECTIVES: The retinal age gap (RAG) is emerging as a potential biomarker for various diseases of the human body, yet its utility depends on machine learning models capable of accurately predicting biological retinal age from fundus images. However, training generalizable models is hindered by potential shortages of diverse training data. To overcome these obstacles, this work develops a novel and computationally efficient distributed learning framework for retinal age prediction. MATERIALS AND METHODS: The proposed framework employs a memory-efficient 8-bit quantized version of RETFound, a cutting-edge foundation model for retinal image analysis, to extract features from fundus images. These features are then used to train an efficient linear regression head model for predicting retinal age. The framework explores federated learning (FL) as well as traveling model (TM) approaches for distributed training of the linear regression head. To evaluate this framework, we simulate a client network using fundus image data from the UK Biobank. Additionally, data from patients with type 1 diabetes from the UK Biobank and the Brazilian Multilabel Ophthalmological Dataset (BRSET) were utilized to explore the clinical utility of the developed methods. RESULTS: Our findings reveal that the developed distributed learning framework achieves retinal age prediction performance on par with centralized methods, with FL and TM providing similar performance (mean absolute error of 3.57 ± 0.18 years for centralized learning, 3.60 ± 0.16 years for TM, and 3.63 ± 0.19 years for FL). Notably, the TM was found to converge with fewer local updates than FL. Moreover, patients with type 1 diabetes exhibited significantly higher RAG values than healthy controls in all models, for both the UK Biobank and BRSET datasets (P < .001). DISCUSSION: The high computational and memory efficiency of the developed distributed learning framework makes it well suited for resource-constrained environments. CONCLUSION: The capacity of this framework to integrate data from underrepresented populations for training of retinal age prediction models could significantly enhance the accessibility of the RAG as an important disease biomarker.
Christopher Nielsen, Raissa Souza, Matthias Wilms, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.2
2024 Towards objective and systematic evaluation of bias in artificial intelligence for medical imaging
abstract
OBJECTIVE: Artificial intelligence (AI) models trained using medical images for clinical tasks often exhibit bias in the form of subgroup performance disparities. However, since not all sources of bias in real-world medical imaging data are easily identifiable, it is challenging to comprehensively assess their impacts. In this article, we introduce an analysis framework for systematically and objectively investigating the impact of biases in medical images on AI models. MATERIALS AND METHODS: Our framework utilizes synthetic neuroimages with known disease effects and sources of bias. We evaluated the impact of bias effects and the efficacy of 3 bias mitigation strategies in counterfactual data scenarios on a convolutional neural network (CNN) classifier. RESULTS: The analysis revealed that training a CNN model on the datasets containing bias effects resulted in expected subgroup performance disparities. Moreover, reweighing was the most successful bias mitigation strategy for this setup. Finally, we demonstrated that explainable AI methods can aid in investigating the manifestation of bias in the model using this framework. DISCUSSION: The value of this framework is showcased in our findings on the impact of bias scenarios and efficacy of bias mitigation in a deep learning model pipeline. This systematic analysis can be easily expanded to conduct further controlled in silico trials in other investigations of bias in medical imaging AI. CONCLUSION: Our novel methodology for objectively studying bias in medical imaging AI can help support the development of clinical decision-support tools that are robust and responsible.
Emma A. M. Stanley, Raissa Souza, Anthony J. Winder, Vedant Gulve, Kimberly Amador, Matthias Wilms, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.2
2024 Identifying Biases in a Multicenter MRI Database for Parkinson's Disease Classification: Is the Disease Classifier a Secret Site Classifier?
abstract
Sharing multicenter imaging datasets can be advantageous to increase data diversity and size but may lead to spurious correlations between site-related biological and non-biological image features and target labels, which machine learning (ML) models may exploit as shortcuts. To date, studies analyzing how and if deep learning models may use such effects as a shortcut are scarce. Thus, the aim of this work was to investigate if site-related effects are encoded in the feature space of an established deep learning model designed for Parkinson's disease (PD) classification based on T1-weighted MRI datasets. Therefore, all layers of the PD classifier were frozen, except for the last layer of the network, which was replaced by a linear layer that was exclusively re-trained to predict three potential bias types (biological sex, scanner type, and originating site). Our findings based on a large database consisting of 1880 MRI scans collected across 41 centers show that the feature space of the established PD model (74% accuracy) can be used to classify sex (75% accuracy), scanner type (79% accuracy), and site location (71% accuracy) with high accuracies despite this information never being explicitly provided to the PD model during original training. Overall, the results of this study suggest that trained image-based classifiers may use unwanted shortcuts that are not meaningful for the actual clinical task at hand. This finding may explain why many image-based deep learning models do not perform well when applied to data from centers not contributing to the training set.
Raissa Souza, Anthony J. Winder, Emma A. M. Stanley, Vibujithan Vigneshwaran, Milton Camacho, Richard Camicioli, Oury Monchi, Matthias Wilms, Nils Daniel Forkert
IEEE J. Biomed. Health Informatics1
2023 Image-encoded biological and non-biological variables may be used as shortcuts in deep learning models trained on multisite neuroimaging data
abstract
OBJECTIVE: This work investigates if deep learning (DL) models can classify originating site locations directly from magnetic resonance imaging (MRI) scans with and without correction for intensity differences. MATERIAL AND METHODS: A large database of 1880 T1-weighted MRI scans collected across 41 sites originally for Parkinson's disease (PD) classification was used to classify sites in this study. Forty-six percent of the datasets are from PD patients, while 54% are from healthy participants. After preprocessing the T1-weighted scans, 2 additional data types were generated: intensity-harmonized T1-weighted scans and log-Jacobian deformation maps resulting from nonlinear atlas registration. Corresponding DL models were trained to classify sites for each data type. Additionally, logistic regression models were used to investigate the contribution of biological (age, sex, disease status) and non-biological (scanner type) variables to the models' decision. RESULTS: A comparison of the 3 different types of data revealed that DL models trained using T1-weighted and intensity-harmonized T1-weighted scans can classify sites with an accuracy of 85%, while the model using log-Jacobian deformation maps achieved a site classification accuracy of 54%. Disease status and scanner type were found to be significant confounders. DISCUSSION: Our results demonstrate that MRI scans encode relevant site-specific information that models could use as shortcuts that cannot be removed using simple intensity harmonization methods. CONCLUSION: The ability of DL models to exploit site-specific biases as shortcuts raises concerns about their reliability, generalization, and deployability in clinical settings.
Raissa Souza, Matthias Wilms, Milton Camacho, G. Bruce Pike, Richard Camicioli, Oury Monchi, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.1
2022 An analysis of the effects of limited training data in distributed learning scenarios for brain age prediction
abstract
OBJECTIVE: Distributed learning avoids problems associated with central data collection by training models locally at each site. This can be achieved by federated learning (FL) aggregating multiple models that were trained in parallel or training a single model visiting sites sequentially, the traveling model (TM). While both approaches have been applied to medical imaging tasks, their performance in limited local data scenarios remains unknown. In this study, we specifically analyze FL and TM performances when very small sample sizes are available per site. MATERIALS AND METHODS: 2025 T1-weighted magnetic resonance imaging scans were used to investigate the effect of sample sizes on FL and TM for brain age prediction. We evaluated models across 18 scenarios varying the number of samples per site (1, 2, 5, 10, and 20) and the number of training rounds (20, 40, and 200). RESULTS: Our results demonstrate that the TM outperforms FL, for every sample size examined. In the extreme case when each site provided only one sample, FL achieved a mean absolute error (MAE) of 18.9 ± 0.13 years, while the TM achieved a MAE of 6.21 ± 0.50 years, comparable to central learning (MAE = 5.99 years). DISCUSSION: Although FL is more commonly used, our study demonstrates that TM is the best implementation for small sample sizes. CONCLUSION: The TM offers new opportunities to apply machine learning models in rare diseases and pediatric research but also allows even small hospitals to contribute small datasets.
Raissa Souza, Pauline Mouches, Matthias Wilms, Anup Tuladhar, Sönke Langner, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.1