Nils Daniel Forkert

dblp:87/1218 · also Nils D. Forkert · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0003-2556-3224ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 18 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
YearPublicationVenuePosition
2026 A novel gradient inversion attack framework to investigate privacy vulnerabilities during retinal image-based federated learning
abstract
Machine learning models trained on retinal images have shown great potential in diagnosing various diseases. However, effectively training these models, especially in resource-limited regions, is often impeded by a lack of diverse data. Federated learning (FL) offers a solution to this problem by utilizing distributed data across a network of clients to enhance the training dataset volume and diversity. Nonetheless, significant privacy concerns have been raised for this approach, notably due to gradient inversion attacks that could expose private patient data used during FL training. Therefore, it is crucial to assess the vulnerability of FL models to such attacks because privacy breaches may discourage data sharing, potentially impacting the models' generalizability and clinical relevance. To tackle this issue, we introduce a novel framework to evaluate the vulnerability of federated deep learning models trained using retinal images to gradient inversion attacks. Importantly, we demonstrate how publicly available data can be used to enhance the quality of reconstructed images through an innovative image-to-image translation technique. The effectiveness of the proposed method was measured by evaluating the similarity between real fundus images and the corresponding reconstructed images using three different convolutional neural network architectures: ResNet-18, VGG-16, and DenseNet-121. Experimental results for the task of retinal age prediction demonstrate that, across all models, over 92 % of the participants in the training set could be identified from their reconstructed retinal vessel structure alone. Furthermore, even with the implementation of differential privacy countermeasures, we show that substantial information can still be extracted from the reconstructed images. Therefore, this work underscores the urgent need for improved defensive strategies to safeguard patient privacy during federated learning.
Christopher Nielsen, Matthias Wilms, Nils Daniel Forkert
Medical Image Anal.3
2025 Synthetic Ground Truth Counterfactuals for Comprehensive Evaluation of Causal Generative Models in Medical Imaging
Emma A. M. Stanley, Vibujithan Vigneshwaran, Erik Y. Ohara, Finn G. Vamosi, Nils Daniel Forkert, Matthias Wilms
MICCAI (8)5
2025 Accounting for population structure in deep learning models for genomic analysis
abstract
BACKGROUND: Deep learning methods are becoming increasingly popular for genotype analyses in recent years. In conventional genomic analyses, it is important to account for confounders to avoid biasing the results. Genetic relatedness is one of the most common confounders in conventional genomic analyses and there is a general consensus that it should be considered in the analysis to prevent distant levels of common ancestry from affecting the identification of causal variants. In contrast, genetic relatedness is not considered or ignored in many of the recently published deep learning models. OBJECTIVE: This study investigates whether the omission of genetic relatedness in deep learning models, common in recent literature, introduces confounding effects similar to those observed in conventional genomic analyses, particularly due to ancestry-related variants. METHODS: We developed and used a deep learning model to perform classifications based on single nucleotide polymorphism data from simulated and real-world datasets to examine whether population structure is confounding the model and potentially causing shortcut learning. RESULTS: The results of this study suggest that population structure may not significantly affect the performance of the deep learning model. However, explainable AI revealed notable differences in the focus between the confounded and unconfounded models when examining SNP feature importance. CONCLUSION: While population structure may not heavily affect model performance, it is important to reduce the models' capabilities of shortcut learning when designing deep learning models for analyzing genomic datasets, by using ancestry-related variants over potentially relevant biomarkers of the disease or disorder in question. The code used to perform these analyses can be found at: https://github.com/notTrivial/populationStructure.
Gabrielle Dagasso, Matthias Wilms, Raissa Souza, Nils Daniel Forkert
J. Biomed. Informatics4
2025 A cross-attention-based deep learning approach for predicting functional stroke outcomes using 4D CTP imaging and clinical metadata
abstract
Acute ischemic stroke (AIS) remains a global health challenge, leading to long-term functional disabilities without timely intervention. Spatio-temporal (4D) Computed Tomography Perfusion (CTP) imaging is crucial for diagnosing and treating AIS due to its ability to rapidly assess the ischemic core and penumbra. Although traditionally used to assess acute tissue status in clinical settings, 4D CTP has also been explored in research for predicting stroke tissue outcomes. However, its potential for predicting functional outcomes, especially in combination with clinical metadata, remains unexplored. Thus, this work aims to develop and evaluate a novel multimodal deep learning model for predicting functional outcomes (specifically, 90-day modified Rankin Scale) in AIS patients by combining 4D CTP and clinical metadata. To achieve this, an intermediate fusion strategy with a cross-attention mechanism is introduced to enable a selective focus on the most relevant features and patterns from both modalities. Evaluated on a dataset comprising 70 AIS patients who underwent endovascular mechanical thrombectomy, the proposed model achieves an accuracy (ACC) of 0.77, outperforming conventional late fusion strategies (ACC = 0.73) and unimodal models based on either 4D CTP (ACC = 0.61) or clinical metadata (ACC = 0.71). The results demonstrate the superior capability of the proposed model to leverage complex inter-modal relationships, emphasizing the value of advanced multimodal fusion techniques for predicting functional stroke outcomes. • Integration of 4D CTP and clinical metadata to predict clinical stroke outcomes. • Novel multimodal fusion strategy using a cross-attention mechanism. • Using both data modalities leads to better results than single-modality models. • Proposed multimodal model outperforms conventional late fusion approaches.
Kimberly Amador, Noah Pinel, Anthony J. Winder, Jens Fiehler, Matthias Wilms, Nils Daniel Forkert
Medical Image Anal.6
2025 Structural network measures reveal the emergence of heavy-tailed degree distributions in lottery ticket multilayer perceptrons
abstract
Artificial neural networks (ANNs) were originally modeled after their biological counterparts, but have since conceptually diverged in many ways. The resulting network architectures are not well understood, and furthermore, we lack the quantitative tools to characterize their structures. Network science provides an ideal mathematical framework with which to characterize systems of interacting components, and has transformed our understanding across many domains, including the mammalian brain. Yet, little has been done to bring network science to ANNs. In this work, we propose tools that leverage and adapt network science methods to measure both global- and local-level characteristics of ANNs. Specifically, we focus on the structures of efficient multilayer perceptrons as a case study, which are sparse and systematically pruned such that they share many characteristics with real-world networks. We use adapted network science metrics to show that the pruning process leads to the emergence of a spanning subnetwork (lottery ticket multilayer perceptrons) with complex architecture. This complex network exhibits global and local characteristics, including heavy-tailed nodal degree distributions and dominant weighted pathways, that mirror patterns observed in human neuronal connectivity. Furthermore, alterations in network metrics precede catastrophic decay in performance as the network is heavily pruned. This network science-driven approach to the analysis of artificial neural networks serves as a valuable tool to establish and improve biological fidelity, increase the interpretability, and assess the performance of artificial neural networks. Significance Statement Artificial neural network architectures have become increasingly complex, often diverging from their biological counterparts in many ways. To design plausible "brain-like" architectures, whether to advance neuroscience research or to improve explainability, it is essential that these networks optimally resemble their biological counterparts. Network science tools offer valuable information about interconnected systems, including the brain, but have not attracted much attention for analyzing artificial neural networks. Here, we present the significance of our work: •We adapt network science tools to analyze the structural characteristics of artificial neural networks. •We demonstrate that organizational patterns similar to those observed in the mammalian brain emerge through the pruning process alone. The convergence on these complex network features in both artificial neural networks and biological brain networks is compelling evidence for their optimality in information processing capabilities. •Our approach is a significant first step towards a network science-based understanding of artificial neural networks, and has the potential to shed light on the biological fidelity of artificial neural networks.
Chris Kang, Jasmine A. Moore, Samuel Robertson, Matthias Wilms, Emma K. Towlson, Nils Daniel Forkert
Neural Networks6
2024 MOX-NET: Multi-stage deep hybrid feature fusion and selection framework for monkeypox classification
Sarmad Maqsood, Robertas Damasevicius, Sana Shahid, Nils Daniel Forkert
Expert Syst. Appl.4
2024 Foundation model-driven distributed learning for enhanced retinal age prediction
abstract
OBJECTIVES: The retinal age gap (RAG) is emerging as a potential biomarker for various diseases of the human body, yet its utility depends on machine learning models capable of accurately predicting biological retinal age from fundus images. However, training generalizable models is hindered by potential shortages of diverse training data. To overcome these obstacles, this work develops a novel and computationally efficient distributed learning framework for retinal age prediction. MATERIALS AND METHODS: The proposed framework employs a memory-efficient 8-bit quantized version of RETFound, a cutting-edge foundation model for retinal image analysis, to extract features from fundus images. These features are then used to train an efficient linear regression head model for predicting retinal age. The framework explores federated learning (FL) as well as traveling model (TM) approaches for distributed training of the linear regression head. To evaluate this framework, we simulate a client network using fundus image data from the UK Biobank. Additionally, data from patients with type 1 diabetes from the UK Biobank and the Brazilian Multilabel Ophthalmological Dataset (BRSET) were utilized to explore the clinical utility of the developed methods. RESULTS: Our findings reveal that the developed distributed learning framework achieves retinal age prediction performance on par with centralized methods, with FL and TM providing similar performance (mean absolute error of 3.57 ± 0.18 years for centralized learning, 3.60 ± 0.16 years for TM, and 3.63 ± 0.19 years for FL). Notably, the TM was found to converge with fewer local updates than FL. Moreover, patients with type 1 diabetes exhibited significantly higher RAG values than healthy controls in all models, for both the UK Biobank and BRSET datasets (P < .001). DISCUSSION: The high computational and memory efficiency of the developed distributed learning framework makes it well suited for resource-constrained environments. CONCLUSION: The capacity of this framework to integrate data from underrepresented populations for training of retinal age prediction models could significantly enhance the accessibility of the RAG as an important disease biomarker.
Christopher Nielsen, Raissa Souza, Matthias Wilms, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.4
2024 Towards objective and systematic evaluation of bias in artificial intelligence for medical imaging
abstract
OBJECTIVE: Artificial intelligence (AI) models trained using medical images for clinical tasks often exhibit bias in the form of subgroup performance disparities. However, since not all sources of bias in real-world medical imaging data are easily identifiable, it is challenging to comprehensively assess their impacts. In this article, we introduce an analysis framework for systematically and objectively investigating the impact of biases in medical images on AI models. MATERIALS AND METHODS: Our framework utilizes synthetic neuroimages with known disease effects and sources of bias. We evaluated the impact of bias effects and the efficacy of 3 bias mitigation strategies in counterfactual data scenarios on a convolutional neural network (CNN) classifier. RESULTS: The analysis revealed that training a CNN model on the datasets containing bias effects resulted in expected subgroup performance disparities. Moreover, reweighing was the most successful bias mitigation strategy for this setup. Finally, we demonstrated that explainable AI methods can aid in investigating the manifestation of bias in the model using this framework. DISCUSSION: The value of this framework is showcased in our findings on the impact of bias scenarios and efficacy of bias mitigation in a deep learning model pipeline. This systematic analysis can be easily expanded to conduct further controlled in silico trials in other investigations of bias in medical imaging AI. CONCLUSION: Our novel methodology for objectively studying bias in medical imaging AI can help support the development of clinical decision-support tools that are robust and responsible.
Emma A. M. Stanley, Raissa Souza, Anthony J. Winder, Vedant Gulve, Kimberly Amador, Matthias Wilms, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.7
2024 Providing clinical context to the spatio-temporal analysis of 4D CT perfusion to predict acute ischemic stroke lesion outcomes
Kimberly Amador, Alejandro Gutierrez, Anthony J. Winder, Jens Fiehler, Matthias Wilms, Nils Daniel Forkert
J. Biomed. Informatics6
2024 Identifying Biases in a Multicenter MRI Database for Parkinson's Disease Classification: Is the Disease Classifier a Secret Site Classifier?
abstract
Sharing multicenter imaging datasets can be advantageous to increase data diversity and size but may lead to spurious correlations between site-related biological and non-biological image features and target labels, which machine learning (ML) models may exploit as shortcuts. To date, studies analyzing how and if deep learning models may use such effects as a shortcut are scarce. Thus, the aim of this work was to investigate if site-related effects are encoded in the feature space of an established deep learning model designed for Parkinson's disease (PD) classification based on T1-weighted MRI datasets. Therefore, all layers of the PD classifier were frozen, except for the last layer of the network, which was replaced by a linear layer that was exclusively re-trained to predict three potential bias types (biological sex, scanner type, and originating site). Our findings based on a large database consisting of 1880 MRI scans collected across 41 centers show that the feature space of the established PD model (74% accuracy) can be used to classify sex (75% accuracy), scanner type (79% accuracy), and site location (71% accuracy) with high accuracies despite this information never being explicitly provided to the PD model during original training. Overall, the results of this study suggest that trained image-based classifiers may use unwanted shortcuts that are not meaningful for the actual clinical task at hand. This finding may explain why many image-based deep learning models do not perform well when applied to data from centers not contributing to the training set.
Raissa Souza, Anthony J. Winder, Emma A. M. Stanley, Vibujithan Vigneshwaran, Milton Camacho, Richard Camicioli, Oury Monchi, Matthias Wilms, Nils Daniel Forkert
IEEE J. Biomed. Health Informatics9
2023 A Flexible Framework for Simulating and Evaluating Biases in Deep Learning-Based Medical Image Analysis
Emma A. M. Stanley, Matthias Wilms, Nils Daniel Forkert
MICCAI (2)3
2023 Image-encoded biological and non-biological variables may be used as shortcuts in deep learning models trained on multisite neuroimaging data
abstract
OBJECTIVE: This work investigates if deep learning (DL) models can classify originating site locations directly from magnetic resonance imaging (MRI) scans with and without correction for intensity differences. MATERIAL AND METHODS: A large database of 1880 T1-weighted MRI scans collected across 41 sites originally for Parkinson's disease (PD) classification was used to classify sites in this study. Forty-six percent of the datasets are from PD patients, while 54% are from healthy participants. After preprocessing the T1-weighted scans, 2 additional data types were generated: intensity-harmonized T1-weighted scans and log-Jacobian deformation maps resulting from nonlinear atlas registration. Corresponding DL models were trained to classify sites for each data type. Additionally, logistic regression models were used to investigate the contribution of biological (age, sex, disease status) and non-biological (scanner type) variables to the models' decision. RESULTS: A comparison of the 3 different types of data revealed that DL models trained using T1-weighted and intensity-harmonized T1-weighted scans can classify sites with an accuracy of 85%, while the model using log-Jacobian deformation maps achieved a site classification accuracy of 54%. Disease status and scanner type were found to be significant confounders. DISCUSSION: Our results demonstrate that MRI scans encode relevant site-specific information that models could use as shortcuts that cannot be removed using simple intensity harmonization methods. CONCLUSION: The ability of DL models to exploit site-specific biases as shortcuts raises concerns about their reliability, generalization, and deployability in clinical settings.
Raissa Souza, Matthias Wilms, Milton Camacho, G. Bruce Pike, Richard Camicioli, Oury Monchi, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.7
2022 A morphometrics approach for inclusion of localised characteristics from medical imaging studies into genome-wide association studies
abstract
Medical images, such as magnetic resonance or computed tomography, are increasingly being used to investigate the genetic architecture of neurological diseases like Alzheimer’s disease, or psychiatric disorders like attention-deficit hyperactivity disorder. The quantified global or regional brain imaging measures are commonly known as imaging-specific or -derived phenotypes (IDPs) when conducting genotype-phenotype association studies. Inclusion of whole medical images rather than derived tabular data as IDPs has been done by either a voxelwise approach or a global approach of whole medical images via principal component analysis. Limitations with multiple testing and inability to isolate high variation regions within the principal components arise with either of these approaches. This work proposes a principal component analysis-like localised approach of dimensionality reduction using diffeomorphic morphometry allowing for the selection of distances to model more regional effects. The main benefit of the proposed method is that it can can reduce the dimensionality of the problem considerably in comparison to the medical image’s variability it is describing while grouping spatial information potentially lost in dimensionality reduction techniques like principal component analyses. Moreover, the approach not only allows to include locality in the analysis but can also be used as a generative model to explore the morphometric changes across an axis of particular components of interest. To demonstrate the feasibility of this pipeline for inclusion in a multivariate genome-wide association study, it was applied to 1,359 subjects from the Adolescent Brain Cognitive Development Study for traits related to attention-deficit disorder. The results show that the proposed method can identify more specific morphometric features associated with genome regions.
Gabrielle Dagasso, Matthias Wilms, Nils Daniel Forkert
BIBM3
2022 Simulating progressive neurodegeneration in silico with deep artificial neural networks
Anup Tuladhar, Jasmine A. Moore, Zahinoor Ismail, Nils Daniel Forkert
CogSci4
2022 Hybrid Spatio-Temporal Transformer Network for Predicting Ischemic Stroke Lesion Outcomes from 4D CT Perfusion Imaging
Kimberly Amador, Anthony J. Winder, Jens Fiehler, Matthias Wilms, Nils Daniel Forkert
MICCAI (3)5
2022 Detecting 3D syndromic faces as outliers using unsupervised normalizing flow models
Jordan J. Bannister, Matthias Wilms, J. David Aponte, David C. Katz, Ophir D. Klein, Francois P. J. Bernier, Richard A. Spritz, Benedikt Hallgrímsson, Nils Daniel Forkert
Artif. Intell. Medicine9
2022 An analysis of the effects of limited training data in distributed learning scenarios for brain age prediction
abstract
OBJECTIVE: Distributed learning avoids problems associated with central data collection by training models locally at each site. This can be achieved by federated learning (FL) aggregating multiple models that were trained in parallel or training a single model visiting sites sequentially, the traveling model (TM). While both approaches have been applied to medical imaging tasks, their performance in limited local data scenarios remains unknown. In this study, we specifically analyze FL and TM performances when very small sample sizes are available per site. MATERIALS AND METHODS: 2025 T1-weighted magnetic resonance imaging scans were used to investigate the effect of sample sizes on FL and TM for brain age prediction. We evaluated models across 18 scenarios varying the number of samples per site (1, 2, 5, 10, and 20) and the number of training rounds (20, 40, and 200). RESULTS: Our results demonstrate that the TM outperforms FL, for every sample size examined. In the extreme case when each site provided only one sample, FL achieved a mean absolute error (MAE) of 18.9 ± 0.13 years, while the TM achieved a MAE of 6.21 ± 0.50 years, comparable to central learning (MAE = 5.99 years). DISCUSSION: Although FL is more commonly used, our study demonstrates that TM is the best implementation for small sample sizes. CONCLUSION: The TM offers new opportunities to apply machine learning models in rare diseases and pediatric research but also allows even small hospitals to contribute small datasets.
Raissa Souza, Pauline Mouches, Matthias Wilms, Anup Tuladhar, Sönke Langner, Nils Daniel Forkert
J. Am. Medical Informatics Assoc.6
2022 Predicting treatment-specific lesion outcomes in acute ischemic stroke from 4D CT perfusion imaging using spatio-temporal convolutional neural networks
Kimberly Amador, Matthias Wilms, Anthony J. Winder, Jens Fiehler, Nils Daniel Forkert
Medical Image Anal.5
2022 A Deep Invertible 3-D Facial Shape Model for Interpretable Genetic Syndrome Diagnosis
abstract
One of the primary difficulties in treating patients with genetic syndromes is diagnosing their condition. Many syndromes are associated with characteristic facial features that can be imaged and utilized by computer-assisted diagnosis systems. In this work, we develop a novel 3D facial surface modeling approach with the objective of maximizing diagnostic model interpretability within a flexible deep learning framework. Therefore, an invertible normalizing flow architecture is introduced to enable both inferential and generative tasks in a unified and efficient manner. The proposed model can be used (1) to infer syndrome diagnosis and other demographic variables given a 3D facial surface scan and (2) to explain model inferences to non-technical users via multiple interpretability mechanisms. The model was trained and evaluated on more than 4700 facial surface scans from subjects with 47 different syndromes. For the challenging task of predicting syndrome diagnosis given a new 3D facial surface scan, age, and sex of a subject, the model achieves a competitive overall top-1 accuracy of 71%, and a mean sensitivity of 43% across all syndrome classes. We believe that invertible models such as the one presented in this work can achieve competitive inferential performance while greatly increasing model interpretability in the domain of medical diagnosis.
Jordan J. Bannister, Matthias Wilms, J. David Aponte, David C. Katz, Ophir D. Klein, Francois P. J. Bernier, Richard A. Spritz, Benedikt Hallgrímsson, Nils Daniel Forkert
IEEE J. Biomed. Health Informatics9
2022 Invertible Modeling of Bidirectional Relationships in Neuroimaging With Normalizing Flows: Application to Brain Aging
abstract
Many machine learning tasks in neuroimaging aim at modeling complex relationships between a brain's morphology as seen in structural MR images and clinical scores and variables of interest. A frequently modeled process is healthy brain aging for which many image-based brain age estimation or age-conditioned brain morphology template generation approaches exist. While age estimation is a regression task, template generation is related to generative modeling. Both tasks can be seen as inverse directions of the same relationship between brain morphology and age. However, this view is rarely exploited and most existing approaches train separate models for each direction. In this paper, we propose a novel bidirectional approach that unifies score regression and generative morphology modeling and we use it to build a bidirectional brain aging model. We achieve this by defining an invertible normalizing flow architecture that learns a probability distribution of 3D brain morphology conditioned on age. The use of full 3D brain data is achieved by deriving a manifold-constrained formulation that models morphology variations within a low-dimensional subspace of diffeomorphic transformations. This modeling idea is evaluated on a database of MR scans of more than 5000 subjects. The evaluation results show that our bidirectional brain aging model (1) accurately estimates brain age, (2) is able to visually explain its decisions through attribution maps and counterfactuals, (3) generates realistic age-specific brain morphology templates, (4) supports the analysis of morphological variations, and (5) can be utilized for subject-specific brain aging simulation.
Matthias Wilms, Jordan J. Bannister, Pauline Mouches, M. Ethan MacDonald, Deepthi Rajashekar, Sönke Langner, Nils Daniel Forkert
IEEE Trans. Medical Imaging7
2021 Bone and joint enhancement filtering: Application to proximal femur segmentation from uncalibrated computed tomography datasets
abstract
Methods for reliable femur segmentation enable the execution of quality retrospective studies and building of robust screening tools for bone and joint disease. An enhance-and-segment pipeline is proposed for proximal femur segmentation from computed tomography datasets. The filter is based on a scale-space model of cortical bone with properties including edge localization, invariance to density calibration, rotation invariance, and stability to noise. The filter is integrated with a graph cut segmentation technique guided through user provided sparse labels for rapid segmentation. Analysis is performed on 20 independent femurs. Rater proximal femur segmentation agreement was 0.21 mm (average surface distance), 0.98 (Dice similarity coefficient), and 2.34 mm (Hausdorff distance). Manual segmentation added considerable variability to measured failure load and volume (CVRMS > 5%) but not density. The proposed algorithm considerably improved inter-rater reproducibility for all three outcomes (CVRMS < 0.5%). The algorithm localized the periosteal surface accurately compared to manual segmentation but with a slight bias towards a smaller volume. Hessian-based filtering and graph cut segmentation localizes the periosteal surface of the proximal femur with comparable accuracy and improved precision compared to manual segmentation.
Bryce A. Besler, Andrew S. Michalski, Michael T. Kuczynski, Aleena Abid, Nils Daniel Forkert, Steven K. Boyd
Medical Image Anal.5
2020 A Kernelized Multi-level Localization Method for Flexible Shape Modeling with Few Training Data
Matthias Wilms, Jan Ehrhardt, Nils Daniel Forkert
MICCAI (4)3
2020 Building machine learning models without sharing patient data: A simulation-based analysis of distributed learning by ensembling
Anup Tuladhar, Sascha Gill, Zahinoor Ismail, Nils Daniel Forkert
J. Biomed. Informatics4
2019 A methodology for generating four-dimensional arterial spin labeling MR angiography virtual phantoms
Renzo Phellan, Thomas Lindner 0002, Michael Helle, Alexandre X. Falcão, Thomas W. Okell, Nils Daniel Forkert
Medical Image Anal.6