Daniel Truhn

dblp:59/5522 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-9605-0728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Beyond benchmarks: Towards robust artificial intelligence bone segmentation in socio-technical systems
abstract
Despite the advances in automated medical image segmentation, AI models still underperform in various clinical settings, posing challenges for integration into real-world workflows. In this pre-registered prospective multicenter evaluation, we analyzed 20 state-of-the-art mandibular segmentation models across 19,218 segmentations of 1,000 clinically resampled CT/CBCT scans. Our results suggest that for a given model, segmentation accuracy can vary by up to 25% in Dice score as socio-technical factors such as voxel size, bone orientation, and patient conditions (e.g., osteosynthesis or pathology) shift from favorable to adverse. Higher sharpness, isotropic smaller voxels, and neutral orientation significantly improved results, while metallic osteosynthesis and anatomical complexity led to significant degradation. Our findings challenge the common view of AI models as “plug-and-play” tools and suggest evidence-based optimization recommendations for both clinicians and developers. This will in turn boost the integration of AI segmentation tools in routine healthcare.
Kunpeng Xie, Lennart Johannes Gruber, Martin Crampen, Elias Tappeiner, Maxime Gillot, Jan Schepers, Jiangchang Xu, Tobias Pankert, Michel Beyer, Negar Shahamiri, Reinier ten Brink, Gauthier Dot, Charlotte Weschke, Niels van Nistelrooij, Pieter-Jan Verhelst, Zhibin Xu, Jonas Bienzeisler, Ashkan Rashad, Tabea Flügge, Ross Cotton, Shankeeth Vinayahalingam, Robert R. Ilesan, Stefan Raith, Dennis Madsen, Constantin Seibold, Tong Xi 0001, Stefaan Bergé, Sven Nebelung, Oldrich Kodym, Osku Sundqvist, Florian M. Thieringer, Hans Lamecker, Antoine Coppens, Thomas Potrusil, Joep Kraeima, Max J. H. Witjes, Guomin Wu, Xiaojun Chen 0003, Adriaan Lambrechts, Stefan Zachow, Alexander Hermans, Daniel Truhn, Victor Alves, Jan Egger, Rainer Röhrig, Frank Hölzle, Behrus Hinrichs-Puladi
Expert Syst. Appl.46
2025 LLM Agents Making Agent Tools
abstract
Georg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic, Jakob Nikolas Kather. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Georg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic, Jakob Nikolas Kather
ACL (1)3
2025 Abnormality-Driven Representation Learning for Radiology Imaging
abstract
Radiology deep learning pipelines predominantly employ end-to-end 3D networks based on models pre-trained on other tasks, which are then fine-tuned on the task at hand. In contrast, adjacent medical fields such as pathology, which focus on 2D images, have effectively adopted task-agnostic foundational models based on self-supervised learning (SSL), combined with weakly-supervised deep learning (DL). However, the field of radiology still lacks task-agnostic representation models due to the computational and data demands of 3D imaging and the anatomical complexity inherent to radiology scans. To address this gap, we propose Clear, a framework for 3D radiology images that uses extracted embeddings from 2D slices along with attention-based aggregation to efficiently predict clinical endpoints. As part of this framework, we introduce Lecl, a novel approach to obtain visual representations driven by abnormalities in 2D axial slices across different locations of the CT scans. Specifically, we trained single-domain contrastive learning approaches using three different architectures: Vision Transformers, Vision State Space Models and Gated Convolutional Neural Networks. We evaluate our approach across three clinical tasks: tumor lesion location, lung disease detection, and patient staging, benchmarking against four state-of-the-art foundation models, including BiomedCLIP. Our findings demonstrate that Clear, using representations learned through Lecl, outperforms existing foundation models, while being substantially more compute- and data-efficient. The code is available at https://github.com/KatherLab/CLEAR .
Marta Ligero, Tim Lenz, Georg Wölflein, Omar S. M. El Nahhas, Daniel Truhn, Jakob Nikolas Kather
MICCAI (4)5
2025 Evaluating the effectiveness of biomedical fine-tuning for large language models on clinical tasks
abstract
OBJECTIVES: Large language models (LLMs) have shown potential in biomedical applications, leading to efforts to fine-tune them on domain-specific data. However, the effectiveness of this approach remains unclear. This study aims to critically evaluate the performance of biomedically fine-tuned LLMs against their general-purpose counterparts across a range of clinical tasks. MATERIALS AND METHODS: We evaluated the performance of biomedically fine-tuned LLMs against their general-purpose counterparts on clinical case challenges from NEJM and JAMA, and on multiple clinical tasks, such as information extraction, document summarization and clinical coding. We used a diverse set of benchmarks specifically chosen to be outside the likely fine-tuning datasets of biomedical models, ensuring a fair assessment of generalization capabilities. RESULTS: Biomedical LLMs generally underperformed compared to general-purpose models, especially on tasks not focused on probing medical knowledge. While on the case challenges, larger biomedical and general-purpose models showed similar performance (eg, OpenBioLLM-70B: 66.4% vs Llama-3-70B-Instruct: 65% on JAMA), smaller biomedical models showed more pronounced underperformance (OpenBioLLM-8B: 30% vs Llama-3-8B-Instruct: 64.3% on NEJM). Similar trends appeared across CLUE benchmarks, with general-purpose models often achieving higher scores in text generation, question answering, and coding. Notably, biomedical LLMs also showed a higher tendency to hallucinate. DISCUSSION: Our findings challenge the assumption that biomedical fine-tuning inherently improves LLM performance, as general-purpose models consistently performed better on unseen medical tasks. Retrieval-augmented generation may offer a more effective strategy for clinical adaptation. CONCLUSION: Fine-tuning LLMs on biomedical data may not yield the anticipated benefits. Alternative approaches, such as retrieval augmentation, should be further explored for effective and reliable clinical integration of LLMs.
Felix J. Dorfner, Amin Dada, Felix Busch, Marcus R. Makowski, Daniel Truhn, Jens Kleesiek, Madhumita Sushil, Lisa Adams, Keno Bressem
J. Am. Medical Informatics Assoc.6
2024 On Instabilities of Unsupervised Denoising Diffusion Models in Magnetic Resonance Imaging Reconstruction
Sven Nebelung, Firas Khader, Jakob Nikolas Kather, Daniel Truhn
MICCAI (7)5
2024 Joint Multi-task Learning Improves Weakly-Supervised Biomarker Prediction in Computational Pathology
Omar S. M. El Nahhas, Georg Wölflein, Marta Ligero, Tim Lenz, Marko van Treeck, Firas Khader, Daniel Truhn, Jakob Nikolas Kather
MICCAI (4)7
2024 Encrypted federated learning for secure decentralized collaboration in cancer image analysis
abstract
Artificial intelligence (AI) has a multitude of applications in cancer research and oncology. However, the training of AI systems is impeded by the limited availability of large datasets due to data protection requirements and other regulatory obstacles. Federated and swarm learning represent possible solutions to this problem by collaboratively training AI models while avoiding data transfer. However, in these decentralized methods, weight updates are still transferred to the aggregation server for merging the models. This leaves the possibility for a breach of data privacy, for example by model inversion or membership inference attacks by untrusted servers. Somewhat-homomorphically-encrypted federated learning (SHEFL) is a solution to this problem because only encrypted weights are transferred, and model updates are performed in the encrypted space. Here, we demonstrate the first successful implementation of SHEFL in a range of clinically relevant tasks in cancer image analysis on multicentric datasets in radiology and histopathology. We show that SHEFL enables the training of AI models which outperform locally trained models and perform on par with models which are centrally trained. In the future, SHEFL can enable multiple institutions to co-train AI models without forsaking data governance and without ever transmitting any decryptable data to untrusted servers.
Daniel Truhn, Soroosh Tayebi Arasteh, Oliver Lester Saldanha, Gustav Mueller-Franzes, Firas Khader, Philip Quirke, Nicholas P. West, Richard Gray 0005, Gordon G. A. Hutchins, Jacqueline A. James, Maurice B. Loughrey, Manuel Salto-Tellez, Hermann Brenner, Alexander Brobeil, Tanwei Yuan, Jenny Chang-Claude, Michael Hoffmeister, Sebastian Foersch, Sebastian Keil, Maximilian Schulze-Hagen, Peter Isfort, Philipp Bruners, Georgios Kaissis, Christiane Kuhl, Sven Nebelung, Jakob Nikolas Kather
Medical Image Anal.1
2024 AIROGS: Artificial Intelligence for Robust Glaucoma Screening Challenge
abstract
The early detection of glaucoma is essential in preventing visual impairment. Artificial intelligence (AI) can be used to analyze color fundus photographs (CFPs) in a cost-effective manner, making glaucoma screening more accessible. While AI models for glaucoma screening from CFPs have shown promising results in laboratory settings, their performance decreases significantly in real-world scenarios due to the presence of out-of-distribution and low-quality images. To address this issue, we propose the Artificial Intelligence for Robust Glaucoma Screening (AIROGS) challenge. This challenge includes a large dataset of around 113,000 images from about 60,000 patients and 500 different screening centers, and encourages the development of algorithms that are robust to ungradable and unexpected input data. We evaluated solutions from 14 teams in this paper and found that the best teams performed similarly to a set of 20 expert ophthalmologists and optometrists. The highest-scoring team achieved an area under the receiver operating characteristic curve of 0.99 (95% CI: 0.98-0.99) for detecting ungradable images on-the-fly. Additionally, many of the algorithms showed robust performance when tested on three other publicly available datasets. These results demonstrate the feasibility of robust AI-enabled glaucoma screening.
Coen de Vente, Koen A. Vermeer, Nicolas Jaccard, He Wang 0016, Hongyi Sun, Firas Khader, Daniel Truhn, Temirgali Aimyshev, Yerkebulan Zhanibekuly, Tien-Dung Le, Adrian Galdran, Miguel Ángel González Ballester, Gustavo Carneiro 0001, Devika R. G., Hrishikesh Panikkasseril Sethumadhavan, Densen Puthussery, Hong Liu 0007, Zekang Yang, Satoshi Kondo, Satoshi Kasai, Ashritha Durvasula, Jónathan Heras, Miguel Ángel Zapata, Teresa Araujo, Guilherme Aresta, Hrvoje Bogunovic, Mustafa Arikan, Yeong Chan Lee, Hyun Bin Cho, Yoon Ho Choi, Abdul Qayyum 0002, Muhammad Imran Razzak, Bram van Ginneken, Hans G. Lemij, Clara I. Sánchez
IEEE Trans. Medical Imaging7
2022 Benchmarking weakly-supervised deep learning pipelines for whole slide classification in computational pathology
Narmin Ghaffari Laleh, Hannah Sophie Muti, Chiara Maria Lavinia Loeffler, Amelie Echle, Oliver Lester Saldanha, Faisal Mahmood 0001, Ming Y. Lu, Christian Trautwein, Rupert Langer, Bastian Dislich, Roman David Bülow, Heike Irmgard Grabsch, Hermann Brenner, Jenny Chang-Claude, Elizabeth Alwers, Titus J. Brinker, Firas Khader, Daniel Truhn, Nadine T. Gaisa, Peter Boor, Michael Hoffmeister, Volkmar Schulz, Jakob Nikolas Kather
Medical Image Anal.18
2022 Erratum to 'Benchmarking weakly-supervised deep learning pipelines for whole slide classification in computational pathology' Medical Image Analysis, Volume 79, July 2022, 102474
Narmin Ghaffari Laleh, Hannah Sophie Muti, Chiara Maria Lavinia Loeffler, Amelie Echle, Oliver Lester Saldanha, Faisal Mahmood 0001, Ming Y. Lu, Christian Trautwein, Rupert Langer, Bastian Dislich, Roman David Bülow, Heike Irmgard Grabsch, Hermann Brenner, Jenny Chang-Claude, Elizabeth Alwers, Titus J. Brinker, Firas Khader, Daniel Truhn, Nadine T. Gaisa, Peter Boor, Michael Hoffmeister, Volkmar Schulz, Jakob Nikolas Kather
Medical Image Anal.18
2019 Multi Scale Curriculum CNN for Context-Aware Breast MRI Malignancy Classification
Christoph Haarburger, Michael Baumgartner 0001, Daniel Truhn, Mirjam Broeckmann, Hannah Schneider, Simone Schrading, Christiane Kuhl, Dorit Merhof
MICCAI (4)3