VLDB 2026 Research / reviewers in the wild / expert
Klaus H. Maier-Hein
dblp:133/0183 · also Klaus H. Fritzsche, Klaus Hermann Maier-Hein, Klaus Maier-Hein
· DBLP profile ↗
69ranked-venue papers
1as first author
50since 2021 · last 2026
0000-0002-6626-2463ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 53 · 1 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 21 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond benchmarks of IUGC: Rethinking requirements of deep learning method for intrapartum ultrasound biometry from fetal ultrasound videos
Jieyun Bai, Yitong Tang, Zhuonan Liang, Jianan Fan, Lisa Mcguire, Jillian Clarke, Tom Weidong Cai, Jacqueline Spurway, Yubo Tang, Shiye Wang, Wenda Shen, Wangwang Yu, Philippe Zhang, Weili Jiang, Salem Muhsin Ali Binqahal Al Nasim, Arsen Abzhanov, Numan Saeed, Mohammad Yaqub, Zunhui Xia, Hongxing Li 0001, Libin Lan, Jayroop Ramesh, Valentin Bacher, Mark Eid, Hoda Kalabizadeh, Christian Rupprecht 0001, Ana I. L. Namburete, Pak-Hei Yeung, Madeleine K. Wyburd, Nicola K. Dinsdale, Assanali Serikbey, Jiankai Li, Sung-Liang Chen, Zicheng Hu, Nana Liu, Yian Deng, Wenfeng Zhang, Mai Tuyet Nhi, Gregor Koehler, Rapheal Stock, Klaus H. Maier-Hein, Marawan Elbatel, Xiaomeng Li 0001, Saad Slimani, Victor M. Campello, Benard Ohene Botwe, Isaac Khobo, Zhenyan Han, Hongying Hou, Di Qiu, Gongning Luo, Dong Ni 0001, Yaosheng Lu, Karim Lekadir, Shuo Li 0001 |
Medical Image Anal. | 47 |
| 2026 | Multi-structure segmentation in CBCT volumes: The ToothFairy2 challengeabstractCone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT. Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana |
Medical Image Anal. | 16 |
| 2026 | Multi-class segmentation of aortic branches and zones in computed tomography angiography: The AortaSeg24 challenge
Muhammad Imran 0013, Jonathan R. Krebs, Vishal Balaji Sivaraman, Amarjeet Kumar, Walker R. Ueland, Michael J. Fassler, Lisheng Wang, Maximilian Rokuss, Michael Baumgartner 0001, Yannick Kirchhof, Klaus H. Maier-Hein, Fabian Isensee, Shuolin Liu, Bong Thanh Nguyen, Dong-jin Shin, Park Ji-Woo, Matthew Choi, Kwang-Hyun Uhm, Sung-Jea Ko, Chanwoong Lee, Jaehee Chun, Yun Gu, Zhaohong Pan, Xiaokun Liang, Markus Tiefenthaler, Enrique Almar-Munoz, Matthias Schwab, Mikhail Kotyushev, Rostislav Epifanov, Marek Wodzinski, Henning Müller, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Zhiwei Wang 0002, Kaixiang Yang 0004, Jintao Ren, Stine Sofia Korreman, Yuchong Gao, Hongye Zeng, Jinghua Yue, Fugen Zhou, Alexander Cosman, Muxuan Liang, Gilbert R. Upchurch Jr., Yuyin Zhou, Michol A. Cooper, Wei Shao 0008 |
Medical Image Anal. | 16 |
| 2026 | Advances in automated fetal brain MRI segmentation and biometry: Insights from the FeTA 2024 challengeabstractAccurate fetal brain tissue segmentation and biometric measurement are essential for monitoring neurodevelopment and detecting abnormalities in utero. The Fetal Tissue Annotation (FeTA) Challenges have established robust multi-center benchmarks for evaluating state-of-the-art segmentation methods. This paper presents the results of the 2024 challenge edition, which introduced three key innovations. First, we introduced a topology-aware metric based on the Euler characteristic difference (ED) to overcome the performance plateau observed with traditional metrics like Dice or Hausdorff distance (HD), as the performance of the best models in segmentation surpassed the inter-rater variability. While the best teams reached similar scores in Dice (0.81-0.82) and HD95 (2.1-2.3 mm), ED provided greater discriminative power: the winning method achieved an ED of 20.9, representing roughly a 50% improvement over the second- and third-ranked teams despite comparable Dice scores. Second, we introduced a new 0.55T low-field MRI test set, which, when paired with high-quality super-resolution reconstruction, achieved the highest segmentation performance across all test cohorts (Dice=0.86, HD95=1.69, ED=6.26). This provides the first quantitative evidence that low-cost, low-field MRI can match or surpass high-field systems in automated fetal brain segmentation. Third, the new biometry estimation task exposed a clear performance gap: although the best model reached a mean average percentage error (MAPE) of 7.72%, most submissions failed to outperform a simple gestational-age-based linear regression model (MAPE=9.56%), and all remained above inter-rater variability with a MAPE of 5.38%. Finally, by analyzing the top-performing models from FeTA 2024 alongside those from previous challenge editions, we identify ensembles of 3D nnU-Net trained on both real and synthetic data with both image- and anatomy-level augmentations as the most effective approaches for fetal brain segmentation. Our quantitative analysis reveals that acquisition site, super-resolution strategy, and image quality are the primary sources of domain shift, informing recommendations to enhance the robustness and generalizability of automated fetal brain analysis methods. Vladyslav Zalevskyi, Thomas Sanchez, Misha P. T. Kaandorp, Margaux Roulet, Diego Fajardo-Rojas, Liu Li 0001, Jana Hutter, Hongwei Li 0004, Matthew J. Barkovich, Luca Wilhelmi, Aline Dändliker, Céline Steger, Mériam Koob, Yvan Gomez, Anton Jakovcic, Melita Klaic, Ana Adzic, Pavel Markovic, Gracia Grabaric, Milan Rados, Jordina Aviles Verdera, Gregor Kasprian, Gregor Dovjak, Raphael Gaubert-Rachmühl, Maurice Aschwanden, Davood Karimi, Denis Peruzzo, Tommaso Ciceri, Giorgio Longari, Rachika E. Hamadache, Amina Bouzid, Xavier Lladó, Simone Chiarella, Gerard Martí-Juan, Miguel Ángel González Ballester, Marco Castellaro, Marco Pinamonti, Valentina Visani, Robin Cremese, Keïn Sam, Fleur Gaudfernau, Param Ahir, Mehul Parikh, Maximilian Zenk, Michael Baumgartner 0001, Klaus H. Maier-Hein, Li Tianhong, Zhao Longfei, Domen Preloznik, Ziga Spiclin, Jae Won Choi, Guotai Wang, Lyuyang Tong, Bo Du 0001, Andrea Gondova, Sungmin You, Kiho Im, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, András Jakab, Roxane Licandro, Kelly Payette, Meritxell Bach Cuadra |
Medical Image Anal. | 48 |
| 2026 | ToothSeg: Robust Tooth Instance Segmentation and Numbering in CBCT Using Deep Learning and Self-CorrectionabstractAccurate interpretation of cone-beam computed tomography (CBCT) scans is critical for oral diagnosis and treatment planning. Existing methods for automated tooth segmentation in CBCT face challenges, such as difficulties in generalizing across imaging artifacts and anatomical variations, as well as requiring manual revisions in many cases. To address these limitations, this study introduces ToothSeg, a fully automated approach for tooth instance segmentation and numbering in CBCT using deep learning and self-correction. ToothSeg combines semantic and instance segmentation into a unified method where their respective strengths are complemented. In particular, self-correction is employed when combining the segmentations, resolving merged or split teeth and determining the optimal sequence of tooth numbers for each dental arch. We conducted a comprehensive evaluation using a diverse in-house dataset (n = 1282, 25+ devices) and the publicly available ToothFairy2 challenge dataset (n = 480, 1 device), including an ablation study, a comparison to state-of-the-art methods, and an analysis of challenging cases. Compared to an optimized semantic segmentation model, including instance segmentation and self-correction consistently improved tooth segmentation (True Positive Dice: 93.6% to 94.3%) and tooth detection and numbering (multiclass instance F1: 94.2% to 95.5%). Furthermore, ToothSeg outperformed the other methods on both datasets (True Positive Dice: $\geq$ +0.4%, multiclass instance F1: $\geq$ +1.8%), particularly for challenging cases. This study provides a promising approach for automated tooth segmentation and numbering in CBCT, which is significant for reducing manual workload and supporting scalable, data-driven research in oral and craniofacial health. Niels van Nistelrooij, Lars Krämer, Steven Kempers, Michel Beyer, Federico Bolelli, Tong Xi 0001, Stefaan Bergé, Max Heiland, Klaus H. Maier-Hein, Shankeeth Vinayahalingam, Fabian Isensee |
IEEE J. Biomed. Health Informatics | 9 |
| 2026 | Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-Ray: Summary of the PENGWIN 2024 ChallengeabstractThe segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and efficiently delineating the bone fragments remains a significant challenge due to complex anatomy and imaging limitations. The PENGWIN challenge, organized as a MICCAI 2024 satellite event, aimed to advance automated fracture segmentation by benchmarking state-of-the-art algorithms on these complex tasks. A diverse dataset of 150 CT scans was collected from multiple clinical centers, and a large set of simulated X-ray images was generated using the DeepDRR method. Final submissions from 16 teams worldwide were evaluated under a rigorous multi-metric testing scheme. The top-performing CT algorithm achieved an average fragment-wise intersection over union (IoU) of 0.930, demonstrating satisfactory accuracy. However, in the X-ray task, the best algorithm achieved an IoU of 0.774, which is promising but not yet sufficient for intra-operative decision-making, reflecting the inherent challenges of fragment overlap in projection imaging. Beyond the quantitative evaluation, the challenge revealed methodological diversity in algorithm design. Variations in instance representation, such as primary-secondary classification versus boundary-core separation, led to differing segmentation strategies. Despite promising results, the challenge also exposed inherent uncertainties in fragment definition, particularly in cases of incomplete fractures. These findings suggest that interactive segmentation approaches, integrating human decision-making with task-relevant information, may be essential for improving model reliability and clinical applicability. Yudi Sang, Yanzhen Liu, Sutuke Yibulayimu, Yunning Wang, Benjamin Killeen, Mingxu Liu, Ping-Cheng Ku, Ole Johannsen, Karol Gotkowski, Maximilian Zenk, Klaus H. Maier-Hein, Fabian Isensee, Peiyan Yue, Yi Wang 0031, Zhaohong Pan, Xiaokun Liang, Daiqi Liu, Fuxin Fan, Artur Jurgas, Andrzej Skalski, Szymon Plotka, Rafal Litka, Yingchun Song, Mathias Unberath, Mehran Armand, Dan Ruan, Shaohua Kevin Zhou, Qiyong Cao, Chunpeng Zhao, Xinbao Wu, Yu Wang 0083 |
IEEE Trans. Medical Imaging | 11 |
| 2025 | LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body ImagingabstractIn this work, we present LesionLocator, a framework for zero-shot longitudinal lesion tracking and segmentation in 3D medical imaging, establishing the first end-to-end model capable of 4D tracking with dense spatial prompts. Our model leverages an extensive dataset of 23,262 annotated medical scans, as well as synthesized longitudinal data across diverse lesion types. The diversity and scale of our dataset significantly enhances model generalizability to real-world medical imaging challenges and addresses key limitations in longitudinal data availability. LesionLocator outperforms all existing promptable models in lesion segmentation by nearly 10 dice points, reaching human-level performance, and achieves state-of-the-art results in lesion tracking, with superior lesion retrieval and segmentation accuracy. LesionLocator not only sets a new benchmark in universal promptable lesion segmentation and automated longitudinal lesion tracking but also provides the first open-access solution of its kind, releasing our synthetic 4D dataset and model to the community, empowering future advancements in medical imaging. Code is available at: www.github.com/MIC-DKFZ/LesionLocator Maximilian Rokuss, Yannick Kirchhoff, Seval Akbal, Balint Kovacs, Saikat Roy, Constantin Ulrich, Tassilo Wald, Lukas T. Rotkopf, Heinz-Peter Schlemmer, Klaus H. Maier-Hein |
CVPR | 10 |
| 2025 | Revisiting MAE Pre-training for 3D Medical Image SegmentationabstractSelf-Supervised Learning (SSL) presents an exciting opportunity to unlock the potential of vast, untapped clinical datasets, for various downstream applications that suffer from the scarcity of labeled data. While SSL has revolutionized fields like natural language processing and computer vision, its adoption in 3D medical image computing has been limited by three key pitfalls: Small pre-training dataset sizes, architectures inadequate for 3D medical image analysis, and insufficient evaluation practices. In this paper, we address these issues by i) leveraging a large-scale dataset of 39k 3D brain MRI volumes and ii) using a Residual Encoder U-Net architecture within the state-of-the-art nnU-Net framework. iii) A robust development framework, incorporating 5 development and 8 testing brain MRI segmentation datasets, allowed performance-driven design decisions to optimize the simple concept of Masked Auto Encoders (MAEs) for 3D CNNs. The resulting model not only surpasses previous SSL methods but also outperforms the strong nnU-Net baseline by an average of approximately 3 Dice points setting a new state-of-the-art. Our code and models are made available here. Tassilo Wald, Constantin Ulrich, Stanislav Lukyanenko, Andrei Goncharov, Alberto Paderno, Maximilian Miller, Leander Maerkisch, Paul F. Jaeger, Klaus H. Maier-Hein |
CVPR | 9 |
| 2025 | An OpenMind for 3D Medical Vision Self-supervised LearningabstractThe field of self-supervised learning (SSL) for 3D medical images lacks consistency and standardization. While many methods have been developed, it is impossible to identify the current state-of-the-art, due to i) varying and small pretraining datasets, ii) varying architectures, and iii) being evaluated on differing downstream datasets. In this paper, we bring clarity to this field and lay the foundation for further method advancements through three key contributions: We a) publish the largest publicly available pre-training dataset comprising 114k 3D brain MRI volumes, enabling all practitioners to pre-train on a large-scale dataset. We b) benchmark existing 3D self-supervised learning methods on this dataset for a state-of-the-art CNN and Transformer architecture, clarifying the state of 3D SSL pre-training. Among many findings, we show that pre-trained methods can exceed a strong from-scratch nnU-Net ResEnc-L baseline. Lastly, we c) publish the code of our pre-training and fine-tuning frameworks and provide the pre-trained models created during the benchmarking process to facilitate rapid adoption and reproduction. Tassilo Wald, Constantin Ulrich, Jonathan Suprijadi, Sebastian Ziegler, Michal Nohel, Robin Peretzke, Gregor Köhler, Klaus H. Maier-Hein |
ICCV | 8 |
| 2025 | ReSi: A Comprehensive Benchmark for Representational Similarity MeasuresabstractMeasuring the similarity of different representations of neural architectures is a fundamental task and an open research challenge for the machine learning community. This paper presents the first comprehensive benchmark for evaluating representational similarity measures based on well-defined groundings of similarity. The representational similarity (ReSi) benchmark consists of (i) six carefully designed tests for similarity measures, (ii) 24 similarity measures, (iii) 14 neural network architectures, and (iv) seven datasets, spanning over the graph, language, and vision domains. The benchmark opens up several important avenues of research on representational similarity that enable novel explorations and applications of neural architectures. We demonstrate the utility of the ReSi benchmark by conducting experiments on various neural network architectures, real world datasets and similarity measures. All components of the benchmark are publicly available and thereby facilitate systematic reproduction and production of research results. The benchmark is extensible, future research can build on and further expand it. We believe that the ReSi benchmark can serve as a sound platform catalyzing future research that aims to systematically evaluate existing and explore novel ways of comparing representations of neural architectures.
ReSi is available at https://github.com/mklabunde/resi. Max Klabunde, Tassilo Wald, Tobias Schumacher 0002, Klaus H. Maier-Hein, Markus Strohmaier, Florian Lemmerich |
ICLR | 4 |
| 2025 | The Missing Piece: A Case for Pre-training in 3D Medical Object Detection
Katharina Eckstein, Constantin Ulrich, Michael Baumgartner 0001, Jessica Kächele, Dimitrios Bounias, Tassilo Wald, Ralf Floca, Klaus H. Maier-Hein |
MICCAI (4) | 8 |
| 2025 | Revisiting 3D Medical Scribble Supervision: Benchmarking Beyond Cardiac Segmentation
Karol Gotkowski, Klaus H. Maier-Hein, Fabian Isensee |
MICCAI (8) | 2 |
| 2025 | Cross-Modality Supervised Prostate Segmentation on CBCT for Adaptive Radiotherapy
Balint Kovacs, Goran Stanic, Fabian Weykamp, Florian Ebert, Dimitrios Bounias, Bouchra Tawk, Martin Niklas, Jakob Liermann, Oliver Jäkel, Klaus H. Maier-Hein, Ralf Floca, Kristina Giske |
MICCAI (6) | 10 |
| 2025 | Enhancing Predictive Imaging Biomarker Discovery Through Treatment Effect AnalysisabstractIdentifying predictive covariates, which forecast individual treatment effectiveness, is crucial for decision-making across different disciplines such as personalized medicine. These covariates, referred to as biomarkers, are extracted from pretreatment data, often within randomized controlled trials, and should be distinguished from prognostic biomarkers, which are independent of treatment assignment. Our study focuses on discovering predictive imaging biomarkers, specific image features, by leveraging pretreatment images to uncover new causal relationships. Unlike laborintensive approaches relying on handcrafted features prone to bias, we present a novel task of directly learning predictive features from images. We propose an evaluation protocol to assess a model's ability to identify predictive imaging biomarkers and differentiate them from purely prognostic ones by employing statistical testing and a comprehensive analysis of image feature attribution. We explore the suitability of deep learning models originally developed for estimating the conditional average treatment effect (CATE) for this task, which have been assessed primarily for their precision of CATE estimation while overlooking the evaluation of imaging biomarker discovery. Our proof-of-concept analysis demonstrates the feasibility and potential of our approach in discovering and validating predictive imaging biomarkers from synthetic outcomes and real-world image datasets. Our code is available at https://github.com/MIC-DKFZ/predictive_image_biomarker_analysis. Shuhan Xiao, Lukas Klein, Jens Petersen, Philipp Vollmuth, Paul F. Jaeger, Klaus H. Maier-Hein |
WACV | 6 |
| 2025 | Including AI in diffusion-weighted breast MRI has potential to increase reader confidence and reduce workloadabstractOBJECTIVES: Breast diffusion-weighted imaging (DWI) has shown potential as a standalone imaging technique for certain indications, eg, supplemental screening of women with dense breasts. This study evaluates an artificial intelligence (AI)-powered computer-aided diagnosis (CAD) system for clinical interpretation and workload reduction in breast DWI. MATERIALS AND METHODS: This retrospective IRB-approved study included: n = 824 examinations for model development (2017-2020) and n = 235 for evaluation (01/2021-06/2021). Readings were performed by three readers using either the AI-CAD or manual readings. BI-RADS-like (Breast Imaging Reporting and Data System) classification was based on DWI. Histopathology served as ground truth. The model was nnDetection-based, trained using 5-fold cross-validation and ensembling. Statistical significance was determined using McNemar's test. Inter-rater agreement was calculated using Cohen's kappa. Model performance was calculated using the area under the receiver operating curve (AUC). RESULTS: The AI-augmented approach significantly reduced BI-RADS-like 3 calls in breast DWI by 29% (P =.019) and increased interrater agreement (0.57 ± 0.10 vs 0.49 ± 0.11), while preserving diagnostic accuracy. Two of the three readers detected more malignant lesions (63/69 vs 59/69 and 64/69 vs 62/69) with the AI-CAD. The AI model achieved an AUC of 0.78 (95% CI: [0.72, 0.85]; P <.001), which increased for women at screening age to 0.82 (95% CI: [0.73, 0.90]; P <.001), indicating a potential for workload reduction of 20.9% at 96% sensitivity. DISCUSSION AND CONCLUSION: Breast DWI might benefit from AI support. In our study, AI showed potential for reduction of BI-RADS-like 3 calls and increase of inter-rater agreement. However, given the limited study size, further research is needed. Dimitrios Bounias, Lina Simons, Michael Baumgartner 0001, Chris Ehring, Peter Neher, Lorenz A. Kapsner, Balint Kovacs, Ralf Floca, Paul F. Jaeger, Jessica Eberle, Dominique Hadler, Frederik B. Laun, Sabine Ohlmeyer, Lena Maier-Hein, Michael Uder, Evelyn Wenkel, Klaus H. Maier-Hein, Sebastian Bickelhaupt |
J. Am. Medical Informatics Assoc. | 17 |
| 2025 | Real-world federated learning in radiology: hurdles to overcome and benefits to gainabstractOBJECTIVE: Federated Learning (FL) enables collaborative model training while keeping data locally. Currently, most FL studies in radiology are conducted in simulated environments due to numerous hurdles impeding its translation into practice. The few existing real-world FL initiatives rarely communicate specific measures taken to overcome these hurdles. To bridge this significant knowledge gap, we propose a comprehensive guide for real-world FL in radiology. Minding efforts to implement real-world FL, there is a lack of comprehensive assessments comparing FL to less complex alternatives in challenging real-world settings, which we address through extensive benchmarking. MATERIALS AND METHODS: We developed our own FL infrastructure within the German Radiological Cooperative Network (RACOON) and demonstrated its functionality by training FL models on lung pathology segmentation tasks across six university hospitals. Insights gained while establishing our FL initiative and running the extensive benchmark experiments were compiled and categorized into the guide. RESULTS: The proposed guide outlines essential steps, identified hurdles, and implemented solutions for establishing successful FL initiatives conducting real-world experiments. Our experimental results prove the practical relevance of our guide and show that FL outperforms less complex alternatives in all evaluation scenarios. DISCUSSION AND CONCLUSION: Our findings justify the efforts required to translate FL into real-world applications by demonstrating advantageous performance over alternative approaches. Additionally, they emphasize the importance of strategic organization, robust management of distributed data and infrastructure in real-world settings. With the proposed guide, we are aiming to aid future FL researchers in circumventing pitfalls and accelerating translation of FL into radiological applications. Markus Bujotzek, Ünal Akünal, Stefan Denner, Peter Neher, Maximilian Zenk, Eric Frodl, Astha Jaiswal, Moon S. Kim 0002, Nicolai R. Krekiehn, Manuel Nickel, Richard Ruppel, Marcus Both, Felix Doellinger, Marcel Opitz, Thorsten Persigehl, Jens Kleesiek, Tobias Penzkofer, Klaus H. Maier-Hein, Andreas Bucher, Rickmer Braren |
J. Am. Medical Informatics Assoc. | 18 |
| 2025 | PSFHS challenge report: Pubic symphysis and fetal head segmentation from intrapartum ultrasound images
Jieyun Bai, Zhanhong Ou, Gregor Köhler, Raphael Stock, Klaus H. Maier-Hein, Marawan Elbatel, Robert Martí, Xiaomeng Li 0001, Yaoyang Qiu, Panjie Gou, Gongping Chen, Lei Zhao 0013, Jianxun Zhang 0002, Yu Dai 0002, Fangyijie Wang, Guénolé C. M. Silvestre, Kathleen M. Curran, Hongkun Sun, Pengzhou Cai, Libin Lan, Dong Ni 0001, Mei Zhong, Gaowen Chen, Víctor M. Campello, Yaosheng Lu, Karim Lekadir |
Medical Image Anal. | 6 |
| 2025 | Corrigendum to "PSFHS challenge report: pubic symphysis and fetal head segmentation from intrapartum ultrasound images" [Medical Image Analysis 99 (2025),103353]
Jieyun Bai, Zhanhong Ou, Gregor Köhler, Raphael Stock, Klaus H. Maier-Hein, Marawan Elbatel, Robert Martí, Xiaomeng Li 0001, Yaoyang Qiu, Panjie Gou, Gongping Chen, Lei Zhao 0013, Jianxun Zhang 0002, Yu Dai 0002, Fangyijie Wang, Guénolé C. M. Silvestre, Kathleen M. Curran, Hongkun Sun, Pengzhou Cai, Libin Lan, Dong Ni 0001, Mei Zhong, Gaowen Chen, Víctor M. Campello, Yaosheng Lu, Karim Lekadir |
Medical Image Anal. | 6 |
| 2025 | SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 25 |
| 2025 | Comparative benchmarking of failure detection methods in medical image segmentation: Unveiling the role of confidence aggregationabstractSemantic segmentation is an essential component of medical image analysis research, with recent deep learning algorithms offering out-of-the-box applicability across diverse datasets. Despite these advancements, segmentation failures remain a significant concern for real-world clinical applications, necessitating reliable detection mechanisms. This paper introduces a comprehensive benchmarking framework aimed at evaluating failure detection methodologies within medical image segmentation. Through our analysis, we identify the strengths and limitations of current failure detection metrics, advocating for the risk-coverage analysis as a holistic evaluation approach. Utilizing a collective dataset comprising five public 3D medical image collections, we assess the efficacy of various failure detection strategies under realistic test-time distribution shifts. Our findings highlight the importance of pixel confidence aggregation and we observe superior performance of the pairwise Dice score (Roy et al., 2019) between ensemble predictions, positioning it as a simple and robust baseline for failure detection in medical image segmentation. To promote ongoing research, we make the benchmarking framework available to the community. Maximilian Zenk, David Zimmerer, Fabian Isensee, Jeremias Traub, Tobias Norajitra, Paul F. Jaeger, Klaus H. Maier-Hein |
Medical Image Anal. | 7 |
| 2025 | Automated Ensemble Multimodal Machine Learning for HealthcareabstractThe application of machine learning in medicine and healthcare has led to the creation of numerous diagnostic and prognostic models. However, despite their success, current approaches generally issue predictions using data from a single modality. This stands in stark contrast with clinician decision-making which employs diverse information from multiple sources. While several multimodal machine learning approaches exist, significant challenges in developing multimodal systems remain that are hindering clinical adoption. In this paper, we introduce a multimodal framework, AutoPrognosis-M, that enables the integration of structured clinical (tabular) data and medical imaging using automated machine learning. AutoPrognosis-M incorporates 17 imaging models, including convolutional neural networks and vision transformers, and three distinct multimodal fusion strategies. In an illustrative application using a multimodal skin lesion dataset, we highlight the importance of multimodal machine learning and the power of combining multiple fusion strategies using ensemble learning. We have open-sourced our framework as a tool for the community and hope it will accelerate the uptake of multimodal machine learning in healthcare and spur further innovation. Fergus Imrie, Stefan Denner, Lucas S. Brunschwig, Klaus H. Maier-Hein, Mihaela van der Schaar |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Segmenting the Inferior Alveolar Canal in CBCTs Volumes: The ToothFairy ChallengeabstractIn recent years, several algorithms have been developed for the segmentation of the Inferior Alveolar Canal (IAC) in Cone-Beam Computed Tomography (CBCT) scans. However, the availability of public datasets in this domain is limited, resulting in a lack of comparative evaluation studies on a common benchmark. To address this scientific gap and encourage deep learning research in the field, the ToothFairy challenge was organized within the MICCAI 2023 conference. In this context, a public dataset was released to also serve as a benchmark for future research. The dataset comprises 443 CBCT scans, with voxel-level annotations of the IAC available for 153 of them, making it the largest publicly available dataset of its kind. The participants of the challenge were tasked with developing an algorithm to accurately identify the IAC using the 2D and 3D-annotated scans. This paper presents the details of the challenge and the contributions made by the most promising methods proposed by the participants. It represents the first comprehensive comparative evaluation of IAC segmentation methods on a common benchmark dataset, providing insights into the current state-of-the-art algorithms and outlining future research directions. Furthermore, to ensure reproducibility and promote future developments, an open-source repository that collects the implementations of the best submissions was released. Federico Bolelli, Luca Lumetti, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Arrigo Pellacani, Kevin Marchesini, Niels van Nistelrooij, Pieter van Lierop, Tong Xi 0001, Yusheng Liu 0001, Rui Xin 0003, Tao Yang 0037, Lisheng Wang, Haoshen Wang, Chenfan Xu, Zhiming Cui 0001, Marek Wodzinski, Henning Müller, Yannick Kirchhoff, Maximilian Rokuss, Klaus H. Maier-Hein, Jae-Hwan Han, Wan Kim, Hong-Gi Ahn, Tomasz Szczepanski, Michal K. Grzeszczyk, Przemyslaw Korzeniowski, Vicent Caselles, Xavier Paolo Burgos-Artizzu, Ferran Prados, Stefaan Bergé, Bram van Ginneken, Alexandre Anesi, Costantino Grana |
IEEE Trans. Medical Imaging | 21 |
| 2024 | Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures
Yannick Kirchhoff, Maximilian Rokuss, Saikat Roy, Balint Kovacs, Constantin Ulrich, Tassilo Wald, Maximilian Zenk, Philipp Vollmuth, Jens Kleesiek, Fabian Isensee, Klaus H. Maier-Hein |
ECCV (77) | 11 |
| 2024 | ValUES: A Framework for Systematic Validation of Uncertainty Estimation in Semantic SegmentationabstractUncertainty estimation is an essential and heavily-studied component for the reliable application of semantic segmentation methods. While various studies exist claiming methodological advances on the one hand, and successful application on the other hand, the field is currently hampered by a gap between theory and practice leaving fundamental questions unanswered: Can data-related and model-related uncertainty really be separated in practice? Which components of an uncertainty method are essential for real-world performance? Which uncertainty method works well for which application? In this work, we link this research gap to a lack of systematic and comprehensive evaluation of uncertainty methods. Specifically, we identify three key pitfalls in current literature and present an evaluation framework that bridges the research gap by providing 1) a controlled environment for studying data ambiguities as well as distribution shifts, 2) systematic ablations of relevant method components, and 3) test-beds for the five predominant uncertainty applications: OoD-detection, active learning, failure detection, calibration, and ambiguity modeling. Empirical results on simulated as well as real-world data demonstrate how the proposed framework is able to answer the predominant questions in the field revealing for instance that 1) separation of uncertainty types works on simulated data but does not necessarily translate to real-world data, 2) aggregation of scores is a crucial but currently neglected component of uncertainty methods, 3) While ensembles are performing most robustly across the different downstream tasks and settings, test-time augmentation often constitutes a light-weight alternative. Code is at: https://github.com/IML-DKFZ/values Kim-Celine Kahl, Carsten T. Lüth, Maximilian Zenk, Klaus H. Maier-Hein, Paul F. Jaeger |
ICLR | 4 |
| 2024 | nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation
Fabian Isensee, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger |
MICCAI (9) | 6 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 17 |
| 2024 | Overcoming Common Flaws in the Evaluation of Selective Classification SystemsabstractSelective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of these systems typically assumes fixed working points based on pre-defined rejection thresholds, methodological progress requires benchmarking the general performance of systems akin to the $\mathrm{AUROC}$ in standard classification. In this work, we define 5 requirements for multi-threshold metrics in selective classification regarding task alignment, interpretability, and flexibility, and show how current approaches fail to meet them. We propose the Area under the Generalized Risk Coverage curve ($\mathrm{AUGRC}$), which meets all requirements and can be directly interpreted as the average risk of undetected failures. We empirically demonstrate the relevance of $\mathrm{AUGRC}$ on a comprehensive benchmark spanning 6 data sets and 13 confidence scoring functions. We find that the proposed metric substantially changes metric rankings on 5 out of the 6 data sets. Jeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner 0001, Klaus H. Maier-Hein, Lena Maier-Hein, Paul F. Jaeger |
NeurIPS | 5 |
| 2024 | Decoupling Semantic Similarity from Spatial Alignment for Neural NetworksabstractWhat representation do deep neural networks learn? How similar are images to each other for neural networks? Despite the overwhelming success of deep learning methods key questions about their internal workings still remain largely unanswered, due to their internal high dimensionality and complexity. To address this, one approach is to measure the similarity of activation responses to various inputs.
Representational Similarity Matrices (RSMs) distill this similarity into scalar values for each input pair.
These matrices encapsulate the entire similarity structure of a system, indicating which input lead to similar responses.
While the similarity between images is ambiguous, we argue that the spatial location of semantic objects does neither influence human perception nor deep learning classifiers. Thus this should be reflected in the definition of similarity between image responses for computer vision systems. Revisiting the established similarity calculations for RSMs we expose their sensitivity to spatial alignment. In this paper we propose to solve this through _semantic RSMs_, which are invariant to spatial permutation. We measure semantic similarity between input responses by formulating it as a set-matching problem. Further, we quantify the superiority of _semantic_ RSMs over _spatio-semantic_ RSMs through image retrieval and by comparing the similarity between representations to the similarity between predicted class probabilities. Tassilo Wald, Constantin Ulrich, Priyank Jaini, Gregor Köhler, David Zimmerer, Stefan Denner, Fabian Isensee, Michael Baumgartner 0001, Klaus H. Maier-Hein |
NeurIPS | 9 |
| 2024 | RecycleNet: Latent Feature Recycling Leads to Iterative Decision RefinementabstractDespite the remarkable success of deep learning systems over the last decade, a key difference still remains between neural network and human decision-making: As humans, we can not only form a decision on the spot, but also ponder, revisiting an initial guess from different angles, distilling relevant information, arriving at a better decision. Here, we propose RecycleNet, a latent feature recycling method, instilling the pondering capability for neural networks to refine initial decisions over a number of recycling steps, where outputs are fed back into earlier network layers in an iterative fashion. This approach makes minimal assumptions about the neural network architecture and thus can be implemented in a wide variety of contexts. Using medical image segmentation as the evaluation environment, we show that latent feature recycling enables the network to iteratively refine initial predictions even beyond the iterations seen during training, converging towards an improved decision. We evaluate this across a variety of segmentation benchmarks and show consistent improvements even compared with top-performing segmentation methods. This allows trading increased computation time for improved performance, which can be beneficial, especially for safety-critical applications. Gregor Köhler, Tassilo Wald, Constantin Ulrich, David Zimmerer, Paul F. Jaeger, Jörg K. H. Franke, Simon Kohl, Fabian Isensee, Klaus H. Maier-Hein |
WACV | 9 |
| 2024 | Fair evaluation of federated learning algorithms for automated breast density classification: The results of the 2022 ACR-NCI-NVIDIA federated learning challenge
Kendall Schmidt, Ben Bearce, Ken Chang, Laura Coombs, Keyvan Farahani, Marawan Elbatel, Kaouther Mouheb, Robert Martí, Ya Zhang 0002, Yanfeng Wang 0001, Yaojun Hu, Haochao Ying, Yuyang Xu, Conrad Testagrose, Mutlu Demirer, Vikash Gupta, Ünal Akünal, Markus Bujotzek, Klaus H. Maier-Hein, Yi Qin 0006, Xiaomeng Li 0001, Jayashree Kalpathy-Cramer, Holger Roth |
Medical Image Anal. | 20 |
| 2023 | Why is the Winner the Best?abstractInternational benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multicenter study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and post-processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work. Matthias Eisenmann, Annika Reinke, Vivienn Weru, Minu Tizabi, Fabian Isensee, Tim Adler, Sharib Ali, Vincent Andrearczyk, Marc Aubreville, Ujjwal Baid, Spyridon Bakas, Niranjan Balu, Sophia Bano, Jorge Bernal, Sebastian Bodenstedt, Alessandro Casella, Veronika Cheplygina, Marie Daum, Marleen de Bruijne, Adrien Depeursinge, Reuben Dorent, Jan Egger, David Gage Ellis, Sandy Engelhardt, Melanie Ganz-Benjaminsen, Noha M. Ghatwary, Gabriel Girard, Patrick Godau, Anubha Gupta, Lasse Hansen, Kanako Harada, Mattias P. Heinrich, Nicholas Heller, Alessa Hering, Arnaud Huaulmé, Pierre Jannin, A. Emre Kavur, Oldrich Kodym, Michal Kozubek 0001, Jianning Li 0002, Hongwei Li 0004, Jun Ma 0016, Carlos Martín-Isla, Bjoern Menze, J. Alison Noble, Valentin Oreiller, Nicolas Padoy, Sarthak Pati, Kelly Payette, Tim Rädsch, Jonathan Rafael-Patino, Vivek Singh Bawa, Stefanie Speidel, Carole H. Sudre, Kimberlin M. H. van Wijnen, Martin Wagner 0001, D. Wei, Amine Yamlahi, Moi Hoon Yap, C. Yuan, Maximilian Zenk, A. Zia, David Zimmerer, Dogu Baran Aydogan, Binod Bhattarai, Louise Bloch, Raphael Brüngel, J. Cho, C. Choi, Qi Dou 0001, Ivan Ezhov, Christoph M. Friedrich, C. Fuller, Rebati Raman Gaire, Adrian Galdran, Álvaro García-Faura, Maria Grammatikopoulou, S. Hong, Mostafa Jahanifar, I. Jang, Abdolrahim Kadkhodamohammadi, I. Kang, Florian Kofler, S. Kondo, Hugo J. Kuijf, M. Luu, Tomaz Martincic, Pedro Morais, Mohamed A. Naser, Bruno Oliveira 0002, David Owen 0001, S. Pang, Szymon Plotka, Élodie Puybareau, Nasir M. Rajpoot, K. Ryu, Numan Saeed, Adam J. Shephard, Dejan Stepec, Ronast Subedi, Guillaume Tochon, Helena R. Torres, Hélène Urien, João L. Vilaça, Kareem A. Wahid, Benedikt Wiestler, Marek Wodzinski, F. Xia, J. Xie, Z. Xiong, Sen Yang 0006, Klaus H. Maier-Hein, Paul F. Jaeger, Annette Kopp-Schneider, Lena Maier-Hein |
CVPR | 122 |
| 2023 | cOOpD: Reformulating COPD Classification on Chest CT Scans as Anomaly Detection Using Contrastive Representations
Silvia Dias Almeida, Carsten T. Lüth, Tobias Norajitra, Tassilo Wald, Marco Nolden, Paul F. Jaeger, Claus Peter Heussel, Jürgen Biederer, Oliver Weinheimer, Klaus H. Maier-Hein |
MICCAI (5) | 10 |
| 2023 | Shape-Based Pose Estimation for Automatic Standard Views of the Knee
Lisa Kausch, Sarina Thomas, Holger Kunze, Jan Siad El Barbari, Klaus H. Maier-Hein |
MICCAI (7) | 5 |
| 2023 | Anatomy-Informed Data Augmentation for Enhanced Prostate Cancer Detection
Balint Kovacs, Nils Netzer, Michael Baumgartner 0001, Carolin Eith, Dimitrios Bounias, Clara Meinzer, Paul F. Jaeger, Kevin S. Zhang, Ralf Floca, Adrian Schrader, Fabian Isensee, Regula Gnirs, Magdalena Görtz, Viktoria Schütz, Albrecht Stenzinger, Markus Hohenfellner, Heinz-Peter Schlemmer, Ivo Wolf, David Bonekamp, Klaus H. Maier-Hein |
MICCAI (7) | 20 |
| 2023 | atTRACTive: Semi-automatic White Matter Tract Segmentation Using Active Learning
Robin Peretzke, Klaus H. Maier-Hein, Jonas Bohn, Yannick Kirchhoff, Saikat Roy, Sabrina Oberli-Palma, Daniela Becker, Pavlina Lenga, Peter Neher |
MICCAI (8) | 2 |
| 2023 | MedNeXt: Transformer-Driven Scaling of ConvNets for Medical Image Segmentation
Saikat Roy, Gregor Köhler, Constantin Ulrich, Michael Baumgartner 0001, Jens Petersen, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein |
MICCAI (4) | 8 |
| 2023 | MultiTalent: A Multi-dataset Approach to Medical Image Segmentation
Constantin Ulrich, Fabian Isensee, Tassilo Wald, Maximilian Zenk, Michael Baumgartner 0001, Klaus H. Maier-Hein |
MICCAI (3) | 6 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 72 |
| 2022 | Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon S. Kim 0002, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen |
BMVC | 8 |
| 2022 | C-arm positioning for standard projections during spinal implant placementabstractFluoroscopy-guided trauma and orthopedic surgeries involve the repeated acquisition of correct anatomy-specific standard projections for guidance, monitoring, and evaluating the surgical result. C-arm positioning is usually performed by hand, involving repeated or even continuous fluoroscopy at a cost of radiation exposure and time. We propose to automate this procedure and estimate the pose update for C-arm repositioning directly from a first X-ray without the need for a patient-specific computed tomography scan (CT) or additional technical equipment. Our method is trained on digitally reconstructed radiographs (DRRs) which uniquely provide ground truth labels for an arbitrary number of training examples. The simulated images are complemented with automatically generated segmentations, landmarks, and with simulated k-wires and screws. To successfully achieve a transfer from simulated to real X-rays, and also to increase the interpretability of results, the pipeline was designed to closely reflect the actual clinical decision-making process followed by spinal neurosurgeons. It explicitly incorporates steps such as region-of-interest (ROI) localization, detection of relevant and view-independent landmarks, and subsequent pose regression. The method was validated on a large human cadaver study simulating a real clinical scenario, including k-wires and screws. The proposed procedure obtained superior C-arm positioning accuracy of dθ=8.8°±4.2° average improvement (pt−test≪0.01), robustness, and generalization capabilities compared to the state-of-the-art direct pose regression framework. Lisa Kausch, Sarina Thomas, Holger Kunze, Tobias Norajitra, André Klein, Leonardo Ayala, Jan Siad El Barbari, Eric Mandelka, Maxim Privalov, Sven Yves Vetter, Andreas H. Mahnken, Lena Maier-Hein, Klaus H. Maier-Hein |
Medical Image Anal. | 13 |
| 2022 | Rapid artificial intelligence solutions in a pandemic - The COVID-19-20 Lung CT Lesion Segmentation Challenge
Holger Roth, Ziyue Xu 0001, Carlos Tor-Díez, Ramon Sánchez-Jacob, Jonathan Zember, Jose Molto, Wenqi Li 0001, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Dong Yang 0005, Ahmed Harouni, Nicola Rieke, Shishuai Hu, Fabian Isensee, Claire Tang, Qinji Yu, Jan Sölter, Vitali Liauchuk, Jan Hendrik Moltz, Bruno Oliveira 0002, Yong Xia 0001, Klaus H. Maier-Hein, Qikai Li, Andreas Husch, Vassili Kovalev, Alessa Hering, João L. Vilaça, Mona Flores, Daguang Xu, Bradford J. Wood, Marius George Linguraru |
Medical Image Anal. | 25 |
| 2022 | MOOD 2020: A Public Benchmark for Out-of-Distribution Detection and Localization on Medical ImagesabstractDetecting Out-of-Distribution (OoD) data is one of the greatest challenges in safe and robust deployment of machine learning algorithms in medicine. When the algorithms encounter cases that deviate from the distribution of the training data, they often produce incorrect and over-confident predictions. OoD detection algorithms aim to catch erroneous predictions in advance by analysing the data distribution and detecting potential instances of failure. Moreover, flagging OoD cases may support human readers in identifying incidental findings. Due to the increased interest in OoD algorithms, benchmarks for different domains have recently been established. In the medical imaging domain, for which reliable predictions are often essential, an open benchmark has been missing. We introduce the Medical-Out-Of-Distribution-Analysis-Challenge (MOOD) as an open, fair, and unbiased benchmark for OoD methods in the medical imaging domain. The analysis of the submitted algorithms shows that performance has a strong positive correlation with the perceived difficulty, and that all algorithms show a high variance for different anomalies, making it yet hard to recommend them for clinical practice. We also see a strong correlation between challenge ranking and performance on a simple toy test set, indicating that this might be a valuable addition as a proxy dataset during anomaly detection algorithm development. David Zimmerer, Peter M. Full, Fabian Isensee, Paul F. Jaeger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Roß, Annika Reinke, Antanas Kascenas, Bjørn Sand Jensen, Alison O'Neil, Jeremy Tan, Benjamin Hou, James Batten, Huaqi Qiu, Bernhard Kainz, Nina Shvetsova, Irina Fedulova, Dmitry V. Dylov, Baolun Yu, Jianyang Zhai, Jingtao Hu, Runxuan Si, Sihang Zhou 0001, Siqi Wang 0001, Xuerun Chen, Yang Zhao 0003, Sergio Naval Marimont, Giacomo Tarroni, Victor Saase, Lena Maier-Hein, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 34 |
| 2021 | nnDetection: A Self-configuring Method for Medical Object Detection
Michael Baumgartner 0001, Paul F. Jaeger, Fabian Isensee, Klaus H. Maier-Hein |
MICCAI (5) | 4 |
| 2021 | C-Arm Positioning for Spinal Standard Projections in Different Intra-operative Settings
Lisa Kausch, Sarina Thomas, Holger Kunze, Tobias Norajitra, André Klein, Jan Siad El Barbari, Maxim Privalov, Sven Yves Vetter, Andreas H. Mahnken, Lena Maier-Hein, Klaus H. Maier-Hein |
MICCAI (4) | 11 |
| 2021 | Continuous-Time Deep Glioma Growth Models
Jens Petersen, Fabian Isensee, Gregor Köhler, Paul F. Jaeger, David Zimmerer, Ulf Neuberger, Wolfgang Wick, Jürgen Debus, Sabine Heiland, Martin Bendszus, Philipp Vollmuth, Klaus H. Maier-Hein |
MICCAI (3) | 12 |
| 2021 | GP-ConvCNP: Better generalization for conditional convolutional Neural Processes on time series dataabstractNeural Processes (NPs) are a family of conditional generative models that are able to model a distribution over functions, in a way that allows them to perform predictions at test time conditioned on a number of context points. A recent addition to this family, Convolutional Conditional Neural Processes (ConvCNP), have shown remarkable improvement in performance over prior art, but we find that they sometimes struggle to generalize when applied to time series data. In particular, they are not robust to distribution shifts and fail to extrapolate observed patterns into the future. By incorporating a Gaussian Process into the model, we are able to remedy this and at the same time improve performance within distribution. As an added benefit, the Gaussian Process reintroduces the possibility to sample from the model, a key feature of other members in the NP family. Jens Petersen, Gregor Köhler, David Zimmerer, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein |
UAI | 6 |
| 2021 | The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge
Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xiaoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan 0002, Guangrui Mu, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiawei Yang 0002, Guangwei Xiong, Jiang Tian, Christopher J. Weight |
Medical Image Anal. | 3 |
| 2021 | CHAOS Challenge - combined (CT-MR) healthy abdominal organ segmentation
A. Emre Kavur, Naciye Sinem Gezer, Mustafa Baris, Sinem Aslan, Pierre-Henri Conze, Vladimir Groza, Duc Duy Pham, Soumick Chatterjee, Philipp Ernst, Savas Özkan, Bora Baydar, Dmitry A. Lachinov, Shuo Han 0001, Josef Pauli, Fabian Isensee, Matthias Perkonigg, Rachana Sathish, Ronnie Rajan, Debdoot Sheet, Gurbandurdy Dovletov, Oliver Speck, Andreas Nürnberger, Klaus H. Maier-Hein, Gozde Bozdagi Akar, Gozde Unal, Oguz Dicle, M. Alper Selver |
Medical Image Anal. | 23 |
| 2021 | Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challengeabstractIntraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts). Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein |
Medical Image Anal. | 32 |
| 2021 | Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation: The M&Ms ChallengeabstractThe emergence of deep learning has considerably advanced the state-of-the-art in cardiac magnetic resonance (CMR) segmentation. Many techniques have been proposed over the last few years, bringing the accuracy of automated segmentation close to human performance. However, these models have been all too often trained and validated using cardiac imaging samples from single clinical centres or homogeneous imaging protocols. This has prevented the development and validation of models that are generalizable across different clinical centres, imaging conditions or scanner vendors. To promote further research and scientific benchmarking in the field of generalizable deep learning for cardiac segmentation, this paper presents the results of the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) Challenge, which was recently organized as part of the MICCAI 2020 Conference. A total of 14 teams submitted different solutions to the problem, combining various baseline models, data augmentation strategies, and domain adaptation techniques. The obtained results indicate the importance of intensity-driven data augmentation, as well as the need for further research to improve generalizability towards unseen scanner vendors or new imaging protocols. Furthermore, we present a new resource of 375 heterogeneous CMR datasets acquired by using four different scanner vendors in six hospitals and three different countries (Spain, Canada and Germany), which we provide as open-access for the community to enable future research in the field. Víctor M. Campello, Polyxeni Gkontra, Cristian Izquierdo, Carlos Martín-Isla, Alireza Sojoudi, Peter M. Full, Klaus H. Maier-Hein, Yao Zhang 0010, Zhiqiang He 0002, Jun Ma 0016, Mario Parreño, Alberto Albiol, Fanwei Kong, Shawn C. Shadden, Jorge Corral Acero, Vaanathi Sundaresan, Mina Saber, Mustafa A. Alattar, Hongwei Li 0004, Bjoern Menze, Firas Khader, Christoph Haarburger, Cian M. Scannell, Mitko Veta, Adam Carscadden, Kumaradevan Punithakumar, Xiao Liu 0037, Sotirios A. Tsaftaris, Xiaoqiong Huang, Xin Yang 0009, Lei Li 0020, Xiahai Zhuang, David Viladés, Martín Luís Descalzo, Andrea Guala 0002, Lucia La Mura, Matthias G. W. Friedrich, Ria Garg, Julie Lebel, Filipe Henriques, Mahir Karakas, Ersin Çavus, Steffen E. Petersen, Sergio Escalera, Santi Seguí, Jose Rodriguez-Palomares, Karim Lekadir |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Deep Probabilistic Modeling of Glioma Growth
Jens Petersen, Paul F. Jaeger, Fabian Isensee, Simon Kohl, Ulf Neuberger, Wolfgang Wick, Jürgen Debus, Sabine Heiland, Martin Bendszus, Philipp Vollmuth, Klaus H. Maier-Hein |
MICCAI (2) | 11 |
| 2019 | Unsupervised Anomaly Localization Using Variational Auto-Encoders
David Zimmerer, Fabian Isensee, Jens Petersen, Simon Kohl, Klaus H. Maier-Hein |
MICCAI (4) | 5 |
| 2019 | MITK-ModelFit: A generic open-source framework for model fits and their exploration in medical imaging - design, implementation and application on the example of DCE-MRIabstractBACKGROUND: Many medical imaging techniques utilize fitting approaches for quantitative parameter estimation and analysis. Common examples are pharmacokinetic modeling in dynamic contrast-enhanced (DCE) magnetic resonance imaging (MRI)/computed tomography (CT), apparent diffusion coefficient calculations and intravoxel incoherent motion modeling in diffusion-weighted MRI and Z-spectra analysis in chemical exchange saturation transfer MRI. Most available software tools are limited to a special purpose and do not allow for own developments and extensions. Furthermore, they are mostly designed as stand-alone solutions using external frameworks and thus cannot be easily incorporated natively in the analysis workflow. RESULTS: We present a framework for medical image fitting tasks that is included in the Medical Imaging Interaction Toolkit MITK, following a rigorous open-source, well-integrated and operating system independent policy. Software engineering-wise, the local models, the fitting infrastructure and the results representation are abstracted and thus can be easily adapted to any model fitting task on image data, independent of image modality or model. Several ready-to-use libraries for model fitting and use-cases, including fit evaluation and visualization, were implemented. Their embedding into MITK allows for easy data loading, pre- and post-processing and thus a natural inclusion of model fitting into an overarching workflow. As an example, we present a comprehensive set of plug-ins for the analysis of DCE MRI data, which we validated on existing and novel digital phantoms, yielding competitive deviations between fit and ground truth. CONCLUSIONS: Providing a very flexible environment, our software mainly addresses developers of medical imaging software that includes model fitting algorithms and tools. Additionally, the framework is of high interest to users in the domain of perfusion MRI, as it offers feature-rich, freely available, validated tools to perform pharmacokinetic analysis on DCE MRI data, with both interactive and automatized batch processing workflows. Charlotte Debus, Ralf Floca, Michael Ingrisch, Ina Kompan, Klaus H. Maier-Hein, Marco Nolden |
BMC Bioinform. | 5 |
| 2019 | Combined tract segmentation and orientation mapping for bundle-specific tractographyabstractWhile the major white matter tracts are of great interest to numerous studies in neuroscience and medicine, their manual dissection in larger cohorts from diffusion MRI tractograms is time-consuming, requires expert knowledge and is hard to reproduce. In previous work we presented tract orientation mapping (TOM) as a novel concept for bundle-specific tractography. It is based on a learned mapping from the original fiber orientation distribution function (FOD) peaks to tract specific peaks, called tract orientation maps. Each tract orientation map represents the voxel-wise principal orientation of one tract. Here, we present an extension of this approach that combines TOM with accurate segmentations of the tract outline and its start and end region. We also introduce a custom probabilistic tracking algorithm that samples from a Gaussian distribution with fixed standard deviation centered on each peak thus enabling more complete trackings on the tract orientation maps than deterministic tracking. These extensions enable the automatic creation of bundle-specific tractograms with previously unseen accuracy. We show for 72 different bundles on high quality, low quality and phantom data that our approach runs faster and produces more accurate bundle-specific tractograms than 7 state of the art benchmark methods while avoiding cumbersome processing steps like whole brain tractography, non-linear registration, clustering or manual dissection. Moreover, we show on 17 datasets that our approach generalizes well to datasets acquired with different scanners and settings as well as with pathologies. The code of our method is openly available at https://github.com/MIC-DKFZ/TractSeg. Jakob Wasserthal, Peter Neher, Dusan Hirjak, Klaus H. Maier-Hein |
Medical Image Anal. | 4 |
| 2018 | Anchor-Constrained Plausibility (ACP): A Novel Concept for Assessing Tractography and Reducing False-Positives
Peter Neher, Bram Stieltjes, Klaus H. Maier-Hein |
MICCAI (3) | 3 |
| 2018 | Tract Orientation Mapping for Bundle-Specific Tractography
Jakob Wasserthal, Peter Neher, Klaus H. Maier-Hein |
MICCAI (3) | 3 |
| 2018 | A Probabilistic U-Net for Segmentation of Ambiguous ImagesabstractMany real-world vision problems suffer from inherent ambiguities. In clinical applications for example, it might not be clear from a CT scan alone which particular region is cancer tissue. Therefore a group of graders typically produces a set of diverse but plausible segmentations. We consider the task of learning a distribution over segmentations given an input. To this end we propose a generative segmentation model based on a combination of a U-Net with a conditional variational autoencoder that is capable of efficiently producing an unlimited number of plausible hypotheses. We show on a lung abnormalities segmentation task and on a Cityscapes segmentation task that our model reproduces the possible segmentation variants as well as the frequencies with which they occur, doing so significantly better than published approaches. These models could have a high impact in real-world applications, such as being used as clinical decision-making algorithms accounting for multiple plausible semantic segmentation hypotheses to provide possible diagnoses and recommend further actions to resolve the present ambiguities. Simon Kohl, Bernardino Romera-Paredes, Clemens Meyer, Jeffrey De Fauw, Joseph R. Ledsam, Klaus H. Maier-Hein, S. M. Ali Eslami, Danilo Jimenez Rezende, Olaf Ronneberger |
NeurIPS | 6 |
| 2018 | Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved?abstractDelineation of the left ventricular cavity, myocardium, and right ventricle from cardiac magnetic resonance images (multi-slice 2-D cine MRI) is a common clinical task to establish diagnosis. The automation of the corresponding tasks has thus been the subject of intense research over the past decades. In this paper, we introduce the "Automatic Cardiac Diagnosis Challenge" dataset (ACDC), the largest publicly available and fully annotated dataset for the purpose of cardiac MRI (CMR) assessment. The dataset contains data from 150 multi-equipments CMRI recordings with reference measurements and classification from two medical experts. The overarching objective of this paper is to measure how far state-of-the-art deep learning methods can go at assessing CMRI, i.e., segmenting the myocardium and the two ventricles as well as classifying pathologies. In the wake of the 2017 MICCAI-ACDC challenge, we report results from deep learning methods provided by nine research groups for the segmentation task and four groups for the classification task. Results show that the best methods faithfully reproduce the expert analysis, leading to a mean value of 0.97 correlation score for the automatic extraction of clinical indices and an accuracy of 0.96 for automatic diagnosis. These results clearly open the door to highly accurate and fully automatic analysis of cardiac CMRI. We also identify scenarios for which deep learning methods are still failing. Both the dataset and detailed results are publicly available online, while the platform will remain open for new submissions. Olivier Bernard 0001, Alain Lalande, Clément Zotti, Frederic Cervenansky, Xin Yang 0009, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara 0001, Miguel Ángel González Ballester, Gerard Sanroma, Sandy Napel, Steffen E. Petersen, Georgios Tziritas, Ilias Grinias, Mahendra Khened, Alex Varghese, Ganapathy Krishnamurthi, Marc-Michel Rohé, Xavier Pennec, Maxime Sermesant, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein, Peter M. Full, Ivo Wolf, Sandy Engelhardt, Christian F. Baumgartner, Lisa M. Koch, Jelmer M. Wolterink, Ivana Isgum, Yeonggul Jang, Yoonmi Hong, Jay Patravali, Shubham Jain 0006, Olivier Humbert, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 24 |
| 2017 | Revealing Hidden Potentials of the q-Space Signal in Breast Cancer
Paul F. Jaeger, Sebastian Bickelhaupt, Frederik B. Laun, Wolfgang Lederer, Heidi Daniel, Tristan Anselm Kuder, Daniel Paech, David Bonekamp, Alexander Radbruch, Stefan Delorme, Heinz-Peter Schlemmer, Franziska Steudle, Klaus H. Maier-Hein |
MICCAI (1) | 13 |
| 2017 | Learn to Track: Deep Learning for Tractography
Philippe Poulin, Marc-Alexandre Côté, Jean-Christophe Houde, Laurent Petit, Peter Neher, Klaus H. Maier-Hein, Hugo Larochelle, Maxime Descoteaux |
MICCAI (1) | 6 |
| 2017 | Physiological Parameter Estimation from Multispectral Images Unleashed
Sebastian J. Wirkert, Anant Suraj Vemuri, Hannes Kenngott, Sara Moccia, Michael Götz, Benjamin F. B. Mayer, Klaus H. Maier-Hein, Daniel S. Elson, Lena Maier-Hein |
MICCAI (3) | 7 |
| 2017 | MiMSeg - an algorithm for automated detection of tumor tissue on NMR apparent diffusion coefficient mapsabstractAlthough there are several data analysis frameworks, both commercial and open source, supporting the detection of tumours on nuclear magnetic resonance (NMR) sequences, none of them gives satisfactory results in the case of low volume tumors. The majority of the frameworks require the detailed analysis of at least two sequences of the examined sample, or give sample specific thresholds distinguishing between the tumor and subtypes of healthy tissue. In this paper, we present a novel algorithm for the automated estimation of tumor specific cut-off values in the domain of the apparent diffusion coefficient (ADC). Once the cut-off characteristics for a particular type of tumor is estimated, their further usage on other independent samples does not require any calculations except for an easy thresholding. The proposed methodology is a combination of classical decomposition of ADC distribution into a Gaussian mixture model (GMM) with k-means clustering subsequently performed on the parameters of mixture model components, leading to the identification of ADC distributions for every tissue type. The maximum conditional probability criterion gives the final threshold estimate. The developed signal analysis pipeline was applied to the problem of Glioblastoma Multiforme grade IV brain tumor segmentation, with a dataset of 119 randomly chosen ADC maps and Leave-One-Out cross-validation procedure for population error estimate. Additionally, a comparison to standard GMM based tumor segmentation algorithms as well as to three other automated segmentation methods was performed and the obtained tumor regions were referenced to the segmentation done by a human expert. The results demonstrate the average MiMSeg similarity to the expert-curated decision measured by the Dice coefficient as equal to 89.2% (with 95% confidence interval 87.7 ÷ 90.6). The MiMSeg algorithm significantly outperforms other techniques in the case of small tumors (of volume less than 10%), obtaining similarity to the expert-curated decision at the level 86.7%, with 44.9% obtained by standard GMM, 79.0% by Self-Organising-Maps algorithm, 68.7% by Murakami’s algorithm, and 78.2% by Kang’s method. Franciszek Binczyk, Bram Stjelties, Christian Weber 0001, Michael Götz, Klaus H. Maier-Hein, Hans-Peter Meinzer, Barbara Bobek-Billewicz, Rafal Tarnawski, Joanna Polanska |
Inf. Sci. | 5 |
| 2017 | 3D Statistical Shape Models Incorporating Landmark-Wise Random Regression Forests for Omni-Directional Landmark Detectionabstract3D Statistical Shape Models (3D-SSM) are widely used for medical image segmentation. However, during segmentation, they typically perform a very limited unidirectional search for suitable landmark positions in the image, relying on weak learners or use-case specific appearance models that solely take local image information into account. As a consequence, segmentation errors arise, and results in general depend on the accuracy of a previous model initialization. Furthermore, these methods become subject to a tedious and use-case dependent parameter tuning in order to obtain optimized results. To overcome these limitations, we propose an extension of 3D-SSM by landmark-wise random regression forests that perform an enhanced omni-directional search for landmark positions, thereby taking rich non-local image information into account. In addition, we provide a long distance model fitting based on a multi-scale approach, that allows an accurate and reproducible segmentation even from distant image positions, thus enabling an application without model initialization. Finally, translation of the proposed method to different organs is straightforward and requires no adaptation of the training process. In segmentation experiments on 45 clinical CT volumes, the proposed omni-directional search significantly increased accuracy and displayed great precision regardless of model initialization. Furthermore, for liver, spleen and kidney segmentation in a competitive multi-organ labeling challenge on publicly available data, the proposed method achieved similar or better results than the state of the art. Finally, liver segmentation results were obtained that successfully compete with specialized state-of-the-art methods from the well-known liver segmentation challenge SLIVER. Tobias Norajitra, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Crowd-Algorithm Collaboration for Large-Scale Endoscopic Image Annotation with Confidence
Lena Maier-Hein, Tobias Roß, Janek Gröhl, Ben Glocker, Sebastian Bodenstedt, Christian Stock, Eric Heim, Michael Götz, Sebastian J. Wirkert, Hannes Kenngott, Stefanie Speidel, Klaus H. Maier-Hein |
MICCAI (2) | 12 |
| 2016 | DALSA: Domain Adaptation for Supervised Learning From Sparsely Annotated MR ImagesabstractWe propose a new method that employs transfer learning techniques to effectively correct sampling selection errors introduced by sparse annotations during supervised learning for automated tumor segmentation. The practicality of current learning-based automated tissue classification approaches is severely impeded by their dependency on manually segmented training databases that need to be recreated for each scenario of application, site, or acquisition setup. The comprehensive annotation of reference datasets can be highly labor-intensive, complex, and error-prone. The proposed method derives high-quality classifiers for the different tissue classes from sparse and unambiguous annotations and employs domain adaptation techniques for effectively correcting sampling selection errors introduced by the sparse sampling. The new approach is validated on labeled, multi-modal MR images of 19 patients with malignant gliomas and by comparative analysis on the BraTS 2013 challenge data sets. Compared to training on fully labeled data, we reduced the time for labeling and training by a factor greater than 70 and 180 respectively without sacrificing accuracy. This dramatically eases the establishment and constant extension of large annotated databases in various scenarios and imaging setups and thus represents an important step towards practical applicability of learning-based approaches in tissue classification. Michael Götz, Christian Weber 0001, Franciszek Binczyk, Joanna Polanska, Rafal Tarnawski, Barbara Bobek-Billewicz, Ullrich Köthe, Jens Kleesiek, Bram Stieltjes, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 10 |
| 2016 | Multi-Objective Memetic Search for Robust Motion and Distortion Correction in Diffusion MRIabstractEffective image-based artifact correction is an essential step in the analysis of diffusion MR images. Many current approaches are based on retrospective registration, which becomes challenging in the realm of high b-values and low signal-to-noise ratio, rendering the corresponding correction schemes more and more ineffective. We propose a novel registration scheme based on memetic search optimization that allows for simultaneous exploitation of different signal intensity relationships between the images, leading to more robust registration results. We demonstrate the increased robustness and efficacy of our method on simulated as well as in vivo datasets. In contrast to the state-of-art methods, the median target registration error (TRE) stayed below the voxel size even for high b-values (3000 s · mm-2and higher) and low SNR conditions. We also demonstrate the increased precision in diffusion-derived quantities by evaluating Neurite Orientation Dispersion and Density Imaging (NODDI) derived measures on a in vivo dataset with severe motion artifacts. These promising results will potentially inspire further studies on metaheuristic optimization in diffusion MRI artifact correction and image registration in general. Jan Hering, Ivo Wolf, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 3 |
| 2015 | A Machine Learning Based Approach to Fiber Tractography Using Classifier Voting
Peter Neher, Michael Götz, Tobias Norajitra, Christian Weber 0001, Klaus H. Maier-Hein |
MICCAI (1) | 5 |
| 2015 | Strengths and weaknesses of state of the art fiber tractography pipelines - A comprehensive in-vivo and phantom evaluation study using Tractometer
Peter Neher, Maxime Descoteaux, Jean-Christophe Houde, Bram Stieltjes, Klaus H. Maier-Hein |
Medical Image Anal. | 5 |
| 2006 | Automated MRI-Based Quantification of the Cerebral Atrophy Providing Diagnostic Information on Mild Cognitive Impairment and Alzheimer's DiseaseabstractAlzheimer’s disease (AD) is a major public health challenge as the median age of the industrialized world’s population is increasing gradually. No cure for this disease has yet been found and the development of new treatments has become a topic of major research interest. This paper aims to propose a sequence of fully automated MRI-based image analysis steps to measure the development stage of atrophy in the brain. The results have been validated on a mixed group of 68 subjects by distinguishing between AD patients, MCIs and health controls using linear classifiers and ANNs. The best classifier identified unseen AD patients correctly in 80% of the cases and control subjects in 85%. Recognizing more than 8 out of 10 MCI subjects, the method also yields an early indication of AD. This simple yet powerful analysis can compete with othermore time-consuming and semi-automaticmethodologies. It could abet an AD diagnosis and provide a tool for measuring the success of therapies. Klaus H. Maier-Hein, Aldo von Wangenheim, Rüdiger Dillmann, Roland Unterhinninghofen |
CBMS | 1 |