EDBT 2026 Demo / reviewers in the wild / expert
Michael Baumgartner 0001
dblp:66/4721-1
· DBLP profile ↗
13ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-4455-9917ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-class segmentation of aortic branches and zones in computed tomography angiography: The AortaSeg24 challenge
Muhammad Imran 0013, Jonathan R. Krebs, Vishal Balaji Sivaraman, Amarjeet Kumar, Walker R. Ueland, Michael J. Fassler, Lisheng Wang, Maximilian Rokuss, Michael Baumgartner 0001, Yannick Kirchhof, Klaus H. Maier-Hein, Fabian Isensee, Shuolin Liu, Bong Thanh Nguyen, Dong-jin Shin, Park Ji-Woo, Matthew Choi, Kwang-Hyun Uhm, Sung-Jea Ko, Chanwoong Lee, Jaehee Chun, Yun Gu, Zhaohong Pan, Xiaokun Liang, Markus Tiefenthaler, Enrique Almar-Munoz, Matthias Schwab, Mikhail Kotyushev, Rostislav Epifanov, Marek Wodzinski, Henning Müller, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Zhiwei Wang 0002, Kaixiang Yang 0004, Jintao Ren, Stine Sofia Korreman, Yuchong Gao, Hongye Zeng, Jinghua Yue, Fugen Zhou, Alexander Cosman, Muxuan Liang, Gilbert R. Upchurch Jr., Yuyin Zhou, Michol A. Cooper, Wei Shao 0008 |
Medical Image Anal. | 14 |
| 2026 | Advances in automated fetal brain MRI segmentation and biometry: Insights from the FeTA 2024 challengeabstractAccurate fetal brain tissue segmentation and biometric measurement are essential for monitoring neurodevelopment and detecting abnormalities in utero. The Fetal Tissue Annotation (FeTA) Challenges have established robust multi-center benchmarks for evaluating state-of-the-art segmentation methods. This paper presents the results of the 2024 challenge edition, which introduced three key innovations. First, we introduced a topology-aware metric based on the Euler characteristic difference (ED) to overcome the performance plateau observed with traditional metrics like Dice or Hausdorff distance (HD), as the performance of the best models in segmentation surpassed the inter-rater variability. While the best teams reached similar scores in Dice (0.81-0.82) and HD95 (2.1-2.3 mm), ED provided greater discriminative power: the winning method achieved an ED of 20.9, representing roughly a 50% improvement over the second- and third-ranked teams despite comparable Dice scores. Second, we introduced a new 0.55T low-field MRI test set, which, when paired with high-quality super-resolution reconstruction, achieved the highest segmentation performance across all test cohorts (Dice=0.86, HD95=1.69, ED=6.26). This provides the first quantitative evidence that low-cost, low-field MRI can match or surpass high-field systems in automated fetal brain segmentation. Third, the new biometry estimation task exposed a clear performance gap: although the best model reached a mean average percentage error (MAPE) of 7.72%, most submissions failed to outperform a simple gestational-age-based linear regression model (MAPE=9.56%), and all remained above inter-rater variability with a MAPE of 5.38%. Finally, by analyzing the top-performing models from FeTA 2024 alongside those from previous challenge editions, we identify ensembles of 3D nnU-Net trained on both real and synthetic data with both image- and anatomy-level augmentations as the most effective approaches for fetal brain segmentation. Our quantitative analysis reveals that acquisition site, super-resolution strategy, and image quality are the primary sources of domain shift, informing recommendations to enhance the robustness and generalizability of automated fetal brain analysis methods. Vladyslav Zalevskyi, Thomas Sanchez, Misha P. T. Kaandorp, Margaux Roulet, Diego Fajardo-Rojas, Liu Li 0001, Jana Hutter, Hongwei Li 0004, Matthew J. Barkovich, Luca Wilhelmi, Aline Dändliker, Céline Steger, Mériam Koob, Yvan Gomez, Anton Jakovcic, Melita Klaic, Ana Adzic, Pavel Markovic, Gracia Grabaric, Milan Rados, Jordina Aviles Verdera, Gregor Kasprian, Gregor Dovjak, Raphael Gaubert-Rachmühl, Maurice Aschwanden, Davood Karimi, Denis Peruzzo, Tommaso Ciceri, Giorgio Longari, Rachika E. Hamadache, Amina Bouzid, Xavier Lladó, Simone Chiarella, Gerard Martí-Juan, Miguel Ángel González Ballester, Marco Castellaro, Marco Pinamonti, Valentina Visani, Robin Cremese, Keïn Sam, Fleur Gaudfernau, Param Ahir, Mehul Parikh, Maximilian Zenk, Michael Baumgartner 0001, Klaus H. Maier-Hein, Li Tianhong, Zhao Longfei, Domen Preloznik, Ziga Spiclin, Jae Won Choi, Guotai Wang, Lyuyang Tong, Bo Du 0001, Andrea Gondova, Sungmin You, Kiho Im, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, András Jakab, Roxane Licandro, Kelly Payette, Meritxell Bach Cuadra |
Medical Image Anal. | 47 |
| 2025 | The Missing Piece: A Case for Pre-training in 3D Medical Object Detection
Katharina Eckstein, Constantin Ulrich, Michael Baumgartner 0001, Jessica Kächele, Dimitrios Bounias, Tassilo Wald, Ralf Floca, Klaus H. Maier-Hein |
MICCAI (4) | 3 |
| 2025 | Including AI in diffusion-weighted breast MRI has potential to increase reader confidence and reduce workloadabstractOBJECTIVES: Breast diffusion-weighted imaging (DWI) has shown potential as a standalone imaging technique for certain indications, eg, supplemental screening of women with dense breasts. This study evaluates an artificial intelligence (AI)-powered computer-aided diagnosis (CAD) system for clinical interpretation and workload reduction in breast DWI. MATERIALS AND METHODS: This retrospective IRB-approved study included: n = 824 examinations for model development (2017-2020) and n = 235 for evaluation (01/2021-06/2021). Readings were performed by three readers using either the AI-CAD or manual readings. BI-RADS-like (Breast Imaging Reporting and Data System) classification was based on DWI. Histopathology served as ground truth. The model was nnDetection-based, trained using 5-fold cross-validation and ensembling. Statistical significance was determined using McNemar's test. Inter-rater agreement was calculated using Cohen's kappa. Model performance was calculated using the area under the receiver operating curve (AUC). RESULTS: The AI-augmented approach significantly reduced BI-RADS-like 3 calls in breast DWI by 29% (P =.019) and increased interrater agreement (0.57 ± 0.10 vs 0.49 ± 0.11), while preserving diagnostic accuracy. Two of the three readers detected more malignant lesions (63/69 vs 59/69 and 64/69 vs 62/69) with the AI-CAD. The AI model achieved an AUC of 0.78 (95% CI: [0.72, 0.85]; P <.001), which increased for women at screening age to 0.82 (95% CI: [0.73, 0.90]; P <.001), indicating a potential for workload reduction of 20.9% at 96% sensitivity. DISCUSSION AND CONCLUSION: Breast DWI might benefit from AI support. In our study, AI showed potential for reduction of BI-RADS-like 3 calls and increase of inter-rater agreement. However, given the limited study size, further research is needed. Dimitrios Bounias, Lina Simons, Michael Baumgartner 0001, Chris Ehring, Peter Neher, Lorenz A. Kapsner, Balint Kovacs, Ralf Floca, Paul F. Jaeger, Jessica Eberle, Dominique Hadler, Frederik B. Laun, Sabine Ohlmeyer, Lena Maier-Hein, Michael Uder, Evelyn Wenkel, Klaus H. Maier-Hein, Sebastian Bickelhaupt |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation
Fabian Isensee, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger |
MICCAI (9) | 4 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 15 |
| 2024 | Overcoming Common Flaws in the Evaluation of Selective Classification SystemsabstractSelective Classification, wherein models can reject low-confidence predictions, promises reliable translation of machine-learning based classification systems to real-world scenarios such as clinical diagnostics. While current evaluation of these systems typically assumes fixed working points based on pre-defined rejection thresholds, methodological progress requires benchmarking the general performance of systems akin to the $\mathrm{AUROC}$ in standard classification. In this work, we define 5 requirements for multi-threshold metrics in selective classification regarding task alignment, interpretability, and flexibility, and show how current approaches fail to meet them. We propose the Area under the Generalized Risk Coverage curve ($\mathrm{AUGRC}$), which meets all requirements and can be directly interpreted as the average risk of undetected failures. We empirically demonstrate the relevance of $\mathrm{AUGRC}$ on a comprehensive benchmark spanning 6 data sets and 13 confidence scoring functions. We find that the proposed metric substantially changes metric rankings on 5 out of the 6 data sets. Jeremias Traub, Till J. Bungert, Carsten T. Lüth, Michael Baumgartner 0001, Klaus H. Maier-Hein, Lena Maier-Hein, Paul F. Jaeger |
NeurIPS | 4 |
| 2024 | Decoupling Semantic Similarity from Spatial Alignment for Neural NetworksabstractWhat representation do deep neural networks learn? How similar are images to each other for neural networks? Despite the overwhelming success of deep learning methods key questions about their internal workings still remain largely unanswered, due to their internal high dimensionality and complexity. To address this, one approach is to measure the similarity of activation responses to various inputs.
Representational Similarity Matrices (RSMs) distill this similarity into scalar values for each input pair.
These matrices encapsulate the entire similarity structure of a system, indicating which input lead to similar responses.
While the similarity between images is ambiguous, we argue that the spatial location of semantic objects does neither influence human perception nor deep learning classifiers. Thus this should be reflected in the definition of similarity between image responses for computer vision systems. Revisiting the established similarity calculations for RSMs we expose their sensitivity to spatial alignment. In this paper we propose to solve this through _semantic RSMs_, which are invariant to spatial permutation. We measure semantic similarity between input responses by formulating it as a set-matching problem. Further, we quantify the superiority of _semantic_ RSMs over _spatio-semantic_ RSMs through image retrieval and by comparing the similarity between representations to the similarity between predicted class probabilities. Tassilo Wald, Constantin Ulrich, Priyank Jaini, Gregor Köhler, David Zimmerer, Stefan Denner, Fabian Isensee, Michael Baumgartner 0001, Klaus H. Maier-Hein |
NeurIPS | 8 |
| 2023 | Anatomy-Informed Data Augmentation for Enhanced Prostate Cancer Detection
Balint Kovacs, Nils Netzer, Michael Baumgartner 0001, Carolin Eith, Dimitrios Bounias, Clara Meinzer, Paul F. Jaeger, Kevin S. Zhang, Ralf Floca, Adrian Schrader, Fabian Isensee, Regula Gnirs, Magdalena Görtz, Viktoria Schütz, Albrecht Stenzinger, Markus Hohenfellner, Heinz-Peter Schlemmer, Ivo Wolf, David Bonekamp, Klaus H. Maier-Hein |
MICCAI (7) | 3 |
| 2023 | MedNeXt: Transformer-Driven Scaling of ConvNets for Medical Image Segmentation
Saikat Roy, Gregor Köhler, Constantin Ulrich, Michael Baumgartner 0001, Jens Petersen, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein |
MICCAI (4) | 4 |
| 2023 | MultiTalent: A Multi-dataset Approach to Medical Image Segmentation
Constantin Ulrich, Fabian Isensee, Tassilo Wald, Maximilian Zenk, Michael Baumgartner 0001, Klaus H. Maier-Hein |
MICCAI (3) | 5 |
| 2021 | nnDetection: A Self-configuring Method for Medical Object Detection
Michael Baumgartner 0001, Paul F. Jaeger, Fabian Isensee, Klaus H. Maier-Hein |
MICCAI (5) | 1 |
| 2019 | Multi Scale Curriculum CNN for Context-Aware Breast MRI Malignancy Classification
Christoph Haarburger, Michael Baumgartner 0001, Daniel Truhn, Mirjam Broeckmann, Hannah Schneider, Simone Schrading, Christiane Kuhl, Dorit Merhof |
MICCAI (4) | 2 |