Constantin Ulrich

dblp:327/3310 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0003-3002-8170ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
YearPublicationVenuePosition
2026 Multi-structure segmentation in CBCT volumes: The ToothFairy2 challenge
abstract
Cone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT.
Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana
Medical Image Anal.15
2026 MBAS2024: A large-scale benchmark for multi-class bi-atrial segmentation in multi-center contrast-enhanced MRIs
abstract
Atrial fibrillation (AF), the most common cardiac arrhythmia, affects one in three adults over 45 years of age. Improving its treatment requires a better understanding of bi-atrial anatomy. Existing benchmarks have focused on the left atrial (LA) cavity, overlooking the fundamental challenges posed by bi-atrial anatomy, most notably the thin atrial walls, which are critical for substrate-guided ablation planning in patients with atrial fibrillation. To address these limitations, the Multi-class Bi-Atrial Segmentation 2024 Challenge (MBAS2024) introduced the first large-scale, multi-class benchmark for simultaneous segmentation of the LA cavity, right atrial (RA) cavity, and bi-atrial walls from late gadolinium-enhanced (LGE) MRI. We systematically evaluated 13 state-of-the-art methods on the world's largest curated bi-atrial dataset, comprising 175 3D multi-center scans with expert-validated annotations, providing a comprehensive assessment of current methodological capabilities and limitations. Key findings include: segmentation of the LA and RA cavities is generally robust to image quality, whereas atrial wall delineation is highly sensitive to image degradation. Performance varies across centers, indicating limited generalization of atrial wall segmentation across different acquisition protocols. Model architecture, rather than hyperparameter tuning, is the primary driver of performance, with U-Net-based models and emerging state-space models (e.g., UMambaBot) achieving higher accuracy at modest computational cost. Segmentation accuracy also varies along the slice dimension, with central slices segmented more reliably. Finally, hybrid labeling strategies-separating LA and RA cavities while merging bi-atrial walls into a single class-consistently improve performance. The MBAS2024 challenge establishes a foundational benchmark for bi-atrial segmentation, providing validated baselines and actionable insights to guide the development of clinically relevant, efficient, and anatomically aware segmentation algorithms to improve targeted ablation in patients with AF.
Fangqiang Xu, James Kennelly, Alexander M. Zolotarev, Caroline H. Roney, Michal Nohel, Constantin Ulrich, Bryan Anenberg, Peter Chang, Yu Hon On, Marta Varela, Claas Thesing, Abhirup Banerjee, Enrique Almar-Munoz, Markus Tiefenthaler, Susana Merino-Caviedes, Emmanuel C. Nnadozie, Abdul Qayyum 0002, Moona Mazher, Waqas Anwaar, Wufeng Xue, Jingsu Kang, Lucas Beveridge, Malitha Gunawardhana, Kunihiko Kiuchi, Martin K. Stiles, Jichao Zhao
Medical Image Anal.6
2025 LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imaging
abstract
In this work, we present LesionLocator, a framework for zero-shot longitudinal lesion tracking and segmentation in 3D medical imaging, establishing the first end-to-end model capable of 4D tracking with dense spatial prompts. Our model leverages an extensive dataset of 23,262 annotated medical scans, as well as synthesized longitudinal data across diverse lesion types. The diversity and scale of our dataset significantly enhances model generalizability to real-world medical imaging challenges and addresses key limitations in longitudinal data availability. LesionLocator outperforms all existing promptable models in lesion segmentation by nearly 10 dice points, reaching human-level performance, and achieves state-of-the-art results in lesion tracking, with superior lesion retrieval and segmentation accuracy. LesionLocator not only sets a new benchmark in universal promptable lesion segmentation and automated longitudinal lesion tracking but also provides the first open-access solution of its kind, releasing our synthetic 4D dataset and model to the community, empowering future advancements in medical imaging. Code is available at: www.github.com/MIC-DKFZ/LesionLocator
Maximilian Rokuss, Yannick Kirchhoff, Seval Akbal, Balint Kovacs, Saikat Roy, Constantin Ulrich, Tassilo Wald, Lukas T. Rotkopf, Heinz-Peter Schlemmer, Klaus H. Maier-Hein
CVPR6
2025 Revisiting MAE Pre-training for 3D Medical Image Segmentation
abstract
Self-Supervised Learning (SSL) presents an exciting opportunity to unlock the potential of vast, untapped clinical datasets, for various downstream applications that suffer from the scarcity of labeled data. While SSL has revolutionized fields like natural language processing and computer vision, its adoption in 3D medical image computing has been limited by three key pitfalls: Small pre-training dataset sizes, architectures inadequate for 3D medical image analysis, and insufficient evaluation practices. In this paper, we address these issues by i) leveraging a large-scale dataset of 39k 3D brain MRI volumes and ii) using a Residual Encoder U-Net architecture within the state-of-the-art nnU-Net framework. iii) A robust development framework, incorporating 5 development and 8 testing brain MRI segmentation datasets, allowed performance-driven design decisions to optimize the simple concept of Masked Auto Encoders (MAEs) for 3D CNNs. The resulting model not only surpasses previous SSL methods but also outperforms the strong nnU-Net baseline by an average of approximately 3 Dice points setting a new state-of-the-art. Our code and models are made available here.
Tassilo Wald, Constantin Ulrich, Stanislav Lukyanenko, Andrei Goncharov, Alberto Paderno, Maximilian Miller, Leander Maerkisch, Paul F. Jaeger, Klaus H. Maier-Hein
CVPR2
2025 An OpenMind for 3D Medical Vision Self-supervised Learning
abstract
The field of self-supervised learning (SSL) for 3D medical images lacks consistency and standardization. While many methods have been developed, it is impossible to identify the current state-of-the-art, due to i) varying and small pretraining datasets, ii) varying architectures, and iii) being evaluated on differing downstream datasets. In this paper, we bring clarity to this field and lay the foundation for further method advancements through three key contributions: We a) publish the largest publicly available pre-training dataset comprising 114k 3D brain MRI volumes, enabling all practitioners to pre-train on a large-scale dataset. We b) benchmark existing 3D self-supervised learning methods on this dataset for a state-of-the-art CNN and Transformer architecture, clarifying the state of 3D SSL pre-training. Among many findings, we show that pre-trained methods can exceed a strong from-scratch nnU-Net ResEnc-L baseline. Lastly, we c) publish the code of our pre-training and fine-tuning frameworks and provide the pre-trained models created during the benchmarking process to facilitate rapid adoption and reproduction.
Tassilo Wald, Constantin Ulrich, Jonathan Suprijadi, Sebastian Ziegler, Michal Nohel, Robin Peretzke, Gregor Köhler, Klaus H. Maier-Hein
ICCV2
2025 The Missing Piece: A Case for Pre-training in 3D Medical Object Detection
Katharina Eckstein, Constantin Ulrich, Michael Baumgartner 0001, Jessica Kächele, Dimitrios Bounias, Tassilo Wald, Ralf Floca, Klaus H. Maier-Hein
MICCAI (4)2
2025 SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001
Medical Image Anal.24
2024 Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures
Yannick Kirchhoff, Maximilian Rokuss, Saikat Roy, Balint Kovacs, Constantin Ulrich, Tassilo Wald, Maximilian Zenk, Philipp Vollmuth, Jens Kleesiek, Fabian Isensee, Klaus H. Maier-Hein
ECCV (77)5
2024 nnU-Net Revisited: A Call for Rigorous Validation in 3D Medical Image Segmentation
Fabian Isensee, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger
MICCAI (9)3
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS14
2024 Decoupling Semantic Similarity from Spatial Alignment for Neural Networks
abstract
What representation do deep neural networks learn? How similar are images to each other for neural networks? Despite the overwhelming success of deep learning methods key questions about their internal workings still remain largely unanswered, due to their internal high dimensionality and complexity. To address this, one approach is to measure the similarity of activation responses to various inputs. Representational Similarity Matrices (RSMs) distill this similarity into scalar values for each input pair. These matrices encapsulate the entire similarity structure of a system, indicating which input lead to similar responses. While the similarity between images is ambiguous, we argue that the spatial location of semantic objects does neither influence human perception nor deep learning classifiers. Thus this should be reflected in the definition of similarity between image responses for computer vision systems. Revisiting the established similarity calculations for RSMs we expose their sensitivity to spatial alignment. In this paper we propose to solve this through _semantic RSMs_, which are invariant to spatial permutation. We measure semantic similarity between input responses by formulating it as a set-matching problem. Further, we quantify the superiority of _semantic_ RSMs over _spatio-semantic_ RSMs through image retrieval and by comparing the similarity between representations to the similarity between predicted class probabilities.
Tassilo Wald, Constantin Ulrich, Priyank Jaini, Gregor Köhler, David Zimmerer, Stefan Denner, Fabian Isensee, Michael Baumgartner 0001, Klaus H. Maier-Hein
NeurIPS2
2024 RecycleNet: Latent Feature Recycling Leads to Iterative Decision Refinement
abstract
Despite the remarkable success of deep learning systems over the last decade, a key difference still remains between neural network and human decision-making: As humans, we can not only form a decision on the spot, but also ponder, revisiting an initial guess from different angles, distilling relevant information, arriving at a better decision. Here, we propose RecycleNet, a latent feature recycling method, instilling the pondering capability for neural networks to refine initial decisions over a number of recycling steps, where outputs are fed back into earlier network layers in an iterative fashion. This approach makes minimal assumptions about the neural network architecture and thus can be implemented in a wide variety of contexts. Using medical image segmentation as the evaluation environment, we show that latent feature recycling enables the network to iteratively refine initial predictions even beyond the iterations seen during training, converging towards an improved decision. We evaluate this across a variety of segmentation benchmarks and show consistent improvements even compared with top-performing segmentation methods. This allows trading increased computation time for improved performance, which can be beneficial, especially for safety-critical applications.
Gregor Köhler, Tassilo Wald, Constantin Ulrich, David Zimmerer, Paul F. Jaeger, Jörg K. H. Franke, Simon Kohl, Fabian Isensee, Klaus H. Maier-Hein
WACV3
2023 MedNeXt: Transformer-Driven Scaling of ConvNets for Medical Image Segmentation
Saikat Roy, Gregor Köhler, Constantin Ulrich, Michael Baumgartner 0001, Jens Petersen, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein
MICCAI (4)3
2023 MultiTalent: A Multi-dataset Approach to Medical Image Segmentation
Constantin Ulrich, Fabian Isensee, Tassilo Wald, Maximilian Zenk, Michael Baumgartner 0001, Klaus H. Maier-Hein
MICCAI (3)1