VLDB 2026 Research / reviewers in the wild / expert
Jens Kleesiek
dblp:63/7927
· DBLP profile ↗
26ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0001-8686-0682ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 12 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can one model fit all? Evaluating foundation models for time series forecasting across clinical medicineabstractArtificial intelligence (AI) is increasingly integrated into clinical medicine, with foundation models emerging as an alternative to task-specific models for forecasting longitudinal healthcare data. These models, pre-trained on large datasets, promise broad applicability across clinical domains, yet their real-world performance and generalizability remain underexplored. To address this gap, we evaluated foundation and task-specific models across diverse clinical use cases, focusing on zero-shot performance, cross-hospital transportability, the impact of fine-tuning, and potential clinical implications. We used data from University Hospital Essen, Germany, two nearby regional hospitals, and the MIMIC-IV database to define six clinical time series use cases, including forecasting of vital signs, laboratory values, and hospital capacity. Transformer-based foundation models were compared in zero-shot and fine-tuned settings to task-specific approaches, including neural networks, gradient boosting, AutoML ensembles, and statistical models. We also assessed predictive value for guiding treatment decisions by dichotomizing forecasts. Zero-shot foundation models frequently approached the performance of optimized task-specific models. Fine-tuning further improved performance, with Chronos and TimesFM ranking among the best-performing models 19 and 18 times, respectively, compared to 21 times for AutoML ensembles. Foundation models showed superior transportability across hospital settings and patient populations. However, variations in forecasting strategies influenced their positive and negative predictive values in clinical decision-making contexts. These results suggest that foundation models are viable for clinical time series forecasting, particularly where generalizability is crucial. Their flexibility and zero-shot capabilities reduce the need for retraining, potentially lowering barriers to adoption and challenging the role of domain-specific models in clinical practice. Gernot Pucher, Amin Dada, Aurel Agbodoyetin, Felix Nensa, Martin Schuler, Hans Christian Reinhardt, Jens Kleesiek, Christopher Martin Sauer |
Artif. Intell. Medicine | 7 |
| 2025 | Little Is Enough: Boosting Privacy by Sharing Only Hard Labels in Federated Semi-Supervised LearningabstractIn many critical applications, sensitive data is inherently distributed and cannot be centralized due to privacy concerns. A wide range of federated learning approaches have been proposed to train models locally at each client without sharing their sensitive data, typically by exchanging model parameters, or probabilistic predictions (soft labels) on a public dataset or a combination of both. However, these methods still disclose private information and restrict local models to those that can be trained using gradient-based methods. We propose a federated co-training (FEDCT) approach that improves privacy by sharing only definitive (hard) labels on a public unlabeled dataset. Clients use a consensus of these shared labels as pseudo-labels for local training. This federated co-training approach empirically enhances privacy without compromising model quality. In addition, it allows the use of local models that are not suitable for parameter aggregation in traditional federated learning, such as gradient-boosted decision trees, rule ensembles, and random forests. Furthermore, we observe that FEDCT performs effectively in federated fine-tuning of large language models, where its pseudo-labeling mechanism is particularly beneficial. Empirical evaluations and theoretical analyses suggest its applicability across a range of federated learning scenarios. Amr Abourayya, Jens Kleesiek, Kanishka Rao, Erman Ayday, R. Bharat Rao, Geoffrey I. Webb, Michael Kamp |
AAAI | 2 |
| 2025 | Every Component Counts: Rethinking the Measure of Success for Medical Semantic Segmentation in Multi-Instance Segmentation TasksabstractWe present Connected-Component (CC)-Metrics, a novel semantic segmentation evaluation protocol, targeted to align existing semantic segmentation metrics to a multi-instance detection scenario in which each connected component matters. We motivate this setup in the common medical scenario of semantic metastases segmentation in a full-body PET/CT. We show how existing semantic segmentation metrics suffer from a bias towards larger connected components contradicting the clinical assessment of scans in which tumor size and clinical relevance are uncorrelated. To rebalance existing segmentation metrics, we propose to evaluate them on a per-component basis thus giving each tumor the same weight irrespective of its size. To match predictions to ground-truth segments, we employ a proximity-based matching criterion, evaluating common metrics locally at the component of interest. Using this approach, we break free of biases introduced by large metastasis for overlap-based metrics such as Dice or Surface Dice. CC-Metrics also improves distance-based metrics such as Hausdorff Distances which are uninformative for small changes that do not influence the maximum or 95th percentile, and avoids pitfalls introduced by directly combining counting-based metrics with overlap-based metrics as it is done in Panoptic Quality. Alexander Jaus, Constantin Seibold, Simon Reiß, Zdravko Marinov, Zeling Ye, Stefan Krieg 0002, Jens Kleesiek, Rainer Stiefelhagen |
AAAI | 8 |
| 2025 | A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks AlignmentabstractHigh computation costs and latency of large language models such as GPT-4 have limited their deployment in clinical settings. Small language models (SLMs) offer a cost-effective alternative, but their limited capacity requires biomedical domain adaptation, which remains challenging. An additional bottleneck is the unavailability and high sensitivity of clinical data. To address these challenges, we propose a novel framework for adapting SLMs into high-performing clinical models. We introduce the MediPhi collection of 3.8B-parameter SLMs developed with our novel framework: pre-instruction tuning of experts on relevant medical and clinical corpora (PMC, Medical Guideline, MedWiki, etc.), model merging, and clinical-tasks alignment. To cover most clinical tasks, we extended the CLUE benchmark to CLUE+, doubling its size. Our expert models deliver relative improvements on this benchmark over the base model without any task-specific fine-tuning: 64.3% on medical entities, 49.5% on radiology reports, and 44% on ICD-10 coding (outperforming GPT-4-0125 by 14%). We unify the expert models into MediPhi via model merging, preserving gains across benchmarks. Furthermore, we built the MediFlow collection, a synthetic dataset of 2.5 million high-quality instructions on 14 medical NLP tasks, 98 fine-grained document types, and JSON format support. Alignment of MediPhi using supervised fine-tuning and direct preference optimization achieves further gains of 18.9% on average. Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila |
ACL (1) | 9 |
| 2025 | Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via GrokkingabstractNeural collapse, i.e., the emergence of highly symmetric, class-wise clustered representations, is frequently observed in deep networks and is often assumed to reflect or enable generalization. In parallel, flatness of the loss landscape has been theoretically and empirically linked to generalization. Yet, the causal role of either phenomenon remains unclear: Are they prerequisites for generalization, or merely by-products of training dynamics? We disentangle these questions using grokking, a training regime in which memorization precedes generalization, allowing us to temporally separate generalization from training dynamics and we find that while both neural collapse and relative flatness emerge near the onset of generalization, only flatness consistently predicts it. Models encouraged to collapse or prevented from collapsing generalize equally well, whereas models regularized away from flat solutions exhibit delayed generalization, resembling grokking, even in architectures and datasets where it does not typically occur. Furthermore, we show theoretically that neural collapse leads to relative flatness under classical assumptions, explaining their empirical co-occurrence. Our results support the view that relative flatness is a potentially necessary and more fundamental property for generalization, and demonstrate how grokking can serve as a powerful probe for isolating its geometric underpinnings. Linara Adilova, Henning Petzka, Jens Kleesiek, Michael Kamp |
NeurIPS | 4 |
| 2025 | Real-world federated learning in radiology: hurdles to overcome and benefits to gainabstractOBJECTIVE: Federated Learning (FL) enables collaborative model training while keeping data locally. Currently, most FL studies in radiology are conducted in simulated environments due to numerous hurdles impeding its translation into practice. The few existing real-world FL initiatives rarely communicate specific measures taken to overcome these hurdles. To bridge this significant knowledge gap, we propose a comprehensive guide for real-world FL in radiology. Minding efforts to implement real-world FL, there is a lack of comprehensive assessments comparing FL to less complex alternatives in challenging real-world settings, which we address through extensive benchmarking. MATERIALS AND METHODS: We developed our own FL infrastructure within the German Radiological Cooperative Network (RACOON) and demonstrated its functionality by training FL models on lung pathology segmentation tasks across six university hospitals. Insights gained while establishing our FL initiative and running the extensive benchmark experiments were compiled and categorized into the guide. RESULTS: The proposed guide outlines essential steps, identified hurdles, and implemented solutions for establishing successful FL initiatives conducting real-world experiments. Our experimental results prove the practical relevance of our guide and show that FL outperforms less complex alternatives in all evaluation scenarios. DISCUSSION AND CONCLUSION: Our findings justify the efforts required to translate FL into real-world applications by demonstrating advantageous performance over alternative approaches. Additionally, they emphasize the importance of strategic organization, robust management of distributed data and infrastructure in real-world settings. With the proposed guide, we are aiming to aid future FL researchers in circumventing pitfalls and accelerating translation of FL into radiological applications. Markus Bujotzek, Ünal Akünal, Stefan Denner, Peter Neher, Maximilian Zenk, Eric Frodl, Astha Jaiswal, Moon S. Kim 0002, Nicolai R. Krekiehn, Manuel Nickel, Richard Ruppel, Marcus Both, Felix Doellinger, Marcel Opitz, Thorsten Persigehl, Jens Kleesiek, Tobias Penzkofer, Klaus H. Maier-Hein, Andreas Bucher, Rickmer Braren |
J. Am. Medical Informatics Assoc. | 16 |
| 2025 | Evaluating the effectiveness of biomedical fine-tuning for large language models on clinical tasksabstractOBJECTIVES: Large language models (LLMs) have shown potential in biomedical applications, leading to efforts to fine-tune them on domain-specific data. However, the effectiveness of this approach remains unclear. This study aims to critically evaluate the performance of biomedically fine-tuned LLMs against their general-purpose counterparts across a range of clinical tasks. MATERIALS AND METHODS: We evaluated the performance of biomedically fine-tuned LLMs against their general-purpose counterparts on clinical case challenges from NEJM and JAMA, and on multiple clinical tasks, such as information extraction, document summarization and clinical coding. We used a diverse set of benchmarks specifically chosen to be outside the likely fine-tuning datasets of biomedical models, ensuring a fair assessment of generalization capabilities. RESULTS: Biomedical LLMs generally underperformed compared to general-purpose models, especially on tasks not focused on probing medical knowledge. While on the case challenges, larger biomedical and general-purpose models showed similar performance (eg, OpenBioLLM-70B: 66.4% vs Llama-3-70B-Instruct: 65% on JAMA), smaller biomedical models showed more pronounced underperformance (OpenBioLLM-8B: 30% vs Llama-3-8B-Instruct: 64.3% on NEJM). Similar trends appeared across CLUE benchmarks, with general-purpose models often achieving higher scores in text generation, question answering, and coding. Notably, biomedical LLMs also showed a higher tendency to hallucinate. DISCUSSION: Our findings challenge the assumption that biomedical fine-tuning inherently improves LLM performance, as general-purpose models consistently performed better on unseen medical tasks. Retrieval-augmented generation may offer a more effective strategy for clinical adaptation. CONCLUSION: Fine-tuning LLMs on biomedical data may not yield the anticipated benefits. Alternative approaches, such as retrieval augmentation, should be further explored for effective and reliable clinical integration of LLMs. Felix J. Dorfner, Amin Dada, Felix Busch, Marcus R. Makowski, Daniel Truhn, Jens Kleesiek, Madhumita Sushil, Lisa Adams, Keno Bressem |
J. Am. Medical Informatics Assoc. | 7 |
| 2025 | FHIR-Former: enhancing clinical predictions through Fast Healthcare Interoperability Resources and large language modelsabstractOBJECTIVE: To address the challenges of data heterogeneity and manual feature engineering in clinical predictive modeling, we introduce FHIR-Former, an open-source framework integrating Fast Healthcare Interoperability Resources (FHIR) with large language models (LLMs) to automate and standardize clinical prediction tasks. MATERIALS AND METHODS: FHIR-Former dynamically processes structured (eg, lab results, medications) and unstructured (eg, clinical notes) data from FHIR resources. The pipeline supports multiple classification tasks, including 30-day readmission, imaging study prediction, and ICD code classification. Leveraging open-source LLMs (GeBERTa), we trained models on 1.1 million data points across ten FHIR resources using retrospective inpatient data (2018-2024). Hyperparameters were optimized via Bayesian methods, and outputs were mapped to FHIR RiskAssessment resources for interoperability. RESULTS: FHIR-Former achieved an F1-score of 70.7% and accuracy of 72.9% for 30-day readmission, 51.8% F1-score (88.1% accuracy) for mortality prediction, and 61% macro F1-score for imaging study classification. The ICD code prediction model attained 94% accuracy. Performance demonstrated promising performance for readmission and showed scalability across tasks without manual feature engineering. DISCUSSION: FHIR-Former eliminates institution-specific preprocessing by adapting to diverse FHIR implementations, enabling seamless integration of multimodal data. Its configurable architecture outperformed prior frameworks reliant on static inputs or limited to unstructured text. Real-time risk scores embedded in FHIR servers enhance clinical workflows without disrupting existing practices. CONCLUSION: By harmonizing FHIR standardization with LLM flexibility, FHIR-Former advances scalable, interoperable predictive modeling in healthcare. The open-source framework facilitates automation, improves resource allocation, and supports personalized decision-making, bridging gaps between AI innovation and clinical practice. Merlin Engelke, Giulia Baldini 0001, Jens Kleesiek, Felix Nensa, Amin Dada |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Comprehensive Study on German Language Models for Clinical and Biomedical Text UnderstandingabstractRecent advances in natural language processing (NLP) can be largely attributed to the advent of pre-trained language models such as BERT and RoBERTa. While these models demonstrate remarkable performance on general datasets, they can struggle in specialized domains such as medicine, where unique domain-specific terminologies, domain-specific abbreviations, and varying document structures are common. This paper explores strategies for adapting these models to domain-specific requirements, primarily through continuous pre-training on domain-specific data. We pre-trained several German medical language models on 2.4B tokens derived from translated public English medical data and 3B tokens of German clinical data. The resulting models were evaluated on various German downstream tasks, including named entity recognition (NER), multi-label classification, and extractive question answering. Our results suggest that models augmented by clinical and translation-based pre-training typically outperform general domain models in medical contexts. We conclude that continuous pre-training has demonstrated the ability to match or even exceed the performance of clinical models trained from scratch. Furthermore, pre-training on clinical data or leveraging translated texts have proven to be reliable methods for domain adaptation in medical NLP tasks. Ahmad Idrissi-Yaghir, Amin Dada, Henning Schäfer, Kamyar Arzideh, Giulia Baldini 0001, Jan Trienes, Max Hasin, Jeanette Bewersdorff, Cynthia Sabrina Schmidt, Marie Bauer, Kaleb E. Smith, Jiang Bian 0001, Yonghui Wu 0001, Jörg Schlötterer, Torsten Zesch, Peter A. Horn, Christin Seifert, Felix Nensa, Jens Kleesiek, Christoph M. Friedrich |
LREC/COLING | 19 |
| 2024 | Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures
Yannick Kirchhoff, Maximilian Rokuss, Saikat Roy, Balint Kovacs, Constantin Ulrich, Tassilo Wald, Maximilian Zenk, Philipp Vollmuth, Jens Kleesiek, Fabian Isensee, Klaus H. Maier-Hein |
ECCV (77) | 9 |
| 2024 | Towards Unifying Anatomy Segmentation: Automated Generation of a Full-Body CT DatasetabstractIn this paper, we present a method for generating automated anatomy segmentation datasets using a sequential process that involves nnU-Net-based pseudo-labeling and anatomy-guided pseudo-label refinement. By combining various fragmented knowledge bases, we generate a dataset of whole-body CT scans with 142 voxel-level labels for 533 volumes providing comprehensive anatomical coverage. We validate its usefulness via Human expert evaluation and medical validity. This dataset enables the analysis of whole-body anatomy segmentation for cancer patients. Besides the DAP Atlas dataset, we release our trained anatomy segmentation models capable of predicting 142 anatomical structures on CT data. Alexander Jaus, Constantin Seibold, Kelsey Hermann, Negar Shahamiri, Alexandra Walter, Kristina Giske, Johannes Haubold, Jens Kleesiek, Rainer Stiefelhagen |
ICIP | 8 |
| 2024 | Anatomy-Guided Pathology Segmentation
Alexander Jaus, Constantin Seibold, Simon Reiß, Lukas Heine, Anton Schily, Moon S. Kim 0002, Fin Hendrik Bahnsen, Ken Herrmann, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (8) | 10 |
| 2024 | GAN-based generation of realistic 3D volumetric data: A systematic review and taxonomyabstractWith the massive proliferation of data-driven algorithms, such as deep learning-based approaches, the availability of high-quality data is of great interest. Volumetric data is very important in medicine, as it ranges from disease diagnoses to therapy monitoring. When the dataset is sufficient, models can be trained to help doctors with these tasks. Unfortunately, there are scenarios where large amounts of data is unavailable. For example, rare diseases and privacy issues can lead to restricted data availability. In non-medical fields, the high cost of obtaining enough high-quality data can also be a concern. A solution to these problems can be the generation of realistic synthetic data using Generative Adversarial Networks (GANs). The existence of these mechanisms is a good asset, especially in healthcare, as the data must be of good quality, realistic, and without privacy issues. Therefore, most of the publications on volumetric GANs are within the medical domain. In this review, we provide a summary of works that generate realistic volumetric synthetic data using GANs. We therefore outline GAN-based methods in these areas with common architectures, loss functions and evaluation metrics, including their advantages and disadvantages. We present a novel taxonomy, evaluations, challenges, and research opportunities to provide a holistic overview of the current state of volumetric GANs. Jianning Li 0002, Kelsey L. Pomykala, Jens Kleesiek, Victor Alves, Jan Egger |
Medical Image Anal. | 4 |
| 2024 | Corrigendum to: GAN-based generation of realistic 3D volumetric data: A systematic review and taxonomy [Medical Image Analysis 93 (2024)]
Jianning Li 0002, Kelsey L. Pomykala, Jens Kleesiek, Victor Alves, Jan Egger |
Medical Image Anal. | 4 |
| 2024 | CellViT: Vision Transformers for precise cell segmentation and classificationabstractNuclei detection and segmentation in hematoxylin and eosin-stained (H&E) tissue images are important clinical tasks and crucial for a wide range of applications. However, it is a challenging task due to nuclei variances in staining and size, overlapping boundaries, and nuclei clustering. While convolutional neural networks have been extensively used for this task, we explore the potential of Transformer-based networks in combination with large scale pre-training in this domain. Therefore, we introduce a new method for automated instance segmentation of cell nuclei in digitized tissue samples using a deep learning architecture based on Vision Transformer called CellViT. CellViT is trained and evaluated on the PanNuke dataset, which is one of the most challenging nuclei instance segmentation datasets, consisting of nearly 200,000 annotated nuclei into 5 clinically important classes in 19 tissue types. We demonstrate the superiority of large-scale in-domain and out-of-domain pre-trained Vision Transformers by leveraging the recently published Segment Anything Model and a ViT-encoder pre-trained on 104 million histological image patches - achieving state-of-the-art nuclei detection and instance segmentation performance on the PanNuke dataset with a mean panoptic quality of 0.50 and an F1-detection score of 0.83. The code is publicly available at https://github.com/TIO-IKIM/CellViT. Fabian Hörst, Moritz Rempe, Lukas Heine, Constantin Seibold, Julius Keyl, Giulia Baldini 0001, Selma Ugurel, Jens T. Siveke, Barbara Grünwald, Jan Egger, Jens Kleesiek |
Medical Image Anal. | 11 |
| 2024 | Deep Interactive Segmentation of Medical Images: A Systematic Review and TaxonomyabstractInteractive segmentation is a crucial research area in medical image analysis aiming to boost the efficiency of costly annotations by incorporating human feedback. This feedback takes the form of clicks, scribbles, or masks and allows for iterative refinement of the model output so as to efficiently guide the system towards the desired behavior. In recent years, deep learning-based approaches have propelled results to a new level causing a rapid growth in the field with 121 methods proposed in the medical imaging domain alone. In this review, we provide a structured overview of this emerging field featuring a comprehensive taxonomy, a systematic review of existing methods, and an in-depth analysis of current practices. Based on these contributions, we discuss the challenges and opportunities in the field. For instance, we find that there is a severe lack of comparison across methods which needs to be tackled by standardized baselines and benchmarks. Zdravko Marinov, Paul F. Jaeger, Jan Egger, Jens Kleesiek, Rainer Stiefelhagen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Model Generalizability Investigation for GFCE-MRI Synthesis in NPC Radiotherapy Using Multi-Institutional Patient-Based Data NormalizationabstractRecently, deep learning has been demonstrated to be feasible in eliminating the use of gadoliniumbased contrast agents (GBCAs) through synthesizing gadolinium-free contrast-enhanced MRI (GFCE-MRI) from contrast-free MRI sequences, providing the community with an alternative to get rid of GBCAs-associated safety issues in patients. Nevertheless, generalizability assessment of the GFCE-MRI model has been largely challenged by the high inter-institutional heterogeneity of MRI data, on top of the scarcity of multi-institutional data itself. Although various data normalization methods have been adopted to address the heterogeneity issue, it has been limited to single-institutional investigation and there is no standard normalization approach presently. In this study, we aimed at investigating generalizability of GFCE-MRI model using data from seven institutions by manipulating heterogeneity of MRI data under five popular normalization approaches. Three state-of-the-art neural networks were applied to map from T1-weighted and T2-weighted MRI to contrast-enhanced MRI (CE-MRI) for GFCE-MRI synthesis in patients with nasopharyngeal carcinoma. MRI data from three institutions were used separately to generate three uni-institution models and jointly for a tri-institution model. The five normalization methods were applied to normalize the data of each model. MRI data from the remaining four institutions served as external cohorts for model generalizability assessment. Quality of GFCE-MRI was quantitatively evaluated against ground-truth CE-MRI using mean absolute error (MAE) and peak signal-to-noise ratio(PSNR). Results showed that performance of all uni-institution models remarkably dropped on the external cohorts. By contrast, model trained using multi-institutional data with Z-Score normalization yielded the best model generalizability improvement. Wen Li 0010, Saikit Lam, Yinghui Wang 0003, Tian Li 0012, Jens Kleesiek, Andy Lai-Yin Cheung, Ying Sun 0015, Francis Kar-ho Lee, Kwok-hung Au, Victor Ho-fun Lee, Jing Cai 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Guiding the Guidance: A Comparative Analysis of User Guidance Signals for Interactive Segmentation of Volumetric Images
Zdravko Marinov, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (3) | 3 |
| 2023 | The HoloLens in medicine: A systematic review and taxonomyabstractThe HoloLens (Microsoft Corp., Redmond, WA), a head-worn, optically see-through augmented reality (AR) display, is the main player in the recent boost in medical AR research. In this systematic review, we provide a comprehensive overview of the usage of the first-generation HoloLens within the medical domain, from its release in March 2016, until the year of 2021. We identified 217 relevant publications through a systematic search of the PubMed, Scopus, IEEE Xplore and SpringerLink databases. We propose a new taxonomy including use case, technical methodology for registration and tracking, data sources, visualization as well as validation and evaluation, and analyze the retrieved publications accordingly. We find that the bulk of research focuses on supporting physicians during interventions, where the HoloLens is promising for procedures usually performed without image guidance. However, the consensus is that accuracy and reliability are still too low to replace conventional guidance systems. Medical students are the second most common target group, where AR-enhanced medical simulators emerge as a promising technology. While concerns about human-computer interactions, usability and perception are frequently mentioned, hardly any concepts to overcome these issues have been proposed. Instead, registration and tracking lie at the core of most reviewed publications, nevertheless only few of them propose innovative concepts in this direction. Finally, we find that the validation of HoloLens applications suffers from a lack of standardized and rigorous evaluation protocols. We hope that this review can advance medical AR research by identifying gaps in the current literature, to pave the way for novel, innovative directions and translation into the medical routine. Christina Schwarz-Gsaxner, Jianning Li 0002, Antonio Pepe 0003, Jens Kleesiek, Dieter Schmalstieg, Jan Egger |
Medical Image Anal. | 5 |
| 2023 | Towards clinical applicability and computational efficiency in automatic cranial implant design: An overview of the AutoImplant 2021 cranial implant design challenge
Jianning Li 0002, David Gage Ellis, Oldrich Kodym, Laurèl Rauschenbach, Christoph Rieß, Ulrich Sure, Karsten H. Wrede, Carlos M. Alvarez, Marek Wodzinski, Mateusz Daniol, Daria Hemmerling, Hamza Mahdi, Allison Clement, Evan Kim, Zachary Fishman, Cari M. Whyne, James G. Mainprize, Michael R. Hardisty, Shashwat Pathak, Chitimireddy Sindhura, Rama Krishna Sai S. Gorthi, Degala Venkata Kiran, Subrahmanyam Gorthi, Artem Kroviakov, Antonio Pepe 0003, Christina Schwarz-Gsaxner, Adam Herout, Victor Alves, Michal Spanel, Michele R. Aizenberg, Jens Kleesiek, Jan Egger |
Medical Image Anal. | 36 |
| 2022 | Reference-Guided Pseudo-Label Generation for Medical Semantic SegmentationabstractProducing densely annotated data is a difficult and tedious task for medical imaging applications. To address this problem, we propose a novel approach to generate supervision for semi-supervised semantic segmentation. We argue that visually similar regions between labeled and unlabeled images likely contain the same semantics and therefore should share their label. Following this thought, we use a small number of labeled images as reference material and match pixels in an unlabeled image to the semantic of the best fitting pixel in a reference set. This way, we avoid pitfalls such as confirmation bias, common in purely prediction-based pseudo-labeling. Since our method does not require any architectural changes or accompanying networks, one can easily insert it into existing frameworks. We achieve the same performance as a standard fully supervised model on X-ray anatomy segmentation, albeit using 95% fewer labeled images. Aside from an in-depth analysis of different aspects of our proposed method, we further demonstrate the effectiveness of our reference-guided learning paradigm by comparing our approach against existing methods for retinal fluid segmentation with competitive performance as we improve upon recent work by up to 15% mean IoU. Constantin Seibold, Simon Reiß, Jens Kleesiek, Rainer Stiefelhagen |
AAAI | 3 |
| 2022 | Detailed Annotations of Chest X-Rays via CT Projection for Report Understanding
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Matthias A. Fink, Victoria Mayer, Jan Sellner, Moon S. Kim 0002, Klaus H. Maier-Hein, Jens Kleesiek, Rainer Stiefelhagen |
BMVC | 9 |
| 2022 | Breaking with Fixed Set Pathology Recognition Through Report-Guided Contrastive Training
Constantin Seibold, Simon Reiß, M. Saquib Sarfraz, Rainer Stiefelhagen, Jens Kleesiek |
MICCAI (5) | 5 |
| 2020 | Self-guided Multiple Instance Learning for Weakly Supervised Disease Classification and Localization in Chest Radiographs
Constantin Seibold, Jens Kleesiek, Heinz-Peter Schlemmer, Rainer Stiefelhagen |
ACCV (5) | 2 |
| 2016 | DALSA: Domain Adaptation for Supervised Learning From Sparsely Annotated MR ImagesabstractWe propose a new method that employs transfer learning techniques to effectively correct sampling selection errors introduced by sparse annotations during supervised learning for automated tumor segmentation. The practicality of current learning-based automated tissue classification approaches is severely impeded by their dependency on manually segmented training databases that need to be recreated for each scenario of application, site, or acquisition setup. The comprehensive annotation of reference datasets can be highly labor-intensive, complex, and error-prone. The proposed method derives high-quality classifiers for the different tissue classes from sparse and unambiguous annotations and employs domain adaptation techniques for effectively correcting sampling selection errors introduced by the sparse sampling. The new approach is validated on labeled, multi-modal MR images of 19 patients with malignant gliomas and by comparative analysis on the BraTS 2013 challenge data sets. Compared to training on fully labeled data, we reduced the time for labeling and training by a factor greater than 70 and 180 respectively without sacrificing accuracy. This dramatically eases the establishment and constant extension of large annotated databases in various scenarios and imaging setups and thus represents an important step towards practical applicability of learning-based approaches in tissue classification. Michael Götz, Christian Weber 0001, Franciszek Binczyk, Joanna Polanska, Rafal Tarnawski, Barbara Bobek-Billewicz, Ullrich Köthe, Jens Kleesiek, Bram Stieltjes, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 8 |
| 2012 | What do Objects Feel Like? - Active Perception for a Humanoid Robot
Jens Kleesiek, Stephanie Badde, Stefan Wermter, Andreas K. Engel |
ICAART (1) | 1 |