VLDB 2026 Research / reviewers in the wild / expert
Bernhard Kainz
dblp:76/5562
· DBLP profile ↗
91ranked-venue papers
8as first author
47since 2021 · last 2026
0000-0002-7813-5023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 7 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 59 · 4 first-author · 30 since 2021Artificial intelligence and machine learning · 14 · 13 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diff-Def: Diffusion-Generated Deformation Fields for Conditional AtlasesabstractAnatomical atlases are widely used for population studies and analysis. Conditional atlases target a specific sub-population defined via certain conditions, such as demographics or pathologies, and allow for the investigation of fine-grained anatomical differences like morphological changes associated with ageing or disease. Existing approaches use either registration-based methods that are often unable to handle large anatomical variations or generative adversarial models, which are challenging to train since they can suffer from training instabilities. Instead of generating atlases directly in as intensities, we propose using latent diffusion models to generate deformation fields, which transform a general population atlas into one representing a specific sub-population. Our approach ensures structural integrity, enhances interpretability and avoids hallucinations that may arise during direct image synthesis by generating this deformation field and regularising it using a neighbourhood of images. We compare our method to several state-of-the-art atlas generation methods using brain MR images from the UK Biobank. Our method generates highly realistic atlases with smooth transformations and high anatomical fidelity, outperforming existing baselines. We demonstrate the quality of these atlases through comprehensive evaluations, including quantitative metrics for anatomical accuracy, perceptual similarity, and qualitative analyses displaying the consistency and realism of the generated atlases. Sophie Starck, Vasiliki Sideri-Lampretsa, Bernhard Kainz, Martin J. Menten, Tamara T. Mueller, Daniel Rueckert |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Image Generation Diversity Issues and How to Tame ThemabstractGenerative methods have reached a level of quality that is almost indistinguishable from real data. However, while individual samples may appear unique, generative models often exhibit limitations in covering the full data distribution. Unlike quality issues, diversity problems within generative models are not easily detected by simply observing single images or generated datasets, which means we need a specific measure to assess the diversity of these models. In this paper, we draw attention to the current lack of diversity in generative models and the inability of common metrics to measure this. We achieve this by framing diversity as an image retrieval problem, where we measure how many real images can be retrieved using synthetic data as queries. This yields the Image Retrieval Score (IRS), an interpretable, hyperparameter-free metric that quantifies the diversity of a generative model’s output. IRS requires only a subset of synthetic samples and provides a statistical measure of confidence. Our experiments indicate that current feature extractors commonly used in generative model assessment are inadequate for evaluating diversity effectively. Consequently, we perform an extensive search for the best feature extractors to assess diversity. Evaluation reveals that current diffusion models converge to limited subsets of the real distribution, with no current state-of-the-art models superpassing 77% of the diversity of the training data. To address this limitation, we introduce Diversity-Aware Diffusion Models (DiADM), a novel approach that improves diversity of unconditional diffusion models without loss of image quality. We do this by disentangling diversity from image quality by using a diversity aware module that uses pseudo-unconditional features as input. We provide a Python package offering unified feature extraction and metric computation to further facilitate the evaluation of generative models https://github.com/MischaD/beyondfid. Mischa Dombrowski, Sarah Cechnicka, Hadrien Reynaud, Bernhard Kainz |
CVPR | 5 |
| 2025 | SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
Stathis Galanakis, Alexander Lattas, Stylianos Moschoglou, Bernhard Kainz, Stefanos Zafeiriou |
ICCV | 4 |
| 2025 | Last Layer Laplacian Pseudocoresets for Robust Medical Image Analysis
Franciskus Xaverius Erick, Johanna P. Müller, Zhe Li 0025, Bernhard Kainz |
MICCAI (8) | 4 |
| 2025 | Mesh4D: A Motion-Aware Multi-view Variational Autoencoder for 3D+t Mesh Reconstruction
Mengyun Qiao, Qiang Ma 0004, Liu Li 0001, Bernhard Kainz, Declan P. O'Regan, Paul M. Matthews, Steven A. Niederer, Wenjia Bai |
MICCAI (16) | 6 |
| 2025 | Multi-agent Reasoning for Cardiovascular Imaging Phenotype Analysis
Mengyun Qiao, Chengqi Zang, Steven A. Niederer, Paul M. Matthews, Wenjia Bai, Bernhard Kainz |
MICCAI (1) | 7 |
| 2025 | NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRIabstractIn many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining.Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging.However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use.We therefore present NOVA, a challenging, real-life _evaluation-only_ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an _extreme_ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B) reveal substantial performance drops, with approximately a 65\% gap in localisation compared to natural-image benchmarks and 40\% and 20\% gaps in captioning and reasoning, respectively, compared to resident radiologists. Therefore, NOVA establishes a testbed for advancing models that can detect, localize, and reason about truly unknown anomalies. Cosmin Bercea, Philipp Raffler, Evamaria O. Riedel, Lena Schmitzer, Angela Kurz, Felix Bitzer, Paula Roßmüller, Julian Canisius, Mirjam L. Beyrle, Che Liu 0002, Wenjia Bai, Bernhard Kainz, Julia A. Schnabel, Benedikt Wiestler |
NeurIPS | 13 |
| 2025 | Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical ImagingabstractRecent progress in vision-language modeling for 3D medical imaging has been fueled by large-scale computed tomography (CT) corpora with paired free-text reports, stronger architectures, and powerful pretrained models. This has enabled applications such as automated report generation and text-conditioned 3D image synthesis. Yet, current approaches struggle with high-resolution, long-sequence volumes: contrastive pretraining often yields vision encoders that are misaligned with clinical language, and slice-wise tokenization blurs fine anatomy, reducing diagnostic performance on downstream tasks. We introduce BTB3D (Better Tokens for Better 3D), a causal convolutional encoder-decoder that unifies 2D and 3D training and inference while producing compact, frequency-aware volumetric tokens. A three-stage training curriculum enables (i) local reconstruction, (ii) overlapping-window tiling, and (iii) long-context decoder refinement, during which the model learns from short slice excerpts yet generalizes to scans exceeding $300$ slices without additional memory overhead. BTB3D sets a new state-of-the-art on two key tasks: it improves BLEU scores and increases clinical F1 by 40\% over CT2Rep, CT-CHAT, and Merlin for report generation; and it reduces FID by 75\% and halves FVD compared to GenerateCT and MedSyn for text-to-CT synthesis, producing anatomically consistent $512\times512\times241$ volumes. These results confirm that precise three-dimensional tokenization, rather than larger language backbones alone, is essential for scalable vision-language modeling in 3D medical imaging. The codebase is available at: https://github.com/ibrahimethemhamamci/BTB3D Ibrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud, Dong Yang 0005, Marc Edgar, Daguang Xu, Bernhard Kainz, Bjoern Menze |
NeurIPS | 9 |
| 2025 | From Self-Check to Consensus: Bayesian Strategic Decoding in Large Language ModelsabstractLarge Language Models exhibit logical inconsistency across multi-turn inference processes, undermining correctness in complex inferential tasks. Challenges arise from ensuring that outputs align with both factual correctness and human intent. Approaches like single-agent reflection and multi-agent debate frequently prioritize consistency, but at the expense of accuracy.
To address this problem, we propose a novel game-theoretic consensus mechanism that enables LLMs to self-check their outputs during the decoding stage of output generation. Our method models the decoding process as a multistage Bayesian Decoding Game, where strategic interactions dynamically converge to a consensus on the most reliable outputs without human feedback or additional training. Remarkably, our game design allows smaller models to outperform much larger models through game mechanisms (e.g., 78.1 LLaMA13B vs. 76.6 PaLM540B). As a model-agnostic method, our approach consistently improves even the latest models, enhancing DeepSeek-7B's performance on MMLU by 12.4%. Our framework effectively balances correctness and consistency, demonstrating that properly designed game-theoretic mechanisms can significantly enhance the self-verification capabilities of language models across various tasks and model architectures. Chengqi Zang, Bernhard Kainz |
NeurIPS | 3 |
| 2025 | Stochastic latent feature distillation: Enhancing dataset distillation via structured uncertainty modelingabstractAs deep learning models continue to scale in complexity and data size, reducing storage and training costs has become increasingly important. Dataset distillation addresses this challenge by synthesizing a small set of synthetic samples that effectively substitute for the original dataset in downstream tasks. Existing approaches typically rely on matching gradients or features either in pixel space or in the latent space of a pretrained generative model. We propose a novel stochastic distillation method that models the joint distribution of latent features using a low-rank multivariate normal distribution, parameterized by a lightweight neural network. This formulation captures spatial correlations in the feature space, which are then projected into class probability space to generate more diverse and informative predictions. The proposed module integrates seamlessly with existing distillation pipelines. Our method achieves state-of-the-art cross-architecture results, improving test accuracy by up to 7.47% in gradient matching and 35.71% in distribution matching over baselines. • Introduce SLFD, a framework that distills data with stochastic latent features. • Model spatial correlations using a low-rank multivariate distribution. • Achieve robust performance on high-resolution ImageNet-1K subsets. • Demonstrate applicability to medical imaging with strong results. Zhe Li 0025, Sarah Cechnicka, Cheng Ouyang, Katharina Breininger, Peter J. Schüffler, Bernhard Kainz |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | The Developing Human Connectome Project: A fast deep learning-based pipeline for neonatal cortical surface reconstructionabstractThe Developing Human Connectome Project (dHCP) aims to explore developmental patterns of the human brain during the perinatal period. An automated processing pipeline has been developed to extract high-quality cortical surfaces from structural brain magnetic resonance (MR) images for the dHCP neonatal dataset. However, the current implementation of the pipeline requires more than 6.5 h to process a single MRI scan, making it expensive for large-scale neuroimaging studies. In this paper, we propose a fast deep learning (DL) based pipeline for dHCP neonatal cortical surface reconstruction, incorporating DL-based brain extraction, cortical surface reconstruction and spherical projection, as well as GPU-accelerated cortical surface inflation and cortical feature estimation. We introduce a multiscale deformation network to learn diffeomorphic cortical surface reconstruction end-to-end from T2-weighted brain MRI. A fast unsupervised spherical mapping approach is integrated to minimize metric distortions between cortical surfaces and projected spheres. The entire workflow of our DL-based dHCP pipeline completes within only 24 s on a modern GPU, which is nearly 1000 times faster than the original dHCP pipeline. The qualitative assessment demonstrates that for 82.5% of the test samples, the cortical surfaces reconstructed by our DL-based pipeline achieve superior (54.2%) or equal (28.3%) surface quality compared to the original dHCP pipeline. Qiang Ma 0004, Kaili Liang, Liu Li 0001, Saga Masui, Yourong Guo, Chiara Nosarti, Emma C. Robinson, Bernhard Kainz, Daniel Rueckert |
Medical Image Anal. | 8 |
| 2025 | Topology Optimization in Medical Image Segmentation With Fast χ Euler CharacteristicabstractDeep learning-based medical image segmentation techniques have shown promising results when evaluated based on conventional metrics such as the Dice score or Intersection-over-Union. However, these fully automatic methods often fail to meet clinically acceptable accuracy, especially when topological constraints should be observed, e.g., continuous boundaries or closed surfaces. In medical image segmentation, the correctness of a segmentation in terms of the required topological genus sometimes is even more important than the pixel-wise accuracy. Existing topology-aware approaches commonly estimate and constrain the topological structure via the concept of persistent homology (PH). However, these methods are difficult to implement for high dimensional data due to their polynomial computational complexity. To overcome this problem, we propose a novel and fast approach for topology-aware segmentation based on the Euler Characteristic ( $\chi $ ). First, we propose a fast formulation for $\chi $ computation in both 2D and 3D. The scalar $\chi $ error between the prediction and ground-truth serves as the topological evaluation metric. Then we estimate the spatial topology correctness of any segmentation network via a so-called topological violation map, i.e., a detailed map that highlights regions with $\chi $ errors. Finally, the segmentation results from the arbitrary network are refined based on the topological violation maps by a topology-aware correction network. Our experiments are conducted on both 2D and 3D datasets and show that our method can significantly improve topological correctness while preserving pixel-wise segmentation accuracy. Liu Li 0001, Qiang Ma 0004, Cheng Ouyang, Johannes C. Paetzold, Daniel Rueckert, Bernhard Kainz |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Trade-Offs in Fine-Tuned Diffusion Models between Accuracy and InterpretabilityabstractRecent advancements in diffusion models have significantly impacted the trajectory of generative machine learning re-search, with many adopting the strategy of fine-tuning pre-trained models using domain-specific text-to-image datasets. Notably, this method has been readily employed for medical applications, such as X-ray image synthesis, leveraging the plethora of associated radiology reports. Yet, a prevailing concern is the lack of assurance on whether these models genuinely comprehend their generated content. With the evolution of text conditional image generation, these models have grown potent enough to facilitate object localization scrutiny. Our research underscores this advancement in the critical realm of medical imaging, emphasizing the crucial role of interpretability. We further unravel a consequential trade-off between image fidelity – as gauged by conventional metrics – and model interpretability in generative diffusion models. Specifically, the adoption of learnable text encoders when fine-tuning results in diminished interpretability. Our in-depth exploration uncovers the underlying factors responsible for this divergence. Consequently, we present a set of design principles for the development of truly interpretable generative models. Code is available at https://github.com/MischaD/chest-distillation. Mischa Dombrowski, Hadrien Reynaud, Johanna P. Müller, Matthew Baugh, Bernhard Kainz |
AAAI | 5 |
| 2024 | Style-Extracting Diffusion Models for Semi-supervised Histopathology Segmentation
Mathias Öttl, Frauke Wilm, Jana Steenpass, Jingna Qiu, Matthias Rübner, Arndt Hartmann, Matthias W. Beckmann, Peter A. Fasching, Andreas K. Maier, Ramona Erber, Bernhard Kainz, Katharina Breininger |
ECCV (75) | 11 |
| 2024 | Arc2Face: A Foundation Model for ID-Consistent Human Faces
Foivos Paraperas Papantoniou, Alexander Lattas, Stylianos Moschoglou, Jiankang Deng, Bernhard Kainz, Stefanos Zafeiriou |
ECCV (37) | 5 |
| 2024 | URCDM: Ultra-Resolution Image Synthesis in Histopathology
Sarah Cechnicka, James Ball, Matthew Baugh, Hadrien Reynaud, Naomi Simmonds, Andrew P. T. Smith, Catherine Horsfield, Candice Roufosse, Bernhard Kainz |
MICCAI (4) | 9 |
| 2024 | Image Distillation for Safe Data Sharing in Histopathology
Zhe Li 0025, Bernhard Kainz |
MICCAI (10) | 2 |
| 2024 | Universal Topology Refinement for Medical Image Segmentation with Polynomial Feature Synthesis
Liu Li 0001, Hanchun Wang, Matthew Baugh, Qiang Ma 0004, Cheng Ouyang, Daniel Rueckert, Bernhard Kainz |
MICCAI (9) | 8 |
| 2024 | Weakly Supervised Learning of Cortical Surface Reconstruction from Segmentations
Qiang Ma 0004, Liu Li 0001, Emma C. Robinson, Bernhard Kainz, Daniel Rueckert |
MICCAI (11) | 4 |
| 2024 | Ensembled Cold-Diffusion Restorations for Unsupervised Anomaly Detection
Sergio Naval Marimont, Vasilis Siomos, Matthew Baugh, Christos Tzelepis, Bernhard Kainz, Giacomo Tarroni |
MICCAI (11) | 5 |
| 2024 | EchoNet-Synthetic: Privacy-Preserving Video Generation for Safe Medical Data Sharing
Hadrien Reynaud, Qingjie Meng, Mischa Dombrowski, Thomas G. Day, Alberto Gómez 0002, Paul Leeson, Bernhard Kainz |
MICCAI (7) | 8 |
| 2024 | Stability and Generalizability in SDE Diffusion Models with Measure-Preserving DynamicsabstractInverse problems describe the process of estimating the causal factors from a set of measurements or data.
Mapping of often incomplete or degraded data to parameters is ill-posed, thus data-driven iterative solutions are required, for example when reconstructing clean images from poor signals.
Diffusion models have shown promise as potent generative tools for solving inverse problems due to their superior reconstruction quality and their compatibility with iterative solvers. However, most existing approaches are limited to linear inverse problems represented as Stochastic Differential Equations (SDEs). This simplification falls short of addressing the challenging nature of real-world problems, leading to amplified cumulative errors and biases.
We provide an explanation for this gap through the lens of measure-preserving dynamics of Random Dynamical Systems (RDS) with which we analyse Temporal Distribution Discrepancy and thus introduce a theoretical framework based on RDS for SDE diffusion models. We uncover several strategies that inherently enhance the stability and generalizability of diffusion models for inverse problems and introduce a novel score-based diffusion framework, the Dynamics-aware SDE Diffusion Generative Model (D^3GM). The Measure-preserving property can return the degraded measurement to the original state despite complex degradation with the RDS concept of stability.
Our extensive experimental results corroborate the effectiveness of D^3GM across multiple benchmarks including a prominent application for inverse problems, magnetic resonance imaging. Chengqi Zang, Liu Li 0001, Sarah Cechnicka, Cheng Ouyang, Bernhard Kainz |
NeurIPS | 6 |
| 2024 | Multitask Weakly Supervised Generative Network for MR-US RegistrationabstractRegistering pre-operative modalities, such as magnetic resonance imaging or computed tomography, to ultrasound images is crucial for guiding clinicians during surgeries and biopsies. Recently, deep-learning approaches have been proposed to increase the speed and accuracy of this registration problem. However, all of these approaches need expensive supervision from the ultrasound domain. In this work, we propose a multitask generative framework that needs weak supervision only from the pre-operative imaging domain during training. To perform a deformable registration, the proposed framework translates a magnetic resonance image to the ultrasound domain while preserving the structural content. To demonstrate the efficacy of the proposed method, we tackle the registration problem of pre-operative 3D MR to transrectal ultrasonography images as necessary for targeted prostate biopsies. We use an in-house dataset of 600 patients, divided into 540 for training, 30 for validation, and the remaining for testing. An expert manually segmented the prostate in both modalities for validation and test sets to assess the performance of our framework. The proposed framework achieves a 3.58 mm target registration error on the expert-selected landmarks, 89.2% in the Dice score, and 1.81 mm 95th percentile Hausdorff distance on the prostate masks in the test set. Our experiments demonstrate that the proposed generative model successfully translates magnetic resonance images into the ultrasound domain. The translated image contains the structural content and fine details due to an ultrasound-specific two-path design of the generative model. The proposed framework enables training learning-based registration methods while only weak supervision from the pre-operative domain is available. Mohammad Farid Azampour, Kristina Mach, Emad Fatemizadeh, Beatrice Demiray, Kay Westenfelder, Katja Steiger, Matthias Eiber, Thomas Wendler 0001, Bernhard Kainz, Nassir Navab |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Foreground-Background Separation through Concept Distillation from Generative Image Foundation ModelsabstractCurating datasets for object segmentation is a difficult task. With the advent of large-scale pre-trained generative models, conditional image generation has been given a significant boost in result quality and ease of use. In this paper, we present a novel method that enables the generation of general foreground-background segmentation models from simple textual descriptions, without requiring segmentation labels. We leverage and explore pre-trained latent diffusion models, to automatically generate weak segmentation masks for concepts and objects. The masks are then used to fine-tune the diffusion model on an inpainting task, which enables fine-grained removal of the object, while at the same time providing a synthetic foreground and background dataset. We demonstrate that using this method beats previous methods in both discriminative and generative performance and closes the gap with fully supervised training while requiring no pixel-wise object labels. We show results on the task of segmenting four different objects (humans, dogs, cars, birds) and a use case scenario in medical image analysis. The code is available at https://github.com/MischaD/fobadiffusion. Mischa Dombrowski, Hadrien Reynaud, Matthew Baugh, Bernhard Kainz |
ICCV | 4 |
| 2023 | Many Tasks Make Light Work: Learning to Localise Medical Anomalies from Multiple Synthetic Tasks
Matthew Baugh, Jeremy Tan, Johanna P. Müller, Mischa Dombrowski, James Batten, Bernhard Kainz |
MICCAI (1) | 6 |
| 2023 | Robust Segmentation via Topology Violation Detection and Feature Synthesis
Liu Li 0001, Qiang Ma 0004, Cheng Ouyang, Zeju Li, Qingjie Meng, Mengyun Qiao, Vanessa Kyriakopoulou, Joseph V. Hajnal, Daniel Rueckert, Bernhard Kainz |
MICCAI (4) | 11 |
| 2023 | Conditional Temporal Attention Networks for Neonatal Cortical Surface Reconstruction
Qiang Ma 0004, Liu Li 0001, Vanessa Kyriakopoulou, Joseph V. Hajnal, Emma C. Robinson, Bernhard Kainz, Daniel Rueckert |
MICCAI (4) | 6 |
| 2023 | Feature-Conditioned Cascaded Video Diffusion Models for Precise Echocardiogram Synthesis
Hadrien Reynaud, Mengyun Qiao, Mischa Dombrowski, Thomas G. Day, Reza Razavi, Alberto Gómez 0002, Paul Leeson, Bernhard Kainz |
MICCAI (10) | 8 |
| 2023 | MoCoSR: Respiratory Motion Correction and Super-Resolution for 3D Abdominal MRI
Berke Doga Basaran, Qingjie Meng, Matthew Baugh, Jonathan K. Stelter, Phillip Lung, Uday Patel, Wenjia Bai, Dimitrios C. Karampinos, Bernhard Kainz |
MICCAI (10) | 10 |
| 2023 | ParaDime: A Framework for Parametric Dimensionality ReductionabstractParaDime is a framework for parametric dimensionality reduction (DR). In parametric DR, neural networks are trained to embed high-dimensional data items in a low-dimensional space while minimizing an objective function. ParaDime builds on the idea that the objective functions of several modern DR techniques result from transformed inter-item relationships. It provides a common interface for specifying these relations and transformations and for defining how they are used within the losses that govern the training process. Through this interface, ParaDime unifies parametric versions of DR techniques such as metric MDS, t-SNE, and UMAP. It allows users to fully customize all aspects of the DR process. We show how this ease of customization makes ParaDime suitable for experimenting with interesting techniques such as hybrid classification/embedding models and supervised DR. This way, ParaDime opens up new possibilities for visualizing high-dimensional data. Andreas P. Hinterreiter, Christina Humer, Bernhard Kainz, Marc Streit |
Comput. Graph. Forum | 3 |
| 2023 | Fast fetal head compounding from multi-view 3D ultrasound
Robert Wright, Alberto Gómez 0002, Veronika A. M. Zimmer, Nicolas Toussaint, Bishesh Khanal, Jacqueline Matthew, Emily Skelton, Bernhard Kainz, Daniel Rueckert, Joseph V. Hajnal, Julia A. Schnabel |
Medical Image Anal. | 8 |
| 2023 | Placenta segmentation in ultrasound imaging: Addressing sources of uncertainty and limited field-of-viewabstractAutomatic segmentation of the placenta in fetal ultrasound (US) is challenging due to the (i) high diversity of placenta appearance, (ii) the restricted quality in US resulting in highly variable reference annotations, and (iii) the limited field-of-view of US prohibiting whole placenta assessment at late gestation. In this work, we address these three challenges with a multi-task learning approach that combines the classification of placental location (e.g., anterior, posterior) and semantic placenta segmentation in a single convolutional neural network. Through the classification task the model can learn from larger and more diverse datasets while improving the accuracy of the segmentation task in particular in limited training set conditions. With this approach we investigate the variability in annotations from multiple raters and show that our automatic segmentations (Dice of 0.86 for anterior and 0.83 for posterior placentas) achieve human-level performance as compared to intra- and inter-observer variability. Lastly, our approach can deliver whole placenta segmentation using a multi-view US acquisition pipeline consisting of three stages: multi-probe image acquisition, image fusion and image segmentation. This results in high quality segmentation of larger structures such as the placenta in US with reduced image artifacts which are beyond the field-of-view of single probes. Veronika A. M. Zimmer, Alberto Gómez 0002, Emily Skelton, Robert Wright, Gavin Wheeler, Shujie Deng, Nooshin Ghavami, Karen Lloyd, Jacqueline Matthew, Bernhard Kainz, Daniel Rueckert, Joseph V. Hajnal, Julia A. Schnabel |
Medical Image Anal. | 10 |
| 2023 | Video-Based Activity Recognition for Automated Motor Assessment of Parkinson's DiseaseabstractOver the last decade, video-enabled mobile devices have become ubiquitous, while advances in markerless pose estimation allow an individual's body position to be tracked accurately and efficiently across the frames of a video. Previous work by this and other groups has shown that pose-extracted kinematic features can be used to reliably measure motor impairment in Parkinson's disease (PD). This presents the prospect of developing an asynchronous and scalable, video-based assessment of motor dysfunction. Crucial to this endeavour is the ability to automatically recognise the class of an action being performed, without which manual labelling is required. Representing the evolution of body joint locations as a spatio-temporal graph, we implement a deep-learning model for video and frame-level classification of activities performed according to part 3 of the Movement Disorder Society Unified PD Rating Scale (MDS-UPDRS). We train and validate this system using a dataset of n = 7310 video clips, recorded at 5 independent sites. This approach reaches human-level performance in detecting and classifying periods of activity within monocular video clips. Our framework could support clinical workflows and patient care at scale through applications such as quality monitoring of clinical data collection, automated labelling of video streams, or a module within a remote self-assessment system. Grzegorz Sarapata, Yuriy Dushin, Gareth Morinan, Joshua Ong, Sanjay Budhdeo, Bernhard Kainz, Jonathan O'Keeffe |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | CortexODE: Learning Cortical Surface Reconstruction by Neural ODEsabstractWe present CortexODE, a deep learning framework for cortical surface reconstruction. CortexODE leverages neural ordinary differential equations (ODEs) to deform an input surface into a target shape by learning a diffeomorphic flow. The trajectories of the points on the surface are modeled as ODEs, where the derivatives of their coordinates are parameterized via a learnable Lipschitz-continuous deformation network. This provides theoretical guarantees for the prevention of self-intersections. CortexODE can be integrated to an automatic learning-based pipeline, which reconstructs cortical surfaces efficiently in less than 5 seconds. The pipeline utilizes a 3D U-Net to predict a white matter segmentation from brain Magnetic Resonance Imaging (MRI) scans, and further generates a signed distance function that represents an initial surface. Fast topology correction is introduced to guarantee homeomorphism to a sphere. Following the isosurface extraction step, two CortexODE models are trained to deform the initial surface to white matter and pial surfaces respectively. The proposed pipeline is evaluated on large-scale neuroimage datasets in various age groups including neonates (25-45 weeks), young adults (22-36 years) and elderly subjects (55-90 years). Our experiments demonstrate that the CortexODE-based pipeline can achieve less than 0.2mm average geometric error while being orders of magnitude faster compared to conventional processing pipelines. Qiang Ma 0004, Liu Li 0001, Emma C. Robinson, Bernhard Kainz, Daniel Rueckert, Amir Alansary |
IEEE Trans. Medical Imaging | 4 |
| 2022 | A variational Bayesian method for similarity learning in non-rigid image registrationabstractWe propose a novel variational Bayesian formulation for diffeomorphic non-rigid registration of medical images, which learns in an unsupervised way a data-specific similarity metric. The proposed framework is general and may be used together with many existing image registration models. We evaluate it on brain MRI scans from the UK Biobank and show that use of the learnt similarity metric, which is parametrised as a neural network, leads to more accurate results than use of traditional functions, e.g. SSD and LCC, to which we initialise the model, without a negative impact on image registration speed or transformation smoothness. In addition, the method estimates the uncertainty associated with the transformation. The code and the trained models are available in a public repository: https://github.com/dgrzech/learnsim. Daniel Grzech, Mohammad Farid Azampour, Ben Glocker, Julia A. Schnabel, Nassir Navab, Bernhard Kainz, Loïc Le Folgoc |
CVPR | 6 |
| 2022 | Natural Synthetic Anomalies for Self-supervised Anomaly Detection and Localization
Hannah M. Schlüter, Jeremy Tan, Benjamin Hou, Bernhard Kainz |
ECCV (31) | 4 |
| 2022 | D'ARTAGNAN: Counterfactual Video Generation
Hadrien Reynaud, Athanasios Vlontzos, Mischa Dombrowski, Ciarán M. Gilligan-Lee, Arian Beqiri, Paul Leeson, Bernhard Kainz |
MICCAI (8) | 7 |
| 2022 | Video Summarization Through Reinforcement Learning With a 3D Spatio-Temporal U-NetabstractIntelligent video summarization algorithms allow to quickly convey the most relevant information in videos through the identification of the most essential and explanatory content while removing redundant video frames. In this paper, we introduce the 3DST-UNet-RL framework for video summarization. A 3D spatio-temporal U-Net is used to efficiently encode spatio-temporal information of the input videos for downstream reinforcement learning (RL). An RL agent learns from spatio-temporal latent scores and predicts actions for keeping or rejecting a video frame in a video summary. We investigate if real/inflated 3D spatio-temporal CNN features are better suited to learn representations from videos than commonly used 2D image features. Our framework can operate in both, a fully unsupervised mode and a supervised training mode. We analyse the impact of prescribed summary lengths and show experimental evidence for the effectiveness of 3DST-UNet-RL on two commonly used general video summarization benchmarks. We also applied our method on a medical video summarization task. The proposed video summarization method has the potential to save storage costs of ultrasound screening videos as well as to increase efficiency when browsing patient video data during retrospective analysis or audit without loosing essential information. Tianrui Liu 0001, Qingjie Meng, Junjie Huang 0001, Athanasios Vlontzos, Daniel Rueckert, Bernhard Kainz |
IEEE Trans. Image Process. | 6 |
| 2022 | MOOD 2020: A Public Benchmark for Out-of-Distribution Detection and Localization on Medical ImagesabstractDetecting Out-of-Distribution (OoD) data is one of the greatest challenges in safe and robust deployment of machine learning algorithms in medicine. When the algorithms encounter cases that deviate from the distribution of the training data, they often produce incorrect and over-confident predictions. OoD detection algorithms aim to catch erroneous predictions in advance by analysing the data distribution and detecting potential instances of failure. Moreover, flagging OoD cases may support human readers in identifying incidental findings. Due to the increased interest in OoD algorithms, benchmarks for different domains have recently been established. In the medical imaging domain, for which reliable predictions are often essential, an open benchmark has been missing. We introduce the Medical-Out-Of-Distribution-Analysis-Challenge (MOOD) as an open, fair, and unbiased benchmark for OoD methods in the medical imaging domain. The analysis of the submitted algorithms shows that performance has a strong positive correlation with the perceived difficulty, and that all algorithms show a high variance for different anomalies, making it yet hard to recommend them for clinical practice. We also see a strong correlation between challenge ranking and performance on a simple toy test set, indicating that this might be a valuable addition as a proxy dataset during anomaly detection algorithm development. David Zimmerer, Peter M. Full, Fabian Isensee, Paul F. Jaeger, Tim Adler, Jens Petersen, Gregor Köhler, Tobias Roß, Annika Reinke, Antanas Kascenas, Bjørn Sand Jensen, Alison O'Neil, Jeremy Tan, Benjamin Hou, James Batten, Huaqi Qiu, Bernhard Kainz, Nina Shvetsova, Irina Fedulova, Dmitry V. Dylov, Baolun Yu, Jianyang Zhai, Jingtao Hu, Runxuan Si, Sihang Zhou 0001, Siqi Wang 0001, Xuerun Chen, Yang Zhao 0003, Sergio Naval Marimont, Giacomo Tarroni, Victor Saase, Lena Maier-Hein, Klaus H. Maier-Hein |
IEEE Trans. Medical Imaging | 17 |
| 2021 | Unsupervised Human Pose Estimation Through Transforming Shape TemplatesabstractHuman pose estimation is a major computer vision problem with applications ranging from augmented reality and video capture to surveillance and movement tracking. In the medical context, the latter may be an important biomarker for neurological impairments in infants. Whilst many methods exist, their application has been limited by the need for well annotated large datasets and the inability to gen-eralize to humans of different shapes and body compositions, e.g. children and infants. In this paper we present a novel method for learning pose estimators for human adults and infants in an unsupervised fashion. We approach this as a learnable template matching problem facilitated by deep feature extractors. Human-interpretable landmarks are estimated by transforming a template consisting of predefined body parts that are characterized by 2D Gaussian distributions. Enforcing a connectivity prior guides our model to meaningful human shape representations. We demonstrate the effectiveness of our approach on two different datasets including adults and infants. Project page: infantmotion.github.io Luca Schmidtke, Athanasios Vlontzos, Simon Ellershaw, Anna Lukens, Tomoki Arichi, Bernhard Kainz |
CVPR | 6 |
| 2021 | Detecting Hypo-plastic Left Heart Syndrome in Fetal Ultrasound via Disease-Specific Atlas Maps
Samuel Budd, Matthew Sinclair, Thomas G. Day, Athanasios Vlontzos, Jeremy Tan, Tianrui Liu 0001, Jacqueline Matthew, Emily Skelton, John M. Simpson, Reza Razavi, Ben Glocker, Daniel Rueckert, Emma C. Robinson, Bernhard Kainz |
MICCAI (7) | 14 |
| 2021 | RATCHET: Medical Transformer for Chest X-ray Diagnosis and Reporting
Benjamin Hou, Georgios Kaissis, Ronald M. Summers, Bernhard Kainz |
MICCAI (7) | 4 |
| 2021 | Ultrasound Video Transformers for Cardiac Ejection Fraction Estimation
Hadrien Reynaud, Athanasios Vlontzos, Benjamin Hou, Arian Beqiri, Paul Leeson, Bernhard Kainz |
MICCAI (6) | 6 |
| 2021 | Detecting Outliers with Poisson Image Interpolation
Jeremy Tan, Benjamin Hou, Thomas G. Day, John M. Simpson, Daniel Rueckert, Bernhard Kainz |
MICCAI (5) | 6 |
| 2021 | Deep radiance caching: Convolutional autoencoders deeper in ray tracing®
Giulio Jiang, Bernhard Kainz |
Comput. Graph. | 2 |
| 2021 | A survey on active learning and human-in-the-loop deep learning for medical image analysis
Samuel Budd, Emma C. Robinson, Bernhard Kainz |
Medical Image Anal. | 3 |
| 2021 | Mutual Information-Based Disentangled Neural Networks for Classifying Unseen Categories in Different Domains: Application to Fetal Ultrasound ImagingabstractDeep neural networks exhibit limited generalizability across images with different entangled domain features and categorical features. Learning generalizable features that can form universal categorical decision boundaries across domains is an interesting and difficult challenge. This problem occurs frequently in medical imaging applications when attempts are made to deploy and improve deep learning models across different image acquisition devices, across acquisition parameters or if some classes are unavailable in new training databases. To address this problem, we propose Mutual Information-based Disentangled Neural Networks (MIDNet), which extract generalizable categorical features to transfer knowledge to unseen categories in a target domain. The proposed MIDNet adopts a semi-supervised learning paradigm to alleviate the dependency on labeled data. This is important for real-world applications where data annotation is time-consuming, costly and requires training and expertise. We extensively evaluate the proposed method on fetal ultrasound datasets for two different image classification tasks where domain features are respectively defined by shadow artifacts and image acquisition devices. Experimental results show that the proposed method outperforms the state-of-the-art on the classification of unseen categories in a target domain with sparsely labeled training data. Qingjie Meng, Jacqueline Matthew, Veronika A. M. Zimmer, Alberto Gómez 0002, David Lloyd 0003, Daniel Rueckert, Bernhard Kainz |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Ultrasound Video Summarization Using Deep Reinforcement Learning
Qingjie Meng, Athanasios Vlontzos, Jeremy Tan, Daniel Rueckert, Bernhard Kainz |
MICCAI (3) | 6 |
| 2019 | Confident Head Circumference Measurement from Ultrasound with Real-Time Feedback for Sonographers
Samuel Budd, Matthew Sinclair, Bishesh Khanal, Jacqueline Matthew, David Lloyd 0003, Alberto Gómez 0002, Nicolas Toussaint, Emma C. Robinson, Bernhard Kainz |
MICCAI (4) | 9 |
| 2019 | Multiple Landmark Detection Using Multi-agent Reinforcement Learning
Athanasios Vlontzos, Amir Alansary, Konstantinos Kamnitsas, Daniel Rueckert, Bernhard Kainz |
MICCAI (4) | 5 |
| 2019 | Complete Fetal Head Compounding from Multi-view 3D Ultrasound
Robert Wright, Nicolas Toussaint, Alberto Gómez 0002, Veronika A. M. Zimmer, Bishesh Khanal, Jacqueline Matthew, Emily Skelton, Bernhard Kainz, Daniel Rueckert, Joseph V. Hajnal, Julia A. Schnabel |
MICCAI (3) | 8 |
| 2019 | Foreword to the special section on the Eurographics Workshop on Visual Computing for Biology and Medicine (VCBM) at Medical Image Computing and Computer Assisted Intervention (MICCAI) 2018
Thomas Schultz 0001, Anna Puig, Bernhard Kainz |
Comput. Graph. | 3 |
| 2019 | Morpho-MNIST: Quantitative Assessment and Diagnostics for Representation LearningabstractRevealing latent structure in data is an active field of research, having introduced exciting technologies such as variational autoencoders and adversarial networks, and is essential to push machine learning towards unsupervised knowledge discovery. However, a major challenge is the lack of suitable benchmarks for an objective and quantitative evaluation of learned representations. To address this issue we introduce Morpho-MNIST, a framework that aims to answer: “to what extent has my model learned to represent specific factors of variation in the data? We extend the popular MNIST dataset by adding a morphometric analysis enabling quantitative comparison of trained models, identification of the roles of latent variables, and characterisation of sample diversity. We further propose a set of quantifiable perturbations to assess the performance of unsupervised and supervised methods on challenging tasks such as outlier detection and domain adaptation. Data and code are available at https://github.com/dccastro/Morpho-MNIST. Daniel C. Castro, Jeremy Tan, Bernhard Kainz, Ender Konukoglu, Ben Glocker |
J. Mach. Learn. Res. | 3 |
| 2019 | Evaluating reinforcement learning agents for anatomical landmark detection
Amir Alansary, Ozan Oktay, Loïc Le Folgoc, Benjamin Hou, Ghislain Vaillant, Konstantinos Kamnitsas, Athanasios Vlontzos, Ben Glocker, Bernhard Kainz, Daniel Rueckert |
Medical Image Anal. | 10 |
| 2019 | Attention gated networks: Learning to leverage salient regions in medical imagesabstractWe propose a novel attention gate (AG) model for medical image analysis that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules when using convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN models such as VGG or U-Net architectures with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed AG models are evaluated on a variety of tasks, including medical image classification and segmentation. For classification, we demonstrate the use case of AGs in scan plane detection for fetal ultrasound screening. We show that the proposed attention mechanism can provide efficient object localisation while improving the overall prediction performance by reducing false positives. For segmentation, the proposed architecture is evaluated on two large 3D CT abdominal datasets with manual annotations for multiple organs. Experimental results show that AG models consistently improve the prediction performance of the base architectures across different datasets and training sizes while preserving computational efficiency. Moreover, AGs guide the model activations to be focused around salient regions, which provides better insights into how model predictions are made. The source code for the proposed AG models is publicly available. Jo Schlemper, Ozan Oktay, Michiel Schaap, Mattias P. Heinrich, Bernhard Kainz, Ben Glocker, Daniel Rueckert |
Medical Image Anal. | 5 |
| 2019 | Weakly Supervised Estimation of Shadow Confidence Maps in Fetal Ultrasound ImagingabstractDetecting acoustic shadows in ultrasound images is important in many clinical and engineering applications. Real-time feedback of acoustic shadows can guide sonographers to a standardized diagnostic viewing plane with minimal artifacts and can provide additional information for other automatic image analysis algorithms. However, automatically detecting shadow regions using learning-based algorithms is challenging because pixel-wise ground truth annotation of acoustic shadows is subjective and time consuming. In this paper, we propose a weakly supervised method for automatic confidence estimation of acoustic shadow regions. Our method is able to generate a dense shadow-focused confidence map. In our method, a shadow-seg module is built to learn general shadow features for shadow segmentation, based on global image-level annotations as well as a small number of coarse pixel-wise shadow annotations. A transfer function is introduced to extend the obtained binary shadow segmentation to a reference confidence map. In addition, a confidence estimation network is proposed to learn the mapping between input images and the reference confidence maps. This network is able to predict shadow confidence maps directly from input images during inference. We use evaluation metrics such as DICE, inter-class correlation, and so on, to verify the effectiveness of our method. Our method is more consistent than human annotation and outperforms the state-of-the-art quantitatively in shadow segmentation and qualitatively in confidence estimation of shadow regions. Furthermore, we demonstrate the applicability of our method by integrating shadow confidence maps into tasks such as ultrasound image classification, multi-view image fusion, and automated biometric measurements. Qingjie Meng, Richard James Housden, Jacqueline Matthew, Daniel Rueckert, Julia A. Schnabel, Bernhard Kainz, Matthew Sinclair, Veronika A. M. Zimmer, Benjamin Hou, Martin Rajchl, Nicolas Toussaint, Ozan Oktay, Jo Schlemper, Alberto Gómez 0002 |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Automatic View Planning with Multi-scale Deep Reinforcement Learning Agents
Amir Alansary, Loïc Le Folgoc, Ghislain Vaillant, Ozan Oktay, Wenjia Bai, Jonathan Passerat-Palmbach, Ricardo Guerrero, Konstantinos Kamnitsas, Benjamin Hou, Steven McDonagh 0001, Ben Glocker, Bernhard Kainz, Daniel Rueckert |
MICCAI (1) | 13 |
| 2018 | 3D Fetal Skull Reconstruction from 2DUS via Deep Conditional Generative Networks
Juan J. Cerrolaza, Carlo Biffi, Alberto Gómez 0002, Matthew Sinclair, Jacqueline Matthew, Caronline Knight, Bernhard Kainz, Daniel Rueckert |
MICCAI (1) | 8 |
| 2018 | Computing CNN Loss and Gradients for Pose Estimation with Riemannian Geometry
Benjamin Hou, Nina Miolane, Bishesh Khanal, Matthew C. H. Lee, Amir Alansary, Steven McDonagh 0001, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz |
MICCAI (1) | 10 |
| 2018 | Fast Multiple Landmark Localisation Using a Patch-Based Iterative Network
Amir Alansary, Juan J. Cerrolaza, Bishesh Khanal, Matthew Sinclair, Jacqueline Matthew, Chandni Gupta, Caroline L. Knight, Bernhard Kainz, Daniel Rueckert |
MICCAI (1) | 9 |
| 2018 | Standard Plane Detection in 3D Fetal Ultrasound Using an Iterative Transformation Network
Bishesh Khanal, Benjamin Hou, Amir Alansary, Juan J. Cerrolaza, Matthew Sinclair, Jacqueline Matthew, Chandni Gupta, Caroline L. Knight, Bernhard Kainz, Daniel Rueckert |
MICCAI (1) | 10 |
| 2018 | Real-Time Prediction of Segmentation Quality
Robert Robinson, Ozan Oktay, Wenjia Bai, Vanya V. Valindria, Mihir Sanghvi, Nay Aung, José Miguel Paiva, Filip Zemrak, Kenneth Fung, Elena Lukaschuk, Aaron M. Lee, Valentina Carapella, Bernhard Kainz, Stefan K. Piechnik, Stefan Neubauer, Steffen E. Petersen, Chris Page, Daniel Rueckert, Ben Glocker |
MICCAI (4) | 14 |
| 2018 | AutoDVT: Joint Real-Time Classification for Vein Compressibility Analysis in Deep Vein Thrombosis Ultrasound Diagnostics
Ryutaro Tanno, Antonios Makropoulos, Salim Arslan, Ozan Oktay, Sven Mischkewitz, Fouad Al-Noor, Jonas Oppenheimer, Ramin Mandegaran, Bernhard Kainz, Mattias P. Heinrich |
MICCAI (2) | 9 |
| 2018 | 3-D Reconstruction in Canonical Co-Ordinate Space From Arbitrarily Oriented 2-D ImagesabstractLimited capture range, and the requirement to provide high quality initialization for optimization-based 2-D/3-D image registration methods, can significantly degrade the performance of 3-D image reconstruction and motion compensation pipelines. Challenging clinical imaging scenarios, which contain significant subject motion, such as fetal in-utero imaging, complicate the 3-D image and volume reconstruction process. In this paper, we present a learning-based image registration method capable of predicting 3-D rigid transformations of arbitrarily oriented 2-D image slices, with respect to a learned canonical atlas co-ordinate system. Only image slice intensity information is used to perform registration and canonical alignment, no spatial transform initialization is required. To find image transformations, we utilize a convolutional neural network architecture to learn the regression function capable of mapping 2-D image slices to a 3-D canonical atlas space. We extensively evaluate the effectiveness of our approach quantitatively on simulated magnetic resonance imaging (MRI), fetal brain imagery with synthetic motion and further demonstrate qualitative results on real fetal MRI data where our method is integrated into a full reconstruction and motion compensation pipeline. Our learning based registration achieves an average spatial prediction error of 7 mm on simulated data and produces qualitatively improved reconstructions for heavily moving fetuses with gestational ages of approximately 20 weeks. Our model provides a general and computationally efficient solution to the 2-D/3-D registration initialization problem and is suitable for real-time scenarios. Benjamin Hou, Bishesh Khanal, Amir Alansary, Steven McDonagh 0001, Alice Davidson, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz |
IEEE Trans. Medical Imaging | 10 |
| 2018 | Anatomically Constrained Neural Networks (ACNNs): Application to Cardiac Image Enhancement and SegmentationabstractIncorporation of prior knowledge about organ shape and location is key to improve performance of image analysis approaches. In particular, priors can be useful in cases where images are corrupted and contain artefacts due to limitations in image acquisition. The highly constrained nature of anatomical objects can be well captured with learning-based techniques. However, in most recent and promising techniques such as CNN-based segmentation it is not obvious how to incorporate such prior knowledge. State-of-the-art methods operate as pixel-wise classifiers where the training objectives do not incorporate the structure and inter-dependencies of the output. To overcome this limitation, we propose a generic training strategy that incorporates anatomical prior knowledge into CNNs through a new regularisation model, which is trained end-to-end. The new framework encourages models to follow the global anatomical properties of the underlying anatomy (e.g. shape, label structure) via learnt non-linear representations of the shape. We show that the proposed approach can be easily adapted to different analysis tasks (e.g. image enhancement, segmentation) and improve the prediction accuracy of the state-of-the-art models. The applicability of our approach is shown on multi-modal cardiac data sets and public benchmarks. In addition, we demonstrate how the learnt deep models of 3-D shapes can be interpreted and used as biomarkers for classification of cardiac pathologies. Ozan Oktay, Enzo Ferrante, Konstantinos Kamnitsas, Mattias P. Heinrich, Wenjia Bai, Jose Caballero, Stuart A. Cook, Antonio M. Simoes Monteiro de Marvao, Timothy Dawes, Declan P. O'Regan, Bernhard Kainz, Ben Glocker, Daniel Rueckert |
IEEE Trans. Medical Imaging | 11 |
| 2017 | Predicting Slice-to-Volume Transformation in Presence of Arbitrary Subject Motion
Benjamin Hou, Amir Alansary, Steven McDonagh 0001, Alice Davidson, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert, Ben Glocker, Bernhard Kainz |
MICCAI (2) | 9 |
| 2017 | PVR: Patch-to-Volume Reconstruction for Large Area Motion Correction of Fetal MRIabstractIn this paper, we present a novel method for the correction of motion artifacts that are present in fetal magnetic resonance imaging (MRI) scans of the whole uterus. Contrary to current slice-to-volume registration (SVR) methods, requiring an inflexible anatomical enclosure of a single investigated organ, the proposed patch-to-volume reconstruction (PVR) approach is able to reconstruct a large field of view of non-rigidly deforming structures. It relaxes rigid motion assumptions by introducing a specific amount of redundant information that is exploited with parallelized patchwise optimization, super-resolution, and automatic outlier rejection. We further describe and provide an efficient parallel implementation of PVR allowing its execution within reasonable time on commercially available graphics processing units, enabling its use in the clinical practice. We evaluate PVR's computational overhead compared with standard methods and observe improved reconstruction accuracy in the presence of affine motion artifacts compared with conventional SVR in synthetic experiments. Furthermore, we have evaluated our method qualitatively and quantitatively on real fetal MRI data subject to maternal breathing and sudden fetal movements. We evaluate peak-signal-to-noise ratio, structural similarity index, and cross correlation with respect to the originally acquired data and provide a method for visual inspection of reconstruction uncertainty. We further evaluate the distance error for selected anatomical landmarks in the fetal head, as well as calculating the mean and maximum displacements resulting from automatic non-rigid registration to a motion-free ground truth image. These experiments demonstrate a successful application of PVR motion compensation to the whole fetal body, uterus, and placenta. Amir Alansary, Martin Rajchl, Steven McDonagh 0001, Maria Deprez, Mellisa Damodaram, David Lloyd 0003, Alice Davidson, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert, Bernhard Kainz |
IEEE Trans. Medical Imaging | 11 |
| 2017 | SonoNet: Real-Time Detection and Localisation of Fetal Standard Scan Planes in Freehand UltrasoundabstractIdentifying and interpreting fetal standard scan planes during 2-D ultrasound mid-pregnancy examinations are highly complex tasks, which require years of training. Apart from guiding the probe to the correct location, it can be equally difficult for a non-expert to identify relevant structures within the image. Automatic image processing can provide tools to help experienced as well as inexperienced operators with these tasks. In this paper, we propose a novel method based on convolutional neural networks, which can automatically detect 13 fetal standard views in freehand 2-D ultrasound data as well as provide a localization of the fetal structures via a bounding box. An important contribution is that the network learns to localize the target anatomy using weak supervision based on image-level labels only. The network architecture is designed to operate in real-time while providing optimal output for the localization task. We present results for real-time annotation, retrospective frame retrieval from saved videos, and localization on a very large and challenging dataset consisting of images and video recordings of full clinical anomaly screenings. We found that the proposed method achieved an average F1-score of 0.798 in a realistic classification experiment modeling real-time detection, and obtained a 90.09% accuracy for retrospective frame retrieval. Moreover, an accuracy of 77.8% was achieved on the localization task. Christian F. Baumgartner, Konstantinos Kamnitsas, Jacqueline Matthew, Tara P. Fletcher, Sandra Smith, Lisa M. Koch, Bernhard Kainz, Daniel Rueckert |
IEEE Trans. Medical Imaging | 7 |
| 2017 | DeepCut: Object Segmentation From Bounding Box Annotations Using Convolutional Neural NetworksabstractIn this paper, we propose DeepCut, a method to obtain pixelwise object segmentations given an image dataset labelled weak annotations, in our case bounding boxes. It extends the approach of the well-known GrabCut [1] method to include machine learning by training a neural network classifier from bounding box annotations. We formulate the problem as an energy minimisation problem over a densely-connected conditional random field and iteratively update the training targets to obtain pixelwise object segmentations. Additionally, we propose variants of the DeepCut method and compare those to a naïve approach to CNN training under weak supervision. We test its applicability to solve brain and lung segmentation problems on a challenging fetal magnetic resonance dataset and obtain encouraging results in terms of accuracy. Martin Rajchl, Matthew C. H. Lee, Ozan Oktay, Konstantinos Kamnitsas, Jonathan Passerat-Palmbach, Wenjia Bai, Mellisa Damodaram, Mary A. Rutherford, Joseph V. Hajnal, Bernhard Kainz, Daniel Rueckert |
IEEE Trans. Medical Imaging | 10 |
| 2017 | Placenta Maps: In Utero Placental Health Assessment of the Human FetusabstractThe human placenta is essential for the supply of the fetus. To monitor the fetal development, imaging data is acquired using (US). Although it is currently the gold-standard in fetal imaging, it might not capture certain abnormalities of the placenta. (MRI) is a safe alternative for the in utero examination while acquiring the fetus data in higher detail. Nevertheless, there is currently no established procedure for assessing the condition of the placenta and consequently the fetal health. Due to maternal respiration and inherent movements of the fetus during examination, a quantitative assessment of the placenta requires fetal motion compensation, precise placenta segmentation and a standardized visualization, which are challenging tasks. Utilizing advanced motion compensation and automatic segmentation methods to extract the highly versatile shape of the placenta, we introduce a novel visualization technique that presents the fetal and maternal side of the placenta in a standardized way. Our approach enables physicians to explore the placenta even in utero. This establishes the basis for a comparative assessment of multiple placentas to analyze possible pathologic arrangements and to support the research and understanding of this vital organ. Additionally, we propose a three-dimensional structure-aware surface slicing technique in order to explore relevant regions inside the placenta. Finally, to survey the applicability of our approach, we consulted clinical experts in prenatal diagnostics and imaging. We received mainly positive feedback, especially the applicability of our technique for research purposes was appreciated. Haichao Miao, Gabriel Mistelbauer, Alexey Karimov, Amir Alansary, Alice Davidson, David Lloyd 0003, Mellisa Damodaram, Lisa Story, Jana Hutter, Joseph V. Hajnal, Mary A. Rutherford, Bernhard Preim, Bernhard Kainz, M. Eduard Gröller |
IEEE Trans. Vis. Comput. Graph. | 13 |
| 2016 | Fast Fully Automatic Segmentation of the Human Placenta from Motion Corrupted MRI
Amir Alansary, Konstantinos Kamnitsas, Alice Davidson, Rostislav Khlebnikov, Martin Rajchl, Christina Malamateniou, Mary A. Rutherford, Joseph V. Hajnal, Ben Glocker, Daniel Rueckert, Bernhard Kainz |
MICCAI (2) | 11 |
| 2016 | Real-Time Standard Scan Plane Detection and Localisation in Fetal Ultrasound Using Fully Convolutional Neural Networks
Christian F. Baumgartner, Konstantinos Kamnitsas, Jacqueline Matthew, Sandra Smith, Bernhard Kainz, Daniel Rueckert |
MICCAI (2) | 5 |
| 2016 | Learning clinically useful information from images: Past, present and future
Daniel Rueckert, Ben Glocker, Bernhard Kainz |
Medical Image Anal. | 3 |
| 2015 | Flexible Reconstruction and Correction of Unpredictable Motion from Stacks of 2D Images
Bernhard Kainz, Amir Alansary, Christina Malamateniou, Kevin Keraudren, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (2) | 1 |
| 2015 | Automated Localization of Fetal Organs in MRI Using Random Forests with Steerable Features
Kevin Keraudren, Bernhard Kainz, Ozan Oktay, Vanessa Kyriakopoulou, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (3) | 2 |
| 2015 | Fast Volume Reconstruction From Motion Corrupted Stacks of 2D SlicesabstractCapturing an enclosing volume of moving subjects and organs using fast individual image slice acquisition has shown promise in dealing with motion artefacts. Motion between slice acquisitions results in spatial inconsistencies that can be resolved by slice-to-volume reconstruction (SVR) methods to provide high quality 3D image data. Existing algorithms are, however, typically very slow, specialised to specific applications and rely on approximations, which impedes their potential clinical use. In this paper, we present a fast multi-GPU accelerated framework for slice-to-volume reconstruction. It is based on optimised 2D/3D registration, super-resolution with automatic outlier rejection and an additional (optional) intensity bias correction. We introduce a novel and fully automatic procedure for selecting the image stack with least motion to serve as an initial registration target. We evaluate the proposed method using artificial motion corrupted phantom data as well as clinical data, including tracked freehand ultrasound of the liver and fetal Magnetic Resonance Imaging. We achieve speed-up factors greater than 30 compared to a single CPU system and greater than 10 compared to currently available state-of-the-art multi-core CPU methods. We ensure high reconstruction accuracy by exact computation of the point-spread function for every input data point, which has not previously been possible due to computational limitations. Our framework and its implementation is scalable for available computational infrastructures and tests show a speed-up factor of 1.70 for each additional GPU. This paves the way for the online application of image based reconstruction methods during clinical examinations. The source code for the proposed approach is publicly available. Bernhard Kainz, Markus Steinberger, Wolfgang Wein, Maria Deprez, Christina Malamateniou, Kevin Keraudren, Thomas Torsney-Weir, Mary A. Rutherford, Paul Aljabar, Joseph V. Hajnal, Daniel Rueckert |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Motion Corrected 3D Reconstruction of the Fetal Thorax from Prenatal MRI
Bernhard Kainz, Christina Malamateniou, Maria Deprez, Kevin Keraudren, Mary A. Rutherford, Joseph V. Hajnal, Daniel Rueckert |
MICCAI (2) | 1 |
| 2014 | Parallel Irradiance Caching for Interactive Monte-Carlo Direct Volume RenderingabstractAbstract We propose a technique to build the irradiance cache for isotropic scattering simultaneously with Monte Carlo progressive direct volume rendering on a single GPU, which allows us to achieve up to four times increased convergence rate for complex scenes with arbitrary sources of light. We use three procedures that run concurrently on a single GPU. The first is the main rendering procedure. The second procedure computes new cache entries, and the third one corrects the errors that may arise after creation of new cache entries. We propose two distinct approaches to allow massive parallelism of cache entry creation. In addition, we show a novel extrapolation approach which outputs high quality irradiance approximations and a suitable prioritization scheme to increase the convergence rate by dedicating more computational power to more complex rendering areas. Rostislav Khlebnikov, Philip Voglreiter, Markus Steinberger, Bernhard Kainz, Dieter Schmalstieg |
Comput. Graph. Forum | 4 |
| 2014 | Parallel generation of architecture on the GPUabstractAbstract In this paper, we present a novel approach for the parallel evaluation of procedural shape grammars on the graphics processing unit (GPU). Unlike previous approaches that are either limited in the kind of shapes they allow, the amount of parallelism they can take advantage of, or both, our method supports state of the art procedural modeling including stochasticity and context‐sensitivity. To increase parallelism, we explicitly express independence in the grammar, reduce inter‐rule dependencies required for context‐sensitive evaluation, and introduce intra‐rule parallelism. Our rule scheduling scheme avoids unnecessary back and forth between CPU and GPU and reduces round trips to slow global memory by dynamically grouping rules in on‐chip shared memory. Our GPU shape grammar implementation is multiple orders of magnitude faster than the standard in CPU‐based rule evaluation, while offering equal expressive power. In comparison to the state of the art in GPU shape grammar derivation, our approach is nearly 50 times faster, while adding support for geometric context‐sensitivity. Markus Steinberger, Michael Kenzel, Bernhard Kainz, Joerg H. Mueller, Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 3 |
| 2014 | On-the-fly generation and rendering of infinite cities on the GPUabstractAbstract In this paper, we present a new approach for shape‐grammar‐based generation and rendering of huge cities in real‐time on the graphics processing unit (GPU). Traditional approaches rely on evaluating a shape grammar and storing the geometry produced as a preprocessing step. During rendering, the pregenerated data is then streamed to the GPU. By interweaving generation and rendering, we overcome the problems and limitations of streaming pregenerated data. Using our methods ofvisibility pruningand adaptive level of detail, we are able to dynamically generate only the geometry needed to render the current view in real‐time directly on the GPU. We also present a robust and efficient way to dynamically update a scene's derivation tree and geometry, enabling us to exploit frame‐to‐frame coherence. Our combined generation and rendering is significantly faster than all previous work. For detailed scenes, we are capable of generating geometry more rapidly than even just copying pregenerated data from main memory, enabling us to render cities with thousands of buildings at up to 100 frames per second, even with the camera moving at supersonic speed. Markus Steinberger, Michael Kenzel, Bernhard Kainz, Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 3 |
| 2013 | Noise-Based Volume Rendering for the Visualization of Multivariate Volumetric DataabstractAnalysis of multivariate data is of great importance in many scientific disciplines. However, visualization of 3D spatially-fixed multivariate volumetric data is a very challenging task. In this paper we present a method that allows simultaneous real-time visualization of multivariate data. We redistribute the opacity within a voxel to improve the readability of the color defined by a regular transfer function, and to maintain the see-through capabilities of volume rendering. We use predictable procedural noise--random-phase Gabor noise--to generate a high-frequency redistribution pattern and construct an opacity mapping function, which allows to partition the available space among the displayed data attributes. This mapping function is appropriately filtered to avoid aliasing, while maintaining transparent regions. We show the usefulness of our approach on various data sets and with different example applications. Furthermore, we evaluate our method by comparing it to other visualization techniques in a controlled user study. Overall, the results of our study indicate that users are much more accurate in determining exact data values with our novel 3D volume visualization method. Significantly lower error rates for reading data values and high subjective ranking of our method imply that it has a high chance of being adopted for the purpose of visualization of multivariate 3D data. Rostislav Khlebnikov, Bernhard Kainz, Markus Steinberger, Dieter Schmalstieg |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | OmniKinect: real-time dense volumetric data acquisition and applicationsabstractReal-time three-dimensional acquisition of real-world scenes has many important applications in computer graphics, computer vision and human-computer interaction. Inexpensive depth sensors such as the Microsoft Kinect allow to leverage the development of such applications. However, this technology is still relatively recent, and no detailed studies on its scalability to dense and view-independent acquisition have been reported. This paper addresses the question of what can be done with a larger number of Kinects used simultaneously. We describe an interference-reducing physical setup, a calibration procedure and an extension to the KinectFusion algorithm, which allows to produce high quality volumetric reconstructions from multiple Kinects whilst overcoming systematic errors in the depth measurements. We also report on enhancing image based visual hull rendering by depth measurements, and compare the results to KinectFusion. Our system provides practical insight into achievable spatial and radial range and into bandwidth requirements for depth data acquisition. Finally, we present a number of practical applications of our system. Bernhard Kainz, Stefan Hauswiesner, Gerhard Reitmayr, Markus Steinberger, Raphaël Grasset, Lukas Gruber, Eduardo E. Veas, Denis Kalkofen, Hartmut Seichter, Dieter Schmalstieg |
VRST | 1 |
| 2012 | Ray prioritization using stylization and visual saliency
Markus Steinberger, Bernhard Kainz, Stefan Hauswiesner, Rostislav Khlebnikov, Denis Kalkofen, Dieter Schmalstieg |
Comput. Graph. | 2 |
| 2012 | Procedural Texture Synthesis for Zoom-Independent Visualization of Multivariate DataabstractAbstract Simultaneous visualization of multiple continuous data attributes in a single visualization is a task that is important for many application areas. Unsurprisingly, many methods have been proposed to solve this task. However, the behavior of such methods during the exploration stage, when the user tries to understand the data with panning and zooming, has not been given much attention. In this paper, we propose a method that uses procedural texture synthesis to create zoom‐independent visualizations of three scalar data attributes. The method is based on random‐phase Gabor noise, whose frequency is adapted for the visualization of the first data attribute. We ensure that the resulting texture frequency lies in the range that is perceived well by the human visual system at any zoom level. To enhance the perception of this attribute, we also apply a specially constructed transfer function that is based on statistical properties of the noise. Additionally, the transfer function is constructed in a way that it does not introduce any aliasing to the texture. We map the second attribute to the texture orientation. The third attribute is color coded and combined with the texture by modifying the value component of the HSV color model. The necessary contrast needed for texture and color perception was determined in a user study. In addition, we conducted a second user study that shows significant advantages of our method over current methods with similar goals. We believe that our method is an important step towards creating methods that not only succeed in visualizing multiple data attributes, but also adapt to the behavior of the user during the data exploration stage. Rostislav Khlebnikov, Bernhard Kainz, Markus Steinberger, Marc Streit, Dieter Schmalstieg |
Comput. Graph. Forum | 2 |
| 2012 | Softshell: dynamic scheduling on GPUsabstractIn this paper we present Softshell, a novel execution model for devices composed of multiple processing cores operating in a single instruction, multiple data fashion, such as graphics processing units (GPUs). The Softshell model is intuitive and more flexible than the kernel-based adaption of the stream processing model, which is currently the dominant model for general purpose GPU computation. Using the Softshell model, algorithms with a relatively low local degree of parallelism can execute efficiently on massively parallel architectures. Softshell has the following distinct advantages: ( 1 ) work can be dynamically issued directly on the device, eliminating the need for synchronization with an external source, i.e ., the CPU; ( 2 ) its three-tier dynamic scheduler supports arbitrary scheduling strategies, including dynamic priorities and real-time scheduling; and ( 3 ) the user can influence, pause, and cancel work already submitted for parallel execution. The Softshell processing model thus brings capabilities to GPU architectures that were previously only known from operating-system designs and reserved for CPU programming. As a proof of our claims, we present a publicly available implementation of the Softshell processing model realized on top of CUDA. The benchmarks of this implementation demonstrate that our processing model is easy to use and also performs substantially better than the state-of-the-art kernel-based processing model for problems that have been difficult to parallelize in the past. Markus Steinberger, Bernhard Kainz, Bernhard Kerbl, Stefan Hauswiesner, Michael Kenzel, Dieter Schmalstieg |
ACM Trans. Graph. | 2 |
| 2011 | Using perceptual features to prioritize ray-based image generationabstractA common challenge in interactive image generation is maintaining high interactivity of the applications that use computationally demanding rendering algorithms. This is usually achieved by sacrificing some of the image quality in order to decrease the rendering time. Most of such algorithms achieve interactive frame rates while trying to preserve as much image quality as possible by applying the reduction steps non-uniformly. However, high-end rendering systems, such as those presented by [Parker et al. 2010], aim to generate highly realistic images of very complex scenes. In such systems ordinary sampling approaches often give visually unacceptable results. In order to allow to optimize the ratio between the sampling rate of the scene and its resulting perceptual quality, we present a new sampling strategy which uses the information about object features that are known to support the comprehension of 3D shape [Cole et al. 2009]. We control the sampling density by exploiting line extraction techniques commonly used in Non-Photorealistic Renderings. Bernhard Kainz, Markus Steinberger, Stefan Hauswiesner, Rostislav Khlebnikov, Denis Kalkofen, Dieter Schmalstieg |
SI3D | 1 |
| 2011 | Crepuscular Rays for Tumor Accessibility PlanningabstractIn modern clinical practice, planning access paths to volumetric target structures remains one of the most important and most complex tasks, and a physician's insufficient experience in this can lead to severe complications or even the death of the patient. In this paper, we present a method for safety evaluation and the visualization of access paths to assist physicians during preoperative planning. As a metaphor for our method, we employ a well-known, and thus intuitively perceivable, natural phenomenon that is usually called crepuscular rays. Using this metaphor, we propose several ways to compute the safety of paths from the region of interest to all tumor voxels and show how this information can be visualized in real-time using a multi-volume rendering system. Furthermore, we show how to estimate the extent of connected safe areas to improve common medical 2D multi-planar reconstruction (MPR) views. We evaluate our method by means of expert interviews, an online survey, and a retrospective evaluation of 19 real abdominal radio-frequency ablation (RFA) interventions, with expert decisions serving as a gold standard. The evaluation results show clear evidence that our method can be successfully applied in clinical practice without introducing substantial overhead work for the acting personnel. Finally, we show that our method is not limited to medical applications and that it can also be useful in other fields. Rostislav Khlebnikov, Bernhard Kainz, Judith Muehl, Dieter Schmalstieg |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Vessel Segmentation for Ablation Treatment Planning and Simulation
Tuomas Alhonnoro, Mika Pollari, Mikko Lilja, Ronan Flanagan, Bernhard Kainz, Judith Muehl, Ursula Mayrhauser, Rupert H. Portugaller, Philipp Stiegler, Karlheinz Tscheliessnigg |
MICCAI (1) | 5 |
| 2009 | Ray casting of multiple volumetric datasets with polyhedral boundaries on manycore GPUsabstractWe present a new GPU-based rendering system for ray casting of multiple volumes. Our approach supports a large number of volumes, complex translucent and concave polyhedral objects as well as CSG intersections of volumes and geometry in any combination. The system (including the rasterization stage) is implemented entirely in CUDA, which allows full control of the memory hierarchy, in particular access to high bandwidth and low latency shared memory. High depth complexity, which is problematic for conventional approaches based on depth peeling, can be handled successfully. As far as we know, our approach is the first framework for multivolume rendering which provides interactive frame rates when concurrently rendering more than 50 arbitrarily overlapping volumes on current graphics hardware. Bernhard Kainz, Markus Grabner, Alexander Bornik, Stefan Hauswiesner, Judith Muehl, Dieter Schmalstieg |
ACM Trans. Graph. | 1 |
| 2009 | In vivo interactive visualization of four-dimensional blood flow patterns
Bernhard Kainz, Ursula Reiter, Gert Reiter, Dieter Schmalstieg |
Vis. Comput. | 1 |
| 2008 | Fast Marker Based C-Arm Pose Estimation
Bernhard Kainz, Markus Grabner, Matthias Rüther |
MICCAI (2) | 1 |