Ender Konukoglu

dblp:45/7041 · DBLP profile ↗
← Back
78ranked-venue papers
11as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 44 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 28 · 1 first-author · 21 since 2021
YearPublicationVenuePosition
2026 Multi-structure segmentation in CBCT volumes: The ToothFairy2 challenge
abstract
Cone-beam computed tomography (CBCT) is widely used for dento-maxillofacial diagnostics and treatment planning, and comprehensive multi-structure segmentation remains time-consuming, limiting large-scale, reproducible research. In this article, we present ToothFairy2, a MICCAI 2024 challenge on multi-structure segmentation in maxillofacial CBCT. The accompanying dataset comprises 530 CBCT volumes (480 public training, 50 hidden test) with expert 3D annotations of 42 classes, including maxilla, mandible, crowns, bridges, implants, inferior alveolar canals, maxillary sinuses, pharynx, and teeth labeled according to the International Tooth Numbering System (FDI). 26 international teams participated in ToothFairy2, and their methods were run and evaluated for voxel-wise multi-class segmentation using a standardized protocol. This report extends the evaluation of teeth to also investigate the current capabilities of tooth detection and FDI numbering. Furthermore, ranking stability was analyzed to assess the robustness of the final challenge outcome. Overall, challenge participants achieved consistently high performance for large, high-contrast structures such as jawbones, pharynx, and most teeth, while maxillary sinuses, dental restorations, and fine structures remain challenging due to class imbalance and metal artifacts. Analysis of tooth-related metrics further revealed that assigning correct FDI numbers was more challenging than delineating individual teeth. By releasing CBCT data, 3D annotations, baseline models, and evaluation code, ToothFairy2 establishes a long-term benchmark to drive the development of automated methods for robust, clinically meaningful multi-structure segmentation in maxillofacial CBCT.
Federico Bolelli, Luca Lumetti, Niels van Nistelrooij, Shankeeth Vinayahalingam, Mattia Di Bartolomeo, Kevin Marchesini, Arrigo Pellacani, Ettore Candeloro, Gabriele Rosati, Tong Xi 0001, Fabian Isensee, Yannick Kirchhoff, Lars Krämer, Maximilian Rokuss, Constantin Ulrich, Klaus H. Maier-Hein, Yuxian Jiang, Yusheng Liu 0001, Lisheng Wang, Haoshen Wang, Zhiming Cui 0001, Zhaohong Pan, Xiaokun Liang, Ender Konukoglu, Marek Wodzinski, Henning Müller, Haipeng Mai, Xiaobing Dang, Shrajan Bhandary, Radu Grosu, Stefaan Bergé, Alexandre Anesi, Costantino Grana
Medical Image Anal.27
2025 A Large-Scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
abstract
3D Gaussian Splatting (3DGS) has become the de facto method of 3D representation in many vision tasks. This calls for the 3D understanding directly in this representation space. To facilitate the research in this direction, we first build a large-scale dataset of 3DGS using the commonly used ShapeNet and ModelNet datasets. Our dataset ShapeSplat consists of 65K objects from 87 unique categories, whose labels are in accordance with the respective datasets. The creation of this dataset utilized the computing equivalent of 2 GPU years on a TITAN XP GPU. We utilize our dataset for unsupervised pretraining and supervised finetuning for classification and segmentation tasks. To this end, we introduce Gaussian-MAE, which highlights the unique benefits of representation learning from Gaussian parameters. Through exhaustive experiments, we provide several valuable insights. In particular, we show that (1) the distribution of the optimized GS centroids significantly differs from the uniformly sampled point cloud (used for initialization) counterpart; (2) this change in distribution results in degradation in classification but improvement in segmentation tasks when using only the centroids; (3) to leverage additional Gaussian parameters, we propose Gaussian feature grouping in a normalized feature space, along with splats pooling layer, offering a tailored solution to effectively group and embed similar Gaussians, which leads to notable improvement in finetuning tasks. Our dataset and model are publicly available at ShapeSplat.
Yue Li 0036, Bin Ren 0005, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Danda Pani Paudel
3DV5
2025 Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
abstract
Generalized few-shot 3D point cloud segmentation (GFS-PCS) adapts models to new classes with few support samples while retaining base class segmentation. Existing GFS-PCS methods enhance prototypes via interacting with support or query features but remain limited by sparse knowledge from few-shot samples. Meanwhile, 3D vision-language models (3D VLMs), generalizing across open-world novel classes, contain rich but noisy novel class knowledge. In this work, we introduce a GFS-PCS framework that synergizes dense but noisy pseudo-labels from 3D VLMs with precise yet sparse few-shot samples to maximize the strengths of both, named GFS-VL. Specifically, we present a prototype-guided pseudo-label selection to filter low-quality regions, followed by an adaptive infilling strategy that combines knowledge from pseudo-label contexts and few-shot samples to adaptively label the filtered, un-labeled areas. Additionally, we design a novel-base mix strategy to embed few-shot samples into training scenes, preserving essential context for improved novel class learning. Moreover, recognizing the limited diversity in current GFS-PCS benchmarks, we introduce two challenging benchmarks with diverse novel classes for comprehensive generalization evaluation. Experiments validate the effectiveness of our framework across models and datasets. Our approach and benchmarks provide a solid foundation for advancing GFS-PCS in the real world. The code is at here.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Junlin Han, Ender Konukoglu, Serge J. Belongie
CVPR6
2025 Exploiting Temporal State Space Sharing for Video Semantic Segmentation
abstract
Video semantic segmentation (VSS) plays a vital role in understanding the temporal evolution of scenes. Traditional methods often segment videos frame-by-frame or in a short temporal window, leading to limited temporal context, redundant computations, and heavy memory requirements. To this end, we introduce a Temporal Video State Space Sharing (TV3S) architecture to leverage Mamba state space models for temporal feature sharing. Our model features a selective gating mechanism that efficiently propagates relevant information across video frames, eliminating the need for a memory-heavy feature pool. By processing spatial patches independently and incorporating shifted operation, TV3S supports highly parallel computation in both training and inference stages, which reduces the delay in sequential state space processing and improves the scalability for long video sequences. Moreover, TV3S incorporates information from prior frames during inference, achieving long-range temporal coherence and superior adaptability to extended sequences. Evaluations on the VSPW and Cityscapes datasets reveal that our approach outperforms current state-of-the-art methods, establishing a new standard for VSS with consistent results across long video sequences. By achieving a good balance between accuracy and efficiency, TV3S shows a significant advancement in spatiotemporal modeling, paving the way for efficient video analysis. The code is publicly available at https://github.com/Ashesham/TV3S.git.
Syed Ariff Syed Hesham, Yun Liu 0011, Guolei Sun, Henghui Ding, Ender Konukoglu, Xue Geng, Xudong Jiang 0001
CVPR6
2025 SceneSplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining
abstract
Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines.
Yue Li 0036, Runyi Yang, Huapeng Li, Mengjiao Ma, Bin Ren 0005, Nikola Popovic 0001, Nicu Sebe, Ender Konukoglu, Theo Gevers, Luc Van Gool, Martin R. Oswald, Danda Pani Paudel
ICCV9
2025 MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
abstract
High-resolution (HR) medical videos are vital for accurate diagnosis, yet are hard to acquire due to hardware limitations and physiological constraints. Clinically, the collected low-resolution (LR) medical videos present unique challenges for video super-resolution (VSR) models, including camera shake, noise, and abrupt frame transitions, which result in significant optical flow errors and alignment difficulties. Additionally, tissues and organs exhibit continuous and nuanced structures, but current VSR models are prone to introducing artifacts and distorted features that can mislead doctors. To this end, we propose MedVSR, a tailored framework for medical VSR. It first employs Cross State-Space Propagation (CSSP) to address the imprecise alignment by projecting distant frames as control matrices within state-space models, enabling the selective propagation of consistent and informative features to neighboring frames for effective alignment. Moreover, we design an Inner State-Space Reconstruction (ISSR) module that enhances tissue structures and reduces artifacts with joint long-range spatial feature learning and large-kernel short-range information aggregation. Experiments across four datasets in diverse medical scenarios, including endoscopy and cataract surgeries, show that MedVSR significantly outperforms existing VSR models in reconstruction performance and efficiency. Code released at https://github.com/CUHK-AIM-Group/MedVSR.
Xinyu Liu 0001, Guolei Sun, Cheng Wang 0043, Yixuan Yuan, Ender Konukoglu
ICCV5
2025 Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
abstract
Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud inputs, overlooking the potential benefits of leveraging multimodal information. In this paper, we address this gap by introducing a multimodal FS-PCS setup, utilizing textual labels and the potentially available 2D image modality. Under this easy-to-achieve setup, we present the MultiModal Few-Shot SegNet (MM-FSS), a model effectively harnessing complementary information from multiple modalities. MM-FSS employs a shared backbone with two heads to extract intermodal and unimodal visual features, and a pretrained text encoder to generate text embeddings. To fully exploit the multimodal information, we propose a Multimodal Correlation Fusion (MCF) module to generate multimodal correlations, and a Multimodal Semantic Fusion (MSF) module to refine the correlations using text-aware semantic guidance. Additionally, we propose a simple yet effective Test-time Adaptive Cross-modal Calibration (TACC) technique to mitigate training bias, further improving generalization. Experimental results on S3DIS and ScanNet datasets demonstrate significant performance improvements achieved by our method. The efficacy of our approach indicates the benefits of leveraging commonly-ignored free modalities for FS-PCS, providing valuable insights for future research. The code is available at github.com/ZhaochongAn/Multimodality-3D-Few-Shot.
Zhaochong An, Guolei Sun, Yun Liu 0011, Runjia Li, Min Wu 0008, Ming-Ming Cheng, Ender Konukoglu, Serge J. Belongie
ICLR7
2025 Uncertainty modeling for fine-tuned implicit functions
abstract
Implicit functions such as Neural Radiance Fields (NeRFs), occupancy networks, and signed distance functions (SDFs) have become pivotal in computer vision for reconstructing detailed object shapes from sparse views. Achieving optimal performance with these models can be challenging due to the extreme sparsity of inputs and distribution shifts induced by data corruptions. To this end, large, noise-free synthetic datasets can serve as shape priors to help models fill in gaps, but the resulting reconstructions must be approached with caution. Uncertainty estimation is crucial for assessing the quality of these reconstructions, particularly in identifying areas where the model is uncertain about the parts it has inferred from the prior. In this paper, we introduce Dropsembles, a novel method for uncertainty estimation in tuned implicit functions. We demonstrate the efficacy of our approach through a series of experiments, starting with toy examples and progressing to a real-world scenario. Specifically, we train a Convolutional Occupancy Network on synthetic anatomical data and test it on low-resolution MRI segmentations of the lumbar spine. Our results show that Dropsembles achieve the accuracy and calibration levels of deep ensembles but with significantly less computational cost.
Anna Susmelj, Mael Macuglia, Natasa Tagasovska, Reto Sutter, Sebastiano Caprara, Jean-Philippe Thiran, Ender Konukoglu
ICLR7
2025 Conformal Forecasting for Surgical Instrument Trajectory
Sara Sangalli, Gary Sarwin, Ertunc Erdil, Carlo Serra, Alessandro Carretta, Victor E. Staartjes, Ender Konukoglu
MICCAI (9)7
2025 Anomaly Detection by Clustering DINO Embeddings Using a Dirichlet Process Mixture
Nico Schulthess, Ender Konukoglu
MICCAI (3)2
2025 SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into three main groups: (i) per-scene optimization-based, (ii) per-scene optimization-free, and (iii) generalizable approach. However, most of them are evaluated only on rendered 2D views of a handful of scenes and viewpoints close to the training views, limiting ability and insight into holistic 3D understanding. To address this gap, we propose the first large-scale benchmark that systematically assesses these three groups of methods directly in 3D space, evaluating on 1060 scenes across three indoor datasets and one outdoor dataset. Benchmark results demonstrate a clear advantage of the generalizable paradigm, particularly in relaxing the scene-specific limitation, enabling fast feed-forward inference on novel scenes, and achieving superior segmentation performance. We further introduce SceneSplat-49K -- a carefully curated 3DGS dataset comprising of around 49K diverse indoor and outdoor scenes trained from multiple sources, with which we demonstrate generalizable approach could harness strong data priors. Our codes, benchmark, and datasets are available.
Mengjiao Ma, Yue Li 0036, Jiahuan Cheng, Runyi Yang, Bin Ren 0005, Nikola Popovic 0001, Mingqiang Wei, Nicu Sebe, Ender Konukoglu, Luc Van Gool, Theo Gevers, Martin R. Oswald, Danda Pani Paudel
NeurIPS10
2025 CamSAM2: Segment Anything Accurately in Camouflaged Videos
abstract
Video camouflaged object segmentation (VCOS), aiming at segmenting camouflaged objects that seamlessly blend into their environment, is a fundamental vision task with various real-world applications. With the release of SAM2, video segmentation has witnessed significant progress. However, SAM2's capability of segmenting camouflaged videos is suboptimal, especially when given simple prompts such as point and box. To address the problem, we propose Camouflaged SAM2 (CamSAM2), which enhances SAM2's ability to handle camouflaged scenes without modifying SAM2's parameters. Specifically, we introduce a decamouflaged token to provide the flexibility of feature adjustment for VCOS. To make full use of fine-grained and high-resolution features from the current frame and previous frames, we propose implicit object-aware fusion (IOF) and explicit object-aware fusion (EOF) modules, respectively. Object prototype generation (OPG) is introduced to abstract and memorize object prototypes with informative details using high-quality features from previous frames. Extensive experiments are conducted to validate the effectiveness of our approach. While CamSAM2 only adds negligible learnable parameters to SAM2, it substantially outperforms SAM2 on three VCOS datasets, especially achieving 12.2 mDice gains with click prompt on MoCA-Mask and 19.6 mDice gains with mask prompt on SUN-SEG-Hard, with Hiera-T as the backbone. The code is available at https://github.com/zhoustan/CamSAM2.
Yuli Zhou, Yawei Li 0001, Yuqian Fu, Luca Benini, Ender Konukoglu, Guolei Sun
NeurIPS5
2025 Generalizable Single-Source Cross-Modality Medical Image Segmentation via Invariant Causal Mechanisms
abstract
Single-source domain generalization (SDG) aims to learn a model from a single source domain that can generalize well on unseen target domains. This is an important task in computer vision, particularly relevant to medical imaging where domain shifts are common. In this work, we consider a challenging yet practical setting: SDG for cross-modality medical image segmentation. We combine causality-inspired theoretical insights on learning domain-invariant representations with recent advancements in diffusion-based augmentation to improve generalization across diverse imaging modalities. Guided by the “intervention-augmentation equivariant” principle, we use controlled diffusion models (DMs) to simulate diverse imaging styles while preserving the content, leveraging rich generative priors in large-scale pretrained DMs to comprehensively perturb the multidimensional style variable. Extensive experiments on challenging cross-modality segmentation tasks demonstrate that our approach consistently outperforms state-of-the-art SDG methods across three distinct anatomies and imaging modalities. The source code is available at https://github.com/ratschlab/ICMSeg.
Boqi Chen, Yuanzhi Zhu 0001, Yunke Ao, Sebastiano Caprara, Reto Sutter, Gunnar Rätsch, Ender Konukoglu, Anna Susmelj
WACV7
2025 Neural Vector Fields for Implicit Surface Representation and Inference
abstract
Abstract Neural implicit fields have recently shown increasing success in representing, learning and analysis of 3D shapes. Signed distance fields and occupancy fields are still the preferred choice of implicit representations with well-studied properties, despite their restriction to closed surfaces. With neural networks, unsigned distance fields as well as several other variations and training principles have been proposed with the goal to represent all classes of shapes. In this paper, we develop a novel and yet a fundamental representation considering unit vectors in 3D space and call it Vector Field (VF). At each point in $$\mathbb {R}^3$$ R 3 , VF is directed to the closest point on the surface. We theoretically demonstrate that VF can be easily transformed to surface density by computing the flux density. Unlike other standard representations, VF directly encodes an important physical property of the surface, its normal. We further show the advantages of VF representation, in learning open, closed, or multi-layered surfaces. We show that, thanks to the continuity property of the neural optimization with VF, a separate distance field becomes unnecessary for extracting surfaces from the implicit field via Marching Cubes. We compare our method on several datasets including ShapeNet where the proposed new neural implicit field shows superior accuracy in representing any type of shape, outperforming other standard methods. Codes are available at https://github.com/edomel/ImplicitVF .
Edoardo Mello Rella, Ajad Chhatkuli, Ender Konukoglu, Luc Van Gool
Int. J. Comput. Vis.3
2025 Learning to segment anatomy and lesions from disparately labeled sources in brain MRI
abstract
Segmenting healthy tissue structures alongside lesions in brain Magnetic Resonance Images (MRI) remains a challenge for today's algorithms due to lesion-caused disruption of the anatomy and lack of jointly labeled training datasets, where both healthy tissues and lesions are labeled on the same images. In this paper, we propose a method that is robust to lesion-caused disruptions and can be trained from disparately labeled training sets, i.e., without requiring jointly labeled samples, to automatically segment both. In contrast to prior work, we decouple healthy tissue and lesion segmentation in two paths to leverage multi-sequence acquisitions and merge information with an attention mechanism. During inference, an image-specific adaptation reduces adverse influences of lesion regions on healthy tissue predictions. During training, the adaptation is taken into account through meta-learning and co-training is used to learn from disparately labeled training images. Our model shows an improved performance on several anatomical structures and lesions on a publicly available brain glioblastoma dataset compared to the state-of-the-art segmentation methods.
Meva Himmetoglu, Ilja Ciernik, Ender Konukoglu
Medical Image Anal.3
2024 Vision-Based Neurosurgical Guidance: Unsupervised Localization and Camera-Pose Prediction
Gary Sarwin, Alessandro Carretta, Victor E. Staartjes, Matteo Zoli, Diego Mazzatenta, Luca Regli, Carlo Serra, Ender Konukoglu
MICCAI (6)8
2024 Implicit Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
abstract
Neural implicit functions have demonstrated significant importance in various areas such as computer vision, graphics. Their advantages include the ability to represent complex shapes and scenes with high fidelity, smooth interpolation capabilities, and continuous representations. Despite these benefits, the development and analysis of implicit functions have been limited by the lack of comprehensive datasets and the substantial computational resources required for their implementation and evaluation. To address these challenges, we introduce "Implicit-Zoo": a large-scale dataset requiring thousands of GPU training days designed to facilitate research and development in this field. Our dataset includes diverse 2D and 3D scenes, such as CIFAR-10, ImageNet-1K, and Cityscapes for 2D image tasks, and the OmniObject3D dataset for 3D vision tasks. We ensure high quality through strict checks, refining or filtering out low-quality data. Using Implicit-Zoo, we showcase two immediate benefits as it enables to: (1) learn token locations for transformer models; (2) Directly regress 3D cameras poses of 2D images with respect to NeRF models. This in turn leads to an \emph{improved performance} in all three task of image classification, semantic segmentation, and 3D pose regression -- thereby unlocking new avenues for research.
Danda Pani Paudel, Ender Konukoglu, Luc Van Gool
NeurIPS3
2024 Editorial for the Special Issue on the 2022 Medical Imaging with Deep Learning Conference
Shadi Albarqouni, Christian F. Baumgartner, Qi Dou 0001, Ender Konukoglu, Bjoern Menze, Archana Venkataraman
Medical Image Anal.4
2024 Federated Feature Augmentation and Alignment
abstract
Federated learning is a distributed paradigm that allows multiple parties to collaboratively train deep learning models without direct exchange of raw data. Nevertheless, the inherent non-independent and identically distributed (non-i.i.d.) nature of data distribution among clients results in significant degradation of the acquired model. The primary goal of this study is to develop a robust federated learning algorithm to addressfeature shiftin clients’ samples, potentially arising from a range of factors such as acquisition discrepancies in medical imaging. To reach this goal, we first propose federated feature augmentation (FedFA$^{l}$), a novel feature augmentation technique tailored for federated learning.FedFA$^{l}$is based on a crucial insight that each client's data distribution can be characterized by first-/second-order statistics (a.k.a., mean and standard deviation) of latent features; and it is feasible to manipulate these local statisticsglobally, i.e., based on information in the entire federation, to let clients have a better sense of the global distribution across clients. Grounded on this insight, we propose to augment each local feature statistic based on a normal distribution, wherein the mean corresponds to the original statistic, and the variance defines the augmentation scope. Central toFedFA$^{l}$is the determination of a meaningful Gaussian variance, which is accomplished by taking into account not only biased data of each individual client, but also underlying feature statistics represented by all participating clients. Beyond consideration oflow-orderstatistics inFedFA$^{l}$, we propose a federated feature alignment component (FedFA$^{h}$) that exploitshigher-orderfeature statistics to gain a more detailed understanding of local feature distribution and enables explicit alignment of augmented features in different clients to promote more consistent feature learning. CombiningFedFA$^{l}$andFedFA$^{h}$yields our full approachFedFA$+$+.FedFA$+$+is non-parametric, incurs negligible additional communication costs, and can be seamlessly incorporated into popular CNN and Transformer architectures. We offer rigorous theoretical analysis, as well as extensive empirical justifications to demonstrate the effectiveness of the algorithm.
Tianfei Zhou, Ye Yuan 0001, Binglu Wang, Ender Konukoglu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Explicitly Minimizing the Blur Error of Variational Autoencoders
Gustav Bredell, Kyriakos Flouris, Krishna Chaitanya, Ertunc Erdil, Ender Konukoglu
ICLR5
2023 FedFA: Federated Feature Augmentation
Tianfei Zhou, Ender Konukoglu
ICLR2
2023 Canonical normalizing flows for manifold learning
abstract
Manifold learning flows are a class of generative modelling techniques that assume a low-dimensional manifold description of the data. The embedding of such a manifold into the high-dimensional space of the data is achieved via learnable invertible transformations. Therefore, once the manifold is properly aligned via a reconstruction loss, the probability density is tractable on the manifold and maximum likelihood can be used to optimize the network parameters. Naturally, the lower-dimensional representation of the data requires an injective-mapping. Recent approaches were able to enforce that the density aligns with the modelled manifold, while efficiently calculating the density volume-change term when embedding to the higher-dimensional space. However, unless the injective-mapping is analytically predefined, the learned manifold is not necessarily an \emph{efficient representation} of the data. Namely, the latent dimensions of such models frequently learn an entangled intrinsic basis, with degenerate information being stored in each dimension. Alternatively, if a locally orthogonal and/or sparse basis is to be learned, here coined canonical intrinsic basis, it can serve in learning a more compact latent space representation. Toward this end, we propose a canonical manifold learning flow method, where a novel optimization objective enforces the transformation matrix to have few prominent and non-degenerate basis functions. We demonstrate that by minimizing the off-diagonal manifold metric elements $\ell_1$-norm, we can achieve such a basis, which is simultaneously sparse and/or orthogonal. Canonical manifold flow yields a more efficient use of the latent space, automatically generating fewer prominent and distinct dimensions to represent data, and consequently a better approximation of target distributions than other manifold flow methods in most experiments we conducted, resulting in lower FID scores.
Kyriakos Flouris, Ender Konukoglu
NeurIPS2
2023 Expert load matters: operating networks at high accuracy and low manual effort
abstract
In human-AI collaboration systems for critical applications, in order to ensure minimal error, users should set an operating point based on model confidence to determine when the decision should be delegated to human experts. Samples for which model confidence is lower than the operating point would be manually analysed by experts to avoid mistakes. Such systems can become truly useful only if they consider two aspects: models should be confident only for samples for which they are accurate, and the number of samples delegated to experts should be minimized. The latter aspect is especially crucial for applications where available expert time is limited and expensive, such as healthcare. The trade-off between the model accuracy and the number of samples delegated to experts can be represented by a curve that is similar to an ROC curve, which we refer to as confidence operating characteristic (COC) curve. In this paper, we argue that deep neural networks should be trained by taking into account both accuracy and expert load and, to that end, propose a new complementary loss function for classification that maximizes the area under this COC curve. This promotes simultaneously the increase in network accuracy and the reduction in number of samples delegated to humans. We perform experiments on multiple computer vision and medical image datasets for classification. Our results demonstrate that the proposed loss improves classification accuracy and delegates less number of decisions to experts, achieves better out-of-distribution samples detection and on par calibration performance compared to existing loss functions.
Sara Sangalli, Ertunc Erdil, Ender Konukoglu
NeurIPS3
2023 Wiener Guided DIP for Unsupervised Blind Image Deconvolution
abstract
Blind deconvolution is an ill-posed problem arising in various fields ranging from microscopy to astronomy. Its ill-posed nature demands adequate priors and initialization to arrive at a desirable solution. Recently, it has been shown that deep networks can serve as an image generation prior (DIP) during unsupervised blind deconvolution optimization, however, DIP’s high frequency artifact suppression ability is not explicitly exploited. We propose to use Wiener-deconvolution to guide DIP during optimization in order to better leverage DIP’s ability for blind image deconvolution. Wiener-deconvolution sharpens an image while introducing high-frequency artifacts, which are reproduced by DIP with a delay compared to low-frequency features and sharp edges, similar to what has been observed for noise. We embed the computational process in a constrained optimization problem together with an automatic kernel initialization method and show that the proposed method yields higher performance and stability across multiple datasets.
Gustav Bredell, Ertunc Erdil, Bruno Weber, Ender Konukoglu
WACV4
2023 Local contrastive loss with pseudo-label based self-training for semi-supervised medical image segmentation
abstract
Supervised deep learning-based methods yield accurate results for medical image segmentation. However, they require large labeled datasets for this, and obtaining them is a laborious task that requires clinical expertise. Semi/self-supervised learning-based approaches address this limitation by exploiting unlabeled data along with limited annotated data. Recent self-supervised learning methods use contrastive loss to learn good global level representations from unlabeled images and achieve high performance in classification tasks on popular natural image datasets like ImageNet. In pixel-level prediction tasks such as segmentation, it is crucial to also learn good local level representations along with global representations to achieve better accuracy. However, the impact of the existing local contrastive loss-based methods remains limited for learning good local representations because similar and dissimilar local regions are defined based on random augmentations and spatial proximity; not based on the semantic label of local regions due to lack of large-scale expert annotations in the semi/self-supervised setting. In this paper, we propose a local contrastive loss to learn good pixel level features useful for segmentation by exploiting semantic label information obtained from pseudo-labels of unlabeled images alongside limited annotated images with ground truth (GT) labels. In particular, we define the proposed contrastive loss to encourage similar representations for the pixels that have the same pseudo-label/GT label while being dissimilar to the representation of pixels with different pseudo-label/GT label in the dataset. We perform pseudo-label based self-training and train the network by jointly optimizing the proposed contrastive loss on both labeled and unlabeled sets and segmentation loss on only the limited labeled set. We evaluated the proposed approach on three public medical datasets of cardiac and prostate anatomies, and obtain high segmentation performance with a limited labeled set of one or two 3D volumes. Extensive comparisons with the state-of-the-art semi-supervised and data augmentation methods and concurrent contrastive learning methods demonstrate the substantial improvement achieved by the proposed method. The code is made publicly available at https://github.com/krishnabits001/pseudo_label_contrastive_training.
Krishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender Konukoglu
Medical Image Anal.4
2023 Volumetric memory network for interactive medical image segmentation
abstract
Despite recent progress of automatic medical image segmentation techniques, fully automatic results usually fail to meet clinically acceptable accuracy, thus typically require further refinement. To this end, we propose a novel Volumetric Memory Network, dubbed as VMN, to enable segmentation of 3D medical images in an interactive manner. Provided by user hints on an arbitrary slice, a 2D interaction network is firstly employed to produce an initial 2D segmentation for the chosen slice. Then, the VMN propagates the initial segmentation mask bidirectionally to all slices of the entire volume. Subsequent refinement based on additional user guidance on other slices can be incorporated in the same manner. To facilitate smooth human-in-the-loop segmentation, a quality assessment module is introduced to suggest the next slice for interaction based on the segmentation quality of each slice produced in the previous round. Our VMN demonstrates two distinctive features: First, the memory-augmented network design offers our model the ability to quickly encode past segmentation information, which will be retrieved later for the segmentation of other slices; Second, the quality assessment module enables the model to directly estimate the quality of each segmentation prediction, which allows for an active learning paradigm where users preferentially label the lowest-quality slice for multi-round refinement. The proposed network leads to a robust interactive segmentation engine, which can generalize well to various types of user annotations (e.g., scribble, bounding box, extreme clicking). Extensive experiments have been conducted on three public medical image segmentation datasets (i.e., MSD, KiTS19, CVC-ClinicDB), and the results clearly confirm the superiority of our approach in comparison with state-of-the-art segmentation models. The code is made publicly available at https://github.com/0liliulei/Mem3D.
Tianfei Zhou, Liulei Li, Gustav Bredell, Jianwu Li, Jan Unkelbach, Ender Konukoglu
Medical Image Anal.6
2023 InOR-Net: Incremental 3-D Object Recognition Network for Point Cloud Representation
abstract
3-D object recognition has successfully become an appealing research topic in the real world. However, most existing recognition models unreasonably assume that the categories of 3-D objects cannot change over time in the real world. This unrealistic assumption may result in significant performance degradation for them to learn new classes of 3-D objects consecutively due to the catastrophic forgetting on old learned classes. Moreover, they cannot explore which 3-D geometric characteristics are essential to alleviate the catastrophic forgetting on old classes of 3-D objects. To tackle the above challenges, we develop a novel Incremental 3-D Object Recognition Network (i.e., InOR-Net), which could recognize new classes of 3-D objects continuously by overcoming the catastrophic forgetting on old classes. Specifically, category-guided geometric reasoning is proposed to reason local geometric structures with distinctive 3-D characteristics of each class by leveraging intrinsic category information. We then propose a novel critic-induced geometric attention mechanism to distinguish which 3-D geometric characteristics within each class are beneficial to overcome the catastrophic forgetting on old classes of 3-D objects while preventing the negative influence of useless 3-D characteristics. In addition, a dual adaptive fairness compensations' strategy is designed to overcome the forgetting brought by class imbalance by compensating biased weights and predictions of the classifier. Comparison experiments verify the state-of-the-art performance of the proposed InOR-Net model on several public point cloud datasets.
Jiahua Dong 0001, Yang Cong, Gan Sun, Lixu Wang, Lingjuan Lyu, Jun Li 0027, Ender Konukoglu
IEEE Trans. Neural Networks Learn. Syst.7
2022 ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image Prior
abstract
Recent works show that convolutional neural network (CNN) architectures have a spectral bias towards lower frequencies, which has been leveraged for various image restoration tasks in the Deep Image Prior (DIP) framework. The benefit of the inductive bias the network imposes in the DIP framework depends on the architecture. Therefore, re-searchers have studied how to automate the search to de-termine the best-performing model. However, common neu-ral architecture search (NAS) techniques are resource and time-intensive. Moreover, best-performing models are de-termined for a whole dataset of images instead of for each image independently, which would be prohibitively expen-sive. In this work, we first show that optimal neural archi-tectures in the DIP framework are image-dependent. Lever-aging this insight, we then propose an image-specific NAS strategy for the DIP framework that requires substantially less training than typical NAS approaches, effectively en-abling image-specific NAS. We justify the proposed strat-egy's effectiveness by (1) demonstrating its performance on a NAS Dataset for DIP that includes 522 models from a particular search space (2) conducting extensive experi-ments on image denoising, inpainting, and super-resolution tasks. Our experiments show that image-specific metrics can reduce the search space to a small cohort of models, of which the best model outperforms current NAS approaches for image restoration. Codes and datasets are available at https://github.com/ozgurkara99/ISNAS-DIP.
Metin Ersin Arican, Özgür Kara, Gustav Bredell, Ender Konukoglu
CVPR4
2022 Rethinking Semantic Segmentation: A Prototype View
abstract
Prevalent semantic segmentation solutions, despite their different network designs (FCN based or attention based) and mask decoding strategies (parametric softmax based or pixel-query based), can be placed in one category, by considering the softmax weights or query vectors as learnable class prototypes. In light of this prototype view, this study uncovers several limitations of such parametric segmentation regime, and proposes a nonparametric alternative based on non-learnable prototypes. Instead of prior methods learning a single weight/query vector for each class in a fully parametric manner, our model represents each class as a set of non-learnable prototypes, relying solely on the mean fea-tures of several training pixels within that class. The dense prediction is thus achieved by nonparametric nearest prototype retrieving. This allows our model to directly shape the pixel embedding space, by optimizing the arrangement between embedded pixels and anchored prototypes. It is able to handle arbitrary number of classes with a constant amount of learnable parameters. We empirically show that, with FCN based and attention based segmentation models (i.e., HR-Net, Swin, SegFormer) and backbones (i.e., ResNet, HRNet, Swin, MiT), our nonparametric framework yields compel-ling results over several datasets (i.e., ADE20K, Cityscapes, COCO-Stuff), and performs well in the large-vocabulary situation. We expect this work will provoke a rethink of the current de facto semantic segmentation model design.
Tianfei Zhou, Wenguan Wang, Ender Konukoglu, Luc Van Gool
CVPR3
2022 Zero Pixel Directional Boundary by Vector Transform
Edoardo Mello Rella, Ajad Chhatkuli, Yun Liu 0011, Ender Konukoglu, Luc Van Gool
ICLR4
2022 Sampling Possible Reconstructions of Undersampled Acquisitions in MR Imaging With a Deep Learned Prior
abstract
Undersampling the k-space during MR acquisitions saves time, however results in an ill-posed inversion problem, leading to an infinite set of images as possible solutions. Traditionally, this is tackled as a reconstruction problem by searching for a single "best" image out of this solution set according to some chosen regularization or prior. This approach, however, misses the possibility of other solutions and hence ignores the uncertainty in the inversion process. In this paper, we propose a method that instead returns multiple images which are possible under the acquisition model and the chosen prior to capture the uncertainty in the inversion process. To this end, we introduce a low dimensional latent space and model the posterior distribution of the latent vectors given the acquisition data in k-space, from which we can sample in the latent space and obtain the corresponding images. We use a variational autoencoder for the latent model and the Metropolis adjusted Langevin algorithm for the sampling. We evaluate our method on two datasets; with images from the Human Connectome Project and in-house measured multi-coil images. We compare to five alternative methods. Results indicate that the proposed method produces images that match the measured k-space data better than the alternatives, while showing realistic structural variability. Furthermore, in contrast to the compared methods, the proposed method yields higher uncertainty in the undersampled phase encoding direction, as expected.
Kerem Can Tezcan, Neerav Karani, Christian F. Baumgartner, Ender Konukoglu
IEEE Trans. Medical Imaging4
2021 Exploring Cross-Image Pixel Contrast for Semantic Segmentation
abstract
Current semantic segmentation methods focus only on mining "local" context, i.e., dependencies between pixels within individual images, by context-aggregation modules (e.g., dilated convolution, neural attention) or structure-aware optimization criteria (e.g., IoU-like loss). However, they ignore "global" context of the training data, i.e., rich semantic relations between pixels across different images. Inspired by recent advance in unsupervised contrastive representation learning, we propose a pixel-wise contrastive algorithm for semantic segmentation in the fully supervised setting. The core idea is to enforce pixel embeddings belonging to a same semantic class to be more similar than embeddings from different classes. It raises a pixel-wise metric learning paradigm for semantic segmentation, by explicitly exploring the structures of labeled pixels, which were rarely explored before. Our method can be effortlessly incorporated into existing segmentation frameworks without extra overhead during testing. We experimentally show that, with famous segmentation models (i.e., DeepLabV3, HRNet, OCR) and backbones (i.e., ResNet, HRNet), our method brings performance improvements across diverse datasets (i.e., Cityscapes, PASCAL-Context, COCO-Stuff, CamVid). We expect this work will encourage our community to rethink the current de facto training paradigm in semantic segmentation.
Wenguan Wang, Tianfei Zhou, Fisher Yu 0001, Jifeng Dai, Ender Konukoglu, Luc Van Gool
ICCV5
2021 High-Resolution Segmentation of Lumbar Vertebrae from Conventional Thick Slice MRI
Federico Turella, Gustav Bredell, Alexander Okupnik, Sebastiano Caprara, Dimitri Graf, Reto Sutter, Ender Konukoglu
MICCAI (1)7
2021 Quality-Aware Memory Network for Interactive Volumetric Image Segmentation
Tianfei Zhou, Liulei Li, Gustav Bredell, Jianwu Li, Ender Konukoglu
MICCAI (2)5
2021 Constrained Optimization to Train Neural Networks on Critical and Under-Represented Classes
abstract
Deep neural networks (DNNs) are notorious for making more mistakes for the classes that have substantially fewer samples than the others during training. Such class imbalance is ubiquitous in clinical applications and very crucial to handle because the classes with fewer samples most often correspond to critical cases (e.g., cancer) where misclassifications can have severe consequences.Not to miss such cases, binary classifiers need to be operated at high True Positive Rates (TPRs) by setting a higher threshold, but this comes at the cost of very high False Positive Rates (FPRs) for problems with class imbalance. Existing methods for learning under class imbalance most often do not take this into account. We argue that prediction accuracy should be improved by emphasizing the reduction of FPRs at high TPRs for problems where misclassification of the positive, i.e. critical, class samples are associated with higher cost.To this end, we pose the training of a DNN for binary classification as a constrained optimization problem and introduce a novel constraint that can be used with existing loss functions to enforce maximal area under the ROC curve (AUC) through prioritizing FPR reduction at high TPR. We solve the resulting constrained optimization problem using an Augmented Lagrangian method (ALM).Going beyond binary, we also propose two possible extensions of the proposed constraint for multi-class classification problems.We present experimental results for image-based binary and multi-class classification applications using an in-house medical imaging dataset, CIFAR10, and CIFAR100. Our results demonstrate that the proposed method improves the baselines in majority of the cases by attaining higher accuracy on critical classes while reducing the misclassification rate for the non-critical class samples.
Sara Sangalli, Ertunc Erdil, Andreas M. Hötker, Olivio Donati, Ender Konukoglu
NeurIPS5
2021 Semi-supervised task-driven data augmentation for medical image segmentation
abstract
Supervised learning-based segmentation methods typically require a large number of annotated training data to generalize well at test time. In medical applications, curating such datasets is not a favourable option because acquiring a large number of annotated samples from experts is time-consuming and expensive. Consequently, numerous methods have been proposed in the literature for learning with limited annotated examples. Unfortunately, the proposed approaches in the literature have not yet yielded significant gains over random data augmentation for image segmentation, where random augmentations themselves do not yield high accuracy. In this work, we propose a novel task-driven data augmentation method for learning with limited labeled data where the synthetic data generator, is optimized for the segmentation task. The generator of the proposed method models intensity and shape variations using two sets of transformations, as additive intensity transformations and deformation fields. Both transformations are optimized using labeled as well as unlabeled examples in a semi-supervised framework. Our experiments on three medical datasets, namely cardiac, prostate and pancreas, show that the proposed approach significantly outperforms standard augmentation and semi-supervised approaches for image segmentation in the limited annotation setting. The code is made publicly available at https://github.com/krishnabits001/task_driven_data_augmentation.
Krishna Chaitanya, Neerav Karani, Christian F. Baumgartner, Ertunc Erdil, Anton S. Becker, Olivio Donati, Ender Konukoglu
Medical Image Anal.7
2021 Normative ascent with local gaussians for unsupervised lesion detection
abstract
Unsupervised abnormality detection is an appealing approach to identify patterns that are not present in training data without specific annotations for such patterns. In the medical imaging field, methods taking this approach have been proposed to detect lesions. The appeal of this approach stems from the fact that it does not require lesion-specific supervision and can potentially generalize to any sort of abnormal patterns. The principle is to train a generative model on images from healthy individuals to estimate the distribution of images of the normal anatomy, i.e., a normative distribution, and detect lesions as out-of-distribution regions. Restoration-based techniques that modify a given image by taking gradient ascent steps with respect to a posterior distribution composed of a normative distribution and a likelihood term recently yielded state-of-the-art results. However, these methods do not explicitly model ascent directions with respect to the normative distribution, i.e. normative ascent direction, which is essential for successful restoration. In this work, we introduce a novel approach for unsupervised lesion detection by modeling normative ascent directions. We present different modelling options based on the defined ascent directions with local Gaussians. We further extend the proposed method to efficiently utilize 3D information, which has not been explored in most existing works. We experimentally show that the proposed method provides higher accuracy in detection and produces more realistic restored images. The performance of the proposed method is evaluated against baselines on publicly available BRATS and ATLAS stroke lesion datasets; the detection accuracy of the proposed method surpasses the current state-of-the-art results.
Xiaoran Chen, Nick Pawlowski, Ben Glocker, Ender Konukoglu
Medical Image Anal.4
2021 Test-time adaptable neural networks for robust medical image segmentation
abstract
Convolutional Neural Networks (CNNs) work very well for supervised learning problems when the training dataset is representative of the variations expected to be encountered at test time. In medical image segmentation, this premise is violated when there is a mismatch between training and test images in terms of their acquisition details, such as the scanner model or the protocol. Remarkable performance degradation of CNNs in this scenario is well documented in the literature. To address this problem, we design the segmentation CNN as a concatenation of two sub-networks: a relatively shallow image normalization CNN, followed by a deep CNN that segments the normalized image. We train both these sub-networks using a training dataset, consisting of annotated images from a particular scanner and protocol setting. Now, at test time, we adapt the image normalization sub-network for each test image, guided by an implicit prior on the predicted segmentation labels. We employ an independently trained denoising autoencoder (DAE) in order to model such an implicit prior on plausible anatomical segmentation labels. We validate the proposed idea on multi-center Magnetic Resonance imaging datasets of three anatomies: brain, heart and prostate. The proposed test-time adaptation consistently provides performance improvement, demonstrating the promise and generality of the approach. Being agnostic to the architecture of the deep CNN, the second sub-network, the proposed design can be utilized with any segmentation network to increase robustness to variations in imaging scanners and protocols. Our code is available at: https://github.com/neerakara/test-time-adaptable-neural-networks-for-domain-generalization.
Neerav Karani, Ertunc Erdil, Krishna Chaitanya, Ender Konukoglu
Medical Image Anal.4
2020 Joint Reconstruction and Bias Field Correction for Undersampled MR Imaging
Mélanie Gaillochet, Kerem Can Tezcan, Ender Konukoglu
MICCAI (2)3
2020 Probabilistic 3D Surface Reconstruction from Sparse MRI Information
Katarína Tóthová, Sarah Parisot, Matthew C. H. Lee, Esther Puyol-Antón, Andrew P. King, Marc Pollefeys, Ender Konukoglu
MICCAI (1)7
2020 Modelling the Distribution of 3D Brain MRI Using a 2D Slice VAE
Anna Volokitin, Ertunc Erdil, Neerav Karani, Kerem Can Tezcan, Xiaoran Chen, Luc Van Gool, Ender Konukoglu
MICCAI (7)7
2020 Contrastive learning of global and local features for medical image segmentation with limited annotations
abstract
A key requirement for the success of supervised deep learning is a large labeled dataset - a condition that is difficult to meet in medical image analysis. Self-supervised learning (SSL) can help in this regard by providing a strategy to pre-train a neural network with unlabeled data, followed by fine-tuning for a downstream task with limited annotations. Contrastive learning, a particular variant of SSL, is a powerful technique for learning image-level representations. In this work, we propose strategies for extending the contrastive learning framework for segmentation of volumetric medical images in the semi-supervised setting with limited annotations, by leveraging domain-specific and problem-specific cues. Specifically, we propose (1) novel contrasting strategies that leverage structural similarity across volumetric medical images (domain-specific cue) and (2) a local version of the contrastive loss to learn distinctive representations of local regions that are useful for per-pixel segmentation (problem-specific cue). We carry out an extensive evaluation on three Magnetic Resonance Imaging (MRI) datasets. In the limited annotation setting, the proposed method yields substantial improvements compared to other self-supervision and semi-supervised learning techniques. When combined with a simple data augmentation technique, the proposed method reaches within 8\% of benchmark performance using only two labeled MRI volumes for training. The code is made public at https://github.com/krishnabits001/domain_specific_cl.
Krishna Chaitanya, Ertunc Erdil, Neerav Karani, Ender Konukoglu
NeurIPS4
2020 Unsupervised lesion detection via image restoration with a normative prior
Xiaoran Chen, Suhang You, Kerem Can Tezcan, Ender Konukoglu
Medical Image Anal.4
2019 PHiSeg: Capturing Uncertainty in Medical Image Segmentation
Christian F. Baumgartner, Kerem Can Tezcan, Krishna Chaitanya, Andreas M. Hötker, Urs J. Muehlematter, Khoschy Schawkat, Anton S. Becker, Olivio Donati, Ender Konukoglu
MICCAI (2)9
2019 A Partially Reversible U-Net for Memory-Efficient Volumetric Image Segmentation
Robin Brügger, Christian F. Baumgartner, Ender Konukoglu
MICCAI (3)3
2019 Morpho-MNIST: Quantitative Assessment and Diagnostics for Representation Learning
abstract
Revealing latent structure in data is an active field of research, having introduced exciting technologies such as variational autoencoders and adversarial networks, and is essential to push machine learning towards unsupervised knowledge discovery. However, a major challenge is the lack of suitable benchmarks for an objective and quantitative evaluation of learned representations. To address this issue we introduce Morpho-MNIST, a framework that aims to answer: “to what extent has my model learned to represent specific factors of variation in the data? We extend the popular MNIST dataset by adding a morphometric analysis enabling quantitative comparison of trained models, identification of the roles of latent variables, and characterisation of sample diversity. We further propose a set of quantifiable perturbations to assess the performance of unsupervised and supervised methods on challenging tasks such as outlier detection and domain adaptation. Data and code are available at https://github.com/dccastro/Morpho-MNIST.
Daniel C. Castro, Jeremy Tan, Bernhard Kainz, Ender Konukoglu, Ben Glocker
J. Mach. Learn. Res.4
2019 An image interpolation approach for acquisition time reduction in navigator-based 4D MRI
Neerav Karani, Lin Zhang 0031, Christine Tanner, Ender Konukoglu
Medical Image Anal.4
2019 MR Image Reconstruction Using Deep Density Priors
abstract
Algorithms for magnetic resonance (MR) image reconstruction from undersampled measurements exploit prior information to compensate for missing k-space data. Deep learning (DL) provides a powerful framework for extracting such information from existing image datasets, through learning, and then using it for reconstruction. Leveraging this, recent methods employed DL to learn mappings from undersampled to fully sampled images using paired datasets, including undersampled and corresponding fully sampled images, integrating prior knowledge implicitly. In this letter, we propose an alternative approach that learns the probability distribution of fully sampled MR images using unsupervised DL, specifically variational autoencoders (VAE), and use this as an explicit prior term in reconstruction, completely decoupling the encoding operation from the prior. The resulting reconstruction algorithm enjoys a powerful image prior to compensate for missing k-space data without requiring paired datasets for training nor being prone to associated sensitivities, such as deviations in undersampling patterns used in training and test time or coil settings. We evaluated the proposed method with T1 weighted images from a publicly available dataset, multi-coil complex images acquired from healthy volunteers ( N=8 ), and images with white matter lesions. The proposed algorithm, using the VAE prior, produced visually high quality reconstructions and achieved low RMSE values, outperforming most of the alternative methods on the same dataset. On multi-coil complex data, the algorithm yielded accurate magnitude and phase reconstruction results. In the experiments on images with white matter lesions, the method faithfully reconstructed the lesions.
Kerem Can Tezcan, Christian F. Baumgartner, Roger Lüchinger, Klaas P. Pruessmann, Ender Konukoglu
IEEE Trans. Medical Imaging5
2018 Visual Feature Attribution Using Wasserstein GANs
abstract
Attributing the pixels of an input image to a certain category is an important and well-studied problem in computer vision, with applications ranging from weakly supervised localisation to understanding hidden effects in the data. In recent years, approaches based on interpreting a previously trained neural network classifier have become the de facto state-of-the-art and are commonly used on medical as well as natural image datasets. In this paper, we discuss a limitation of these approaches which may lead to only a subset of the category specific features being detected. To address this problem we develop a novel feature attribution technique based on Wasserstein Generative Adversarial Networks (WGAN), which does not suffer from this limitation. We show that our proposed method performs substantially better than the state-of-the-art for visual attribution on a synthetic dataset and on real 3D neuroimaging data from patients with mild cognitive impairment (MCI) and Alzheimer's disease (AD). For AD patients the method produces compellingly realistic disease effect maps which are very close to the observed effects.
Christian F. Baumgartner, Lisa M. Koch, Kerem Can Tezcan, Jia Xi Ang, Ender Konukoglu
CVPR5
2018 A Lifelong Learning Approach to Brain MR Segmentation Across Scanners and Protocols
Neerav Karani, Krishna Chaitanya, Christian F. Baumgartner, Ender Konukoglu
MICCAI (1)4
2018 Deep EndoVO: A recurrent convolutional neural network (RCNN) based visual odometry approach for endoscopic capsule robots
abstract
Ingestible wireless capsule endoscopy is an emerging minimally invasive diagnostic technology for inspection of the GI tract and diagnosis of a wide range of diseases and pathologies. Medical device companies and many research groups have recently made substantial progresses in converting passive capsule endoscopes to active capsule robots, enabling more accurate, precise, and intuitive detection of the location and size of the diseased areas. Since a reliable real time pose estimation functionality is crucial for actively controlled endoscopic capsule robots, in this study, we propose a monocular visual odometry (VO) method for endoscopic capsule robot operations. Our method lies on the application of the deep recurrent convolutional neural networks (RCNNs) for the visual odometry task, where convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are used for the feature extraction and inference of dynamics across the frames, respectively. Detailed analyses and evaluations made on a real pig stomach dataset proves that our system achieves high translational and rotational accuracies for different types of endoscopic capsule robot trajectories.
Mehmet Turan, Yasin Almalioglu, Helder Araújo, Ender Konukoglu, Metin Sitti
Neurocomputing4
2018 Sparse-then-dense alignment-based 3D map reconstruction method for endoscopic capsule robots
abstract
Despite significant progress achieved in the last decade to convert passive capsule endoscopes to actively controllable robots, robotic capsule endoscopy still has some challenges. In particular, a fully dense three-dimensional (3D) map reconstruction of the explored organ remains an unsolved problem. Such a dense map would help doctors detect the locations and sizes of the diseased areas more reliably, resulting in more accurate diagnoses. In this study, we propose a comprehensive medical 3D reconstruction method for endoscopic capsule robots, which is built in a modular fashion including preprocessing, keyframe selection, sparse-then-dense alignment-based pose estimation, bundle fusion, and shading-based 3D reconstruction. A detailed quantitative analysis is performed using a non-rigid esophagus gastroduodenoscopy simulator, four different endoscopic cameras, a magnetically activated soft capsule robot, a sub-millimeter precise optical motion tracker, and a fine-scale 3D optical scanner, whereas qualitative ex-vivo experiments are performed on a porcine pig stomach. To the best of our knowledge, this study is the first complete endoscopic 3D map reconstruction approach containing all of the necessary functionalities for a therapeutically relevant 3D map reconstruction.
Mehmet Turan, Yusuf Yigit Pilavci, Ipek Ganiyusufoglu, Helder Araújo, Ender Konukoglu, Metin Sitti
Mach. Vis. Appl.5
2018 Reducing Navigators in Free-Breathing Abdominal MRI via Temporal Interpolation Using Convolutional Neural Networks
abstract
Navigated 2-D multi-slice dynamic magnetic resonance imaging (MRI) acquisitions are essential for MR guided therapies. This technique yields time-resolved volumetric images during free-breathing, which are ideal for visualizing and quantifying breathing induced motion. To achieve this, navigated dynamic imaging requires acquiring multiple navigator slices. Reducing the number of navigator slices would allow for acquiring more data slices in the same time, and hence, increasing through-plane resolution or alternatively the overall acquisition time can be reduced while keeping resolution unchanged. To this end, we propose temporal interpolation of navigator slices using convolutional neural networks (CNNs). Our goal is to acquire fewer navigators and replace the missing ones with interpolation. We evaluate the proposed method on abdominal navigated dynamic MRI sequences acquired from 14 subjects. Investigations with several CNN architectures and training loss functions show favorable results for cost and a simple feed-forward network with no skip connections. When compared with interpolation by non-linear registration, the proposed method achieves higher interpolation accuracy on average as quantified in terms of root mean square error and residual motion. Analysis of the differences shows that the better performance is due to more accurate interpolation at peak exhalation and inhalation positions. Furthermore, the CNN-based approach requires substantially lower execution times than that of the registration-based method. At last, experiments on dynamic volume reconstruction reveal minimal differences between reconstructions with acquired and interpolated navigator slices.
Neerav Karani, Christine Tanner, Sebastian Kozerke, Ender Konukoglu
IEEE Trans. Medical Imaging4
2017 Temporal Interpolation of Abdominal MRIs Acquired During Free-Breathing
Neerav Karani, Christine Tanner, Sebastian Kozerke, Ender Konukoglu
MICCAI (2)4
2016 Correction of Fat-Water Swaps in Dixon MRI
Ben Glocker, Ender Konukoglu, Ioannis Lavdas, Juan Eugenio Iglesias, Eric O. Aboagye, Andrea G. Rockall, Daniel Rueckert
MICCAI (3)2
2015 The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS)
abstract
In this paper we report the set-up and results of the Multimodal Brain Tumor Image Segmentation Benchmark (BRATS) organized in conjunction with the MICCAI 2012 and 2013 conferences. Twenty state-of-the-art tumor segmentation algorithms were applied to a set of 65 multi-contrast MR scans of low- and high-grade glioma patients-manually annotated by up to four raters-and to 65 comparable scans generated using tumor image simulation software. Quantitative evaluations revealed considerable disagreement between the human raters in segmenting various tumor sub-regions (Dice scores in the range 74%-85%), illustrating the difficulty of this task. We found that different algorithms worked best for different sub-regions (reaching performance comparable to human inter-rater variability), but that no single algorithm ranked in the top for all sub-regions simultaneously. Fusing several good algorithms using a hierarchical majority vote yielded segmentations that consistently ranked above all individual algorithms, indicating remaining opportunities for further methodological improvements. The BRATS image data and manual annotations continue to be publicly available through an online evaluation system as an ongoing benchmarking resource.
Bjoern Menze, András Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin S. Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth R. Gerstner, Marc-André Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buendia, D. Louis Collins, Nicolas Cordier, Jason J. Corso, Antonio Criminisi, Tilak Das, Hervé Delingette, Çagatay Demiralp, Christopher R. Durst, Michel Dojat, Senan Doyle, Joana Festa, Florence Forbes, Ezequiel Geremia, Ben Glocker, Polina Golland, Xiaotao Guo, Andac Hamamci, Khan M. Iftekharuddin, Raj Jena, Nigel M. John, Ender Konukoglu, Danial Lashkari, José Antonio Mariz, Raphael Meier, Sérgio Pereira, Doina Precup, Stephen J. Price, Tammy Riklin-Raviv, Syed M. S. Reza, Michael T. Ryan, Duygu Sarikaya, Lawrence H. Schwartz, Hoo-Chang Shin, Jamie Shotton, Carlos A. Silva 0002, Nuno J. Sousa, Nagesh K. Subbanna, Gábor Székely, Thomas J. Taylor, Owen M. Thomas, Nicholas J. Tustison, Gozde Unal, Flor Vasseur, Max Wintermark, Dong Hye Ye, Liang Zhao 0018, Binsheng Zhao, Darko Zikic, Marcel Prastawa, Mauricio Reyes 0001, Koenraad Van Leemput
IEEE Trans. Medical Imaging38
2013 Vertebrae Localization in Pathological Spine CT via Dense Classification from Sparse Annotations
Ben Glocker, Darko Zikic, Ender Konukoglu, David R. Haynor, Antonio Criminisi
MICCAI (2)3
2013 Is Synthesizing MRI Contrast Useful for Inter-modality Analysis?
Juan Eugenio Iglesias, Ender Konukoglu, Darko Zikic, Ben Glocker, Koenraad Van Leemput, Bruce Fischl
MICCAI (1)2
2013 Example-Based Restoration of High-Resolution Magnetic Resonance Image Acquisitions
Ender Konukoglu, André J. W. van der Kouwe, Mert R. Sabuncu, Bruce Fischl
MICCAI (1)1
2013 Modality Propagation: Coherent Synthesis of Subject-Specific Scans with Data-Driven Regularization
Dong Hye Ye, Darko Zikic, Ben Glocker, Antonio Criminisi, Ender Konukoglu
MICCAI (1)5
2013 Regression forests for efficient anatomy detection and localization in computed tomography scans
Antonio Criminisi, Duncan P. Robertson, Ender Konukoglu, Jamie Shotton, Sayan D. Pathak, Khan M. Siddiqui
Medical Image Anal.3
2013 Neighbourhood approximation using randomized forests
Ender Konukoglu, Ben Glocker, Darko Zikic, Antonio Criminisi
Medical Image Anal.1
2013 WESD-Weighted Spectral Distance for Measuring Shape Dissimilarity
abstract
This paper presents a new distance for measuring shape dissimilarity between objects. Recent publications introduced the use of eigenvalues of the Laplace operator as compact shape descriptors. Here, we revisit the eigenvalues to define a proper distance, called Weighted Spectral Distance (WESD), for quantifying shape dissimilarity. The definition of WESD is derived through analyzing the heat trace. This analysis provides the proposed distance with an intuitive meaning and mathematically links it to the intrinsic geometry of objects. We analyze the resulting distance definition, present and prove its important theoretical properties. Some of these properties include: 1) WESD is defined over the entire sequence of eigenvalues yet it is guaranteed to converge, 2) it is a pseudometric, 3) it is accurately approximated with a finite number of eigenvalues, and 4) it can be mapped to the [0,1) interval. Last, experiments conducted on synthetic and real objects are presented. These experiments highlight the practical benefits of WESD for applications in vision and medical image analysis.
Ender Konukoglu, Ben Glocker, Antonio Criminisi, Kilian M. Pohl
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Joint Classification-Regression Forests for Spatially Structured Multi-object Segmentation
Ben Glocker, Olivier Pauly, Ender Konukoglu, Antonio Criminisi
ECCV (4)3
2012 Temporal Shape Analysis via the Spectral Signature
Elena Bernardis, Ender Konukoglu, Yangming Ou, Dimitris N. Metaxas, Benoit Desjardins, Kilian M. Pohl
MICCAI (2)2
2012 Automatic Localization and Identification of Vertebrae in Arbitrary Field-of-View CT Scans
Ben Glocker, Johannes Feulner, Antonio Criminisi, David R. Haynor, Ender Konukoglu
MICCAI (3)5
2012 Neighbourhood Approximation Forests
abstract
Methods that leverage neighbourhood structures in high-dimensional image spaces have recently attracted attention. These approaches extract information from a new image using its "neighbours" in the image space equipped with an application-specific distance. Finding the neighbourhood of a given image is challenging due to large dataset sizes and costly distance evaluations. Furthermore, automatic neighbourhood search for a new image is currently not possible when the distance is based on ground truth annotations. In this article we present a general and efficient solution to these problems. "neighbourhood approximation forests" (NAF) is a supervised learning algorithm that approximates the neighbourhood structure resulting from an arbitrary distance. As NAF uses only image intensities to infer neighbours it can also be applied to distances based on ground truth annotations. We demonstrate NAF in two scenarios: (i) choosing neighbours with respect to a deformation-based distance, and (ii) age prediction from brain MRI. The experiments show NAF's approximation quality, computational advantages and use in different contexts.
Ender Konukoglu, Ben Glocker, Darko Zikic, Antonio Criminisi
MICCAI (3)1
2012 Decision Forests for Tissue-Specific Segmentation of High-Grade Gliomas in Multi-channel MR
Darko Zikic, Ben Glocker, Ender Konukoglu, Antonio Criminisi, Çagatay Demiralp, Jamie Shotton, Owen M. Thomas, Tilak Das, Raj Jena, Stephen J. Price
MICCAI (3)3
2012 Discriminative Segmentation-Based Evaluation Through Shape Dissimilarity
abstract
Segmentation-based scores play an important role in the evaluation of computational tools in medical image analysis. These scores evaluate the quality of various tasks, such as image registration and segmentation, by measuring the similarity between two binary label maps. Commonly these measurements blend two aspects of the similarity: pose misalignments and shape discrepancies. Not being able to distinguish between these two aspects, these scores often yield similar results to a widely varying range of different segmentation pairs. Consequently, the comparisons and analysis achieved by interpreting these scores become questionable. In this paper, we address this problem by exploring a new segmentation-based score, called normalized Weighted Spectral Distance (nWSD), that measures only shape discrepancies using the spectrum of the Laplace operator. Through experiments on synthetic and real data we demonstrate that nWSD provides additional information for evaluating differences between segmentations, which is not captured by other commonly used scores. Our results demonstrate that when jointly used with other scores, such as Dice's similarity coefficient, the additional information provided by nWSD allows richer, more discriminative evaluations. We show for the task of registration that through this addition we can distinguish different types of registration errors. This allows us to identify the source of errors and discriminate registration results which so far had to be treated as being of similar quality in previous evaluation studies.
Ender Konukoglu, Ben Glocker, Dong Hye Ye, Antonio Criminisi, Kilian M. Pohl
IEEE Trans. Medical Imaging1
2011 A multi-front eikonal model of cardiac electrophysiology for interactive simulation of radio-frequency ablation
Erik Pernod, Maxime Sermesant, Ender Konukoglu, Jatin Relan, Hervé Delingette, Nicholas Ayache
Comput. Graph.3
2010 Spatial Decision Forests for MS Lesion Segmentation in Multi-Channel MR Images
Ezequiel Geremia, Bjoern Menze, Olivier Clatz, Ender Konukoglu, Antonio Criminisi, Nicholas Ayache
MICCAI (1)4
2010 Extrapolating glioma invasion margin in brain magnetic resonance images: Suggesting new irradiation margins
Ender Konukoglu, Olivier Clatz, Pierre-Yves Bondiau, Hervé Delingette, Nicholas Ayache
Medical Image Anal.1
2010 Image Guided Personalization of Reaction-Diffusion Type Tumor Growth Models Using Modified Anisotropic Eikonal Equations
abstract
Reaction-diffusion based tumor growth models have been widely used in the literature for modeling the growth of brain gliomas. Lately, recent models have started integrating medical images in their formulation. Including different tissue types, geometry of the brain and the directions of white matter fiber tracts improved the spatial accuracy of reaction-diffusion models. The adaptation of the general model to the specific patient cases on the other hand has not been studied thoroughly yet. In this paper, we address this adaptation. We propose a parameter estimation method for reaction-diffusion tumor growth models using time series of medical images. This method estimates the patient specific parameters of the model using the images of the patient taken at successive time instances. The proposed method formulates the evolution of the tumor delineation visible in the images based on the reaction-diffusion dynamics; therefore, it remains consistent with the information available. We perform thorough analysis of the method using synthetic tumors and show important couplings between parameters of the reaction-diffusion model. We show that several parameters can be uniquely identified in the case of fixing one parameter, namely the proliferation rate of tumor cells. Moreover, regardless of the value the proliferation rate is fixed to, the speed of growth of the tumor can be estimated in terms of the model parameters with accuracy. We also show that using the model-based speed, we can simulate the evolution of the tumor for the specific patient case. Finally, we apply our method to two real cases and show promising preliminary results.
Ender Konukoglu, Olivier Clatz, Bjoern Menze, Bram Stieltjes, Marc-André Weber, Emmanuel Mandonnet, Hervé Delingette, Nicholas Ayache
IEEE Trans. Medical Imaging1
2007 Towards an Identification of Tumor Growth Parameters from Time Series of Images
Ender Konukoglu, Olivier Clatz, Pierre-Yves Bondiau, Maxime Sermesant, Hervé Delingette, Nicholas Ayache
MICCAI (1)1
2007 HDF: Heat diffusion fields for polyp detection in CT colonography
Ender Konukoglu, Burak Acar
Signal Process.1
2007 Polyp Enhancing Level Set Evolution of Colon Wall: Method and Pilot Study
abstract
Computer aided detection (CAD) in computed tomography colonography (CTC) aims at detecting colonic polyps that are the precursors of colon cancer. In this work, we propose a colon wall evolution algorithm polyp enhancing level sets (PELS) based on the level-set formulation that regularizes and enhances polyps as a preprocessing step to CTC CAD algorithms. The underlying idea is to evolve the polyps towards spherical protrusions on the colon wall while keeping other structures, such as haustral folds, relatively unchanged and, thereby, potentially improve the performance of CTC CAD algorithms, especially for smaller polyps. To evaluate our methods, we conducted a pilot study using an arbitrarily chosen CTC CAD method, the surface normal overlap (SNO) CAD algorithm, on a nine patient CTC data set with 47 polyps of sizes ranging from 2.0 to 17.0 mm in diameter. PELS increased the maximum sensitivity by 8.1% (from 21/37 to 24/37) for small polyps of sizes ranging from 5.0 to 9.0 mm in diameter. This is accompanied by a statistically significant separation between small polyps and false positives. PELS did not change the CTC CAD performance significantly for larger polyps.
Ender Konukoglu, Burak Acar, David S. Paik, Christopher F. Beaulieu, Jarrett Rosenberg, Sandy Napel
IEEE Trans. Medical Imaging1
2006 Extrapolating Tumor Invasion Margins for Physiologically Determined Radiotherapy Regions
Ender Konukoglu, Olivier Clatz, Pierre-Yves Bondiau, Hervé Delingette, Nicholas Ayache
MICCAI (1)1
2006 Shape-Based Hand Recognition
abstract
The problem of person recognition and verification based on their hand images has been addressed. The system is based on the images of the right hands of the subjects, captured by a flatbed scanner in an unconstrained pose at 45 dpi. In a pre-processing stage of the algorithm, the silhouettes of hand images are registered to a fixed pose, which involves both rotation and translation of the hand and, separately, of the individual fingers. Two feature sets have been comparatively assessed, Hausdorff distance of the hand contours and independent component features of the hand silhouette images. Both the classification and the verification performances are found to be very satisfactory as it was shown that, at least for groups of about five hundred subjects, hand-based recognition is a viable secure access control scheme.
Erdem Yörük, Ender Konukoglu, Bülent Sankur, Jérôme Darbon
IEEE Trans. Image Process.2