Fons van der Sommen

dblp:132/2209 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-3593-2356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Symmetrical Flow Matching: Unified Image Generation, Segmentation, and Classification with Score-Based Generative Models
abstract
Flow Matching has emerged as a powerful framework for learning continuous transformations between distributions, enabling high-fidelity generative modeling. This work introduces Symmetrical Flow Matching (SymmFlow), a new formulation that unifies semantic segmentation, classification, and image generation within a single model. Using a symmetric learning objective, SymmFlow models forward and reverse transformations jointly, ensuring bi-directional consistency, while preserving sufficient entropy for generative diversity. A new training objective is introduced to explicitly retain semantic information across flows, featuring efficient sampling while preserving semantic structure, allowing for one-step segmentation and classification without iterative refinement. Unlike previous approaches that impose strict one-to-one mapping between masks and images, SymmFlow generalizes to flexible conditioning, supporting both pixel-level and image-level class labels. Experimental results on various benchmarks demonstrate that SymmFlow achieves state-of-the-art performance on semantic image synthesis, obtaining FID scores of 11.9 on CelebAMask-HQ and 7.0 on COCO-Stuff with only 25 inference steps. Additionally, it delivers competitive results on semantic segmentation and shows promising capabilities in classification tasks.
Francisco Caetano, Christiaan G. A. Viviers, Peter H. N. de With, Fons van der Sommen
AAAI4
2026 Scaling up self-supervised learning for improved surgical foundation models
abstract
• Demonstration of effectiveness of SSL for surgical computer vision using the largest dataset reported to date. • Strong generalization and robust evaluation are shown across six surgical datasets, four procedures, and three tasks, outperforming current SOTA foundation models. • Providing insights into large-scale SSL for surgical computer vision in terms of scaling, pretraining time, dataset composition, and model architecture. • Release of the models and a curated dataset of 2.1 million surgical video frames, establishing a critical resource for advancing surgical foundation model training Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer vision has been limited. This study addresses this gap by introducing SurgeNetXL, a novel surgical foundation model that sets a new benchmark in surgical computer vision. Trained on the largest reported surgical dataset to date, comprising over 4.7 million video frames, SurgeNetXL achieves consistent top-tier performance across six datasets spanning four surgical procedures and three tasks, including semantic segmentation, surgical phase recognition, and critical view of safety (CVS) classification. Compared with the best-performing surgical foundation model, SurgeNetXL shows mean improvements of 4.0%, 8.9%, and 11.4% for semantic segmentation, phase recognition, and CVS classification, respectively. Additionally, SurgeNetXL outperforms ImageNet1k by 16.1%, 8.0%, and 4.3% for the respective tasks. In addition to advancing model performance, this study provides key insights into scaling pretraining datasets, extending training durations, and optimizing model architectures specifically for surgical computer vision. These findings pave the way for improved generalization and robustness in data-scarce scenarios, offering a comprehensive framework for future research in this domain. All models and a subset of the SurgeNetXL dataset, including over 2 million video frames, are publicly available at: https://github.com/TimJaspers0801/SurgeNet .
Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li 0002, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Alkhalil, Fons van der Sommen
Medical Image Anal.15
2025 DisCoPatch: Taming Adversarially-Driven Batch Statistics for Improved Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection holds significant importance across many applications. While semantic and domain-shift OOD problems are well-studied, this work focuses on covariate shifts - subtle variations in the data distribution that can degrade machine learning performance. We hypothesize that detecting these subtle shifts can improve our understanding of in-distribution boundaries, ultimately improving OOD detection. In adversarial discriminators trained with Batch Normalization (BN), real and adversarial samples form distinct domains with unique batch statistics - a property we exploit for OOD detection. We introduce DisCoPatch, an unsupervised Adversarial Variational Autoencoder (VAE) framework that harnesses this mechanism. During inference, batches consist of patches from the same image, ensuring a consistent data distribution that allows the model to rely on batch statistics. DisCoPatch uses the VAE's suboptimal outputs (generated and reconstructed) as negative samples to train the discriminator, thereby improving its ability to delineate the boundary between in-distribution samples and covariate shifts. By tightening this boundary, DisCoPatch achieves state-of-the-art results in public OOD detection benchmarks. The proposed model not only excels in detecting covariate shifts, achieving 95.5% AUROC on ImageNet-1K(-C), but also outperforms all prior methods on public Near-OOD (95.0%) benchmarks. With a compact model size of 25MB, it achieves high OOD detection performance at notably lower latency than existing methods, making it an efficient and practical solution for real-world OOD detection applications. The code is available at github.com/caetas/DisCoPatch.
Francisco Caetano, Christiaan G. A. Viviers, Luis Albert Zavala-Mondragón, Peter H. N. de With, Fons van der Sommen
ICCV5
2025 SemiVT-Surge: Semi-supervised Video Transformer for Surgical Phase Recognition
Yiping Li 0002, Ronald L. P. D. de Jong, Sahar Nasirihaghighi, Tim J. M. Jaspers, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Fons van der Sommen, Jelle P. Ruurda, Marcel Breeuwer, Yasmina Alkhalil
MICCAI (10)8
2025 Learning to recognize correctly completed procedure steps in egocentric assembly videos through spatio-temporal modeling
abstract
Procedure step recognition (PSR) aims to identify all correctly completed steps and their sequential order in videos of procedural tasks. The existing state-of-the-art models rely solely on detecting assembly object states in individual video frames. By neglecting temporal features, model robustness and accuracy are limited, especially when objects are partially occluded. To overcome these limitations, we propose Spatio-Temporal Occlusion-Resilient Modeling for Procedure Step Recognition (STORM-PSR), a dual-stream framework for PSR that leverages both spatial and temporal features. The assembly state detection stream operates effectively with unobstructed views of the object, while the spatio-temporal stream captures both spatial and temporal features to recognize step completions even under partial occlusion. This stream includes a spatial encoder, pre-trained using a novel weakly supervised approach to capture meaningful spatial representations, and a transformer-based temporal encoder that learns how these spatial features relate over time. STORM-PSR is evaluated on the MECCANO and IndustReal datasets, reducing the average delay between actual and predicted assembly step completions by 11.2% and 26.1%, respectively, compared to prior methods. We demonstrate that this reduction in delay is driven by the spatio-temporal stream, which does not rely on unobstructed views of the object to infer completed steps. The code for STORM-PSR, along with the newly annotated MECCANO labels, is made publicly available at https://timschoonbeek.github.io/stormpsr . • Spatio-temporal features reduce prediction delay for procedure step recognition (PSR). • Weakly-supervised key-frame selection enables meaningful spatial features. • Key-clip aware sampling significantly improves training of the temporal encoder. • PSR annotations and benchmark on the MECCANO dataset to stimulate further research.
Tim J. Schoonbeek, Shao-Hsuan Hung, Dan Lehman, Hans Onvlee, Jacek Kustra, Peter H. N. de With, Fons van der Sommen
Comput. Vis. Image Underst.7
2025 Will Transformers change gastrointestinal endoscopic image analysis? A comparative analysis between CNNs and Transformers, in terms of performance, robustness and generalization
abstract
Gastrointestinal endoscopic image analysis presents significant challenges, such as considerable variations in quality due to the challenging in-body imaging environment, the often-subtle nature of abnormalities with low interobserver agreement, and the need for real-time processing. These challenges pose strong requirements on the performance, generalization, robustness and complexity of deep learning-based techniques in such safety-critical applications. While Convolutional Neural Networks (CNNs) have been the go-to architecture for endoscopic image analysis, recent successes of the Transformer architecture in computer vision raise the possibility to update this conclusion. To this end, we evaluate and compare clinically relevant performance, generalization and robustness of state-of-the-art CNNs and Transformers for neoplasia detection in Barrett's esophagus. We have trained and validated several top-performing CNNs and Transformers on a total of 10,208 images (2,079 patients), and tested on a total of 7,118 images (998 patients) across multiple test sets, including a high-quality test set, two internal and two external generalization test sets, and a robustness test set. Furthermore, to expand the scope of the study, we have conducted the performance and robustness comparisons for colonic polyp segmentation (Kvasir-SEG) and angiodysplasia detection (Giana). The results obtained for featured models across a wide range of training set sizes demonstrate that Transformers achieve comparable performance as CNNs on various applications, show comparable or slightly improved generalization capabilities and offer equally strong resilience and robustness against common image corruptions and perturbations. These findings confirm the viability of the Transformer architecture, particularly suited to the dynamic nature of endoscopic video analysis, characterized by fluctuating image quality, appearance and equipment configurations in transition from hospital to hospital. The code is made publicly available at: https://github.com/BONS-AI-VCA-AMC/Endoscopy-CNNs-vs-Transformers.
Carolus H. J. Kusters, Tim J. M. Jaspers, T. G. W. Boers, Martijn R. Jong, Jelmer Jukema, Kiki Fockens, Albert Jeroen de Groof, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With
Medical Image Anal.9
2025 Investigating and Improving Latent Density Segmentation Models for Aleatoric Uncertainty Quantification in Medical Imaging
abstract
Data uncertainties, such as sensor noise, occlusions or limitations in the acquisition method can introduce irreducible ambiguities in images, which result in varying, yet plausible, semantic hypotheses. In Machine Learning, this ambiguity is commonly referred to as aleatoric uncertainty. In image segmentation, latent density models can be utilized to address this problem. The most popular approach is the Probabilistic U-Net (PU-Net), which uses latent Normal densities to optimize the conditional data log-likelihood Evidence Lower Bound. In this work, we demonstrate that the PU-Net latent space is severely sparse and heavily under-utilized. To address this, we introduce mutual information maximization and entropy-regularized Sinkhorn Divergence in the latent space to promote homogeneity across all latent dimensions, effectively improving gradient-descent updates and latent space informativeness. Our results show that by applying this on public datasets of various clinical segmentation problems, our proposed methodology receives up to 11% performance gains compared against preceding latent variable models for probabilistic segmentation on the Hungarian-Matched Intersection over Union. The results indicate that encouraging a homogeneous latent space significantly improves latent density modeling for medical image segmentation.
M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen
IEEE Trans. Medical Imaging5
2024 Retaining Informative Latent Variables in Probabilistic Segmentation
abstract
Conditional latent-variable models can successfully quantify annotation variability in segmentation. Training such models involves tuning the dimensionality of the latent space to optimally capture the inherent data ambiguity. Nevertheless, we discover after careful tuning, that the latent space does not always reflect this. In fact, some latent dimensions are completely neglected. For such segmentation models the latent dimensionality is often poorly motivated or based on computational constraints. In this paper, we offer an information-theoretic approach to optimally leverage all latent dimensions. We adapt and improve the Probabilistic U-Net to maximize the mutual information between the latent and output variables, leading to improved latent space properties and higher segmentation performance.
M. M. Amaan Valiuddin, Christiaan G. A. Viviers, Ruud van Sloun, Peter H. N. de With, Fons van der Sommen
ICASSP5
2024 IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting
abstract
Although action recognition for procedural tasks has received notable attention, it has a fundamental flaw in that no measure of success for actions is provided. This limits the applicability of such systems especially within the industrial domain, since the outcome of procedural actions is often significantly more important than the mere execution. To address this limitation, we define the novel task of procedure step recognition (PSR), focusing on recognizing the correct completion and order of procedural steps. Alongside the new task, we also present the multi-modal IndustReal dataset. Unlike currently available datasets, IndustReal contains procedural errors (such as omissions) as well as execution errors. A significant part of these errors are exclusively present in the validation and test sets, making IndustReal suitable to evaluate robustness of algorithms to new, unseen mistakes. Additionally, to encourage reproducibility and allow for scalable approaches trained on synthetic data, the 3D models of all parts are publicly available. Annotations and benchmark performance are provided for action recognition and assembly state detection, as well as the new PSR task. IndustReal, along with the code and model weights, is available at: https://github.com/TimSchoonbeek/IndustReal.
Tim J. Schoonbeek, Tim Houben, Hans Onvlee, Peter H. N. de With, Fons van der Sommen
WACV5
2024 Foundation models in gastrointestinal endoscopic AI: Impact of architecture, pre-training approach and data efficiency
abstract
Pre-training deep learning models with large data sets of natural images, such as ImageNet, has become the standard for endoscopic image analysis. This approach is generally superior to training from scratch, due to the scarcity of high-quality medical imagery and labels. However, it is still unknown whether the learned features on natural imagery provide an optimal starting point for the downstream medical endoscopic imaging tasks. Intuitively, pre-training with imagery closer to the target domain could lead to better-suited feature representations. This study evaluates whether leveraging in-domain pre-training in gastrointestinal endoscopic image analysis has potential benefits compared to pre-training on natural images. To this end, we present a dataset comprising of 5,014,174 gastrointestinal endoscopic images from eight different medical centers (GastroNet-5M), and exploit self-supervised learning with SimCLRv2, MoCov2 and DINO to learn relevant features for in-domain downstream tasks. The learned features are compared to features learned on natural images derived with multiple methods, and variable amounts of data and/or labels (e.g. Billion-scale semi-weakly supervised learning and supervised learning on ImageNet-21k). The effects of the evaluation is performed on five downstream data sets, particularly designed for a variety of gastrointestinal tasks, for example, GIANA for angiodyplsia detection and Kvasir-SEG for polyp segmentation. The findings indicate that self-supervised domain-specific pre-training, specifically using the DINO framework, results into better performing models compared to any supervised pre-training on natural images. On the ResNet50 and Vision-Transformer-small architectures, utilizing self-supervised in-domain pre-training with DINO leads to an average performance boost of 1.63% and 4.62%, respectively, on the downstream datasets. This improvement is measured against the best performance achieved through pre-training on natural images within any of the evaluated frameworks. Moreover, the in-domain pre-trained models also exhibit increased robustness against distortion perturbations (noise, contrast, blur, etc.), where the in-domain pre-trained ResNet50 and Vision-Transformer-small with DINO achieved on average 1.28% and 3.55% higher on the performance metrics, compared to the best performance found for pre-trained models on natural images. Overall, this study highlights the importance of in-domain pre-training for improving the generic nature, scalability and performance of deep learning for medical image analysis. The GastroNet-5M pre-trained weights are made publicly available in our repository: huggingface.co/tgwboers/GastroNet-5M_Pretrained_Weights.
T. G. W. Boers, Kiki Fockens, Joost van der Putten, Tim J. M. Jaspers, Carolus H. J. Kusters, Jelmer Jukema, Martijn R. Jong, Maarten R. Struyvenberg, Jeroen de Groof, Jacques J. Bergman, Peter H. N. de With, Fons van der Sommen
Medical Image Anal.12
2024 Robustness evaluation of deep neural networks for endoscopic image analysis: Insights and strategies
abstract
Computer-aided detection and diagnosis systems (CADe/CADx) in endoscopy are commonly trained using high-quality imagery, which is not representative for the heterogeneous input typically encountered in clinical practice. In endoscopy, the image quality heavily relies on both the skills and experience of the endoscopist and the specifications of the system used for screening. Factors such as poor illumination, motion blur, and specific post-processing settings can significantly alter the quality and general appearance of these images. This so-called domain gap between the data used for developing the system and the data it encounters after deployment, and the impact it has on the performance of deep neural networks (DNNs) supportive endoscopic CAD systems remains largely unexplored. As many of such systems, for e.g. polyp detection, are already being rolled out in clinical practice, this poses severe patient risks in particularly community hospitals, where both the imaging equipment and experience are subject to considerable variation. Therefore, this study aims to evaluate the impact of this domain gap on the clinical performance of CADe/CADx for various endoscopic applications. For this, we leverage two publicly available data sets (KVASIR-SEG and GIANA) and two in-house data sets. We investigate the performance of commonly-used DNN architectures under synthetic, clinically calibrated image degradations and on a prospectively collected dataset including 342 endoscopic images of lower subjective quality. Additionally, we assess the influence of DNN architecture and complexity, data augmentation, and pretraining techniques for improved robustness. The results reveal a considerable decline in performance of 11.6% (±1.5) as compared to the reference, within the clinically calibrated boundaries of image degradations. Nevertheless, employing more advanced DNN architectures and self-supervised in-domain pre-training effectively mitigate this drop to 7.7% (±2.03). Additionally, these enhancements yield the highest performance on the manually collected test set including images with lower subjective quality. By comprehensively assessing the robustness of popular DNN architectures and training strategies across multiple datasets, this study provides valuable insights into their performance and limitations for endoscopic applications. The findings highlight the importance of including robustness evaluation when developing DNNs for endoscopy applications and propose strategies to mitigate performance loss.
Tim J. M. Jaspers, T. G. W. Boers, Carolus H. J. Kusters, Martijn R. Jong, Jelmer Jukema, Albert Jeroen de Groof, Jacques J. Bergman, Peter H. N. de With, Fons van der Sommen
Medical Image Anal.9
2024 Advancing 6-DoF Instrument Pose Estimation in Variable X-Ray Imaging Geometries
abstract
Accurate 6-DoF pose estimation of surgical instruments during minimally invasive surgeries can substantially improve treatment strategies and eventual surgical outcome. Existing deep learning methods have achieved accurate results, but they require custom approaches for each object and laborious setup and training environments often stretching to extensive simulations, whilst lacking real-time computation. We propose a general-purpose approach of data acquisition for 6-DoF pose estimation tasks in X-ray systems, a novel and general purpose YOLOv5-6D pose architecture for accurate and fast object pose estimation and a complete method for surgical screw pose estimation under acquisition geometry consideration from a monocular cone-beam X-ray image. The proposed YOLOv5-6D pose model achieves competitive results on public benchmarks whilst being considerably faster at 42 FPS on GPU. In addition, the method generalizes across varying X-ray acquisition geometry and semantic image complexity to enable accurate pose estimation over different domains. Finally, the proposed approach is tested for bone-screw pose estimation for computer-aided guidance during spine surgeries. The model achieves a 92.41% by the 0.1·d ADD-S metric, demonstrating a promising approach for enhancing surgical precision and patient outcomes. The code for YOLOv5-6D is publicly available at https://github.com/cviviers/YOLOv5-6D-Pose.
Christiaan G. A. Viviers, Lena Filatova, Maurice Termeer, Peter H. N. de With, Fons van der Sommen
IEEE Trans. Image Process.5
2022 Block-Level Surrogate Models for Inference Time Estimation in Hardware-Aware Neural Architecture Search
Kurt Stolle, Sebastian Vogel, Fons van der Sommen, Willem P. Sanberg
ECML/PKDD (5)3
2022 Depth estimation from a single SEM image using pixel-wise fine-tuning with multimodal data
abstract
Abstract To support the ongoing size reduction in integrated circuits, the need for accurate depth measurements of on-chip structures becomes increasingly important. Unfortunately, present metrology tools do not offer a practical solution. In the semiconductor industry, critical dimension scanning electron microscopes (CD-SEMs) are predominantly used for 2D imaging at a local scale. The main objective of this work is to investigate whether sufficient 3D information is present in a single SEM image for accurate surface reconstruction of the device topology. In this work, we present a method that is able to produce depth maps from synthetic and experimental SEM images. We demonstrate that the proposed neural network architecture, together with a tailored training procedure, leads to accurate depth predictions. The training procedure includes a weakly supervised domain adaptation step, which is further referred to as pixel-wise fine-tuning. This step employs scatterometry data to address the ground-truth scarcity problem. We have tested this method first on a synthetic contact hole dataset, where a mean relative error smaller than 6.2% is achieved at realistic noise levels. Additionally, it is shown that this method is well suited for other important semiconductor metrics, such as top critical dimension (CD), bottom CD and sidewall angle. To the extent of our knowledge, we are the first to achieve accurate depth estimation results on real experimental data, by combining data from SEM and scatterometry measurements. An experiment on a dense line space dataset yields a mean relative error smaller than 1%.
Tim Houben, Thomas Huisman, Maxim Pisarenco, Fons van der Sommen, Peter H. N. de With
Mach. Vis. Appl.4
2022 Noise Reduction in CT Using Learned Wavelet-Frame Shrinkage Networks
abstract
Encoding-decoding (ED) CNNs have demonstrated state-of-the-art performance for noise reduction over the past years. This has triggered the pursuit of better understanding the inner workings of such architectures, which has led to the theory of deep convolutional framelets (TDCF), revealing important links between signal processing and CNNs. Specifically, the TDCF demonstrates that ReLU CNNs induce low-rankness, since these models often do not satisfy the necessary redundancy to achieve perfect reconstruction (PR). In contrast, this paper explores CNNs that do meet the PR conditions. We demonstrate that in these type of CNNs soft shrinkage and PR can be assumed. Furthermore, based on our explorations we propose the learned wavelet-frame shrinkage network, or LWFSN and its residual counterpart, the rLWFSN. The ED path of the (r)LWFSN complies with the PR conditions, while the shrinkage stage is based on the linear expansion of thresholds proposed Blu and Luisier. In addition, the LWFSN has only a fraction of the training parameters (<1%) of conventional CNNs, very small inference times, low memory footprint, while still achieving performance close to state-of-the-art alternatives, such as the tight frame (TF) U-Net and FBPConvNet, in low-dose CT denoising.
Luis Albert Zavala-Mondragón, Peter M. J. Rongen, Javier Oliván Bescós, Peter H. N. de With, Fons van der Sommen
IEEE Trans. Medical Imaging5
2021 Evaluating Self-Supervised Learning Methods for Downstream Classification of Neoplasia in Barrett's Esophagus
abstract
A major problem in applying machine learning for the medical domain is the scarcity of labeled data, which results in the demand for methods that enable high-quality models trained with little to no labels. Self-supervised learning methods present a plausible solution to this problem, enabling the use of large sets of unlabeled data for model pretraining. In this study, multiple of these methods and training strategies are employed on a large dataset of endoscopic images from the gastrointestinal tract (GastroNet). The suitability of these methods is assessed for an intra-domain downstream classification task on a small endoscopic dataset, involving neoplasia in Barrett’s esophagus. The classification performances are compared against pretraining on ImageNet and training from scratch. This yields promising results for domain-specific self-supervised methods, where super-resolution outperforms pretraining on ImageNet with a mean classification accuracy of 83.8% (cf. 79.2%). This implies that the large amounts of unlabeled data in hospitals could be employed in combination with self-supervised learning methods to improve models for downstream tasks.
Stefan Cornelissen, Joost van der Putten, T. G. W. Boers, Jelmer Jukema, Kiki Fockens, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With
ICIP7
2021 Automatic image and text-based description for colorectal polyps using BASIC classification
abstract
Colorectal polyps (CRP) are precursor lesions of colorectal cancer (CRC). Correct identification of CRPs during in-vivo colonoscopy is supported by the endoscopist's expertise and medical classification models. A recent developed classification model is the Blue light imaging Adenoma Serrated International Classification (BASIC) which describes the differences between non-neoplastic and neoplastic lesions acquired with blue light imaging (BLI). Computer-aided detection (CADe) and diagnosis (CADx) systems are efficient at visually assisting with medical decisions but fall short at translating decisions into relevant clinical information. The communication between machine and medical expert is of crucial importance to improve diagnosis of CRP during in-vivo procedures. In this work, the combination of a polyp image classification model and a language model is proposed to develop a CADx system that automatically generates text comparable to the human language employed by endoscopists. The developed system generates equivalent sentences as the human-reference and describes CRP images acquired with white light (WL), blue light imaging (BLI) and linked color imaging (LCI). An image feature encoder and a BERT module are employed to build the AI model and an external test set is used to evaluate the results and compute the linguistic metrics. The experimental results show the construction of complete sentences with an established metric scores of BLEU-1 = 0.67, ROUGE-L = 0.83 and METEOR = 0.50. The developed CADx system for automatic CRP image captioning facilitates future advances towards automatic reporting and may help reduce time-consuming histology assessment.
Roger Fonolla, Quirine E. W. van der Zander, Ramon-Michel Schreuder, Sharmila Subramaniam, Pradeep Bhandari, Ad A. M. Masclee, Erik J. Schoon, Fons van der Sommen, Peter H. N. de With
Artif. Intell. Medicine8
2021 Image Noise Reduction Based on a Fixed Wavelet Frame and CNNs Applied to CT
abstract
Radiation exposure in CT imaging leads to increased patient risk. This motivates the pursuit of reduced-dose scanning protocols, in which noise reduction processing is indispensable to warrant clinically acceptable image quality. Convolutional Neural Networks (CNNs) have received significant attention as an alternative for conventional noise reduction and are able to achieve state-of-the art results. However, the internal signal processing in such networks is often unknown, leading to sub-optimal network architectures. The need for better signal preservation and more transparency motivates the use of Wavelet Shrinkage Networks (WSNs), in which the Encoding-Decoding (ED) path is the fixed wavelet frame known as Overcomplete Haar Wavelet Transform (OHWT) and the noise reduction stage is data-driven. In this work, we considerably extend the WSN framework by focusing on three main improvements. First, we simplify the computation of the OHWT that can be easily reproduced. Second, we update the architecture of the shrinkage stage by further incorporating knowledge of conventional wavelet shrinkage methods. Finally, we extensively test its performance and generalization, by comparing it with the RED and FBPConvNet CNNs. Our results show that the proposed architecture achieves similar performance to the reference in terms of MSSIM (0.667, 0.662 and 0.657 for DHSN2, FBPConvNet and RED, respectively) and achieves excellent quality when visualizing patches of clinically important structures. Furthermore, we demonstrate the enhanced generalization and further advantages of the signal flow, by showing two additional potential applications, in which the new DHSN2 is used as regularizer: (1) iterative reconstruction and (2) ground-truth free training of the proposed noise reduction architecture. The presented results prove that the tight integration of signal processing and deep learning leads to simpler models with improved generalization.
Luis Albert Zavala-Mondragón, Peter H. N. de With, Fons van der Sommen
IEEE Trans. Image Process.3
2020 Multi-stage domain-specific pretraining for improved detection and localization of Barrett's neoplasia: A comprehensive clinically validated study
abstract
Patients suffering from Barrett's Esophagus (BE) are at an increased risk of developing esophageal adenocarcinoma and early detection is crucial for a good prognosis. To aid the endoscopists with the early detection for this preliminary stage of esophageal cancer, this work concentrates on the development and extensive evaluation of a state-of-the-art computer-aided classification and localization algorithm for dysplastic lesions in BE. To this end, we have employed a large-scale endoscopic data set, consisting of 494,355 images, in combination with a novel semi-supervised learning algorithm to pretrain several instances of the proposed neural network architecture. Next, several Barrett-specific data sets that are increasingly closer to the target domain with significantly more data compared to other related work, were used in a multi-stage transfer learning strategy. Additionally, the algorithm was evaluated on two prospectively gathered external test sets and compared against 53 medical professionals. Finally, the model was also evaluated in a live setting without interfering with the current biopsy protocol. Results from the performed experiments show that the proposed model improves on the state-of-the-art on all measured metrics. More specifically, compared to the best performing state-of-the-art model, the specificity is improved by more than 20% points while simultaneously preserving high sensitivity and reducing the false positive rate substantially. Our algorithm yields similar scores on the localization metrics, where the intersection of all experts is correctly indicated in approximately 92% of the cases. Furthermore, the live pilot study shows great performance in a clinical setting with a patient level accuracy, sensitivity, and specificity of 90%. Finally, the proposed algorithm outperforms each individual medical expert by at least 5% and the average assessor by more than 10% over all assessor groups with respect to accuracy.
Joost van der Putten, Jeroen de Groof, Maarten R. Struyvenberg, T. G. W. Boers, Kiki Fockens, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Fons van der Sommen, Peter H. N. de With
Artif. Intell. Medicine9
2020 Modeling clinical assessor intervariability using deep hypersphere encoder-decoder networks
abstract
Abstract In medical imaging, a proper gold-standard ground truth as, e.g., annotated segmentations by assessors or experts is lacking or only scarcely available and suffers from large intervariability in those segmentations. Most state-of-the-art segmentation models do not take inter-observer variability into account and are fully deterministic in nature. In this work, we propose hypersphere encoder–decoder networks in combination with dynamic leaky ReLUs, as a new method to explicitly incorporate inter-observer variability into a segmentation model. With this model, we can then generate multiple proposals based on the inter-observer agreement. As a result, the output segmentations of the proposed model can be tuned to typical margins inherent to the ambiguity in the data. For experimental validation, we provide a proof of concept on a toy data set as well as show improved segmentation results on two medical data sets. The proposed method has several advantages over current state-of-the-art segmentation models such as interpretability in the uncertainty of segmentation borders. Experiments with a medical localization problem show that it offers improved biopsy localizations, which are on average 12% closer to the optimal biopsy location.
Joost van der Putten, Fons van der Sommen, Jeroen de Groof, Maarten R. Struyvenberg, Svitlana Zinger, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Peter H. N. de With
Neural Comput. Appl.2
2019 Image Features for Automated Colorectal Polyp Classification Based on Clinical Prediction Models
abstract
Accurate endoscopic differentiation on resection of colorectal polyps (CRPs) (resect-discard or diagnose-leave strategies) increases cost-efficiency and reduces patient risk. We aim to develop a classification algorithm for automated differentiation of CRPs, by following the validated clinical Work-group serrAted polypS and Polyposis (WASP) classification scheme. Quantitative image features are investigated for each individual WASP criterion and classification is performed by conventional SVM. The technical WASP model results in areas under the curve of 0.87-0.95 and accuracies of 78-89%. Predicting polyp histology using model-based learning out-performs medical experts (accuracy, 87-93% vs 86 87%). Direct classification predicts more premalignant polyps-as being benign, compared to the automated WASP scheme. These errors do not occur when including ROC characteristics to the WASP model. The proposed WASP model is the first automated system, competing with medical expert classification.
Michelle C. A. van Grinsven, Thom Scheeve, Ramon-Michel Schreuder, Fons van der Sommen, Erik J. Schoon, Peter H. N. de With
ICIP4
2019 Informative Frame Classification of Endoscopic Videos Using Convolutional Neural Networks and Hidden Markov Models
abstract
The goal of endoscopic analysis is to find abnormal lesions and determine further therapy from the obtained information. For example, in case of Barrett's esophagus, the objective of endoscopy is to timely detect dysplastic lesions, before endoscopic resection is no longer possible. However, the procedure produces a variety of non-informative frames and lesions can be missed due to poor video quality. Especially when analyzing entire endoscopic videos made by non-expert endoscopists, informative frame classification is crucial to e.g. video quality grading. This analysis involves classification problems such as polyp detection or dysplasia detection in Barrett's Esophagus. This work concentrates on the design of an automated indication of informativeness of video frames. We propose an algorithm consisting of state-of-the-art deep learning techniques, to initialize frame-based classification, followed by a hidden Markov model to incorporate temporal information and control consistent decision making. Results from the performed experiments show that the proposed model improves on the state-of-the-art with an F1-score of 91%, and a substantial increase in sensitivity of 10%, thereby indicating improved labeling consistency. Additionally, the algorithm is capable of processing 261 frames per second, which is multiple times faster compared to other informative frame classification algorithms, thus enabling real-time computation.
Joost van der Putten, Jeroen de Groof, Fons van der Sommen, Maarten R. Struyvenberg, Svitlana Zinger, Wouter L. Curvers, Erik J. Schoon, Jacques J. Bergman, Peter H. N. de With
ICIP3
2018 Automatic Detection of Early Esophageal Cancer with CNNS Using Transfer Learning
abstract
The incidence of Esophageal Adenocarcinoma (EAC), a form of esophageal cancer, has rapidly increased in recent years. Dysplastic tissue can be removed endoscopically at an early stage, and since survival chances of patients are limited at later stages of the disease, early detection is of key impor- tance. Recently, several CAD systems for HD endoscopic images have been proposed, but these are computationally expensive, making them unfit for clinical use requiring real- time analysis. In this paper, we present a novel approach for early esophageal cancer detection using Transfer Learning with CNNs. Given the small amount of annotated data, CNN Codes are applied, where intermediate layers of the net- work are used as features for conventional classifiers. Various classifiers are combined with four of the most widely-used networks. Additionally, sliding windows are used to obtain a coarse-grained annotation indicating any possible cancerous regions. This approach outperforms the current state-of-the-art with a frame-based AUC of 0.92, while allowing both near real-time prediction and annotation at 2 fps, in a MATLAB-based framework.
Sjors van Riel, Fons van der Sommen, Svitlana Zinger, Erik J. Schoon, Peter H. N. de With
ICIP2
2018 How to Exploit Weaknesses in Biomedical Challenge Design and Organization
Annika Reinke, Matthias Eisenmann, Sinan Onogur, Marko Stankovic 0002, Patrick Godau, Peter M. Full, Hrvoje Bogunovic, Bennett A. Landman, Oskar Maier, Bjoern Menze, Gregory C. Sharp, Korsuk Sirinukunwattana, Stefanie Speidel, Fons van der Sommen, Guoyan Zheng, Henning Müller, Michal Kozubek 0001, Tal Arbel, Andrew P. Bradley, Pierre Jannin, Annette Kopp-Schneider, Lena Maier-Hein
MICCAI (4)14
2017 Improved Barrett's Cancer Detection in Volumetric Laser Endomicroscopy Scans Using Multiple-Frame Voting
abstract
This paper explores the feasibility of using multiframe analysis to increase the classification performance of machine learning methods for cancer detection in Volumetric Laser Endomicroscopy (VLE). VLE is a novel and promising modality for the detection of neoplasia in patients with Baretts Esophagus (BE). It produces hundreds of high-resolution, cross-sectional images of the esophagus and offers considerable advantages compared to current methods. While some recent studies have proposed cancer detection algorithms for single VLE frames, the study described in this paper is the first to make use of VLE volumes for the differentiation between dysplastic and non-dysplastic tissue. We explore the use of various voting schemes for a broad range of features and classification methods. Our results demonstrate that multi-frame analysis leads to superior performance, irrespective of the chosen feature-classifier combination. By using multi-frame analysis with straightforward voting methods, the Area Under the receiver operating Curve (AUC) is increased by an average of over 12% compared to using single VLE frames. When only considering methods that achieve expert performance or higher (AUC≥0.81), an even larger performance improvement of up to 16.9% is observed. Furthermore, with many feature/classifier combinations showing AUC values ranging from 0.90 to 0.98, our experiments indicate that computeraided methods can considerably outperform medical experts, who demonstrate an AUC of 0.81 using a recently proposed clinical prediction model.
Alexandros Rikos, Fons van der Sommen, Anne-Fré Swager, Svitlana Zinger, Erik J. Schoon, Wouter L. Curvers, Jacques J. Bergman, Peter H. N. de With
CBMS2
2015 Real-time semantic context labeling for image understanding
abstract
The use of context information in a scene is an important aid for full semantic scene understanding in security and surveillance applications. To this end, this paper presents an innovative semantic context-labeling algorithm for three context classes, trading-off quality and real-time execution. Our system consists of three consecutive stages: image segmentation, region-based feature extraction and classification. We propose the joint use of the features color in HSV space, texture from Gabor filters and spatial context, in combination with the Directional Nearest Neighbor (DNN) method for constructing the undirected graph for segmentation. Compared to recent literature, this combination is over 35 times faster and achieves a coverability rate that is 65% higher.
Martin A. R. Pieck, Fons van der Sommen, Svitlana Zinger, Peter H. N. de With
ICIP2
2014 Supportive automatic annotation of early esophageal cancer using local gabor and color features
Fons van der Sommen, Svitlana Zinger, Erik J. Schoon, Peter H. N. de With
Neurocomputing1