EDBT 2026 Demo / reviewers in the wild / expert
Tom Vercauteren
dblp:99/4387
· DBLP profile ↗
106ranked-venue papers
8as first author
39since 2021 · last 2026
0000-0003-1794-0456ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 84 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 7 since 2021Systems, architecture and hardware · 7 · 1 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionabstractEnabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexplored in surgical videos, with existing approaches struggling to generalize due to reliance on static visual cues and predefined instrument names. In this work, we introduce SurgRef, a novel motion-guided framework that grounds free-form language expressions in instrument motion, capturing how tools move and interact across time, rather than what they look like. This allows models to understand and segment instruments even under occlusion, ambiguity, or unfamiliar terminology. To train and evaluate SurgRef, we present Ref-IMotion, a diverse, multi-institutional video dataset with dense spatiotemporal masks and rich motion-centric expressions. SurgRef achieves state-of-the-art accuracy and generalization across surgical procedures, setting a new benchmark for robust, language-driven surgical video segmentation. Kun Yuan 0004, Long Bai 0008, Nassir Navab, Hongliang Ren 0001, Hong Joo Lee 0001, Tom Vercauteren, Nicolas Padoy |
AAAI | 9 |
| 2026 | A Framework for Real-Time Surgical Phase Recognition with Application to Robot-Assisted Partial NephrectomyabstractSurgical practice has increasingly integrated advanced technologies to improve procedural outcomes, efficiency, and safety in modern operating rooms. Within this evolving landscape, Automated Surgical Phase Recognition (SPR) leverages Artificial Intelligence to temporally segment surgical workflows into key events, thereby supporting both real-time decision-making and off-line analysis. Despite the potential of SPR, previous research focused on short and linear surgeries, paying limited attention to the development, assessment, and deployment of real-time systems for complex surgical workflows. This work addresses these gaps by targeting the highly-complex and non linear workflow of Robot-Assisted Partial Nephrectomy (RAPN). We develop a real-time SPR system trained on 143 annotated RAPN surgical videos spanning 15 phases. The system incorporates a trainable canonical calibration error estimator combined with Viterbi decoding for more reliable outcomes. Additionally, we introduce a novel assessment framework designed to simultaneously evaluate offline, real-time, and averaged SPR performance, synthesising historical phase predictions over time. For deployment, we implement the SPR pipeline as an end-to-end application using the NVIDIA Holoscan platform. The system was successfully tested during three live RAPN cases on human patients in a collaborating hospital, achieving an average inference latency of 16.65 ms and an accuracy of 68.2%. Results indicate that Viterbi decoding boosts performance in this complex surgery, while canonical calibration does not significantly increase overall performance but enhances classification reliability. We show the feasibility of deploying a real-time SPR pipeline for RAPN, which holds promise for optimising OR planning. The application is available at https://github.com/nvidia-holoscan/holohub/tree/main/applications/orsi Marco Mezzina, Tom Vercauteren, Tinne Tuytelaars, Matthew B. Blaschko |
WACV | 2 |
| 2026 | Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challengeabstractReliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding. Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm |
Medical Image Anal. | 32 |
| 2026 | OOD-SEG: Exploiting out-of-distribution detection techniques for learning image segmentation from sparse multi-class positive-only annotationsabstract• We propose a segmentation framework for learning from sparse positive-only multi-class annotations without background labels. • We model implicit background via pixel-wise OOD detection, enabling separation of in-distribution data from negative/unknown regions. • We propose a two-level cross-validation strategy that addressing the scarcity of OOD datasets and the lack of established evaluation metrics in segmentation. • We introduce a novel convolutional adaptation of an established OOD detection method originally designed for classification purposes. • We demonstrate robustness and generalisation through experiments on multi-class hyperspectral and RGB surgical imaging datasets. Despite significant advancements, segmentation based on deep neural networks in medical and surgical imaging faces several challenges, two of which we aim to address in this work. First, acquiring complete pixel-level segmentation labels for medical images is time-consuming and requires domain expertise. Second, typical segmentation pipelines cannot detect out-of-distribution (OOD) pixels, leaving them prone to spurious outputs during deployment. In this work, we propose a novel segmentation approach which broadly falls within the positive-unlabelled (PU) learning paradigm and exploits tools from OOD detection techniques. Our framework learns only from sparsely annotated pixels from multiple positive-only classes and does not use any annotation for the background class. These multi-class positive annotations naturally fall within the in-distribution (ID) set. Unlabelled pixels may contain positive classes but also negative ones, including what is typically referred to as background in standard segmentation formulations. To the best of our knowledge, this work is the first to formulate multi-class segmentation with sparse positive-only annotations as a pixel-wise PU learning problem and to address it using OOD detection techniques. Here, we forgo the need for background annotation and consider these together with any other unseen classes as part of the OOD set. Our framework can integrate, at a pixel-level, any OOD detection approaches designed for classification tasks. To address the lack of existing OOD datasets and established evaluation metric for medical image segmentation, we propose a cross-validation strategy that treats held-out labelled classes as OOD. Extensive experiments on both multi-class hyperspectral and RGB surgical imaging datasets demonstrate the robustness and generalisation capability of our proposed framework. Junwen Wang, Oscar MacCormac, Jonathan Shapey, Tom Vercauteren |
Medical Image Anal. | 5 |
| 2026 | Average Calibration Losses for Reliable Uncertainty in Medical Image SegmentationabstractDeep neural networks for medical image segmentation are often overconfident, compromising both reliability and clinical utility. In this work, we propose differentiable formulations of marginal L1 Average Calibration Error (mL1-ACE) as an auxiliary loss that can be computed on a per-image basis. We compare both hard-and soft-binning approaches to directly improve pixel-wise calibration. Our experiments on four datasets (ACDC, AMOS, KiTS, BraTS) demonstrate that incorporating mL1-ACE significantly reduces calibration errors, particularly Average Calibration Error (ACE) and Maximum Calibration Error (MCE), while largely maintaining high Dice Similarity Coefficients (DSCs). We find that the soft-binned variant yields the greatest improvements in calibration, over the DSC plus cross-entropy loss baseline, but often compromises segmentation performance, with hard-binned mL1-ACE maintaining segmentation performance, albeit with weaker calibration improvement. To gain further insight into calibration performance and its variability across an imaging dataset, we introduce dataset reliability histograms, an aggregation of per-image reliability diagrams. The resulting analysis highlights improved alignment between predicted confidences and true accuracies. Overall, our approach provides practitioners with explicit control over the calibration-accuracy trade-off, enabling more reliable integration of deep learning methods into clinical workflows. We share our code here: https://github.com/ cai4cai/Average-Calibration-Losses. Theodore Barfoot, Luis C. García-Peraza-Herrera, Samet Akcay, Ben Glocker, Tom Vercauteren |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Tree-Based Semantic Losses: Application to Sparsely-Supervised Large Multi-class Hyperspectral Segmentation
Junwen Wang, Oscar MacCormac, William Rochford, Aaron Kujawa, Jonathan Shapey, Tom Vercauteren |
MICCAI (8) | 6 |
| 2025 | Multitask learning in minimally invasive surgical vision: A reviewabstractMinimally invasive surgery (MIS) has revolutionized many procedures and led to reduced recovery time and risk of patient injury. However, MIS poses additional complexity and burden on surgical teams. Data-driven surgical vision algorithms are thought to be key building blocks in the development of future MIS systems with improved autonomy. Recent advancements in machine learning and computer vision have led to successful applications in analysing videos obtained from MIS with the promise of alleviating challenges in MIS videos. Surgical scene and action understanding encompasses multiple related tasks that, when solved individually, can be memory-intensive, inefficient, and fail to capture task relationships. Multitask learning (MTL), a learning paradigm that leverages information from multiple related tasks to improve performance and aid generalization, is well-suited for fine-grained and high-level understanding of MIS data. This review provides a narrative overview of the current state-of-the-art MTL systems that leverage videos obtained from MIS. Beyond listing published approaches, we discuss the benefits and limitations of these MTL systems. Moreover, this manuscript presents an analysis of the literature for various application fields of MTL in MIS, including those with large models, highlighting notable trends, new directions of research, and developments. Oluwatosin Alabi, Tom Vercauteren, Miaojing Shi |
Medical Image Anal. | 2 |
| 2025 | LoViT: Long Video Transformer for surgical phase recognitionabstractOnline surgical phase recognition plays a significant role towards building contextual tools that could quantify performance and oversee the execution of surgical workflows. Current approaches are limited since they train spatial feature extractors using frame-level supervision that could lead to incorrect predictions due to similar frames appearing at different phases, and poorly fuse local and global features due to computational constraints which can affect the analysis of long videos commonly encountered in surgical interventions. In this paper, we present a two-stage method, called Long Video Transformer (LoViT), emphasizing the development of a temporally-rich spatial feature extractor and a phase transition map. The temporally-rich spatial feature extractor is designed to capture critical temporal information within the surgical video frames. The phase transition map provides essential insights into the dynamic transitions between different surgical phases. LoViT combines these innovations with a multiscale temporal aggregator consisting of two cascaded L-Trans modules based on self-attention, followed by a G-Informer module based on ProbSparse self-attention for processing global temporal information. The multi-scale temporal head then leverages the temporally-rich spatial features and phase transition map to classify surgical phases using phase transition-aware supervision. Our approach outperforms state-of-the-art methods on the Cholec80 and AutoLaparo datasets consistently. Compared to Trans-SVNet, LoViT achieves a 2.4 pp (percentage point) improvement in video-level accuracy on Cholec80 and a 3.1 pp improvement on AutoLaparo. Our results demonstrate the effectiveness of our approach in achieving state-of-the-art performance of surgical phase recognition on two datasets of different surgical procedures and temporal sequencing characteristics. The project page is available at https://github.com/MRUIL/LoViT. Yang Liu 0271, Maxence Boels, Luis C. García-Peraza-Herrera, Tom Vercauteren, Prokar Dasgupta, Alejandro Granados, Sébastien Ourselin |
Medical Image Anal. | 4 |
| 2025 | UM-CAM: Uncertainty-weighted multi-resolution class activation maps for weakly-supervised segmentationabstractWeakly-supervised medical image segmentation methods utilizing image-level labels have gained attention for reducing the annotation cost. They typically use Class Activation Maps (CAM) from a classification network but struggle with incomplete activation regions due to low-resolution localization without detailed boundaries. Differently from most of them that only focus on improving the quality of CAMs, we propose a more unified weakly-supervised segmentation framework with image-level supervision. Firstly, an Uncertainty-weighted Multi-resolution Class Activation Map (UM-CAM) is proposed to generate high-quality pixel-level pseudo-labels. Subsequently, a Geodesic distance-based Seed Expansion (GSE) strategy is introduced to rectify ambiguous boundaries in the UM-CAM by leveraging contextual information. To train a final segmentation model from noisy pseudo-labels, we introduce a Random-View Consensus (RVC) training strategy to suppress unreliable pixel/voxels and encourage consistency between random-view predictions. Extensive experiments on 2D fetal brain segmentation and 3D brain tumor segmentation tasks showed that our method significantly outperforms existing weakly-supervised methods. Code is available at: https://github.com/HiLab-git/UM-CAM . Guotai Wang, Qiang Yue 0005, Tom Vercauteren, Sébastien Ourselin, Shaoting Zhang 0001 |
Pattern Recognit. | 5 |
| 2024 | A self-supervised and adversarial approach to hyperspectral demosaicking and RGB reconstruction in surgical imaging
Peichao Li, Oscar MacCormac, Jonathan Shapey, Tom Vercauteren |
BMVC | 4 |
| 2024 | Average Calibration Error: A Differentiable Loss for Improved Reliability in Image SegmentationabstractDeep neural networks for medical image segmentation often produce overconfident results misaligned with empirical observations. Such miscalibration challenges their clinical translation. We propose to use marginal L1 average calibration error (mL1-ACE) as a novel auxiliary loss function to improve pixel-wise calibration without compromising segmentation quality. We show that this loss, despite using hard binning, is directly differentiable, bypassing the need for approximate but differentiable surrogate or soft binning approaches. Our work also introduces the concept of dataset reliability histograms which generalises standard reliability diagrams for refined visual assessment of calibration in semantic segmentation aggregated at the dataset level. Using mL1-ACE, we reduce average and maximum calibration error by 45% and 55% respectively, maintaining a Dice score of 87% on the BraTS 2021 dataset. We share our code here: https://github.com/cai4cai/ACE-DLIRIS . Theodore Barfoot, Luis C. García-Peraza-Herrera, Ben Glocker, Tom Vercauteren |
MICCAI (9) | 4 |
| 2024 | Transferring Relative Monocular Depth to Surgical Vision with Temporal ConsistencyabstractRelative monocular depth, inferring depth correct up to a shift and scale from a single image, is an active research topic. Recent deep learning models, trained on large and varied meta-datasets, now provide excellent performance in the domain of natural images. However, few datasets exist which provide ground truth depth for endoscopic images, making training such models from scratch unfeasible. This work investigates the transfer of these models into the surgical domain, and presents an effective and simple way to improve on standard supervision through the use of temporal consistency self-supervision. We show temporal consistency significantly improves supervised training alone when transferring to the low-data regime of endoscopy, and outperforms the prevalent self-supervision technique for this task. In addition we show our method drastically outperforms the state-of-the-art method from within the domain of endoscopy. We also release our code, models, and ensembled meta-dataset, Meta-MED, establishing a strong benchmark for future work. Charlie Budd, Tom Vercauteren |
MICCAI (6) | 2 |
| 2024 | Label Merge-and-Split: A Graph-Colouring Approach for Memory-Efficient Brain Parcellation
Aaron Kujawa, Reuben Dorent, Sébastien Ourselin, Tom Vercauteren |
MICCAI (9) | 4 |
| 2024 | Nonrigid Reconstruction of Freehand Ultrasound Without a TrackerabstractReconstructing 2D freehand Ultrasound (US) frames into 3D space without using a tracker has recently seen advances with deep learning. Predicting good frame-to-frame rigid transformations is often accepted as the learning objective, especially when the ground-truth labels from spatial tracking devices are inherently rigid transformations. Motivated by a) the observed nonrigid deformation due to soft tissue motion during scanning, and b) the highly sensitive prediction of rigid transformation, this study investigates the methods and their benefits in predicting nonrigid transformations for reconstructing 3D US. We propose a novel co-optimisation algorithm for simultaneously estimating rigid transformations among US frames, supervised by ground-truth from a tracker, and a nonrigid deformation, optimised by a regularised registration network. We show that these two objectives can be either optimised using meta-learning or combined by weighting. A fast scattered data interpolation is also developed for enabling frequent reconstruction and registration of non-parallel US frames, during training. With a new data set containing over 357,000 frames in 720 scans, acquired from 60 subjects, the experiments demonstrate that, due to an expanded thus easier-to-optimise solution space, the generalisation is improved with the added deformation estimation, with respect to the rigid ground-truth. The global pixel reconstruction error (assessing accumulative prediction) is lowered from 18.48 to 16.51 mm, compared with baseline rigid-transformation-predicting methods. Using manually identified landmarks, the proposed co-optimisation also shows potentials in compensating nonrigid tissue motion at inference, which is not measurable by tracker-provided ground-truth. The code and data used in this paper are made publicly available at https://github.com/QiLi111/NR-Rec-FUS . Qi Li 0030, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, Yipeng Hu |
MICCAI (4) | 6 |
| 2024 | MONAI Label: A framework for AI-assisted interactive labeling of 3D medical images
Andres Diaz-Pinto, Sachidanand Alle, Vishwesh Nath, Yucheng Tang, Alvin Ihsani, Muhammad Asad 0001, Fernando Pérez-García, Pritesh Mehta, Wenqi Li 0001, Mona Flores, Holger Roth, Tom Vercauteren, Daguang Xu, Prerna Dogra, Sébastien Ourselin, Andrew Feng, Manuel Jorge Cardoso |
Medical Image Anal. | 12 |
| 2024 | Generating multi-pathological and multi-modal images and labels for brain MRIabstractThe last few years have seen a boom in using generative models to augment real datasets, as synthetic data can effectively model real data distributions and provide privacy-preserving, shareable datasets that can be used to train deep learning models. However, most of these methods are 2D and provide synthetic datasets that come, at most, with categorical annotations. The generation of paired images and segmentation samples that can be used in downstream, supervised segmentation tasks remains fairly uncharted territory. This work proposes a two-stage generative model capable of producing 2D and 3D semantic label maps and corresponding multi-modal images. We use a latent diffusion model for label synthesis and a VAE-GAN for semantic image synthesis. Synthetic datasets provided by this model are shown to work in a wide variety of segmentation tasks, supporting small, real datasets or fully replacing them while maintaining good performance. We also demonstrate its ability to improve downstream performance on out-of-distribution data. Virginia Fernandez, Walter H. L. Pinaya, Pedro Borges, Mark S. Graham, Petru-Daniel Tudosiu, Tom Vercauteren, Manuel Jorge Cardoso |
Medical Image Anal. | 6 |
| 2024 | A Dempster-Shafer Approach to Trustworthy AI With Application to Fetal Brain MRI SegmentationabstractDeep learning models for medical image segmentation can fail unexpectedly and spectacularly for pathological cases and images acquired at different centers than training images, with labeling errors that violate expert knowledge. Such errors undermine the trustworthiness of deep learning models for medical image segmentation. Mechanisms for detecting and correcting such failures are essential for safely translating this technology into clinics and are likely to be a requirement of future regulations on artificial intelligence (AI). In this work, we propose a trustworthy AI theoretical framework and a practical system that can augment any backbone AI system using a fallback method and a fail-safe mechanism based on Dempster-Shafer theory. Our approach relies on an actionable definition of trustworthy AI. Our method automatically discards the voxel-level labeling predicted by the backbone AI that violate expert knowledge and relies on a fallback for those voxels. We demonstrate the effectiveness of the proposed trustworthy AI approach on the largest reported annotated dataset of fetal MRI consisting of 540 manually annotated fetal brain 3D T2w MRIs from 13 centers. Our trustworthy AI method improves the robustness of four backbone AI models for fetal brain MRIs acquired across various centers and for fetuses with various brain abnormalities. Lucas Fidon, Michael Aertsen, Florian Kofler, Andrea Bink, Anna L. David, Thomas Deprest, Doaa Emam, Frédéric Guffens, András Jakab, Gregor Kasprian, Patric Kienast, Andrew Melbourne, Bjoern Menze, Nada Mufti, Ivana Pogledic, Daniela Prayer, Marlene Stuempflen, Esther Van Elslander, Sébastien Ourselin, Jan Deprest, Tom Vercauteren |
IEEE Trans. Pattern Anal. Mach. Intell. | 21 |
| 2023 | OpTaS: An Optimization-based Task Specification Library for Trajectory Optimization and Model Predictive ControlabstractThis paper presents OpTaS, a task specification Python library for Trajectory Optimization (TO) and Model Predictive Control (MPC) in robotics. Both TO and MPC are increasingly receiving interest in optimal control and in particular handling dynamic environments. While a flurry of software libraries exists to handle such problems, they either provide interfaces that are limited to a specific problem formulation (e.g. TracIK, CHOMP), or are large and statically specify the problem in configuration files (e.g. EXOTica, eTaSL). OpTaS, on the other hand, allows a user to specify custom nonlinear constrained problem formulations in a single Python script allowing the controller parameters to be modified during execution. The library provides interface to several open source and commercial solvers (e.g. IPOPT, SNOPT, KNITRO, SciPy) to facilitate integration with established workflows in robotics. Further benefits of OpTaS are highlighted through a thorough comparison with common libraries. An additional key advantage of OpTaS is the ability to define optimal control tasks in the joint-space, task-space, or indeed simultaneously. The code for OpTaS is easily installed via pip, and the source code with examples can be found at github.com/cmower/optas. Christopher E. Mower, João Moura 0003, Nazanin Zamani Behabadi, Sethu Vijayakumar, Tom Vercauteren, Christos Bergeles |
ICRA | 5 |
| 2023 | Adaptive Multi-scale Online Likelihood Network for AI-Assisted Interactive Segmentation
Muhammad Asad 0001, Helena Williams, Indrajeet Mandal, Sarim Ather, Jan Deprest, Jan D'hooge, Tom Vercauteren |
MICCAI (2) | 7 |
| 2023 | Deep Reinforcement Learning Based System for Intraoperative Hyperspectral Video AutofocusingabstractHyperspectral imaging (HSI) captures a greater level of spectral detail than traditional optical imaging, making it a potentially valuable intraoperative tool when precise tissue differentiation is essential. Hardware limitations of current optical systems used for handheld real-time video HSI result in a limited focal depth, thereby posing usability issues for integration of the technology into the operating room. This work integrates a focus-tunable liquid lens into a video HSI exoscope, and proposes novel video autofocusing methods based on deep reinforcement learning. A first-of-its-kind robotic focal-time scan was performed to create a realistic and reproducible testing dataset. We benchmarked our proposed autofocus algorithm against traditional policies, and found our novel approach to perform significantly ( $$p<0.05$$ ) better than traditional techniques ( $$0.070\pm .098$$ mean absolute focal error compared to $$0.146\pm .148$$ ). In addition, we performed a blinded usability trial by having two neurosurgeons compare the system with different autofocus policies, and found our novel approach to be the most favourable, making our system a desirable addition for intraoperative HSI. Charlie Budd, Jianrong Qiu, Oscar MacCormac, Christopher E. Mower, Mirek Janatka, Théo Trotouin, Jonathan Shapey, Mads S. Bergholt, Tom Vercauteren |
MICCAI (9) | 10 |
| 2023 | Unified Brain MR-Ultrasound Synthesis Using Multi-modal Hierarchical RepresentationsabstractWe introduce MHVAE, a deep hierarchical variational autoencoder (VAE) that synthesizes missing images from various modalities. Extending multi-modal VAEs with a hierarchical latent structure, we introduce a probabilistic formulation for fusing multi-modal images in a common latent representation while having the flexibility to handle incomplete image sets as input. Moreover, adversarial learning is employed to generate sharper images. Extensive experiments are performed on the challenging problem of joint intra-operative ultrasound (iUS) and Magnetic Resonance (MR) synthesis. Our model outperformed multi-modal VAEs, conditional GANs, and the current state-of-the-art unified method (ResViT) for synthesizing missing images, demonstrating the advantage of using a hierarchical latent representation and a principled probabilistic fusion operation. Our code is publicly available. Reuben Dorent, Nazim Haouchine, Fryderyk Victor Kögl, Samuel Joutard, Parikshit Juvekar, Erickson Torio, Alexandra J. Golby, Sébastien Ourselin, Sarah F. Frisken, Tom Vercauteren, Tina Kapur, William M. Wells III |
MICCAI (10) | 10 |
| 2023 | Deep Homography Prediction for Endoscopic Camera Motion Imitation Learning
Sébastien Ourselin, Christos Bergeles, Tom Vercauteren |
MICCAI (9) | 4 |
| 2023 | Text Promptable Surgical Instrument Segmentation with Vision-Language ModelsabstractIn this paper, we propose a novel text promptable surgical instrument segmentation approach to overcome challenges associated with diversity and differentiation of surgical instruments in minimally invasive surgeries. We redefine the task as text promptable, thereby enabling a more nuanced comprehension of surgical instruments and adaptability to new instrument types. Inspired by recent advancements in vision-language models, we leverage pretrained image and text encoders as our model backbone and design a text promptable mask decoder consisting of attention- and convolution-based prompting schemes for surgical instrument segmentation prediction. Our model leverages multiple text prompts for each surgical instrument through a new mixture of prompts mechanism, resulting in enhanced segmentation performance. Additionally, we introduce a hard instrument area reinforcement module to improve image feature comprehension and segmentation precision. Extensive experiments on several surgical instrument segmentation datasets demonstrate our model's superior performance and promising generalization capability. To our knowledge, this is the first implementation of a promptable approach to surgical instrument segmentation, offering significant potential for practical application in the field of robotic-assisted surgery. Code is available at https://github.com/franciszzj/TP-SIS. Zijian Zhou 0002, Oluwatosin Alabi, Tom Vercauteren, Miaojing Shi |
NeurIPS | 4 |
| 2023 | TISS-net: Brain tumor image synthesis and segmentation using cascaded dual-task networks and error-prediction consistencyabstractAccurate segmentation of brain tumors from medical images is important for diagnosis and treatment planning, and it often requires multi-modal or contrast-enhanced images. However, in practice some modalities of a patient may be absent. Synthesizing the missing modality has a potential for filling this gap and achieving high segmentation performance. Existing methods often treat the synthesis and segmentation tasks separately or consider them jointly but without effective regularization of the complex joint model, leading to limited performance. We propose a novel brain Tumor Image Synthesis and Segmentation network (TISS-Net) that obtains the synthesized target modality and segmentation of brain tumors end-to-end with high performance. First, we propose a dual-task-regularized generator that simultaneously obtains a synthesized target modality and a coarse segmentation, which leverages a tumor-aware synthesis loss with perceptibility regularization to minimize the high-level semantic domain gap between synthesized and real target modalities. Based on the synthesized image and the coarse segmentation, we further propose a dual-task segmentor that predicts a refined segmentation and error in the coarse segmentation simultaneously, where a consistency between these two predictions is introduced for regularization. Our TISS-Net was validated with two applications: synthesizing FLAIR images for whole glioma segmentation, and synthesizing contrast-enhanced T1 images for Vestibular Schwannoma segmentation. Experimental results showed that our TISS-Net largely improved the segmentation accuracy compared with direct segmentation from the available modalities, and it outperformed state-of-the-art image synthesis-based segmentation methods. Jianghao Wu 0001, Lu Wang 0002, Shuojue Yang, Yuanjie Zheng, Jonathan Shapey, Tom Vercauteren, Sotirios Bisdas, Robert Bradford, Shakeel R. Saeed, Neil Kitchen, Sébastien Ourselin, Shaoting Zhang 0001, Guotai Wang |
Neurocomputing | 7 |
| 2023 | CrossMoDA 2021 challenge: Benchmark of cross-modality domain adaptation techniques for vestibular schwannoma and cochlea segmentationabstractDomain Adaptation (DA) has recently been of strong interest in the medical imaging community. While a large variety of DA techniques have been proposed for image segmentation, most of these techniques have been validated either on private datasets or on small publicly available datasets. Moreover, these datasets mostly addressed single-class problems. To tackle these limitations, the Cross-Modality Domain Adaptation (crossMoDA) challenge was organised in conjunction with the 24th International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI 2021). CrossMoDA is the first large and multi-class benchmark for unsupervised cross-modality Domain Adaptation. The goal of the challenge is to segment two key brain structures involved in the follow-up and treatment planning of vestibular schwannoma (VS): the VS and the cochleas. Currently, the diagnosis and surveillance in patients with VS are commonly performed using contrast-enhanced T1 (ceT1) MR imaging. However, there is growing interest in using non-contrast imaging sequences such as high-resolution T2 (hrT2) imaging. For this reason, we established an unsupervised cross-modality segmentation benchmark. The training dataset provides annotated ceT1 scans (N=105) and unpaired non-annotated hrT2 scans (N=105). The aim was to automatically perform unilateral VS and bilateral cochlea segmentation on hrT2 scans as provided in the testing set (N=137). This problem is particularly challenging given the large intensity distribution gap across the modalities and the small volume of the structures. A total of 55 teams from 16 countries submitted predictions to the validation leaderboard. Among them, 16 teams from 9 different countries submitted their algorithm for the evaluation phase. The level of performance reached by the top-performing teams is strikingly high (best median Dice score — VS: 88.4%; Cochleas: 85.7%) and close to full supervision (median Dice score — VS: 92.5%; Cochleas: 87.7%). All top-performing methods made use of an image-to-image translation approach to transform the source-domain images into pseudo-target-domain images. A segmentation network was then trained using these generated images and the manual annotations provided for the source image. Reuben Dorent, Aaron Kujawa, Marina Ivory, Spyridon Bakas, Nicola Rieke, Samuel Joutard, Ben Glocker, Manuel Jorge Cardoso, Marc Modat, Kayhan Batmanghelich, Arseniy Belkov, Maria G. Baldeon Calisto, Jae Won Choi, Benoit M. Dawant, Hexin Dong, Sergio Escalera, Yubo Fan, Lasse Hansen, Mattias P. Heinrich, Smriti Joshi, Victoriya Kashtanova, Hyeongyu Kim, Satoshi Kondo, Christian N. Kruse, Susana K. Lai-Yuen, Hao Li 0108, Buntheng Ly, Ipek Oguz, Hyungseob Shin, Boris Shirokikh, Zixian Su, Guotai Wang, Jianghao Wu 0001, Yanwu Xu 0001, Li Zhang 0047, Sébastien Ourselin, Jonathan Shapey, Tom Vercauteren |
Medical Image Anal. | 40 |
| 2023 | Fetal brain tissue annotation and segmentation challenge resultsabstractIn-utero fetal MRI is emerging as an important tool in the diagnosis and analysis of the developing human brain. Automatic segmentation of the developing fetal brain is a vital step in the quantitative analysis of prenatal neurodevelopment both in the research and clinical context. However, manual segmentation of cerebral structures is time-consuming and prone to error and inter-observer variability. Therefore, we organized the Fetal Tissue Annotation (FeTA) Challenge in 2021 in order to encourage the development of automatic segmentation algorithms on an international level. The challenge utilized FeTA Dataset, an open dataset of fetal brain MRI reconstructions segmented into seven different tissues (external cerebrospinal fluid, gray matter, white matter, ventricles, cerebellum, brainstem, deep gray matter). 20 international teams participated in this challenge, submitting a total of 21 algorithms for evaluation. In this paper, we provide a detailed analysis of the results from both a technical and clinical perspective. All participants relied on deep learning methods, mainly U-Nets, with some variability present in the network architecture, optimization, and image pre- and post-processing. The majority of teams used existing medical imaging deep learning frameworks. The main differences between the submissions were the fine tuning done during training, and the specific pre- and post-processing steps performed. The challenge results showed that almost all submissions performed similarly. Four of the top five teams used ensemble learning methods. However, one team's algorithm performed significantly superior to the other submissions, and consisted of an asymmetrical U-Net network architecture. This paper provides a first of its kind benchmark for future automatic multi-tissue segmentation algorithms for the developing human brain in utero. Kelly Payette, Hongwei Li 0004, Priscille de Dumast, Roxane Licandro, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Yuchen Pei, Lisheng Wang, Juanying Xie, Huiquan Zhang, Guiming Dong, Hao Fu 0014, Guotai Wang, ZunHyan Rieu, Hyun Gi Kim, Davood Karimi, Ali Gholipour, Helena R. Torres, Bruno Oliveira 0002, João L. Vilaça, Netanell Avisdris, Ori Ben-Zvi, Dafna Ben-Bashat, Lucas Fidon, Michael Aertsen, Tom Vercauteren, Daniel Sobotka, Georg Langs, Mireia Alenyà, Maria Inmaculada Villanueva, Oscar Camara 0001, Bella Specktor-Fadida, Leo Joskowicz, Liao Weibin, Lv Yi, Xuesong Li 0003, Moona Mazher, Abdul Qayyum 0002, Domenec Puig, Hamza Kebiri, KuanLun Liao, YiXuan Wu, JinTai Chen, Yunzhi Xu, Lana Vasung, Bjoern Menze, Meritxell Bach Cuadra, András Jakab |
Medical Image Anal. | 32 |
| 2023 | Editorial for special issue on explainable and generalizable deep learning methods for medical image computing
Guotai Wang, Shaoting Zhang 0001, Sharon X. Huang, Tom Vercauteren, Dimitris N. Metaxas |
Medical Image Anal. | 4 |
| 2023 | UPL-SFDA: Uncertainty-Aware Pseudo Label Guided Source-Free Domain Adaptation for Medical Image SegmentationabstractDomain Adaptation (DA) is important for deep learning-based medical image segmentation models to deal with testing images from a new target domain. As the source-domain data are usually unavailable when a trained model is deployed at a new center, Source-Free Domain Adaptation (SFDA) is appealing for data and annotation-efficient adaptation to the target domain. However, existing SFDA methods have a limited performance due to lack of sufficient supervision with source-domain images unavailable and target-domain images unlabeled. We propose a novel Uncertainty-aware Pseudo Label guided (UPL) SFDA method for medical image segmentation. Specifically, we propose Target Domain Growing (TDG) to enhance the diversity of predictions in the target domain by duplicating the pre-trained model's prediction head multiple times with perturbations. The different predictions in these duplicated heads are used to obtain pseudo labels for unlabeled target-domain images and their uncertainty to identify reliable pseudo labels. We also propose a Twice Forward pass Supervision (TFS) strategy that uses reliable pseudo labels obtained in one forward pass to supervise predictions in the next forward pass. The adaptation is further regularized by a mean prediction-based entropy minimization term that encourages confident and consistent results in different prediction heads. UPL-SFDA was validated with a multi-site heart MRI segmentation dataset, a cross-modality fetal brain segmentation dataset, and a 3D fetal tissue segmentation dataset. It improved the average Dice by 5.54, 5.01 and 6.89 percentage points for the three tasks compared with the baseline, respectively, and outperformed several state-of-the-art SFDA methods. Jianghao Wu 0001, Guotai Wang, Ran Gu, Wentao Zhu 0002, Tom Vercauteren, Sébastien Ourselin, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Robust joint registration of multiple stains and MRI for multimodal 3D histology reconstruction: Application to the Allen human brain atlas
Adrià Casamitjana, Marco Lorenzi, Sebastiano Ferraris, Loïc Peter, Marc Modat, Allison Stevens, Bruce Fischl, Tom Vercauteren, Juan Eugenio Iglesias |
Medical Image Anal. | 8 |
| 2022 | Cross-Modality Image Registration Using a Training-Time Privileged Third ModalityabstractIn this work, we consider the task of pairwise cross-modality image registration, which may benefit from exploiting additional images available only at training time from an additional modality that is different to those being registered. As an example, we focus on aligning intra-subject multiparametric Magnetic Resonance (mpMR) images, between T2-weighted (T2w) scans and diffusion-weighted scans with high b-value (DWI [Formula: see text]). For the application of localising tumours in mpMR images, diffusion scans with zero b-value (DWI [Formula: see text]) are considered easier to register to T2w due to the availability of corresponding features. We propose a learning from privileged modality algorithm, using a training-only imaging modality DWI [Formula: see text], to support the challenging multi-modality registration problems. We present experimental results based on 369 sets of 3D multiparametric MRI images from 356 prostate cancer patients and report, with statistical significance, a lowered median target registration error of 4.34 mm, when registering the holdout DWI [Formula: see text] and T2w image pairs, compared with that of 7.96 mm before registration. Results also show that the proposed learning-based registration networks enabled efficient registration with comparable or better accuracy, compared with a classical iterative algorithm and other tested learning-based methods with/without the additional modality. These compared algorithms also failed to produce any significantly improved alignment between DWI [Formula: see text] and T2w in this challenging application. Qianye Yang, David Atkinson, Yunguan Fu, Tom Syer, Wen Yan 0005, Shonit Punwani, Matthew J. Clarkson, Dean C. Barratt, Tom Vercauteren, Yipeng Hu |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Inter Extreme Points Geodesics for End-to-End Weakly Supervised Image Segmentation
Reuben Dorent, Samuel Joutard, Jonathan Shapey, Aaron Kujawa, Marc Modat, Sébastien Ourselin, Tom Vercauteren |
MICCAI (2) | 7 |
| 2021 | Label-Set Loss Functions for Partial Supervision: Application to Fetal Brain 3D MRI Parcellation
Lucas Fidon, Michael Aertsen, Doaa Emam, Nada Mufti, Frédéric Guffens, Thomas Deprest, Philippe Demaerel, Anna L. David, Andrew Melbourne, Sébastien Ourselin, Jan Deprest, Tom Vercauteren |
MICCAI (2) | 12 |
| 2021 | Interactive Segmentation via Deep Learning and B-Spline Explicit Active Surfaces
Helena Williams, João Pedrosa, Laura Cattani, Susanne Housmans, Tom Vercauteren, Jan Deprest, Jan D'hooge |
MICCAI (1) | 5 |
| 2021 | Learning joint segmentation of tissues and brain lesions from task-specific hetero-modal domain-shifted datasetsabstractBrain tissue segmentation from multimodal MRI is a key building block of many neuroimaging analysis pipelines. Established tissue segmentation approaches have, however, not been developed to cope with large anatomical changes resulting from pathology, such as white matter lesions or tumours, and often fail in these cases. In the meantime, with the advent of deep neural networks (DNNs), segmentation of brain lesions has matured significantly. However, few existing approaches allow for the joint segmentation of normal tissue and brain lesions. Developing a DNN for such a joint task is currently hampered by the fact that annotated datasets typically address only one specific task and rely on task-specific imaging protocols including a task-specific set of imaging modalities. In this work, we propose a novel approach to build a joint tissue and lesion segmentation model from aggregated task-specific hetero-modal domain-shifted and partially-annotated datasets. Starting from a variational formulation of the joint problem, we show how the expected risk can be decomposed and optimised empirically. We exploit an upper bound of the risk to deal with heterogeneous imaging modalities across datasets. To deal with potential domain shift, we integrated and tested three conventional techniques based on data augmentation, adversarial learning and pseudo-healthy generation. For each individual task, our joint approach reaches comparable performance to task-specific and fully-supervised models. The proposed framework is assessed on two different types of brain lesions: White matter lesions and gliomas. In the latter case, lacking a joint ground-truth for quantitative assessment purposes, we propose and use a novel clinically-relevant qualitative assessment methodology. Reuben Dorent, Thomas C. Booth, Wenqi Li 0001, Carole H. Sudre, Sina Kafiabadi, Manuel Jorge Cardoso, Sébastien Ourselin, Tom Vercauteren |
Medical Image Anal. | 8 |
| 2021 | Imitation learning for improved 3D PET/MR attenuation correctionabstractThe assessment of the quality of synthesised/pseudo Computed Tomography (pCT) images is commonly measured by an intensity-wise similarity between the ground truth CT and the pCT. However, when using the pCT as an attenuation map (μ-map) for PET reconstruction in Positron Emission Tomography Magnetic Resonance Imaging (PET/MRI) minimising the error between pCT and CT neglects the main objective of predicting a pCT that when used as μ-map reconstructs a pseudo PET (pPET) which is as similar as possible to the gold standard CT-derived PET reconstruction. This observation motivated us to propose a novel multi-hypothesis deep learning framework explicitly aimed at PET reconstruction application. A convolutional neural network (CNN) synthesises pCTs by minimising a combination of the pixel-wise error between pCT and CT and a novel metric-loss that itself is defined by a CNN and aims to minimise consequent PET residuals. Training is performed on a database of twenty 3D MR/CT/PET brain image pairs. Quantitative results on a fully independent dataset of twenty-three 3D MR/CT/PET image pairs show that the network is able to synthesise more accurate pCTs. The Mean Absolute Error on the pCT (110.98 HU ± 19.22 HU) compared to a baseline CNN (172.12 HU ± 19.61 HU) and a multi-atlas propagation approach (153.40 HU ± 18.68 HU), and subsequently lead to a significant improvement in the PET reconstruction error (4.74% ± 1.52% compared to baseline 13.72% ± 2.48% and multi-atlas propagation 6.68% ± 2.06%). Kerstin Kläser 0002, Thomas Varsavsky, Pawel J. Markiewicz, Tom Vercauteren, Alexander Hammers, David Atkinson, Kris Thielemans, Brian F. Hutton, Manuel Jorge Cardoso, Sébastien Ourselin |
Medical Image Anal. | 4 |
| 2021 | MIDeepSeg: Minimally interactive segmentation of unseen objects from medical images using deep learning
Xiangde Luo, Guotai Wang, Tao Song 0002, Jingyang Zhang, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren, Shaoting Zhang 0001 |
Medical Image Anal. | 8 |
| 2021 | Image Compositing for Segmentation of Surgical Tools Without Manual AnnotationsabstractProducing manual, pixel-accurate, image segmentation labels is tedious and time-consuming. This is often a rate-limiting factor when large amounts of labeled images are required, such as for training deep convolutional networks for instrument-background segmentation in surgical scenes. No large datasets comparable to industry standards in the computer vision community are available for this task. To circumvent this problem, we propose to automate the creation of a realistic training dataset by exploiting techniques stemming from special effects and harnessing them to target training performance rather than visual appeal. Foreground data is captured by placing sample surgical instruments over a chroma key (a.k.a. green screen) in a controlled environment, thereby making extraction of the relevant image segment straightforward. Multiple lighting conditions and viewpoints can be captured and introduced in the simulation by moving the instruments and camera and modulating the light source. Background data is captured by collecting videos that do not contain instruments. In the absence of pre-existing instrument-free background videos, minimal labeling effort is required, just to select frames that do not contain surgical instruments from videos of surgical interventions freely available online. We compare different methods to blend instruments over tissue and propose a novel data augmentation approach that takes advantage of the plurality of options. We show that by training a vanilla U-Net on semi-synthetic data only and applying a simple post-processing, we are able to match the results of the same network trained on a publicly available manually labeled real dataset. Luis C. García-Peraza-Herrera, Lucas Fidon, Claudia D'Ettorre, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin |
IEEE Trans. Medical Imaging | 5 |
| 2021 | CA-Net: Comprehensive Attention Convolutional Neural Networks for Explainable Medical Image SegmentationabstractAccurate medical image segmentation is essential for diagnosis and treatment planning of diseases. Convolutional Neural Networks (CNNs) have achieved state-of-the-art performance for automatic medical image segmentation. However, they are still challenged by complicated conditions where the segmentation target has large variations of position, shape and scale, and existing CNNs have a poor explainability that limits their application to clinical decisions. In this work, we make extensive use of multiple attentions in a CNN architecture and propose a comprehensive attention-based CNN (CA-Net) for more accurate and explainable medical image segmentation that is aware of the most important spatial positions, channels and scales at the same time. In particular, we first propose a joint spatial attention module to make the network focus more on the foreground region. Then, a novel channel attention module is proposed to adaptively recalibrate channel-wise feature responses and highlight the most relevant feature channels. Also, we propose a scale attention module implicitly emphasizing the most salient feature maps among multiple scales so that the CNN is adaptive to the size of an object. Extensive experiments on skin lesion segmentation from ISIC 2018 and multi-class segmentation of fetal MRI found that our proposed CA-Net significantly improved the average segmentation Dice score from 87.77% to 92.08% for skin lesion, 84.79% to 87.08% for the placenta and 93.20% to 95.88% for the fetal brain respectively compared with U-Net. It reduced the model size to around 15 times smaller with close or even better accuracy compared with state-of-the-art DeepLabv3+. In addition, it has a much higher explainability than existing networks by visualizing the attention weight maps. Our code is available at https://github.com/HiLab-git/CA-Net. Ran Gu, Guotai Wang, Tao Song 0002, Rui Huang 0001, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren, Shaoting Zhang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Zero-Shot Super-Resolution With a Physically-Motivated Downsampling Kernel for EndomicroscopyabstractSuper-resolution (SR) methods have seen significant advances thanks to the development of convolutional neural networks (CNNs). CNNs have been successfully employed to improve the quality of endomicroscopy imaging. Yet, the inherent limitation of research on SR in endomicroscopy remains the lack of ground truth high-resolution (HR) images, commonly used for both supervised training and reference-based image quality assessment (IQA). Therefore, alternative methods, such as unsupervised SR are being explored. To address the need for non-reference image quality improvement, we designed a novel zero-shot super-resolution (ZSSR) approach that relies only on the endomicroscopy data to be processed in a self-supervised manner without the need for ground-truth HR images. We tailored the proposed pipeline to the idiosyncrasies of endomicroscopy by introducing both: a physically-motivated Voronoi downscaling kernel accounting for the endomicroscope's irregular fibre-based sampling pattern, and realistic noise patterns. We also took advantage of video sequences to exploit a sequence of images for self-supervised zero-shot image quality improvement. We run ablation studies to assess our contribution in regards to the downscaling kernel and noise simulation. We validate our methodology on both synthetic and original data. Synthetic experiments were assessed with reference-based IQA, while our results for original images were evaluated in a user study conducted with both expert and non-expert observers. The results demonstrated superior performance in image quality of ZSSR reconstructions in comparison to the baseline method. The ZSSR is also competitive when compared to supervised single-image SR, especially being the preferred reconstruction technique by experts. Agnieszka Barbara Szczotka, Dzhoshkun I. Shakir, Matthew J. Clarkson, Stephen P. Pereira, Tom Vercauteren |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Deep Placental Vessel Segmentation for Fetoscopic Mosaicking
Sophia Bano, Francisco Vasconcelos 0001, Luke M. Shepherd, Emmanuel B. Vander Poorten, Tom Vercauteren, Sébastien Ourselin, Anna L. David, Jan Deprest, Danail Stoyanov |
MICCAI (3) | 5 |
| 2020 | An Unsupervised Approach to Ultrasound Elastography with End-to-end Strain Regularisation
Rémi Delaunay, Yipeng Hu, Tom Vercauteren |
MICCAI (3) | 3 |
| 2020 | Scribble-Based Domain Adaptation via Co-segmentation
Reuben Dorent, Samuel Joutard, Jonathan Shapey, Sotirios Bisdas, Neil Kitchen, Robert Bradford, Shakeel R. Saeed, Marc Modat, Sébastien Ourselin, Tom Vercauteren |
MICCAI (1) | 10 |
| 2020 | Uncertainty-Guided Efficient Interactive Refinement of Fetal Brain Segmentation from Stacks of MRI Slices
Guotai Wang, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren, Shaoting Zhang 0001 |
MICCAI (4) | 5 |
| 2020 | Longitudinal Image Registration with Temporal-Order and Subject-Specificity Discrimination
Qianye Yang, Yunguan Fu, Francesco Giganti, Nooshin Ghavami, Qingchao Chen, J. Alison Noble, Tom Vercauteren, Dean C. Barratt, Yipeng Hu |
MICCAI (3) | 7 |
| 2020 | Refractive Two-View Reconstruction for Underwater 3D VisionabstractRecovering 3D geometry from cameras in underwater applications involves the Refractive Structure-from-Motion problem where the non-linear distortion of light induced by a change of medium density invalidates the single viewpoint assumption. The pinhole-plus-distortion camera projection model suffers from a systematic geometric bias since refractive distortion depends on object distance. This leads to inaccurate camera pose and 3D shape estimation. To account for refraction, it is possible to use the axial camera model or to explicitly consider one or multiple parallel refractive interfaces whose orientations and positions with respect to the camera can be calibrated. Although it has been demonstrated that the refractive camera model is well-suited for underwater imaging, Refractive Structure-from-Motion remains particularly difficult to use in practice when considering the seldom studied case of a camera with a flat refractive interface. Our method applies to the case of underwater imaging systems whose entrance lens is in direct contact with the external medium. By adopting the refractive camera model, we provide a succinct derivation and expression for the refractive fundamental matrix and use this as the basis for a novel two-view reconstruction method for underwater imaging. For validation we use synthetic data to show the numerical properties of our method and we provide results on real data to demonstrate its practical application within laboratory settings and for medical applications in fluid-immersed endoscopy. We demonstrate our approach outperforms classic two-view Structure-from-Motion method relying on the pinhole-plus-distortion camera model. François Chadebecq, Francisco Vasconcelos 0001, Rene M. Lacher, Efthymios Maneas, Adrien E. Desjardins, Sébastien Ourselin, Tom Vercauteren, Danail Stoyanov |
Int. J. Comput. Vis. | 7 |
| 2020 | Image computing for fibre-bundle endomicroscopy: A review
Antonios Perperidis, Kevin Dhaliwal, Steve McLaughlin 0001, Tom Vercauteren |
Medical Image Anal. | 4 |
| 2020 | CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer-Assisted InterventionsabstractData-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their interpretation and exploitation in clinical practice. While similar benefits are longed for in the field of interventional imaging, this ambition is challenged by a much higher heterogeneity. Clinical workflows within interventional suites and operating theatres are extremely complex and typically rely on poorly integrated intra-operative devices, sensors, and support infrastructures. Taking stock of some of the most exciting developments in machine learning and artificial intelligence for computer assisted interventions, we highlight the crucial need to take context and human factors into account in order to address these challenges. Contextual artificial intelligence for computer assisted intervention, or CAI4CAI, arises as an emerging opportunity feeding into the broader field of surgical data science. Central challenges being addressed in CAI4CAI include how to integrate the ensemble of prior knowledge and instantaneous sensory information from experts, sensors and actuators; how to create and communicate a faithful and actionable shared representation of the surgery among a mixed human-AI actor team; how to design interventional systems and associated cognitive shared control schemes for online uncertainty-aware collaborative decision making ultimately producing more precise and reliable interventions. Tom Vercauteren, Mathias Unberath, Nicolas Padoy, Nassir Navab |
Proc. IEEE | 1 |
| 2019 | Robotic Control of a Multi-Modal Rigid Endoscope Combining Optical Imaging with All-Optical UltrasoundabstractFetoscopy is a technically challenging surgery, due to the dynamic environment and low diameter endoscopes often resulting in a limited field of view. In this paper, we report on the design and operation of a robotic multimodal endoscope with optical ultrasound and white light stereo camera. The manufacture and control of the endoscope is presented, along with large area (80 mm ×80 mm) surface visualisations of a placenta phantom using the optical ultrasound sensor. The repeatability of the surface visualisations was found to be 0. 446 ± 0.139 mm and 0. 267 ± 0.017 mm for a raster and spiral scan, respectively. George Dwyer, Richard J. Colchester, Erwin J. Alles, Efthymios Maneas, Sébastien Ourselin, Tom Vercauteren, Jan Deprest, Emmanuel B. Vander Poorten, Paolo De Coppi, Adrien E. Desjardins, Danail Stoyanov |
ICRA | 6 |
| 2019 | Macro-Micro Multi-Arm Robot for Single-Port Access SurgeryabstractMinimally invasive surgery is now a well established field in surgery but continuous efforts are made to reduce invasiveness even further. This paper proposes a novel concept of small-diameter multi-arm robot for SinglePort Access Surgery. The concept introduces a combination of backbone and actuation principles in a macro-micro fashion to achieve an excellent decoupling of the triangulation platform (macro) and of the end-effectors (micro). Concentric tube robots are used for the triangulation platform, while compliant fluidic-actuated bending segments are used as end-effectors. The fluidic actuation is advantageous as it minimally interferes with the triangulation platform. The triangulation platform on the other hand provides a stable base for the end-effectors such that large distal actuation bandwidth can be achieved. A specific embodiment for Spina Bifida repair is developed and proposed. The surgical and technical requirements as well as the mechanical design are presented in details. A first prototype is built and characterization experiments are conducted to evaluate its performance. T. Vandebroek, Mouloud Ourak, Caspar Gruijthuijsen, Allan Javaux, Julie Legrand, Tom Vercauteren, Sébastien Ourselin, Jan Deprest, Emmanuel B. Vander Poorten |
IROS | 6 |
| 2019 | Deep Sequential Mosaicking of Fetoscopic Videos
Sophia Bano, Francisco Vasconcelos 0001, Marcel Tella-Amo, George Dwyer, Caspar Gruijthuijsen, Jan Deprest, Sébastien Ourselin, Emmanuel B. Vander Poorten, Tom Vercauteren, Danail Stoyanov |
MICCAI (1) | 9 |
| 2019 | Hetero-Modal Variational Encoder-Decoder for Joint Modality Completion and Segmentation
Reuben Dorent, Samuel Joutard, Marc Modat, Sébastien Ourselin, Tom Vercauteren |
MICCAI (2) | 5 |
| 2019 | Incompressible Image Registration Using Divergence-Conforming B-Splines
Lucas Fidon, Michael Ebner, Luis C. García-Peraza-Herrera, Marc Modat, Sébastien Ourselin, Tom Vercauteren |
MICCAI (2) | 6 |
| 2019 | Improved Placental Parameter Estimation Using Data-Driven Bayesian Modelling
Dimitra Flouri, David Owen 0001, Rosalind Aughwane, Nada Mufti, Magdalena J. Sokolska, David Atkinson, Giles S. Kendall, Alan Bainbridge, Tom Vercauteren, Anna L. David, Sébastien Ourselin, Andrew Melbourne |
MICCAI (3) | 9 |
| 2019 | Conditional Segmentation in Lieu of Image Registration
Yipeng Hu, Eli Gibson, Dean C. Barratt, Mark Emberton, J. Alison Noble, Tom Vercauteren |
MICCAI (2) | 6 |
| 2019 | Permutohedral Attention Module for Efficient Non-local Neural Networks
Samuel Joutard, Reuben Dorent, Amanda Isaac, Sébastien Ourselin, Tom Vercauteren, Marc Modat |
MICCAI (6) | 5 |
| 2019 | Automatic Segmentation of Vestibular Schwannoma from T2-Weighted MRI by Deep Spatial Attention with Hardness-Weighted Loss
Guotai Wang, Jonathan Shapey, Wenqi Li 0001, Reuben Dorent, Alex Demitriadis, Sotirios Bisdas, Ian Paddick, Robert Bradford, Shaoting Zhang 0001, Sébastien Ourselin, Tom Vercauteren |
MICCAI (2) | 11 |
| 2019 | Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networksabstractDespite the state-of-the-art performance for medical image segmentation, deep convolutional neural networks (CNNs) have rarely provided uncertainty estimations regarding their segmentation outputs, e.g., model (epistemic) and image-based (aleatoric) uncertainties. In this work, we analyze these different types of uncertainties for CNN-based 2D and 3D medical image segmentation tasks at both pixel level and structure level. We additionally propose a test-time augmentation-based aleatoric uncertainty to analyze the effect of different transformations of the input image on the segmentation output. Test-time augmentation has been previously used to improve segmentation accuracy, yet not been formulated in a consistent mathematical framework. Hence, we also propose a theoretical formulation of test-time augmentation, where a distribution of the prediction is estimated by Monte Carlo simulation with prior distributions of parameters in an image acquisition model that involves image transformations and noise. We compare and combine our proposed aleatoric uncertainty with model uncertainty. Experiments with segmentation of fetal brains and brain tumors from 2D and 3D Magnetic Resonance Images (MRI) showed that 1) the test-time augmentation-based aleatoric uncertainty provides a better uncertainty estimation than calculating the test-time dropout-based model uncertainty alone and helps to reduce overconfident incorrect predictions, and 2) our test-time augmentation outperforms a single-prediction baseline and dropout-based multiple predictions. Guotai Wang, Wenqi Li 0001, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren |
Neurocomputing | 6 |
| 2019 | Adversarial training with cycle consistency for unsupervised super-resolution in endomicroscopyabstractIn recent years, endomicroscopy has become increasingly used for diagnostic purposes and interventional guidance. It can provide intraoperative aids for real-time tissue characterization and can help to perform visual investigations aimed for example to discover epithelial cancers. Due to physical constraints on the acquisition process, endomicroscopy images, still today have a low number of informative pixels which hampers their quality. Post-processing techniques, such as Super-Resolution (SR), are a potential solution to increase the quality of these images. SR techniques are often supervised, requiring aligned pairs of low-resolution (LR) and high-resolution (HR) images patches to train a model. However, in our domain, the lack of HR images hinders the collection of such pairs and makes supervised training unsuitable. For this reason, we propose an unsupervised SR framework based on an adversarial deep neural network with a physically-inspired cycle consistency, designed to impose some acquisition properties on the super-resolved images. Our framework can exploit HR images, regardless of the domain where they are coming from, to transfer the quality of the HR images to the initial LR images. This property can be particularly useful in all situations where pairs of LR/HR are not available during the training. Our quantitative analysis, validated using a database of 238 endomicroscopy video sequences from 143 patients, shows the ability of the pipeline to produce convincing super-resolved images. A Mean Opinion Score (MOS) study also confirms this quantitative image quality assessment. Daniele Ravì, Agnieszka Barbara Szczotka, Stephen P. Pereira, Tom Vercauteren |
Medical Image Anal. | 4 |
| 2019 | DeepIGeoS: A Deep Interactive Geodesic Framework for Medical Image SegmentationabstractAccurate medical image segmentation is essential for diagnosis, surgical planning and many other applications. Convolutional Neural Networks (CNNs) have become the state-of-the-art automatic segmentation methods. However, fully automatic results may still need to be refined to become accurate and robust enough for clinical use. We propose a deep learning-based interactive segmentation method to improve the results obtained by an automatic CNN and to reduce user interactions during refinement for higher accuracy. We use one CNN to obtain an initial automatic segmentation, on which user interactions are added to indicate mis-segmentations. Another CNN takes as input the user interactions with the initial segmentation and gives a refined result. We propose to combine user interactions with CNNs through geodesic distance transforms, and propose a resolution-preserving network that gives a better dense prediction. In addition, we integrate user interactions as hard constraints into a back-propagatable Conditional Random Field. We validated the proposed framework in the context of 2D placenta segmentation from fetal MRI and 3D brain tumor segmentation from FLAIR images. Experimental results show our method achieves a large improvement from automatic CNNs, and obtains comparable and even higher accuracy with fewer user interventions and less time compared with traditional interactive methods. Guotai Wang, Maria A. Zuluaga, Wenqi Li 0001, Rosalind Pratt, Premal A. Patel, Michael Aertsen, Tom Doel, Anna L. David, Jan Deprest, Sébastien Ourselin, Tom Vercauteren |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2018 | MRI Measurement of Placental Perfusion and Fetal Blood Oxygen Saturation in Normal Pregnancy and Placental Insufficiency
Rosalind Aughwane, Magdalena J. Sokolska, Alan Bainbridge, David Atkinson, Giles S. Kendall, Jan Deprest, Tom Vercauteren, Anna L. David, Sébastien Ourselin, Andrew Melbourne |
MICCAI (2) | 7 |
| 2018 | An Automated Localization, Segmentation and Reconstruction Framework for Fetal Brain MRI
Michael Ebner, Guotai Wang, Wenqi Li 0001, Michael Aertsen, Premal A. Patel, Rosalind Aughwane, Andrew Melbourne, Tom Doel, Anna L. David, Jan Deprest, Sébastien Ourselin, Tom Vercauteren |
MICCAI (1) | 12 |
| 2018 | Adversarial Deformation Regularization for Training Image Registration Neural Networks
Yipeng Hu, Eli Gibson, Nooshin Ghavami, Ester Bonmati, Caroline M. Moore, Mark Emberton, Tom Vercauteren, J. Alison Noble, Dean C. Barratt |
MICCAI (1) | 7 |
| 2018 | Model-Based Refinement of Nonlinear Registrations in 3D Histology Reconstruction
Juan Eugenio Iglesias, Marco Lorenzi, Sebastiano Ferraris, Loïc Peter, Marc Modat, Allison Stevens, Bruce Fischl, Tom Vercauteren |
MICCAI (2) | 8 |
| 2018 | Weakly-supervised convolutional neural networks for multimodal image registrationabstractOne of the fundamental challenges in supervised learning for multimodal image registration is the lack of ground-truth for voxel-level spatial correspondence. This work describes a method to infer voxel-level transformation from higher-level correspondence information contained in anatomical labels. We argue that such labels are more reliable and practical to obtain for reference sets of image pairs than voxel-level correspondence. Typical anatomical labels of interest may include solid organs, vessels, ducts, structure boundaries and other subject-specific ad hoc landmarks. The proposed end-to-end convolutional neural network approach aims to predict displacement fields to align multiple labelled corresponding structures for individual image pairs during the training, while only unlabelled image pairs are used as the network input for inference. We highlight the versatility of the proposed strategy, for training, utilising diverse types of anatomical labels, which need not to be identifiable over all training image pairs. At inference, the resulting 3D deformable image registration algorithm runs in real-time and is fully-automated without requiring any anatomical labels or initialisation. Several network architecture variants are compared for registering T2-weighted magnetic resonance images and 3D transrectal ultrasound images from prostate cancer patients. A median target registration error of 3.6 mm on landmark centroids and a median Dice of 0.87 on prostate glands are achieved from cross-validation experiments, in which 108 pairs of multimodal images from 76 patients were tested with high-quality anatomical labels. Yipeng Hu, Marc Modat, Eli Gibson, Wenqi Li 0001, Nooshin Ghavami, Ester Bonmati, Guotai Wang, Steven Bandula, Caroline M. Moore, Mark Emberton, Sébastien Ourselin, J. Alison Noble, Dean C. Barratt, Tom Vercauteren |
Medical Image Anal. | 14 |
| 2018 | Joint registration and synthesis using a probabilistic model for alignment of MRI and histological sectionsabstractNonlinear registration of 2D histological sections with corresponding slices of MRI data is a critical step of 3D histology reconstruction algorithms. This registration is difficult due to the large differences in image contrast and resolution, as well as the complex nonrigid deformations and artefacts produced when sectioning the sample and mounting it on the glass slide. It has been shown in brain MRI registration that better spatial alignment across modalities can be obtained by synthesising one modality from the other and then using intra-modality registration metrics, rather than by using information theory based metrics to solve the problem directly. However, such an approach typically requires a database of aligned images from the two modalities, which is very difficult to obtain for histology and MRI. Here, we overcome this limitation with a probabilistic method that simultaneously solves for deformable registration and synthesis directly on the target images, without requiring any training data. The method is based on a probabilistic model in which the MRI slice is assumed to be a contrast-warped, spatially deformed version of the histological section. We use approximate Bayesian inference to iteratively refine the probabilistic estimate of the synthesis and the registration, while accounting for each other's uncertainty. Moreover, manually placed landmarks can be seamlessly integrated in the framework for increased performance and robustness. Experiments on a synthetic dataset of MRI slices show that, compared with mutual information based registration, the proposed method makes it possible to use a much more flexible deformation model in the registration to improve its accuracy, without compromising robustness. Moreover, our framework also exploits information in manually placed landmarks more efficiently than mutual information: landmarks constrain the deformation field in both methods, but in our algorithm, it also has a positive effect on the synthesis - which further improves the registration. We also show results on two real, publicly available datasets: the Allen and BigBrain atlases. In both of them, the proposed method provides a clear improvement over mutual information based registration, both qualitatively (visual inspection) and quantitatively (registration error measured with pairs of manually annotated landmarks). Juan Eugenio Iglesias, Marc Modat, Loïc Peter, Allison Stevens, Roberto Annunziata, Tom Vercauteren, Ed S. Lein, Bruce Fischl, Sébastien Ourselin |
Medical Image Anal. | 6 |
| 2018 | Interactive Medical Image Segmentation Using Deep Learning With Image-Specific Fine TuningabstractConvolutional neural networks (CNNs) have achieved state-of-the-art performance for automatic medical image segmentation. However, they have not demonstrated sufficiently accurate and robust results for clinical use. In addition, they are limited by the lack of image-specific adaptation and the lack of generalizability to previously unseen object classes (a.k.a. zero-shot learning). To address these problems, we propose a novel deep learning-based interactive segmentation framework by incorporating CNNs into a bounding box and scribble-based segmentation pipeline. We propose image-specific fine tuning to make a CNN model adaptive to a specific test image, which can be either unsupervised (without additional user interactions) or supervised (with additional scribbles). We also propose a weighted loss function considering network and interaction-based uncertainty for the fine tuning. We applied this framework to two applications: 2-D segmentation of multiple organs from fetal magnetic resonance (MR) slices, where only two types of these organs were annotated for training and 3-D segmentation of brain tumor core (excluding edema) and whole brain tumor (including edema) from different MR sequences, where only the tumor core in one MR sequence was annotated for training. Experimental results show that: 1) our model is more robust to segment previously unseen objects than state-of-the-art CNNs; 2) image-specific fine tuning with the proposed weighted loss function significantly improves segmentation accuracy; and 3) our method leads to accurate results with fewer user interactions and less user time than traditional interactive segmentation methods. Guotai Wang, Wenqi Li 0001, Maria A. Zuluaga, Rosalind Pratt, Premal A. Patel, Michael Aertsen, Tom Doel, Anna L. David, Jan Deprest, Sébastien Ourselin, Tom Vercauteren |
IEEE Trans. Medical Imaging | 11 |
| 2017 | Refractive Structure-from-Motion Through a Flat Refractive InterfaceabstractRecovering 3D scene geometry from underwater images involves the Refractive Structure-from-Motion (RSfM) problem, where the image distortions caused by light refraction at the interface between different propagation media invalidates the single view point assumption. Direct use of the pinhole camera model in RSfM leads to inaccurate camera pose estimation and consequently drift. RSfM methods have been thoroughly studied for the case of a thick glass interface that assumes two refractive interfaces between the camera and the viewed scene. On the other hand, when the camera lens is in direct contact with the water, there is only one refractive interface. By explicitly considering a refractive interface, we develop a succinct derivation of the refractive fundamental matrix in the form of the generalised epipolar constraint for an axial camera. We use the refractive fundamental matrix to refine initial pose estimates obtained by assuming the pinhole model. This strategy allows us to robustly estimate underwater camera poses, where other methods suffer from poor noise-sensitivity. We also formulate a new four view constraint enforcing camera pose consistency along a video which leads us to a novel RSfM framework. For validation we use synthetic data to show the numerical properties of our method and we provide results on real data to demonstrate performance within laboratory settings and for applications in endoscopy. François Chadebecq, Francisco Vasconcelos 0001, George Dwyer, Rene M. Lacher, Sébastien Ourselin, Tom Vercauteren, Danail Stoyanov |
ICCV | 6 |
| 2017 | ToolNet: Holistically-nested real-time segmentation of robotic surgical toolsabstractReal-time tool segmentation from endoscopic videos is an essential part of many computer-assisted robotic surgical systems and of critical importance in robotic surgical data science. We propose two novel deep learning architectures for automatic segmentation of non-rigid surgical instruments. Both methods take advantage of automated deep-learning-based multi-scale feature extraction while trying to maintain an accurate segmentation quality at all resolutions. The two proposed methods encode the multi-scale constraint inside the network architecture. The first proposed architecture enforces it by cascaded aggregation of predictions and the second proposed network does it by means of a holistically-nested architecture where the loss at each scale is taken into account for the optimization process. As the proposed methods are for real-time semantic labeling, both present a reduced number of parameters. We propose the use of parametric rectified linear units for semantic labeling in these small architectures to increase the regularization of the network while maintaining the segmentation accuracy. We compare the proposed architectures against state-of-the-art fully convolutional networks. We validate our methods using existing benchmark datasets, including ex vivo cases with phantom tissue and different robotic surgical instruments present in the scene. Our results show a statistically significant improved Dice Similarity Coefficient over previous instrument segmentation methods. We analyze our design choices and discuss the key drivers for improving accuracy. Luis C. García-Peraza-Herrera, Wenqi Li 0001, Lucas Fidon, Caspar Gruijthuijsen, Alain Devreker, George Attilakos, Jan Deprest, Emmanuel B. Vander Poorten, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin |
IROS | 10 |
| 2017 | Body wall force sensor for simulated minimally invasive surgery: Application to fetal surgeryabstractSurgical interventions are increasingly executed minimal invasively. Surgeons insert instruments through tiny incisions in the body and pivot slender instruments to treat organs or tissue below the surface. While a blessing for patients, surgeons need to pay extra attention to overcome the fulcrum effect, reduced haptic feedback and deal with lost hand-eye coordination. The mental load makes it difficult to pay sufficient attention to the forces that are exerted on the body wall. In delicate procedures such as fetal surgery, this might be problematic as irreparable damage could cause premature delivery. As a first attempt to quantify the interaction forces applied on the patient's body wall, a novel 6 degrees of freedom force sensor was developed for an ex-vivo set up. The performance of the sensor was characterised. User experiments were conducted by 3 clinicians on a set up simulating a fetal surgical intervention. During these simulated interventions, the interaction forces were recorded and analysed when a normal instrument was employed. These results were compared with a session where a flexible instrument under haptic guidance was used. The conducted experiments resulted in interesting insights in the interaction forces and stresses that develop during such difficult surgical intervention. The results also implicated that haptic guidance schemes and the use of flexible instruments rather than rigid ones could have a significant impact on the stresses that occur at the body wall. Allan Javaux, Laure Esteveny, David Bouget, Caspar Gruijthuijsen, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin, Dominiek Reynaerts, Kathleen Denis, Jan Deprest, Emmanuel B. Vander Poorten |
IROS | 6 |
| 2017 | Scalable Multimodal Convolutional Networks for Brain Tumour Segmentation
Lucas Fidon, Wenqi Li 0001, Luis C. García-Peraza-Herrera, Jinendra Ekanayake, Neil Kitchen, Sébastien Ourselin, Tom Vercauteren |
MICCAI (3) | 7 |
| 2017 | Intraoperative Organ Motion Models with an Ensemble of Conditional Generative Adversarial Networks
Yipeng Hu, Eli Gibson, Tom Vercauteren, Hashim Uddin Ahmed, Mark Emberton, Caroline M. Moore, J. Alison Noble, Dean C. Barratt |
MICCAI (2) | 3 |
| 2016 | Bilateral Weighted Adaptive Local Similarity Measure for Registration in Neurosurgery
Martin Kochan, Marc Modat, Tom Vercauteren, Mark White 0001, Laura Mancini, Gavin Winston, Andrew W. McEvoy, John S. Thornton, Tarek A. Yousry, John S. Duncan, Sébastien Ourselin, Danail Stoyanov |
MICCAI (3) | 3 |
| 2016 | Dynamically Balanced Online Random Forests for Interactive Scribble-Based Segmentation
Guotai Wang, Maria A. Zuluaga, Rosalind Pratt, Michael Aertsen, Tom Doel, Maria Klusmann, Anna L. David, Jan Deprest, Tom Vercauteren, Sébastien Ourselin |
MICCAI (2) | 9 |
| 2016 | From computer-assisted intervention research to clinical impact: The need for a holistic approach
Sébastien Ourselin, Mark Emberton, Tom Vercauteren |
Medical Image Anal. | 3 |
| 2016 | Slic-Seg: A minimally interactive segmentation of the placenta from sparse and motion-corrupted fetal MRI in multiple viewsabstractSegmentation of the placenta from fetal MRI is challenging due to sparse acquisition, inter-slice motion, and the widely varying position and shape of the placenta between pregnant women. We propose a minimally interactive framework that combines multiple volumes acquired in different views to obtain accurate segmentation of the placenta. In the first phase, a minimally interactive slice-by-slice propagation method called Slic-Seg is used to obtain an initial segmentation from a single motion-corrupted sparse volume image. It combines high-level features, online Random Forests and Conditional Random Fields, and only needs user interactions in a single slice. In the second phase, to take advantage of the complementary resolution in multiple volumes acquired in different views, we further propose a probability-based 4D Graph Cuts method to refine the initial segmentations using inter-slice and inter-image consistency. We used our minimally interactive framework to examine the placentas of 16 mid-gestation patients from MRI acquired in axial and sagittal views respectively. The results show the proposed method has 1) a good performance even in cases where sparse scribbles provided by the user lead to poor results with the competitive propagation approaches; 2) a good interactivity with low intra- and inter-operator variability; 3) higher accuracy than state-of-the-art interactive segmentation methods; and 4) an improved accuracy due to the co-segmentation based refinement, which outperforms single volume or intensity-based Graph Cuts. Guotai Wang, Maria A. Zuluaga, Rosalind Pratt, Michael Aertsen, Tom Doel, Maria Klusmann, Anna L. David, Jan Deprest, Tom Vercauteren, Sébastien Ourselin |
Medical Image Anal. | 9 |
| 2015 | Fluidic actuation for intra-operative in situ imagingabstractA novel fluidic actuation system has been developed for in situ imaging of anatomic tissues. The actuator consists of a micromachined superelastic tool guide driven by a pair of pneumatic artificial muscles. Two additional working channels allow easy interchange of instruments or sensing equipment. This paper describes the design and construction of the actuation system. Experimental results are also reported indicating a bending repeatability of 0.1 degrees and an operational bandwidth exceeding 8Hz. To show-case the performance of the device, the actuator was loaded with an all-optical ultrasound imaging probe. First scanned images of human placental tissue surface using an all-optical ultrasound probe are presented. While a model has been developed to estimate the probe position in space as function of the input pressure, in future work, this model will be complemented with additional sensor measurements of the bending probe taking into account the hysteretic behaviour of both muscles and nitinol structure. Alain Devreker, Benoit Rosa, Adrien E. Desjardins, Erwin J. Alles, Luis C. García-Peraza-Herrera, Efthymios Maneas, Danail Stoyanov, Anna L. David, Tom Vercauteren, Jan Deprest, Sébastien Ourselin, Dominiek Reynaerts, Emmanuel B. Vander Poorten |
IROS | 9 |
| 2015 | Scale Factor Point Spread Function Matching: Beyond Aliasing in Image Resampling
Manuel Jorge Cardoso, Marc Modat, Tom Vercauteren, Sébastien Ourselin |
MICCAI (2) | 3 |
| 2015 | Motion-Aware Mosaicing for Confocal Laser Endomicroscopy
Jessie Mahé, Nicolas Linard, Marzieh Kohandani Tafreshi, Tom Vercauteren, Nicholas Ayache, François Lacombe, Rémi Cuingnet |
MICCAI (1) | 4 |
| 2015 | A Registration Approach to Endoscopic Laser Speckle Contrast Imaging for Intrauterine Visualisation of Placental Vessels
Gustavo Sato dos Santos, Efthymios Maneas, Daniil I. Nikitichev, Anamaria Barburas, Anna L. David, Jan Deprest, Adrien E. Desjardins, Tom Vercauteren, Sébastien Ourselin |
MICCAI (1) | 8 |
| 2015 | Slic-Seg: Slice-by-Slice Segmentation Propagation of the Placenta in Fetal MRI Using One-Plane Scribbles and Online Learning
Guotai Wang, Maria A. Zuluaga, Rosalind Pratt, Michael Aertsen, Anna L. David, Jan Deprest, Tom Vercauteren, Sébastien Ourselin |
MICCAI (3) | 7 |
| 2015 | Interventional Photoacoustic Imaging of the Human Placenta with Ultrasonic Tracking for Minimally Invasive Fetal Surgeries
Wenfeng Xia 0001, Efthymios Maneas, Daniil I. Nikitichev, Charles A. Mosse, Gustavo Sato dos Santos, Tom Vercauteren, Anna L. David, Jan Deprest, Sébastien Ourselin, Paul C. Beard, Adrien E. Desjardins |
MICCAI (1) | 6 |
| 2014 | Semi-automated Query Construction for Content-Based Endomicroscopy Video Retrieval
Marzieh Kohandani Tafreshi, Nicolas Linard, Barbara André, Nicholas Ayache, Tom Vercauteren |
MICCAI (1) | 5 |
| 2013 | A Viterbi Approach to Topology Inference for Large Scale Endomicroscopy Video Mosaicing
Jessie Mahé, Tom Vercauteren, Benoit Rosa, Julien Dauguet |
MICCAI (1) | 2 |
| 2012 | Scanning the surface of soft tissues with a micrometer precision thanks to endomicroscopy based visual servoingabstractProbe-based confocal laser endomicroscopy is a recent tissue imaging technology that requires placing a probe in contact with the tissue to be imaged and provides real time images with a microscopic resolution. Additionally, generating adequate probe movements to sweep the tissue surface can be used to reconstruct a wide mosaic of the scanned region while increasing the resolution which is appropriate for anatomico-pathological cancer diagnosis. However, properly controlling the motion along the scanning trajectory is a major problem. Indeed, the tissue exhibits deformations under friction forces exerted by the probe leading to deformed mosaics. In this paper we propose a visual servoing approach for controlling the probe movements relative to the tissue while rejecting the tissue deformation disturbance. The probe displacement with respect to the tissue is firstly estimated using the confocal images and an image registration real-time algorithm. Secondly, from this real-time image-based position measurement, the probe motion is controlled thanks to a simple proportional-integral compensator and a feedforward term. Ex vivo experiments using a Stäubli TX40 robot and a Mauna Kea Technologies Cellvizio imaging device demonstrate the effectiveness of the approach on liver and muscle tissue. Benoit Rosa, Mustafa Suphi Erden, Tom Vercauteren, Jérôme Szewczyk, Guillaume Morel |
IROS | 3 |
| 2012 | Online Blind Calibration of Non-uniform Photodetectors: Application to Endomicroscopy
Nicolas Savoire, Barbara André, Tom Vercauteren |
MICCAI (3) | 3 |
| 2012 | Re-localisation of a biopsy site in endoscopic images and characterisation of its uncertainty
Baptiste Allain, Mingxing Hu, Laurence B. Lovat, Richard J. Cook, Tom Vercauteren, Sébastien Ourselin, David J. Hawkes |
Medical Image Anal. | 5 |
| 2012 | Learning Semantic and Visual Similarity for Endomicroscopy Video RetrievalabstractContent-based image retrieval (CBIR) is a valuable computer vision technique which is increasingly being applied in the medical community for diagnosis support. However, traditional CBIR systems only deliver visual outputs, i.e., images having a similar appearance to the query, which is not directly interpretable by the physicians. Our objective is to provide a system for endomicroscopy video retrieval which delivers both visual and semantic outputs that are consistent with each other. In a previous study, we developed an adapted bag-of-visual-words method for endomicroscopy retrieval, called "Dense-Sift," that computes a visual signature for each video. In this paper, we present a novel approach to complement visual similarity learning with semantic knowledge extraction, in the field of in vivo endomicroscopy. We first leverage a semantic ground truth based on eight binary concepts, in order to transform these visual signatures into semantic signatures that reflect how much the presence of each semantic concept is expressed by the visual words describing the videos. Using cross-validation, we demonstrate that, in terms of semantic detection, our intuitive Fisher-based method transforming visual-word histograms into semantic estimations outperforms support vector machine (SVM) methods with statistical significance. In a second step, we propose to improve retrieval relevance by learning an adjusted similarity distance from a perceived similarity ground truth. As a result, our distance learning method allows to statistically improve the correlation with the perceived similarity. We also demonstrate that, in terms of perceived similarity, the recall performance of the semantic signatures is close to that of visual signatures and significantly better than those of several state-of-the-art CBIR methods. The semantic signatures are thus able to communicate high-level medical knowledge while being consistent with the low-level visual signatures and much shorter than them. In our resulting retrieval system, we decide to use visual signatures for perceived similarity learning and retrieval, and semantic signatures for the output of an additional information, expressed in the endoscopist own language, which provides a relevant semantic translation of the visual retrieval outputs. Barbara André, Tom Vercauteren, Anna M. Buchner, Michael B. Wallace, Nicholas Ayache |
IEEE Trans. Medical Imaging | 2 |
| 2011 | Retrieval Evaluation and Distance Learning from Perceived Similarity between Endomicroscopy Videos
Barbara André, Tom Vercauteren, Anna M. Buchner, Michael B. Wallace, Nicholas Ayache |
MICCAI (3) | 2 |
| 2011 | A smart atlas for endomicroscopy using automated video retrieval
Barbara André, Tom Vercauteren, Anna M. Buchner, Michael B. Wallace, Nicholas Ayache |
Medical Image Anal. | 2 |
| 2011 | Evaluation of Registration Methods on Thoracic CT: The EMPIRE10 ChallengeabstractEMPIRE10 (Evaluation of Methods for Pulmonary Image REgistration 2010) is a public platform for fair and meaningful comparison of registration algorithms which are applied to a database of intrapatient thoracic CT image pairs. Evaluation of nonrigid registration techniques is a nontrivial task. This is compounded by the fact that researchers typically test only on their own data, which varies widely. For this reason, reliable assessment and comparison of different registration algorithms has been virtually impossible in the past. In this work we present the results of the launch phase of EMPIRE10, which comprised the comprehensive evaluation and comparison of 20 individual algorithms from leading academic and industrial research groups. All algorithms are applied to the same set of 30 thoracic CT pairs. Algorithm settings and parameters are chosen by researchers expert in the configuration of their own method and the evaluation is independent, using the same criteria for all participants. All results are published on the EMPIRE10 website (http://empire10.isi.uu.nl). The challenge remains ongoing and open to new participants. Full results from 24 algorithms have been published at the time of writing. This paper details the organization of the challenge, the data and evaluation methods and the outcome of the initial launch with 20 algorithms. The gain in knowledge and future work are discussed. Keelin Murphy, Bram van Ginneken, Joseph M. Reinhardt, Sven Kabus, Kai Ding 0003, Kunlin Cao, Kaifang Du, Gary E. Christensen, Vincent Garcia, Tom Vercauteren, Nicholas Ayache, Olivier Commowick, Grégoire Malandain, Ben Glocker, Nikos Paragios, Nassir Navab, Vladlena Gorbunova, Jon Sporring, Marleen de Bruijne, Xiao Han 0011, Mattias P. Heinrich, Julia A. Schnabel, Mark Jenkinson, Cristian Lorenz, Marc Modat, Jamie McClelland, Sébastien Ourselin, Sascha E. A. Muenzing, Max A. Viergever, Dante De Nigris, D. Louis Collins, Tal Arbel, Marta Peroni, Rui Li 0053, Gregory C. Sharp, Alexander Schmidt-Richberg, Jan Ehrhardt, René Werner, Dirk Smeets, Dirk Loeckx, Gang Song, Nicholas J. Tustison, Brian B. Avants, James C. Gee, Marius Staring, Stefan Klein 0001, Berend C. Stoel, Martin Urschler, Manuel Werlberger, Jef Vandemeulebroucke, Simon Rit, David Sarrut, Josien P. W. Pluim |
IEEE Trans. Medical Imaging | 11 |
| 2010 | A System for Biopsy Site Re-targeting with Uncertainty in Gastroenterology and Oropharyngeal Examinations
Baptiste Allain, Mingxing Hu, Laurence B. Lovat, Richard J. Cook, Tom Vercauteren, Sébastien Ourselin, David J. Hawkes |
MICCAI (2) | 5 |
| 2010 | An Image Retrieval Approach to Setup Difficulty Levels in Training Systems for Endomicroscopy Diagnosis
Barbara André, Tom Vercauteren, Anna M. Buchner, Muhammad Waseem Shahid, Michael B. Wallace, Nicholas Ayache |
MICCAI (2) | 2 |
| 2010 | Spherical Demons: Fast Diffeomorphic Landmark-Free Surface RegistrationabstractWe present the Spherical Demons algorithm for registering two spherical images. By exploiting spherical vector spline interpolation theory, we show that a large class of regularizors for the modified Demons objective function can be efficiently approximated on the sphere using iterative smoothing. Based on one parameter subgroups of diffeomorphisms, the resulting registration is diffeomorphic and fast. The Spherical Demons algorithm can also be modified to register a given spherical image to a probabilistic atlas. We demonstrate two variants of the algorithm corresponding to warping the atlas or warping the subject. Registration of a cortical surface mesh to an atlas mesh, both with more than 160 k nodes requires less than 5 min when warping the atlas and less than 3 min when warping the subject on a Xeon 3.2 GHz single processor machine. This is comparable to the fastest nondiffeomorphic landmark-free surface registration algorithms. Furthermore, the accuracy of our method compares favorably to the popular FreeSurfer registration algorithm. We validate the technique in two different applications that use registration to transfer segmentation labels onto a new image 1) parcellation of in vivo cortical surfaces and 2) Brodmann area localization in ex vivo cortical surfaces. B. T. Thomas Yeo, Mert R. Sabuncu, Tom Vercauteren, Nicholas Ayache, Bruce Fischl, Polina Golland |
IEEE Trans. Medical Imaging | 3 |
| 2010 | Learning Task-Optimal Registration Cost Functions for Localizing Cytoarchitecture and Function in the Cerebral CortexabstractImage registration is typically formulated as an optimization problem with multiple tunable, manually set parameters. We present a principled framework for learning thousands of parameters of registration cost functions, such as a spatially-varying tradeoff between the image dissimilarity and regularization terms. Our approach belongs to the classic machine learning framework of model selection by optimization of cross-validation error. This second layer of optimization of cross-validation error over and above registration selects parameters in the registration cost function that result in good registration as measured by the performance of the specific application in a training data set. Much research effort has been devoted to developing generic registration algorithms, which are then specialized to particular imaging modalities, particular imaging targets and particular postregistration analyses. Our framework allows for a systematic adaptation of generic registration cost functions to specific applications by learning the "free" parameters in the cost functions. Here, we consider the application of localizing underlying cytoarchitecture and functional regions in the cerebral cortex by alignment of cortical folding. Most previous work assumes that perfectly registering the macro-anatomy also perfectly aligns the underlying cortical function even though macro-anatomy does not completely predict brain function. In contrast, we learn 1) optimal weights on different cortical folds or 2) optimal cortical folding template in the generic weighted sum of squared differences dissimilarity measure for the localization task. We demonstrate state-of-the-art localization results in both histological and functional magnetic resonance imaging data sets. B. T. Thomas Yeo, Mert R. Sabuncu, Tom Vercauteren, Daphne J. Holt, Katrin Amunts, Karl Zilles, Polina Golland, Bruce Fischl |
IEEE Trans. Medical Imaging | 3 |
| 2009 | Asymmetric Image-Template RegistrationabstractA natural requirement in pairwise image registration is that the resulting deformation is independent of the order of the images. This constraint is typically achieved via a symmetric cost function and has been shown to reduce the effects of local optima. Consequently, symmetric registration has been successfully applied to pairwise image registration as well as the spatial alignment of individual images with a template. However, recent work has shown that the relationship between an image and a template is fundamentally asymmetric. In this paper, we develop a method that reconciles the practical advantages of symmetric registration with the asymmetric nature of image-template registration by adding a simple correction factor to the symmetric cost function. We instantiate our model within a log-domain diffeomorphic registration framework. Our experiments show exploiting the asymmetry in image-template registration improves alignment in the image coordinates. Mert R. Sabuncu, B. T. Thomas Yeo, Koenraad Van Leemput, Tom Vercauteren, Polina Golland |
MICCAI (1) | 4 |
| 2009 | DT-REFinD: Diffusion Tensor Registration With Exact Finite-Strain DifferentialabstractIn this paper, we propose the DT-REFinD algorithm for the diffeomorphic nonlinear registration of diffusion tensor images. Unlike scalar images, deforming tensor images requires choosing both a reorientation strategy and an interpolation scheme. Current diffusion tensor registration algorithms that use full tensor information face difficulties in computing the differential of the tensor reorientation strategy and consequently, these methods often approximate the gradient of the objective function. In the case of the finite-strain (FS) reorientation strategy, we borrow results from the pose estimation literature in computer vision to derive an analytical gradient of the registration objective function. By utilizing the closed-form gradient and the velocity field representation of one parameter subgroups of diffeomorphisms, the resulting registration algorithm is diffeomorphic and fast. We contrast the algorithm with a traditional FS alternative that ignores the reorientation in the gradient computation. We show that the exact gradient leads to significantly better registration at the cost of computation time. Independently of the choice of Euclidean or Log-Euclidean interpolation and sum of squared differences dissimilarity measure, the exact gradient achieves better alignment over an entire spectrum of deformation penalties. Alignment quality is assessed with a battery of metrics including tensor overlap, fractional anisotropy, inverse consistency and closeness to synthetic warps. The improvements persist even when a different reorientation scheme, preservation of principal directions, is used to apply the final deformations. B. T. Thomas Yeo, Tom Vercauteren, Pierre Fillard, Jean-Marc Peyrat, Xavier Pennec, Polina Golland, Nicholas Ayache, Olivier Clatz |
IEEE Trans. Medical Imaging | 2 |
| 2008 | Symmetric Log-Domain Diffeomorphic Registration: A Demons-Based Approach
Tom Vercauteren, Xavier Pennec, Aymeric Perchant, Nicholas Ayache |
MICCAI (1) | 1 |
| 2008 | Spherical Demons: Fast Surface Registration
B. T. Thomas Yeo, Mert R. Sabuncu, Tom Vercauteren, Nicholas Ayache, Bruce Fischl, Polina Golland |
MICCAI (1) | 3 |
| 2007 | Non-parametric Diffeomorphic Image Registration with the Demons Algorithm
Tom Vercauteren, Xavier Pennec, Aymeric Perchant, Nicholas Ayache |
MICCAI (2) | 1 |
| 2006 | Robust mosaicing with correction of motion distortions and tissue deformations for in vivo fibered microscopy
Tom Vercauteren, Aymeric Perchant, Grégoire Malandain, Xavier Pennec, Nicholas Ayache |
Medical Image Anal. | 1 |
| 2006 | Adaptive Optimization of IEEE 802.11 DCF Based on Bayesian Estimation of the Number of Competing TerminalsabstractThe performance of the distributed coordination function (DCF) of the IEEE 802.11 protocol has been shown to heavily depend on the number of terminals accessing the distributed medium. The DCF uses a carrier sense multiple access scheme with collision avoidance (CSMA/CA), where the backoff parameters are fixed and determined by the standard. While those parameters were chosen to provide a good protocol performance, they fail to provide an optimum utilization of the channel in many scenarios. In particular, under heavy load scenarios, the utilization of the medium can drop tenfold. Most of the optimization mechanisms proposed in the literature are based on adapting the DCF backoff parameters to the estimate of the number of competing terminals in the network. However, existing estimation algorithms are either inaccurate or too complex. In this paper, we propose an enhanced version of the IEEE 802.11 DCF that employs an adaptive estimator of the number of competing terminals based on sequential Monte Carlo methods. The algorithm uses a Bayesian approach, optimizing the backoff parameters of the DCF based on the predictive distribution of the number of competing terminals. We show that our algorithm is simple yet highly accurate even at small time scales. We implement our proposed new DCF in the ns-2 simulator and show that it outperforms existing methods. We also show that its accuracy can be used to improve the results of the protocol even when the terminals are not in saturation mode. Moreover, we show that there exists a Nash equilibrium strategy that prevents rogue terminals from changing their parameters for their own benefit, making the algorithm safely applicable in a complete distributed fashion Alberto López Toledo, Tom Vercauteren, Xiaodong Wang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2005 | Optimizing IEEE 802.11 DCF using Bayesian estimators of the network stateabstractThe optimization mechanisms proposed in the literature for the distributed coordination function (DCF) of the IEEE 802.11 protocol are often based on adapting the backoff parameters to the estimate of the number of competing terminals in the network. However, existing estimation algorithms are either inaccurate or too complex. In this paper we propose an enhanced version of the IEEE 802.11 DCF that employs an estimator of the number of competing terminals based on a sequential Monte Carlo (SMC) or a approximate maximum a posteriori (MAP) approach. The algorithm uses a Bayesian framework, optimizing the backoff parameters of the DCF based on the predictive distribution of the number of competing terminals. We show that our algorithm is simple yet highly accurate even at small time scales. We implement our proposed new DCF in the ns-2 simulator and show that it outperforms existing methods. We also show that its accuracy can be used to improve the results of the protocol even when the nodes are not in saturation mode. Alberto López Toledo, Tom Vercauteren, Xiaodong Wang 0001 |
ICASSP (5) | 2 |
| 2005 | Online Bayesian estimation of hidden Markov models with unknown transition matrix and applications to IEEE 802.11 networksabstractWe develop online Bayesian signal processing algorithms to estimate the state and parameters of a hidden Markov model (HMM) with unknown transition matrix. The first online estimator is based on the sequential Monte Carlo (SMC) technique and uses a set of sufficient statistics to carry the information about the transition matrix. A deterministic variant of the SMC estimator is then developed, which is simpler to implement and offers superior performance. Finally, a novel approximate maximum a posteriori (MAP) algorithm is proposed. These algorithms offer a solution to the problem of estimating the number of competing terminals in an IEEE 802.11 network where better performance can be expected if the backoff parameters are adapted to the number of active users. Realistic simulations using the ns-2 network simulator are provided to demonstrate the excellent performance of the proposed estimators. Tom Vercauteren, Alberto López Toledo, Xiaodong Wang 0001 |
ICASSP (4) | 1 |
| 2005 | Mosaicing of Confocal Microscopic In Vivo Soft Tissue Video Sequences
Tom Vercauteren, Aymeric Perchant, Xavier Pennec, Nicholas Ayache |
MICCAI | 1 |
| 2005 | Joint multiple target tracking and classification in collaborative sensor networksabstractWe address the problem of jointly tracking and classifying several targets within a sensor network where false detections are present. In order to meet the requirements inherent to sensor networks such as distributed processing and low-power consumption, a collaborative signal processing algorithm is presented. At any time, for a given tracked target, only one sensor is active. This leader node is focused on a single target but takes into account the possible existence of other targets. It is assumed that the motion model of a given target belongs to one of several classes. This class-target dynamic association is the basis of our classification criterion. We propose an algorithm based on the sequential Monte Carlo (SMC) filtering of jump Markov systems to track the dynamic of the system and make the corresponding estimates. A novel class-based resampling scheme is developed in order to get a robust classification of the targets. Furthermore, an optimal sensor selection scheme based on the maximization of the expected mutual information is integrated naturally within the SMC target tracking framework. Simulation results are presented to illustrate the excellent performance of the proposed multitarget tracking and classification scheme in a collaborative sensor network. Tom Vercauteren, Dong Guo 0003, Xiaodong Wang 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2004 | Joint multiple target tracking and classification in collaborative sensor networksabstractWe address the problem of jointly tracking and classifying several targets within a sensor network where false detections are present. A collaborative signal processing algorithm where multiple targets are dynamically associated with leader nodes is presented. It is assumed that each target belongs to one of several classes and that the class information leads to the motion model of a target. We propose an algorithm based on sequential Monte Carlo (SMC) filtering of jump Markov systems to jointly track the system dynamic and classify the targets. Furthermore, an optimal sensor selection scheme based on the maximization of the expected mutual information is integrated naturally within the SMC tracking framework. Simulation results have illustrated the excellent performance of the proposed scheme. Tom Vercauteren, Dong Guo 0003, Xiaodong Wang 0001 |
ISIT | 1 |