Danail Stoyanov

dblp:53/3543 · DBLP profile ↗
← Back
132ranked-venue papers
10as first author
65since 2021 · last 2026
0000-0002-0980-3227ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 94 · 7 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 9 first-author · 22 since 2021Artificial intelligence and machine learning · 33 · 1 first-author · 16 since 2021Systems, architecture and hardware · 20 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Surgical AI Copilot: Energy-Based Fourier Gradient Low-Rank Adaptation for Surgical LLM Agent Reasoning and Planning
abstract
Image-guided surgery demands adaptive, real-time decision support, yet static AI models struggle with structured task planning and providing interactive guidance. Large language models (LLMs)-powered agents offer a promising solution by enabling dynamic task planning and predictive decision support. Despite recent advances, the absence of surgical agent datasets and robust parameter-efficient fine-tuning techniques limits the development of LLM agents capable of complex intraoperative reasoning. In this paper, we introduce Surgical AI Copilot, an LLM agent for image-guided pituitary surgery, capable of conversation, planning, and task execution in response to queries involving tasks such as MRI tumor segmentation, endoscope anatomy segmentation, overlaying preoperative imaging with intraoperative views, instrument tracking, and surgical visual question answering (VQA). To enable structured agent planning, we develop the PitAgent dataset, a surgical context-aware planning dataset covering surgical tasks like workflow analysis, instrument localization, anatomical segmentation, and query-based reasoning. Additionally, we propose DEFT-GaLore, a Deterministic Energy-based Fourier Transform (DEFT) gradient projection technique for efficient low-rank adaptation of recent LLMs (e.g., LLaMA 3.2, Qwen 2.5), enabling their use as surgical agent planners. We extensively validate our agent's performance and the proposed adaptation technique against other state-of-the-art low-rank adaptation methods on agent planning and prompt generation tasks, including a zero-shot surgical VQA benchmark, demonstrating the significant potential for truly efficient and scalable surgical LLM agents in real-time operative settings.
Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Danail Stoyanov, Hani J. Marcus, Linzhe Jiang, Matthew J. Clarkson, Mobarak I. Hoque
AAAI5
2026 SurgflowNet: Leveraging unannotated video for consistent endoscopic pituitary surgery workflow recognition
abstract
-score and 13.4% in Edit Score over the SOTA, SurgflowNetdemonstrates a significant improvement in workflow recognition for endoscopic pituitary surgery.
Anjana Wijekoon, Adrito Das, Zhehua Mao, Danyal Z. Khan, John G. Hanrahan, Danail Stoyanov, Hani J. Marcus, Sophia Bano
Artif. Intell. Medicine6
2026 Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
abstract
Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding.
Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm
Medical Image Anal.10
2026 TDN-PPO: An Automatic Control Framework of a Surgical Robot for Posterior Segment Ophthalmic Surgery
abstract
In ophthalmic surgery, particularly the procedures involving the posterior segment, clinicians face significant challenges in maintaining precise control of handheld instruments to peel fragile membranes without damaging the surrounding healthy fundus tissue, even for seasoned clinicians employing specialised ophthalmic surgical robots. The implementation of autonomous control in robot-assisted surgical systems holds promise for overcoming these obstacles and simplifying intricate surgical tasks. This paper introduces an autonomous control framework, integrating a Tip Detection Network (TDN) with a Proximal Policy Optimization (PPO) network, designed to autonomously navigate the Tip of the Surgical Instrument (ToSI) towards the intended lesion site in a real-world scenario. Results indicate that the accuracy of the TDN module in detecting the ToSI position in images of varying sizes can be reliably maintained within a 4.6-pixel range. The autonomous control deviation for the PPO module ranges between [0.6585μm, 7.995μm], with an average discrepancy of 5.118μm. The communication frequency across the modules is maintained at 35.7 Hz. In physical environments, the TDN-PPO framework adeptly navigates the ToSI to autonomously and precisely converge on the preset Target Lesion (TL) macular hole, maintaining a tip-to-target distance error within a margin of 2.094 pixels (38μm). Throughout the autonomous navigation phase, the maximal contact force between the ToSI and the TL is capped at 41.1 mN, aligning with the upper threshold for contact force between the ToSI and tissue prescribed in clinical surgical settings.
Ning Wang 0042, Sophia Bano, Danail Stoyanov, Ziting Liang, Agostino Stilli
IEEE Trans Autom. Sci. Eng.4
2026 PitVQA++: Vector Matrix-Low-Rank Adaptation for Open-Ended Visual Question Answering in Pituitary Surgery
abstract
Vision-Language Models (VLMs) in visual question answering (VQA) offer a unique opportunity to enhance intra-operative decision-making, promote intuitive interactions, and significantly advance surgical education. However, the development of VLMs for surgical VQA is challenging due to limited datasets and the risk of overfitting and catastrophic forgetting during full fine-tuning of pretrained weights. While parameter-efficient techniques like Low-Rank Adaptation (LoRA) and Matrix of Rank Adaptation (MoRA) address adaptation challenges, their uniform parameter distribution overlooks the feature hierarchy in deep networks, where earlier layers, that learn general features, require more parameters than later ones. This work introduces PitVQA++ with an Open-ended PitVQA dataset and vector matrix-low-rank adaptation (Vector-MoLoRA), an innovative VLM fine-tuning approach for adapting GPT-2 to pituitary surgery. Open-Ended PitVQA comprises 109,173 frames from 25 procedural videos with 795,270 question-answer sentence pairs, covering key surgical elements such as phase and step recognition, context understanding, tool detection, localization, and interactions recognition. Vector-MoLoRA incorporates the principles of LoRA and MoRA to develop a matrix-low-rank adaptation strategy that employs rank vectors to allocate more parameters to earlier layers, gradually reducing them in the later layers. Our approach, validated on the Open-Ended PitVQA and EndoVis18-VQA datasets, effectively mitigates catastrophic forgetting while significantly enhancing performance over recent baselines. Performance-rejection analysis further highlights Vector-MoLoRA's enhanced reliability and trustworthiness in handling uncertain predictions. Our source code and dataset is available at https://github.com/HRL-Mike/PitVQA-Plus.
Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque
IEEE Trans. Medical Imaging5
2026 AnatoDiff: Synthesizing Anatomically Truthful Radiographs With Limited Training Images
abstract
Rapid advancements in diffusion models have enabled synthesis of realistic and anonymized imagery in radiography. However, due to their complexity, these models typically require large training volumes, often exceeding 10,000 images. Pre-training on natural images can partly mitigate this issue, but often fails to generate anatomically accurate shapes due to the significant domain gap. This prohibits applications in specialized medical conditions with limited data. We propose AnatoDiff, a diffusion model synthesizing high-quality X-Ray images with accurate anatomical shapes using only 500 to 1,000 training samples. AnatoDiff incorporates a Shape Prototype Module and Anatomical Fidelity loss, allowing for smaller training volumes through targeted supervision. We extensively validate AnatoDiff across three open-source datasets from distinct anatomical regions: Neonatal Abdomen (1,000 images); Adult Chest (500 images); and Humerus (500 images). Results demonstrate significant benefits, with an average improvement of 14.9% in Fréchet Inception Distance, 9.7% in Improved Precision, and 2.3% in Improved Recall compared to state-of-the-art (SOTA) few-shot and data-limited natural image synthesis methods. Unlike other models, AnatoDiff consistently generates anatomically correct images with accurate shapes. Additionally, a ResNet-50 classifier trained on AnatoDiff-generated images shows a 2.1% to 5.3% increase in F1-score, compared to being trained on SOTA diffusion images, across 500 to 10,000 samples. A survey with 10 medical professionals reveals that images generated by AnatoDiff are challenging to distinguish from real ones, with a Matthews correlation coefficient of 0.277 and Fleiss' Kappa of 0.126, highlighting the effectiveness of AnatoDiff in generating high-quality, anatomically accurate radiographs. Our code is available at https://github.com/KawaiYung/AnatoDiff.
Ka-Wai Yung, Jayaram Sivaraj, Lodovico di Giura, Simon Eaton, Paolo De Coppi, Danail Stoyanov, Stavros Loukogeorgakis, Evangelos B. Mazomenos
IEEE Trans. Medical Imaging6
2025 Tracking Everything in Robotic-Assisted Surgery
abstract
Accurate tracking of tissues and instruments in videos is crucial for Robotic-Assisted Minimally Invasive Surgery (RAMIS), as it enables the robot to comprehend the surgical scene with precise locations and interactions of tissues and tools. Traditional keypoint-based sparse tracking is limited by featured points, while flow-based dense two-view matching suffers from long-term drifts. Recently, the Tracking Any Point (TAP) algorithm was proposed to overcome these limitations and achieve dense accurate long-term tracking. However, its efficacy in surgical scenarios remains untested, largely due to the lack of a comprehensive surgical tracking dataset for evaluation. To address this gap, we introduce a new annotated surgical tracking dataset for benchmarking tracking methods for surgical scenarios, comprising real-world surgical videos with complex tissue and instrument motions. We extensively evaluate state-of-the-art (SOTA) TAP-based algorithms on this dataset and reveal their limitations in challenging surgical scenarios, including fast instrument motion, severe occlusions, and motion blur, etc. Furthermore, we propose a new tracking method, namely SurgMotion, to solve the challenges and further improve the tracking performance. Our proposed method outperforms most TAP-based algorithms in surgical instruments tracking, and especially demonstrates significant improvements over baselines in challenging medical videos. Our code and dataset are available at https://github.com/zhanbh1019/SurgicalMotion.
Bohan Zhan, Yi Fang 0006, Francisco Vasconcelos 0001, Danail Stoyanov, Daniel S. Elson, Baoru Huang
ICRA6
2025 SurgicalGS: Dynamic 3D Gaussian Splatting for Accurate Robotic-Assisted Surgical Scene Reconstruction
Jialei Chen 0007, Mobarak I. Hoque, Francisco Vasconcelos 0001, Danail Stoyanov, Daniel S. Elson, Baoru Huang
MICCAI (11)5
2025 Endo-FASt3r: Endoscopic Foundation Model Adaptation for Structure from Motion
abstract
Accurate depth and camera pose estimation is essential for achieving high-quality 3D visualisations in robotic-assisted surgery. Despite recent advancements in foundation model adaptation to monocular depth estimation of endoscopic scenes via self-supervised learning (SSL), no prior work has explored their use for pose estimation. These methods rely on low rank-based adaptation approaches, which constrain model updates to a low-rank space. We propose Endo-FASt3r, the first monocular SSL depth and pose estimation framework that uses foundation models for both tasks. We extend the Reloc3r relative pose estimation foundation model by designing Reloc3rX, introducing modifications necessary for convergence in SSL. We also present DoMoRA, a novel adaptation technique that enables higher-rank updates and faster convergence. Experiments on the SCARED dataset show that Endo-FASt3r achieves a substantial $$10\%$$ improvement in pose estimation and a $$2\%$$ improvement in depth estimation over prior work. Similar performance gains on the Hamlyn and StereoMIS datasets reinforce the generalisability of Endo-FASt3r across different datasets. Our code is available at: https://github.com/Mona-ShZeinoddin/Endo_FASt3r.git .
Mona Sheikh Zeinoddin, Mobarak I. Hoque, Zafer Tandogdu, Greg Shaw, Matthew J. Clarkson, Evangelos B. Mazomenos, Danail Stoyanov
MICCAI (11)7
2025 3D Acetabular Surface Reconstruction from 2D Pre-operative X-Ray Images Using SRVF Elastic Registration and Deformation Graph
abstract
Accurate and reliable selection of the appropriate acetabular cup size is crucial for restoring joint biomechanics in total hip arthroplasty (THA). This paper proposes a novel framework integrating square-root velocity function (SRVF)-based elastic shape registration technique with an embedded deformation (ED) graph approach to reconstruct the 3D articular surface of the acetabulum by fusing multiple views of 2D pre-operative pelvic X-ray images and a hemispherical surface model. The SRVF-based elastic registration establishes 2D-3D correspondences between the parametric hemispherical model and X-ray images, and the ED framework incorporates the SRVF-derived correspondences as constraints to optimize the 3D acetabular surface reconstruction using nonlinear least-squares optimization. Validations using both simulation and real patient datasets are performed to demonstrate the robustness and the potential clinical value of the proposed algorithm. The reconstruction result can assist surgeons in selecting the correct acetabular cup on the first attempt in THA, minimising the need for revision surgery. Code and data are available at: https://github.com/zsustc/3D-ASR .
Shuai Zhang 0029, Sujith Konandetails, Danail Stoyanov, Evangelos B. Mazomenos
MICCAI (16)5
2025 PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgery
abstract
The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686.
Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano
Medical Image Anal.31
2025 Vision-Based Automatic Control of a Surgical Robot for Posterior Segment Ophthalmic Surgery
abstract
In ophthalmic surgery, especially in posterior segment procedures, clinicians face significant challenges, like the inherent tremor of the surgeon’s arm, restricted visibility, and heavy reliance on the surgeon’s skills for precise control of hand-held tools during micro-surgical movements. Automatic control of robotic-assisted ophthalmic surgical systems has the potential to overcome these challenges, simplifying complex surgical procedures. This paper proposes a novel image-guided automatic control method for an Ophthalmic micro-Surgical Robot (OmSR), specifically designed for posterior segment eye surgery. The method relies on forceps shadow tracking. The paper introduces a tip detection network (Net-SR), which accurately calculates the coordinates of the Tips of Surgical Forceps (ToSF) and Tips of Shadow (ToS) to enable automatic navigation. Additionally, through the Non-Uniform Rational B-Spline (NURBS) curve interpolation and speed look-ahead algorithm, dense and time-continuous data points are obtained to improve control accuracy and smoothness. The accuracy of the Net-SR network and motion of the ToSF, and the effectiveness of the proposed automatic controller are experimentally evaluated. Results demonstrate a significant 98.21% improvement in the Net-SR network accuracy over the normal keypoint detection network. The use of the speed look-ahead algorithm leads to a notable 41.7% improvement in optimal speed, and the ToSF successfully reaches the target lesion with vision-based navigation and no overscale motion. Note to Practitioners—The practical problem that motivated this research is the need for safer and more efficient surgical procedures, focusing on minimizing the risk of fundus tissue damage associated with intraoperative surgical instruments. To overcome challenges related to handheld and tele-operated control, we explore automatic control as a promising solution. In this paper, the tip of the instrument can consistently and accurately reach the target lesion with high precision and no overscale motion, allowing for deskilling of complex and repetitive tasks. This capability holds potential for the clinical needle insertion operation and membrane peeling operation. The proposed control methods can also be extended to other surgical procedures.
Ning Wang 0042, Sophia Bano, Danail Stoyanov, Agostino Stilli
IEEE Trans Autom. Sci. Eng.4
2024 Spatial-Temporal NAS for Fast Surgical Segmentation
Felix J. S. Bragman, Ricardo Sanchez-Matilla, Imanol Luengo, Danail Stoyanov
BMVC5
2024 Miniaturisation and Evaluation of the SoftSCREEN System in Colon Phantoms
abstract
Screening of the lower gastrointestinal (GI) tract is of paramount importance for the early detection of precancerous lesions in the intestine, with an impact on reducing the high death rate of patients affected by cancer worldwide. Colonoscopy, i.e. standard procedure for screening the colon, is effective in reducing the incidence of colorectal cancer worldwide, nonetheless, this procedure remains an invasive method of screening, that typically causes discomfort and requires sedation for the patient. The SoftSCREEN system, a tethered robotic capsule designed for colonoscopy, aims to enable minimally invasive diagnosis of intestinal diseases through its innovative design that incorporates elastic tracks for locomotion and inflatable toroidal chambers for adaptable geometry to match the local lumen of the GI tract. After demonstrating the viability of the proposed design in a large-scale proof of concept in our previous work, the authors present here a miniaturised version of the SoftSCREEN system. We assess its performance in multiple phantom tests and evaluate the effect of pressure regulation on its locomotion. The conducted extensive tests demonstrate the capability of the soft robot to move inside intricate passages, capture internal images, and adjust its geometry to optimise traction. The results underscore the potential of the proposed design, offering promising advancements in the development of a robotic platform for efficient front-wheel locomotion and accurate intestinal screening.
Vanni Consumi, Neri Niccolò Dei, Gastone Ciuti, Danail Stoyanov, Agostino Stilli
IROS4
2024 Harnessing Symmetry Breaking in Soft Robotics: A Novel Approach for Underactuated Fingers
abstract
Soft robotics, an emerging domain in modern robotics, introduces innovative possibilities alongside challenges in controllability, particularly with multi-degree inflatable actuators. We present a novel manipulation method using underactuated soft fingers that addresses these challenges by harnessing symmetry breaking. Central to our approach is the mechanism of self-organization within a ring actuator equipped with five fingers. Typically considered a drawback, we exploit the actuator’s buckling behavior to facilitate in-hand manipulation. This strategic utilization enables object motion in both clockwise and counterclockwise directions via system perturbations and adjustments in frequency and duty cycle parameters. Employing the self-organizing properties of our actuator, our method is empirically validated through simulations and real actuator experiments, demonstrating the system’s ability in manipulating objects by leveraging the inherent flexibility and morphological advantages. The design enables two degrees of freedom with minimal input, allowing objects to rotate due to the actuator’s self-organizing actions. This simplification of control mechanisms is essential for soft robotics manipulation. Our findings indicate that control systems in soft robotics can be significantly simplified, harnessing the adaptable behavior inherent in its morphology.
Ryman Hashem, Toby Howison, Agostino Stilli, Danail Stoyanov, Weiliang Xu 0001, Fumiya Iida
IROS4
2024 An MR Safe Double-Arch Needle Insertion Robot with Scissor-Folding Mechanism for Abdominal Percutaneous Interventions*
abstract
Tumors affecting abdominal organs rank among the deadliest malignancies. In this context, Magnetic Resonance Imaging (MRI) serves as an effective diagnostic tool with a strong potential to support image-guided minimally invasive interventions for treating these tumours, offering an ionizing-radiation-free medical modality. MRI provides exceptional soft tissue contrast and multi-angle imaging, enabling accurate intraoperative localisation of target tumours within these vital organs. Nevertheless, MRI-guided minimally invasive interventions still encounter significant challenges due to the strong magnetic field environment and the narrow and deep bore of MRI machines. This paper proposes a novel MR safe 5-degrees-of-freedom (DoFs) parallel table-mounted double-arch needle insertion robot with a scissor-folding mechanism (SFM) for abdominal interventions. The proposed robot is designed to fit a standard 70-cm MRI bore. Initial evaluation experiments indicate mean errors of 3.14 mm for the proposed robotic arch and 2.23 mm for the full needle insertion robot, respectively. Additionally, preliminary testing of the system in an MRI environment resulted in unaltered MRI imaging output, with negligible artefacts associated with the presence of the robot within the bore.
Ziting Liang, Chuang Lu, Haoqian Yang, Ryman Hashem, Mohamed E. M. K. Abdelaziz, Lukas Lindenroth, Steven Bandula, Danail Stoyanov, Agostino Stilli
IROS8
2024 Local Path Planning among Pushable Objects based on Reinforcement Learning
abstract
In this paper, we introduce a method to tackle the problem of robot local path planning among pushable objects –an open problem in robotics. In particular, we simultaneously train multiple agents in a physics-based simulation environment, utilizing an Advantage Actor-Critic algorithm coupled with a deep neural network. The developed online policy enables these agents to push obstacles in ways that are not limited to axial alignments, adapt to unforeseen changes in obstacle dynamics instantaneously, and effectively tackle local path planning in confined areas. We tested the method in various simulated environments to prove the adaptation effectiveness to various unseen scenarios in unfamiliar settings. Moreover, we have successfully applied this policy on an actual quadruped robot, confirming its capability to handle the unpredictability and noise associated with real-world sensors and the inherent uncertainties in unexplored object-pushing tasks.
Linghong Yao, Valerio Modugno, Andromachi Maria Delfaki, Yuanchang Liu, Danail Stoyanov, Dimitrios Kanoulas
IROS5
2024 HUP-3D: A 3D Multi-view Synthetic Dataset for Assisted-Egocentric Hand-Ultrasound-Probe Pose Estimation
abstract
We present HUP-3D, a 3D multiview multimodal synthetic dataset for hand ultrasound (US) probe pose estimation in the context of obstetric ultrasound. Egocentric markerless 3D joint pose estimation has potential applications in mixed reality medical education. The ability to understand hand and probe movements opens the door to tailored guidance and mentoring applications. Our dataset consists of over 31k sets of RGB, depth, and segmentation mask frames, including pose-related reference data, with an emphasis on image diversity and complexity. Adopting a camera viewpoint-based sphere concept allows us to capture a variety of views and generate multiple hand grasps poses using a pre-trained network. Additionally, our approach includes a software-based image rendering concept, enhancing diversity with various hand and arm textures, lighting conditions, and background images. We validated our proposed dataset with state-of-the-art learning models and we obtained the lowest hand-object keypoint errors. The supplementary material details the parameters for sphere-based camera view angles and the grasp generation and rendering pipeline configuration. The source code for our grasp generation and rendering pipeline, along with the dataset, is publicly available at https://manuelbirlo.github.io/HUP-3D/ .
Manuel Birlo, Razvan Caramalau, Philip J. Edwards, Brian Dromey, Matthew J. Clarkson, Danail Stoyanov
MICCAI (1)6
2024 Gaussian Pancakes: Geometrically-Regularized 3D Gaussian Splatting for Realistic Endoscopic Reconstruction
Sierra Bonilla, Shuai Zhang 0029, Dimitris Psychogyios, Danail Stoyanov, Francisco Vasconcelos 0001, Sophia Bano
MICCAI (6)4
2024 PitVQA: Image-Grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
Runlong He, Mengya Xu, Adrito Das, Danyal Z. Khan, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarakol Islam
MICCAI (6)7
2024 Region-Specific Retrieval Augmentation for Longitudinal Visual Question Answering: A Mix-and-Match Paradigm
abstract
Visual Question Answering (VQA) has advanced in recent years, inspiring adaptations to radiology for medical diagnosis. Longitudinal VQA, which requires an understanding of changes in images over time, can further support patient monitoring and treatment decision-making. This work introduces RegioMix, a retrieval augmented paradigm for longitudinal VQA, formulating a novel approach that generates retrieval objects through a mix-and-match technique, utilizing different regions from various retrieved images. Furthermore, this process generates a pseudo-difference description based on the retrieved pair, by leveraging available reports from each retrieved region. To align such statements to both the posed question and input image pair, we introduce a Dual Alignment module. Experiments on the MIMIC-Diff-VQA X-ray dataset demonstrate our method’s superiority, outperforming the state-of-the-art by 77.7 in CIDEr score and $$8.3\%$$ in BLEU-4, while relying solely on the training dataset for retrieval, showcasing the effectiveness of our approach. Code is available at https://github.com/KawaiYung/RegioMix .
Ka-Wai Yung, Jayaram Sivaraj, Danail Stoyanov, Stavros Loukogeorgakis, Evangelos B. Mazomenos
MICCAI (5)3
2024 Placental vessel segmentation and registration in fetoscopy: Literature review and MICCAI FetReg2021 challenge findings
abstract
Fetoscopy laser photocoagulation is a widely adopted procedure for treating Twin-to-Twin Transfusion Syndrome (TTTS). The procedure involves photocoagulation pathological anastomoses to restore a physiological blood exchange among twins. The procedure is particularly challenging, from the surgeon's side, due to the limited field of view, poor manoeuvrability of the fetoscope, poor visibility due to amniotic fluid turbidity, and variability in illumination. These challenges may lead to increased surgery time and incomplete ablation of pathological anastomoses, resulting in persistent TTTS. Computer-assisted intervention (CAI) can provide TTTS surgeons with decision support and context awareness by identifying key structures in the scene and expanding the fetoscopic field of view through video mosaicking. Research in this domain has been hampered by the lack of high-quality data to design, develop and test CAI algorithms. Through the Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge, which was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge, we released the first large-scale multi-center TTTS dataset for the development of generalized and robust semantic segmentation and video mosaicking algorithms with a focus on creating drift-free mosaics from long duration fetoscopy videos. For this challenge, we released a dataset of 2060 images, pixel-annotated for vessels, tool, fetus and background classes, from 18 in-vivo TTTS fetoscopy procedures and 18 short video clips of an average length of 411 frames for developing placental scene segmentation and frame registration for mosaicking techniques. Seven teams participated in this challenge and their model performance was assessed on an unseen test dataset of 658 pixel-annotated images from 6 fetoscopic procedures and 6 short clips. For the segmentation task, overall baseline performed was the top performing (aggregated mIoU of 0.6763) and was the best on the vessel class (mIoU of 0.5817) while team RREB was the best on the tool (mIoU of 0.6335) and fetus (mIoU of 0.5178) classes. For the registration task, overall the baseline performed better than team SANO with an overall mean 5-frame SSIM of 0.9348. Qualitatively, it was observed that team SANO performed better in planar scenarios, while baseline was better in non-planner scenarios. The detailed analysis showed that no single team outperformed on all 6 test fetoscopic videos. The challenge provided an opportunity to create generalized solutions for fetoscopic scene understanding and mosaicking. In this paper, we present the findings of the FetReg2021 challenge, alongside reporting a detailed literature review for CAI in TTTS fetoscopy. Through this challenge, its analysis and the release of multi-center fetoscopic data, we provide a benchmark for future research in this field.
Sophia Bano, Alessandro Casella, Francisco Vasconcelos 0001, Abdul Qayyum 0002, Abdessalam Benzinou, Moona Mazher, Fabrice Mériaudeau, Chiara Lena, Ilaria A. Cintorrino, Gaia Romana De Paolis, Jessica Biagioli, Daria Grechishnikova, Jing Jiao, Bizhe Bai, Yanyan Qiao, Binod Bhattarai, Rebati Raman Gaire, Ronast Subedi, Eduard Vazquez, Szymon Plotka, Aneta Lisowska, Arkadiusz Sitek, George Attilakos, Ruwan Wimalasundera, Anna L. David, Dario Paladini, Jan Deprest, Elena De Momi, Leonardo S. Mattos, Sara Moccia, Danail Stoyanov
Medical Image Anal.31
2024 Guided image generation for improved surgical image segmentation
abstract
The lack of large datasets and high-quality annotated data often limits the development of accurate and robust machine-learning models within the medical and surgical domains. In the machine learning community, generative models have recently demonstrated that it is possible to produce novel and diverse synthetic images that closely resemble reality while controlling their content with various types of annotations. However, generative models have not been yet fully explored in the surgical domain, partially due to the lack of large datasets and due to specific challenges present in the surgical domain such as the large anatomical diversity. We propose Surgery-GAN, a novel generative model that produces synthetic images from segmentation maps. Our architecture produces surgical images with improved quality when compared to early generative models thanks to the combination of channel- and pixel-level normalization layers that boost image quality while granting adherence to the input segmentation map. While state-of-the-art generative models often generate overfitted images, lacking diversity, or containing unrealistic artefacts such as cartooning; experiments demonstrate that Surgery-GAN is able to generate novel, realistic, and diverse surgical images in three different surgical datasets: cholecystectomy, partial nephrectomy, and radical prostatectomy. In addition, we investigate whether the use of synthetic images together with real ones can be used to improve the performance of other machine-learning models. Specifically, we use Surgery-GAN to generate large synthetic datasets which we then use to train five different segmentation models. Results demonstrate that using our synthetic images always improves the mean segmentation performance with respect to only using real images. For example, when considering radical prostatectomy, we can boost the mean segmentation performance by up to 5.43%. More interestingly, experimental results indicate that the performance improvement is larger in the set of classes that are under-represented in the training sets, where the performance boost of specific classes reaches up to 61.6%.
Emanuele Colleoni, Ricardo Sanchez-Matilla, Imanol Luengo, Danail Stoyanov
Medical Image Anal.4
2024 SimCol3D - 3D reconstruction during colonoscopy challenge
abstract
Colorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question.
Anita Rau, Sophia Bano, Yueming Jin, Pablo Azagra, Javier Morlana, Rawen Kader, Edward Sanderson, Bogdan J. Matuszewski, Erez Posner, Netanel Frank, Varshini Elangovan, Sista Raviteja, Zhengwen Li, Jiquan Liu, Seenivasan Lalithkumar, Mobarakol Islam, Hongliang Ren 0001, Laurence B. Lovat, J. M. M. Montiel, Danail Stoyanov
Medical Image Anal.22
2024 MonoLoT: Self-Supervised Monocular Depth Estimation in Low-Texture Scenes for Automatic Robotic Endoscopy
abstract
The self-supervised monocular depth estimation framework is well-suited for medical images that lack ground-truth depth, such as those from digestive endoscopes, facilitating navigation and 3D reconstruction in the gastrointestinal tract. However, this framework faces several limitations, including poor performance in low-texture environments, limited generalisation to real-world datasets, and unclear applicability in downstream tasks like visual servoing. To tackle these challenges, we propose MonoLoT, a self-supervised monocular depth estimation framework featuring two key innovations: point matching loss and batch image shuffle. Extensive ablation studies on two publicly available datasets, namely C3VD and SimCol, have shown that methods enabled by MonoLoT achieve substantial improvements, with accuracies of 0.944 on C3VD and 0.959 on SimCol, surpassing both depth-supervised and self-supervised baselines on C3VD. Qualitative evaluations on real-world endoscopic data underscore the generalisation capabilities of our methods, outperforming both depth-supervised and self-supervised baselines. To demonstrate the feasibility of using monocular depth estimation for visual servoing, we have successfully integrated our method into a proof-of-concept robotic platform, enabling real-time automatic intervention and control in digestive endoscopy. In summary, our method represents a significant advancement in monocular depth estimation for digestive endoscopy, overcoming key challenges and opening promising avenues for medical applications.
Qi He 0008, Sophia Bano, Danail Stoyanov, Siyang Zuo
IEEE J. Biomed. Health Informatics4
2024 DIPO: Differentiable Parallel Operation Blocks for Surgical Neural Architecture Search
abstract
Deep learning has been used across a large number of computer vision tasks, however designing the network architectures for each task is time consuming. Neural Architecture Search (NAS) promises to automatically build neural networks, optimised for the given task and dataset. However, most NAS methods are constrained to a specific macro-architecture design which makes it hard to apply to different tasks (classification, detection, segmentation). Following the work in Differentiable NAS (DNAS), we present a simple and efficient NAS method, Differentiable Parallel Operation (DIPO), that constructs a local search space in the form of a DIPO block, and can easily be applied to any convolutional network by injecting it in-place of the convolutions. The DIPO block's internal architecture and parameters are automatically optimised end-to-end for each task. We demonstrate the flexibility of our approach by applying DIPO to 4 model architectures (U-Net, HRNET, KAPAO and YOLOX) across different surgical tasks (surgical scene segmentation, surgical instrument detection, and surgical instrument pose estimation) and evaluated across 5 datasets. Results show significant improvements in surgical scene segmentation (+10.5% in CholecSeg8K, +13.2% in CaDIS), instrument detection (+1.5% in ROBUST-MIS, +5.3% in RoboKP), and instrument pose estimation (+9.8% in RoboKP).
Ricardo Sanchez-Matilla, Danail Stoyanov, Imanol Luengo
IEEE J. Biomed. Health Informatics3
2024 FedDP: Dual Personalization in Federated Medical Image Segmentation
abstract
Personalized federated learning (PFL) addresses the data heterogeneity challenge faced by general federated learning (GFL). Rather than learning a single global model, with PFL a collection of models are adapted to the unique feature distribution of each site. However, current PFL methods rarely consider self-attention networks which can handle data heterogeneity by long-range dependency modeling and they do not utilize prediction inconsistencies in local models as an indicator of site uniqueness. In this paper, we propose FedDP, a novel fed erated learning scheme with d ual p ersonalization, which improves model personalization from both feature and prediction aspects to boost image segmentation results. We leverage long-range dependencies by designing a local query (LQ) that decouples the query embedding layer out of each local model, whose parameters are trained privately to better adapt to the respective feature distribution of the site. We then propose inconsistency-guided calibration (IGC), which exploits the inter-site prediction inconsistencies to accommodate the model learning concentration. By encouraging a model to penalize pixels with larger inconsistencies, we better tailor prediction-level patterns to each local site. Experimentally, we compare FedDP with the state-of-the-art PFL methods on two popular medical image segmentation tasks with different modalities, where our results consistently outperform others on both tasks. Our code and models are available at https://github.com/jcwang123/PFL-Seg-Trans.
Jiacheng Wang 0002, Yueming Jin, Danail Stoyanov, Liansheng Wang 0002
IEEE Trans. Medical Imaging3
2023 ASPnet: Action Segmentation with Shared-Private Representation of Multiple Data Sources
abstract
Most state-of-the-art methods for action segmentation are based on single input modalities or naïve fusion of multiple data sources. However, effective fusion of complementary information can potentially strengthen segmentation models and make them more robust to sensor noise and more accurate with smaller training datasets. In order to improve multimodal representation learning for action segmentation, we propose to disentangle hidden features of a multi-stream segmentation model into modality-shared components, containing common information across data sources, and private components; we then use an attention bottleneck to capture long-range temporal dependencies in the data while preserving disentanglement in consecutive processing layers. Evaluation on 50salads, Breakfast and RARP45 datasets shows that our multimodal approach outperforms different data fusion baselines on both multiview and multimodal data sources, obtaining competitive or better results compared with the state-of-the-art. Our model is also more robust to additive sensor noise and can achieve performance on par with strong video baselines even with less training data.
Beatrice van Amsterdam, Abdolrahim Kadkhodamohammadi, Imanol Luengo, Danail Stoyanov
CVPR4
2023 CycleSTTN: A Learning-Based Temporal Model for Specular Augmentation in Endoscopy
abstract
Feature detection and matching is a computer vision problem that underpins different computer assisted techniques in endoscopy, including anatomy and lesion recognition, camera motion estimation, and 3D reconstruction. This problem is made extremely challenging due to the abundant presence of specular reflections. Most of the solutions proposed in the literature are based on filtering or masking out these regions as an additional processing step. There has been little investigation into explicitly learning robustness to such artefacts with single-step end-to-end training. In this paper, we propose an augmentation technique (CycleSTTN) that adds temporally consistent and realistic specularities to endoscopic videos. Such videos can act as ground truth data with known texture occluded behind the added specularities. We demonstrate that our image generation technique produces better results than a standard CycleGAN model. Additionally, we leverage this data augmentation to re-train a deep-learning based feature extractor (SuperPoint) and show that it improves. CycleSTTN code is made available here .
Rema Daher, Oscar León Barbed, Ana Cristina Murillo, Francisco Vasconcelos 0001, Danail Stoyanov
MICCAI (10)5
2023 A Multi-task Network for Anatomy Identification in Endoscopic Pituitary Surgery
Adrito Das, Danyal Z. Khan, Simon C. Williams, John G. Hanrahan, Anouk Borg, Neil L. Dorward, Sophia Bano, Hani J. Marcus, Danail Stoyanov
MICCAI (9)9
2023 Realistic Endoscopic Illumination Modeling for NeRF-Based Data Generation
abstract
Expanding training and evaluation data is a major step towards building and deploying reliable localization and 3D reconstruction techniques during colonoscopy screenings. However, training and evaluating pose and depth models in colonoscopy is hard as available datasets are limited in size. This paper proposes a method for generating new pose and depth datasets by fitting NeRFs in already available colonoscopy datasets. Given a set of images, their associated depth maps and pose information, we train a novel light source location-conditioned NeRF to encapsulate the 3D and color information of a colon sequence. Then, we leverage the trained networks to render images from previously unobserved camera poses and simulate different camera systems, effectively expanding the source dataset. Our experiments show that our model is able to generate RGB images and depth maps of a colonoscopy sequence from previously unobserved poses with high accuracy. Code and trained networks can be accessed at https://github.com/surgical-vision/REIM-NeRF .
Dimitris Psychogyios, Francisco Vasconcelos 0001, Danail Stoyanov
MICCAI (9)3
2023 Regressing Simulation to Real: Unsupervised Domain Adaptation for Automated Quality Assessment in Transoesophageal Echocardiography
Jialang Xu, Yueming Jin, Bruce Martin, Andrew P. T. Smith, Susan Wright, Danail Stoyanov, Evangelos B. Mazomenos
MICCAI (9)6
2023 3D Generative Model Latent Disentanglement via Local Eigenprojection
abstract
Designing realistic digital humans is extremely complex. Most data-driven generative models used to simplify the creation of their underlying geometric shape do not offer control over the generation of local shape attributes. In this paper, we overcome this limitation by introducing a novel loss function grounded in spectral geometry and applicable to different neural-network-based generative models of 3D head and body meshes. Encouraging the latent variables of mesh variational autoencoders (VAEs) or generative adversarial networks (GANs) to follow the local eigenprojections of identity attributes, we improve latent disentanglement and properly decouple the attribute creation. Experimental results show that our local eigenprojection disentangled (LED) models not only offer improved disentanglement with respect to the state-of-the-art, but also maintain good generation capabilities with training times comparable to the vanilla implementations of the models. Our code and pre-trained models are available at github.com/simofoti/LocalEigenprojDisentangled.
Simone Foti, Bongjin Koo, Danail Stoyanov, Matthew J. Clarkson
Comput. Graph. Forum3
2023 Histogram of Oriented Gradients meet deep learning: A novel multi-task deep network for 2D surgical image semantic segmentation
abstract
We present our novel deep multi-task learning method for medical image segmentation. Existing multi-task methods demand ground truth annotations for both the primary and auxiliary tasks. Contrary to it, we propose to generate the pseudo-labels of an auxiliary task in an unsupervised manner. To generate the pseudo-labels, we leverage Histogram of Oriented Gradients (HOGs), one of the most widely used and powerful hand-crafted features for detection. Together with the ground truth semantic segmentation masks for the primary task and pseudo-labels for the auxiliary task, we learn the parameters of the deep network to minimize the loss of both the primary task and the auxiliary task jointly. We employed our method on two powerful and widely used semantic segmentation networks: UNet and U2Net to train in a multi-task setup. To validate our hypothesis, we performed experiments on two different medical image segmentation data sets. From the extensive quantitative and qualitative results, we observe that our method consistently improves the performance compared to the counter-part method. Moreover, our method is the winner of FetReg Endovis Sub-challenge on Semantic Segmentation organised in conjunction with MICCAI 2021. Code and implementation details are available at:https://github.com/thetna/medical_image_segmentation.
Binod Bhattarai, Ronast Subedi, Rebati Raman Gaire, Eduard Vazquez, Danail Stoyanov
Medical Image Anal.5
2023 A Temporal Learning Approach to Inpainting Endoscopic Specularities and Its Effect on Image Correspondence
abstract
Video streams are utilised to guide minimally-invasive surgery and diagnosis in a wide range of procedures, and many computer-assisted techniques have been developed to automatically analyse them. These approaches can provide additional information to the surgeon such as lesion detection, instrument navigation, or anatomy 3D shape modelling. However, the necessary image features to recognise these patterns are not always reliably detected due to the presence of irregular light patterns such as specular highlight reflections. In this paper, we aim at removing specular highlights from endoscopic videos using machine learning. We propose using a temporal generative adversarial network (GAN) to inpaint the hidden anatomy under specularities, inferring its appearance spatially and from neighbouring frames, where they are not present in the same location. This is achieved using in-vivo data from gastric endoscopy (Hyper Kvasir) in a fully unsupervised manner that relies on the automatic detection of specular highlights. System evaluations show significant improvements to other methods through direct comparison and ablation studies that depict the importance of the network's temporal and transfer learning components. The generalisability of our system to different surgical setups and procedures was also evaluated qualitatively on in-vivo data of gastric endoscopy and ex-vivo porcine data (SERV-CT, SCARED). We also assess the effect of our method in comparison to other methods on computer vision tasks that underpin 3D reconstruction and camera motion estimation, namely stereo disparity, optical flow, and sparse point feature matching. These are evaluated quantitatively and qualitatively and results show a positive effect of our specular inpainting method on these tasks in a novel comprehensive analysis. Our code and dataset are made available at https://github.com/endomapper/Endo-STTN.
Rema Daher, Francisco Vasconcelos 0001, Danail Stoyanov
Medical Image Anal.3
2023 Robust endoscopic image mosaicking via fusion of multimodal estimation
abstract
We propose an endoscopic image mosaicking algorithm that is robust to light conditioning changes, specular reflections, and feature-less scenes. These conditions are especially common in minimally invasive surgery where the light source moves with the camera to dynamically illuminate close range scenes. This makes it difficult for a single image registration method to robustly track camera motion and then generate consistent mosaics of the expanded surgical scene across different and heterogeneous environments. Instead of relying on one specialised feature extractor or image registration method, we propose to fuse different image registration algorithms according to their uncertainties, formulating the problem as affine pose graph optimisation. This allows to combine landmarks, dense intensity registration, and learning-based approaches in a single framework. To demonstrate our application we consider deep learning-based optical flow, hand-crafted features, and intensity-based registration, however, the framework is general and could take as input other sources of motion estimation, including other sensor modalities. We validate the performance of our approach on three datasets with very different characteristics to highlighting its generalisability, demonstrating the advantages of our proposed fusion framework. While each individual registration algorithm eventually fails drastically on certain surgical scenes, the fusion approach flexibly determines which algorithms to use and in which proportion to more robustly obtain consistent mosaics.
Liang Li 0010, Evangelos B. Mazomenos, James Henry Chandler, Keith Obstein, Pietro Valdastri, Danail Stoyanov, Francisco Vasconcelos 0001
Medical Image Anal.6
2023 Minimum resolution requirements of digital pathology images for accurate classification
abstract
Digitization of pathology has been proposed as an essential mitigation strategy for the severe staffing crisis facing most pathology departments. Despite its benefits, several barriers have prevented widespread adoption of digital workflows, including cost and pathologist reluctance due to subjective image quality concerns. In this work, we quantitatively determine the minimum image quality requirements for binary classification of histopathology images of breast tissue in terms of spatial and sampling resolution. We train an ensemble of deep learning classifier models on publicly available datasets to obtain a baseline accuracy and computationally degrade these images according to our derived theoretical model to identify the minimum resolution necessary for acceptable diagnostic accuracy. Our results show that images can be degraded significantly below the resolution of most commercial whole-slide imaging systems while maintaining reasonable accuracy, demonstrating that macroscopic features are sufficient for binary classification of stained breast tissue. A rapid low-cost imaging system capable of identifying healthy tissue not requiring human assessment could serve as a triage system for reducing caseloads and alleviating the significant strain on the current workforce.
Lydia Neary-Zajiczek, Linas Beresna, Benjamin Razavi, Vijay Pawar, Michael J. Shaw, Danail Stoyanov
Medical Image Anal.6
2023 CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu 0009, Armine Vardazaryan, Fangfang Xia, Tong Xia, Fucang Jia, Yuxuan Yang 0007, Hao Wang 0081, Derong Yu, Guoyan Zheng, Xiaotian Duan, Neil Getty, Ricardo Sanchez-Matilla, Maria Robu, Li Zhang 0040, Huabin Chen, Jiacheng Wang 0002, Liansheng Wang 0002, Beerend G. A. Gerats, Sista Raviteja, Rachana Sathish, Rong Tao, Satoshi Kondo, Winnie Pang, Hongliang Ren 0001, Julian Ronald Abbing, Mohammad Hasan Sarhan, Sebastian Bodenstedt, Nithya Bhasker, Bruno Oliveira 0002, Helena R. Torres, Finn Gaida, Tobias Czempiel, João L. Vilaça, Pedro Morais, Jaime C. Fonseca 0001, Ruby Mae Egging, Inge Nicole Wijma, Chen Qian 0006, Guibin Bian, Zhen Li 0026, Velmurugan Balasubramanian, Debdoot Sheet, Imanol Luengo, Yuanbo Zhu, Shuai Ding 0001, Jakob-Anton Aschenbrenner, Nicolas Elini van der Kar, Mengya Xu, Mobarakol Islam, Seenivasan Lalithkumar, Alexander Jenke, Danail Stoyanov, Didier Mutter, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Nicolas Padoy
Medical Image Anal.57
2022 MoBYv2AL: Self-supervised Active Learning for Image Classification
Razvan Caramalau, Binod Bhattarai, Danail Stoyanov, Tae-Kyun Kim 0001
BMVC3
2022 3D Shape Variational Autoencoder Latent Disentanglement via Mini-Batch Feature Swapping for Bodies and Faces
abstract
Learning a disentangled, interpretable, and structured latent representation in 3D generative models of faces and bodies is still an open problem. The problem is particularly acute when control over identity features is required. In this paper, we propose an intuitive yet effective self-supervised approach to train a 3D shape variational autoencoder (VAE) which encourages a disentangled latent representation of identity features. Curating the mini-batch generation by swapping arbitrary features across different shapes allows to define a loss function leveraging known differences and similarities in the latent representations. Experimental results conducted on 3D meshes show that state-of-the-art methods for latent disentanglement are not able to disentangle identity features of faces and bodies. Our proposed method properly decouples the generation of such features while maintaining good representation and reconstruction capabilities. Our code and pretrained models are available at github.com/simofoti/3DVAE-SwapDisentangled.
Simone Foti, Bongjin Koo, Danail Stoyanov, Matthew J. Clarkson
CVPR3
2022 Generalized Product-of-Experts for Learning Multimodal Representations in Noisy Environments
abstract
A real-world application or setting involves interaction between different modalities (e.g., video, speech, text). In order to process the multimodal information automatically and use it for an end application, Multimodal Representation Learning (MRL) has emerged as an active area of research in recent times. MRL involves learning reliable and robust representations of information from heterogeneous sources and fusing them. However, in practice, the data acquired from different sources are typically noisy. In some extreme cases, a noise of large magnitude can completely alter the semantics of the data leading to inconsistencies in the parallel multimodal data. In this paper, we propose a novel method for multimodal representation learning in a noisy environment via the generalized product of experts technique. In the proposed method, we train a separate network for each modality to assess the credibility of information coming from that modality, and subsequently, the contribution from each modality is dynamically varied while estimating the joint distribution. We evaluate our method on two challenging benchmarks from two diverse domains: multimodal 3D hand-pose estimation and multimodal surgical video segmentation. We attain state-of-the-art performance on both benchmarks. Our extensive quantitative and qualitative evaluations show the advantages of our method compared to previous approaches.
Abhinav Joshi, Jinang Shah, Binod Bhattarai, Ashutosh Modi, Danail Stoyanov
ICMI6
2022 Localization of Interaction using Fibre-Optic Shape Sensing in Soft-Robotic Surgery Tools
abstract
Minimally invasive surgery requires real-time tool tracking to guide the surgeon where depth perception and visual occlusion present navigational challenges. Although vision-based and external sensor-based tracking methods exist, fibre-optic sensing can overcome their limitations as they can be integrated directly into the device, are biocompatible, small, robust and geometrically versatile. In this paper, we integrate a fibre Bragg grating-based shape sensor into a soft robotic device. The soft robot is the pneumatically attachable flexible (PAF) rail designed to act as a soft interface between manipulation tools and intra-operative imaging devices. We demonstrate that the shape sensing fibre can detect the location of the tools paired with the PAF rail, by exploiting the change in curvature sensed by the fibre when a strain is applied to it. We then validate this with a series of grasping tasks and continuous US swipes, using the system to detect in real-time the location of the tools interacting with the PAF rail. The overall location-sensing accuracy of the system is 64.6%, with a margin of error between predicted location and actual location of 3.75 mm.
Solène Dietsch, Aoife McDonald-Bowyer, Emmanouil Dimitrakakis, Joanna M. Coote, Lukas Lindenroth, Agostino Stilli, Danail Stoyanov
IROS7
2022 Navigation Among Movable Obstacles with Object Localization using Photorealistic Simulation
abstract
While mobile navigation has been focused on obstacle avoidance, Navigation Among Movable Obstacles (NAMO) via interaction with the environment, is a problem that is still open and challenging. This paper, presents a novel system integration to handle NAMO using visual feedback. In order to explore the capabilities of our introduced system, we explore the solution of the problem via graph-based path planning in a photorealistic simulator (NVIDIA Isaac Sim), in order to identify if the simulation-to-reality (sim2real) problem in robot navigation can be resolved. We consider the case where a wheeled robot navigates in a warehouse, in which movable boxes are common obstacles. We enable online real-time object localization and obstacle movability detection, to either avoid objects or, if it is not possible, to clear them out from the robot planned path by using pushing actions. We firstly test the integrated system in photorealistic environments, and we then validate the method on a real-world mobile wheeled robot (UCL MPPL) and its on-board sensory and computing system.
Kirsty Ellis, Henry Zhang, Danail Stoyanov, Dimitrios Kanoulas
IROS3
2022 BiometryNet: Landmark-based Fetal Biometry Estimation from Standard Ultrasound Planes
Netanell Avisdris, Leo Joskowicz, Brian Dromey, Anna L. David, Donald Peebles, Danail Stoyanov, Dafna Ben-Bashat, Sophia Bano
MICCAI (4)6
2022 Retrieval of Surgical Phase Transitions Using Reinforcement Learning
Sophia Bano, Ann-Sophie Page, Jan Deprest, Danail Stoyanov, Francisco Vasconcelos 0001
MICCAI (8)5
2022 Utility of optical see-through head mounted displays in augmented reality-assisted surgery: A systematic review
abstract
This article presents a systematic review of optical see-through head mounted display (OST-HMD) usage in augmented reality (AR) surgery applications from 2013 to 2020. Articles were categorised by: OST-HMD device, surgical speciality, surgical application context, visualisation content, experimental design and evaluation, accuracy and human factors of human-computer interaction. 91 articles fulfilled all inclusion criteria. Some clear trends emerge. The Microsoft HoloLens increasingly dominates the field, with orthopaedic surgery being the most popular application (28.6%). By far the most common surgical context is surgical guidance (n=58) and segmented preoperative models dominate visualisation (n=40). Experiments mainly involve phantoms (n=43) or system setup (n=21), with patient case studies ranking third (n=19), reflecting the comparative infancy of the field. Experiments cover issues from registration to perception with very different accuracy results. Human factors emerge as significant to OST-HMD utility. Some factors are addressed by the systems proposed, such as attention shift away from the surgical site and mental mapping of 2D images to 3D patient anatomy. Other persistent human factors remain or are caused by OST-HMD solutions, including ease of use, comfort and spatial perception issues. The significant upward trend in published articles is clear, but such devices are not yet established in the operating room and clinical studies showing benefit are lacking. A focused effort addressing technical registration and perceptual factors in the lab coupled with design that incorporates human factors considerations to solve clear clinical problems should ensure that the significant current research efforts will succeed.
Manuel Birlo, Philip J. Edwards, Matthew J. Clarkson, Danail Stoyanov
Medical Image Anal.4
2022 SERV-CT: A disparity dataset from cone-beam CT for validation of endoscopic 3D reconstruction
abstract
In computer vision, reference datasets from simulation and real outdoor scenes have been highly successful in promoting algorithmic development in stereo reconstruction. Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties and the presence of blood and smoke. These issues present difficulties for both stereo reconstruction itself and also for standardised dataset production. Previous datasets have been produced using computed tomography (CT) or structured light reconstruction on phantom or ex vivo models. We present a stereo-endoscopic reconstruction validation dataset based on cone-beam CT (SERV-CT). Two ex vivo small porcine full torso cadavers were placed within the view of the endoscope with both the endoscope and target anatomy visible in the CT scan. Subsequent orientation of the endoscope was manually aligned to match the stereoscopic view and benchmark disparities, depths and occlusions are calculated. The requirement of a CT scan limited the number of stereo pairs to 8 from each ex vivo sample. For the second sample an RGB surface was acquired to aid alignment of smooth, featureless surfaces. Repeated manual alignments showed an RMS disparity accuracy of around 2 pixels and a depth accuracy of about 2 mm. A simplified reference dataset is provided consisting of endoscope image pairs with corresponding calibration, disparities, depths and occlusions covering the majority of the endoscopic image and a range of tissue types, including smooth specular surfaces, as well as significant variation of depth. We assessed the performance of various stereo algorithms from online available repositories. There is a significant variation between algorithms, highlighting some of the challenges of surgical endoscopic images. The SERV-CT dataset provides an easy to use stereoscopic validation for surgical applications with smooth reference disparities and depths covering the majority of the endoscopic image. This complements existing resources well and we hope will aid the development of surgical endoscopic anatomical reconstruction algorithms.
Philip J. Edwards, Dimitris Psychogyios, Stefanie Speidel, Lena Maier-Hein, Danail Stoyanov
Medical Image Anal.5
2022 Surgical data science - from concepts toward clinical translation
abstract
Recent developments in data science in general and machine learning in particular have transformed the way experts envision the future of surgery. Surgical Data Science (SDS) is a new research field that aims to improve the quality of interventional healthcare through the capture, organization, analysis and modeling of data. While an increasing number of data-driven approaches and clinical applications have been studied in the fields of radiological and clinical data science, translational success stories are still lacking in surgery. In this publication, we shed light on the underlying reasons and provide a roadmap for future advances in the field. Based on an international workshop involving leading researchers in the field of SDS, we review current practice, key achievements and initiatives as well as available standards and tools for a number of topics relevant to the field, namely (1) infrastructure for data acquisition, storage and access in the presence of regulatory constraints, (2) data annotation and sharing and (3) data analytics. We further complement this technical perspective with (4) a review of currently available SDS products and the translational progress from academia and (5) a roadmap for faster clinical translation and exploitation of the full potential of SDS, based on an international multi-round Delphi process.
Lena Maier-Hein, Matthias Eisenmann, Duygu Sarikaya, Keno März, Toby Collins, Anand Malpani, Johannes Fallert, Hubertus Feußner, Stamatia Giannarou, Pietro Mascagni, Hirenkumar Nakawala, Adrian Park 0001, Carla M. Pugh, Danail Stoyanov, S. Swaroop Vedula, Kevin Cleary, Gabor Fichtinger, Germain Forestier, Bernard Gibaud, Teodor P. Grantcharov, Makoto Hashizume, Doreen Heckmann-Nötzel, Hannes Kenngott, Ron Kikinis, Lars Mündermann, Nassir Navab, Sinan Onogur, Tobias Roß, Raphael Sznitman, Russell H. Taylor, Minu Tizabi, Martin Wagner 0001, Gregory D. Hager, Thomas Neumuth, Nicolas Padoy, Justin Collins, Ines Gockel, Jan Goedeke, Daniel A. Hashimoto, Luc Joyeux, Kyle Lam, Daniel Richard Leff, Amin Madani, Hani J. Marcus, Ozanan R. Meireles, Alexander Seitel, Dogu Teber, Frank Ückert, Beat P. Müller-Stich, Pierre Jannin, Stefanie Speidel
Medical Image Anal.14
2022 Polyp detection on video colonoscopy using a hybrid 2D/3D CNN
abstract
Colonoscopy is the gold standard for early diagnosis and pre-emptive treatment of colorectal cancer by detecting and removing colonic polyps. Deep learning approaches to polyp detection have shown potential for enhancing polyp detection rates. However, the majority of these systems are developed and evaluated on static images from colonoscopies, whilst in clinical practice the treatment is performed on a real-time video feed. Non-curated video data remains a challenge, as it contains low-quality frames when compared to still, selected images often obtained from diagnostic records. Nevertheless, it also embeds temporal information that can be exploited to increase predictions stability. A hybrid 2D/3D convolutional neural network architecture for polyp segmentation is presented in this paper. The network is used to improve polyp detection by encompassing spatial and temporal correlation of the predictions while preserving real-time detections. Extensive experiments show that the hybrid method outperforms a 2D baseline. The proposed architecture is validated on videos from 46 patients and on the publicly available SUN polyp database. A higher performance and increased generalisability indicate that real-world clinical implementations of automated polyp detection can benefit from the hybrid algorithm and the inclusion of temporal information.
Juana González-Bueno Puyal, Patrick Brandao, Omer F. Ahmad, Kanwal K. Bhatia, Daniel Toth 0001, Rawen Kader, Laurence B. Lovat, Peter Mountney, Danail Stoyanov
Medical Image Anal.9
2022 Robot-Assisted Minimally Invasive Surgery - Surgical Robotics in the Data Age
abstract
Telesurgical robotics, as a technical solution for robot-assisted minimally invasive surgery (RAMIS), has become the first domain within medicosurgical robotics that achieved a true global clinical adoption. Its relative success (still at a low single-digit percentile total market penetration) roots in the particular human-in-the-loop control, in which the trained surgeon is always kept responsible for the clinical outcome achieved by the robot-actuated invasive tools. Nowadays, this paradigm is challenged by the need for improved surgical performance, traceability, and safety reaching beyond the human capabilities. Partially due to the technical complexity and the financial burden, the adoption of telesurgical robotics has not reached its full potential, by far. Apart from the absolutely market-dominating da Vinci surgical system, there are already 60+ emerging RAMIS robot types, out of which 15 have already achieved some form of regulatory clearance. This article aims to connect the technological advancement with the principles of commercialization, particularly looking at engineering components that are under development and have the potential to bring significant advantages to the clinical practice. Current RAMIS robots often do not exceed the functionalities deriving from their mechatronics, due to the lack of data-driven assistance and smart human–machine collaboration. Computer assistance is gradually gaining more significance within emerging RAMIS systems. Enhanced manipulation capabilities, refined sensors, advanced vision, task-level automation, smart safety features, and data integration mark together the inception of a new era in telesurgical robotics, infiltrated by machine learning (ML) and artificial intelligence (AI) solutions. Observing other domains, it is definite that a key requirement of a robust AI is the good quality data, derived from proper data acquisition and sharing to allow building solutions in real time based on ML. Emerging RAMIS technologies are reviewed both in a historical and a future perspective.
Tamás Haidegger, Stefanie Speidel, Danail Stoyanov, Richard M. Satava
Proc. IEEE3
2022 Initial Responses to False Positives in AI-Supported Continuous Interactions: A Colonoscopy Case Study
abstract
The use of artificial intelligence (AI) in clinical support systems is increasing. In this article, we focus on AI support for continuous interaction scenarios. A thorough understanding of end-user behaviour during these continuous human-AI interactions, in which user input is sustained over time and during which AI suggestions can appear at any time, is still missing. We present a controlled lab study involving 21 endoscopists and an AI colonoscopy support system. Using a custom-developed application and an off-the-shelf videogame controller, we record participants’ navigation behaviour and clinical assessment across 14 endoscopic videos. Each video is manually annotated to mimic an AI recommendation, being either true positive or false positive in nature. We find that time between AI recommendation and clinical assessment is significantly longer for incorrect assessments. Further, the type of medical content displayed significantly affects decision time. Finally, we discover that the participant’s clinical role plays a large part in the perception of clinical AI support systems. Our study presents a realistic assessment of the effects of imperfect and continuous AI support in a clinical scenario.
Niels van Berkel, Jeremy Opie, Omer F. Ahmad, Laurence B. Lovat, Danail Stoyanov, Ann Blandford
ACM Trans. Interact. Intell. Syst.5
2022 Gesture Recognition in Robotic Surgery With Multimodal Attention
abstract
Automatically recognising surgical gestures from surgical data is an important building block of automated activity recognition and analytics, technical skill assessment, intra-operative assistance and eventually robotic automation. The complexity of articulated instrument trajectories and the inherent variability due to surgical style and patient anatomy make analysis and fine-grained segmentation of surgical motion patterns from robot kinematics alone very difficult. Surgical video provides crucial information from the surgical site with context for the kinematic data and the interaction between the instruments and tissue. Yet sensor fusion between the robot data and surgical video stream is non-trivial because the data have different frequency, dimensions and discriminative capability. In this paper, we integrate multimodal attention mechanisms in a two-stream temporal convolutional network to compute relevance scores and weight kinematic and visual feature representations dynamically in time, aiming to aid multimodal network training and achieve effective sensor fusion. We report the results of our system on the JIGSAWS benchmark dataset and on a new in vivo dataset of suturing segments from robotic prostatectomy procedures. Our results are promising and obtain multimodal prediction sequences with higher accuracy and better temporal structure than corresponding unimodal solutions. Visualization of attention scores also gives physically interpretable insights on network understanding of strengths and weaknesses of each sensor.
Beatrice van Amsterdam, Isabel Funke, Philip J. Edwards, Stefanie Speidel, Justin Collins, Ashwin Sridhar, John D. Kelly, Matthew J. Clarkson, Danail Stoyanov
IEEE Trans. Medical Imaging9
2022 SSIS-Seg: Simulation-Supervised Image Synthesis for Surgical Instrument Segmentation
abstract
Surgical instrument segmentation can be used in a range of computer assisted interventions and automation in surgical robotics. While deep learning architectures have rapidly advanced the robustness and performance of segmentation models, most are still reliant on supervision and large quantities of labelled data. In this paper, we present a novel method for surgical image generation that can fuse robotic instrument simulation and recent domain adaptation techniques to synthesize artificial surgical images to train surgical instrument segmentation models. We integrate attention modules into well established image generation pipelines and propose a novel cost function to support supervision from simulation frames in model training. We provide an extensive evaluation of our method in terms of segmentation performance along with a validation study on image quality using evaluation metrics. Additionally, we release a novel segmentation dataset from real surgeries that will be shared for research purposes. Both binary and semantic segmentation have been considered, and we show the capability of our synthetic images to train segmentation models compared with the latest methods from the literature.
Emanuele Colleoni, Dimitris Psychogyios, Beatrice van Amsterdam, Francisco Vasconcelos 0001, Danail Stoyanov
IEEE Trans. Medical Imaging5
2022 Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene Segmentation
abstract
Automatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL.
Yueming Jin, Yang Yu 0070, Cheng Chen 0013, Pheng-Ann Heng, Danail Stoyanov
IEEE Trans. Medical Imaging6
2022 MSDESIS: Multitask Stereo Disparity Estimation and Surgical Instrument Segmentation
abstract
Reconstructing the 3D geometry of the surgical site and detecting instruments within it are important tasks for surgical navigation systems and robotic surgery automation. Traditional approaches treat each problem in isolation and do not account for the intrinsic relationship between segmentation and stereo matching. In this paper, we present a learning-based framework that jointly estimates disparity and binary tool segmentation masks. The core component of our architecture is a shared feature encoder which allows strong interaction between the aforementioned tasks. Experimentally, we train two variants of our network with different capacities and explore different training schemes including both multi-task and single-task learning. Our results show that supervising the segmentation task improves our network's disparity estimation accuracy. We demonstrate a domain adaptation scheme where we supervise the segmentation task with monocular data and achieve domain adaptation of the adjacent disparity task, reducing disparity End-Point-Error and depth mean absolute error by 77.73% and 61.73% respectively compared to the pre-trained baseline model. Our best overall multi-task model, trained with both disparity and segmentation data in subsequent phases, achieves 89.15% mean Intersection-over-Union in RIS and 3.18 millimetre depth mean absolute error in SCARED test sets. Our proposed multi-task architecture is real-time, able to process ( 1280×1024 ) stereo input and simultaneously estimate disparity maps and segmentation masks at 22 frames per second. The model code and pre-trained models are made available: https://github.com/dimitrisPs/msdesis.
Dimitris Psychogyios, Evangelos B. Mazomenos, Francisco Vasconcelos 0001, Danail Stoyanov
IEEE Trans. Medical Imaging4
2021 Deep Reinforcement Learning for Concentric Tube Robot Control with a Goal-Based Curriculum
abstract
Concentric Tube Robots (CTRs), a type of continuum robot, are a collection of concentric, pre-curved tubes composed of super elastic nickel titanium alloy. CTRs can bend and twist from the interactions between neighboring tubes causing the kinematics and therefore control of the end-effector to be very challenging to model. In this paper, we develop a control scheme for a CTR end-effector in Cartesian space with no prior kinematic model using a deep reinforcement learning (DRL) approach with a goal-based curriculum reward strategy. We explore the use of curricula by changing the goal tolerance through training with constant, linear and exponential decay functions. Also, relative and absolute joint representations as a way of improving training convergence are explored. Quantitative comparisons for combinations of curricula and joint representations are performed and the exponential decay relative approach is used for training a robust policy in a noise-induced simulation environment. Compared to a previous DRL approach, our new method reduces training time and employs a more complex simulation environment. We report mean Cartesian errors of 1.29 mm and a success rate of 0.93 with a relative decay curriculum. In path following, we report mean errors of 1.37 mm in a noise-induced path following task. Albeit in simulation, these results indicate the promise of using DRL in model free control of continuum robots and CTRs in particular.
Keshav Iyengar, Danail Stoyanov
ICRA2
2021 Autonomous object harvesting using synchronized optoelectronic microrobots
abstract
Optoelectronic tweezer-driven microrobots (OETdMs) are a versatile micromanipulation technology based on the application of light induced dielectrophoresis to move small dielectric structures (microrobots) across a photoconductive substrate. The microrobots in turn can be used to exert forces on secondary objects and carry out a wide range of micromanipulation operations, including collecting, transporting and depositing microscopic cargos. In contrast to alternative (direct) micromanipulation techniques, OETdMs are relatively gentle, making them particularly well suited to interacting with sensitive objects such as biological cells. However, at present such systems are used exclusively under manual control by a human operator. This limits the capacity for simultaneous control of multiple microrobots, reducing both experimental throughput and the possibility of cooperative multi-robot operations. In this article, we describe an approach to automated targeting and path planning to enable open-loop control of multiple microrobots. We demonstrate the performance of the method in practice, using microrobots to simultaneously collect, transport and deposit silica microspheres. Using computational simulations based on real microscopic image data, we investigate the capacity of microrobots to collect target cells from within a dissociated tissue culture. Our results indicate the feasibility of using OETdMs to autonomously carry out micromanipulation tasks within complex, unstructured environments.
Christopher Bendkowski, Laurent Mennillo, Mohamed Elsayed 0008, Filip Stojic, Shuailong Zhang, Cindi Morshead, Vijay Pawar, Aaron R. Wheeler, Danail Stoyanov, Michael J. Shaw
IROS11
2021 AutoFB: Automating Fetal Biometry Estimation from Standard Ultrasound Planes
Sophia Bano, Brian Dromey, Francisco Vasconcelos 0001, Raffaele Napolitano, Anna L. David, Donald Peebles, Danail Stoyanov
MICCAI (7)7
2021 Detection of Critical Structures in Laparoscopic Cholecystectomy Using Label Relaxation and Self-supervision
David Owen 0001, Maria Grammatikopoulou, Imanol Luengo, Danail Stoyanov
MICCAI (4)4
2021 Scalable Joint Detection and Segmentation of Surgical Instruments with Weak Supervision
Ricardo Sanchez-Matilla, Maria Robu, Imanol Luengo, Danail Stoyanov
MICCAI (2)4
2021 Designing Visual Markers for Continuous Artificial Intelligence Support: A Colonoscopy Case Study
abstract
Colonoscopy, the visual inspection of the large bowel using an endoscope, offers protection against colorectal cancer by allowing for the detection and removal of pre-cancerous polyps. The literature on polyp detection shows widely varying miss rates among clinicians, with averages ranging around 22%--27%. While recent work has considered the use of AI support systems for polyp detection, how to visualise and integrate these systems into clinical practice is an open question. In this work, we explore the design of visual markers as used in an AI support system for colonoscopy. Supported by the gastroenterologists in our team, we designed seven unique visual markers and rendered them on real-life patient video footage. Through an online survey targeting relevant clinical staff ( N = 36), we evaluated these designs and obtained initial insights and understanding into the way in which clinical staff envision AI to integrate in their daily work-environment. Our results provide concrete recommendations for the future deployment of AI support systems in continuous, adaptive scenarios.
Niels van Berkel, Omer F. Ahmad, Danail Stoyanov, Laurence B. Lovat, Ann Blandford
ACM Trans. Comput. Heal.3
2021 Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy
abstract
The Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques.
Sharib Ali, Mariia Dmitrieva, Noha M. Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Bogdan J. Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Lê Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian 0004, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher
Medical Image Anal.28
2021 CaDIS: Cataract dataset for surgical RGB-image segmentation
abstract
Video feedback provides a wealth of information about surgical procedures and is the main sensory cue for surgeons. Scene understanding is crucial to computer assisted interventions (CAI) and to post-operative analysis of the surgical procedure. A fundamental building block of such capabilities is the identification and localization of surgical instruments and anatomical structures through semantic segmentation. Deep learning has advanced semantic segmentation techniques in the recent years but is inherently reliant on the availability of labelled datasets for model training. This paper introduces a dataset for semantic segmentation of cataract surgery videos complementing the publicly available CATARACTS challenge dataset. In addition, we benchmark the performance of several state-of-the-art deep learning models for semantic segmentation on the presented dataset. The dataset is publicly available at https://cataracts-semantic-segmentation2020.grand-challenge.org/.
Maria Grammatikopoulou, Evangello Flouty, Abdolrahim Kadkhodamohammadi, Gwenolé Quellec, Andre Chow, Jean Nehme, Imanol Luengo, Danail Stoyanov
Medical Image Anal.8
2021 Corrigendum to Dual-modality endoscopic probe for tissue surface shape reconstruction and hyperspectral imaging enabled by deep neural networks [Medical Image Analysis 48 (2018) 162-176/2018.06.004]
Jianyu Lin, Neil Clancy, Taran Tatla, Danail Stoyanov, Lena Maier-Hein, Daniel S. Elson
Medical Image Anal.6
2021 Image Compositing for Segmentation of Surgical Tools Without Manual Annotations
abstract
Producing manual, pixel-accurate, image segmentation labels is tedious and time-consuming. This is often a rate-limiting factor when large amounts of labeled images are required, such as for training deep convolutional networks for instrument-background segmentation in surgical scenes. No large datasets comparable to industry standards in the computer vision community are available for this task. To circumvent this problem, we propose to automate the creation of a realistic training dataset by exploiting techniques stemming from special effects and harnessing them to target training performance rather than visual appeal. Foreground data is captured by placing sample surgical instruments over a chroma key (a.k.a. green screen) in a controlled environment, thereby making extraction of the relevant image segment straightforward. Multiple lighting conditions and viewpoints can be captured and introduced in the simulation by moving the instruments and camera and modulating the light source. Background data is captured by collecting videos that do not contain instruments. In the absence of pre-existing instrument-free background videos, minimal labeling effort is required, just to select frames that do not contain surgical instruments from videos of surgical interventions freely available online. We compare different methods to blend instruments over tissue and propose a novel data augmentation approach that takes advantage of the plurality of options. We show that by training a vanilla U-Net on semi-synthetic data only and applying a simple post-processing, we are able to match the results of the same network trained on a publicly available manually labeled real dataset.
Luis C. García-Peraza-Herrera, Lucas Fidon, Claudia D'Ettorre, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin
IEEE Trans. Medical Imaging4
2020 Predicting Visual Overlap of Images Through Interpretable Non-metric Box Embeddings
Anita Rau, Guillermo Garcia-Hernando, Danail Stoyanov, Gabriel J. Brostow, Daniyar Turmukhambetov
ECCV (5)3
2020 Multi-Task Recurrent Neural Network for Surgical Gesture Recognition and Progress Prediction
abstract
Surgical gesture recognition is important for surgical data science and computer-aided intervention. Even with robotic kinematic information, automatically segmenting surgical steps presents numerous challenges because surgical demonstrations are characterized by high variability in style, duration and order of actions. In order to extract discriminative features from the kinematic signals and boost recognition accuracy, we propose a multi-task recurrent neural network for simultaneous recognition of surgical gestures and estimation of a novel formulation of surgical task progress. To show the effectiveness of the presented approach, we evaluate its application on the JIGSAWS dataset, that is currently the only publicly available dataset for surgical gesture recognition featuring robot kinematic data. We demonstrate that recognition performance improves in multi-task frameworks with progress estimation without any additional manual labelling and training.
Beatrice van Amsterdam, Matthew J. Clarkson, Danail Stoyanov
ICRA3
2020 Deep Placental Vessel Segmentation for Fetoscopic Mosaicking
Sophia Bano, Francisco Vasconcelos 0001, Luke M. Shepherd, Emmanuel B. Vander Poorten, Tom Vercauteren, Sébastien Ourselin, Anna L. David, Jan Deprest, Danail Stoyanov
MICCAI (3)9
2020 Synthetic and Real Inputs for Tool Segmentation in Robotic Surgery
Emanuele Colleoni, Philip J. Edwards, Danail Stoyanov
MICCAI (3)3
2020 Surgical Video Motion Magnification with Suppression of Instrument Artefacts
Mirek Janatka, Hani J. Marcus, Neil L. Dorward, Danail Stoyanov
MICCAI (3)4
2020 Endoscopic Polyp Segmentation Using a Hybrid 2D/3D CNN
Juana González-Bueno Puyal, Kanwal K. Bhatia, Patrick Brandao, Omer F. Ahmad, Daniel Toth 0001, Rawen Kader, Laurence B. Lovat, Peter Mountney, Danail Stoyanov
MICCAI (6)9
2020 Refractive Two-View Reconstruction for Underwater 3D Vision
abstract
Recovering 3D geometry from cameras in underwater applications involves the Refractive Structure-from-Motion problem where the non-linear distortion of light induced by a change of medium density invalidates the single viewpoint assumption. The pinhole-plus-distortion camera projection model suffers from a systematic geometric bias since refractive distortion depends on object distance. This leads to inaccurate camera pose and 3D shape estimation. To account for refraction, it is possible to use the axial camera model or to explicitly consider one or multiple parallel refractive interfaces whose orientations and positions with respect to the camera can be calibrated. Although it has been demonstrated that the refractive camera model is well-suited for underwater imaging, Refractive Structure-from-Motion remains particularly difficult to use in practice when considering the seldom studied case of a camera with a flat refractive interface. Our method applies to the case of underwater imaging systems whose entrance lens is in direct contact with the external medium. By adopting the refractive camera model, we provide a succinct derivation and expression for the refractive fundamental matrix and use this as the basis for a novel two-view reconstruction method for underwater imaging. For validation we use synthetic data to show the numerical properties of our method and we provide results on real data to demonstrate its practical application within laboratory settings and for medical applications in fluid-immersed endoscopy. We demonstrate our approach outperforms classic two-view Structure-from-Motion method relying on the pinhole-plus-distortion camera model.
François Chadebecq, Francisco Vasconcelos 0001, Rene M. Lacher, Efthymios Maneas, Adrien E. Desjardins, Sébastien Ourselin, Tom Vercauteren, Danail Stoyanov
Int. J. Comput. Vis.8
2020 Surgical spectral imaging
abstract
Recent technological developments have resulted in the availability of miniaturised spectral imaging sensors capable of operating in the multi- (MSI) and hyperspectral imaging (HSI) regimes. Simultaneous advances in image-processing techniques and artificial intelligence (AI), especially in machine learning and deep learning, have made these data-rich modalities highly attractive as a means of extracting biological information non-destructively. Surgery in particular is poised to benefit from this, as spectrally-resolved tissue optical properties can offer enhanced contrast as well as diagnostic and guidance information during interventions. This is particularly relevant for procedures where inherent contrast is low under standard white light visualisation. This review summarises recent work in surgical spectral imaging (SSI) techniques, taken from Pubmed, Google Scholar and arXiv searches spanning the period 2013-2019. New hardware, optimised for use in both open and minimally-invasive surgery (MIS), is described, and recent commercial activity is summarised. Computational approaches to extract spectral information from conventional colour images are reviewed, as tip-mounted cameras become more commonplace in MIS. Model-based and machine learning methods of data analysis are discussed in addition to simulation, phantom and clinical validation experiments. A wide variety of surgical pilot studies are reported but it is apparent that further work is needed to quantify the clinical value of MSI/HSI. The current trend toward data-driven analysis emphasises the importance of widely-available, standardised spectral imaging datasets, which will aid understanding of variability across organs and patients, and drive clinical translation.
Neil Clancy, Geoffrey Jones, Lena Maier-Hein, Daniel S. Elson, Danail Stoyanov
Medical Image Anal.5
2019 Weakly Supervised Recognition of Surgical Gestures
abstract
Kinematic trajectories recorded from surgical robots contain information about surgical gestures and potentially encode cues about surgeon's skill levels. Automatic segmentation of these trajectories into meaningful action units could help to develop new metrics for surgical skill assessment as well as to simplify surgical automation. State-of-the-art methods for action recognition relied on manual labelling of large datasets, which is time consuming and error prone. Unsupervised methods have been developed to overcome these limitations. However, they often rely on tedious parameter tuning and perform less well than supervised approaches, especially on data with high variability such as surgical trajectories. Hence, the potential of weak supervision could be to improve unsupervised learning while avoiding manual annotation of large datasets. In this paper, we used at a minimum one expert demonstration and its ground truth annotations to generate an appropriate initialization for a GMM-based algorithm for gesture recognition. We showed on real surgical demonstrations that the latter significantly outperforms standard task-agnostic initialization methods. We also demonstrated how to improve the recognition accuracy further by redefining the actions and optimising the inputs.
Beatrice van Amsterdam, Hirenkumar Nakawala, Elena De Momi, Danail Stoyanov
ICRA4
2019 Robotic Control of a Multi-Modal Rigid Endoscope Combining Optical Imaging with All-Optical Ultrasound
abstract
Fetoscopy is a technically challenging surgery, due to the dynamic environment and low diameter endoscopes often resulting in a limited field of view. In this paper, we report on the design and operation of a robotic multimodal endoscope with optical ultrasound and white light stereo camera. The manufacture and control of the endoscope is presented, along with large area (80 mm ×80 mm) surface visualisations of a placenta phantom using the optical ultrasound sensor. The repeatability of the surface visualisations was found to be 0. 446 ± 0.139 mm and 0. 267 ± 0.017 mm for a raster and spiral scan, respectively.
George Dwyer, Richard J. Colchester, Erwin J. Alles, Efthymios Maneas, Sébastien Ourselin, Tom Vercauteren, Jan Deprest, Emmanuel B. Vander Poorten, Paolo De Coppi, Adrien E. Desjardins, Danail Stoyanov
ICRA11
2019 RCM-SLAM: Visual localisation and mapping under remote centre of motion constraints
abstract
In robotic surgery the motion of instruments and the laparoscopic camera is constrained by their insertion ports, i. e. a remote centre of motion (RCM). We propose a Simultaneous Localisation and Mapping (SLAM) approach that estimates laparoscopic camera motion under RCM constraints. To achieve this we derive a minimal solver for the absolute camera pose given two 2D-3D point correspondences (RCM-PnP) and also a bundle adjustment optimiser that refines camera poses within an RCM-constrained parameterisation. These two methods are used together with previous work on relative pose estimation under RCM [1] to assemble a SLAM pipeline suitable for robotic surgery. Our simulations show that RCM-PnP outperforms conventional PnP for a wide noise range in the RCM position. Results with video footage from a robotic prostatectomy show that RCM constraints significantly improve camera pose estimation.
Francisco Vasconcelos 0001, Evangelos B. Mazomenos, John D. Kelly, Danail Stoyanov
ICRA4
2019 Semi-Autonomous Interventional Manipulation using Pneumatically Attachable Flexible Rails
abstract
During laparoscopic surgery, tissues frequently need to be retracted and mobilized for manipulation or visualisation. State-of-the-art robotic platforms for minimally invasive surgery (MIS) typically rely on rigid tools to interact with soft tissues. Such tools offer a very narrow contact surface thus applying relatively large forces that can lead to tissue damage, posing a risk for the success of the procedure and ultimately for the patient. In this paper, we show how the use of Pneumatically Attachable Flexible (PAF) rail, a vacuum-based soft attachment for laparoscopic applications, can reduce such risk by offering a larger contact surface between the tool and the tissue. Ex vivo experiments are presented investigating the short- and long-term effects of different levels of vacuum pressure on the tissues surface. These experiments aim at evaluating the best trade-off between applied pressure, potential damage, task duration and connection stability. A hybrid control system has been developed to perform and investigate the organ repositioning task using the proposed system. The task is only partially automated allowing the surgeon to be part of the control loop. A gradient-based planning algorithm is integrated with learning from teleoperation algorithm which allows the robot to improve the learned trajectory. The use of Similar Smooth Path Repositioning (SSPR) algorithm is proposed to improve a demonstrated trajectory based on a known cost function. The results obtained show that a smoother trajectory allows to decrease the minimum level of pressure needed to guarantee active suction during PAF positioning and placement.
Claudia D'Ettorre, Agostino Stilli, George Dwyer, Joana B. Neves, Maxine Tran, Danail Stoyanov
IROS6
2019 Deep Sequential Mosaicking of Fetoscopic Videos
Sophia Bano, Francisco Vasconcelos 0001, Marcel Tella-Amo, George Dwyer, Caspar Gruijthuijsen, Jan Deprest, Sébastien Ourselin, Emmanuel B. Vander Poorten, Tom Vercauteren, Danail Stoyanov
MICCAI (1)10
2019 Whole-Sample Mapping of Cancerous and Benign Tissue Properties
Lydia Neary-Zajiczek, Clara Essmann, Neil Clancy, Aiman Haider, Elena Miranda, Michael J. Shaw, Amir Gander, Brian R. Davidson, Delmiro Fernandez-Reyes, Vijay Pawar, Danail Stoyanov
MICCAI (1)11
2019 Patch-based adaptive weighting with segmentation and scale (PAWSS) for visual tracking in surgical video
abstract
Vision-based tracking in an important component for building computer assisted interventions in minimally invasive surgery as it facilitates estimation of motion for instruments and anatomical targets. Tracking-by-detection algorithms are widely used for visual tracking, where the problem is treated as a classification task and a tracking target appearance model is updated over time using online learning. In challenging conditions, like surgical scenes, where tracking targets deform and vary in scale, the update step is prone to include background information in model appearance or to lack the ability to estimate change of scale, which degrades the performance of classifier. In this paper, we propose a Patch-based Adaptive Weighting with Segmentation and Scale (PAWSS) tracking framework that tackles both scale and background problems. A simple but effective colour-based segmentation model is used to suppress background information and multi-scale samples are extracted to enrich the training pool, which allows the tracker to handle both incremental and abrupt scale variations between frames. Experimentally, we evaluate our approach on Online Tracking Benchmark (OTB) dataset and Visual Object Tracking (VOT) challenge datasets, showing that our approach outperforms recent state-of-the-art trackers, and it especially improves successful rate score on OTB dataset, while on VOT datasets, PAWSS ranks among the top trackers while operating at real-time frame rates. Focusing on the application of PAWSS to surgical scenes, we evaluate on MICCAI 2015 challenge instrument tracking challenge and in vivo datasets, showing that our approach performs the best among all submitted methods and also has promising performance on in vivo surgical instrument tracking.
Xiaofei Du 0001, Maximilian Allan, Sebastian Bodenstedt, Lena Maier-Hein, Stefanie Speidel, Alessio Dore, Danail Stoyanov
Medical Image Anal.7
2019 Nonrigid reconstruction of 3D breast surfaces with a low-cost RGBD camera for surgical planning and aesthetic evaluation
abstract
Accounting for 26% of all new cancer cases worldwide, breast cancer remains the most common form of cancer in women. Although early breast cancer has a favourable long-term prognosis, roughly a third of patients suffer from a suboptimal aesthetic outcome despite breast conserving cancer treatment. Clinical-quality 3D modelling of the breast surface therefore assumes an increasingly important role in advancing treatment planning, prediction and evaluation of breast cosmesis. Yet, existing 3D torso scanners are expensive and either infrastructure-heavy or subject to motion artefacts. In this paper we employ a single consumer-grade RGBD camera with an ICP-based registration approach to jointly align all points from a sequence of depth images non-rigidly. Subtle body deformation due to postural sway and respiration is successfully mitigated leading to a higher geometric accuracy through regularised locally affine transformations. We present results from 6 clinical cases where our method compares well with the gold standard and outperforms a previous approach. We show that our method produces better reconstructions qualitatively by visual assessment and quantitatively by consistently obtaining lower landmark error scores and yielding more accurate breast volume estimates.
Rene M. Lacher, Francisco Vasconcelos 0001, Norman R. Williams, Gerrit Rindermann, John H. Hipwell, David J. Hawkes, Danail Stoyanov
Medical Image Anal.7
2019 Widening siamese architectures for stereo matching
abstract
Computational stereo is one of the classical problems in computer vision. Numerous algorithms and solutions have been reported in recent years focusing on developing methods for computing similarity, aggregating it to obtain spatial support and finally optimizing an energy function to find the final disparity. In this paper, we focus on the feature extraction component of stereo matching architecture and we show standard CNNs operation can be used to improve the quality of the features used to find point correspondences. Furthermore, we use a simple space aggregation that hugely simplifies the correlation learning problem, allowing us to better evaluate the quality of the features extracted. Our results on benchmark data are compelling and show promising potential even without refining the solution.
Patrick Brandao, Evangelos B. Mazomenos, Danail Stoyanov
Pattern Recognit. Lett.3
2018 SurReal: enhancing Surgical simulation Realism using style transfer
Imanol Luengo, Evangello Flouty, Petros Giataganas, Piyamate Wisanuvej, Jean Nehme, Danail Stoyanov
BMVC6
2018 Automated Pick-Up of Suturing Needles for Robotic Surgical Assistance
abstract
Robot-assisted laparoscopic prostatectomy (RALP) is a treatment for prostate cancer that involves complete or nerve sparing removal prostate tissue that contains cancer. After removal the bladder neck is successively sutured directly with the urethra. The procedure is called urethrovesical anastomosis and is one of the most dexterity demanding tasks during RALP. Two suturing instruments and a pair of needles are used in combination to perform a running stitch during urethrovesical anastomosis. While robotic instruments provide enhanced dexterity to perform the anastomosis, it is still highly challenging and difficult to learn. In this paper, we presents a vision-guided needle grasping method for automatically grasping the needle that has been inserted into the patient prior to anastomosis. We aim to automatically grasp the suturing needle in a position that avoids hand-offs and immediately enables the start of suturing. The full grasping process can be broken down into: a needle detection algorithm; an approach phase where the surgical tool moves closer to the needle based on visual feedback; and a grasping phase through path planning based on observed surgical practice. Our experimental results show examples of successful autonomous grasping that has the potential to simplify and decrease the operational time in RALP by assisting a small component of urethrovesical anastomosis.
Claudia D'Ettorre, George Dwyer, Xiaofei Du 0001, François Chadebecq, Francisco Vasconcelos 0001, Elena De Momi, Danail Stoyanov
ICRA7
2018 Higher Order of Motion Magnification for Vessel Localisation in Surgical Video
Mirek Janatka, Ashwin Sridhar, John D. Kelly, Danail Stoyanov
MICCAI (4)4
2018 Automated Performance Assessment in Transoesophageal Echocardiography with Convolutional Neural Networks
Evangelos B. Mazomenos, Kamakshi Bansal, Bruce Martin, Andrew P. T. Smith, Susan Wright, Danail Stoyanov
MICCAI (4)6
2018 DeepPhase: Surgical Phase Recognition in CATARACTS Videos
Odysseas Zisimopoulos, Evangello Flouty, Imanol Luengo, Petros Giataganas, Jean Nehme, Andre Chow, Danail Stoyanov
MICCAI (4)7
2018 Dual-modality endoscopic probe for tissue surface shape reconstruction and hyperspectral imaging enabled by deep neural networks
Jianyu Lin, Neil Clancy, Taran Tatla, Danail Stoyanov, Lena Maier-Hein, Daniel S. Elson
Medical Image Anal.6
2018 Long Term Safety Area Tracking (LT-SAT) with online failure detection and recovery for robotic minimally invasive surgery
Veronica Penza, Xiaofei Du 0001, Danail Stoyanov, Antonello Forgione, Leonardo S. Mattos, Elena De Momi
Medical Image Anal.3
2018 3-D Pose Estimation of Articulated Instruments in Robotic Minimally Invasive Surgery
abstract
Estimating the 3-D pose of instruments is an important part of robotic minimally invasive surgery for automation of basic procedures as well as providing safety features, such as virtual fixtures. Image-based methods of 3-D pose estimation provide a non-invasive low cost solution compared with methods that incorporate external tracking systems. In this paper, we extend our recent work in estimating rigid 3-D pose with silhouette and optical flow-based features to incorporate the articulated degrees-of-freedom (DOFs) of robotic instruments within a gradient-based optimization framework. Validation of the technique is provided with a calibrated ex-vivo study from the da Vinci Research Kit (DVRK) robotic system, where we perform quantitative analysis on the errors each DOF of our tracker. Additionally, we perform several detailed comparisons with recently published techniques that combine visual methods with kinematic data acquired from the joint encoders. Our experiments demonstrate that our method is competitively accurate while relying solely on image data.
Maximilian Allan, Sébastien Ourselin, David J. Hawkes, John D. Kelly, Danail Stoyanov
IEEE Trans. Medical Imaging5
2018 Articulated Multi-Instrument 2-D Pose Estimation Using Fully Convolutional Networks
abstract
Instrument detection, pose estimation, and tracking in surgical videos are an important vision component for computer-assisted interventions. While significant advances have been made in recent years, articulation detection is still a major challenge. In this paper, we propose a deep neural network for articulated multi-instrument 2-D pose estimation, which is trained on detailed annotations of endoscopic and microscopic data sets. Our model is formed by a fully convolutional detection-regression network. Joints and associations between joint pairs in our instrument model are located by the detection subnetwork and are subsequently refined through a regression subnetwork. Based on the output from the model, the poses of the instruments are inferred using maximum bipartite graph matching. Our estimation framework is powered by deep learning techniques without any direct kinematic information from a robot. Our framework is tested on single-instrument RMIT data, and also on multi-instrument EndoVis and in vivo data with promising results. In addition, the data set annotations are publicly released along with our code and model.
Xiaofei Du 0001, Thomas Kurmann, Ping-Lin Chang, Maximilian Allan, Sébastien Ourselin, Raphael Sznitman, John D. Kelly, Danail Stoyanov
IEEE Trans. Medical Imaging8
2017 Refractive Structure-from-Motion Through a Flat Refractive Interface
abstract
Recovering 3D scene geometry from underwater images involves the Refractive Structure-from-Motion (RSfM) problem, where the image distortions caused by light refraction at the interface between different propagation media invalidates the single view point assumption. Direct use of the pinhole camera model in RSfM leads to inaccurate camera pose estimation and consequently drift. RSfM methods have been thoroughly studied for the case of a thick glass interface that assumes two refractive interfaces between the camera and the viewed scene. On the other hand, when the camera lens is in direct contact with the water, there is only one refractive interface. By explicitly considering a refractive interface, we develop a succinct derivation of the refractive fundamental matrix in the form of the generalised epipolar constraint for an axial camera. We use the refractive fundamental matrix to refine initial pose estimates obtained by assuming the pinhole model. This strategy allows us to robustly estimate underwater camera poses, where other methods suffer from poor noise-sensitivity. We also formulate a new four view constraint enforcing camera pose consistency along a video which leads us to a novel RSfM framework. For validation we use synthetic data to show the numerical properties of our method and we provide results on real data to demonstrate performance within laboratory settings and for applications in endoscopy.
François Chadebecq, Francisco Vasconcelos 0001, George Dwyer, Rene M. Lacher, Sébastien Ourselin, Tom Vercauteren, Danail Stoyanov
ICCV7
2017 ToolNet: Holistically-nested real-time segmentation of robotic surgical tools
abstract
Real-time tool segmentation from endoscopic videos is an essential part of many computer-assisted robotic surgical systems and of critical importance in robotic surgical data science. We propose two novel deep learning architectures for automatic segmentation of non-rigid surgical instruments. Both methods take advantage of automated deep-learning-based multi-scale feature extraction while trying to maintain an accurate segmentation quality at all resolutions. The two proposed methods encode the multi-scale constraint inside the network architecture. The first proposed architecture enforces it by cascaded aggregation of predictions and the second proposed network does it by means of a holistically-nested architecture where the loss at each scale is taken into account for the optimization process. As the proposed methods are for real-time semantic labeling, both present a reduced number of parameters. We propose the use of parametric rectified linear units for semantic labeling in these small architectures to increase the regularization of the network while maintaining the segmentation accuracy. We compare the proposed architectures against state-of-the-art fully convolutional networks. We validate our methods using existing benchmark datasets, including ex vivo cases with phantom tissue and different robotic surgical instruments present in the scene. Our results show a statistically significant improved Dice Similarity Coefficient over previous instrument segmentation methods. We analyze our design choices and discuss the key drivers for improving accuracy.
Luis C. García-Peraza-Herrera, Wenqi Li 0001, Lucas Fidon, Caspar Gruijthuijsen, Alain Devreker, George Attilakos, Jan Deprest, Emmanuel B. Vander Poorten, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin
IROS9
2017 Body wall force sensor for simulated minimally invasive surgery: Application to fetal surgery
abstract
Surgical interventions are increasingly executed minimal invasively. Surgeons insert instruments through tiny incisions in the body and pivot slender instruments to treat organs or tissue below the surface. While a blessing for patients, surgeons need to pay extra attention to overcome the fulcrum effect, reduced haptic feedback and deal with lost hand-eye coordination. The mental load makes it difficult to pay sufficient attention to the forces that are exerted on the body wall. In delicate procedures such as fetal surgery, this might be problematic as irreparable damage could cause premature delivery. As a first attempt to quantify the interaction forces applied on the patient's body wall, a novel 6 degrees of freedom force sensor was developed for an ex-vivo set up. The performance of the sensor was characterised. User experiments were conducted by 3 clinicians on a set up simulating a fetal surgical intervention. During these simulated interventions, the interaction forces were recorded and analysed when a normal instrument was employed. These results were compared with a session where a flexible instrument under haptic guidance was used. The conducted experiments resulted in interesting insights in the interaction forces and stresses that develop during such difficult surgical intervention. The results also implicated that haptic guidance schemes and the use of flexible instruments rather than rigid ones could have a significant impact on the stresses that occur at the body wall.
Allan Javaux, Laure Esteveny, David Bouget, Caspar Gruijthuijsen, Danail Stoyanov, Tom Vercauteren, Sébastien Ourselin, Dominiek Reynaerts, Kathleen Denis, Jan Deprest, Emmanuel B. Vander Poorten
IROS5
2017 DejaVu: Intra-operative Simulation for Surgical Gesture Rehearsal
Nazim Haouchine, Danail Stoyanov, Frédérick Roy, Stephane Cotin
MICCAI (2)2
2017 Fast Estimation of Haemoglobin Concentration in Tissue Via Wavelet Decomposition
Geoffrey Jones, Neil Clancy, Xiaofei Du 0001, Maria Robu, Simon R. Arridge, Daniel S. Elson, Danail Stoyanov
MICCAI (2)7
2017 Simultaneous Recognition and Pose Estimation of Instruments in Minimally Invasive Surgery
Thomas Kurmann, Pablo Márquez-Neila, Xiaofei Du 0001, Pascal Fua, Danail Stoyanov, Sebastian Wolf 0005, Raphael Sznitman
MICCAI (2)5
2017 A Comparative Study of Breast Surface Reconstruction for Aesthetic Outcome Assessment
Rene M. Lacher, Francisco Vasconcelos 0001, David Bishop, Norman R. Williams, Mohammed Keshtgar, David J. Hawkes, John H. Hipwell, Danail Stoyanov
MICCAI (2)8
2017 Endoscopic Depth Measurement and Super-Spectral-Resolution Imaging
Jianyu Lin, Neil Clancy, Taran Tatla, Danail Stoyanov, Lena Maier-Hein, Daniel S. Elson
MICCAI (2)6
2017 Vision-based and marker-less surgical tool detection and tracking: a review of the literature
David Bouget, Maximilian Allan, Danail Stoyanov, Pierre Jannin
Medical Image Anal.3
2017 Comparative Validation of Polyp Detection Methods in Video Colonoscopy: Results From the MICCAI 2015 Endoscopic Vision Challenge
abstract
Colonoscopy is the gold standard for colon cancer screening though some polyps are still missed, thus preventing early disease detection and treatment. Several computational systems have been proposed to assist polyp detection during colonoscopy but so far without consistent evaluation. The lack of publicly available annotated databases has made it difficult to compare methods and to assess if they achieve performance levels acceptable for clinical use. The Automatic Polyp Detection sub-challenge, conducted as part of the Endoscopic Vision Challenge (http://endovis.grand-challenge.org) at the international conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) in 2015, was an effort to address this need. In this paper, we report the results of this comparative evaluation of polyp detection methods, as well as describe additional experiments to further explore differences between methods. We define performance metrics and provide evaluation databases that allow comparison of multiple methodologies. Results show that convolutional neural networks are the state of the art. Nevertheless, it is also demonstrated that combining different methodologies can lead to an improved overall performance.
Jorge Bernal, Nima Tajkbaksh, Francisco Javier Sánchez, Bogdan J. Matuszewski, Hao Chen 0011, Lequan Yu, Quentin Angermann, Olivier Romain, Bjorn Rustad, Ilangko Balasingham, Konstantin Pogorelov, Sungbin Choi, Quentin Debard, Lena Maier-Hein, Stefanie Speidel, Danail Stoyanov, Patrick Brandao, Henry Córdova, Cristina Sánchez-Montes, Suryakanth R. Gurudu, Gloria Fernández-Esparrach, Xavier Dray, Jianming Liang, Aymeric Histace
IEEE Trans. Medical Imaging16
2017 Bayesian Estimation of Intrinsic Tissue Oxygenation and Perfusion From RGB Images
abstract
Multispectral imaging (MSI) can potentially assist the intra-operative assessment of tissue structure, function and viability, by providing information about oxygenation. In this paper, we present a novel technique for recovering intrinsic MSI measurements from endoscopic RGB images without custom hardware adaptations. The advantage of this approach is that it requires no modification to existing surgical and diagnostic endoscopic imaging systems. Our method uses a radiometric color calibration of the endoscopic camera's sensor in conjunction with a Bayesian framework to recover a per-pixel measurement of the total blood volume (THb) and oxygen saturation (SO2) in the observed tissue. The sensor's pixel measurements are modeled as weighted sums over a mixture of Poisson distributions and we optimize the variables SO2and THb to maximize the likelihood of the observations. To validate our technique, we use synthetic images generated from Monte Carlo physics simulation of light transport through soft tissue containing sub-surface blood vessels. We also validate our method on in vivo data by comparing it to a MSI dataset acquired with a hardware system that sequentially images multiple spectral bands without overlap. Our results are promising and show that we are able to provide surgeons with additional relevant information by processing endoscopic images with our modeling and inference framework.
Geoffrey Jones, Neil Clancy, Yusuf Helo, Simon R. Arridge, Daniel S. Elson, Danail Stoyanov
IEEE Trans. Medical Imaging6
2016 Similarity Registration Problems for 2D/3D Ultrasound Calibration
Francisco Vasconcelos 0001, Donald Peebles, Sébastien Ourselin, Danail Stoyanov
ECCV (6)4
2016 Hand-eye calibration for robotic assisted minimally invasive surgery without a calibration object
abstract
In a robot mounted camera arrangement, hand-eye calibration estimates the rigid relationship between the robot and camera coordinate frames. Most hand-eye calibration techniques use a calibration object to estimate the relative transformation of the camera in several views of the calibration object and link these to the forward kinematics of the robot to compute the hand-eye transformation. Such approaches achieve good accuracy for general use but for applications such as robotic assisted minimally invasive surgery, acquiring a calibration sequence multiple times during a procedure is not practical. In this paper, we present a new approach to tackle the problem by using the robotic surgical instruments as the calibration object with well known geometry from CAD models used for manufacturing. Our approach removes the requirement of a custom sterile calibration object to be used in the operating room and it simplifies the process of acquiring calibration data when the laparoscope is constrained to move around a remote centre of motion. This is the first demonstration of the feasibility to perform hand-eye calibration using components of the robotic system itself and we show promising validation results on synthetic data as well as data acquired with the da Vinci Research Kit.
Krittin Pachtrachai, Maximilian Allan, Vijay Pawar, Stephen Hailes, Danail Stoyanov
IROS5
2016 Bilateral Weighted Adaptive Local Similarity Measure for Registration in Neurosurgery
Martin Kochan, Marc Modat, Tom Vercauteren, Mark White 0001, Laura Mancini, Gavin Winston, Andrew W. McEvoy, John S. Thornton, Tarek A. Yousry, John S. Duncan, Sébastien Ourselin, Danail Stoyanov
MICCAI (3)12
2016 Probe-Based Rapid Hybrid Hyperspectral and Tissue Surface Imaging Aided by Fully Convolutional Networks
abstract
Tissue surface shape and reflectance spectra provide rich intra-operative information useful in surgical guidance. We propose a hybrid system which displays an endoscopic image with a fast joint inspection of tissue surface shape using structured light (SL) and hyperspectral imaging (HSI). For SL a miniature fibre probe is used to project a coloured spot pattern onto the tissue surface. In HSI mode standard endoscopic illumination is used, with the fibre probe collecting reflected light and encoding the spatial information into a linear format that can be imaged onto the slit of a spectrograph. Correspondence between the arrangement of fibres at the distal and proximal ends of the bundle was found using spectral encoding. Then during pattern decoding, a fully convolutional network (FCN) was used for spot detection, followed by a matching propagation algorithm for spot identification. This method enabled fast reconstruction (12 frames per second) using a GPU. The hyperspectral image was combined with the white light image and the reconstructed surface, showing the spectral information of different areas. Validation of this system using phantom and ex vivo experiments has been demonstrated.
Jianyu Lin, Neil Clancy, Xueqing Sun, Mirek Janatka, Danail Stoyanov, Daniel S. Elson
MICCAI (3)6
2015 Fluidic actuation for intra-operative in situ imaging
abstract
A novel fluidic actuation system has been developed for in situ imaging of anatomic tissues. The actuator consists of a micromachined superelastic tool guide driven by a pair of pneumatic artificial muscles. Two additional working channels allow easy interchange of instruments or sensing equipment. This paper describes the design and construction of the actuation system. Experimental results are also reported indicating a bending repeatability of 0.1 degrees and an operational bandwidth exceeding 8Hz. To show-case the performance of the device, the actuator was loaded with an all-optical ultrasound imaging probe. First scanned images of human placental tissue surface using an all-optical ultrasound probe are presented. While a model has been developed to estimate the probe position in space as function of the input pressure, in future work, this model will be complemented with additional sensor measurements of the bending probe taking into account the hysteretic behaviour of both muscles and nitinol structure.
Alain Devreker, Benoit Rosa, Adrien E. Desjardins, Erwin J. Alles, Luis C. García-Peraza-Herrera, Efthymios Maneas, Danail Stoyanov, Anna L. David, Tom Vercauteren, Jan Deprest, Sébastien Ourselin, Dominiek Reynaerts, Emmanuel B. Vander Poorten
IROS7
2015 Image Based Surgical Instrument Pose Estimation with Multi-class Labelling and Optical Flow
Maximilian Allan, Ping-Lin Chang, Sébastien Ourselin, David J. Hawkes, Ashwin Sridhar, John D. Kelly, Danail Stoyanov
MICCAI (1)7
2015 Tissue Surface Reconstruction Aided by Local Normal Information Using a Self-calibrated Endoscopic Structured Light System
Jianyu Lin, Neil Clancy, Danail Stoyanov, Daniel S. Elson
MICCAI (1)3
2014 Comparative Validation of Single-Shot Optical Techniques for Laparoscopic 3-D Surface Reconstruction
abstract
Intra-operative imaging techniques for obtaining the shape and morphology of soft-tissue surfaces in vivo are a key enabling technology for advanced surgical systems. Different optical techniques for 3-D surface reconstruction in laparoscopy have been proposed, however, so far no quantitative and comparative validation has been performed. Furthermore, robustness of the methods to clinically important factors like smoke or bleeding has not yet been assessed. To address these issues, we have formed a joint international initiative with the aim of validating different state-of-the-art passive and active reconstruction methods in a comparative manner. In this comprehensive in vitro study, we investigated reconstruction accuracy using different organs with various shape and texture and also tested reconstruction robustness with respect to a number of factors like the pose of the endoscope as well as the amount of blood or smoke present in the scene. The study suggests complementary advantages of the different techniques with respect to accuracy, robustness, point density, hardware complexity and computation time. While reconstruction accuracy under ideal conditions was generally high, robustness is a remaining issue to be addressed. Future work should include sensor fusion and in vivo validation studies in a specific clinical context. To trigger further research in surface reconstruction, stereoscopic data of the study will be made publically available at www.open-CAS.com upon publication of the paper.
Lena Maier-Hein, Anja Groch, Adrien Bartoli, Sebastian Bodenstedt, G. Boissonnat, Ping-Lin Chang, Neil Clancy, Daniel S. Elson, Sven Haase, Eric Heim, Joachim Hornegger, Pierre Jannin, Hannes Kenngott, Thomas Kilgus, Beat P. Müller-Stich, D. Oladokun, Sebastian Röhl, Thiago R. dos Santos, Heinz-Peter Schlemmer, Alexander Seitel, Stefanie Speidel, Martin Wagner 0001, Danail Stoyanov
IEEE Trans. Medical Imaging23
2013 Real-Time Dense Stereo Reconstruction Using Convex Optimisation with a Cost-Volume for Image-Guided Robotic Surgery
Ping-Lin Chang, Danail Stoyanov, Andrew J. Davison, Philip J. Edwards
MICCAI (1)2
2013 Optical techniques for 3D surface reconstruction in computer-assisted laparoscopic surgery
Lena Maier-Hein, Peter Mountney, Adrien Bartoli, Haytham Elhawary, Daniel S. Elson, Anja Groch, Andreas Kolb 0001, Marcos A. Rodrigues 0001, Jonathan M. Sorger, Stefanie Speidel, Danail Stoyanov
Medical Image Anal.11
2012 Metric depth recovery from monocular images using Shape-from-Shading and specularities
abstract
Despite recent advances in modeling the Shape-from-Shading (SFS) problem and its numerical solution, practical applications have been limited. This is primarily due to the lack of perspective SFS models without the assumption of a light source at the camera centre and the non-metric spatial localisation of the reconstructed shape. In this work, we propose a novel formulation of the SFS problem that allows the reconstruction of surfaces lit by a near point light source away from the camera centre. We also show how knowledge of the light source position can enable the recovery of depth information in a metric space by triangulating specular highlights. Validation of the proposed technique is reported on synthetic and endoscopic images.
Marco Visentini Scarzanella, Danail Stoyanov, Guang-Zhong Yang
ICIP2
2012 Stereoscopic Scene Flow for Robotic Assisted Minimally Invasive Surgery
Danail Stoyanov
MICCAI (1)1
2011 Dense Surface Reconstruction for Enhanced Navigation in MIS
Johannes Totz, Peter Mountney, Danail Stoyanov, Guang-Zhong Yang
MICCAI (1)3
2010 Dynamic Guidance for Robotic Surgery Using Image-Constrained Biomechanical Models
Philip Pratt, Danail Stoyanov, Marco Visentini Scarzanella, Guang-Zhong Yang
MICCAI (1)2
2010 Real-Time Stereo Reconstruction in Robotically Assisted Minimally Invasive Surgery
Danail Stoyanov, Marco Visentini Scarzanella, Philip Pratt, Guang-Zhong Yang
MICCAI (1)1
2010 Tracking of Irregular Graphical Structures for Tissue Deformation Recovery in Minimally Invasive Surgery
Marco Visentini Scarzanella, Robert D. Merrifield, Danail Stoyanov, Guang-Zhong Yang
MICCAI (3)3
2009 Illumination position estimation for 3D soft-tissue reconstruction in robotic minimally invasive surgery
abstract
For robotic assisted minimally invasive surgery, recovering the 3D soft-tissue shape and morphology in vivo is important for providing image-guidance, motion compensation and applying dynamic active constraints. In this paper, we propose a practical method for calibrating the illumination source position in monocular and stereoscopic laparoscopes. The method relies on using the geometric constraints from specular reflections obtained during the laparoscope camera calibration process. By estimating the light source position, the method forgoes the common assumption of coincidence with the camera centre and can be used to obtain constraints on the normal of the surface geometry during surgery from specularities. We demonstrate the effectiveness of the proposed approach with numerical simulations and by qualitative analysis of real stereo-laparoscopic calibrations.
Danail Stoyanov, Daniel S. Elson, Guang-Zhong Yang
IROS1
2009 i-BRUSH: A Gaze-Contingent Virtual Paintbrush for Dense 3D Reconstruction in Robotic Assisted Surgery
Marco Visentini Scarzanella, George P. Mylonas, Danail Stoyanov, Guang-Zhong Yang
MICCAI (1)3
2008 Belief Propagation for Depth Cue Fusion in Minimally Invasive Surgery
Benny P. L. Lo, Marco Visentini Scarzanella, Danail Stoyanov, Guang-Zhong Yang
MICCAI (2)3
2008 Gaze-Contingent 3D Control for Focused Energy Ablation in Robotic Assisted Surgery
Danail Stoyanov, George P. Mylonas, Guang-Zhong Yang
MICCAI (2)1
2007 A Probabilistic Framework for Tracking Deformable Soft Tissue in Minimally Invasive Surgery
Peter Mountney, Benny P. L. Lo, Surapa Thiemjarus, Danail Stoyanov, Guang-Zhong Yang
MICCAI (2)4
2007 Assessment of Perceptual Quality for Gaze-Contingent Motion Stabilization in Robotic Assisted Minimally Invasive Surgery
George P. Mylonas, Danail Stoyanov, Ara Darzi, Guang-Zhong Yang
MICCAI (2)2
2007 Stabilization of Image Motion for Robotic Assisted Beating Heart Surgery
Danail Stoyanov, Guang-Zhong Yang
MICCAI (1)1
2006 Simultaneous Stereoscope Localization and Soft-Tissue Mapping for Minimal Invasive Surgery
Peter Mountney, Danail Stoyanov, Andrew J. Davison, Guang-Zhong Yang
MICCAI (1)2
2005 Removing specular reflection components for robotic assisted laparoscopic surgery
abstract
In this paper, we propose a practical method for removing specular artifacts on the epicardial surface of the heart in robotic laparoscopic surgery while preserving the underlying image structure. We use freeform temporal registration of the non-rigid surface motion to recover chromatic information saturated by highlights. The diffuse and specular image components are then separated by shifting pixel intensities with respect to chromaticity gathered from the spatio-temporal volume. Results on in vivo data and reconstructions of 3D structure from the diffuse images show the potential value of the technique.
Danail Stoyanov, Guang-Zhong Yang
ICIP (3)1
2005 Gaze-Contingent Soft Tissue Deformation Tracking for Minimally Invasive Robotic Surgery
George P. Mylonas, Danail Stoyanov, Fani Deligianni, Ara Darzi, Guang-Zhong Yang
MICCAI2
2005 Laparoscope Self-calibration for Robotic Assisted Minimally Invasive Surgery
Danail Stoyanov, Ara Darzi, Guang-Zhong Yang
MICCAI (2)1
2005 Soft-Tissue Motion Tracking and Structure Estimation for Robotic Assisted MIS Procedures
Danail Stoyanov, George P. Mylonas, Fani Deligianni, Ara Darzi, Guang-Zhong Yang
MICCAI (2)1
2004 Dense 3D Depth Recovery for Soft Tissue Deformation During Robotically Assisted Laparoscopic Surgery
Danail Stoyanov, Ara Darzi, Guang-Zhong Yang
MICCAI (2)1
2003 Current Issues of Photorealistic Rendering for Virtual and Augmented Reality in Minimally Invasive Surgery
abstract
In surgery, virtual and augmented reality are increasingly being used as new ways of training, preoperative planning, diagnosis and surgical navigation. Further development of virtual and augmented reality in medicine is moving towards photorealistic rendering and patient specific modeling, permitting high fidelity visual examination and user interaction. This coincides with the current development in computer vision and graphics where image information is used directly to render novel views of a scene. These techniques require extensive use of geometric information about the scene and provide a comprehensive review of the underlying techniques required for building patient specific models with photorealstic rendering. It also highlights some of the opportunities that image based modeling and rendering techniques can offer in the context of minimally invasive surgery.
Danail Stoyanov, Mohamed A. ElHelw, Benny P. L. Lo, Adrian James Chung, Fernando Bello, Guang-Zhong Yang
IV1