EDBT 2026 Demo / reviewers in the wild / expert
Matthew J. Clarkson
dblp:79/140 · also Matthew John Clarkson
· DBLP profile ↗
37ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-5565-1252ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Surgical AI Copilot: Energy-Based Fourier Gradient Low-Rank Adaptation for Surgical LLM Agent Reasoning and PlanningabstractImage-guided surgery demands adaptive, real-time decision support, yet static AI models struggle with structured task planning and providing interactive guidance. Large language models (LLMs)-powered agents offer a promising solution by enabling dynamic task planning and predictive decision support. Despite recent advances, the absence of surgical agent datasets and robust parameter-efficient fine-tuning techniques limits the development of LLM agents capable of complex intraoperative reasoning. In this paper, we introduce Surgical AI Copilot, an LLM agent for image-guided pituitary surgery, capable of conversation, planning, and task execution in response to queries involving tasks such as MRI tumor segmentation, endoscope anatomy segmentation, overlaying preoperative imaging with intraoperative views, instrument tracking, and surgical visual question answering (VQA). To enable structured agent planning, we develop the PitAgent dataset, a surgical context-aware planning dataset covering surgical tasks like workflow analysis, instrument localization, anatomical segmentation, and query-based reasoning. Additionally, we propose DEFT-GaLore, a Deterministic Energy-based Fourier Transform (DEFT) gradient projection technique for efficient low-rank adaptation of recent LLMs (e.g., LLaMA 3.2, Qwen 2.5), enabling their use as surgical agent planners. We extensively validate our agent's performance and the proposed adaptation technique against other state-of-the-art low-rank adaptation methods on agent planning and prompt generation tasks, including a zero-shot surgical VQA benchmark, demonstrating the significant potential for truly efficient and scalable surgical LLM agents in real-time operative settings. Jiayuan Huang, Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Danail Stoyanov, Hani J. Marcus, Linzhe Jiang, Matthew J. Clarkson, Mobarak I. Hoque |
AAAI | 8 |
| 2026 | PitVQA++: Vector Matrix-Low-Rank Adaptation for Open-Ended Visual Question Answering in Pituitary SurgeryabstractVision-Language Models (VLMs) in visual question answering (VQA) offer a unique opportunity to enhance intra-operative decision-making, promote intuitive interactions, and significantly advance surgical education. However, the development of VLMs for surgical VQA is challenging due to limited datasets and the risk of overfitting and catastrophic forgetting during full fine-tuning of pretrained weights. While parameter-efficient techniques like Low-Rank Adaptation (LoRA) and Matrix of Rank Adaptation (MoRA) address adaptation challenges, their uniform parameter distribution overlooks the feature hierarchy in deep networks, where earlier layers, that learn general features, require more parameters than later ones. This work introduces PitVQA++ with an Open-ended PitVQA dataset and vector matrix-low-rank adaptation (Vector-MoLoRA), an innovative VLM fine-tuning approach for adapting GPT-2 to pituitary surgery. Open-Ended PitVQA comprises 109,173 frames from 25 procedural videos with 795,270 question-answer sentence pairs, covering key surgical elements such as phase and step recognition, context understanding, tool detection, localization, and interactions recognition. Vector-MoLoRA incorporates the principles of LoRA and MoRA to develop a matrix-low-rank adaptation strategy that employs rank vectors to allocate more parameters to earlier layers, gradually reducing them in the later layers. Our approach, validated on the Open-Ended PitVQA and EndoVis18-VQA datasets, effectively mitigates catastrophic forgetting while significantly enhancing performance over recent baselines. Performance-rejection analysis further highlights Vector-MoLoRA's enhanced reliability and trustworthiness in handling uncertain predictions. Our source code and dataset is available at https://github.com/HRL-Mike/PitVQA-Plus. Runlong He, Danyal Z. Khan, Evangelos B. Mazomenos, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarak I. Hoque |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Endo-FASt3r: Endoscopic Foundation Model Adaptation for Structure from MotionabstractAccurate depth and camera pose estimation is essential for achieving high-quality 3D visualisations in robotic-assisted surgery. Despite recent advancements in foundation model adaptation to monocular depth estimation of endoscopic scenes via self-supervised learning (SSL), no prior work has explored their use for pose estimation. These methods rely on low rank-based adaptation approaches, which constrain model updates to a low-rank space. We propose Endo-FASt3r, the first monocular SSL depth and pose estimation framework that uses foundation models for both tasks. We extend the Reloc3r relative pose estimation foundation model by designing Reloc3rX, introducing modifications necessary for convergence in SSL. We also present DoMoRA, a novel adaptation technique that enables higher-rank updates and faster convergence. Experiments on the SCARED dataset show that Endo-FASt3r achieves a substantial $$10\%$$ improvement in pose estimation and a $$2\%$$ improvement in depth estimation over prior work. Similar performance gains on the Hamlyn and StereoMIS datasets reinforce the generalisability of Endo-FASt3r across different datasets. Our code is available at: https://github.com/Mona-ShZeinoddin/Endo_FASt3r.git . Mona Sheikh Zeinoddin, Mobarak I. Hoque, Zafer Tandogdu, Greg Shaw, Matthew J. Clarkson, Evangelos B. Mazomenos, Danail Stoyanov |
MICCAI (11) | 5 |
| 2025 | An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challengeabstractAugmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain. Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli |
Medical Image Anal. | 9 |
| 2025 | Controllable illumination invariant GAN for diverse temporally-consistent surgical video synthesis
Long Chen 0019, Mobarak I. Hoque, Zhe Min, Matthew J. Clarkson, Thomas Dowrick |
Medical Image Anal. | 4 |
| 2025 | Competing for Pixels: A Self-Play Algorithm for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS) methods, reliant on image-level labels indicating object presence, lack explicit correspondence between labels and regions of interest (ROIs), posing a significant challenge. Despite this, WSSS methods have attracted attention due to their much lower annotation costs compared to fully-supervised segmentation. Leveraging reinforcement learning (RL) self-play, we propose a novel WSSS method that gamifies image segmentation of a ROI. We formulate segmentation as a competition between two agents that compete to select ROI-containing patches until exhaustion of all such patches. The score at each time-step, used to compute the reward for agent training, represents likelihood of object presence within the selection, determined by an object presence detector pre-trained using only image-level binary classification labels of object presence. Additionally, we propose a game termination condition that can be called by either side upon exhaustion of all ROI-containing patches, followed by the selection of a final patch from each. Upon termination, the agent is incentivised if ROI-containing patches are exhausted or disincentivised if a ROI-containing patch is found by the competitor. This competitive setup ensures minimisation of over- or under-segmentation, a common problem with WSSS methods. Extensive experimentation across four datasets demonstrates significant performance improvements over recent state-of-the-art methods. Shaheer U. Saeed, Shiqi Huang 0001, João Ramalhinho, Iani J. M. B. Gayo, Nina Montaña Brown, Ester Bonmati, Stephen P. Pereira, Brian R. Davidson, Dean C. Barratt, Matthew J. Clarkson, Yipeng Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2024 | HUP-3D: A 3D Multi-view Synthetic Dataset for Assisted-Egocentric Hand-Ultrasound-Probe Pose EstimationabstractWe present HUP-3D, a 3D multiview multimodal synthetic dataset for hand ultrasound (US) probe pose estimation in the context of obstetric ultrasound. Egocentric markerless 3D joint pose estimation has potential applications in mixed reality medical education. The ability to understand hand and probe movements opens the door to tailored guidance and mentoring applications. Our dataset consists of over 31k sets of RGB, depth, and segmentation mask frames, including pose-related reference data, with an emphasis on image diversity and complexity. Adopting a camera viewpoint-based sphere concept allows us to capture a variety of views and generate multiple hand grasps poses using a pre-trained network. Additionally, our approach includes a software-based image rendering concept, enhancing diversity with various hand and arm textures, lighting conditions, and background images. We validated our proposed dataset with state-of-the-art learning models and we obtained the lowest hand-object keypoint errors. The supplementary material details the parameters for sphere-based camera view angles and the grasp generation and rendering pipeline configuration. The source code for our grasp generation and rendering pipeline, along with the dataset, is publicly available at https://manuelbirlo.github.io/HUP-3D/ . Manuel Birlo, Razvan Caramalau, Philip J. Edwards, Brian Dromey, Matthew J. Clarkson, Danail Stoyanov |
MICCAI (1) | 5 |
| 2024 | PitVQA: Image-Grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
Runlong He, Mengya Xu, Adrito Das, Danyal Z. Khan, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarakol Islam |
MICCAI (6) | 8 |
| 2024 | Nonrigid Reconstruction of Freehand Ultrasound Without a TrackerabstractReconstructing 2D freehand Ultrasound (US) frames into 3D space without using a tracker has recently seen advances with deep learning. Predicting good frame-to-frame rigid transformations is often accepted as the learning objective, especially when the ground-truth labels from spatial tracking devices are inherently rigid transformations. Motivated by a) the observed nonrigid deformation due to soft tissue motion during scanning, and b) the highly sensitive prediction of rigid transformation, this study investigates the methods and their benefits in predicting nonrigid transformations for reconstructing 3D US. We propose a novel co-optimisation algorithm for simultaneously estimating rigid transformations among US frames, supervised by ground-truth from a tracker, and a nonrigid deformation, optimised by a regularised registration network. We show that these two objectives can be either optimised using meta-learning or combined by weighting. A fast scattered data interpolation is also developed for enabling frequent reconstruction and registration of non-parallel US frames, during training. With a new data set containing over 357,000 frames in 720 scans, acquired from 60 subjects, the experiments demonstrate that, due to an expanded thus easier-to-optimise solution space, the generalisation is improved with the added deformation estimation, with respect to the rigid ground-truth. The global pixel reconstruction error (assessing accumulative prediction) is lowered from 18.48 to 16.51 mm, compared with baseline rigid-transformation-predicting methods. Using manually identified landmarks, the proposed co-optimisation also shows potentials in compensating nonrigid tissue motion at inference, which is not measurable by tracker-provided ground-truth. The code and data used in this paper are made publicly available at https://github.com/QiLi111/NR-Rec-FUS . Qi Li 0030, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, Yipeng Hu |
MICCAI (4) | 5 |
| 2024 | Can engineers represent surgeons in usability studies? Comparison of results from evaluating augmented reality guidance for laparoscopic surgeryabstractObtaining feedback from time-constrained end-users is a major challenge in evaluating novel systems for specialised applications. The performance and feedback of engineers and surgeons was evaluated through an experiment where participants were asked to identify tumour locations within an anatomically realistic silicon liver model across three different conditions of an Augmented Reality (AR) prototype system (Baseline, Split AR and Full AR). Our findings show that engineers and surgeons share some similarities in their performance, feedback and behaviour, particularly when reliance on the AR system is high for both groups. However, engineers typically focus more on accuracy of the image alignment and are more accurate in their responses when supported by AR. Senior surgeons typically perform faster and use AR as supplementary information, while the performance of junior surgeons is more closely aligned to the performance of engineers. We conclude that engineers could be involved in preliminary evaluations of a surgical system or in evaluations of systems which are aimed at training junior surgeons, but that it is essential to involve surgeons in later evaluations, where ecological validity is a more important consideration. SooJeong Yoo, João Ramalhinho, Thomas Dowrick, Murali Somasundaram, Kurinchi Gurusamy, Brian R. Davidson, Matthew J. Clarkson, Ann Blandford |
Comput. Graph. | 7 |
| 2024 | Active learning using adaptable task-based prioritisationabstractSupervised machine learning-based medical image computing applications necessitate expert label curation, while unlabelled image data might be relatively abundant. Active learning methods aim to prioritise a subset of available image data for expert annotation, for label-efficient model training. We develop a controller neural network that measures priority of images in a sequence of batches, as in batch-mode active learning, for multi-class segmentation tasks. The controller is optimised by rewarding positive task-specific performance gain, within a Markov decision process (MDP) environment that also optimises the task predictor. In this work, the task predictor is a segmentation network. A meta-reinforcement learning algorithm is proposed with multiple MDPs, such that the pre-trained controller can be adapted to a new MDP that contains data from different institutes and/or requires segmentation of different organs or structures within the abdomen. We present experimental results using multiple CT datasets from more than one thousand patients, with segmentation tasks of nine different abdominal organs, to demonstrate the efficacy of the learnt prioritisation controller function and its cross-institute and cross-organ adaptability. We show that the proposed adaptable prioritisation metric yields converging segmentation accuracy for a new kidney segmentation task, unseen in training, using between approximately 40% to 60% of labels otherwise required with other heuristic or random prioritisation metrics. For clinical datasets of limited size, the proposed adaptable prioritisation offers a performance improvement of 22.6% and 10.2% in Dice score, for tasks of kidney and liver vessel segmentation, respectively, compared to random prioritisation and alternative active sampling strategies. Shaheer U. Saeed, João Ramalhinho, Mark A. Pinnock, Ziyi Shen, Yunguan Fu, Nina Montaña Brown, Ester Bonmati, Dean C. Barratt, Stephen P. Pereira, Brian R. Davidson, Matthew J. Clarkson, Yipeng Hu |
Medical Image Anal. | 11 |
| 2023 | SARAMIS: Simulation Assets for Robotic Assisted and Minimally Invasive SurgeryabstractMinimally-invasive surgery (MIS) and robot-assisted minimally invasive (RAMIS) surgery offer well-documented benefits to patients such as reduced post-operative pain and shorter hospital stays.However, the automation of MIS and RAMIS through the use of AI has been slow due to difficulties in data acquisition and curation, partially caused by the ethical considerations of training, testing and deploying AI models in medical environments.We introduce \texttt{SARAMIS}, the first large-scale dataset of anatomically derived 3D rendering assets of the human abdominal anatomy.Using previously existing, open-source CT datasets of the human anatomy, we derive novel 3D meshes, tetrahedral volumes, textures and diffuse maps for over 104 different anatomical targets in the human body, representing the largest, open-source dataset of 3D rendering assets for synthetic simulation of vision tasks in MIS+RAMIS, increasing the availability of openly available 3D meshes in the literature by three orders of magnitude.We supplement our dataset with a series of GPU-enabled rendering environments, which can be used to generate datasets for realistic MIS/RAMIS tasks.Finally, we present an example of the use of \texttt{SARAMIS} assets for an autonomous navigation task in colonoscopy from CT abdomen-pelvis scans for the first time in the literature.\texttt{SARAMIS} is publically made available at https://github.com/NMontanaBrown/saramis/, with assets released under a CC-BY-NC-SA license. Nina Montaña Brown, Shaheer U. Saeed, Ahmed Abdulaal, Thomas Dowrick, Yakup Kilic, Sophie Wilkinson, Jack Gao, Meghavi Mashar, Chloe He 0002, Alkisti Stavropoulou, Emma Thomson, Zachary Baum, Simone Foti, Brian R. Davidson, Yipeng Hu, Matthew J. Clarkson |
NeurIPS | 16 |
| 2023 | 3D Generative Model Latent Disentanglement via Local EigenprojectionabstractDesigning realistic digital humans is extremely complex. Most data-driven generative models used to simplify the creation of their underlying geometric shape do not offer control over the generation of local shape attributes. In this paper, we overcome this limitation by introducing a novel loss function grounded in spectral geometry and applicable to different neural-network-based generative models of 3D head and body meshes. Encouraging the latent variables of mesh variational autoencoders (VAEs) or generative adversarial networks (GANs) to follow the local eigenprojections of identity attributes, we improve latent disentanglement and properly decouple the attribute creation. Experimental results show that our local eigenprojection disentangled (LED) models not only offer improved disentanglement with respect to the state-of-the-art, but also maintain good generation capabilities with training times comparable to the vanilla implementations of the models. Our code and pre-trained models are available at github.com/simofoti/LocalEigenprojDisentangled. Simone Foti, Bongjin Koo, Danail Stoyanov, Matthew J. Clarkson |
Comput. Graph. Forum | 4 |
| 2023 | Prototypical few-shot segmentation for cross-institution male pelvic structures with spatial registrationabstractThe prowess that makes few-shot learning desirable in medical image analysis is the efficient use of the support image data, which are labelled to classify or segment new classes, a task that otherwise requires substantially more training images and expert annotations. This work describes a fully 3D prototypical few-shot segmentation algorithm, such that the trained networks can be effectively adapted to clinically interesting structures that are absent in training, using only a few labelled images from a different institute. First, to compensate for the widely recognised spatial variability between institutions in episodic adaptation of novel classes, a novel spatial registration mechanism is integrated into prototypical learning, consisting of a segmentation head and an spatial alignment module. Second, to assist the training with observed imperfect alignment, support mask conditioning module is proposed to further utilise the annotation available from the support images. Extensive experiments are presented in an application of segmenting eight anatomical structures important for interventional planning, using a data set of 589 pelvic T2-weighted MR images, acquired at seven institutes. The results demonstrate the efficacy in each of the 3D formulation, the spatial registration, and the support mask conditioning, all of which made positive contributions independently or collectively. Compared with the previously proposed 2D alternatives, the few-shot segmentation performance was improved with statistical significance, regardless whether the support data come from the same or different institutes. Yunguan Fu, Iani J. M. B. Gayo, Qianye Yang, Zhe Min, Shaheer U. Saeed, Wen Yan 0005, J. Alison Noble, Mark Emberton, Matthew J. Clarkson, Henkjan J. Huisman, Dean C. Barratt, Victor Adrian Prisacariu, Yipeng Hu |
Medical Image Anal. | 11 |
| 2023 | The value of Augmented Reality in surgery - A usability study on laparoscopic liver surgeryabstractAugmented Reality (AR) is considered to be a promising technology for the guidance of laparoscopic liver surgery. By overlaying pre-operative 3D information of the liver and internal blood vessels on the laparoscopic view, surgeons can better understand the location of critical structures. In an effort to enable AR, several authors have focused on the development of methods to obtain an accurate alignment between the laparoscopic video image and the pre-operative 3D data of the liver, without assessing the benefit that the resulting overlay can provide during surgery. In this paper, we present a study that aims to assess quantitatively and qualitatively the value of an AR overlay in laparoscopic surgery during a simulated surgical task on a phantom setup. We design a study where participants are asked to physically localise pre-operative tumours in a liver phantom using three image guidance conditions - a baseline condition without any image guidance, a condition where the 3D surfaces of the liver are aligned to the video and displayed on a black background, and a condition where video see-through AR is displayed on the laparoscopic video. Using data collected from a cohort of 24 participants which include 12 surgeons, we observe that compared to the baseline, AR decreases the median localisation error of surgeons on non-peripheral targets from 25.8 mm to 9.2 mm. Using subjective feedback, we also identify that AR introduces usability improvements in the surgical task and increases the perceived confidence of the users. Between the two tested displays, the majority of participants preferred to use the AR overlay instead of navigated view of the 3D surfaces on a separate screen. We conclude that AR has the potential to improve performance and decision making in laparoscopic surgery, and that improvements in overlay alignment accuracy and depth perception should be pursued in the future. João Ramalhinho, SooJeong Yoo, Thomas Dowrick, Bongjin Koo, Murali Somasundaram, Kurinchi Gurusamy, David J. Hawkes, Brian R. Davidson, Ann Blandford, Matthew J. Clarkson |
Medical Image Anal. | 10 |
| 2022 | 3D Shape Variational Autoencoder Latent Disentanglement via Mini-Batch Feature Swapping for Bodies and FacesabstractLearning a disentangled, interpretable, and structured latent representation in 3D generative models of faces and bodies is still an open problem. The problem is particularly acute when control over identity features is required. In this paper, we propose an intuitive yet effective self-supervised approach to train a 3D shape variational autoencoder (VAE) which encourages a disentangled latent representation of identity features. Curating the mini-batch generation by swapping arbitrary features across different shapes allows to define a loss function leveraging known differences and similarities in the latent representations. Experimental results conducted on 3D meshes show that state-of-the-art methods for latent disentanglement are not able to disentangle identity features of faces and bodies. Our proposed method properly decouples the generation of such features while maintaining good representation and reconstruction capabilities. Our code and pretrained models are available at github.com/simofoti/3DVAE-SwapDisentangled. Simone Foti, Bongjin Koo, Danail Stoyanov, Matthew J. Clarkson |
CVPR | 4 |
| 2022 | Utility of optical see-through head mounted displays in augmented reality-assisted surgery: A systematic reviewabstractThis article presents a systematic review of optical see-through head mounted display (OST-HMD) usage in augmented reality (AR) surgery applications from 2013 to 2020. Articles were categorised by: OST-HMD device, surgical speciality, surgical application context, visualisation content, experimental design and evaluation, accuracy and human factors of human-computer interaction. 91 articles fulfilled all inclusion criteria. Some clear trends emerge. The Microsoft HoloLens increasingly dominates the field, with orthopaedic surgery being the most popular application (28.6%). By far the most common surgical context is surgical guidance (n=58) and segmented preoperative models dominate visualisation (n=40). Experiments mainly involve phantoms (n=43) or system setup (n=21), with patient case studies ranking third (n=19), reflecting the comparative infancy of the field. Experiments cover issues from registration to perception with very different accuracy results. Human factors emerge as significant to OST-HMD utility. Some factors are addressed by the systems proposed, such as attention shift away from the surgical site and mental mapping of 2D images to 3D patient anatomy. Other persistent human factors remain or are caused by OST-HMD solutions, including ease of use, comfort and spatial perception issues. The significant upward trend in published articles is clear, but such devices are not yet established in the operating room and clinical studies showing benefit are lacking. A focused effort addressing technical registration and perceptual factors in the lab coupled with design that incorporates human factors considerations to solve clear clinical problems should ensure that the significant current research efforts will succeed. Manuel Birlo, Philip J. Edwards, Matthew J. Clarkson, Danail Stoyanov |
Medical Image Anal. | 3 |
| 2022 | Gesture Recognition in Robotic Surgery With Multimodal AttentionabstractAutomatically recognising surgical gestures from surgical data is an important building block of automated activity recognition and analytics, technical skill assessment, intra-operative assistance and eventually robotic automation. The complexity of articulated instrument trajectories and the inherent variability due to surgical style and patient anatomy make analysis and fine-grained segmentation of surgical motion patterns from robot kinematics alone very difficult. Surgical video provides crucial information from the surgical site with context for the kinematic data and the interaction between the instruments and tissue. Yet sensor fusion between the robot data and surgical video stream is non-trivial because the data have different frequency, dimensions and discriminative capability. In this paper, we integrate multimodal attention mechanisms in a two-stream temporal convolutional network to compute relevance scores and weight kinematic and visual feature representations dynamically in time, aiming to aid multimodal network training and achieve effective sensor fusion. We report the results of our system on the JIGSAWS benchmark dataset and on a new in vivo dataset of suturing segments from robotic prostatectomy procedures. Our results are promising and obtain multimodal prediction sequences with higher accuracy and better temporal structure than corresponding unimodal solutions. Visualization of attention scores also gives physically interpretable insights on network understanding of strengths and weaknesses of each sensor. Beatrice van Amsterdam, Isabel Funke, Philip J. Edwards, Stefanie Speidel, Justin Collins, Ashwin Sridhar, John D. Kelly, Matthew J. Clarkson, Danail Stoyanov |
IEEE Trans. Medical Imaging | 8 |
| 2022 | Voice-Assisted Image Labeling for Endoscopic Ultrasound Classification Using Neural NetworksabstractUltrasound imaging is a commonly used technology for visualising patient anatomy in real-time during diagnostic and therapeutic procedures. High operator dependency and low reproducibility make ultrasound imaging and interpretation challenging with a steep learning curve. Automatic image classification using deep learning has the potential to overcome some of these challenges by supporting ultrasound training in novices, as well as aiding ultrasound image interpretation in patient with complex pathology for more experienced practitioners. However, the use of deep learning methods requires a large amount of data in order to provide accurate results. Labelling large ultrasound datasets is a challenging task because labels are retrospectively assigned to 2D images without the 3D spatial context available in vivo or that would be inferred while visually tracking structures between frames during the procedure. In this work, we propose a multi-modal convolutional neural network (CNN) architecture that labels endoscopic ultrasound (EUS) images from raw verbal comments provided by a clinician during the procedure. We use a CNN composed of two branches, one for voice data and another for image data, which are joined to predict image labels from the spoken names of anatomical landmarks. The network was trained using recorded verbal comments from expert operators. Our results show a prediction accuracy of 76% at image level on a dataset with 5 different labels. We conclude that the addition of spoken commentaries can increase the performance of ultrasound image classification, and eliminate the burden of manually labelling large EUS datasets necessary for deep learning applications. Ester Bonmati, Yipeng Hu, Alex Grimwood, Gavin J. Johnson, George Goodchild, Margaret G. Keane, Kurinchi Gurusamy, Brian R. Davidson, Matthew J. Clarkson, Stephen P. Pereira, Dean C. Barratt |
IEEE Trans. Medical Imaging | 9 |
| 2022 | Cross-Modality Image Registration Using a Training-Time Privileged Third ModalityabstractIn this work, we consider the task of pairwise cross-modality image registration, which may benefit from exploiting additional images available only at training time from an additional modality that is different to those being registered. As an example, we focus on aligning intra-subject multiparametric Magnetic Resonance (mpMR) images, between T2-weighted (T2w) scans and diffusion-weighted scans with high b-value (DWI [Formula: see text]). For the application of localising tumours in mpMR images, diffusion scans with zero b-value (DWI [Formula: see text]) are considered easier to register to T2w due to the availability of corresponding features. We propose a learning from privileged modality algorithm, using a training-only imaging modality DWI [Formula: see text], to support the challenging multi-modality registration problems. We present experimental results based on 369 sets of 3D multiparametric MRI images from 356 prostate cancer patients and report, with statistical significance, a lowered median target registration error of 4.34 mm, when registering the holdout DWI [Formula: see text] and T2w image pairs, compared with that of 7.96 mm before registration. Results also show that the proposed learning-based registration networks enabled efficient registration with comparable or better accuracy, compared with a classical iterative algorithm and other tested learning-based methods with/without the additional modality. These compared algorithms also failed to produce any significantly improved alignment between DWI [Formula: see text] and T2w in this challenging application. Qianye Yang, David Atkinson, Yunguan Fu, Tom Syer, Wen Yan 0005, Shonit Punwani, Matthew J. Clarkson, Dean C. Barratt, Tom Vercauteren, Yipeng Hu |
IEEE Trans. Medical Imaging | 7 |
| 2021 | Registration of Untracked 2D Laparoscopic Ultrasound to CT Images of the Liver Using Multi-Labelled Content-Based Image RetrievalabstractLaparoscopic Ultrasound (LUS) is recommended as a standard-of-care when performing laparoscopic liver resections as it images sub-surface structures such as tumours and major vessels. Given that LUS probes are difficult to handle and some tumours are iso-echoic, registration of LUS images to a pre-operative CT has been proposed as an image-guidance method. This registration problem is particularly challenging due to the small field of view of LUS, and usually depends on both a manual initialisation and tracking to compose a volume, hindering clinical translation. In this paper, we extend a proposed registration approach using Content-Based Image Retrieval (CBIR), removing the requirement for tracking or manual initialisation. Pre-operatively, a set of possible LUS planes is simulated from CT and a descriptor generated for each image. Then, a Bayesian framework is employed to estimate the most likely sequence of CT simulations that matches a series of LUS images. We extend our CBIR formulation to use multiple labelled objects and constrain the registration by separating liver vessels into portal vein and hepatic vein branches. The value of this new labeled approach is demonstrated in retrospective data from 5 patients. Results show that, by including a series of 5 untracked images in time, a single LUS image can be registered with accuracies ranging from 5.7 to 16.4 mm with a success rate of 78%. Initialisation of the LUS to CT registration with the proposed framework could potentially enable the clinical translation of these image fusion techniques. João Ramalhinho, Henry F. J. Tregidgo, Kurinchi Gurusamy, David J. Hawkes, Brian R. Davidson, Matthew J. Clarkson |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Zero-Shot Super-Resolution With a Physically-Motivated Downsampling Kernel for EndomicroscopyabstractSuper-resolution (SR) methods have seen significant advances thanks to the development of convolutional neural networks (CNNs). CNNs have been successfully employed to improve the quality of endomicroscopy imaging. Yet, the inherent limitation of research on SR in endomicroscopy remains the lack of ground truth high-resolution (HR) images, commonly used for both supervised training and reference-based image quality assessment (IQA). Therefore, alternative methods, such as unsupervised SR are being explored. To address the need for non-reference image quality improvement, we designed a novel zero-shot super-resolution (ZSSR) approach that relies only on the endomicroscopy data to be processed in a self-supervised manner without the need for ground-truth HR images. We tailored the proposed pipeline to the idiosyncrasies of endomicroscopy by introducing both: a physically-motivated Voronoi downscaling kernel accounting for the endomicroscope's irregular fibre-based sampling pattern, and realistic noise patterns. We also took advantage of video sequences to exploit a sequence of images for self-supervised zero-shot image quality improvement. We run ablation studies to assess our contribution in regards to the downscaling kernel and noise simulation. We validate our methodology on both synthetic and original data. Synthetic experiments were assessed with reference-based IQA, while our results for original images were evaluated in a user study conducted with both expert and non-expert observers. The results demonstrated superior performance in image quality of ZSSR reconstructions in comparison to the baseline method. The ZSSR is also competitive when compared to supervised single-image SR, especially being the preferred reconstruction technique by experts. Agnieszka Barbara Szczotka, Dzhoshkun I. Shakir, Matthew J. Clarkson, Stephen P. Pereira, Tom Vercauteren |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Multi-Task Recurrent Neural Network for Surgical Gesture Recognition and Progress PredictionabstractSurgical gesture recognition is important for surgical data science and computer-aided intervention. Even with robotic kinematic information, automatically segmenting surgical steps presents numerous challenges because surgical demonstrations are characterized by high variability in style, duration and order of actions. In order to extract discriminative features from the kinematic signals and boost recognition accuracy, we propose a multi-task recurrent neural network for simultaneous recognition of surgical gestures and estimation of a novel formulation of surgical task progress. To show the effectiveness of the presented approach, we evaluate its application on the JIGSAWS dataset, that is currently the only publicly available dataset for surgical gesture recognition featuring robot kinematic data. We demonstrate that recognition performance improves in multi-task frameworks with progress estimation without any additional manual labelling and training. Beatrice van Amsterdam, Matthew J. Clarkson, Danail Stoyanov |
ICRA | 2 |
| 2020 | Dimensions of ecological validity for usability evaluations in clinical settings
Niels van Berkel, Matthew J. Clarkson, Guofang Xiao, Eren Dursun, Moustafa Allam, Brian R. Davidson, Ann Blandford |
J. Biomed. Informatics | 2 |
| 2020 | An Enhanced Visualization of DBT Imaging Using Blind Deconvolution and Total Variation Minimization RegularizationabstractDigital Breast Tomosynthesis (DBT) presents out-of-plane artifacts caused by features of high intensity. Given observed data and knowledge about the point spread function (PSF), deconvolution techniques recover data from a blurred version. However, a correct PSF is difficult to achieve and these methods amplify noise. When no information is available about the PSF, blind deconvolution can be used. Additionally, Total Variation (TV) minimization algorithms have achieved great success due to its virtue of preserving edges while reducing image noise. This work presents a novel approach in DBT through the study of out-of-plane artifacts using blind deconvolution and noise regularization based on TV minimization. Gradient information was also included. The methodology was tested using real phantom data and one clinical data set. The results were investigated using conventional 2D slice-by-slice visualization and 3D volume rendering. For the 2D analysis, the artifact spread function (ASF) and Full Width at Half Maximum (FWHMMASF) of the ASF were considered. The 3D quantitative analysis was based on the FWHM of disks profiles at 90°, noise and signal to noise ratio (SNR) at 0° and 90°. A marked visual decrease of the artifact with reductions of FWHMASF (2D) and FWHM90° (volume rendering) of 23.8% and 23.6%, respectively, was observed. Although there was an expected increase in noise level, SNR values were preserved after deconvolution. Regardless of the methodology and visualization approach, the objective of reducing the out-of-plane artifact was accomplished. Both for the phantom and clinical case, the artifact reduction in the z was markedly visible. Ana M. Mota, Matthew J. Clarkson, Pedro Almeida 0001, Nuno Matela |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Generating Large Labeled Data Sets for Laparoscopic Image Processing Tasks Using Unpaired Image-to-Image Translation
Micha Pfeiffer, Isabel Funke, Maria Robu, Sebastian Bodenstedt, Leon Strenger, Sandy Engelhardt, Tobias Roß, Matthew J. Clarkson, Kurinchi Gurusamy, Brian R. Davidson, Lena Maier-Hein, Carina Riediger, Thilo Welsch, Jürgen Weitz, Stefanie Speidel |
MICCAI (5) | 8 |
| 2018 | Automatic Multi-Organ Segmentation on Abdominal CT With Dense V-NetworksabstractAutomatic segmentation of abdominal anatomy on computed tomography (CT) images can support diagnosis, treatment planning, and treatment delivery workflows. Segmentation methods using statistical models and multi-atlas label fusion (MALF) require inter-subject image registrations, which are challenging for abdominal images, but alternative methods without registration have not yet achieved higher accuracy for most abdominal organs. We present a registration-free deep-learning-based segmentation algorithm for eight organs that are relevant for navigation in endoscopic pancreatic and biliary procedures, including the pancreas, the gastrointestinal tract (esophagus, stomach, and duodenum) and surrounding organs (liver, spleen, left kidney, and gallbladder). We directly compared the segmentation accuracy of the proposed method to the existing deep learning and MALF methods in a cross-validation on a multi-centre data set with 90 subjects. The proposed method yielded significantly higher Dice scores for all organs and lower mean absolute distances for most organs, including Dice scores of 0.78 versus 0.71, 0.74, and 0.74 for the pancreas, 0.90 versus 0.85, 0.87, and 0.83 for the stomach, and 0.76 versus 0.68, 0.69, and 0.66 for the esophagus. We conclude that the deep-learning-based segmentation represents a registration-free method for multi-organ abdominal CT segmentation whose accuracy can surpass current methods, potentially supporting image-guided navigation in gastrointestinal endoscopy procedures. Eli Gibson, Francesco Giganti, Yipeng Hu, Ester Bonmati, Steven Bandula, Kurinchi Gurusamy, Brian R. Davidson, Stephen P. Pereira, Matthew J. Clarkson, Dean C. Barratt |
IEEE Trans. Medical Imaging | 9 |
| 2017 | Towards Image-Guided Pancreas and Biliary Endoscopy: Automatic Multi-organ Segmentation on Abdominal CT with Dense Dilated Networks
Eli Gibson, Francesco Giganti, Yipeng Hu, Ester Bonmati, Steven Bandula, Kurinchi Gurusamy, Brian R. Davidson, Stephen P. Pereira, Matthew J. Clarkson, Dean C. Barratt |
MICCAI (1) | 9 |
| 2015 | Database-Based Estimation of Liver Deformation under Pneumoperitoneum for Surgical Image-Guidance and Simulation
Stian Flage Johnsen, Stephen A. Thompson, Matthew J. Clarkson, Marc Modat, Johannes Totz, Kurinchi Gurusamy, Brian R. Davidson, Zeike A. Taylor, David J. Hawkes, Sébastien Ourselin |
MICCAI (2) | 3 |
| 2012 | Cortical Folding Analysis on Patients with Alzheimer's Disease and Mild Cognitive Impairment
David M. Cash, Andrew Melbourne, Marc Modat, Manuel Jorge Cardoso, Matthew J. Clarkson, Nick C. Fox, Sébastien Ourselin |
MICCAI (3) | 5 |
| 2011 | Longitudinal Cortical Thickness Estimation Using Khalimsky's Cubic Complex
Manuel Jorge Cardoso, Matthew J. Clarkson, Marc Modat, Sébastien Ourselin |
MICCAI (2) | 2 |
| 2010 | A Framework for Using Diffusion Weighted Imaging to Improve Cortical Parcellation
Matthew J. Clarkson, Ian B. Malone, Marc Modat, Kelvin K. Leung, Natalie S. Ryan, Daniel C. Alexander, Nick C. Fox, Sébastien Ourselin |
MICCAI (1) | 1 |
| 2010 | Increasing Power to Predict Mild Cognitive Impairment Conversion to Alzheimer's Disease Using Hippocampal Atrophy Rate and Statistical Shape Models
Kelvin K. Leung, Kai-Kai Shen, Josephine Barnes, Gerard R. Ridgway, Matthew J. Clarkson, Jurgen Fripp, Olivier Salvado, Fabrice Mériaudeau, Nick C. Fox, Pierrick Bourgeat |
MICCAI (2) | 5 |
| 2009 | Improved Maximum a Posteriori Cortical Segmentation by Iterative Relaxation of Priors
Manuel Jorge Cardoso, Matthew J. Clarkson, Gerard R. Ridgway, Marc Modat, Nick C. Fox, Sébastien Ourselin |
MICCAI (1) | 2 |
| 2001 | Using Photo-Consistency to Register 2D Optical Images of the Human Face to a 3D Surface ModelabstractThe authors propose a novel method to register two or more optical images to a 3D surface model. The potential applications of such a registration method could be in medicine for example, in image guided interventions, surveillance and identification, industrial inspection, or telemanipulation in remote or hostile environments. Registration is performed by optimizing a similarity measure with respect to the transformation parameters. We propose a novel similarity measure based on "photo-consistency." For each surface point, the similarity measure computes how consistent the corresponding optical image information in each view is with a lighting model. The relative pose of the optical images must be known. We validate the system using data from an optical-based surface reconstruction system and surfaces derived from magnetic resonance (MR) images of the human face. We test the accuracy and robustness of the system with respect to the number of video images, video image noise, errors in surface location and area, and complexity of the matched surfaces. We demonstrate the algorithm working on 10 further optical-based reconstructions of the human head and skin surfaces derived from MR images of the heads of five volunteers. Matching four optical images to a surface model produced a 3D error of between 1.45 and 1.59 mm, at a success rate of 100 percent, where the initial misregistration was up to 16 mm or degrees from the registration position. Matthew J. Clarkson, Daniel Rueckert, Derek L. G. Hill, David J. Hawkes |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Volume and Shape Preservation of Enhancing Lesions when Applying Non-rigid Registration to a Time Series of Contrast Enhancing MR Breast Images
Christine Tanner, Julia A. Schnabel, Daniel Chung, Matthew J. Clarkson, Daniel Rueckert, Derek L. G. Hill, David J. Hawkes |
MICCAI | 4 |
| 1999 | Registration of Video Images to Tomographic Images by Optimising Mutual Information Using Texture Mapping
Matthew J. Clarkson, Daniel Rueckert, Andrew P. King, Philip J. Edwards, Derek L. G. Hill, David J. Hawkes |
MICCAI | 1 |