VLDB 2026 Research / reviewers in the wild / expert
Nicolas Padoy
dblp:60/6834
· DBLP profile ↗
72ranked-venue papers
7as first author
43since 2021 · last 2026
0000-0002-5010-4137ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 51 · 3 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 22 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Where It Moves, It Matters: Referring Surgical Instrument Segmentation via MotionabstractEnabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexplored in surgical videos, with existing approaches struggling to generalize due to reliance on static visual cues and predefined instrument names. In this work, we introduce SurgRef, a novel motion-guided framework that grounds free-form language expressions in instrument motion, capturing how tools move and interact across time, rather than what they look like. This allows models to understand and segment instruments even under occlusion, ambiguity, or unfamiliar terminology. To train and evaluate SurgRef, we present Ref-IMotion, a diverse, multi-institutional video dataset with dense spatiotemporal masks and rich motion-centric expressions. SurgRef achieves state-of-the-art accuracy and generalization across surgical procedures, setting a new benchmark for robust, language-driven surgical video segmentation. Kun Yuan 0004, Long Bai 0008, Nassir Navab, Hongliang Ren 0001, Hong Joo Lee 0001, Tom Vercauteren, Nicolas Padoy |
AAAI | 10 |
| 2026 | EndoChat: Grounded multimodal large language model for endoscopic surgeryabstractRecently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a deficiency of MLLMs specialized for surgical scene understanding in endoscopic procedures. To this end, we present EndoChat, an MLLM tailored to address various dialogue paradigms and subtasks in understanding endoscopic procedures. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and seven surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, who provide positive feedback on the majority of conversation cases generated by EndoChat. Overall, these results demonstrate that EndoChat has the potential to advance training and automation in robotic-assisted surgery. Our dataset and model are publicly available at https://github.com/gkw0010/EndoChat. Guankun Wang, Long Bai 0008, Kun Yuan 0004, Zhen Li 0026, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001, Fan Zhang 0016, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 14 |
| 2025 | Medical Multimodal Model Stealing Attacks via Adversarial Domain AlignmentabstractMedical multimodal large language models (MLLMs) are becoming an instrumental part of healthcare systems, assisting medical personnel with decision making and results analysis. Models for radiology report generation are able to interpret medical imagery, thus reducing the workload of radiologists. As medical data is scarce and protected by privacy regulations, medical MLLMs represent valuable intellectual property. However, these assets are potentially vulnerable to model stealing, where attackers aim to replicate their functionality via black-box access. So far, model stealing for the medical domain has focused on image classification; however, existing attacks are not effective against MLLMs. In this paper, we introduce Adversarial Domain Alignment (ADA-Steal), the first stealing attack against medical MLLMs. ADA-Steal relies on natural images, which are public and widely available, as opposed to their medical counterparts. We show that data augmentation with adversarial noise is sufficient to overcome the data distribution gap between natural images and the domain-specific distribution of the victim MLLM. Experiments on the IU X-RAY and MIMIC-CXR radiology datasets demonstrate that Adversarial Domain Alignment enables attackers to steal the medical MLLM without any access to medical data. Yaling Shen, Zhixiong Zhuang, Kun Yuan 0004, Maria-Irina Nicolae, Nassir Navab, Nicolas Padoy, Mario Fritz |
AAAI | 6 |
| 2025 | Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging ScenesabstractMulti-view person association is a fundamental step towards multi-view analysis of human activities. Although the person re-identification features have been proven effective, they become unreliable in challenging scenes where persons share similar appearances. Therefore, cross-view geometric constraints are required for a more robust association. However, most existing approaches are either fully-supervised using ground-truth identity labels or require calibrated camera parameters that are hard to obtain. In this work, we investigate the potential of learning from synchronization, and propose a self-supervised uncalibrated multi-view person association approach, Self-MVA, without using any annotations. Specifically, we propose a self-supervised learning framework, consisting of an encoder-decoder model and a self-supervised pretext task, cross-view image synchronization, which aims to distinguish whether two images from different views are captured at the same time. The model encodes each person’s unified geometric and appearance features, and we train it by utilizing synchronization labels for supervision after applying Hungarian matching to bridge the gap between instance-wise and image-wise distances. To further reduce the solution space, we propose two types of self-supervised linear constraints: multi-view re-projection and pairwise edge association. Extensive experiments on three challenging public benchmark datasets (WILDTRACK, MVOR, and SOLDIERS) show that our approach achieves state-of-the-art results, surpassing existing unsupervised and fully-supervised approaches. Code is available at https://github.com/CAMMA-public/Self-MVA. Vinkle Srivastav, Didier Mutter, Nicolas Padoy |
CVPR | 4 |
| 2025 | CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical ToolsabstractTool tracking in surgical videos is essential for advancing computer-assisted interventions, such as skill assessment, safety zone estimation, and human-machine collaboration. However, the lack of context-rich datasets limits AI applications in this field. Existing datasets rely on overly generic tracking formalizations that fail to capture surgical-specific dynamics, such as tools moving out of the camera’s view or exiting the body. This results in less clinically relevant trajectories and a lack of flexibility for real-world surgical applications. Methods trained on these datasets often struggle with visual challenges such as smoke, reflection, and bleeding, further exposing the limitations of current approaches. We introduce CholecTrack20, a specialized dataset for multi-class, multi-tool tracking in surgical procedures. It redefines tracking formalization with three perspectives: (1) intraoperative, (2) intracorporeal, and (3) visibility, enabling adaptable and clinically meaningful tool trajectories. The dataset comprises 20 full-length surgical videos, annotated at 1 fps, yielding over 35K frames and 65K labeled tool instances. Annotations include spatial location, category, identity, operator, phase, and visual challenges. Benchmarking state-of-the-art methods on Cholec-Track20 reveals significant performance gaps, with current approaches (< 45% HOTA) failing to meet the accuracy required for clinical translation. These findings motivate the need for advanced and intuitive tracking algorithms and establish CholecTrack20 as a foundation for developing robust AI-driven surgical assistance systems. Chinedu Innocent Nwoye, Kareem Elgohary, Anvita Srinivas, Fauzan Zaid, Joël L. Lavanchy, Nicolas Padoy |
CVPR | 6 |
| 2025 | OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language PretrainingabstractSurgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining framework specifically designed for ophthalmic surgical workflow understanding. OphCLIP leverages the OphVL dataset we constructed, a large-scale and comprehensive collection of over 375K hierarchically structured video-text pairs with tens of thousands of different combinations of attributes (surgeries, phases/operations/actions, instruments, medications, as well as more advanced aspects like the causes of eye diseases, surgical objectives, and postoperative recovery recommendations, etc). These hierarchical video-text correspondences enable OphCLIP to learn both fine-grained and long-term visual representations by aligning short video clips with detailed narrative descriptions and full videos with structured titles, capturing intricate surgical details and high-level procedural insights, respectively. Our OphCLIP also designs a retrieval-augmented pretraining framework to leverage the underexplored large-scale silent surgical procedure videos, automatically retrieving semantically relevant content to enhance the representation learning of narrative videos. Evaluation across 11 datasets for phase recognition and multi-instrument identification shows OphCLIP's robust generalization and superior performance. Kun Yuan 0004, Yaling Shen, Xiaohao Xu, Wei Li 0320, Zhongxing Xu, Zelin Peng, Siyuan Yan, Vinkle Srivastav, Diping Song, Tianbin Li, Danli Shi, Jin Ye 0002, Nicolas Padoy, Nassir Navab, Junjun He, ZongYuan Ge |
ICCV | 17 |
| 2025 | Multi-modal Representations for Fine-Grained Multi-Label Critical View of Safety Recognition
Britty Baby, Vinkle Srivastav, Pooja P. Jain, Kun Yuan 0004, Pietro Mascagni, Nicolas Padoy |
MICCAI (11) | 6 |
| 2025 | Feature Mixing Approach for Detecting Intraoperative Adverse Events in Laparoscopic Roux-en-Y Gastric Bypass Surgery
Rupak Bose, Chinedu Innocent Nwoye, Jorge F. Lazo, Joël L. Lavanchy, Nicolas Padoy |
MICCAI (11) | 5 |
| 2025 | SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Kun Yuan 0004, Guankun Wang, Mobarak I. Hoque, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
MICCAI (9) | 7 |
| 2025 | Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
Soham Walimbe, Britty Baby, Vinkle Srivastav, Nicolas Padoy |
MICCAI (11) | 4 |
| 2025 | Recognizing Surgical Phases Anywhere: Few-Shot Test-Time Adaptation and Task-Graph Guided Refinement
Kun Yuan 0004, Tingxuan Chen, Joël L. Lavanchy, Christian Heiliger, Ege Özsoy, Yiming Huang 0007, Long Bai 0008, Nassir Navab, Vinkle Srivastav, Hongliang Ren 0001, Nicolas Padoy |
MICCAI (9) | 12 |
| 2025 | Learning from Sparse Point Labels for Dense Carcinosis Localization in Advanced Ovarian Cancer Assessment
Farahdiba Zarin, Riccardo Oliva, Vinkle Srivastav, Armine Vardazaryan, Andrea Rosati, Alice Zampolini Faustini, Giovanni Scambia, Anna Fagotti, Pietro Mascagni, Nicolas Padoy |
MICCAI (11) | 10 |
| 2025 | SurgiTrack: Fine-grained multi-class multi-tool tracking in surgical videosabstractAccurate tool tracking is essential for the success of computer-assisted intervention. Previous efforts often modeled tool trajectories rigidly, overlooking the dynamic nature of surgical procedures, especially tracking scenarios like out-of-body and out-of-camera views. Addressing this limitation, the new CholecTrack20 dataset provides detailed labels that account for multiple tool trajectories in three perspectives : (1) intraoperative, (2) intracorporeal, and (3) visibility, representing the different types of temporal duration of tool tracks. These fine-grained labels enhance tracking flexibility but also increase the task complexity. Re-identifying tools after occlusion or re-insertion into the body remains challenging due to high visual similarity, especially among tools of the same category. This work recognizes the critical role of the tool operators in distinguishing tool track instances, especially those belonging to the same tool category. The operators’ information are however not explicitly captured in surgical videos. We therefore propose SurgiTrack, a novel deep learning method that leverages YOLOv7 for precise tool detection and employs an attention mechanism to model the originating direction of the tools, as a proxy to their operators, for tool re-identification. To handle diverse tool trajectory perspectives, SurgiTrack employs a harmonizing bipartite matching graph, minimizing conflicts and ensuring accurate tool identity association. Experimental results on CholecTrack20 demonstrate SurgiTrack’s effectiveness, outperforming baselines and state-of-the-art methods with real-time inference capability. This work sets a new standard in surgical tool tracking, providing dynamic trajectories for more adaptable and precise assistance in minimally invasive surgeries. • Formalize multi-perspective tool tracking: intraoperative, intracorporeal, visibility • Benchmark state-of-the-art detection & tracking methods on the CholecTrack20 dataset • Develop SurgiTrack: direction-based re-ID & harmonizing graph-matching tracking model • SurgiTrack outperforms existing methods in metrics, quantitatively and qualitatively • In-depth evaluation across perspectives, frame rates, and surgical visual challenges Chinedu Innocent Nwoye, Nicolas Padoy |
Medical Image Anal. | 2 |
| 2025 | SAF-IS: A spatial annotation free framework for instance segmentation of surgical tools
Luca Sestini, Benoit Rosa, Elena De Momi, Giancarlo Ferrigno, Nicolas Padoy |
Medical Image Anal. | 5 |
| 2025 | Learning multi-modal representations by watching hundreds of surgical video lecturesabstractRecent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. SurgVLP constructs a new contrastive learning objective to align video clip embeddings with the corresponding multiple text embeddings by bringing them together within a joint latent space. To effectively demonstrate the representational capability of the learned joint latent space, we introduce several vision-and-language surgical tasks and evaluate various vision-only tasks specific to surgery, e.g., surgical tool, phase, and triplet recognition. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The code is available at https://github.com/CAMMA-public/SurgVLP. Kun Yuan 0004, Vinkle Srivastav, Tong Yu 0009, Joël L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy |
Medical Image Anal. | 8 |
| 2025 | Rethinking data imbalance in class incremental surgical instrument segmentationabstractIn surgical instrument segmentation, the increasing variety of instruments over time poses a significant challenge for existing neural networks, as they are unable to effectively learn such incremental tasks and suffer from catastrophic forgetting. When learning new data, the model experiences a sharp performance drop on previously learned data. Although several continual learning methods have been proposed for incremental understanding tasks in surgical scenarios, the issue of data imbalance often leads to a strong bias in the segmentation head, resulting in poor performance. Data imbalance can occur in two forms: (i) class imbalance between new and old data, and (ii) class imbalance within the same time point of data. Such imbalances often cause the dominant classes to take over the training process of continual semantic segmentation (CSS). To address this issue, we propose SurgCSS, a novel plug-and-play CSS framework for surgical instrument segmentation under data imbalance. Specifically, we generate realistic surgical backgrounds through inpainting and blend instrument foregrounds with the generated backgrounds in a class-aware manner to balance the data distribution in various scenarios. We further propose the Class Desensitization Loss by employing contrastive learning to correct edge biases caused by data imbalance. Moreover, we dynamically fuse the weight parameters of the old and new models to achieve a better trade-off between the biased and unbiased model weights. To investigate the data imbalance problem in surgical scenarios, we construct a new benchmark for surgical instrument CSS by integrating four public datasets: EndoVis 2017, EndoVis 2018, CholecSeg8k, and SAR-RAPR50. Extensive experiments demonstrate the effectiveness of the proposed framework, achieving significant performance improvement against existing baselines. Our method demonstrates excellent potential for clinical applications. The code is publicly available at github.com/Zzsf11/SurgCSS. Shifang Zhao, Long Bai 0008, Kun Yuan 0004, Feng Li 0034, Jieming Yu, Wenzhen Dong, Guankun Wang, Mobarakol Islam, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 9 |
| 2025 | Surgical text-to-image generationabstractAcquiring surgical data for research and development is significantly hindered by high annotation costs and practical and ethical constraints. Synthetically generated images present a valuable alternative. In this work, we explore adapting text-to-image generative models for the surgical domain using the CholecT50 dataset, which provides surgical images annotated with action triplets (instrument, verb, target). We investigate several language models and find T5 to offer more distinct features for differentiating surgical actions on triplet-based textual inputs, and showcasing stronger alignment between long and triplet-based captions. To address challenges in training text-to-image models solely on triplet-based captions without additional input signals, we discover that triplet text embeddings are instrument-centric in the latent space. Leveraging this insight, we design an instrument-based class balancing technique to counteract data imbalance and skewness, improving training convergence. Extending Imagen, a diffusion-based generative model, we develop Surgical Imagen to generate photorealistic and activity-aligned surgical images from triplet-based textual prompts. We assess the model on quality, alignment, reasoning, and knowledge, achieving FID and CLIP scores of 3.7 and 26.8% respectively. Human expert survey shows that participants were highly challenged by the realistic characteristics of the generated samples, demonstrating Surgical Imagen’s effectiveness as a practical alternative to real data collection. • Present Surgical Imagen for photorealistic surgical images from text. • Utilize triplet-based labels (instrument verb target) for accurate image generation. • Investigate various language models; T5 proves best for triplet alignment. • Introduce instrument-based class balancing to address dataset imbalances. • Validate model effectiveness through expert surveys and automated metrics. Chinedu Innocent Nwoye, Rupak Bose, Kareem Elgohary, Lorenzo Arboit, Giorgio Carlino, Joël L. Lavanchy, Pietro Mascagni, Nicolas Padoy |
Pattern Recognit. Lett. | 8 |
| 2024 | SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationabstractWe present a new self-supervised approach, SelfPose3d, for estimating 3d poses of multiple persons from multiple camera views. Unlike current state-of-the-art fully-supervised methods, our approach does not require any 2d or 3d ground-truth poses and uses only the multiview input images from a calibrated camera setup and 2d pseudo poses generated from an off-the-shelf 2d human pose estimator. We propose two self-supervised learning objectives: self-supervised person localization in 3d space and self-supervised 3d pose estimation. We achieve self-supervised 3d person localization by training the model on synthetically generated 3d points, serving as 3d person root positions, and on the projected root-heatmaps in all the views. We then model the 3d poses of all the localized persons with a bottleneck representation, map them onto all views obtaining 2d joints, and render them using 2d Gaussian heatmaps in an end-to-end differentiable manner. Afterwards, we use the corresponding 2d joints and heatmaps from the pseudo 2d poses for learning. To alleviate the intrinsic inaccuracy of the pseudo labels, we propose an adaptive supervision attention mechanism to guide the self-supervision. Our experiments and analysis on three public benchmark datasets, including Panoptic, Shelf, and Campus, show the effectiveness of our approach, which is comparable to fully-supervised methods. Code is available at https://github.com/CAMMA-public/SelfPose3D. Vinkle Srivastav, Nicolas Padoy |
CVPR | 3 |
| 2024 | Jumpstarting Surgical Computer Vision
Deepak Alapatt, Aditya Murali, Vinkle Srivastav, Pietro Mascagni, Nicolas Padoy |
MICCAI (6) | 5 |
| 2024 | Towards Real-Time Intrahepatic Vessel Identification in Intraoperative Ultrasound-Guided Liver Surgery
Karl-Philippe Beaudet, Alexandros Karargyris, Sidaty El Hadramy, Stephane Cotin, Jean-Paul Mazellier, Nicolas Padoy, Juan Verde |
MICCAI (6) | 6 |
| 2024 | Enhancing Gait Video Analysis in Neurodegenerative Diseases by Knowledge Augmentation in Vision Language Model
Diwei Wang, Kun Yuan 0004, Candice Müller, Nicolas Padoy, Hyewon Seo |
MICCAI (5) | 5 |
| 2024 | HecVL: Hierarchical Video-Language Pretraining for Zero-Shot Surgical Phase Recognition
Kun Yuan 0004, Vinkle Srivastav, Nassir Navab, Nicolas Padoy |
MICCAI (6) | 4 |
| 2024 | Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge AugmentationabstractSurgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data. This study aims to bridge the gap by addressing issues regarding textual information loss in surgical lecture videos and the spatial-temporal challenges of surgical VLP. To tackle these issues, we propose a hierarchical knowledge augmentation approach and a novel Procedure-Encoded Surgical Knowledge-Augmented Video-Language Pretraining (PeskaVLP) framework. The proposed knowledge augmentation approach uses large language models (LLM) to refine and enrich surgical concepts, thus providing comprehensive language supervision and reducing the risk of overfitting. The PeskaVLP framework combines language supervision with visual self-supervision, constructing hard negative samples and employing a Dynamic Time Warping (DTW) based loss function to effectively comprehend the cross-modal procedural alignment. Extensive experiments on multiple public surgical scene understanding and cross-modal retrieval datasets show that our proposed method significantly improves zero-shot transferring performance and offers a generalist visual repre- sentation for further advancements in surgical scene understanding. The source code will be available at https://github.com/CAMMA-public/PeskaVLP. Kun Yuan 0004, Vinkle Srivastav, Nassir Navab, Nicolas Padoy |
NeurIPS | 4 |
| 2024 | Latent Graph Representations for Critical View of Safety AssessmentabstractAssessing the critical view of safety in laparoscopic cholecystectomy requires accurate identification and localization of key anatomical structures, reasoning about their geometric relationships to one another, and determining the quality of their exposure. Prior works have approached this task by including semantic segmentation as an intermediate step, using predicted segmentation masks to then predict the CVS. While these methods are effective, they rely on extremely expensive ground-truth segmentation annotations and tend to fail when the predicted segmentation is incorrect, limiting generalization. In this work, we propose a method for CVS prediction wherein we first represent a surgical image using a disentangled latent scene graph, then process this representation using a graph neural network. Our graph representations explicitly encode semantic information - object location, class information, geometric relations - to improve anatomy-driven reasoning, as well as visual features to retain differentiability and thereby provide robustness to semantic errors. Finally, to address annotation cost, we propose to train our method using only bounding box annotations, incorporating an auxiliary image reconstruction objective to learn fine-grained object boundaries. We show that our method not only outperforms several baseline methods when trained with bounding box annotations, but also scales effectively when trained with segmentation masks, maintaining state-of-the-art performance. Aditya Murali, Deepak Alapatt, Pietro Mascagni, Armine Vardazaryan, Alain Garcia, Nariaki Okamoto, Didier Mutter, Nicolas Padoy |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Why is the Winner the Best?abstractInternational benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multicenter study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and post-processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work. Matthias Eisenmann, Annika Reinke, Vivienn Weru, Minu Tizabi, Fabian Isensee, Tim Adler, Sharib Ali, Vincent Andrearczyk, Marc Aubreville, Ujjwal Baid, Spyridon Bakas, Niranjan Balu, Sophia Bano, Jorge Bernal, Sebastian Bodenstedt, Alessandro Casella, Veronika Cheplygina, Marie Daum, Marleen de Bruijne, Adrien Depeursinge, Reuben Dorent, Jan Egger, David Gage Ellis, Sandy Engelhardt, Melanie Ganz-Benjaminsen, Noha M. Ghatwary, Gabriel Girard, Patrick Godau, Anubha Gupta, Lasse Hansen, Kanako Harada, Mattias P. Heinrich, Nicholas Heller, Alessa Hering, Arnaud Huaulmé, Pierre Jannin, A. Emre Kavur, Oldrich Kodym, Michal Kozubek 0001, Jianning Li 0002, Hongwei Li 0004, Jun Ma 0016, Carlos Martín-Isla, Bjoern Menze, J. Alison Noble, Valentin Oreiller, Nicolas Padoy, Sarthak Pati, Kelly Payette, Tim Rädsch, Jonathan Rafael-Patino, Vivek Singh Bawa, Stefanie Speidel, Carole H. Sudre, Kimberlin M. H. van Wijnen, Martin Wagner 0001, D. Wei, Amine Yamlahi, Moi Hoon Yap, C. Yuan, Maximilian Zenk, A. Zia, David Zimmerer, Dogu Baran Aydogan, Binod Bhattarai, Louise Bloch, Raphael Brüngel, J. Cho, C. Choi, Qi Dou 0001, Ivan Ezhov, Christoph M. Friedrich, C. Fuller, Rebati Raman Gaire, Adrian Galdran, Álvaro García-Faura, Maria Grammatikopoulou, S. Hong, Mostafa Jahanifar, I. Jang, Abdolrahim Kadkhodamohammadi, I. Kang, Florian Kofler, S. Kondo, Hugo J. Kuijf, M. Luu, Tomaz Martincic, Pedro Morais, Mohamed A. Naser, Bruno Oliveira 0002, David Owen 0001, S. Pang, Szymon Plotka, Élodie Puybareau, Nasir M. Rajpoot, K. Ryu, Numan Saeed, Adam J. Shephard, Dejan Stepec, Ronast Subedi, Guillaume Tochon, Helena R. Torres, Hélène Urien, João L. Vilaça, Kareem A. Wahid, Benedikt Wiestler, Marek Wodzinski, F. Xia, J. Xie, Z. Xiong, Sen Yang 0006, Klaus H. Maier-Hein, Paul F. Jaeger, Annette Kopp-Schneider, Lena Maier-Hein |
CVPR | 47 |
| 2023 | Trackerless Volume Reconstruction from Intraoperative Ultrasound Images
Sidaty El Hadramy, Juan Verde, Karl-Philippe Beaudet, Nicolas Padoy, Stephane Cotin |
MICCAI (10) | 4 |
| 2023 | Intraoperative CT Augmentation for Needle-Based Liver Interventions
Sidaty El Hadramy, Juan Verde, Nicolas Padoy, Stephane Cotin |
MICCAI (9) | 3 |
| 2023 | Encoding Surgical Videos as Latent Spatiotemporal Graphs for Object and Anatomy-Driven Reasoning
Aditya Murali, Deepak Alapatt, Pietro Mascagni, Armine Vardazaryan, Alain Garcia, Nariaki Okamoto, Didier Mutter, Nicolas Padoy |
MICCAI (9) | 8 |
| 2023 | Surgical Action Triplet Detection by Mixed Supervised Learning of Instrument-Tissue Interactions
Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy |
MICCAI (9) | 4 |
| 2023 | Self-distillation for Surgical Action Recognition
Amine Yamlahi, Thuy Nuong Tran, Patrick Godau, Melanie Schellenberg, Dominik Michael, Finn-Henri Smidt, Jan-Hinrich Nölke, Tim Adler, Minu Tizabi, Chinedu Innocent Nwoye, Nicolas Padoy, Lena Maier-Hein |
MICCAI (9) | 11 |
| 2023 | Editorial for the MEDIA MICCAI special issue 2021
Marleen de Bruijne, Philippe C. Cattin, Stephane Cotin, Nicolas Padoy, Stefanie Speidel, Yefeng Zheng 0001, Caroline Essert |
Medical Image Anal. | 4 |
| 2023 | CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu 0009, Armine Vardazaryan, Fangfang Xia, Tong Xia, Fucang Jia, Yuxuan Yang 0007, Hao Wang 0081, Derong Yu, Guoyan Zheng, Xiaotian Duan, Neil Getty, Ricardo Sanchez-Matilla, Maria Robu, Li Zhang 0040, Huabin Chen, Jiacheng Wang 0002, Liansheng Wang 0002, Beerend G. A. Gerats, Sista Raviteja, Rachana Sathish, Rong Tao, Satoshi Kondo, Winnie Pang, Hongliang Ren 0001, Julian Ronald Abbing, Mohammad Hasan Sarhan, Sebastian Bodenstedt, Nithya Bhasker, Bruno Oliveira 0002, Helena R. Torres, Finn Gaida, Tobias Czempiel, João L. Vilaça, Pedro Morais, Jaime C. Fonseca 0001, Ruby Mae Egging, Inge Nicole Wijma, Chen Qian 0006, Guibin Bian, Zhen Li 0026, Velmurugan Balasubramanian, Debdoot Sheet, Imanol Luengo, Yuanbo Zhu, Shuai Ding 0001, Jakob-Anton Aschenbrenner, Nicolas Elini van der Kar, Mengya Xu, Mobarakol Islam, Seenivasan Lalithkumar, Alexander Jenke, Danail Stoyanov, Didier Mutter, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Nicolas Padoy |
Medical Image Anal. | 62 |
| 2023 | CholecTriplet2022: Show me a tool and tell me the triplet - An endoscopic vision challenge for surgical action triplet detection
Chinedu Innocent Nwoye, Tong Yu 0009, Saurav Sharma, Aditya Murali, Deepak Alapatt, Armine Vardazaryan, Kun Yuan 0004, Jonas Hajek, Wolfgang Reiter, Amine Yamlahi, Finn-Henri Smidt, Xiaoyang Zou, Guoyan Zheng, Bruno Oliveira 0002, Helena R. Torres, Satoshi Kondo, Satoshi Kasai, Felix Holm, Ege Özsoy, Shuangchun Gui, Sista Raviteja, Rachana Sathish, Pranav Poudel, Binod Bhattarai, Ziheng Wang 0003, Guo Rui, Melanie Schellenberg, João L. Vilaça, Tobias Czempiel, Zhenkun Wang 0001, Debdoot Sheet, Shrawan Kumar Thapa, Max Berniker, Patrick Godau, Pedro Morais, Sudarshan Regmi, Thuy Nuong Tran, Jaime C. Fonseca 0001, Jan-Hinrich Nölke, Estevão Lima, Eduard Vazquez, Lena Maier-Hein, Nassir Navab, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Didier Mutter, Nicolas Padoy |
Medical Image Anal. | 49 |
| 2023 | Dissecting self-supervised learning methods for surgical computer visionabstractThe field of surgical computer vision has undergone considerable breakthroughs in recent years with the rising popularity of deep neural network-based methods. However, standard fully-supervised approaches for training such models require vast amounts of annotated data, imposing a prohibitively high cost; especially in the clinical domain. Self-Supervised Learning (SSL) methods, which have begun to gain traction in the general computer vision community, represent a potential solution to these annotation costs, allowing to learn useful representations from only unlabeled data. Still, the effectiveness of SSL methods in more complex and impactful domains, such as medicine and surgery, remains limited and unexplored. In this work, we address this critical need by investigating four state-of-the-art SSL methods (MoCo v2, SimCLR, DINO, SwAV) in the context of surgical computer vision. We present an extensive analysis of the performance of these methods on the Cholec80 dataset for two fundamental and popular tasks in surgical context understanding, phase recognition and tool presence detection. We examine their parameterization, then their behavior with respect to training data quantities in semi-supervised settings. Correct transfer of these methods to surgery, as described and conducted in this work, leads to substantial performance gains over generic uses of SSL - up to 7.4% on phase recognition and 20% on tool presence detection - as well as state-of-the-art semi-supervised phase recognition approaches by up to 14%. Further results obtained on a highly diverse selection of surgical datasets exhibit strong generalization properties. The code is available at https://github.com/CAMMA-public/SelfSupSurg. Sanat Ramesh, Vinkle Srivastav, Deepak Alapatt, Tong Yu 0009, Aditya Murali, Luca Sestini, Chinedu Innocent Nwoye, Idris Hamoud, Saurav Sharma, Antoine Fleurentin, Georgios Exarchakis, Alexandros Karargyris, Nicolas Padoy |
Medical Image Anal. | 13 |
| 2023 | FUN-SIS: A Fully UNsupervised approach for Surgical Instrument Segmentation
Luca Sestini, Benoit Rosa, Elena De Momi, Giancarlo Ferrigno, Nicolas Padoy |
Medical Image Anal. | 5 |
| 2023 | Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmarkabstractPURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery. Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt |
Medical Image Anal. | 16 |
| 2023 | Live laparoscopic video retrieval with compressed uncertainty
Tong Yu 0009, Pietro Mascagni, Juan Verde, Jacques Marescaux, Didier Mutter, Nicolas Padoy |
Medical Image Anal. | 6 |
| 2023 | Federated Cycling (FedCy): Semi-Supervised Federated Learning of Surgical PhasesabstractRecent advancements in deep learning methods bring computer-assistance a step closer to fulfilling promises of safer surgical procedures. However, the generalizability of such methods is often dependent on training on diverse datasets from multiple medical institutions, which is a restrictive requirement considering the sensitive nature of medical data. Recently proposed collaborative learning methods such as Federated Learning (FL) allow for training on remote datasets without the need to explicitly share data. Even so, data annotation still represents a bottleneck, particularly in medicine and surgery where clinical expertise is often required. With these constraints in mind, we propose FedCy, a federated semi-supervised learning (FSSL) method that combines FL and self-supervised learning to exploit a decentralized dataset of both labeled and unlabeled videos, thereby improving performance on the task of surgical phase recognition. By leveraging temporal patterns in the labeled data, FedCy helps guide unsupervised training on unlabeled data towards learning task-specific features for phase recognition. We demonstrate significant performance gains over state-of-the-art FSSL methods on the task of automatic recognition of surgical phases using a newly collected multi-institutional dataset of laparoscopic cholecystectomy videos. Furthermore, we demonstrate that our approach also learns more generalizable features when tested on data from an unseen domain. Hasan Kassem, Deepak Alapatt, Pietro Mascagni, Alexandros Karargyris, Nicolas Padoy |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Weakly Supervised Temporal Convolutional Networks for Fine-Grained Surgical Activity RecognitionabstractAutomatic recognition of fine-grained surgical activities, called steps, is a challenging but crucial task for intelligent intra-operative computer assistance. The development of current vision-based activity recognition methods relies heavily on a high volume of manually annotated data. This data is difficult and time-consuming to generate and requires domain-specific knowledge. In this work, we propose to use coarser and easier-to-annotate activity labels, namely phases, as weak supervision to learn step recognition with fewer step annotated videos. We introduce a step-phase dependency loss to exploit the weak supervision signal. We then employ a Single-Stage Temporal Convolutional Network (SS-TCN) with a ResNet-50 backbone, trained in an end-to-end fashion from weakly annotated videos, for temporal activity segmentation and recognition. We extensively evaluate and show the effectiveness of the proposed method on a large video dataset consisting of 40 laparoscopic gastric bypass procedures and the public benchmark CATARACTS containing 50 cataract surgeries. Sanat Ramesh, Diego Dall'Alba, Cristians Gonzalez, Tong Yu 0009, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Paolo Fiorini, Nicolas Padoy |
IEEE Trans. Medical Imaging | 9 |
| 2022 | Surgical data science - from concepts toward clinical translationabstractRecent developments in data science in general and machine learning in particular have transformed the way experts envision the future of surgery. Surgical Data Science (SDS) is a new research field that aims to improve the quality of interventional healthcare through the capture, organization, analysis and modeling of data. While an increasing number of data-driven approaches and clinical applications have been studied in the fields of radiological and clinical data science, translational success stories are still lacking in surgery. In this publication, we shed light on the underlying reasons and provide a roadmap for future advances in the field. Based on an international workshop involving leading researchers in the field of SDS, we review current practice, key achievements and initiatives as well as available standards and tools for a number of topics relevant to the field, namely (1) infrastructure for data acquisition, storage and access in the presence of regulatory constraints, (2) data annotation and sharing and (3) data analytics. We further complement this technical perspective with (4) a review of currently available SDS products and the translational progress from academia and (5) a roadmap for faster clinical translation and exploitation of the full potential of SDS, based on an international multi-round Delphi process. Lena Maier-Hein, Matthias Eisenmann, Duygu Sarikaya, Keno März, Toby Collins, Anand Malpani, Johannes Fallert, Hubertus Feußner, Stamatia Giannarou, Pietro Mascagni, Hirenkumar Nakawala, Adrian Park 0001, Carla M. Pugh, Danail Stoyanov, S. Swaroop Vedula, Kevin Cleary, Gabor Fichtinger, Germain Forestier, Bernard Gibaud, Teodor P. Grantcharov, Makoto Hashizume, Doreen Heckmann-Nötzel, Hannes Kenngott, Ron Kikinis, Lars Mündermann, Nassir Navab, Sinan Onogur, Tobias Roß, Raphael Sznitman, Russell H. Taylor, Minu Tizabi, Martin Wagner 0001, Gregory D. Hager, Thomas Neumuth, Nicolas Padoy, Justin Collins, Ines Gockel, Jan Goedeke, Daniel A. Hashimoto, Luc Joyeux, Kyle Lam, Daniel Richard Leff, Amin Madani, Hani J. Marcus, Ozanan R. Meireles, Alexander Seitel, Dogu Teber, Frank Ückert, Beat P. Müller-Stich, Pierre Jannin, Stefanie Speidel |
Medical Image Anal. | 35 |
| 2022 | Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos
Chinedu Innocent Nwoye, Tong Yu 0009, Cristians Gonzalez, Barbara Seeliger, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Nicolas Padoy |
Medical Image Anal. | 8 |
| 2022 | Unsupervised domain adaptation for clinician pose estimation and instance segmentation in the operating room
Vinkle Srivastav, Afshin Gangi, Nicolas Padoy |
Medical Image Anal. | 3 |
| 2021 | A generalizable approach for multi-view 3D human pose regression
Abdolrahim Kadkhodamohammadi, Nicolas Padoy |
Mach. Vis. Appl. | 2 |
| 2020 | Encode the Unseen: Predictive Video Hashing for Scalable Mid-stream Retrieval
Tong Yu 0009, Nicolas Padoy |
ACCV (5) | 2 |
| 2020 | Recognition of Instrument-Tissue Interactions in Endoscopic Videos via Action TripletsabstractRecognition of surgical activity is an essential component to develop context-aware decision support for the operating room. In this work, we tackle the recognition of fine-grained activities, modeled as action triplets representing the tool activity. To this end, we introduce a new laparoscopic dataset, CholecT40, consisting of 40 videos from the public dataset Cholec80 in which all frames have been annotated using 128 triplet classes. Furthermore, we present an approach to recognize these triplets directly from the video data. It relies on a module called Class Activation Guide (CAG), which uses the instrument activation maps to guide the verb and target recognition. To model the recognition of multiple triplets in the same frame, we also propose a trainable 3D Interaction Space, which captures the associations between the triplet components. Finally, we demonstrate the significance of these contributions via several ablation studies and comparisons to baselines on CholecT40. Chinedu Innocent Nwoye, Cristians Gonzalez, Tong Yu 0009, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Nicolas Padoy |
MICCAI (3) | 7 |
| 2020 | Self-supervision on Unlabelled or Data for Multi-person 2D/3D Human Pose Estimation
Vinkle Srivastav, Afshin Gangi, Nicolas Padoy |
MICCAI (1) | 3 |
| 2020 | CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer-Assisted InterventionsabstractData-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their interpretation and exploitation in clinical practice. While similar benefits are longed for in the field of interventional imaging, this ambition is challenged by a much higher heterogeneity. Clinical workflows within interventional suites and operating theatres are extremely complex and typically rely on poorly integrated intra-operative devices, sensors, and support infrastructures. Taking stock of some of the most exciting developments in machine learning and artificial intelligence for computer assisted interventions, we highlight the crucial need to take context and human factors into account in order to address these challenges. Contextual artificial intelligence for computer assisted intervention, or CAI4CAI, arises as an emerging opportunity feeding into the broader field of surgical data science. Central challenges being addressed in CAI4CAI include how to integrate the ensemble of prior knowledge and instantaneous sensory information from experts, sensors and actuators; how to create and communicate a faithful and actionable shared representation of the surgery among a mixed human-AI actor team; how to design interventional systems and associated cognitive shared control schemes for online uncertainty-aware collaborative decision making ultimately producing more precise and reliable interventions. Tom Vercauteren, Mathias Unberath, Nicolas Padoy, Nassir Navab |
Proc. IEEE | 3 |
| 2020 | Future-State Predicting LSTM for Early Surgery Type RecognitionabstractThis work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the development of "smart" OR systems that can provide automatic context-aware assistance, and also enable quick database indexing. The task is however ridden with challenges specific to videos belonging to the domain of laparoscopy, such as high visual similarity across surgeries and large variations in video durations. To capture the spatio-temporal dependencies in these videos, we choose as our model a combination of a convolutional neural network (CNN) and long short-term memory (LSTM) network. We then propose two complementary approaches for improving early recognition performance. The first approach is a CNN fine-tuning method that encourages surgeries to be distinguished based on the initial frames of laparoscopic videos. The second approach, referred to as " Future-State Predicting LSTM," trains an LSTM to predict information related to future frames, which helps in distinguishing between the different types of surgeries. We evaluate our approaches on a large dataset of 425 laparoscopic videos containing nine types of surgeries (Laparo425), and achieve on average an accuracy of 75% having observed only the first 10 min of a surgery. These results are quite promising from a practical standpoint and also encouraging for other types of image-guided surgeries. Siddharth Kannan, Gaurav Yengera, Didier Mutter, Jacques Marescaux, Nicolas Padoy |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Self-Supervised Surgical Tool Segmentation using Kinematic InformationabstractSurgical tool segmentation in endoscopic images is the first step towards pose estimation and (sub-)task automation in challenging minimally invasive surgical operations. While many approaches in the literature have shown great results using modern machine learning methods such as convolutional neural networks, the main bottleneck lies in the acquisition of a large number of manually-annotated images for efficient learning. This is especially true in surgical context, where patient-to-patient differences impede the overall generalizability. In order to cope with this lack of annotated data, we propose a self-supervised approach in a robot-assisted context. To our knowledge, the proposed approach is the first to make use of the kinematic model of the robot in order to generate training labels. The core contribution of the paper is to propose an optimization method to obtain good labels for training despite an unknown hand-eye calibration and an imprecise kinematic model. The labels can subsequently be used for fine-tuning a fully-convolutional neural network for pixel-wise classification. As a result, the tool can be segmented in the endoscopic images without needing a single manually-annotated image. Experimental results on phantom and in vivo datasets obtained using a flexible robotized endoscopy system are very promising. Cristian da Costa Rocha, Nicolas Padoy, Benoit Rosa |
ICRA | 2 |
| 2019 | Human Pose Estimation on Privacy-Preserving Low-Resolution Depth Images
Vinkle Srivastav, Afshin Gangi, Nicolas Padoy |
MICCAI (5) | 3 |
| 2019 | RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual AnnotationsabstractAccurate surgery duration estimation is necessary for optimal OR planning, which plays an important role in patient comfort and safety as well as resource optimization. It is, however, challenging to preoperatively predict surgery duration since it varies significantly depending on the patient condition, surgeon skills, and intraoperative situation. In this paper, we propose a deep learning pipeline, referred to as RSDNet, which automatically estimates the remaining surgery duration (RSD) intraoperatively by using only visual information from laparoscopic videos. The previous state-of-the-art approaches for RSD prediction are dependent on manual annotation, whose generation requires expensive expert knowledge and is time-consuming, especially considering the numerous types of surgeries performed in a hospital and the large number of laparoscopic videos available. A crucial feature of RSDNet is that it does not depend on any manual annotation during training, making it easily scalable to many kinds of surgeries. The generalizability of our approach is demonstrated by testing the pipeline on two large datasets containing different types of surgeries: 120 cholecystectomy and 170 gastric bypass videos. The experimental results also show that the proposed network significantly outperforms a traditional method of estimating RSD without utilizing manual annotation. Further, this paper provides a deeper insight into the deep learning network through visualization and interpretation of the features that are automatically learned. Andru Putra Twinanda, Gaurav Yengera, Didier Mutter, Jacques Marescaux, Nicolas Padoy |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Pose optimization of a C-arm imaging device to reduce intraoperative radiation exposure of staff and patient during interventional proceduresabstractMinimally-invasive (MI) procedures are becoming more popular and frequent due to their benefits such as reduced patient trauma and hospitalization time. However, several common types of MI interventions are performed under X-ray guidance, which exposes both patients and staff to harmful ionizing radiation. Radiation exposure has therefore become a major concern for the medical community. Yet, few efforts to actively reduce it by exploiting the robotic capabilities of the devices present in the surgical suite have been performed. The propagation of radiation highly depends on the X-ray source positioning. Hence, we propose an approach to optimize the imaging device's pose in order to reduce the exposure to radiation of both patient and staff, while preserving the visibility of the targeted anatomical structure in the acquired image. Our method is based on the optimization of a cost function, which takes the current context and device parameters into account to compute the overall radiation exposure. It relies on GPU-accelerated Monte Carlo methods to simulate radiation propagation and performs the optimization in quasi real-time. When evaluated on a set of standard imaging configurations, our approach is able to recommend a device's pose in a few seconds, for which the delivered dose is reduced. Such an approach can contribute to lower the probability of appearance and severity of long-term negative effects due to radiation exposure and improve overall radiation safety. Nicolas Loy Rodas, Julien Bert, Dimitris Visvikis, Michel de Mathelin, Nicolas Padoy |
ICRA | 5 |
| 2017 | Deep Neural Networks Predict Remaining Surgery Duration from Cholecystectomy Videos
Ivan Aksamentov, Andru Putra Twinanda, Didier Mutter, Jacques Marescaux, Nicolas Padoy |
MICCAI (2) | 5 |
| 2017 | A Multi-view RGB-D Approach for Human Pose Estimation in Operating RoomsabstractMany approaches have been proposed for human pose estimation in single and multi-view RGB images. However, some environments, such as the operating room, are still very challenging for state-of-the-art RGB methods. In this paper, we propose an approach for multi-view 3D human pose estimation from RGB-D images and demonstrate the benefits of using the additional depth channel for pose refinement beyond its use for the generation of improved features. The proposed method permits the joint detection and estimation of the poses without knowing a priori the number of persons present in the scene. We evaluate this approach on a novel multi-view RGB-D dataset acquired during live surgeries and annotated with ground truth 3D poses. Abdolrahim Kadkhodamohammadi, Afshin Gangi, Michel de Mathelin, Nicolas Padoy |
WACV | 4 |
| 2017 | Articulated clinician detection using 3D pictorial structures on RGB-D data
Abdolrahim Kadkhodamohammadi, Afshin Gangi, Michel de Mathelin, Nicolas Padoy |
Medical Image Anal. | 4 |
| 2017 | EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic VideosabstractSurgical workflow recognition has numerous potential medical applications, such as the automatic indexing of surgical video databases and the optimization of real-time operating room scheduling, among others. As a result, surgical phase recognition has been studied in the context of several kinds of surgeries, such as cataract, neurological, and laparoscopic surgeries. In the literature, two types of features are typically used to perform this task: visual features and tool usage signals. However, the used visual features are mostly handcrafted. Furthermore, the tool usage signals are usually collected via a manual annotation process or by using additional equipment. In this paper, we propose a novel method for phase recognition that uses a convolutional neural network (CNN) to automatically learn features from cholecystectomy videos and that relies uniquely on visual information. In previous studies, it has been shown that the tool usage signals can provide valuable information in performing the phase recognition task. Thus, we present a novel CNN architecture, called EndoNet, that is designed to carry out the phase recognition and tool presence detection tasks in a multi-task manner. To the best of our knowledge, this is the first work proposing to use a CNN for multiple recognition tasks on laparoscopic videos. Experimental comparisons to other methods show that EndoNet yields state-of-the-art results for both tasks. Andru Putra Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel de Mathelin, Nicolas Padoy |
IEEE Trans. Medical Imaging | 6 |
| 2015 | Pictorial Structures on RGB-D Images for Human Pose Estimation in the Operating Room
Abdolrahim Kadkhodamohammadi, Afshin Gangi, Michel de Mathelin, Nicolas Padoy |
MICCAI (1) | 4 |
| 2015 | Marker-Less AR in the Hybrid Room Using Equipment Detection for Camera Relocalization
Nicolas Loy Rodas, Fernando Barrera, Nicolas Padoy |
MICCAI (1) | 3 |
| 2014 | Piecewise Planar Decomposition of 3D Point Clouds Obtained from Multiple Static RGB-D CamerasabstractIn this paper, we address the problem of segmenting a 3D point cloud obtained from several RGB-D cameras into a set of 3D piecewise planar regions. This is a fundamental problem in computer vision, whose solution is helpful for further scene analysis, such as support inference and object localisation. In existing planar segmentation approaches for point clouds, the point cloud originates from a single RGB-D view. There is however a growing interest to monitor environments with computer vision setups that contain a set of calibrated 3D cameras located around the scene. To fully exploit the multi-view aspect of such setups, we propose in this paper a novel approach to perform the planar piecewise segmentation directly in 3D. This approach, called Voxel-MRF (V-MRF), is based on discrete 3D Markov random fields, whose nodes correspond to scene voxels and whose labels represent 3D planes. The voxelization of the scene permits to cope with noisy depth measurements, while the MRF formulation provides a natural handling of the 3D spatial constraints during the optimisation. The approach results in a decomposition of the scene into a set of 3D planar patches. A by-product of the method is also a joint planar segmentation of the original images into planar regions with consistent labels across the views. We demonstrate the advantages of our approach using a benchmark dataset of objects with known geometry. We also present qualitative results on challenging data acquired by a multi-camera system installed in two operating rooms. Fernando Barrera, Nicolas Padoy |
3DV | 2 |
| 2014 | 3D Global Estimation and Augmented Reality Visualization of Intra-operative X-ray Dose
Nicolas Loy Rodas, Nicolas Padoy |
MICCAI (1) | 2 |
| 2014 | Fisher Kernel Based Task Boundary Retrieval in Laparoscopic Database with Single Video Query
Andru Putra Twinanda, Michel de Mathelin, Nicolas Padoy |
MICCAI (3) | 3 |
| 2012 | Deformable Tracking of Textured Curvilinear ObjectsabstractThreads and wires are deformable 3-dimensional (3D) and curvilinear objects which are commonly manipulated by humans in various medical and manufacturing tasks. Several applications, including computerassisted evaluation, augmented reality guidance, and autonomous robotic manipulation [2, 3] would benefit from the real-time estimation of the 3D shapes of these deformable objects from images. This estimation is however challenging due to multiple factors: 1) little information is available within an image to visually detect and distinguish a curvilinear object due to its thin and usually uniform appearance; 2) different 3D shapes may lead to the same visual perception, even in a stereo setting in case portions of the objects lie in an epipolar plane; and 3) the motions and deformations can be large, depending on the stiffness of the object. Additionally, a tracking approach that can consistently track specific points along the object defined by their arclength, such as the extremities or midpoint, would be particularly useful in the aforementioned applications. To deal with visual ambiguities such as drift along the curve, we propose to texture the object with a coarse pattern of alternating colors and formulate the shape estimation as a deformable 1D template tracking problem. Tracking is expressed as an energy minimization over a set of control pointsQ parameterizing a 3D NURBS C3D modeling the object: Nicolas Padoy, Gregory D. Hager |
BMVC | 1 |
| 2012 | Statistical modeling and recognition of surgical workflow
Nicolas Padoy, Tobias Blum, Seyed-Ahmad Ahmadi, Hubertus Feußner, Marie-Odile Berger, Nassir Navab |
Medical Image Anal. | 1 |
| 2011 | Human-Machine Collaborative surgery using learned modelsabstractIn the future of surgery, tele-operated robotic assistants will offer the possibility of performing certain commonly occurring tasks autonomously. Using a natural division of tasks into subtasks, we propose a novel surgical Human-Machine Collaborative (HMC) system in which portions of a surgical task are performed autonomously under complete surgeon's control, and other portions manually. Our system automatically identifies the completion of a manual subtask, seamlessly executes the next automated task, and then returns control back to the surgeon. Our approach is based on learning from demonstration. It uses Hidden Markov Models for the recognition of task completion and temporal curve averaging for learning the executed motions. We demonstrate our approach using a da Vinci tele-surgical robot. We show on two illustrative tasks where such human-machine collaboration is intuitive that automated control improves the usage of the master manipulator workspace. Because such a system does not limit the traditional use of the robot, but merely enhances its capabilities while leaving full control to the surgeon, it provides a safe and acceptable solution for surgical performance enhancement. Nicolas Padoy, Gregory D. Hager |
ICRA | 1 |
| 2011 | 3D thread tracking for robotic assistance in tele-surgeryabstractRemote tele-manipulation tasks can be both long and exhausting. The operative workload can however be reduced through contextual systems, in which routine or dexterous actions are performed automatically. In this paper, we investigate this idea in tele-surgery by proposing automatic scissors, namely the possibility for a surgeon to invoke a third robotic arm to come and automatically cut the thread that he/she is holding. In particular, we address the problem of tracking deformable 3-dimensional (3D) curvilinear objects from stereo images. We propose an approach based on discrete Markov random field (MRF) optimization to track, in 3D, a thread modeled by a non-uniform rational B-spline (NURBS). We evaluate its accuracy off-line on synthetic and real data and illustrate its use for an automatic scissors command within an assistance system based on the da Vinci tele-surgical robot. Nicolas Padoy, Gregory D. Hager |
IROS | 1 |
| 2011 | Spatio-Temporal Registration of Multiple Trajectories
Nicolas Padoy, Gregory D. Hager |
MICCAI (1) | 1 |
| 2009 | Wavelet energy map: A robust support for multi-modal registration of medical imagesabstractMulti-modal registration is the task of aligning images from an object acquired with different imaging systems, sensors or parameters. The current gold standard for medical images is the maximization of mutual information by computing the joint intensity distribution. However intensities are highly sensitive to various kinds of noise and denoising is a very challenging task often involving a priori knowledge and parameter tuning. We propose to perform registration on a novel robust information support: the wavelet energy map, giving a measure of local energy for each pixel. This spatial feature is derived from local spectral components computed with a redundant wavelet transform. The multi-frequential aspect of our method is particularly adapted to robust registration of images showing ambiguities such as tissues, complex textures and multiple interfaces. We show the benefits of the wavelet energy map approach in comparison to the classical framework in 2D and 3D rigid registration experiments on synthetic and real data. Olivier Pauly, Nicolas Padoy, Holger Poppert, Lorena Esposito, Nassir Navab |
CVPR | 2 |
| 2008 | On-line Recognition of Surgical Activity for Monitoring in the Operating Room
Nicolas Padoy, Tobias Blum, Hubertus Feußner, Marie-Odile Berger, Nassir Navab |
AAAI | 1 |
| 2008 | Modeling and Online Recognition of Surgical Phases Using Hidden Markov Models
Tobias Blum, Nicolas Padoy, Hubertus Feußner, Nassir Navab |
MICCAI (2) | 2 |
| 2007 | A Boosted Segmentation Method for Surgical Workflow Analysis
Nicolas Padoy, Tobias Blum, Irfan A. Essa, Hubertus Feußner, Marie-Odile Berger, Nassir Navab |
MICCAI (1) | 1 |
| 2006 | New CTA Protocol and 2D-3D Registration Method for Liver Catheterization
Martin Groher, Nicolas Padoy, Tobias F. Jakobs, Nassir Navab |
MICCAI (1) | 2 |
| 2006 | Messages Scheduling for Parallel Data Redistribution between ClustersabstractWe study the problem of redistributing data between clusters interconnected by a backbone. We suppose that at most k communications can be performed at the same time (the value of k depending on the characteristics of the platform). Given a set of messages, we aim at minimizing the total communication time assuming that communications can be preempted and that preemption comes with an extra cost. Our problem, called k-preemptive bipartite scheduling (KPBS) is proven to be NP-hard. We study its lower bound. We propose two 8/3-approximation algorithms with low complexity and fast heuristics. Simulation results show that both algorithms perform very well compared to the optimal solution and to the heuristics. Experimental results, based on an MPI implementation of these algorithms, show that both algorithms outperform a brute-force TCP-based solution, where no scheduling of the messages is performed Johanne Cohen, Emmanuel Jeannot, Nicolas Padoy, Frédéric Wagner |
IEEE Trans. Parallel Distributed Syst. | 3 |