Jacques Marescaux

dblp:18/1992 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-1400-6230ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 since 2021Artificial intelligence and machine learning · 4Systems, architecture and hardware · 3Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2025 SAMUSA: Segment Anything Model 2 for UltraSound Annotation
Baptiste Podvin, Toby Collins, Güinther Saibro, Chiara Innocenzi, Flavio Milana, Yvonne Keeza, Grace Ufitinema, Florien Ujemurwego, Guido Torzilli, Jacques Marescaux, Daniel George, Alexandre Hostettler
MICCAI (11)11
2025 Learning multi-modal representations by watching hundreds of surgical video lectures
abstract
Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. SurgVLP constructs a new contrastive learning objective to align video clip embeddings with the corresponding multiple text embeddings by bringing them together within a joint latent space. To effectively demonstrate the representational capability of the learned joint latent space, we introduce several vision-and-language surgical tasks and evaluate various vision-only tasks specific to surgery, e.g., surgical tool, phase, and triplet recognition. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The code is available at https://github.com/CAMMA-public/SurgVLP.
Kun Yuan 0004, Vinkle Srivastav, Tong Yu 0009, Joël L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy
Medical Image Anal.5
2023 Live laparoscopic video retrieval with compressed uncertainty
Tong Yu 0009, Pietro Mascagni, Juan Verde, Jacques Marescaux, Didier Mutter, Nicolas Padoy
Medical Image Anal.4
2023 Weakly Supervised Temporal Convolutional Networks for Fine-Grained Surgical Activity Recognition
abstract
Automatic recognition of fine-grained surgical activities, called steps, is a challenging but crucial task for intelligent intra-operative computer assistance. The development of current vision-based activity recognition methods relies heavily on a high volume of manually annotated data. This data is difficult and time-consuming to generate and requires domain-specific knowledge. In this work, we propose to use coarser and easier-to-annotate activity labels, namely phases, as weak supervision to learn step recognition with fewer step annotated videos. We introduce a step-phase dependency loss to exploit the weak supervision signal. We then employ a Single-Stage Temporal Convolutional Network (SS-TCN) with a ResNet-50 backbone, trained in an end-to-end fashion from weakly annotated videos, for temporal activity segmentation and recognition. We extensively evaluate and show the effectiveness of the proposed method on a large video dataset consisting of 40 laparoscopic gastric bypass procedures and the public benchmark CATARACTS containing 50 cataract surgeries.
Sanat Ramesh, Diego Dall'Alba, Cristians Gonzalez, Tong Yu 0009, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Paolo Fiorini, Nicolas Padoy
IEEE Trans. Medical Imaging7
2022 Automatic Detection of Steatosis in Ultrasound Images with Comparative Visual Labeling
Güinther Saibro, Michele Diana, Benoît Sauer, Jacques Marescaux, Alexandre Hostettler, Toby Collins
MICCAI (3)4
2022 Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos
Chinedu Innocent Nwoye, Tong Yu 0009, Cristians Gonzalez, Barbara Seeliger, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Nicolas Padoy
Medical Image Anal.7
2020 Recognition of Instrument-Tissue Interactions in Endoscopic Videos via Action Triplets
abstract
Recognition of surgical activity is an essential component to develop context-aware decision support for the operating room. In this work, we tackle the recognition of fine-grained activities, modeled as action triplets representing the tool activity. To this end, we introduce a new laparoscopic dataset, CholecT40, consisting of 40 videos from the public dataset Cholec80 in which all frames have been annotated using 128 triplet classes. Furthermore, we present an approach to recognize these triplets directly from the video data. It relies on a module called Class Activation Guide (CAG), which uses the instrument activation maps to guide the verb and target recognition. To model the recognition of multiple triplets in the same frame, we also propose a trainable 3D Interaction Space, which captures the associations between the triplet components. Finally, we demonstrate the significance of these contributions via several ablation studies and comparisons to baselines on CholecT40.
Chinedu Innocent Nwoye, Cristians Gonzalez, Tong Yu 0009, Pietro Mascagni, Didier Mutter, Jacques Marescaux, Nicolas Padoy
MICCAI (3)6
2020 Future-State Predicting LSTM for Early Surgery Type Recognition
abstract
This work presents a novel approach for the early recognition of the type of a laparoscopic surgery from its video. Early recognition algorithms can be beneficial to the development of "smart" OR systems that can provide automatic context-aware assistance, and also enable quick database indexing. The task is however ridden with challenges specific to videos belonging to the domain of laparoscopy, such as high visual similarity across surgeries and large variations in video durations. To capture the spatio-temporal dependencies in these videos, we choose as our model a combination of a convolutional neural network (CNN) and long short-term memory (LSTM) network. We then propose two complementary approaches for improving early recognition performance. The first approach is a CNN fine-tuning method that encourages surgeries to be distinguished based on the initial frames of laparoscopic videos. The second approach, referred to as " Future-State Predicting LSTM," trains an LSTM to predict information related to future frames, which helps in distinguishing between the different types of surgeries. We evaluate our approaches on a large dataset of 425 laparoscopic videos containing nine types of surgeries (Laparo425), and achieve on average an accuracy of 75% having observed only the first 10 min of a surgery. These results are quite promising from a practical standpoint and also encouraging for other types of image-guided surgeries.
Siddharth Kannan, Gaurav Yengera, Didier Mutter, Jacques Marescaux, Nicolas Padoy
IEEE Trans. Medical Imaging4
2019 RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual Annotations
abstract
Accurate surgery duration estimation is necessary for optimal OR planning, which plays an important role in patient comfort and safety as well as resource optimization. It is, however, challenging to preoperatively predict surgery duration since it varies significantly depending on the patient condition, surgeon skills, and intraoperative situation. In this paper, we propose a deep learning pipeline, referred to as RSDNet, which automatically estimates the remaining surgery duration (RSD) intraoperatively by using only visual information from laparoscopic videos. The previous state-of-the-art approaches for RSD prediction are dependent on manual annotation, whose generation requires expensive expert knowledge and is time-consuming, especially considering the numerous types of surgeries performed in a hospital and the large number of laparoscopic videos available. A crucial feature of RSDNet is that it does not depend on any manual annotation during training, making it easily scalable to many kinds of surgeries. The generalizability of our approach is demonstrated by testing the pipeline on two large datasets containing different types of surgeries: 120 cholecystectomy and 170 gastric bypass videos. The experimental results also show that the proposed network significantly outperforms a traditional method of estimating RSD without utilizing manual annotation. Further, this paper provides a deeper insight into the deep learning network through visualization and interpretation of the features that are automatically learned.
Andru Putra Twinanda, Gaurav Yengera, Didier Mutter, Jacques Marescaux, Nicolas Padoy
IEEE Trans. Medical Imaging4
2018 Soft-Body Registration of Pre-operative 3D Models to Intra-operative RGBD Partial Body Scans
Richard Modrzejewski, Toby Collins, Adrien Bartoli, Alexandre Hostettler, Jacques Marescaux
MICCAI (4)5
2017 Deep Neural Networks Predict Remaining Surgery Duration from Cholecystectomy Videos
Ivan Aksamentov, Andru Putra Twinanda, Didier Mutter, Jacques Marescaux, Nicolas Padoy
MICCAI (2)4
2017 EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos
abstract
Surgical workflow recognition has numerous potential medical applications, such as the automatic indexing of surgical video databases and the optimization of real-time operating room scheduling, among others. As a result, surgical phase recognition has been studied in the context of several kinds of surgeries, such as cataract, neurological, and laparoscopic surgeries. In the literature, two types of features are typically used to perform this task: visual features and tool usage signals. However, the used visual features are mostly handcrafted. Furthermore, the tool usage signals are usually collected via a manual annotation process or by using additional equipment. In this paper, we propose a novel method for phase recognition that uses a convolutional neural network (CNN) to automatically learn features from cholecystectomy videos and that relies uniquely on visual information. In previous studies, it has been shown that the tool usage signals can provide valuable information in performing the phase recognition task. Thus, we present a novel CNN architecture, called EndoNet, that is designed to carry out the phase recognition and tool presence detection tasks in a multi-task manner. To the best of our knowledge, this is the first work proposing to use a CNN for multiple recognition tasks on laparoscopic videos. Experimental comparisons to other methods show that EndoNet yields state-of-the-art results for both tasks.
Andru Putra Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel de Mathelin, Nicolas Padoy
IEEE Trans. Medical Imaging4
2016 Automatic localization of endoscope in intraoperative CT image: A simple approach to augmented reality guidance in laparoscopic surgery
abstract
The use of augmented reality in minimally invasive surgery has been the subject of much research for more than a decade. The endoscopic view of the surgical scene is typically augmented with a 3D model extracted from a preoperative acquisition. However, the organs of interest often present major changes in shape and location because of the pneumoperitoneum and patient displacement. There have been numerous attempts to compensate for this distortion between the pre- and intraoperative states. Some have attempted to recover the visible surface of the organ through image analysis and register it to the preoperative data, but this has proven insufficiently robust and may be problematic with large organs. A second approach is to introduce an intraoperative 3D imaging system as a transition. Hybrid operating rooms are becoming more and more popular, so this seems to be a viable solution, but current techniques require yet another external and constraining piece of apparatus such as an optical tracking system to determine the relationship between the intraoperative images and the endoscopic view. In this article, we propose a new approach to automatically register the reconstruction from an intraoperative CT acquisition with the static endoscopic view, by locating the endoscope tip in the volume data. We first describe our method to localize the endoscope orientation in the intraoperative image using standard image processing algorithms. Secondly, we highlight that the axis of the endoscope needs a specific calibration process to ensure proper registration accuracy. In the last section, we present quantitative and qualitative results proving the feasibility and the clinical potential of our approach.
Sylvain Bernhardt, Stéphane Nicolau, Vincent Agnus, Luc Soler, Christophe Doignon, Jacques Marescaux
Medical Image Anal.6
2013 Inter-operative Trajectory Registration for Endoluminal Video Synchronization: Application to Biopsy Site Re-localization
Anant Suraj Vemuri, Stéphane Nicolau, Nicholas Ayache, Jacques Marescaux, Luc Soler
MICCAI (1)4
2012 Simulation of Pneumoperitoneum for Laparoscopic Surgery Planning
Jordan Bano, Alexandre Hostettler, Stéphane Nicolau, Stephane Cotin, Christophe Doignon, H. S. Wu, M. H. Huang, Luc Soler, Jacques Marescaux
MICCAI (1)9
2009 An augmented reality system for liver thermal ablation: Design and evaluation on clinical cases
Stéphane Nicolau, Xavier Pennec, Luc Soler, Xavier Buy, Afshin Gangi, Nicholas Ayache, Jacques Marescaux
Medical Image Anal.7
2006 A Modular and Evolutive Software for Patient Modeling Using Components, Design Patterns and a Formal XML-Based Component Management System
abstract
This paper deals with the design aspect of a software for modeling the anatomical and pathological structures of patients from medical images, for diagnosis purposes. In terms of functionalities, it allows to combine image processing algorithms, and to visualize and manipulate 3D models and images. The proposed software uses appropriate design patterns, specific extensible and reusable components and a system managing their combination, thanks to a formal XML-based description of their interfaces. This architecture facilitates the dynamic integration of new functionalities, in particular in terms of image processing algorithms. We describe the structural and behavioral aspects of the proposed component-based architecture.
Jean-Baptiste Fasquel, Guillaume Brocker, Johan Moreau, Vincent Agnus, Nicolas Papier, Christophe Koehl, Luc Soler, Jacques Marescaux
CBMS8
2006 A Hierarchical Topological Knowledge Based Image Segmentation Approach Optimizing the use of Contextual Regions of Interest : Illustration for Medical Image Analysis
abstract
This paper concerns image segmentation and presents a method to automically determine optimal regions of interest (ROI) according to topological information. The use of ROI avoids the processing of irrelevant image points, therefore improving and accelerating segmentations. ROI determination is based on the optimal use of both the a priori knowledge about topological structure of an image and the contextual information. Contextual information concerns the nature of already segmented regions in the case of the hierarchical segmentation approach we consider. We describe this general purpose method and propose a formulation for the optimal determination of ROIs according to both informations. Then, we illustrate the use and the implementation of such a method in the particular case of medical image segmentation.
Jean-Baptiste Fasquel, Vincent Agnus, Luc Soler, Jacques Marescaux
ICIP4
2006 Computational Models for Image-Guided Robot-Assisted and Simulated Medical Interventions
abstract
Medical image analysis plays a crucial role in the diagnosis, planning, control, and follow-up of therapy. To be combined efficiently with medical robotics, medical image analysis can be supported by the development of specific computational models of the human body operating at various levels. We describe a hierarchy of these computational models, including the geometrical, physical, and physiological levels, and illustrate their potential use in a number of advanced medical applications including image-guided robot-assisted and simulated medical interventions. We conclude with scientific perspectives.
Hervé Delingette, Xavier Pennec, Luc Soler, Jacques Marescaux, Nicholas Ayache
Proc. IEEE4
2005 Active filtering of physiological motion in robotized surgery using predictive control
abstract
This work presents a predictive-control approach to active mechanical filtering of complex, periodic motions of organs induced by respiration or heart beating in robotized surgery. Two different predictive-control schemes are proposed for the compensation of respiratory motions or cardiac motions. For respiratory motions, the periodic property of the disturbance has been included into the input-output model of the controlled system so as to have the robotic system learn and anticipate perturbation motions. A new cost function is proposed for the unconstrained generalized predictive controller (GPC), where reference tracking is decoupled from the rejection of predictable periodic motions. Cardiac motions are more complex, since they are the combination of two periodic nonharmonic components. An adaptive disturbance predictor is proposed which outputs future predicted disturbance values. These predicted values are used to anticipate the disturbance by using the predictive feature of a regular GPC. Experimental results are presented on a laboratory testbed and in vivo on pigs. They demonstrate the effectiveness of the two proposed methods to compensate complex physiological motion.
Romuald Ginhoux, Jacques Gangloff, Michel de Mathelin, Luc Soler, Maria Mara Arenas Sanchez, Jacques Marescaux
IEEE Trans. Robotics6
2004 Beating Heart Tracking in Robotic Surgery using 500 Hz Visual Servoing, Model Predictive Control and an Adaptive Observer
abstract
This work presents first in-vivo results of beating heart tracking with a surgical robot arm in off-pump cardiac surgery. The tracking is performed in a 2D visual servoing scheme using a 500 frame per second video camera. Heart motion is measured by means of active optical markers that are put onto the heart surface. Amplitude of the motion is evaluated along the two axis of the image reference frame. This is a complex and fast motion that mainly reflects the influence of both the respiratory motion and the electro-mechanical activity of the myocardium. A model predictive controller is setup to track the two degrees of freedom of the observed motion by computing velocities for two of the robot joints. The servoing scheme takes advantage of the ability of predictive control to anticipate over future references provided they are known or they can be predicted. An adaptive observer is defined along with a simple cardiac model to estimate the two components of the heart motion. The predictions are then fed into the controller references and it is shown that the tracking behaviour is greatly improved.
Romuald Ginhoux, Jacques Gangloff, Michel de Mathelin, Luc Soler, Maria Mara Arenas Sanchez, Jacques Marescaux
ICRA6
2004 Virtual Reality and Augmented Reality in Digestive Surgery
abstract
Medical image processing led to a major improvement of patient care: the 3D modeling of patients from their CT-scan or MRI provides an improved surgical planning and simulation allows to train the surgical gesture before carrying it out. These two preoperative steps can be used intra-operatively with the development of augmented reality (AR). In this paper, we present the tools we developed to provide our first prototypal AR guiding system for abdominal surgery.
Luc Soler, Stéphane Nicolau, Jérôme Schmid, Christophe Koehl, Jacques Marescaux, Xavier Pennec, Nicholas Ayache
ISMAR5
2003 A 500 Hz predictive visual servoing scheme to mechanically filter complex repetitive organ motions in robotized surgery
abstract
Periodic deformations of organs and soft tissues are complex, repetitive disturbances for surgeons manipulating robotic interfaces in computer-assisted surgery. They are due to respiratory movements or heart beats, and they have to be manually compensated for by the surgeon whenever accurate gestures are needed, as it is the case in cardiac or robotized laparoscopic surgery. This work presents a repetitive model predictive control scheme for the cancellation of fast periodic motions by a robot arm, which is controlled by visual servoing at 500 Hz by means of a high-speed camera. The problem we address is to keep a constant distance in the camera images from a surgical tool's tip to the organ surface. Contributions of the control input to reference tracking and to the fast-disturbance rejection are split and computed separately to ensure that the surgeon's interaction on the robot bas no influence on the cancellation performance. The system is tested in a laboratory experiment with an experimental surgical arm and in in vivo conditions on a living pig with a standard surgical robot. Results show the effectiveness and the potential of the proposed control scheme.
Romuald Ginhoux, Jacques Gangloff, Michel de Mathelin, Luc Soler, Joël Leroy, Jacques Marescaux
IROS6
2003 RF-Sim: a Treatment Planning Tool for Radiofrequency Ablation of Hepatic Tumors
abstract
With recent advancements of technology, radiofrequency ablation has become one of the most used techniques to treat liver tumors. But radiologists still have to face the difficulty of planning their treatment while only relying on 2D slices acquired from CT-scan. We present a tool called RF-Sim, being part of a complete 3D reconstruction and visualization project, and including both a realistic radiofrequency ablation simulator for training and rehearsal, and an automatic treatment planner taking into account tumor's environment. They help radiologists to have a better visualization of patients' anatomic structures and pathologies, and allow them to easily find an adequate treatment. They run on a common laptop and can be used in the operating room.
Caroline Essert, Luc Soler, Nicolas Papier, Vincent Agnus, Afshin Gangi, Didier Mutter, Jacques Marescaux
IV7
2003 Autonomous 3-D positioning of surgical instruments in robotized laparoscopic surgery using visual servoing
abstract
This paper presents a robotic vision system that automatically retrieves and positions surgical instruments during robotized laparoscopic surgical operations. The instrument is mounted on the end-effector of a surgical robot which is controlled by visual servoing. The goal of the automated task is to safely bring the instrument at a desired three-dimensional location from an unknown or hidden position. Light-emitting diodes are attached on the tip of the instrument, and a specific instrument holder fitted with optical fibers is used to project laser dots on the surface of the organs. These optical markers are detected in the endoscopic image and allow localizing the instrument with respect to the scene. The instrument is recovered and centered in the image plane by means of a visual servoing algorithm using feature errors in the image. With this system, the surgeon can specify a desired relative position between the instrument and the pointed organ. The relationship between the velocity screw of the surgical instrument and the velocity of the markers in the image is estimated online and, for safety reasons, a multistages servoing scheme is proposed. Our approach has been successfully validated in a real surgical environment by performing experiments on living tissues in the surgical training room of the Institut de Recherche sur les Cancers de l'Appareil Digestif (IRCAD), Strasbourg, France.
Alexandre Krupa, Jacques Gangloff, Christophe Doignon, Michel de Mathelin, Guillaume Morel, Joël Leroy, Luc Soler, Jacques Marescaux
IEEE Trans. Robotics Autom.8
2002 Autonomous Retrieval and Positioning of Surgical Instruments in Robotized Laparoscopic Surgery using Visual Servoing and Laser Pointers
abstract
This paper presents a robotic vision system that automatically retrieves and positions surgical instruments in robotized laparoscopic surgery. The surgical instrument is mounted on the end-effector of a surgical robot which can be controlled by automatic visual feedback. The goal of the automated task is to bring the instrument at a desired location from an unknown or hidden position. To achieve this task, a special instrument-holder is designed with optical fibers and collimators. This instrument-holder projects laser dot patterns onto the organ surface which are seen in the endoscopic images. Then, the instrument is retrieved and centered in the image plane using a visual servoing algorithm. With this system, the surgeon can also specify a desired position for the instrument in the image. Our approach is successfully validated in a real surgical environment by performing experiments on living animals in the surgical training room of IRCAD.
Alexandre Krupa, Jacques Gangloff, Michel de Mathelin, Christophe Doignon, Guillaume Morel, Luc Soler, Joël Leroy, Jacques Marescaux
ICRA8
2002 Automatic 3-D Positioning of Surgical Instruments during Robotized Laparoscopic Surgery Using Automatic Visual Feedback
Alexandre Krupa, Michel de Mathelin, Christophe Doignon, Jacques Gangloff, Guillaume Morel, Luc Soler, Joël Leroy, Jacques Marescaux
MICCAI (1)8
2001 Development of Semi-autonomous Control Modes in Laparoscopic Surgery Using Automatic Visual Servoing
Alexandre Krupa, Michel de Mathelin, Christophe Doignon, Jacques Gangloff, Guillaume Morel, Luc Soler, Jacques Marescaux
MICCAI7