Omid Mohareri

dblp:82/11496 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 2 since 2021Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SurgLaVi: Large-scale hierarchical dataset for surgical vision-language representation learning
Alejandra Pérez, Chinedu Innocent Nwoye, Ramtin Raji Kermani, Omid Mohareri, Muhammad Abdullah Jamal
Medical Image Anal.4
2025 Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD Datasets
abstract
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the model using contrastive learning to learn cross-modal representations. In the second stage, we further pre-train the model using masked autoencoding and denoising/noise prediction used in diffusion models. Masked autoencoding focuses on reconstructing the missing patches in the input modality using local spatial correlations, while denoising learns high frequency components of the input data. Moreover, it incorporates global distillation in the second stage by leveraging the knowledge acquired in stage one. Our approach is scalable, robust and suitable for pre-training RGB-D datasets. Extensive experiments on multiple datasets such as ScanNet, NYUv2 and SUN RGB-D show the efficacy and superior performance of our approach. Specifically, we show an improvement of +1.3% mIoU against Mask3D on ScanNet semantic segmentation. We further demonstrate the effectiveness of our approach in low-data regime by evaluating it for semantic segmentation task against the state-of-the-art methods.
Muhammad Abdullah Jamal, Omid Mohareri
CVPR2
2024 M33D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding
abstract
We present a new pre-training strategy called M33D (Multi-Modal Masked 3D) built based on Multi-modal masked autoencoders that can leverage 3D priors and learned cross-modal representations in RGB-D data. We integrate two major self-supervised learning frameworks; Masked Image Modeling (MIM) and contrastive learning; aiming to effectively embed masked 3D priors and modality complementary features to enhance the correspondence between modalities. In contrast to recent approaches which are either focusing on specific downstream tasks or require multi-view correspondence, we show that our pre-training strategy is ubiquitous, enabling improved representation learning that can transfer into improved performance on various downstream tasks such as video action recognition, video action detection, 2D semantic segmentation and depth estimation. Experiments show that M33D outperforms the existing state-of-the-art approaches on ScanNet, NYUv2, UCF-101 and OR-AR, particularly with an improvement of +1.3% mIoU against Mask3D on ScanNet semantic segmentation. We further evaluate our method on low-data regime and demonstrate its superior data efficiency compared to current state-of-the-art approaches.
Muhammad Abdullah Jamal, Omid Mohareri
WACV2
2024 Tracking and mapping in medical computer vision: A review
abstract
As computer vision algorithms increase in capability, their applications in clinical systems will become more pervasive. These applications include: diagnostics, such as colonoscopy and bronchoscopy; guiding biopsies, minimally invasive interventions, and surgery; automating instrument motion; and providing image guidance using pre-operative scans. Many of these applications depend on the specific visual nature of medical scenes and require designing algorithms to perform in this environment. In this review, we provide an update to the field of camera-based tracking and scene mapping in surgery and diagnostics in medical computer vision. We begin with describing our review process, which results in a final list of 515 papers that we cover. We then give a high-level summary of the state of the art and provide relevant background for those who need tracking and mapping for their clinical applications. After which, we review datasets provided in the field and the clinical needs that motivate their design. Then, we delve into the algorithmic side, and summarize recent developments. This summary should be especially useful for algorithm designers and to those looking to understand the capability of off-the-shelf methods. We maintain focus on algorithms for deformable environments while also reviewing the essential building blocks in rigid tracking and mapping since there is a large amount of crossover in methods. With the field summarized, we discuss the current state of the tracking and mapping methods along with needs for future algorithms, needs for quantification, and the viability of clinical applications. We then provide some research directions and questions. We conclude that new methods need to be designed or combined to support clinical applications in deformable environments, and more focus needs to be put into collecting datasets for training and evaluation.
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Michael C. Yip, Tim Salcudean
Medical Image Anal.2
2024 Surgical Tattoos in Infrared: A Dataset for Quantifying Tissue Tracking and Mapping
abstract
Quantifying performance of methods for tracking and mapping tissue in endoscopic environments is essential for enabling image guidance and automation of medical interventions and surgery. Datasets developed so far either use rigid environments, visible markers, or require annotators to label salient points in videos after collection. These are respectively: not general, visible to algorithms, or costly and error-prone. We introduce a novel labeling methodology along with a dataset that uses said methodology, Surgical Tattoos in Infrared (STIR). STIR has labels that are persistent but invisible to visible spectrum algorithms. This is done by labelling tissue points with IR-fluorescent dye, indocyanine green (ICG), and then collecting visible light video clips. STIR comprises hundreds of stereo video clips in both in vivo and ex vivo scenes with start and end points labelled in the IR spectrum. With over 3,000 labelled points, STIR will help to quantify and enable better analysis of tracking and mapping methods. After introducing STIR, we analyze multiple different frame-based tracking methods on STIR using both 3D and 2D endpoint error and accuracy metrics. STIR is available at https://dx.doi.org/10.21227/w8g4-g548.
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean
IEEE Trans. Medical Imaging2
2023 SENDD: Sparse Efficient Neural Depth and Deformation for Tissue Tracking
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean
MICCAI (9)2
2022 Fast Graph Refinement and Implicit Neural Representation for Tissue Tracking
abstract
Tracking of tissue in the surgical environment is often done via locating frame-to-frame keypoint correspondences, and then using these correspondences to warp a prior underlying model such as a spline, mesh, or embedded deformation. We introduce a novel learned model which takes keypoint correspondences as input and enables a prior-free estimation of deformation at any location. For fast point tracking, our model allows for sparse queries, unlike dense grid based CNNs, which run on full images. Our model begins with a novel graph-based point refinement scheme which refines matched keypoints, updating their features and movement instead of discarding possible outliers. Then, we use these refined matches to learn a novel neural implicit representation for estimating movement of any location given its k-nearest neighbor (k-NN) keypoints. We name our implicit deformation model KINFlow (k-NN implicit neural flow). We demonstrate the performance of KINFlow photometrically on three different datasets. KINFlow is the first model to use a graph network to estimate flow of arbitrary query points, and can estimate movement of 1024 points in under 3 ms.
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean
ICRA2
2022 Multi-modal Unsupervised Pre-training for Surgical Operating Room Workflow Analysis
Muhammad Abdullah Jamal, Omid Mohareri
MICCAI (8)2
2022 Adaptation of Surgical Activity Recognition Models Across Operating Rooms
Ali Mottaghi, Aidean Sharghi, Serena Yeung-Levy, Omid Mohareri
MICCAI (8)4
2022 Recurrent Implicit Neural Graph for Deformable Tracking in Endoscopic Videos
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean
MICCAI (4)2
2021 Multi-view Surgical Video Action Detection via Mixed Global View Attention
Adam Schmidt, Aidean Sharghi, Helene Haugerud, Daniel Oh, Omid Mohareri
MICCAI (4)5
2020 Automatic Operating Room Surgical Activity Recognition for Robot-Assisted Surgery
Aidean Sharghi, Helene Haugerud, Daniel Oh, Omid Mohareri
MICCAI (3)4
2020 A partial augmented reality system with live ultrasound and registered preoperative MRI for guiding robot-assisted radical prostatectomy
Golnoosh Samei, Keith Tsang, Claudia Kesch, Julio Lobo, Soheil Hor, Omid Mohareri, Silvia D. Chang, Larry Goldenberg, Peter C. Black, Tim Salcudean
Medical Image Anal.6
2018 Real-Time FEM-Based Registration of 3-D to 2.5-D Transrectal Ultrasound Images
abstract
We present a novel technique for real-time deformable registration of 3-D to 2.5-D transrectal ultrasound (TRUS) images for image-guided, robot-assisted laparoscopic radical prostatectomy (RALRP). For RALRP, a pre-operatively acquired 3-D TRUS image is registered to thin-volumes comprised of consecutive intra-operative 2-D TRUS images, where the optimal transformation is found using a gradient descent method based on analytical first and second order derivatives. Our method relies on an efficient algorithm for real-time extraction of arbitrary slices from a 3-D image deformed given a discrete mesh representation. We also propose and demonstrate an evaluation method that generates simulated models and images for RALRP by modeling tissue deformation through patient-specific finite-element models (FEM). We evaluated our method on in-vivo data from 11 patients collected during RALRP and focal therapy interventions. In the presence of an average landmark deformation of 3.89 and 4.62 mm, we achieved accuracies of 1.15 and 0.72 mm, respectively, on the synthetic and in-vivo data sets, with an average registration computation time of 264 ms, using MATLAB on a conventional PC. The results show that the real-time tracking of the prostate motion and deformation is feasible, enabling a real-time augmented reality-based guidance system for RALRP.].
Golnoosh Samei, Orcun Goksel, Julio Lobo, Omid Mohareri, Peter C. Black, Robert Rohling, Tim Salcudean
IEEE Trans. Medical Imaging4
2016 Bimanual teleoperation with heart motion compensation on the da Vinci® Research Kit: Implementation and preliminary experiments
abstract
This paper describes the implementation of a heart motion compensation system on the da Vinci surgical system (Intuitive Surgical Inc.) with the da Vinci Research Kit, for the purpose of simulating minimally invasive coronary artery bypass surgery on the beating heart. A Novint Falcon device is used to simulate the motion of the coronary bypass site. The 3D position of the heart simulator is measured optically in real time using an NDI Optotrak Certus tracker. The NDI measurements are used to command the da Vinci patient side manipulators to track, from a fixed distance, the simulated heart surface. Visual stabilization is achieved by having a stereo endoscopic camera carried by another patient-side manipulator that is also tracking the heart surface. The system performance is evaluated and discussed. In a preliminary evaluation study, surgeons were asked to perform a bimanual teleoperation suturing task on the simulated heart, with the gold standard being a suturing task on a stationary surface. When motion compensation was enabled, the median completion time dropped from 1.24 to 1.07 of the gold standard, the number of errors was reduced, and subjective measures show higher preference for the use of motion compensation in the da Vinci controllers.
Angelica Ruszkowski, Caitlin Schneider, Omid Mohareri, Tim Salcudean
ICRA3
2015 On the feasibility of heart motion compensation on the daVinci® surgical robot for coronary artery bypass surgery: Implementation and user studies
abstract
This paper describes the implementation of a heart motion compensation system on the da Vinci surgical system (Intuitive Surgical Inc.) for coronary artery bypass surgery. By introducing a robot-assisted solution, this surgery could be performed completely minimally invasively and on a beating heart. In this work we describe the development of open loop controllers based on spectral line decomposition and the assumption of a periodic trajectory. This allows the da Vinci patient-side manipulators to track an actual heart trajectory with sub-millimetre error. Further, to simulate a virtually stabilized environment, we present the novel concept of maintaining the camera fixed relative to the heart target, effectively decoupling the vision tracking and arm tracking challenges. Finally, we executed preliminary experiments to evaluate surgeons' ability to perform simulated suturing and peg transfer tasks on a moving target. Performance for the simulated suturing was evaluated based on task completion time, accuracy of needle placement, and number of errors. For the suture task, the number of missed targets decreased from 37% to 13% when compensation was enabled, the number of hit targets increased from 26% to 41%, and completion time decreased. For the peg transfer tasks, again completion time and number of errors were measured. Though the margin for error was larger, there was less perceived difficulty of the task when compensation was enabled.
Angelica Ruszkowski, Omid Mohareri, Samuel Victor Lichtenstein, Richard Cook, Tim Salcudean
ICRA2
2015 A retrofit eye gaze tracker for the da Vinci and its integration in task execution using the da Vinci Research Kit
abstract
The integration of eye-gaze tracking at the console of surgical robots has the potential to add both speed and functionality to the human-robot interface. In this paper, we present a novel eye gaze tracker that can be integrated as a simple retrofit to the da Vinci console. In particular, the eye tracker can be used with the da Vinci Research Kit (dVRK) to control the patient side manipulators. We present the eye-tracker design and calibration. First, a 2D calibration is carried out to estimate the gaze on the da Vinci's stereoscopic display, followed by a 3D “hand-eye” calibration to estimate the gaze in the surgical scene. Using the dVRK, we demonstrated the use of the eye-tracker to perform a gaze-assisted task, in which the estimated gaze is used as a set point for the robot and the user is guided towards the point of gaze through haptic feedback. The task performed was peg placement and was evaluated in a preliminary user study with five users.
Irene Tong, Omid Mohareri, Samuel Tatasurya, Craig Hennessey, Tim Salcudean
IROS2
2015 A System for MR-Ultrasound Guidance during Robot-Assisted Laparoscopic Radical Prostatectomy
Omid Mohareri, Guy Nir, Julio Lobo, Richard Savdie, Peter C. Black, Tim Salcudean
MICCAI (1)1
2014 Bimanual telerobotic surgery with asymmetric force feedback: A daVinci® surgical system implementation
abstract
This paper describes the applicability of an asymmetric force feedback control framework for bimanual robot-assisted surgery using the da Vinci surgical system (Intuitive Surgical Inc.). The core idea of this method, previously presented in [1], is that when completing two-handed tasks involving an action and a reaction force, the forces applied on the environment by the action hand are not transferred back to the same hand, but rather to the reaction hand. Such a method provides an intuitive way of feeling the force, while avoiding the instability issues, since the control loop in not closed from the slave to the master of the same hand. In the introductory paper [1], the technique was implemented using game controllers with simple tasks. In this paper, the technique was implemented on the da Vinci surgical system (Classic version) using the da Vinci Research Kit (dVRK) controllers that enable complete access to all control levels of the da Vinci robot manipulators via custom mechatronics and open-source software. The implementation involved a full re-write of a teleoperation controller based on kinematic correspondence with gravity compensation, as well as torque control functions for force rendering on the da Vinci master manipulators. A series of suture knot tying and haptic exploration experiments were conducted in which a small group of users, both surgeons (N=3) and novices (N=6) evaluated the system. The results show that the proposed technique has some promise when implemented in a realistic 14 degrees of freedom system, but further work is necessary to make the system fully usable.
Omid Mohareri, Caitlin Schneider, Tim Salcudean
IROS1
2014 Multi-parametric 3D Quantitative Ultrasound Vibro-Elastography Imaging for Detecting Palpable Prostate Tumors
Omid Mohareri, Angelica Ruszkowski, Julio Lobo, Joseph Ischia, Ali Baghani, Guy Nir, Hani Eskandari, Edward C. Jones, Ladan Fazli, Larry Goldenberg, Mehdi Moradi, Tim Salcudean
MICCAI (1)1
2013 Asymmetric force feedback control framework for teleoperated robot-assisted surgery
abstract
Lack of haptic feedback in teleoperated robot-assisted surgery (RAS) is known to be detrimental in many surgical tasks. While performing a class of force sensitive tasks, surgeons commonly use both hands. Oftentimes, one hand is used to exert tension/compression forces, and the other to hold the suture knot or tissue. A novel control framework to accomplish haptic force feedback for two-handed tasks in teleoperated RAS is presented in this paper. The force applied on the surgical environment by the action hand is not transferred back to the same hand, but rather to the other hand. In two-handed tasks that involve an action and a reaction force, this provides an intuitive way of feeling the action. Because the loop in not closed from the slave back to the master of the same hand, it does not have a destabilizing effect. The technique can be easily implemented using a variable-structure controller that combines two PD controllers and a switch. It has been evaluated with an experimental setup consisting of four haptic devices with promising results.
Omid Mohareri, Tim Salcudean, Christopher Y. Nguan
ICRA1
2012 Indirect adaptive tracking control of a nonholonomic mobile robot via neural networks
Omid Mohareri, Rached Dhaouadi, Ahmad B. Rad
Neurocomputing1