Alejandro Martin-Gomez

dblp:171/9397 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-9341-3477ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Extend Your Horizon: A Device-Agnostic Surgical Tool Tracking Framework with Multi-View Optimization for Augmented Reality
abstract
Surgical navigation has proven to be an effective approach for providing real-time guidance and visualization of relevant information by estimating the pose of the patient’s anatomy and surgical tools. During navigated surgery, instruments are commonly equipped with fiducial markers and tracked by stationary optical tracking systems (OTS) to provide accurate navigation cues. Augmented Reality (AR) has been adopted for intuitive visual guidance, even motivating several efforts to enable surgical instrument tracking through built-in sensors on Head-Mounted Displays (HMDs). However, existing tracking methods typically require a direct line-of-sight to instruments, which is challenging to maintain in dynamic surgical environments due to frequent occlusions caused by moving medical equipment, surgical tools, and personnel. To address this challenge, this work introduces a novel framework capable of tracking surgical instruments even under occlusion by fusing different sensors in a dynamic scene graph representation. Our framework uniquely combines tracking systems with varying degrees of accuracy, providing real-time assessments of tracking reliability to the user. Unlike conventional sensor fusion approaches that are heavily dependent on specific sensor modalities, the proposed method is agnostic to tracking device modality and robust to their motion states (e.g., stationary OTS versus dynamic AR-HMD). Experimental results demonstrate that our dynamic scene graph framework successfully integrates and optimizes measurements from multiple tracking sources, significantly enhancing AR visualization consistency and accuracy with robustness under occlusion.
Mingxu Liu, Hongchao Shu, Ruixing Liang, Yihao Liu 0004, Ojas Taskar, Amir Kheradmand, Mehran Armand, Alejandro Martin-Gomez
VR9
2024 On the Fly Robotic-Assisted Medical Instrument Planning and Execution Using Mixed Reality
abstract
Robotic-assisted medical systems (RAMS) have gained significant attention for their advantages in alleviating surgeons’ fatigue and improving patients’ outcomes. These systems comprise a range of human-computer interactions, including medical scene monitoring, anatomical target planning, and robot manipulation. However, despite its versatility and effectiveness, RAMS demands expertise in robotics, leading to a high learning cost for the operator. In this work, we introduce a novel framework using mixed reality technologies to ease the use of RAMS. The proposed framework achieves real-time planning and execution of medical instruments by providing 3D anatomical image overlay, human-robot collision detection, and robot programming interface. These features, integrated with an easy-to-use calibration method for head-mounted display, improve the effectiveness of human-robot interactions. To assess the feasibility of the framework, two medical applications are presented in this work: 1) coil placement during transcranial magnetic stimulation and 2) drill and injector device positioning during femoroplasty. Results from these use cases demonstrate its potential to extend to a wider range of medical scenarios.
Letian Ai, Yihao Liu 0004, Mehran Armand, Amir Kheradmand, Alejandro Martin-Gomez
ICRA5
2024 Uncertainty-Aware Shape Estimation of a Surgical Continuum Manipulator in Constrained Environments using Fiber Bragg Grating Sensors
abstract
Continuum Dexterous Manipulators (CDMs) are well-suited tools for minimally invasive surgery due to their inherent dexterity and reachability. Nonetheless, their flexible structure and non-linear curvature pose significant challenges for shape-based feedback control. The use of Fiber Bragg Grating (FBG) sensors for shape sensing has shown great potential in estimating the CDM’s tip position and subsequently reconstructing the shape using optimization algorithms. This optimization, however, is under-constrained and may be ill-posed for complex shapes, falling into local minima. In this work, we introduce a novel method capable of directly estimating a CDM’s shape from FBG sensor wavelengths using a deep neural network. In addition, we propose the integration of uncertainty estimation to address the critical issue of uncertainty in neural network predictions. Neural network predictions are unreliable when the input sample is outside the training distribution or corrupted by noise. Recognizing such deviations is crucial when integrating neural networks within surgical robotics, as inaccurate estimations can pose serious risks to the patient. We present a robust method that not only improves the precision upon existing techniques for FBG-based shape estimation but also incorporates a mechanism to quantify the models’ confidence through uncertainty estimation. We validate the uncertainty estimation through extensive experiments, demonstrating its effectiveness and reliability on out-of-distribution (OOD) data, adding an additional layer of safety and precision to minimally invasive surgical robotics.
Alexander Schwarz, Arian Mehrfard, Golchehr Amirkhani, Henry Phalen, Justin H. Ma, Robert B. Grupp, Alejandro Martin-Gomez, Mehran Armand
ICRA7
2024 Calibration of Augmented Reality Headset with External Tracking System Using AX=YB
abstract
In Augmented Reality, a robust virtual-to-real calibration that aligns the virtual and real spaces is crucial to ensure accurate overlays. Inspired by popular methods in robotics, we propose establishing the virtual-to-real calibration as a hand-eye/robot-world calibration problem using an $A X=Y B$ formulation. This formulation uses both the self-localization of a head-mounted display and the tracking functionality of an external tracking system. Additional techniques are also provided to address the data synchronization issue between the two measurement systems. To further improve the results, we integrate a post-acquisition outlier filter based on the $A X-Y B$ Frobenius norm. Improvements resulting from the filter were first validated by simulation. For the assessment of the complete pipeline, both subjective evaluation based on human perception and objective evaluation based on computer vision methods were used. The results show that the proposed method is accurate and noise-resistant. With only 100 measurement samples collected, the overlay of a tracked object at the farthest distance (1300 mm) in front of the tracking system has an average rotation error of $1.77 \pm 0.54^{\circ}$ and an average translation error of $4.82 \pm 1.71 \mathrm{~mm}$.
Letian Ai, Yihao Liu 0004, Mehran Armand, Alejandro Martin-Gomez
ISMAR4
2024 ARthroNeRF: Field of View Enhancement of Arthroscopic Surgeries using Augmented Reality and Neural Radiance Fields
abstract
Arthroscopy is a minimally invasive orthopedic procedure commonly used to treat joints such as the shoulder, hip, or knee. A major difficulty for surgeons in arthroscopic procedures is the simultaneous coordination of the surgical tools used for manipulation and the arthroscopic camera. To further complicate this task, the narrow space where arthroscopic procedures are performed limits the ability to move the arthroscope inside the patient’s body, restricting the field of view. In this work, to overcome these limitations, we introduce ARthroNeRF, a novel framework that combines Neural Radiance Fields (NeRF) and Augmented Reality (AR). This framework allows for the generation of synthetic viewpoints from the perspective of surgical tools without the need for an additional camera. To evaluate the feasibility of the proposed framework, we conducted a user study with 18 participants. In this study, participants were tasked to touch hidden targets assisted by a synthetic view generated from the perspective of a surgical tool. The results of this study demonstrate that ARthroNeRF provides accurate supplementary visual information and suggest that ARthroNeRF has the potential to streamline the learning process in arthroscopic surgery. In addition, we built a system capable of presenting the reconstructed scenes using the Microsoft HoloLens 2. The incorporation of AR, overlaying synthesized images alongside the original arthroscopic footage and within the surgeon’s visual field, represents a viable alternative to enhance perception in spatially constrained scenarios.
Xinrui Zou, Alexander Schwarz, Mehran Armand, Alejandro Martin-Gomez
ISMAR5
2024 STTAR: Surgical Tool Tracking Using Off-the-Shelf Augmented Reality Head-Mounted Displays
abstract
The use of Augmented Reality (AR) for navigation purposes has shown beneficial in assisting physicians during the performance of surgical procedures. These applications commonly require knowing the pose of surgical tools and patients to provide visual information that surgeons can use during the performance of the task. Existing medical-grade tracking systems use infrared cameras placed inside the Operating Room (OR) to identify retro-reflective markers attached to objects of interest and compute their pose. Some commercially available AR Head-Mounted Displays (HMDs) use similar cameras for self-localization, hand tracking, and estimating the objects' depth. This work presents a framework that uses the built-in cameras of AR HMDs to enable accurate tracking of retro-reflective markers without the need to integrate any additional electronics into the HMD. The proposed framework can simultaneously track multiple tools without having previous knowledge of their geometry and only requires establishing a local network between the headset and a workstation. Our results show that the tracking and detection of the markers can be achieved with an accuracy of$0.09\pm 0.06\ mm$on lateral translation,$0.42 \pm 0.32\ mm$on longitudinal translation and$0.80 \pm 0.39^\circ$for rotations around the vertical axis. Furthermore, to showcase the relevance of the proposed framework, we evaluate the system's performance in the context of surgical procedures. This use case was designed to replicate the scenarios of k-wire insertions in orthopedic procedures. For evaluation, seven surgeons were provided with visual navigation and asked to perform 24 injections using the proposed framework. A second study with ten participants served to investigate the capabilities of the framework in the context of more general scenarios. Results from these studies provided comparable accuracy to those reported in the literature for AR-based navigation procedures.
Alejandro Martin-Gomez, Tianyu Song 0002, Guangzhi Wang, Hui Ding 0003, Nassir Navab, Zhe Zhao 0005, Mehran Armand
IEEE Trans. Vis. Comput. Graph.1
2023 Robotic Navigation Autonomy for Subretinal Injection via Intelligent Real-Time Virtual iOCT Volume Slicing
abstract
In the last decade, various robotic platforms have been introduced that could support delicate retinal surgeries. Concurrently, to provide semantic understanding of the surgical area, recent advances have enabled microscope-integrated intraoperative Optical Coherent Tomography (iOCT) with high-resolution 3D imaging at near video rate. The combination of robotics and semantic understanding enables task autonomy in robotic retinal surgery, such as for subretinal injection. This procedure requires precise needle insertion for best treatment outcomes. However, merging robotic systems with iOCT intro-duces new challenges. These include, but are not limited to high demands on data processing rates and dynamic registration of these systems during the procedure. In this work, we propose a framework for autonomous robotic navigation for subretinal injection, based on intelligent real-time processing of iOCT volumes. Our method consists of an instrument pose estimation method, an online registration between the robotic and the iOCT system, and trajectory planning tailored for navigation to an injection target. We also introduce intelligent virtual B-scans, a volume slicing approach for rapid instrument pose estimation, which is enabled by Convolutional Neural Networks (CNNs). Our experiments on ex-vivo porcine eyes demonstrate the precision and repeatability of the method. Finally, we discuss identified challenges in this work and suggest potential solutions to further the development of such systems.
Shervin Dehghani, Michael Sommersperger, Peiyao Zhang, Alejandro Martin-Gomez, Benjamin Busam, Peter Gehlbach, Nassir Navab, M. Ali Nasseri, Iulian Iordachita
ICRA4
2023 A Closer Look at Dynamic Medical Visualization Techniques
abstract
In navigated surgery, physicians perform complex tasks assisted by virtual representations of anatomical structures and surgical tools. Integrating Augmented Reality (AR) in these scenarios enriches the information presented to the surgeon through a range of visualization techniques. Their selection is a crucial task as they represent the primary interface between the system and the surgeon.In this work, we present a novel approach to conveying augmented content using dynamic visualization techniques, allowing users to gather depth and shape information from both pictorial and kinetic cues. We conducted user studies comparing two novel dynamic methods – Object Flow and Wave Propagation – and three state-of-the-art static visualization techniques among medical experts. Our studies provide a detailed comparison of the visualization techniques’ efficacy in conveying shape and depth information from medical data, as well as task load and usability reported by the participants and post hoc analyses. We found that kinetic cues can assist users in understanding complex anatomical structures in medical AR.
Alejandro Martin-Gomez, Felix Merkl, Alexander Winkler, Christian Heiliger, Ulrich Eck, Konrad Karcz, Nassir Navab
ISMAR1
2023 Injured Avatars: The Impact of Embodied Anatomies and Virtual Injuries on Well-Being and Performance
abstract
Human cognition relies on embodiment as a fundamental mechanism. Virtual avatars allow users to experience the adaptation, control, and perceptual illusion of alternative bodies. Although virtual bodies have medical applications in motor rehabilitation and therapeutic interventions, their potential for learning anatomy and medical communication remains underexplored. For learners and patients, anatomy, procedures, and medical imaging can be abstract and difficult to grasp. Experiencing anatomies, injuries, and treatments virtually through one's own body could be a valuable tool for fostering understanding. This work investigates the impact of avatars displaying anatomy and injuries suitable for such medical simulations. We ran a user study utilizing a skeleton avatar and virtual injuries, comparing to a healthy human avatar as a baseline. We evaluate the influence on embodiment, well-being, and presence with self-report questionnaires, as well as motor performance via an arm movement task. Our results show that while both anatomical representation and injuries increase feelings of eeriness, there are no negative effects on embodiment, well-being, presence, or motor performance. These findings suggest that virtual representations of anatomy and injuries are suitable for medical visualizations targeting learning or communication without significantly affecting users' mental state or physical control within the simulation.
Constantin Kleinbeck, Hannah Schieber, Julian Kreimeier, Alejandro Martin-Gomez, Mathias Unberath, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.4
2022 OCT-guided Robotic Subretinal Needle Injections: A Deep Learning-Based Registration Approach
abstract
Subretinal injection (SI) is an ophthalmic surgical procedure that allows for the direct injection of therapeutic substances into the subretinal space to treat vitreoretinal disorders. Although this treatment has grown in popularity, various factors contribute to its difficulty. These include the retina’s fragile, nonregenerative tissue, as well as hand tremor and poor visual depth perception. In this context, the usage of robotic devices may reduce hand tremors and facilitate gradual and controlled SI. For the robot to successfully move to the target area, it needs to understand the spatial relationship between the attached needle and the tissue. The development of optical coherence tomography (OCT) imaging has resulted in a substantial advancement in visualizing retinal structures at micron resolution. This paper introduces a novel foundation for an OCT-guided robotic steering framework that enables a surgeon to plan and select targets within the OCT volume. At the same time, the robot automatically executes the trajectories necessary to achieve the selected targets. Our contribution consists of a novel combination of existing methods, creating an intraoperative OCT-Robot registration pipeline. We combined straightforward affine transformation computations with robot kinematics and a deep neural network-determined tool-tip location in OCT. We evaluate our framework’s capability in a cadaveric pig eye open-sky procedure and using an aluminum target board. Targeting the subretinal space of the pig eye produced encouraging results with a mean Euclidean error of 23.8μm.
Kristina Mach, Shuwen Wei, Ji Woong Kim, Alejandro Martin-Gomez, Peiyao Zhang, Jin U. Kang, M. Ali Nasseri, Peter Gehlbach, Nassir Navab, Iulian Iordachita
BIBM4
2022 The Impact of Focus and Context Visualization Techniques on Depth Perception in Optical See-Through Head-Mounted Displays
abstract
Estimating the depth of virtual content has proven to be a challenging task in Augmented Reality (AR) applications. Existing studies have shown that the visual system makes use of multiple depth cues to infer the distance of objects, occlusion being one of the most important ones. The ability to generate appropriate occlusions becomes particularly important for AR applications that require the visualization of augmented objects placed below a real surface. Examples of these applications are medical scenarios in which the visualization of anatomical information needs to be observed within the patient's body. In this regard, existing works have proposed several focus and context (F+C) approaches to aid users in visualizing this content using Video See-Through (VST) Head-Mounted Displays (HMDs). However, the implementation of these approaches in Optical See-Through (OST) HMDs remains an open question due to the additive characteristics of the display technology. In this article, we, for the first time, design and conduct a user study that compares depth estimation between VST and OST HMDs using existing in-situ visualization methods. Our results show that these visualizations cannot be directly transferred to OST displays without increasing error in depth perception tasks. To tackle this gap, we perform a structured decomposition of the visual properties of AR F+C methods to find best-performing combinations. We propose the use of chromatic shadows and hatching approaches transferred from computer graphics. In a second study, we perform a factorized analysis of these combinations, showing that varying the shading type and using colored shadows can lead to better depth estimation when using OST HMDs.
Alejandro Martin-Gomez, Jakob Weiss, Andreas Keller, Ulrich Eck, Daniel Roth 0001, Nassir Navab
IEEE Trans. Vis. Comput. Graph.1
2020 Gain A New Perspective: Towards Exploring Multi-View Alignment in Mixed Reality
abstract
Manufacturing, maintenance, assembly, and training tasks represent some of the human activities that have captured special interest for Mixed Reality (MR) applications. For most of these scenarios, accurate object alignment constitutes a requirement to ensure the desired outcome. This task has proved to be especially challenging in egocentric approaches, frequently leading to estimation errors in depth. While traditional MR methods provide virtual guides such as text, arrows, or animations to assist users during alignment, this work explores the feasibility of using additional views generated by virtual cameras and mirrors. Presenting additional views from different perspectives can help to mitigate the estimation errors and show information that is not directly visible to users.To explore the benefits of using additional views for alignment tasks, and to collect reliable data and diminish external factors, we conducted a user study in a controlled virtual environment where participants aligned objects supported by additional views from a top-down camera and virtual mirrors. Data regarding alignment error, time to completion, user's interaction and attention, distance traveled, average head velocity, usability, and mental effort were collected. Our results show that using additional views reduces the mental effort and distance traveled by users and increases acceptance without negatively affecting the alignment accuracy. Therefore, we believe that users will also benefit from integrating these techniques during alignment tasks in MR environments.
Alejandro Martin-Gomez, Javad Fotouhi, Ulrich Eck, Nassir Navab
ISMAR1
2020 Augmented Mirrors
abstract
A recurrent problem in egocentric Augmented Reality (AR) applications is the misestimation of depth. Providing alternative views from non-egocentric perspectives can convey useful information for applications that require the correct judgment of depth as it is in the case of placement and alignment of virtual and real content, but also for exploration and visualization tasks.In this paper, we introduce Augmented Mirrors. Through the integration of a real mirror, our approach is capable to reflect changes of the real and virtual content of an AR application while users benefit from the perceptual advantages of using mirrors. Our concept, simple yet effective, only requires tracking the user and mirror poses with the accuracy demanded by a specific application. To showcase the potential and flexibility of the Augmented Mirrors, we present and discuss multiple examples ranging from alignment, exploration, spatial understanding, and selective content visualization using different AR-enabled devices and tracking technologies. We envision the Augmented Mirrors as a new and valuable concept that can be used in applications that benefit from additional viewpoints and require the simultaneous visualization of real and virtual content.
Alejandro Martin-Gomez, Alexander Winkler, Daniel Roth 0001, Ulrich Eck, Nassir Navab
ISMAR1
2019 Visualization Techniques for Precise Alignment in VR: A Comparative Study
abstract
Many studies explored the effectiveness of augmented, virtual, and mixed reality for object placement tasks. Two main approaches for assisting users during object alignment exist: static visualization techniques and interactive guides. This paper presents a comparative evaluation of four static visualization techniques used to render virtual objects when precise alignment in 6 degrees of freedom (DoF) is required. The selection of these techniques is based on the amount of occlusion caused by the visual guides during the alignment task. To the best of our knowledge, no previous work exists that evaluates which visualization technique is most suitable to support users while precisely aligning objects in virtual environments. We designed a virtual reality scenario considering two conditions -with and without time constraints- in which users aligned pairs of objects. To evaluate the users performance, quantitative and qualitative scores were collected. Our results suggest that visualization techniques with low levels of occlusion can improve alignment performance and increase user acceptance.
Alejandro Martin-Gomez, Ulrich Eck, Nassir Navab
VR1