VLDB 2026 Research / reviewers in the wild / expert
Mathias Unberath
dblp:165/8137
· DBLP profile ↗
54ranked-venue papers
2as first author
38since 2021 · last 2026
0000-0002-0055-9950ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Compensation Can Enhance User Engagement by Triggering Sensitivity to Financial Losses in Crowd-sourced StudiesabstractParticipation in crowd-sourced user studies is often driven by monetary incentives. However, standard payment schemes that reward completion unless responses are of poor quality may not invoke sufficient accountability. By compromising user engagement, a lack of accountability can affect data quality and the study’s ecological validity. Here, we investigate alternative compensation strategies that manipulate payment framing and evaluate their impact on engagement through task effort, outcomes, and perception. We compared a standard scheme with implicit rejection risk to a reinforced accountability condition with explicit performance-linked deductions, and two dynamic conditions that unexpectedly switched strategies. In a study with 106 Prolific participants on an image captioning task, we found that only shifting from implicit risk to reinforced accountability significantly increased engagement, likely due to loss aversion after participants had already invested time. The reverse shift decreased effort as observed in the standard group. Our results highlight the importance of carefully designing compensation schemes. Catalina Gomez, Mung Yao Jia, Sue Min Cho, Chien-Ming Huang 0001, Mathias Unberath |
CHI | 5 |
| 2026 | Causality-Driven Audits of Model RobustnessabstractRobustness audits of deep neural networks (DNN) provide a means to uncover model sensitivities to the challenging real-world imaging conditions that significantly degrade DNN performance in-the-wild. Such conditions are often the result of multiple interacting factors inherent to the environment, sensor, or processing pipeline and may lead to complex image distortions that are not easily categorized. When robustness audits are limited to a set of isolated imaging effects or distortions, the results cannot be (easily) transferred to real-world conditions where image corruptions may be more complex or nuanced. To address this challenge, we present a new alternative robustness auditing method that uses causal inference to measure DNN sensitivities to the factors of the imaging process that cause complex distortions. Our approach uses causal models to explicitly encode assumptions about the domain-relevant factors and their interactions. Then, through extensive experiments on natural and rendered images across multiple vision tasks, we show that our approach reliably estimates causal effects of each factor on DNN performance using only observational domain data. These causal effects directly tie DNN sensitivities to observable properties of the imaging pipeline in the domain of interest towards reducing the risk of unexpected DNN failures when deployed in that domain. Nathan Drenkow, William Paul, Chris Ribaudo, Mathias Unberath |
WACV | 4 |
| 2026 | Reasoning Segmentation for Images and Videos: A Survey
Yiqing Shen 0003, Chenjia Li, Jeong-O. Jeong, Tianpeng Wang, Michael Latman, Mathias Unberath |
Int. J. Comput. Vis. | 7 |
| 2026 | An interactive and explainable AI approach to improve human-machine teaming in cancer subtyping from digital cytopathology
Haomin Chen, Catalina Gomez, Zelia Correa, Tin Y. A. Liu, Tatyana Milman, Maya Eiger-Moscovich, Patricia Chévez-Barrios, Diva Salomao, Mathias Unberath |
Medical Image Anal. | 9 |
| 2026 | Benchmark of Segmentation Techniques for Pelvic Fracture in CT and X-Ray: Summary of the PENGWIN 2024 ChallengeabstractThe segmentation of pelvic fracture fragments in CT and X-ray images is crucial for trauma diagnosis, surgical planning, and intraoperative guidance. However, accurately and efficiently delineating the bone fragments remains a significant challenge due to complex anatomy and imaging limitations. The PENGWIN challenge, organized as a MICCAI 2024 satellite event, aimed to advance automated fracture segmentation by benchmarking state-of-the-art algorithms on these complex tasks. A diverse dataset of 150 CT scans was collected from multiple clinical centers, and a large set of simulated X-ray images was generated using the DeepDRR method. Final submissions from 16 teams worldwide were evaluated under a rigorous multi-metric testing scheme. The top-performing CT algorithm achieved an average fragment-wise intersection over union (IoU) of 0.930, demonstrating satisfactory accuracy. However, in the X-ray task, the best algorithm achieved an IoU of 0.774, which is promising but not yet sufficient for intra-operative decision-making, reflecting the inherent challenges of fragment overlap in projection imaging. Beyond the quantitative evaluation, the challenge revealed methodological diversity in algorithm design. Variations in instance representation, such as primary-secondary classification versus boundary-core separation, led to differing segmentation strategies. Despite promising results, the challenge also exposed inherent uncertainties in fragment definition, particularly in cases of incomplete fractures. These findings suggest that interactive segmentation approaches, integrating human decision-making with task-relevant information, may be essential for improving model reliability and clinical applicability. Yudi Sang, Yanzhen Liu, Sutuke Yibulayimu, Yunning Wang, Benjamin Killeen, Mingxu Liu, Ping-Cheng Ku, Ole Johannsen, Karol Gotkowski, Maximilian Zenk, Klaus H. Maier-Hein, Fabian Isensee, Peiyan Yue, Yi Wang 0031, Zhaohong Pan, Xiaokun Liang, Daiqi Liu, Fuxin Fan, Artur Jurgas, Andrzej Skalski, Szymon Plotka, Rafal Litka, Yingchun Song, Mathias Unberath, Mehran Armand, Dan Ruan, Shaohua Kevin Zhou, Qiyong Cao, Chunpeng Zhao, Xinbao Wu, Yu Wang 0083 |
IEEE Trans. Medical Imaging | 29 |
| 2025 | Online Reasoning Video Segmentation with Just-in-Time Digital TwinsabstractReasoning segmentation (RS) aims to identify and segment objects of interest based on implicit text queries. As such, RS is a catalyst for embodied AI agents, enabling them to interpret high-level commands without requiring explicit step-by-step guidance. However, current RS approaches rely heavily on the visual perception capabilities of multimodal large language models (LLMs), leading to several major limitations. First, they struggle with queries that require multiple steps of reasoning or those that involve complex spatial/temporal relationships. Second, they necessitate LLM fine-tuning, which may require frequent updates to maintain compatibility with contemporary LLMs and may increase risks of catastrophic forgetting during fine-tuning. Finally, being primarily designed for static images or offline video processing, they scale poorly to online video data. To address these limitations, we propose an agent framework that disentangles perception and reasoning for online video RS without LLM fine-tuning. Our innovation is the introduction of a just-in-time digital twin concept, where -- given an implicit query -- a LLM plans the construction of a low-level scene representation from high-level video using specialist vision models. We refer to this approach to creating a digital twin as "just-in-time" because the LLM planner will anticipate the need for specific information and only request this limited subset instead of always evaluating every specialist model. The LLM then performs reasoning on this digital twin representation to identify target objects. To evaluate our approach, we introduce a new comprehensive video reasoning segmentation benchmark comprising 200 videos with 895 implicit text queries. The benchmark spans three reasoning categories (semantic, spatial, and temporal) with three different reasoning chain complexity. Yiqing Shen 0003, Chenjia Li, Seenivasan Lalithkumar, Mathias Unberath |
ICCV | 5 |
| 2025 | ConceptAgent: LLM-Driven Precondition Grounding and Tree Search for Robust Task Planning and ExecutionabstractRobotic planning and execution in open-world environments is a complex problem due to the vast state spaces and high variability of task embodiment. Recent advances in perception algorithms, combined with Large Language Models (LLMs) for planning, offer promising solutions to these challenges, as the common sense reasoning capabilities of LLMs provide a strong heuristic for efficiently searching the action space. However, prior work fails to address the possibility of hallucinations from LLMs, which results in failures to execute the planned actions largely due to logical fallacies at high-or low-levels. To contend with automation failure due to such hallucinations, we introduce ConceptAgent, a natural language-driven robotic platform designed for task execution in unstructured environments. With a focus on scalability and reliability of LLM-based planning in complex state and action spaces, we present innovations designed to limit these shortcomings, including 1) Predicate Grounding to prevent and recover from infeasible actions, and 2) an embodied version of LLM-guided Monte Carlo Tree Search with self reflection. ConceptAgent combines these planning enhancements with dynamic language aligned 3d scene graphs, and large multi-modal pretrained models to perceive, localize, and interact with its environment, enabling reliable task completion. In simulation experiments, ConceptAgent achieved a 19% task completion rate across three room layouts and 30 easy level embodied tasks outperforming other state-of-the-art LLM-driven reasoning baselines that scored 10.26% and 8.11% on the same benchmark. Additionally, ablation studies on moderate to hard embodied tasks revealed a 20% increase in task completion from the baseline agent to the fully enhanced ConceptAgent, highlighting the individual and combined contributions of Predicate Grounding and LLM-guided Tree Search to enable more robust automation in complex state and action spaces. Additionally, in real-world mobile manipulation trials, conducted in randomized, low-clutter environments, a ConceptAgent-driven Spot robot achieved a 40% task completion rate, demonstrating the performance of our perception system in real-world scenarios. Corban Rivera, Grayson Byrd, William Paul, Tyler Feldman, Meghan Booker, Emma Holmes, David Handelman, Bethany Kemp, Andrew Badger, Aurora Schmidt, Krishna Murthy Jatavallabhula, Celso de Melo, Seenivasan Lalithkumar, Mathias Unberath, Rama Chellappa |
ICRA | 14 |
| 2025 | Feeling the Stakes: Realism and Ecological Validity in User Research for Computer-Assisted Interventions
Sue Min Cho, Winnie Wu, Ethan Kilmer, Russell H. Taylor, Mathias Unberath |
MICCAI (14) | 5 |
| 2025 | FluoroSAM: A Language-Promptable Foundation Model for Flexible X-Ray Image Segmentation
Benjamin Killeen, Liam J. Wang, Blanca Iñígo, Mehran Armand, Russell H. Taylor, Greg Osgood, Mathias Unberath |
MICCAI (7) | 8 |
| 2025 | Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins
Yiqing Shen 0003, Chenjia Li, Cheng-Yi Li, Tito Porras, Mathias Unberath |
MICCAI (9) | 6 |
| 2025 | Interpretable Severity Scoring of Pelvic Trauma Through Automated Fracture Detection and Bayesian InferenceabstractPelvic ring disruptions result from blunt injury mechanisms and are potentially lethal mainly due to associated injuries and massive pelvic hemorrhage. The severity of pelvic fractures in trauma victims is frequently assessed by grading the fracture according to the Tile AO/OTA classification in whole-body Computed Tomography (CT) scans. Due to the high volume of whole-body CT scans generated in trauma centers, the overall information content of a single whole-body CT scan and low manual CT reading speed, an automatic approach to Tile classification would provide substantial value, e.g., to prioritize the reading sequence of the trauma radiologists or enable them to focus on other major injuries in multi-trauma patients. In such a high-stakes scenario, an automated method for Tile grading should ideally be transparent such that the symbolic information provided by the method follows the same logic a radiologist or orthopedic surgeon would use to determine the fracture grade. This paper introduces an automated yet interpretable pelvic trauma decision support system to assist radiologists in fracture detection and Tile grading. To achieve interpretability despite processing high-dimensional whole-body CT images, we design a neurosymbolic algorithm that operates similarly to human interpretation of CT scans. The algorithm first detects relevant pelvic fractures on CTs with high specificity using Faster-RCNN. To generate robust fracture detections and associated detection (un)certainties, we perform test-time augmentation of the CT scans to apply fracture detection several times in a self-ensembling approach. The fracture detections are interpreted using a structural causal model based on clinical best practices to infer an initial Tile grade. We apply a Bayesian causal model to recover likely co-occurring fractures that may have been rejected initially due to the highly specific operating point of the detector, resulting in an updated list of detected fractures and corresponding final Tile grade. Our method is transparent in that it provides fracture location and types, as well as information on important counterfactuals that would invalidate the system's recommendation. Our approach achieves an AUC of 0.89/0.74 for translational and rotational instability,which is comparable to radiologist performance. Despite being designed for human-machine teaming, our approach does not compromise on performance compared to previous black-box methods. Haomin Chen, David Dreizin, Catalina Gomez, Anna Zapaishchykova, Mathias Unberath |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Misjudging the Machine: Gaze May Forecast Human-Machine Team Performance in Surgery
Sue Min Cho, Russell H. Taylor, Mathias Unberath |
MICCAI (6) | 3 |
| 2024 | FastSAM3D: An Efficient Segment Anything Model for 3D Volumetric Medical Images
Yiqing Shen 0003, Jingxing Li, Xinyuan Shao, Blanca Iñígo, Ankush Jindal, David Dreizin, Mathias Unberath |
MICCAI (12) | 7 |
| 2024 | RobustCLEVR: A Benchmark and Framework for Evaluating Robustness in Object-centric LearningabstractObject-centric representation learning offers the potential to overcome limitations of image-level representations by explicitly parsing image scenes into their constituent components. While image-level representations typically lack robustness to natural image corruptions, the robustness of object-centric methods remains largely untested. To address this gap, we present the RobustCLEVR benchmark dataset and evaluation framework. Our framework takes a novel approach to evaluating robustness by enabling the specification of causal dependencies in the image generation process grounded in expert knowledge and capable of producing a wide range of image corruptions unattainable in existing robustness evaluations. Using our framework, we define several causal models of the image corruption process which explicitly encode assumptions about the causal relationships and distributions of each corruption type. We generate dataset variants for each causal model on which we evaluate state-of-the-art object-centric methods. Overall, we find that object-centric methods are not inherently robust to image corruptions. Our causal evaluation approach exposes model sensitivities not observed using conventional evaluation processes, yielding greater insight into robustness differences across algorithms. Lastly, while conventional robustness evaluations view corruptions as out-of-distribution, we use our causal framework to show that even training on in-distribution image corruptions does not guarantee increased model robustness. This work provides a step towards more concrete and substantiated understanding of model performance and deterioration under complex corruption processes of the real-world.1 Nathan Drenkow, Mathias Unberath |
WACV | 2 |
| 2024 | Vessel-targeted compensation of deformable motion in interventional cone-beam CT
Alexander Lu, Heyuan Huang, Wojciech Zbijewski, Mathias Unberath, Jeffrey H. Siewerdsen, Clifford R. Weiss, Alejandro Sisniega |
Medical Image Anal. | 5 |
| 2024 | A Fully Differentiable Framework for 2D/3D Registration and the Projective Spatial TransformersabstractImage-based 2D/3D registration is a critical technique for fluoroscopic guided surgical interventions. Conventional intensity-based 2D/3D registration approa- ches suffer from a limited capture range due to the presence of local minima in hand-crafted image similarity functions. In this work, we aim to extend the 2D/3D registration capture range with a fully differentiable deep network framework that learns to approximate a convex-shape similarity function. The network uses a novel Projective Spatial Transformer (ProST) module that has unique differentiability with respect to 3D pose parameters, and is trained using an innovative double backward gradient-driven loss function. We compare the most popular learning-based pose regression methods in the literature and use the well-established CMAES intensity-based registration as a benchmark. We report registration pose error, target registration error (TRE) and success rate (SR) with a threshold of 10mm for mean TRE. For the pelvis anatomy, the median TRE of ProST followed by CMAES is 4.4mm with a SR of 65.6% in simulation, and 2.2mm with a SR of 73.2% in real data. The CMAES SRs without using ProST registration are 28.5% and 36.0% in simulation and real data, respectively. Our results suggest that the proposed ProST network learns a practical similarity function, which vastly extends the capture range of conventional intensity-based 2D/3D registration. We believe that the unique differentiable property of ProST has the potential to benefit related 3D medical imaging research applications. The source code is available at https://github.com/gaocong13/Projective-Spatial-Transformers. Cong Gao 0003, Anqi Feng, Xingtong Liu, Russell H. Taylor, Mehran Armand, Mathias Unberath |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Neuralangelo: High-Fidelity Neural Surface ReconstructionabstractNeural surface reconstruction has been shown to be powerful for recovering dense 3D surfaces via image-based neural rendering. However, current methods struggle to recover detailed structures of real-world scenes. To address the issue, we present Neuralangelo, which combines the representation power of multiresolution 3D hash grids with neural surface rendering. Two key ingredients enable our approach: (1) numerical gradients for computing higher-order derivatives as a smoothing operation and (2) coarse-to-fine optimization on the hash grids controlling different levels of details. Even without auxiliary inputs such as depth, Neuralangelo can effectively recover dense 3D surface structures from multiview images with fidelity significantly surpassing previous methods, enabling detailed large-scale scene reconstruction from RGB video captures. Zhaoshuo Li, Thomas Müller 0013, Alex Evans, Russell H. Taylor, Mathias Unberath, Ming-Yu Liu 0001, Chen-Hsuan Lin 0001 |
CVPR | 5 |
| 2023 | TransNuSeg: A Lightweight Multi-task Transformer for Nuclei Segmentation
Zhenqi He, Mathias Unberath, Jing Ke, Yiqing Shen 0003 |
MICCAI (4) | 2 |
| 2023 | Pelphix: Surgical Phase Recognition from X-Ray Images in Percutaneous Pelvic Fixation
Benjamin Killeen, Jan Mangulabnan, Mehran Armand, Russell H. Taylor, Greg Osgood, Mathias Unberath |
MICCAI (9) | 7 |
| 2023 | Data AUDIT: Identifying Attribute Utility- and Detectability-Induced Bias in Task Models
Mitchell Pavlak, Nathan Drenkow, Nicholas Petrick, Mohammad Mehdi Farhangi, Mathias Unberath |
MICCAI (3) | 5 |
| 2023 | Temporally Consistent Online Depth Estimation in Dynamic ScenesabstractTemporally consistent depth estimation is crucial for online applications such as augmented reality. While stereo depth estimation has received substantial attention as a promising way to generate 3D information, there is relatively little work focused on maintaining temporal stability. Indeed, based on our analysis, current techniques still suffer from poor temporal consistency. Stabilizing depth temporally in dynamic scenes is challenging due to concurrent object and camera motion. In an online setting, this process is further aggravated because only past frames are available. We present a framework named Consistent Online Dynamic Depth (CODD) to produce temporally consistent depth estimates in dynamic scenes in an online setting. CODD augments per-frame stereo networks with novel motion and fusion networks. The motion network accounts for dynamics by predicting a per-pixel SE3 transformation and aligning the observations. The fusion network improves temporal depth consistency by aggregating the current and past estimates. We conduct extensive experiments and demonstrate quantitatively and qualitatively that CODD outperforms competing methods in terms of temporal consistency and performs on par in terms of per-frame accuracy. Zhaoshuo Li, Dilin Wang, Francis X. Creighton, Russell H. Taylor, Ganesh Venkatesh, Mathias Unberath |
WACV | 7 |
| 2023 | Mapping DNN Embedding Manifolds for Network Generalization PredictionabstractDeep Neural Networks(DNN) often fail in surprising ways, and predicting how well a trained DNN will generalize in a new, external operating domain is essential for deploying DNNs in safety critical applications, e.g., perception for self-driving vehicles or medical image analysis. Recently, the task of Network Generalization Prediction (NGP) has been proposed to predict how a DNN will generalize in an external operating domain. Previous NGP approaches have leveraged multiple labeled test sets or labeled metadata. In this study, we propose an embedding map, the first NGP approach that predicts DNN performance based on how unlabeled images from an external operating domain map in the DNN embedding space. We evaluate our proposed Embedding Map and other recently proposed NGP approaches for pedestrian, melanoma, and animal classification tasks. We find that our embedding map has the best average NGP performance, and that our embedding map is effective at modeling complex, non-linear embedding space structures. Molly O'Brien, Brett Wolfinger, Julia V. Bukowski, Mathias Unberath, Aria Pezeshk, Gregory D. Hager |
WACV | 4 |
| 2023 | Mitigating knowledge imbalance in AI-advised decision-making through collaborative user involvementabstractIntegrating artificial intelligence (AI) systems into decision-making tasks attempts to assist people by augmenting or complementing their abilities and ultimately improve task performance. However, when considering recommendations from modern “black box” intelligent systems, users are confronted with the decision of accepting or overriding AI’s recommendations. These decisions are even more challenging to make when there exists a significant knowledge imbalance between the users and the AI system—namely, when people lack necessary task knowledge and are therefore unable to accurately complete the task on their own. In this work, we aim to understand people’s behavior in AI-assisted decision-making tasks when faced with the challenge of knowledge imbalance and explore whether involving users in an AI’s prediction generation process makes them more willing to follow the AI’s recommendations and enhances their perception of collaboration. Our empirical study reveals that the involvement of users in generating AI recommendations during a task with notable knowledge imbalance causes them to be more willing to agree with the AI’s suggestions and to perceive the AI agent and their collaboration as a team more positively. Catalina Gomez, Mathias Unberath, Chien-Ming Huang 0001 |
Int. J. Hum. Comput. Stud. | 2 |
| 2023 | Injured Avatars: The Impact of Embodied Anatomies and Virtual Injuries on Well-Being and PerformanceabstractHuman cognition relies on embodiment as a fundamental mechanism. Virtual avatars allow users to experience the adaptation, control, and perceptual illusion of alternative bodies. Although virtual bodies have medical applications in motor rehabilitation and therapeutic interventions, their potential for learning anatomy and medical communication remains underexplored. For learners and patients, anatomy, procedures, and medical imaging can be abstract and difficult to grasp. Experiencing anatomies, injuries, and treatments virtually through one's own body could be a valuable tool for fostering understanding. This work investigates the impact of avatars displaying anatomy and injuries suitable for such medical simulations. We ran a user study utilizing a skeleton avatar and virtual injuries, comparing to a healthy human avatar as a baseline. We evaluate the influence on embodiment, well-being, and presence with self-report questionnaires, as well as motor performance via an arm movement task. Our results show that while both anatomical representation and injuries increase feelings of eeriness, there are no negative effects on embodiment, well-being, presence, or motor performance. These findings suggest that virtual representations of anatomy and injuries are suitable for medical visualizations targeting learning or communication without significantly affecting users' mental state or physical control within the simulation. Constantin Kleinbeck, Hannah Schieber, Julian Kreimeier, Alejandro Martin-Gomez, Mathias Unberath, Daniel Roth 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002 |
ECCV (32) | 6 |
| 2022 | SAGE: SLAM with Appearance and Geometry Prior for Endoscopyabstract., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry of the observed anatomy from a monocular endoscopic video. To this end, we develop a Simultaneous Localization and Mapping system by combining the learning-based appearance and optimizable geometry priors and factor graph optimization. The appearance and geometry priors are explicitly learned in an end-to-end differentiable training pipeline to master the task of pair-wise image alignment, one of the core components of the SLAM system. In our experiments, the proposed SLAM system is shown to robustly handle the challenges of texture scarceness and illumination variation that are commonly seen in endoscopy. The system generalizes well to unseen endoscopes and subjects and performs favorably compared with a state-of-the-art feature-based SLAM system. The code repository is available at https://github.com/lppllppl920/SAGE-SLAM.git. Xingtong Liu, Zhaoshuo Li, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
ICRA | 6 |
| 2022 | CaRTS: Causality-Driven Robot Tool Segmentation from Vision and Kinematics Data
Hao Ding 0021, Jintan Zhang, Peter Kazanzides, Jie Ying Wu, Mathias Unberath |
MICCAI (8) | 5 |
| 2022 | AR-Loupe: Magnified Augmented Reality by Combining an Optical See-Through Head-Mounted Display and a LoupeabstractHead-mounted loupes can increase the user's visual acuity to observe the details of an object. On the other hand, optical see-through head-mounted displays (OST-HMD) are able to provide virtual augmentations registered with real objects. In this article, we propose AR-Loupe, combining the advantages of loupes and OST-HMDs, to offer augmented reality in the user's magnified field-of-vision. Specifically, AR-Loupe integrates a commercial OST-HMD, Magic Leap One, and binocular Galilean magnifying loupes, with customized 3D-printed attachments. We model the combination of user's eye, screen of OST-HMD, and the optical loupe as a pinhole camera. The calibration of AR-Loupe involves interactive view segmentation and an adapted version of stereo single point active alignment method (Stereo-SPAAM). We conducted a two-phase multi-user study to evaluate AR-Loupe. The users were able to achieve sub-millimeter accuracy ( 0.82 mm) on average, which is significantly ( ) smaller compared to normal AR guidance ( 1.49 mm). The mean calibration time was 268.46 s. With the increased size of real objects through optical magnification and the registered augmentation, AR-Loupe can aid users in high-precision tasks with better visual acuity and higher accuracy. Tianyu Song 0002, Mathias Unberath, Peter Kazanzides |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Feasibility of a Cannula-Mounted Piezo Robot for Image-Guided Vertebral Augmentation: Toward a Low Cost, Semi-Autonomous ApproachabstractVertebral compression fractures (VCFs), the most common fragility fractures secondary to osteoporosis, affect more than 200 million individuals worldwide. Percutaneous vertebral augmentation is an effective interventional treatment option that is routinely performed across the world. Because fluoroscopy-guided vertebral augmentation is a well-established and safe minimally invasive technique, automating its delivery is among the most important next steps. In this work, we describe the design and evaluation of a novel cannula mounted vertebral augmentation robot in a simulated X-ray environment as a first step toward autonomous vertebral augmentation. The cannula robot employs a piezo stack with inchworm control to place surgical tools within the vertebral body, while X-ray imaging verifies the robot does not interfere with imaging. Finite element analysis of the robot confirms that radiolucent materials were rigid enough to be used in the robot design as expected deformations for the cannula drive, accessory drive, and locking mechanisms$(1.299 \pm 0.034 \ um, 1.280 \pm 0.027\ um$, and$1.960 \pm 0.218\ um$, respectively) did not exceed the stroke lengths of the piezo stacks. An in silico clinical trial based on a human anatomy model suffering from VCF validates that the cannula robot does not impede visualization of the critical anatomy and tool-to-tissue positioning. Together these results demonstrate the feasibility of a cannula mounted robot for vertebral augmentation. Justin D. Opfermann, Benjamin Killeen, Christopher R. Bailey, Ali Uneri, Kensei Suzuki, Mehran Armand, Ferdinand Hui, Axel Krieger, Mathias Unberath |
BIBE | 10 |
| 2021 | Neighborhood Normalization for Robust Geometric Feature LearningabstractExtracting geometric features from 3D models is a common first step in applications such as 3D registration, tracking, and scene flow estimation. Many hand-crafted and learning-based methods aim to produce consistent and distinguishable geometric features for 3D models with partial overlap. These methods work well in cases where the point density and scale of the overlapping 3D objects are similar, but struggle in applications where 3D data are obtained independently with unknown global scale and scene overlap. Unfortunately, instances of this resolution mismatch are common in practice, e.g., when aligning data from multiple sensors. In this work, we introduce a new normalization technique, Batch-Neighborhood Normalization, aiming to improve robustness to mean-std variation of local feature distributions that presumably can happen in samples with varying point density. We empirically demonstrate that the presented normalization method’s performance compares favorably to comparison methods in indoor and outdoor environments, and on a clinical dataset, on common point registration benchmarks in both standard and, particularly, resolution-mismatch settings. The source code and clinical dataset are available at https://github.com/lppllppl920/NeighborhoodNormalization-Pytorch. Xingtong Liu, Benjamin Killeen, Ayushi Sinha, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
CVPR | 7 |
| 2021 | Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective with TransformersabstractStereo depth estimation relies on optimal correspondence matching between pixels on epipolar lines in the left and right images to infer depth. In this work, we revisit the problem from a sequence-to-sequence correspondence perspective to replace cost volume construction with dense pixel matching using position information and attention. This approach, named STereo TRansformer (STTR), has several advantages: It 1) relaxes the limitation of a fixed disparity range, 2) identifies occluded regions and provides confidence estimates, and 3) imposes uniqueness constraints during the matching process. We report promising results on both synthetic and real-world datasets and demonstrate that STTR generalizes across different domains, even without fine-tuning. Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy S. Ding, Francis X. Creighton, Russell H. Taylor, Mathias Unberath |
ICCV | 7 |
| 2021 | Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic SurgeryabstractAutomatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html. Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001 |
ICRA | 5 |
| 2021 | E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-Based Stereoscopic Depth Perception
Yonghao Long 0001, Zhaoshuo Li, Chi Hang Yee, Chi-Fai Ng, Russell H. Taylor, Mathias Unberath, Qi Dou 0001 |
MICCAI (4) | 6 |
| 2021 | An Interpretable Approach to Automated Severity Scoring in Pelvic Trauma
Anna Zapaishchykova, David Dreizin, Zhaoshuo Li, Jie Ying Wu, Shahrooz Faghih Roohi, Mathias Unberath |
MICCAI (3) | 6 |
| 2021 | Exploring partial intrinsic and extrinsic symmetry in 3D medical imaging
Javad Fotouhi, Giacomo Taylor, Mathias Unberath, Alex Johnson, Sing Chun Lee, Greg Osgood, Mehran Armand, Nassir Navab |
Medical Image Anal. | 3 |
| 2021 | Reconstruction of Orthographic Mosaics From Perspective X-Ray ImagesabstractImage stitching is a prominent challenge in medical imaging, where the limited field-of-view captured by single images prohibits holistic analysis of patient anatomy. The barrier that prevents straight-forward mosaicing of 2D images is depth mismatch due to parallax. In this work, we leverage the Fourier slice theorem to aggregate information from multiple transmission images in parallax-free domains using fundamental principles of X-ray image formation. The details of the stitched image are subsequently restored using a novel deep learning strategy that exploits similarity measures designed around frequency, as well as dense and sparse spatial image content. Our work provides evidence that reconstruction of orthographic mosaics is possible with realistic motions of the C-arm involving both translation and rotation. We also show that these orthographic mosaics enable metric measurements of clinically relevant quantities directly on the 2D image plane. Javad Fotouhi, Xingtong Liu, Mehran Armand, Nassir Navab, Mathias Unberath |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Development and Pre-Clinical Analysis of Spatiotemporal-Aware Augmented Reality in Orthopedic InterventionsabstractSuboptimal interaction with patient data and challenges in mastering 3D anatomy based on ill-posed 2D interventional images are essential concerns in image-guided therapies. Augmented reality (AR) has been introduced in the operating rooms in the last decade; however, in image-guided interventions, it has often only been considered as a visualization device improving traditional workflows. As a consequence, the technology is gaining minimum maturity that it requires to redefine new procedures, user interfaces, and interactions. The main contribution of this paper is to reveal how exemplary workflows are redefined by taking full advantage of head-mounted displays when entirely co-registered with the imaging system at all times. The awareness of the system from the geometric and physical characteristics of X-ray imaging allows the exploration of different human-machine interfaces. Our system achieved an error of 4.76 ± 2.91mm for placing K-wire in a fracture management procedure, and yielded errors of 1.57 ± 1.16° and 1.46 ± 1.00° in the abduction and anteversion angles, respectively, for total hip arthroplasty (THA). We compared the results with the outcomes from baseline standard operative and non-immersive AR procedures, which had yielded errors of [4.61mm, 4.76°, 4.77°] and [5.13mm, 1.78°, 1.43°], respectively, for wire placement, and abduction and anteversion during THA. We hope that our holistic approach towards improving the interface of surgery not only augments the surgeon's capabilities but also augments the surgical team's experience in carrying out an effective intervention with reduced complications and provide novel approaches of documenting procedures for training purposes. Javad Fotouhi, Arian Mehrfard, Tianyu Song 0002, Alex Johnson, Greg Osgood, Mathias Unberath, Mehran Armand, Nassir Navab |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Pose-Dependent Weights and Domain Randomization for Fully Automatic X-Ray to CT RegistrationabstractFully automatic X-ray to CT registration requires a solid initialization to provide an initial alignment within the capture range of existing intensity-based registrations. This work addresses that need by providing a novel automatic initialization, which enables end to end registration. First, a neural network is trained once to detect a set of anatomical landmarks on simulated X-rays. A domain randomization scheme is proposed to enable the network to overcome the challenge of being trained purely on simulated data and run inference on real X-rays. Then, for each patient CT, a fully-automatic patient-specific landmark extraction scheme is used. It is based on backprojecting and clustering the previously trained network's predictions on a set of simulated X-rays. Next, the network is retrained to detect the new landmarks. Finally the combination of network and 3D landmark locations is used to compute the initialization using a perspective-n-point algorithm. During the computation of the pose, a weighting scheme is introduced to incorporate the confidence of the network in detecting the landmarks. The algorithm is evaluated on the pelvis using both real and simulated x-rays. The mean (± standard deviation) target registration error in millimetres is 4.1 ± 4.3 for simulated X-rays with a success rate of 92% and 4.2 ± 3.9 for real X-rays with a success rate of 86.8%, where a success is defined as a translation error of less than 30 mm . Matthias Grimm, Javier Esteban, Mathias Unberath, Nassir Navab |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Extremely Dense Point Correspondences Using a Learned Feature DescriptorabstractHigh-quality 3D reconstructions from endoscopy video play an important role in many clinical applications, including surgical navigation where they enable direct video-CT registration. While many methods exist for general multi-view 3D reconstruction, these methods often fail to deliver satisfactory performance on endoscopic video. Part of the reason is that local descriptors that establish pair-wise point correspondences, and thus drive reconstruction, struggle when confronted with the texture-scarce surface of anatomy. Learning-based dense descriptors usually have larger receptive fields enabling the encoding of global information, which can be used to disambiguate matches. In this work, we present an effective self-supervised training scheme and novel loss design for dense descriptor learning. In direct comparison to recent local and dense descriptors on an in-house sinus endoscopy dataset, we demonstrate that our proposed dense descriptor can generalize to unseen patients and scopes, thereby largely improving the performance of Structure from Motion (SfM) in terms of model density and completeness. We also evaluate our method on a public dense optical flow dataset and a small-scale SfM public dataset to further demonstrate the effectiveness and generality of our method. The source code is available at https://github.com/lppllppl920/DenseDescriptorLearning-Pytorch. Xingtong Liu, Yiping Zheng, Benjamin Killeen, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
CVPR | 7 |
| 2020 | Generalizing Spatial Transformers to Projective Geometry with Applications to 2D/3D Registration
Cong Gao 0003, Xingtong Liu, Wenhao Gu, Benjamin Killeen, Mehran Armand, Russell H. Taylor, Mathias Unberath |
MICCAI (3) | 7 |
| 2020 | Reconstructing Sinus Anatomy from Endoscopic Video - Towards a Radiation-Free Approach for Quantitative Longitudinal Assessment
Xingtong Liu, Maia Stiber, Jindan Huang, Masaru Ishii, Gregory D. Hager, Russell H. Taylor, Mathias Unberath |
MICCAI (3) | 7 |
| 2020 | CAI4CAI: The Rise of Contextual Artificial Intelligence in Computer-Assisted InterventionsabstractData-driven computational approaches have evolved to enable extraction of information from medical images with a reliability, accuracy and speed which is already transforming their interpretation and exploitation in clinical practice. While similar benefits are longed for in the field of interventional imaging, this ambition is challenged by a much higher heterogeneity. Clinical workflows within interventional suites and operating theatres are extremely complex and typically rely on poorly integrated intra-operative devices, sensors, and support infrastructures. Taking stock of some of the most exciting developments in machine learning and artificial intelligence for computer assisted interventions, we highlight the crucial need to take context and human factors into account in order to address these challenges. Contextual artificial intelligence for computer assisted intervention, or CAI4CAI, arises as an emerging opportunity feeding into the broader field of surgical data science. Central challenges being addressed in CAI4CAI include how to integrate the ensemble of prior knowledge and instantaneous sensory information from experts, sensors and actuators; how to create and communicate a faithful and actionable shared representation of the surgery among a mixed human-AI actor team; how to design interventional systems and associated cognitive shared control schemes for online uncertainty-aware collaborative decision making ultimately producing more precise and reliable interventions. Tom Vercauteren, Mathias Unberath, Nicolas Padoy, Nassir Navab |
Proc. IEEE | 2 |
| 2020 | Dense Depth Estimation in Monocular Endoscopy With Self-Supervised Learning MethodsabstractWe present a self-supervised approach to training convolutional neural networks for dense depth estimation from monocular endoscopy data without a priori modeling of anatomy or shading. Our method only requires monocular endoscopic videos and a multi-view stereo method, e.g., structure from motion, to supervise learning in a sparse manner. Consequently, our method requires neither manual labeling nor patient computed tomography (CT) scan in the training and application phases. In a cross-patient experiment using CT scans as groundtruth, the proposed method achieved submillimeter mean residual error. In a comparison study to recent self-supervised depth estimation methods designed for natural video on in vivo sinus endoscopy data, we demonstrate that the proposed approach outperforms the previous methods by a large margin. The source code for this work is publicly available online at https://github.com/lppllppl920/EndoscopyDepthEstimation-Pytorch. Xingtong Liu, Ayushi Sinha, Masaru Ishii, Gregory D. Hager, Austin Reiter, Russell H. Taylor, Mathias Unberath |
IEEE Trans. Medical Imaging | 7 |
| 2019 | Towards Fully Automatic X-Ray to CT Registration
Javier Esteban, Matthias Grimm, Mathias Unberath, Guillaume Zahnd, Nassir Navab |
MICCAI (6) | 3 |
| 2019 | LumiPath - Towards Real-Time Physically-Based Rendering on Embedded Devices
Laura Fink, Sing Chun Lee, Jie Ying Wu, Xingtong Liu, Tianyu Song 0002, Yordanka Velikova, Marc Stamminger, Nassir Navab, Mathias Unberath |
MICCAI (5) | 9 |
| 2019 | Learning to Avoid Poor Images: Towards Task-aware C-arm Cone-beam CT Trajectories
Jan-Nico Zaech, Cong Gao 0003, Bastian Bier, Russell H. Taylor, Andreas K. Maier, Nassir Navab, Mathias Unberath |
MICCAI (5) | 7 |
| 2018 | X-ray-transform Invariant Anatomical Landmark Detection for Pelvic Trauma Surgery
Bastian Bier, Mathias Unberath, Jan-Nico Zaech, Javad Fotouhi, Mehran Armand, Greg Osgood, Nassir Navab, Andreas K. Maier |
MICCAI (4) | 2 |
| 2018 | Exploiting Partial Structural Symmetry for Patient-Specific Image Augmentation in Trauma Interventions
Javad Fotouhi, Mathias Unberath, Giacomo Taylor, Arash Ghaani Farashahi, Bastian Bier, Russell H. Taylor, Greg Osgood, Mehran Armand, Nassir Navab |
MICCAI (4) | 2 |
| 2018 | Closing the Calibration Loop: An Inside-Out-Tracking Paradigm for Augmented Reality in Orthopedic Surgery
Jonas Hajek, Mathias Unberath, Javad Fotouhi, Bastian Bier, Sing Chun Lee, Greg Osgood, Andreas K. Maier, Mehran Armand, Nassir Navab |
MICCAI (4) | 2 |
| 2018 | Double Your Views - Exploiting Symmetry in Transmission Imaging
Alexander Preuhs, Andreas K. Maier, Michael Manhart 0001, Javad Fotouhi, Nassir Navab, Mathias Unberath |
MICCAI (1) | 6 |
| 2018 | DeepDRR - A Catalyst for Machine Learning in Fluoroscopy-Guided Procedures
Mathias Unberath, Jan-Nico Zaech, Sing Chun Lee, Bastian Bier, Javad Fotouhi, Mehran Armand, Nassir Navab |
MICCAI (4) | 1 |
| 2018 | Prior-Free Respiratory Motion Estimation in Rotational AngiographyabstractRotational coronary angiography using C-arm angiography systems enables intra-procedural 3-D imaging that is considered beneficial for diagnostic assessment and interventional guidance. Despite previous efforts, rotational angiography was not yet successfully established in clinical practice for coronary artery procedures due to challenges associated with substantial intra-scan respiratory and cardiac motion. While gating handles cardiac motion during reconstruction, respiratory motion requires compensation. State-of-the-art algorithms rely on 3-D / 2-D registration that requires an uncompensated reconstruction of sufficient quality. To overcome this limitation, we investigate two prior-free respiratory motion estimation methods based on the optimization of: 1) epipolar consistency conditions (ECCs) and 2) a task-based auto-focus measure (AFM). The methods assess redundancies in projection images or impose favorable properties of 3-D space, respectively, and are used to estimate the respiratory motion of the coronary arteries within rotational angiograms. We evaluate our algorithms on the publicly available CAVAREV benchmark and on clinical data. We quantify reductions in error due to respiratory motion compensation using a dedicated reconstruction domain metric. Moreover, we study the improvements in image quality when using an analytic and a novel temporal total variation regularized algebraic reconstruction algorithm. We observed substantial improvement in all figures of merit compared with the uncompensated case. Improvements in image quality presented as a reduction of double edges, blurring, and noise. Benefits of the proposed corrections were notable even in cases suffering little corruption from respiratory motion, translating to an improvement in the vessel sharpness of (6.08 ± 4.46)% and (14.7 ± 8.80)% when the ECC-based and the AFM-based compensation were applied. On the CAVAREV data, our motion compensation approach exhibits an improvement of (27.6 ± 7.5)% and (97.0 ± 17.7)% when the ECC and AFM were used, respectively. At the time of writing, our method based on AFM is leading the CAVAREV scoreboard. Both motion estimation strategies are purely image-based and accurately estimate the displacements of the coronary arteries due to respiration. While current evidence suggests the superior performance of AFM, future work will further investigate the use of ECC in the context of angiography as they solely rely on geometric calibration and projection-domain images. Mathias Unberath, Oliver Taubmann, André Aichert, Stephan Achenbach, Andreas K. Maier |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Deep Learning Computed Tomography: Learning Projection-Domain Weights From Image Domain in Limited Angle ProblemsabstractIn this paper, we present a new deep learning framework for 3-D tomographic reconstruction. To this end, we map filtered back-projection-type algorithms to neural networks. However, the back-projection cannot be implemented as a fully connected layer due to its memory requirements. To overcome this problem, we propose a new type of cone-beam back-projection layer, efficiently calculating the forward pass. We derive this layer's backward pass as a projection operation. Unlike most deep learning approaches for reconstruction, our new layer permits joint optimization of correction steps in volume and projection domain. Evaluation is performed numerically on a public data set in a limited angle setting showing a consistent improvement over analytical algorithms while keeping the same computational test-time complexity by design. In the region of interest, the peak signal-to-noise ratio has increased by 23%. In addition, we show that the learned algorithm can be interpreted using known concepts from cone beam reconstruction: the network is able to automatically learn strategies such as compensation weights and apodization windows. Tobias Würfl, Mathis Hoffmann, Vincent Christlein, Katharina Breininger, Yixing Huang, Mathias Unberath, Andreas K. Maier |
IEEE Trans. Medical Imaging | 6 |
| 2017 | Handling multiple materials for exposure of digital forgeries using 2-D lighting environments
Christian Riess, Mathias Unberath, Farzad Naderi, Sven Pfaller, Marc Stamminger, Elli Angelopoulou |
Multim. Tools Appl. | 2 |