Daniel Roth 0001

dblp:63/6882-1 · DBLP profile ↗
← Back
48ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0001-5175-1566ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 8 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 20 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hybrid Foveated Path Tracing with Peripheral Gaussians for Immersive Anatomy
abstract
Volumetric medical imaging offers great potential for understanding complex pathologies. Yet, traditional 2D slices provide little support for interpreting spatial relationships, forcing users to mentally reconstruct anatomy into three dimensions. Direct volumetric path tracing and VR rendering can improve perception but are computationally expensive, while precomputed representations, like Gaussian Splatting, require planning ahead. Both approaches limit interactive use.We propose a hybrid rendering approach for high-quality, interactive, and immersive anatomical visualization. Our method combines streamed foveated path tracing with a lightweight Gaussian Splatting approximation of the periphery. The peripheral model generation is optimized with volume data and continuously refined using foveal renderings, enabling interactive updates. Depth-guided reprojection further improves robustness to latency and allows users to balance fidelity with refresh rate. We compare our method against direct path tracing and Gaussian Splatting. Our results highlight how their combination can preserve strengths in visual quality while re-generating the peripheral model in under a second, eliminating extensive preprocessing and approximations. This opens new options for interactive medical visualization.
Constantin Kleinbeck, Luisa Theelke, Hannah Schieber, Ulrich Eck, Rüdiger von Eisenhart-Rothe, Daniel Roth 0001
VR6
2026 Grand Challenges in Cross Reality
abstract
Cross Reality (CR) is a new emerging field based on the current developments in Mixed Reality hardware, especially supported by the broad market penetration of video-based see-through Head-Mounted Displays. It refers to applications that span across different stages (real, Augmented Reality, Augmented Virtuality, Virtual Reality) of the reality-virtuality continuum, where users are interconnected between different stages and/or are able to transition between these stages. This publication follows the concept of other grand challenges publications and reflects the discussion of various researchers invested in CR. After an initial discussion at the 1stJoint Workshop on Cross Reality at IEEE ISMAR 2023, six topic groups have been identified, leading to 22 challenges, which were discussed in groups over the period of multiple months. The discussion of these challenges should act as a road map for future research in the area of CR.
Christoph Anthes, Mark Billinghurst, Uwe Gruenefeld, Hans-Christian Jetter, Hai-Ning Liang, Frank Maurer, David Aigner, Craig Anslow, Guillaume Bataille, Abraham G. Campbell, Judith Friedl-Knirsch, Alexander Gall, Renan Luigi Martins Guarese, Sebastian Hubenschmid, Yue Li 0023, Fabian Pointecker, Andreas Riegler, Daniel Roth 0001, Rishi Vanukuru, Nanjia Wang, Lingyun Yu 0001, Johannes Zagermann, Daniel Zielasko
IEEE Trans. Vis. Comput. Graph.18
2026 MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects
abstract
Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or objects, across multiple views. While commonly cameras are calibrated in an initial step or updated through the constant use of markers, another option is to leverage information already present in the scene, like known objects. Another downside of marker-based tracking is that markers have to be tracked inside the field-of-view (FoV) of the cameras. To overcome these limitations, we propose a constant dynamic camera pose estimation leveraging spatiotemporal FoV overlaps of known objects on the fly. To achieve that, we enhance the state-of-the-art object pose estimator to update our spatiotemporal scene graph, enabling a relation even among non-overlapping FoV cameras. To evaluate our approach, we introduce a multi-camera, multi-object pose estimation dataset with temporal FoV overlap, including static and dynamic cameras. Furthermore, in FoV overlapping scenarios, we outperform the state-of-the-art on the widely used YCB-V and T-LESS dataset in camera pose accuracy. Our performance on both previous and our proposed datasets validates the effectiveness of our marker-less approach for AR applications. The code and dataset are available on https://github.com/roth-hex-lab/IEEE-VR-2026-MultiCam.
Shiyu Li 0003, Hannah Schieber, Kristoffer Waldow, Benjamin Busam, Julian Kreimeier, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.6
2026 The Influence of Environmental Fidelity on Virtual Presence, Intrinsic Motivation, Cognitive Load and Learning Outcomes in Medical VR
abstract
Immersive virtual reality learning environments (IVRLEs) are increasingly used in medical education, yet the role of environmental fidelity-particularly scene design-remains underexplored. This study examines how varying levels of fidelity and contextualization affect motivational and cognitive outcomes. Eighty-seven medical students were randomly assigned to one of three scene conditions: a minimalistic "Blank Scene," a "Reconstructed Classroom Scene", or an "Inside-Human Scene". All students used a custom-developed application to learn about embryonic heart development. We measured virtual presence, intrinsic motivation, cognitive load, learning outcomes, and usability. Results showed that scene design influenced virtual presence, selected aspects of intrinsic motivation, cognitive load, and learning outcomes. The Inside-Human Scene elicited higher physical and self-presence as well as higher comprehension scores compared to the Reconstructed Classroom Scene. The Reconstructed Classroom Scene was associated with higher extraneous cognitive load. Intrinsic cognitive load was rated higher in the Inside-Human Scene, while germane cognitive load did not differ between conditions. No significant differences were found for task performance or factual recall. Overall, the findings indicate that scene design in IVRLEs affects how learners engage with complex content and may support deeper understanding when perceptual and contextual properties are coherent, while visually detailed environments may increase extraneous cognitive demands without improving learning.
Danny Schott, Matthias Kunz, Claudia Schrader, Elias Ringler, Alexander Schwadtke, Jonas Mandel, Constantin Kleinbeck, Daniel Roth 0001, Anne Albrecht, Rüdiger Braun-Dullaeus, Christian Hansen 0001
IEEE Trans. Vis. Comput. Graph.9
2026 Evaluating Cutout Rendering Techniques for Pass-Through Embodiment Using a Real-Mirror Metaphor
abstract
A convincing sense of embodiment in virtual reality (VR) is crucial for creating immersive and engaging experiences, as it shapes how users perceive and interact with their virtual bodies. The sense of embodiment is thereby, among others, affected by the shape, appearance, and fidelity of the virtual body. However, achieving convincing avatar appearance remains a challenge for VR applications. One promising solution is Pass-Through Embodiment (PTE), which enables users to see their real bodies in VR. PTE combines depth-based segmentation with the pass-through video stream of video-see-through displays to effectively visualize photon-captured representations of their own bodies. Despite the source-fidelity of the representation, the resolution of integrated depth sensors in Head Mounted Displays (HMD) can produce artifacts at segmentation boundaries, leading to visible aliasing. The perceptual impact of these artifacts on the VR experience remains unexplored. Therefore, in this paper we compare three edge-rendering techniques designed to reduce artifacts without compromising performance. Aside from a soft gradient, we introduce two new methods with a hard and dithered edge. The latter aims to balance the sharpness of hard masks with the smoothness of gradient transitions, without relying on alpha blending. To evaluate those methods in a PTE context, we conducted a within-subjects study that introduces a novel real-mirror paradigm, using an actual physical mirror as reference for reflection. We found significant results in measured presence and embodiment. Subsequent analysis revealed that our dithered cutout approach significantly outperforms hard masks, while no significant difference was found between soft condition. These results suggest a perceptual continuum where dithering and soft blending both effectively reduce visual artifacts through gradient representation. Together with high overall ratings on presence and embodiment across all conditions, these findings confirm PTE as a robust method for supporting embodiment and presence, while highlighting the potential of dithering as a computationally efficient yet perceptually comparable alternative to smooth blending.
Kristoffer Waldow, Arnulph Fuhrmann, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.3
2025 Sonify Anything: Towards Context-Aware Sonic Interactions in AR
abstract
In Augmented Reality (AR), virtual objects interact with real objects. However, the lack of physicality of virtual objects leads to the absence of natural sonic interactions. When virtual and real objects collide, either no sound or a generic sound is played. Both lead to an incongruent multisensory experience reducing interaction and object realism. Unlike in Virtual Reality (VR) and games, where predefined scenes and interactions allow for the playback of prerecorded sound samples, AR requires real-time sound synthesis that dynamically adapts to novel contexts and objects to provide audiovisual congruence during interaction. To enhance real-virtual object interactions in AR, we propose a framework for context-aware sounds using methods from computer vision to recognize and segment the materials of real objects. The material's physical properties and the impact dynamics of the interaction are used to generate material-based sounds in real-time using physical modelling synthesis. In a user study with 24 participants, we compared our congruent material-based sounds to a generic sound effect, mirroring the current standard of non-context-aware sounds in AR applications. The results showed that material-based sounds led to significantly more realistic sonic interactions. Material-based sounds also enabled participants to distinguish visually similar materials with significantly greater accuracy and confidence. These findings show that context-aware, material-based sonic interactions in AR foster a stronger sense of realism and enhance our perception of real-world surroundings.
Laura Schütz, Sasan Matinfar, Ulrich Eck, Daniel Roth 0001, Nassir Navab
ISMAR4
2025 Semantics-Controlled Gaussian Splatting for Outdoor Scene Reconstruction and Rendering in Virtual Reality
abstract
Advancements in 3D rendering like Gaussian Splatting (GS) allow novel view synthesis and real-time rendering in virtual reality (VR). However, GS-created 3D environments are often difficult to edit. For scene enhancement or to incorporate 3D assets, segmenting Gaussians by class is essential. Existing segmentation approaches are typically limited to certain types of scenes, e.g., "circular" scenes, to determine clear object boundaries. However, this method is ineffective when removing large objects in non-"circling" scenes such as large outdoor scenes.We propose Semantics-Controlled GS (SCGS), a segmentation-driven GS approach, enabling the separation of large scene parts in uncontrolled, natural environments. SCGS allows scene editing and the extraction of scene parts for VR. Additionally, we introduce a challenging outdoor dataset, overcoming the "circling" setup. We outperform the state-of-the-art in visual quality on our dataset and in segmentation quality on the 3D-OVS dataset. We conducted an exploratory user study, comparing a 360-video, plain GS, and SCGS in VR with a fixed viewpoint. In our subsequent main study, users were allowed to move freely, evaluating plain GS and SCGS. Our main study results show that participants clearly prefer SCGS over plain GS. We overall present an innovative approach that surpasses the state-of-the-art both technically and in user experience.
Hannah Schieber, Jacob Young, Tobias Langlotz, Stefanie Zollmann, Daniel Roth 0001
VR5
2025 Multi-Layer Gaussian Splatting for Immersive Anatomy Visualization
abstract
In medical image visualization, path tracing of volumetric medical data like computed tomography (CT) scans produces lifelike three-dimensional visualizations. Immersive virtual reality (VR) displays can further enhance the understanding of complex anatomies. Going beyond the diagnostic quality of traditional 2D slices, they enable interactive 3D evaluation of anatomies, supporting medical education and planning. Rendering high-quality visualizations in real-time, however, is computationally intensive and impractical for compute-constrained devices like mobile headsets. We propose a novel approach utilizing Gaussian Splatting (GS) to create an efficient but static intermediate representation of CT scans. We introduce a layered GS representation, incrementally including different anatomical structures while minimizing overlap and extending the GS training to remove inactive Gaussians. We further compress the created model with clustering across layers. Our approach achieves interactive frame rates while preserving anatomical structures, with quality adjustable to the target hardware. Compared to standard GS, our representation retains some of the explorative qualities initially enabled by immersive path tracing. Selective activation and clipping of layers are possible at rendering time, adding a degree of interactivity to otherwise static GS models. This could enable scenarios where high computational demands would otherwise prohibit using path-traced medical volumes.
Constantin Kleinbeck, Hannah Schieber, Klaus Engel, Ralf Gutjahr, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.5
2025 Investigating the Impact of Video Pass-Through Embodiment on Presence and Performance in Virtual Reality
abstract
Creating a compelling sense of presence and embodiment can enhance the user experience in virtual reality (VR). One method to accomplish this is through self-representation with embodied personalized avatars or video self-avatars. However, these approaches require external hardware and primarily evaluate hand representations in VR across various tasks. We therefore present in this paper an alternative approach: video Pass-Through Embodiment (PTE), which utilizes the per-eye real-time depth map from Head-Mounted Displays (HMDs) traditionally used for Augmented Reality features. This method allows the user's real body to be cut out of the pass-through video stream and be represented in the VR environment without the need for additional hardware. To evaluate our approach, we conducted a between-subjects study involving 40 participants who completed a seated object sorting task using either PTE or a customized avatar. The results show that PTE, despite its limited depth resolution that leads to some visual artifacts, significantly enhances the user's sense of presence and embodiment. In addition, PTE does not negatively affect task performance, cognitive load, or cause VR sickness. These findings imply that video pass-through embodiment offers a practical and efficient alternative to traditional avatar-based methods in VR.
Kristoffer Waldow, Constantin Kleinbeck, Arnulph Fuhrmann, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.4
2024 HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic Scenarios
abstract
Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches, research is shifting towards category-level pose estimation for practical applications. Current category-level datasets, however, fall short in annotation quality and pose variety. Addressing this, we introduce HouseCat6D, a new category-level 6D pose dataset. It features 1) multi-modality with Polarimetric RGB and Depth (RGBD+P), 2) encompasses 194 diverse objects across 10 household cat-egories, including two photometrically challenging ones, and 3) provides high-quality pose annotations with an error range of only 1.35 mm to 1.74 mm. The dataset also includes 4) 41 large-scale scenes with comprehensive view-point and occlusion coverage,5) a checkerboard-free en-vironment, and 6) dense 6D parallel-jaw robotic grasp annotations. Additionally, we present benchmark results for leading category-level pose estimation networks.
Patrick Ruhkamp, Guangyao Zhai, Hannah Schieber, Giulia Rizzoli, Pengyuan Wang 0002, Hongcheng Zhao, Lorenzo Garattoni, Daniel Roth 0001, Sven Meier, Nassir Navab, Benjamin Busam
CVPR10
2024 ASDF: Assembly State Detection Utilizing Late Fusion by Integrating 6D Pose Estimation
abstract
In medical and industrial domains, providing guidance for assembly processes can be critical to ensure efficiency and safety. Errors in assembly can lead to significant consequences such as extended surgery times and prolonged manufacturing or maintenance times in industry. Assembly scenarios can benefit from in-situ augmented reality visualization, i.e., augmentations in close proximity to the target object, to provide guidance, reduce assembly times, and minimize errors. In order to enable in-situ visualization, 6D pose estimation can be leveraged to identify the correct location for an augmentation. Existing 6D pose estimation techniques primarily focus on individual objects and static captures. However, assembly scenarios have various dynamics, including occlusion during assembly and dynamics in the appearance of assembly objects. Existing work focus either on object detection combined with state detection, or focus purely on the pose estimation. To address the challenges of 6D pose estimation in combination with assembly state detection, our approach ASDF builds upon the strengths of YOLOv8, a real-time capable object detection framework. We extend this framework, refine the object pose, and fuse pose knowledge with network-detected pose information. Utilizing our late fusion in our Pose2State module results in refined 6D pose estimation and assembly state detection. By combining both pose and state information, our Pose2State module predicts the final assembly state with precision. The evaluation of our ASDF dataset shows that our Pose2State module leads to an improved assembly state detection and that the improvement of the assembly state further leads to a more robust 6D pose estimation. Moreover, on the GBOT dataset, we outperform the pure deep learning-based network and even outperform the hybrid and pure tracking-based approaches.
Hannah Schieber, Shiyu Li 0003, Niklas Corell, Philipp Beckerle, Julian Kreimeier, Daniel Roth 0001
ISMAR6
2024 GBOT: Graph-Based 3D Object Tracking for Augmented Reality-Assisted Assembly Guidance
abstract
Guidance for assemblable parts is a promising field for augmented reality. Augmented reality assembly guidance requires 6D object poses of target objects in real time. Especially in time-critical medical or industrial settings, continuous and markerless tracking of individual parts is essential to visualize instructions superimposed on or next to the target object parts. In this regard, occlusions by the user’s hand or other objects and the complexity of different assembly states complicate robust and real-time markerless multi-object tracking. To address this problem, we present Graph-based Object Tracking (GBOT), a novel graph-based single-view RGB-D tracking approach. The real-time markerless multi-object tracking is initialized via 6D pose estimation and updates the graph-based assembly poses. The tracking through various assembly states is achieved by our novel multi-state assembly graph. We update the multi-state assembly graph by utilizing the relative poses of the individual assembly parts. Linking the individual objects in this graph enables more robust object tracking during the assembly process. For evaluation, we introduce a synthetic dataset of publicly available and 3D printable assembly assets as a benchmark for future work. Quantitative experiments in synthetic data and further qualitative study in real test data show that GBOT can outperform existing work towards enabling context-aware augmented reality assembly guidance. Dataset and code will be made publically available.****https://github.com/roth-hex-lab/gbot
Shiyu Li 0003, Hannah Schieber, Niklas Corell, Bernhard Egger 0001, Julian Kreimeier, Daniel Roth 0001
VR6
2024 Neural Motion Tracking: Formative Evaluation of Zero Latency Rendering
abstract
Low motion-to-photon latencies between physical movement and rendering updates are crucial for an immersive virtual reality (VR) experience and to avoidusers’ discomfort and sickness. Current methods aim to minimize the delay between the motion measurement and rendering at the cost of increasing technical complexity and possibly decreasing accuracy. By relying on capturing physical motion, these strategies will, by nature, not result in zero latency rendering or will be based on prediction and resulting uncertainty. This paper presents and evaluates a novel alternative and proof of principle for VR motion tracking that could enable motion-to-photon latencies of zero and below zero in time. We termed our concept Neural Motion Tracking, which we define as the sensing and assessment of motion through human neural activation of the somatic nervous system. In contrast to measuring physical activity, the key principle is that we aim to utilize the physiological timeframe between a user’s intention and the execution of motion. We aim to foresee upcoming motion ahead of the physical movement, by sampling preceding electromyographic signals before the muscle activation. The electromechanical delay (EMD) between potential change in the muscle activation and actual physical movement opens a gap in which measurement can be taken and evaluated before the physical motion. In a first proof of principle, we evaluated the concept with two activities, arm bending and head rotation, measured with a binary activation measure. Our results indicate that it is possible to predict movement and update a rendering up to 2 ms before its physical execution, which is assessed by optical tracking after approximately 4 ms. However, to make the best use of this advantage, electromyography (EMG) sensor data should be as high quality as possible (i.e., low noise and from muscle-near electrodes). Our results empirically quantify this characteristic for the first time when compared to state-of-the-art optical tracking systems for VR. We discuss our results and potential pathways to motivate further work toward marker- and latency-less motion tracking.
Daniel Roth 0001, Valentin Bräutigam, Nidhi Joshi, Constantin Kleinbeck, Hannah Schieber, Julian Kreimeier
VRST1
2024 A Systematic Literature Review of User Evaluation in Immersive Analytics
abstract
Abstract User evaluation is a common and useful tool for systematically generating knowledge and validating novel approaches in the domain of Immersive Analytics. Since this research domain centres around users, user evaluation is of extraordinary relevance. Additionally, Immersive Analytics is an interdisciplinary field of research where different communities bring in their own methodologies. It is vital to investigate and synchronise these different approaches with the long‐term goal to reach a shared evaluation framework. While there have been several studies focusing on Immersive Analytics as a whole or on certain aspects of the domain, this is the first systematic review of the state of evaluation methodology in Immersive Analytics. The main objective of this systematic literature review is to illustrate methodologies and research areas that are still underrepresented in user studies by identifying current practice in user evaluation in the domain of Immersive Analytics in coherence with the PRISMA protocol. (see https://www.acm.org/publications/class-2012 )
Judith Friedl-Knirsch, Fabian Pointecker, S. Pfistermüller, Christian Stach, Christoph Anthes, Daniel Roth 0001
Comput. Graph. Forum6
2024 NeRFtrinsic Four: An end-to-end trainable NeRF jointly optimizing diverse intrinsic and extrinsic camera parameters
abstract
Novel view synthesis using neural radiance fields (NeRF) is the state-of-the-art technique for generating high-quality images from novel viewpoints. Existing methods require a priori knowledge about extrinsic and intrinsic camera parameters. This limits their applicability to synthetic scenes, or real-world scenarios with the necessity of a preprocessing step. Current research on the joint optimization of camera parameters and NeRF focuses on refining noisy extrinsic camera parameters and often relies on the preprocessing of intrinsic camera parameters. Further approaches are limited to cover only one single camera intrinsic. To address these limitations, we propose a novel end-to-end trainable approach called NeRFtrinsic Four. We utilize Gaussian Fourier features to estimate extrinsic camera parameters and dynamically predict varying intrinsic camera parameters through the supervision of the projection error. Our approach outperforms existing joint optimization methods on LLFF and BLEFF. In addition to these existing datasets, we introduce a new dataset called iFF with varying intrinsic camera parameters. NeRFtrinsic Four is a step forward in joint optimization NeRF-based view synthesis and enables more realistic and flexible rendering in real-world scenarios with varying camera parameters. • A dynamic joint end-to-end trainable optimization framework, capable of handling diverse cameras. • A pose-multilayer perceptron (MLP), using Gaussian Fourier features for the handling of challenging poses. • Our novel iFF dataset focusing on the challenge of diverse cameras, on which we demonstrate the advantages of NeRFtrinsic Four.
Hannah Schieber, Fabian Deuser, Bernhard Egger 0001, Norbert Oswald, Daniel Roth 0001
Comput. Vis. Image Underst.5
2024 Indoor Synthetic Data Generation: A Systematic Review
abstract
Deep learning-based object recognition, 6D pose estimation, and semantic scene understanding require a large amount of training data to achieve generalization. Time-consuming annotation processes, privacy, and security aspects lead to a scarcity of real-world datasets. To overcome this lack of data, synthetic data generation has been proposed, including multiple facets in the area of domain randomization to extend the data distribution. The objective of this review is to identify methods applied for synthetic data generation aiming to improve 6D pose estimation, object recognition, and semantic scene understanding in indoor scenarios. We further review methods used to extend the data distribution and discuss best practices to bridge the gap between synthetic and real-world data. We adhered to the guidelines of the systematic PRISMA technique. Three databases, IEEE Xplore, Springer Link, and ACM, and an additional manual search were conducted. In total, we identified 241 studies and included 34 in our systematic review. In summary, synthetic data generation has been performed using crop-out methods, graphic APIs, 3D modeling or authoring tools, or game engine-based methods. To extend the data distribution, varying scene parameters, i.e., lighting conditions or textures and the use of distracting objects in the scene are promising.
Hannah Schieber, Kubilay Can Demir, Constantin Kleinbeck, Seung-Hee Yang, Daniel Roth 0001
Comput. Vis. Image Underst.5
2024 A Study on Collaborative Visual Data Analysis in Augmented Reality with Asymmetric Display Types
abstract
Collaboration is a key aspect of immersive visual data analysis. Due to its inherent benefit of seeing co-located collaborators, augmented reality is often useful in such collaborative scenarios. However, to enable the augmentation of the real environment, there are different types of technology available. While there are constant developments in specific devices, each of these device types provide different premises for collaborative visual data analysis. In our work we combine handheld, optical see-through and video see-through displays to explore and understand the impact of these different device types in collaborative immersive analytics. We conducted a mixed-methods collaborative user study where groups of three performed a shared data analysis task in augmented reality with each user working on a different device, to explore differences in collaborative behaviour, user experience and usage patterns. Both quantitative and qualitative data revealed differences in user experience and usage patterns. For collaboration, the different display types influenced how well participants could participate in the collaborative data analysis, nevertheless, there was no measurable effect in verbal communication.
Judith Friedl-Knirsch, Christian Stach, Fabian Pointecker, Christoph Anthes, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.5
2023 Investigating the Effects of Selective Information Presentation in Intensive Care Units Using Virtual Reality
abstract
Medical personnel working in intensive care units (ICUs) are continuously exposed to a multitude of alarms emanating from various monitoring devices, such as cardiac monitors, ventilators, or infusion pumps. The sheer volume of alarms, coupled with high false positive rates, can lead to alarm fatigue. This phenomenon compromises patient safety and places an additional burden on nurses who must diligently prioritize and respond to alarms in the highly dynamic environment. While the testing of stress-reducing strategies in a real ICU is challenging, virtual reality (VR) represents a powerful tool and methodology to simulate an ICU environment and test optimization scenarios for alarm display strategies. For example, redistributing alarms to responsible individuals (personalized information presentation) has been proposed as a solution, but testing in real ICU environments is not applicable due to critical patient safety. In this paper, we present a VR simulation of an ICU to simulate comparable stress situations, as well as to assess the impact of a selective and personalized alarm representation strategy in an evaluation study in two conditions. A stress condition mirrors the current ubiquitous audible alarm distribution in most ICUs, where alarms are heard non-patient-specific throughout the ward. In an experimental condition, alarms are filtered patient-specific to reduce information overload and noise pollution. Our user study with medical personnel and novices shows that stress levels can be simulated with our system as indicated by physiological responses. Further, we show that the perceived task load can be reduced with selective information presentation. We discuss the potential benefits of ICU simulations as a methodology and personalized alarm distribution as a first potential strategy for future technologies in ICUs.
Luisa Theelke, Fynn-Lennardt Metzler, Julian Kreimeier, Christopher Hauer, Johannes Binder, Daniel Roth 0001
ISMAR6
2023 Mixed Reality 3D Teleconsultation for Emergency Decompressive Craniotomy: An Evaluation with Medical Residents
abstract
Enabling collaborative telepresence in healthcare, especially surgical procedures, presents a critical challenge. The decompressive craniotomy procedure stands out as particularly complex and time-sensitive. The current teleconsultation approach relies on 2D color cameras, often offering only a fixed view and limited visual capabilities between experts and surgeons. However, teleconsultation can be addressed with Mixed Reality and immersive technology to potentially enable a better consultation of the procedure. We conducted an extensive user study focusing on decompressive craniotomy to investigate the advantages and challenges of our 3D teleconsultation system compared to a 2D video-based consultation system. Our 3D teleconsultation system leverages real-time 3D reconstruction of the patient and environment to empower experts to provide guidance and create virtual 3D annotations. The study utilized 3D-printed head models to perform a lifelike surgical intervention. It involved 14 medical residents and demonstrated an in-vitro 17% improvement in accurately describing the incision size on the patient’s head, contributing to potentially improved patient outcomes.
Daniel Roth 0001, Robin Strak, Frieder Pankratz, Julia Reichling, Clemens Kraetsch, Simon Weidert, Marc Lazarovici, Nassir Navab, Ulrich Eck
ISMAR2
2023 Deep Learning in Surgical Workflow Analysis: A Review of Phase and Step Recognition
abstract
OBJECTIVE: In the last two decades, there has been a growing interest in exploring surgical procedures with statistical models to analyze operations at different semantic levels. This information is necessary for developing context-aware intelligent systems, which can assist the physicians during operations, evaluate procedures afterward or help the management team to effectively utilize the operating room. The objective is to extract reliable patterns from surgical data for the robust estimation of surgical activities performed during operations. The purpose of this article is to review the state-of-the-art deep learning methods that have been published after 2018 for analyzing surgical workflows, with a focus on phase and step recognition. METHODS: Three databases, IEEE Xplore, Scopus, and PubMed were searched, and additional studies are added through a manual search. After the database search, 343 studies were screened and a total of 44 studies are selected for this review. CONCLUSION: The use of temporal information is essential for identifying the next surgical action. Contemporary methods used mainly RNNs, hierarchical CNNs, and Transformers to preserve long-distance temporal relations. The lack of large publicly available datasets for various procedures is a great challenge for the development of new and robust models. As supervised learning strategies are used to show proof-of-concept, self-supervised, semi-supervised, or active learning methods are used to mitigate dependency on annotated data. SIGNIFICANCE: The present study provides a comprehensive review of recent methods in surgical workflow analysis, summarizes commonly used architectures, datasets, and discusses challenges.
Kubilay Can Demir, Hannah Schieber, Tobias Weise, Daniel Roth 0001, Matthias S. May, Andreas K. Maier, Seung-Hee Yang
IEEE J. Biomed. Health Informatics4
2023 Injured Avatars: The Impact of Embodied Anatomies and Virtual Injuries on Well-Being and Performance
abstract
Human cognition relies on embodiment as a fundamental mechanism. Virtual avatars allow users to experience the adaptation, control, and perceptual illusion of alternative bodies. Although virtual bodies have medical applications in motor rehabilitation and therapeutic interventions, their potential for learning anatomy and medical communication remains underexplored. For learners and patients, anatomy, procedures, and medical imaging can be abstract and difficult to grasp. Experiencing anatomies, injuries, and treatments virtually through one's own body could be a valuable tool for fostering understanding. This work investigates the impact of avatars displaying anatomy and injuries suitable for such medical simulations. We ran a user study utilizing a skeleton avatar and virtual injuries, comparing to a healthy human avatar as a baseline. We evaluate the influence on embodiment, well-being, and presence with self-report questionnaires, as well as motor performance via an arm movement task. Our results show that while both anatomical representation and injuries increase feelings of eeriness, there are no negative effects on embodiment, well-being, presence, or motor performance. These findings suggest that virtual representations of anatomy and injuries are suitable for medical visualizations targeting learning or communication without significantly affecting users' mental state or physical control within the simulation.
Constantin Kleinbeck, Hannah Schieber, Julian Kreimeier, Alejandro Martin-Gomez, Mathias Unberath, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.6
2022 The Effects of Avatar and Environment Design on Embodiment, Presence, Activation, and Task Load in a Virtual Reality Exercise Application
abstract
The development of embodied Virtual Reality (VR) systems involves multiple central design choices. These design choices affect the user perception and therefore require thorough consideration. This article reports on two user studies investigating the influence of common design choices on relevant intermediate factors (sense of embodiment, presence, motivation, activation, and task load) in a VR application for physical exercises. The first study manipulated the avatar fidelity (abstract, partial body vs. anthropomorphic, full-body) and the environment (with vs. without mirror). The second study manipulated the avatar type (healthy vs. injured) and the environment type (beach vs. hospital) and, hence, the avatar-environment congruence. The full-body avatar significantly increased the sense of embodiment and decreased mental demand. Interestingly, the mirror did not influence the dependent variables. The injured avatar significantly increased the temporal demand. The beach environment significantly reduced the tense activation. On the beach, participants felt more present in the incongruent condition embodying the injured avatar.
Andrea Bartl, Christian Merz, Daniel Roth 0001, Marc Erich Latoschik
ISMAR3
2022 The Impact of Focus and Context Visualization Techniques on Depth Perception in Optical See-Through Head-Mounted Displays
abstract
Estimating the depth of virtual content has proven to be a challenging task in Augmented Reality (AR) applications. Existing studies have shown that the visual system makes use of multiple depth cues to infer the distance of objects, occlusion being one of the most important ones. The ability to generate appropriate occlusions becomes particularly important for AR applications that require the visualization of augmented objects placed below a real surface. Examples of these applications are medical scenarios in which the visualization of anatomical information needs to be observed within the patient's body. In this regard, existing works have proposed several focus and context (F+C) approaches to aid users in visualizing this content using Video See-Through (VST) Head-Mounted Displays (HMDs). However, the implementation of these approaches in Optical See-Through (OST) HMDs remains an open question due to the additive characteristics of the display technology. In this article, we, for the first time, design and conduct a user study that compares depth estimation between VST and OST HMDs using existing in-situ visualization methods. Our results show that these visualizations cannot be directly transferred to OST displays without increasing error in depth perception tasks. To tackle this gap, we perform a structured decomposition of the visual properties of AR F+C methods to find best-performing combinations. We propose the use of chromatic shadows and hatching approaches transferred from computer graphics. In a second study, we perform a factorized analysis of these combinations, showing that varying the shading type and using colored shadows can lead to better depth estimation when using OST HMDs.
Alejandro Martin-Gomez, Jakob Weiss, Andreas Keller, Ulrich Eck, Daniel Roth 0001, Nassir Navab
IEEE Trans. Vis. Comput. Graph.5
2022 Stereopsis Only: Validation of a Monocular Depth Cues Reduced Gamified Virtual Reality with Reaction Time Measurement
abstract
The visual depth perception is composed of monocular and binocular depth cues. Studies show that in absence of binocular depth cues the performance of visuomotor tasks like pointing to or grasping objects is limited. Thus, binocular depth cues are of great importance for motor control required in everyday life. However, binocular depth cues like retinal disparity (basis for stereopsis) might be influenced due to developmental disorders of the visual system. For example, amblyopia in which one eye's visual input is not processed leads to loss of stereopsis. The primary amblyopia treatment is occlusion of the healthy eye to force the amblyopic eye to train. However, improvements in stereopsis are poor. Therefore, binocular treatments arose that equilibrate both eyes' visual input to enable binocular vision. However, most approaches rely on divided stimuli which do not account for loss of stereopsis. We created a Virtual Reality (VR) with reduced monocular depth cues in which a stereoscopic task is shown to both eyes simultaneously, consisting of two balls jumping towards the user. One ball appears closer to the user which must be identified. To evaluate the task performance the reaction time is measured. We validated our approach with 18 participants with stereopsis under three contrast settings including one leading to monocular vision. The number of correct responses reduces from 90% under binocular vision to 52% under monocular vision corresponding to random guessing. Our results indicate that it is possible to disable monocular depth cues and create a dynamic stereoscopic task inside a VR.
Wolfgang A. Mehringer, Markus Wirth, Daniel Roth 0001, Georg Michelson, Björn M. Eskofier
IEEE Trans. Vis. Comput. Graph.3
2022 A Virtual Reality Based System for the Screening and Classification of Autism
abstract
Autism - also known as Autism Spectrum Disorders or Autism Spectrum Conditions - is a neurodevelopmental condition characterized by repetitive behaviours and differences in communication and social interaction. As a consequence, many autistic individuals may struggle in everyday life, which sometimes manifests in depression, unemployment, or addiction. One crucial problem in patient support and treatment is the long waiting time to diagnosis, which was approximated to thirteen months on average. Yet, the earlier an intervention can take place the better the patient can be supported, which was identified as a crucial factor. We propose a system to support the screening of Autism Spectrum Disorders based on a virtual reality social interaction, namely a shopping experience, with an embodied agent. During this everyday interaction, behavioral responses are tracked and recorded. We analyze this behavior with machine learning approaches to classify participants from an autistic participant sample in comparison to a typically developed individuals control sample with high accuracy, demonstrating the feasibility of the approach. We believe that such tools can strongly impact the way mental disorders are assessed and may help to further find objective criteria and categorization.
Marta Robles, Negar Namdarian, Julia Otto, Evelyn Wassiljew, Nassir Navab, Christine M. Falter-Wagner, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.7
2021 The Impact of Implicit and Explicit Feedback on Performance and Experience during VR-Supported Motor Rehabilitation
abstract
This paper examines the impact of implicit and explicit feedback in Virtual Reality (VR) on performance and user experience during motor rehabilitation. In this work, explicit feedback consists of visual and auditory cues provided by a virtual trainer, compared to traditional feedback provided by a real physiotherapist. Implicit feedback was generated by the walking motion of the virtual trainer accompanying the patient during virtual walks. Here, the potential synchrony of movements between the trainer and trainee is intended to create an implicit visual affordance of motion adaption. We hypothesize that this will stimulate the activation of mirror neurons, thus fostering neuroadaptive processes. We conducted a clinical user study in a rehabilitation center employing a gait robot. We investigated the performance outcome and subjective experience of four resulting VR-supported rehabilitation conditions: with/without explicit feedback, and with/without implicit (synchronous motion) stimulation by a virtual trainer. We further included two baseline conditions reflecting the current NonVR procedure in the rehabilitation center. Our results show that additional feedback generally resulted in better patient performance, objectively assessed by the necessary applied support force of the robot. Additionally, our VR-supported rehabilitation procedure improved enjoyment and satisfaction, while no negative impacts could be observed. Implicit feedback and adapted motion synchrony by the virtual trainer led to higher mental demand, giving rise to hopes of increased neural activity and neuroadaptive stimulation.
Negin Hamzeheinejad, Daniel Roth 0001, Samantha Monty, Julian Breuer, Anuschka Rodenbergc, Marc Erich Latoschik
VR2
2021 The Impact of Avatar Appearance, Perspective and Context on Gait Variability and User Experience in Virtual Reality
abstract
Gait supervision plays an important role in the diagnosis, analysis and rehabilitation of motor impairments and neurodegenerative disorders. For example, in Parkinson's disease, gait assessment is used for progression observation and medication guidance. Previous work has presented the potential of virtual reality (VR) supported gait applications. While virtual environments and user representation strategies are used for gait applications, the influence of appearance and context cues on gait performance is not extensively researched. In this paper, we analyzed the influence of avatar appearance, environment awareness, and camera perspective on gait parameters relevant for clinical application. Four different avatar appearances, varying in abstraction, two environmental settings, as well as an egocentric and exocentric camera perspective were compared in three walking tasks on a treadmill. Our results show that variability, as an indicator for gait stability, is significantly impacted by VR exposure in comparison to a real world (in vivo) baseline. Further, our results revealed that walking tasks influence gait behavior significantly different in VR compared to in vivo. Overall, these findings suggest that particular care has to be taken when assessing gait characteristics acquired from subjects immersed in VR and that equivalence of results with in vivo may not be blindly assumed.
Markus Wirth, Stefan Gradl, Georg Prosinger, Felix Kluge 0001, Daniel Roth 0001, Björn M. Eskofier
VR5
2021 Magnoramas: Magnifying Dioramas for Precise Annotations in Asymmetric 3D Teleconsultation
abstract
When users create hand-drawn annotations in Virtual Reality they often reach their physical limits in terms of precision, especially if the region to be annotated is small. One intuitive solution employs magnification beyond natural scale. However, scaling the whole environment results in wrong assumptions about the coherence between physical and virtual space. In this paper, we introduce Mag-noramas, a novel interaction method for selecting and extracting a region of interest that the user can subsequently scale and transform inside the virtual space. Our technique enhances the user's capabilities to perform supernaturally precise virtual annotations on virtual objects. We explored our technique in a user study within asimplified clinical scenario of a teleconsultation-supported craniectomy procedure that requires accurate annotations on a human head. Teleconsultation was performed asymmetrically between a remote expert in Virtual Reality that collaborated with a local user through Augmented Reality. The remote expert operates inside a reconstructed environment, captured from RGB-D sensors at the local site, and is embodied by an avatar to establish co-presence. The results show that Magnoramas significantly improve the precision of annotations while preserving usability and perceived presence measures compared to the baseline method. By hiding the 3D reconstruction while keeping the Magnorama, users can intentionally choose to lower their perceived social presence and focus on their tasks.
Alexander Winkler, Frieder Pankratz, Marc Lazarovici, Dirk Wilhelm, Ulrich Eck, Daniel Roth 0001, Nassir Navab
VR7
2021 Avatars for Teleconsultation: Effects of Avatar Embodiment Techniques on User Perception in 3D Asymmetric Telepresence
abstract
A 3D Telepresence system allows users to interact with each other in a virtual, mixed, or augmented reality (VR, MR, AR) environment, creating a shared space for collaboration and communication. There are two main methods for representing users within these 3D environments. Users can be represented either as point cloud reconstruction-based avatars that resemble a physical user or as virtual character-based avatars controlled by tracking the users' body motion. This work compares both techniques to identify the differences between user representations and their fit in the reconstructed environments regarding the perceived presence, uncanny valley factors, and behavior impression. Our study uses an asymmetric VR/AR teleconsultation system that allows a remote user to join a local scene using VR. The local user observes the remote user with an AR head-mounted display, leading to facial occlusions in the 3D reconstruction. Participants perform a warm-up interaction task followed by a goal-directed collaborative puzzle task, pursuing a common goal. The local user was represented either as a point cloud reconstruction or as a virtual character-based avatar, in which case the point cloud reconstruction of the local user was masked. Our results show that the point cloud reconstruction-based avatar was superior to the virtual character avatar regarding perceived co-presence, social presence, behavioral impression, and humanness. Further, we found that the task type partly affected the perception. The point cloud reconstruction-based approach led to higher usability ratings, while objective performance measures showed no significant difference. We conclude that despite partly missing facial information, the point cloud-based reconstruction resulted in better conveyance of the user behavior and a more coherent fit into the simulation context.
Gleb Gorbachev, Ulrich Eck, Frieder Pankratz, Nassir Navab, Daniel Roth 0001
IEEE Trans. Vis. Comput. Graph.6
2020 Augmented Mirrors
abstract
A recurrent problem in egocentric Augmented Reality (AR) applications is the misestimation of depth. Providing alternative views from non-egocentric perspectives can convey useful information for applications that require the correct judgment of depth as it is in the case of placement and alignment of virtual and real content, but also for exploration and visualization tasks.In this paper, we introduce Augmented Mirrors. Through the integration of a real mirror, our approach is capable to reflect changes of the real and virtual content of an AR application while users benefit from the perceptual advantages of using mirrors. Our concept, simple yet effective, only requires tracking the user and mirror poses with the accuracy demanded by a specific application. To showcase the potential and flexibility of the Augmented Mirrors, we present and discuss multiple examples ranging from alignment, exploration, spatial understanding, and selective content visualization using different AR-enabled devices and tracking technologies. We envision the Augmented Mirrors as a new and valuable concept that can be used in applications that benefit from additional viewpoints and require the simultaneous visualization of real and virtual content.
Alejandro Martin-Gomez, Alexander Winkler, Daniel Roth 0001, Ulrich Eck, Nassir Navab
ISMAR4
2020 Construction of the Virtual Embodiment Questionnaire (VEQ)
abstract
User embodiment is important for many virtual reality (VR) applications, for example, in the context of social interaction, therapy, training, or entertainment. However, there is no data-driven and validated instrument to empirically measure the perceptual aspects of embodiment, necessary to reliably evaluate this important phenomenon. To provide a method to assess components of virtual embodiment in a reliable and consistent fashion, we constructed a Virtual Embodiment Questionnaire (VEQ). We reviewed previous literature to identify applicable constructs and questionnaire items, and performed a confirmatory factor analysis (CFA) on the data from three experiments ( N=196). The analysis confirmed three factors: (1) ownership of a virtual body, (2) agency over a virtual body, and (3) the perceived change in the body schema. A fourth study ( N=22) was conducted to confirm the reliability and validity of the scale, by investigating the impacts of latency and latency jitter present in the simulation. We present the proposed scale and study results and discuss resulting implications.
Daniel Roth 0001, Marc Erich Latoschik
IEEE Trans. Vis. Comput. Graph.1
2019 Physiological Effectivity and User Experience of Immersive Gait Rehabilitation
abstract
Gait impairments from neurological injuries require repeated and exhaustive physical exercises for rehabilitation. Prolonged physical training in clinical environments can easily become frustrating and de-motivating for various reasons which in turn risks to decrease efficiency during the healing process. This paper introduces an immersive VR system for gait rehabilitation which targets user experience and increase of motivation while evoking comparable physiological responses needed for successful training effects. The system provides a virtual environment consisting of open fields, forest, mountains, waterfalls, animals, and a beach for inspiring strolls and is able to include a virtual trainer as a companion during the walks. We evaluated the ecological validity of the system with healthy subjects before performing the clinical trial. We assessed the system's target qualities with a longitudinal study with 45 healthy participants in three consecutive days in comparison to a baseline non-VR condition. The system was able to evoke similar physiological responses. The workload was increased for the VR condition but the system also elicited a higher enjoyment and motivation which was the main goal. The latter benefits slightly decreased over time (as did workload) while they were still higher than in the non-VR condition. The virtual trainer did not show to be beneficial, the corresponding implications are discussed. Overall, the approach shows promising results which renders the system a viable alternative for the given use case while it motivates interesting direction for future work.
Negin Hamzeheinejad, Daniel Roth 0001, Daniel Götz, Franz Weilbach, Marc Erich Latoschik
VR2
2019 Technologies for Social Augmentations in User-Embodied Virtual Reality
abstract
Technologies for Virtual, Mixed, and Augmented Reality (VR, MR, and AR) allow to artificially augment social interactions and thus to go beyond what is possible in real life. Motivations for the use of social augmentations are manifold, for example, to synthesize behavior when sensory input is missing, to provide additional affordances in shared environments, or to support inclusion and training of individuals with social communication disorders. We review and categorize augmentation approaches and propose a software architecture based on four data layers. Three components further handle the status analysis, the modification, and the blending of behaviors. We present a prototype (injectX) that supports behavior tracking (body motion, eye gaze, and facial expressions from the lower face), status analysis, decision-making, augmentation, and behavior blending in immersive interactions. Along with a critical reflection, we consider further technical and ethical aspects.
Daniel Roth 0001, Gary Bente, Peter Kullmann, David Mal, Christian Felix Purps, Kai Vogeley, Marc Erich Latoschik
VRST1
2019 Sick Moves! Motion Parameters as Indicators of Simulator Sickness
abstract
We explore motion parameters, more specifically gait parameters, as an objective indicator to assess simulator sickness in Virtual Reality (VR). We discuss the potential relationships between simulator sickness, immersion, and presence. We used two different camera pose (position and orientation) estimation methods for the evaluation of motion tasks in a large-scale VR environment: a simple model and an optimized model that allows for a more accurate and natural mapping of human senses. Participants performed multiple motion tasks (walking, balancing, running) in three conditions: a physical reality baseline condition, a VR condition with the simple model, and a VR condition with the optimized model. We compared these conditions with regard to the resulting sickness and gait, as well as the perceived presence in the VR conditions. The subjective measures confirmed that the optimized pose estimation model reduces simulator sickness and increases the perceived presence. The results further show that both models affect the gait parameters and simulator sickness, which is why we further investigated a classification approach that deals with non-linear correlation dependencies between gait parameters and simulator sickness. We argue that our approach could be used to assess and predict simulator sickness based on human gait parameters and we provide implications for future research.
Tobias Feigl, Daniel Roth 0001, Stefan Gradl, Markus Wirth, Marc Erich Latoschik, Björn M. Eskofier, Michael Philippsen, Christopher Mutschler
IEEE Trans. Vis. Comput. Graph.2
2018 Rethinking real-time strategy games for virtual reality
abstract
Recent improvements in virtual reality (VR) technology promise the opportunity to redesign established game genres, such as real-time strategy (RTS) games. In this work we have a look at a taxonomy of RTS games and apply it to RTS titles for VR. Hereby, we identify possible difficulties such as the need for novel means of navigation in VR. We discuss conceivable solutions and illustrate them by referring to relevant work by others and by means of AStar0ID, an exploratory prototype VR RTS science fiction game. Our main contribution is the systematic inspection and discussion of foundational RTS aspects in the context of VR and, thus, to provide a substantial basis to further rethink and evolve the RTS genre in this new light.
Samuel Truman, Nicolas Rapp, Daniel Roth 0001, Sebastian von Mammen
FDG3
2018 Beyond Replication: Augmenting Social Behaviors in Multi-User Virtual Realities
abstract
This paper presents a novel approach for the augmentation of social behaviors in virtual reality (VR). We designed three visual transformations for behavioral phenomena crucial to everyday social interactions: eye contact, joint attention, and grouping. To evaluate the approach, we let users interact socially in a virtual museum using a large-scale multi-user tracking environment. Using a between-subject design (N = 125) we formed groups of five participants. Participants were represented as simplified avatars and experienced the virtual museum simultaneously, either with or without the augmentations. Our results indicate that our approach can significantly increase social presence in multi-user environments and that the augmented experience appears more thought-provoking. Furthermore, the augmentations seem also to affect the actual behavior of participants with regard to more eye contact and more focus on avatars/objects in the scene. We interpret these findings as first indicators for the potential of social augmentations to impact social perception and behavior in VR.
Daniel Roth 0001, Constantin Kleinbeck, Tobias Feigl, Christopher Mutschler, Marc Erich Latoschik
VR1
2018 Simulator Sick but Still Immersed: A Comparison of Head-Object Collision Handling and Their Impact on Fun, Immersion, and Simulator Sickness
abstract
We compared three techniques for handling head-object collisions in room-scale virtual reality (VR). We developed a game whose mechanics induce such collisions which we either addressed (1) not at all, (2) by fading the screen information to black, or (3) by restricting translation, i.e. correcting the virtual offset in such a way that no penetration occurred. We measured these conditions' impact on simulator sickness, fun, and immersion perception. We found that the translation-restricted method yielded the greatest immersion value but also contributed the most to simulator sickness.
Peter Ziegler, Daniel Roth 0001, Andreas Knote, Michael Kreuzer, Sebastian von Mammen
VR2
2018 The Impact of Avatar Personalization and Immersion on Virtual Body Ownership, Presence, and Emotional Response
abstract
This article reports the impact of the degree of personalization and individualization of users' avatars as well as the impact of the degree of immersion on typical psychophysical factors in embodied Virtual Environments. We investigated if and how virtual body ownership (including agency), presence, and emotional response are influenced depending on the specific look of users' avatars, which varied between (1) a generic hand-modeled version, (2) a generic scanned version, and (3) an individualized scanned version. The latter two were created using a state-of-the-art photogrammetry method providing a fast 3D-scan and post-process workflow. Users encountered their avatars in a virtual mirror metaphor using two VR setups that provided a varying degree of immersion, (a) a large screen surround projection (L-shape part of a CAVE) and (b) a head-mounted display (HMD). We found several significant as well as a number of notable effects. First, personalized avatars significantly increase body ownership, presence, and dominance compared to their generic counterparts, even if the latter were generated by the same photogrammetry process and hence could be valued as equal in terms of the degree of realism and graphical quality. Second, the degree of immersion significantly increases the body ownership, agency, as well as the feeling of presence. These results substantiate the value of personalized avatars resembling users' real-world appearances as well as the value of the deployed scanning process to generate avatars for VR-setups where the effect strength might be substantial, e.g., in social Virtual Reality (VR) or in medical VR-based therapies relying on embodied interfaces. Additionally, our results also strengthen the value of fully immersive setups which, today, are accessible for a variety of applications due to the widely available consumer HMDs.
Thomas Waltemate, Dominik Gall, Daniel Roth 0001, Mario Botsch, Marc Erich Latoschik
IEEE Trans. Vis. Comput. Graph.3
2017 Socially immersive avatar-based communication
abstract
In this paper, we present SIAM-C, an avatar-mediated communication platform to study socially immersive interaction in virtual environments. The proposed system is capable of tracking, transmitting, representing body motion, facial expressions, and voice via virtual avatars and inherits the transmission of human behaviors that are available in real-life social interactions. Users are immersed using active stereoscopic rendering projected onto a life-size projection plane, utilizing the concept of “fish tank” virtual reality (VR). Our prototype connects two separate rooms and allows for socially immersive avatar-mediated communication in VR.
Daniel Roth 0001, Kristoffer Waldow, Marc Erich Latoschik, Arnulph Fuhrmann, Gary Bente
VR1
2017 The effect of avatar realism in immersive social virtual realities
abstract
This paper investigates the effect of avatar realism on embodiment and social interactions in Virtual Reality (VR). We compared abstract avatar representations based on a wooden mannequin with high fidelity avatars generated from photogrammetry 3D scan methods. Both avatar representations were alternately applied to participating users and to the virtual counterpart in dyadic social encounters to examine the impact of avatar realism on self-embodiment and social interaction quality. Users were immersed in a virtual room via a head mounted display (HMD). Their full-body movements were tracked and mapped to respective movements of their avatars. Embodiment was induced by presenting the users' avatars to themselves in a virtual mirror. Afterwards they had to react to a non-verbal behavior of a virtual interaction partner they encountered in the virtual space. Several measures were taken to analyze the effect of the appearance of the users' avatars as well as the effect of the appearance of the others' avatars on the users. The realistic avatars were rated significantly more human-like when used as avatars for the others and evoked a stronger acceptance in terms of virtual body ownership (VBO). There also was some indication of a potential uncanny valley. Additionally, there was an indication that the appearance of the others' avatars impacts the self-perception of the users.
Marc Erich Latoschik, Daniel Roth 0001, Dominik Gall, Jascha Achenbach, Thomas Waltemate, Mario Botsch
VRST2
2016 FaceBo: Real-time face and body tracking for faithful avatar synthesis
abstract
This paper introduces a low-cost framework capable of combining both real-time markerless face and body tracking for faithful avatar embodiment in Virtual Reality (VR). We discuss suitable hardware and software solutions and present a first prototype. This work lays the technological basis for further research on the importance of the appearance and behavioral realism of avatars, e.g., for the illusion of virtual body ownership, for social interactions in VR, as well as for VR entertainment applications (immersive games or movies).
Jean-Luc Lugrin, David Zilch, Daniel Roth 0001, Gary Bente, Marc Erich Latoschik
VR3
2016 A simplified inverse kinematic approach for embodied VR applications
abstract
In this paper, we compare a full body marker set with a reduced rigid body marker set supported by inverse kinematics. We measured system latency, illusion of virtual body ownership, and task load in an applied scenario for inducing acrophobia. While not showing a significant change in body ownership or task performance, results do show that latency and task load are reduced when using the rigid body inverse kinematics solution. The approach therefore has the potential to improve virtual reality experiences.
Daniel Roth 0001, Jean-Luc Lugrin, Julia Buser, Gary Bente, Arnulph Fuhrmann, Marc Erich Latoschik
VR1
2016 Avatar realism and social interaction quality in virtual reality
abstract
In this paper, we describe an experimental method to investigate the effects of reduced social information and behavioral channels in immersive virtual environments with full-body avatar embodiment. We compared physical-based and verbal-based social interactions in real world (RW) and virtual reality (VR). Participants were represented by abstract avatars that did not display gaze, facial expressions or social cues from appearance. Our results show significant differences in terms of presence and physical performance. However, differences in effectiveness in the verbal task were not present. Participants appear to efficiently compensate for missing social and behavioral cues by shifting their attentions to other behavioral channels.
Daniel Roth 0001, Jean-Luc Lugrin, Dmitri Galakhov, Arvid Hofmann, Gary Bente, Marc Erich Latoschik, Arnulph Fuhrmann
VR1
2016 Breaking bad behavior: immersive training of class room management
abstract
This article presents a fully immersive portable low-cost Virtual Reality system to train classroom management skills. An instructor controls the simulation of a virtual classroom populated with 24 semi-autonomous virtual agents via a desktop-based graphical user interface (GUI). The GUI provides behavior control and trainee evaluation widgets alongside a non-immersive view of the class and the trainee. The trainee's interface uses an Head-Mounted Display (HMD) and earphones for output. A depth camera and the HMD's built-in motion sensors are used for tracking the trainee and for avatar animation. An initial evaluation of both interfaces confirms the system's usefulness, specifically its capability to successfully simulate critical aspects of classroom management.
Marc Erich Latoschik, Jean-Luc Lugrin, Michael Habel, Daniel Roth 0001, Christian Seufert, Silke Grafe
VRST4
2016 FakeMi: a fake mirror system for avatar embodiment studies
abstract
This paper introduces a fake mirror system as a research tool to study the effect of avatar embodiment with non-visually immersive virtual environments. The system combines marker-less face and body tracking to animate the individual avatars seen in a stereoscopic display with a correct perspective projection. The display dimensions match typical dimensions of a real physical mirror and the animated avatars are rendered based on a geometrically correct reflection as expected from a real mirror including correct body and face animations. The first evaluation of the system reveals the high acceptance of the setup as well as a convincing illusion of a real mirror with different types of avatars.
Marc Erich Latoschik, Jean-Luc Lugrin, Daniel Roth 0001
VRST3
2016 Audio feedback and illusion of virtual body ownership in mixed reality
abstract
This paper presents an exploratory experiment measuring the role of audio feedback on the illusion of virtual body ownership (IVBO) under non-immersive mixed reality (MR) settings with Human and Non-Human avatars. Our preliminary results revealed that all avatars elicited a similar level of IVBO, despite the addition of audio feedback.
Jean-Luc Lugrin, David Obremski, Daniel Roth 0001, Marc Erich Latoschik
VRST3
2016 Avatar anthropomorphism and acrophobia
abstract
In this paper, we investigate the impact of avatar anthropomorphism on the fear of heights, when using full body avatar embodiment under an immersive virtual reality (VR) setting. Clear differences could be found in perceived anthropomorphism, but preliminary results do not show differences in stress level between Human and Non-Human avatars, although a high level of perceived secureness was reported with Non-Human avatars.
Jean-Luc Lugrin, Ivan Polyschev, Daniel Roth 0001, Marc Erich Latoschik
VRST3
2016 SIAMC: a socially immersive avatar mediated communication platform
abstract
In this paper, we present a avatar-mediated communication platform for socially immersive interaction in virtual reality (VR). Our approach is based on the combination of body tracking, facial expression tracking and "fishtank" VR. Our prototype enables two remote users to communicate via avatars.
Daniel Roth 0001, Kristoffer Waldow, Felix Stetter, Gary Bente, Marc Erich Latoschik, Arnulph Fuhrmann
VRST1