EDBT 2026 Demo / reviewers in the wild / expert
Henry Fuchs
dblp:f/HenryFuchs
· DBLP profile ↗
125ranked-venue papers
20as first author
19since 2021 · last 2025
0000-0002-8834-4638ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 107 · 17 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 66 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coherent 3D Portrait Video Reconstruction via Triplane FusionabstractRecent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user’s appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image but fail to faithfully preserve the user’s per-frame appearance (e.g., instantaneous facial expressions and lighting). As a result, neither of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user’s per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets. Shengze Wang 0002, Chao Liu 0064, Matthew A. Chan 0001, Michael Stengel, Henry Fuchs, Shalini De Mello, Koki Nagano |
CVPR | 6 |
| 2025 | BLADE: Single-view Body Mesh Estimation through Accurate Depth EstimationabstractSingle-Image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve both accurate 3D pose and 2D alignment at the same time. Error is mainly introduced by inaccurate perspective projection heuristically derived from orthographic parameters. To resolve this long-standing challenge, we present our method BLADE which accurately recovers perspective parameters from a single image without heuristic assumptions. We start from the inverse relationship between perspective distortion and the person’s Z-translation Tz, and we show that Tzcan be reliably estimated from the image. We then discuss the important role of Tzfor accurate human mesh recovery estimated from closerange images. Finally, we show that, once Tzand the 3D human mesh are estimated, one can accurately recover the focal length and full 3D translation. Extensive experiments on standard benchmarks and real-world close-range images show that our method accurately recovers projection parameters from a single image, and consequently attains state-of-the-art accuracy on both 3D pose estimation and 2D alignment for a wide range of images. Shengze Wang 0002, Tianye Li, Ye Yuan 0007, Henry Fuchs, Koki Nagano, Shalini De Mello, Michael Stengel |
CVPR | 5 |
| 2025 | HoloZip: High Hologram Compression via Latent-of-Latent CodingabstractHolographic displays are gaining increasing popularity, particularly as a holy grail solution to augmented and virtual reality (AR/VR) wearable displays. However, the generation of holograms is computationally intensive for AR/VR edge devices, which are expected to be compact and lightweight with low power consumption and heat dissipation to last all day. To address this, we propose a distributed hologram generation framework, dubbed HoloZip, to jointly generate, compress, transmit, and decode computer-generated holograms across the cloud-edge devices. Specifically, we perform compute intensive hologram generation via a vision transformer based backbone on the cloud, compress and transmit before decoding on a local edge device by a lightweight model. Our HoloZip framework supports both both 2D/3D holograms as well as holographic videos, achieving a reconstruction quality of over 30 dB in PSNR (lossy compression standard) with bit rates as low as 0.6 bits per pixel as validated by experiments. Huaizhi Qu, Ruichen Zhang 0004, Hengyu Lian, Mufan Qiu, Samarjit Chakraborty, Henry Fuchs, Tianlong Chen 0001, Praneeth Chakravarthula |
ICCP | 7 |
| 2025 | Investigating Encoding and Perspective for Augmented Reality Motion GuidanceabstractAugmented reality (AR) offers promising opportunities to support movement-based activities, such as personal training or physical therapy, with real-time, spatially-situated visual cues. While many approaches leverage AR to guide motion, existing design guidelines focus on simple, upper-body movements within the user's field of view. We lack evidence-based design recommendations for guiding more diverse scenarios involving movements with varying levels of visibility and direction. We conducted an experiment to investigate how different visual encodings and perspectives affect motion guidance performance and usability, using three exercises that varied in visibility and planes of motion. Our findings reveal significant differences in preference and performance across designs. Notably, the best perspective varied depending on motion visibility and showing more information about the overall motion did not necessarily improve motion execution. We provide empirically-grounded guidelines for designing immersive, interactive visualizations for motion guidance to support more effective AR systems. Jade Kandel, Sriya Kasumarthi, Spiros Tsalikis, Chelsea Duppen, Daniel Szafir, Michael Lewek, Henry Fuchs, Danielle Albers Szafir |
ISMAR | 7 |
| 2025 | EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart GlassesabstractAll-day smart glasses are likely to emerge as platforms capable of continuous contextual sensing, uniquely positioning them for unprecedented assistance in our daily lives. Integrating the multi-modal AI agents required for human memory enhancement while performing continuous sensing, however, presents a major energy efficiency challenge for all-day usage. Achieving this balance requires intelligent, context-aware sensor management. Our approach, EgoTrigger, leverages audio cues from the microphone to selectively activate power-intensive cameras, enabling efficient sensing while preserving substantial utility for human memory enhancement. EgoTrigger uses a lightweight audio model (YAMNet) and a custom classification head to trigger image capture from hand-object interaction (HOI) audio cues, such as the sound of a drawer opening or a medication bottle being opened. In addition to evaluating on the QA-Ego4D dataset, we introduce and evaluate on the Human Memory Enhancement Question-Answer (HME-QA) dataset. Our dataset contains 340 human-annotated first-person QA pairs from full-length Ego4D videos that were curated to ensure that they contained audio, focusing on HOI moments critical for contextual understanding and memory. Our results show EgoTrigger can use 54% fewer frames on average, significantly saving energy in both power-hungry sensing components (e.g., cameras) and downstream operations (e.g., wireless transmission), while achieving comparable performance on datasets for an episodic memory task. We believe this context-aware triggering strategy represents a promising direction for enabling energy-efficient, functional smart glasses capable of all-day use - supporting applications like helping users recall where they placed their keys or information about their routine activities (e.g., taking medications). Akshay Paruchuri, Sinan Hersek, Lavisha Aggarwal, Xin Liu 0034, Achin Kulshrestha, Andrea Colaco, Henry Fuchs, Ishan Chatterjee |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | Learning View Synthesis for Desktop Telepresence With Few RGBD CamerasabstractRecent telepresence systems have shown significant improvements in quality compared to prior systems. However, they struggle to achieve both low cost and high quality at the same time. In this work, we envision a future where telepresence systems become a commodity and can be installed on typical desktops. To this end, we present a high-quality view synthesis method that uses a cost-effective capture system that consists of commodity hardware accessible to the general public. We propose a neural renderer that uses a few RGBD cameras as input to synthesize novel views of a user and their surroundings. At the core of the renderer is Multi-Layer Point Cloud (MPC), a novel 3D representation that improves reconstruction accuracy by removing non-linear biases in depth cameras. Our temporally-aware renderer further improves the stability of synthesized videos by conditioning on past information. Additionally, we propose Spatial Skip Connections (SSC) to improve image upsampling under limited GPU memory. Experimental results show that our renderer outperforms recent methods in terms of view synthesis quality. Our method generalizes to new users and challenging content (e.g. hand gestures and clothing deformation) without costly per-video optimization, object templates, or heavy pre-processing. The code and dataset will be made available. Shengze Wang 0002, Ryan Schmelzle, Liujie Zheng, Youngjoong Kwon, Roni Sengupta, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Dynamic Eyebox Steering for Improved Pinlight AR Near-Eye DisplaysabstractAn optical-see-through near-eye display (NED) for augmented reality (AR) allows the user to perceive virtual and real imagery simultaneously. Existing technologies for optical-see-through AR NEDs involve trade-offs between key metrics such as field of view (FOV), eyebox size, form factor, etc. We have enhanced an existing compact wide-FOV pinlight AR NED design with real-time 3D pupil localization in order to dynamically steer and thus effectively enlarge the usable eyebox. This is achieved with a dual-camera rig that captures stereoscopic views of the pupils. The 3D pupil location is used to dynamically calculate a display pattern that spatio-temporally modulates the light entering the wearer's eyes. We have built a demonstrable compact prototype and have conducted a user study that indicates the effectiveness of our eyebox steering method (e.g., without eyebox steering, in 10.5% of our tests, users were unable to perceive the test pattern correctly before experiment timeout; with eyebox steering, that fraction decreased dramatically to 1.25%). This is a small yet crucial step in making simple wide-FOV pinlight NEDs usable for human users and not just as demonstration prototypes filmed with a precisely positioned camera standing in for the user's eye. Further contributions of this paper include a detailed description of display design, calibration technique, and user study design, all of which may benefit other NED research. Xinxing Xia, Zheye Yu, Dongyu Qiu, Andrei State, Tat-Jen Cham, Frank Guan, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | PD-Insighter: A Visual Analytics System to Monitor Daily Actions for Parkinson's Disease TreatmentabstractPeople with Parkinson's Disease (PD) can slow the progression of their symptoms with physical therapy. However, clinicians lack insight into patients' motor function during daily life, preventing them from tailoring treatment protocols to patient needs. This paper introduces PD-Insighter, a system for comprehensive analysis of a person's daily movements for clinical review and decision-making. PD-Insighter provides an overview dashboard for discovering motor patterns and identifying critical deficits during activities of daily living and an immersive replay for closely studying the patient's body movements with environmental context. Developed using an iterative design study methodology in consultation with clinicians, we found that PD-Insighter's ability to aggregate and display data with respect to time, actions, and local environment enabled clinicians to assess a person's overall functioning during daily life outside the clinic. PD-Insighter's design offers future guidance for generalized multiperspective body motion analytics, which may significantly improve clinical decision-making and slow the functional decline of PD and other medical conditions. Jade Kandel, Chelsea Duppen, Qian Zhang 0066, Howard Jiang, Angelos Angelopoulos, Ashley Paula-Ann Neall, Pranav Wagh, Daniel Szafir, Henry Fuchs, Michael Lewek, Danielle Albers Szafir |
CHI | 9 |
| 2023 | Neural Image-based Avatars: Generalizable Radiance Fields for Human Avatar Modeling
Youngjoong Kwon, Dahun Kim, Duygu Ceylan, Henry Fuchs |
ICLR | 4 |
| 2023 | DELIFFAS: Deformable Light Fields for Fast Avatar SynthesisabstractGenerating controllable and photorealistic digital human avatars is a long-standing and important problem in Vision and Graphics. Recent methods have shown great progress in terms of either photorealism or inference speed while the combination of the two desired properties still remains unsolved. To this end, we propose a novel method, called DELIFFAS, which parameterizes the appearance of the human as a surface light field that is attached to a controllable and deforming human mesh model. At the core, we represent the light field around the human with a deformable two-surface parameterization, which enables fast and accurate inference of the human appearance. This allows perceptual supervision on the full image compared to previous approaches that could only supervise individual pixels or small patches due to their slow runtime. Our carefully designed human representation and supervision strategy leads to state-of-the-art synthesis results and inference time. The video results and code are available at https://vcai.mpi-inf.mpg.de/projects/DELIFFAS. Youngjoong Kwon, Lingjie Liu, Henry Fuchs, Marc Habermann, Christian Theobalt |
NeurIPS | 3 |
| 2022 | Geometry-Aware Eye Image-To-Image TranslationabstractRecently, image-to-image translation (I2I) has met with great success in computer vision, but few works have paid attention to the geometric changes that occur during translation. The geometric changes are necessary to reduce the geometric gap between domains at the cost of breaking correspondence between translated images and original ground truth. We propose a novel geometry-aware semi-supervised method to preserve this correspondence while still allowing geometric changes. The proposed method takes a synthetic image-mask pair as input and produces a corresponding real pair. We also utilize an objective function to ensure consistent geometric movement of the image and mask through the translation. Extensive experiments illustrate that our method yields a 11.23% higher mean Intersection-Over-Union than the current methods on the downstream eye segmentation task. The generated image has a 15.9% decrease in Frechet Inception Distance indicating higher image quality. Conny Lu, Qian Zhang 0066, Kapil Krishnakumar, Jixu Chen, Henry Fuchs, Sachin S. Talathi |
ETRA | 5 |
| 2022 | Keynote SpeakersabstractHenry Fuchs is the Federico Gil Distinguished Professor of Computer Science and Adjunct Professor of Biomedical Engineering at UNC Chapel Hill. He has been active in computer graphics since the early 1970s, with rendering algorithms (BSP Trees), hardware (Pixel-Planes and PixelFlow), virtual environments, tele-immersion systems and medical applications. He received a Ph.D. in 1975 from the University of Utah. From 1975 to 1978 he was an assistant professor at the University of Texas at Dallas. Since 1978, he’s been on the faculty at UNC Chapel Hill. Henry Fuchs |
ISMAR | 1 |
| 2022 | Neural 3D Gaze: 3D Pupil Localization and Gaze Tracking based on Anatomical Eye Model and Neural Refraction CorrectionabstractEye tracking has already made its way to current commercial wearable display devices, and is becoming increasingly important for virtual and augmented reality applications. However, the existing model-based eye tracking solutions are not capable of conducting very accurate gaze angle measurements, and may not be sufficient to solve challenging display problems such as pupil steering or eyebox expansion. In this paper, we argue that accurate detection and localization of pupil in 3D space is a necessary intermediate step in model-based eye tracking. Existing methods and datasets either ignore evaluating the accuracy of 3D pupil localization or evaluate it only on synthetic data. To this end, we capture the first 3D pupilgaze-measurement dataset using a high precision setup with head stabilization and release it as the first benchmark dataset to evaluate both 3D pupil localization and gaze tracking methods. Furthermore, we utilize an advanced eye model to replace the commonly used oversimplified eye model. Leveraging the eye model, we propose a novel 3D pupil localization method with a deep learning-based corneal refraction correction. We demonstrate that our method outperforms the state-of-the-art works by reducing the 3D pupil localization error by 47.5% and the gaze estimation error by 18.7%. Our dataset and codes can be found here: link. Conny Lu, Praneeth Chakravarthula, Kaihao Liu, Xixiang Liu, Henry Fuchs |
ISMAR | 6 |
| 2022 | Tailor Me: An Editing Network for Fashion Attribute Shape ManipulationabstractFashion attribute editing aims to manipulate fashion images based on a user-specified attribute, while preserving the details of the original image as intact as possible. Recent works in this domain have mainly focused on direct manipulation of the raw RGB pixels, which only allows to perform edits involving relatively small shape changes (e.g., sleeves). The goal of our Virtual Personal Tailoring Network (VPTNet) is to extend the editing capabilities to much larger shape changes of fashion items, such as cloth length. To achieve this goal, we decouple the fashion attribute editing task into two conditional stages: shape-then-appearance editing. To this aim, we propose a shape editing network that employs a semantic parsing of the fashion image as an interface for manipulation. Compared to operating on the raw RGB image, our parsing map editing enables performing more complex shape editing operations. Second, we introduce an appearance completion network that takes the previous stage results and completes the shape difference regions to produce the final RGB image. Qualitative and quantitative experiments on the DeepFashion-Synthesis dataset confirm that VPTNet outperforms state-of-the-art methods for both small and large shape attribute editing. Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Viswanathan (Vishy) Swaminathan, Henry Fuchs |
WACV | 6 |
| 2022 | Hogel-Free HolographyabstractHolography is a promising avenue for high-quality displays without requiring bulky, complex optical systems. While recent work has demonstrated accurate hologram generation of 2D scenes, high-quality holographic projections of 3D scenes has been out of reach until now. Existing multiplane 3D holography approaches fail to model wavefronts in the presence of partial occlusion while holographic stereogram methods have to make a fundamental tradeoff between spatial and angular resolution. In addition, existing 3D holographic display methods rely on heuristic encoding of complex amplitude into phase-only pixels which results in holograms with severe artifacts. Fundamental limitations of the input representation, wavefront modeling, and optimization methods prohibit artifact-free 3D holographic projections in today’s displays. To lift these limitations, we introduce hogel-free holography which optimizes for true 3D holograms, supporting both depth- and view-dependent effects for the first time. Our approach overcomes the fundamental spatio-angular resolution tradeoff typical to stereogram approaches. Moreover, it avoids heuristic encoding schemes to achieve high image fidelity over a 3D volume. We validate that the proposed method achieves 10 dB PSNR improvement on simulated holographic reconstructions. We also validate our approach on an experimental prototype with accurate parallax and depth focus effects. Praneeth Chakravarthula, Ethan Tseng, Henry Fuchs, Felix Heide |
ACM Trans. Graph. | 3 |
| 2022 | VGTC Virtual Reality Awards Program Chair MessageabstractThe IEEE VGTC Virtual Reality Awards program was expanded this year with additional areas and with the establishment of the Virtual Reality Academy, the biggest expansion since the VR Awards started in 2005: Henry Fuchs, Federico Gil |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance RenderingabstractIn this paper, we aim at synthesizing a free-viewpoint video of an arbitrary human performance using sparse multi-view cameras. Recently, several works have addressed this problem by learning person-specific neural radiance fields (NeRF) to capture the appearance of a particular human. In parallel, some work proposed to use pixel-aligned features to generalize radiance fields to arbitrary new scenes and objects. Adopting such generalization approaches to humans, however, is highly challenging due to the heavy occlusions and dynamic articulations of body parts. To tackle this, we propose Neural Human Performer, a novel approach that learns generalizable neural radiance fields based on a parametric human body model for robust performance capture. Specifically, we first introduce a temporal transformer that aggregates tracked visual features based on the skeletal body motion over time. Moreover, a multi-view transformer is proposed to perform cross-attention between the temporally-fused features and the pixel-aligned features at each time step to integrate observations on the fly from multiple views. Experiments on the ZJU-MoCap and AIST datasets show that our method significantly outperforms recent generalizable NeRF methods on unseen identities and poses. Youngjoong Kwon, Dahun Kim, Duygu Ceylan, Henry Fuchs |
NeurIPS | 4 |
| 2021 | Mobile. Egocentric Human Body Motion Reconstruction Using Only Eyeglasses-mounted Cameras and a Few Body-worn Inertial SensorsabstractWe envision a convenient telepresence system available to users anywhere, anytime. Such a system requires displays and sensors embedded in commonly worn items such as eyeglasses, wristwatches, and shoes. To that end, we present a standalone real-time system for the dynamic 3D capture of a person, relying only on cameras embedded into a head-worn device, and on Inertial Measurement Units (IMUs) worn on the wrists and ankles. Our prototype system egocentrically reconstructs the wearer's motion via learning-based pose estimation, which fuses inputs from visual and inertial sensors that complement each other, overcoming challenges such as inconsistent limb visibility in head-worn views, as well as pose ambiguity from sparse IMUs. The estimated pose is continuously re-targeted to a prescanned surface model, resulting in a high-fidelity 3D reconstruction. We demonstrate our system by reconstructing various human body movements and show that our visual-inertial learning-based method, which runs in real time, outperforms both visual-only and inertial-only approaches. We captured an egocentric visual-inertial 3D human pose dataset publicly available at https://sites.google.com/site/youngwooncha/egovip for training and evaluating similar methods. Young-Woon Cha, Husam Shaik, Qian Zhang 0066, Andrei State, Adrian Ilie, Henry Fuchs |
VR | 7 |
| 2021 | Gaze-Contingent Retinal Speckle Suppression for Perceptually-Matched Foveated Holographic DisplaysabstractComputer-generated holographic (CGH) displays show great potential and are emerging as the next-generation displays for augmented and virtual reality, and automotive heads-up displays. One of the critical problems harming the wide adoption of such displays is the presence of speckle noise inherent to holography, that compromises its quality by introducing perceptible artifacts. Although speckle noise suppression has been an active research area, the previous works have not considered the perceptual characteristics of the Human Visual System (HVS), which receives the final displayed imagery. However, it is well studied that the sensitivity of the HVS is not uniform across the visual field, which has led to gaze-contingent rendering schemes for maximizing the perceptual quality in various computer-generated imagery. Inspired by this, we present the first method that reduces the "perceived speckle noise" by integrating foveal and peripheral vision characteristics of the HVS, along with the retinal point spread function, into the phase hologram computation. Specifically, we introduce the anatomical and statistical retinal receptor distribution into our computational hologram optimization, which places a higher priority on reducing the perceived foveal speckle noise while being adaptable to any individual's optical aberration on the retina. Our method demonstrates superior perceptual quality on our emulated holographic display. Our evaluations with objective measurements and subjective studies demonstrate a significant reduction of the human perceived noise. Praneeth Chakravarthula, Zhan Zhang 0009, Okan Tarhan Tursun, Piotr Didyk, Qi Sun 0003, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video
Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Eunbyung Park, Viswanathan (Vishy) Swaminathan, Henry Fuchs |
ECCV (4) | 7 |
| 2020 | Stimulating the Human Visual System Beyond Real World Performance in Future Augmented Reality DisplaysabstractNew augmented-reality near-eye displays provide capabilities for enriching real-world visual experiences with digital content. Most current research focuses on improving both hardware and software to provide digital content that seamlessly blends with the real world. This is believed to not only contribute to the visual experience but also increase human task performance. In this work, we take a step further and ask the question of whether the capabilities of current and future display designs combined with efficient perception-inspired content optimizations can be used to improve human task performance beyond the human capabilities in the natural world. Based on an in-depth analysis of previous literature, we hypothesize here that such enhancements can be achieved when the human visual system is provided with content that optimizes the oculomotor responses. To further investigate possible gains, we present a series of perceptual experiments that built upon this idea. More specifically, we focus on speeding up accommodation response, which significantly contributes to the eye-adaptation when a new stimulus is shown. Through our experiments, we demonstrate that such speedups canbe achieved, and more importantly, they can lead to significant improvements in human task performance. While not all of our results give definite answers, we believe that they reveal plentiful opportunities for further enhancing the human experience and task performance when using new augmented-reality displays. David Dunn, Okan Tarhan Tursun, Hyeonseung Yu, Piotr Didyk, Karol Myszkowski, Henry Fuchs |
ISMAR | 6 |
| 2020 | Improved vergence and accommodation via Purkinje Image tracking with multiple cameras for AR glassesabstractWe present a personalized, comprehensive eye-tracking solution based on tracking higher-order Purkinje images, suited explicitly for eyeglasses-style AR and VR displays. Existing eye-tracking systems for near-eye applications are typically designed to work for an on-axis configuration and rely on pupil center and corneal reflections (PCCR) to estimate gaze with an accuracy of only about 0.5°to 1°. These are often expensive, bulky in form factor, and fail to estimate monocular accommodation, which is crucial for focus adjustment within the AR glasses.Our system independently measures the binocular vergence and monocular accommodation using higher-order Purkinje reflections from the eye, extending the PCCR based methods. We demonstrate that these reflections are sensitive to both gaze rotation and lens accommodation and model the Purkinje images' behavior in simulation. We also design and fabricate a user-customized eye tracker using cheap off-the-shelf cameras and LEDs. We use an end-to-end convolutional neural network (CNN) for calibrating the eye tracker for the individual user, allowing for robust and simultaneous estimation of vergence and accommodation. Experimental results show that our solution, specifically catering to individual users, outperforms state-of-the-art methods for vergence and depth estimation, achieving an accuracy of 0.3782° and 1.108cm respectively. Conny Lu, Praneeth Chakravarthula, Yujie Tao, Steven Chen, Henry Fuchs |
ISMAR | 5 |
| 2020 | Towards Eyeglass-style Holographic Near-eye Displays with StaticallyabstractHolography is perhaps the only method demonstrated so far that can achieve a wide field of view (FOV) and a compact eyeglass-style form factor for augmented reality (AR) near-eye displays (NEDs). Unfortunately, the eyebox of such NEDs is impractically small (~ <; 1mm). In this paper, we introduce and demonstrate a design for holographic NEDs with a practical, wide eyebox of ~ 10mm and without any moving parts, based on holographic lenslets. In our design, a holographic optical element (HOE) based on a lenslet array was fabricated as the image combiner with expanded eyebox. A phase spatial light modulator (SLM) alters the phase of the incident laser light projected onto the HOE combiner such that the virtual image can be perceived at different focus distances, which can reduce the vergence-accommodation conflict (VAC). We have successfully implemented a bench-top prototype following the proposed design. The experimental results show effective eyebox expansion to a size of ~ 10mm. With further work, we hope that these design concepts can be incorporated into eyeglass-size NEDs. Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Tat-Jen Cham, Henry Fuchs |
ISMAR | 6 |
| 2020 | Rotationally-Consistent Novel View Synthesis for HumansabstractHuman novel view synthesis aims to synthesize target views of a human subject given input images taken from one or more reference viewpoints. Despite significant advances in model-free novel view synthesis, existing methods present two major limitations when applied to complex shapes like humans. First, these methods mainly focus on simple and symmetric objects, e.g., cars and chairs, limiting their performances to fine-grained and asymmetric shapes. Second, existing methods cannot guarantee visual consistency across different adjacent views of the same object. To solve these problems, we present in this paper a learning framework for the novel view synthesis of human subjects, which explicitly enforces consistency across different generated views of the subject. Specifically, we introduce a novel multi-view supervision and an explicit rotational loss during the learning process, enabling the model to preserve detailed body parts and to achieve consistency between adjacent synthesized views. To show the superior performance of our approach, we present qualitative and quantitative results on the Multi-View Human Action (MVHA) dataset we collected (consisting of 3D human models animated with different Mocap sequences and captured from 54 different viewpoints), the Pose-Varying Human Model (PVHM) dataset, and ShapeNet. The qualitative and quantitative results demonstrate that our approach outperforms the state-of-the-art baselines in both per-view synthesis quality, and in preserving rotational consistency and complex shapes (e.g. fine-grained details, challenging poses) across multiple adjacent views in a variety of scenarios, for both humans and rigid objects. Youngjoong Kwon, Stefano Petrangeli, Dahun Kim, Henry Fuchs, Viswanathan (Vishy) Swaminathan |
ACM Multimedia | 5 |
| 2020 | Learned hardware-in-the-loop phase retrieval for holographic near-eye displaysabstractHolography is arguably the most promising technology to provide wide field-of-view compact eyeglasses-style near-eye displays for augmented and virtual reality. However, the image quality of existing holographic displays is far from that of current generation conventional displays, effectively making today's holographic display systems impractical. This gap stems predominantly from the severe deviations in the idealized approximations of the "unknown" light transport model in a real holographic display, used for computing holograms. In this work, we depart from such approximate "ideal" coherent light transport models for computing holograms. Instead, we learn the deviations of the real display from the ideal light transport from the images measured using a display-camera hardware system. After this unknown light propagation is learned, we use it to compensate for severe aberrations in real holographic imagery. The proposed hardware-in-the-loop approach is robust to spatial, temporal and hardware deviations, and improves the image quality of existing methods qualitatively and quantitatively in SNR and perceptual quality. We validate our approach on a holographic display prototype and show that the method can fully compensate unknown aberrations and erroneous and non-linear SLM phase delays, without explicitly modeling them. As a result, the proposed method significantly outperforms existing state-of-the-art methods in simulation and experimentation - just by observing captured holographic images. Praneeth Chakravarthula, Ethan Tseng, Tarun Srivastava, Henry Fuchs, Felix Heide |
ACM Trans. Graph. | 4 |
| 2019 | StereoDRNet: Dilated Residual StereoNetabstractWe propose a system that uses a convolution neural network (CNN) to estimate depth from a stereo pair followed by volumetric fusion of the predicted depth maps to produce a 3D reconstruction of a scene. Our proposed depth refinement architecture, predicts view-consistent disparity and occlusion maps that helps the fusion system to produce geometrically consistent reconstructions. We utilize 3D dilated convolutions in our proposed cost filtering network that yields better filtering while almost halving the computational cost in comparison to state of the art cost filtering architectures. For feature extraction we use the Vortex Pooling architecture. The proposed method achieves state of the art results in KITTI 2012, KITTI 2015 and ETH 3D stereo benchmarks. Finally, we demonstrate that our system is able to produce high fidelity 3D scene reconstructions that outperforms the state of the art stereo system. Rohan Chabra, Julian Straub, Chris Sweeney, Richard A. Newcombe, Henry Fuchs |
CVPR | 5 |
| 2019 | Wirtinger holography for near-eye displaysabstractNear-eye displays using holographic projection are emerging as an exciting display approach for virtual and augmented reality at high-resolution without complex optical setups --- shifting optical complexity to computation. While precise phase modulation hardware is becoming available, phase retrieval algorithms are still in their infancy, and holographic display approaches resort to heuristic encoding methods or iterative methods relying on various relaxations. In this work, we depart from such existing approximations and solve the phase retrieval problem for a hologram of a scene at a single depth at a given time by revisiting complex Wirtinger derivatives, also extending our framework to render 3D volumetric scenes. Using Wirtinger derivatives allows us to pose the phase retrieval problem as a quadratic problem which can be minimized with first-order optimization methods. The proposed Wirtinger Holography is flexible and facilitates the use of different loss functions, including learned perceptual losses parametrized by deep neural networks, as well as stochastic optimization methods. We validate this framework by demonstrating holographic reconstructions with an order of magnitude lower error, both in simulation and on an experimental hardware prototype. Praneeth Chakravarthula, Yifan Peng 0001, Joel S. Kollin, Henry Fuchs, Felix Heide |
ACM Trans. Graph. | 4 |
| 2019 | Manufacturing Application-Driven Foveated Near-Eye DisplaysabstractTraditional optical manufacturing poses a great challenge to near-eye display designers due to large lead times in the order of multiple weeks, limiting the abilities of optical designers to iterate fast and explore beyond conventional designs. We present a complete near-eye display manufacturing pipeline with a day lead time using commodity hardware. Our novel manufacturing pipeline consists of several innovations including a rapid production technique to improve surface of a 3D printed component to optical quality suitable for near-eye display application, a computational design methodology using machine learning and ray tracing to create freeform static projection screen surfaces for near-eye displays that can represent arbitrary focal surfaces, and a custom projection lens design that distributes pixels non-uniformly for a foveated near-eye display hardware design candidate. We have demonstrated untethered augmented reality near-eye display prototypes to assess success of our technique, and show that a ski-goggles form factor, a large monocular field of view$(30^{o}\times 55^{o})$, and a resolution of 12 cycles per degree can be achieved. Kaan Aksit, Praneeth Chakravarthula, Kishore Rathinavel, Youngmo Jeong, Rachel A. Albert, Henry Fuchs, David P. Luebke |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | Implementation and Evaluation of a 50 kHz, $28μs Motion-to-Pose Latency Head Tracking InstrumentabstractThis paper presents the implementation and evaluation of a 50,000-pose-sample-per-second, 6-degree-of-freedom optical head tracking instrument with motion-to-pose latency of 28μs and dynamic precision of 1-2 arcminutes. The instrument uses high-intensity infrared emitters and two duo-lateral photodiode-based optical sensors to triangulate pose. This instrument serves two purposes: it is the first step towards the requisite head tracking component in sub- 100μs motion-to-photon latency optical see-through augmented reality (OST AR) head-mounted display (HMD) systems; and it enables new avenues of research into human visual perception - including measuring the thresholds for perceptible real-virtual displacement during head rotation and other human research requiring high-sample-rate motion tracking. The instrument's tracking volume is limited to about 120×120×250 but allows for the full range of natural head rotation and is sufficient for research involving seated users. We discuss how the instrument's tracking volume is scalable in multiple ways and some of the trade-offs involved therein. Finally, we introduce a novel laser-pointer-based measurement technique for assessing the instrument's tracking latency and repeatability. We show that the instrument's motion-to-pose latency is 28μs and that it is repeatable within 1-2 arcminutes at mean rotational velocities (yaw) in excess of 500°/sec. Alex Blate, Mary C. Whitton, Montek Singh, Greg Welch, Andrei State, Turner Whitted, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2019 | Varifocal Occlusion-Capable Optical See-through Augmented Reality Display based on Focus-tunable OpticsabstractOptical see-through augmented reality (AR) systems are a next-generation computing platform that offer unprecedented user experiences by seamlessly combining physical and digital content. Many of the traditional challenges of these displays have been significantly improved over the last few years, but AR experiences offered by today's systems are far from seamless and perceptually realistic. Mutually consistent occlusions between physical and digital objects are typically not supported. When mutual occlusion is supported, it is only supported for a fixed depth. We propose a new optical see-through AR display system that renders mutual occlusion in a depth-dependent, perceptually realistic manner. To this end, we introduce varifocal occlusion displays based on focus-tunable optics, which comprise a varifocal lens system and spatial light modulators that enable depth-corrected hard-edge occlusions for AR experiences. We derive formal optimization methods and closed-form solutions for driving this tunable lens system and demonstrate a monocular varifocal occlusion-capable optical see-through AR display capable of perceptually realistic occlusion across a large depth range. Kishore Rathinavel, Gordon Wetzstein, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2019 | Towards a Switchable AR/VR Near-eye Display with Accommodation-Vergence and Eyeglass Prescription SupportabstractIn this paper, we present our novel design for switchable AR/VR near-eye displays which can help solve the vergence-accommodation-conflict issue. The principal idea is to time-multiplex virtual imagery and real-world imagery and use a tunable lens to adjust focus for the virtual display and the see-through scene separately. With this novel design, prescription eyeglasses for near- and far-sighted users become unnecessary. This is achieved by integrating the wearer's corrective optical prescription into the tunable lens for both virtual display and see-through environment. We built a prototype based on the design, comprised of micro-display, optical systems, a tunable lens, and active shutters. The experimental results confirm that the proposed near-eye display design can switch between AR and VR and can provide correct accommodation for both. Xinxing Xia, Frank Guan, Andrei State, Praneeth Chakravarthula, Kishore Rathinavel, Tat-Jen Cham, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | Towards Efficient 3D Calibration for Different Types of Multi-view Autostereoscopic 3D DisplaysabstractA novel and efficient 3D calibration method for different types of autostereoscopic multi-view 3D displays is presented in this paper. In our method, a camera is placed at different locations within the viewing volume of a 3D display to capture a series of images that relate to the subset of light rays emitted by the 3D display and arriving at each of the camera positions. Gray code patterns modulate the images shown on the 3D display, helping to significantly reduce the number of images captured by the camera and thereby accelerate the process of calculating the correspondence relationship between the pixels on the 3D display and the locations of the capturing camera. The proposed calibration method has been successfully tested on two different types of multi-view 3D displays and can be easily generalized for calibrating other types of such displays. The experimental results show that this novel 3D calibration method can also be used to improve the image quality by reducing the frequently observed crosstalk that typically exists when multiple users are simultaneously viewing multi-view 3D displays from a range of viewing positions. Xinxing Xia, Frank Guan, Andrei State, Tat-Jen Cham, Henry Fuchs |
CGI | 5 |
| 2018 | Real-time 3D Face-Eye Performance Capture of a Person Wearing VR HeadsetabstractTeleconference or telepresence based on virtual reality (VR) head-mount display (HMD) device is a very interesting and promising application since HMD can provide immersive feelings for users. However, in order to facilitate face-to-face communications for HMD users, real-time 3D facial performance capture of a person wearing HMD is needed, which is a very challenging task due to the large occlusion caused by HMD. The existing limited solutions are very complex either in setting or in approach as well as lacking the performance capture of 3D eye gaze movement. In this paper, we propose a convolutional neural network (CNN) based solution for real-time 3D face-eye performance capture of HMD users without complex modification to devices. To address the issue of lacking training data, we generate massive pairs of HMD face-label dataset by data synthesis as well as collecting VR-IR eye dataset from multiple subjects. Then, we train a dense-fitting network for facial region and an eye gaze network to regress 3D eye model parameters. Extensive experimental results demonstrate that our system can efficiently and effectively produce in real time a vivid personalized 3D avatar with the correct identity, pose, expression and eye motion corresponding to the HMD user. Guoxian Song, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng, Juyong Zhang, Henry Fuchs |
ACM Multimedia | 6 |
| 2018 | Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn CamerasabstractWe propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation. Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 15 |
| 2018 | FocusAR: Auto-focus Augmented Reality Eyeglasses for both Real World and Virtual ImageryabstractWe describe a system which corrects dynamically for the focus of the real world surrounding the near-eye display of the user and simultaneously the internal display for augmented synthetic imagery, with an aim of completely replacing the user prescription eyeglasses. The ability to adjust focus for both real and virtual stimuli will be useful for a wide variety of users, but especially for users over 40 years of age who have limited accommodation range. Our proposed solution employs a tunable-focus lens for dynamic prescription vision correction, and a varifocal internal display for setting the virtual imagery at appropriate spatially registered depths. We also demonstrate a proof of concept prototype to verify our design and discuss the challenges to building an auto-focus augmented reality eyeglasses for both real and virtual. Praneeth Chakravarthula, David Dunn, Kaan Aksit, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | An Extended Depth-at-Field Volumetric Near-Eye Augmented Reality DisplayabstractWe introduce an optical design and a rendering pipeline for a full-color volumetric near-eye display which simultaneously presents imagery with near-accurate per-pixel focus across an extended volume ranging from 15cm (6.7 diopters) to 4M (0.25 diopters), allowing the viewer to accommodate freely across this entire depth range. This is achieved using a focus-tunable lens that continuously sweeps a sequence of 280 synchronized binary images from a high-speed, Digital Micromirror Device (DMD) projector and a high-speed, high dynamic range (HDR) light source that illuminates the DMD images with a distinct color and brightness at each binary frame. Our rendering pipeline converts 3-D scene information into a 2-D surface of color voxels, which are decomposed into 280 binary images in a voxel-oriented manner, such that 280 distinct depth positions for full-color voxels can be displayed. Kishore Rathinavel, Hanpeng Wang, Alex Blate, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | A VR-based user study on the effects of vision impairments on recognition distances of escape-route signs in buildingsabstractIn workplaces or publicly accessible buildings, escape routes are signposted according to official norms or international standards that specify distances, angles and areas of interest for the positioning of escape-route signs. In homes for the elderly, in which the residents commonly have degraded mobility and suffer from vision impairments caused by age or eye diseases, the specifications of current norms and standards may be insufficient. Quantifying the effect of symptoms of vision impairments like reduced visual acuity on recognition distances is challenging, as it is cumbersome to find a large number of user study participants who suffer from exactly the same form of vision impairments. Hence, we propose a new methodology for such user studies: By conducting a user study in virtual reality (VR), we are able to use participants with normal or corrected sight and simulate vision impairments graphically. The use of standardized medical eyesight tests in VR allows us to calibrate the visual acuity of all our participants to the same level, taking their respective visual acuity into account. Since we primarily focus on homes for the elderly, we accounted for their often limited mobility by implementing a wheelchair simulation for our VR application. Katharina Krösl, Dominik Bauer, Michael Schwärzler, Henry Fuchs, Georg Suter, Michael Wimmer 0001 |
Vis. Comput. | 4 |
| 2017 | Supporting free walking in a large virtual environment: imperceptible redirected walking with an immersive distractorabstractRedirected walking, a technique in which the user's orientation in the physical space is constantly and imperceptibly changed from their orientation in the virtual world, has been shown to be an effective technique when only a limited physical space is available. Unfortunately, previous efforts have restricted redirected walking applications to operate under the constant supervision of researchers to prevent the users from leaving the tracked area. In addition, Virtual Environments (VE) used in these applications were often limited to narrow hallways, mazes or predefined waypoints, while the performance of redirected walking in a large, open VE is not well explored. In this paper, we introduce the idea and implementation for an imperceptible redirected walking system that supports the illusion of free walking in a large, open virtual environment with minimal amount of physical interventions, by integrating the distractor into the user's main immersive activity in the VE. We demonstrate this new approach with two user studies of an immersive interactive game. Our study indicates that for the majority of the subjects, the illusion is maintained of unconstrained walking in a very large area (a full-size basketball court, 50 feet × 95 feet), even while they were limited to a physical area of a mere 6% of the size of the basketball court (16 feet × 16 feet tracked area). Our result demonstrates that the illusion of free walking is created, since a majority of them was not interrupted by the researchers and did not realize they were redirected, and the 34 subjects took vastly different routes to reach the distant goal (See Figure 2). We believe that this technique demonstrates a more immersive way of designing redirected walking application and shows possibility of bringing redirected walking applications out of the monitored lab environments. Our result may provide insights for the designers of immersive experiences to create other redirected walking applications for consumer VR systems with room-size tracking. Haiwei Chen, Henry Fuchs |
CGI | 2 |
| 2017 | Towards imperceptible redirected walking: integrating a distractor into the immersive experienceabstractPhysically walking in a virtual world has been repeatedly demonstrated to be superior to navigation with game controllers and the like. The problem is that virtual worlds are often much larger than the available physical space. Redirected walking, a technique in which the users orientation in the physical space is constantly and imperceptibly changed from their orientation in the virtual world, has been shown to be an effective technique for walking in a limited physical space [Razzaque et al. 2001], but it needs frequent and rapid head rotation to keep the reorientation changes imperceptible. A rapidly moving object in the environment, a distractor, has been shown to be effective at inspiring the users to rapidly rotate their head, but such a distractor can be distracting to the user's main activity in the virtual world. We introduce the notion of the distractor being integrated into the user's main activity in the virtual world - in our example, a fire-breathing dragon into an immersive adventure game. With such an integrated character, the users do not need any special instructions about redirected walking; they are simply performing their intended activities. We report the results of a small (N=24) user study which indicates that for the majority of subjects (17 of 24) the illusion is maintained of unconstrained walking in a very large area (a full-sized basketball court, 45 × 90 feet) even while they were limited to a small, 16 × 16 feet region. We speculate that this technique may extend to many other applications in which the distractor can be integrated into the major immersive activity and thus enable the illusion of natural, unconstrained walking in large virtual worlds even when only a small physical space is available. Haiwei Chen, Henry Fuchs |
I3D | 2 |
| 2017 | Scene-adaptive high dynamic range display for low latency augmented realityabstractFor generalized augmented reality to be feasible, the augmenting elements must be visible in varied environments and under rapidly changing, high dynamic range lighting, from bright sunlight to deep shadows. We present a high dynamic range, optical see-through, augmented reality display that dynamically adjusts the brightness of the virtual imagery to match the current brightness of the real scene. Critical components include the spatial brightness sensor array and the positional brightness image intensity matcher. The color, scene-adaptive HDR display system is based on a high-rate (15 kHz) DMD projector using a high-speed RGB LED illuminator, each color with independent 16 bit intensity control for each binary DMD frame. The critical input to the intensity matching algorithm is the output of an array of high sensitivity light sensors. This paper discusses the implementation of the system and reports performance via still and video demonstrations under a variety of lighting conditions. Peter Lincoln, Alex Blate, Montek Singh, Andrei State, Mary C. Whitton, Turner Whitted, Henry Fuchs |
I3D | 7 |
| 2017 | Optimizing placement of commodity depth cameras for known 3D dynamic scene captureabstractCommodity depth cameras, such as the Microsoft Kinect®, have been widely used for the capture and reconstruction of the 3D structure of room-sized dynamic scenes. Camera placement and coverage during capture significantly impact the quality of the resulting reconstruction. In particular, dynamic occlusions and sensor interference have been shown to result in poor resolution and holes in the reconstruction results. This paper presents a novel algorithmic framework and a method for off-line optimization of depth cameras placements for a given 3D dynamic scene, simulated using virtual 3D models. We derive a fitness metric for a particular configuration of sensors by combining factors such as visibility and resolution of the entire dynamic scene with probabilities of interference between sensors. We employ this fitness metric both in a greedy algorithm that determines the number of depth cameras needed to cover the scene, and in a simulated annealing algorithm that optimizes the placements of those sensors. We compare our algorithm's optimized placements with manual sensor placements for a real dynamic scene. We present quantitative assessments using our fitness metric, as well as qualitative assessments to demonstrate that our algorithm not only enhances the resolution and total coverage of the reconstruction, but also fills in voids by avoiding occlusions and sensor interference when compared with the reconstruction of the same scene using mual sensor placement. Rohan Chabra, Adrian Ilie, Nicholas Rewkowski, Young-Woon Cha, Henry Fuchs |
VR | 5 |
| 2017 | Real-world occlusion in optical see-through AR displaysabstractIn this work we describe a system composed by an optical see-through AR headset---a Microsoft HoloLens---, stereo projectors and shutter glasses. Projectors are used to add to the device the capability of occluding real-world surfaces to make the virtual objects to appear more solid and less transparent. Giovanni Avveduto, Franco Tecchia, Henry Fuchs |
VRST | 3 |
| 2017 | Wide Field Of View Varifocal Near-Eye Display Using See-Through Deformable Membrane MirrorsabstractAccommodative depth cues, a wide field of view, and ever-higher resolutions all present major hardware design challenges for near-eye displays. Optimizing a design to overcome one of these challenges typically leads to a trade-off in the others. We tackle this problem by introducing an all-in-one solution - a new wide field of view, gaze-tracked near-eye display for augmented reality applications. The key component of our solution is the use of a single see-through, varifocal deformable membrane mirror for each eye reflecting a display. They are controlled by airtight cavities and change the effective focal power to present a virtual image at a target depth plane which is determined by the gaze tracker. The benefits of using the membranes include wide field of view (100° diagonal) and fast depth switching (from 20 cm to infinity within 300 ms). Our subjective experiment verifies the prototype and demonstrates its potential benefits for near-eye see-through displays. David Dunn, Cary Tippets, Kent Torell, Petr Kellnhofer, Kaan Aksit, Piotr Didyk, Karol Myszkowski, David P. Luebke, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2016 | Faster feedback for remote scene viewing with pan-tilt stereo cameraabstractWe demonstrate a remote scene viewing system for telepresence purposes. The system is based on a pan-tilt stereo camera that captures stereo video and transfers it to a remote user over a network. On the user end, the live stereo video is processed and displayed in a Head-Mounted Display. Faster feedback can be achieved through latency compensation. Using a wider field-of-view, higher resolution camera, the appropriate subset of the image is selected and displayed. We introduce the hardware configuration and software framework for the system and a method to calculate the homography between the camera image space and user head image space. Our perceived latency of the system is estimated to be 50-100ms. Henry Fuchs |
VR | 2 |
| 2016 | From Motion to Photons in 80 Microseconds: Towards Minimal Latency for Virtual and Augmented RealityabstractWe describe an augmented reality, optical see-through display based on a DMD chip with an extremely fast (16 kHz) binary update rate. We combine the techniques of post-rendering 2-D offsets and just-in-time tracking updates with a novel modulation technique for turning binary pixels into perceived gray scale. These processing elements, implemented in an FPGA, are physically mounted along with the optical display elements in a head tracked rig through which users view synthetic imagery superimposed on their real environment. The combination of mechanical tracking at near-zero latency with reconfigurable display processing has given us a measured average of 80 µs of end-to-end latency (from head motion to change in photons from the display) and also a versatile test platform for extremely-low-latency display systems. We have used it to examine the trade-offs between image quality and cost (i.e. power and logical complexity) and have found that quality can be maintained with a fairly simple display modulation scheme. Peter Lincoln, Alex Blate, Montek Singh, Turner Whitted, Andrei State, Anselmo Lastra, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2015 | 3D scanning deformable objects with a single RGBD sensorabstractWe present a 3D scanning system for deformable objects that uses only a single Kinect sensor. Our work allows considerable amount of nonrigid deformations during scanning, and achieves high quality results without heavily constraining user or camera motion. We do not rely on any prior shape knowledge, enabling general object scanning with freeform deformations. To deal with the drift problem when nonrigidly aligning the input sequence, we automatically detect loop closures, distribute the alignment error over the loop, and finally use a bundle adjustment algorithm to optimize for the latent 3D shape and nonrigid deformation parameters simultaneously. We demonstrate high quality scanning results in some challenging sequences, comparing with state of art nonrigid techniques, as well as ground truth data. Mingsong Dou, Jonathan Taylor 0001, Henry Fuchs, Andrew W. Fitzgibbon, Shahram Izadi |
CVPR | 3 |
| 2014 | Minimizing latency for augmented reality displays: Frames considered harmfulabstractWe present initial results from a new image generation approach for low-latency displays such as those needed in head-worn AR devices. Avoiding the usual video interfaces, such as HDMI, we favor direct control of the internal display technology. We illustrate our new approach with a bench-top optical see-through AR proof-of-concept prototype that uses a Digital Light Processing (DLPTM) projector whose Digital Micromirror Device (DMD) imaging chip is directly controlled by a computer, similar to the way random access memory is controlled. We show that a perceptually-continuous-tone dynamic gray-scale image can be efficiently composed from a very rapid succession of binary (partial) images, each calculated from the continuous-tone image generated with the most recent tracking data. As the DMD projects only a binary image at any moment, it cannot instantly display this latest continuous-tone image, and conventional decomposition of a continuous-tone image into binary time-division-multiplexed values would induce just the latency we seek to avoid. Instead, our approach maintains an estimate of the image the user currently perceives, and at every opportunity allowed by the control circuitry, sets each binary DMD pixel to the value that will reduce the difference between that user-perceived image and the newly generated image from the latest tracking data. The resulting displayed binary image is “neither here nor there,” but always approaches the moving target that is the constantly changing desired image, even when that image changes every 50μs. We compare our experimental results with imagery from a conventional DLP projector with similar internal speed, and demonstrate that AR overlays on a moving object are more effective with this kind of low-latency display device than with displays of similar speed that use a conventional video interface. Turner Whitted, Anselmo Lastra, Peter Lincoln, Andrei State, Andrew Maimone, Henry Fuchs |
ISMAR | 7 |
| 2014 | Temporally enhanced 3D capture of room-sized dynamic scenes with commodity depth camerasabstractIn this paper, we introduce a system to capture the enhanced 3D structure of a room-sized dynamic scene with commodity depth cameras such as Microsoft Kinects. It is challenging to capture the entire dynamic room. First, the raw data from depth cameras are noisy due to the conflicts of the room's large volume and cameras' limited optimal working distance. Second, the severe occlusions between objects lead to dramatic missing data in the captured 3D. Our system incorporates temporal information to achieve a noise-free and complete 3D capture of the entire room. More specifically, we pre-scan the static parts of the room offline, and track their movements online. For the dynamic objects, we perform non-rigid alignment between frames and accumulate data over time. Our system also supports the topology changes of the objects and their interactions. We demonstrate the success of our system with various situations. Mingsong Dou, Henry Fuchs |
VR | 2 |
| 2014 | Telepresence: Soon not just a dream and a promiseabstractDreams of telepresence are fed by special effects in movies, on stage, and in TV programs. These illusions regularly fool the audiences of these productions, but they do not achieve telepresence for the actual participants in two or more distant locations. Some of the illusions used in these productions have been known and exploited for centuries. Why, then, is telepresence so diffcult to achieve? This talk will explore some of the classic illusions, review recent technical advances in the component technologies of 3D acquisition and 3D display, and suggest directions for future development. We will also examine the impact of such recent developments as Microsoft Kinect and Google Glass, and how they may dramatically improve the chances that some forms of telepresence may become commercially viable in the near future. Henry Fuchs |
VR | 1 |
| 2014 | Pinlight displays: wide field of view augmented reality eyeglasses using defocused point light sourcesabstractWe present a novel design for an optical see-through augmented reality display that offers a wide field of view and supports a compact form factor approaching ordinary eyeglasses. Instead of conventional optics, our design uses only two simple hardware components: an LCD panel and an array of point light sources (implemented as an edge-lit, etched acrylic sheet) placed directly in front of the eye, out of focus. We code the point light sources through the LCD to form miniature see-through projectors. A virtual aperture encoded on the LCD allows the projectors to be tiled, creating an arbitrarily wide field of view. Software rearranges the target augmented image into tiled sub-images sent to the display, which appear as the correct image when observed out of the viewer's accommodation range. We evaluate the design space of tiled point light projectors with an emphasis on increasing spatial resolution through the use of eye tracking. We demonstrate feasibility through software simulations and a real-time prototype display that offers a 110° diagonal field of view in the form factor of large glasses and discuss remaining challenges to constructing a practical display. Andrew Maimone, Douglas Lanman, Kishore Rathinavel, Kurtis Keller, David P. Luebke, Henry Fuchs |
ACM Trans. Graph. | 6 |
| 2013 | Scanning and tracking dynamic objects with commodity depth camerasabstractThe 3D data collected using state-of-the-art algorithms often suffers from various problems, such as incompletion and inaccuracy. Using temporal information has been proven effective for improving the reconstruction quality; for example, KinectFusion [21] shows significant improvements for static scenes. In this work, we present a system that uses commodity depth and color cameras, such as Microsoft Kinects, to fuse the 3D data captured over time for dynamic objects to build a complete and accurate model, and then tracks the model to match later observations. The key ingredients of our system include a nonrigid matching algorithm that aligns 3D observations of dynamic objects by using both geometry and texture measurements, and a volumetric fusion algorithm that fuses noisy 3D data. We demonstrate that the quality of the model improves dramatically by fusing a sequence of noisy and incomplete depth data of human and that by deforming this fused model to later observations, noise-and-hole-free 3D models are generated for the human moving freely. Mingsong Dou, Henry Fuchs, Jan-Michael Frahm |
ISMAR | 2 |
| 2013 | Computational augmented reality eyeglassesabstractIn this paper we discuss the design of an optical see-through head-worn display supporting a wide field of view, selective occlusion, and multiple simultaneous focal depths that can be constructed in a compact eyeglasses-like form factor. Building on recent developments in multilayer desktop 3D displays, our approach requires no reflective, refractive, or diffractive components, but instead relies on a set of optimized patterns to produce a focused image when displayed on a stack of spatial light modulators positioned closer than the eye accommodation distance. We extend existing multilayer display ray constraint and optimization formulations while also purposing the spatial light modulators both as a display and as a selective occlusion mask. We verify the design on an experimental prototype and discuss challenges to building a practical display. Andrew Maimone, Henry Fuchs |
ISMAR | 2 |
| 2013 | The 2013 virtual reality career awardabstractThe 2013 virtual reality career award goes to Henry Fuchs, The University of North Carolina at Chapel Hill, for his lifetime contributions to research and practice in virtual environments, telepresence, and medical applications. Since the 1970s, Henry Fuchs has made pioneer contributions to many of the technologies needed to enable virtual and augmented reality: automatic construction of 3D models and scenes, fast rendering algorithms (BSP Trees), graphics hardware (Pixel-Planes), large-area tracking systems (HiBall), optical and video see-through head-mounted displays. Many of these advances were inspired by demanding applications such as augmenting visualization for surgical assistance (by merging real and virtual objects) and teleimmersion for remote medical consultation. He continues to innovate within a multi-national telepresence research center with sites at ETH Zurich, NTU Singapore, and UNC Chapel Hill. The IEEE VGTC is pleased to award Henry Fuchs the 2013 Virtual Reality Career Award. Henry Fuchs |
VR | 1 |
| 2013 | General-purpose telepresence with head-worn optical see-through displays and projector-based lightingabstractIn this paper we propose a general-purpose telepresence system design that can be adapted to a wide range of scenarios and present a framework for a proof-of-concept prototype. The prototype system allows users to see remote participants and their surroundings merged into the local environment through the use of an optical see-through head-worn display. Real-time 3D acquisition and head tracking allows the remote imagery to be seen from the correct point of view and with proper occlusion. A projector-based lighting control system permits the remote imagery to appear bright and opaque even in a lit room. Immersion can be adjusted across the VR continuum. Our approach relies only on commodity hardware; we also experiment with wider field of view custom displays. Andrew Maimone, Xubo Yang, Nate Dierk, Andrei State, Mingsong Dou, Henry Fuchs |
VR | 6 |
| 2013 | Focus 3D: Compressive accommodation displayabstractWe present a glasses-free 3D display design with the potential to provide viewers with nearly correct accommodative depth cues, as well as motion parallax and binocular cues. Building on multilayer attenuator and directional backlight architectures, the proposed design achieves the high angular resolution needed for accommodation by placing spatial light modulators about a large lens: one conjugate to the viewer's eye, and one or more near the plane of the lens. Nonnegative tensor factorization is used to compress a high angular resolution light field into a set of masks that can be displayed on a pair of commodity LCD panels. By constraining the tensor factorization to preserve only those light rays seen by the viewer, we effectively steer narrow high-resolution viewing cones into the user's eyes, allowing binocular disparity, motion parallax, and the potential for nearly correct accommodation over a wide field of view. We verify the design experimentally by focusing a camera at different depths about a prototype display, establish formal upper bounds on the design's accommodation range and diffraction-limited performance, and discuss practical limitations that must be overcome to allow the device to be used with human observers. Andrew Maimone, Gordon Wetzstein, Matthew Hirsch, Douglas Lanman, Ramesh Raskar, Henry Fuchs |
ACM Trans. Graph. | 6 |
| 2013 | The 2013 Virtual Reality Career AwardabstractPresents the recipient of the 2013 Virtual Reality Career Award. Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | Room-sized informal telepresence systemabstractWe present a room-sized telepresence system for informal gatherings rather than conventional meetings. Unlike conventional systems which constrain participants to sit in fixed positions, our system aims to facilitate casual conversations between people in two sites. The system consists of a wall of large flat displays at each of the two sites, showing a panorama of the remote scene, constructed from a multiplicity of color and depth cameras. The main contribution of this paper is a solution that ameliorates the eye contact problem during conversation in typical scenarios while still maintaining a consistent view of the entire room for all participants. We achieve this by using two sets of cameras - a cluster of ”Panorama Cameras” located at the center of the display wall and are used to capture a panoramic view of the entire room, and a set of ”Personal Cameras” distributed along the display wall to capture front views of nearby participants. A robust segmentation algorithm with the assistance of depth cameras and an image synthesis algorithm work together to generate a consistent view of the entire scene. In our experience this new approach generates fewer distracting artifacts than conventional 3D reconstruction methods, while effectively correcting for eye gaze. Mingsong Dou, Jan-Michael Frahm, Henry Fuchs, Bill Mauchly, Mod Marathe |
VR | 4 |
| 2012 | Reducing interference between multiple structured light depth sensors using motionabstractWe present a method for reducing interference between multiple structured light-based depth sensors operating in the same spectrum with rigidly attached projectors and cameras. A small amount of motion is applied to a subset of the sensors so that each unit sees its own projected pattern sharply, but sees a blurred version of the patterns of other units. If high spacial frequency patterns are used, each sensor sees its own pattern with higher contrast than the patterns of other units, resulting in simplified pattern disambiguation. An analysis of this method is presented for a group of commodity Microsoft Kinect color-plus-depth sensors with overlapping views. We demonstrate that applying a small vibration with a simple motor to a subset of the Kinect sensors results in reduced interference, as manifested as holes and noise in the depth maps. Using an array of six Kinects, our system reduced interference-related missing data from from 16.6% to 1.4% of the total pixels. Another experiment with three Kinects showed an 82.2% percent reduction in the measurement error introduced by interference. A side-effect is blurring in the color images of the moving units, which is mitigated with post-processing. We believe our technique will allow inexpensive commodity depth sensors to form the basis of dense large-scale capture systems. Andrew Maimone, Henry Fuchs |
VR | 2 |
| 2012 | Enhanced personal autostereoscopic telepresence system using commodity depth cameras
Andrew Maimone, Jonathan Bidwell, Henry Fuchs |
Comput. Graph. | 4 |
| 2012 | The Design and Evaluation of a Large-Scale Real-Walking Locomotion InterfaceabstractRedirected Free Exploration with Distractors (RFEDs) is a large-scale real-walking locomotion interface developed to enable people to walk freely in Virtual Environments (VEs) that are larger than the tracked space in their facility. This paper describes the RFED system in detail and reports on a user study that evaluated RFED by comparing it to Walking-in-Place (WIP) and Joystick (JS) interfaces. The RFED system is composed of two major components, redirection and distractors. This paper discusses design challenges, implementation details, and lessons learned during the development of two working RFED systems. The evaluation study examined the effect of the locomotion interface on users' cognitive performance on navigation and wayfinding measures. The results suggest that participants using RFED were significantly better at navigating and wayfinding through virtual mazes than participants using walking-in-place and joystick interfaces. Participants traveled shorter distances, made fewer wrong turns, pointed to hidden targets more accurately and more quickly, and were able to place and label targets on maps more accurately, and more accurately estimate the virtual environment size. Tabitha C. Peck, Henry Fuchs, Mary C. Whitton |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | Encumbrance-free telepresence system with real-time 3D capture and display using commodity depth camerasabstractThis paper introduces a proof-of-concept telepresence system that offers fully dynamic, real-time 3D scene capture and continuous-viewpoint, head-tracked stereo 3D display without requiring the user to wear any tracking or viewing apparatus. We present a complete software and hardware framework for implementing the system, which is based on an array of commodity Microsoft Kinect™color-plus-depth cameras. Novel contributions include an algorithm for merging data between multiple depth cameras and techniques for automatic color calibration and preserving stereo quality even with low rendering rates. Also presented is a solution to the problem of interference that occurs between Kinect cameras with overlapping views. Emphasis is placed on a fully GPU-accelerated data processing and rendering pipeline that can apply hole filling, smoothing, data merger, surface generation, and color correction at rates of up to 100 million triangles/sec on a single PC and graphics board. Also presented is a Kinect-based marker-less tracking system that combines 2D eye recognition with depth information to allow head-tracked stereo views to be rendered for a parallax barrier autostereoscopic display. Our system is affordable and reproducible, offering the opportunity to easily deliver 3D telepresence beyond the researcher's lab. Andrew Maimone, Henry Fuchs |
ISMAR | 2 |
| 2011 | Continual surface-based multi-projector blending for moving objectsabstractWe introduce a general technique for blending imagery from multiple projectors on a tracked, moving, non-planar object. Our technique continuously computes visibility of pixels over the surfaces of the object and dynamically computes the per-pixel weights for each projector. This approach supports smooth transitions between areas of the object illuminated by different number of projectors, down to the illumination contribution of individual pixels within each polygon. To achieve real-time performance, we take advantage of graphics hardware, implementing much of the technique with a custom dynamic blending shader program within the GPU associated with each projector. We demonstrate the technique with some tracked objects being illuminated by three projectors. Peter Lincoln, Greg Welch, Henry Fuchs |
VR | 3 |
| 2011 | An evaluation of navigational ability comparing Redirected Free Exploration with Distractors to Walking-in-Place and joystick locomotio interfacesabstractWe report on a user study evaluating Redirected Free Exploration with Distractors (RFED), a large-scale, real-walking, locomotion interface, by comparing it to Walking-in-Place (WIP) and Joystick (JS), two common locomotion interfaces. The between-subjects study compared navigation ability in RFED, WIP, and JS interfaces in VEs that are more than two times the dimensions of the tracked space. The interfaces were evaluated based on navigation and wayfinding metrics and results suggest that participants using RFED were significantly better at navigating and wayfinding through virtual mazes than participants using walking-in-place and joystick interfaces. Participants traveled shorter distances, made fewer wrong turns, pointed to hidden targets more accurately and more quickly, and were able to place and label targets on maps more accurately. Moreover, RFED participants were able to more accurately estimate VE size. Tabitha C. Peck, Henry Fuchs, Mary C. Whitton |
VR | 2 |
| 2010 | Augmenting reality for medicine, training, presence and telepresenceabstractAt least since Sutherland's 1968 head-mounted display, augmented reality systems have inspired (and frustrated) generations of developers, users, and enthusiasts. These inspirations have led to decades of effort that yielded major innovations in technologies for 3D capture, 3D displays, tracking, and real-time image generation. Some of these component technologies, such as head-mounted displays, have proven much more difficult to bring to widespread adoption than many of us expected. Others, such as real-time image generation, have become phenomenally successful, with major societal impacts extending far beyond the initial applications. In this talk, I will describe several developments in these areas, and illustrate the resulting systems that exploited them: 1) enhancing 3D scene acquisition by laser scanning, or with structured light, or with multiple acquisition cameras; 2) augmenting a physician's view of her patient with registered internal imagery; 3) augmenting a user's surroundings with projection onto multiple nearby surfaces; 4) augmenting a user's remote presence using a human-sized avatar that mimics appearance, pose, and gestures; 5) augmenting tabletop displays with multi-user autostereoscopic capabilities. The common goal of all these systems is to enrich users' immediate surroundings with computer-generated or -controlled enhancements, which can run the gamut from mere virtual imagery to full-fledged robotic androids. I will also speculate about the possible paths to progress in the coming years. The future is encouraging, as the decreasing cost of the necessary components lowers the barriers to entry and encourages ever-increasing participation in innovation, development and use. Coupled with the rapidly advancing technologies in sensors, cameras, displays, robotics, and networks, this should enable us to accelerate bringing the visions of augmenting reality to daily life. Henry Fuchs |
ISMAR | 1 |
| 2010 | A practical multi-viewer tabletop autostereoscopic displayabstractThis paper introduces a multi-user autostereoscopic tabletop display and its associated real-time rendering methods. Tabletop displays that support both multiple viewers and autostereoscopy have been extremely difficult to construct. Our new system is inspired by the “Random Hole Display” design that modified the pattern of openings in a barrier mounted in front of a flat panel display from thin slits to a dense pattern of tiny, pseudo-randomly placed holes. This allows viewers anywhere in front of the display to see a different subset of the display's native pixels through the random-hole screen. However, a fraction of the visible pixels will be observable by more than a single viewer. Thus the main challenge is handling these “conflicting” pixels, which ideally must show different colors to each viewer. We introduce several solutions to this problem and describe in detail the current method of choice, a combination of color blending and approximate error diffusion, performing in real time in our GPU-based implementation. The easily reproducible design uses a pattern film barrier affixed to the display by means of a transparent polycarbonate layer spacer. We use a commercial optical tracker for viewers' locations and synthesize the appropriate image (or a stereoscopic image pair) for each viewer. The system supports graceful degradation with increasing number of simultaneous views, and graceful improvement as the number of views decreases. Gu Ye, Andrei State, Henry Fuchs |
ISMAR | 3 |
| 2010 | Improved Redirection with Distractors: A large-scale-real-walking locomotion interface and its effect on navigation in virtual environmentsabstractUsers in virtual environments often find navigation more difficult than in the real world. Our new locomotion interface, Improved Redirection with Distractors (IRD), enables users to walk in larger-than-tracked space VEs without predefined waypoints. We compared IRD to the current best interface, really walking, by conducting a user study measuring navigational ability. Our results show that IRD users can really walk through VEs that are larger than the tracked space and can point to targets and complete maps of VEs no worse than when really walking. Tabitha C. Peck, Henry Fuchs, Mary C. Whitton |
VR | 2 |
| 2009 | Animatronic Shader Lamps AvatarsabstractApplications such as telepresence and training involve the display of real or synthetic humans to multiple viewers. When attempting to render the humans with conventional displays, non-verbal cues such as head pose, gaze direction, body posture, and facial expression are difficult to convey correctly to all viewers. In addition, a framed image of a human conveys only a limited physical sense of presence - primarily through the display's location. While progress continues on articulated robots that mimic humans, the focus has been on the motion and behavior of the robots. We introduce a new approach for robotic avatars of real people: the use of cameras and projectors to capture and map the dynamic motion and appearance of a real person onto a humanoid animatronic model. We call these devices Animatronic Shader Lamps Avatars (SLA).We present a proof-of-concept prototype comprised of a camera, a tracking system, a digital projector, and a life-sized styrofoam head mounted on a pan-tilt unit. The system captures imagery of a moving, talking user and maps the appearance and motion onto the animatronic SLA, delivering a dynamic, real-time representation of the user to multiple viewers. Peter Lincoln, Greg Welch, Andrew Nashel, Adrian Ilie, Andrei State, Henry Fuchs |
ISMAR | 6 |
| 2009 | A Distributed Cooperative Framework for Continuous Multi-Projector Pose EstimationabstractWe present a novel calibration framework for multi-projector displays that achieves continuous geometric calibration by estimating and refining the poses of all projectors in an ongoing fashion during actual display use. Our framework provides scalability by operating as a distributed system of "intelligent" projector units: projectors augmented with rigidly-mounted cameras, and paired with dedicated computers. Each unit interacts asynchronously with its peers, leveraging their combined computational power to cooperatively estimate the poses of all of the projectors. In cases where the projection surface is static, our system is able to continuously refine all of the projector poses, even when they change simultaneously. Tyler Johnson, Greg Welch, Henry Fuchs, Eric La Force, Herman Towles |
VR | 3 |
| 2009 | Evaluation of Reorientation Techniques and Distractors for Walking in Large Virtual EnvironmentsabstractVirtual Environments (VEs) that use a real-walking locomotion interface have typically been restricted in size to the area of the tracked lab space. Techniques proposed to lift this size constraint, enabling real walking in VEs that are larger than the tracked lab space, all require reorientation techniques (ROTs) in the worst-case situation-when a user is close to walking out of the tracked space. We propose a new ROT using visual and audial distractors-objects in the VE that the user focuses on while the VE rotates-and compare our method to current ROTs through three user studies. ROTs using distractors were preferred and ranked more natural by users. Our findings also suggest that improving visual realism and adding sound increased a user's feeling of presence. Users were also less aware of the rotating VE when ROTs with distractors were used. Our findings also suggest that improving visual realism and adding sound increased a user's feeling of presence. Tabitha C. Peck, Henry Fuchs, Mary C. Whitton |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2008 | Evaluation of Reorientation Techniques for Walking in Large Virtual EnvironmentsabstractVirtual environments (VEs) that use a real-walking locomotion interface have typically been restricted in size to the area of the tracked lab space. Techniques proposed to lift this size constraint, enabling real walking in VEs that are larger than the tracked lab space, all require reorientation techniques (ROTs) in the worst-case situation–when a user is close to walking out of the tracked space. We propose a new ROT using distractors–objects in the VE for the user to focus on while the VE rotates –and compare our method to current ROTs through two user studies. Our findings show ROTs using distractors were preferred and ranked more natural by users. Users were also less aware of the rotating VE, when ROTs with distractors were used. Tabitha C. Peck, Mary C. Whitton, Henry Fuchs |
VR | 3 |
| 2008 | Exploring the potential of video technologies for collaboration in emergency medical care: Part II. Task performanceabstractAbstract We conducted an experiment with a posttest, between‐subjects design to evaluate the potential of emerging 3D telepresence technology to support collaboration in emergency health care. 3D telepresence technology has the potential to provide richer visual information than do current 2D video conferencing techniques. This may be of benefit in diagnosing and treating patients in emergency situations where specialized medical expertise is not locally available. The experimental design and results concerning information behavior are presented in the article “Exploring the Potential of Video Technologies for Collaboration in Emergency Medical Care: Part I. Information Sharing” (Sonnenwald et al., this issue). In this article, we explore paramedics' task performance during the experiment as they diagnosed and treated a trauma victim while working alone or in collaboration with a physician via 2D videoconferencing or via a 3D proxy. Analysis of paramedics' task performance shows that paramedics working with a physician via a 3D proxy performed the fewest harmful interventions and showed the least variation in task performance time. Paramedics in the 3D proxy condition also reported the highest levels of self‐efficacy. Interview data confirm these statistical results. Overall, the results indicate that 3D telepresence technology has the potential to improve paramedics' performance of complex medical tasks and improve emergency trauma health care if designed and implemented appropriately. Hanna M. Söderholm, Diane H. Sonnenwald, James E. Manning, Bruce Cairns, Greg Welch, Henry Fuchs |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2008 | Exploring the potential of video technologies for collaboration in emergency medical care: Part I. Information sharingabstractAbstract We are investigating the potential of 3D telepresence, or televideo, technology to support collaboration among geographically separated medical personnel in trauma emergency care situations. 3D telepresence technology has the potential to provide richer visual information than current 2D videoconferencing techniques. This may be of benefit in diagnosing and treating patients in emergency situations where specialized medical expertise is not locally available. The 3D telepresence technology does not yet exist, and there is a need to understand its potential before resources are spent on its development and deployment. This poses a complex challenge. How can we evaluate the potential impact of a technology within complex, dynamic work contexts when the technology does not yet exist? To address this challenge, we conducted an experiment with a posttest, between‐subjects design that takes the medical situation and context into account. In the experiment, we simulated an emergency medical situation involving practicing paramedics and physicians, collaborating remotely via two conditions: with today's 2D videoconferencing and a 3D telepresence proxy. In this article, we examine information sharing between the attending paramedic and collaborating physician. Postquestionnaire data illustrate that the information provided by the physician was perceived to be more useful by the paramedic in the 3D proxy condition than in the 2D condition; however, data pertaining to the quality of interaction and trust between the collaborating physician and paramedic show mixed results. Postinterview data help explain these results. Diane H. Sonnenwald, Hanna M. Söderholm, James E. Manning, Bruce Cairns, Greg Welch, Henry Fuchs |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2007 | Real-Time Projector Tracking on Complex Geometry Using Ordinary ImageryabstractUsing cameras to geometrically calibrate projector-based displays has been widely reported in the literature over the last decade. Most systems project structured-light patterns during a setup phase before starting the application in order to evaluate the geometric mapping from projector pixels to locations on the display surface. If this mapping changes e.g. the projector is moved, the application must be stopped and calibration re-initiated. Our pose estimation technique is based on detecting a set of image correspondences between the moving projector and a static camera using only the imagery displayed by the projector. We use an up-front calibration process to establish the geometry of the display surface, the pose of the camera, and the initial pose of the projector. Tyler Johnson, Henry Fuchs |
CVPR | 2 |
| 2007 | Real-Time Projector Tracking on Complex Geometry Using Ordinary Imagery
Tyler Johnson, Henry Fuchs |
CVPR | 2 |
| 2007 | The potential impact of 3d telepresence technology on task performance in emergency trauma careabstractEmergency trauma is a major health problem worldwide. To evaluate the potential of emerging 3D telepresence technology for facilitating paramedic - physician collaboration while providing emergency medical trauma care we conducted a between-subjects post-test experimental lab study. During a simulated emergency situation 60 paramedics diagnosed and treated a trauma victim while working alone or in collaboration with a physician via 2D video or a 3D proxy. Analysis of paramedics' task performance shows that the fewest harmful procedures occurred in the 3D proxy condition. Paramedics in the 3D proxy condition also reported higher levels of self-efficacy. These results indicate 3D telepresence technology has potential to improve paramedics' performance of complex emergency medical tasks and improve emergency trauma health care when designed appropriately. Hanna M. Söderholm, Diane H. Sonnenwald, Bruce Cairns, James E. Manning, Greg Welch, Henry Fuchs |
GROUP | 6 |
| 2007 | Temporally Consistent Reconstruction from Multiple Video Streams Using Enhanced Belief PropagationabstractWe present an approach for 3D reconstruction from multiple video streams taken by static, synchronized and calibrated cameras that is capable of enforcing temporal consistency on the reconstruction of successive frames. Our goal is to improve the quality of the reconstruction by finding corresponding pixels in subsequent frames of the same camera using optical flow, and also to at least maintain the quality of the single time-frame reconstruction when these correspondences are wrong or cannot be found. This allows us to process scenes with fast motion, occlusions and self- occlusions where optical flow fails for large numbers of pixels. To this end, we modify the belief propagation algorithm to operate on a 3D graph that includes both spatial and temporal neighbors and to be able to discard messages from outlying neighbors. We also propose methods for introducing a bias and for suppressing noise typically observed in uniform regions. The bias encapsulates information about the background and aids in achieving a temporally consistent reconstruction and in the mitigation of occlusion related errors. We present results on publicly available real video sequences. We also present quantitative comparisons with results obtained by other researchers. E. Scott Larsen, Philippos Mordohai, Marc Pollefeys, Henry Fuchs |
ICCV | 4 |
| 2007 | VR Support of Clinical Applications: Collaboration, Politics, and Ethics
Grigore C. Burdea, Zohara A. Cohen, Henry Fuchs, Richard M. Satava |
VR | 3 |
| 2007 | A Personal Surround Environment: Projective Display with Correction for Display Surface Geometry and Extreme Lens DistortionabstractProjectors equipped with wide-angle lenses can have an advantage over traditional projectors in creating immersive display environments since they can be placed very close to the display surface to reduce user shadowing issues while still producing large images. However, wide-angle projectors exhibit severe image distortion requiring the image generator to correctively pre-distort the output image. In this paper, we describe a new technique based on Raskar's (1998) two-pass rendering algorithm that is able to correct for both arbitrary display surface geometry and the extreme lens distortion caused by fisheye lenses. We further detail how the distortion correction algorithm can be implemented in a real-time shader program running on a commodity GPU to create low-cost, personal surround environments Tyler Johnson, Florian Gyarfas, Richard Skarbez, Herman Towles, Henry Fuchs |
VR | 5 |
| 2006 | RANSAC-Assisted Display Model Reconstruction for Projective DisplayabstractUsing projectors to create perspectively correct imagery on arbitrary display surfaces requires geometric knowledge of the display surface shape, the projector calibration, and the user’s position in a common coordinate system. Prior solutions have most commonly modeled the display surface as a tessellated mesh derived from the 3D-point cloud acquired during system calibration. In this paper we describe a method for functional reconstruction of the display surface, which takes advantage of the knowledge that most interior display spaces (e.g. walls, floors, ceilings, building columns) are piecewise planar. Using a RANSAC algorithm to recursively fit planes to a 3D-point cloud sampling of the surface, followed by a conversion of the plane definitions into simple planar polygon descriptions, we are able to create a geometric model which is less complex than a dense tessellated mesh and offers a simple method for accurately modeling the corners of rooms. Planar models also eliminate subtle, but irritating, texture distortion often seen in tessellated mesh approximations to planar surfaces. Patrick Quirk, Tyler Johnson, Richard Skarbez, Herman Towles, Florian Gyarfas, Henry Fuchs |
VR | 6 |
| 2005 | Simulation-Based Design and Rapid Prototyping of a Parallax-Free, Orthoscopic Video See-Through Head-Mounted DisplayabstractWe built a video see-through head-mounted display with zero eye offset from commercial components and a mount fabricated via rapid prototyping. The orthoscopic HMD's layout was created and optimized with a software simulator. We describe simulator and HMD design, we show the HMD in use and demonstrate zero parallax. Andrei State, Kurtis Keller, Henry Fuchs |
ISMAR | 3 |
| 2005 | Immersive Integration for Virtual and Human-Centered EnvironmentsabstractSummary form only given. We envision future work and play environments that are more effectively human-centered with the user's computing interface being more closely integrated with the physical surroundings than today's conventional computer display screens and keyboards. We are working toward realizable versions of such environments, in which multiple video projectors and digital cameras enable every visible surface to be both measured in 3D and used for display. If the 3D surface positions are transmitted to a distant location, they may also enable distant collaborations to become more like working in adjacent offices connected by large windows. In one prototype, depth maps are calculated from streams of video images and the resulting 3D surface points are displayed to the user in head-tracked stereo. Another prototype allows direct "painting" onto movable objects - a dollhouse, for example. One long-term goal is advanced training for trauma surgeons by immersive replay of recorded procedures. More generally, we hope to demonstrate that the principal interface of a future computing environment need not be limited to a screen the size of one or two sheets of paper. Just as a useful physical environment is all around us, so too can the increasingly ubiquitous computing environment be all around us - becoming more effectively human-centered and integrated into our physical surroundings. Henry Fuchs |
VL/HCC | 1 |
| 2005 | Adaptive Instant Displays: Continuously Calibrated Projections Using Per-Pixel Light ControlabstractWe present a framework for achieving user-defined on-demand displays in setups containing bricks of movable cameras and DLP-projectors. A dynamic calibration procedure is introduced, which handles cameras and projectors in a unified way and allows continuous flexible setup changes, while seamless projection alignment and blending is performed simultaneously. For interaction, an intuitive laser pointer based technique is developed, which can be combined with real-time 3D information acquired from the scene. All these tasks can be performed concurrently with the display of a user-chosen application in a non-disturbing way. This is achieved by using an imperceptible structured light approach enabling pixel-based surface light control suited for a wide range of computer graphics and vision algorithms. To ensure scalability of light control in the same working space, multiple projectors are multiplexed. Daniel Cotting, Henry Fuchs, Remo Ziegler, Markus Gross 0001 |
Comput. Graph. Forum | 2 |
| 2004 | Embedding Imperceptible Patterns into Projected Images for Simultaneous Acquisition and DisplayabstractWe introduce a method to imperceptibly embed arbitrary binary patterns into ordinary color images displayed by unmodified off-the-shelf digital light processing (DLP) projectors. The encoded images are visible only to cameras synchronized with the projectors and exposed for a short interval, while the original images appear only minimally degraded to the human eye. To achieve this goal, we analyze and exploit the micro-mirror modulation pattern used by the projection technology to generate intensity levels for each pixel and color channel. Our real-time embedding process maps the user's original color image values to the nearest values whose camera-perceived intensities are the ones desired by the binary image to be embedded. The color differences caused by this mapping process are compensated by error-diffusion dithering. The non-intrusive nature of our approach allows simultaneous (immersive) display and acquisition under controlled lighting conditions, as defined on a pixel level by the binary patterns. We therefore introduce structured light techniques into human-inhabited mixed and augmented reality environments, where they previously often were too intrusive. Daniel Cotting, Martin Näf, Markus Gross 0001, Henry Fuchs |
ISMAR | 4 |
| 2004 | Immersive Integration of Physical and Virtual EnvironmentsabstractAbstract We envision future work and play environments in which the user's computing interface is more closely integrated with the physical surroundings than today's conventional computer display screens and keyboards.We are working toward realizable versions of such environments, in which multiple video projectors and digital cameras enable every visible surface to be both measured in 3D and used for display. If the 3D surface positions were transmitted to a distant location, they may also enable distant collaborations to become more like working in adjacent offices connected by large windows. With collaborators at the University of Pennsylvania, Brown University, Advanced Network and Services, and the Pittsburgh Supercomputing Center, we at Chapel Hill have been working to bring these ideas to reality. In one system, depth maps are calculated from streams of video images and the resulting 3D surface points are displayed to the user in head‐tracked stereo. Among the applications we are pursuing for this tele‐presence technology, is advanced training for trauma surgeons by immersive replay of recorded procedures. Other applications display onto physical objects, to allow more natural interaction with them ``painting'' a dollhouse, for example. More generally, we hope to demonstrate that the principal interface of a future computing environment need not be limited to a screen the size of one or two sheets of paper. Just as a useful physical environment is all around us, so too can the increasingly ubiquitous computing environment be all around us ‐integrated seamlessly with our physical surroundings. Henry Fuchs |
Comput. Graph. Forum | 1 |
| 2004 | Introduction to the Special Issue on Immersive TelecommunicationsabstractS.285-287 Oliver Schreer, Henry Fuchs, Wijnand A. IJsselsteijn, Hiroshi Yasuda |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Real-time compression for dynamic 3D environmentsabstractThe goal of tele-immersion has long been to enable people at remote locations to share a sense of presence. A tele-immersion system acquires the 3D representation of a collaborator's environment remotely and sends it over the network where it is rendered in the user's environment. Acquisition, reconstruction, transmission, and rendering all have to be done in real-time to create a sense of presence. With added commodity hardware resources, parallelism can increase the acquisition volume and reconstruction data quality while maintaining real-time performance. However this is not as easy for rendering since all of the data need to be combined into a single display.In this paper we present an algorithm to compress data from such 3D environments in real-time to solve this imbalance. We expect the compression algorithm to scale comparably to the acquisition and reconstruction, reduce network transmission bandwidth, and reduce the rendering requirement for real-time performance. We have tested the algorithm using a synthetic office data set and have achieved a 5 to 1 compression for 22 depth streams. Sang-Uok Kum, Ketan Mayer-Patel, Henry Fuchs |
ACM Multimedia | 3 |
| 2003 | Acquisition of large-scale surface light fieldsabstractAbstract: This sketch reports the development of an acquisition methodology to acquire high-resolution surface light fields of an office-size environment. 1 Wei-Chao Chen, Lars S. Nyland, Anselmo Lastra, Henry Fuchs |
SIGGRAPH | 4 |
| 2002 | Augmented reality guidance for needle biopsies: An initial randomized, controlled trial in phantoms
Michael Rosenthal, Andrei State, Joohi Lee, Gentaro Hirota, Jeremy Ackerman, Kurtis Keller, Etta D. Pisano, Michael R. Jiroutek, Keith E. Muller, Henry Fuchs |
Medical Image Anal. | 10 |
| 2001 | An implicit finite element method for elastic solids in contactabstractFocuses on the simulation of mechanical contact between nonlinearly elastic objects such as the components of the human body. The computation of the reaction forces that act on the contact surfaces (contact forces) is the key for designing a reliable contact handling algorithm. In traditional methods, contact forces are often defined as discontinuous functions of deformation, which leads to poor convergence characteristics. This problem becomes especially serious in areas with complicated self contact such as skin folds. We introduce a novel penalty finite element formulation based on the concept of material depth, the distance between a particle inside an object and the object's boundary. By linearly interpolating pre-computed material depths at node points, contact forces can be analytically integrated over contact surfaces without increasing the computational cost. The continuity achieved by this formulation supports an efficient and reliable solution of the nonlinear system. This algorithm is implemented as part of our implicit finite element program for static, quasistatic and dynamic analysis of nonlinear viscoelastic solids. We demonstrate its effectiveness on an animation showing realistic effects such as folding skin and sliding contacts of the tissues involved in knee flexion. The finite element model of the leg and its internal structures was derived from the Visible Human data set. Gentaro Hirota, Susan Fisher, Andrei State, Henry Fuchs |
CA | 5 |
| 2001 | Augmented Reality Guidance for Needle Biopsies: A Randomized, Controlled Trial in Phantoms
Michael Rosenthal, Andrei State, Joohi Lee, Gentaro Hirota, Jeremy Ackerman, Kurtis Keller, Etta D. Pisano, Michael R. Jiroutek, Keith E. Muller, Henry Fuchs |
MICCAI | 10 |
| 2001 | Intraoperative Tracking of Anatomical Structures Using Fluoroscopy and a Vascular Balloon Catheter
Michael Rosenthal, Susan Weeks, Stephen R. Aylward, Elizabeth Bullitt, Henry Fuchs |
MICCAI | 5 |
| 2001 | Life-sized projector-based dioramasabstractWe introduce an idea and some preliminary results for a new projector-based approach to re-creating real and imagined sites. Our goal is to achieve re-creations that are both visually and spatially realistic, providing a small number of relatively unencumbered users with a strong sense of immersion as they jointly walkaround the virtual site.Rather than using head-mounted or general-purpose projector-based displays, our idea builds on previous projector-based work on spatially-augmented realityand shader lamps. Using simple white building blocks we construct a static physical model that approximates the size, shape, and spatial arrangementof the site. We then project dynamic imagery onto the blocks, transforming the lifeless physical model into a visually faithful reproduction of the actual site. Some advantages of this approach include wide field-of-view imagery, real walking around the site, reduced sensitivity to tracking errors, reduced sensitivity to system latency, auto-stereoscopic vision, the natural addition of augmented virtualityand the provision of haptics.In addition to describing the major challenges to (and limitations of) this vision, in this paper we describe some short-term solutions and practical methods, and we present some proof-of-concept results. Kok-Lim Low, Greg Welch, Anselmo Lastra, Henry Fuchs |
VRST | 4 |
| 2000 | Toward a compelling sensation of telepresence: demonstrating a portal to a distant (static) officeabstractIn 1998 we introduced the idea for a project we call the Office of the Future. Our long-term vision is to provide a better every-day working environment, with high-fidelity scene reconstruction for life-sized 3D tele-collaboration. In particular, we want a true sense of presence with our remote collaborator and their real surroundings. The challenges related to this vision are enormous and involve many technical tradeoffs. This is true in particular for scene reconstruction. Researchers have been striving to achieve real-time approaches, and while they have made respectable progress, the limitations of conventional technologies relegate them to relatively low resolution in a restricted volume. We present a significant step toward our ultimate goal, via a slightly different path. In lieu of low-fidelity dynamic scene modeling we present an exceedingly high fidelity reconstruction of a real but static office. By assembling the best of available hardware and software technologies in static scene acquisition, modeling algorithms, rendering, tracking and stereo projective display, we are able to demonstrate a portal to a real office, occupied today by a mannequin, and in the future by a real remote collaborator. We now have both a compelling sense of just how good it could be, and a framework into which we will later incorporate dynamic scene modeling, as we continue to head toward our ultimate goal of 3D collaborative telepresence. Wei-Chao Chen, Herman Towles, Lars S. Nyland, Greg Welch, Henry Fuchs |
IEEE Visualization | 5 |
| 1999 | Immersive teleconferencing: a new algorithm to generate seamless panoramic video imageryabstractThis paper presents a new algorithm for immersive teleconferencing, which addresses the problem of registering and blending multiple images together to create a single seamless panorama. In the immersive teleconference paradigm, one frame of the teleconference is a panorama that is constructed from a compound-image sensing device. These frames are rendered at the remote site on a projection surface that surrounds the user, creating an immersive feeling of presence and participation in the teleconference. Our algorithm efficiently creates panoramic frames for a teleconference session that are both geometrically registered and intensity blended. We demonstrate a prototype that is able to capture images from a compound-image sensor, register them into a seamless panoramic frame, and render those panoramic frames on a projection surface at 30 frames per second. Aditi Majumder, W. Brent Seales, Meenakshisundaram Gopi, Henry Fuchs |
ACM Multimedia (1) | 4 |
| 1999 | Geometrically correct imagery for teleconferencingabstractCurrent camera-monitor teleconferencing applications produce unrealistic imagery and break any sense of presence for the participants. Other capture/display technologies can be used to provide more compelling teleconferencing. However, complex geometries in capture/display systems make producing geometrically correct imagery difficult. It is usually impractical to detect, model and compensate for all effects introduced by the capture/display system. Most applications simply ignore these issues and rely on the user acceptance of the camera-monitor paradigm. Ruigang Yang, Michael S. Brown, W. Brent Seales, Henry Fuchs |
ACM Multimedia (1) | 4 |
| 1999 | Multi-Projector Displays Using Camera-Based RegistrationabstractConventional projector-based display systems are typically designed around precise and regular configurations of projectors and display surfaces. While this results in rendering simplicity and speed, it also means painstaking construction and ongoing maintenance. In previously published work, we introduced a vision of projector-based displays constructed from a collection of casually-arranged projectors and display surfaces. In this paper, we present flexible yet practical methods for realizing this vision, enabling low-cost mega-pixel display systems with large physical dimensions, higher resolution, or both. The techniques afford new opportunities to build personal 3D visualization systems in offices, conference rooms, theaters, or even your living room. As a demonstration of the simplicity and effectiveness of the methods that we continue to perfect, we show in the included video that a 10-year old child can construct and calibrate a two-camera, two-projector, head-tracked display system, all in about 15 minutes. Ramesh Raskar, Michael S. Brown, Ruigang Yang, Wei-Chao Chen, Greg Welch, Herman Towles, W. Brent Seales, Henry Fuchs |
IEEE Visualization | 8 |
| 1998 | Augmented Reality Visualization for Laparoscopic Surgery
Henry Fuchs, Mark A. Livingston, Ramesh Raskar, D'nardo Colucci, Kurtis Keller, Andrei State, Jessica R. Crawford, Paul Rademacher, Samuel H. Drake, Anthony A. Meyer |
MICCAI | 1 |
| 1998 | The Office of the Future: A Unified Approach to Image-based Modeling and Spatially Immersive DisplaysabstractWe introduce ideas, proposed technologies, and initial results for an office of the future that is based on a unified application of computer vision and computer graphics in a system that combines and builds upon the notions of the CAVE™, tiled display systems, and image-based modeling .The basic idea is to use real-time computer vision techniques to dynamically extract per-pixel depth and reflectance information for the visible surfaces in the office including walls, furniture, objects, and people, and then to either project images on the surfaces, render images of the surfaces , or interpret changes in the surfaces.In the first case, one could designate every-day (potentially irregular) real surfaces in the office to be used as spatially immersive display surfaces, and then project high-resolution graphics and text onto those surfaces.In the second case, one could transmit the dynamic image-based models over a network for display at a remote site.Finally, one could interpret dynamic changes in the surfaces for the purposes of tracking, interaction, or augmented reality applications.To accomplish the simultaneous capture and display we envision an office of the future where the ceiling lights are replaced by computer controlled cameras and "smart" projectors that are used to capture dynamic image-based models with imperceptible structured light techniques, and to display high-resolution images on designated display surfaces.By doing both simultaneously on the designated display surfaces, one can dynamically adjust or autocalibrate for geometric, intensity, and resolution variations resulting from irregular or changing display surfaces, or overlapped projector images.Our current approach to dynamic image-based modeling is to use an optimized structured light scheme that can capture per-pixel depth and reflectance at interactive rates.Our system implementation is not yet imperceptible, but we can demonstrate the approach in the laboratory.Our approach to rendering on the designated (potentially irregular) display surfaces is to employ a two-pass projective texture scheme to generate images that when projected onto the surfaces appear correct to a moving headtracked observer.We present here an initial implementation of the overall vision, in an office-like setting, and preliminary demonstrations of our dynamic modeling and display techniques. Ramesh Raskar, Greg Welch, Matthew D. Cutts, Adam T. Lake, Lev Stesin, Henry Fuchs |
SIGGRAPH | 6 |
| 1997 | Building Telepresence Systems: Translating Science Fiction Ideas into RealityabstractMany people feel that at some time in the distant future it will be possible to see and interact with remote individuals so realistically that they will appear to be standing next to us, that it will also be possible for physicians to look inside their patients and see organs and tumors, as if they possessed Superman’s X‐ray vision. Although we are far from the realization of such dreams, we are witnessing some encouraging progress toward them. For example, we understand that two aspects common to these systems are a) the acquisition and synthesis of complex 3D information (whether from cameras at a distant scene or medical imaging devices inside a patient), and b) reconstruction and presentation of the 3D information to the observer (whether a distant collaborator or a nearby physician). Early results from several institutions are encouraging. It is now possible to walk around distant scenes, although the visual data still needs to be preprocessed, the scene still needs to be static and the reconstruction still has gaps. It is possible to look inside patients, although with crude ultrasound imaging and with very limited visualization, or not naturally from the physician’s own point of view. It is not yet possible to view any of this kind of information with high degree of immersive 3D realism without wearing cumbersome visualization aids. Truly compelling realization of these long‐held dreams will take careful analysis of the remaining problems, creative thinking about new approaches, and innovative and sustained development of the required acquisition and display technologies. Henry Fuchs |
Comput. Graph. Forum | 1 |
| 1997 | Conveying the 3D Shape of Smoothly Curving Transparent Surfaces via TextureabstractTransparency can be a useful device for depicting multiple overlapping surfaces in a single image. The challenge is to render the transparent surfaces in such a way that their 3D shape can be readily understood and their depth distance from underlying structures clearly perceived. This paper describes our investigations into the use of sparsely-distributed discrete, opaque texture as an artistic device for more explicitly indicating the relative depth of a transparent surface and for communicating the essential features of its 3D shape in an intuitively meaningful and minimally occluding way. The driving application for this work is the visualization of layered surfaces in radiation therapy treatment planning data, and the technique is illustrated on transparent isointensity surfaces of radiation dose. We describe the perceptual motivation and artistic inspiration for defining a stroke texture that is locally oriented in the direction of greatest normal curvature (and in which individual strokes are of a length proportional to the magnitude of the curvature in the direction they indicate), and we discuss two alternative methods for applying this texture to isointensity surfaces defined in a volume. We propose an experimental paradigm for objectively measuring observers' ability to judge the shape and depth of a layered transparent surface, in the course of a task which is relevant to the needs of radiotherapy treatment planning, and use this paradigm to evaluate the practical effectiveness of our approach through a controlled observer experiment based on images generated from actual clinical data. Victoria Interrante, Henry Fuchs, Stephen M. Pizer |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 1996 | Technologies for Augmented Reality Systems: Realizing Ultrasound-Guided Needle BiopsiesabstractWe present a real-time stereoscopic video-see-through augmented reality (AR) system applied to the medical procedure known as ultrasound-guided needle biopsy of the breast.The AR system was used by a physician during procedures on breast models and during non-invasive examinations of human subjects.The system merges rendered live ultrasound data and geometric elements with stereo images of the patient acquired through head-mounted video cameras and presents these merged images to the physician in a head-mounted display.The physician sees a volume visualization of the ultrasound data directly under the ultrasound probe, properly registered within the patient and with the biopsy needle.Using this system, a physician successfully guided a needle into an artificial tumor within a training phantom of a human breast.We discuss the construction of the AR system and the issues and decisions which led to the system architecture and the design of the video see-through head-mounted display.We designed methods to properly resolve occlusion of the real and synthetic image elements.We developed techniques for realtime volume visualization of time-and position-varying ultrasound data.We devised a hybrid tracking system which achieves improved registration of synthetic and real imagery and we improved on previous techniques for calibration of a magnetic tracker. Andrei State, Mark A. Livingston, William F. Garrett, Gentaro Hirota, Mary C. Whitton, Etta D. Pisano, Henry Fuchs |
SIGGRAPH | 7 |
| 1996 | Real-Time Incremental Visualization of Dynamic Ultrasound Volumes Using Parallel BSP TreesabstractWe present a method for producing real-time volume visualizations of continuously captured, arbitrarily-oriented 2D arrays (slices) of data. Our system constructs a 3D representation on-the-fly from incoming 2D ultrasound slices by modeling and rendering the slices as planar polygons with translucent surface textures. We use binary space partition (BSP) tree data structures to provide non-intersecting, visibility-ordered primitives for accurate opacity accumulation images. New in our system is a method of using parallel, time-shifted BSP trees to efficiently manage the continuously captured ultrasound data and to decrease the variability in image generation time between output frames. This technique is employed in a functioning real-time augmented reality system that a physician has used to examine human patients prior to breast biopsy procedures. We expect the technique can be used for real-time visualization of any 2D data being collected from a tracked sensor moving along an arbitrary path. William F. Garrett, Henry Fuchs, Mary C. Whitton, Andrei State |
IEEE Visualization | 2 |
| 1996 | Illustrating Transparent Surfaces with Curvature-Directed StrokesabstractTransparency can be a useful device for simultaneously depicting multiple superimposed layers of information in a single image. However, in computer-generated pictures-as in photographs and in directly viewed actual objects-it can often be difficult to adequately perceive the three-dimensional shape of a layered transparent surface or its relative depth distance from underlying structures. Inspired by artists' use of line to show shape, we have explored methods for automatically defining a distributed set of opaque surface markings that intend to portray the three-dimensional shape and relative depth of a smoothly curving layered transparent surface in an intuitively meaningful (and minimally occluding) way. This paper describes the perceptual motivation, artistic inspiration and practical implementation of an algorithm for "texturing" a transparent surface with uniformly distributed opaque short strokes, locally oriented in the direction of greatest normal curvature, and of length proportional to the magnitude of the surface curvature in the stroke direction. The driving application for this work is the visualization of layered surfaces in radiation therapy treatment planning data, and the technique is illustrated on transparent isointensity surfaces of radiation dose. Victoria Interrante, Henry Fuchs, Stephen M. Pizer |
IEEE Visualization | 2 |
| 1995 | Interactive Volume Visualization on a Heterogeneous Message-Passing MulticomputerabstractThis paper describes VOL2, an interactive general-purpose volume renderer based on ray casting and implemented on Pixel-Planes 5, a distributed-memory, message-passing multicomputer. VOL2 is a pipelined renderer using image-space task parallelism and object-space data partitioning. We describe the parallelization and load balancing techniques used in order to achieve interactive response and near-real-time frame rates. We also present a number of applications for our system and derive some general conclusions about operation of image-order rendering algorithms on message-passing multicomputers. Andrei State, Jonathan McAllister, Ulrich Neumann, Tim J. Cullip, David T. Chen, Henry Fuchs |
SI3D | 7 |
| 1995 | Enhancing Transparent Skin Surfaces with Ridge and Valley LinesabstractThere are many applications that can benefit from the simultaneous display of multiple layers of data. The objective in these cases is to render the layered surfaces in a such way that the outer structures can be seen and seen through at the same time. The paper focuses on the particular application of radiation therapy treatment planning, in which physicians need to understand the three dimensional distribution of radiation dose in the context of patient anatomy. We describe a promising technique for communicating the shape and position of the transparent skin surface while at the same time minimally occluding underlying isointensity dose surfaces and anatomical objects: adding a sparse, opaque texture comprised of a small set of carefully chosen lines. We explain the perceptual motivation for explicitly drawing ridge and valley curves on a transparent surface, describe straightforward mathematical techniques for detecting and rendering these lines, and propose a small number of reasonably effective methods for selectively emphasizing the most perceptually relevant lines in the display. Victoria Interrante, Henry Fuchs, Stephen M. Pizer |
IEEE Visualization | 2 |
| 1994 | Frameless rendering: double buffering considered harmfulabstractThe use of double-buffered displays, in which the previous image is displayed until the next image is complete, can impair the interactivity of systems that require tight coupling between the human user and the computer. We are experimenting with an alternate rendering strategy that computes each pixel based on the most recent input (i.e., view and object positions) and immediately updates the pixel on the display. We avoid the image tearing normally associated with single-buffered displays by randomizing the order in which pixels are updated. The resulting image sequences give the impression of moving continuously, with a rough approximation of motion blur, rather than jerking between discrete positions. Gary Bishop, Henry Fuchs, Leonard McMillan, Ellen J. Scher Zagier |
SIGGRAPH | 2 |
| 1994 | Case Study: Observing a Volume Rendered Fetus within a Pregnant PatientabstractAugmented reality systems with see-through headmounted displays have been used primarily for applications that are possible with today's computational capabilities. We explore possibilities for a particular application-in-place, real-time 3D ultrasound visualization-without concern for such limitations. The question is not "How well could we currently visualize the fetus in real time," but "How well could we see the fetus if we had sufficient compute power?" Our video sequence shows a 3D fetus within a pregnant woman's abdomen-the way this would look to a HMD user. Technical problems in making the sequence are discussed. This experience exposed limitations of current augmented reality systems; it may help define the capabilities of future systems needed for applications as demanding as real-time medical visualization.> Andrei State, David T. Chen, Chris Tector, Andrew Brandt, Ryutarou Ohbuchi, Michael Bajura, Henry Fuchs |
IEEE Visualization | 8 |
| 1992 | A Demonstrated Optical Tracker with Scalable Work Area for Head-Hounted Display SystemsabstractAn optoelectronic head-tracking system for head-mounted displays is described. The system features a scalable work area that currently measures 10' x 12', a measurement update rate of 20-100 Hz with 20-60 ms of delay, and a resolution specification of 2 mm and 0.2 degrees. The sensors consist of four head-mounted imaging devices that view infrared lightemitting diodes (LEDs) mounted in a 10' x 12' grid of modular 2' x 2' suspended ceiling panels. Photogrammetric techniques allow the head's location to be expressed as a function of the known LED positions and their projected images on the sensors. The work area is scaled by simply adding panels to the ceiling's grid. Discontinuities that occurred when changing working sets of LEDs were reduced by carefully managing all error sources, including LED placement tolerances, and by adopting an overdetermined mathematical model for the computation of head position: space resecfion by collinearity. The working system was demonstrated in the Tomorrow's Realities gallery at the ACM SIGGRAPH '91 conference. Mark Ward, Ronald T. Azuma, Robert Bennett, Stefan Gottschalk, Henry Fuchs |
SI3D | 5 |
| 1992 | Merging virtual objects with the real world: seeing ultrasound imagery within the patientabstractWe describe initial results which show "live" ultrasound echography data visualized within a pregnant human subject.The visualization is achieved by using a small video camera mounted in front of a conventional head-mounted display worn by an observer.The camera's video images are composite with computer-generated ones that contain one or more 2D ultrasound images properly transformed to the observer's current viewing position.As the observer walks around the subject.the ultrasound images appear stationary in 3-space within the subject.This kind of enhancement of the observer's vision may have many other applications, e.g., image guided surgical procedures and on location 3D interactive architecture preview. Michael Bajura, Henry Fuchs, Ryutarou Ohbuchi |
SIGGRAPH | 2 |
| 1991 | Achieving Direct Volume Visualization with Interactive Semantic Region SelectionabstractThe authors have achieved rates as high as 15 frames per second for interactive direct visualization of 3D data by trading some function for speed, while volume rendering with a full complement of ramp classification capabilities is performed at 1.4 frames per second. These speeds have made the combination of region selection with volume rendering practical for the first time. Semantic-driven selection, rather than geometric clipping, has proved to be a natural means of interacting with 3D data. Internal organs in medical data or other regions of interest can be built from preprocessed region primitives. The resulting combined system has been applied to real 3D medical data with encouraging results.> Terry S. Yoo, Ulrich Neumann, Henry Fuchs, Stephen M. Pizer, Tim J. Cullip, John Rhoades, Ross T. Whitaker |
IEEE Visualization | 3 |
| 1990 | A real-time optical 3D tracker for head-mounted display systemsabstractIn this paper, a new optical system for real-time, three-dimensional position tracking is described. This system adopts an "inside-out" tracking paradigm. The working environment is a room where the ceiling is lined with a regular pattern of infrared LEDs flashing under the system's control. Three cameras are mounted on a helmet which the user wears. Each camera uses a lateral effect photodiode as the recording surface. The 2D positions of the LED images inside the field of view of the cameras are detected and reported in real time. The measured 2D image positions and the known 3D positions of the LEDs are used to compute the position and orientation of the camera assembly in space.We have designed an iterative algorithm to estimate the 3D position of the camera assembly in space. The algorithm is a generalized version of the Church's method, and allows for multiple cameras with nonconvergent nodal points. Several equations are formulated to predict the system's error analytically. The requirements of accuracy, speed, adequate working volume, light weight and small size of the tracker are also addressed.A prototype was designed and built to demonstrate the integration and coordination of all essential components of the new tracker. This prototype uses off-the-shelf components and can be easily duplicated. Our results indicate that the new system significantly out-performs other existing systems. The new tracker provides more than 200 updates per second, registers 0.1-degree rotational movements and 2-millimeter translational movements, and processes a working volume about 1,000 ft3 (10 ft on each side). Jih-fang Wang, Vernon L. Chi, Henry Fuchs |
I3D | 3 |
| 1989 | Pixel-planes 5: a heterogeneous multiprocessor graphics system using processor-enhanced memoriesabstractThis paper introduces the architecture and initial algorithms for Pixel-Planes 5, a heterogeneous multi-computer designed both for high-speed polygon and sphere rendering (1M Phong-shaded triangles/second) and for supporting algorithm and application research in interactive 3D graphics. Techniques are described for volume rendering at multiple frames per second, font generation directly from conic spline descriptions, and rapid calculation of radiosity form-factors. The hardware consists of up to 32 math-oriented processors, up to 16 rendering units, and a conventional 1280 × 1024-pixel frame buffer, interconnected by a 5 gigabit ring network. Each rendering unit consists of a 128 × 128-pixel array of processors-with-memory with parallel quadratic expression evaluation for every pixel. Implemented on 1.6 micron CMOS chips designed to run at 40MHz, this array has 208 bits/pixel on-chip and is connected to a video RAM memory system that provides 4,096 bits of off-chip memory. Rendering units can be independently reasigned to any part of the screen or to non-screen-oriented computation. As of April 1989, both hardware and software are still under construction, with initial system operation scheduled for fall 1989. Henry Fuchs, John Poulton, John G. Eyles, Trey Greer, Jack Goldfeather, David A. Ellsworth, Steven E. Molnar, Greg Turk, Brice Tebbs, Laura Israel |
SIGGRAPH | 1 |
| 1988 | Parallel processing for computer vision and display (panel session)
Peter M. Dew, Henry Fuchs, Tosiyasu L. Kunii, Michael J. Wozny |
SIGGRAPH | 2 |
| 1987 | Issues from the 1986 workshop on interactive 3D graphics (panel)
Henry Fuchs, Stuart K. Card, Franklin C. Crow, Stephen M. Pizer |
CHI | 1 |
| 1986 | Image rendering by adaptive refinementabstractThis paper describes techniques for improving the performance of image rendering on personal workstations by using CPU cycles going idle while the user is examining a static image on the screen. In that spirit, we believe that a renderer's work is never done. Our goal is to convey the most information to the user as early as possible, with image quality constantly improving with time. We do this by first generating a crude image rapidly and then adaptively refining it where necessary as long as the user does not change viewing parameters. The renderer operates in a succession of phases, first displaying only vertices of polygons, next polygon edges, then flat shading polygons, then shadowing polygons, then Gouraud shading polygons, then Phong shading polygons, and finally anti-aliasing. Performance is enhanced by each phase using results from previous phases and trimming the amount of data needed by the next phase. In this way, only a fraction of the pixels in an image may be Phong shaded while the rest may be Gouraud or flat shaded. Similarly anti-aliasing is performed only on pixels around which there is significant color change. The system features fast response to user intervention, encourages user intervention at any moment, and makes useful the idle cycles in a personal computer. Larry Bergman, Henry Fuchs, Eric D. Grant, Susan Spach |
SIGGRAPH | 2 |
| 1986 | Fast constructive-solid geometry display in the pixel-powers graphics systemabstractWe present two algorithms for the display of CSG-defined objects on Pixel-Powers, an extension of the Pixel-Planes logic-enhanced memory architecture, which calculates for each and every pixel on the screen (in parallel) the value of any quadratic function in the screen coordinates (x,y). The first algorithm restructures any CSG tree into an equivalent, but possibly larger, tree whose display can be achieved by the second algorithm. The second algorithm traverses the restructured tree and generates quadratic coefficients and opcodes for Pixel-Powers. These opcodes instruct Pixel-Powers to generate the boundaries of primitives and perform set operations using the standard Z-buffer algorithm.Several externally-supplied CSG data sets have been processed with the new tree-traversal algorithm and an associated Pixel-Powers simulator. The resulting images indicate that good results can be obtained very rapidly with the new system. For example, the commonly used MBB test part (at right) with 24 primitives is translated into approximately 1900 quadratic equations. On a Pixel-Powers system running at 10MHz (the speed at which our current Pixel-Planes memories run), the image should be rendered in about 7.5 milliseconds. Jack Goldfeather, Jeff P. Hultquist, Henry Fuchs |
SIGGRAPH | 3 |
| 1985 | Fast spheres, shadows, textures, transparencies, and imgage enhancements in pixel-planesabstractPixel-planes is a logic-enhanced memory system for raster graphics and imaging. Although each pixel-memory is enhanced with a one-bit ALU, the system's real power comes from a tree of one-bit adders that can evaluate linear expressions Ax+By+C for every pixel (x,y) simultaneously, as fast as the ALUs and the memory circuits can accept the results. We and others have begun to develop a variety of algorithms that exploit this fast linear expression evaluation capability. In this paper we report some of those results. Illustrated in this paper is a sample image from a small working prototype of the Pixel-planes hardware and a variety of images from simulations of a full-scale system. Timing estimates indicate that 30,000 smooth shaded triangles can be generated per second, or 21,000 smooth-shaded and shadowed triangles can be generated per second, or over 25,000 shaded spheres can be generated per second. Image-enhancement by adaptive histogram equalization can be performed within 4 seconds on a 512x512 image. Henry Fuchs, Jack Goldfeather, Jeff P. Hultquist, Susan Spach, John D. Austin, Frederick P. Brooks Jr., John G. Eyles, John Poulton |
SIGGRAPH | 1 |
| 1984 | Trends in semiconductor hardware for graphics systems (Panel)abstractDuring the next 5 - 10 years, text and graphic systems will tend to merge because of the demand for a more productive man/machine interface, failing memory costs and the availability of higher performance VLSI controllers. This panel will discuss video controllers, memory components and their architectures, graphic systems configurations and the evolution of enhanced system performance versus reduced system cost. Henry Fuchs |
SIGGRAPH | 1 |
| 1983 | Near real-time shaded display of rigid objectsabstractDescribed is a visible surface algorithm and an implementation that generates shaded display of objects with hundreds of polygons rapidly enough for interactive use — several images per second. The basic algorithm, introduced in [Fuchs, Kedem and Naylor, 1980], is designed to handle rigid objects and scenes by preprocessing the object data base to minimize visibility computation cost. The speed of the algorithm is further enhanced by its simplicity, which allows it to be implemented within the internal graphics processor of a general purpose raster system. Henry Fuchs, Greg Abram, Eric D. Grant |
SIGGRAPH | 1 |
| 1980 | Trends in high performance graphic systems(Panel Session)abstractAccompanying the rapid development of integrated circuit fabrication technology has been a parallel, but slower, development of IC design techniques and systems. Recent approaches to IC design enable individual designers to consider developing their own VLSI circuits. Such capability may open the door to a more varied set of system designs than could previously be considered in most design environments. Henry Fuchs, D. Cohen, Robert F. Sproull, James H. Clark 0001, Frederic I. Parke |
SIGGRAPH | 1 |
| 1980 | On visible surface generation by a priori tree structuresabstractThis paper describes a new algorithm for solving the hidden surface (or line) problem, to more rapidly generate realistic images of 3-D scenes composed of polygons, and presents the development of theoretical foundations in the area as well as additional related algorithms. As in many applications the environment to be displayed consists of polygons many of whose relative geometric relations are static, we attempt to capitalize on this by preprocessing the environment's database so as to decrease the run-time computations required to generate a scene. This preprocessing is based on generating a “binary space partitioning” tree whose in order traversal of visibility priority at run-time will produce a linear order, dependent upon the viewing position, on (parts of) the polygons, which can then be used to easily solve the hidden surface problem. In the application where the entire environment is static with only the viewing-position changing, as is common in simulation, the results presented will be sufficient to solve completely the hidden surface problem. Henry Fuchs, Zvi M. Kedem, Bruce F. Naylor |
SIGGRAPH | 1 |
| 1979 | An Expanded Multiprocessor Architecture for Video GraphicsabstractPresented is the design of a flexible expandable multi-processor system for video graphics and image processing. The design involves a central controller which broadcasts data to a variable number of independently executing processing units, each of which in turn controls a variable number of memory units among which the video (frame buffer) image is distributed. An interleaved addressing organization of the video memories guarantees both an even workload distribution as well as maintenance of image coherence for each processing element. Execution speed and image resolution can be independently altered (at any time) by varying the number of processing and memory units. Sample applications of the system—for rapid line drawing and “electronic scene generation” (visible surface algorithms)—are described. Variations of the design for low cost and for powerful, real-time configurations are outlined. Henry Fuchs, Brian W. Johnson |
ISCA | 1 |
| 1979 | Generating smooth 2-D monocolor line drawings on video displaysabstractOne of the major drawbacks of video display systems for line drawing applications has been the poor image quality they usually produce—“jaggy”, “staircased” line edges, moire patterns in regions of closely spaced lines, even, with some systems, lines disappearing (“falling in”) between pixels. Correcting these effects, with appropriate area-sampling techniques, has generally been too computationally expensive to adopt. José Barros, Henry Fuchs |
SIGGRAPH | 2 |
| 1979 | Predetermining visibility priority in 3-D scenes (Preliminary Report)abstractThe principal calculation performed by all visible surface algorithms is the determination of the visible polygon at each pixel in the image. Of the many possible speedups and efficiencies found for this problem, only one published algorithm (developed almost a decade ago by a group at General Electric) took advantage of an observation that many visibility calculations could be performed without knowledge of the eventual viewing position and orientation—once for all possible images. The method is based on a “potential obscuration” relation between polygons in the simulated environment. Unfortunately, the method worked only for certain objects; unmanagable objects had to be manually (and expertly!) subdivided into managable pieces. Henry Fuchs, Zvi M. Kedem, Bruce F. Naylor |
SIGGRAPH | 1 |
| 1977 | Optimal surface reconstruction from planar contoursabstractNo abstract available. Henry Fuchs, Zvi M. Kedem, Samuel P. Uselton |
SIGGRAPH | 1 |