Michael Stengel

dblp:18/9656 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
5since 2021 · last 2025
0009-0008-6234-936XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 Coherent 3D Portrait Video Reconstruction via Triplane Fusion
abstract
Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user’s appearance. On the other hand, self-reenactment methods can render coherent 3D portraits by driving a 3D avatar built from a single reference image but fail to faithfully preserve the user’s per-frame appearance (e.g., instantaneous facial expressions and lighting). As a result, neither of these two frameworks is an ideal solution for democratized 3D telepresence. In this work, we address this dilemma and propose a novel solution that maintains both coherent identity and dynamic per-frame appearance to enable the best possible realism. To this end, we propose a new fusion-based method that takes the best of both worlds by fusing a canonical 3D prior from a reference view with dynamic appearance from per-frame input views, producing temporally stable 3D videos with faithful reconstruction of the user’s per-frame appearance. Trained only using synthetic data produced by an expression-conditioned 3D GAN, our encoder-based method achieves both state-of-the-art 3D reconstruction and temporal consistency on in-studio and in-the-wild datasets.
Shengze Wang 0002, Chao Liu 0064, Matthew A. Chan 0001, Michael Stengel, Henry Fuchs, Shalini De Mello, Koki Nagano
CVPR5
2025 BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation
abstract
Single-Image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve both accurate 3D pose and 2D alignment at the same time. Error is mainly introduced by inaccurate perspective projection heuristically derived from orthographic parameters. To resolve this long-standing challenge, we present our method BLADE which accurately recovers perspective parameters from a single image without heuristic assumptions. We start from the inverse relationship between perspective distortion and the person’s Z-translation Tz, and we show that Tzcan be reliably estimated from the image. We then discuss the important role of Tzfor accurate human mesh recovery estimated from closerange images. Finally, we show that, once Tzand the 3D human mesh are estimated, one can accurately recover the focal length and full 3D translation. Extensive experiments on standard benchmarks and real-world close-range images show that our method accurately recovers projection parameters from a single image, and consequently attains state-of-the-art accuracy on both 3D pose estimation and 2D alignment for a wide range of images.
Shengze Wang 0002, Tianye Li, Ye Yuan 0007, Henry Fuchs, Koki Nagano, Shalini De Mello, Michael Stengel
CVPR8
2023 Real-Time Radiance Fields for Single-Image Portrait View Synthesis
abstract
We present a one-shot method to infer and render a photorealistic 3D representation from a single unposed image (e.g., face portrait) in real-time. Given a single RGB input, our image encoder directly predicts a canonical triplane representation of a neural radiance field for 3D-aware novel view synthesis via volume rendering. Our method is fast (24 fps) on consumer hardware, and produces higher quality results than strong GAN-inversion baselines that require test-time optimization. To train our triplane encoder pipeline, we use only synthetic data, showing how to distill the knowledge from a pretrained 3D GAN into a feedforward encoder. Technical contributions include a Vision Transformer-based triplane encoder, a camera data augmentation strategy, and a well-designed loss function for synthetic data training. We benchmark against the state-of-the-art methods, demonstrating significant improvements in robustness and image quality in challenging real-world settings. We showcase our results on portraits of faces (FFHQ) and cats (AFHQ), but our algorithm can also be applied in the future to other categories with a 3D-aware image generator.
Alex Trevithick, Matthew A. Chan 0001, Michael Stengel, Eric R. Chan, Chao Liu 0064, Zhiding Yu, Sameh Khamis, Manmohan Krishna Chandraker, Ravi Ramamoorthi, Koki Nagano
ACM Trans. Graph.3
2022 QuadStream: A Quad-Based Scene Streaming Architecture for Novel Viewpoint Reconstruction
abstract
Streaming rendered 3D content over a network to a thin client device, such as a phone or a VR/AR headset, brings high-fidelity graphics to platforms where it would not normally possible due to thermal, power, or cost constraints. Streamed 3D content must be transmitted with a representation that is both robust to latency and potential network dropouts. Transmitting a video stream and reprojecting to correct for changing viewpoints fails in the presence of disocclusion events; streaming scene geometry and performing high-quality rendering on the client is not possible on limited-power mobile GPUs. To balance the competing goals of disocclusion robustness and minimal client workload, we introduce QuadStream , a new streaming content representation that reduces motion-to-photon latency by allowing clients to efficiently render novel views without artifacts caused by disocclusion events. Motivated by traditional macroblock approaches to video codec design, we decompose the scene seen from positions in a view cell into a series of quad proxies , or view-aligned quads from multiple views. By operating on a rasterized G-Buffer, our approach is independent of the representation used for the scene itself; the resulting QuadStream is an approximate geometric representation of the scene that can be reconstructed by a thin client to render both the current view and nearby adjacent views. Our technical contributions are an efficient parallel quad generation, merging, and packing strategy for proxy views covering potential client movement in a scene; a packing and encoding strategy that allows masked quads with depth information to be transmitted as a frame-coherent stream; and an efficient rendering approach for rendering our QuadStream representation into entirely novel views on thin clients. We show that our approach achieves superior quality compared both to video data streaming methods, and to geometry-based streaming.
Jozef Hladky, Michael Stengel, Nicholas Vining, Bernhard Kerbl, Hans-Peter Seidel, Markus Steinberger
ACM Trans. Graph.2
2021 A distributed, decoupled system for losslessly streaming dynamic light probes to thin clients
abstract
We present a networked, high-performance graphics system that combines dynamic, high-quality, ray traced global illumination computed on a server with direct illumination and primary visibility computed on a client. This approach provides many of the image quality benefits of real-time ray tracing on low-power and legacy hardware, while maintaining a low latency response and mobile form factor.
Michael Stengel, Alexander Majercik, Ben Boudaoud, Morgan McGuire
MMSys1
2020 Optical Gaze Tracking with Spatially-Sparse Single-Pixel Detectors
abstract
Gaze tracking is an essential component of next generation displays for virtual reality and augmented reality applications. Traditional camera-based gaze trackers used in next generation displays are known to be lacking in one or multiple of the following metrics: power consumption, cost, computational complexity, estimation accuracy, latency, and form-factor. We propose the use of discrete photodiodes and light-emitting diodes (LEDs) as an alternative to traditional camera-based gaze tracking approaches while taking all of these metrics into consideration. We begin by developing a rendering-based simulation framework for understanding the relationship between light sources and a virtual model eyeball. Findings from this framework are used for the placement of LEDs and photodiodes. Our first prototype uses a neural network to obtain an average error rate of 2.67° at 400 Hz while demanding only 16 mW. By simplifying the implementation to using only LEDs, duplexed as light transceivers, and more minimal machine learning model, namely a light-weight supervised Gaussian process regression algorithm, we show that our second prototype is capable of an average error rate of 1.57° at 250 Hz using 800 mW.
Richard Li 0002, Eric Whitmire, Michael Stengel, Ben Boudaoud, Jan Kautz, David P. Luebke, Shwetak N. Patel, Kaan Aksit
ISMAR3
2020 Toward Standardized Classification of Foveated Displays
abstract
Emergent in the field of head mounted display design is a desire to leverage the limitations of the human visual system to reduce the computation, communication, and display workload in power and form-factor constrained systems. Fundamental to this reduced workload is the ability to match display resolution to the acuity of the human visual system, along with a resulting need to follow the gaze of the eye as it moves, a process referred to as foveation. A display that moves its content along with the eye may be called a Foveated Display, though this term is also commonly used to describe displays with non-uniform resolution that attempt to mimic human visual acuity. We therefore recommend a definition for the term Foveated Display that accepts both of these interpretations. Furthermore, we include a simplified model for human visual Acuity Distribution Functions (ADFs) at various levels of visual acuity, across wide fields of view and propose comparison of this ADF with the Resolution Distribution Function of a foveated display for evaluation of its resolution at a particular gaze direction. We also provide a taxonomy to allow the field to meaningfully compare and contrast various aspects of foveated displays in a display and optical technology-agnostic manner.
Josef B. Spjut, Ben Boudaoud, Jonghyun Kim 0006, Trey Greer, Rachel A. Albert, Michael Stengel, Kaan Aksit, David P. Luebke
IEEE Trans. Vis. Comput. Graph.6
2019 NVGaze: An Anatomically-Informed Dataset for Low-Latency, Near-Eye Gaze Estimation
abstract
Quality, diversity, and size of training data are critical factors for learning-based gaze estimators. We create two datasets satisfying these criteria for near-eye gaze estimation under infrared illumination: a synthetic dataset using anatomically-informed eye and face models with variations in face shape, gaze direction, pupil and iris, skin tone, and external conditions (2M images at 1280x960), and a real-world dataset collected with 35 subjects (2.5M images at 640x480). Using these datasets we train neural networks performing with sub-millisecond latency. Our gaze estimation network achieves 2.06(±0.44)° of accuracy across a wide 30°×40° field of view on real subjects excluded from training and 0.5° best-case accuracy (across the same FOV) when explicitly trained for one real subject. We also train a pupil localization network which achieves higher robustness than previous methods.
Joohwan Kim, Michael Stengel, Alexander Majercik, Shalini De Mello, David Dunn, Samuli Laine, Morgan McGuire, David P. Luebke
CHI2
2019 RetroTracker: Upgrading Existing Virtual Reality Tracking Systems
abstract
Virtual reality systems often make use of spatially tracked handheld props in the form of controllers or specialized objects to add realism and interaction. Tracking these objects today relies on the use of expensive, bulky, and power-consuming trackers that must be attached to an object. We propose a passive tracking technique that works with existing low-cost, off-the-shelf optical tracking components and is capable of turning any object into a tracked virtual reality prop. Our method utilizes paper-thin retro-reflective markers that can be placed in any free-form on everyday objects. The proof-of-concept prototype acts as a simple add-on for an existing tracking system and requires only a minimal amount of compute overhead. We demonstrate that our method allows bringing physical real-world objects to virtual worlds with ease, and provides an object identification technique using patterned retro-reflective markers.
Kylee M. Krzanich, Eric Whitmire, Michael Stengel, Michael Kass, Kaan Aksit, David P. Luebke
VR3
2019 Near-Eye Display and Tracking Technologies for Virtual and Augmented Reality
abstract
Abstract Virtual and augmented reality (VR/AR) are expected to revolutionise entertainment, healthcare, communication and the manufacturing industries among many others. Near‐eye displays are an enabling vessel for VR/AR applications, which have to tackle many challenges related to ergonomics, comfort, visual quality and natural interaction. These challenges are related to the core elements of these near‐eye display hardware and tracking technologies. In this state‐of‐the‐art report, we investigate the background theory of perception and vision as well as the latest advancements in display engineering and tracking technologies. We begin our discussion by describing the basics of light and image formation. Later, we recount principles of visual perception by relating to the human visual system. We provide two structured overviews on state‐of‐the‐art near‐eye display and tracking technologies involved in such near‐eye displays. We conclude by outlining unresolved research questions to inspire the next generation of researchers.
George Alex Koulieris, Kaan Aksit, Michael Stengel, Rafal Mantiuk, Katerina Mania, Christian Richardt
Comput. Graph. Forum3
2019 Foveated AR: dynamically-foveated augmented reality display
abstract
We present a near-eye augmented reality display with resolution and focal depth dynamically driven by gaze tracking. The display combines a traveling microdisplay relayed off a concave half-mirror magnifier for the high-resolution foveal region, with a wide field-of-view peripheral display using a projector-based Maxwellian-view display whose nodal point is translated to follow the viewer's pupil during eye movements using a traveling holographic optical element. The same optics relay an image of the eye to an infrared camera used for gaze tracking, which in turn drives the foveal display location and peripheral nodal point. Our display supports accommodation cues by varying the focal depth of the microdisplay in the foveal region, and by rendering simulated defocus on the "always in focus" scanning laser projector used for peripheral display. The resulting family of displays significantly improves on the field-of-view, resolution, and form-factor tradeoff present in previous augmented reality designs. We show prototypes supporting 30, 40 and 60 cpd foveal resolution at a net 85° × 78° field of view per eye.
Jonghyun Kim 0006, Youngmo Jeong, Michael Stengel, Kaan Aksit, Rachel A. Albert, Ben Boudaoud, Trey Greer, Joohwan Kim, Ward Lopes, Alexander Majercik, Peter Shirley, Josef B. Spjut, Morgan McGuire, David P. Luebke
ACM Trans. Graph.3
2017 Subtle gaze guidance for immersive environments
abstract
Immersive displays allow presentation of rich video content over a wide field of view. We present a method to boost visual importance for a selected - possibly invisible - scene part in a cluttered virtual environment. This desirable feature enables to unobtrusively guide the gaze direction of a user to any location within the immersive 360° surrounding. Our method is based on subtle gaze direction which did not include head rotations in previous work. For covering the full 360° environment and wide field of view, we contribute an approach for dynamic stimulus positioning and shape variation based on eccentricity to compensate for visibility differences across the visual field. Our approach is calibrated in a perceptual study for a head-mounted display with binocular eye tracking. An additional study validates the method within an immersive visual search task.
Steve Grogorick, Michael Stengel, Elmar Eisemann, Marcus A. Magnor
SAP2
2017 Perception-driven Accelerated Rendering
abstract
Advances in computer graphics enable us to create digital images of astonishing complexity and realism. However, processing resources are still a limiting factor. Hence, many costly but desirable aspects of realism are often not accounted for, including global illumination, accurate depth of field and motion blur, spectral effects, etc. especially in real-time rendering. At the same time, there is a strong trend towards more pixels per display due to larger displays, higher pixel densities or larger fields of view. Further observable trends in current display technology include more bits per pixel (high dynamic range, wider color gamut/fidelity), increasing refresh rates (better motion depiction), and an increasing number of displayed views per pixel (stereo, multi-view, all the way to holographic or lightfield displays). These developments cause significant unsolved technical challenges due to aspects such as limited compute power and bandwidth. Fortunately, the human visual system has certain limitations, which mean that providing the highest possible visual quality is not always necessary. In this report, we present the key research and models that exploit the limitations of perception to tackle visual quality and workload alike. Moreover, we present the open problems and promising future research targeting the question of how we can minimize the effort to compute and display only the necessary pixels while still offering a user full visual experience.
Martin Weier, Michael Stengel, Thorsten Roth, Piotr Didyk, Elmar Eisemann, Martin Eisemann, Steve Grogorick, André Hinkenjann, Ernst Kruijff, Marcus A. Magnor, Karol Myszkowski, Philipp Slusallek
Comput. Graph. Forum2
2016 Adaptive Image-Space Sampling for Gaze-Contingent Real-time Rendering
abstract
With ever-increasing display resolution for wide field-of-view displays—such as head-mounted displays or 8k projectors—shading has become the major computational cost in rasterization. To reduce computational effort, we propose an algorithm that only shades visible features of the image while cost-effectively interpolating the remaining features without affecting perceived quality. In contrast to previous approaches we do not only simulate acuity falloff but also introduce a sampling scheme that incorporates multiple aspects of the human visual system: acuity, eye motion, contrast (stemming from geometry, material or lighting properties), and brightness adaptation. Our sampling scheme is incorporated into a deferred shading pipeline to shade the image's perceptually relevant fragments while a pull-push algorithm interpolates the radiance for the rest of the image. Our approach does not impose any restrictions on the performed shading. We conduct a number of psycho-visual experiments to validate scene- and task-independence of our approach. The number of fragments that need to be shaded is reduced by 50 % to 80 %. Our algorithm scales favorably with increasing resolution and field-of-view, rendering it well-suited for head-mounted displays and wide-field-of-view projection.
Michael Stengel, Steve Grogorick, Martin Eisemann, Marcus A. Magnor
Comput. Graph. Forum1
2015 An Affordable Solution for Binocular Eye Tracking and Calibration in Head-mounted Displays
abstract
Immersion is the ultimate goal of head-mounted displays (HMD) for Virtual Reality (VR) in order to produce a convincing user experience. Two important aspects in this context are motion sickness, often due to imprecise calibration, and the integration of a reliable eye tracking. We propose an affordable hard- and software solution for drift-free eye-tracking and user-friendly lens calibration within an HMD. The use of dichroic mirrors leads to a lean design that provides the full field-of-view (FOV) while using commodity cameras for eye tracking. Our prototype supports personalizable lens positioning to accommodate for different interocular distances. On the software side, a model-based calibration procedure adjusts the eye tracking system and gaze estimation to varying lens positions. Challenges such as partial occlusions due to the lens holders and eye lids are handled by a novel robust monocular pupil-tracking approach. We present four applications of our work: Gaze map estimation, foveated rendering for depth of field, gaze-contingent level-of-detail, and gaze control of virtual avatars.
Michael Stengel, Steve Grogorick, Martin Eisemann, Elmar Eisemann, Marcus A. Magnor
ACM Multimedia1
2015 Non-obscuring binocular eye tracking for wide field-of-view head-mounted-displays
abstract
We present a complete hardware and software solution for integrating binocular eye tracking into current state-of-the-art lens-based Head-mounted Displays (HMDs) without affecting the user's wide field-of-view off the display. The system uses robust and efficient new algorithms for calibration and pupil tracking and allows realtime eye tracking and gaze estimation. Estimating the relative gaze direction of the user opens the door to a much wider spectrum of virtual reality applications and games when using HMDs. We show a 3d-printed prototype of a low-cost HMD with eye tracking that is simple to fabricate and discuss a variety of VR applications utilizing gaze estimation.
Michael Stengel, Steve Grogorick, Martin Eisemann, Elmar Eisemann, Marcus A. Magnor
VR1
2015 Temporal Video Filtering and Exposure Control for Perceptual Motion Blur
abstract
We propose the computation of a perceptual motion blur in videos. Our technique takes the predicted eye motion into account when watching the video. Compared to traditional motion blur recorded by a video camera our approach results in a perceptual blur that is closer to reality. This postprocess can also be used to simulate different shutter effects or for other artistic purposes. It handles real and artificial video input, is easy to compute and has a low additional cost for rendered content. We illustrate its advantages in a user study using eye tracking.
Michael Stengel, Pablo Bauszat, Martin Eisemann, Elmar Eisemann, Marcus A. Magnor
IEEE Trans. Vis. Comput. Graph.1
2014 Garment Replacement in Monocular Video Sequences
abstract
We present a semi-automatic approach to exchange the clothes of an actor for arbitrary virtual garments in conventional monocular video footage as a postprocess. We reconstruct the actor's body shape and motion from the input video using a parameterized body model. The reconstructed dynamic 3D geometry of the actor serves as an animated mannequin for simulating the virtual garment. It also aids in scene illumination estimation, necessary to realistically light the virtual garment. An image-based warping technique ensures realistic compositing of the rendered virtual garment and the original video. We present results for eight real-world video sequences featuring complex test cases to evaluate performance for different types of motion, camera settings, and illumination conditions.
Lorenz Rogge, Felix Klose, Michael Stengel, Martin Eisemann, Marcus A. Magnor
ACM Trans. Graph.3
2013 Optimizing Apparent Display Resolution Enhancement for Arbitrary Videos
abstract
Display resolution is frequently exceeded by available image resolution. Recently, apparent display resolution enhancement (ADRE) techniques show how characteristics of the human visual system can be exploited to provide super-resolution on high refresh rate displays. In this paper, we address the problem of generalizing the ADRE technique to conventional videos of arbitrary content. We propose an optimization-based approach to continuously translate the video frames in such a way that the added motion enables apparent resolution enhancement for the salient image region. The optimization considers the optimal velocity, smoothness, and similarity to compute an appropriate trajectory. In addition, we provide an intuitive user interface that allows to guide the algorithm interactively and preserves important compositions within the video. We present a user study evaluating apparent rendering quality and show versatility of our method on a variety of general test scenes.
Michael Stengel, Martin Eisemann, Stephan Wenger, Benjamin Hell, Marcus A. Magnor
IEEE Trans. Image Process.1
2011 View infinity: a zoomable interface for feature-oriented software development
abstract
Software product line engineering provides efficient means to develop variable software. To support program comprehension of software product lines (SPLs), we developed View Infinity, a tool that provides seamless and semantic zooming of different abstraction layers of an SPL. First results of a qualitative study with experienced SPL developers are promising and indicate that View Infinity is useful and intuitive to use.
Michael Stengel, Mathias Frisch, Sven Apel, Janet Siegmund, Christian Kästner, Raimund Dachselt
ICSE1